Vision Transformer, from scratch
A ViT trained on 1,935 photographs of Bengali terracotta temples — a dataset I had to build first — to test the paper's claim about pre-training. It failed, exactly as predicted.
1 post
A ViT trained on 1,935 photographs of Bengali terracotta temples — a dataset I had to build first — to test the paper's claim about pre-training. It failed, exactly as predicted.