Transformers and Vision Transformers
Recorded for Prof. Bruce Maxwell’s CS5330 (Pattern Recognition and Computer Vision) at Northeastern. The goal was to walk through the Transformer architecture from Vaswani et al. (2017) end-to-end, then show how Dosovitskiy et al. (2021) ported the whole thing to images with the Vision Transformer.
- Vaswani et al. (2017). Attention Is All You Need. arXiv:1706.03762
- Dosovitskiy et al. (2021). An Image Is Worth 16×16 Words. arXiv:2010.11929
- Johnson, J. (2022). EECS 498/598 Lecture 18: Vision Transformers. University of Michigan
- Bugnot (2025); Mazurek (2021); Arora / Bukhari (2021) — additional explainer resources listed in the deck