less than 1 minute read

Recorded for Prof. Bruce Maxwell’s CS5330 (Pattern Recognition and Computer Vision) at Northeastern. The goal was to walk through the Transformer architecture from Vaswani et al. (2017) end-to-end, then show how Dosovitskiy et al. (2021) ported the whole thing to images with the Vision Transformer.

  • Vaswani et al. (2017). Attention Is All You Need. arXiv:1706.03762
  • Dosovitskiy et al. (2021). An Image Is Worth 16×16 Words. arXiv:2010.11929
  • Johnson, J. (2022). EECS 498/598 Lecture 18: Vision Transformers. University of Michigan
  • Bugnot (2025); Mazurek (2021); Arora / Bukhari (2021) — additional explainer resources listed in the deck