← Attention & Transformers
Free
Self-Attention Mechanics
Queries, keys, values, and scaled dot-product attention.
Cheatsheet — Attention
- Scale by (\sqrt{d_k}) to keep softmax stable
- Causal masks for autoregressive decoding
- Complexity ~ sequence length² (motivation for efficient attention)
- LayerNorm + residuals are load-bearing