Multi-Head Attention

h = 4 Heads × 16D = 64D

Run 4 attention heads in parallel across 16D subspaces to capture distinct linguistic and semantic relationships.

h: 4
d_k: 16
d_model: 64
Plain Language Intuition

A single attention head can only focus on one type of relationship at a time. Multiple heads let the model attend to different linguistic features simultaneously - one head tracks grammatical neighbors, another tracks subject-verb bindings, and another tracks poetic metaphors.

Production Real-World Context

GPT-3 uses 96 attention heads (each 128D, totaling 12,288D). Head pruning experiments demonstrate that distinct heads specialize in syntax, rare vocabulary tokens, and long-range coreference.

Multi-Head Attention

Intuition: Rather than a single attention pattern, run multiple attention heads in parallel across smaller 16D subspaces, then merge their insights with an output projection W^O.

Inspect Individual Attention Heads

4 Subspaces (16D each)

Head 0: Syntactic Alignment (Adjacent tokens)

Attends heavily to the immediate previous word / grammatical neighbor.

W_O Output Projection: 64D → 64D
"سحر"
"با"
"باد"
"می‌گفتم"
"حدیث"
"سحر"
100%
0%
0%
0%
0%
"با"
70%
30%
0%
0%
0%
"باد"
10%
80%
10%
0%
0%
"می‌گفتم"
5%
15%
70%
10%
0%
"حدیث"
5%
5%
20%
60%
10%
Multi-Head Attention - PyTorch Reference Implementationpython
39 lines
Key Takeaways & Core Rules
  • Multi-Head Attention pipeline: project -> parallel attention -> concatenate -> output projection W^O.
  • Each head operates on a lower-dimensional slice: d_k = d_model / h = 64 / 4 = 16 dimensions.
  • More heads enable diverse syntactic and semantic patterns, while keeping compute cost identical to single-head 64D attention.
  • The output matrix W^O blends all 4 head vectors back into a single unified 64D token representation.
Try This Experiment:

Click through Heads 0, 1, 2, and 3 above. Notice how Head 0 focuses on adjacent words while Head 1 binds long-range poetic subject-rhymes.