flowchart LR
N[Random noise] --> G[Generator<br/>the forger]
G --> F[Fake image]
R[Real images] --> D[Discriminator<br/>the detective]
F --> D
D --> O{Real or fake?}
Session 1 · 3 hours · Lecture + live-coding lab
Session 1 · 3 hours
From input embeddings to output tokens: what is a decoder only block in a transformer ?
B, T, C) like a production habitProduced: a shape-verified decoder block
No generative component is working until it is scored on a held-out golden set with cost and latency tracked.
An impressive demo on the training distribution = the canonical self-deception.
| In scope | Out of scope |
|---|---|
| Working with and around models | Training foundation models |
| Architecture enough to debug systems | Transformer-internals research |
You will ship RAG, agents, and eval harnesses later.
Without architecture intuition you cannot answer:
Today = the mental model for every later cost and failure mode.
Kokoro:
Audio8:
| Family | How it generates | Famous for | Today |
|---|---|---|---|
| RNN / LSTM (1986 / 1997) | Reads and writes one token at a time, carrying a memory | Google Translate (2016), Siri, autocomplete | Comeback as Mamba, xLSTM |
| GAN (2014) | A generator tries to fool a detector | Fake faces, deepfakes, StyleGAN | Fast image upscaling, still used inside image pipelines |
| Energy-based (1982 →) | Scores how “plausible” an input is, then slides toward low energy | Hopfield nets, Boltzmann machines | Ideas live on in diffusion models and LeCun’s JEPA |
| Transformer (2017) | Attention over all previous tokens | GPT, Claude, Llama | Dominant for text, also inside image and video models |
flowchart LR
N[Random noise] --> G[Generator<br/>the forger]
G --> F[Fake image]
R[Real images] --> D[Discriminator<br/>the detective]
F --> D
D --> O{Real or fake?}
🎲 Fun fact (but I’m not sure it’s true)
Ian Goodfellow had the idea in a Montreal bar in 2014 and coded it the same night
🎲 Fun fact
The 2024 Nobel Prize in Physics went to John Hopfield and Geoffrey Hinton, for the Hopfield network (1982) and the Boltzmann machine (1985). Both are energy-based models borrowed from statistical physics.
🎲 Fun fact
Karpathy’s tiny character-level RNN produced fake Shakespeare, fake Wikipedia pages and fake Linux source code that almost compiles. And RNNs are back: Mamba (2023) and xLSTM (2024, from Hochreiter’s own team) process very long sequences with a fixed-size memory.
This idea has been around for a long time now, statistical models were developped with information theory concepts decades ago !
from transformers import AutoModel, AutoTokenizer
tok = AutoTokenizer.from_pretrained("gpt2")
model = AutoModel.from_pretrained("gpt2")
ids = tok("The capital of France", return_tensors="pt").input_ids # [[464, 3139, 286, 4881]]
E = model.get_input_embeddings() # Embedding(50257, 768)
vectors = E(ids) # shape (1, 4, 768)This is what we call a decoder-only language model
GPT2: embeddings have a dimension of 768 GPT3: embedding dimension of 12,288 Claude Fable: not publicly available, probably around 16,384
Here is a schema of a typical model:
Concrete prompt:
For a conversationnal model, what you want is:
But what is concretely this Transformer Block with attention and FFN ?
Example of next token probabilities
Nothing seemed to really crack the problem at hand, until…


Attention = soft routing: each token builds a weighted summary of other tokens.
For one query position (“is” in The capital of France is):
Intuition: the model learns to put high weight on France when predicting Paris.
\[ \mathrm{Attention}(Q,K,V) = \mathrm{softmax}\Big(\frac{QK^\top}{\sqrt{d_k}}\Big) V \]
| Symbol | Role | Shape (per head) |
|---|---|---|
Q |
Queries | (B, T, d_k) |
K |
Keys | (B, T, d_k) |
V |
Values | (B, T, d_k) |
QKᵀ |
Similarity scores | (B, T, T) ← this is the T² cost |
| softmax | Weights sum to 1 | still (B, T, T) |
Why divide by √d_k? Keeps scores from exploding → softmax stays usable.
Practice: long context hurts latency because of that (T, T) matrix (plus memory).
Sequence: ["The", "cat", "sat"] — predicting after "sat".
Illustrative attention weights for query "sat":
| Key → | The | cat | sat |
|---|---|---|---|
| weight | 0.1 | 0.6 | 0.3 |
Output ≈ 0.1·V_The + 0.6·V_cat + 0.3·V_sat
The model “looks at” cat most — useful for guessing on the mat vs in the car.
Source: jalammar.github.io
Embeddings alone do not know that token 1 is before token 5.
| Approach | Idea | Practice note |
|---|---|---|
| Absolute (sin/cos or learned) | Add a vector per index 0..T-1 |
Simple; length limits feel hard |
| Relative / RoPE (modern LMs) | Encode relative distance in attention | Better length extrapolation story |
You need the idea, not the research paper: every production LM injects position somehow.
Product translation: RAG stuffing 50 chunks into context only helps if attention actually routes to the right spans (rerankers exist for this — Session 5).

One head = one soft routing pattern.
Several heads = several patterns in parallel.
| Why multiple heads? | Practice consequence |
|---|---|
| Different relationships at once | n_heads is a hyperparameter in every LLM config |
d_model split across heads |
d_k = d_model / n_heads must divide evenly |
Config literacy: when a model card says n_heads=32, d_model=4096, you now know what that means.
Source: jalammar.github.io
A transformer block is not “replace x with attention.”
It is add a refinement:
| Piece | Job | Why you care |
|---|---|---|
Residual + |
Keep original signal; train deep stacks | Explains why deep LMs are trainable |
| LayerNorm / RMSNorm | Stabilise activations | Pre-norm (modern) vs post-norm (classic) |
| FFN / MLP | Per-token nonlinear mix (C → 4C → C) |
Often most parameters in the block |
Practice: LoRA (Session 3) often targets attention and/or FFN projections — those matrices are where adaptation lives.
Decoder-only LMs must not see the future.
Mask sets illegal scores to −∞ before softmax → weight 0.
Lab goal: prove with a probe that future positions get zero attention weight.
If you only remember one diagram from today, remember the lower-triangular mask.
distilbert-base-uncased is an encoder: great for embeddings, it cannot generate text🎲 Fun fact
The 2017 Attention is All You Need transformer was an encoder-decoder built for translation (English → German and French). GPT (2018) kept only the decoder, BERT (2018) only the encoder.
Concrete:
| Decoding choice | Behaviour | Practice |
|---|---|---|
| Greedy argmax | Deterministic, dull | Good for structured tasks |
| Temperature sampling | Creative / risky | Chat, images-of-text prompts |
| Top-p / top-k | Truncate tail | Production default family |
Per token: 2 (K and V) × layers × KV heads × head dim × 2 bytes (fp16)
| Model | Layers | KV heads × dim | Per token | Full context |
|---|---|---|---|---|
| GPT-2 small | 12 | 12 × 64 | 36 KB | 1k tokens → 36 MB |
| Qwen2.5-0.5B | 24 | 2 × 64 | 12 KB | 32k tokens → 384 MB |
| Llama 3.1 8B | 32 | 8 × 128 | 128 KB | 128k tokens → 16 GB |
The tokenizer fixes the length T before the model even runs, and you pay per token.
💼 Business corner
Serving customers in Russian with an English-centric tokenizer: 6.6× the tokens for the same message, so 6.6× the API bill, slower answers and a context window that fills 6.6× faster. Petrov et al. (2023) measured gaps of up to 15× between languages. Check token counts in your languages before signing a pricing plan.
Your guess: how many tokens for each line?
The question is not “how many characters?” but “how many vocabulary pieces?” URLs and code are expensive; plain English is cheap.
The vocabulary size V sits at both ends of the model: the embedding table (V × C) on the way in, the logits (B × T × V) on the way out.
| Model | V | C | Embedding table | Share of all parameters |
|---|---|---|---|---|
| GPT-2 small | 50,257 | 768 | 38.6M | 31% |
| Qwen2.5-0.5B | 151,936 | 896 | 136M | 28% |
| Llama 3 8B | 128,256 | 4,096 | 525M (×2: in + out) | 13% |
🎲 Fun fact
SolidGoldMagikarp (2023): GPT-3’s vocabulary contained Reddit usernames that almost never appeared in its training text. Their embedding rows were never really trained, and asking the model to repeat ” SolidGoldMagikarp” produced insults, evasions or the word “distribute”.
🎲 Fun fact
Golden Gate Claude (Anthropic, 2024): researchers found an internal feature for the Golden Gate Bridge and turned it up. The model then claimed to be the bridge. Interpretability works on features, not just on attention weights. See here
\[\text{LayerNorm}(x) = \gamma \, \frac{x - \text{mean}(x)}{\sqrt{\text{var}(x) + \varepsilon}} + \beta\]
StandardScaler, but per token, across its C features, at every layer
transformers librarydistilbert-base-uncased only has 66,362,880 trainable parameters!Build a single causal attention head using basic pytorch objects (no nn.Transformer ;))
Compute the attention score using Q, K matrices:
Compute scaled scores
normalise them with softmax, multiply by V
concatenate the heads and apply the output projection.
At every stage, compare the observed shape with the expected shape written on the slide.
Use pre-normalisation:
x = x + attention(LayerNorm(x))
x = x + MLP(LayerNorm(x))
The value returned by the block must have the same B × T × C shape as its input
| Symptom | Likely cause |
|---|---|
softmax sees NaNs |
Forgot to mask with -inf (used 0 instead) |
Shape (B, T, T) vs heads |
Missing reshape/transpose for multi-head |
| Loss of residual | Overwrote x instead of x = x + … |
d_model % n_heads != 0 |
Illegal head split |
| Later topic | Hook from Session 1 |
|---|---|
| BPE & Chinchilla (S2) | Embedding dim and vocabulary size are economic variables |
| LoRA (S3) | adapters on Q/K/V/FFN matrices |
| Structured decoding (S4) | softmax over constrained token sets |
| RAG (S5) | large embedding dim → cost; attention may ignore your chunk |
| Agents (S6) | each thought/tool step = more forward passes |
| Eval (S8) | architecture bugs look like “model quality” |
Before you leave, you can explain:
Albert School · Transformer Architecture