Disclaimer: The opinions expressed in this article are my own and do not represent the views of Google. This content is based solely on publicly available information.
TL;DR of Parts 1-5. Words become embeddings (Part 1). RNNs/LSTMs read them sequentially with vanishing-gradient pain (Part 2). Self-attention reads all positions in parallel via Q, K, V (Part 3). Multi-head attention adds richer representational capacity, and positional encoding restores word order (Part 4). The Encoder Layer wraps multi-head self-attention with residuals, LayerNorm, and a feed-forward network — same shape in, same shape out, stackable (Part 5). We have the understanding half. To generate text we need the second half: the decoder. Part 5 built the Encoder Layer — the building block we now extend.
Series navigation
| Part | Topic |
|---|---|
| 1 | Introduction & Word Embeddings |
| 2 | From RNNs to LSTMs |
| 3 | The Attention Mechanism |
| 4 | Multi-Head Attention & Positional Encoding |
| 5 | The Transformer Encoder |
| 6 (this article) | The Decoder & Full Transformer |
| 7 | Pre-trained Models & Tokenization |
| 8 | Supervised Fine-Tuning |
| 9 | LoRA & QLoRA |
| 10 | DPO Alignment |
Part 5 gave us the complete Transformer Encoder Layer: multi-head self-attention wrapped with residuals, LayerNorm, and a feed-forward network — a shape-preserving block we can stack to any depth. That half of the architecture is excellent at building rich contextual representations of an input sequence, but it has no way to produce new tokens one at a time. This part adds the missing half: the DecoderLayer, which introduces two new mechanisms — masked self-attention so the decoder cannot cheat by peeking at future tokens, and cross-attention so it can read the encoder’s output — then assembles both halves into a complete nn.Transformer.
Why the decoder is different
To understand what the decoder layer must do, it helps to contrast it directly with the encoder we just built. The encoder saw the full source sentence at once and could freely mix information in both directions; the decoder operates under strict constraints that mirror how generation actually works.
The encoder’s job is to contextualize every position in the source sequence. It can freely attend to every other source token because the entire source sequence is available at once.
The decoder’s job is to generate the target sequence one token at a time. At inference time we don’t know the future tokens — they are what we are generating. At training time we do know them (we are learning to predict them), but we must train under the same constraint as inference: when predicting token t, the model can only see tokens 0 through t-1.
We enforce this with a causal (look-ahead) mask — a triangular matrix that blocks attention to future positions. The same masking trick at every position keeps generation honest at training time without requiring us to generate one step at a time.
The decoder also needs to condition on the source sentence — otherwise it would just generate generic English, not a translation of the specific source. That conditioning happens via cross-attention: a second attention block in every decoder layer where the queries come from the decoder’s own representations but the keys and values come from the encoder’s output.
Building the causal mask
The central mechanism that enforces the “no peeking at future tokens” constraint is the causal mask — a simple matrix added to attention scores before softmax. Getting this right is what makes training and autoregressive inference behave identically.
The causal mask is an upper-triangular matrix of -inf values:
import torch
def causal_mask(size: int) -> torch.Tensor:
return torch.triu(torch.full((size, size), float("-inf")), diagonal=1)
mask = causal_mask(6)
print(mask)
Output:
tensor([[0., -inf, -inf, -inf, -inf, -inf],
[0., 0., -inf, -inf, -inf, -inf],
[0., 0., 0., -inf, -inf, -inf],
[0., 0., 0., 0., -inf, -inf],
[0., 0., 0., 0., 0., -inf],
[0., 0., 0., 0., 0., 0.]])
Adding this matrix to the attention scores before softmax pushes every “future” score to negative infinity, which softmax turns into zero. Past and current positions are untouched (their entries are 0, an additive identity).
At inference time, we generate token by token: feed in all tokens produced so far, extract the last position’s logits, sample the next token, append it to the sequence, repeat. The causal mask ensures training and inference are consistent.
The Decoder Layer
With the causal mask in hand, we can now wire up the full decoder layer. The structure mirrors the encoder — residuals and LayerNorm around every sub-block — but adds the cross-attention block that connects decoder to encoder output.
Inside one decoder layer, six operations run:
- Masked multi-head self-attention (queries, keys, values all from target sequence)
- Add & Norm
- Cross-attention (queries from decoder, keys + values from encoder output)
- Add & Norm
- Position-wise FFN
- Add & Norm
import torch
import torch.nn as nn
class DecoderLayer(nn.Module):
def __init__(self, d_model, num_heads, dff, dropout=0.1):
super().__init__()
self.self_attn = nn.MultiheadAttention(d_model, num_heads,
dropout=dropout, batch_first=True)
self.cross_attn = nn.MultiheadAttention(d_model, num_heads,
dropout=dropout, batch_first=True)
self.norm1 = nn.LayerNorm(d_model)
self.norm2 = nn.LayerNorm(d_model)
self.norm3 = nn.LayerNorm(d_model)
self.ffn = nn.Sequential(
nn.Linear(d_model, dff),
nn.ReLU(),
nn.Linear(dff, d_model),
)
self.drop = nn.Dropout(dropout)
def forward(self, tgt, memory, tgt_mask=None):
# 1. Masked self-attention on the target
sa, _ = self.self_attn(tgt, tgt, tgt, attn_mask=tgt_mask,
need_weights=False)
tgt = self.norm1(tgt + self.drop(sa))
# 2. Cross-attention: queries from tgt, keys+values from encoder memory
ca, _ = self.cross_attn(tgt, memory, memory, need_weights=False)
tgt = self.norm2(tgt + self.drop(ca))
# 3. Feed-forward
ff = self.ffn(tgt)
return self.norm3(tgt + self.drop(ff))
Two key differences from the EncoderLayer:
- Masked self-attention — we pass
attn_mask=tgt_maskso future positions are blocked. - Cross-attention —
cross_attn(tgt, memory, memory, ...). Notice thattgtis the query butmemory(encoder output) is both key and value. This is how the decoder reads from the encoder.
layer = DecoderLayer(d_model=32, num_heads=4, dff=64)
n_params = sum(p.numel() for p in layer.parameters())
print(layer)
print(f"\nTotal parameters: {n_params:,}")
DecoderLayer(
(self_attn): MultiheadAttention(
(out_proj): NonDynamicallyQuantizableLinear(in_features=32, out_features=32, bias=True)
)
(cross_attn): MultiheadAttention(
(out_proj): NonDynamicallyQuantizableLinear(in_features=32, out_features=32, bias=True)
)
(norm1): LayerNorm((32,), eps=1e-05, elementwise_affine=True)
(norm2): LayerNorm((32,), eps=1e-05, elementwise_affine=True)
(norm3): LayerNorm((32,), eps=1e-05, elementwise_affine=True)
(ffn): Sequential(
(0): Linear(in_features=32, out_features=64, bias=True)
(1): ReLU()
(2): Linear(in_features=64, out_features=32, bias=True)
)
(drop): Dropout(p=0.1, inplace=False)
)
Total parameters: 12,832
Three LayerNorms in the decoder vs. two in the encoder, because there is one more residual sub-block (the cross-attention). The decoder has ~50% more parameters per layer than the encoder at the same d_model.
Putting both halves together
We now have all the pieces: an encoder stack that reads the source, and a decoder stack that generates the target step by step. Connecting them is straightforward — PyTorch’s nn.Transformer does exactly this, and its print(model) output lets us verify every sub-module is where we expect.
The full Transformer is just an encoder stack feeding into a decoder stack. Here is the nn.Transformer from torch.nn itself, configured exactly the way we have been describing:
import torch
import torch.nn as nn
transformer = nn.Transformer(
d_model=32,
nhead=4,
num_encoder_layers=2,
num_decoder_layers=2,
dim_feedforward=64,
dropout=0.1,
batch_first=True,
)
n_params = sum(p.numel() for p in transformer.parameters())
print(transformer)
print(f"\nTotal parameters: {n_params:,}")
Transformer(
(encoder): TransformerEncoder(
(layers): ModuleList(
(0-1): 2 x TransformerEncoderLayer(
(self_attn): MultiheadAttention(
(out_proj): NonDynamicallyQuantizableLinear(in_features=32, out_features=32, bias=True)
)
(linear1): Linear(in_features=32, out_features=64, bias=True)
(dropout): Dropout(p=0.1, inplace=False)
(linear2): Linear(in_features=64, out_features=32, bias=True)
(norm1): LayerNorm((32,), eps=1e-05, elementwise_affine=True)
(norm2): LayerNorm((32,), eps=1e-05, elementwise_affine=True)
(dropout1): Dropout(p=0.1, inplace=False)
(dropout2): Dropout(p=0.1, inplace=False)
)
)
(norm): LayerNorm((32,), eps=1e-05, elementwise_affine=True)
)
(decoder): TransformerDecoder(
(layers): ModuleList(
(0-1): 2 x TransformerDecoderLayer(
(self_attn): MultiheadAttention(
(out_proj): NonDynamicallyQuantizableLinear(in_features=32, out_features=32, bias=True)
)
(multihead_attn): MultiheadAttention(
(out_proj): NonDynamicallyQuantizableLinear(in_features=32, out_features=32, bias=True)
)
(linear1): Linear(in_features=32, out_features=64, bias=True)
(dropout): Dropout(p=0.1, inplace=False)
(linear2): Linear(in_features=64, out_features=32, bias=True)
(norm1): LayerNorm((32,), eps=1e-05, elementwise_affine=True)
(norm2): LayerNorm((32,), eps=1e-05, elementwise_affine=True)
(norm3): LayerNorm((32,), eps=1e-05, elementwise_affine=True)
(dropout1): Dropout(p=0.1, inplace=False)
(dropout2): Dropout(p=0.1, inplace=False)
(dropout3): Dropout(p=0.1, inplace=False)
)
)
(norm): LayerNorm((32,), eps=1e-05, elementwise_affine=True)
)
)
Total parameters: 42,880
The encoder’s TransformerEncoderLayer has one MultiheadAttention sub-module (self-attention) and a feed-forward block. The decoder’s TransformerDecoderLayer has two — self_attn (the masked one) and multihead_attn (cross-attention) — plus the feed-forward block.
flowchart TD
subgraph ENC["Encoder Stack (N ×)"]
direction TB
E_SA["Multi-Head Self-Attention"]
E_AN1["Add & Norm"]
E_FF["Feed-Forward"]
E_AN2["Add & Norm"]
E_SA --> E_AN1 --> E_FF --> E_AN2
end
subgraph DEC["Decoder Stack (N ×)"]
direction TB
D_MSA["Masked Multi-Head Self-Attention"]
D_AN1["Add & Norm"]
D_CA["Multi-Head Cross-Attention\n(Q from decoder · K,V from encoder)"]
D_AN2["Add & Norm"]
D_FF["Feed-Forward"]
D_AN3["Add & Norm"]
D_MSA --> D_AN1 --> D_CA --> D_AN2 --> D_FF --> D_AN3
end
SRC["Source tokens"] --> SE["Input Embedding + Positional Encoding"]
SE --> ENC
ENC -->|"Encoder output\n(K and V)"| D_CA
TGT["Target tokens (shifted right)"] --> TE["Output Embedding + Positional Encoding"]
TE --> DEC
DEC --> LIN["Linear → vocab size"]
LIN --> SM["Softmax → next-token probabilities"]
Figure 1: Full Transformer — encoder stack on the left feeding the decoder stack on the right.
The diagram shows the two halves running in parallel. Source tokens flow through an embedding plus positional encoding into the encoder stack; target tokens flow through their own embedding plus positional encoding into the decoder stack. The single arrow between the stacks is the only coupling: each decoder layer reads the encoder’s final output as K and V during cross-attention. Below the decoder stack, a linear projection to vocabulary size and a softmax turn hidden states into next-token probabilities — that final pair is everything that separates “representations” from “predictions.”
A working seq2seq model
The nn.Transformer block handles the architecture, but a real training model also needs embedding layers, positional encodings, and a projection head that turns the decoder’s hidden states into token probabilities. Putting these together gives us a fully functional seq2seq model.
Here is what a full seq2seq training model looks like end-to-end:
import torch
import torch.nn as nn
class Seq2Seq(nn.Module):
def __init__(self, src_vocab, tgt_vocab, d_model, nhead, n_layers, dff, max_len=512):
super().__init__()
self.src_emb = nn.Embedding(src_vocab, d_model)
self.tgt_emb = nn.Embedding(tgt_vocab, d_model)
self.pos = nn.Embedding(max_len, d_model) # learned PE
self.transformer = nn.Transformer(
d_model=d_model, nhead=nhead,
num_encoder_layers=n_layers, num_decoder_layers=n_layers,
dim_feedforward=dff, batch_first=True,
)
self.out = nn.Linear(d_model, tgt_vocab)
def forward(self, src, tgt):
L_s, L_t = src.size(1), tgt.size(1)
pos_s = torch.arange(L_s, device=src.device).unsqueeze(0)
pos_t = torch.arange(L_t, device=tgt.device).unsqueeze(0)
s_emb = self.src_emb(src) + self.pos(pos_s)
t_emb = self.tgt_emb(tgt) + self.pos(pos_t)
causal = nn.Transformer.generate_square_subsequent_mask(L_t,
device=src.device)
h = self.transformer(s_emb, t_emb, tgt_mask=causal)
return self.out(h) # (B, L_t, tgt_vocab)
# Tiny model: 10-word source vocab, 12-word target vocab
model = Seq2Seq(src_vocab=10, tgt_vocab=12, d_model=32,
nhead=4, n_layers=2, dff=64)
n = sum(p.numel() for p in model.parameters())
print(f"Seq2Seq parameters: {n:,}") # 60,364
# Example forward pass
src = torch.randint(0, 10, (2, 7)) # batch=2, src_len=7
tgt = torch.randint(0, 12, (2, 5)) # batch=2, tgt_len=5 (shifted right)
logits = model(src, tgt)
print(f"Logits shape: {logits.shape}") # (2, 5, 12)
Output:
Seq2Seq parameters: 60,364
Logits shape: torch.Size([2, 5, 12])
Train this with cross-entropy on parallel sentence pairs and you have a working translation model — the same architecture that produced Attention Is All You Need’s WMT-2014 results. Modern translation systems are scaled-up versions of this exact recipe.
Encoder-only, decoder-only, encoder-decoder
The full encoder-decoder we just built is one of three architectural families that emerged from the original Transformer paper. Knowing which variant a given model uses tells you immediately what it can and cannot do.
The 2017 paper used the full encoder-decoder for translation. Three variants emerged and dominate different use cases today:
| Variant | Examples | Best for | Cannot do |
|---|---|---|---|
| Encoder-only | BERT, RoBERTa, ModernBERT | Classification, NER, QA, embeddings | Generation |
| Decoder-only | GPT family, Llama, Qwen, Gemma, Claude | Text generation, chat, coding | Bidirectional understanding at training time |
| Encoder-decoder | T5, BART, mT5, Whisper | Translation, summarization, structured generation | Purely causal generation without conditioning |
For the next four parts of this series, we shift to the decoder-only variant — the architecture used by every chat LLM — and start working with real pre-trained weights through Hugging Face. We are done building from scratch. Now we use what we built to understand how the real systems work.
Decoder-only models: what changes
Because Parts 7–10 of this series focus entirely on the decoder-only family — GPT, Llama, Qwen, Claude — it is worth spelling out exactly what gets removed when you drop the encoder and cross-attention. The resulting architecture is simpler and, paradoxically, more powerful for pure generation tasks.
A decoder-only model is architecturally simpler than the full encoder-decoder:
- No encoder stack — the model processes only one sequence
- No cross-attention — each decoder layer has only the masked self-attention and FFN sub-layers
- Causal mask always on — every layer sees only past and current tokens
import torch
import torch.nn as nn
class DecoderOnly(nn.Module):
"""
Minimal decoder-only Transformer (GPT-style).
Causal mask applied at every layer so autoregressive generation is trivial.
"""
def __init__(self, vocab_size, d_model, num_heads, n_layers, dff, max_len=512):
super().__init__()
self.emb = nn.Embedding(vocab_size, d_model)
self.pos = nn.Embedding(max_len, d_model)
layer = nn.TransformerEncoderLayer(
d_model, num_heads, dff, dropout=0.1, batch_first=True
)
self.blocks = nn.TransformerEncoder(layer, num_layers=n_layers)
self.head = nn.Linear(d_model, vocab_size)
def forward(self, x):
T = x.size(1)
p = torch.arange(T, device=x.device).unsqueeze(0)
h = self.emb(x) + self.pos(p)
# Generate the causal mask for this sequence length
causal = nn.Transformer.generate_square_subsequent_mask(T, device=x.device)
h = self.blocks(h, mask=causal)
return self.head(h)
gpt_tiny = DecoderOnly(vocab_size=50257, d_model=768, num_heads=12,
n_layers=12, dff=3072)
n = sum(p.numel() for p in gpt_tiny.parameters())
print(f"GPT-2 small scale parameters: {n:,}")
Output:
GPT-2 small scale parameters: 162,692,689
This is in the same ballpark as GPT-2 small (124M). Our toy build over-counts because we use a separate output Linear head (~38M parameters) instead of tying it to the input embedding the way real GPT-2 does. The architecture — 12 stacked masked-attention + FFN blocks at d_model=768 and nhead=12 — is exactly what we have built across Parts 1-6; only the weight-sharing trick is missing.
The O(n²) attention bottleneck
Before we leave the ground-up architecture and shift to pre-trained models in Part 7, there is one fundamental limitation in everything we have built that is worth naming — it shows up in every real deployment decision about context length, inference cost, and memory usage.
Everything we have built in Parts 3–6 has one common bottleneck: self-attention is quadratic in sequence length. For a sequence of length n, the attention weight matrix is n × n. Each row requires a dot product with all n keys.
import time, torch
def timed_attn(n, d=64, warmup=2, reps=5):
Q = torch.randn(1, n, d)
K = torch.randn(1, n, d)
V = torch.randn(1, n, d)
for _ in range(warmup): # warm caches
_ = ((Q @ K.transpose(-2, -1)) / d**0.5).softmax(-1) @ V
times = []
for _ in range(reps):
t = time.time()
scores = (Q @ K.transpose(-2, -1)) / d**0.5
weights = scores.softmax(dim=-1)
_ = weights @ V
times.append(time.time() - t)
return min(times) # best-of-N
for n in [128, 512, 2048, 8192]:
print(f"n={n:5d}: {timed_attn(n)*1000:.2f} ms")
Output (CPU-only, exact ms vary by machine):
n= 128: 0.04 ms
n= 512: 0.18 ms
n= 2048: 2.86 ms
n= 8192: 46.50 ms
Time scales roughly with n². Going from 512 to 8192 tokens (16×) costs about 260× more compute on this run — even more than the theoretical 256× because the larger matrices spill out of cache. This is why the practical context window of early Transformer LLMs was 2K–4K tokens, and why techniques like Flash Attention, sliding-window attention, and linear attention emerged to push that limit to 128K+ tokens in modern LLMs.
The architectural foundation we have built is identical to what those systems use — they just implement the attention computation more efficiently. Flash Attention, for instance, computes the same result as our softmax(QK^T/√d)V formula but tiles the computation to stay in L2 cache rather than writing the full n × n matrix to HBM, cutting memory by O(n).
Up next — Part 7: Pre-trained Models & Tokenization
Building from scratch was the long way around — and a fantastic way to internalize what is happening inside an LLM. From here on we work the way real teams work in 2026: we download a pre-trained model and adapt it.
In Part 7: Pre-trained Models & Tokenization we will load a real, small pre-trained LLM with the Hugging Face transformers library, look at its sub-word tokenizer (which is much smarter than the word-level mapping we have implicitly assumed so far), and generate text — both deterministically (greedy) and stochastically (sampling). The tokenization comparison will show how a real BPE tokenizer slices up English vs Portuguese, and we will see why every chat-tuned model insists on a specific input format called a chat template.