AI as a White Box — Trace a Tiny Language Model
AI as a White Box — Trace a Tiny Language Model
Reviewed claim by claim on 2026-09-09 against the original Transformer paper, the GPT-2 technical report, official PyTorch documentation, and the executable learning resource identified below. “White box” means the computation can be inspected; it does not mean every learned representation has a complete human explanation. S1S3S7
Learning promise
After this lesson, you should be able to follow one token through a small decoder-only language model, name the shape and purpose of each important tensor, explain how a next-token mistake changes weights, and distinguish inspecting a mechanism from explaining what every internal feature means.
Black box, white box, and the boundary between them
| View | What you can see | Question you can answer |
|---|---|---|
| Black box | Prompt and response | “What behavior did I observe?” |
| Glass box | Observable inputs, context, retrieval, tool calls, configuration, outputs, and run traces | “What observable path produced this result?” |
| White box | Architecture, parameters, forward pass, loss, gradients, optimizer, and generation loop | “How is this model computed and changed?” |
| Interpretability research | Interventions, probes, circuits, and causal tests | “What internal features or pathways contribute to a behavior?” |
In this lesson, white box is a learning stance, not a claim that a neural network becomes fully understandable. Source code exposes the operations. Checkpoints expose learned numbers. Hooks expose activations and gradients. Those observations still require experiments before they support a semantic or causal explanation. Attention weights alone are not a complete explanation of a prediction, and the strength of that conclusion depends on the test being performed. S6
The whole machine in one picture
flowchart LR
A["Text"] --> B["Tokenizer: token IDs"]
B --> C["Token + position embeddings"]
C --> D["Repeated decoder blocks"]
D --> E["Vocabulary logits"]
E --> F["Next-token distribution"]
F --> G["Select one token"]
G --> B
E --> H["Cross-entropy loss during training"]
H --> I["Backpropagation: gradients"]
I --> J["Optimizer updates parameters"]
J --> CThe upper loop is autoregressive inference: predict one token, append it, and run again. The lower loop is training: compare predictions with known next tokens, compute gradients, and update parameters. S1S3
Notation: a few shape symbols are enough
| Symbol | Meaning | Tiny example |
|---|---|---|
B | Batch size: sequences processed together | 2 |
T | Sequence length in tokens | 8 |
C | Model width, also called d_model | 32 |
V | Vocabulary size | 65 |
H | Number of attention heads | 4 |
D | Width of each head, normally C / H | 8 |
A tensor shape describes its axes, not its meaning. For example, [B, T, C] means one C-number vector for every token position in every sequence.
White-box forward pass
Assume the input tensor contains integer token IDs with shape [B, T].
1. Turn IDs into vectors
The token-embedding table has shape [V, C]. Looking up every ID produces token embeddings [B, T, C]. Positional information with compatible shape is added so the model can distinguish ordering. These values are learned parameters in common GPT-style implementations. S1
token IDs [B,T]
-> token lookup [B,T,C]
+ position vectors [T,C]
= initial residual stream [B,T,C]
An embedding is not a stored word definition. It is a learned vector that becomes useful through the rest of the network.
2. Let each position read permitted earlier positions
Each attention head projects the residual stream into queries, keys, and values:
Q, K, V: [B,H,T,D]
scores = (Q @ transpose(K)) / sqrt(D): [B,H,T,T]
masked scores -> softmax -> attention weights: [B,H,T,T]
attention weights @ V: [B,H,T,D]
The causal mask sets attention to future positions to an unusable value before softmax. Therefore position t can use positions 0..t, but not the answer at t+1. The heads are concatenated back to [B, T, C], projected, and added to the residual stream. S1
The sqrt(D) scaling keeps dot products from growing too sharply as head width increases. Softmax makes each permitted row non-negative and sum to approximately one. S1
3. Transform each position with an MLP
The feed-forward network applies learned linear transformations and a nonlinear activation independently at each token position. It changes the channel dimension internally, then returns to [B, T, C]. A residual connection adds the result back to the stream. Normalization placement varies by architecture; inspect the implementation rather than assuming one universal order. S1
A decoder block is therefore not “attention only”:
residual stream
-> normalization -> causal self-attention -> residual addition
-> normalization -> MLP -> residual addition
4. Produce scores for every possible next token
After the final block and normalization, a linear projection maps [B, T, C] to logits [B, T, V]. A logit is an unnormalized score. For generation, softmax converts the final position’s V logits into a probability distribution. Temperature, top-k, or other decoding rules may alter how a token is selected; they do not retrain the model. S2
How the model learns
Training uses shifted copies of the same sequence:
input: [the, cat, sat]
target: [cat, sat, down]
For every input position, cross-entropy compares the model’s raw vocabulary logits with the correct next-token ID. PyTorch’s CrossEntropyLoss combines log-softmax with negative log-likelihood, so training code should normally pass raw logits rather than applying softmax first. S4
One training step is:
- Forward: compute logits and loss.
- Clear gradients: remove gradients left from the previous step.
- Backward: automatic differentiation applies the chain rule from the loss back through every differentiable operation.
- Update: the optimizer changes parameters using their gradients and its update rule.
- Repeat: new batches produce new errors and updates. S3
A gradient says how a small parameter change would affect the current loss. It is not a stored explanation, a sentence, or the model’s private reasoning. The optimizer does not decide what a concept means; it follows a numerical update rule that tends to reduce the training objective. S3
How generation differs from training
| Training | Inference |
|---|---|
| Known text supplies both inputs and next-token targets | Only the prompt and previously generated tokens are known |
| Loss and gradients are computed | Gradients are normally disabled |
| Optimizer changes parameters | Parameters normally remain fixed |
| Many positions can be scored in parallel under a causal mask | New tokens are selected sequentially |
| Randomness may come from batching, dropout, and training configuration | Randomness may come from sampling policy |
At inference time, generation is a loop:
while not finished:
crop to allowed context
logits = model(token_ids)
next_token_logits = logits[:, -1, :]
probabilities = softmax(next_token_logits / temperature)
next_id = select(probabilities)
append next_id
Six-session agenda
| Session | Focus | Inspect or build | Evidence of understanding |
|---|---|---|---|
| 1 | Data path | Tokenize a short corpus; print IDs, vocabulary, batch inputs, and shifted targets | Explain why target position t is input position t+1 |
| 2 | Forward pass | Print every major tensor shape from embeddings to logits | Reconstruct [B,T] -> [B,T,V] without notes |
| 3 | Causal attention | Display one mask and one attention matrix; test row sums and future positions | Explain why future-token leakage makes training invalid |
| 4 | Loss and learning | Run one batch; inspect loss, one parameter, its gradient, and its value after step() | Distinguish parameter, activation, gradient, and optimizer state |
| 5 | Generation | Implement greedy selection, then temperature sampling | Explain why different output does not imply different weights |
| 6 | Interventions | Zero one head, shuffle position IDs, or remove the mask; compare measured effects | Make a narrow causal claim supported by the intervention |
Use 60–90 minutes per session. Stop at a tiny character- or subword-level model that fits on one machine. The goal is inspectability, not competitive language quality.
Hands-on path
Primary executable resource
Use Andrej Karpathy’s build-nanogpt as the guided implementation. Its commit history and linked lecture build a GPT-style model and training loop incrementally, making it more suitable for this lesson than starting from a production framework. S8
agentctl research record On 2026-09-09, agentctl compared build-nanogpt, nanoGPT, and transformers-from-scratch for provenance, inspectability, dependency complexity, and coverage of training plus generation. It selected build-nanogpt because it is a lecture-linked, stepwise reconstruction with a small PyTorch-centered implementation. The resource URL and supporting source URLs were then independently checked. Earlier bounded agentctl calls to agy, codex, and claude timed out and contributed no claims.
Lab: make the invisible state visible
Work through the resource, but add your own inspection checkpoints. Do not merely run the finished code.
- Use a very small dataset and configuration: small
V, shortT, one or two blocks, a few heads, and narrowC. - Fix the random seed and save the initial configuration.
- Print the shapes of token IDs, embeddings,
Q/K/V, attention scores, attention output, residual stream, and logits. - Assert that every attention row sums to about one and every future position receives about zero probability after masking.
- Before an optimizer step, record one parameter value, its gradient, and the loss. Record the parameter again after the step.
- Overfit one tiny batch. Confirm that its loss falls; do not confuse this with generalization.
- Generate from the same prompt with greedy selection and at least two temperatures. Keep weights fixed.
- Perform one intervention—remove positional information, disable one head, or deliberately break the causal mask—and record the measured result.
- Write a short conclusion that separates observation, inference, and unresolved questions.
The lab is complete only when you can point to the exact line that creates the causal mask, the loss, the backward pass, the optimizer update, and the next-token selection.
Debugging invariants
| Check | Expected result | Likely fault if it fails |
|---|---|---|
| Logits shape | [B,T,V] | Output projection or reshape |
| Attention score shape | [B,H,T,T] | Head split or transpose |
| Permitted attention row sum | Approximately 1 | Wrong softmax axis |
| Attention above causal diagonal | Approximately 0 | Missing or inverted mask |
| Gradient exists after backward | Non-None for trainable used parameters | Detached graph, disabled gradient, or unused parameter |
| Parameter changes after step | Small numerical change | Missing backward(), step(), gradient, or learning rate |
| Tiny-batch loss declines | Downward trend | Data shift, loss shape, update, or capacity problem |
Common misconceptions
- “White box means fully explained.” Visible code and tensors expose mechanism, not a complete semantic account. S6
- “One neuron or weight stores one fact.” Do not infer a one-fact-to-one-weight map from activation or parameter inspection without a causal test.
- “Attention shows what the model was thinking.” Attention is an internal computation that can be measured, but an attention map alone does not establish a faithful causal explanation. S6
- “Softmax chooses the token.” Softmax creates a distribution; the decoding rule selects from it.
- “Backpropagation updates weights.” Backpropagation computes gradients; the optimizer uses them to update weights. S3
- “Lower training loss proves intelligence.” It only shows better fit to the specified training objective on the measured data.
- “From scratch means no libraries.” In this agenda it means no high-level transformer abstraction; PyTorch still supplies tensor kernels and automatic differentiation.
- “The model searches the internet while generating.” A bare decoder transforms its current token context with fixed parameters. External retrieval requires a surrounding system.
Exit check
You are ready to move on when you can do all of the following without hand-waving:
- Draw and label the complete path from text to next-token ID.
- State the shapes of
Q,K, attention scores, residual stream, and logits. - Explain why the causal mask is required during parallel training.
- Show how shifted targets, cross-entropy, gradients, and an optimizer connect.
- Point to one parameter before and after an update and explain the difference.
- Change temperature without claiming that the model learned.
- Design an intervention that tests one narrow causal hypothesis.
- Say what remains unexplained after every tensor is visible.
Source map
| Source | Role in this lesson |
|---|---|
| S1 Original Transformer paper | Scaled dot-product attention, masking, multi-head attention, residual connections, normalization, and feed-forward blocks |
| S2 GPT-2 technical report | Decoder-style autoregressive language modeling and next-token generation |
| S3 Official PyTorch autograd tutorial | Computation graphs, gradients, and backward differentiation |
S4 Official PyTorch CrossEntropyLoss documentation | Raw logits, class targets, and loss semantics |
| S5 Official PyTorch optimization tutorial | Training-loop order and optimizer updates |
| S6 Jain and Wallace | Empirical warning against treating attention weights as explanations by default |
| S7 Wiegreffe and Pinter | Qualifications and tests needed for claims about attention as explanation |
S8 build-nanogpt | Primary executable learning path selected through agentctl research |
Next
← AI as a Glass Box · AI Foundations · Apply this model-level understanding in AI Harness Engineering, where the model becomes one component inside a controlled system.