AI as a White Box — Trace a Tiny Language Model

AI as a White Box — Trace a Tiny Language Model

Fact-check status

Reviewed claim by claim on 2026-09-09 against the original Transformer paper, the GPT-2 technical report, official PyTorch documentation, and the executable learning resource identified below. “White box” means the computation can be inspected; it does not mean every learned representation has a complete human explanation. S1S3S7

Learning promise

After this lesson, you should be able to follow one token through a small decoder-only language model, name the shape and purpose of each important tensor, explain how a next-token mistake changes weights, and distinguish inspecting a mechanism from explaining what every internal feature means.

Black box, white box, and the boundary between them

View What you can see Question you can answer
Black box Prompt and response “What behavior did I observe?”
Glass box Observable inputs, context, retrieval, tool calls, configuration, outputs, and run traces “What observable path produced this result?”
White box Architecture, parameters, forward pass, loss, gradients, optimizer, and generation loop “How is this model computed and changed?”
Interpretability research Interventions, probes, circuits, and causal tests “What internal features or pathways contribute to a behavior?”

In this lesson, white box is a learning stance, not a claim that a neural network becomes fully understandable. Source code exposes the operations. Checkpoints expose learned numbers. Hooks expose activations and gradients. Those observations still require experiments before they support a semantic or causal explanation. Attention weights alone are not a complete explanation of a prediction, and the strength of that conclusion depends on the test being performed. S6

The whole machine in one picture

flowchart LR
    A["Text"] --> B["Tokenizer: token IDs"]
    B --> C["Token + position embeddings"]
    C --> D["Repeated decoder blocks"]
    D --> E["Vocabulary logits"]
    E --> F["Next-token distribution"]
    F --> G["Select one token"]
    G --> B
    E --> H["Cross-entropy loss during training"]
    H --> I["Backpropagation: gradients"]
    I --> J["Optimizer updates parameters"]
    J --> C

The upper loop is autoregressive inference: predict one token, append it, and run again. The lower loop is training: compare predictions with known next tokens, compute gradients, and update parameters. S1S3

Notation: a few shape symbols are enough

Symbol Meaning Tiny example
B Batch size: sequences processed together 2
T Sequence length in tokens 8
C Model width, also called d_model 32
V Vocabulary size 65
H Number of attention heads 4
D Width of each head, normally C / H 8

A tensor shape describes its axes, not its meaning. For example, [B, T, C] means one C-number vector for every token position in every sequence.

White-box forward pass

Assume the input tensor contains integer token IDs with shape [B, T].

1. Turn IDs into vectors

The token-embedding table has shape [V, C]. Looking up every ID produces token embeddings [B, T, C]. Positional information with compatible shape is added so the model can distinguish ordering. These values are learned parameters in common GPT-style implementations. S1

token IDs [B,T]
    -> token lookup [B,T,C]
    + position vectors [T,C]
    = initial residual stream [B,T,C]

An embedding is not a stored word definition. It is a learned vector that becomes useful through the rest of the network.

2. Let each position read permitted earlier positions

Each attention head projects the residual stream into queries, keys, and values:

Q, K, V: [B,H,T,D]
scores = (Q @ transpose(K)) / sqrt(D): [B,H,T,T]
masked scores -> softmax -> attention weights: [B,H,T,T]
attention weights @ V: [B,H,T,D]

The causal mask sets attention to future positions to an unusable value before softmax. Therefore position t can use positions 0..t, but not the answer at t+1. The heads are concatenated back to [B, T, C], projected, and added to the residual stream. S1

The sqrt(D) scaling keeps dot products from growing too sharply as head width increases. Softmax makes each permitted row non-negative and sum to approximately one. S1

3. Transform each position with an MLP

The feed-forward network applies learned linear transformations and a nonlinear activation independently at each token position. It changes the channel dimension internally, then returns to [B, T, C]. A residual connection adds the result back to the stream. Normalization placement varies by architecture; inspect the implementation rather than assuming one universal order. S1

A decoder block is therefore not “attention only”:

residual stream
  -> normalization -> causal self-attention -> residual addition
  -> normalization -> MLP                  -> residual addition

4. Produce scores for every possible next token

After the final block and normalization, a linear projection maps [B, T, C] to logits [B, T, V]. A logit is an unnormalized score. For generation, softmax converts the final position’s V logits into a probability distribution. Temperature, top-k, or other decoding rules may alter how a token is selected; they do not retrain the model. S2

How the model learns

Training uses shifted copies of the same sequence:

input:   [the, cat, sat]
target:  [cat, sat, down]

For every input position, cross-entropy compares the model’s raw vocabulary logits with the correct next-token ID. PyTorch’s CrossEntropyLoss combines log-softmax with negative log-likelihood, so training code should normally pass raw logits rather than applying softmax first. S4

One training step is:

  1. Forward: compute logits and loss.
  2. Clear gradients: remove gradients left from the previous step.
  3. Backward: automatic differentiation applies the chain rule from the loss back through every differentiable operation.
  4. Update: the optimizer changes parameters using their gradients and its update rule.
  5. Repeat: new batches produce new errors and updates. S3

A gradient says how a small parameter change would affect the current loss. It is not a stored explanation, a sentence, or the model’s private reasoning. The optimizer does not decide what a concept means; it follows a numerical update rule that tends to reduce the training objective. S3

How generation differs from training

Training Inference
Known text supplies both inputs and next-token targets Only the prompt and previously generated tokens are known
Loss and gradients are computed Gradients are normally disabled
Optimizer changes parameters Parameters normally remain fixed
Many positions can be scored in parallel under a causal mask New tokens are selected sequentially
Randomness may come from batching, dropout, and training configuration Randomness may come from sampling policy

At inference time, generation is a loop:

while not finished:
    crop to allowed context
    logits = model(token_ids)
    next_token_logits = logits[:, -1, :]
    probabilities = softmax(next_token_logits / temperature)
    next_id = select(probabilities)
    append next_id

Six-session agenda

Session Focus Inspect or build Evidence of understanding
1 Data path Tokenize a short corpus; print IDs, vocabulary, batch inputs, and shifted targets Explain why target position t is input position t+1
2 Forward pass Print every major tensor shape from embeddings to logits Reconstruct [B,T] -> [B,T,V] without notes
3 Causal attention Display one mask and one attention matrix; test row sums and future positions Explain why future-token leakage makes training invalid
4 Loss and learning Run one batch; inspect loss, one parameter, its gradient, and its value after step() Distinguish parameter, activation, gradient, and optimizer state
5 Generation Implement greedy selection, then temperature sampling Explain why different output does not imply different weights
6 Interventions Zero one head, shuffle position IDs, or remove the mask; compare measured effects Make a narrow causal claim supported by the intervention

Use 60–90 minutes per session. Stop at a tiny character- or subword-level model that fits on one machine. The goal is inspectability, not competitive language quality.

Hands-on path

Primary executable resource

Use Andrej Karpathy’s build-nanogpt as the guided implementation. Its commit history and linked lecture build a GPT-style model and training loop incrementally, making it more suitable for this lesson than starting from a production framework. S8

agentctl research record

On 2026-09-09, agentctl compared build-nanogpt, nanoGPT, and transformers-from-scratch for provenance, inspectability, dependency complexity, and coverage of training plus generation. It selected build-nanogpt because it is a lecture-linked, stepwise reconstruction with a small PyTorch-centered implementation. The resource URL and supporting source URLs were then independently checked. Earlier bounded agentctl calls to agy, codex, and claude timed out and contributed no claims.

Lab: make the invisible state visible

Work through the resource, but add your own inspection checkpoints. Do not merely run the finished code.

  1. Use a very small dataset and configuration: small V, short T, one or two blocks, a few heads, and narrow C.
  2. Fix the random seed and save the initial configuration.
  3. Print the shapes of token IDs, embeddings, Q/K/V, attention scores, attention output, residual stream, and logits.
  4. Assert that every attention row sums to about one and every future position receives about zero probability after masking.
  5. Before an optimizer step, record one parameter value, its gradient, and the loss. Record the parameter again after the step.
  6. Overfit one tiny batch. Confirm that its loss falls; do not confuse this with generalization.
  7. Generate from the same prompt with greedy selection and at least two temperatures. Keep weights fixed.
  8. Perform one intervention—remove positional information, disable one head, or deliberately break the causal mask—and record the measured result.
  9. Write a short conclusion that separates observation, inference, and unresolved questions.

The lab is complete only when you can point to the exact line that creates the causal mask, the loss, the backward pass, the optimizer update, and the next-token selection.

Debugging invariants

Check Expected result Likely fault if it fails
Logits shape [B,T,V] Output projection or reshape
Attention score shape [B,H,T,T] Head split or transpose
Permitted attention row sum Approximately 1 Wrong softmax axis
Attention above causal diagonal Approximately 0 Missing or inverted mask
Gradient exists after backward Non-None for trainable used parameters Detached graph, disabled gradient, or unused parameter
Parameter changes after step Small numerical change Missing backward(), step(), gradient, or learning rate
Tiny-batch loss declines Downward trend Data shift, loss shape, update, or capacity problem

Common misconceptions

Exit check

You are ready to move on when you can do all of the following without hand-waving:

Source map

Source Role in this lesson
S1 Original Transformer paper Scaled dot-product attention, masking, multi-head attention, residual connections, normalization, and feed-forward blocks
S2 GPT-2 technical report Decoder-style autoregressive language modeling and next-token generation
S3 Official PyTorch autograd tutorial Computation graphs, gradients, and backward differentiation
S4 Official PyTorch CrossEntropyLoss documentation Raw logits, class targets, and loss semantics
S5 Official PyTorch optimization tutorial Training-loop order and optimizer updates
S6 Jain and Wallace Empirical warning against treating attention weights as explanations by default
S7 Wiegreffe and Pinter Qualifications and tests needed for claims about attention as explanation
S8 build-nanogpt Primary executable learning path selected through agentctl research

Next

← AI as a Glass Box · AI Foundations · Apply this model-level understanding in AI Harness Engineering, where the model becomes one component inside a controlled system.