AI as a Glass Box — Observe the Complete Run
AI as a Glass Box — Observe the Complete Run
Learning promise
After this lesson, you should be able to turn one AI interaction into a traceable run: reconstruct what entered the system, what context and tools it used, what configuration controlled it, what it returned, and which claims the evidence does not justify.
You do not need model weights, training data, source code, activations, gradients, or private chain-of-thought. You need observable events, stable identifiers, controlled experiments, and careful boundaries between evidence and inference.
Why a glass box belongs between black box and white box
| View | Access | Strongest question it answers |
|---|---|---|
| Black box | Prompt and final response | “What happened from the user’s point of view?” |
| Glass box | Observable request, context, retrieval, tools, configuration, timing, usage, validation, and output | “What path through the system produced this result?” |
| White box | Architecture, parameters, forward pass, activations, loss, gradients, and optimizer | “How did the model compute and learn numerically?” |
| Interpretability research | Probes, interventions, circuit analysis, and causal tests | “Which internal mechanisms contribute to this behavior?” |
A hosted model can be opaque internally while the surrounding application is highly observable. Conversely, owning an open-weight model does not automatically give you a good record of which prompt version, retrieved documents, tool results, or policy decisions produced a particular answer.
Glass-box rule: record the observable path, then make only the claims that path supports.
The observable run
flowchart LR
A["User request"] --> B["Instruction and context assembly"]
B --> C["Retrieval"]
C --> D["Model call"]
D --> E["Tool proposal"]
E --> F["Authorization and execution"]
F --> D
D --> G["Validation"]
G --> H["Final response"]
B -. "trace" .-> I["Run record"]
C -. "provenance" .-> I
D -. "model config and usage" .-> I
F -. "tool result" .-> I
G -. "checks" .-> I
H -. "outcome" .-> IOpenTelemetry models a trace as connected spans and defines developing semantic conventions for generative-AI operations. W3C PROV supplies a stable model for describing entities, activities, agents, and their lineage. Together they provide useful vocabulary for a glass-box run record. S1
Three evidence windows
1. Before the model call: what was assembled?
Inspect the actual request sent to the model, not only the text typed by the user.
| Evidence | Why it matters |
|---|---|
| User input or a protected reference to it | Distinguishes the original request from later additions |
| System and developer instruction versions | Explains constraints invisible in the chat box |
| Conversation slice | Reveals what history was included, summarized, or omitted |
| Retrieved document IDs, versions, chunks, and scores | Establishes provenance and retrieval behavior |
| Tool definitions and permissions | Shows which actions the model could propose |
| Model identifier and generation settings | Makes comparisons reproducible enough to interpret |
| Input token or usage count | Helps explain truncation, latency, and cost |
The important artifact is the assembled context. A polished answer cannot tell you afterward which hidden instruction, stale document, or missing message shaped it.
2. During the run: what path did the system take?
Represent each meaningful step as an event or span with a parent, start and end time, status, and sanitized attributes. A run may contain retrieval, reranking, several model calls, validation, human approval, and tools. OpenTelemetry’s GenAI conventions cover model and agent operations, but the specification warns that the conventions are still evolving. S1
For each tool call, separate:
- What the model proposed.
- What the application validated and authorized.
- What code actually executed.
- What result or error was returned to the model.
- What user-visible effect was committed.
This distinction prevents a tool-shaped model output from being mistaken for a completed action.
3. After the call: what evidence supports the result?
Record the final response, finish reason, usage, latency, validation results, citations, human decisions, and user-visible outcome. When an answer cites sources, verify that each source exists and supports the nearby claim; a retrieval score or citation string does not prove correctness.
The record should let you answer:
- Which exact configuration and context produced this response?
- Which documents and tool results were available to the model?
- Did an action execute, fail, retry, or remain only a proposal?
- Which checks passed or failed?
- What changed between this run and another run?
Minimum glass-box run record
The representation can be JSON, database rows, or trace spans. The conceptual fields matter more than this exact schema.
{
"run_id": "run-0187",
"started_at": "2026-09-09T14:32:10Z",
"task": "answer_with_sources",
"versions": {
"workflow": "sha256:...",
"instructions": "v4",
"model": "provider/model-version",
"retrieval_index": "2026-09-08"
},
"input": {
"content_ref": "protected://inputs/0187",
"token_count": 842
},
"retrieval": [
{"document_id": "doc-42", "version": "7", "chunk_id": "c3", "score": 0.81}
],
"steps": [
{"span_id": "s1", "type": "model", "status": "ok", "latency_ms": 920},
{"span_id": "s2", "type": "tool", "name": "search", "status": "ok"}
],
"validation": {
"schema": "pass",
"citation_support": "manual-review"
},
"output": {
"content_ref": "protected://outputs/0187",
"token_count": 311,
"finish_reason": "stop"
}
}
Use stable run, trace, span, document, tool-call, and effect identifiers so records can be connected without copying sensitive payloads into every log. W3C PROV’s entity/activity/agent relationships are useful when lineage must cross systems or organizations. S2
Observation is not explanation
| You observed | You may conclude | You may not conclude |
|---|---|---|
| A document chunk was placed in context | The model had access to that chunk | The model faithfully used every claim in it |
| A tool call was proposed | The model emitted a structured request | The action executed or was authorized |
| A tool returned a value | That value entered the recorded run | The value was correct or trusted |
| Output changed when temperature changed | Decoding configuration affected this trial | Temperature was the only possible cause unless other variables were fixed |
| Two identical requests produced different text | The observed system is not textually deterministic under those conditions | The provider changed the model |
| A citation appears in the answer | The response contains a source pointer | The source supports the claim |
| Latency rose | This run took longer | The model itself was necessarily the bottleneck |
A trace narrows hypotheses. A controlled experiment tests them.
Controlled glass-box experiments
Change one observable variable at a time and repeat trials when generation is stochastic.
| Experiment | Hold fixed | Change | Measure |
|---|---|---|---|
| Instruction test | Input, model, context, tools | Instruction version | Task success and failure mode |
| Retrieval test | Input, model, instructions | Retrieved chunks or reranker | Citation support and answer completeness |
| Decoding test | Prompt and model version | Temperature or selection rule | Output diversity and task accuracy |
| Tool test | Task and model | Tool available versus unavailable | Completion, errors, latency, and fallback behavior |
| Model comparison | Dataset, rubric, harness | Model identifier | Quality, latency, usage, refusals, and cost |
| Regression test | Fixed evaluation set and rubric | One release candidate | Newly introduced failures |
Do not read causality from a single before/after example when random sampling, service changes, retries, or hidden provider behavior could also differ. Record what you controlled, what you could not control, and how many trials you ran.
Five-session agenda
| Session | Focus | Build or inspect | Evidence of understanding |
|---|---|---|---|
| 1 | Reconstruct one call | Capture assembled instructions, context, model configuration, and output | Explain why the chat message is only part of the model input |
| 2 | Trace the path | Give retrieval, model, tool, and validation steps linked IDs and timing | Reconstruct the run in causal order |
| 3 | Establish provenance | Connect each retrieved chunk and tool result to its versioned source | Answer “where did this evidence come from?” |
| 4 | Run controlled trials | Change one variable, repeat trials, and compare results | Separate observation, inference, and causal claim |
| 5 | Protect and operationalize | Add redaction, access, retention, and an incident-ready run summary | Debug a failure without exposing unnecessary content |
Use 45–75 minutes per session. Any model interface is suitable if you can wrap it with your own recorder; optional provider telemetry can add evidence but is not required.
Hands-on lab: explain two different answers
- Choose a harmless question with a short, authoritative source document.
- Create a run ID and record the instruction version, model identifier, configuration, timestamps, and assembled context.
- Run the request once with the source document and once without it. Keep every other visible setting fixed.
- Record retrieval provenance, model usage, latency, final answer, citations, and validation results.
- Compare the runs claim by claim. Label each difference as observed, inferred, or unresolved.
- Repeat each condition at least three times if sampling is enabled.
- Introduce one failed tool call and verify that the trace distinguishes proposal, attempted execution, error, and fallback.
- Produce a one-paragraph incident summary using only the run record.
- Remove or restrict raw payloads and confirm that the remaining metadata still supports your analysis.
The lab is complete when another person can reconstruct the observable path and reproduce the comparison without relying on your memory.
Privacy and security boundary
Observability creates another sensitive data system. Prompts, retrieved text, tool arguments, outputs, and even URLs may contain personal data, credentials, proprietary content, or adversarial text. NIST’s Generative AI Profile treats privacy, information security, confabulation, and related risks as lifecycle concerns rather than model-only problems. S3
Apply these defaults:
- Collect the minimum content needed for a stated purpose.
- Prefer protected payload storage plus references in general traces.
- Redact secrets and personal data before ordinary logging.
- Do not use an unsalted hash as protection for guessable sensitive text; use access-controlled storage or a keyed construction where matching is required.
- Separate production access from authorized debugging access.
- Set retention and deletion rules.
- Record policy and schema versions.
- Test the redaction path itself.
- Never treat hidden chain-of-thought as required telemetry; record observable decisions, tool proposals, validations, and outcomes.
Common misconceptions
- “Glass box means open weights.” Open weights belong to white-box access; a glass box observes the run boundary and system path.
- “A trace is the truth.” A trace is evidence produced by instrumentation, which can be incomplete, misconfigured, or wrong.
- “Logging the user prompt is enough.” The model may receive additional instructions, history, retrieved data, tool definitions, and prior results.
- “Retrieved means used.” Retrieval proves availability, not faithful influence.
- “Tool call means action.” A proposal, authorization, execution, and committed effect are different events.
- “More logging is always better.” Excess collection increases privacy, security, retention, and access risk. S3
- “One changed output explains the cause.” Stochastic generation and uncontrolled variables require repeated, bounded experiments.
- “The trace reveals the model’s thoughts.” It exposes recorded operations and artifacts, not a privileged transcript of internal cognition.
Exit check
You are ready for the white-box lesson when you can:
- Draw the observable path from user request to final response.
- Reconstruct one run using identifiers, versions, timestamps, and provenance.
- Distinguish model input from user input.
- Distinguish tool proposal, authorization, execution, result, and effect.
- Explain what a retrieval record and a citation do—and do not—prove.
- Design a one-variable experiment with repeated trials.
- Diagnose a difference between two runs without inventing hidden causes.
- Protect sensitive content while preserving enough evidence to debug.
- State which questions require white-box access instead.
Source map
| Source | Role in this lesson |
|---|---|
| S1 OpenTelemetry GenAI semantic conventions | Traces, spans, metrics, events, model operations, usage, and the warning that exact conventions are still developing |
| S2 W3C PROV Overview | Stable provenance concepts connecting entities, activities, agents, use, generation, and derivation |
| S3 NIST AI 600-1 | Generative-AI lifecycle risk framing for measurement, monitoring, privacy, information security, and confabulation |
agentctl research record On 2026-09-09, agentctl evaluated these sources for a hosted-model glass-box lesson. It identified OpenTelemetry as the most specific observability vocabulary, W3C PROV as the stronger stable lineage model, and NIST AI 600-1 as risk guidance rather than a telemetry schema. Claims here preserve those narrower roles.
Next
Continue to AI as a White Box to open the model itself: tensors, attention, logits, loss, gradients, optimization, and autoregressive generation.