AI as a Glass Box — Observe the Complete Run

AI as a Glass Box — Observe the Complete Run

Fact-check status

Reviewed on 2026-09-09 against OpenTelemetry’s generative-AI semantic conventions, the W3C PROV family, and NIST AI 600-1. OpenTelemetry’s GenAI conventions are marked Development, so exact attribute names are version-sensitive; the underlying lesson is provider-neutral. S1S3

Learning promise

After this lesson, you should be able to turn one AI interaction into a traceable run: reconstruct what entered the system, what context and tools it used, what configuration controlled it, what it returned, and which claims the evidence does not justify.

You do not need model weights, training data, source code, activations, gradients, or private chain-of-thought. You need observable events, stable identifiers, controlled experiments, and careful boundaries between evidence and inference.

Why a glass box belongs between black box and white box

View Access Strongest question it answers
Black box Prompt and final response “What happened from the user’s point of view?”
Glass box Observable request, context, retrieval, tools, configuration, timing, usage, validation, and output “What path through the system produced this result?”
White box Architecture, parameters, forward pass, activations, loss, gradients, and optimizer “How did the model compute and learn numerically?”
Interpretability research Probes, interventions, circuit analysis, and causal tests “Which internal mechanisms contribute to this behavior?”

A hosted model can be opaque internally while the surrounding application is highly observable. Conversely, owning an open-weight model does not automatically give you a good record of which prompt version, retrieved documents, tool results, or policy decisions produced a particular answer.

Glass-box rule: record the observable path, then make only the claims that path supports.

The observable run

flowchart LR
    A["User request"] --> B["Instruction and context assembly"]
    B --> C["Retrieval"]
    C --> D["Model call"]
    D --> E["Tool proposal"]
    E --> F["Authorization and execution"]
    F --> D
    D --> G["Validation"]
    G --> H["Final response"]

    B -. "trace" .-> I["Run record"]
    C -. "provenance" .-> I
    D -. "model config and usage" .-> I
    F -. "tool result" .-> I
    G -. "checks" .-> I
    H -. "outcome" .-> I

OpenTelemetry models a trace as connected spans and defines developing semantic conventions for generative-AI operations. W3C PROV supplies a stable model for describing entities, activities, agents, and their lineage. Together they provide useful vocabulary for a glass-box run record. S1

Three evidence windows

1. Before the model call: what was assembled?

Inspect the actual request sent to the model, not only the text typed by the user.

Evidence Why it matters
User input or a protected reference to it Distinguishes the original request from later additions
System and developer instruction versions Explains constraints invisible in the chat box
Conversation slice Reveals what history was included, summarized, or omitted
Retrieved document IDs, versions, chunks, and scores Establishes provenance and retrieval behavior
Tool definitions and permissions Shows which actions the model could propose
Model identifier and generation settings Makes comparisons reproducible enough to interpret
Input token or usage count Helps explain truncation, latency, and cost

The important artifact is the assembled context. A polished answer cannot tell you afterward which hidden instruction, stale document, or missing message shaped it.

2. During the run: what path did the system take?

Represent each meaningful step as an event or span with a parent, start and end time, status, and sanitized attributes. A run may contain retrieval, reranking, several model calls, validation, human approval, and tools. OpenTelemetry’s GenAI conventions cover model and agent operations, but the specification warns that the conventions are still evolving. S1

For each tool call, separate:

  1. What the model proposed.
  2. What the application validated and authorized.
  3. What code actually executed.
  4. What result or error was returned to the model.
  5. What user-visible effect was committed.

This distinction prevents a tool-shaped model output from being mistaken for a completed action.

3. After the call: what evidence supports the result?

Record the final response, finish reason, usage, latency, validation results, citations, human decisions, and user-visible outcome. When an answer cites sources, verify that each source exists and supports the nearby claim; a retrieval score or citation string does not prove correctness.

The record should let you answer:

Minimum glass-box run record

The representation can be JSON, database rows, or trace spans. The conceptual fields matter more than this exact schema.

{
  "run_id": "run-0187",
  "started_at": "2026-09-09T14:32:10Z",
  "task": "answer_with_sources",
  "versions": {
    "workflow": "sha256:...",
    "instructions": "v4",
    "model": "provider/model-version",
    "retrieval_index": "2026-09-08"
  },
  "input": {
    "content_ref": "protected://inputs/0187",
    "token_count": 842
  },
  "retrieval": [
    {"document_id": "doc-42", "version": "7", "chunk_id": "c3", "score": 0.81}
  ],
  "steps": [
    {"span_id": "s1", "type": "model", "status": "ok", "latency_ms": 920},
    {"span_id": "s2", "type": "tool", "name": "search", "status": "ok"}
  ],
  "validation": {
    "schema": "pass",
    "citation_support": "manual-review"
  },
  "output": {
    "content_ref": "protected://outputs/0187",
    "token_count": 311,
    "finish_reason": "stop"
  }
}

Use stable run, trace, span, document, tool-call, and effect identifiers so records can be connected without copying sensitive payloads into every log. W3C PROV’s entity/activity/agent relationships are useful when lineage must cross systems or organizations. S2

Observation is not explanation

You observed You may conclude You may not conclude
A document chunk was placed in context The model had access to that chunk The model faithfully used every claim in it
A tool call was proposed The model emitted a structured request The action executed or was authorized
A tool returned a value That value entered the recorded run The value was correct or trusted
Output changed when temperature changed Decoding configuration affected this trial Temperature was the only possible cause unless other variables were fixed
Two identical requests produced different text The observed system is not textually deterministic under those conditions The provider changed the model
A citation appears in the answer The response contains a source pointer The source supports the claim
Latency rose This run took longer The model itself was necessarily the bottleneck

A trace narrows hypotheses. A controlled experiment tests them.

Controlled glass-box experiments

Change one observable variable at a time and repeat trials when generation is stochastic.

Experiment Hold fixed Change Measure
Instruction test Input, model, context, tools Instruction version Task success and failure mode
Retrieval test Input, model, instructions Retrieved chunks or reranker Citation support and answer completeness
Decoding test Prompt and model version Temperature or selection rule Output diversity and task accuracy
Tool test Task and model Tool available versus unavailable Completion, errors, latency, and fallback behavior
Model comparison Dataset, rubric, harness Model identifier Quality, latency, usage, refusals, and cost
Regression test Fixed evaluation set and rubric One release candidate Newly introduced failures

Do not read causality from a single before/after example when random sampling, service changes, retries, or hidden provider behavior could also differ. Record what you controlled, what you could not control, and how many trials you ran.

Five-session agenda

Session Focus Build or inspect Evidence of understanding
1 Reconstruct one call Capture assembled instructions, context, model configuration, and output Explain why the chat message is only part of the model input
2 Trace the path Give retrieval, model, tool, and validation steps linked IDs and timing Reconstruct the run in causal order
3 Establish provenance Connect each retrieved chunk and tool result to its versioned source Answer “where did this evidence come from?”
4 Run controlled trials Change one variable, repeat trials, and compare results Separate observation, inference, and causal claim
5 Protect and operationalize Add redaction, access, retention, and an incident-ready run summary Debug a failure without exposing unnecessary content

Use 45–75 minutes per session. Any model interface is suitable if you can wrap it with your own recorder; optional provider telemetry can add evidence but is not required.

Hands-on lab: explain two different answers

  1. Choose a harmless question with a short, authoritative source document.
  2. Create a run ID and record the instruction version, model identifier, configuration, timestamps, and assembled context.
  3. Run the request once with the source document and once without it. Keep every other visible setting fixed.
  4. Record retrieval provenance, model usage, latency, final answer, citations, and validation results.
  5. Compare the runs claim by claim. Label each difference as observed, inferred, or unresolved.
  6. Repeat each condition at least three times if sampling is enabled.
  7. Introduce one failed tool call and verify that the trace distinguishes proposal, attempted execution, error, and fallback.
  8. Produce a one-paragraph incident summary using only the run record.
  9. Remove or restrict raw payloads and confirm that the remaining metadata still supports your analysis.

The lab is complete when another person can reconstruct the observable path and reproduce the comparison without relying on your memory.

Privacy and security boundary

Observability creates another sensitive data system. Prompts, retrieved text, tool arguments, outputs, and even URLs may contain personal data, credentials, proprietary content, or adversarial text. NIST’s Generative AI Profile treats privacy, information security, confabulation, and related risks as lifecycle concerns rather than model-only problems. S3

Apply these defaults:

Common misconceptions

Exit check

You are ready for the white-box lesson when you can:

Source map

Source Role in this lesson
S1 OpenTelemetry GenAI semantic conventions Traces, spans, metrics, events, model operations, usage, and the warning that exact conventions are still developing
S2 W3C PROV Overview Stable provenance concepts connecting entities, activities, agents, use, generation, and derivation
S3 NIST AI 600-1 Generative-AI lifecycle risk framing for measurement, monitoring, privacy, information security, and confabulation
agentctl research record

On 2026-09-09, agentctl evaluated these sources for a hosted-model glass-box lesson. It identified OpenTelemetry as the most specific observability vocabulary, W3C PROV as the stronger stable lineage model, and NIST AI 600-1 as risk guidance rather than a telemetry schema. Claims here preserve those narrower roles.

Next

Continue to AI as a White Box to open the model itself: tensors, attention, logits, loss, gradients, optimization, and autoregressive generation.