Skip to content
By Dharit Shah, Hector Martinez, Adam Scerra
Picture of Dharit Shah
Dharit Shah
Picture of Hector Martinez
Hector Martinez
Picture of Adam Scerra
Adam Scerra

Tracing Internals ​

Fullsend's distributed tracing system records structured telemetry for every agent run. This document explains how the tracing implementation works and how to extend it. It is aimed at contributors modifying the telemetry package or the span instrumentation in the run command.

For enabling tracing on an installation, see How To Emit Traces. For span schemas and attribute definitions, see the Tracing Reference. The design is specified in ADR 0050.

How the tracing system is structured ​

The tracing code is split between two locations:

  • internal/telemetry/ - the TracerProvider, exporters, and W3C traceparent helpers.
  • internal/cli/run.go - span creation, attribute assignment, and trace context propagation to child scripts and the in-sandbox runtime.

The telemetry package owns the provider and exporters. run.go owns the span lifecycle: it decides when spans start and end, and which attributes they carry.

Package layout ​

internal/telemetry/
├── telemetry.go          Setup, TracerProvider wiring, env parsing
├── fileexporter.go       Synchronous OTLP JSON file exporter
├── trace.go              W3C traceparent formatting helpers
├── main_test.go
├── telemetry_test.go
├── fileexporter_test.go
├── otlpsink_test.go
└── trace_test.go

TracerProvider setup ​

telemetry.Setup(dir, version) returns a trace.Tracer and a cleanup function. It creates a TracerProvider with two span processors:

  1. fileExporter via SimpleSpanProcessor: writes every completed span as one OTLP JSON line to run-telemetry.jsonl. Calls f.Sync() after each write so spans survive a process crash.

  2. otlptracehttp via BatchSpanProcessor wrapped in parentSampledProcessor: exports to a remote backend when an OTEL_EXPORTER_OTLP_*ENDPOINT env var is set.

If neither exporter can be created (bad directory, SDK disabled), Setup returns a noop tracer. Telemetry failures never affect the run.

telemetry.Setup()
│
├── OTEL_SDK_DISABLED check   → noop tracer if "true"
├── os.OpenFile(jsonl)        → noop tracer on error
│
├── SimpleSpanProcessor(fileExporter)           ← always present
│
└── OTEL_EXPORTER_OTLP_*ENDPOINT check
    ├── validateEndpoints()   → stderr warning, skip
    ├── newOTLPExporter()     → stderr warning, skip
    └── parentSampledProcessor(BatchSpanProcessor(otlpExporter))

Why parentSampledProcessor exists ​

The OTel SDK's AlwaysSample sampler records all spans locally, which is needed for the file exporter. But AlwaysSample ignores an upstream unsampled decision. The parentSampledProcessor wraps the OTLP batch processor and suppresses the entire trace from export when the root span's remote parent has the W3C sampled flag unset (-00). It tracks suppressed trace IDs in a sync.Map so child spans under the same trace are also dropped.

Result: the file exporter always writes all spans; the OTLP exporter respects upstream sampling.

Endpoint validation ​

validateEndpoints rejects non-http(s) URLs and unsupported protocols before creating the exporter. A malformed endpoint produces a stderr warning; the SDK's default localhost:4318 fallback is never used silently.

Span lifecycle in run.go ​

run.go creates four span types arranged in a parent-child hierarchy:

run (root)
├── sandbox_create    (gen_ai.operation.name=create_agent)
└── agent             (one per iteration; gen_ai.operation.name=invoke_agent)
    └── execute_tool  (one per id-bearing tool call; gen_ai.operation.name=execute_tool)

Root span ​

Created by resolveTraceIdentity(). When an inbound TRACEPARENT env var is present, the W3C propagator extracts a remote span context and the root span is SpanKindConsumer. Otherwise it is SpanKindInternal.

Start attributes include: fullsend.agent, fullsend.work_item_id, gen_ai.operation.name, gen_ai.agent.name.

End attributes (set in a deferred cleanup): exit_code, fullsend.cost_usd, fullsend.num_turns, fullsend.tool_calls, fullsend.iterations.

For the full attribute list (including fullsend.security_trace_id and fullsend.prescript.*), see the Tracing reference attribute table.

sandbox_create span ​

Started before sandbox creation, ended after bootstrap completes. Carries gen_ai.operation.name=create_agent.

agent spans ​

One per iteration (validation loop iterations included). Started before Exec(), ended after output extraction.

The helper functions agentSpanStartAttrs() and agentSpanEndAttrs() build the attribute slices. Start attributes: iteration, gen_ai.operation.name, gen_ai.agent.name. End attributes: iteration, exit_code, gen_ai.system and gen_ai.provider.name (same serving-endpoint value), model, token counts, fullsend.cost_usd, fullsend.runtime, fullsend.tool_calls. Multi-provider runtimes resolve the provider from the effective model via runtime.GenAISystemFor; System() is only the fallback.

execute_tool spans ​

One per id-bearing tool call the runtime reports (Claude Code today), up to 1,024 per iteration, a child of that iteration's agent span, named execute_tool <tool name>. toolSpanTracker (internal/cli/tool_spans.go) starts the span when it handles the ToolUseEvent and ends it when it handles the matching ToolResultEvent. iterationEventHandler runs the console renderer first, then the tracker, then the Level 3 collector, so both timestamps are runner-side receipt instants on one clock, each trailing the sandbox by the pipe latency and the parser's decode of the line: the start also trails receipt by the renderer's output for the call (the renderer prints nothing for a result), and neither waits on the collector's redaction pass over its own event (events are handled one at a time, so with several calls open an instant can still trail the collector's pass over an earlier event, such as another call's large result) — the start is arguments-complete, not execution start. Attributes: gen_ai.operation.name=execute_tool, gen_ai.tool.name, gen_ai.tool.call.id; a result flagged is_error sets error.type=tool_error and status Error. A result whose stream line exceeded the parser's 1 MiB bound arrives as ToolResultEvent{Oversized} (the parser salvages the call id from the line's retained prefix) and ends its span at receipt marked fullsend.tool.result_oversized, status Unset and no error.type, since is_error was never decoded. A call still open when the stream ends — the runtime was stopped, or an over-long result line showed no id in that prefix — is closed as error.type=unanswered by Finish(), which runAgent calls once rt.Run returns, before content assembly; a call superseded by a second tool_use with the same id is ended the same way at the reuse; a result with no matching call (its tool_use line was skipped) becomes a near-zero-duration span marked fullsend.tool.unmatched=true. Events without an id — pi and codex emit none — produce no span, so the child count can be below fullsend.tool_calls, which counts every reported call, id or not; a server_tool_use block on an assistant line produces no event at all (its result never arrives as a tool_result), so it appears in neither count. The name passes through redactText — the runner env literal pass on both sides of security.OutputPipeline(), the redaction span content gets — and is bounded to 256 bytes before it becomes the attribute; the span name keeps at most 128 bytes of it. The call id goes through the same redactText and dropped from the span on any finding — never substituted, since a masked id could collide with another call's — while the raw bounded id still keys the open-call map, so correlation is unaffected. The tracker records at most maxToolSpansPerIteration (1,024) spans per iteration and reports the overflow, which runAgent records as fullsend.tool_spans.dropped on the agent span — a burst of agent-controlled calls must not fill the OTLP batch queue and evict the agent span, which ends after Finish() and content assembly. These spans are metadata: they carry no tool content and are emitted whether or not the Level 3 gate is on.

Level 3 content on agent spans ​

When the content-capture gate is on (telemetry.ContentCaptureEnabled()), runAgent constructs one contentCollector per iteration — iteration and agent span are 1:1, so a run-scoped collector would repeat earlier iterations' content on later spans — and tees the runtime's normalized event stream to it through RunParams.OnEvent.

The tee trap: supplying any OnEvent replaces the runtime's default console renderer (internal/runtime/claude.go), so the handler built by iterationEventHandler always calls the renderer first, then the tool-span tracker, then the collector — the tracker stamps span instants when it handles an event, so it must not wait on the collector's redaction pass over that event. The handler is always set — tool spans are emitted with the gate off — and with the gate off the collector is nil and inert, so console output stays byte-identical to the default renderer path.

The collector (internal/cli/content_collector.go) coalesces contiguous text/reasoning deltas, maps tool use to tool_call parts and tool results to tool_call_response parts (only the Claude parser emits ToolResultEvent, call ids and ToolUseEvent.Arguments today — pi and codex emit none of them, #7414; the schema's required result field is response), redacts every part — security.OutputPipeline(), with replaceEnvSecrets for the values of sensitive runner env keys and of providerOnlyKeys, in one pass, on both sides of it (redactText) — at assembly (redaction runs before the size budget — truncating first could split a secret past recognition), enforces a 256 KiB ordered-suffix budget (the ending survives — the final answer is what consumers judge) plus an 8 KiB per-tool-result bound (tail-kept, redacted before the cut, the part marked fullsend.truncated) and an 8 KiB per-call arguments bound (over it the arguments are dropped whole and the part marked, since a cut object is not JSON; over four times it as written they are dropped before they are decoded, so no walk does work the bound does not limit), with exact dropped-byte accounting across content, tool names, summaries, responses, part ids, and dropped arguments (counted as re-encoded JSON, or as redacted text when they were not a JSON value or were dropped before they were decoded; on a key collision each member a later key (in sorted order) replaced is not counted), then holds the marshaled string to maxEncodedContentBytes (255,000 — just under the one size the pilot backend is proven to accept) by trimming the oldest content again, measured on the encoding itself and still charged in raw bytes (a separate, earlier boundary — the parser's 1 MiB stream-line cap — skips oversized lines; an oversized tool_result line still yields an empty part marked fullsend.truncated). None of the four content bounds is a measured backend limit; the constants' comments and Size limits name what blocks raising them. The collector emits gen_ai.output.messages JSON following the GenAI output-messages schema, including the schema-required finish_reason from the iteration outcome. attachContent records the content and its marker attributes on the span before either finalizeAgentSpan path can end it, so failed iterations keep their content.

Arguments that parse as JSON are not redacted as one serialized text. The redactor's patterns are written for plain text, and over JSON they miss an assignment that opens a string or follows an escaped newline, a value behind escaped quotes, and JSON nested in a string; Unicode folding can also turn a fullwidth quotation mark into one that closes the string. toolArguments decodes the value, redacts each string and object key on its own (a number as its digits; one that redacts becomes the redacted string), and encodes the result again. A string value or key that holds one JSON object, array or string literal once scanned — what is exported — is judged decoded as well (heldSecret). The scan reads the string whole, since a secret can span a document's strings, but as text: it misses a secret-named member nested deeper than the value right after the key, and a secret only decoding spells out (an assignment that opens a string, a token or runner environment value written with an escape). So the document is walked the same way, its output and findings discarded; when the walk raises a finding — any but the normalizer's, save its removal of an escape sequence or of tag characters, which can carry text no pattern sees (digits as a CSI parameter, a token inside an OSC sequence); a secret the scan masked in place no longer shows — the string is masked *** whole, with one held_document finding. So is a document that cannot be judged: held more than maxHeldDepth (4) deep, each level being scanned once more; naming a member twice, since decoding keeps only the last (RFC 7493 forbids it); or shaped like a document — it begins and ends like one, white space and invisible characters aside, and quotes something — but not parsing, whether the scan broke it, folding or masking it, or it never parsed; or one the normalizer changes at all, as written: folding can move a value out of its secret-named member while the document still parses (fullwidth quotation marks), so structure the normalizer makes, or moves, is not trusted. That masks whole text with no secret in it too — a document with a repeated name, one a pattern misreads once decoded (a notebook line key = …), one holding an escaped terminal colour code, JSON with comments or trailing commas, an object literal or a Python dict with a quoted value, an Edit fragment shaped {…} with one — and a mask the walk matches again (a connection string's password of ten bytes or more, the runner environment marker in an assignment, a mask of eight bytes or more under a secret-named member) masks its string whole and counts its secret twice. A string the normalizer stripped an escape sequence or tag characters from — a value or a key — is masked *** whole: a colour code ends at the next letter, so the stripping can take a token's first letter and leave the rest, and a title code's payload is text no pattern sees. A member is named as written as well as as scanned, and a key masked whole names a secret whatever it was called — its name is not to be had from its mask, and stripping can eat the first letter of one folding spells — so the values under it are judged as a secret-named member's. Each level — the arguments, and each held document — is also judged as a reader of its keys, strings and numbers in the order written sees them (spanning): joined by a line break, and by nothing (a line of a notebook cell keeps its own line break), each as decoded and as the normalizer renders it, before any is masked, for a secret that can hold a line break — the private key block over the lines of an array or over a key and a value, an exact value (a runtime secret, a runner environment value) over the lines of one. The arguments are then dropped and charged as encoded; a held document is masked whole (its text was scanned whole first, so a block that scan masked in place stays in place). Every other pattern's secret stops at white space or at a quote, so a match over the break has only run its context into the next string, which is left unjudged. Not judged: structure in text not shaped like a document (several documents, YAML, a document inside code); an assignment, a header or a connection string whose value begins in the next string; a token or an exact value split into pieces between strings; a member-name pair in single quotes over two strings' apostrophes. Each string and number is also scanned once more, already redacted, beside the nearest enclosing key the member-name pattern (json_field) names, and masked whole when that pattern matches the pair — a pattern keyed on a member name has no other way to see the pair; that scan's other findings are discarded, since both strings were already scanned. It runs the pattern stage alone: the normalizer is not idempotent over escape sequences, and a second pass over the pair could strip the value or the key's keyword. A value in an array or a nested object under a secret-named member is judged as that member's own value would be — the pattern's eight-character floor included, so a short count stays; booleans and nulls are left as they are. Keys under such a member are judged the same way, since a credential can be the key: a key of eight characters or more is masked ***, field names included, and two masked alike collide (below). Arguments that are not one JSON value cannot be walked that way: they are scanned as text — with the misses above — so their findings count, then dropped and charged. So are arguments in which two keys of one object redact to the same string, where keeping either member would misreport the call. A call the stream reports without a name carries no arguments. This happens once, when the event is handled; eviction and Result do not rescan it. A name that redacts to nothing at Result takes the arguments with it: they are charged to the dropped bytes, and the part — kept only when it has a summary — is marked.

The input message does not come from the stream. When runAgent composes a retry prompt (buildFeedbackPrompt, under feedback_mode: append), attachInput redacts it with the collector's pipeline — redactFeedback ran before the feedback was sanitized and framed, and sanitizing can join a token that scan saw split, so the recorded copy is pattern-scanned again; the prompt the agent is sent is not changed, and the mask the first scan leaves for a connection-string password of ten or more bytes matches its own pattern and counts a finding — and sets gen_ai.input.messages on the span before the runtime starts, so each finalize path carries it. Its findings join the iteration's, and its encoded size is charged against maxEncodedContentBytes: the two content attributes of one span together stay within the proven size.

Consumer contract (for eval scorers and other readers of run-telemetry.jsonl): parse the gen_ai.output.messages attribute as JSON; check fullsend.content.truncated / fullsend.content.dropped_bytes before treating content as complete; masked secrets appear as the redactor's mask tokens, or as [REDACTED:<key>] for a runner env or provider-only value, and are counted in fullsend.content.redactions. A tool_call part's arguments, when present, is a JSON value (an object for the tools seen so far); a tool_call part marked fullsend.truncated had arguments that were dropped whole — on that part type the marker never means a partial value, as it does on a tool result. gen_ai.input.messages is absent unless the iteration is a retry that carried validation feedback — absence is the normal case, not a gap. Masks redactFeedback left in the feedback (abcd..., ***, [REDACTED:<ENV_KEY>]) are in gen_ai.input.messages too, but that function discards its findings, so they are not counted in fullsend.content.redactions; the connection-string mask described above is the known exception. The attribute names and shapes above are the consumption contract — see the Tracing reference.

Trace identity and TRACEPARENT propagation ​

resolveTraceIdentity() handles W3C trace context propagation in three steps:

  1. Extracts TRACEPARENT and TRACESTATE from env via the W3C propagator.
  2. Starts the root span (Consumer if remote parent, Internal otherwise).
  3. Computes the propagated traceparent with flag preservation: if the inbound parent was valid, remote, and unsampled, the outbound traceparent keeps the unsampled flag instead of the local AlwaysSample flag. This prevents child runs from re-advertising as sampled when the parent trace opted out.

The resulting TRACEPARENT string is passed to pre-scripts and post-scripts via childScriptEnv(). That function strips any inherited TRACEPARENT from os.Environ() and runner_env before appending fullsend's own value (issue #2779). That value is the run-root span.

A second, complementary path reaches the in-sandbox agent process. Before every iteration, runAgent starts the per-iteration agent span and writes that span's TRACEPARENT (same trace ID, agent span ID, flags from resolveTraceIdentity) into .fullsend/iteration.env via writeIterationEnv. The sandbox .env sources that file last, so the runtime sees the agent-span parent rather than the run-root value childScriptEnv() gives host-side scripts. An inbound unsampled parent stays -00 here too, so a runtime that honours W3C sampling does not export when the parent trace opted out.

File exporter output format ​

fileExporter writes OTLP JSON with hex-encoded trace/span IDs (per the OTLP JSON spec, not base64). Each line is a complete TracesData message. Non-finite floats (NaN, Infinity) are encoded as proto3 JSON strings per the protobuf spec.

buildResourceSpans() groups SDK spans by resource and instrumentation scope, preserving insertion order. This matches the structure a backend receives via OTLP/HTTP.

Prerequisites ​

  • A local clone of the fullsend repo
  • Go toolchain (see go.mod for the minimum version)
  • Podman or Docker, if testing with a local tracing backend

How to add a new span attribute ​

  1. Add the attribute.String / attribute.Int / attribute.Float64 call in run.go at the appropriate point. Use start attributes for values known when the span opens; use end attributes (set before span.End()) for values computed during the span.

  2. Choose the attribute key namespace:

  3. Update the attribute tables in the Tracing Reference.

  4. No exporter changes are needed; both exporters pick up new attributes automatically.

How to test ​

Unit tests ​

bash
go test ./internal/telemetry/...

Tests use t.Setenv for OTEL env vars and t.TempDir for the file exporter. No external backend is required.

Test seam for the OTLP exporter ​

newOTLPExporter is a package-level var that tests override to spy on or stub out exporter creation:

go
orig := newOTLPExporter
defer func() { newOTLPExporter = orig }()
newOTLPExporter = func(_ context.Context) (sdktrace.SpanExporter, error) {
    // spy or stub
}

Testing with a local backend ​

Start a Jaeger instance and point the exporter at it:

bash
podman run -d --name jaeger \
  -p 16686:16686 \
  -p 4318:4318 \
  jaegertracing/jaeger

export OTEL_EXPORTER_OTLP_ENDPOINT="http://localhost:4318"
go run ./cmd/fullsend run triage ...

See Running Agents Locally for full flags. View traces at http://localhost:16686.

Replaying collected traces ​

hack/upload-traces.sh replays run-telemetry.jsonl files into any OTLP backend using otelcol-contrib. This is useful for inspecting traces from CI runs or from another machine without re-running the agent.

bash
# Replay a single file
hack/upload-traces.sh run/agent-run-123/run-telemetry.jsonl \
  --endpoint http://localhost:4318

# Replay all .jsonl files under a directory
hack/upload-traces.sh run/ --endpoint http://localhost:4318

The collector runs continuously and watches for new data; press Ctrl+C when done.

Prerequisites: otelcol-contrib >= 0.120.0 on PATH (releases).

Authentication headers: otelcol-contrib ignores the standard OTEL_EXPORTER_OTLP_HEADERS env var. Edit hack/upload-traces-otelcol-config.yaml and uncomment the headers: block to set auth tokens or routing headers (e.g. x-mlflow-experiment-id).

Compatible backends ​

Any OTLP/HTTP-capable backend works. LLM-aware backends recognize the gen_ai.* attributes and surface GenAI dashboards (token cost rollups, prompt/completion inspection, agent-specific views) without CLI-side changes.

BackendLocal quickstartUI
Jaegerpodman run -p 16686:16686 -p 4318:4318 jaegertracing/jaegerlocalhost:16686
Arize Phoenixpodman run -p 6006:6006 -p 4318:4318 arizephoenix/phoenixlocalhost:6006
MLflow >= 3.6uvx mlflow serverlocalhost:5000
otel-guipodman run -p 4318:4318 ghcr.io/metafab/otel-gui:latestlocalhost:4318

See also ​