Traces Are Not Explanations
- BY
- ROOT TEAM
- PUBLISHED
- AUGUST 31, 2026
- READING TIME
- 5 MIN READ
Every observability tool for agents can show you a thousand spans. Not one of them can point at the sentence that caused the failure. This post is about the missing discipline of failure attribution.
A team we know spent two days last month debugging an agent. They had a beautiful trace. Every tool call, every prompt, every token count, timestamps down to the millisecond. The trace was complete. It was also useless, because none of it said which single sentence in a forty page context window convinced the agent to take the wrong path. They found it eventually. By reading. All forty pages. By hand.
That story is becoming the default experience of anyone running agents in production, and it is worth being precise about why.
Observability answered the wrong question
The observability wave for agents did real work. It took debugging from a black box to a glass box. You can watch every step now. You can see the inputs and the outputs and the costs. This was the easy half of the problem and to be fair, most of the industry treated it as the whole problem.
But a trace is a recording, not a reason. It shows you the order things happened in. It does not show you which thing caused which. Those are different questions, and only one of them helps you fix anything.
The decision is made of context
Here is the part that makes this tractable, and it is the part most tooling skips. You do not need to open the model to find a cause. The model did not invent its behavior out of nothing. At every step, the agent decided based on three things only. The instruction it was given, the context it had accumulated, and the result of its last tool call.
Record those three things at every step and the decision becomes reconstructable. The why stops being a mystery inside a neural network and becomes an artifact you can walk backward through, the same way you walk a stack trace.
Symptom: the agent refunded the wrong customer
why
Decision: it matched on a name that appeared in two accounts
why
Context: a retrieved support ticket blurred the two accounts together
why
Origin: the retrieval pipeline injected a stale document nobody had reviewed
Notice where this lands. The cause was never the agent being clever or dumb. The cause was a pipeline feeding it bad context with no review step. That is a fixable system problem. But you only reach it if you can walk the decision backward, and walking it backward requires that the context was captured at the moment of the decision. Not summarized afterward. Captured.
What failure attribution actually means
Failure attribution is the discipline of ranking the parts of a prompt, a memory, or a tool output by how much they contributed to a bad decision. It has no standard tooling today, which is strange, because engineers do an informal version of it constantly. When you bisect a git history you are attributing a failure to a commit. When you comment out lines one by one you are attributing it to a line. Agents need the same instinct and the same mechanics.
Three mechanics make it real.
Capture the full state at each decision point. The exact context, not a reconstruction of it.
Checkpoint and replay. If you can restore the agent to the moment before a bad decision, you can change one input and watch whether the decision changes. That is the agent version of bisecting.
Compare branches. Run the path that worked next to the path that failed. The diff between the two paths is not just a debugging aid. It is the explanation, stated in the only language that cannot lie, which is the difference between what happened and what could have happened.
The uncomfortable conclusion
If your debugging process for agents ends at the trace viewer, your process ends before the diagnosis does. A trace tells you the agent did something wrong. Attribution tells you what to change so it does not happen again. Teams without the second thing start distrusting agents after every incident, and eventually they quietly stop using them. Not because the agents failed, but because the failures had no addressable cause.
We think the next generation of agent tooling will not be measured by how many spans it can show you. It will be measured by whether it can produce one honest sentence about which input caused the wrong decision. That sentence is the product. Everything else is decoration.
Try it this week
Take one failed agent run. Pull the full context at the step where it went wrong, not the summary. Read the context and underline the specific sentence that pushed the decision. Then ask why that sentence was there. The answer to that second why is usually the real bug, and it is usually not in the agent at all.