Here is a failure mode in AI agent evaluation that will not show up in your test suite if you are only checking final answers. An agent gets the right result by pure luck. It called the wrong tool, misread the output, then stumbled into a correct answer through a second mistake that happened to cancel the first. Your eval marks it green. In production, on a slightly different input, those two mistakes will not cancel, and now you have a confident wrong answer and no idea why.
This is the core problem with evaluating agents the way we evaluated plain LLM calls. A single model call has one output, so you grade the output. An agent has a whole path of decisions, tool calls, retries and reasoning steps between the request and the response. If you only look at where it ended up, you are blind to how it got there, and how it got there is exactly what determines whether it will keep working.
Trajectory is the thing you are actually testing
The trajectory is the sequence of steps an agent takes: which tools it invoked, in what order, with what arguments, what came back, and how it responded to each result. Good AI agent evaluation treats that trajectory as a first-class object to test, not just as debug logs you read after something breaks.
Once you look at the trajectory, a set of questions opens up that final-answer grading cannot reach. Did the agent pick the right tool for the job, or did it reach for a search when it already had the data? Did it call tools in a sensible order, or waste three steps before doing the obvious first thing? Did it pass correct arguments? Did it recover gracefully when a tool errored, or did it spiral? Did it stop when it had enough, or keep going and burn tokens? Each of those is a distinct way an agent can be quietly bad while still, sometimes, producing the right answer.
The AI agent evaluation metrics worth tracking
A few trajectory-level measures tend to earn their place. None of them is exotic, and together they tell you far more than an accuracy score alone.
Tool selection accuracy: how often the agent chooses the correct tool for the step it is on. This is often the single most diagnostic metric, because most agent failures trace back to reaching for the wrong capability.
Trajectory match or overlap: how close the agent’s actual path is to a reference path an expert would take. You can score this loosely, by whether the right tools appeared at all, or strictly, by exact order. Loose is usually more useful, because there is often more than one reasonable path.
Step efficiency: how many steps the agent took versus the minimum it needed. A rising step count on unchanged tasks is an early warning that a prompt or model change has made the agent less decisive.
Error recovery: when a tool returns an error or an empty result, does the agent adapt or fall apart? This is where brittle agents reveal themselves, and it is almost never tested by final-answer evals.
Completion and stopping: did the agent actually finish the task, and did it know when to stop? Agents that never stop are as much a problem as agents that stop too early.
How to grade a trajectory without going mad
AI agent evaluation gives you three broad options here, and mature setups use all three at different points.
Deterministic checks are the cheapest and most reliable where they apply. If a task should always call a specific tool, or should never touch a particular system, you can assert that directly against the trajectory. Use these wherever the correct behaviour is unambiguous, because they never flake.
Reference trajectories work when you can write down what a good path looks like for a task. You compare the agent’s run against the reference and score the overlap. This is powerful for a curated test set, though it costs effort to build and maintain the references.
Model-graded, or LLM-as-judge, evaluation is what you reach for when the quality of a step is a matter of judgement rather than a hard rule. A judge model reads the trajectory and rates whether each decision was reasonable. It scales well and handles nuance, but it introduces its own noise. So you calibrate it against human judgement on a sample before you trust it, and you keep checking that it still agrees.
Start smaller than you think
You do not need all of this on day one. Trying to build it all at once is how teams end up with an eval system more complicated than the agent. It also raises the build versus buy question: which parts of this evaluation harness to own and which to rent. Start with the failure that would hurt most in production, usually the agent confidently taking a wrong action, and write the trajectory checks that would catch it. Add tool selection accuracy, because it is cheap and diagnostic. Grow the rest as real failures teach you what you were not measuring.
The mindset shift is the whole point. Stop asking only whether the agent was right, and start asking whether it was right for the right reasons. An agent that reaches good answers through sound decisions will keep doing so as inputs shift. An agent that reaches good answers through luck is a production incident that has not happened yet.
When we help teams put agents into production at humaineeti, this is usually where the real work sits. It is building the trajectory-level evaluation that tells you an agent is genuinely reliable, not just occasionally correct. It is less glamorous than the agent itself, and it is the part that decides whether the thing survives contact with real users.





Comments 1