A support agent can sound convincing and still make the wrong decision. This experiment makes the evaluation loop visible before anyone trusts it with a real customer.
Replay the conversation, not just the answer
Each fixture follows a realistic support request through intent detection, policy retrieval, tool use and the final response.
The reviewer can see where the agent made its decision, what context it used and whether it crossed a permission boundary.
Turn failure into a useful signal
A human reviewer marks the trajectory against a small set of criteria: policy correctness, customer safety, action accuracy and explanation quality.
- Pass when the response and action agree with policy.
- Flag when the agent reaches a plausible but unsafe conclusion.
- Escalate when the evidence is incomplete or ambiguous.
The loop is the product
The output is not a leaderboard. It is a decision about what to change next: a prompt, a policy, a tool contract or an evaluation case.
Support agent evaluation
Review the cases your agent can’t finish alone.