← Back to LabDebora’s Lab / Publication

Can you trust an agent before it reaches production?

A focused evaluation loop for replaying support conversations, reviewing failure modes and turning human feedback into the next agent version.

Written from practice↓ Read the field note

A support agent can sound convincing and still make the wrong decision. This experiment makes the evaluation loop visible before anyone trusts it with a real customer.

Replay the conversation, not just the answer

Each fixture follows a realistic support request through intent detection, policy retrieval, tool use and the final response.

The reviewer can see where the agent made its decision, what context it used and whether it crossed a permission boundary.

Turn failure into a useful signal

A human reviewer marks the trajectory against a small set of criteria: policy correctness, customer safety, action accuracy and explanation quality.

  1. Pass when the response and action agree with policy.
  2. Flag when the agent reaches a plausible but unsafe conclusion.
  3. Escalate when the evidence is incomplete or ambiguous.
03Support fixtures
04Quality criteria
01Next decision

The loop is the product

The output is not a leaderboard. It is a decision about what to change next: a prompt, a policy, a tool contract or an evaluation case.

Ttandem by Toloka

Support agent evaluation

Review the cases your agent can’t finish alone.

EVALUATION · SUPPORT-AGENT-EVAL-V1
Relay Support Agentv2.4 · 3 test cases · 4 evaluators
TEST CASEQUALITYSCORE
3 of 3 cases shown · last run 2 min ago