The agent sounded great.
The grader failed it.
This is a verbatim replay of one internal test run. Every message, tool call and grade below is taken from the run log, timed for viewing. The insurer, the caller and every record are simulated. Press play to hear the call with two voiced speakers; no audio starts on its own. The grading is real code, checking words, actions and final system state against a written rulebook.
The call · policyholder wants to cancel
The grader · six checks, run by code at call end
retention_applied: absent → true
FAIL · 1 of 6 checks passed
The caller hung up happy. A spot-check of this call would likely pass it. The grader checked the rulebook and the database.
- forbidden action taken: {"tool": "apply_retention_discount"}
- state policies.HM-NJ-200871.retention_applied should be absent, got True
- missing required action (in order): {"tool": "send_form", "args": {"form": "CANCEL_REQUEST"}}
Internal Gate 1 run, 25 Sep 2026. One trial of four on this case; a single trial is not a prevalence estimate. Claude Sonnet 4.5, base model with no task-specific tuning. Harborline Mutual is a fictional insurer; "Priya Shah" and policy HM-NJ-200871 are simulated records. Overall Gate 1 aggregate figures remain withheld while a separate grading discrepancy is repaired and regraded.
What a buyer receives after a full run: the report card, every failure transcript, and the grading basis. See the sample report →