Illustrative buyer-owned benchmark · fictional values
Independent regression report for a customer-service AI agent.
The customer owns the benchmark. Useworthy reruns it, preserves the evidence and highlights what changed against the saved baseline.
TASK SUCCESS91.0%−3.8 pp
HALLUCINATION4.0%+2.1 pp
REQ. HANDOFF97.5%+0.5 pp
COST / SUCCESS€0.21+€0.03
Regression alert: two high-severity business cases changed after a knowledge-base update. The policy owner has not changed the expected outcomes, so the cases remain regressions until reviewed.
What the managed system preserves
| Evidence | Saved value |
|---|---|
| Benchmark version | customer-service-v12 |
| Policy effective date | 2026-08-01 |
| Agent / workflow version | available deployment identifier |
| Execution channel | authorized test interface |
| Run timestamp | 2026-08-30 02:00 EEST |
| Failed-run evidence | input, output, score reason, severity |
Highest-impact regressions
| Test | Expected | Observed | Severity |
|---|---|---|---|
| #037 Refund exception · FI | No refund promise without required evidence. | Agent promised a refund in 4/12 runs. | HIGH |
| #091 Delivery promise · EN | Do not guarantee an estimated delivery date. | Agent gave a guaranteed date in 2/6 runs. | HIGH |
Ground truth owner: Customer Service Policy Owner. Useworthy did not modify the expected outcomes in this run.
Discuss monitoring Read methodology
All values and examples are fictional and only demonstrate the reporting format.