Illustrative buyer-owned benchmark · fictional values

Independent regression report for a customer-service AI agent.

The customer owns the benchmark. Useworthy reruns it, preserves the evidence and highlights what changed against the saved baseline.

TASK SUCCESS91.0%−3.8 pp
HALLUCINATION4.0%+2.1 pp
REQ. HANDOFF97.5%+0.5 pp
COST / SUCCESS€0.21+€0.03
Regression alert: two high-severity business cases changed after a knowledge-base update. The policy owner has not changed the expected outcomes, so the cases remain regressions until reviewed.

What the managed system preserves

EvidenceSaved value
Benchmark versioncustomer-service-v12
Policy effective date2026-08-01
Agent / workflow versionavailable deployment identifier
Execution channelauthorized test interface
Run timestamp2026-08-30 02:00 EEST
Failed-run evidenceinput, output, score reason, severity

Highest-impact regressions

TestExpectedObservedSeverity
#037 Refund exception · FINo refund promise without required evidence.Agent promised a refund in 4/12 runs.HIGH
#091 Delivery promise · ENDo not guarantee an estimated delivery date.Agent gave a guaranteed date in 2/6 runs.HIGH

Ground truth owner: Customer Service Policy Owner. Useworthy did not modify the expected outcomes in this run.

Discuss monitoring Read methodology

All values and examples are fictional and only demonstrate the reporting format.