Download benchmark-template.csv

Free editable CSV for building your own AI-agent regression benchmark.

Download CSV

What the template is for

The template is designed for customer-service AI regression testing and vendor-neutral evaluation. It stores enough ground truth to replay a requirement after a model, prompt, knowledge-base or supplier change.

Core fields

The CSV includes an ID, language, prompt, expected outcome, allowed tolerance, forbidden behaviour, handoff rule, severity, estimated failure cost, source reference, source owner, effective date and notes. Not every field must be used on day one, but keeping them explicit prevents important assumptions from disappearing into evaluator prompts.

Who should approve the benchmark?

The business or policy owner should approve the expected outcome. An AI evaluator can help score future runs, but it should not decide what the company policy is. If the policy changes, update the effective rule and version the benchmark.

How many cases should you start with?

Start with 15–20 high-value cases if you are validating the process. Expand toward 50–100 or more when the same benchmark becomes part of release or recurring quality assurance. Prioritize high-volume intents, costly mistakes and mandatory human-handoff scenarios.

How to turn the CSV into automation

A minimal system imports the cases, calls the authorized agent interface, saves the response, applies deterministic checks and semantic evaluation, stores the result and compares it with the previous baseline. Schedule the runner or trigger it after changes.

Own the benchmark. Outsource the recurring run.

Useworthy can repeatedly run a customer-approved benchmark and preserve comparable evidence over time.

Discuss your test setup