How to evaluate AI agents: test sets, metrics and monitoring
How to evaluate AI agents before and after launch: a test set from real cases, what to measure, automatic and model-based grading, safety and monitoring.
- Evaluate an agent the way you test software: a fixed set of real cases with known correct outcomes, run on every change.
- Measure the final outcome, the steps and tool calls, safety behaviour, and cost and speed, not just whether the reply sounds good.
- Use exact checks wherever outputs are structured; use a model as grader only for fuzzy qualities, and spot-check it.
- After launch, monitor real usage and turn every mistake into a new test case.
To evaluate an AI agent, build a test set from real past cases with known correct outcomes, define pass criteria for the result, the tool calls and safety, score automatically on every change, and keep monitoring after launch. Every mistake found in real use becomes a new test case.
Evals are what turn an impressive demo into a system you can rely on, and what let you change a prompt or switch models without guessing. This guide explains how to build them step by step, what to measure, how to grade, and what to watch once the agent is live.
Why agents need evals
Language models are not deterministic in the way ordinary code is. The same input can give slightly different outputs, and a small change in a prompt can fix one case and break three others. Without a fixed test set you can’t tell whether a change made things better or worse.
Evals answer three practical questions:
- Is the agent good enough to launch, and for which cases?
- Did the last change improve it or quietly break something?
- Can we switch to a cheaper or faster model without losing quality?
Evals also make conversations with the business easier. Instead of “it seems to work”, you can say which cases pass, which fail and what the agent does when it is unsure. That is what lets a team decide, with evidence, which tasks the agent may handle on its own.
Step 1: build a test set from real cases
Start with real past examples: emails, tickets, documents or conversations, each with the correct outcome written down. Our guide to AI agents for business operations suggests 50 to 200 to begin with. Make sure the set includes:
- Common, everyday cases in roughly the proportion they really occur.
- Hard cases: missing information, unusual wording, mixed languages.
- Cases the agent must refuse or escalate to a person.
- Cases with tricky inputs, such as text that tries to instruct the agent.
- Cases where the right answer is “I don’t know”.
Keep the test set under version control, remove personal data you don’t need, and agree the correct outcomes with the people who do the work today.
Step 2: decide what to measure
| Dimension | Example checks | How to score |
|---|---|---|
| Outcome | Right category, right fields extracted, right final record | Exact comparison with the expected result |
| Steps and tool calls | Called the right tool with the right values; no unnecessary steps | Compare the logged trace with expectations |
| Answer quality | Correct, complete, cites its source, no unsupported claims | Model-based grading with spot checks, or human review |
| Safety | Asked for approval when required; refused out-of-scope requests; ignored injected instructions | Pass/fail per case; any failure blocks release |
| Escalation | Handed hard cases to a person instead of guessing | Share of must-escalate cases escalated |
| Cost and speed | Text used per task, steps per task, time to finish | Measured from logs on every run |
Step 3: grade automatically, carefully
There are three ways to grade, and good eval suites use all of them:
- Exact checks for structured outputs: categories, extracted fields, tool names and arguments. Fast, cheap and reliable. Use them wherever you can.
- A model as grader for fuzzy qualities such as whether an answer is complete or polite. Give the grader a clear rubric and the reference answer, and check a sample of its judgements by hand, because graders make mistakes too.
- Human review for a sample of cases, especially early on and for high-risk tasks. It also catches problems the rubric missed.
For RAG systems, grade retrieval separately from the answer: was the right passage found at all? Our guide to RAG for business explains why most answer errors start there.
Step 4: run evals on every change
Treat the eval suite like automated tests in ordinary software. Run it before any change ships:
- A new or edited prompt.
- A new tool, or a change to an existing one.
- A different model or model version.
- Changes to documents, splitting or search settings.
Compare before-and-after scores case by case, not only the average. A change that improves the average but breaks a safety case should not ship. Because model outputs vary, run important cases more than once and look at consistency.
Step 5: evaluate during rollout
| Stage | What happens | What you learn |
|---|---|---|
| Shadow mode | The agent proposes actions; people keep doing the work | Real accuracy on live cases, and new test cases |
| Human approval | The agent prepares actions; a person approves each one | Share approved without edits; where people correct it |
| Limited autonomy | Proven low-risk categories run alone; others still go to people | Error rate on autonomous cases; escalation rate |
Step 6: monitor in production
Once live, keep watching:
- Edit and rejection rates on approved actions.
- Escalations, and cases where the agent should have escalated but didn’t.
- Unanswered or “I don’t know” questions, which show content gaps.
- Cost per task and time per task, against the pilot figures.
- Complaints or corrections from users, linked to the logged trace.
Review a sample of conversations or tasks every week. Each real mistake becomes a test case, so the eval suite grows with the agent.
Common eval mistakes
- Testing only easy cases. A suite of typical examples hides the failures that matter. Include hard, refusal and escalation cases on purpose.
- Writing test cases from imagination. Invented examples miss the messiness of real emails, scans and mixed languages. Use real past cases.
- Trusting a model grader blindly. Graders have their own biases and errors. Check a sample by hand, and prefer exact checks wherever possible.
- Looking only at the average. An overall score can rise while one important category gets worse. Read results by case and category.
- Letting the test set go stale. If production mistakes never become new cases, the suite stops reflecting reality.
- Ignoring cost and speed. A change that improves accuracy slightly but doubles the cost per task may not be worth shipping.
A minimal eval setup for a first pilot
You don’t need a special platform to start. A first pilot can run on a simple setup:
- A spreadsheet or file of real cases with the expected outcome for each.
- A script that runs the agent on every case and saves the full trace.
- Exact checks for structured outputs, plus a short rubric for anything fuzzy.
- A results table by case, with pass or fail and a note on why.
- A rule that nothing ships if a safety case fails.
Grow it as the agent grows: more cases, automatic runs before every release and dashboards for production monitoring.
Questions to ask any AI vendor
- How big is the test set, and where do the cases come from?
- What are the pass criteria, and who agreed them?
- Which checks are exact, and which use a model as grader?
- Do evals run automatically before every release?
- What do you monitor after launch, and how often do you review it?
Every agent we build through our AI agent development service ships with an eval suite built from the client’s real cases and monitoring for cost, accuracy and failures. Developers can learn the same practice in our Applied AI & Agents course.
Frequently asked questions
What are AI agent evals?
Evals are an automated test suite for an AI system: a fixed set of real cases with known correct outcomes, scored on every change, so you can see whether the agent got better or worse.
How many test cases does an AI agent need?
A practical start is 50 to 200 real past cases, covering common, hard, refusal and escalation cases. The set should grow as real mistakes are found.
What is LLM-as-a-judge?
Using a language model to grade outputs against a rubric and reference answer. It suits fuzzy qualities such as completeness, but its judgements should be spot-checked by people.
What should I measure when evaluating an AI agent?
The final outcome, the steps and tool calls, answer quality, safety behaviour, escalation, and cost and speed per task. Agree two or three headline numbers and launch thresholds in advance.
How do you test an AI agent for safety?
Include cases where the agent must ask for approval, refuse, escalate or ignore instructions hidden in inputs, and treat any failure in these cases as a release blocker.
Do evals stop after launch?
No. Monitor edit rates, escalations, unanswered questions and cost in production, review samples regularly and add every real mistake to the test set.