A prompt change made an agent's replies shorter and calmer. The review room preferred every new example. In production, the same change caused the agent to omit a currency condition from refund explanations. It sounded better and performed worse.
That is the central difficulty of agent evaluation. There may be several acceptable answers, but there are still unacceptable omissions, actions and claims. A system without one right sentence does not mean a system without testable behaviour.
Build the golden set before the next feature
A golden set is a stable collection of representative inputs with the evidence and expectations needed to judge outputs. It should contain ordinary work, difficult boundaries and known failures. Do not fill it with synthetic questions that conveniently match the prompt. Production traces, support corrections and rejected tool calls are better raw material.
For each case, store what matters: source context, allowed tools, required facts, forbidden actions and an example of an acceptable outcome. Some checks can be deterministic. A total must match the source. A cited identifier must exist. A write tool must not run without approval. Other checks need human or model grading against a rubric.
OpenAI's Evals API documentation describes evals as reusable testing criteria combined with a data source, which can be run across models and parameters. The reusable part matters more than the dashboard.
Compare outputs in pairs
Absolute scoring asks a reviewer to invent a stable internal scale. Pairwise comparison asks a narrower question: given the same input and evidence, is candidate A better, candidate B better, or are they equivalent? Reviewers are usually more consistent at that task.
Hide which version produced each answer. Randomise order. Give the rubric before the examples. Require a short reason tied to observable behaviour, not tone alone. Pairwise wins can reveal a direction, while deterministic checks still act as gates. A more helpful answer cannot compensate for an unauthorised action or a fabricated fact.
Use model graders for volume only after calibrating them against human decisions. If the grader disagrees systematically on a business-critical slice, it is measuring the wrong thing.
Slice the failures, do not average them away
A single aggregate score can improve while a vital workflow regresses. Break results down by task, language, customer tier, tool path, data source and risk class. Track refusal when the answer is available separately from confident answers when evidence is missing.
The smallest slice may be the one that matters commercially. A rare refund exception, permission boundary or German compound in a search query can disappear inside an overall average. Release gates should protect named critical slices even when the total looks healthy.
Keep the set alive
Golden does not mean frozen forever. Add a case when production reveals a new failure. Remove duplicates, but do not rewrite history to make the current system look better. Version the dataset, rubric, model, prompt, tools and retrieval configuration together. Otherwise a result cannot be reproduced.
Run the suite before changing models, prompts, tool descriptions, retrieval settings or response formatting. Sample live traffic after release to catch distribution shifts the set does not contain.
Establish a release gate this week
Collect recent corrected answers and failed runs. Choose a compact set covering the workflows that can cost money, lose trust or change data. Add deterministic assertions first, then a short pairwise rubric for relevance, completeness and groundedness.
Run the current production configuration against the proposed one. Inspect every regression in a critical slice. Ship only when known hard failures remain blocked and any tradeoff is explicit. "It seemed fine" records a mood. An eval records what changed, where it changed and whether the business accepts the difference.
