A good sentence can still fail the assignment.
Source: We cancelled the Monday recording because the guest was ill. The replacement session happened on Thursday.
Task: Give one exact quote explaining the cancellation.
Failing answer: “Production delays forced us to reschedule Monday’s recording.”
The answer invents a reason and puts words in quotation marks that the source never said. Smooth prose does not offset either error. A passing response could quote: “We cancelled the Monday recording because the guest was ill.”
An accurate paraphrase would still fail this particular task if presented as a verbatim quote. Grade against the requested job as well as the underlying facts.
Build a small set with different failure modes.
The downloadable test cases contain five original exercises: exact quotation, receipt arithmetic, conflicting schedule records, a required source gap, and an instruction embedded inside research material. Each includes the input and explicit checks. These are teaching fixtures, not reported benchmark results.
Run the cases through the same steps your team will actually use. If the system retrieves documents, test retrieval as well as the final draft. If a producer has already selected every source and written the answer into the prompt, record that preparation as part of the workflow.
Use a check that fits the failure.
Exact amounts and required fields can often be checked directly. Editorial usefulness needs judgment: does the answer help the next person make the intended decision? A quote checker can find matching text, but a reviewer still needs to assess attribution and context.
Anthropic’s evaluation guide distinguishes tasks, repeated trials and graders, and explains the roles of code, model and human review. For these content exercises, record factual checks separately from the editorial assessment. Calibrate any model-based judge against examples a human has reviewed.
Decide which errors block use before looking at results. A fabricated quotation should not disappear into an average score because the structure and tone were good.
Keep a record you can compare later.
Save the source bundle, prompt, system version, raw output and review decision for each attempt. Repeat tasks to see whether failures recur. Use the same inputs and review criteria when comparing a change, while recording model or retrieval changes that prevent a clean comparison.
Keep some real tasks aside from prompt development. Once the team has tuned against an example, it is useful as a regression check but weaker evidence of performance on unfamiliar work. The five public fixtures are a starting point; passing them alone does not establish readiness for your client material.
Count the work after the first answer.
The review sheet separates input preparation, human checking, repairs, machine elapsed time and direct tool costs. A two-minute generation followed by an hour of correction is a different operating result from a usable brief in two minutes.
Finish the evaluation with a specific decision: usable for this task with named review, revise and retest, or unsuitable for the current use. Record the remaining failure and who owns it. Widen the scope only when the additional kinds of work have their own evidence.
Questions about the method.
How many examples do we need?
Start with actual jobs and the failures that would make their output unusable. These five exercises teach the method; they are not a statistically sufficient sample for every workflow. Expand coverage as you find new task types and failure modes.
Can another AI grade the writing?
It can assist with a clear rubric, but inspect its agreement with human judgments and preserve the source evidence. Keep mechanical checks for facts that can be checked directly, and do not let a style score cancel a factual failure.
Does this evaluate a model or a whole system?
The intended unit is the workflow your team uses: inputs, retrieval, instructions, model, tools and human review. Record the components so a change can be traced and the result can be interpreted.
Tell us what you want to make.
Tell us about the channel or idea, the team you have and where you need help. We’ll discuss the scope on a 30-minute call.
Book a 30-minute call rishwajeet@machinehouse.media