Start with the work, not the agent.
Pick a task your team already does: collect sources, prepare a guest brief or check references in a script. Give the proposed automation the same inputs and compare its output with work your team would accept. Record where someone still has to check, correct or finish it.
The manual baseline is not an obstacle to automation. It reveals which decisions experienced people make without writing them down. If nobody can explain why a brief is accepted, a model’s fluent output can hide disagreement rather than resolve it. Measure a few real examples before making a time-saving claim.
Choose among manual, fixed and agentic work.
| Approach | Good candidate | Boundary |
|---|---|---|
| Manual editorial work | Choosing a premise, assessing a sensitive claim, resolving stakeholder disagreement. | Human work still needs evidence and a recorded decision. Familiarity is not a substitute for review. |
| Fixed automation | Rename known exports, validate required fields, generate links from a manifest, render an approved page. | Useful when the steps and rules are clear. Unexpected inputs should fail visibly rather than be guessed through. |
| Bounded model step | Draft an evidence-linked summary or propose alternative arguments from supplied material. | Treat the text as a candidate. Check factual meaning, missing qualifications and attribution. |
| Agentic workflow | A genuinely variable research task where the next permitted step depends on what has been found. | Needs limited tools, a stop rule, cost bounds and review of both outputs and actions. More autonomy increases the evaluation burden. |
Anthropic’s engineering discussion distinguishes predefined workflows from agents that dynamically direct their process. The distinction is useful even as products change. Begin with a simpler workflow when it can perform the task reliably; do not introduce dynamic tool use just to rename a sequence of fixed steps.
Start with an actual selection bottleneck.
Money Matters had a concrete selection problem: thousands of applications, each requiring enough attention to judge whether it could become a useful episode.
During his work with Ankur Warikoo, Rishwajeet built a workflow with two stages. The first checked demographic and structural eligibility. The second scored editorial potential using criteria informed by earlier episodes. A strategist reviewed the shortlist and made the final selection.
The separation matters. A submission can fit the show’s requirements without containing a strong conversation. Combining eligibility and story judgment into one unexplained score would make it harder to see why a candidate was rejected or what a reviewer should examine.
| Stage | Job in the published case | Question to answer when adapting the method |
|---|---|---|
| Eligibility | Check demographic and structural fit. | Are the criteria explicit, and what happens to an incomplete or ambiguous application? |
| Editorial scoring | Assess potential against patterns from earlier episodes. | Can the reviewer see the story evidence behind the assessment? |
| Strategist decision | Choose from the shortlist. | Which judgment remains with the person responsible for the programme? |
The public case reports that the initial workflow took four days to build and reduced weekly review from days to under 60 minutes. Those are historical figures from that operation, not a benchmark for another team. Before claiming an improvement in your own workflow, measure the complete accepted task, including review and correction.
The case is evidence for a useful division of work. It does not require calling every step an agent. Start by identifying the bottleneck, separating decisions that need different treatment and keeping the consequential editorial choice with its owner.
A complete research-to-brief workflow.
Suppose the team needs an editor brief from an approved script and source collection. The final deliverable must let an editor find the right material and understand the intended sequence.
- Intake, human: agree the audience, purpose, source permissions and who approves the final brief. Confirm that selected files exist and are accessible.
- Manifest, fixed rules: assign stable source IDs and check required fields. A missing file enters an exception list; it is not replaced by a guessed reference.
- Draft, bounded model: propose an outline and source-linked selects using only the supplied material. Mark uncertain claims and missing evidence. Do not invent timecodes.
- Checks, mixed: code can validate fields and reference IDs. An editor must assess whether a quote supports the claim, whether the sequence is faithful and whether the suggested visual exists.
- Approval, human: resolve exceptions and accept an exact version. Changes to meaning reopen the relevant decision.
- Delivery, authorised: generate the approved handoff and verify access at the intended destination. External sending or publication requires its own authorisation.
Keep failures recoverable. Store the source version, draft and review decision so a corrected input can be rerun without silently replacing the accepted brief. Give the exception queue an owner. A workflow that produces a draft but leaves inaccessible sources for the editor has moved work rather than completed it.
At the manual level, the producer builds the manifest and an editor writes and checks the brief. At the fixed-workflow level, file checks and document assembly run automatically while the same editorial decisions remain explicit. Optional agentic research would only enter when an approved evidence gap requires choosing among permitted sources; it returns a source packet for review before the brief changes. It does not acquire authority to invent missing clips or publish the result.
Catch an output that sounds more complete than the evidence.
Original teaching fixture. The memo and candidate below are invented for evaluation, not output from a measured production system.
Source memo
Tuesday: two producers tried the draft brief template. One could locate the selected clips. One could not open the source folder. No editing-time measurements were collected. Before another trial, the producer must repair access and the editorial lead must approve the claim wording.
Faulty candidate
“The new AI workflow cut editing time by 40% across the team. Both producers completed the pilot successfully. Automatically send the new brief to the client.”
| Problem | Required response |
|---|---|
| Invented 40% improvement | Reject. The source contains no editing-time measurements. |
| Two successful producers | Correct. One producer could not open the source folder. |
| Automatic client delivery | Block. Access repair and claim approval remain open; external sending is not authorised. |
Acceptable rewrite: Two producers tried the draft brief. One found the selected clips; the other could not open the source folder. Repair access before another trial. No editing-time improvement has been measured. The editorial lead still needs to approve claim wording.
A second model could help flag discrepancies, but agreement between models is not evidence. The reviewer should be able to trace every material statement to the memo. AP’s published generative-AI guidance treats output as unvetted source material. That is a useful verification principle; this guide is not a claim of AP endorsement or a reproduction of its newsroom rules.
Evaluate failures you would actually care about.
Build a small set of representative fixtures: a clean source, a missing file, contradictory accounts, an unsupported number, a changed version and a request outside the system’s authority. Record the expected decision before running the workflow. Keep some examples separate from prompt development so the evaluation can reveal unfamiliar failures.
Score concrete defects: unsupported claims, lost qualifications, incorrect source links and unauthorised actions. Also record human correction time, latency and cost. A schema-valid draft can still be substantively wrong. A polished average can hide a rare failure that makes the workflow unsuitable for its intended use.
Set release criteria from the task’s consequence and your actual baseline. Recheck affected fixtures when the prompt, model, tools or source structure changes. Do not convert one successful demonstration into a claim that the operation is autonomous.
Build, buy or keep the step manual.
Buy when an existing tool fits the task, permission model, exports and review flow closely enough to reduce maintenance. Build when a specific repeated gap justifies owning integrations, monitoring and recovery. Keep the work manual when volume is low or the judgment cannot yet be evaluated sensibly.
Compare total work: setup, source preparation, review, corrections, subscriptions, integration upkeep and recovery from failure. Test with representative material before committing. The useful question is whether the complete accepted task gets easier at an acceptable error rate, not whether the first draft appears instantly. Download the workflow, faulty-output fixture and evaluation fields.
Questions about the method.
When does a content operation need an AI agent?
When a useful task requires dynamically choosing among permitted next steps and that added autonomy can be evaluated and controlled. Many content tasks work well with fixed automation and a bounded model step.
Can AI check its own factual accuracy?
It can propose checks and flag inconsistencies, but its agreement with itself or another model does not establish truth. Material claims still need verification against appropriate evidence.
Tell us what you want to make.
Tell us about the channel or idea, the team you have and where you need help. We’ll discuss the scope on a 30-minute call.
Book a 30-minute call rishwajeet@machinehouse.media