7 stages. 2 adversarial loops. Nothing reaches you unverified, and nothing runs until you approve the plan.
This is the process for handing an agent real work and trusting what comes back. One prompt makes you the quality control, because every guess the agent makes is yours to catch. This workflow moves the checking inside the process: the plan is proven against your intent before the build, and the work is proven against the plan before you see it. It costs 2 to 3 times the tokens, and it buys back the hours you spend catching mistakes.
The simple version is a planner and a checker. This is the version with teeth, and it is the shape my own engines run every day.
Say what you want, rough is fine. The agent interviews you into the real thing: outcome, boundaries, and a done-when list. That list is the contract every stage below answers to.
Intent restated, every assumption tagged (told, inferred, or guessed), the steps, done-when carried through. No work yet.
Findings fold into a revised plan. Clean pass or 3 rounds.
One page: what it believes you want, what it assumed, what done means. You approve it or correct it, and nothing runs before that. The highest-value minute in the whole process.
Executes the approved plan, journals as it goes. Any deviation stops and asks. Deviating silently is the one sin the workflow never forgives.
Fresh eyes that never touched the build, ordered to refute it. Done-when list walked line by line, every claim traced to a source. P0, P1, or PASS.
Every finding gets fixed against the approved plan and your done-when list, and the work goes back to a fresh verifier. When it passes clean, it reaches you with receipts: the checks it ran, anything it caught, and the fixes it made.
The checker never grades its own homework. New context, every time.
Told to review, it approves. Told to refute, it finds the problem.
Clean pass or 3 rounds. Nothing polishes forever on your token bill.
Both loops in the workflow are this same machine. Only the thing being attacked changes: the plan in loop 1, the finished work in loop 2.
A hallucination is a guess nobody challenged. This workflow gives every guess 3 walls to die against.
Your context files already carry the business. The brief carries what they can't: what you want this time, and what done means. Every gap between the two gets filled with a guess, written down as fact, and built on. Ten minutes here outranks every loop downstream.
What this run produces, exactly.
The checks a finished result must pass. Write them, or the agent invents its own.
The lines it must not cross, however the work goes.
Anything true today that your context files don't know yet.
The whole workflow, ready to paste into any capable agent. Say what you want in your own words and stage 1 sharpens it with you.
I want: [say what you want in your own words. Rough is fine. Stage 1 turns it into a real brief with you.] Run this as a workflow. Do not skip a stage. HARD RULE. Every review below runs as a separate subagent with a clean context. No agent may review work it produced or helped produce, in any stage, ever. If you cannot spawn subagents, simulate this: quote the work back cold and review it with no memory of why you made each choice. 1. BRIEF. Interview me, one question at a time, until you can write the brief: the exact outcome, who it is for, what it must include, what it must never do, and a done-means list of checks the finished result must pass. If my opening message already covers all of that, skip the questions and go straight to the brief. Show me the brief and wait for my OK. 2. PLAN. My intent restated in your own words, every assumption you are making tagged (I told you, you inferred it, or you guessed), the steps, and the done-means list carried through. No work yet. 3. COUNCIL. Send the brief and the plan to 3 independent subagents, clean context each, one lens each. None of them saw the planning: - Intent: does this build what I asked, and nothing else? - Assumptions: find every inference dressed up as a fact, including any the plan failed to tag. - Failure: where does this plan break? Each pass ends in a verdict: P0 (wrong, stop), P1 (fix this named thing), or PASS. Fold findings into a revised plan and review again. Maximum 3 rounds. 4. MY APPROVAL. Show me the final plan as 1 page: what you believe I want, what you assumed, the steps, what done means. Do nothing further until I approve it. 5. BUILD. Execute the approved plan exactly. Keep a short journal. If a step needs to deviate from the plan, stop and ask. Never deviate silently. 6. VERDICT. Spawn an independent subagent that took no part in the plan or the build. Its only job is to refute. Walk the done-means list line by line. Trace every claim to a source. Flag anything the plan never authorised. P0, P1, or PASS. 7. FIX. Take every finding and fix the work against the approved plan and the done-means list. Send the result to a fresh verifier, stage 6 again. Repeat until it passes clean, maximum 3 rounds, then tell me if anything still fails. 8. DELIVER. Present the work with receipts: the checks you ran, anything you caught, and the fixes you made.
The agent does the work. You own it. Every plan waits for your approval, and every result answers to a done-when list you wrote. That holds at every level of autonomy, no matter how good the models get.
Ω