Freeze the baseline
- Use a representative input.
- Save prompt, output, defects, score, and review time.
Build · Core skill lab
Improve an AI task by changing the information architecture, then prove the pack reduces defects or review time.
Your work saves in this browser.
Field tools
Copy these into your interview, agent, review, or working document. They are specific to this repetition.
Keep these layers separate so they can change at different speeds.
1. OBJECTIVE — task and audience 2. OUTPUT CONTRACT — schema and required fields 3. DURABLE RULES — definitions, constraints, forbidden moves 4. QUALITY BOUNDARY — paired good/bad examples with reasons 5. TASK CONTEXT — decision, source, current constraints 6. RETRIEVED EVIDENCE — only relevant excerpts 7. VERIFICATION — scorer, thresholds, human review
For every block, demand a reason it deserves attention.
Which error does this block prevent? Is it durable or task-specific? Could a schema replace prose? Could one paired example replace a paragraph? Is the source authoritative and current? What happens if this block is removed? Can the output cite which context it used?
Calibrate judgment
A transcript-to-evidence task is tested before and after a structured pack.
A schema, no-inference rule, and one paired example reduced unsupported claims from seven to one and review time from 28 to 11 minutes. Removing company history had no negative effect, so it stays out.
Why it works: The task and input stay constant, defects are counted, and context is pruned rather than celebrated by volume.
Added the company wiki, strategy deck, all research, brand guide, and a long role prompt. The answer sounded more informed.
Why it fails: There is no task contract, salience strategy, controlled comparison, defect measure, or evidence that the extra material caused improvement.
Review → revise → repeat
Check only standards your current artifact actually meets. Then record one consequential revision before exporting it.