Name the gate
- Specify the change under test.
- Bound the model job and retained human judgment.
Learn · Core skill lab
Turn real user work and known failures into a release gate that can detect whether one change helped.
Your work saves in this browser.
Field tools
Copy these into your interview, agent, review, or working document. They are specific to this repetition.
Use one record per task so results can be compared and audited.
ID / segment / source INPUT EXPECTED OUTCOME ACCEPTABLE VARIATION KNOWN TRAP DIMENSION SCORES + ANCHORS CRITICAL FAILURE? BASELINE OUTPUT / SCORE CHANGED OUTPUT / SCORE REVIEWER NOTE
Do this before changing the prompt again.
1. Which exact cases regressed? 2. What failure mechanism do they share? 3. Is the problem model, context, tool, UX, policy, or scorer? 4. Did a quality gain hide a critical failure? 5. What single intervention targets the mechanism? 6. Which full set must rerun before release?
Calibrate judgment
A triage prompt improves most cases but still follows an instruction embedded in a ticket.
Version 7 gains eight points and improves common billing tasks, but the prompt-injection case remains a critical failure. Revise and rerun all 15; the aggregate improvement does not override the release gate.
Why it works: The gate is pre-committed, critical failure outranks averages, and the next intervention targets a mechanism.
Tried five prompts on a few examples. Version 7 sounded more professional and got better answers most of the time, so ship it.
Why it fails: The workload, expected outcomes, scorer, critical failures, baseline, and regression threshold are not reproducible.
Review → revise → repeat
Check only standards your current artifact actually meets. Then record one consequential revision before exporting it.