Prompt Craft
A prompt that works once may fail on missing information or contradictory notes. A small repeatable test set makes revisions easier to judge.

Collect representative cases
Choose a normal case, a short incomplete case, a long case, a contradictory case and a case that should produce an explicit unknown. Use synthetic or approved material. Record the expected behaviour before looking at outputs, so you do not redefine success to match an attractive answer.
Write pass criteria in plain language
For meeting extraction, a pass might require all explicit commitments, no invented owner and an evidence reference for every row. These are more useful than “sounds professional”. Some criteria are binary; others require judgement. Keep those categories separate and write down why a borderline answer passed or failed.
Change one variable
Run the same cases with the existing prompt and the proposed revision. Keep the tool and relevant settings as consistent as you can, noting the date and model label shown by the product. Repeated runs can vary, so one win is not a stable measurement. Review whether a change fixes one case while damaging another.
Store failures as examples
When a prompt assigns a deadline that is absent from the notes, preserve the case and explain the failure. Add it to the set before another revision. A short collection of realistic failures is more useful than a growing pile of elaborate instructions with no evidence that they help. Retest after a meaningful product or workflow change.
An example to adapt
Evaluate this output against the following checklist. For each item, return pass, fail or uncertain with a short supporting excerpt. Do not improve the output yet. Checklist: [Insert task-specific criteria.] Input: [Insert original case.] Output: [Insert result.] A human will make the final judgement.
Illustrative starting point. Adapt it to your inputs, permissions and review process.
Before you use it
- Expected behaviour was written first.
- All variants use the same cases.
- Failures are retained.
- Results are not presented as universal model benchmarks.