A Five-Case Test Set for Your Reusable Prompt

Prompt Craft

A prompt that works once may fail on missing information or contradictory notes. A small repeatable test set makes revisions easier to judge.

Prompt evaluation worksheet with normal, edge and failure cases
Original editorial diagram by Prompt Task Journal. Illustrative worksheet, not an application screenshot.

Collect representative cases

Choose a normal case, a short incomplete case, a long case, a contradictory case and a case that should produce an explicit unknown. Use synthetic or approved material. Record the expected behaviour before looking at outputs, so you do not redefine success to match an attractive answer.

Write pass criteria in plain language

For meeting extraction, a pass might require all explicit commitments, no invented owner and an evidence reference for every row. These are more useful than “sounds professional”. Some criteria are binary; others require judgement. Keep those categories separate and write down why a borderline answer passed or failed.

Change one variable

Run the same cases with the existing prompt and the proposed revision. Keep the tool and relevant settings as consistent as you can, noting the date and model label shown by the product. Repeated runs can vary, so one win is not a stable measurement. Review whether a change fixes one case while damaging another.

Store failures as examples

When a prompt assigns a deadline that is absent from the notes, preserve the case and explain the failure. Add it to the set before another revision. A short collection of realistic failures is more useful than a growing pile of elaborate instructions with no evidence that they help. Retest after a meaningful product or workflow change.

An example to adapt

Evaluate this output against the following checklist. For each item, return pass, fail or uncertain with a short supporting excerpt. Do not improve the output yet.
Checklist: [Insert task-specific criteria.]
Input: [Insert original case.]
Output: [Insert result.]
A human will make the final judgement.

Illustrative starting point. Adapt it to your inputs, permissions and review process.

Before you use it

  • Expected behaviour was written first.
  • All variants use the same cases.
  • Failures are retained.
  • Results are not presented as universal model benchmarks.