A specification says what must be true. It does not, by itself, say what evidence will be inspected or how a finding will be reached. That missing layer is where confident-looking AI output can escape serious scrutiny.
A verification plan translates each material acceptance condition into an observable check. NASA defines product verification as proof of compliance with specifications, using test, analysis, demonstration, inspection, or a combination of methods [1]. The same principle works for a report, estimate, risk register, or decision brief: define the evidence before the result exists.
This post stands alone. You can use the method once for a single deliverable, even if no improvement loop follows.
Design the Check Before Production
For every acceptance condition, specify the evidence, permitted source, method, possible finding, and any judgment reserved for a person. The three findings are deliberately limited:
MET: the available evidence supports the condition.
NOT MET: the evidence shows the condition failed.
UNRESOLVED: the evidence is insufficient or the decision belongs to a person.
UNRESOLVED is not a softer form of failure. It protects the review from false certainty when sources conflict, calculations cannot be reproduced, or professional judgment is required.
NIST recommends rigorous testing, uncertainty measures, benchmark comparisons, documented results, and independent review where appropriate. It also notes that independent review can reduce internal bias and conflicts of interest [2]. In practice, that means separating the producing role from the verifying role or, at minimum, starting the verification in a fresh context that does not inherit the producer’s self-assessment.
Use an Example for Form, Not Truth
An approved past example can show the expected structure, terminology, and level of detail. Language models can adapt from examples supplied in the prompt [3], which makes a good reference valuable.
But a reference is not evidence that the new work is correct. The verifier must check the new version against its approved specification and permitted sources. If the example conflicts with them, the conflict should be reported rather than copied.
The second AI is also not a final authority. Research on LLM judges documents position, verbosity, self-enhancement, and reasoning biases [4]. Independent AI review adds disciplined challenge; it does not replace accountable human approval.
Prompt 1: Design the Verification Plan
Design a concise verification plan for the work product defined in the attached
approved specification. Complete this before production begins.
For each material acceptance condition, define:
Condition | Observable evidence | Permitted source | Verification method |
Meaning of MET, NOT MET, and UNRESOLVED | Required human judgment
Use the attached approved example only to match the expected format, structure,
terminology, and level of detail. Do not treat it as evidence for this project.
Flag any conflict between the example, specification, or current sources.
Use UNRESOLVED when the evidence cannot support a finding or when a person must
make the decision. Ask me to correct or approve the plan. After approval, issue
it as a named version and stop before producing the work product.Prompt 2: Apply the Approved Plan
Independently apply the attached approved verification plan to this completed
work product version. Do not revise it and do not rely on the producer’s
self-assessment.
For each acceptance condition, report:
Evidence inspected | MET, NOT MET, or UNRESOLVED | Finding | Required correction
or human decision
Cite the exact source, calculation, test, or passage supporting each finding.
Do not treat polish or similarity to the format example as evidence of
correctness. Report overall MET only when every required condition is MET, NOT
MET when any required condition is NOT MET, and UNRESOLVED when any material
condition remains unresolved. End with all required human decisions.Keep Findings Separate From Actions
Verification reports what the evidence supports. It should not quietly rewrite the work product, relax the acceptance conditions, or decide whether a candidate should replace an earlier version.
For a one-time deliverable, a person can use the report to approve, revise, or reject the work. In a later bounded loop, the same findings can feed explicit KEEP, DISCARD, or HALT decisions. Keeping those decisions outside verification makes the review reusable and easier to audit.
For higher-consequence work, add requirement identifiers, source owners, calculation files, test environments, reviewer independence rules, and retained evidence. The basic method stays the same: condition, evidence, method, finding.
Questions
Must the verifier be a different AI model?
No. A different model can add diversity, but a separate role or fresh context is the minimum practical boundary. Critical findings still need the professional review required by the task.
Can the verifier fix problems as it finds them?
Not during the same review. Mixing correction and verification makes it difficult to know which version was actually checked. Report the finding first; revise and verify a new version separately.
What if a condition cannot be checked automatically?
Define the available evidence and return UNRESOLVED with the required human decision. A clear escalation is more useful than a fabricated pass.
References
1. National Aeronautics and Space Administration. (2016). NASA systems engineering handbook: Verification and validation definitions.
2. National Institute of Standards and Technology. (2023). Artificial Intelligence Risk Management Framework Core: Measure.
3. Brown, T. B., et al. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877–1901.
4. Zheng, L., et al. (2023). Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. Advances in Neural Information Processing Systems, 36.
Notes
Join the EPM Network to access insights, influence our research, and connect with a community shaping the industry’s future.
Support us by sharing this article with your friends and colleagues, or over social media.


