A&A INSIGHTS
How to verify the business result after AI says it is done
A verification design for AI integrators that distinguishes generated files, approval and actual destination state instead of relying on success messages.
日本語で読む
THE STARTING POINT
For an integration team, an AI completion message or a normal tool exit should be one piece of evidence. This hypothetical monthly-report workflow defines what to check from input through recipient access, how to recover safely and what evidence to retain.
Define completion in the recipient’s terms
A&A perspective
This article is for the person integrating AI document generation and delivery. Our starting point is a completion statement about the intended recipient. For example: the designated readers can open the approved report for the intended month at the expected destination. That statement gives the team distinct acceptance conditions for generation, approval, placement and access. It also makes partial completion describable without calling the entire workflow finished.
A&A perspective
Before a trial, prepare a controlled request with a specific period, organization, version, approver and destination. Carry the same artifact identifier from creation through verification. In this design, a similarly named file is insufficient: compare content and version as well. Use synthetic inputs and a destination controlled by the team where the check does not require real customer data.
Assign each success signal a limited meaning
Hypothetical example
Consider hypothetical failures. A generation tool exits normally and AI says the report is shared, but the new file is in a working folder while the shared destination still contains last month’s version. In another case, a delivery request is accepted but the controlled inbox has not received it. Our design records tool completion and request acceptance with those limited meanings, then performs a separate destination check.
A&A perspective
Do not treat a model’s description of a tool call as the call record itself. Compare it with the actual result, target and execution time. Include destination and version in the verification record so a previous run or another environment cannot satisfy the present request accidentally. Retain logs for diagnosis, while tying acceptance to destination readback and the defined approval conditions.
Compare producer and destination using the same version
A&A perspective
For the report example, first compare calculated results with the approval artifact. Verify that approval refers to that version, place it at the intended destination, and read it back with permissions representative of the intended reader. Compare period, key values and version. Where publication or sending is involved, stay within the authorized scope; verification should not create messages to unrelated recipients.
A&A perspective
If the destination cannot be inspected, preserve that limitation in the result: request accepted, delivery unverified. Define the numerator and denominator of any success metric accordingly. The proposal does not require equating unknown with failure; it requires keeping unknown separate from confirmed arrival. Agree with the business owner on the evidence available before starting routine operation.
Exercise three different assembly failures
A&A perspective
When changing how conditions or commentary enter a report, test inclusion, precedence and preservation separately. Inclusion checks all intended stores and notes. Precedence checks that this month’s approved definition wins even when an obsolete definition is present. Preservation checks that adding commentary does not rewrite approved figures or historical records. Each test targets a different failure mode rather than repeating the generated text as an expectation.
A&A perspective
Set expected results independently from the generator, using small examples approved for the workflow. Exercise missing-store inputs, competing old and new definitions, and historical records that must remain unchanged. Evaluate readable explanation separately. Correct figures with misleading commentary need revision; fluent commentary about the wrong version does not pass. This separation gives maintainers a more useful diagnosis than one overall quality score.
Recover the remaining work without duplicate delivery
A&A perspective
Start with a simple state record: generated, awaiting approval, placed and destination verified. If checking stops partway through, read the current destination before retrying. If the correct version is already there, no additional delivery is needed. If a revision is required, follow the agreed approval route. Set a retry bound and escalation condition rather than looping until the model emits a success phrase.
A&A perspective
To connect this work to operating value, measure correction and duplicate-handling effort alongside verification time. In the trial, include retries in total workflow time and count cases left unverified. The design does not promise zero failures. It aims to make the completed boundary and remaining work explicit enough that the team can assess the burden and hand off recovery accurately.
Hand off evidence together with unresolved checks
A&A perspective
A handoff record should name target and version, the operation performed, the destination observation, unresolved checks and the next owner. For example, approved version placed, access verification pending tells the next person whether generation needs repeating. Keep sensitive evidence where authorized operators can inspect it, while public reports state only the relevant status. Retain enough evidence to explain an incorrect completion decision later. If the verification method changes, preserve what earlier success records actually meant.
Treat an AI completion message as a prompt to check the result. Match target, version, approval and destination, then preserve any unverified portion explicitly. This makes the integration’s success claim assessable against what the intended user can actually access.
Sources & editorial note
Primary pages read for this article. Publication dates below belong to the sources; access dates record our research.
- Demystifying evals for AI agents
Anthropic · 2026-01-09
Accessed 2026-09-14 - Building effective agents
Anthropic · 2024-12-19
Accessed 2026-09-14
AI-assisted editorial production
A&A uses AI for research, writing, translation and editorial checks. Source facts, our analysis and hypothetical examples are labeled separately.
Editorial check: 2026-09-14