How to Check AI-Generated Software Against Acceptance Criteria
AI-generated code can look complete and still miss the intended behavior. Check it against the ratified requirements, criterion by criterion, and leave acceptance to an authorized person.
An AI coding agent adds pause and resume to a subscription service. The change includes an API, billing logic, audit events, confirmation email behavior, and tests. It is a substantial candidate, but that is not enough to accept it.
A product owner or engineering leader accountable for acceptance must decide whether the candidate meets the ratified requirements and acceptance criteria.
The acceptance decision depends on details the implementation cannot settle by itself. An overdue account must not be able to pause. A failed confirmation email must not undo a scheduled pause. The ratified contract version also requires a production response-time measurement after release.
Start the review with the ratified requirements and acceptance criteria rather than the diff or agent transcript. Loft keeps the ratified contract version, candidate revision, and criterion-level evidence connected so an authorized person can decide whether to accept the result.
Establish the acceptance basis before reviewing the code
A reviewer needs a stable account of what the change is meant to do. A product contract is a human-ratified, structured specification of what a bounded product area should do, the decisions and constraints that govern it, who may change it, and what evidence matters for acceptance.
For the pause-and-resume change, that contract covers more than the happy path. It says when a pause starts, who is eligible, what happens to billing, how early resume behaves, which audit events are written, and what the customer is told. It also records the production latency target and the conditions under which it must be measured.
Without that ratified basis, review tends to ask whether the generated code is plausible. Plausibility is too weak for acceptance. Several implementations can be plausible while handling billing dates, overdue accounts, or failed email delivery differently.
The Loft product keeps proposals separate from the ratified contract and records who may change it. This gives the coding agent a bounded implementation scope without giving the agent authority over product intent.
Keep implementation and evaluation tied to the same version
Requirements, tests, and code all change. A review becomes ambiguous when the agent implements one ratified contract version and the reviewer evaluates another.
The ratified contract version should travel with the work. The implementation scope names that version and the affected requirements. The candidate revision identifies the code under review. Any later change to product intent becomes a proposed contract change rather than an informal correction inside the implementation.
Linking the candidate to that version identifies what the agent was asked to build and shows when a passing test is credited against a criterion that changed later.
In the fictional example, the coding agent receives the ratified contract version for pause and resume, plus the affected requirements. Each candidate revision is then checked against that version.
Map evidence to every acceptance criterion
A green test suite is useful, but it is not an acceptance record. Tests written with the implementation can repeat the same mistaken assumption, omit a required failure path, or say nothing about behavior that appears only in production.
Loft calls this comparison reconciliation: the implementation revision, tests, measurements, and other evidence are checked against the ratified contract version, then each in-scope acceptance criterion or quality target is reported as supported, failed, missing evidence, or unresolved.
The four statuses are:
- Supported: a reviewed test for the criterion passed at that candidate revision.
- Failed: a reviewed scenario contradicts the criterion.
- Missing: no reviewed evidence maps to the criterion yet.
- Unresolved: the required evidence cannot exist yet under the conditions named in the ratified contract version.
A missing failure-path test may be added before the acceptance decision. A production latency target may remain unresolved until the software runs under the conditions named in the ratified contract version. Calling both of them "not tested" would hide what can be fixed now and what must remain open.
Keep each item of evidence attached to the criterion it addresses. A broad statement such as "all tests passed" forces the reviewer to reconstruct coverage. A criterion-level map shows which reviewed test, measurement, or other record bears on each requirement, and it leaves gaps visible.
A candidate with substantial test coverage may still need another iteration
Sample Co., the subscription service, its roles, code, and values are fictional. The specification and review are real Loft records, and the candidate revisions, test runs, and reconciliation records are real runs in a sample repository. The first candidate includes a deliberately planted overdue-account defect. The complete record is available in Loft's worked product example.
The first candidate implements much of the pause-and-resume behavior. Reconciliation records three different findings.
One acceptance criterion fails: an account marked overdue by the billing job can still pause. Another criterion is missing evidence: nothing shows that a failed confirmation email leaves the pause scheduled. The production response-time target can be measured only from production traffic.
The product owner sends the candidate back with the failed review scenario and the missing evidence requirement. The coding agent produces a second revision. A reviewed test for every in-scope acceptance criterion now passes, and the product owner accepts Candidate 2. The production quality target remains unresolved and is recorded as follow-up work after release.
Preserve human authority at both ends of the review
Agents can inspect code, implement changes, generate tests, and prepare evidence. Assurance specialists can add independent scenarios and judge whether a test addresses the intended behavior. Other specialists contribute constraints and decisions within their authority.
The Loft loop is Author → Ratify → Implement → Reconcile → Decide. People designated by the organization ratify the contract; an authorized person accepts, rejects, or reopens the implemented result. Loft records that path without establishing product truth or making either decision.
The earlier article When Every Role Gets Faster, Why Doesn't Software Delivery? examines the handoff problem across an organization. The focus here is narrower: the evidence and authority needed for one acceptance decision.
Use five checks before accepting AI-generated software
For one candidate revision, a reviewer should be able to answer five questions:
- Which ratified contract version governed the implementation?
- Which requirements and acceptance criteria were in scope?
- What reviewed evidence maps to each criterion?
- Which criteria failed, lack evidence, or remain unresolved?
- Who has authority to accept, reject, or reopen the result?
If any answer is unclear, the review record is incomplete. The right response may be another test, a product decision, a contract change, a measurement after release, or another implementation candidate. These are different kinds of work, and the record should keep them distinct.
Start with one consequential change
Choose one consequential, reversible change with behavior that a named person can ratify and a result that an authorized person can accept or reject. Record the governing requirements, constraints, and acceptance criteria. Give the coding agent that ratified contract version and a bounded scope. Review the candidate evidence criterion by criterion, preserving failures, gaps, and unresolved claims.
See the complete fictional example, including both candidate revisions and the recorded decisions. To apply the same method to a product of your own, take one product area through the Loft loop.