A passing test suite answers a precise question: does the code satisfy the checks we wrote? It does not automatically prove that the request was complete, that the solution respects the architecture or that a maintainer would be willing to support it over time. With coding agents, this gap matters more because producing many patches has become cheaper while the attention required to assess them remains limited.

What a passing test actually proves

A passing test provides useful evidence against a specific regression. Its strength, however, depends on coverage, the chosen inputs and the independence of the oracle that determines the correct result. A suite may not exercise an authorisation branch, a race condition or the interaction with a real version of an external service.

If the same agent changes both the implementation and the tests, there is also a risk that the two will agree on the wrong behaviour. New, deleted, skipped or weakened tests should therefore be reviewed as part of the patch, not treated as neutral evidence of its quality.

When the maintainer supplies the judgement that is missing

In 2026, METR asked four active maintainers of scikit-learn, Sphinx and pytest to assess 296 agent-generated pull requests that had passed the automated SWE-bench Verified grader. The report estimates that roughly half of these test-passing patches would not have been merged into the main branch, even after adjusting for variability in maintainer decisions.

The result should not be treated as a universal measure. The study covers three repositories, a single harness and models available up to 2025; the review was static, with no CI and no opportunity for the agent to revise its patch after receiving feedback. The authors therefore do not present this as a fundamental limit, but as a warning against equating benchmark scores with usefulness without oversight.

The benchmark can measure the wrong target too

OpenAI stopped reporting SWE-bench Verified scores in frontier model launches and considers the benchmark no longer suitable for reliably measuring their progress, after documenting problems with the tasks and a risk of contamination: descriptions, patches and public repositories may already have appeared in data available to the models. In some cases, the tests were too narrow or covered requirements that were not stated in the problem.

A subsequent audit of SWE-bench Pro estimated that roughly 30% of the tasks examined had substantial problems, with different rates for the assisted pipeline and human annotators. This too is an estimate for a specific dataset, not a verdict on every benchmark. The operational lesson is simpler: before interpreting a score, assess task quality, test coverage, the environment and the evaluation criteria.

The questions that do not fit into an assertion

The maintainer checks whether the patch fulfils the intent, limits its scope and uses the existing abstractions. They assess compatibility, naming, error handling, documentation, operational cost and ease of removal. They may reject working code because it duplicates a function, introduces a disproportionate dependency or weakens an architectural boundary.

Security and data require even more rigorous scrutiny. A local fix may bypass an authorisation check, expose information in logs or turn a harmless migration on test data into a prolonged lock in production. These risks emerge by relating the diff to the system’s history and environment, not merely by running the suite.

Design an acceptance pipeline, not a patch factory

Give agents well-scoped tasks with explicit acceptance criteria and sensitive files. Require small diffs, an explanation of decisions, tests linked to requirements and a list of the checks performed. CI should flag changes to workflows, coverage thresholds, dependencies, migrations and assertions; critical areas need a competent, independent approver.

Measure time to an accepted patch, rework cycles, escaped defects and review load, not merely the number of pull requests or passing tests. A coding agent creates value when it reduces the total effort required to reach a maintainable change; if it shifts the cost to the reviewer or the next incident, the apparent speed is illusory.

OFFICIAL SOURCES

Further reading.