In OpenAI’s account of harness engineering, Codex wrote every line of a new product: the application, tests, CI, documentation, observability and internal tools. The most interesting result is not the stated volume of code, but the infrastructure built around the agent. People, the repository and controls were redesigned so that intent became verifiable and errors produced reusable feedback.

What the OpenAI case does—and does not—show

OpenAI reports that, starting from an empty repository, a small team guided Codex over five months to a code base of around one million lines and approximately 1,500 pull requests. The product had daily internal users and external alpha testers; the company estimates that it took roughly one tenth of the time required for manual development.

These are self-reported figures for a single product, not a controlled comparison. Lines of code alone do not measure value or quality, and the team had access to uncommon expertise, models and infrastructure. OpenAI also notes that behaviour depends heavily on that repository and that it is not yet known how coherence will evolve over a period of years. The case should be read as an engineering experience, not a promise of performance.

Make the repository readable to the agent

To an agent, a decision left in a chat or in someone’s memory is effectively invisible. The team therefore brought product specifications, architectural principles, execution plans, decisions and technical debt into the repository, all versioned and linked to the code.

The main instruction file remained a concise map rather than a monolithic manual: it points to deeper sources when needed. This progressive disclosure limits noise and makes attributes such as document owner, status and freshness verifiable. It also helps people who are new to the project, not just the model.

Translate the architecture into mechanical constraints

Documentation describes the direction; linting and structural tests prevent the costliest deviations. In the case described by OpenAI, domains follow permitted layers and dependencies, while automated checks verify structured logging, conventions, file size and reliability requirements.

These constraints do not prescribe every implementation detail. They define invariants—for example, validating data at boundaries or not reversing a dependency—and leave the agent room in how it satisfies them. The team thereby moves recurring judgement from individual reviews into executable rules. Human responsibility remains with priorities, acceptance criteria, constraint configuration and outcome validation, even when operational review is delegated to other agents.

Close the loop with feedback and recovery

An effective harness allows the agent to reproduce the problem, change the code, run tests and checks, observe the application and prepare a pull request. Build failures, review comments and user-reported bugs should feed back into the system as tests, documentation or new guardrails instead of remaining isolated fixes.

An exit route is also essential: reproducible environments, logs, small changes and rollback reduce the cost of a wrong trajectory. OWASP nevertheless recommends a human owner for every generated change and explicit approval before merging. An agent evaluating its own work provides another signal, but does not constitute independent verification in critical areas.

Optimise attention, not the number of prompts

When code arrives faster, the bottleneck becomes the human capacity to specify priorities, assess results and keep the system coherent. Start with one workflow, make setup and tests reliable, codify recurring defects and only then run more agents in parallel. Otherwise, patches and reviews multiply without increasing the value delivered.

DORA describes AI as an amplifier of the existing organisational system. Harness engineering makes this idea concrete: a clear repository, ownership, constraints and recovery turn model speed into team capability. The prompt remains important, but the advantage that compounds lies in the method that outlives any single conversation.

OFFICIAL SOURCES

Further reading.