Data published by OpenAI shows more code, experiments and long tasks with agents. See metrics, limits and lessons for enterprise R&D.
Direct answer
OpenAI reported that programming agents already account for 3.1 agent journeys for every human journey in its research organization and that August 2026 recorded the highest number of experiments per active researcher since measurement began. The gain, however, does not eliminate supervision: more than half of successful four- to eight-hour tasks still required human intervention.
What the data shows
OpenAI describes a rapid shift in research work: agents are used daily, including in concurrent sessions, to write code, operate infrastructure, run evaluations, and analyze experiments. The company associates adoption with more code contributions and a greater volume of experiments, without claiming isolated causality.
Productivity is not total autonomy
The figure of 3.1 agent journeys per human journey measures equivalent computational time, not replacement of researchers. People continue to set priorities, judge evidence, and decide whether an experiment should move forward, be paused, or discarded.
The intervention remains relevant
In successful tasks estimated at four to eight hours of human labor, more than half required at least one intervention. For companies, this indicates that completion rate should be read along with rework, corrections, review time and impact of errors.
How to apply in business R&D
Start with delimited and verifiable tasks: data preparation, exploratory analysis, reproducing results, and test automation. Record hypothesis, version, tools, evidence and human decision. Expansion should only occur when quality and traceability remain stable.
Nexus Reading
Agents expand the capacity to experiment, but they can also accelerate fragile conclusions. The competitive advantage comes from combining parallelism with clear stopping criteria, evaluations and responsible parties.
FAQ
Do agents already replace researchers?
No. The publication itself highlights direction, judgment and human decisions throughout the process.
Which metric to use in a pilot?
Time to validated result, success rate, interventions, rework and total cost per useful conclusion.
Do more experiments mean more innovation?
Not necessarily. Volume needs to be accompanied by quality, reproducibility and impact of discoveries.
Essential guides to delve deeper into the decision
This editorial analysis was produced by Nexus from the official sources below, consulted on September 9, 2026. The text is original and interprets practical implications for companies.
- OpenAI — Research acceleration: The view inside OpenAI: data, methodology, limitations and role of human supervision.
- OpenAI — Research publications: official record of the publication and its date.
Date reported by the main source: September 6, 2026.

