The report measures how much of the research is already AI-led, display coverage, and compute usage. See how to adapt indicators to business governance.

Direct answer

On September 17, 2026, Anthropic proposed three groups of metrics to provide visibility into frontier AI development: AI participation in research itself, oversight of agent actions, and computation allocation. In the August portrait, the company states that Claude led 26% of measured AI R&D tasks, without operating fully autonomously in any subset. For companies, the most applicable model is to measure monitoring coverage, review time and escalation rate, always with explicit methodology and limitations.

Automation needs an operational scale

Separating assistance, collaboration, leadership and autonomy avoids treating any use of AI as a full replacement. The classification must consider the complete task, including exceptions and final approval.

Monitoring coverage is the first denominator

Before discussing incidents, the company needs to know which portion of actions undergo online or offline controls. Unobserved gaps make any safety rating misleading.

Latency defines whether supervision arrives on time

Irreversible actions require blocking or review before execution. Reversible activities can be analyzed later, as long as there is a deadline, priority and person responsible for scheduling.

Scaling rate needs context

Few alerts may indicate safe behavior or a weak monitor; many may reveal a risk or excess of false positives. Reviewed samples and independent testing complete the interpretation.

Nexus Reading

The most useful governance transforms principles into observable indicators by agent, tool and process. Persistent identity, action path and interruption criteria must exist before expanding autonomy.

FAQ

Can Anthropic's index be directly compared to another company?

Not yet safely. The source itself points out differences in methodology, classification and validation between organizations.

What metrics can a company adopt now?

Monitoring coverage, review time, blocking or escalation rate, and autonomy level per task.

Does a low block rate prove the agent is safe?

No. You need to know coverage, monitor quality, false negatives and the nature of the actions evaluated.

Essential guides to delve deeper into the decision

Primary sources and references

This editorial analysis was produced by Nexus from the official sources below, consulted on September 18, 2026. The text is original and interprets practical implications for companies.

Date reported by the main source: September 17, 2026.