On September 17, Anthropic published an internal assessment: as of August, Claude leads 26% of its AI research and development work under human supervision. Work at collaboration level or above exceeds 90%; no measured subcategory is fully autonomous. The index weights work using approximate person-time, and Claude also produces the ratings.
The announcement raises an interesting question: how does model development change when the tool participates in the research itself? Readers should distinguish the amount of delegated work, the authority to make decisions and responsibility for the outcome.
Why one percentage cannot describe a researcher’s whole job
The Epoch AI approach cited by Anthropic breaks research into specific tasks. Its authors propose more than 60 tasks across six categories and a 0–5 automation scale. Their argument is that success on an isolated benchmark misses the complexity of actual work, such as coordinating several projects without clear success criteria. They present the taxonomy as something to combine with measurements, rather than a finished forecast.
That suggests a practical question for any automation claim: what exactly is included in the measured process? Writing a piece of code, choosing a research hypothesis and deciding whether the evidence justifies releasing a product involve different decisions. An aggregate figure cannot explain how responsibility is distributed among them.

Does this mean AI already builds itself
In its explanation of recursive self-improvement, Anthropic describes a stronger scenario: a model develops its successor fully autonomously. The company explicitly says this has not been achieved and is not inevitable. It also distinguishes running a well-defined experiment from research judgment: choosing a direction and understanding which results really matter.
Turning a measure of tool involvement into a percentage of “people replaced” therefore answers a different question with the wrong evidence. Nor can a single snapshot establish a date when a system will independently manage an entire development cycle.
Who will check this work
The following day, Anthropic announced a partnership with Accenture for evaluators embedded in the development process. Anthropic will directly fund Accenture’s work. The company acknowledges that access, reporting and funding rules still need to be established. Announcing the partnership is therefore not a completed independent audit of these figures.
How to read the next “AI building AI” claim
- Look for the task description: what did the system handle, and what remained with a person?
- Distinguish executing a task independently from deciding its objective independently.
- Check who measured the outcome and what evidence an external evaluator can inspect.
- Look at errors and review time, as well as the number of completed actions.
When choosing an assistant for your own work, start with a concrete use case. Our ChatGPT, Gemini and Claude comparison can help frame the criteria; a laboratory’s internal figures do not, by themselves, establish which service suits you best.

Join the conversation
Stay on topic and respect other readers. Your first comment may appear after editorial review.