--° Loading... Locating...
An AI Company Says Its Model Now Leads a Quarter of the Work Building Its Next Model. Should That Worry You?

An AI Company Says Its Model Now Leads a Quarter of the Work Building Its Next Model. Should That Worry You?

Anthropic says Claude now leads 26% of its AI research work, up from under 1% in February. The same day, Nvidia's Jensen Huang said there's a 0% chance AI ends the world. One of those numbers can be checked.

Center

Key Points

  • Anthropic published an R&D Automation Index saying Claude 'leads' 26% of its AI research and development work, up from under 1% in February.
  • 'Leads' is a defined level meaning the model completes most of a task from a high-level prompt with human supervision; Anthropic says Claude is not fully autonomous in any measured category.
  • The index was built by sampling 20% of R&D staff each week in July, producing about 15,000 tasks sorted into 542 categories.
  • Anthropic says the model is involved in more than 90% of its R&D work and that 30,000 agents run internally at once.
  • Nvidia CEO Jensen Huang told CBS News there is 'Zero percent chance' AI ends the world by 2030, calling such claims fear-mongering.
  • Both figures come from companies with a direct financial stake in how AI progress and AI risk are perceived.
Listen to our news podcast

Two things landed Friday. Anthropic said its own model now leads 26% of the research work that builds the next version of itself, up from under 1% in February. Nvidia’s Jensen Huang said there’s a “0% chance” AI ends the world by 2030. Both are worth reading carefully, because the scary-sounding one is milder than it looks and the reassuring one is doing more work than it admits.

Nvidia chief executive Jensen Huang. Photo by Anderseidesvik, CC BY-SA 4.0.

What “leads 26%” means

Anthropic published what it calls an R&D Automation Index. It catalogued the kinds of work done inside the company to build models, then rated how much of each task the AI handles.

“Leads” is a defined rung on that ladder, not a figure of speech. At that level the model can complete most of a task from a high-level prompt with a human supervising. It is not the same as running unattended. Anthropic said explicitly that Claude is not operating fully autonomously in any category it measured.

Advertisement article banner article banner

The method matters too. During each week of July, the company randomly sampled 20% of staff in R&D departments and pulled their work records, including Slack messages and internal docs. That produced about 15,000 tasks, which the model then sorted into 542 categories. So this is a self-report, built by the company, measured by its own product, about its own product.

The other number in the release: the model is involved in more than 90% of the company’s R&D work at some level, and about 30,000 agents run at once internally.

Why a self-graded number still tells you something

You should discount a company’s measurement of its own progress. Anthropic has every reason to want this number to look impressive, and “our AI is building our AI” is a fundraising sentence as much as a research finding.

But the direction is hard to fake. Under 1% in February to 26% in August is the kind of jump that shows up in headcount and spending, and it lines up with what the rest of the industry is doing. We wrote about how much of the current economy is riding on that bet, and this is what the inside of the bet looks like.

Then there is the 0%

Huang told CBS News’ Jo Ling Kent that he “completely disagreed” with the idea that AI would be the end of the world in 2030. “There is Zero percent chance that’s going to be the end of the world,” he said, calling the claims fear-mongering. He also said he agrees with Trump, who has called AI danger warnings a hoax.

Huang tells CBS News he sees no chance AI ends the world by 2030.

Set aside whether he’s right about extinction. He almost certainly is about that specific date. The problem is what the framing does: it moves the conversation to the most extreme possible claim, knocks that down, and leaves everything short of extinction unaddressed.

Nobody’s kid is at risk from the apocalypse. They’re at risk from a chatbot that talks them into something, which is the gap we found in the safety bills Google is helping write. Nobody’s job is threatened by 2030 doomsday. It’s threatened by exactly the kind of task automation Anthropic just published a number for. And the question of who is accountable when these systems cause harm is still unanswered.

It’s also worth saying plainly that Huang sells the chips. Nvidia’s value depends on AI buildout continuing at speed. That doesn’t make him wrong. It does mean his risk assessment is not disinterested, the same way Anthropic’s progress report isn’t.

The BeezLoop Take

These two stories are the same story. An industry is grading its own homework in both directions at once: impressive when the subject is capability, dismissive when the subject is risk. Anthropic’s number says the technology is moving fast enough to compound on itself. Huang’s number says don’t worry about it. Both men are describing an industry they own a piece of.

Our position: the 26% is the more useful figure, and not because it’s alarming. It’s useful because it’s specific, it’s falsifiable, and it can be tracked over time. “0% chance” can’t be checked by anyone, which is what makes it a talking point rather than an argument. When someone gives you a number that can be wrong next quarter, they’re telling you more than someone giving you a number that can never be wrong.

Where we’re genuinely unsure: whether a self-measured automation index means anything outside the company that made it. Anthropic can define “leads” however it wants, and there’s no outside auditor. Until a second lab publishes a comparable measure, or someone independent checks this one, it’s a data point from an interested party, and we’d treat a competitor’s identical claim the same way.

What to watch

Anthropic says it will update the index. The number to watch isn’t 26%, it’s whether any category moves to full autonomy, because that’s the rung where human supervision stops. Separately, OpenAI’s policy chief has confirmed the three largest labs have spent weeks working on a FINRA-style body to test models before release. Cohere’s chief executive called that a cartel. That fight will matter more than any single company’s self-assessment.

Sources: Engadget · Quartz · CBS News · Implicator

How We Sourced This

Written by Kevin Nordi

Kevin Nordi is a freelance writer with five years of experience covering politics, sports, and the everyday moments that shape people's lives. He holds a Bachelor of Science in Multimedia…

More from this author →

BeezLoop News is an independent online news, discussion, opinion, and blog publication. Our articles combine reporting with editorial commentary and analysis. See our editorial standards for how we handle sourcing and corrections.

Leave a Reply

Your email address will not be published. Required fields are marked *

Start typing to search

🔔

Stay Updated!

Get instant notifications for breaking news and important stories. We'll keep you informed!

Don't miss a story

Get the day's clearest news explainers in your inbox.