For decades, one of the stranger ideas surrounding artificial intelligence has been recursive self-improvement: build a machine capable enough to help improve the process that builds machines, then allow each generation to contribute to the development of the next. Science fiction usually turns that idea into a clean dramatic event. The machine becomes sufficiently intelligent, rewrites itself, produces something smarter, and somewhere shortly afterward humanity begins discovering that perhaps connecting everything to everything was not our finest decision. Actual technological change rarely works that neatly. It usually arrives through a collection of ordinary improvements, boring interfaces and productivity tools until somebody finally notices that the underlying system no longer resembles the one that existed a few years earlier.
That appears to be what is beginning to happen inside frontier artificial-intelligence research. Anthropic says Claude now leads roughly 26 percent of the AI research and development work the company has measured internally, while more than 90 percent of that work involves Claude at a level the company describes as collaboration or greater. On Anthropic's most heavily used internal agent platform, approximately 30,000 AI agents are performing research and engineering work at any given time. None of this means Claude has been handed the laboratory keys and told to build its own successor. Anthropic says humans remain responsible for the research process, and none of the categories it measured qualify as completely autonomous AI research. The interesting part is how quickly the balance appears to be changing. In February, Anthropic measured the portion of work Claude was leading at below one percent. By August, it was 26 percent.
That number is more significant than another benchmark showing that a new model can code slightly better, answer harder questions or beat another model on a test most people will forget the name of by next Tuesday. What Anthropic is describing is a change in the machinery used to develop artificial intelligence itself. Claude is no longer simply the product being developed. It is increasingly becoming one of the tools used to develop the next product. Better AI can help researchers build better AI, which can then return to the same research environment with better programming, reasoning, debugging and analytical abilities. The feedback loop people have discussed theoretically for years does not require an autonomous machine secretly redesigning itself in a server room. A supervised version can emerge much more quietly, with human researchers directing the process while AI performs an increasing share of the work between the original question and the final answer.
Frontier AI development involves far more than launching an enormous training run and waiting for a new model to emerge. Researchers spend enormous amounts of time writing code, designing evaluations, inspecting failed experiments, comparing results, searching technical literature, debugging infrastructure, studying model behavior, building tests, analyzing logs and chasing ideas that look promising until someone discovers they were produced by a broken evaluation or a misplaced decimal point. The glamorous breakthroughs sit on top of a mountain of ordinary technical work. That mountain is exactly where AI agents can become useful long before they are capable of replacing the scientist.
A researcher can ask one agent to review relevant papers while another modifies experimental code. Another can inspect the results, another can search for a regression and another can generate alternative implementations. One system might spend twenty minutes on a task while another remains active for hours. They do not need to resemble human scientists to increase the amount of research an organization can perform. They only need to remove enough repetitive or time-consuming work from the researchers who would otherwise have to do it themselves.
This is the reason the figure of roughly 30,000 agents deserves more attention than it will probably receive outside technical circles. Thirty thousand AI agents are not equivalent to thirty thousand human employees, and pretending they are would badly distort what these systems currently do. They do not possess thirty thousand independent careers, educations or lifetimes of experience. What they provide is parallelism. Work that once had to move through a relatively narrow human pipeline can be divided into smaller tasks and attacked simultaneously.
That changes the economics of experimentation. Research is partly a problem of searching through possibilities, and most possibilities are useless. Most experiments fail. Most clever ideas eventually reveal themselves to be less clever than advertised. Progress often comes from trying enough approaches to find the small number that survive contact with reality. If two laboratories contain researchers of roughly equal ability but one can investigate five times as many serious approaches during the same period, it gains an advantage without necessarily having better ideas at the beginning. It simply gets more opportunities to discover something useful.
Computing has been doing this to science for decades. Simulation allowed engineers to test aircraft designs before building physical prototypes. Automated software testing allowed programmers to check thousands of cases overnight. High-performance computing allowed scientists to explore problems that would have been impossible to calculate manually. AI agents extend that process into areas of work that were previously harder to automate. They can write portions of the experiment, interpret preliminary results, search documentation, compare approaches and prepare the next round of work. The researcher remains responsible for determining what matters, but the machinery between deciding on a question and receiving useful evidence becomes faster.
That is where the discussion of recursive AI development becomes considerably more grounded. Suppose Claude helps Anthropic researchers develop a technique that improves a future version of Claude. The next model becomes somewhat better at writing experimental software, investigating failures and analyzing results. Anthropic then deploys that improved model inside the same research environment. It can now contribute somewhat more effectively to the work necessary to produce another generation. Humans can remain involved at every stage and the process is still recursive in an important sense: the technology is participating in the process used to improve the technology.
There is nothing particularly mystical about that. Computers already help engineers design faster computers. Semiconductor equipment is developed using chips created by earlier generations of semiconductor equipment. Modern aircraft are designed using simulation and computational tools that exist because earlier generations of aerospace engineers created the machines that made those tools possible. Technology has always helped build better technology. Artificial intelligence makes the relationship unusual because the thing being improved is increasingly able to perform some of the intellectual work involved in its own improvement. The hammer is beginning to study metallurgy.
That does not automatically mean intelligence suddenly accelerates beyond human control, and there is little value in pretending otherwise. The more immediate consequence is that the speed limit imposed by human labor begins to move. If AI can handle increasing amounts of implementation, debugging, analysis and information retrieval, researchers can spend more of their time deciding what problems should be investigated. The bottleneck shifts from performing every technical step to choosing which steps are worth performing.
That shift could make human judgment more valuable rather than less. A system capable of running ten thousand experiments is extremely useful when someone knows which experiments should be run. It is merely an efficient way to generate ten thousand useless results when nobody does. If software becomes cheap to generate, architecture matters more. If technical literature becomes cheap to search, recognizing which contradictory result deserves investigation matters more. If an AI can create twenty implementations of an idea, understanding which implementation is valid becomes the scarce skill.
The human researcher gradually moves upward through the process. Instead of manually performing every operation, researchers define objectives, allocate resources, interpret results and decide where the next round of work should go. There is something mildly absurd about spending decades trying to build machines capable of human intellectual labor only to discover that one of the first consequences may be turning scientists into managers, but that is increasingly what large agent systems resemble. A human provides an objective, software breaks the objective into pieces, agents execute those pieces, results return, problems are escalated and the human decides what happens next.
Once agents begin delegating work to other agents and communicating with one another, the structure becomes even more interesting. It stops looking like a person using a sophisticated software tool and starts looking like a primitive organization. Some systems perform research. Others write code. Others evaluate the results. Monitoring systems watch the agents performing those tasks. Humans remain at the top of the process, but the amount of machine activity underneath them can grow enormously without a comparable increase in the number of employees.
That possibility makes headcount a less useful measure of research capacity. A laboratory with several hundred exceptional researchers and a sufficiently capable population of agents could eventually perform work that once required a much larger human organization. The size of the biological workforce would no longer tell us how much intellectual labor the institution can produce. Compute, model capability and agent infrastructure would become part of the effective workforce.
That creates another problem immediately: somebody has to supervise all of it. Tens of thousands of agents can generate an enormous number of actions, far too many for humans to inspect individually. Anthropic describes automated monitoring systems that examine agent behavior before and after actions. The company says 100 percent of actions on the measured platform pass through an online monitor before execution and are later ingested by an offline monitor. In August, it analyzed more than a billion decisions, with about 0.002 percent blocked online. The offline system reportedly flags roughly 100,000 transcripts each week for additional classifier review and escalates about fifty of the highest-priority cases to people.
An undesirable event that occurs once in tens of thousands of actions sounds extremely rare when a person performs a few dozen meaningful actions in a day. When software performs hundreds of millions or billions of actions, statistically rare events become operationally routine. Cybersecurity, aviation and high-volume financial systems have dealt with versions of this problem for years. Reliability is never just a percentage. The number of opportunities for failure matters too.
Large-scale AI infrastructure therefore requires a second automated system whose job is watching the first one. Agents perform tasks while automated monitors evaluate their actions. Other systems inspect logs and search for suspicious patterns. Some actions are blocked automatically. Other cases are escalated for human review. The oversight machinery grows alongside the productive machinery because human beings cannot manually inspect everything being generated.
This introduces one of the most difficult problems in the entire model: AI can produce work faster than people can verify it. Anyone who has used a capable coding agent on a sufficiently large project has already encountered a small version. Asking a model to write a short function produces something a programmer can inspect quickly. Asking the same system to alter dozens of files, refactor interfaces, migrate data structures and repair tests can produce changes in minutes that take considerably longer to understand. Generation starts outrunning verification.
Now move that same problem into scientific research. An AI system could eventually read hundreds of papers, modify several experimental systems, run thousands of tests and return an analysis before a human researcher could independently reconstruct everything it had done. That is an enormous productivity improvement as long as the answer is correct. When something goes wrong, locating the mistake becomes much harder.
The obvious solution is to automate verification as well. Models review model output. Agents inspect work produced by other agents. Tests validate generated software. Evaluators compare results against known behavior. Monitoring systems look for actions outside the expected range. All of those techniques are useful, but they introduce the possibility of correlated mistakes. Two independent human researchers may approach a problem differently enough that one notices something the other missed. Two systems built from closely related models may share exactly the same blind spot.
Software engineers already know how treacherous that can be. A test built around the same mistaken assumption as the program it evaluates can confidently announce that everything works. Nothing works. The test and the software simply agree about the wrong thing. AI can reproduce the same failure at the level of reasoning, where several systems appear to confirm one another because they inherited the same underlying assumptions.
Anthropic's own automation measurement contains a version of this problem. The company used Claude agents to map roughly 15,000 granular tasks into a hierarchy and another Claude judge to assign automation levels. Staff ratings were used as a check, and Anthropic reports that model and human judgments were usually within one level, but the result remains a self-measurement built partly with the technology being measured. That does not make the result meaningless. It makes independent replication and common definitions necessary before anyone compares laboratories as if 26 percent were a universally calibrated number.
The economics surrounding this are equally important. In a conventional research organization, headcount and budgets provide a rough indication of where effort is being spent. In an AI-heavy laboratory, compute increasingly becomes another form of labor allocation. Anthropic reports that during one sampled week in July, about six percent of AI R&D compute went to safety work and about twelve percent of compute used specifically for AI-driven AI R&D went to safety. The company calls those estimates conservative and acknowledges that the classification boundary is imperfect.
More processors assigned to research agents can mean more experiments, implementations and analysis. More compute assigned to evaluation means more machine-generated work can be checked. Resources serving customers cannot simultaneously run every research experiment, and resources used to develop new capabilities cannot simultaneously be devoted to monitoring them. The allocation of machine intelligence becomes an organizational decision.
Competition complicates everything. Suppose one frontier laboratory concludes that AI-assisted AI development is becoming dangerous and deliberately limits how much independent work its agents can perform. Another laboratory makes a more aggressive decision. The cautious organization may reduce certain risks, but it may also slow the rate at which its researchers improve their models. If the aggressive competitor develops a stronger model first, that model can then be returned to its own research process and potentially increase the difference further.
Nobody involved has to be reckless for this dynamic to exist. It emerges from ordinary competitive pressure. Every organization can privately prefer caution while collectively accelerating because no organization wants to be the one that slows down alone. That is one reason measuring AI-assisted AI research matters. Society needs more information than another leaderboard showing which company has the smartest model this month. We need some understanding of the machinery producing the next generation.
How many agents are operating inside frontier laboratories? What percentage of research work do they perform? What tools can they access? How independently can they operate? What actions require human authorization? How frequently do monitoring systems intervene? How much compute is being devoted to AI-driven research? Most importantly, how quickly are all of those numbers changing? Those measurements describe something different from model capability. Benchmarks describe the machine. These measurements describe the factory.
The distinction matters because the architecture being developed for AI research will not remain confined to AI research. A system capable of reading technical literature, writing software, analyzing data, coordinating complicated tasks and interpreting experimental results can be directed toward biology, chemistry, materials science, robotics, aerospace, electronics and energy research. The methods used to accelerate artificial-intelligence development can become general methods for accelerating scientific development.
Anthropic has already expanded Claude's role in scientific work, including biomolecular modeling, and Reuters reports that the company has established a physical biology laboratory in the San Francisco Bay Area as part of its life-sciences effort. Connecting AI to physical experimentation is a different threshold from letting it manipulate code or information. A bad simulation can be rerun. A physical experiment involves equipment, material, expense and safety. But the connection is clear: the same agent architecture being developed to accelerate AI research is also an architecture for accelerating research itself.
This is why the public-facing chatbot may eventually turn out to be one of the least consequential forms of artificial intelligence. Chat interfaces are what most people see, so they dominate the discussion. Someone types a question and receives an answer. Sometimes the machine is extraordinarily useful. Sometimes it invents something with such confidence that you almost admire the commitment. Either way, the interaction encourages people to think of AI as another software product.
Inside research organizations, AI is becoming infrastructure. Infrastructure does not need to be spectacular. It becomes important by disappearing into everything. Electric motors changed factories not because one giant motor sat in the middle of the building doing all the work, but because smaller motors spread into pumps, conveyors, machine tools, ventilation systems and hundreds of other processes. Each motor solved a specific problem. Together they changed the architecture of manufacturing.
AI agents may be doing something similar to intellectual work. One coding assistant is a tool. One research agent is a tool. Tens of thousands of interconnected agents operating continuously inside an organization begin to constitute a system. The transition from tool to infrastructure happens gradually enough that we may not recognize it until the new arrangement has already become ordinary.
There probably will not be a clean moment when the first automated research institution officially arrives. Nobody will cut a ribbon outside an intelligence factory. The percentage of machine-led work will simply keep changing. Agent populations will grow. Models will improve. Researchers will spend less time performing individual technical operations and more time directing systems that perform them. Eventually the structure of the laboratory will be sufficiently different from what came before that calling AI an “assistant” will sound strangely inadequate.
That is why the most important number in Anthropic's report is not 26 percent. It is whatever number comes after it. If the percentage stabilizes around its present level, this period may eventually look like the point where artificial intelligence became a powerful but bounded engineering tool. If it continues climbing, we may be watching the early construction of increasingly automated research institutions. If capability and agent populations rise together, the effective research capacity of frontier laboratories could grow much faster than their human workforce.
For most of technological history, machines expanded what people could physically accomplish. Engines multiplied muscle. Electricity made energy easier to move and control. Computers multiplied calculation. Networks accelerated the movement of information. Artificial intelligence is beginning to amplify something different: the process by which new technology is developed. Nobody has turned the laboratory over to Claude, and nobody needs to. The quieter, more believable transition is already underway. The machines have jobs there now.
Anthropic's numbers do not prove autonomous recursive self-improvement. They establish something more immediate and measurable: AI has become part of the production system for future AI. The governing metric is no longer only model capability. It is capability multiplied by agent count, tool access, research responsibility and the organization's ability to verify what the resulting factory produces.

