What Happens When the World’s Best Researchers Are No Longer Working Alone?

Anthropic recently gave an unreleased research version of Claude an almost unreasonable challenge:

Take a real stab at the Riemann Hypothesis.

It did not solve it.

But that may not be the most interesting part of the story.

While working on the problem, Claude found a way to make progress on a related question in analytic number theory, improving a longstanding lower bound for the proportion of zeros of the Riemann zeta function known to lie on the critical line — from 41.6% to 67.2%.

Anthropic says the result was subsequently examined by mathematicians, compared against previous research and translated into a formally verifiable proof.

But what caught my attention was not simply the mathematical result.

It was how the work was done.

From answering questions to exploring problems

For several years, most of us have interacted with AI through a relatively simple pattern:

Question → Answer

We ask for an analysis.

A summary.

Some code.

A recommendation.

A solution to a problem.

But the Claude experiment looks very different.

Initially, the model explored hundreds of possible ideas. They failed.

It then spent more than a day coordinating around 60 AI subagents, which investigated different approaches in parallel.

Together they ran thousands of computational checks, wrote hundreds of Python scripts, tested hypotheses against known mathematical results, challenged one another’s reasoning and searched for counterexamples.

According to Anthropic, the process generated 31 million output tokens and involved approximately 2,400 shell commands.

Once a promising result emerged, Claude continued testing it.

Agents attempted to independently reproduce the finding.

Others tried to disprove it.

Claude searched dozens of existing research papers to determine whether the result had already appeared in the mathematical literature.

Eventually, Claude itself recommended that human mathematicians examine the result.

That is a very different model of Artificial Intelligence.

It is no longer simply:

“Give me the answer.”

It becomes:

Objective → Exploration → Experimentation → Critique → Verification → Human Review

And that changes something fundamental.

Failure may no longer mean failure

Claude failed at the original objective.

The Riemann Hypothesis remains unsolved.

Yet something potentially valuable emerged during the attempt.

This is actually very familiar to science.

Researchers frequently begin by trying to answer one question and discover something important somewhere else.

They follow a hypothesis.

It fails.

They investigate why.

Another possibility appears.

A side observation becomes more interesting than the original question.

Scientific progress has always contained this element of exploration and unexpected discovery.

What AI changes is the scale at which that exploration may become possible.

A human researcher has limited time.

There are only so many papers one person can read.

Only so many calculations one person can perform.

Only so many alternative hypotheses one research team can investigate simultaneously.

Now imagine changing that constraint.

Imagine a researcher being able to say:

Explore these approaches.

Test the assumptions.

Search the literature.

Run the calculations.

Try to disprove the strongest ideas.

Let independent agents examine the most promising results.

Then bring me back the few possibilities that survived.

The scientist is not removed from the process.

The scientist becomes capable of exploring dramatically more of the problem space.

The research organisation around one human

This may ultimately be more important than the debate about whether AI can independently “do science.”

Consider another possibility.

What happens when one exceptional mathematician effectively gains an entire AI research organisation?

Or a physicist.

A chemist.

A medical researcher.

An engineer.

Dozens of specialised agents could explore parallel approaches, search different bodies of literature, run simulations, test assumptions, compare results and challenge one another continuously.

The human researcher would still contribute something extraordinarily important:

judgment.

intuition.

context.

domain expertise.

the ability to recognise why a result matters.

And ultimately, responsibility for deciding what deserves to be trusted.

But AI could radically expand the territory that researcher is capable of exploring.

That distinction matters.

The most interesting future of AI in science may not be machines replacing scientists.

It may be scientists gaining capabilities that previously required entire teams — or years of human effort.

Research no longer has to take a lifetime

For generations, some scientific problems have consumed significant parts of a researcher’s career.

Not necessarily because the ideas were impossible to understand, but because exploring the enormous number of possible paths simply takes time.

Reading.

Calculating.

Testing.

Rejecting.

Starting again.

Searching for related work.

Checking whether somebody else has already explored the same direction.

Reproducing results.

Arguing with colleagues.

Running another experiment.

And another.

AI does not eliminate the need for those activities.

It may dramatically compress the time required to perform them.

That is why the most interesting part of the Claude experiment may not ultimately be the number 67.2%.

It may be the process that produced it.

Hundreds of failed ideas.

Dozens of parallel agents.

Millions of tokens of exploration.

Thousands of computational operations.

Repeated attempts to challenge the result.

Then human review.

We are beginning to see what happens when intellectual exploration is no longer constrained by the working hours of a single human being.

And that raises a much bigger question.

Are we ready for the speed of AI-augmented research?

If AI systems become increasingly capable of sustained exploration, scientific discovery may begin to operate at a very different tempo.

The timeline between:

question → exploration → insight → validation

could become dramatically shorter.

That will bring enormous opportunities.

It will also make human judgment, verification and governance more important — not less.

Because accelerating discovery also means accelerating the production of hypotheses, claims and possible answers.

The challenge will not simply be generating them.

It will be knowing which ones deserve to survive.

Perhaps, then, the future is not a choice between human intelligence and Artificial Intelligence.

It is a new research model in which human intelligence directs, evaluates and gives meaning to an unprecedented scale of machine-assisted exploration.

The Riemann Hypothesis was not solved.

But the experiment may have demonstrated something equally significant about where scientific research is heading.

AI may not only help us answer questions faster.

It may allow us to explore far more of the space in which new questions — and new discoveries — are hiding.

And perhaps, for the next generation of researchers:

Research no longer has to take a lifetime.

Next
Next

Why We Do Not Lack Knowledge — We Lack Measure