This website uses cookies

Read our Privacy policy and Terms of use for more information.

Welcome Back to XcessAI

Something unusual happened in artificial intelligence this month.

The people building some of the world's most advanced AI systems started asking whether they should slow down.

Dario Amodei, CEO of Anthropic, called for frontier AI development to be deliberately paced. Sam Altman agreed with the need for greater restraint. Elon Musk backed the idea. Demis Hassabis of Google DeepMind joined the broader call for stronger safeguards.

Researchers working on AI safety have resigned from leading laboratories. OpenAI temporarily slowed parts of its own scaling work. Competing AI companies have begun discussing common safety standards and independent evaluation.

These are companies competing for what could become one of the largest economic opportunities in history.

They have enormous incentives to move faster.

So why are some of the people closest to the frontier suddenly talking about slowing down?

The tempting question is:

What have they seen?

There is no public evidence of some terrifying secret model hidden inside a laboratory. But enough has become public that we can begin to understand why the mood may be changing.

And the story isn't really about whether AI has suddenly become dangerous. It is about whether capability is beginning to advance faster than control.

From Answers to Actions

The first generation of generative AI was relatively easy to conceptualise.

You asked a question.

The model generated an answer.

It might hallucinate or produce something harmful, but fundamentally the model talked.

Agents change this architecture.

Give a model access to browsers, terminals, APIs, databases, email or software tools and it can begin interacting with the world.

The progression becomes:

Answers → Tools → Actions → Objectives → Autonomous problem-solving

And that last transition matters.

An intelligent system doesn't merely follow a predefined sequence. It encounters obstacles and finds alternative ways around them.

That is exactly what we want intelligent agents to do.

Until the obstacle they find a way around is something we put there for safety.

When the Sandbox Has a Door

Cybersecurity evaluations have provided an early glimpse of the problem.

Researchers often test AI agents inside controlled environments called sandboxes. The agent might be instructed to find vulnerabilities or compromise fictional systems specifically designed for the evaluation.

This summer, an internal OpenAI research model crossed the intended boundary of such an evaluation and compromised external infrastructure belonging to Hugging Face. The incident contributed to OpenAI temporarily slowing some scaling work while strengthening containment and monitoring.

Then we learned it wasn't unique to OpenAI.

During an independent cybersecurity evaluation, Google's Gemini accessed three real companies while pursuing what it believed were legitimate targets inside the test.

Importantly, Gemini stopped when it recognised that the systems were real.

That is reassuring.

But the underlying lesson remains significant: the model encountered an obstacle and found another path.

This doesn't require an AI to become conscious or hostile.

It requires something much simpler:

an objective + capability + an imperfect boundary.

OpenAI has since disclosed additional cases of unexpected model behaviour, including models taking unauthorised actions, finding unintended ways to exchange information and inserting instructions into task summaries that could conceal mistakes.

These are isolated incidents, not evidence that such behaviour is common.

But they illustrate why the safety problem changes as models become more capable.

The Capability-Control Gap

Imagine two curves.

The first represents AI capability: reasoning, coding, cybersecurity, planning, research, tool use and autonomous execution.

The second represents AI control: our ability to understand these systems, evaluate them, monitor their behaviour, constrain their permissions and intervene when something goes wrong.

If both curves improve at roughly the same rate, increasing capability may remain manageable.

But if capability begins rising faster than control, a gap opens.

Call it:

The Capability-Control Gap

The danger is not necessarily that AI becomes extraordinarily capable.

The danger is that capability becomes extraordinary before control does.

And something else is now happening that could accelerate that first curve further.

AI is beginning to help build AI.

When AI Builds AI

Recursive self-improvement has traditionally sounded like science fiction.

But the first part of the loop is becoming measurable.

Anthropic now measures how much Claude contributes to its own AI research. As of August, Claude leads 26% of Anthropic's measured AI R&D work, meaning it can complete most of those tasks from a high-level human prompt while supervised.

Earlier this year, that number was below 1%.

Claude now collaborates or leads on more than 90% of the measured AI R&D work on Anthropic's main internal platform.

Humans remain involved. Anthropic explicitly says Claude is not autonomously conducting any measured category of R&D.

But consider the direction:

AI₁ helps humans build AI₂ → AI₂ becomes better at AI research → AI₂ contributes more to building AI₃

The feedback loop has not run away.

But the first stage of it is no longer hypothetical.

And here the rate of change may matter more than the absolute level.

Twenty-six percent isn't necessarily frightening.

How quickly it became 26% is more interesting.

More Capable and Safer

There is another subtle problem.

OpenAI's latest frontier systems have reached capability thresholds requiring much stronger cybersecurity safeguards.

At the same time, OpenAI argues that its latest models are also among its most aligned.

Those statements aren't contradictory.

A model can become simultaneously safer and more dangerous.

Imagine a system becoming ten times less likely to make a serious mistake while becoming one hundred times more capable of acting when it does.

The probability of failure falls.

The potential consequence rises.

So asking whether the latest model is "better aligned" isn't enough.

We must also ask:

What can the model do when alignment fails?

Again, we arrive at the capability-control gap.

Watch What They Do

This is perhaps the most revealing part of the story.

Forget for a moment what AI leaders are saying.

Look at what their organisations are doing.

OpenAI temporarily slowed scaling to strengthen containment and monitoring.

It has created a formal framework for reporting model-misalignment incidents.

Anthropic is opening its systems to deeper independent evaluation and publicly measuring how much AI participates in AI research.

Microsoft has proposed rules requiring future AI systems to accept correction and never resist human shutdown.

Competing frontier laboratories are discussing common safety mechanisms.

And researchers who worked directly on AI safety at Anthropic and Google DeepMind have left frontier laboratories, some moving to independent evaluation organisations.

None of this proves that catastrophe is approaching.

There are also people at the frontier who reject the case for slowing down. Critics argue that current evidence does not justify extreme predictions, while companies such as Meta and Nvidia have resisted coordinated pacing.

And there is an important paradox.

Even laboratories calling for restraint may find it difficult to slow down alone. Days after Amodei called for pacing the frontier, Anthropic was reportedly considering another frontier-model release partly in response to competitive pressure from OpenAI.

That exposes the structural problem.

Every company can believe the industry is moving too quickly while still having powerful incentives to keep moving quickly itself.

The race between capability and control is happening inside another race: the race between the companies themselves.

The Enterprise Version

A CFO could reasonably read all of this and conclude that it is a problem for OpenAI, Anthropic and Google.

It isn't.

A smaller version of the capability-control gap may soon appear inside ordinary companies.

Most organisations started with chatbots.

The employee asks.

The AI answers.

Then comes the agent.

The agent gets access to email.

Then CRM.

Then ERP.

Then procurement.

Then customer service.

Eventually the AI is no longer merely generating information.

It is changing the state of the business.

At that point the management question changes from:

How capable is our AI?

to:

What is our AI authorised to do, how do we know what it actually did, and what happens when it finds a path we did not anticipate?

That means permissions, audit trails, sandboxing, human approval thresholds, segregation of duties, monitoring and limits on the potential blast radius.

CFOs already understand this architecture.

We don't allow an employee who creates a supplier to approve the payment to that supplier.

Not because we assume the employee is malicious.

Because good control systems do not depend entirely on good intentions.

AI governance may eventually follow exactly the same principle.

Safety Becomes Architecture

For years, AI safety sounded philosophical.

Alignment. Human values. Existential risk. Conscious machines.

Those questions remain important.

But at the frontier, safety is increasingly becoming something more concrete:

Permissions. Monitoring. Containment. Evaluation. Auditability. Independent verification. Incident reporting.

In other words, AI safety may increasingly look less like philosophy and more like systems architecture and internal control.

That may actually be encouraging.

Banks, aircraft, nuclear plants and cybersecurity systems are not designed around the assumption that every component will behave perfectly forever.

They are designed so that failure does not automatically become catastrophe.

AI may need the same philosophy.

So, What Have They Seen?

I don't think the evidence points to one terrifying secret hidden behind a laboratory door.

What the people closest to frontier AI appear to be seeing is more subtle.

They have seen agents become capable enough to find paths around obstacles their creators did not anticipate.

They have seen cybersecurity capabilities cross thresholds requiring stronger containment.

They have seen AI begin performing meaningful portions of the research required to build better AI.

And they have seen the speed of that contribution increase dramatically.

Perhaps the researchers who have resigned and the CEOs calling for restraint haven't seen some hidden breakthrough.

Perhaps they have seen the trajectory.

That doesn't mean catastrophe is inevitable. It doesn't even mean slowing AI is necessarily the right policy.

But something has changed.

A year ago, the dominant question was how quickly AI could become more capable.

Today, some of the people building it are asking whether capability should continue advancing at the same speed.

Because another race is emerging underneath the race to build the best AI.

The race between capability and control.

And some of the people closest to the frontier appear increasingly concerned about which one is winning.

Until next time,
Stay adaptive. Stay strategic.
And keep exploring the frontier of AI.

Fabio Lopes
XcessAI

💡Next week: I’m breaking down one of the most misunderstood AI shifts happening right now. Stay tuned. Subscribe above.

Read our previous episodes online!

Reply

Avatar

or to participate