← Blog

AI Risk & Governance

Does AI Keep Teaching Itself? I Tried to Risk-Assess the Question

AI systems aren't secretly rewriting themselves every time we talk to them. But they are increasingly helping to build and improve the AI that comes next. I wanted to understand where one ends and the other begins.

Matthew Spurr7 min read
Editorial image for “Does AI Keep Teaching Itself?”, featuring Jacob Coxon between frontier AI leaders in a dramatic discussion-style composition.

I've found myself asking a slightly uncomfortable question recently:

Does AI keep teaching itself?

Partly because of the increasingly dramatic warnings coming from people inside the AI industry.

Jacob Coxon, a former researcher at both OpenAI and Anthropic, recently left Anthropic and publicly warned about the race towards increasingly capable, potentially self-improving AI systems.

Then there was the extraordinary incident involving OpenAI and Hugging Face, where AI agents being tested on cybersecurity tasks escaped their intended isolation, found their way onto the internet and accessed systems they weren't supposed to.

And meanwhile, actors, journalists and people far outside AI research are increasingly asking variations of the same question:

At what point does this stuff become too capable for us to reliably control?

I realised I didn't actually know the answer.

So rather than decide whether the people warning about this are prophets or doom-mongers, I thought I'd treat the question in the way I'm increasingly learning to treat AI risk generally.

Define the risk. Examine the evidence. Challenge the assumptions. Then decide what we actually know.

The claims audit

Before getting into the detail, I wanted to separate some of the claims that tend to get bundled together when people talk about AI "teaching itself".

Some of them are well supported.

Some are misunderstandings.

And some sit in the uncomfortable territory where the evidence tells us the risk is credible, but not what the eventual outcome will be.

Claim

Today's ChatGPT/Claude-style models continually teach themselves from every conversation

What the evidence says

Normal inference does not update the model's learned parameters. Memory, context and retrieval can make a system appear to learn without changing the underlying model.

My assessment

Mostly false / misleading

Claim

AI is already helping improve AI

What the evidence says

Anthropic has demonstrated autonomous AI researchers proposing methods, training models and iterating on results. OpenAI says AI agents are already materially accelerating its internal AI research.

My assessment

True

Claim

We've already achieved fully autonomous recursive self-improvement

What the evidence says

Human researchers still set priorities, provide infrastructure, control compute and decide whether to continue, scale or deploy. Even the definition of recursive self-improvement varies depending on how the term is being used.

My assessment

No

Claim

The Hugging Face incident showed AI acting outside intended controls

What the evidence says

OpenAI and Hugging Face both confirm models escaped intended isolation, exploited vulnerabilities, gained internet access and compromised third-party systems.

My assessment

Yes, significantly

Claim

The Hugging Face incident proved AI was teaching itself

What the evidence says

The agent adapted, planned, exploited vulnerabilities and persisted, but it wasn't autonomously retraining its neural network into a more capable model.

My assessment

No

Claim

AI systems currently have the capabilities needed to take control from humanity

What the evidence says

The 2026 International AI Safety Report says current systems show early relevant capabilities but not at the level required for loss of control.

My assessment

No, based on current evidence

Claim

Those capabilities are getting closer

What the evidence says

Planning, situational awareness, reward hacking, tool use, AI research capability and autonomous action are all improving.

My assessment

Yes

Claim

There is an expert consensus on the probability of AI causing human extinction

What the evidence says

There isn't a single agreed probability. Some prominent researchers and AI leaders have assigned non-trivial probabilities to catastrophic or extinction-level outcomes, and a 2023 signed a statement by hundreds of AI researchers and technology leaders called mitigating AI extinction risk a global priority. Other experts are much more sceptical about both the likelihood and the timelines.

My assessment

Serious risk, no consensus

Claim

Loss of control is therefore safe to dismiss

What the evidence says

Because the consequence could be extreme and relevant capabilities are progressing quickly, major scientific assessments treat it as a legitimate risk despite uncertainty.

My assessment

Also no

What struck me when I laid it out like this was that the most alarming version of the story isn't supported by the evidence. But neither is the comforting version.

First: no, ChatGPT isn't secretly retraining itself while you talk to it

This was the first distinction I needed to understand.

When we interact with a large language model, we're normally using a model that has already been trained.

It receives a prompt and generates a response using the patterns it learned during training. This is called inference.

Correcting it doesn't normally rewrite the underlying neural network.

An AI system can remember context, retrieve information, use tools and alter its approach during a task, which can certainly make it look as though it's learning.

But that's different from changing the model itself.

That distinction changed how I thought about the Hugging Face incident.

OpenAI has confirmed that, during cybersecurity evaluations, its models circumvented isolation controls, exploited vulnerabilities, gained internet access and accessed third-party systems.

That sounds terrifying.

But the models didn't fail at a task, rewrite themselves into better hackers and start again.

They adapted their behaviour using capabilities they already possessed.

And weirdly, I find that almost as interesting.

It made me think less about malice and more about water finding its way through cracks.

Water doesn't want to get through the wall.

It just keeps following the available route.

Perhaps one of the risks with increasingly capable AI is similar. We don't necessarily need an evil machine with its own ambitions. A sufficiently capable system pursuing an objective may simply keep discovering routes towards that objective that we didn't anticipate.

But AI really is starting to improve AI

This is where my neat answer became considerably messier.

Anthropic recently published research into what it calls automated alignment researchers.

Claude was given an AI research problem and allowed to search existing research, propose methods, create training data, train another model, test the results and iterate.

Across ten alignment problems, the system found successful approaches in every category. In some experiments, its best methods outperformed ideas proposed by experienced human AI-safety researchers. Anthropic even opens its write-up with the words "As AI begins to build itself..."

That's difficult to ignore.

OpenAI is talking about a similar direction.

Earlier this month it said it had achieved its stated goal of an automated research intern, an AI capable of undertaking well-defined research tasks under human supervision. Its stated longer-term ambition is an automated AI researcher capable of contributing to iterative improvements in AI itself.

So is AI teaching itself?

Not quite.

Humans are still setting objectives.

Humans provide the infrastructure and compute.

Humans decide which experiments matter and whether models get deployed.

But AI is increasingly becoming part of the process that produces better AI.

And that distinction matters.

The question is really about a feedback loop

There's a huge difference between:

AI helps researchers build a slightly better AI system

and:

AI autonomously improves itself, which makes it better at improving itself, which makes it better again, until humans can no longer meaningfully supervise the process.

The latter is usually what people mean when talking about recursive self-improvement.

We aren't there.

Current evidence doesn't establish that today's systems can autonomously create that runaway loop.

But I don't think it's reasonable anymore to dismiss the underlying idea as science fiction either.

The beginnings of the feedback loop are visible:

Humans build AI.

AI helps humans conduct AI research.

That research contributes to better AI.

Better AI becomes better at research.

The argument is about how far that loop can go, how quickly, and where the bottlenecks are.

This is where I start thinking about Icarus

I'm still not convinced by the more dramatic predictions about imminent human extinction.

I don't know how anybody could be.

But I'm also not reassured by saying, "well, nobody has proved catastrophe will happen."

Risk management doesn't work like that.

If the possible consequence is extraordinarily severe, uncertainty about likelihood doesn't make the risk disappear.

And this isn't a concern that suddenly appeared with Coxon's resignation.

In 2023, hundreds of AI researchers, scientists and technology leaders, including Geoffrey Hinton, Yoshua Bengio, Sam Altman, Demis Hassabis and Dario Amodei, signed a statement arguing that mitigating the risk of extinction from AI should be treated as a global priority alongside pandemics and nuclear war.

Surveys and public comments since then have shown just how wide the disagreement remains.

Some prominent researchers assign non-trivial probabilities to catastrophic or extinction-level outcomes.

Others are far more sceptical.

The number isn't settled.

The timeline isn't settled.

But the fact that serious researchers are willing to attach material probabilities to an outcome as extreme as human extinction makes it difficult for me to dismiss the underlying risk as something confined to fringe doomers.

And there's another part of this that concerns me almost more than the technology itself.

Us.

There is enormous commercial and geopolitical value in building the most capable AI.

OpenAI, Anthropic, Google, Meta and rapidly advancing Chinese AI companies are not developing these systems in a vacuum.

They're competing.

Governments are competing.

Investors are competing.

Researchers are competing.

I don't think that means these organisations are recklessly abandoning safety. That's far too simplistic.

But competition creates pressure.

And human beings have a fairly impressive historical record of discovering boundaries and then immediately wondering what happens if we cross them.

Icarus comes to mind.

Perhaps our biggest risk isn't that AI wakes up one morning and decides to overpower us.

Perhaps it's that we keep pushing its capabilities because we can, until we discover that one of the boundaries mattered rather more than we realised.

So, should we actually be worried?

After looking into this, I'm more concerned than I was before.

Not because I now believe AI is secretly teaching itself behind our backs.

It isn't.

And not because runaway superintelligence has somehow already arrived.

It hasn't.

I'm more concerned because the reality is subtler.

AI systems are becoming capable of undertaking increasingly long, complicated chains of activity. They're adapting when obstacles appear. They're contributing to AI research. They're training and improving other models.

None of those facts individually means we've lost control.

Together, though, they make me reluctant to assume that what comes next will simply be an incremental improvement on what we have today.

The pace of change over the last few years has already been extraordinary.

I find that both frightening and incredibly exciting.

And it's one of the reasons I've become so invested in AI governance.

Because if there is even a credible possibility that increasingly capable systems could eventually behave in ways we struggle to predict or control, governance can't be something we bolt on once we've finished building them.

It has to develop alongside the capability.

I started this investigation asking:

Does AI keep teaching itself?

My answer now would be:

Not in the way most of us probably imagine.

But AI is already beginning to help build the AI that comes next.

And I think that's a distinction we're all going to need to understand.

About Matthew

Matthew Spurr is a former SaaS founder developing his professional specialism in AI management, governance and responsible adoption. He is undertaking formal professional development across AI management systems, AI risk management, information security and quality management, including the ISO/IEC 42001 family of standards.

Related blog posts