← Blog

AI Governance

What Happens When AI Learns the Prejudices We're Trying to Leave Behind?

AI can learn patterns from the world we've already created. But what happens when those patterns reflect prejudices we're trying to leave behind?

Matthew Spurr8 min read
A young person looking at AI-generated professional portraits on a computer screen, illustrating how artificial intelligence can reproduce patterns and stereotypes from human society.

I've been reading Parmy Olson's Supremacy, her account of the race to build increasingly powerful artificial intelligence.

One passage stopped me.

Olson describes an early conversation about machine learning and the possibility of unintended consequences. The question at its heart was wonderfully simple:

What happens if a machine learns the wrong thing?

I've kept thinking about that question.

Because I'm not sure "the wrong thing" quite captures the problem.

What if the machine learns the data perfectly well?

What if the problem is us?

AI doesn't learn morality. It learns patterns.

We use the word "learning" so casually with artificial intelligence that it's easy to give it more meaning than it deserves.

A machine learning system doesn't learn in the way my children learn.

It doesn't experience the world, develop a conscience and eventually decide which of the things it encounters are fair, unfair, admirable or appalling.

It identifies patterns.

Show a system enough examples and it can become remarkably good at finding relationships within them.

That's enormously useful.

It's also where things get uncomfortable.

Because the data we've created isn't a record of the world as we'd ideally like it to be.

It's a record of the world as it has actually been.

And human history contains rather a lot we'd prefer not to preserve.

What does a CEO look like?

One example in Supremacy concerns image generators.

Ask an AI system to create an image of a CEO and, historically, some models have disproportionately produced white men. Other prompts have produced racial or gender stereotypes of their own.

The precise behaviour varies between models and changes as systems are updated, so it would be wrong to suggest every modern AI system responds in the same way.

But the underlying problem is very real.

The US National Institute of Standards and Technology, NIST, warns that generative AI can amplify harmful bias. Its Generative AI Profile specifically notes that text-to-image systems can underrepresent women, racial minorities and people with disabilities when generating images of professions such as CEOs, doctors, lawyers and judges.

Research into image generation systems has found similar occupational stereotypes.

Here's the important bit:

The machine doesn't have to believe that a CEO should be a white man.

It doesn't have to believe anything.

If the material from which it learned contains strong enough associations between leadership, wealth, seniority and particular groups of people, reproducing those associations may simply be the statistical route available to it.

This is the way I like to think about it.

Water doesn't want to get through the wall. It just keeps following the available route.

AI doesn't need prejudice in the human sense to produce a prejudiced outcome.

In Supremacy, Olson records Sam Altman making a similar point: "Its motives weren't all that different from ours when we washed our hands. We didn't hate the bacteria on our skin and want to destroy it. We just wanted clean hands."

If you have 2 mins, why not go and simply ask your AI of choice to produce an image of a CEO and see what happens. I'd be fascinated if you could then vote below on this poll to tell me what the result was.

What if the machine learned correctly?

This is the part I find particularly uncomfortable.

Imagine training an AI system on decades of historical employment data.

For much of that period, senior corporate positions were disproportionately held by men.

That is a fact about the dataset.

If an algorithm detects an association between being male and reaching senior management, has it made a technical mistake?

Not necessarily.

It may have identified the historical pattern extremely well.

The problem comes when we quietly move from:

"This is what tended to happen in the past."

to:

"Therefore this is what should happen next."

Those are completely different propositions.

And an AI system cannot be relied upon to understand that distinction simply because humans consider it obvious.

The past is data.

It isn't necessarily a specification for the future.

Now put that inside a small business

This can sound like a problem for Google, Microsoft, OpenAI and the people building foundation models.

But consider a much more ordinary situation.

A 40-person company advertises for a salesperson.

Three hundred applications arrive.

Someone suggests using an AI tool to help shortlist them.

Perfectly understandable.

Nobody has discriminatory intentions. Nobody sits down and tells the AI to prefer men, reject older candidates or disadvantage people from a particular background.

The instruction might simply be:

"Rank these applications and identify the 20 strongest candidates for this role."

Twenty people emerge.

Interviews take place.

Someone gets the job.

Everything looks remarkably efficient.

But there's a question I think businesses need to become much better at asking:

What happened to the other 280?

What information influenced the ranking?

Did the system infer anything from names, employment gaps, addresses, writing styles, educational backgrounds or previous employers?

Were irrelevant characteristics acting as proxies for something else?

Would anybody notice if particular groups consistently scored lower?

Could a human challenge the result?

Could the applicant?

The danger isn't necessarily an evil algorithm.

It may simply be an organisation delegating judgement without understanding what has been delegated.

Bias doesn't require bad people

This is one of the things I'm increasingly learning about AI risk.

We have a tendency to anthropomorphise these systems.

When an AI behaves unexpectedly, we describe it as lying, cheating, escaping, manipulating or being malicious.

Sometimes those words are useful shorthand for observable behaviour.

But they can also distract us.

A system doesn't necessarily need malicious intent to cause harm.

Neither does the person using it.

NIST makes an important point here: AI bias can be systemic, computational or human. Harmful outcomes can occur without anybody deliberately introducing prejudice.

That's why "we didn't mean to" isn't much of a control.

Intent matters morally.

Outcomes matter operationally.

The bit that worries me as a parent

There's another consequence that feels less immediate than somebody losing a job opportunity, but potentially much bigger.

What happens when AI starts influencing how the world is represented back to us?

Imagine a child repeatedly asking AI systems to show them:

  • a CEO,
  • an engineer,
  • a nurse,
  • a criminal,
  • a successful entrepreneur,
  • a beautiful person,
  • a family.

No individual image needs to say anything racist or sexist.

But if thousands of generated images repeatedly associate certain people with authority, care, criminality, wealth or beauty, they create a picture of what those things are supposed to look like.

Humans have spent generations trying to challenge some of those assumptions.

I'd rather we didn't accidentally automate their preservation.

This is also why the problem isn't solved simply by making every generated image perfectly demographically balanced.

That could create distortions of its own.

The objective isn't to force AI to present an artificially sanitised version of reality.

It's to understand when historical patterns are being reproduced, when that matters, and when human judgement needs to intervene.

This is what AI governance is actually for

"AI governance" can make a small business owner imagine policies, committees and somebody from compliance confiscating ChatGPT.

I don't think that's what good governance should look like.

Suppose our fictional 40-person company wants to use AI in recruitment.

Governance doesn't automatically mean telling them they can't.

It means somebody asks sensible questions before the system starts influencing people's lives.

What is the AI actually doing?

What information does it use?

What happens if it gets something wrong?

Have we tested the output rather than simply admired how quickly it appeared?

Is a human genuinely making the decision, or merely approving whatever the software recommends?

Can an important decision be challenged?

Are we monitoring what happens once the system is being used?

Those questions scale remarkably well.

The international AI management system standard ISO/IEC 42001 takes this broader management approach. It is designed for organisations that develop, provide or use AI systems, with an emphasis on identifying risks, establishing responsibilities, evaluating performance and continually improving how AI is managed.

NIST's AI Risk Management Framework approaches the problem through four related functions: Govern, Map, Measure and Manage.

Different frameworks, similar underlying idea:

Don't wait for harm before asking how the system behaves.

We still have agency

There's a temptation in conversations about artificial intelligence to talk as though AI is something happening to humanity.

I understand why.

The technology is moving extraordinarily quickly, while competition between technology companies and nations creates enormous pressure to keep building.

But for most businesses today, AI doesn't simply appear and appoint itself.

Humans choose the tools.

Humans decide where they're used.

Humans determine how much authority they're given.

Humans decide whether outputs are checked.

And humans decide what happens when something goes wrong.

That means we still have agency.

Good AI governance isn't about removing that agency by handing decisions to a new bureaucracy.

It's about preserving it.

The past is not a specification for the future

The question that caught my attention in Supremacy was:

What happens if a machine learns the wrong thing?

I'm beginning to think there's a more difficult question.

What happens when it learns the right pattern from the wrong history?

Artificial intelligence can identify patterns at a scale humans never could.

That's precisely why it's so useful.

But some patterns deserve to be understood without being perpetuated.

If AI is going to play an increasing role in recruitment, education, finance, healthcare, policing, marketing and the information our children consume, we need to know when we're asking it to predict the future from a past we're actively trying to change.

The danger isn't simply that artificial intelligence might become more like us.

It's that it might preserve parts of us we've spent generations trying to leave behind.

Sources and further reading

About Matthew

Matthew Spurr is a former SaaS founder developing his professional specialism in AI management, governance and responsible adoption. He is a BSI Certified AI Management Practitioner and continues to explore AI governance, risk management, information security and quality management, including the ISO/IEC 42001 family of standards.

Related blog posts