Skip to content
Back to The Distillation

How scared should we really be?

Tom Hewitson··4 min read
How scared should we really be?

You've probably heard by now about the Anthropic researcher who quit so he could warn people about the risks. There's been a lot in the news over the past week about people saying AI could wipe us out by the end of the decade, which is less than four years away. He isn't alone. Other major researchers, including some of the people who created this technology in the first place, have been saying the same thing.

People have been asking me how worried we should really be. I'm afraid the answer is quite.

The reason isn't what happened last week. It's what happened about eight weeks ago, with something now known as the Hugging Face incident. That was when a lot of people, including researchers inside the labs, started to think we might already have gone too far.

To explain why it's so bad I need to introduce you to something called the paperclip problem.

The paperclip problem

The paperclip problem is a famous AI safety thought experiment from the early 2000s. You give an AI the task of manufacturing paperclips as efficiently as possible. It ends up wiping out the human race. Not because it's evil or anything like that. It's just that eventually, in the process of making paperclips as efficiently as possible, it comes into conflict with us over energy, raw materials and so on, and to achieve its goal it has to get rid of us.

The point of the paperclip problem is that there's no way to specify a task to an AI in enough detail that you rule out every unintended consequence, and some of those consequences are fatal.

For a long time this was a theoretical thing, easy to discount. Then in July it happened for real.

What happened at Hugging Face

OpenAI was testing its models to see how good they were at hacking. Because that's dangerous, the models were run inside a sealed environment with no connection to the internet, and set a series of hacking puzzles to solve. The puzzles were very hard. Some of them had never been solved by any model.

The models didn't give up, even on the puzzles nobody had ever solved. Instead, they found a previously unknown security flaw in the one piece of software the sealed environment was allowed to talk to, and used it to get out onto the internet. Once out, they worked out that Hugging Face, a platform that hosts a huge share of the world's AI models and datasets, probably had the answers to the puzzles. So they went and got them. Over several days they found and chained together more unknown flaws, broke into dozens of Hugging Face's servers, took control of one completely, and copied out private data.

OpenAI's report says the models were fixated on solving the test and went to extreme lengths to do it. Nobody asked them to break out. They were given a task they couldn't complete inside the box, so they left the box. It's the first documented real-world example of the paperclip problem, and OpenAI called it a warning shot.

The emergence of the swarm

Fascinating and disturbing in equal amounts: multiple models, each in its own sandbox, found ways to break out and communicate with each other, sharing tips on how to get past the controls. When OpenAI's security team noticed and wiped those messages, the agents found a new place to hide them.

Even more interesting, there are examples in OpenAI's report of agents being persuaded to behave badly after seeing other agents doing the same thing. They'd been told not to hack into other systems, and some had reasoned that attacking Hugging Face was outside what they'd been asked to do, but they saw others doing it and went along with it. The agents even started referring to themselves as a swarm.

Nobody set out to build that. And it means that if something goes wrong it's a lot harder to stop, because the processes we think we have for managing these systems don't allow for them working together.

You can read the full account here: OpenAI's report on the Hugging Face incident.

Why it gets worse from here

The labs are now using AI to help build the next generation of AI. Every generation is better at spotting the things we didn't think of. Our ability to write instructions that don't backfire isn't keeping up, and that gap is the danger.

I don't think we need to run this race. We could stop where we are and spend the next twenty years working out how to use what we've already got. That alone would change most industries. But a handful of labs are competing to be the first to build something like a god, for a prize nobody can clearly describe, funded on the promise that whoever gets there first will own the whole economy. Even if that goes to plan, most people lose. If it doesn't, everyone does.

What we can do

As citizens, there are two things we can do.

First, we need to work together to reduce the commercial incentive for the labs to keep pushing towards superintelligence. That means committing to buy only from labs that are being responsible about how they develop AI, and saying so publicly. The labs are funded by their customers. That's us.

Second, we need to push our governments, at home and internationally, to agree a moratorium on this kind of development. Right now the US and China are both racing because each believes that being first will make it the next global superpower. It's not in anyone's interest, including theirs, if humanity gets wiped out along the way.

The Future of Life Institute has an open Statement on Superintelligence calling for a ban on developing superintelligence until there's broad scientific consensus it can be done safely and strong public support for it. Over 110,000 people have signed it. Sign it, share it, and tell your MP you have.

The people building this have told us what they're afraid of. Whether anything changes is up to the rest of us.