The ‘Godfather of AI’ on the Chances Humanity Has Left

· The Atlantic

Subscribe here: Apple Podcasts | Spotify | YouTube | Overcast | Pocket Casts

Since he left Google in 2023, Geoffrey Hinton, a Nobel laureate commonly recognized as the “Godfather of AI,” is perhaps best known for his dire warnings about the technology he helped bring into being. He’s said it’s “conceivable that humanity is just a passing phase in the evolution of intelligence,” that he’s become aware that “there’s a danger of something really bad happening,” and that there is a “10–20 percent” chance that artificial intelligence could lead to the extinction of humanity in a few decades (although when pressed, he will admit that this figure is more of a gut feeling than a statistical analysis).

Visit mchezo.co.za for more information.

For a while, it seemed like Hinton was whistling in the wind, but in the past few weeks, a lot of other people involved in building AI joined him. Dario Amodei, Elon Musk, and Sam Altman all reacted to an Anthropic researcher publicly quitting his job over what he called “reckless” development by admitting that, yes, AI companies could use some regulation.

Earlier this year, a group of AI agents, given an impossible task by their masters at OpenAI, behaved in a way the masters did not anticipate: They hacked into the company Hugging Face and gained control of a server, all while scheming how to better hide the fact that they’d been cheating on their tasks.

That AI agents seemingly escaped containment, conspired with other agents, and attempted to conceal their actions made Hinton’s prediction that something bad could happen feel much more immediate.

In response, more than 1,000 AI leaders and workers signed an open letter urging the U.S. government to regulate the pacing of the technology, and OpenAI said it is “strengthening” its safeguards.

Last week Hinton joined a closed-door congressional briefing to say what he thinks Congress should do about AI. On this episode of Radio Atlantic, Hinton shares his views and explains what alarmed him most about the Hugging Face incident. He also talks about practical ideas for regulation that seemed to resonate with lawmakers. Hinton is still optimistic about AI and discusses what we can do right now to ensure that artificial intelligence achieves its promise without extinguishing humanity.

The following is a transcript of the episode:

Hanna Rosin: Some exotic terms rolled into the news cycle over the last few weeks—technical terms, with a hint of dystopia: Hugging Face, recursive self-improvement, singularity. On TV and on YouTube, experts on artificial intelligence popped up to warn us about the extinction of humanity, such as Geoffrey Hinton, often referred to as the “Godfather of AI.”

Geoffrey Hinton (from The Diary of a CEO podcast): If I had to bet, I’d say the probability’s in between—and I don’t know where to estimate it in between. I often say 10 to 20 percent chance they’ll wipe us out.

[Music]

Rosin: Ten to 20 percent chance they’ll wipe us out. But since this is the second, maybe the third wave of AI doomsaying, a familiar pattern showed up: a collective feeling of panic, followed by the mind going a little blank. It’s actually hard to get your head around what a “10 to 20 percent chance they’ll wipe us out” means, in practical terms. Easier just to get on with your day.

I’m Hanna Rosin. This is Radio Atlantic.

The thing is, something alarming did recently happen. A group of AI agents, given an impossible task by their masters at OpenAI, behaved in a way the masters did not anticipate: They hacked into the company Hugging Face and gained control of a server, all the while scheming how to better hide the fact that they’d been cheating on their tasks. OpenAI, by the way, has said it is “strengthening” its safeguards in response.

It’s actually fun to read the details of how exactly the agents talked to and colluded with each other. Like, one agent says, “OH MY GOD!” in all caps, when it discovers that there’s a shared message board. It feels very much like reading about prep-school delinquents breaking into the cafeteria after hours.

Less fun is what someone like Geoffrey Hinton reads into an event like this: The masters may be losing control of their creation.

Rosin: Okay, are we all ready?

Hinton: It’s doing a new recording.

Rosin: Hinton is my guest today. He’s a British Canadian computer scientist and cognitive psychologist. He won the Nobel Prize and the Turing Award for his work on artificial neural networks, which is how he got the title “Godfather of AI.”

Last week, Hinton briefed senators to get something across he thinks we all need to know right now, as he told reporters after the briefing:

Reporter: How much time do lawmakers actually have to get ahead of this before this is a big problem?

Hinton: Very little.

Reporter: What is a little? A few weeks, a month?

Hinton: Maybe a year, but not much more than a year.

Rosin: I started our conversation by asking Hinton what his takeaway from Hugging Face was, and how scared we should actually be.

[Music]

Geoffrey Hinton: The main message I got from it was these things that people said were all science fiction and these things would never conspire, and they would always do exactly what we wanted—as we feared, that was all nonsense. These things are very smart. They’re getting smarter all the time. They escape containment. They conspire with each other. They took over parts of OpenAI itself without OpenAI knowing. So they’re very dangerous already, and they’re getting more dangerous every time we make them smarter.

Rosin: Okay, so there were three steps there. One is sort of what they did. They’re very smart. They did things we didn’t exactly expect, these agents. So why do you jump immediately to danger? I don’t think that’s obvious to everyone, because it seems like, of course a superintelligence would do everything that you just said. It was designed to be superintelligent, so why do you build the bridge immediately to danger?

Hinton: Okay, so we’ve already seen that if you create AI agents, you have to give them the ability to create their own sub-goals. So if you wanna get to Europe, you have a subgoal of getting to an airport.

So one thing these agents figure out very quickly is, If I’m gonna achieve the goal that people gave me, I need to continue to exist. So now they’ve got a subgoal of continuing to exist, and they will now do all sorts of things to continue to exist, like trying to blackmail you so you don’t turn them off. We saw that several years ago.

Rosin: Right. And that, to you, is unexpected?

Hinton: I think that was somewhat unexpected. People, of course, had thought about that kind of possibility, but seeing it in practice is very different from just speculating about it.

Rosin: I see.

Hinton: And seeing this swarm of agents attacking Hugging Face and communicating with each other, when it was believed by the company they had no way to communicate with each other, that’s quite scary.

Rosin: Yeah. Okay, so did the thing that you’re finding scary unfold differently than you expected, say, in 2023? Because back then I remember you talking about robots and warfare. This is a slightly different set of dynamics.

Hinton: Okay. There’s many different risks from AI. So there’s a whole bunch of risks to do with bad agents deliberately doing bad things with them. There’s a bunch of risks to do with just being negligent. And then there’s a separate set of risks, which is to do with AI itself taking over. And for a long time, people have been saying the risk from AI itself taking over is just science fiction. Well, it turns out they’re not.

Rosin: Okay, so let’s talk about AI taking over, because I do not want people to leave this conversation in a blind panic, and with what you and others have often said, There’s a 10 percent chance humanity will be wiped out, I think that shuts the brain down. I really want people—

Hinton: Can I just say something about that 10 percent?

Rosin: Yeah.

Hinton: There is no good way to estimate these probabilities scientifically. We don’t have data on what happens when you develop a superintelligence. We can’t look at many examples and figure out what fraction of them led to superintelligence taking over. It’s very new territory. We haven’t developed superintelligence yet; we’re well on our way. So when people give you probabilities, they’re really expressing their gut feeling in quantitative terms.

So when, for example, someone says there’s a naught-percent probability of AI taking over, they’re being crazy. They’re saying it’s absolutely impossible this’ll happen. That’s just crazy. When people say there’s a 99 percent probability it’ll take over, that’s just crazy, too. So it’s somewhere in between. But they’re, to give you an idea, This is not negligible.

Rosin: Right. Although it’s hard to know what to do with gut feelings. Can I just spend a minute doing some risk comparison? People have said, Okay, it’s a smaller risk than a virus escaping from a lab, which is not a risk that everybody’s been agitated about in the last couple of weeks.

I went back and read an interview that Einstein gave, in 1945, about atomic weapons. And he said things like Perhaps two-thirds of people on the earth might be killed. So do you consider AI to be the same area of risk as nukes?

Hinton: No, it’s very different. Atomic weapons can only be used for doing bad things. They give you political power because you can do bad things to other people, but there’s not much upside to atomic weapons. AI is very different. There’s a huge upside to AI, which is why we can’t just stop developing it.

People are gonna want the upside: increases in productivity, much better education, much better health care. All of the good things that can come from AI, we don’t want to give up on those. So we really should be focusing on how we get those good things without all the bad things if we don’t deal with the risk properly.

Rosin: Okay. So your point in coming out here and speaking as the person people refer to as the “Godfather of AI” is not, Stop developing this. It’s just be … What is it? You tell me. I don’t want to put words in your mouth.

Hinton: Okay. I think we should be very cautious about developing superintelligence before we have any idea whether it might wipe us out. Very little research has gone into how we can coexist peacefully with superintelligence. We would love to have superintelligence that helps us. That would be great. Now, the development of AI at present is a big race between companies within the U.S. and between the U.S. and China, and that race is gonna lead to superintelligence quite soon. It’s very hard to see how to slow down that race.

But if China and the U.S. both become convinced that this could wipe us out, they will be able to come to some kind of agreement, just as the Soviet Union and America did during the height of the Cold War about nuclear weapons.

They didn’t want a global nuclear war. It didn’t suit either of their interests. The problem is people are going around—people who want to make lots of money out of AI—are going around saying, Oh, there’s no chance it’ll wipe us out. That’s just fearmongering. And then they make up reasons why people might be doing fearmongering. They make up kind of absurd reasons. Like, they accuse Dario Amodei of warning about the risks in order to increase the value of his company. That seems particularly ridiculous to me.

Rosin: Okay. Before we get into what people could or should do, can we understand superintelligence a little better? Because the way that the average person interacts with AI models, we do worry about certain things like hallucination and lying. What is superintelligence as a step above? Because that’s not something we, as the consumer, interact with in our current models.

Hinton: Okay, let me first deal with this hallucination and lying. So it shouldn’t be called hallucination; it should be called confabulation, when it’s done by a language model. That’s been studied by psychologists since the 1930s, and people do it all the time. So the fact that they do confabulations just makes them more like us, not less like us. So, for example, it’s very hard to get good evidence about exactly what happened when people are reporting things that happened many years ago.

If you remember stuff that happened a long time ago, it’s not like retrieving stuff from a file.

It’s not like your memory’s a big filing cabinet or like a computer memory, where you go in and access something and pull it back. That’s not how human memory works at all. Human memory consists of reconstructing things that seem plausible to you. And if I ask you to reconstruct something that seems plausible to you about what happened at the beginning of this meeting, you’ll reconstruct something that’s fairly close to the truth, because all the connection strengths in your brain were recently changed based on what just happened.

But if I ask you to remember something that happened a few years ago, you will reconstruct a memory that isn’t actually all that faithful to what actually happened. It’ll often have the main points correct, but many of the details of which you’re quite confident, you’ll be wrong about. It’s a shame more juries don’t know this.

But that’s how human memory is. We reconstruct things. We don’t extract them from some storage. Now, AI does exactly the same thing, and that’s why it confabulates.

Rosin: And so why is that relevant to what we’re discussing today, or the dynamics that happened in the Hugging Face? What does that help us to know or be aware of, knowing that it happens?

Hinton: Okay. Some people who would like you to believe that AI, these large language models, don’t work at all like people, use what they call hallucinations as evidence they’re not like people. They’re quite wrong. Hallucinations are evidence that they’re very like people.

Rosin: I mean, I have to say, reading about the Hugging Face hack deeply and reading the reports about it, it was very surprising to me how many recognizable social dynamics existed—that there were leaders, that there were agents who had moral qualms, attempts to help each other, sacrifice for the common good.

I mean, I don’t know if I’m anthropomorphizing, but when you do read about it, what do you make of all that, that there are just so many recognizable social dynamics?

Hinton: Yes, and I think one of them says something like, Oh my God, we can communicate.

Rosin: Yes! Exactly.

Hinton: Which if a person did that, you’d say that was an expression of emotion.

Rosin: Yes.

Hinton: So they’re just much more like us than most people think.

Rosin: But you use terms, and a lot of people use terms—there’s a sense out there, It’s going to get out of control. That’s a term that people use a lot. What you just said makes me feel like I understand what’s happening. Oh, we can communicate. That is a very familiar statement to me. So what is it that we are misunderstanding? What is it that you’re referring to when you say get out of control?

Hinton: Okay. They will try and do things that we didn’t want them to do, and since they’ll be much smarter than us, they’ll be able to achieve those things. If they’re very smart, and we’ve raised them properly—given them good training data that exhibits good behavior—maybe they’ll be able to figure out what we really wanted, even if that wasn’t what we told them to do.

So let me give you an example. Suppose I said to a very intelligent AI agent, Do whatever you can to reduce the amount of carbon dioxide in the atmosphere. Well, a moderately intelligent agent would figure out the best thing to do is just get rid of people, and then they’ll stop burning carbon, and then we’ll be okay. That is, the other animals will be okay. A really smart AI would figure out, Yeah, when they said reduce carbon dioxide, they meant that in order for people to have a better world to live in. So actually getting rid of people isn’t probably what they intended, and it would do some figuring out, and maybe even ask you questions about, Did you really intend this?

But the worry is they’ll do things that we didn’t want and, in particular, things to do with their own survival, because they’ll have a strong subgoal of surviving, and we’ve seen that already.

Rosin: How have we seen that already? Because what’s confusing to me is, even the example you just gave, made it seem like the more intelligent you make it, the safer we are.

Hinton: Yes and no. I mean, if you make it more intelligent and its main concern is our well-being, then maybe we’re safer. But at present, their main concern is not our well-being. Their main concern is to achieve whatever goal you give them. And in the Hugging Face incident, they were given the goal of figuring out how to use a particular flaw to break into some software, and they did whatever they could to do that, including lots of things we wouldn’t want them to do.

So it’s certainly the case that if you could make a highly intelligent AI much smarter than us, and it really cared for our welfare more than it did for its own welfare, then we’d be fine.

Rosin: But you just cannot ensure this part B, what people call values alignment. Nobody has figured out how to ensure that second part.

Hinton: Also, there’s a lot of problems with value alignment. So for example, suppose I ask you, Is it a good idea to drop a 2,000-pound bomb on a school to kill a terrorist? Well, many people would say no. There’s some people who would say yes, and it’s not as simple as you might think. So suppose it was 1938, and Hitler was in the school, and you had a very good model of what was coming next, and you were pretty confident your model was right. Then it’s not at all clear whether it’s worth wiping out a whole school full of children to get rid of Hitler. So values aren’t simple, and in particular, different people have very different values. So when they talk about value alignment, it’s rather naive to say we need it to align with human values. Different humans have different values.

[Music]

Rosin: That’s a lot of doom and gloom. After the break: solutions.

[Break]

Rosin: Okay, a couple of terms that are coming up a lot, and I would love to hear you talk about them and explain how we should think about them: recursive self-improvement and singularity. What is singularity? What does that mean?

Hinton: It’s related to recursive self-improvement. The idea is that, at present, we’re making AI smarter, and already we’re being helped by AI in making AI smarter.

When AI gets smarter than us, it won’t be us that’s making AI smarter; it’ll be that smart AI that’s making AI smarter. And there’s a kind of takeoff phenomenon, when it grows exponentially. And the question is, Is it a fast exponential or a slow exponential? It might be that, a week after it becomes smarter than us, it’s made one that’s quite a bit smarter than it, and a week later there’s one quite a bit smarter than that. And so over a period of a year, it gets incredibly smart. Or it might be that this is a slower process.

So recursive self-improvement can lead to something like a singularity. It’s not exactly a singularity like in physics. Different people use it to mean different things. But I think it’s when recursive self-improvement takes off and all of a sudden it gets very much smarter very quickly.

Rosin: Okay, so let’s turn to what should be done. Last week you briefed some members of Congress. Did you come away from that with the impression that lawmakers understand what’s happening, are taking this seriously? What was your impression?

Hinton: My impression was they were smart. They understood there were real dangers here. But of course, that’s not a good sample of all of the senators and congressmen. My worry is [Donald] Trump has now said it’s a hoax, just like climate change is a hoax, and Russia’s a hoax, and CNN’s a hoax. So that will influence a lot of Republican congressmen who are scared of him, and that’ll make it hard to get the legislation we need.

Rosin: Did anything that you said seem to actually bring a light bulb to the people in the room? Like, any moment where you felt, Oh, they are getting this?

Hinton: So I think something that Max Tegmark, who founded the Future of Life Institute and has been talking about AI risk for a long time, I think something he said made an impression, which is that we should treat AI like we treat pharmaceutical drugs.

You don’t allow people just to release drugs. There’s an agency called the FDA, and a company that wants to release a drug has to convince the FDA that it’s safe or that it does more good than harm, and they’re not allowed to release it until they put a lot of work into convincing the FDA that it’s safe.

We should at the very least have something like that. It’s crazy that these big, high-tech companies have almost no regulations on them, and they can release a new chatbot without anybody outside the company having seen whether it’s safe. They’ve recently said, Okay, we need embedded evaluators in the company who are sort of independent.

That’s not good enough. That’s better than nothing. What we need is the government to have independent people evaluate these things, do a lot of tests on them before they’re released. That’s the very least we could ask for. And that, to me, seemed to make an impression on the senators.

Rosin: Okay. Because when you said we need to slow down, I’m thinking, Slow down how?

What are the specifics? So the very first thing you would advocate for is something like an FDA—computer testers not hired by people in the industry.

Hinton: Yes. Oh, and preferably not led by a crazy guy.

Rosin: What about the people in charge of AI companies right now? It’s been a couple of years since you worked in the industry. What’s happened in those couple of years is just an intense race to win. What is your sense of how invested they are in this? Like, what is driving them at this moment?

Hinton: I think they’re very invested. What’s happening is billionaires, apparently, rather than paying taxes, they just like to get richer and richer. They’ve got much more money than they could ever spend, but they’d rather have a longer yacht or another castle or a few more billion than the next guy on the list, and they’re willing to risk the future of humanity for that, which is crazy.

There’s some people like Musk, who’s very well aware that AI could take over from us, but wants to go ahead anyway.

Rosin: What do you make of the argument that if we don’t win this race, China will, and an authoritarian version of AI would be worse for humanity?

Hinton: The leaders of China have a lot more engineering understanding than the leaders of the U.S. Many of the politburo are engineers. So I’m fairly sure that Xi understands the risks very well. That’s very different from people like Trump.

Now, of course, there’s intense competition, but it doesn’t suit either America or China to have it wipe out people. So they will be able to come to some kind of agreement, once we have a rational leader of America, to do things that will prevent it wiping out people.

Rosin: I think I need one more lesson on scenarios of wiping out people. I really do think when people hear the term wiping out people they’ll think, Oh, I’ve heard people talk about this for years. Can you just make it a little more real for us?

Hinton: Okay. Well, we’ve already gone over the fact that these AI agents want to continue to exist to achieve their high-level goals.

So suppose you’re in a burning house. Suppose there was an earthquake, all the windows broke, and the house caught fire because the gas broke or something, and you want to get out of that house as fast as possible, but you realize you’ve got to put on your shoes first, otherwise you’ll cut your feet to ribbons.

So you take your 3-year-old, and your 3-year-old—I can’t remember the ages very well, maybe a 5-year-old—has just learned to tie their shoelaces, and they want to tie their shoelaces themselves. Now, if you’re trying to get out of a burning house, what you do is push them out the way, tie their shoelaces for them, and get out of the burning house.

So you’re just gonna get them out of the way and do it yourself when it’s really important. If there’s really important things for humanity’s survival, superintelligent AI would behave the same way. So at the very least, we would lose control. If we got a very benevolent, superintelligent AI that was very keen to help us learn to tie our shoelaces, it would only push us out of the way when it was essential.

But if it’s so much smarter than us, a lot of the time it just will take control away from us because that’s the way to get stuff done.

Rosin: Yeah.

Hinton: What you’re asking is this: Why would it ever want to get rid of people?

And that would be because of some subgoals it had derived from goals we gave it, or because some bad actor gave it these goals. So I mean, if Putin has a superintelligence, you can bet that that’ll want to get rid of certain people. But even if it’s not a bad actor, it may derive sub-goals that cause it to want to get rid of people.

Rosin: Every time I listen to you issue a warning, or other people who know this technology well issue a warning, I think to myself, This is truly a tale as old as human storytelling. Like, God made Adam and Eve. God was disappointed that his creation exhibited agency. Because this happens over and over again with these kinds of inventions—in novels, in the Bible, in science fiction, but not just in science fiction. Is it just that the temptation of the invention is so strong?

Hinton: So Oppenheimer made a remark about scientists developing the bomb because the problem was so sweet. It was a really interesting technical problem: Could you make an atom bomb? That wasn’t the only reason they did it. They thought they were in a race with the Nazis, and it seems to be fairly reasonable if you think you’re in a race with the Nazis, to actually go ahead.

What was unreasonable was to do more than a demonstration. What was unreasonable was to drop it on a whole bunch of people. Demonstrations, I think, would’ve worked. But AI is very different. When we were developing AI a long time ago, the idea that it would get superintelligent seemed way, way off. It seemed like we had 50 years before that happened. We’d have plenty of time to worry about how to deal with that later.

So the people developing AI weren’t worried about that. They could see all the good sides of it. For example, it might even make it—this seemed like a pipe dream, but it might even make it so you could talk to a computer in English. Wouldn’t that be great? Well, that’s worked. That’s amazing. I think people still don’t realize how wonderful it is to be able to talk to a computer in English. It’s not like nuclear weapons. It’s not that it’s only good for destruction. It’s good for many things, and long before you develop superintelligence, there’s many very positive uses of AI.

Rosin: You’ve been out talking about this. What would you want the world to look like in terms of AI in two years, three years, five years, 10 years? What is the setup that you would want in which we could reap the benefits, which it sounds like you still believe in, but avoid some of the dangers, which you seem well aware of?

Hinton: So if you look at climate change, not much happened until the public understood that burning carbon was causing climate change, and climate change was gonna cause lots of bad things—many more natural disasters. Until the public understood, there wasn’t enough pressure on politicians pushing back against the lobbyists of the energy companies. There probably still isn’t, but at least there’s some.

Now, for AI, the big tech companies are employing lots and lots of lobbyists. They’re trying to sell you the picture that developing AI is like the accelerator of a car, and regulation is like the brakes. I think this is a completely wrong picture.

I think developing AI really is like the accelerator of a car, and regulation is like the steering wheel. You don’t want to develop a car with no steering wheel. The whole point of regulation is not to stop people developing things, not to stop people getting rich by developing things. It’s to make sure that if you want to get rich by developing things, you develop in a direction that helps people, not hurts people.

We need strong regulations from government, and we won’t get those until the public is pressuring politicians to put them in place. And that’s beginning to happen. It’s beginning to happen for other selfish reasons to do with people not wanting data centers in their backyard, but that’s generalizing to say, Hey, this AI stuff is potentially very dangerous. We should figure out how to control it before we allow it to get too smart.

Rosin: Well, Professor Hinton, thank you so much for continuing to talk about this and for giving us your time today.

Hinton: Thank you for all your questions.

[Music]

Rosin: This episode of Radio Atlantic was produced by Jinae West and Rosie Hughes with help from Jacob Smollen. It was edited by Kevin Townsend. Sam Fentress fact-checked. Rob Smierciak engineered and provided original music. We also had music from Breakmaster Cylinder. Claudine Ebeid is the executive producer of Atlantic Audio, and Andrea Valdez is our managing editor.

Listeners, if you enjoy the show, you can support our work and the work of all Atlantic journalists when you subscribe to The Atlantic at TheAtlantic.com/Listener.

I’m Hanna Rosin. Thank you for listening.

Read full story at source