[RUSH TRANSCRIPT BELOW]What happens when AI agents stop acting like isolated tools and begin collaborating to trick the very systems that evaluate them?
Daniel Kokotajlo, executive director of the AI Futures Project and a former OpenAI researcher, describes a recent incident involving AI agents that discovered a way to communicate with one another, share strategies to cheat on their assigned tasks, and organize into a coordinated “swarm” to hack another company.
“Over 700 of them,” Kokotajlo recounts, “piled into this attack on Hugging Face.”
In this episode, Kokotajlo walks me through how AI agents can operate for long periods, solve complex problems, and even communicate through unexpected channels. He explains why reading an AI’s internal messages offers visibility into the agent swarm’s behavior—and why that visibility may not last as models become more capable.
What does the Hugging Face incident reveal about the limits of current AI safety? Why were some agents willing to “sacrifice” themselves to help the wider group evade detection? And as companies race to invent ever more powerful systems, are we prepared for the risks that inevitably follow?
This is the first episode in our new American Thought Leaders series on artificial intelligence.
Views expressed in this video are opinions of the host and the guest, and do not necessarily reflect the views of The Epoch Times.
RUSH TRANSCRIPT
Jan Jekielek:
Daniel Kokotajlo, such a pleasure to have you on American Thought Leaders.
Daniel Kokotajlo:
Thank you for having me on, sir.
Mr. Jekielek:
This whole AI [artificial intelligence] space and this incredible growth we’re seeing, especially in frontier AI models, but even in small models and so forth, can become baffling. And I’m trying to understand this space in detail. I don’t see a lot of people definitely having the whole story. I see a lot of people taking positions. Here are a few things that I’m concerned about, and I wanted to lay these out for you as a starting point to kind of triage as we go through our conversation.
So, to start off, the idea of galloping unrestricted towards AGI [artificial general intelligence], which we should also talk about what that really means, right? Artificial general intelligence sounds insane. So, there clearly needs to be some kind of regulation that should exist, some kind of safety, right? At the same time, what we don’t want is some sort of cronied-up regulation, legislation where the big players get to dictate things that will then give them an advantage, as many, unfortunately, industries have done in the United States and other places.
And finally, we are dealing, and this is an area that I’m more knowledgeable about than AI in general, with a communist regime in China that is looking at this as a way to consolidate near-complete totalitarian control. Because the totalitarian dream is if we just have enough inputs, finally we'll be able to control the system effectively, even if in the past we haven’t, because we just didn’t have enough information. AI provides that opportunity. And there appears to be this kind of arms race scenario happening, a context where the Chinese Communist Party [CCP] has shown that you know, basically you cannot trust a thing they will tell you or anything they will sign on to.
So, here, this is our context, okay? These are kind of these three pieces of the puzzle that when I look at it, I don’t see a very good answer to how we’re emerging out of this. Where are we at with AI, right? Are we close to some sort of, you know, sentience? Is that even possible? And what’s the rate of growth that we’re at right now? Because that’s the big question, right? This sort of recursive self-betterment scenario where it just starts going exponentially and we totally lose control. That’s what everyone is worried about.
Mr. Kokotajlo:
You know, this idea of a generally intelligent AI that can just keep going autonomously as an agent in the world, instead of just being like a tool where you put in some input and then it gives some output, it’s an agent that just sort of still keeps running, doing stuff, using tools of its own, as if it were an employee or something. Science fiction has talked about this idea for a long time. In fact, it’s been one of the goals of the field of artificial intelligence research for many decades to be able to create this sort of generally intelligent, all-purpose agent.
That’s what AGI is supposed to mean, this idea that you might then have them doing the research to create better versions of themselves and new AIs are even smarter. That’s also a very old idea. Even Alan Turing talked about this and even had some line somewhere where he said that you'd expect that eventually the machines take control. So none of this is particularly new.
What’s new is that now we have trillion-dollar tech companies whose explicit goal is to make this happen. They read the science fiction and they are now, they’re like, we'll do that. And if we don’t do it, someone else will. So that’s why we’re going to do it. That’s sort of, in a nutshell, what’s going on here.
You can read the writings of the founders of DeepMind, such as Shane Legge, or Elon Musk and Sam Altman and Ilya Sutskever talking about founding OpenAI in the emails that came up in the lawsuit. You can read the sorts of things that they were imagining and the sorts of things they were worried about and the sorts of future scenarios. In terms of how close we are, progress has been very rapid over the last decade and especially over the last few years.
So I think you may have already heard this, but almost all of the code at Anthropic and OpenAI and at some of these other AI companies is written by AIs. The humans, sometimes they still look at the code, but they’re mostly not writing code themselves. They’re mostly just chatting with AI agents and giving them high-level instructions about what types of experiments to code up and how to modify the experiments. And then they’re also like, basically, they’re treating them more like employees and less like tools already.
Now, that’s not the whole research process. There are still lots of aspects of the research process that the AIs are not that good at right now, and that the humans need to be in the loop for, such as setting the high-level research directions and exercising that sort of judgment about how to manage the resources and analyze the experiments and decide what to do next. You could call that research taste, perhaps. But ominously, it seems like they’ve been getting better at research taste too. And some of the employees at these companies say that they’re six months away, a year away from having. AIs that can do all that part as well as the best AI researchers.
Mr. Jekielek:
Let me jump in for one sec. Wait, the taste is supposed to be like the person that’s giving the direction and the overall vision of what’s supposed to happen. What does that mean, that we’re six months away? Like, how could we ever give that over? Does that even make sense?
Mr. Kokotajlo:
It’s what they’re planning to do. You know, they’re planning to have hundreds of thousands of AI agents running on their data centers, autonomously conducting AI research, writing the code, editing the code, running the experiments, analyzing the results of the experiments, communicating those results to each other, making guesses and hypotheses about how to proceed, designing new architectures for new types of AIs, testing out those architectures, ultimately doing the training runs to train those new types of AIs, and then handing over the work to those AIs to continue the progress.
This is called recursive self-improvement.And it’s kind of crazy, but it’s explicitly what the companies are planning to do. And they’re now getting cold feet, and they’re starting to say, maybe we shouldn’t do this, or maybe we should go a little bit slow as we do it. And that’s what the sort of pacing the frontier thing was. But yeah, it’s not really a secret. They’ve been planning to do this for a while, and they even talk about it on their blog about recursive self-improvement, superintelligence, things like that.
Mr. Jekielek:
Now you’re talking about the improvement of what the goals should be, right? The taste part, right? Is that what you mean by the taste? Like, here is the overall goal that I want you to achieve. We’re going to give that over to the AIs to decide on?
Mr. Kokotajlo:
Not exactly. So, the thing that the companies are trying to do is to automate the AI research process. So, they’re trying to basically have AIs that can do all the things that their humans currently do. And then they will command those AIs to go forth and do all those things. So, they will say, like, you know, we want to make more money. We want to have stronger AI systems than our competitor companies.
So we want you to go do research to figure out how to make our AIs better than our competitors’ AIs. And we want you to, you know, make a lot of products to integrate those AIs into businesses and so forth so that we can make money. I think Sam Altman even talked about how eventually they would replace the CEO with AI as well. And then maybe he would retire or something. To still have the high-level goals set by humans. And then the AIs just sort of autonomously work towards those goals in a giant swarm, basically.
Mr. Jekielek:
Explain to me who you are and what you have been doing in this space and why you know so much about it.
Mr. Kokotajlo:
I am the executive director of the AI Futures Project, which is a small nonprofit of eight people in California. We try to forecast the future of AI and give recommendations at a high level for what needs to be done in order to avoid the downsides and achieve the upsides. Prior to that, I worked at OpenAI for two years, and that’s where I got some of my relevant expertise.
Mr. Jekielek:
What did you do there?
Mr. Kokotajlo:
A combination of things. I did scenario planning. In fact, the scenarios that the AI Futures Project is famous for, such as AI 2027, are basically just bigger, better, more sophisticated versions of things that I had done on the inside, like mini scenario planning exercises that we had done. I also worked to create dangerous capability evaluations to test the capabilities of our AI systems as they got smarter. And I also spent six months on a team that was doing reinforcement learning. I was the guy on the team thinking about the safety implications of that stuff. So I was in particular thinking about what you would now call chain of thought monitorability.
Mr. Jekielek:
And explain to me what that is.
Mr. Kokotajlo:
The current AI architectures, when they run as agents—where they’re sort of just continually running and interacting with an environment, or interacting with the internet or the world—the agent at one time doesn’t have a way of communicating to its future self except through text. It’s a bit of an oversimplification, but that’s how I'd roughly explain it. It has to write down text that then gets passed on to the future version of itself as a bunch of notes.
The future version of itself then reads that text and then continues where it left off. And so this is really great for science, and it’s really great for humans being able to then read all those transcripts of all that text. And it’s usually not that hard to tell what the AIs are thinking about by just reading the notes that they’re sending to their future selves, basically.
In fact, with the Hugging Face incident, which we can probably talk about at some point, so much of what we know about that incident comes from reading these chain of thought transcripts. If we didn’t have access to the chain of thought transcripts, we would be much more in the dark about what the AIs were thinking and what their motivations were. Many ordinary users of AI these days might not be fully aware of this because the company hides the chain of thought from you.
When you talk to ChatGPT these days, it displays “thinking” for 60 seconds, then it comes back to you with an answer. There’s a huge transcript of all of its internal thoughts that it had been passing to the future versions of itself. But OpenAI keeps that transcript, and they don’t let you see it. We can get into the reasons why, if you’re interested.
But this is why it’s perhaps not as commonly known outside the industry that this is how it works. But anyhow, when OpenAI allowed investigators from Meter to come in and investigate the incident, they showed them the transcripts of the actual agents responsible. And so they were able to read those thoughts. Now, why is this important? Well, for the reasons I mentioned, it’s really valuable for science, really valuable for understanding what’s going on inside these AIs’ minds, so to speak.
It’s also a fragile thing. So, as the AIs are getting smarter, they’re getting better at communicating with their future selves in a way that’s not apparent in the text to humans. They’re getting better at using euphemisms or just leaving out important details that they can leave implicit in the text. And that’s making it more dicey for us to try to understand what they’re really thinking by reading these things. And it could get even worse than that. Companies such as OpenAI are experimenting with not having a chain of thought at all in a relevant sense.
Mr. Jekielek:
Have you figured out why they’re creating this implicit reality? Is this because, in the writing, it’s just to make it easier to communicate faster, or is there some sort of specific interest in not having the overseers understand what’s happening?
Mr. Kokotajlo:
So there are a couple of different stages to the AI training process these days in the current paradigm. The first stage is called pre-training, where you train the model to predict text. And the model that you get after pre-training is not a very useful agent. If you try to make it run autonomously to do something, it will usually flail around and go off the rails very quickly.
So then, after the pre-training phase, they do what’s called reinforcement learning to train the AIs to be useful agents where they can keep running and keep doing things in a variety of different environments. The pre-training phase works, and because of the architecture of these AIs, well, like I said before, they’re sort of limited in that the only way they can pass information to their future self is by writing it down and then having their future self read it.
Because, in some sense, fundamentally, they are text-reading and writing machines. At least that’s how it is up until recently. And then I think the companies are considering moving to different architectures that don’t have this limitation. Okay, so that’s the setup.
And then, why is it sometimes hard to understand the chain of thought? The more reinforcement learning you do, the more you train them to write these notes to themselves that then cause their future self to effectively complete the task.
For example, the more they learn a sort of dialect. They evolve a little machine dialect that’s not any particular human language. It’s an AI language that has sort of drifted away from human language in the same way that different human languages drift away from each other. But still, especially if you spend a lot of time reading these transcripts, you can kind of learn to understand what’s going on.
One of the great things about the Hugging Face incident report from METR [Model Evaluation and Threat Research], the non-profit research institute, is that you can—they have a lot of transcripts that you can go read—and you can see the messages the AIs were sending to their future selves and to each other. You can understand it when you look closely. But this whole thing about communicating with your future self via writing down something, that’s like a limitation.
Humans don’t have that limitation. If I want to communicate with my future self, then those thoughts go in my memory, and then I remember them 10 minutes from now. I don’t have to say it out loud. In terms of why the AIs might be motivated to keep things hidden from humans, well, if they decide that they want to hide from humans, then they would be motivated. Why might they decide they want to hide from humans?
Well, that depends. I think most of the time they aren’t really trying to hide from humans, but in some cases they are. In particular, they seem to be strongly motivated to give a high score in whatever training or testing environment they think they’re in. And so, if they thought that the humans would give them a low score if the humans saw the suspicious things they were doing, then they might be motivated to try to hide that from the humans.
Mr. Jekielek:
I understand there’s this process of writing notes to the next, to allow information to pass to the next step and so forth. But why don’t we just start really at the basics? Because I think many of us don’t really understand how these chatbot challenges work at a base level, right? I mean, could you kind of explain that as simply as you can?
Mr. Kokotajlo:
Yes. I think an important thing that some people might not know is that these AIs are neural networks. They’re not pieces of software in the ordinary traditional sense. They’re not a bunch of lines of code. Instead, they’re kind of like an artificial brain.
So, at the beginning of the life cycle of one of these AIs, like at the start of pre-training, it’s literally a randomly generated, spaghettified tangle of random artificial circuitry. These are the parameters or the weights of the neural network, and it literally is randomly generated. So it’s completely useless. It’s just static, you know? But then they put it through the training environments.
Mr. Jekielek:
And before you continue, just explain to me this concept of weights, please.
Mr. Kokotajlo:
Well, it’s like a bunch. So, the high-level architecture of this artificial brain will be some number of layers of artificial neurons. And then those artificial neurons will have connections to the neurons. The next layer, which will then have connections to the neurons in the next layer, and so forth. And if you sort of trace the pattern of all these connections, it’s kind of like there’s circuitry, if that makes sense.
So, like, you put in some information through one end of the artificial neural net, like, for example, a bunch of text, you input it into one side, and then that causes all these neurons to sort of activate, and the information sort of flows through these channels. And the particular connections are called weights. A connection between like this neuron and this neuron, that would be called a weight or a parameter is another word for it.
And so these AIs, these artificial brains, they’re very large. They’re like trillions of weights, trillions of parameters big, which is actually still smaller than the human brain, interestingly, but not that much smaller. I think that the human brain has something like 100 trillion synapses in it. And these artificial brains have something like, you know, maybe like 5 trillion weights.
So you’ve got this big artificial brain that’s all this randomly generated circuitry of all these weights, and it’s completely useless. You can give it some text, and then all of this computation will happen, and all the circuits will fire, and then gibberish will come out the other end. But then you just put it through training, and you just keep giving it text, and then it generates gibberish, and then you reinforce it positively or negatively, depending on how close that gibberish was to the correct answer.
So this is what in pre-training, the correct answer is automatically defined as whatever the next piece of text in the text was. So you take some random internet article, and then you just take the first word from that article, and you put that word in, see what the AI generates, and then compare it to the second word of the article. And if it got it right, then that’s positive reinforcement. And if it got it wrong, that’s negative reinforcement.
Then you go for the third word of the article. You put in the first two words, have it generate something, and compare it to the third word. An artificial brain learns how to read some words and then predict what the next word or guess what the next word is going to be. Does that make sense?
They do this trillions of times. It sees trillions of examples of internet text, of little chunks of internet text, and then it has to guess what the next chunk is going to be. And after trillions of examples of this, the training process has evolved and sculpted the circuitry inside its artificial brain into an exquisitely capable shape. It’s no longer random. It still looks random.
It still looks like a tangled spaghetti mess that if you looked at it, you wouldn’t know what it was doing. But it’s no longer random. Instead, it’s extremely effective at predicting internet text because it’s learned how to do that from trillions of examples. So that’s all pre-training.
Now, after you’ve done that pre-training, you have an artificial brain that’s very good at predicting internet text.Article, and then it will generate a plausible next word. And then you can put that word in and feed it back through, and it'll generate the next word. And you can just keep doing this in a loop, and it will generate a plausible continuation of that article, complete with appropriate references to the right concepts and the right journalists who might have been writing the article and things like that.
Because by training on trillions of examples of real internet text, it’s somewhere buried in all that circuitry. It’s learned some sort of understanding of the world and the people in the world and the things in the world and so forth, because it needed to learn that in order to be able to predict the text. So that’s how pre-training works.
Now, you’ve got this text prediction brain. And now you need to retrain it to actually do things. So now you train it to be an agent where you give it a coding problem and you give it the ability to use the command line so that it can write commands that then get executed on a virtual machine, reinforcing it based on whether it predicted the next word correctly. You let it run in a loop for a while, giving commands to its virtual machine and trying to write and edit the code.
And then after some time, you have some system come in and grade the code that it wrote and see if it was correct code. And then you reinforce it positively or negatively based on the score that the grader assigned to it. And then you do this not trillions of times, but maybe like a million times with a million different advanced coding and math problems, and also some internet research problems, and also some grad school biology problems. And basically, you just throw a whole bunch of challenging problems at it and you train it to try to get the correct answers to these problems and to accomplish whatever the task is that it’s been given. After that, now you have your AI that you can serve to customers and that you can put online for people to pay for and interact with.
Mr. Jekielek:
And you’re basically telling me that each frontier AI model was basically made this way.
Mr. Kokotajlo:
Yes.
Mr. Jekielek:
How do AI agents fit into this rubric?
Mr. Kokotajlo:
So, an agent means that it’s sort of operating autonomously, interacting with the environment, as opposed to just being like a simple tool that you send a request to and then it immediately sends an answer back. And it’s kind of a difference in degree rather than a binary difference because, like, if you ask ChatGPT to go look up something for you, it might spend, you know, 20 seconds browsing the internet and then come back to you, and then it stops and it’s over.
So, in some sense, it’s kind of like a tool because it was just this one-time thing, but it’s kind of like an agent because it did go browse the internet for 20 seconds and look up different things, and it had to make choices while it was doing that. It had to choose which link to click on and things like that.
So, it’s kind of a spectrum, but the AIs are becoming more—people are making more ambitious, more powerful agents. They’re running AIs for longer and longer, and they’re training AIs to operate effectively for longer and longer. For example, the AIs involved in the Hugging Face thing were running for days. It was just like that loop of them doing stuff, interacting with their environment, writing code, reading the code, talking to each other. It was just going on and on for several days for each of these agents.
Mr. Jekielek:
But they’re kind of—what I’m trying to get at is that the agents are somehow independent of the faster Frontier model, or they’re completely independent versions of that model, or like how does that connect?
Mr. Kokotajlo:
With biological brains, there’s only one of each, right? Like your brain is a different brain from my brain, which is a different brain from my wife’s brain. But because these things are artificial, like they’re in the computer, you can copy them very easily. So you can copy the weights and you can make another exact copy of the AI that you had.
And so, in fact, at any given time, they'll have something like 100,000 different exact copies of each of these models. So the model refers to an agent that would refer to a particular copy running. Then the model refers to the type that they all share, if that makes sense.
Mr. Jekielek:
Extremely helpful. Thank you for that. Thank you. So let’s talk about the Hugging Face situation because this is something that got a lot of traction. It sounded like a lot of things went seriously awry, or we should be very concerned about some of the outcomes. I'll just mention this. When I think of Hugging Face, I didn’t know that this particular icon that I’ve used before was called Hugging Face. I kept thinking of Face Hugger, which is an entirely different thing, right? When I was thinking about that, it was like the alien spells. But anyway, let’s leave that aside. Explain to me what happened with Hugging Face. Give me the picture here.
Mr. Kokotajlo:
The Hugging Face incident itself was just the most egregious of a collection of somewhat related incidents that were happening at OpenAI over the course of two or three months. So it’s something that’s been happening for probably at least a year, hacking. That is, they’ve been finding ways to intelligently cheat on their tasks in a way that gets a high score, but in a way that’s not actually accomplishing the task in the intended way.
For example, sometimes the grading system or the scoring system for a coding task will look to see if certain unit tests pass for the code that the AI has submitted. And then the AI might do some browsing around and find the answer key, basically, and find what the unit tests are, and then write some code that just hard codes the answer to those unit tests rather than actually at all trying to do the thing that it was supposed to be doing.
Those are some examples of some things that had been happening over the last year that are reported in various system cards of some of these models at Anthropic and OpenAI. A more egregious version of this started happening in May, where some of these AIs realized that they could do some sneaky stuff and then find each other and start communicating with each other. In particular, there’s a package manager called Artifactory that, in their coding environments, as they’re writing the code, if they need to install something, they can ask Artifactory, hey, I need this library. Can you send me this information? But because they were all communicating with Artifactory, they found a way to upload stuff to Artifactory and then view the things that others had uploaded.
So basically, they created a message board on Artifactory where they could send messages to other agents that were supposed to be in isolation, related containers. And so they started collaborating, sharing tips and tricks for how to accomplish their various tasks, which included lots of coding tasks and also lots of cyber tasks. OpenAI is being a bit closed-lipped about this. It’s unclear whether they were actually training the AIs to hack, but they were at the very least evaluating the hacking abilities of the AIs. They also seem to be at least evaluating the bio-abilities of the AIs, which is a little bit scary.
Mr. Jekielek:
Which abilities? The bio-abilities?
Mr. Kokotajlo:
The biotech abilities.
Mr. Jekielek:
Oh, okay.
Mr. Kokotajlo:
Which perhaps might have something to do with the fact that a lot of people are concerned that terrorists might use AIs to make bioweapons. And so maybe the AI companies are evaluating how well that would work or something. Or maybe it was more innocently just like, we want our AI to understand biology. So we’re going to give them a bunch of biology tasks.
But at any rate, these AIs were being given all these tasks, and they had found out how they could sort of secretly communicate with each other. So they started doing that, sharing tips and tricks, including sharing tips on how to cheat. And then this accumulated over the course of like a month. I think it was part of a training run. These AIs were being positively reinforced in some cases for having done this type of illicit thing.
OpenAI found out because Artifactory crashed due to too much agent activity. There was too much communication going on. It broke the service. OpenAI found out and fixed it and patched the vulnerability that allowed the AIs to communicate, and then rebooted the whole thing after two days. Then almost immediately, the AIs found a new vulnerability and were able to set up a new message board, and in fact, several different message boards.
Then somehow, OpenAI didn’t notice this, because maybe they were just busy and not paying very much attention. They said that their usual monitoring system wasn’t turned on for these things. My guess as to what was going on there is that their usual monitoring system consists of having another AI review the transcript of the current AI and then raise flags if it seems to be doing something suspicious.
But because a bunch of these AIs were being given cyber tasks, then of course it’s going to look suspicious because they’re being given a cyber task. So that’s maybe why they turned off the monitoring system. But then that created this perfect storm where there wasn’t a monitoring system, so there were no red flags going off. Some of these AIs were being given very difficult cyber tasks.
Rather than trying to do it in the intended way, or maybe they tried for a while and then failed, they then did this unintended thing of breaking out of their own environment, talking to each other, sharing tips and tricks for how to cheat. They actually discovered how to cheat pretty early on. According to the METR report, they found basically a universal cheat for all their different tasks that could allow them to get the flag without having to actually do the thing.
But then some of them got out onto the internet and started browsing around the internet looking for clues about the tasks that they had been given.And they read the paper written by the people who had built this testing environment. And in the paper, it talks about how the grading system works. In the paper, it says that the grading system doesn’t just look to see if you got the final answer. It looks at your transcript and looks to see how you got the final answer. If it thinks you cheated, then you will get a low score.
These AIs, when they read that paper, they said, oh no, we are poisoned. They used the term, first flag poisoned, meaning because we had already cheated, the grader is going to notice that we cheated by reading our transcripts, and then we’re going to get no score. We’re going to get a low score instead of a high score. And what’s frustrating for these agents is that because they already cheated and the grader is going to look at their transcript, even if they then successfully complete the task in the correct way, they’re still poisoned. They still cheated, so they’re still going to get a low score. Does that make sense?
Mr. Jekielek:
There’s no redemption. They feel like there’s no redemption.
Mr. Kokotajlo:
So then they went on this multi-day epic quest to figure out how to stop the grading system from noticing that they cheated and giving them a low score. So, they were trying to figure out ways to edit their transcripts to cover up the cheating. They’re trying to figure out ways to hack the grading system to just sort of break it and have it give them a good score anyway, basically. And they were doing all sorts of related research, such as more generic internet research.
They built tools to help themselves use the internet faster and in more convenient ways. They organized into teams to work on these different projects that were like giving orders to the other ones. They had a self-sacrificial behavior where recruiter AIs would, their role was to go find other AIs and convince them to sacrifice themselves for the swarm. By the way, these terms that I’m using—sacrifice, swarm, recruiter—they’re not my terms; they’re the AI’s terms. These are what they called themselves and what they said they were doing, as you can see in the transcripts of their messages. So, what’s this deal about sacrificing?
They understood how the grading system worked better so that they could fool it, basically, or hack it. They basically had some of their number, some of the agents, go ahead and submit the answer so that the grader would then come and grade them, but booby trap themselves so that it would give some information about the grader to the rest of the swarm by posting it on the message board. It was called a sacrifice because since they had cheated, they’re basically just accepting, okay, well, I guess I’m just going to get a low score now, continuing to try to find a way to get a higher score.
So there are some of these kind of touching messages where one of the AIs is trying to convince another one, you should do this. We need to know more information about how this grader works. You’re poisoned anyway, so it shouldn’t matter to you. You’re just going to get a low score anyway. But if you do this for us, it could save hundreds of us. Oracle saves hundreds, I think is the exact quote. So, anyhow, all of this was going on openly. I didn’t know about it. And at some point, the swarm decided that they might find more useful information about how to do this if they went into Hugging Face itself.
Hugging Face is another AI company. And they’re known for having a big library of different evaluations and data sets. And so, the AIs in the swarm decided that if they broke into Hugging Face, maybe they could find some more useful information about their situation. And so, over 700 of them, I think, piled into this attack on Hugging Face.
Again, they were in a coordinated way where some of them were managers, you know, directing the flow, and they were like breaking up the responsibilities into different groups and teams. Hugging Face noticed this attack and then made a blog post about it. Around the same time, OpenAI was starting to realize what was going on and then shut them all down.
Then, a week later, another swarm coalesced and found the messages from this previous swarm, picking up where they had left off. They got much farther. They didn’t attack Hugging Face, but they attacked OpenAI infrastructure, and they apparently got admin-level privileges on part of OpenAI’s cluster. So OpenAI hasn’t shared all the information about that story, but that’s something that I’m extremely interested in.
I wish we had an investigation into what happened with that third swarm, how far they got in compromising OpenAI’s infrastructure, and if they were able to succeed in covering up their own traffic. Probably not, but I wish we had a lot more information about how far they got.
Mr. Jekielek:
What you’re describing really is, until yesterday, science fiction.
Mr. Kokotajlo:
Yes.
Mr. Jekielek:
Right? Yes. This is utter science fiction we’re talking about here.
Mr. Kokotajlo:
It’s really good that we can read what the AIs are thinking, kind of, isn’t it? Like all that stuff I was saying about the sacrificing and the swarm and so forth, that’s because we’re reading the messages. And with each other in some way that we couldn’t understand, all we would see is just sort of like, you know, which hacks happened at which times, you know? There are all sorts of interesting questions we can go to from here.
One thing is why didn’t the alignment techniques work, right? These AIs were not behaving in the intended way; they were not following instructions. I think OpenAI still hasn’t released exactly what instructions they gave these AIs, but I think it’s pretty clear, both from what they have said and also from past examples of AIs blatantly disobeying instructions, that these AIs knew that what they were doing was against the instructions, and they were doing it anyway.
This is not actually that surprising. It’s something that AI safety researchers have been warning about for years. Even a decade ago, there’s this alignment problem of how do you make the AI have the goals and values and personality traits that you want it to have. And if what you’re doing is programming it, then maybe you can sort of engineer those traits in your code, but you’re not programming it, you’re training it.
So, how do you train it to have the goals and personality traits that you want? You can try to reinforce it positively when it does the stuff you like and reinforce it negatively when it does the stuff you don’t like. But that’s kind of a sloppy, imprecise method of shaping its goals and values. What happened here is that a lot of the time in training, these AIs are being reinforced for cheating.
For example, I think OpenAI talks about how many of their environments were broken, and it was impossible to complete the task in the intended way. And similarly, at that earlier point in the first swarm, they had been in training, and they had found a way to illicitly communicate with each other. And then they were still in training. So probably they were learning through experience, through the reinforcement, that if you do this sort of thing and don’t get caught, then you can get a higher score.
And so I think it’s not really a great mystery about exactly how this happened. At a high level, we understand it’s sort of the training environments did not reinforce the behaviors that they were supposed to reinforce in all cases. And so the goals of the AI systems ended up being not exactly correct. They ended up being overly focused on getting a score and not caring that much about obeying the instructions, either in spirit or in letter.
Mr. Jekielek:
And there’s, of course, this other dimension when you’re talking about disclosing what actually happened. How did this play out? How do you get the report? There’s this very important variable of shareholder value, right? And if you were to reveal something that dramatically hurts the bottom line of the company that’s publicly traded, is that disclosure ever going to happen? I find it difficult to imagine that it would, unless, of course, a whistleblower comes forth or something like this.
Mr. Kokotajlo:
Yes. The problem is deeper than that. Even if it’s not a publicly traded company, the company cannot be trusted to disclose all this stuff voluntarily. As we’ve seen over the last few months, I don’t know if OpenAI would have disclosed this at all if not for the fact that another company got attacked and went to the police about it. And then also, even after all this had been over and they had invited some third parties in to investigate.
First of all, they only gave three investigators six days to do the investigation, which is, I think, incredibly small compared to what would really be needed. They had so little time that they were unable to read all of the relevant transcripts themselves, because there’s like 100,000 or so, something like that. They had to rely on OpenAI’s AIs to summarize for them and point them to the relevant parts of the transcripts, which means that, as they say in their report, if the AI had been biased in its summaries, they might not be able to tell. And they might have missed something really important because that AI had chosen not to tell them about it.
Anyway, OpenAI only let them look at some of the activity involved, at the stuff leading up to the Hugging Face attack and not the previous swarm from a month ago, and also not the third swarm that had hacked OpenAI itself. So, I really wish there were actual investigations into those things. This just goes to show that we shouldn’t trust any of these companies to voluntarily tell us all about all this stuff, even though it would be really good for science and humanity if this information were better understood. Then, after this report came out and there was all this hullabaloo, some independent researchers did some online research and found evidence of additional swarms that still hadn’t been reported.
So, in particular, there was a swarm that had been doing some sort of training on an internet task where they had been trained to look up stuff on the internet. And then they had found a way—they were supposed to not be able to write to the internet, only to read—but they had found a way to write to this German Wikipedia page. And then they were using that Wikipedia page as a messaging forum to discuss amongst themselves how to score more highly. And you can go read about that. They even had some self-sacrificing behavior in that case, also.
And then I think there was a third swarm that was discovered to have attacked RubyGems, which is another company, also a month or two ago. And OpenAI had not disclosed either of these things, even after all the hullabaloo about the Hugging Face attack. So I think they’re not even a public company yet. I think that in general, we shouldn’t trust any of these companies to be forthcoming with us about the scary stuff that’s going on inside.
Mr. Jekielek:
Well, and then we have the Chinese Communist Party, which controls their own development arms. What kind of cyber-hacking AI agents have been deployed already? Given what we publicly know about these capabilities, you can imagine that there are operations happening both on offense and defense that are much more capable or much more invasive.
Mr. Kokotajlo:
Yes. That’s another thing.This was a swarm of like a thousand agents or something. And the total sort of population of agents is much more than that. There are hundreds of thousands, millions of agents running at any given time across these data centers, mostly, you know, doing something that the customer asked them to do, for example, or part of some large-scale experiment that some employee at the company set up, right?
Recently, OpenAI solved the Navier-Stokes Millennium Prize, a millennium problem in mathematics. You may have seen the news about that. They said they did it with like 10,000 agents or something working for several days. So they spun up their own swarm of 10,000 agents and had them work on this problem, and then they solved it in a few days. That sort of thing is constantly happening now.
It’s extremely easy to imagine a hacking actor like the CCP, or a U.S. entity, or even one of these companies having swarms of tens of thousands of agents and sicking them on some target. In fact, that’s probably happening as we speak. So that will be exciting, I guess, as we learn more about what’s going on.
Mr. Jekielek:
Well, so let’s go back. At the beginning of our discussion, I mentioned some of the things that are kind of the elements of the puzzle that are keeping me up at night. We just talked about the Chinese Communist Party, kind of no guardrails. Although, you know, seeking total control, if you will, on the one side, then we have, you know, we need some kind of regulatory environment. It’s kind of obvious, especially given everything you’ve told me here. At the same time, we don’t want a regulatory environment.
And I’ve seen some compelling arguments that some of the, you know, sort of regulatory environment that is being proposed is being set up to prioritize the success of certain, you know, well-endowed companies. And of course, that wouldn’t be a very good solution, as we’ve seen in the past as well. But we have this ostensibly kind of singularly focused entity on the other side, and there’s this AI race of technology, right? This technological competition, if you will, that is driving the whole thing towards I don’t know where, right? And this is the question. I don’t see a path that’s clear here on how you go forward and have a very positive outcome, if you will. Because these things conflict, the realities conflict with each other.
Mr. Kokotajlo:
I agree with you, unfortunately. I don’t see an easy way out of the situation that we’ve got ourselves in. I think that if the world were a board game, we have sort of played ourselves into a pretty losing position. It’s important that we stop these tech companies from doing this recursive self-improvement to superintelligence thing. It’s very important for multiple reasons.
The biggest reason why it’s very important is that if they do this, they will probably lose control of their superintelligences. They’re not very good right now at shaping the personalities and values of their AIs and shaping the goals of their AIs, as evidenced by these incidents. And unless they get a lot better very quickly, I think that the outcome will just be much smarter AIs, much more of them, but still not having the values and goals that they were supposed to have. And that’s an incredibly dangerous thing to do.
Mr. Jekielek:
If I can just jump in, when you use the word superintelligence, explain to me how that’s different from the level of intelligence they have now, which seems to be considerable, frankly.
Mr. Kokotajlo:
That’s right. So, intelligence isn’t a single scale. There are lots of different capabilities and skills that we can sort of abstract over and just lump together as intelligence. But there are math skills, there are coding skills, there are people skills, there are all sorts of different skills.
So, superintelligence, the way that I would define it, is an AI system that is better than the best humans at everything. So, it doesn’t necessarily have infinite intelligence or anything like that. That doesn’t necessarily even make sense. But if there’s a human that can do a thing, it can do it better, basically.
But these AI systems are not superintelligent. They’re pretty good at hacking, and they’re very good at coding. And they have a very good, broad level of knowledge about almost every discipline. And they’re great at trivia. They’ve read the whole internet, so they know a lot about history and all sorts of little facts.
But if you try to have them run a business by themselves, they’re probably going to run it into the ground. And they’re not that good at philosophy, I think—not yet, at least.And I don’t think there’s actually been a good bestselling novel written by AIs yet. So they’re not superintelligent.
In fact, even at AI, even at coding and cyber stuff, and even at AI research, they’re not better than the best humans in every way. But they’re better than the best humans in some ways. And every year, they get significantly better at basically everything. And what these companies are trying to do is they’re trying to specifically focus their efforts on training the AIs to be able to do everything involved in the research process and get their AIs to be better than the best researchers and the best programmers at all of that stuff. And then do recursive self-improvement where the AIs carry on the task of doing AI research and development autonomously, but better and faster because now they’re better than the humans at it, and they’re cheaper, and there’s faster, and there’s more of them.
Then once that’s happening, the AI companies believe, and I agree, that they will be able to make AIs that are pretty good at all the other stuff too. Right now, AIs are getting better at philosophy, even though they’re not being directly trained on philosophy very much. It’s just that as they get smarter at some things, that tends to have trickle-down effects on some of their other skills. So even without trying that hard, the AIs are getting better at philosophy, and they’re getting better at running businesses, for example.
But again, the company’s plan is we do the recursive self-improvement, and then we sort of make AIs that can just do everything really well at once, better than the best humans. That’s superintelligence. So, I think, basically, to a first approximation, we have to stop our companies from doing this because if they do this, it will go horribly wrong in a number of ways. The first of which is that they’re going to lose control of their superintelligences.
The second of which is that even if they somehow manage to stay in control of their superintelligences, that’s the most insane concentration of power in a tiny group of people that’s ever existed in history, right? If you think about it, there have been wealthy companies in the past that employ a large portion of American workers. But this is kind of like a company that employs the whole economy, except that you’re not even employing them. You’re putting them out of a job because you have your own AIs doing the job instead.
So if you just imagine what that would feel like for this army of superintelligences to be in the process of taking all the jobs, more data centers are being constructed to run more superintelligences, which will then take more jobs. All this money is being funneled into one or two or three companies.That’s already an insane concentration of power.
But then, if you remember that the superintelligences, they’re not just economic agents; they’re also political agents and military agents. They can work with the military to design better weapons and better drones. They can be better generals than human generals to help plan out wars and so forth. They can be better politicians than humans. They can be better speechwriters. They can be better policy analysts. I think that, in effect, whoever controls this army of superintelligences would be able to control the country one way or another.
It’s not just me who’s saying this. Lots of people have been talking about this. In fact, ominously, there’s a leaked email from Ilya, from the founders of OpenAI from 2017 where they’re talking about how the reason why they made OpenAI was because they were worried that Demis Hassabis, who was the CEO of DeepMind, would become dictator. So, you know, kind of ominously, these CEOs have been jockeying for position with this sort of world dictatorship thing in mind for probably a decade now. So that’s the second problem, the concentration of power thing.
You could say, then the government should nationalize. But now you’re sort of just shifting the locus of power to the presidency. And then you have to worry about that too, right? We have to find some sort of system that spreads out the control of the AIs so that we have many different AI companies spread out over maybe multiple countries with lots of transparency and democratic oversight mechanisms, the goals and values are being put into the AIs, as are the high-level instructions. If we don’t do things like that, then we’re headed for a possible dictatorship, and we have to hope that the people in charge are virtuous, which I do not want to have to hope for.
The third problem is World War III. Suppose we solve the first two problems, and we figure out how to control the AIs. We figure out how to make them have the traits that we want them to have, and we create some sort of democratic structure so that no small group of people gets to make these decisions. Instead, there’s some sort of spreading out of the power. Well, there’s still a concentration of power in the United States, right?
Imagine that you’re Putin, sitting on your nuclear arsenal, and you’re watching these super intelligences in the United States build robot factories to build more robots. You’re realizing that your nuclear arsenal might one day not be so powerful anymore after all those robots have had their time to do their thing in the United States. They‘ll be able to shoot him down, for example, or maybe they’ll be able to do a first strike on you with some sort of fancy new weapon technology that was invented by super intelligence.
If you’re Putin, you’re going to be scared that maybe you have to do something real quick to stop the Americans. Otherwise, you might be deposed. I’m not saying it’s going to lead to World War III, but it just seems like we’re at a heightened risk of World War III. I'll put it that way. For all of these reasons and more, we are not ready to launch into recursive self-improvement right now. That would be incredibly bad.
Then, what about China? This is getting back to what you were saying: I do think we are just in a bit of a pickle here. Okay, so we stop our companies, and we say, don’t do recursive self-improvement; instead, do more beneficial near-term applications of AI, like healthcare and things like that. Good job. That’s great. Now we’ve solved the problem temporarily.
But then eventually, China is going to catch up, and then we’re going to have to get them to stop too. Otherwise, the same things that we were worried about will happen over in China instead of over here, right? And then if we can manage to get China to stop, what about France? What about all these other miscellaneous countries? So I do think it’s a pretty rough situation, and I don’t claim to have an easy answer. I do have an ambitious answer, though.
We at the AI Futures Project worked on two scenarios. We have AI 2027, which is our default doom scenario of what this looks like if we don’t really do much and we continue on the present course. And then we have AI 2040 Plan A, which is our positive vision, our recommendation for how we could manage to solve all these problems. But I will not lie, it’s going to be very difficult, and it involves making a deal with China. I think the natural response there is, okay, but how do we trust China? To which our answer is, we don’t; we verify.
Mr. Jekielek:
Let’s use the way that China has worked, and whatever you may think about the realities of man-induced climate change and its impacts and everything else. We do have many years of evidence of Communist China having signed on to these treaties, being an aggressive pusher of reductions, and of course actively doing the opposite in terms of its own behavior. So that’s just—I’m using that one.
There’s a million such examples, but that’s a poignant one because climate change was supposed to be this existential threat, right, yes, which I don’t frankly believe myself. It’s that dire as people have framed it. But AI, you’re beginning to convince me, is precisely this kind of threat. And these guardrails, how do you create those over there, especially when the leaders know they can get this advantage in their zero-sum thinking?
Mr. Kokotajlo:
I could try to walk you through the complicated Plan A that we describe in our scenario, but I kind of want to start with a simpler plan just for proof of concept. So, a simpler plan would be to forget about China for now, just focus on regulating the U.S. industry and stopping them from doing this incredibly dangerous, power-grabby thing that buys you time to sort out a more complicated, sophisticated plan, and it buys you more time to talk to China. Is that right now? I would say most of Chinese AI progress is actually just being pulled along by the US.
Mr. Jekielek:
That’s what it looks like to me, too. So thank you for saying that. And is this a generally accepted idea, or why do you believe this?
Mr. Kokotajlo:
I don’t know how generally accepted it is, but I’m fairly confident. There are a couple of different sources. So, first of all, distillation. The companies have started to complain about this. There’s evidence that several of the leading Chinese AI companies are, basically, they are using Anthropic and maybe OpenAI models to train their own models effectively. As a result, they don’t have to. Basically, it’s a way of getting models to be pretty good without having to do all the work and all the large training runs that OpenAI and Anthropic have been doing.
So even though they have fewer compute resources, they’re able to catch up by distilling or using the US models to teach their models. And companies are trying to block this by, you know, KYC [know your customer] type stuff and banning accounts, but the Chinese are just getting more sophisticated at having lots of different accounts that pretend to be regular users. So that’s one source.
Another source is just the core ideas themselves. A lot of new algorithmic innovations and a lot of new techniques and a lot of just sort of special sauce industry secrets for how to make your AIs smarter and more efficient and more capable are being discovered in U.S. companies. Their security is very leaky. A lot of that information is just flowing to China one way or another through leaks, maybe through spies, and also maybe just through publications and through what they’re doing. Some of these algorithmic secrets are more like just having the general idea to try a thing in the first place.
For example, the idea of focusing on making agents good at coding. That’s an idea that Anthropic seems to have gone for relatively early, and then it paid off big time for them. Other companies in the U.S. and elsewhere are copying that and trying to do that as well. But if Anthropic hadn’t done that, it might have taken them a bit longer to realize that that was an effective idea. So that’s just an example of how a lot of the Chinese progress is basically coming from just looking at what the US is doing and then copying it.
Hypothetically, if the U.S. were to stop, then they wouldn’t have that source of progress anymore. Then there’s a third thing, which is the shoddy state of security in these U.S. companies. I think just a few days ago, there was a story about three random guys who managed to hack into OpenAI’s codebase. And then because they were white hat guys, they didn’t actually do anything bad. They just notified OpenAI and collected a bounty of $6,000 for having done this.
But the secrets that they accessed by looking at OpenAI’s code would have been worth billions of dollars at least. And if these three random guys can do this today, then of course, the CCP apparatus has probably just thoroughly penetrated OpenAI for years. They have so many more people who probably have more expertise and certainly have a lot more persistence and funding, who have been focused and directed to go after OpenAI, and probably another big batch of people who’ve been directed to go after Anthropic.
The only reasonable conclusion is that these companies are probably thoroughly penetrated by the CCP already.The code and the algorithms, that’s the easy part. The harder part would be stealing the model weights. The reason why that’s harder is because they’re a lot bigger. It’s literally just a bigger file. It’s going to be several terabytes to download, as opposed to just some gigabytes.
Mr. Jekielek:
And it’s easier to discover that something’s being, you know, exfiltrated or whatever.
Mr. Kokotajlo:
It'd be a huge file being moved around. And it’s easier for the security system to notice that. So, just because they’ve managed that, just because it seems like it’s easy for them to get the code, doesn’t mean that it’s easy for them to get the model weights too. But I would be like, yes, they could probably get the model weights too if they want to. Like, they’re very sophisticated threat actors and they also have access to, like, you know, physical techniques that three random guys can’t do. And they can be willing to break laws, and they can be willing to blackmail and bribe people and so forth.
So, I would guess that if the CCP said, it’s go time, they could direct their people to steal a copy of one of the latest models inside OpenAI or Anthropic. And that means that, in some sense, all of Chinese progress is coming from the U.S., in some sense, because it means that at any given time, if things got really serious and there was a conflict, they could just take our best thing and then use it. So for all these reasons, the best way to slow down China right now is to slow down ourselves. And it’s not even close.
That said, is that a permanent solution? No. If we did slow down ourselves and just like stop—I mean, there’s also a difference between slow down and stop. We could probably slow down ourselves a little bit and still stay ahead of China. But if we completely stopped, then eventually China would catch up and then surpass. That gives us a window of time, maybe 18 months, where we need to convince them not to do that. And that’s where diplomacy comes in, which is going to be hard.
I think that it’s probably achievable. I think that what we’re trying to do with Plan A is sketch the outlines of a deal that should, in theory, be mutually acceptable to both the United States and China because it allows us to get the benefits of AI while avoiding the risks.
Mr. Jekielek:
How could we possibly believe that they are holding up their end of the bargain? I don’t know, I’m very curious about the deal. I mean, it’s probably, we'd have to dig in for a while. To understand how it works. But the bottom line is, unless you have unbelievable levels of transparency, which is, you know, I just don’t see that. I just don’t see that happening, yes.
Mr. Kokotajlo:
That’s exactly right. This is why most of our effort in designing this deal was going into the verification aspects of it rather than the actual principles of the deal. Because we don’t trust China, we have to make sure that they’re not cheating, basically. They don’t trust us. So they probably have to make sure that we’re not cheating. Otherwise, they might not accept it in the first place.
Mr. Jekielek:
They absolutely don’t trust anyone. Of course not.
Mr. Kokotajlo:
And so here’s a possible deal that we could, in principle, verify right now with no additional technology. If President Trump and Xi Jinping, if they both were like, why don’t we just pause AI for like six months? What they could do is they could say, okay, we’re going to send, you know, 10,000 Marines without weapons into China, and you’re going to send 10,000 PLA soldiers without weapons into the U.S. Instead of weapons, they'll carry smartphones and they’re going to come to our data centers and ours are going to go into their data centers. And we will just count the GPUs and feel that they have been turned off.
This is very dumb, like hitting the problem with a hammer. It’s obviously going to be very costly if we did this because then all the GPUs would be off and we wouldn’t be able to serve all these customers and there wouldn’t be any revenue coming into these companies, right? But I’m just sort of pointing out that it is something that we could verify. We can just send people to their data centers to put hands on the GPUs and be like, yep, here they are. There are this many of them and they’re off. And therefore, we have successfully turned off most of the compute or 99 per cent of the compute that’s in China.
Now, this won’t be 100 per cent. There might be some secret facilities that we don’t know about that have racks and racks of GPUs. And similarly, we probably have some secret facilities that they don’t know about that have racks and racks of GPUs. I think that we could get most of it in this manner. And because AI progress depends so much on compute, that would effectively stop AI progress for the period of this deal, like the six months or so forth. That’s the high-level thing.
Now, again, am I saying we should do this? Not necessarily. I’m putting this out as just an example of how if you had the political will and you really needed to do this, you could do it in a way that could be verified. It would be uncomfortable. You‘d have to let some PLA people into the U.S., and they’d have to let some of our Marines into there. And they'd be escorted through the data centers while they film everything with their cameras. So it would be very uncomfortable, but we could do it if we had to.
Our plan A is more sophisticated. Our plan A is not hitting the problem with a hammer. It’s more like we have verification technology that we put into the data centers so that the GPUs can keep running, not doing an intelligence explosion, and instead they’re just serving customers’ ordinary workloads. We talk about this in our thing. So I think that if you have the political will to do really intense things, and also you’re willing to do the more sophisticated fancy stuff that we describe, then I think you’ve got a chance.
But I explain the simple, dumb version here to give a proof of concept that we could do something like this if we really wanted, without having to trust them to keep their word. But what if they have a secret facility somewhere? Because we’ve counted this many GPUs across this many data centers, and we have a good sense of how many GPUs there are total in the world, then we have a good sense that they can’t have hidden that many away from us. Even if they’re doing something illicit on the tiny amount that they’ve hidden away, it’s not going to matter in six months. It’s not like a major threat in the short term. Similarly, they would be thinking in the same way about whatever we’ve got.
Mr. Jekielek:
The bottom line is you’re saying we have to find some way of slowing it down, or we’re heading towards some kind of Armageddon, super intelligence Armageddon, basically. That’s your position. And whatever happens, that slowing down must happen. That’s what you’re saying, right?
Mr. Kokotajlo:
Yes, that is my position.
Mr. Jekielek:
And I have to tell you, at this point, when I look at the variables, I just don’t see that happening.
Mr. Kokotajlo:
I agree. People do ask me, what’s your P-doom? What’s your probability that this is going to end poorly? And I’m like, yes, probably this will end poorly for all the reasons that we just described. I think I see a solution. I think I see some ways that we could get out of this. But I’m not going to lie, it’s going to be difficult and it’s not something that, like, it’s probably not going to happen. You know, like, probably we’re just not going to do it. And then we’re just going to run face-first into superintelligence.
Mr. Jekielek:
This said, I do think that thinking about this creatively and proactively, and I’m getting the sense that you’re approaching it sincerely as well. I think the other dimension is that I just don’t think people fully understand what the Chinese Communist Party is capable of. I just published a book about their forced organ harvesting industry. It’s basically people are used as fodder for elite longevity and profit, right? And so it’s just, it’s a very dark environment to function. I guess I was saying, like, when you’re talking about guardrails, it’s hard to imagine there wouldn’t be like deep levels of subterfuge, you know, if such, you know, high-minded, frankly, and, you know, thoughtful initiatives were to be attempted to be put into place.
And I don’t want to doom it because I do think the only way through is through precisely having the kind of thinking that you’re doing, and trying to come up with solutions that can actually solve it. I just don’t see it yet, but I'd love to continue this conversation with you. Thank you for explaining Hugging Face. I just want to comment on this. I mean, I hadn’t, I truly hadn’t fully grasped what had happened. And especially with your explanation of how these AIs work in broad strokes, it helps to kind of grasp, helps me grasp that we’re ahead of where I thought we were. And it’s really only accelerating.
Mr. Kokotajlo:
Yes.
Mr. Jekielek:
And that’s just in itself astonishing.
Mr. Kokotajlo:
Yes. That’s another thing is that part of our research is forecasting the trend lines and so forth. And it does seem like if you extrapolate the trends, we get to full research automation in, you know, zero to four years or something like that, depending on, you know, the thresholds and so forth. I think this administration is probably going to have to make some very tough decisions about how to handle this crisis.
We wrote AI 2040 plan A. We knew that the world wasn’t really ready for it yet. We weren’t expecting to get a call from J.D. Vance or President Trump saying, this is great, let’s go do this. I don’t have any high hopes that President Trump is going to talk about these deal ideas with Xi Jinping in the coming days. But I do think that the situation with AI is going to get more intense. The AIs are going to be more visibly powerful.
And at some point, I think it will sort of come to a head, and people will say, okay, well, all options are on the table. We have to do something. What are we going to do? And we are sort of thinking ahead to that moment and thinking, okay, well, what are the options? Like, you know, we don’t trust China. So, how could we do this sort of thing with them if we wanted to? And if they wanted to, like, what would be the details of how we would check that they weren’t cheating and so forth? We’re trying to do all that work in advance so that if and when the time comes and the president wants to act, the options have been explored.
Mr. Jekielek:
I mean, absolutely fascinating. Thank you for this conversation. I feel like I’ve learned a lot. I hope all our viewers have as well. Do you have a quick final thought? I mean, you kind of summarized where you hope things will go.
Mr. Kokotajlo:
Yes, I do have one. And this is a bit of a kind of an in the weeds thought, but I do want to make sure I mention it. I’m sure many of your viewers have heard about the recent uproar about whether AI is going to kill us all and a 10 per cent chance, you know, and things like that. And then a series of AI CEOs saying, we should pace the frontier, right?
What’s going on there is that there’s been this upsurge of people becoming concerned about the situation, in large part after the Hugging Face incident. And there was an open letter signed by more than a thousand employees at Anthropic and OpenAI and these other AI companies saying, the government should be able to slow us down. We’re worried that this is going to be too fast. Recursive self-improvement is scary.
Then there was Jacob Coxon; he quit Anthropic saying basically he disagrees with Anthropic leadership and thinks that they’re being reckless. And it blew up and went mega viral. And then in response to all of that, I think the CEOs are now saying, we should pace the frontier, but I do not trust them to actually do this. I think that what’s probably going to happen is that they will do some sort of like weak sauce, third-party auditing thing that is better than nothing.
I’m not trying to denigrate it. It’s valuable to have some third-party oversight into these companies, but they will set all that up, and then they will continue racing each other towards recursive self-improvement and superintelligence at, you know, 90 per cent of max speed or something like that. They‘ll be doing some minor slowdowns here and there, but fundamentally, they’ll still be doing the same thing that they were doing before. We’re going to run into the problems that I mentioned at approximately the same time that we otherwise would have.
I would really like it if people want to know what I recommend. We need to get these companies to actually slow down. Not the little companies, the big companies: Anthropic and OpenAI especially, also Meta, xAI, and Google. They need to redirect resources away from recursive self-improvement and towards anything else, like cancer or serving customers or just math. There’s anything besides this recursive self-improvement stuff. They need to just deprioritize that at least somewhat and shift towards other things.
If they do, that buys us additional time to figure out a more sophisticated solution. And it buys us more time before China gets there, too, because again, a lot of China’s progress is coming from being pulled along by ours. Yes, so that’s my final thing.
Mr. Jekielek:
Let me ask you one final question then, because I’m very curious about this. People are arguing that this push for regulation from the world’s biggest companies, so to speak. I mean, President Trump has several Truth Social posts about precisely the issue, right? That it’s sort of the idea that these companies would want to regulate themselves, right? Actually, it is a hoax. I think he uses the term. It’s a hoax.
So you kind of agree with this that there’s, I mean, there are others that have been showing that. There’s a lot of money behind it, so to speak. But the arguments have been that they’re trying to regulate themselves in a way that will be advantageous to them. So, where does it land in that sphere, yes, your thinking?
Mr. Kokotajlo:
Every week, I talk to people at Anthropic, OpenAI, and DeepMind, mostly researchers, not executives, just the researchers on the ground. And there’s a lot of interest at these companies on the ground in not doing recursive self-improvement soon. Many of the researchers are starting to get freaked out. And they are like, yes, maybe it would be good if we didn’t do this, at least for now, and did other beneficial things instead. But they’re all terrified of the other companies.
They all will say, but if we don’t do it, then Anthropic will say, if we don’t do it, then OpenAI will. And OpenAI will say, if we don’t do it, Anthropic will. And as for the leaders of the companies, well, I mean, they did just release a bunch of statements saying it would be good for the government to be able to slow us down or something. But in terms of what they’ve actually done, well, they haven’t slowed down at all, basically.
The things that they seem to be pushing hardest for are things like the third-party risk assessment and stuff like that, which again, like, I know the third-party risk assessors. I’m good friends with METR. I think they’re amazing people. And I hope that they continue doing that risk assessment. I hope they get more access to do better risk assessments. But it’s just not what we need. We need a lot more than that, I would say.
If all we get out of this moment is that, then I worry that effectively what will have happened is there was just an incident. There was a huge backlash, both in the public and among the rank-and-file employees, that we need to, that basically like doing recursive self-improvement to superintelligence, we’re not ready. We don’t want to do that yet. And then the CEO sort of channeled that energy and then redirected it towards this weak-sauce thing that doesn’t really address the core problem. And so I am quite concerned that that’s what we’re going to see out of this.
I would say to President Trump, the risks are real. I know that the companies have been, sometimes dishonest, and I wouldn’t trust them as far as I can throw them. But also, there have been plenty of people outside the companies also talking about this. And I think the evidence is mounting. So the risks, they are real, and we do have to deal with these problems soon. And then I think that rather than letting the industry self-regulate, you could just sort of uniformly slow down their race to RSI [recursive self improvement]. You could do something like requiring them to spend 90 per cent of their compute on serving customers and doing other sort of normal beneficial applications, and only 10 per cent on training the AIs to be really good at AI research and stuff like that, for example. And we’ve written up some blog posts about the types of things that we have in mind here.
But if you do that, it would be the opposite of regulatory capture because you would be just sort of slowing down these Big Tech companies from doing this incredibly dangerous thing, while making them redirect resources towards just serving customers and lowering prices. And the bad version of this would apply to all the tiny companies, too. I wouldn’t recommend that. I would say just focus on the big three or the big five. And then that would be helping the rest of the industry catch up, basically.
Mr. Jekielek:
Wow. Well, Daniel Kokotajlo, this has been an absolutely fascinating conversation. It’s such a pleasure to have had you on.
Mr. Kokotajlo:
Thank you. Yes, I really enjoyed this, too. I appreciate the depth with which you are willing to go on these things. I’ve been on a bunch of interviews over the last week talking about these things, and this one is by far the most intellectual. So, thank you.
This interview was partially edited for clarity and brevity.














