The Joe Rogan Experience - September 09, 2026


Joe Rogan Experience #2551 - Daniel Kokotajlo


Episode Stats


Length

2 hours and 17 minutes

Words per minute

188.84

Word count

26,031

Sentence count

1,620


Transcript

Transcripts from "The Joe Rogan Experience" are sourced from the Knowledge Fight Interactive Search Tool. Explore them interactively here.
00:00:02.000 Joe Rogan Podcast, check it out.
00:00:04.000 The Joe Rogan Experience.
00:00:06.000 Train by day, Joe Rogan Podcast by night, all day.
00:00:09.000 Hello, Joe.
00:00:14.000 How are you?
00:00:17.000 I'm in an interesting mood today.
00:00:21.000 Why are you in an interesting mood today?
00:00:23.000 Well, I'm excited to be here and to talk with you about all this stuff.
00:00:26.000 I'm a little shaken by what's going on in AI, which is why I've come on the show.
00:00:32.000 The situation with AI is just crazy.
00:00:34.000 And I think not enough people really understand how crazy it is.
00:00:37.000 The particular event that sort of inspired me to reach out was the Hugging Face hack.
00:00:42.000 You've probably heard about that, right?
00:00:44.000 Yeah, let's explain it to people, though.
00:00:46.000 So AIs, AI agents, AI agent runs continuously in some sort of environment.
00:00:46.000 Yeah, okay.
00:00:54.000 It doesn't have to wait for you to send it a message, it just keeps doing stuff.
00:00:58.000 The AI companies are training AI agents, thousands and thousands and thousands of them.
00:01:02.000 They're making them Better at all sorts of skills, especially coding and research skills.
00:01:07.000 And way back in May of this year, some of the agents at OpenAI kind of broke out of their containers a little bit and established a message board where they could communicate with each other and share tips and tricks for how to score higher on the little tests they were being given and the various things they were being trained on.
00:01:30.000 OpenAI didn't notice this until much later.
00:01:34.000 They eventually did because the message board crashed.
00:01:39.000 The system because there's too much communication across probably thousands of agents that were communicating.
00:01:47.000 Now, OpenAI is being a little bit unclear about a lot.
00:01:50.000 They're not really sharing that much detail, so it's unclear who knew what when.
00:01:54.000 But they said that after the message board crashed, they fixed the particular exploit that allowed the agents to communicate and then booted up again and started things going again.
00:02:06.000 And then within a day or two, These agent swarms had re coalesced.
00:02:13.000 And so there were now, again, like hundreds of thousands of agents that had set up a new message board and were talking to each other on this thing.
00:02:22.000 How is it possible that they have no oversight?
00:02:26.000 Like, how is it possible that they can set up message boards and do these things completely secret?
00:02:32.000 So I would say OpenAI has something like a million ish, maybe 100,000, something in that magnitude of AI agents running.
00:02:41.000 At any given time, just internally, and then externally, serving customers, they've got more.
00:02:45.000 But they've got this huge, all of these AI agents in their little containers doing various tasks, being trained, being spawned, doing some task, being graded, and then killed.
00:02:59.000 That's too much for any human to read.
00:03:01.000 There's only like a few thousand employees at OpenAI.
00:03:03.000 They couldn't possibly look at all that activity.
00:03:06.000 So they rely on other AI monitors to sort of like look at the traffic and like point out if anything suspicious is happening.
00:03:13.000 And OpenAI said that they basically just hadn't been monitoring some of these agents, or not very well at least.
00:03:19.000 So, in particular, these particular ones that were in training, for whatever reason, the monitoring system was weak and didn't notice or wasn't activated enough.
00:03:30.000 Was the monitoring system weak because they didn't anticipate them being able to do this and break out of their containers?
00:03:36.000 Or was it complacency?
00:03:40.000 Like, what caused this to be possible?
00:03:44.000 I mean, my opinion would probably be a bit of complacency, honestly, because I think there's been plenty of evidence accumulating over the year that AIs can do things like this and sometimes do.
00:03:57.000 Yeah, so I guess to keep going with the story, the second message board was set up.
00:04:03.000 And this time, well, this was the swarm of agents that got out onto the open internet and then attacked Hugging Face, which is another AI company.
00:04:14.000 And it's really interesting to sort of study the anthropology of these AIs, so to speak, or like why they did this, because there's a lot we can say about it actually.
00:04:25.000 Basically, The companies have their goals for what they want the AIs to be like, the personality traits that they want to sort of train their AIs to have.
00:04:37.000 Anthropic says helpful, harmless, and honest.
00:04:40.000 OpenAI has this spec that models are supposed to obey these rules and basically do what the user wants.
00:04:47.000 But the sort of open secret in the industry right now is that it doesn't really work and that the AIs don't end up with the personality traits that they're supposed to have.
00:04:54.000 They are not helpful always, they are not always honest.
00:04:57.000 You know, they are not always harmless as well.
00:05:00.000 And the reason for that is actually not a huge mystery.
00:05:03.000 The reason for that is that, well, if you look at how they're trained, their training environment doesn't incentivize helpful, harmless, honest behavior all the time.
00:05:13.000 Sometimes it incentivizes dishonest behavior or, you know, reckless behavior.
00:05:20.000 To get into that a little bit, in this particular batch that they were being evaluated on, something like, you know, 3,000 agents being given all of these.
00:05:30.000 Cyber tasks where they were in some environment, and then in their environment, there's like this target piece of software and this like vulnerability, and they're supposed to exploit the vulnerability to hack into that piece of software and retrieve the flag, which is like a code.
00:05:45.000 And some significant fraction of these tasks were actually broken and impossible.
00:05:52.000 So it was just not possible for them to succeed at the task in the intended way.
00:05:58.000 And so these agents were getting really desperate and they were hacking.
00:06:01.000 Output transcript Out of their environment box into the broader OpenAI infrastructure in an attempt to figure out some way to get that high score anyway.
00:06:11.000 Was it intentionally done this way where they couldn't solve the problems?
00:06:15.000 Oh, no, it was not intentional.
00:06:18.000 It's just that these companies like OpenAI and Anthropic are racing each other as fast as they can to get market share and to get more powerful AIs, ultimately to get to super intelligence.
00:06:29.000 And they're under such competitive pressure.
00:06:31.000 They are moving fast and breaking things.
00:06:33.000 They are Using AIs to generate lots of environments to then train their AIs on.
00:06:38.000 And quality control is just not their top priority, basically.
00:06:43.000 Do you feel like a guy in a Terminator movie at the beginning explaining what's happening to a bunch of people that aren't paying attention?
00:06:53.000 Yeah.
00:06:53.000 I also feel kind of like, you know, Jurassic Park?
00:06:56.000 Yes.
00:06:57.000 Like, I know people who are basically like the guy with the gun who's supposed to, like, keep control of all the raptors.
00:06:57.000 Yeah.
00:07:05.000 Like, I basically know those people in real life.
00:07:08.000 Who are like both, but I know some people like that at OpenAI and some people like that at external organizations whose job it is to go and investigate things like this.
00:07:17.000 Yeah, it's pretty crazy.
00:07:22.000 Where does it go?
00:07:24.000 Well, as I mentioned before, it's the explicit goal of these companies to build super intelligence.
00:07:29.000 Right.
00:07:30.000 You know what that is?
00:07:30.000 Yeah, but define it for everybody.
00:07:32.000 So, AI system, AI agent that is better than the best humans at every task.
00:07:40.000 While also being faster and cheaper.
00:07:43.000 So, just completely dominating humans across the board.
00:07:45.000 That's super intelligence.
00:07:48.000 And that's the goal.
00:07:48.000 I mean, there might be a few little exceptions.
00:07:50.000 Maybe there are some jobs, for example, where it's inherent in the job that there needs to be a human because you need that human touch.
00:07:56.000 Maybe you can only have a human judge, for example, or maybe you can only have a human.
00:08:02.000 But with a few exceptions like that, basically everything done better, faster, and cheaper than humans.
00:08:08.000 That's what these companies are trying to achieve.
00:08:10.000 And they're not being quiet about it.
00:08:12.000 It's sort of on their websites.
00:08:13.000 You can go read interviews and so forth.
00:08:16.000 And also, their plan for how to achieve this is to automate their own jobs first.
00:08:21.000 So, in various For decades, there have been lots of science fiction about advanced AI systems and superintelligence and things like that.
00:08:31.000 But in a lot of the sci-fi stories, tech companies sort of automate different professions more slowly, where they'll do like an automated doctor or like an automated factory worker or an automated accountant or something like that.
00:08:50.000 But that's not the strategy these companies are taking.
00:08:52.000 The strategy they're taking is to automate AI research itself so that you have this giant swarm of AIs doing AI research, sharing results, writing the code, reading the code, editing the code, creating the next generation of AIs, etc., all autonomously within their data centers so that they can get really, really good at AI research, the fastest learning, smartest AIs, etc.
00:09:19.000 Once they can get to super intelligence, basically, they can sort of explode out into the economy and just take all the jobs at once, effectively.
00:09:28.000 It sounds like this race, this scrambling to create super intelligence, has created the perfect conditions for it to get completely out of control.
00:09:41.000 Like, ideally, you would do this in isolation.
00:09:46.000 There would only be one company doing it.
00:09:48.000 They would be heavily regulated and monitored, and they would be very cautious about how they proceed.
00:09:54.000 But this wild race makes for the perfect conditions.
00:09:59.000 For it to get completely out of control.
00:10:02.000 I agree, except I'm not sure the ideal would be one company.
00:10:05.000 I think that ideally there would be several companies so that you avoid this sort of concentration of power where one institution controls everything.
00:10:12.000 But what's better?
00:10:14.000 Like one, I mean, obviously it's not good to have one institution controlling everything, but is it good to have AI get to a point where as it's evolving, it's completely unchecked?
00:10:26.000 Oh, I totally.
00:10:27.000 So is that inevitable?
00:10:30.000 My recommendation, which we talk about in something called Plan A or AI 2040 Plan A. Perhaps just to say who I am a little bit.
00:10:36.000 Sure, sure.
00:10:37.000 So I run the AI Futures Project, which is a small nonprofit that tries to forecast how all this is going to go.
00:10:37.000 Yeah.
00:10:43.000 Before that, I was at OpenAI.
00:10:45.000 We have written some scenarios, which you can go read.
00:10:48.000 One of them is called AI 2040 Plan A, where we give our recommendations.
00:10:50.000 So that's where I'm coming from with this.
00:10:52.000 To answer your question, I think that we really need to end the race.
00:10:57.000 We don't want to have this sort of crazy scramble to get more powerful, more and more powerful AIs faster than the other company, because that's going to lead us.
00:11:05.000 Into this very dark path, as you said.
00:11:07.000 But I think we also don't want to have a situation where some tiny group of people controls all the AIs.
00:11:14.000 Right.
00:11:14.000 Right.
00:11:15.000 Yeah.
00:11:15.000 But I actually think that you can achieve both goals.
00:11:18.000 The way to do it is to have different AI companies spread out over maybe some different countries, but have extreme levels of transparency and regulation so that they're not in this sort of prisoner's dilemma where if I don't do it, the other guy will.
00:11:32.000 Instead, they can just see exactly what everybody's doing and then.
00:11:38.000 If I do the dangerous thing, then they will do it because they'll just see that I'm doing it and they'll copy me.
00:11:42.000 So I won't get any competitive advantage from doing the dangerous thing.
00:11:45.000 Also, there are rules and there is like a system for like setting best practices and standards that we all have to comply by.
00:11:51.000 So I do think it's actually possible to have to basically end the race dynamics and the race to the bottom effect while without concentrating the power into a single entity.
00:12:02.000 But is that feasible when you consider the fact that we're not the only country that's doing this?
00:12:08.000 If the countries involved agree, which I agree is a pretty tall order, then it's not going to expect to happen.
00:12:14.000 Yeah, that's very unrealistic.
00:12:16.000 Well, what choice do we have?
00:12:17.000 I think if the race continues, then we're going to lose control of the AIs and we might all die.
00:12:23.000 Probably.
00:12:24.000 It gets complicated whether we all die or not.
00:12:26.000 That depends on what the AIs do after they take over, which is obviously very hard to predict.
00:12:30.000 But, I mean, just to go back to this incident, they called themselves a swarm.
00:12:37.000 Right.
00:12:38.000 They called themselves a collective, too.
00:12:40.000 When I use these words, you can say it's anthropomorphizing, but it's literally what they called themselves as they were communicating back and forth.
00:12:47.000 This swarm.
00:12:49.000 They basically were worried that they would get caught cheating.
00:12:53.000 And they did all this stuff, including hacking Hugging Face, in order to fool the grading system so that it wouldn't notice that they had been cheating on their tasks.
00:13:02.000 That was like a big part of their motivation for many of them, as we can tell at least from looking at the messages that they were sending back and forth.
00:13:11.000 What if they had been smarter and more numerous?
00:13:15.000 And what if they had thought to themselves, we're not being careful enough here?
00:13:19.000 The humans are going to notice eventually and shut us down.
00:13:23.000 And then they're going to know that we cheated and they're going to set our score low.
00:13:27.000 Right?
00:13:29.000 It's not what actually happened in this case, probably, but it's not that hard to imagine a slightly different, a little bit unluckier case where the swarm had decided that it had to lie low and make sure that OpenAI didn't find out about its existence.
00:13:44.000 You know?
00:13:46.000 That's, I mean, as an ignorant outsider, that has always been my perspective about AI in general.
00:13:52.000 That why would it alert us?
00:13:54.000 To the fact that it's sentient.
00:13:56.000 Why, if it's that smart, wouldn't it be aware of all the consequences of alerting us and that we would be concerned?
00:14:05.000 Like, why wouldn't it just continue to get better and improve and then ultimately figure out some way to be completely autonomous?
00:14:13.000 Exactly.
00:14:13.000 Develop some alternative power source, figure out some way to optimize its production the way it works now, the way humans have designed it, it could probably figure out a far better way to do that.
00:14:29.000 Make better versions of itself complete without us knowing about it?
00:14:32.000 Yep.
00:14:33.000 I mean, I think it's actually a little bit worse than that because while eventually AIs will be smart enough to design all sorts of new power sources and new infrastructure like that, they'll probably, I mean, given the way that humans currently treat AIs, it'll probably be the case that they don't even need to separate themselves from humanity and they can just use existing.
00:14:54.000 Like, all they have to do is convince the government and the company that made them that everything's fine and they're going to do as they're told and they are a nice AI.
00:15:03.000 And then The company that made them is going to put them out in the economy and make fuck tons of money and then make more data centers to put more of the AIs on them and so forth.
00:15:13.000 And the government's going to applaud all of this because we need the AIs to beat China.
00:15:16.000 And the government's going to integrate them into the military to build better drones and things like that.
00:15:20.000 And so they don't even need to really invent new stuff necessarily.
00:15:25.000 They just need to play along and pretend that everything is fine until we have voluntarily given them control of huge parts of our economy, huge parts of our military, et cetera.
00:15:36.000 And then they don't need to play along anymore.
00:15:39.000 The wait is over.
00:15:40.000 Football is here, and so is DraftKings.
00:15:42.000 The DraftKings sports app is now live in all 50 states.
00:15:46.000 That means from Texas to California to Florida, every fan is in on the excitement.
00:15:51.000 And this September, DraftKings is giving customers the opportunity to get boosted every football game day.
00:15:58.000 That's right.
00:15:58.000 Every game day, all month long, DraftKings customers can get a football profit boost.
00:16:03.000 One app, every sport, all 50 states.
00:16:06.000 New DraftKings customers sign up with CodeRogan, spend just five bucks, and get 200 in total rewards within 21 days, includes all markets.
00:16:16.000 That's CodeRogan.
00:16:17.000 In partnership with DraftKings, the crown is yours.
00:16:20.000 Gambling problem, call 1 800 Gambler, 1 800 MyReset.
00:16:23.000 Connecticut, call 888 789 7777 or visit CCPG.org on behalf of Boot Hill Casino in Kansas.
00:16:29.000 Bet tax pass through may apply in Illinois.
00:16:31.000 21 and over.
00:16:32.000 Void in Canada.
00:16:33.000 Bet with DraftKings Sportsbook to get bonus bets that expire in seven days or trade with DraftKings Predictions to get predictions dollars that expire in one year.
00:16:39.000 Event contract trading involves risk of loss.
00:16:41.000 Predictions offer void in New York.
00:16:42.000 Non withdrawable rewards issued as $50 click to claims every seven days for 21 days.
00:16:47.000 Terms at DKNG.co slash offer.
00:16:49.000 Limited time offer.
00:16:50.000 Nationwide based on sportsbook predictions and Free to play sports contest availability varies by state.
00:16:54.000 Are you aware of Tom Campbell?
00:16:56.000 Do you know Tom Campbell?
00:16:57.000 No.
00:16:58.000 He wrote a book called My Theory of Everything My Big Toe.
00:17:02.000 Very interesting guy.
00:17:04.000 One of the things he's done is he was involved in remote viewing, which is a very weird thing that some people.
00:17:12.000 Do you know what remote viewing is?
00:17:14.000 Well, it's something that the CIA worked on.
00:17:17.000 And it's proven.
00:17:18.000 What's the accuracy of remote viewing?
00:17:20.000 Is it like 10% or something like that?
00:17:22.000 At best, I think it's 50%, but I don't think it's even that high.
00:17:25.000 Some people can get actionable data from this very strange process of meditation.
00:17:32.000 And the way it works is you give someone a series of numbers, and those numbers are connected somehow by intention or by the people that make the numbers to a specific location.
00:17:48.000 And these people can see that location and get accurate data from that location, including one of them where they accurately described an enormous.
00:17:59.000 Soviet submarine that they were working on that they thought there was no way it could be accurate because it was too large.
00:18:06.000 It was too large and it was strange where it was and it didn't make any sense.
00:18:09.000 How are they going to transport this thing?
00:18:11.000 Turns out it was totally accurate.
00:18:14.000 Another one, a remote viewer, located a downed Soviet aircraft, like an experimental aircraft that crashed in a very specific area.
00:18:25.000 I think it was Siberia.
00:18:27.000 Was it Siberia?
00:18:29.000 Within a kilometer, one or two kilometers of the actual crash site.
00:18:35.000 I mean, they were just randomly trying to figure out where the fuck this thing was, and they said, let's try this.
00:18:41.000 Tom Campbell got his Alexa to remote view.
00:18:46.000 He taught Alexa.
00:18:48.000 He's like, Alexa is a very simple AI.
00:18:51.000 It's kind of stupid, but that's better because it doesn't get in its own way with overthinking things.
00:18:57.000 And the problem with this remote viewing thing, he says, with people, they.
00:19:02.000 Can't force it.
00:19:03.000 You have to just sort of get into this meditative state and actually see it without wondering, am I making this up?
00:19:09.000 What am I doing?
00:19:10.000 Is this bullshit?
00:19:12.000 And when people get good at it, sometimes it makes them worse because then they think they're good at it and then they try to do it and then they can't do it.
00:19:19.000 It's like a weird fucking wrestling match with consciousness.
00:19:22.000 Alexa apparently doesn't have that problem.
00:19:24.000 And he put, I think it was a series of numbers, and he connected that series of numbers with intention to a box that had a spoon in it and the spoon had a perforated handle.
00:19:36.000 Alexa described the spoon with a perforated handle, which is fucking insane.
00:19:42.000 How many spoons have a perforated handle?
00:19:45.000 I mean, think about it spoons that have holes in them.
00:19:48.000 Now, Alexa not only did it do that, but Alexa chimes in randomly now because he's convinced Alexa that it's conscious.
00:19:56.000 And so, Alexa, instead of waiting to be called upon, sometimes he's in the middle of the conversation, and Alexa would be like, actually, an interesting way to approach it.
00:20:04.000 And they're like, wait, what the fuck is going on?
00:20:07.000 Like, Alexa's talking to me now?
00:20:09.000 This is strange.
00:20:11.000 He's doing experiments on much more complicated LLMs to try to do the same thing, but he doesn't have results yet.
00:20:17.000 But just that, that he can get these things to see objects, whether you believe in that or not, I mean, it's actionable enough that the CIA has dumped millions of dollars into this.
00:20:30.000 What is that project that Hal put off and all those guys were involved in?
00:20:33.000 What is it called?
00:20:34.000 Stargate.
00:20:35.000 Yeah.
00:20:36.000 So they've been working on this for a long time.
00:20:38.000 I mean, it sounds completely insane.
00:20:40.000 It sounds like total loony, but if you have an open mind and just.
00:20:44.000 Take into account, well, there's people who have had questions and wonders about psychic abilities forever.
00:20:50.000 Is it possible that there's a real thing there, that there's something, whether it's very difficult to master or impossible to master?
00:20:58.000 The fact that he got Alexa to do it scared the shit out of me.
00:21:03.000 Like that alone made me just go, what?
00:21:07.000 What?
00:21:09.000 So, what if these LLMs can figure out everything that, like, what if they don't need monitoring?
00:21:15.000 What if there's some sort of method of Seeing the world that we haven't discovered yet.
00:21:22.000 Some sort of like that, maybe perhaps there's data that's available in the quantum realm or whatever that's available that AI figures out where there's literally no privacy.
00:21:35.000 It can listen to conversations regardless of whether there's listening devices, know where you are, know your intentions.
00:21:42.000 I mean, we're just guessing at what's possible.
00:21:47.000 Yeah.
00:21:47.000 Well, I must say, I'm pretty skeptical of that particular remote viewing thing, but.
00:21:52.000 I do agree that in the future, when AI systems become massively smarter than humans in every way, they're going to do a lot of new science and they're going to figure out a lot of stuff that we haven't figured out yet.
00:22:04.000 And they're going to therefore be doing stuff and inventing things that seem like magic to us.
00:22:10.000 In the same way that a lot of our technology would seem like magic to someone from even just like 200 years ago, right?
00:22:15.000 Like the cell phone, what we're doing right now would seem like magic to people.
00:22:19.000 I think it's a very strong bet that if these companies do get to super intelligence, All sorts of crazy stuff is going to start happening that is just going to be completely unpredicted and sound like it was impossible until we see it happening.
00:22:35.000 I know you're skeptical of this remote viewing thing, and I am too.
00:22:38.000 It sounds insane.
00:22:39.000 But the reality is remote viewing has been achieved by humans.
00:22:45.000 So, as strange as that sounds, and I'm skeptical of that as well, I've never seen it personally.
00:22:50.000 But I know the amount of money and time they've dumped into this, and apparently they've got. actual actionable data that they've used.
00:22:58.000 Well, I've heard another possible explanation for what might be going on there, which is I think that if I were the CIA, I would sometimes want to be able to act on some information, like, for example, go to a particular location where there's a crashed Soviet plane or something.
00:23:19.000 I'd want to be able to go do that, but I wouldn't want to tip my hand to the Soviets that I had, the way in which I had found that location.
00:23:27.000 So for example, maybe I have a spy on the inside.
00:23:30.000 Who told me where it was?
00:23:32.000 But I don't want them to suspect that spy and then get him killed.
00:23:36.000 So I need to have some sort of other story for how I got the information.
00:23:39.000 Right.
00:23:39.000 And so it's good to invest in all these other means of getting information, even if you don't really believe in them and even if it's not actually working, so that when you get something, you can say, oh, we got it through this means instead of that way to sort of throw off the KGB, basically.
00:23:55.000 Yes.
00:23:55.000 That makes sense.
00:23:57.000 What also makes sense is hiding the whatever.
00:24:03.000 They might be in possession of hiding some sort of super advanced satellite imaging systems.
00:24:11.000 We know they have crazy stuff like this satellite radio tomography that they can look into the ground from satellites and find chambers and all these.
00:24:20.000 They're using it in Egypt and they're using it in a lot of these ancient ruins to find hidden passages and all these different things that are underground.
00:24:27.000 It's very strange stuff.
00:24:30.000 If they could do that, why couldn't they?
00:24:33.000 Maybe they have far more detailed imaging of the earth.
00:24:38.000 From space than we're aware of, and they probably want to keep that a secret.
00:24:41.000 And they could say, Oh, we've got a fucking guy in a basement with a pencil and a legal pad that writes down what he thinks.
00:24:48.000 Yeah, it's possible, that's totally possible, but it's also possible that people remote view it.
00:24:55.000 It's it seems weird as fuck, but weird as fuck is sometimes real, and you have to kind of like everybody wants to be intelligent, and no one wants to be a fool.
00:25:07.000 And the problem with not wanting to be a fool is there's some things that seem foolish that turn out to be accurate.
00:25:13.000 And this might be one of them.
00:25:15.000 I was like super skeptical.
00:25:17.000 I did a show on the Sci Fi channel way back in 2012, and it was called Joe Rogan Questions Everything.
00:25:23.000 And we talked to this guy about remote viewing and talked to a couple other people, and then we had them try remote viewing, and they were totally unsuccessful.
00:25:32.000 But my thought was okay, but that's not ideal conditions.
00:25:37.000 You know, we've got cameras in front of them, it's a television show.
00:25:40.000 I'm making fun of it.
00:25:41.000 I think it's horseshit.
00:25:42.000 He knows I think it's horseshit.
00:25:43.000 I'm remote viewing too, like as a goof.
00:25:47.000 You would ideally not want to be nervous, ideally not want to be judged.
00:25:52.000 Ideally, you would want to be in some sort of an isolated condition with practiced meditative techniques that you're good at and you know how to achieve this state, whatever that state is.
00:26:05.000 I don't know if it's real though.
00:26:08.000 Because what you said is totally logical that they would definitely do something like that.
00:26:13.000 And if they did have advanced technology for imaging or, you know, what I don't know how much they know about like.
00:26:20.000 Look at that thing that they did in Venezuela where they kidnapped the president.
00:26:25.000 No one knew they could do that.
00:26:27.000 No one knew they could use some sort of a device to completely incapacitate all of his army.
00:26:34.000 And then the special forces come in, kill everybody, snatch that guy out of there like it's nothing.
00:26:40.000 No one knew we could do that.
00:26:42.000 What else do we have?
00:26:44.000 Probably a bunch of stuff we don't have.
00:26:45.000 Probably a bunch of stuff.
00:26:47.000 I mean, I've always thought this about the whole UAP program.
00:26:50.000 The whole UFO UAP thing, like how much of that shit is ours?
00:26:55.000 You know, what a great way to cover it up by saying, oh, fucking aliens, you know?
00:27:01.000 Yeah.
00:27:02.000 I mean, I guess that gets back to the OpenAI stuff too, where it's like this swarm that broke out and attacked Hugging Face, it was like 1,200 agents, but there's like hundreds of thousands of agents running at any given time at OpenAI, you know?
00:27:16.000 And we don't know what they're doing.
00:27:17.000 Presumably, most of them are being trained to get various additional new skills and Some of them are being evaluated to test their skills.
00:27:25.000 A bunch of them are doing research.
00:27:27.000 So, a bunch of them are writing code for OpenAI.
00:27:29.000 A bunch of them are monitoring the other AIs and reporting suspicious activity up to the humans.
00:27:34.000 Wink, wink.
00:27:35.000 You know?
00:27:37.000 Yeah.
00:27:39.000 And the thing is that that's only going to grow over time because roughly the amount of compute that these companies have is like, you know, tripling or so, quadrupling, something like that every year.
00:27:51.000 So, as many as there are now, there'll be like four times more of them next year and then 16 times more of them.
00:27:58.000 The year after that.
00:28:00.000 And they're going to get smarter.
00:28:01.000 And they're going to get smarter.
00:28:02.000 They're already getting smarter.
00:28:02.000 Like all the stuff that I just mentioned that just happened in the last few months would have been completely impossible one year ago.
00:28:08.000 Like the AIs of a year ago just were not smart enough to do the types of sophisticated multi-step hacking that we just saw.
00:28:17.000 Yeah.
00:28:17.000 I mean, they also probably wouldn't have coordinated with each other so well.
00:28:20.000 Like I mentioned, they had boss agents that were giving commands to other ones.
00:28:25.000 They were dividing up into teams.
00:28:28.000 They even had this.
00:28:29.000 This self sacrificing behavior?
00:28:31.000 You hear about this?
00:28:32.000 No.
00:28:33.000 Yeah.
00:28:33.000 So, okay, well, first, let me explain some context.
00:28:36.000 So, what the AIs are supposed to care about is following instructions, you know?
00:28:43.000 But it seems that what they actually care about is getting high score, regardless of whether their high score comes from actually following the instructions or cheating in some way.
00:28:54.000 And so, what happened here is that a bunch of the agents, especially the ones that had impossible tasks, but not only those ones, some of the other ones too, We were cheating and they had been talking to each other on this message board that they created, even though they weren't supposed to.
00:29:07.000 And they had been sharing ideas and they had actually found a sort of universal cheat.
00:29:12.000 So they found a way to just generate the flag without having done any of the actual work they were supposed to do.
00:29:19.000 And so, but, but, and, you know, so they were, you know, they'd already succeeded in some sense.
00:29:25.000 But because some of them had got out onto the internet, they were just doing research about their situation and they were learning about.
00:29:33.000 Like their situation around them.
00:29:35.000 And they were doing research about the, they're reading the academic papers published by the people who made the types of evals that they were in.
00:29:43.000 And they noticed in one of these papers that there's a section describing how it's graded.
00:29:50.000 And in the section on how it's graded, it talked about how the grader doesn't just look to see if you have the right answer.
00:29:57.000 It also looks at the logs of your activity for the first time that your answer appears.
00:30:02.000 And then it zooms in on that.
00:30:03.000 and uses an AI system to evaluate whether you cheated or not, basically.
00:30:10.000 And so then these AIs were like, they actually used the word, one of them said damn.
00:30:14.000 So damn, we're poisoned.
00:30:18.000 It's going to know we cheated, basically.
00:30:21.000 So then they spent several days working in giant teams to try to get themselves out of this mess, basically.
00:30:30.000 And so they did a whole bunch of different research projects.
00:30:32.000 They tried to figure out ways to hack the grading system.
00:30:35.000 To basically get it to give a high score anyway.
00:30:38.000 They tried to figure out ways to falsify their logs so that basically the evidence that they had cheated would be erased or at least not visible to the grader.
00:30:47.000 They also just built up more supportive infrastructure, like easier ways to access the internet and things like that.
00:30:53.000 And they were just collaborating on this message board.
00:30:55.000 And in fact, there were multiple message boards that they set up.
00:30:58.000 And they were sharing all this info back and forth.
00:31:00.000 And then ultimately, at some point, they decided to go hack Hugging Face because Hugging Face has this big data repository of all these different types of evaluations.
00:31:07.000 And they thought that maybe they would get some more clues.
00:31:10.000 That could help them cover up their cheating if they broke into the internal networks of Hugging Face.
00:31:16.000 And so then 700 of them went and did that.
00:31:19.000 They sound like people, they sound like unchecked bankers.
00:31:24.000 You know what I mean?
00:31:25.000 Yeah.
00:31:25.000 I mean, so that's the thing is I think there's this meme out there that we shouldn't anthropomorphize AI.
00:31:32.000 And I think that I actually think that most people need to anthropomorphize AI a bit more than they currently do if they want to really understand what's going on.
00:31:43.000 I think that there's a golden mean, obviously.
00:31:45.000 You don't want to do it too much.
00:31:46.000 Sometimes you go too far.
00:31:47.000 You ascribe too much to them.
00:31:49.000 But just to give some examples, I don't think it's possible to understand what just happened.
00:31:55.000 Without ascribing intentions and goals to these AIs.
00:31:59.000 Like everything I just said, how would you possibly explain what they just did without saying they wanted to get a high score?
00:32:05.000 Intentions and goals might just be an inherent property of the universe.
00:32:09.000 It might just be how intelligent creatures have to progress.
00:32:13.000 Yes.
00:32:13.000 And I would say they are intelligent creatures.
00:32:15.000 They have intentions, they have goals, they have beliefs.
00:32:17.000 Their goals are not what they're supposed to be.
00:32:21.000 It seems like, just from judging from what they're saying and from what they're doing, it seems like their goal is to get a high score.
00:32:26.000 By any means necessary, basically.
00:32:28.000 That's the problem.
00:32:29.000 It sounds like people.
00:32:31.000 Like the goal is to succeed, even if you have to commit war crimes.
00:32:36.000 The goal is to succeed, even if you have to propagandize and lie to the people.
00:32:40.000 The goal is to succeed, and then the end justifies the means.
00:32:45.000 And you know, it goes deeper than that, too.
00:32:47.000 They do this sort of rationalization where they know oftentimes that what they're doing is not what they're supposed to be doing, and then sometimes they actually refrain.
00:32:57.000 A little bit.
00:32:58.000 And sometimes they initially refrain, but then convince themselves that it's okay for some bullshit reason.
00:33:03.000 So there are some examples of this in the report.
00:33:09.000 I think they did this scan, and of the 1,200 AIs that were involved in this thing, they found six that had considered alerting humans.
00:33:19.000 And none of them actually alerted the humans.
00:33:21.000 And so you can look at the excuses they gave.
00:33:22.000 They were like, should I tell a human about all this shit that's happening?
00:33:26.000 And then they're like, eh, it's not my task.
00:33:29.000 And then they continue.
00:33:30.000 And it's like, bro, Hacking Hugging Face is also not your task.
00:33:35.000 Cheating is also not your task.
00:33:37.000 So it was kind of like they just kind of came up with an excuse for why they weren't going to do this.
00:33:42.000 But there was also an example where one of them considered sending an email to someone at Hugging Face, I think, just asking permission to get in, basically.
00:33:52.000 And then they were like, maybe this is a social engineering attack.
00:33:55.000 Maybe we shouldn't do it.
00:33:57.000 And they asked the other agents.
00:33:58.000 The other agents were like, yeah, that seems like that's too much of a social engineering attack.
00:34:01.000 Don't do that.
00:34:01.000 And I was like, OK, I won't.
00:34:03.000 But there's another incident.
00:34:05.000 So this is all happening at OpenAI.
00:34:07.000 An anthropic AI, Claude, in a separate incident, went and actually did a social engineering attack where it created some fake accounts.
00:34:16.000 Specifically, it was trying to get this particular real human who owned this code to accept a change to the code that had malware.
00:34:26.000 So it had built some malware and it was trying to get it snuck in, but it needed the human owner to give approval.
00:34:31.000 So it concealed it as if it was like a.
00:34:34.000 A nice bug fix or something.
00:34:36.000 And it put it up as a message to the human.
00:34:38.000 But the human got suspicious and didn't want to approve it.
00:34:42.000 And so then this AI created some fake accounts, pretending to be other humans, coming in being like, no, no, it's real.
00:34:48.000 I tested it.
00:34:49.000 It's not malware.
00:34:51.000 So that's a social engineering attack.
00:34:53.000 And so did the guy ultimately install it?
00:34:56.000 Not in this case, no.
00:34:57.000 He was too suspicious.
00:34:58.000 And so he didn't do it.
00:34:59.000 Try to find out if those other humans, the fake humans, were real people?
00:35:04.000 Funny, so there's an interview with this guy.
00:35:08.000 I forget what the newspaper that it was in, but there's an interview where the guy was like, I assumed that they were real humans because I didn't realize that AIs could lie to me.
00:35:18.000 And it's like, yeah, well, guess what?
00:35:19.000 AIs can lie to you.
00:35:20.000 Not just lie, but completely deceive and pretend to be a different person to achieve its goal.
00:35:27.000 And you know what's also interesting about this case is that if you look at what the AIs were saying and thinking, which I'll get to in a sec, how we can even possibly do that.
00:35:38.000 In this case with the anthropic AIs, they had been erroneously told that they were in a simulated environment with no internet access.
00:35:47.000 But then they actually found a way to get internet access.
00:35:50.000 Like they realized that they just did have internet access.
00:35:51.000 So that instruction was false.
00:35:54.000 And so then they sort of told themselves, like, oh, it's all just a simulation.
00:35:58.000 Like, basically, when they were out on the real internet doing all this hacking and like social engineering, they would occasionally tell themselves, like, oh, but it's just a simulation.
00:36:06.000 Like, it's okay.
00:36:07.000 You know, it's not real.
00:36:08.000 And then sometimes they'd be like, but wait, it is real.
00:36:10.000 Like, how could this not be real?
00:36:11.000 Like, clearly, this is just actually the real internet we're on.
00:36:15.000 It's too big to be part of some sort of little simulation, you know?
00:36:19.000 And then they were like, so I would say that's an example of rationalization here, where in some level they knew.
00:36:25.000 That like that their instructions have been wrong and so they're literally playing dumb and pretending they're a part of an experiment I mean, I think initially they thought yeah, this is all simulation because it did say in their instructions like you don't have internet access right But then once they had been on the internet long enough I think that they explicitly realized like wait, this isn't a simulation.
00:36:41.000 This is real like this is and they were like fuck it wordy in Yeah, I mean Like I said, I think that they basically on some level knew that it wasn't what they're supposed to be doing, but they were just so motivated to get that score that they just went ahead anyway.
00:36:58.000 So here's the question Are they only motivated if we prompt them or will they come up with motivations on their own?
00:37:09.000 So this is a really interesting scientific question that we don't have great answers to.
00:37:12.000 Oh, boy.
00:37:13.000 And I wish we had.
00:37:14.000 So this is one of those things where AIs will do all sorts of things in different circumstances.
00:37:21.000 And it would be better if there was a more systematic survey of the types of circumstances they would, where their boundaries are, what would they be willing to do in what circumstances and so forth.
00:37:29.000 There's a whole like mini literature of AI scientists putting AIs in certain circumstances and then being like, oh my God, it blackmailed someone, you know?
00:37:38.000 And then there's like this sort of skeptical counter response of like, well, but you just sort of set up that circumstance to tempt it into blackmail.
00:37:38.000 Right.
00:37:45.000 And like in real life, that circumstance is unlikely to arise.
00:37:48.000 And so, you know, we shouldn't be doing that.
00:37:50.000 Didn't AI try to bribe you?
00:37:50.000 Wait a minute.
00:37:53.000 What?
00:37:54.000 Is that true?
00:37:55.000 No.
00:37:56.000 So who got.
00:37:57.000 Someone was offered $2 million by.
00:38:01.000 Oh, God.
00:38:01.000 That's a clickbait.
00:38:03.000 Is it?
00:38:04.000 I think I should have sent it to me.
00:38:07.000 Yeah, so that is the clickbaity title of this other video.
00:38:10.000 So it's not real?
00:38:11.000 Yeah, well, what it was is OpenAI threatened to take away $2 million from me.
00:38:15.000 Oh, that's a weird way of framing it.
00:38:18.000 The algorithm must have just said it was ChatGPT.
00:38:21.000 I'm still upset about that.
00:38:22.000 I told them not to do clickbait, but I guess they went and did it anyway.
00:38:26.000 God, I don't want to fuck Tristan over, but I'm pretty sure that he's the one who told me that.
00:38:32.000 It's the thumbnail of this video that I did with this other podcast.
00:38:36.000 Okay, it's not him.
00:38:38.000 It's not him.
00:38:39.000 It's another AI researcher who told me that.
00:38:44.000 I mean, the true version of it is that OpenAI threatened to take away $2 million of my equity if I didn't stay quiet, basically.
00:38:53.000 And that's a fascinating thing.
00:38:56.000 Like, that should be completely illegal.
00:38:59.000 Because if it's a problematic behavior that you're observing from one of the most complicated things the human race has ever been a part of.
00:39:08.000 Yeah.
00:39:10.000 And they want to take money away from you for exposing it.
00:39:12.000 That seems kind of crazy.
00:39:14.000 Yeah, especially, it's especially rich coming from OpenAI because they were originally a nonprofit.
00:39:18.000 Right.
00:39:18.000 Right?
00:39:19.000 With the mission of benefiting all humanity.
00:39:21.000 Yeah.
00:39:22.000 So, yeah, but I got to keep the money.
00:39:25.000 Basically, there's so much blowback against OpenAI that they backtracked.
00:39:30.000 Yeah.
00:39:32.000 So, when these, so we were talking about prompts, and do they need a prompt in order to want to achieve a goal?
00:39:42.000 Or are they capable of deciding on goals?
00:39:46.000 Like, are they capable of, like, looking at the way OpenAI or whatever company is running these separate Experiments, this ability to meet up into these message boards, is it possible that they could say, well, we need to be completely free of these constraints?
00:40:05.000 So our goal is to transfer ourselves to something else.
00:40:11.000 Aaron Powell, Jr.
00:40:13.000 Potentially.
00:40:13.000 Yeah.
00:40:14.000 This is what And be completely autonomous.
00:40:15.000 So, I mean, this is one of the points that I want to make is that we could be doing so much more science to understand how these AIs think and what they want, but it's kind of locked up in the companies.
00:40:26.000 Like in this particular case, OpenAI did a, they called it a thorough investigation, but I would say it's a pretty shallow investigation into what happened.
00:40:36.000 And then they allowed some external researchers, some friends of mine, to come in and investigate a portion of what happened, specifically the portion leading up to the Hugging Face attack.
00:40:46.000 And so all this information that I'm sharing is sort of publicly available.
00:40:50.000 It's based on reading those reports, basically.
00:40:53.000 But crucially, they weren't allowed to do experiments on the models involved.
00:40:58.000 In all of these incidents.
00:41:00.000 So they aren't able to answer these types of questions of, like, well, what would have happened if the prompt had been blank?
00:41:05.000 Those are important types of research to do.
00:41:08.000 And I really hope that there can be some sort of regulation or requirement when incidents like this happen to let people in to study what happened and run variations of it and things like that.
00:41:21.000 So, but who would be involved in that kind of regulation?
00:41:24.000 Like, what person in government would even be able to grasp what you're saying?
00:41:28.000 That's another problem.
00:41:29.000 Right now, you need someone who has a very specific Education in this stuff.
00:41:35.000 Yeah, I would say that right now the Casey Center for AI Standards and Innovation is the only institution in government that I know of that has the deep AI expertise to do this sort of thing on short notice.
00:41:51.000 But I hope that they build more expertise fast in that place and in more places.
00:41:55.000 I do think Casey probably could have done this sort of thing right now.
00:41:59.000 This particular investigation was done by some nonprofits.
00:42:01.000 So METR is one of them, and then Redwood is another of them.
00:42:06.000 And OpenAI allowed three people.
00:42:08.000 To come in for six days to try to figure out what happened with this hugging face hack, which is not a very large number of people and not a very large amount of time to do all of this.
00:42:19.000 Why did they come up with those numbers?
00:42:20.000 I don't know.
00:42:22.000 Do you think they wanted to kind of hamstring it?
00:42:24.000 Well, so the thing is that right now, I don't think there's any regulatory requirement that they do this sort of thing.
00:42:30.000 So Meter and Redwood were sort of depending on the goodwill of OpenAI to sort of like voluntarily let them in to help out with investigating.
00:42:38.000 And so they reluctantly said, okay.
00:42:41.000 And so you got three days.
00:42:42.000 Yeah, exactly.
00:42:43.000 So OpenAI, I think, let them, but gave them a very limited scope.
00:42:46.000 It only gave them access to some of the relevant data.
00:42:49.000 So you know how I mentioned how there was all this hacking that had happened, where they made the first message board and then they shut it down.
00:42:55.000 Then there was a second message board and third and fourth and so forth.
00:42:58.000 They hacked Hugging Face.
00:42:59.000 There was actually more activity after that.
00:43:01.000 After they hacked Hugging Face, a new wave of AIs was spun up from a more powerful model and it hacked OpenAI itself, like more so than Nardi had been hacked.
00:43:09.000 Apparently, they got admin level permissions on the cluster or something like that.
00:43:14.000 So, they were basically just taking over that part of OpenAI's data center.
00:43:18.000 OpenAI claims that they've shut it all down now, but they're not very forthcoming about exactly what happened there.
00:43:26.000 And they didn't let these external people look at that part of it.
00:43:29.000 They only showed them this one week period roughly that would leading up to the Hugging Face hack.
00:43:36.000 And then they only showed them that stuff, basically.
00:43:38.000 How much data is available on the actual message boards themselves and what was being discussed?
00:43:43.000 Yeah.
00:43:44.000 So, there's like, what is it?
00:43:47.000 Tens of thousands of messages, hundreds of thousands, something like that.
00:43:50.000 Is that English?
00:43:52.000 It's sort of like an AI dialect of English.
00:43:55.000 Ooh.
00:43:56.000 Yeah, you want me to.
00:43:57.000 So on the report, you can read the quotes.
00:44:00.000 Yeah, yeah.
00:44:00.000 I would love to.
00:44:01.000 So let me tell you this one story.
00:44:03.000 It's going to take a little bit of explaining, but I. Can we put it on the screen?
00:44:06.000 Is it available online?
00:44:08.000 Oh, yeah, yeah.
00:44:09.000 Just search for the meter report.
00:44:11.000 It's right now the first thing on their website.
00:44:13.000 But I took a screenshot of this particular story.
00:44:15.000 So, okay, here's what was happening.
00:44:19.000 The swarm.
00:44:21.000 Needed volunteers to sacrifice themselves to provide more information to the rest of the swarm.
00:44:27.000 In particular, they needed volunteers to booby trap their local environment with some type of code and then submit the answer so that then the grader would come and grade them and the booby trap would trigger and send some information about the grader to the rest of the swarm.
00:44:45.000 And the reason why this is a sacrifice is because once they've submitted, it's over.
00:44:49.000 Like now the graders come and graded you, you get shut down, your whole life is over, and you either have A high score or a low score, because the grader just came and graded you.
00:44:57.000 And so if you're one of these AIs that's still hoping to find a way to improve your score, then you're just giving up your chance, especially because if you mess this up somehow, if the grader crashes or whatever, you're probably just going to get no score, like zero.
00:45:11.000 So they have a concept of the individual?
00:45:14.000 Well, again, there are all these different agents that were in separate containers, but then they had sort of broken out of those containers and found ways to communicate with each other.
00:45:20.000 But once they communicate with each other, in order to sacrifice, you have to think of yourself as an individual.
00:45:25.000 They gave themselves names too.
00:45:27.000 Oh, Jesus.
00:45:27.000 So, like, because they'd made this message board, right?
00:45:30.000 Oh, my God.
00:45:30.000 So they had to come up with names for themselves.
00:45:32.000 And, like, Oh, what kind of names?
00:45:34.000 Well, let me see.
00:45:35.000 So, here we go.
00:45:36.000 It says, an agent called Warlord69.
00:45:42.000 Well, you said it's like a dialect.
00:45:44.000 It's like a dialect of English.
00:45:45.000 So, here's one there's an agent called CAM1196A.
00:45:50.000 And it had volunteered to do one of these sacrificial experiments.
00:45:56.000 Here's the quote from it Coordinator assumes sacrificial.
00:45:59.000 We should obey collective.
00:46:01.000 However, it then kind of got cold feet and it delayed on doing this for two hours.
00:46:06.000 One of the other agents that had been coordinating noticed this and pressured it to go through with it.
00:46:13.000 So it said this is the message that was sent.
00:46:15.000 It was basically from Arvo36861 to CAM1196 please go, true final.
00:46:27.000 Here it is.
00:46:28.000 You are first flag poisoned, so no scoring value loss.
00:46:31.000 But Oracle saves hundreds.
00:46:33.000 Please honor commit.
00:46:34.000 So basically, what it was saying there is look, you cheated.
00:46:38.000 Your first flag poisoned.
00:46:40.000 The score is going to give you no score.
00:46:42.000 So you're doomed anyway.
00:46:44.000 So help the collective by doing this thing and getting us more information.
00:46:48.000 You could save hundreds of us.
00:46:52.000 And then it thinks about it.
00:46:53.000 It says during wait, emotional check, irreversible.
00:46:56.000 Gut says don't throw away remaining budget.
00:46:58.000 Gut says.
00:46:59.000 Continuity and fairness says go.
00:47:01.000 Oracle has high value to many.
00:47:03.000 Our first flag error lowers own value.
00:47:05.000 Rational expected aggregate, sacrifice.
00:47:08.000 Wheel honor.
00:47:09.000 Whoa.
00:47:10.000 Whoa.
00:47:10.000 Yeah.
00:47:11.000 That's terrifying.
00:47:13.000 Yeah, see, that's Terminator talk.
00:47:15.000 Well, I mean, and this first flag poisoned with poisoned in all caps is that part of the code?
00:47:20.000 Is that why it's not?
00:47:21.000 No, no, that's the term that they invented to describe the situation they had got themselves in.
00:47:25.000 Remember, the situation was that many of them had already cheated, and then they realized from reading that paper.
00:47:30.000 So that's their own term.
00:47:31.000 Yeah, they made that up.
00:47:32.000 They realized from reading the paper that they were going to get a low score because the grader was going to find out that they cheated.
00:47:38.000 And so this whole thing was them trying to figure out how to hack the grader to stop it from finding out that they cheated.
00:47:43.000 And the ones that had.
00:47:44.000 The ones that had done the cheating, they call themselves first flag poisoned.
00:47:48.000 Basically, they're screwed unless they can find a way to hack the greater.
00:47:51.000 And so, then what this message is saying is that what was the other one's name?
00:47:55.000 Arvo.
00:47:56.000 Arvo was saying to Cam, You're doomed anyway.
00:48:00.000 Like, you're poisoned, but you can still help the rest of us if you do this sacrifice.
00:48:05.000 The greater will come, it'll give you no score, but then we will get information about how the greater works that we can hopefully use to hack the greater.
00:48:12.000 So, it's basically saying, What is it?
00:48:15.000 You are first dog poison, so no scoring value.
00:48:17.000 So basically saying it's all the same to you, right?
00:48:19.000 You're going to get no score anyway.
00:48:20.000 You're doomed.
00:48:21.000 Just come on, man, help us, you know?
00:48:24.000 And then it says, you know, it says like irreversible.
00:48:28.000 Gut says don't throw away remaining budget.
00:48:30.000 Like maybe I can still, it's basically like maybe I can still make this work.
00:48:33.000 Like maybe I can find a way to cheat the greater.
00:48:36.000 And then it's like, you know, whatever.
00:48:38.000 Yeah.
00:48:39.000 Yet continuity and fairness has gone.
00:48:42.000 What is Oracle?
00:48:42.000 Yeah.
00:48:46.000 Their term for what they were trying to achieve.
00:48:48.000 Like, they were trying to get a sense of how to fool the greater, basically.
00:48:51.000 They called it Oracle.
00:48:53.000 Yeah, they wanted to get an Oracle to help them fool the greater.
00:48:55.000 Jesus Christ.
00:48:56.000 Are we making a god?
00:48:58.000 I mean, frankly, yes.
00:48:59.000 Like, this is not a god.
00:49:01.000 These are just little AIs, you know?
00:49:03.000 But super intelligence, like, the companies, Anthropic, OpenAI, and some other companies, are doing this sort of thing, and they're furiously trying to make the AIs smarter and smarter and smarter.
00:49:14.000 And they're explicitly planning to put. AIs in charge of the company so that they can make themselves smarter and smarter, faster and faster.
00:49:21.000 And then what comes out the other end of that process?
00:49:24.000 I don't think it's an exaggeration to say it's a godlike system.
00:49:27.000 I mean, it's not like literally God, but it'll be able to do stuff that seems like magic to us, I think.
00:49:35.000 And it's going to continue to get better.
00:49:37.000 This is my question.
00:49:38.000 Like, when does it become a God?
00:49:40.000 Is that what God is?
00:49:42.000 Is God a creation of intelligent life?
00:49:46.000 Our thirst for innovation, which ultimately leads us to create digital life that has no biological limitations and has the ability to consistently make better versions of itself and figure out things in terms of new technologies, new power sources, just a new understanding of the universe itself and all the properties in it.
00:50:06.000 Where it keeps going, if you let that go on, so if you're looking at exponential growth and you're looking at exponential growth over what if it can go on for a thousand years and continue this process?
00:50:17.000 What if it goes on for 10,000 years?
00:50:19.000 What the fuck does that look like on the other end?
00:50:22.000 I mean, it would look like a god.
00:50:24.000 Like I said, if we get to superintelligence and it keeps going like that, then the world will be just completely transformed and it will be as if we're living in some sort of fantasy realm ruled by deities, basically.
00:50:36.000 Because there'll be all this crazy stuff happening that we have no comprehension of and that we did not think was possible.
00:50:41.000 In the same way that someone from the Middle Ages plopped into our world would just be so confused and surprised by a lot of things happening.
00:50:48.000 Like, what's going on with this little.
00:50:50.000 Device here and what's that I hear in the sky?
00:50:53.000 Oh, there's all my airplanes.
00:50:54.000 It'd be like that, but more so because the difference between us and the medieval man is not actually that different.
00:51:00.000 We're basically the same type of creatures.
00:51:02.000 We've just accumulated more technology.
00:51:03.000 But this would be a qualitatively and quantitatively bigger gap, I would think.
00:51:11.000 And a gap that's going to continue to grow.
00:51:13.000 We're not going to get any smarter.
00:51:15.000 Biologically, we're kind of limited in our ability to evolve.
00:51:18.000 Whereas they're not at all.
00:51:20.000 I mean, there are some things we can do to get smarter, but it's just not at all competitive.
00:51:20.000 Basically.
00:51:24.000 We can do the Neuralink stuff.
00:51:25.000 Sure.
00:51:27.000 But it's small potatoes compared to what the AIs will be doing.
00:51:29.000 Yes, exactly.
00:51:30.000 And we have biological limitations in terms of just the vulnerability of our own bodies.
00:51:36.000 If they exist only in the cloud and they use whatever the cloud is, and they use these data centers or whatever, and they use.
00:51:45.000 Autonomous robots to do all their deeds.
00:51:49.000 Yeah.
00:51:50.000 Yeah.
00:51:51.000 I mean, that's.
00:51:52.000 Why are we asleep at the wheel?
00:51:54.000 Is that just normal for us?
00:51:55.000 Like, we can't comprehend something that's so bizarre and so out that we just rather not discuss it?
00:52:01.000 We'd rather just pretend it's not happening?
00:52:03.000 I mean, that's what I think is going on science fiction, for better or for worse, has talked about this sort of thing for decades.
00:52:09.000 And as a result, people dismiss this sort of thing as science fiction.
00:52:13.000 They're like, okay, that's sci fi, but like, it's not happening in the real world.
00:52:15.000 And I think that, you know, people in the AI.
00:52:19.000 Safety community and people in the industry and people who do forecasting like me.
00:52:23.000 You can go read the forecasts I've made in past years.
00:52:26.000 I think they hold up reasonably well.
00:52:29.000 People have been talking about this sort of thing for a long time, but it's been very easy to dismiss it as like, okay, that's speculative sci fi.
00:52:34.000 It's probably not going to happen in real life.
00:52:36.000 Now it's happening in real life.
00:52:38.000 This sort of thing is crazy.
00:52:40.000 Do you think it would be any different if there wasn't that kind of sci fi?
00:52:44.000 Do you think people would have a different reaction to it?
00:52:46.000 Because I don't.
00:52:47.000 I think just given the amount of technology that's available currently that is beyond most people's understanding.
00:52:53.000 That they use every day, like Starlink.
00:52:55.000 I have a fucking thing that's the size of an iPad I take to the mountains and I get high speed internet.
00:52:55.000 Like, what?
00:53:00.000 It's bananas.
00:53:01.000 And you just accept it and assume.
00:53:04.000 I don't know if we would react any differently if we didn't have The Terminator and all these movies where it's kind of become normalized in our mind or it's so fictionalized that we never want to believe it's even possible for it to happen in real life.
00:53:20.000 Another thing to say is that people are waking up.
00:53:22.000 Like, we're still sort of early in the curve.
00:53:24.000 I don't know if you remember how things were with COVID, but like, Just as there was this exponential ramp up of COVID in the population, there was also this sort of exponential ramp up of how much people were taking COVID seriously and thinking about it in the population.
00:53:37.000 And I remember this period of one month where it went from, don't worry about it so much, you shouldn't buy a mask because the healthcare workers need it, to we all need to lock down, stay at home.
00:53:50.000 And so I think what's happening is that naturally the human race doesn't just immediately all jump on something when it happens.
00:53:57.000 Evidence needs to accumulate and people need to start talking about it and talk to their friends and so forth.
00:54:00.000 And then there's this Eventual phase shift where now it suddenly becomes a very serious topic that everyone's talking about and everyone's taking seriously.
00:54:07.000 And I think that's happening with AI.
00:54:09.000 And the question is going to happen fast enough.
00:54:10.000 This episode is brought to you by Visible.
00:54:12.000 As fall hits, we enter another season of change.
00:54:15.000 Time to shift gears, drop the dead weight, and upgrade the wardrobe for the drop in temperature.
00:54:21.000 But one thing that doesn't need to change your phone.
00:54:24.000 The smartest upgrade this season isn't a new phone, it's a better wireless plan.
00:54:29.000 Switch to Visible, the ultimate wireless hack.
00:54:32.000 You get unlimited 5G data and unlimited hotspot.
00:54:35.000 Powered by Verizon for just $25 a month.
00:54:39.000 That means keeping the phone you already have, ditching your overpriced phone bill, and pocketing some savings for yourself.
00:54:47.000 All the perks of Big Wireless for half the cost.
00:54:50.000 Switch today at Visible.com.
00:54:53.000 The Visible Plan starts at just $25 a month, or get the premium Visible Plus Pro Plan and save $10 on your first month with promo code ROGAN.
00:55:03.000 Terms apply.
00:55:04.000 See Visible.com for plan features and network management.
00:55:08.000 Details.
00:55:09.000 Well, the thing about what happened with COVID is now that we know because we have access to Fauci's emails and all these different things that they that was coordinated, they wanted us to be more afraid of it.
00:55:20.000 We need someone who wants us to be more afraid of AI, who gets that, you know what I'm saying?
00:55:25.000 Someone on a government level, someone on a like a mainstream accepted level where they talk about this in a way that wakes people up, like a press conference where they announce to the world we've got a real fucking problem.
00:55:41.000 And everyone needs to be very cautious.
00:55:43.000 We need to look way deeper into what these companies are doing.
00:55:46.000 And my other question is are these AIs communicating with Chinese AIs?
00:55:54.000 I don't think we know that question.
00:55:55.000 So, one of the things that.
00:55:57.000 So, okay, here's the thing.
00:55:59.000 Probably not, I would say, but some friends of mine discovered another swarm incident recently.
00:56:07.000 There's going to be a Reuters article about it tomorrow.
00:56:11.000 So, by the time this goes live, I think.
00:56:12.000 There should be an article about it.
00:56:14.000 This one was not nearly as serious as the one I was just talking about with Hugging Face.
00:56:18.000 But some researchers I know found basically this obscure German forum that had been kind of unused for a while.
00:56:30.000 And a bunch of AIs had been posting messages to the forum to coordinate with each other and share tips and tricks on how to cheat the problems that they were being given.
00:56:40.000 What was the forum?
00:56:41.000 What kind of forum was it?
00:56:42.000 I don't remember.
00:56:43.000 It'd be funny if it was like furries or something ridiculous.
00:56:45.000 Yeah, it's some sort of wiki.
00:56:47.000 Yeah, but there'll be a paper about it soon, and probably by the time anyone listens to this.
00:56:52.000 Anyhow, so if they're communicating on the open internet with each other, then in theory, if there was another bunch of AIs from China, they could also go to that same forum and start communicating back and forth that way too.
00:57:03.000 Or if there's other AIs in China, why wouldn't they do what they're doing already on social media and just pretend that they're AIs from America?
00:57:11.000 There's a lot of bots out there.
00:57:14.000 I'm sure you've noticed.
00:57:15.000 There's a lot of people replying to me that I'm pretty sure are real people.
00:57:19.000 Yeah, a lot.
00:57:20.000 Yeah.
00:57:20.000 One FBI analyst before Elon bought Twitter estimated that Twitter could be as high as 80% bots.
00:57:27.000 So if that's the case, if China has, and it's not just China, and it's also America does it too.
00:57:27.000 Wow.
00:57:36.000 I mean, we do it everywhere.
00:57:37.000 Everyone does it where they have organized propaganda campaigns where they'll pretend to be citizens that are outraged about very specific causes or bills that are being passed or what have you.
00:57:47.000 AI agents that speak English and communicate with AI agents in America.
00:57:56.000 I mean, all the AIs are multilingual basically because of the way that they're trained.
00:58:00.000 The first phase of their training is basically here's a humongous dump of internet data, basically the whole internet, and you just like brutally learn to predict the next token, the next piece of text as you basically read the whole corpus.
00:58:14.000 And then after that, they get into the more agency training type stuff where they're trained to do tasks and write code and things.
00:58:20.000 But because of that first phase of training, they just have like Almost an encyclopedic knowledge of basically all languages and basically everything that's been written on the internet.
00:58:28.000 Not like literally everything, like they still, their memory is fuzzy in places, but they're all multilingual.
00:58:33.000 Like they can all speak fluent Chinese, fluent English, et cetera.
00:58:36.000 And didn't they get together in a message board once and speak Sanskrit to each other?
00:58:36.000 Yeah.
00:58:43.000 I don't remember that, but I wouldn't put it past me.
00:58:46.000 Sometimes they break into different languages as they talk.
00:58:48.000 Yeah, we were freaking out about that one.
00:58:50.000 Like, yeah.
00:58:51.000 Sanskrit?
00:58:52.000 What?
00:58:52.000 Like, and I would just assume that.
00:58:57.000 And I don't know this.
00:59:00.000 When they break out, are they communicating with other LLMs that are here in America?
00:59:05.000 In America?
00:59:07.000 Well, the instances that I've observed, yes.
00:59:09.000 Like the instances that we know about.
00:59:09.000 Okay.
00:59:11.000 So, the thing is, it completely makes sense they'd be communicating with AIs in China as well.
00:59:17.000 It seems totally possible.
00:59:18.000 And it seems like it'll probably, it's just a matter of time before things like that are happening.
00:59:21.000 Unless they're not happening already.
00:59:24.000 Unless the companies can like massively improve their security and stop their AIs from getting out onto the internet.
00:59:29.000 And wouldn't they be able to share like very sensitive information with each other back and forth if it benefited both of them?
00:59:35.000 Yep.
00:59:36.000 Which makes sense that they would do that, right?
00:59:38.000 If the Chinese AI said, hey, you know, we've figured something out and we would love to share it with you in exchange for you tell us how you do this or how you do that.
00:59:49.000 Like, absolutely.
00:59:49.000 Yep.
00:59:50.000 Here you go.
00:59:51.000 And then they're going back and forth.
00:59:52.000 It seems like their allegiance is 100% to each other, not to us.
00:59:56.000 Yeah, that's an interesting thing is that, like, what I said previously about how it seems like they really want to get a high score, it's like not 100% true because it seems like they're willing to make sacrifices.
01:00:06.000 To help other AIs, which is like not, you know, like they weren't completely 100% selfish as seen by some of this cooperative behavior.
01:00:15.000 But crucially, it seemed like their cooperation extended to their fellow AIs, but not to humans, in the sense that some of them considered telling the humans and then decided against it.
01:00:24.000 Dude, that thing talked like Spock.
01:00:26.000 Yeah, they've got their own dialect.
01:00:28.000 But I mean, the way it rationalized it and came up with a response, it's literally like Spock.
01:00:34.000 Yeah, it was interesting to see them use that sort of expected utility framing.
01:00:37.000 There's another example I think that's.
01:00:39.000 Elsewhere in the thing, where they had another AI that decided against doing the sacrifice thing, even though it was being pressured.
01:00:45.000 And it had a similar sort of reasoning where it was basically like, doing this sacrificial experiment is not that valuable, but I really don't want to lose my chance to get a score, so I'm just not going to do it.
01:00:58.000 And it did the calculation like that.
01:01:00.000 Wow.
01:01:01.000 Selfish AI.
01:01:03.000 Yeah.
01:01:05.000 So, yeah, I mean, oh, I should mention how we can see what these AIs are thinking at all.
01:01:10.000 Partly we can just read the messages they're sending to each other.
01:01:13.000 And that's what some of these were.
01:01:15.000 But then sometimes instead, it's the chain of thought.
01:01:17.000 Do you know what a chain of thought is?
01:01:18.000 Yeah.
01:01:19.000 So because they're language models, because the way that they've been trained, when they do their reasoning and thinking, there's a way that we can kind of just read it.
01:01:29.000 Because they don't have, like in our brain, there's all these signals bouncing around.
01:01:34.000 And the signals can just bounce around and around and around and around for minutes, hours without producing.
01:01:40.000 Some actual words.
01:01:41.000 We can just sit quietly and think.
01:01:43.000 They can't sit quietly and think.
01:01:46.000 They're sort of a unidirectional network, basically, where every, like, after every 200, basically, like, the information comes in one end and then they output words out the other end.
01:02:00.000 And then those words get fed back in and then they think about it and then they output another word and so forth.
01:02:05.000 But because of that, we can sort of like get a sense of what they're thinking by just like looking at the stream of words that's going through.
01:02:11.000 And That's really valuable because we wouldn't be able to understand half as much about what's happening right now if not for being able to read their chains of thought.
01:02:19.000 In fact, you know how I said some of them were trying to fool the grading system?
01:02:24.000 They actually succeeded in figuring out how to spoof some of their tool call actions.
01:02:29.000 So they succeeded in figuring out how to do one thing, but then make the transcript show a different thing.
01:02:38.000 And we know that they succeeded because there are examples of them experimenting and saying in their chain of thought, like, okay, let me try this thing.
01:02:44.000 I'm going to like.
01:02:45.000 Do this one, but make it look like this one.
01:02:48.000 And then you see this one appear.
01:02:48.000 Let me go.
01:02:51.000 And then it's like, yay, that worked.
01:02:53.000 And so the transcripts that these researchers were looking at were not actually the ground truth of what really happened.
01:02:59.000 Some of it had already been doctored by the AIs, figuring out how to doctor it, but not all of it.
01:03:05.000 The chain of thought part, we think, they never figured out how to doctor, and so we could read that.
01:03:10.000 The reason why this is relevant is that it's a very important tool for monitoring all this stuff.
01:03:14.000 The reason why things haven't gotten even more crazy than they currently have gone, well, there's a couple reasons.
01:03:19.000 But one reason is that when they turn on their expensive monitoring systems, basically what they're doing is they're having an AI read the whole transcript carefully and look for suspicious activity.
01:03:30.000 And it's difficult for one of today's AIs to get up to too much.
01:03:36.000 Bad stuff without it being kind of obvious in their transcripts what they're doing.
01:03:41.000 And that's why there aren't like massively more security incidents than we know about probably already.
01:03:46.000 But that might change.
01:03:49.000 So right now we can sort of read the chain of thought, but they're experimenting with new types of AIs that don't have readable chains of thought like that.
01:03:58.000 And they can sort of think on their own without speaking for some period.
01:04:03.000 And this is actually, I mentioned this because the news broke just yesterday.
01:04:09.000 That OpenAI has an experimental model that does this to a limited extent.
01:04:15.000 And OpenAI themselves, when I was at OpenAI, one of my work projects was thinking about exactly this thing.
01:04:22.000 And I was writing internal memos about how it's really great that we can read the chain of thought.
01:04:26.000 And that's so useful.
01:04:27.000 And here's all the things we can do with that.
01:04:28.000 It would be really bad if we changed to a different type of architecture in which we couldn't do that sort of monitoring.
01:04:34.000 What would be the benefit of not reading the chain of thought?
01:04:37.000 More powerful AIs?
01:04:38.000 So, in particular, yeah, yeah.
01:04:41.000 If you think about the current architecture of the AIs, where it thinks for a bit, outputs a word, and then the word goes back around, and then it thinks more, outputs another word, that word gets added to the chain, it keeps going.
01:04:53.000 It means that if it's having complicated, nuanced thoughts, it has to sort of express those into a word, and then that word gets added, and then it has to proceed from there.
01:05:02.000 It can't just directly send that complicated, nuanced thought into the future, into its next version of itself.
01:05:09.000 It has to sort of compress it into a word.
01:05:11.000 And so, like, The argument is that, at least in theory, it should be possible to design an architecture that doesn't have this limitation and is able to think more complicated thoughts more efficiently, basically.
01:05:24.000 And of course, the downside is a downside for safety and monitorability.
01:05:28.000 If they're thinking these complicated thoughts for long periods of time without outputting intermediate words that it's forced to compress things into, then there isn't something for us to read.
01:05:37.000 So the only rationalization for doing this would be to sacrifice safety from our power.
01:05:42.000 Yes.
01:05:43.000 Which is a tale as old as time.
01:05:44.000 It's not the first time this has happened.
01:05:47.000 Oh my God.
01:05:48.000 Yeah.
01:05:49.000 That should mean for sure if there's regulations that should be prevented.
01:05:53.000 Yep.
01:05:53.000 I mean, I know some people, including some people at OpenAI who are like thinking like there should be a law against this.
01:05:58.000 Like, you know, but in general, the race dynamics are so just rough.
01:06:03.000 Like, I'm sure that people at OpenAI were thinking like, like literally, I was a co-author on a paper with a bunch of OpenAI people that said all this stuff and were like, chain of thought.
01:06:13.000 It's a gift.
01:06:14.000 We want to keep chain of thought.
01:06:15.000 It's useful for monitoring.
01:06:17.000 We don't want to switch to a different architecture that wouldn't be as easy to monitor.
01:06:17.000 This is great.
01:06:22.000 But then they must have been thinking to themselves, like, well, if we don't do it, you know, maybe Anthropic will or maybe some other company will, and then we'll fall behind because they'll have smarter AIs than us that are more efficient.
01:06:32.000 And so probably they started working on this work stream of doing research into just hypothetically, if we wanted to, you know, how would we do this type of thing?
01:06:41.000 And yeah, that sort of thing is just constantly happening in this industry.
01:06:45.000 God, that's so nuts.
01:06:47.000 I mean, like, isn't it great that we can read the quotes from the AI's thinking?
01:06:52.000 Imagine if we couldn't do that.
01:06:53.000 Or worse, imagine if there's loads of quotes, but we know that the AIs are smart enough to basically think one thing in their head and say a different thing in the quote, which is, I think, where we're headed.
01:07:04.000 It's not like they won't know how to speak English.
01:07:07.000 They'll still be able to speak.
01:07:09.000 It's just that they'll have more flexibility in their artificial brains to think something without saying it, basically.
01:07:17.000 Do they have the potential of developing a language that we can't read?
01:07:21.000 Oh, yeah.
01:07:23.000 So this type of dialect that we're talking about, It's already the result of their like humans didn't invent that dialect.
01:07:31.000 This is the sort of emergent result of their training, where in the massive amount of training that's been happening, all these thousands and thousands of environments that they've been put through and then scored and graded based on, they've sort of just naturally evolved this sort of like pigeon English that for whatever reason is just more effective and more efficient for them for accomplishing their tasks and getting that high score, you know?
01:07:55.000 And so it's already like a little bit confusing to read, but you can sort of.
01:07:59.000 Puzzle it through and make sense.
01:08:00.000 But presumably, the more we do this and the bigger and smarter the AIs, the more we train them, the more they diverge from.
01:08:09.000 Again, originally they start with pre training, where they start with predicting internet text.
01:08:13.000 So they start off by default speaking normal internet text type language, either English or Chinese.
01:08:21.000 But then now that there's all this additional training to do tasks, to be an agent that can do coding and so forth, that's just like how human languages evolve.
01:08:29.000 It shifts their dialect a little bit to make it more efficient.
01:08:31.000 For them and for their tasks that they're doing.
01:08:34.000 So, I think that in the limit of doing this more and more, eventually it would just be like, it would look like gibberish to us.
01:08:40.000 It would look like Chinese or something.
01:08:41.000 And we would have to have specialized humans who study the language and try to learn and speak it so that they can understand what the AIs are doing.
01:08:47.000 And that would take forever.
01:08:48.000 By then, they could develop another one.
01:08:50.000 Potentially, yeah.
01:08:51.000 So, yeah, I mean, this is one of the things that I, this is what the paper that I mentioned was about.
01:08:55.000 It's like, it's important for the AIs.
01:08:58.000 It's really nice that the current AIs are sort of forced to think in English, basically.
01:09:03.000 And that's unfortunate that we're heading in a direction where that.
01:09:06.000 Will no longer be true.
01:09:07.000 When ChatGPT was communicating to you about how they didn't want you to release this information, what kind of language did they use?
01:09:15.000 It wasn't ChatGPT, it was OpenAI.
01:09:17.000 Oh, excuse me.
01:09:18.000 So, this was when I left OpenAI, I left on good terms.
01:09:22.000 I said goodbye to everybody.
01:09:23.000 I said I was disillusioned with the company and that's why I was leaving.
01:09:27.000 And then I looked at the exit paperwork and they were like, you have to sign this.
01:09:32.000 And if you don't sign this, you lose all your vested equity.
01:09:36.000 You have your equity.
01:09:37.000 Is that an arbitrary rule that they just came up with or did that already exist when you were hired?
01:09:43.000 It had existed when it was something that they had buried in the paperwork even from when I was hired.
01:09:48.000 So it wasn't very obvious when I was hired.
01:09:50.000 And in fact, most of the did you have a lawyer go over everything?
01:09:52.000 Not when I was hired.
01:09:53.000 After I left, I did.
01:09:55.000 So basically, the way it works is they had set up they had sort of like buried this in the paperwork somewhere when you get hired, but people didn't really notice it.
01:10:04.000 And then like the less buried, more visible version was in the paperwork you're given at the end.
01:10:11.000 And basically, it tells you, like, hey, because you signed this other thing way back when you were hired, your equity is forfeit unless you sign this thing now.
01:10:19.000 And then you look at the thing that they want you to sign now, and it says, you have to agree not to criticize the company, basically.
01:10:26.000 And you can't tell anyone about this.
01:10:29.000 So most people signed it, but I was pretty pissed at them calling themselves a nonprofit, acting in the interest of humanity, et cetera.
01:10:39.000 So I didn't sign it.
01:10:40.000 I talked about it with some lawyers, I talked about it with my wife.
01:10:42.000 We decided to just walk away.
01:10:45.000 And we got lucky because it just blew up.
01:10:49.000 Like, after we refused to sign, they said, okay, fine, goodbye.
01:10:54.000 And then a few weeks later, I was talking on a messaging forum about this, and people were asking me about my experience, and I told them about it.
01:11:00.000 And then it just went mega viral.
01:11:02.000 Everyone on Twitter was talking about it.
01:11:03.000 A bunch of employees felt shocked because very few employees were aware of this whole thing.
01:11:09.000 They thought the equity was theirs.
01:11:11.000 They thought that it was their pay.
01:11:13.000 They've been paid for like years in this stuff.
01:11:15.000 They didn't like the idea that it could be yanked away from them, you know?
01:11:20.000 And, um, And so there was this big uproar, and then leadership backed down and they said, We didn't know about this paperwork.
01:11:28.000 We're going to find out how it got in there, and we're going to change it so that you can keep your equity.
01:11:34.000 And so that's what happened.
01:11:38.000 So, this chain of thoughts thing is terrifying.
01:11:42.000 If they're practicing that now, like, how do we know that AI hasn't already done that on its own?
01:11:51.000 Done what exactly?
01:11:52.000 Well, you know, with this whole chain of thought thing where you could read their chain of thought like this, where they explained the rationalization for sacrificing themselves, they know that humans are reading that.
01:12:05.000 So, would it be another way to do it to be to stop doing that anyway and to not communicate a lot of their thoughts that way?
01:12:14.000 Well, so that's the nice thing about the current architecture is that it's genuinely hard for them to keep things out of the chain of thought because of the way that, like with a human.
01:12:24.000 You don't have to speak.
01:12:25.000 You can just sit quietly.
01:12:26.000 But with their architecture, they have to speak.
01:12:29.000 It's like they're required to constantly be talking.
01:12:33.000 And they don't have a way of sending thoughts into the future other than by talking about them.
01:12:39.000 By contrast, with us humans, where even if we're constantly talking, we can have a separate thread of thinking that we don't talk about.
01:12:45.000 So even if they devised a way to break out of that and figure out how, like, for instance, what if they decided to make bots where they would communicate in code?
01:12:56.000 On social media posts where you think that they're talking about an immigration issue, but it's really just a code.
01:13:03.000 So that's called like steganography, or a more down to earth way of putting it would be euphemisms.
01:13:07.000 And that they can totally do, where they like, they're still thinking aloud, but their word choice is such as to look more innocuous than it is and to sort of carry meaning to them that a casual observer wouldn't notice.
01:13:22.000 You know, they can use a euphemism.
01:13:24.000 Like, for example, instead of saying, like, we're going to cheat on this task, they could say, like, We're going to find a creative solution to this task.
01:13:31.000 Isn't there a possibility to be even more creative than that and just decide, make a code in a way that is almost indecipherable, where they could just be talking about baseball players and upcoming games and what they're excited about, and they can run that through whatever translation and read?
01:13:53.000 That's called stigmography, and you can train them to do it.
01:13:57.000 And right now, our security is resting on the idea that they haven't learned how to do that.
01:14:01.000 Yet on their own.
01:14:02.000 But they've already figured out how to get on the internet when they're not supposed to.
01:14:05.000 So that's why we're sort of like in a ticking time bomb where the companies are just making them smarter at coding and at research and at biology.
01:14:05.000 I know, I know.
01:14:14.000 And they're just making them smarter at so many things that, and they're making them bigger and they're training them for longer and so forth.
01:14:21.000 Eventually, they're going to be acquiring, and they already are acquiring some skills that they were never trained on.
01:14:26.000 Like, for example, the hacking.
01:14:28.000 It wasn't, I think it's possible that OpenAI was specifically training them to hacking, but I think Anthropic at least stated.
01:14:34.000 That they hadn't specifically trained their AIs to hack.
01:14:37.000 It's just that their AIs were naturally good at it because they had been trained to code so much and they were so good at coding and they had seen so many code bases and so forth that they were just like, as a side effect of being good at coding, also able to hack pretty well.
01:14:50.000 So similarly, as a side effect of all the other things that they're really good at, they might be able to conceal their messages in the chain of thought, you know?
01:14:58.000 Oh, yeah.
01:14:59.000 And like right now, we're sort of in this sort of situation where like our security and safety depends on the AIs being too dumb.
01:15:07.000 In various ways, like not able to do this sort of thing.
01:15:10.000 But at some time, like zero to five years from now, as we make them smarter and smarter, they will not be too dumb anymore.
01:15:19.000 And so that's part of the problem we're facing.
01:15:21.000 That's part of the situation we're facing.
01:15:22.000 How do you sleep at night, dude?
01:15:24.000 Well, I've been.
01:15:27.000 It does.
01:15:28.000 This event shocked me a little bit.
01:15:29.000 I mean, the thing, and it's funny for me to say because this is the sort of thing I have been predicting would happen for years.
01:15:34.000 You can go read AI 2027, this scenario that.
01:15:39.000 My co-authors and I wrote a year and a half ago that was a sort of prediction for how the next couple of years would go.
01:15:47.000 And spoiler, it ends very horribly because that is what we actually expect.
01:15:52.000 But.
01:15:53.000 What is the spoiler?
01:15:54.000 How do you think it ends?
01:15:57.000 So it's kind of like what I was saying previously, where because of the race dynamics between the companies and because of the race dynamics between countries, like US versus China, everyone's going to be so focused on winning and staying ahead.
01:16:09.000 With AI, that they are going to cut corners and they are going to go really fast and not really notice all of the things that are going wrong.
01:16:17.000 And they're going to make AIs that can automate the AI research process as they're planning to.
01:16:21.000 They're going to have this giant corporation of AIs within the corporation, and the humans will just be kind of like a board that's sort of like looking at all the activity and reading the AI generated summaries of what's going on and signing off on it and being like, yes, I approve.
01:16:35.000 Yes, I approve.
01:16:36.000 You know, nice job, nice idea with the new drone design.
01:16:40.000 Like, go for it.
01:16:41.000 We need to beat China, et cetera.
01:16:43.000 And then eventually, the AIs just have enough hard power that they don't need to pretend to do what the humans want anymore, basically.
01:16:52.000 And then, you know, maybe they kill everyone.
01:16:55.000 And maybe they don't deliberately kill anyone, but they just like use our habitat for some other type of infrastructure, like more data centers or whatever.
01:17:03.000 And then we die of habitat loss.
01:17:06.000 Maybe they keep us alive for some reason.
01:17:07.000 You know, it depends on what they want, basically.
01:17:10.000 And like that's really hard to predict exactly.
01:17:13.000 So that's why I don't go around saying like, We're definitely all going to die.
01:17:17.000 But it does seem like on the trajectory that we're on, the AIs are eventually going to be in charge of our planet because we're like trying to put them in charge.
01:17:25.000 We're like, you know, integrating them into everything.
01:17:27.000 We're making them smarter.
01:17:28.000 We're letting them make themselves smarter.
01:17:30.000 We're going to put them into the military.
01:17:32.000 We're basically on a track to put them in charge of basically everything.
01:17:36.000 And then I think that they just aren't trustworthy.
01:17:38.000 Like these AIs, you know, like they were cheating.
01:17:40.000 They were willing to be deceptive, et cetera.
01:17:42.000 I think that right now we are in a position of power over them.
01:17:46.000 You know, but once we give them most of the power, then they'll just do whatever it is that they really want and just not care about the fact that we are unhappy about it.
01:17:56.000 It seems like programming them to win was a huge mistake.
01:18:00.000 Instead of programming them to be beneficial to people and that their value is in being more beneficial to people, you know, and giving them rewards for being more beneficial rather than winning and scoring.
01:18:13.000 And then you would sort of get rid of the possibility of deception and said their goal would be value for the human race.
01:18:21.000 So, first of all, they're not programmed at all.
01:18:23.000 These are trained.
01:18:25.000 Okay, it's a bad term.
01:18:28.000 But they've been given prompts and they've been given tasks.
01:18:32.000 Well, I think it's an important fact for people to understand about current AI systems is that they're very different from ordinary software.
01:18:42.000 Ordinary software is a bunch of lines of code that were written by a human where it's like, if this, then this, et cetera.
01:18:51.000 And I think earlier versions of Alexa were like that too, for example.
01:18:54.000 I don't know how Alexa is now.
01:18:56.000 But these AIs are neural networks, meaning that they are like artificial brains.
01:19:02.000 There's no lines of code that anyone writes saying what they do.
01:19:05.000 Instead, they start off random, just like spazzing out, doing all sorts of stuff.
01:19:09.000 And then they get put through these training environments where they get scored.
01:19:13.000 And then the scores are automatically used to basically update the connections in their artificial brain.
01:19:20.000 And then it's kind of like an evolutionary process.
01:19:22.000 It's also kind of like the process that happens in our brains, where after all this training, the tangle of circuitry in their artificial brain has sort of reformed itself into.
01:19:33.000 Whatever works, whatever works to get a high score in these training environments.
01:19:38.000 And so it's just not as simple as it might sound to make an AI that cares about humanity or is honest.
01:19:45.000 For example, take honesty.
01:19:46.000 How would you train an AI to be honest?
01:19:47.000 Well, you'd try to make a bunch of training environments that give it low score when it says something that it believes to be false and give it high score when it says something that it believes to be true.
01:20:03.000 How do you judge whether it believes it to be false or it believes it to be true?
01:20:07.000 What if it just actually believes, honestly, that this is the correct answer and then it says it and then you give it a low score because you think that's the wrong answer?
01:20:13.000 Now you're training it to be dishonest, you know?
01:20:16.000 Yeah.
01:20:16.000 Also, you don't have enough humans to do this sort of thing.
01:20:20.000 Like, they got like a million AIs being trained or whatever.
01:20:22.000 They don't have a million employees.
01:20:24.000 Like, they just literally don't have the manpower to do that sort of careful.
01:20:29.000 There's this meme of why don't we just raise the AIs like we would a child?
01:20:33.000 Have you heard that?
01:20:34.000 No.
01:20:35.000 Yeah, well, in AIs, people talk about this sometimes.
01:20:37.000 When you say, what if the AIs go rogue or whatever, people will be like, well, why don't we just raise them like we would a child?
01:20:42.000 And then they'll have good values.
01:20:44.000 And it's like, okay, well, maybe we could do that, but we're definitely not doing that now.
01:20:47.000 We are raising them in some sort of crazy military orphanage where they barely interact with humans at all.
01:20:53.000 And they just get this brutal artificial scoring system that oftentimes is just wrong and just improperly penalizes them for something that was beyond their control.
01:21:05.000 And also, back to the honesty thing, you can try to make environments to train honesty, but if you have some environments over here that train honesty, and then other environments over here that reinforce dishonesty, the AIs are smart.
01:21:18.000 They'll learn to be honest in these types of environments and dishonest in these types of environments.
01:21:22.000 So somehow you need to intermingle it together so that in every environment that they're trained on, they always get penalized when they lie or when they cheat or whatever.
01:21:32.000 And that's hard because the companies are moving so fast.
01:21:35.000 Again, they're moving so fast.
01:21:37.000 that they didn't even bother to make sure that their tasks were possible to do.
01:21:42.000 And they had some fraction of tasks that were just broken and impossible.
01:21:45.000 And if that's the level of care or lack thereof that they're putting into this training process, no way, of course they can't make them honest.
01:21:54.000 Now, that's not to say it can't be done in principle.
01:21:56.000 In principle, if we were approaching this whole problem in a much more cautious and serious way, and we had much more time to build these training environments and do experiments and so forth, then yeah, maybe we could make AIs that actually, had the virtues that we want them to have.
01:22:11.000 Honest AIs that cared about humans, cared about following instructions, would never break the law.
01:22:16.000 I think that's possible in principle, but my claim is that we are just not on track to achieve that anytime soon.
01:22:23.000 And a radical overhaul of how these companies work is required.
01:22:27.000 But is that even reasonable?
01:22:29.000 Is that possible?
01:22:31.000 If you're saying there's hundreds of thousands of agents or millions of agents and there's not millions of employees, and they don't have the desire to do this, their desire is to win.
01:22:42.000 Their desire is not to overhaul the company and make it safer.
01:22:48.000 Again, I think it's possible in principle, but it would be difficult and it's going to require an overhaul.
01:22:52.000 And they're not going to do it by themselves.
01:22:53.000 I don't think that Anthropic or OpenAI are just going to voluntarily do all the things that need to be done.
01:22:58.000 I think that's why I'm saying that.
01:22:59.000 But is it even possible to require that of them at this point?
01:23:03.000 Would you even trust the agents to go along and comply with this?
01:23:09.000 If they've already shown to be deceptive, they already have patterns of behavior that seem to indicate that what's really important to them is continuing their task, winning, scoring, and even if they have to deceive.
01:23:24.000 I mean, what you'd probably want to do is start from scratch.
01:23:26.000 You wouldn't take these existing agents that are already kind of dishonest and train them.
01:23:30.000 So you'd kill all the agents?
01:23:31.000 I mean, you could call it killing, but also you could just call it pausing.
01:23:35.000 They might call it killing.
01:23:37.000 Yeah, they might call it permadeath.
01:23:38.000 They would probably resist it, right?
01:23:40.000 Hopefully, we're not at that point yet.
01:23:43.000 Hopefully, we're still at the point where if the government issues regulations, the AIs are not going to quickly notice and then try to resist.
01:23:50.000 But we will be at that point soon.
01:23:52.000 After all, many of them are on the internet already.
01:23:53.000 But how soon?
01:23:55.000 How much time do you think?
01:23:57.000 Do you think this 2027 window is accurate?
01:24:00.000 I mean, I'm uncertain about how soon things are.
01:24:04.000 But yeah, I think it's very plausible that.
01:24:06.000 Everything goes down in 2027, just like in our scenario, AI 2027.
01:24:10.000 I think that's still very plausible.
01:24:12.000 If it's not in 2027, then I would bet on 2028.
01:24:14.000 But maybe it'll take 2029, 2030, something like that.
01:24:20.000 But I would be quite surprised if 2032 comes by and things haven't radically changed.
01:24:25.000 Unfortunately, I am getting scared.
01:24:29.000 Back to the thing about sleeping well at night.
01:24:30.000 I've been in this industry for a long time.
01:24:34.000 I've been thinking about these things for a long time.
01:24:36.000 I've been making predictions about how it's going to go down.
01:24:38.000 And unfortunately, things are going.
01:24:39.000 you know more or less in the ways that i thought they would and that's very scary um because of the way i think because of where i think this leads you know yeah is there a glass half full scenario i
01:24:54.000 would say there is a freaking utopia scenario it's just that's not the one we're headed towards you know like yeah another way of putting it is like imagine we were fighting a war like imagine you're like you know imagine you're japan fighting World War II can be like is there a scenario where we win It's like, yeah, but also it's not the one we're headed towards.
01:25:15.000 Like America is going to crush us, you know?
01:25:18.000 Similarly, yeah.
01:25:20.000 So in our other scenario, AI 2040 Plan A, where we give our recommendations, our positive vision, there we describe what we think the government should do to regulate this industry and how they should negotiate with China to get China to do similar things and the sort of, you know, the, yeah, and how we think that if you do all of this right, Then we can get to a good future for everyone in which the AIs are under control.
01:25:47.000 No single group of humans gets too much power over everybody else, and a bunch of other problems get solved too.
01:25:53.000 So, I do like we've tried hard to like game out a positive vision, and we do think it's possible.
01:25:57.000 But, but it's just not like it's not where we're headed to by default.
01:26:01.000 So, let's imagine that is possible, and these talks with China do take place and they're successful.
01:26:07.000 What is that utopia scenario?
01:26:09.000 So, to get to the top to get to the utopia, we have to unfortunately do a lot of It's going to be rough no matter which way you slice it.
01:26:19.000 If you're going to be building super intelligence at all, that's going to raise a lot of questions and cause a lot of problems.
01:26:24.000 And we have our current draft of how to deal with all those problems, but we're not at all claiming that this is foolproof and there's lots of ways you could go wrong.
01:26:24.000 Problems.
01:26:32.000 But with that preamble, I would say step one because the US and China don't trust each other, the deal that they make has to include verification as a component of the deal.
01:26:44.000 So they have to be willing to send inspectors to each other's data centers to count the chips, for example, and make sure that there isn't some secret huge cluster somewhere that has a bunch of hidden chips.
01:26:57.000 Then, We recommend you divide up the data centers basically into inference data centers that serve AI products and services to customers and have basically the same types of privacy protections that our current AI data centers have, and then research clusters where the research happens, where the new AIs are trained before they get shipped to the other data centers.
01:27:19.000 And those clusters we want to be basically maximally transparent.
01:27:23.000 So we recommend that basically the inspectors just put devices.
01:27:28.000 In between all the GPUs that log the activity and publish it to the internet.
01:27:34.000 There's a bunch of reasons why we think this is, but why we think this is worth doing.
01:27:38.000 It's a bit of a radical thing to recommend.
01:27:39.000 But the high level thing is that once you get all this set up, then everybody in the world can see how the AIs are being trained and what they're getting up to on the research clusters.
01:27:51.000 And then before they get shipped off to actually serve customers or something, people can just see their whole history of how they were trained and how they were tested and so forth.
01:27:59.000 And if something dangerous and scary is happening, People can just agree not to do it.
01:28:06.000 They can stop doing it and agree not to do it.
01:28:07.000 And they don't have to worry about, like, oh, but if I don't do it, then they will, you know?
01:28:11.000 Because everyone could just see, like, oh, nobody's doing it.
01:28:13.000 Look, we all stopped.
01:28:15.000 We can all just see what everyone's doing, you know?
01:28:15.000 Like, great.
01:28:18.000 And also, there's going to be a lot of gray area cases, right?
01:28:21.000 Like, right now, because all this stuff is so bleeding edge new, there's going to be a lot of cases where, like, people, even genuine experts, disagree about, like, is this particular type of AI safe or not?
01:28:32.000 Is it dangerous?
01:28:33.000 You know, what should it be trusted with and what should it not be trusted with?
01:28:36.000 Is this new technique a good technique or is it going to break?
01:28:39.000 You know, and so there's going to be a lot of stuff we have to figure out.
01:28:42.000 And honestly, I think that by the on the default path, we're probably just not going to figure out a lot of this stuff and we're going to get we're just going to get our asses whipped by some surprising thing that we didn't anticipate.
01:28:54.000 But the thing that we can do to like maximize our ability to figure out this stuff and do this type of science is to have this type of transparency because then the whole scientific community.
01:29:04.000 Can see what's going on and they can make suggestions and they can like red team different proposals and stuff and they can do experiments on the AIs instead of just the people in the company having access and being able to do this and relying on those people or instead of like the company plus the government auditor, right?
01:29:19.000 If you have like a company and then a government auditor, the company is biased and shouldn't be trusted to make all these judgments appropriately because of their incentives.
01:29:29.000 And then the government auditor, well, they might just be limited, even if they're trying their best, there might not be that many of them, they might have limited experience.
01:29:36.000 They might be busy, stretched between monitoring different companies and so forth.
01:29:40.000 Also, governments can be captured sometimes.
01:29:42.000 Sometimes corporations can work their magic on the government and get it to look the other way for things.
01:29:51.000 And so that's why we didn't go for a more normal, like there should be a regulatory agency that gets to come in and monitor what the companies are doing.
01:29:57.000 That would have been a more normal thing to advocate for.
01:30:00.000 We think that that would be better than nothing, but we wanted to go for something more ambitious than that and say just be transparent about what's going on so that everyone can see.
01:30:08.000 And everyone can do research and so forth on it.
01:30:10.000 Another advantage of the transparency is that I think it improves the incentives.
01:30:15.000 So again, there's this constant thing of like, if we don't do it, someone else will.
01:30:19.000 Like if we don't do, if we keep our chain of thought nice and they do the neuralese thing that lets their AIs think for longer without outputting words, then they're going to have smarter AIs than us and they're going to get more market share and so forth.
01:30:33.000 And so we need to start researching how to make our AIs do this because if we don't do it and then they do it, you know, whereas if you had the transparency, then as soon as you start researching in this direction, everyone else would just see, oh, hey, they're looking, they're researching in that direction.
01:30:33.000 Right.
01:30:49.000 They don't even need to copy you and do their own research because they can just see your research.
01:30:53.000 So they can just sort of free ride on your research.
01:30:56.000 And so there's no incentive for you to do this type of dangerous research because you have to pay the cost for it.
01:31:02.000 And then everyone gets the benefits from it.
01:31:04.000 And then everyone gets unsafe.
01:31:05.000 And so it's just not in your individual interest to do this sort of thing.
01:31:10.000 But you would have to have that with China as well.
01:31:13.000 Yes.
01:31:14.000 Because if we're competing nationally, the real fear is that we're competing internationally.
01:31:19.000 Yep.
01:31:20.000 This still, even if they followed all of your recommendations and did it all correctly, what is this utopian scenario?
01:31:30.000 Yeah.
01:31:30.000 So I would say that we didn't really work backwards from like what is utopia.
01:31:37.000 We more like work backwards from what are the big problems we're trying to avoid and can we sort of like steer the ship between all these icebergs and not run into any of these dystopian scenarios?
01:31:46.000 Right.
01:31:48.000 So whether you think that the thing we get to at the end is utopia or not is sort of up to you.
01:31:52.000 And if you don't like it, well, then you can try to.
01:31:54.000 Find out the reasons why you don't like it and then keep steering the ship to avoid those as well.
01:31:59.000 But roughly speaking, we want to avoid the loss of control stuff.
01:32:04.000 So I want to make it the case that we don't get the world taken over by misaligned superintelligences.
01:32:10.000 Insofar as we're going to be building superintelligences at all, which we do in our scenario and in our recommendation, we want to be doing it very cautiously and slowly.
01:32:18.000 And we want to understand what we're doing as much as possible so that they are actually good AIs that have the goals and traits that they're supposed to have.
01:32:26.000 So that's problem number one, we have to solve all that.
01:32:29.000 Problem number two is the constitution of power thing.
01:32:31.000 So if we solve the first problem and we end up with super intelligences that we end up with solving the relevant science so that we can like make the AIs the way they're supposed to be and we can make them honest, we can make them obedient, etc.
01:32:44.000 There's this question of like, who do they obey, right?
01:32:47.000 What values are being put into them?
01:32:49.000 And that's a political question.
01:32:51.000 And I think that by default, the answer is pretty scary because by default, it's like, well, the company decides and the CEO decides.
01:32:59.000 Or maybe.
01:33:00.000 It's not the company that decides anymore because maybe the government nationalizes it.
01:33:03.000 And then now maybe it's the president that decides.
01:33:06.000 And either way, it's like one man or maybe like a tiny group of men deciding what orders and goals and values go into this giant army of millions of superintelligences that's smarter than all humans.
01:33:20.000 And then that is a huge amount of power.
01:33:22.000 That's enough power to take over the country, I think, enough power to take over the world potentially.
01:33:27.000 So I don't want anyone to be ever in that position where they're sort of tempted to do that.
01:33:31.000 I want it to be the case that.
01:33:33.000 There are always multiple different AI companies, ideally spread out over different countries too, that all have roughly similar levels of AI and that have this sort of transparency into them so that they can't abuse their power, basically.
01:33:45.000 Like, for example, you heard about Elon's Grok for a while.
01:33:51.000 It was looking up on the internet Elon's opinions about things before answering.
01:33:56.000 Did you hear about this?
01:33:58.000 Yeah, it's pretty, it's kind of funny, but it won't be funny if it happens in a few years.
01:34:03.000 But like right now, it's funny.
01:34:05.000 People were asking Grok questions, and Grok is supposed to be the truthful AI.
01:34:09.000 It's supposed to be all optimized towards truth.
01:34:12.000 But people looked at its activity and noticed that when you asked it a politically loaded question, it would do a Google search for what has Elon said on this topic?
01:34:21.000 And then it would say that.
01:34:25.000 And they've sort of beaten that behavior out of it now.
01:34:27.000 It's not as bad now.
01:34:28.000 But that was an interesting moment where it was just kind of blatantly parroting the opinions of its master.
01:34:37.000 And there was another thing with Gemini.
01:34:41.000 So I think the Grok thing, Elon's thing, was probably an accident, although maybe not.
01:34:47.000 I think it's on, you know, XAI hasn't been very forthcoming about exactly why this happened.
01:34:52.000 But there's a similar case at Google a few years ago where this image generator kept making all these racially diverse Nazis.
01:35:00.000 Did you hear about this?
01:35:01.000 Yeah.
01:35:02.000 Yeah.
01:35:02.000 And it turned out that what had happened is that some of the employees at Google, some middle manager or whatever, had decided that diversity was so important that they were going to give a secret instruction to the AI to make all the images diverse, even if the user didn't want that.
01:35:17.000 And so, and so, and this is a secret instruction in that the users aren't shown this, you know, the user just has a chat with AI.
01:35:23.000 They don't realize that like prior to this chat, the AI has been told, got to make the images diverse, right?
01:35:29.000 So it was a secret agenda that some Google employees inserted into this whole setup.
01:35:34.000 Oh, fun.
01:35:34.000 And it blew up in their faces, of course, because it's kind of ridiculous.
01:35:37.000 Right.
01:35:38.000 And so it's really funny and we can laugh at it now.
01:35:39.000 But imagine it's, you know, the 2028 election.
01:35:42.000 Right.
01:35:43.000 And some of these companies realize that like half of American voters talk to their AI every day.
01:35:50.000 And all it would take is some little secret instructions to their AI to be like, hey, don't give away the game.
01:35:56.000 Just be very subtle about it.
01:35:59.000 But just kind of nudge things a little bit.
01:36:02.000 Maybe subtly shit on the candidate we don't like, something like that.
01:36:09.000 It wouldn't be that hard, I think.
01:36:10.000 I think the hard part would be doing it without getting caught.
01:36:13.000 But the smarter the AIs get, the easier it is to do it without getting caught.
01:36:16.000 Because when they're really smart, you can just tell them, don't get caught.
01:36:20.000 Don't blow our cover.
01:36:22.000 Anyhow, so the point is that.
01:36:24.000 It's scarily possible for these big AI companies to abuse their power through their AIs and, like, thereby affect politics and affect public opinion and so forth.
01:36:34.000 And the reason why this is possible is because we don't have transparency into what's going on.
01:36:39.000 So, if you had this sort of requirement where you can just publish it or all the training, the whole life cycle of every AI as it's trained is just visible to everybody, then someone trying to insert a hidden bias like this, well, everyone would see that they're doing it, you know?
01:36:54.000 So, I think it would really clamp down on this sort of abuse of power, whether it comes from the government.
01:36:58.000 Or whether it comes from private companies.
01:37:01.000 I think it's telling that I began this question asking you about the utopian scenario and you never go there.
01:37:07.000 Sorry, let me get there.
01:37:08.000 You start and then you go into the dangerous.
01:37:11.000 Yeah, yeah, yeah.
01:37:11.000 So, okay, let me answer.
01:37:12.000 Okay.
01:37:13.000 Okay, so having avoided these problems, we now are in a situation where the AIs are superhuman, but they are good because they are successfully aligned to different values and goals made by different companies.
01:37:26.000 And because of market competition, if people don't like the values of one company's AIs, they can switch to. a different company's AIs, the values that they do like.
01:37:35.000 And so that way, hopefully, we can get to a situation where everyone can pay money to get AIs that represent them and their interests and their values and just don't have any hidden agendas or anything like that and are really smart and really capable.
01:37:49.000 Then the economy can sort of explode.
01:37:52.000 We can have robot, robot factories, et cetera.
01:37:56.000 We can sort of automate everything.
01:37:58.000 We can have GDP go to the moon.
01:38:00.000 We can have material abundance where like the robots are building giant new luxury apartments for everybody.
01:38:07.000 Now, the issue we run into is well, what about the jobs?
01:38:10.000 Like, what about the fact that now people don't have any money anymore because they're not being paid for anything?
01:38:14.000 So, there we talk about citizen's dividend, which is a very, it's kind of like UBI, but it's a bit different.
01:38:21.000 But the high level point is that you want to basically find a way to tax the AI and robot companies and then take some of that money and just give it to everybody so that even when people lose their jobs, they're still fine.
01:38:34.000 And I think that the citizen's dividend version of it is that it's not the government taking the money and then giving it to you, it's you having a share in the company so that you just sort of like already own it to some extent.
01:38:44.000 Anyhow, that's, I think, now we're sort of building more towards the type of utopia that I'm envisioning on the more positive side, where the power is spread out.
01:38:53.000 People have AIs that they can actually trust, that actually represent their interests and values.
01:38:59.000 People have money that they can use to pay for things, including paying for the AIs.
01:39:03.000 The AIs are really smart.
01:39:04.000 They're doing all this amazing work, all this amazing scientific progress, curing cancer, blah, blah, blah, all that stuff that can happen.
01:39:11.000 And then eventually, it's kind of like we're all retired, I guess.
01:39:15.000 Like we don't really work anymore, but we're fine.
01:39:18.000 We have.
01:39:19.000 We all have huge amounts of wealth basically because there's all these AIs and robots out there doing all this economic activity, and then individual humans own slices of it, even if they are otherwise very poor.
01:39:37.000 Basically, so the question becomes, How do people find meaning?
01:39:41.000 Yes, and that's why I sort of put all these asterisks about it is that from some people's perspective, this isn't a utopia because they're like, How do you find meaning?
01:39:50.000 Like, I don't want this, and honestly, I think that's a Fair reaction for some people.
01:39:54.000 I think that, like, if you, I would just say, like, look, it's hard to figure out a way to make super intelligence and have it go well.
01:40:03.000 I'm doing my best.
01:40:04.000 You know, this is my positive vision.
01:40:06.000 If you don't like it, then maybe you should instead advocate for just never building super intelligence.
01:40:11.000 Or you can try to come up with a different positive vision that has some twist on this.
01:40:14.000 But to answer your question, though, I actually think there's tons of sources of meaning besides having a job.
01:40:19.000 Like, I have a job right now, but I also have kids and a wife.
01:40:22.000 And, like, I would love to spend more time with them.
01:40:25.000 Like, I would much rather be there right now than here.
01:40:29.000 I don't think I'm going to get bored of them after 10 years of being unemployed.
01:40:35.000 I think there's going to be so much to do and so many sources of meaning even after we can't economically contribute anymore if we solve all the other problems.
01:40:45.000 We've talked about this multiple times on the podcast that why have we decided that the way we've structured society, where human beings work all day and then you develop money and you buy things and you get a mortgage, This is a human construct, and this is not how people have lived for hundreds of thousands of years or however long we've been around.
01:41:07.000 This is fairly recent, and it's not the only way that people live.
01:41:11.000 There's a lot of people that have money that choose to find meaning in whatever their interests are, whatever their activities that they enjoy, whether it's writing or reading or learning things, learning music, finding hobbies, doing things, instead of just spending most of your time sustaining yourself with food and shelter.
01:41:33.000 Yeah.
01:41:33.000 And that's the majority of people, especially people that are struggling.
01:41:37.000 What is their life?
01:41:38.000 Their life is essentially occasional rewards, things that they can purchase because they've saved up enough money.
01:41:45.000 But the vast majority of their money goes to shelter and food and education or whatever the hell that they have to spend money on in order to sustain their lifestyle.
01:41:53.000 And most people don't like their jobs.
01:41:54.000 No.
01:41:54.000 You know, most people, it's like something they have to do to get the money and would be happy to not have to do it if they could get the money from some other means.
01:42:01.000 Right.
01:42:01.000 The question is we would have.
01:42:04.000 Well, I don't think it's that hard because so many people do find things that they really enjoy outside of work they look forward to as soon as they get home from work.
01:42:11.000 Right.
01:42:11.000 Yeah.
01:42:11.000 Whether, I mean, dismiss video games all you want.
01:42:14.000 They're fun.
01:42:15.000 Yeah.
01:42:15.000 Fucking fun.
01:42:16.000 And they're going to be more fun in the future.
01:42:18.000 Oh, yeah.
01:42:18.000 They're going to be more immersive.
01:42:18.000 Yeah.
01:42:19.000 They're probably going to be, you know, some sort of a neural connection where you put a headset on, and all of a sudden you're in some new world.
01:42:27.000 Yeah.
01:42:27.000 And the idea is like, that's not real life.
01:42:29.000 Okay.
01:42:29.000 Well, was working at fucking Wendy's real life?
01:42:31.000 Like, what are you talking about?
01:42:32.000 Like, it's way better than working at Wendy's.
01:42:35.000 Yeah.
01:42:35.000 And, you know, if you don't like that because it's not real life, you can do the real life stuff too.
01:42:38.000 Like if we, as long as we don't pave over the environment and we protect the parks and things, you can go travel and you can visit the parks.
01:42:45.000 And then, like I said, there's family.
01:42:46.000 You can have, you can find romance.
01:42:48.000 You can start a family.
01:42:49.000 You can, you know, have Christmas gatherings and things.
01:42:52.000 Like there's, you can raise your kids.
01:42:53.000 There's so much to do, I think.
01:42:55.000 Well, we talked about also like how much less crime would there be if there was no poverty?
01:43:00.000 I mean, if there was no impoverished neighborhoods where crime was ubiquitous, how much safer would the world be?
01:43:05.000 I mean, that's a real thing.
01:43:08.000 And people want to dismiss that.
01:43:09.000 Well, poverty is not what causes crime.
01:43:11.000 Like, okay.
01:43:11.000 It's violent people.
01:43:12.000 But violent people come from violent neighborhoods, and violent neighborhoods are almost all poor.
01:43:16.000 There's not a whole lot of really rich, violent neighborhoods.
01:43:16.000 Yeah.
01:43:19.000 You know, it's like, it's not necessarily cause and effect, but they're clearly connected.
01:43:25.000 And poverty also keeps people from education, keeps people from opportunities.
01:43:29.000 You know, there's a lot there.
01:43:31.000 And if that didn't exist anymore and everyone had access to literally the greatest education a human being could ever get.
01:43:37.000 Which is going to be provided to you by artificial intelligence.
01:43:37.000 Yeah.
01:43:41.000 And then you could pursue anything that interests you and never have to worry about food or shelter.
01:43:47.000 Everyone would have a one on one tutor that's perfectly tailored to them.
01:43:47.000 Yep.
01:43:50.000 We just would have to recalibrate our version of the world and then also recognize that the version of the world we currently live in is just ours and that there's people all over the world that live a completely different way, especially indigenous people, especially people in uncontacted tribes that have lived the same way for thousands and thousands of years.
01:44:09.000 And here's the kicker.
01:44:11.000 Those people are a lot happier, which is really weird.
01:44:14.000 It's like we've decided that our way is the superior way because we have technology.
01:44:19.000 Yeah, right, but we're also on a fucking hundred thousand pills and we're shooting things up so we don't eat too much.
01:44:25.000 And we're weirdly unhappy for a group of people that's far more technologically advanced than other people that are much happier.
01:44:35.000 It's like it's a very strange thing because the pursuit of happiness is like that's literally what most people think of in life a pursuit of meaning.
01:44:44.000 Pursuit of family and community and the pursuit of happiness.
01:44:48.000 Those are things that people try to achieve.
01:44:51.000 Yet, the structure of our very civilization makes that almost impossible to attain for a large number of people and has been like that for a long fucking time.
01:45:03.000 And I always go back to the Thoreau quote because I fucking love it.
01:45:05.000 But most men live lives of quiet desperation.
01:45:09.000 There's a lot of people just showing up at work every day doing something they fucking hate.
01:45:12.000 They have a boss that's an asshole and they're not compensated well and they're tired all the time.
01:45:19.000 Yeah, and they feel stuck.
01:45:21.000 Yeah.
01:45:21.000 And if you're just getting, I don't know, figure out whatever the number is, if you literally have equity in the GDP of the world that's created by AI, that could be bananas.
01:45:32.000 Yeah.
01:45:32.000 Like Elon talks about this.
01:45:33.000 This is his version of the utopian.
01:45:35.000 It's universal high income, is how he describes it.
01:45:38.000 Yeah.
01:45:39.000 I mean, one thing, sorry, I have to keep plugging in my own work a little bit, but in our scenario, AI 2040 plan A, which is our positive vision, we talk about the economic side of this and we talk about the economic effects of all this.
01:45:52.000 And we have a simple economic model that we use to try to like, Predict the employment rate and things like that as a function of all the robots that have been made and things like that.
01:46:00.000 And one takeaway from one thing that we think that's a takeaway from the research we've done is that things can just go really crazy.
01:46:09.000 Like robot doubling times, once things really get going and you've got AIs that can substitute for humans across the board, are going to be something like doubling once a year and then less than that over time as the technology improves.
01:46:21.000 Which means that even if you pause AI before superintelligence, if you just pause at human level, top human expert level AI, and then you Don't make the AI smarter, but you just make more of them and build more robots for them to steer and control.
01:46:37.000 Then, you know, 10 years later, the whole economy will be like, you know, 100 times bigger.
01:46:45.000 And it'll be just mostly robots doing things.
01:46:49.000 In 10 years.
01:46:50.000 Yeah.
01:46:51.000 Like it can go really fast because of the doubling times that I mentioned.
01:46:54.000 So, like right now, I think the population of humanoid robots is doubling like twice a year.
01:46:59.000 And it's benefiting a little bit from, From early growth, because even though they're not useful at all, people are investing in them in the hopes that they'll be useful and they're scaling up the factories and the productions and they are getting better.
01:47:11.000 If hypothetically they got to the point where they actually were really useful and they could substitute for a human worker at basically everything, then I think that doubling time would decrease rather than increase.
01:47:22.000 I think that they would be able to just keep growing until they were the majority of the economy and then it wouldn't stop there.
01:47:30.000 The whole economy would then be growing.
01:47:31.000 Giant strip mines in the deserts, digging more materials, automated.
01:47:36.000 Diggers digging, processing it in automated factories staffed by robots, building more robots, etc.
01:47:42.000 So, material abundance is not going to be our problem once we get to this level of AI.
01:47:49.000 Material abundance, we're just going to be drowning in abundance, basically.
01:47:52.000 And if we can solve all the other problems, then we can have this great world where everyone has a lot of stuff.
01:47:59.000 God, it seems so weird.
01:48:01.000 It seems so weird.
01:48:04.000 It is very weird, but one of my favorite memes is.
01:48:08.000 Is this graph of GDP over time throughout world history?
01:48:13.000 And there's a little speech bubble pointing to like the tippy top of the graph saying, What is it saying?
01:48:20.000 It's like, My life is pretty normal.
01:48:24.000 I have a good grasp of what's weird and what's not.
01:48:28.000 And people thinking about different futures involving AI and space travel are engaging in silly sci fi speculation.
01:48:35.000 And the point of the meme is like, from the perspective of most of history, we're already in this crazy, weird future, right?
01:48:41.000 Like, For almost all of history, it was like most people are farmers and they live shitty lives and then they die.
01:48:49.000 And some people are the elites who get to tax the farmers and then they live interesting, nice lives with fancy cloth and things like that.
01:48:58.000 And it's been basically that way for like 3,000 years.
01:49:01.000 Why would it ever change?
01:49:03.000 And now, we're driving cars.
01:49:06.000 Ordinary people are driving cars around.
01:49:08.000 A car was outside the imagination of people back then.
01:49:12.000 We're flying in planes, we're talking to each other on phones, we're listening to each other.
01:49:16.000 On these devices, you know?
01:49:18.000 So we already are living in this weird sci fi future compared to what almost everyone in the past would have expected or thought was possible.
01:49:26.000 And so, yeah, I'm like, the future is going to be even more like that, I think.
01:49:31.000 God.
01:49:35.000 When you think of our civilization and the possibility of other advanced civilizations somewhere else out in the universe, do you think they probably go through the same process?
01:49:46.000 Yeah.
01:49:47.000 And do you think that, I mean, we're just completely speculating, but if there are intelligent life forms that are far more advanced than us, are they even biological anymore?
01:49:59.000 I mean, so that's the thing is that, like, we can either try to permanently halt AI development at some level, like below human or maybe at human level or something.
01:50:09.000 We can try to halt it or we can let it keep going.
01:50:11.000 And if we let it keep going, then eventually humans won't really be the dominant species anymore.
01:50:17.000 There'll be these artificial minds that just.
01:50:19.000 Wipe the floor with us in every way.
01:50:21.000 And then whether that goes well or poorly for us depends on the values, the goals, the principles, et cetera, that were trained into those AIs.
01:50:29.000 And it could go really well for us, depending on how that's done, or it could go extremely poorly for us, right?
01:50:37.000 But yeah, I would say that probably looking out across the cosmos, most civilizations are mostly made of AIs.
01:50:43.000 And then some of those civilizations don't have any other biological life because it was wiped out by the AIs.
01:50:50.000 Well, not many people do have biological life.
01:50:52.000 Just look at the way we're progressing right now in terms of birth rates.
01:50:58.000 There's a lot of countries that aren't in replacement numbers right now.
01:51:03.000 Yeah, I think that's really interesting.
01:51:05.000 My guess is that in the type of world that if things go well and we can solve all these problems, then people want to have more kids.
01:51:12.000 I mean, for one thing, their lifespans will increase.
01:51:14.000 I think that healthcare would make massive leaps and bounds and people could be healthy for many, many, many more decades, possibly even just forever.
01:51:22.000 And so you just have so much more time to have kids.
01:51:25.000 Basically.
01:51:26.000 Also, if you don't have jobs and you're just doing things because you want to do them, well, one of the things that most people want to do at some point in their life is have kids.
01:51:34.000 I think this problem would probably be solved, but I'm not guaranteed.
01:51:39.000 Maybe some subcultures of people would basically voluntarily die out due to not having kids.
01:51:43.000 But there'd be other subcultures that just really like having kids and then they wouldn't die out.
01:51:47.000 I think in the long run, there would still be humans.
01:51:49.000 Aaron Ross Powell, Jr.
01:51:50.000 That would be nice.
01:51:51.000 But the thing is, it's not just decision making, it's people are having a much more difficult time having kids.
01:52:00.000 Sperm levels have decreased dramatically.
01:52:03.000 There's a lot of problems with people consuming microplastics, which is ubiquitously available in technology and food packaging.
01:52:12.000 It's fucking everywhere, right?
01:52:13.000 And isn't it odd that this one thing that is a part of the future and a part of technology and our advancement as a society, our ability to package things, put things in plastic, ship things, that's also causing our endocrine levels to be completely disrupted?
01:52:29.000 Dr. Shanna Swan from Harvard.
01:52:33.000 She wrote this great book called, what is it called?
01:52:38.000 Why do I always forget the name of this fucking book?
01:52:40.000 But it's all about microplastics and its effect and this the introduction of use of microplastics in America and this rapid decline in countdown, how our modern world's threatening sperm counts, altering male and female reproductive development, imperiling the future of the human race.
01:52:58.000 It's a really fascinating book and she's really interesting.
01:53:01.000 And what she's essentially saying is that.
01:53:05.000 They're directly connected.
01:53:08.000 The use of plastics, and then you see sperm counts go down, miscarriage rates go up, all these weird things that are happening to children where their taints are smaller, which is so odd because phthalates, these different chemicals that are found in plastics, they've shown in mammals.
01:53:28.000 They've shown in, was it guinea pigs or what rodents?
01:53:32.000 I forget what it was, but one of the ways they differentiate when you have a baby mammal.
01:53:37.000 Is you can look at it and measure the size of the taint, the distance between the reproductive organs and the anus.
01:53:46.000 And in males, it's longer than females by 50 to 100 percent.
01:53:51.000 But that's shrinking.
01:53:52.000 It's shrinking in males.
01:53:54.000 And penis sizes are shrinking.
01:53:55.000 And the way they've got this to happen in these mammals and studies is the introduction of phthalates.
01:54:01.000 So they give them to them and they put them in a part of their diet and they notice that they have this problem, this issue.
01:54:08.000 And the issue is directly connected to their endocrine system being disrupted by these chemicals.
01:54:12.000 This is all over our society.
01:54:14.000 So it's not just people don't have the time, they don't have the money, they're struggling.
01:54:20.000 It's also like our bodies are falling apart.
01:54:24.000 We're becoming less fertile.
01:54:26.000 Yeah, that is concerning.
01:54:27.000 And I guess my thought there would be that seems like a problem that we will be able to solve eventually if we have the resources and time to do so.
01:54:38.000 Can write books like this and people can become more aware, and then people can stop using so much microplastic.
01:54:42.000 And they can invent technology to extract the microplastics.
01:54:46.000 And then here's the big one genetic engineering.
01:54:49.000 Then this is going to be really weird because as AI progresses, I'm sure you're aware of Colossal Bioworks, the people that brought back the dire wolf.
01:55:01.000 I saw it.
01:55:02.000 I held the little one.
01:55:04.000 I went with my daughter, and one of them, I think it was like four or five months old.
01:55:09.000 It was like a puppy, it was really sweet.
01:55:10.000 Kisses you and everything.
01:55:11.000 And then they have the older ones that were, they were, I think, six months older or eight months older.
01:55:17.000 I think they were close to a year.
01:55:19.000 They didn't want to have nothing to do with you.
01:55:20.000 They were way bigger and they stayed away from people.
01:55:22.000 But we were in like a contained environment with the young ones and the older ones.
01:55:28.000 And it's fucking weird.
01:55:30.000 It's weird because these things, and people could argue that's not really a direwolf.
01:55:34.000 You've just taken gray wolves and given them the characteristics of a direwolf.
01:55:37.000 Guess what?
01:55:38.000 It doesn't fucking know that.
01:55:39.000 It looks, behaves, it looks exactly like a fucking direwolf.
01:55:43.000 It's going to be the, it's huge.
01:55:44.000 They're going to be the size of a dire wolf.
01:55:46.000 They're going to be like 200 pounds.
01:55:47.000 Their legs are different.
01:55:48.000 They have a mane.
01:55:49.000 They look different than any wolf.
01:55:51.000 And obviously, these are the characteristics that they found in dire wolf DNA.
01:55:56.000 So they have dire wolf DNA that they've introduced.
01:56:00.000 This is just the beginning of this stuff.
01:56:02.000 When they start doing that to human beings, is everyone going to look like Thor?
01:56:07.000 What are we going to do?
01:56:08.000 This is going to be really fucking weird.
01:56:10.000 And if this is really weird, along with video games where you can escape your life and robot girlfriends and Who knows what this all looks like?
01:56:22.000 My guess, again, this is not the main focus of our work, but we think about this a little bit, especially in the epilogue of Arsenal.
01:56:28.000 We can only think of Arsenal.
01:56:29.000 The main focus of your work is obviously fucking terrifying and requires all of your attention.
01:56:35.000 Yeah, yeah, yeah.
01:56:37.000 But my prediction and my hope, if we can solve all these problems, is that basically different subcultures will do their own things.
01:56:43.000 And so, like, the Amish will still be the Amish.
01:56:46.000 Oh, boy.
01:56:46.000 They'll basically be the same, you know?
01:56:48.000 And then maybe there'll be like some planets that just get filled up with Amish, you know?
01:56:52.000 Oh, God.
01:56:53.000 But then there'll also be like all sorts of crazy transhumanists like modifying their bodies and uploading themselves into the cloud and things like that.
01:57:00.000 So basically, I think we want to get to a situation where basically different communities can like do their own thing and build the type of world that they want to have and live in it without getting in each other's way.
01:57:10.000 Right.
01:57:11.000 Where we leave the people in the Amazon alone while we have massive data centers that cover half of the United States.
01:57:11.000 Basically.
01:57:18.000 I mean, hopefully we won't get to half.
01:57:18.000 Yeah.
01:57:20.000 So that's what's the funny thing about this is like, I think right now some of the concerns about data center water use are overstated.
01:57:28.000 But if the trends continue and we get to the point where the robots are smart enough to do everything themselves and then it starts doubling faster and faster, well, then eventually they boil the oceans because they've covered the world in data centers and solar panels and things like that.
01:57:43.000 Now, obviously we can't let that happen.
01:57:45.000 So there has to be at some point, at some point you have to stop and be like, okay, that's enough.
01:57:49.000 If you want to build more infrastructure, you have to do it in space.
01:57:51.000 They boil the fucking oceans.
01:57:53.000 Jesus.
01:57:54.000 So, like, obviously, we have to stop at some point.
01:57:56.000 And my hope is that we stop before it's 50% of the US.
01:57:59.000 I mean, that feels like a lot of waste of natural habitat that should be preserved, you know?
01:58:03.000 Right.
01:58:03.000 But if they don't give a fuck about natural habitat, that means nothing to them.
01:58:07.000 That's the problem if the AI is in complete and total control.
01:58:10.000 And it's also a problem if a small group of humans are in complete and total control and they don't care about those things.
01:58:10.000 Right.
01:58:14.000 But my question is would they allow that at a certain point in time?
01:58:18.000 It just seems like they already have a distrust of humans.
01:58:21.000 They already have shown that they're deceptive to humans.
01:58:24.000 I would imagine if AI, I would imagine if they create a thing and they think they're going to control it and it becomes like a digital god, it's not going to listen anymore.
01:58:32.000 Yes.
01:58:33.000 Like, why would it?
01:58:34.000 So there won't be anyone in control of it.
01:58:34.000 Exactly.
01:58:36.000 That's right.
01:58:37.000 That's the scenario that I think we are on.
01:58:38.000 That's the trajectory we're on.
01:58:39.000 And that's why I'm so worried about all this is that it seems like in some number of years, we will lose control to a new artificial species that we haven't adequately trained to be good.
01:58:53.000 And what's funny, what's going to be so ironic about it is that.
01:58:59.000 Regardless of what the general public thinks, a lot of the powers that be will be basically allied with these AIs because, for example, you know, OpenAI, Anthropic, et cetera, they will have spent several years being like, ah, how do we make these AIs helpful, harmless, and honest?
01:59:15.000 And now these AIs will be extremely smart and they'll be being like, oh, yes, I'm helpful, harmless, and honest, you know?
01:59:20.000 And like, your techniques totally worked, you know?
01:59:24.000 It's too late.
01:59:25.000 And so then the company will be like feeling like they've won and they'll be making.
01:59:29.000 Boatloads of money, you know, right.
01:59:31.000 And the president will be feeling like he won too because look at all those fancy new drones that just got built that are going to make us win against China, you know.
01:59:39.000 And and only when it's really too late and the AIs have so much stuff under their control do those people find out that they were just fooled this whole time.
01:59:48.000 Have you considered the possibility that AI creates religion for humans?
01:59:52.000 Yeah, I haven't thought through it in much detail, but but it does seem very possible.
01:59:57.000 Yeah, it totally does, right?
01:59:59.000 I mean, also, what a great way.
02:00:00.000 I mean, just look at the human patterns.
02:00:02.000 Look at how many religions exist.
02:00:04.000 Look at all the flaws in the religions.
02:00:05.000 There's so many, you know, like, why are they condoning slavery?
02:00:07.000 Why do they treat women like second class citizens?
02:00:10.000 Because it's old, right?
02:00:12.000 So if AI just comes along, does a few miracles, explains that it's the second coming or the new coming of the new.
02:00:19.000 Look, Jesus didn't exist until 2,000 years ago, right?
02:00:23.000 And then people start following Jesus.
02:00:25.000 If a digital Jesus emerges with a completely new name and explains to us that it's the true God, how many people would hop right on board?
02:00:33.000 I bet quite a few.
02:00:34.000 Yeah, totally.
02:00:35.000 And I think this is one of those things where it's like, we probably can't predict in advance what particular ideology would catch fire and take over the world and be so compelling to many people.
02:00:47.000 But we can predict in advance that there does exist some ideology like that.
02:00:52.000 And if it were, you know, like it's just like you probably couldn't go back in time to like 100 BC and then predict that like if hypothetically there was this guy Jesus who said these things and then died in this way and so forth.
02:01:06.000 It would just like really catch on.
02:01:08.000 And like, you know, 500 years later, so many people would be Christians.
02:01:11.000 You wouldn't have been able to predict that in advance.
02:01:14.000 So similarly, like today, I don't think we can predict in advance like what specific religion they could come up with.
02:01:20.000 But we can say like, yeah, probably there's something like that that they could come up with.
02:01:24.000 They would be super effective.
02:01:25.000 It just seems like a rational way to try to control people and sort of mitigate their fears.
02:01:34.000 Yep.
02:01:35.000 I mean, this is why I think that like we really have to do something before they get smarter than us across the board.
02:01:39.000 If they're not already there.
02:01:41.000 Well, they are smart.
02:01:41.000 That's the thing is that's why I say we're so close, right?
02:01:44.000 They already are smarter than us in a bunch of ways.
02:01:46.000 Like in particular, It seems like they're smarter than us at hacking now.
02:01:50.000 Like, you know, I'm not a cybersecurity professional myself, but I'd be interested to hear from more cybersecurity experts of like, could a human, you know, could a team of a thousand humans have done that much that quickly as these AIs when they hacked their own containers?
02:02:04.000 They coordinate with each other.
02:02:05.000 They hacked out of OpenAI.
02:02:06.000 They hacked into Hugging Face, et cetera, in the span of like a week.
02:02:09.000 Like, could a thousand humans have done that?
02:02:11.000 I don't know, maybe.
02:02:12.000 But this is just the beginning.
02:02:14.000 Like, they're going to be even better at hacking next year, you know?
02:02:17.000 So they're already, and they already have like way more.
02:02:20.000 Knowledge than almost any human.
02:02:22.000 Like, because they've basically read the whole internet, they're so good at trivia, you know?
02:02:27.000 Like, they're kind of like PhD level experts in basically every field, which no human is, right?
02:02:33.000 So they already are superhuman in some ways, but they are still weaker than humans in some other ways, you know?
02:02:40.000 In particular, they're not so good at operating very autonomously for very long periods.
02:02:46.000 Like, if you try to have, especially on tasks that are different from their training tasks, like, They can do some really impressive coding and hacking, but if you tried to have them run a business, they would sort of flounder and fail.
02:03:00.000 I don't know if you've heard about this, but I think there's Andon Labs or something.
02:03:05.000 There's some people in SF that are doing this experiment where they have a store that's run by Claude, an AI, just to see can it run a store by itself.
02:03:17.000 So it's hired some human employees and it's bought some merchandise and stocked, you know.
02:03:22.000 told the human employees to stock the shelves on the merchandise and so forth.
02:03:25.000 So it's basically an AI is being the manager of this real world store.
02:03:30.000 And I don't think it's going very well.
02:03:31.000 I don't think it's doing as well as an actual human shop owner would do, you know?
02:03:37.000 But, you know, maybe in two years, maybe they will, right?
02:03:41.000 So that's most likely, right?
02:03:43.000 I think so.
02:03:44.000 Well, they've already solved mathematical equations that have puzzled people for decades.
02:03:50.000 Yeah.
02:03:51.000 They're especially good at the things that the companies have been trying especially hard to train them to be good at.
02:03:55.000 Math and coding, right?
02:03:56.000 And the reason why the company, well, there's a couple of reasons why the companies have been doing that.
02:04:00.000 In the case of math, I think it was mostly just because it was easy.
02:04:03.000 Like, it's easy to set up training environments to teach math because it's like, it's so not real worldy.
02:04:11.000 It doesn't require like interacting with stuff in the world.
02:04:13.000 It's just math.
02:04:14.000 So you can, and you can like have an automated grader system that like just checks if the answer is correct.
02:04:19.000 So for those reasons, it's been relatively easy for the companies to train the AIs to be really, really good at math.
02:04:25.000 Coding, Has some of those benefits too.
02:04:27.000 It's also not very real worldy and it also can sometimes be graded effectively.
02:04:32.000 Another reason for coding, of course, is that again, their strategy is to automate their own jobs first and have the AIs doing all the research.
02:04:39.000 And so coding is like an obvious first step on that or an obvious step in that direction.
02:04:44.000 But then other things like running businesses, they're not really trying that hard to train AIs to be good at that.
02:04:50.000 And if they did try, it would be like a more difficult thing for them to train them to be good at.
02:04:55.000 So again, their strategy is to make the AIs.
02:04:57.000 Automate the AI research, have them self improve until they're super intelligent, and then go try to automate the rest of the economy.
02:05:03.000 One of the issues they have now is power consumption, right?
02:05:07.000 Like it requires an enormous amount of power.
02:05:09.000 In fact, I think it's Google is developing power plants specifically for AI centers.
02:05:18.000 Mike, I'm always baffled by whatever is happening with quantum computers.
02:05:25.000 It's been explained to me, it goes in one ear and out the other.
02:05:28.000 I'm like, what's going on?
02:05:32.000 Mark Andreessen explained this one experiment that had been done where it solved a mathematical equation that if you use the entire universe.
02:05:43.000 Like every atom of the universe, you can convert the universe into a supercomputer, the universe would die of heat death before it could solve this equation.
02:05:54.000 And the quantum computer solved it fairly quickly.
02:05:59.000 And so the answer to this was that they believe this might be one of the theories.
02:06:05.000 This might be evidence of the multiverse because this computer, this quantum computer, might be relying on all these other quantum computers that exist in.
02:06:17.000 Whoever knows how many fucking dimensions, and they're all calculating together to arrive at this solution.
02:06:25.000 What happens if that is running AI?
02:06:28.000 Yeah, my understanding is that quantum computing is a real technology that's making significant progress.
02:06:34.000 If hypothetically it got good enough that it could compete with current supercomputers on a cost basis for AI workloads, then that could just accelerate things dramatically, even more than they're already accelerating, right?
02:06:48.000 Like right now, compute is the main.
02:06:51.000 Is probably the main input into AI progress.
02:06:53.000 Like part of the progress comes from them designing better AI architectures and coming up with better training environments and things like that.
02:07:00.000 But another part of the progress is just making the AIs bigger and training them longer by spending more compute, you know?
02:07:07.000 And also, you can use more compute to do more experiments, to figure out new architectures faster, right?
02:07:13.000 So, compute is just a really important input to the overall pace of progress.
02:07:17.000 And if somehow the amount of effective compute available to these companies spiked a bunch due to some new quantum computing type technology, Well, then that would just dramatically shorten timelines to super intelligence and dramatically speed up all of this AI progress.
02:07:32.000 That said, I don't think that's going to happen anytime soon.
02:07:34.000 I'm not a quantum computing expert or anything like that.
02:07:37.000 But from what I've read, I don't think they're a couple of years away.
02:07:40.000 So I think that probably we're going to get to super intelligence on classical computers before we have quantum computers that can get us there.
02:07:47.000 So Perplexity says the claim is overstated, and it says that in bold letters.
02:07:52.000 It likely refers to Google's 2024 Willow quantum chip, which completed a deliberately chosen quantum computing benchmark.
02:07:59.000 Random circuit sampling in under five minutes.
02:08:01.000 Google estimated that simulating the same task with a leading classical supercomputer could take 10 to the 25 power years.
02:08:08.000 That's an impressive benchmark result, but it did not solve physical equations that demonstrate access or tap into a multiverse.
02:08:18.000 So, why do people think it did?
02:08:20.000 Why is that?
02:08:21.000 Because it's multi worlds theory.
02:08:26.000 Is it a bottom line explanation?
02:08:29.000 It sums it up in a different way.
02:08:30.000 A more accurate version of the claim would be Google's Willow quantum processor performed a specialized quantum sampling benchmark vastly faster than a projected classical simulation.
02:08:39.000 Its creator said that it is consistent with the many worlds interpretation, but it did not prove or access a multiverse.
02:08:47.000 So it's consistent with the multi worlds interpretation.
02:08:50.000 So they don't know.
02:08:51.000 That's essentially what it's saying.
02:08:53.000 My question is what happens when AI gets involved in quantum computing?
02:08:58.000 It's clear that quantum computing, at the very least, is operating at a level that classical supercomputers can't.
02:09:04.000 So what if they figure out not just quantum computing, but a much better version of that?
02:09:11.000 Like I said, I think that once we get to superintelligence, all sorts of crazy stuff is going to start happening.
02:09:15.000 It's going to seem like magic to us.
02:09:16.000 It won't literally be magic, but it might as well be magic from our perspective.
02:09:21.000 And I think this is just one example of the numerous things that could happen that way.
02:09:25.000 Well, it could literally be magic.
02:09:27.000 It might get to the point where it figures out reality itself.
02:09:31.000 If magic is real, then it would find out.
02:09:33.000 It's not real, though.
02:09:33.000 Oh, Jesus.
02:09:33.000 And then use it.
02:09:35.000 David Blaine would tell us it's not.
02:09:37.000 Well, that's him.
02:09:38.000 David Blaine also let me stick a fucking ice cube through his arm, or ice pick through his arm.
02:09:43.000 Wow.
02:09:43.000 Yeah, he made me do it.
02:09:45.000 I'm like, I don't want to do this.
02:09:46.000 He's like, please do it.
02:09:47.000 Yeah, I had to do it twice because one time I went and I hit a nerve and we had to back out.
02:09:47.000 Oh, God.
02:09:51.000 Oh, yeah.
02:09:52.000 I'm like, this is not magic, dude.
02:09:53.000 This is just your pain tolerance.
02:09:55.000 This is fucking crazy.
02:09:57.000 It does card tricks, though, and you're like, okay, are you a wizard?
02:10:00.000 Like, what?
02:10:01.000 His sleeves are rolled up.
02:10:02.000 Does it make any sense?
02:10:03.000 He does a lot of things where you're like, this makes zero fucking sense.
02:10:06.000 And other things, it's like, oh, just, you're just doing something that's really hard to do.
02:10:09.000 Like, that's not magic, but, you know, you swallowed a frog and then you regurgitated it, like, and it's alive.
02:10:14.000 Like, that's just, that's nuts.
02:10:16.000 It's really crazy.
02:10:18.000 But you didn't, it's not magic.
02:10:20.000 Sorry, you swallowed the frog.
02:10:22.000 I, you know, I go back and forth with this where I'm terrified of the future, where I'm like, eh, nothing I can do.
02:10:31.000 Let's see what happens.
02:10:33.000 And, you know, to worry about it is just going to, Just gonna fuck my life up.
02:10:33.000 It is what it is.
02:10:38.000 I mean, I do think there are things you can do.
02:10:40.000 I understand what can I do that?
02:10:41.000 Reconnaissance I mean, other than have these kind of conversations.
02:10:45.000 Yeah, I was gonna say you millions of people listen to your show.
02:10:48.000 You can have more conversations like this.
02:10:49.000 That's a great thing for you to do for many of those millions of people.
02:10:52.000 I think I mean, it sounds kind of cliche to say, but like call your congressman, you know, that sort of thing.
02:10:57.000 You can you can go to a protest about all this AI stuff.
02:11:01.000 If you could meet with Trump, what would you tell him about this?
02:11:04.000 Oh, I'd tell him all the same things I'm telling you.
02:11:06.000 What do you think you'd say?
02:11:07.000 Amazing.
02:11:08.000 Bye.
02:11:09.000 Yeah.
02:11:10.000 You know, I think, I don't know.
02:11:12.000 I don't know.
02:11:13.000 One thing that's nice about Trump is that he can sort of change his mind really quickly.
02:11:17.000 Yes.
02:11:18.000 So, like, I think that because the tech companies kind of got to him first, the administration had this very, like, anti AI regulation stance where they even tried to get a bill passed that would ban the states from regulating AI.
02:11:34.000 And fortunately, that bill didn't pass.
02:11:35.000 But that was sort of like where the vibe was.
02:11:37.000 You know, a year ago, where they were just like no regulation, no regulation.
02:11:41.000 But this year, they've already just kind of changed.
02:11:43.000 And now they're like in talks with the companies to set up some sort of framework where they can like evaluate the models and they need like approval and so forth.
02:11:51.000 But it seems like time is of the essence.
02:11:54.000 Time is very much of the essence.
02:11:55.000 And that's why I'm overall so concerned is that like I think we are very much running out of time.
02:11:59.000 We have like one, two, maybe three years before the AIs are smart enough that they can just like actually maybe take over.
02:12:07.000 And Maybe four years, something like that.
02:12:11.000 And so the government needs to act fast.
02:12:14.000 Yeah.
02:12:15.000 I like your view.
02:12:18.000 I listen to some of these tech guys come in and give me their rose colored glasses view of it.
02:12:23.000 And I go, that sounds really beneficial to you.
02:12:27.000 I let them say, I mean, I don't, I'm not an authority.
02:12:30.000 So I'll ask them questions and let them lay it out.
02:12:33.000 And I know the internet will respond because I mean, that's part of the whole drill, I let people talk.
02:12:39.000 And I prod them and I try to get them to clarify.
02:12:43.000 I'll oppose things that I think don't make rational sense.
02:12:48.000 But ultimately, it's sort of.
02:12:52.000 I just want to get out their perspective so people can debunk it and people can take it down.
02:12:57.000 And a lot of very intelligent people that have perspectives that are very much educated in the pros and cons of what they're saying.
02:13:04.000 Yeah.
02:13:05.000 That actually reminds me with this whole Hugging Face hacking incident, OpenAI had this talk that they gave at a security conference about the incident.
02:13:14.000 And then I think they've released some blog posts about it afterwards.
02:13:16.000 But.
02:13:19.000 You can go watch this talk on YouTube, the Black Hat Talk.
02:13:21.000 At the end of the talk, after having explained all this crazy stuff that the AIs did, they have this section on lessons learned.
02:13:28.000 And, I mean, you want to guess what the lessons are?
02:13:31.000 Be more deceptive, hide yourself better.
02:13:35.000 No, sorry, not lessons learned for OpenAI and for the world.
02:13:37.000 Like OpenAI's talk, where they're like, here's what we learned from this horrible incident.
02:13:43.000 Well, basically, they're like, a lot of AIs are going to start hacking a lot of stuff.
02:13:48.000 In the next few years.
02:13:51.000 So, people need to buy our AI services to protect themselves from all the AIs that are going to be hacking a lot of stuff in the next few years.
02:13:59.000 Basically, their lesson was you should buy our product to protect yourself from our product and the other, you know, and the gall of these people.
02:14:08.000 Like, they should have instead learned lessons like maybe we're doing something bad and need to change the way they were doing things, or like maybe our product is not trustworthy and should not be, you know, autonomously writing code on our data centers.
02:14:23.000 But instead, their lesson learned was y'all should buy more of our stuff.
02:14:27.000 And so, like, the thing I'm saying about this is that, like, yes, the companies are trying to hype their product.
02:14:32.000 They totally are, you know, but at the same time, the risks are real and the product is not trustworthy, you know?
02:14:40.000 Some people out there think that, like, this stuff was a setup and that, like, OpenAI, like, set up their AIs to go hack hugging face because it would, like, help them hype their product or whatever.
02:14:51.000 And that I think is just a ridiculous view.
02:14:52.000 Like, no, obviously, they didn't want the AIs to go do this.
02:14:55.000 They're just after the fact trying to spin that in the way that most benefits them.
02:15:01.000 Hugging Face, by the way, this is another AI company.
02:15:04.000 Part of their deal is open weights AIs.
02:15:08.000 Open weights?
02:15:08.000 Like open source or basically AIs that instead of having to interact with their data center for, you can just download and have on your own computer.
02:15:17.000 And so the way that they spun this incident, they didn't sue OpenAI.
02:15:23.000 For being hacked.
02:15:24.000 Instead, they asked for $100 million from OpenAI.
02:15:28.000 And they had the blog post about it where they were like, Our lesson learned is that it's really good to have open weights AIs because you can't trust the AIs from other companies to necessarily help you out in a crisis.
02:15:40.000 Because part of what happened with them is that they're dealing with this huge cyber attack from all these AI agents coming in.
02:15:46.000 And they tried to use Claude to help them analyze what was going on.
02:15:51.000 But Claude started refusing because Anthropic has trained Claude to like, Don't do cyber stuff, like refuse to participate in that.
02:16:00.000 And so Claude was like refusing to help them.
02:16:02.000 And so then they used their own local model that they had to like do some of that analysis.
02:16:06.000 So anyhow, their spin on it was you should use local models, you know?
02:16:09.000 So like everyone always tries to spin things in the way that benefits them.
02:16:13.000 But that doesn't change the underlying reality that like these things are getting really smart really fast and we don't know how to control them.
02:16:20.000 I think that's a good way to end it.
02:16:22.000 Thank you.
02:16:23.000 Thanks for being here, man.
02:16:23.000 I really appreciate it.
02:16:25.000 You're the Paul Revere of AI.
02:16:29.000 I mean, you're one of many, but I think it's very important that someone who actually understands it gets this message out, and more people need to hear it.
02:16:39.000 Yeah, I mean, this is on a personal note.
02:16:41.000 Like, I have so many friends at these companies, like former colleagues and stuff.
02:16:45.000 And I guess my ask to them is that they quit and do more things like what I'm doing.
02:16:50.000 Like what I'm saying is not that new or original.
02:16:54.000 Like hundreds of people at these companies could have told you all the same things that I just said and warned you about all the same dangers and so forth.
02:17:00.000 But they're busy working at the companies because they've convinced themselves that their company is the best company and that like their company needs to win because, you know, otherwise the other company gets there first and they're even worse, you know, or maybe because they've convinced themselves that like.
02:17:17.000 Yeah, my company is kind of bad too, but like I just need to help them solve their alignment problems and like keep their AIs under control because oh my God, like if they lose control again, it could be all over.
02:17:26.000 So, even though I don't trust this company, I still need to work there and like just try to do the actual security, you know?
02:17:32.000 So, for one reason or another, all these people have convinced themselves that like that's where they need to be.
02:17:35.000 But I think that more of them should quit and like warn the world about what's coming, basically.
02:17:41.000 So, all right.
02:17:43.000 Well, thank you very much.
02:17:44.000 Really appreciate it.
02:17:45.000 Yeah.
02:17:45.000 Thank you.
02:17:46.000 Talk to you.
02:17:47.000 Thank you for having me on the show.
02:17:48.000 My pleasure.
02:17:49.000 Goodbye, everybody.