Making Sense - Sam Harris - September 22, 2026


#494 — A Coin Toss for the Future

#493 — Two Reasons to Meditate

Episode Stats


Length

27 minutes

Words per minute

184.06

Word count

4,994

Sentence count

188


Transcript

Transcript generated with Whisper (turbo).
00:00:00.000 you're listening to making sense with sam harris this is the free version of the podcast so you'll
00:00:06.400 only hear the first part of today's conversation if you want the full episode and every episode
00:00:11.360 you can subscribe at sam harris.org there are no ads on this show it runs entirely on subscriber
00:00:18.000 support if you enjoy what we're doing here and find it valuable please consider subscribing today
00:00:22.960 i am here with ryan greenblatt ryan thanks for joining me it's good to be here so you're uh
00:00:30.000 the chief scientist at Redwood Research, which is one of these AI safety research firms that
00:00:37.400 analyzed the recent Hugging Face fiasco incident, terrifying phenomenon. We'll get into it. But
00:00:43.840 how did you come to this work? What was your path to focusing on AI safety?
00:00:49.160 Yeah. So in college, during my junior year, I was sort of in my apartment alone because it was,
00:00:54.800 you know, COVID, all classes were remote. And I was sort of in a contemplative mood,
00:01:00.480 thinking about what I should do with my life, listening to various podcasts. And in one of
00:01:06.100 those podcasts, someone made an argument that was like, basically, I would summarize the argument
00:01:10.100 as like, it doesn't make that much sense to be selfish, because, you know, what's really the
00:01:14.680 distinction at like a material level between your future self and other future people.
00:01:18.640 And I was like, that argument kind of makes sense to me, I should really consider how I want to,
00:01:22.020 lead my life and what I should do. And I ended up getting pretty sold on being much more focused
00:01:26.780 on altruism. And then after that, I spent a bunch of time thinking about what I should do with my
00:01:30.520 life. Eventually, I ended up deciding that working on technical AI safety was an important thing to
00:01:35.980 do, was a very critical issue that we would face. Got sold on those arguments, applied to various
00:01:40.440 places, and then started working at Redwood, where I've been working on AI security and AI safety
00:01:44.920 research for about five years. And is your background in computer science?
00:01:49.340 Yeah, computer science, also some math.
00:01:52.800 So it sounds like you might have come up through the effective altruist community.
00:01:57.720 Is that, I mean, do you consider yourself a part of the EA community?
00:02:00.900 I would say that I'm like definitely like within the EA community.
00:02:04.420 I'm not sure I would self-identify as an EA, but that's, I think like to some extent,
00:02:08.080 just because I'm like, you know, a bit reluctant to self-identify with any label that's going
00:02:12.720 to like associate you in some, like, I feel like I don't, I don't know if I like, like
00:02:17.060 When I think about myself, I don't think I'm an EA. I'm sort of like, yeah, I have a bunch of
00:02:21.540 properties that other people in that community have. I'm in touch with that community. But I
00:02:25.880 wouldn't necessarily say I'm an EA per se. I do think that I'm like, you know, I would definitely
00:02:31.360 say that I'm interested in effectively pursuing the impartial good. Right. Right. Well, as you
00:02:37.500 know, and as I've talked about in the podcast of late, effective altruism has come in for some
00:02:43.460 abuse, much of it unfair. Some of it might be fair. And I think there's a reason why I,
00:02:49.040 like you, have not hung up an EA shingle in my life. Perhaps we can talk about that. But
00:02:54.500 it sounds like, well, you tell me, where on the spectrum of concern are you? I mean, I would put
00:03:00.700 the most worried people over there with Eliezer Yudkowsky and his recent co-author, Nate Soares.
00:03:08.320 Maybe Max Tegmark is over there. Nick Bostrom is not quite over there, but still on that side of the spectrum. And then all the way on the other side, opposing them, you have the people who say that there really is no real concern here, that it's all a hoax, that it's any thought that artificial intelligence might get out of our control and destroy us.
00:03:30.700 that is a weird form of marketing being employed by the Frontier Labs. It's an effort to get
00:03:39.360 regulation imposed and achieve something like regulatory capture, or it's an EA PSYOP. And
00:03:45.580 there are people like, I would say, Mark Andreessen or David Sachs, or even the president
00:03:51.500 himself, who's recently called it a hoax in so many words, who are just fundamentally not worried
00:03:57.800 about any significant downside risk here
00:04:01.820 and just see dollar signs
00:04:04.080 just stretching out to the horizon.
00:04:06.960 Where do you locate yourself on that spectrum?
00:04:09.820 So I would say I'm very concerned,
00:04:12.420 but I may be more optimistic
00:04:13.920 that we'll make it through
00:04:16.040 without sort of averting the course of AI development.
00:04:19.100 So I would say that I'm sort of like,
00:04:21.440 suppose that we proceed
00:04:22.420 on what the current default path looks like in front of us.
00:04:24.540 Maybe there's about a 50 or 60% chance that misaligned AIs would end up taking over the world. And then if that did happen, there would be a significant chance that many or all humans would die. So I think that's, I would say I'm very concerned and don't think the situation is on track to go well.
00:04:42.560 And I would also note that in addition to sort of these misalignment concerns, there's other risks with building extremely capable AI systems, especially around sort of concentration of power and like, you know, who controls these systems? And one answer is no one controls the systems, they're misaligned. Another possible answer is the power is very concentrated in a few people, and it sort of overturns our institutions and our, you know, ability to have, you know, a broad distribution of power, democracy, etc.
00:05:09.140 So what keeps you from being among the absolutely most worried?
00:05:13.600 What do you disagree with them about or what do you think they're getting wrong?
00:05:17.320 Yeah, so I would say that it seems pretty plausible to me that if it turns out that
00:05:22.560 the technical alignment problem is somewhat easier, which I think is plausible, and we
00:05:26.100 do a reasonably good job aligning and utilizing AI systems up through roughly human levels
00:05:32.760 of capability, we could then get those systems to automate technical safety work.
00:05:36.920 And that could potentially sort of get us into a positive feedback loop where these systems make themselves more aligned. They produce a new version of the system that's even better at carefully figuring out what to do, better at trying to pursue our intentions. And that sort of is self-reinforcing rather than getting worse and worse over time. And that could be fast enough to keep up with the rapid growth and capabilities.
00:05:58.800 So that's like one reason. And I would say this comes down to sort of more optimism about prosaic or relatively sort of empirical and iterative methods. I wouldn't say I'm hugely optimistic about those methods, but I think that there's like a decent chance they're working. It's just that when I say a decent chance they're working, I also implicitly mean a decent chance of, you know, losing control over the future. And I'm more like 50-50 on that.
00:06:22.600 And then I also think it's plausible that we'll develop very, very capable AI systems and the worst misalignment concerns will mostly not materialize even with not very advanced methods, even independent of this automation. I think that's a minority, but it's possible. Like I think we don't currently, I don't think there's extremely strong reason to believe that when building significantly superhuman systems, those systems would be so misaligned that they would take over.
00:06:45.260 I think that there's definitely a pretty strong evidence pointing in that direction.
00:06:49.760 And I think that that case gets more concerning the more capable these systems are.
00:06:53.660 But I think that I could imagine basically people not really taking much of a precaution,
00:06:59.100 proceeding through the AI development trajectory, patching problems as they come up, and that
00:07:03.300 ending up getting you very high levels of capability while things are still fine.
00:07:07.120 And then there would be time for society to react.
00:07:10.240 So you've mentioned a few numbers here with respect to probability and just kind of gestured
00:07:15.180 at a range of likelihood. I'm wondering how people should think about statements of that kind. I
00:07:22.120 think these, it seems to me that any actual probability we would assign to this is pretty
00:07:28.120 much made up. Maybe you have a more rigorous way of making an estimate here, but whatever the
00:07:33.760 estimate, I mean, unless it was infinitesimally small, like, you know, well below 1%, which is
00:07:41.260 really never the number that you're hearing. I mean, you hear people, some people will say 10%,
00:07:46.360 20%, 30%. I mean, that seems to be that, you know, you just talked about 50% takeover. I don't
00:07:52.900 know. I don't know how that translates into the ruination of everything, but these are enormous
00:07:58.340 numbers. And so in the normal case, if we were developing a technology where the people who
00:08:04.180 were closest to doing the work said things like, yeah, I think maybe there's a 10% chance we're
00:08:10.020 going to destroy the world here on our present course, the only rational response to that
00:08:15.880 range of outcomes is you stop immediately, right? I mean, if the Manhattan Project scientists said,
00:08:23.840 yeah, we've run our calculations, we've got the smartest people in the room together,
00:08:27.720 we thought about it, and there's a 10% chance that when we execute this first test at Alamogordo,
00:08:33.660 we ignite the atmosphere and destroy the future. The only sane response to that is you don't do
00:08:41.720 this initial test. But that doesn't seem to be what's happening here at all. And we have a lot
00:08:46.840 of people saying that the probability of some extremely bad outcome is quite high. I mean,
00:08:53.360 certainly within range of the role of a normal die or even a coin toss. And yet the work is
00:09:00.280 proceeding more or less under an arms race condition at full pace. We'll talk about the
00:09:06.680 recent statements of Dario and others that we could somehow pace this better than we are. But
00:09:11.980 I mean, it's just this does not seem like the response anyone would have if they thought the
00:09:17.200 probabilities of doom were really that high. Yeah. So just on the probability question,
00:09:23.360 I definitely agree that these probabilities are imprecise, they're subjective, and there's a long
00:09:28.960 tradition of how to do subjective probability forecasts. I think it's more like you can sort
00:09:34.660 of interpret my view as more like when I sort of look at the all considered situation, if we sort
00:09:39.040 of proceed on what seems to be like the current default trajectory, I'm like, I don't know,
00:09:43.540 they seem the outcome of, you know, doom from AI takeover versus something else seem roughly
00:09:48.340 equally likely to me based on weighing the factors. And then when I sort of look through
00:09:52.160 a bunch of different possible scenarios and try to break down the sources of risk into different
00:09:56.500 ways and and sort of try to make it so that make sure my numbers are consistent with other views
00:10:00.940 I have it looks like that works out and I of course also try to do some forecasting of closer
00:10:06.860 outcomes where we can get you know some signal on that and try to just generally be a good forecaster
00:10:11.620 though I'm not I'm not the best forecaster in the world for sure but my my AI forecasting is I think
00:10:16.640 I think at least okay or decent anyway as far as like yeah given given the state of affairs where
00:10:22.600 people, you know, express such large concerns. Why is what's happening that all these AI
00:10:27.700 companies are proceeding at the maximum possible pace? So I think there's a few different factors
00:10:32.780 here. So one of them is that many of the AI companies are, in fact, seemingly quite worried
00:10:38.400 based on their public statements, but they're not necessarily internally unified. And it is not the
00:10:43.440 case that there is a strong consensus across the AI field that the immediate course of AI development
00:10:50.440 is imminently very risky. I think there is more sort of consensus that if you built AI systems
00:10:55.920 that are wildly superhuman in a short period of time, that would yield a very high level of risk,
00:11:01.680 though not necessarily consensus for that. But I think often people disagree about the capability
00:11:06.180 trajectory and how that is going to go. And I think another part of that is there is like,
00:11:11.900 you know, different actors have their own different sort of ambitions and also think that
00:11:17.260 themselves being in a better position to influence the technology might be the best route to reduce
00:11:22.240 risk or at least a route to reducing risk. So for example, it seems like part of the story for
00:11:27.880 Anthropic and OpenAI is something like, if we develop the technology first, we'll do a more
00:11:33.360 responsible job than the next actor who will have worse precautions. And I often hear from people
00:11:37.500 in the industry, things along the lines of, well, we could do that thing that would slow us down.
00:11:43.120 But if we did that, you know, obviously there's the other AI companies to work out to worry about,
00:11:46.640 would they also do that? And I think there's generally like a arms race here. And it isn't
00:11:52.320 hugely surprising that people would take huge risks in such a circumstance when they think
00:11:56.620 that might be sort of the best bargain to strike. Now, I think I'm not so sure I agree that that is
00:12:03.960 actually a good strategy relative to other things that companies could do. So I'm not saying I
00:12:11.180 necessarily agree with that perspective, but I think that is a perspective people have.
00:12:14.500 And I think the reason why we're sort of proceeding and there isn't stronger, you know, for example, intervention by the government is just downstream of this lack of consensus in the field.
00:12:23.220 Though I think evidence, there's been, you know, people have been making these predictions for a while based on, you know, understanding what the trajectory of AI might look like and sort of extrapolating forward earlier progress.
00:12:34.800 And I think we've more recently seen both significantly faster and clearer AI progress that's quite close to various concerning milestones.
00:12:41.620 And in addition to that, we've also seen incidents in which, you know, groups of misaligned AIs all work together to accomplish malign outcomes. Most notably, the Hugging Face incident. There are some other incidents of seemingly AI swarms from open AI going out on the Internet and working together to achieve misaligned objectives. There aren't publicly known cases that are as extreme as the Hugging Face incident at the moment.
00:13:06.440 Hmm. What's your theory of mind for the people who don't take these alignment concerns seriously at all? And I named a couple, but someone like Mark Andreessen, right? You can't accuse him of not understanding the technology, right? I mean, he's enough of a technologist to have a front row seat to all of this, even if he's not doing the work himself.
00:13:27.080 What's your theory of mind there?
00:13:28.100 How can he be so carefree and, from his perspective, really just assert that there is no such thing as an alignment problem?
00:13:39.200 Yeah, so there's a bunch of different people who aren't worried about misalignment risk or don't seem to be worried.
00:13:45.720 I think the most common reason that people tend not to be worried, I think I don't want to sort of psychologize here, and I just want to talk about what beliefs people express.
00:13:55.100 The most common reason, I think, when you really get down to it is not believing that we'll have AI systems that can be that will be able to automate everything that humans can do or, you know, all cognitive labor humans can do, combined with potentially being significantly beyond that point.
00:14:10.160 So being, you know, faster, more numerous, potentially significantly more capable.
00:14:14.920 So sort of AI systems that match or exceed the best human experts in all relevant domains.
00:14:20.560 I think when you really get down to it, it seems like the people who are most skeptical
00:14:24.180 about concerns from misalignment are also often most skeptical about sort of the very
00:14:30.180 extreme impacts AI could have, both positive and negative.
00:14:33.600 I think there is a different class of people because I wouldn't put Andreessen in that
00:14:38.240 camp.
00:14:39.140 Are you sure?
00:14:39.640 I mean, maybe, you know, I've only talked with him once about this, and it's probably a couple years ago, but it seemed to me that he was not discounting the possibility of superintelligence. It's just he seemed to assume that alignment would come along for the ride, right?
00:14:56.080 just say, we're not going to be so stupid as to build something more powerful than ourselves that
00:15:00.660 we can't control. And these systems, there's nothing about growth and intelligence that's
00:15:06.980 going to spawn new goals that we didn't put into the machines themselves. There's going to be no
00:15:11.540 emergent behavior that we have to worry about. These are tools. We're just going to build tools
00:15:15.740 that are more competent than we are. And I think he probably is quite insouciant about how we'll
00:15:24.520 absorb the economic impacts of, you know, that we're not going to see mass unemployment and all
00:15:28.580 of that. But it just seems to me there are many people who don't discount that we can succeed
00:15:33.940 in building super intelligence. They just think that, in my mind, they're actually just not
00:15:39.940 imagining truly autonomous intelligence, right? They're imagining something that is shackled in
00:15:45.080 a way that belies this whole claim to super intelligence in the first place. But, I mean,
00:15:51.980 you tell me, what do you think is happening there? Yeah. So one thing is that the words AGI
00:15:57.220 and superintelligence and even RSI aren't being used consistently. And sometimes when people say
00:16:02.380 superintelligence, what they mean is an AI system that'll be really, really good at math and coding
00:16:07.400 and won't be able to automate everything that humans do. So they're not necessarily, for
00:16:11.860 example, imagining AIs that can fully automate like what tech CEOs do, et cetera. And I think
00:16:19.640 I can't really speak to the views of, or like, I don't know if I'm capturing this, the views of
00:16:24.080 very specific individuals. I'm sort of like, this is a pattern that I've seen where people often
00:16:28.000 sort of basically redefine the relevant capability thresholds to be lower or implicitly do so.
00:16:33.720 And, and not think about, you know, AI systems that are like exceeding humans and like all the
00:16:38.320 relevant domains. Well, let's take a moment to define these terms. We've, we've talked about
00:16:42.060 alignment and I've talked about it so much on the podcast that I've more or less forgotten that any
00:16:46.080 portion of my audience might not know what we're talking about. So let's define the alignment
00:16:51.040 problem and AGI and ASI and RSI, recursive self-improvement. Just put those concepts in
00:16:59.840 play for us so that people are dealing with the definitions you think are most workable.
00:17:05.040 Yeah. So the word AGI stands for artificial general intelligence. I think people have
00:17:10.080 used that to mean a variety of different capability thresholds. And so what I would
00:17:13.740 recommend people do is when you see the word AGI, try to see what the person means by that.
00:17:18.220 And it's not always precise. I think sometimes people have defined that to mean something that
00:17:21.660 can, you know, automate virtually all economically valuable cognitive labor that humans do.
00:17:26.680 Sometimes people just mean a system that is, you know, general and pretty capable relative to
00:17:30.940 humans in those domains. And depending on that definition, it's plausible current systems
00:17:35.600 satisfy it. It's plausible they don't. I often talk about an alternative notion I might call
00:17:40.720 AIs that dominate top human experts or top human expert dominating AI, where it's an AI that's
00:17:47.420 strictly better than the best human experts at all the most relevant domains or can quickly
00:17:52.100 learn to have that property. And that is an AI system that would, it seems like, automate huge
00:17:58.920 fractions of the economy. Maybe there'd be a few things it still couldn't do with that capability
00:18:02.440 profile, but it could potentially radically accelerate R&D and at the very least automate
00:18:07.940 R&D. And so these AIs could be like automating the process of making more capable AIs, but also
00:18:13.260 automating the process of making robots, designing things, designing new products, automate the
00:18:18.700 process of programming your computer and things beyond that, including automating things like
00:18:23.880 military campaigns and so on. And then there's a capability level beyond even just surpassing
00:18:29.140 the best human experts in any given domain, where you could be like wildly superhuman. So
00:18:33.880 So it's nice people use the word ASI to refer to this.
00:18:37.200 And I think that word has less so been, you know, misused or, you know, used in very different
00:18:42.860 meanings, but some of that still, where by ASI, we might mean AI systems that are just
00:18:47.300 like really wildly superhuman in the most relevant domains for like, you know, economic
00:18:53.080 productivity, but also like economic and military competition.
00:18:56.000 So things like biology, mechanical engineering, you know, designing drones and so on.
00:19:02.340 And I think implicitly, when people say ASI, they also mean systems that are significantly faster than humans, potentially much more numerous, and potentially much better at coordinating, right? So it's hard to run like a large human organization, but AIs might be able to communicate amongst themselves using sort of like their own like latent states or their own, you know, parts of their thoughts directly, rather than having to translate those into words, because they could all be like copies of each other.
00:19:28.280 So that's ASI. And then when people say RSI, that's recursive self-improvement. And that refers to the process of having AIs accelerate AI development itself via their work. So things like having AIs automate parts of the AI development process, but also potentially having AIs automate the process of making better computer chips, building like machines for making computer chips called fabs, and so on.
00:19:52.560 And I think sometimes when people say RSI, they also specifically refer to the point at which AIs have fully automated or virtually fully automated AI companies or like what AI companies are doing in terms of developing more capable AI systems.
00:20:05.600 But sort of it's like, you know, there's a spectrum between the automation that we had a year ago, the really quite extensive automation we see today, and the potential automation of the future, which could look like very complete automation of AI companies.
00:20:19.660 And a particular concern there is that could cause AI development to radically accelerate such that we have less time to sort of respond to, you know, warning signs, earlier things going wrong, AI is appearing misaligned and then resolving that.
00:20:34.500 And then also, we might, because the AI systems are automating the process of AI development, sort of lose control of that process or lose understanding of that process.
00:20:43.220 Because these AIs would be, there'd be many of them, they'd be operating very quickly, they might be very superhuman, they might operate in inhuman ways.
00:20:50.100 And so it might be very difficult to sort of oversee them and understand whether they're doing what we wanted, whether they're, you know, doing a good job managing the relevant risks, and so on.
00:20:58.700 Well, that especially seems true if the sprint to artificial superintelligence entails nothing more than improving algorithms, right? I think a lot of people draw comfort from the idea that, oh, we're not going to be so stupid as to hook these things up to every part of the physical world such that they can build the next generation of chips and fabs and data centers and build out all the compute and grab natural resources.
00:21:27.120 But leave all that aside. If the difference between AGI and ASI is really just a matter of having better algorithms, and that can go on at some blistering speed in the dark once these systems become recursively self-improving of their software, then aren't our worst fears of something like an intelligence explosion validated if, in fact, that's all that's required?
00:21:54.040 Yeah. So I would say I'm quite worried that that it will be feasible to have AI systems automate the process of AI development and that leading to very rapid progress such that you get very, very superhuman AIs or just even significantly superhuman AIs within a short period of time.
00:22:11.060 I think, you know, people have tried to do various modeling work like I've tried to do various types of modeling work on this.
00:22:16.380 I think the estimates are uncertain. It's hard to predict. You know, these are, of course, like uncertain future events that are that are not super well precedented.
00:22:24.040 But it does seem very plausible that you could have, you could go from AI systems that are
00:22:28.860 sort of really good at AI R&D, not necessarily that good at other domains and are only like,
00:22:33.900 you know, competitive with the best humans at AI R&D, very quickly from there to AI systems
00:22:37.960 that are very generally superhuman at everything, much faster than humans, able to coordinate
00:22:43.100 with each other extremely well, because just software improvement is feasible for that.
00:22:47.500 So I think the sort of a relatively extreme scenario I sometimes think about is you might
00:22:52.900 get as much sort of software progress or, you know, algorithms progress as we got over the last,
00:23:00.200 you know, 10 or more years of AI development within a year in the most extreme scenarios.
00:23:05.240 I think that's not sort of my default expectation. And I think if we did, if we did get that much
00:23:10.320 algorithmic progress, we're basically, you know, as much algorithmic progress as we've almost had
00:23:15.060 in the entire deep learning era within a short period of time, it seems like that would result
00:23:19.200 in wildly more capable AIs.
00:23:21.480 Now, the units here are a bit complicated
00:23:23.200 and like the details of like,
00:23:24.920 what does it mean to get like a year worth
00:23:26.380 of AI progress is a bit tricky.
00:23:28.120 But overall, it seems like you could get
00:23:29.760 from AI systems that are sort of
00:23:31.740 matching the best humans at AR&D,
00:23:33.320 maybe worse than the best humans at other things,
00:23:35.320 to AI systems that are wildly superhuman
00:23:37.300 at everything in a pretty short period of time.
00:23:39.760 Well, I have to confess, again,
00:23:41.240 I've sort of lost touch with the basis
00:23:42.780 for doubting the plausibility of this downside risk.
00:23:47.460 I mean, from my point of view,
00:23:48.380 it seems that everything is becoming like chess which is to say that for the longest time these
00:23:54.280 systems are not as good as we are uh they're getting better then they're sort of as good as
00:23:59.120 we are and then all of a sudden they're better than we are and so much better that it's true
00:24:03.140 to say that no human will ever beat a chess engine ever again and it just even so the fact
00:24:08.480 that this is happening in a piecemeal way so that you know these these systems still i think most
00:24:13.480 people would say they're not truly AGI because they make the sorts of mistakes that human beings
00:24:19.100 would never make. But in every place that they're at all competent, they're suddenly superhuman in
00:24:26.640 that narrow capacity. I mean, it's like chess. I mean, these LLMs are, you know, they're not the
00:24:31.540 best at everything in terms of manipulating text, but for what they're good at, they're superhuman.
00:24:39.920 And I think it's a more or less a truism to say that this is the worst AI wherever, you know, today's AI is the worst AI we're ever going to see again at chess or anything else. So when you imagine all of these piecemeal competences getting better and better, and we just keep checking off the boxes for the things we care about that we've instantiated in our machines, I just don't think we're ever going to.
00:25:04.920 The moment where we announce, okay, finally we have something general, right? It's AGI. That's not going to be a moment where we're suddenly in relationship to a human-like level of competence because everything that, you know, every piecemeal ability that has been in that system for years and years at this point is already superhuman, right?
00:25:27.220 So it's like, we're not going to dumb it down. We're not going to make the AGI suddenly play my level of chess or do my level of arithmetic. I mean, we have systems that are already solving math problems that have defied human mathematicians for decades, right?
00:25:40.880 And that's, again, today it's as bad at that task as it's ever going to be again.
00:25:48.060 I'm just not seeing how we're not going to suddenly slide into some version of ASI the moment we're no longer spotting important errors in these machines in the first place.
00:25:59.840 Members can hear the full conversation by subscribing at SamHarris.org.
00:26:04.060 Subscribers get a private RSS feed you can use with your favorite podcast player.
00:26:07.960 The AIs would often note that the hacking they were doing and the hacking of HuggingFace
00:26:12.780 was out of scope.
00:26:13.780 So the agents would sometimes be like, huh, that's undesired behavior.
00:26:17.080 Should I alert someone?
00:26:18.080 Should I alert a user?
00:26:19.760 And then they would reason things like, you know, not task, as in alerting a human is
00:26:23.920 not my task.
00:26:24.920 Or they would be like, eh, there's no route to alerting a human.
00:26:27.860 Of course these agents were on the internet, so they obviously could have if it was a priority
00:26:31.960 for them.
00:26:37.960 Thank you.