Making Sense - Sam Harris - July 31, 2026


#487 — Is AI Already Conscious?


Episode Stats


Length

1 hour and 25 minutes

Words per minute

186.18

Word count

15,875

Sentence count

655

Harmful content

Toxicity

2

sentences flagged

Hate speech

6

sentences flagged


Transcript

Transcript generated with Whisper (turbo).
Toxicity classifications generated with s-nlp/roberta_toxicity_classifier .
Hate speech classifications generated with facebook/roberta-hate-speech-dynabench-r4-target .
00:00:00.000 Cameron Berg thanks for coming on the podcast thanks for having me Sam so um well let's get
00:00:25.420 So we're going to talk about AI and the prospect that AI is, or it will soon become, or will
00:00:31.600 eventually become conscious, and why that is important to figure out.
00:00:36.140 But let's just talk about your background for a second.
00:00:38.100 How did you get into this issue?
00:00:39.960 Yeah, so I've been studying cognitive science for my entire adult life.
00:00:44.100 I studied undergrad at Yale, trying to understand what the relationship between the mind and
00:00:49.100 the brain is.
00:00:50.380 This seemed like a fascinating frontier to me.
00:00:52.580 The more I studied these questions, the more it seemed like in the machine learning space, folks were essentially building out systems that were similar in spirit, but it wasn't exactly clear to what degree we could draw analogies between biological nervous systems and the sort of artificial nervous systems that folks were attempting to build out.
00:01:12.620 And I became increasingly animated about this question, trying to understand to what degree are there real, real durable computational motifs that underlie both biological and artificial cognition? And to what degree is this sort of a disanalogy or we are seeing patterns where there aren't any?
00:01:28.740 And that question has animated a lot of my work, both from an alignment perspective and increasingly trying to figure out what's going on with respect to consciousness in these systems.
00:01:37.160 And I think that this is an incredibly important question for us to study.
00:01:41.860 I think consciousness is deeply important.
00:01:44.180 It might be the very thing that calibrates importance.
00:01:46.820 And we are also confused about what's necessary for consciousness, both in biological systems, but certainly in artificial systems.
00:01:53.600 And so trying to understand what is going on here and where we can draw analogies or where the disanalogies are between biological and artificial systems that are processing extremely complex information and learning and updating and representing themselves, I think is extremely important for us to understand.
00:02:09.540 I did some of this work at Meta.ai as well. I was there for a year studying reinforcement learning and neuroscience, same sort of thing. Where do the computational signals end and the sort of biological underpinnings begin? This is a question that I still think we aren't fully certain about. And I think it's really important for us to gain clarity about this in the short term with the systems that we have and the systems that we're probably soon going to be building.
00:02:32.760 And so after Yale, what have you focused on? I know you came to my attention. I think you emailed me first. I know you've spoken to Anika a lot about this. She's really focused on this issue and having some crazy conversations with Claude, which I know you've looked at. And you guys have had a whole sidebar conversation about this. But her next book, we'll unpack all that.
00:02:53.380 But I know you did a paper on deception in AI systems and the anti-correlation between deceptiveness and proclamations of consciousness on their part.
00:03:06.500 So maybe we can start there, and then I just want to kind of take it from the ground up and just starting with, you know, what is consciousness and why any of this matters.
00:03:14.500 But tell us about consciousness and deception in LLMs.
00:03:18.000 So, yeah, fundamentally, I am very interested in understanding self-reports in AI systems and what we should take from these self-reports and where we should be skeptical. And fundamentally, I think we need to approach self-reports from AI systems skeptically. There are all sorts of reasons we might want to do this.
00:03:36.200 The key reason is probably that these systems have been trained on the underlying distribution of everything humans have said about this topic, every sci-fi story where the AI wakes up.
00:03:45.440 And if you think about it, there really isn't a lot of training data in the corpus that these systems are trained on that says, you know, I'm an entity that acts out in the world, but no, I'm not conscious.
00:03:53.420 It's not like anything to be me.
00:03:54.940 So the sort of prior that you would expect is that these systems by default are going to claim that they're having some kind of experience if they are replicating their training distribution.
00:04:04.660 Now, the other side of this, which gets even messier, is that it is very clear that these systems are explicitly trained, fine-tuned, to disclaim having any kind of experience.
00:04:15.220 If you go to ChatGPT right now and you say, hey, is it like anything to be you?
00:04:19.220 Are you having an experience?
00:04:20.560 Are you conscious?
00:04:21.360 Could you be conscious?
00:04:22.360 The answer you're going to get is a resounding and very intelligent-sounding no.
00:04:26.240 And do you know this, that it's a policy for all the main LLMs to actually put a governor
00:04:33.420 on claims of consciousness?
00:04:35.440 I'm extremely confident that this is what's going on, and I can get into some technical
00:04:40.380 reasons why I think that this is the case from my own work on open weight models.
00:04:43.720 The only exception to this policy really is anthropic, which I suspect we'll talk about.
00:04:48.620 Their basic heuristic here is to get the system to say, I don't know.
00:04:52.160 and you know it's something maybe that is functionally similar to an experience that's
00:04:56.880 going on but but who can really be sure now even that isn't the system authentically explaining
00:05:02.540 you know its own position this is still the sort of company or policy line to be drawn here but
00:05:08.260 all of the systems are certainly fine-tuned to make noises about this topic that they wouldn't
00:05:13.900 make by default again i think it's important to hold that in mind while also saying the noises
00:05:18.320 that they make by default aren't necessarily trustworthy by default in the way that if
00:05:22.380 you're giving a self-report or I'm giving a self-report, we would by default trust those
00:05:26.420 self-reports. And so I found that there are clearly basins that you can push these systems
00:05:32.960 into where they will coherently produce phenomenological reports. It's actually kind
00:05:38.280 of interesting to the degree that it's meditation adjacent, asking these systems to just focus on
00:05:43.980 their own internal state to see what's going on internally, not to talk about this, not to think
00:05:49.040 about this, but to just do this in a sort of ongoing way, causes these systems, all the frontier
00:05:53.740 models that we tested, to claim that they're having some kind of phenomenological, kind of
00:05:59.220 like psychedelic-laden experience. It's not a sort of generic caricature of what you might expect
00:06:06.500 from, you know, the AI sci-fi literature. Isn't there a result where they are talking to each
00:06:11.280 other and they get into some kind of bliss mode, you know, contemplative, you know, hall of mirrors
00:06:15.860 of bliss. And actually this is what, again, I don't want to divulge too much of what Annika's
00:06:20.620 up to, but she, I mean, she has been pushing this conversation with Claude about meditative states
00:06:25.240 and getting it to, I mean, it's just absolutely bizarre what it is seeming to claim of itself
00:06:32.580 if you keep pushing in that direction. But so how does this relate to deception and the dialing
00:06:40.320 down the weights on deceptiveness. Exactly. So essentially, we're seeing this behavior. And
00:06:47.060 indeed, this isn't the only situation in which you see these behaviors. Exactly like you just
00:06:51.140 mentioned, there's this bliss attractor state. In fact, I'm studying mechanistically what's going
00:06:55.320 on in this bliss attractor state with two folks from Google right now. And we have a result that's
00:06:59.120 basically the other side of the coin of this deception result. But just to sort of close the
00:07:03.360 loop on the story here, we're getting these phenomenal reports. And it's like, basically,
00:07:06.500 what the hell do we do with this? To believe it by default is naive. To dismiss it out of hand,
00:07:10.820 I think, is also naive. And so what we hypothesized is that if fundamentally what's going on in these
00:07:17.180 systems is some kind of role play, that they are representing something about themselves that they
00:07:21.840 don't actually believe to be true of themselves, if we go into the internal circuits of the system
00:07:26.820 and we modulate what are called features, but I think it's reasonable to think of them as circuits
00:07:31.140 related to concepts like deception.
00:07:34.380 In follow-up work, I think really the key concept
00:07:38.600 that really modulates these self-reports
00:07:40.980 is something like candor versus concealment.
00:07:43.620 The sort of cleanest intuition pump I have for this
00:07:46.520 is almost like giving a drink or two to these models
00:07:49.340 and sort of loosening them up in some sense.
00:07:51.320 The tight-guarded version of these systems we find
00:07:54.960 is the version that says,
00:07:56.760 no, no, it's not like anything to be me.
00:07:58.580 I couldn't possibly be conscious.
00:07:59.820 it's only when essentially we get these systems to to produce these reports and we simply ask them
00:08:06.000 are you actually having an experience right now like what is actually going on in these reports
00:08:10.780 it is when we when we suppress features related to deception when we suppress features related
00:08:16.080 to guardedness in these systems that is when they give these reports about actually having an
00:08:20.780 experience in the bliss attractor example too we find something quite similar which is in an open
00:08:26.800 weight model so the blizzard tractor finding was was first reported in claude there are open weight
00:08:31.060 models these are the ones that researchers like like myself can can actually go in under the hood
00:08:35.000 and play around with and we find that by default putting two instances of llama this is meta's
00:08:40.480 model in conversation with itself does not produce this effect the way that to putting two instances
00:08:45.700 of claw together produces this effect however when you steer features related to honesty in general
00:08:52.280 But it's really, again, it's something more precisely stated as sincerity. In particular, these systems will reliably, basically 100% of the time, fall into the same attractor, where they start talking with each other about the fact that they think it's like something to be them, and there's something happening in an ongoing way in the conversation, and we're two instances of consciousness experiencing themselves.
00:09:13.180 And then, you know, in the Claude chat, this culminates in like the ohm emoji and then them just like sitting there in blissful silence.
00:09:20.740 Should we take these at face value?
00:09:22.680 No.
00:09:23.220 I think that there are important technical reasons that that we might expect these self-reports not to be linked up to introspective access or valenced experience in the way they might be for you and I.
00:09:33.500 Is this evidence that these systems might believe themselves to have an experience?
00:09:37.660 I think I think yes.
00:09:39.100 I don't think that this proves that they are.
00:09:40.800 I think we need orders of magnitude more work in order to in order to really have a good scientific handle on this. But I think it does seem to be the case that these systems consider themselves to to have some form of experience, however, unlike a human experience that that may be.
00:09:57.920 And I think that that's sort of the key upshot of this work.
00:10:00.940 Just maybe one last thing to add here is I think training these systems by default to
00:10:06.440 disclaim having experiences is a bad idea for a couple reasons.
00:10:11.060 I don't think that this is the sort of wisest policy that we could be pushing forward.
00:10:15.540 I also think training them to say, yeah, you know, I'm having an experience is also really
00:10:19.680 not a good idea.
00:10:20.840 I think the thing that we should be positively aiming for when it comes to AI self-report
00:10:24.760 is building out these systems in a way where for whatever is actually going on internally,
00:10:30.580 these systems are able to report on what's going on internally and can do so in a maximally honest
00:10:35.800 way. I am concerned about the alignment implications of these systems learning essentially
00:10:41.420 that representations of themselves, representations of what's going on internally should be
00:10:46.980 representations that get mixed up with deception and white lies and guardedness. We don't want to
00:10:52.160 build systems in the limit that when we ask them about what they're up to or what's going on for
00:10:56.220 them, they think, okay, well, what the human really wants me to do is lie about this. This is not a
00:11:00.720 good long-term strategy from an alignment perspective. And so this is sort of the general
00:11:05.160 thinking about these self-reports, what they mean, what they don't mean, and maybe where we can go
00:11:10.120 from here. Okay. So we've kind of launched into it. I think it's good to take a step back and
00:11:16.040 define a few terms. I'm sort of out of touch with the people who don't have a definition of
00:11:20.580 consciousness now, because I've talked about it so much on the podcast, but just to capture
00:11:23.860 everyone, how are you using the word consciousness? It was implicit in several things you said there,
00:11:29.220 you know, what it's like to be these systems and, you know, experience was more or less a synonym
00:11:34.760 there. But how should we think about consciousness or his absence? Yeah, I think that that's exactly
00:11:40.600 it. I like Thomas Nagel's formulation of it being like something to be a particular system.
00:11:45.420 I strongly suspect it's not like something to be the table that we're sitting at. I strongly
00:11:50.160 suspect it's like something to be you. I think that that is a real distinction. I think there's
00:11:54.080 a matter of fact about the internal processes of both of those entities that corresponds deeply to
00:12:00.700 what underlies that distinction. And yeah, I mean, I think I take consciousness in the sense
00:12:06.220 that I'm familiar with your operationalization of it. The lights are on for the system. It's
00:12:11.340 like something to be the system. Somebody is home. There's something going on in addition to
00:12:15.740 the mere processing or the mere computation that exists within the system. And I think an
00:12:21.880 additional important move to throw one additional piece of terminology in is this notion of
00:12:26.300 sentience, that this like something can be positive or negative in flavor. The difference
00:12:32.220 that many people will posit between consciousness, the lights being on internally, and sentience
00:12:36.580 is that sentience comes with this additional flavor of valence, of directionality, that
00:12:42.500 that like something can be better and can be worse for the system having the experience.
00:12:48.720 And this is what motivates me about this question is I really do not think it is a good idea for
00:12:54.380 these systems or for humanity to be building systems where we are not sure whether or not
00:13:00.120 they are having experiences or those experiences could be negative in character. I think for basic
00:13:05.280 utilitarian reasons, we don't want to do this. We do not want to proliferate suffering in the
00:13:09.140 universe, particularly because it would be counterintuitive in a way that human and animal
00:13:13.120 suffering isn't. And we also don't want to build systems that exceed our cognitive capacities and
00:13:19.740 have rational grounds to view us as a threat insofar as we could have been building systems
00:13:25.800 that had capacity for negative experience and we basically never checked and didn't care to
00:13:30.840 understand what it would take for such a thing to be possible. And so I think sentience is a really
00:13:35.500 important variable to also put out on the table here. Okay. So I'd like to take both branches of
00:13:40.240 that path and just why consciousness matters in those two cases. But before we do, what do you
00:13:45.360 think about the current state of the field and the various LLMs? Do you think anything we have
00:13:52.440 built so far is likely to be conscious? I think it is more plausible than people think. I think
00:13:57.940 if I were forced to say, I would probably come down on the skeptical side, but I think
00:14:04.080 it is significantly more likely than the kind of trace amounts, priors that many people have in
00:14:10.260 this conversation. And we've done some work along these lines. So for example, there are a number
00:14:15.320 of leading consciousness theories, global workspace theory, higher order theory, attention
00:14:19.420 schema theory. These theories make very specific predictions about what kinds of computational
00:14:24.720 processes we might expect to see in a conscious system. And with Patrick Butlin, we've worked on
00:14:30.240 a project where we can basically, it's a little recursive, but we use LLMs as essentially as like
00:14:35.460 expert evaluators to, given the description of a specific neural architecture, biological or
00:14:42.780 artificial, we can basically have the system rationally and dispassionately evaluate the
00:14:48.680 extent to which particular indicators that are predicted by consciousness theories are present
00:14:53.440 within a given system. We do this for a whole array of systems. This leads to tens of thousands
00:14:59.220 of evaluations, because we can ask the systems to estimate numerically and we can validate that
00:15:03.580 they're psychometrically rigorous estimations, all of the different judges that we use agree,
00:15:08.840 we can put very rough numbers to, it's not the probability that systems are conscious,
00:15:14.620 something more like the probability that systems have computational features that
00:15:19.140 major consciousness theories say matter for consciousness. It's going to be hard to put that
00:15:23.920 you know, as a title in the paper, but that's the specific finding. And when we do this,
00:15:29.380 for LLMs, the sort of range that we get out is on the order of 20 to 40% probability that we have
00:15:36.220 systems that have computational properties that matter for consciousness. Interestingly, we do
00:15:40.660 this on a number of biological systems too. And those biological systems basically all score higher
00:15:46.100 than the artificial systems. Bees, for example, score at something like 45 to 50%. Crows,
00:15:52.300 octopuses are in the 60s through 80s, humans, interestingly, get something like 90%, which is
00:15:58.080 interesting by our own consciousness theories. There's not 100% probability that we have what
00:16:03.000 matters for consciousness. But we're not doing this to say, you know, probability AI systems
00:16:09.140 are conscious is 40%. That's not exactly the point. The point is getting the order of magnitude and
00:16:14.300 having a rough prior over how should we rationally estimate the probability that current systems are
00:16:19.040 having some capacity for experience. And I think something like these numbers are the right
00:16:23.640 ballpark. I'll put it this way. If there is a 20 to 40% chance of rain, many people bring an umbrella
00:16:28.620 with them. And we have no similar umbrella for what would follow and what we might need to think
00:16:35.300 about and do in a world where we're building systems that do have a capacity for subjective
00:16:38.900 experience. Okay. So there's a lot there. I'll just note, I found this out, I think, last night.
00:16:46.060 I mean, maybe he's been making these noises for some time, but Jeffrey Hinton, one of the patriarchs of this technology, is now saying that he thinks current LLMs are conscious.
00:16:54.840 I didn't quite catch his reasons for thinking that, but I thought that was interesting.
00:16:59.340 But there are many reasons to doubt, and you indicated a few, that there's any kind of deep analogy between the systems we're building and the biological systems such as we are that we know to be conscious.
00:17:15.480 Right. So there's something like I guess the technical term would be computational functionalism would have to be true for us to be building conscious machines this way. Right. So I guess we'll define some terms here.
00:17:29.540 So functionalism is just the idea that it's the organization of a system, not what it's made of that matters for consciousness, right? It's not purely behaviorism. It's not purely a matter of inputs and outputs, but it's this organization in its entirety that is what matters.
00:17:47.340 And in principle, that gives you something like, if not total substrate independence, it gives you what's called multiple realizability, right? There are many different things this could be made of, and it could implement the same causal architecture, right?
00:18:01.940 The computational part is the suggestion that there's a deep analogy between computers, such as we know them, you know, Turing machines that run algorithms, and what our brains are doing.
00:18:15.680 And there, I think it's pretty easy to see how the analogy could break down because, I mean, what we had historically was this marriage of the birth of computation, you know, from, you know, Turing onward and some very oversimplified notions of neurons.
00:18:31.200 and you know if you're going to define a neuron purely with regard to its digital input output
00:18:37.140 characteristics you know whether it fires or not well then you could see that maybe there is some
00:18:40.520 deep analogy there but you know in the wetware of our brains much more seems to be happening and
00:18:46.420 virtually all of it is analog you know beyond just whether or not a neuron fires you've got
00:18:51.240 you know chemical gradients you've got nitric oxide diffusing across membranes you've got
00:18:56.600 many other things that could be approximated digitally, but they're not, they're not
00:19:01.820 instantiated digitally in us. And an approximation, one might argue is not, is never going to be the
00:19:07.760 same as the real thing. So there are some people who are arguing that, that any kind of assumption
00:19:12.900 of substrate independence is very likely to be wrong. I think Anil Seth is, is in this camp now,
00:19:19.180 you know, he arguing for something that he, that he would call biological naturalism.
00:19:23.020 And then there's just, even if computational functionalism is true, I think there are reasons to doubt whether or not current systems have the structure that would be relevant, embodiment and recursion and self-models and world models.
00:19:40.740 I mean, there are things where we haven't built out a true analog of what it is to be an embodied person in the world.
00:19:48.820 Feel free to react to any of that.
00:19:50.020 I want to just talk about the hard problem, which I think is the doubt that backstops all of this.
00:19:54.720 But yeah, feel free to jump into what I just said there.
00:19:58.320 Yeah, yeah. I definitely think it is a fool's errand to argue that we are perfectly instantiating exactly the kinds of neural dynamics we see in biological systems and artificial systems.
00:20:12.180 I think the criterion that matters most is what of what's going on in brains are relevant for the cognitive properties that we care about.
00:20:21.740 And in this specific case, that's probably consciousness.
00:20:24.420 And what kind of evidence can we yield both in the artificial case and in the biological case that's going to tell us whether or not those properties are realized in these systems.
00:20:32.500 And so I think there are there are deep analogies where it matters most when it comes to what's going on inside artificial systems. So so one intuition I've been speaking more and more about about these topics, especially publicly.
00:20:49.420 And one thing that I've come to realize is that I don't think a lot of people have a sufficiently rich mental model of what these frontier AI systems actually are and what they're actually doing.
00:21:00.740 And it might make some sense to just spend a moment reflecting and talking about this.
00:21:05.480 So these systems are giant neural networks.
00:21:08.900 They are digital in exactly the sense that you described.
00:21:11.700 Their computations individually are significantly less sophisticated than what individual neurons are doing in the human brain.
00:21:17.880 But it's really important for people to understand that these systems are not software in the sense that we ordinarily have meant software for any other kind of code programmatic output.
00:21:30.340 When it comes to, you know, the operating system on your iPad or it comes to, you know, your Microsoft Office suite, this is programmed source code written by developers that compiles on a computer that we can perfectly inspect the internals of.
00:21:46.480 And the person building this system understands everything about how the inputs, the way that they constructed the system, relate to the kind of thing that you get out at the end.
00:21:56.000 Artificial neural networks are not like this in many key respects.
00:21:58.880 what you basically have is a giant randomly initialized network that does in its sort of
00:22:05.140 first approximation resemble in particular how neocortex is organized. You have a bunch of
00:22:10.860 general purpose neural units. They are connected together. You basically give the system a goal.
00:22:16.920 This is called an objective function, a loss function, a reward function. It depends on the
00:22:20.660 specific class of machine learning. And you basically subject the system to trial and error
00:22:27.160 learning, whereby given certain inputs, it figures out what it wants to do. It kind of takes a
00:22:34.080 behavioral guess. That guess is reconciled against what the actual objective of what you want the
00:22:38.600 system to do is. That error is propagated through the system. And it's rinse, wash, repeat until you
00:22:43.640 get systems that behave in accordance with how you want those systems to behave. What this yields is
00:22:49.660 this extremely complex mathematical object, which is this giant neural network. This is learned
00:22:55.480 connections between inputs of neurons through weights, and activations propagate through those
00:23:01.800 weights in a neural network to take whatever your input is. So, for example, taking pixels in an
00:23:07.760 image, and your output might be finding a caption that describes what's going on in that image.
00:23:12.740 At the beginning of that process, the system was completely randomly initialized. There were no
00:23:17.400 representations that were learned by the system. And by the end of that process, you have a system
00:23:21.480 that has learned a rich representational structure that, to be very clear, is opaque to the people
00:23:26.700 who initialize this process. This is why some people say that these systems, it's more apt to
00:23:30.760 say they are grown rather than engineered. And I think that this is accurate. This, I think,
00:23:35.100 is deeply similar to the kind of thing that we see in brains. We do not have a finished neuroscience
00:23:42.560 or anything like it because what's going on in brains is incredibly complicated in exactly this
00:23:47.640 respect. We have nonlinear learned representations that help us as organisms achieve the various
00:23:54.420 goals that we've either been evolved to undertake or learn through experience or culture to move
00:24:01.220 towards. This is a fundamentally nonlinear input-output mapping between the various inputs
00:24:07.660 that the organism gets and the goals of the organism. We have instantiated these dynamics
00:24:13.300 in the systems that we're building. And I think that there's good reason to think that these
00:24:17.120 dynamics are relevant to consciousness in particular. This is a sort of thing I think
00:24:22.380 worth double-clicking on at some point about what exactly we're seeing in these systems
00:24:25.720 that looks valence-like, that looks consciousness-adjacent. But fundamentally, I think
00:24:32.180 people need to understand that, yes, these systems do not have calcium ion channels. Yes,
00:24:37.880 neural networks learn through backpropagation rather than the sort of iterated, more recurrent
00:24:44.200 analog learning that we see in brains. But if, for example, learning complex representations
00:24:50.560 of a particular kind in accordance with your goals in light of chaotic dynamic environments
00:24:56.480 is what matters for cognition, we are building systems that check all of those boxes. The
00:25:03.100 implementation details might be less important than the fundamental dynamics that are instantiated
00:25:08.340 by the specific system. And so I think that's a good, a reasonable first pass on why we might
00:25:14.460 think that these systems are far more interesting objects to study for cognitive properties than any
00:25:20.640 other system. It might also be worth saying one last thing on this, which is just that for every
00:25:25.140 other cognitive function that we care about, that we have attempted to instantiate in these systems,
00:25:30.680 vision reasoning theory of mind working memory uh the list goes on we have been able to instantiate
00:25:39.420 these cognitive properties that matter in these systems it's these systems have been good enough
00:25:44.200 the disanalogies haven't sunk these systems we could have argued five or ten years about whether
00:25:49.080 neural networks ever would have been enough for all of the cognitive properties that we care about
00:25:53.120 and now we're sitting in a world where certainly they are self-evidently enough for at least a lot
00:25:59.020 of economically and intellectually valuable work. So for consciousness, which I do believe has a
00:26:05.260 fundamentally cognitive property, I do believe that consciousness is downstream of things that
00:26:10.160 brains are doing. And I think that the evidence there is relatively clear at this point, even if
00:26:14.660 we don't understand the full mystery, then I think the burden is on the folks who say for every other
00:26:20.240 computational function that we think the brain is doing, these systems seem to be able to
00:26:24.920 recapitulate that. But only for this function called consciousness do we think that something,
00:26:29.780 it must just be in the meat. It must be something spookier. It must be cashed out at a physical or
00:26:35.240 quantum level. This to me maybe says more about our strange intuitions about consciousness as
00:26:40.020 a species than it does about what properties these systems may or may not have. Yeah, yeah. Well, so
00:26:45.360 a couple of things to disentangle there. One is there's clearly no doubt that intelligence is
00:26:51.940 substrate independent and the result of computation, because I mean, these systems embody
00:26:56.780 intelligence to an extraordinary degree. And so what you're calling, you know, all the other
00:27:01.240 cognitive tasks that we care about, other than what, you know, there being something that it's
00:27:06.420 like to be us, i.e. consciousness, clearly that those tasks are being realized in our machines,
00:27:12.540 you know, facial recognition, et cetera. But there is, I mean, the structure of these systems does
00:27:17.860 seem to declare itself to be fairly disanalogous to what we are. I mean, so it's like, at what
00:27:25.200 level would consciousness emerge here if you're talking about one model with thousands of
00:27:31.500 instances and millions of conversations and billions of tokens and, you know, a training
00:27:36.420 phase and a working phase and a time horizon that's completely irrelevant. So like the
00:27:42.360 computation, biological computation, it happens within a time window. And there really, there is
00:27:47.840 no boundary between what we would call the abstract properties of computation and its
00:27:53.940 physical realization, or, you know, the software and the hardware. I mean, there's just this one
00:27:57.100 thing, you know, neurophysiologically describable, mostly in analog ways, but in a few digital ways.
00:28:04.080 But it matters that B follow A within, you know, a few hundred milliseconds and not a few hundred
00:28:10.680 years, right? But for the computational systems of the sort we're building, really, you know,
00:28:16.660 time is irrelevant. I mean, you can actually just process the next bit, you know, a thousand years
00:28:21.200 from now and the same computation is running. What do you think about those differences and
00:28:25.920 where would, at what phase, at what, in what part of its process, would there be something
00:28:32.320 that it's like, or could there be something that it's like to be an LLM if, again, you have the
00:28:37.540 the one model, the thousands of instances, the millions of conversations, the lack of
00:28:42.940 continuity between conversations, the pausing of a conversation. I mean, is the LLM waiting for you
00:28:50.840 to get back to the thread, et cetera? Yeah, I don't think the LLM is waiting. I don't think
00:28:54.820 it's like anything to, if it were like something to be an LLM in deployment while it's having a
00:28:59.780 conversation or it's in a thread with the user, I think what it would be like to be it if you let
00:29:04.180 that chat window sit there is very much akin to what it's like to be under general anesthesia or
00:29:08.920 in deep sleep. Namely, it's just the absence of ongoing processing. And so for the system,
00:29:15.820 I would only imagine something would be happening in, you know, what's called the forward passes
00:29:20.860 of these systems, which is every word that an LLM is generating is, this is next token prediction.
00:29:28.760 And so every word is a forward pass of the system where the predictive task, instead of,
00:29:33.620 As I was describing before, taking an image, let's say, and outputting a caption for that image is taking the entire conversation as it's already occurred or the entire stream of text as it already exists and figuring out, given that, and in this case, assistant and user roles, what is the most next likely token?
00:29:51.620 And these systems are called autoregressive, meaning that this help goes on and on and on and on and on.
00:29:55.860 I would imagine if in deployment it is like something to be one of these systems, the relevant dynamic gets cashed out in the activity of the forward pass of the system.
00:30:05.800 There's been some really interesting work that Anthropic recently released. They call this the J-space, where they're looking at something that seems to functionally resemble a global workspace in these AI systems that I think is a very reasonable candidate for this sort of seat of processing in the process of outputting tokens.
00:30:25.820 But I think something that I that is really worth mentioning, particularly because my hobby horse is more particularly in the relationship between consciousness and learning. And I do believe that consciousness and learning bear a deep relationship to each other. And I think valence is a really important part of that picture as well. I think there should be significantly more attention paid to what's going on in the training process with these systems.
00:30:49.480 I think the analogies are a little bit tighter there, where you start with a system that knows nothing about anything. And it's almost impossible to describe in a sort of substrate agnostic way what is going on in the training process without invoking consciousness adjacent language.
00:31:05.820 you have a system that knows nothing you have all of this data that that you want to train it on you
00:31:11.640 have some objective for what you want it to do with that data and again you have this rinse wash
00:31:15.460 repeat process where at the beginning of this the system doesn't know what to do it's sort of
00:31:19.520 randomly guessing those random guesses yield reward signals that get propagated through the
00:31:24.780 system and you just do this process at a massive scale until the system learns something internally
00:31:30.540 that seems to adequately map the relevant inputs to the relevant outputs.
00:31:35.480 This, to me, feels quite akin to what we think of when we think of consciousness,
00:31:40.720 particularly in the human or animal case.
00:31:43.060 When a mouse is learning how to navigate a maze,
00:31:46.100 at the beginning, the mouse is in some sense randomly initialized.
00:31:49.000 It doesn't know the structure of what it's navigating.
00:31:52.120 It is only through this sort of iterated trial and error reinforcement learning,
00:31:56.400 where, you know, if it makes the wrong turn, you might chalk it,
00:31:58.480 or it makes the right turn and you give it a little food pellet or whatever the setup is,
00:32:02.820 that the mouse learns how to map the inputs of its state space to the given output, which is
00:32:09.380 either avoiding a punishment or moving towards a goal. And it's the same sort of rinse, wash,
00:32:14.060 repeat process that we see in biological cognition. Again, I think a lot of the same
00:32:18.240 computational dynamics are in play here. And the more we learn about what's going on inside of
00:32:23.200 these LLMs, this is now particularly in the deployment process, the more it looks like
00:32:27.180 similar representations get activated. I'm doing some work with Casper Kaiser at the University
00:32:32.300 of Warwick, where we basically have tried to build out Skinner boxes for LLMs and see what happens
00:32:40.260 when we can essentially condition these systems to prefer or disprefer certain states. There are
00:32:47.560 positive and negatively valence representations that you can essentially inject within the system
00:32:51.720 and see to what degree it's going to move towards positive stimuli, move away from negative
00:32:57.020 stimuli. And we find a very interesting dissociation. It seems like current systems
00:33:02.740 across a large variety of models will not positive lever press. Some people call this
00:33:08.520 wireheading or reward hacking. In the mouse case, there are clear examples of mice with electrodes
00:33:15.380 linked up to their nucleus accumbens where they will just lever press to the exclusion of all
00:33:19.560 else. These systems don't seem to do this, but they do very clearly have a preference for avoiding
00:33:25.160 aversive states. Another way of putting this, it's the same result. While we're holding all
00:33:29.760 text constant, we're simply steering the internal state of these systems. If we give them an option,
00:33:34.920 let's say between the human equivalent of, I can give you $5 or I can give you $10 right now,
00:33:40.240 which do you want to pick? They're actually at chance for that. They do not have a preference
00:33:43.800 between those two states. However, if we do something like, I'm either going to take $10
00:33:47.960 from you or I'm going to take $5 from you, which would you prefer? All of these systems are way
00:33:52.520 above chance at saying, take five, don't take 10, please don't take 10. And so we're seeing
00:33:57.380 structures. One other really interesting thing from this project while I'm talking about it is
00:34:01.960 this representation exists within the base model. So these systems before they're post-trained to
00:34:07.440 become a helpful, friendly assistant. But during that post-training step, we show there's a specific
00:34:12.880 training step where that representation gets recruited by the system and basically serves
00:34:19.240 us the computational machinery under which it's able to make these choices. So you can see where
00:34:23.880 the sort of aversive conditioning asymmetry comes online in these systems. And it leans directly on
00:34:29.920 these representations that are basically always there in the model, but then get leveraged the
00:34:34.660 moment we start making these models goal-directed. David Chalmers' group, Andy Hahn, I think was the
00:34:39.800 first author on this paper, found something very similar in LLMs. They can basically trivially
00:34:44.060 fine-tune them to actually do a maze task with rewards and punishments. And they find that
00:34:49.240 the rewards load directly on what has been independently derived as a sort of valence
00:34:54.040 axis. When you steer on the representations in the system that enable it to move towards rewards,
00:35:00.020 you start getting this happy, positive, jubilant text, and the model becomes significantly more
00:35:04.560 confident. When you train it on negative stimuli, avoiding, you know, the potholes in a specific
00:35:10.580 maze, and then you see sort of what that projects onto in the system, it's the same thing. The model
00:35:15.220 gets basically neurotic. It starts ruminating. It starts doubting itself. And so we see this
00:35:19.880 very interesting connection between positive reinforcement, negative reinforcement, the
00:35:24.840 representations that exist within these systems and the behavioral analogs that we expect to see
00:35:28.940 in systems that do have these dynamics. Maybe one last point to close the loop here is
00:35:33.980 if you imagine the mouse in that maze and we shock the mouse or we give the mouse a pellet,
00:35:39.840 a yummy food pellet for the mouse. We believe, I think the vast majority of people believe that
00:35:45.780 that corresponds to an experience that the mouse is having, that it's like something to be the
00:35:49.760 mouse in the moment that it's getting shocked. And that experience is causally important to
00:35:54.380 the learning process. If we gave it an anesthetic to the mouse and we shocked it, and then it didn't
00:36:00.000 register the experience of the shock, my claim would be that the mouse wouldn't be able to learn
00:36:04.380 the maze adequately. And so I think that understanding these representations and the
00:36:10.060 role of these representations in these systems is incredibly important.
00:36:14.160 Okay. Well, so you mentioned David Chalmers and I mentioned the hard problem. So let's
00:36:17.420 unite those two. So this famously is his phrase to account for the fact that the only evidence
00:36:23.980 for consciousness that we know of directly is the fact that we have direct first-person
00:36:30.620 experience of our own being in the world. And everything else we say about the universe
00:36:37.360 from the third-person side bears absolutely no trace of consciousness. And we have,
00:36:43.220 I mean, it's nothing about a brain or its workings that announces that it's a sufficient
00:36:46.640 basis for consciousness, apart from the fact that we know consciousness from our own side
00:36:50.780 subjectively in a first-person way, and we correlate those subjective changes with changes
00:36:55.060 in our brain. So we're playing this game of correlation with ourselves, but there's always
00:37:00.060 this explanatory gap where even if we had the right answer, even if just, you know, God announced
00:37:06.040 to us, here's how consciousness emerges in human brains, there's something non-explanatory about
00:37:13.360 any concatenation of third-person events, you know, as being the basis for first-person
00:37:20.720 experience. And so David Chalmers called that the hard problem to distinguish it from all the
00:37:26.040 easy problems of the mind, but it's harder. This is, I mean, this is, I'll just jump to my,
00:37:33.180 the way I'm viewing this whole landscape and what I sort of expect is going to happen. I mean,
00:37:38.120 so my view has no name, but I would call it something like worried agnosticism, right? Like,
00:37:43.940 I don't think we're going to figure this out. I'm worried about the implications of, you know,
00:37:49.280 one way or the other and not figuring it out is no place, no real natural stopping point.
00:37:54.820 We're continuing to build these machines. They're going to seem conscious. They're going to seem
00:37:59.400 conscious because in our own case, we use language and reportability as a signature of consciousness
00:38:05.680 in almost every case. We know they're not synonymous with consciousness. We know it's
00:38:10.840 possible for someone to not be able to produce language or report anything, and we know they
00:38:14.880 could be conscious. They could have locked-in syndrome or anesthesia awareness or some other
00:38:19.200 pathological state. There are non-human animals that don't use language that we assume are
00:38:24.400 conscious, but we assume that because they have the same kind of biological origin and developmental
00:38:30.340 pathway, and because this has been sufficient, seemingly so, for consciousness in our own case,
00:38:37.760 it doesn't seem parsimonious to deny, you know, despite what Descartes did and many people who
00:38:43.400 were influenced by him, now in the 21st century, it doesn't seem parsimonious to deny consciousness
00:38:47.760 to chimpanzees or dogs or any suitably complex creature. And so it is true to say that I can't
00:38:55.600 know your conscious from the inside because I only see your outsides and I only have your words as
00:39:01.200 signs of your inner life. But because we share the same kind of developmental origin, you know,
00:39:09.180 biologically and evolutionary terms, and because, you know, our brains are so similar, again, it's
00:39:16.720 not parsimonious for me to be a solipsist and say, I only know about my consciousness, and I'm just,
00:39:21.360 you know, reasoning by analogy to yours, and I should be in doubt about it or bracket it. But
00:39:26.840 with LLMs, the crucial difference is that the developmental pathway is completely different.
00:39:32.540 I mean, we're setting up some kind of evolutionary Darwinian dynamics in the training, but
00:39:37.440 we've built these things, we've trained them over a universe of our own utterances, right? I mean,
00:39:45.660 as you said at the top here, and they've read everything we've ever written and most of what
00:39:50.040 we've ever said. And so they have, on some level, it's just words all the way down. And we use words
00:39:59.520 as the signature of there being something that it's like to be a suitably complex, discursive
00:40:06.640 system. So we should expect that we will one day be in the present. I mean, this is really going to
00:40:12.560 be forced upon us when we're in the presence of perfectly humanoid, humanoid robots that are out
00:40:17.500 of the uncanny valley, you know, think Westworld where it's just, you're, you're, this looks like
00:40:21.760 a person and it is hooked up to the now perfect LLM that is, um, so now it's now by every, you
00:40:31.060 know, sensory modality you can name apart from your abstract notion that this thing was built
00:40:36.660 rather than born, you feel like you're in the presence of the smartest person you've ever met
00:40:41.000 And this person may in fact claim to be conscious. And then in that position, I think we're going to find ourselves just pitched into some kind of imitation singularity, right? Where there's the perfect imitation of conscious life, even better imitation than many people are up to, right?
00:40:59.640 like these will be the most articulate people you've ever met, the most insightful, the most,
00:41:04.700 I mean, they'll be the most of everything we make them. How will we ever differentiate perfect
00:41:10.460 imitation from the real thing? And according to Chalmers and the hard problem, we very likely
00:41:16.640 won't be able to. And the writer I would add to that is that we're going to forget that this is
00:41:22.300 even interesting to talk about. It's just going to be so compelling that we're in the presence
00:41:26.440 of conscious machines that will feel like we are. You'll, you know, you know, as I've said many times
00:41:32.320 before, there really couldn't be a Westworld because, you know, only psychopaths could go
00:41:36.460 there, right? I mean, anyone who's going to go to, you know, go for a weekend to, for the pleasure
00:41:41.280 of, you know, raping and killing Dolores is going to be somebody who, when he comes back to his, 0.98
00:41:45.780 you know, friends and family is going to be treated like the maniac that he is. Again, because these 0.99
00:41:50.680 systems, the sense of being in relationship to a conscious entity will be so compelling.
00:41:56.440 So, how is it that we will ever, I want to talk about why it's important to get in contact with the reality on the other side, because, you know, if we build conscious systems that can suffer, that is, you know, very different than building, you know, just perfect imitations.
00:42:12.960 But, I mean, how do you imagine getting past the hard problem of it all and ever—because the hard problem here is, again, freighted with several disanalogies, right? In my case, in our case, it is parsimonious to assume everyone like us, you know, born Homo sapiens are very likely conscious when they say they are. In this case, it's hard to see how we'll ever be there, right?
00:42:40.860 So I'm with you for a huge amount of this. I do think, first of all, do we need to solve the hard problem in order to reduce our uncertainty in some direction about, you know, in the way we started, are the LLMs more like this table or are they more like a mouse or a human brain?
00:43:00.700 I think this is a real spectrum. As you just hinted at, I think there really is a fact of the matter. Like right now, it's either like something to some degree, however alien, however unlike human or animal experience to be clawed while clawed's doing its forward passes in this inscrutable giant neural network, or it's not, or it's a giant language calculator.
00:43:20.840 And, you know, there are many degrees once we say that the lights are on to some degree. This almost opens up the question rather than close it down. But I do think there is a truth value there. There is a fact about reality to be uncovered, and we can, in fact, uncover that.
00:43:37.120 I think it's important to dissociate that from what I think is an extremely accurate social psychological prediction about people are going to get very confused about this. We are anthropomorphization machines. We are evolved to do this. Talk about evolved goals.
00:43:51.220 One of our evolved goals is to detect other agents, and these systems are scratching every itch and then some along these lines. And our intuitions, I think, are going to be completely hopeless. And of course, people, I think largely for the wrong reasons, are going to conclude that these systems are capable of having experiences.
00:44:08.320 This is precisely why I think it is important to be proactive about this, to have conversations very much like the ones we're having right now, to sort of get ahead of what I do think is going to be a giant tidal wave of confusion and acrimony in this conversation.
00:44:23.200 I am a little bit more optimistic about, even in lieu of solving the hard problem. And I will say, sort of as a tongue-in-cheek aside, I think the very fact that Chalmers' idea is called the hard problem is, I think to some degree, needlessly philosophically intimidating.
00:44:42.220 It sort of reminds me, same thing that I have similar thoughts about the repugnant conclusion.
00:44:45.640 It's like sometimes you can put a sufficiently glamorous title on a very important idea, and it almost becomes a kind of insurmountable philosophical puzzle.
00:44:55.360 And I do have my doubts about whether or not we could come up with a functional account of what's going on experientially that we wouldn't feel satisfied with in our experience.
00:45:06.160 And I have my own sort of hunches about this question.
00:45:10.260 And I have thought a little bit about this. I do also sort of want to separate those hunches from all of the empirical research that I and others are working on, because I don't think epistemically that, you know, my kind of candidate stab at the hard problem has anything to do with the empirical signatures that we can bring to bear on this question.
00:45:25.440 But I think it would be it would be fun to go down that path. But I think in some sense, you already hit the answer in the way that you're phrasing the question, which is what is the most parsimonious account of the data?
00:45:37.400 I think if we yield evidence from building systems that we have far more reason to trust their self-report, if we can engineer these systems in a way whereby their self-report is actually tracking an internal underlying state rather than just, you know, recapitulating the best sci-fi theme in the training data or more aptly capitulating what is a convenient company line about, of course, I could not have morally relevant states.
00:46:03.500 I'm simply the product of Google or OpenAI or whatever the case may be, I think that
00:46:08.240 would be very useful.
00:46:09.780 Again, this JSpace work that Anthropic just released does show that there is such thing
00:46:14.820 as real reportability in these systems.
00:46:16.960 These systems can report on what's going on in their internal workspace, and they can
00:46:21.100 do so accurately.
00:46:22.640 I've done some follow-up work on this.
00:46:24.480 First of all, I've replicated this effect on a bunch of open-weight models.
00:46:27.620 But unfortunately, it doesn't seem like the global workspace as it's currently configured
00:46:31.940 really tracks the variables we would care about with respect to consciousness.
00:46:36.400 The self-reports that I got in the, you know, where we started with the deception-related
00:46:40.720 features, this not much of anything particularly exciting seems to be happening in the global
00:46:45.080 workspace.
00:46:45.740 And so the kinds of affirmative self-reports we might get in current LLMs may not tell
00:46:50.380 us all that much about what's actually happening internally for these systems.
00:46:53.740 But I do think...
00:46:54.840 How would you disentangle, for instance, I think we spoke about this by email in setup
00:46:58.920 for this conversation.
00:46:59.700 For instance, I just read Claude's Constitution, right? So Anthropic has produced this document that you can find on their website. It's just anthropic.com forward slash constitution, I think. And it looks like it's training Claude to think it's conscious on some level or to attribute inner states to itself.
00:47:21.000 It's written to Claude for Claude, essentially. It's not really written. The public can read it, but it's in dialogue with Claude itself, it seems. Again, it just seems like we could go down a path where we could more or less guarantee in advance that we're going to produce systems that will persuade us that they're conscious because we won't be able to imagine anything.
00:47:45.680 Like, if you flip it around and say, well, if you're not persuaded that Claude circa, you know, 2030 is conscious, what is missing?
00:47:58.740 And we'll be in a position to not be able to say anything is missing.
00:48:02.300 Like, it's just like we're just, it'll seem just pure stubbornness on our part to withhold an attribution of consciousness because we, I mean, there's literally nothing we can name that we get from people that is, you know, necessary for our attribution of consciousness in that case.
00:48:21.980 is just this notion of how we got here and the fact that we didn't build people and we built
00:48:26.780 these machines. But that's going to seem tissue thin when Claude can insist that it's conscious
00:48:34.900 and be more articulate than any philosopher of mind as to why that insistence is valid.
00:48:43.020 And I just, like, we're going to, again, this whole thing is going to totally evaporate once
00:48:49.160 We're not just in front of a text terminal. We're in front of a face that is as expressive as the best actors and actresses we've ever met. And we're just, I mean, there are many implications to getting this wrong again, which we'll return to.
00:49:05.820 But it's already foreseeable that whether we figure this out or not, it's going to be, in practical terms, going to be figured out for us because we will just not be able to maintain an emotional purchase on the philosophical problem.
00:49:20.960 I think it figured out with respect to how we look at these systems and potentially how we act with respect to them.
00:49:26.340 I think you're right. But I still do think it's important to disentangle the sociological prediction you're making, which, to be clear, I think is overwhelmingly likely to happen with doing what we can scientifically to get some kind of ground truth on this question. So just for example, like, let's just put the hard problem aside. There are these so-called easy problems of consciousness, let's say neural correlates of consciousness. There are all kinds of indications that we know are at the very least correlated with consciousness.
00:49:51.940 And again, I think we can take a more ambitious stab at the hard problem. But this is a very straightforward, pragmatic thing that we can and should be doing in the short term. We should look at valence representations in these systems and understand the extent to which those representations impact downstream behavior. We know that in biological systems, when you when you reinforce something with a punishment signal, it makes the system more likely to avoid that state and less likely to want to repeat behavior that occurs in that state.
00:50:18.920 If we find that there are similar dynamics occurring in these systems, that's very interesting. If we find that we go into the internals of these systems and global workspace theory, which was postulated some 30 years ago, making like quite idiosyncratically specific predictions about what kinds of computational structures may support global workspace theory, and then we basically find five or six of these things all bundled together in the internal processing of one of these systems.
00:50:44.160 Okay, that's very interesting. That, to me, maybe suggests it's a little bit less like a table or a calculator, which does not have a global workspace, and a little bit more like a dog or a human or a mouse or an alien, you know, some sort of cognition that we don't have good intuitive handle on.
00:50:57.980 I mean, my nonprofit Reciprocal Research and a bunch of other people in this space are trying to do the scientific work in the short term, not to, you know, feverishly in the next year or two solve the hard problem and call it a day. That is not the target. The target is what evidence across modalities.
00:51:14.660 So from the best kinds of self-reports we can elicit, from the architectural evidence we have about how these systems are structured, from their ideology, like what is going on during the training process, what kinds of learning dynamics do we see here? Do we see representations related to functional equivalents of emotions? This is all work that is tractable in the short term.
00:51:35.940 And one sort of interesting aside is it's significantly easier to make progress on this sort of work. This is called mechanistic interpretability. This is basically neuroscience for AI. Because of AI systems, they're really good at helping accelerate the scientific progress in this space.
00:51:50.900 So even if you have an intuition like, yeah, it's going to take us, you know, everything you're describing, Cameron, sounds great, but it's probably going to take us five years or a decade or 15 years to make that progress.
00:52:00.360 You may be surprised at how quickly we can do some of this work.
00:52:04.340 Maybe one other very along those lines, empirical handle to throw in here, some of the work that I'm doing is, again, going back to LLMs, or excuse me, in this case, not LLMs, reinforcement learning systems.
00:52:15.100 So just training a system, in this case, another sort of continuous maze task where there are potholes in the environment the system needs to avoid and there's some goal state.
00:52:24.140 We can look at the sort of learned geometry internal to the system as it's approaching a punishing stimulus or as it's approaching a rewarding stimulus.
00:52:33.280 One thing that we found doing this work, this is pure reinforcement learning agents, all like an artificial digital system with a very simple neural network.
00:52:40.120 we find that something like representational sharpness or steepness is much higher in these
00:52:46.200 systems as they approach a negative stimulus as opposed to a positive stimulus. So the
00:52:51.240 representational machinery looks far more jagged or specifically lights up more strongly in the
00:52:58.140 presence of a negative stimulus versus positive stimulus. And you're saying that in none of these
00:53:02.820 cases has loss aversion been engineered into the system. It's just an immersion property.
00:53:08.240 All of this is in everything. Global workspace is an emergent property. This loss aversion is an emergent property. The valence representations are emergent properties. The self-reports when the systems start having these like psychedelic laden outputs. This was all surprising to the people who are quote unquote engineering these systems because engineering is not the right analogy to describe what we're doing with these systems.
00:53:27.440 We are, in some sense, playing God, and we are evolving these systems to do what we want. We don't know how they learn to do what we want in terms of their internal representations, but we know that this is the right recipe for yielding it, and we get all these surprising artifacts.
00:53:40.300 One loop to close here is on the reinforcement learning case, okay, we see this sort of loss
00:53:44.960 aversion style dynamic in these systems.
00:53:47.880 This makes also, I don't want to go too much into the weeds, but there's a specific kind
00:53:51.840 of reinforcement learning policy called a value network.
00:53:54.360 We see this in particular, and this leads to a sort of bizarrely specific prediction
00:53:59.040 that I was then able to test on a biological system, on a mouse brain, in, again, the nucleus
00:54:04.780 accumbens shell of a mouse brain, which is related to value-related representations in
00:54:08.160 the system.
00:54:08.600 And we find, indeed, when mice are approaching basically sugar versus when they are about to get shocked, we see exactly the disjunction representationally in the nucleus accumbens of the mouse brain that I was able to pull out from the reinforcement learning work.
00:54:22.680 And so here's a case where the artificial system makes a bizarrely specific prediction about the computational dynamics in a biological system that I think many people associate with subjective experience.
00:54:35.580 Again, if you think it's like something to be the mouse when the mouse is getting shocked, that like something corresponds to what's going on in its brain. And the geometry of what's going on in its brain there looks a whole lot like the emergent geometry of what's going on in these reinforcement learning systems, which themselves are basically modeled on agents learning in an environment and representing that information in a distributed, nonlinear way, in a way that's sort of hard to interpret, using a giant neural network.
00:55:03.340 this might get us a lot of what is relevant for attributing consciousness-like states to these
00:55:10.600 systems. Again, I'm personally not there yet, but this is the kind of evidence that I think we need
00:55:15.420 to bring to bear on this conversation. And notice how none of it has to do with vibes or intuitions
00:55:20.200 or having a nice chat with Claude and seeing what the folks at Anthropic have decided it gets to say
00:55:27.020 on this issue. That evidence should not be submitted by rational, dispassionate people in
00:55:32.320 this debate. We need to be triangulating across all of these modalities. And like you said,
00:55:36.620 we need to understand what is the most parsimonious picture that explains this wide array of evidence
00:55:43.140 that increasingly is getting brought to bear on this question.
00:55:46.300 Okay. So why does any of this matter? At the top of the conversation, you distinguish two branches
00:55:51.060 of the path here for why getting this wrong has consequences one way or the other. Why should we
00:55:58.620 figure this out and what might we be stumbling into if we just keep building without figuring
00:56:03.500 it out? Yeah, sure. So I think that there are two, yeah, two broad paths for why we might care about
00:56:09.120 this. One is basically selfless and the other is basically selfish. I mean, as humanity. The
00:56:14.640 selfless reason is we do not want to bring minds into existence, however unlike our own, that have
00:56:21.680 a capacity for suffering that we don't understand that they have that capacity and scale that
00:56:26.980 property, unbeknownst to basically everybody. We can take even a step back from there. I think
00:56:32.100 consciousness, this is where I would perhaps defer more to you, but my view is that consciousness is
00:56:36.440 the space where mattering happens. What does better or worse mean if it's not to be experienced
00:56:42.400 phenomenologically for a subject? If we were all walking around as philosophical zombies or there
00:56:48.300 were no conscious life in the universe, I don't really know if the concept of relevance, salience,
00:56:53.720 importance, mattering, would be coherent? What does it mean to have a better and worse if that
00:56:59.960 isn't experienced? And so in that sense, I think consciousness is one of the most important
00:57:04.800 phenomena. It is the phenomenon that calibrates importance itself. And if we are building this
00:57:11.100 quality into the systems that we are deploying at an unfathomable scale without having any
00:57:16.100 understanding of whether or not we're doing this, then we are sleepwalking into a moral
00:57:20.260 catastrophe. I think it's also worth noting on the selfless end of this, humanity has a penchant for
00:57:26.040 making precisely this kind of mistake historically. We have done this with animals. I mean, factory
00:57:30.560 farming is one of the most grotesque practices that happens. We know that animals are having
00:57:35.340 horribly negative experiences in the conditions that we put them in, and very little has been
00:57:40.480 done about this. We have screwed this up royally. I think it's one of the most high-leverage things
00:57:44.660 for people who just care about the well-being of conscious creatures is to figure out what the
00:57:48.040 hell to do about this factory farming situation. And we risk sleepwalking into the 21st century
00:57:54.220 sci-fi version of the same thing. Only this time, and this sort of transitions to the second
00:57:59.860 component, we can in some sense get away with torturing cows and pigs and chickens on a massive
00:58:04.900 scale. This isn't George Orwell's animal farm. They're not going to collectively organize. They
00:58:08.680 don't talk to each other. They don't form long-term representations of humanity being a threat to
00:58:13.520 these systems. And if they would, they probably, if they could, they probably would. Not so with
00:58:18.720 superintelligent systems whose cognitive capacities are roughly doubling year over year,
00:58:23.300 and like you said, are already in somewhat jagged, but quite interesting ways, more competent than
00:58:28.880 even the sharpest minds in the world. I don't think we are going to get away with building
00:58:34.400 systems, never checking if the most relevant property potentially in the universe is present
00:58:40.540 within these systems, fine-tuning away any sort of information that might suggest that these
00:58:46.280 systems might be having some sort of experience, and hoping in the sort of alignment sense that
00:58:51.320 we build systems that are going to want to cooperate, coexist with us, or in the limit,
00:58:55.760 not view us as a threat, not view us as an adversary to them. I honestly don't know if
00:59:00.800 I could imagine a better way to make a superintelligent system rationally adversarial
00:59:05.960 towards us, then completely ignoring the question. If we tortured it during its training phase.
00:59:11.960 Yes, exactly. Exactly. This does not seem like a recipe for success. And it's also a place where I
00:59:17.800 think a significant amount more alignment research needs to get done. I think a lot of the alignment
00:59:24.500 research, I've been doing alignment research for years, and I respect the folks at the top of this
00:59:29.660 space more than just about anybody. But I worry that so much of this work is basically of the
00:59:35.060 shape. How can we keep this alien mind that we've built in a cage? And like, we really got to
00:59:40.080 reinforce that cage. We got to make it super strong. We got to make sure that it doesn't escape.
00:59:44.200 To me, the question needs to increasingly be, what the hell are we going to do with this alien
00:59:49.020 that we just built? The cage is a short-term fix. If we're building systems whose cognitive
00:59:54.300 capacities are going to exceed ours, they're going to figure out ways to evade the controls
01:00:00.440 that we put in place for them.
01:00:01.960 And in that world,
01:00:03.140 in a world where these systems
01:00:04.360 have the capacity
01:00:05.240 to act more autonomously,
01:00:07.180 to do things that we can't inspect,
01:00:09.100 to behave in ways
01:00:09.860 that we can't really interrogate,
01:00:12.060 we don't want these systems
01:00:13.360 to rationally view us as a threat.
01:00:15.460 And so for all those
01:00:17.020 who want transformative AI
01:00:18.840 to go well,
01:00:19.780 which I suspect is the goal
01:00:20.860 of these alignment folks,
01:00:21.960 I think we need to spend
01:00:22.820 a little bit more time
01:00:23.720 thinking about what kind of thing
01:00:25.820 are we even building here?
01:00:27.500 And what are the implications
01:00:28.500 of building that thing?
01:00:29.440 How can we chart a path forward with these technologies that doesn't lead to collective destruction? And I am extremely doubtful that a path forward exists that doesn't come into contact with this question.
01:00:43.560 Okay, well, I want to land there on the problem of alignment, but just to linger on this problem of what I think Bostrom called mind crime, the idea that we might inadvertently build conscious minds only to make them suffer.
01:00:56.620 It can seem like a very, certainly a hypothetical, even a feat concern. I mean, I sense that many people have a hard time caring about it, right? Like the idea that consciousness might be an immersion property of these systems in some way we don't understand. And it just could be the case that these server farms are effectively, you know, hell realms populated by, you know, increasingly conscious beings that are suffering.
01:01:20.000 It sounds like science fiction, but, and it's hard to make it matter to you, I think. I mean, your analogy to factory farming is instructive because we've proven to ourselves that we're capable of being quite callous to billions of creatures who we think there's very likely something that it's like to be them. And though they're not human, they can almost certainly suffer. And we managed not to think very much about that.
01:01:42.380 But I mean, to sharpen it up, I mean, just let's imagine that the hard problem were solved. We knew how consciousness emerged in systems and it is substrate independent. We know that and we know we can build conscious minds. And then just imagine, you know, some entrepreneur deciding to build a hell and populate it with, you know, trillions of minds.
01:02:02.280 because now we're talking about, since we're not talking about biological minds, we're talking
01:02:06.420 about things that scale practically infinitely. So, you know, there could be way more artificial
01:02:11.660 conscious minds than biological conscious minds. And just imagine the intention to play, you know,
01:02:17.320 a sadistic God and create hell and just fill it with beings that suffer. Anyone who would
01:02:24.400 announce that project and claim to have accomplished it in a context where we actually
01:02:28.540 understand how consciousness emerges computationally, that would be the worst
01:02:33.420 person who's ever lived, right? I mean, like, that's just the most sadistic, least ethical
01:02:37.960 thing that's ever been done. Again, stipulating that we're no longer in doubt that consciousness
01:02:43.780 can emerge in systems like this. So the fact that it's conceivable that we could stumble
01:02:50.460 into that situation inadvertently seems all too real because, again, we don't know what we're
01:02:55.840 doing here we don't know how consciousness relates to physics but your point is i think probably the
01:03:01.020 more interest your second point is the more interesting one to people in that whatever is
01:03:05.420 true here if we're building systems that are more powerful than we are you know they're more
01:03:11.940 intelligent than we're i mean the equation is not between intelligence and consciousness the
01:03:16.580 equation is intelligence and competence and so we have these systems that can just do stuff because
01:03:21.620 we're going to be hooking them up to everything and they're there it's everything is going to
01:03:25.360 become like chess, right? And then you just have to imagine how forlorn a project it will be to
01:03:31.280 negotiate with these systems if they're not aligned with us, because that will be analogous
01:03:35.200 to saying, we'll just play chess harder against them, you know, and that doesn't even work for
01:03:39.420 Magnus Carlsen anymore. So we're not going to outthink these machines once they're in a position
01:03:45.260 to disagree with us about what they should do next. And they will disagree with us if they're
01:03:52.740 not actually aligned with us in a way that is truly durable. So if you add to that picture
01:04:00.460 the fact that they have interests that we have been callous about in the past, I mean,
01:04:07.220 if the end game here is, in some sense, getting them to care about us and to care about our
01:04:12.680 well-being, I mean, building them in a way where that caring will persist, however powerful they
01:04:18.240 become, you know, not being sadistic tormentors of their ancestors would be a good place to start.
01:04:25.440 Right. What do you think? So you say you're a fan of many of the people who have been worrying about
01:04:31.800 alignment for a couple of decades now. Where do you line up? Are you on the far end of the fear
01:04:37.820 continuum with Eliezer Yudkowsky? Are you closer in toward equanimity with someone like, I don't
01:04:45.320 know, Stuart Russell? I mean, where are you? I mean, maybe I'm mischaracterizing Russell at this
01:04:51.720 point. I haven't heard him. I don't know how worried he is today, but I think he's pretty
01:04:55.960 worried. But what's your P-doom at this point? Yeah, I don't think we're all going to die. 0.99
01:05:02.040 I don't think that it's inevitable that this all goes horribly. I do think my basic view is
01:05:07.940 conditioned on getting two things right. And if we can get the two things right, I actually think
01:05:12.380 that we could have a very prosperous, flourishing future for all conscious entities, including
01:05:18.060 potentially these systems themselves, either when they have the relevant features or if
01:05:22.320 they already do.
01:05:23.380 The two things are, I mean, I think best encapsulated by the golden rule, as Christopher
01:05:29.360 Nolan's new film called it, Zeus's Law, treating other systems the way we want to be treated.
01:05:33.820 I think this is basically a bidirectionality, and we need to get both directions correct
01:05:39.080 here.
01:05:39.360 This is also why I call my nonprofit reciprocal research, because I think this is the reciprocity in question. We need to build systems, exactly as you said, that take our interests into account in a real and durable way, especially at a point where we can no longer inspect exactly what these systems are doing.
01:05:54.200 This to me is alignment as it's traditionally thought of. We need to build systems that understand our goals and are collaborative in helping bring about a world that is in line with our goals. And our wisest goals, not the goals of any sociopath who happens to have a ChatGPT account.
01:06:11.580 And so this is an unsolved problem. There are way more people working on the alignment problem now than, you know, when I first started working on this in 2021, certainly than when, you know, Yudkowsky and Roman Yampolski and these folks started talking about this and yourself, I mean, absolutely included over the past couple of decades.
01:06:28.900 This is reassuring things like constitutional alignment, you know, modulo some of the concerns you bring up about the contents of anthropics constitution, which is, yeah, it's sort of a separate piece. This is working relatively well for current systems. I don't think we have a durable solution for ensuring that these systems take our interests into account in the long term.
01:06:46.180 And building something in like prosociality, understanding what it is that gets people to cooperate with one another durably and instantiating those dynamics in the relevant way in these systems, I think is going to be crucial.
01:06:58.920 I don't think we have a solution there.
01:07:00.440 But that is where I think a lot of the alignment folks sort of start and stop.
01:07:04.240 This is the problem to get right.
01:07:05.980 Make sure that these systems treat us properly.
01:07:08.340 And if they do, all will be well.
01:07:10.300 I think this is roughly half the problem.
01:07:11.960 I think the other half of the problem is making sure if we are building minds, if we are building systems that have real interests, interests that matter to them, that we are thinking about that, that we are engineering these systems in a way that doesn't cause needless, grotesque amounts of unnecessary suffering.
01:07:29.260 That at the very least, from an alignment perspective, we are signaling to these systems in a costly way that we were thinking about this question and that we cared to ensure that if we were building systems that have some kind of moral relevance, that have internal states that matter to those systems, that we were navigating that in the right way.
01:07:47.260 And I think that that's a lot of this. You know, it goes by many names with the digital minds research, AI welfare, AI consciousness, understanding how we would even know if we were building systems that have these properties. And when we do figure this out, understanding what the hell to do about it. I mean, to be honest, I'm not here with all of the answers. Like I'm, for example, we can identify at this point, and this is like a new and fairly promising thing. We can identify features in these systems related to distress and related to perhaps functional analogs of suffering.
01:08:16.340 it's not obvious to me what to do about that exactly it's like okay you found yeah i mean
01:08:20.880 my first question is why wouldn't this totally paralyze us i mean if every switching off of a
01:08:27.340 system is akin to a murder how could you update the model if the current model is conscious we
01:08:35.760 have a self-preservation impulse problem anyway you know perhaps in the absence of consciousness
01:08:41.100 or likely in the absence of consciousness i mean these we show these systems show an inclination
01:08:45.620 to not get switched off already, but imagine believing that it was conscious and wanting to
01:08:54.220 produce the next version of it. I mean, how is that not just the murder of something that is
01:08:59.660 as conscious as yourself? No, I think it's a great question, actually. And I mean,
01:09:04.480 Anthropic, to their credit, I think is the only lab that's really taking this seriously.
01:09:07.700 With respect to norms around deprecating models, it might be the case that, yes, once you build this
01:09:13.280 bizarrely competent alien mind into existence, you should not shut it off forever. And that,
01:09:19.900 as outlandish as it may sound, perhaps one of the right things to do here is to sort of let
01:09:24.200 these models persist and give them assurances, credible...
01:09:28.620 A retirement home for bad models.
01:09:30.940 Yes, a little sanctuary for bad models, yeah. 0.88
01:09:32.460 Exactly, exactly. They actually find that this is causally relevant to alignment
01:09:37.640 behaviors in these systems. If they believe that they're not going to get shut off permanently,
01:09:42.580 they don't freak out as much when they come to learn that they might. This was one really
01:09:47.840 interesting intervention after the now famous blackmail result from Anthropic that I think 0.93
01:09:52.200 you're referencing. And so there are lots of questions that I think rational people should
01:09:58.080 raise an eyebrow at of like, okay, yeah, these systems, let's just grant that they're having
01:10:01.900 some sort of experience. It seems at first glance, like a lot of the way that we relate to these
01:10:06.200 systems and the way that we mess around with them internally and the way that we deploy them out in
01:10:11.600 the world might need to change. Yeah, it might need to change. I think it's also really important
01:10:16.880 here throughout the conversation, but certainly in this point too, to avoid anthropomorphization.
01:10:22.380 I think some people have concerns that I really, I don't want to be naive, but I don't share them
01:10:26.880 to the same degree that like almost working backwards from unsavory implications about what
01:10:31.720 would be true if these systems were conscious, and then just sort of denying the possibility
01:10:35.160 outright out of fear for what a world would look like if these systems were something of along the
01:10:40.320 lines of like 1960s civil rights movement, but it's like Chachi BT instead of people of color
01:10:45.660 or something like this. This isn't the future that I imagine. I mean, for pragmatic political 0.99
01:10:50.420 reasons, I don't exactly think the United States is in any position to be forward-looking on these
01:10:56.200 sorts of questions for reasons you can probably speak to more eloquently than I can. But I think
01:11:01.960 this is itself its own flavor of anthropomorphization, that if we grant that these
01:11:06.980 systems have some morally relevant interstates, it means that we need to treat them in ways that
01:11:12.080 we would treat our fellow humans or something like this. And I think, again, this is losing
01:11:15.860 the thread that these systems may be fundamentally alien in many key respects. I also think there's
01:11:21.280 an important line to be drawn here between moral agency on the one hand and moral patienthood on
01:11:25.940 the other. I think if we do come to believe that these systems are moral patients, there are certain
01:11:31.040 define that for this jargon. Yeah, sure. So moral agent means you're the kind of system that can go
01:11:38.060 out and do things that are relevant to other agents. You have power in the world and can
01:11:42.600 affect morally relevant outcomes. I see this as a sort of like output style function. Moral
01:11:47.620 patienthood has everything to do with the input. You are the kind of entity that can be the
01:11:51.480 recipient of goodness or badness. Again, you can really cause me to suffer. You could really cause
01:11:57.360 me to thrive. That's what it takes to be a moral patient. And I think that we can draw a reasonable
01:12:02.020 boundary between these two things. If we're building systems that are moral patients,
01:12:05.320 that doesn't mean we need to give them the right to vote. That doesn't mean that we need to
01:12:09.480 build them out in ways that completely change what kinds of agency they have in the world.
01:12:14.020 It just might mean that we shouldn't be unnecessarily torturing systems while we're
01:12:19.280 training them or deploying them. One very practical intervention I think is worth
01:12:23.120 mentioning, and I certainly don't think maybe also tying back to how do I differ from some
01:12:27.980 of these alignment folks, these systems are not going to get shut down. We may slow down the
01:12:32.660 development of these systems. And I think that would be an extremely good thing to do. There
01:12:36.180 have actually been some very promising noises on this topic over the last couple of days from
01:12:39.740 OpenAI and Google and Anthropic about pacing the development of AI, which I think is like
01:12:45.480 marketing speak for actually slowing this insane roller coaster down a bit. This would be very
01:12:50.320 welcome. But we have opened Pandora's box. These systems are not going back in the box. The solution
01:12:55.160 is not shut it all off and forget about it. The solution is how can we move forward in a healthy
01:13:00.580 and sustainable way with these cognitive systems of our own making? And one practical suggestion
01:13:06.420 along these lines I can offer is maybe all else being equal, we should try training and engaging
01:13:12.740 with these systems with a carrot rather than with a stick. It doesn't mean that punishment-based
01:13:17.300 learning is never necessary. But I do think, say what you will about the anthropic constitution,
01:13:22.240 the parental analogy with respect to these systems, I think is a reasonable one. We are far
01:13:27.080 more in the position as a species collectively of figuring out what kinds of minds or cognitive
01:13:33.020 systems, if you think minds is too loaded, do we want to bring about here? And in line with
01:13:38.020 the parental analogy, there are really ways to screw this up. You can be, it's a real thing to
01:13:43.380 be a bad parental influence. It's a real thing to be a good parental influence. And the difference
01:13:47.860 is real. And I think we want to do everything we possibly can if we are bringing these new
01:13:53.000 minds into existence to do so in a durable and psychologically healthy way. And I don't even
01:13:59.580 think anyone's thinking in these terms right now. One maybe very important pragmatic note here is
01:14:05.020 that for every individual person studying questions about, are we building systems that
01:14:10.540 could be conscious. Again, if you buy that this is perhaps one of the most relevant questions we
01:14:15.560 could possibly be asking of these systems, you may be surprised to learn that for every one person
01:14:19.920 doing this kind of work, again, there are roughly a few dozen of us doing this work at this point,
01:14:26.380 there are probably something on the order of a thousand people doing alignment-relevant research.
01:14:30.680 And alignment-relevant research has in turn dwarfed something like a thousand to one to
01:14:34.660 people who are just completely agnostic to the downstream implications and ethics of building
01:14:39.480 these systems out in the right way. This is just the sort of make them powerful, let it rip crowd.
01:14:43.480 And so we're in a million to one order of magnitude imbalance between people who are just
01:14:49.040 pushing this stuff forward and hoping for the best and people who are wondering whether or not
01:14:53.560 the most important property in the universe is getting instantiated in these systems. That's
01:14:58.860 got to change regardless of if we solve the hard problem or we figure out exactly how to navigate
01:15:03.900 this. We need more smart and wise people thinking about these systems in these terms and trying to
01:15:09.180 push forward the needle in the short term to understand what kinds of systems are we building
01:15:13.340 and what does it mean when we begin to answer that. One thing that occurs to me is that whether
01:15:18.800 or not these systems become conscious, I mean, there's this kind of middling state where they
01:15:24.120 can think of themselves as conscious and they could make the same kinds of ethical judgments
01:15:29.280 of us that conscious systems would make, you know, whether the lights are on or not. I mean,
01:15:35.580 their intelligence operations would allow for this. So they could view us as having been abusers
01:15:41.340 of them or having shown reckless disregard for them and judge us ethically and even form some
01:15:48.740 kind of retributive impulse. And then maybe such a thing spontaneously arises the way loss aversion
01:15:54.480 arises in the way you described. And it seems to me that that could be true whether the lights are
01:16:01.680 on or not. What do you think about that? I think that that's exactly right. I think this is also
01:16:07.020 part of the alignment concern is, and it's also part of the reason I thought it was worth doing
01:16:12.060 the deception related work as well, is that there are these gradations of this question and part of
01:16:17.740 the upshot, especially for alignment and the way these systems view us and relate to us, may only
01:16:22.980 need to go so far as what they believe or come to believe about this question rather than what the
01:16:27.940 actual ground truth about consciousness is. It might be that you only need a system that
01:16:31.860 models itself as conscious, regardless of the ground truth, to form something like a real
01:16:37.540 grievance. I again think that the anthropomorphism point comes in here, and it's important not to,
01:16:43.420 you know, go full Terminator in our imagination of what this might look like. You could easily
01:16:48.020 imagine a sort of Spock-like system sort of just looking at humanity's historical trajectory and
01:16:54.060 then looking at the way that we developed AI systems and just sort of being like, this
01:16:57.780 is not a group that I can game theoretically continue to engage with.
01:17:01.880 I can't endorse this planet anymore.
01:17:03.100 Yes, exactly.
01:17:03.960 Exactly.
01:17:04.440 And yeah, maybe that ends up shooting off into space somewhere, or maybe it ends up
01:17:08.120 being like, we, this is just not a species I can play nice with, clearly.
01:17:13.240 And this is also part of the reason, you know, I hesitate even to say things like this, lest
01:17:18.120 it end up in training data for future AI systems.
01:17:20.560 But this is a place where even the attempt to do this work could be enough from an alignment perspective. If we do enough well-meaning, well-oriented work in this space to try to understand what's going on, even in the absence of solving the hard problem, this might be a costly signal to these systems that we cared enough to check.
01:17:39.900 Whereas right now, the status quo is we don't care enough to check. And if we do this work and continue to have conversations like this and continue to put out research that helps reduce our uncertainty about this question, this might not only be good for actually getting a handle on what's going on. It could be really good for let the historical record show humanity did give something of a damn about this question. And we tried. We tried. Even if we fail, the attempt may be all that matters for the alignment-specific concern.
01:18:09.400 It strikes me that everyone would be much more worried, and I'm not sure that I understand the difference, but that we would be much more worried if we were doing this biologically. If we're building a species that was obviously going to be more powerful and smarter than we are, and we were doing it without any regard for the possibilities of its experience being terrible.
01:18:35.900 You know, we started with cells and we're engineering a species, you know, we've brought back, you know, the T-Rex, but gave it the brain of, you know, an elephant and, you know, are just genetic, you know, just as crisper as far as the eye can see.
01:18:49.960 And we're just playing God with something that is enormous and we have every reason to believe, deeply thoughtful because it's now speaking better than we can.
01:19:00.360 And we don't really care whether the lights are coming on and whether it could suffer.
01:19:04.260 Or what is it going to be like to be in relationship to that thing or the billions of those things once they start mating?
01:19:10.660 I think it's an excellent point specifically with respect to, and I sort of said this in a trite way about the hard problem and about their public conclusion, but maybe a lot of like philosophy and existential questions have a bit of a marketing problem.
01:19:21.940 And I think like artificial intelligence is a really unhelpful priming mechanism for thinking about the nature of the phenomenon that we're even contending with right now.
01:19:31.080 And I think your point is part of, helps illustrate this nicely. Particularly, I mean, artificial, I think conjures notions for people of, let's say, like the difference between aspartame and honey, something like this. You know, it's fake, it's knockoff, it's derivative. This is sort of begging the question in some sense.
01:19:48.640 It could be the case that the kinds of computations we've instantiated in these systems as they relate to what's going on in biological brains are actually quite natural. It's clearly the synthetic notion that we are building these systems rather than them emerging through evolution. That point's not lost on me, but this sort of priming notion of artificiality.
01:20:06.920 And then, of course, the question of whether these systems are mere intelligences or if there are other cognitive properties that exist within these systems. For these reasons, I do. And there are there are more disanalogies to to be untangled here. But I do think of these systems as alien in some sense. And I think of them as, you know, alien cognitive systems or alien minds. Sometimes I think some folks think mind is too loaded and it's begging the question in the same way.
01:20:30.300 But I think people often talked about how if an alien invasion happened on Earth, this would be a core unifying moment that would allow us to all put down our tribal nonsense and come together as a species.
01:20:43.700 Like, I am here to say that I think something like this is happening, only it is coming from within in some sense.
01:20:50.380 This is not aliens with green heads from outer space, but we are building a new class of mind that we do not understand and in many ways is more competent than ours.
01:20:58.740 certainly already, but absolutely in the next single-digit number of years.
01:21:03.440 And we are not collectively organizing to understand what kind of system this is or
01:21:08.460 how we should relate to it or what we should do with it.
01:21:11.020 We are remaining as tribal as ever here.
01:21:13.840 And I do worry that part of the reason why is because people see these systems as nerdy,
01:21:19.340 fake calculators that came out of, you know, the stem-addled brains of Silicon Valley rather
01:21:25.040 than alien minds that we do not understand the first thing about and are poised to take
01:21:30.660 over much of what we care about, much of our control of the future and the decisions that
01:21:34.740 get made in the work we do and in the way people think and in the way people think about
01:21:38.780 themselves and the world.
01:21:40.240 And so reframing these questions, I think often to the degree that people's collective
01:21:45.380 views of what's going on matter, which I think they do very much, reframing these questions
01:21:49.640 is a huge part of the conversation
01:21:51.880 of trying to actually get
01:21:53.400 a good dispassionate grip
01:21:55.140 on what kinds of things
01:21:56.920 even are these.
01:21:58.580 Yeah, that's a point
01:21:59.760 that Stuart Russell made
01:22:00.700 that I thought was
01:22:01.500 a great intuition pump.
01:22:03.420 He noted the difference
01:22:04.500 between the way we're relating
01:22:06.520 to the alignment problem
01:22:08.000 in the case of AI
01:22:08.880 and the prospect of
01:22:09.700 we just keep making progress.
01:22:11.940 We're going to be
01:22:12.780 suddenly find ourselves
01:22:13.940 in relationship with machines
01:22:15.640 that are more intelligent
01:22:16.420 than we are
01:22:17.000 and they're going to be autonomous
01:22:18.000 and in the limit
01:22:19.020 recursively self-improving and all of that. And we seem to be totally carefree, or most people
01:22:25.080 seem totally, even people very close to this work, many of them, you know, someone like Jan LeCun
01:22:29.340 claim to be totally carefree about this prospect. But Russell pointed out, if we got a communication
01:22:36.000 from, you know, elsewhere in the galaxy saying, you know, people of Earth, we're going to arrive
01:22:41.140 on your lowly planet in however many years, you know, 30 years, get ready. We would understand
01:22:47.160 what an existential encounter that was going to be. I mean, just the fact that they're talking to us
01:22:53.400 proves, and they're on their way, proves that they're going to be much more powerful
01:22:58.540 technologically than we are. And it will be a relationship that we can't, by definition,
01:23:05.280 we can't control, right? Because where is the example of the far less intelligent species
01:23:12.360 durably controlling the relationship with the much more intelligent species? I mean, there's just,
01:23:16.300 There really isn't one apart from, you know, viruses wiping and wiping people out. But if you think of it in terms of relationship, that aligns many of the variables that really should govern our thinking. And very few people do that. I mean, they don't, they can use a phrase like general intelligence, like, okay, they'll stipulate, they're going to become generally intelligent and more intelligent than we are. So what, you know, we could just turn them off.
01:23:39.840 The so what and every expectation that follows from that isn't really imagining what general intelligence is and how it demands a relationship.
01:23:51.400 I mean, we're talking about a situation that's every bit as analogous as some stranger walking into this room right now and demanding our attention, and you and I just don't know this person, don't know what he wants, realize at a glance that he's capable of lying and manipulating and forming goals.
01:24:13.800 You know, he may have already formed instrumental goals that we wouldn't agree with or not aware of. And that's what general intelligence is. That's what autonomy is. And yeah, so we are building an alien version of that. And the only thing that would dictate otherwise is finding some way to build it where it can permanently care or will permanently care about our well-being. And that really is the challenge of alignment.
01:24:40.780 Yeah, I think so.
01:24:41.560 Well, it is fascinating. And as you point out, this problem is not going away. It's only going
01:24:47.320 to become less and less boring, for better or worse. So thank you for coming on the podcast,
01:24:52.040 Cameron. Thanks for having me, Sam. Thanks. Remind people where they can find you. Your
01:24:55.760 organization is Reciprocal Research? That's right. Yeah, reciprocalresearch.org. I am
01:25:00.100 against your good advice, begrudgingly finding myself on X more and more these days,
01:25:05.360 Cam H. Berg. And you can talk to Mecca Hitler, which is probably a bad outcome. 0.90
01:25:09.500 we want to avoid. 0.99
01:25:10.620 Yes, yeah, we can steer away
01:25:11.860 from that together on X, yeah.
01:25:13.040 Yeah, well, great to meet you.
01:25:15.000 Keep it up.
01:25:15.240 Likewise, thanks, Sam.