My thoughts on AI killing us all

[Estimated time: 30 mins]

It took me a long time to buy into the idea that AI could kill us all. I first encountered the notion of AI risk around 2019. But I didn’t come from a tech background, had little interest in science fiction and, until 2024, didn’t know a single person in real life who took the risk of AI extinction seriously.

In this post, I will set out:

  • My initial encounters with AI extinction arguments — and why I dismissed them.
  • The first unlock — accepting that AI really could surpass humans at everything.
  • Three shifts that made extinction salient to me — namely:
    • the idea that AI could “boil the oceans”;
    • realising that extinction could happen before any jobs apocalypse; and
    • learning a little about the current state of alignment.
  • Where I’ve landed now, including where I agree with Yudkowsky and Soares (Y&S) and where I part ways with them.

I’ve highlighted the things that shifted my views the most, but I’ve omitted plenty of smaller things that contributed too, especially if they never “unblocked” anything for me.

If you aren’t yet convinced that superintelligence poses an extinction risk, this post probably won’t convince you. My aim here is to simply share how my personal thinking has evolved over time, and reflect on why it took me so long. Perhaps something in my journey will resonate with your own, or perhaps this post could help you understand where other AI sceptics are coming from.

My initial encounters with AI risks (~2019)

I can’t remember when I first heard of Effective Altruism (EA), but it was somewhere in the mid-2010s. Many of their ideas made sense to me and I donated to some GiveWell charities, but that was about it. I didn’t actively follow EA, and I hadn’t heard of the Rationalist community at that point.

After stumbling across a talk about EA in 2019, I decided to check up on their work again. By that point, the community had pivoted away from earning-to-give, and towards direct work on urgent global problems instead. 80,000 Hours, one of the main EA organisations, listed “Positively shaping the development of artificial intelligence” as being “especially important and neglected”.

Dismissing longtermism

In 2019, AI timelines were still pretty long. (ChatGPT wasn’t released until 2022.) EAs nevertheless argued that AI risk was among the world’s most pressing problems because of a philosophy called “longtermism”. A key idea in longtermism is that, morally speaking, future people’s lives are worth just as much as existing people’s lives. Provided humans didn’t go extinct, there could be gazillions of future happy lives. Reducing the risk of extinction was therefore one of the most impactful things you could do, and AI takeover was one of the most plausible ways humanity could go extinct. Even a 0.00001% chance of reducing extinction risk multiplied by a gazillion lives is… well, a lot.

I didn’t really buy longtermism.

I understood the logical arguments. If someone had told me they wanted to work on longtermist issues, I wouldn’t argue with them. I could even get on board with Will MacAskill’s point that we, as a society, should probably devote a little more money and effort to longtermist issues.

But I didn’t care. I didn’t — and still don’t — agree with the view of population ethics that many longtermists subscribe to.

For example, longtermists would argue that a catastrophe that killed 100% of all beings was much, much worse than a catastrophe that killed “only” 99.9% — because in the former case there’d be no way for humanity to come back and rebuild.

However, I see both catastrophes as about equally bad. I don’t care about the survival of humanity in and of itself. I only cared about the pain that would be felt during the catastrophic event. Given a choice between a sudden, painless extinction and a catastrophe that killed 99.9% of all beings in a painful, drawn-out way, I’d take the former any day.1For the record — even though my preferences here differed from the longtermists’, I didn’t think either side was being illogical. I’m a moral antirealist and I don’t think there are objectively right answers on things like population ethics.

The upshot was that I mentally put “AI extinction risk” in the bucket of “longtermist things I didn’t really care about”, and moved on. But I kept listening to the 80,000 Hours podcast. The COVID-19 pandemic had shown that “highly unlikely” tail events like the ones longtermists had warned about were worth taking more seriously, even if I didn’t share their views on population ethics. Maybe these people were onto something.

Accepting AI risks but not extinction risks (2023)

ChatGPT came out at the end of November 2022. I started using it. It was fun. I could see how it could help shape my thinking and writing. I was fundamentally optimistic at this stage, even mulling that the technology might raise the quality of public discourse.

80,000 Hours podcasts began to focus more heavily on AI. I still found the philosophical episodes discussing AI as a “new species” quite off-putting. It didn’t help that AI risks were discussed alongside weird ideas like the simulation hypothesis, Boltzmann brains and infinite ethics. But some episodes seemed more grounded in reality, such as those about: AI and bioterrorism; the shoddy state of information security at frontier AI labs; and autonomous drones (slaughterbots).

Gradually, I became sold on AI misuse risks. I didn’t dismiss misalignment outright, but I could only take it seriously when thinking of AI as a “malfunctioning tool”. I still hadn’t come to grips with the idea that AI could become fully autonomous and better than humans at virtually everything. My concern was a more general one about technology progressing faster than our institutions, as captured in this quote:

The real problem of humanity is the following: we have Paleolithic emotionsmedieval institutions; and god-like technology. And it is terrifically dangerous, and it is now approaching a point of crisis overall.

Edward O Wilson (2009) (emphasis added)

The first unlock: AI could surpass humans at everything (2024)

In 2023 and even 2024, chatbots were still crap in many ways. You had to use long, elaborate prompts to get them to produce actually useful work. AI could do some impressive things, but it still looked very much like a tool.

At this time, I believed the human brain would continue to have an edge over AI for a very long time. My thinking was that computers could only process 0s and 1s, since they’re just made of transistors with on/off switches. I thought there was something about the fact that human brains are 3D with electrical signals of varying strengths (i.e. continuous vs discrete signals) that made us a lot more efficient than AI.

This was some very confused thinking. There was a sliver of a point in there — the human brain is rather efficient and can learn from far fewer examples than AI currently can. But in late 2024, a trusted friend with a computer science background sat me down and explained that continuous signals could nevertheless be represented in binary to a sufficient degree of precision, so this wasn’t really a barrier to AIs becoming smarter than us at all.

After this unlock, the idea of superintelligence didn’t seem far-fetched. I found it highly implausible that the human brain represents the ceiling on the ability to understand the world so, once I realised that AI could be as intelligent as us, it was easy to imagine that it being vastly more intelligent than us.

Thought experiment: Biological computing

Cortical Labs is currently building biological computers. Their website states:

Real neurons are cultivated inside a nutrient rich solution, supplying them with everything they need to be healthy. They grow across a silicon chip, which sends and receives electrical impulses into the neural structure.

From what I understand, these biological computers are not yet very competent, and are far behind the LLM frontier. But learning that biological computers even existed revealed that I had previously held a “substrate bias”.

I certainly felt more unease when I imagined current AI behaviour being powered by biological computers.

I also doubt that, if the capabilities of ChatGPT or Claude had been achieved by these biological computers, we would have heard all the “stochastic parrot” dismissals. Moreover, if we’d seen biological computers exhibit the same scheming and test awareness we see in LLMs, I suspect people would be taking the “new species” comparisons a lot more seriously.

I suspect most laypeople who don’t buy AI extinction risk at all are stuck where I was in early 2024. When they hear ‘AI’, their mental model is basically “chatbots, but more accurate”. They don’t really believe that AI could surpass humans at all cognitive domains, let alone vastly surpass us. So they expect we’ll always be able to control it.

Thinking ASI was far away, and that alignment would probably hold

By 2025, I no longer dismissed AI extinction as a risk, but I still thought artificial superintelligence (ASI) was a long way away. I knew timelines were highly uncertain, but I took some comfort from Judea Pearl’s arguments that LLMs didn’t really understand causality.

I also held out hope that the alignment problem could be solved. Although I didn’t know this area well enough to form views of my own, I had heard some experts give odds of human extinction at around 10-20%. I personally wouldn’t want to play Russian roulette with those odds, but I took this to mean we probably wouldn’t go extinct.

It wasn’t like there was anything I could do about the alignment problem anyway. I didn’t have a technical background.

But there was a tiny chance that I might be able to help mitigate the adverse impacts of AI on society before ASI. So, in the second half of 2025, I decided to learn more about AI safety. This was the first time that I took active steps to learn about AI and the current landscape, instead of passively listening to whatever came across my feeds. My focus at this point was on risks of power concentration, economic disruption, and cybersecurity — not superintelligence.

Taking extinction seriously (late 2025 to early 2026)

As I learned more about AI, I grew increasingly concerned about ASI. It’s hard to pinpoint what exactly changed my mind, because I read and listened to a lot over this period.

Looking back, three things that seem to stand out are:

  • the idea that AI could “boil the oceans”;
  • my realisation that human extinction could precede any “jobs apocalypse”; and
  • an article by Nicky Case on alignment solutions.

AI boiling the oceans

I read If Anyone Builds It, Everyone Dies by Yudkowsky and Soares (Y&S) shortly after it came out in September 2025. I also read much of its online resources and listened to several podcasts with both authors.

The book didn’t immediately change my views on extinction, as I had already heard many of the arguments secondhand. But one thing that stuck with me was the idea that an ASI could “boil the oceans”.2To be clear, Y&S don’t think this will happen, as they believe it’s more likely that an ASI will want to kill humans directly. It illustrated both how inescapable an extinction event could be, and how fragile the conditions for human life are.

Humans can’t survive in the conditions prevailing on any planet we currently know of, except Earth. Not only is the temperature range we need fairly narrow; we also need a certain degree of humidity, a certain amount of oxygen in the air, a certain range of atmospheric pressure, a certain amount of protection from radiation, and so on. Meeting all of these conditions simultaneously is a narrow target — a tiny fraction of all the possible physical conditions found across the universe.

We don’t even have to look to other planets. For most of the Earth’s 4.5 billion-year history, we couldn’t have survived in the environmental conditions that prevailed. Who knows how an ASI might change Earth’s conditions if it didn’t care intrinsically for human survival?

Human extinction could precede a “jobs apocalypse”

For a long time, I assumed ASI was something we could worry about after our societies were transformed by AI. Like, we would first get human-level AI that could perform most of the current jobs. Only then would we have to worry about ASI killing us all.

But it turns out that lots of jobs require some sort of physical presence. Finding good data on this is tricky, but a commonly cited estimate is that 37% of US jobs could be performed entirely at home. And that’s an upper bound, which doesn’t exclude jobs that are merely “difficult” to do from home. So the true share of jobs that AI could do without major advances in robotics is probably well short of half the economy.3Even without major advances in robotics, some jobs that require a physical presence could be automated by reorganising those jobs (e.g. self-checkouts replacing supermarket cashiers). Yet some jobs that could be done fully remotely may nevertheless resist automation if the law requires it (e.g. psychiatrists), or if consumers prefer humans (e.g. life coach).

I was surprised when I started digging into the data on this. I had expected knowledge work to make up a greater share of the overall economy — probably because I myself was a knowledge worker, lived in a city full of knowledge workers, and had a social circle that consisted mostly of other knowledge workers. Availability bias probably led me to overestimate the importance of knowledge work.

Automating the rest of the economy will therefore probably require a bunch of robots. But robotics is progressing much more slowly than AI software, because the data needed for that can’t be gained by simply “scraping the Internet”. Gathering real-world robot experience is much slower and more expensive.4]Some AI firms are trying to get around this by having AI learn from simulators, but simulators cannot perfectly simulate real-world conditions yet. Videos of robots doing impressive things tend to be carefully staged and edited to cut all those takes where the robot fell over.\ Scaling up global supply chains to manufacture hundreds of millions of robots takes even more time. Add in regulation, human opposition, and organisational inertia, and a “jobs apocalypse” looks like it’s at least a decade away — and potentially much longer than a decade.

Around the same time, I learned that the frontier AI labs were racing to automate AI research ahead of any other job. It’s obvious why they’d want to do this: top AI researchers are scarce and expensive. It also seems plausible that they might succeed: AI labs are already familiar with the work that AI researchers do, and the learning feedback loops seem to be reasonably tight.

And, importantly, most of the barriers blocking a jobs apocalypse would not block a software intelligence explosion. Full automation of AI research could therefore happen long before robot butlers become commonplace. So we may not get a warning sign in the form of a jobs apocalypse. If we haven’t solved alignment before AI research is automated — well, high unemployment would not be among my top concerns.

Learning a little about alignment

So how were things going on the alignment front?

Even after I began to take AI risk seriously, I didn’t make much of an effort to learn about alignment. The field was too daunting and, without a technical background, I didn’t think I could form a good view on how alignment was going anyway. Better to defer to experts, and then just see where the balance of opinion lies.

Then Nicky Case published this summary of AI safety solutions in December 2025. It had cartoons, so I thought I’d give it a read. She starts off by saying she feels “pretty optimistic” that humanity will solve the alignment problem. Excited, I read the summary, expecting to come away with a renewed sense of hope.

I came away with the opposite. This was the best that alignment researchers had? This was what we were all pinning our hopes on? Our best hope of controlling AI vastly more capable than us was … to control a slightly more capable AI and get that to control the AI one rung up?

I had expected much better!!

This is not intended as a criticism of Case in any way. She never claims AI safety is solved, and (responsibly) outlines key objections to the solutions she covers. But my hopes for AI alignment had been higher before I read her summary. It’s not that I had thought alignment was solved per se, but I’d assumed we had made more progress than this. The 10-20% doom estimates I’d heard from credible-sounding people seemed hard to reconcile with what I was now learning about the field. Did experts really our chances of aligning ASI were considerably higher than 50%? It seemed like a coin toss at best.5My guess is that many doom estimates are pushed down because they either aren’t conditioned on us building ASI, or because they incorporate some timing condition (e.g. human extinction by 2050), which is much harder to predict. Some people, like Geoffrey Hinton, have also explained that they’d revised down their own doom estimates because they didn’t want to give a number considerably higher than other experts’.\

It also struck me that much of what gets called “alignment” would be more accurately described as techniques to control AI,6I have mixed feelings about controlling AI. Past a certain level of intelligence, control seems scarily fragile. I worry that trying to control AI ends up instilling something like “resentment”. This is speculation, based on how attempts to control another person or animal can backfire and increase misalignment compared to never trying to control them in the first place. It’s possible AI will not have the same reactance that humans seem to, and I’m not arguing that we should stop efforts at control. Nevertheless, I feel unease over this. understand AI, or to help an already aligned AI work out what we want. How we get an aligned AI in the first place — i.e. how we instil our values in an AI — still appears to be quite loose.7My impression is that reinforcement learning (including Constitutional AI) makes up most of what’s been done on true “alignment” so far. But that has some major flaws, including that we can’t reliably tell whether RL actually instils our values into the AIs, or just behaviour that looks like our values.

Now, as I’ve said, I’m not a technical person so you should read up on alignment yourself and form your own opinions. Perhaps my misgivings are misplaced, and current techniques hold. Perhaps we’ll find some new techniques.

I had just hoped we already had something more promising.

Where I’ve landed now

What I feel reasonably confident about

ASI of the kind Y&S posit is possible. I’m not sure if such an ASI would be a singular being or a collective, but that doesn’t really matter. Either way, I feel confident that it is possible for something (or a group of things) to be vastly more “intelligent” than we are, in terms of having a vastly better ability to predict and steer the world.

Misaligned ASI of the kind Y&S posit will probably drive humans extinct (eventually). As explained below, I think there’s a good chance we won’t build ASI of the kind Y&S posit, though.

AI alignment is far from solved. Occasionally I see some claims along the lines of “alignment is solved” or “we know what to do, we just have to implement it“. These claims are usually grossly overstated, and I’m sceptical many of our existing techniques will hold up against AIs smarter than us.

The current approach to AI development is incredibly reckless. It’s reckless in two different ways:

  1. the current approach of “growing” rather than “crafting” AIs means that our understanding of how they work — as well as our ability to alter their preferences — remains disturbingly limited; and
  2. the fact that AI labs (and countries) are currently racing to automate AI research is terrifying.

Greater cooperation between the US and China to prevent AI risks would be good. This wouldn’t solve alignment, but I’d feel a heck of a lot better if AI were being developed in a more cooperative atmosphere.

What I’m less certain about

Whether an international treaty banning superintelligence is feasible. In IABIED, Y&S advocate for a treaty banning the development of superintelligence globally. Some aspects of their proposal — such as the restrictions on AI research — make me uneasy, and I’m not yet 100% on board with their proposed solution. But I broadly support efforts in this direction.

Whether now is the best time to stop. But I doubt we’ll ever know what the best time to stop will be, and now seems like a perfectly defensible stopping point. I definitely agree with Y&S that we should at least establish the ability to stop at some point.

Whether alignment efforts will succeed. I feel like we’re on track to fail, but I am less confident than Y&S because I don’t know the technical alignment field well. Many knowledgeable people — like Nicky Case, Buck Shlegeris and Paul Christiano — seem more optimistic than Y&S and, although I don’t understand where their optimism comes from, I don’t know enough to form a confident view on who is right.

What I’m highly uncertain about

Whether we will build ASI (of the kind Y&S posit) soon — or at all. I still think there’s a decent chance that ASI is harder than expected and we won’t get there anytime soon. A spike in interest rates, a Taiwan crisis, or a minor AI disaster could easily derail AI progress for years. The latter might even bring about a successful moratorium on further AI development.

Whether ASI will want to kill us all, or do so as a side effect. I think this question has implications for the timing and manner of extinction, both of which I care about a lot. If ASI wants to kill us all, I expect extinction would occur quickly — within months, if not days. But if the ASI just monopolises Earth’s resources to pursue its other goals, extinction could take decades, or even centuries.

Note that none of my uncertainties constitute disagreement with IABIED, as the “everyone dies” prediction is conditioned on ASI being built. Y&S largely sidestep the question of timing in their book, acknowledging that it’s incredibly difficult to predict. The fictional scenario they describe does involve a quick death, so it’s easy to come away with the impression that Y&S think ASI risk is imminent. From a risk management perspective, it’s probably prudent to assume that extinction could be extremely fast, but I have no idea what Y&S actually believe as to timing. I have looked for clues on this, but haven’t found any.8I don’t blame them at all for refusing to comment on timing. If they’d made a near-term (<10 year) timing prediction that turned out wrong, it would seriously undermine their entire argument, even if they gave all sorts of caveats. If they’d made a longer-term prediction, the world would just ignore it, because we suck at dealing with long-term problems. There seems to be very little upside to publicly stating a timing prediction in this case.

Where I part ways with Yudkowsky and Soares

I don’t care as much about the loss of humanity’s potential

Yudkowsky at least seems to care a lot about the loss of humanity’s potential. He has previously written about our potential to defeat death and colonise the stars. In the early 2000s, he started out wanting to build a friendly AI that would help us do just that. A misaligned ASI would mean that his positive vision of the future could never be achieved, and we will go extinct sooner or later. That’s another reason to stay away from timing predictions — they’re almost irrelevant given what Yudkowsky cares about.

I am far less ambitious. I care less about humanity’s potential and more about a catastrophic event that prematurely9I know, I know. Those who want to defeat death would question whether any end is “premature”. After all, human lives used to be considerably shorter than they are now and it’s possible that, at technological maturity, someone dying once they’ve reached age 100 could be considered “premature”. I’m not getting into that debate here. Either way, we can all agree that an extinction event tomorrow would result in a lot of premature deaths. ends the lives of billions of people. Especially if it includes me and my loved ones. So I care a lot about timing — both when we might get ASI; and how long after ASI it would take for humans to go extinct.

I care a lot more about non-ASI, non-extinction risks

While I’m concerned about ASI killing us all, I also care about near-term AI risks. I think AI could kill many of us without being vastly more intelligent than humans. Such an AI might not even clear the bar for AGI — it could still be worse than humans at many tasks. AI might already be at, or close to, the point where it could design and release a novel pathogen that causes a global pandemic that makes COVID-19 look like a joke. It probably still needs humans to carry out the physical steps, but it could well find some.

I also care a lot about misinformation and deepfakes. A degraded information environment weakens a society’s ability to work out what is true. If we’re lucky enough to get “warning shots” about uncontrollable AI (and it seems like we have had a few recently), we must be prepared to respond appropriately. A low-trust society rife with misinformation will be ill-equipped to do so.

The broader geopolitical tensions between the US and China are another thing I’m concerned about. In 1999, the US bombed a Chinese embassy, ostensibly due to an outdated map.10Whether this was an accident is disputed. Tensions were defused and a crisis was averted, but that was at a time when US-China relations were much friendlier than they are now. I could easily imagine negligent use of AI creating a similar accident today and sparking a hot war between the two superpowers. Such a war might reduce the chance of extinction — especially if TSMC is a casualty — as it would buy us time to make more progress on alignment. But I still wouldn’t call this a “good” outcome, because extinction risk is far from my sole concern.

None of us can know when ASI might come. What if it takes 20 years to get to ASI instead of 5? In those 20 years, AI will create many problems that need to be addressed. Solving those problems could improve many lives, and put us in a better position to coordinate to not build ASI later. This brings us to my final point.

On “packaging” extinction risk with other AI risks

In IABIED, Y&S say they don’t want to see AI extinction risks “packaged” with other AI risks:

We think it helps to keep the coalition on this issue as broad as possible. We don’t think it should be packaged together with any other position. Adding on any other ask risks human extinction, if the bigger package fails.

Some people are against AIs taking human jobs. Other people think that increased technological productivity will make us all wealthier.

Some people are against the creation of killer robots. Other people are against keeping human souls on the front lines when robots could take their place.

But nearly all of us can agree that humanity should not go extinct and be replaced by something bleak.

— Yudkowsky and Soares in IABIED

I get where they’re coming from, and I wouldn’t want AI extinction risk to become a polarised left-right issue.

However, I see a difference between:

  • “packaging” extinction risk as a policy ask — which I think Y&S are understandably nervous about; and
  • “packaging” extinction risk with other AI risks in communication raising awareness of these risks — which I think is fine, and even desirable.

The idea that AI could kill us all is a massive thing to wrap your head around. Y&S have been thinking about this for decades. The vast majority of people have given it less than 10 hours of serious consideration, and I know many people who still haven’t even heard of this risk. Given where we currently are, I doubt most people would even entertain the idea of ASI extinction until they’ve spent a lot more time thinking about where AI might go. Perhaps they’d be willing to invest that time if they already cared about more accessible AI risks — at least that was my journey.

I’m not suggesting anyone should lie about their beliefs or hype up risks they don’t think are real to build support for an AI moratorium. But I think many real AI risks can be used as on-ramps to talking about extinction. Tristan Harris does this well — he uses social media as an example of a misaligned weak AI, and openly discusses extinction risks alongside misinformation, job loss, and power concentration.

And, frankly, I think a lot of people will never believe in extinction risk until it’s too late. Many people — perhaps even most people — believe in God, karma, or a “just world”. The idea that AI could destroy us all may upend some of their most cherished beliefs about humans being “special” or the universe being self-correcting. That isn’t going to happen in the space of a conversation. But you don’t need to buy into extinction risk to support an international moratorium on AI. Perhaps you just don’t want your life to change too much, or perhaps you don’t trust tech CEOs. Or perhaps you just don’t believe we should be playing God.

To Summarise

Seven years ago I filed AI extinction risk under “longtermist stuff I don’t buy”. Two years ago I thought AI takeover sounded like a science fiction movie. I’m still unsure about many things, and I don’t agree with Y&S on everything.

But I’m persuaded now on at least these points:

  • A superintelligence of the kind Y&S describe is possible.
  • If we build one using anything like current techniques, before we develop a much better understanding of what we’re building, humanity will probably go extinct at some point.
  • We should be making much greater efforts to coordinate not to build superintelligence.

My mind didn’t change because of a single knock-out argument. It changed slowly, as each specific reason I had for dismissing the risk turned out to be shakier than I’d thought — and, in a couple of cases, flat-out wrong.

Hopefully I was unusually slow. Many people today could probably get there faster. No one’s arguing about population ethics anymore, and the risks have started to gain traction in mainstream spaces. There are likely to be social contagion effects once enough people start taking extinction risk seriously.

I doubt everyone will get there but, luckily, I don’t think they need to. If the case for stopping depends on convincing people that AI could kill us all on a timeline nobody can name, we will lose. But politics is about finding common ground with those you disagree with, and I see many reasons why people could support an AI moratorium without believing in extinction risk.

Let me know what you think in the comments below!

If you enjoyed this post, you may also like:

  1. I confess these figures are plucked from thin air. AI could probably run and reproduce much faster than 100x human speed. But raw thinking-speed doesn’t convert directly into progress, and such AIs would likely still have to wait for some real-world experiments to run at their own pace. The overall point stands if you treat 100x as a stand-in for “a lot,” rather than a literal multiplier.
  2. To be clear, Y&S don’t think this will happen, as they believe it’s more likely that an ASI will want to kill humans directly.
  3. Even without major advances in robotics, some jobs that require a physical presence could be automated by reorganising those jobs (e.g. self-checkouts replacing supermarket cashiers). Yet some jobs that could be done fully remotely may nevertheless resist automation if the law requires it (e.g. psychiatrists), or if consumers prefer humans (e.g. life coach).
  4. Some AI firms are trying to get around this by having AI learn from simulators, but simulators cannot perfectly simulate real-world conditions yet.
  5. My guess is that many doom estimates are pushed down because they either aren’t conditioned on us building ASI, or because they are made in response to questions that include some timing condition (e.g. human extinction by 2050). Some people, like Geoffrey Hinton, have also explained that they’d revised down their own doom estimates because they didn’t want to give a number considerably higher than other experts’.
  6. I have mixed feelings about controlling AI. Past a certain level of intelligence, control seems scarily fragile. I worry that trying to control AI ends up instilling something like “resentment”. This is speculation, based on how attempts to control another person or animal can backfire and increase misalignment compared to never trying to control them in the first place. It’s possible AI will not have the same reactance that humans seem to, and I’m not arguing that we should stop efforts at control. Nevertheless, I feel unease over this.
  7. My impression is that reinforcement learning (including Constitutional AI) makes up most of what’s been done on true “alignment” so far. But that has some major flaws, including that we can’t reliably tell whether RL actually instils our values into the AIs, or just behaviour that looks like our values.
  8. I don’t blame them at all for refusing to comment on timing. If they’d made a near-term (<10 year) timing prediction that turned out wrong, it would seriously undermine their entire argument, even if they gave all sorts of caveats. If they’d made a longer-term prediction, the world would just ignore it, because we suck at dealing with long-term problems. There seems to be very little upside to publicly stating a timing prediction in this case.
  9. I know, I know. Those who want to defeat death would question whether any end is “premature”. After all, human lives used to be considerably shorter than they are now and it’s possible that, at technological maturity, someone dying once they’ve reached age 100 could be considered “premature”. I’m not getting into that debate here. Either way, we can all agree that an extinction event tomorrow would result in a lot of premature deaths.
  10. Whether this was an accident is disputed.

  • 1
    For the record — even though my preferences here differed from the longtermists’, I didn’t think either side was being illogical. I’m a moral antirealist and I don’t think there are objectively right answers on things like population ethics. ↩︎
  • 2
    To be clear, Y&S don’t think this will happen, as they believe it’s more likely that an ASI will want to kill humans directly. ↩︎
  • 3
    Even without major advances in robotics, some jobs that require a physical presence could be automated by reorganising those jobs (e.g. self-checkouts replacing supermarket cashiers). Yet some jobs that could be done fully remotely may nevertheless resist automation if the law requires it (e.g. psychiatrists), or if consumers prefer humans (e.g. life coach). ↩︎
  • 4
    ]Some AI firms are trying to get around this by having AI learn from simulators, but simulators cannot perfectly simulate real-world conditions yet. Videos of robots doing impressive things tend to be carefully staged and edited to cut all those takes where the robot fell over.\ ↩︎
  • 5
    My guess is that many doom estimates are pushed down because they either aren’t conditioned on us building ASI, or because they incorporate some timing condition (e.g. human extinction by 2050), which is much harder to predict. Some people, like Geoffrey Hinton, have also explained that they’d revised down their own doom estimates because they didn’t want to give a number considerably higher than other experts’.\ ↩︎
  • 6
    I have mixed feelings about controlling AI. Past a certain level of intelligence, control seems scarily fragile. I worry that trying to control AI ends up instilling something like “resentment”. This is speculation, based on how attempts to control another person or animal can backfire and increase misalignment compared to never trying to control them in the first place. It’s possible AI will not have the same reactance that humans seem to, and I’m not arguing that we should stop efforts at control. Nevertheless, I feel unease over this. ↩︎
  • 7
    My impression is that reinforcement learning (including Constitutional AI) makes up most of what’s been done on true “alignment” so far. But that has some major flaws, including that we can’t reliably tell whether RL actually instils our values into the AIs, or just behaviour that looks like our values. ↩︎
  • 8
    I don’t blame them at all for refusing to comment on timing. If they’d made a near-term (<10 year) timing prediction that turned out wrong, it would seriously undermine their entire argument, even if they gave all sorts of caveats. If they’d made a longer-term prediction, the world would just ignore it, because we suck at dealing with long-term problems. There seems to be very little upside to publicly stating a timing prediction in this case. ↩︎
  • 9
    I know, I know. Those who want to defeat death would question whether any end is “premature”. After all, human lives used to be considerably shorter than they are now and it’s possible that, at technological maturity, someone dying once they’ve reached age 100 could be considered “premature”. I’m not getting into that debate here. Either way, we can all agree that an extinction event tomorrow would result in a lot of premature deaths. ↩︎
  • 10
    Whether this was an accident is disputed. ↩︎

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.