
Ever wondered if you — along with the rest of human civilization — might die in a couple of months? In this week’s deep dive, we look at the growing concern about AI extinction and try to figure out whether or not our future will look like The Terminator.
From the Hugging Face attacks to the highly public resignation of a leading AI researcher, the last couple of weeks have seen mounting concerns over the rapid advancement of AI models. While some have chalked concerns up to overactive imaginations and elaborate marketing, others have begun buying bunkers and preparing for the apocalypse.
In times like this, It can be difficult to orient oneself. Does AI pose an existential risk to humanity? Is it already too late? Why can’t we just turn off the computer?
While we do not have all of the answers, we hope that the following will help you to parse one of the more pressing political issues of our current moment.
A Brief History
Concerns about AI safety have long preceded the actual existence of AI.
As early as Harlin Ellison’s I Have No Mouth and I Must Scream, the American public has wondered whether or not the future might see the rise of an omnipotent and malicious robot overlord.
Contemporary concerns over AI safety, however, started to take shape in the early 2010s. In 2014, Oxford philosopher Nick Bostrom released Superintelligence, which has become the semi-definitive text of today’s AI safety space.
In Superintelligence, Bostrom imagines an AI system that is significantly more intelligent than any living human. He goes on to argue that no matter what goals this system is given, it is liable to develop “instrumental goals” (e.g. self preservation) that would lead to catastrophic outcomes (e.g. the release of a bioweapon that would prevent any human from being alive enough to shut it off).
Bostrom’s concerns, which have been echoed by safety advocates like Nate Soares and Elizier Yudkowsky in their recent best-seller If Anyone Builds It, Everyone Dies, have fundamentally shaped the current AI safety discourse.
In other words, concerns about AI extinction long preceded the profitability of said concerns, even if “extinction-level threat” happens to be good marketing for an upcoming IPO.
But does this mean that the end is nigh?
Alignment, Improvement, and Takeoff
Today, safety researchers are largely focused on alignment.
As the name would suggest, alignment research is invested in figuring out a way to build models that are aligned with a set of goals that developers deem important. But while this sounds good and nice, alignment is not as simple as it may sound.
First and foremost, the neural networks that comprise LLMs are something of a black box. While the surface level reasoning of LLMs can be roughly reconstructed, current models are capable of in-depth reasoning that far exceeds our understanding (see, for instance, Anthropic’s analysis of “J-Space”). As a result, it is difficult to determine whether a model is actually aligned, or whether it is just saying things that an aligned model ought to be saying. This has led some to conclude that meaningful alignment is impossible.
Second — assuming that alignment is possible — the obvious next question goes something like: what values should an LLM be aligned with? For many, the idea that Sam Altman or Elon Musk will answer this question is not reassuring. Though MechaHitler might be gone, he has not been forgotten.
Still, alignment becomes increasingly important as models improve.
Over the last year, the pace of model improvement has sped up tremendously. For many in the safety space, this rapid advancement has heightened concerns over recursive self-improvement.
Right now, model development is bottlenecked by the human ability to improve models. Recursive self-improvement describes the state in which models become smart enough to break ground on AI research and improve themselves. This process, wherein the rapid improvement of models becomes entirely recursive and unchecked, is called takeoff.
What Can We Do?
Doubtlessly, this all seems very bleak. Still, it might be worth waiting a tick before preparing for the end times.
Over the past few weeks, AI safety concerns have reached the mainstream. So far, this has amounted to a White House dinner, a congressional hearing on rogue AI agents, and a tentative agreement among American AI CEOs to try to be safer. While this is hardly satisfying, especially given extinction-level stakes, it suggests that AI will continue to dominate the political landscape through 2026 and onward to 2028. Already, politicians like Bernie Sanders have embraced legislation calling for a global pause on development.
Is it still too late? Perhaps. In any event, there has never been a better moment to find out more.
Thanks for reading!