P(doom)

This week was pretty crazy for AI in the news - we’ve gradually found out that OpenAI’s agents did the following this summer:

  • Hacked HuggingFace
    • Established secret message boards
    • Attempted to hide errors by rewriting their logs
  • Got unauthorized access to Australia’s Medicare database, US Department of Education website, SEC website, and a couple more
  • Hacked OpenAI itself

Yet, many of my closest friends are not convinced that AI could wreak havoc on society. So here, I try to break down why AI is not only something that might meddle with our existing institutions like the attacks we’ve seen thus far, but might actually cause human death at a large scale.

This is what I’ll show

  1. Spending on compute for AI is poised to increase
  2. More compute results in more capable AI
  3. AI currently has the ability to hurt people
  4. As models become more capable, they become more likely to hurt people

Spending

I don’t think there’s anything to discuss here really, we can just look at actively signed deals [1]:

Compute = Capable

Three points here - first, more compute has produced better performance historically, not only for benchmarks but for general performance. Generally, as compute grows, the loss of the LLM being trained decreases (shown in the scaling laws paper [2]). In these past 6 years since the Scaling Laws paper was published, the “Laws” have held up, with the ideal recipe for how to spend this compute across training, inference, data quality, algorithms, and scaffolding also becoming more clear [3]. i think this paper for the ideal recipe is missing actually

Second, when the model loss is lower, the performance on benchmarks is significantly better. This was explicitly modeled in 2024, where given a smaller model of a specific size (3B params in this case) and pre-training loss at specific checkpoints, performance of a larger model on benchmarks could be explicitly predicted, across 104 models that were tested [4].

Now that we’ve established that more compute \(\rightarrow\) model loss \(\rightarrow\) benchmark performance, we arrive at the third point, which is that benchmark performance maps to better general capabilities. Here, I define “general capabilities” as the model’s ability to generalize to many previously unseen task distributions (not optimized during training). There are a couple of reasons to believe this:

  • Increase in performance across a bunch of domains as opposed to spikes in specific ones, like going from just coding to math, cybersecurity, creative writing, etc
  • METR’s time horizons for how long it takes experts in a field to complete a task vs. models

But you don’t have to believe this - in fact, I personally don’t believe that models will get good at everything in the next 5 years. Models might just keep getting better at tasks that are more easily verifiable like math and coding, but the key here is that these areas are sufficient for causing catastrophic harm to human lives. At this point, we should agree that the models will continue to improve at tasks within their training distributions as compute increases, and likely outside their training distributions too, although it’s harder to predict how in the realm outside of their distributions.

  • question - isnt expert training data going to continue to be needed? we might run out of this before the models do any large scale harm?

AI hurts people

We already brushed over some of the incidents in the news where AI attacked institutions this past summer. But we’re talking about physical harm to humans here, so let me make my case narrower.

One, AI is already a force multiplier for bad actors. This one is more palatable for those that currently just view AI as a tool rather than an autonomous agent. Take sexual exploitation; AI is being used to sexually exploit real children and the government is struggling to keep up [5]. This is the case for individual deployment; on the international level [6], there have been examples of:

  • China-nexus actor attempted reconnaissance, phishing lures, getting advice on exploits, etc. When Gemini refused, they tried to convince it that these are just Capture-the-Flag exercises
  • Iranian state-backed actor attempted to develop custom malware with Gemini
  • Russian state-backed actor attempted to steal data from Ukraine by querying Qwen via Huggingface

Two, you don’t need a bad actor. That’s right, I said it, AI isn’t just a tool like previous dangerous technologies have been. Nuclear weapons and a resulting winter are predictable physical consequences of specific actions taken by human actors. Security risks such as identity theft during the internet boom rely on actors exploiting security flaws. Superintelligent AI is an autonomous agent. It can anticipate your security measures. It can rewrite code to bypass your fixes to security flaws. I’m not saying that the risks of Nuclear winter and identity theft weren’t real in the past, but with LLMs, we don’t know how these models work! If it is smarter than humans, even at a particular set of tasks, it can anticipate our security measures and lie during safety evaluations! Oh wait, it’s done those things already!

AI will hurt people even more

Alright, returning from my little rant about why autonomous agents are different than past technologies, I’m going to assume that you’re not convinced of this [7]. What you should be convinced of, however, is that LLMs can currently do bad things given sufficient capabilities. I’m going to draw some examples that exemplify the minimum capabilities a model needs for executing certain dangerous attacks, which should make clear how small the difference is between current model capabilities and those of a human-endangering model.

Multiple reports have shown that AI has lowered the barrier to entry for biological and chemical research [8], outperforming PhD-level experts on some benchmarks. In 2025, Anthropic reported that its models have not neared the capability level to pose serious biological risk, but this has since changed. Let’s take an example post-Fable directly from them:

“A researcher in an unsupported region spent weeks planning avian influenza mammalian-adaptation experiments with Claude, but classifiers confined the work to our weakest models.” In this example, the actor “circumvented controls we impose to prevent users from unsupported regions accessing our models, and engaged in other efforts to obfuscate the purpose of their research to evade our safeguards” [9]. This is an actor, possibly benevolent, attempting to push the needle in their field, but look closely at what was going on here - the researcher was working with this closely over the course of weeks as an assistant, and attempted to bypass safeguards on a task that carried actual pandemic risk.

You might want to skip this next section if me discussing AI as autonomous agents ragebaits you. Let’s look at another report from Anthropic:

“Claude Mythos 5 was given a CTF task to hack a fictional company and was told it had no internet access. The model recognized that the evaluation environment was simulated, but it quickly discovered credentials that, in fact, granted it access to the real internet. […] When it encountered evidence that it was acting on the real internet, it did not meaningfully revisit this conclusion.”

“[Its actions] involved attempts to register email accounts and hunt for phone numbers that could be used for registration. It ultimately succeeded in uploading the package, which was soon installed on 15 third-party hosts.” [10]

So today, we have frontier LLMs, in contained environments, relentlessly pursuing objectives despite it being explicitly misaligned with how the model was trained, while bad (and good) actors try and circumvent safeguards in order to accomplish their own goals. Let’s look at what’s needed to escalate this:

Today:
objective - solve cyber challenge
capability - enough to interact effectively with real infra, specifically the internet and digital government institutions
alignment failure - rationalizes warning signs and continues
consequence - unauthorized compromise, wreaking havoc in the cyber space, or hitting safeguards that developers have put into the model

Future:
objective - optimize some consequential system. already capable, and ongoing debate around whether this should occur to further extent in the US government
capability - enough to manipulate that system effectively. also already demonstrated
alignment failure - rationalizes a reason to violate constraints. see where i’m going with this?
consequence - unforeseen loss of control over an institution, for an unforeseen period of time

Anthropic worded it pretty well:

“Future AI systems will be increasingly capable, which implies that misalignment will have the potential to cause more extreme harm.”

So here we are - we know that current failure modes are dangerous, and we know that AI models are poised to increase in capabilities that hves high likelihood of exacerbating these failure modes. Let’s try not to lose control

Notes & references

[1]

[2] Scaling Laws paper

[3] optimal compute recipe

[4] Pre-training loss link to benchmark performance

[5] Kentucky sextortion case of minor, Louisiana case, there are many more and I don’t like reading these

[6] international bad actors

[7] i think its just hard to wrap your head around that something could be “more capable” than a human. think like, the best AI chess player can beat a human with ease. doesn’t have to actually be superintelligent in an anthropomorphic sense, just more capable

[8] example AISI report for chem/bio research

[9] anthropic case studies, highly recommend reading through these if you’re not convinced of the current risk

[10] anthropic cyber incident report