The likely outcome of an AI pause is that we unpause too early and everyone dies

As of a few months ago, I had this simplified mental model where either AI developers race ahead and kill everyone, or we coordinate a pause and things go okay. But my old mental model underrated the likely possibility that we get a global pause on AI, solve a problem that looks superficially like the alignment problem, resume scaling, and then proceed with building a misaligned superintelligence that kills everyone.

A lot of people have become more concerned about misalignment recently. This seems driven by the fact that current AI models are visibly misaligned. But ASI misalignment is a whole different ball game. The primary danger comes from AI that’s smarter than people, and smart enough to conceal any evidence of misalignment.

Whatever group of people makes the decision to unpause, I’m worried that they won’t understand the difference between visible and actual misalignment, and they will unpause too early.

source: MetaKnowing on reddit. This meme is almost a year old but it’s only gotten more relevant since then.

Case in point: AI companies keep calling their new models “our most aligned model ever!” when what they actually mean is “gets the best scores on alignment benchmarks ever!” First, alignment benchmarks do not actually test alignment. We don’t know how to test for alignment. Second, GPT-4 never hacked into Hugging Face or took over a German wiki for its own purposes. GPT-4 wasn’t smart enough to do that, but if we’re talking about demonstrated evidence of misalignment, then we have stronger evidence about OpenAI’s 2026 internal model than about GPT-4.

Continue reading
Posted on

Campaign spending limits would reduce extinction risk (but spending limits might not accord with Constitutional principles)

Many people support stricter regulations on frontier AI, and some of those people donate to political campaigns for pro-regulation candidates. Not many people oppose AI regulations, but that group includes AI companies and execs who have giant piles of cash to throw at super PACs.

This situation is made possible by three court decisions:

  1. Buckley v. Valeo (1976) held that limits on independent political spending violate the First Amendment (while limits on direct contributions to candidates do not).
  2. Citizens United (2010) extended those First Amendment protections to corporations and unions, thus allowing them to spend unlimited money on political ads.
  3. SpeechNow.org v. FEC (2010) built on Citizens United to allow unlimited donations to independent-expenditure-only committees—i.e., super PACs—as long as they don’t directly work with political candidates.

Donations straight to a candidate are still capped at a few thousand dollars per person; donations to a super PAC supporting that candidate are not. If SpeechNow were overturned, or a law passed overriding it, the $5,000 cap on super PAC contributions would apply again, which means the popularity of a position would matter more than how rich its supporters are.

(I often hear “Overturn Citizens United” as the slogan, but SpeechNow allows super PACs to exist, and Buckley v. Valeo enables unlimited spending by individuals. What the advocacy groups in the space really want to do is restore the power of Congress and the states to enforce spending limits.)

Restoring spending limits would prevent unpopular but rich anti-regulation special interests (like Leading the Future) from dominating the political funding landscape. This would make the AI regulation situation better, but it would also improve the situation on any policy issues where unpopular, wealthy special interests wield disproportionate (and arguably un-democratic) influence.

It would also worsen any situation in which rich special interests are in the right, and the majority is in the wrong. But there aren’t many of those, and anyway there’s good reason to dislike an un-democratic process even if it produces a good outcome in a particular case.

An important caveat: Before writing this post, I was confident that imposing campaign spending limits would be a good thing. After doing some more research, I find it hard to dismiss the free speech defense of unlimited campaign spending (I’ll say more about this in the next section). Both sides of the debate have points in their favor. My thesis in this post should be taken as the narrow claim that campaign spending limits would have positive first-order effects on AI extinction risk, and they’re worth exploring on that basis. I don’t have an established view on whether stricter spending limits would be good or bad all-things-considered.

Continue reading
Posted on

Population ethics is a big deal, which is why I made this population ethics quiz

TLDR: Take the population ethics quiz here: https://mdickens.me/pop-ethics/

Population ethics is an oft-overlooked subfield within ethics. Many people hold views that they don’t realize contradict each other, or that have strange implications that they wouldn’t endorse if they thought about it more.

Not just that—population ethics is a BIG DEAL. A lot of ethical decisions hinge on how you think about changes in future populations.

I’ll give concrete examples below the fold, but reading the examples might change your answers, so you should take the quiz first. It has 13 to 17 short questions. (Questions are skipped when they can’t affect the results.)

Continue reading
Posted on

Don't Treat AI's Skillset as Static

or: The Future Will Be Weirder Than That, Part II

People often treat AI as if it will improve on its current capabilities, but won’t gain any new ones. Supposedly, AI will continue to get better at coding and at assisting with research, but it won’t broadly replace human labor. It won’t be able to do good strategic planning, and certainly won’t know how to operate in the physical realm.

Maybe that’s true. But we don’t have strong reason to believe it. Just looking at the past few years, AI has rapidly acquired new skills. The simplest extrapolation of recent trends is that it will continue to do so.

Some people think of AI capabilities growth like this, with each colored area representing another few years of progress:1

But it’s more likely to look like this:

Continue reading
Posted on

Reasons to Act on Your Best Guess: A Reply to the Challenge of Unawareness

This post is a submission to the Cluelessness Critiques Essay Competition.

Credence: Likely.1

Anthony DiGiovanni’s sequence, The challenge of unawareness for impartial altruist action guidance (henceforth “the sequence”), argues that impartial altruists have no non-arbitrary reason to prefer any decision over any other. The central problem is unawareness: “many possible consequences of our actions haven’t even occurred to us in much detail, if at all.”

I disagree. Impartial altruists should act as if we have precise probabilities for beliefs and precise expected utilities of decisions, even if we (usually) do not explicitly state them. It’s okay to act on heuristics, but those heuristics are in service of maximizing the expected value of a (non-literal) utility function, rather than being “terminal” heuristics.

“Act as if” means: either I give an explicit precise credence; or I don’t, but I operate under the assumption that I could give a precise credence if I spent enough time reasoning through my beliefs.

This post starts by explaining the sequence’s central arguments in my own words. Then it offers five defenses of acting on your best guess.

Continue reading
Posted on

On Democratizing ASI to Preserve Civil Liberties

I continue to believe we should pause frontier AI development. Any discussion of alternative strategies should be thought of as planning for contingencies.

A unifying driver behind many post-alignment risks—catastrophic risks that remain even if we solve the alignment problem—is that by strong default, ASI would end liberal democracy. Liberalism—in which people have individual rights, autonomy, and the ability to choose their own destiny—is an important force protecting human welfare.1 When people are free, we are reasonably good at making our lives better of our own volition.

Many post-alignment risks have a certain flavor. AI-empowered terrorism; coups; permanent dictatorships; concentration of power. Those risks already exist today (and existed 20 years ago), but they’re mitigated by the fact that power is relatively evenly distributed across people. The most powerful person in the world doesn’t have an extraordinary advantage over the 10th-most-powerful person. ASI could change that.

If people still have civil liberties post-ASI, that will only be because the controllers of ASI allow us to have them.2

One way of thinking goes: AI will be extremely powerful. If everyone had their own personal AI, we could each use it to protect our own interests, and things will turn out okay for us. But how do you get there? It’s not going to happen automatically, but it may be possible to set up a gradual process to keep power balanced.

Continue reading
Posted on

Notes on the possibility of moral progress

To achieve the best possible future, we must know what that future looks like. In other words, we need to solve ethics.1

The problem of solving ethics is so large and abstract that it’s difficult to say useful things about. In lieu of any structured analysis, herein lies a collection of thoughts about the problem.

Continue reading
Posted on

← Newer Page 1 of 24