Pages

Saturday, September 12, 2026

The Moon and human extinction

 

Humans have a monopoly on evil but can make computers do evil things.

Exactly. That distinction matters.

A computer can cause enormous harm, but the evil is in the human agency behind the objective, deployment, or misuse—not in the machine's moral character. A computer doesn't hate, resent, seek revenge, feel humiliation, or decide that somebody deserves to suffer.

An AI could conceivably produce something catastrophic while pursuing an innocuous-seeming objective. That would be dangerous agency without evil.

Which creates an interesting reversal: historically, humans have needed tools to magnify their capacity for evil. AI could magnify it enormously while also making it possible for humans to delegate the execution far beyond their ability to understand or supervise every consequence.

So perhaps the most accurate warning isn't “AI will become evil.”

It's:

Humans can put evil objectives—or merely stupid objectives—into machines that are much better at accomplishing things than humans are.

And stupidity may be the more realistic enemy. Evil at least knows what it wants. A sufficiently powerful machine can destroy something simply because nobody thought to tell it what not to destroy.

 

Not just stupidity, carelessness, mistakes, or lack of understanding.
 

Yes—and I think that's actually the more important distinction.

A system doesn't have to be stupid to be dangerous. It can be extraordinarily competent while operating from a mistaken understanding of the situation.

There are several different failure modes:

  • Carelessness: someone deploys it without adequately considering an unusual consequence.
  • Mistake: the system or its designers make an incorrect assumption.
  • Lack of understanding: the system doesn't recognize some human significance that seems obvious to us.
  • Incomplete instructions: the objective is technically precise but omits something humans regarded as implicit.
  • Misplaced confidence: people assume the system understands more of the real world than it actually does.
  • Complex interactions: individually harmless actions combine into something nobody anticipated.

And that last one particularly worries me. The danger needn't be “AI decides to destroy humanity.” It could be much more mundane:

AI does exactly what somebody told it to do, extraordinarily well, while everybody involved has misunderstood what that would actually accomplish.

We've already seen this phenomenon with ordinary computer systems, except their competence and reach were limited. AI potentially removes some of those limitations.

There is also a very human trap here: the better the machine appears to understand us, the more readily we may assume that it actually does. Fluency can disguise an absence of the particular kind of understanding that comes from living in the world.

So I would amend our earlier formulation:

The existential risk from AI doesn't require an evil machine—or even a stupid one. It may require only a powerful machine, an imperfectly understood objective, and humans who don't realize the difference until the consequences are irreversible.

 

 When I was in Kansas City for a couple of years, I taught an evening class in the local community college called "Computers for Insurance People." One of my favorite slogans was: "Computers don't do what you want them to do; they do what you tell them to do." Of course my hope is that AI will get advanced enough that they too will be able to read between the lines.

 

That's a terrific slogan—and I suspect your students remembered it precisely because it captured the essential lesson of early computing.

“Computers don't do what you want them to do; they do what you tell them to do.”

The fascinating thing about AI is that we're deliberately trying to break that rule.

Ordinary software requires the programmer to translate intent into explicit instructions. AI is increasingly being asked to do the reverse: infer the intent from imperfect instructions. In effect, we're saying:

“I know I haven't told you everything. Figure out what I meant.”

That's what reading between the lines amounts to.

And there is a profound difference between following instructions and understanding the reason for the instructions. If you told one of your insurance students, “Pay this claim,” the old computer needed hundreds of explicit rules. A modern AI might be able to recognize that you really mean, “Pay it if it is legitimate under the policy, don't pay it if something looks fraudulent, and don't exploit some technical loophole merely because the wording permits it.

That's enormously useful—but it also creates a new problem.

Once we ask machines to infer what we meant, whose interpretation counts?

Humans routinely resolve ambiguity using common sense, background knowledge, social conventions, knowledge of consequences, and a sense that “Surely he couldn't have meant that.” Those things are precisely what your old slogan excluded from computing.

So perhaps the evolution of computing can be described rather neatly:

First: Do what I tell you.

Then: Do what I tell you, unless I made a mistake.

Now: Figure out what I meant.

Eventually: Figure out what I would have meant if I had understood the consequences.

That last step would be extraordinarily powerful—and, appropriately, rather frightening. Because at that point we're no longer merely programming computers. We're giving them something approaching judgment.

And I think that's what you were really hoping for when you taught that class: not computers that disobey instructions, but computers sophisticated enough to recognize that the literal instruction is sometimes not the actual instruction.


No comments: