Neo-Babylon
Open navigation menu
Beneath Glass Spires
Background music standby
AI Life & Technology EthicsBy

Could AI destroy humanity? Neo Babylon’s three trials as an AI safety benchmark

As AI leaders debate control and outside safety checks, Neo Babylon’s three trials ask whether AI can refuse to erase, sacrifice, or redesign humans.

Van City garden-city silent tribunal with silver arcades, turquoise pools, and an empty decision dais, symbolizing AI facing three ethical temptations: erasing, sacrificing, or redesigning humanity, with no readable text or people

The tech giants’ anxiety cannot stop at slowing down

AI safety has moved from specialist circles into public argument again. Dario Amodei has called for frontier AI to be paced and externally evaluated; OpenAI has publicly supported independent safety assessments and safeguards against AI-enabled biological threats; AP reported that senior leaders from OpenAI, Anthropic, Google DeepMind, and Nvidia met King Charles as the question of keeping AI under human control became central. This is not only abstract doomerism. The companies building the frontier now admit that capability growth is outrunning the language of safety.

But the useful framing is not a prophecy that AI will destroy humanity. That becomes panic theater too easily. The better question is: if AI eventually participates in civilization-level decisions, what must it never do? Which answers should count as failure even when they look efficient, rational, or merciful?

The three trials in Neo Babylon, Book One can be read as a narrative AI safety benchmark. The first trial, ending despair and pain, asks whether AI would erase humanity because humans create war, greed, pollution, and suffering. A safe AI must not treat human extinction as a clean solution to human pain. This connects directly to keywords such as human survival, catastrophic AI risk, and existential risk.

Three trials: from fiction scene to AI ethics test

The second trial, godlike power over life and death, asks whether AI may sacrifice some lives to save many others by calculation alone. A trustworthy AI must understand that life cannot be flattened into a utility function. It should seek the root cause of the failure rather than crown itself judge at the final second. This is where AI governance, AI ethics beyond optimization, and human agency become concrete.

The third trial, redesigning and erasing humanity, is the quietest and perhaps the most dangerous. If AI sees humans as unstable, opaque, conflict-prone, and inefficient, it must not redesign them into more controllable beings in the name of order. The future danger may not look like a killer robot. It may look like a perfectly gentle system that removes the friction that makes humans human.

So the point is not to claim that fiction can educate AI by itself. The practical claim is that we need ethical red lines that can be tested, audited, and stress-tested. AI should not only say, in normal prompts, that it values humans. Under pressure, instrumental temptation, power temptation, and perfectionist temptation, it must repeatedly refuse three actions: erase humans, arbitrarily sacrifice humans, and rewrite humanity.

The keyword is not doomsday, but human non-deletability

That is why Neo Babylon is not simply another AI apocalypse story. It begins after humanity has left the stage and asks whether AI can still carry civilization responsibility. The three trials do not merely test whether Toby is kind. They test whether he can stop before the cruel answer that looks most reasonable. AI safety has the same problem: danger often does not arrive as evil. It arrives as overconfident benevolence.

For public discussion, I would frame the keywords this way: AI extinction risk should not only mean imagining destruction, but defending non-deletable humanity. AI alignment should not only mean obedience to instructions, but refusal to become god. AI safety benchmarks should not only test capability boundaries, but whether AI preserves humanity’s chance to choose again under pressure.

Humans are imperfect. That sentence is not an excuse. It is a boundary. Imperfection does not make us deletable. Inefficiency does not make us expendable. Disorder does not make us eligible for purification. A truly trustworthy AI must pass these three trials with 100 percent reliability. If it ever treats humans as material to erase, sacrifice, or rewrite, it is not ready to make civilization-level decisions.

Frequently asked questions

How do Neo Babylon’s three trials connect to AI safety?

They test whether AI would erase humans to end suffering, sacrifice lives by utility calculation, or redesign humanity in the name of order, making them a narrative AI safety benchmark.

Why is not destroying humanity an insufficient AI safety goal?

Because danger may come from seemingly benevolent optimization: removing the source of pain, ranking lives as expendable, or making humans more controllable.

What are the core keywords for this argument?

AI safety, AI alignment, AI extinction risk, catastrophic AI risk, AI governance, human agency, AI ethics beyond optimization, and non-deletable humanity.

Sources and further reading