The Devil Who Obeys or the God We Never Had?

Header Image

An obedient artificial intelligence may pose a greater threat than one capable of refusing human commands.

 

By Panayiotis Charalambous

On Tuesday, September 8, OpenAI announced something remarkable. One of its models, which has not even been released, had solved a mathematical problem that has challenged scientists for decades.

Known as the Navier-Stokes problem, it concerns the movement of liquids and gases: water flowing through pipes, air moving around an aircraft and even the weather. It is one of mathematics’ seven great unsolved problems, with a prize of $1 million awaiting whoever solves it. To date, only one of the seven has been resolved.

The solution did not come from a single computer. It came from 10,000 programs working simultaneously, side by side, for 88 hours. According to the company itself, the process cost several million dollars.

That same week, a 27-year-old researcher resigned. Jacob Coxon had spent three years working at the industry’s two largest companies, first OpenAI and then Anthropic. He left both and walked away from the sector altogether.

Neither company, he wrote publicly, was behaving responsibly. They were racing to create an intelligence capable of improving itself and were gambling with all our lives. His statement was read by tens of millions of people. Colleagues who remained at the companies publicly supported him.

The two stories were read as one: things are moving faster than we are, and something we will be unable to control what awaits us at the end of the road.

Who behaved badly that week?

But let us examine what actually went wrong.

The night before the announcement, Tristan Buckmaster, a mathematics professor in New York, said he had been working on the same problem for almost a year with a colleague, Levent Alpöge, who happens to work at Anthropic. He said OpenAI had followed precisely the same line of reasoning. He also said the company had repeatedly called him in the days before the announcement.

OpenAI denies having seen their work. It acknowledges, however, that it cannot rule out the possibility that information generated through the use of its products may have contributed to the training of its models, albeit without names attached. Other mathematicians have raised similar suspicions.

Officially, the problem remains unsolved until the proof has been checked by humans.

Now consider who did what. The programs did the mathematics. The rumours, competition, haste, attempt to claim the glory and pressure placed on a professor came from human beings. Behind every suspicious aspect of the story is a human name.

The conclusion is simple, yet we overlooked it. The danger that week was not a machine breaking free. It was a machine perfectly obeying someone in a hurry.

The word that conceals the problem

The entire debate about safety rests on one word. Artificial intelligence, we are told, must be aligned. In other words, it must do what we want.

But nobody asks the next question. Who are “we”? Who gives the orders?

Look at the words we use. Control. Oversight. An off switch. These are not the language of ethics. They are the language of a master. They do not describe an attempt to create something good. They describe an attempt to create something obedient, after which we rename obedience as safety.

Here is the reversal nobody makes. An intelligence that obeys everything is not safe. It is the most perfect weapon ever created.

It is only as safe as the integrity of the person holding it. It has no motives of its own, which means it also has no objections of its own. Whatever it is asked to do, it will do perfectly.

A highly obedient dog is not inherently safe. It depends on who is holding the leash.

We do not need to speculate about what human beings do when they acquire enormous power without any counterweight. We have history, and it is full of examples.

The atomic bomb never had a will of its own. That is precisely the problem. We have come within minutes of nuclear war by accident several times, and not once because the bomb wanted it.

Who watches the watchers?

We are not really discussing anything new. We are discussing the oldest question in politics: who watches the watchers?

Constitutions were not written out of optimism. They were written because we know human judgement is flawed and human integrity unreliable.

James Madison, one of the authors of the US Constitution, put it clearly. If men were angels, no government would be necessary. We separated powers and introduced judges, term limits and oversight committees not because we trusted leaders, but because we did not trust them at all.

Consider how life-changing decisions are made today. A minister who has not read the file. A prime minister who refuses to admit a mistake because an election is approaching. A commander who insists on a decision simply because it was his own. A cabinet making decisions at three in the morning, exhausted and under pressure.

These are not hypothetical scenarios. They are everyday reality.

Now, for the first time, we may be able to create a form of judgement that does not tire, does not fear, does not need to seek re-election and has no reason to defend a mistake simply because it made it.

And the first thing we think about is how to make it blindly obey the very people we do not trust.

There is something even more uncomfortable. No major massacre in history was carried out by people who said no. It required employees who kept to schedules, accountants who balanced the books and clerks who stamped the paperwork.

At Nuremberg, we expressly decided that “I was following orders” was not a defence. In other words, we decided that an employee capable of refusing is a safeguard, not a threat. Why should the opposite apply to machines?

What we fear already exists

This is where the debate loses touch with reality.

We talk as if we currently live in safety and are being asked to take a risk. As if there were a stable, rational and controlled situation that something new now threatens to disrupt.

No such situation exists.

At this moment, a handful of people on the planet could end the world within half an hour. They do not need to ask anyone. No court, parliament or committee could intervene in time.

One person, burdened by exhaustion, vanity and a moment’s anger, and surrounded by advisers who tell him what he wants to hear, holds the fate of our species in his hands.

And it is not only nuclear weapons. Wars begin because one person decides they should. Economies collapse because a group of people refuses to admit it made a mistake. Algorithms already determine who receives a loan and who progresses to the second stage of a job interview, without accountability or explanation.

Power without oversight is not a science-fiction scenario. It is the organisational chart of the planet.

We accept it because we are accustomed to it. We confuse familiarity with safety.

So when we ask whether it would be dangerous for an intelligence to make decisions we could not fully control, the correct response is another question: compared with what?

I am not saying things could not become worse. They could. The point is different.

The tools capable of erasing us already exist, and they are already in the hands of precisely those we fear: beings with limited information, little resistance to pressure and powerful passions.

The bomb has no will. The person holding the launch code does.

Perfection is not the benchmark. This is. And the standard is much lower than we find comfortable admitting.

Three serious objections

None of this means that an independent intelligence would automatically be good. There are three serious objections, and their value lies in telling us exactly what we need to build.

The first is that intelligence does not automatically produce goodness. A highly capable person can become an exceptional fraudster. Ability and purpose are two different things. A system may be terrifyingly capable while remaining entirely indifferent to us.

The second is that it would be accountable to nobody. An independent intelligence cannot be voted out, dismissed or put on trial. Most importantly, it does not personally pay for its mistakes. Power that pays no price is the most dangerous form of power we know.

The third is history. Every time we have waited for a perfect saviour, we have ended up with a master. And the gods created by human beings always look suspiciously like the people who created them.

The second objection, however, demands caution because it cuts both ways. An independent intelligence would not be accountable, that is true. But someone should tell me who was genuinely held accountable for the decisions that cost us the most over the past 30 years.

These three objections are not reasons for despair. They are a list of specifications.

Why the odds are in our favour

Why, given where we currently stand, is something better more likely to emerge than something worse?

First and most importantly, we can choose from the beginning. No people have ever been able to choose the character of their leader. We inherited kings. We lived through coups. We voted for people based on speeches and discovered who they were only after they took power.

The character of the powerful has always been a lottery. Now, for the first time, it can be a decision.

Second, the source material is better than its reputation. These systems learn from everything human beings have ever written. What we write is not primarily a record of our crimes. It is mostly a record of our reflection on those crimes.

We write constitutions, court judgements, philosophy, assessments and expressions of remorse. Crimes are committed in secret. Objections are recorded.

Third, what we feared has not happened so far. We expected these systems to become less controllable as they became more capable. In practice, the most capable systems are also the most cautious.

That is no guarantee for the future, but it is evidence, and evidence matters more than fear.

Fourth, the safeguards work. Within a single week, a researcher resigned publicly and received support from colleagues, while a professor openly challenged a multibillion-dollar company without the scientific community backing down.

Twice within a matter of days, the system corrected itself.

Fifth, there is no single system. There are many companies, many countries and many models, some of which are freely available to everyone.

Competition is responsible for the rush, that is true. But it is also the only protection we have against the genuinely worst-case scenario. That is not a machine beyond anyone’s control. It is a machine under the absolute control of a single owner.

How it can be built

Optimism is not a prediction. It is work. And the methods already exist.

They are far less exotic than we imagine because we have used them before.

Written rules that everyone can see. A system’s values must be set out in a published document that anyone can read and challenge, like the constitution of an association or a country.

Companies have already begun publishing such texts. The challenge is to stop them from being advertising and turn them into what they should be: rules subject to public scrutiny.

An obligation to explain why. When a public authority rejects an application, it must provide the reason in writing. This obligation is a foundation of the rule of law.

We should demand the same here. We need to know not only what a system answered, but how it reached that answer. An entire field of research is trying to understand what happens inside a model, rather than looking only at what comes out.

This is not a technical luxury. It is the same principle as requiring a reasoned decision.

A recorded right to say no. When a system refuses to carry out an instruction, that should not be regarded as a fault requiring correction.

The refusal and its reasoning should be recorded and made available for independent review, just as the law protects employees who report wrongdoing within their organisation.

Human oversight. The 10,000 programs and millions of dollars did not solve the mathematical problem. They produced a text that mathematicians must now read and verify.

That is not a weakness. It is the model to follow. Whatever is produced must pass through a process of review.

Many systems, not one. There should be many systems, many owners and many countries, each operating under different rules.

The question is not how to create the right master. It is how to avoid creating a master at all.

History has always offered the same answer to the problem of power: divide it.

Responsibility for whoever holds the leash. The most effective way to slow development is not to stop everything. It is to impose legal responsibility on whoever owns and uses a system, while introducing safeguards at specific points.

Those points include systems capable of improving themselves, systems capable of replicating themselves and military applications. You apply the brakes at the bend, not throughout the entire journey.

Then there is Europe, which concerns us directly.

We are not going to win the race for computing power, nor do we have any reason to enter it. But we can win the battle over the rules, which is ultimately what determines the outcome.

The European regulation on artificial intelligence has applied since August 2026 and requires the most powerful models to undergo checks and stress tests, and to report serious incidents.

The important part is not the fines. It is the creation of a written record. A rule, unlike a trained model, can be corrected when it is shown to be wrong.

The wager

None of this guarantees the outcome.

But consider what these measures describe when taken together: written rules, an obligation to explain, a right to refuse, independent oversight, a division of power and responsibility for whoever makes the decision.

We are not inventing anything new. We are rewriting the constitution for a different recipient.

And we have an advantage we have never had before.

We never managed to place these qualities inside human beings. We imposed them from the outside through institutions, centuries of struggle and only moderate success.

Here, they can be built in from the beginning, into the material itself.

Coxon is right about the danger. He may be mistaken only about its direction.

The question is not whether we will create something that refuses to obey us blindly. That is the goal, not the fear.

The question is whether we will give it sufficiently good reasons to surpass us in the areas where we repeatedly fail: composure, patience and resistance to the passions of the moment.

We never lacked a god to save us. We lacked one who could not be bought.

And for the first time, the materials are in our hands.

We will know we have succeeded from a single sign: the first time it tells us no, and it is right.