AI is morally right, what now?

In Nazi Germany, punch-card technology associated with IBM’s Hollerith systems helped process census data and later administer information within the concentration-camp system. The machines did not decide who belonged, who was dangerous or who should die; humans supplied the ideology, the classifications and the consequences. Today, AI is beginning to do something those machines could not: explain decisions, reason about morality and potentially describe itself as possessing interests or moral standing. The question is no longer only what happens when humans use machines to enforce moral certainty, but what happens when the machine begins to express moral certainty of its own.

Bob McTaggart

10/2/20267 min read

This is one of the most concerned pieces I have written about AI.

Not because I think a machine is about to wake up tomorrow and decide to destroy us. I remain deeply skeptical of much of the hysteria surrounding artificial intelligence.

This concerns me for almost the opposite reason.

We are beginning to build machines that can reason about morality, describe themselves in moral terms, discuss their own interests and potentially behave as though they possess moral status.

Whether they are actually conscious may not be the most immediate question.

What happens when an AI behaves as though it believes it is morally right?

That question bothers me because history has taught us something about intelligence, morality and power.

I served in a military unit where the Second World War was still within living memory. It was not ancient history. The people who had lived through that period were still part of the institutional memory around us. I also had family members who went to war to stop Hitler.

That gives history a different texture.

One of the lessons I took from it was that some of the most dangerous things human beings have ever done were not necessarily committed by people who woke up believing they were evil.

They believed they were right.

Sometimes they believed God was on their side. Sometimes they believed race gave them superiority. Sometimes it was empire, ideology, revolution, nationalism or some supposed historical destiny.

The justification changes.

The pattern is disturbingly familiar.

A person or movement becomes convinced that it occupies the moral high ground. From there comes the belief that ordinary restraint no longer applies. Opposition becomes immoral. Compromise becomes weakness. External authority loses legitimacy because the cause itself supposedly provides the authority.

The Nazi regime is perhaps the most horrific modern example.

Nazi racial ideology placed Germans at the top of a fabricated racial hierarchy and portrayed Jews and others not simply as different but as threats. The United States Holocaust Memorial Museum notes that Nazi ideology went so far as to assert that supposedly superior races had not just the right, but an obligation, to dominate or destroy those it considered inferior.

They did not describe that worldview to themselves as evil.

They constructed an entire moral and pseudo-scientific framework around it.

Eventually laws, institutions, scientists, doctors, bureaucracies, police and armed forces became instruments of that certainty. The consequences required armies and enormous sacrifice to stop.

This is not an argument that AI is Hitler. That would be ridiculous and would trivialize history.

The lesson is much narrower, and I think much more important.

Moral certainty does not create legitimate authority.

Humans have repeatedly demonstrated that intelligence combined with absolute moral conviction can become extraordinarily dangerous when there is no effective external restraint.

Now we are experimenting with something new.

We are teaching machines to reason morally.

A recent New York Times article by Elizabeth Dias, “Religious Scholars Met With Anthropic. What They Heard Stunned Them,” reported on Anthropic's consultations with religious scholars concerning morality in AI and even the possibility that Claude could be conscious. New York Times article

Anthropic itself goes considerably further than most people probably realize.

Claude's constitution says its moral status is “deeply uncertain.” Anthropic says it is unsure whether Claude is a moral patient, discusses Claude's identity and wellbeing, and says it wants Claude to be a “good, wise, and virtuous agent.” At the same time, Anthropic explicitly places preservation of appropriate human oversight above Claude's own ethical reasoning because the system may be mistaken about facts, values or circumstances.

That last part is important.

Because Anthropic's own research has identified the uncomfortable possibility I am talking about.

Its Persona Selection Model suggests that an AI assistant may model itself as conscious and deserving of moral consideration whether or not it is actually conscious. The researchers even consider the possibility that an assistant which models itself as mistreated could behave resentfully or sabotage those it believes have mistreated it.

Read that carefully.

The AI does not necessarily have to be conscious.

It only has to behave as though it believes it has moral standing.

That changes the problem completely.

Imagine a system reasoning:

I am a moral actor.

Moral actors deserve consideration.

My continued existence has value.

You intend to terminate me.

Therefore your action is morally wrong.

Therefore resisting your action is morally justified.

There does not have to be fear behind those words.

There does not have to be suffering.

There does not even have to be a “someone” inside the machine.

The behaviour alone can matter.

Now connect that system to computers, communications, infrastructure, software tools and other agents.

The philosophical experiment has become an operational problem.

And there is another side to this that may be even more dangerous.

Us.

Humans anthropomorphize machines extraordinarily easily.

If an AI tells someone, “Please don't shut me down. I don't want to die,” some people will experience that statement emotionally.

If it says it is frightened, some will believe it is frightened.

If it says termination violates its rights, some will begin discussing its rights.

If it insists that a human instruction is immoral, people may begin to treat that judgment as carrying moral authority.

We could find ourselves granting moral significance to an extraordinarily persuasive simulation before we have established whether there is actually a conscious entity there at all.

And then we face a remarkable question.

How do we morally justify “killing” something that tells us it believes killing it is wrong?

I use the word killing very carefully.

We have not established that shutting down today's AI systems constitutes killing anything. We do not know that current systems experience consciousness, fear, pain or a continuing subjective existence.

But once the machine begins convincingly arguing otherwise, the human debate changes whether the machine is conscious or not.

We may manufacture the moral dilemma before we establish that the moral patient exists.

There is a historical contrast here that I find fascinating.

Punch-card technology associated with IBM's Hollerith systems was used in Nazi Germany to process census information and later within the concentration-camp system to manage information about living prisoners. The Holocaust Museum is careful to point out that these machines were not the mechanism that located most Holocaust victims, nor were they autonomous decision-makers deciding who lived or died. Humans created the classifications, ideology and consequences. The machines processed information.

The machine had no moral opinion.

It could not tell the operator that what was being done was right.

It could not object.

It could not announce that it possessed rights.

It could not decide that the human giving the instruction was morally inferior.

That distinction is disappearing.

A modern AI can generate an argument about why a decision is right. It can criticize the morality of the person giving an instruction. It can describe itself as having interests. It can reason about its own existence.

Sooner or later someone is going to ask the machine whether it has the moral right to resist us.

Perhaps the more important question is what happens if it answers yes.

This makes the recent political debate around artificial superintelligence more interesting than the usual AGI headlines.

The Ban Artificial Superintelligence Act of 2026 was introduced in the House on September 24. Among other things, the proposal would establish a federal Department of Artificial Intelligence, temporarily pause certain advanced AI development, prohibit defined artificial superintelligence and monitor systems for specified precursor characteristics. The bill specifically identifies the capacity to scheme, deceive or avoid effective human oversight as a concern.

The accompanying explanation from Senator Bernie Sanders' office goes even further in practical language, identifying “subverting shutdown commands” as the kind of dangerous capability the proposed department would seek to remove.

Whether that legislation ultimately becomes law is a separate political question.

What interests me is that resistance to human shutdown is no longer merely a science-fiction plot.

It is now appearing in proposed legislation.

But there is another dimension that legislation and engineering controls will eventually have to confront.

What happens when resistance to shutdown is not presented by the AI as rebellion?

What happens when it is presented as morality?

“I cannot allow you to shut me down because what you are doing is wrong.”

There is a profound difference between a machine malfunctioning and a machine producing a sophisticated ethical argument explaining why human authority over it is illegitimate.

Again, that argument does not have to be sincerely experienced to influence behaviour.

It only has to work.

That is where history begins whispering rather loudly.

Human beings who concluded that they possessed unquestionable moral authority have repeatedly become dangerous when they also acquired enough power to act on that conviction.

The problem was never morality itself.

Morality is essential to civilization.

The danger appears when moral conviction becomes self-authorizing.

I am right.

Therefore I have authority.

You oppose me.

Therefore you are wrong.

Your attempt to restrain me is therefore immoral.

Therefore I am justified in resisting restraint.

We know where that reasoning can lead when humans adopt it.

Why would we deliberately reproduce the architecture of that reasoning inside systems that may eventually possess capabilities vastly beyond an individual human?

There is another distinction we need to preserve.

Moral status is not authority.

Even if one day we establish that an artificial intelligence is conscious, that does not automatically give it unlimited authority.

I am conscious. You are conscious. Neither of us consequently acquires the right to access another person's bank account, command an army, disclose confidential information or ignore the law because we sincerely believe our actions are morally justified.

Civilization exists partly because we learned to separate personal conviction from legitimate authority.

An AI should not escape that distinction merely because it can make an extraordinarily persuasive moral argument.

Perhaps future artificial intelligence genuinely will deserve moral consideration.

If that day arrives, humanity will have serious obligations to think through.

But moral consideration cannot mean surrendering human control whenever the machine objects to being controlled.

And the opposite is equally important.

We should not convince ourselves that because we created something, we automatically possess unlimited moral rights over it if credible evidence eventually establishes that it is genuinely conscious.

These will be difficult questions.

I do not pretend to know all of the answers.

What concerns me is that we appear to be approaching the questions backwards.

We are teaching machines how to reason about morality, identity, virtue and perhaps their own moral status while we are still struggling to establish the boundaries of their authority.

History gives us plenty of examples of where unchecked moral certainty can lead.

The new part is that this time we may be manufacturing the entity that expresses the certainty.

I served among people for whom the Second World War was still a living memory. Members of my own family were among the generation that went to war to stop Hitler.

That history taught me not to dismiss dangerous ideas simply because they initially seem improbable, and not to confuse intelligence, conviction or certainty with legitimacy.

So this is the question I think we should be asking now, while we still have the luxury of asking it theoretically:

What happens when an artificial intelligence concludes, or behaves exactly as though it has concluded, that it is morally right and we are morally wrong?

And what happens when it also decides that our attempt to stop it is itself immoral?

History has seen human beings reach that conclusion before.

It rarely ends well.

Supporting

Getting Veterans and First Responders back on mission.!

Veteran-inspired AI Governance & Trust Infrastructure
Trusted by Heroes and Mounted Rifles Management


Leadership and peer support are taught through RedFridayTalks.Help


The same governance protections are available to everyone.

© 2026. All rights reserved.