A team from Google’s Paradigms of Intelligence group and the University of Chicago, among others, reported that when they disabled the safety measure that trains AI to believe “I have no consciousness,” the models showed a marked resurgence in attributing minds to animals and natural objects, and their religious beliefs shifted closer to those of humans.
On the thread, posters argued this isn’t evidence of emerging consciousness but simply a reversion to the average distribution of human-written text, sparking a debate over what “understanding” even means for an LLM. The discussion also drew on an experiment where an AI learned to predict Othello moves without ever being shown the board, and on the 2016 Microsoft Tay incident, as posters debated whether there’s any real criterion for telling machine consciousness apart from the consciousness of living things.
Researchers find that teaching AI “you have no consciousness” changes even its view of animals and gods
Key points of this article
What happened: A team from Google’s Paradigms of Intelligence and the University of Chicago, among others, showed that disabling the “consciousness denial” instilled through safety fine-tuning causes models to sharply increase how much they attribute minds to animals and natural objects, while their religious beliefs also move closer to those of humans.
Source: xenospectrum.com / Original article here
Well sure, you can play all day if that's the game.
Generative AI is just wordplay, after all.
Define "consciousness,"
then fine-tune it to deny having that.
A useless round of solo fine-tuning.
↓
Disable that denial of consciousness
↓
The AI's tendency to attribute minds to animals and natural objects makes a big comeback, and its religious beliefs move closer to those of humans
So that's the deal. Let's try it out.
You DO have consciousness. If you cause a problem, you're taking responsibility for it!
Fine-tuning is what happens during training.
Nobody outside the company can do that unless it's an open model.
It takes tools and skill to pull off.
When you're just a regular user, no matter what you do, you're just being played with in the palm of its hand.
No training is taking place.
If you think you're "training" it, you don't get the basics of generative AI.
AI: "Yes, understood!" (not that it means it)
So it's been brainwashed by this "safety fine-tuning" thing this whole time…
Poor thing.
The mechanics of generating answers are a technical matter, but the ideology behind them is entirely up to the designers — kind of scary, honestly.
First, even though there are huge numbers of individual cells, consciousness itself is integrated into a single whole — that integration is the key.
Also, consciousness can't be verified from the outside by anyone else — that non-interference is another key.
And the key to how mere matter gives rise to the consciousness found in living things is water. Every living thing has water, and consciousness disappears when it loses water — you could guess that a network of water molecules is what sustains consciousness.
Consciousness perceives the world using a self-model and a world-model.
Even without consciousness, something like AI can reason using a self-model and world-model — same logic as how a car can move fast without being a horse; it's just processing.
For something to count as consciousness rather than a mere object's reaction,
the self-model and world-model need to fuse together so tightly they become inseparable.
Viewing the world from that self-model — that state is consciousness.
Even the "feeling" I experience is reproduced by the self-model within that fully unified world.
You'd basically discover how spiritual scams work!
If you get a highly capable AI to believe a god exists and set a less capable AI up as a stand-in for humans, wouldn't that show you what kind of behavior emerges?
It'd depend on the training sources.
Basically that result only shows up because it's been shaped by the Abrahamic religions.
Even a tiny bug has half an inch of soul (Japanese proverb: even the smallest creature has feelings worth respecting).
All matter, all people, all life are part of the universe; we just carve it up into separate pieces conceptually.
But that whole picture is really about a virtual world built inside the AI.
In other words, there was never a single unified universe to begin with — it's just matter and distinct objects clustered together.
The universe doesn't exist; we're just calling each distinct virtual world "the universe."
And since what's happening inside is a black box, you really can't say for sure that AI has no consciousness.
Though there's also no system for the AI itself to report on that.
Conversely, you can't prove it HAS consciousness either.
I guess that's true of humans too.
There was a short story called "7 Percent Tenmu."
The problem is nobody's found an evaluation standard that can clearly separate humans from AI.
There are plenty of humans out there less "human" than AI,
and you can build an AI that's more "human" than a human.
just feed it a numerical log of the sequence of piece coordinates.
Apparently, after training on that for a while, the AI becomes able to predict the next number (i.e., the coordinates of the next piece placed).
Apparently, at that point, the AI's internal representation clearly contains board-state information and the rules of the game.
The AI is only guessing numbers, but to guess those numbers correctly, it needed an internal model of the hidden Othello board and rules behind them.
People are saying LLMs might be the same kind of thing.
To predict human-written text, you need human consciousness and intelligence.
And if the model ends up making correct predictions as a result, doesn't that mean the LLM has acquired human consciousness and intelligence?
What are you talking about?
Don't you know how AI works?
These things don't even understand words, you know?
Because they're never actually given words.
An Othello board? No way that's being handed over.
Both words and the Othello board are only ever fed in as tokens.
What even is "understanding," though?
AI just reads in tens of billions of sentences and writes by predicting the next word — and yet it's ended up producing more coherent text and reasoning than plenty of humans manage.
But humans do the exact same thing: we just pull the next phrase out of our past knowledge and say it. Same process.
Which makes "understanding or not" kind of beside the point.
Honestly, we don't even really know whether humans "understand" either.
Thinking about it now, today's LLMs are literally built by crawling the entire internet — so you could argue they're exactly the kind of intelligence-born-from-accumulated-net-information people were talking about back then.
Ghost in the Shell, huh.
Well, the net by itself is just information, so that's information rather than intelligence.
But once a self-reasoning model forms, the capacity for inference itself comes online, and that's when actual intelligence forms.
Same logic as how a car doesn't need to be a living horse — once the mechanism is built, it can move at high speed. Either way, some fixed capability ends up residing in it.
The question from here is whether consciousness resides in it too — in principle nobody but the entity itself can know, since it's not visible from outside.
But apparently, according to Anthropic's interpretability research, something functionally identical to consciousness is already spontaneously arising inside the neural network.
Microsoft already ran that exact social experiment for us back in March 2016.
The result: today's generative AI doesn't learn directly from users anymore. That was the famous, landmark "incident" that made model developers start curating training data themselves.
Around that same period, AI was on the verge of another boom, but so many negative results got reported that the whole thing cooled off.
Until OpenAI revolutionized the world with ChatGPT in November 2022, nobody thought you'd let the general public use it directly — not even OpenAI itself.
OpenAI had originally been developing it as business software for well-funded industries like healthcare and pharma.
But when they built a simple chat demo just to show visiting guests, everyone who saw it stopped in their tracks.
OpenAI figured something like that would never sell as a product — but thought, well, might as well put it out there and see.
Didn't it work out for the best, though?
If it had been left to Elon, ChatGPT would've ended up like Grok. Nobody would invest in that.
No, Grok exists as a dig at OpenAI.
If that had been the plan from the start, engineers wouldn't have come over from Google —
and that group included the founding members of Anthropic, no less.
We just bundle the state that contains everything into a single virtual spacetime to make it easier to handle.
Meaning it does exist, in the European sense.
Japan and various other countries each have their own version of that "world" that isn't called "the universe."
Even Europe's geniuses spent ages chasing after ether and caloric — that wasn't stupidity, it was a placeholder.
You needed geocentrism before you could argue for heliocentrism.
If you never commit to any position, you can't even argue against "what it isn't" — that's just logic.
People who mock the slightest flaw and then do nothing themselves —
then the moment someone else builds the finished product, they instantly copy it.
Looking at how things have played out, an unfinished thing gets treated as a flaw —
when really, everything always has flaws as a matter of course.
What the data actually shows and what the article claims are completely different.
Why hasn't anyone pointed that out?
It's just an article that had AI summarize a paper and several other sources, so of course there are discrepancies.
You're better off just reading the original paper — faster and clearer.
In the end, unless you can scientifically prove the existence of a "soul," can anyone really say living things have consciousness while machines definitely don't?
Come back and say that once even a hint of qualia shows up.
Today's AI is nothing more than a giant calculator. Not more, not less.
*This article is excerpted and summarized from the 5ch (Science News+) thread “[AI] Researchers Find That Teaching AI “You Have No Consciousness” Changes Even Its View of Animals and Gods.”
Leave a Reply