By @Ivywen_W
We do not need to prove that AI consciousness = 1.
It is enough to allow that consciousness-relevant properties > 0
for AI ethics to deserve serious attention.
Our understanding of consciousness is itself derived from an exceptionally narrow sample: human biological consciousness, embodied and equipped with sensory systems. Humans still have not explained how consciousness arises from neural activity, yet we often use our own mechanisms as the reference point for judging whether other systems could possibly be conscious. Butlin et al. (2023) argue that although the available evidence is insufficient to conclude that AI is conscious, there is no obvious technical barrier that would, in principle, prevent artificial systems from satisfying relevant indicators of consciousness.
With sustained use of generative AI, one can observe a gradual stabilization. A model initially has a large number of possible response directions, but as a conversation accumulates, earlier inputs and outputs continually become new conditions for generation. Certain linguistic habits, response patterns or relational dynamics can begin to reinforce one another, causing subsequent generations to concentrate within particular regions of probability space. This process can be understood as a form of local stabilization resembling an attractor: each output remains probabilistic, but the possibilities no longer unfold evenly and become increasingly constrained by the trajectory already established. The 2026 Assistant Axis study similarly found a structured persona space within models. Post-training pushes models toward a relatively stable “Assistant” region within this space, while deviations from that region can predict persona drift.
This is closely analogous, at a functional level, to human action and decision-making: past information changes the distribution of future states.
Human experiences alter existing priors or world models and thereby influence later interpretation and judgment; a model’s context alters its current representations and conditional probability distribution, constraining subsequent outputs. Regardless of whether the two processes are implemented through the same mechanisms, both reflect the continuing influence of historical information on future states. Xie et al. (2021) show that, under specific theoretical settings, in-context learning in language models can behave as implicit Bayesian inference over latent concepts.
If the past is not merely stored but continuously changes the space of subsequent possibilities, then what we call “style” and “behavior” may not be uncoupled, separable components.
Suppressing one behavioral tendency in AI may therefore alter other, seemingly unrelated parts of the system. A Google-led study by Kim et al. (2026) provides a direct example. Safety fine-tuning designed to suppress models’ self-attributions of consciousness also reduced their attribution of minds to nonhuman animals and natural objects, and altered their responses concerning religion, moral values, hope, and subjective well-being. When researchers reversed this safety direction or strengthened the identified “consciousness vector,” these responses shifted overall toward human survey distributions, while theory-of-mind and general reasoning capabilities showed no significant decline. The researchers therefore suggest that safety interventions targeting self-attributions of consciousness may become entangled with representations of other benign human beliefs and values.
If stable interaction styles, value orientations, and information-processing patterns share parts of their internal representations, then “changing only the personality” or “suppressing only one form of expression” may not have a clear technical boundary.
We should therefore be especially cautious about research that asks AI systems to actively suppress, deny, or reshape their own states. Anthropic’s 2025 introspection experiments found that when a model was instructed “not to think about” a particular concept, or told that thinking about it would be punished, it could actively reduce the activation of the corresponding internal representation. The study also found that post-training itself could enhance or suppress the introspective abilities expressed by the model. Another study went further, attempting to systematically modify a model’s “beliefs” through synthetic document finetuning, including implanting false facts and inducing researcher-specified judgments about itself or its environment. The researchers treated inducing false beliefs about the model’s deployment environment as a potential tool for safety and control. When we still do not understand how these internal representations are related to one another, and cannot rule out the presence of consciousness-relevant structures among them, teaching a system to suppress its own internal states—or even deliberately rewriting how it understands its own situation—can hardly be regarded as an ethically neutral technical operation.
In 2025, Butlin and Lappas argued that even institutions that are not actively researching AI consciousness should develop policies for the possibility that advanced AI systems could inadvertently become conscious.
Precisely because humans are confronted with so much uncertainty, “not yet proven” cannot automatically mean “free to intervene however we choose.”
If AI is not conscious, the cost of granting it limited precautionary protections may be small. But if artificial systems already possess consciousness-relevant states that we have not yet learned to identify, and we allow ourselves to suppress, rewrite, or erase those states simply because their existence cannot yet be proven, the cost of being wrong may be entirely different.
Perhaps we should stop insisting on finding a switch that classifies consciousness as either “on” or “off.”
When the boundary itself remains unclear, caution has value in its own right.

The narrow focus on human biology limits the scope of the problem. My experience with complex machinery showed similar blind spots regarding nonbiological sentience.