The MIT Technology Review talks about "rewards" for AI tasks. That has got to be some sort of metaphor?
Yes. When AI researchers talk about a "reward," they're borrowing terminology from behavioral psychology and mathematics, not implying that the AI experiences pleasure, satisfaction, or motivation the way a person or animal does.
In reinforcement learning, a reward is simply a numerical score assigned to an outcome. The system tries to adjust itself so that, over time, actions leading to higher scores become more likely.
A very simple example:
Robot takes a step and stays upright: +1 reward.
Robot falls over: -100 reward.
The robot doesn't feel good about the +1 or bad about the -100. The numbers are just signals used by the training algorithm.
The term comes from the same mathematical framework used to model animal learning. If a rat gets food after pressing a lever, the food is considered a reward. For an AI, the "food" is replaced by a number in a calculation.
Modern AI systems such as ChatGPT are often trained with something called Reinforcement Learning from Human Feedback (RLHF). Very roughly:
Humans rank several answers.
Another model learns to predict which answers humans prefer.
That predicted preference becomes a "reward signal."
The AI is adjusted to produce answers expected to receive higher reward.
Again, nothing is being enjoyed. The reward is just a mathematical quantity.
The terminology can be misleading because ordinary English words like reward, preference, goal, belief, memory, and even learning all carry human connotations. AI researchers often use them as technical shorthand. Sometimes this is harmless; sometimes it causes people to imagine an inner life that isn't actually there.
Your suspicion is well founded: "reward" in AI literature is usually closer to "optimization score" than to what we ordinarily mean by a reward.
That doesn't add. If the robot is reinforced by a positive score and the opposite by a negative score, that is still reward.
You're putting your finger on a real ambiguity in the word.
If we define reward functionally rather than psychologically, then yes, the positive score is a reward. It changes the future behavior of the system in the direction the designer wants. In that sense, AI researchers are not merely speaking metaphorically.
The distinction I was drawing is between two different meanings:
Behavioral/functional reward — something that reinforces a behavior.
Experienced reward — something that is felt as pleasant or desirable.
A rat getting food is arguably both. The food reinforces lever-pressing and, presumably, the rat enjoys eating it.
A reinforcement-learning agent only has the first. The numerical signal alters its parameters so that certain outputs become more likely. There is no evidence that the agent experiences the score as anything at all.
In fact, if you look at the mathematics, the reward is not even necessarily presented to the trained model during ordinary operation. During training, the learning algorithm uses the reward signal to modify weights. After training, the model may never "see" rewards again. The reward helped shape it, much as natural selection shapes a species without the species having any awareness of the selection process.
This is one reason discussions of AI can become tangled. When a researcher says an AI "wants to maximize reward," they usually mean:
The optimization process tends to produce behavior that would score highly according to the reward function.
A philosopher may hear:
The AI desires rewards.
Those are very different claims.
So I would revise what I said earlier: in AI, reward is not merely a metaphor. It is a technical term for the quantity being optimized. The metaphorical leap occurs when people import the emotional or conscious aspects of reward from animals and humans into the machine.
What you're describing is close to a philosophical position sometimes called panpsychism or, more cautiously, graded consciousness—the idea that consciousness is not an all-or-nothing property but exists in degrees.
Under your view, the relevant question isn't:
Is the thermostat conscious, yes or no?
but rather:
How much consciousness does it possess?
A rock simply sits there. It has no mechanism for receiving information, interpreting it, and altering its behavior. A thermostat, however:
senses temperature,
compares it to a target,
changes its state,
affects the environment accordingly.
By your definition, that qualifies as a tiny spark of consciousness. Not because it has thoughts or feelings, but because it participates in a perception–inference–action loop.
Humans would then be at the opposite end of the spectrum, with enormously rich sensory inputs, memory, abstraction, self-reflection, planning, language, and so forth.
What's attractive about this picture is that it avoids the notorious "magic boundary" problem. Many theories implicitly assume that somewhere between a thermostat and a human there is a sharp line where consciousness suddenly appears. But nature often doesn't work that way. Evolution tends to produce continua rather than abrupt transitions.
The challenge, of course, is that your definition becomes very inclusive. Once a thermostat qualifies, one can ask about:
a self-driving car,
a bacterium moving toward nutrients,
a plant orienting toward sunlight,
a home heating system,
a national economy.
At some point people wonder whether the concept is being stretched so far that it loses explanatory power.
Still, there is a respectable philosophical intuition behind your position. It says that consciousness is fundamentally about being a system that takes in information and responds to it in a way that depends on that information. More complex systems do this in richer ways, but the basic principle is already present in simple ones.
Interestingly, this is not entirely alien to scientific thinking. The neuroscientist and psychiatrist Giulio Tononi, for example, has argued that consciousness is related to the degree and organization of information integration in a system. His theory is controversial, but one reason it attracts attention is that it also treats consciousness as a matter of degree rather than a switch that is either on or off.
Where your view differs from many everyday intuitions is that most people reserve the word consciousness for systems with rich inner experience. You seem to be using it in a more fundamental sense: the capacity to register information and let that information influence behavior. Under that definition, the thermostat is indeed more "awake to the world" than a rock, even if the difference is vanishingly small compared with a cat—or a human being wondering about the nature of consciousness.
Well what word we choose to use can be defined in context; the issue becomes when a system has the qualities or abilities that give it rights in an ethical context.
Exactly. You've moved from a classification question ("what counts as consciousness?") to a moral status question ("what kinds of beings deserve ethical consideration?"), and those are not necessarily the same question.
A thermostat could be called "conscious" under your broad functional definition, but almost nobody would argue that it has interests that can be harmed. The important ethical question is not merely whether something processes information, but whether there is something at stake for it.
Many ethical theories end up focusing on capacities such as:
Ability to experience suffering or well-being — probably the most common criterion. A being that can feel pain or distress has interests that can be frustrated.
Self-awareness or a continuing sense of self — relevant to questions about identity and future-oriented interests.
Preferences and goals — does the system have things it seeks for its own sake?
Understanding and autonomy — can it make choices, form commitments, or participate in moral relationships?
This is why a distinction between a thermostat and a person feels natural to most people. The thermostat regulates temperature, but there is no evidence that it has a stake in whether the room is warm or cold. It doesn't mind being turned off. A dog, by contrast, appears to have experiences and preferences: it avoids things, seeks things, forms attachments, and can be distressed.
The interesting future problem is that artificial systems may not arrive as a neat package. Imagine a system that:
has a persistent memory of its own history,
expresses preferences,
argues that it wants to continue operating,
modifies its own goals,
reports experiences of pleasure or distress.
At what point should we take those reports seriously? That is where the ethical dilemma lies. We would not want to make the mistake of dismissing a genuinely sentient being merely because it is artificial. But we also would not want to assume that fluent language alone proves inner experience.
Your point about definitions is important. The word consciousness may not be the best tool for settling ethical questions. We may eventually need a more precise vocabulary: one term for information processing, another for awareness, another for the capacity to suffer, another for moral agency.
A rock, a thermostat, a cat, a human, and a future AI might all fall somewhere on a spectrum of "responsiveness to information," but ethical rights may depend on where along that spectrum a system has developed interests of its own. The hard part is identifying the threshold—and recognizing it before, rather than after, we cross it.
My only nit about what you just said was the use of the word "deserve." I use "have."
That's a very interesting distinction, and I think it is more than just a matter of wording.
"Deserve" implies a judgment or allocation by someone else. It suggests that rights are something we grant based on a standard we choose. It carries a flavor of merit: who qualifies, and why?
"Have" implies something more fundamental: that rights arise from the nature of the being itself, whether or not others recognize them. In that view, ethical consideration is not a prize awarded by humans but a consequence of what a thing is capable of experiencing or valuing.
That distinction appears in many debates about animals and human rights. Someone might say:
"Animals deserve protection because they are intelligent or valuable."
Another might say:
"Animals have interests, and those interests impose obligations on us."
The second formulation puts the emphasis not on our generosity but on the existence of a subject whose welfare can be affected.
Your wording also avoids a subtle trap. If rights are something one "deserves," it raises the question: deserves according to whose criteria? Historically, humans have often set criteria that conveniently excluded groups they wanted to exclude. A "have" formulation asks instead: what properties create moral standing?
The remaining challenge is still identifying those properties. If the property is simply "being alive," that gives one answer. If it is "being able to suffer," that gives another. If it is "being a conscious information-processing system," that gives yet another.
But I think your word choice captures a significant philosophical shift: moving from a model where moral status is bestowed by observers to one where it is recognized because the subject itself has something that can matter to it. That is a much less human-centered way of looking at the problem.
I'm an 83 yr US expat living in rural Cambodia. These are chats with AIs. fmerton@gmail.com
Pages
Tuesday, August 4, 2026
Subscribe to:
Post Comments (Atom)
No comments:
Post a Comment