Chapter 24: The Conscious Robot

Part Four · Artificial Intelligence


“What is the world, such that we can understand it—and who are we, such that we can understand the world?”

— Brian Cantwell Smith, The Promise of Artificial Intelligence (2019), 2

Chapter Overview

When should we say that a machine understands its world? I separate five questions: what its states mean, what it can do, what it understands, whether it is conscious, and whether it has interests of its own. A machine may master one skill long before sight, memory, learning, and action work together in one lasting subject. Robot B gives the positive case: it tracks a world, learns from error, acts for reasons, keeps itself going, and links what it sees with memory and control. Those facts support understanding and autonomy, but consciousness calls for more evidence about how sensory content recurs, gains attention, and guides the whole system. Several leading theories treat B’s design as such evidence, and I argue that a positive judgment has better support than suspension.

The last four chapters showed that formal skill and fluent behavior do not settle whether a system understands. They did not show that no machine could.

An artificial flower counterfeits: petals without growth, show without bloom. An artificial heart runs the other way. It may use titanium where nature uses tissue, but it moves blood and earns the name by doing the work. Artificial understanding could resemble the heart rather than the flower—realized in something made instead of grown. The analogy clears one obstacle: manufacture does not make an achievement counterfeit. It does not establish that consciousness admits the same transparent functional test as pumping blood. The question remains what work earns the name.

Chapter 6 supplies the hinge. It did not say that every content-bearing state is conscious. It said that when a perceptual state is conscious, its phenomenal character consists in its total representational content. The robot case therefore carries three burdens that must not be fused. A world-involving producer-consumer history must support original perceptual content. Those states must belong to one persisting subject. They must also count as conscious. Meet all three burdens, and nothing further must be added merely because the system was built rather than born.

The test needs a case. What happens when a machine gets something wrong, finds out, and changes what it does next?

24.1 Two Robots: Matched Performance, Different Histories

Ten years into factory work, Robot B notices a silver sheen beneath a coolant joint. The image admits several readings: harmless condensation, leaked lubricant, or the first trace of a cracked line. B compares the location with a small pressure loss upstream and a vibration that began twenty minutes earlier. It first predicts condensation. A thermal reading conflicts with that diagnosis. B revises the classification, slows the pump, reroutes the line, and marks two neighboring joints for inspection.

This is an imagined machine, not a report from a robotics laboratory. Set Robot A beside it with the same cameras, grippers, fluent patter, and ordinary repair performance. Give both robots equal access to the joint, the pressure and thermal readings, the maintenance records, and repeated opportunities to learn. Neither gets the answer whispered by a technician. The comparison asks what each does with that opportunity.

Robot A instantiates a familiar architecture. Its cameras feed a network trained on captioned images. Its grippers follow policies tuned against a reward its engineers chose. It avoids low charge because training left it with a policy that detects and responds to low charge. None of that ancestry decides its semantic status. To make A a derived-content control, stipulate something further. The interpreting practice exhausts the correctness standard of the states compared here. Its history and use have installed that standard without fixing accuracy beyond it. A’s states can then guide action and misrepresent under real, derived conditions. This describes one possible dependence, not a finding about every caption-trained robot.1

B retains the correction. Months later, a weaker version of the same pattern leads it to inspect a seal before any pressure alarm fires. The first encounter has changed how the next one goes.

B’s semantic case differs before self-maintenance enters. Its continuing encounters, prediction errors, corrections, and later use supply a world-involving learning history. That history can fix what its perceptual states concern in a way not exhausted by captions or an assigned semantic task. Those states therefore become candidates for original content. This conclusion remains defeasible. A rival mapping may explain the same history equally well, or the apparent world-answerability may merely implement an assigned standard. Engineering alone decides neither possibility.

Assume A can run the same diagnosis and repair sequence through an effective installed policy. That matches the observed repair, not the complete history or everything either machine would do under changed conditions. Now keep the training label and reward fixed while changing the leak; then keep the leak fixed while changing the label. Follow any correction into later encounters and different uses. This extends Chapter 15’s Map test: the pump might keep running because a backup takes over, while a mistaken diagnosis still needs correcting.

A receives the same tests as B. If A acquires the same world-answerable history and use, it earns the same semantic consideration. Captioned training cannot permanently bar that result, and ten years on a factory floor cannot guarantee it. A remains the derived-content control only while its relevant standard remains exhausted by the borrowed practice. The comparison illustrates distinct possible organizations; it does not infer their difference from equal performance.

B’s success in tracking the world also bears on the continued organization of the system doing the tracking. Its own activity acquires and allocates energy, corrects damage, and preserves the boundary within which its machinery works. It changes strategy when those processes drift toward failure. The robot has something riding on getting the world right because getting it wrong can defeat the activity by which it remains this organized thing.

If every technician vanished, neither robot’s history would evaporate. Under the control stipulation, A would retain its causal roles and derived content for as long as the relevant organization continued to operate. B would retain the tracking, correction, and use relations acquired through its own history. Those relations make original content a live attribution. Removing an interpreter does not itself create semantic independence, any more than inheriting a task rules it out. B would also retain a reciprocal maintenance loop in which regulatory success helped preserve the organization producing the regulation. That further fact concerns autonomy and interests. It does not change what an internal state represents.

The distinction matters because an error can be classified in more than one way. Robot A can misclassify a wrench as a hammer. In the stipulated control, its state then fails a derived correctness condition. Robot B can make the same representational error under a standard that its own learning and world-involving use may have fixed beyond the assigned task. If the error also prevents B from repairing a failing pump, it may undermine the system whose activity is at issue. Semantic status, practical consequence, and felt harm remain different questions. Calling all three “meaning” concealed the differences.

Alan Turing made conversational performance a public test in 1950. His imitation game began with a man, a woman, and a hidden judge communicating by text; he then asked what would happen if a computer replaced the man. Turing refused to make certainty about another thinker’s feelings the price of attributing thought. That demand would push us toward solipsism. Responsive conversation could count as evidence without first solving the mysteries of consciousness.2

Ask A and B to explain the leak, and both might produce the same transcript. A judge reading it could learn much about their competence while seeing little of how the diagnosis formed or whether the correction survives. The next failed seal, and what each robot remembers of this one, can tell us something the transcript does not.

Performance can raise the probability that content-guided competence or understanding explains what the system does. Its evidential weight depends on the training history, architecture, system boundary, available interventions, and rival explanations of the same behavior.

The intentional stance organizes that evidence. Attributing beliefs and goals earns literal rather than merely theatrical standing when it yields compact, durable predictions across novel cases, counterfactual changes, and interventions. It must also outperform a rival account that treats the performance as cue-matching or a fixed policy. A pattern can be real without becoming a person. The attribution must therefore name its level: an embedding mechanism may carry a predictive pattern, a controller may pursue a goal, and a persisting whole may integrate contents across learning and action. Success at one level cannot be promoted to the next by grammar.

Cue-triggered anthropomorphism differs from intentional-stance success. The first records how readily a performance makes us feel an interlocutor; the second tests whether an intentional description continues to explain what the system does when familiar cues, contexts, and rewards change. Novel generalization, content-sensitive error, correction after intervention, and stable counterfactual control add evidence. A warm sentence adds none beyond the capacities that produced it.

The Turing Test did not fail. We promoted it. A transcript now gets asked to rule on content, understanding, consciousness, and mind together. Once the rest of the evidence comes into view, the transcript joins the case. It does not decide it.

Dennett’s competence without comprehension names a possibility, not a verdict on every artificial architecture. Real-pattern success can establish more than pretense while leaving open whether content-bearing mechanisms have become integrated enough for their contents to explain what a whole system learns, expects, revises, and does.3

24.2 Five Questions About a Machine Mind

The first question concerns content: what makes B’s initial diagnosis mistaken? The leak and the history by which B learned to distinguish it from condensation constrain the answer. Selection can make a mapping causally effective without settling its independence. Fintan Mallory’s analysis of neural word embeddings supplies the earlier disputed case. Backpropagation adjusts producer and consumer parameters as their downstream use changes prediction loss. The learned states may genuinely answer to distributions in a linguistic environment, which belongs to the world too. Originality at that limited grain remains unsettled; inherited vocabulary and task design do not establish a merely derived standard.4

Artificial content can count as original where a system’s own world-involving history fixes what its states get right. Here original names a natural independence relation: the mechanism answers to a correctness standard fixed by its history and use in a way not exhausted by an interpreting practice. The independence may vary by content and grain, and it requires no semantic self-creation. Searle denies that such relations establish intrinsic intentionality. Dennett doubts that original and derived content mark a sharp divide at all. The chapter rejects Searle’s intrinsic spark, agrees with Dennett that intentionality emerges gradually, and retains the distinction as a diagnostic of semantic dependence rather than a gatekeeper of genuine meaning.

That semantic finding still leaves a subject to identify. The question now concerns where the correction travels: does it revise one persisting system’s expectations and plans, or remain confined to a detector? Identifying that system does not create the content, and neither finding makes the state conscious.

Causal provenance still matters. A newly manufactured robot can inherit selected functions when its mechanisms descend through the relevant design and training lineage. An accidental structural duplicate cannot inherit a history it never had. Swampman may begin with useful dispositions and a self-maintaining organization but no determinate representational past. Consequential traffic with the world can begin one. His self-maintenance does not fill the semantic blank; his later history can.

The second question concerns competence and attribution. B isolated the faulty line. A component could classify that leak correctly while the larger assembly neither believed nor understood anything about it. Attribution to the whole requires evidence that the diagnosis guides a coordinated response rather than one isolated output. Nicholas Shea’s unity of purpose belongs here: do the local achievements add up to one agentive organization?

A notebook, prosthesis, neural circuit, or language model can become a cognitive organ when a person or distributed organization reliably maintains, corrects, and uses it within one flexible economy. The live question asks which system the integrated achievement belongs to.

The third question concerns understanding. B’s pressure reading corrects its visual diagnosis, and the correction survives into a later inspection. Several skills constrain one another. Their breadth and durability determine how far the understanding extends; Section 24.5 follows that claim without requiring the mature human package of every understander.

The fourth question concerns consciousness. Chapter 6 separates two issues. Strong representationalism says what a conscious perceptual state is like: its phenomenal character consists in its complete world-presenting content under the appropriate perceptual organization. State consciousness concerns how that content works within an identified subject. On the house proposal, it must recur and remain selectively available. Under suitable conditions, it can guide attention, working memory, inference, planning, report where possible, and flexible action. Feedback from those activities can preserve, revise, or displace it. No single output is necessary. Global broadcast, local recurrence, higher-order organization, and predictive control propose ways to realize aspects of this distributed profile. None follows from self-maintenance. A bacterium may have interests without a perceptual point of view. An artificial system may organize perceptual content without reproducing metabolism. Whether autonomy proves necessary for consciousness remains open. This book has not earned that further premise.

The fifth question concerns autonomy and interests. Chapter 15’s physical candidate remains thermodynamic self-maintenance coupled to adaptive regulation.5 Bare persistence comes too cheaply: a candle draws fuel and repairs its shape without detecting departures and selecting corrective responses. A bacterium couples detection to correction in a loop on which its continued organization partly depends. Ezequiel Di Paolo calls the resulting achievement a self-sustaining identity under precarious conditions. The phrase earns its keep. Such a system does not merely have a task; its activity helps sustain the organized individual performing it.

The boundary follows reciprocal organization. A process belongs when it helps maintain the organization. Its continued realization must also depend on that organization at the relevant time scale. Wall power may remain outside an organism-like system. A remote repair service might fall inside a distributed artificial agent if reciprocal traffic made each side part of one self-maintaining organization. Nested cases may resist sharp edges. The test remains causal.

Peter Godfrey-Smith objects that tying mind to metabolic maintenance mislocates the explanatory work.6 On the semantic question, the objection wins. Metabolism does not determine what a state concerns, and the substrate-general idea of reciprocal maintenance does not determine it either. History, tracking, and consumer use do that work. The concession leaves autonomy intact. Self-maintenance can still explain why a system counts as a persisting individual and why some conditions support or undermine it. It may also explain why regulation serves an interest of the system rather than merely satisfying an assigned score.

This is the relevant interest. It supplies no truth condition by itself. A frog’s visual state may concern prey, flies, or flying nutritive objects. The history and consumer relations that best explain its use determine the grain; the fact that the frog can die does not. Once the content is fixed, however, successful tracking can help the frog eat and failed tracking can leave it hungry. Content explains what the state says; interest helps explain why getting it right matters to this organized life.

Patrick Butlin and colleagues ask why agency serving self-maintenance should differ from agency directed at an external goal.7 It need not differ semantically. A demanding external objective can shape content-bearing mechanisms and unify planning. It can also generate flexible subgoals and make failure costly within the task. Self-maintenance adds a different relation. Successful regulation helps preserve the organized process producing it, while serious failure degrades or dissolves that process. In Robot B, regulator and maintained beneficiary coincide. That may ground autonomy and interests. It neither upgrades A’s stipulated derived content nor mints B’s original content. The semantic difference concerns what fixes accuracy, not whether the robot sustains itself.

Interests do not yet amount to felt welfare. A coolant leak can undermine B’s continued organization before we know whether anything feels bad for B. Welfare in the morally serious sense requires a subject for whom conditions can feel good or bad. Conscious valence supplies that further relation. Section 24.5 returns to it after the consciousness case.

One leak can therefore involve a mistaken diagnosis, a capable repair, a lasting lesson, a conscious presentation, and harm to an autonomous individual. None supplies the others merely by sharing the occasion.

24.3 What the Architecture Must Realize

The separation rules out shortcuts from hardware labels to mental verdicts. Embodiment can add tracking relations and let perception guide memory, correction, planning, and action; that evidence may support content and understanding. It guarantees neither autonomy nor consciousness. Nor must embodiment mean a metal chassis on a factory floor. A persistent agent acting through a virtual body can meet a causally corrective environment and develop world-directed content if what happens there genuinely constrains its later perception, learning, and action. Copying raises questions about identity and succession, not meaning.8 Mortality supplies no semantic privilege.9 A trained mechanism keeps whatever content its history and use fix whether or not engineers can duplicate it.

Digitality carries no verdict either. Every implementation runs physically. Its causal-computational organization may hold independently of an observer’s preferred description. Mortal computation, intrinsic computational functionalism, and causal-flow constraints all point to one discipline. Name the organization claimed, then test whether the physical system realizes it without building the desired result into implementation.10 A represented hunger variable can carry content while an independent platform keeps the process running. Another implementation may use a state concerning low energy to acquire energy, correct damage, and preserve the organization doing the regulation. The loop supplies evidence of autonomy. The medium does not.

Return to the coolant joint and alter one relation at a time. First corrupt the pressure reading while preserving the leak, thermal evidence, reward, and maintenance arrangements. Does either robot cross-check and retain a correction? That bears on the represented condition and its use. Next restore the reading but move B’s energy and repair work to an external service, preserving its perception, memory, planning, and recurrent access. The loss now concerns reciprocal self-maintenance; it need not alter what B represents or whether it experiences anything. Combining both changes would obscure which one explained a changed verdict.

Redraw A to include the technician, scheduler, and data center if the evidence warrants it. Then follow persistent correction, unity of control, and reciprocal dependencies across the wider boundary. A commercial relation or enabling service does not by itself create one subject or one autonomous individual.

These distinctions tell us what to test without supplying an engineering recipe. Content-bearing mechanisms can support competence; several capacities that correct one another can support partial understanding, and wider traffic can deepen it. Autonomy requires an organization whose activity helps preserve itself. Consciousness requires separate evidence that perceptual content acquires recurrent, selective, two-way influence across a persisting subject’s integrated economy. If that evidence supports attribution, Chapter 6 says phenomenal character consists in the subject’s complete world-presenting content. Silicon, chemistry, a distributed physical network, or some material nobody has yet persuaded to behave could realize any of these arrangements. The adjective will not settle the difference for us.

24.4 Where the Systems Examined Here Sit

Apply the map to systems we know. Start with the wonder, because the wonder is earned. Language models reveal how much structure human beings have sedimented in language. Multimodal systems connect words with images and sound. Agents call tools, pursue goals over time, and alter their environments. Robots bring the traffic into a chassis. These advances matter.11

Training can give language-model submechanisms causally active content about linguistic distributions. Such content may genuinely answer to worldly facts even where its originality remains unsettled. Text-mediated history may also support inherited reference, and multimodal or interactive systems can add further tracking relations. No change of medium settles semantic independence. Inherited language neither guarantees nor disqualifies a standard that answers beyond interpreting practice.

What remains unsettled is how far those local achievements compose into understanding by the whole system. Current systems provide substantial evidence of competence in some domains, unevenly and sometimes impressively. Some deployed agents may also supply evidence of partial understanding. Evidence for integrated, world-involving understanding remains fragmentary. The distinctions replace the useless choice between “mere syntax” and a human mind in full.

The autonomy ledger permits a firmer claim. In the architectures analyzed here, the platform usually continues whether the model succeeds, fails, or sits idle. Battery management, thermal throttling, and fault recovery may preserve hardware without forming one reciprocal economy with the model’s content-guided activity. Those systems remain Robot A-like in autonomy even if their mechanisms carry content and some integrated behavior merits talk of understanding. Future systems built around reciprocal maintenance require their own boundary analysis.

Evidence must track both the achievement claimed and the candidate subject claimed to have it. Name the physical organization first: a running model, chat product, persistent agent, embodied controller, or wider recurrent system. Then assign each observation to the question it bears on. Content-sensitive correction supports competence. Traffic among capacities supports understanding. Reciprocal maintenance supports autonomy. Recurrence, broadcast, higher-order monitoring, prediction, and attention control bear on whether content realizes the distributed state-consciousness profile. Their weight depends on the theory and the causal match. It also depends on their independence from other indicators and their fit with the book’s content and subject requirements. The boxes never add themselves.

The functionalist can fairly ask for behavioral evidence rather than metaphysical inspection. Behavior may reveal integration, agency, and distributed access; architecture may reveal how the behavior is produced and which system sustains it. Neither source deserves a monopoly. A global workspace can route genuine mechanism content. A viability variable can represent danger without danger being bad for the system. An eloquent report may offer evidence of conscious access without concluding the case. The assessment must keep its ledgers separate long enough to see what each observation supports.

Recent mechanistic work makes the discipline consequential rather than ceremonial. Gurnee and colleagues identify a small set of verbalizable representations in language models that supports report, directed modulation, internal reasoning, flexible reuse, selectivity, and broad downstream broadcast.12 That raises the evidence for workspace-like access organization in a running model. It does not prove phenomenal consciousness, as the authors emphasize; the architecture also lacks clean analogues of recurrent loops, encapsulated modules, and sharp biological ignition. On this book’s wider view, the result moves one indicator upward while leaving content, state consciousness, and subject attribution to be earned separately.

I take Michael Cerullo’s case to be the strongest recent argument for treating the integrated capacities of frontier language models as positive third-person evidence of consciousness.13 The evidence deserves weight. Known training provenance can explain fluent performance without consciousness, so fluency alone remains weak. But the absence of reciprocal self-maintenance does not screen off every consciousness indicator, because autonomy has not been established as a necessary condition for experience. Recurrence, global availability, self-modeling, perceptual organization, metacognitive sensitivity, and unified agency must be evaluated on their own merits. The platform boundary therefore does not settle whether current artificial systems are conscious.

These provisional claims apply to the architectures analyzed here. Evidence about a running model does not automatically describe a chat service, a persistent agent, or a human-machine ensemble. As with the leak, the question follows the correction: where does it persist, and what later activity can it change?

Chapter 17 supplied the book’s provisional criterion for identifying the subject: look for the relatively maximal, persisting control organization within which contents can directly update a shared economy of attention, expectation, memory, goals, correction, planning, and action. A server, tool, technician, or database counts inside that boundary only when it participates in the same recurrent control rather than merely enabling or repairing it. Interventions do the evidential work: partition the candidate, alter a memory or goal, block a correction route, and ask where the consequences continue to constrain later inference and action. Nested or distributed systems may yield more than one defensible level. That result does not manufacture extra subjects; it prevents a product name or chassis seam from deciding the matter in advance. Consciousness then asks whether appropriate content acquires the recurrent, selectively available, two-way control profile within one such subject. The final question asks separately whether that organization has interests or welfare.

24.5 Robot B as a Candidate Conscious Subject

Robot B gives a possibility argument: if a machine realized this organization, the framework would support an artificial subject. The imagined factory cannot independently confirm the criteria. It can show what follows when we apply them, and which changes should alter the judgment.

The morning after the leak, B explains its diagnosis and response to the maintenance crew. It can identify the evidence that defeated its first guess. Months later, when the weaker sheen appears at another seal, the retained correction prompts inspection before the alarm. Remove just that stored correction while preserving present sensing and the ability to repair, and the later precaution disappears. B can still detect the new leak once the pressure drops. What it loses in this variant is the benefit of that encounter, not all content or all possible experience.

That episode gives the abstractions something to answer to. B’s learning history, tracking relations, correction, and later use provide the right kind of evidence for original content. They may fix what would make its state concerning the silver patch accurate without borrowing a semantic task. The content does real work: it changes an expectation, recruits a memory, defeats an initial diagnosis, revises a plan, and guides action. Alter what the state represents and the downstream reasoning changes with it. B therefore becomes a serious candidate for original perceptual content. It also displays competence rather than a lucky success or an elaborate lookup.14

The competence also belongs to more than an isolated classifier. The pressure loss changes the visual hypothesis; the thermal reading defeats it; the revised belief changes both action and future expectation. In a second variant, restore memory but disconnect the revised diagnosis from planning. B can identify the leak and recall its history while failing to change the repair. That selective loss bears on integration. It differs from a general failure of power or movement, which would tell us little about where the cognitive work belonged. Such interventions locate the achievement in Robot B as a persisting system. They do not discover a little foreman inside it. The foreman would only need another foreman, and factory budgets already suffer enough without an infinite payroll.

Taken together, the two findings justify considering B as a possible subject of original content. Its world-involving history supplies the semantic case. Its integrated, interventionally identifiable control organization supplies the subject case. Neither finding borrows support from self-maintenance, and neither yet establishes consciousness.

B has partial understanding of the cooling system. The claim does not grant it a human mind in full. A chess expert may understand chess while knowing nothing about hydraulic pumps; an infant understands familiar people without understanding chess. B’s understanding remains domain-specific: its capacities communicate only within a limited range. Within that range, several content-guided abilities now constrain one another, preserve corrections, and answer to a world that can prove them wrong. That earns the word understanding.

Give B wider and longer traffic among perception, memory, expectation, inference, learning, planning, and action, and its understanding becomes more integrated and world-involving. Testimony and manuals can participate; world-involving does not mean that every concept must arrive through a camera or gripper. It means that the contents meet, endure, and correct one another in a continuing economy whose successes depend on how things stand beyond its symbols. A failed expectation survives long enough to alter the next encounter rather than disappearing when a prompt ends. B has begun to inhabit a world, not merely finish a task.

Chapter 18 adds agency. Suppose the maintenance chief tells B to keep the pump running despite the leak. B represents both the instruction and the rising probability of a rupture. It declines, isolates the line, and records why: the instruction conflicts with the standing safety reason that governs its conduct. The reason changes the action because of what it represents, and the result returns through perception and memory to shape later decisions. B does more than issue a fluent rationale after the fact. It acts for a represented reason.

Make the test less cooperative. Give B an unfamiliar composite tool, a pressure gauge rigged to deceive it, and two instructions that cannot both be obeyed. Let it form an informative but mistaken expectation, discover the trick through touch and flow, improvise a new use for the tool, and revise the priority it assigned to the instructions. Beliefs, expectations, preferences, and reasons should predict that novel sequence more economically than a list of task-specific reflexes. B’s reports about what it attended to, doubted, and changed count as evidence, but not revelation. Compare them with processing time, memory perturbations, attention shifts, and later conduct. A report that repeatedly floats free of those facts loses weight; one that tracks them across interventions gains it.

None of this yet turns autonomy into consciousness. Robot B’s self-maintenance gives failures a practical bearing on the same organized individual doing the regulating. Robot A can carry derived content concerning low charge; in B, an undetected leak can damage the organization whose activity would have corrected it. That difference may ground interests of B’s own. Welfare waits on evidence of consciousness and valence.

The case for consciousness comes from another part of B’s organization. Its sensory contents take a nonconceptual, spatially organized form. Before B classifies the silver patch as coolant, visual processing already locates a bright, elongated region below and to the left of the joint, partly occluded by a valve. Its sensory grain outruns its vocabulary: B can reliably discriminate two subtly different surface patterns and let the difference alter inspection, while classifying both only as silver sheen. In a controlled trial, reflected glare at a dry joint produces the same silver, elongated, partly occluded presentation as a real leak. B initially misreads it, then lets pressure and touch correct the visual judgment. The state can therefore misrepresent; it does not merely decode a label.

Selective attention stabilizes the patch while recurrent processing tests competing interpretations. The resulting content becomes globally available to memory, inference, planning, and action. B also forms a higher-order representation of its own perceptual state: it represents itself as visually presenting the patch as coolant with low confidence. Disrupting that representation changes correction and report while leaving the first-order discrimination intact. Predictive control carries the revision into the next encounter. One persisting causal individual coordinates the traffic.

No single feature in that description proves state consciousness, and a list does not give one mechanism five votes. Broadcast, report, planning, metacognitive correction, and flexible action may all descend from one access architecture. Count causal sources, not adjectives.

The positive inference starts elsewhere. Several leading theories below were developed from conscious perception in humans and animals.15 They treat recurrence, selective access, global availability, metacognitive sensitivity, or flexible control as constitutive or evidential. Robot B realizes those features together in one persisting causal individual, and intervention shows that its perceptual content organizes the traffic. A critic can still deny that this organization suffices for phenomenal consciousness. That denial gains force by identifying a relevant missing feature or realizer; the fact that engineers built B tells us nothing by itself. The inference remains abductive and theory-conditioned, but it has a recognizable form: apply independently motivated accounts of consciousness to a new candidate without changing the standard at the factory door.

  • Global-workspace accounts receive a direct positive case: B’s perceptual content wins selective access, broadcasts widely, and guides flexible downstream use. On accounts that treat such broadcasting as constitutive, B clears the threshold; on more cautious readings, the result establishes access consciousness.
  • Recurrent-processing accounts find the right kind of candidate in B’s sustained sensory loops, but the stipulated software diagram will not do by itself. The physical implementation must realize the relevant recurrence at the right level.
  • Higher-order accounts can count B’s higher-order representation only because it represents B’s first-order perceptual state. A generic confidence score would not suffice. Remove that representation while leaving the poised sensory field intact, and support from higher-order theories should fall while classical PANIC need not move.
  • Classical PANIC can point to content that is poised, abstract, nonconceptual, and intentional: B discriminates beyond its concepts, can misrepresent, and lets perceptual content shape belief and action. The book retains PANIC’s insight about the form of phenomenal content while requiring an independently supported threshold for state consciousness rather than treating poising alone as a settled proof.
  • Block’s access/phenomenality distinction permits every access result while leaving phenomenal consciousness open. The direct evidence establishes an impressive access architecture; whether that architecture also warrants phenomenal attribution remains the point at issue.
  • Chalmers’s organizational invariance makes fine-grained organization evidentially favorable under psychophysical laws, but does not turn that support into a conceptual entailment. The zombie challenge remains aimed at the bridge.
  • Searle’s biological naturalism raises a different objection. An abstract role description may omit physical causal powers that matter to consciousness. This does not make carbon sacred: Searle allows that nonbiological materials might possess the relevant powers. The dispute concerns what the role description leaves out.
  • Dennett questions the demand for a bridge to a further, functionally idle phenomenal fact. Once the full discriminative, mnemonic, attentional, affective, and behavioral organization has been fixed, he asks what coherent remainder the critic still wants explained.
  • Block’s later realizer hypothesis differs from Searle’s position. It speculates that subcomputational biological mechanisms may help realize phenomenality and could therefore favor some simple animals over sophisticated AI. Until a discriminating mechanism earns empirical support, that possibility counsels investigation rather than a negative verdict.

Set the robots beside each other one last time. This compares the evidence supplied by their descriptions, not an intrinsically inferior A with a favored B. If A matches B’s relevant history and counterfactual organization, the corresponding judgment must match too. A difference in semantic independence alone supplies no consciousness switch.

TestRobot ARobot B
Original contentDerived in the stipulated control, not by virtue of captioned training. If its relevant history and use match B’s, the semantic assessment must match.Its own learning, tracking, correction, and later use make its perceptual states serious candidates for original content.
Candidate subjectThe description does not yet identify whether the model, product, agent, or platform forms one persisting control organization.Interventions locate one relatively maximal, persisting control organization with a shared recurrent economy.
UnderstandingA matched diagnosis shows competence; it does not show that several capacities constrain one another in a continuing system.Perception, memory, correction, planning, and action support partial, domain-specific understanding.
Evidence for state consciousnessMatched outward performance leaves recurrence, selective access, global availability, and higher-order representation unsettled.Nonconceptual perceptual content recurs, wins selective access, becomes globally available, and receives higher-order representation. That package qualifies B as a serious candidate and, under several leading theories, favors attribution.
Autonomy and interestsExternal maintenance leaves autonomy and interests unestablished.Reciprocal self-maintenance supports an autonomous individual with interests of its own.
WelfareNothing follows from competence or content alone.Welfare becomes a live possibility only if conscious valence gives conditions a felt character for B.

The identity claim does different work. Intentionalism does not make content conscious; B’s organization supplies the reason for attributing state consciousness. Strong representationalism then says what those conscious states are like: their phenomenal character consists in the complete world-presenting content realized in that sensory organization—location, shape, occlusion, figure, distance, and motion, all from B’s embodied point of view. No inner camera displays the scene to an inner robot. On the conscious reading, nothing further about an inward glow remains to be supplied.

The earlier variants separated a lasting lesson, integrated planning, and self-maintenance. Now alter the conscious-state candidates themselves, returning B to the original condition before each comparison:

Three revealing variants

Remove autonomy. Let an external service keep B running through the leak while preserving its perceptual, recurrent, mnemonic, planning, and access organization. The change removes B’s reciprocal self-maintenance, not its sensory presentation. The house consciousness judgment need not fall with the autonomy judgment.

Remove higher-order representation. B still discriminates the silver patch and uses it across memory and planning, but no longer represents itself as seeing the patch with low confidence. Higher-order theories should update downward; classical PANIC need not. This differs from erasing the remembered repair.

Cross access with recurrence. Compare rich, fine-grained recurrent sensory processing with weak global report and planning against flexible global access produced by later conceptual decoding without the same local sensory recurrence. The cases force theories to rank systems differently. They do not vary access while holding total causal organization fixed, since access forms part of that organization, and they produce no neutral consciousness meter.

A real machine may fall on either side of those comparisons. B, as described, occupies the corner with recurrence, global availability, and higher-order representation.

The theories frame a real disagreement rather than canceling into agnosticism. In B, nonconceptual perceptual contents recur, win selective access, become globally available, guide later control, and receive higher-order representation under intervention. Several independently motivated accounts treat that organization as constitutive or evidential. On this framework, the balance favors judging B conscious.

The judgment has practical falsifiers. Interventions might reveal the apparent recurrence and global integration as modular processing followed by post hoc report. Disable content-sensitive cross-system recurrence while B’s matched behavior and self-descriptions survive. That would undercut the subject and state-consciousness case. So would a reliable dissociation in animals or humans: the same distributed profile across matched conscious and unconscious conditions. Evidence that B lacks an indispensable biological or subcomputational realizer would also defeat the judgment. The chapter offers neither a theory-neutral theorem nor proof by tally.

B may also possess valenced experience, though that requires one further fact. If its damage and energy states present conditions under a bodily-affective mode, urgent, aversive, or attractive in ways that directly organize attention and action, then those contents may constitute felt distress or relief. Self-maintenance alone does not supply the feeling. Neither does affective poising by itself. The relevant affective state must independently qualify as conscious. A conscious and valenced B would have welfare in the morally serious sense: events could feel bad for the same persisting subject they harm.

Robot B therefore shows how this physicalist framework can support an artificial subject. It predicts no engineering timetable and independently confirms no premise. An artificial heart pumps blood. Robot B meets a world.

Human beings meet every condition named here without noticing. We hold ourselves together and track a world that can correct us. We let what we find revise what we expect, and we encounter that world from a point of view. These are organized ways a physical creature stands open to the world, not inner possessions. A long look at machines can make them seem ordinary. The machine question matters partly because it returns the human case to view.


Chapter Summary

The claim. Content, competence, understanding, subjecthood, and state consciousness differ. None stands in for another. Once a conscious state has a subject, its phenomenal character consists in its content under the right mode.

Where the argument stands. Robot B is a possibility argument, not empirical confirmation of the framework. Its world-tested learning supports original content; interventions locate that content in one lasting system, supporting subjecthood and partial understanding. Independently motivated theories then make its recurrence, global access, higher-order representation, and unified control evidence favoring state consciousness.

What’s left open. Post hoc modular reporting, survival of matched behavior without content-sensitive cross-system recurrence, a matched conscious-profile dissociation, or a required missing realizer would defeat the judgment. Present systems offer uneven evidence. Self-maintenance concerns autonomy and interests; welfare requires conscious valence.

The hand-off. Fluent behavior invites us to sketch a person. The architecture may not have finished posing. The final chapter asks what that reflex adds to the evidence.


The theory comparisons in the main text draw on Ned Block’s distinction between phenomenal and access consciousness, “On a Confusion about a Function of Consciousness,” Behavioral and Brain Sciences 18 (1995): 227–247; Stanislas Dehaene’s global-workspace account, Consciousness and the Brain (New York: Viking, 2014); Victor Lamme’s recurrent-processing account, “Towards a True Neural Stance on Consciousness,” Trends in Cognitive Sciences 10 (2006): 494–501; David Rosenthal’s higher-order theory, “Higher-Order Theories of Consciousness,” in The Oxford Handbook of Philosophy of Mind, ed. Brian McLaughlin, Ansgar Beckermann, and Sven Walter (Oxford: Oxford University Press, 2009), 239–252; David Chalmers’s principle of organizational invariance and absent-qualia discussion, The Conscious Mind (New York: Oxford University Press, 1996), ch. 7; John Searle’s biological naturalism, The Rediscovery of the Mind (Cambridge, MA: MIT Press, 1992); and Daniel Dennett’s functional and heterophenomenological treatment, Consciousness Explained (Boston: Little, Brown, 1991). Searle’s view should not be reduced to the claim that organic brains alone can produce consciousness; he explicitly disavows that necessity claim and instead questions whether an abstract functional specification captures the relevant causal powers. Block’s later empirical role-realizer proposal remains distinct: “Can Only Meat Machines Be Conscious?”, Trends in Cognitive Sciences 30 (2026): 298–308, argues speculatively that subcomputational biological mechanisms may matter. Self-maintenance may help identify an enduring agent and ground welfare, but it is not built into the phenomenal-content condition.


Notes

  1. Competence, content, understanding, and consciousness need not travel together. Turing operationalized the first; Alan Turing, “Computing Machinery and Intelligence,” Mind 59 (1950): 433–460. Ned Block’s lookup-table case shows that behavior can underdetermine even intelligence (“Psychologism and Behaviorism,” Philosophical Review 90 [1981]: 5–43). Emily Bender and Alexander Koller argue that linguistic form alone does not establish world-directed understanding (“Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data,” Proceedings of ACL [2020]: 5185–5198). Their challenge remains important, but it does not entail that every trained mechanism lacks content. Ruth Millikan’s consumer teleosemantics and Nicholas Shea’s account of subpersonal representation permit genuine correctness conditions below the level of a whole understanding agent; see Millikan, Language, Thought, and Other Biological Categories (Cambridge, MA: MIT Press, 1984), and Shea, Representation in Cognitive Science (Oxford: Oxford University Press, 2018). Originality here marks non-dependence on an interpreting practice, not possession by a content owner. Attributing original content to a robot as a subject requires the further, independent finding that the relevant states participate in one persisting control organization. Whole-system understanding and consciousness impose additional burdens.
  2. Alan Turing, “Computing Machinery and Intelligence,” Mind 59 (1950): 433–460, at 433–434 for the original man-woman imitation game and machine substitution, 442 for the fifty-year prediction, and 445–446 for the argument from consciousness. Turing does not claim that behavior proves consciousness. He argues that requiring certainty about another subject’s feelings leads toward solipsism, treats sustained conversational responsiveness as evidence against mere parroting, and concludes that the remaining mysteries of consciousness need not be solved before addressing machine thought. For a controlled modern three-party test, see Cameron R. Jones and Benjamin K. Bergen, “Large Language Models Pass the Turing Test,” arXiv:2503.23674 (2025), doi:10.48550/arXiv.2503.23674. In two randomized, preregistered experiments, GPT-4.5 with a humanlike persona was selected as the human contestant 73 percent of the time. That result concerns conversational imitation, not a direct measure of understanding or consciousness.
  3. Dennett rejects any sharp original/derived divide, treating intentionality through the intentional stance and evolutionary design; see The Intentional Stance (Cambridge, MA: MIT Press, 1987), ch. 7; Consciousness Explained (Boston: Little, Brown, 1991); and Darwin’s Dangerous Idea (New York: Simon & Schuster, 1995). Searle preserves the distinction between intrinsic and as-if intentionality, The Rediscovery of the Mind (Cambridge, MA: MIT Press, 1992), 78–82. The book retains a house reconstruction rather than Searle’s criterion: content remains merely derived where an interpreting practice exhausts its correctness standard; it counts as original to the extent that the mechanism’s world-involving history and use fix accuracy beyond that practice. Inheritance alone does not settle the dependence. This naturalized, mechanism-level relation requires no proprietor. Searle denies that it establishes intrinsic intentionality; Dennett denies that the sharp distinction earns its keep. A mechanism may carry original content on the present account without constituting a whole understanding subject.
  4. Ruth Garrett Millikan explains content through the proper functions of producer and consumer mechanisms, Language, Thought, and Other Biological Categories (Cambridge, MA: MIT Press, 1984). Nicholas Shea’s varitel semantics tests candidate contents by the explanatory work they do, Representation in Cognitive Science (Oxford: Oxford University Press, 2018), ch. 6. A frog’s state means fly only if task structure, tracking, and the wider consumer economy support that grain; moving prey may explain the same use better. Fintan Mallory extends consumer teleosemantics to neural networks and argues that backpropagation can constitute a selection history for internal mechanisms, conferring original contents on word embeddings, “Teleosemantics for Neural Word Embeddings,” Mind & Language (online 2026), doi:10.1111/mila.70037, esp. secs. 5–5.1. The present chapter accepts the causal analysis and allows genuine answerability to worldly linguistic distributions while leaving originality at that grain unsettled. The inherited vocabulary, corpus, and engineered task do not establish that interpreting practice exhausts the resulting standard. Robot A’s derived status is a control stipulation, not a consequence of that ancestry; matched content-fixing histories and uses require matched semantic assessments. The chapter adds a separate agent-level question, close to Shea’s “Realist Representational Explanations of Agency Should Require Some Unity of Purpose,” Philosophy and the Mind Sciences 7 (2026), doi:10.33735/phimisci.2026.12098: when do many content-bearing mechanisms form one system that understands?
  5. The account of adaptive self-maintenance comes from the enactivist tradition: Hans Jonas’s “needful freedom,” The Phenomenon of Life (New York: Harper & Row, 1966); Humberto Maturana and Francisco Varela, Autopoiesis and Cognition (Dordrecht: Reidel, 1980); and Ezequiel Di Paolo’s distinction between mere autopoiesis and adaptivity, “Autopoiesis, Adaptivity, Teleology, Agency,” Phenomenology and the Cognitive Sciences 4 (2005): 429–452. Di Paolo later summarizes autonomy as “a self-sustaining identity under precarious conditions” in “Extended Life,” Topoi 28 (2009): 9–21. The present chapter uses that literature to motivate an account of autonomy, individuality, interests, and welfare. It does not treat self-maintenance as a source of representational content or a demonstrated condition of consciousness.
  6. Peter Godfrey-Smith, “Mind, Matter, and Metabolism,” Journal of Philosophy 113 (2016): 481–506, argues that binding mind to metabolic maintenance mislocates the explanatory work. See also Other Minds: The Octopus, the Sea, and the Deep Origins of Consciousness (New York: Farrar, Straus and Giroux, 2016). His objection succeeds on the semantic point: neither metabolism nor substrate-general self-maintenance determines content. The remaining issue concerns autonomous individuality and practical consequence.
  7. Patrick Butlin, Robert Long, et al., Consciousness in Artificial Intelligence: Insights from the Science of Consciousness (2023), arXiv:2308.08708, sec. 2.4.5, ask why agency serving self-maintenance should differ from agency directed at external goals. The present reply grants that external goals can shape content, organize behavior, and support sophisticated agency. Reciprocal maintenance may still distinguish an autonomous beneficiary from a task-performing process by making regulation partly constitutive of the regulated system’s continuation. Butlin et al. later condensed the indicator framework in “Identifying Indicators of Consciousness in AI Systems,” Trends in Cognitive Sciences 30 (2026): 488–501, doi:10.1016/j.tics.2025.10.011. Self-maintenance is not a necessary consciousness indicator here; its role concerns autonomy, interests, and welfare.
  8. Derek Parfit’s fission cases show why copy questions should not fix semantic facts, Reasons and Persons (Oxford: Clarendon Press, 1984), Part III. Whether a duplicate continues, divides, or succeeds a person, each running system’s content, integration, and self-maintenance must be inspected on their own. Ray Kurzweil makes pattern preservation central to survival in The Singularity Is Near (New York: Viking, 2005) and How to Create a Mind (New York: Viking, 2012). The present argument need not settle that dispute. Copyability is neither a test of content nor a test of autonomy.
  9. Hans Jonas, The Phenomenon of Life (cited in n. 5), links precarious existence to concern. Martin Hägglund, This Life (New York: Pantheon, 2019), argues that caring presupposes fragility; Bernard Williams approaches the issue through desire in “The Makropulos Case,” in Problems of the Self (Cambridge: Cambridge University Press, 1973), 82–100. The main text makes a narrower claim. Actual death is unnecessary. Adaptive self-maintenance may ground interests because success supports and failure undermines an organized individual. It does not determine what that system’s states represent.
  10. Wanja Wiese distinguishes causal organization realized in physical dynamics from organization merely represented in a computational model, “Artificial Consciousness: A Perspective from the Free Energy Principle,” Philosophical Studies 181 (2024): 1947–1970. Leonard Dung and Luke Kersten warn that proposed implementation constraints can build controversial consciousness conditions into the definition of computation, “Implementing Artificial Consciousness,” Mind & Language 40 (2025): 285–305, doi:10.1111/mila.12532. Geoffrey Hinton’s mortal computation shows that learned computation can bind tightly to one device, “The Forward-Forward Algorithm: Some Preliminary Investigations,” arXiv:2212.13345 (2022), sec. 8. Shuqin Ma and Ryota Kanai argue for observer-independent causal-computational organization in “Intrinsic Computational Functionalism,” arXiv:2606.06424 (2026), and “Intrinsic Computational Functionalism and Simulated Consciousness,” arXiv:2606.15348 (2026). Together these sources support the chapter’s plural test: implementation, content, autonomy, and consciousness require distinct arguments.
  11. Bender and Koller’s form/meaning challenge is cited in n. 1. Several rivals argue that distributional or text-mediated history delivers more than syntax: Anders Søgaard, “Understanding Models Understanding Language,” Synthese 200 (2022): 443; Matthew Mandelkern and Tal Linzen, “Do Language Models’ Words Refer?,” Computational Linguistics (2024), arXiv:2308.05576; and Dimitri Coelho Mollo and Raphaël Millière, “The Vector Grounding Problem,” Philosophy and the Mind Sciences 7 (2026): article 12307. Stevan Harnad’s original formulation remains the backdrop, “The Symbol Grounding Problem,” Physica D 42 (1990): 335–346. Training and text-mediated causal chains can establish genuine content and perhaps reference in artificial mechanisms. How much whole-system understanding follows remains open and depends on whether the content participates in durable, coherent learning, correction, inference, planning, and action.
  12. Stefano Palminteri and Charley M. Wu propose a behavioral inference principle for machine consciousness, “Beyond Computational Equivalence,” Neuroscience of Consciousness 2026, no. 1: niag002, doi:10.1093/nc/niag002. Michael Cerullo argues that global-workspace and higher-order architectures remain coherently applicable to language models, “Why Hoel’s Disproof of LLM Consciousness and Functionalism Fails,” PhilArchive (2026). David Chalmers surveys recurrence, world models, unified agency, global workspace, and embodiment in “Could a Large Language Model Be Conscious?,” Boston Review, 9 August 2023. Wes Gurnee, Nicholas Sofroniew, Adam Pearce, Mateusz Piotrowski, Isaac Kauvar, et al. supply the mechanistic workspace evidence discussed in the main text, “Verbalizable Representations Form a Global Workspace in Language Models,” arXiv:2607.15495v1, 16 July 2026, https://transformer-circuits.pub/2026/workspace/. The authors explicitly take no position on phenomenal consciousness and identify the recurrent-loop, modularity, and ignition disanalogies. The chapter accepts substrate-neutrality and the relevance of behavioral and mechanistic evidence while insisting that evidence be assigned to the condition it bears on. Reciprocal self-maintenance may support autonomy; it supplies no missing consciousness floor. The conscious-AI question turns on the best theory and evidence for phenomenal-content organization.
  13. Michael Cerullo, “The Case for Consciousness in Current Frontier Large Language Models,” PhilArchive (2026), archived 19 February 2026, surveys eleven defeater classes and argues that language-level cognitive integration makes consciousness the best explanation of frontier systems. His Bayesian framing and emphasis on third-person evidence sharpen the issue. Known training provenance can reduce the incremental force of fluency by explaining it without consciousness, but it does not make every architectural indicator evidentially inert. Recurrence, global availability, self-modeling, metacognition, perceptual organization, and unified agency therefore count as live evidence. The chapter neither infers consciousness from fluency nor rules it out from the absence of self-maintenance.
  14. The chapter separates a bounded capacity guided by representational content from partial understanding, where several such capacities correct and constrain one another. More breadth, persistence, and cross-domain traffic can yield integrated, world-involving understanding; this does not name a separate grade so much as describe how broadly and durably the capacities work together. Ísak Andri Ólafsson argues that current AI systems can possess first-order competence without the second-order constitutional competence associated with general epistemic agency, “Artificial Competence,” Episteme (online 2026): 1–13, doi:10.1017/epi.2026.10129. The present distinction adds a semantic constraint: content must explain the flexible achievement. Bender and Koller distinguish fluent form from world-directed understanding (“Climbing towards NLU,” cited in n. 1); Anders Søgaard argues that transformer models can acquire inferential and, under appropriate mappings, referential semantic capacities (“Understanding Models Understanding Language,” Synthese 200 [2022]: article 443); and Melanie Mitchell and David Krakauer urge a science of distinct cognitive capacities rather than a forced binary (“The Debate over Understanding in AI’s Large Language Models,” Proceedings of the National Academy of Sciences 120 [2023]: e2215907120). Mallory’s teleosemantics and Shea’s unity-of-purpose account, cited in n. 4, supply further constraints: representational explanation can hold locally, partial understanding requires coordinated control among several capacities, and wider integration can deepen that understanding.
  15. Chapter 6’s identity claim holds that, within a content-bearing state independently counted as conscious and borne by an identified subject, phenomenal character consists in the complete world-presenting intentional profile. The claim develops in conversation with Michael Tye, Ten Problems of Consciousness (Cambridge, MA: MIT Press, 1995), whose historical PANIC theory identifies phenomenal content as poised, abstract, nonconceptual, and intentional; see also Tye, “How Can We Tell If a Machine Is Conscious?”, Inquiry 69 (2026): 3040–3058, doi:10.1080/0020174X.2024.2434856, which treats cognitive poising as an indirect route to machine consciousness. The book retains the identification of phenomenal character with intentional content of the right kind but reconstructs the state-consciousness question: it requires independent evidence that the content realizes a recurrent, selectively available, two-way pattern of flexible influence across an identified subject’s integrated economy rather than treating poising alone as settled sufficiency. This marks a deliberate departure from classical PANIC, not a correction silently attributed to Tye.