The Turing Test is Alan Turing’s proposed replacement for the question “Can machines think?”: a conversation game in which a machine counts as thinking if human judges cannot reliably tell it from a person by typed exchange alone. This page sets out the test, what Turing did and did not claim for it, the standard objections, and where the arrival of large language models leaves it.
The Test
Alan Turing opened “Computing Machinery and Intelligence” (1950) by refusing the question he had been handed. “Can machines think?” he judged “too meaningless to deserve discussion” — the word think carried too much disputed freight to settle anything. So he proposed a substitute he could actually run: the imitation game.
A human judge sits at a terminal and exchanges typed messages with two hidden partners — one human, one machine. Each tries to convince the judge that it is the human. The judge’s job is to tell them apart. Turing’s wager was that a machine which reliably defeats the judge (one that cannot be distinguished from a person by conversation alone) should be credited with thinking. We need not peer inside it; we need only watch what it can do.
The move is deliberately and self-consciously behavioral. It trades an unanswerable question about inner states for an answerable question about performance. And it is restricted, by design, to language: the terminal hides the body, the voice, the face, so that nothing but the quality of the conversation can move the judge. Turing predicted that by the year 2000 machines would play the game well enough that “an average interrogator will not have more than 70 per cent chance of making the right identification after five minutes of questioning.”
What Turing Was — and Was Not — Claiming
The test is routinely misremembered as a theory of consciousness. It is nothing of the kind, and Turing said so. Replying to “the argument from consciousness” — the objection that a machine could never feel anything — he declined to settle whether the machine has any inner life at all. He bracketed the question. His test asks whether a system thinks in the operational sense of conversing indistinguishably from a person; it does not ask, and does not pretend to answer, whether there is anything it is like to be that system.1 He offered an operational criterion for a deflated notion of thinking; his successors inflated a passed test into proof of an inner life — exactly the inference Turing himself withheld.
Turing also anticipated, and answered in turn, a battery of objections he saw coming: the theological objection (thinking requires a soul), the mathematical objection (Gödel’s results limit what any formal machine can prove), and “Lady Lovelace’s objection” — that a machine “can only do what we know how to order it to perform,” and so can never originate anything. His replies were sharp enough that the paper still reads as the founding document of the field.
The Standard Objections
Searle’s Chinese Room (1980). The most famous reply is built as a deliberate dual of the imitation game. Searle imagines himself locked in a room, manipulating Chinese symbols by a rulebook well enough to pass a Turing Test for Chinese — and understanding not a word. If a system can pass the test while understanding nothing, then passing the test is not sufficient for understanding. Behavioral indistinguishability and genuine semantics come apart. (See The Chinese Room.)
Block’s psychologism (1981). Ned Block presses a different lever. Imagine a machine that passes the test by sheer brute lookup — a vast table pairing every possible conversational input, up to some length, with a sensible canned reply. Such a machine could in principle win the imitation game, yet it computes nothing, understands nothing, and possesses no intelligence in any sense we care about; it merely indexes a finite list someone else wrote. Block’s point is that the Turing Test is behaviorist — it fixes on the input-output profile and is therefore blind to how that profile is produced. But intelligence is partly a matter of internal organization, not just external performance. A test that cannot tell a thinker from a jukebox tests the wrong thing.2
The general charge. Both objections converge on one diagnosis. The imitation game is a criterion of seeming, and seeming is cheap. It can be manufactured by understanding, or faked by lookup, or stumbled into by statistics — and the test, watching only the output, cannot distinguish the cases. Whatever we most wanted to know when we asked “Can machines think?” is precisely what an operational test of conversation is structured to leave out.
The Large-Language-Model Moment
For half a century the test sat as a thought experiment. It no longer does. Contemporary language models often produce text that human judges struggle to distinguish from human writing; on important bounded versions, the game is substantially won. This has not settled the philosophical question — it has sharpened it.
A base model trained on text predicts tokens from a corpus descended from human linguistic practice. Training can make derived content causally internal, and deployed systems may add memory, tools, sensors, and world-sensitive feedback. Bender and Koller’s point remains: linguistic form alone does not establish independently grounded meaning or distal reference.3 Passing the imitation game therefore provides evidence of conversational competence while leaving the provenance and scope of content, understanding, consciousness, and the bearer unsettled. The test succeeded as an engineering target and failed as a self-sufficient criterion of mind.
My View
The Turing Test sets exactly the bar my view rejects — and it is worth being precise about where the fault lies, because it does not lie with Turing.
In 1950, with behaviorism ascendant and no science of the inner life in prospect, an operational test was a reasonable, clarifying move, and Turing was scrupulous: he bracketed consciousness rather than pretending his test disposed of it. Give him that. The error enters with his successors, who forget the bracketing and brandish a passed test as proof of understanding or experience. The test’s own author declined that inference. So do I.
My objection to the criterion is the objection that sank behaviorism, of which the test is the cleanest single instrument. Behavior alone underdetermines the organization producing it. Conversational performance can supply defeasible evidence of content-guided competence, but a terminal transcript cannot by itself identify the bearer, reveal content-fixing history, show that several capacities compose understanding, or establish phenomenal organization. Those questions require evidence about the system behind the exchange. See The Symbol Grounding Problem and Semantic Externalism.
The LLM moment therefore vindicates the objection without making performance worthless. Winning the game and having a mind remain different achievements; success at one can still count as evidence relevant to the other. The anthropomorphic reflex makes fluent talk feel more conclusive than it is, while provenance and architecture determine how much evidential weight survives. Turing named the modern question well. His readers erred when they turned one behavioral measure into a verdict about every dimension of mind.
Related Concepts
- Behaviorism — the Turing Test is behaviorism’s cleanest instrument; it inherits behaviorism’s decisive failure, the inability to distinguish a genuine inner state from perfect mimicry
- The Chinese Room — Searle’s argument is built as a dual of the imitation game: a system that passes the test for Chinese while understanding nothing
- Functionalism — the test measures input-output profile; Block’s objection (right behavior, wrong or absent internal organization) is the same lever that troubles functionalism
- The Symbol Grounding Problem — what a terminal-bound test screens off: the causal contact between symbols and the world that meaning requires
- Semantic Externalism — meaning depends on world-involving relations a conversation test cannot detect, present or absent
- Computation and the Church-Turing Thesis — Turing the logician proved which functions a machine can compute; the test is a separate, later, and far weaker claim about how to tell a thinker from a non-thinker
- The Hard Problem of Consciousness — Turing bracketed consciousness; the question he set aside is the one that does not go away