The phone
A world of software in your palm—but only after you reach, unlock, look, and operate.
First it lived in apps. Then it moved onto our wrists, fingers, glasses, and roads. Now it is beginning to navigate rooms and work beside us. This is not five unrelated gadgets. It is one steady migration: from digital capability to useful presence in the physical world.
Nearly three centuries of bodies waiting for a capable mind.
Plate: the Turk's cabinet opened, copper engraving, Joseph Racknitz, 1789. Public domain, Humboldt University Library.
In 1770 a machine sat down at a chessboard and beat half of Europe. The intelligence inside was a hidden human.
Wolfgang von Kempelen's Mechanical Turk toured the world's courts for eight decades and reportedly played Napoleon Bonaparte and Benjamin Franklin. It was a sensation because everyone wanted it to be true: a machine that perceives the world and acts in it. The body was convincing. The mind was a chess master crouched in the cabinet.[1]
That has been the shape of the failure ever since. Vaucanson's celebrated duck of 1739 flapped and ate and digested nothing.[2] Factory arms of the 1960s moved with superhuman precision and understood none of it. We could build bodies. We could not build the part that knows what the room means.
Then the order reversed. The last decade produced minds without bodies: models that read, reason, plan, and write software, trapped behind a text box. For the first time in the 287-year history of this dream, the missing piece is not the intelligence. It is the hardware catching up to it. That is what makes this moment different, and it is why the timeline below bends so hard after 2023.
Each step asked a little less attention from us—and learned to do a little more with its surroundings.
The important story is not that hardware became smaller. It is that intelligence gained better senses, richer context, and eventually a way to act.
Mechanical bodies created the appearance of intelligence long before machinery could supply the real thing.
It flapped, drank, ate, and appeared to digest. The performance was mechanical theater; the intelligence was projection.[2]
A chess-playing automaton toured Europe. Its brilliant “machine mind” was a human operator concealed in the cabinet.[1]
DARPA’s autonomous-vehicle challenges turned physical reasoning into a measurable engineering problem.
Fifteen finalists entered the first Grand Challenge. None completed the desert route.[3]
Stanford won the second challenge; five vehicles completed 132 autonomous desert miles. A failure became a field in eighteen months.[3]
Six vehicles finished DARPA’s Urban Challenge while interacting with manned and unmanned traffic.[4]
The interface traveled from desk, to hand, to wrist, to a sensor that could disappear into the background.
A full software platform arrived in a small handheld device, operated directly with our fingers.[5]
Computing shifted to the wrist: timely information, health signals, and small interactions designed to be checked at a glance.[6]
The screen disappeared. A device could learn quietly from sleep, temperature, heart rate, and movement, then explain the pattern later.[7]
Robots left the lab while foundation models learned to convert language and vision into movement.
Boston Dynamics opened commercial sales after early adopters used Spot in construction, energy, factories, and hazardous environments.[8]
Google DeepMind’s vision-language-action model trained on web and robot data, then emitted control instructions for a physical robot.[9]
Capital, manufacturing, and real deployment—not polished demo reels—become the proof.
Figure announced more than $1 billion in Series C funding at a $39 billion post-money valuation to scale Helix, manufacturing, and deployments.[13]
NEO opened preorders at $20,000, with U.S. priority delivery stated for 2026. Its advertised tasks include tidying, fetching, opening doors, and turning off lights.[14]
Investment is moving from “can a humanoid walk?” toward manufacturing, field operations, and repeatable work.[15]
Embodiment begins before a machine has arms and legs.
A system becomes physically useful when it can sense the same moment you are in, remember what matters, and respond without pulling you out of the room.
A world of software in your palm—but only after you reach, unlock, look, and operate.
The interface becomes a moment: a tap, a direction, a pulse, the next thing on your calendar.
No screen at all. It pays attention continuously, then turns a stream of signals into a useful pattern.
Even G2 puts prep notes, quiet cues, directions, and summaries into a discreet display while your hands stay free.[16]
The shift is subtle: from a device you visit to intelligence that shares the moment with you.
Autonomous driving is embodied intelligence at metropolitan scale: perception, prediction, planning, and motion in one loop. Through March 2026, Waymo reported 220.6 million rider-only miles and materially lower crash rates than its human-driver benchmarks. Those are company-reported figures, but the scale is the point: physical AI is already operating outside the lab.[17]
A personal corpus is the record you choose to keep.
The phone put compute in your pocket. Watches and rings made passive self-data normal. The next layer is a private, permissioned record of the moments that matter: signals from your body, words from a conversation, and eventually the first-person context that helps an assistant understand your day.
Health platforms already make a personal data layer practical: with permission, apps can read and write health and fitness data while the person remains in control of access and deletion.
From: sleep, heart rate, movement, and recovery.[21]
A recorder such as PLAUD can be intentionally started for a meeting, a thought, or an interview, then turn that moment into a searchable transcript and summary. It extends memory without claiming to record a life by default.
From: conversations you choose to capture.[22]
Meta’s Ego4D research shows the direction, not a released 24/7 consumer product: AI learning to locate an answer within past first-person video and understand the ongoing visual, audio, and motion context of daily life.
From: research into episodic memory and wearable perception.[23]
Useful context must be visible, consented, controllable, and yours to remove.
The body, the model, and the training world are arriving at the same time.
Humanoid demos are not new. What changed is that movement can now be paired with models that generalize, simulation that multiplies experience, and hardware designed for deployment.
RT-2 showed that web knowledge and robot demonstrations can live in one vision-language-action model. π0 pushed the idea toward a generalist policy across robots and dexterous tasks.
Change: actions can be learned as a reusable vocabulary, not hand-coded one task at a time.
Modern robotics stacks train and test in simulation, then transfer policies into physical machines. NVIDIA’s Isaac platform and GR00T package that loop for humanoid developers.
Change: a robot can accumulate more practice than the calendar alone permits.
Electric actuators, better batteries, dexterous hands, and maturing supply chains are turning spectacular prototypes into products designed for factories and homes.
Change: the success metric becomes a useful shift—not a flawless demo.
Do not confuse a long-range scenario with a near-term market.
The headline numbers span different years and include different things. They are not apples-to-apples. Read them as evidence of widening possibility—not a promise that a robot will unload your dishwasher next Tuesday.
Goldman’s 2024 estimate projects a $38 billion total addressable market in 2035.[18] Morgan Stanley’s 2025 scenario reaches beyond robot sales into supply chains, maintenance, and support, and explicitly expects adoption to remain relatively slow until the mid-2030s.[19] Citi’s 2024 model puts almost 650 million humanoids and a potential $7 trillion market in 2050.[20]
The first useful body may be yours.
That is why ambient systems matter. They let intelligence practice the hard parts of physical life—timing, context, memory, restraint—without pretending today’s hardware can do everything.
It listens when invited, remembers the context you choose to give it, and surfaces the next useful thing while you remain the actor.
It perceives the shared environment, plans within real constraints, and acts through a car, arm, mobile platform, or humanoid body.
The future is not intelligence replacing daily life. It is intelligence learning how to fit inside it.
Primary sources first. Forecasts labeled as forecasts.
Product capabilities are the manufacturers’ stated claims unless independent performance data is explicitly noted. Market estimates use different definitions and horizons.
Research reviewed July 19, 2026. Dates above refer to the source publication or the data cutoff stated by the source. Company announcements describe company claims; forecast figures describe analyst scenarios, not observed outcomes.