Intelligence is moving beyond the screen. Into the rhythm of everyday life.

First it lived in apps. Then it moved onto our wrists, fingers, glasses, and roads. Now it is beginning to navigate rooms and work beside us. This is not five unrelated gadgets. It is one steady migration: from digital capability to useful presence in the physical world.

287 years of trying to build a body
5 surfaces on one arc
Every claim cited below

1. The dream is old. The mind is new.

Nearly three centuries of bodies waiting for a capable mind.

Plate: the Turk's cabinet opened, copper engraving, Joseph Racknitz, 1789. Public domain, Humboldt University Library.

In 1770 a machine sat down at a chessboard and beat half of Europe. The intelligence inside was a hidden human.

Wolfgang von Kempelen's Mechanical Turk toured the world's courts for eight decades and reportedly played Napoleon Bonaparte and Benjamin Franklin. It was a sensation because everyone wanted it to be true: a machine that perceives the world and acts in it. The body was convincing. The mind was a chess master crouched in the cabinet.[1]

That has been the shape of the failure ever since. Vaucanson's celebrated duck of 1739 flapped and ate and digested nothing.[2] Factory arms of the 1960s moved with superhuman precision and understood none of it. We could build bodies. We could not build the part that knows what the room means.

Then the order reversed. The last decade produced minds without bodies: models that read, reason, plan, and write software, trapped behind a text box. For the first time in the 287-year history of this dream, the missing piece is not the intelligence. It is the hardware catching up to it. That is what makes this moment different, and it is why the timeline below bends so hard after 2023.

2. The path was not screen to robot. It was screen to world.

Each step asked a little less attention from us—and learned to do a little more with its surroundings.

The important story is not that hardware became smaller. It is that intelligence gained better senses, richer context, and eventually a way to act.

1739–1770

The imitation

Mechanical bodies created the appearance of intelligence long before machinery could supply the real thing.

1739
Vaucanson’s duck body

It flapped, drank, ate, and appeared to digest. The performance was mechanical theater; the intelligence was projection.[2]

1770
The Mechanical Turk body

A chess-playing automaton toured Europe. Its brilliant “machine mind” was a human operator concealed in the cabinet.[1]

2004–2007

The world becomes navigable

DARPA’s autonomous-vehicle challenges turned physical reasoning into a measurable engineering problem.

2004
No car finishes mind

Fifteen finalists entered the first Grand Challenge. None completed the desert route.[3]

2005
Five cars finish mind

Stanford won the second challenge; five vehicles completed 132 autonomous desert miles. A failure became a field in eighteen months.[3]

2007
Traffic, people, intersections mind

Six vehicles finished DARPA’s Urban Challenge while interacting with manned and unmanned traffic.[4]

2007–2015

The intelligence moves closer

The interface traveled from desk, to hand, to wrist, to a sensor that could disappear into the background.

2007
iPhone held

A full software platform arrived in a small handheld device, operated directly with our fingers.[5]

2014–15
Apple Watch glanced

Computing shifted to the wrist: timely information, health signals, and small interactions designed to be checked at a glance.[6]

2015
Oura Ring sensed

The screen disappeared. A device could learn quietly from sleep, temperature, heart rate, and movement, then explain the pattern later.[7]

2020–2024

The model learns action

Robots left the lab while foundation models learned to convert language and vision into movement.

2020
Spot goes on sale body

Boston Dynamics opened commercial sales after early adopters used Spot in construction, energy, factories, and hazardous environments.[8]

2023
RT-2 translates language into action mind

Google DeepMind’s vision-language-action model trained on web and robot data, then emitted control instructions for a physical robot.[9]

2024
GR00T, electric Atlas, π0 convergence

NVIDIA announced a humanoid foundation model; Boston Dynamics showed a commercial electric Atlas program; Physical Intelligence published a generalist robot policy.[10][11][12]

2025–2026

The product race

Capital, manufacturing, and real deployment—not polished demo reels—become the proof.

2025
Figure funds embodied intelligence scale

Figure announced more than $1 billion in Series C funding at a $39 billion post-money valuation to scale Helix, manufacturing, and deployments.[13]

2025–26
1X sells the home promise home

NEO opened preorders at $20,000, with U.S. priority delivery stated for 2026. Its advertised tasks include tidying, fetching, opening doors, and turning off lights.[14]

2026
Apptronik’s Series A passes $935M scale

Investment is moving from “can a humanoid walk?” toward manufacturing, field operations, and repeatable work.[15]

3. The wearable is the on-ramp, not the destination.

Embodiment begins before a machine has arms and legs.

A system becomes physically useful when it can sense the same moment you are in, remember what matters, and respond without pulling you out of the room.

Held · 2007

The phone

A world of software in your palm—but only after you reach, unlock, look, and operate.

Your attentionhigh
Its actionlow
Glanced · 2015

The watch

The interface becomes a moment: a tap, a direction, a pulse, the next thing on your calendar.

Your attentionmedium
Its actionlow
Passive · 2015

The ring

No screen at all. It pays attention continuously, then turns a stream of signals into a useful pattern.

Your attentionlow
Its actioninsight
Ambient · now

The glasses

Even G2 puts prep notes, quiet cues, directions, and summaries into a discreet display while your hands stay free.[16]

Your attentionlower
Its actioncontext

The shift is subtle: from a device you visit to intelligence that shares the moment with you.

The road got there first. sensepredictdecideact A physical environment · continuous uncertainty · real consequences

Autonomous driving is embodied intelligence at metropolitan scale: perception, prediction, planning, and motion in one loop. Through March 2026, Waymo reported 220.6 million rider-only miles and materially lower crash rates than its human-driver benchmarks. Those are company-reported figures, but the scale is the point: physical AI is already operating outside the lab.[17]

4. A day can become context.

A personal corpus is the record you choose to keep.

The phone put compute in your pocket. Watches and rings made passive self-data normal. The next layer is a private, permissioned record of the moments that matter: signals from your body, words from a conversation, and eventually the first-person context that helps an assistant understand your day.

01Signals

Health platforms already make a personal data layer practical: with permission, apps can read and write health and fitness data while the person remains in control of access and deletion.

From: sleep, heart rate, movement, and recovery.[21]

02Words

A recorder such as PLAUD can be intentionally started for a meeting, a thought, or an interview, then turn that moment into a searchable transcript and summary. It extends memory without claiming to record a life by default.

From: conversations you choose to capture.[22]

03First-person context

Meta’s Ego4D research shows the direction, not a released 24/7 consumer product: AI learning to locate an answer within past first-person video and understand the ongoing visual, audio, and motion context of daily life.

From: research into episodic memory and wearable perception.[23]

The useful question is not “what did it watch?”It is “what did I choose to preserve, and can I retrieve it?” A personal corpus could help answer a narrow question—what did I have for breakfast, what did we decide, where did I leave something—only from data its owner permitted it to keep.

Useful context must be visible, consented, controllable, and yours to remove.

5. Three curves finally meet.

The body, the model, and the training world are arriving at the same time.

Humanoid demos are not new. What changed is that movement can now be paired with models that generalize, simulation that multiplies experience, and hardware designed for deployment.

01A model that can generalize

RT-2 showed that web knowledge and robot demonstrations can live in one vision-language-action model. π0 pushed the idea toward a generalist policy across robots and dexterous tasks.

Change: actions can be learned as a reusable vocabulary, not hand-coded one task at a time.

02A world that can be rehearsed

Modern robotics stacks train and test in simulation, then transfer policies into physical machines. NVIDIA’s Isaac platform and GR00T package that loop for humanoid developers.

Change: a robot can accumulate more practice than the calendar alone permits.

03A body built for work

Electric actuators, better batteries, dexterous hands, and maturing supply chains are turning spectacular prototypes into products designed for factories and homes.

Change: the success metric becomes a useful shift—not a flawless demo.

GENERAL MODELSlanguage · vision · action SIMULATED EXPERIENCEpractice at machine speed CAPABLE HARDWAREbalance · hands · batteries physicalintelligence

6. The forecasts disagree. The direction does not.

Do not confuse a long-range scenario with a near-term market.

The headline numbers span different years and include different things. They are not apples-to-apples. Read them as evidence of widening possibility—not a promise that a robot will unload your dishwasher next Tuesday.

Goldman Sachs Research2035 · humanoid TAM
$38B
Morgan Stanley Research2050 · market + ecosystem
$5T+
Citi scenario2050 · potential market
$7T
nearer horizonbroader + longer horizon →

Goldman’s 2024 estimate projects a $38 billion total addressable market in 2035.[18] Morgan Stanley’s 2025 scenario reaches beyond robot sales into supply chains, maintenance, and support, and explicitly expects adoption to remain relatively slow until the mid-2030s.[19] Citi’s 2024 model puts almost 650 million humanoids and a potential $7 trillion market in 2050.[20]

The honest near-term thesis is smaller and more useful: factories first, bounded tasks first, assistance before autonomy, and a long learning curve inside the home.

7. Before intelligence works around us, it learns to work with us.

The first useful body may be yours.

That is why ambient systems matter. They let intelligence practice the hard parts of physical life—timing, context, memory, restraint—without pretending today’s hardware can do everything.

Stage one · now

Intelligence beside you

It listens when invited, remembers the context you choose to give it, and surfaces the next useful thing while you remain the actor.

  • Prepares you before a conversation
  • Keeps key context in your line of sight
  • Captures commitments without breaking presence
  • Leaves judgment and action with you
Stage two · emerging

Intelligence around you

It perceives the shared environment, plans within real constraints, and acts through a car, arm, mobile platform, or humanoid body.

  • Navigates roads and rooms
  • Manipulates ordinary objects
  • Learns new physical tasks
  • Earns autonomy through bounded reliability
Presence over interruption.The system should help you stay in the room, not pull you into another screen.
Assistance before spectacle.A quiet useful cue matters more than a cinematic demo.
Earned autonomy.The closer software gets to the physical world, the more reliability, consent, and restraint matter.

The future is not intelligence replacing daily life. It is intelligence learning how to fit inside it.

COS Glasses is the stage-one experiment.Context in your line of sight, with you still firmly in the loop.

8. Read the trail.

Primary sources first. Forecasts labeled as forecasts.

Product capabilities are the manufacturers’ stated claims unless independent performance data is explicitly noted. Market estimates use different definitions and horizons.

  1. Smithsonian Institution. “Can a Machine Think?” Archive essay on Kempelen’s Turk. repository.si.edu
  2. Museum of Modern Art. The Machine as Seen at the End of the Mechanical Age, catalogue entry on Vaucanson’s duck. moma.org
  3. DARPA. Grand Challenge innovation timeline: no 2004 finisher; five 2005 finishers. darpa.mil
  4. DARPA. Urban Challenge record, 2007. darpa.mil
  5. Apple. “Apple Reinvents the Phone with iPhone,” January 9, 2007. apple.com
  6. Apple. “Apple Unveils Apple Watch,” September 9, 2014. apple.com
  7. ŌURA. “The Origins and Evolution of Oura Ring,” February 10, 2025. ouraring.com
  8. Boston Dynamics. Commercial sales of Spot, June 16, 2020. bostondynamics.com
  9. Google DeepMind. “RT-2: New model translates vision and language into action,” July 28, 2023. deepmind.google
  10. NVIDIA. Project GR00T and Isaac robotics platform announcement, March 18, 2024. nvidia.com
  11. Boston Dynamics. “An Electric New Era for Atlas,” April 2024. bostondynamics.com
  12. Physical Intelligence. “π0: Our First Generalist Policy,” October 31, 2024. pi.website
  13. Figure. Series C and $39B post-money valuation announcement, September 16, 2025. figure.ai
  14. 1X. NEO launch, capabilities, pricing, and delivery plan, October 28, 2025. 1x.tech
  15. Apptronik. Series A funding update, February 11, 2026. apptronik.com
  16. Even Realities. Even G2 product capabilities and design. evenrealities.com
  17. Waymo. Safety Impact Data Hub, rider-only miles through March 2026. waymo.com
  18. Goldman Sachs Research. Humanoid robot TAM forecast to 2035, February 27, 2024. goldmansachs.com
  19. Morgan Stanley Research. “Humanoids: A $5 Trillion Market,” May 14, 2025. morganstanley.com
  20. Citi. “The Rise of AI Robots,” November 4, 2024. citigroup.com
  21. Apple Developer. HealthKit documentation: health and fitness data shared with apps only with the user’s permission, and managed by the user. developer.apple.com
  22. PLAUD. Voice recorder workflow: intentionally start a recording, then search the resulting transcript and summary. plaud.ai
  23. Meta AI. Ego4D research on first-person video understanding and episodic-memory queries; this is research, not a claim of a released always-on consumer camera. ai.meta.com

Research reviewed July 19, 2026. Dates above refer to the source publication or the data cutoff stated by the source. Company announcements describe company claims; forecast figures describe analyst scenarios, not observed outcomes.