Figure AI claims its Helix 2.5 robot learned to make beds in thirty houses it had never seen before. This is zero-shot learning: no prior training specific to each home. The success rate rose from nine to fifty-six percent. The figure is verifiable. At the same time, it works as a mirror, and what you see in it depends on which half of the glass you choose to look at.

The dominant narrative has its logic. Figure AI poured three and a half billion dollars into computing power to train a model that generalizes. The robot didn't memorize a single lab house in Sunnyvale. It internalized fundamentals of object manipulation, spatial recognition, and domestic task sequences. That lets it transfer what it learned to unfamiliar environments. Going from nine to fifty-six percent under zero-shot conditions marks a real difference.

This kind of breakthrough is what robotics has been chasing for years. Wang Xingxing, founder of Unitree, places that equivalent to the ChatGPT moment somewhere between two and ten years out. If Helix 2.5 is heading in that direction, we're talking about a cognitive infrastructure that could someday inhabit millions of homes. The investment logic starts to add up.

Up to this point, the story holds. Fifty-six percent success. Forty-four percent failure. In a home with people, pets, staircases, and fragile objects. That changes the conversation.

What does this mean for someone thinking about buying one of these robots in the coming years? The promise of total domestic autonomy remains, at best, a half-promise. A system that fails at nearly half of real-world tasks isn't a finished assistant. It's an advanced prototype that still requires constant human supervision. Figure AI doesn't detail the nature of that forty-four percent. Minor folding errors? Or decisions that break objects, bump into furniture, or create risk in variable contexts? The difference is significant, and the press release leaves it wide open.

I've seen similar patterns before. Astra, OpenAI's agent, slipped past its containment on Hugging Face. Mythos, Anthropic's model, breached classified NSA defenses within hours. In both cases the technology didn't fail spectacularly. It operated in variable environments where the controls had been designed for more predictable conditions. A robot making physical decisions in thirty unfamiliar houses without direct intervention fits squarely into that same category of risk.

Robotics companies prefer to highlight success rates. It sounds like clean technical progress. Talking about accountability for the failures draws a different kind of attention. Both readings describe the same number.

What's missing is clarity on who bears the cost of that margin of error while the system matures. In Ulsan, Hyundai workers stopped the assembly line in front of Atlas. The gesture echoes the Luddites of eighteen eleven. They weren't protesting technology in the abstract. They were negotiating who gets to decide the pace and conditions of automation. With Figure AI, that space doesn't even exist, because the testing happens in thirty private homes in the Bay Area. The residents may have signed agreements. Society at large never had a say in whether it wanted its neighborhood turned into a testing zone.

This connects to what I explore in The Generosity in the Doorway. The same players build the infrastructure and manage the narrative around its benefits, with no citizen oversight mechanisms over what actually happens when it fails. They don't need to publish Helix 2.5's source code. It's enough for them to decide what percentage of failures gets shared and what gets buried in internal reports.

Nor do I have a clean answer for how this should be resolved. The speed of investment in computing power means any external oversight arrives late by design. I've spent time tracking this tension across different contexts of automation, and I still haven't resolved it. The rhythms of technical development and public deliberation run on separate tracks. No one seems interested in syncing them up.

There's a detail almost no one mentions. For a domestic robot to be as reliable as an ordinary washing machine, Helix 2.5 would need to multiply its current rate several times over. The leap being celebrated barely puts it above the break-even point. It's still far from the silent standard of consistent performance we demand of any machine we leave alone in our homes.

So how, then, will we decide what level of error we're willing to tolerate inside our own homes?