WRITINGS/HUMANOID-GAMES-PROCUREMENT-ROUND

The Humanoid Games are a procurement round

·11 MIN READ

Two thousand robots ran, fought, and folded laundry in Beijing last week, and several of them caught fire doing it. The events that were held tell you what China thinks it can sell by 2029. The events that were not held tell you what nobody has solved yet, and that list is the more useful document.

01The weekend

The second World Humanoid Robot Games ran August 22 to 26 at Beijing's National Speed Skating Oval. Some numbers, because the numbers are the story: 2,056 robots, 666 teams, sixteen countries, around 1,300 matches over five days. That is roughly 260 a day, run in parallel across a floor built for the 2022 Winter Olympics. Last year's edition had 500 robots and 280 teams, so entries tripled in twelve months, in a category that did not exist as a category three years ago.

Roughly fifty events, split almost evenly between sports and work. On the sports side: sprints, distance, hurdles, high jump, football, gymnastics, boxing, table tennis, dance. On the work side: retail, food service, assembly, warehousing, hospitality, household chores, firefighting, search and rescue. The dexterity events were oddly specific in a way that repays attention. Power tool assembly. Powder weighing. Bricklaying. Picking beans out of a dish with tweezers.

The headline result was a 100 meters run in 8.86 seconds, well inside Usain Bolt's 9.58. A 400 in 38.16. A high jump at 2.88 meters, forty centimeters over the human record. Worth noting that at least three different times circulated for that hundred meters depending on which heat and which press release you read, which is its own small data point about how carefully the results were kept.

Also, in the same five days, robots tripped, came apart on the track, and in at least one case caught fire. Those clips traveled further than any of the records did. There is no column on the results sheet for a machine that stopped working.

Both sets of facts are true and neither is the interesting part.

02The teleoperation line

The interesting part is buried in the rules.

Some events required full autonomy. No human in the loop, no operator with a controller: flat-ground running, football, gymnastics, dance, boxing, table tennis. Other events permitted teleoperation. Hurdles. Jumps. Weightlifting. Kickboxing.

Read that split again, because it is the most honest capability disclosure anyone published this year. A machine can sprint a hundred meters faster than the fastest human who ever lived, entirely on its own, and then needs a person on a controller to get over a hurdle.

That reads as a contradiction and is not one. It is a fairly precise map of where the frontier sits. Flat-ground running is a controls problem on a known surface with a fixed gait, and controls problems on known surfaces are solved. Actuators are cheap and strong now. Batteries are adequate. The 8.86 is a hardware achievement dressed as an athletic one.

A hurdle is a different animal. It requires seeing an obstacle, estimating its distance while your own head is oscillating, committing to a takeoff point two strides out, and being unrecoverably wrong if the estimate is off by ten centimeters. Closed-loop perception at speed. Kickboxing and weightlifting are worse again, because they add contact with something that pushes back and does not hold still.

So the line between the autonomous events and the teleoperated ones is the line between dynamics you can model in advance and dynamics you have to resolve in the moment. Everything commercially interesting about humanoid robots lives on the wrong side of that line. Nobody is paying for a machine that can sprint. They are paying for a machine that can pick a part out of a bin it has not seen, in a lighting condition nobody planned for, four thousand times a shift, without dropping enough of them to matter.

The teleoperation list is the todo list.

03What was not on the schedule

Benchmarks reveal their designers. You put on the schedule what you believe your entrants can survive, and the absences say more than the entries.

Nothing on stairs. Nothing on unstructured terrain, so no gravel, no cabling underfoot, no wet floor. Nothing on working alongside a human in shared space. Nothing on following a spoken instruction, which is remarkable in the year that language models became the default interface to everything else. And nothing, anywhere, on the two measures a buyer actually cares about: what it costs to complete the task, and how long the machine runs before it needs a person.

The robot that caught fire is not an embarrassment. It is the only reliability data the Games produced, and it produced it by accident.

Consider the household service events in that light. Household is the hardest environment that exists: unmapped, cluttered, poorly lit, full of soft objects and a child and a dog, and it changes every day. A humanoid that can run a hospitality scenario in a speed skating oval and a humanoid that can operate in an actual house are separated by roughly the entire remaining problem. Including household on the schedule at this stage tells you the schedule has a promotional job to do alongside its measurement job.

Which is fine. Every benchmark suite in a young field is half marketing. The trick is knowing which half you are reading.

04The 96 percent

Sixteen countries entered. About 96 percent of the entries were Chinese.

John Koetsier's read in Forbes is that the event schedule functions as a market map, and that the work-scenario categories are a public statement of where China expects humanoids to be deployed first. That is right, and I would push it one step further. This is not primarily a competition. It is a supplier qualification round with a starting pistol.

Look at what it accomplishes for a state that has named embodied intelligence a strategic industry. It produces a ranked, public list of domestic vendors across fifty capability categories. It creates a shared spec that six hundred teams now optimize against, which is how you get component standardisation without writing a standard. It generates a spectacular volume of comparable performance data. And it does all of this in front of cameras, at a venue built for the Winter Olympics, in a format that reads to a general audience as national capability rather than industrial policy.

Wang Peng of the Beijing Academy of Social Sciences put it to the press this way: “Last year it was about whether they could run to the finish line. This year it's about whether they can get the job done.” No sports commentator talks like that. That is how a procurement officer talks.

The Western programs are running the same play with the volume down. Figure has Figure 03 units on the line at BMW's Spartanburg plant, billed to the customer at around twenty-five dollars per robot-operating-hour. Boston Dynamics started shipping a mass-production Atlas this year and sold out the 2026 allocation. Tesla has somewhere between a thousand and twelve hundred Optimus units working inside its own factories and has sold precisely zero of them to anyone, which is the same move Netflix made with post-production tooling: keep the cost advantage, decline the software business. Unitree shipped more than five thousand units last year and is aiming at ten to twenty thousand this year, at roughly a tenth of what a Western unit costs.

Nobody at the Games published a cost per completed task. Figure published a price per robot-hour. One of those two numbers is a product and the other is a demo.

05The number nobody printed

Here is the trap, and it is a trap the software AI world walked into first and only recently walked out of. In 2024, every frontier lab published benchmark charts. Bar after bar, a few points of improvement on graduate-level reasoning, competitive coding, multilingual comprehension. The charts went up and to the right and were mostly honest. They were also almost useless for deciding whether to put a model into production, because nothing on them measured the two things that determined whether a deployment survived contact with a real business: cost per completed task, and the failure rate on the long tail of inputs nobody thought to test.

The teams that shipped working systems in 2025 and 2026 were not the ones that picked the model at the top of the chart. They were the ones that built an eval on their own data, priced the whole loop including retries and human review, and found out that a cheaper model with tighter scaffolding beat the leaderboard winner on the only axis that paid.

The Humanoid Games are the 2024 benchmark chart with actuators. Eight point eight six seconds is a beautiful number on an axis nobody buys on.

The axis people buy on has four values.

  • Cost per completed task, all in. Amortized hardware, power, maintenance, the integration engineer, and the teleoperator you still quietly employ.
  • Mean time between interventions. Where an intervention is any moment a human has to touch the machine.
  • Success rate on the ugly tail. The bins that were loaded wrong and the box that arrived crushed.
  • Cycle time against the human baseline. Which for most warehouse tasks is a bar that is a lot higher than people expect, because an experienced picker is fast.

None of the four appeared on a scoreboard in Beijing.

06Where this goes

What follows is dated so it can be checked later. The order these things happen in is the part I hold with any confidence; the calendar is a guess.

2027

The third Games adds stairs and adds at least one reliability measure, probably a duration or endurance format, because the absence became too conspicuous to keep. Watch the teleoperation list: if hurdles move to the autonomous column, closed-loop perception at speed has broken open and everything below moves eighteen months earlier. My guess is hurdles go autonomous and adversarial contact events, kickboxing and anything with a resisting opponent, do not. The first credible third-party cost-per-task figures get published this year and they are ugly, somewhere in the forty to eighty dollars an hour range fully loaded for tasks a person does for eighteen. Deployments continue anyway, because early deployments are bought out of R&D budgets, not operations budgets, and R&D budgets do not care about unit economics.

2028

The humanoid form factor starts losing its own competitions. In warehousing and retail scenarios, the fastest entries are wheeled bases with two good arms, because legs are a tax you pay only for stairs and for spaces built around human hips. Legs stay in construction, inspection, and anywhere the environment cannot be modified. The market splits the way cloud split: a cheap Chinese hardware tier under twenty thousand dollars a unit, and a Western robot-hour service tier where nobody owns the machine. The service tier wins the enterprise, because a plant manager can sign a per-hour number and cannot sign a capex line for an unproven machine. Somebody gets seriously hurt this year, there is a lawsuit, and the insurance market rather than the regulator sets the actual deployment pace. That is how it went with autonomous vehicles and it is how it will go here.

2029

The moat turns out to be data, and specifically the data nobody could scrape. Ten thousand units running eight hours a day, three hundred days a year, is twenty-four million hours of contact-rich manipulation, force feedback, grasp failure, and recovery. That corpus does not exist on the internet and cannot be synthesized convincingly, and the first fleet to accumulate it compounds away from everyone else. This, not the medals, is what the Games are actually recruiting for: six hundred teams pointed at the same task list generate a shared data substrate. Deployment concentrates hard into three verticals, warehouse tote handling, commercial food preparation, and back-of-house hospitality, and stays out of the household almost entirely.

2030

The word humanoid falls out of procurement language the way cloud did, and for the same reason: it stops describing a decision anyone makes. You buy picking hours. By then the bottleneck has moved off intelligence entirely and onto hands, which wear out, and service networks, which do not exist yet. The company that wins this decade may well turn out to be a maintenance business.

07The ways this is wrong

Three things would break the timeline, and I hold all of the above loosely enough to say so.

A general manipulation model that transfers zero-shot across bodies, so that a policy trained on one robot's arm works on a different manufacturer's arm without retraining, would collapse the 2028 and 2029 lines into 2027. There are early signs of this and I do not know how to price them. If it lands, everything gets faster than anyone's plan.

Export controls in either direction remove the cheap hardware tier from Western markets, and everything above stretches by about three years. That is a customs question rather than a technology one, and it is probably the single largest variance in the forecast.

And a fatality in year one, in a warehouse, on video, freezes enterprise deployment for eighteen months regardless of what the machines can do. Capability has never been what gates this class of technology. Liability has.

08If you are not buying robots

Most people reading this are not going to buy a humanoid. The transferable part is the reading method, and it applies directly to the software agents already in your business.

The Games are what a field looks like when it is still measuring capability because capability is the thing it can measure. The transition from demo to deployment happens at the exact moment the field switches to measuring cost per completed task and rate of human intervention, and that transition is worth watching for, because it is when the buying starts.

Agentic software crossed that line about a year ago. The teams still comparing benchmark scores are running demos. The teams that instrumented cost per completed task, tracked how often a human had to step in, and priced the retries are the ones with systems in production that survived their second quarter. Same discipline, five years earlier, less on fire.

If you want to know which of your workflows is actually ready for an agent, the question is not what the model can do. It is what one completed unit of that work costs you today, and how often it goes wrong in a way somebody has to fix. Tell us what the work is and what it costs you, and we will tell you whether there is a system in it.