Humanoid Loco-Manipulation With Discrete VLA Model
One vocabulary for language and whole-body actions, decoded in real time.
One vocabulary for language and whole-body actions, decoded in real time.
“Grab the cart handle, push the cart forward and stop at the desk, then grasp the coffee cup and place it on the desk.”
“Pick up the green bowl from the counter, turn right, walk over to the sink and place the bowl in the sink.”
“Pick up the bottle, walk over to the trash bin and place the bottle in the trash bin.”
SIMPLE: 6 tasks × 3 randomisation levels × 10 episodes = 180. Generalist: one policy for all tasks; specialist: one per task.
Level-2 successes in simulation
| Method | XMovePick | BendPick | Handover | Mobile P&P | Grasp | XMoveBendPick | Overall |
|---|---|---|---|---|---|---|---|
| Holo-M | 10/10/10 | 10/10/8 | 10/8/10 | 8/9/8 | 9/9/9 | 8/8/9 | 163/180 |
| Holo-M AR | 10/10/7 | 9/10/9 | 10/8/9 | 8/8/8 | 9/10/7 | 8/7/10 | 157/180 |
| Ψ0 | 10/10/6 | 10/10/10 | 7/7/10 | 7/5/6 | 10/10/8 | 10/9/9 | 154/180 |
| ACT | 10/10/5 | 10/9/9 | 7/7/10 | 5/5/5 | 10/10/8 | 9/10/10 | 149/180 |
| DreamZero | 10/10/10 | 9/9/8 | 7/8/9 | 5/3/3 | 9/10/7 | 0/0/1 | 118/180 |
| π0.5 | 7/5/1 | 10/10/8 | 5/4/5 | 3/3/3 | 10/10/8 | 0/0/0 | 92/180 |
| GR00T N1.6 | 10/10/7 | 7/7/6 | 1/3/3 | 0/0/0 | 9/9/7 | 4/4/1 | 88/180 |
| DP | 3/3/2 | 10/8/6 | 3/2/4 | 4/0/0 | 8/9/8 | 0/0/0 | 70/180 |
| EgoVLA | 0/1/2 | 7/5/8 | 0/4/3 | 0/0/0 | 10/10/7 | 3/5/4 | 69/180 |
| InternVLA-M1 | 0/0/0 | 5/5/0 | 0/0/0 | 0/0/0 | 0/0/0 | 3/5/7 | 25/180 |
| H-RDT | 0/0/2 | 0/0/1 | 0/1/0 | 0/0/0 | 0/0/0 | 0/0/0 | 4/180 |
| Method | XMovePick | BendPick | Handover | Mobile P&P | Grasp | XMoveBendPick | Overall |
|---|---|---|---|---|---|---|---|
| Holo-M | 9/7/3 | 9/9/8 | 7/7/10 | 8/9/6 | 8/9/8 | 9/8/9 | 143/180 |
| Holo-M AR | 9/5/8 | 10/8/8 | 9/6/8 | 5/6/7 | 8/9/8 | 7/8/9 | 138/180 |
| Ψ0 | 9/10/9 | 4/2/4 | 10/10/9 | 6/6/3 | 8/7/6 | 4/4/3 | 114/180 |
| ACT | 0/0/0 | 10/9/10 | 0/0/0 | 6/8/9 | 5/6/8 | 0/5/3 | 79/180 |
| DreamZero | 0/0/0 | 0/0/0 | 7/7/6 | 0/0/0 | 7/5/6 | 0/0/0 | 38/180 |
| π0.5 | 0/0/3 | 0/0/0 | 6/1/6 | 0/0/0 | 6/4/2 | 0/0/0 | 28/180 |
Holo-M is a discrete vision-language-action model for humanoids. A unified tokenizer turns whole-body actions into four groups of tokens that extend the vocabulary of a vision-language model, so one model learns from humanoid teleoperation, egocentric human video and simulation. Grouped discrete diffusion decodes the tokens in real time. On the SIMPLE benchmark, Holo-M is the best generalist and the best specialist.
Vision-language-action (VLA) models using discrete action tokens have proven effective for controling robotic arms on manipulation tasks. For a humanoid, however, the whole-body action space – legs, torso, arms, and hands – is far higher-dimensional and heterogeneous, raising tokenization, training, and real-time inference challenges that the previous VLA models do not address. We present Holo-M, to our knowledge the first discrete VLA model for humanoid loco-manipulation that intrinsically exploits the language model by extending its vocabulary with action tokens. In this model, we devise a unified action tokenizer that decomposes the humanoid action space into four body-part-specific tokenizers – end-effector, body, hand, and kinematics – enabling training across drastically different embodiments and data sources, including humanoid teleoperation, ego-centric human video, and simulation. By extending the language model’s vocabulary with these action tokens, we avoid the knowledge-insulation problem inherent to the models that use separate continuous action experts. To meet real-time control requirements, we decode each body part’s action tokens through grouped discrete diffusion decoding, rather than using autoregression on the action tokens. We have conducted extensive experiments on the SIMPLE humanoid loco-manipulation benchmark, in which Holo-M achieves the highest success rates in both the generalist and specialist evaluations, leading the second best by significant margins. We will release all the code and model weights.
A new observation every 500 ms; the robot keeps executing the current chunk while the next one is decoded.

Inference latency per one-second chunk
mean latencyhold budget
@misc{shao2026holom,
title = {Humanoid Loco-Manipulation With Discrete VLA Model},
author = {Wenxin Shao and Siqi Chai and Kun Li and Kerou Zhang and Xinzhou Jiang and Wei Xu and Qiang Liu},
year = {2026},
eprint = {2609.35709},
archivePrefix = {arXiv},
primaryClass = {cs.RO},
url = {https://arxiv.org/abs/2609.35709}
}