Holo-M

Humanoid Loco-Manipulation With Discrete VLA Model

One vocabulary for language and whole-body actions, decoded in real time.

📄 arXiv GitHub · coming soon
01
Actions as language
To our knowledge the first discrete VLA for humanoid loco-manipulation: action tokens extend the VLM's own vocabulary.
02
Unified action tokenizer
Four body-part tokenizers (end-effector, body, hand, kinematics) share one discrete vocabulary.
03
Cross-embodiment data
Humanoid teleoperation, egocentric human video and simulation train one model.
04
Grouped discrete diffusion
Parallel within a body part, sequential across parts: 32 decoding steps instead of 208.
05
Progressive training
Four stages, from cross-embodiment pre-training to task specialists, on one backbone.
06
Best on SIMPLE
163 / 180 as specialist and 143 / 180 as generalist, the highest of all compared models.

Real-robot rollouts

Task 1

Serving coffee with a cart

“Grab the cart handle, push the cart forward and stop at the desk, then grasp the coffee cup and place it on the desk.”

real time
Task 2

Bowl to sink

“Pick up the green bowl from the counter, turn right, walk over to the sink and place the bowl in the sink.”

real time
Task 3

Bottle to trash bin

“Pick up the bottle, walk over to the trash bin and place the bottle in the trash bin.”

real time

Benchmark results

SIMPLE: 6 tasks × 3 randomisation levels × 10 episodes = 180. Generalist: one policy for all tasks; specialist: one per task.

Level-2 successes in simulation

Mobile P&P
XMovePick
XMoveBendPick
BendPick
Handover
Grasp
Specialist · overall out of 180
060120180Holo-M163Holo-M AR157Ψ0154ACT149DreamZero118π0.592GR00T N1.688DP70EgoVLA69InternVLA-M125H-RDT4
Specialist · cells: successes out of 10 at level 0 / 1 / 2; underlined = best in column, bold = ours
MethodXMovePickBendPickHandoverMobile P&PGraspXMoveBendPickOverall
Holo-M10/10/1010/10/810/8/108/9/89/9/98/8/9163/180
Holo-M AR10/10/79/10/910/8/98/8/89/10/78/7/10157/180
Ψ010/10/610/10/107/7/107/5/610/10/810/9/9154/180
ACT10/10/510/9/97/7/105/5/510/10/89/10/10149/180
DreamZero10/10/109/9/87/8/95/3/39/10/70/0/1118/180
π0.57/5/110/10/85/4/53/3/310/10/80/0/092/180
GR00T N1.610/10/77/7/61/3/30/0/09/9/74/4/188/180
DP3/3/210/8/63/2/44/0/08/9/80/0/070/180
EgoVLA0/1/27/5/80/4/30/0/010/10/73/5/469/180
InternVLA-M10/0/05/5/00/0/00/0/00/0/03/5/725/180
H-RDT0/0/20/0/10/1/00/0/00/0/00/0/04/180
Generalist · overall out of 180
060120180Holo-M143Holo-M AR138Ψ0114ACT79DreamZero38π0.528
Generalist · cells: successes out of 10 at level 0 / 1 / 2; underlined = best in column, bold = ours
MethodXMovePickBendPickHandoverMobile P&PGraspXMoveBendPickOverall
Holo-M9/7/39/9/87/7/108/9/68/9/89/8/9143/180
Holo-M AR9/5/810/8/89/6/85/6/78/9/87/8/9138/180
Ψ09/10/94/2/410/10/96/6/38/7/64/4/3114/180
ACT0/0/010/9/100/0/06/8/95/6/80/5/379/180
DreamZero0/0/00/0/07/7/60/0/07/5/60/0/038/180
π0.50/0/30/0/06/1/60/0/06/4/20/0/028/180
Holo-MHolo-M AR (autoregressive ablation)baselines

Method

Holo-M is a discrete vision-language-action model for humanoids. A unified tokenizer turns whole-body actions into four groups of tokens that extend the vocabulary of a vision-language model, so one model learns from humanoid teleoperation, egocentric human video and simulation. Grouped discrete diffusion decodes the tokens in real time. On the SIMPLE benchmark, Holo-M is the best generalist and the best specialist.

Full abstract

Vision-language-action (VLA) models using discrete action tokens have proven effective for controling robotic arms on manipulation tasks. For a humanoid, however, the whole-body action space – legs, torso, arms, and hands – is far higher-dimensional and heterogeneous, raising tokenization, training, and real-time inference challenges that the previous VLA models do not address. We present Holo-M, to our knowledge the first discrete VLA model for humanoid loco-manipulation that intrinsically exploits the language model by extending its vocabulary with action tokens. In this model, we devise a unified action tokenizer that decomposes the humanoid action space into four body-part-specific tokenizers – end-effector, body, hand, and kinematics – enabling training across drastically different embodiments and data sources, including humanoid teleoperation, ego-centric human video, and simulation. By extending the language model’s vocabulary with these action tokens, we avoid the knowledge-insulation problem inherent to the models that use separate continuous action experts. To meet real-time control requirements, we decode each body part’s action tokens through grouped discrete diffusion decoding, rather than using autoregression on the action tokens. We have conducted extensive experiments on the SIMPLE humanoid loco-manipulation benchmark, in which Holo-M achieves the highest success rates in both the generalist and specialist evaluations, leading the second best by significant margins. We will release all the code and model weights.

Overview of the Holo-M framework
Four-stage progressive training

Deployment

A new observation every 500 ms; the robot keeps executing the current chunk while the next one is decoded.

Unitree G1-comptwo Dex3-1 handshead camera 640×360 at 30 Hzone RTX 5090one-second action chunks
Real-time decoding schedule

Inference latency per one-second chunk

2 steps176 ms
4 steps276 ms
8 steps477 ms

mean latencyhold budget

Citation

@misc{shao2026holom,
  title         = {Humanoid Loco-Manipulation With Discrete VLA Model},
  author        = {Wenxin Shao and Siqi Chai and Kun Li and Kerou Zhang and Xinzhou Jiang and Wei Xu and Qiang Liu},
  year          = {2026},
  eprint        = {2609.35709},
  archivePrefix = {arXiv},
  primaryClass  = {cs.RO},
  url           = {https://arxiv.org/abs/2609.35709}
}