CogWAM: Aligning Semantic Cognition with World Action Modeling via Event-Driven Interfaces

A persistent Semantic State that a robot policy rewrites only when the task actually moves on — not on a fixed clock, and not by re-planning from scratch every step.

Sen Wang1,2, Liu Liu2,‡, Xinjiang Wang2, Zequn Chen2, Haoyi Jiang3, Taojun Ding2, Tingyang Xiao2, Zhizhong Su2,✉, Jie Wang4, Sanping Zhou1,✉
‡ Project Lead · ✉ Corresponding Authors
1Xi'an Jiaotong University · 2Horizon Robotics · 3Huazhong University of Science and Technology · 4University of Science and Technology of China
CogWAM overview: event-triggered semantic-state updates, the execution loop, and headline results on RoboDojo, BiCoord, and real-world deployment.
Event-triggered Semantic State updates avoid both the redundant recomputation of synchronous updating and the staleness of periodic asynchronous updating, and translate into consistent gains across simulation and real-world deployment.
01 / Idea

What is CogWAM?

Long-horizon manipulation needs more than the current camera frame: a policy has to remember what it already finished and what it is doing now. CogWAM keeps that as an explicit, persistent Semantic State — completed events plus the active subtask — and updates it only when a lightweight KEEP / UPDATE decision says the task has actually moved on. That state conditions two query streams, WORLD and ACTION, that share one backbone but specialize toward future-latent prediction and control respectively. Future prediction is a training-time objective only: at inference the World stream is dropped, and actions are generated directly from the observation and the maintained Semantic State — not from an imagined future.

02 / Method

Event-driven semantic interface

A five-minute tour of the two mechanisms that make the Semantic State useful: how it decides when to change, and how it reaches the physical model.

CogWAM architecture: event-driven semantic planning producing a persistent Semantic State, which conditions World and Action queries inside a shared multi-modal self-attention backbone.
Figure 1. CogWAM architecture. Left: the Semantic State is read from the previous step, and a KEEP/UPDATE decision determines whether it advances. Right: WORLD and ACTION queries condition on the same Semantic State and current DINO features but attend through separate cross-attention and FFN blocks — future-latent prediction is a training-time branch only.

Event-Triggered Semantic State

At every replanning step the model predicts KEEP or UPDATE from a single forward pass. KEEP costs nothing further. UPDATE additionally generates a memory increment and the new active subtask, autoregressively, and only then. The state can persist across many action chunks — it changes when the task does, not on a fixed clock.

Progress-Conditioned World & Action Queries

Two learnable query sets are appended after the observation, instruction, and Semantic State in one causal sequence. Both read the same task-progress context, but their hidden states are routed to different objectives — WORLD toward future multi-view DINO latents, ACTION toward the control chunk — so the two branches share semantics while keeping objective-specific representations.

Boundary-Aware Training

Semantic transitions are rare, so training oversamples UPDATE events and the near-boundary "hard KEEP" cases that look almost like a transition. A scheduled probability also swaps in the model's own predicted state during training, closing the gap to closed-loop inference where the annotated state is never available.

03 / Real Robot

Real-world deployment

Six multi-stage tabletop tasks on a dual-arm setup, plus four controlled distribution shifts, evaluated closed-loop.

Real-world setup: dual-arm AgileX PIPER platform with head-mounted and wrist-mounted RealSense D435 cameras, and the six basic tasks plus four generalization settings.
Real-world setup: a dual-arm AgileX PIPER platform with two 6-DoF manipulators and parallel grippers, using synchronized RGB from one head-mounted and two wrist-mounted Intel RealSense D435 cameras. Six basic tasks (top) and four generalization settings — spatial location, object appearance, distractors, novel objects (bottom-right pair shown for two tasks).
Setup: PiPER dual-arm · head + 2× wrist RealSense D435 · closed-loop control
04 / RoboDojo

Simulation rollouts

RoboDojo evaluates 42 tasks across five capability dimensions. Clips below are real rollouts of the checkpoint below — not cherry-picked renders.

Checkpoint: dino_multilayer_50k · step 50,000 · no prior robot-data pre-training
05 / Results

Reported results

Numbers below are exactly as reported in the paper. CogWAM uses no prior embodied robot-data pre-training in any of these tables.

RoboDojo — multi-task evaluation

Each cell is Score / Success Rate (%), averaged over five capability dimensions.

BiCoord — long-horizon bimanual manipulation

Six of 18 tasks with the longest expert trajectories are shown; Average is over all 18. Each cell is Stage-wise Success Rate / Success Rate (%).

Real-world evaluation

Success counts, 20 trials per basic task (120 total) and 30 trials per generalization dimension across all six tasks (120 total).

Semantic State update strategy — deployment cost

Measured on the real robot at a 30 Hz control loop with a 333 ms replanning budget (NVIDIA RTX 5090).

Ablation: w/o Semantic Update vs Synchronous vs Periodic Async vs KEEP/UPDATE, on RoboDojo Long-Horizon/Memory and on BiCoord.
Effect of the Semantic State and its update strategy on RoboDojo (Long-Horizon, Memory) and BiCoord (18-task average). KEEP/UPDATE is CogWAM's event-triggered strategy.
This project page reports results for the RoboDojo, BiCoord, and real-world experiments exactly as printed in the paper. It does not re-aggregate or re-score any rollout, and does not present any number the paper does not report.
06 / Citation

Citation

@article{wang2026cogwam,
  title={CogWAM: Aligning Semantic Cognition with World Action Modeling via Event-Driven Interfaces},
  author={Wang, Sen and Liu, Liu and Wang, Xinjiang and Chen, Zequn and Jiang, Haoyi and Ding, Taojun and Xiao, Tingyang and Su, Zhizhong and Wang, Jie and Zhou, Sanping},
  journal={arXiv preprint arXiv:2609.37721},
  year={2026}
}