EVO-WAM

Evolving World Action Models
through Video-Action Verification

Explore
Shiyang Zhou1,2,‡Xionghao Wu3,‡Wenbo Li4,†Shenghe Zheng5Jiyao Zhang6Songsong Yu7Yijun Yang5,8Jianhui Liu9Haoze Sun10Senqiao Yang11Li Jiang12,2Jingyong Su1Haoyang Huang4Zhuotao Tian1,2,*
1HITSZ2SLAI3THU4JD5HKUST6PKU7SJTU8HKUSTGZ9HKU10UBC11CUHK12CUHKSZ

shiyangzhou@stu.hit.edu.cnfenglinglwb@gmail.comtianzhuotao@hit.edu.cn

‡ Equal contribution† Project lead* Corresponding author

01 / RESULTS

Imagination that
improves performance.

Mean task success (%) · R0–R4

SIMULATION Cosmos3

26.9→68.0%

+41.1 percentage points

Cosmos3 simulation success rates: R0 26.9%, R1 58.3%, R2 66.6%, R3 63.6%, R4 68.0%. A gain of 41.1 percentage points from R0 to R4.

7 unseen RoboTwin 2.0 tasks

SIMULATION DreamZero

28.5→46.4%

+17.9 percentage points

DreamZero simulation success rates: R0 28.5%, R1 36.5%, R2 42.3%, R3 45.1%, R4 46.4%. A gain of 17.9 percentage points from R0 to R4.

7 unseen RoboTwin 2.0 tasks

REAL WORLD Cosmos3

20.0→76.7%

+56.7 percentage points

Cosmos3 real-robot success rates: R0 20.0%, R1 60.0%, R2 76.7%, R3 73.3%, R4 76.7%. A gain of 56.7 percentage points from R0 to R4.

3 unseen long-horizon composite tasks

Results & evaluation ↗

02 / THE STORY

See self-evolution unfold.

Imagine the task.
Verify the experience.
Improve the policy.

EVO-WAM learns from its own generated video–action trajectories, retaining prefixes that pass task-completion and action-consistency checks.

Self-training uses no additional expert demonstrations or external execution of candidate actions.

Method

03 / IN ACTION

From imagined experience
to better execution.

Place Ducks

BeforeR1
Task incompleteDownload
AfterR2
Task completed Download

Separate trials, similar layouts. Inference pauses removed.

04 / THE METHOD

Better experience.
One round at a time.

Duck placement

Candidate 1
Candidate 2
Candidate 3
Candidate 4
Full framework EVO-WAM framework: autoregressive video-action-state generation, visual goal evaluation and inverse-dynamics consistency verification, followed by supervised fine-tuning with accepted prefixes and original data.

Paper, Figure 2.

05 / RESOURCES

Explore EVO-WAM.

Paper
arXivComing soon
CodeComing soon