Omni-modal humanoid motion

Movens-ZeroScaling Humanoid Motion Foundation Models with Large-Scale Human Videos

Wendong Bu1,3,4,*Tinghong Ye1,*Qizhou Wang2,3,*Jiacheng Li5Zhiqi Ge1Yuze Lin1Zhongqi Yue6

Yi Su2Jingyu Wang1Kai Shen1Jingdong Liu1Chengyuan Deng1Xiaoxuan Chen1Xin Tan7

Jiaming Ji8Wenqiao Zhang1Chenghong Jiang4Tuoyu Li1Siliang Tang1,3Jun Xiao1

Tat-Seng Chua9Yueting Zhuang1Xingxing Wang2Juncheng Li1,3,4,✉

1Zhejiang University2Unitree Robotics3Joint Laboratory of Embodied Intelligence, ZJU & Unitree Robotics

4Bit Oasis5HiThink Research6Microsoft Research7East China Normal University

8Peking University9National University of Singapore

* Equal contribution✉ Corresponding author

Watch full demoExplore Interactive 3D ↗arXivPaper Coming SoonDataset Movens-5MBenchmark Movens-Intent
Code GitHub ↗
Explore the project
15MHuman motions
5MExecutable motions
10K+Hours of behavior
1.08B+Motion frames
4Input modalities

01 / Full film

See Movens-Zero in motion.

A five-minute real-robot film spanning live and Internet video, spoken instructions, music, and text prompts.

Full demo · 05:16Watch on YouTube

02 / Why Movens-Zero

Human behavior, translated for humanoids.

Motion capture is precise but difficult to scale. Internet video is diverse, yet reconstructed motion can be noisy, physically invalid, or impossible for a humanoid to execute. Movens-Zero bridges that gap.

01

Web-scale data engine

Recover, repair, retarget, and verify motion from raw video.

02

Omni-modal supervision

Text, RGB video, speech, and music share one motion space.

03

Closed-loop generation

Replan from the latest condition and robot state in real time.

Movens-Zero overview showing multimodal inputs, continuous humanoid behavior, dataset scale, and model scaling
Movens-ZeroInterleaved multimodal intent becomes continuous whole-body behavior.

03 / Dataset

Human behavior at humanoid scale.

The Movens dataset is built from raw Internet video. Its data engine filters and annotates clips, reconstructs human motion, repairs physical artifacts, retargets behavior, and verifies execution before training.

297Fine-grained motion types
16Scene groups
6Physical checks
4Tracking policies

Data engine

From raw video to executable motion

A five-stage data engine starts from raw Internet video, curates coherent clips, filters motion and visual quality, aligns multimodal semantics, recovers and retargets human motion, and finally refines it through closed-loop execution. Each stage narrows the gap between visually plausible human behavior and motion a humanoid can reliably track.

  1. 01Clip generationCoherent clips from segmented and filtered web video
  2. 02Quality filteringHuman-motion and visual-quality screening
  3. 03Semantic annotationVideo, text, speech, and music alignment
  4. 04Motion recoveryHuman reconstruction, repair, and retargeting
  5. 05Physical refinementClosed-loop optimization and execution verification
Five-stage Movens data pipeline from web video curation to physical humanoid execution verification

Inside one Movens record

From a human video to executable humanoid data.

Each showcase example pairs the source clip with recovered human motion, a retargeted humanoid reference, and a tracking rollout. Together, the eight records illustrate diverse motions, scenes, and objects.

01Original video02SMPL rendering03G1 retargeted04G1 tracking

Dataset profile

Scale is visible in the distribution.

Four views summarize temporal scale, behavior taxonomy, scene coverage, and the relationship between motions and their contexts.

Profile 0110,000+ hours · 1.08B+ frames

Duration at scale

Distribution of Movens clip durations

A median duration of 16.0 seconds preserves meaningful temporal behavior while supporting web-scale processing.

Profile 02297 motion types · 7 families

A hierarchical motion vocabulary

Hierarchical distribution of Movens motion types

Performance, dance, daily activity, locomotion, fitness, gesture, and sport form a broad long-tail taxonomy.

Profile 0316 scene groups

Behavior in context

Distribution of Movens scene groups

Studio, home, outdoor, stage, fitness, and public settings keep motion grounded in diverse visual contexts.

Profile 04Motion × scene association

Diversity beyond raw counts

Association between Movens motion and scene groups

The same motion families appear across multiple environments, reducing dependence on a single visual setting.

Public release

Explore Movens-5M.

The public Movens-5M release is available on Hugging Face. The examples above illustrate the source-to-execution data structure independently of any dataset viewer row.

Open dataset on Hugging Face

04 / Method

One model, many ways to express intent.

A frozen omni-modal context encoder conditions a shared flow-based action expert. The latest robot state keeps generation grounded in what the hardware is actually doing.

Movens-Zero architecture with multimodal context encoder, flow-based motion transformer, robot motion representation, and tracking controller
Shared prior

Task-mixed flow matching

Every modality trains one reusable motion prior instead of an isolated policy.

Responsive control

Multimodal real-time chunking

Overlapping predictions enable continuous replanning without abrupt transitions.

Physical grounding

State-conditioned generation

Each chunk uses the newest robot state before a high-frequency controller tracks it.

05 / Capability showcase

Five ways to express intent. One shared motion model.

Explore a curated set of real-robot demonstrations across live video, Internet video, speech, music, and text.

Video / Live

Real-robot demonstrations

Follow a performer as the motion unfolds.

Movens-Zero uses the current video stream and the robot’s latest state to continuously replan executable whole-body motion.

Real-time video input · 01
Real-time video input · 02
Real-time video input · 03
Real-time video input · 04
Real-time video input · 05
Real-time video input · 06

Video / In the wild

Real-robot demonstrations

Bring human motion from the Internet onto the robot.

In-the-wild references provide locomotion, gesture, rhythm, and direction changes that the model translates into continuous humanoid behavior.

Internet video input · 01
Internet video input · 02
Internet video input · 03
Internet video input · 04
Internet video input · 05
Internet video input · 06

Audio / Spoken instruction

Real-robot demonstrations

Say what to do—and how to do it.

Spoken instructions specify actions, body side, repetition, order, and duration, from left-leg lunges to alternating waves.

Speech input · 01“Open your arms wide then cross them in front of your chest”
Speech input · 02“Do ten left-leg lunges”
Speech input · 03“Salute with your left hand”
Speech input · 04“Strike a few bodybuilding poses”
Speech input · 05“Alternate waving your left and right hands up to your head”
Speech input · 06“Wave your left hand for seven seconds”

Audio / Music

Real-robot demonstrations

Let sound shape the movement.

Music conditions the timing and character of coordinated footwork, balance shifts, turns, and upper-body motion.

Music input · 01
Music input · 02
Music input · 03
Music input · 04
Music input · 05
Music input · 06

Language / Text

Real-robot demonstrations

Turn a prompt into a whole-body behavior.

Natural-language prompts can describe an action, a path, or a style—from counted squats and counterclockwise walking to salsa and zombie motion.

Text input · 01“Dance some hip-hop”
Text input · 02“Do three squats”
Text input · 03“Perform a salsa dance”
Text input · 04“Walk counterclockwise”
Text input · 05“Walk forward and wave”
Text input · 06“Act like a zombie”

06 / Scaling study

Scaling data, models, and execution together.

The study varies data from 1% to 100% of Movens-5M and the action expert from 95.6M to 1.30B parameters under a shared execution protocol.

Dataset fraction
1%3%10%30%100%
Action expert
95.6M277M717M1.30B

Quantitative tables in the current technical report are still being finalized; no unreleased accuracy claims are shown here.

07 / Citation

Build on Movens-Zero.

Paper and model links will be added when the technical report is released. Code and the public dataset are available now.

@article{bu2026movenszero,
  title   = {Movens-Zero: Scaling Humanoid Motion Foundation Models with
             Large-Scale Human Videos},
  author  = {Wendong Bu and Tinghong Ye and Qizhou Wang and Jiacheng Li and
             Zhiqi Ge and Yuze Lin and Zhongqi Yue and Yi Su and Jingyu Wang and
             Kai Shen and Jingdong Liu and Chengyuan Deng and Xiaoxuan Chen and
             Xin Tan and Jiaming Ji and Wenqiao Zhang and Chenghong Jiang and
             Tuoyu Li and Siliang Tang and Jun Xiao and Tat-Seng Chua and
             Yueting Zhuang and Xingxing Wang and Juncheng Li},
  journal = {Technical Report},
  year    = {2026}
}