Anonymous Submission · ICRA 2027 (under review)

3D Scene Flow as a World Action Model Intermediate for Learning from Human Video

World action models typically predict RGB video. We argue the right target is instead 3D motion. FloMo renders 3D scene flow into the latent space of a pretrained video generation model and fine-tunes a single backbone to predict motion and action. Predicting motion instead of pixels gives better in-distribution success, far stronger out-of-distribution generalization, higher sample efficiency, and the best success on every RoboTwin simulation task we evaluate.

Authors omitted for double-blind review

Paper Code (coming soon) See the method

Architecture

Motion mediates transfer from video to action

FloMo architecture: text, image, scene flow, and action feed a shared video-generation backbone
FloMo architecture. From an image and text prompt, a pretrained video backbone (Wan 2.2 5B DiT) jointly flow-matches dense 3D motion and robot actions.

Motion tokenization

3D motion in the video latent space

Scene flow keeps task motion and discards the appearance details that do not determine action.

The action that produces a future state depends on how the scene moves, not how it looks. 2D optical flow captures motion but conflates camera ego-motion with task-relevant motion, so FloMo predicts 3D scene flow in a canonical frame: the cumulative 3D displacement of each tracked point, expressed in the first frame's camera.

We render this as an RGB video and encode it with the frozen Wan 2.2 VAE. Rendered flow lands in the same latent space as video, so the pretrained DiT handles it with no new tokenizer. We fine-tune only the backbone (LoRA, rank 64).

Pixels vs. motion

RGB observation · 2D tracks · 3D scene flow

RGB observation 2D tracks 3D scene flow
Lift board
RGB observation · 2D tracks · 3D scene flow.
Clear weeds
RGB observation · 2D tracks · 3D scene flow.
Place ring in cup
RGB observation · 2D tracks · 3D scene flow.
Wrap gift box
RGB observation · 2D tracks · 3D scene flow.

Method

Two-stage co-training

FloMo is co-trained on a mixture of egocentric human video and bimanual robot teleoperation.

Stage 1
Mid-training

Train on the full mixture. Human video supervises scene flow only; robot data supervises scene flow and action. This builds a broad motion prior at scale.

Stage 2
Fine-tuning

Train on robot video/action data only, dropping human video. This sharpens action prediction without diluting the gradient with action-free samples.

At inference, from an image and a language instruction, FloMo denoises for 10 flow-matching steps, then executes a chunk of 16 actions before re-planning.

Contributions

What we show

01

Motion prediction as RGB video prediction

We render 3D scene flow as RGB and encode it with the frozen video VAE, so motion shares the video latent space. A single backbone flow-matches motion and action with no new tokenizer.

02

Motion is the most effective target

On real-robot manipulation, predicting motion gives higher in-distribution success and out-of-distribution generalization than predicting video; adding video prediction interferes with action learning. FloMo also leads every task on the RoboTwin simulation benchmark.

03

Generalization from human video

FloMo transfers to objects, motion primitives, and tasks seen only in action-free human video and never present in robot training.

Results

Predicting motion beats predicting pixels

In-distribution success

Three contact-rich bimanual tasks (Oven, Toolbox, Corn), 15 real-robot rollouts per task. The VMA (video + motion + action) and VA (video + action) variants share FloMo's architecture and training protocol and differ only in prediction target.

In-distribution success rates for FloMo, VMA, and VA on Oven, Toolbox, and Corn
In-distribution success. Across 15 rollouts FloMo performs best on every task; the matched video–motion–action (VMA) and video–action (VA) variants lower success.
Average in-distribution success (%)

FloMo wins every task, averaging 91% against 51% for VMA and 40% for VA. Adding video prediction to the motion–action objective costs 40 points; replacing motion with video costs 51. The gap is largest on Toolbox (87% vs. 13% for both video variants), where opening the lid requires locating a narrow crevice and executing precise, contact-rich motion.

In-distribution rollouts

FloMo successes on the same three in-distribution evaluation tasks.

Oven rollout poster
In-distribution

Open the oven, retrieve the croissant, place it on a plate.

OvenIn-dist
Open the oven, retrieve the croissant, place it on a plate.
ToolboxIn-dist
Open the hinged lid, grasp the screwdriver, place it inside.
CornIn-dist
Grasp the corn cob and place it in the bowl.

Out-of-distribution generalization

FloMo is the only method that generalizes consistently across all three OOD axes: an unseen primitive (Close Cabinet), task (Pull String), and object (Highlighter). Every method that predicts pixels, through video or image-space 2D tracks, degrades sharply.

Method Close Cabinet Pull String Highlighter Aggregate
SR ↑TP ↑ SR ↑TP ↑ SR ↑TP ↑ SR ↑TP ↑
VMA (video + motion)0/1000/1002/10452/3015
VA (video)0/10201/10101/10152/3015
DreamZero5/10751/10150/1006/3030
AMPLIFY6/10700/1000/1006/3023
FloMo (ours)6/10707/10705/106518/3068

10 rollouts per condition (30 total)  ·  SR = successful rollouts / total  ·  TP = avg. fraction of task stages completed (%). FloMo averages 60% success and 68% task progress, vs. 20% SR for AMPLIFY and DreamZero and 7% for VMA and VA.

Out-of-distribution rollouts

FloMo successes across the unseen primitive, task, and object evaluations.

OOD · unseen primitive

Push the cabinet door shut.

Out-of-distribution rollouts

Close CabinetUnseen primitive
Push the cabinet door shut.
Pull StringUnseen task
Pull the string on the toy in the air.
HighlighterUnseen object
Place the unseen highlighter in the bowl.

Simulation: RoboTwin benchmark

Single-task adaptation on five RoboTwin tasks spanning bimanual grasping, sequential placement, articulated objects, precision placement, and bimanual transfer. Each policy is trained on the 50 official demonstrations (5,000 gradient steps, batch size 256) and evaluated on 100 Easy-setting episodes per task and method.

RoboTwin Easy success rates per task for FloMo, pi 0.5, FAST-WAM, and BC
RoboTwin single-task evaluation. Easy-setting success over 100 episodes per task and method. FloMo performs best on every task.

FloMo averages 63.6% success against 44.8% for the pretrained VLA π0.5, 41.0% for the video-predicting world action model FAST-WAM (same rank-64 Wan-backbone LoRA as FloMo), and 19.2% for a 20M-parameter diffusion-policy BC baseline without video pretraining.

Method Pick Dual BottlesStack Two BowlsOpen LaptopHang MugHandover MicAverage
BC (Diffusion Policy)250660519.2
FAST-WAM20276858541.0
π0.549465817044.8
FloMo (ours)76727179263.6

Success rate (%) over 100 RoboTwin Easy episodes per task and method.

Why motion works

The action signal is both easier to learn and easier to find

Two controlled probes point to the same conclusion: motion is a more direct bridge from observation to action than pixels.

22%lower action error
with only 1% of the data
Inverse-dynamics scaling
Motion contains the information necessary to predict actions. It reaches the same performance ceiling as video despite containing less information, and is slightly more sample-efficient.
3×more action attention
to motion than video
Action attention by modality
Action queries seek out motion. Attention to motion rises through the DiT trunk and dominates attention to video, revealing where the model finds action-relevant evidence.

Generalization from human video

Scene flow transfers from human video

3D scene flow is embodiment-agnostic: a human hand and a robot gripper produce a similar motion field. FloMo learns motions from action-free human video and carries them to the robot.

Ablation · remove only the human video
FloMo
60%
with human video
FloMo
7%
no human video

Same model, same motion target. Removing only the human video drops average OOD success by 53 points.

Training data Close Cabinet Pull String Highlighter Average
SRTPSRTPSRTPSRTP
FloMo6/10707/10705/106518/3068
FloMo (no human video)0/5201/5500/5201/1530

Both models predict 3D motion and differ only in whether the ~81 h EgoVerse subset and 0.7 h of internal action-free human demonstrations are in training. FloMo uses 10 rollouts per condition; the no-human-video ablation uses 5. SR = successful rollouts / total; TP = avg. fraction of task stages completed (%).

Same model, same motion target. Removing only the human video drops average OOD success from 60% to 7%. Without it, the model can begin each task but rarely completes the novel motion: it contacts the cabinet on 2 of 5 trials but never closes it, and grasps the string on 4 of 5 but pulls it on only one. The robot data contains no cabinets, strings, or markers, while the human video covers all three motion classes. The generalization comes from human video, made transferable by the motion representation.

Side by side

Pull String (OOD): FloMo vs. baselines

FloMo succeeds on 7/10 Pull String rollouts; no baseline exceeds 1/10. Prompt: Pull the string on the toy in the air.

FloMoSuccess
DreamZeroFailure
BC (no pretraining)Failure
FloMoSuccess
DreamZeroFailure
BC (no pretraining)Failure

Code

Code & resources

Code (coming soon)

Source released upon acceptance.