Embodied world modeling

KineWorldAction-Induced Transport Fields for Embodied World Modeling

  • Ziying Song1,*
  • Yuchen Liu2,*
  • Zhuoran Xu3,*
  • Ziyang Liu4
  • Jian Jin5
  • Jiangtao Su1
  • Haibao Yu6
  • Lei Yang1,†
  • Yuanpei Chen7,‡
  • 1 Nanyang Technological University
  • 2 North University of China
  • 3 psibot. Inc
  • 4 Tsinghua University
  • 5 China Academy of Information and Communications Technology
  • 6 University of Hong Kong
  • 7 Peking University

* Co-first authors† Corresponding author‡ Project lead

Predicting the visual consequences of robot actions with camera-aligned motion and transport-aware training.

Manuscript figures and selected video samples. See the linked repositories for current artifact availability.

Robot arm interacting with objects on a tabletop
Selected video frame 312 / 1000
Robot gripper above a red tabletop object
Sample 001
01KTLCamera-aligned kinematic transport
02TAWDTransport-aware RGB supervision
03VideoAction-conditioned future frames
The approach / 02

From motion to visual consequence.

KineWorld uses commanded robot motion to guide video generation and to focus training on regions affected by interaction.

KineWorld framework showing KTL transport conditioning and TAWD supervision Figure 01 KineWorld framework from the V4 manuscript Open full size ↗
01 / KTL

Kinematic Transport Lifting

Robot-only renderings and masks support camera-aligned image-plane transport from commanded actions. The resulting field conditions future video generation.

Conditioning · training and inference
02 / TAWD

Transport-Aware World Diffusion

Transport support is calibrated on the video latent grid and converted to normalized token weights for future-RGB supervision.

Loss weighting · training only
Appendix figure A1

A closer look at the pipeline.

Actions drive robot renderings, transport provides a camera-aligned condition, and the resulting support changes the allocation of future-RGB training loss. The 3D field in this drawing is schematic; KTL estimates image-plane transport with RAFT between rendered robot states.

Conceptual KineWorld pipeline with input rendering, transport lifting, and transport-aware weightingAppendix figure A1 Detailed conceptual pipeline Open full size ↗
Camera-aligned transport / 03

Follow the action through the image.

The manuscript illustration connects robot motion, image-plane transport, and the support used by the training objective. It is a conceptual visualization, not a measured rollout.

Open figure at full size ↗
Conceptual transport support construction across two robot manipulation scenes
Figure 02 · Transport-support illustration from the V4 manuscript.
Training signal / 04

Allocate attention where actions matter.

The paper's analytical plots illustrate how a constructed transport support changes token-weight allocation. These curves are illustrative, not measured training statistics.

From the manuscript / 05

A qualitative rollout view.

A retained manuscript figure places a Wan2.2 example and a KineWorld example side by side at sampled times. This visual comparison does not establish a benchmark score or identify the source of the video samples on this page.

Manuscript qualitative rollout comparison at input and five future time pointsFigure 05 Retained qualitative comparison from the V4 manuscript Open full size ↗
Appendix gallery / 06

Fifteen more rollouts, frame by frame.

The V4 appendix collects 15 single-view episode pairs across three plates. Each pair starts with the decoded input and follows with five sampled future frames, allowing the scene, robot, and objects to be inspected through time.

Read the rowsWan2.2 above, KineWorld below
Read the columnsInput, then five future frames
InterpretationVisual examples, not a matched ablation
Plate 01 / 03

Objects on the table.

Episodes 45, 125, 212, 300, and 336 range from simple block layouts to food-handling scenes. Follow the robot's appearance and the position of objects across each six-frame row.

The appendix describes visible rollout behavior; a few examples cannot establish task completion or trajectory accuracy.

Open the full plate ↗
Wan2.2 and KineWorld future-frame comparisons for episodes 45, 125, 212, 300, and 336Episodes 45 · 125 · 212 · 300 · 336
Plate 02 / 03

Changing scenes, changing objects.

Episodes 391, 488, 586, 704, and 732 include sparse white tabletops and textured workspaces. Compare what moves, what stays in view, and whether the robot setting remains consistent.

Frames are sampled from each encoded video; the display is a qualitative sequence, not a controlled component comparison.

Open the full plate ↗
Wan2.2 and KineWorld future-frame comparisons for episodes 391, 488, 586, 704, and 732Episodes 391 · 488 · 586 · 704 · 732
Plate 03 / 03

Across interaction types.

Episodes 780, 865, 938, 976, and 982 show several different object arrangements and tasks. Look across the five future frames for changes in robot embodiment, object identity, and scene composition.

These are retained manuscript examples. They are separate from the three playable videos on this page and are not attributed to the public step-500 checkpoint.

Open the full plate ↗
Wan2.2 and KineWorld future-frame comparisons for episodes 780, 865, 938, 976, and 982Episodes 780 · 865 · 938 · 976 · 982
Three-camera validation / 07

One event. Three camera views.

These V4 appendix plates compare reference and generated frames at eleven shared frame indices. Each plate places the left wrist, head, and right wrist views side by side, exposing where a rollout agrees across views and where it diverges.

Across columnsLeft wrist · head · right wrist
Within each cellReference at left, KineWorld at right
Across rowsEleven shared frame indices
Episode 001 / Bottle

Follow a bottle lift.

Compare the bottle's position during the reference lift with the generated head and left-wrist views. The plate makes differences in close-up object visibility easy to inspect.

Validation diagnostic from the manuscript; not a scored multi-view test result.

Open all eleven rows ↗
Reference and KineWorld bottle manipulation frames across left wrist, head, and right wrist camerasBottle manipulation · episode 1
Episode 050 / Placement

Compare the tray and wrists.

Follow the food-placement sequence in all three views. The tray contents and wrist-camera detail give concrete places to inspect differences between reference and generated frames.

The views are synchronized for visual inspection; this is separate from the paper's scored test outputs.

Open all eleven rows ↗
Reference and KineWorld food-placement frames across left wrist, head, and right wrist camerasFood placement · episode 50
Episode 100 / Switch

Watch the event timing.

Compare when motion appears in the head view and how objects enter the right-wrist view. The static left-wrist scene provides a useful contrast within the same episode.

A stable background alone is not evidence that object transport agrees across cameras.

Open all eleven rows ↗
Reference and KineWorld switch-interaction frames across left wrist, head, and right wrist camerasSwitch interaction · episode 100
Watch / 08

Video samples in motion.

Three playable clips from the supplied collection. The complete 1,000-video archive has a separate dataset repository ↗.

01 / 03 Sample 001
02 / 03 Sample 312
03 / 03 Sample 1000
Explore / 09

Project resources.

Code, model files, and the video collection are published separately. Check each repository for current file availability.

The public source and released artifacts are separate from the full V4 experimental setup. This page makes no numerical benchmark claim or checkpoint attribution for the sample videos. A public manuscript link will be added when available.

Cite / 10

BibTeX.

If you find KineWorld useful, please cite our work.

Download .bib
@misc{song2026kineworld,
  title  = {{KineWorld}: Action-Induced Transport Fields for Embodied World Modeling},
  author = {Song, Ziying and Liu, Yuchen and Xu, Zhuoran and Liu, Ziyang and
            Jin, Jian and Su, Jiangtao and Yu, Haibao and Yang, Lei and Chen, Yuanpei},
  year   = {2026},
  url    = {https://modaxiansheng.github.io/KineWorld/}
}