Wonder: Video World Model Done Better

Jiacong Xu*, Hanwen Jiang, Zhixin Shu, Kalyan Sunkavalli, Vishal Patel, Yiqun Mei*†

Adobe Research, Johns Hopkins University

* Equal Contribution Project Lead

I2V+V2V Multimodality
0.5s Constant Latency
16 FPS Throughput
6-DoF camera control

Overview

Abstract

Given an image or video, Wonder builds a persistent interactive world that users can navigate by moving the camera, revealing unseen regions, and returning to previously observed areas with constant latency for up to one minute. This requires system-level co-design across control representation, memory mechanism, and training strategy. Wonder turns camera motion into dense visual evidence, retrieves relevant full-fidelity history through sparse attention, and uses a rectified distillation pipeline to preserve control and long-term consistency. Together, these components enable diverse 16 FPS rollouts with coherent geometry, appearance, dynamics, and interactive control across image-to-video and video-conditioned generation.

Method

System-Level Co-Design

01

Control Signal

Rendering a synthetic camera space with a 3D scaffold and environment map, turning translation and rotation into dense, frame-aligned visual evidence.

02

Memory Mechanism

Keeping full-fidelity history KV caches and adaptively selecting relevant entries via sparse attention to preserve long-horizon memory at constant latency.

03

Training Strategy

Improving student capacity with a Mixture-of-Students design and resolving camera drift during distillation via GAN Control Regularization.

Demos

Image-to-Video Game Worlds

Demos

Image-to-Video Cartoon Worlds

Demos

Image-to-Video Real Worlds

Demos

Complex Camera Movement

Demos

Video-to-Video Real Worlds

Demos

Video-to-Video Cartoon & Game Worlds

Comparison

Compare with SOTA Models

Citation

@article{xu2026wonder,
  title={Wonder: Video World Model Done Better},
  author={Xu, Jiacong and Jiang, Hanwen and Shu, Zhixin and Sunkavalli, Kalyan and Patel, Vishal M. and Mei, Yiqun},
  journal={arXiv preprint arXiv:2607.26037},
  year={2026},
  url={https://arxiv.org/abs/2607.26037}
}

Notice: All images and videos on this page are used solely for research and demonstration purposes. Copyrights remain with their respective owners. If you believe any content infringes your rights, please contact us, and we will promptly remove the corresponding example.