LongTake

Long-Horizon Continuation
in Autoregressive Video Diffusion

1KAIST 2Georgia Institute of Technology

† Equal advising

Coherent scene evolution and sustained dynamics
over extended durations.

30 & 60 second rolloutsBounded KV-cache4-step sampling

01 / METHOD

Learning to sustain dynamics
through two-stage training.

LongTake learns from curated real long videos, then trains a few-step generator under student self-rollout. Long-Horizon TF provides a strong direct initialization for self-rollout DMD, without a separate few-step distillation stage.

01 / TEACHER GAP

Long-Horizon Teacher Forcing

Short clips provide no direct supervision for prediction beyond their training horizon. Long-Horizon TF addresses this teacher gap by learning to predict later frames from long ground-truth video prefixes with a bounded KV-cache.

02 / STUDENT GAP

Self-rollout DMD

At inference, the KV-cache comes from generated frames rather than ground-truth videos. Self-rollout DMD addresses this student gap, starting directly from the Long-Horizon TF model. Optional Hybrid DMD reuses the AR teacher for conditional supervision of later frames, while retaining bidirectional joint supervision over the initial window.

Two-stage LongTake training pipeline. Stage 1: Long-Horizon TF. Stage 2: self-rollout DMD, optionally extended with Hybrid DMD for conditional supervision over later chunks.

02 / TEACHER INITIALIZATION

Long-Horizon TF provides a stronger initialization for DMD.

Short TF, Short TF + CD, and Long TF under the same 5s self-rollout DMD training, evaluated with 30s rollouts.

All three use joint DMD only (λ = 0), isolating teacher initialization from the optional conditional guidance in Hybrid DMD.

30 seconds · MovieGen prompt 056
All videos
Subject Consistency, Dynamic Degree, and Aesthetic Quality: Short TF 98.19, 51.30, 63.90; Short TF + CD 96.04, 81.72, 58.82; Long TF 96.83, 90.73, 62.51.
0:000:30

Read the full prompt

03 / COMPARISONS

Long-horizon continuation with sustained dynamics.

Compare 30s and 60s autoregressive rollouts with the same prompt and seed.

All videos
0:000:30

Read the full prompt

Effect of the conditional DMD weight30-second ablation

λ controls the conditional DMD weight. We use λ = 0.2 for the main comparisons and gallery.

Same prompt & seed · 30 s
All videos
0:000:30

Read the full prompt

04 / VIDEO GALLERY

30s and 60s autoregressive rollouts.

Selected qualitative examples of long-horizon continuation, coherent scene evolution, and sustained dynamics.