Long-horizon autoregressive video generation is limited by a finite context window. When an object or scene falls out of context, its fine-grained visual details may be lost and difficult to recover upon reappearance. The simplest solution would be to keep the entire generated history active, but this quickly becomes impractical: as the rollout grows, so do both the KV cache and the cost of attending to it. The challenge is therefore not to remember everything, but to recover a small set of historical KV needed for the current continuation.
To retain access to such visual details, we introduce MosaiChunk, a spatio-temporal memory mechanism that composes a mosaic of selected historical key–value (KV) entries across space and time. Our approach is motivated by the observation that a frozen video generator can directly consume such non-contiguous historical KV and recover the corresponding visual content. We therefore keep the generator fixed and learn only a lightweight router that determines which historical sections to include in the mosaic under a fixed active-memory budget.
Qualitative results
Within each scene, methods share the prompt and noise seed; I2V methods also share the conditioning frame and input camera trajectory. Comparisons therefore differ only in the memory supplied to the backbone. Across T2V and I2V examples, MosaiChunk better preserves visual details when previously seen objects or scenes return to view. More results in Video Viewer.
Prompt (excerpt)Full prompt
01 / METHOD
Can a frozen backbone reuse a non-contiguous subset of historical KV?
Autoregressive video models generate one chunk at a time, attending to cached keys and values (KV) from earlier chunks. Under sliding-window inference, the backbone is conditioned on the most recent chunks, as well as a fixed attention sink used to stabilize generation; we omit this additional sink from the explanations below for simplicity. To condition generation on more distant memory, existing retrieval methods select past chunks as additional context. But are whole chunks necessary?
We test whether a video generator can be conditioned on non-contiguous chunks, providing targeted KV sections as additional context. We select only the KV entries that represent the relevant content as far memory. We ask if these entries can steer the generation and bring that content back. We retain a copy of an earlier chunk's KV before eviction and manually mark the cookie region in one frame. We map this region to latent positions and select the corresponding cached keys and values. We supply them alongside the sliding window when the tin opens again.
Memory architecture
Our goal is to automatically identify and compose the right historical KV entries under a fixed active-cache budget.
Stage 1: Encode and store. We partition each chunk's unrotated KV into equal-size groups, called sections, using balanced k-means. We observe that the KV entries for each section exhibit strong temporal redundancy: sections from neighboring chunks are always similar. Therefore, we train a lightweight descriptor encoder that maps the pooled keys of each section s to a descriptor d(s). The descriptor space learns an anti-recency bias so that similarity over descriptors reflects semantic relatedness rather than temporal proximity. The section bank stores these sections of verbatim KV entries and their descriptors in CPU memory.
Stage 2: Score historical sections. For retrieval when generating chunk cₜ, the sections of the latest generated chunk cₜ₋₁ form the query set Q, and the historical sections outside the sliding window form the candidate set H. The score measures a candidate's highest descriptor similarity to any query section.
Stage 3: Extract and compose. The router ranks all candidate sections by their scores and selects the global top-N. Since sections have equal size, the far-memory budget determines N. A softmax over all candidate scores gives a weight for each section. We normalize the selected weights to unit mean and use them to scale the selected sections' values. We then concatenate their KV into a MosaiChunk, which serves as far memory for chunk cₜ. The frozen DiT reads the MosaiChunk alongside the sliding window.
Self-distillation
We use RAVEN-adapted MiniMax-H3 (H3-AR) for text-to-video (T2V) and LingBot-World-Infinity for image-to-video (I2V). We train the descriptor encoder through self-distillation to match the same frozen generator's predictions under richer far memory. Teacher and student use the same frozen DiT, noisy latent, denoising step, and sliding window; only far memory differs. The teacher receives whole historical chunks that fully cover the historical content to be redrawn. The student receives the router's MosaiChunk under a smaller memory budget. We minimize the mean squared error (MSE) between teacher and student predictions.
02 / BENCHMARK
Existing benchmarks such as WBench, WorldMark, and PersistBench evaluate video quality and consistency, but do not systematically test long-horizon revisits in which frames establishing the target content's appearance are evicted from the model's sliding window. We therefore introduce RememBench, a benchmark with two splits for evaluating consistency across such revisits in autoregressive video models. Revisits are driven by prompts in the T2V split and by camera motion in the I2V split.
The split contains 100 samples with scenarios disjoint from router training. Each model input extends Ring Forcing's three-stage appear–disappear–reappear design to four prompt segments, with a separate segment keeping the object out of sight. The nominal prompt transitions occur at 3.6, 6.2, and 11.2 seconds. The third segment therefore requests five seconds with the object out of sight, exceeding the largest sliding-window baseline's approximately three seconds of recent context.
Model input: Prompt
1Appears
2Disappears
3Out of sight
4Reappears
Input camera trajectory · schematic
Visual diversity
Model-generated T2V frames when the prompt first reveals the object; 16 randomly sampled scenes.
03 / RESULTS
We evaluate whether MosaiChunk provides useful fine-grained historical information beyond a sliding window and compare it with existing retrieval mechanisms that condition on whole chunks, testing whether section-level memory better preserves visual details. We conduct both comparisons on the T2V and I2V splits of RememBench. We compare MosaiChunk with the sliding-window baseline (Base) and Mixture of Contexts (MoC), a whole-chunk retrieval mechanism. We adapt MoC on each frozen backbone.
Quantitative results
Let F_c denote the far memory used to generate chunk c, and ‖F_c‖ its budget in video-chunk equivalents. For example, ‖F_c‖ = 1 means that the size of the far memory is equivalent to one video chunk. Retrieval methods keep their sliding window fixed, whereas Base matches the total active KV cache budget by retaining additional recent chunks.
CLIP similarity and LPIPS assess consistency between departure and revisit frames. T2V pairs are manually marked; I2V pairs use the conditioning frame and a revisit selected from Pi3X-reconstructed camera poses. We report medians over scenes.
Revisit consistency vs. memory budget
CLIP ↑
Higher is betterLPIPS ↓
Lower is betterHover or focus a point for its method, memory budget, and score.
Our experiments show that MosaiChunk consistently improves revisit consistency over both sliding-window inference and whole-chunk retrieval under matched memory budgets, across both T2V and I2V settings. The advantage also holds with less far memory: on T2V, MosaiChunk with one chunk exceeds MoC with two (0.899 vs. 0.841 CLIP); on I2V, half a chunk already does so (0.815 vs. 0.807). These comparisons demonstrate both effective retrieval at small budgets and improved revisit consistency when more memory is available.