01 · Overview
Faster backbones, better rollouts
Abstract
Interactive video world models need to generate each video chunk efficiently while responding faithfully to user controls. Many systems use chunk-wise autoregressive generation with few-step denoising, but each chunk still requires several costly denoising iterations. Training-free caching can reduce this cost, yet existing policies make reuse decisions primarily from model-internal denoising dynamics and do not explicitly account for control transitions. Interactive generation exposes a signal they do not use: the controls for a chunk arrive before it is denoised, so a schedule derived from them costs no forward pass.
We analyze adjacent chunks under different control regimes and find that structural similarity drops around action changes, while low-frequency structure remains more persistent than high-frequency detail. Motivated by these observations, we propose CtrlCache, a training-free control-aware caching framework. Its action-aware scheduling and refresh policy labels each chunk as initial, transition, turning, or steady; at one selected interior denoising step, initial and transition chunks retain full computation, while turning and steady chunks reuse the transformer residual. A frequency-mixed history prior guidance further transfers complementary information from the preceding clean latent without an additional DiT forward pass. On Matrix-Game 2.0 and LingBot-World v1/v2, CtrlCache achieves 1.21×–1.41× DiT-backbone speedups while improving WBench Overall on all three models.
02 · Motivation
Controls tell us when the scene will change
We measure the cosine similarity between the clean latents of adjacent chunks under different control regimes.
Similarity drops at action changes. Chunks where the control switches differ most from their predecessor, so cached computation is least trustworthy exactly there.
Low frequencies persist. Coarse layout stays similar across chunks much longer than fine detail, so it is safe to carry forward during steady interaction.

Cosine similarity between the clean latents of adjacent chunks for three representative sequences, split into full latent, low-frequency and high-frequency components. Shading marks the control regime of each chunk: non-turning (steady), sustained turning, and action change (transition).
03 · Method
A cache schedule read straight from the controls
Each chunk is classified from its control sequence alone (see the example at the top of the page). The class decides what happens at one selected interior denoising step, while the first and final steps are always fully computed.
Action-aware scheduling & refresh
Action changes are detected across chunk boundaries and within a chunk. Turning and steady chunks skip the transformer block stack at the selected step and apply the residual from the most recent full step of the same chunk; initial and transition chunks keep full computation, refreshing the cache exactly when the scene is expected to change.
Frequency-mixed history prior guidance
In steady chunks, the preceding chunk's last clean latent is split by a 3×3 average pool into low- and high-frequency parts. A prior that keeps the layout, attenuates detail, and decays with temporal distance is blended into the first-step clean latent estimate, at no extra DiT forward.
Training free — model weights, sampler, control interface and persistent context are unchanged.

Full pipeline. (a) Chunk-wise autoregressive denoising with residual reuse at the selected step; (b) the action-aware scheduling and refresh policy; (c) the frequency-mixed history prior guidance.
04 · Results
Quality and efficiency on WBench
Navigation track. vs. Original measures fidelity to the unmodified model under identical prompts, controls and initial frames. Latency is mean DiT-backbone denoising time per case on one A100. Bold marks the best result within each model.
| WBench | vs. Original | Efficiency | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Model | Method | Quality↑ | Cons.↑ | Inter.↑ | Setting↑ | Physics↑ | Overall↑ | PSNR↑ | SSIM↑ | LPIPS↓ | Latency (s)↓ | Speedup↑ |
| Matrix-Game 2.0 | Original | 0.7360 | 0.6838 | 0.8261 | 0.5129 | 0.5127 | 0.6543 | – | – | – | 14.99 | 1.00× |
| TeaCache | 0.7252 | 0.6820 | 0.8145 | 0.4982 | 0.5275 | 0.6495 | 13.89 | 0.3344 | 0.4810 | 10.33 | 1.45× | |
| EasyCache | 0.7245 | 0.6803 | 0.8093 | 0.5464 | 0.5518 | 0.6625 | 13.46 | 0.3152 | 0.5021 | 10.17 | 1.47× | |
| CtrlCache (ours) | 0.7370 | 0.6867 | 0.8265 | 0.5527 | 0.5900 | 0.6786 | 14.51 | 0.3504 | 0.4757 | 10.67 | 1.41× | |
| LingBot-World v1 | Original | 0.8073 | 0.8715 | 0.8088 | 0.7868 | 0.6338 | 0.7816 | – | – | – | 104.54 | 1.00× |
| TeaCache | 0.8052 | 0.8755 | 0.8306 | 0.7844 | 0.7023 | 0.7996 | 15.54 | 0.4638 | 0.3855 | 80.31 | 1.30× | |
| EasyCache | 0.8074 | 0.8751 | 0.7907 | 0.7910 | 0.7287 | 0.7986 | 15.26 | 0.4506 | 0.3946 | 79.80 | 1.31× | |
| CtrlCache (ours) | 0.8073 | 0.8737 | 0.8200 | 0.8246 | 0.7379 | 0.8127 | 17.27 | 0.5014 | 0.3451 | 83.03 | 1.26× | |
| LingBot-World v2 | Original | 0.8290 | 0.8716 | 0.8389 | 0.7692 | 0.7124 | 0.8042 | – | – | – | 66.93 | 1.00× |
| TeaCache | 0.8307 | 0.8848 | 0.8279 | 0.8078 | 0.7171 | 0.8137 | 14.98 | 0.4176 | 0.3538 | 51.54 | 1.30× | |
| EasyCache | 0.8303 | 0.8824 | 0.8296 | 0.8047 | 0.6983 | 0.8091 | 14.90 | 0.4161 | 0.3586 | 51.12 | 1.31× | |
| CtrlCache (ours) | 0.8278 | 0.8789 | 0.8334 | 0.8306 | 0.7298 | 0.8201 | 18.35 | 0.5238 | 0.2832 | 55.18 | 1.21× | |
05 · Videos
Video comparisons
All methods share the same initial frame, prompt, control sequence and seed. Frames in the paper are sampled from these videos. Use Slider to compare any two methods, or All four to see them side by side.
06 · Citation
BibTeX
@article{song2026ctrlcache,
title = {CtrlCache: Accelerating Interactive Video World Models with Control-Aware Caching},
author = {Song, Shangye and Gong, Dong and Jia, Hong and Koh, Yun Sing and Zhang, Xinyu},
journal = {arXiv preprint arXiv:2610.08777},
year = {2026}
}