NewTraining-free acceleration for interactive world models

CtrlCache: Accelerating Interactive Video World Models with Control-Aware Caching

Shangye Song1 Dong Gong2 Hong Jia1 Yun Sing Koh1 Xinyu Zhang1,*

1University of Auckland2UNSW Sydney

*Corresponding author

TL;DR — In interactive generation, the next controls are known before a chunk is denoised. CtrlCache uses them to decide when cached computation can be reused and when it must be refreshed, giving 1.21–1.41× faster DiT backbones with higher WBench scores and no retraining.

01 · Overview

Faster backbones, better rollouts

Matrix-Game 2.0
1.41×speedup
WBench Overall 0.6543 → 0.6786
LingBot-World v1
1.26×speedup
WBench Overall 0.7816 → 0.8127
LingBot-World v2
1.21×speedup
WBench Overall 0.8042 → 0.8201

Abstract

Interactive video world models need to generate each video chunk efficiently while responding faithfully to user controls. Many systems use chunk-wise autoregressive generation with few-step denoising, but each chunk still requires several costly denoising iterations. Training-free caching can reduce this cost, yet existing policies make reuse decisions primarily from model-internal denoising dynamics and do not explicitly account for control transitions. Interactive generation exposes a signal they do not use: the controls for a chunk arrive before it is denoised, so a schedule derived from them costs no forward pass.

We analyze adjacent chunks under different control regimes and find that structural similarity drops around action changes, while low-frequency structure remains more persistent than high-frequency detail. Motivated by these observations, we propose CtrlCache, a training-free control-aware caching framework. Its action-aware scheduling and refresh policy labels each chunk as initial, transition, turning, or steady; at one selected interior denoising step, initial and transition chunks retain full computation, while turning and steady chunks reuse the transformer residual. A frequency-mixed history prior guidance further transfers complementary information from the preceding clean latent without an additional DiT forward pass. On Matrix-Game 2.0 and LingBot-World v1/v2, CtrlCache achieves 1.21×–1.41× DiT-backbone speedups while improving WBench Overall on all three models.

02 · Motivation

Controls tell us when the scene will change

We measure the cosine similarity between the clean latents of adjacent chunks under different control regimes.

1

Similarity drops at action changes. Chunks where the control switches differ most from their predecessor, so cached computation is least trustworthy exactly there.

2

Low frequencies persist. Coarse layout stays similar across chunks much longer than fine detail, so it is safe to carry forward during steady interaction.

Adjacent-chunk latent similarity under steady, turning and transition regimes

Cosine similarity between the clean latents of adjacent chunks for three representative sequences, split into full latent, low-frequency and high-frequency components. Shading marks the control regime of each chunk: non-turning (steady), sustained turning, and action change (transition).

03 · Method

A cache schedule read straight from the controls

Each chunk is classified from its control sequence alone (see the example at the top of the page). The class decides what happens at one selected interior denoising step, while the first and final steps are always fully computed.

Action-aware scheduling & refresh

Action changes are detected across chunk boundaries and within a chunk. Turning and steady chunks skip the transformer block stack at the selected step and apply the residual from the most recent full step of the same chunk; initial and transition chunks keep full computation, refreshing the cache exactly when the scene is expected to change.

Frequency-mixed history prior guidance

In steady chunks, the preceding chunk's last clean latent is split by a 3×3 average pool into low- and high-frequency parts. A prior that keeps the layout, attenuates detail, and decays with temporal distance is blended into the first-step clean latent estimate, at no extra DiT forward.

Training free — model weights, sampler, control interface and persistent context are unchanged.

CtrlCache pipeline

Full pipeline. (a) Chunk-wise autoregressive denoising with residual reuse at the selected step; (b) the action-aware scheduling and refresh policy; (c) the frequency-mixed history prior guidance.

04 · Results

Quality and efficiency on WBench

Navigation track. vs. Original measures fidelity to the unmodified model under identical prompts, controls and initial frames. Latency is mean DiT-backbone denoising time per case on one A100. Bold marks the best result within each model.

WBenchvs. OriginalEfficiency
ModelMethodQuality↑Cons.↑Inter.↑Setting↑Physics↑Overall↑ PSNR↑SSIM↑LPIPS↓Latency (s)↓Speedup↑
Matrix-Game 2.0Original0.73600.68380.82610.51290.51270.6543–––14.991.00×
TeaCache0.72520.68200.81450.49820.52750.649513.890.33440.481010.331.45×
EasyCache0.72450.68030.80930.54640.55180.662513.460.31520.502110.171.47×
CtrlCache (ours)0.73700.68670.82650.55270.59000.678614.510.35040.475710.671.41×
LingBot-World v1Original0.80730.87150.80880.78680.63380.7816–––104.541.00×
TeaCache0.80520.87550.83060.78440.70230.799615.540.46380.385580.311.30×
EasyCache0.80740.87510.79070.79100.72870.798615.260.45060.394679.801.31×
CtrlCache (ours)0.80730.87370.82000.82460.73790.812717.270.50140.345183.031.26×
LingBot-World v2Original0.82900.87160.83890.76920.71240.8042–––66.931.00×
TeaCache0.83070.88480.82790.80780.71710.813714.980.41760.353851.541.30×
EasyCache0.83030.88240.82960.80470.69830.809114.900.41610.358651.121.31×
CtrlCache (ours)0.82780.87890.83340.83060.72980.820118.350.52380.283255.181.21×

05 · Videos

Video comparisons

All methods share the same initial frame, prompt, control sequence and seed. Frames in the paper are sampled from these videos. Use Slider to compare any two methods, or All four to see them side by side.

06 · Citation

BibTeX

@article{song2026ctrlcache,
  title   = {CtrlCache: Accelerating Interactive Video World Models with Control-Aware Caching},
  author  = {Song, Shangye and Gong, Dong and Jia, Hong and Koh, Yun Sing and Zhang, Xinyu},
  journal = {arXiv preprint arXiv:2610.08777},
  year    = {2026}
}