LOCI: Spatial Linear Memory
for Streaming World Models

Ji Xia1 · Tingting Liao1 · Xuezhi Liang1 · Hao Li2,3 · Guangyi Liu1,†

1 Institute of Foundation Models, Mohamed bin Zayed University of Artificial Intelligence
2 Mohamed bin Zayed University of Artificial Intelligence   3 Pinscreen
† Corresponding author

Inference code & model weights coming soon

1

Key idea of LOCI

When a camera revisits a previously observed region, a video world model should reproduce what was there before. LOCI combines a compact recurrent memory with a cache of visual observations, using projective camera geometry to guide memory reads and writes.

LOCI observes a scene, accumulates spatial context in recurrent memory, focuses on co-visible history and recovers a previously observed view.
Figure 1.Projectively conditioned recurrent memory integrates historical context, while retained visual observations preserve detail. Spatial context helps retrieve relevant visual history. Attention weights are schematic; click the diagram to enlarge.
Spatial recurrent memory.
Camera geometry enters both memory addressing and stored content in recurrent linear-attention blocks.
Retained visual detail.
Cache-backed attention blocks retain past observations and read them using accumulated scene context.
Bounded streaming.
A bounded bank of retained observations supports long rollouts at constant memory, alongside recurrent state that summarises history.
2

Long-range recall

Recorded trajectory · GT history as memory

The model is given a recorded history of the scene, then generates a continuation along the recorded camera path. The video below compares the matching ground-truth continuation, LOCI and a full-softmax model.

Figure 2.Tokyo, long-range recall. 75.1 s ground-truth history as memory → 44.9 s generated continuation. The full recorded trajectory is 120.1 s; only the continuation is shown here.
3

Paper & code

The paper is available now. Inference code, model weights and examples are being prepared for release.

Paper PDFRead the paper ↗ Code repositoryREADME preview ↗
Inference code & model weightsComing soon
4

Citation

@article{xia2026loci,
  title={LOCI: Spatial Linear Memory for Streaming World Models},
  author={Xia, Ji and Liao, Tingting and Liang, Xuezhi and Li, Hao and Liu, Guangyi},
  year={2026}
}