LOCI: Spatial Linear Memory
for Streaming World Models
1 Institute of Foundation Models, Mohamed bin Zayed University of Artificial Intelligence
2 Mohamed bin Zayed University of Artificial Intelligence 3 Pinscreen
† Corresponding author
Inference code & model weights coming soon
Key idea of LOCI
When a camera revisits a previously observed region, a video world model should reproduce what was there before. LOCI combines a compact recurrent memory with a cache of visual observations, using projective camera geometry to guide memory reads and writes.
- Spatial recurrent memory.
- Camera geometry enters both memory addressing and stored content in recurrent linear-attention blocks.
- Retained visual detail.
- Cache-backed attention blocks retain past observations and read them using accumulated scene context.
- Bounded streaming.
- A bounded bank of retained observations supports long rollouts at constant memory, alongside recurrent state that summarises history.
Long-range recall
Recorded trajectory · GT history as memoryThe model is given a recorded history of the scene, then generates a continuation along the recorded camera path. The video below compares the matching ground-truth continuation, LOCI and a full-softmax model.
Paper & code
The paper is available now. Inference code, model weights and examples are being prepared for release.
Citation
@article{xia2026loci,
title={LOCI: Spatial Linear Memory for Streaming World Models},
author={Xia, Ji and Liao, Tingting and Liang, Xuezhi and Li, Hao and Liu, Guangyi},
year={2026}
}