IMPACT Attention Is the Interaction Map forScalable Interaction-Aware World Model Training

Rongze Tang1,2,*, Jianjie Fang3,*, Zhaolu Wang3, Ziyou Wang3, Xvyuan Liu3, Haisheng Su4, Xin Zhang4, Wei Wu4, Chen Gao2,3,†, Yong Li3, Zhibo Chen1,2,†

1USTC   2Zhongguancun Academy   3Tsinghua University   4Manifold AI

* Equal contribution   † Corresponding authors

Abstract

World models have made remarkable progress in action-conditioned future prediction for embodied agents, yet still struggle to model physically plausible interactions. Existing approaches address this limitation by constraining generation with external representations encoding motion, geometry, or semantics, whose construction depends on auxiliary estimators or manual annotations.

We identify a supervision-allocation mismatch under the globally averaged MSE denoising objective: prevalent static content dominates optimization, while sparse dynamic-object regions are disproportionately under-supervised. IMPACT uses manipulated-object cross-attention as an internal spatiotemporal prior, calibrates sampled candidates with detached local prediction errors, and targets denoising supervision with the resulting interaction map.

IMPACT overview
Contributions

Attention Is the Interaction Map.

01 — PROBLEM

Uniform supervision spends gradient on what is abundant.

Standard MSE spreads its signal across the frame. IMPACT reallocates supervision toward sparse object motion and contact.

Ground TruthStandard MSEIMPACT
Robot armStacking blocks
Ground truth stacking blocks
Standard MSE stacking blocks
IMPACT stacking blocks
Human handWhite box
Ground truth white box
Standard MSE white box
IMPACT white box
Method

Calibrate in the forward pass.
Target in the backward pass.

IMPACT method framework
Object-token groundingIdentify “blue bowl” as the manipulated object.
Results

Consistent across backbones and control signals.
Effective across embodiments.

WorldArena62.53EWMScore · Cosmos + IMPACT
WorldArena+3.81EWMScore · Wan 2.2-AC
EgoDex110.94FVD ↓
EgoDex0.772Hand IoU ↑

Robot-arm manipulation

WorldArena
ModelEWMScoreVisualMotionContentPhysics3DControl
Wan 2.661.8661.6268.3160.3642.3175.8860.85
CtrlWorld59.7055.3350.2863.9954.8986.3054.65
WoW54.8852.9850.0264.5238.1174.7749.89
Wan 2.2-AC58.6556.5748.2360.5553.4186.1654.40
Wan 2.2-AC + IMPACT62.4660.6053.2860.2955.8792.5660.00
Cosmos-Predict 2.5 (action)55.9157.8738.8968.7342.2382.5349.51
Cosmos-Predict 2.5 (action) + IMPACT62.5356.3268.2257.9041.8277.9971.19

Human-hand manipulation

EgoDex
ModelFVD ↓FID ↓CLIP-Hand ↑Hand IoU ↑
VACE358.4250.650.8950.493
Wan 2.2-AC366.1244.710.9210.693
Wan 2.2-AC + IMPACT110.945.790.9520.772

Component ablation

WorldArena
ConfigurationEWMScoreVisualMotionPhysics
Wan 2.2-AC58.6556.5748.2353.41
+ IWS61.5460.4449.1655.16
+ IWS + ADS62.4660.6053.2855.87
Gallery

Contact stays coherent.
Objects follow the action.

Citation
@article{tang2026impact,
  title  = {IMPACT: Attention Is the Interaction Map for Scalable Interaction-Aware World Model Training},
  author = {Tang, Rongze and Fang, Jianjie and Wang, Zhaolu and Wang, Ziyou and Liu, Xvyuan and Su, Haisheng and Zhang, Xin and Wu, Wei and Gao, Chen and Li, Yong and Chen, Zhibo},
  journal = {arXiv preprint arXiv:2609.00161},
  year    = {2026},
  url     = {https://arxiv.org/abs/2609.00161}
}