Uniform supervision spends gradient on what is abundant.
Standard MSE spreads its signal across the frame. IMPACT reallocates supervision toward sparse object motion and contact.






World models have made remarkable progress in action-conditioned future prediction for embodied agents, yet still struggle to model physically plausible interactions. Existing approaches address this limitation by constraining generation with external representations encoding motion, geometry, or semantics, whose construction depends on auxiliary estimators or manual annotations.
We identify a supervision-allocation mismatch under the globally averaged MSE denoising objective: prevalent static content dominates optimization, while sparse dynamic-object regions are disproportionately under-supervised. IMPACT uses manipulated-object cross-attention as an internal spatiotemporal prior, calibrates sampled candidates with detached local prediction errors, and targets denoising supervision with the resulting interaction map.
Standard MSE spreads its signal across the frame. IMPACT reallocates supervision toward sparse object motion and contact.






Object-token grounding is the base unit. Optical flow, point tracking, depth, segmentation, and 3D reconstruction add 18×–142× preprocessing cost.
Each video presents the complete path from input and attention prior to the ADS interaction map and IWS relative gain.
| Model | EWMScore | Visual | Motion | Content | Physics | 3D | Control |
|---|---|---|---|---|---|---|---|
| Wan 2.6 | 61.86 | 61.62 | 68.31 | 60.36 | 42.31 | 75.88 | 60.85 |
| CtrlWorld | 59.70 | 55.33 | 50.28 | 63.99 | 54.89 | 86.30 | 54.65 |
| WoW | 54.88 | 52.98 | 50.02 | 64.52 | 38.11 | 74.77 | 49.89 |
| Wan 2.2-AC | 58.65 | 56.57 | 48.23 | 60.55 | 53.41 | 86.16 | 54.40 |
| Wan 2.2-AC + IMPACT | 62.46 | 60.60 | 53.28 | 60.29 | 55.87 | 92.56 | 60.00 |
| Cosmos-Predict 2.5 (action) | 55.91 | 57.87 | 38.89 | 68.73 | 42.23 | 82.53 | 49.51 |
| Cosmos-Predict 2.5 (action) + IMPACT | 62.53 | 56.32 | 68.22 | 57.90 | 41.82 | 77.99 | 71.19 |
| Model | FVD ↓ | FID ↓ | CLIP-Hand ↑ | Hand IoU ↑ |
|---|---|---|---|---|
| VACE | 358.42 | 50.65 | 0.895 | 0.493 |
| Wan 2.2-AC | 366.12 | 44.71 | 0.921 | 0.693 |
| Wan 2.2-AC + IMPACT | 110.94 | 5.79 | 0.952 | 0.772 |
| Configuration | EWMScore | Visual | Motion | Physics |
|---|---|---|---|---|
| Wan 2.2-AC | 58.65 | 56.57 | 48.23 | 53.41 |
| + IWS | 61.54 | 60.44 | 49.16 | 55.16 |
| + IWS + ADS | 62.46 | 60.60 | 53.28 | 55.87 |
@article{tang2026impact,
title = {IMPACT: Attention Is the Interaction Map for Scalable Interaction-Aware World Model Training},
author = {Tang, Rongze and Fang, Jianjie and Wang, Zhaolu and Wang, Ziyou and Liu, Xvyuan and Su, Haisheng and Zhang, Xin and Wu, Wei and Gao, Chen and Li, Yong and Chen, Zhibo},
journal = {arXiv preprint arXiv:2609.00161},
year = {2026},
url = {https://arxiv.org/abs/2609.00161}
}