ICLR · 2025
MotionClone: Training-Free Motion Cloning for Controllable Video Generation
Key takeaway
MotionClone transfers motion from a reference video without training by using sparse temporal attention as guidance for text-to-video and image-to-video generation. arXiv abstract · v6
Abstract
Motion-based controllable video generation offers the potential for creating captivating visual content. Existing methods typically necessitate model training to encode particular motion cues or incorporate fine-tuning to inject certain motion patterns, resulting in limited flexibility and generalization. In this work, we propose MotionClone, a training-free framework that enables motion cloning from reference videos to versatile motion-controlled video generation, including text-to-video and image-to-video. Based on the observation that the dominant components in temporal-attention maps drive motion synthesis, while the rest mainly capture noisy or very subtle motions, MotionClone utilizes sparse temporal attention weights as motion representations for motion guidance, facilitating diverse motion transfer across varying scenarios. Meanwhile, MotionClone allows for the direct extraction of motion representation through a single denoising step, bypassing the cumbersome inversion processes and thus promoting both efficiency and flexibility. Extensive experiments demonstrate that MotionClone exhibits proficiency in both global camera motion and local object motion, with notable superiority in terms of motion fidelity, textual alignment, and temporal consistency.
Author abstract · arXiv abstract · v6
Publication
International Conference on Learning Representations (ICLR), 2025
Paper and resources
Research topics
MotionClone · Motion-controlled video generation · Training-free motion transfer · Temporal attention · Camera motion · Object motion · Text-to-video · Image-to-video
Research problem and approach
Motion-controlled generation often requires specialized training or fine-tuning. MotionClone identifies dominant temporal-attention components as useful motion representations and extracts them in a single denoising step without a full inversion process. arXiv abstract · v6
Main contributions
- Represents reference motion with sparse temporal-attention weights. arXiv abstract · v6
- Supports motion transfer across scenarios with efficient single-step motion extraction. arXiv abstract · v6
Method comparison
| Approach | Key difference |
|---|---|
| Training- or fine-tuning-based motion control | Learns or injects specific motion patterns through model updates. |
| MotionClone | Extracts sparse temporal attention from a reference video and guides generation without model training. |
Selected results
- MotionClone obtains textual alignment 0.3187 and temporal consistency 0.9621, compared with VMC’s 0.3134 and 0.9614. These automatic metrics are distinct from the human-rating rows in the table. Table 1 · automatic motion-cloning evaluation · arXiv v6
- MotionClone’s user-study scores are 3.69 for motion preservation, 4.31 for appearance diversity and 4.28 for temporal consistency, compared with VMC’s 2.59, 3.51 and 2.85. These are rating scores, not percentages. Table 1 · human evaluation · arXiv v6
Cite this paper
Pengyang Ling, Jiazi Bu, Pan Zhang, Xiaoyi Dong, Yuhang Zang, Tong Wu, Huaian Chen, Jiaqi Wang, Yi Jin. MotionClone: Training-Free Motion Cloning for Controllable Video Generation. International Conference on Learning Representations, 2025, 2025, pp. 75579–75601.
@inproceedings{arxiv240605338,
title = {{MotionClone: Training-Free Motion Cloning for Controllable Video Generation}},
author = {Pengyang Ling and Jiazi Bu and Pan Zhang and Xiaoyi Dong and Yuhang Zang and Tong Wu and Huaian Chen and Jiaqi Wang and Yi Jin},
booktitle = {International Conference on Learning Representations},
year = {2025},
volume = {2025},
pages = {75579--75601},
url = {https://proceedings.iclr.cc/paper_files/paper/2025/hash/bc82dbfbfa43232be85b8d9838f49c3e-Abstract-Conference.html}
}