ICLR · 2025

MotionClone: Training-Free Motion Cloning for Controllable Video Generation

Pengyang Ling, Jiazi Bu, Pan Zhang, Xiaoyi Dong, Yuhang Zang, Tong Wu, Huaian Chen, Jiaqi Wang, Yi Jin

Key takeaway

MotionClone transfers motion from a reference video without training by using sparse temporal attention as guidance for text-to-video and image-to-video generation. arXiv abstract · v6

Abstract

Motion-based controllable video generation offers the potential for creating captivating visual content. Existing methods typically necessitate model training to encode particular motion cues or incorporate fine-tuning to inject certain motion patterns, resulting in limited flexibility and generalization. In this work, we propose MotionClone, a training-free framework that enables motion cloning from reference videos to versatile motion-controlled video generation, including text-to-video and image-to-video. Based on the observation that the dominant components in temporal-attention maps drive motion synthesis, while the rest mainly capture noisy or very subtle motions, MotionClone utilizes sparse temporal attention weights as motion representations for motion guidance, facilitating diverse motion transfer across varying scenarios. Meanwhile, MotionClone allows for the direct extraction of motion representation through a single denoising step, bypassing the cumbersome inversion processes and thus promoting both efficiency and flexibility. Extensive experiments demonstrate that MotionClone exhibits proficiency in both global camera motion and local object motion, with notable superiority in terms of motion fidelity, textual alignment, and temporal consistency.

Author abstract · arXiv abstract · v6

Publication

International Conference on Learning Representations (ICLR), 2025

Paper and resources

Research topics

MotionClone · Motion-controlled video generation · Training-free motion transfer · Temporal attention · Camera motion · Object motion · Text-to-video · Image-to-video

Research problem and approach

Motion-controlled generation often requires specialized training or fine-tuning. MotionClone identifies dominant temporal-attention components as useful motion representations and extracts them in a single denoising step without a full inversion process. arXiv abstract · v6

Main contributions

Method comparison

ApproachKey difference
Training- or fine-tuning-based motion controlLearns or injects specific motion patterns through model updates.
MotionCloneExtracts sparse temporal attention from a reference video and guides generation without model training.

arXiv abstract · v6

Selected results

  • MotionClone obtains textual alignment 0.3187 and temporal consistency 0.9621, compared with VMC’s 0.3134 and 0.9614. These automatic metrics are distinct from the human-rating rows in the table. Table 1 · automatic motion-cloning evaluation · arXiv v6
  • MotionClone’s user-study scores are 3.69 for motion preservation, 4.31 for appearance diversity and 4.28 for temporal consistency, compared with VMC’s 2.59, 3.51 and 2.85. These are rating scores, not percentages. Table 1 · human evaluation · arXiv v6

Cite this paper

Pengyang Ling, Jiazi Bu, Pan Zhang, Xiaoyi Dong, Yuhang Zang, Tong Wu, Huaian Chen, Jiaqi Wang, Yi Jin. MotionClone: Training-Free Motion Cloning for Controllable Video Generation. International Conference on Learning Representations, 2025, 2025, pp. 75579–75601.

@inproceedings{arxiv240605338,
  title     = {{MotionClone: Training-Free Motion Cloning for Controllable Video Generation}},
  author    = {Pengyang Ling and Jiazi Bu and Pan Zhang and Xiaoyi Dong and Yuhang Zang and Tong Wu and Huaian Chen and Jiaqi Wang and Yi Jin},
  booktitle = {International Conference on Learning Representations},
  year      = {2025},
  volume    = {2025},
  pages     = {75579--75601},
  url       = {https://proceedings.iclr.cc/paper_files/paper/2025/hash/bc82dbfbfa43232be85b8d9838f49c3e-Abstract-Conference.html}
}