CVPR · 2025
WildAvatar: Learning In-the-wild 3D Avatars from the Web
Key takeaway
WildAvatar scales 3D human avatar data beyond laboratory capture by automatically annotating and filtering web videos with more than 10,000 subjects and scenes. arXiv abstract · v4
Abstract
Existing research on avatar creation is typically limited to laboratory datasets, which require high costs against scalability and exhibit insufficient representation of the real world. On the other hand, the web abounds with off-the-shelf real-world human videos, but these videos vary in quality and require accurate annotations for avatar creation. To this end, we propose an automatic annotating pipeline with filtering protocols to curate these humans from the web. Our pipeline surpasses state-of-the-art methods on the EMDB benchmark, and the filtering protocols boost verification metrics on web videos. We then curate WildAvatar, a web-scale in-the-wild human avatar creation dataset extracted from YouTube, with 10000+ different human subjects and scenes. WildAvatar is at least 10× richer than previous datasets for 3D human avatar creation and closer to the real world. To explore its potential, we demonstrate the quality and generalizability of avatar creation methods on WildAvatar. We will publicly release our code, data source links and annotations to push forward 3D human avatar creation and other related fields for real-world applications.
Author abstract · arXiv abstract · v4
Publication
IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2025
Paper and resources
Research topics
WildAvatar · 3D human avatars · In-the-wild reconstruction · Web video datasets · Human motion annotation · Avatar generalization · Automatic data curation · Human reconstruction
Research problem and approach
Laboratory avatar datasets are expensive and poorly represent real-world variation. WildAvatar combines automatic annotation with filtering protocols to turn diverse YouTube footage into data for in-the-wild avatar creation. arXiv abstract · v4
Main contributions
- Develops an annotation and quality-filtering pipeline for unconstrained human videos. arXiv abstract · v4
- Curates a large in-the-wild dataset and evaluates avatar reconstruction quality and generalization. arXiv abstract · v4
Method comparison
| Approach | Key difference |
|---|---|
| Laboratory avatar datasets | Requires expensive controlled capture and offers limited representation of unconstrained scenes. |
| WildAvatar | Automatically annotates and filters diverse web videos to scale in-the-wild avatar training data. |
Selected results
- The pipeline filters 465,801 candidate clips to 10,647 qualified clips. Across filtering and final refinement, PCK at threshold 0.1 rises from 0.282 to 0.921, while the out-of-mask SMPL overlap measure falls from 0.760 to 0.028. Table 3 · web-video annotation pipeline · arXiv v4
- For the Gaussian-Human method in the novel-pose evaluation, replacing HMR2.0 annotations with WildAvatar annotations raises PSNR from 24.73 to 25.89 dB. Every one of the seven avatar methods listed in the table gains PSNR with the proposed annotations. Table 4 · downstream novel-pose synthesis · arXiv v4
Cite this paper
Zihao Huang, Shoukang Hu, Guangcong Wang, Tianqi Liu, Yuhang Zang, Zhiguo Cao, Wei Li, Ziwei Liu. WildAvatar: Learning In-the-wild 3D Avatars from the Web. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025, pp. 15963–15975.
@inproceedings{arxiv240702165,
title = {{WildAvatar: Learning In-the-wild 3D Avatars from the Web}},
author = {Zihao Huang and Shoukang Hu and Guangcong Wang and Tianqi Liu and Yuhang Zang and Zhiguo Cao and Wei Li and Ziwei Liu},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
month = {June},
year = {2025},
pages = {15963--15975},
url = {https://openaccess.thecvf.com/content/CVPR2025/html/Huang_WildAvatar_Learning_In-the-wild_3D_Avatars_from_the_Web_CVPR_2025_paper.html}
}