ICLR · 2026

STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence

Zihan Liu, Zhikang Niu, Qiuyang Xiao, Zhisheng Zheng, Ruoqi Yuan, Yuhang Zang(corresponding author), Yuhang Cao, Xiaoyi Dong, Jianze Liang, Xie Chen, Leilei Sun, Dahua Lin, Jiaqi Wang(corresponding author)

Corresponding author

Key takeaway

STAR-Bench probes audio reasoning that text captions cannot adequately replace, exposing weaknesses in understanding sound dynamics across time and three-dimensional space. arXiv abstract · v2

Abstract

Despite rapid progress in Multi-modal Large Language Models and Large Audio-Language Models, existing audio benchmarks largely test semantics that can be recovered from text captions, masking deficits in fine-grained perceptual reasoning. We formalize audio 4D intelligence that is defined as reasoning over sound dynamics in time and 3D space, and introduce STAR-Bench to measure it. STAR-Bench combines a Foundational Acoustic Perception setting (six attributes under absolute and relative regimes) with a Holistic Spatio-Temporal Reasoning setting that includes segment reordering for continuous and discrete processes and spatial tasks spanning static localization, multi-source relations, and dynamic trajectories. Our data curation pipeline uses two methods to ensure high-quality samples. For foundational tasks, we use procedurally synthesized and physics-simulated audio. For holistic data, we follow a four-stage process that includes human annotation and final selection based on human performance. Unlike prior benchmarks where caption-only answering reduces accuracy slightly, STAR-Bench induces far larger drops (-31.5% temporal, -35.2% spatial), evidencing its focus on linguistically hard-to-describe cues. Evaluating 19 models reveals substantial gaps compared with humans and a capability hierarchy: closed-source models are bottlenecked by fine-grained perception, while open-source models lag across perception, knowledge, and reasoning. Our STAR-Bench provides critical insights and a clear path forward for developing future models with a more robust understanding of the physical world.

Author abstract · arXiv abstract · v2

Publication

International Conference on Learning Representations (ICLR), 2026

Paper and resources

Research topics

STAR-Bench · Audio 4D intelligence · Spatial audio · Temporal audio reasoning · Acoustic perception · Sound localization · Audio-language models · Multimodal evaluation

Research problem and approach

Semantic audio benchmarks can conceal perceptual deficits because captions already contain enough information to answer. STAR-Bench targets acoustic attributes, temporal ordering and spatial relations that demand listening. arXiv abstract · v2

Main contributions

  • Combines foundational acoustic perception with holistic temporal and spatial reasoning tasks. arXiv abstract · v2
  • Uses procedural synthesis, physical simulation and human-validated curation for complementary task families. arXiv abstract · v2

Method comparison

ApproachKey difference
Caption-solvable audio evaluationEmphasizes semantic information that can often be recovered without listening to the original sound.
STAR-BenchTests acoustic attributes and temporal-spatial relations that are difficult to convey fully through captions.

arXiv abstract · v2

Selected results

  • Gemini-2.5-Pro achieves 49.59% macro accuracy on STAR-Bench versus the human reference of 79.11%. Its temporal overall accuracy is 58.52% and spatial overall accuracy is 43.62%, compared with 88.00% and 73.72% for humans. Table 2 · audio reasoning evaluation · arXiv v2
  • For temporal reasoning, Gemini-2.5-Pro’s average accuracy is 58.52%, but its all-correct rate across repeated runs is 34.89%. The two metrics distinguish average correctness from reliable answers on the same examples. Table 5 · repeated-run consistency · arXiv v2

Cite this paper

Zihan Liu, Zhikang Niu, Qiuyang Xiao, Zhisheng Zheng, Ruoqi Yuan, Yuhang Zang, Yuhang Cao, Xiaoyi Dong, Jianze Liang, Xie Chen, Leilei Sun, Dahua Lin, Jiaqi Wang. STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence. International Conference on Learning Representations, 2026, 2026, pp. 134703–134731.

@inproceedings{arxiv251024693,
  title     = {{STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence}},
  author    = {Zihan Liu and Zhikang Niu and Qiuyang Xiao and Zhisheng Zheng and Ruoqi Yuan and Yuhang Zang and Yuhang Cao and Xiaoyi Dong and Jianze Liang and Xie Chen and Leilei Sun and Dahua Lin and Jiaqi Wang},
  booktitle = {International Conference on Learning Representations},
  year      = {2026},
  volume    = {2026},
  pages     = {134703--134731},
  url       = {https://proceedings.iclr.cc/paper_files/paper/2026/hash/d9f8b5abc8e0926539ecbb492af7b2f1-Abstract-Conference.html}
}