Binaural reference renderings of ST-AudioQA scenes. Use headphones and disable device-level spatial audio or head tracking.
Music approaching
Front-right → front-left · 1.7 m → 1.3 mAmbulance siren approaching
Front-right → front-left · 1.9 m → 1.3 mHair dryer moving away
Front-left → front-right · 1.2 m → 2.2 mThrobbing sound moving away
Front-left → front-right · 1.3 m → 1.7 mSound events are entities with semantic identities, locations, and trajectories, but current audio-language models usually reason about clips as global event content. Conversely, sound event localization models track source directions over time but offer limited semantic coverage for language reasoning. To address this gap, we introduce ST-AudioQA, a spatio-temporal audio QA dataset and benchmark built from first-order ambisonic (FOA) renderings of static and moving sound sources. Each scene provides source identity, activity, direction, distance, and motion metadata, enabling dense trajectory supervision and questions about what is sounding, where it is, how it moves, and how sources relate. We further propose ST-Audio Encoder, a time-resolved FOA audio encoder that learns event semantics together with source trajectories, and ST-AudioLM, which connects the audio tokens from the encoder to an LLM for spatio-temporal audio QA. Experiments show that this representation improves the semantic-localization tradeoff and yields stronger reasoning performance than static spatial and localization-oriented baselines.
ST-Audio Encoder preserves broad event semantics while learning time-resolved activity, direction, and distance trajectories from FOA features. ST-AudioLM exposes one semantic token and 40 trajectory-aware temporal tokens to the LLM for source-time-space question answering.
| Encoder | Sem. mAP ↑ | DoA MAE ↓ | Dist. MAE ↓ | Traj. Acc@20 ↑ |
|---|---|---|---|---|
| Intensity-based DoA estimator | — | 23.9 | — | 33.0 |
| Spatial-AST-FOA (temporal crop) | 45.2 | 41.4 | 0.64 | — |
| PSELDNets-mACCDOA + sem./dist. heads | 29.7 | 13.7 | 0.38 | 61.4 |
| ST-Audio Encoder | 62.8 | 13.8 | 0.32 | 62.3 |
| Model | Sem. mAP | Sem. Y/N | DoA | Dist. | ΔDoA | ΔDist. | Move |
|---|---|---|---|---|---|---|---|
| Random | 0.8 | 50.0 | 22.9 | 10.0 | 50.0 | 50.0 | 50.0 |
| Qwen2-Audio | 8.8 | 76.3 | 15.9 | 0.0 | 51.2 | 34.3 | 37.2 |
| BAT | 1.0 | 64.6 | 48.5 | 29.8 | 49.7 | 52.8 | 99.4 |
| PSELDNets-mACCDOA + OLMo2 | 9.1 | 89.9 | 72.0 | 39.7 | 51.3 | 62.6 | 73.4 |
| Spatial-AST-FOA + OLMo2 | 25.3 | 93.3 | 71.4 | 44.9 | 78.4 | 82.4 | 99.6 |
| ST-AudioLM | 27.6 | 93.7 | 81.6 | 51.5 | 91.1 | 87.8 | 99.8 |
| Model | Sem. mAP | Sem. Y/N | Src. DoA | Src. Dist. | Src. ΔDoA | Src. ΔDist. | Src. Move | Ground Loc. | Ground Chg. | Ground Move |
|---|---|---|---|---|---|---|---|---|---|---|
| Random | 1.5 | 50.0 | 22.9 | 10.0 | 50.0 | 50.0 | 50.0 | 25.0 | 25.0 | 25.0 |
| Qwen2-Audio | 4.7 | 69.1 | 20.5 | 0.0 | 48.8 | 36.3 | 36.9 | 3.6 | 19.4 | 4.7 |
| BAT | 1.8 | 60.5 | 35.6 | 21.0 | 49.8 | 50.4 | 64.7 | 39.4 | 37.4 | 57.0 |
| PSELDNets-mACCDOA + OLMo2 | 5.4 | 76.1 | 51.2 | 28.2 | 50.5 | 58.3 | 56.1 | 46.1 | 32.8 | 30.1 |
| Spatial-AST-FOA + OLMo2 | 12.6 | 83.2 | 51.0 | 30.9 | 65.2 | 68.8 | 73.9 | 48.8 | 40.2 | 70.0 |
| ST-AudioLM | 14.3 | 83.8 | 54.2 | 32.8 | 71.6 | 70.0 | 66.2 | 49.7 | 47.4 | 59.3 |
| Model | Temporal relation | Movement-conditioned spatial relation | Cross-source trajectory relation | Average |
|---|---|---|---|---|
| Random | 50.0 | 50.0 | 50.0 | 50.0 |
| Qwen2-Audio | 51.6 | 35.9 | 45.6 | 44.4 |
| BAT | 76.3 | 51.6 | 55.3 | 61.1 |
| PSELDNets-mACCDOA + OLMo2 | 86.3 | 50.3 | 51.4 | 62.7 |
| Spatial-AST-FOA + OLMo2 | 80.4 | 55.2 | 54.3 | 63.3 |
| ST-AudioLM | 86.0 | 55.8 | 60.6 | 67.5 |
| Model | Input | Static localization | Relation | Trajectory | Overall |
|---|---|---|---|---|---|
| Random | — | 33.3 | 33.3 | 33.3 | 33.3 |
| Qwen2-Audio | 2 channels as 2 inputs | 36.6 | 52.8 | 26.7 | 37.4 |
| Audio Flamingo 3 | Mono | 42.2 | 22.2 | 13.3 | 37.9 |
| Spatial-Omni | FOA | 46.2 | 41.7 | 10.0 | 42.8 |
| ST-AudioLM | FOA | 47.9 | 63.9 | 60.0 | 50.4 |
ST-Audio Encoder preserves broad event semantics while remaining competitive with localization-oriented trajectory tracking. ST-AudioLM improves dynamic perception and grounding, achieves the best overall Type-C score, and transfers to real recordings. QA entries are percentages and higher is better; DoA denotes direction of arrival.
@article{hyun2026spatio,
title={Spatio-Temporal Audio Language Modeling for Dynamic Sound Sources},
author={Hyun-Bin, Oh and Shimada, Kazuki and Takida, Yuhta and Sung-Bin, Kim and Uesaka, Toshimitsu and Shibuya, Takashi and Lee, Kyeongyoon and Oh, Tae-Hyun and Mitsufuji, Yuki},
journal={arXiv preprint arXiv:2606.14141},
year={2026}
}