OmniEcho:具身智慧體空間音訊理解研究與基準
為何重要
The capability to process spatial audio represents a critical step towards multimodal embodied intelligence (e.g., robots interpreting sounds to navigate). OmniEcho validates spatial audio as a complementary signal to vision, particularly valuable in occluded or dark environments. For the industry, this establishes a new evaluation standard (OmniEchoBench) for sensory integration, shifting focus f
Human can effortlessly localize and integrate sound with vision, a capability still lacking in embodied agents. To address this, a research team introduces the OmniEcho model and OmniEchoBench, a benchmark designed to evaluate spatial audio-visual perception in embodied settings.
- OmniEchoBench comprises 197 real-world scenes, 2,972 question-answer pairs, and 900 navigation samples.
- The dataset includes first-order ambisonics (FOA) audio collected from 30 real-world environments.
- OmniEcho uses an FOA spatial encoder and a pretrained semantic audio pathway, achieving performance in sound-guided navigation close to traditional vision-language navigation levels.