Spatial Audio Rendering for Video Conferencing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current communication systems, such as videoconferencing and teleconferencing, face challenges in providing spatial audio, leading to reduced collaboration quality, speech intelligibility, and overall user experience due to spatial collisions caused by low angular separation of estimated directions of arrival of multiple speakers.
Innovation Solution
A communication system that includes a first computing device with an array of audio output devices and a processor, which receives transmitted speech data and metadata describing the estimated direction of arrival of speech from a microphone array at a second computing device, and renders audio at an array of audio devices to eliminate spatial collisions by repositioning audio based on the estimated directions of arrival.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If audio from multiple speakers is rendered using estimated directions of arrival, then spatial audio positioning is improved, but spatial collisions occur due to low angular separation causing reduced speech intelligibility
Solution Approach 1:
The patent transitions from 2D angular separation to 3D spherical coordinate representation of speaker positions. By using azimuth and elevation angles together with distance information, speakers that are angularly close in 2D can be spatially separated in 3D space, eliminating collisions while preserving speech intelligibility through accurate spatial positioning.
Solution Approach 2:
The system dynamically adjusts audio rendering parameters based on real-time estimated directions of arrival. When speakers are detected with low angular separation, the system adaptively modifies their spatial positioning and audio characteristics to prevent collisions, maintaining both spatial accuracy and speech intelligibility under varying conference conditions.
2Reliability
If audio rendering maintains accurate spatial positioning of multiple speakers, then collaboration quality is improved, but spatial collisions reduce overall user experience
Solution Approach 1:
The patent extracts collision detection and resolution as separate processing steps from the audio rendering pipeline. By identifying overlapping spatial positions beforehand and resolving them through coordinate transformation and positional adjustment, the system eliminates harmful spatial collisions while preserving accurate spatial positioning for high-quality collaboration.
Solution Approach 2:
The system performs preliminary spatial collision detection before audio rendering by analyzing estimated directions of arrival and predicting potential overlaps. It proactively adjusts speaker positioning and applies corrective rendering techniques to prevent spatial collisions before they degrade user experience, maintaining collaboration quality throughout the conference.
Data Source
AI summary
A communication system may include, in an example, a first computing device communicatively coupled, via a network, to at least a second computing device maintained at a geographically distinct location than the first computing device; the first computing device including: an array of audio output devices and a processor to receive transmitted speech data and metadata describing an estimated direction of arrival (DOA) of speech from a plurality of speakers at an array of microphones at the second computing device and render audio at the array of audio output devices associated with the first computing device by eliminating spatial collision during rendering; said spatial collision arising due to the low angular separation of the estimated DOA of a plurality of speakers.


