Spatial Audio Rendering via Source Separation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing spatial audio technologies face challenges in accurately rendering audio signals based on user device position within a virtual space, as they struggle to fully separate individual sound sources from composite audio signals, leading to degraded audio quality and immersion issues when movement is not seamlessly translated into audio scene changes.
Innovation Solution
The system employs spatial audio capture apparatuses to receive composite audio signals, identifies user device positions, and renders audio differently based on successful separation of individual sound sources within predetermined areas, using measures like correlation between composite and reference signals to determine separation success, enabling volumetric audio rendering and six degrees-of-freedom movement.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If individual audio signals are separated from composite audio signals to enable accurate spatial rendering, then audio quality and immersion are improved, but separation success rate deteriorates when multiple sound sources are present
Solution Approach 1:
The system attempts to separate individual audio signals from composite signals when possible (partial action), but does not require complete separation success for all sound sources. When separation succeeds for some sources, those are rendered with high accuracy while others fall back to alternative methods, achieving useful audio rendering without requiring perfect separation in all cases.
Solution Approach 2:
The system dynamically changes rendering parameters based on separation success. When separation succeeds, parameters are set for high-accuracy individual source rendering; when separation fails, parameters switch to alternative rendering methods. This parameter adaptation resolves the contradiction by adjusting the approach based on actual signal conditions.
2Adaptability or versatility
If volumetric audio rendering is implemented with six degrees-of-freedom movement tracking, then user immersion is improved, but device complexity increases
Solution Approach 1:
The system segments the virtual space into multiple predetermined areas, each with its own spatial audio capture apparatus. This segmentation allows complex volumetric rendering to be implemented in a modular way, where each area can be processed independently, reducing overall system complexity while maintaining immersive capabilities.
Solution Approach 2:
The system implements a universal rendering framework that handles multiple types of audio signals (separated individual sources and composite signals) and multiple user movements (six degrees-of-freedom) through a single integrated apparatus. This multi-functionality achieves high adaptability without proportionally increasing complexity, as the same infrastructure handles diverse rendering scenarios.
3Manufacturing precision
If audio rendering is adjusted based on separation success to maintain quality, then audio consistency is improved, but processing time increases
Solution Approach 1:
The system performs audio signal separation and success evaluation in advance before final rendering. By preprocessing the composite signals and determining separation outcomes beforehand, the system avoids time-consuming processing during actual rendering, thus maintaining audio quality consistency while reducing real-time processing delays.
Data Source
AI summary
An apparatus is disclosed, configured to receive, from first and second spatial audio capture apparatuses, respective first and second composite audio signals comprising components derived from one or more sound sources in a capture space. The apparatus is further configured to identify a position of a user device corresponding to one of first and second areas respectively associated with the positions of the first and second spatial audio capture apparatuses, and to render audio representing the one or more sound sources to the user device, the rendering being performed differently dependent on, for the spatial audio capture apparatus associated with the identified first or second area, whether or not individual audio signals from each of the one or more sound sources can be successfully separated from its composite signal.


