Spatial Audio Processing Using Hybrid Direct Diffuse Component Separation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing spatial audio processing technologies face challenges in achieving high spatial resolution with a limited number of microphones and processor bandwidth, often resulting in processing artifacts, especially under reverberant conditions.
Innovation Solution
A hybrid approach combining linear and parametric renderers to separate direct and diffuse sound components, where the direct component is processed by a parametric renderer and the diffuse component by a linear renderer, with band-splitting and DDR/DRR adjustment to minimize artifacts and enhance spatial resolution.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a high number of microphones is used to achieve high spatial resolution, then spatial resolution is improved, but device complexity and cost increase
Solution Approach 1:
The patent changes the processing parameters by separating the sound field into direct and diffuse components, then applying different rendering approaches (parametric vs linear) to each component. This allows achieving high spatial resolution with fewer microphones by intelligently processing different sound components rather than relying solely on increasing microphone count
Solution Approach 2:
The patent segments the sound field into direct sound components and diffuse sound components, processing each separately with appropriate rendering techniques. This segmentation allows the system to achieve high spatial resolution for direct sounds while managing computational load and microphone requirements
2Measurement precision
If parametric processing is used to improve spatial resolution with low microphone count, then spatial resolution is improved, but processing artifacts increase under reverberant conditions
Solution Approach 1:
The patent segments the sound field into direct and diffuse components, applying parametric processing only to the direct component where it is most effective. The diffuse component is processed separately, preventing artifacts from contaminating the overall output under reverberant conditions
Solution Approach 2:
The patent applies different processing qualities to different sound components: parametric processing with high spatial resolution is applied locally to the direct sound component, while linear processing is applied to the diffuse component. This local quality approach optimizes performance for each component type while avoiding artifact generation
3Productivity
If linear processing is used to reduce computational load, then processing cost is reduced, but spatial resolution is severely restrained
Solution Approach 1:
The patent changes the processing parameters by using linear processing for the diffuse sound component (which has lower spatial resolution requirements) and parametric processing for the direct component. This parameter differentiation allows the system to reduce overall computational load while maintaining high spatial resolution where it matters most
Solution Approach 2:
The patent applies different processing qualities to different sound components: linear processing with lower computational requirements is applied to the diffuse component, while parametric processing with high spatial resolution is applied to the direct component. This local quality approach optimizes the balance between processing efficiency and spatial resolution
Data Source
AI summary
Processing input audio channels for generating spatial audio can include receiving a plurality of microphone signals that capture a sound field. Each microphone signal can be transformed into a frequency domain signal. From each frequency domain signal, a direct component and a diffuse component can be extracted. The direct component can be processed with a parametric renderer. The diffuse component can be processed with a linear renderer. The components can be combined, resulting in a spatial audio output. The levels of the components can be adjusted to match a direct to diffuse ratio (DDR) of the output with the DDR of the captured sound field. Other aspects are also described and claimed.


