Spatial Audio Filtering for Multi-Source Direction Separation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing spatial audio capture methods struggle to accurately distinguish and filter multiple simultaneous sound sources in complex audio environments, leading to inefficient amplification and attenuation of desired and undesired sound directions, which affects the quality of the perceived audio experience.
Innovation Solution
Implementing a method that estimates two sound source directions and their energy ratios per frequency band, using multiple microphone configurations, to generate improved filtering gains and attenuations, incorporating previous frame's DOA estimates and energy ratios for enhanced spatial filtering.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional spatial audio capture methods are used, then the system complexity remains low, but the ability to distinguish and filter multiple simultaneous sound sources deteriorates
Solution Approach 1:
The patent segments the audio signal processing into multiple frequency bands and applies separate direction of arrival estimation for each band. This allows the system to handle multiple simultaneous sound sources by processing them independently in different frequency ranges, improving measurement precision without overwhelming the system with monolithic processing complexity.
Solution Approach 2:
The patent extends the processing from single-band to multi-band frequency domain analysis, adding a frequency dimension to the spatial filtering process. This enables the system to distinguish multiple sound sources that may be overlapping in the time domain by separating them across frequency bands, thereby improving source separation accuracy.
2Reliability
If spatial filtering is applied to enhance desired sound directions, then the audio quality improvement increases, but the risk of filter leakage and undesired attenuation increases
Solution Approach 1:
The patent applies different filtering characteristics to different frequency bands based on the local acoustic environment and sound source characteristics. By estimating direction of arrival separately for each frequency band and applying band-specific filtering, the system achieves reliable audio quality enhancement while minimizing filter leakage through localized, adaptive filtering rather than uniform global filtering.
Solution Approach 2:
The patent implements dynamic filtering that adapts to changing acoustic environments by continuously updating direction of arrival estimates and adjusting filter parameters accordingly. This dynamic approach allows the system to maintain reliable audio quality while reducing filter leakage as conditions change, rather than using static filtering that may cause harmful artifacts.
3Productivity
If multiple sound source directions are estimated per frequency band, then the audio zooming effect improves, but the processing time and computational load increase
Solution Approach 1:
The patent divides the computational task into segments by processing different frequency bands independently and in parallel. This segmentation allows the system to estimate multiple sound source directions efficiently by distributing the computational load across frequency bands, improving overall processing efficiency while maintaining the ability to provide accurate audio zooming effects.
Data Source
AI summary
An apparatus including circuitry configured to: obtain two or more audio signals from respective two or more microphones; determine, in one or more frequency band of the two or more audio signals, a first sound source direction parameter and first sound source energy parameter based on processing of the two or more audio signals; determine, in the one or more frequency band of the two or more audio signals, a second sound source direction parameter and second sound source energy parameter based on processing of the two or more audio signals; obtain a region defining a direction and/or range for a filter; and generate the filter to be applied to the two or more audio signals, wherein filter gain/attenuation parameters are generated based on the region in relation to the first sound source direction parameter, the first sound source energy parameter, the second sound source direction parameter and the second sound source energy parameter.


