Microphone Array Audio Source Separation via Sector-Based Decoupling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio processing technologies, such as those using directional microphones and microphone arrays, are inefficient and computationally intensive, especially in hand-held devices and consumer electronics, when dealing with multiple overlapping sound sources, as they struggle to accurately separate and cancel out specific sound sources.
Innovation Solution
The method involves determining propagation vectors and beam former outputs for each sound source using microphone arrays, followed by the creation of a decoupling matrix to decorrelate audio data, allowing for effective separation of sound sources without degrading audio quality, even when sources overlap.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional microphone direction detection techniques are used to analyze correlation between signals from different microphones, then sound source direction can be determined, but the computational complexity increases significantly and robustness decreases
Solution Approach 1:
The patent divides the sound field into multiple non-overlapping sectors around the microphone array. Each sector is independently processed to identify sound sources within that specific region. This segmentation approach reduces the overall computational complexity by breaking down the complex global correlation analysis into simpler local sector-based analyses, while maintaining accurate sound source direction determination within each sector.
2Measurement precision
If directional microphones are used to create narrow listening zones for filtering out sounds, then sound separation can be achieved, but the device complexity increases when dealing with multiple sound sources
Solution Approach 1:
Instead of using multiple directional microphones in different spatial dimensions to achieve sound separation, the patent introduces a temporal dimension by processing signals from a single microphone array through sequential sector analysis. The system divides the 360-degree space into sectors and processes each sector's audio data in time sequence, effectively adding a temporal processing dimension that replaces the need for multiple physical directional microphone dimensions.
3Ease of operation
If sector-based sound capture is used with sectors extending from the microphone array, then sound capture is simplified, but sound sources within the same sector cannot be separated
Solution Approach 1:
The patent further segments each sector by dividing it into multiple sub-sectors. This hierarchical segmentation allows the system to maintain the simplicity of sector-based capture while enabling separation of multiple sound sources within the same sector. By creating finer-grained sub-sectors, the system can assign different listening zones to different sound sources even when they are located within the broader boundaries of a single sector.
Solution Approach 2:
The patent implements dynamic adjustment of listening zones within sectors. The system can adaptively modify the boundaries and characteristics of listening zones based on the detected positions and characteristics of sound sources. This dynamic behavior allows the same sector to accommodate multiple sound sources at different times or positions, maintaining operational simplicity while achieving precise sound source separation when needed.
Data Source
AI summary
A system and method for decorrelating audio data. A method includes determining a plurality of propagation vectors for each of a plurality of sound sources based on audio data captured by a plurality of sound capturing devices and a location of each of the plurality of sound sources, wherein the plurality of sound sources and the plurality of sound capturing devices are deployed in a space, wherein the audio data is captured by the plurality of sound capturing devices based on sounds emitted by the plurality of sound sources in the space; determining a plurality of beam former outputs, wherein each beam former output is determined for one of the plurality of sound sources; determining a decoupling matrix based on the plurality of beam former outputs and the propagation vectors; and decorrelating audio data captured by the plurality of sound capturing devices based on the decoupling matrix.


