Audio Source Separation Using Angular Location
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio source separation technologies face challenges in identifying and isolating a target signal from a mixture of voices in real-world scenarios, particularly in environments with multiple concurrent speakers, where the target speaker's voice may not be the loudest or most prominent.
Innovation Solution
A deep learning-based system utilizing a neural network that determines the angle of arrival of the target signal from a microphone array, allowing for the separation and enhancement of the desired audio signal by steering a virtual microphone direction towards the selected speaker, while canceling out other sounds through spectral and spatial characteristics analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional audio source separation methods are used, then the system can process mixed audio signals, but it fails to accurately identify and isolate the target speaker's voice when multiple speakers are present simultaneously
Solution Approach 1:
The patent introduces spatial dimension (angle of arrival) as an additional feature dimension for source separation. By incorporating angular information from microphone arrays, the system transforms the problem from purely spectral analysis to spatio-spectral analysis, enabling accurate identification of target speakers in multi-speaker scenarios through directional discrimination
Solution Approach 2:
The patent employs angle of arrival estimation as an intermediary step between raw microphone signals and final source separation. This intermediate spatial parameter serves as a mediator that guides the separation process by indicating the directional location of the target speaker, thereby improving identification accuracy in complex acoustic environments
2Reliability
If the system enhances the target speaker's voice, then communication clarity improves, but the system complexity increases due to the need for angle of arrival determination and spatial filtering
Solution Approach 1:
The patent divides the source separation task into distinct functional segments: angle of arrival estimation module, spectral analysis module, and spatial filtering module. This segmentation allows each component to specialize in a specific aspect of the problem, improving overall reliability while making the complex system more manageable and modular
Solution Approach 2:
The patent replaces traditional mechanical beamforming approaches with deep learning-based spatial filtering. The neural network learns optimal spatial filters directly from data, substituting complex mechanical signal processing with adaptive computational methods that achieve similar or better results with improved flexibility
Data Source
AI summary
Systems and methods for audio source separation. A deep learning-based system uses an azimuth angle location to separate an audio signal originating from a selected location from other sound. Techniques are disclosed for steering a virtual direction of a microphone towards a selected speaker. A deep-learning based audio regression method, which can be implemented as a neural network, learns to separate out various speakers by leveraging spectral and spatial characteristics of all sources. The neural network can focus on multiple sources in multiple respective target directions, and cancel out other sounds. A user can choose which source to listen to. The network can use the time-domain signal and a frequency-domain signal to separate out the target signal and generate a separated audio output. The direction of the selected speaker relative to the microphone array can be input to the system as a vector.


