Sound Source Separation via Spatial Filtering and Regularization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing sound source and speech separation technologies, such as beamformers and independent component analysis, are imprecise and face challenges in real-world scenarios with multiple interfering sources and varying environments, and combining these technologies has not provided significant improvements.
Innovation Solution
A multiple phase process that combines spatial filtering with regularization, using beamformers and nullformers to separate audio signals in the frequency domain, followed by independent component analysis with multi-tap filters to enhance separation, and additional nonlinear spatial filters for improved results.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If beamformers are used for sound source separation, then the system is simple and converges quickly, but the separation precision is insufficient in real-world scenarios with reflections from multiple angles
Solution Approach 1:
The patent segments the sound source separation process into two distinct phases: a spatial filtering phase using beamformers for quick convergence, and a regularization phase using independent component analysis for precise separation. This segmentation allows each phase to optimize for its specific strength while compensating for the other's weaknesses.
Solution Approach 2:
The patent merges beamforming and independent component analysis into a unified two-phase system where beamformer outputs serve as inputs to the ICA mechanism. This combination leverages the fast convergence of beamformers and the high separation precision of ICA to achieve superior overall performance.
2Measurement precision
If independent component analysis is used for sound source separation, then the separation precision is high, but the system is complex and difficult to converge
Solution Approach 1:
The patent applies preliminary spatial filtering using beamformers before independent component analysis. This preliminary action pre-processes the signals to reduce complexity and provide better initial conditions for the ICA mechanism, facilitating faster and more reliable convergence.
Solution Approach 2:
The beamformer acts as an intermediary between the microphone array and the independent component analysis mechanism. It transforms the raw microphone signals into spatially filtered signals that are more suitable for ICA processing, thereby simplifying the overall system complexity while maintaining high separation precision.
3Measurement precision
If independent component analysis is used for sound source separation, then the separation precision is high, but the system takes time to learn coefficients and is sensitive to initial conditions
Solution Approach 1:
The beamformer performs preliminary signal processing that provides better initial conditions for the independent component analysis mechanism. This reduces the learning time required for ICA to converge and makes the system less sensitive to initial conditions by pre-organizing the signal structure.
Solution Approach 2:
The patent uses multi-tap filters that incorporate historical information from previous frames, creating a continuous learning process. This continuity allows the system to build upon previous learning results rather than starting from scratch, significantly reducing the overall learning time while maintaining high separation precision.
4Device complexity
If traditional sound source separation methods are used, then the processing is simple, but the system cannot effectively track moving sources and suppress noise in varying environments
Solution Approach 1:
The patent implements dynamic adaptation by using multi-tap filters that incorporate historical information and allow the system to adapt to changing acoustic environments. The independent component analysis mechanism continuously learns and adjusts to track moving sources, making the system reliable in varying environments despite increased processing complexity.
Solution Approach 2:
The system uses feedback from previous frames through multi-tap filters to continuously refine source separation. This feedback mechanism enables the system to track moving sources effectively and adapt to changing environments by incorporating historical information into the current processing, improving reliability despite increased complexity.
Data Source
AI summary
Described is a multiple phase process/system that combines spatial filtering with regularization to separate sound from different sources such as the speech of two different speakers. In a first phase, frequency domain signals corresponding to the sensed sounds are processed into separated spatially filtered signals including by inputting the signals into a plurality of beamformers (which may include nullformers) followed by nonlinear spatial filters. In a regularization phase, the separated spatially filtered signals are input into an independent component analysis mechanism that is configured with multi-tap filters, followed by secondary nonlinear spatial filters. Separated audio signals are the provided via an inverse-transform.


