Audio Source Separation via Recognition-Driven Demixing Matrix Adjustment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio source separation methods, such as adaptive beamformers and constrained blind source separation, struggle to adapt to spatial variations of target signals, leading to performance degradation in speech recognition systems.
Innovation Solution
A method and device that apply a demixing matrix to received signals, generate recognition scores, and adjust the matrix using spatial or mask constraints to improve audio source separation by adapting to spatial variations, enhancing the separation of target signals.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If adaptive beamformer technology is used to perform spatial filtering, then audio/speech signals from a specific direction can be enhanced, but the steering direction becomes incorrect due to DoA estimation error
Solution Approach 1:
The patent implements a feedback mechanism where the system performs recognition operations on separated results to generate recognition scores, uses these scores to generate spatial constraints, and then adjusts the demixing matrix accordingly. This closed-loop feedback allows the system to correct initial separation errors and adapt to actual signal characteristics, resolving the reliability issue caused by initial DoA estimation errors.
Solution Approach 2:
The patent transforms the static demixing matrix into a dynamic one that can be adjusted based on recognition scores and spatial constraints. The system continuously updates the demixing matrix to adapt to spatial variations of target signals, making the steering direction more reliable even when initial DoA estimation is inaccurate.
2Measurement precision
If constrained blind source separation method is used to generate the demixing matrix, then multiple audio sources can be separated and permutation problem can be solved, but the method cannot adapt to spatial variation of target signals
Solution Approach 1:
The patent makes the demixing matrix dynamic by allowing it to be adjusted based on recognition scores and generated spatial constraints. This enables the CBSS method to adapt to spatial variations of target signals while maintaining its ability to separate multiple audio sources and solve permutation problems.
Solution Approach 2:
The system performs self-adjustment by using its own recognition scores to generate spatial constraints that automatically refine the demixing matrix. This self-service mechanism allows the system to adapt to spatial variations without external intervention, improving both separation performance and spatial adaptability.
3Adaptability or versatility
If the demixing matrix is adjusted dynamically based on recognition scores and spatial constraints, then adaptability to spatial variation improves, but computational complexity increases
Solution Approach 1:
The system uses its own recognition scores as the basis for generating spatial constraints, eliminating the need for external spatial information or complex environmental modeling. This self-service approach achieves spatial adaptability through internally generated constraints rather than externally complex processing.
Solution Approach 2:
The patent changes the parameters of the demixing matrix based on recognition scores and spatial constraints rather than completely recalculating the matrix. This parameter adjustment approach reduces computational complexity compared to full matrix recalculations while maintaining adaptability to spatial variations.
Data Source
AI summary
A method of audio source separation includes steps of applying a demixing matrix on a plurality of received signals to generate a plurality of separated results; performing a recognition operation on the plurality of separated results to generate a plurality of recognition scores; generating a constraint according to the plurality of recognition scores; and adjusting the demixing matrix according to the constraint; where the adjusted demixing matrix is applied to the plurality of received signals to generate a plurality of updated separated results from the plurality of received signals.


