Target Speech Extraction Using Nullformer and ICA
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition systems face performance degradation due to high calculation overhead and noise interference, particularly in real environments where the learning and operational environments differ, and current methods like ICA, BSSA, and SBSE either overburden calculations or fail to accurately extract target speech signals.
Innovation Solution
A target speech signal extraction method that uses information on the direction of arrival of the target speech source to generate a nullformer for removing the target speech signal and estimating noise, employing independent component analysis (ICA) to minimize dependency between real and dummy outputs, thereby reducing calculation load and improving extraction accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If independent component analysis (ICA) is used to extract target speech signal from multiple input signals, then the target speech signal can be extracted, but a process of identifying the direction of each output signal is required which increases calculation amount and degrades performance
Solution Approach 1:
The patent applies preliminary action by using direction of arrival (DOA) information to pre-identify which output signal corresponds to the target speech source before performing full ICA analysis. This preliminary identification based on spatial information reduces the subsequent calculation burden and avoids the need for exhaustive direction identification of all output signals.
Solution Approach 2:
The patent extracts and utilizes only the necessary component (DOA information) from the input signals to identify the target speech source direction. By focusing on the spatial characteristic of the target signal rather than analyzing all output signal components, the calculation amount is reduced while maintaining extraction accuracy.
2Reliability
If blind spatial subtraction array (BSSA) method is used to remove target speech signal output and estimate noise, then noise power spectrum can be estimated, but the target speech signal output still includes noise and estimation cannot be perfect leading to degraded speech recognition performance
Solution Approach 1:
The patent introduces an intermediary approach by using a nullformer as a mediator between the ICA output and the final target speech signal. The nullformer, designed based on DOA information, selectively nulls the target speech direction while preserving other signals, and this intermediate processing step improves both noise estimation accuracy and target signal extraction quality compared to direct BSSA method.
3Measurement precision
If semi-blind source estimation (SBSE) method is used with preliminary direction information, then the target speech signal can be more accurately separated, but additional transformation of input mixing vectors increases calculation amount
Solution Approach 1:
The patent applies partial action by using only the necessary DOA information for nullformer design rather than performing complete transformation of all input mixing vectors as required by full SBSE method. This partial application of preliminary information achieves accurate target speech separation while avoiding the excessive calculation burden of complete vector transformation.
4Measurement precision
If real-time independent vector analysis (IVA) method is used to overcome permutation problem, then permutation problem across frequency bins is resolved, but one target speech signal still needs to be selected from output signals creating ICA-related problems
Solution Approach 1:
The patent resolves the signal selection problem by performing preliminary action with DOA-based identification before IVA processing. By pre-identifying which output corresponds to the target speech source using direction information, the system avoids the complexity of selecting the correct signal from IVA outputs, while still benefiting from IVA's permutation problem solution across frequency bins.
Data Source
AI summary
Provided is a target speech signal extraction method for robust speech recognition including: receiving information on a direction of arrival of the target speech source with respect to the microphones; generating a nullformer by using the information on the direction of arrival of the target speech source to remove the target speech signal from the input signals and to estimate noise; setting a real output of the target speech source using an adaptive vector as a first channel and setting a dummy output by the nullformer as a remaining channel; setting a cost function for minimizing dependency between the real output of the target speech source and the dummy output using the nullformer by performing independent component analysis (ICA) or independent vector analysis (IVA); setting an auxiliary function to the cost function; and estimating the target speech signal by using the cost function and the auxiliary function.


