Target Speech Extraction Using Auxiliary Function Optimization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition systems face performance degradation due to high calculation overhead and noise robustness issues, particularly in real environments, where accurate target speech extraction is hindered by errors in estimating the direction of arrival and noise power spectrum.
Innovation Solution
A target speech signal extraction method that initializes steering vectors and adaptive vectors using independent component analysis (ICA) or independent vector analysis (IVA), minimizing dependency between real and dummy outputs, and updates these vectors to accurately extract the target speech signal with reduced calculation, leveraging information on the direction of arrival of the target speech source.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If independent component analysis (ICA) is used to extract target speech signal from multiple input signals, then the target speech signal can be extracted, but a process of identifying the direction of each output signal is required which increases calculation amount and degrades performance due to estimation errors
Solution Approach 1:
The patent applies preliminary action by pre-initializing the separating matrix using direction of arrival (DOA) information or steering vectors before performing ICA. This preliminary initialization provides a better starting point for the iterative optimization, reducing the number of iterations needed and thereby reducing the overall calculation amount while maintaining extraction accuracy.
Solution Approach 2:
The patent substitutes the mechanical process of identifying directions for each output signal with a mathematical approach using auxiliary functions. Instead of physically or computationally tracing each signal's direction, the method uses an auxiliary function that directly guides the optimization toward the target speech source based on DOA information, replacing a complex mechanical identification process with a more efficient mathematical substitution.
2Object-affected harmful factors
If blind spatial subtraction array (BSSA) method is used to remove target speech signal output and subtract noise power spectrum, then noise can be removed, but the target speech signal output still includes noise and noise power spectrum estimation cannot be perfect which degrades speech recognition performance
Solution Approach 1:
The patent replaces the mechanical subtraction process of BSSA with an auxiliary function-based optimization approach. Instead of subtracting estimated noise power spectra from the signals, the method uses an auxiliary function that directly optimizes the separating matrix to extract the target speech signal while inherently suppressing noise, achieving better noise removal without the degradation caused by imperfect noise estimation.
Solution Approach 2:
The patent implements feedback through the iterative optimization process using the auxiliary function. The algorithm continuously refines the separating matrix by comparing the extracted signals with the expected target speech characteristics, using this feedback to improve the extraction accuracy and noise suppression performance, thereby maintaining high speech recognition performance even in noisy environments.
3Measurement precision
If semi-blind source estimation (SBSE) method is used with preliminary direction information, then the target speech signal can be more accurately separated, but additional transformation of input mixing vectors is required which increases calculation amount
Solution Approach 1:
The patent extracts and utilizes only the essential DOA information or steering vectors needed for initialization, rather than performing additional transformations on the entire input mixing vectors. By taking out and using only the critical directional information for initializing the separating matrix, the method achieves accurate separation without the excessive calculation burden of transforming all input vectors.
4Reliability
If real-time independent vector analysis (IVA) method is used to overcome permutation problem, then permutation problem across frequency bins can be solved, but one target speech signal still needs to be selected from output signals which creates problems in ICA
Solution Approach 1:
The patent applies preliminary action by initializing the separating matrix with DOA information before IVA processing. This preliminary initialization provides a consistent reference frame across all frequency bins, which helps prevent permutation problems from occurring in the first place, rather than just solving them after they occur. This reduces the complexity of selecting the correct target signal from outputs.
Data Source
AI summary
A target speech signal extraction method for robust speech recognition includes: initializing a steering vector for a target speech source and an adaptive vector, setting a real output channel of the target speech source as an output by the adaptive vector, initializing adaptive vectors for a noise and setting a dummy channel as an output by the adaptive vectors for the noise; setting a cost function for minimizing dependency between a real output for the target speech source and a dummy output for the noise; setting an auxiliary function to the cost function, and updating the adaptive vector for the target speech source and the adaptive vectors for the noise by using the auxiliary function and the steering vector; estimating the target speech signal by using the adaptive vector thereby extracting the target speech signal from the input signals; and updating the steering vector for the target speech source.


