Target Speech Extraction via Nullformer and ICA
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition systems face performance degradation due to high calculation overhead and noise contamination, particularly when estimating the direction of arrival of target speech signals, leading to inefficient target speech extraction and increased processing time.
Innovation Solution
A method that utilizes information on the direction of arrival of the target speech source to generate a nullformer for removing the target speech signal and estimating noise, employing independent component analysis (ICA) with a cost function to minimize dependency between real and dummy outputs, thereby extracting the target speech signal efficiently.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If independent component analysis is performed to extract target speech signal from multiple input signals, then the target speech signal can be extracted, but the calculation amount increases and processing time is extended
Solution Approach 1:
The patent applies preliminary action by using direction of arrival (DOA) information to pre-identify which output signal corresponds to the target speech source before performing full ICA processing. This preliminary identification step allows the system to skip unnecessary calculation and directly select the target speech signal, thereby reducing processing time while maintaining extraction accuracy.
2Measurement precision
If direction information is used to generate separating matrix for target signal estimation, then the target speech signal can be more accurately separated, but the calculation amount increases
Solution Approach 1:
The patent applies local quality by using direction of arrival information specifically for identifying the target speech output signal rather than applying complex processing to all signals uniformly. This localized approach focuses computational resources only where needed (in the direction of the target source) while using simpler selection criteria for other signals, thereby reducing overall calculation complexity while maintaining separation accuracy.
3Object-affected harmful factors
If noise power spectrum is estimated by ICA using projection-back method and subtracted, then noise can be removed, but the performance of speech recognition is degraded because the target speech signal output still includes noise
Solution Approach 1:
The patent applies the taking out principle by directly extracting the target speech signal from the mixed input signals using ICA and DOA information, rather than attempting to remove noise through subtraction methods. This extraction approach separates the target speech from noise at the source level, providing a cleaner signal that improves speech recognition performance compared to noise subtraction methods that leave residual noise.
4Measurement precision
If additional transformation of input mixing vectors is performed in SBSE method, then the target speech signal can be more accurately estimated, but the calculation amount increases in comparison with other methods
Solution Approach 1:
The patent applies partial action by performing only the necessary transformation of input mixing vectors using direction of arrival information, rather than applying the full SBSE transformation process. This partial transformation approach achieves sufficient target signal estimation accuracy while significantly reducing the calculation amount and improving processing efficiency compared to complete SBSE implementation.
Data Source
AI summary
Provided is a target speech signal extraction method for robust speech recognition including: (a) receiving information on a direction of arrival of the target speech source with respect to the microphones; (b) generating a nullformer by using the information on the direction of arrival of the target speech source to remove the target speech signal from the input signals and to estimate noise; (c) setting a real output of the target speech source using an adaptive vector w(k) as a first channel and setting a dummy output by the nullformer as a remaining channel; (d) setting a cost function for minimizing dependency between the real output of the target speech source and the dummy output using the nullformer by performing independent component analysis (ICA); and (e) estimating the target speech signal by using the cost function, thereby extracting the target speech signal from the input signals.


