Target Speech Extraction Using Auxiliary Function Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems face performance degradation due to high calculation overhead and noise robustness issues, particularly in real environments, where accurate target speech extraction is hindered by errors in estimating the direction of arrival and noise power spectrum.

Innovation Solution

A target speech signal extraction method that initializes steering vectors and adaptive vectors using independent component analysis (ICA) or independent vector analysis (IVA), minimizing dependency between real and dummy outputs, and updates these vectors to accurately extract the target speech signal with reduced calculation, leveraging information on the direction of arrival of the target speech source.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If independent component analysis (ICA) is used to extract target speech signal from multiple input signals, then the target speech signal can be extracted, but a process of identifying the direction of each output signal is required which increases calculation amount and degrades performance due to estimation errors

Engineering Contradiction:
Improvetarget speech extraction accuracyVSAvoidcalculation amount
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by pre-initializing the separating matrix using direction of arrival (DOA) information or steering vectors before performing ICA. This preliminary initialization provides a better starting point for the iterative optimization, reducing the number of iterations needed and thereby reducing the overall calculation amount while maintaining extraction accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent substitutes the mechanical process of identifying directions for each output signal with a mathematical approach using auxiliary functions. Instead of physically or computationally tracing each signal's direction, the method uses an auxiliary function that directly guides the optimization toward the target speech source based on DOA information, replacing a complex mechanical identification process with a more efficient mathematical substitution.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Object-affected harmful factors

If blind spatial subtraction array (BSSA) method is used to remove target speech signal output and subtract noise power spectrum, then noise can be removed, but the target speech signal output still includes noise and noise power spectrum estimation cannot be perfect which degrades speech recognition performance

Engineering Contradiction:
Improvenoise removalVSAvoidspeech recognition performance
Core Design Contradiction:
Object-affected harmful factorsVSReliability

Solution Approach 1:

The patent replaces the mechanical subtraction process of BSSA with an auxiliary function-based optimization approach. Instead of subtracting estimated noise power spectra from the signals, the method uses an auxiliary function that directly optimizes the separating matrix to extract the target speech signal while inherently suppressing noise, achieving better noise removal without the degradation caused by imperfect noise estimation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent implements feedback through the iterative optimization process using the auxiliary function. The algorithm continuously refines the separating matrix by comparing the extracted signals with the expected target speech characteristics, using this feedback to improve the extraction accuracy and noise suppression performance, thereby maintaining high speech recognition performance even in noisy environments.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If semi-blind source estimation (SBSE) method is used with preliminary direction information, then the target speech signal can be more accurately separated, but additional transformation of input mixing vectors is required which increases calculation amount

Engineering Contradiction:
Improvetarget speech signal separation accuracyVSAvoidcalculation amount
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts and utilizes only the essential DOA information or steering vectors needed for initialization, rather than performing additional transformations on the entire input mixing vectors. By taking out and using only the critical directional information for initializing the separating matrix, the method achieves accurate separation without the excessive calculation burden of transforming all input vectors.

Inventive Principle:
Principle #2Taking out (Extraction)

4Reliability

If real-time independent vector analysis (IVA) method is used to overcome permutation problem, then permutation problem across frequency bins can be solved, but one target speech signal still needs to be selected from output signals which creates problems in ICA

Engineering Contradiction:
Improvepermutation problem resolutionVSAvoidoutput signal selection complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by initializing the separating matrix with DOA information before IVA processing. This preliminary initialization provides a consistent reference frame across all frequency bins, which helps prevent permutation problems from occurring in the first place, rather than just solving them after they occur. This reduces the complexity of selecting the correct target signal from outputs.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11694707B2Online target-speech extraction method based on auxiliary function for robust automatic speech recognition
Publication Date: 2023.07.04 IND UNIV COOP FOUND SOGANG UNIV
  • US11694707B2 patent drawing
  • US11694707B2 patent drawing
  • US11694707B2 patent drawing

AI summary

A target speech signal extraction method for robust speech recognition includes: initializing a steering vector for a target speech source and an adaptive vector, setting a real output channel of the target speech source as an output by the adaptive vector, initializing adaptive vectors for a noise and setting a dummy channel as an output by the adaptive vectors for the noise; setting a cost function for minimizing dependency between a real output for the target speech source and a dummy output for the noise; setting an auxiliary function to the cost function, and updating the adaptive vector for the target speech source and the adaptive vectors for the noise by using the auxiliary function and the steering vector; estimating the target speech signal by using the adaptive vector thereby extracting the target speech signal from the input signals; and updating the steering vector for the target speech source.