Target Speech Extraction Using Nullformer and ICA

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems face performance degradation due to high calculation overhead and noise interference, particularly in real environments where the learning and operational environments differ, and current methods like ICA, BSSA, and SBSE either overburden calculations or fail to accurately extract target speech signals.

Innovation Solution

A target speech signal extraction method that uses information on the direction of arrival of the target speech source to generate a nullformer for removing the target speech signal and estimating noise, employing independent component analysis (ICA) to minimize dependency between real and dummy outputs, thereby reducing calculation load and improving extraction accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If independent component analysis (ICA) is used to extract target speech signal from multiple input signals, then the target speech signal can be extracted, but a process of identifying the direction of each output signal is required which increases calculation amount and degrades performance

Engineering Contradiction:
Improvetarget speech signal extraction accuracyVSAvoidcalculation amount
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by using direction of arrival (DOA) information to pre-identify which output signal corresponds to the target speech source before performing full ICA analysis. This preliminary identification based on spatial information reduces the subsequent calculation burden and avoids the need for exhaustive direction identification of all output signals.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent extracts and utilizes only the necessary component (DOA information) from the input signals to identify the target speech source direction. By focusing on the spatial characteristic of the target signal rather than analyzing all output signal components, the calculation amount is reduced while maintaining extraction accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

2Reliability

If blind spatial subtraction array (BSSA) method is used to remove target speech signal output and estimate noise, then noise power spectrum can be estimated, but the target speech signal output still includes noise and estimation cannot be perfect leading to degraded speech recognition performance

Engineering Contradiction:
Improvenoise estimation accuracyVSAvoidtarget speech signal extraction accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent introduces an intermediary approach by using a nullformer as a mediator between the ICA output and the final target speech signal. The nullformer, designed based on DOA information, selectively nulls the target speech direction while preserving other signals, and this intermediate processing step improves both noise estimation accuracy and target signal extraction quality compared to direct BSSA method.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If semi-blind source estimation (SBSE) method is used with preliminary direction information, then the target speech signal can be more accurately separated, but additional transformation of input mixing vectors increases calculation amount

Engineering Contradiction:
Improvetarget speech signal separation accuracyVSAvoidcalculation amount
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies partial action by using only the necessary DOA information for nullformer design rather than performing complete transformation of all input mixing vectors as required by full SBSE method. This partial application of preliminary information achieves accurate target speech separation while avoiding the excessive calculation burden of complete vector transformation.

Inventive Principle:
Principle #16Partial or excessive action

4Measurement precision

If real-time independent vector analysis (IVA) method is used to overcome permutation problem, then permutation problem across frequency bins is resolved, but one target speech signal still needs to be selected from output signals creating ICA-related problems

Engineering Contradiction:
Improvefrequency bin permutation consistencyVSAvoidtarget speech signal selection complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent resolves the signal selection problem by performing preliminary action with DOA-based identification before IVA processing. By pre-identifying which output corresponds to the target speech source using direction information, the system avoids the complexity of selecting the correct signal from IVA outputs, while still benefiting from IVA's permutation problem solution across frequency bins.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10991362B2Online target-speech extraction method based on auxiliary function for robust automatic speech recognition
Publication Date: 2021.04.27 IND UNIV COOP FOUND SOGANG UNIV
  • US10991362B2 patent drawing
  • US10991362B2 patent drawing
  • US10991362B2 patent drawing

AI summary

Provided is a target speech signal extraction method for robust speech recognition including: receiving information on a direction of arrival of the target speech source with respect to the microphones; generating a nullformer by using the information on the direction of arrival of the target speech source to remove the target speech signal from the input signals and to estimate noise; setting a real output of the target speech source using an adaptive vector as a first channel and setting a dummy output by the nullformer as a remaining channel; setting a cost function for minimizing dependency between the real output of the target speech source and the dummy output using the nullformer by performing independent component analysis (ICA) or independent vector analysis (IVA); setting an auxiliary function to the cost function; and estimating the target speech signal by using the cost function and the auxiliary function.