Target Speech Extraction via Nullformer and ICA

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems face performance degradation due to high calculation overhead and noise contamination, particularly when estimating the direction of arrival of target speech signals, leading to inefficient target speech extraction and increased processing time.

Innovation Solution

A method that utilizes information on the direction of arrival of the target speech source to generate a nullformer for removing the target speech signal and estimating noise, employing independent component analysis (ICA) with a cost function to minimize dependency between real and dummy outputs, thereby extracting the target speech signal efficiently.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If independent component analysis is performed to extract target speech signal from multiple input signals, then the target speech signal can be extracted, but the calculation amount increases and processing time is extended

Engineering Contradiction:
Improvetarget speech extraction accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by using direction of arrival (DOA) information to pre-identify which output signal corresponds to the target speech source before performing full ICA processing. This preliminary identification step allows the system to skip unnecessary calculation and directly select the target speech signal, thereby reducing processing time while maintaining extraction accuracy.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If direction information is used to generate separating matrix for target signal estimation, then the target speech signal can be more accurately separated, but the calculation amount increases

Engineering Contradiction:
Improvetarget speech separation accuracyVSAvoidcalculation complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies local quality by using direction of arrival information specifically for identifying the target speech output signal rather than applying complex processing to all signals uniformly. This localized approach focuses computational resources only where needed (in the direction of the target source) while using simpler selection criteria for other signals, thereby reducing overall calculation complexity while maintaining separation accuracy.

Inventive Principle:
Principle #3Local quality

3Object-affected harmful factors

If noise power spectrum is estimated by ICA using projection-back method and subtracted, then noise can be removed, but the performance of speech recognition is degraded because the target speech signal output still includes noise

Engineering Contradiction:
Improvenoise contaminationVSAvoidspeech recognition performance
Core Design Contradiction:
Object-affected harmful factorsVSReliability

Solution Approach 1:

The patent applies the taking out principle by directly extracting the target speech signal from the mixed input signals using ICA and DOA information, rather than attempting to remove noise through subtraction methods. This extraction approach separates the target speech from noise at the source level, providing a cleaner signal that improves speech recognition performance compared to noise subtraction methods that leave residual noise.

Inventive Principle:
Principle #2Taking out (Extraction)

4Measurement precision

If additional transformation of input mixing vectors is performed in SBSE method, then the target speech signal can be more accurately estimated, but the calculation amount increases in comparison with other methods

Engineering Contradiction:
Improvetarget signal estimation accuracyVSAvoidprocessing efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent applies partial action by performing only the necessary transformation of input mixing vectors using direction of arrival information, rather than applying the full SBSE transformation process. This partial transformation approach achieves sufficient target signal estimation accuracy while significantly reducing the calculation amount and improving processing efficiency compared to complete SBSE implementation.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10657958B2Online target-speech extraction method for robust automatic speech recognition
Publication Date: 2020.05.19 SOGANG UNIV RES FOUND
  • US10657958B2 patent drawing
  • US10657958B2 patent drawing
  • US10657958B2 patent drawing

AI summary

Provided is a target speech signal extraction method for robust speech recognition including: (a) receiving information on a direction of arrival of the target speech source with respect to the microphones; (b) generating a nullformer by using the information on the direction of arrival of the target speech source to remove the target speech signal from the input signals and to estimate noise; (c) setting a real output of the target speech source using an adaptive vector w(k) as a first channel and setting a dummy output by the nullformer as a remaining channel; (d) setting a cost function for minimizing dependency between the real output of the target speech source and the dummy output using the nullformer by performing independent component analysis (ICA); and (e) estimating the target speech signal by using the cost function, thereby extracting the target speech signal from the input signals.