Speech Discrimination via Frequency Band Weighting for Multi-Channel Echo Suppression

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems struggle to accurately extract user speech from acoustic signals that include system sounds with multiple channels, such as stereo or 5.1ch audio, due to residual echoes, which are not effectively handled by current echo cancellation methods.

Innovation Solution

A speech discrimination apparatus that assigns weights to frequency bands based on the system sound's amplitude, suppressing the main elements of the system sound and extracting features from the acoustic signal, thereby reducing the influence of residual echoes, using a weight assignment unit, feature extraction unit, and speech/non-speech discrimination unit.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If frequency spectrum exclusion method is used to reduce echo influence, then speech discrimination accuracy is improved for single-channel system sounds, but the method becomes ineffective for multi-channel system sounds such as stereo music

Engineering Contradiction:
Improvespeech discrimination accuracyVSAvoidapplicability to multi-channel system sounds
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent segments the multi-channel system sound into individual channel components and processes each channel separately through echo cancellation. By dividing the complex multi-channel signal into manageable single-channel segments, the system can apply effective echo cancellation to each channel independently, thereby resolving the contradiction between maintaining speech discrimination accuracy and achieving adaptability to multi-channel sounds.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent merges the processed single-channel signals back into a composite multi-channel output after individual echo cancellation. This combining approach allows the system to maintain the effectiveness of single-channel processing while achieving the goal of handling multi-channel system sounds, thus resolving the contradiction between measurement precision and adaptability.

Inventive Principle:
Principle #5Merging (Combining)

2Adaptability or versatility

If echo cancellation is performed on multi-channel system sounds, then adaptability is improved, but residual echoes remain that degrade speech discrimination performance

Engineering Contradiction:
Improvecapability to handle stereo and 5.1ch audioVSAvoidspeech extraction accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent applies local quality by making the echo cancellation process adaptive to each specific channel's characteristics. Instead of applying a uniform cancellation approach to all channels, the system adjusts the cancellation parameters locally for each channel based on its specific echo properties, thereby reducing residual echoes while maintaining adaptability to different multi-channel formats.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent implements dynamic echo cancellation that adapts to the changing characteristics of multi-channel system sounds in real-time. The system dynamically adjusts cancellation parameters based on the detected echo properties of each channel, enabling effective speech discrimination even when handling diverse multi-channel audio formats with varying echo characteristics.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS9330682B2Apparatus and method for discriminating speech, and computer readable medium
Publication Date: 2016.05.03 TOSHIBA DIGITAL SOLUTIONS CORP
  • US9330682B2 patent drawing
  • US9330682B2 patent drawing
  • US9330682B2 patent drawing

AI summary

According to one embodiment, an apparatus for discriminating speech/non-speech of a first acoustic signal includes a weight assignment unit, a feature extraction unit, and a speech/non-speech discrimination unit. The first acoustic signal includes a user's speech and a reproduced sound. The reproduced sound is a system sound having a plurality of channels reproduced from a plurality of speakers. The weight assignment unit is configured to assign a weight to each frequency band based on the system sound. The feature extraction unit is configured to extract a feature from a second acoustic signal based on the weight of each frequency band. The second acoustic signal is the first acoustic signal in which the reproduced sound is suppressed. The speech/non-speech discrimination unit is configured to discriminate speech/non-speech of the first acoustic signal based on the feature.