Audio Output Masking for Speech Recognition Echo

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Automatic speech recognition systems face interference from acoustic echoes and residual noise, which can degrade the accuracy of speech recognition, especially when music or environmental noise is present.

Innovation Solution

Implementing frequency band masking techniques to filter out specific frequency bands from audio output signals, such as music, to reduce acoustic echo interference, and using complementary filters on the input signal to enhance the detection of user utterances, along with training acoustic models for specific output masks based on genres or user characteristics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Object-affected harmful factors

If acoustic echo cancellation is used to remove echo from microphone signal, then echo interference is reduced, but residual noise remains that interferes with speech recognition

Engineering Contradiction:
Improveacoustic echo interferenceVSAvoidspeech recognition accuracy
Core Design Contradiction:
Object-affected harmful factorsVSReliability

Solution Approach 1:

The patent segments the audio frequency spectrum into multiple frequency bands. By dividing the spectrum, the system can selectively mask specific frequency bands in the output signal that correspond to residual echo frequencies, while preserving other bands. This segmentation allows targeted noise reduction without affecting the entire audio spectrum, thereby improving speech recognition accuracy while maintaining audio quality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different quality characteristics to different parts of the audio signal by implementing frequency-specific masking. Certain frequency bands are masked in the output signal to prevent echo reinforcement, while other bands remain unmasked to preserve audio fidelity. This local quality approach ensures that noise reduction is applied only where necessary, maintaining overall speech recognition reliability.

Inventive Principle:
Principle #3Local quality

2Object-affected harmful factors

If frequency band masking is applied to output signal, then acoustic echo interference is reduced, but audio quality may be degraded

Engineering Contradiction:
Improveacoustic echo interferenceVSAvoidaudio quality
Core Design Contradiction:
Object-affected harmful factorsVSLoss of information

Solution Approach 1:

The patent applies partial masking by selectively masking only specific frequency bands in the output signal rather than applying uniform masking across the entire spectrum. This partial action approach reduces acoustic echo interference in problematic frequency ranges while preserving audio quality in other frequency bands, thus minimizing information loss while achieving the desired noise reduction effect.

Inventive Principle:
Principle #16Partial or excessive action

3Adaptability or versatility

If adaptive algorithms continuously adjust echo estimates, then echo cancellation adapts to environment changes, but computational complexity increases

Engineering Contradiction:
Improveecho cancellation adaptationVSAvoidcomputational complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent performs preliminary frequency band masking on the output signal before audio playback. By pre-identifying and masking the frequency bands that are most likely to contain echo information, the system reduces the computational burden on adaptive algorithms. The adaptive algorithms then only need to adjust estimates within the remaining unmasked bands, reducing overall computational complexity while maintaining adaptability to environmental changes.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9704478B1Audio output masking for improved automatic speech recognition
Publication Date: 2017.07.11 AMAZON TECH INC
  • US9704478B1 patent drawing
  • US9704478B1 patent drawing
  • US9704478B1 patent drawing

AI summary

Features are disclosed for filtering portions of an output audio signal in order to improve automatic speech recognition on an input signal which may include a representation of the output signal. A signal that includes audio content can be received, and a frequency or band of frequencies can be selected to be filtered from the signal. The frequency band may correspond to a desired frequency band for speech recognition. An input signal can be obtained comprising audio data corresponding to a user utterance and presentation of the output signal. Automatic speech recognition can be performed on the input signal. In some cases, an acoustic model trained for use with such frequency band filtering may be used to perform speech recognition.