Audio Output Masking for Speech Recognition Echo
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Automatic speech recognition systems face interference from acoustic echoes and residual noise, which can degrade the accuracy of speech recognition, especially when music or environmental noise is present.
Innovation Solution
Implementing frequency band masking techniques to filter out specific frequency bands from audio output signals, such as music, to reduce acoustic echo interference, and using complementary filters on the input signal to enhance the detection of user utterances, along with training acoustic models for specific output masks based on genres or user characteristics.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If acoustic echo cancellation is used to remove echo from microphone signal, then echo interference is reduced, but residual noise remains that interferes with speech recognition
Solution Approach 1:
The patent segments the audio frequency spectrum into multiple frequency bands. By dividing the spectrum, the system can selectively mask specific frequency bands in the output signal that correspond to residual echo frequencies, while preserving other bands. This segmentation allows targeted noise reduction without affecting the entire audio spectrum, thereby improving speech recognition accuracy while maintaining audio quality.
Solution Approach 2:
The patent applies different quality characteristics to different parts of the audio signal by implementing frequency-specific masking. Certain frequency bands are masked in the output signal to prevent echo reinforcement, while other bands remain unmasked to preserve audio fidelity. This local quality approach ensures that noise reduction is applied only where necessary, maintaining overall speech recognition reliability.
2Object-affected harmful factors
If frequency band masking is applied to output signal, then acoustic echo interference is reduced, but audio quality may be degraded
Solution Approach 1:
The patent applies partial masking by selectively masking only specific frequency bands in the output signal rather than applying uniform masking across the entire spectrum. This partial action approach reduces acoustic echo interference in problematic frequency ranges while preserving audio quality in other frequency bands, thus minimizing information loss while achieving the desired noise reduction effect.
3Adaptability or versatility
If adaptive algorithms continuously adjust echo estimates, then echo cancellation adapts to environment changes, but computational complexity increases
Solution Approach 1:
The patent performs preliminary frequency band masking on the output signal before audio playback. By pre-identifying and masking the frequency bands that are most likely to contain echo information, the system reduces the computational burden on adaptive algorithms. The adaptive algorithms then only need to adjust estimates within the remaining unmasked bands, reducing overall computational complexity while maintaining adaptability to environmental changes.
Data Source
AI summary
Features are disclosed for filtering portions of an output audio signal in order to improve automatic speech recognition on an input signal which may include a representation of the output signal. A signal that includes audio content can be received, and a frequency or band of frequencies can be selected to be filtered from the signal. The frequency band may correspond to a desired frequency band for speech recognition. An input signal can be obtained comprising audio data corresponding to a user utterance and presentation of the output signal. Automatic speech recognition can be performed on the input signal. In some cases, an acoustic model trained for use with such frequency band filtering may be used to perform speech recognition.


