Acoustic Echo Cancellation Using Speech Recognition Feedback
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In voice communications, the near-to-far ratio is compromised due to sound from the loudspeaker reaching the microphones simultaneously with the near-side talker's sound, leading to poor bi-directional communication performance during double talk scenarios.
Innovation Solution
An acoustic-echo processing module, enhanced by feedback from an automatic speech recognition engine, processes signals to cancel echo and adjust processing parameters, such as halting adaptive filter adjustments and adjusting volume or frequency, to improve echo cancellation during double talk.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the loudspeaker plays far-end signal, then the far-end talker can be heard by near-side talker, but sound from loudspeaker reaches microphones causing echo and decreasing near-to-far ratio
Solution Approach 1:
The system uses feedback from the speech recognizer to the acoustic-echo processing module. The speech recognizer detects near-side talker speech and provides feedback signals that trigger echo cancellation processing, allowing the system to adaptively respond to double-talk conditions and maintain communication quality
Solution Approach 2:
The acoustic-echo processing module acts as an intermediary between the loudspeaker output and microphone input. It processes the microphone signals to cancel echo components before they are used for further processing, effectively mediating the harmful interaction between loudspeaker and microphone paths
2Manufacturing precision
If adaptive filter continuously adjusts to cancel echo, then echo cancellation performance improves, but during double talk it may diverge and degrade performance
Solution Approach 1:
The speech recognizer provides feedback about the presence of near-side talker speech to the acoustic-echo processing module. This feedback allows the system to detect double-talk conditions and adjust adaptive filter behavior accordingly, preventing divergence while maintaining echo cancellation effectiveness
Solution Approach 2:
The system dynamically adjusts the behavior of the adaptive filter based on real-time conditions. When double-talk is detected via speech recognition feedback, the system modifies filter adjustment rates or parameters to prevent divergence, allowing the filter to be both precise under normal conditions and stable during double-talk
Data Source
AI summary
An automatic speech recognition engine receives an acoustic-echo processed signal from an acoustic-echo processing (AEP) module, where said echo processed signal contains mainly the speech from the near-end talker. The automatic speech recognition engine analyzes the content of the acoustic-echo processed signal to determine whether words or keywords are present. Based upon the results of this analysis, the automatic speech recognition engine produces a value reflecting the likelihood that some words or keywords are detected. Said value is provided to the AEP module. Based upon the value, the AEP module determines if there is double talk and processes the incoming signals accordingly to enhance its performance.


