Reference-less Echo Mitigation Using Machine Learning Audio Area Estimation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio communication devices face challenges in suppressing interference signals, particularly echoes from local media playback during video or audio conference calls, due to the lack of a reference signal for echo cancellation in shared media playback modes.
Innovation Solution
A two-stage reference-less echo mitigation technique using a machine learning model, such as a deep neural network, to estimate a target audio area and distinguish between target speech and interference signals, combined with an adaptive level-based suppressor to attenuate residual echoes without relying on a reference copy of the interference signal.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional echo cancellation using reference signals is used, then echo suppression is effective, but it fails in shared media playback mode where no reference copy is available
Solution Approach 1:
The patent creates a synthetic reference signal by copying and processing the local media playback audio through an echo path model. This synthetic reference is then used for echo cancellation in the shared media playback mode, effectively replacing the need for a actual reference copy from the remote end.
Solution Approach 2:
The patent introduces an intermediary echo path model that simulates the acoustic characteristics between the local playback device and the remote participant. This intermediary model enables echo cancellation without requiring direct reference signals from the remote end, bridging the gap in shared media playback scenarios.
2Reliability
If media content is used as reference signal for echo cancellation, then echo suppression is enabled, but transmission delays and non-linearity in communication channel prevent effective usage
Solution Approach 1:
The patent performs preliminary action by capturing and processing the local media playback audio signal before it is transmitted to the remote end. The echo path model is pre-configured with acoustic characteristics, allowing the system to generate accurate synthetic reference signals in advance, eliminating the need to wait for delayed remote references.
Solution Approach 2:
The patent substitutes the mechanical communication channel (which introduces delays and non-linearity) with a computational echo path model. This model mathematically represents the acoustic transmission characteristics, replacing the physical delay-prone communication channel with an instantaneous computational approximation.
3Object-affected harmful factors
If audio signals from outside target audio area are suppressed, then echo from local playback is reduced, but target speech intelligibility may be affected
Solution Approach 1:
The patent applies local quality by creating a target audio area with specific acoustic characteristics that are unique to the local environment. The system suppresses only audio signals that do not match these local characteristics, thereby selectively removing echoes while preserving target speech that conforms to the expected acoustic profile of the target area.
Solution Approach 2:
The patent changes parameters such as acoustic characteristics, reverberation patterns, and signal timing to distinguish between target speech and local playback echoes. By analyzing changes in these parameters, the system can identify and suppress echoes while maintaining target speech intelligibility, as the two signal types exhibit different parameter profiles.
Data Source
AI summary
Disclosed is a reference-less echo mitigation or cancellation technique. The technique enables suppression of echoes from an interference signal when a reference version of the interference signal conventionally used for echo mitigation may not be available. A first stage of the technique may use a machine learning model to model a target audio area surrounding a device so that a target audio signal estimated as originating from within the target audio area may be accepted. In contrast, audio signals such as playback of media content on a TV or other interfering signals estimated as originating from outside the target audio area may be suppressed. A second stage of the technique may be a level-based suppressor that further attenuates the residual echo from the output of the first stage based on an audio level threshold. Side information may be provided to adjust the target audio area or the audio level threshold.


