Voice Processing Attenuating Remote Echo in Double-Talk
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing echo cancellation systems fail to effectively suppress residual echoes in double-talk states, leading to weakened near-end voice signals and compromised communication quality due to non-linear processor (NLP) limitations and automatic gain control (AGC) amplification of residual echoes.
Innovation Solution
A method and device that detect the communication system's working state, attenuate remote-end voice signals, perform echo processing to obtain near-end and residual echo signals, apply non-linear suppression, and control gain to effectively suppress residual echoes, thereby improving communication quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If NLP is used to suppress residual echoes, then echo suppression capability is improved, but near-end voice signal is weakened in double-talk state
Solution Approach 1:
The system performs preliminary classification of the working state (single-talk, double-talk, or silent state) before applying NLP processing. By detecting the working state in advance and only enabling NLP in single-talk states, the system avoids weakening near-end voice signals during double-talk while still suppressing residual echoes when appropriate.
Solution Approach 2:
The NLP processing is dynamically enabled or disabled based on the detected working state. The system transitions between different processing modes (NLP enabled/disabled) according to real-time state detection, making the echo suppression adaptive to current communication conditions rather than applying a fixed processing approach.
2Manufacturing precision
If AGC is added to amplify near-end voice signal, then voice signal quality is improved, but residual echoes are also amplified
Solution Approach 1:
The system performs preliminary working state detection to determine whether the current state is single-talk or double-talk before enabling AGC. By classifying the state in advance, the system ensures AGC is only activated when near-end speech is present and echoes are not simultaneously generated, preventing echo amplification.
Solution Approach 2:
The system uses feedback from the working state detector to control AGC activation. The detection result feeds back to the AGC module, creating a closed-loop control system that enables or disables gain control based on real-time analysis of the communication state, thereby preventing residual echo amplification.
3Object-affected harmful factors
If self-adapting filter is used for echo cancelation, then linear echo suppression is improved, but non-linear residual echoes remain
Solution Approach 1:
The echo cancellation process is segmented into two distinct stages: first the self-adapting filter handles linear echoes, then NLP processes non-linear residual echoes. This segmentation allows each processing stage to specialize in different types of echo components, improving overall cancellation effectiveness.
Solution Approach 2:
NLP acts as an intermediary processing stage between the self-adapting filter and the final output. It receives the signal after linear echo cancellation and specifically targets non-linear residual echoes, serving as a mediator that bridges the gap left by linear filtering alone.
Data Source
AI summary
Provided in the present disclosure are a voice processing method, an apparatus, an electronic device, and a storage medium, the method comprising: detecting the working state of a current call system, and when the working state is a two-end speaking state or a remote-end speaking state, performing compression processing on a subsequent remote-end voice signal, acquiring a near-end voice signal by means of a microphone, performing echo processing on the basis of the near-end voice signal and the compression-processed remote-end voice signal to obtain an echo-processed near-end voice signal and a remaining echo signal, performing non-linear suppression processing on the near-end voice signal and the remaining echo signal, and performing gain control on the suppression-processed near-end voice signal.


