Voice Signal Processing Echo Suppression via Coherence Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current echo cancellation technologies in VOIP applications face challenges such as limited adaptive filtering convergence, nonlinear distortion, environmental background noise, and delay jitter, leading to residual echo leakage and poor sound quality.
Innovation Solution
A voice signal processing method that calculates an initial echo loss using a derivative signal and adjusts it with a status coefficient to achieve a target echo loss, while also calculating a proximal-distal coherence value to determine if the echo loss matches a single-talk echo state, thereby updating the status coefficient and performing volume control.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-generated harmful factors
If adaptive filter is used for echo cancellation, then echo suppression is achieved, but residual echo leakage occurs due to limited convergence and environmental noise
Solution Approach 1:
The patent introduces a proximal signal and a distal reference signal as intermediary elements to improve echo estimation. The proximal signal (captured by a proximal microphone) and distal reference signal (captured by a distal microphone) are used to calculate a coherence value that serves as a mediator to detect single-talk echo states, enabling more accurate residual echo suppression than traditional adaptive filtering alone.
Solution Approach 2:
The patent replaces the traditional mechanical adaptive filtering system with a signal processing approach based on coherence calculation. Instead of relying solely on adaptive filter convergence, the system uses coherence between proximal and distal signals to detect echo states, substituting the mechanical filtering process with a more robust statistical signal analysis method.
2Stability of the object's composition
If dynamic range controller is cascaded behind residual echo suppressor, then sound volume stability is improved, but incorrect volume adjustments occur during single-talk with echo leakage
Solution Approach 1:
The patent implements a feedback mechanism where the coherence value calculated from proximal and distal signals feeds back to the dynamic range controller. This feedback allows the system to distinguish between actual voice signals and residual echo, preventing incorrect volume adjustments during single-talk scenarios while maintaining sound volume stability during normal operation.
Solution Approach 2:
The patent performs preliminary coherence calculation and single-talk detection before the dynamic range controller adjusts volume. By预先 determining the echo state using coherence analysis, the system prevents incorrect preliminary actions (volume adjustments) that would otherwise occur when residual echo is mistakenly identified as voice.
3Object-generated harmful factors
If residual echo suppressor is used, then echo suppression is enhanced, but sound quality deteriorates due to rapid volume gain changes
Solution Approach 1:
The coherence value acts as an intermediary that mediates between residual echo suppression and sound quality maintenance. By using coherence to detect true voice presence, the system avoids aggressive echo suppression that would otherwise cause rapid volume gain changes and sound quality deterioration.
Solution Approach 2:
The patent introduces dynamic coherence-based detection that adapts to different acoustic environments. The system dynamically adjusts its behavior based on the calculated coherence value, being more aggressive in echo suppression when coherence indicates echo presence, and more conservative when coherence indicates single-talk, thereby maintaining sound quality while suppressing echo.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Provided is a voice signal processing method. The method comprises the following steps: acquiring a status coefficient, a target voice signal to be processed, and a derivative signal of the target voice signal (S101); using the derivative signal to respectively calculate an initial echo loss corresponding to each signal frame in the target voice signal, then using the status coefficient to adjust the initial echo loss to acquire a target echo loss (S102); using the derivative signal to calculate a proximal-distal coherence value of the target voice signal (S 103); judging whether the target echo loss matches with a single-talk echo state (S 104); if so, recording the proximal-distal coherence value in the status statistical data , and updating statistical information (S105); and using the statistical information to update the status coefficient, and performing volume control on the target voice signal after the status coefficient is updated (S 106). The method ensures better sound quality, and further improves user experience. Also provided are a voice signal processing device, an apparatus, and a readable storage medium, having corresponding technical effects.