Adaptive Echo Suppression Using Signal-Level Mask Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing echo suppression devices fail to detect speech when it is small and may excessively suppress the voice of the near-end speaker, leading to voice disappearance.
Innovation Solution
An echo suppression device that includes a mask storage unit, a mask selection unit, and an echo suppressor, which generate and select optimal masks based on the magnitude of reception signals to accurately detect speech and suppress echoes, even when speech is small.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a mask is generated assuming a large signal in the receiving signal path, then echo suppression is effective for large reception signals, but the near-end speaker's voice may disappear when the speech is small and reception signal is large
Solution Approach 1:
The patent generates multiple base masks with different magnitudes by changing the magnitude of the learning signal. The mask selection unit then selects or generates an optimal mask whose magnitude corresponds to the actual reception signal level, rather than using a single mask designed for large signals. This parameter change approach resolves the contradiction by adapting the mask magnitude to match the actual signal conditions, ensuring reliable echo suppression without causing near-end voice disappearance.
Solution Approach 2:
The patent dynamically adjusts the mask magnitude based on the actual reception signal level. Instead of using a static mask generated for large signals, the system generates or selects masks with varying magnitudes that correspond to different reception signal levels. This dynamic adaptation allows the echo suppressor to effectively handle both large and small speech signals without causing voice disappearance, resolving the reliability-precision contradiction.
2Device complexity
If a single base mask is used, then device complexity is reduced, but the system cannot adapt to varying reception signal levels
Solution Approach 1:
The patent generates multiple base masks with different magnitudes by systematically varying the learning signal magnitude parameter. This creates a set of masks adapted to different reception signal levels without requiring complex adaptive algorithms. The mask selection unit simply selects or generates the appropriate mask based on the current signal level, achieving adaptability while maintaining relatively simple device complexity.
3Adaptability or versatility
If multiple base masks are generated and stored, then adaptability to different signal levels is improved, but memory requirements and processing complexity increase
Solution Approach 1:
The patent generates multiple base masks by changing the magnitude parameter of a single learning signal rather than storing completely independent masks for each signal level. This parameter-based generation approach reduces the quantity of mask data needed while maintaining adaptability across different signal levels. The system stores a manageable set of base masks with varying magnitudes that collectively cover the full range of reception signal levels.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Even when a speech is small, the speech is allowed to be detected and an echo is allowed to be appropriately suppressed. Whenever a sample point of a reception signal transmitted through a receiving signal path that transmits a signal to a speaker is acquired, an optimal mask is sequentially generated or selected from base masks as one or a plurality of masks generated based on a learning signal based on a reception signal acquired within a predetermined period before a time point at which the sample point was acquired. Whenever the optimal mask is selected, whether a double-talk state is present is sequentially detected based on a result of comparing an input signal with the optimal mask. When detecting that a speech is not input to a microphone and the reception signal includes a speech, a process of suppressing an echo is sequentially performed on the input signal.