Voice Activity Detector Noise Adaptation Dynamics
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech enhancement systems struggle to accurately estimate noise, especially in dynamic environments with sudden changes, leading to reduced speech clarity due to time-varying noise characteristics and fluctuations, which can result in misidentification of noise during speech and poor echo cancellation performance.
Innovation Solution
A method and system that improve noise estimation by classifying background noise and speech through spectral analysis and temporal variability, using multiple frequency resolutions, and modifying noise adaptation rates based on estimated noise characteristics, with adaptive logic that adjusts noise estimates quickly to sudden changes and stabilizes during voiced segments, while maintaining low computational complexity and memory requirements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If noise is monitored during pauses in speech and average noise condition is recorded, then noise estimation is simplified, but the system cannot identify sudden noise changes that occur during speech
Solution Approach 1:
The system dynamically adjusts the noise adaptation rate based on speech activity detection. When speech is detected, the adaptation rate is reduced to prevent noise tracking during voiced segments. When no speech is present, the adaptation rate increases to quickly track sudden noise changes. This dynamic adjustment resolves the contradiction by making the estimation process adaptive to current conditions.
Solution Approach 2:
The system changes the noise adaptation rate parameter based on speech activity. The adaptation rate is modified from a fixed value to a variable parameter that depends on whether speech is currently active. This parameter change enables the system to balance between quick noise tracking and stable noise estimation during different operational states.
2Measurement precision
If minimum noise threshold is adjusted to match sudden noise level changes, then noise tracking improves, but speech may be incorrectly removed during echo cancellation
Solution Approach 1:
The system uses speech activity detection as feedback to control the noise estimation process. The speech activity detector provides feedback about whether speech is currently active, which then modulates the noise adaptation rate. This feedback mechanism prevents the system from tracking noise during speech, thereby avoiding incorrect speech removal while still enabling quick noise tracking during non-speech periods.
Solution Approach 2:
The noise adaptation rate is made dynamic and speech-dependent. During speech activity, the adaptation rate decreases to stabilize noise estimation and prevent speech removal. During non-speech periods, the adaptation rate increases to quickly track noise changes. This dynamic behavior resolves the reliability issue while maintaining noise tracking accuracy.
3Speed
If noise adaptation rate is increased to quickly track sudden noise changes, then noise estimation responsiveness improves, but computational complexity increases
Solution Approach 1:
The system segments the noise estimation process into two distinct modes: a fast adaptation mode for tracking sudden noise changes and a stable estimation mode for maintaining accuracy during speech. By segmenting the operational states and applying different adaptation rates to each, the system achieves fast response when needed without continuously maintaining high computational complexity.
Data Source
Figure 1
Figure 2
Figure 3~4
AI summary
An enhancement system improves the estimate of noise from a received signal. The system includes a spectrum monitor that divides a portion of the signal at more than one frequency resolution. Adaptation logic derives a noise adaptation factor of the received signal. A plurality of devices tracks the characteristics of an estimated noise in the received signal and modifies multiple noise adaptation rates. Weighting logic applies the modified noise adaptation rates derived from the signal divided at a first frequency resolution to the signal divided at a second frequency resolution.