Adaptive Voice Activity Detection for Speech Recognition in Noise
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing noise cancellation techniques, such as Spectral Subtraction and Voice Activity Detection, are ineffective in environments with high levels of stationary and non-stationary noise, leading to increased error rates in speech recognition systems, especially when desired and undesired audio are simultaneously present.
Innovation Solution
The system employs dual microphone channels with an adaptive noise cancellation unit and a voice activity detector that adjusts its threshold based on estimated background noise levels, using adaptive finite impulse response filters and Wiener filters to enhance signal-to-noise ratios and improve noise cancellation performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If Spectral Subtraction is used to reduce noise in speech recognition algorithms, then noise reduction is achieved in low noise environments, but acceptable error rates cannot be maintained when the magnitude of undesired audio becomes large
Solution Approach 1:
The patent applies dynamics by making the noise cancellation system adaptive to varying noise conditions. The voice activity detector dynamically adjusts its operation based on detected voice activity in reference channels, and the noise cancellation unit adapts its filtering characteristics based on estimated background noise levels and detected speech presence, allowing the system to maintain reliability across different noise magnitudes
Solution Approach 2:
The patent changes parameters by adjusting the voice activity detection threshold based on estimated background noise levels. When noise levels are high, the threshold is adjusted to prevent false detection, and when noise levels are low, the threshold is adjusted to improve detection sensitivity. This parameter adaptation allows the system to maintain acceptable error rates across varying noise conditions
2Measurement precision
If traditional Voice Activity Detection is used to detect desired speech and treat undesired audio as noise, then voice activity detection works well for single sound source or stationary noise with small magnitude, but performance deteriorates in noisy environments with large magnitude noise
Solution Approach 1:
The patent segments the audio processing by using multiple reference channels to detect voice activity separately, then combining these detections. The noise cancellation system is divided into multiple units that process different channels independently before integrating results, allowing more accurate voice activity detection in noisy environments by leveraging information from multiple sources
Solution Approach 2:
The patent makes the voice activity detector universal by enabling it to operate effectively across different noise conditions and environments. The detector is designed to function whether the dominant noise is stationary or non-stationary, whether noise magnitude is small or large, and whether single or multiple sound sources are present, achieving multi-functional adaptability
3Ease of operation
If dual microphone VAD systems compare energy level ratio between main and reference microphones with preset threshold, then voice activity detection is performed, but when background noise level changes the preset threshold fails to detect desired voice activity or accepts undesired audio as desired voice activity
Solution Approach 1:
The patent applies dynamics by replacing the static preset threshold with a dynamic threshold that adapts to changing background noise conditions. The voice activity detector automatically adjusts its threshold based on estimated background noise levels, enabling continuous accurate operation as environmental conditions change without manual intervention
Solution Approach 2:
The patent implements feedback by using the output of the voice activity detector and noise estimation to continuously adjust the detection threshold. The system monitors background noise levels and feeds this information back to the threshold adjustment mechanism, creating a closed-loop system that maintains detection accuracy under varying conditions
4Object-affected harmful factors
If noise cancellation approaches are employed to reduce noise from stationary and non-stationary sources, then noise reduction is achieved in relatively low noise environments, but the system cannot maintain effectiveness when magnitude of noise is large relative to desired audio
Solution Approach 1:
The patent applies dynamics by making the noise cancellation unit adaptive to varying noise conditions. The system dynamically adjusts its noise cancellation characteristics based on detected voice activity and estimated background noise levels, allowing it to maintain effectiveness whether noise magnitude is small or large relative to desired audio
Solution Approach 2:
The patent changes parameters by adjusting the noise cancellation filter characteristics based on the ratio of noise magnitude to desired audio magnitude. When noise is small relative to speech, stronger cancellation is applied, while when noise is large relative to speech, the cancellation is adjusted to preserve speech quality, maintaining reliability across different signal-to-noise ratios
Data Source
AI summary
Systems, apparatuses, and methods are described to increase a signal-to-noise ratio difference between a main channel and reference channel. The increased signal-to-noise ratio difference is accomplished with an adaptive threshold for a desired voice activity detector (DVAD) and shaping filters. The DVAD includes averaging an output signal of a reference microphone channel to provide an estimated average background noise level. A threshold value is selected from a plurality of threshold values based on the estimated average background noise level. The threshold value is used to detect desired voice activity on a main microphone channel.


