Adaptive Spatial VAD for Non-Stationary Noise Discrimination
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional Voice Activity Detection (VAD) systems struggle to differentiate between target speech and interfering speech in noisy environments, leading to false positives and decreased performance, especially in scenarios with non-stationary noise sources like TVs or radios, and require significant processing resources, making them impractical for low-power devices.
Innovation Solution
The implementation of an adaptive spatial VAD system that uses a sub-band analysis module, a constrained minimum variance adaptive filter, and a mask estimator to discriminate between target and interfering speech by minimizing output variance and estimating a relative transfer function of noise sources, allowing for efficient noise reduction and improved speech recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional VAD techniques are used, then speech detection can be performed, but processing resources and memory requirements become excessively large for low-power devices
Solution Approach 1:
The audio signal is divided into multiple frequency sub-bands, and VAD is performed independently in each sub-band. This segmentation allows the system to focus computational resources only on sub-bands containing speech, significantly reducing overall processing requirements while maintaining detection accuracy.
Solution Approach 2:
Different processing strategies are applied to different frequency sub-bands based on their local characteristics. Sub-bands with speech activity receive detailed analysis, while sub-bands without speech use simpler detection methods, optimizing the balance between accuracy and computational cost.
2Reliability
If conventional VAD techniques are used, then basic speech detection is possible, but the system cannot differentiate between target speech and interfering speech in noisy environments
Solution Approach 1:
The system transitions from single-channel time-domain analysis to multi-channel frequency-domain analysis. By examining speech activity across multiple frequency sub-bands and comparing patterns across channels, the system can distinguish target speech from interfering speech and noise more effectively.
Solution Approach 2:
A spatial filtering mechanism acts as an intermediary between the raw audio signals and the VAD decision. The filter processes the multi-channel input to enhance target speech while suppressing interfering speech and noise, providing a cleaner signal for accurate detection.
3Productivity
If conventional VAD techniques are used, then processing can be performed, but false positives increase in scenarios with non-stationary noise sources like TVs or radios
Solution Approach 1:
The system dynamically adapts to changing noise conditions by continuously monitoring sub-band energy patterns and adjusting detection thresholds. This dynamic adaptation allows the VAD to maintain high accuracy even when non-stationary noise sources like TVs or radios are present, reducing false positives while preserving detection speed.
Data Source
AI summary
Systems and methods include a first voice activity detector operable to detect speech in a frame of a multichannel audio input signal and output a speech determination, a constrained minimum variance adaptive filter operable to receive the multichannel audio input signal and the speech determination and minimize a signal variance at the output of the filter, thereby producing an equalized target speech signal, a mask estimator operable to receive the equalized target speech signal and the speech determination and generate a spectral-temporal mask to discriminate a target speech from noise and interference speech, and a second activity voice detector operable to detect voice in a frame of the speech discriminated signal. An audio input sensor array including a plurality of microphones, each microphone generating a channel of the multichannel audio input signal. A sub-band analysis module operable to decompose each of the channels into a plurality of frequency sub-bands.


