Speech Features Voice Activity Detection for Transient Noise Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional Voice Activity Detection (VAD) systems in motor vehicle environments face challenges with transient noises and low Signal-to-Noise ratio, leading to false indications and suboptimal noise reduction, which affects Automatic Speech Recognition (ASR) systems.
Innovation Solution
A Speech Features-based Voice Activity Detection (SFVAD) system that uses time-frequency masks to estimate speech and noise indications, employing Cepstral-based pitch detection and Centrum calculation to prevent contamination, and adapts to spectral changes, reducing false detection rates and improving noise reduction capabilities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional energy-based VAD algorithms are used, then the system is simple to implement, but performance is degraded in transient noise and low SNR scenarios
Solution Approach 1:
The patent transforms the VAD approach from energy-domain parameters to spectral parameters by applying Short-Time Fourier Transform (STFT) to obtain magnitude spectra. The system then uses spectral subtraction and speech presence probability calculations in the frequency domain, fundamentally changing the parameter space from time-domain energy to frequency-domain spectral characteristics, which improves robustness against transient noise
Solution Approach 2:
The patent introduces spectral subtraction as an intermediary process between the raw audio signal and the VAD decision. By first estimating and subtracting noise spectrum from the input signal to obtain an enhanced speech estimate, the system creates an intermediate representation that removes transient noise contamination before making VAD decisions, thereby improving detection accuracy
2Reliability
If speech presence probability calculation is used, then false detection rates are reduced, but computational requirements increase
Solution Approach 1:
The patent applies spectral subtraction selectively only during identified noise segments rather than continuously processing all audio frames. The system uses a noise detector to identify when transient noise is present, then applies the computationally intensive spectral subtraction only in those specific cases, reducing overall computational burden while maintaining improved detection accuracy when needed
Solution Approach 2:
The patent performs preliminary noise estimation and spectral subtraction before the actual VAD decision-making process. By pre-enhancing the speech signal and removing transient noise components in advance, the system simplifies subsequent detection steps and reduces the computational complexity of the final VAD algorithm
3Adaptability or versatility
If the system adapts to spectral changes, then noise reduction capability is improved, but system complexity increases
Solution Approach 1:
The patent implements dynamic noise spectrum estimation that adapts to changing acoustic environments. The system continuously updates noise spectrum estimates based on recent audio frames, allowing it to track and adapt to spectral changes in transient noise characteristics. This dynamic adaptation enables the system to maintain effective noise reduction across varying noise conditions without requiring manual reconfiguration
Data Source
AI summary
The single-channel, Speech Features-Based Voice Activity Detection (SFVAD) system is a robust, low-latency system that generates per-frame speech and noise indications, along with calculating a pair of speech and noise time-frequency masks. The SFVAD system controls an adaptation mechanism for a Beam-Forming system control module and improves the speech quality and noise reduction capabilities of Automatic Speech Recognition applications, such as Virtual Assistance (VA) and Hands-Free (HF) calls, by robustly handling transient noises. The system extracts speech-like patterns from an input audio signal and it is invariant to the power-level of the input audio signal. Noise calculation is controlled by a pair of speech features-based detectors (voiced and unvoiced). A Cepstral-based pitch detector and a Centrum calculation method are used to prevent contamination of the calculated noise by speech content. The SFVAD system robustly handles instant changes of background noise level and has dramatically lower false detection rates.


