Voice Activity Detection Using Adaptive State Switching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice activity detection (VAD) systems face inefficiencies in detecting speech offsets, leading to mis-detection and decreased performance due to the hard hangover process, and while soft hangover processing improves efficiency, it still falls short in enhancing overall VAD performance.
Innovation Solution
A VAD apparatus with multiple working states, utilizing non-linearly processed sub-band segmental signal-to-noise ratio (SNR) parameters to differentiate between normal and offset states, allowing for adaptive threshold adjustments and flexible processing algorithms to improve detection accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a hard hangover process is applied to ensure correct detection of speech offsets, then the possibility of mis-detection at speech offsets is diminished, but many real inactive frames are unnecessarily forced to active thus decreasing the VAD overall performance
Solution Approach 1:
The patent implements multiple working states (normal state and offset state) that dynamically switch based on detection conditions. The system transitions from a static hard hangover approach to a dynamic state-based approach where processing parameters change according to the current working state, resolving the contradiction by applying different strategies for different operational contexts
Solution Approach 2:
The system changes processing parameters based on working state: in normal state it uses standard VAD parameters, while in offset state it uses parameters specifically optimized for offset detection. This parameter adaptation allows the system to maintain high detection accuracy at speech offsets while avoiding the performance degradation caused by hard hangover forcing all frames to active
2Device complexity
If conventional VAD parameters are used, then the processing is simple, but the parameters exhibit a weak speech characteristic at the offsets of speech bursts thus increasing the possibility of mis-detecting speech offsets
Solution Approach 1:
The patent segments the VAD processing into distinct working states (normal state and offset state) with different parameter sets. By dividing the processing into segments tailored to specific operational conditions, the system achieves both simplicity through modular design and precision through state-specific parameter optimization
Solution Approach 2:
The system applies different parameter characteristics to different working states: standard parameters for normal operation and specialized parameters with strong speech characteristics for offset detection. This local quality approach ensures each state uses the most appropriate parameters for its specific detection needs
Data Source
Figure 1
Figure 2
AI summary
A voice activity detection apparatus (1) for determining a voice activity detection decision (VADD) for an input audio signal, wherein the voice activity detection apparatus (1) comprises a state detector (2) adapted to determine a current working state (WS) of at least two different working states of the voice activity detection apparatus (1) dependent on the input audio signal wherein each of the at least two different working states (WS) is associated with a corresponding working state parameter decision set (WSPDS) including at least one voice activity decision parameter (VADP) and a voice activity calculator (3) adapted to calculate a voice activity detection parameter value for the at least one VADP of the working state parameter decision set (WSPDS) associated with the current working state (WS) and to determine the voice activity detection decision (VADD) by comparing the calculated voice activity detection parameter value of the respective voice activity decision parameter (VADP) with a threshold.