Voice Activity Detector Sub-Frame Energy Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Voice activity detectors (VADs) are less accurate in noisy environments, leading to false alarms and mis-detections, and require significant time to stabilize when strong noise is present, which can cause erroneous radio transmissions and computational inefficiencies.
Innovation Solution
A VAD system that employs sub-frame and frame processing blocks for rapid signal energy analysis, using energy level estimation, noise elimination, and self-adapting thresholds to differentiate between speech and noise, reducing stabilization time to 250 milliseconds or less and eliminating short interfering impulses.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If power stationarity test is used to detect noise transitions, then noise threshold tracking accuracy is improved, but stabilization time increases to 1-3 seconds
Solution Approach 1:
The patent divides the signal processing into sub-frames (shorter duration) and frames (longer duration), analyzing energy levels at multiple time scales. This segmentation allows rapid detection of noise transitions in sub-frames while using frames for broader context, reducing the stabilization time from 1-3 seconds to under 250 milliseconds without sacrificing measurement precision.
Solution Approach 2:
The patent performs preliminary energy level estimation on sub-frames before conducting full frame analysis. By pre-processing and identifying potential noise transitions in shorter sub-frames first, the system can react more quickly to noise level changes while maintaining accurate threshold tracking through subsequent frame-level verification.
2Productivity
If VAD operates in noisy environments, then speech detection coverage is improved, but false alarm rate increases
Solution Approach 1:
The patent applies different processing qualities to different time scales: sub-frames use rapid energy level estimation for quick noise transition detection, while frames use more comprehensive analysis for accurate speech presence determination. This local differentiation of processing quality allows the system to maintain high speech detection coverage while reducing false alarms through multi-level verification.
Solution Approach 2:
The patent dynamically adjusts the noise threshold based on detected noise transitions and adapts the VAD decision criteria based on the current operating conditions. This dynamic adaptation allows the system to maintain reliable operation across varying noise environments, reducing false alarms while preserving speech detection coverage.
3Device complexity
If conventional VAD algorithms are used, then implementation simplicity is maintained, but adaptability to different speech codecs decreases
Solution Approach 1:
The patent implements a universal energy-based detection mechanism that operates independently of specific speech codec characteristics. By focusing on fundamental energy level analysis rather than codec-specific features, the VAD becomes adaptable to multiple speech codecs while maintaining relatively simple implementation through standardized energy estimation and threshold comparison operations.
Data Source
AI summary
A voice activity detector (100) includes a frame divider (201) for dividing frames of an input signal into consecutive sub-frames, an energy level estimator (202) for estimating an energy level of the input signal in each of the consecutive sub-frames, a noise eliminator (203) for analyzing the estimated energy levels of sets of the sub-frames to detect and eliminate from enhancement noise sub-frames and to indicate remaining sub-frames as speech sub-frames, and an energy level enhancer (205) for enhancing the estimated energy level for each of the indicated speech sub-frames by an amount which relates to a detected change of the estimated energy level for a current speech sub-frame relative to that for neighboring speech sub-frames.


