Adaptive Voice Activity Detection for Audio Encoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice activity detection (VAD) algorithms in audio encoding are conservative, leading to inefficient use of radio resources and audible artifacts due to high voice activity factor, especially in challenging background noise conditions, and result in increased clipping when switching between speech and comfort noise.
Innovation Solution
A method that dynamically adjusts categorization parameters based on the encoding mode to reduce the voice activity factor, using energy, signal-to-noise, pitch, and tone information to categorize segments as active or non-active, allowing for adaptive encoding that decreases the number of active segments, particularly in low quality coding modes, and considers network traffic to optimize bitrates.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conservative VAD algorithms are used to ensure reliable speech detection, then speech quality is maintained, but voice activity factor increases leading to inefficient radio resource usage
Solution Approach 1:
The patent applies dynamics by making the VAD algorithm adaptive rather than static. The algorithm dynamically adjusts its detection sensitivity and parameters based on real-time analysis of background noise characteristics, signal energy levels, and speech probability estimates. This allows the system to optimize between reliable detection and resource efficiency under varying operational conditions.
Solution Approach 2:
The patent changes multiple parameters including energy thresholds, signal-to-noise ratio thresholds, and detection sensitivity levels based on the operating conditions. By adjusting these parameters dynamically, the system can reduce voice activity factor when background noise is high or speech is clear, while maintaining reliable detection when conditions are challenging.
2Productivity
If VAD algorithms characterize fewer segments as active to improve resource efficiency, then radio resource usage improves, but clipping increases causing audible artifacts
Solution Approach 1:
The patent applies preliminary action by performing advance analysis of signal characteristics before making switching decisions. The system calculates speech probability, analyzes energy trends, and evaluates background noise characteristics in advance to predict when speech is likely to occur. This allows the system to maintain active encoding slightly beyond what conservative algorithms would use, preventing clipping artifacts while still improving resource efficiency.
Solution Approach 2:
The patent provides beforehand cushioning by using comfort noise parameters and smooth transition techniques to mitigate the effects of potential clipping. The system prepares comfort noise representations in advance and uses them to mask or smooth transitions, reducing the audible impact of any clipping that may occur when switching between active and inactive states.
3Productivity
If VAD algorithms reduce voice activity factor in noisy conditions, then resource efficiency improves, but switching between speech and comfort noise becomes more audible
Solution Approach 1:
The patent applies periodic action by using smooth, gradual transitions between active speech encoding and comfort noise encoding rather than abrupt switches. The system uses overlapping analysis windows and gradual parameter changes to create smooth transitions that are less audible. This periodic, smoothed approach reduces the harshness of switching artifacts while maintaining resource efficiency benefits.
4Reliability
If encoding is performed in high quality modes, then voice quality is maintained, but the impact of clipping and switching artifacts is more noticeable
Solution Approach 1:
The patent applies parameter changes by dynamically adjusting the encoding quality mode based on the detected speech activity and background noise conditions. When the system detects clear speech or low noise conditions where clipping would be more noticeable, it may temporarily increase encoding quality or adjust VAD sensitivity. When background noise is high and would mask artifacts, the system can use lower encoding modes or more aggressive VAD, reducing the visibility of artifacts while maintaining acceptable quality.
Data Source
AI summary
Encoding audio signals with selecting an encoding mode for encoding the signal categorizing the signal into active segments having voice activity and non-active segments having substantially no voice activity by using categorization parameters depending on the selected encoding mode and encoding at least the active segments using the selected encoding mode.


