Adaptive Voice Activity Detection for Audio Encoding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice activity detection (VAD) algorithms in audio encoding are conservative, leading to inefficient use of radio resources and audible artifacts due to high voice activity factor, especially in challenging background noise conditions, and result in increased clipping when switching between speech and comfort noise.

Innovation Solution

A method that dynamically adjusts categorization parameters based on the encoding mode to reduce the voice activity factor, using energy, signal-to-noise, pitch, and tone information to categorize segments as active or non-active, allowing for adaptive encoding that decreases the number of active segments, particularly in low quality coding modes, and considers network traffic to optimize bitrates.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conservative VAD algorithms are used to ensure reliable speech detection, then speech quality is maintained, but voice activity factor increases leading to inefficient radio resource usage

Engineering Contradiction:
Improvespeech detection reliabilityVSAvoidradio resource efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent applies dynamics by making the VAD algorithm adaptive rather than static. The algorithm dynamically adjusts its detection sensitivity and parameters based on real-time analysis of background noise characteristics, signal energy levels, and speech probability estimates. This allows the system to optimize between reliable detection and resource efficiency under varying operational conditions.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes multiple parameters including energy thresholds, signal-to-noise ratio thresholds, and detection sensitivity levels based on the operating conditions. By adjusting these parameters dynamically, the system can reduce voice activity factor when background noise is high or speech is clear, while maintaining reliable detection when conditions are challenging.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If VAD algorithms characterize fewer segments as active to improve resource efficiency, then radio resource usage improves, but clipping increases causing audible artifacts

Engineering Contradiction:
Improveradio resource efficiencyVSAvoidclipping artifacts
Core Design Contradiction:
ProductivityVSObject-generated harmful factors

Solution Approach 1:

The patent applies preliminary action by performing advance analysis of signal characteristics before making switching decisions. The system calculates speech probability, analyzes energy trends, and evaluates background noise characteristics in advance to predict when speech is likely to occur. This allows the system to maintain active encoding slightly beyond what conservative algorithms would use, preventing clipping artifacts while still improving resource efficiency.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent provides beforehand cushioning by using comfort noise parameters and smooth transition techniques to mitigate the effects of potential clipping. The system prepares comfort noise representations in advance and uses them to mask or smooth transitions, reducing the audible impact of any clipping that may occur when switching between active and inactive states.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

3Productivity

If VAD algorithms reduce voice activity factor in noisy conditions, then resource efficiency improves, but switching between speech and comfort noise becomes more audible

Engineering Contradiction:
Improveresource efficiencyVSAvoidswitching artifacts
Core Design Contradiction:
ProductivityVSObject-generated harmful factors

Solution Approach 1:

The patent applies periodic action by using smooth, gradual transitions between active speech encoding and comfort noise encoding rather than abrupt switches. The system uses overlapping analysis windows and gradual parameter changes to create smooth transitions that are less audible. This periodic, smoothed approach reduces the harshness of switching artifacts while maintaining resource efficiency benefits.

Inventive Principle:
Principle #19Periodic action

4Reliability

If encoding is performed in high quality modes, then voice quality is maintained, but the impact of clipping and switching artifacts is more noticeable

Engineering Contradiction:
Improvevoice qualityVSAvoidartifact visibility
Core Design Contradiction:
ReliabilityVSObject-generated harmful factors

Solution Approach 1:

The patent applies parameter changes by dynamically adjusting the encoding quality mode based on the detected speech activity and background noise conditions. When the system detects clear speech or low noise conditions where clipping would be more noticeable, it may temporarily increase encoding quality or adjust VAD sensitivity. When background noise is high and would mask artifacts, the system can use lower encoding modes or more aggressive VAD, reducing the visibility of artifacts while maintaining acceptable quality.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS8645133B2Adaptation of voice activity detection parameters based on encoding modes
Publication Date: 2014.02.04 NOKIA TECHNOLOGIES OY
  • US8645133B2 patent drawing
  • US8645133B2 patent drawing
  • US8645133B2 patent drawing

AI summary

Encoding audio signals with selecting an encoding mode for encoding the signal categorizing the signal into active segments having voice activity and non-active segments having substantially no voice activity by using categorization parameters depending on the selected encoding mode and encoding at least the active segments using the selected encoding mode.