Adaptive Voice Activity Detector with Dynamic Noise State Updates

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing Voice Activity Detectors (VADs) face issues such as prematurely cutting off voice signals, misinterpreting high-level tone signals as background noise, and improper initialization or update of noise states, leading to undesirable performance and bandwidth wastage in varying environments.

Innovation Solution

A method and system for voice activity detection that adaptively extends the active voice mode after transitioning to an inactive voice mode, uses adaptive thresholds and energy-based calculations to distinguish between voice and noise, and implements an adaptive noise state update mechanism to ensure accurate noise characterization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of energy

If a Voice Activity Detector (VAD) is used to detect silence/background noise for higher compression ratio, then bandwidth efficiency is improved, but voice signals may be prematurely cut off or misinterpreted as background noise

Engineering Contradiction:
Improvebandwidth efficiencyVSAvoidvoice signal detection accuracy
Core Design Contradiction:
Loss of energyVSReliability

Solution Approach 1:

The VAD implements dynamic adaptation by continuously updating noise state estimates and adjusting detection thresholds based on changing environmental conditions. The system transitions from static threshold detection to dynamic threshold adaptation, where the decision thresholds evolve with the noise characteristics to maintain reliable voice detection while preserving bandwidth efficiency.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes detection parameters adaptively by modifying noise state estimates and threshold values based on the statistical properties of the input signal. By monitoring signal energy, zero-crossing rates, and spectral characteristics over time, the VAD adjusts its operating parameters to distinguish between voice and background noise more accurately, preventing premature cutoff while maintaining compression benefits.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If adaptive noise state update is implemented to improve noise characterization, then detection accuracy is improved, but computational complexity increases

Engineering Contradiction:
Improvenoise state estimation accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The noise state estimation mechanism serves itself by using the input signal statistics to automatically update its own parameters without external intervention. The VAD continuously monitors the signal and self-adjusts the noise state estimates based on observed patterns, eliminating the need for manual calibration or complex external control systems while maintaining high estimation accuracy.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system applies partial adaptation by updating noise state estimates only when necessary based on signal characteristics. Rather than continuously recalculating all parameters, the VAD selectively updates noise state information when the input signal indicates a change in environmental conditions, reducing computational overhead while maintaining sufficient accuracy for reliable detection.

Inventive Principle:
Principle #16Partial or excessive action

3Loss of energy

If dual-mode speech coding is used to achieve higher compression ratio during silence, then bandwidth usage is reduced, but tone signals may be misinterpreted as inactive voice

Engineering Contradiction:
Improvebandwidth usageVSAvoidtone signal identification accuracy
Core Design Contradiction:
Loss of energyVSMeasurement precision

Solution Approach 1:

The VAD implements feedback mechanisms by continuously monitoring the coded signal and using this information to adjust subsequent detection decisions. When tone signals are detected, the system feeds back this information to modify the noise state estimates and detection thresholds, preventing misinterpretation of active tone signals as inactive background noise while maintaining the bandwidth savings of dual-mode coding.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary analysis of signal characteristics before making mode selection decisions. By examining zero-crossing rates, spectral density, and energy distribution in advance, the VAD pre-classifies signals as likely voice, tone, or background noise, allowing it to select the appropriate coding mode more accurately and prevent tone signals from being misinterpreted as inactive voice.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS7983906B2Adaptive voice mode extension for a voice activity detector
Publication Date: 2011.07.19 MACOM TECH SOLUTIONS HLDG INC
  • US7983906B2 patent drawing
  • US7983906B2 patent drawing
  • US7983906B2 patent drawing

AI summary

There is provided a voice activity detection method for indicating an active voice mode and an inactive voice mode. The method comprises receiving a first portion of an input signal; determining that the first portion of the input signal includes an active voice signal; indicating the active voice mode in response to the determining that the first portion of the input signal includes the active voice signal; receiving a second portion of the input signal immediately following the first portion of the input signal; determining that the second portion of the input signal includes an inactive voice signal; extending the indicating the active voice mode for a period of time after determining that the second portion of the input signal includes the inactive voice signal, wherein the period of time varies based on one or more conditions; and indicating the inactive voice mode after expiration of the period of time.