Voice Activity Detection Using Adaptive SNR Band Weighting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice activity detectors (VADs) face challenges in accurately distinguishing between background noise and active speech, especially in dynamic noise environments with low signal-to-noise ratios, leading to false speech detection and missed speech segments, which affects processing and transmission quality.

Innovation Solution

The solution involves adjusting SNR values in bands through outlier filtering and adaptive weighting, where SNR outlier filtering sorts instantaneous SNR values and sets weights to zero for outlier bands, and adaptive weighting applies weights based on noise level and type to improve the accuracy of voice activity detection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If simple energy measures are used for voice activity detection, then processing is efficient and bit rate is minimized, but detection accuracy degrades significantly at low signal-to-noise ratios

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidvoice activity detection accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent divides the frequency spectrum into multiple bands and performs voice activity detection independently in each band. By segmenting the detection process across different frequency bands, the system can identify speech components that may be masked by noise in other bands, thereby improving detection accuracy at low SNR while maintaining processing efficiency through parallel band-wise operations.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from single-band energy measurement to multi-band spectral analysis. By adding the frequency band dimension to the detection process, the system can exploit spectral differences between speech and noise signals, enabling more accurate voice activity detection in noisy environments without significantly increasing computational complexity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If a conservative voice activity detector is used, then false speech detection increases, but if an aggressive VAD is used, then speech segments are missed

Engineering Contradiction:
Improvefalse speech detection rateVSAvoidmissed speech segments
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent applies different detection thresholds and criteria to different frequency bands based on their local characteristics. By adapting the detection sensitivity to each band's specific noise and speech content, the system avoids the binary choice between conservative and aggressive single-threshold VAD, reducing both false detections and missed speech segments through localized decision-making.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent incorporates feedback mechanisms where detection results from one band inform the detection process in other bands. The system uses inter-band correlation and consistency checks to validate speech detections, allowing the detector to be more confident in true speech segments while rejecting false detections, thereby balancing reliability and information preservation.

Inventive Principle:
Principle #23Feedback

3Stability of the object's composition

If long-term SNR is used to estimate VAD threshold, then the threshold is smoothed, but accuracy decreases under fast-varying non-stationary noise

Engineering Contradiction:
Improvethreshold smoothingVSAvoidVAD threshold accuracy
Core Design Contradiction:
Stability of the object's compositionVSMeasurement precision

Solution Approach 1:

The patent implements dynamic threshold adaptation where the VAD threshold is adjusted in real-time based on current noise conditions in each frequency band. Rather than using a fixed long-term smoothed threshold, the system continuously updates thresholds to track non-stationary noise variations, maintaining both stability through adaptive smoothing and accuracy through responsiveness to current conditions.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the parameters used for threshold estimation by incorporating multiple time constants and adapting the smoothing factor based on noise stationarity detection. This allows the system to switch between more and less smoothed thresholds dynamically, preserving threshold stability during stationary periods while enabling rapid adaptation during non-stationary noise conditions.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9099098B2Voice activity detection in presence of background noise
Publication Date: 2015.08.04 QUALCOMM INC
  • US9099098B2 patent drawing
  • US9099098B2 patent drawing
  • US9099098B2 patent drawing

AI summary

In speech processing systems, compensation is made for sudden changes in the background noise in the average signal-to-noise ratio (SNR) calculation. SNR outlier filtering may be used, alone or in conjunction with weighting the average SNR. Adaptive weights may be applied on the SNRs per band before computing the average SNR. The weighting function can be a function of noise level, noise type, and/or instantaneous SNR value. Another weighting mechanism applies a null filtering or outlier filtering which sets the weight in a particular band to be zero. This particular band may be characterized as the one that exhibits an SNR that is several times higher than the SNRs in other bands.