Voice Activity Detector Sub-Frame Energy Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Voice activity detectors (VADs) are less accurate in noisy environments, leading to false alarms and mis-detections, and require significant time to stabilize when strong noise is present, which can cause erroneous radio transmissions and computational inefficiencies.

Innovation Solution

A VAD system that employs sub-frame and frame processing blocks for rapid signal energy analysis, using energy level estimation, noise elimination, and self-adapting thresholds to differentiate between speech and noise, reducing stabilization time to 250 milliseconds or less and eliminating short interfering impulses.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If power stationarity test is used to detect noise transitions, then noise threshold tracking accuracy is improved, but stabilization time increases to 1-3 seconds

Engineering Contradiction:
Improvenoise threshold tracking accuracyVSAvoidstabilization time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent divides the signal processing into sub-frames (shorter duration) and frames (longer duration), analyzing energy levels at multiple time scales. This segmentation allows rapid detection of noise transitions in sub-frames while using frames for broader context, reducing the stabilization time from 1-3 seconds to under 250 milliseconds without sacrificing measurement precision.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary energy level estimation on sub-frames before conducting full frame analysis. By pre-processing and identifying potential noise transitions in shorter sub-frames first, the system can react more quickly to noise level changes while maintaining accurate threshold tracking through subsequent frame-level verification.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If VAD operates in noisy environments, then speech detection coverage is improved, but false alarm rate increases

Engineering Contradiction:
Improvespeech detection coverageVSAvoidfalse alarm rate
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent applies different processing qualities to different time scales: sub-frames use rapid energy level estimation for quick noise transition detection, while frames use more comprehensive analysis for accurate speech presence determination. This local differentiation of processing quality allows the system to maintain high speech detection coverage while reducing false alarms through multi-level verification.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent dynamically adjusts the noise threshold based on detected noise transitions and adapts the VAD decision criteria based on the current operating conditions. This dynamic adaptation allows the system to maintain reliable operation across varying noise environments, reducing false alarms while preserving speech detection coverage.

Inventive Principle:
Principle #15Dynamics

3Device complexity

If conventional VAD algorithms are used, then implementation simplicity is maintained, but adaptability to different speech codecs decreases

Engineering Contradiction:
Improveimplementation simplicityVSAvoidcodec compatibility
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent implements a universal energy-based detection mechanism that operates independently of specific speech codec characteristics. By focusing on fundamental energy level analysis rather than codec-specific features, the VAD becomes adaptable to multiple speech codecs while maintaining relatively simple implementation through standardized energy estimation and threshold comparison operations.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS8909522B2Voice activity detector based upon a detected change in energy levels between sub-frames and a method of operation
Publication Date: 2014.12.09 MOTOROLA SOLUTIONS INC
  • US8909522B2 patent drawing
  • US8909522B2 patent drawing
  • US8909522B2 patent drawing

AI summary

A voice activity detector (100) includes a frame divider (201) for dividing frames of an input signal into consecutive sub-frames, an energy level estimator (202) for estimating an energy level of the input signal in each of the consecutive sub-frames, a noise eliminator (203) for analyzing the estimated energy levels of sets of the sub-frames to detect and eliminate from enhancement noise sub-frames and to indicate remaining sub-frames as speech sub-frames, and an energy level enhancer (205) for enhancing the estimated energy level for each of the indicated speech sub-frames by an amount which relates to a detected change of the estimated energy level for a current speech sub-frame relative to that for neighboring speech sub-frames.