Voice Activity Detection Using Dual-Stage Feature Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Voice activity detection systems face challenges in achieving low latency and high sensitivity while minimizing false alarms and lost speech, due to overlapping characteristics of speech and noise in short time frames, leading to inefficiencies in communication systems.

Innovation Solution

The system determines voice activity by combining high sensitivity short-term detection with recent high specificity feature determinations and state-related information, using a history of feature computations to make decisions on audio signal commencement or termination, and adjusts gain levels based on nuisance characteristics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If voice activity detection uses short time frame analysis to achieve low latency, then detection speed is improved, but false alarms increase due to overlapping speech and noise characteristics

Engineering Contradiction:
Improvedetection speedVSAvoidfalse alarm rate
Core Design Contradiction:
SpeedVSReliability

Solution Approach 1:

The patent segments the detection process into two distinct parts: a short-term high-sensitivity detector operating on individual frames for low-latency detection, and a long-term high-specificity feature aggregation over multiple frames for accurate classification. This segmentation allows each component to optimize for its specific function without compromise.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary high-sensitivity detection on the current frame first, then aggregates features from recent frames to confirm or correct the detection. This preliminary action ensures low latency while the subsequent aggregation reduces false alarms through contextual verification.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If voice activity detection increases sensitivity to detect all speech, then speech detection capability is improved, but false alarms increase due to noise misclassification

Engineering Contradiction:
Improvespeech detection capabilityVSAvoidfalse alarm rate
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent applies different quality requirements to different parts of the detection system: the short-term detector uses high sensitivity (loose criteria) for local frame analysis, while the long-term feature aggregation uses high specificity (strict criteria) for overall classification. Each part has optimized quality characteristics suited to its function.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The aggregation of features from multiple recent frames acts as an intermediary layer between the high-sensitivity short-term detection and the final classification decision. This intermediary process smooths out transient noise spikes while preserving genuine speech patterns, reconciling sensitivity with reliability.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If voice activity detection uses long time frame analysis to achieve high specificity, then classification accuracy is improved, but latency increases

Engineering Contradiction:
Improveclassification accuracyVSAvoiddetection latency
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The detection system is segmented into immediate short-term analysis for low-latency response and extended long-term feature aggregation for high-accuracy classification. The short-term component provides rapid initial detection while the long-term component refines accuracy without causing excessive latency.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs partial long-term analysis by aggregating features from a limited number of recent frames rather than analyzing the entire signal history. This partial action provides sufficient contextual information for accurate classification while maintaining acceptable latency constraints.

Inventive Principle:
Principle #16Partial or excessive action

4Productivity

If transmission control ceases transmission during voice inactivity to save bandwidth, then communication efficiency is improved, but speech quality deteriorates due to incorrect silence detection

Engineering Contradiction:
Improvecommunication efficiencyVSAvoidspeech quality
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The transmission control system uses feedback from the dual-stage voice activity detection (short-term sensitivity combined with long-term specificity) to make informed decisions about transmission cessation. The aggregated feature history provides feedback that confirms genuine silence versus transient noise, preventing premature transmission termination and maintaining speech quality while improving efficiency.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS9373343B2Method and system for signal transmission control
Publication Date: 2016.06.21 DOLBY LABORATORIES LICENSING CORP
  • US9373343B2 patent drawing
  • US9373343B2 patent drawing
  • US9373343B2 patent drawing

AI summary

An audio signal with a temporal sequence of blocks or frames is received or accessed. Features are determined as characterizing aggregately the sequential audio blocks/frames that have been processed recently, relative to current time. The feature determination exceeds a specificity criterion and is delayed, relative to the recently processed audio blocks/frames. Voice activity indication is detected in the audio signal. VAD is based on a decision that exceeds a preset sensitivity threshold and is computed over a brief time period, relative to blocks/frames duration, and relates to current block/frame features. The VAD and the recent feature determination are combined with state related information, which is based on a history of previous feature determinations that are compiled from multiple features, determined over a time prior to the recent feature determination time period. Decisions to commence or terminate the audio signal, or related gains, are outputted based on the combination.