Noise-Added VAD Model for Accurate Voice Signal Boundary Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional voice activity detection methods fail to accurately recognize the start and end of voice signals, leading to false or missed determinations, which increases power consumption and delays in voice recognition systems.

Innovation Solution

A method and apparatus for voice activity detection that divides audio files into frames, extracts acoustic features, and uses a noise-added VAD model trained with noise and voice segments to determine the start and end of voice signals by inputting these features to a deep neural network for probability-based classification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional VAD methods based on signal processing or deep learning models are used, then the system can process audio files, but the accuracy of recognizing voice signal start and end points deteriorates

Engineering Contradiction:
Improveaccuracy of voice signal start and end recognitionVSAvoidfalse or missed determination rate
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent applies preliminary action by pre-training the VAD model with noise-added voice segments before actual voice activity detection. The model is prepared in advance with augmented training data that includes various noise conditions, enabling it to accurately distinguish voice from noise during actual operation. This pre-preparation resolves the contradiction by improving recognition accuracy without compromising reliability during real-time detection.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements dynamics by using a probabilistic output mechanism where the VAD model outputs probability values for each audio frame indicating voice or noise classification. Instead of fixed threshold-based decisions, the system dynamically adjusts classification based on probability distributions, allowing flexible and accurate determination of voice signal boundaries while maintaining high reliability across varying audio conditions.

Inventive Principle:
Principle #15Dynamics

2Loss of energy

If conventional VAD methods are used to determine voice signal boundaries, then processing can be performed, but power consumption increases due to false or missed determinations

Engineering Contradiction:
Improvepower consumptionVSAvoidaccuracy of voice signal boundary recognition
Core Design Contradiction:
Loss of energyVSMeasurement precision

Solution Approach 1:

By pre-training the VAD model with noise-added voice segments, the system achieves accurate voice boundary detection from the outset, eliminating the need for repeated processing and corrections that would consume additional power. The preliminary training ensures the model makes correct determinations immediately, reducing energy waste associated with false positives and missed detections.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces traditional signal processing methods with a deep learning-based VAD model that processes audio frames probabilistically. This substitution enables more accurate and efficient voice activity detection, reducing the computational overhead and power consumption associated with conventional methods that require multiple processing passes to achieve similar accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Loss of time

If conventional VAD methods are used, then voice recognition processing can occur, but response time increases due to waiting delays

Engineering Contradiction:
Improveresponse timeVSAvoidaccuracy of voice signal end detection
Core Design Contradiction:
Loss of timeVSReliability

Solution Approach 1:

The dynamic probabilistic output mechanism allows the system to make real-time voice activity decisions based on probability thresholds. This enables quick response to voice signal boundaries without requiring extended analysis periods, reducing response time while maintaining accurate detection of voice end points through confidence-based classification.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

Replacing conventional signal processing with the deep learning VAD model enables faster and more reliable voice boundary detection. The model processes audio frames sequentially and outputs probability-based classifications immediately, eliminating the delays associated with traditional methods that require multiple analysis passes to achieve comparable accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11127416B2Method and apparatus for voice activity detection
Publication Date: 2021.09.21 BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
  • US11127416B2 patent drawing
  • US11127416B2 patent drawing
  • US11127416B2 patent drawing

AI summary

A method and an apparatus for voice activity detection provided in embodiments of the present disclosure allow for dividing a to-be-detected audio file into frames to obtain a first sequence of audio frames, extracting an acoustic features of each audio frame in the first sequence of audio frames, and then inputting the acoustic feature of each audio frame to a noise-added VAD model in chronological order to obtain a probability value of each audio frame in the first sequence of audio frames; and then determining, by an electronic device, a start and an end of the voice signal according to the probability value of each audio frame. During the VAD detection, the start and the end of a voice signal in an audio are recognized with a noise-added VAD model to realize the purpose of accurately recognizing the start and the end of the voice signal.