Voice Activity Detection Using Spectral Subtraction and Probability Modeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional voice activity detection methods struggle to accurately distinguish speech signals from noise in low signal-to-noise ratio (SNR) environments, leading to poor performance and errors in speech recognition and compression.

Innovation Solution

The proposed solution involves converting input signals into the frequency domain, generating a spectral subtraction signal by subtracting a noise spectrum, and applying this signal to a probability distribution model, such as the Rayleigh-Laplace distribution, to determine the presence of speech signals, even in low SNR conditions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional voice activity detection methods are used, then the system can operate with simple processing, but the detection accuracy deteriorates in low SNR environments

Engineering Contradiction:
Improvevoice activity detection accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The detection process is segmented into distinct functional modules: a domain conversion module that transforms input signals to frequency domain, a subtracted-spectrum-generation module that removes noise spectra, a modeling module that applies probability distribution models, and a speech-detection module that determines speech presence. This segmentation allows each module to specialize in a specific task, improving overall detection accuracy in low SNR conditions while maintaining organized complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediate processing steps between raw signal input and final detection: spectral subtraction acts as an intermediary to remove noise before analysis, and probability distribution modeling serves as another intermediary layer that bridges the cleaned spectrum and binary speech detection. These intermediaries enhance detection precision by progressively refining the signal representation.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If spectral subtraction and probability distribution modeling are applied, then detection accuracy improves in low SNR, but processing complexity increases

Engineering Contradiction:
Improvespeech signal distinction accuracyVSAvoidalgorithm complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary spectral subtraction before the main detection process. By pre-removing estimated noise spectra from the input signal, the subsequent probability distribution modeling operates on cleaner data, which improves accuracy. This preliminary action prepares the signal in advance, making the final detection more robust against noise without requiring overly complex real-time processing.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent transforms the signal from time domain to frequency domain, changing the representation parameters. This domain conversion allows spectral analysis and subtraction to be performed more effectively. Additionally, the system changes the detection parameter from simple energy thresholding to probability distribution-based decision-making, which provides better discrimination in low SNR conditions.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If noise spectrum subtraction is performed, then speech signal clarity improves, but distribution estimation errors may increase due to noise amplification

Engineering Contradiction:
Improvespeech noise distinction capabilityVSAvoiddistribution estimation reliability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The system employs feedback mechanisms where the detection results and signal characteristics are used to continuously refine the noise spectrum estimation. This adaptive approach allows the noise model to be updated based on actual signal conditions, preventing over-subtraction and reducing distribution estimation errors. The feedback loop ensures that spectral subtraction enhances clarity without introducing excessive distortion.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent applies smoothing and regularization techniques as cushioning measures before and during the spectral subtraction process. These preprocessing steps prevent extreme values and outliers from being amplified during noise removal. By cushioning the signal processing against potential distortions in advance, the system maintains more reliable distribution estimates even when aggressive noise subtraction is required.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

Data Source

PatentUS7711558B2Apparatus and method for detecting voice activity period
Publication Date: 2010.05.04 SAMSUNG ELECTRONICS CO LTD
  • US7711558B2 patent drawing
  • US7711558B2 patent drawing
  • US7711558B2 patent drawing

AI summary

An apparatus and method for detecting a voice activity period. The apparatus for detecting a voice activity period includes a domain conversion module that converts an input signal into a frequency domain signal in the unit of a frame obtained by dividing the input signal at predetermined intervals, a subtracted-spectrum-generation module that generates a spectral subtraction signal which is obtained by subtracting a predetermined noise spectrum from the converted frequency domain signal, a modeling module that applies the spectral subtraction signal to a predetermined probability distribution model, and a speech-detection module that determines whether a speech signal is present in a current frame through a probability distribution calculated by the modeling module.