Voice Activity Detection Using Spectral Subtraction and Probability Modeling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice activity detection methods struggle to accurately distinguish speech signals from noise in low signal-to-noise ratio (SNR) environments, leading to poor performance and errors in speech recognition and compression.
Innovation Solution
The proposed solution involves converting input signals into the frequency domain, generating a spectral subtraction signal by subtracting a noise spectrum, and applying this signal to a probability distribution model, such as the Rayleigh-Laplace distribution, to determine the presence of speech signals, even in low SNR conditions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional voice activity detection methods are used, then the system can operate with simple processing, but the detection accuracy deteriorates in low SNR environments
Solution Approach 1:
The detection process is segmented into distinct functional modules: a domain conversion module that transforms input signals to frequency domain, a subtracted-spectrum-generation module that removes noise spectra, a modeling module that applies probability distribution models, and a speech-detection module that determines speech presence. This segmentation allows each module to specialize in a specific task, improving overall detection accuracy in low SNR conditions while maintaining organized complexity.
Solution Approach 2:
The patent introduces intermediate processing steps between raw signal input and final detection: spectral subtraction acts as an intermediary to remove noise before analysis, and probability distribution modeling serves as another intermediary layer that bridges the cleaned spectrum and binary speech detection. These intermediaries enhance detection precision by progressively refining the signal representation.
2Measurement precision
If spectral subtraction and probability distribution modeling are applied, then detection accuracy improves in low SNR, but processing complexity increases
Solution Approach 1:
The system performs preliminary spectral subtraction before the main detection process. By pre-removing estimated noise spectra from the input signal, the subsequent probability distribution modeling operates on cleaner data, which improves accuracy. This preliminary action prepares the signal in advance, making the final detection more robust against noise without requiring overly complex real-time processing.
Solution Approach 2:
The patent transforms the signal from time domain to frequency domain, changing the representation parameters. This domain conversion allows spectral analysis and subtraction to be performed more effectively. Additionally, the system changes the detection parameter from simple energy thresholding to probability distribution-based decision-making, which provides better discrimination in low SNR conditions.
3Measurement precision
If noise spectrum subtraction is performed, then speech signal clarity improves, but distribution estimation errors may increase due to noise amplification
Solution Approach 1:
The system employs feedback mechanisms where the detection results and signal characteristics are used to continuously refine the noise spectrum estimation. This adaptive approach allows the noise model to be updated based on actual signal conditions, preventing over-subtraction and reducing distribution estimation errors. The feedback loop ensures that spectral subtraction enhances clarity without introducing excessive distortion.
Solution Approach 2:
The patent applies smoothing and regularization techniques as cushioning measures before and during the spectral subtraction process. These preprocessing steps prevent extreme values and outliers from being amplified during noise removal. By cushioning the signal processing against potential distortions in advance, the system maintains more reliable distribution estimates even when aggressive noise subtraction is required.
Data Source
AI summary
An apparatus and method for detecting a voice activity period. The apparatus for detecting a voice activity period includes a domain conversion module that converts an input signal into a frequency domain signal in the unit of a frame obtained by dividing the input signal at predetermined intervals, a subtracted-spectrum-generation module that generates a spectral subtraction signal which is obtained by subtracting a predetermined noise spectrum from the converted frequency domain signal, a modeling module that applies the spectral subtraction signal to a predetermined probability distribution model, and a speech-detection module that determines whether a speech signal is present in a current frame through a probability distribution calculated by the modeling module.


