Voice Activity Detection for Audio Noise Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional noise reduction techniques in audio processing systems are computationally demanding and exhibit latency, struggling to accurately distinguish background noise from speech during periods of low volume or silence, often leading to incomplete noise removal and potential clipping of speech.
Innovation Solution
The implementation of a voice activity detector that analyzes audio data in the frequency domain, using a feature extractor and machine learning models to generate noise and speech masks, which are then processed to isolate and remove noise from audio data without affecting speech, utilizing a time-frequency converter for efficient noise reduction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If conventional noise reduction techniques analyze waveforms to identify and remove background noise, then noise removal capability is improved, but computational demand increases and latency occurs
Solution Approach 1:
The patent transforms the audio signal from time-domain waveform analysis to frequency-domain spectrogram analysis. This parameter transformation enables more efficient noise identification by representing audio as time-frequency matrices, reducing computational complexity while maintaining noise removal effectiveness
Solution Approach 2:
The patent replaces conventional mechanical waveform analysis with machine learning-based spectrogram analysis. Neural networks process the frequency-domain representation to identify and remove noise, achieving lower computational demand and latency compared to traditional signal processing methods
2Object-affected harmful factors
If conventional techniques attenuate background noise during low volume or silence periods, then noise removal is improved, but speech may be inadvertently eliminated causing clipping
Solution Approach 1:
The patent applies different processing strategies to different regions of the audio signal based on voice activity detection. During silence periods, noise is removed while preserving potential speech; during speech periods, speech is protected from removal. This localized quality adjustment prevents speech clipping while maintaining noise removal effectiveness
Solution Approach 2:
The patent implements voice activity detection that continuously monitors the audio signal and provides feedback to control noise removal intensity. When speech is detected, the system reduces noise attenuation to prevent speech clipping; during silence, it increases noise removal. This feedback mechanism resolves the contradiction between noise removal and speech preservation
Data Source
AI summary
In various examples, a noise reduction may be performed based at least on determining that audio data encoding sound includes undesirable sound or lacks desirable sound. A frequency is determined for audio data based at least on value(s) associated with frequency(ies) within a frequency band and used to determine that sound encoded in the audio data includes undesirable sound or lacks desirable sound.


