Voice Activity Detection for Audio Noise Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional noise reduction techniques in audio processing systems are computationally demanding and exhibit latency, struggling to accurately distinguish background noise from speech during periods of low volume or silence, often leading to incomplete noise removal and potential clipping of speech.

Innovation Solution

The implementation of a voice activity detector that analyzes audio data in the frequency domain, using a feature extractor and machine learning models to generate noise and speech masks, which are then processed to isolate and remove noise from audio data without affecting speech, utilizing a time-frequency converter for efficient noise reduction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Object-affected harmful factors

If conventional noise reduction techniques analyze waveforms to identify and remove background noise, then noise removal capability is improved, but computational demand increases and latency occurs

Engineering Contradiction:
Improvenoise removal capabilityVSAvoidcomputational demand
Core Design Contradiction:
Object-affected harmful factorsVSPower

Solution Approach 1:

The patent transforms the audio signal from time-domain waveform analysis to frequency-domain spectrogram analysis. This parameter transformation enables more efficient noise identification by representing audio as time-frequency matrices, reducing computational complexity while maintaining noise removal effectiveness

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces conventional mechanical waveform analysis with machine learning-based spectrogram analysis. Neural networks process the frequency-domain representation to identify and remove noise, achieving lower computational demand and latency compared to traditional signal processing methods

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Object-affected harmful factors

If conventional techniques attenuate background noise during low volume or silence periods, then noise removal is improved, but speech may be inadvertently eliminated causing clipping

Engineering Contradiction:
Improvenoise removal accuracyVSAvoidspeech preservation accuracy
Core Design Contradiction:
Object-affected harmful factorsVSReliability

Solution Approach 1:

The patent applies different processing strategies to different regions of the audio signal based on voice activity detection. During silence periods, noise is removed while preserving potential speech; during speech periods, speech is protected from removal. This localized quality adjustment prevents speech clipping while maintaining noise removal effectiveness

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent implements voice activity detection that continuously monitors the audio signal and provides feedback to control noise removal intensity. When speech is detected, the system reduces noise attenuation to prevent speech clipping; during silence, it increases noise removal. This feedback mechanism resolves the contradiction between noise removal and speech preservation

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20240304203A1Noise reduction using voice activity detection in audio processing systems and applications
Publication Date: 2024.09.12 NVIDIA CORP
  • US20240304203A1 patent drawing
  • US20240304203A1 patent drawing
  • US20240304203A1 patent drawing

AI summary

In various examples, a noise reduction may be performed based at least on determining that audio data encoding sound includes undesirable sound or lacks desirable sound. A frequency is determined for audio data based at least on value(s) associated with frequency(ies) within a frequency band and used to determine that sound encoded in the audio data includes undesirable sound or lacks desirable sound.