Speech Features Voice Activity Detection for Transient Noise Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional Voice Activity Detection (VAD) systems in motor vehicle environments face challenges with transient noises and low Signal-to-Noise ratio, leading to false indications and suboptimal noise reduction, which affects Automatic Speech Recognition (ASR) systems.

Innovation Solution

A Speech Features-based Voice Activity Detection (SFVAD) system that uses time-frequency masks to estimate speech and noise indications, employing Cepstral-based pitch detection and Centrum calculation to prevent contamination, and adapts to spectral changes, reducing false detection rates and improving noise reduction capabilities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional energy-based VAD algorithms are used, then the system is simple to implement, but performance is degraded in transient noise and low SNR scenarios

Engineering Contradiction:
ImproveVAD detection accuracyVSAvoidalgorithm complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent transforms the VAD approach from energy-domain parameters to spectral parameters by applying Short-Time Fourier Transform (STFT) to obtain magnitude spectra. The system then uses spectral subtraction and speech presence probability calculations in the frequency domain, fundamentally changing the parameter space from time-domain energy to frequency-domain spectral characteristics, which improves robustness against transient noise

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces spectral subtraction as an intermediary process between the raw audio signal and the VAD decision. By first estimating and subtracting noise spectrum from the input signal to obtain an enhanced speech estimate, the system creates an intermediate representation that removes transient noise contamination before making VAD decisions, thereby improving detection accuracy

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If speech presence probability calculation is used, then false detection rates are reduced, but computational requirements increase

Engineering Contradiction:
Improvefalse detection rateVSAvoidcomputational power
Core Design Contradiction:
ReliabilityVSPower

Solution Approach 1:

The patent applies spectral subtraction selectively only during identified noise segments rather than continuously processing all audio frames. The system uses a noise detector to identify when transient noise is present, then applies the computationally intensive spectral subtraction only in those specific cases, reducing overall computational burden while maintaining improved detection accuracy when needed

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent performs preliminary noise estimation and spectral subtraction before the actual VAD decision-making process. By pre-enhancing the speech signal and removing transient noise components in advance, the system simplifies subsequent detection steps and reduces the computational complexity of the final VAD algorithm

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If the system adapts to spectral changes, then noise reduction capability is improved, but system complexity increases

Engineering Contradiction:
Improvespectral adaptationVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements dynamic noise spectrum estimation that adapts to changing acoustic environments. The system continuously updates noise spectrum estimates based on recent audio frames, allowing it to track and adapt to spectral changes in transient noise characteristics. This dynamic adaptation enables the system to maintain effective noise reduction across varying noise conditions without requiring manual reconfiguration

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20240355351A1Speech features-based single channel voice activity detection method and system for reducing noise from an audio signal
Publication Date: 2024.10.24 GM GLOBAL TECHNOLOGY OPERATIONS LLC
  • US20240355351A1 patent drawing
  • US20240355351A1 patent drawing
  • US20240355351A1 patent drawing

AI summary

The single-channel, Speech Features-Based Voice Activity Detection (SFVAD) system is a robust, low-latency system that generates per-frame speech and noise indications, along with calculating a pair of speech and noise time-frequency masks. The SFVAD system controls an adaptation mechanism for a Beam-Forming system control module and improves the speech quality and noise reduction capabilities of Automatic Speech Recognition applications, such as Virtual Assistance (VA) and Hands-Free (HF) calls, by robustly handling transient noises. The system extracts speech-like patterns from an input audio signal and it is invariant to the power-level of the input audio signal. Noise calculation is controlled by a pair of speech features-based detectors (voiced and unvoiced). A Cepstral-based pitch detector and a Centrum calculation method are used to prevent contamination of the calculated noise by speech content. The SFVAD system robustly handles instant changes of background noise level and has dramatically lower false detection rates.