Adaptive Spatial VAD for Non-Stationary Noise Discrimination

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional Voice Activity Detection (VAD) systems struggle to differentiate between target speech and interfering speech in noisy environments, leading to false positives and decreased performance, especially in scenarios with non-stationary noise sources like TVs or radios, and require significant processing resources, making them impractical for low-power devices.

Innovation Solution

The implementation of an adaptive spatial VAD system that uses a sub-band analysis module, a constrained minimum variance adaptive filter, and a mask estimator to discriminate between target and interfering speech by minimizing output variance and estimating a relative transfer function of noise sources, allowing for efficient noise reduction and improved speech recognition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional VAD techniques are used, then speech detection can be performed, but processing resources and memory requirements become excessively large for low-power devices

Engineering Contradiction:
Improvespeech detection accuracyVSAvoidprocessing resource consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The audio signal is divided into multiple frequency sub-bands, and VAD is performed independently in each sub-band. This segmentation allows the system to focus computational resources only on sub-bands containing speech, significantly reducing overall processing requirements while maintaining detection accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Different processing strategies are applied to different frequency sub-bands based on their local characteristics. Sub-bands with speech activity receive detailed analysis, while sub-bands without speech use simpler detection methods, optimizing the balance between accuracy and computational cost.

Inventive Principle:
Principle #3Local quality

2Reliability

If conventional VAD techniques are used, then basic speech detection is possible, but the system cannot differentiate between target speech and interfering speech in noisy environments

Engineering Contradiction:
Improvespeech detection accuracyVSAvoidinterfering speech and noise
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The system transitions from single-channel time-domain analysis to multi-channel frequency-domain analysis. By examining speech activity across multiple frequency sub-bands and comparing patterns across channels, the system can distinguish target speech from interfering speech and noise more effectively.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

A spatial filtering mechanism acts as an intermediary between the raw audio signals and the VAD decision. The filter processes the multi-channel input to enhance target speech while suppressing interfering speech and noise, providing a cleaner signal for accurate detection.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If conventional VAD techniques are used, then processing can be performed, but false positives increase in scenarios with non-stationary noise sources like TVs or radios

Engineering Contradiction:
Improvedetection speedVSAvoidfalse positive rate
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system dynamically adapts to changing noise conditions by continuously monitoring sub-band energy patterns and adjusting detection thresholds. This dynamic adaptation allows the VAD to maintain high accuracy even when non-stationary noise sources like TVs or radios are present, reducing false positives while preserving detection speed.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11257512B2Adaptive spatial VAD and time-frequency mask estimation for highly non-stationary noise sources
Publication Date: 2022.02.22 SYNAPTICS INC
  • US11257512B2 patent drawing
  • US11257512B2 patent drawing
  • US11257512B2 patent drawing

AI summary

Systems and methods include a first voice activity detector operable to detect speech in a frame of a multichannel audio input signal and output a speech determination, a constrained minimum variance adaptive filter operable to receive the multichannel audio input signal and the speech determination and minimize a signal variance at the output of the filter, thereby producing an equalized target speech signal, a mask estimator operable to receive the equalized target speech signal and the speech determination and generate a spectral-temporal mask to discriminate a target speech from noise and interference speech, and a second activity voice detector operable to detect voice in a frame of the speech discriminated signal. An audio input sensor array including a plurality of microphones, each microphone generating a channel of the multichannel audio input signal. A sub-band analysis module operable to decompose each of the channels into a plurality of frequency sub-bands.