Voice Activity Detection Unit Using Spectro-Spatial Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice activity detection systems in portable electronic devices and hearing aids face challenges in accurately distinguishing speech from noise, especially in noisy environments, and struggle to identify the direction of speech sources amidst diffuse background noise.

Innovation Solution

A voice activity detection unit that analyzes time-frequency representations of input signals from multiple microphones, using spectro-spatial characteristics to differentiate between target speech and noise, and estimates the direction of speech sources by combining spectro-temporal and spectro-spatial features.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional voice activity detection methods are used, then the system can detect speech in simple environments, but it fails to accurately distinguish speech from noise in noisy environments

Engineering Contradiction:
Improvespeech detection accuracyVSAvoidnoise interference
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent transitions from traditional single-dimension voice activity detection to a multi-dimensional approach by incorporating spectro-spatial characteristics. The system analyzes speech signals in both time-frequency domain and spatial domain simultaneously, using microphone arrays to capture directional information. This dimensional expansion enables the system to distinguish speech from noise by examining spectral patterns and spatial distribution together, achieving robust speech detection in noisy environments.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent employs parameter changes by transforming the voice activity detection problem into different parameter spaces. Instead of relying on single threshold-based detection, the system transforms signals into time-frequency representations and spatial spectral representations, then combines these transformed parameters to make detection decisions. This parameter transformation approach allows the system to adapt to varying noise conditions and improve detection accuracy.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If multiple microphones are used to improve speech detection in noise, then speech source direction can be estimated, but the device complexity increases

Engineering Contradiction:
Improvespeech source direction estimationVSAvoidmicrophone array processing
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the complex speech detection task into distinct processing stages: time-frequency transformation, spatial spectral analysis, and combined decision-making. The system segments the audio signal processing into frequency bins and time frames, then independently analyzes spatial characteristics for each segment. This segmentation reduces computational complexity compared to processing the entire signal as a whole, while still achieving accurate direction estimation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements multi-functionality by designing a voice activity detection system that simultaneously performs multiple functions: speech presence detection, speech source direction estimation, and noise suppression. The same spectro-spatial analysis framework used for voice activity detection also provides directional information, eliminating the need for separate processing chains and reducing overall system complexity despite using multiple microphones.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If spectro-spatial characteristics are analyzed to improve speech detection, then speech intelligibility is enhanced, but the computational requirements increase

Engineering Contradiction:
Improvespeech intelligibilityVSAvoidcomputational energy consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent applies partial action by selectively processing only the most relevant frequency bins and time frames for speech detection. Instead of uniformly processing all frequency components and time segments with equal computational effort, the system identifies and focuses computational resources on regions of the time-frequency spectrum where speech is most likely to occur. This selective processing maintains speech intelligibility while reducing overall computational energy consumption.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentEP3300078B1A voice activitity detection unit and a hearing device comprising a voice activity detection unit
Publication Date: 2020.12.30 OTICON
  • EP3300078B1 patent drawingFigure 1A~1B
  • EP3300078B1 patent drawingFigure 2A~2B
  • EP3300078B1 patent drawingFigure 3A~3B

AI summary

A voice activity detection unit is configured to receive at least two electric input signals in a number of frequency bands and a number of time instances, k and m being frequency band and time indices, respectively, (k, m) defining a specific time-frequency tile of said electric input signal. The voice activity detection unit is configured to provide a resulting voice activity detection estimate comprising one or more parameters indicative of whether or not a given time-frequency tile contains or to what extent it comprises a target speech signal. The voice activity detection unit comprises a) a first detector for analyzing the time-frequency representation of the electric input signals and identifying spectro-spatial characteristics of said electric input signals, and b) and is configured for providing said resulting voice activity detection estimate in dependence of said spectro-spatial characteristics. The invention may be used in hearing aids, table microphones, speakerphones, etc.