Voice Activity Detection Unit Using Spectro-Spatial Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice activity detection systems in portable electronic devices and hearing aids face challenges in accurately distinguishing speech from noise, especially in noisy environments, and struggle to identify the direction of speech sources amidst diffuse background noise.
Innovation Solution
A voice activity detection unit that analyzes time-frequency representations of input signals from multiple microphones, using spectro-spatial characteristics to differentiate between target speech and noise, and estimates the direction of speech sources by combining spectro-temporal and spectro-spatial features.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional voice activity detection methods are used, then the system can detect speech in simple environments, but it fails to accurately distinguish speech from noise in noisy environments
Solution Approach 1:
The patent transitions from traditional single-dimension voice activity detection to a multi-dimensional approach by incorporating spectro-spatial characteristics. The system analyzes speech signals in both time-frequency domain and spatial domain simultaneously, using microphone arrays to capture directional information. This dimensional expansion enables the system to distinguish speech from noise by examining spectral patterns and spatial distribution together, achieving robust speech detection in noisy environments.
Solution Approach 2:
The patent employs parameter changes by transforming the voice activity detection problem into different parameter spaces. Instead of relying on single threshold-based detection, the system transforms signals into time-frequency representations and spatial spectral representations, then combines these transformed parameters to make detection decisions. This parameter transformation approach allows the system to adapt to varying noise conditions and improve detection accuracy.
2Measurement precision
If multiple microphones are used to improve speech detection in noise, then speech source direction can be estimated, but the device complexity increases
Solution Approach 1:
The patent applies segmentation by dividing the complex speech detection task into distinct processing stages: time-frequency transformation, spatial spectral analysis, and combined decision-making. The system segments the audio signal processing into frequency bins and time frames, then independently analyzes spatial characteristics for each segment. This segmentation reduces computational complexity compared to processing the entire signal as a whole, while still achieving accurate direction estimation.
Solution Approach 2:
The patent implements multi-functionality by designing a voice activity detection system that simultaneously performs multiple functions: speech presence detection, speech source direction estimation, and noise suppression. The same spectro-spatial analysis framework used for voice activity detection also provides directional information, eliminating the need for separate processing chains and reducing overall system complexity despite using multiple microphones.
3Reliability
If spectro-spatial characteristics are analyzed to improve speech detection, then speech intelligibility is enhanced, but the computational requirements increase
Solution Approach 1:
The patent applies partial action by selectively processing only the most relevant frequency bins and time frames for speech detection. Instead of uniformly processing all frequency components and time segments with equal computational effort, the system identifies and focuses computational resources on regions of the time-frequency spectrum where speech is most likely to occur. This selective processing maintains speech intelligibility while reducing overall computational energy consumption.
Data Source
Figure 1A~1B
Figure 2A~2B
Figure 3A~3B
AI summary
A voice activity detection unit is configured to receive at least two electric input signals in a number of frequency bands and a number of time instances, k and m being frequency band and time indices, respectively, (k, m) defining a specific time-frequency tile of said electric input signal. The voice activity detection unit is configured to provide a resulting voice activity detection estimate comprising one or more parameters indicative of whether or not a given time-frequency tile contains or to what extent it comprises a target speech signal. The voice activity detection unit comprises a) a first detector for analyzing the time-frequency representation of the electric input signals and identifying spectro-spatial characteristics of said electric input signals, and b) and is configured for providing said resulting voice activity detection estimate in dependence of said spectro-spatial characteristics. The invention may be used in hearing aids, table microphones, speakerphones, etc.