Acoustic Voice Activity Detection Using Virtual Microphones
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing noise suppression systems in acoustic applications face challenges in accurately distinguishing voiced and unvoiced speech from background noise, especially in noisy environments, due to reliance on single microphone data and susceptibility to noise interference.
Innovation Solution
The use of a dual omnidirectional microphone array (DOMA) with virtual microphones configured to have similar noise responses and dissimilar speech responses, employing adaptive filtering to enhance voice activity detection (VAD) accuracy, and incorporating non-acoustic sensors for additional noise suppression.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If single microphone data is used for voice activity detection, then device complexity is reduced, but measurement precision deteriorates due to noise interference
Solution Approach 1:
The patent divides the speech signal into voiced and unvoiced segments using dual microphone arrays with different directional characteristics. The first microphone array captures speech with certain directional sensitivity while the second array captures speech with complementary directional sensitivity, allowing segmentation of speech components for more accurate voice activity detection in noisy environments.
Solution Approach 2:
The patent combines signals from multiple microphone arrays with different directional responses to achieve improved voice activity detection. By merging the complementary information from both arrays, the system overcomes the limitations of single microphone data and achieves higher measurement precision without excessive complexity increase.
2Measurement precision
If dual omnidirectional microphone array with virtual microphones is used, then voice activity detection accuracy is improved, but device complexity increases
Solution Approach 1:
The patent introduces virtual microphones as an intermediary computational layer between the physical dual omnidirectional microphone arrays and the voice activity detection algorithm. These virtual microphones are formed through signal processing that combines outputs from the physical arrays, providing enhanced directional sensitivity and noise rejection capabilities without requiring additional physical microphones, thus managing device complexity while improving detection accuracy.
3Measurement precision
If adaptive filtering is employed to enhance VAD accuracy, then measurement precision is improved, but device complexity and computational requirements increase
Solution Approach 1:
The patent employs adaptive filtering that dynamically adjusts its parameters based on the statistical characteristics of the received signals. The filter adapts to changing noise conditions and speech patterns in real-time, improving voice activity detection accuracy by optimizing the weighting and combination of signals from different microphone arrays according to current environmental conditions.
4Reliability
If non-acoustic sensors are incorporated for noise suppression, then reliability is improved, but device complexity increases
Solution Approach 1:
The patent integrates non-acoustic sensors that serve multiple functions: they provide additional noise suppression capabilities, enable alternative voice activity detection pathways when acoustic sensors are insufficient, and contribute to overall system reliability. This multi-functional approach justifies the increased device complexity by providing robust performance across diverse operating conditions.
Data Source
AI summary
Acoustic Voice Activity Detection (AVAD) methods and systems are described. The AVAD methods and systems, including corresponding algorithms or programs, use microphones to generate virtual directional microphones which have very similar noise responses and very dissimilar speech responses. The ratio of the energies of the virtual microphones is then calculated over a given window size and the ratio can then be used with a variety of methods to generate a VAD signal. The virtual microphones can be constructed using either an adaptive or a fixed filter.


