Voice Activity Detection Using MVDR and Delay-and-Subtract Signal Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice activity detection (VAD) systems face reduced performance when the audio system is near acoustically reflective environments, such as nearby walls or the user's hands, due to acoustic reflections that contaminate the reference signal, leading to unreliable detection of user speech.
Innovation Solution
The method involves combining microphone signals using a minimum-variance distortionless response (MVDR) and delay-and-subtract combinations to produce primary and reference signals, which are then compared to determine voice activity, with additional processing in specific frequency bands to account for boundary interference, including adding and subtracting these signals to enhance detection accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional VAD systems use a single reference signal for voice activity detection, then the system structure remains simple, but the detection reliability deteriorates in acoustically reflective environments due to signal contamination
Solution Approach 1:
The patent segments the reference signal into multiple components: a first reference signal from a first microphone and a second reference signal from a second microphone. This segmentation allows the system to process and combine multiple reference signals to improve detection reliability in reflective environments without creating a single complex processing path.
Solution Approach 2:
The patent combines multiple reference signals (first reference signal and second reference signal) to create an enhanced reference signal. This merging of multiple signal sources improves the reliability of voice activity detection by providing a more robust reference that is less susceptible to acoustic reflections from any single direction.
2Measurement precision
If the system processes signals without accounting for acoustic reflections, then the processing speed remains fast, but the measurement precision of voice activity detection deteriorates
Solution Approach 1:
The patent applies preliminary processing to the reference signals by combining the first and second reference signals before using them for voice activity detection. This preliminary combination action prepares an enhanced reference signal that accounts for acoustic reflections, improving measurement precision without requiring complex real-time processing during detection.
Solution Approach 2:
The patent introduces an intermediary enhanced reference signal that mediates between the raw microphone signals and the final detection process. This intermediary signal has been processed to account for acoustic reflections, serving as a refined reference that improves detection accuracy while keeping the overall system manageable.
3Reliability
If the system uses multiple microphone signals with different combinations, then the voice separation performance improves, but the device complexity increases
Solution Approach 1:
The patent segments the microphone signals into different groups: a first set of microphone signals used to generate a primary signal, and a second set used to generate reference signals. This segmentation allows the system to apply different processing combinations to different signal groups, improving voice separation through diverse processing approaches while managing complexity through structured organization.
Solution Approach 2:
The patent creates a multi-functional signal processing system where the same set of microphone signals serves multiple purposes: generating primary signals for voice capture, generating reference signals for noise cancellation, and enabling voice activity detection. This multi-functionality improves voice separation performance by utilizing the signals in various combinations without requiring entirely separate processing paths for each function.
Data Source
AI summary
Audio systems, methods, and processor instructions are provided that detect voice activity of a user and provide an output voice signal. The systems, methods, and instructions receive a plurality of microphone signals and combine the plurality of microphone signals according to a first combination and a second combination. The first combination produces a primary signal having enhanced response in the direction of the user's mouth, and the second combination produces a reference signal having reduced response in the direction of the user's mouth. The primary signal and the reference signal are added and subtracted to produce a voice-enhanced signal and a voice-reduced signal, respectively. The voice-enhanced signal and the voice-reduced signal are compares and an output voice signal is provided based upon the comparison.


