Voice Activity Detection Using Dual Microphone and Vibration Sensors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing noise suppression systems face challenges in accurately identifying voiced and unvoiced speech in noisy environments, particularly with portable communication devices, due to noise pollution and uncertainties in signal content, leading to poor performance in speech recognition and speaker verification.
Innovation Solution
A voice activity detection system combining acoustic and vibration sensors to improve noise suppression by using a dual omnidirectional microphone array and adaptive filtering, which forms virtual microphones with distinct noise and speech responses, effectively reducing false positives and enhancing signal-to-noise ratio.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If single microphone data is used for voice activity detection, then device complexity is reduced, but measurement precision deteriorates due to noise and signal content uncertainties
Solution Approach 1:
The system segments the acoustic signal processing by using multiple omnidirectional microphones to capture different spatial information. Each microphone provides independent signal data that is processed separately through adaptive filtering algorithms, allowing the system to distinguish speech from noise by comparing multiple segmented signal sources rather than relying on a single microphone data stream.
Solution Approach 2:
The patent transitions from single-microphone spatial detection to multi-microphone spatial detection, adding dimensional information through the microphone array configuration. By arranging omnidirectional microphones in specific geometric patterns and processing their combined signals with adaptive filters, the system creates virtual microphone positions that provide enhanced spatial discrimination capability for voice activity detection.
2Object-affected harmful factors
If noise suppression algorithms are applied to speech signals, then noise reduction is improved, but speech distortion increases when voice activity detection is inaccurate
Solution Approach 1:
The system implements feedback mechanisms where the output of adaptive filtering processes is continuously monitored and fed back into the detection algorithm. The voice activity detection system uses the processed signal information to refine its detection decisions, adjusting thresholds and parameters based on the actual speech-noise separation performance, thereby reducing speech distortion while maintaining noise suppression effectiveness.
Solution Approach 2:
The patent dynamically changes detection parameters such as energy thresholds, time windows, and frequency band selections based on the detected noise level and speech activity patterns. By adapting these parameters in real-time according to environmental conditions, the system optimizes the balance between noise suppression and speech preservation, preventing over-suppression that would cause speech distortion.
3Reliability
If robust voice activity detection methods are implemented, then noise suppression performance is improved, but device complexity increases
Solution Approach 1:
The adaptive filtering system performs self-adjustment by automatically adapting its filter coefficients based on the statistical properties of the input signals. The algorithm autonomously learns the characteristics of speech and noise components from the multi-microphone inputs and adjusts its parameters without external intervention, providing robust voice activity detection while minimizing the need for complex manual configuration and system control mechanisms.
Data Source
AI summary
A voice activity detector (VAD) combines the use of an acoustic VAD and a vibration sensor VAD as appropriate to the conditions a host device is operated. The VAD includes a first detector receiving a first signal and a second detector receiving a second signal. The VAD includes a first VAD component coupled to the first and second detectors. The first VAD component determines that the first signal corresponds to voiced speech when energy resulting from at least one operation on the first signal exceeds a first threshold. The VAD includes a second VAD component coupled to the second detector. The second VAD component determines that the second signal corresponds to voiced speech when a ratio of a second parameter corresponding to the second signal and a first parameter corresponding to the first signal exceeds a second threshold.


