Bimodal Microphone Voice Activity Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Single modal microphones for wearable devices are not reliable for distinguishing between wearer voice activity and ambient audio, leading to undesired activations, such as unlocking a device when someone else speaks.
Innovation Solution
The use of bimodal microphones, combining air and bone conduction or in-ear microphones, to capture and process audio signals, leveraging the relative transfer function between modalities to determine whether the sound source is the wearer or the environment, employing a neural network classifier trained on diverse voice and background data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If single modal microphones are used for voice activity detection, then device complexity is reduced, but reliability of distinguishing wearer voice from ambient audio deteriorates
Solution Approach 1:
The audio detection function is segmented into two separate modalities: air conduction microphone for ambient sound and bone conduction microphone for wearer's voice. Each microphone captures different types of sound waves through different transmission paths, allowing the system to distinguish between wearer voice and ambient audio by analyzing the characteristics of each modality separately and then fusing the results.
2Reliability
If bimodal microphones are used to capture air and bone conduction signals, then reliability of wearer voice detection is improved, but device complexity increases
Solution Approach 1:
The system merges two different types of microphones (air conduction and bone conduction) into a unified voice activity detection system. By combining the signals from both modalities and processing them together through a fusion algorithm, the system achieves more reliable wearer voice detection than either modality could provide alone, while managing the increased complexity through integrated signal processing.
3Measurement precision
If bimodal signal processing with neural network classifier is implemented, then measurement precision of voice origin identification is improved, but computational complexity increases
Solution Approach 1:
The system replaces traditional mechanical or rule-based voice detection methods with a neural network classifier. The neural network automatically learns the complex patterns and features that distinguish wearer voice from ambient audio by training on labeled data, substituting manual feature engineering and decision rules with an adaptive machine learning model that achieves higher precision in voice origin identification.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Effectively differentiates between wearer voice and background noise, enhancing the reliability of voice activity detection and reducing false activations, with robust performance across various speakers and environments.
Implementation Method 1
another microphone where the audio signal is transmitted through a non-air medium (example, bone conduction or in-ear mics)
Data Source
AI summary
Embodiments include a wearable device, such as a head-worn device. The wearable device includes a first microphone to receive a first sound signal from a wearer of the wearable device; a second microphone to receive a second sound signal from the wearer of the wearable device; and a processor to process the first sound signals and the second sound signals to determine that the first and second sound signals originate from the wearer of the wearable device.


