Bimodal Microphone Voice Activity Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Single modal microphones for wearable devices are not reliable for distinguishing between wearer voice activity and ambient audio, leading to undesired activations, such as unlocking a device when someone else speaks.

Innovation Solution

The use of bimodal microphones, combining air and bone conduction or in-ear microphones, to capture and process audio signals, leveraging the relative transfer function between modalities to determine whether the sound source is the wearer or the environment, employing a neural network classifier trained on diverse voice and background data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If single modal microphones are used for voice activity detection, then device complexity is reduced, but reliability of distinguishing wearer voice from ambient audio deteriorates

Engineering Contradiction:
Improvemicrophone configurationVSAvoidvoice activity detection accuracy
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The audio detection function is segmented into two separate modalities: air conduction microphone for ambient sound and bone conduction microphone for wearer's voice. Each microphone captures different types of sound waves through different transmission paths, allowing the system to distinguish between wearer voice and ambient audio by analyzing the characteristics of each modality separately and then fusing the results.

Inventive Principle:
Principle #1Segmentation

2Reliability

If bimodal microphones are used to capture air and bone conduction signals, then reliability of wearer voice detection is improved, but device complexity increases

Engineering Contradiction:
Improvewearer voice detection accuracyVSAvoidmicrophone system structure
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system merges two different types of microphones (air conduction and bone conduction) into a unified voice activity detection system. By combining the signals from both modalities and processing them together through a fusion algorithm, the system achieves more reliable wearer voice detection than either modality could provide alone, while managing the increased complexity through integrated signal processing.

Inventive Principle:
Principle #5Merging (Combining)

3Measurement precision

If bimodal signal processing with neural network classifier is implemented, then measurement precision of voice origin identification is improved, but computational complexity increases

Engineering Contradiction:
Improvevoice origin detection accuracyVSAvoidsignal processing system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system replaces traditional mechanical or rule-based voice detection methods with a neural network classifier. The neural network automatically learns the complex patterns and features that distinguish wearer voice from ambient audio by training on labeled data, substituting manual feature engineering and decision rules with an adaptive machine learning model that achieves higher precision in voice origin identification.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Effectively differentiates between wearer voice and background noise, enhancing the reliability of voice activity detection and reducing false activations, with robust performance across various speakers and environments.

Implementation Method 1

another microphone where the audio signal is transmitted through a non-air medium (example, bone conduction or in-ear mics)

Methodology Applied
Scientific EffectBone conduction:

Data Source

PatentUS9978397B2Wearer voice activity detection
Publication Date: 2018.05.22 GOOGLE LLC
  • US9978397B2 patent drawing
  • US9978397B2 patent drawing
  • US9978397B2 patent drawing

AI summary

Embodiments include a wearable device, such as a head-worn device. The wearable device includes a first microphone to receive a first sound signal from a wearer of the wearable device; a second microphone to receive a second sound signal from the wearer of the wearable device; and a processor to process the first sound signals and the second sound signals to determine that the first and second sound signals originate from the wearer of the wearable device.