Head-Mounted Speech Segmentation Using Audio-Vibration Correlation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice interfaces in wearable devices face challenges in accurately determining when to activate and deactivate, as they struggle to differentiate between wearer speech and other audio sources, leading to inefficient speech segmentation.

Innovation Solution

The integration of vibration sensors, such as accelerometers and bone-conducting microphones, with audio microphones to create a vibration channel, allowing correlation with audio data to determine causal relationships and identify wearer speech, thereby activating or deactivating the voice interface appropriately.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If voice interface activation relies solely on audio data from microphones, then the device can detect speech inputs, but it cannot accurately differentiate between wearer speech and other audio sources

Engineering Contradiction:
Improvespeech detection accuracyVSAvoidinability to identify speech source
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent segments the audio detection function into two separate sensing channels: an audio channel using microphones for airborne sound detection, and a vibration channel using accelerometers and bone-conducting microphones for structural vibration detection. This segmentation allows the system to distinguish between audio sources based on which channel detects the vibration, thereby identifying whether speech originates from the wearer.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces vibration sensors (accelerometers and bone-conducting microphones) as intermediary devices that detect mechanical vibrations transmitted through the head structure. These intermediaries provide additional information about the physical source of speech, enabling the system to differentiate wearer speech from external audio sources that would not produce corresponding vibrations in the head.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If the voice interface remains continuously active to capture all potential speech inputs, then no speech is missed, but unnecessary audio processing occurs reducing efficiency

Engineering Contradiction:
Improvespeech capture completenessVSAvoidaudio processing efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent performs preliminary filtering by continuously monitoring vibration channel data to detect vibrations consistent with speech production before fully activating the voice interface. This preliminary action allows the system to prepare for speech processing in advance, ensuring reliable capture while avoiding full activation for non-speech audio events, thus improving processing efficiency.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements a feedback mechanism where vibration channel data continuously informs the state of the voice interface activation. The system uses real-time vibration detection as feedback to dynamically adjust whether the audio channel should be actively processed, creating an efficient loop that maintains reliability while optimizing productivity by avoiding unnecessary processing.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If multiple sensor types are integrated to improve speech identification, then wearer speech detection accuracy improves, but device complexity increases

Engineering Contradiction:
Improvewearer speech identification accuracyVSAvoidsensor integration complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent makes the vibration sensors serve multiple functions: accelerometers detect both speech vibrations and head movements, while bone-conducting microphones detect both speech vibrations and environmental sounds. This multi-functionality allows the system to use the same hardware for multiple detection purposes, improving wearer speech identification without proportionally increasing device complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent combines traditional airborne sound microphones with vibration-based sensors (accelerometers and bone-conducting microphones) into a unified detection system. By merging these different sensing approaches into a single integrated architecture that processes both audio and vibration channels, the system achieves improved speech identification accuracy while managing complexity through shared processing infrastructure.

Inventive Principle:
Principle #5Merging (Combining)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach enhances the accuracy of voice interface activation and deactivation, ensuring that only wearer speech is processed, reducing unnecessary audio input and improving the overall efficiency of speech recognition systems.

Implementation Method 1

vibration data from sensors like accelerometers and bone-conducting microphones

Methodology Applied
Scientific EffectBone conduction:

Data Source

PatentUS9779758B2Augmenting speech segmentation and recognition using head-mounted vibration and/or motion sensors
Publication Date: 2017.10.03 GOOGLE LLC
  • US9779758B2 patent drawing
  • US9779758B2 patent drawing
  • US9779758B2 patent drawing

AI summary

Example methods and systems use multiple sensors to determine whether a speaker is speaking. Audio data in an audio-channel speech band detected by a microphone can be received. Vibration data in a vibration-channel speech band representative of vibrations detected by a sensor other than the microphone can be received. The microphone and the sensor can be associated with a head-mountable device (HMD). It is determined whether the audio data is causally related to the vibration data. If the audio data and the vibration data are causally related, an indication can be generated that the audio data contains HMD-wearer speech. Causally related audio and vibration data can be used to increase accuracy of text transcription of the HMD-wearer speech. If the audio data and the vibration data are not causally related, an indication can be generated that the audio data does not contain HMD-wearer speech.