Head-Mounted Speech Segmentation Using Audio-Vibration Correlation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice interfaces in wearable devices face challenges in accurately determining when to activate and deactivate, as they struggle to differentiate between wearer speech and other audio sources, leading to inefficient speech segmentation.
Innovation Solution
The integration of vibration sensors, such as accelerometers and bone-conducting microphones, with audio microphones to create a vibration channel, allowing correlation with audio data to determine causal relationships and identify wearer speech, thereby activating or deactivating the voice interface appropriately.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If voice interface activation relies solely on audio data from microphones, then the device can detect speech inputs, but it cannot accurately differentiate between wearer speech and other audio sources
Solution Approach 1:
The patent segments the audio detection function into two separate sensing channels: an audio channel using microphones for airborne sound detection, and a vibration channel using accelerometers and bone-conducting microphones for structural vibration detection. This segmentation allows the system to distinguish between audio sources based on which channel detects the vibration, thereby identifying whether speech originates from the wearer.
Solution Approach 2:
The patent introduces vibration sensors (accelerometers and bone-conducting microphones) as intermediary devices that detect mechanical vibrations transmitted through the head structure. These intermediaries provide additional information about the physical source of speech, enabling the system to differentiate wearer speech from external audio sources that would not produce corresponding vibrations in the head.
2Reliability
If the voice interface remains continuously active to capture all potential speech inputs, then no speech is missed, but unnecessary audio processing occurs reducing efficiency
Solution Approach 1:
The patent performs preliminary filtering by continuously monitoring vibration channel data to detect vibrations consistent with speech production before fully activating the voice interface. This preliminary action allows the system to prepare for speech processing in advance, ensuring reliable capture while avoiding full activation for non-speech audio events, thus improving processing efficiency.
Solution Approach 2:
The patent implements a feedback mechanism where vibration channel data continuously informs the state of the voice interface activation. The system uses real-time vibration detection as feedback to dynamically adjust whether the audio channel should be actively processed, creating an efficient loop that maintains reliability while optimizing productivity by avoiding unnecessary processing.
3Measurement precision
If multiple sensor types are integrated to improve speech identification, then wearer speech detection accuracy improves, but device complexity increases
Solution Approach 1:
The patent makes the vibration sensors serve multiple functions: accelerometers detect both speech vibrations and head movements, while bone-conducting microphones detect both speech vibrations and environmental sounds. This multi-functionality allows the system to use the same hardware for multiple detection purposes, improving wearer speech identification without proportionally increasing device complexity.
Solution Approach 2:
The patent combines traditional airborne sound microphones with vibration-based sensors (accelerometers and bone-conducting microphones) into a unified detection system. By merging these different sensing approaches into a single integrated architecture that processes both audio and vibration channels, the system achieves improved speech identification accuracy while managing complexity through shared processing infrastructure.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach enhances the accuracy of voice interface activation and deactivation, ensuring that only wearer speech is processed, reducing unnecessary audio input and improving the overall efficiency of speech recognition systems.
Implementation Method 1
vibration data from sensors like accelerometers and bone-conducting microphones
Data Source
AI summary
Example methods and systems use multiple sensors to determine whether a speaker is speaking. Audio data in an audio-channel speech band detected by a microphone can be received. Vibration data in a vibration-channel speech band representative of vibrations detected by a sensor other than the microphone can be received. The microphone and the sensor can be associated with a head-mountable device (HMD). It is determined whether the audio data is causally related to the vibration data. If the audio data and the vibration data are causally related, an indication can be generated that the audio data contains HMD-wearer speech. Causally related audio and vibration data can be used to increase accuracy of text transcription of the HMD-wearer speech. If the audio data and the vibration data are not causally related, an indication can be generated that the audio data does not contain HMD-wearer speech.


