Bone Conduction Speech Detection in Earbuds
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech detection systems in headsets, particularly earbuds, face challenges in distinguishing user speech from background noise due to limited microphone placement and occlusion by the ear anatomy, leading to poor noise reduction and signal quality in voice capture scenarios.
Innovation Solution
A device comprising a bone conducted signal sensor and a processor that determines speech metrics and noise estimates, applying signal attenuation factors to enhance speech level estimation and noise suppression, using techniques like MCRA and FFT to improve noise reduction in earbuds.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If microphones are positioned widely apart for beam forming or sidelobe cancellation, then noise reduction capability is improved, but the device geometry is heavily constrained and form factor increases
Solution Approach 1:
The patent introduces bone conduction sensors as an intermediary mechanism to capture speech signals through the skull bone, bypassing the need for traditional air conduction microphones. This mediator enables speech capture without requiring microphone placement near the user's mouth or wide spacing between microphones, thus resolving the geometric constraint while maintaining noise reduction capability
Solution Approach 2:
The patent replaces the mechanical acoustic system (microphones capturing air-borne sound waves) with a biomechanical system (bone conduction sensors detecting vibrations transmitted through skull bone). This substitution eliminates the need for traditional microphone positioning and enables effective speech capture in compact earbud form factors
2Measurement precision
If microphones are positioned near the user's mouth to improve signal to noise ratio, then speech capture quality is improved, but the device cannot be implemented in earbuds due to form factor constraints
Solution Approach 1:
The bone conduction sensor acts as an intermediary that captures speech vibrations directly through the skull bone, eliminating the need for physical proximity between the microphone and the user's mouth. This mediator enables high signal-to-noise ratio speech capture in compact earbud devices
Solution Approach 2:
The patent replaces the acoustic field-based microphone system with a vibration-based bone conduction sensor system, enabling effective speech capture without requiring the microphone to be positioned near the user's mouth, thus accommodating small earbud form factors
3Reliability
If the ear canal and pinna occlude the acoustic signal path to provide in-ear positioning, then private audio delivery is improved, but speech detection from external microphones becomes more difficult
Solution Approach 1:
Instead of trying to detect external speech through the occluded ear canal (which diffuses and attenuates the signal), the patent inverts the approach by detecting speech vibrations directly through the skull bone via bone conduction sensors. This inversion bypasses the occlusion problem entirely while maintaining the private audio delivery benefit of in-ear positioning
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The solution effectively reduces background noise and enhances speech capture quality by accurately estimating speech presence and applying adaptive noise suppression, improving Mean Opinion Score (MOS) results in noisy environments.
Implementation Method 1
at least one signal input component for receiving a bone conducted signal from a bone conducted signal sensor of an earbud
Data Source
AI summary
Embodiments generally relate to a device comprising at least one signal input component for receiving a bone conducted signal from a bone conducted signal sensor of an earbud; memory storing executable code; and a processor configured to access the memory and execute the executable code. Executing the executable code causes the processor to: receive the bone conducted signal; determine at least one speech metric for the received bone conducted signal, wherein the speech metric is based on the input level of the bone conducted signal and a noise estimate for the bone conducted signal; based at least in part on comparing the speech metric to a speech metric threshold, update a speech certainty indicator indicative of a level of certainty of a presence of speech in the bone conducted signal; update at least one signal attenuation factor based on the speech certainty indicator; and generate an updated speech level estimate output by applying the signal attenuation factor to a speech level estimate.


