Ear-Worn Optical Silent Speech Detection in Noisy Settings
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech communication technologies face challenges in enabling silent speech detection, particularly in noisy environments, and require users to divert attention from their surroundings for input, raising privacy concerns and complicating communication in public spaces.
Innovation Solution
A non-contact sensing device using coherent light to detect secondary speckle patterns from facial movements, combined with optional skin contact electrodes, processes these movements to generate speech outputs without vocalization, integrated into wearable forms like headphones or smartphones.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If acoustic sensors are used to detect speech, then speech can be detected, but the system cannot detect silent speech and is susceptible to noise interference
Solution Approach 1:
The system segments the detection task into multiple independent sensing modalities: acoustic sensors detect audible speech, while imaging sensors (camera, TOF sensor, or infrared sensor) detect facial muscle movements and thermal changes. Each sensor type handles a specific aspect of speech detection, allowing the system to detect both spoken and silent speech independently and fuse the results for improved accuracy.
2Adaptability or versatility
If the system uses multiple sensor types, then detection capability is improved, but device complexity increases
Solution Approach 1:
The system employs imaging sensors that serve multiple functions: they detect facial muscle movements for silent speech recognition, capture thermal radiation for silent speech detection, and can also monitor visual speech patterns. This multi-functionality reduces the need for entirely separate sensor systems while maintaining comprehensive detection capability across different speech modes.
3Device complexity
If acoustic sensors are used, then the system is simple, but noise from other sources interferes with detection
Solution Approach 1:
The system introduces imaging sensors as an intermediary detection mechanism that captures facial muscle movements and thermal changes as intermediate signals. These intermediate signals serve as proxies for vocal cord activity, allowing the system to infer speech intent without directly capturing acoustic signals that are susceptible to noise interference from environmental sources.
4Ease of manufacture
If traditional speech recognition is used, then processing is simple, but it cannot recognize silent speech
Solution Approach 1:
The system replaces the traditional acoustic-based mechanical speech recognition process with an imaging-based detection system. Instead of processing acoustic waveforms, the system captures visual images of facial muscle movements and thermal radiation patterns, then processes these image data to recognize silent speech. This substitution enables recognition of silent speech while maintaining computational processing capabilities.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Enables silent speech communication that is imperceptible to others, immune to ambient noise, and maintains user attention on surroundings, providing reliable speech output through synthesized audio or text.
Implementation Method 1
an imaging sensor to detect radiation from your face
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A sensing device (20, 60) includes a bracket (22) configured to fit an ear of a user (24) of the device. An optical sensing head (28) is held by the bracket in a location in proximity to a face of the user and senses light reflected from the face and to output a signal in response to the detected light. Processing circuitry (70, 75) processes the signal to generate a speech output.