Sub-vocal Speech Recognition Using EMG and IMU Sensors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Individuals with speech impairments, such as those with degenerative muscle diseases or post-stroke conditions, face challenges using existing speech assistance devices like electrolarynx, which require vibration for sound production, limiting their ability to communicate effectively.
Innovation Solution
A sub-vocal speech acquisition device (SVSA) utilizing EMG sensors and an IMU to detect muscle activity and jaw position, processing these signals to recognize intended speech without audible vibration, and translating them into audible output through a neural network and language model for digital devices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If electrolarynx is used to assist speech production, then speech clarity is improved, but the device requires vibration capability which excludes users who cannot create vibrations
Solution Approach 1:
The patent replaces the mechanical vibration-based electrolarynx system with an electromagnetic sensing system. EMG sensors detect electrical signals from muscle activity during sub-vocal speech, and an IMU captures jaw movement patterns. This substitution eliminates the need for audible vibration while maintaining speech recognition capability, allowing users who cannot produce vibrations to communicate effectively.
Solution Approach 2:
The patent introduces sensors as intermediaries between the user's sub-vocal speech mechanisms and the digital device. The EMG sensors and IMU act as mediators that capture subtle muscle and jaw movements without requiring audible sound production. This intermediary system bridges the gap between sub-vocal intent and digital output, enabling communication for users who cannot produce traditional speech sounds.
2Measurement precision
If sensors detect sub-vocal muscle activity and jaw position, then speech recognition accuracy is improved, but system complexity increases
Solution Approach 1:
The patent employs a multi-functional sensor system where EMG sensors serve dual purposes: detecting both muscle electrical activity and jaw movement patterns. The IMU similarly functions to capture both position and motion data. This multi-functionality reduces the need for separate specialized sensors, thereby managing system complexity while maintaining high measurement precision for speech recognition.
3Loss of information
If the system processes sub-vocal signals without audible vibration, then privacy and silent communication are improved, but signal processing complexity increases
Solution Approach 1:
The patent extracts only the essential features needed for speech recognition from the complex sensor data stream. The system processes EMG and IMU signals to identify key muscle activation patterns and jaw movement characteristics, discarding redundant information. This extraction approach enables private silent communication while managing processing complexity by focusing only on discriminative features.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Enables users to convey information silently and effectively, overcoming the limitations of traditional speech devices by accurately recognizing sub-vocal speech patterns and converting them into audible speech or text, enhancing communication for a wider range of speech-impaired individuals.
Implementation Method 1
Electromyography (EMG) sensors measure electrical activity in response to nerve stimulation of the muscles.
Implementation Method 2
an electrode pair plus Inertial Measurement Unit (IMU) mounted under the chin. The IMU may contain accelerometers, gyroscopes, or magnetometers.
Data Source
AI summary
A sub-vocal speech recognition (SVSR) apparatus includes a headset that is worn over an ear and electromyography (EMG) electrodes and an Inertial Measurement Unit (IMU) in contact with a user's skin in a position over the neck, under the chin and behind the ear. When a user speaks or mouths words, the EMG and IMU signals are recorded by sensors and amplified and filtered, before being divided in multi-millisecond time windows. These time windows are then transmitted to the interface computing device for Mel Frequency Cepstral Coefficients (MFCC) conversion into aggregated vector representation (AVR). The AVR is the input to the SVSR system, which utilizes a neural network, CTC function, and language model to classify the phoneme. The phonemes are then combined into words and sent back to the interface computing device, where they are played either as audible output, such as from a speaker, or non-audible output, such as text.


