Wearable Silent Speech Device Using EMG and Audio Sensors

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional communication systems rely on voiced speech, typing, or selection methods, which may not be accurate or efficient over extended periods and under varying user conditions, particularly for silent or sub-vocalized speech inputs.

Innovation Solution

A wearable silent speech device using EMG and microphone sensors to record speech signals, adjust a silent speech machine learning model based on these signals, and condition it for improved recognition of silent speech inputs, allowing for continuous and accurate interaction with devices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional voice-based systems are used for speech recognition, then they can process voiced speech inputs, but they fail to accurately recognize silent or sub-vocalized speech inputs

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidcompatibility with silent speech inputs
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system divides speech signal processing into separate channels: an audio channel for voiced speech and an EMG channel for silent speech. Each channel is processed by dedicated neural networks that are trained independently on their respective modalities, allowing the system to accurately recognize both voiced and silent speech inputs without interference between modalities

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The wearable device is designed to handle multiple speech modalities (voiced, silent, and sub-vocalized) through a unified architecture that combines audio and EMG processing. The system can adapt to different user speaking styles and conditions by processing signals from both sensors, making it universally applicable to various speech input types

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Quantity of substance

If a machine learning model is trained on limited data, then the device can operate with smaller datasets, but the model accuracy and adaptability to changing user conditions deteriorate

Engineering Contradiction:
Improvetraining data volumeVSAvoidmodel accuracy over extended periods
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The system implements continuous feedback loops where the wearable device collects new speech data from users during normal operation. This feedback data is used to periodically retrain and fine-tune the neural network models, allowing the system to adapt to changing user conditions, accents, and speaking styles while maintaining high accuracy over extended periods

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary training on diverse speech datasets before deployment, creating robust initial models. During operation, it continuously collects and stores speech samples for later use in model refinement, preparing the system in advance to handle various user conditions without requiring extensive real-time training

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If multiple sensors are used to record speech signals, then the system can capture comprehensive speech data, but the device complexity and data processing requirements increase

Engineering Contradiction:
Improvespeech signal completenessVSAvoidsensor integration and processing
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system segments the speech recognition task into separate processing streams: one for audio signals and another for EMG signals. Each sensor type is processed by dedicated neural networks that operate independently, reducing the computational complexity of integrating multiple sensor modalities while maintaining comprehensive speech data capture

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system uses an intermediary processing layer that combines audio and EMG features before feeding them to the final speech recognition model. This intermediary layer processes and reconciles the different sensor modalities, simplifying the overall system architecture while preserving the comprehensive information from both sensors

Inventive Principle:
Principle #24Intermediary (Mediator)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enables effective and continuous communication through silent speech recognition, enhancing user interaction with devices over time and across changing conditions without the limitations of traditional voice-based systems.

Implementation Method 1

The first sensor is an EMG sensor and the second sensor is a microphone

Methodology Applied
Scientific EffectElectromyography (EMG):

Implementation Method 2

The first sensor is an EMG sensor and the second sensor is a microphone

Methodology Applied
Scientific EffectAcoustic detection:

Data Source

PatentUS20240296833A1Wearable silent speech device, systems, and methods for adjusting a machine learning model
Publication Date: 2024.09.05 WISPR AI INC
  • US20240296833A1 patent drawing
  • US20240296833A1 patent drawing
  • US20240296833A1 patent drawing

AI summary

The present disclosure relates to methods and systems for adjusting a silent speech machine learning model for use with a wearable silent speech device. In some embodiments, a method may include recording speech signals from a user, using a first sensor and a second sensor of a wearable silent speech device. The method may include providing for a silent speech machine learning model for use with the wearable silent speech device, determining whether the silent speech machine learning model is to be adjusted, and in response to determining the silent speech machine learning model is to be adjusted, adjusting the silent speech machine learning model based on at least the speech signals recorded using the first sensor and the second sensor.