Wearable EMG Sensors Augment Speech Recognition Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems lack the ability to effectively incorporate contextual information from user movements and muscle activations, leading to suboptimal performance in accuracy and speed.

Innovation Solution

The integration of neuromuscular signals, recorded using electromyography (EMG) sensors, to augment speech data, allowing the system to interpret musculo-skeletal representations and modify operations such as formatting, punctuation, and interaction modes, thereby enhancing speech recognition performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If only speech data is used as input to the speech recognition system, then the system structure remains simple, but the speech recognition accuracy and speed are suboptimal

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidsystem structure
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines speech data with neuromuscular signals (EMG data) from wearable sensors to create a hybrid input system. The speech recognizer processes both audio signals and neuromuscular signals simultaneously, merging multiple data sources to improve recognition accuracy while managing system complexity through integrated processing architecture.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces neuromuscular signals as an intermediary that bridges the gap between user intent and speech recognition. The wearable EMG sensors detect muscle activations related to speech production, providing complementary information that enhances the speech recognizer's ability to accurately interpret speech, especially in challenging acoustic environments.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If only speech data is used as input, then the system operation remains simple, but the speech recognition speed is limited

Engineering Contradiction:
Improvespeech recognition speedVSAvoidsystem operation
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system merges speech data processing with neuromuscular signal processing to accelerate recognition. By parallel processing of both data types and combining their outputs, the system achieves faster and more accurate speech recognition compared to using speech data alone, as the neuromuscular signals provide additional constraints that reduce processing ambiguity.

Inventive Principle:
Principle #5Merging (Combining)

3Measurement precision

If neuromuscular signals are integrated to augment speech data, then speech recognition accuracy improves, but the device complexity increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidsystem structure
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the speech recognition system into distinct functional modules: wearable EMG sensors for neuromuscular signal acquisition, signal processing components for extracting relevant features, and a speech recognizer that integrates multiple data sources. This modular segmentation manages complexity by allowing each component to be optimized independently while maintaining overall system accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The wearable device with EMG sensors serves multiple functions: detecting muscle activations related to speech production, providing contextual information about user state, and potentially controlling other device functions. This multi-functionality justifies the added complexity by providing multiple benefits from a single integrated system.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Measurement precision

If neuromuscular signals are used for hybrid input modes, then control precision over speech recognition processes improves, but the ease of operation decreases

Engineering Contradiction:
Improvecontrol precisionVSAvoiduser interaction
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system uses neuromuscular signals that occur naturally during speech production, requiring no additional user effort or training. The EMG sensors automatically detect muscle activations that accompany speech, providing control precision without increasing the ease of operation, as the system leverages existing physiological signals rather than requiring new user actions.

Inventive Principle:
Principle #25Self-service

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach improves speech recognition accuracy and speed by leveraging contextual information from user movements and muscle activations, enabling more precise control over speech recognition processes and hybrid input modes.

Implementation Method 1

a plurality of neuromuscular sensors arranged on one or more wearable devices. The plurality of neuromuscular sensors is configured to continuously record a plurality of neuromuscular signals from the user

Methodology Applied
Scientific EffectElectromyography (EMG):

Data Source

PatentUS11036302B1Wearable devices and methods for improved speech recognition
Publication Date: 2021.06.15 META PLATFORMS TECHNOLOGIES LLC
  • US11036302B1 patent drawing
  • US11036302B1 patent drawing
  • US11036302B1 patent drawing

AI summary

Systems and methods for using neuromuscular information to improve speech recognition. The system includes a plurality of neuromuscular sensors, arranged on one or more wearable devices, wherein the plurality of neuromuscular sensors is configured to continuously record a plurality of neuromuscular signals from a user, at least one storage device configured to store one or more trained statistical models, and at least one computer processor programmed to provide as an input to the one or more trained statistical models, the plurality of neuromuscular signals or signals derived from the plurality of neuromuscular signals, determine based, at least in part, on an output of the one or more trained statistical models, at least one instruction for modifying an operation of a speech recognizer, and provide the at least one instruction to the speech recognizer.