Silent Speech Recognition via Facial Skin Strain

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition technologies face challenges in noisy environments and are ineffective for individuals with pronunciation or speech impairments, and silent speech recognition methods require multiple sensor positions, increasing cost and complexity.

Innovation Solution

A computing device uses a position optimization model to select optimal positions on the face for facial skin strain data collection, extracting features and classifying voice through a speech classification model, iteratively updating the models to improve recognition accuracy and reduce the number of required positions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If multiple sensor positions are used for silent speech recognition, then recognition accuracy is improved, but device complexity and cost increase

Engineering Contradiction:
Improverecognition accuracyVSAvoidnumber of sensor positions
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by pre-training a position optimization model offline to identify optimal sensor positions before actual speech recognition. The model is trained using facial skin strain data from multiple positions, and the optimized positions are stored for later use. During actual recognition, only the pre-identified optimal positions are used, eliminating the need to collect and process data from all possible positions in real-time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes the parameter of sensor position selection from a fixed multi-position requirement to a dynamically optimized single or few-position configuration. The position optimization model adjusts which positions are selected based on the specific user's facial characteristics and speech patterns, transforming the system from using a predetermined number of positions to using an optimized subset that maintains accuracy while reducing complexity.

Inventive Principle:
Principle #35Parameter changes

2Speed

If sound-based speech recognition is used, then recognition speed is improved, but it fails in noisy environments and for people with speech impairments

Engineering Contradiction:
Improverecognition speedVSAvoidperformance in noisy environments
Core Design Contradiction:
SpeedVSReliability

Solution Approach 1:

The patent replaces the acoustic field (sound waves detected by microphones) with the mechanical field (facial skin strain detected by sensors). Instead of measuring air pressure variations from speech sounds, the system measures mechanical deformation of facial skin tissues during speech production. This substitution eliminates dependence on sound transmission, enabling reliable speech recognition in noisy environments and for individuals with speech impairments while maintaining recognition speed through direct physiological measurement.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11810549B2Speech recognition using facial skin strain data
Publication Date: 2023.11.07 SAMSUNG ELECTRONICS CO LTD
  • US11810549B2 patent drawing
  • US11810549B2 patent drawing
  • US11810549B2 patent drawing

AI summary

A computing device trains a position optimization model for determining, from among a plurality of positions, one or more optimal positions on a face based on a training data set including facial skin strain data at the plurality of positions. The computing device trains a speech classification model for classifying a voice from the facial skin strain data based on the training data at the one or more optimal positions determined by the position optimization model among the training data set.