Silent Speech Recognition via Facial Skin Strain
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition technologies face challenges in noisy environments and are ineffective for individuals with pronunciation or speech impairments, and silent speech recognition methods require multiple sensor positions, increasing cost and complexity.
Innovation Solution
A computing device uses a position optimization model to select optimal positions on the face for facial skin strain data collection, extracting features and classifying voice through a speech classification model, iteratively updating the models to improve recognition accuracy and reduce the number of required positions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple sensor positions are used for silent speech recognition, then recognition accuracy is improved, but device complexity and cost increase
Solution Approach 1:
The patent applies preliminary action by pre-training a position optimization model offline to identify optimal sensor positions before actual speech recognition. The model is trained using facial skin strain data from multiple positions, and the optimized positions are stored for later use. During actual recognition, only the pre-identified optimal positions are used, eliminating the need to collect and process data from all possible positions in real-time.
Solution Approach 2:
The patent changes the parameter of sensor position selection from a fixed multi-position requirement to a dynamically optimized single or few-position configuration. The position optimization model adjusts which positions are selected based on the specific user's facial characteristics and speech patterns, transforming the system from using a predetermined number of positions to using an optimized subset that maintains accuracy while reducing complexity.
2Speed
If sound-based speech recognition is used, then recognition speed is improved, but it fails in noisy environments and for people with speech impairments
Solution Approach 1:
The patent replaces the acoustic field (sound waves detected by microphones) with the mechanical field (facial skin strain detected by sensors). Instead of measuring air pressure variations from speech sounds, the system measures mechanical deformation of facial skin tissues during speech production. This substitution eliminates dependence on sound transmission, enabling reliable speech recognition in noisy environments and for individuals with speech impairments while maintaining recognition speed through direct physiological measurement.
Data Source
AI summary
A computing device trains a position optimization model for determining, from among a plurality of positions, one or more optimal positions on a face based on a training data set including facial skin strain data at the plurality of positions. The computing device trains a speech classification model for classifying a voice from the facial skin strain data based on the training data at the one or more optimal positions determined by the position optimization model among the training data set.


