Subvocalized Speech Recognition via Multi-Modal Sensor Fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition systems fail to accurately detect and interpret non-audible and subvocalized speech in noisy environments and private settings, such as military and police applications, where maintaining low acoustic signatures is crucial.
Innovation Solution
A system utilizing a combination of sensors, including an air microphone, camera, and in-ear proximity sensor, integrated into an eyeglass-type near-eye-display platform, processes data using machine learning techniques like frame-based, sequence-based, and sequence-to-sequence phoneme classification to detect and recognize mouthed and whispered commands.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional speech recognition systems are used, then audible speech can be detected, but they fail in noisy environments and cannot detect subvocalized speech
Solution Approach 1:
The patent segments the speech detection task into multiple modalities: audible speech detection, subvocalized speech detection, and silence detection. Each modality uses specialized sensors (microphones, cameras, proximity sensors) that process different types of input signals independently, then combines results to achieve robust speech recognition in noisy environments
Solution Approach 2:
The system implements multi-functional sensors that can detect multiple types of speech inputs. The proximity sensor and camera work together to detect both audible and subvocalized speech, while the microphone array handles both quiet and noisy environments, creating a universal speech recognition system
2Reliability
If audible speech is used for interaction, then speech recognition is accurate, but it disrupts privacy and increases acoustic signature
Solution Approach 1:
The patent introduces subvocalized speech as an intermediary input mode between silent mouthing and audible speech. This intermediary mode allows users to communicate with the system without producing audible sound, maintaining privacy and low acoustic signature while still providing sufficient signal for recognition through proximity sensors and cameras
Solution Approach 2:
The system changes the detection parameters from acoustic field measurements (microphones) to optical and proximity field measurements (cameras and proximity sensors) when detecting subvocalized speech. This parameter change enables detection of speech without acoustic emission, reducing the acoustic signature to near-zero levels
3Measurement precision
If multiple sensors are used for subvocalized speech detection, then detection accuracy improves, but device complexity increases
Solution Approach 1:
The patent merges multiple sensor types (microphones, cameras, proximity sensors) into an integrated eyewear platform. The sensors are spatially distributed across the glasses frame and temple, with their functions combined through a unified processing system that correlates data from all sensors to detect subvocalized speech patterns
Solution Approach 2:
The system uses the user's own anatomical features (ear canal, jaw movement, skin vibrations) as natural sensors. The proximity sensor detects jaw movement during subvocalized speech, while skin-mounted sensors detect vibrations from vocal cord activity, allowing the system to leverage the user's body as part of the sensing mechanism
Data Source
AI summary
A system for subvocalized speech recognition includes a plurality of sensors, a controller and a processor. The sensors are coupled to a near-eye display (NED) and configured to capture non-audible and subvocalized commands provided by a user wearing the NED. The controller interfaced with the plurality of sensors is configured to combine data acquired by each of the plurality of sensors. The processor coupled to the controller is configured to extract one or more features from the combined data, compare the one or more extracted features with a pre-determined set of commands, and determine a command of the user based on the comparison.


