Subvocalized Speech Recognition via Multi-Modal Sensor Fusion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems fail to accurately detect and interpret non-audible and subvocalized speech in noisy environments and private settings, such as military and police applications, where maintaining low acoustic signatures is crucial.

Innovation Solution

A system utilizing a combination of sensors, including an air microphone, camera, and in-ear proximity sensor, integrated into an eyeglass-type near-eye-display platform, processes data using machine learning techniques like frame-based, sequence-based, and sequence-to-sequence phoneme classification to detect and recognize mouthed and whispered commands.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional speech recognition systems are used, then audible speech can be detected, but they fail in noisy environments and cannot detect subvocalized speech

Engineering Contradiction:
Improvespeech detection accuracyVSAvoidnoise interference
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent segments the speech detection task into multiple modalities: audible speech detection, subvocalized speech detection, and silence detection. Each modality uses specialized sensors (microphones, cameras, proximity sensors) that process different types of input signals independently, then combines results to achieve robust speech recognition in noisy environments

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system implements multi-functional sensors that can detect multiple types of speech inputs. The proximity sensor and camera work together to detect both audible and subvocalized speech, while the microphone array handles both quiet and noisy environments, creating a universal speech recognition system

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If audible speech is used for interaction, then speech recognition is accurate, but it disrupts privacy and increases acoustic signature

Engineering Contradiction:
Improvespeech recognition reliabilityVSAvoidacoustic signature
Core Design Contradiction:
ReliabilityVSObject-generated harmful factors

Solution Approach 1:

The patent introduces subvocalized speech as an intermediary input mode between silent mouthing and audible speech. This intermediary mode allows users to communicate with the system without producing audible sound, maintaining privacy and low acoustic signature while still providing sufficient signal for recognition through proximity sensors and cameras

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system changes the detection parameters from acoustic field measurements (microphones) to optical and proximity field measurements (cameras and proximity sensors) when detecting subvocalized speech. This parameter change enables detection of speech without acoustic emission, reducing the acoustic signature to near-zero levels

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If multiple sensors are used for subvocalized speech detection, then detection accuracy improves, but device complexity increases

Engineering Contradiction:
Improvesubvocalized speech detection accuracyVSAvoidsensor system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges multiple sensor types (microphones, cameras, proximity sensors) into an integrated eyewear platform. The sensors are spatially distributed across the glasses frame and temple, with their functions combined through a unified processing system that correlates data from all sensors to detect subvocalized speech patterns

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system uses the user's own anatomical features (ear canal, jaw movement, skin vibrations) as natural sensors. The proximity sensor detects jaw movement during subvocalized speech, while skin-mounted sensors detect vibrations from vocal cord activity, allowing the system to leverage the user's body as part of the sensing mechanism

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS10665243B1Subvocalized speech recognition
Publication Date: 2020.05.26 META PLATFORMS TECHNOLOGIES LLC
  • US10665243B1 patent drawing
  • US10665243B1 patent drawing
  • US10665243B1 patent drawing

AI summary

A system for subvocalized speech recognition includes a plurality of sensors, a controller and a processor. The sensors are coupled to a near-eye display (NED) and configured to capture non-audible and subvocalized commands provided by a user wearing the NED. The controller interfaced with the plurality of sensors is configured to combine data acquired by each of the plurality of sensors. The processor coupled to the controller is configured to extract one or more features from the combined data, compare the one or more extracted features with a pre-determined set of commands, and determine a command of the user based on the comparison.