Multimodal Speech and Gesture Recognition via Contextual Vocabulary Narrowing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current mobile computing devices lack efficient methods for accurately recognizing user inputs, particularly in noisy environments and when using touch-sensitive screens without a full-sized keyboard, leading to reduced accuracy in speech and handwriting recognition.

Innovation Solution

The implementation of a speech and gesture recognition enhancement technique that uses user-specific supplementary data context to narrow the vocabulary of recognition subsystems, combining speech, touch, and vision inputs to enhance command recognition and interaction with computing devices in a natural and efficient manner.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a full-sized physical keyboard is used for input, then input accuracy is improved, but device portability and compactness deteriorate

Engineering Contradiction:
Improveinput recognition accuracyVSAvoiddevice size
Core Design Contradiction:
Measurement precisionVSVolume of moving object

Solution Approach 1:

The patent replaces the mechanical keyboard system with alternative input modalities including speech recognition (acoustic field substitution) and touch-sensitive screen gestures (optical/capacitive field substitution). This eliminates the need for physical keyboard components while maintaining input functionality, resolving the contradiction between input accuracy and device compactness.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces intermediary recognition subsystems (speech recognition software, handwriting recognition software, gesture recognition software) that mediate between the user and the computing device. These intermediaries translate various input forms into commands, providing keyboard-like functionality without requiring physical keyboard components.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Volume of moving object

If speech recognition is used for input, then device portability is improved, but recognition accuracy in noisy environments deteriorates

Engineering Contradiction:
Improvedevice sizeVSAvoidspeech recognition accuracy
Core Design Contradiction:
Volume of moving objectVSMeasurement precision

Solution Approach 1:

The patent merges multiple input modalities (speech, touch, handwriting, gestures) into a unified recognition system. By combining these different input methods, the system can cross-validate and supplement each other, improving overall recognition accuracy in noisy environments while maintaining device portability.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent implements feedback mechanisms where the recognition subsystem provides confidence scores and alternative interpretations to the user. This allows users to confirm, correct, or rephrase inputs, effectively creating a feedback loop that improves recognition accuracy even in challenging acoustic environments.

Inventive Principle:
Principle #23Feedback

3Volume of moving object

If touch-sensitive screen gestures are used for input, then device portability is improved, but input precision and accuracy deteriorate

Engineering Contradiction:
Improvedevice sizeVSAvoidhandwriting recognition accuracy
Core Design Contradiction:
Volume of moving objectVSMeasurement precision

Solution Approach 1:

The patent implements dynamic adaptation of the recognition system based on the detected input modality. The system dynamically adjusts parameters such as vocabulary selection, recognition thresholds, and processing algorithms based on whether it detects speech, handwriting, or gesture input, optimizing accuracy for each specific input type while maintaining compact device design.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes recognition parameters dynamically based on context and input type. This includes adjusting vocabulary lists based on user-specific supplementary data context, modifying sensitivity thresholds for touch gestures, and adapting speech recognition parameters based on environmental noise levels, thereby improving accuracy without requiring a physical keyboard.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9093072B2Speech and gesture recognition enhancement
Publication Date: 2015.07.28 MICROSOFT TECHNOLOGY LICENSING LLC
  • US9093072B2 patent drawing
  • US9093072B2 patent drawing
  • US9093072B2 patent drawing

AI summary

The recognition of user input to a computing device is enhanced. The user input is either speech, or handwriting data input by the user making screen-contacting gestures, or a combination of one or more prescribed words that are spoken by the user and one or more prescribed screen-contacting gestures that are made by the user, or a combination of one or more prescribed words that are spoken by the user and one or more prescribed non-screen-contacting gestures that are made by the user.