Multimodal Speech and Gesture Recognition via Contextual Vocabulary Narrowing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current mobile computing devices lack efficient methods for accurately recognizing user inputs, particularly in noisy environments and when using touch-sensitive screens without a full-sized keyboard, leading to reduced accuracy in speech and handwriting recognition.
Innovation Solution
The implementation of a speech and gesture recognition enhancement technique that uses user-specific supplementary data context to narrow the vocabulary of recognition subsystems, combining speech, touch, and vision inputs to enhance command recognition and interaction with computing devices in a natural and efficient manner.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a full-sized physical keyboard is used for input, then input accuracy is improved, but device portability and compactness deteriorate
Solution Approach 1:
The patent replaces the mechanical keyboard system with alternative input modalities including speech recognition (acoustic field substitution) and touch-sensitive screen gestures (optical/capacitive field substitution). This eliminates the need for physical keyboard components while maintaining input functionality, resolving the contradiction between input accuracy and device compactness.
Solution Approach 2:
The patent introduces intermediary recognition subsystems (speech recognition software, handwriting recognition software, gesture recognition software) that mediate between the user and the computing device. These intermediaries translate various input forms into commands, providing keyboard-like functionality without requiring physical keyboard components.
2Volume of moving object
If speech recognition is used for input, then device portability is improved, but recognition accuracy in noisy environments deteriorates
Solution Approach 1:
The patent merges multiple input modalities (speech, touch, handwriting, gestures) into a unified recognition system. By combining these different input methods, the system can cross-validate and supplement each other, improving overall recognition accuracy in noisy environments while maintaining device portability.
Solution Approach 2:
The patent implements feedback mechanisms where the recognition subsystem provides confidence scores and alternative interpretations to the user. This allows users to confirm, correct, or rephrase inputs, effectively creating a feedback loop that improves recognition accuracy even in challenging acoustic environments.
3Volume of moving object
If touch-sensitive screen gestures are used for input, then device portability is improved, but input precision and accuracy deteriorate
Solution Approach 1:
The patent implements dynamic adaptation of the recognition system based on the detected input modality. The system dynamically adjusts parameters such as vocabulary selection, recognition thresholds, and processing algorithms based on whether it detects speech, handwriting, or gesture input, optimizing accuracy for each specific input type while maintaining compact device design.
Solution Approach 2:
The patent changes recognition parameters dynamically based on context and input type. This includes adjusting vocabulary lists based on user-specific supplementary data context, modifying sensitivity thresholds for touch gestures, and adapting speech recognition parameters based on environmental noise levels, thereby improving accuracy without requiring a physical keyboard.
Data Source
AI summary
The recognition of user input to a computing device is enhanced. The user input is either speech, or handwriting data input by the user making screen-contacting gestures, or a combination of one or more prescribed words that are spoken by the user and one or more prescribed screen-contacting gestures that are made by the user, or a combination of one or more prescribed words that are spoken by the user and one or more prescribed non-screen-contacting gestures that are made by the user.


