Multimodal Speech Gesture Interaction System for Distant Displays
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current multimodal interfaces require users to hold devices or perform explicit actions to combine speech and gesture inputs, making interactions with distant screen displays cumbersome and inefficient.
Innovation Solution
A system that continuously recognizes speech and tracks gestures without explicit user activation, using wireless microphones and motion sensors like Kinect, to process audio and gesture input streams and generate multimodal commands for interaction with distant displays.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If users hold remote devices or touch screens to provide speech and gesture inputs, then input recognition accuracy is improved, but operation complexity and user effort increase
Solution Approach 1:
The system automatically detects and processes speech and gesture inputs without requiring users to hold devices or perform explicit activation actions. The microphone array and motion sensors continuously monitor the environment, and the system self-determines when to process inputs based on detected speech events and gesture patterns, eliminating the need for users to remember activation sequences or hold devices.
Solution Approach 2:
The patent replaces mechanical interaction (holding remotes, touching screens) with acoustic and optical fields. Microphone arrays capture speech waves and motion sensors detect gesture movements through electromagnetic fields, converting physical actions into digital signals without requiring direct mechanical contact with input devices.
2Loss of information
If explicit activation actions (button presses, touch inputs) are required to signal input processing, then input intent is clearly identified, but interaction time and complexity increase
Solution Approach 1:
The system performs preliminary continuous monitoring of speech and gesture streams before actual input is needed. Microphones and motion sensors are always active, capturing data in advance so that when a user speaks or gestures, the system can immediately process the input without waiting for activation signals, reducing interaction time while maintaining intent clarity through contextual analysis.
Solution Approach 2:
The system uses multimodal feedback to determine input intent by analyzing the combination and timing of speech and gesture events. When speech and gesture inputs occur in close temporal proximity, the system feedback mechanism identifies this as a deliberate multimodal input sequence, clearly identifying user intent without requiring explicit activation actions.
3Ease of operation
If continuous speech and gesture monitoring is implemented without explicit activation, then interaction naturalness is improved, but system complexity and processing load increase
Solution Approach 1:
The system segments the continuous monitoring task into separate specialized components: microphone arrays for speech detection, motion sensors for gesture tracking, and distinct processing modules for each modality. This segmentation allows each component to be optimized independently and reduces overall system complexity by dividing the complex continuous monitoring function into manageable, specialized subsystems.
Solution Approach 2:
The system employs universal sensors and processing algorithms that can detect and interpret multiple types of inputs (speech, gestures, and their combinations) using the same hardware infrastructure. The microphone array and motion sensors serve multiple functions, detecting both speech and gesture events, while the processing system universally handles different input types through integrated algorithms, reducing the need for separate specialized systems.
Data Source
AI summary
Disclosed herein are systems, methods, and non-transitory computer-readable storage media for processing multimodal input. A system configured to practice the method continuously monitors an audio stream associated with a gesture input stream, and detects a speech event in the audio stream. Then the system identifies a temporal window associated with a time of the speech event, and analyzes data from the gesture input stream within the temporal window to identify a gesture event. The system processes the speech event and the gesture event to produce a multimodal command. The gesture in the gesture input stream can be directed to a display, but is remote from the display. The system can analyze the data from the gesture input stream by calculating an average of gesture coordinates within the temporal window.


