Multimodal Speech Gesture Interaction System for Distant Displays

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current multimodal interfaces require users to hold devices or perform explicit actions to combine speech and gesture inputs, making interactions with distant screen displays cumbersome and inefficient.

Innovation Solution

A system that continuously recognizes speech and tracks gestures without explicit user activation, using wireless microphones and motion sensors like Kinect, to process audio and gesture input streams and generate multimodal commands for interaction with distant displays.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If users hold remote devices or touch screens to provide speech and gesture inputs, then input recognition accuracy is improved, but operation complexity and user effort increase

Engineering Contradiction:
Improveinput recognition accuracyVSAvoiduser effort
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system automatically detects and processes speech and gesture inputs without requiring users to hold devices or perform explicit activation actions. The microphone array and motion sensors continuously monitor the environment, and the system self-determines when to process inputs based on detected speech events and gesture patterns, eliminating the need for users to remember activation sequences or hold devices.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces mechanical interaction (holding remotes, touching screens) with acoustic and optical fields. Microphone arrays capture speech waves and motion sensors detect gesture movements through electromagnetic fields, converting physical actions into digital signals without requiring direct mechanical contact with input devices.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Loss of information

If explicit activation actions (button presses, touch inputs) are required to signal input processing, then input intent is clearly identified, but interaction time and complexity increase

Engineering Contradiction:
Improveinput intent clarityVSAvoidinteraction time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system performs preliminary continuous monitoring of speech and gesture streams before actual input is needed. Microphones and motion sensors are always active, capturing data in advance so that when a user speaks or gestures, the system can immediately process the input without waiting for activation signals, reducing interaction time while maintaining intent clarity through contextual analysis.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses multimodal feedback to determine input intent by analyzing the combination and timing of speech and gesture events. When speech and gesture inputs occur in close temporal proximity, the system feedback mechanism identifies this as a deliberate multimodal input sequence, clearly identifying user intent without requiring explicit activation actions.

Inventive Principle:
Principle #23Feedback

3Ease of operation

If continuous speech and gesture monitoring is implemented without explicit activation, then interaction naturalness is improved, but system complexity and processing load increase

Engineering Contradiction:
Improveinteraction naturalnessVSAvoidsystem complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system segments the continuous monitoring task into separate specialized components: microphone arrays for speech detection, motion sensors for gesture tracking, and distinct processing modules for each modality. This segmentation allows each component to be optimized independently and reduces overall system complexity by dividing the complex continuous monitoring function into manageable, specialized subsystems.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system employs universal sensors and processing algorithms that can detect and interpret multiple types of inputs (speech, gestures, and their combinations) using the same hardware infrastructure. The microphone array and motion sensors serve multiple functions, detecting both speech and gesture events, while the processing system universally handles different input types through integrated algorithms, reducing the need for separate specialized systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11189288B2System and method for continuous multimodal speech and gesture interaction
Publication Date: 2021.11.30 MICROSOFT TECHNOLOGY LICENSING LLC
  • US11189288B2 patent drawing
  • US11189288B2 patent drawing
  • US11189288B2 patent drawing

AI summary

Disclosed herein are systems, methods, and non-transitory computer-readable storage media for processing multimodal input. A system configured to practice the method continuously monitors an audio stream associated with a gesture input stream, and detects a speech event in the audio stream. Then the system identifies a temporal window associated with a time of the speech event, and analyzes data from the gesture input stream within the temporal window to identify a gesture event. The system processes the speech event and the gesture event to produce a multimodal command. The gesture in the gesture input stream can be directed to a display, but is remote from the display. The system can analyze the data from the gesture input stream by calculating an average of gesture coordinates within the temporal window.