Multimodal Signal Analysis for Imprecise Gaze and Gesture Commands

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems for detecting and responding to natural human movements and conversational queries are limited by their reliance on specific domains and locations, and often require precise cues, failing to effectively utilize naturalistic human behaviors for general purposes.

Innovation Solution

A method and system that utilize multimodal signal analysis to process commands and queries by combining gaze and gesture data, even with imprecise cues, to identify and act upon entities of interest, using a variety of sensors and databases to provide accurate and timely responses without requiring constrained interfaces.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If systems use specific domain restrictions and precise cue requirements, then measurement precision is improved, but adaptability deteriorates

Engineering Contradiction:
Improvecue precisionVSAvoiddomain adaptability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system employs multiple signal modalities (gaze tracking, gesture recognition, speech processing) that can function across different domains and contexts. Each modality serves multiple purposes: gaze data can identify objects of interest, determine user attention, or indicate reading direction; gesture data can complement speech commands or independently control system functions. This multi-functional approach allows the system to adapt to various domains while maintaining acceptable precision through redundant measurement channels.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system merges data from multiple independent signal modalities to compensate for the imprecision of individual cues. By combining gaze direction data with gesture data and speech recognition, the system creates a more robust and adaptable command interpretation framework. The fusion of these modalities allows the system to maintain measurement precision across diverse domains by cross-validating signals from different sources.

Inventive Principle:
Principle #5Merging (Combining)

2Device complexity

If systems require constrained interfaces and specific camera angles, then device complexity is reduced, but ease of operation deteriorates

Engineering Contradiction:
Improvesystem complexityVSAvoidinterface ease
Core Design Contradiction:
Device complexityVSEase of operation

Solution Approach 1:

The system automatically processes natural human behaviors (gaze movements, spontaneous gestures, conversational speech) without requiring users to learn specialized interface protocols. The multimodal signal processing framework self-adapts to capture and interpret these natural cues, eliminating the need for constrained interfaces while maintaining manageable complexity through automated signal fusion and interpretation algorithms.

Inventive Principle:
Principle #25Self-service

3Adaptability or versatility

If systems use multiple signal modalities, then adaptability is improved, but device complexity deteriorates

Engineering Contradiction:
Improvesignal modality versatilityVSAvoidprocessing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system segments the complex task of natural language command interpretation into distinct processing stages: gaze signal acquisition and processing, gesture signal acquisition and processing, speech recognition, temporal correlation analysis, and final command synthesis. Each module handles a specific aspect of signal processing independently, then integrates results through temporal correlation. This segmentation manages complexity by creating modular, independently testable components while maintaining high adaptability through the comprehensive multimodal approach.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentEP3485351B1Command processing using multimodal signal analysis
Publication Date: 2023.07.12 APPLE INC
  • EP3485351B1 patent drawingFigure 1
  • EP3485351B1 patent drawingFigure 2
  • EP3485351B1 patent drawingFigure 3

AI summary

A first set of signals corresponding to a first signal modality (such as the direction of a gaze) during a time interval is collected from an individual. A second set of signals corresponding to a different signal modality (such as hand-pointing gestures made by the individual) is also collected. In response to a command, where the command does not identify a particular object to which the command is directed, the first and second set of signals is used to identify candidate objects of interest, and an operation associated with a selected object from the candidates is performed.