Multi-Modal Virtual Keyboard Input Using Gaze and Trackpad

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional artificial reality input interfaces face challenges such as inaccuracy, slowness, user fatigue, and ergonomic discomfort due to reliance on single input modalities like tracked motion, eye gaze, or touch.

Innovation Solution

The implementation of a virtual interface that combines multiple input modalities, including user gaze and hand input via a controller device's trackpad, to dynamically move a cursor on a virtual keyboard and resolve character selections efficiently.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If single input modality (tracked motion, eye gaze, or touch) is used, then device complexity is reduced, but input accuracy and speed deteriorate

Engineering Contradiction:
Improveinput accuracyVSAvoidinput interface complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines multiple input modalities (gaze tracking, hand tracking, and controller input) into a unified input system. The gaze cursor provides visual feedback of intended input location, while hand tracking and controller inputs enable precise selection and manipulation. This merging of modalities resolves the contradiction by achieving high input accuracy through multi-modal integration while managing complexity through coordinated system design.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The input system is designed to accept and process multiple types of input modalities through a unified framework. The system can operate with gaze input alone, hand tracking alone, or in combination with controller inputs, making the system universally adaptable to different user needs and situations. This multi-functionality enables high accuracy across diverse input methods while maintaining a consistent interface model.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Productivity

If single input modality is used, then ease of operation is improved, but productivity deteriorates

Engineering Contradiction:
Improveinput speedVSAvoiduser fatigue
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The system dynamically adapts between different input modalities based on user needs and context. Gaze tracking enables rapid navigation and selection without physical movement, reducing fatigue for prolonged use. Hand tracking and controller inputs provide precise control when needed. This dynamic switching between modalities maintains high input speed while minimizing user fatigue by using the least effortful modality for each task.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The gaze cursor acts as an intermediary between the user's visual attention and the input selection mechanism. It provides real-time visual feedback of where input will be registered, enabling users to navigate and select items rapidly without precise hand movements. This intermediary mechanism significantly increases input speed while reducing the physical effort and fatigue associated with traditional pointing and clicking.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If multiple input modalities are combined, then input accuracy and speed are improved, but device complexity increases

Engineering Contradiction:
Improveinput speedVSAvoidinput interface complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The input system is segmented into distinct functional components: gaze tracking for navigation and intent detection, hand tracking for gesture recognition, and controller inputs for precise manipulation. Each component operates independently but contributes to the unified input process. This segmentation enables high input speed through specialized optimization of each modality while managing complexity by keeping components modular and independently manageable.

Inventive Principle:
Principle #1Segmentation

4Ease of operation

If multiple input modalities are combined, then user fatigue is reduced, but device complexity increases

Engineering Contradiction:
Improveuser fatigueVSAvoidinput interface complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system automatically selects and switches between input modalities based on the current task context and user behavior patterns. It self-adjusts to use gaze tracking for navigation when appropriate, hand tracking for gestures when needed, and controller inputs for precise selections. This self-service capability reduces user fatigue by automatically using the least effortful modality for each task while managing complexity through automated decision-making algorithms.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12242666B2Artificial reality input using multiple modalities
Publication Date: 2025.03.04 META PLATFORMS TECHNOLOGIES LLC
  • US12242666B2 patent drawing
  • US12242666B2 patent drawing
  • US12242666B2 patent drawing

AI summary

Aspects of the disclosure are directed to an interface for receiving input using multiple modalities in an artificial reality environment. The interface can be a virtual keyboard displayed in an artificial reality environment that includes characters arranged as elements. Implementations include an artificial reality device/system for displaying the artificial reality environment and receiving user input in a first modality, and a controller device for receiving user input in an additional input modality. For example, the artificial reality system can be configured to receive user gaze input as a first input modality and the controller device can be configured to receive input in a second modality, such as touch input received at a trackpad. An interface manager can process input in the multiple modalities to control an indicator on the virtual interface. The interface manager can also resolve character selections from the virtual interface according to the input.