Gaze-and-Hand Tracking for Intent-Aware 3D Interface Inputs

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods and interfaces for processing inputs in three-dimensional environments, such as augmented and virtual reality, are cumbersome, inefficient, and prone to errors, leading to a significant cognitive burden on users and inefficient energy usage, particularly in battery-operated devices.

Innovation Solution

The implementation of improved user interfaces that utilize gaze and hand tracking, along with predefined configurations, to provide consistent input identifiers and differentiate between direct and indirect manipulation, while selectively providing gesture information based on user intent and object behavior parameters, enhancing interaction efficiency and reducing unnecessary inputs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If conventional input processing methods are used in three-dimensional environments, then the system can handle basic user inputs, but the interface becomes cumbersome and creates significant cognitive burden on users

Engineering Contradiction:
Improveuser interaction easeVSAvoidinterface complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent segments the input processing into distinct phases: gaze detection phase, hand configuration detection phase, and input confirmation phase. This segmentation allows the system to process different types of inputs (gaze-only, gaze-plus-hand, voice) independently, reducing the cognitive burden on users by making each interaction step clear and manageable.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary processing layer that sits between the raw input signals and the application logic. This intermediary layer disambiguates between different input types, manages the complexity of multiple input mechanisms, and presents a simplified interface to both users and applications, thereby reducing perceived complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If the system provides comprehensive support for multiple input mechanisms, then input versatility is improved, but processing time increases and energy is wasted

Engineering Contradiction:
Improveinput mechanism versatilityVSAvoidinput processing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent performs preliminary detection of gaze input and hand configuration before confirming the actual input. By pre-processing and pre-validating input signals, the system can quickly determine whether an input is intentional or accidental, reducing overall processing time while maintaining support for multiple input mechanisms.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements a tiered input recognition approach where the system processes inputs at different levels of detail based on context. For example, gaze-only inputs may be processed more quickly with fewer validation steps, while gaze-plus-hand inputs receive more thorough processing. This partial action approach reduces average processing time while maintaining versatility.

Inventive Principle:
Principle #16Partial or excessive action

3Reliability

If the system processes all gaze inputs as potential user inputs, then no input is missed, but false positives increase and user mistakes occur

Engineering Contradiction:
Improveinput recognition accuracyVSAvoidfalse positive inputs
Core Design Contradiction:
ReliabilityVSObject-generated harmful factors

Solution Approach 1:

The patent incorporates feedback mechanisms where the system provides visual or haptic confirmation when it detects a potential input. This feedback loop allows users to confirm or cancel unintended inputs, significantly reducing false positives while maintaining high reliability in recognizing genuine user intentions.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent implements a buffer or threshold mechanism that cushions against false positives by requiring multiple conditions to be met simultaneously (gaze direction, hand configuration, dwell time) before registering an input. This beforehand cushioning prevents accidental inputs from being misinterpreted as intentional user actions.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

4Reliability

If the system continuously monitors all input devices and user gestures, then input detection completeness is improved, but energy consumption increases significantly

Engineering Contradiction:
Improveinput detection completenessVSAvoiddevice energy consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent implements periodic sampling of input devices rather than continuous monitoring. The system activates full monitoring only when triggered by specific events (such as detected gaze patterns or hand configurations), and enters lower-power states between these events. This periodic action maintains input detection completeness while dramatically reducing average energy consumption.

Inventive Principle:
Principle #19Periodic action

Solution Approach 2:

The patent employs intelligent power management where the system itself determines when full input monitoring is necessary based on contextual cues. Rather than relying on external power management, the system autonomously adjusts its monitoring intensity based on detected user engagement levels, achieving both complete detection and energy efficiency.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12449947B2Devices, methods, and graphical user interfaces for processing inputs to a three-dimensional environment
Publication Date: 2025.10.21 APPLE INC
  • US12449947B2 patent drawing
  • US12449947B2 patent drawing
  • US12449947B2 patent drawing

AI summary

While a view of an environment is visible via a display generation component of a computer system, the computer system detects a gaze input directed to a first location, corresponding to a first user interface element, in the environment. In response to detecting the gaze input: if a user's hand is in a predefined configuration during the gaze input, the computer system: provides, to the first user interface element, information about the gaze input; and then, in response to detecting the gaze input moving to a different, second location in the environment while the user's hand is maintained in the predefined configuration, provides, to a second user interface element that corresponds to the second location, information about the gaze input. If the user's hand is not in the predefined configuration during the gaze input, the computer system forgoes providing, to the first user interface element, information about the gaze input.