Gaze-Based GUI Interaction for Lower-Cognitive-Burden AR/VR Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for interacting with graphical user interfaces in augmented and virtual reality environments are cumbersome, inefficient, and place a significant cognitive burden on users, often requiring multiple inputs and providing insufficient feedback, leading to errors and energy wastage, particularly in battery-operated devices.

Innovation Solution

Implementing computer systems with gaze-tracking sensors and display generation components that allow for intuitive interaction through gaze detection, enabling operations based on user gaze direction and reducing the need for manual inputs by displaying and hiding interface objects accordingly.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If conventional input devices (cameras, controllers, joysticks, touch-sensitive surfaces) are used to interact with virtual/augmented reality environments, then the system can process user inputs, but the interaction becomes cumbersome, inefficient, and creates significant cognitive burden on users

Engineering Contradiction:
Improveuser interaction efficiencyVSAvoidinterface complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent replaces mechanical input devices (controllers, joysticks, touch surfaces) with a gaze-tracking system that uses optical detection of eye movements and pupillary responses. This substitution eliminates the need for physical manipulation of input devices, allowing users to interact with virtual environments through natural eye movements and physiological responses, thereby improving ease of operation while reducing cognitive burden

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system automatically detects and interprets user intent through gaze direction and pupillary dilation without requiring explicit commands or manual inputs. The computer system serves itself by monitoring its own display output and correlating it with measured pupillary responses to determine user reactions and preferences, eliminating the need for complex user input mechanisms

Inventive Principle:
Principle #25Self-service

2Productivity

If multiple inputs are required to achieve desired outcomes in augmented reality environments, then the system can perform complex operations, but the interaction takes longer than necessary and wastes energy

Engineering Contradiction:
Improveoperation speedVSAvoiddevice energy consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The system continuously monitors gaze direction and pupillary responses in advance to predict user intent and preferences before explicit actions are required. By pre-processing gaze data and correlating it with displayed content, the system can anticipate user reactions and prepare appropriate responses, reducing the time and energy needed for subsequent interactions

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system establishes a closed-loop feedback mechanism where pupillary responses are measured and correlated with displayed content to determine user reactions. This feedback is used to automatically adjust the user interface, select appropriate content, or modify system behavior without requiring additional user inputs, thereby increasing productivity while minimizing energy consumption

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12481357B2Devices, methods, for interacting with graphical user interfaces
Publication Date: 2025.11.25 APPLE INC
  • US12481357B2 patent drawing
  • US12481357B2 patent drawing
  • US12481357B2 patent drawing

AI summary

In some embodiments, the present disclosure includes techniques and user interfaces for interacting with graphical user interfaces using gaze. In some embodiments, the present disclosure includes techniques and user interfaces for repositioning virtual objects. In some embodiments, the present disclosure includes techniques and user interfaces for transitioning modes of a camera capture user interface.