3D User Interface Interaction Using Gaze and Hand Tracking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for interacting with virtual, augmented, and extended reality environments are cumbersome, inefficient, and place a significant cognitive burden on users, requiring multiple inputs and lacking sufficient feedback, leading to energy waste and complex manipulation of virtual objects.
Innovation Solution
A computer system with integrated gaze and hand-tracking inputs, allowing for intuitive interaction by detecting gaze and hand movements to activate user interface objects, switch between interface groups, and adjust display properties, enhancing user interface efficiency and reducing the need for multiple inputs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If conventional input devices (cameras, controllers, joysticks, touch-sensitive surfaces) are used to interact with virtual/augmented/extended reality environments, then the system can detect user inputs, but the interaction becomes cumbersome, inefficient, and complex
Solution Approach 1:
The patent combines multiple input modalities (gaze tracking, hand tracking, voice recognition) into a unified interaction system. The system integrates these different input methods to work together seamlessly, allowing users to interact with virtual environments using natural human behaviors rather than separate devices for each function.
Solution Approach 2:
The system creates a universal input mechanism that can handle multiple types of user inputs through a single integrated framework. This multi-functional approach allows the same system to process gaze data, hand gestures, and voice commands without requiring separate specialized devices for each input type.
2Productivity
If multiple inputs are required to achieve a desired outcome in virtual/augmented/extended reality environments, then the system can perform complex operations, but the cognitive burden on users increases and interaction time is extended
Solution Approach 1:
The system performs preliminary processing of user inputs by continuously tracking gaze and hand movements in advance of actual interaction needs. This allows the system to pre-process and prepare for upcoming user actions, reducing the time required to execute commands and decreasing overall interaction time.
Solution Approach 2:
The system provides real-time feedback to users about their inputs and system responses, creating a more efficient interaction loop. This feedback mechanism helps users understand the connection between their inputs and device responses, reducing cognitive burden and enabling faster, more intuitive interactions.
3Loss of energy
If conventional user interfaces require a series of inputs to achieve a desired outcome, then the system can perform detailed operations, but energy is wasted and battery life is reduced
Solution Approach 1:
The system uses periodic action by activating intensive processing only when needed based on detected user inputs. Instead of continuously processing all input modalities at full capacity, the system periodically engages processing resources in response to actual user interactions, thereby reducing energy consumption while maintaining ease of operation.
4Adaptability or versatility
If gaze tracking and hand tracking are integrated to detect user inputs, then the system can provide more intuitive interaction, but the processing complexity and computational requirements increase
Solution Approach 1:
The patent segments the processing of different input modalities into separate modules. Gaze tracking, hand tracking, and voice recognition are processed as distinct segments that can be independently optimized and managed. This segmentation allows the system to handle multiple input types flexibly while managing processing complexity through modular architecture.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A computer system displays a first view of a three-dimensional environment. While displaying the first view, the computer system detects a gaze input directed to a first position in the three-dimensional environment corresponding to a location of a user's hand in a physical environment in conjunction with detecting a movement of the user's hand in the physical environment that meets preset criteria. In response, the computer system displays a plurality of user interface objects at respective second positions that are away from the first position in the three-dimensional environment corresponding to the location of the user's hand in the physical environment, wherein a respective user interface object, when activated, causes display of a corresponding computer-generated experience in the three-dimensional environment.