3D Interface Invocation Using Gaze and Hand Pose Cues
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for interacting with system user interfaces in augmented and virtual reality environments are cumbersome, inefficient, and place a significant cognitive burden on users, requiring extensive input and providing insufficient feedback, leading to energy waste in battery-operated devices.
Innovation Solution
The system uses attention-based methods to invoke controls and interfaces, such as directing attention to a hand location, detecting hand orientations and gestures, and utilizing torso direction for ergonomic menu display, to reduce the number and nature of user inputs, and provides efficient visual and tactile feedback.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If conventional input methods are used to invoke system user interfaces in augmented reality environments, then the system can process user inputs, but the interaction becomes cumbersome and requires extensive input from the user
Solution Approach 1:
The system performs preliminary action by automatically detecting and invoking system user interfaces based on user gaze direction and hand pose before the user explicitly requests them. The interface anticipation module proactively determines when to display interfaces by monitoring user attention and hand positioning, eliminating the need for users to manually invoke interfaces through extensive input sequences.
2Reliability
If multiple input steps are required to display system user interfaces, then the system can ensure accurate user intent, but the cognitive burden on the user increases significantly
Solution Approach 1:
The system implements feedback by providing real-time visual indicators that show the detected user gaze direction and hand pose, and by displaying anticipated interfaces that users can confirm or reject. This feedback loop allows the system to accurately interpret user intent while keeping the interaction simple, as users can see what the system understands and correct it if needed without undergoing complex multi-step input processes.
3Stability of the object's composition
If traditional user interface invocation methods are used, then the system maintains consistent interface display, but energy consumption increases due to prolonged processing and display time
Solution Approach 1:
The system applies periodic action by continuously monitoring user gaze and hand pose at regular intervals to detect when interface invocation is needed, rather than continuously displaying interfaces or processing potential inputs. The interface anticipation module periodically assesses user intent based on sustained gaze direction and hand positioning, enabling energy-efficient processing while maintaining consistent and appropriate interface display.
4Measurement precision
If extensive user input is required to interact with virtual objects, then the system can ensure precise control, but the interaction becomes tedious and error-prone
Solution Approach 1:
The system implements self-service by automatically interpreting user intent through gaze and hand pose detection, and by providing contextual assistance through anticipated interfaces that reduce the need for extensive manual input. The system serves itself by monitoring its own state and user behavior to proactively present relevant controls and options, allowing users to interact with virtual objects more naturally while maintaining precise control through the anticipated interface elements.
Data Source
AI summary
While a view of an environment is visible via one or more display generation components of a computer system, and while the view of the environment includes a respective object that moves as a hand of a user moves, the computer system detects, via one or more input devices, a respective input. In response to detecting the respective input: in accordance with a determination that first criteria are met, the first criteria including a requirement that the hand of the user is holding a controller, the computer system displays, via the one or more display generation components, a first user interface object at a first location relative to the respective object; and, in accordance with a determination that the first criteria are not met, the computer system forgoes displaying the first user interface object at the first location.


