Eye Tracking Context for AR Generative AI Intent Inference
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional smart glasses and AR devices struggle to accurately interpret user intent and provide contextually relevant responses due to the lack of nuanced information from traditional input methods, leading to inefficiencies in data processing and user interaction.
Innovation Solution
An AR device equipped with eye tracking technology processes eye gaze information and contextual data using a generative machine learning model to generate relevant outputs, reducing the need for multiple queries and optimizing computational resources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional input methods are used for user interaction, then device complexity is reduced, but user interaction efficiency deteriorates due to lack of nuanced information
Solution Approach 1:
The system segments user interaction into multiple data dimensions: eye gaze information, head pose data, and explicit input. Each segment captures different aspects of user intent, with eye tracking providing nuanced attention information that complements traditional input methods, thereby improving overall interaction efficiency without losing informational depth
Solution Approach 2:
Eye tracking data serves as an intermediary that bridges the gap between traditional input methods and the system's understanding of user intent. By capturing subtle attention patterns and focus areas, the eye tracking intermediary provides nuanced information that enhances the quality of interaction without requiring complex additional hardware interfaces
2Measurement precision
If multiple queries are processed to achieve desired results, then information accuracy is improved, but time consumption increases
Solution Approach 1:
The system performs preliminary analysis by continuously monitoring eye gaze patterns and head pose to pre-identify areas of user interest and potential query intent. This preliminary action allows the system to anticipate user needs and prepare relevant information in advance, reducing the number of iterative queries needed and thereby decreasing time consumption while maintaining information accuracy
Solution Approach 2:
The system uses eye tracking feedback to continuously refine its understanding of user intent. By analyzing gaze duration, fixation patterns, and saccade movements, the system receives real-time feedback about what information the user is seeking, allowing it to adjust its responses and reduce the number of back-and-forth queries needed to achieve accurate information delivery
3Reliability
If comprehensive data processing is performed to understand user intent, then response relevance is improved, but computational load increases
Solution Approach 1:
The system applies partial processing by focusing computational resources only on the most relevant aspects of eye tracking data, such as fixation duration and gaze position, rather than processing every possible eye movement parameter. This selective approach maintains response relevance by capturing the most informative signals while reducing overall computational load and energy consumption
Solution Approach 2:
The system applies different levels of processing intensity to different types of data based on their informational value. Eye gaze data receiving high computational attention for intent inference, while other auxiliary data receives lighter processing. This local quality differentiation ensures response relevance is maintained for critical parameters while minimizing unnecessary computational expenditure on less informative data
Data Source
AI summary
Examples relate to systems and methods for enhancing generative AI outputs using eye tracking data. An eye tracking system accesses eye gaze information associated with a field of view of a head-wearable apparatus and generates contextual information associated with the field of view of the head-wearable apparatus based on the eye gaze information. The eye tracking system processes, by a generative machine learning model, the contextual information and at least one image of the field of view of the head-wearable apparatus to generate an output and presents on a display of the head-wearable apparatus the output generated by the generative machine learning model.


