Gaze-Based Intent Data for More Accurate LLM Responses
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current large language models (LLMs) struggle to accurately understand user intent due to reliance on textual, audio, or visual prompts, leading to inefficient resource consumption and suboptimal interactions.
Innovation Solution
A user device supplements LLM prompts with biometric-based intent data using gaze data, dwell time, and pupil behavior to generate intent data, which is then used by the LLM to provide more aligned responses.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If LLMs rely on textual, audio, or visual prompts to understand user intent, then the system can process user inputs, but the understanding of user intent becomes inaccurate and resource consumption increases
Solution Approach 1:
The patent introduces biometric data (gaze tracking, pupil behavior, dwell time) as an intermediary layer between the user and the LLM. These biometric signals serve as a mediator that conveys user intent more accurately than traditional prompts, reducing the computational burden on the LLM while improving intent understanding accuracy.
Solution Approach 2:
The system segments the user intent detection process into multiple components: gaze data collection, dwell time measurement, pupil behavior analysis, and intent data generation. This segmentation allows each component to be optimized independently, improving overall accuracy while managing computational resources efficiently.
2Adaptability or versatility
If LLMs process only explicit textual prompts, then the system maintains simplicity in operation, but the ability to capture nuanced user intent deteriorates
Solution Approach 1:
The system enables the user device to automatically collect and process biometric data without requiring additional user actions. The gaze tracking, dwell time measurement, and pupil behavior analysis occur passively as the user naturally interacts with the interface, eliminating the need for users to manually express their intent through complex prompts.
Solution Approach 2:
The patent adds a new dimension to user input by incorporating biometric data from physiological responses. Instead of relying solely on textual, audio, or visual prompts, the system now processes multi-dimensional data including gaze coordinates, dwell time duration, and pupil diameter changes, significantly enhancing intent interpretation capability.
3Productivity
If LLMs require explicit and precise prompts from users, then the system ensures accurate input interpretation, but the interaction efficiency and user accessibility decrease
Solution Approach 1:
The system performs preliminary analysis of biometric data before the user even formulates a prompt. By continuously monitoring gaze patterns, dwell time, and pupil behavior, the system pre-processes intent signals, allowing it to anticipate user needs and provide responses before explicit prompts are required, thereby dramatically improving interaction efficiency.
Solution Approach 2:
The biometric-based intent detection system serves multiple functions simultaneously: it tracks user attention, determines intent, validates prompts, and provides contextual information to the LLM. This multi-functionality eliminates the need for users to manually articulate their intent through complex prompts, making the system accessible to users with various communication abilities.
Data Source
AI summary
A user device may receive a user interface that includes content, and may provide the user interface for display to a user of the user device. The user device may receive a user interaction with the user interface, and may calculate, based on the user interaction, gaze data identifying a gaze of the user, a dwell time of the gaze, and an eye behavior of the user relative to the content. The user device may generate intent data based on the gaze data, and may provide the intent data and one or more prompts to a large language model (LLM) system. The user device may receive one or more responses from the LLM system based on providing the intent data and the one or more prompts to the LLM system.


