Gaze-Based Image Augmentation for Automatic User Intent Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In augmented reality applications, existing technologies fail to automatically recognize user intentions and provide relevant information without explicit input, as they lack the ability to accurately determine the object of interest, its situation, and the task being performed based on user gaze analysis.
Innovation Solution
A method and apparatus that recognize the object of interest, situation, and task of a user by generating an image sequence from partial regions of an input image based on user gaze, using neural networks for object and task recognition, and determining relevant information to visually augment the image with additional information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If gaze-based recognition is used to automatically identify object of interest and user task, then automation level is improved, but recognition accuracy deteriorates due to lack of explicit user input
Solution Approach 1:
The patent divides the input image into multiple partial regions based on gaze trajectory, creating a segmented view that focuses on areas of user interest. This segmentation allows the system to process and recognize objects and tasks more accurately within the context of user attention, resolving the contradiction between automation and precision by analyzing only the relevant portions of the image that the user is actually observing.
Solution Approach 2:
The patent introduces temporal information by tracking gaze trajectory over time, transforming the static image analysis into a dynamic multi-dimensional approach. By analyzing gaze patterns across multiple time points and combining this with spatial information from the segmented image regions, the system achieves more accurate task recognition while maintaining automation, thus resolving the accuracy-automation tradeoff.
2Adaptability or versatility
If multiple neural networks are applied to recognize objects and tasks from segmented image regions, then recognition capability is improved, but processing time increases
Solution Approach 1:
The patent performs preliminary image segmentation based on gaze trajectory before applying neural networks for recognition. By pre-dividing the image into relevant regions and ordering the processing steps (segmentation first, then object recognition, then task recognition), the system optimizes the workflow to minimize redundant computations and reduce overall processing time while maintaining comprehensive recognition capability through multiple specialized neural networks.
3Measurement precision
If gaze trajectory and temporal information are encoded to enhance task prediction, then prediction accuracy is improved, but data complexity increases
Solution Approach 1:
The patent extracts and isolates specific temporal features from the gaze trajectory data, such as fixation duration, saccade velocity, and temporal patterns, separating these meaningful features from the raw complex gaze data. By extracting only the relevant temporal information needed for task prediction and representing it in a simplified format, the system achieves improved prediction accuracy while reducing data processing complexity through feature selection and transformation.
Data Source
AI summary
A method with image augmentation includes recognizing, based on a gaze of the user corresponding to the input image, any one or any combination of any two or more of an object of interest of a user, a situation of the object of interest, and a task of the user from partial regions of an input image determining relevant information indicating an intention of the user, based on any two or any other combination of the object of interest of the user, the situation of the object of interest, and the task of the user, and generating a visually augmented image by visually augmenting the input image based on the relevant information.


