User Intention Prediction via Gaze Sequence Visual Encoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing augmented reality (AR) applications lack effective methods to predict user intentions without explicit input, relying on gaze analysis to infer user needs, which is challenging due to varying cognitive tasks and visual attention patterns.
Innovation Solution
A method and apparatus that acquire a gaze sequence, generate a coded image by visually encoding temporal information, and predict user intentions using feature vectors extracted from both the input image and coded image, incorporating gaze trajectory, velocity, fixation duration, and recurrence patterns.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If gaze analysis is used to predict user intentions in AR applications, then user interaction accuracy is improved, but system complexity increases due to varying cognitive tasks and visual attention patterns
Solution Approach 1:
The system segments the complex gaze analysis task into multiple processing stages: acquiring raw gaze sequences, extracting temporal information features, encoding gaze patterns into structured representations, and finally predicting user intentions. This segmentation allows each module to handle specific aspects of the problem independently, reducing overall system complexity while maintaining prediction accuracy
Solution Approach 2:
The patent introduces an intermediary encoding module that transforms complex temporal gaze information into standardized visual representations. This intermediary representation serves as a bridge between raw gaze data and intention prediction algorithms, simplifying the relationship between input data and output predictions while preserving critical temporal patterns
2Loss of information
If temporal information from gaze sequences is encoded into visual representations, then information retention is improved, but data processing complexity increases
Solution Approach 1:
The system creates visual copies of temporal gaze information by encoding gaze sequences into image-like representations. These visual copies preserve the essential temporal patterns and spatial trajectories of gaze movements while being compatible with standard image processing algorithms, thereby retaining information without requiring complex temporal analysis pipelines
Solution Approach 2:
The encoding process transforms temporal parameters (gaze coordinates over time) into spatial parameters (visual patterns in encoded images). By changing the parameter domain from temporal to spatial, the system retains gaze information in a format that can be processed using efficient spatial algorithms rather than complex temporal analysis
3Measurement precision
If multiple gaze features (trajectory, velocity, fixation duration) are incorporated into prediction, then prediction accuracy is improved, but computational requirements increase
Solution Approach 1:
The patent merges multiple gaze features (trajectory, velocity, fixation duration, recurrence patterns) into a unified visual encoding representation. By combining these diverse features into a single integrated visual structure, the system maintains comprehensive information for accurate prediction while avoiding the computational overhead of processing separate feature streams independently
Data Source
AI summary
A method and apparatus for predicting an intention acquires a gaze sequence of a user, acquires an input image corresponding to the gaze sequence, generates a coded image by visually encoding temporal information included in the gaze sequence to the input image, and predicts an intention of the user corresponding to the gaze sequence based on the input image and the coded image.


