Visual Attention Tracking With Gaze-Triggered Scene Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing visual content capture systems lack personalization and are computationally and energy-intensive, leading to inefficiencies in resource-constrained platforms, while eye-tracking systems suffer from measurement noise and transient gaze drift, making them inadequate for reliable visual attention tracking.
Innovation Solution
A system that combines eye movement and gaze direction data with visual content analysis using a unified neural network to accurately detect and record scenes of interest, triggered by saccade-smooth pursuit transitions, reducing computational and energy consumption by focusing analysis on regions of interest.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If visual content analysis is used to determine content of interest, then content identification accuracy is improved, but computational complexity and energy consumption increase
Solution Approach 1:
The system performs preliminary eye tracking to identify potential regions of interest before conducting full visual content analysis. This preliminary action filters the scene into candidate regions, so that computationally intensive visual analysis is only applied to these limited regions rather than the entire scene, reducing overall energy consumption while maintaining accurate content identification.
2Adaptability or versatility
If eye tracking is used to identify content of interest, then personalization is improved, but measurement accuracy deteriorates due to noise and gaze drift
Solution Approach 1:
The system merges eye tracking data with visual content analysis results to determine content of interest. By combining these two approaches, the system leverages the personalization capability of eye tracking while compensating for its measurement inaccuracies through the objective visual content analysis, achieving both personalized and accurate content identification.
Solution Approach 2:
The system uses feedback from visual content analysis to correct and refine eye tracking measurements. When visual analysis identifies salient objects, this information feeds back to adjust the interpretation of gaze data, filtering out noise and transient drift to improve the accuracy of determining what content the user is actually interested in.
3Reliability
If full scene analysis is performed continuously, then content detection reliability is improved, but energy consumption increases
Solution Approach 1:
The system segments the visual scene into multiple regions and uses eye tracking to identify which regions should be analyzed in detail. Instead of continuously analyzing the entire scene, the system divides the workload by focusing computational resources only on the regions where the user's gaze is directed, maintaining detection reliability while significantly reducing energy consumption.
Solution Approach 2:
The system performs visual content analysis periodically based on eye tracking triggers rather than continuously. When the eye tracking system detects sustained gaze on a region, this triggers a periodic visual analysis of that region. This periodic action maintains reliable content detection while avoiding the continuous energy expenditure of full-scene analysis.
Data Source
AI summary
A method for detecting content of interest to a user includes obtaining a first data stream indicative of eye movement and/or gaze direction of the user as the user is viewing a scene in a field of view of the user, obtaining a second data stream indicative of visual content in the field of view of the user, determining, based on the first data stream and the second data stream, that content of interest to the user is present in the scene in the field of view of the user, and, in response to determining that content of interest to the user is present in the scene in the field of view of the user, triggering, with the processor, an operation to be performed with respect to the scene in the field of view of the user.


