Object Recognition via Eye Fixation Correlation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for automatic object recognition in dynamic environments, such as those using FixTag and SLAM algorithms, require accurate pre-mapping of environments and are computationally intensive, limiting their applicability in novel or changing conditions.
Innovation Solution
An object recognition processing apparatus that determines fixation locations in image sequences, correlates them with field of view images, and classifies features using stored measurement values, reducing the need for prior knowledge of the environment and operator intervention, and allowing for efficient processing of naturally occurring image sequences.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If FixTag or SLAM algorithms are used for automatic object recognition, then object recognition accuracy is improved, but the computational complexity and requirement for pre-mapped environments increase
Solution Approach 1:
The system performs preliminary actions by capturing eye fixation data before object recognition processing. By determining fixation locations in advance and using them to guide subsequent image processing, the system reduces the computational burden of analyzing entire images while maintaining recognition accuracy.
Solution Approach 2:
The system segments the visual processing task by dividing it into distinct components: eye fixation determination, fixation location correlation with image regions, and targeted object recognition. This segmentation allows the system to focus computational resources only on relevant regions identified by eye fixation, rather than processing entire images.
2Reliability
If comprehensive environmental mapping is performed before analysis, then object recognition reliability is improved, but the system becomes brittle to scene layout changes and requires careful calibration
Solution Approach 1:
The system implements dynamics by allowing the environmental model to be constructed and updated during the analysis process rather than requiring a static pre-mapped environment. The correlation between fixation locations and image regions adapts to scene changes, enabling the system to handle dynamic environments without brittle dependencies on fixed calibrations.
Solution Approach 2:
The system performs self-service by automatically constructing environmental models and correlating fixation data with image content without requiring external calibration or pre-mapping. The algorithm independently establishes the relationship between eye fixation locations and corresponding image regions, making the system self-adapting to new environments.
3Loss of information
If all data from the entire input video stream is presented for classification, then completeness of analysis is improved, but the amount of data requiring operator classification and processing time increase significantly
Solution Approach 1:
The system extracts only the most relevant information for classification by identifying eye fixation locations and correlating them with corresponding image regions. Instead of presenting all video data to operators, the system extracts and presents only the specific regions that received visual attention, significantly reducing data volume while maintaining analytical completeness.
Solution Approach 2:
The system introduces an intermediary layer of eye fixation analysis between raw video data and operator classification. This intermediary process filters and prioritizes data by using fixation locations to identify which regions warrant operator attention, thereby reducing the burden on operators while ensuring comprehensive analysis of relevant content.
Data Source
AI summary
A method, non-transitory computer readable medium, and apparatus that assist with object recognition includes determining when at least one eye of an observer fixates on a location in one or more of a sequence of fixation tracking images. The determined fixation location in the one or more of the sequence of images is correlated to a corresponding one of one or more sequence of field of view images. At least the determined fixation location in each of the correlated sequence of field of view images is classified based on at least one of a classification input or a measurement and comparison of one or more features of the determined fixation location in each of the correlated sequence of field of view images against one or more stored measurement feature values. The determined classification of the fixation location in each of the correlated sequence of field of view images is output.


