Object Recognition via Eye Fixation Correlation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for automatic object recognition in dynamic environments, such as those using FixTag and SLAM algorithms, require accurate pre-mapping of environments and are computationally intensive, limiting their applicability in novel or changing conditions.

Innovation Solution

An object recognition processing apparatus that determines fixation locations in image sequences, correlates them with field of view images, and classifies features using stored measurement values, reducing the need for prior knowledge of the environment and operator intervention, and allowing for efficient processing of naturally occurring image sequences.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If FixTag or SLAM algorithms are used for automatic object recognition, then object recognition accuracy is improved, but the computational complexity and requirement for pre-mapped environments increase

Engineering Contradiction:
Improveobject recognition accuracyVSAvoidcomputational complexity and environmental requirements
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by capturing eye fixation data before object recognition processing. By determining fixation locations in advance and using them to guide subsequent image processing, the system reduces the computational burden of analyzing entire images while maintaining recognition accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system segments the visual processing task by dividing it into distinct components: eye fixation determination, fixation location correlation with image regions, and targeted object recognition. This segmentation allows the system to focus computational resources only on relevant regions identified by eye fixation, rather than processing entire images.

Inventive Principle:
Principle #1Segmentation

2Reliability

If comprehensive environmental mapping is performed before analysis, then object recognition reliability is improved, but the system becomes brittle to scene layout changes and requires careful calibration

Engineering Contradiction:
Improveobject recognition reliabilityVSAvoidadaptability to dynamic environments
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system implements dynamics by allowing the environmental model to be constructed and updated during the analysis process rather than requiring a static pre-mapped environment. The correlation between fixation locations and image regions adapts to scene changes, enabling the system to handle dynamic environments without brittle dependencies on fixed calibrations.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system performs self-service by automatically constructing environmental models and correlating fixation data with image content without requiring external calibration or pre-mapping. The algorithm independently establishes the relationship between eye fixation locations and corresponding image regions, making the system self-adapting to new environments.

Inventive Principle:
Principle #25Self-service

3Loss of information

If all data from the entire input video stream is presented for classification, then completeness of analysis is improved, but the amount of data requiring operator classification and processing time increase significantly

Engineering Contradiction:
Improvecompleteness of analysisVSAvoidprocessing time and operator workload
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system extracts only the most relevant information for classification by identifying eye fixation locations and correlating them with corresponding image regions. Instead of presenting all video data to operators, the system extracts and presents only the specific regions that received visual attention, significantly reducing data volume while maintaining analytical completeness.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system introduces an intermediary layer of eye fixation analysis between raw video data and operator classification. This intermediary process filters and prioritizes data by using fixation locations to identify which regions warrant operator attention, thereby reducing the burden on operators while ensuring comprehensive analysis of relevant content.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS9785835B2Methods for assisting with object recognition in image sequences and devices thereof
Publication Date: 2017.10.10 ROCHESTER INSTITUTE OF TECHNOLOGY
  • US9785835B2 patent drawing
  • US9785835B2 patent drawing
  • US9785835B2 patent drawing

AI summary

A method, non-transitory computer readable medium, and apparatus that assist with object recognition includes determining when at least one eye of an observer fixates on a location in one or more of a sequence of fixation tracking images. The determined fixation location in the one or more of the sequence of images is correlated to a corresponding one of one or more sequence of field of view images. At least the determined fixation location in each of the correlated sequence of field of view images is classified based on at least one of a classification input or a measurement and comparison of one or more features of the determined fixation location in each of the correlated sequence of field of view images against one or more stored measurement feature values. The determined classification of the fixation location in each of the correlated sequence of field of view images is output.