Eye Gaze Tracking for Document Information Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current Natural Language Processing (NLP) based information extraction methods struggle to identify specific instances of information, such as the most recent place of employment from a document, as they lack the ability to understand the ordering or context, whereas humans can easily identify this information using eye gaze tracking.

Innovation Solution

A method utilizing eye gaze tracking to extract information from documents by receiving image data and sensor measurements to determine the regions of the document being read, identifying fixations and saccades, and extracting textual information based on calibration measurements and probability calculations to pinpoint the relevant information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If NLP based information extraction is used, then all instances of information can be extracted from documents, but the ability to identify specific instances (e.g., most recent place of employment) deteriorates

Engineering Contradiction:
Improvequantity of information extractedVSAvoidprecision of specific information identification
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent introduces eye gaze tracking as an intermediary mechanism between the document and the information extraction system. The eye tracker captures human gaze patterns which serve as a mediator to guide the extraction process, allowing the system to identify specific instances of information based on human attention patterns rather than relying solely on NLP algorithms.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system incorporates feedback loops where eye gaze data is continuously monitored and fed back into the information extraction process. The gaze patterns provide real-time feedback about where the human reader is focusing attention, enabling the system to dynamically adjust its extraction criteria to prioritize specific instances of information over others.

Inventive Principle:
Principle #23Feedback

2Loss of information

If standard NLP extraction algorithms are used, then comprehensive information can be retrieved, but the ability to understand document ordering and context deteriorates

Engineering Contradiction:
Improvecompleteness of information retrievalVSAvoidunderstanding of document ordering and context
Core Design Contradiction:
Loss of informationVSAdaptability or versatility

Solution Approach 1:

The patent adds a temporal dimension to the information extraction process by incorporating eye gaze data over time. Instead of merely extracting information based on static text patterns, the system analyzes the temporal sequence of gaze movements to understand document ordering and context, capturing how readers naturally progress through the document.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The system segments the document into multiple regions based on eye gaze patterns, dividing the continuous text flow into discrete areas of interest. This segmentation allows the system to process different portions of the document independently, maintaining both comprehensive information retrieval and an understanding of contextual relationships between segments.

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If human readers manually identify specific information, then precise selection of instances is achieved, but processing time and labor requirements increase

Engineering Contradiction:
Improveaccuracy of specific information identificationVSAvoidtime required for information extraction
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system leverages the human reader's own eye gaze patterns as the guiding mechanism for information extraction. By using the reader's natural attention patterns, the system eliminates the need for separate manual annotation or selection processes, allowing the extraction to occur automatically based on the reader's inherent focus patterns.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical processes of information selection with an automated eye tracking system. Instead of requiring humans to manually highlight or select specific information, the system uses optical sensors to capture gaze patterns and automatically identifies the relevant information, substituting manual operations with automated optical measurement and processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enables precise extraction of specific information from documents by leveraging human gaze patterns, improving the accuracy and efficiency of information retrieval beyond standard NLP methods.

Implementation Method 1

Eye trackers measure gaze position by shining infrared light into the eye, thereby creating reflections off the cornea

Methodology Applied
Scientific EffectReflection: Reflection

Data Source

PatentUS20240304012A1Method and system for extracting information from documents via eye gaze tracking
Publication Date: 2024.09.12 JPMORGAN CHASE BANK NA
  • US20240304012A1 patent drawing
  • US20240304012A1 patent drawing
  • US20240304012A1 patent drawing

AI summary

A method and system for using eye gaze tracking to extract information in textual form from documents is provided. The method includes: receiving an image that corresponds to a document; receiving, from an eye-tracking sensor configured to detect a sequence of eye-gaze positions on the document as a function of time, a sequence of measurements that correspond to a human reading of the document; determining, based on the received sequence of measurements, a region of the document that is being read by a human; and extracting the textual information that corresponds to the region.