Eye-Tracking Sensor Document Region Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for determining salient regions in documents are ineffective, particularly in tasks like title detection and table extraction, due to the lack of a universal standard for document structure, and struggle to replicate human perception for computer-based detection.

Innovation Solution

The use of attention tracking sensors, such as eye-tracking sensors, to detect and segment salient regions in documents by receiving sequences of measurements from human readings, determining and demarcating these regions, and outputting a displayable version of the document with identified salient areas, which can include titles, section headers, tables, and graphs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional NLP techniques are used to detect salient regions, then the method is simple and widely applicable, but the detection accuracy fails particularly in title detection and table extraction

Engineering Contradiction:
Improvedetection accuracyVSAvoidmethod complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces eye-tracking sensors as an intermediary to capture human reading behavior data. This mediator bridges the gap between human perception and computer detection, providing empirical data about which regions humans actually focus on when reading documents. The eye-tracking data serves as a bridge that translates subjective human attention into objective measurable signals that can guide automated detection.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements feedback by using eye-tracking measurements to continuously refine and adjust the detection of salient regions. The eye-tracking data provides real-time feedback about human attention patterns, allowing the system to learn from actual reading behavior and improve its detection accuracy iteratively. This feedback mechanism enables the system to adapt to different document types and structures dynamically.

Inventive Principle:
Principle #23Feedback

2Adaptability or versatility

If no universal standard for document structure is assumed, then the method can handle diverse document formats, but determining salient regions becomes difficult without structural guidance

Engineering Contradiction:
Improvedocument format flexibilityVSAvoidsalient region determination difficulty
Core Design Contradiction:
Adaptability or versatilityVSDifficulty of detecting and measuring

Solution Approach 1:

The system employs self-service by allowing eye-tracking data to automatically reveal the salient structure of each document without requiring pre-defined templates or universal standards. Each document's salient regions are identified based on the actual reading behavior observed for that specific document, enabling the system to adapt to diverse formats naturally. The data-driven approach allows the system to discover document structure autonomously.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent changes the fundamental parameter for detecting salient regions from structural features to human attention metrics. Instead of relying on document structure parameters like headings or formatting, the system uses eye-tracking parameters such as fixation duration, saccade patterns, and attention heatmaps. This parameter transformation enables the system to handle diverse document formats effectively by focusing on the universal aspect of human reading behavior rather than format-specific structures.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If human perception is directly translated to computer detection, then salient regions can be accurately identified, but the implementation requires complex sensor integration and data processing

Engineering Contradiction:
Improvesalient region identification accuracyVSAvoidsensor integration complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies universality by designing a framework where eye-tracking sensors can be integrated into multiple types of devices (computers, tablets, mobile devices) and applied to various document types. The system processes different input formats (video streams, coordinate data, fixation patterns) through a unified analysis framework, making the solution broadly applicable while maintaining accuracy. The multi-functional design allows the same core algorithm to handle different sensor configurations and document formats.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enables accurate detection and extraction of salient document regions, facilitating text summarization and key phrase identification, thereby improving the ability of computers to mimic human perception in document analysis.

Implementation Method 1

The sensor may include an eye-tracking sensor configured to detect a sequence of eye-gaze positions on the document as a function of time.

Methodology Applied
Scientific EffectEye-tracking:

Data Source

PatentUS12112563B2Method of detecting, segmenting and extracting salient regions in documents using attention tracking sensors
Publication Date: 2024.10.08 JPMORGAN CHASE BANK NA
  • US12112563B2 patent drawing
  • US12112563B2 patent drawing
  • US12112563B2 patent drawing

AI summary

A method and system for detecting, segmenting, and extracting salient regions in documents by using attention tracking sensors is provided. The method includes: receiving an image that corresponds to a document; receiving, from a sensor, a sequence of measurements that correspond to a human reading of the document; determining, based on the sequence of measurements, at least one region of the document as being a salient document region; demarcating the salient document region in an electronically displayable manner; and outputting a file that includes a displayable version of the document with the demarcated document region. The salient document region may include a title, a section header, and/or a table. The sensor may be an eye-tracking sensor that detects a sequence of eye-gaze positions on the document as a function of time.