Eye Gaze Trained Layout Interpretation Model
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing text interpretation technologies, such as OCR and TTS, fail to accurately process visual text that deviates from standard layouts, leading to incorrect outputs and difficulties in understanding non-conventional text formats, particularly in advertisements, signs, and other visually complex documents.
Innovation Solution
A system that utilizes eye gaze data to train machine learning models to interpret text layouts, allowing for the determination and correction of non-standard text arrangements, enabling accurate processing and output of text in various formats across different contexts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If OCR and TTS techniques are used to process visual text, then text can be converted to editable format or audible reading, but the techniques fail when text does not follow default layout or order, producing incorrect outputs
Solution Approach 1:
The system performs preliminary action by training the layout interpretation model in advance using gaze data that captures natural reading patterns. This pre-trained model then automatically determines text layout order before OCR processing, ensuring accurate text extraction even from non-standard layouts without requiring manual intervention during actual text processing.
Solution Approach 2:
The layout interpretation model serves as an intermediary between the visual text input and the OCR/TTS processing. It analyzes eye gaze patterns to determine the correct reading order and layout structure, then provides this structured information to the OCR engine, which uses it to accurately process and output the text in the correct sequence.
2Measurement precision
If eye gaze data is collected and used to train layout interpretation models, then text layout understanding improves, but data collection and model training require additional time and resources
Solution Approach 1:
The system employs self-service by automatically collecting gaze data during normal reading activities without requiring separate data collection sessions. The model trains itself using this naturally accumulated data, and the entire process occurs in the background during regular device usage, eliminating the need for dedicated training time and making the system progressively smarter over time without user intervention.
Data Source
AI summary
Gaze data collected from eye gaze tracking performed while training text was read may be used to train at least one layout interpretation model. In this way, the at least one layout interpretation model may be trained to determine current text that includes words arranged according to a layout, process the current text with the at least one layout interpretation model to determine the layout, and output the current text with the words arranged according to the layout.


