Raster Image Table of Contents Generation Without OCR
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing automated table of contents tools for documents require conversion of scanned or print-ready text into electronic text format using OCR, which is resource-intensive, inaccurate, and loses graphic elements.
Innovation Solution
Methods and devices that automatically identify and rank topical items in raster images based on graphical features, associate them with topics and subtopics, and create a cropped-image index without the need for OCR, allowing for dynamic generation of tables of contents directly from raster images.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If OCR processing is used to convert scanned text to electronic text format, then text recognition capability is improved, but resource consumption increases and accuracy is limited
Solution Approach 1:
The patent extracts and processes only the essential graphical features (pixel patterns, shapes, colors) needed for table of contents generation, rather than converting the entire raster image to electronic text. This selective extraction approach reduces computational resources while maintaining accuracy for the specific task of identifying topical items.
Solution Approach 2:
The patent replaces the OCR mechanical system (which converts images to text characters) with a graphical feature recognition system that directly analyzes pixel patterns, shapes, and colors in raster images. This substitution eliminates the need for character-by-character recognition, reducing resource consumption while preserving accuracy for topical item identification.
2Adaptability or versatility
If OCR processing is used to convert scanned text to electronic text format, then text becomes searchable and editable, but graphic elements are lost
Solution Approach 1:
The patent creates a graphical copy or representation of the raster content by analyzing pixel patterns, shapes, and colors directly from the scanned image. This copying approach preserves the visual and graphical elements of the original document while extracting sufficient information to identify topical items and generate a table of contents, avoiding the information loss inherent in OCR conversion.
3Productivity
If automated table of contents tools analyze electronic text representation, then table of contents generation is efficient, but the process requires cumbersome OCR conversion first
Solution Approach 1:
The patent inverts the conventional approach by not converting raster images to electronic text first. Instead, it directly analyzes the raster image content through graphical feature recognition to identify topical items and generate the table of contents. This inversion eliminates the cumbersome OCR conversion step while maintaining generation efficiency through direct image analysis.
Data Source
AI summary
Methods and devices receive a document comprising raster images, using an optical scanner. These methods and devices automatically identify topical items within the raster images based on raster content in the raster images, using a processor. Further, these methods and devices automatically associate the topical items with topics in the document based on previously established rules for identifying topical sections, and automatically crop the topical items from the raster images to produce cropped portions of the raster images, using the processor. These methods and devices then automatically create an index for the document by combining the cropped portions of the raster images organized by the topics, using the processor, and output the index from the processor.


