Scanning System Page Number Alignment via OCR
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing scanning systems face difficulties in accurately aligning page numbers between document images and electronically managed page information, leading to inconvenient searching within documents that include covers, prefaces, and other unnumbered pages.
Innovation Solution
A scanning system and information processing program that includes a document reading unit, recognition unit, and difference elimination unit to detect and eliminate page number discrepancies by generating table-of-contents information, ensuring accurate page information alignment and facilitating easier document navigation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If page numbers are sequentially assigned to image data starting from the first page, then the scanning system can process documents automatically, but page number discrepancies occur when documents include unnumbered pages like covers and prefaces
Solution Approach 1:
The system performs preliminary character recognition on the first page to identify table of contents information before assigning page numbers. By detecting chapter names and their corresponding page numbers in advance, the system can adjust the page number assignment to match the actual document structure, preventing discrepancies caused by unnumbered introductory pages.
Solution Approach 2:
The system uses feedback from character recognition results to correct page number assignments. By comparing the sequentially assigned page numbers with the actual page numbers found in the table of contents through OCR, the system identifies and eliminates discrepancies, ensuring accurate page number alignment throughout the document.
2Device complexity
If the scanning system assigns page numbers sequentially to all pages, then processing is simplified, but users experience inconvenience when searching for locations in the main text
Solution Approach 1:
The system performs preliminary character recognition on the first page to extract table of contents information before finalizing page number assignments. This allows the system to pre-determine the correct page number mapping for the main text, ensuring that when users navigate to a chapter in the table of contents, they are directed to the correct page without confusion from mismatched page numbers.
Solution Approach 2:
The system introduces an intermediary correction process that mediates between the sequential page number assignment and the actual document structure. By using character recognition results as an intermediary to identify and correct page number discrepancies, the system maintains simple processing while ensuring accurate document navigation through the table of contents.
Data Source
AI summary
A scanning system includes: a document reading unit configured to read a document and generate first image data of a plurality of pages read from the document; a recognition unit configured to recognize a character included in image data of the plurality of pages; a difference elimination unit configured to detect, based on a recognition result obtained by the recognition unit, a difference between first page information obtained from the recognition result and second page information sequentially assigned to image data; and an output unit configured to output second image data including table-of-contents information in which the difference is eliminated.


