Scanning System Page Number Alignment via OCR

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing scanning systems face difficulties in accurately aligning page numbers between document images and electronically managed page information, leading to inconvenient searching within documents that include covers, prefaces, and other unnumbered pages.

Innovation Solution

A scanning system and information processing program that includes a document reading unit, recognition unit, and difference elimination unit to detect and eliminate page number discrepancies by generating table-of-contents information, ensuring accurate page information alignment and facilitating easier document navigation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If page numbers are sequentially assigned to image data starting from the first page, then the scanning system can process documents automatically, but page number discrepancies occur when documents include unnumbered pages like covers and prefaces

Engineering Contradiction:
Improveautomatic page number assignmentVSAvoidpage number accuracy
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The system performs preliminary character recognition on the first page to identify table of contents information before assigning page numbers. By detecting chapter names and their corresponding page numbers in advance, the system can adjust the page number assignment to match the actual document structure, preventing discrepancies caused by unnumbered introductory pages.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses feedback from character recognition results to correct page number assignments. By comparing the sequentially assigned page numbers with the actual page numbers found in the table of contents through OCR, the system identifies and eliminates discrepancies, ensuring accurate page number alignment throughout the document.

Inventive Principle:
Principle #23Feedback

2Device complexity

If the scanning system assigns page numbers sequentially to all pages, then processing is simplified, but users experience inconvenience when searching for locations in the main text

Engineering Contradiction:
Improvepage number assignment processVSAvoiddocument navigation
Core Design Contradiction:
Device complexityVSEase of operation

Solution Approach 1:

The system performs preliminary character recognition on the first page to extract table of contents information before finalizing page number assignments. This allows the system to pre-determine the correct page number mapping for the main text, ensuring that when users navigate to a chapter in the table of contents, they are directed to the correct page without confusion from mismatched page numbers.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system introduces an intermediary correction process that mediates between the sequential page number assignment and the actual document structure. By using character recognition results as an intermediary to identify and correct page number discrepancies, the system maintains simple processing while ensuring accurate document navigation through the table of contents.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20240282138A1Scanning system and information processing program
Publication Date: 2024.08.22 SEIKO EPSON CORP
  • US20240282138A1 patent drawing
  • US20240282138A1 patent drawing
  • US20240282138A1 patent drawing

AI summary

A scanning system includes: a document reading unit configured to read a document and generate first image data of a plurality of pages read from the document; a recognition unit configured to recognize a character included in image data of the plurality of pages; a difference elimination unit configured to detect, based on a recognition result obtained by the recognition unit, a difference between first page information obtained from the recognition result and second page information sequentially assigned to image data; and an output unit configured to output second image data including table-of-contents information in which the difference is eliminated.