Scanned Page Data Division Using OCR Layout Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing image forming apparatuses struggle with efficiently processing and aggregating multiple types of form documents, particularly in dividing scanned page data into manageable units while maintaining accuracy and efficiency.
Innovation Solution
An information processing apparatus and method that employs optical character recognition (OCR) for character and layout analysis, combined with machine learning algorithms, to classify and divide scanned page data into page units based on page ordering rules, utilizing a rule setting unit, a rule order unit, and a machine learning order unit to enhance accuracy and automation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual rule setting is used for dividing scanned page data, then processing accuracy can be maintained, but user intervention time and operational complexity increase
Solution Approach 1:
The system performs automatic OCR recognition and page division without requiring user intervention. The apparatus independently extracts text, determines page boundaries, and divides scanned data into separate pages, enabling the system to serve itself rather than requiring manual operation.
Solution Approach 2:
Manual rule setting and page division operations are replaced by automated OCR recognition and algorithmic processing. The mechanical/manual process of reviewing and setting division rules is substituted with optical character recognition and automatic page boundary detection systems.
2Productivity
If automated division is implemented without OCR, then processing speed increases, but division accuracy and reliability deteriorate
Solution Approach 1:
OCR recognition is performed in advance before page division. The system pre-processes the scanned data by recognizing text and extracting layout information, which then serves as the basis for accurate page boundary detection and division, ensuring both speed and reliability.
Solution Approach 2:
The OCR recognition and page division processes are integrated into a continuous automated workflow. Once OCR is initiated, the system continuously processes the recognized text to determine page boundaries and execute division without interruption, maintaining both high speed and reliability throughout the process.
3Measurement precision
If complex page ordering rules are applied, then classification precision improves, but device complexity and processing time increase
Solution Approach 1:
The page division process is segmented into distinct automated stages: OCR text recognition, layout analysis, page boundary detection, and final division execution. This segmentation allows complex classification rules to be applied systematically without overwhelming system complexity.
Data Source
AI summary
Provided is an information processing apparatus that divides a plurality of scanned page data with high accuracy. The OCR unit performs optical character recognition for in a page for each of the plurality of page data. The rule order unit classifies each of the plurality of page data based on a page ordering rule according to the characters and the layout recognized by the OCR unit, and it divides the plurality of page data into page units.


