Adaptive Text-Block Estimation for Accurate Document OCR
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing techniques for extracting character strings from scanned documents, such as business forms, require users to manually set rules for determining the number of rows and character string positions, which is time-consuming and burdensome.
Innovation Solution
An image processing apparatus that estimates text blocks based on pre-registered document rules, allowing users to modify and derive conditions for accurate extraction without heavy manual effort, using a multi-function peripheral (MFP) with integrated OCR and document analysis capabilities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual rule setting for determining text block positions and row numbers is implemented, then extraction accuracy is improved, but user burden and time consumption increase
Solution Approach 1:
The system performs preliminary actions by automatically generating extraction rules based on registered document templates before actual character string extraction. The rule generation unit creates initial extraction rules by analyzing the layout and structure of registered documents, so that when a new document is processed, the extraction can proceed with pre-established guidelines rather than requiring manual rule creation from scratch.
Solution Approach 2:
The system enables self-service by allowing the rule generation unit to automatically create extraction rules without user intervention. The apparatus uses its own registered document templates to generate extraction rules autonomously, and the determination unit automatically applies these rules to extract character strings, reducing the need for manual rule setting while maintaining high extraction accuracy.
2Productivity
If fixed extraction rules are used for all documents of the same type, then processing efficiency is improved, but adaptability to layout variations deteriorates
Solution Approach 1:
The system implements dynamics by making extraction rules flexible rather than fixed. The determination unit dynamically adjusts the number of rows and text block positions based on the actual document content and layout, even within the same document type. This allows the system to maintain high processing efficiency while adapting to variations in document layouts through automated rule application and adjustment.
Solution Approach 2:
The system applies parameter changes by modifying extraction rule parameters such as the number of rows and text block positions based on the specific document being processed. The determination unit changes these parameters dynamically according to the actual document layout and content, enabling the system to handle layout variations while maintaining efficient automated processing.
3Reliability
If comprehensive rule setting for each document type is performed, then extraction reliability is improved, but system complexity increases
Solution Approach 1:
The system achieves universality by creating a multi-functional rule management architecture. The rule generation unit serves multiple purposes: it generates initial extraction rules from registered templates, stores these rules in the rule storage unit, and provides them to the determination unit for various document types. This unified approach ensures extraction reliability across different document types while avoiding the complexity of separate rule-setting mechanisms for each document type.
Solution Approach 2:
The system uses copying by replicating proven extraction rules from registered document templates to new documents of the same type. Instead of creating complex unique rules for each document, the system copies and adapts rules from the template, ensuring reliable extraction while keeping the system simple. The rule storage unit maintains these copied rules for efficient reuse.
Data Source
AI summary
To make it possible to extract a value with a high accuracy without imposing a heavy burden on a user even in a case where the character string row of a value corresponding to a certain item within a document changes. Based on a value extraction rule of a registered document whose type is the same as that of an input document, a text block corresponding to a value is estimated from among text blocks included in the scanned image of the input document and the character string that is the value is extracted. Then, after a user modifies the text block corresponding to the extracted character string, a rule is derived for estimating a value block so that it is possible to estimate the modified text block as the text block corresponding to the value.


