Document Type Identification Using Segmented Determination Units
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for extracting character strings from scanned images often incorrectly identify unregistered documents with similar layouts, leading to erroneous overwriting of registered document information, which can result in failed extraction of desired character strings in subsequent scans.
Innovation Solution
An image processing apparatus with an obtaining unit, a first determination unit, an extraction unit, a second determination unit, and a display control unit to accurately identify document types, extract relevant character strings, and prompt users to overwrite or register new documents based on similarity analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If a single determination method is used to identify document types, then the processing speed is fast, but the accuracy of distinguishing between registered and unregistered documents deteriorates
Solution Approach 1:
The patent divides the document identification process into two separate determination methods: a first determination unit that performs initial document type identification, and a second determination unit that performs verification. This segmentation allows the system to maintain both speed (through the first determination) and accuracy (through the second determination), resolving the contradiction between processing speed and identification accuracy.
2Ease of operation
If automatic overwriting is performed based on layout similarity, then the ease of operation is improved, but the reliability of information integrity deteriorates
Solution Approach 1:
The patent implements a feedback mechanism where the second determination unit verifies whether the extracted character string actually corresponds to the target document type before allowing overwriting. This feedback loop prevents erroneous overwriting of unregistered documents while maintaining automatic operation, thus resolving the contradiction between ease of operation and information integrity.
3Device complexity
If a single determination method is used, then the device complexity is low, but the measurement precision of document format similarity deteriorates
Solution Approach 1:
The patent segments the determination process into two distinct units with different functions: the first determination unit handles initial document type identification, while the second determination unit handles verification of document format similarity. This segmentation improves measurement precision without creating excessive complexity, as each unit has a specific, well-defined role.
Solution Approach 2:
The patent applies different determination methods to different aspects of document analysis: the first determination unit uses one method for initial identification, while the second determination unit uses a different method for verification. This local quality approach allows each unit to be optimized for its specific task, improving overall precision while maintaining reasonable system complexity.
Data Source
AI summary
The image processing apparatus has an obtaining unit configured to obtain a scanned image, a first determination unit configured to determine a document type of a document format similar to a document format of the scanned image based on information on each registered document type, an extraction unit configured to extract a character string corresponding to a predetermined item, a second determination unit configured to use a different method for determining whether the document format indicated by the scanned image is similar to the document format of the document type determined by the first determination unit in a case where a user modifies the extracted character string, and a display control unit configured to display a screen prompting the user to perform overwriting in a case where the second determination unit determines that the document format is similar.


