Metadata Entry Automation for Document Image OCR Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing techniques for associating metadata with image data using OCR results are inefficient, requiring users to repeatedly select character regions and entry fields, increasing operational burden as the number of metadata items increases.
Innovation Solution
An information processing apparatus that displays a document image alongside entry fields for metadata, allowing users to select character regions for OCR results, with automatic identification and selection of the next blank entry field, reducing the need for frequent mouse pointer movement between the image and entry fields.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If users manually select character regions and entry fields for each metadata item, then metadata can be accurately associated with document images, but operational burden increases significantly as the number of metadata items increases
Solution Approach 1:
The system performs preliminary OCR recognition on the entire document image to extract all character strings and their position information before the user enters metadata. This pre-processing creates a ready-to-use pool of recognized text data that can be automatically associated with metadata entry fields, eliminating the need for users to manually select character regions for each metadata item and significantly reducing operational burden while maintaining accurate association
Solution Approach 2:
The system implements automatic identification of the next blank entry field and automatic association of OCR results with appropriate metadata fields. The system serves itself by automatically matching recognized character strings with corresponding metadata fields based on position, sequence, or other criteria, reducing the need for continuous manual intervention and allowing users to focus only on reviewing and confirming the auto-associated metadata
2Ease of operation
If users frequently move the mouse pointer between the document image and entry fields, then they can select character regions and input metadata, but time consumption increases
Solution Approach 1:
The system merges the document image display area with the metadata entry interface by overlaying or closely positioning entry fields directly on or near the corresponding character regions in the document image. This spatial integration allows users to select character regions and input metadata without frequent mouse pointer movement between separate windows or distant interface elements, significantly reducing time consumption while maintaining full metadata entry capability
Solution Approach 2:
The system replaces manual mechanical mouse pointer movement with automatic programmatic identification and selection of the next blank entry field. The system uses coordinate information from selected character regions to automatically determine and activate the appropriate metadata entry field, substituting repetitive manual mouse operations with automated computational logic that executes instantaneously
Data Source
AI summary
According to an exemplary embodiment of the present disclosure, a screen including a first pane in which a document image is displayed, and a plurality of entry fields in which metadata is to be entered is displayed. In a case where one of character regions in the document image displayed in the first pane is selected by a user, a character recognition result of the selected character region is entered in an entry field that is identified as an input destination among the plurality of entry fields. In a case where the plurality of entry fields includes at least one blank entry field, one of the at least one blank entry field is automatically identified as a next input destination. Accordingly, operability can be improved in entering metadata using character recognition results of character regions selected on the document image.


