Document OCR Region Specification via Mark Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Users face difficulties in effectively performing character recognition on documents when they need to specify regions for OCR processing, especially when the symbol for instruction is hard to write in a small space or overlaps with other characters, leading to degraded usability compared to handwritten instructions.
Innovation Solution
An information processing apparatus that extracts pre-specified marks from a document image and performs character recognition on regions associated with these marks, allowing for improved region specification and usability by using an extraction table to define marks, reading positions, and associated applications.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If a user handwrites a symbol in a document to instruct OCR processing on a region near the symbol, then the user can specify a region for character recognition, but the symbol may overlap with other characters or be too small to write, leading to failed character recognition
Solution Approach 1:
The patent introduces a display device as an intermediary between the user and the document. The user specifies the region for OCR on the displayed image of the document rather than writing directly on the physical document. This mediator allows for easier and more reliable region specification through digital interaction (clicking, dragging, or using input devices) without the physical constraints of paper space, thereby resolving the contradiction between ease of operation and recognition reliability
Solution Approach 2:
The patent creates a digital copy of the document image on the display device. Instead of manipulating the physical document directly, the user interacts with the copied digital representation. This copy allows for flexible region specification without affecting the original document and enables precise control over the OCR region through digital interfaces, improving both ease of use and accuracy
2Measurement precision
If a user specifies a region for character recognition while viewing a document image on a screen, then the user can indicate the region, but usability is degraded compared to handwriting a symbol directly in the document
Solution Approach 1:
The patent segments the document processing into distinct phases: displaying the document image, receiving region specification input, and performing OCR on the specified region. This segmentation allows the system to handle each task optimally - display for visualization, input device for precise region selection, and OCR engine for text recognition. The separation of concerns improves both precision and usability compared to requiring a single handwritten symbol to convey all information
Data Source
AI summary
An information processing apparatus includes a processor configured to extract a mark specified in advance from an image of a document; and acquire a character string by performing character recognition on a region located in a particular direction with respect to a position of the mark, the direction being associated in advance with the mark.


