Document Image Information Extraction via Dynamic Search Regions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing document image processing technologies face difficulties in extracting necessary information from documents with complicated layouts, especially when no table structure is present, as they require predefined formats and struggle with arbitrary document formatting.
Innovation Solution
An information processing apparatus that performs optical character recognition, evaluates character strings as item names or values, groups adjacent item names, sets search regions, and extracts item values based on their relationship, allowing for effective information extraction regardless of the presence of a table structure.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a table structure is used to extract information from document images, then the extraction accuracy for structured data is improved, but the method cannot handle documents without table structures
Solution Approach 1:
The patent segments the document image into multiple search regions based on layout analysis, dividing the complex extraction task into manageable regions. Each region can be processed independently, allowing the system to handle both tabular and non-tabular documents effectively.
Solution Approach 2:
The patent employs dynamic search region setting that adapts to the document structure. The search regions are not fixed but are determined dynamically based on the actual layout characteristics of each document, enabling versatile handling of different document types.
2Device complexity
If predefined document formats are used for information extraction, then the extraction process is simplified, but the method cannot handle documents with arbitrary formats
Solution Approach 1:
The patent performs preliminary layout analysis to identify potential search regions before actual information extraction. This preliminary action prepares the system to handle various formats by pre-segmenting the document structure, simplifying the subsequent extraction process while maintaining versatility.
Solution Approach 2:
The patent changes the parameters of search regions dynamically based on document characteristics. By adjusting region boundaries, sizes, and positions according to the actual document layout, the system maintains simple extraction processes while adapting to arbitrary formats.
3Measurement precision
If search regions are set based on item name groups and adjacent character strings, then the extraction accuracy for complicated layouts is improved, but the processing complexity increases
Solution Approach 1:
The patent segments the document into search regions based on item name groups, creating discrete processing units. This segmentation improves extraction accuracy by focusing on relevant regions while the modular approach keeps processing complexity manageable through systematic region evaluation.
Solution Approach 2:
The patent evaluates potential search regions systematically, considering item name groups and their adjacent character strings. By applying evaluation criteria to identify relevant regions without processing the entire document uniformly, the system achieves high accuracy while controlling processing complexity through selective region evaluation.
Data Source
AI summary
According to a form of the present disclosure, character strings are obtained from a document image by performing an optical character recognition process on the document image, whether each of the obtained character strings is an item name or an item value is evaluated, character strings determined to be the item names and continuously present in a horizontal or vertical direction are grouped as one item name group, a search region is set by combining the one item name group and a region where one or more character strings each determined to be the item value and adjacent to the item name group are continuously present, and the item value is extracted based on the relationship between the item name and the item value for each of the search regions.


