Document Image Information Extraction via Key Point Alignment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for extracting information from document images, particularly tables, face challenges in accurately and efficiently extracting information due to variations in orientation, tilt, and layout, leading to poor structured effects, especially in scenarios with large handwritten content or offsetting of needle punching.
Innovation Solution
A method and apparatus that acquire a location template corresponding to a document image category, determine key point locations, generate a transformation matrix based on these locations, and use it to determine and extract information locations, thereby standardizing and simplifying the extraction process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional OCR and structuralization methods are used to extract information from document images, then information extraction can be performed, but the structured effect is poor when dealing with variations in orientation, tilt, and layout
Solution Approach 1:
The patent applies preliminary action by performing orientation and skew correction before information extraction. The system detects the orientation of document images and tables, then pre-processes them to correct skew and alignment issues. This preliminary correction ensures that subsequent OCR and structuralization operations work on properly oriented images, significantly improving extraction accuracy for documents with various orientations and tilts.
Solution Approach 2:
The patent employs parameter changes by dynamically adjusting orientation parameters based on detected document and table angles. The system calculates rotation angles from detected key points and applies transformation matrices to change the orientation parameters of images and tables. This allows the extraction system to adapt to different document orientations and layouts, improving both accuracy and versatility.
2Measurement precision
If table range reconstruction and cell-by-cell OCR are performed, then information can be extracted from tables, but the process becomes complex and time-consuming
Solution Approach 1:
The patent applies segmentation by dividing the document processing into distinct components: document-level orientation detection, table detection, cell segmentation, and information extraction. The system first detects overall document orientation, then identifies table regions, segments cells within tables, and finally extracts information from each cell. This hierarchical segmentation improves accuracy while managing complexity through organized processing stages.
Solution Approach 2:
The patent applies preliminary action by performing table detection and range reconstruction before detailed cell processing. The system detects table boundaries and reconstructs table structures in advance, then uses this pre-processed information to guide subsequent cell-by-cell extraction. This preliminary table structure establishment reduces the complexity of cell processing and improves overall extraction efficiency.
3Measurement precision
If orientation and skew correction is applied to document images, then extraction accuracy improves, but additional processing time is required
Solution Approach 1:
The patent employs parameter changes by calculating optimal rotation angles from detected key points and applying transformation matrices only when needed. The system determines orientation parameters from document and table detections, then applies corrective transformations only to images requiring adjustment. This selective parameter adjustment improves accuracy while minimizing unnecessary processing time for already-properly-oriented documents.
4Measurement precision
If comprehensive table detection and reconstruction is performed, then information extraction completeness improves, but device complexity increases
Solution Approach 1:
The patent applies segmentation by separating table detection, range reconstruction, cell segmentation, and information extraction into distinct processing modules. Each module handles a specific aspect of table processing, making the overall complex system manageable through functional decomposition. This segmentation allows comprehensive table processing while organizing complexity into manageable, independent components that can be processed in sequence.
Data Source
AI summary
Embodiments of the present disclosure disclose a method and apparatus for extracting information, a device and a storage medium, relate to the field of image processing technology. The method may include: acquiring a location template corresponding to a category of a target document image; determining key point locations on the target document image; generating a transformation matrix based on the key point locations on the target document image and key point locations on the location template; determining locations of information corresponding to the target document image, based on locations of information on the location template and the transformation matrix; and extracting information at the locations of information corresponding to the target document image to obtain information in the target document image.


