Document Image Information Extraction via Key Point Alignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for extracting information from document images, particularly tables, face challenges in accurately and efficiently extracting information due to variations in orientation, tilt, and layout, leading to poor structured effects, especially in scenarios with large handwritten content or offsetting of needle punching.

Innovation Solution

A method and apparatus that acquire a location template corresponding to a document image category, determine key point locations, generate a transformation matrix based on these locations, and use it to determine and extract information locations, thereby standardizing and simplifying the extraction process.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional OCR and structuralization methods are used to extract information from document images, then information extraction can be performed, but the structured effect is poor when dealing with variations in orientation, tilt, and layout

Engineering Contradiction:
Improveinformation extraction accuracyVSAvoidhandling of orientation and layout variations
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent applies preliminary action by performing orientation and skew correction before information extraction. The system detects the orientation of document images and tables, then pre-processes them to correct skew and alignment issues. This preliminary correction ensures that subsequent OCR and structuralization operations work on properly oriented images, significantly improving extraction accuracy for documents with various orientations and tilts.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent employs parameter changes by dynamically adjusting orientation parameters based on detected document and table angles. The system calculates rotation angles from detected key points and applies transformation matrices to change the orientation parameters of images and tables. This allows the extraction system to adapt to different document orientations and layouts, improving both accuracy and versatility.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If table range reconstruction and cell-by-cell OCR are performed, then information can be extracted from tables, but the process becomes complex and time-consuming

Engineering Contradiction:
Improvetable information extraction accuracyVSAvoidinformation extraction speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent applies segmentation by dividing the document processing into distinct components: document-level orientation detection, table detection, cell segmentation, and information extraction. The system first detects overall document orientation, then identifies table regions, segments cells within tables, and finally extracts information from each cell. This hierarchical segmentation improves accuracy while managing complexity through organized processing stages.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies preliminary action by performing table detection and range reconstruction before detailed cell processing. The system detects table boundaries and reconstructs table structures in advance, then uses this pre-processed information to guide subsequent cell-by-cell extraction. This preliminary table structure establishment reduces the complexity of cell processing and improves overall extraction efficiency.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If orientation and skew correction is applied to document images, then extraction accuracy improves, but additional processing time is required

Engineering Contradiction:
Improveextraction accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent employs parameter changes by calculating optimal rotation angles from detected key points and applying transformation matrices only when needed. The system determines orientation parameters from document and table detections, then applies corrective transformations only to images requiring adjustment. This selective parameter adjustment improves accuracy while minimizing unnecessary processing time for already-properly-oriented documents.

Inventive Principle:
Principle #35Parameter changes

4Measurement precision

If comprehensive table detection and reconstruction is performed, then information extraction completeness improves, but device complexity increases

Engineering Contradiction:
Improveinformation extraction completenessVSAvoidprocessing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies segmentation by separating table detection, range reconstruction, cell segmentation, and information extraction into distinct processing modules. Each module handles a specific aspect of table processing, making the overall complex system manageable through functional decomposition. This segmentation allows comprehensive table processing while organizing complexity into manageable, independent components that can be processed in sequence.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11468655B2Method and apparatus for extracting information, device and storage medium
Publication Date: 2022.10.11 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US11468655B2 patent drawing
  • US11468655B2 patent drawing
  • US11468655B2 patent drawing

AI summary

Embodiments of the present disclosure disclose a method and apparatus for extracting information, a device and a storage medium, relate to the field of image processing technology. The method may include: acquiring a location template corresponding to a category of a target document image; determining key point locations on the target document image; generating a transformation matrix based on the key point locations on the target document image and key point locations on the location template; determining locations of information corresponding to the target document image, based on locations of information on the location template and the transformation matrix; and extracting information at the locations of information corresponding to the target document image to obtain information in the target document image.