Visual-Word Codebooks for Document Field Detection in Unknown Layouts

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for extracting information from document images face challenges due to complex layouts and the need for large labeled datasets, which are often impractical, especially for documents with unknown structures.

Innovation Solution

A novel approach using a codebook of visual words generated from document images, optimized through maximizing mutual information, allows for the detection of fields in documents with unknown layouts based on spatial structure, without requiring extensive labeled data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional methods are used for extracting information from document images, then extraction accuracy can be achieved, but the system requires large labeled datasets and complex layouts which makes it impractical for documents with unknown structures

Engineering Contradiction:
Improveadaptability to documents with unknown layoutsVSAvoidamount of labeled data required
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The system performs self-service by automatically generating the codebook from the document images themselves without requiring external labeled datasets. The codebook is built by extracting keypoint regions, calculating local descriptors, and clustering them to create visual words that naturally represent the document structure, allowing the system to adapt to unknown layouts autonomously

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The invention changes the approach from using traditional labeled data parameters to using an optimized codebook parameter system. By maximizing mutual information between visual words and field positions, the system transforms the representation parameters to enable accurate field detection without labeled data, adapting to any document layout through this parameter optimization

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If codebook optimization is performed by maximizing mutual information, then field detection accuracy is improved, but computational complexity increases

Engineering Contradiction:
Improvefield position detection accuracyVSAvoidcomputational complexity of codebook optimization
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary action by pre-extracting keypoint regions and pre-calculating local descriptors for all document images before the optimization process. This preliminary processing creates a ready-to-use codebook that can be optimized efficiently, reducing the computational burden during the mutual information maximization step and enabling accurate field detection without excessive complexity

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250329187A1Detecting fields in document images
Publication Date: 2025.10.23 ABBYY DEVELOPMENT INC
  • US20250329187A1 patent drawing
  • US20250329187A1 patent drawing
  • US20250329187A1 patent drawing

AI summary

A method of detecting fields in document images includes: receiving, by a processing device, a codebook comprising a set of visual words, each visual word corresponding to a center of a cluster of local descriptors, wherein each local descriptor is associated with a respective keypoint region of a first set of document images; calculating, based on a second set of document images, for each visual word of the codebook, a respective frequency distribution of a field position of a specified field with respect to the visual word; loading a document image for extraction of target fields; and detecting fields in the document image based on the calculated frequency distributions.