Probabilistic Model for Repeated Structure Data Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data extraction methods for documents with repeated structure are inefficient due to variability in field content, layout, and scanning distortions, leading to high manual processing costs and limitations in automated solutions that assume periodic structures and separated columns.

Innovation Solution

A probabilistic model-based system that identifies a reference record and fields, generates candidate fields and records, and selects optimal records using perceptual cues like alignment and saliency, allowing for flexible extraction of repeated structure across documents with varying layouts and interlacing columns.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional data extraction methods are used, then extraction speed is improved, but extraction accuracy deteriorates due to document variability

Engineering Contradiction:
Improveextraction speedVSAvoidextraction accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system dynamically adjusts extraction parameters such as field boundary definitions, spatial relationship thresholds, and matching criteria based on the specific document characteristics detected during analysis. This allows the extraction process to adapt to variations in document layouts, fonts, and structures while maintaining high accuracy across different document types.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system performs self-calibration by automatically learning from annotated training documents and refining its extraction models without requiring manual reconfiguration for each document type. The probabilistic model continuously improves its parameters based on training data, enabling accurate extraction across variable document formats while maintaining high productivity.

Inventive Principle:
Principle #25Self-service

2Loss of energy

If automated extraction systems are implemented, then processing cost is reduced, but reliability deteriorates due to inability to handle layout variations

Engineering Contradiction:
Improveprocessing costVSAvoidextraction reliability
Core Design Contradiction:
Loss of energyVSReliability

Solution Approach 1:

The extraction system transitions from static, pre-defined extraction rules to dynamic, adaptive processing that automatically adjusts to different document layouts. The system uses probabilistic models that can handle interlacing columns, varying field positions, and unexpected document structures, maintaining reliability while reducing processing costs through full automation.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system incorporates feedback mechanisms where extraction results are validated and used to refine the probabilistic models. Training documents with known correct answers allow the system to learn from errors and improve its reliability over time, while the automated nature maintains cost efficiency.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If probabilistic model with multiple cues is used, then extraction accuracy is improved, but system complexity increases

Engineering Contradiction:
Improveextraction accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The complex probabilistic model is segmented into multiple independent cue-generating components, each responsible for extracting specific features (spatial relationships, text content, visual patterns). These modular cues are then combined probabilistically, which manages complexity while maintaining high extraction accuracy through systematic integration of multiple independent analysis streams.

Inventive Principle:
Principle #1Segmentation

4Measurement precision

If wrapping approach with manual subgraph matching is used, then extraction accuracy is improved, but productivity deteriorates due to manual specification requirements

Engineering Contradiction:
Improveextraction accuracyVSAvoidprocessing throughput
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system automatically generates and refines extraction rules through training on annotated documents, eliminating the need for manual subgraph matching specification. This self-learning capability maintains extraction accuracy while dramatically improving productivity by enabling automated processing of large document volumes without manual configuration for each document type.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS8625886B2Finding repeated structure for data extraction from document images
Publication Date: 2014.01.07 XEROX CORP
  • US8625886B2 patent drawing
  • US8625886B2 patent drawing
  • US8625886B2 patent drawing

AI summary

Methods and system employing the same for finding repeated structure for data extraction from document images are provided. A reference record and one or more reference fields thereof are identified from a document image. One or more candidate fields are generated for each of the reference fields. One or more best candidate records from the candidate fields are selected using a probabilistic model and an optimal record set is determined from the best candidate records.