Synthetic Training Data Generation for Document Key-Value Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The availability of diverse and large-volume training data for machine learning models to extract key-value pairs from document images is limited due to the tedious and resource-intensive manual annotation process, leading to deficient model performance.

Innovation Solution

Automated techniques generate synthetic training data, including synthetic document images and annotation data, using optical character recognition and key-value databases to produce a large number of diverse training datapoints without human intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual annotation is used to create training data, then annotation accuracy can be maintained, but the process becomes tedious and time-consuming

Engineering Contradiction:
Improveannotation accuracyVSAvoidannotation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent uses OCR technology to create copies of text content from document images and generates synthetic training data by combining these copied text elements with template images. This automated copying and重组 process replaces manual annotation while maintaining data quality, thereby reducing annotation time without sacrificing accuracy.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs self-annotation by automatically extracting text from document images using OCR, identifying key-value pairs through algorithmic processing, and generating annotated training data without human intervention. This self-service approach eliminates the need for manual annotation while maintaining consistent quality standards.

Inventive Principle:
Principle #25Self-service

2Reliability

If more training data is generated to improve model accuracy, then model performance increases, but the manual annotation process becomes more resource-intensive

Engineering Contradiction:
Improvemodel accuracyVSAvoiddata generation efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent generates multiple synthetic training samples by copying and recombining extracted text content with various template images. This automated copying process can generate large volumes of training data efficiently, improving model accuracy through increased data diversity without proportionally increasing resource consumption.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system varies parameters such as text content, template selection, and image characteristics to generate diverse synthetic training data. By changing these parameters automatically, the system can produce numerous training samples with different characteristics, improving model robustness and accuracy while maintaining high generation efficiency.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If diverse training data covering various document types is created, then model versatility improves, but the complexity of manual annotation increases

Engineering Contradiction:
Improvemodel versatilityVSAvoidannotation complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent employs a universal template-based approach where the same OCR and text extraction processes can handle multiple document types. By using configurable templates that can represent different document formats, the system achieves multi-functionality in data generation, improving model versatility without increasing annotation complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system segments the training data generation process into independent modules: OCR text extraction, key-value pair identification, template selection, and image synthesis. This segmentation allows each module to be optimized independently and facilitates the handling of diverse document types through modular template configurations, maintaining simplicity while improving versatility.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250322683A1Generating synthetic training data including document images with key-value pairs
Publication Date: 2025.10.16 ORACLE INT CORP
  • US20250322683A1 patent drawing
  • US20250322683A1 patent drawing
  • US20250322683A1 patent drawing

AI summary

Automated techniques are for generating a large volume of diverse training data that can be used for training machine learning models to extract KV pairs from document images. Given a single input document image and associated annotation data, a large number of diverse synthetic training datapoints are automatically generated by a synthetic data generation system, each datapoint including a synthetic document image and associated annotation data. The generated synthetic training datapoints can be used to train and improve the performance of ML models for extracting KV pairs from document images. In certain implementations, multiple synthetic datapoints are generated by varying the values associated with a key for a content item within the input document image.