Synthetic Document Generation for Deep Learning Data Augmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine and deep learning models face challenges in accurately extracting data from electronic documents with unique layouts, as they require large volumes of annotated training documents and can be biased towards frequently seen labels, leading to inefficiencies and potential failure when encountering structurally different documents.

Innovation Solution

The method involves generating synthetic electronic documents with semantic and structural variance using macro and micro augmentation operations, allowing for the creation of a training data set tailored to a specific document layout, which can be used to train deep learning models effectively.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a learning model is trained on commonly-used electronic documents with standard layouts, then the model can accurately extract data from those standard documents, but the model fails to accurately extract data from electronic documents with unique or structurally different layouts

Engineering Contradiction:
Improvedata extraction accuracyVSAvoidlayout adaptability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent applies geometric transformations (rotations, flips, scaling) and semantic modifications (text replacements, format changes) to the template document, systematically varying parameters to generate diverse synthetic documents that maintain structural relationships while presenting different visual appearances, thereby training the model to recognize patterns across layout variations

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent creates multiple synthetic copies of a single annotated template document through augmentation operations, generating a training dataset where each synthetic document is a transformed copy that preserves the underlying data structure and semantic relationships while varying the visual presentation, enabling the model to learn from diverse layout examples without requiring multiple manually annotated documents

Inventive Principle:
Principle #26Copying

2Measurement precision

If manually annotated training documents are used to train the learning model, then the model can be trained on specific document structures, but the process requires large time and cost investment

Engineering Contradiction:
Improvedata extraction accuracyVSAvoidannotation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent creates multiple synthetic copies of a single annotated template document through augmentation operations, generating a training dataset where each synthetic document is a transformed copy that preserves the underlying data structure and semantic relationships while varying the visual presentation, enabling the model to learn from diverse layout examples without requiring multiple manually annotated documents

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs self-augmentation by automatically generating synthetic training documents from a single annotated template using geometric and semantic transformations, eliminating the need for external human annotators to create diverse training data, thereby significantly reducing annotation time and cost while maintaining model training effectiveness

Inventive Principle:
Principle #25Self-service

3Adaptability or versatility

If additional training data is provided to improve model performance on structurally-different documents, then more training documents are available, but this does not guarantee improvement and may lead to average performance across all documents

Engineering Contradiction:
Improvedocument structure adaptabilityVSAvoiddata extraction accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent applies different augmentation operations to different regions of the document (header, body, footer sections) with region-specific transformation parameters, allowing the model to learn layout-specific patterns while maintaining focus on critical data extraction areas, thereby improving performance on structurally-different documents without diluting accuracy through uniform random augmentations

Inventive Principle:
Principle #3Local quality

4Productivity

If a global learning model is used to handle commonly-used documents, then the model can process standard documents effectively, but the model struggles with documents that have unique layouts requiring custom training

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidlayout flexibility
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal training template that can generate synthetic documents representing multiple document types and layout variations through parameterized augmentation operations, allowing a single trained model to handle diverse document structures effectively, thereby eliminating the need for separate custom models for different document types while maintaining processing efficiency

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20230334309A1Augmenting electronic documents to generate synthetic training data sets
Publication Date: 2023.10.19 SAP SE
  • US20230334309A1 patent drawing
  • US20230334309A1 patent drawing
  • US20230334309A1 patent drawing

AI summary

Systems, methods, and computer-readable media for generating a synthetic training data set from an original unstructured electronic document are disclosed. The synthetic training data set may be used to train a deep learning model to extract data from the original electronic document. The original electronic document may comprise annotated data fields. Each annotated data field may comprise a bounding box and a label. The original electronic document may comprise a header, a table, and a footer. Macro augmentation operations may be applied to the original electronic document to create sub-templates representative of distinct page layouts in the original electronic document. The synthetic training data set may be generated by applying geometric and semantic data augmentations to the sub-templates and the original electronic documents. The synthetic training data set may then be provided the deep learning model for training.