Synthetic Document Generation for Deep Learning Data Augmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine and deep learning models face challenges in accurately extracting data from electronic documents with unique layouts, as they require large volumes of annotated training documents and can be biased towards frequently seen labels, leading to inefficiencies and potential failure when encountering structurally different documents.
Innovation Solution
The method involves generating synthetic electronic documents with semantic and structural variance using macro and micro augmentation operations, allowing for the creation of a training data set tailored to a specific document layout, which can be used to train deep learning models effectively.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a learning model is trained on commonly-used electronic documents with standard layouts, then the model can accurately extract data from those standard documents, but the model fails to accurately extract data from electronic documents with unique or structurally different layouts
Solution Approach 1:
The patent applies geometric transformations (rotations, flips, scaling) and semantic modifications (text replacements, format changes) to the template document, systematically varying parameters to generate diverse synthetic documents that maintain structural relationships while presenting different visual appearances, thereby training the model to recognize patterns across layout variations
Solution Approach 2:
The patent creates multiple synthetic copies of a single annotated template document through augmentation operations, generating a training dataset where each synthetic document is a transformed copy that preserves the underlying data structure and semantic relationships while varying the visual presentation, enabling the model to learn from diverse layout examples without requiring multiple manually annotated documents
2Measurement precision
If manually annotated training documents are used to train the learning model, then the model can be trained on specific document structures, but the process requires large time and cost investment
Solution Approach 1:
The patent creates multiple synthetic copies of a single annotated template document through augmentation operations, generating a training dataset where each synthetic document is a transformed copy that preserves the underlying data structure and semantic relationships while varying the visual presentation, enabling the model to learn from diverse layout examples without requiring multiple manually annotated documents
Solution Approach 2:
The system performs self-augmentation by automatically generating synthetic training documents from a single annotated template using geometric and semantic transformations, eliminating the need for external human annotators to create diverse training data, thereby significantly reducing annotation time and cost while maintaining model training effectiveness
3Adaptability or versatility
If additional training data is provided to improve model performance on structurally-different documents, then more training documents are available, but this does not guarantee improvement and may lead to average performance across all documents
Solution Approach 1:
The patent applies different augmentation operations to different regions of the document (header, body, footer sections) with region-specific transformation parameters, allowing the model to learn layout-specific patterns while maintaining focus on critical data extraction areas, thereby improving performance on structurally-different documents without diluting accuracy through uniform random augmentations
4Productivity
If a global learning model is used to handle commonly-used documents, then the model can process standard documents effectively, but the model struggles with documents that have unique layouts requiring custom training
Solution Approach 1:
The patent creates a universal training template that can generate synthetic documents representing multiple document types and layout variations through parameterized augmentation operations, allowing a single trained model to handle diverse document structures effectively, thereby eliminating the need for separate custom models for different document types while maintaining processing efficiency
Data Source
AI summary
Systems, methods, and computer-readable media for generating a synthetic training data set from an original unstructured electronic document are disclosed. The synthetic training data set may be used to train a deep learning model to extract data from the original electronic document. The original electronic document may comprise annotated data fields. Each annotated data field may comprise a bounding box and a label. The original electronic document may comprise a header, a table, and a footer. Macro augmentation operations may be applied to the original electronic document to create sub-templates representative of distinct page layouts in the original electronic document. The synthetic training data set may be generated by applying geometric and semantic data augmentations to the sub-templates and the original electronic documents. The synthetic training data set may then be provided the deep learning model for training.


