Synthetic Document Image Augmentation for Deep Learning Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Automated recognition of fields and associated values in business documents with varying formats and image quality is challenging due to scanning artifacts and limited availability of diverse document datasets for training machine learning systems.

Innovation Solution

A computerized system that synthetically generates distorted document images using augmentation modules to simulate scanning effects, allowing for precise training of deep learning networks and testing of preprocessing components.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If machine learning systems are trained on real scanned document images, then the system learns to handle real-world scanning artifacts and distortions, but obtaining large diverse datasets is challenging due to confidentiality and cost

Engineering Contradiction:
Improvetraining data qualityVSAvoiddataset size
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent creates synthetic copies of document images by applying various distortion transformations (skew, perspective, noise, compression artifacts) to clean template documents. This allows generating large volumes of training data without needing actual scanned documents, resolving the contradiction between data quality and data quantity

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system varies multiple parameters in the synthetic image generation process including distortion strength, noise levels, compression ratios, and transformation types. By systematically changing these parameters, the system generates diverse training samples that cover the range of real-world variations without requiring additional physical documents

Inventive Principle:
Principle #35Parameter changes

2Reliability

If traditional computer vision approaches are used for preprocessing, then the system can handle document distortions, but the approach is less effective compared to deep learning methods

Engineering Contradiction:
Improvedistortion correction accuracyVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies preprocessing distortions to the training data before the deep learning model is trained. This preliminary action allows the model to learn distortion correction capabilities directly during training, eliminating the need for separate traditional computer vision preprocessing steps and achieving better accuracy

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If deep learning systems are trained on diverse document formats and qualities, then recognition accuracy improves, but the variability in formatting and image quality makes automated processing challenging

Engineering Contradiction:
Improvefield recognition accuracyVSAvoidprocessing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system generates synthetic training data with controlled variations in formatting parameters, image quality parameters, and distortion parameters. This systematic parameter variation allows the model to learn robust recognition across diverse formats while maintaining manageable system complexity through automated synthetic data generation

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10984284B1Synthetic augmentation of document images
Publication Date: 2021.04.20 AUTOMATION ANYWHERE INC
  • US10984284B1 patent drawing
  • US10984284B1 patent drawing
  • US10984284B1 patent drawing

AI summary

A computerized method and system for adding distortions to a computer-generated image of a document stored in an image file. An original computer-generated image file is selected and is processed to generate one or more distorted image files for each original computer-generated image file by selecting one or more augmentation modules from a set of augmentation modules to form an augmentation sub-system. The original computer-generated image file is processed with the augmentation sub-system to generate an augmented image file by altering the original computer-generated image file to add distortions that simulate distortions introduced during scanning of a paper-based representation of a document represented in the original computer-generated image file.