Synthetic Form Image Generation for ML Training Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current synthetic data generation methods are inadequate for form document images, particularly in handling numerical data and considering field dependencies, which are crucial for machine learning algorithms to process and understand form content, due to the lack of effective synthetic form image generation techniques.

Innovation Solution

A method and system for generating synthetic form images by classifying field value data into categories, learning statistical distributions for categorical and numerical data, sampling data elements randomly, and assembling the synthetic data with field labels to produce labeled textual data sets, which are then rendered over a structured form layout to create synthetic form images.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If synthetic data generation is performed at character or word level using previous methods, then data generation can be achieved, but numerical valued data which constitutes more than 50% of field values cannot be synthesized

Engineering Contradiction:
Improvequantity of synthesizable field valuesVSAvoidability to handle numerical data
Core Design Contradiction:
Quantity of substanceVSAdaptability or versatility

Solution Approach 1:

The patent changes the data generation approach from character/word level to statistical distribution level. It learns statistical distributions (mean, standard deviation, min, max) for numerical fields and uses these parameters to generate realistic numerical values, thereby enabling synthesis of numerical valued data that constitutes more than 50% of form fields.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces manual data entry and verification mechanisms with an automated synthetic data generation system. It substitutes human involvement with machine learning algorithms that learn from existing form images and automatically generate labeled synthetic data, eliminating the need for expensive manual processes.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If real form images are used for training machine learning algorithms, then high quality training data can be obtained, but expensive human verification and manual field level redaction are required due to sensitive nature of some fields

Engineering Contradiction:
Improvequality of training dataVSAvoidcost and complexity of data preparation
Core Design Contradiction:
ReliabilityVSEase of manufacture

Solution Approach 1:

The patent creates synthetic copies of real form images by rendering generated data onto template form layouts. These synthetic copies preserve the visual structure and format of real forms while containing artificially generated data, eliminating the need to use actual sensitive forms for training purposes.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent introduces synthetic data as an intermediary between real sensitive forms and machine learning training requirements. Instead of directly using real forms with sensitive information, the system generates intermediate synthetic representations that maintain form structure and visual characteristics while removing sensitive data concerns.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Device complexity

If synthetic form images are generated without considering field dependencies, then generation process is simpler, but form field labels required for information extraction cannot be provided

Engineering Contradiction:
Improvecomplexity of generation processVSAvoidloss of field dependency information
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The patent performs preliminary classification of form fields into data types (personally identifiable information, categorical data, numerical data) before generating synthetic values. This preliminary organization enables the system to apply appropriate generation strategies for each field type while maintaining field labels and dependencies required for information extraction tasks.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10546054B1System and method for synthetic form image generation
Publication Date: 2020.01.28 INTUIT INC
  • US10546054B1 patent drawing
  • US10546054B1 patent drawing
  • US10546054B1 patent drawing

AI summary

A method and system for generating synthetic form image involves obtaining a multitude of field value data and associated field labels for a chosen type of form document from an electronic data source, classifying the multitude of field value data into a multitude of data categories, where the multitude of data categories, learning statistical data distributions for categorical and numerical data types using the classified categorical and numerical data, and sampling data elements randomly using the learned data distributions to generate synthetic data for categorical and numerical data. The method also involves assembling the synthetic data for the multitude of data categories with the associated field labels to generate a labeled synthetic textual data set, rendering the labeled synthetic textual data set over a structured form layout image to produce a synthetic form image, and storing the synthetic form image and the labeled synthetic textual data set.