Synthetic Form Image Generation for ML Training Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current synthetic data generation methods are inadequate for form document images, particularly in handling numerical data and considering field dependencies, which are crucial for machine learning algorithms to process and understand form content, due to the lack of effective synthetic form image generation techniques.
Innovation Solution
A method and system for generating synthetic form images by classifying field value data into categories, learning statistical distributions for categorical and numerical data, sampling data elements randomly, and assembling the synthetic data with field labels to produce labeled textual data sets, which are then rendered over a structured form layout to create synthetic form images.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If synthetic data generation is performed at character or word level using previous methods, then data generation can be achieved, but numerical valued data which constitutes more than 50% of field values cannot be synthesized
Solution Approach 1:
The patent changes the data generation approach from character/word level to statistical distribution level. It learns statistical distributions (mean, standard deviation, min, max) for numerical fields and uses these parameters to generate realistic numerical values, thereby enabling synthesis of numerical valued data that constitutes more than 50% of form fields.
Solution Approach 2:
The patent replaces manual data entry and verification mechanisms with an automated synthetic data generation system. It substitutes human involvement with machine learning algorithms that learn from existing form images and automatically generate labeled synthetic data, eliminating the need for expensive manual processes.
2Reliability
If real form images are used for training machine learning algorithms, then high quality training data can be obtained, but expensive human verification and manual field level redaction are required due to sensitive nature of some fields
Solution Approach 1:
The patent creates synthetic copies of real form images by rendering generated data onto template form layouts. These synthetic copies preserve the visual structure and format of real forms while containing artificially generated data, eliminating the need to use actual sensitive forms for training purposes.
Solution Approach 2:
The patent introduces synthetic data as an intermediary between real sensitive forms and machine learning training requirements. Instead of directly using real forms with sensitive information, the system generates intermediate synthetic representations that maintain form structure and visual characteristics while removing sensitive data concerns.
3Device complexity
If synthetic form images are generated without considering field dependencies, then generation process is simpler, but form field labels required for information extraction cannot be provided
Solution Approach 1:
The patent performs preliminary classification of form fields into data types (personally identifiable information, categorical data, numerical data) before generating synthetic values. This preliminary organization enables the system to apply appropriate generation strategies for each field type while maintaining field labels and dependencies required for information extraction tasks.
Data Source
AI summary
A method and system for generating synthetic form image involves obtaining a multitude of field value data and associated field labels for a chosen type of form document from an electronic data source, classifying the multitude of field value data into a multitude of data categories, where the multitude of data categories, learning statistical data distributions for categorical and numerical data types using the classified categorical and numerical data, and sampling data elements randomly using the learned data distributions to generate synthetic data for categorical and numerical data. The method also involves assembling the synthetic data for the multitude of data categories with the associated field labels to generate a labeled synthetic textual data set, rendering the labeled synthetic textual data set over a structured form layout image to produce a synthetic form image, and storing the synthetic form image and the labeled synthetic textual data set.


