Convolutional Neural Document Conversion Model Using Synthetic Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing OCR, HTR, and OMR systems face limitations in computational efficiency and operational reliability due to limited training data availability and the computationally expensive use of language models, which hinder effective character-level conversion of document images.

Innovation Solution

The use of synthetically generated machine printed and handwritten text data, along with greedy decoding, to train convolutional neural document conversion models, eliminating the need for language models and enhancing training efficiency and reliability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If language models are used for document conversion, then conversion accuracy may be improved, but computational complexity and resource requirements increase significantly

Engineering Contradiction:
Improveconversion accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts and removes the language model component from the document conversion system, retaining only the essential convolutional neural network for character-level conversion. This eliminates the computationally expensive language model while preserving the core conversion functionality through synthetic data training.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent employs synthetic training data that can be generated quickly and discarded after training, replacing the need for extensive real-world annotated data and complex language models. The synthetic data serves as a temporary training resource that enables efficient model development without long-term computational burden.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

2Measurement precision

If extensive real-world training data is collected and annotated, then model training accuracy improves, but data collection time and annotation costs increase

Engineering Contradiction:
Improvemodel training accuracyVSAvoiddata collection time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary generation of synthetic training data that mimics real-world document variations before actual model training begins. This pre-generated data includes diverse scenarios (different fonts, layouts, languages) that would otherwise require extensive manual collection and annotation, thereby eliminating time-consuming data preparation steps.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates synthetic copies of training data through algorithmic generation rather than manual collection. These synthetic data copies replicate the statistical properties and visual characteristics of real document images, providing sufficient training material without requiring physical data collection or human annotation efforts.

Inventive Principle:
Principle #26Copying

3Productivity

If synthetic training data is used, then training efficiency and computational resource requirements improve, but potential loss of real-world data variability occurs

Engineering Contradiction:
Improvetraining efficiencyVSAvoiddata variability
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent systematically varies parameters in synthetic data generation including font types, sizes, styles, document layouts, languages, and image quality characteristics. By controlling and randomizing these parameters, the system generates diverse synthetic data that covers the full range of real-world document variations, maintaining adaptability while achieving training efficiency.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent designs the synthetic data generation system to produce multi-functional training data that serves multiple purposes simultaneously: training for different languages, document types, and visual conditions. This universal approach allows a single synthetic data pipeline to replace multiple specialized data collection efforts, maintaining versatility while improving productivity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12046064B2Predictive document conversion
Publication Date: 2024.07.23 OPTUM TECH INC
  • US12046064B2 patent drawing
  • US12046064B2 patent drawing
  • US12046064B2 patent drawing

AI summary

There is a need for more effective and efficient predictive document conversion. This need can be addressed by, for example, solutions for performing document conversion using a trained convolutional neural document conversion machine learning. In one example, the trained convolutional neural document conversion machine learning model is associated with a preprocessing block having a plurality of preprocessing subblocks, one or more main processing blocks each having a plurality of main processing subblocks, and a plurality of postprocessing subblocks each having one or more postprocessing subblocks, and the trained convolutional neural document conversion machine learning model is further associated with a preprocessing subblock repetition count hyper-parameter that defines a preprocessing subblock count of the plurality of preprocessing subblocks.