Probabilistic Synthetic Document Templates for Format-Adaptive Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning models struggle to effectively recognize and extract information from diverse business document formats due to variations in formatting and layout across different businesses, necessitating extensive customization for each pattern.

Innovation Solution

A system generates templates for synthetic document generation, utilizing randomization functions to select and position entities based on probabilistic models of document characteristics, allowing for the creation of diverse training datasets that mimic real-world document formats.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If machine learning models are trained to recognize patterns in business documents, then the models can extract information from documents following those patterns, but the models require extensive customization for each different document pattern

Engineering Contradiction:
Improveinformation extraction accuracyVSAvoidmodel customization complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent creates synthetic document datasets by copying and transforming template documents with randomized elements. These synthetic copies mimic real document patterns while providing controlled variation, allowing models to learn multiple patterns without requiring separate customization for each pattern type.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system randomizes various parameters in document templates including layout configurations, entity positions, formatting styles, and content values. By systematically varying these parameters across synthetic training documents, the model learns to handle diverse document patterns through parameter adaptation rather than structural customization.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If machine learning models are trained with diverse document formats, then the models become more adaptable to different business document patterns, but training data collection and preparation becomes more complex

Engineering Contradiction:
Improvedocument format adaptabilityVSAvoidtraining data preparation complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by pre-defining document templates with structured placeholders and configurable parameters before generating training data. This upfront preparation of template structures enables systematic generation of diverse training documents without complex ad-hoc data collection for each pattern.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The template-based generation system serves multiple functions: it creates diverse training documents, controls data distribution, ensures proper formatting, and facilitates efficient data generation. This universal approach replaces multiple specialized data preparation processes with a single multi-functional template system.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If synthetic document datasets are generated using templates and randomization, then diverse training data can be created efficiently, but the generation process requires sophisticated template management

Engineering Contradiction:
Improvetraining data generation efficiencyVSAvoidtemplate management complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The document templates are segmented into distinct components including layout structures, entity placeholders, formatting parameters, and content templates. This segmentation allows independent management and randomization of each component, simplifying the overall template management while enabling efficient generation of diverse training documents.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12380719B2Generating templates for use in synthetic document generation processes
Publication Date: 2025.08.05 ORACLE INT CORP
  • US12380719B2 patent drawing
  • US12380719B2 patent drawing
  • US12380719B2 patent drawing

AI summary

The system generates templates to be used for use in synthetic document generation. Generating a template includes selecting entities to be included in the template and/or selecting characteristics of entities in the template. The system may execute randomization functions to determine which entities, of a candidate set of entities, are to be included in a template. The randomization functions may accept, as input, probabilities associated with the entities to compute the inclusion or exclusion of entities in the template. Different entities may be associated with different corresponding probabilities. The system may execute randomization functions to determine characteristics of entities that are included in a template. The characteristics may be selected from a specific candidate set of characteristics or a range of characteristic values.