HTML Table Rendering Pipeline for Synthetic Training Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The scarcity of quality training data instances for machine learning models used in computer vision services, particularly for table recognition, is exacerbated by high costs and regulatory barriers, leading to inefficiencies in model training.

Innovation Solution

A pipeline comprising an HTML-based and a template-based approach for generating diverse training instances, including HTML code generation, image rendering, and annotation, to create training datasets for machine learning models, addressing various table layouts, backgrounds, and text variations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If quality training data instances are obtained from existing sources, then model training accuracy can be improved, but costs increase and regulatory barriers prevent data usage

Engineering Contradiction:
Improvemodel training accuracyVSAvoiddata acquisition feasibility
Core Design Contradiction:
ReliabilityVSEase of manufacture

Solution Approach 1:

The patent creates synthetic copies of real table images through HTML rendering. Instead of using actual copyrighted documents or sensitive data, the system generates artificial table images that replicate the visual characteristics, layouts, and structures of real tables. This allows unlimited replication of training data without incurring additional costs or facing regulatory restrictions on data usage.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs self-service by automatically generating its own training data through HTML code conversion to images. The pipeline includes automated HTML generation, rendering to images, annotation creation, and quality filtering - all without requiring external data sources, manual annotation, or human intervention. This eliminates dependency on external data providers and regulatory approvals.

Inventive Principle:
Principle #25Self-service

2Reliability

If more training data instances are generated to improve model accuracy, then training quality increases, but data scarcity and regulatory barriers persist

Engineering Contradiction:
Improvetraining data qualityVSAvoidavailable training data volume
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The system performs preliminary action by pre-generating large volumes of diverse training data before model training begins. The HTML-based generation pipeline can create unlimited variations of table structures, layouts, and content in advance. This preliminary data preparation eliminates the constraint of data scarcity during the actual model training phase, allowing the model to be trained on abundant synthetic data without needing to access limited real-world sources.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If synthetic training data is generated using traditional methods, then some training instances can be created, but the quality and diversity are insufficient for effective model training

Engineering Contradiction:
Improvetraining instance generation efficiencyVSAvoidtraining data quality
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system employs parameter changes by systematically varying HTML attributes such as table layouts, cell structures, text content, fonts, and styling during synthetic data generation. By changing these parameters across multiple generations, the pipeline produces highly diverse training instances that cover various table formats and complexities. This parameter variation ensures both high productivity in data generation and high reliability in training quality, as the model encounters diverse examples during training.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12494002B2Synthetic table generation pipeline for training deep table extraction models
Publication Date: 2025.12.09 ORACLE INT CORP
  • US12494002B2 patent drawing
  • US12494002B2 patent drawing
  • US12494002B2 patent drawing

AI summary

Techniques are described for HTML-based image generation. An example, method can include generating hypertext markup language (HTML) code for a table comprising a table structure of a set of rows and columns. The method can further include generating HTML code for a text to populate a cell of the table. The method can further include generating a rendered image of the table using the HTML code. The method can further include detecting a first pixel of the rendered image comprising the first color, and a second pixel of the rendered image comprising the second color. The method can further include detecting the text on the rendered image. The method can further include generating a bounding box, surrounding the detected text. The method can further include generating annotation comprising a bounding box parameter and a text parameter.