Synthetic Data Generation Using Correlated Column Modeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Machine learning models often perform poorly due to limited or inappropriate data, leading to inaccurate predictions and increased time and expense in data procurement, which can result in competitive disadvantages and delays in service release.

Innovation Solution

A data management system generates synthetic data using a large language model to create correlated and uncorrelated data columns based on source data, merging them to train machine learning models effectively.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If real-world data is procured to train machine learning models, then model accuracy and performance are improved, but time and expense increase significantly

Engineering Contradiction:
Improvemodel accuracyVSAvoiddata procurement time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent generates synthetic data that copies the statistical properties, correlations, and patterns of real-world data without requiring actual data collection. The system creates artificial datasets that replicate the characteristics of target domain data, enabling model training to proceed without time-consuming real data procurement while maintaining training effectiveness

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs preliminary data preparation by generating synthetic training data in advance before actual model development begins. By pre-generating realistic synthetic datasets that capture domain-specific patterns and correlations, the approach eliminates delays that would occur during later data collection phases

Inventive Principle:
Principle #10Preliminary action

2Reliability

If real-world data is procured to train machine learning models, then model performance is improved, but expense increases significantly

Engineering Contradiction:
Improvemodel performanceVSAvoiddata procurement cost
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The patent creates synthetic data copies that replicate the essential characteristics and statistical properties of expensive real-world data. By generating artificial datasets that maintain realistic patterns, correlations, and distributions, the system eliminates or reduces the need to purchase or collect costly real data while preserving model training quality

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system generates inexpensive synthetic data that can be created on-demand without the high costs associated with real data acquisition. The synthetic datasets serve as disposable training materials that can be regenerated as needed, replacing expensive real data procurement cycles

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

3Loss of time

If limited available data is used for training, then data procurement time and expense are reduced, but model accuracy decreases due to blind spots

Engineering Contradiction:
Improvedata procurement timeVSAvoidmodel accuracy
Core Design Contradiction:
Loss of timeVSReliability

Solution Approach 1:

The patent uses synthetic data copying to replicate not just the structure but also the diverse scenarios and edge cases present in comprehensive real-world datasets. The generated synthetic data includes rare events and boundary conditions that might be missing from limited available data, thereby eliminating blind spots without requiring extensive real data collection

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system transforms limited available data into expanded synthetic datasets by applying parameter variations and transformations. This allows the generation of diverse training examples from limited source data, covering a broader range of scenarios and reducing model blind spots while maintaining the time efficiency of using existing data sources

Inventive Principle:
Principle #35Parameter changes

4Quantity of substance

If data generated for other scenarios in other domains is used, then data availability increases, but model training effectiveness decreases for the target domain

Engineering Contradiction:
Improvedata availabilityVSAvoidmodel training effectiveness
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent applies local quality by generating synthetic data with domain-specific characteristics tailored to the target domain's unique patterns, correlations, and requirements. Rather than using generic cross-domain data, the system creates locally-optimized synthetic datasets that reflect the specific context and nuances of the target application area, ensuring training effectiveness

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20250390720A1Guided intelligent synthetic data generation
Publication Date: 2025.12.25 ORACLE INT CORP
  • US20250390720A1 patent drawing
  • US20250390720A1 patent drawing
  • US20250390720A1 patent drawing

AI summary

A data management system receives a request to generate synthetic data based on source data, and determines a set of columns of the source data that satisfy a correlation condition and at least one other column of the source data that does not satisfy the correlation condition. The data management system prompts a large language model to generate synthetic data for the set of columns based at least in part on first source data values for the set of columns, and generates synthetic data for the at least one other column based at least in part on a distribution of second source data values for the at least one other column. The data management system merges the synthetic data to generate a resulting synthetic set of data of the plurality of columns. The data management system stores the resulting synthetic set of data in a repository of training data, and uses the training data to train a machine learning model.