Synthetic Data Generation Using Correlated Column Modeling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Machine learning models often perform poorly due to limited or inappropriate data, leading to inaccurate predictions and increased time and expense in data procurement, which can result in competitive disadvantages and delays in service release.
Innovation Solution
A data management system generates synthetic data using a large language model to create correlated and uncorrelated data columns based on source data, merging them to train machine learning models effectively.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If real-world data is procured to train machine learning models, then model accuracy and performance are improved, but time and expense increase significantly
Solution Approach 1:
The patent generates synthetic data that copies the statistical properties, correlations, and patterns of real-world data without requiring actual data collection. The system creates artificial datasets that replicate the characteristics of target domain data, enabling model training to proceed without time-consuming real data procurement while maintaining training effectiveness
Solution Approach 2:
The system performs preliminary data preparation by generating synthetic training data in advance before actual model development begins. By pre-generating realistic synthetic datasets that capture domain-specific patterns and correlations, the approach eliminates delays that would occur during later data collection phases
2Reliability
If real-world data is procured to train machine learning models, then model performance is improved, but expense increases significantly
Solution Approach 1:
The patent creates synthetic data copies that replicate the essential characteristics and statistical properties of expensive real-world data. By generating artificial datasets that maintain realistic patterns, correlations, and distributions, the system eliminates or reduces the need to purchase or collect costly real data while preserving model training quality
Solution Approach 2:
The system generates inexpensive synthetic data that can be created on-demand without the high costs associated with real data acquisition. The synthetic datasets serve as disposable training materials that can be regenerated as needed, replacing expensive real data procurement cycles
3Loss of time
If limited available data is used for training, then data procurement time and expense are reduced, but model accuracy decreases due to blind spots
Solution Approach 1:
The patent uses synthetic data copying to replicate not just the structure but also the diverse scenarios and edge cases present in comprehensive real-world datasets. The generated synthetic data includes rare events and boundary conditions that might be missing from limited available data, thereby eliminating blind spots without requiring extensive real data collection
Solution Approach 2:
The system transforms limited available data into expanded synthetic datasets by applying parameter variations and transformations. This allows the generation of diverse training examples from limited source data, covering a broader range of scenarios and reducing model blind spots while maintaining the time efficiency of using existing data sources
4Quantity of substance
If data generated for other scenarios in other domains is used, then data availability increases, but model training effectiveness decreases for the target domain
Solution Approach 1:
The patent applies local quality by generating synthetic data with domain-specific characteristics tailored to the target domain's unique patterns, correlations, and requirements. Rather than using generic cross-domain data, the system creates locally-optimized synthetic datasets that reflect the specific context and nuances of the target application area, ensuring training effectiveness
Data Source
AI summary
A data management system receives a request to generate synthetic data based on source data, and determines a set of columns of the source data that satisfy a correlation condition and at least one other column of the source data that does not satisfy the correlation condition. The data management system prompts a large language model to generate synthetic data for the set of columns based at least in part on first source data values for the set of columns, and generates synthetic data for the at least one other column based at least in part on a distribution of second source data values for the at least one other column. The data management system merges the synthetic data to generate a resulting synthetic set of data of the plurality of columns. The data management system stores the resulting synthetic set of data in a repository of training data, and uses the training data to train a machine learning model.


