Domain-Guided Synthetic Data Generation for Private ML Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The lack of suitable training datasets for machine learning models, particularly in domains where specific labeled numerical datasets are scarce, small, or not publicly available, hinders innovation and experimentation, due to privacy concerns and legal issues with data sharing, leading to the generation of unstructured and ineffective training data.
Innovation Solution
A synthetic data generation system that incorporates domain expertise and knowledge to create ML-ready datasets using a guided user interface and rule engine, enabling the definition of input-output relationships and correlations, supported by a flexible data schema language, to generate datasets that approximate real-world patterns.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If real world data is collected and shared to create training datasets, then the quality and authenticity of training data is improved, but privacy concerns and legal issues arise that prevent data sharing
Solution Approach 1:
The patent creates synthetic copies of real-world data that replicate the statistical properties, relationships, and patterns of actual data without containing any real sensitive information. The synthetic data generation system produces artificial datasets that preserve the structural characteristics and correlations of real data while eliminating privacy risks and legal constraints associated with sharing actual customer or operational data.
2Productivity
If synthetic data is generated without domain knowledge, then data generation speed is improved, but the generated data lacks meaningful patterns and relationships
Solution Approach 1:
The patent introduces domain experts as intermediaries who provide knowledge about the specific industry or application domain. This domain knowledge is captured through interviews, surveys, or documentation and then integrated into the synthetic data generation process. The domain expert acts as a bridge between the data generation system and the real-world context, ensuring that the synthetic data incorporates meaningful patterns, relationships, and constraints specific to the target domain while maintaining rapid generation capabilities.
3Measurement precision
If large amounts of real data are collected to train accurate ML models, then model accuracy is improved, but the time and resources required for data collection increase
Solution Approach 1:
The patent performs preliminary actions by pre-defining the statistical distributions, relationships, and constraints of the target domain through domain expert input before actual data generation begins. The synthetic data generation system is configured in advance with the necessary domain knowledge, statistical parameters, and relationship definitions, allowing it to rapidly produce high-quality training data without the time-consuming process of collecting and curating real-world data sets.
Data Source
AI summary
A method and system for synthetic data generation are provided that receive a schema configuration file in a synthetic data set request from a client application, create a set of worker processes to generate the synthetic data set based on the schema configuration file, upload the generated synthetic data to an analytics platform, and enable the client application to utilize the generated synthetic data in prediction models for the analytics platform.


