Synthetic Data Generation Using PGM and ABM for Privacy-Safe ML Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine learning model training is hindered by the lack of sufficient data, particularly in scenarios where data is protected by privacy regulations or does not cover rare events, and conventional generative models are difficult for developers to use effectively.
Innovation Solution
The generation of synthetic datasets using probabilistic graphical models (PGM) and agent-based models (ABM) to create factual and counterfactual data, allowing for the customization of data distributions and correlations, and the use of generative models to scrub and validate data without revealing underlying content.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If real data is used for training machine learning models, then the models can learn from actual scenarios, but privacy regulations and security concerns prevent access to sufficient data volumes
Solution Approach 1:
The patent creates synthetic copies of real data through generative models that replicate the statistical properties and relationships of actual data without containing sensitive information. These synthetic datasets serve as substitutes for real data, enabling model training while preserving privacy and security requirements.
Solution Approach 2:
The patent transforms real data into synthetic data by changing key parameters such as privacy protection levels, data sensitivity classifications, and generative model configurations. This allows the same underlying data to be used in multiple contexts with different privacy guarantees and training requirements.
2Quantity of substance
If conventional generative models are used to create synthetic data, then data volume can be increased, but the models are difficult for developers to use and modify
Solution Approach 1:
The patent creates a universal platform that handles multiple functions including data generation, validation, privacy assessment, and model training within a single system. This multi-functional approach eliminates the need for developers to separately manage complex generative models, validation pipelines, and privacy compliance checks.
Solution Approach 2:
The system performs automatic validation and quality assessment of synthetic data without requiring manual intervention from developers. The generative models self-adjust parameters based on validation feedback, and the system automatically ensures privacy compliance, reducing the operational burden on developers.
3Object-affected harmful factors
If data scrubbing processes are applied to protect privacy, then security concerns are reduced, but the process is time-consuming and may inadvertently release real data
Solution Approach 1:
The patent performs privacy protection and data validation in advance during the synthetic data generation process itself, rather than as a separate post-processing step. The generative models are trained to inherently preserve privacy while maintaining data utility, eliminating the need for time-consuming scrubbing operations afterward.
Solution Approach 2:
The system introduces synthetic data as an intermediary between real sensitive data and machine learning models. This intermediary layer provides the necessary privacy protection while maintaining data functionality, avoiding the need for direct manipulation of sensitive real data through risky scrubbing processes.
4Measurement precision
If machine learning models are trained only on factual data from existing environments, then the training data is accurate to real scenarios, but the models cannot handle rare or hypothetical scenarios
Solution Approach 1:
The patent enables dynamic switching between factual and counterfactual data generation modes. The system can adaptively generate synthetic data that matches real-world statistics when accuracy is needed, or introduces hypothetical variations when scenario coverage is required, making the training data flexible and adaptable to different model needs.
Solution Approach 2:
The patent creates composite training datasets that combine factual real-world data with counterfactual synthetic data. This composite approach integrates the accuracy benefits of real data with the scenario-diversity benefits of synthetic data, producing training sets that are both accurate and versatile.
Data Source
AI summary
A system, method, and computer-readable medium for generating factual and/or counterfactual data are described. This may have the effect of improving the complexity of data available for training machine learning models. The models may include, but not limited to, a probabilistic graphical model (PGM) and/or an agent-based model (ABM). Further aspects may provide for scrubbing actual data to create a data model that does not reveal the content of the underlying source data. Yet further aspects may provide for validating a data model.


