Synthetic Data Generation Using PGM and ABM for Privacy-Safe ML Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning model training is hindered by the lack of sufficient data, particularly in scenarios where data is protected by privacy regulations or does not cover rare events, and conventional generative models are difficult for developers to use effectively.

Innovation Solution

The generation of synthetic datasets using probabilistic graphical models (PGM) and agent-based models (ABM) to create factual and counterfactual data, allowing for the customization of data distributions and correlations, and the use of generative models to scrub and validate data without revealing underlying content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If real data is used for training machine learning models, then the models can learn from actual scenarios, but privacy regulations and security concerns prevent access to sufficient data volumes

Engineering Contradiction:
Improvemodel training accuracyVSAvoiddata volume
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent creates synthetic copies of real data through generative models that replicate the statistical properties and relationships of actual data without containing sensitive information. These synthetic datasets serve as substitutes for real data, enabling model training while preserving privacy and security requirements.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent transforms real data into synthetic data by changing key parameters such as privacy protection levels, data sensitivity classifications, and generative model configurations. This allows the same underlying data to be used in multiple contexts with different privacy guarantees and training requirements.

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If conventional generative models are used to create synthetic data, then data volume can be increased, but the models are difficult for developers to use and modify

Engineering Contradiction:
Improvesynthetic data volumeVSAvoiddeveloper usability
Core Design Contradiction:
Quantity of substanceVSEase of operation

Solution Approach 1:

The patent creates a universal platform that handles multiple functions including data generation, validation, privacy assessment, and model training within a single system. This multi-functional approach eliminates the need for developers to separately manage complex generative models, validation pipelines, and privacy compliance checks.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system performs automatic validation and quality assessment of synthetic data without requiring manual intervention from developers. The generative models self-adjust parameters based on validation feedback, and the system automatically ensures privacy compliance, reducing the operational burden on developers.

Inventive Principle:
Principle #25Self-service

3Object-affected harmful factors

If data scrubbing processes are applied to protect privacy, then security concerns are reduced, but the process is time-consuming and may inadvertently release real data

Engineering Contradiction:
Improveprivacy security riskVSAvoiddata preparation time
Core Design Contradiction:
Object-affected harmful factorsVSLoss of time

Solution Approach 1:

The patent performs privacy protection and data validation in advance during the synthetic data generation process itself, rather than as a separate post-processing step. The generative models are trained to inherently preserve privacy while maintaining data utility, eliminating the need for time-consuming scrubbing operations afterward.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system introduces synthetic data as an intermediary between real sensitive data and machine learning models. This intermediary layer provides the necessary privacy protection while maintaining data functionality, avoiding the need for direct manipulation of sensitive real data through risky scrubbing processes.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Measurement precision

If machine learning models are trained only on factual data from existing environments, then the training data is accurate to real scenarios, but the models cannot handle rare or hypothetical scenarios

Engineering Contradiction:
Improvedata accuracy to real scenariosVSAvoidmodel scenario coverage
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent enables dynamic switching between factual and counterfactual data generation modes. The system can adaptively generate synthetic data that matches real-world statistics when accuracy is needed, or introduces hypothetical variations when scenario coverage is required, making the training data flexible and adaptable to different model needs.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent creates composite training datasets that combine factual real-world data with counterfactual synthetic data. This composite approach integrates the accuracy benefits of real data with the scenario-diversity benefits of synthetic data, producing training sets that are both accurate and versatile.

Inventive Principle:
Principle #40Composite materials

Data Source

PatentUS12566950B2Generation of secure synthetic data based on true-source datasets
Publication Date: 2026.03.03 CAPITAL ONE SERVICES LLC
  • US12566950B2 patent drawing
  • US12566950B2 patent drawing
  • US12566950B2 patent drawing

AI summary

A system, method, and computer-readable medium for generating factual and/or counterfactual data are described. This may have the effect of improving the complexity of data available for training machine learning models. The models may include, but not limited to, a probabilistic graphical model (PGM) and/or an agent-based model (ABM). Further aspects may provide for scrubbing actual data to create a data model that does not reveal the content of the underlying source data. Yet further aspects may provide for validating a data model.