Domain-Guided Synthetic Data Generation for Private ML Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The lack of suitable training datasets for machine learning models, particularly in domains where specific labeled numerical datasets are scarce, small, or not publicly available, hinders innovation and experimentation, due to privacy concerns and legal issues with data sharing, leading to the generation of unstructured and ineffective training data.

Innovation Solution

A synthetic data generation system that incorporates domain expertise and knowledge to create ML-ready datasets using a guided user interface and rule engine, enabling the definition of input-output relationships and correlations, supported by a flexible data schema language, to generate datasets that approximate real-world patterns.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If real world data is collected and shared to create training datasets, then the quality and authenticity of training data is improved, but privacy concerns and legal issues arise that prevent data sharing

Engineering Contradiction:
Improvequality of training dataVSAvoidprivacy concerns and legal issues
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent creates synthetic copies of real-world data that replicate the statistical properties, relationships, and patterns of actual data without containing any real sensitive information. The synthetic data generation system produces artificial datasets that preserve the structural characteristics and correlations of real data while eliminating privacy risks and legal constraints associated with sharing actual customer or operational data.

Inventive Principle:
Principle #26Copying

2Productivity

If synthetic data is generated without domain knowledge, then data generation speed is improved, but the generated data lacks meaningful patterns and relationships

Engineering Contradiction:
Improvedata generation speedVSAvoidquality of synthetic data
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent introduces domain experts as intermediaries who provide knowledge about the specific industry or application domain. This domain knowledge is captured through interviews, surveys, or documentation and then integrated into the synthetic data generation process. The domain expert acts as a bridge between the data generation system and the real-world context, ensuring that the synthetic data incorporates meaningful patterns, relationships, and constraints specific to the target domain while maintaining rapid generation capabilities.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If large amounts of real data are collected to train accurate ML models, then model accuracy is improved, but the time and resources required for data collection increase

Engineering Contradiction:
Improvemodel accuracyVSAvoiddata collection time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary actions by pre-defining the statistical distributions, relationships, and constraints of the target domain through domain expert input before actual data generation begins. The synthetic data generation system is configured in advance with the necessary domain knowledge, statistical parameters, and relationship definitions, allowing it to rapidly produce high-quality training data without the time-consuming process of collecting and curating real-world data sets.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12386919B2Synthetic data generation for machine learning model simulation
Publication Date: 2025.08.12 SALESFORCE INC
  • US12386919B2 patent drawing
  • US12386919B2 patent drawing
  • US12386919B2 patent drawing

AI summary

A method and system for synthetic data generation are provided that receive a schema configuration file in a synthetic data set request from a client application, create a set of worker processes to generate the synthetic data set based on the schema configuration file, upload the generated synthetic data to an analytics platform, and enable the client application to utilize the generated synthetic data in prediction models for the analytics platform.