Conditional Synthetic Dataset Generation for Privacy-Preserving AI Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing anonymization techniques for personal information in large datasets result in information loss, making it challenging to use such data for AI training without compromising privacy.
Innovation Solution
A method for generating synthetic datasets by converting original data items using condition data items as conditions, employing a conditional generative model to preserve data characteristics while anonymizing personal information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If anonymization techniques are used to protect personal information in large datasets, then privacy is protected, but information loss occurs making the data less useful for AI training
Solution Approach 1:
The patent creates synthetic copies of real data that preserve the statistical characteristics and relationships of the original data while containing no actual personal information. The generative model learns from real data and produces synthetic instances that replicate data patterns, distributions, and correlations without copying sensitive information directly.
Solution Approach 2:
The patent transforms real data into synthetic data by changing fundamental parameters - replacing actual personal information with generated values while maintaining the structural and statistical properties. This parameter transformation allows the data to retain utility for AI training while eliminating privacy risks associated with the original information.
2Quantity of substance
If multiple original datasets from different sources are integrated, then data comprehensiveness improves, but data quality and consistency become harder to maintain
Solution Approach 1:
The patent processes each source dataset independently through the generative model, creating separate synthetic datasets that preserve the characteristics of each source. This segmentation approach allows independent quality control for each source while maintaining the ability to integrate them later, as each synthetic dataset retains its own data quality properties.
Solution Approach 2:
The patent creates a universal synthetic data generation framework that can process multiple different source datasets through the same generative model architecture. This multi-functional approach allows consistent quality standards to be applied across diverse data sources while preserving the unique characteristics of each source in its respective synthetic output.
Data Source
AI summary
Proposed is a method for generating a synthetic dataset including a plurality of data items. The method includes receiving, by a data processing device, an original dataset including a plurality of original data items, selecting, by the data processing device, at least one original data item among the plurality of original data items as a condition data item, converting, by the data processing device, an original data item of the remaining original data items, excluding the condition data item, into a first synthetic data item using the condition data item as a condition and converting, by the data processing device, an original data item of the unsynthesized original data items among the plurality of data items, into a second synthetic data item using at least one of the previously generated synthetic data items as a new condition data item.


