Synthetic Data Generation via GANs for Privacy-Preserving Analytics
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data from disparate sources is often siloed and cannot be shared due to legal restrictions or privacy concerns, limiting the ability of systems to combine data for improved predictive analytics without actually transmitting the data between systems.
Innovation Solution
The use of generative adversarial networks (GANs) to generate statistically-representative sample data, allowing systems to train and exchange sample-data generators that mimic the features of other systems' data without transmitting actual data, thereby enabling improved predictive models without data sharing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If data is shared between systems to improve predictive analytics, then predictive accuracy is improved, but data privacy and security are compromised
Solution Approach 1:
The patent creates synthetic copies of data that replicate the statistical properties and patterns of original data without containing actual sensitive information. A generative model learns from the original data and generates synthetic samples that preserve analytical value while eliminating privacy risks, allowing multiple systems to use identical synthetic copies without compromising source data security
Solution Approach 2:
The patent introduces a synthetic data generation system as an intermediary between source systems and analytical systems. This intermediary creates a bridge by generating synthetic data that mediates the information flow, allowing predictive analytics to proceed without direct access to sensitive original data, thus resolving the contradiction between data utilization and privacy protection
2Object-affected harmful factors
If data is siloed to maintain privacy, then data security is improved, but predictive analytics capability deteriorates
Solution Approach 1:
Instead of sharing original data across silos, the patent generates synthetic copies within each silo that capture the essential patterns and relationships needed for predictive analytics. Each system maintains its data security while possessing synthetic data copies that enable analytical capabilities comparable to having access to combined data
Solution Approach 2:
The patent transforms the state of data from original sensitive information to synthetic parameterized representations. By changing the parameters from actual data values to statistical distributions and patterns, the system maintains security while preserving analytical utility, allowing siloed systems to perform predictive analytics without breaking data isolation
3Productivity
If actual data is transmitted between systems, then data utilization is improved, but data transmission security requirements increase
Solution Approach 1:
The patent replaces data transmission with synthetic data generation and exchange. Instead of transmitting actual data between systems requiring secure channels, systems transmit or share synthetic copies that contain no sensitive information, dramatically reducing transmission security requirements while maintaining data utilization benefits
Solution Approach 2:
The patent extracts only the essential statistical properties and patterns from original data, separating them from the actual sensitive content. This extraction creates synthetic data that contains only the analytical value needed for productivity improvement, eliminating the need for complex security infrastructure required for transmitting complete original datasets
Data Source
AI summary
Systems and methods for statistically-representative sample data generation are disclosed. For example, a sample-data generator and/or a data discriminator may be received by a system, which may utilize the sample-data generator to generate sample data. The data discriminator may be utilized to train the sample-data generator until the data discriminator cannot discriminate between data received from the sample-data generator and data received by a database associated with the system. The trained sample-data generator may be sent to other systems, which may generate and utilize, such as for prediction model training, statistically-representative sample data generated by the trained sample-data generator.


