GAN Synthetic Data Generation for Accuracy and Privacy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems fail to generate synthetic data that accurately represents real-world computing data while protecting sensitive information, leading to insufficient accuracy and privacy concerns in machine learning applications.
Innovation Solution
Utilizing generative adversarial networks (GANs) to create synthetic data that meets both accuracy and privacy thresholds, ensuring that sensitive data is not retrievable and that machine learning models perform similarly to those trained on real data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If real data is used for training machine learning models, then model accuracy is improved, but sensitive information privacy is compromised
Solution Approach 1:
The patent creates synthetic data copies that replicate the statistical properties and patterns of real data without containing actual sensitive information. The synthetic data is generated through a pipeline that transforms real data into anonymized representations, allowing machine learning models to be trained on these copies rather than the original sensitive data, thus maintaining model accuracy while eliminating privacy risks
Solution Approach 2:
The patent introduces synthetic data as an intermediary between real data and machine learning models. Instead of directly training models on sensitive real data, the system uses synthetic data as a mediator that preserves the necessary patterns and relationships while removing sensitive information, enabling privacy-preserving machine learning
2Object-affected harmful factors
If synthetic data is generated to protect privacy, then data security is improved, but data accuracy is reduced
Solution Approach 1:
The patent implements feedback mechanisms where the synthetic data generation system continuously refines its output based on quality metrics and validation results. The system evaluates the synthetic data against real data distributions and adjusts the generation parameters to improve accuracy while maintaining privacy, ensuring that the synthetic data meets both security and accuracy requirements
Solution Approach 2:
The patent employs parameter transformation techniques where real data is converted to synthetic data through controlled parameter changes and transformations. The system adjusts various parameters during the synthesis process to preserve statistical properties, relationships, and patterns while removing sensitive information, thereby maintaining data accuracy despite the synthetic nature of the data
3Ease of operation
If existing data sharing methods are used, then data accessibility is improved, but sensitive information leakage occurs
Solution Approach 1:
The patent enables data sharing by providing copies of data in the form of synthetic datasets that replicate the essential characteristics of real data without containing sensitive information. Organizations can share and access synthetic data for analysis and modeling while the underlying sensitive information remains protected in the original data sources
Solution Approach 2:
The patent uses synthetic data as an intermediary medium for data sharing between organizations. Instead of directly sharing sensitive real data which causes information leakage, the system shares synthetic data that serves as a safe intermediary, allowing access and analysis while preventing sensitive information from leaving the controlled environment
Data Source
AI summary
A method generates synthetic data by a data collection system where the synthetic data meets a first threshold for accuracy and a second threshold for protecting sensitive data from recovery from the synthetic data. The method includes collecting data including sensitive data and non-sensitive data, executing a first machine learning model to generate the synthetic data from the collected data where the synthetic data meets the first threshold, executing a second machine learning model to update the synthetic data to meet the second threshold, checking whether the updated synthetic data meets the first threshold, releasing the updated synthetic data where the first threshold is met, and re-executing the first machine learning model and second machine learning model to update the synthetic data where the first threshold is not met during the checking.


