Synthetic Data Generator for Secure Multi-Organization Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data analysis methods face challenges in securely and efficiently generating synthetic data from combined datasets due to regulatory constraints and the complexity of de-identification processes, particularly in sharing personal information across organizations.
Innovation Solution
An apparatus and method for generating synthetic data using a synthetic data generator that decrypts and combines ciphertext from multiple data providing apparatuses within a trusted execution environment, employing homomorphic encryption and machine learning-based models to produce data satisfying differential privacy, thereby enabling secure and efficient data analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If de-identification techniques are used to protect data privacy, then data sharing becomes possible, but the risk of exposure increases after combining data
Solution Approach 1:
The patent creates synthetic copies of original data that preserve statistical characteristics and analytical value while containing no actual personal information. The synthetic data generator produces artificial datasets that mimic the structure and patterns of real data without replicating sensitive information, thereby enabling data sharing without exposure risks.
Solution Approach 2:
The patent introduces synthetic data as an intermediary between original sensitive data and analysis needs. Instead of directly sharing or combining original data, the system generates synthetic representations that serve as a safe medium for data analysis, eliminating the need for direct data exposure while maintaining analytical utility.
2Reliability
If encrypted data is used for data privacy protection, then data security is improved, but a method for satisfying technology has to be devised according to analysis query, increasing time and complexity
Solution Approach 1:
The patent performs data synthesis in advance, generating synthetic datasets before any analysis queries are executed. This preliminary generation of synthetic data eliminates the need to devise complex decryption and re-encryption methods for each analysis query, as the synthetic data is already prepared and can be freely analyzed without security constraints.
Solution Approach 2:
The patent extracts the essential statistical characteristics and patterns from original data while leaving behind the sensitive personal information. By taking out only the analytical value and creating synthetic representations, the system separates the useful analytical content from the sensitive data, allowing straightforward analysis without complex security protocols.
3Productivity
If multiple organizations share original data for data combining, then analysis performance is improved, but regulatory constraints prevent data sharing
Solution Approach 1:
The patent enables each organization to generate synthetic copies of their own data that preserve the statistical properties needed for analysis. These synthetic copies can be freely shared and combined across organizations without regulatory barriers, allowing multi-organization collaboration while maintaining data sovereignty and compliance with privacy regulations.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Disclosed are an apparatus and method for generating synthetic data. The apparatus for generating synthetic data according to an embodiment includes a synthetic data generator configured to generate synthetic data corresponding to combined data obtained by combining original data held by each of a plurality of data providing apparatuses, and a synthetic data provider configured to provide the synthetic data to a data using apparatus.