Synthetic Data Calibration via Latent Space Sampling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge lies in securely sharing sensitive data for analysis while maintaining confidentiality, as unauthorized access can lead to fraud and identity theft, and existing methods fail to preserve data security and multivariate structure.
Innovation Solution
The method involves generating synthetic data using generative models, specifically autoencoders, to create a calibrated latent space that mirrors the structure of remote source data, allowing for secure data analysis without sharing the original data, by optimizing weighted selection probabilities and latent space augmentation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If sensitive source data is shared for analysis, then analytical insights can be obtained, but data security and confidentiality are compromised
Solution Approach 1:
The patent generates synthetic data that replicates the multivariate structure, statistical properties, and relationships of the source data without containing actual sensitive information. This copy enables analytical insights while eliminating security risks associated with sharing real data.
Solution Approach 2:
The synthetic data acts as an intermediary between the source data and analytical processes. It transfers the essential structural and statistical properties needed for analysis while blocking access to actual sensitive information, thus mediating between data utility and security requirements.
2Object-affected harmful factors
If synthetic data is generated without calibration, then data security is maintained, but the multivariate structure and statistical properties of source data are not preserved
Solution Approach 1:
The calibration process adjusts key parameters of the synthetic data including mean, standard deviation, correlation coefficients, and distribution shapes to match those of the source data. This ensures synthetic data preserves multivariate structure and statistical properties while maintaining security.
Solution Approach 2:
The system uses iterative feedback loops where synthetic data is generated, compared against source data statistics, and refined through recalibration. This continuous feedback ensures the synthetic data progressively converges to match the multivariate structure and statistical properties of the source data.
3Loss of information
If source data is transferred physically for analysis, then comprehensive analysis can be performed, but data exposure and security risks increase
Solution Approach 1:
Instead of transferring actual source data, the system creates and transfers synthetic data copies that contain all necessary structural and statistical information for analysis. This eliminates physical data transfer risks while maintaining analytical completeness.
4Loss of information
If source data is transferred physically for analysis, then comprehensive analysis can be performed, but data exposure and security risks increase
Solution Approach 1:
The system generates synthetic data copies that replicate the multivariate structure and statistical properties of source data, enabling comprehensive analysis without exposing actual sensitive information. This copying approach maintains analytical completeness while eliminating data exposure risks.
Solution Approach 2:
Synthetic data serves as an intermediary that enables complete analytical processing without direct access to source data. It mediates between the need for comprehensive analysis and the requirement to prevent data exposure, allowing analysts to work with data that has identical structural properties but contains no actual sensitive information.
Data Source
AI summary
A method, a system, and a computer program product for calibrating synthetic data. A synthetic data is generated based on one or more source data using one or more generative models. The generative models are used to generate a latent space based on one or more source data. One or more latent space vectors associated with the generated latent space are determined in accordance with one or more data profiles associated with the one or more source data. The latent space vectors associated with the generated latent space are sampled. Based on the sampling, an optimized synthetic data is generated by comparing the sampled latent space vectors with one or more baseline data associated with one or more data profiles.


