Synthetic Dataset Generation via Data Perturbation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data anonymization techniques, such as masking applications, are limited in effectively protecting confidential data items as they can be reconstructed from non-confidential data items and do not comprehensively anonymize datasets, especially when all data items are confidential, hindering analysis and sharing of datasets.
Innovation Solution
A computer-implemented method generates a new dataset by perturbing data items to create test datasets that are characterized by property values substantially similar to the original dataset, ensuring that the new dataset conveys aspects of the original dataset without revealing confidential data items, using a dataset generation application with an iteration controller, perturbation engine, and consistency engine.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If masking applications are used to remove personally-identifying data items, then patient privacy is protected, but the remaining data items can still be used to reconstruct confidential information
Solution Approach 1:
The patent creates synthetic copies of datasets that replicate statistical properties and relationships without containing actual confidential data. Instead of merely masking sensitive fields, the system generates artificial data items that preserve analytical value while eliminating reconstructability of original confidential information through perturbation and synthetic generation techniques
Solution Approach 2:
The patent applies parameter changes by perturbing data items through statistical transformations that modify individual data points while preserving overall dataset characteristics. This involves changing parameters such as adding noise, applying differential privacy mechanisms, and transforming data distributions to prevent reconstruction while maintaining analytical utility
2Reliability
If masking applications remove all confidential data items, then confidentiality is maintained, but the dataset becomes useless for analysis
Solution Approach 1:
The patent creates synthetic copies of datasets that replicate statistical properties and relationships without containing actual confidential data. Instead of merely masking sensitive fields, the system generates artificial data items that preserve analytical value while eliminating reconstructability of original confidential information through perturbation and synthetic generation techniques
Solution Approach 2:
The patent applies parameter changes by perturbing data items through statistical transformations that modify individual data points while preserving overall dataset characteristics. This involves changing parameters such as adding noise, applying differential privacy mechanisms, and transforming data distributions to prevent reconstruction while maintaining analytical utility
3Reliability
If masking applications are fine-tuned for specific data types, then those data types are anonymized, but other data types remain vulnerable
Solution Approach 1:
The patent implements a universal anonymization framework that can process multiple data types simultaneously through a single synthetic data generation system. The perturbation mechanisms and synthetic generation techniques are designed to work across diverse data structures, relationships, and formats, providing comprehensive anonymization coverage without requiring separate masking applications for each data type
Data Source
AI summary
In various embodiments, a dataset generation application generates a new dataset based on an original dataset. The dataset generation engine perturbs a first data item included in the original dataset to generate a second data item. The dataset generation application then generates a test dataset based on the original dataset and the second data item. The test dataset includes the second data item instead of the first data item. Subsequently, the dataset generation application determines that the test dataset is characterized by a first property value that is substantially similar to a second property value that characterizes the original dataset. The first property value and the second property value are associated with the same property. Finally, the dataset generation application generates a new dataset based on the test dataset. The new dataset conveys aspect(s) of the original dataset without revealing the first data item.


