Synthetic Data Generation via ML Constraints
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data anonymization techniques are insufficient in preventing malicious entities from re-identifying individuals even after personally identifiable information (PII) removal and de-identification, leading to privacy risks and leakage, and there is a need for improved methods to generate synthetic data that accurately reflects real-world trends and patterns.
Innovation Solution
The use of a machine learning (ML) model to generate synthetic data by reproducing identified attributes from microdata while applying constraints to prevent rare attribute combinations, along with a user interface for filtering and comparing synthesized data to pre-computed aggregate counts to ensure accuracy and privacy preservation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional de-identification techniques are used to remove PII, then data sharing capability is improved, but privacy protection deteriorates because malicious entities can still re-identify individuals through background information correlation
Solution Approach 1:
The patent generates synthetic data that copies the statistical properties, distributions, and relationships of the original microdata without copying actual individual records. This synthetic replica allows data sharing while preventing re-identification because the synthetic data contains no real PII or identifiable information about specific individuals.
Solution Approach 2:
The patent introduces synthetic data as an intermediary between the original microdata and any analysis needs. This intermediary preserves the analytical utility of the data while breaking the direct link to individual identities, thereby preventing privacy leakage while maintaining data sharing capability.
2Object-affected harmful factors
If aggregate data is used instead of microdata, then privacy protection is improved, but measurement precision deteriorates because aggregate statistics cannot capture individual-level patterns and trends
Solution Approach 1:
The patent applies different levels of data detail to different needs: synthetic data records maintain individual-level granularity for precise analysis, while the synthetic generation process ensures privacy protection at the population level. This allows both individual-level precision and population-level privacy protection simultaneously.
3Productivity
If de-identified microdata is released for analysis, then data utility is improved, but reliability deteriorates because the data may not accurately reflect real-world distributions due to over-anonymization
Solution Approach 1:
The patent uses the original microdata's statistical properties as feedback to guide the synthetic data generation process. By matching distributions, correlations, and patterns from the original data, the synthetic data maintains reliability and representativeness while preserving privacy through the synthetic nature of the records.
Data Source
AI summary
Techniques for synthesizing and analyzing data are disclosed. A ML model anonymizes microdata to generate synthesized data. This anonymizing is performed by reproducing attributes identified within microdata and by applying constraints to prevent rare attribute combinations from being reproduced in the synthesized data. User input selects attributes to filter the synthesized data, thereby generating a subset of records. A UI displays a synthesized aggregate count representing how many records are in the subset. Pre-computed aggregate counts are accessed to indicate how many records in the microdata embody certain attributes. Based on the user input, there is an attempt to identify a particular count from the pre-computed aggregate counts. This count reflects how many records of the microdata would remain if the selected attributes were used to filter the microdata. That count is displayed along with the synthesized aggregate count. The two counts are juxtaposed next to one another.


