Identity Disclosure Risk Evaluation in Synthetic Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing identity disclosure assessment models for synthetic data are inadequate, as they assume partially synthetic data and do not consider all possible generalizations an adversary may use to identify individuals, leading to increased identification risks, especially with fully synthetic data where no direct mapping exists between synthetic and real records.
Innovation Solution
A method to determine identity disclosure risk in synthetic data by evaluating the probability of matching synthetic records with real records, using a generalization lattice to consider all possible generalizations of quasi-identifier variables, and adjusting for verification and error rates to assess the risk of identity disclosure and new information learning.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If previous identity disclosure assessment models are used, then assessment can be performed under assumption of partially synthetic data, but identification risk increases when applied to fully synthetic data where no direct mapping exists
Solution Approach 1:
The patent changes the fundamental parameters of the assessment model by introducing a generalization lattice that operates on all possible generalizations of quasi-identifier variables rather than assuming direct mappings. This transforms the model from being applicable only to partially synthetic data to being applicable to both partially and fully synthetic data, resolving the contradiction between adaptability and reliability.
Solution Approach 2:
The patent segments the identification risk assessment into multiple levels of generalization using a lattice structure. Instead of treating synthetic data as a single homogeneous type, it divides the assessment into different generalization levels (from specific to general), allowing accurate risk evaluation across different synthetic data types including fully synthetic data where no direct mapping exists.
2Device complexity
If previous attack models are used that do not consider all possible generalizations, then assessment process is simpler, but identification risk substantially increases due to unconsidered attack vectors
Solution Approach 1:
The patent creates a universal assessment model that handles all possible attack vectors through the generalization lattice. The lattice structure universally covers all generalizations of quasi-identifier variables, making the model multi-functional against different attack types (direct matching, generalization-based matching, subset matching) without requiring separate models for each attack vector.
Solution Approach 2:
The patent adds a new dimension to the assessment model by introducing the generalization lattice that operates across multiple levels of variable generalization. This transforms the assessment from a single-dimensional direct matching check to a multi-dimensional evaluation across different generalization levels, capturing attack vectors that were previously invisible.
3Productivity
If direct matching approach is used between synthetic and real records, then assessment is computationally simpler, but accuracy decreases when no direct mapping exists in fully synthetic data
Solution Approach 1:
The patent performs preliminary action by pre-computing the generalization lattice structure and all possible generalizations of quasi-identifier variables before conducting the actual matching assessment. This preliminary preparation enables efficient querying and accurate probability calculation during the assessment phase, resolving the contradiction between computational efficiency and accuracy.
Solution Approach 2:
The patent introduces the generalization lattice as an intermediary structure between synthetic records and real sample records. Instead of directly comparing synthetic records with real records (which fails when no direct mapping exists), the lattice serves as a mediator that evaluates match probabilities across all possible generalizations, maintaining accuracy while enabling computation for fully synthetic data.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Although synthetic data synthesized from real sample data may not have a direct matching between synthetic data and individuals, there may still be a risk with identity disclosure. The identity disclosure risks associated with fully synthetic data may be assessed.