Identity Disclosure Risk Evaluation in Synthetic Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing identity disclosure assessment models for synthetic data are inadequate, as they assume partially synthetic data and do not consider all possible generalizations an adversary may use to identify individuals, leading to increased identification risks, especially with fully synthetic data where no direct mapping exists between synthetic and real records.

Innovation Solution

A method to determine identity disclosure risk in synthetic data by evaluating the probability of matching synthetic records with real records, using a generalization lattice to consider all possible generalizations of quasi-identifier variables, and adjusting for verification and error rates to assess the risk of identity disclosure and new information learning.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If previous identity disclosure assessment models are used, then assessment can be performed under assumption of partially synthetic data, but identification risk increases when applied to fully synthetic data where no direct mapping exists

Engineering Contradiction:
Improveapplicability to different synthetic data typesVSAvoididentification risk accuracy
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent changes the fundamental parameters of the assessment model by introducing a generalization lattice that operates on all possible generalizations of quasi-identifier variables rather than assuming direct mappings. This transforms the model from being applicable only to partially synthetic data to being applicable to both partially and fully synthetic data, resolving the contradiction between adaptability and reliability.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent segments the identification risk assessment into multiple levels of generalization using a lattice structure. Instead of treating synthetic data as a single homogeneous type, it divides the assessment into different generalization levels (from specific to general), allowing accurate risk evaluation across different synthetic data types including fully synthetic data where no direct mapping exists.

Inventive Principle:
Principle #1Segmentation

2Device complexity

If previous attack models are used that do not consider all possible generalizations, then assessment process is simpler, but identification risk substantially increases due to unconsidered attack vectors

Engineering Contradiction:
Improveassessment model complexityVSAvoididentification risk
Core Design Contradiction:
Device complexityVSObject-affected harmful factors

Solution Approach 1:

The patent creates a universal assessment model that handles all possible attack vectors through the generalization lattice. The lattice structure universally covers all generalizations of quasi-identifier variables, making the model multi-functional against different attack types (direct matching, generalization-based matching, subset matching) without requiring separate models for each attack vector.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent adds a new dimension to the assessment model by introducing the generalization lattice that operates across multiple levels of variable generalization. This transforms the assessment from a single-dimensional direct matching check to a multi-dimensional evaluation across different generalization levels, capturing attack vectors that were previously invisible.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Productivity

If direct matching approach is used between synthetic and real records, then assessment is computationally simpler, but accuracy decreases when no direct mapping exists in fully synthetic data

Engineering Contradiction:
Improveassessment computation efficiencyVSAvoidmatch probability accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent performs preliminary action by pre-computing the generalization lattice structure and all possible generalizations of quasi-identifier variables before conducting the actual matching assessment. This preliminary preparation enables efficient querying and accurate probability calculation during the assessment phase, resolving the contradiction between computational efficiency and accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces the generalization lattice as an intermediary structure between synthetic records and real sample records. Instead of directly comparing synthetic records with real records (which fails when no direct mapping exists), the lattice serves as a mediator that evaluates match probabilities across all possible generalizations, maintaining accuracy while enabling computation for fully synthetic data.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentEP3901791A1Systems and method for evaluating identity disclosure risks in synthetic personal data
Publication Date: 2021.10.27 REPLICA ANALYTICS
  • EP3901791A1 patent drawingFigure 1
  • EP3901791A1 patent drawingFigure 2
  • EP3901791A1 patent drawingFigure 3

AI summary

Although synthetic data synthesized from real sample data may not have a direct matching between synthetic data and individuals, there may still be a risk with identity disclosure. The identity disclosure risks associated with fully synthetic data may be assessed.