Synthetic Data Privacy Risk Quantification via Anonymeter

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing synthetic data methods fail to effectively quantify and mitigate residual privacy risks, particularly in high-accuracy applications, due to overfitting and the need for frequent recalibration, which can lead to identity disclosure and compliance issues with data protection regulations like GDPR.

Innovation Solution

Combining synthetic data with statutory pseudonymization and the Anonymeter framework, which jointly quantifies privacy risks through attack-based evaluations for singling out, linkability, and inference risks, providing a flexible protection level that balances usability and security.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If synthetic data is generated to preserve statistical properties of original data, then data utility is improved, but privacy risk increases due to overfitting and residual privacy leaks

Engineering Contradiction:
Improvedata utilityVSAvoidprivacy risk
Core Design Contradiction:
Loss of informationVSObject-affected harmful factors

Solution Approach 1:

The patent combines synthetic data generation with statutory pseudonymization techniques to create a hybrid approach. Synthetic data preserves statistical properties for utility, while pseudonymization adds layers of protection by replacing identifiers with pseudonyms, thereby reducing identity disclosure risks while maintaining analytical value

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces pseudonymized data as an intermediary layer between original sensitive data and synthetic data. This intermediary preserves statistical relationships needed for utility while breaking direct links to identifiable individuals, thus mitigating privacy risks associated with overfitting

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If synthetic data is used to share sensitive information, then data sharing capability is improved, but compliance with data protection regulations deteriorates due to residual privacy risks

Engineering Contradiction:
Improvedata sharing capabilityVSAvoidregulatory compliance
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent applies pseudonymization as a preliminary action before generating synthetic data. By replacing identifiers with pseudonyms in advance, the system proactively mitigates privacy risks before data sharing occurs, ensuring compliance with GDPR and other regulations while maintaining data sharing capabilities

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements a feedback mechanism through the Anonymeter framework that quantifies privacy risks in synthetic data. This feedback loop allows continuous monitoring and adjustment of privacy protection levels, ensuring regulatory compliance while optimizing data sharing utility

Inventive Principle:
Principle #23Feedback

3Measurement precision

If attack-based evaluations are implemented to quantify privacy risks, then privacy assessment accuracy is improved, but system complexity increases

Engineering Contradiction:
Improveprivacy assessment accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the privacy risk assessment into three distinct attack-based evaluations: singling out attacks, linkability attacks, and inference attacks. Each attack type targets specific privacy risks independently, allowing precise measurement of different privacy dimensions while managing system complexity through modular evaluation

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20230401336A1Enhanced Synthetic Data and a Unified Framework for Quantifying Privacy Risk in Synthetic Data
Publication Date: 2023.12.14 ANONOS INNOVATIONS LLC
  • US20230401336A1 patent drawing
  • US20230401336A1 patent drawing
  • US20230401336A1 patent drawing

AI summary

Embodiments disclosed herein improve data privacy and security by combining synthetic data and statutory pseudonymization to create protected data that is more effectively disconnected from the original source data. By bringing synthetic data and statutory pseudonymization techniques together, a flexible level of protection may be applied to data that strikes an appropriate balance between the ease of use of cleartext data and the aggressive protection of statutory pseudonymization. Further embodiments disclosed herein improve data privacy and security by providing a novel statistical framework that jointly quantifies different types of privacy risks in synthetic datasets and that includes attack-based evaluations for the singling out, linkability, and inference risks. According to other embodiments, the modular nature of the framework facilitates the future integration of new and potentially stronger attacks for evaluating privacy risks. The framework separates the evaluation of the success rate of the privacy attacks from the calculation of the reported privacy risks.