Reversible Data Anonymization for Machine-Learning Traceback Mitigation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data anonymization techniques are vulnerable to de-anonymization by third parties, as they fail to adequately protect identifying information, particularly when combined with publicly available data.

Innovation Solution

A system employing machine learning models and unique identifiers to anonymize data reversibly, using a reverse anonymity data sharing service that generates unique identifiers for sensitive information and employs machine learning to identify and mitigate potential de-anonymization risks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional data anonymization techniques are used, then data can be shared more easily, but the data becomes vulnerable to de-anonymization attacks

Engineering Contradiction:
Improvedata shareabilityVSAvoidanonymization security
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system performs preliminary actions by proactively generating synthetic data that mimics the statistical properties of real data before sharing occurs. This synthetic data serves as a preemptive protective measure, establishing anonymity safeguards in advance rather than attempting to protect anonymity after de-anonymization risks have materialized.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces synthetic data as an intermediary between real data and data sharing. Instead of directly sharing real data or completely anonymizing it, the system uses synthetic data as a mediator that preserves statistical utility while breaking direct links to identifiable individuals, thus resolving the contradiction between shareability and security.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Object-affected harmful factors

If identifying information is removed to protect privacy, then data utility for analysis is reduced

Engineering Contradiction:
Improveprivacy protectionVSAvoiddata utility
Core Design Contradiction:
Object-affected harmful factorsVSLoss of information

Solution Approach 1:

The system changes parameters by transforming real data into synthetic data with modified statistical parameters. The synthetic data preserves key statistical properties (means, variances, correlations) necessary for analysis while altering parameters that could lead back to identifiable individuals, thus protecting privacy without sacrificing data utility.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent discards identifying information during the synthetic data generation process while recovering and preserving statistical properties needed for analysis. The synthetic data retains the essential analytical value of the original data by maintaining statistical relationships and patterns, even though specific identifying details are discarded.

Inventive Principle:
Principle #34Discarding and recovering

Data Source

PatentUS12417313B1System and methods for data anonymity trackback mitigation
Publication Date: 2025.09.16 UNITED SERVICES AUTOMOBILE ASSOCIATION (USAA)
  • US12417313B1 patent drawing
  • US12417313B1 patent drawing
  • US12417313B1 patent drawing

AI summary

The techniques provided herein may be used to reversibly anonymize data and to mitigate possible de-anonymization the data by third parties. Specifically, a reverse anonymity data sharing service is used to anonymize personal data by replacing confidential information found in the data with unique identifiers. A reverse anonymity data store is used to store the unique identifiers with the information they replaced such that the anonymization may be reversed if needed. Machine learning models are used to prevent traceback, using publicly available data sources, of the anonymized data to it's the individual who generated the data or whom the data is about.