Longitudinal Dataset Re-identification Risk Measurement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current de-identification methods fail to adequately assess and mitigate the risk of re-identification of personal data, particularly in datasets with longitudinal information, due to the complexity of quasi-identifiers and adversary knowledge models, leading to potential data breaches.

Innovation Solution

A computer-implemented method for re-identification risk measurement and suppression that retrieves and processes datasets with personally identifiable information, reduces multiple occurrences, groups individuals based on quasi-identifiers, orders features by identifying power, subsamples features, determines similarity measures, and combines them to assess overall risk, while incorporating date-linked knowledge and range-based counting to manage adversary power.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If de-identification removes direct identifiers (names and addresses), then ease of operation is improved, but re-identification risk increases

Engineering Contradiction:
Improveease of de-identificationVSAvoidre-identification risk
Core Design Contradiction:
Ease of operationVSObject-affected harmful factors

Solution Approach 1:

The patent replaces manual de-identification processes with an automated computer-implemented system that uses algorithms to measure re-identification risk and apply suppression. The system automatically retrieves datasets, measures risk using similarity measures, determines suppression needs, and validates results - substituting human judgment with a systematic computational approach that simultaneously addresses both ease of operation and risk reduction.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent implements a feedback loop where the system measures re-identification risk, determines suppression requirements, applies suppression, and then validates whether the risk is reduced to acceptable levels. This iterative feedback process allows the system to automatically adjust de-identification strategies to balance operational ease with risk mitigation.

Inventive Principle:
Principle #23Feedback

2Measurement precision

If comprehensive risk assessment is performed on longitudinal data, then measurement precision is improved, but device complexity increases

Engineering Contradiction:
Improverisk assessment accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the complex risk assessment process into distinct modular components: retrieving dataset, measuring re-identification risk, determining suppression requirements, applying suppression, and validating results. Each module handles a specific aspect of the assessment, making the overall complex system manageable and maintainable while achieving comprehensive and precise risk evaluation.

Inventive Principle:
Principle #1Segmentation

3Reliability

If suppression is applied to reduce re-identification risk, then reliability is improved, but loss of information increases

Engineering Contradiction:
Improveprivacy protectionVSAvoiddata quality
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent applies partial suppression - only suppressing the minimum necessary data elements to reduce re-identification risk to acceptable levels rather than applying blanket suppression to all data. The system identifies specific quasi-identifiers that contribute most to re-identification risk and targets suppression at those elements, preserving as much useful information as possible while achieving adequate privacy protection.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS9990515B2Method of re-identification risk measurement and suppression on a longitudinal dataset
Publication Date: 2018.06.05 PRIVACY ANALYTICS
  • US9990515B2 patent drawing
  • US9990515B2 patent drawing
  • US9990515B2 patent drawing

AI summary

In longitudinal datasets, it is usually unrealistic that an adversary would know the value of every quasi-identifier. De-identifying a dataset under this assumption results in high levels of generalization and suppression as every patient is unique. Adversary power gives an upper bound on the number of values an adversary knows about a patient. Considering all subsets of quasi-identifiers with the size of the adversary power is computationally infeasible. A method is provided to assess re-identification risk by determining a representative risk which can be used as a proxy for the overall risk measurement and enable suppression of identifiable quasi-identifiers.