Longitudinal Dataset Re-identification Risk Measurement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current de-identification methods fail to adequately assess and mitigate the risk of re-identification of personal data, particularly in datasets with longitudinal information, due to the complexity of quasi-identifiers and adversary knowledge models, leading to potential data breaches.
Innovation Solution
A computer-implemented method for re-identification risk measurement and suppression that retrieves and processes datasets with personally identifiable information, reduces multiple occurrences, groups individuals based on quasi-identifiers, orders features by identifying power, subsamples features, determines similarity measures, and combines them to assess overall risk, while incorporating date-linked knowledge and range-based counting to manage adversary power.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If de-identification removes direct identifiers (names and addresses), then ease of operation is improved, but re-identification risk increases
Solution Approach 1:
The patent replaces manual de-identification processes with an automated computer-implemented system that uses algorithms to measure re-identification risk and apply suppression. The system automatically retrieves datasets, measures risk using similarity measures, determines suppression needs, and validates results - substituting human judgment with a systematic computational approach that simultaneously addresses both ease of operation and risk reduction.
Solution Approach 2:
The patent implements a feedback loop where the system measures re-identification risk, determines suppression requirements, applies suppression, and then validates whether the risk is reduced to acceptable levels. This iterative feedback process allows the system to automatically adjust de-identification strategies to balance operational ease with risk mitigation.
2Measurement precision
If comprehensive risk assessment is performed on longitudinal data, then measurement precision is improved, but device complexity increases
Solution Approach 1:
The patent segments the complex risk assessment process into distinct modular components: retrieving dataset, measuring re-identification risk, determining suppression requirements, applying suppression, and validating results. Each module handles a specific aspect of the assessment, making the overall complex system manageable and maintainable while achieving comprehensive and precise risk evaluation.
3Reliability
If suppression is applied to reduce re-identification risk, then reliability is improved, but loss of information increases
Solution Approach 1:
The patent applies partial suppression - only suppressing the minimum necessary data elements to reduce re-identification risk to acceptable levels rather than applying blanket suppression to all data. The system identifies specific quasi-identifiers that contribute most to re-identification risk and targets suppression at those elements, preserving as much useful information as possible while achieving adequate privacy protection.
Data Source
AI summary
In longitudinal datasets, it is usually unrealistic that an adversary would know the value of every quasi-identifier. De-identifying a dataset under this assumption results in high levels of generalization and suppression as every patient is unique. Adversary power gives an upper bound on the number of values an adversary knows about a patient. Considering all subsets of quasi-identifiers with the size of the adversary power is computationally infeasible. A method is provided to assess re-identification risk by determining a representative risk which can be used as a proxy for the overall risk measurement and enable suppression of identifiable quasi-identifiers.


