Asymmetric Re-identification Risk Model for Anonymized Cohorts
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current risk assessment methods for re-identification in datasets are inadequate as they fail to differentiate between various types of re-identification scenarios, leading to inaccurate risk estimation and potential over-perturbation or under-perturbation of data, which can compromise patient privacy and data quality in medical studies.
Innovation Solution
A system and method that assesses re-identification risk by distinguishing between public and acquaintance quasi-identifiers, calculating separate risk measures for each type of attack, and perturbing data to maintain a predetermined threshold of risk, thereby producing an anonymized cohort with reduced re-identification risk.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional risk measures (k-anonymity, expected number of correct re-identification) are used to assess re-identification risk, then a single uniform risk threshold is applied to all records, but this fails to account for different types of re-identification scenarios (population-to-sample vs. sample-to-population attacks) and leads to inaccurate risk estimation
Solution Approach 1:
The patent segments the re-identification risk assessment into two distinct models: one for population-to-sample attacks and another for sample-to-population attacks. Each model uses appropriate risk measures tailored to its specific attack scenario, allowing for more accurate risk estimation without requiring a single overly complex unified model.
Solution Approach 2:
The patent dynamically selects which risk assessment model to apply based on the specific re-identification scenario being evaluated. The system adapts its assessment approach depending on whether the attack vector is population-to-sample or sample-to-population, enabling precise risk measurement while maintaining model simplicity through context-appropriate selection.
2Reliability
If data perturbation is applied to reduce re-identification risk, then patient privacy is protected, but data quality degrades and medical study reliability is compromised
Solution Approach 1:
The patent applies perturbation selectively based on the calculated re-identification risk for each specific record. Records with high re-identification risk undergo perturbation to protect privacy, while records with low risk remain unchanged to preserve data quality. This localized approach ensures privacy protection is applied only where necessary.
Solution Approach 2:
The patent changes the perturbation parameters (such as perturbation magnitude and method) based on the specific risk assessment results. By adjusting perturbation intensity according to the calculated risk level, the system protects privacy for high-risk records while minimizing data quality degradation for lower-risk records.
3Ease of operation
If uniform risk threshold is applied to all anonymized records, then the assessment process is simple, but it cannot differentiate between records with different re-identification risks in different attack scenarios
Solution Approach 1:
The patent segments the risk assessment into distinct models for different attack scenarios, maintaining simplicity within each segment while achieving precise differentiation across scenarios. Each segmented model uses appropriate risk measures for its specific context, avoiding the need for a single complex threshold system.
Solution Approach 2:
The patent dynamically adjusts the risk threshold and assessment criteria based on the specific attack scenario being evaluated. This allows the system to maintain operational simplicity by using scenario-appropriate thresholds while achieving precise risk differentiation through adaptive threshold selection.
Data Source
AI summary
System and method to produce an anonymized cohort, members of the cohort having less than a predetermined risk of re-identification. The system includes a user-facing communication interface to receive an anonymized cohort request comprising traits to include in members of the cohort; a data source-facing communication channel to query a data source, to find anonymized records that possess at least some of the requested traits; and a processor programmed to carry out the instructions of: forming a dataset from at least some of the anonymized records; calculating a risk of re-identification of the anonymized records in the dataset based upon the data query; perturbing anonymized records in the dataset that exceed a predetermined risk of re-identification, until the risk of re-identification is not greater than the pre-determined threshold, to produce the anonymized cohort; and providing, via a user-facing communication channel, the anonymized cohort.


