Anonymized Data Re-Identification Risk Scoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing anonymization methods lack a universal algorithm that effectively balances privacy protection and data utility, and evaluating re-identification risk is challenging due to varying applicability contexts and the evolution of information technologies.
Innovation Solution
A method for evaluating re-identification risk using a deterministic search based on external information sources and a distance-based search, involving data transformation into a Euclidean space and using k-NN to quantify re-identification failure and protection degree.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If anonymization algorithms modify data by deleting, generalizing or replacing personal information, then privacy protection is improved, but information content and data utility are lost
Solution Approach 1:
The patent changes the parameter of anonymization from binary (anonymized/not anonymized) to a continuous spectrum by introducing a risk score that quantifies re-identification probability. This allows selective application of anonymization intensity based on individual record risk levels, preserving information content where risk is low while maintaining privacy where risk is high.
Solution Approach 2:
The patent applies partial anonymization by computing risk scores for individual records and applying anonymization only to records exceeding a threshold. This partial action approach avoids excessive anonymization of low-risk records, thereby preserving information content while still protecting high-risk records.
2Ease of manufacture
If a unique anonymization algorithm is used across all contexts, then implementation simplicity is improved, but adaptability to different data types and use cases deteriorates
Solution Approach 1:
The patent creates a universal risk assessment framework that can evaluate re-identification risk across different data types, anonymization methods, and attack scenarios. The framework serves multiple functions: risk evaluation, parameter optimization, and effectiveness verification, making it adaptable to various contexts without requiring context-specific algorithms.
Solution Approach 2:
The patent introduces dynamic adaptability by allowing the risk assessment model to learn from data characteristics and adjust risk scores accordingly. The framework can adapt to different data types, anonymization techniques, and threat models, providing context-appropriate risk evaluation while maintaining a single unified approach.
3Productivity
If existing anonymization methods are applied without risk evaluation, then processing speed is improved, but the effectiveness of privacy protection deteriorates
Solution Approach 1:
The patent performs preliminary risk assessment before finalizing anonymization parameters. By computing risk scores in advance and using them to guide anonymization intensity, the system ensures that privacy protection effectiveness is optimized without requiring repeated trial-and-error processing, thus maintaining processing speed.
Solution Approach 2:
The patent implements a feedback mechanism where risk scores are computed based on data characteristics and anonymization parameters, and these scores feed back into the anonymization process to adjust parameters. This closed-loop approach ensures that privacy protection effectiveness is continuously optimized while maintaining efficient processing through automated parameter adjustment.
Data Source
AI summary
The method delivers a degree of protection (txP3) representative of the risk of re-identification of data in the case of a correspondence search attack including a deterministic search based on an external information source and a correspondence search based on a distance. The method comprises steps of E) consolidating a set of original individuals (EDO) and a set of anonymous individuals (IA); F) identifying, in the set of original individuals, individuals at risk (IOrs) via the deterministic correspondence search; G) evaluating a degree of failure of re-identification (txP1) for the sets of original individuals and of anonymous individuals, on the basis of the correspondence search based on distance; H) computing the degree of protection as a function of a total number of individuals in the original dataset, of a number (RS) of individuals at risk identified in step B) and of the degree of failure of re-identification (txP1).


