Simulated Risk Contribution for Data De-identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for de-identifying datasets struggle to effectively assess and mitigate the risk of re-identification, particularly with quasi-identifiers that are not directly obvious.
Innovation Solution
A system and method that estimate disclosure risk by retrieving a population distribution, assigning information scores to quasi-identifying fields, aggregating these scores, and calculating anonymity and confidentiality values to determine the re-identification risk of individual data subjects.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If de-identification removes only direct identifiers (names and addresses), then data utility is maintained, but re-identification risk increases due to quasi-identifiers
Solution Approach 1:
The patent segments the de-identification process into multiple stages: identifying direct identifiers, identifying quasi-identifiers, and applying different protection levels to each type. This allows selective de-identification that maintains data utility while addressing re-identification risks through targeted protection of quasi-identifiers.
Solution Approach 2:
The patent applies local quality by differentiating between direct identifiers and quasi-identifiers, applying stronger protection to quasi-identifiers while maintaining direct identifiers for data utility. The system assigns different de-identification strategies to different data elements based on their identification strength.
2Reliability
If comprehensive de-identification is applied to all fields, then re-identification risk is reduced, but data utility and accessibility decrease
Solution Approach 1:
The patent implements local quality by applying differential de-identification strategies to different field types. Direct identifiers receive basic protection while quasi-identifiers receive enhanced protection, allowing the system to maintain data utility where appropriate while protecting against re-identification where necessary.
Solution Approach 2:
The system segments the data fields into distinct categories (direct identifiers, quasi-identifiers, and other fields) and applies appropriate de-identification measures to each segment, avoiding unnecessary protection that would reduce data utility.
3Measurement precision
If traditional risk assessment methods are used, then disclosure risk can be evaluated, but computational resources and processing time increase significantly
Solution Approach 1:
The patent applies preliminary action by pre-calculating and storing risk metrics for different quasi-identifier combinations and population distributions before actual de-identification occurs. This allows rapid risk assessment during data processing without requiring computationally intensive real-time analysis of all possible re-identification scenarios.
Solution Approach 2:
The system uses partial action by focusing computational resources on the most critical risk factors - specifically on quasi-identifiers and their combinations - rather than exhaustively analyzing all possible data fields and population scenarios, thereby reducing overall computational burden.
4Measurement precision
If detailed analysis of all quasi-identifiers is performed, then disclosure risk is accurately measured, but processing time increases
Solution Approach 1:
The patent implements preliminary action by pre-computing risk metrics for common quasi-identifier combinations and storing them in lookup tables. During actual processing, the system queries these pre-computed values rather than performing exhaustive calculations, significantly reducing processing time while maintaining measurement precision for the most common scenarios.
Solution Approach 2:
The system applies partial action by focusing computational effort on the most impactful quasi-identifiers and their combinations, rather than uniformly analyzing all fields. This prioritized approach maintains accurate risk measurement for critical factors while reducing overall processing time.
Data Source
AI summary
Computing devices utilizing computer-readable media implementing methods arranged for deriving risk contribution models from a dataset are presented herein. Rather than inspect the entire data model to identify all identifying fields, the computing device develops a list of a subset of identifying fields. For each such field, the computing device creates a distribution of values/information values from other sources. Then, when risk measurement is performed, simulated values (or information values) are generated for these fields. These are incorporated into the overall risk measurement and utilized in an anonymization process.


