De-identified Data Access Control via Re-identification Risk Scoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Re-identification of de-identified data poses significant security and privacy concerns, making it difficult to determine the risk of data being used for re-identification, especially in large datasets.
Innovation Solution
A data set management system calculates a re-identification risk score for quasi-identifiers in de-identified data sets and selectively outputs actual or synthetic data based on this score to manage access and reduce re-identification risk.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If de-identified data is made available for analysis, then data utility and research value are improved, but re-identification risk increases
Solution Approach 1:
The patent introduces a risk assessment system as an intermediary between the de-identified data and users. This system calculates re-identification risk scores based on various factors (dataset size, quasi-identifier combinations, external data availability) and uses these scores to control data access, thereby mediating between data utility and privacy protection
Solution Approach 2:
The patent changes the parameter of data representation by generating synthetic data that mimics the statistical properties of real de-identified data but without containing actual identifiable records. This parameter change allows data to maintain utility while eliminating re-identification risk
2Reliability
If strict privacy controls are applied to prevent re-identification, then security is improved, but data accessibility and usability deteriorate
Solution Approach 1:
The patent implements dynamic access control where data accessibility is adjusted based on the calculated re-identification risk score. For low-risk datasets, full access is permitted; for high-risk datasets, synthetic data is provided instead. This dynamic approach maintains security while preserving accessibility where safe
Solution Approach 2:
The patent creates synthetic copies of de-identified data that replicate the statistical characteristics and analytical value of real data without containing actual personal information. These copies can be freely distributed and accessed, improving data accessibility while maintaining security
3Object-affected harmful factors
If comprehensive risk assessment is performed on all data, then re-identification risk is reduced, but computational complexity and processing time increase
Solution Approach 1:
The patent segments the risk assessment process into distinct components: identifying quasi-identifiers, calculating risk scores for individual attributes, combining risks for attribute sets, and determining overall dataset risk. This segmentation allows for systematic and efficient processing of large datasets
Data Source
Figure 1A
Figure 1B
Figure 1C
AI summary
A system may receive, from one or more data sources, one or more de-identified data sets that include de-identified personal data. The system may receive a request for a feature set of the one or more de-identified data sets, wherein the feature set includes a set of quasi-identifiers included in the de-identified personal data. The system may calculate a re-identification risk score for the set of quasi-identifiers. The system may selectively output, based on the re-identification risk score, one of: actual data, from the one or more de-identified data sets, of the feature set if the re-identification risk score satisfies a condition, or synthetic data, generated by the device from the one or more de-identified data sets, for the feature set, or a combination of the synthetic data and the actual data for the feature set, if the re-identification risk score does not satisfy the condition.