Data Risk Assessment Using Sample and Population Re-Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for determining data risk rely on a single attack type assumption, leading to inaccurate data risk estimation due to the relationship between the sample dataset and the prevalent population, and do not consider multiple attack types.
Innovation Solution
A method that determines data risk by consolidating re-identification probabilities of each record in a sample dataset and statistical information of a population, using weighted sums and various approaches like entropy, Bayesian, and hypothesis testing to balance and improve accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If data risk is determined based on a single attack type assumption (either sample dataset or population), then the determination process is simple, but the accuracy of data risk estimation deteriorates
Solution Approach 1:
The patent combines multiple attack type assumptions (sample dataset-based and population-based) into a unified data risk determination framework. By merging the results from different attack scenarios through weighted aggregation, the system achieves more accurate and comprehensive data risk estimation while maintaining manageable process complexity.
2Ease of operation
If data risk determination considers only sample dataset information, then the process is straightforward, but the reliability deteriorates due to potential mismatch with prevalent population
Solution Approach 1:
The patent introduces population-level statistical information as an intermediary component that bridges the gap between sample dataset analysis and real-world data risk assessment. This intermediary population information compensates for potential mismatches between the sample dataset and prevalent population, thereby improving the reliability of data risk estimation without significantly complicating the determination process.
3Measurement precision
If multiple attack types are considered in data risk determination, then the accuracy improves, but the device complexity increases
Solution Approach 1:
The patent manages the complexity of considering multiple attack types by parameterizing the determination process. It uses configurable weights for different attack scenarios and employs standardized calculation methodologies for each attack type, allowing the system to accurately assess multiple attack vectors while maintaining controlled complexity through parameter-driven flexibility.
Data Source
Figure 1A~1B
Figure 2
Figure 3
AI summary
Example embodiments of the present disclosure relate to a method for determining a data risk, an electronic device, and a computer-readable medium. In the solution, a first data risk value is determined based on a re-identification probability of each record in the sample dataset, and a second data risk value is determined based on statistical information of a population which includes the sample dataset. In addition, a data risk may be determined based on the first data risk value and the second data risk value. Embodiments of the present disclosure consider both a re-identification probability of each record in the sample dataset and statistical information of a population for determining the data risk. Therefore, the determined data risk will be more accurate, thereby an accurate re-identification risk may be obtained accordingly.