Big Data De-identification via Flexible Grouping and Occurrence Rates
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current big data de-identification methods face challenges in preventing re-identification of individuals and ensuring the reliability of statistical analysis, as they often result in reduced accuracy due to exclusion of records with fewer than N same abstraction reference field values, leading to potential personal information leakage and incomplete data abstraction.
Innovation Solution
A de-identification method that generates abstraction records by selecting N records with the same abstraction reference field values, allocating statistical function values for numerical attributes and occurrence rates for category attributes, and grouping records with fewer than N same values to minimize data exclusion, ensuring the reliability of statistical analysis by maintaining the original meaning of the data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If records with fewer than N same abstraction reference field values are excluded from abstraction, then the de-identification reliability is improved, but the data loss increases and statistical analysis accuracy deteriorates
Solution Approach 1:
The patent changes the parameter of group size threshold from a fixed value N to a flexible range [N, M], allowing records to be grouped even when the group size is between N and M. This parameter change enables inclusion of previously excluded records while maintaining de-identification reliability through occurrence rate tracking.
Solution Approach 2:
The patent introduces occurrence rate as a feedback mechanism that tracks how many original records correspond to each abstraction record. This feedback allows the system to maintain de-identification reliability by monitoring and controlling the grouping process, ensuring that even when groups have sizes between N and M, the occurrence rate information preserves the relationship between abstracted and original records.
2Reliability
If records with fewer than N same abstraction reference field values are excluded from abstraction, then the de-identification reliability is improved, but the statistical analysis accuracy deteriorates
Solution Approach 1:
The patent changes the parameter of group size threshold from a fixed value N to a flexible range [N, M], allowing inclusion of records that would otherwise be excluded. This enables more complete data utilization for statistical analysis while maintaining de-identification reliability through the occurrence rate mechanism.
Solution Approach 2:
The occurrence rate serves as feedback that preserves statistical information about the relationship between abstracted and original records. By tracking occurrence rates, the system maintains the ability to perform accurate statistical analysis on the abstracted data while ensuring de-identification reliability is not compromised.
3Productivity
If more records are included in abstraction groups, then the data utilization is improved, but the risk of re-identification increases
Solution Approach 1:
The occurrence rate acts as a feedback control mechanism that monitors and controls the grouping process. By tracking the number of original records that map to each abstraction record, the system can include more records in abstraction groups for better data utilization while maintaining re-identification protection through occurrence rate verification.
Solution Approach 2:
The patent changes the grouping parameter from a strict threshold N to a flexible range [N, M], enabling broader data inclusion. The occurrence rate parameter provides continuous monitoring to ensure that even with larger group sizes, re-identification risk remains controlled through the ability to verify and validate abstraction integrity.
Data Source
AI summary
Provided is a de-identification method for big data, for anonymizing the big data so that the big data may be freely distributed to an external system without concern about personal information leakage and enabling a statistical value calculated from the distributed data to be maximally close to a statistical value of original data to thereby secure the reliability of statistical analysis. Records in which values of abstraction reference fields are all the same and the number thereof is less than or equal to N are separately grouped without being excluded from being abstracted, and a connection-type attribute value including an occurrence rate value of a corresponding category attribute value in a group is allocated as an attribute value of an abstracted record to minimize abstraction missing data, so that the statistical value calculated from the distributed data becomes maximally close to the statistical value of the original data.


