Incremental Data Generalization for Continuous Anonymization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing anonymization methods require complete datasets before anonymization can be applied, leading to delays in data utilization for secondary purposes due to the need for k-anonymity, t-closeness, or l-diversity, which cannot be applied until all data is collected.
Innovation Solution
A method and device for anonymizing data using generalization techniques that allow anonymization to be performed on partial datasets, with subsequent refinements as more data is collected, ensuring k-anonymity and other requirements are met at each stage, using a processing device to generate and refine assignment ranges for quasi-identifiers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If anonymization is performed using traditional methods (k-anonymity, t-closeness, l-diversity), then data protection requirements are met, but data processing must wait until complete datasets are collected, causing delays in utilization
Solution Approach 1:
The patent applies preliminary action by performing anonymization on partial datasets before complete data is available. The system generates an initial generalization for the first datasets at the first time point, enabling early processing and utilization while maintaining anonymity. This preliminary anonymization allows data to be processed for secondary purposes without waiting for the complete dataset, thus improving productivity while preserving reliability through subsequent refinement when complete data becomes available.
2Productivity
If anonymization is performed on partial datasets early, then data processing speed improves, but anonymization quality may be insufficient without complete data context
Solution Approach 1:
The patent applies dynamics by making the generalization adaptive and evolving over time. The system generates an initial generalization for partial datasets, then refines it when complete datasets are available. The refinement process adjusts the generalization based on the additional context from complete data, ensuring that anonymization quality improves while maintaining the ability to process data early. This dynamic approach allows the system to balance processing speed with anonymization quality throughout the data collection timeline.
3Measurement precision
If data is collected completely before anonymization, then anonymization accuracy is maximized, but time is lost waiting for complete data collection
Solution Approach 1:
The patent applies preliminary action by performing anonymization on partial datasets before complete data is available. The system generates an initial generalization for the first datasets at the first time point, enabling early processing and utilization while maintaining anonymity. This preliminary anonymization allows data to be processed for secondary purposes without waiting for the complete dataset, thus improving productivity while preserving reliability through subsequent refinement when complete data becomes available.
Solution Approach 2:
The patent applies feedback by using the complete datasets available at the second time point to refine the initial generalization. The system compares the initial anonymization results with the complete data context and adjusts the generalization accordingly, improving anonymization accuracy while maintaining the time benefits of early processing. This feedback mechanism ensures that the final anonymization is as accurate as if performed on complete data, while still allowing earlier utilization.
Data Source
AI summary
A computer-implemented method for anonymizing data via generalization, wherein the data includes a first number of first datasets at a first time point and a second number of second datasets at a second time point. The first datasets are a subset of the second datasets. The computer-implemented method comprises: generating a first generalization for the first datasets that fulfills a required anonymization, wherein the first generalization includes a first group of assignment ranges by which values of a quasi-identifier of the data are generalized; and generating a second generalization for the second datasets that fulfills the required anonymization, wherein the second generalization includes a second group of assignment ranges by which values of the quasi-identifier are generalized. The second group includes more assignment ranges than the first group.


