Data Anonymization Hierarchy Control for Minimal Record Loss
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data anonymization methods face challenges in balancing the need to minimize information loss while ensuring k-anonymity, leading to degraded data analysis accuracy due to excessive record deletion.
Innovation Solution
An information processing device that automatically determines anonymization granularity by classifying records into sets, calculating the number of records and their ratios, and adjusting hierarchy levels based on predefined priorities to achieve k-anonymity with minimal data loss.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If data is anonymized with fine granularity to minimize information loss per record, then information quality is improved, but the number of records to be deleted increases
Solution Approach 1:
The patent applies dynamics by making the anonymization granularity adjustable and adaptable. The system dynamically selects the appropriate hierarchy level for each masking target item based on calculated ratios and priority settings, rather than using a fixed granular level for all data. This allows the anonymization process to optimize between information loss and record deletion by flexibly adjusting the granularity level.
Solution Approach 2:
The patent changes the parameter of anonymization granularity by introducing hierarchy levels (first through fourth levels) for masking target items. By adjusting which hierarchy level is selected for each item, the system can control the degree of anonymization. The calculation unit computes the ratio of records that would be deleted at each granularity level, and the priority setting unit allows operators to adjust parameters to find the optimal balance between maintaining information quality and preserving record quantity.
2Quantity of substance
If data is anonymized with rough granularity to minimize record deletion, then the number of records retained is improved, but information loss per record increases
Solution Approach 1:
The system dynamically adjusts anonymization granularity based on operator priorities and calculated ratios. When operators prioritize retaining records over information quality, the system automatically selects coarser hierarchy levels. Conversely, when information quality is prioritized, finer granularity is applied. This dynamic adaptation resolves the contradiction by allowing flexible trade-off adjustment.
Solution Approach 2:
The patent applies partial action by allowing selective application of different anonymization granularities to different masking target items. Rather than uniformly applying rough or fine granularity across all data, the system calculates the optimal hierarchy level for each item individually based on its characteristics and the operator's priority settings, achieving partial optimization across the dataset.
3Ease of operation
If manual determination of anonymization granularity is used to control record deletion, then operator control is improved, but processing time and complexity increase
Solution Approach 1:
The patent applies preliminary action by having the calculation unit pre-calculate the ratio of records that would be deleted at each hierarchy level before the operator makes the final decision. This preliminary computation provides the operator with informed choices and automates the tedious calculation work, reducing both processing time and complexity while maintaining operator control over the final granularity selection.
Solution Approach 2:
The system performs self-service by automatically calculating the impact of different anonymization granularities and presenting options to the operator. The calculation unit independently computes deletion ratios for each hierarchy level, and the priority setting unit automatically adjusts settings based on operator preferences, minimizing the operator's manual workload while preserving strategic control.
Data Source
AI summary
With respect to an information processing device which anonymizes data composed of records including one or more items through statistical processing, the information processing device includes a memory, and a processor configured to classify respective records constituting the data into one or more first sets, based on masking target items, a dictionary, and a selected hierarchy level, classify the respective records into one or more second sets with respect to a number of records belonging to each of the one or more first sets, and calculate a number of records of each of the one or more second sets and a ratio of records belonging to each of the one or more second sets to the records constituting the data, change the selected hierarchy level based on the ratio and priority set in advance, and create a statistically processed record by statistically processing records belonging to a same first set.


