Privacy-Preserving Dataset Masking via K-Anonymity Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for preserving dataset privacy are inadequate as they require human intervention and are not flexible enough to meet user needs while maintaining data accuracy, and existing techniques do not effectively mask quasi-identifiers to prevent privacy invasion.
Innovation Solution
A method that determines k-anonymity and l-diversity values to cluster datasets, merges groups based on these values, and masks quasi-identifiers to enhance privacy preservation, using a decision-tree algorithm to categorize and merge data entries while minimizing data utility impact.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional anonymization methods are used to preserve privacy, then privacy protection is improved, but data accuracy deteriorates due to human intervention limitations and inability to anticipate future analysis purposes
Solution Approach 1:
The patent applies preliminary action by pre-calculating and storing multiple masking schemes for quasi-identifiers before data analysis occurs. The system prepares different masking levels in advance, allowing flexible selection based on actual analysis needs without requiring human intervention at the time of analysis, thus preserving both privacy and data accuracy.
Solution Approach 2:
The patent implements dynamics by making the masking strategy adaptable and changeable based on different analysis purposes. The system can dynamically select which quasi-identifiers to mask and to what extent, transforming the static anonymization process into a flexible, purpose-driven approach that maintains data utility while protecting privacy.
2Reliability
If more fields are masked to increase privacy robustness, then privacy preservation is improved, but data utility deteriorates
Solution Approach 1:
The patent applies local quality by masking different numbers of quasi-identifier fields in different data groups based on their specific characteristics and the privacy requirements of each group. Instead of uniformly masking all fields across the entire dataset, the system selectively applies masking to specific fields in specific groups, preserving data utility where possible while maintaining privacy robustness where needed.
3Reliability
If conventional methods require human intervention to determine relative and irrelative fields, then privacy protection is improved, but device complexity and operation difficulty increase
Solution Approach 1:
The patent implements self-service by enabling the system to automatically determine which fields are relative or irrelative to analysis purposes without human intervention. The system uses algorithms to autonomously identify quasi-identifiers, evaluate their relevance to different analysis scenarios, and apply appropriate masking strategies, eliminating the need for human experts to manually analyze and classify fields.
4Ease of manufacture
If conventional methods cannot anticipate future analysis purposes, then privacy protection is improved through simpler processes, but adaptability to future needs deteriorates
Solution Approach 1:
The patent applies preliminary action by pre-computing and storing multiple masking configurations for quasi-identifiers that cover various potential analysis purposes. This allows the system to quickly adapt to future analysis needs by selecting from pre-prepared masking schemes without requiring complex re-analysis or human intervention, thus maintaining both process simplicity and high adaptability.
Data Source
AI summary
A method and a system for preserving privacy of a dataset are provided. In the method, a k-anonymity value with respect to a sensitive data field is determined according to at least one first quasi-identifier. Data entries in each group have the same value in the one or more fields of the first quasi-identifier and data entries in different groups have different values in the one or more fields of the first quasi-identifier. A first group and a second group among the plurality of groups are determined according to the reference number Kr, where the first group and the second group are merged into a merging group. The number of data entries in the merging group is not less than a reference number Kr. One or more fields of at least one first quasi-identifier is masked for the merging group.


