Privacy-Preserving Dataset Masking via K-Anonymity Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for preserving dataset privacy are inadequate as they require human intervention and are not flexible enough to meet user needs while maintaining data accuracy, and existing techniques do not effectively mask quasi-identifiers to prevent privacy invasion.

Innovation Solution

A method that determines k-anonymity and l-diversity values to cluster datasets, merges groups based on these values, and masks quasi-identifiers to enhance privacy preservation, using a decision-tree algorithm to categorize and merge data entries while minimizing data utility impact.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional anonymization methods are used to preserve privacy, then privacy protection is improved, but data accuracy deteriorates due to human intervention limitations and inability to anticipate future analysis purposes

Engineering Contradiction:
Improveprivacy protectionVSAvoiddata accuracy
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent applies preliminary action by pre-calculating and storing multiple masking schemes for quasi-identifiers before data analysis occurs. The system prepares different masking levels in advance, allowing flexible selection based on actual analysis needs without requiring human intervention at the time of analysis, thus preserving both privacy and data accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements dynamics by making the masking strategy adaptable and changeable based on different analysis purposes. The system can dynamically select which quasi-identifiers to mask and to what extent, transforming the static anonymization process into a flexible, purpose-driven approach that maintains data utility while protecting privacy.

Inventive Principle:
Principle #15Dynamics

2Reliability

If more fields are masked to increase privacy robustness, then privacy preservation is improved, but data utility deteriorates

Engineering Contradiction:
Improveprivacy robustnessVSAvoiddata utility
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent applies local quality by masking different numbers of quasi-identifier fields in different data groups based on their specific characteristics and the privacy requirements of each group. Instead of uniformly masking all fields across the entire dataset, the system selectively applies masking to specific fields in specific groups, preserving data utility where possible while maintaining privacy robustness where needed.

Inventive Principle:
Principle #3Local quality

3Reliability

If conventional methods require human intervention to determine relative and irrelative fields, then privacy protection is improved, but device complexity and operation difficulty increase

Engineering Contradiction:
Improveprivacy protectionVSAvoidprocess complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements self-service by enabling the system to automatically determine which fields are relative or irrelative to analysis purposes without human intervention. The system uses algorithms to autonomously identify quasi-identifiers, evaluate their relevance to different analysis scenarios, and apply appropriate masking strategies, eliminating the need for human experts to manually analyze and classify fields.

Inventive Principle:
Principle #25Self-service

4Ease of manufacture

If conventional methods cannot anticipate future analysis purposes, then privacy protection is improved through simpler processes, but adaptability to future needs deteriorates

Engineering Contradiction:
Improveprocess simplicityVSAvoidflexibility to analysis purposes
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent applies preliminary action by pre-computing and storing multiple masking configurations for quasi-identifiers that cover various potential analysis purposes. This allows the system to quickly adapt to future analysis needs by selecting from pre-prepared masking schemes without requiring complex re-analysis or human intervention, thus maintaining both process simplicity and high adaptability.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS8812524B2Method and system for preserving privacy of a dataset
Publication Date: 2014.08.19 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US8812524B2 patent drawing
  • US8812524B2 patent drawing
  • US8812524B2 patent drawing

AI summary

A method and a system for preserving privacy of a dataset are provided. In the method, a k-anonymity value with respect to a sensitive data field is determined according to at least one first quasi-identifier. Data entries in each group have the same value in the one or more fields of the first quasi-identifier and data entries in different groups have different values in the one or more fields of the first quasi-identifier. A first group and a second group among the plurality of groups are determined according to the reference number Kr, where the first group and the second group are merged into a merging group. The number of data entries in the merging group is not less than a reference number Kr. One or more fields of at least one first quasi-identifier is masked for the merging group.