Anonymizing Set-Valued Attributes Without Taxonomy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current anonymization technologies face challenges in processing set-valued attributes without taxonomy in datasets containing both single-valued and set-valued attributes, leading to increased information loss and inefficiencies, especially when multiple attributes coexist, as they often require extensive processing patterns and fail to scale effectively.

Innovation Solution

An information processing device and method that acquire cluster information with anonymized set-valued attributes, disclose and refine attribute values to divide clusters that satisfy predetermined anonymity, allowing for scalable anonymization without requiring taxonomy for set-valued attributes, using a top-down approach to minimize processing and maintain data utility.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If k-anonymization is applied to single-valued attributes, then privacy protection is improved, but it cannot handle set-valued attributes without taxonomy effectively

Engineering Contradiction:
Improveprivacy protectionVSAvoidhandling capability for set-valued attributes
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent extends the k-anonymization framework to handle both single-valued and set-valued attributes uniformly. The anonymization process is designed to work with diverse attribute types without requiring separate处理方法, achieving multi-functionality that resolves the limitation of existing single-valued attribute approaches

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Device complexity

If existing anonymization technologies are applied to datasets with mixed attribute types, then processing complexity increases, but information loss increases

Engineering Contradiction:
Improveprocessing complexityVSAvoidinformation loss
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The patent segments the anonymization process into distinct phases: initial anonymization of single-valued attributes, followed by targeted handling of set-valued attributes. This segmentation allows each attribute type to be processed with appropriate methods, reducing unnecessary complexity and minimizing information loss that would occur with uniform processing approaches

Inventive Principle:
Principle #1Segmentation

3Reliability

If comprehensive anonymization processing is applied to all attributes, then privacy protection is improved, but processing time increases

Engineering Contradiction:
Improveprivacy protectionVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies partial anonymization by focusing computational resources on set-valued attributes that require special handling, while using more efficient methods for single-valued attributes. This partial action approach achieves adequate privacy protection without the excessive processing time that would result from applying comprehensive anonymization to all attribute types uniformly

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS9940473B2Information processing device, information processing method and medium
Publication Date: 2018.04.10 NEC CORP
  • US9940473B2 patent drawing
  • US9940473B2 patent drawing
  • US9940473B2 patent drawing

AI summary

An information processing device of the present invention includes: a cluster information acquisition unit which acquires information indicating a cluster which is a set of records in an anonymized state in which at least a portion of attribute values of set-valued attributes, which can include one value or a plurality of values included in the records, is removed from the cluster which is a set of records including an attribute value so that the cluster satisfies a predetermined anonymity; and a set-valued attribute refinement unit which discloses at least a portion of attribute values from among removed attribute values of the set-valued attributes of records included in the cluster acquired by the cluster acquisition, and divides the cluster into clusters which satisfy the predetermined anonymity based on the disclosed attribute values.