Identity De-identification Device Using Frequency-Based Hierarchy Trees

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing personal information anonymization technologies face challenges in automating the selection of dominant anonymous data, making it difficult to determine the availability and efficiency of anonymization processes, which increases operational costs and compromises data protection.

Innovation Solution

A personal information anonymization device that automatically generates a generalization hierarchy tree for each attribute, using frequency data to represent dominant concepts and recode attribute values, ensuring that a threshold number of attribute value combinations are met, thereby enhancing data protection and reducing operational costs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If all candidates reaching the threshold value are output in anonymization processing, then data protection is improved, but device complexity and operational cost increase due to difficulty in automating dominance determination

Engineering Contradiction:
Improvedata protectionVSAvoidanonymization process complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system automatically determines the dominance of anonymous data candidates by computing information quantities and comparing them against thresholds, eliminating the need for manual evaluation. The frequency obtaining unit and information quantity computing unit enable the system to self-evaluate and select optimal anonymization candidates automatically.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The invention introduces information quantity as a quantitative parameter to evaluate anonymous data candidates. By computing and comparing information quantities of different candidates against a threshold, the system objectively determines dominance, replacing subjective manual assessment with automated parameter-based evaluation.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If manual selection of dominant anonymous data is performed, then data protection quality is improved, but productivity decreases due to increased operational time and cost

Engineering Contradiction:
Improveanonymization qualityVSAvoidanonymization throughput
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system performs automatic evaluation and selection of dominant anonymous data candidates through automated computation of information quantities and threshold comparisons. This eliminates manual intervention while maintaining evaluation quality, thereby increasing processing throughput and productivity.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The invention replaces manual mechanical evaluation processes with automated computational systems. The frequency obtaining unit, information quantity computing unit, and threshold comparison mechanism substitute human operators, enabling high-speed automated processing while maintaining evaluation rigor.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If generalization hierarchy trees are separately defined for every attribute, then measurement precision is improved, but device complexity increases making automation difficult

Engineering Contradiction:
Improveattribute obfuscation precisionVSAvoidhierarchy tree management complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system uses a unified threshold value that applies across all attributes for determining dominance of anonymous data candidates. This universal threshold mechanism simplifies the management of generalization hierarchy trees while maintaining precise control over anonymization quality across different attribute types.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The invention introduces information quantity as a universal parameter for evaluating anonymous data candidates across all attributes. By using this standardized parameter with a universal threshold, the system maintains measurement precision while enabling automated evaluation without requiring complex attribute-specific management.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP2573699B1Identity information de-identification device
Publication Date: 2017.06.07 HITACHI LTD
  • EP2573699B1 patent drawingFigure 1
  • EP2573699B1 patent drawingFigure 2~3
  • EP2573699B1 patent drawingFigure 4

AI summary

A de-identification device is provided for automatically configuring a general hierarchy tree of attribute values in protection technology of identity information. In addition, the provided de-identification device quantitively evaluates the amount of information which is lost when generalizing an attribute value, and can thereby automatically assess priorities between de-identified data and between data that is being de-identified. Information of each person includes attribute values of the person for a plurality of attributes. De-identification is achieved by obfuscating the attribute values, and a structure in which attribute values to be obfuscated are expressed in a tree structure according to the level of obfuscation is called a general hierarchy tree. The disclosed identity information de-identification device achieves automatic configuration by configuring a tree using frequency information of attribute values. In addition, by defining a lost information amount metric means, using the general hierarchy tree, in formation amount loss between two de-identified data or between data being de-identified is quantitively assessed.