Big Data De-identification via Flexible Grouping and Occurrence Rates

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current big data de-identification methods face challenges in preventing re-identification of individuals and ensuring the reliability of statistical analysis, as they often result in reduced accuracy due to exclusion of records with fewer than N same abstraction reference field values, leading to potential personal information leakage and incomplete data abstraction.

Innovation Solution

A de-identification method that generates abstraction records by selecting N records with the same abstraction reference field values, allocating statistical function values for numerical attributes and occurrence rates for category attributes, and grouping records with fewer than N same values to minimize data exclusion, ensuring the reliability of statistical analysis by maintaining the original meaning of the data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If records with fewer than N same abstraction reference field values are excluded from abstraction, then the de-identification reliability is improved, but the data loss increases and statistical analysis accuracy deteriorates

Engineering Contradiction:
Improvede-identification reliabilityVSAvoiddata loss
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent changes the parameter of group size threshold from a fixed value N to a flexible range [N, M], allowing records to be grouped even when the group size is between N and M. This parameter change enables inclusion of previously excluded records while maintaining de-identification reliability through occurrence rate tracking.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces occurrence rate as a feedback mechanism that tracks how many original records correspond to each abstraction record. This feedback allows the system to maintain de-identification reliability by monitoring and controlling the grouping process, ensuring that even when groups have sizes between N and M, the occurrence rate information preserves the relationship between abstracted and original records.

Inventive Principle:
Principle #23Feedback

2Reliability

If records with fewer than N same abstraction reference field values are excluded from abstraction, then the de-identification reliability is improved, but the statistical analysis accuracy deteriorates

Engineering Contradiction:
Improvede-identification reliabilityVSAvoidstatistical analysis accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent changes the parameter of group size threshold from a fixed value N to a flexible range [N, M], allowing inclusion of records that would otherwise be excluded. This enables more complete data utilization for statistical analysis while maintaining de-identification reliability through the occurrence rate mechanism.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The occurrence rate serves as feedback that preserves statistical information about the relationship between abstracted and original records. By tracking occurrence rates, the system maintains the ability to perform accurate statistical analysis on the abstracted data while ensuring de-identification reliability is not compromised.

Inventive Principle:
Principle #23Feedback

3Productivity

If more records are included in abstraction groups, then the data utilization is improved, but the risk of re-identification increases

Engineering Contradiction:
Improvedata utilizationVSAvoidre-identification risk
Core Design Contradiction:
ProductivityVSObject-affected harmful factors

Solution Approach 1:

The occurrence rate acts as a feedback control mechanism that monitors and controls the grouping process. By tracking the number of original records that map to each abstraction record, the system can include more records in abstraction groups for better data utilization while maintaining re-identification protection through occurrence rate verification.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent changes the grouping parameter from a strict threshold N to a flexible range [N, M], enabling broader data inclusion. The occurrence rate parameter provides continuous monitoring to ensure that even with larger group sizes, re-identification risk remains controlled through the ability to verify and validate abstraction integrity.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11941153B2De-identification method for big data
Publication Date: 2024.03.26 BOALA CO LTD
  • US11941153B2 patent drawing
  • US11941153B2 patent drawing
  • US11941153B2 patent drawing

AI summary

Provided is a de-identification method for big data, for anonymizing the big data so that the big data may be freely distributed to an external system without concern about personal information leakage and enabling a statistical value calculated from the distributed data to be maximally close to a statistical value of original data to thereby secure the reliability of statistical analysis. Records in which values of abstraction reference fields are all the same and the number thereof is less than or equal to N are separately grouped without being excluded from being abstracted, and a connection-type attribute value including an occurrence rate value of a corresponding category attribute value in a group is allocated as an attribute value of an abstracted record to minimize abstraction missing data, so that the statistical value calculated from the distributed data becomes maximally close to the statistical value of the original data.