Big Data Anonymization via Surrogate Synchronization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional anonymization methods for big data fail to prevent re-identification of personal information, making it difficult to freely distribute and combine anonymized data across systems without risking personal information leakage.

Innovation Solution

A method involving a data server with a communication, processing, and storage unit that anonymizes personal information by converting synchronization target personal identification attribute values into surrogate values, applying error values to group them into cells, and assigning synchronization attributes based on section information, making re-identification extremely hard and enabling secure distribution and combination of data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional anonymization methods (masking, substitution, semi-identification, categorization) are applied to big data, then personal information is anonymized to some extent, but re-identification is still possible and data combination across systems remains difficult

Engineering Contradiction:
Improveanonymization effectivenessVSAvoiddata combinability
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent segments the anonymization process into multiple distinct steps: generating surrogate values from original personal identifiers, creating synchronization dictionaries that map original values to surrogates, applying cell sectionalization to group surrogate values, and adding error values to create final anonymized attributes. This multi-stage segmentation ensures thorough anonymization while preserving combinability through the synchronization dictionary.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The synchronization dictionary acts as an intermediary between original personal identification data and anonymized data. It contains the mapping relationships that enable data combination across systems without exposing actual personal identifiers. The dictionary serves as a mediator that allows re-identification only by authorized parties with access to the original data, while preventing unauthorized re-identification of anonymized data.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Object-affected harmful factors

If personal identification attributes are removed from data sets to prevent re-identification, then privacy protection is improved, but the ability to combine records of the same person across different data sets is lost

Engineering Contradiction:
Improvepersonal information leakage riskVSAvoiddata combination capability
Core Design Contradiction:
Object-affected harmful factorsVSEase of operation

Solution Approach 1:

The patent creates a copy of the personal identification attribute in the form of a synchronization attribute that derives from the original identifier but is transformed through surrogate values and cell sectionalization. This copied attribute maintains the ability to link records of the same person across data sets while being sufficiently transformed to prevent direct re-identification. The synchronization dictionary preserves the mapping relationship for authorized re-identification when needed.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent transforms the personal identification attribute through multiple parameter changes: converting original identifiers to surrogate values, applying cell sectionalization to group values into ranges, and adding error values to create final anonymized attributes. These parameter changes ensure that the anonymized attribute cannot be directly reverse-engineered to the original identifier, yet maintains combinability through the synchronization dictionary.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If big data is distributed to external systems for analysis, then data utility and analysis capability are improved, but the risk of personal information leakage increases

Engineering Contradiction:
Improvedata analysis efficiencyVSAvoidprivacy protection
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent applies anonymization processing in advance before data distribution to external systems. The synchronization dictionary is created and stored securely, while anonymized data sets are distributed with transformed attributes. This preliminary anonymization action enables external systems to perform analysis on distributed data without exposure to raw personal identifiers, maintaining privacy protection while enabling data utility.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11501020B2Method for anonymizing personal information in big data and combining anonymized data
Publication Date: 2022.11.15 BOALA CO LTD
  • US11501020B2 patent drawing
  • US11501020B2 patent drawing
  • US11501020B2 patent drawing

AI summary

A method for anonymizing and combining personal information in big data, by which big data is anonymized to be freely distributed to an external system without fear of personal information leakage and distributed data can be combined with each other is proposed.