Big Data Anonymization via Surrogate Synchronization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional anonymization methods for big data fail to prevent re-identification of personal information, making it difficult to freely distribute and combine anonymized data across systems without risking personal information leakage.
Innovation Solution
A method involving a data server with a communication, processing, and storage unit that anonymizes personal information by converting synchronization target personal identification attribute values into surrogate values, applying error values to group them into cells, and assigning synchronization attributes based on section information, making re-identification extremely hard and enabling secure distribution and combination of data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional anonymization methods (masking, substitution, semi-identification, categorization) are applied to big data, then personal information is anonymized to some extent, but re-identification is still possible and data combination across systems remains difficult
Solution Approach 1:
The patent segments the anonymization process into multiple distinct steps: generating surrogate values from original personal identifiers, creating synchronization dictionaries that map original values to surrogates, applying cell sectionalization to group surrogate values, and adding error values to create final anonymized attributes. This multi-stage segmentation ensures thorough anonymization while preserving combinability through the synchronization dictionary.
Solution Approach 2:
The synchronization dictionary acts as an intermediary between original personal identification data and anonymized data. It contains the mapping relationships that enable data combination across systems without exposing actual personal identifiers. The dictionary serves as a mediator that allows re-identification only by authorized parties with access to the original data, while preventing unauthorized re-identification of anonymized data.
2Object-affected harmful factors
If personal identification attributes are removed from data sets to prevent re-identification, then privacy protection is improved, but the ability to combine records of the same person across different data sets is lost
Solution Approach 1:
The patent creates a copy of the personal identification attribute in the form of a synchronization attribute that derives from the original identifier but is transformed through surrogate values and cell sectionalization. This copied attribute maintains the ability to link records of the same person across data sets while being sufficiently transformed to prevent direct re-identification. The synchronization dictionary preserves the mapping relationship for authorized re-identification when needed.
Solution Approach 2:
The patent transforms the personal identification attribute through multiple parameter changes: converting original identifiers to surrogate values, applying cell sectionalization to group values into ranges, and adding error values to create final anonymized attributes. These parameter changes ensure that the anonymized attribute cannot be directly reverse-engineered to the original identifier, yet maintains combinability through the synchronization dictionary.
3Productivity
If big data is distributed to external systems for analysis, then data utility and analysis capability are improved, but the risk of personal information leakage increases
Solution Approach 1:
The patent applies anonymization processing in advance before data distribution to external systems. The synchronization dictionary is created and stored securely, while anonymized data sets are distributed with transformed attributes. This preliminary anonymization action enables external systems to perform analysis on distributed data without exposure to raw personal identifiers, maintaining privacy protection while enabling data utility.
Data Source
AI summary
A method for anonymizing and combining personal information in big data, by which big data is anonymized to be freely distributed to an external system without fear of personal information leakage and distributed data can be combined with each other is proposed.


