Personal Information Anonymization via Grouped Data Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current anonymization methods for personal information lack objective evaluation and are time-consuming, especially when dealing with large datasets, as they rely on visual inspection by experts, making it difficult to determine the level of anonymization and potentially leaving non-anonymized data undetected.
Innovation Solution
A method and system for personal information anonymization that classifies data into groups based on commonality, performs anonymization processes, calculates ratios, and outputs recommended anonymization values, using a server-client architecture to facilitate efficient evaluation and enhancement of anonymization levels through graphical representations and automated processes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If visual inspection by experts is used to evaluate anonymization, then evaluation accuracy may be improved, but time consumption and operational complexity increase significantly
Solution Approach 1:
The patent replaces the mechanical visual inspection process by experts with an automated computer-based evaluation system. The system uses algorithms to automatically analyze anonymized data, calculate anonymity metrics, and generate evaluation reports, substituting human visual inspection with computational analysis that is both accurate and efficient.
Solution Approach 2:
The evaluation system enables the anonymization process to self-evaluate by automatically assessing the degree of anonymization without requiring external expert intervention. The system performs self-checks on the anonymized data quality, calculates relevant metrics, and provides feedback for improvement, making the evaluation process autonomous and efficient.
2Measurement precision
If visual inspection by experts is used to evaluate anonymization, then evaluation accuracy may be improved, but operational complexity and expertise requirements increase
Solution Approach 1:
The patent replaces the complex expert-based visual inspection process with an automated computer system that performs evaluation through algorithms. This substitution eliminates the need for specialized human expertise while maintaining evaluation accuracy, as the system automatically handles data analysis, metric calculation, and result generation.
Solution Approach 2:
The system creates standardized evaluation templates and algorithms that can be repeatedly applied to different anonymization cases. By copying proven evaluation methods into automated computational routines, the system maintains consistent evaluation standards without requiring expert knowledge for each new case, thereby simplifying operation.
3Reliability
If detailed anonymization evaluation is performed on large datasets, then evaluation thoroughness is improved, but processing time and computational resources increase
Solution Approach 1:
The patent divides the evaluation process into distinct segments or modules, each handling specific aspects of anonymization assessment. The system segments data processing into manageable units, applies evaluation algorithms to each segment, and aggregates results, enabling thorough evaluation of large datasets through systematic division of labor that maintains both completeness and efficiency.
Solution Approach 2:
The system performs preliminary processing and preparation of data before detailed evaluation, such as data cleaning, standardization, and preliminary anonymization checks. By preparing data in advance and organizing it into suitable formats, the system reduces the computational burden during the actual evaluation phase, enabling thorough analysis without excessive processing time.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A personal information anonymization method is disclosed. Each of a plurality of data including personal information is classified into any one of a plurality of groups based on a degree of commonality of the personal information. An anonymization process, that standardizes the personal information of each of data belonging to each of the groups, is performed for each of the groups. A total number of the data belonging to each of the groups is calculated for each of the groups. The plurality of the groups are classified based on the total number of the sets of the data. A classification result is output.