Rule-Based Dataset Anonymization via Privacy and Consistency Metrics
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data anonymization techniques rely on human evaluation, which is tedious, unreliable, and inconsistent, often compromising between privacy and utility, and may lead to insufficient protection, resulting in re-identification of sensitive information.
Innovation Solution
A system and method utilizing a processor with a data privacy evaluator and rules engine that evaluates anonymized datasets based on privacy and consistency metrics, generating a final output to determine the extent of variation from the original dataset, allowing automated access or feedback to ensure desired privacy protection and utility.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data anonymization is applied to protect privacy, then privacy protection is improved, but data utility deteriorates
Solution Approach 1:
The system changes the parameters of anonymization by introducing multiple granularity levels (coarse, medium, fine) that allow dynamic adjustment of the degree of anonymization. This enables optimization of the trade-off between privacy protection and data utility by selecting appropriate granularity levels for different data elements and use cases.
Solution Approach 2:
The system implements dynamic evaluation of anonymized datasets using automated metrics and models that assess both privacy protection effectiveness and data utility preservation. This dynamic assessment allows for adaptive adjustment of anonymization strategies to achieve optimal balance between the two competing objectives.
2Measurement precision
If manual evaluation is used to assess anonymization effectiveness, then evaluation accuracy may be improved, but time consumption and complexity increase
Solution Approach 1:
The system replaces manual mechanical evaluation with automated computational evaluation using processors that execute predefined metrics and models. This substitution maintains evaluation accuracy through systematic assessment while dramatically reducing time consumption and eliminating human subjectivity and inconsistency.
Solution Approach 2:
The system enables self-service evaluation where the anonymization system automatically assesses its own output using built-in metrics and models. This self-evaluation capability provides continuous, consistent, and scalable assessment without requiring external manual intervention, thereby maintaining accuracy while reducing time and resource requirements.
3Object-affected harmful factors
If strong anonymization is applied to prevent re-identification, then privacy protection is improved, but data value for modeling deteriorates
Solution Approach 1:
The system applies different levels of anonymization to different parts of the dataset based on their sensitivity and use requirements. By implementing local quality variations through multiple granularity levels, the system protects sensitive information that requires strong anonymization while preserving the value of less sensitive data that can support meaningful modeling, thereby reducing re-identification risk without unduly compromising modeling capability.
Data Source
AI summary
Embodiments herein facilitate a rule-based anonymization of an original dataset. The system may include a processor including a data privacy evaluator and a rules engine. The data privacy evaluator may receive at least one anonymized dataset corresponding to a predefined strategy of anonymization. The at least one anonymized dataset may include a variation from the original dataset by at least one of a privacy metric and a consistency metric. The data privacy evaluator may evaluate the at least one anonymized dataset and may generate a final output value based on a first output and a second output. The processor may assess the final output value with respect to a predefined threshold through the rules engine. If the final output value may be equal or higher than the predefined threshold, the system may permit an access to the anonymized dataset.


