Anonymized Data Repository via Cluster-Based PII Transformation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
As data repositories, especially in cloud computing, face increasing amounts of Personally Identifiable Information (PII), existing technologies struggle to efficiently anonymize and manage this data, making it difficult to analyze and report effectively while ensuring compliance with privacy regulations.
Innovation Solution
The implementation of an improved anonymization algorithm that applies one-way data transformations, such as data masking and morphing, and cryptographic hash functions to create anonymized data repositories, independent of the K value, using a cluster-based process to transform PII into non-identifiable information, thereby ensuring the anonymized data cannot be reversed.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional anonymization methods are used on PII data, then data privacy is protected, but data utility for analysis and reporting is degraded
Solution Approach 1:
The patent creates anonymized copies of the original PII data that can be used for analysis and reporting without compromising the security of the original data. The anonymized data repository serves as a surrogate that preserves statistical properties while removing identifying information, allowing multiple users to access and analyze data without exposing sensitive PII.
Solution Approach 2:
The system transforms PII data by changing its parameters through anonymization techniques such as generalization, suppression, and perturbation. These transformations modify the data characteristics to remove identifying information while maintaining the underlying statistical patterns and relationships needed for analysis.
2Productivity
If PII data is stored in cloud-based repositories, then data accessibility and processing capability are improved, but compliance with privacy regulations becomes more difficult
Solution Approach 1:
The system performs anonymization as a preliminary action before data is stored in or accessed from the cloud repository. By pre-anonymizing the data, the system ensures that PII is removed before cloud storage, simplifying compliance efforts and allowing the cloud infrastructure to process data without privacy concerns.
Solution Approach 2:
The anonymized data repository acts as an intermediary layer between the original PII data source and the cloud-based analysis tools. This intermediary preserves the necessary data characteristics for processing while blocking direct access to sensitive PII, thus enabling cloud productivity while maintaining privacy compliance.
3Reliability
If data is anonymized using existing technologies, then privacy protection is achieved, but the anonymization process is time-consuming and computationally intensive
Solution Approach 1:
The system implements automated anonymization processes that operate autonomously on incoming data streams, reducing manual intervention and processing time. The anonymized data repository can be automatically updated as new data arrives, eliminating the need for repeated manual anonymization efforts.
Solution Approach 2:
The system establishes anonymization templates and rules in advance, allowing data to be quickly anonymized using pre-defined parameters. This preliminary preparation of anonymization strategies significantly reduces the computational burden and time required for actual data processing.
Data Source
AI summary
A computing system includes an anonymizer server. The anonymizer server is communicatively coupled to a data repository configured to store a personal identification information (PII) data. The anonymizer server is configured to perform operations including receiving an anonymized data request, and creating an anonymized data repository based on the anonymized data request. The anonymizer server is also configured to perform operations including anonymizing the PII data to create an anonymized data by applying a cluster-based process, and storing the anonymized data in the anonymized data repository.


