Layered Stochastic Anonymization for Data Security and Utility
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine learning technologies face challenges in anonymizing data without de-anonymization, which is essential for secure and efficient data analysis, especially in cloud computing environments where digital data is often associated with user identities.
Innovation Solution
The implementation of layered stochastic anonymization using generative models, such as variational autoencoders, to generate example datasets with a degree of similarity to the original data, combined with classifier models to evaluate and provide confidence scores for anonymization, facilitating improved data security and analysis while reducing costs and increasing accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional data anonymization methods are used, then data security is improved, but data utility and analysis accuracy deteriorate
Solution Approach 1:
The patent creates synthetic copies of original data through generative models that preserve statistical properties and relationships while removing identifying information. The generated data maintains utility for analysis while ensuring security, resolving the contradiction between data security and analysis accuracy
Solution Approach 2:
The system transforms data parameters by adjusting the degree of anonymization and similarity control. By changing parameters such as noise levels, transformation intensity, and generation constraints, the system optimizes both security and utility simultaneously
2Reliability
If complex anonymization processes are applied, then data security is improved, but processing time and computational cost increase
Solution Approach 1:
The system performs preliminary actions by pre-processing data to identify sensitive features and pre-configuring anonymization parameters before the actual anonymization process. This preparation reduces computational overhead during execution, improving processing speed while maintaining security
Solution Approach 2:
The patent replaces traditional mechanical anonymization techniques with machine learning-based generative models that automatically learn data patterns and generate anonymized versions. This substitution reduces manual intervention and optimizes processing efficiency while enhancing security
3Measurement precision
If high similarity to original data is maintained, then data utility is improved, but risk of de-anonymization increases
Solution Approach 1:
The system dynamically adjusts similarity parameters and anonymization intensity to optimize the balance between data utility and security. By controlling parameters such as feature preservation levels and noise injection rates, the system maintains high utility while preventing de-anonymization attacks
Data Source
AI summary
Techniques that facilitate layered stochastics anonymization of data are provided. In one example, a system includes a machine learning component and an evaluation component. The machine learning component performs a machine learning process for first data associated with one or more features to generate second data indicative of one or more example datasets within a degree of similarity to the first data. The first data and the second data comprise a corresponding data format. The evaluation component evaluates the second data for a particular feature from the one or more features and generates third data indicative of a confidence score for the second data.


