Layered Stochastic Anonymization for Data Security and Utility

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning technologies face challenges in anonymizing data without de-anonymization, which is essential for secure and efficient data analysis, especially in cloud computing environments where digital data is often associated with user identities.

Innovation Solution

The implementation of layered stochastic anonymization using generative models, such as variational autoencoders, to generate example datasets with a degree of similarity to the original data, combined with classifier models to evaluate and provide confidence scores for anonymization, facilitating improved data security and analysis while reducing costs and increasing accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional data anonymization methods are used, then data security is improved, but data utility and analysis accuracy deteriorate

Engineering Contradiction:
Improvedata securityVSAvoidanalysis accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent creates synthetic copies of original data through generative models that preserve statistical properties and relationships while removing identifying information. The generated data maintains utility for analysis while ensuring security, resolving the contradiction between data security and analysis accuracy

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system transforms data parameters by adjusting the degree of anonymization and similarity control. By changing parameters such as noise levels, transformation intensity, and generation constraints, the system optimizes both security and utility simultaneously

Inventive Principle:
Principle #35Parameter changes

2Reliability

If complex anonymization processes are applied, then data security is improved, but processing time and computational cost increase

Engineering Contradiction:
Improvedata securityVSAvoidprocessing speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system performs preliminary actions by pre-processing data to identify sensitive features and pre-configuring anonymization parameters before the actual anonymization process. This preparation reduces computational overhead during execution, improving processing speed while maintaining security

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces traditional mechanical anonymization techniques with machine learning-based generative models that automatically learn data patterns and generate anonymized versions. This substitution reduces manual intervention and optimizes processing efficiency while enhancing security

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If high similarity to original data is maintained, then data utility is improved, but risk of de-anonymization increases

Engineering Contradiction:
Improvedata utilityVSAvoidde-anonymization risk
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The system dynamically adjusts similarity parameters and anonymization intensity to optimize the balance between data utility and security. By controlling parameters such as feature preservation levels and noise injection rates, the system maintains high utility while preventing de-anonymization attacks

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11763188B2Layered stochastic anonymization of data
Publication Date: 2023.09.19 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11763188B2 patent drawing
  • US11763188B2 patent drawing
  • US11763188B2 patent drawing

AI summary

Techniques that facilitate layered stochastics anonymization of data are provided. In one example, a system includes a machine learning component and an evaluation component. The machine learning component performs a machine learning process for first data associated with one or more features to generate second data indicative of one or more example datasets within a degree of similarity to the first data. The first data and the second data comprise a corresponding data format. The evaluation component evaluates the second data for a particular feature from the one or more features and generates third data indicative of a confidence score for the second data.