In-Memory Database On-the-Fly Anonymization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data anonymization techniques are computationally intensive, limiting their application to offline or small datasets, and preventing real-time processing and anonymization of large volumes of data, which is necessary for advanced business analytics and real-time data services.
Innovation Solution
Implementing an in-memory database system that determines privacy risks and anonymizes datasets in real-time by using anonymization algorithms to mask sensitive information, allowing for on-the-fly anonymization of data retrieved from databases.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional anonymization techniques are used, then privacy protection is improved, but processing speed and real-time capability deteriorate
Solution Approach 1:
The system pre-computes and stores anonymization mappings, generalization hierarchies, and privacy risk thresholds in the in-memory database before queries are executed. This preliminary preparation enables the anonymization engine to perform real-time anonymization by simply applying pre-defined rules and mappings during query execution, rather than computing anonymization algorithms on-the-fly for each query.
Solution Approach 2:
The system dynamically adjusts anonymization parameters such as generalization levels, k-anonymity thresholds, and privacy risk tolerance based on query characteristics and data sensitivity. By changing these parameters adaptively, the system optimizes the balance between privacy protection strength and processing speed for different query scenarios.
2Reliability
If anonymization is performed on large datasets, then privacy risk reduction is improved, but computational resource consumption increases
Solution Approach 1:
The system segments the anonymization process into distinct modular components: query analysis, privacy risk assessment, anonymization rule selection, and data transformation. Each component operates independently on specific portions of the data or query results, enabling parallel processing and reducing overall computational resource consumption while maintaining comprehensive privacy protection.
Solution Approach 2:
The system creates and maintains copies of anonymization rules, mappings, and processed data structures in in-memory caches. Instead of repeatedly computing the same anonymization transformations on large datasets, the system reuses pre-computed mappings and rules from memory, dramatically reducing computational resource consumption for subsequent queries on the same or similar data.
3Productivity
If real-time anonymization is implemented, then data utility for business applications is improved, but system complexity increases
Solution Approach 1:
The system introduces an intermediary anonymization engine layer between the database and business applications. This intermediary automatically handles privacy risk assessment and anonymization transformation without requiring complex integration with each application. The engine translates diverse query requirements into standardized anonymization operations, simplifying the overall system architecture while enabling real-time protected data access.
Solution Approach 2:
The anonymization engine is designed as a universal system that handles multiple types of queries, data formats, and privacy requirements through a single unified framework. It supports various anonymization techniques (generalization, suppression, perturbation, k-anonymity) and can adapt to different privacy policies and regulatory requirements, reducing system complexity by avoiding the need for separate specialized systems for each use case.
Data Source
AI summary
The method includes determining, using an in-memory database, a privacy risk associated with a resultant dataset of a query, returning, by the in-memory database, an anonymized dataset if the privacy risk is above a threshold value, the anonymized dataset being based on an anonymization, by the in-memory database, of the resultant dataset, and returning, by the in-memory database, the resultant dataset if the privacy risk is below a threshold value.


