AI Pseudonymization Recommendation for Consistent Column Masking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The selection of pseudonymization techniques is heavily biased towards individual user experience and preferences, leading to burdensome job handovers and inefficient use of human resources due to changes in personnel, and requires cumbersome historical data retrieval.
Innovation Solution
An automatic pseudonymization technique recommendation method using artificial intelligence, involving training a first learning model with numeric vectors generated from column names, data types, and industry classifications to recommend pseudonymization techniques based on a trained decision tree model and word embedding model.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If pseudonymization techniques are selected based on individual user experience and preferences, then the selection reflects user expertise, but the process becomes burdensome during personnel changes and consumes excessive time and resources
Solution Approach 1:
The patent creates a digital copy of pseudonymization expertise by training an AI model on historical pseudonymization decisions, column characteristics, and technique selections. This digital knowledge base replicates the decision-making patterns of experienced users, allowing the system to recommend techniques automatically without requiring human experts to manually review or transfer their knowledge during personnel changes
Solution Approach 2:
The system performs preliminary action by pre-processing column data into numeric vectors and pre-training the AI model with historical pseudonymization data before actual pseudonymization tasks begin. This advance preparation enables the system to quickly retrieve and recommend appropriate techniques during personnel transitions without requiring time-consuming manual analysis of processing history
2Loss of information
If manual lookup of processing history is required to find past pseudonymization techniques, then accurate historical information can be retrieved, but the process becomes cumbersome and inefficient
Solution Approach 1:
The patent replaces the mechanical manual lookup process with an automated AI-based recommendation system. Instead of requiring users to manually search through processing history and pseudonymization records, the system automatically processes column numeric vectors through the trained AI model to retrieve and recommend appropriate past pseudonymization techniques, significantly improving ease of operation while maintaining access to historical information
3Ease of manufacture
If pseudonymization technique selection relies on individual user preferences, then the process is simple for experienced users, but it creates dependency on specific personnel and reduces adaptability to organizational changes
Solution Approach 1:
The patent creates a universal pseudonymization recommendation system that serves multiple users and purposes. The AI model is trained on aggregated historical data from multiple users and scenarios, enabling it to provide consistent, high-quality recommendations across different personnel and organizational contexts. This universal system eliminates dependency on any single user's preferences while maintaining the simplicity and expertise-based quality of technique selection
Data Source
AI summary
Provided is an automatic pseudonymization technique recommendation method using artificial intelligence, including receiving a training dataset including a plurality of pieces of data including a column name, a data type, a pseudonymization technique recommendation, adding, a data type vector obtained from the data type to a word vector obtained corresponding to the column name to obtain a numeric vector, and labeling the numeric vector with the pseudonymization technique recommendation to generate a plurality of pieces of training data, training a first learning model using the plurality of pieces of training data so that the first learning model outputs a pseudonymization technique recommendation in response to an input of the numeric vector, obtaining a numeric vector corresponding to each column of the dataset to be pseudonymized, and inputting the numeric vector obtained into the trained first learning model to obtain a pseudonymization technique recommendation for each column of the dataset.


