ML Identifier Generation for Data Leakage Protection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The protection of user data in network environments is a challenge, particularly during content recommendation services, where direct transmission of user data to service providers increases the risk of data leakage and security breaches.
Innovation Solution
A computer-implemented method and system that generates an identifier for a user device's data record using a machine learning model, which is then sent to the service provider instead of the actual data record, thereby protecting user data from leakage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If user data is directly transmitted to service providers for content recommendation, then content recommendation accuracy is improved, but data security and risk of data leakage deteriorates
Solution Approach 1:
The patent introduces an intermediary identifier generation system that processes user data without exposing the raw data to service providers. The system generates identifiers that mediate between user data and recommendation services, allowing accurate recommendations while preventing direct access to sensitive user information.
Solution Approach 2:
The patent creates a copy or representation of user data in the form of generated identifiers. These identifiers capture the essential characteristics needed for recommendation accuracy while being safe to transmit and store, effectively replacing the transmission of actual sensitive user data.
2Object-affected harmful factors
If user data is protected from direct transmission, then data security is improved, but content recommendation personalization deteriorates
Solution Approach 1:
The patent replaces the mechanical system of direct data transmission with a computational system that generates identifiers through processing. This substitution allows the system to maintain personalization capabilities through algorithmic processing of identifier patterns rather than direct data access.
Solution Approach 2:
The patent transforms user data into different parameter representations through identifier generation. The generated identifiers contain encoded information that preserves personalization capabilities while changing the form and parameters of the data to eliminate security risks.
3Object-affected harmful factors
If identifiers are generated using machine learning models, then data protection is improved, but system complexity deteriorates
Solution Approach 1:
The patent performs preliminary actions by pre-training machine learning models to generate identifiers. This preliminary training phase separates the complexity of model development from runtime operations, allowing the system to protect data effectively while keeping runtime system complexity manageable.
Data Source
AI summary
Computer technology for protecting data security in a computerized system for recommending content to users where, a processing unit generates an identifier for a first data record relating to a user device based on a first machine learning model. Then, the processing unit sends the identifier to a service provider, and the service provider uses the identifier to determine one or more contents to be sent to the user device. Creating and using a decision tree machine learning (ML) model and a cluster ML model with training records and a transformed records.


