ML Pipeline Obfuscation Using Endpoint Embeddings
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine learning-based data discovery and classification systems require transferring sensitive data in clear text, posing security risks and non-compliance with data privacy regulations.
Innovation Solution
A system and platform that classifies documents and discovers sensitive entities using a machine learning agent on an endpoint machine, generating embeddings that are ingested into a centralized server for model training, thereby obfuscating clear data and ensuring compliance with security policies.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If sensitive data is transferred in clear text for machine learning model training, then model training can be performed centrally, but data security is compromised and data privacy regulations are violated
Solution Approach 1:
The patent applies preliminary action by performing data embedding and obfuscation at the source endpoint before data leaves the secure environment. The ML agent embeds sensitive data locally and transforms it into obfuscated representations, ensuring data is protected before transmission. This prevents clear text data from ever being exposed during transit or storage in centralized servers.
Solution Approach 2:
The patent introduces an intermediary embedding layer that acts as a mediator between the sensitive data source and the centralized ML training system. This embedding layer transforms sensitive data into intermediate representations that retain useful features for training while removing personally identifiable information and sensitive content, thus protecting data privacy while enabling model training.
2Reliability
If data is obfuscated using embedding in the ML pipeline, then data security and privacy compliance are improved, but system complexity increases
Solution Approach 1:
The patent extracts the embedding and obfuscation functionality into a separate ML agent component that runs independently at the endpoint. This separates the complex data protection logic from the centralized training system, allowing the core ML training infrastructure to remain simple while the embedding agent handles security requirements locally.
Solution Approach 2:
The patent uses pre-trained embedding models that can be copied and deployed across multiple endpoint agents. Instead of implementing complex embedding logic in each location, standardized embedding models are replicated and executed locally, simplifying the overall system architecture while maintaining consistent data protection across distributed endpoints.
Data Source
AI summary
Provided is a system and platform for Machine Learning (ML) based Data Discovery and Classification. The system and platform comprising components of a user console, a ML agent, and a ML data engine. By way of a ML pipeline, sensitive data is obfuscated that would otherwise by in the clear when transmitted to a centralized server. The ML model pipeline decouples embedding from model training. In a first step, the ML Agent runs on data endpoint machine or proxy to convert clear text data to embedding vectors. In a second step, the ML data engine runs on a centralized server to train models using the embedding vectors. The separation of pipeline components and respective handling of workflow requests and messages associated therewith prevents the transfer of clear data in the open. Other embodiments disclosed.


