ML Pipeline Obfuscation Using Endpoint Embeddings

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning-based data discovery and classification systems require transferring sensitive data in clear text, posing security risks and non-compliance with data privacy regulations.

Innovation Solution

A system and platform that classifies documents and discovers sensitive entities using a machine learning agent on an endpoint machine, generating embeddings that are ingested into a centralized server for model training, thereby obfuscating clear data and ensuring compliance with security policies.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If sensitive data is transferred in clear text for machine learning model training, then model training can be performed centrally, but data security is compromised and data privacy regulations are violated

Engineering Contradiction:
Improvemodel training efficiencyVSAvoiddata security
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent applies preliminary action by performing data embedding and obfuscation at the source endpoint before data leaves the secure environment. The ML agent embeds sensitive data locally and transforms it into obfuscated representations, ensuring data is protected before transmission. This prevents clear text data from ever being exposed during transit or storage in centralized servers.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary embedding layer that acts as a mediator between the sensitive data source and the centralized ML training system. This embedding layer transforms sensitive data into intermediate representations that retain useful features for training while removing personally identifiable information and sensitive content, thus protecting data privacy while enabling model training.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If data is obfuscated using embedding in the ML pipeline, then data security and privacy compliance are improved, but system complexity increases

Engineering Contradiction:
Improvedata securityVSAvoidML pipeline complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts the embedding and obfuscation functionality into a separate ML agent component that runs independently at the endpoint. This separates the complex data protection logic from the centralized training system, allowing the core ML training infrastructure to remain simple while the embedding agent handles security requirements locally.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent uses pre-trained embedding models that can be copied and deployed across multiple endpoint agents. Instead of implementing complex embedding logic in each location, standardized embedding models are replicated and executed locally, simplifying the overall system architecture while maintaining consistent data protection across distributed endpoints.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS12468983B2Machine Learning (ML) model pipeline with obfuscation to protect sensitive data therein
Publication Date: 2025.11.11 THALES DIS CPL USA INC
  • US12468983B2 patent drawing
  • US12468983B2 patent drawing
  • US12468983B2 patent drawing

AI summary

Provided is a system and platform for Machine Learning (ML) based Data Discovery and Classification. The system and platform comprising components of a user console, a ML agent, and a ML data engine. By way of a ML pipeline, sensitive data is obfuscated that would otherwise by in the clear when transmitted to a centralized server. The ML model pipeline decouples embedding from model training. In a first step, the ML Agent runs on data endpoint machine or proxy to convert clear text data to embedding vectors. In a second step, the ML data engine runs on a centralized server to train models using the embedding vectors. The separation of pipeline components and respective handling of workflow requests and messages associated therewith prevents the transfer of clear data in the open. Other embodiments disclosed.