Mediation Application PII Detection via Pre-trained ML Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional Enterprise Service Buses (ESBs) do not effectively handle personally identifiable information (PII) during data transfer, as they lack mechanisms to identify and manage PII, and they miss opportunities to enrich data in transit, leading to potential security risks and compliance issues.
Innovation Solution
An Enhanced Data Orchestration (EDO) system that utilizes pre-trained machine learning models to convert, process, and reassemble messages, identifying PII and attaching metadata to ensure proper handling, while enabling in-flight analysis and enrichment of data, using a unified runtime environment and a graphical user interface for building data orchestration pipelines.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional ESB is used for data transfer, then system simplicity is maintained, but PII identification and management capabilities are lost
Solution Approach 1:
The patent embeds machine learning models and data orchestration capabilities within the existing ESB architecture. The EDO system nests multiple processing layers including format conversion applications, mediation applications with ML models, and message reassembly applications, allowing PII identification and management to be integrated without replacing the entire ESB system.
Solution Approach 2:
The patent introduces an Enhanced Data Orchestration (EDO) layer as an intermediary between the conventional ESB and message sources/destinations. This EDO layer acts as a mediator that intercepts messages, applies machine learning models for PII identification, performs format conversion, and reassembles messages with metadata, thereby adding PII management capabilities without directly modifying the core ESB.
2Reliability
If machine learning models are integrated into ESB, then PII identification capability is improved, but processing time increases
Solution Approach 1:
The patent employs pre-trained machine learning models that have already been trained on PII data before deployment. This preliminary training action allows the models to immediately identify PII in transit messages without requiring real-time training, significantly reducing processing time while maintaining high identification accuracy and compliance assurance.
Solution Approach 2:
The patent replaces manual or rule-based PII identification methods with automated machine learning-based detection. This substitution enables the system to process messages through intelligent models that can identify complex PII patterns faster and more accurately than conventional mechanical or manual approaches, reducing overall processing time while improving compliance assurance.
3Adaptability or versatility
If data format conversion is performed, then data compatibility is improved, but processing complexity increases
Solution Approach 1:
The patent implements a unified data format conversion application within the EDO layer that handles multiple format conversion scenarios through a single multi-functional component. This conversion application can transform various message formats (XML, JSON, CSV, etc.) into a standardized internal format and back, providing universal compatibility across different data sources and destinations while consolidating conversion logic to manage complexity.
4Loss of information
If metadata is attached to messages, then data enrichment is improved, but message size increases
Solution Approach 1:
The patent extracts only the essential metadata elements related to PII identification and data orchestration from the message processing pipeline. Rather than attaching comprehensive metadata about every message attribute, the system selectively extracts and attaches only the critical PII-related metadata (such as PII type, confidence score, and handling requirements), thereby enriching the data with necessary information while minimizing the increase in message size.
Data Source
AI summary
A method of enhanced data orchestration (EDO). The method comprises receiving a first message by a mediation application executing on a computer system and analyzing the first message by the mediation application based at least in part on invoking a machine learning (ML) model by the mediation application, wherein the analyzing determines that a feature of the first message is a probable first component of an item of personally identifiable information (PII). The method further comprises receiving a second message by the mediation application, determining by the mediation application that a feature of the second message when combined with the feature of the first message constitutes an item of PII, and treating the first message and the second message in accordance with predefined PII handling protocols by the mediation application.


