Enterprise Data Masking for Secure External Gen AI Use
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The vulnerability of enterprise data when shared with externally hosted generative artificial intelligence models poses significant security and privacy risks, as it can be used to train AI models, posing a threat to business entities.
Innovation Solution
A processor-implemented method and system that preprocesses data using filtering operations such as prompt injection detection, profanity detection, toxicity and bias detection, and masking, followed by unmasking and further filtering to ensure secure data transmission and reception, while utilizing a cloud-agnostic framework for integration with AI applications.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If enterprise data is passed to externally hosted AI models to leverage their power, then the capability and functionality of Gen AI applications is improved, but data security and privacy are compromised
Solution Approach 1:
The patent introduces an intermediary system that sits between the enterprise data and externally hosted AI models. This intermediary performs preprocessing operations including masking sensitive information, detecting prompt injections, filtering profanity and toxicity, and validating outputs. The intermediary enables enterprise data to interact with external AI models while maintaining security boundaries, thus resolving the contradiction between leveraging external AI capabilities and protecting data security.
2Reliability
If data is masked to protect privacy, then data security is improved, but the usefulness and accuracy of AI model responses deteriorates
Solution Approach 1:
The patent applies masking selectively rather than uniformly across all data. The system identifies and masks only specific sensitive information (such as personally identifiable information, proprietary data) while leaving non-sensitive portions of the data unmasked. This localized approach to masking preserves the confidentiality of sensitive data while maintaining the contextual information needed for accurate AI model responses.
Solution Approach 2:
The system performs preliminary masking and filtering operations on input data before it reaches the AI model, and then performs unmasking and validation operations on the model's output. This preliminary action ensures that sensitive information is protected from the start, while the subsequent unmasking of outputs restores useful information for the enterprise, thus maintaining both confidentiality and accuracy.
3Reliability
If multiple filtering operations are performed on data and outputs, then data security and quality are improved, but processing time and system complexity increase
Solution Approach 1:
The patent divides the filtering and processing operations into distinct modular components: prompt injection detection, profanity detection, toxicity and bias detection, masking operations, and output validation. Each filtering operation is implemented as a separate module that can be independently configured and executed. This segmentation allows the system to perform comprehensive filtering while maintaining efficient processing through modular architecture and enables selective activation of different filtering operations based on specific needs.
Data Source
AI summary
The present disclosure herein addresses the problem of data security and privacy by providing a system and method for preserving privacy and security of enterprise data for generative artificial intelligence enabled applications. The system of the present disclosure enables removing all personal identifiable information (PII) and sensitive data by masking them in outgoing data keeping meaning of context and instructions same. A masked output is obtained as a response from external models. The response is further unmasked and an actual output is obtained for an end user. In this way, the system of the present disclosure protects an enterprise data from going out and keeps them secure and confidential. The system of the present disclosure also takes care of prompt injection, checks for truthfulness of the answers based on a given context, checks for malicious code in external model response and performs model scanning for not being compromised.


