Enterprise AI Data Masking for Secure External Model Use
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The vulnerability of enterprise data when shared with externally hosted generative artificial intelligence models poses significant security and privacy risks, as it can be used to train AI models, compromising business entities.
Innovation Solution
A processor-implemented method and system that preprocesses data using filtering operations and masking techniques, including prompt injection detection, profanity and toxicity/bias detection, and contextual correctness checks, to ensure secure data handling within an enterprise network, while integrating with externally hosted AI models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If enterprise data is passed to externally hosted AI models to leverage advanced AI capabilities, then AI model performance and functionality is improved, but data security and privacy risk increases
Solution Approach 1:
The patent introduces an intermediary security processing layer between the enterprise data and the externally hosted AI model. This intermediary layer performs filtering operations to remove sensitive information, masking operations to anonymize data, and validation operations to ensure data safety before transmission to the external AI model, thus resolving the contradiction by enabling AI functionality while preventing data exposure
Solution Approach 2:
The patent applies preliminary security actions by performing filtering, masking, and validation operations on enterprise data before it is transmitted to the externally hosted AI model. These preliminary actions prepare the data in advance to remove or protect sensitive information, allowing the data to be safely processed by the AI model without exposing security risks
2Object-affected harmful factors
If filtering and masking operations are performed on enterprise data before AI processing, then data privacy and security is improved, but data processing complexity increases
Solution Approach 1:
The patent segments the data processing pipeline into distinct functional modules: filtering operations module, masking operations module, and validation operations module. Each module performs a specific security function independently, making the complex security processing manageable and modular while maintaining effective data privacy protection through coordinated operation of these segmented functions
3Reliability
If multiple filtering operations are performed on data and AI outputs, then security coverage is improved, but processing time increases
Solution Approach 1:
The patent implements continuous security validation throughout the entire data processing lifecycle, applying filtering operations to input data, masking operations to protect sensitive information, and validation operations to both input and output data. This continuous approach ensures comprehensive security coverage without requiring separate discrete security steps, thereby maintaining processing efficiency while enhancing reliability
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The present disclosure herein addresses the problem of data security and privacy by providing a system and method for preserving privacy and security of enterprise data for generative artificial intelligence enabled applications. The system of the present disclosure enables removing all personal identifiable information (PII) and sensitive data by masking them in outgoing data keeping meaning of context and instructions same. A masked output is obtained as a response from external models. The response is further unmasked and an actual output is obtained for an end user. In this way, the system of the present disclosure protects an enterprise data from going out and keeps them secure and confidential. The system of the present disclosure also takes care of prompt injection, checks for truthfulness of the answers based on a given context, checks for malicious code in external model response and performs model scanning for not being compromised.