Data Redaction via Machine Learning Detection and Masking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing computing systems fail to securely protect sensitive data as it is transmitted across various network layers, leaving it vulnerable to interception and resulting in potential fraud, identity theft, and consumer distrust.
Innovation Solution
Implementing a system that uses machine learning models to detect sensitive data at user devices and routers, applying masking techniques such as encryption and tokenization before transmission, and redacting sensitive data from display to ensure secure data protection across all network layers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If sensitive data is transmitted across network layers without masking, then data transmission efficiency is maintained, but security and reliability deteriorate due to vulnerability to interception
Solution Approach 1:
The system applies masking techniques (encryption, tokenization) to sensitive data before transmission across network layers. This preliminary action ensures data is protected before entering the network environment, preventing interception and maintaining security throughout transmission without requiring complex security infrastructure at each network layer.
Solution Approach 2:
The patent introduces an intermediary masking layer that sits between the data source and network transmission. This intermediary applies encryption and tokenization transformations, acting as a buffer that protects sensitive data from direct exposure during network transmission while maintaining system architecture simplicity.
2Reliability
If masking techniques are applied to sensitive data before transmission, then security is improved, but processing time and system complexity increase
Solution Approach 1:
Masking techniques are applied in advance before data enters the network transmission pipeline. By performing encryption and tokenization beforehand, the system avoids time-consuming security operations during active transmission, reducing latency while maintaining protection.
3Reliability
If sensitive data is redacted from display at user interface, then privacy protection is improved, but data accessibility and usability deteriorate
Solution Approach 1:
The system applies different quality levels of data protection to different contexts: full redaction for public displays, partial masking for authorized views, and complete visibility for authorized operations. This local differentiation maintains privacy where needed while preserving data accessibility for legitimate users.
Solution Approach 2:
The redaction level is dynamically adjusted based on user authorization, context, and operational needs. Authorized users can access unredacted data for legitimate purposes, while unauthorized users see redacted versions, making the system adaptive rather than static.
4Measurement precision
If machine learning models are deployed at multiple network layers, then detection accuracy is improved, but device complexity and computational requirements increase
Solution Approach 1:
The machine learning detection system is segmented and distributed across multiple network layers, with each layer having specialized detection capabilities. This segmentation allows targeted detection at appropriate levels without requiring all layers to have full detection functionality, reducing overall complexity while maintaining accuracy.
Data Source
AI summary
Methods, media, and systems are provided for centralized and decentralized protection of sensitive data. Sensitive data, for example, may include personal user data or data protected by data privacy regulations. A router may receive data from a user device. In an embodiment, the data may be received from an application that is managed by a Kubernetes cluster. In some embodiments, the application provides an active user interface that includes one or more sensitive fields for entering sensitive data. A machine learning model may be used to detect that the data being received by the router includes sensitive data. Additionally, the machine learning model may detect an entry of sensitive data at one or more sensitive fields on one or more active user interfaces. The machine learning model may be trained using a plurality of application programming interface requests. A masking technique may be applied to the sensitive data.


