AI Data Redaction With Text Mapping for Document Integrity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data redaction techniques struggle with inconsistent and resource-intensive removal of sensitive information from structured and unstructured data sets, often compromising document integrity due to varying data conversions.
Innovation Solution
An automated data redaction system using artificial intelligence to identify, map, and redact sensitive information in real-time from structured and unstructured data sets, maintaining document integrity through machine learning models and image processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If conventional data redaction techniques are used to handle large quantities of structured and unstructured data, then data redaction can be performed, but the process becomes resource-intensive and time-consuming
Solution Approach 1:
The patent replaces conventional mechanical data processing systems with an AI-based system that uses machine learning models to automatically detect, identify, and redact sensitive information. The system employs natural language processing and computer vision algorithms to handle both structured and unstructured data, eliminating the need for manual review and significantly reducing computational resource consumption while maintaining high accuracy.
Solution Approach 2:
The system performs self-service by automatically learning from training data to identify patterns of sensitive information without requiring human intervention. The machine learning models continuously improve their accuracy through feedback mechanisms, enabling the system to autonomously handle diverse data types and adapt to new redaction requirements without external assistance.
2Reliability
If conventional data redaction techniques are applied to ensure consistency across different data sets, then redaction can be standardized, but document integrity is compromised due to repeated data transformations
Solution Approach 1:
The patent segments the redaction process into distinct stages: data ingestion, AI-based analysis, sensitivity detection, and selective redaction. This segmentation allows each component to specialize in specific tasks, maintaining document integrity while ensuring consistent redaction application. The system processes different data types (structured and unstructured) through separate but coordinated pathways, preventing the need for repeated transformations.
Solution Approach 2:
The system changes the approach by using AI models that directly analyze data in its original format rather than requiring standardization to predetermined schemas. The machine learning models adapt to various data structures and formats, maintaining document integrity while achieving consistent redaction results across different data types without repeated transformations.
3Measurement precision
If manual data redaction is performed to ensure accuracy, then sensitive information can be identified, but the process becomes time-consuming and cannot keep pace with large data volumes
Solution Approach 1:
The patent replaces manual redaction processes with AI-based systems that use machine learning models trained to accurately identify sensitive information patterns. The system employs natural language processing for text data and computer vision for image data, achieving human-level or superior accuracy while processing millions of records simultaneously, thereby eliminating the time consumption associated with manual review.
Solution Approach 2:
The system maintains continuous operation by processing data streams in real-time without interruption. The machine learning models continuously analyze incoming data, making immediate redaction decisions, which ensures both high accuracy and rapid processing speed. The system can operate 24/7 without human fatigue, maintaining consistent performance across large data volumes.
Data Source
AI summary
A method for providing automated data redaction is disclosed. The method includes receiving a data set from a source, the data set including an unstructured electronic document and a structured electronic document; extracting, by using a model, texts from the data set; mapping each of the texts to a location in the data set, the location relating to a position of each of the texts on a page; determining, by using the model, whether the texts include restricted information; redacting the texts from the data set based on the mapping when the texts are determined to include restricted information; and generating a redacted data set based on a result of the redacting.


