AI Training Document Redaction Preserving Topology
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing document redaction techniques fail to effectively remove sensitive information while maintaining the document's topology and file size, making it difficult to train artificial intelligence models without compromising security.
Innovation Solution
A method that parses unredacted documents to identify sensitive information using a predetermined rule set, substituting it with placeholder information to generate a redacted document with a similar topology and file size, allowing AI models to be trained on the redacted data without compromising security.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If existing document redaction techniques are used to remove sensitive information, then security is improved, but the document topology and file size are altered making AI training difficult
Solution Approach 1:
The patent creates a redacted copy of the original document that preserves the topological structure, formatting, and file size while replacing sensitive information with placeholder data. This copying approach allows AI models to be trained on the redacted document without exposing sensitive information, while maintaining the document's structural integrity for effective training.
Solution Approach 2:
The patent changes the content parameters of the document by substituting sensitive information with placeholder text, while deliberately maintaining other parameters such as document topology, file size, and formatting. This selective parameter modification resolves the contradiction by preserving structural parameters needed for AI training while changing content parameters to ensure security.
2Reliability
If sensitive information is removed from documents, then security is improved, but AI model training effectiveness is reduced
Solution Approach 1:
The patent creates a redacted copy of the original document that preserves the topological structure, formatting, and file size while replacing sensitive information with placeholder data. This copying approach allows AI models to be trained on the redacted document without exposing sensitive information, while maintaining the document's structural integrity for effective training.
Solution Approach 2:
The patent introduces placeholder information as an intermediary element that occupies the space where sensitive information would normally appear. This intermediary maintains the document's structure and file size, allowing AI models to learn from the document's topology and patterns without actually processing sensitive data, thus bridging the gap between security requirements and training effectiveness.
3Productivity
If document structure is preserved for AI training, then AI model training effectiveness is improved, but sensitive information may be exposed
Solution Approach 1:
The patent applies different treatments to different parts of the document: sensitive information is replaced with placeholder text while non-sensitive structural elements are preserved. This local differentiation allows the document to maintain its topological quality for AI training while removing harmful sensitive information from specific locations within the document.
Solution Approach 2:
The patent creates a redacted copy of the original document that preserves the topological structure, formatting, and file size while replacing sensitive information with placeholder data. This copying approach allows AI models to be trained on the redacted document without exposing sensitive information, while maintaining the document's structural integrity for effective training.
Data Source
AI summary
Systems and methods are provided herein for redaction of artificial intelligence (AI) training documents. Data comprising an unredacted document is received. The unredacted document comprises a plurality of objects arranged according to a first topology. The unredacted document is parsed to identify objects either directly or relationally containing user sensitive information using a predetermined rule set based on the first topology. The user sensitive information within the unredacted document is substituted with placeholder information to generate a redacted document having a second topology. The second topology is substantially identical to the first topology. In some variations, the redacted document is provided to an AI model for training.


