Electronic Document Redaction Markup Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for redacting sensitive data from electronic documents do not effectively remove all instances of the data, including those hidden in the markup, leaving them accessible to unauthorized users.
Innovation Solution
A system and method that identifies and replaces both visible and invisible instances of sensitive data items in the markup of electronic documents with neutral data items, maintaining the document's file format and functionality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If only visible instances of sensitive data items are removed from the rendered version of the ED, then the redaction process is simple and fast, but unauthorized users can still access sensitive data by viewing the markup
Solution Approach 1:
The patent segments the redaction process into two distinct phases: (1) identifying and redacting visible instances of sensitive data in the rendered document, and (2) identifying and redacting hidden instances of sensitive data in the markup. This segmentation allows the system to comprehensively remove all sensitive data while maintaining a structured, manageable process that doesn't overwhelm the system with excessive complexity.
Solution Approach 2:
The patent applies preliminary action by first identifying all instances of sensitive data items in the markup before performing the actual redaction. The system scans the markup to locate hidden instances, prepares the redaction plan, and then executes the redaction. This preliminary identification ensures that no sensitive data is missed during redaction while maintaining process efficiency.
2Reliability
If all instances of sensitive data items in the markup are replaced, then complete redaction is achieved, but the document processing time and computational resources increase
Solution Approach 1:
The patent applies partial action by focusing redaction efforts specifically on identified instances of sensitive data items rather than processing the entire document uniformly. The system identifies only the relevant portions of the markup containing sensitive data and applies redaction selectively to those areas, achieving complete redaction of sensitive information while minimizing unnecessary processing of other document content.
Solution Approach 2:
The system performs self-service by automatically scanning the markup to identify hidden instances of sensitive data items without requiring manual intervention. The automated identification and redaction process eliminates the need for users to manually review and redact each instance, significantly reducing processing time while ensuring comprehensive redaction of all sensitive data.
3Productivity
If sensitive data items are removed from the rendered version only, then the redaction process is quick, but the document markup still contains accessible sensitive data
Solution Approach 1:
The patent applies preliminary anti-action by proactively identifying and neutralizing hidden instances of sensitive data in the markup before they can be accessed by unauthorized users. The system scans the markup structure, detects sensitive data items that are not visible in the rendered version, and redacts them in advance, preventing potential unauthorized access while maintaining efficient processing speeds.
Solution Approach 2:
The patent introduces an intermediary process that bridges the gap between visible and hidden data instances. The system uses markup analysis as an intermediary layer to detect hidden sensitive data items, allowing it to extend the redaction capability from only visible instances to all instances including hidden ones, thereby preventing unauthorized access while maintaining processing efficiency.
Data Source
AI summary
A method for redacting an electronic document (ED) having a file format, including: obtaining a request to redact a sensitive data item in the ED; identifying a first and a second instance of the sensitive data item in a markup of the ED, where the second instance of the sensitive data item is not visible in a rendered version of the ED; and generating a redacted ED having the file format by replacing the first and the second instance of the sensitive data item with a neutral data item.


