Document Image Sanitization for ML Training Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Machine learning systems require large amounts of data from paper-based business documents, but processing these documents poses privacy concerns due to sensitive information that needs to be protected.
Innovation Solution
A system and method that utilize an execution module to identify, delineate, and replace sensitive data in document images, with user validation and feedback used to refine the module's training, ensuring privacy preservation while enabling the use of real-world documents for training data sets.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If real world documents are used to train machine learning systems, then the quantity and quality of training data is improved, but privacy and security are worsened due to exposure of sensitive information
Solution Approach 1:
The patent extracts only the non-sensitive portions of documents for training while removing sensitive information through redaction. The execution module identifies and removes sensitive data elements (PII, financial data, etc.) from documents before using them for training machine learning models, thus separating the useful training data from the harmful sensitive information.
Solution Approach 2:
The patent introduces an intermediary redaction system that processes documents between the source and the training data creation. This intermediary module (execution module with submodules) acts as a mediator that sanitizes documents by identifying and removing sensitive information, allowing safe transfer of data from real-world sources to training datasets.
2Object-affected harmful factors
If sensitive data is removed from documents, then privacy is improved, but the utility of the document for training is worsened due to loss of information
Solution Approach 1:
The patent applies local quality by selectively redacting only specific sensitive portions of documents while preserving non-sensitive content. Different regions of the document are treated differently: sensitive areas (names, addresses, financial data) are removed or masked, while non-sensitive areas (document structure, business logic, contextual information) are preserved for training purposes.
Solution Approach 2:
The patent changes the parameter of data sensitivity through intelligent identification and selective removal. The execution module analyzes document content to determine which portions exceed sensitivity thresholds and applies transformation (redaction, masking, or replacement) only to those portions, maintaining the overall utility of the document for training while achieving privacy protection.
3Productivity
If automated redaction is used, then productivity is improved, but accuracy is worsened due to potential misidentification of sensitive data
Solution Approach 1:
The patent implements feedback mechanisms where users review and validate the automated redaction results. The validation module allows users to correct misidentifications and refine the redaction process. This feedback loop enables the system to learn from user corrections and improve its automated detection accuracy over time, balancing speed with precision.
Solution Approach 2:
The patent applies partial action by providing automated redaction as an initial pass that covers most cases, then allowing selective manual review only for uncertain or critical portions. This approach maintains high productivity through automation while ensuring accuracy through targeted human validation where needed, rather than requiring full manual review of all documents.
Data Source
AI summary
Systems and methods relating to the replacement or removal of sensitive data in images of documents. An initial image of a document with sensitive data is received at an execution module and changes are made based on the execution module's training. The changes include replacing or effectively removing the sensitive data from the image of the document. The resulting sanitized image is then sent to a user for validation of the changes. The feedback from the user is then used in training the execution module to refine its behaviour when applying changes to other initial images of documents. To train the execution module, training data sets of document images with sensitive data manually tagged by users are used. The execution module thus learns to identify sensitive data and its submodules replace that sensitive data with suitable replacement data. The feedback from the user works to improve the resulting sanitized images from the execution module.


