Document Image Sanitization for ML Training Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Machine learning systems require large amounts of data from paper-based business documents, but processing these documents poses privacy concerns due to sensitive information that needs to be protected.

Innovation Solution

A system and method that utilize an execution module to identify, delineate, and replace sensitive data in document images, with user validation and feedback used to refine the module's training, ensuring privacy preservation while enabling the use of real-world documents for training data sets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If real world documents are used to train machine learning systems, then the quantity and quality of training data is improved, but privacy and security are worsened due to exposure of sensitive information

Engineering Contradiction:
Improvetraining data quantityVSAvoidprivacy risk
Core Design Contradiction:
Quantity of substanceVSObject-affected harmful factors

Solution Approach 1:

The patent extracts only the non-sensitive portions of documents for training while removing sensitive information through redaction. The execution module identifies and removes sensitive data elements (PII, financial data, etc.) from documents before using them for training machine learning models, thus separating the useful training data from the harmful sensitive information.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces an intermediary redaction system that processes documents between the source and the training data creation. This intermediary module (execution module with submodules) acts as a mediator that sanitizes documents by identifying and removing sensitive information, allowing safe transfer of data from real-world sources to training datasets.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Object-affected harmful factors

If sensitive data is removed from documents, then privacy is improved, but the utility of the document for training is worsened due to loss of information

Engineering Contradiction:
Improveprivacy protectionVSAvoiddata utility
Core Design Contradiction:
Object-affected harmful factorsVSLoss of information

Solution Approach 1:

The patent applies local quality by selectively redacting only specific sensitive portions of documents while preserving non-sensitive content. Different regions of the document are treated differently: sensitive areas (names, addresses, financial data) are removed or masked, while non-sensitive areas (document structure, business logic, contextual information) are preserved for training purposes.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent changes the parameter of data sensitivity through intelligent identification and selective removal. The execution module analyzes document content to determine which portions exceed sensitivity thresholds and applies transformation (redaction, masking, or replacement) only to those portions, maintaining the overall utility of the document for training while achieving privacy protection.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If automated redaction is used, then productivity is improved, but accuracy is worsened due to potential misidentification of sensitive data

Engineering Contradiction:
Improveprocessing speedVSAvoidredaction accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent implements feedback mechanisms where users review and validate the automated redaction results. The validation module allows users to correct misidentifications and refine the redaction process. This feedback loop enables the system to learn from user corrections and improve its automated detection accuracy over time, balancing speed with precision.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent applies partial action by providing automated redaction as an initial pass that covers most cases, then allowing selective manual review only for uncertain or critical portions. This approach maintains high productivity through automation while ensuring accuracy through targeted human validation where needed, rather than requiring full manual review of all documents.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12182308B2Removal of sensitive data from documents for use as training sets
Publication Date: 2024.12.31 SERVICENOW INC
  • US12182308B2 patent drawing
  • US12182308B2 patent drawing
  • US12182308B2 patent drawing

AI summary

Systems and methods relating to the replacement or removal of sensitive data in images of documents. An initial image of a document with sensitive data is received at an execution module and changes are made based on the execution module's training. The changes include replacing or effectively removing the sensitive data from the image of the document. The resulting sanitized image is then sent to a user for validation of the changes. The feedback from the user is then used in training the execution module to refine its behaviour when applying changes to other initial images of documents. To train the execution module, training data sets of document images with sensitive data manually tagged by users are used. The execution module thus learns to identify sensitive data and its submodules replace that sensitive data with suitable replacement data. The feedback from the user works to improve the resulting sanitized images from the execution module.