Automated Sensitive Data Replacement in Documents

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Machine learning systems require large amounts of data from paper-based business documents, but processing these documents poses privacy concerns due to sensitive information, necessitating a method to preserve user privacy while allowing document usage in training data sets with minimal human intervention.

Innovation Solution

A system comprising pre-processing, initial processing, and data replacement stages to identify and replace sensitive data in documents, using image pre-processing, clustering, data type determination, and data replacement modules to generate and insert replacement data, ensuring privacy protection with minimal human intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If real-world documents containing sensitive information are used to train machine learning systems, then the quality and realism of training data is improved, but user privacy is compromised due to exposure of sensitive information

Engineering Contradiction:
Improvetraining data qualityVSAvoidprivacy exposure
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent extracts and removes sensitive information from documents while preserving the overall document structure and non-sensitive content. The system identifies sensitive data fields (such as personal names, addresses, phone numbers, financial information) and extracts them for replacement, allowing the document to retain its realistic appearance and structure for training purposes while eliminating privacy risks.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces an intermediary processing system that acts as a mediator between the original sensitive document and the training data set. This intermediary system automatically detects, redacts, and replaces sensitive information with placeholder or synthetic data, creating an intermediate version of the document that preserves structural integrity and non-sensitive information while protecting user privacy.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Object-affected harmful factors

If manual review and redaction of sensitive data is performed, then privacy protection is improved, but processing time and human intervention requirements increase

Engineering Contradiction:
Improveprivacy protectionVSAvoidprocessing time
Core Design Contradiction:
Object-affected harmful factorsVSLoss of time

Solution Approach 1:

The patent implements a self-service automated system that performs sensitive data detection and redaction without requiring manual human review. The system uses machine learning models and pattern recognition algorithms to automatically identify sensitive information types, locate them in documents, and perform appropriate redaction or replacement, enabling the process to serve itself and eliminating the need for time-consuming manual intervention.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual process of reviewing and redacting sensitive data with an automated computational system. Instead of human operators manually scanning documents and redacting information, the system uses optical character recognition, natural language processing, and trained classification models to automatically detect and redact sensitive information, dramatically reducing processing time while maintaining or improving privacy protection effectiveness.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Object-affected harmful factors

If all data in documents is replaced to ensure privacy, then privacy protection is improved, but the usefulness of training data for machine learning is reduced

Engineering Contradiction:
Improveprivacy protectionVSAvoidtraining data usefulness
Core Design Contradiction:
Object-affected harmful factorsVSLoss of information

Solution Approach 1:

The patent applies local quality by selectively replacing only the sensitive portions of documents while preserving the rest of the content. Instead of blanket replacement of all data, the system identifies specific sensitive fields (such as personal identifiers, financial information, health data) and replaces only those areas, leaving non-sensitive structural elements, business logic, and contextual information intact, thereby maintaining both privacy protection and training data usefulness.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent inverts the traditional approach by not removing or obscuring sensitive data entirely, but rather replacing it with synthetic or placeholder data that maintains the document's structural integrity and data relationships. This inversion allows the training system to learn from the document structure and patterns without actually exposing real sensitive information, achieving privacy protection while preserving training value.

Inventive Principle:
Principle #13The other way round (Inversion)

Data Source

PatentUS12111953B2Sensitive data detection and replacement
Publication Date: 2024.10.08 SERVICENOW INC
  • US12111953B2 patent drawing
  • US12111953B2 patent drawing
  • US12111953B2 patent drawing

AI summary

Systems and methods for privacy and sensitive data protection. An image of a document is received at a pre-processing stage and image pre-processing is applied to the image to ensure that the resulting image is sufficient for further processing. Pre-processing may involve processing relating to image quality and image orientation. The image is then passed to an initial processing stage. At the initial processing stage, the relevant data in the document are located and bounding boxes are placed around the data. The resulting image is then passed to a processing stage. At this stage, the type of data within the bounding boxes is determined and suitable replacement data is generated. The replacement data is then inserted into the image to thereby remove and replace the sensitive data in the image.