Document Anonymization via Placeholder Substitution

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for protecting user privacy by removing sensitive content from documents either fail to preserve the usefulness of the documents for analysis or do not adequately anonymize the information, leading to incomplete privacy protection.

Innovation Solution

A computer-implemented technique that replaces sensitive content with generic placeholder characters while preserving formatting and structure, and identifies properties like grammatical characteristics, to ensure privacy and enhance the accuracy of machine-learning analysis engines.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Object-affected harmful factors

If sensitive content is removed from documents using traditional methods, then privacy protection is improved, but the usefulness of documents for machine-learning analysis deteriorates

Engineering Contradiction:
Improveprivacy protectionVSAvoidusefulness for analysis
Core Design Contradiction:
Object-affected harmful factorsVSLoss of information

Solution Approach 1:

The patent introduces an intermediary layer between the original sensitive content and the analysis system. Generic placeholder characters serve as this intermediary, replacing sensitive information while preserving structural and formatting patterns that machine-learning engines need for training and analysis, thus resolving the contradiction between privacy protection and analytical usefulness

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent applies different treatment to different parts of the document: sensitive content is replaced with generic placeholders while non-sensitive structural elements (formatting, layout, patterns) are preserved. This local differentiation allows the document to simultaneously protect privacy in specific locations while maintaining overall analytical value

Inventive Principle:
Principle #3Local quality

2Object-affected harmful factors

If sensitive content is encrypted or sanitized, then privacy protection is improved, but the document structure and formatting are lost

Engineering Contradiction:
Improveprivacy protectionVSAvoiddocument structure
Core Design Contradiction:
Object-affected harmful factorsVSShape

Solution Approach 1:

The patent creates a copy of the document's structural and formatting properties while replacing only the sensitive content. The generic placeholder characters replicate the position, length, and formatting characteristics of the original sensitive content, preserving the document's shape and structure while achieving privacy protection

Inventive Principle:
Principle #26Copying

3Object-affected harmful factors

If all personal information is removed from documents, then privacy protection is improved, but machine-learning analysis accuracy deteriorates

Engineering Contradiction:
Improveprivacy protectionVSAvoidanalysis accuracy
Core Design Contradiction:
Object-affected harmful factorsVSMeasurement precision

Solution Approach 1:

The patent changes the parameter of the sensitive content from specific identifying information to generic placeholder characters that preserve structural parameters. This parameter transformation maintains the document's format, length, and positioning characteristics while removing personally identifiable information, thereby preserving analysis accuracy without compromising privacy

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP3655877B1Removing sensitive content from documents while preserving their usefulness for subsequent processing
Publication Date: 2021.07.21 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP3655877B1 patent drawingFigure 1
  • EP3655877B1 patent drawingFigure 2
  • EP3655877B1 patent drawingFigure 3

AI summary

A computer-implemented technique is described herein for removing sensitive content from documents in a manner that preserves the usefulness of the documents for subsequent analysis. For instance, the technique obscures sensitive content in the documents, while retaining meaningful information in the documents for subsequent processing by a machine-learning engine or other machine-implemented analysis mechanisms. According to one illustrative aspect, the technique removes sensitive content from documents using a modification strategy that is chosen based on one or more selection factors. One selection factor pertains to the nature of the processing that is to be performed on the documents after they have been anonymized.