Document Anonymization via Placeholder Substitution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for protecting user privacy by removing sensitive content from documents either fail to preserve the usefulness of the documents for analysis or do not adequately anonymize the information, leading to incomplete privacy protection.
Innovation Solution
A computer-implemented technique that replaces sensitive content with generic placeholder characters while preserving formatting and structure, and identifies properties like grammatical characteristics, to ensure privacy and enhance the accuracy of machine-learning analysis engines.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If sensitive content is removed from documents using traditional methods, then privacy protection is improved, but the usefulness of documents for machine-learning analysis deteriorates
Solution Approach 1:
The patent introduces an intermediary layer between the original sensitive content and the analysis system. Generic placeholder characters serve as this intermediary, replacing sensitive information while preserving structural and formatting patterns that machine-learning engines need for training and analysis, thus resolving the contradiction between privacy protection and analytical usefulness
Solution Approach 2:
The patent applies different treatment to different parts of the document: sensitive content is replaced with generic placeholders while non-sensitive structural elements (formatting, layout, patterns) are preserved. This local differentiation allows the document to simultaneously protect privacy in specific locations while maintaining overall analytical value
2Object-affected harmful factors
If sensitive content is encrypted or sanitized, then privacy protection is improved, but the document structure and formatting are lost
Solution Approach 1:
The patent creates a copy of the document's structural and formatting properties while replacing only the sensitive content. The generic placeholder characters replicate the position, length, and formatting characteristics of the original sensitive content, preserving the document's shape and structure while achieving privacy protection
3Object-affected harmful factors
If all personal information is removed from documents, then privacy protection is improved, but machine-learning analysis accuracy deteriorates
Solution Approach 1:
The patent changes the parameter of the sensitive content from specific identifying information to generic placeholder characters that preserve structural parameters. This parameter transformation maintains the document's format, length, and positioning characteristics while removing personally identifiable information, thereby preserving analysis accuracy without compromising privacy
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A computer-implemented technique is described herein for removing sensitive content from documents in a manner that preserves the usefulness of the documents for subsequent analysis. For instance, the technique obscures sensitive content in the documents, while retaining meaningful information in the documents for subsequent processing by a machine-learning engine or other machine-implemented analysis mechanisms. According to one illustrative aspect, the technique removes sensitive content from documents using a modification strategy that is chosen based on one or more selection factors. One selection factor pertains to the nature of the processing that is to be performed on the documents after they have been anonymized.