Document Data Redaction via Segmentation and Mediator Principles
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge lies in analyzing and improving document data collections while maintaining user privacy, as human reviewers are prohibited from accessing private information, complicating the process of tuning algorithms and evaluating service processes that utilize these collections.
Innovation Solution
A method involving a data processing apparatus that receives a document data collection with fixed phrases not posing a personal information exposure risk, extracts candidate phrases from user-donated documents with explicit permission, generates a redacted collection by obfuscating or removing non-matching phrases, and allows human reviewers to examine only the matched phrases, thereby ensuring privacy protection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If access to document data collection is precluded to protect user privacy, then user privacy is protected, but the ability to analyze and improve the quality of the document data collection deteriorates
Solution Approach 1:
The patent segments the document data collection into two distinct parts: fixed phrases (reusable, non-private content patterns) and variable content (user-specific private information). By separating these components, the system enables human reviewers to access and analyze the fixed phrases portion while automatically redacting or masking the variable content, thus resolving the contradiction between privacy protection and analytical accessibility
Solution Approach 2:
The patent introduces an intermediary automated processing system that acts as a mediator between the private document data collection and human reviewers. This intermediary automatically identifies, extracts, and redacts personal information from documents, generating a sanitized version that human reviewers can safely access and analyze without exposure to private data, thereby enabling analysis while maintaining privacy protection
2Reliability
If all personal information is removed from document data collection, then privacy protection is ensured, but the quality and usefulness of the data collection for service improvement deteriorates
Solution Approach 1:
The patent segments information into fixed phrases (non-private, reusable patterns) and variable content (user-specific data). This segmentation allows the system to retain and utilize fixed phrases for service improvement and quality enhancement while removing only the variable personal information, thus preventing information loss of useful content while maintaining privacy protection
Solution Approach 2:
The patent applies different quality treatments to different parts of the document data collection: fixed phrases are preserved in their original form for analysis and reuse, while variable content is redacted or masked. This local differentiation of quality treatment allows useful non-private information to be retained and leveraged for service improvement while privacy-sensitive portions are appropriately protected
Data Source
AI summary
Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for redacting data from a document collection generated for a set of documents that include personal information. The redaction of the data is based in part on a comparison of the document collection to a set of a personal documents of users for which the users have provided explicit approval to use in the processing of the document collection.


