Document Data Redaction via Segmentation and Mediator Principles

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The challenge lies in analyzing and improving document data collections while maintaining user privacy, as human reviewers are prohibited from accessing private information, complicating the process of tuning algorithms and evaluating service processes that utilize these collections.

Innovation Solution

A method involving a data processing apparatus that receives a document data collection with fixed phrases not posing a personal information exposure risk, extracts candidate phrases from user-donated documents with explicit permission, generates a redacted collection by obfuscating or removing non-matching phrases, and allows human reviewers to examine only the matched phrases, thereby ensuring privacy protection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If access to document data collection is precluded to protect user privacy, then user privacy is protected, but the ability to analyze and improve the quality of the document data collection deteriorates

Engineering Contradiction:
Improveuser privacy protectionVSAvoidability to analyze and improve document data collection
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent segments the document data collection into two distinct parts: fixed phrases (reusable, non-private content patterns) and variable content (user-specific private information). By separating these components, the system enables human reviewers to access and analyze the fixed phrases portion while automatically redacting or masking the variable content, thus resolving the contradiction between privacy protection and analytical accessibility

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary automated processing system that acts as a mediator between the private document data collection and human reviewers. This intermediary automatically identifies, extracts, and redacts personal information from documents, generating a sanitized version that human reviewers can safely access and analyze without exposure to private data, thereby enabling analysis while maintaining privacy protection

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If all personal information is removed from document data collection, then privacy protection is ensured, but the quality and usefulness of the data collection for service improvement deteriorates

Engineering Contradiction:
Improveprivacy protectionVSAvoiduseful information in document data collection
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent segments information into fixed phrases (non-private, reusable patterns) and variable content (user-specific data). This segmentation allows the system to retain and utilize fixed phrases for service improvement and quality enhancement while removing only the variable personal information, thus preventing information loss of useful content while maintaining privacy protection

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different quality treatments to different parts of the document data collection: fixed phrases are preserved in their original form for analysis and reuse, while variable content is redacted or masked. This local differentiation of quality treatment allows useful non-private information to be retained and leveraged for service improvement while privacy-sensitive portions are appropriately protected

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS9734148B2Information redaction from document data
Publication Date: 2017.08.15 GOOGLE LLC
  • US9734148B2 patent drawing
  • US9734148B2 patent drawing
  • US9734148B2 patent drawing

AI summary

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for redacting data from a document collection generated for a set of documents that include personal information. The redaction of the data is based in part on a comparison of the document collection to a set of a personal documents of users for which the users have provided explicit approval to use in the processing of the document collection.