Document Sanitization via Hypernym Replacement and Confidence Scoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current redaction methods either remove too much information or not enough, and fail to appropriately divide labor between humans and machines, leading to inadequate protection of sensitive information in documents.
Innovation Solution
A system that determines privacy risk for terms in a document by calculating confidence measures and providing a user interface to indicate information utility and privacy loss or gain, allowing for the replacement of sensitive terms with less sensitive alternatives while preserving document utility, using data mining and linguistic parsing technologies.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If sensitive information is redacted by blacking-out or obscuring, then privacy protection is improved, but information utility is lost and the document becomes meaningless
Solution Approach 1:
The system changes the parameter of redaction from complete obscuration to selective replacement. Instead of blacking out terms entirely, it replaces them with hypernym terms that maintain contextual meaning while reducing privacy risk. The confidence measure cs(t) serves as a parameter to determine which terms require replacement, enabling precise control over the balance between privacy protection and information utility preservation.
Solution Approach 2:
The system introduces hypernym terms as intermediaries between the original sensitive terms and complete redaction. These hypernym terms act as mediators that preserve document utility while reducing privacy risk. For example, replacing a specific person's name with a generic role term maintains the document's informational value while protecting the individual's privacy.
2Productivity
If automated redaction methods are used, then productivity is improved, but reliability deteriorates due to inability to discriminate among potential terms
Solution Approach 1:
The system implements feedback through the confidence measure cs(t), which quantifies the privacy risk of each term. This feedback mechanism allows the automated system to discriminate among potential terms for redaction based on calculated risk levels. The system can adjust its redaction strategy based on the confidence scores, improving both accuracy and reliability while maintaining high productivity.
3Reliability
If manual redaction methods are used, then reliability is improved through human judgment, but productivity deteriorates due to heavy burden on humans
Solution Approach 1:
The system applies partial automation by using automated methods for terms with high confidence scores (clear privacy risks) while leaving terms with low confidence scores for manual review. This partial action approach maintains high productivity for obvious cases while ensuring reliability for ambiguous cases, optimally dividing labor between machines and humans.
4Object-affected harmful factors
If too much information is redacted, then privacy protection is improved, but information utility is lost
Solution Approach 1:
The system applies local quality by treating different terms differently based on their individual confidence scores. Instead of uniformly redacting all potential sensitive terms, it selectively replaces only those terms with high privacy risk (high cs(t) values). This localized approach preserves information utility for terms that don't require redaction while protecting privacy for terms that do.
Data Source
AI summary
One embodiment provides a system for facilitating sanitizing a modified version of a document relative to one or more sensitive topics. During operation, the system determines a privacy risk for a term in the modified version relative to the sensitive topics, wherein the privacy risk measures the extent to which the sensitive topic(s) can be inferred based on the term. Next, the system determines an information utility and privacy loss or gain for the modified version, where the information utility reflects the extent to which the modified version has changed and the privacy loss or gain reflects the extent to which the modified version is reduced in sensitivity.


