Anonymizing Electronic Documents via Equivalence Classes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current automated data extraction methods for Web email require human intervention for debugging and privacy protection, which is inefficient and violates user privacy due to restricted access to personal information.
Innovation Solution
Implementing a method to anonymize electronic documents by grouping structurally similar documents into equivalence classes using a tailored hash function, ensuring anonymity and productivity measurement for auditors while preserving user privacy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If human auditors directly access personal information for auditing, then auditing accuracy and debugging capability are improved, but user privacy is compromised and access must be restricted
Solution Approach 1:
The patent introduces an intermediary anonymization system that sits between the original electronic documents and the auditors. This system applies k-anonymity techniques to generalize personal information while preserving auditing capabilities, allowing auditors to work with anonymized data that protects user privacy yet maintains sufficient detail for accurate auditing and debugging
Solution Approach 2:
The patent creates anonymized copies of original electronic documents that can be freely accessed by auditors without compromising the original data. These copies contain generalized personal information that preserves document structure and content for auditing purposes while removing or masking sensitive identifiers, enabling unlimited auditor access without privacy risks
2Productivity
If automated data extraction methods are used, then processing efficiency is improved, but human intervention is still required for debugging and evaluation
Solution Approach 1:
The patent applies anonymization techniques in advance to create standardized, anonymized document copies before auditing begins. This preliminary processing automates the removal of sensitive information and generalization of personal data, eliminating the need for manual intervention in debugging and evaluation while maintaining full auditing capability
3Object-affected harmful factors
If k-anonymity is enforced over the entire audit lifetime, then user privacy is protected, but auditor productivity measurement becomes more complex
Solution Approach 1:
The patent segments the auditing process into distinct phases: anonymization processing, auditing execution, and productivity measurement. By separating these functions, the system can enforce k-anonymity throughout the audit lifetime while independently measuring auditor productivity through metrics such as anonymized documents processed, audit findings generated, and processing time, without the measurements becoming inextricably complex
Data Source
AI summary
Methods, systems, and computer-readable media for anonymizing electronic documents. In accordance with one or more embodiments, structurally-similar electronic documents can be identified among a group of electronic documents (e.g., e-mail messages, documents containing HTML formatting, etc.). A hash function can be specifically tailored to identify the similarly structured documents. The structurally-similar electronic documents can be grouped into a same equivalence class. Masked anonymized document samples can be generated from the structurally-similar electronic documents utilizing the same equivalence class, thereby ensuring that the anonymized document samples when viewed as a part of an audit remain anonymous. An online process is provided to guarantee k-anonymity of the users over the entire lifetime of the auditing process. An auditor's productivity can be measured based on the amount of content revealed to the auditor within the samples he is assigned. The auditor's productivity is maximized while ensuring anonymization over the lifetime of the audit.


