PII Redaction via Multi-Source Entity Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current document redaction technologies often resort to generic category-based methods, which can inadvertently redact information unrelated to the intended individual, failing to effectively protect sensitive information while maintaining document accessibility.
Innovation Solution
A system and method that utilize advanced techniques such as entity analysis and multiple data source retrieval to identify and redact personally identifiable information (PII) specifically associated with an individual, thereby avoiding generic category-based redaction and ensuring only the intended PII is removed, using tools like IBM Identity Insight, NetOwl, and Rosette Entity Extractor.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If generic category-based redaction methods are used, then redaction coverage is broad, but precision of identifying the correct individual's PII deteriorates
Solution Approach 1:
The system performs preliminary actions by retrieving known PII about the individual from multiple data sources before conducting the redaction process. This preliminary retrieval of identifying information (names, addresses, social security numbers, etc.) enables the system to precisely identify which PII belongs to the specific individual whose document is being accessed, thereby resolving the contradiction between identification precision and system complexity.
2Reliability
If generic category-based redaction is applied, then all PII in a category is removed, but unrelated information is inadvertently redacted
Solution Approach 1:
The system applies local quality by making different parts of the document have different redaction treatments based on their association with the individual. Instead of applying uniform category-based redaction to all PII, the system selectively redacts only those PII items that match the retrieved identifying information about the specific individual, while preserving other non-sensitive or unrelated information. This resolves the contradiction between redaction reliability and information loss.
3Measurement precision
If manual redaction review is performed, then accuracy of PII identification is improved, but processing time increases
Solution Approach 1:
The system implements self-service by automatically retrieving known PII from multiple data sources and autonomously identifying which PII in the document belongs to the individual, without requiring manual review. The system serves itself by using the retrieved identifying information to automatically determine redaction targets, thereby maintaining high accuracy while avoiding the time consumption of manual processes.
Data Source
AI summary
Personal information is retrieved from at least one data source and personal information associated with a first individual is identified. A document is generated that is a version of a first document, wherein the personal information associated with the first individual cannot be discerned.


