PII Redaction via Multi-Source Entity Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current document redaction technologies often resort to generic category-based methods, which can inadvertently redact information unrelated to the intended individual, failing to effectively protect sensitive information while maintaining document accessibility.

Innovation Solution

A system and method that utilize advanced techniques such as entity analysis and multiple data source retrieval to identify and redact personally identifiable information (PII) specifically associated with an individual, thereby avoiding generic category-based redaction and ensuring only the intended PII is removed, using tools like IBM Identity Insight, NetOwl, and Rosette Entity Extractor.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If generic category-based redaction methods are used, then redaction coverage is broad, but precision of identifying the correct individual's PII deteriorates

Engineering Contradiction:
Improveprecision of PII identificationVSAvoidcomplexity of redaction system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by retrieving known PII about the individual from multiple data sources before conducting the redaction process. This preliminary retrieval of identifying information (names, addresses, social security numbers, etc.) enables the system to precisely identify which PII belongs to the specific individual whose document is being accessed, thereby resolving the contradiction between identification precision and system complexity.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If generic category-based redaction is applied, then all PII in a category is removed, but unrelated information is inadvertently redacted

Engineering Contradiction:
Improveaccuracy of PII redactionVSAvoidloss of non-sensitive information
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The system applies local quality by making different parts of the document have different redaction treatments based on their association with the individual. Instead of applying uniform category-based redaction to all PII, the system selectively redacts only those PII items that match the retrieved identifying information about the specific individual, while preserving other non-sensitive or unrelated information. This resolves the contradiction between redaction reliability and information loss.

Inventive Principle:
Principle #3Local quality

3Measurement precision

If manual redaction review is performed, then accuracy of PII identification is improved, but processing time increases

Engineering Contradiction:
Improveaccuracy of PII identificationVSAvoiddocument processing speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system implements self-service by automatically retrieving known PII from multiple data sources and autonomously identifying which PII in the document belongs to the individual, without requiring manual review. The system serves itself by using the retrieved identifying information to automatically determine redaction targets, thereby maintaining high accuracy while avoiding the time consumption of manual processes.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS9904798B2Focused personal identifying information redaction
Publication Date: 2018.02.27 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US9904798B2 patent drawing
  • US9904798B2 patent drawing
  • US9904798B2 patent drawing

AI summary

Personal information is retrieved from at least one data source and personal information associated with a first individual is identified. A document is generated that is a version of a first document, wherein the personal information associated with the first individual cannot be discerned.