De-identifying Unstructured Healthcare Data with Case-Type Tags
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional healthcare data systems lack automated solutions for de-identifying protected health information (PHI) and personal identification information (PII) from unstructured data, leading to limitations in data accessibility and compliance with privacy regulations like HIPAA, as manual removal methods are inefficient and incomplete.
Innovation Solution
A dictionary-based system that uses tunable blacklists and whitelists to identify and remove PHI/PII from unstructured data, replacing removed information with case-type tags to maintain data coherence and enable secure merging of de-identified data sets across disparate sources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If PHI/PII is removed from unstructured data to comply with privacy regulations, then data privacy and security are improved, but data coherence and analytical utility deteriorate
Solution Approach 1:
The patent introduces case-type tags as intermediary elements that replace removed PHI/PII. These tags serve as mediators that preserve the structural and semantic coherence of the original data while ensuring no actual personal information is exposed. The tags maintain the positional and contextual relationships of the removed information without containing the sensitive data itself.
Solution Approach 2:
The system creates a copy of the data structure and relationships without copying the actual PHI/PII values. By replacing sensitive information with standardized case-type tags that replicate the structural role of the original data, the system preserves analytical utility while eliminating privacy risks.
2Reliability
If manual methods are used to remove PHI/PII from unstructured data, then data privacy is improved, but productivity and efficiency deteriorate
Solution Approach 1:
The system performs self-service de-identification by automatically identifying, flagging, and removing PHI/PII from unstructured data without requiring manual human intervention. The automated process analyzes text patterns, applies privacy rules, and executes redaction independently, dramatically improving productivity while maintaining consistent privacy protection.
Solution Approach 2:
The patent replaces the mechanical manual process of PHI removal with an automated computational system. Instead of human operators manually reviewing and redacting sensitive information, the system uses algorithmic processing to identify and remove PHI/PII, replacing labor-intensive manual work with efficient automated processing.
3Reliability
If PHI/PII is removed from healthcare data sets, then compliance with HIPAA regulations is improved, but the ability to match and merge data sets deteriorates
Solution Approach 1:
The system uses case-type tags as intermediary placeholders that preserve the structural information needed for data matching and merging. These tags maintain the positional and relational context of removed PHI/PII, enabling datasets to be linked and merged based on their de-identified structures without requiring exposure of actual personal identifiers.
Data Source
AI summary
A method and apparatus for identifying personally identifiable information (PII) and protected health information (PHI) within unstructured data, removing the PII and PHI from the unstructured data, and replacing the removed information with case-type tags that allows the user to understand what information was removed and to tune the level of information removal in future data sets.


