Geography Bricks for Healthcare Data De-identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies face challenges in effectively de-identifying sensitive information, particularly in the healthcare industry, due to stringent regulations that require protecting patient confidentiality and compliance with data protection laws, such as HIPAA and EU data protection directives, which restrict the use of personally identifiable data.
Innovation Solution
A modified N-Tree approach is employed to break down geographic regions recursively into smaller segments, ensuring that each segment contains a minimum threshold of individuals, allowing for the aggregation of data without revealing individual identities, using polygon-shaped regions like rectangles, trapezoids, and hexagons, and updating these regions to account for population migration.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data is aggregated at a coarse geographic level (e.g., state level), then data protection is improved, but data utility and analytical precision deteriorate
Solution Approach 1:
The patent applies segmentation by dividing the geographic area into multiple hierarchical levels (state, county, census tract, block group, and custom micro-geographic zones). This allows data to be aggregated at the finest granular level possible while still meeting the minimum threshold requirement, thereby maximizing data utility while maintaining protection. The system creates custom geographic boundaries that can be dynamically adjusted to achieve optimal balance between precision and protection.
2Measurement precision
If geographic regions are divided into smaller segments to increase data precision, then data utility improves, but the risk of re-identification increases
Solution Approach 1:
The patent changes the parameter of geographic granularity dynamically based on the data characteristics and regulatory requirements. The system calculates the appropriate minimum threshold and divides the geographic area into segments that meet this threshold, creating custom micro-geographic zones that are smaller than traditional boundaries but still provide sufficient aggregation to prevent re-identification. This parameter adjustment allows optimization of both data utility and protection.
Solution Approach 2:
The patent applies local quality by creating custom micro-geographic zones with varying boundaries based on local population density and data characteristics. Instead of using uniform geographic boundaries, the system adjusts the granularity locally to match the specific needs of different regions, allowing finer segmentation in high-density areas and coarser segmentation in low-density areas, thereby optimizing the balance between data utility and re-identification risk.
3Ease of operation
If traditional geographic boundaries (ZIP codes, counties) are used for aggregation, then ease of operation is improved, but the ability to meet specific business rule criteria deteriorates
Solution Approach 1:
The patent applies dynamics by creating a flexible, dynamic geographic segmentation system that can adapt to different business rules and data characteristics. Instead of being locked into traditional static boundaries, the system dynamically generates custom micro-geographic zones based on real-time calculations of population density, minimum thresholds, and regulatory requirements. This dynamic approach allows the system to meet specific business rule criteria while maintaining ease of operation through automated processing.
Data Source
AI summary
Techniques of the described subject matter employ a break down algorithm in which a population of individuals is broken down into segments that have a greater number of individuals than a threshold minimum. Information on aggregated individuals may then be used to accomplish a variety of tasks, such as consumer purchasing preferences, market data analysis, sales force allocation, etc., without revealing the specific identity of any individuals or permitting others to determine, from the data, the identity of any individuals.


