Patient Data De-identification via Multi-Dimensional Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods fail to aggregate patient medical information in compliance with HIPAA regulations and do not effectively reduce the risk of re-identification, particularly in small geographic areas or for healthcare research and marketing purposes.

Innovation Solution

A method and system that de-identifies patient data by aggregating based on geographic proximity and medical information similarity, using clustering and coding processes to create de-identified geographic areas and medical characteristic hierarchies, while ensuring compliance with HIPAA regulations and reducing re-identification risk.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If patient data is aggregated using HIPAA safe harbor regulations, then compliance with HIPAA standards is achieved, but re-identification risk is not sufficiently reduced and useful information is lost

Engineering Contradiction:
ImproveHIPAA complianceVSAvoiduseful information
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent changes the aggregation parameters from fixed geographic boundaries (ZIP codes) to flexible clustering based on multiple dimensions including geographic proximity, medical information similarity, and population density. This allows dynamic adjustment of aggregation levels to maintain HIPAA compliance while preserving more useful information for research and marketing purposes.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent segments patient data into multiple hierarchical levels of aggregation, allowing different degrees of detail for different purposes. Data can be viewed at coarse-grained levels for compliance and fine-grained levels for analysis, resolving the contradiction between compliance and information utility.

Inventive Principle:
Principle #1Segmentation

2Productivity

If patient data is aggregated in small geographic areas, then local healthcare research and marketing value is improved, but re-identification risk increases

Engineering Contradiction:
Improvelocal healthcare research valueVSAvoidre-identification risk
Core Design Contradiction:
ProductivityVSObject-affected harmful factors

Solution Approach 1:

The patent merges multiple small geographic areas into clusters that meet minimum population thresholds for HIPAA compliance. By combining geographically proximate areas with similar medical characteristics, the system maintains local relevance while achieving the statistical anonymity required to reduce re-identification risk.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces an intermediary clustering layer between individual patient data and aggregated statistics. This clustering mechanism acts as a mediator that preserves local healthcare value while enforcing privacy protections, allowing research and marketing activities without direct access to identifiable patient information.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If traditional ZIP code aggregation is used, then HIPAA compliance is achieved, but re-identification risk remains high due to granular geographic information

Engineering Contradiction:
ImproveHIPAA complianceVSAvoidre-identification risk
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent adds additional dimensions to the aggregation process beyond simple geographic boundaries. By incorporating medical information similarity, population density, and healthcare utilization patterns, the system creates multi-dimensional clusters that are less susceptible to re-identification while maintaining compliance with HIPAA standards.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS8447738B1Computer system and method for de-identification of patient and/or individual health and/or medical related information, such as patient micro-data
Publication Date: 2013.05.21 EXPRESS SCRIPTS STRATEGIC DEVELOPMENT INC
  • US8447738B1 patent drawing
  • US8447738B1 patent drawing
  • US8447738B1 patent drawing

AI summary

A computer-implemented method de-identifies data collected for patients. IN at least one embodiment, the method comprises the sequential, non-sequential and/or sequence independent steps of providing information representative of at least one patient, at least one medical characteristic associated with at least one patient thereto, and a geographic area of the at least one patient, and providing at least one organizational structure for organizing medical characteristics. The method also includes associating the at least one organizational structure with at least one geographical area and at least one medical characteristic, and aggregating, in the at least one organizational structure, said information by medical characteristic and the at least one geographic area therein. Various alternative embodiments are additionally disclosed.