De-identification Risk Assessment via Probability Modeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for de-identifying patient data in medical texts do not accurately assess the risk of re-identification, as they focus on recall rather than the inherent risk of re-identification and do not account for the transformation or retention of personally identifying information (PII), leading to inconsistent and often inaccurate evaluations of de-identification effectiveness.

Innovation Solution

A system and method that calculate the probability of re-identification by modeling risks from an adversary's perspective, accounting for both direct and quasi-identifiers, and incorporating uncertainty in estimates, using a combination of risk measurements from leaked and caught PII, with a focus on precise determination of acceptable re-identification probabilities through benchmark comparisons.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If de-identification tools mask direct identifiers and perturb quasi-identifiers, then patient identifying information is removed, but the risk of re-identification remains inaccurately assessed

Engineering Contradiction:
Improveaccuracy of re-identification risk assessmentVSAvoidprecision of re-identification probability measurement
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent changes the measurement parameters from simple recall metrics to comprehensive risk assessment that incorporates multiple factors including direct identifier leakage probability, quasi-identifier re-identification risk, and their interactions. This transforms the assessment from a single-dimensional recall measurement to a multi-parameter probability model that accurately reflects re-identification risk.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces an intermediary risk assessment layer that sits between the de-identification process and the final data release. This intermediary layer calculates probabilities of direct identifier leakage and quasi-identifier re-identification, then combines these to produce an overall risk assessment that guides whether data release is acceptable.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If evaluation focuses on recall metrics, then de-identification tool performance is measured, but inherent re-identification risk is not captured

Engineering Contradiction:
Improveefficiency of de-identification evaluationVSAvoidaccuracy of risk assessment
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent segments the risk assessment into distinct components: direct identifier leakage risk, quasi-identifier re-identification risk, and their combined effect. By evaluating each segment separately and then integrating them, the system maintains evaluation efficiency while capturing comprehensive risk information that simple recall metrics miss.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds a new dimension to the evaluation by incorporating probability calculations and risk thresholds beyond traditional recall metrics. This dimensional expansion transforms the assessment from measuring what was removed to measuring the actual risk remaining, providing a more reliable basis for data release decisions.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Object-affected harmful factors

If de-identification transforms PII, then patient privacy is protected, but evaluation of transformation effectiveness is inconsistent

Engineering Contradiction:
Improveprotection level against re-identificationVSAvoidprecision of de-identification effectiveness measurement
Core Design Contradiction:
Object-affected harmful factorsVSMeasurement precision

Solution Approach 1:

The patent implements feedback mechanisms where risk assessment results are used to adjust and improve the de-identification process. By continuously measuring the actual risk rather than assuming transformation effectiveness, the system provides feedback that guides parameter adjustments to achieve target risk levels.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent changes the measurement parameters from transformation completeness to risk outcome. Instead of measuring whether all PII was transformed, it measures the actual re-identification risk remaining, allowing for more precise and meaningful evaluation of de-identification effectiveness.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10395059B2System and method to reduce a risk of re-identification of text de-identification tools
Publication Date: 2019.08.27 PRIVACY ANALYTICS
  • US10395059B2 patent drawing
  • US10395059B2 patent drawing
  • US10395059B2 patent drawing

AI summary

A computer-implemented system and method to reduce re-identification risk of a data set. The method includes the steps of retrieving, via a database-facing communication channel, a data set from a database communicatively coupled to the processor, the data set selected to include patient medical records that meet a predetermined criteria; identifying, by a processor coupled to a memory, direct identifiers in the data set; identifying, by the processor, quasi-identifiers in the data set; calculating, by the processor, a first probability of re-identification from the direct identifiers; calculating, by the processor, a second probability of re-identification from the quasi-direct identifiers; perturbing, by the processor, the data set if one of the first probability or second probability exceeds a respective predetermined threshold, to produce a perturbed data set; and providing, via a user-facing communication channel, the perturbed data set to the requestor.