Journalist Risk Estimation via Equivalence Class Distribution

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for de-identifying personal data in databases are inadequate in assessing the risk of re-identification, particularly for quasi-identifiers that are not obvious, which can compromise privacy when data is shared with third parties.

Innovation Solution

A computer-implemented method and system that estimate the journalist risk of a dataset by determining the sample equivalence class distribution, equating it to the population equivalence class distribution, and calculating the probability that an equivalence class in the dataset came from the population, using combinatorial calculations and Bayes' theorem to assess the risk of re-identification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If direct identifiers (names and addresses) are removed from personal data, then ease of data sharing and research value are improved, but privacy protection deteriorates because quasi-identifiers can still enable re-identification

Engineering Contradiction:
Improvedata sharing efficiencyVSAvoidre-identification risk
Core Design Contradiction:
ProductivityVSObject-affected harmful factors

Solution Approach 1:

The patent replaces traditional manual or simple automated de-identification methods with a sophisticated computational system that uses equivalence class distribution estimation and probabilistic modeling to assess re-identification risk, enabling more accurate privacy protection while maintaining data utility

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the approach from binary de-identification (removed/not removed) to a continuous probabilistic assessment by calculating the probability that an equivalence class in the dataset came from the population, allowing for nuanced risk evaluation and threshold-based decision making

Inventive Principle:
Principle #35Parameter changes

2Loss of information

If quasi-identifiers are retained in the dataset, then data utility and research value are improved, but privacy protection deteriorates due to increased risk of re-identification through combination of demographic attributes

Engineering Contradiction:
Improvedata utilityVSAvoidre-identification risk
Core Design Contradiction:
Loss of informationVSObject-affected harmful factors

Solution Approach 1:

The patent implements a feedback mechanism where the system calculates re-identification risk probabilities and uses this information to determine whether de-identification thresholds are met, allowing data custodians to adjust the level of de-identification based on the calculated risk levels and organizational risk tolerance

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent performs preliminary risk assessment by estimating population equivalence class distribution from the sample dataset before data sharing occurs, enabling data custodians to evaluate and mitigate re-identification risks in advance rather than after a breach occurs

Inventive Principle:
Principle #10Preliminary action

3Reliability

If comprehensive de-identification assessment is performed including equivalence class distribution estimation, then privacy protection is improved, but computational complexity and processing time increase

Engineering Contradiction:
Improveprivacy protection accuracyVSAvoidrisk assessment system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the complex risk assessment problem into manageable components: (1) determining sample equivalence class distribution, (2) equating population EC distribution to sample EC distribution, (3) calculating probability that an EC came from population, and (4) calculating journalist risk measurement, making the system implementable and maintainable

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11664098B2Determining journalist risk of a dataset using population equivalence class distribution estimation
Publication Date: 2023.05.30 PRIVACY ANALYTICS
  • US11664098B2 patent drawing
  • US11664098B2 patent drawing
  • US11664098B2 patent drawing

AI summary

Methods and systems to de-identify a longitudinal dataset of personal records based on journalistic risk computed from a sample set of the personal records, including determining a similarity distribution of the sample set based on quasi-identifiers of the respective personal records, converting the similarity distribution of the sample set to an equivalence class distribution, and computing journalistic risk based on the equivalence distribution. In an embodiment, multiple similarity measures are determined for a personal record based on comparisons with multiple combinations of other personal records of the sample set, and an average of the multiple similarity measures is rounded. In an embodiment, similarity measures are determined for a subset of the sample set and, for each similarity measure, the number of records having the similarity measure is projected to the subset of personal records. Journalistic risk may be computed for multiple types of attacks.