Data Risk Assessment Using Sample and Population Re-Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for determining data risk rely on a single attack type assumption, leading to inaccurate data risk estimation due to the relationship between the sample dataset and the prevalent population, and do not consider multiple attack types.

Innovation Solution

A method that determines data risk by consolidating re-identification probabilities of each record in a sample dataset and statistical information of a population, using weighted sums and various approaches like entropy, Bayesian, and hypothesis testing to balance and improve accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If data risk is determined based on a single attack type assumption (either sample dataset or population), then the determination process is simple, but the accuracy of data risk estimation deteriorates

Engineering Contradiction:
Improvedetermination process complexityVSAvoiddata risk estimation accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent combines multiple attack type assumptions (sample dataset-based and population-based) into a unified data risk determination framework. By merging the results from different attack scenarios through weighted aggregation, the system achieves more accurate and comprehensive data risk estimation while maintaining manageable process complexity.

Inventive Principle:
Principle #5Merging (Combining)

2Ease of operation

If data risk determination considers only sample dataset information, then the process is straightforward, but the reliability deteriorates due to potential mismatch with prevalent population

Engineering Contradiction:
Improvedetermination process easeVSAvoiddata risk estimation reliability
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent introduces population-level statistical information as an intermediary component that bridges the gap between sample dataset analysis and real-world data risk assessment. This intermediary population information compensates for potential mismatches between the sample dataset and prevalent population, thereby improving the reliability of data risk estimation without significantly complicating the determination process.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If multiple attack types are considered in data risk determination, then the accuracy improves, but the device complexity increases

Engineering Contradiction:
Improvedata risk estimation accuracyVSAvoiddetermination system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent manages the complexity of considering multiple attack types by parameterizing the determination process. It uses configurable weights for different attack scenarios and employs standardized calculation methodologies for each attack type, allowing the system to accurately assess multiple attack vectors while maintaining controlled complexity through parameter-driven flexibility.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP4600857A1Determination of data risk
Publication Date: 2025.08.13 KONINKLIJKE PHILIPS NV
  • EP4600857A1 patent drawingFigure 1A~1B
  • EP4600857A1 patent drawingFigure 2
  • EP4600857A1 patent drawingFigure 3

AI summary

Example embodiments of the present disclosure relate to a method for determining a data risk, an electronic device, and a computer-readable medium. In the solution, a first data risk value is determined based on a re-identification probability of each record in the sample dataset, and a second data risk value is determined based on statistical information of a population which includes the sample dataset. In addition, a data risk may be determined based on the first data risk value and the second data risk value. Embodiments of the present disclosure consider both a re-identification probability of each record in the sample dataset and statistical information of a population for determining the data risk. Therefore, the determined data risk will be more accurate, thereby an accurate re-identification risk may be obtained accordingly.