Patient Data Subset Sizing with Anonymity Probability Curves

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The inferencing capability of machine learning models is limited by the scope of the training data, and updating these models in deployment environments is hindered by contractual restrictions and privacy issues that prevent sharing of patient data between vendor and client sites.

Innovation Solution

Techniques are developed to determine specific portions of patient data, such as subsets of pixels from medical images, that ensure a low probability of uniquely identifying the patient, allowing for data sharing while maintaining privacy, thereby enabling model updates and refinements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If patient data is shared between vendor and client sites for model retraining, then model performance can be continuously improved, but patient privacy is compromised and contractual restrictions are violated

Engineering Contradiction:
Improvemodel improvement rateVSAvoidprivacy risk
Core Design Contradiction:
ProductivityVSObject-affected harmful factors

Solution Approach 1:

The patent segments patient data into distinct feature components and selectively shares only de-identified portions with the vendor. This segmentation allows the client site to maintain privacy controls while enabling the vendor to access sufficient data for model retraining and performance improvement.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary data processing layer that de-identifies and transforms patient data before transmission to the vendor. This intermediary mechanism enables data sharing for model improvement while protecting patient privacy through automated de-identification processes.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Object-affected harmful factors

If complete patient data is retained at the client site for privacy reasons, then patient privacy is protected, but model retraining and performance improvement are hindered

Engineering Contradiction:
Improveprivacy protectionVSAvoidmodel update capability
Core Design Contradiction:
Object-affected harmful factorsVSProductivity

Solution Approach 1:

The patent segments patient data into identifiable and de-identified portions, retaining complete data at the client site for privacy protection while sharing segmented de-identified portions with the vendor to enable model retraining and performance improvement.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts only the necessary de-identified portions of patient data from the complete dataset, allowing the client site to retain full data for privacy reasons while providing extracted de-identified portions to the vendor for model updates.

Inventive Principle:
Principle #2Taking out (Extraction)

3Productivity

If data subsets are shared for model retraining, then model performance can be improved, but determining appropriate subset sizes that maintain privacy is complex

Engineering Contradiction:
Improvemodel retraining efficiencyVSAvoiddata subset selection complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent performs preliminary de-identification and data subset preparation at the client site before sharing with the vendor. This preliminary action simplifies the subsequent model retraining process by providing pre-processed, privacy-compliant data subsets without requiring complex real-time analysis at the vendor site.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent transforms data parameters by applying de-identification techniques and selecting specific feature subsets, converting complex patient data into simplified, privacy-compliant formats that are easier to share and process for model retraining.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12411988B2Privacy driven data subset sizing
Publication Date: 2025.09.09 GE PRECISION HEALTHCARE LLC
  • US12411988B2 patent drawing
  • US12411988B2 patent drawing
  • US12411988B2 patent drawing

AI summary

Techniques are described for maintaining patient privacy in association with obtaining patient data for machine learning applications. In an example, a method can comprise accessing, by a system comprising a processor, a training dataset associated with a machine learning model, the training dataset comprising of data samples respectively comprising unique characteristics of subjects. The method further comprises determining, by the system, characteristic curve information that correlates different data portions extracted from the data samples to respective probabilities of matching the different data portions to respective ones of the data samples from which they are extracted. The method further comprises controlling, by the system, collection of new data portions extracted from new data samples corresponding to the data samples based on the new data portions conforming to criteria that defines a target data portion, wherein the criteria comprise a probability of the respective probabilities that satisfies an anonymity criterion.