Dataset Augmentation via De-identified Data Authorization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Researchers face challenges in creating larger and richer datasets for studies due to difficulties in finding participants willing to contribute their personal data, and existing methods struggle to effectively incorporate de-identified data from other sources that are relevant to the research purpose.

Innovation Solution

A computer system determines regions of interestingness within authorized data records and identifies relevant de-identified data from other sources, requesting authorization from corresponding entities to augment the dataset, using a multidimensional model to analyze and de-identify data while ensuring compliance with privacy regulations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If researchers collect personal data from entities with consent, then data quality and relevance for specific research purposes are improved, but the quantity of available data is limited due to difficulty in finding enough participants

Engineering Contradiction:
Improvequantity of dataVSAvoidease of data collection
Core Design Contradiction:
Quantity of substanceVSEase of manufacture

Solution Approach 1:

The patent combines two previously separate data sources into a unified dataset: (1) personal data collected with specific consent for research purposes, and (2) de-identified data from other sources. This merging allows researchers to achieve larger dataset quantities while maintaining data quality through the complementary nature of these data sources.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces de-identified data as an intermediary element that bridges the gap between limited personally-identifiable data and the need for larger datasets. By using de-identified data from recommendation systems and other sources, researchers can expand data quantity without directly contacting additional participants.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If researchers use de-identified data from other sources, then the quantity and diversity of data are improved, but data relevance and quality for specific research objectives may deteriorate

Engineering Contradiction:
Improvequantity of dataVSAvoiddata relevance
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent applies local quality by treating different data sources differently based on their characteristics. Personally-identifiable data receives full contextual processing and validation, while de-identified data undergoes specific relevance filtering and matching processes. This differentiated approach ensures each data type is processed appropriately to maintain overall data relevance.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent performs preliminary actions by pre-filtering and pre-matching de-identified data against research objectives before integration. The system identifies and selects only those de-identified records that are relevant to the specific research study, ensuring data relevance is established before the data is combined with personally-identifiable data.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If researchers augment datasets with de-identified data, then data utility for research studies is improved, but complexity of data management and authorization processes increases

Engineering Contradiction:
Improvedata utilityVSAvoidcomplexity of data management
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent creates a universal data management framework that handles multiple data types (personally-identifiable and de-identified) through a single integrated process. The system performs multiple functions including authorization management, relevance filtering, data matching, and dataset integration within one unified architecture, reducing the need for separate complex systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent implements self-service mechanisms where the system automatically manages the complexity of integrating de-identified data. The framework autonomously handles authorization requests, relevance assessment, and data matching without requiring manual intervention, thereby improving data utility while containing management complexity through automation.

Inventive Principle:
Principle #25Self-service

4Reliability

If researchers request authorization from entities for de-identified data, then data privacy and compliance are improved, but time and resources required for data collection increase

Engineering Contradiction:
Improvedata privacy complianceVSAvoidtime for authorization
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary authorization actions by obtaining consent during the initial data collection phase. When users provide personally-identifiable data for research, they simultaneously grant permission for their de-identified data to be used in future studies. This preliminary authorization eliminates the need for repeated consent requests, maintaining privacy compliance while reducing time delays.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent establishes continuity of useful action by creating an ongoing authorization framework where initial consent enables continuous use of de-identified data across multiple research studies. This continuous authorization model maintains privacy and compliance requirements while eliminating repetitive authorization processes that would consume time and resources.

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS10892042B2Augmenting datasets using de-identified data and selected authorized records
Publication Date: 2021.01.12 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10892042B2 patent drawing
  • US10892042B2 patent drawing
  • US10892042B2 patent drawing

AI summary

A computer system utilizes a dataset to support a research study. One or more regions of interestingness are determined within a model of a first set of data records that are authorized for the research study by associated entities. A second set of data records is represented within the model, wherein the second set of data records are relevant for supporting objectives of the research study after de-identification. Records from the second dataset that are particularly useful for supporting objectives of the research study are identified, and authorization is requested from the corresponding entities of the identified data records from the second set of data records. After receiving authorization, those records are included with the first set to generate a resulting dataset. Embodiments of the present invention further include a method and program product for processing requests for health information in substantially the same manner described above.