Augmenting Datasets with De-identified Records

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Researchers face challenges in creating larger and richer datasets for studies due to difficulties in obtaining sufficient personal data from willing participants, while other entities have made their data available for de-identification, which can enhance research but remains underutilized.

Innovation Solution

A computer system generates a dataset by identifying regions of interest within a first dataset authorized for a research study and augments it with de-identified data from a second dataset, ensuring compliance with de-identification requirements to create a larger, richer dataset for analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If researchers collect personal data from willing participants, then the dataset size is limited by participant availability, but data privacy and consent requirements restrict the ability to include more data

Engineering Contradiction:
Improvedataset sizeVSAvoiddata availability flexibility
Core Design Contradiction:
Quantity of substanceVSAdaptability or versatility

Solution Approach 1:

The patent segments the dataset into two distinct components: (1) personally identifiable information (PII) data from participants who provided explicit consent, and (2) de-identified data from broader sources. This segmentation allows the system to combine datasets with different privacy characteristics, thereby increasing overall dataset size while respecting participant consent boundaries.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces de-identification as an intermediary process that transforms raw personal data into anonymized form. This intermediary step enables the inclusion of additional data sources that would otherwise be restricted by privacy concerns, effectively bridging the gap between data availability and privacy protection requirements.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If researchers use de-identified data from other entities, then dataset richness increases, but ensuring compliance with de-identification requirements becomes more complex

Engineering Contradiction:
Improvedata volumeVSAvoidcompliance complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent applies de-identification procedures in advance, before the data is combined with the primary consented dataset. By performing de-identification as a preliminary action on source data, the system ensures compliance requirements are met before integration, simplifying the overall compliance process rather than attempting to manage complexity during data combination.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent works with copies of data that have been de-identified, rather than original personal data. This copying approach allows researchers to use abundant de-identified data from multiple sources while maintaining privacy compliance, as the copies no longer contain personally identifiable information that would trigger complex consent and authorization requirements.

Inventive Principle:
Principle #26Copying

3Reliability

If researchers restrict data to only consented participants, then privacy compliance is maintained, but the dataset lacks sufficient volume for robust research conclusions

Engineering Contradiction:
Improveprivacy complianceVSAvoiddata volume
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent merges two distinct data sources with different privacy characteristics: (1) PII data from participants with explicit consent for specific research purposes, and (2) de-identified data from broader populations. This merging strategy allows the system to achieve sufficient data volume for robust research conclusions while maintaining privacy compliance through the complementary nature of the two data sources.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a composite dataset that combines personally identifiable information data and de-identified data in a structured manner. This composite approach allows the system to leverage the strengths of both data types—the contextual richness of PII data from consented participants and the volume of de-identified data—while managing privacy risks through clear differentiation and appropriate use restrictions.

Inventive Principle:
Principle #40Composite materials

Data Source

PatentUS11093646B2Augmenting datasets with selected de-identified data records
Publication Date: 2021.08.17 MERATIVE US LP
  • US11093646B2 patent drawing
  • US11093646B2 patent drawing
  • US11093646B2 patent drawing

AI summary

A computer system utilizes a dataset to support a research study. Regions of interestingness are determined within a model of data records of a first dataset that are authorized for the research study by associated entities. Data records from a second dataset are represented within the model, wherein the data records from the second dataset are relevant for supporting objectives of the research study. Data records from the second dataset that fail to satisfy de-identification requirements are removed. A resulting dataset is generated that including the first dataset records within a selected region of interestingness and selected records of the second dataset within the same region. The second dataset records within the resulting dataset are de-identified based on the de-identification requirements. Embodiments of the present invention further include a method and program product for utilizing a dataset to support a research study in substantially the same manner described above.