Metadata-Driven Health Data Normalization and De-identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current healthcare systems face challenges in processing and analyzing large volumes of diverse health data due to structural and terminological inconsistencies, which can lead to biased insights and compromised patient care, and conventional de-identification techniques often remove too much information, limiting data utility.

Innovation Solution

A health data platform utilizing artificial intelligence, machine learning, and natural language processing to automate data normalization, combining syntactic and semantic normalization, and employing metadata-driven approaches to convert diverse data into a unified format, while ensuring data security and privacy through de-identification processes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional de-identification techniques are used to protect patient privacy, then data security and privacy are improved, but data utility is reduced due to excessive information removal

Engineering Contradiction:
Improvedata security and privacyVSAvoiddata utility
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent applies parameter changes by adjusting the de-identification process to remove only specific identifying parameters (names, addresses, phone numbers, dates) while preserving clinically valuable parameters. This selective parameter removal maintains data utility for research purposes while ensuring patient privacy protection, resolving the contradiction between security and information retention

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent segments the de-identification process into distinct stages: initial de-identification that removes direct identifiers, followed by normalization that standardizes remaining data structures, and finally semantic mapping that preserves meaningful clinical concepts. This segmented approach allows selective removal of only necessary identifying information while maintaining data utility

Inventive Principle:
Principle #1Segmentation

2Quantity of substance

If data from multiple health systems is aggregated to increase dataset diversity, then research insights are improved, but processing complexity increases due to structural and terminological inconsistencies

Engineering Contradiction:
Improvedataset diversity and volumeVSAvoidprocessing complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent implements a universal normalization layer that processes diverse data structures from multiple health systems through a common framework. This normalization component translates various source formats into a standardized intermediate representation, enabling consistent processing of heterogeneous data while maintaining the ability to handle different data types and structures through a single multi-functional system

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces an intermediary normalization layer between raw incoming data and the research analysis pipeline. This intermediary component performs semantic mapping and structural standardization, acting as a mediator that reconciles terminological and structural differences from various health systems before data enters the research workflow, thereby reducing processing complexity downstream

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If extensive data processing and normalization are performed to ensure data quality, then measurement precision is improved, but processing time increases

Engineering Contradiction:
Improvedata quality and accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary normalization and de-identification processing as data enters the system, before research analysis begins. By pre-processing data to establish consistent structures and remove identifiers upfront, the system eliminates the need for repeated processing during analysis, thereby maintaining high data quality while reducing overall processing time for research workflows

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240370404A1Systems and methods for metadata driven normalization
Publication Date: 2024.11.07 TRUVETA INC
  • US20240370404A1 patent drawing
  • US20240370404A1 patent drawing
  • US20240370404A1 patent drawing

AI summary

Techniques for metadata driven normalization are disclosed. In some examples, the disclosed technology includes receiving a plurality of configuration data structures, each configuration data structure comprising one or more mappings, each mapping comprising a reference to instructions for transforming data. Metadata associated with a received data set is identified and compared to metadata information associated with the configuration data structures. A configuration data structure is selected based on this comparison and referenced instructions are used to generate and then execute code for transforming the received data set.