Graph-Based Metadata Structuring for Medical Data Categorization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Manual curation of medical imaging data sets is time-consuming, expensive, and non-scalable, making it challenging to structure medical data effectively for machine learning applications.

Innovation Solution

An improved method of data structuring uses simple text-based properties in metadata attributes to categorize medical data, determining pairwise similarity between medical records using string metrics like edit distance, and identifying groups of similar records to facilitate scalable data categorization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual curation of medical imaging data sets is performed, then data can be properly labeled and structured, but the process becomes time consuming and expensive

Engineering Contradiction:
Improvedata structuring qualityVSAvoiddata curation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical data curation with an automated computational system that uses graph-based algorithms and machine learning to structure medical imaging data. The system automatically extracts metadata, builds similarity graphs, and clusters records without human intervention, eliminating the time-consuming manual process while maintaining data quality.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables data to structure itself by automatically extracting metadata from medical imaging records, computing pairwise similarities, and organizing records into coherent groups without requiring manual curation. The algorithmic process self-manages the entire data structuring workflow.

Inventive Principle:
Principle #25Self-service

2Manufacturing precision

If manual data cleaning and structuring is performed on large-scale medical data, then data quality improves, but the process becomes prohibitively slow and expensive

Engineering Contradiction:
Improvedata cleaning precisionVSAvoiddata processing throughput
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent replaces slow manual data cleaning with high-speed automated algorithms that process medical imaging records computationally. The graph-based similarity computation and clustering algorithms execute rapidly on large datasets, achieving both high precision in data cleaning and high throughput in processing capacity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system transforms the data cleaning problem from a sequential manual process into a parallel computational problem by representing data relationships as graphs with nodes and edges. This dimensional transformation enables simultaneous processing of multiple record comparisons, dramatically increasing throughput while maintaining precision.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Measurement precision

If machine learning models use large numbers of weak input features from pixel determination, then model evaluation capability improves, but large curated data sets are required for training

Engineering Contradiction:
Improvemodel evaluation accuracyVSAvoidtraining data volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent extracts meaningful features directly from metadata attributes of medical imaging records, pulling out clinically relevant information such as patient demographics, imaging parameters, and diagnostic findings. This extraction of high-quality features from structured data reduces the need for large training datasets compared to models relying on raw pixel features.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system transforms the feature representation from raw pixel values to structured metadata parameters through automated extraction and graph-based clustering. This parameter transformation creates more informative features that require less training data to achieve the same model evaluation accuracy.

Inventive Principle:
Principle #35Parameter changes

4Stability of the object's composition

If semi-structured medical data with site-specific naming schemes is processed manually, then data can be cleaned and standardized, but the process becomes error prone and non-scalable

Engineering Contradiction:
Improvedata standardizationVSAvoiddata processing scalability
Core Design Contradiction:
Stability of the object's compositionVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal graph-based framework that handles diverse medical imaging data from multiple clinical sites with different naming schemes and formats. The system uniformly processes various data types and structures through the same algorithmic approach, achieving both standardization and scalability across heterogeneous datasets.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system replaces error-prone manual standardization with automated computational processes that consistently apply naming conventions and data formatting rules. The algorithmic approach eliminates human errors while scaling to handle large volumes of semi-structured data from multiple sources.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20250046065A1Graph based metadata structuring algorithm to enable machine learning
Publication Date: 2025.02.06 TEMPUS AI INC
  • US20250046065A1 patent drawing
  • US20250046065A1 patent drawing
  • US20250046065A1 patent drawing

AI summary

In the disclosed systems and methods for categorizing medical data, a computer system obtains, in electronic form, a plurality of medical records. Each medical record includes corresponding medical data from a respective medical evaluation and corresponding metadata comprising a plurality of attributes about the respective medical evaluation. Each respective attribute comprises a corresponding string of text. The computer system determines, for each respective pair of medical records consisting of a first medical record and a second medical record, a corresponding pairwise similarity between, for each respective attribute in a set of attributes, the corresponding string of text for the first medical record and the corresponding string of text for the second medical record. The computer system identifies a first subset of the plurality of medical records. Each respective medical record in the first subset is connected to each other through pairwise similarities that each satisfies a similarity threshold.