Phenotype Embedding via Hierarchical Knowledge Graph

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for generating patient embeddings from Electronic Health Records (EHR) data do not effectively utilize medical knowledge expressed via phenotypical attributes, leading to irrelevant and insufficient representations, especially when cohort sizes are small, limiting the robustness and applicability of deep learning models.

Innovation Solution

The use of multi-task and transfer learning with sparse gating mechanisms and domain knowledge to generate pheno-embeddings through a hierarchical knowledge graph, expanding sparse data sets and incorporating medical ontologies to create specialized embeddings for phenotypes, allowing for scalable and robust patient representation learning.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional deep learning methods are used to generate patient embeddings from EHR data, then computational processing can be performed, but the embeddings lack relevance and robustness due to insufficient utilization of medical knowledge and small cohort sizes

Engineering Contradiction:
Improverobustness of patient embeddingsVSAvoidmedical knowledge expressed via phenotypical attributes
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent introduces a knowledge graph as an intermediary structure that encodes medical knowledge and phenotypical attributes. This knowledge graph serves as a mediator between the EHR data and the deep learning model, allowing the model to leverage structured medical knowledge even when cohort sizes are small. The knowledge graph captures relationships between medical concepts, phenotypes, and attributes, which are then integrated into the embedding generation process through graph neural networks.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent performs preliminary construction of the knowledge graph and phenotypic embedding models before applying them to specific patient cohorts. By pre-encoding medical knowledge, phenotypical attributes, and relationships in the knowledge graph structure, the system prepares the informational framework in advance. This preliminary action allows the model to leverage pre-encoded medical knowledge during inference, improving robustness without requiring large amounts of cohort-specific training data.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If deep learning models are trained on small cohorts, then computational resources are conserved, but the applicability and robustness of the models are reduced

Engineering Contradiction:
Improvecomputational efficiencyVSAvoidapplicability of patient representations
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal knowledge graph structure that can be applied across different cohorts, diseases, and medical domains. The knowledge graph encodes general medical knowledge and phenotypical relationships that are transferable across different applications. This universal structure allows the same model framework to be applied to various cohorts and medical tasks without requiring retraining on large datasets, thereby maintaining computational efficiency while improving adaptability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent moves from traditional flat EHR data representation to a multi-dimensional knowledge graph structure that incorporates hierarchical relationships, phenotypical attributes, and medical concept connections. By adding these dimensional layers (graph structure, phenotypic dimensions, hierarchical levels), the model gains richer contextual information from small cohorts, improving applicability without proportionally increasing computational requirements.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Device complexity

If simple data representations are used for medical concepts, then computational complexity is reduced, but latent relationships in medical data cannot be captured

Engineering Contradiction:
Improvecomplexity of data representation modelVSAvoidlatent relationships in medical concepts
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The patent implements a nested structure where the knowledge graph contains hierarchical nesting of medical concepts, phenotypes, and attributes. The knowledge graph itself is nested within the deep learning architecture, with graph neural networks nested within the overall embedding generation pipeline. This nested organization allows the model to capture complex latent relationships at multiple levels (from individual concepts to phenotypes to patient profiles) while maintaining a structured approach that manages computational complexity through hierarchical organization.

Inventive Principle:
Principle #7Nested doll (Nesting)

Solution Approach 2:

The patent creates a composite representation by combining multiple data sources and structural elements: EHR data, knowledge graph structure, phenotypical attributes, and graph neural network computations. This composite approach integrates heterogeneous information types into a unified embedding representation, capturing latent relationships that would be missed by simple representations while managing complexity through the structured integration of multiple components.

Inventive Principle:
Principle #40Composite materials

Data Source

PatentUS11681726B2System for generating specialized phenotypical embedding
Publication Date: 2023.06.20 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11681726B2 patent drawing
  • US11681726B2 patent drawing
  • US11681726B2 patent drawing

AI summary

Systems and methods that use multi-tasking and transfer learning with sparse gating mechanisms and domain knowledge to generate pheno-embeddings in a scalable manner that can improve the relevance of the patient embeddings from Electronic Health Records. A system, comprises at least one processor that executes the following computer executable components stored in memory: a structural pheno-embedding model that employs a hierarchical knowledge graph; a data augmentation component that expands on a sparse data set associated with the knowledge graph; and an embedding component that generates a specialized embedding for phenotypes using the structural pheno-embedding model and the augmented data set for a selected cohort.