Phenotype Embedding via Hierarchical Knowledge Graph
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for generating patient embeddings from Electronic Health Records (EHR) data do not effectively utilize medical knowledge expressed via phenotypical attributes, leading to irrelevant and insufficient representations, especially when cohort sizes are small, limiting the robustness and applicability of deep learning models.
Innovation Solution
The use of multi-task and transfer learning with sparse gating mechanisms and domain knowledge to generate pheno-embeddings through a hierarchical knowledge graph, expanding sparse data sets and incorporating medical ontologies to create specialized embeddings for phenotypes, allowing for scalable and robust patient representation learning.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional deep learning methods are used to generate patient embeddings from EHR data, then computational processing can be performed, but the embeddings lack relevance and robustness due to insufficient utilization of medical knowledge and small cohort sizes
Solution Approach 1:
The patent introduces a knowledge graph as an intermediary structure that encodes medical knowledge and phenotypical attributes. This knowledge graph serves as a mediator between the EHR data and the deep learning model, allowing the model to leverage structured medical knowledge even when cohort sizes are small. The knowledge graph captures relationships between medical concepts, phenotypes, and attributes, which are then integrated into the embedding generation process through graph neural networks.
Solution Approach 2:
The patent performs preliminary construction of the knowledge graph and phenotypic embedding models before applying them to specific patient cohorts. By pre-encoding medical knowledge, phenotypical attributes, and relationships in the knowledge graph structure, the system prepares the informational framework in advance. This preliminary action allows the model to leverage pre-encoded medical knowledge during inference, improving robustness without requiring large amounts of cohort-specific training data.
2Productivity
If deep learning models are trained on small cohorts, then computational resources are conserved, but the applicability and robustness of the models are reduced
Solution Approach 1:
The patent creates a universal knowledge graph structure that can be applied across different cohorts, diseases, and medical domains. The knowledge graph encodes general medical knowledge and phenotypical relationships that are transferable across different applications. This universal structure allows the same model framework to be applied to various cohorts and medical tasks without requiring retraining on large datasets, thereby maintaining computational efficiency while improving adaptability.
Solution Approach 2:
The patent moves from traditional flat EHR data representation to a multi-dimensional knowledge graph structure that incorporates hierarchical relationships, phenotypical attributes, and medical concept connections. By adding these dimensional layers (graph structure, phenotypic dimensions, hierarchical levels), the model gains richer contextual information from small cohorts, improving applicability without proportionally increasing computational requirements.
3Device complexity
If simple data representations are used for medical concepts, then computational complexity is reduced, but latent relationships in medical data cannot be captured
Solution Approach 1:
The patent implements a nested structure where the knowledge graph contains hierarchical nesting of medical concepts, phenotypes, and attributes. The knowledge graph itself is nested within the deep learning architecture, with graph neural networks nested within the overall embedding generation pipeline. This nested organization allows the model to capture complex latent relationships at multiple levels (from individual concepts to phenotypes to patient profiles) while maintaining a structured approach that manages computational complexity through hierarchical organization.
Solution Approach 2:
The patent creates a composite representation by combining multiple data sources and structural elements: EHR data, knowledge graph structure, phenotypical attributes, and graph neural network computations. This composite approach integrates heterogeneous information types into a unified embedding representation, capturing latent relationships that would be missed by simple representations while managing complexity through the structured integration of multiple components.
Data Source
AI summary
Systems and methods that use multi-tasking and transfer learning with sparse gating mechanisms and domain knowledge to generate pheno-embeddings in a scalable manner that can improve the relevance of the patient embeddings from Electronic Health Records. A system, comprises at least one processor that executes the following computer executable components stored in memory: a structural pheno-embedding model that employs a hierarchical knowledge graph; a data augmentation component that expands on a sparse data set associated with the knowledge graph; and an embedding component that generates a specialized embedding for phenotypes using the structural pheno-embedding model and the augmented data set for a selected cohort.


