Gene Target Prioritization via Knowledge Graph Embeddings
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data warehousing systems fail to effectively utilize large volumes of human genetic and disease-related data due to lack of proper data quality screening and contextual relationship information, limiting their ability to efficiently identify gene targets for diseases.
Innovation Solution
A system that extracts datasets from multiple databases and stores them in a data lake using graph-based datasets, generating a knowledge graph to represent links between genes and diseases, enabling efficient identification of target genes associated with diseases through machine learning analytics.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data are aggregated into large data warehouses without proper data quality screening and contextual relationship information, then the volume of stored information increases, but the usefulness and analytical value of the data decreases
Solution Approach 1:
The patent segments the monolithic data warehouse into multiple specialized knowledge graphs, each representing specific domains (genetic mutations, gene expressions, drug interactions, diseases). This segmentation allows for targeted data quality screening and contextual relationship modeling in each domain while maintaining the overall volume of stored information.
Solution Approach 2:
The patent introduces knowledge graphs as intermediary structures between raw data aggregation and analytical applications. These knowledge graphs serve as mediators that enforce data quality standards, add contextual relationships, and structure information in a way that preserves analytical value while maintaining data volume.
2Ease of operation
If traditional string matching mechanisms are used for searching enterprise data, then the search implementation is simple, but the ability to provide complete and contextual queried data is limited
Solution Approach 1:
The patent changes the search parameter from simple string matching to multi-relational graph queries that traverse knowledge graphs. This allows searches to incorporate contextual relationships and data quality metrics while maintaining ease of operation through standardized query interfaces.
Solution Approach 2:
The patent adds a new dimension to data search by transitioning from one-dimensional string matching to multi-dimensional graph traversal across knowledge graphs. This enables queries to consider multiple relationships and contextual factors simultaneously while preserving ease of use through standardized query mechanisms.
3Quantity of substance
If most stored data is not easily searchable or available for machine learning analytics, then the data storage capacity is maximized, but the productivity of data analysis and drug discovery decreases
Solution Approach 1:
The patent performs preliminary actions by pre-processing and structuring data into knowledge graphs with defined schemas, relationships, and quality metrics before analytical applications are executed. This preliminary structuring makes data immediately accessible and ready for machine learning analytics while maintaining full storage capacity.
Solution Approach 2:
The patent changes the accessibility parameter of stored data by transforming unstructured or semi-structured data into standardized knowledge graph formats with defined query interfaces. This transformation maintains data volume while dramatically improving searchability and readiness for machine learning analytics.
Data Source
AI summary
Systems and methods enable the discovery of new relationships between diseases and genes by prioritizing the selection of gene targets for a disease using an embedding space generated from a knowledge graph by mapping datasets collected from various data sources using a graph schema, modeling disease and gene associations with link weightings, analyzing the data with several machine learning models, and scoring predictions.


