Gene Target Prioritization via Knowledge Graph Embeddings

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data warehousing systems fail to effectively utilize large volumes of human genetic and disease-related data due to lack of proper data quality screening and contextual relationship information, limiting their ability to efficiently identify gene targets for diseases.

Innovation Solution

A system that extracts datasets from multiple databases and stores them in a data lake using graph-based datasets, generating a knowledge graph to represent links between genes and diseases, enabling efficient identification of target genes associated with diseases through machine learning analytics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If data are aggregated into large data warehouses without proper data quality screening and contextual relationship information, then the volume of stored information increases, but the usefulness and analytical value of the data decreases

Engineering Contradiction:
Improvevolume of stored dataVSAvoidloss of contextual information
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The patent segments the monolithic data warehouse into multiple specialized knowledge graphs, each representing specific domains (genetic mutations, gene expressions, drug interactions, diseases). This segmentation allows for targeted data quality screening and contextual relationship modeling in each domain while maintaining the overall volume of stored information.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces knowledge graphs as intermediary structures between raw data aggregation and analytical applications. These knowledge graphs serve as mediators that enforce data quality standards, add contextual relationships, and structure information in a way that preserves analytical value while maintaining data volume.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If traditional string matching mechanisms are used for searching enterprise data, then the search implementation is simple, but the ability to provide complete and contextual queried data is limited

Engineering Contradiction:
Improveease of data searchVSAvoidloss of data completeness
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent changes the search parameter from simple string matching to multi-relational graph queries that traverse knowledge graphs. This allows searches to incorporate contextual relationships and data quality metrics while maintaining ease of operation through standardized query interfaces.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent adds a new dimension to data search by transitioning from one-dimensional string matching to multi-dimensional graph traversal across knowledge graphs. This enables queries to consider multiple relationships and contextual factors simultaneously while preserving ease of use through standardized query mechanisms.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Quantity of substance

If most stored data is not easily searchable or available for machine learning analytics, then the data storage capacity is maximized, but the productivity of data analysis and drug discovery decreases

Engineering Contradiction:
Improvedata storage capacityVSAvoidproductivity of drug discovery
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The patent performs preliminary actions by pre-processing and structuring data into knowledge graphs with defined schemas, relationships, and quality metrics before analytical applications are executed. This preliminary structuring makes data immediately accessible and ready for machine learning analytics while maintaining full storage capacity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes the accessibility parameter of stored data by transforming unstructured or semi-structured data into standardized knowledge graph formats with defined query interfaces. This transformation maintains data volume while dramatically improving searchability and readiness for machine learning analytics.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20210174906A1Systems And Methods For Prioritizing The Selection Of Targeted Genes Associated With Diseases For Drug Discovery Based On Human Data
Publication Date: 2021.06.10 ACCENTURE GLOBAL SOLUTIONS LTD
  • US20210174906A1 patent drawing
  • US20210174906A1 patent drawing
  • US20210174906A1 patent drawing

AI summary

Systems and methods enable the discovery of new relationships between diseases and genes by prioritizing the selection of gene targets for a disease using an embedding space generated from a knowledge graph by mapping datasets collected from various data sources using a graph schema, modeling disease and gene associations with link weightings, analyzing the data with several machine learning models, and scoring predictions.