Knowledge Graph Embeddings for Multi-Property Compound Discovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for new drug discovery using machine learning are resource-intensive and inefficient due to the need for multiple output layers and loss functions when training variational autoencoders (VAEs) with multiple properties, leading to improper training and incorrect identification of new drugs.
Innovation Solution
A knowledge transfer system that utilizes knowledge graph embeddings to represent all properties of SMILE data in a knowledge graph, training a latent space efficiently by combining graph embeddings and latent spaces, and decoding to identify new compounds, thereby conserving computing resources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple output layers and loss functions are used to train VAEs with multiple properties, then the model can capture more compound properties, but resource consumption increases and training efficiency decreases
Solution Approach 1:
The patent merges multiple output layers and loss functions into a single unified loss function that processes multiple compound properties simultaneously. By integrating the training objective into one cohesive framework, the system maintains the ability to handle multiple properties while significantly reducing computing resource consumption and eliminating the inefficiencies of separate training processes.
2Adaptability or versatility
If multiple output layers and loss functions are used to train VAEs with multiple properties, then the model can capture more compound properties, but training efficiency decreases leading to incorrect identification of new drugs
Solution Approach 1:
The patent combines multiple training objectives into a single unified loss function, which streamlines the training process and improves training efficiency. This merging eliminates the computational overhead and coordination complexity associated with multiple separate loss functions, enabling faster and more accurate identification of new drugs while maintaining comprehensive property analysis.
3Loss of information
If traditional VAE training with multiple properties is used, then compound properties can be analyzed, but the process is time consuming and expensive
Solution Approach 1:
The patent changes the training parameter structure by consolidating multiple properties into a single unified loss function with a unified set of parameters. This parameter consolidation transforms the training process from a multi-stage, time-consuming procedure into a single efficient operation, dramatically reducing drug discovery time while preserving all compound property information through the integrated framework.
Data Source
AI summary
A device may receive a knowledge graph and SMILE data identifying compounds, and may train embeddings based on the knowledge graph. The device may generate graph embeddings for the SMILE data based on the embeddings, and may encode the SMILE data into a latent space. The device may combine the graph embeddings and the latent space to generate a combined latent-embedding space, and may decode the combined latent-embedding space to generate decoded SMILE data. The device may utilize the decoded SMILE data to train an encoder, and may process source SMILE data, with the trained encoder, to generate a source combined latent-embedding space. The device may search the source combined latent-embedding space to identify new SMILE data, and may decode the new SMILE data to generate decoded new SMILE data. The device may evaluate the decoded new SMILE data to identify particular SMILE data associated with a new compound.


