Knowledge Graph Embeddings for Multi-Property Compound Discovery

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for new drug discovery using machine learning are resource-intensive and inefficient due to the need for multiple output layers and loss functions when training variational autoencoders (VAEs) with multiple properties, leading to improper training and incorrect identification of new drugs.

Innovation Solution

A knowledge transfer system that utilizes knowledge graph embeddings to represent all properties of SMILE data in a knowledge graph, training a latent space efficiently by combining graph embeddings and latent spaces, and decoding to identify new compounds, thereby conserving computing resources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple output layers and loss functions are used to train VAEs with multiple properties, then the model can capture more compound properties, but resource consumption increases and training efficiency decreases

Engineering Contradiction:
Improveability to handle multiple propertiesVSAvoidcomputing resource consumption
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The patent merges multiple output layers and loss functions into a single unified loss function that processes multiple compound properties simultaneously. By integrating the training objective into one cohesive framework, the system maintains the ability to handle multiple properties while significantly reducing computing resource consumption and eliminating the inefficiencies of separate training processes.

Inventive Principle:
Principle #5Merging (Combining)

2Adaptability or versatility

If multiple output layers and loss functions are used to train VAEs with multiple properties, then the model can capture more compound properties, but training efficiency decreases leading to incorrect identification of new drugs

Engineering Contradiction:
Improveability to handle multiple propertiesVSAvoidtraining efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent combines multiple training objectives into a single unified loss function, which streamlines the training process and improves training efficiency. This merging eliminates the computational overhead and coordination complexity associated with multiple separate loss functions, enabling faster and more accurate identification of new drugs while maintaining comprehensive property analysis.

Inventive Principle:
Principle #5Merging (Combining)

3Loss of information

If traditional VAE training with multiple properties is used, then compound properties can be analyzed, but the process is time consuming and expensive

Engineering Contradiction:
Improvecompound property informationVSAvoiddrug discovery time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent changes the training parameter structure by consolidating multiple properties into a single unified loss function with a unified set of parameters. This parameter consolidation transforms the training process from a multi-stage, time-consuming procedure into a single efficient operation, dramatically reducing drug discovery time while preserving all compound property information through the integrated framework.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12406773B2Transferring information through knowledge graph embeddings
Publication Date: 2025.09.02 ACCENTURE GLOBAL SOLUTIONS LTD
  • US12406773B2 patent drawing
  • US12406773B2 patent drawing
  • US12406773B2 patent drawing

AI summary

A device may receive a knowledge graph and SMILE data identifying compounds, and may train embeddings based on the knowledge graph. The device may generate graph embeddings for the SMILE data based on the embeddings, and may encode the SMILE data into a latent space. The device may combine the graph embeddings and the latent space to generate a combined latent-embedding space, and may decode the combined latent-embedding space to generate decoded SMILE data. The device may utilize the decoded SMILE data to train an encoder, and may process source SMILE data, with the trained encoder, to generate a source combined latent-embedding space. The device may search the source combined latent-embedding space to identify new SMILE data, and may decode the new SMILE data to generate decoded new SMILE data. The device may evaluate the decoded new SMILE data to identify particular SMILE data associated with a new compound.