SAFE Representation Generative Chemistry Model

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current machine learning models in drug discovery face challenges in optimizing compounds for specific protein targets due to discrete latent spaces, which hinder gradient-based optimization and tasks like scaffold decoration and linker design, limiting their effectiveness in enhancing binding affinity and selectivity.

Innovation Solution

A generative chemistry model based on SAFE representations is developed, using an encoder-decoder transformer model and a variational autoencoder to create a continuous latent space, enabling gradient-based optimization and advanced optimization tasks such as scaffold decoration and morphing through techniques like masking and motif reconstruction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If transformer models with discrete latent spaces are used for compound generation, then the ability to generate valid compounds is improved, but gradient-based optimization for specific protein targets becomes impossible

Engineering Contradiction:
Improvecompound validityVSAvoidgradient-based optimization
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent transforms the discrete latent space parameters into continuous parameters, enabling gradient-based optimization while preserving compound validity through the VAE framework that maintains the encoder-decoder architecture's ability to generate valid SMILES strings

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If linear representations like SMILES are used for training, then broad chemical structure generation is achieved, but advanced optimization tasks like scaffold decoration and linker design become difficult

Engineering Contradiction:
Improvechemical structure generationVSAvoidscaffold decoration precision
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent segments the molecular representation into distinct components (scaffold, linkers, functional groups) within the SAFE representation framework, allowing independent optimization of each component while maintaining overall molecular validity and enabling precise control over scaffold decoration and linker design tasks

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250014689A1Method and system for developing generative chemistry model based on sequential attachment-based fragment embedding (SAFE) representation
Publication Date: 2025.01.09 QUANTIPHI INC
  • US20250014689A1 patent drawing
  • US20250014689A1 patent drawing
  • US20250014689A1 patent drawing

AI summary

A method and system for developing a generative chemistry model based on SAFE representations is disclosed. The method includes encoding one or more chemical compounds into the SAFE representations. The method may include training an encoder-decoder transformer model based on the SAFE representations using one or more masking techniques. The encoder-decoder transformer model creates a latent space to represent encoded SAFE representations of the chemical compounds. The method may further include training a variational Autoencoder (VAE) to generate a continuous latent space by compressing the latent space of the encoder-decoder transformer model. The method may further include optimizing the continuous latent space to generate a plurality of chemical compounds with specific properties by decoding the SAFE representations of the chemical compounds.