A phenotype and target point synergistic guidance molecule generation method based on multi-objective reinforcement learning

By integrating phenotypic and target information through multi-objective reinforcement learning, the molecular generation process is optimized, solving the problems of long time consumption and high cost in traditional drug discovery. This generates candidate molecules with high targeting affinity and good drug-like properties, promoting the precision drug design for complex diseases such as cancer.

CN122494029APending Publication Date: 2026-07-31NANJING TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING TECH UNIV
Filing Date
2026-03-12
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Traditional drug discovery processes are time-consuming, costly, and have low success rates. Especially when treating complex diseases such as cancer, target-directed methods ignore cellular system responses, leading to off-target effects, while phenotype-directed methods ignore structural diversity, making it difficult for generated molecules to achieve the desired effects.

Method used

By employing multi-objective reinforcement learning to combine phenotypic and target information, and integrating transcriptomic phenotypic features and protein structures through a dual-channel VAE architecture, a compound reward mechanism and multi-objective reinforcement learning are introduced to optimize the molecule generation process, ensuring that the generated molecules have high affinity for the target and meet phenotypic requirements.

Benefits of technology

Generate candidate molecules with high targeting affinity and good drug-like properties to improve drug design efficiency and stability, applicable to precision drug design for complex diseases such as cancer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122494029A_ABST
    Figure CN122494029A_ABST
Patent Text Reader

Abstract

This invention relates to the field of deep learning and artificial intelligence drug design technology, specifically to a method and system for phenotypic and target-guided molecular generation based on multi-objective reinforcement learning. The method includes: pre-training a molecular generation network conditioned on expression profiles using drug-induced transcriptome expression profile data and drug molecule datasets to establish a mapping relationship between biological phenotypic features and molecular chemical space; constructing a multi-objective reward function by integrating target protein docking affinity, molecular drug-likeness, and synthetic accessibility scores; and optimizing the ranking of generated molecule pairs based on molecular attribute scores. This invention effectively bridges the gap between phenotypic screening and target validation by synergistically integrating phenotypic features and target-specific information. The generated candidate molecules maintain consistency with the target phenotype while possessing significantly improved target affinity and excellent physicochemical properties, enabling the generation of drug-like molecules with high efficacy, strong diversity, and specific biological effects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of deep learning and artificial intelligence drug design technology, specifically to a molecular generation method based on multi-objective reinforcement learning. Background Technology

[0002] In the medical field, drug discovery is a core process in treating diseases, especially complex ones such as cancer, cardiovascular disease, and neurodegenerative diseases. Traditional drug discovery relies on high-throughput screening and clinical trials, but this process is time-consuming, costly, and has a low success rate. Cancer, as one of the leading causes of death worldwide, faces even greater challenges in treatment: strong tumor heterogeneity, high drug resistance, and complex mechanisms lead to the failure of many candidate drugs in clinical trials. Medical research emphasizes "precision medicine," which involves designing personalized drugs based on a patient's genotype, phenotype, and molecular targets to improve efficacy and reduce side effects. In medical applications, target-guided generative methods utilize protein structure information to optimize the binding affinity between molecules and targets, but they neglect cellular system-level responses, potentially leading to off-target effects or in vivo failure. The efficacy of drug molecules against their targets is heavily influenced by the complex cellular environment, and drugs designed to bind to targets do not always elicit the expected changes at the cellular phenotype level. Phenotypic molecular generation methods, which do not rely on specific targets, focus on utilizing cellular-level response information to guide models in generating molecules capable of regulating disease states. They are particularly suitable for situations where disease mechanisms are complex or unclear, offering the advantage of capturing systemic responses. However, they are costly to experiment with, difficult to standardize, and challenging to trace molecular mechanisms. Furthermore, phenotype-guided methods often neglect the diversity and constraints at the structural level. During the generation process, most methods do not consider the structural matching between molecules and proteins, potentially leading to generated molecules with phenotypic driving potential that fail to achieve the expected effects in actual binding. Summary of the Invention

[0003] The purpose of this invention is to develop a phenotypic and target-guided molecular generation method that combines phenotypic and target-guided approaches for de novo molecular design. By integrating multimodal data fusion of gene expression profiles, protein structures, and molecular properties, and combining this with a multi-objective reinforcement learning model, accurate molecular generation is achieved, further optimizing the molecule's affinity for the target and its phenotypic effects. This method not only generates molecules with high affinity for specific targets but also improves generation efficiency and stability, thereby promoting the widespread application of precision drug design in complex diseases such as cancer.

[0004] This invention is based on multi-objective reinforcement learning, achieving targeted generation of drug-like molecules through the synergistic integration of transcriptomic phenotypic features and target structural information. 1) Data preparation: Collecting large-scale small molecule drug datasets, drug-induced transcriptomic expression profile datasets, and 3D structural data of specific target proteins. 2) Phenotypic-guided generator pre-training: Drug molecule sequences are input into a molecular variational autoencoder for structural feature learning. Drug-induced expression profiles are input into an expression profile variational autoencoder for phenotypic feature encoding. 3) Dual-channel latent space construction: A dual-channel VAE architecture is used to map molecular structure and expression profile features to a joint latent distribution space, establishing a correlation between phenotypic conditions and chemical structures. 4) A molecular docking engine is used to calculate the binding affinity score between the generated molecule and the target protein. Drug-likeness quantitative assessment is used to ensure that the generated molecules possess excellent physicochemical properties. 5) Using the pre-trained phenotypic generator as a prior model, a reinforcement learning algorithm is used to fine-tune the generating agent, guiding it to explore high-reward regions. 6) Introducing a ranking loss: Samples within a batch are ranked based on attribute scores, enhancing the model's ability to identify high-performance structures. Prior likelihood regularization is introduced to ensure that the generated molecules conform to phenotypic constraints and do not deviate from the effective chemical space. Entropy regularization is introduced to maintain the randomness of the generation strategy and improve the structural diversity of the generated molecules. 7) Integrate strategy gradient loss, ranking loss, prior loss and entropy loss to achieve joint optimization of molecular quality, target specificity and phenotypic correlation.

[0005] The specific technical solution of this invention is: a phenotype and target synergistic guided molecular generation method based on multi-objective reinforcement learning, comprising the following steps:

[0006] The model requires data including large-scale drug molecule datasets, drug-induced differential transcriptome expression profile datasets, 3D structural data of the target protein, and a library of known ligand molecules. This invention uses the transcriptome profile obtained after 24 hours of 10 μM drug exposure as phenotypic feature input and pre-calculates the binding pocket region of the target protein.

[0007] The drug molecule encoder uses a GRU-based recurrent neural network to extract drug molecule sequence features. The expression profile encoder uses a multilayer perceptron to extract transcriptomic features of drug perturbations. A mapping between phenotypic conditions and chemical space is established through a dual-channel VAE. The expression profile decoder is responsible for reconstructing transcriptomic perturbation features, and the molecular decoder is responsible for reconstructing the drug molecule structure based on the latent representation.

[0008] A composite reward mechanism is defined, including target binding affinity reward and drug-likeness reward. Multi-objective reinforcement learning is used for model fine-tuning, with a pre-trained model as the initial agent, and policy gradient updates driven by multi-objective rewards. A ranking loss is introduced, dynamically ranking samples based on their attribute scores within a batch to enhance the model's ability to search the high-performance chemical space. Prior likelihood regularization is integrated into the loss function to prevent mode collapse, and an entropy maximization term is added to improve the structural diversity of generated molecules. By jointly optimizing the policy gradient loss, ranking loss, and regularization loss, precise control over the quality of generated molecules is achieved.

[0009] The present invention has the following beneficial effects:

[0010] 1. It can generate candidate molecules with high targeting affinity and good drug-like properties for multiple key targets related to cancer.

[0011] 2. Drug-induced transcriptomic phenotypic data can be directly used to generate molecules that can induce phenotypic changes in target cells without prior knowledge of the target.

[0012] 3. In the new drug screening stage, it can simultaneously optimize multiple drug properties, including high targeting binding ability, good drug-like properties and high synthetic feasibility, avoiding the problem of drug development failure caused by optimizing a single property. Attached Figure Description

[0013] Figure 1 This is a diagram of the model algorithm framework.

[0014] Figure 2 The model is compared with other algorithms in terms of effectiveness, uniqueness, and novelty. Detailed Implementation

[0015] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0016] Reference Figure 1 A phenotypic and target-guided molecular generation method based on multi-objective reinforcement learning includes the following steps:

[0017] S1: Pre-training phase. Using a large-scale drug molecule dataset and a drug-induced transcriptome expression profile dataset, a dual-channel variational autoencoder establishes a mapping between phenotype and molecular structure. Phenotypic features are extracted using the expression profile encoder, and drug molecules are reconstructed using the molecular decoder, resulting in a priori generative model with phenotype-guided capabilities.

[0018] S2: Reward Function Construction Stage. A multi-objective evaluation system is constructed for the target protein. A molecular docking engine is used to calculate the affinity score between the molecule and the target protein, and simultaneously calculate the quantitative drug-likeness score of the molecule. The multi-objective reward calculation formula is as follows:

[0019] Reward(s) = DS(s) × QED(s)

[0020] Here, 's' represents a molecule, DS is the processed result of the docking score calculated by the molecular docking engine, and the QED score is the calculated quantitative drug-likeness score of the molecule. It's worth noting that when the generated molecule is unreasonable or does not meet the requirements, DS is directly set to 0 as a penalty. Simultaneously, we introduce a rescaling parameter k to normalize the scoring function to the range [0, 1]. The complete definition of DS is as follows:

[0021]

[0022] S3: Multi-objective reinforcement learning fine-tuning stage. The pre-trained model is used as the initial agent and fine-tuned using a policy gradient algorithm. A ranking loss is introduced to enhance the sampling probability of high-quality molecules by comparing the attribute scores of samples within a batch; the expected property score of a molecule is represented as AS(s), which is defined based on docking scores and related chemical properties. Given a pair of molecules (s... i s j If AS(s) is satisfied i )>AS(s j We expect the agent model to assign higher likelihood probabilities to molecules with high property scores. The ranking loss is defined as follows:

[0023]

[0024] Where, γ ij = (ji)·γ, representing the ranking difference between two molecules multiplied by a hyperparameter boundary value. The log-likelihood probability is estimated by an agent model with parameter θ. Simultaneously, prior likelihood regularization is introduced to constrain the distribution of generated molecules from deviating from the effective chemical space.

[0025] S4: Comprehensive Optimization Phase. The network is trained using a multi-task joint loss function, defined as follows:

[0026]

[0027] in To enhance the gradient loss of the learning strategy, This indicates a ranking loss. This represents the prior likelihood regularization term. Let be the entropy to maintain molecular diversity, and λ be the adjustment weight of each loss term.

[0028] In the network model:

[0029] Phenotypic Guided Generator: Receives the target transcriptome spectrum as conditional input and generates candidate molecular sequences that conform to specific biological effects.

[0030] Multi-target evaluation module: Real-time evaluation of target affinity, drug-likeness, and phenotypic consistency of generated molecules, providing reward feedback for reinforcement learning.

[0031] Dynamic ranking mechanism: The generated samples are ranked by performance through ranking loss, capturing subtle features of high-performance chemical structures and improving the quality of generated molecules.

[0032] In this embodiment, the experiment used a dataset containing expression profiles of 978 key genes. To evaluate the performance of this invention in generating drug-like molecules, we validated it against 10 challenging protein targets and compared its performance with other state-of-the-art methods: Figure 2 This demonstrates that the model maintains a high level in terms of effectiveness, uniqueness, and novelty.

Claims

1. A phenotypic and target-guided molecular generation method based on multi-objective reinforcement learning, characterized in that: Includes the following steps: 1) Pre-trained phenotypic-guided generative model: Using drug-induced transcriptome expression profile data and drug molecule datasets, a molecular generative network conditioned on expression profiles is constructed and pre-trained to establish a mapping relationship from biological phenotypic features to molecular chemical space; 2) Construct a multi-objective reward function: integrate the docking affinity score of the target protein, the molecular drug efficacy assessment index and the molecular synthesis accessibility score to form a multi-dimensional feedback evaluation mechanism; 3) Perform multi-objective reinforcement learning fine-tuning: Use the pre-trained generator network as the initial agent, and fine-tune and optimize it through the policy gradient algorithm under the drive of the multi-objective reward function; 4) Introduce sorting loss constraint: During fine-tuning, sort the generated molecules according to their attribute scores within the batch, calculate the sorting loss and update the model parameters to enhance the model's sampling probability for high-performance molecules.

2. The phenotypic and target-guided molecular generation method based on multi-objective reinforcement learning according to claim 1, characterized in that: In step 1), the molecular generation network adopts a dual-channel VAE architecture, including an expression profile encoder for extracting transcriptome features and a drug molecule encoder / decoder for extracting and generating chemical structures.

3. The phenotypic and target-guided molecular generation method based on multi-objective reinforcement learning according to claim 1, characterized in that: In step 2), the docking affinity score of the target protein is the binding energy value obtained by simulating the flexible docking of the generated molecular ligand with the 3D structure of the target protein.

4. The phenotypic and target-guided molecular generation method based on multi-objective reinforcement learning according to claim 1, characterized in that: The ranking loss calculation logic in step 4) is as follows: extract molecular pairs from the preset molecular set, sort the molecular pairs according to the expected property scores of each molecular, and constrain the generation network through the loss function so that the difference between the logarithmic generation probability of the molecular with higher property scores and the logarithmic generation probability of the molecular with lower property scores tends to be greater than a preset ranking difference threshold.

5. The phenotypic and target-guided molecular generation method based on multi-objective reinforcement learning according to claim 1, characterized in that: During the fine-tuning phase of reinforcement learning, the total loss function also includes a prior likelihood regularization term, which is used to constrain the fine-tuned model from deviating from the effective chemical space established by pre-training and to prevent mode collapse.

6. The phenotypic and target-guided molecular generation method based on multi-objective reinforcement learning according to claim 1, characterized in that: The total loss function also includes an entropy maximization term, which is used to maintain the randomness of the generation strategy to ensure the chemical diversity of the output molecular structure.