Molecular generation optimization system based on reinforcement learning

Through a molecular generation optimization system based on reinforcement learning, using multi-stage learning and course learning strategies, the molecular generation strategy is dynamically adjusted, and the problems of molecular generation in the existing technology are difficult to meet the diversity, chemical rationality and drug properties optimization, and efficient and quality-optimized drug design is achieved.

CN120089239APending Publication Date: 2025-06-03BEIJING ANGOPRO TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510399779.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The prior art is difficult to meet the diversity of molecular generation, chemical rationality and drug properties optimization in drug design, especially in terms of dynamic adjustment of generation strategies and multi-objective optimization.

Method used

A molecular generation optimization system based on reinforcement learning is adopted, through multi-stage learning and course learning strategies, the molecular generation strategy is dynamically adjusted to optimize the drug properties of the molecules, while maintaining the diversity and chemical rationality of the generated molecules. The system includes a parameter input module, a multi-stage learning module, a course learning module, a dynamic reward function adjustment module and a molecular generation module.

Benefits of technology

It achieves the diversity and chemical rationality of the generated molecules while meeting specific drug design goals, significantly improves the efficiency and quality of molecular generation, and brings significant application benefits to drug design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120089239A_ABST
    Figure CN120089239A_ABST
Patent Text Reader

Abstract

The invention discloses a molecule generation optimization system based on reinforcement learning, and the system comprises a parameter input module which is used for inputting an initial molecular structure and a drug design target; the multi-stage learning module is used for gradually adjusting a molecule generation strategy, and the multi-stage learning module comprises at least two learning stages; the course learning module is used for gradually increasing the complexity and difficulty of molecule generation according to a preset course plan; wherein the curriculum plan comprises a plurality of curriculum stages, and different molecular generation tasks and learning targets are set in each curriculum stage; the dynamic reward function adjustment module is used for dynamically adjusting a reward function according to a molecule generation result of the current learning stage so as to optimize drug properties of molecules; and the molecule generation module is used for generating a molecular structure meeting a drug design target according to the output of the multi-stage learning module and the course learning module. The method not only improves the efficiency and quality of molecule generation, but also brings significant application benefits to drug design.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of drug design, and particularly relates to a molecular generation optimization system based on reinforcement learning. Background Art

[0002] In the field of molecular generation, traditional techniques mainly rely on rule-based expert systems and statistics-based machine learning methods. These methods are widely used in drug design and usually require a large amount of prior knowledge and data support. In recent years, with the development of artificial intelligence technology, molecular generation methods based on deep learning and reinforcement learning have gradually attracted attention. For example, the ClickGen model proposed by the team of Professor Hou Tingjun at Zhejiang University combines modular reactions and reinforcement learning techniques and can generate molecules with high diversity and synthetic accessibility. In addition, Patent CN117671350A proposes a method based on a deep conditional generation model, which generates target images of specified categories through an improved conditional generative adversarial network (CGAN), demonstrating the potential of deep generative models in complex data generation tasks.

[0003] Although the existing technologies have improved the efficiency and quality of molecular generation to a certain extent, there are still many limitations. Molecules generated by traditional methods often fail to meet specific drug design goals, such as drug activity, selectivity, metabolic stability, etc., and there are significant deficiencies in maintaining molecular diversity and chemical rationality. For example, although ClickGen can generate molecules with high synthetic accessibility, it still needs to be improved in dynamically adjusting the generation strategy to adapt to different drug design goals. Although Patent CN117671350A demonstrates the powerful capabilities of deep generative models, in the specific application in the field of drug design, how to combine multi-objective optimization of drug design remains a challenge. Summary of the Invention

[0004] To solve the above technical problems, the present invention proposes a molecular generation optimization system based on reinforcement learning, which dynamically adjusts the molecular generation strategy through multi-stage learning and curriculum learning strategies to optimize the drug properties of molecules while maintaining the diversity and chemical rationality of the generated molecules.

[0005] To achieve the above object, in a first aspect, the present invention provides a molecular generation optimization system based on reinforcement learning, comprising:

[0006] A parameter input module for inputting an initial molecular structure and a drug design goal;

[0007] A multi-stage learning module for gradually adjusting the molecular generation strategy, and the multi-stage learning module includes at least two learning stages;

[0008] A course learning module for gradually increasing the complexity and difficulty of molecular generation according to a preset course plan; wherein, the course plan includes multiple course stages, and different molecular generation tasks and learning objectives are set for each course stage;

[0009] A dynamic reward function adjustment module for dynamically adjusting the reward function according to the molecular generation results of the current learning stage to optimize the drug properties of the molecules;

[0010] A molecular generation module for generating a molecular structure that meets the drug design goal according to the outputs of the multi-stage learning module and the course learning module.

[0011] Preferably, each learning stage in the multi-stage learning module includes:

[0012] An initial molecular generation strategy for generating an initial molecular structure, and the initial molecular generation strategy adopts a preset molecular library or a random generation algorithm;

[0013] A strategy optimization unit for optimizing the initial molecular generation strategy according to the reward function of the current stage, and the strategy optimization unit adopts a reinforcement learning algorithm.

[0014] Preferably, the course learning module includes:

[0015] A course plan setting unit for presetting the molecular generation goals and difficulties of different learning stages, and the course plan setting unit is customized according to the drug design requirements;

[0016] A course progress control unit for adjusting the course progress according to the completion status of the current learning stage, and the course progress control unit can dynamically adjust the switching timing of the learning stages.

[0017] Preferably, the dynamic reward function adjustment module includes:

[0018] A molecular property evaluation unit for evaluating the drug properties of the generated molecular structure, and the molecular property evaluation unit adopts quantum chemical calculations or a machine learning model;

[0019] A reward function update unit for dynamically updating the reward function according to the molecular property evaluation results, and the reward function update unit adopts an adaptive adjustment algorithm.

[0020] Preferably, the molecular generation module includes:

[0021] A molecular structure generation unit for generating a molecular structure according to the optimized generation strategy, and the molecular structure generation unit adopts a deep generation model;

[0022] A diversity preservation unit, which is used to ensure the diversity of the generated molecular structures, and the diversity preservation unit is implemented by introducing a diversity reward mechanism;

[0023] A chemical rationality verification unit, which is used to verify the chemical rationality of the generated molecular structures, and the chemical rationality verification unit adopts chemical rules and constraints;

[0024] A result analysis unit, which is used to analyze the generated molecular structures and output a list of molecules that meet specific drug design goals, and the result analysis unit adopts statistical analysis methods and machine learning classification models.

[0025] Preferably, the system further includes a data storage module, which is used to store the molecular generation data and the reward function adjustment records of each learning stage, and the data storage module adopts a database management system.

[0026] In a second aspect, the present invention also discloses an optimization method for molecular generation based on reinforcement learning, which is used to implement the system described in the first aspect. The method includes the following steps:

[0027] Input an initial molecular structure and a drug design goal;

[0028] Gradually adjust the molecular generation strategy through a multi-stage learning module, and the multi-stage learning module includes at least two learning stages;

[0029] According to a preset curriculum plan, gradually increase the complexity and difficulty of molecular generation through a curriculum learning module; wherein, the curriculum plan includes multiple curriculum stages, and different molecular generation tasks and learning goals are set for each curriculum stage;

[0030] Dynamically adjust the reward function according to the molecular generation results of the current learning stage to optimize the drug properties of the molecules;

[0031] Generate molecular structures that meet the drug design goals according to the outputs of the multi-stage learning module and the curriculum learning module.

[0032] In a third aspect, the present invention also discloses a computer device, which includes a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the method described in the first aspect.

[0033] In a fourth aspect, the present invention also discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.

[0034] In a fifth aspect, the present invention also discloses a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.

[0035] Compared with the prior art, the present invention has the following advantages and technical effects:

[0036] The present invention provides a molecular generation optimization system based on reinforcement learning, including: a parameter input module for inputting an initial molecular structure and a drug design target; a multi-stage learning module for gradually adjusting the molecular generation strategy, where the multi-stage learning module includes at least two learning stages; a curriculum learning module for gradually increasing the complexity and difficulty of molecular generation according to a preset curriculum plan, where the curriculum plan includes multiple curriculum stages, and different molecular generation tasks and learning objectives are set for each curriculum stage; a dynamic reward function adjustment module for dynamically adjusting the reward function according to the molecular generation results of the current learning stage to optimize the drug properties of the molecule; and a molecular generation module for generating a molecular structure that meets the drug design target according to the outputs of the multi-stage learning module and the curriculum learning module.

[0037] By introducing multi-stage learning and curriculum learning strategies, the method of the present invention can gradually adjust the molecular generation strategy and dynamically optimize the reward function, so as to maintain the diversity and chemical rationality of the generated molecules while meeting specific drug design targets. This method not only improves the efficiency and quality of molecular generation, but also brings significant application benefits to drug design and has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0039] Figure 1 is the basic framework diagram of the system according to the embodiment of the present invention;

[0040] Figure 2 is the flowchart of multi-stage learning according to the embodiment of the present invention;

[0041] Figure 3 is the flowchart of curriculum learning according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine the embodiments to detail this application.

[0043] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0044] Example 1

[0045] In this embodiment, a molecular generation optimization system based on reinforcement learning is provided, including:

[0046] A parameter input module for inputting an initial molecular structure and a drug design goal;

[0047] A multi-stage learning module for gradually adjusting the molecular generation strategy. The multi-stage learning module includes at least two learning stages, and each learning stage corresponds to different molecular generation goals and reward functions. The multi-stage learning is carried out in a preset order to gradually optimize the molecular generation strategy;

[0048] Furthermore, each learning stage in the multi-stage learning module includes:

[0049] An initial molecular generation strategy for generating an initial molecular structure. The initial molecular generation strategy uses a preset molecular library or a random generation algorithm;

[0050] A strategy optimization unit for optimizing the initial molecular generation strategy according to the reward function of the current stage. The strategy optimization unit uses a reinforcement learning algorithm, such as Q-learning or Policy Gradient.

[0051] A curriculum learning module for gradually increasing the complexity and difficulty of molecular generation according to a preset curriculum plan. The curriculum plan includes multiple curriculum stages, and each curriculum stage sets different molecular generation tasks and learning goals;

[0052] Furthermore, the curriculum learning module includes:

[0053] A curriculum plan setting unit for presetting the molecular generation goals and difficulties of different learning stages. The curriculum plan setting unit is customized according to the drug design requirements;

[0054] A curriculum progress control unit for adjusting the curriculum progress according to the completion status of the current learning stage. The curriculum progress control unit can dynamically adjust the switching timing of the learning stages.

[0055] In this embodiment, the multi-stage learning module and the curriculum learning module are coordinated through a cooperative control mechanism to ensure the continuity and stability of the molecular generation process. The cooperative control mechanism uses a state feedback control and an adaptive adjustment strategy.

[0056] A dynamic reward function adjustment module for dynamically adjusting the reward function according to the molecular generation results of the current learning stage to optimize the drug properties of the molecule. The dynamic adjustment includes correcting the reward value based on the molecular property evaluation results;

[0057] Furthermore, the dynamic reward function adjustment module includes:

[0058] A molecular property evaluation unit for evaluating the drug properties of the generated molecular structure, and the molecular property evaluation unit is based on quantum chemical calculations or machine learning models;

[0059] A reward function update unit for dynamically updating the reward function according to the molecular property evaluation results, and the reward function update unit adopts an adaptive adjustment algorithm.

[0060] In this embodiment, the dynamic reward function adjustment module can adjust the reward function in real time according to external feedback information to meet different drug design requirements, and the external feedback information includes experimental data and expert evaluation results.

[0061] A molecular generation module for generating a molecular structure that meets specific drug design goals according to the outputs of the multi-stage learning module and the curriculum learning module. The molecular generation module can maintain the diversity and chemical rationality of the generated molecules and is controlled by a diversity maintenance unit and a chemical rationality verification unit.

[0062] Furthermore, the molecular generation module includes:

[0063] A molecular structure generation unit for generating a molecular structure according to the optimized generation strategy, and the molecular structure generation unit adopts a deep generation model, such as a generative adversarial network (GAN) or a variational autoencoder (VAE);

[0064] A diversity maintenance unit for ensuring the diversity of the generated molecular structures, and the diversity maintenance unit is realized by introducing a diversity reward mechanism;

[0065] A chemical rationality verification unit for verifying the chemical rationality of the generated molecular structures, and the chemical rationality verification unit adopts chemical rules and constraints.

[0066] A result analysis unit for analyzing the generated molecular structures and outputting a list of molecules that meet specific drug design goals, and the result analysis unit adopts statistical analysis methods and machine learning classification models.

[0067] A data storage module for storing the molecular generation data and reward function adjustment records of each learning stage, and the data storage module adopts a database management system;

[0068] This embodiment is applied to the field of drug design and specifically includes:

[0069] A drug target selection unit for selecting specific drug targets, and the drug target selection unit adopts disease-related gene and protein information;

[0070] A molecular screening unit, which is used to screen out molecular structures that meet the requirements of specific drug targets from the generated molecular list. The molecular screening unit adopts virtual screening and biological activity testing.

[0071] The basic framework diagram of the molecular generation optimization system based on reinforcement learning is as Figure 1 shown.

[0072] 1. Initialization of the environment module:

[0073] The environment module is used to provide the initial state and feedback information for molecular generation. The initial state includes a set of predefined molecular structures and a chemical property database.

[0074] The environment module conducts data interaction with the agent module through an interface, transmits the current state information, and receives the generation actions of the agent module.

[0075] 2. Setting of the agent module:

[0076] The agent module adopts a deep neural network model, specifically a generative adversarial network (GAN) or a variational autoencoder (VAE), to generate new molecular structures.

[0077] The agent module receives the state information of the environment module and generates new molecular structures according to the internal policy, and outputs them to the environment module for evaluation.

[0078] 3. Design of the reward function module:

[0079] The reward function module dynamically adjusts the reward function according to the drug design goal. The reward function comprehensively considers the drug properties, chemical rationality, and diversity of the molecule.

[0080] Specific reward functions include, but are not limited to, indicators such as the binding affinity, toxicity, and solubility of the molecule, and the comprehensive reward value is calculated by weighted summation.

[0081] 4. Multi-stage learning and curriculum learning:

[0082] Multi-stage learning (Staged Learning) divides the molecular generation process into multiple stages, and different learning goals and reward functions are set for each stage.

[0083] Curriculum learning (Curriculum Learning) gradually increases the difficulty of the learning tasks according to the performance of the agent module, and gradually transitions from simple molecular structures to complex molecular structures.

[0084] 5. Optimization module:

[0085] The optimization module receives the feedback from the reward function module and updates the parameters of the agent module using reinforcement learning algorithms (such as PPO, DQN) to optimize the generation strategy.

[0086] During the optimization process, through multiple iterations, the quality and diversity of the generated molecules are gradually improved.

[0087] Figure 2 This is the multi-stage learning flow chart of the present invention, showing the learning objectives and reward function settings at different stages during the molecule generation process. It includes the following steps:

[0088] 1. The first stage: basic structure generation:

[0089] Learning objective: Generate a molecular structure that conforms to basic chemical rules.

[0090] Reward function: Mainly consider the chemical rationality and diversity of the molecule. The reward function is R 1 = w 1 · Rationality + w 2 · Diversity.

[0091] In the first stage: The agent module generates a basic molecular structure, the environment module evaluates its chemical rationality and diversity, and the optimization module updates the parameters.

[0092] 2. The second stage: drug property optimization:

[0093] Learning objective: Optimize the drug properties of the molecule based on the basic structure.

[0094] Reward function: Add drug property indicators such as binding affinity and toxicity. The reward function is R 2 = w 1 · Rationality + w 2 · Diversity + w 3 · Drug properties.

[0095] In the second stage: The agent module optimizes the drug properties based on the basic structure, the environment module evaluates the new indicators, and the optimization module continues to update the parameters.

[0096] 3. The third stage: comprehensive optimization:

[0097] Learning objective: Comprehensively optimize all indicators of the molecule to achieve the drug design goal.

[0098] Reward function: Comprehensively consider all indicators. The reward function is R 3 = w 1 · Rationality + w 2 · Diversity + w 3 · Drug properties + w 4 · Other indicators.

[0099] In the third stage: The agent module comprehensively optimizes all indicators, the environment module conducts a comprehensive evaluation, and the optimization module finally adjusts the parameters.

[0100] Figure 3 This is the flowchart of the course learning for the present invention, which shows the process of gradually increasing the difficulty of learning tasks according to the performance of the agent module. It includes the following steps:

[0101] 1. Initial task setting:

[0102] Set simple molecular generation tasks, such as generating small molecule structures.

[0103] The reward function is relatively simple and mainly considers chemical rationality.

[0104] Initial stage: The agent module starts with simple tasks, the environment module evaluates its performance, and the optimization module updates the parameters.

[0105] 2. Task difficulty increment:

[0106] According to the performance of the agent module, gradually increase the task difficulty, such as generating molecular structures of medium complexity.

[0107] The reward function adds a diversity metric.

[0108] Incremental stage: According to the performance of the agent module, gradually increase the task difficulty, and the environment module and the optimization module make corresponding adjustments.

[0109] 3. Advanced task challenge:

[0110] When the performance of the agent module is stable, set advanced tasks, such as generating complex drug molecules.

[0111] The reward function synthesizes all metrics for comprehensive optimization.

[0112] Advanced stage: The agent module challenges advanced tasks, the environment module conducts a comprehensive evaluation, and the optimization module finally adjusts the parameters.

[0113] The following is a specific example:

[0114] Input: The initial molecular structure is "C6H6", and the drug design goal is to improve anti-inflammatory activity.

[0115] Multi-stage learning:

[0116] First stage: Generate various benzene ring derivatives.

[0117] Second stage: Screen and optimize molecules with potential anti-inflammatory activity.

[0118] Third stage: Ensure the generation of various anti-inflammatory molecules with different structures.

[0119] Course learning:

[0120] Initial stage: Learn to generate simple benzene ring derivatives.

[0121] Intermediate stage: Learn to generate anti-inflammatory molecules of medium complexity.

[0122] Advanced stage: Generate highly optimized anti-inflammatory molecules.

[0123] Dynamic adjustment of the reward function: Dynamically adjust the reward function according to the test results of the anti-inflammatory activity of the generated molecules to further optimize the generation strategy.

[0124] Output: Generate a series of molecular structures with high anti-inflammatory activity.

[0125] Through the above technical solutions, this embodiment can effectively optimize the drug properties of molecules, while maintaining the diversity and chemical rationality of the generated molecules, and significantly improve the efficiency and effectiveness of drug design.

[0126] Embodiment 2

[0127] Based on the same inventive concept, this embodiment also provides a method for optimizing molecular generation based on reinforcement learning, including the following steps:

[0128] Input the initial molecular structure and the drug design goal;

[0129] Gradually adjust the molecular generation strategy through a multi-stage learning module, where the multi-stage learning module includes at least two learning stages;

[0130] According to a preset curriculum plan, gradually increase the complexity and difficulty of molecular generation through a curriculum learning module; where the curriculum plan includes multiple curriculum stages, and different molecular generation tasks and learning goals are set for each curriculum stage;

[0131] Dynamically adjust the reward function according to the molecular generation results of the current learning stage to optimize the drug properties of the molecules;

[0132] Generate a molecular structure that meets the drug design goal according to the outputs of the multi-stage learning module and the curriculum learning module.

[0133] A method for optimizing molecular generation based on reinforcement learning provided by this embodiment has all the advantages of a system for optimizing molecular generation based on reinforcement learning provided by Embodiment 1.

[0134] Embodiment 3

[0135] This embodiment also discloses a computer device, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the method described in Embodiment 1.

[0136] Embodiment 4

[0137] This embodiment also discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in Embodiment 1 are implemented.

[0138] Embodiment 5

[0139] This embodiment also discloses a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the method described in Embodiment 1 are implemented.

[0140] The above are only the preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A molecular generation optimization system based on reinforcement learning, characterized in that: include: Parameter input module, used to input initial molecular structure and drug design goals; A multi-stage learning module for gradually adjusting the molecule generation strategy, wherein the multi-stage learning module includes at least two learning stages; A course learning module, which is used to gradually increase the complexity and difficulty of molecule generation according to a preset course plan; wherein the course plan includes multiple course stages, each of which sets different molecule generation tasks and learning objectives; A dynamic reward function adjustment module, which is used to dynamically adjust the reward function according to the molecule generation results in the current learning phase to optimize the drug properties of the molecule; The molecule generation module is used to generate molecular structures that meet the drug design goals based on the outputs of the multi-stage learning module and the curriculum learning module.

2. The system according to claim 1, characterized in that Each learning stage in the multi-stage learning module includes: An initial molecule generation strategy, used to generate an initial molecular structure, wherein the initial molecule generation strategy adopts a preset molecule library or a random generation algorithm; A strategy optimization unit is used to optimize the initial molecule generation strategy according to the reward function of the current stage, and the strategy optimization unit adopts a reinforcement learning algorithm.

3. The system according to claim 1, characterized in that The course learning modules include: A course plan setting unit, used to preset the molecular generation goals and difficulty of different learning stages, and the course plan setting unit is customized according to drug design requirements; The course progress control unit is used to adjust the course progress according to the completion status of the current learning stage. The course progress control unit can dynamically adjust the switching timing of the learning stage.

4. The system according to claim 1, characterized in that The dynamic reward function adjustment module includes: A molecular property evaluation unit, used to evaluate the drug properties of the generated molecular structure, wherein the molecular property evaluation unit adopts quantum chemical calculation or machine learning model; The reward function updating unit is used to dynamically update the reward function according to the molecular property evaluation result, and the reward function updating unit adopts an adaptive adjustment algorithm.

5. The system according to claim 1, characterized in that The molecule generation module comprises: A molecular structure generation unit, used to generate a molecular structure according to the optimized generation strategy, wherein the molecular structure generation unit adopts a deep generation model; A diversity maintenance unit, used to ensure that the generated molecular structures have diversity, and the diversity maintenance unit is implemented by introducing a diversity reward mechanism; A chemical rationality verification unit, used to verify the chemical rationality of the generated molecular structure, wherein the chemical rationality verification unit adopts chemical rules and constraints; The result analysis unit is used to analyze the generated molecular structures and output a list of molecules that meet specific drug design goals. The result analysis unit adopts statistical analysis methods and machine learning classification models.

6. The system according to claim 1, characterized in that It also includes a data storage module for storing the molecule generation data and reward function adjustment records of each learning stage, and the data storage module adopts a database management system.

7. A molecular generation optimization method based on reinforcement learning, characterized in that: For implementing the system according to any one of claims 1 to 6, the method comprises the following steps: Input initial molecular structure and drug design goals; gradually adjusting the molecule generation strategy through a multi-stage learning module, wherein the multi-stage learning module includes at least two learning stages; According to the preset course plan, the complexity and difficulty of molecule generation are gradually increased through the course learning module; wherein the course plan includes multiple course stages, and each course stage sets different molecule generation tasks and learning objectives; Dynamically adjust the reward function based on the molecule generation results of the current learning phase to optimize the drug properties of the molecule; Generate molecular structures that meet drug design goals based on the outputs of multi-stage learning modules and course learning modules.

8. A computer device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method of claim 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to claim 7 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to claim 7 are implemented.

Citation Information

Patent Citations

  • SAR target image generation method based on depth condition generation model

    CN117671350A