Drug molecule design system combining curriculum learning and reinforcement learning

Through a molecular design system combining curriculum learning and reinforcement learning, the problems of low efficiency, insufficient diversity and premature convergence in the existing technology are solved, efficient and diverse molecular generation is achieved, and the efficiency and economic benefits of drug research and development are significantly improved.

CN120108560APending Publication Date: 2025-06-06BEIJING ANGOPRO TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510400838.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing molecular design methods are inefficient, lack of diversity and are prone to convergence to the local optimal solution prematurely.

Method used

Combining the drug molecular design system of course learning and reinforcement learning, the molecular design process is gradually guided through the collaborative work of data preprocessing, course learning modules and reinforcement learning modules to avoid premature convergence.

Benefits of technology

It significantly improves the efficiency and diversity of molecular generation, avoids the trap of local optimal solution, shortens the drug development cycle, and reduces costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108560A_ABST
    Figure CN120108560A_ABST
Patent Text Reader

Abstract

The invention discloses a course learning and reinforcement learning combined drug molecule design system, which comprises a data preprocessing module used for performing cleaning, standardization and feature extraction on input molecular data; the course learning module is used for gradually guiding the learning process of molecular design according to a preset course strategy, reinforcing the learning process and providing molecular design tasks with different difficulties; the reinforcement learning module is connected with the course learning module and is used for performing molecular generation and optimization under the guidance of the course learning module; the molecule generation module is connected with the reinforcement learning module and is used for generating a target molecule structure according to the output of the reinforcement learning module; and the evaluation module is connected with the molecule generation module and is used for evaluating the performance of the target molecule structure and feeding back an evaluation result to the reinforcement learning module so as to optimize the learning process. By optimizing the molecular design process, the screening efficiency and success rate of the drug candidate molecules are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of molecular design, and in particular relates to a drug molecule design system combining course learning and reinforcement learning. Background Art

[0002] In the field of drug discovery, traditional molecular design methods mainly rely on experimental screening and rule-based expert systems. These methods usually require a large amount of experimental data and expertise, and the design process is time-consuming and inefficient. In addition, traditional methods are prone to falling into local optimal solutions when generating new molecules, making it difficult to explore molecular structures with high diversity and novelty, thus limiting the breadth and depth of drug research and development.

[0003] In response to the above problems, some patents have proposed some improved methods in recent years. For example, patent CN109523819A discloses a molecular generation method based on deep learning. This method generates new molecular structures by constructing a deep neural network model and using a large amount of known molecular data for training. However, this method tends to converge prematurely during the training process, and the generated molecules are not diverse enough. Another patent, CN110999082A, proposes a reinforcement learning-driven molecular design method that guides the molecular generation process by defining a reward function. However, in practical applications, this method is sensitive to initial parameters and requires a lot of computing resources, making it difficult to run efficiently on large-scale data sets.

[0004] Although existing patents have improved the efficiency and diversity of molecular design to a certain extent, there are still problems such as premature convergence and high consumption of computing resources. Summary of the invention

[0005] In order to solve the problems of low efficiency, insufficient diversity and premature convergence to local optimal solutions in existing molecular design methods, the present invention provides a drug molecule design system combining curriculum learning and reinforcement learning, including:

[0006] Data preprocessing module, used for cleaning, standardization and feature extraction of input molecular data;

[0007] A course learning module, connected to the data preprocessing module, is used to gradually guide the learning process of molecular design according to a preset course strategy, strengthen the learning process, and provide molecular design tasks of different difficulty levels;

[0008] A reinforcement learning module, connected to the course learning module, for performing molecule generation and optimization under the guidance of the course learning module;

[0009] A molecule generation module, connected to the reinforcement learning module, for generating a target molecular structure according to an output of the reinforcement learning module;

[0010] An evaluation module, connected to the molecule generation module, is used to evaluate the performance of the target molecular structure and feed the evaluation result back to the reinforcement learning module to optimize the learning process.

[0011] Preferably, the course learning module includes a course strategy formulation unit and a course progress control unit;

[0012] The course strategy formulation unit is used to formulate a course strategy according to preset rules, and dynamically adjust the course strategy according to the difficulty of molecular design and prior knowledge;

[0013] The course progress control unit is used to control the learning progress of molecular design according to the course strategy.

[0014] Preferably, the course strategy formulation unit includes a task grading unit and a strategy guiding unit;

[0015] The task classification unit is used to classify the molecular design tasks into multiple levels according to the complexity and design difficulty of the molecules;

[0016] The strategy guidance unit is used to provide tasks of relatively low difficulty in the initial stage, and gradually increase the difficulty of the tasks as the learning progresses, so that the reinforcement learning module gradually adapts to complex molecular design.

[0017] Preferably, the reinforcement learning module includes a state representation unit, an action selection unit, a reward function calculation unit, and an optimization learning algorithm unit;

[0018] The state representation unit is used to define the state space in the molecular design process, and to represent the molecular design problem as a state in reinforcement learning; the state space includes molecular structure characteristics and chemical properties;

[0019] The action selection unit is used to define possible operations and then select an action to generate a molecule according to the current state; the operation includes adding, deleting or modifying atoms or bonds in the molecule;

[0020] The reward function calculation unit is used to design a reward function, calculate the performance reward value of the generated molecule according to the reward function, and evaluate the quality of the generated molecular structure through the performance reward value; the reward function comprehensively considers factors such as the structural complexity, biological activity and synthetic feasibility of the molecule to calculate the reward value.

[0021] Preferably, the molecule generation module includes a molecule structure generation unit and a molecule property prediction unit;

[0022] The molecular structure generating unit is used to generate a molecular structure according to the output of the reinforcement learning module;

[0023] The molecular property prediction unit is used to predict the properties of the generated molecules.

[0024] Preferably, the molecular structure generating unit includes a regular structure generating unit and a legality checking unit;

[0025] The rule structure generating unit is used to generate a specific molecular structure according to chemical rules based on the output of the reinforcement learning module;

[0026] The legality checking unit is used to ensure that the generated molecular structure complies with chemical legality; the molecular structure includes valence bond rules and stereochemistry;

[0027] The molecular property prediction unit uses a deep learning model to predict the physicochemical properties and biological activities of the generated molecules.

[0028] Preferably, the evaluation module includes a performance indicator calculation unit and a feedback adjustment unit;

[0029] The performance index calculation unit is used to define and calculate the performance index of the generated molecule; the performance index includes the biological activity, stability, and synthetic feasibility of the molecule;

[0030] The feedback adjustment unit is used to adjust the parameters and learning strategy of the reinforcement learning module according to the performance index to optimize the generation process.

[0031] Preferably, the performance index calculation unit includes a plurality of submodules for respectively calculating different performance indexes of the molecule and feeding back the results to the reinforcement learning module;

[0032] The feedback adjustment unit adopts an adaptive adjustment strategy to dynamically adjust the learning parameters and learning strategy of the reinforcement learning module according to the feedback performance indicators.

[0033] Preferably, the course learning module and the reinforcement learning module communicate with each other via a data interface;

[0034] The data interface includes a data transmission unit and a data synchronization unit;

[0035] The data transmission unit is used to transmit data between the course learning module and the reinforcement learning module;

[0036] The data synchronization unit is used to ensure data consistency between the course learning module and the reinforcement learning module.

[0037] Preferably, the system further comprises a database module and a visualization module;

[0038] The database module is used to store data generated during the molecular design process;

[0039] The visualization module is used to display the molecular design process and results in a graphical manner.

[0040] Compared with the prior art, the present invention has the following advantages and technical effects:

[0041] The present invention proposes a molecular design system that combines curriculum learning and reinforcement learning. Through curriculum learning, the reinforcement learning process is gradually guided, which not only effectively improves the efficiency and diversity of molecule generation, but also avoids premature convergence to the local optimal solution. The application of this system in drug discovery can significantly shorten the R&D cycle, reduce R&D costs, and bring significant economic and social benefits to the field of drug research and development. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0043] Figure 1 A schematic diagram of the system structure of an embodiment of the present invention;

[0044] Figure 2 The figure is a schematic diagram of the application process of an embodiment of the present invention in drug discovery. DETAILED DESCRIPTION

[0045] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0046] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0047] Embodiment 1

[0048] like Figure 1-2 As shown, this embodiment provides a drug molecule design system combining curriculum learning and reinforcement learning, including:

[0049] Data preprocessing module, used for cleaning, standardization and feature extraction of input molecular data;

[0050] The course learning module is connected with the data preprocessing module, which is used to gradually guide the learning process of molecular design according to the preset course strategy, strengthen the learning process, and provide molecular design tasks of different difficulty levels;

[0051] The reinforcement learning module is connected to the course learning module and is used to generate molecules under the guidance of the course learning module; specifically, the molecule generation and optimization are performed based on the tasks provided by the course learning module.

[0052] A molecule generation module, connected to the reinforcement learning module, for generating a target molecular structure according to the output of the reinforcement learning module;

[0053] The evaluation module is connected to the molecular generation module and is used to evaluate the performance of the target molecular structure. The evaluation results are fed back to the reinforcement learning module to optimize the learning process. By gradually increasing the complexity of the molecular design, the reinforcement learning module is guided to avoid premature convergence to the local optimal solution.

[0054] Furthermore, the course learning module includes a course strategy formulation unit and a course progress control unit;

[0055] The course strategy formulation unit is used to formulate course strategies according to preset rules and dynamically adjust course strategies according to the difficulty of molecular design and prior knowledge;

[0056] A course progress control unit is used to control the learning progress of molecular design according to the course strategy.

[0057] Furthermore, the curriculum strategy formulation unit includes the task grading unit and the strategy guidance unit;

[0058] The task classification unit is used to classify the molecular design tasks into multiple levels according to the complexity and design difficulty of the molecules.

[0059] The strategy guidance unit is used to provide tasks of lower difficulty in the initial stage and gradually increase the difficulty of tasks as learning progresses, ensuring that the reinforcement learning module can gradually adapt to more complex molecular designs.

[0060] The course strategy formulation unit gradually guides the learning process of the reinforcement learning module according to the preset course strategy. The specific steps are as follows:

[0061] Initial task selection: Select simple molecular structures as the initial learning task.

[0062] Increasing task difficulty: According to the learning progress, the task difficulty is gradually increased, such as increasing the complexity of the molecular structure.

[0063] Learning progress monitoring: Monitor the learning progress of the reinforcement learning module in real time and adjust the course strategy.

[0064] Furthermore, the reinforcement learning module includes a state representation unit, an action selection unit, a reward function calculation unit, and an optimization learning algorithm unit;

[0065] The state representation unit is used to define the state space in the molecular design process and represent the molecular design problem as a state in reinforcement learning; the state space includes molecular structure characteristics, chemical properties, etc.

[0066] The action selection unit is used to define possible operations, such as adding, deleting, or modifying atoms or bonds in a molecule, and then select the action to generate the molecule based on the current state;

[0067] The reward function calculation unit is used to design the reward function, calculate the performance reward value of the generated molecule according to the reward function, and evaluate the quality of the generated molecular structure through the performance reward value; the reward function comprehensively considers factors such as the structural complexity, biological activity and synthetic feasibility of the molecule to calculate the reward value.

[0068] Further, the molecule generation module includes a molecule structure generation unit and a molecule property prediction unit;

[0069] A molecular structure generation unit, used to generate a molecular structure according to the output of the reinforcement learning module;

[0070] The molecular property prediction unit is used to predict the properties of the generated molecules.

[0071] Furthermore, the molecular structure generation unit includes a rule structure generation unit and a legality checking unit;

[0072] The rule structure generation unit is used to generate a specific molecular structure according to chemical rules based on the output of the reinforcement learning module.

[0073] The legality checking unit is used to ensure that the generated molecular structure complies with chemical legality, such as valence bond rules, stereochemistry, etc.

[0074] The molecular property prediction unit uses a deep learning model to predict the physicochemical properties and biological activities of the generated molecules.

[0075] The molecular generation module generates new molecular structures based on the output of the reinforcement learning module. The specific steps are as follows:

[0076] Structure initialization: Generate the initial molecular structure based on the initial task.

[0077] Structural optimization: Based on the feedback from the reinforcement learning module, the molecular structure is gradually optimized.

[0078] Structural output: Output the final generated molecular structure.

[0079] Further, the evaluation module includes a performance indicator calculation unit and a feedback adjustment unit;

[0080] The performance index calculation unit is used to define and calculate the performance indexes of the generated molecules; the performance indexes include the biological activity, stability, synthetic feasibility, etc. of the molecules.

[0081] The feedback adjustment unit is used to adjust the parameters and learning strategies of the reinforcement learning module according to the performance indicators and optimize the generation process.

[0082] Furthermore, the performance index calculation unit includes multiple submodules, which respectively calculate different performance indexes of the molecule and feed back the comprehensive results to the reinforcement learning module.

[0083] Furthermore, the feedback adjustment unit adopts an adaptive adjustment strategy to dynamically adjust the learning parameters and learning strategy of the reinforcement learning module according to the feedback performance indicators.

[0084] Furthermore, the course learning module and the reinforcement learning module communicate with each other through a data interface, and the data interface includes a data transmission unit and a data synchronization unit;

[0085] A data transmission unit, used for transmitting data between the course learning module and the reinforcement learning module;

[0086] Data synchronization unit, used to ensure data consistency between the course learning module and the reinforcement learning module.

[0087] Furthermore, the system also includes a database module and a visualization module;

[0088] Database module, used to store data generated during the molecular design process;

[0089] Visualization module, used to display the molecular design process and results in a graphical way.

[0090] Furthermore, the database module supports distributed storage, ensuring efficient management and access to large-scale molecular design data.

[0091] Furthermore, the visualization module provides a variety of visualization methods, including molecular structure diagrams, performance index curves, etc., to facilitate users to intuitively understand the molecular design process and results.

[0092] For further optimization, the workflow of the system in this embodiment is as follows:

[0093] Initialization: Set the initial task difficulty and initialize the reinforcement learning model.

[0094] Course learning: Provide corresponding molecular design tasks according to the difficulty of the current task.

[0095] Reinforcement Learning: Based on the current task, the molecular structure is generated through the reinforcement learning module.

[0096] Molecule Generation: The molecule generator generates specific molecules based on the output of the reinforcement learning module.

[0097] Evaluation feedback: The evaluation module evaluates the generated molecules and feeds the results back to the reinforcement learning module.

[0098] Iterative optimization: According to the feedback results, adjust the task difficulty and learning strategy, and perform iterative optimization until the preset optimization goal is achieved.

[0099] Among them, the difficulty of tasks is gradually increased through course learning, making the learning process more efficient and stable.

[0100] Based on reinforcement learning theory, through reward mechanism and iterative optimization, the model can autonomously learn and optimize molecular design strategies.

[0101] In specific implementation, the course learning module can adopt the following strategies:

[0102] Initial task: Design simple small molecules, such as monocyclic compounds.

[0103] Intermediate tasks: Design molecules of medium complexity, such as polycyclic compounds.

[0104] Advanced tasks: Design of complex macromolecules, such as multifunctional compounds.

[0105] The reinforcement learning module can adopt a deep Q network (DQN) algorithm, where the state space includes the feature vector of the molecular structure, the action space includes operations of adding, deleting or modifying atoms and bonds, and the reward function comprehensively considers the stability and biological activity of the molecule.

[0106] This embodiment proposes a drug molecule design system that combines course learning and reinforcement learning, aiming to solve the problems of low efficiency, insufficient diversity, and premature convergence to local optimal solutions in traditional drug molecule design methods. By gradually guiding the reinforcement learning process through course learning, the system can efficiently generate drug candidate molecules with high diversity and excellent performance, while avoiding the limitations of traditional methods. Through the above technical scheme, the present invention can effectively improve the efficiency and diversity of molecule generation, while avoiding premature convergence to local optimal solutions, and has important application value. Specifically including:

[0107] 1. Gradually guide learning to avoid premature convergence;

[0108] This system gradually guides the reinforcement learning process through the course learning module, and gradually increases the difficulty of the molecular design task according to the preset course strategy. The course learning module includes a task grading unit and a strategy guidance unit, which can dynamically adjust the course strategy according to the difficulty of molecular design and prior knowledge. This step-by-step guidance method enables the reinforcement learning module to start from simple tasks and gradually adapt to more complex molecular designs, thus avoiding the problem of premature convergence common in traditional reinforcement learning. For example, in the initial stage, the system can provide simple monocyclic compound design tasks, and gradually transition to the design of polycyclic compounds and multifunctional compounds as learning progresses. This strategy not only improves learning efficiency, but also significantly increases the diversity and complexity of generated molecules.

[0109] 2. Improve the diversity and novelty of molecular generation;

[0110] Through the combination of curriculum learning and reinforcement learning, the system is able to generate molecular structures with high diversity and novelty. The reinforcement learning module dynamically adjusts the generation strategy by defining a reward function and comprehensively considering factors such as the structural complexity, biological activity and synthetic feasibility of the molecule. This enables the system to explore a wider range of chemical space and generate more molecules with potential medicinal value. Compared with traditional rule-based molecular design methods, this system can break through the limitations of fixed rules and generate more innovative molecular structures, thereby significantly improving the success rate of drug discovery.

[0111] 3. Efficient optimization and performance evaluation;

[0112] The evaluation module in the system can conduct a comprehensive performance evaluation of the generated molecules, including biological activity, stability, and synthetic feasibility. The performance index calculation unit calculates different performance indicators of the molecule through multiple submodules, and feeds the results back to the reinforcement learning module. The feedback adjustment unit adopts an adaptive adjustment strategy to dynamically adjust the learning parameters and strategies of the reinforcement learning module according to the feedback performance indicators. This closed-loop optimization mechanism ensures that the generated molecules not only conform to chemical rules in structure, but also have excellent performance in drug performance, significantly improving the screening efficiency and success rate of drug candidate molecules.

[0113] 4. Reduce computing resource consumption;

[0114] Through the gradual guidance of the course learning module, the system can use computing resources more efficiently. In the initial stage, the system converges quickly by processing simple tasks, reducing unnecessary computing overhead. As the difficulty of the task gradually increases, the system gradually adapts to complex molecular design, avoiding the waste of computing resources caused by premature processing of complex tasks in traditional methods. In addition, the system also ensures data consistency between the course learning module and the reinforcement learning module through the data interface, further improving the operating efficiency of the system.

[0115] 5. User-friendly and visual support;

[0116] The system is equipped with a database module and a visualization module, which can store the data generated during the molecular design process and display the design process and results in a graphical manner. The database module supports distributed storage to ensure efficient management and access to large-scale molecular design data. The visualization module provides a variety of visualization methods, including molecular structure diagrams, performance index curves, etc., to facilitate users to intuitively understand the molecular design process and results. This user-friendly design enables researchers to design drug molecules more efficiently, while facilitating monitoring and adjustment of the design process.

[0117] 6. Dynamic adjustment and adaptive learning;

[0118] The curriculum learning module and reinforcement learning module of the system communicate with each other through a data interface, and can dynamically adjust the learning strategy based on the feedback performance indicators. For example, when the generated molecules perform poorly on a certain performance indicator, the system can optimize the generation strategy by adjusting the reward function or learning parameters. This dynamic adjustment mechanism enables the system to adaptively respond to different design tasks and goals, further improving the flexibility and adaptability of the system.

[0119] In summary, the present invention significantly improves the efficiency and diversity of drug molecule design by combining course learning and reinforcement learning. The system can not only avoid premature convergence to the local optimal solution, but also generate drug candidate molecules with high biological activity and synthetic feasibility. In addition, the system's dynamic adjustment and visualization support functions further enhance user experience and design efficiency. Through these technical means, the present invention provides a more efficient, flexible and innovative molecular design solution for the field of drug discovery, which is expected to significantly shorten the R&D cycle, reduce R&D costs, and bring significant economic and social benefits to drug research and development.

[0120] Embodiment 2: Specific application scenario;

[0121] 1. Application in drug discovery

[0122] like Figure 2 As shown, the application process of the molecular design system of this embodiment in drug discovery is as follows:

[0123] Target disease selection: Select the type of disease to be studied, such as cancer, diabetes, etc.

[0124] Target identification: Identify biological targets associated with the disease of interest.

[0125] Molecular design: Use this framework to design molecular structures targeting specific targets.

[0126] Activity testing: The designed molecules are tested for activity in vitro or in vivo.

[0127] Optimization iteration: Based on the test results, optimize the molecular structure until it meets the drug development requirements.

[0128] 2. Specific steps of molecular design

[0129] Initial task setting: Set the initial molecular structure task based on the target characteristics.

[0130] Course learning guidance: Through the course learning module, the reinforcement learning module is gradually guided to generate complex molecular structures.

[0131] Reinforcement learning optimization: Optimize the molecule generation strategy through the reinforcement learning module.

[0132] Molecular generation and evaluation: Generate candidate molecular structures and evaluate them through the evaluation module.

[0133] Technical details:

[0134] The course learning modules adopt the following strategies:

[0135] Difficulty increasing strategy: According to the learning progress, gradually increase the complexity of the molecular structure, such as gradually transitioning from simple small molecules to complex large molecules.

[0136] Feedback adjustment strategy: Dynamically adjust the course difficulty based on the feedback from the reinforcement learning module to ensure a smooth transition in the learning process.

[0137] The reinforcement learning module uses the following algorithms:

[0138] Deep Q-Network (DQN): Use deep neural networks to approximate the Q function and optimize the molecule generation strategy.

[0139] Policy Gradient Method (PG): Improve the efficiency and diversity of molecule generation by directly optimizing the policy function.

[0140] The evaluation module uses the following indicators:

[0141] Structural rationality: Evaluate whether the generated molecular structure conforms to chemical rules.

[0142] Activity prediction: Using machine learning models to predict the biological activity of molecules.

[0143] Diversity assessment: Evaluate the diversity of generated molecular structures to avoid premature convergence to local optimal solutions.

[0144] The above are only preferred specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A drug molecule design system combining curriculum learning and reinforcement learning, characterized in that: include: Data preprocessing module, used for cleaning, standardization and feature extraction of input molecular data; A course learning module, connected to the data preprocessing module, is used to gradually guide the learning process of molecular design according to a preset course strategy, strengthen the learning process, and provide molecular design tasks of different difficulty levels; A reinforcement learning module, connected to the course learning module, for performing molecule generation and optimization under the guidance of the course learning module; A molecule generation module, connected to the reinforcement learning module, for generating a target molecular structure according to an output of the reinforcement learning module; An evaluation module, connected to the molecule generation module, is used to evaluate the performance of the target molecular structure and feed the evaluation result back to the reinforcement learning module to optimize the learning process.

2. The system according to claim 1, characterized in that The course learning module includes a course strategy formulation unit and a course progress control unit; The course strategy formulation unit is used to formulate a course strategy according to preset rules, and dynamically adjust the course strategy according to the difficulty of molecular design and prior knowledge; The course progress control unit is used to control the learning progress of molecular design according to the course strategy.

3. The system according to claim 2, characterized in that The course strategy formulation unit includes a task classification unit and a strategy guidance unit; The task classification unit is used to classify the molecular design tasks into multiple levels according to the complexity and design difficulty of the molecules; The strategy guidance unit is used to provide tasks of relatively low difficulty in the initial stage, and gradually increase the difficulty of the tasks as the learning progresses, so that the reinforcement learning module gradually adapts to complex molecular design.

4. The system according to claim 1, characterized in that The reinforcement learning module includes a state representation unit, an action selection unit, a reward function calculation unit, and an optimization learning algorithm unit; The state representation unit is used to define the state space in the molecular design process, and to represent the molecular design problem as a state in reinforcement learning; the state space includes molecular structure characteristics and chemical properties; The action selection unit is used to define possible operations and then select an action to generate a molecule according to the current state; the operation includes adding, deleting or modifying atoms or bonds in the molecule; The reward function calculation unit is used to design a reward function, calculate the performance reward value of the generated molecule according to the reward function, and evaluate the quality of the generated molecular structure through the performance reward value; the reward function comprehensively considers factors such as the structural complexity, biological activity and synthetic feasibility of the molecule to calculate the reward value.

5. The system according to claim 1, characterized in that The molecule generation module includes a molecule structure generation unit and a molecule property prediction unit; The molecular structure generating unit is used to generate a molecular structure according to the output of the reinforcement learning module; The molecular property prediction unit is used to predict the properties of the generated molecules.

6. The system according to claim 5, characterized in that The molecular structure generation unit includes a rule structure generation unit and a legality checking unit; The rule structure generating unit is used to generate a specific molecular structure according to chemical rules based on the output of the reinforcement learning module; The legality checking unit is used to ensure that the generated molecular structure complies with chemical legality; the molecular structure includes valence bond rules and stereochemistry; The molecular property prediction unit uses a deep learning model to predict the physicochemical properties and biological activities of the generated molecules.

7. The system according to claim 1, characterized in that The evaluation module includes a performance indicator calculation unit and a feedback adjustment unit; The performance index calculation unit is used to define and calculate the performance index of the generated molecule; the performance index includes the biological activity, stability, and synthetic feasibility of the molecule; The feedback adjustment unit is used to adjust the parameters and learning strategy of the reinforcement learning module according to the performance index to optimize the generation process.

8. The system according to claim 7, characterized in that The performance index calculation unit includes a plurality of submodules, which are used to calculate different performance indexes of molecules respectively, and feed back the results to the reinforcement learning module; The feedback adjustment unit adopts an adaptive adjustment strategy to dynamically adjust the learning parameters and learning strategy of the reinforcement learning module according to the feedback performance indicators.

9. The system according to claim 1, characterized in that The course learning module and the reinforcement learning module communicate with each other via a data interface; The data interface includes a data transmission unit and a data synchronization unit; The data transmission unit is used to transmit data between the course learning module and the reinforcement learning module; The data synchronization unit is used to ensure data consistency between the course learning module and the reinforcement learning module.

10. The system according to claim 1, characterized in that The system also includes a database module and a visualization module; The database module is used to store data generated during the molecular design process; The visualization module is used to display the molecular design process and results in a graphical manner.

Citation Information

Patent Citations

  • Passenger IC card data and stop matching method based on bus arrival and departure

    CN109523819A

  • Filter device and communication device

    CN110999082A