A Chemical Retro-Synthesis Analysis Method and System

By using the method of combining Monte Carlo search tree and neural network in the CASP program, the inverse synthesis analysis network is generated, which solves the problem of high demand for external chemical reaction data in the prior art and realizes high-quality synthetic route design.

CN114627980BActive Publication Date: 2025-06-10NANJING UNIV +2
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210335947.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-31
Publication Date
2025-06-10
Estimated Expiration
2042-03-31

AI Technical Summary

Technical Problem

Existing CASP programs have high demand for external chemical reaction data in synthetic route design and are difficult to generate results consistent with expert judgments.

Method used

Using the method of combining Monte Carlo search tree and neural network, the synthesis tree is input to the synthesis tree by training compounds, and the training data of the inverse synthesis analysis network is generated, and the inverse synthesis analysis network is used to train the inverse synthesis analysis network to reduce dependence on external data.

Benefits of technology

It realizes the generation of relatively good synthetic routes while reducing external data requirements, which improves the flexibility and accuracy of synthetic route design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114627980B_ABST
    Figure CN114627980B_ABST
Patent Text Reader

Abstract

The present invention relates to a chemical retrosynthesis analysis method and system, and relates to the field of compound analysis. The method includes: inputting training compounds into a synthesis tree, performing a search using a Monte Carlo search tree to obtain training data for a retrosynthesis analysis network; training the retrosynthesis analysis network using the training data to obtain a trained retrosynthesis analysis network; obtaining a compound to be synthesized and inputting the compound to be synthesized into the trained retrosynthesis analysis network to obtain a synthesis route for the compound to be synthesized. The present invention can more flexibly utilize chemical reaction data from various sources, reduce the demand for external data, and still obtain a synthesis route with relatively good quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of compound analysis, and particularly to a chemical retrosynthesis analysis method and system. Background Art

[0002] A Computer Assisted Synthesis Planning Program (hereinafter referred to as the CASP program) is a type of program that can automatically design a synthesis route for a given organic compound by combining computer algorithms with a chemical database. In such programs, the most important and core technology is a method for calculating the synthesis route of a given compound according to certain user requirements based on data of known reaction templates, data of basic building blocks, and other data or information related to synthesis route design. Early CASP programs mostly used very simple search and sorting algorithms, so they could not obtain results comparable to those of chemists. In recent research and practice, using artificial intelligence methods for synthesis route design has gradually become the mainstream in developing CASP programs. However, in such algorithms, the sorting part is very sensitive to the external data used, so it is usually difficult to give results highly consistent with expert evaluations or laboratory results. Therefore, in the algorithm design idea, more emphasis should be placed on its search ability to make it easy to generate a sufficient number of synthesis routes and provide greater reference value for manual design. In addition, the data-driven nature of artificial intelligence methods determines that high-quality chemical reaction data is the determining factor for improving its performance. High-quality chemical reaction data is still scarce and expensive to date. Therefore, it is necessary to develop and implement a technical solution for designing a synthesis route that can more flexibly utilize chemical reaction data from various sources, reduce the demand for external data, and still obtain relatively good-quality synthesis routes. Summary of the Invention

[0003] The object of the present invention is to provide a chemical retrosynthesis analysis method and system to more flexibly utilize chemical reaction data from various sources, reduce the demand for external data, and still obtain relatively good-quality synthesis routes.

[0004] To achieve the above object, the present invention provides the following solution:

[0005] A chemical retrosynthesis analysis method, comprising:

[0006] Input the training compounds into the synthesis tree, and use the Monte Carlo search tree for searching to obtain the training data for the retrosynthesis analysis network; the training data includes the basic molecules of the training compounds and the synthesis routes of the training compounds starting from the basic molecules; the Monte Carlo search tree is updated according to the action value estimation and the state value estimation;

[0007] Use the training data to train the retrosynthesis analysis network to obtain a trained retrosynthesis analysis network;

[0008] Obtain the compound to be synthesized and input the compound to be synthesized into the trained retrosynthesis analysis network to obtain the synthesis route of the compound to be synthesized.

[0009] Optionally, the step of inputting the training compounds into the synthesis tree and using the Monte Carlo search tree for searching to obtain the training data for the retrosynthesis analysis network specifically includes:

[0010] Use the training compounds as the root nodes of the synthesis tree, and use the Monte Carlo search tree to determine the retrosynthesis template probabilities of the training compounds and the retrosynthesis template probabilities of the intermediates after the decomposition of the training chemicals;

[0011] Calculate the state label of the root node;

[0012] Judge whether the state label is -1 to obtain a first judgment result; if the first judgment result is yes, the construction of the synthesis tree fails, reset all nodes except the nodes with the state label of 1 to -1, and determine the training data according to the leaf nodes and edges of the synthesis tree;

[0013] If the first judgment result is no, then judge whether there is a node in the synthesis tree whose depth reaches the first set node depth to obtain a second judgment result;

[0014] If the second judgment result is no, then judge whether there is a node with the state label of 0 in the synthesis tree to obtain a third judgment result; if the third judgment result is yes, traverse all the leaf nodes with the state label of 0 in the synthesis tree, and add child nodes to the leaf nodes with the state label of 0 according to the retrosynthesis template probabilities of the training compounds and the retrosynthesis template probabilities of the intermediates after the decomposition of the training chemicals; determine the state labels of the child nodes and return to the step of "judging whether the state label is -1 to obtain a first judgment result";

[0015] If the third judgment result is no, the construction of the synthesis tree is successful, reset all node labels to 1 and determine the training data according to the leaf nodes and edges of the synthesis tree;

[0016] If the result of the second judgment is yes, determine whether there is a node with a status flag of 0 in the synthesis tree to obtain a fourth judgment result; if the fourth judgment result is no, the synthesis tree is successfully constructed, all node flags are reset to 1, and training data is determined according to the leaf nodes and edges of the synthesis tree;

[0017] If the fourth judgment result is no, the synthesis tree construction fails. Reset all nodes except those with a status flag of 1 to -1, and determine training data according to the leaf nodes and edges of the synthesis tree.

[0018] Optionally, the expression of the loss function of the retrosynthesis analysis network is:

[0019] L = π T ln p+(z - v) 2 +λ||ω|| 2

[0020] where L is the loss function, ω is the retrosynthesis analysis network parameter, and ||ω|| 2 is the sum of the squares of all parameters; λ is the L2 regularization parameter, and π T is the Monte Carlo policy, p is the neural network policy, v is the node value estimate, and z is the status flag.

[0021] Optionally, the expression of the reverse synthesis template probability is:

[0022]

[0023] where π(A|S 0 ) is the reverse synthesis template probability, s 0 is the root node of the Monte Carlo search tree, a is the reverse synthesis template, N(s 0 , a) is the total number of visits to the edge storing the reverse synthesis template, and τ is the temperature coefficient.

[0024] A chemical retrosynthesis analysis system, comprising:

[0025] A training data acquisition module, configured to input a training compound into a synthesis tree, search using a Monte Carlo search tree, and obtain training data for the retrosynthesis analysis network; the training data includes the basic molecule of the training compound and the training compound synthesis route starting from the basic molecule; the Monte Carlo search tree is updated according to action value estimation and state value estimation;

[0026] A training module, configured to train the retrosynthesis analysis network using the training data to obtain a trained retrosynthesis analysis network;

[0027] The retrosynthesis analysis module is used to obtain the compound to be synthesized and input the compound to be synthesized into the trained retrosynthesis analysis network to obtain the synthesis route of the compound to be synthesized.

[0028] Optionally, the training data acquisition module specifically includes:

[0029] The retrosynthesis template probability determination module is used to take the training compound as the root node of the synthesis tree and use the Monte Carlo search tree to determine the retrosynthesis template probability of the training compound and the retrosynthesis template probability of the intermediate after the decomposition of the training chemical;

[0030] The state label determination module is used to calculate the state label of the root node;

[0031] The first judgment module is used to judge whether the state label is -1 to obtain a first judgment result; if the first judgment result is yes, the synthesis tree construction fails, reset all nodes except the nodes with state label 1 to -1, and determine the training data according to the leaf nodes and edges of the synthesis tree;

[0032] The second judgment module is used to, if the first judgment result is no, judge whether there is a node in the synthesis tree whose depth reaches the first set node depth to obtain a second judgment result;

[0033] The third judgment module is used to, if the second judgment result is no, judge whether there is a node with state label 0 in the synthesis tree to obtain a third judgment result;

[0034] The addition and return module is used to, if the third judgment result is yes, traverse all leaf nodes with state label 0 in the synthesis tree, and add child nodes to the leaf nodes with state label 0 according to the retrosynthesis template probability of the training compound and the retrosynthesis template probability of the intermediate after the decomposition of the training chemical; determine the state label of the child nodes and return to the step "judge whether the state label is -1 to obtain a first judgment result";

[0035] The first reset module is used to, if the third judgment result is no, the synthesis tree construction is successful, reset all node labels to 1 and determine the training data according to the leaf nodes and edges of the synthesis tree;

[0036] The fourth judgment module is used to, if the second judgment result is yes, judge whether there is a node with state label 0 in the synthesis tree to obtain a fourth judgment result; if the fourth judgment result is no, the synthesis tree construction is successful, reset all node labels to 1 and determine the training data according to the leaf nodes and edges of the synthesis tree;

[0037] The second reset module is used to determine that the synthesis tree construction fails if the fourth judgment result is negative, reset all nodes except the nodes with the state flag of 1 to -1, and determine the training data according to the leaf nodes and edges of the synthesis tree.

[0038] Optionally, the expression of the loss function of the retrosynthesis analysis network is:

[0039] L = π T ln p+(z - v) 2 + λ||ω|| 2

[0040] where L is the loss function, ω is the retrosynthesis analysis network parameter, and ||ω|| 2 is the sum of the squares of all parameters; λ is the L2 regularization parameter, and π T is the Monte Carlo policy, p is the neural network policy, v is the node value estimate, and z is the state flag.

[0041] Optionally, the expression of the reverse synthesis template probability is:

[0042]

[0043] where π(A|S 0 ) is the reverse synthesis template probability, s 0 is the root node of the Monte Carlo search tree, a is the reverse synthesis template, and N(s 0 , a) is the total number of visits to the edge storing the reverse synthesis template, and τ is the temperature coefficient.

[0044] According to the specific embodiments provided by the present invention, the following technical effects are disclosed:

[0045] The present invention inputs the training compounds into the synthesis tree, performs search using the Monte Carlo search tree to obtain the training data of the retrosynthesis analysis network; the training data includes the training molecular synthesis routes; uses the training data to train the retrosynthesis analysis network to obtain a trained retrosynthesis analysis network; obtains the compound to be synthesized and inputs the compound to be synthesized into the trained retrosynthesis analysis network to obtain the synthesis route of the compound to be synthesized, thereby realizing more flexible utilization of chemical reaction data from various sources, reducing the demand for external data, and still being able to obtain relatively good-quality synthesis routes. Description of the Drawings

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the following described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0047] Figure 1 Flow chart of the chemical retrosynthesis analysis method provided by the present invention;

[0048] Figure 2 Schematic diagram of the synthesis tree provided by the present invention;

[0049] Figure 3 Flow chart for training the retrosynthesis analysis network;

[0050] Figure 4 Architecture diagram of the retrosynthesis analysis network. Detailed implementation manners

[0051] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0052] The purpose of the present invention is to provide a chemical retrosynthesis analysis method and system to realize more flexible utilization of chemical reaction data from various sources, reduce the demand for external data, and still obtain relatively good-quality synthesis routes.

[0053] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0054] As Figure 1 shown, a chemical retrosynthesis analysis method provided by the present invention includes:

[0055] Step 101: Input the training compounds into the synthesis tree, and use the Monte Carlo search tree for searching to obtain the training data of the retrosynthesis analysis network; the training data includes the basic molecules of the training compounds and the synthesis routes of the training compounds starting from the basic molecules; the Monte Carlo search tree is updated according to the action value estimation and the state value estimation.

[0056] Step 102: Use the training data to train the retrosynthesis analysis network to obtain a trained retrosynthesis analysis network.

[0057] Step 103: Obtain the compound to be synthesized and input the compound to be synthesized into the trained retrosynthesis analysis network to obtain the synthesis route of the compound to be synthesized.

[0058] Step 101 specifically includes:

[0059] Taking the training compound as the root node of the synthesis tree, use Monte Carlo search tree to determine the reverse synthesis template probability of the training compound and the reverse synthesis template probability of the intermediate after decomposition of the training chemical.

[0060] Calculate the state label of the root node.

[0061] Judge whether the state label is -1 to obtain the first judgment result; if the first judgment result is yes, the synthesis tree construction fails, reset all nodes except those with state label 1 to -1, and determine the training data according to the leaf nodes and edges of the synthesis tree.

[0062] If the first judgment result is no, then judge whether there is a node in the synthesis tree whose depth reaches the first set node depth to obtain the second judgment result.

[0063] If the second judgment result is no, then judge whether there is a node with state label 0 in the synthesis tree to obtain the third judgment result; if the third judgment result is yes, traverse all leaf nodes with state label 0 in the synthesis tree, and add child nodes to the leaf nodes with state label 0 according to the reverse synthesis template probability of the training compound and the reverse synthesis template probability of the intermediate after decomposition of the training chemical; determine the state label of the child nodes and return to the step "Judge whether the state label is -1 to obtain the first judgment result".

[0064] If the third judgment result is no, the synthesis tree construction is successful, reset all node labels to 1 and determine the training data according to the leaf nodes and edges of the synthesis tree.

[0065] If the second judgment result is yes, then judge whether there is a node with state label 0 in the synthesis tree to obtain the fourth judgment result; if the fourth judgment result is no, the synthesis tree construction is successful, reset all node labels to 1 and determine the training data according to the leaf nodes and edges of the synthesis tree;

[0066] If the fourth judgment result is yes, the synthesis tree construction fails, reset all nodes except those with state label 1 to -1 and determine the training data according to the leaf nodes and edges of the synthesis tree.

[0067] In practical applications, the expression of the loss function of the reverse synthesis analysis network is:

[0068] L = π T ln p+(z - v) 2 +λ||ω|| 2

[0069] where L is the loss function, ω is the reverse synthesis analysis network parameter, ||ω|| 2is the sum of the squares of all parameters; λ is the L2 regularization parameter, and π T is the Monte Carlo policy, p is the neural network policy, v is the node value estimate, and z is the state label.

[0070] In practical applications, the expression for the reverse synthesis template probability is:

[0071]

[0072] where π(A|S 0 ) is the reverse synthesis template probability, s 0 is the root node of the Monte Carlo search tree, a is the reverse synthesis template, and N(s 0 , a) is the total number of visits to the edge storing the reverse synthesis template, and τ is the temperature coefficient.

[0073] The chemical reverse synthesis analysis system provided by the present invention includes:

[0074] A training data acquisition module, configured to input training compounds into a synthesis tree, search using a Monte Carlo search tree, and obtain training data for the reverse synthesis analysis network; the training data includes the basic molecules of the training compounds and the synthesis routes of the training compounds starting from the basic molecules; the Monte Carlo search tree is updated according to the action value estimate and the state value estimate.

[0075] A training module, configured to train the reverse synthesis analysis network using the training data to obtain a trained reverse synthesis analysis network.

[0076] A reverse synthesis analysis module, configured to obtain a compound to be synthesized and input the compound to be synthesized into the trained reverse synthesis analysis network to obtain the basic molecule and the basic molecule synthesis route.

[0077] Among them, the training data acquisition module specifically includes:

[0078] A reverse synthesis template probability determination module, configured to use the training compound as the root node of the synthesis tree, and use the Monte Carlo search tree to determine the reverse synthesis template probability of the training compound and the reverse synthesis template probability of the intermediate after decomposition of the training chemical;

[0079] A state label determination module, configured to calculate the state label of the root node;

[0080] A first judgment module, configured to judge whether the state label is -1 to obtain a first judgment result; if the first judgment result is yes, the synthesis tree construction fails, reset all nodes except the nodes with a state label of 1 to -1, and determine the training data according to the leaf nodes and edges of the synthesis tree;

[0081] The second judgment module is used to, if the first judgment result is negative, judge whether there is a node in the synthesis tree whose depth reaches the first set node depth, and obtain a second judgment result;

[0082] The third judgment module is used to, if the second judgment result is negative, judge whether there is a node with a status flag of 0 in the synthesis tree, and obtain a third judgment result;

[0083] The addition and return module is used to, if the third judgment result is positive, traverse all leaf nodes with a status flag of 0 in the synthesis tree, and add child nodes to the leaf nodes with a status flag of 0 according to the reverse synthesis template probability of the training compound and the reverse synthesis template probability of the intermediate after the decomposition of the training chemical; determine the status flag of the child nodes and return to the step of "judging whether the status flag is -1 to obtain the first judgment result";

[0084] The first reset module is used to, if the third judgment result is negative, the synthesis tree construction is successful, reset all node flags to 1 and determine the training data according to the leaf nodes and edges of the synthesis tree;

[0085] The fourth judgment module is used to, if the second judgment result is positive, judge whether there is a node with a status flag of 0 in the synthesis tree, and obtain a fourth judgment result; if the fourth judgment result is negative, the synthesis tree construction is successful, reset all node flags to 1 and determine the training data according to the leaf nodes and edges of the synthesis tree;

[0086] The second reset module is used to, if the fourth judgment result is negative, the synthesis tree construction fails, reset all nodes except those with a status flag of 1 to -1 and determine the training data according to the leaf nodes and edges of the synthesis tree. Among them, the expression of the loss function of the reverse synthesis analysis network is:

[0087] L = π T ln p+(z - v) 2 +λ||ω|| 2

[0088] Among them, L is the loss function, ω is the reverse synthesis analysis network parameter, ||ω|| 2 is the sum of the squares of all parameters; λ is the L2 regularization parameter, π T is the Monte Carlo strategy, p is the neural network strategy, v is the node value estimation, and z is the status flag.

[0089] Among them, the expression of the reverse synthesis template probability is:

[0090]

[0091] Among them, π(A|S 0) is the probability of the retrosynthesis template, s 0 is the root node of the Monte Carlo search tree, a is the retrosynthesis template, N(s 0 , a) is the total number of visits to the edge storing the retrosynthesis template, and τ is the temperature coefficient.

[0092] There are two calculation modes in the software: the model construction mode and the model application mode. The model construction mode constructs a model that can be used for synthetic route design in this software through the basic molecular database, chemical reaction template database, chemical reaction template classification database, and other necessary parameters given by the user through the reinforcement learning method; the model application mode uses the model obtained in the model construction mode to design the synthetic route for the compound to be synthesized given by the user.

[0093] The basic molecular database is a database composed of commercial compounds or readily available compounds obtained through a certain source, or can also be a database composed of a series of compounds with simple structures, easy to synthesize or purchase defined by the user.

[0094] The basic molecule is a molecule in the basic molecular database or a molecule that meets the conditions listed in (5).

[0095] The target compound database is a database composed of the products of all reactions extracted from the chemical reaction database obtained by the software from open-source channels.

[0096] The target compound is the compound molecule in the target compound database, which is provided to the software during the model iteration process for calculating the synthetic route. The calculated synthetic route is used to be converted into the training data used in model training.

[0097] Chemical informatics tools refer to open-source software tools that can process standardized chemical information, chemical codes, and chemical databases, including but not limited to RDKit, Indigo, Openbabel, etc.

[0098] A chemical reaction template is extracted from several similar chemical reaction SMARTS expressions through cheminformatics tools, and it is a SMARTS expression encoding a class of chemical reactions. A chemical reaction template represents some chemical reactions with similar structural transformations, and it has a part encoding the structural features of reactants and a part encoding the structural features of products, which are respectively called the reactant template and the product template. Both of these two templates are extracted from the reactants and products of the chemical reactions represented by the chemical reaction template. If the reactant template in the chemical reaction template is extracted from the reactants (products) of the original chemical reaction and the product template is extracted from the products (reactants) of the original chemical reaction, then this chemical reaction template is called a forward (reverse) synthesis template. Through cheminformatics tools, the SMILES expression of a chemical molecule that matches the reactant template in the chemical reaction template can be transformed into the SMILES expression of another molecule, that is, this chemical reaction can be abstractly carried out in the computer. The matching between the chemical reaction template and the chemical molecule is determined by cheminformatics tools. In fact, it is to judge whether there is a substructure in the molecular structure that matches the reactant template in the chemical reaction template. Applying the matching template to a molecule always results in a new molecule, while applying a mismatched template to a molecule always gives an empty result.

[0099] A chemical reaction template database is a database composed of SMARTS expressions encoding chemical reaction templates, which can be user-defined. The software uses cheminformatics tools to extract a simple chemical reaction template database from the chemical reaction database obtained from open-source channels and can be used when the user does not provide this database. It should be noted that there is always an empty template in the chemical reaction template database. Applying this template to any molecule does not give any result (the result is empty), which is for dealing with the situation where all chemical reaction templates in the database cannot match a certain molecule. Unless otherwise specified below, it is default that the chemical reaction template database consists of reverse synthesis templates.

[0100] The chemical reaction template classification database is the general name of three new chemical reaction template databases obtained after a certain classification of the chemical reaction template database. The acquisition method and application method of this database are shown in (4).

[0101] The inverse synthesis strategy, for a certain chemical molecule and a database of chemical reaction templates (composed of retrosynthetic templates), is the probability that the molecule applies each retrosynthetic template in the database. This probability is the probability normalized according to the size of the chemical reaction template database, that is, the sum of the probabilities that any molecule applies each template in the database is 1. In the present invention, there are two inverse synthesis strategies, which have different sources and are used in different positions. The inverse synthesis strategy calculated by the neural network is called the neural network strategy, denoted by the symbol p; the strategy obtained by MCTS is called the MCTS strategy, denoted by the symbol π.

[0102] The MC search tree (fully known as the Monte Carlo search tree) is a tree-like data structure constructed by the MCTS process and generated when calculating the inverse synthesis strategy of a single compound. A tree-like data structure contains two structures: nodes and edges. Nodes are connected to each other by directed edges, that is, an edge has a starting node and a terminating node. The starting node of an edge is called the parent node of the terminating node, and the terminating node of an edge is called the child node of the starting node. In the context of the present invention, the terminating node of an edge may not be unique, but multiple nodes exist as terminating nodes, and these terminating nodes are all regarded as the child nodes of the starting node. A unique root node is defined in the tree-like data structure, and it is not the child node of any other node. Leaf nodes are also defined in the tree-like data structure, and they are not the parent node of any other node. The definition of the tree-like data structure requires that for any node other than the root node, there is exactly one edge with this node as the terminating node. Nodes are denoted by the symbol s, edges are denoted by the symbol a, and (s, a) represents the edge a with the node s as the starting node. The depth d(s) of a node in the tree-like data structure is defined as the number of edges between the root node and this node, and the depth of the root node is generally defined as 0; the depth d(a) of an edge is defined as the depth of its starting node + 1. For example, an edge with the root node as the starting node has a depth of 1, and so on. The nodes and edges of the tree-like data structure can also be used as any abstract data structure to store different types of information.

[0103] The termination state is a special type of leaf node in the MC search tree. Such a leaf node satisfies one of the following three conditions: 1. The depth reaches a certain maximum value, that is, maxMCTSDepth of the maximum search depth of MCTS in (1); 2. The chemical molecule stored in it is a basic molecule; 3. The chemical molecule stored in it is not a basic molecule and cannot be matched with any chemical reaction template.

[0104] In the MC search tree involved in the present invention, the information stored in a node includes: 1. The depth d(s) of the node; the chemical molecule SMILES expression S corresponding to the node; 2. The neural network policy p(S), which is the neural network policy of molecule S calculated by the neural network; 3. The node value estimate v, which represents the difficulty of synthesizing the molecule on this node and is calculated by the neural network with the chemical molecule on the node as the input (non-terminal state), or obtained by determining the terminal state condition (terminal state). In the terminal state, the node value estimate v stored in the node will be an exact value rather than an estimated value, and the model calculation is not used. As long as the node meets the terminal state condition 3, it is considered that v = -1; as long as the node meets the terminal state condition 2, it is considered that v = 1; in other cases, if the node meets the terminal state condition 1, it is considered that v = -1; 4. N(s), which means the total number of times the node has been visited currently. The information stored in an edge includes: 1. The chemical reaction template SMARTS expression A corresponding to the edge; 2. N(s,a), which means the total number of times the edge a starting from node s has been visited currently; 3. q(s,a), which means the edge value estimate of the edge a starting from node s.

[0105] The model is a neural network model constructed by pytorch in the software. The function of the model is to calculate the neural network policy and the node value estimate in the MC search tree according to the input chemical molecule (the molecule stored in this node is exactly the molecule input to the model).

[0106] MCTS (fully known as Monte Carlo Tree Search) is a search algorithm. Each time this search algorithm is executed, it needs to complete multiple simulation processes, and constructs an MC search tree as the simulation progresses. Finally, it gives the solution to the search problem according to the information stored in the search tree. In the present invention, each simulation process includes three stages: the selection stage, the expansion and evaluation stage, and the update stage. The specific solutions for each stage are as described in (1).

[0107] A synthesis tree is a tree - shaped data structure constructed by software for a compound to be synthesized. To distinguish it from the MC search tree, nodes in the synthesis tree are denoted by m, edges are denoted by r, and (m, r) represents the edge r starting from node m. The main difference between the synthesis tree and the MC search tree is that the information stored in each node and edge of the synthesis tree is different. The information stored in each node of the synthesis tree includes: the depth d(m) of the node; the chemical molecule SMILES expression M corresponding to the node; the status flag z of the node. The information stored in each edge of the synthesis tree includes: the depth d(r) of the edge; the chemical reaction template R (which is a retrosynthesis template) executed by the chemical molecule on the starting node connected by the edge. Among them, the status flag z of the nodes in the synthesis tree is determined according to the following rules: if the chemical molecule corresponding to the node is a basic molecule, then z = 1; if the chemical molecule corresponding to the node appears in the synthesis tree for the second time and later, the flag value is z = 2, and the software will not expand this node anymore; if the chemical molecule corresponding to the node is not a basic molecule, but the template selected by the software cannot be executed, or the template selected by the software is an empty template, then z = - 1; in other cases, z = 0.

[0108] When the software designs a synthesis route for a certain chemical molecule, it may obtain a successful or failed result. The way to judge success is: within the maximum depth allowed for constructing the synthesis tree, all the chemical molecules stored on the leaf nodes in the constructed synthesis tree are basic molecules. Any other situation is judged as failure.

[0109] The success rate is the proportion of molecules for which the software can generate a complete synthesis route when designing synthesis routes for a certain number of molecules.

[0110] The “+=” operator is defined as: if x is a variable, then x += 1 is equivalent to x = x + 1, that is, increment the current value of variable x by 1 and then assign it to variable x.

[0111] (1) A method for calculating the retrosynthesis strategy of a compound using the combination of MCTS and neural network

[0112] Input data: the compound to be searched, the model, the chemical reaction template database, the basic molecule database, the chemical reaction template classification database.

[0113] Method parameters: the total number of MCTS simulations MCTSSims, the temperature coefficient temp (written as τ in the formula), the maximum search depth of MCTS maxMCTSDepth, the maximum time limit for a single simulation timeBudget.

[0114] Output result: the MCTS strategy π of the compound to be searched

[0115] Method description:

[0116] In the selection stage, the algorithm needs to construct a path from the root node to a leaf node. A path is also a tree-like data structure and is a subset of all nodes and edges in the MC search tree. Nodes and edges in the path, in addition to satisfying the definition of the tree-like data structure, also additionally satisfy: the edges starting from any node are unique; the leaf nodes in the path must be leaf nodes in the currently constructed MC search tree, unless it does not belong to the current MC search tree. The method for constructing the path is as follows:

[0117] (1) Add the root node s 0 to the path.

[0118] (2) If this is the first-round simulation and the MC search tree only contains the root node at this time, then the path has been constructed (without any edges because there are no edges in the MC search tree at this time). If this is not the first-round simulation, then according to the neural network policy p(S 0 ), where S 0 is the chemical molecule (compound to be searched) on the root node, calculate the chemical reaction template A 0 that should be stored on the edge (s 0 , a 1 ) starting from s 1 according to the following formula:

[0119]

[0120] A 1 = argmax A∈RS {q(s 0 , A) + U(s 0 , A)}

[0121] In the formula, RS is the chemical reaction template database. In the formula, all templates matching s 0 in RS will be calculated, but the number of templates matching s 0 stored on the molecule S 0 can be greatly reduced by the method in (IV); A is a certain retrosynthesis template in the database; p(S 0 , A) is the probability that the retrosynthesis template A is applied in the neural network policy p(S 0 ). Here, the subscripts of each s, S, a, and A represent the depth of the node and edge, and the same applies hereinafter. c puct is a constant in the formula used to control the balance between exploration and exploitation of the MCTS algorithm. In the present invention, c puct = 1 and remains unchanged. U(s 0 , A) is based on the usage situation of the chemical reaction template A in the molecule s 0 in the current MC search tree, for the edge storing this template (s 0, a) An improvement in edge value estimation. q(s 0 , A) is the edge (s where template A is stored 0 , a) The current edge value estimation. If the q value cannot be obtained from the existing MC search tree, it means this edge is visited for the first time. At this time, take q(s 0 , A) = 0. For the edge visited for the first time, N(s 0 , a) value cannot be obtained from the MC search tree either. At this time, take N(s 0 , a) = 0. Calculate the A that makes the value in the curly brackets of the second formula the maximum 1 After that, select the edge corresponding to A 1 (s 0 , a 1 ) and add it to the path;

[0122] (3) Apply A 1 to the molecule S 0 to obtain a set of molecules and create nodes corresponding one-to-one with the molecules in the set and add them to the path as the child nodes of s 0 .

[0123] (4) Starting from d = 1, loop through the following operations ① to ⑤:

[0124] ① Select all the nodes that belong to the MC search tree but are not the leaf nodes of the MC search tree from all the leaf nodes at depth d on the current path, and form a set

[0125] ② Starting from k = 1, traverse each node in the set . For the k-th node in the set , calculate the chemical reaction template that should be stored on the edge starting from according to a similar formula where,

[0126]

[0127]

[0128] Among them, is the probability that the reverse synthesis template A in the neural network policy is applied; is an improvement in the edge value estimation of the edge storing this template based on the usage of the chemical reaction template A in the current MC search tree on the molecule ; is an improvement in the edge value estimation; is the total number of times the node has been visited so far (excluding this time); is the edge to the total number of times visited so far (excluding this time); is the edge storing template A current estimated value of the edge

[0129] Note that at this time, it is still possible to encounter an edge in the case of the first visit, and still take In the formula, all templates in RS that match will be calculated, but the method in (4) can greatly reduce the number of templates matching the molecules stored on ;

[0130] ③ Add the edge storing template to the path, and apply to to obtain the molecule set Construct nodes corresponding one-to-one with the molecule set and add them to the path;

[0131] ④ Go back to ② and continue to process the (k + 1)-th node until all molecules in the set are processed, and then jump out of this loop;

[0132] ⑤ d += 1, then go back to ① and recalculate the set until all leaf nodes on the path that belong to the MC search tree are the original leaf nodes in the MC search tree, and then jump out of this loop;

[0133] In the expansion and evaluation phase, the algorithm analyzes the information stored in all leaf nodes in the path of the algorithm and completes the calculation of the node value estimation. When analyzing each leaf node, it first determines whether the leaf node is in a terminal state. If it is in a terminal state, the exact value of v is calculated according to the conditions satisfied by the terminal state; if it is not in a terminal state, it then determines whether the leaf node is visited for the first time. A leaf node being visited for the first time means that the molecule stored on the leaf node has never had its neural network policy p and node value estimation v calculated, that is, this node does not currently belong to the current MC search tree. For a leaf node visited for the first time, its neural network policy p and node value estimation v are calculated using the model. For a leaf node not visited for the first time, in the current MC search tree, according to the already calculated p, the chemical reaction template corresponding to the component with the highest probability is selected and applied to the molecule stored on the leaf node, and then corresponding new edges and child nodes are added to this leaf node, the MC search tree is expanded, and the v values of these child nodes are evaluated and stored on the nodes. The terminal states among these nodes will have their exact v values calculated according to the conditions satisfied by the terminal state, and the non-terminal states will have their node value estimations calculated using the model. The non-terminal states also additionally calculate the neural network policy p and store p and v on the nodes. The newly added nodes here are also added to the path at the same time, and they have actually become new leaf nodes in the path but will not be further analyzed. When the analysis of all leaf nodes is completed, it enters the update phase. If the total time of a round of simulation reaches the maximum time limit timeBudget during the analysis, the analysis of the remaining leaf nodes is immediately stopped and the update phase is entered immediately, and the node value estimations of the remaining leaf nodes are forced to be recorded as v = 0 (even without considering the terminal state).

[0134] In the update phase, the node value estimations of all leaf nodes on the path are involved in the update. These leaf nodes include the terminal state leaf nodes, the leaf nodes visited for the first time, and the newly added nodes on the path. If the maximum time limit for a single simulation is reached in the expansion and evaluation phase, then the nodes with node value estimations recorded as 0 also participate in the update (they are actually also leaf nodes in the path). The algorithm updates the edge value estimations q(s, a) of all edges visited in this simulation according to the above node value estimation v, and updates the total visit counts N(s) and N(s, a) of all nodes and edges in the path selected in this simulation. The specific update method proceeds according to the following process:

[0135] First, update the visit counts of each node and edge. For the total visit counts of a node s visited for the first time and an edge (s, a) visited for the first time:

[0136] N(s) = 0, N(s, a) = 1

[0137] For the total visit counts of a node s not visited for the first time and an edge (s, a) visited for the first time:

[0138] N(s) += 1, N(s, a) += 1

[0139] For the update of the edge value estimation, there are two ways to choose: Way 1 is called avg, and Way 2 is called min. The difference between the two ways lies in the different methods of calculating the update amount G of the edge value estimation. The calculation of the update amount G starts from the node value estimations v of all the leaf nodes obtained above and is iteratively calculated in the direction from the child nodes to the parent nodes, that is, first calculate the update amount of the edge connecting the leaf node and its parent node in the connection path, then calculate the update amount of the edge connecting the parent node and the parent node of the parent node, and so on, until the update amounts of all the edges connected to the root node are calculated. For the edge (s, a) directly connected to the leaf node, its update amount G(s, a) calculated according to Way 1 is:

[0140] If

[0141] G(s, a) = -1 if

[0142] G(s, a) calculated according to Way 2 is:

[0143] If

[0144] G(s, a) = -1 if

[0145] where s′ represents each leaf node that is the termination node of the edge (s, a), v(s′) represents their node value estimations, and n(s′, a, s) represents the total number of termination nodes of the edge (s, a). This quantity is not necessarily 1 because a compound may not have only one synthesis precursor.

[0146] For the edge (s, a) not directly connected to the leaf node, its update amount G(s, a) is related to the update amounts G(s′, a′) of the edges (s′, a′) with all the termination nodes s′ of this edge as the starting nodes. Calculated according to Way 1 is:

[0147] If

[0148] G(s, a) = -1 if

[0149] G(s, a) calculated according to Way 2 is:

[0150] If

[0151] G(s, a) = -1 if

[0152] After calculating the edge value estimation update amount G(s, a) of all edges on the path, update the edge value estimation q(s, a) according to the following formula:

[0153]

[0154] When this formula is used for the first visited edge, the value of q(s, a) on the right side is not yet defined. At this time, take q(s, a) = 0 on the right side.

[0155] The above is the process of a complete simulation in MCTS. When the algorithm completes the specified number of simulation times (MCTSSims), the simulation process of the algorithm will end, and the solution to the problem will be obtained according to the information stored in the constructed MC search tree. Specifically, from all the edges (s 0 , a) connected to the root node s of the MC search tree, obtain the total number of visits N(s 0 , a) of each edge, and then calculate the probability of applying a certain retrosynthesis template A to the compound S to be searched according to the following formula: 0 0 The probability of applying a certain retrosynthesis template A:

[0156]

[0157] Among them, N(s 0 , a) in the numerator represents the total number of visits of the edge storing template A, and the denominator is the sum of taking the power of the total number of visits of all edges connected to the root node with non-zero total number of visits (edges with zero total number of visits obviously have no impact on the result). τ is called the temperature coefficient in the formula. This parameter controls the difference in the probabilities of applying each chemical reaction template in the MCTS strategy. The larger the temperature coefficient, the closer the probabilities of applying each chemical reaction template are, and the MCTS strategy is closer to a random strategy with equal probabilities of applying each chemical reaction template; the smaller the temperature coefficient, the MCTS strategy is closer to a greedy strategy that only makes the probability of the edge with the most visits tend to 1. After calculating the probabilities corresponding to the templates in each chemical reaction template database, use these probabilities as components to form the final MCTS strategy π. This strategy will be used in the process (2) of constructing the compound synthesis tree.

[0158] (2) A method for using the above-mentioned calculated retrosynthesis strategy to design the optimal synthesis route or multiple synthesis routes of the compound to be synthesized by constructing the synthesis tree of the compound to be synthesized

[0159] Input data: compound to be synthesized, model, chemical reaction template database, basic molecule database

[0160] ​Input parameters: maximum length of the synthesis route depth, maximum ratio of non-decomposable compounds unavailable_ratio_threshold

[0161] Method description:

[0162] The method for the software to design the synthesis route of the compound to be synthesized is to construct the synthesis tree of the compound and then give the specific synthesis route according to the synthesis tree.

[0163] 1. Design the optimal route

[0164] First, add the root node m to the synthesis tree 0 , m 0 stores the molecule M to be synthesized 0 . Here, the subscript of m represents the depth of the node, and the subscript of M represents the depth of the node m where M is located. The same applies hereinafter. Calculate the MCTS strategy π(M 0 ) of the compound M to be synthesized using the method in (1). Add an edge (m 0 , r 0 ) to the synthesis tree, and select the reverse reaction template R 1 corresponding to the component with the highest probability in π(M 0 ) and add it to the edge (m 1 , r 0 ). Here, the subscript of r represents the depth of the edge, and the subscript of R represents the depth of the template on the edge r. The same applies hereinafter. Apply R 1 to M 0 to obtain the molecule set . Here, the superscript of each M represents the serial number of each molecule generated after applying R 1 ; n 1 is the total number of molecules (maximum serial number) generated after applying R 1 to M 0 , and the subscript of n represents the depth of the edge r where the applied template R is located. Finally, add the node as the child node of m 1 to the synthesis tree. The molecule stored on the i-th node is and calculate the state label z of these nodes. If z = -1 appears in these state labels, the construction of the synthesis tree stops immediately, and it is determined that the synthesis design of the target compound has failed; if z = -1 does not appear in all state labels, enter the following loop process:

[0165] (1) Initialize d = 1. Determine whether there is a node in the synthesis tree whose depth has reached the maximum value d 0 (1) Initialize d = 1. Determine whether there is a node in the synthesis tree whose depth has reached the maximum value d max(i.e., the maximum length of the synthesis route depth), that is, the maximum length of the synthesis route depth. If so, break out of the loop; then determine whether there are still leaf nodes with z = 0 in the synthesis tree. If there are no such leaf nodes, break out of the loop and stop constructing the synthesis tree;

[0166] (2) If the loop is not broken, traverse all the leaf nodes in the current synthesis tree with z = 0 and depth d. For the k-th such leaf node Execute the following loop:

[0167] ① Calculate the MCTS strategy of the molecule stored in it Then add an edge to this leaf node The edge stores The reverse reaction template corresponding to the component with the highest probability in

[0168] ② Apply the reverse reaction template to the molecule To obtain a set of molecules i d+1,k Is the serial number of each molecule in the set of molecules obtained by applying the template To the molecule The total number of molecules in the set of molecules is This is also the maximum value of i d+1,k ;

[0169] ③ Add To the node Sub-nodes (These are all the termination nodes of the edge ), the i-th d+1,k Sub-node The molecule stored on it is Then calculate the state flag z of these sub-nodes. If z = -1 appears among these nodes, break out of all loops and stop constructing the synthesis tree. This synthesis design is determined to fail;

[0170] ④ Return to ① and continue to process the next node Until all the leaf nodes with z = 0 in the current synthesis tree are traversed;

[0171] (3) d += 1, then return to (1) until the condition is met and break out of the loop;

[0172] After jumping out of the loop, the software will obtain the synthesized tree that has been constructed. According to the synthesized tree, the software will determine whether the synthesis is successful, and the method has been described in the glossary. If the design is successful, the software will output the molecule M stored on each node in the synthesized tree and the retro-synthesis template R stored on each edge in a certain format as the synthesis route of the compound to be synthesized; if the design fails, the software will also output the molecule M stored on each node in the synthesized tree and the retro-synthesis template R stored on each edge in a certain format, but this does not represent a complete synthesis route at this time.

[0173] A synthesized tree that has been constructed and describes a successful synthesis route design is shown in Figure 2 as follows. Note that in the figure, it is assumed that 2 molecules are generated after each application of the retro-synthesis template for the convenience of drawing. The synthesized tree in general cases will be more complex. The basic molecule symbol in the figure is changed to B to distinguish it from the intermediate of non-basic molecules. d here is not the depth but the maximum length depth of the synthesis route.

[0174] For the above method of designing the synthesis route of the compound to be synthesized, the software uses it as the method for designing the optimal synthesis route of the target compound during the model training stage; during the model application stage, the software uses it as the method for designing the optimal synthesis route of the compound to be synthesized given by the user.

[0175] 2. Design multiple different routes

[0176] During the model application stage, the software can also design multiple synthesis routes for the compound to be synthesized given by the user. This method is mostly the same as the above method for designing the optimal synthesis route, only with differences in selecting R through π. Note that this method is not applied during the model training stage. The specific steps of the method are briefly described as follows:

[0177] First, add a root node m 0 to the synthesized tree. The molecule stored on m 0 is the compound to be synthesized M 0 . Here, the subscript of m represents the depth of the node, and the subscript of M represents the depth of the node where M is located. The same applies hereinafter. Use the method in (1) to calculate the MCTS strategy π(M 0 ) of the compound to be synthesized M 0 . Add an edge (m 0 , r 1 ) to the synthesized tree. Randomly select one component corresponding to the retro-reaction template R 0 from the top K components with the highest probabilities in π(M 1 ) and add it to the edge (m 0 , r 1 ). Here, the subscript of r represents the depth of the edge, and the subscript of R represents the depth of the template where R is located. The same applies hereinafter. Apply R to M 0 ​1 , a set of molecules is obtained Here, the superscript of each M represents the serial number of each molecule generated after applying R 1 . After that, n 1 is the total number of molecules (the maximum serial number) generated after applying R to M 0 . The subscript of n represents the depth of the side r where the applied template R is located. Finally, the node 1 is added as a child node of m to the synthesis tree. The molecule stored on the i-th node 0 is and the state label z of these nodes is calculated. If the proportion of the number of nodes with z = -1 in these state labels among all the nodes in the current synthesis tree exceeds the unavailable_ratio_threshold, the construction of the synthesis tree stops immediately, and at the same time, it is determined that the synthesis design of the target compound has failed; if there is no z = -1 in all the state labels, the following loop process is entered:

[0178] (1) Initialize d = 1. Determine whether there is a node in the synthesis tree whose depth has reached the maximum value d max , that is, the maximum length depth of the synthesis route. If so, jump out of the loop; then determine whether there are still leaf nodes with z = 0 in the synthesis tree. If there are no such leaf nodes, jump out of the loop and stop constructing the synthesis tree;

[0179] (2) If the loop is not jumped out, traverse all the leaf nodes in the current synthesis tree with z = 0 and depth d. For the k-th such leaf node the following steps are executed:

[0180] ① Calculate the MCTS strategy of the molecule stored in it Then add an edge to this leaf node from randomly select one of the reverse reaction templates corresponding to the K components with the highest probability from the and store it on the edge ;

[0181] ② Apply the reverse reaction template to the molecule to obtain a set of molecules i is the serial number of each molecule in the set of molecules obtained after applying the template d+1,k to the molecule . The total number of molecules in the set of molecules is This is also the maximum value of i ; d+1,k

[0182] ③ Add to the node​​ Add sub - node (These are all the end - nodes of the edges .) For the i - th d+1,k sub - node the molecule stored on it is Then calculate the state label z of these sub - nodes. If the proportion of the number of nodes with z = - 1 among all the nodes in the current synthesis tree exceeds the unavailable_ratio_threshold, then break out of all loops and stop constructing the synthesis tree. The current synthesis design is determined to fail;

[0183] ④ Go back to ① and continue to process the next node until all the leaf nodes with z = 0 in the current synthesis tree are traversed;

[0184] (3) d+ = 1, then go back to (1) until the condition is met and then break out of the loop;

[0185] After breaking out of the loop, the software will obtain the constructed synthesis tree. According to the synthesis tree, the software will determine whether the synthesis is successful. If the design is successful, the software will output the molecules M stored on each node in the synthesis tree and the reverse - synthesis templates R stored on each edge in a certain format as the synthesis route of the compound to be synthesized; if the design fails, the software will also output the molecules M stored on each node in the synthesis tree and the reverse - synthesis templates R stored on each edge in a certain format, but this does not represent a complete synthesis route at this time.

[0186] Since the reverse - synthesis templates applied to each molecule in the synthesis tree are random, multiple parallel synthesis - route designs for the same compound to be synthesized can obtain different synthesis routes for the same compound to be synthesized.

[0187] (III) A reinforcement - learning method for model - building mode, used to obtain a model that can design compound synthesis routes in software.

[0188] Method parameters: the total number of model - iteration rounds, the number of target compounds N used for generating data in each round train , the number of target compounds N used for evaluating the model in each round test , the maximum number of rounds numItersForTrainExamplesHistory for storing historical training data. The default values of the parameters are: N train = 40, N test = 20, numItersForTrainExamplesHistory = 10. The total number of model - iteration rounds is set by the user himself.

[0189] In the model construction mode, the software starts from a model with randomly initialized parameters, performs multiple model iteration rounds, and finally obtains a model that can be used to design compound synthesis routes in the software. Each iteration round is divided into three stages: data generation, model training, and model evaluation. The flowchart of the model construction mode is as Figure 3 shown.

[0190] First, the architecture of the model is described, as Figure 4 shown. The input molecule first calculates its 1024-dimensional ECFP4 molecular fingerprint using cheminformatics tools, and then the fingerprint is input into the first fully connected layer with the ReLU activation function. After batch normalization and a Dropout layer with a p value (not the p in Figure 4 here) of 0.3, it enters the second fully connected layer with 512 neurons, with the ReLU activation function. After batch normalization and a Dropout layer with p = 0.3, finally, a LogSoftmax layer with 512 neurons calculates the logarithm ln p of the neural network policy for the current input molecule, and a Tanh layer with 512 neurons calculates the node value estimate v(s) of the current input molecule. p is a vector with a dimension equal to the total number of chemical reaction templates + 1, and the last dimension corresponds to the so-called empty template. The logarithm ln p of p is equivalent to the result of taking the logarithm of each component in p; v is a real number between 0 and 1. Note here that the neural network policy is the result obtained by re-exponentiating the ln p output by the neural network. Taking the logarithm is only an intermediate process in model training, purely for numerical stability considerations. To more reasonably display the function of the model and avoid cumbersome descriptions, in the text except this part, it is still considered that the model outputs the neural network policy p.

[0191] In the data generation stage, the software randomly extracts a certain number of target compounds (software built-in parameter, denoted by N train ) from the target compound database, designs the optimal synthesis route using the method in (ii), generates the synthesis tree for each target compound, and records whether the synthesis design corresponding to the synthesis tree is successful. If the synthesis route corresponding to the synthesis tree is designed successfully, then the status marker z on all nodes in the synthesis tree is reset to z = 1; otherwise, it is all reset to z = -1. If it is at the first model iteration round at this time, then the historical best success rate is the proportion of successfully designed target compounds among all the target compounds for which the synthesis route is designed; if it is the kth model iteration round (k ≥ 2) at this time, then the update of the historical success rate depends on the result of the model evaluation in the previous iteration round. If the model after the previous round of training is accepted in the previous round of model evaluation, then the historical success rate is updated to the weighted average of the success rate of designing the synthesis route of the target compounds in this round and the success rate of designing the synthesis route of the target compounds during the model evaluation in the previous round, that is

[0192]

[0193] where acc best,k represents the historical best success rate of the k-th round of model iteration; acc train,k represents the success rate of designing the synthesis routes of all target compounds in the data generation stage of the k-th round of iteration; acc eval,k-1 represents the success rate of designing the synthesis routes of all target compounds in the model evaluation stage of the (k - 1)-th round of model iteration; N test is the total number of synthesis routes of target compounds to be designed in the model evaluation stage. If the model after the previous round of training is not accepted in the previous round of model evaluation, the historical success rate is updated to the weighted average of the historical success rate of the previous round and the success rate of designing the synthesis routes of all target compounds in the current round of data generation stage, that is:

[0194]

[0195] where acc best,k-1 is the historical best success rate in the previous iteration round, and N cur is the number of consecutive rounds that the model evaluation has not accepted the model after training since the model after training was accepted in the previous model evaluation. Each time the model evaluation accepts the model after training, N cur is reset to 0; each time the model evaluation does not accept the model after training, N cur is incremented by 1.

[0196] To obtain the training dataset of the model, for each synthesis tree, the molecular SMILES expression M, MCTS policy π, and node state label z stored at each node are selected to form a tuple [M, π, z]. These tuples will form the training dataset of the model, where M is used as the input of the neural network after being transformed into a 1024-dimensional ECFP4 molecular fingerprint by cheminformatics tools, π is the label when the model outputs the neural network policy p, and z is the label when the model outputs the node value estimate v. The training dataset generated in each round of model iteration will be saved until the number of model iteration rounds exceeds numItersForTrainExamplesHistory. At this time, only the training datasets generated in the previous numItersForTrainExamplesHistory rounds including the training dataset generated in the current round are saved.

[0197] In the model training stage, the software optimizes the parameters of the model by minimizing the loss function in the following formula according to all the currently saved training datasets:

[0198] L = π T ln p + (z - v) 2 + λ||ω|| 2

[0199] where ω is the parameter of the model, and ||ω|| 2 is the sum of the squares of all parameters; λ is the L2 regularization parameter. In the model training stage, the MCTS strategy is used to improve the neural network strategy, and the success and failure of the synthetic design are used to improve the node value estimation in the MC search tree. Therefore, this method of constructing the model is a reinforcement learning method.

[0200] In the model evaluation stage, the software extracts a certain number of target compounds (the parameter N mentioned above test , the built-in parameter of the software) from the target compound database, and designs the optimal synthetic route using the method in (ii). If the success rate is higher than the historical best success rate acc best,k of this round by a certain value (a small amount ∈, the built-in parameter of the software, ∈ = 0.05), the model after this round of training is accepted, replacing the model used in the generation of this round of data and used for the design of the synthetic route of the target compound in the next round of data generation; otherwise, the model after training is discarded, and the model used in this round is continued to be used for the design of the synthetic route of the target compound in the next round of data generation.

[0201] When the software completes all model iterations set by the total number of model iteration rounds, the software runs to termination, and the user obtains a model that can design the synthetic route of compounds in this software.

[0202] (iv) Method for extracting chemical reaction template classification data from chemical reaction template data and using this data to accelerate the MCTS process in (i)

[0203] This method classifies the templates into three categories according to the number of product templates on the right side of the chemical reaction template: 1. Those with a single product template; 2. Those with two product templates; 3. Those with three or more product templates. When applying a chemical reaction template to a molecule, there may be substructures at multiple different positions in the molecule that match the template. When a chemoinformatics tool performs an operation, it cannot distinguish the substructures at different positions in the molecule, but instead performs an operation on each matching substructure once. If a chemical reaction template contains two product templates and the molecule that matches it contains three matching substructures at different positions, then applying the template to the molecule will generate six new molecules. If there are more matching positions, the number of new molecules will increase rapidly, and the MC search tree will become too complex, resulting in a very slow search speed. Therefore, in the expansion and evaluation stage in (I), for all the templates that match a certain molecule, it is necessary to remove the chemical reaction templates that match too many sites and have too many product templates according to the number of matching positions in the molecule and the number of product templates of the template itself. For the first type of template, it is required that at most three sites in the molecule match the template; for the second type of template, it is required that at most two sites in the molecule match, and one of the two sites must correspond to an intramolecular reaction when applying the rule, otherwise only one site is allowed to match; for the third type, no restrictions are imposed because the number of such templates is very small. After making such restrictions, each node of the MC search tree almost always has at most two child nodes (in the case where the third type of template is not applied), which greatly reduces the complexity of the MC search tree, reduces the total number of nodes and edges, and accelerates the search process. This method classifies the chemical reaction template database through regular expressions to obtain sets of chemical reaction templates containing different numbers of product templates, which are provided for use in (I).

[0204] (V) To expand the coverage of basic molecules, in addition to the basic molecules in the basic molecule database, this method sets other conditions for judging basic molecules. Molecules that meet the conditions but do not belong to the basic molecule database will still be considered basic molecules. These conditions include: 1. If the number of C atoms in a molecule does not exceed 6, it can be considered a simple organic compound and is automatically determined to be a basic molecule; 2. If a molecule contains any metal atoms other than Li, Mg, Cu, Zn, Sn, Pb, and Bi, then without calculating the number of C atoms, it is directly determined to be a basic molecule.

[0205] The CASP program developed according to this solution can, in principle, use chemical reaction template data from any source, requires a very small number of chemical reaction templates, and can maintain a high calculation success rate. This is because this solution is a technical solution based on the reinforcement learning algorithm. In the above method (i), a suitable policy function and reward function (i.e., the model neural network) are effectively utilized to guide the MCTS search process; in the above method (iii), all the synthetic route information generated by the learning program itself can be learned, and the parameters of the model neural network can be continuously improved, thereby acting on the MCTS process in (i) and improving the accuracy of the search. The two methods complement each other and promote each other, enhancing data adaptability, reducing data requirements, and maintaining the calculation success rate through the reinforcement learning method.

[0206] The CASP program developed according to this solution can design a brand-new synthetic route for the target compound that has never been reported in the current literature or patents, provide direction guidance for the synthesis design of new compounds, or design a new synthetic route for the compounds with existing synthetic routes. This is because in this technical solution, only the information about chemical reaction templates is allowed to be learned by the model, and the correlation information of the chemical reactions that generate these templates is not added to the learning process. Therefore, the model will not design the synthetic route according to the chemical reactions in the existing literature or patents, but only use the information learned during training. This is a direct manifestation of the use of the reinforcement learning method in this solution;

[0207] The CASP program developed according to this solution can design different synthetic routes for the same target compound. This is because in the above method (ii) 2, random sampling can be performed from the retrosynthetic strategy, so as to avoid generating duplicate routes as much as possible.

[0208] The present invention separates the MCTS process from the construction of the synthetic route, laying the foundation for improving the performance of the neural network in MCTS through the synthetic route design result; the separation of the MCTS process and the construction of the synthetic route is reflected in the independent construction of the MC search tree and the synthetic tree (parts (i) and (ii)). Only by doing so can a complete MCTS calculation be performed for each intermediate molecule (non-basic molecule) in the synthetic tree (or the synthetic route), so that the retrosynthetic strategy calculation for each intermediate molecule is the most accurate. If the MC search tree and the synthetic tree are combined into one, not every molecule in the tree will experience a complete MCTS process. This is because they are not the root nodes of the tree, and the paths during each simulation may not include them, so the total number of times they are simulated may be less than or even much less than the total number of MCTS simulations, and the number of simulations between nodes is also very unbalanced (some are more and some are less). In this way, it is impossible to obtain the accurate retrosynthetic strategy for each intermediate molecule on the synthetic route.

[0209] The present invention applies a neural network model to the MCTS algorithm in a brand-new way, guides the MCTS search process, and replaces the simulation stage with strong randomness, thereby solving the chemical synthesis design problem, greatly reducing the randomness of the synthesis design, and making the optimal route of the compound reproducible. If the neural network is not used to replace the simulation to calculate the node value estimation, then the calculation of the node value estimation requires random sampling, which easily leads to the MCTS search being too random and ultimately unable to find a suitable synthesis route, so a large number of parallel calculations are required to obtain the expected results. After using the neural network, the node value estimation will be a definite value and will no longer be generated by a random process, so the reproducibility is very good, and generally only one calculation is needed.

[0210] The present invention extracts training data from the synthesis routes self-designed by the software to improve the performance of the neural network used in MCTS, thereby enhancing the search ability of the MCTS algorithm and improving the success rate and quality of the synthesis routes designed by the software.

[0211] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and reference can be made to the description in the method part for the relevant parts.

[0212] In this article, specific examples are used to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. At the same time, for those of ordinary skill in the art, there will be changes in the specific implementation manner and application scope according to the idea of the present invention. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A chemical retrosynthesis analysis method, characterized in that, comprising: Inputting training compounds into a synthesis tree, and using a Monte Carlo search tree for searching to obtain training data for the retrosynthesis analysis network; The training data includes the basic molecules of the training compounds and the synthesis routes of the training compounds starting from the basic molecules; The Monte Carlo search tree is updated according to action value estimation and state value estimation; Using the training data to train the retrosynthesis analysis network to obtain a trained retrosynthesis analysis network; Obtaining a compound to be synthesized and inputting the compound to be synthesized into the trained retrosynthesis analysis network to obtain the synthesis route of the compound to be synthesized; The step of inputting training compounds into a synthesis tree and using a Monte Carlo search tree for searching to obtain training data for the retrosynthesis analysis network specifically includes: Taking the training compound as the root node of the synthesis tree, and using the Monte Carlo search tree to determine the reverse synthesis template probability of the training compound and the reverse synthesis template probability of the intermediate after decomposition of the training chemical; Calculating the state label of the root node; Judging whether the state label is -1 to obtain a first judgment result; if the first judgment result is yes, the synthesis tree construction fails, reset all nodes except the nodes with state label 1 to -1, and determine the training data according to the leaf nodes and edges of the synthesis tree; If the first judgment result is no, then judging whether there is a node in the synthesis tree whose depth reaches a first set node depth to obtain a second judgment result; If the second judgment result is no, then judging whether there is a node with a state label of 0 in the synthesis tree to obtain a third judgment result; if the third judgment result is yes, traverse all leaf nodes with a state label of 0 in the synthesis tree, and add child nodes to the leaf nodes with a state label of 0 according to the reverse synthesis template probability of the training compound and the reverse synthesis template probability of the intermediate after decomposition of the training chemical; determine the state label of the child nodes and return to the step "judging whether the state label is -1 to obtain a first judgment result"; If the third judgment result is no, the synthesis tree construction is successful, reset all node labels to 1 and determine the training data according to the leaf nodes and edges of the synthesis tree; If the second judgment result is yes, then judging whether there is a node with a state label of 0 in the synthesis tree to obtain a fourth judgment result; if the fourth judgment result is no, the synthesis tree construction is successful, reset all node labels to 1 and determine the training data according to the leaf nodes and edges of the synthesis tree; If the fourth judgment result is yes, the synthesis tree construction fails, reset all nodes except the nodes with state label 1 to -1 and determine the training data according to the leaf nodes and edges of the synthesis tree.

2. The chemical retrosynthesis analysis method according to claim 1, characterized in that, The expression of the loss function of the retrosynthesis analysis network is: L = π T ln p+(z - v) 2 + λ||ω|| 2 Among them, L is the loss function, ω is the inverse synthesis analysis network parameter, and ||ω|| 2 is the sum of the squares of all parameters; λ is the L2 regularization parameter, and π T is the Monte Carlo strategy, p is the neural network strategy, v is the node value estimation, and z is the state marker.

3. The chemical retrosynthesis analysis method according to claim 1, characterized in that, The expression of the reverse synthesis template probability is: Among them, π(A|S 0 ) is the reverse synthesis template probability, s 0 is the root node of the Monte Carlo search tree, a is the reverse synthesis template, N(s 0 , a) is the total number of visits to the edge storing the reverse synthesis template, and τ is the temperature coefficient.

4. A chemical retrosynthesis analysis system, characterized in that, comprising: A training data acquisition module, configured to input a training compound into a synthesis tree, and perform a search using a Monte Carlo search tree to obtain training data for an inverse synthesis analysis network; The training data includes a basic molecule of the training compound and a synthesis route of the training compound starting from the basic molecule; The Monte Carlo search tree is updated according to action value estimation and state value estimation; A training module, configured to train the inverse synthesis analysis network using the training data to obtain a trained inverse synthesis analysis network; An inverse synthesis analysis module, configured to obtain a compound to be synthesized and input the compound to be synthesized into the trained inverse synthesis analysis network to obtain a synthesis route of the compound to be synthesized; The training data acquisition module specifically includes: A reverse synthesis template probability determination module, configured to use the training compound as the root node of the synthesis tree, and use the Monte Carlo search tree to determine the reverse synthesis template probability of the training compound and the reverse synthesis template probability of the intermediate after decomposition of the training chemical; A state label determination module, configured to calculate the state label of the root node; A first judgment module, configured to judge whether the state label is -1 to obtain a first judgment result; if the first judgment result is yes, the synthesis tree construction fails, reset all nodes except the node with a state label of 1 to -1, and determine the training data according to the leaf nodes and edges of the synthesis tree; A second judgment module, configured to, if the first judgment result is no, judge whether there is a node in the synthesis tree whose depth reaches a first set node depth to obtain a second judgment result; A third judgment module, configured to, if the second judgment result is no, judge whether there is a node with a state label of 0 in the synthesis tree to obtain a third judgment result; An adding and returning module, configured to, if the third judgment result is yes, traverse all leaf nodes with a state label of 0 in the synthesis tree, and add child nodes to the leaf nodes with a state label of 0 according to the reverse synthesis template probability of the training compound and the reverse synthesis template probability of the intermediate after decomposition of the training chemical; determine the state label of the child nodes and return to the step "judge whether the state label is -1 to obtain a first judgment result"; A first reset module, configured to, if the third judgment result is no, the synthesis tree construction is successful, reset all node labels to 1 and determine the training data according to the leaf nodes and edges of the synthesis tree; A fourth judgment module, configured to, if the second judgment result is yes, judge whether there is a node with a state label of 0 in the synthesis tree to obtain a fourth judgment result; if the fourth judgment result is no, the synthesis tree construction is successful, reset all node labels to 1 and determine the training data according to the leaf nodes and edges of the synthesis tree; A second reset module, configured to, if the fourth judgment result is yes, the synthesis tree construction fails, reset all nodes except the node with a state label of 1 to -1 and determine the training data according to the leaf nodes and edges of the synthesis tree.

5. The chemical inverse synthesis analysis system according to claim 4, wherein, The expression of the loss function of the inverse synthesis analysis network is: L = π T ln p+(z - v) 2 + λ||ω|| 2 Among them, L is the loss function, ω is the inverse synthesis analysis network parameter, and ||ω|| 2 is the sum of the squares of all parameters; λ is the L2 regularization parameter, and π T is the Monte Carlo policy, p is the neural network policy, v is the node value estimation, and z is the state marker.

6. The chemical retrosynthesis analysis system according to claim 4, wherein, the expression of the reverse synthesis template probability is: Among them, π(A|S 0 ) is the reverse synthesis template probability, s 0 is the root node of the Monte Carlo search tree, a is the reverse synthesis template, N(s 0 , a) is the total number of visits to the edge storing the reverse synthesis template, and τ is the temperature coefficient.

Citation Information

Patent Citations

  • Reactant derivation method and reverse synthesis derivation method based on Monte Carlo tree

    CN113782109A