A method and system for planning molecular retrosynthetic routes
By modeling multi-step inverse synthesis as a tree search problem, using the end-to-end Transformer model and evolution algorithm to optimize population iteration, pruning technology reduces the search space, solving the problem of inefficiency in multi-step inverse synthesis, and achieving more efficient molecular inverse synthesis path planning.
Patent Information
- Application Number
- CN202411253711.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-09
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-09-09
AI Technical Summary
The existing multi-step molecular inverse synthesis method has invalid nodes in the search space, resulting in inefficiency. The evaluation function depends on the probability output of the Transformer model rather than the principle of molecular synthesis, affecting the population iteration effect.
Multi-step inverse synthesis is modeled as a tree search problem, and single-step inverse synthesis is used using the end-to-end Transformer model to design a coding strategy and build a probability model. The search space is reduced through pruning technology, combined with evolutionary algorithms to optimize population iteration, and use suitable genetic operators and coding methods.
The search efficiency and accuracy of multi-step inverse synthesis are improved, invalid nodes are reduced, and the efficiency and quality of multi-step inverse synthesis path planning is improved.
Smart Images

Figure CN119207637B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of molecular retrosynthesis, and relates to a method and a system for planning a molecular retrosynthesis route. Background Art
[0002] Molecular retrosynthesis is a very important task, especially in drug synthesis and other related fields. Its goal is to identify the retrosynthesis route from the target product to the original reactants. The retrosynthesis process includes single-step retrosynthesis and multi-step retrosynthesis. Single-step retrosynthesis involves decomposing the target product into one or more molecules that can be used to synthesize it. Multi-step retrosynthesis takes a decomposed molecule as the new target product and repeats the single-step retrosynthesis process until a molecule existing in the raw material dataset or that can be obtained through simple synthesis is found.
[0003] In recent years, the Transformer model has become a key technology in the field of artificial intelligence. It has demonstrated excellent performance in various tasks, such as machine translation and computer vision. In the molecular retrosynthesis task, which can be regarded as a machine translation problem, the Transformer architecture has made substantial progress in the single-step retrosynthesis process, bringing an important breakthrough to this field.
[0004] The planning of multi-step retrosynthesis routes has received significant attention. This method usually utilizes search algorithms, such as Monte Carlo tree search and A* algorithm, etc. In this process, the single-step retrosynthesis model plays a crucial role, and its quality greatly affects the effectiveness of multi-step search. The single-step model learns reaction rules from the dataset, and then within the multi-step search framework, iteratively calls the single-step model to infer and identify the most suitable simulation route. This iterative process improves the overall quality and efficiency of multi-step retrosynthesis. Some scholars have also adopted Monte Carlo tree search (MCTS) as the multi-step search algorithm and the Transformer architecture as the single-step retrosynthesis model to plan the retrosynthesis route. This method effectively predicts the feasible routes for different target products. MCTS is a search algorithm known for its powerful search ability and is suitable for complex problems. However, the efficiency of MCTS depends on the complexity of the problem and the number of simulations required. On the other hand, the Transformer architecture performs well on different datasets.
[0005] Generally speaking, after calling the single-step model in the multi-step search process, it is possible to search for the retrosynthesis route and solve the problem of automatic retrosynthesis path planning for organic molecules.
[0006] Currently, there are methods for the pioneering application of evolutionary algorithms (EAs) to the molecular retrosynthesis problem, introducing innovative approaches to this field. Although this method has significantly advanced the application of evolutionary algorithms in molecular synthesis, it has also highlighted areas that urgently need further improvement. First, this method combines a single-step model with a multi-step search algorithm, where the output of the single-step model is discretely encoded, but the genetic operators are continuously encoded. Within the framework of the evolutionary algorithm, this encoding method may not fully utilize the potential efficiency of the algorithm. Second, the individual evaluation function relies on the probability values of the elements within each individual and the similarity between the leaf nodes and the molecules in the raw material database to drive population iteration. Considering that these probabilities are generated based on the output of a Transformer model under previous labeling conditions, this evaluation function may benefit more from the principles of molecular synthesis rather than the output probability values of the Transformer. In addition, the molecular complexity difference between the original molecules in the database and the more complex branched molecules on the leaf nodes often leads to lower similarity scores, which may hinder effective population iteration. Finally, this method defines a large search space within which many nodes correspond to invalid molecular expressions, significantly affecting the efficiency of the search. For example, Patent CN114822703A discloses a molecular retrosynthesis method. First, the target molecule is used as the root node, and then this node is expanded to obtain the second node, and so on. The entire search space is searched through recursive traversal to determine the retrosynthesis path of the target molecule, improving the accuracy of compound molecule retrosynthesis prediction. Summary of the Invention
[0007] To address the deficiencies of the above technologies, the present invention provides a method for planning a molecular retrosynthesis route, including the following steps:
[0008] Step 1: Train an end-to-end Transformer model through an existing dataset, and the Transformer model is used to achieve single-step retrosynthesis;
[0009] Step 2: Design an encoding strategy according to the characteristics of the retrosynthesis problem;
[0010] Step 3: Initialize the population according to the encoding strategy, and the number of individuals is n;
[0011] Step 4: Construct a probability model, where the input of the probability model is the population and the output is a probability model;
[0012] Step 5: Sample the population through the probability model to obtain n new solutions;
[0013] Step 6: Combine the original population and the obtained new solutions, evaluate them through f(x), and obtain the top n solutions with the highest scores. If the last element of the individual is in the raw material library, save the current route;
[0014] Step 7: Construct a probability model for these n solutions, and repeat Step 4 until the population converges or the maximum number of iterations is reached and then stop.
[0015] In Step 1 of the present invention, the data set comes from the data sets USPTO_50K, USPTO_MIT, and Pistachio.
[0016] In the single-step retrosynthesis process of the present invention, the molecular expression is represented by a SMILES expression.
[0017] Preferably, if there are multiple reactants, they are separated by a period ('.'); the placeholder <RX_T> represents the reaction of the T-th type.
[0018] In Step 2 of the present invention, the encoding strategy is that each individual in the population is encoded using two arrays. The first array represents the sorting of the outputs of the single-step models from the root node to the leaf nodes, and the second array represents the possible branching situations that each element in the first array can choose.
[0019] The multi-step retrosynthesis process of the present invention is modeled as a tree search problem in a k-ary tree with q layers; the multi-step retrosynthesis process includes multiple single-step reactions, and the goal is to find a feasible path from the root node to the leaf node; the root node of the tree is the target product, and it is gradually expanded into a k-ary tree through a single-step model; the search tree will be expanded q layers downward, and the leaf nodes of the last layer represent the terminal reactants in the retrosynthesis.
[0020] When the present invention expands the search tree through a single-step model, invalid nodes are removed.
[0021] In Step 4 of the present invention, a probability model is constructed according to the formula in the mutation method:
[0022] Mutation method: Construction of the probability model. For the variable in the i-th dimension, there are k different results, and the number of times this result appears is defined as c i,j , j ∈ [0, k - 1]; where, c i,j represents the number of times j appears in the i-th variable, and is defined as:
[0023]
[0024] where, represents the value of the m-th individual in the i-th variable, and n represents the number of individuals in the population; is an indicator function, and is defined as:
[0025]
[0026] The probability matrix of the t-th generation is defined as:
[0027]
[0028] Among them, τ represents a minimum value, and the probability model is constructed as:
[0029]
[0030] Among them, α represents the attenuation coefficient and is defined as:
[0031]
[0032] Among them, t represents the current iteration number, and t max represents the maximum iteration number; the new solution x is sampled through the probability model P i,j The probability matrix is normalized to ensure that the sum of probabilities in each column is equal to 1, making the matrix suitable as a probability distribution in random sampling; the probability model is defined as:
[0033]
[0034] After the present invention generates a new solution, the new solution will be combined with the current solution, and then their similarity is calculated through the evaluation function f(x), and the top n individuals are selected to enter the next round of iteration. This process will continue until the stopping criterion is met or the maximum iteration number is reached.
[0035] Based on the above method, the present invention also proposes a molecular retrosynthesis route planning system, including: a memory and a processor; a computer program is stored on the memory, and when the computer program is executed by the processor, the above method is implemented.
[0036] The present invention also proposes a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above method is implemented.
[0037] Compared with the prior art, the present invention models the multi-branch molecular retrosynthesis problem and models this problem as a tree search problem. Then, the evolutionary algorithm is used to encode this optimization problem, a suitable encoding strategy is selected, and matching genetic operators are designed, and a certain pruning technique is adopted to reduce the search space, improving the efficiency of the multi-step search algorithm. The invention has the characteristics of being transferable and learnable. The invention innovatively uses the problem information for encoding, innovatively uses genetic operators matching the encoding method, and innovatively uses pruning techniques in this application problem. Description of the Drawings
[0038] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0039] Figure 1 is the schematic diagram of the complete reverse synthesis process of the present invention.
[0040] Figure 2 is the schematic diagram of the single-step model of the present invention.
[0041] Figure 3 is the schematic diagram of the reverse synthesis problem modeling and pruning of the present invention.
[0042] Figure 4 is the flow chart of the evolutionary algorithm of the present invention. Detailed implementation manners
[0043] Combined with the following specific embodiments and drawings, the present invention will be further described in detail. The processes, conditions, experimental methods, etc. for implementing the present invention, except for the specifically mentioned content below, are all common knowledge and well-known common sense in the art, and the present invention has no particularly restricted content.
[0044] 1. Definition of multi-branch reverse synthesis problem
[0045] In the reverse synthesis process, as Figure 1 shown, the conversion from the target product P to the reactant R constitutes a single-step reverse synthesis process. In this process, the molecular expression is represented by the SMILES expression, and this process is similar to the machine translation task. Therefore, transformers commonly used in the machine translation task are used for the single-step reverse synthesis work. As Figure 2 shown, the input of this task is the SMILES expression of the reaction product (for example, <RX_T>NC(=O)c1cn(Cc2c(F)cccc2F)nn1), and the expected output is the SMILES expressions of k reactants. If there are multiple reactants, they are separated by a period ('.') (for example, Fc1cccc(F)c1CBr.[N-]=[N+]=[N-]). The placeholder <RX_T> represents the T-th type of reaction.
[0046] The multi-step reverse synthesis process includes multiple single-step reactions, effectively tracing from the target product P to the intermediate molecule R, and then to the raw material W, as Figure 1 shown. The goal of the multi-step search process is to find a feasible path from the root node to the leaf node. As Figure 3As shown, the root node of the tree is the target product, which is gradually expanded downward into a k-ary tree through a single-step model. Assume that the multi-step retrosynthesis involves q single-step retrosynthesis processes, x = (x 1 , x 2 ,..., x q ), and x i ∈ {0, 1,..., k - 1} represents the x i -th possible output of the i-th single-step retrosynthesis. The search tree will be expanded downward by q layers, and the leaf nodes of the last layer represent the terminal reactants in the retrosynthesis. The retrosynthesis problem is now modeled as a tree search problem in a k-ary tree with q layers.
[0047] When expanding the search tree through the single-step model, not all k SMILES expressions of the model output are valid. Therefore, the invalid nodes will be removed, and this process effectively prunes the search tree, reduces the search space, and ensures that all nodes in the search space are valid, as Figure 3 shown.
[0048] To verify the feasibility of the route, the present invention compares the leaf nodes with the raw material library and calculates the objective function, which is defined as follows:
[0049] f(x) = (-1) * DSC(x i , δ) (1)
[0050] where DSC represents calculating the molecular structure similarity between x i and δ. is the raw material library of organic molecules.
[0051] 2. Evolutionary optimization to solve the tree search problem
[0052] Coding method: Since the retrosynthesis problem is a discrete problem and each node may be a multi-branch problem, the coding method of each individual in the population is carried out using two arrays, which are set as x and y respectively. The dimensions of x and y are the same, where y represents the v-th branch of each node in the search tree, y = (y 1 , y 2 ,..., y q ), and y i ∈ {0, 1,..., v - 1}. The present invention currently solves a 2-branch problem, so v = 2.
[0053] Mutation method: Construction of the probability model. The variable of the i-th dimension may have k different results, and the number of times this result appears is defined as c i,j , j ∈ [0, k - 1]. Where c i,j represents the number of times j appears in the i-th variable and is defined as:
[0054]
[0055] Among them, represents the value of the m-th individual in the i-th variable, and n represents the number of individuals in the population. is an indicator function, which is defined as:
[0056]
[0057] Therefore, the probability matrix of the t-th generation can be defined as:
[0058]
[0059] Among them, τ represents a minimum value, then the probability model can be constructed as:
[0060]
[0061] Among them, α represents the attenuation coefficient, which is defined as:
[0062]
[0063] Among them, t represents the current iteration number, and t max represents the maximum iteration number. Therefore, the new solution x can be sampled through the probability model P i,j The probability matrix is normalized to ensure that the sum of probabilities in each column is equal to 1, making the matrix suitable as the probability distribution in random sampling. The probability model is defined as:
[0064]
[0065] Decoding and evaluation: After generating new solutions, these new solutions will be combined with the current solutions, and then their similarity will be calculated through the evaluation function f(x). The top n individuals are selected to enter the next round of iteration, and this process will continue until the stopping criterion is met or the maximum iteration number is reached.
[0066] The molecular inverse synthesis method proposed by the present invention includes the following steps:
[0067] Step 1: Train an end-to-end transformer model through the existing dataset to achieve single-step inverse synthesis work;
[0068] Step 2: Design an encoding strategy according to the characteristics of the inverse synthesis problem, that is: each individual in the population is encoded using two arrays. The first array represents the sorting of the single-step model outputs from the root node to the leaf nodes, and the second array represents the possible branching situations for each element in the first array;
[0069] Step 3: Initialize the population according to the above coding strategy, with the number of individuals being n;
[0070] Step 4: Construct a probability model according to the formula in the above mutation method. The input of the probability model is the population, and the output is a probability model;
[0071] Step 5: Sample the population through the probability model to obtain n new solutions;
[0072] Step 6: Combine the original population and the obtained new solutions, evaluate them through f(x), and obtain the top n solutions with the highest scores. If the last element of the individual is in the raw material library, save the current route;
[0073] Step 7: Construct a probability model for these n solutions, and repeat Step 4 until the population converges or the maximum number of iterations is reached and stop.
[0074] As Figure 4 shown, during the coding process, both arrays in the individual are searched in the Figure 4 way.
[0075] Generally speaking, the present invention fully overcomes the discreteness and multi-branch problems of the retrosynthesis problem, finds a suitable coding method and matching genetic operators, and finally realizes the search for multi-branch retrosynthesis paths.
[0076] Example 1
[0077] Step 1: The formats of the target molecule and the generated molecule in the existing dataset are as follows, appearing in pairs, which are the input and output of the single-step end-to-end model respectively:
[0078] <RX_1>c1ccc(Cn2ccc3ccccc32)cc1 and ClCc1ccccc1.c1ccc2[nH]ccc2c1;
[0079] <RX_6>CCOc1ccc2ccccc2c1-c1c[nH]c(N)n1 and
[0080] CCOc1ccc2ccccc2c1-c1c[nH]c(NC(C)=O)n1;
[0081] Send a large number of datasets similar to the above characters into the transformer model to realize single-step synthesis work.
[0082] Step 2: According to the binary-branch characteristics of each node, encode the individuals in the population. Assume the individual length is 4 and the output of the single-step model has 10 values. Then the encoding results of the individuals are as follows: [[2,3,1,5],[0,1,0,0]]. The first array represents the sorting of the single-step model outputs from the root node to the leaf node, and the second array represents the possible branch situations that each element in the first array can choose.
[0083] Step 3: Assume the population size is 5. Initialize the population according to the above individual encoding strategy:
[0084] [[2,3,1,5],[0,1,0,0]]
[0085] [[1,3,2,6],[1,1,0,0]]
[0086] [[9,2,1,7],[0,0,0,1]]
[0087] [[6,3,8,4],[1,1,0,1]]
[0088] [[2,5,3,1],[0,0,1,0]]
[0089] Step 4: According to the methods of Formula 2 to Formula 7, use the population initialized in Step 3 as the input to construct a probability model, and use this probability model as the way to generate new solutions in the evolutionary algorithm. The probability model is constructed through the following code:
[0090] import numpy as np
[0091] def update(self,Xs):
[0092] self.t += 1
[0093] gamma = self.cal_gamma(self.t)
[0094] [No,Dim] = Xs.shape
[0095] for d in range(Dim):
[0096] counts = np.bincount(Xs[:,d],minlength = self.M)#counts dimension is M*1
[0097] col_sum = np.sum(counts)+1e-10#add a very small positive number
[0098] self.Prob[:, d] = (1 - gamma) * self.Prob[:, d] + gamma * counts.T / col_sum
[0099] self.Prob[:, d] = np.clip(self.Prob[:, d], 0, 1) # self.Prob's dimension is M * D
[0100] Step 5: Sample the population through the probability model, and the sampling method is as follows:
[0101] def sample(self, N):
[0102] col_probs = self.Prob / np.sum(self.Prob, axis = 0)
[0103] Xs = np.apply_along_axis(lambda p: np.random.choice(self.M, size = N, p = p), axis = 0, arr = col_probs)
[0104] return Xs
[0105] Five new solutions are obtained by sampling, as follows:
[0106] [[3, 1, 2, 5], [1, 0, 0, 0]]
[0107] [[1, 2, 6, 3], [1, 1, 0, 1]]
[0108] [[5, 1, 3, 8], [0, 1, 0, 1]]
[0109] [[2, 7, 4, 6], [1, 0, 0, 1]]
[0110] [[3, 1, 0, 8], [0, 0, 1, 0]]
[0111] Step 6: Combine the original population and the obtained new solutions to get the following 10 individuals
[0112] [[2, 3, 1, 5], [0, 1, 0, 0]]
[0113] [[1, 3, 2, 6], [1, 1, 0, 0]]
[0114] [[9, 2, 1, 7], [0, 0, 0, 1]]
[0115] [[6, 3, 8, 4], [1, 1, 0, 1]]
[0116] [[2,5,3,1],[0,0,1,0]]
[0117] [[3,1,2,5],[1,0,0,0]]
[0118] [[1,2,6,3],[1,1,0,1]]
[0119] [[5,1,3,8],[0,1,0,1]]
[0120] [[2,7,4,6],[1,0,0,1]]
[0121] [[3,1,0,8],[0,0,1,0]]
[0122] These 10 individuals are evaluated by f(x). During the evaluation process, the last molecule of the route is decoded to obtain the molecular SMILES expression, which is then compared with the molecules in the raw material database. The top 5 solutions with the highest scores are:
[0123] [[2,3,1,5],[0,1,0,0]]
[0124] [[1,3,2,6],[1,1,0,0]]
[0125] [[1,2,6,3],[1,1,0,1]]
[0126] [[5,1,3,8],[0,1,0,1]]
[0127] [[3,1,0,8],[0,0,1,0]]
[0128] Among them, [[5,1,3,8],[0,1,0,1]] satisfies the last element of the individual in the raw material library and saves the current route.
[0129] Step 7: Take these 5 solutions as input and repeat step 4 until the population converges or the maximum number of iterations is reached.
[0130] The protection content of the present invention is not limited to the above embodiments. Without departing from the spirit and scope of the inventive concept, changes and advantages that can be thought of by those skilled in the art are included in the present invention and are protected by the attached claims.
Claims
1. A method for planning a molecular retrosynthetic route, characterized in that, The following steps are involved: Step 1: Train an end-to-end transformer model using an existing data set, where the transformer model is used to achieve single-step retrosynthesis; Step 2: The retrosynthesis problem is a discrete problem. A coding strategy is designed according to the characteristics of the retrosynthesis problem. The coding strategy is to encode each individual in the population using two arrays. The first array represents the order of the single-step model output from the root node to the leaf node, and the second array represents the branch selected by each element in the first array. Step 3: Initialize the population according to the encoding strategy, with the number of individuals being n; Step 4: Construct a probability model. The input of the probability model is the population, and the output is a probability model. In step 4, a probability model is constructed according to the formula in the mutation method: Mutation method: construction of probability model. The variable in the i-th dimension has k different results, and the number of times this result appears is defined as c i,j , j ∈ [0, k - 1]; where c i,j represents the number of times j appears in the i-th variable and is defined as: Among them, represents the value of the m-th individual in the i-th variable, and n represents the number of individuals in the population; is an indicator function, which is defined as: The probability matrix of generation t is defined as: Among them, τ represents a minimum value, and the probability model is constructed as follows: Where α represents the attenuation coefficient, which is defined as: where t represents the current iteration number, and t max represents the maximum number of iterations; the new solution x is sampled through the probability model P i,j The probability matrix is normalized to ensure that the sum of probabilities in each column equals 1, making the matrix suitable as a probability distribution in random sampling; the probability model is defined as: Step 5: Sample the population through the probability model to obtain n new solutions; Step 6: Put the original population and the new solution together, evaluate them through f(x), and get the top n solutions with the highest scores. If the last element of the individual is in the raw material library, save the current route. Step 7: Use these n solutions to construct a probability model and repeat step 4 until the population converges or the maximum number of iterations is reached; The multi-step retrosynthesis process is modeled as a tree search problem in a k-ary tree with q layers; the multi-step retrosynthesis process includes multiple single-step reactions, and the goal is to find a feasible path from a root node to a leaf node; the root node of the tree is the target product, and is gradually expanded downward into a k-ary tree through a single-step model; the search tree will expand downward q layers, and the leaf node of the last layer represents the terminal reactant in the retrosynthesis; When expanding the search tree through the single-step model, invalid nodes will be removed, effectively pruning the search tree, reducing the search space, and ensuring that all nodes in the search space are valid.
2. The molecular retrosynthesis route planning method according to claim 1, characterized in that In step 1, the data set comes from the data sets USPTO_50K, USPTO_MIT and Pistachio.
3. The molecular retrosynthesis route planning method according to claim 1, wherein In the single-step retrosynthesis process, the molecular expression is represented by a SMILES expression.
4. The molecular retrosynthetic route planning method according to claim 1, wherein, If there are multiple reactants, they are separated by periods; placeholders<RX_T> Represents the Tth type of reaction.
5. The molecular retrosynthesis route planning method according to claim 1, characterized in that After generating a new solution, the new solution will be combined with the current solution, and then their similarity will be calculated by evaluating the function f(x). The first n individuals will be selected to enter the next round of iteration. This process will continue until the stopping criterion is met or the maximum number of iterations is reached.
6. A molecular retrosynthesis route planning system, characterized in that, include: Memory and processor; The memory stores a computer program, and when the computer program is executed by the processor, the method according to any one of claims 1 to 5 is implemented.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.