Single-step inverse synthesis reaction prediction method and system

By using the mutual distillation dual Transformer model method in single-step inverse synthesis reaction prediction, the problem that the prior art fails to effectively evaluate the prediction accuracy of each reaction type is solved, and the model is achieved with higher accuracy and more reliable prediction effects on each reaction type.

CN120199355AActive Publication Date: 2025-06-24HUNAN UNIV

Patent Information

Application Number
CN202510677150.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-06-24
Estimated Expiration
2045-05-26

AI Technical Summary

Technical Problem

When evaluating the accuracy of single-step inverse synthesis reaction prediction models, the prior art fails to effectively consider the prediction accuracy under each reaction type, resulting in the model's mastery ability on different reaction types not being fully evaluated.

Method used

The dual Transformer model method based on mutual distillation is used to train the model through temperature sampling parameters, calculate the accuracy on each reaction type, and perform model training through cross entropy loss and distillation loss, and adjust the distillation weight to optimize the model performance.

Benefits of technology

It improves the accuracy of the model on each reaction type, solves the problem that the prediction accuracy is difficult to improve after training of a single Transformer model, and provides a more reliable inverse synthesis reaction prediction tool.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199355A_ABST
    Figure CN120199355A_ABST
Patent Text Reader

Abstract

The invention discloses a single-step inverse synthesis reaction prediction method and system, and the method comprises the steps: obtaining a plurality of pieces of inverse synthesis reaction data, carrying out the preprocessing of all reaction data sets, obtaining ten updated reaction data sets, inputting the preprocessed product molecules into a pre-trained inverse synthesis model, and carrying out the prediction of a single-step inverse synthesis reaction. And obtaining a result of a reactant molecule list corresponding to the product molecule. According to the method, the problem that existing inverse synthesis reaction prediction does not pay attention to differential expression on different reaction types can be solved, and the performance of the model on overall reaction prediction and reaction prediction on partial scarce data reaction types is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of deep learning and mutual distillation, and more specifically, relates to a single-step retrosynthesis reaction prediction method and system. Background Art

[0002] Single-step retrosynthesis aims to efficiently plan the synthesis route of a target compound by strategically decomposing a complex product molecule into simple precursor reactants. Its methods are mainly divided into three categories: template-based, template-free, and semi-template. The template-based method is one of the common single-step retrosynthesis methods. This method extracts reaction templates from a reaction database and learns the mapping relationship between product molecules and reactants. The template-free method attempts to get rid of the dependence on artificially extracted templates and directly models the transformation from reactants to product molecules. Most template-free methods adopt the Transformer architecture and regard single-step retrosynthesis as a language translation task. The semi-template method attempts to combine the advantages of the template-based and template-free methods. First, it locates the reaction center on the product molecule, then breaks the chemical bond to obtain synthons, and finally predicts the reactants based on the synthons. This method to a certain extent simulates the retrosynthetic analysis process of chemists.

[0003] Most of the existing works based on the above three methods mostly consider the overall accuracy of the model for the USPTO-50K dataset under known and unknown reaction types, without discussing the prediction accuracy of the model for each product molecule under each reaction type, which can to a certain extent reflect the model's mastery ability for different reaction types. Summary of the Invention

[0004] In view of the above defects or improvement requirements of the existing work, the present invention provides a single-step retrosynthesis reaction prediction method and system based on mutual distillation, aiming to solve the technical problem that most of the existing work considers the overall accuracy of the model for the USPTO-50K dataset under known and unknown reaction types, without discussing the prediction accuracy of the model for each product molecule under each reaction type.

[0005] To achieve the above object, according to one aspect of the present invention, a single-step retrosynthesis reaction prediction method and system are provided, including the following steps:

[0006] Step 1: Obtain a retrosynthesis reaction dataset, construct the first to tenth reaction data sets according to the reaction type values and process them according to a predetermined process; construct two Transformer models, namely model 1 and model 2, whose parameters are and ; initialize the total number of epochs of model training, learning rate , temperature sampling parameter and and the distillation weight parameters of Model 1 and Model 2 and ;

[0007] Step 2: Sample the reaction data set in Step 1 according to the temperature sampling parameter to obtain the product molecule p1 and the true reactant list r1. Input p1 into Model 1 and Model 2 to obtain the predicted reactant lists R 11 and R 12 ;

[0008] Step 3: Calculate the cross-entropy loss and distillation loss between the reactant lists R 11 and R 12 obtained in Step 2 and the true reactant list r1, and perform backpropagation training on Model 1;

[0009] Step 4: Sample the reaction data set according to the temperature sampling parameter to obtain the product molecule p2 and the true reactant list r2. Input p2 into Model 2 and the trained Model 1 in Step 3 to obtain the predicted reactant lists R 22 and R 21 ;

[0010] Step 5: Calculate the cross-entropy loss and distillation loss between the reactant lists R 22 and R 21 obtained in Step 4 and the true reactant list r2, and perform backpropagation training on Model 2;

[0011] Step 6: In each round of training, repeat Steps 2 to 5 above for iterative training of Model 1 and Model 2; after the current round of training ends, adjust the distillation weight parameters and to change the preference coefficients of the distillation loss and cross-entropy loss in Model 1 and Model 2, and adjust the loss calculation equations of Model 1 and Model 2 in the next round;

[0012] Step 7: When the current training round is greater than the total number of rounds epoch, the training ends. Test the performance of Model 1 and Model 2 on each reaction type to obtain the trained Model 1 and Model 2.

[0013] Preferably, Step 1 is specifically that the Transformer model is mainly divided into an encoder and a decoder. First, define the attention mechanism equation of the Transformer encoder. The attention score of the i-th attention head , where , , , , X is the embedding information of each product molecule in the training set, Q is the sequence matrix, K is the key matrix, V is the value matrix, represents the matrix weight for Q in the i-th attention head, represents the matrix weight for K in the i-th attention head, represents the matrix weight for V in the i-th attention head, d model is the model dimension, represents the dimension of a single attention head, represents for taking the square root, the attention head index i ranges from 1 to 8, softmax represents the normalized exponential function. Then the outputs of all attention heads are concatenated, and the multi-head matrix ,

[0014] where is the weight of the linear layer, h is the number of attention heads, Concat represents the concatenation operation for the matrices in the parentheses, is the abbreviation of the multi-head matrix, and then residual connection and layer normalization are performed, , and the obtained result Z1 is then input into the feed-forward neural network layer for calculation, , where FFN is the abbreviation of the feed-forward neural network, W1, b1, W2, b2 are the weights of the two-layer linear layer, ReLU represents the rectified linear unit function, and residual chain connection and layer normalization are performed again. The output of the i-th layer of the encoder , where the layer number i ranges from 1 to 6, and LayerNorm is layer normalization. After the input is processed by the multi-head self-attention layer and the feed-forward neural network layer of the first encoder layer, the output of the first encoder layer is obtained, which will be used as the input of the next encoder layer or as the final output of the encoder. This process is repeated until the output of the sixth encoder layer is used as the final output encoder_out of the encoder. Subsequently, the masked attention mechanism equation of the Transformer decoder is defined. The attention score of the i-th attention head , where , Y is the word embedding vector of the target sequence at time t and before time t, M is the mask matrix, which is used to mark the results of future time steps as negative infinity, ensuring that at time t, only elements with time less than or equal to t can be attended to during calculation, and elements in the future cannot be attended to. Then the outputs of all attention heads are concatenated, and the multi-head matrix , where is the weight of the linear layer, h is the number of attention heads, and then residual connection and layer normalization are performed. The output of the masked attention layer . Then the cross-attention mechanism equation of the Transformer decoder is defined. The cross-attention score of the i-th head , where , encoder_out is the final output of the encoder; then the outputs of all attention heads are concatenated and passed through a linear layer, followed by residual connection and layer normalization, and then input into a feed-forward neural network layer for calculation, and residual connection and layer normalization are performed again to obtain the output decoder_out1 of the first decoder layer, which is input into the next decoder layer for the same processing until the output of the sixth decoder layer is obtained, which is used as the final output decoder_out of the decoder, and then the output distribution at the current moment is obtained through a linear layer and a softmax function;

[0015] The calculation formula of the loss function in step 3 is: , where denotes the summation operation on the values calculated from the first moment to the r-th moment for t, x is the product molecular representation, denotes all target sequence representations before the t-th moment, is the prediction probability of the correct word by model 1 at the t-th moment, denotes that model 1 gives the correct output according to the current parameters and the target sequence before time step t, denotes the cross-entropy loss; , where denotes the probability distribution of a certain output given by model 2 at time step t, denotes the log probability of a certain output given by model 1 at time step t; , where is the distillation weight of model 1 for the learning degree of the output result of model 2 in reaction type i, and finally the total loss is minimized by optimizing the parameters of model 1,

[0016] The calculation formula of the loss function in step 5 is: , where x is the product molecular representation, denotes all target sequence representations before the t-th moment, is the prediction probability of the correct word by model 2 at the t-th moment, denotes the cross-entropy loss function; , where denotes the probability distribution of a certain output given by model 1 at time step t, denotes the probability distribution of a certain output The logarithmic probability, represents the knowledge distillation loss function; , where is the distillation weight of model 2 The degree of learning for the output result of model 1 in reaction type i, and finally the total loss is minimized represents the parameters for optimizing model 2 , aligning the reactant list predicted by model 2 with the true reactant list and the reactant list predicted by model 1.

[0017] Step 6 is specifically to search space O k which has three update strategies: reducing the distillation weight, keeping the distillation weight unchanged, and increasing the distillation weight. For model 1, the parameters of model 1 and the distillation weight are saved to temporary variables for subsequent restoration. Traverse the update strategies in O k , and under the current update strategy, adjust the distillation weights for all reaction types in . Continue to train for one epoch according to the adjusted distillation weights, then verify model 1, calculate the cross-entropy loss with the true reactant list, and store the results in matrix R; when all update strategies are traversed, for each reaction type of model 1, select the candidate weight corresponding to the minimum verification loss from matrix R as the updated distillation weight; then, traverse all strategies for model 2 and verify, updating the distillation weight in model 2.

[0018] Preferably, step 1 is specifically as follows: First, download reaction data from the open-source dataset USPTO-50K. The reaction data is mainly divided into two columns. One column is the reaction type value of the synthesis reaction, and the other column is the RXN reaction equation. The framework of this equation is like reactant1.reactant2 >> product molecule. The reactants and product molecules are characterized in the Simplified Molecular-Input Line-Entry System (SMILES) format, which is the text representation of the molecule. Then, divide the reaction data into 10 reaction data sets according to the reaction type value. Preprocess the chemical reaction data of each reaction data set. Read the CSV file or TXT file containing the reaction data, extract the SMILES of the reactants and the SMILES of the product molecules, perform a quality check on the SMILES of the reactants and product molecules, and filter out some unqualified data, such as the reactant or product molecule being empty, the SMILES of the reactant or product molecule being invalid, the molecule being too small, or the atomic mapping of the reactant and product molecules being inconsistent. Subsequently, adopt the root alignment mode, randomly select an atom from the product molecule as the root atom, and then adjust the atomic order of the reactants according to the atomic mapping to align them with the product molecule. Repeat multiple times to generate augmented data until the data is amplified to 20 times the original. Then perform word segmentation on the SMILES. The segmentation is achieved through regular expressions. Consider the parts that conform to the following rules as tokens, such as atoms within square brackets, element symbols, special symbols, two-digit % plus numbers, and single digits. Convert the reaction expression into tokens separated by spaces. Finally, save the processed source data and target data to the source data file and target data file respectively; Finally, extract the information of the product molecule and reactant list from the ten reaction data sets to obtain the updated reaction data sets.

[0019] Preferably, the reactant list obtained in step 1 predicts the corresponding multiple reactant molecule SMILES from the SMILES of the augmented product molecules, and the SMILES are connected by ".";

[0020] The retrosynthesis reaction prediction model is implemented based on the Transformer model. Both Transformer models consist of an encoder, a decoder, and an output linear layer;

[0021] Preferably, the encoder is mainly composed of an embedding layer, a multi-head attention layer, and a feed-forward neural network layer. Its specific structure is as follows:

[0022] The embedding layer is a linear network. The input of the embedding layer is the molecular information of a batch of product molecules corresponding to the reaction data set under a certain reaction category, and the output is a set of embedding vectors with a size of Batch_size×Seq_length×512, where Batch_size×Seq_length×512 is the dimension of the latent space;

[0023] The embedding layer includes a word embedding layer and a position encoding layer; The embedding layer includes a word embedding layer and a position encoding layer; The input of the word embedding layer is the molecular information of a batch of molecules corresponding to the reaction data set under a certain reaction type, that is, a matrix of Batch_size×Seq_length×Len, and the output is a set of word embedding vectors of Batch_size×Seq_length×512; The input of the position encoding layer is the molecular information of a batch of molecules corresponding to the reaction data set under a certain reaction type, that is, a matrix of Batch_size×Seq_length×Len, and the output is a set of word embedding vectors of Batch_size×Seq_length×512.

[0024] The specific structure of the feed-forward neural network layer and the processing process of the input are exactly the same as those of the feed-forward neural network layer in the encoder; The decoder layer is composed of a masked multi-head attention layer, a multi-head attention layer and a feed-forward neural network layer, and is stacked 6 layers in total; the input of the decoder layer is the set of word embedding vectors of the result x t at time t, and the output is a set of molecular representation vectors decoder_out; The input of the output linear layer is the output decoder_out of the decoder, and the output is a matrix vector of Batch_size×Seq_length×V, where V is the number of the molecular vocabulary; During the process of predicting reactants, the product molecule SMILES is input into the encoder to obtain the encoder output encoder_out. Starting from the t = 0 moment, the first symbol of the output data is the start symbol, and the rest are masked. The data passes through the decoder and the output linear layer to obtain the probability distribution of the next token. The index of this token is extracted according to the probability and used as known information. The above process is repeated continuously until the index of the end symbol is extracted after the data passes through the decoder and the output linear layer, and the prediction process ends, obtaining the reactant list.

[0025] Specifically, the word embedding layer first uses the torch.nn.Embedding module in the deep learning pytorch framework to increase the dimension of the molecular information of a batch of molecules corresponding to the input product molecular information set, so as to obtain a set of product molecular feature vectors. Subsequently, each element of the embedding vector is multiplied by a predefined scaling factor to obtain a set of scaled feature vectors. The positional encoding layer also processes the molecular information of this batch of molecules. The SinusoidalPositionalEmbedding module in the fairseq framework of the sequence modeling toolkit is used to generate a corresponding set of positional vectors according to the positions of the molecular information. Subsequently, the set of molecular feature vectors and the set of positional vectors are subjected to an element-wise addition operation to obtain a new embedding vector x. The dimension of the vector x remains unchanged. The torch.nn.LayerNorm module in the pytorch framework is used again to perform layer normalization on this intermediate vector set. Finally, the torch.nn.Dropout module in the pytorch framework is used to randomly discard the embedding vector x to output a set of word embedding vectors of size Batch_size×Seq_length×512.

[0026] Preferably, the input of the multi-head self-attention layer is the set of word embedding vectors output by the embedding layer, that is, a matrix of Batch_size×Seq_length×512, and the output is a set of attention vectors of Batch_size×Seq_length×512;

[0027] Specifically, the multi-head attention layer first uses the Linear linear layer in the PyTorch framework to perform a linear transformation on the set of word embedding vectors to obtain an intermediate vector set, and then uses the attention mechanism to process this intermediate vector set to obtain a set of attention vectors. The residual connection is used to add the set of attention vectors to its input residually and the torch.nn.LayerNorm module in the pytorch framework is used for layer normalization to output a vector set of size Batch_size×Seq_length×512.

[0028] Preferably, the input of the feed-forward neural network layer is the set of vectors of size Batch_size×Seq_length×512 output by the multi-head attention layer, and the output is a set of molecular representation vectors of size Batch_size×Seq_length×512;

[0029] Specifically, the feed-forward neural network layer first performs dimensionality increase processing on the vector set using the Linear layer in the PyTorch framework, then uses the rectified linear unit (ReLU) to perform non-linear processing on the vector after dimensionality increase, uses the Linear layer in the PyTorch framework again for dimensionality reduction, and finally uses residual connection to add the attention vector set and its input residually and uses the torch.nn.LayerNorm module in the PyTorch framework for layer normalization to output a molecular representation vector set of size Batch_size×Seq_length×512.

[0030] The encoder layer consists of a multi-head self-attention layer and a feed-forward neural network layer, and is stacked 6 layers in total. The input of the encoder layer is a set of word embedding vectors, and the output is a set of molecular representation vectors encoder_out.

[0031] Specifically, the word embedding vectors obtain the output encoder_out1 of the first encoder layer after passing through the first encoder layer, and then encoder_out1 is input into the second encoder layer to obtain the output encoder_out2 of the second encoder layer. Then, for the obtained output encoder_out2 of the second encoder layer, it is continuously input into the next encoder layer, repeating the above process until it is input into the sixth encoder layer, and the output encoder_out6 of the sixth encoder layer is obtained as the final output of the entire encoder.

[0032] Preferably, the decoder mainly consists of an embedding layer, a masked multi-head attention layer, a multi-head attention layer, and a feed-forward neural network layer;

[0033] The embedding layer is a linear neural network, and its input is the known result x from the start to the current time t t , and outputs a set of word embedding vectors of size Batch_size×Seq_length×512;

[0034] The specific structure of the embedding layer and the processing process of the input are exactly the same as those of the embedding layer module of the encoder.

[0035] The input of the masked multi-head attention layer module is the set of word embedding vectors output by the decoder embedding layer, and the output is a set of masked attention vectors of size Batch_size×Seq_length×512.

[0036] Specifically, the masked multi-head attention layer first linearly transforms the set of word embedding vectors using the Linear layer in the PyTorch framework to obtain an intermediate set of vectors. Then, it processes this intermediate set of vectors using the attention mechanism to ensure that future information is not used during prediction, resulting in an attention set of vectors. The attention set of vectors is added to its input using residual connection and layer normalization is performed using the torch.nn.LayerNorm module in the PyTorch framework to output a set of vectors of size Batch_size×Seq_length×512.

[0037] The input of the multi-head attention layer module is the set of vectors output by the decoder masked multi-head attention layer, and the output is an attention set of vectors of size Batch_size×Seq_length×512.

[0038] Specifically, the multi-head attention layer linearly transforms the molecular representation encoder_out output by the encoder and the set of vectors output by the masked multi-head attention layer using the Linear layer in the PyTorch framework to obtain an intermediate set of vectors. Then, it processes this intermediate set of vectors using the attention mechanism to obtain an attention set of vectors. The attention set of vectors is added to its input using residual connection and layer normalization is performed using the torch.nn.LayerNorm module in the PyTorch framework to output a set of vectors of size Batch_size×Seq_length×512.

[0039] The specific structure of the feed-forward neural network layer and the processing process of the input are exactly the same as those of the feed-forward neural network layer in the encoder.

[0040] The decoder layer consists of a masked multi-head attention layer, a multi-head attention layer, and a feed-forward neural network layer, and is stacked 6 layers in total. The input of the decoder layer is the set of word embedding vectors of the result x t at time t, and the output is the set of molecular representation vectors decoder_out.

[0041] Specifically, the word embedding vectors of x t obtain the output decoder_out1 of the first decoder layer after passing through the first decoder layer. Then, decoder_out1 is input into the second decoder layer to obtain the output decoder_out2 of the second decoder layer. Then, for the obtained output decoder_out2 of the second decoder layer, it is continuously input into the next decoder layer, repeating the above process until it is input into the sixth decoder layer, and the output decoder_out6 of the sixth decoder layer is obtained as the final output decoder_out of the decoder.

[0042] The input of the output linear layer is the output of the decoder decoder_out. The output size maps a latent vector of Batch_size×Seq_length×512 dimensions to a vector of the vocabulary size, and then a probability distribution is obtained through the softmax method.

[0043] During the process of predicting reactants, the product molecule SMILES is input into the encoder to obtain the encoder output encoder_out. Starting from time t = 0, the first symbol of the output data is the start symbol, and the rest are masked. The data passes through the decoder and the output linear layer to obtain the probability distribution of the next token. The index of this token is extracted according to the probability and used as known information. The above process is continuously repeated until the index of the end symbol is extracted after the data passes through the decoder and the output linear layer, and the prediction process ends, obtaining a list of reactants.

[0044] According to another aspect of the present invention, there is provided an inverse synthesis reaction prediction system, including:

[0045] The first module obtains a plurality of synthesis reaction data, each synthesis reaction data has a reaction type value, and all reaction data with reaction type values from 1 to 10 are correspondingly constructed into the first to tenth reaction data sets. The above all reaction data sets are preprocessed to obtain ten updated reaction data sets;

[0046] The second module inputs the product molecule information set in the ten reaction data information sets obtained by the first module into a pre-trained inverse synthesis reaction prediction model to obtain the reactant list results required for generating each product molecule.

[0047] Generally speaking, compared with the prior art through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:

[0048] 1. Since the present invention adopts steps 2 to 6, by performing mutual distillation learning on two models based on Transformer and training the models in the way of temperature sampling for data of different reaction types, the accuracy rate on each reaction type can be calculated, solving the problem in the existing work that the analysis accuracy rate on each reaction type is not considered;

[0049] 2. Since the present invention adopts steps 2 to 6, by constructing two different sampling distributions to train these two models and enabling mutual distillation learning between the two models, the accuracy rate of the models on each reaction type is improved, so the technical problem that it is difficult to further improve the prediction accuracy after training a single Transformer model can be solved;

[0050] 3. The method of the present invention can solve the problem of single-step retrosynthesis reaction prediction, improve the performance and accuracy of the model, and provide more reliable tools and methods for single-step retrosynthesis reaction prediction in the medical field;

[0051] 4. The present invention can be widely applied to various retrosynthesis reaction prediction tasks and improve the performance of model training. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 is the overall flowchart of the single-step retrosynthesis reaction prediction method based on mutual distillation of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0053] The principle of the present invention will be described below in conjunction with the drawings and examples. The provided examples are only for illustrating and explaining the present invention, rather than limiting the present invention.

[0054] The basic idea of the present invention is to use mutual distillation and two Transformer models to achieve retrosynthesis reaction prediction. According to the sequence of the product molecule, the sequence is preprocessed to obtain a sequence representation, and then the sequence representation is input into a pre-trained encoder to obtain a molecular encoding. Finally, the molecular encoding is input into a decoder to generate the text sequence of the reactant list.

[0055] As Figure 1 shown, the present invention provides a single-step retrosynthesis reaction prediction method, including the following steps:

[0056] Step 1: Obtain a retrosynthesis reaction data set, construct the first to tenth reaction data sets according to the reaction type values and process them according to a predetermined process; construct two Transformer models, respectively called Model 1 and Model 2, whose parameters are and ; initialize the total number of epochs of model training, the learning rate , the temperature sampling parameter and as well as the distillation weight parameters and of Model 1 and Model 2;

[0057] Specifically, the molecules obtained in this step are all text data, and the molecular information is stored in text form. In this step, reaction data is first downloaded from the open-source dataset USPTO-50K. The reaction data is mainly divided into two columns. One column is the reaction type value of the synthesis reaction, and the other column is the RXN reaction equation. The framework of this equation is like reactant1.reactant2>>product molecule. The reactants and product molecules are characterized in the Simplified Molecular-Input Line-Entry System (SMILES) format, that is, the text representation of the molecule. Then, according to the reaction type value, the reaction data is divided into 10 reaction data sets. The chemical reaction data of each reaction data set is preprocessed. Read the CSV file or TXT file containing the reaction data, extract the SMILES of the reactants and the SMILES of the product molecules, perform quality checks on the SMILES of the reactants and product molecules, and filter out some unqualified data, such as the reactant or product molecule being empty, the SMILES of the reactant or product molecule being invalid, the molecule being too small, or the atomic mapping of the reactants and product molecules being inconsistent. Subsequently, in the root alignment mode, a random atom is selected from the product molecule as the root atom, and then the atomic order of the reactants is adjusted according to the atomic mapping to align it with the product molecule. Repeat multiple times to generate augmented data until the data is amplified to 20 times the original. Then, perform word segmentation on the SMILES. The segmentation is implemented through regular expressions. The parts that conform to the following rules are regarded as tokens, such as atoms within square brackets, element symbols, special symbols, two-digit % plus numbers, and single digits. Convert the reaction expression into tokens separated by spaces. Finally, save the processed source data and target data into the source data file and target data file respectively; the final inverse synthesis reaction prediction result is to predict the corresponding multiple reactant molecule SMILES for the preprocessed product molecule SMILES, and the SMILES are connected by "."; finally, extract the information of the product molecules and reactant lists for the ten reaction data sets to obtain the corresponding reaction data information sets. The value range of the total number of iterations epoch in this step is between 100 and 200, preferably 150; the learning rate ranges from 0 to 1, preferably 0.0015; the temperature sampling parameter ranges from 0 - 10, preferably 1; the temperature sampling parameter ranges from 0 - 10, preferably 5.

[0058] Specifically, the Transformer model is mainly divided into an encoder and a decoder. First, define the attention mechanism equation of the Transformer encoder. The attention score of the i-th attention head, where , , , , X is the embedding information of each product molecule in the training set, Q is the sequence matrix, K is the key matrix, V is the value matrix, represents the matrix weight for Q in the i-th attention head, represents the matrix weight for K in the i-th attention head, represents the matrix weight for V in the i-th attention head, d model is the model dimension, represents the dimension of a single attention head, represents for taking the square root, the attention head index i ranges from 1 to 8, softmax represents the normalized exponential function. Then the outputs of all attention heads are concatenated, and the multi-head matrix , where is the weight of the linear layer, h is the number of attention heads, Concat represents the concatenation operation for the matrices in the parentheses, is the abbreviation of the multi-head matrix, and then residual connection and layer normalization are performed, , and the obtained result Z1 is then input into the feed-forward neural network layer for calculation, , where FFN is the abbreviation of the feed-forward neural network, W1, b1, W2, b2 are the weights of the two-layer linear layer, ReLU represents the rectified linear unit function, and residual chain connection and layer normalization are performed again. The output of the i-th layer of the encoder , where the layer number i ranges from 1 to 6, and LayerNorm is layer normalization. After the input is processed by the multi-head self-attention layer and the feed-forward neural network layer of the first encoder layer, the output of the first encoder layer is obtained, which will be used as the input of the next encoder layer or as the final output of the encoder. This process is repeated until the output of the sixth encoder layer is used as the final output encoder_out of the encoder. Subsequently, the masked attention mechanism equation of the Transformer decoder is defined. The attention score of the i-th attention head , where , Y is the word embedding vector of the target sequence at time t and before time t, M is the mask matrix, which is used to mark the results of future moments as negative infinity to ensure that at time t, only elements with a time less than or equal to t can be attended to, and elements in the future cannot be attended to. Then the outputs of all attention heads are concatenated, and the multi-head matrix , where is the weight of the linear layer, h is the number of attention heads, and then residual connection and layer normalization are performed. The output of the masked attention layer . Then the cross-attention mechanism equation of the Transformer decoder is defined. The cross-attention score of the i-th head , where , encoder_out is the final output of the encoder; then the outputs of all attention heads are concatenated and passed through a linear layer, followed by residual connection and layer normalization, and then input into a feed-forward neural network layer for calculation, and residual connection and layer normalization are performed again to obtain the output decoder_out1 of the first decoder layer, which is input into the next decoder layer for the same processing until the output of the sixth decoder layer is obtained, which is used as the final output decoder_out of the decoder, and then the output distribution at the current moment is obtained through a linear layer and a softmax function.

[0059] Step 2: Sample the reaction data set in step 1 according to the temperature sampling parameter to obtain the product molecule p1 and the true reactant list r1, and input p1 into Model 1 and Model 2 to obtain the predicted reactant lists R 11 and R 12 ;

[0060] Step 3: Calculate the cross-entropy loss and the distillation loss between the reactant lists R 11 and R 12 obtained in step 2 and the true reactant list r1, and perform backpropagation training on Model 1.

[0061] Specifically, the calculation formula of the loss function in this step is: , where denotes the summation operation for the values calculated from the first moment to the r-th moment of t, x is the product molecule representation, denotes all target sequence representations before the t-th moment, is the predicted probability of the correct word by Model 1 at the t-th moment, denotes the probability that Model 1 gives the correct output according to the current parameters and the target sequence before the time step t, denotes the cross-entropy loss; , where denotes the probability distribution that Model 2 gives a certain output at the time step t, denotes the log probability that Model 1 gives a certain output at the time step t; , where is the distillation weight of Model 1 for the learning degree of the output result of Model 2 in the reaction type i, and finally the total loss is minimized to represent optimizing the parameters of Model 1, so that the reactant list predicted by Model 1 aligns with the true reactant list and the reactant list predicted by Model 2.

[0062] Step 4: Sample the reaction data set according to the temperature sampling parameter Obtain the product molecule p2 and the true reactant list r2, and input p2 into Model 1 and Model 2 to obtain the predicted reactant lists R 21 and R 22 ;

[0063] Step 5: Calculate the cross-entropy loss and the distillation loss between the reactant lists R 21 and R 22 obtained in Step 4 and the true reactant list r2, and perform backpropagation training on Model 2;

[0064] Specifically, the calculation formula of the loss function in this step is: , where x is the product molecule representation, represents all target sequence representations before time t, is the predicted probability of the correct word by Model 2 at time t, represents the cross-entropy loss function; , where represents the probability distribution of a certain output given by Model 1 at time step t, represents the log probability of a certain output given by Model 2 at time step t, represents the knowledge distillation loss function; , where is the distillation weight of Model 2 for the learning degree of the output result of Model 1 in reaction type i. Finally, minimize the total loss represents optimizing the parameters of Model 2 , so that the reactant list predicted by Model 2 aligns with the true reactant list and the reactant list predicted by Model 1.

[0065] Step 6: In each round of training process, repeat Steps 2 to 5 above for iterative training of Model 1 and Model 2; after the current round of training ends, adjust the distillation weight parameters and , change the preference coefficients of the distillation loss and the cross-entropy loss in Model 1 and Model 2, and adjust the loss calculation equations of Model 1 and Model 2 in the next round.

[0066] Specifically, there are three update strategies in the search space O k : reducing the distillation weight, keeping the distillation weight unchanged, and increasing the distillation weight. For Model 1, save the parameters of Model 1 and the distillation weight to temporary variables for subsequent recovery, and traverse O kThe update strategy in, under the current update strategy, for All the distillation weights of the reaction types in are adjusted. According to the adjusted distillation weights, continue to train for one epoch, then verify Model 1, calculate the cross-entropy loss with the true reactant list, and store the results in matrix R; when all update strategies are traversed, for each reaction type of Model 1, select the candidate weight corresponding to the minimum verification loss from matrix R as the updated distillation weight; then, similarly traverse all strategies for Model 2 and verify, and update the distillation weights in Model 2.

[0067] Step 7: When the current training round is greater than the total number of rounds epoch, the training ends. Test the performance of Model 1 and Model 2 on each reaction type to obtain the trained Model 1 and Model 2.

[0068] The inverse synthesis reaction prediction model of the present invention is implemented based on the Transformer model, as Figure 1 shown, each of the two Transformer models includes an encoder, a decoder, and an output linear layer.

[0069] The encoder is mainly composed of an embedding layer, a multi-head attention layer, and a feed-forward neural network layer. Its specific structure is as follows:

[0070] The embedding layer is a linear network. The input of the embedding layer is the molecular information of a batch of product molecules corresponding to the reaction data set under a certain reaction category, and the output is a set of embedding vectors of size Batch_size×Seq_length×512, where Batch_size×Seq_length×512 is the dimension of the latent space;

[0071] The embedding layer includes a word embedding layer and a position encoding layer.

[0072] The input of the word embedding layer is the molecular information of a batch of molecules corresponding to the reaction data set under a certain reaction type, that is, a set of word embedding vectors of size Batch_size×Seq_length×512.

[0073] The input of the position encoding layer is the molecular information of a batch of molecules corresponding to the reaction data set under a certain reaction type, that is, a set of word embedding vectors of size Batch_size×Seq_length×512.

[0074] Specifically, the word embedding layer first uses the torch.nn.Embedding module in the deep learning pytorch framework to increase the dimension of the molecular information of a batch of molecules corresponding to the input product molecular information set, so as to obtain a set of product molecular feature vectors. Subsequently, each element of the embedding vector is multiplied by a predefined scaling factor to obtain a scaled set of feature vectors. The position encoding layer also processes the molecular information of this batch of molecules, and uses the SinusoidalPositionalEmbedding module in the fairseq framework of the sequence modeling toolkit to generate a corresponding set of position vectors according to the position of the molecular information. Subsequently, the set of molecular feature vectors and the set of position vectors are subjected to an element-wise addition operation to obtain a new embedding vector x. The dimension of the vector x remains unchanged. The torch.nn.LayerNorm module in the pytorch framework is used again to perform layer normalization on this intermediate vector set. Finally, the torch.nn.Dropout module in the pytorch framework is used to randomly discard the embedding vector x to output a set of word embedding vectors of size Batch_size×Seq_length×512.

[0075] The input of the multi-head self-attention layer is the set of word embedding vectors output by the embedding layer, that is, a matrix of Batch_size×Seq_length×512, and the output is a set of attention vectors of Batch_size×Seq_length×512;

[0076] Specifically, the multi-head attention layer first uses the Linear linear layer in the PyTorch framework to perform a linear transformation on the set of word embedding vectors to obtain an intermediate vector set, and then uses the attention mechanism to process this intermediate vector set to obtain a set of attention vectors. The residual connection is used to add the set of attention vectors to its input residually and the torch.nn.LayerNorm module in the pytorch framework is used for layer normalization to output a vector set of size Batch_size×Seq_length×512.

[0077] The input of the feed-forward neural network layer is the set of vectors of size Batch_size×Seq_length×512 output by the multi-head attention layer, and the output is a set of molecular representation vectors of size Batch_size×Seq_length×512.

[0078] Specifically, the feed-forward neural network layer first performs dimensionality increase processing on the vector set using the Linear layer in the PyTorch framework, then uses the rectified linear unit (ReLU) to perform non-linear processing on the vector after dimensionality increase, uses the Linear layer in the PyTorch framework again for dimensionality reduction, and finally uses residual connection to add the attention vector set and its input residually and uses the torch.nn.LayerNorm module in the PyTorch framework for layer normalization to output a molecular representation vector set of size Batch_size×Seq_length×512.

[0079] The encoder layer consists of a multi-head self-attention layer and a feed-forward neural network layer, and is stacked 6 layers in total. The input of the encoder layer is a set of word embedding vectors, and the output is a set of molecular representation vectors encoder_out.

[0080] Specifically, after passing through the first encoder layer, the word embedding vectors obtain the output encoder_out1 of the first encoder layer, and then encoder_out1 is input into the second encoder layer to obtain the output encoder_out2 of the second encoder layer. Then, for the obtained output of the second encoder layer, it is continuously input into the next encoder layer, repeating the above process until it is input into the sixth encoder layer, and the output encoder_out6 of the sixth encoder layer is obtained as the final output of the entire encoder.

[0081] The decoder mainly consists of an embedding layer, a masked multi-head attention layer, a multi-head attention layer, and a feed-forward neural network layer.

[0082] The embedding layer is a linear neural network, and its input is the known result x from the start to the current time t t , and outputs a set of word embedding vectors of size Batch_size×Seq_length×512.

[0083] The specific structure of the embedding layer and the processing process of the input are exactly the same as those of the embedding layer module of the encoder;

[0084] The input of the masked multi-head attention layer module is the set of word embedding vectors output by the decoder embedding layer, and the output is a set of masked attention vectors of size Batch_size×Seq_length×512.

[0085] Specifically, the masked multi-head attention layer first performs a linear transformation on the set of word embedding vectors using the Linear layer in the PyTorch framework to obtain a set of intermediate vectors. Then, it uses the attention mechanism to process this set of intermediate vectors to ensure that future information is not used during prediction, resulting in a set of attention vectors. The set of attention vectors is added to its input using residual connection and layer normalization is performed using the torch.nn.LayerNorm module in the PyTorch framework to output a set of vectors of size Batch_size×Seq_length×512.

[0086] The input of the multi-head attention layer module is the set of vectors output by the decoder masked multi-head attention layer, and the output is a set of attention vectors of size Batch_size×Seq_length×512.

[0087] Specifically, the multi-head attention layer performs a linear transformation on the molecular representation encoder_out output by the encoder and the set of vectors output by the masked multi-head attention layer using the Linear layer in the PyTorch framework to obtain a set of intermediate vectors. Then, it uses the attention mechanism to process this set of intermediate vectors to obtain a set of attention vectors. The set of attention vectors is added to its input using residual connection and layer normalization is performed using the torch.nn.LayerNorm module in the PyTorch framework to output a set of vectors of size Batch_size×Seq_length×512.

[0088] The specific structure of the feed-forward neural network layer and the process of processing the input are exactly the same as those of the feed-forward neural network layer in the encoder;

[0089] The decoder layer consists of a masked multi-head attention layer, a multi-head attention layer, and a feed-forward neural network layer, and is stacked 6 layers. The input of the decoder layer is the set of word embedding vectors of the result x t at time t, and the output is the set of molecular representation vectors decoder_out.

[0090] Specifically, the word embedding vectors of x t obtain the output decoder_out1 of the first decoder layer after passing through the first decoder layer. Then, decoder_out1 is input into the second decoder layer to obtain the output decoder_out2 of the second decoder layer. Then, for the obtained output decoder_out2 of the second decoder layer, it is continuously input into the next decoder layer, repeating the above process until it is input into the sixth decoder layer, and the output decoder_out6 of the sixth decoder layer is obtained as the final output decoder_out of the decoder.

[0091] The input of the output linear layer is the output of the decoder decoder_out. The output size maps a latent vector of Batch_size×Seq_length×512 dimensions to a vector of the vocabulary size, and then a probability distribution is obtained through the Softmax method.

[0092] During the process of predicting reactants, the product molecule SMILES is input into the encoder to obtain the encoder output encoder_out. Starting from time t = 0, the first symbol of the output data is the start symbol, and the rest are masked. The data passes through the decoder and the output linear layer to obtain the probability distribution of the next token. The index of this token is extracted according to the probability and used as known information. The above process is repeated continuously until the index of the end symbol is extracted after the data passes through the decoder and the output linear layer, and the prediction process ends to obtain the reactant list.

[0093] Test results

[0094] The test environment of the present invention: Under the Ubuntu 20.04 operating system, the CPU is an Intel(R) Xeon(R) Gold6330 CPU @ 2.00GHz, the GPU is 1 NVIDIA RTX 3090 24GB, and the algorithm of the present invention is implemented using Python 3.8 programming.

[0095] To illustrate the effectiveness of the method of the present invention and the improvement of the prediction effect for retro-synthesis reactions, it is tested on a test set from the dataset USPTO-50K. The test results obtained by the present invention are compared with current advanced methods, and the evaluation results are shown in Table 1.

[0096] According to the test results on the test set from the dataset USPTO-50K recorded in Table 1, it can be seen that the single-step retro-synthesis reaction prediction method based on mutual distillation proposed in the present invention is superior to existing methods in the four retro-synthesis prediction indicators of top-1, top-3, top-5, and top-10.

[0097] Table 1 Comparison of retro-synthesis reaction prediction results

[0098]

[0099] Although the above is a description of the specific implementation process of the present invention in combination with the drawings, the present invention is not limited to the above specific solutions. The above examples and descriptions are only used to illustrate the principle of the present invention. Within the spirit and principle of the present invention, any modifications, equivalent replacements, or improvements made should fall within the protection scope of the present invention.

Claims

1. A method for predicting a single-step retro-synthesis reaction, characterized in that, including the following steps, Step 1: Obtain the retro-synthesis reaction dataset, construct the first to tenth reaction data sets according to the reaction type values and process them according to a predetermined process; construct two Transformer models called Model 1 and Model 2 respectively, and their parameters are and ; Initialize the total number of epochs of model training, learning rate , temperature sampling parameter and as well as the distillation weight parameters of Model 1 and Model 2 and ; Step 2: Sampling the reaction data set in Step 1 according to the temperature sampling parameter to obtain the product molecule p1 and the true reactant list r1, and input p1 into Model 1 and Model 2 to obtain the predicted reactant lists R 11 and R 12 ; Step 3: Take the reactant list R obtained in Step 2 11 and R 12 Calculate the cross-entropy loss and the distillation loss with the true reactant list r1, and perform backpropagation training on Model 1; Step 4: Sample the reaction data set according to the temperature sampling parameter to obtain the product molecule p2 and the true reactant list r2, and input p2 into Model 2 and the trained Model 1 in Step 3 to obtain the predicted reactant lists R 22 and R 21 ; Step 5: Take the reactant list R obtained in Step 4 21 and R 22 Calculate the cross-entropy loss and the distillation loss with the true reactant list r2, and perform backpropagation training on Model 2; Step 6: During each round of training, repeat the above Steps 2 to 5 to iteratively train Model 1 and Model 2; after the current round of training ends, adjust the distillation weight parameters and to change the preference coefficients of the distillation loss and the cross-entropy loss in Model 1 and Model 2, and adjust the loss calculation equations for the next round of Model 1 and Model 2; Step 7: When the current training round is greater than the total number of rounds epoch, the training ends; test the performance of Model 1 and Model 2 on each reaction type to obtain the trained Model 1 and Model 2.

2. The single-step retrosynthesis reaction prediction method according to claim 1, characterized in that, Step 1 is specifically as follows. The Transformer model is mainly divided into an encoder and a decoder. First, define the attention mechanism equation of the Transformer encoder. The attention score of the i-th attention head , where , , , , X is the embedding information of each product molecule in the training set, Q is the sequence matrix, K is the key matrix, V is the value matrix, represents the matrix weight taken for Q in the i-th attention head, represents the matrix weight taken for K in the i-th attention head, represents the matrix weight taken for V in the i-th attention head, d model is the model dimension, represents the dimension of a single attention head, represents taking the square root. The attention head index i ranges from 1 to 8, and softmax represents the normalized exponential function. Then, concatenate the outputs of all attention heads. The multi-head matrix , where is the weight of the linear layer, h is the number of attention heads, and Concat represents the concatenation operation on the matrices in the parentheses, is the abbreviation of the multi-head matrix. Subsequently, perform residual connection and layer normalization, , and input the obtained result Z1 into the feed-forward neural network layer for calculation, , where FFN is the abbreviation of the feed-forward neural network, W1, b1, W2, and b2 are the weights of two linear layers, ReLU represents the rectified linear unit function, and perform residual chain connection and layer normalization again. The output of the i-th layer of the encoder , where the layer number i ranges from 1 to 6, and LayerNorm is layer normalization. After the input is processed by the multi-head self-attention layer and the feed-forward neural network layer of the first encoder layer, the output of the first encoder layer is obtained, which will be used as the input of the next encoder layer or as the final output of the encoder. Repeat this process until the output of the sixth encoder layer is used as the final output encoder_out of the encoder. Subsequently, define the masked attention mechanism equation of the Transformer decoder. The attention score of the i-th attention head , where , Y is the word embedding vector of the target sequence at time t and before time t, M is the mask matrix, which is used to mark the results of future time steps as negative infinity to ensure that at time t, only elements with a time less than or equal to t can be attended to, and elements in the future cannot be attended to. Then, concatenate the outputs of all attention heads. The multi-head matrix , where is the weight of the linear layer, h is the number of attention heads, followed by residual connection and layer normalization, and the masked attention layer outputs . Then, define the cross-attention mechanism equation of the Transformer decoder. The cross-attention score of the i-th head , where , encoder_out is the final output of the encoder; then concatenate the outputs of all attention heads and pass through a linear layer, perform residual connection and layer normalization, then input to the feed-forward neural network layer for calculation, and perform residual connection and layer normalization again to obtain the output decoder_out1 of the first decoder layer, input it to the next decoder layer for the same processing until obtaining the output of the sixth decoder layer, take it as the final output decoder_out of the decoder, and then obtain the output distribution at the current moment through a linear layer and the softmax function; The calculation formula of the loss function in step 3 is as follows: , where represents the summation operation on the values calculated from the 1st moment to the rth moment for t, x is the product molecule representation, represents all target sequence representations before the t-th moment, is the predicted probability of the correct word by Model 1 at the t-th moment, represents that Model 1 gives the correct output according to the current parameters and the target sequence before the time step t, represents the cross-entropy loss; , where represents the probability distribution of a certain output given by Model 2 at the time step t , represents the logarithmic probability of a certain output given by Model 1 at the time step t ; , where is the distillation weight of Model 1 for the learning degree of the output result of Model 2 in the reaction type i. Finally, the total loss is minimized to represent optimizing the parameters of Model 1, so that the list of reactants predicted by Model 1 is aligned with the true list of reactants and the list of reactants predicted by Model 2; The calculation formula of the loss function in step 5 is as follows: , where x is the product molecular representation, represents all target sequence representations before time t, is the predicted probability of the correct word by model 2 at time t, represents the cross-entropy loss function; , where represents the probability distribution of a certain output given by model 1 at time step t , represents the log probability of a certain output given by model 2 at time step t , represents the knowledge distillation loss function; , where is the distillation weight of model 2 in reaction type i for the learning degree of the output result of model 1. Finally, the total loss is minimized represents optimizing the parameters of model 2 , so that the reactant list predicted by model 2 is aligned with the true reactant list and the reactant list predicted by model 1; Step 6 specifically involves searching space O k which has three update strategies: reducing the distillation weight, keeping the distillation weight unchanged, and increasing the distillation weight. For model 1, the parameters of model 1 and the distillation weight are saved in temporary variables for subsequent recovery. Then, traverse the update strategies in O k and adjust the distillation weights of all reaction types in under the current update strategy. Continue to train for one epoch according to the adjusted distillation weights, then verify model 1, calculate the cross-entropy loss with the true reactant list, and store the results in matrix R; when all update strategies have been traversed, for each reaction type of model 1, select the candidate weight corresponding to the minimum verification loss from matrix R as the updated distillation weight; then, traverse all strategies for model 2 and verify it, and update the distillation weights in model 2.

3. The single-step retrosynthesis reaction prediction method according to claim 1, characterized in that Step 1 is specifically as follows: First, download reaction data from the open-source dataset USPTO-50K. The reaction data is mainly divided into two columns. One column is the reaction type value of the synthesis reaction, and the other column is the RXN reaction equation. The framework of this equation is like reactant1.reactant2>>product molecule. The reactants and product molecules are represented in the Simplified Molecular Input Line Entry Specification (SMILES) format, that is, the text representation of the molecule. Then, according to the reaction type value, the reaction data is divided into 10 reaction data sets. Preprocess the chemical reaction data of each reaction data set. Read the CSV file or TXT file containing the reaction data, extract the SMILES of the reactants and the SMILES of the product molecules, perform quality checks on the SMILES of the reactants and product molecules, and filter out some unqualified data, such as the reactant or product molecule being empty, the SMILES of the reactant or product molecule being invalid, the molecule being too small, or the atom mapping of the reactant and product molecules being inconsistent. Subsequently, adopt the root alignment mode, randomly select an atom from the product molecule as the root atom, and then adjust the atomic order of the reactants according to the atom mapping to align them with the product molecule. Repeat multiple times to generate augmented data until the data is amplified to 20 times the original. Then perform tokenization on the SMILES. The segmentation is achieved through regular expressions. The parts that conform to the following rules are regarded as tokens, such as atoms within square brackets, element symbols, special symbols, two-digit % followed by a number, and single digits. Convert the reaction expression into space-separated tokens. Finally, save the processed source data and target data to the source data file and target data file respectively; Finally, extract the information of the product molecules and reactant lists for the ten reaction data sets to obtain the updated reaction data sets.

4. The single-step retrosynthesis reaction prediction method according to claim 2, characterized in that, The reactant list obtained in Step 1 is to predict the SMILES of multiple corresponding reactant molecules from the SMILES of the augmented product molecules, and the SMILES are connected by "."; The retrosynthesis reaction prediction model is implemented based on the Transformer model. Both Transformer models consist of an encoder, a decoder, and an output linear layer.

5. The single-step retrosynthesis reaction prediction method according to claim 4, characterized in that, The encoder is mainly composed of an embedding layer, a multi-head attention layer, and a feed-forward neural network layer. Its specific structure is as follows: The embedding layer is a linear network. The input of the embedding layer is the molecular information of a batch of product molecules corresponding to the reaction data set under a certain reaction category, and the output is a set of embedding vectors of Batch_size×Seq_length×512, where Batch_size represents the batch size, Seq_length represents the maximum number of tokens used by each molecule in the current batch, and 512 is the dimension of the latent space; The embedding layer includes a word embedding layer and a position encoding layer; The input of the word embedding layer is the molecular information of a batch of molecules corresponding to the reaction data set under a certain reaction type, that is, a matrix of Batch_size×Seq_length×Len, and the output is a set of word embedding vectors of Batch_size×Seq_length×512; The input of the position encoding layer is the molecular information of a batch of molecules corresponding to the reaction data set under a certain reaction type, that is, a matrix of Batch_size×Seq_length×Len, and the output is a set of word embedding vectors of Batch_size×Seq_length×512.

6. The single-step retrosynthesis reaction prediction method according to claim 5, wherein The input of the multi-head self-attention layer is the set of word embedding vectors output by the embedding layer, that is, a matrix of Batch_size×Seq_length×512, and the output is a set of attention vectors of Batch_size×Seq_length×512.

7. The single-step retrosynthesis reaction prediction method according to claim 6, wherein The input of the feed-forward neural network layer is the set of vectors of Batch_size×Seq_length×512 output by the multi-head attention layer, and the output is a set of molecular representation vectors of Batch_size×Seq_length×512; The encoder layer is composed of a multi-head self-attention layer and a feed-forward neural network layer, and is stacked 6 layers in total; the input of the encoder layer is the set of word embedding vectors, and the output is the set of molecular representation vectors encoder_out.

8. The single-step retrosynthesis reaction prediction method according to claim 7, wherein The decoder is mainly composed of an embedding layer, a masked multi-head attention layer, a multi-head attention layer, and a feed-forward neural network layer; The embedding layer is a linear neural network, and its input is the known result x from the beginning to the current time t. t The output is a set of word embedding vectors of Batch_size × Seq_length × 512. The specific structure of the embedding layer and the processing process of the input are exactly the same as those of the embedding layer module of the encoder; The input of the masked multi-head attention layer module is the set of word embedding vectors output by the decoder embedding layer, and the output is a set of masked attention vectors of Batch_size×Seq_length×512; The input of the multi-head attention layer module is the set of vectors output by the decoder masked multi-head attention layer, and the output is a set of attention vectors of Batch_size×Seq_length×512; The specific structure of the feed-forward neural network layer and the processing process of the input are exactly the same as those of the feed-forward neural network layer in the encoder; The decoder layer consists of a masked multi-head attention layer, a multi-head attention layer, and a feed-forward neural network layer, and is stacked 6 layers in total; The input of the decoder layer is the result x at time t t and the set of word embedding vectors, and the output is the set of molecular representation vectors decoder_out; The input of the output linear layer is the output decoder_out of the decoder, and the output is a matrix vector of Batch_size×Seq_length×V, where V is the number of the molecular vocabulary; During the process of predicting reactants, the product molecule SMILES is input into the encoder to obtain the encoder output encoder_out. Starting from time t = 0, the first symbol of the output data is the start symbol, and the rest are masked. The data passes through the decoder and the output linear layer to obtain the probability distribution of the next token. The index of the token is extracted according to the probability and used as known information. The above process is repeated continuously until the index of the end symbol is extracted after the data passes through the decoder and the output linear layer, and the prediction process ends to obtain the reactant list.

9. The single-step retro-synthesis reaction prediction method according to claim 7, wherein In step 1, the predetermined process is executed as follows: the reaction data sets of the first to tenth categories are processed according to the division of the training set, test set, and validation set in the data set USPTO-50K to obtain the training set, test set, and validation set of each reaction type; all the above sets are preprocessed to obtain the updated reaction data sets respectively, and the information of each product molecule and reactant list is extracted from each reaction data set respectively, and then the information of all product molecules and reactant lists in the reaction data sets of the first to tenth categories are respectively constructed into the updated reaction data sets of the first to tenth categories.

10. A single-step retrosynthesis reaction prediction system based on the method according to any one of claims 1-9, characterized in that, Including: The first module obtains a plurality of synthesis reaction data, and each synthesis reaction data has a reaction type value. All reaction data with reaction type values of 1 to 10 are correspondingly constructed into reaction data sets of the first to tenth categories. All the above reaction data sets are preprocessed to obtain ten updated reaction data sets; The second module inputs the product molecule information set in the ten reaction data information sets obtained by the first module into a pre-trained inverse synthesis reaction prediction model to obtain the reactant list results required for generating each product molecule.

Citation Information

Patent Citations

  • Deep learning-based inverse synthesis prediction method and device, medium and equipment

    CN114220496A

  • Diversified inverse synthesis analysis model evaluation method and device

    CN117972531A

  • Improved Transform Blind-Chinese conversion method based on knowledge distillation and continuous writing pre-training

    CN119150887A

  • Automatic compression method and platform for multilevel knowledge distillation-based pre-trained language model

    WO2022126797A1

Cited By

  • Semi-template drug molecule inverse synthesis prediction method based on graph-sequence pre-training

    CN120823906A