A method and device for training a reaction condition prediction model
By applying the reaction center masking strategy and multi-head attention mechanism on the language representation network layer, combined with the reaction condition decoder, the reaction condition prediction model is trained, and the problem of low prediction accuracy in the existing technology is solved, achieving a more efficient prediction effect.
Patent Information
- Application Number
- CN202211579416.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-08
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-12-08
AI Technical Summary
When the prior art performs reaction condition prediction on large-scale general chemical reaction data sets, the prediction accuracy is low and the model interpretability is poor.
The language characterization network layer based on the reaction center masking strategy is used for pre-training, combined with the reaction condition decoder, and through the multi-head attention mechanism and cross-entropy loss optimization, the reaction condition prediction model is trained.
The accuracy of reaction condition prediction is significantly improved, allowing the model to more accurately predict catalysts, solvents, reagents and temperatures on large data sets.
Smart Images

Figure CN116227485B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular, to a method and device for training a reaction condition prediction model. Background Art
[0002] The successful implementation of synthetic planning depends on the selection of accurate reaction conditions. A reliable reaction condition prediction algorithm can help researchers optimize chemical reactions more efficiently, thereby obtaining target molecules faster. Although some reaction condition recommendation algorithms have been proposed by researchers, most of them are prediction schemes for single reactions or single conditions and cannot be directly applied to the synthetic planning process.
[0003] Currently, some machine learning algorithms have been applied to the reaction condition prediction problem. However, most of the condition recommendation algorithms using advanced molecular representations only recommend for single reactions and cannot be trained on large-scale general chemical reaction datasets, and there are quite some difficulties when applied to reaction planning algorithm applications. The currently commonly used model algorithm uses the idea similar to RNN, uses molecular fingerprints as the input of the neural network, first predicts the most important catalyst, and uses the conditions predicted at each step as the input composition for the next step condition prediction, and sequentially predicts the catalyst, two solvents, two reagents, and temperature. The interpretability is poor and the model prediction accuracy is relatively low. Summary of the Invention
[0004] The technical problem to be solved by the embodiments of the present application is to provide a method and device for training a reaction condition prediction model to improve the prediction accuracy on a large-scale reaction condition prediction dataset.
[0005] In a first aspect, an embodiment of the present application provides a method for training a reaction condition prediction model, the method including:
[0006] Obtain a pre-training sample of a language representation network layer based on a reaction center masking strategy, where the pre-training sample includes: the SMILES sequence of a chemical reaction expression;
[0007] Identify the chemical reaction center in the SMILES sequence of the chemical reaction expression, take a masking rate of 0.5 for the corresponding vocabulary, take a masking probability of 0.15 for the remaining vocabulary, call the language representation network layer to predict the masked vocabulary, calculate the cross-entropy loss, and optimize the language representation network layer;
[0008] Obtain a model training sample of the reaction condition prediction model, where the model training sample includes: the SMILES sequence of a chemical reaction expression and a reaction condition sequence;
[0009] Input the model training samples into the reaction condition prediction model to be trained, where the reaction condition prediction model to be trained includes: a pre-trained language representation network layer and a reaction condition decoder;
[0010] Call the language representation network layer to process the SMILES sequence of the chemical reaction expression, and obtain hidden tensors representing the chemical reaction sequence with attention weight information corresponding to each word in the SMILES sequence of the chemical reaction expression;
[0011] Call the masked multi-head attention network in the reaction condition decoder to perform attention representation on the reaction condition sequence to obtain the corresponding condition sequence features, and based on the hidden tensor and the condition sequence feature tensor, obtain the predicted reaction condition sequence;
[0012] Based on the reaction condition sequence and the predicted reaction condition sequence, calculate the loss value of the reaction condition prediction model to be trained;
[0013] When the loss value is within a preset range, use the trained reaction condition prediction model to be trained as the final reaction condition prediction model.
[0014] Optionally, the reaction condition prediction model to be trained further includes: a feature extraction network layer and a vector conversion layer,
[0015] Before calling the language representation network layer to process the SMILES sequence of the chemical reaction expression and obtain the hidden tensors representing the chemical reaction sequence with attention weight information corresponding to each word in the SMILES sequence of the chemical reaction expression, it further includes:
[0016] Call the feature extraction network layer to extract each word in the SMILES sequence of the chemical reaction expression and the position feature of each word in the SMILES sequence of the chemical reaction expression;
[0017] Call the vector conversion layer to perform vector conversion processing on each word and the position feature respectively, and obtain the word vector corresponding to each word and the position vector corresponding to the position feature.
[0018] Optionally, the language representation network layer includes: a multi-head attention mechanism layer, a layer normalization and Dropout layer, a feed-forward neural network, and a language representation encoder,
[0019] The step of calling the language representation network layer to process the SMILES sequence of the chemical reaction expression and obtain the hidden tensors representing the chemical reaction sequence with attention weight information corresponding to each word in the SMILES sequence of the chemical reaction expression includes:
[0020] Call the multi - head attention mechanism layer to process the sum - normalized result of the word vector and the position vector, and obtain the correlation relationship between each word;
[0021] Call the layer normalization and Dropout layer to process the word vector and the position vector according to the correlation relationship, and obtain a hidden tensor with weights between each word;
[0022] Pass the hidden tensor with weights between each word through a feed - forward neural network to obtain a feed - forward neural network hidden tensor;
[0023] Call the language representation output layer to fuse and process the weights between each word and the forward processing result, and obtain and output the hidden tensor of the language representation encoder;
[0024] When the internal tensor of the language representation encoder flows, it includes multiple residual connections, namely the residual connection from the input embedding layer to the encoder self - output layer and the residual connection from the self - output layer to the language representation output layer.
[0025] Optionally, the reaction condition decoder includes: a masked multi - head attention network, a multi - head attention mechanism layer, and a feed - forward neural network,
[0026] The step of calling the masked multi - head attention network in the reaction condition decoder to perform attention representation on the reaction condition sequence to obtain the corresponding conditional sequence features, and based on the hidden tensor and the conditional sequence feature tensor to obtain the predicted reaction condition sequence includes:
[0027] Call the masked multi - head attention network to learn the attention weights of the reaction condition sequence to obtain masked conditional sequence features;
[0028] Call the multi - head attention mechanism layer to perform attention learning on the masked conditional sequence features based on the attention weights to obtain initial predicted conditional sequence features;
[0029] Call the feed - forward neural network to perform forward learning on the initial predicted conditional sequence features to obtain the predicted reaction condition sequence.
[0030] Optionally, the step of calculating the loss value of the to - be - trained reaction condition prediction model based on the reaction condition sequence and the predicted reaction condition sequence includes:
[0031] Calculate the cross - entropy loss value based on the reaction condition sequence and the predicted reaction condition sequence, and use the cross - entropy loss value as the loss value of the to - be - trained reaction condition prediction model.
[0032] Optionally, the reaction condition prediction model to be trained further includes: a temperature decoder,
[0033] After the masked multi-head attention network in the reaction condition decoder is called to perform attention representation on the reaction condition sequence to obtain the corresponding conditional sequence features, and based on the hidden tensor and the conditional sequence feature tensor, a predicted reaction condition sequence is obtained, the following steps are further included:
[0034] Call the temperature decoder to process the predicted reaction condition sequence according to the attention weights to obtain the predicted reaction temperature corresponding to the SMILES sequence of the chemical reaction expression.
[0035] Optionally, calculating the loss value of the reaction condition prediction model to be trained based on the reaction condition sequence and the predicted reaction condition sequence includes:
[0036] Calculating a cross-entropy loss value based on the reaction condition sequence and the predicted reaction condition sequence;
[0037] Calculating a mean square error loss value based on the predicted reaction temperature;
[0038] Calculating the loss value of the reaction condition prediction model to be trained based on the cross-entropy loss value, the mean square error loss value, and a balance coefficient.
[0039] In a second aspect, an embodiment of the present application provides a reaction condition prediction model training device, and the device includes:
[0040] A pre-training sample acquisition module, configured to acquire pre-training samples of a language representation network layer based on a reaction center masking strategy, where the pre-training samples include: SMILES sequences of chemical reaction expressions;
[0041] A language representation network layer optimization module, configured to identify chemical reaction centers in the SMILES sequence of the chemical reaction expression, take a masking rate of 0.5 for the corresponding vocabulary, and a masking probability of 0.15 for the remaining vocabulary, call the language representation network layer to predict the masked vocabulary, calculate the cross-entropy loss, and optimize the language representation network layer;
[0042] A model training sample acquisition module, configured to acquire model training samples of a reaction condition prediction model, where the model training samples include: SMILES sequences of chemical reaction expressions and reaction condition sequences;
[0043] A model training sample input module, configured to input the model training samples into the reaction condition prediction model to be trained, where the reaction condition prediction model to be trained includes: a pre-trained language representation network layer and a reaction condition decoder;
[0044] A hidden tensor acquisition module, which is used to call the language representation network layer to process the SMILES sequence of the chemical reaction expression, and obtain a hidden tensor representing the chemical reaction sequence with attention weight information corresponding to each word in the SMILES sequence of the chemical reaction expression;
[0045] A predicted reaction condition sequence acquisition module, which is used to call the masked multi-head attention network in the reaction condition decoder to perform attention representation on the reaction condition sequence to obtain the corresponding condition sequence features, and based on the hidden tensor and the condition sequence feature tensor, obtain a predicted reaction condition sequence;
[0046] A loss value calculation module, which is used to calculate the loss value of the reaction condition prediction model to be trained based on the reaction condition sequence and the predicted reaction condition sequence;
[0047] A reaction condition prediction model acquisition module, which is used to use the trained reaction condition prediction model to be trained as the final reaction condition prediction model when the loss value is within a preset range.
[0048] Optionally, the reaction condition prediction model to be trained further includes: a feature extraction network layer and a vector conversion layer,
[0049] The device further includes:
[0050] A position feature acquisition module, which is used to call the feature extraction network layer to extract each word in the SMILES sequence of the chemical reaction expression and the position feature of each word in the SMILES sequence of the chemical reaction expression;
[0051] A position vector acquisition module, which is used to call the vector conversion layer to perform vector conversion processing on each word and the position feature respectively, and obtain a word vector corresponding to each word and a position vector corresponding to the position feature.
[0052] Optionally, the language representation network layer includes: a multi-head attention mechanism layer, a layer normalization and Dropout layer, a feed-forward neural network, and a language representation encoder,
[0053] The hidden tensor acquisition module includes:
[0054] An association relationship acquisition unit, which is used to call the multi-head attention mechanism layer to process the sum-normalized result of the word vector and the position vector, and obtain the association relationship between each word;
[0055] A hidden tensor acquisition unit, which is used to call the layer normalization and Dropout layer to process the word vector and the position vector according to the association relationship, and obtain a hidden tensor with weights between each word;
[0056] A feedforward network tensor obtaining unit, configured to obtain a feedforward neural network hidden tensor by passing the hidden tensor with weights between each word through a feedforward neural network;
[0057] An encoder tensor obtaining unit, configured to call the language representation output layer to perform a fusion process on the weights between each word and the forward processing result, and obtain and output the hidden tensor of the language representation encoder;
[0058] When tensors flow inside the language representation encoder, it includes multiple residual connections, namely the residual connection from the input embedding layer to the encoder self-output layer and the residual connection from the self-output layer to the language representation output layer.
[0059] Optionally, the reaction condition decoder includes: a masked multi-head attention network, a multi-head attention mechanism layer, and a feedforward neural network,
[0060] The predicted reaction condition sequence obtaining module includes:
[0061] A mask sequence feature obtaining unit, configured to call the masked multi-head attention network to perform attention weight learning on the reaction condition sequence, and obtain a masked condition sequence feature;
[0062] An initial sequence feature obtaining unit, configured to call the multi-head attention mechanism layer to perform attention learning on the masked condition sequence feature based on the attention weights, and obtain an initial predicted condition sequence feature;
[0063] A predicted sequence feature obtaining unit, configured to call the feedforward neural network to perform forward learning on the initial predicted condition sequence feature, and obtain the predicted reaction condition sequence.
[0064] Optionally, the loss value calculation module includes:
[0065] A first loss value calculation unit, configured to calculate a cross-entropy loss value based on the reaction condition sequence and the predicted reaction condition sequence, and use the cross-entropy loss value as the loss value of the reaction condition prediction model to be trained.
[0066] Optionally, the reaction condition prediction model to be trained further includes: a temperature decoder,
[0067] The apparatus further includes:
[0068] A predicted reaction temperature obtaining module, configured to call the temperature decoder to process the predicted reaction condition sequence according to the attention weights, and obtain the predicted reaction temperature corresponding to the SMILES sequence of the chemical reaction formula.
[0069] Optionally, the loss value calculation module includes:
[0070] A cross - entropy loss value calculation unit, configured to calculate a cross - entropy loss value based on the reaction condition sequence and the predicted reaction condition sequence;
[0071] A mean squared error loss calculation unit, configured to calculate a mean squared error loss value based on the predicted reaction temperature;
[0072] A second loss value calculation unit, configured to calculate the loss value of the reaction condition prediction model to be trained based on the cross - entropy loss value, the mean squared error loss value, and a balance coefficient.
[0073] In a third aspect, an embodiment of the present application provides an electronic device, including:
[0074] A processor, a memory, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, the reaction condition prediction model training method described in any one of the above is implemented.
[0075] In a fourth aspect, an embodiment of the present application provides a computer - readable storage medium, when instructions in the storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the reaction condition prediction model training method described in any one of the above.
[0076] Compared with the prior art, the embodiments of the present application include the following advantages:
[0077] In the embodiments of the present application, by obtaining pre-training samples of the language representation network layer based on a reaction center masking strategy, the pre-training samples include: SMILES sequences of chemical reaction expressions; identifying the chemical reaction centers in the SMILES sequences of the chemical reaction expressions, taking a masking rate of 0.5 for the corresponding vocabulary, and taking a masking probability of 0.15 for the remaining vocabulary, calling the language representation network layer to predict the masked vocabulary, calculating the cross-entropy loss, and optimizing the language representation network layer; obtaining model training samples of the reaction condition prediction model, the model training samples include: SMILES sequences of chemical reaction expressions and reaction condition sequences; inputting the model training samples into the reaction condition prediction model to be trained, the reaction condition prediction model to be trained includes: a pre-trained language representation network layer and a reaction condition decoder; calling the language representation network layer to process the SMILES sequences of the chemical reaction expressions to obtain hidden tensors representing the chemical reaction sequences with attention weight information corresponding to each word in the SMILES sequences of the chemical reaction expressions; calling the masked multi-head attention network in the reaction condition decoder to perform attention representation on the reaction condition sequences to obtain corresponding condition sequence features, and based on the hidden tensors and the condition sequence feature tensors, obtaining a predicted reaction condition sequence; calculating a loss value of the reaction condition prediction model to be trained based on the reaction condition sequence and the predicted reaction condition sequence; in the case where the loss value is within a preset range, taking the trained reaction condition prediction model to be trained as the final reaction condition prediction model. In the embodiments of the present application, by regarding the prediction of the reaction background as a sequence-to-sequence translation task (i.e., the reaction condition sequence formed by the catalyst, solvent 1, solvent 2, reagent 1, and reagent 2), using an Attention-based prediction model to represent chemical reactions using sequences, and respectively predicting the catalyst, solvent, reagent, and temperature, the prediction accuracy can be significantly improved on a large reaction condition prediction dataset.
[0078] It should be understood that the above general description and subsequent detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] Figure 1 is a flowchart of the steps of a reaction condition prediction model training method provided by an embodiment of the present application;
[0080] Figure 2 is a schematic diagram of the prediction process of a reaction condition prediction model provided by an embodiment of the present application;
[0081] Figure 3 is a schematic diagram of the pre-training process of a masked reaction center modeling provided by an embodiment of the present application;
[0082] Figure 4Schematic structural diagram of a reaction condition prediction model training device provided by an embodiment of the present application;
[0083] Figure 5 Schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0084] To make the above objects, features, and advantages of the present application more apparent and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific implementation manners.
[0085] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "the", and "said" used in the embodiments of the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.
[0086] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or terminal including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or terminal. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or terminal including the element.
[0087] To solve the problems of poor interpretability and low model prediction accuracy when using existing model algorithms for reaction condition prediction. An interpretable pre-training-based reaction condition predictor (Parrot) is proposed in the embodiments of the present application. The model uses Bert as an encoder to extract reaction features from reaction SMILES, and uses a transformer decoder with fewer parameters to calculate the hidden layer representation of reaction background conditions. Finally, a classifier is used to predict reaction conditions, and a regression layer (temperature decoder) is used to calculate the temperature. The structure of the model provided by the embodiments of the present application can be as Figure 1As shown in the figure, the prediction of chemical background (catalyst, solvent 1, solvent 2, reagent 1, reagent 2) can be regarded as a sequence-to-sequence translation task. The conditions predicted later also consider the conditions that have been predicted, but the target sequence has a fixed length (length 6). The information contained in the memory tensor of the encoder and the output tensor of the decoder is used to predict the temperature corresponding to five reaction conditions. Each of these tensors is transformed by a feed-forward neural network and fed into a third feed-forward neural network after tensor concatenation to calculate the temperature.
[0088] Next, the technical solutions of the embodiments of the present application will be described in detail as follows in combination with specific embodiments.
[0089] Referring to Figure 1 , a step flow chart of a reaction condition prediction model training method provided by an embodiment of the present application is shown. As Figure 1 shown, the reaction condition prediction model training method may include the following steps:
[0090] Step 101: Obtain pre-training samples of the language representation network layer based on the reaction center masking strategy. The pre-training samples include: the SMILES sequence of the chemical reaction expression.
[0091] Embodiments of the present application can be applied to a Transformer variant model based on Attention, using sequences to represent chemical reactions, and predicting scenarios of catalysts, solvents, reagents, and temperature respectively.
[0092] Embodiments of the present application design two pre-training strategies, Masked Language Modeling (Masked ML) and Masked Reaction Center Modeling (Masked RCM) including chemical reaction domain knowledge. The reaction data sets used in the two pre-training strategies are approximately 1.3 million reaction SMILES obtained by cleaning USPTO 1976 - 2016sep. These data have cleared all reaction conditions and only include reactants and products (reactants >> products), maintaining the same input format and content as the reaction condition prediction task. A Masked RCM training strategy with domain knowledge is designed. The implementation schematic diagram of this strategy is as Figure 3 shown. In the Masked RCM strategy, in order to strengthen the model's understanding of the reaction center, the masking probability of the reaction center label can be increased to 0.5 instead of 0.15. Through Masked RCM, the model can pay more attention to the prediction and embedding of the vocabulary representing the reaction center.
[0093] In this embodiment, first, pre-training samples of the language representation network layer based on the reaction center masking strategy can be obtained to optimize the language representation network layer of the reaction condition prediction model. Among them, the pre-training samples may include: the SMILES sequence of the chemical reaction expression.
[0094] After obtaining the pre-training samples of the language representation network layer based on the reaction center masking strategy, step 102 is executed.
[0095] Step 102: Identify the chemical reaction center in the SMILES sequence of the chemical reaction expression, set the masking rate of the corresponding vocabulary to 0.5, and set the masking probability of the remaining vocabulary to 0.15. Call the language representation network layer to predict the masked vocabulary, calculate the cross-entropy loss, and optimize the language representation network layer.
[0096] After obtaining the pre-training samples of the language representation network layer based on the reaction center masking strategy, the chemical reaction center in the SMILES sequence of the chemical reaction expression can be identified, the masking rate of the corresponding vocabulary can be set to 0.5, and the masking probability of the remaining vocabulary can be set to 0.15. Call the language representation network layer to predict the masked vocabulary, calculate the cross-entropy loss, and optimize the language representation network layer.
[0097] After optimizing the language representation network layer, step 103 is executed.
[0098] Step 103: Obtain the model training samples of the reaction condition prediction model, where the model training samples include: the SMILES sequence of the chemical reaction expression and the reaction condition sequence.
[0099] The model training sample refers to the sample used to train the reaction condition prediction model to be trained. In this example, the model training sample can include: the SMILES sequence of the chemical reaction expression and the reaction condition sequence, where the reaction condition sequence can be a sequence formed by five reaction conditions: catalyst, solvent, solvent, reagent, and reagent in sequence.
[0100] When training the reaction condition prediction model to be trained, the model training samples can be obtained. In a specific implementation, the model training samples can be the SMILES sequences of chemical reaction expressions obtained by cleaning USPTO 1976 - 2016sep.
[0101] After obtaining the model training samples of the reaction condition prediction model, step 104 is executed.
[0102] Step 104: Input the model training samples into the reaction condition prediction model to be trained, where the reaction condition prediction model to be trained includes: a pre-trained language representation network layer and a reaction condition decoder.
[0103] After obtaining the model training samples, the model training samples can be input into the reaction condition prediction model to be trained. In this example, the reaction condition prediction model to be trained can include: a pre-trained language representation network layer and a reaction condition decoder, such asFigure 2 As shown, the language representation network layer is: Bert Layer, and the reaction condition decoder is: ConditionDecoder.
[0104] After inputting the model training sample into the reaction condition prediction model to be trained, step 105 is executed.
[0105] Step 105: Invoke the language representation network layer to process the SMILES sequence of the chemical reaction expression, and obtain a hidden tensor representing the chemical reaction sequence with attention weight information corresponding to each word in the SMILES sequence of the chemical reaction expression.
[0106] After inputting the model training sample into the reaction condition prediction model to be trained, the language representation network layer can be invoked to process the SMILES sequence of the chemical reaction expression, so as to obtain a hidden tensor representing the chemical reaction sequence with attention weight information corresponding to each word in the SMILES sequence of the chemical reaction expression.
[0107] In this embodiment, the reaction condition prediction model to be trained may further include: a feature extraction network layer and a vector conversion layer. Among them, the feature extraction network layer can extract each word in the SMILES sequence of the chemical reaction expression, and the position feature of each word in the SMILES sequence of the chemical reaction expression. The implementation process can be described in detail in combination with the following specific implementation manners.
[0108] In a specific implementation manner of the present application, before the above step 105, it may further include:
[0109] Step A1: Invoke the feature extraction network layer to extract each word in the SMILES sequence of the chemical reaction expression, and the position feature of each word in the SMILES sequence of the chemical reaction expression.
[0110] In this embodiment, after inputting the model training sample into the reaction condition prediction model to be trained, the feature extraction network layer can be invoked to extract each word in the SMILES sequence of the chemical reaction expression, and the position feature of each word in the SMILES sequence of the chemical reaction expression.
[0111] After invoking the feature extraction network layer to extract each word in the SMILES sequence of the chemical reaction expression, and the position feature of each word in the SMILES sequence of the chemical reaction expression, step A2 is executed.
[0112] Step A2: Invoke the vector conversion layer to perform vector conversion processing on each word and the position feature respectively, and obtain a word vector corresponding to each word and a position vector corresponding to the position feature.
[0113] After calling the feature extraction network layer to extract each word in the SMILES sequence of the chemical reaction expression and the position features of each word in the SMILES sequence of the chemical reaction expression, the vector conversion layer can be called to perform vector conversion processing on each word and position feature respectively to obtain the word vector corresponding to each word and the position vector corresponding to the position feature. As Figure 2 shown, the SMILES sequence of the chemical reaction expression can be input into the model structure InputEmbedding in the left part, which includes Token Embedding (word feature extraction network layer and vector conversion layer) to extract each word in the SMILES sequence of the chemical reaction expression and perform vector conversion to obtain the word vector corresponding to each word. The model structure in the left part can also include: Positional Embedding (i.e., position feature extraction layer and vector conversion layer) to extract the position features of each word in the SMILES sequence of the chemical reaction expression and perform vector conversion to obtain the position vector corresponding to each position feature.
[0114] After obtaining the sum of the position vector and the word vector, the subsequent model processing process can be carried out. That is, call the language representation network layer to process the sum of the word vector and the position vector, and this processing process can be described in detail in combination with the following specific implementation manners.
[0115] In another specific implementation manner of the present application, step 105 above may include:
[0116] Sub-step B1: Call the multi-head attention mechanism layer to process the normalized sum of the word vector and the position vector to obtain the correlation relationship between each word.
[0117] In this embodiment, after obtaining the position vector and the word vector, the multi-attention mechanism layer can be called to process the normalized sum of the word vector and the position vector to obtain the correlation relationship between each word. As Figure 2 shown, Input Embedding can output the word vector and the position vector, and then the Bert Self Attention can be called to process the normalized sum of the word vector and the position vector to output the correlation relationship between each word.
[0118] After calling the multi-head attention mechanism layer to process the normalized sum of the word vector and the position vector to obtain the correlation relationship between each word, sub-step B2 is executed.
[0119] Sub-step B2: Call the layer normalization and Dropout layers to process the word vectors and the position vectors according to the association relationship, and obtain a hidden tensor with weights between each word.
[0120] After calling the multi-head attention mechanism layer to process the sum-normalized result of the word vectors and the position vectors to obtain the association relationship between each word, the layer normalization and Dropout layers can be called to process the word vectors and the position vectors according to the association relationship, and obtain a hidden tensor with weights between each word. As Figure 2 shown, Bert SelfAttention outputs the association relationship between each word, and Bert Self Output can process the word vectors and the position vectors output by Input Embedding according to the association relationship between each word output by Bert Self Attention, so as to obtain a hidden tensor with weights between each word. For example, there are 10 words in the SMILES sequence of a chemical reaction expression, and a hidden tensor with weights between these 10 words can be obtained through Bert Self Output, etc.
[0121] It can be understood that the above examples are only examples listed for better understanding of the technical solutions of the embodiments of the present application, and do not serve as the sole limitation of this embodiment.
[0122] After calling the layer normalization and Dropout layers to process the word vectors and the position vectors according to the association relationship, and obtaining a hidden tensor with weights between each word, sub-step B3 is executed.
[0123] Sub-step B3: Pass the hidden tensor with weights between each word through a feed-forward neural network to obtain a feed-forward neural network hidden tensor.
[0124] After calling the layer normalization and Dropout layers to process the word vectors and the position vectors according to the association relationship, and obtaining a hidden tensor with weights between each word, the hidden tensor with weights between each word can be passed through a feed-forward neural network to obtain a feed-forward neural network hidden tensor. As Figure 2 shown, the output of Bert Self Output can be used as the input of Bert Intermediate. The hidden tensor with weights between each word output by Bert Self Output passes through a feed-forward neural network, and a feed-forward neural network hidden tensor can be obtained.
[0125] After passing the hidden tensor with weights between each word through a feed-forward neural network to obtain a feed-forward neural network hidden tensor, sub-step B4 is executed.
[0126] Sub-step B4: Call the language representation output layer to fuse the weights between the words and the forward processing result, and obtain and output the hidden tensor of the language representation encoder.
[0127] After obtaining the hidden tensor with weights between words through a feed-forward neural network to get the feed-forward neural network hidden tensor, the language representation output layer can be called to fuse the weights between the words and the forward processing result, and obtain and output the hidden tensor of the language representation encoder. As Figure 2 shown, the output of Bert Intermediate and the output of BertSelf Output can be used as the input of Bert Output. Calling Bert Output can fuse the weights of each word output by Bert SelfOutput and the forward processing result output by Bert Intermediate, so as to obtain and output the hidden tensor of the language representation encoder.
[0128] After calling the language representation output layer to fuse the weights between the words and the forward processing result, and obtaining and outputting the hidden tensor of the language representation encoder, step 106 is executed.
[0129] Step 106: Call the masked multi-head attention network in the reaction condition decoder to perform attention representation on the reaction condition sequence to obtain the corresponding conditional sequence features, and based on the hidden tensor and the conditional sequence feature tensor, obtain the predicted reaction condition sequence.
[0130] The predicted reaction condition sequence refers to the reaction conditions corresponding to the reaction molecule expressions predicted by the reaction condition prediction model to be trained. That is, the predicted reaction condition sequence can be a sequence formed by: catalyst, solvent, solvent, reagent, reagent.
[0131] After inputting the model training samples into the reaction condition prediction model to be trained, the masked multi-head attention network in the reaction condition decoder can be called to perform attention representation on the reaction condition sequence to obtain the corresponding conditional sequence features, and based on the hidden tensor and the conditional sequence feature tensor, obtain the predicted reaction condition sequence. This implementation process can be described in detail in combination with the following specific implementation manners.
[0132] In a specific implementation manner of the present application, the above step 105 may include:
[0133] Sub-step C1: Call the masked multi-head attention network to learn the attention weights of the reaction condition sequence to obtain the masked conditional sequence features.
[0134] In this embodiment, the masked multi-head attention network includes: a masked multi-head attention network, a multi-head attention mechanism layer, and a feed-forward neural network, as shown in Figure 2 the right part of the model structure. The masked multi-head attention network is the ConditionDecoder (i.e., the conditional decoder). The Condition Decoder may include: Masked Multi-Head Attention, Multi-Head Attention, and Feed Forward.
[0135] After inputting the reaction condition sequence into the reaction condition prediction model to be trained, the masked multi-head attention network can be called to learn the attention weights of the reaction condition sequence to obtain the masked condition sequence features. As shown in Figure 2 the figure, Masked Multi-Head Attention can be called to perform attention learning on the reaction condition sequence, so as to obtain masked condition sequence features, etc.
[0136] After calling the masked multi-head attention network to learn the attention weights of the reaction condition sequence to obtain the masked condition sequence features, sub-step C2 is executed.
[0137] Sub-step C2: Call the multi-head attention mechanism layer to perform attention learning on the masked condition sequence features based on the attention weights to obtain the initial predicted condition sequence features.
[0138] After calling the masked multi-head attention network to learn the attention weights of the reaction condition sequence to obtain the masked condition sequence features, the multi-head attention mechanism layer can be called to perform attention learning on the masked condition sequence features based on the attention weights to obtain the initial predicted condition sequence features. As shown in Figure 2 the figure, the output of Masked Multi-Head Attention and the output of the Bert Layer can be used as the input of Multi-Head Attention. Calling Multi-Head Attention can perform attention learning on the masked condition sequence features output by Masked Multi-Head Attention according to the attention weights output by the Bert Layer to obtain the initial predicted condition sequence features, etc.
[0139] In this example, for the interpretability analysis part of the Parrot reaction condition prediction model, the attention weight a of the model can be used. eAssociate atom w with chemical reaction condition C. Calculate the attention weights through the tensor calculation between the encoder and decoder of the model. These decoder layers all contain multiple heads, and each head learns an attention matrix Attention ∈ R N×M , and this attention weight represents the embedding tensor X of each token in the input reaction sequence X with a length of N i to the embedding tensor Y of each reaction condition sequence vocabulary in the reaction condition sequence Y with an output length of M j . Therefore, each element Attention ij is the attention weight connecting X i to Y j .
[0140] In each head of the multi-head attention layer, the vector representation of each vocabulary X i or Y j will first be converted into key (Q), query (K), and value (V) vectors using the following operations.
[0141] K i = W k X i , Q j = W q Y j , V i = W v X i (1)
[0142] where W k , W q , W v are learnable parameters. A i can be regarded as the correlation probability vector of X i to Y, and is calculated according to the following equation:
[0143]
[0144] In this embodiment, the attention weights of the vocabulary can be converted into atomic attention weights. And the average value of the attention weights of each head is used as the attention weight a for analysis.
[0145] After invoking the multi-head attention mechanism layer to perform attention learning on the masked condition sequence features based on the attention weights to obtain the initial predicted condition sequence features, sub-step C2 is executed.
[0146] Sub-step C3: Invoke the feed-forward neural network to perform forward learning on the initial predicted condition sequence features to obtain the predicted reaction condition sequence.
[0147] After calling the multi-head attention mechanism layer to perform attention learning on the masked conditional sequence features based on the attention weights to obtain the initial predicted conditional sequence features, a feed-forward neural network can be called to perform forward learning on the initial predicted conditional sequence features to obtain the predicted reaction conditional sequence. As Figure 2 shown, the output of the Multi-Head Attention can be used as the input of the FeedForward, and the Feed Forward can perform forward learning on the initial predicted conditional sequence features output by the Multi-Head Attention, so as to obtain the predicted reaction conditional sequence.
[0148] After obtaining the predicted reaction conditional sequence, step 107 is executed.
[0149] Step 107: Calculate the loss value of the reaction condition prediction model to be trained based on the reaction condition sequence and the predicted reaction condition sequence.
[0150] After obtaining the predicted reaction conditional sequence, the loss value of the reaction condition prediction model to be trained can be calculated based on the reaction condition sequence and the predicted reaction condition sequence.
[0151] In this embodiment, when the reaction condition prediction model to be trained does not involve the temperature prediction function, the loss function only includes the classification loss function. Specifically, the cross-entropy loss value can be calculated based on the reaction condition sequence and the predicted reaction condition sequence, and the cross-entropy loss value can be used as the loss value of the reaction condition prediction model to be trained.
[0152] When the reaction condition prediction model to be trained involves the temperature prediction function, the information contained in the memory tensor of the encoder and the decoder output tensor can be used to predict the temperature corresponding to the five reaction conditions. Each of these tensors is deformed by a feed-forward neural network and is fed into a third feed-forward neural network after tensor concatenation to calculate the temperature. The implementation process can be described in detail in combination with the following specific implementation manners.
[0153] In a specific implementation manner of the present application, after the above step 106, the following may further be included:
[0154] Step D1: Call the temperature decoder to process the predicted reaction condition sequence according to the attention weights to obtain the predicted reaction temperature corresponding to the SMILES sequence of the chemical reaction expression.
[0155] In this embodiment, as Figure 2As shown, the temperature decoder is the Temperature Decoder. The output of the Bert Layer and the output of the Condition Decoder can be used as the input of the Temperature Decoder. By calling the Temperature Decoder, the predicted response condition sequence can be processed according to the attention weights to obtain the predicted reaction temperature corresponding to the SMILES sequence of the chemical reaction expression.
[0156] After obtaining the predicted reaction temperature, the loss value of the reaction condition prediction model to be trained can be calculated based on the predicted reaction condition sequence and the predicted reaction temperature. The implementation process can be described in detail in combination with the following specific implementation manners.
[0157] In another specific implementation manner of the present application, step 107 above may include:
[0158] Sub-step E1: Calculate the cross-entropy loss value based on the reaction condition sequence and the predicted reaction condition sequence.
[0159] In this embodiment, after obtaining the predicted reaction condition sequence, the cross-entropy loss value can be calculated based on the reaction condition sequence and the predicted reaction condition sequence.
[0160] Sub-step E2: Calculate the mean squared error loss value based on the predicted reaction temperature.
[0161] After obtaining the predicted reaction temperature, the mean squared error loss value can be calculated based on the predicted reaction temperature.
[0162] Sub-step E3: Calculate the loss value of the reaction condition prediction model to be trained based on the cross-entropy loss value, the mean squared error loss value and the balance coefficient.
[0163] After obtaining the cross-entropy loss value and the mean squared error loss value, the loss value of the reaction condition prediction model to be trained can be calculated based on the cross-entropy loss value, the mean squared error loss value and the balance coefficient. The specific calculation formula is as follows:
[0164]
[0165] In the above formula (3), I is the chemical background condition number, c i is the predicted label of the i-th condition, is the true label of the i-th condition, t is the predicted temperature, is the true temperature. In this embodiment, I = 6 (including 5 chemical background conditions and an end marker).
[0166] Step 108: When the loss value is within a preset range, use the trained reaction condition prediction model to be trained as the final reaction condition prediction model.
[0167] After calculating the loss value of the reaction condition prediction model to be trained, it can be determined whether the loss value is within a preset range.
[0168] If the loss value is not within the preset range, it means that the reaction condition prediction model to be trained has not converged. At this time, the model parameters of the reaction condition prediction model to be trained can be updated according to the calculated loss value, and the reaction condition prediction model to be trained can be continuously trained until the model converges.
[0169] If the loss value is within the preset range, it means that the reaction condition prediction model to be trained has converged. At this time, the trained reaction condition prediction model to be trained can be used as the final reaction condition prediction model for subsequent reaction condition prediction scenarios.
[0170] The reaction condition prediction model training method provided by the embodiments of the present application obtains pre-training samples of the language representation network layer based on the reaction center masking strategy. The pre-training samples include: the SMILES sequence of the chemical reaction expression; identify the chemical reaction center in the SMILES sequence of the chemical reaction expression, and set the masking rate of the corresponding vocabulary to 0.5, and the masking probability of the remaining vocabulary to 0.15. Call the language representation network layer to predict the masked vocabulary, calculate the cross-entropy loss, and optimize the language representation network layer; obtain the model training samples of the reaction condition prediction model. The model training samples include: the SMILES sequence of the chemical reaction expression and the reaction condition sequence; input the model training samples into the reaction condition prediction model to be trained. The reaction condition prediction model to be trained includes: a pre-trained language representation network layer and a reaction condition decoder; call the language representation network layer to process the SMILES sequence of the chemical reaction expression to obtain a hidden tensor representing the chemical reaction sequence with attention weight information corresponding to each word in the SMILES sequence of the chemical reaction expression; call the masked multi-head attention network in the reaction condition decoder to perform attention representation on the reaction condition sequence to obtain the corresponding condition sequence features, and based on the hidden tensor and the condition sequence feature tensor, obtain the predicted reaction condition sequence; based on the reaction condition sequence and the predicted reaction condition sequence, calculate the loss value of the reaction condition prediction model to be trained; when the loss value is within a preset range, use the trained reaction condition prediction model to be trained as the final reaction condition prediction model. The embodiments of the present application regard the prediction of the reaction background as a sequence-to-sequence translation task (i.e., the reaction condition sequence formed by the catalyst, solvent 1, solvent 2, reagent 1, and reagent 2), and use an Attention-based prediction model to represent chemical reactions using sequences, and predict the catalyst, solvent, reagent, and temperature respectively, which can significantly improve the prediction accuracy on a large reaction condition prediction dataset.
[0171] Referring to Figure 4 , a schematic structural diagram of a reaction condition prediction model training device provided by an embodiment of the present application is shown. As Figure 4 shown, the reaction condition prediction model training device 400 may include the following modules:
[0172] The pre-training sample acquisition module 410 is used to obtain pre-training samples of the language representation network layer based on the reaction center masking strategy. The pre-training samples include: the SMILES sequence of the chemical reaction expression;
[0173] The language representation network layer optimization module 420 is used to identify the chemical reaction center in the SMILES sequence of the chemical reaction expression, take a masking rate of 0.5 for the corresponding vocabulary, and take a masking probability of 0.15 for the remaining vocabulary, call the language representation network layer to predict the masked vocabulary, calculate the cross-entropy loss, and optimize the language representation network layer;
[0174] The model training sample acquisition module 430 is used to acquire the model training samples of the reaction condition prediction model, and the model training samples include: the SMILES sequence of the chemical reaction expression and the reaction condition sequence;
[0175] The model training sample input module 440 is used to input the model training samples into the reaction condition prediction model to be trained, and the reaction condition prediction model to be trained includes: a pre-trained language representation network layer and a reaction condition decoder;
[0176] The hidden tensor acquisition module 450 is used to call the language representation network layer to process the SMILES sequence of the chemical reaction expression, and obtain the hidden tensor representing the chemical reaction sequence with attention weight information corresponding to each word in the SMILES sequence of the chemical reaction expression;
[0177] The predicted reaction condition sequence acquisition module 460 is used to call the masked multi-head attention network in the reaction condition decoder to perform attention representation on the reaction condition sequence to obtain the corresponding condition sequence features, and based on the hidden tensor and the condition sequence feature tensor, obtain the predicted reaction condition sequence;
[0178] The loss value calculation module 470 is used to calculate the loss value of the reaction condition prediction model to be trained based on the reaction condition sequence and the predicted reaction condition sequence;
[0179] The reaction condition prediction model acquisition module 480 is used to use the trained reaction condition prediction model to be trained as the final reaction condition prediction model when the loss value is within a preset range.
[0180] Optionally, the reaction condition prediction model to be trained further includes: a feature extraction network layer and a vector conversion layer,
[0181] The device further includes:
[0182] The position feature acquisition module is used to call the feature extraction network layer to extract each word in the SMILES sequence of the chemical reaction expression and the position feature of each word in the SMILES sequence of the chemical reaction expression;
[0183] A position vector acquisition module, configured to call the vector conversion layer to perform vector conversion processing on each word and the position feature respectively, so as to obtain a word vector corresponding to each word and a position vector corresponding to the position feature.
[0184] Optionally, the language representation network layer includes: a multi-head attention mechanism layer, a layer normalization and Dropout layer, a feed-forward neural network, and a language representation encoder.
[0185] The hidden tensor acquisition module includes:
[0186] An association relationship acquisition unit, configured to call the multi-head attention mechanism layer to process the sum-normalized result of the word vector and the position vector, so as to obtain the association relationship between each word.
[0187] A hidden tensor acquisition unit, configured to call the layer normalization and Dropout layer to process the word vector and the position vector according to the association relationship, so as to obtain a hidden tensor with weights between each word.
[0188] A feed-forward network tensor acquisition unit, configured to pass the hidden tensor with weights between each word through the feed-forward neural network to obtain a feed-forward neural network hidden tensor.
[0189] An encoder tensor acquisition unit, configured to call the language representation output layer to fuse the weights between each word and the forward processing result, so as to obtain and output the hidden tensor of the language representation encoder.
[0190] When the internal tensor of the language representation encoder flows, it includes multiple residual connections, namely the residual connection from the input embedding layer to the encoder self-output layer and the residual connection from the self-output layer to the language representation output layer.
[0191] Optionally, the reaction condition decoder includes: a masked multi-head attention network, a multi-head attention mechanism layer, and a feed-forward neural network.
[0192] The predicted reaction condition sequence acquisition module includes:
[0193] A mask sequence feature acquisition unit, configured to call the masked multi-head attention network to learn the attention weights of the reaction condition sequence, so as to obtain a masked condition sequence feature.
[0194] An initial sequence feature acquisition unit, configured to call the multi-head attention mechanism layer to perform attention learning on the masked condition sequence feature based on the attention weights, so as to obtain an initial predicted condition sequence feature.
[0195] A prediction sequence feature acquisition unit, configured to call the feedforward neural network to perform forward learning on the initial prediction condition sequence feature to obtain the predicted reaction condition sequence.
[0196] Optionally, the loss value calculation module includes:
[0197] A first loss value calculation unit, configured to calculate a cross-entropy loss value based on the reaction condition sequence and the predicted reaction condition sequence, and use the cross-entropy loss value as the loss value of the reaction condition prediction model to be trained.
[0198] Optionally, the reaction condition prediction model to be trained further includes: a temperature decoder,
[0199] The apparatus further includes:
[0200] A predicted reaction temperature acquisition module, configured to call the temperature decoder to process the predicted reaction condition sequence according to the attention weights to obtain the predicted reaction temperature corresponding to the SMILES sequence of the chemical reaction expression.
[0201] Optionally, the loss value calculation module includes:
[0202] A cross-entropy loss value calculation unit, configured to calculate a cross-entropy loss value based on the reaction condition sequence and the predicted reaction condition sequence;
[0203] A mean square error loss calculation unit, configured to calculate a mean square error loss value based on the predicted reaction temperature;
[0204] A second loss value calculation unit, configured to calculate the loss value of the reaction condition prediction model to be trained based on the cross-entropy loss value, the mean square error loss value, and a balance coefficient.
[0205] The reaction condition prediction model training device provided by the embodiment of the present application obtains pre-training samples of the language representation network layer based on the reaction center masking strategy. The pre-training samples include: the SMILES sequence of the chemical reaction expression; identify the chemical reaction center in the SMILES sequence of the chemical reaction expression, and take a masking rate of 0.5 for the corresponding vocabulary, and a masking probability of 0.15 for the remaining vocabulary. Call the language representation network layer to predict the masked vocabulary, calculate the cross-entropy loss, and optimize the language representation network layer; obtain the model training samples of the reaction condition prediction model. The model training samples include: the SMILES sequence of the chemical reaction expression and the reaction condition sequence; input the model training samples into the reaction condition prediction model to be trained. The reaction condition prediction model to be trained includes: a pre-trained language representation network layer and a reaction condition decoder; call the language representation network layer to process the SMILES sequence of the chemical reaction expression to obtain a hidden tensor representing the chemical reaction sequence with attention weight information corresponding to each word in the SMILES sequence of the chemical reaction expression; call the masked multi-head attention network in the reaction condition decoder to perform attention representation on the reaction condition sequence to obtain the corresponding condition sequence features, and based on the hidden tensor and the condition sequence feature tensor, obtain the predicted reaction condition sequence; based on the reaction condition sequence and the predicted reaction condition sequence, calculate the loss value of the reaction condition prediction model to be trained; in the case where the loss value is within a preset range, use the trained reaction condition prediction model to be trained as the final reaction condition prediction model. By regarding the prediction of the reaction background as a sequence-to-sequence translation task (i.e., the reaction condition sequence formed by the catalyst, solvent 1, solvent 2, reagent 1, and reagent 2), and using an Attention-based prediction model to represent the chemical reaction with a sequence, the catalyst, solvent, reagent, and temperature can be predicted respectively, which can significantly improve the prediction accuracy on a large reaction condition prediction data set.
[0206] Embodiment III
[0207] The embodiment of the present application provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the above-mentioned reaction condition prediction model training method is implemented.
[0208] Figure 5 Fig. shows a schematic structural diagram of an electronic device 500 according to an embodiment of the present invention. As Figure 5As shown, the electronic device 500 includes a central processing unit (CPU) 501, which can execute various appropriate actions and processes according to the computer program instructions stored in the read-only memory (ROM) 502 or the computer program instructions loaded from the storage unit 508 into the random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 can also be stored. The CPU 501, ROM 502, and RAM 503 are connected to each other via a bus 504. The input / output (I / O) interface 505 is also connected to the bus 504.
[0209] Multiple components in the electronic device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, mouse, microphone, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, optical disk, etc.; and a communication unit 509, such as a network card, modem, wireless communication transceiver, etc. The communication unit 509 allows the electronic device 500 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0210] Each of the processes and processes described above can be executed by the processing unit 501. For example, the method of any of the above embodiments can be implemented as a computer software program, which is tangibly included in a computer-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the CPU 501, one or more actions in the method described above can be executed.
[0211] Embodiment Four
[0212] The embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned reaction condition prediction model training method is implemented.
[0213] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same and similar parts among the embodiments can be referred to each other.
[0214] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, apparatuses, or computer program products. Therefore, the embodiments of the present application can take the form of all-hardware embodiments, all-software embodiments, or embodiments combining software and hardware aspects. Moreover, the embodiments of the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0215] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, terminals (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminals to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminals generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0216] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminals to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0217] These computer program instructions can also be loaded onto a computer or other programmable data processing terminals, so that a series of operation steps are executed on the computer or other programmable terminals to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable terminals provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0218] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.
[0219] The above has introduced in detail a method, device, electronic device, and computer-readable storage medium for training a reaction condition prediction model. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A method for training a reaction condition prediction model, characterized in that The method includes: Obtaining pre-training samples of the language representation network layer based on a reaction center masking strategy, where the pre-training samples include: SMILES sequences of chemical reaction expressions; Identifying the chemical reaction centers in the SMILES sequences of the chemical reaction expressions, taking the corresponding vocabulary with a masking rate of 0.5, and the remaining vocabulary with a masking probability of 0.15, calling the language representation network layer to predict the masked vocabulary, calculating the cross-entropy loss, and optimizing the language representation network layer; Obtaining model training samples of the reaction condition prediction model, where the model training samples include: SMILES sequences of chemical reaction expressions and reaction condition sequences; Inputting the model training samples into the reaction condition prediction model to be trained, where the reaction condition prediction model to be trained includes: a pre-trained language representation network layer and a reaction condition decoder; Calling the language representation network layer to process the SMILES sequences of the chemical reaction expressions to obtain hidden tensors representing the chemical reaction sequences with attention weight information corresponding to each word in the SMILES sequences of the chemical reaction expressions; Calling the masked multi-head attention network in the reaction condition decoder to perform attention representation on the reaction condition sequences to obtain corresponding condition sequence features, and based on the hidden tensors and the condition sequence feature tensors, obtaining predicted reaction condition sequences; Calculating the loss value of the reaction condition prediction model to be trained based on the reaction condition sequences and the predicted reaction condition sequences; When the loss value is within a preset range, taking the trained reaction condition prediction model to be trained as the final reaction condition prediction model.
2. The method according to claim 1, wherein The reaction condition prediction model to be trained further includes: a feature extraction network layer and a vector conversion layer, Before the step of calling the language representation network layer to process the SMILES sequences of the chemical reaction expressions to obtain hidden tensors representing the chemical reaction sequences with attention weight information corresponding to each word in the SMILES sequences of the chemical reaction expressions, it further includes: Calling the feature extraction network layer to extract each word in the SMILES sequences of the chemical reaction expressions and the position features of each word in the SMILES sequences of the chemical reaction expressions; Calling the vector conversion layer to perform vector conversion processing on each word and the position features respectively to obtain word vectors corresponding to each word and position vectors corresponding to the position features.
3. The method according to claim 2, wherein The language representation network layer includes: a multi-head attention mechanism layer, a layer normalization and Dropout layer, a feed-forward neural network, and a language representation encoder, The step of calling the language representation network layer to process the SMILES sequences of the chemical reaction expressions to obtain hidden tensors representing the chemical reaction sequences with attention weight information corresponding to each word in the SMILES sequences of the chemical reaction expressions includes: Calling the multi-head attention mechanism layer to process the sum-normalized results of the word vectors and the position vectors to obtain the correlation relationships between each word; Call the layer normalization and Dropout layers to process the word vectors and the position vectors according to the association relationship, and obtain a hidden tensor with weights between each word; Pass the hidden tensor with weights between each word through a feed-forward neural network to obtain a feed-forward neural network hidden tensor; Call the language representation output layer to perform a fusion process on the weights between each word and the forward processing result, and obtain and output the hidden tensor of the language representation encoder; When the internal tensor of the language representation encoder flows, it includes multiple residual connections, namely the residual connection from the input embedding layer to the encoder self-output layer and the residual connection from the self-output layer to the language representation output layer.
4. The method according to claim 1, wherein The reaction condition decoder includes: a masked multi-head attention network, a multi-head attention mechanism layer, and a feed-forward neural network, The calling of the masked multi-head attention network in the reaction condition decoder to perform attention representation on the reaction condition sequence to obtain the corresponding conditional sequence features, and based on the hidden tensor and the conditional sequence feature tensor, obtaining the predicted reaction condition sequence includes: Call the masked multi-head attention network to perform attention weight learning on the reaction condition sequence to obtain masked conditional sequence features; Call the multi-head attention mechanism layer to perform attention learning on the masked conditional sequence features based on the attention weights to obtain initial predicted conditional sequence features; Call the feed-forward neural network to perform forward learning on the initial predicted conditional sequence features to obtain the predicted reaction condition sequence.
5. The method according to claim 1, wherein The calculating of the loss value of the to-be-trained reaction condition prediction model based on the reaction condition sequence and the predicted reaction condition sequence includes: Based on the reaction condition sequence and the predicted reaction condition sequence, calculate a cross-entropy loss value, and use the cross-entropy loss value as the loss value of the to-be-trained reaction condition prediction model.
6. The method according to claim 1, wherein The to-be-trained reaction condition prediction model further includes: a temperature decoder, After the calling of the masked multi-head attention network in the reaction condition decoder to perform attention representation on the reaction condition sequence to obtain the corresponding conditional sequence features, and based on the hidden tensor and the conditional sequence feature tensor, obtaining the predicted reaction condition sequence, it further includes: Call the temperature decoder to process the predicted reaction condition sequence according to the attention weights to obtain the predicted reaction temperature corresponding to the SMILES sequence of the chemical reaction expression.
7. The method according to claim 6, wherein The calculating of the loss value of the to-be-trained reaction condition prediction model based on the reaction condition sequence and the predicted reaction condition sequence includes: Based on the reaction condition sequence and the predicted reaction condition sequence, calculate a cross-entropy loss value; Based on the predicted reaction temperature, calculate a mean squared error loss value; Based on the cross-entropy loss value, the mean squared error loss value, and a balance coefficient, calculate the loss value of the to-be-trained reaction condition prediction model.
8. A reaction condition prediction model training device, characterized in that, The device includes: A pre-training sample acquisition module, configured to acquire pre-training samples of a language representation network layer based on a reaction center masking strategy, where the pre-training samples include: SMILES sequences of chemical reaction expressions; The language representation network layer optimization module is used to identify the chemical reaction centers in the SMILES sequence of the chemical reaction expression, take a masking rate of 0.5 for the corresponding vocabulary, and a masking probability of 0.15 for the remaining vocabulary, call the language representation network layer to predict the masked vocabulary, calculate the cross-entropy loss, and optimize the language representation network layer; The model training sample acquisition module is used to acquire the model training samples of the reaction condition prediction model, and the model training samples include: the SMILES sequence of the chemical reaction expression and the reaction condition sequence; The model training sample input module is used to input the model training samples into the reaction condition prediction model to be trained, and the reaction condition prediction model to be trained includes: a pre-trained language representation network layer and a reaction condition decoder; The hidden tensor acquisition module is used to call the language representation network layer to process the SMILES sequence of the chemical reaction expression, and obtain the hidden tensor representing the chemical reaction sequence with attention weight information corresponding to each word in the SMILES sequence of the chemical reaction expression; The predicted reaction condition sequence acquisition module is used to call the masked multi-head attention network in the reaction condition decoder to perform attention representation on the reaction condition sequence to obtain the corresponding conditional sequence features, and based on the hidden tensor and the conditional sequence feature tensor, obtain the predicted reaction condition sequence; The loss value calculation module is used to calculate the loss value of the reaction condition prediction model to be trained based on the reaction condition sequence and the predicted reaction condition sequence; The reaction condition prediction model acquisition module is used to, when the loss value is within a preset range, use the trained reaction condition prediction model to be trained as the final reaction condition prediction model.
9. The device according to claim 8, characterized in that, The reaction condition prediction model to be trained further includes: a feature extraction network layer and a vector conversion layer, The device further includes: The position feature acquisition module is used to call the feature extraction network layer to extract each word in the SMILES sequence of the chemical reaction expression and the position feature of each word in the SMILES sequence of the chemical reaction expression; The position vector acquisition module is used to call the vector conversion layer to perform vector conversion processing on each word and the position feature respectively, and obtain the word vector corresponding to each word and the position vector corresponding to the position feature.
10. The device according to claim 9, characterized in that, The language representation network layer includes: a multi-head attention mechanism layer, a layer normalization and Dropout layer, a feed-forward neural network, and a language representation encoder, The hidden tensor acquisition module includes: The association relationship acquisition unit is used to call the multi-head attention mechanism layer to process the sum-normalized result of the word vector and the position vector, and obtain the association relationship between each word; The hidden tensor acquisition unit is used to call the layer normalization and Dropout layer to process the word vector and the position vector according to the association relationship, and obtain the hidden tensor with the weights between each word; A feed-forward network tensor acquisition unit, configured to obtain a feed-forward neural network hidden tensor by passing the hidden tensor with weights between each word through a feed-forward neural network; An encoder tensor acquisition unit, configured to call the language representation output layer to perform a fusion process on the weights between each word and the forward processing result, and obtain and output the hidden tensor of the language representation encoder; When tensors flow inside the language representation encoder, it includes multiple residual connections, namely the residual connection from the input embedding layer to the encoder self-output layer and the residual connection from the self-output layer to the language representation output layer.
Citation Information
Patent Citations
Neural machine translation method based on pre-training bilingual word vector
CN113297841A
Chemical reaction yield prediction method based on causal discovery and multi-structure information coding
CN113470758A