A reaction condition recommendation method and device for predicting ranking by yield

By constructing a reaction condition and yield prediction model based on graph convolutional networks and multilayer nonlinear neural networks, the problem of low accuracy in reaction condition prediction in existing technologies is solved, and high-yield reaction conditions are recommended, thereby improving the efficiency and accuracy of chemical synthesis.

CN119324001BActive Publication Date: 2025-11-18烟台国工智能科技有限公司
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411854251.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-11-18
Estimated Expiration
2044-12-17

AI Technical Summary

Technical Problem

Existing methods for predicting reaction conditions have low accuracy, which leads to the need for extensive experimental verification during chemical synthesis, increasing time and resource costs.

Method used

By constructing a reaction condition and yield prediction model based on graph convolutional networks and multilayer nonlinear neural networks, and using cheminformatics tools to process molecular features, reaction condition datasets and yield datasets are generated for training and prediction, and the reaction conditions with the highest yield are recommended.

Benefits of technology

It improves the accuracy of reaction condition prediction, reduces the number of experimental verifications, saves experimental time and resources, and improves the efficiency of chemical synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119324001B_ABST
    Figure CN119324001B_ABST
Patent Text Reader

Abstract

The application discloses a reaction condition recommendation method and device for yield prediction ranking, which comprises the following steps: processing an original data set to generate a sample data set; processing samples with several reaction conditions in the sample data set to retain samples with the maximum yield and generate a reaction condition data set; retaining all samples to generate a reaction yield data set; constructing a reaction condition prediction model based on a graph convolution network and a nonlinear neural network and training the model; constructing a reaction yield prediction model based on a multilayer nonlinear neural network and training the model; inputting a molecular SMILES of a target reaction into the reaction condition prediction model to obtain a reaction condition combination; inputting the molecular SMILES and the molecular SMILES in the reaction condition combination into the reaction yield prediction model to obtain a yield corresponding to the reaction condition combination and sort the yield, and taking the reaction condition combination with the maximum yield as a recommendation result. The application provides a reaction condition with high yield for a user, and simultaneously gives a yield value to indicate the yield upper limit of the reaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of reaction condition prediction technology, and more specifically to a method and apparatus for recommending reaction conditions by ranking based on yield prediction. Background Technology

[0002] In chemical synthesis, target compounds are often obtained through a series of chemical synthesis steps. Improving the yield is crucial during synthesis; only high-yield synthetic routes are accepted and used by chemists. To increase reaction yield, chemists empirically try different reaction conditions until the desired yield is achieved. This process often requires lengthy experiments and significant resources in terms of raw materials, manpower, and time. Furthermore, some reactions cannot achieve high yields despite numerous experiments with various reaction conditions, and such routes are abandoned. If a method existed that could predict the reaction conditions and achievable yield values, it would save chemists considerable experimental time and improve the overall efficiency of chemical synthesis.

[0003] With the development of deep learning, methods for reaction condition prediction have made significant progress. However, existing reaction condition prediction methods have accuracy errors, and they often recommend multiple reaction conditions to improve model hit rate. But users don't know which of these multiple reaction conditions is the best, so multiple experiments are still needed to confirm the prediction results.

[0004] Therefore, how to invent a method to recommend high-yield reaction conditions and predict their corresponding yield values ​​has become an urgent problem to be solved. Summary of the Invention

[0005] Therefore, the present invention provides a method and apparatus for recommending reaction conditions by ranking based on yield prediction. The method ranks the recommended reaction conditions by yield prediction, provides users with high-yield reaction conditions, and gives yield values ​​to illustrate the upper limit of the reaction yield.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for recommending reaction conditions based on yield prediction ranking, comprising:

[0007] The original dataset is filtered and processed to extract the molecular smiles, reaction conditions and yield of the reaction, and a sample dataset is generated.

[0008] The sample dataset containing samples of the same reaction with several reaction conditions is processed, and the sample with the highest yield is retained to generate a reaction condition dataset.

[0009] The samples in the sample dataset that have the same reaction under several reaction conditions are processed, and all samples are retained to generate a reaction yield dataset.

[0010] A reaction condition prediction model is constructed based on graph convolutional networks and nonlinear neural networks; the reaction condition prediction model is trained using the reaction condition dataset to obtain a trained reaction condition prediction model.

[0011] A reaction yield prediction model is constructed based on a multi-layer nonlinear neural network; the reaction yield prediction model is trained using the reaction yield dataset to obtain a trained reaction yield prediction model.

[0012] The target reaction molecules (SMILES) are input into the trained reaction condition prediction model, and the model is used for prediction to obtain the reaction condition combination for the target reaction. The target reaction molecules (SMILES) and the molecules (SMILES) in the reaction condition combination are then input one by one into the trained reaction yield prediction model, and the model is used for prediction to obtain the yield corresponding to the reaction condition combination. The yields of the reaction condition combinations are sorted, and the reaction condition combination with the highest yield is selected as the recommended result.

[0013] As a preferred method for recommending reaction conditions by predicting and ranking based on yield, the reaction condition dataset is generated by taking the reaction molecules SMILES as input and the reaction conditions as labels.

[0014] As a preferred method for recommending reaction conditions by predicting and ranking based on yield, the reaction yield dataset is generated by taking the molecular SMILES of the reaction and reaction conditions as input and the reaction yield as a label.

[0015] As a preferred embodiment of a reaction condition recommendation method that ranks reactions based on yield prediction, the prediction steps of the reaction condition prediction model are as follows:

[0016] The reactant and product molecules in the target reaction are processed using cheminformatics tools to generate cheminformatics features of the reactant molecules and the product molecules; the cheminformatics features of the reactant molecules and the product molecules are then fused to generate a feature representation of the reaction.

[0017] Molecular graphs of reactants and products are constructed, and several molecular graphs are spliced ​​together using an adjacency matrix to obtain the molecular graph of the reaction; the chemical features of each atom in the molecule are extracted and a feature matrix corresponding to the molecular graph of the reaction is formed.

[0018] The cheminformatics features of molecules are processed through a multi-layer nonlinear neural network to obtain a global representation based on molecules.

[0019] The molecular graph of the reaction and the feature matrix are input into a multi-layer graph convolutional network. Local neighbor information between atoms in the molecular graph of the reaction is learned by aggregation. The aggregated atomic representation is processed by an average pooling layer to obtain a local representation based on atoms.

[0020] The global representation and the local representation are combined and processed through several prediction layers to output several reaction conditions; the several reaction conditions are arranged and combined to generate several reaction condition combinations.

[0021] The characteristic representation of the reaction is processed through a multi-layer nonlinear neural network to output a temperature prediction value.

[0022] Several combinations of the aforementioned reaction conditions are input into a multi-layer nonlinear neural network for processing, and the temperature corresponding to each combination of the aforementioned reaction conditions is output to obtain several target reaction conditions for the target reaction.

[0023] As a preferred embodiment of a reaction condition recommendation method that ranks reactions based on yield prediction, the prediction steps of the reaction yield prediction model are as follows:

[0024] The chemical reaction fingerprint is obtained by calculating the molecular SMILES of the target reaction and the SMILES of the reaction conditions.

[0025] Normalize the yield;

[0026] The chemical reaction fingerprint is input into a multilayer nonlinear neural network to obtain the predicted yield.

[0027] The present invention also provides a reaction condition recommendation device based on yield prediction ranking, which, based on the above-mentioned reaction condition recommendation method based on yield prediction ranking, includes:

[0028] The raw dataset processing module is used to filter and process the raw dataset, extract the molecular smiles, reaction conditions and yield of the reaction, and generate a sample dataset.

[0029] The reaction condition dataset generation module is used to process samples with several reaction conditions for the same reaction in the sample dataset, retain the sample with the highest yield, and generate a reaction condition dataset.

[0030] The reaction yield dataset generation module is used to process samples with several reaction conditions for the same reaction in the sample dataset, retain all samples, and generate a reaction yield dataset.

[0031] The reaction condition prediction model construction and training module is used to construct a reaction condition prediction model based on graph convolutional networks and nonlinear neural networks; and to train the reaction condition prediction model using the reaction condition dataset to obtain a trained reaction condition prediction model.

[0032] The reaction yield prediction model construction and training module is used to construct a reaction yield prediction model based on a multi-layer nonlinear neural network; and to train the reaction yield prediction model using the reaction yield dataset to obtain a trained reaction yield prediction model.

[0033] The reaction condition recommendation acquisition module is used to input the molecular SMILES of the target reaction into the trained reaction condition prediction model, perform prediction processing through the trained reaction condition prediction model, and obtain the reaction condition combination of the target reaction; input the molecular SMILES of the target reaction and the molecular SMILES in the reaction condition combination one by one into the trained reaction yield prediction model, perform prediction processing through the trained reaction yield prediction model, and obtain the yield corresponding to the reaction condition combination; sort the yields corresponding to the reaction condition combinations, and take the reaction condition combination with the highest yield as the recommendation result.

[0034] As a preferred embodiment of a reaction condition recommendation device that ranks reactions by yield prediction, the reaction condition dataset generation module takes the reaction molecules SMILES as input and the reaction conditions as labels during the generation of the reaction condition dataset.

[0035] As a preferred embodiment of a reaction condition recommendation device that ranks reactions by yield prediction, the reaction yield dataset generation module takes the molecular SMILES of the reaction and reaction conditions as input and the reaction yield as a label to generate the reaction yield dataset during the process of generating the reaction yield dataset.

[0036] As a preferred embodiment of a reaction condition recommendation device that ranks reactions based on yield prediction, the reaction condition recommendation acquisition module includes a prediction submodule of the reaction condition prediction model comprising:

[0037] The feature representation generation submodule is used to process reactant molecules and product molecules in the target reaction using cheminformatics tools to generate cheminformatics features of the reactant molecules and cheminformatics features of the product molecules; and to fuse the cheminformatics features of the reactant molecules and the cheminformatics features of the product molecules to generate a feature representation of the reaction.

[0038] The reaction molecular diagram and feature matrix generation submodule is used to construct molecular diagrams of reactants and products, and to obtain the molecular diagram of the reaction by splicing several molecular diagrams through an adjacency matrix; it also extracts the chemical features of each atom in the molecule and forms a feature matrix corresponding to the molecular diagram of the reaction.

[0039] The global representation acquisition submodule is used to process the cheminformatics features of molecules through a multi-layer nonlinear neural network to obtain a global representation based on molecules.

[0040] The local representation acquisition submodule is used to input the molecular graph of the reaction and the feature matrix into a multilayer graph convolutional network, and learn the local neighbor information between atoms in the molecular graph of the reaction by aggregation; the aggregated atomic representation is processed by an average pooling layer to obtain an atom-based local representation.

[0041] The reaction condition combination generation submodule is used to combine the global representation and the local representation, process them through several prediction layers, and output several reaction conditions; and to arrange and combine the several reaction conditions to generate several reaction condition combinations.

[0042] The temperature prediction value acquisition submodule is used to process the feature representation of the reaction through a multi-layer nonlinear neural network and output the temperature prediction value.

[0043] The target reaction condition acquisition submodule is used to input several combinations of the reaction conditions into a multi-layer nonlinear neural network for processing, and output the temperature corresponding to each combination of the reaction conditions to obtain several target reaction conditions for the target reaction.

[0044] As a preferred embodiment of a reaction condition recommendation device that ranks reactions based on yield prediction, the reaction condition recommendation acquisition module includes a prediction submodule of the reaction yield prediction model comprising:

[0045] The chemical reaction fingerprint acquisition submodule is used to calculate and obtain the chemical reaction fingerprint based on the molecular SMILES of the target reaction and the SMILES of the reaction conditions.

[0046] The yield normalization submodule is used to normalize the yield.

[0047] The yield prediction acquisition submodule is used to input the chemical reaction fingerprint into a multilayer nonlinear neural network to obtain the predicted yield.

[0048] This invention has the following advantages: It extracts the molecular samples, reaction conditions, and yields of a reaction by filtering the original dataset, generating a sample dataset; it processes samples of the same reaction with multiple reaction conditions within the sample dataset, retaining the sample with the highest yield, generating a reaction condition dataset; it processes samples of the same reaction with multiple reaction conditions within the sample dataset, retaining all samples, generating a reaction yield dataset; it constructs a reaction condition prediction model based on graph convolutional networks and nonlinear neural networks; it trains the reaction condition prediction model using the reaction condition dataset, obtaining a trained reaction condition prediction model; and it constructs a reaction yield prediction model based on a multilayer nonlinear neural network. The method involves training a reaction yield prediction model using the reaction yield dataset to obtain a trained reaction yield prediction model. The target reaction's molecular SMILES are input into the trained reaction condition prediction model, and prediction processing is performed to obtain the combination of reaction conditions for the target reaction. The target reaction's molecular SMILES and the molecular SMILES in the reaction condition combinations are then input one by one into the trained reaction yield prediction model, and prediction processing is performed to obtain the yield corresponding to the reaction condition combination. The yields corresponding to the reaction condition combinations are ranked, and the reaction condition combination with the highest yield is used as the recommendation result. This invention proposes a reaction condition recommendation method based on yield prediction ranking, solving the problem of low reaction condition prediction accuracy leading to an increase in the number of experiments required by users. If the predicted yields for all reaction conditions are low, it indicates that the synthesis rate of the reaction is poor, and the corresponding route can be abandoned directly. This greatly reduces the process that users need to experimentally verify the results predicted by the reaction condition model one by one. This invention models molecular graphs and molecular fingerprints using graph convolutional networks and neural networks respectively, taking into account the local changes and global properties of the reaction, which is beneficial to improving the accuracy of reaction conditions. By using chemical reaction descriptor features as input features for the yield prediction model, it can effectively capture and inform the changes between different reaction conditions. Attached Figure Description

[0049] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings in the following description are merely exemplary, and those skilled in the art can derive other embodiments based on the provided drawings without creative effort.

[0050] The structures, proportions, sizes, etc. illustrated in this specification are only for the purpose of assisting those skilled in the art in understanding and reading the content disclosed herein, and are not intended to limit the conditions under which the present invention can be implemented. Therefore, they have no substantial technical significance. Any modifications to the structure, changes in the proportions, or adjustments to the size, without affecting the effects and objectives that the present invention can produce, should still fall within the scope of the technical content disclosed in the present invention.

[0051] Figure 1 This is a schematic diagram of a reaction condition recommendation method based on yield prediction ranking provided in Embodiment 1 of the present invention;

[0052] Figure 2 This is a schematic diagram of the reaction condition prediction model prediction process in the reaction condition recommendation method based on yield prediction ranking provided in Embodiment 1 of the present invention.

[0053] Figure 3 This is a schematic diagram illustrating the specific implementation framework of the reaction condition prediction model in the reaction condition recommendation method based on yield prediction ranking provided in Embodiment 1 of the present invention.

[0054] Figure 4 This is a schematic diagram of the reaction yield prediction model prediction process in the reaction condition recommendation method that ranks reactions by yield prediction provided in Embodiment 1 of the present invention.

[0055] Figure 5 This is a schematic diagram illustrating the specific implementation framework of the reaction yield prediction model in the reaction condition recommendation method based on yield prediction ranking provided in Embodiment 1 of the present invention.

[0056] Figure 6 This is a schematic diagram of the target reaction in one possible embodiment provided in Embodiment 1 of the present invention;

[0057] Figure 7 This is a schematic diagram of a reaction condition recommendation device architecture based on yield prediction ranking provided in Embodiment 2 of the present invention;

[0058] Figure 8 This is a schematic diagram of the reaction condition prediction model architecture in a reaction condition recommendation device that ranks reactions by yield prediction, as provided in Embodiment 2 of the present invention.

[0059] Figure 9 This is a schematic diagram of the reaction yield prediction model architecture in a reaction condition recommendation device that ranks reactions by yield prediction, as provided in Embodiment 2 of the present invention. Detailed Implementation

[0060] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0061] Example 1

[0062] See Figure 1 Embodiment 1 of the present invention provides a method for recommending reaction conditions based on yield prediction ranking, comprising the following steps:

[0063] S1. Filter the original dataset to extract the molecular SMILES, reaction conditions and yield of the reaction, and generate a sample dataset.

[0064] S2. Process the samples in the sample dataset that have several reaction conditions for the same reaction, retain the sample with the highest yield, and generate a reaction condition dataset.

[0065] S3. Process the samples in the sample dataset that have several reaction conditions for the same reaction, retain all samples, and generate a reaction yield dataset.

[0066] S4. Construct a reaction condition prediction model based on graph convolutional networks and nonlinear neural networks; train the reaction condition prediction model using the reaction condition dataset to obtain a trained reaction condition prediction model.

[0067] S5. Construct a reaction yield prediction model based on a multi-layer nonlinear neural network; train the reaction yield prediction model using the reaction yield dataset to obtain a trained reaction yield prediction model.

[0068] S6. Input the target reaction molecules SMILES into the trained reaction condition prediction model, and perform prediction processing through the trained reaction condition prediction model to obtain the reaction condition combination of the target reaction; input the target reaction molecules SMILES and the molecules SMILES in the reaction condition combination one by one into the trained reaction yield prediction model, and perform prediction processing through the trained reaction yield prediction model to obtain the yield corresponding to the reaction condition combination; sort the yields corresponding to the reaction condition combinations, and take the reaction condition combination with the highest yield as the recommended result.

[0069] In this embodiment, in step S1, the original dataset is filtered to extract the smears, reaction conditions and yield of the reaction, and a sample dataset is generated.

[0070] Specifically, reactions and reaction conditions were extracted from the publicly available coupling reaction dataset from AstraZeneca. These reaction conditions included catalyst, solvent, base, and temperature. The RDKit toolkit was used to filter out unreasonable reactions. Samples were created from the data, including molecular samples, reaction conditions, and yields, and duplicates were removed. Samples with a yield of 0 were deleted to generate a new sample dataset.

[0071] In this embodiment, in step S2, samples with several reaction conditions for the same reaction in the sample dataset are processed, and the sample with the highest yield is retained to generate a reaction condition dataset.

[0072] Specifically, in the sample dataset, multiple reaction conditions for a reaction are counted, and only the sample with the highest yield is retained; the reaction molecules SMILES are used as input, and the reaction conditions are used as labels to obtain the reaction condition dataset.

[0073] In this embodiment, in step S3, samples with several reaction conditions for the same reaction in the sample dataset are processed, all samples are retained, and a reaction yield dataset is generated.

[0074] Specifically, in the sample dataset, samples of the same reaction under different reaction conditions are retained; the corresponding molecular SMIELS are found according to the name of the compound in the reaction condition, and the molecular SMILES of the reaction and reaction conditions are used as input, and the reaction yield is used as the label to generate a reaction yield dataset.

[0075] In this embodiment, in step S4, a reaction condition prediction model is constructed based on graph convolutional networks and nonlinear neural networks; the reaction condition prediction model is trained using the reaction condition dataset to obtain a trained reaction condition prediction model.

[0076] Specifically, a reaction condition prediction model is constructed based on graph convolutional networks and nonlinear neural networks. The reaction condition dataset is divided into a training set and a test set in a ratio of 8:2. The model is trained on the training set for 100 iterations with a batch size of 32. The model loss is the sum of the cross-entropy loss of reaction condition prediction and the mean squared error loss of temperature prediction. The learning rate is set to 0.01. The model with the highest accuracy on the test set during the iteration process is saved as the trained reaction condition prediction model.

[0077] In this embodiment, in step S5, a reaction yield prediction model is constructed based on a multi-layer nonlinear neural network; the reaction yield prediction model is trained using the reaction yield dataset to obtain a trained reaction yield prediction model.

[0078] Specifically, a reaction yield prediction model is constructed based on a multi-layer nonlinear neural network. The reaction yield dataset is divided into a training set and a test set in a ratio of 8:2. The training set is trained for 100 iterations with a batch size of 32. The mean squared error loss is calculated and the model parameters are optimized. The learning rate is set to 0.01. The model with the lowest loss on the test set during the iterations is retained as the trained reaction yield prediction model.

[0079] In this embodiment, in step S6, the molecular SMILES of the target reaction are input into the trained reaction condition prediction model, and the reaction condition combination of the target reaction is obtained through prediction processing by the trained reaction condition prediction model; the molecular SMILES of the target reaction and the molecular SMILES in the reaction condition combination are input one by one into the trained reaction yield prediction model, and the yield corresponding to the reaction condition combination is obtained through prediction processing by the trained reaction yield prediction model; the yields corresponding to the reaction condition combinations are sorted, and the reaction condition combination with the highest yield is taken as the recommended result.

[0080] Specifically, such as Figure 2 and Figure 3 As shown, the prediction steps of the reaction condition prediction model are as follows:

[0081] T1. Process the reactant molecules and product molecules in the target reaction using cheminformatics tools to generate cheminformatics features of the reactant molecules and cheminformatics features of the product molecules; fuse the cheminformatics features of the reactant molecules and cheminformatics features of the product molecules to generate a feature representation of the reaction;

[0082] Specifically, cheminformatics tools are used to process reactant and product molecules in the target reaction to generate molecular fingerprints of the reactants and products. The sum of the molecular fingerprints of the reactants is then spliced ​​with the molecular fingerprints of the products to obtain the molecular fingerprint of the reaction.

[0083] T2. Construct molecular diagrams of reactants and products, and combine several molecular diagrams through adjacency matrices to obtain the molecular diagram of the reaction; extract the chemical features of each atom in the molecule and form a feature matrix corresponding to the molecular diagram of the reaction.

[0084] Specifically, a molecular graph of reactants and products is constructed, represented by an adjacency matrix; the adjacency matrix is ​​then concatenated diagonally to obtain the adjacency matrix of the reaction molecular graph. This generates the chemical characteristics of each atom in the molecule, forming a characteristic matrix corresponding to the adjacency matrix. ;

[0085] T3. The chemical informatics features of molecules are processed through a multi-layer nonlinear neural network to obtain a global representation based on molecules;

[0086] Specifically, the molecular fingerprint is processed through two layers of nonlinear neural networks to obtain a global representation based on molecules;

[0087] The expression for the global representation based on molecules is:

[0088] ;

[0089] In the formula, This is a global representation based on molecules; and These are all parameter matrices of a neural network; and These are all bias vectors of the neural network. It is the ReLU activation function; This is the molecular fingerprint of the reaction.

[0090] T4. Input the molecular graph of the reaction and the feature matrix into a multilayer graph convolutional network, and learn the local neighbor information between atoms in the molecular graph of the reaction by aggregation; process the aggregated atom representation through an average pooling layer to obtain a local representation based on atoms;

[0091] Specifically, the adjacency matrix and atomic feature matrix representing the molecular graph are processed... Layered graph convolutional networks learn local neighbor information between atoms in a molecular graph by aggregating this information. The convolution of the layers is shown below:

[0092] ;

[0093] In the formula, For the first Convolution of layers; For the first -1 layer of convolution; This represents the adjacency matrix after self-linking. It is the identity matrix; for The degree matrix, and They represent the first The parameter matrix and bias matrix of the layer.

[0094] All aggregated atomic representations are passed through an average pooling layer to obtain an atom-based local representation; the expression for the atom-based local representation is:

[0095] ;

[0096] In the formula, This is a local representation based on atoms; This represents the average pooling layer.

[0097] T5. Combine the global representation and the local representation, process them through several prediction layers, and output several reaction conditions; arrange and combine the several reaction conditions to generate several reaction condition combinations.

[0098] Specifically, the local and global representations are concatenated and passed through three prediction layers to predict the catalyst, solvent, and bases. The vector length output by each prediction layer corresponds to the number of compounds in the corresponding chemical reaction conditions. The prediction vector is then obtained by passing the sigmoid function.

[0099] ;

[0100] ;

[0101] ;

[0102] In the formula, , and These represent the prediction vectors for catalyst, base, and solvent, respectively. , and Here are the parameter matrices for the three networks; , and Here are the bias vectors for the three networks. This indicates vector concatenation, where each vector exceeds a specified threshold. , and The compounds corresponding to the positions are the predicted results of chemical environmental reaction conditions. In this embodiment, the three thresholds are set to 0.3, 0.5, and 0.5, respectively.

[0103] T6. The characteristic representation of the reaction is processed through a multi-layer nonlinear neural network to output a predicted temperature value;

[0104] Specifically, the molecular fingerprint of the reaction is processed through a multi-layer nonlinear neural network to output the molecular fingerprint processing result; the expression of the molecular fingerprint processing result is:

[0105] ;

[0106] In the formula, This is the result of molecular fingerprint processing; and These are all parameter matrices of a neural network; and These are all bias vectors of the neural network;

[0107] The one-hot encoding of the reaction conditions is processed through a multi-layer nonlinear neural network to output the encoding result; the expression of the encoding result is:

[0108] ;

[0109] In the formula, The result of the encoding process; and This is the parameter matrix of the neural network; and This is the bias vector of the neural network; , and These are the one-hot encoded vectors for the catalyst, base, and solvent, respectively.

[0110] The molecular fingerprint processing result and the encoding processing result are concatenated, and the temperature prediction value is obtained through a temperature prediction layer; the expression for the temperature prediction value is:

[0111] ;

[0112] In the formula, This is a predicted temperature value; and These represent the parameter matrix and the bias vector, respectively.

[0113] T7. Input the various combinations of reaction conditions into a multi-layer nonlinear neural network for processing, and output the temperature corresponding to each combination of reaction conditions to obtain several target reaction conditions for the target reaction.

[0114] Specifically, the various combinations of reaction conditions are input into the temperature prediction section to predict the temperature corresponding to each combination of reaction conditions, thereby obtaining multiple possible sets of reaction conditions for the target reaction.

[0115] In this embodiment, as Figure 4 and Figure 5 As shown, the prediction steps of the reaction yield prediction model are as follows:

[0116] M1. Calculate the chemical reaction fingerprint based on the molecular SMILES of the target reaction and the SMILES of the reaction conditions.

[0117] Specifically, the chemical reaction fingerprint DRFP is calculated based on the molecular SMILES of the reaction and the SMILES of the chemical environment reaction conditions.

[0118] M2, normalize the yield;

[0119] Specifically, the yield is normalized, and the yield value divided by 100 is used as the label for training the model.

[0120] M3. Input the chemical reaction fingerprint into a multilayer nonlinear neural network to obtain the predicted yield.

[0121] Specifically, the chemical reaction fingerprint DRFP is input into a three-layer nonlinear neural network to obtain the predicted yield:

[0122] ;

[0123] In the formula, The predicted yield; , and These are the weight matrices; , and These are the paranoia vectors.

[0124] In one possible embodiment, an example of a response condition recommendation based on yield prediction ranking is provided as follows:

[0125] The target response to be predicted is as follows Figure 6 As shown, the reaction SMILES is: CN(C)C(=O)c1cc(C2CCCN2)c2oc(N3CCOCC3)cc(=O)c2c1.N#Cc1cc(F)cc(Br)c1>>CN(C)C(=O)c1cc(C2CCCN2c2cc(F)cc(C#N)c2)c2oc(N3CCOCC3)cc(=O)c2c1;

[0126] The reaction smiles were input into the reaction condition prediction model, and the prediction results of a set of reaction condition combinations are shown in Table 1:

[0127] Table 1. Prediction results of reaction condition combinations

[0128]

[0129] The SMILES and different reaction conditions were sequentially input into the yield prediction model, and the results are shown in Table 2:

[0130] Table 2. Reaction Yield Prediction Results

[0131]

[0132] Finally, the optimal reaction conditions for the target reaction were found to be Xantphos catalyst, palladium acetate, cesium carbonate, hexaoxane, and 102.3 °C, respectively, with an achievable yield of 53.6%. Users can choose whether to use the reaction based on whether the maximum yield meets the requirements; if so, they can conduct experiments according to the optimal reaction conditions.

[0133] In summary, this invention generates a sample dataset by filtering and processing the original dataset to extract the molecular samples, reaction conditions, and yields of the reaction; processing samples of the same reaction with multiple reaction conditions within the sample dataset and retaining the sample with the highest yield to generate a reaction condition dataset; processing samples of the same reaction with multiple reaction conditions within the sample dataset and retaining all samples to generate a reaction yield dataset; constructing a reaction condition prediction model based on graph convolutional networks and nonlinear neural networks; training the reaction condition prediction model using the reaction condition dataset to obtain a trained reaction condition prediction model; and constructing a reaction yield prediction model based on a multilayer nonlinear neural network. The reaction yield prediction model is trained using the reaction yield dataset to obtain a trained reaction yield prediction model. The molecular SMILES of the target reaction are input into the trained reaction condition prediction model, and prediction processing is performed to obtain the reaction condition combinations for the target reaction. The molecular SMILES of the target reaction and the molecular SMILES in the reaction condition combinations are input one by one into the trained reaction yield prediction model, and prediction processing is performed to obtain the yield corresponding to the reaction condition combinations. The yields corresponding to the reaction condition combinations are ranked, and the reaction condition combination with the highest yield is used as the recommendation result. This invention proposes a reaction condition recommendation method based on yield prediction ranking, solving the problem of low reaction condition prediction accuracy leading to an increase in the number of experiments required by users. If the predicted yields for all reaction conditions are low, it indicates that the synthesis rate of the reaction is poor, and the corresponding route can be abandoned directly. This greatly reduces the process that users need to experimentally verify the results predicted by the reaction condition model one by one. This invention models molecular graphs and molecular fingerprints using graph convolutional networks and neural networks respectively, taking into account the local changes and global properties of the reaction, which is beneficial to improving the accuracy of reaction conditions. By using chemical reaction descriptor features as input features for the yield prediction model, it can effectively capture and inform the changes between different reaction conditions.

[0134] It should be noted that the method of this disclosure embodiment can be executed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method of this disclosure embodiment, and the multiple devices will interact with each other to complete the method described.

[0135] It should be noted that the above description describes some embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0136] Example 2

[0137] See Figure 7 Embodiment 2 of the present invention also provides a reaction condition recommendation device for yield prediction ranking, comprising:

[0138] The raw dataset processing module 001 is used to filter and process the raw dataset, extract the molecular SMILES, reaction conditions and yield of the reaction, and generate a sample dataset.

[0139] The reaction condition dataset generation module 002 is used to process samples with several reaction conditions for the same reaction in the sample dataset, retain the sample with the highest yield, and generate a reaction condition dataset.

[0140] The reaction yield dataset generation module 003 is used to process samples with several reaction conditions for the same reaction in the sample dataset, retain all samples, and generate a reaction yield dataset.

[0141] The reaction condition prediction model construction and training module 004 is used to construct a reaction condition prediction model based on graph convolutional networks and nonlinear neural networks; and to train the reaction condition prediction model using the reaction condition dataset to obtain a trained reaction condition prediction model.

[0142] The reaction yield prediction model construction and training module 005 is used to construct a reaction yield prediction model based on a multi-layer nonlinear neural network; and to train the reaction yield prediction model using the reaction yield dataset to obtain a trained reaction yield prediction model.

[0143] The reaction condition recommendation acquisition module 006 is used to input the molecular SMILES of the target reaction into the trained reaction condition prediction model, perform prediction processing through the trained reaction condition prediction model, and obtain the reaction condition combination of the target reaction; input the molecular SMILES of the target reaction and the molecular SMILES in the reaction condition combination one by one into the trained reaction yield prediction model, perform prediction processing through the trained reaction yield prediction model, and obtain the yield corresponding to the reaction condition combination; sort the yields corresponding to the reaction condition combinations, and take the reaction condition combination with the highest yield as the recommendation result.

[0144] In this embodiment, the reaction condition dataset generation module 002 takes the reaction molecules SMILES as input and the reaction conditions as labels during the generation of the reaction condition dataset.

[0145] In this embodiment, the reaction yield dataset generation module 003 takes the molecular SMILES of the reaction and reaction conditions as input and the reaction yield as a label to generate the reaction yield dataset during the process of generating the reaction yield dataset.

[0146] In this embodiment, as Figure 8 As shown, in the reaction condition recommendation acquisition module 006, the prediction submodule of the reaction condition prediction model includes:

[0147] The feature representation generation submodule 061 is used to process reactant molecules and product molecules in the target reaction using cheminformatics tools to generate cheminformatics features of the reactant molecules and cheminformatics features of the product molecules; and to fuse the cheminformatics features of the reactant molecules and the cheminformatics features of the product molecules to generate a feature representation of the reaction.

[0148] The reaction molecular diagram and feature matrix generation submodule 062 is used to construct molecular diagrams of reactants and products, and to obtain the molecular diagram of the reaction by splicing several molecular diagrams through an adjacency matrix; it also extracts the chemical features of each atom in the molecule and forms a feature matrix corresponding to the molecular diagram of the reaction.

[0149] The global representation acquisition submodule 063 is used to process the cheminformatics features of molecules through a multi-layer nonlinear neural network to obtain a global representation based on molecules.

[0150] The local representation acquisition submodule 064 is used to input the molecular graph of the reaction and the feature matrix into a multilayer graph convolutional network, and learn the local neighbor information between atoms in the molecular graph of the reaction by aggregation; the aggregated atomic representation is processed by an average pooling layer to obtain an atom-based local representation.

[0151] The reaction condition combination generation submodule 065 is used to combine the global representation and the local representation, process them through several prediction layers, and output several reaction conditions; and to arrange and combine the several reaction conditions to generate several reaction condition combinations.

[0152] Temperature prediction value acquisition submodule 066 is used to process the feature representation of the reaction through a multi-layer nonlinear neural network and output the temperature prediction value.

[0153] The target reaction condition acquisition submodule 067 is used to input several combinations of the reaction conditions into a multi-layer nonlinear neural network for processing, and output the temperature corresponding to each combination of the reaction conditions to obtain several target reaction conditions for the target reaction.

[0154] In this embodiment, as Figure 9 As shown, in the reaction condition recommendation acquisition module 006, the prediction submodule of the reaction yield prediction model includes:

[0155] The chemical reaction fingerprint acquisition submodule 068 is used to calculate and obtain the chemical reaction fingerprint based on the molecular SMILES of the target reaction and the SMILES of the reaction conditions.

[0156] The yield normalization submodule 069 is used to normalize the yield.

[0157] The predicted yield acquisition submodule 0610 is used to input the chemical reaction fingerprint into a multilayer nonlinear neural network to obtain the predicted yield.

[0158] It should be noted that the information interaction and execution process between the modules of the above system are based on the same concept as the method embodiment in Embodiment 1 of this application, and the resulting technical effects are the same as those in the method embodiment of this application. For details, please refer to the description in the method embodiment shown above in this application, and it will not be repeated here.

[0159] Example 3

[0160] Embodiment 3 of the present invention provides a non-transitory computer-readable storage medium storing program code for a reaction condition recommendation method that ranks by yield prediction. The program code includes instructions for executing the reaction condition recommendation method that ranks by yield prediction according to Embodiment 1 or any possible implementation thereof.

[0161] Computer-readable storage media can be any available medium that a computer can access, or a data storage device such as a server or data center that integrates one or more available media. The available medium can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).

[0162] Example 4

[0163] Embodiment 4 of the present invention provides an electronic device, including: a memory and a processor;

[0164] The processor and the memory communicate with each other via a bus; the memory stores program instructions that can be executed by the processor, and the processor can call the program instructions to execute a reaction condition recommendation method for yield prediction ranking, as described in Embodiment 1 or any possible implementation thereof.

[0165] Specifically, a processor can be implemented in hardware or software. When implemented in hardware, the processor can be a logic circuit, an integrated circuit, etc. When implemented in software, the processor can be a general-purpose processor that reads software code stored in memory. This memory can be integrated into the processor or located outside the processor and exist independently.

[0166] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable system. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means.

[0167] It is obvious to those skilled in the art that the modules or steps of the present invention described above can be implemented using general-purpose computing systems. They can be centralized on a single computing system or distributed across a network of multiple computing systems. Optionally, they can be implemented using program code executable by a computing system, thereby storing them in a storage system for execution by the computing system. In some cases, the steps shown or described can be performed in a different order than those presented herein, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.

[0168] Although the present invention has been described in detail above with general descriptions and specific embodiments, modifications or improvements can be made to it, which will be obvious to those skilled in the art. Therefore, all such modifications or improvements made without departing from the spirit of the present invention fall within the scope of protection claimed by the present invention.

Claims

1. A method for recommending reaction conditions based on yield prediction ranking, characterized in that, include: The original dataset is filtered and processed to extract the molecular smiles, reaction conditions and yield of the reaction, and a sample dataset is generated. The reactions and reaction conditions were extracted from the publicly available coupling reaction dataset from AstraZeneca. The reaction conditions included catalyst, solvent, base, and temperature. Unreasonable reactions were filtered out using the RDKit toolkit. Samples were formed by extracting the molecular smiles, reaction conditions, and yields of the reactions from the data, and duplicates were removed. Samples with a yield of 0 were deleted to generate a sample dataset. The sample dataset is processed to obtain a reaction condition dataset by retaining the sample with the highest yield for the same reaction. In the sample dataset, multiple reaction conditions for a reaction are counted, and only the sample with the highest yield is retained. The reaction molecule SMILES is used as input and the reaction conditions are used as labels to obtain the reaction condition dataset. The sample dataset is processed to include samples of the same reaction with several reaction conditions. All samples are retained to generate a reaction yield dataset. In the sample dataset, samples of the same reaction with different reaction conditions are retained. The corresponding molecular SMILES are found according to the name of the compound in the reaction conditions. The molecular SMILES of the reaction and reaction conditions are used as input, and the reaction yield is used as the label to generate a reaction yield dataset. A reaction condition prediction model is constructed based on graph convolutional networks and nonlinear neural networks. The reaction condition prediction model is trained using the reaction condition dataset to obtain a trained reaction condition prediction model. The reaction condition dataset is divided into training and test sets according to a certain ratio. The model is iteratively trained on the training set. The model loss is the sum of the cross-entropy loss of reaction condition prediction and the mean squared error loss of temperature prediction. The model with the highest accuracy on the test set during the iteration process is saved as the trained reaction condition prediction model. A reaction yield prediction model is constructed based on a multi-layer nonlinear neural network. The reaction yield prediction model is trained using the reaction yield dataset to obtain a well-trained reaction yield prediction model. The reaction yield dataset is divided into a training set and a test set according to a certain ratio. Iterative training is performed on the training set, the mean squared error loss is calculated and the model parameters are optimized. The model with the lowest loss on the test set during the iteration is retained as the well-trained reaction yield prediction model. The target reaction's molecular SMILES are input into the trained reaction condition prediction model, and the model is used for prediction to obtain the combination of reaction conditions for the target reaction. The target reaction's molecular SMILES and the molecular SMILES in the reaction condition combinations are then input one by one into the trained reaction yield prediction model, and the model is used for prediction to obtain the yield corresponding to the reaction condition combination. The yields corresponding to the reaction condition combinations are sorted, and the reaction condition combination with the highest yield is used as the recommended result. The prediction steps of the reaction condition prediction model are as follows: T1. Process the reactant molecules and product molecules in the target reaction using cheminformatics tools to generate cheminformatics features of the reactant molecules and cheminformatics features of the product molecules; fuse the cheminformatics features of the reactant molecules and cheminformatics features of the product molecules to generate a feature representation of the reaction; Specifically, cheminformatics tools are used to process reactant and product molecules in the target reaction to generate molecular fingerprints of reactants and products. The sum of the molecular fingerprints of reactants is then spliced ​​with the molecular fingerprints of products to obtain the molecular fingerprint of the reaction. T2. Construct molecular diagrams of reactants and products, and combine several molecular diagrams through adjacency matrices to obtain the molecular diagram of the reaction; extract the chemical features of each atom in the molecule and form a feature matrix corresponding to the molecular diagram of the reaction. Specifically, molecular graphs of reactants and products are constructed, represented by adjacency matrices; these adjacency matrices are then concatenated diagonally to obtain the adjacency matrix of the reaction molecular graph. This generates the chemical characteristics of each atom in the molecule, forming a characteristic matrix corresponding to the adjacency matrix. ; T3. The chemical informatics features of molecules are processed through a multi-layer nonlinear neural network to obtain a global representation based on molecules; Specifically, the molecular fingerprint is processed through two layers of nonlinear neural networks to obtain a global representation based on molecules; The expression for the global representation based on molecules is: ; In the formula, This is a global representation based on molecules; and These are all parameter matrices of a neural network; and These are all bias vectors of the neural network. It is the ReLU activation function; The molecular fingerprint of the reaction; T4. Input the molecular graph of the reaction and the feature matrix into a multilayer graph convolutional network, and learn the local neighbor information between atoms in the molecular graph of the reaction by aggregation; process the aggregated atom representation through an average pooling layer to obtain a local representation based on atoms; Specifically, the adjacency matrix and atomic feature matrix representing the molecular graph are processed... Layered graph convolutional networks learn local neighbor information between atoms in a molecular graph by aggregating this information. The convolution of the layers is shown below: ; In the formula, For the first Convolution of layers; For the first -1 layer of convolution; This represents the adjacency matrix after self-linking. It is the identity matrix; for The degree matrix, and They represent the first The parameter matrix and bias matrix of the layer; All aggregated atomic representations are passed through an average pooling layer to obtain an atom-based local representation; the expression for the atom-based local representation is: ; In the formula, This is a local representation based on atoms; Indicates the average pooling layer; T5. Combine the global representation and the local representation, process them through several prediction layers, and output several reaction conditions; arrange and combine the several reaction conditions to generate several reaction condition combinations. Specifically, the local and global representations are concatenated and passed through three prediction layers to predict the catalyst, solvent, and bases. The vector length output by each prediction layer corresponds to the number of compounds in the corresponding chemical environment reaction conditions. The prediction vector is then obtained by passing the sigmoid function. ; ; ; In the formula, , and These represent the prediction vectors for catalyst, base, and solvent, respectively. , and Here are the parameter matrices for the three networks; , and Here are the bias vectors for the three networks. This indicates vector concatenation, where each vector exceeds a specified threshold. , and The compounds corresponding to the positions are the predicted results of chemical environmental reaction conditions; the three thresholds are set to 0.3, 0.5 and 0.5 respectively; T6. The characteristic representation of the reaction is processed through a multi-layer nonlinear neural network to output a predicted temperature value; Specifically, the molecular fingerprint of the reaction is processed through a multi-layer nonlinear neural network to output the molecular fingerprint processing result; the expression of the molecular fingerprint processing result is: ; In the formula, This is the result of molecular fingerprint processing; and These are all parameter matrices of a neural network; and These are all bias vectors of the neural network; The one-hot encoding of the reaction conditions is processed through a multi-layer nonlinear neural network to output the encoding result; the expression of the encoding result is: ; In the formula, The result of the encoding process; and This is the parameter matrix of the neural network; and This is the bias vector of the neural network; , and These are the one-hot encoded vectors for the catalyst, base, and solvent, respectively. The molecular fingerprint processing result and the encoding processing result are concatenated, and the temperature prediction value is obtained through a temperature prediction layer; the expression for the temperature prediction value is: ; In the formula, This is a predicted temperature value; and These represent the parameter matrix and the bias vector, respectively. T7. Input the various combinations of reaction conditions into a multi-layer nonlinear neural network for processing, and output the temperature corresponding to each combination of reaction conditions to obtain several target reaction conditions for the target reaction. Specifically, the various combinations of reaction conditions are input into the temperature prediction section to predict the temperature corresponding to each combination of reaction conditions, thereby obtaining multiple possible sets of reaction conditions for the target reaction. The prediction steps of the reaction yield prediction model are as follows: M1. Calculate the chemical reaction fingerprint based on the molecular SMILES of the target reaction and the SMILES of the reaction conditions. Specifically, the chemical reaction fingerprint DRFP is calculated based on the molecular SMILES of the reaction and the SMILES of the chemical environment reaction conditions. M2, normalize the yield; Specifically, the yield is normalized, and the yield value is divided by 100 as the label for training the model; M3. Input the chemical reaction fingerprint into a multilayer nonlinear neural network to obtain the predicted yield; Specifically, the chemical reaction fingerprint DRFP is input into a three-layer nonlinear neural network to obtain the predicted yield: ; In the formula, The predicted yield; , and These are the weight matrices; , and These are the paranoia vectors.

2. A reaction condition recommendation device for ranking based on yield prediction, employing the reaction condition recommendation method for ranking based on yield prediction as described in claim 1, characterized in that, include: The raw dataset processing module is used to filter and process the raw dataset, extract the molecular smiles, reaction conditions and yield of the reaction, and generate a sample dataset. The reaction condition dataset generation module is used to process samples with several reaction conditions for the same reaction in the sample dataset, retain the sample with the highest yield, and generate a reaction condition dataset. The reaction yield dataset generation module is used to process samples with several reaction conditions for the same reaction in the sample dataset, retain all samples, and generate a reaction yield dataset. The reaction condition prediction model construction and training module is used to construct a reaction condition prediction model based on graph convolutional networks and nonlinear neural networks; and to train the reaction condition prediction model using the reaction condition dataset to obtain a trained reaction condition prediction model. The reaction yield prediction model construction and training module is used to construct a reaction yield prediction model based on a multi-layer nonlinear neural network; and to train the reaction yield prediction model using the reaction yield dataset to obtain a trained reaction yield prediction model. The reaction condition recommendation acquisition module is used to input the molecular SMILES of the target reaction into the trained reaction condition prediction model, perform prediction processing through the trained reaction condition prediction model, and obtain the reaction condition combination of the target reaction; input the molecular SMILES of the target reaction and the molecular SMILES in the reaction condition combination one by one into the trained reaction yield prediction model, perform prediction processing through the trained reaction yield prediction model, and obtain the yield corresponding to the reaction condition combination; sort the yields corresponding to the reaction condition combinations, and take the reaction condition combination with the highest yield as the recommendation result.

3. The reaction condition recommendation device for yield prediction and ranking according to claim 2, characterized in that, In the reaction condition dataset generation module, during the generation of the reaction condition dataset, the reaction molecules SMILES are used as input and the reaction conditions are used as labels to generate the reaction condition dataset.

4. The reaction condition recommendation device for yield prediction and ranking according to claim 3, characterized in that, In the reaction yield dataset generation module, during the process of generating the reaction yield dataset, the molecular SMILES of the reaction and reaction conditions are used as input, and the reaction yield is used as the label to generate the reaction yield dataset.

5. The reaction condition recommendation device for yield prediction and ranking according to claim 4, characterized in that, In the reaction condition recommendation acquisition module, the prediction submodule of the reaction condition prediction model includes: The feature representation generation submodule is used to process reactant molecules and product molecules in the target reaction using cheminformatics tools to generate cheminformatics features of the reactant molecules and cheminformatics features of the product molecules; and to fuse the cheminformatics features of the reactant molecules and the cheminformatics features of the product molecules to generate a feature representation of the reaction. The reaction molecular diagram and feature matrix generation submodule is used to construct molecular diagrams of reactants and products, and to obtain the molecular diagram of the reaction by splicing several molecular diagrams through an adjacency matrix; it also extracts the chemical features of each atom in the molecule and forms a feature matrix corresponding to the molecular diagram of the reaction. The global representation acquisition submodule is used to process the cheminformatics features of molecules through a multi-layer nonlinear neural network to obtain a global representation based on molecules. The local representation acquisition submodule is used to input the molecular graph of the reaction and the feature matrix into a multilayer graph convolutional network, and learn the local neighbor information between atoms in the molecular graph of the reaction by aggregation; the aggregated atomic representation is processed by an average pooling layer to obtain an atom-based local representation. The reaction condition combination generation submodule is used to combine the global representation and the local representation, process them through several prediction layers, and output several reaction conditions; and to arrange and combine the several reaction conditions to generate several reaction condition combinations. The temperature prediction value acquisition submodule is used to process the feature representation of the reaction through a multi-layer nonlinear neural network and output the temperature prediction value. The target reaction condition acquisition submodule is used to input several combinations of the reaction conditions into a multi-layer nonlinear neural network for processing, and output the temperature corresponding to each combination of the reaction conditions to obtain several target reaction conditions for the target reaction.

6. The reaction condition recommendation device for yield prediction and ranking according to claim 5, characterized in that, In the reaction condition recommendation acquisition module, the prediction submodule of the reaction yield prediction model includes: The chemical reaction fingerprint acquisition submodule is used to calculate and obtain the chemical reaction fingerprint based on the molecular SMILES of the target reaction and the SMILES of the reaction conditions. The yield normalization submodule is used to normalize the yield. The yield prediction acquisition submodule is used to input the chemical reaction fingerprint into a multilayer nonlinear neural network to obtain the predicted yield.

Citation Information

Patent Citations

  • Synthetic route development method and device integrating inverse synthesis, condition and reaction prediction

    CN118136141A