A method, model and model construction method for predicting backbone reactivity
By building a backbone reactant prediction model, filtering out reagent compounds and using AI deep neural network training, the problems of high repeatability and low diversity of chemical reactant prediction results in existing technologies are solved, and more efficient chemical retrosynthesis route design is achieved.
Patent Information
- Application Number
- CN202211518548.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-30
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2042-11-30
AI Technical Summary
Existing chemical reactant prediction models have problems with high repeatability, small differences, and low diversity in the retrosynthesis process, resulting in the failure to achieve the diversity effect expected by chemists.
By filtering out the reagent compounds in the original chemical reaction formula, using the reagent compound restriction data table and atom mapping technology, a backbone reactant prediction model is constructed, and AI deep neural networks such as Transformer, GPT2, and GNN are used for training to generate highly diverse backbone reactant prediction results.
At the same resource cost, the diversity of the backbone reactant prediction model was significantly improved, the number of candidate reactions generated increased by about 2 times, and the diversity of route search results of the AI chemical retrosynthesis system was improved.
Smart Images

Figure CN115862757B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the interdisciplinary field of chemistry and artificial intelligence, and in particular to a prediction method, a model and a model construction method for backbone reactants. Background Art
[0002] The retrosynthesis prediction model (reactant prediction model with target product as input) proposed in the prior art (such as CN202110846347.8) includes but is not limited to the use of technical solutions based on AI deep neural networks such as Transformer, GPT2, and GNN. In the practical process of chemical reactant retrosynthesis prediction, there are many problems such as high repeatability, small differences, and low diversity in the prediction results. Specifically, taking the patent CN202110846347.8 (invention name “Method and device for synthesizing target products by using neural networks”) as an example, the reactant prediction model trained by the technical solution described in the patent is usually given a target product as input. The model predicts the output of the Top-15 candidate reactant combinations. After trunk deduplication, the remaining candidate reactants with different trunks are only 2-3 on average. Due to this limitation, the chemical retrosynthesis route design software and devices further implemented based on such technical solutions have very limited differences between the several candidate routes finally generated, and usually fail to achieve the diversity effect expected by chemists. Therefore, solving the problem of low diversity of prediction results from the root of the reactant prediction model is the key to improving the diversity of retrosynthesis prediction results. Summary of the Invention
[0003] The technical solutions commonly used in existing chemical reactant prediction technologies, including but not limited to AI deep neural network solutions such as Transformer, GPT2, and GNN, have a relatively prominent problem of insufficient diversity in practice. The main reason for this is that among several groups of reactant combinations predicted by the model, there are situations where the backbone reactants are repeated in large quantities and only a small number of reagent compounds have slight differences. These different prediction results are not much different in the chemical sense and are usually regarded as the same reaction in the eyes of chemists. Therefore, in order to improve the diversity of the prediction results of the chemical reactant prediction model and more efficiently generate more different candidate reactions for AI chemical retrosynthesis at the same resource cost, the present invention proposes a prediction method, model and model construction method for backbone reactants to automatically obtain backbone reactants with high diversity.
[0004] In a first aspect, the present invention discloses a method for predicting backbone reactants, which uses a backbone reactant prediction model to perform single-step reverse prediction of backbone reactants, that is, inputting a target product into the backbone reactant prediction model, and the backbone reactant prediction model outputs a predicted backbone reactant after prediction;
[0005] The backbone reactant prediction model is an AI model obtained through training and verification using an AI deep neural network solution and SMILES of several backbone chemical reaction formulas as training and verification sets;
[0006] The SMILES of the backbone chemical reaction formula is obtained by the following method: filtering the SMILES of the reagent compounds in the reagent compound restriction data table contained in the SMILES of the original complete chemical reaction formula to obtain the SMILES of the preliminary backbone chemical reaction formula; filtering out compounds that do not contribute atoms in the chemical reaction by atomic mapping comparison between reactants and products in the SMILES of the preliminary backbone chemical reaction formula, removing the SMILES of the compounds that do not contribute atoms from the SMILES of the preliminary backbone chemical reaction formula, and obtaining the SMILES of the backbone chemical reaction formula containing only backbone reactants and products; the reagent compound restriction data table is a collection of SMILES of several reagent compounds.
[0007] In some embodiments, the number of SMILES of the backbone chemical reaction formula is 1 million data points or more, 3 million data points or more, or 5 million data points or more; the number of SMILES of the reagent compound is 1 million data points or more, 3 million data points or more, or 5 million data points or more.
[0008] In some embodiments, the AI deep neural network solution includes but is not limited to one or more of Transformer, GPT2, and GNN.
[0009] In some embodiments, SMILES include, but are not limited to, one or more of randomized SMILES, canonical SMILES, and Root-aligned SMILES (R-SMILES).
[0010] In a second aspect, the present invention further discloses a backbone reactant prediction model, which uses an AI deep neural network solution and adopts SMILES of several backbone chemical reaction formulas as training sets and validation sets. The AI model obtained through training and validation is used for single-step reverse prediction of backbone reactants, that is, the target product is input and predicted by the backbone reactant prediction model, and the predicted backbone reactant is output;
[0011] The SMILES of the backbone chemical reaction formula is obtained by the following method: filtering the SMILES of the reagent compounds in the reagent compound restriction data table contained in the SMILES of the original complete chemical reaction formula to obtain the SMILES of the preliminary backbone chemical reaction formula; filtering out compounds that do not contribute atoms in the chemical reaction by atomic mapping comparison between reactants and products in the SMILES of the preliminary backbone chemical reaction formula, removing the SMILES of the compounds that do not contribute atoms from the SMILES of the preliminary backbone chemical reaction formula, and obtaining the SMILES of the backbone chemical reaction formula containing only backbone reactants and products; the reagent compound restriction data table is a collection of SMILES of several reagent compounds.
[0012] In some embodiments, the number of SMILES of the backbone chemical reaction formula is 1 million data points or more, 3 million data points or more, or 5 million data points or more; the number of SMILES of the reagent compound is 1 million data points or more, 3 million data points or more, or 5 million data points or more.
[0013] In some embodiments, the AI deep neural network solution includes but is not limited to one or more of Transformer, GPT2, and GNN.
[0014] In a third aspect, the present invention further discloses a method for constructing a backbone reactant prediction model, comprising the following steps:
[0015] Filtering the SMILES of reagent compounds in the reagent compound restriction data table contained in the SMILES of the original complete chemical reaction formula to obtain a SMILES of a preliminary backbone chemical reaction formula; filtering out compounds that do not contribute atoms to the chemical reaction by comparing atomic mappings between reactants and products in the SMILES of the preliminary backbone chemical reaction formula, and removing the SMILES of the compounds that do not contribute atoms from the SMILES of the preliminary backbone chemical reaction formula to obtain a SMILES of a backbone chemical reaction formula containing only backbone reactants and products; the reagent compound restriction data table is a collection of SMILES of several reagent compounds;
[0016] Using an AI deep neural network solution, several SMILES of the backbone chemical reaction formulas were used as training sets and validation sets, and the backbone reactant prediction model was obtained through training and validation.
[0017] In some embodiments, the number of SMILES of the backbone chemical reaction formula is 1 million data points or more, 3 million data points or more, or 5 million data points or more; the number of SMILES of the reagent compound is 1 million data points or more, 3 million data points or more, or 5 million data points or more.
[0018] In some embodiments, the AI deep neural network solution includes but is not limited to one or more of Transformer, GPT2, and GNN.
[0019] In a fourth aspect, the present invention provides a chip, comprising: a processor for calling and running a computer program from a memory, so that a device equipped with the chip executes: the backbone reactant prediction method as described above.
[0020] In a fifth aspect, the present invention further provides a computer-readable storage medium having a computer program stored thereon, which implements the functions of the backbone reactant prediction method as described above when executed by a processor.
[0021] The present invention filters out the reagent compounds in the original complete chemical reaction formula through two methods: reagent compound restriction data table and atom mapping, so that the interference of reagent compounds is avoided during the training of the backbone reactant prediction model, and the focus can be on the prediction of backbone reactants, thereby effectively improving the backbone diversity of the final reactant prediction results.
[0022] The present invention enables the trained backbone reactant prediction model to generate more candidate reactions with different backbones while consuming the same resource cost. This significantly and efficiently enhances the diversity of single-step reverse reaction predictions and further route search results generated by the AI chemical retrosynthesis system by approximately twice the average of existing technologies, making the AI chemical retrosynthesis system more helpful to chemists in designing synthetic routes. The backbone reactant prediction model of the present invention is an AI model within the AI chemical retrosynthesis system.
[0023] The concept, specific structure and technical effects of the present invention will be further described below in conjunction with the accompanying drawings to fully understand the purpose, characteristics and effects of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 It is a schematic diagram of the overall process of the main chemical reaction extraction.
[0025] Figure 2 This is an example of generating a backbone chemical reaction formula from an original complete chemical reaction formula.
[0026] Figure 3It is a schematic diagram comparing the technical effects of the backbone reactant prediction model involved in the present invention and the common inverse prediction model.
[0027] Figure 4 This is an example of the prediction results of a common inverse prediction model.
[0028] Figure 5 This is an example of the prediction results of the backbone reactant prediction model involved in the present invention. DETAILED DESCRIPTION
[0029] In order to make the technical means, creative features, objectives and effects of the invention easier to understand, the invention is further described below with reference to specific diagrams. However, the invention is not limited to the following implementation cases.
[0030] It should be noted that the structures, proportions, sizes, etc. illustrated in the drawings in this specification are only used to match the contents disclosed in the specification so that people familiar with this technology can understand and read them. They are not used to limit the conditions under which the present invention can be implemented. Therefore, they have no substantive technical significance. Any modification of the structure, change in the proportion relationship or adjustment of the size should still fall within the scope of the technical content disclosed in the present invention without affecting the efficacy and purpose that can be achieved by the present invention.
[0031] SMILES (simplified molecular-input line-entry specification) is a string representation of molecules generated by a depth-first traversal of a molecular graph. Due to its readability and ease of use, it has been widely used in the field of reaction prediction. Because SMILES is generated by a depth-first traversal, a molecule can often be represented by multiple valid SMILES representations through enumeration, which are called randomized SMILES. Canonical SMILES uses a canonicalization algorithm to ensure that a molecule can only be represented by a single canonical SMILES. R-SMILES uses SMILES representation by aligning the root atoms of the input and output.
[0032] In this article, "reactants" should be understood as reactants in a broad sense, including not only the main reactants involved in the composition of product molecules, but also catalyst compounds and reagent compounds.
[0033] Example 1 Extraction of the main chemical reaction formula
[0034] The backbone chemical reaction formula is extracted using the backbone reaction extraction module. This module preprocesses all chemical reaction data before model training, including reagent compound restriction data tables and atom mapping equipment / tools. The following example illustrates how the backbone chemical reaction formula is extracted.
[0035] Figure 1 The extraction process of the backbone chemical reaction formula of the present invention is shown. ABcd->E represents an original complete chemical reaction formula, wherein A and B represent two backbone reactants, and c and d represent two reagent compounds. A, B, c, and d can react to produce the target product E. First, through the reagent compound restriction data table, the reactants A, B, c, and d in the original complete chemical reaction formula are compared one by one with the reagent compound restriction data table, thereby identifying that c is a known reagent compound and filtering it. At this point, the remaining chemical reaction formula is ABd->E, wherein compound d is not present in the reagent compound restriction data table, so it cannot be successfully identified and filtered. Therefore, in order to make up for such deficiencies that may exist in the reagent compound restriction data table, atomic mapping machine equipment / tool software is further used to calculate and compare the mapping relationship of the main atoms between the reactants and the target product before and after the chemical reaction, and identify compounds in the reactants that do not contribute important atoms to the reaction. These compounds are also regarded as non-backbone reactants (usually a reagent compound) and filtered. Finally, we can get the main chemical reaction formula containing only the main reactants and the target products, which is AB->E in this example. Figure 1 As shown, after preprocessing all chemical reaction data (original complete chemical reaction formulas) as described above, the target product is used as the input of the machine learning model (or AI model), and the backbone reactants are used as the output of the machine learning model for model training, ultimately obtaining a highly diverse backbone reactant prediction model. The chemical reaction formulas and compounds here are all represented in SMILES format to facilitate the extraction of backbone chemical reaction formulas and the training and verification of machine learning models. Machine learning models include AI translation models (Transformer), GPT2, and GNN.
[0036] Figure 2 The extraction process and results of the backbone chemical reaction formula of a specific original complete chemical reaction formula are presented in the form of reverse reaction. Figure 2 In the original complete chemical reaction formula (hereinafter referred to as the original reaction) includes products, backbone reactants, reagents / solvents, etc. Using the reagent compound restriction data table and atom mapping technology, the reagents / solvents are identified and filtered out from the original reaction and removed. The backbone chemical reaction formula (hereinafter referred to as the backbone reaction) is used as data for subsequent backbone reactant prediction model training and validation.
[0037] Example 2 Training and Verification of Backbone Reactant Prediction Model
[0038] The training of the backbone reactant prediction model is achieved through a single-step reverse prediction model training module. The single-step reverse prediction model training module is a module that trains the backbone reactant prediction model by taking the target product of the backbone reaction as the model input and the backbone reactant as the model prediction output on the basis of the preprocessing of the chemical reaction data by the backbone reaction extraction module. In terms of technical solutions, AI deep neural network solutions including but not limited to Transformer, GPT2, GNN, etc. can be used. Taking the Transformer-based translation model as an example, the SMILES text sequence of the target product is used as the model input. After model prediction, the SMILES text sequence of the backbone reactant is obtained.
[0039] The model is trained using millions of chemical reaction data as the training set, and verified using hundreds of chemical reactions as the test set.
[0040] Example 3 Comparison of the Backbone Reactant Prediction Model and the Conventional Inverse Prediction Model
[0041] Millions of chemical reaction data are used for model training, and the control variable method is used to fairly compare the different effects of the backbone inverse prediction model proposed in this invention, which is trained using backbone chemical reaction data (SMILES of backbone chemical reaction formulas), and the ordinary inverse prediction model, which is trained using complete chemical reaction data (SMILES of original complete chemical reaction formulas), on the diversity of prediction results.
[0042] Specifically, first use the complete chemical reaction data including solvents or reagents (such as Figure 2 A standard reverse prediction model was trained using the original reaction data (shown in the original reaction data) as a control group. Based on this data, without any data filtering, addition, or other processing, the backbone chemical reaction formulas were extracted using the method described in Example 1. Each chemical reaction was processed and replaced with the extracted backbone chemical reaction data. All model training parameters were kept consistent, and the backbone reverse prediction model was trained again as the experimental group.
[0043] After the control group and experimental group models are trained, they are evaluated using a test set consisting of several chemical reactions that have not appeared in the training set. Figure 3 As shown in the figure, taking a chemical reaction as an example, the control group model and the experimental group model are used respectively. Based on the input target product, the candidate reactions that can synthesize the target product are reversely predicted. It should be pointed out that for the convenience of explanation, Figure 3In this example, each model only predicts the top 5 candidate reactions (i.e., only the top 5 candidate reactions with the highest scores). However, in our actual evaluation process, which will be described later, we predict the top 15 candidate reactions. However, the principles for comparing and grouping the top 5 and top 15 are the same. Figure 3 The T in the equation represents the product, the capital letters A, B, and C represent different backbone reactants, and r represents the reagent. The different subscripts of A, B, C, and r indicate that these different reactants or reagents can replace each other, and the chemical reactions are all feasible. It can be seen that after grouping, the inverse prediction model of the control group produced a total of 3 different feasible reactions, but because two of the reactions only differed in the reagents and had the same backbone reactants, they were grouped together and regarded as the same reaction, ultimately obtaining 2 different feasible reactions. At the same time, the backbone inverse prediction model of the experimental group produced a total of 4 different feasible reactions. Since the backbone inverse prediction model tends not to predict reagent compounds, it largely avoids the possibility of predicting 2 reactions with only minor differences in reagent compounds, as in the control group, thereby more efficiently predicting 4 chemical reactions with different backbones.
[0044] In order to further illustrate the effectiveness of the method for predicting backbone reactants proposed in the present invention, we further illustrate it with a more specific example. Figure 4 and Figure 5 The results show the evaluation results of the control group and experimental group models on a real response in our test set. Figure 4 is the control group; the prediction results of the ordinary inverse prediction model, Figure 5 is the prediction result of the backbone inverse prediction model. The following is an explanation:
[0045] Control group: Figure 4After the control group's inverse prediction model generated 15 raw candidate chemical reactions, we first performed a filtering step called "backbone deduplication." This involves extracting the backbone chemical formulas from two reactions using the aforementioned method and comparing them to see if they match. If so, the backbones of the two reactions are duplicated. In this case, "backbone deduplication" discards one reaction and retains the other. After "backbone deduplication," only four distinct reactions remained from the 15 original candidate reactions, demonstrating a high rate of backbone duplication in the control group's standard inverse prediction model. Next, chemists were invited to further review these four distinct reactions, using their expertise to determine whether any of them could be grouped together. Ultimately, the chemists grouped the four reactions into three distinct groups. Therefore, the control group's standard inverse prediction model demonstrated its performance by generating three distinct viable candidate reactions in the top 15 predictions.
[0046] Experimental group: Figure 5 Similarly, after the experimental group's backbone reverse prediction model predicted 15 original candidate reactions, it also first performed "backbone deduplication" and obtained 13 candidate chemical reactions with different backbone chemical reaction formulas. This shows that the backbone reverse prediction model's prediction results have a very low backbone duplication rate. Then, after similar review and grouping by chemists, they ultimately obtained 6 different feasible candidate chemical reactions, twice as many as the control group.
[0047] The above is an explanation of some specific examples. The results are not isolated cases. In actual application scenarios, we selected hundreds of orders of magnitude of chemical reactions to form a test set for fair comparative testing, and randomly sampled the results and submitted them to chemists for review. The final comparison results are: Compared with ordinary reverse prediction models, the backbone reverse prediction model predicts candidate reactions for target products and performs trunk deduplication on the results under the premise of the same parameter configuration and resource consumption. An average of 2.34 times the number of candidate reactions with different trunks can be obtained. Furthermore, the number of correct reactions confirmed by chemists can reach 1.93 times that of the latter. Therefore, the backbone reactant prediction method proposed in the present invention can significantly improve the diversity of the final backbone reverse prediction results without increasing additional computing costs.
[0048] The preferred embodiments of the present invention have been described in detail above. It should be understood that numerous modifications and variations based on the concepts of the present invention are possible without inventive effort by those skilled in the art. Therefore, any technical solution that can be derived by one skilled in the art through logical analysis, reasoning, or limited experimentation based on the concepts of the present invention and the prior art should be within the scope of protection defined by the claims.
Claims
1. A method for predicting backbone reactants, characterized in that: Single-step reverse prediction of backbone reactants using a backbone reactant prediction model, i.e., inputting a target product into the backbone reactant prediction model, and the backbone reactant prediction model outputting a predicted backbone reactant after prediction; The backbone reactant prediction model is an AI model obtained through training and verification using an AI deep neural network solution and SMILES of several backbone chemical reaction formulas as training and verification sets; The SMILES of the backbone chemical reaction formula is obtained by the following method: filtering the SMILES of the reagent compounds in the reagent compound restriction data table contained in the SMILES of the original complete chemical reaction formula to obtain the SMILES of the preliminary backbone chemical reaction formula; Filtering out compounds that do not contribute atoms to the chemical reaction by comparing atomic mappings between reactants and products in the SMILES of the preliminary backbone chemical reaction formula, and removing the SMILES of the compounds that do not contribute atoms from the SMILES of the preliminary backbone chemical reaction formula to obtain a SMILES of the backbone chemical reaction formula containing only backbone reactants and products; The reagent compound restriction data table is a collection of SMILES of several reagent compounds.
2. The backbone reactant prediction method according to claim 1, wherein: The number of SMILES of the backbone chemical reaction formula is 1 million data points or more; the number of SMILES of the reagent compound is 1 million data points or more.
3. The backbone reactant prediction method according to claim 1, wherein: The AI deep neural network solution is selected from one of Transformer, GPT2, and GNN.
4. A backbone reactant prediction model, characterized in that: To use an AI deep neural network solution, SMILES of several backbone chemical reaction formulas are used as training and validation sets. The AI model obtained through training and validation is used to predict backbone reactants in a single step. That is, the target product is input and predicted by the backbone reactant prediction model, and the predicted backbone reactant is output; The SMILES of the backbone chemical reaction formula is obtained by the following method: filtering the SMILES of the reagent compounds in the reagent compound restriction data table contained in the SMILES of the original complete chemical reaction formula to obtain the SMILES of the preliminary backbone chemical reaction formula; Filtering out compounds that do not contribute atoms to the chemical reaction by comparing atomic mappings between reactants and products in the SMILES of the preliminary backbone chemical reaction formula, and removing the SMILES of the compounds that do not contribute atoms from the SMILES of the preliminary backbone chemical reaction formula to obtain a SMILES of the backbone chemical reaction formula containing only backbone reactants and products; The reagent compound restriction data table is a collection of SMILES of several reagent compounds.
5. The backbone reactant prediction model according to claim 4, wherein: The number of SMILES of the backbone chemical reaction formula is 1 million data points or more; the number of SMILES of the reagent compound is 1 million data points or more.
6. The backbone reactant prediction model according to claim 4, wherein: The AI deep neural network solution is selected from one of Transformer, GPT2, and GNN.
7. A method for constructing a backbone reactant prediction model, characterized in that: The following steps are involved: Filter the SMILES of the reagent compounds in the reagent compound restriction data table contained in the SMILES of the original complete chemical reaction formula to obtain the SMILES of the preliminary backbone chemical reaction formula; By comparing the atomic mappings between the reactants and products in the SMILES of the preliminary backbone chemical reaction formula, compounds that do not contribute atoms to the chemical reaction are filtered out, and the SMILES of the compounds that do not contribute atoms are removed from the SMILES of the preliminary backbone chemical reaction formula to obtain a SMILES of the backbone chemical reaction formula containing only the backbone reactants and products; the reagent compound restriction data table is a collection of SMILES of several reagent compounds; Using an AI deep neural network solution, several SMILES of the backbone chemical reaction formulas were used as training sets and validation sets, and the backbone reactant prediction model was obtained through training and validation.
8. The method for constructing a backbone reactant prediction model according to claim 7, wherein: The number of SMILES of the backbone chemical reaction formula is 1 million data points or more; the number of SMILES of the reagent compound is 1 million data points or more.
9. A chip, characterized in that: The invention comprises a processor configured to call and run a computer program from a memory, so that a device equipped with the chip executes the backbone reactant prediction method according to any one of claims 1 to 3.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the function of the backbone reactant prediction method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Method and apparatus for synthesizing target products by using neural networks
CN114093430A
Method for constructing reagent compound prediction model, and method and device for automatically predicting and complementing chemical reaction reagents
CN113990405A
Deep learning-based inverse synthesis prediction method and device, medium and equipment
CN114220496A