A Small Molecule One-Step Retro-Synthesis Prediction Method Based on an Atomic Feature Transfer Network

By designing a deep learning model RetroAFPNN based on atomic feature transfer network, the problem that existing inverse synthesis prediction tools cannot predict molecules outside the template is solved, and a high-accuracy single-step inverse synthesis prediction is achieved, avoiding the tedious work of template updates.

CN115966263BActive Publication Date: 2025-07-01NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211668855.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-21
Publication Date
2025-07-01
Estimated Expiration
2042-12-21

AI Technical Summary

Technical Problem

Existing reverse synthesis prediction tools cannot predict the reverse synthesis path of drug molecules outside the template, and require frequent update of the template, which is cumbersome.

Method used

A deep learning model RetroAFPNN based on atomic feature transfer network and contrast learning is designed to analyze chemical bonds that are prone to breakage in the target molecule, and then complete single-step inverse synthesis prediction.

Benefits of technology

The inverse synthesis path prediction of molecules outside the template is realized, which improves the accuracy of prediction and avoids the tedious work of template updates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115966263B_ABST
    Figure CN115966263B_ABST
Patent Text Reader

Abstract

The present invention discloses a small molecule single-step retrosynthesis prediction method based on an atomic feature transfer network, designs a molecular single-step retrosynthesis prediction model RetroAFPNN, proposes an identification model for the cleavage site of a target molecule based on the atomic feature transfer network AFPNN, and proposes a reactant recommendation model SR-FC based on a fully connected layer. After testing, the model has a high accuracy rate and can recommend relatively accurate reactants for the target molecule.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer-aided drug research and development, and specifically relates to a small molecule single-step retrosynthesis prediction method based on an atomic feature transfer network. Background Art

[0002] The research and development of drugs for treating specific diseases, from initial laboratory research, clinical trials to final market launch, is a high-investment, high-risk, long-cycle project. Modern drug development aims to accelerate the intermediate process and reduce costs by using machine learning techniques such as target identification and validation, virtual screening, and lead optimization in the drug discovery stage and preclinical stage. Despite significant progress in the past few decades, organic synthesis remains a challenge in drug discovery. In the early days, this task was always completed by experts in the chemical field relying on their senior experience, which required very high background requirements for them. Moreover, limited by the insufficient computing power of the human brain, it took at least 3 hours to recommend a synthesis route relying on expert experience. The purpose of the retrosynthesis plan is to transform the target molecule into more readily available precursors and find effective synthesis routes.

[0003] In recent years, the rapid development of computer-assisted synthetic planning (CASP), especially retrosynthesis prediction, has received extensive attention. It only takes 5-10 minutes to design the corresponding synthesis route for each target molecule; moreover, it can also recommend multiple synthesis routes that require different substrates at the same time, and researchers can make specific selections according to their experimental conditions and needs. However, the most major limitation of current CASP tools is that the single-step retrosynthesis prediction strategies they apply are all template-based prediction methods. This paradigm encodes the rules of chemical reactions into the computer, lacks generalization, and cannot predict the retrosynthesis routes of drug molecules outside the template. Moreover, with the discovery of new knowledge, these templates need to be updated frequently, which is also a very cumbersome task. Therefore, the research and development of template-free single-step retrosynthesis prediction algorithms are more important for the future drug research and development field. The main purpose of the present invention is to develop a template-free single-step retrosynthesis prediction model to assist in the design of future drug retrosynthesis routes.

[0004] Contrary to forward reaction prediction, retrosynthesis is a reverse extrapolation from product molecules to inexpensive and accessible reactants. Retrosynthetic analysis can effectively solve the synthesis problems of complex molecules and promote the development of the science of organic synthesis. In addition, with the progress of experimental techniques in systems biology and the continuous accumulation of experimental data, a large amount of biomedical data has emerged, providing impetus for rationally partitioning biosynthetic design. Deep learning (DL), a subfield of artificial intelligence machine learning (ML), can directly understand and learn the inherent laws and complex representations from raw data. Therefore, new attempts applying deep learning (DL) have gradually come onto the stage, opening up a new paradigm for chemical synthesis research.

[0005] As the development of machine translation has received increasing attention and made it possible to be template-free, some researchers have found that the analogy between machine translation and retrosynthesis is obvious. Currently, most studies on template-free single-step retrosynthesis tasks have developed models for molecular translation based on seq2seq algorithms such as LSTM and Transformer (and their variants). Such methods ignore an important issue, that is, the drug molecule itself, as a graph, contains rich structural information. Therefore, many GNN-based studies have emerged in retrosynthesis technology. It can learn the representation of each atom by recursively passing the information of the molecular graph to aggregate the representation of each atom. However, such generation-based models also have a relatively obvious shortcoming, that is, they lack consideration of the properties of the atoms themselves and the influence brought by the surrounding atoms. However, chemical reactions between molecules occur precisely because some key atoms play an important enough role.

[0006] In the field of organic synthesis, a unique and crucial piece of knowledge is to find the chemical bonds in the target molecule that are easy to break. On the other hand, bond energy is a physical quantity that measures the strength of chemical bonds from an energy factor. Therefore, when distinguishing breakable chemical bonds from other bonds, this is also an important indicator that needs to be considered and cannot be ignored when designing the model. Based on the above two points and the deficiencies in other studies, this application designs a deep learning model RetroAFPNN based on an atomic feature transfer network and contrastive learning to analyze the breakable chemical bonds in the target molecule and then complete its single-step retrosynthesis prediction. Summary of the Invention

[0007] The present invention provides a method for predicting single-step retrosynthesis of small molecules based on an atomic feature transfer network, using a template-free single-step retrosynthesis model RetroAFPNN, which solves the problem that general retrosynthesis tools cannot predict molecules outside the template. And compared with generation-based models, the present invention takes into account the problem of insufficient attention to atoms and achieves a higher accuracy performance.

[0008] To achieve the above object, the technical solution provided by the present invention is as follows:

[0009] A small molecule single-step retrosynthesis prediction method based on an atomic feature transfer network, characterized in that it includes the following steps:

[0010] 1) Use the target molecule cleavage site recognition model to predict the cleavage site

[0011] 1.1) Construct the target molecule cleavage site recognition model

[0012] The target molecule cleavage site recognition model includes two atomic feature transfer network layers and a fully connected layer;

[0013] 1.2) Train the target molecule cleavage site recognition model constructed in step 1.1

[0014] 1.2.1) Data collection

[0015] Collect the chemical reaction data required in the training and testing processes of the target molecule cleavage site recognition model, and divide it into a training set and a testing set according to a ratio;

[0016] 1.2.2) Data processing

[0017] Process all the chemical reaction data obtained in step 1.2.1) into Smiles type data;

[0018] 1.2.3) Construct the initial features of atoms

[0019] For each chemical molecule in the data obtained in step 1.2.2), construct the initial features of each atom in the molecule;

[0020] 1.2.4) Use two layers of atomic feature transfer network layers (Atomic Feature Passing Neural Network, AFPNN) to reconstruct the initial features of the atoms obtained in 1.2.3

[0021] Construct the topological structure diagram of the target molecule. Through two layers of atomic feature transfer network layers (Atomic Feature Passing Neural Network, AFPNN), aggregate the features between other atoms connected to each atom around it to reconstruct the features of the atom, and obtain the reconstructed features of the atom; here the main function of AFPNN is to aggregate the features between other atoms connected to each atom around it to reconstruct the features of the atom.

[0022] 1.2.5) Construct bond features

[0023] Construct the features of all bonds by summing the features of the atoms at both ends of each bond. Each bond forms a sample. Finally, obtain the features of all samples in all molecules and label the samples with positive and negative labels y. Here, it means that for each chemical bond in each molecule, the method of summing the features of the atoms at both ends of the bond is used to construct the features of this bond. Each bond is a sample, some are positive samples and some are negative samples. The judgment basis is whether the chemical bond is a broken bond. If so, it is a positive sample; otherwise, it is a negative sample.

[0024] 1.2.6) Map the bond features to a one-dimensional space through a fully connected layer model

[0025] Use a fully connected layer (FC) to map the bond features constructed in step 1.2.5) to 1 dimension, and obtain the feature results after all bond features are mapped to 1 dimension

[0026] 1.2.7) Negative feedback regulation

[0027] Use the cross-entropy loss function to calculate the feature results obtained in step 1.2.6) The loss between the feature results obtained in step 1.2.6) and the label y obtained in step 1.2.5), and then update the trainable parameters in the target molecule cleavage site recognition model through negative feedback regulation. After multiple trainings, obtain the final target molecule cleavage site recognition model;

[0028] During model training, it is necessary to repeat continuously to update the parameters in the model, so that the gap between the label predicted for the chemical bond in the training set and its true label is minimized to complete the model training.

[0029] 1.3) Use the trained target molecule cleavage site recognition model in step 1.2) to predict the cleavage site of the target molecule;

[0030] 2) Use the synthon-to-reactant conversion model SR-FC to recommend the corresponding reactants

[0031] 2.1) For the target molecule, take the broken bond predicted in step 1) as the center and obtain a substructure with a topological depth of s as the core structure representing the target molecule;

[0032] 2.2) Break the target molecule at the correct cleavage position through a function in Rdkit to form synthons;

[0033] 2.3) Compare the synthons obtained in step 2.2) with their corresponding reactants, count the difference structures between the two, and construct a database of additional groups required when converting synthons to reactants;

[0034] 2.4) Combine the extra groups obtained in step 2.3) in pairs and perform One-Hot encoding to form multiple groups of labels;

[0035] 2.5) Extract the molecular fingerprint features of the core structure of the target molecule obtained in step 2.1) through MACCSkeys, and then construct the function mapping relationship between it and the labels obtained in step 2.4) through two fully connected layers. After iterative training, obtain the conversion model SR-FC from synthons to reactants;

[0036] 2.6) Use the conversion model SR-FC from synthons to reactants obtained in step 2.5) to recommend the corresponding reactants and complete the retrosynthesis prediction.

[0037] Further, in step 1.2.1), there are 50K pieces of chemical reaction data.

[0038] Further, in step 1.2.2), use the algorithm for reading chemical reactions in Rdkit to organize the chemical reaction data collected in step 1.2.1), and process all chemical reaction data into Smiles type data of a unified standard.

[0039] Further, the method for constructing the initial atomic features in step 1.2.3) is specifically as follows:

[0040] For each atom in the chemical molecule, extract the following features and concatenate all the features into a feature vector;

[0041] ① Represent the type of each atom using One-Hot encoding, with a feature length of 23 dimensions;

[0042] ② Calculate the degree of each atom, with a feature length of 1 dimension;

[0043] ③ Determine whether the atom belongs to an aromatic ring, represented by 0 and 1, with a feature length of 1 dimension;

[0044] ④ Calculate the number of hydrogen atoms connected to the atom, with a feature length of 1 dimension;

[0045] ⑤ Calculate the charge number carried by the atom, with a feature length of 1 dimension;

[0046] ⑥ Statistically calculate the atomic mass of the atom, with a feature length of 1 dimension;

[0047] After concatenation, the feature vector length of each atom is the sum of the lengths of the above features, which is 28 dimensions.

[0048] Further, the specific method for reconstruction in step 1.2.4) is as follows:

[0049] A1. Mathematical modeling

[0050] In step 1.2.3), the original features of the atoms in the target molecule are constructed, and the length of the feature vector is 28 dimensions; among them, the i-th atom is represented by A i denoted as, A i =[f1, f2,..., f 28 ;

[0051] The target molecule D = {A1, A2,..., A n}; where n is the number of atoms in the target molecule;

[0052] Use e ij to represent the bond energy between the i-th atom A i and the j-th atom A j ; therefore, if A i and A j are not bonded (i.e., there is no edge connection), then e ij = 0;

[0053] A2. Normalization

[0054] To avoid the drowning of data features, normalize the original features of the constructed atoms:

[0055]

[0056] where n is the total number of atoms in the target molecule;

[0057] A3. Construction of the atomic feature transfer network layer

[0058] Construct an atomic feature transfer network layer to update the original features of all atoms in the target molecule D,

[0059]

[0060] where tanh is the activation function;

[0061] (·, ·) represents the connection relationship;

[0062] W1 represents a trainable weight matrix;

[0063] b represents a trainable bias;

[0064] N(i) is the neighbor of the atom A i in the target molecule D.

[0065] Use a two-layer atomic feature transfer network layer to update the atomic features, and the features of the updated atoms are:

[0066]

[0067] where relu is the activation function.

[0068] Further, the specific method for constructing the bond feature in step 1.2.5) is as follows:

[0069] In the target molecule D, for the bond K ij the feature representation is as follows:

[0070]

[0071] wherein, represents the feature after reconstruction of the i-th atom in molecule D;

[0072] represents the feature after reconstruction of the j-th atom in molecule D;

[0073] represents the concatenation of vectors;

[0074] W2 represents the trainable parameter;

[0075] e ij represents the bond energy magnitude between the i-th atom A i and the j-th atom A j ;

[0076] Further, the cross-entropy loss function adopted in step 1.2.7) is:

[0077]

[0078] wherein, m is the number of samples.

[0079] Further, in step 2.1), the topological depth S takes 1;

[0080] In step 2.5), Rdkit is used to represent the features of the core structure in binary. The MACCSkeys fingerprint developed by MDL company has a total of 166 features, but the total length of MACCSkeys is 167 bits. The 0th bit is a placeholder, and the 1st - 166th bits are molecular feature bits; the molecular fingerprint of the core structure with a topological depth of 1 around the broken bond of the target molecule is extracted by this method, and the length of each molecular fingerprint is 167 bits.

[0081] Meanwhile, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and the special feature is that: when the computer program is executed by a processor, the steps of the above method are implemented.

[0082] An electronic device, the special feature is that: it includes a processor and a computer-readable storage medium; the computer program is stored on the computer-readable storage medium, and when the computer program is run by the processor, the steps of the above method are executed.

[0083] Principle of the present invention:

[0084] The present invention utilizes the Atomic Feature Propagation Neural Network (AFPNN) to design a single-step retrosynthesis prediction model, RetroAFPNN, which can infer the retrosynthesis of complex drug molecules. In the first part, for the identification of the cleavage sites of the target molecule, after reconstructing the features of the atoms in the target molecule using the AFPNN, the features of the bonds are then constructed, and finally, through a fully connected layer, the prediction result of the cleavage sites is output. In the second part, for the recommendation of reactants, a mapping relationship between the core structure with a topological depth of n around the cleavage bond of the target molecule and the reactants is constructed. After the target molecule is cleaved into synthons, the corresponding reactants are recommended.

[0085] Advantages of the present invention:

[0086] The present invention mainly uses the Atomic Feature Propagation Neural Network (AFPNN) to capture the features of atoms and uses bond energy as the weight perturbation between them. Then, it is used to identify the most easily cleaved chemical bonds in the target molecule. Finally, a conversion model from synthons to reactants, SR-FC, is constructed to complete the single-step retrosynthesis prediction of the target molecule. The present invention incorporates professional domain knowledge into machine learning, deeply considers the rules in retrosynthesis from the perspective of the ideas of researchers in completing retrosynthesis work, has a short prediction time, and a higher accuracy. Description of the Drawings

[0087] Figure 1 is the overall prediction flowchart of the present invention;

[0088] Figure 2 is the construction flowchart of the cleavage site recognition model for the target molecule in the present invention;

[0089] Figure 3 is the construction flowchart of the reactant recommendation model SR-FC in the present invention. Detailed Embodiment

[0090] The following further describes the content of the present invention in detail with reference to the drawings and specific embodiments:

[0091] Based on the characteristics of the Atomic Feature Propagation Neural Network, the present invention proposes a method for single-step retrosynthesis prediction of small molecules based on the Atomic Feature Propagation Neural Network. This method is divided into two stages. The first stage is the identification stage of the cleavage sites of the target molecule, and the second stage is the inference stage of the reactants. The present invention has been tested to have high accuracy, providing a certain basis for the development of subsequent retrosynthesis prediction work.

[0092] A specific embodiment of the method for single-step retrosynthesis prediction of small molecules based on the Atomic Feature Propagation Neural Network proposed according to the present invention is as follows:

[0093] The model training proposed by the present invention uses the chemical reaction data set containing 50,000 chemical reactions. According to the ratio of 9:1, the data set is divided into a training set and a test set.

[0094] The present invention can be used for the prediction of single-step retrosynthesis of small molecules. For the original data, the present invention adopts the algorithm for reading chemical reactions in Rdkit to sort out the collected chemical reaction data, processes all chemical reaction data into Smiles type data of a unified standard, and then converts the compound into graph data according to the topological relationship between atoms in the target molecule. At the same time, the original features of each atom are constructed.

[0095] For the graph data of each molecule, combined with the original features of its atoms, the Atomic Feature Propagation Neural Network Layer AFPNN is used to reconstruct the initial features of the atoms.

[0096] Using the reconstructed atoms, the method of feature summation is adopted to construct the features of the corresponding bonds, and the labels of positive and negative samples are marked at the same time.

[0097] Then, a fully connected layer is used to map the features of the bonds to obtain the calculation result.

[0098] Finally, the cross-entropy loss function is used to calculate the loss between the calculation result and the label, and the parameters in the model are trained through negative feedback adjustment according to the loss residual.

[0099] After the training is completed, a small molecule cleavage site recognition model based on the atomic feature transfer network is obtained.

[0100] In order to evaluate the performance of the model, the present invention uses the method of five-fold cross-validation to calculate the accuracy of cleavage site recognition and the accuracy of reactant recommendation on the test set. The test results are shown in Table 1:

[0101] Table 1 Performance display of the cleavage site recognition model

[0102]

[0103] Subsequently, the substructures with a topological depth of 1 centered on the cleavage bond in the target molecules in the training set are statistically analyzed, and the differences between synthons and reactants are statistically analyzed. Then, a mapping function from the core substructure to the additional group is constructed by using two fully connected layers. After training, a conversion model SR-FC from synthons to reactants is obtained.

[0104] In order to evaluate the performance of the model, the present invention tests the accuracy of the model on the test set data. After testing, the recognition rate of SR-FC for the mapping from the core substructure of the target molecule to the reactant reaches 0.895.

[0105] Finally, the accuracy rate of combining the "fracture site recognition model" and the "conversion model from synthons to reactants" was statistically calculated on the test set of the present invention, and it was compared with other relatively advanced models under the same standard. The model developed by the method of the present invention showed excellent performance. The results are shown in Table 2:

[0106] Table 2 Performance display of the small molecule single-step retrosynthesis comprehensive model RetroAFPNN

[0107]

[0108] It can be seen from the test results that the present invention has a relatively high prediction accuracy rate for the small molecule single-step retrosynthesis path. Among them, the recommended accuracy rate of Top5 reached 0.880, showing a significant effect.

[0109] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present invention.

Claims

1. A single-step retrosynthesis prediction method for small molecules based on an atomic feature transfer network, characterized by comprising the following steps: 1) Predict the cleavage sites using a target molecule cleavage site recognition model 1.1) Construct a target molecule cleavage site recognition model The target molecule cleavage site recognition model includes two atomic feature transfer network layers and a fully connected layer; 1.2) Train the target molecule cleavage site recognition model constructed in step 1.1) 1.2.1) Data collection Collect the chemical reaction data required in the training and testing processes of the cleavage site recognition model for target molecules, and divide it into a training set and a testing set according to a certain proportion; 1.2.2) Data processing Process all the chemical reaction data obtained in step 1.2.1) into Smiles type data; 1.2.3) Construct the initial features of atoms For each chemical molecule in the data obtained in step 1.2.2), construct the initial features of each atom in the molecule; 1.2.4) Reconstruct the initial features of atoms obtained in 1.2.3) using two-layer atom feature transfer network layers Construct the topological structure diagram of the target molecule. Through two-layer atom feature transfer network layers, aggregate the features between other atoms connected to each atom around it to reconstruct the features of the atom, and obtain the reconstructed features of the atom; 1.2.5) Construct bond features Construct the features of all bonds by summing the reconstructed features of the atoms at both ends of each bond. Each bond forms a sample. Finally, obtain the features of all samples in all molecules and label the samples with positive and negative labels y; 1.2.6) Map the bond features to a one-dimensional space through a fully connected layer model Use a fully connected layer to map the key features constructed in step 1.2.5) to 1 dimension, and obtain the feature result after all key features are mapped to 1 dimension 1.2.7) Negative feedback regulation Calculate the loss between the feature results obtained in step 1.2.6) using the cross-entropy loss function and the label y obtained in step 1.2.5), and then update the trainable parameters in the target molecule cleavage site recognition model through negative feedback regulation. After multiple trainings, the final target molecule cleavage site recognition model is obtained; 1.3) Use the cleavage site recognition model of the target molecule trained in step 1.2) to predict the cleavage site of the target molecule; 2) Use the synthon-to-reactant conversion model SR-FC to recommend corresponding reactants 2.1) For the target molecule, take the predicted cleavage bond in step 1) as the center, and obtain the substructure with a topological depth of s as the core structure representing the target molecule; 2.2) Break the target molecule at the correct cleavage position through the function in Rdkit to form synthons; 2.3) Compare the synthons obtained in step 2.2) with their corresponding reactants, count the difference structures between the two, and construct a database of additional groups required for the conversion from synthons to reactants; 2.4) Combine the additional groups obtained in step 2.3) in pairs and perform One-Hot encoding to form multiple groups of labels; 2.5) Extract the molecular fingerprint features of the core structure of the target molecule obtained in step 2.1) through MACCSkeys, and then construct the function mapping relationship between it and the labels obtained in step 2.4) through two-layer fully connected layers. After iterative training, obtain the synthon-to-reactant conversion model SR-FC; 2.6) Use the synthon-to-reactant conversion model SR-FC obtained in step 2.5) to recommend corresponding reactants and complete the retrosynthesis prediction.

2. The one-step retrosynthesis prediction method for small molecules based on an atom feature transfer network according to claim 1, wherein: In step 1.2.1), there are 50K pieces of the chemical reaction data.

3. The one-step retrosynthesis prediction method for small molecules based on an atom feature transfer network according to claim 1 or 2, wherein: In step 1.2.2), use the algorithm for reading chemical reactions in Rdkit to organize the chemical reaction data collected in step 1.2.1), and process all the chemical reaction data into Smiles type data of a unified standard.

4. The small molecule one-step retrosynthesis prediction method based on an atomic feature transfer network according to claim 3, wherein The method for constructing the initial atomic features in step 1.2.3) is as follows: For each atom in the chemical molecule, the following features are extracted and all the features are concatenated into a feature vector; ① The type of each atom is represented in the One-Hot encoding manner, and the feature length is 23 dimensions; ② Calculate the degree of each atom, and the feature length is 1 dimension; ③ Determine whether the atom belongs to an aromatic ring, which is represented by 0 and 1, and the feature length is 1 dimension; ④ Calculate the number of hydrogen atoms connected to the atom, and the feature length is 1 dimension; ⑤ Calculate the charge number carried by the atom, and the feature length is 1 dimension; ⑥ Statistically calculate the atomic mass of the atom, and the feature length is 1 dimension; After concatenation, the feature vector length of each atom is the sum of the lengths of the above features, which is 28 dimensions.

5. The small molecule one-step retrosynthesis prediction method based on an atomic feature transfer network according to claim 4, wherein The specific method for reconstruction in step 1.2.4) is as follows: A1. Mathematical modeling The i-th atom is represented by A i where A i = [f1, f2,..., f 28 ; Target molecule D = {A1, A2,..., A n}, where n is the number of atoms in the target molecule; Adopt e ij Indicates the bond energy magnitude between the i-th atom A i and the j-th atom A j ; A2. Normalization Perform normalization processing on the original features of the constructed atoms: where n is the total number of atoms in the target molecule; A3. Construction of the atomic feature transfer network layer Construct an atomic feature transfer network layer to update the original features of all atoms in the target molecule D, where tanh is the activation function; (·, ·) represents the connection relationship; W1 represents a trainable weight matrix; b represents a trainable bias; N(i) is a neighbor of atom A in the target molecule D i ; Use a two-layer atomic feature transfer network layer to update the atomic features. The features of the updated atoms are as follows: where relu is the activation function.

6. The one-step retrosynthesis prediction method for small molecules based on an atomic feature transfer network according to claim 5, wherein, The specific method for constructing the bond features in step 1.2.5) is as follows: In the target molecule D, the bond K ij is characterized as follows: Among them, represents the reconstructed feature of the i-th atom in molecule D; Denote the feature after reconstruction of the j-th atom in molecule D; Denotes the concatenation of vectors; W2 represents a trainable parameter; e ij represents the bond energy between the \(i\)-th atom \(A\) i and the \(j\)-th atom \(A\). j The magnitude of the bond energy is as follows.

7. The small molecule one-step retrosynthesis prediction method based on an atomic feature transfer network according to claim 6, characterized in that, The cross-entropy loss function adopted in step 1.2.7) is: where m is the number of samples.

8. A small molecule single-step retrosynthesis prediction method based on an atomic feature transfer network according to claim 7, characterized in that: In step 2.1), the topological depth S is taken as 1.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.

10. An electronic device, characterized in that: It includes a processor and a computer-readable storage medium; A computer program is stored on the computer-readable storage medium, and when the computer program is run by the processor, it executes the steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Artificial intelligence-based retrosynthesis prediction method and device, equipment and storage medium

    CN111524557A

  • Single-step inverse synthesis method and system based on multi-semantic network

    CN114496105A