Drug molecule inverse synthesis method and system for generating expected synthesis path
By converting the target product molecules and reactant molecules into two-dimensional molecular maps, extracting characteristic information and determining candidate reactant molecules, the problem of difficult to generate reasonable molecular structure in the retrosynthesis of drug molecules in the prior art is solved, and an interpretable reverse synthesis method of drug molecules is realized to generate a synthetic path that meets expectations.
Patent Information
- Application Number
- CN202510010302.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-30
AI Technical Summary
Existing drug molecule reverse synthesis techniques are difficult to generate reasonable and expected molecular structures and properties, and the reverse synthesis reasoning process lacks transparency and explanatory nature.
By converting the target product molecules and reactant molecules into two-dimensional molecular maps, characteristic information is extracted, candidate reactant molecules are determined, and the expected reaction path is generated through decision algorithms to achieve reverse synthesis of drug molecules.
An interpretable method for reverse synthesis of drug molecules is provided, which can generate synthesis paths that meet expectations, and improve the transparency and efficiency of the reverse synthesis process.
Smart Images

Figure CN120072106A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of drug molecule retrosynthesis, and in particular, to a drug molecule retrosynthesis method and system for generating an expected synthesis path. Background Art
[0002] As the core of organic chemistry, retrosynthetic analysis provides chemists with a way to obtain novel molecules that are difficult to synthesize in the fields of materials and drug manufacturing. Traditional computer-aided synthesis methods rely on rules or expert knowledge, with high labor costs and limited search scope. Existing technologies usually adopt single-step retrosynthesis or multi-step retrosynthesis methods to achieve drug molecule retrosynthesis. However, these two methods lack effective constraints on reactant molecules, resulting in the inability to generate molecular structures and properties that are both reasonable and meet expectations. Moreover, the reasoning process of retrosynthesis also lacks transparency and interpretability. Summary of the Invention
[0003] The purpose of the present invention is to provide a drug molecule retrosynthesis method and system for generating an expected synthesis path to improve the above problems.
[0004] To achieve the above purpose, the embodiments of the present application provide the following technical solutions:
[0005] On the one hand, the embodiments of the present application provide a drug molecule retrosynthesis method for generating an expected synthesis path, the method comprising:
[0006] Obtaining a target product molecule and a reactant molecule corresponding to the target product molecule;
[0007] Converting the target product molecule and the reactant molecule corresponding to the target product molecule into corresponding two-dimensional molecular graphs to obtain a first molecular graph and a second molecular graph;
[0008] Performing feature extraction on the first molecular graph and the second molecular graph to obtain first feature information and second feature information;
[0009] Determining candidate reactant molecules according to the first feature information and the second feature information;
[0010] Determining the retrosynthesis path of the drug molecule according to the first molecular graph and the two-dimensional molecular graph corresponding to the candidate reactant molecules.
[0011] On the second hand, the embodiments of the present application provide a drug molecule retrosynthesis system for generating an expected synthesis path, the system comprising:
[0012] An obtaining module, configured to obtain a target product molecule and a reactant molecule corresponding to the target product molecule;
[0013] The first processing module is configured to convert the target product molecule and the reactant molecule corresponding to the target product molecule into corresponding two-dimensional molecular graphs, obtaining a first molecular graph and a second molecular graph;
[0014] The feature extraction module is configured to perform feature extraction on the first molecular graph and the second molecular graph, obtaining first feature information and second feature information;
[0015] The second processing module is configured to determine candidate reactant molecules according to the first feature information and the second feature information;
[0016] The third processing module is configured to determine the retrosynthetic pathway of the drug molecule according to the first molecular graph and the two-dimensional molecular graph corresponding to the candidate reactant molecule.
[0017] In a third aspect, an embodiment of the present application provides a device. The drug molecule retrosynthesis device for generating a desired synthesis pathway includes a memory and a processor. The memory is used to store a computer program; the processor is configured to implement the steps of the above-mentioned drug molecule retrosynthesis method for generating a desired synthesis pathway when executing the computer program.
[0018] In a fourth aspect, an embodiment of the present application provides a readable storage medium. A computer program is stored on the readable storage medium, and when the computer program is executed by a processor, the steps of the above-mentioned drug molecule retrosynthesis method for generating a desired synthesis pathway are implemented.
[0019] The beneficial effects of the present invention are as follows:
[0020] The present invention converts the target product molecule and the reactant molecule corresponding to the target product molecule into corresponding two-dimensional molecular graphs, obtaining a first molecular graph and a second molecular graph; further extracts features in the molecular graphs respectively to obtain first feature information and second feature information; determines candidate reactant molecules according to the first feature information and the second feature information, and then determines the similar part between the molecular graph of the candidate reactant molecule and the first molecular graph as synthons in the retrosynthetic process, adds a leaving group at the reaction center to perfect the synthons as product molecules; finally obtains the final expected reaction pathway through a decision algorithm, providing an interpretable drug molecule retrosynthesis method.
[0021] Other features and advantages of the present invention will be described in the subsequent specification, and, in part, will become apparent from the specification or be understood by implementing the embodiments of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in the written specification and the accompanying drawings. Description of the Drawings
[0022] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0023] Figure 1 It is a schematic flowchart of a drug molecule retrosynthesis method for generating an expected synthesis path described in the embodiments of the present invention.
[0024] Figure 2 It is a schematic structural diagram of a drug molecule retrosynthesis system for generating an expected synthesis path described in the embodiments of the present invention.
[0025] Figure 3 It is a schematic structural diagram of a drug molecule retrosynthesis device for generating an expected synthesis path described in the embodiments of the present invention.
[0026] Figure 4 It is a retrosynthesis path diagram of the target product molecule.
[0027] In the figure: 800, a drug molecule retrosynthesis device for generating an expected synthesis path; 801, a processor; 802, a memory; 803, a multimedia component; 804, an I / O interface; 805, a communication component; 901, an acquisition module; 902, a first processing module; 903, a feature extraction module; 904, a second processing module; 905, a third processing module. Specific Embodiments
[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents the selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0029] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are only used for distinguishing descriptions and cannot be construed as indicating or implying relative importance.
[0030] Embodiment 1:
[0031] This embodiment provides a method for inverse synthesis of drug molecules for generating a desired synthetic path. It can be understood that in this embodiment, a scenario can be set up, for example, a scenario where a desired inverse synthesis path of drug molecules needs to be generated based on a target molecule.
[0032] Referring to Figure 1 , the figure shows that this method includes step S1, step S2, step S3, step S4, and step S5.
[0033] Step S1: Obtain a target product molecule and a reactant molecule corresponding to the target product molecule;
[0034] Step S2: Convert the target product molecule and the reactant molecule corresponding to the target product molecule into corresponding two-dimensional molecular graphs to obtain a first molecular graph and a second molecular graph;
[0035] In this step, the target product molecule and the reactant molecule corresponding to the target product molecule are converted from their SMILES string representations to two-dimensional molecular graph representations using the RDKit tool to obtain a first molecular graph and a second molecular graph.
[0036] Step S3: Extract features from the first molecular graph and the second molecular graph to obtain first feature information and second feature information;
[0037] In step S3, there are also included step S31, step S32, and step S33, which specifically include:
[0038] Step S31: Assign an initial feature vector to each atom and an edge feature vector in the first molecular graph;
[0039] Step S32: In the message passing stage, update the feature vector of the atom using the edge feature vector of each atom to obtain an updated feature vector;
[0040] In this step, the process of updating the feature vector is specifically as follows:
[0041]
[0042] In the above formula, is the neighbor message of atom v after the (t + 1)-th iteration; Mt Represents the message function; Is the eigenvector of atom v after t iterations; e vw Is the eigenvector of the edge; Is the eigenvector of neighbor atom w of atom v after t iterations; N(v) is the set of neighbor atoms of atom v; Represents the eigenvector of atom v after t+1 iterations; U t Is the node update function; f 1 And f 2 Are feature transformation functions; σ is the sigmoid function; tanh is the non-linear feature transformation, h v Is the updated eigenvector.
[0043] Step S33: Send the updated eigenvector to a preset weight assignment function to obtain the first feature information.
[0044] In this step, a weight assignment function is defined to determine whether atom v belongs to functional group g to assign weights, and then the functional group information and reaction condition information (i.e., temperature, pressure, catalyst, solvent, pH value) are fused through the attention mechanism to obtain the first feature information. The specific process is as follows:
[0045]
[0046] In the above formula, h is the first feature information; Is the number of atomic nodes; Assoc(v,g) is the weight assignment function; h v Is the updated eigenvector; softmax is the normalized exponential function; Is the eigenvector f of the functional group g And the eigenvector f of its corresponding reaction condition con Of the concatenation, both of these eigenvectors are obtained through the autoencoder, + represents concatenation, W k 、W q And W v Are learnable weights, Is the scaling factor.
[0047] In this embodiment, the first feature information extracted by combining the weight assignment score with the attention mechanism can not only capture the local features of the molecule, but also consider the global properties of the functional group, thus providing more comprehensive information for retrosynthetic analysis. It should be noted that the process of extracting the first feature information is the same as that of the second feature information, so it will not be elaborated here.
[0048] Step S4: Determine the candidate reactant molecules according to the first feature information and the second feature information;
[0049] In step S4, steps S41, S42, S43, and S44 are further included, specifically including:
[0050] Step S41: Obtain the preset first constraint condition information and second constraint condition information. The first constraint condition information includes constraints on molecular size, polarity, and stereochemistry, and the second constraint condition information includes heuristic rules based on expert experience.
[0051] In this step, a series of chemical constraint conditions need to be defined to evaluate whether the reactant molecules meet the expected chemical properties and structural characteristics. Among them, the first constraint condition information is specifically:
[0052]
[0053] In the above formula, l represents the number of constraint conditions, λ k is the weight of constraint condition k, C k is a function used to evaluate the matching degree of constraint condition k between the target product molecule and the reactant molecule corresponding to the target product molecule, and h P , h R respectively represent the molecular feature information of the reactant molecule and the target product molecule, that is, the first feature information and the second feature information.
[0054] Among them, the second constraint condition information is specifically:
[0055]
[0056] Among them, n represents the number of heuristic rules, μ j is the weight of expert heuristic rule j, and the function H j is used to evaluate the matching degree of expert rule j between the target product molecule and the reactant molecule corresponding to the target product molecule.
[0057] Step S42: Calculate according to the first constraint condition information, the second constraint condition information, the first feature information, and the second feature information to obtain score information, which is used to evaluate whether each reactant molecule meets the expected chemical properties and structural characteristics.
[0058] In this step, the specific calculation process of the score information is:
[0059] f score (h R , h P ) = w 1 ·S constraint (h R , h P ) + w 2 ·S expert (hR , h P )
[0060] In the above formula, f score (h R , h P ) represents fractional information; w 1 and w 2 are the relative importance weights of the constraint conditions and expert knowledge; S constraint (h R , h P ) and S expert (h R , h P ) represent the first constraint condition information and the second constraint condition information respectively.
[0061] Step S43: Sort the reactant molecules according to the fractional information to obtain the sorted fractional information;
[0062] Step S44: Screen the sorted fractional information to obtain the candidate reactant molecules.
[0063] In this step, the reactant molecules are sorted from high to low by the fractional information to screen out the candidate reactant molecules that are most likely to produce the desired target product.
[0064] In this embodiment, a first scoring model that comprehensively considers chemical constraints and expert knowledge is established, and the appropriate candidate reactant molecules are effectively screened according to the predefined threshold, thereby linking the target product molecules and the candidate reactant molecules to construct different reaction types. This reaction type also includes two attributes: reaction conditions and chemical reagents, laying a foundation for subsequent retrosynthesis steps. In addition, through the combination of chemical constraint conditions and expert heuristic rules, reasonable reactant molecules can be screened out more effectively, and a more satisfactory synthetic route can be generated. It should be noted that the first loss function of the first scoring model is:
[0065]
[0066] Among them, c represents the type of chemical reaction, M represents the total number of reaction types, w c is the weight of reaction type c. Smaller weights are set for common reaction types, and larger weights are set for uncommon reaction types, so that more novel reactants can be obtained. y o,c is a binary indicator (0 or 1), which is 1 if class c is the correct classification of observation o, and 0 otherwise. p o,c is the probability that the model predicts that observation o belongs to class c.
[0067] Step S5. Determine the retrosynthetic route of the drug molecule based on the first molecular graph and the two-dimensional molecular graph corresponding to the candidate reactant molecule.
[0068] In step S5, it further includes step S51, step S52, and step S53, which specifically include:
[0069] Step S51. Determine the common substructure between the reactant molecule and the target product molecule based on the first molecular graph and the two-dimensional molecular graph corresponding to the candidate reactant molecule, and obtain the common substructure information.
[0070] In step S51, it further includes step S511, step S512, and step S513, which specifically include:
[0071] Step S511. Obtain a preset scoring function.
[0072] In this step, the preset scoring function is specifically:
[0073]
[0074] where S env is a function for evaluating the similarity of the chemical environments of two atoms, p is an atom in the target product molecule, u is an atom of the same type in the candidate reactant molecule, h p is the first feature information, and h u is the third feature information obtained by feature extraction from the two-dimensional molecular graph corresponding to the candidate reactant molecule; G 1 , G 2 are the first molecular graph and the two-dimensional molecular graph corresponding to the candidate reactant molecule respectively; S sub (G 1 , G 2 ) represents the scoring function.
[0075] Step S512. Calculate the matching degree between the first molecular graph and the two-dimensional molecular graph corresponding to the candidate reactant molecule according to the preset scoring function, and obtain the matching degree information.
[0076] Step S513. Determine the common substructure of the first molecular graph and the two-dimensional molecular graph corresponding to the candidate reactant molecule according to the matching degree information, and obtain the common substructure information.
[0077] In this step, select an optimal substructure from all possible substructure mappings, that is, the optimal substructure corresponding to the maximum value of the scoring function.
[0078] In this embodiment, establishing the second scoring model through the scoring function can effectively extract the common substructures from the candidate reactant molecules and the target product molecules, providing important information for retrosynthetic analysis. This process is of great significance for designing synthetic routes and predicting reaction mechanisms. It should be noted that the second loss function of the second scoring model is:
[0079]
[0080] In the above formula, N is the number of samples in the batch, sim() is a function for calculating the similarity between two subgraphs, and f pre and f ac are the predicted substructure and the true substructure respectively.
[0081] Step S52: Optimize the common substructure information to obtain a perfect reactant molecule;
[0082] In this step, in order to generate a new reactant molecule that meets the synthesis requirements, on the basis of retaining the substructures similar to the product molecule, it is necessary to carry out necessary optimization and improvement on the reactant molecule.
[0083] In step S52, it further includes steps S521, S522, S523, S524, and S525, which specifically include:
[0084] Step S521: Determine the reaction center in the common substructure information to obtain the key atoms in the common substructure information;
[0085] In this step, the specific calculation process for determining the reaction center in the common substructure information is as follows:
[0086]
[0087] In the above formula, ΔE i represents the change in the chemical environment of atom v; x represents the number of atoms in the optimal substructure; f * represents the optimal substructure, and its third loss function is:
[0088]
[0089] Among them, y v is a binary label indicating whether atom v has changed, and p v is the probability distribution of the change of atom v predicted by the model.
[0090] Step S522: Obtain the leaving group set;
[0091] Step S523: Use the trained multi-layer perceptron to predict the probability that the leaving groups in the set of leaving groups are added to the key atoms, and obtain a probability feature vector;
[0092] In this step, training the multi-layer perceptron is a well-known technical solution to those skilled in the art, so it will not be elaborated here. It should be noted that the fourth loss function of the trained multi-layer perceptron is:
[0093]
[0094] where, is the probability distribution of the model predicting that the target molecule G 1 matches the leaving group I G i .
[0095] Step S524: Screen the set of leaving groups according to the probability feature vector to obtain the screened leaving groups;
[0096] Step S525: Use the screened leaving groups to optimize the key atoms to obtain a refined reactant molecule.
[0097] In this step, the screened leaving group is the leaving group with the largest probability value in the probability feature vector. Adding it to the key atom can make necessary modifications and improvements to the reactant molecule while retaining the substructure similar to the target product molecule, thereby generating a new reactant molecule that may meet the synthesis requirements. This step is crucial for designing an effective synthesis route and improving the success rate of retrosynthetic analysis.
[0098] It should be noted that the total loss function is:
[0099] L = λ 1 L t + λ 2 L stru + λ 3 L c + λ 4 L IG
[0100] In the above formula, λ 1 , λ 2 , λ 3 , λ 4 are the weights of different task losses, and L t , L stru , L c , L IG are the first loss function, the second loss function, the third loss function, and the fourth loss function respectively.
[0101] Step S53: Determine the retrosynthetic pathway of the drug molecule based on the refined reactant molecules.
[0102] In this step, a reinforcement learning strategy is used to determine a relatively more feasible reaction pathway based on the properties of the reactant molecules. The specific algorithm steps are as follows:
[0103] 1. Establish and initialize the value function: For each state-action pair (S t , A t ), initialize Q(S t , A t ) as a small random number;
[0104] 2. For each time step t:
[0105] · Starting from the current state S t , select an action A t according to the exploration-exploitation strategy.
[0106] · Execute the action A t , observe the result, and obtain the reward R t+1 and the new state S t+1 .
[0107] · Calculate the temporal difference (TD) error L(S t , A t ):
[0108] L(S t , A t ) = Q(S t , A t ) - (R t+1 + γ · max a Q(S t+1 , a))
[0109] In the above formula, S t is the state at time step t, A t is the action taken at time step t, R t+1 is the immediate reward obtained from the environment after taking the action A t . γ is the discount factor, which determines the importance of future rewards relative to current rewards. a is a possible action, Q is the value function, and the Q value is updated through the temporal difference error, specifically:
[0110] Q(S t+1 , A t+1 ) ← Q(S t , A t ) + α · L(S t , A t )
[0111] Among them, α is the learning rate, which is used to control the update step size.
[0112] 3. Repeat step 2 until the Q value converges or reaches a predetermined number of iterations.
[0113] After completing the above algorithm steps, the reaction path of the target product can be obtained. That is, the reaction type and its corresponding reactants are regarded as a node. For example, Figure 4 As shown, from the root node to each leaf node is a reaction path, and reinforcement learning will automatically decide a relatively more feasible reaction path according to the molecular properties. If the policy pays more attention to the simplicity of the reactants, then Figure 4 the left path of the tree is better. If the policy pays more attention to a shorter reaction path, then Figure 4 the right path is better. Different strategies need to be adjusted in different scenarios. Therefore, the present invention does not limit the strategy.
[0114] Different from the traditional black box model, the inverse synthesis process of the present invention is similar to the manual steps of inverse synthesis. Compared with the lack of transparency and interpretability in the inference process of inverse synthesis in the prior art, the present invention helps to understand the inference process of the model, increases the credibility of the model, and thus provides the ability to provide valuable insights for chemical experts.
[0115] Example 2:
[0116] As Figure 2 shown, this embodiment provides a drug molecule inverse synthesis system for generating an expected synthesis path. The system includes an acquisition module 901, a first processing module 902, a feature extraction module 903, a second processing module 904, and a third processing module 905, which specifically include:
[0117] The acquisition module 901 is used to acquire the target product molecule and the reactant molecule corresponding to the target product molecule;
[0118] The first processing module 902 is used to convert the target product molecule and the reactant molecule corresponding to the target product molecule into corresponding two-dimensional molecular graphs to obtain a first molecular graph and a second molecular graph;
[0119] The feature extraction module 903 is used to extract features from the first molecular graph and the second molecular graph to obtain first feature information and second feature information;
[0120] The second processing module 904 is used to determine candidate reactant molecules according to the first feature information and the second feature information;
[0121] The third processing module 905 is used to determine the inverse synthesis path of the drug molecule according to the first molecular graph and the two-dimensional molecular graph corresponding to the candidate reactant molecule.
[0122] In a specific implementation manner of the present disclosure, the feature extraction module further includes a first processing unit, a second processing unit, and a third processing unit, specifically including:
[0123] The first processing unit is configured to assign an initial feature vector to each atom and an edge feature vector in the first molecular graph;
[0124] The second processing unit is configured to update the feature vector of the atom by using the edge feature vector of each atom during the message passing stage to obtain an updated feature vector;
[0125] The third processing unit is configured to send the updated feature vector to a preset weight assignment function to obtain first feature information.
[0126] In a specific implementation manner of the present disclosure, the second processing module further includes a first acquisition unit, a fourth processing unit, a fifth processing unit, and a sixth processing unit, specifically including:
[0127] The first acquisition unit is configured to acquire preset first constraint condition information and second constraint condition information, where the first constraint condition information includes constraints on molecular size, polarity, and stereochemistry, and the second constraint condition information includes heuristic rules based on expert experience;
[0128] The fourth processing unit is configured to perform calculations based on the first constraint condition information, the second constraint condition information, the first feature information, and the second feature information to obtain score information, where the score information is used to evaluate whether each reactant molecule meets the expected chemical properties and structural characteristics;
[0129] The fifth processing unit is configured to sort the reactant molecules according to the score information to obtain sorted score information;
[0130] The sixth processing unit is configured to screen the sorted score information to obtain the candidate reactant molecules.
[0131] In a specific implementation manner of the present disclosure, the third processing module further includes a seventh processing unit, an eighth processing unit, and a ninth processing unit, specifically including:
[0132] The seventh processing unit is configured to determine a common substructure between the reactant molecule and the target product molecule according to the first molecular graph and the two-dimensional molecular graph corresponding to the candidate reactant molecule to obtain common substructure information;
[0133] The eighth processing unit is configured to optimize the common substructure information to obtain a refined reactant molecule;
[0134] A ninth processing unit for determining an inverse synthesis path of a drug molecule based on the refined reactant molecule.
[0135] In a specific embodiment of the present disclosure, the seventh processing unit further includes a second acquisition unit, a tenth processing unit, and an eleventh processing unit, which specifically include:
[0136] The second acquisition unit is used to acquire a preset scoring function;
[0137] The tenth processing unit is used to calculate the matching degree between the first molecular graph and the two-dimensional molecular graph corresponding to the candidate reactant molecule according to the preset scoring function, and obtain matching degree information;
[0138] The eleventh processing unit is used to determine the common substructure of the first molecular graph and the two-dimensional molecular graph corresponding to the candidate reactant molecule according to the matching degree information, and obtain common substructure information.
[0139] In a specific embodiment of the present disclosure, the eighth processing unit further includes a twelfth processing unit, a third acquisition unit, a thirteenth processing unit, a fourteenth processing unit, and a fifteenth processing unit, which specifically include:
[0140] The twelfth processing unit is used to determine the reaction center in the common substructure information to obtain the key atoms in the common substructure information;
[0141] The third acquisition unit is used to acquire a leaving group set;
[0142] The thirteenth processing unit is used to predict the probability that a leaving group in the leaving group set is added to a key atom by using a trained multi-layer perceptron, and obtain a probability feature vector;
[0143] The fourteenth processing unit is used to screen the leaving group set according to the probability feature vector to obtain the screened leaving group;
[0144] The fifteenth processing unit is used to optimize the key atom by using the screened leaving group to obtain a refined reactant molecule.
[0145] It should be noted that regarding the system in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0146] Example 3:
[0147] Corresponding to the above method embodiments, a drug molecule retrosynthesis device for generating a desired synthetic path is further provided in this embodiment. A drug molecule retrosynthesis device for generating a desired synthetic path described below can be correspondingly referred to in relation to a drug molecule retrosynthesis method for generating a desired synthetic path described above.
[0148] Figure 3 FIG. is a block diagram of a drug molecule retrosynthesis device 800 for generating a desired synthetic path shown according to an exemplary embodiment. As Figure 3 shown, the drug molecule retrosynthesis device 800 for generating a desired synthetic path may include: a processor 801, a memory 802. The drug molecule retrosynthesis device 800 for generating a desired synthetic path may further include one or more of a multimedia component 803, an I / O interface 804, and a communication component 805.
[0149] Among them, the processor 801 is used to control the overall operation of the drug molecule retrosynthesis device 800 for generating an expected synthetic path to complete all or part of the steps in the above-mentioned drug molecule retrosynthesis method for generating an expected synthetic path. The memory 802 is used to store various types of data to support the operation of the drug molecule retrosynthesis device 800 for generating an expected synthetic path. These data may include, for example, instructions for any application or method operating on the drug molecule retrosynthesis device 800 for generating an expected synthetic path, as well as application-related data, such as contact data, sent and received messages, pictures, audio, video, and so on. The memory 802 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disc. The multimedia component 803 may include a screen and an audio component. The screen can be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signal can be further stored in the memory 802 or sent through the communication component 805. The audio component also includes at least one speaker for outputting audio signals. The I / O interface 804 provides an interface between the processor 801 and other interface modules, and the above other interface modules can be a keyboard, a mouse, buttons, etc. These buttons can be virtual buttons or physical buttons. The communication component 805 is used for wired or wireless communication between the drug molecule retrosynthesis device 800 for generating an expected synthetic path and other devices. Wireless communication, such as Wi-Fi, Bluetooth, near field communication (NFC), 2G, 3G, or 4G, or a combination of one or more of them. Accordingly, the communication component 805 may include: a Wi-Fi module, a Bluetooth module, and an NFC module.
[0150] In an exemplary embodiment, the drug molecule retrosynthesis device 800 for generating a desired synthetic route may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components, and is used to execute the above-mentioned drug molecule retrosynthesis method for generating a desired synthetic route.
[0151] In another exemplary embodiment, a computer-readable storage medium including program instructions is further provided. When the program instructions are executed by a processor, the steps of the above-mentioned drug molecule retrosynthesis method for generating a desired synthetic route are implemented. For example, the computer-readable storage medium may be the above-mentioned memory 802 including program instructions, and the above program instructions may be executed by the processor 801 of the drug molecule retrosynthesis device 800 for generating a desired synthetic route to complete the above-mentioned drug molecule retrosynthesis method for generating a desired synthetic route.
[0152] Embodiment 4:
[0153] Corresponding to the above method embodiment, a readable storage medium is further provided in this embodiment. A readable storage medium described below can be correspondingly referred to with a drug molecule retrosynthesis method for generating a desired synthetic route described above.
[0154] A readable storage medium stores a computer program. When the computer program is executed by a processor, the steps of the drug molecule retrosynthesis method for generating a desired synthetic route in the above method embodiment are implemented.
[0155] The readable storage medium may specifically be various readable storage media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.
[0156] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
[0157] As described above, these are only the specific implementation manners of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or replacements, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A method for the reverse synthesis of drug molecules for generating a desired synthetic pathway, characterized in that: include: Obtaining target product molecules and reactant molecules corresponding to the target product molecules; Converting the target product molecule and the reactant molecule corresponding to the target product molecule into corresponding two-dimensional molecular graphs to obtain a first molecular graph and a second molecular graph; Performing feature extraction on the first molecular graph and the second molecular graph to obtain first feature information and second feature information; Determine a candidate reactant molecule according to the first characteristic information and the second characteristic information; The retrosynthetic pathway of the drug molecule is determined based on the first molecular graph and the two-dimensional molecular graph corresponding to the candidate reactant molecule.
2. The method for reverse synthesis of drug molecules for generating a desired synthesis pathway according to claim 1, characterized in that: Extracting features from the first molecular graph and the second molecular graph includes: Assigning an initialized eigenvector and an edge eigenvector to each atom in the first molecular graph; In the message passing phase, the feature vector of the atom is updated using the feature vector of each atomic edge to obtain an updated feature vector; The updated feature vector is sent to a preset weight distribution function to obtain first feature information.
3. The method for reverse synthesis of drug molecules for generating a desired synthesis pathway according to claim 1, characterized in that: Determining a candidate reactant molecule according to the first characteristic information and the second characteristic information includes: Acquiring preset first constraint information and second constraint information, wherein the first constraint information includes constraints on molecular size, polarity, and stereochemistry, and the second constraint information includes heuristic rules based on expert experience; Calculating according to the first constraint information, the second constraint information, the first feature information, and the second feature information to obtain score information, wherein the score information is used to evaluate whether each reactant molecule meets expected chemical properties and structural characteristics; Sort the reactant molecules according to the score information to obtain sorted score information; The ranked score information is screened to obtain the candidate reactant molecules.
4. The method for reverse synthesis of drug molecules for generating a desired synthesis pathway according to claim 1, characterized in that: Determining the retrosynthetic pathway of the drug molecule according to the first molecular graph and the two-dimensional molecular graph corresponding to the candidate reactant molecule includes: Determine the common substructure between the reactant molecule and the target product molecule according to the first molecular graph and the two-dimensional molecular graph corresponding to the candidate reactant molecule, and obtain the common substructure information; Optimizing the common substructure information to obtain a perfect reactant molecule; The retrosynthetic pathway of the drug molecule is determined based on the perfected reactant molecules.
5. The method for reverse synthesis of drug molecules for generating a desired synthesis route according to claim 4, characterized in that: Determining the common substructure between the reactant molecule and the target product molecule according to the first molecular graph and the two-dimensional molecular graph corresponding to the candidate reactant molecule to obtain the common substructure information includes: Get the preset scoring function; Calculating the matching degree between the first molecular graph and the two-dimensional molecular graph corresponding to the candidate reactant molecule according to the preset scoring function to obtain matching degree information; The common substructure of the first molecular graph and the two-dimensional molecular graph corresponding to the candidate reactant molecule is determined according to the matching degree information to obtain common substructure information.
6. A drug molecule retrosynthesis system for generating a desired synthesis pathway, characterized in that: include: An acquisition module, used to acquire target product molecules and reactant molecules corresponding to the target product molecules; A first processing module, used for converting the target product molecule and the reactant molecule corresponding to the target product molecule into corresponding two-dimensional molecular graphs to obtain a first molecular graph and a second molecular graph; A feature extraction module, used to extract features from the first molecular graph and the second molecular graph to obtain first feature information and second feature information; A second processing module, configured to determine a candidate reactant molecule according to the first characteristic information and the second characteristic information; The third processing module is used to determine the retrosynthetic pathway of the drug molecule according to the first molecular graph and the two-dimensional molecular graph corresponding to the candidate reactant molecule.
7. The drug molecule retrosynthesis system for generating a desired synthesis route according to claim 6, characterized in that: The feature extraction module comprises: A first processing unit, configured to assign an initialized feature vector and an edge feature vector to each atom in the first molecular graph; A second processing unit is used to update the feature vector of the atom using the feature vector of each atomic edge in the message passing phase to obtain an updated feature vector; The third processing unit is used to send the updated feature vector to a preset weight distribution function to obtain the first feature information.
8. The drug molecule retrosynthesis system for generating a desired synthesis route according to claim 6, characterized in that: The second processing module comprises: A first acquisition unit, used to acquire preset first constraint information and second constraint information, wherein the first constraint information includes constraints on molecular size, polarity, and stereochemistry, and the second constraint information includes heuristic rules based on expert experience; a fourth processing unit, configured to perform calculation according to the first constraint information, the second constraint information, the first feature information, and the second feature information to obtain score information, wherein the score information is used to evaluate whether each reactant molecule meets expected chemical properties and structural characteristics; a fifth processing unit, configured to sort the reactant molecules according to the score information to obtain sorted score information; The sixth processing unit is used to screen the sorted score information to obtain the candidate reactant molecules.
9. The drug molecule retrosynthesis system for generating a desired synthesis pathway according to claim 6, characterized in that: The third processing module comprises: A seventh processing unit, configured to determine a common substructure between the reactant molecule and the target product molecule according to the first molecular graph and the two-dimensional molecular graph corresponding to the candidate reactant molecule, and obtain common substructure information; an eighth processing unit, configured to optimize the common substructure information to obtain a perfect reactant molecule; A ninth processing unit is used to determine a retrosynthetic pathway of the drug molecule based on the perfected reactant molecule.
10. The drug molecule retrosynthesis system for generating a desired synthesis route according to claim 9, characterized in that: The seventh processing unit comprises: A second acquisition unit, used to acquire a preset scoring function; a tenth processing unit, configured to calculate the degree of matching between the first molecular graph and the two-dimensional molecular graph corresponding to the candidate reactant molecule according to the preset scoring function to obtain matching degree information; An eleventh processing unit is used to determine a common substructure of the first molecular graph and the two-dimensional molecular graph corresponding to the candidate reactant molecule according to the matching degree information to obtain common substructure information.