Reverse synthesis method, apparatus, electronic device, and storage medium
By identifying the reaction center in the molecular diagram of the target product and dividing it into a connection tree of structural units, the problem of handling complex target products in the prior art is solved, and a more efficient retrosynthetic method is realized.
Patent Information
- Application Number
- CN202211230121.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-08
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2042-10-08
AI Technical Summary
Existing retrosynthetic methods are difficult to handle structurally complex target products and have a narrow range of applications.
By identifying the reaction center of the target product, its molecular graph is segmented to obtain the molecular graph of the structural unit, which is then converted into a connection tree. Target nodes are added to obtain the connection tree of the reactants.
It simplifies the processing procedure, improves processing efficiency, and expands the scope of application, especially suitable for complex molecular structures.
Smart Images

Figure CN117012296B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the technical field of computer, in particular to a reverse synthesis method and device, electronic equipment and storage medium. BACKGROUND
[0002] The reverse synthesis technology refers to a technology of searching for a reactant for synthesizing a target product. A reverse synthesis method is provided in the related technology. The method first determines a reaction center of the target product, indicating that the reactant has a chemical reaction at the reaction center, thereby generating the target product. Then, the target product is disconnected at the reaction center to obtain at least one structural unit, and a molecular graph of the at least one structural unit is completed, so that a molecular graph of at least one reactant is obtained. However, the above method is difficult to process a target product with a relatively complex structure, and has a narrow application range. SUMMARY
[0003] Embodiments of the present application provide a reverse synthesis method and device, electronic equipment and storage medium, which can process a connection tree of a structural unit to obtain a connection tree of a reactant, and the processing flow is relatively simple. The technical solution is as follows:
[0004] In one aspect, a reverse synthesis method is provided, and the method comprises:
[0005] determining a reaction center of a target product in a molecular graph of the target product;
[0006] segmenting the molecular graph of the target product based on the reaction center to obtain a molecular graph of at least one structural unit;
[0007] converting the molecular graph of the at least one structural unit into a connection tree;
[0008] determining a target node of each structural unit based on the connection tree of each structural unit, the target node being a new node to be added to the connection tree of the structural unit;
[0009] adding the target node of each structural unit to the connection tree of each structural unit to obtain a connection tree of a reactant corresponding to each structural unit, wherein the reactant is a substance for synthesizing the target product.
[0010] In another aspect, a reverse synthesis device is provided, and the device comprises:
[0011] a reaction center determination module configured to determine a reaction center of a target product in a molecular graph of the target product;
[0012] The segmentation module is configured to segment a molecular graph of the target product based on the reaction center, to obtain a molecular graph of at least one structural unit.
[0013] The conversion module is configured to convert the molecular graph of the at least one structural unit into a connection tree.
[0014] The completion module is configured to determine a target node of each structural unit based on the connection tree of the structural unit, the target node being a new node that needs to be added to the connection tree of the structural unit.
[0015] The completion module is configured to add the target node of each structural unit to the connection tree of the structural unit, to obtain a connection tree of a reactant corresponding to the structural unit, the reactant being a substance used to synthesize the target product.
[0016] Optionally, the completion module is configured to:
[0017] encode the connection tree of each structural unit to obtain a feature of the connection tree of each structural unit;
[0018] determine the target node of each structural unit based on the feature of the connection tree of each structural unit.
[0019] Optionally, the completion module is configured to:
[0020] determine a target type from a plurality of preset types based on the feature of the reaction center, the feature of the connection tree of the structural unit, and the features of the plurality of preset types, each preset type representing a node type;
[0021] determine a node belonging to the target type as the target node of the structural unit.
[0022] Optionally, the completion module comprises:
[0023] The determination unit is configured to determine a first connection position and a second connection position based on the connection tree of the structural unit and the target node of the structural unit, the first connection position being a position on the connection tree at which the target node needs to be connected, the second connection position being a position on the target node at which the connection tree needs to be connected, and the first connection position and the second connection position being of the same type.
[0024] The merging unit is configured to merge the first connection position on the connection tree of the structural unit and the second connection position on the target node, to obtain the connection tree of the reactant.
[0025] Optionally, the determination unit is configured to:
[0026] determining a plurality of position pairs, wherein each position pair comprises a first position on the connection tree and a second position on the target node, and the types of the two positions in the same position pair are the same;
[0027] determining a first parameter of each position pair based on the features of the first position and the second position in the position pair, wherein the first parameter represents a probability of merging the first position and the second position in the position pair;
[0028] selecting a target position pair with the largest first parameter from the plurality of position pairs, determining the first position in the target position pair as the first connection position, and determining the second position in the target position pair as the second connection position.
[0029] Optionally, the apparatus further comprises:
[0030] a molecular graph completion module configured to add a molecular structure corresponding to the target node of each structural unit in the molecular graph of the structural unit to obtain a molecular graph of each reactant.
[0031] Optionally, the apparatus further comprises:
[0032] a second parameter determination module configured to determine a second parameter based on the features of the reaction center and the features of the connection tree of the at least one structural unit, wherein the second parameter represents a probability of completing the connection tree;
[0033] the completion module is configured to execute the steps of determining the target node of each structural unit and adding the target node of each structural unit on the connection tree of the structural unit when the second parameter is greater than a preset threshold value.
[0034] a connection tree determination module configured to determine the connection tree of each structural unit as the connection tree of one reactant when the second parameter is not greater than the preset threshold value.
[0035] Optionally, the steps of determining the target node of each structural unit and adding the target node of each structural unit on the connection tree of the structural unit are executed based on an inverse synthesis model, and the apparatus further comprises:
[0036] a sample acquisition module configured to acquire a first connection tree of a sample structural unit and a second connection tree of a sample reactant, wherein the sample structural unit is obtained by disconnecting a sample product at the reaction center, and the sample product is a product obtained after a chemical reaction of the sample reactant;
[0037] a training module configured to train the inverse synthetic model based on the first connection tree and the second connection tree, so that the trained inverse synthetic model is used to determine the connection tree of the reactants based on the connection tree of any structural unit.
[0038] Optionally, the training module is configured to:
[0039] determine a loss value of the inverse synthetic model based on the first connection tree and the second connection tree, the loss value comprising at least one of a first loss value, a second loss value and a third loss value;
[0040] train the inverse synthetic model based on the loss value, so that the loss value of the trained inverse synthetic model is reduced;
[0041] wherein the first loss value is a loss value of first connection positions and second connection positions obtained by the inverse synthetic model after the first connection tree is input into the inverse synthetic model, the first connection positions are positions on the first connection tree where new nodes need to be connected, and the second connection positions are positions on the new nodes where the first connection tree needs to be connected;
[0042] the second loss value is a loss value of a second parameter obtained by the inverse synthetic model after the first connection tree is input into the inverse synthetic model, the second parameter representing a probability of completing the connection tree;
[0043] the third loss value is a loss value of a target type obtained by the inverse synthetic model after the first connection tree is input into the inverse synthetic model, the target type being a type of new nodes that need to be added on the first connection tree.
[0044] Optionally, the apparatus further comprises:
[0045] a connection tree pair acquisition module configured to select at least one connection tree pair from a data set, each connection tree pair comprising connection trees of two structural units, and the two connection trees in the same connection tree pair satisfying a similarity condition, the data set comprising connection trees of a plurality of structural units;
[0046] the training module is configured to train the inverse synthetic model based on the at least one connection tree pair, so that the inverse synthetic model distinguishes different positions in the two connection trees in the same connection tree pair.
[0047] Optionally, the step of determining the reaction center of the target product is performed based on a reaction center identification model, and the apparatus further comprises:
[0048] a sample acquisition module configured to acquire a molecular graph of a sample product and a position identifier, the position identifier representing a position of a reaction center of the sample product in the molecular graph;
[0049] The training module is configured to train the reaction center recognition model based on the molecular graph and the position identifier, so that the trained reaction center recognition model is used to determine the reaction center in the molecular graph of any product.
[0050] Optionally, the apparatus further comprises:
[0051] The connected tree pair acquisition module is configured to select at least one connected tree pair from a data set, each connected tree pair comprising two connected trees of structural units, and the two connected trees in the same connected tree pair satisfying a similarity condition, the data set comprising connected trees of multiple structural units.
[0052] The training module is configured to train the reaction center recognition model based on the at least one connected tree pair, so that the reaction center recognition model distinguishes different positions in the two connected trees in the same connected tree pair.
[0053] In another aspect, an electronic device is provided, which comprises a processor and a memory, the memory storing at least one computer program, the at least one computer program being loaded and executed by the processor to implement the operations performed by the retrosynthetic method according to the above aspect.
[0054] In another aspect, a computer readable storage medium is provided, which stores at least one computer program, the at least one computer program being loaded and executed by a processor to implement the operations performed by the retrosynthetic method according to the above aspect.
[0055] In another aspect, a computer program product is provided, which comprises a computer program, the computer program being loaded and executed by a processor to implement the operations performed by the retrosynthetic method according to the above aspect.
[0056] The embodiments of the present application provide a retrosynthetic scheme, which splits the molecular graph of a target product based on a reaction center to obtain at least one molecular graph of a structural unit, simulates a scenario of disconnecting the target product from the reaction center, and then no longer directly processes the molecular graph, but converts the molecular graph of the structural unit into a connected tree to represent the structural unit in the form of a connected tree, so as to process the connected tree of the structural unit to obtain a connected tree of a reactant. Since the connected tree has a relatively simple form of expression, it is more intuitive and concise than the molecular graph, and thus the processing flow is relatively simple, and is still applicable even in the case that the molecular structure of the target product is relatively complex. Therefore, the processing efficiency of the embodiments of the present application is higher, and the scope of application is wider. BRIEF DESCRIPTION OF DRAWINGS
[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort on the basis of these drawings.
[0058] Figure 1 is a schematic diagram of an implementation environment provided by an embodiment of the present application;
[0059] Figure 2 is a flowchart of a reverse synthesis method provided by an embodiment of the present application;
[0060] Figure 3 is a schematic diagram of converting a molecular graph into a connection tree provided by an embodiment of the present application;
[0061] Figure 4 is a flowchart of another reverse synthesis method provided by an embodiment of the present application;
[0062] Figure 5 is a flowchart of a reverse synthesis model training method provided by an embodiment of the present application;
[0063] Figure 6 is a flowchart of a reaction center recognition model training method provided by an embodiment of the present application;
[0064] Figure 7 is a flowchart of training a reverse synthesis model provided by an embodiment of the present application;
[0065] Figure 8 is a structural schematic diagram of a reverse synthesis device provided by an embodiment of the present application;
[0066] Figure 9 is a structural schematic diagram of an electronic device 900 provided by an embodiment of the present application. DETAILED DESCRIPTION
[0067] In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the drawings.
[0068] It can be understood that the terms "first", "second", and the like used in the present application can be used herein to describe various concepts, but unless specifically stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the present application, a first structural unit can be referred to as a second structural unit, and similarly, a second structural unit can be referred to as a first structural unit.
[0069] At least one refers to one or more than one, for example, at least one structural unit can be one structural unit, two structural units, three structural units, or any integer greater than or equal to one structural unit. Multiple refers to two or more than two, for example, multiple structural units can be two structural units, three structural units, or any integer greater than or equal to two structural units. Each refers to each of at least one, for example, each structural unit refers to each of the multiple structural units, if the multiple structural units are three structural units, each structural unit refers to each of the three structural units.
[0070] It can be understood that in the embodiments of the present application, data related to user information and the like are involved, and when the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of countries and regions.
[0071] Artificial intelligence (AI) is the use of digital computers or digital computer-controlled machines to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which tries to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making.
[0072] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, both hardware and software technologies. Artificial intelligence basic technologies generally include technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics, etc. Artificial intelligence software technology includes computer vision technology and machine learning / deep learning.
[0073] Computer Vision (CV) is a science that studies how to make machines "see". More specifically, it refers to using cameras and computers to replace human eyes to identify and measure targets, and further perform image processing to make the computer processing become images more suitable for human eyes to observe or transmitted to instruments for detection. As a scientific discipline, computer vision researches related theories and technologies, and attempts to establish artificial intelligence systems that can obtain information from images or multidimensional data. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and other technologies. It also includes common face recognition, fingerprint recognition and other biometric identification technologies.
[0074] Machine Learning (ML) is a multi-disciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It is a subject that studies how computers simulate or implement human learning behavior to acquire new knowledge or skills, and reorganize existing knowledge structure to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental approach to making computers intelligent. Its applications are widespread in various fields of artificial intelligence. Machine learning and deep learning include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and teaching learning.
[0075] The embodiments of the present application are based on the above artificial intelligence technology, and provide a reverse synthesis method. The method can be applied to reverse synthesis of a given target product to find reactants for synthesizing the target product.
[0076] The embodiments of the present application are applied to electronic devices, which are terminals or servers, or can also be other electronic devices. The terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, etc., but is not limited thereto. The server can be a standalone physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and basic cloud computing services such as big data and artificial intelligence platforms.
[0077] Figure 1 is a schematic diagram of an implementation environment provided by the embodiments of the present application, as Figure 1As shown, the implementation environment includes: a terminal 101 and a server 102. The terminal 101 and the server 102 can be directly or indirectly connected through wired or wireless communication, which is not limited in the present application.
[0078] In the embodiment of the present application, the terminal 101 sends a reverse synthesis request for a target product to the server 102, and the server 102 determines a connection tree of at least one reactant based on a molecular graph of the target product in response to the reverse synthesis request, and represents at least one reactant for synthesizing the target product in the form of the connection tree.
[0079] In a possible implementation manner, the server 102 stores product identifiers of various products and provides the terminal 101 to display, and the user selects a target product identifier of a target product from the product identifiers of various products on the terminal 101, and the reverse synthesis request sent by the terminal 101 carries the target product identifier of the target product, and the server 102 obtains the molecular graph of the target product based on the target product identifier, thereby performing reverse synthesis based on the molecular graph of the target product.
[0080] In another possible implementation manner, the user inputs a molecular graph of a target product in the terminal 101, and the reverse synthesis request sent by the terminal 101 carries the molecular graph of the target product, and the server 102 performs reverse synthesis based on the molecular graph of the target product.
[0081] After the reverse synthesis is completed, the server 102 can send the connection tree of the at least one reactant to the terminal 101, and the terminal 101 can display the connection tree of the at least one reactant, or convert the connection tree of the at least one reactant into a molecular graph and display the molecular graph of the at least one reactant for the user to view.
[0082] Optionally, the terminal 101 installs an application provided by the server 102, and the terminal 101 and the server 102 can interact through the application, thereby realizing the reverse synthesis method.
[0083] Figure 2 is a flowchart of a reverse synthesis method provided by an embodiment of the present application, as shown in the figure, the method is executed by an electronic device, and the method includes: Figure 2
[0084] 201, the electronic device determines a reaction center of a target product in a molecular graph of the target product.
[0085] In the chemical field, substances can undergo chemical reactions to form new substances, where the substances undergoing chemical reactions are called reactants, and the new substances formed after chemical reactions are called products, and the reactants are substances used to synthesize products. For example, substances A and B undergo oxidation to obtain substance C, and substances A and B are called reactants, and substance C is called product.
[0086] Wherein, the process of chemical reaction can be called synthesis, and the process of determining the reactants given the product can be called reverse synthesis. The embodiments of the present application take the target product as an example to illustrate the reverse synthesis method, and the target product can be any product in the chemical field, and the processing process of other products is similar to the embodiments of the present application, which will not be described here.
[0087] In addition, in the chemical field, a molecule is the smallest unit of a substance that can exist independently, is relatively stable, and maintains the physical and chemical properties of the substance. Molecules are composed of atoms, which are combined into molecules in a certain order and arrangement by certain forces. The molecular structure is used to describe the connection relationship of atoms in the molecule, and the molecular structure greatly affects the reactivity, polarity, phase shape, color, magnetism and biological activity of the molecule.
[0088] And in the embodiments of the present application, the target product is a molecule, which can be composed of atoms. The molecular graph of the target product draws the molecular structure of the target product, that is, the molecular graph is used to describe the atoms that make up the target product and the connection relationship between these atoms, which can be a chemical bond or ring structure. In the molecular graph, atoms are represented by nodes, and the connection lines between nodes represent the chemical bonds between atoms, and the molecular graph is a two-dimensional image, but the three-dimensional connection relationship between each atom in the target product can be described in the molecular graph, so that the three-dimensional chemical structure of the target product can be obtained according to the molecular graph.
[0089] In the process of chemical reaction, the chemical bond of the reactant at a certain position is broken and fused with the chemical bond at a certain position of other reactants, thereby forming the target product, and a reaction center will be formed in the target product, which corresponds to the position of the reactant that undergoes chemical reaction. Therefore, in order to determine the reactant, it is necessary to first determine the reaction center of the target product based on the molecular graph of the target product.
[0090] Optionally, the reaction center is a chemical bond, which can be a single bond, or a double bond, or other forms of molecular connection structure.
[0091] Optionally, a SMILES (Simplified Molecular Input Line Entry System) character of the target product is acquired, the SMILES character is a string of characters converted from a three-dimensional chemical structure of the target product based on a simplified molecular linear input specification, and the SMILES character can represent the structure of the target product. The molecular graph of the target product can be generated based on the SMILES character of the target product.
[0092] 202. The electronic device splits the molecular graph of the target product based on the reaction center to obtain a molecular graph of at least one structural unit.
[0093] The splitting process can simulate a scenario in which the target product is disconnected from the reaction center to obtain at least one structural unit. The structural unit refers to a molecular fragment obtained after the reaction center of the target product is disconnected, which can be an ion, and the structural unit can also be referred to as a synthon.
[0094] For example, the reaction center in the target product is connected to two different parts. After the target product is disconnected from the reaction center, two structural units are obtained, which are independent of each other and represent different parts of the target product. Alternatively, the target product has a ring structure, and after the target product is disconnected from the reaction center, only one structural unit is obtained. Accordingly, after the molecular graph of the target product is split based on the reaction center, the molecular graph of at least one structural unit is obtained.
[0095] 203. The electronic device converts the molecular graph of at least one structural unit into a connection tree.
[0096] In the related art, the molecular graph of the structural unit is directly processed to obtain the molecular graph of the reactant. However, this approach can lead to difficulty in processing high-dimensional information. If the molecular structure of the target product is complex, the processing complexity is high, which affects the processing efficiency. Therefore, in the embodiments of the present application, the molecular graph of the structural unit is first converted into a connection tree, and then the connection tree is processed.
[0097] Optionally, the molecular graph of the structural unit includes a plurality of elements, the elements include nodes, rings, and bonds, and the element mapping table includes nodes corresponding to each element. The molecular graph of the structural unit is converted into a connection tree by: searching the element mapping table to determine the node corresponding to each element in the molecular graph, and connecting the nodes corresponding to each element according to the position of each element in the molecular graph to obtain the connection tree of the structural unit.
[0098] The connection tree can include at least two nodes and connection lines between the nodes, and is different from the molecular graph in that the connection tree does not have a ring, and thus is simpler in structure and more intuitive and concise in expression than the molecular graph. For example, at least two nodes in the molecular graph are combined into one node, and the ring originally present is eliminated, so that the ring originally present is changed into a node, and the molecular graph is changed into the connection tree. Figure 3 A schematic diagram of converting a molecular graph into a connection tree is shown in FIG. 1. Figure 3 As shown in FIG. 1, four Cl (chlorine) atoms in the molecular graph are converted into four nodes in the connection tree, the No. 1 Cl atom corresponds to the No. 1 node in the connection tree, the No. 2 Cl atom corresponds to the No. 2 node in the connection tree, and so on. The No. 5 ring structure in the molecular graph is converted into the No. 5 node in the connection tree, and the connection lines between the No. 1, No. 2 and No. 3 Cl atoms and the connection lines between the connection lines and the No. 5 ring structure are also converted into two nodes in the connection tree. As can be seen, the connection tree is simpler in expression and more intuitive and concise in expression than the molecular graph.
[0099] In the embodiments of the present application, each structure unit is a part of the target product, and the structure unit must come from the reactant and also belongs to a part of the reactant. Since the structure unit is expressed by the connection tree, the connection tree based on the structure unit can determine the connection tree of the reactant, so as to express which substance is used to synthesize the target product by the connection tree. Moreover, since the connection tree is simpler in expression, the processing flow is simpler, and the above processing mode is still applicable even if the molecular structure of the target product is complex. The process of determining the connection tree of the reactant based on the connection tree of the structure unit is described in detail in steps 204-205.
[0100] 204. The electronic device determines the target node of each structure unit based on the connection tree of each structure unit, the target node being a new node to be added on the connection tree of the structure unit.
[0101] In the process of synthesizing the target product by chemical change of at least one reactant, after the chemical bond at a position of the reactant is broken, the chemical bond at a position of another reactant is fused, so that in the reverse synthesis process, after the target product is broken from the reaction center, a "gap" can exist at the broken position, resulting in that the obtained structure unit is not a complete reactant. Correspondingly, the connection tree of the structure unit is not a complete connection tree.
[0102] Therefore, it is necessary to complete the connection tree of the structural unit, so as to restore the connection tree of the reactant. In order to complete the connection tree of the structural unit, it is necessary to determine the new node, i.e. the target node, to be added to the connection tree of the structural unit. The format of the target node is similar to that of the node in the connection tree. The target node can include one atom or multiple atoms. In the case that the target node includes multiple atoms, the target node can further include the connection relationship between the multiple atoms.
[0103] Optionally, the step 204 comprises: encoding the connection tree of each structural unit to obtain the feature of the connection tree of each structural unit. Based on the feature of the connection tree of each structural unit, the target node of each structural unit is determined, so as to add the corresponding target node to the connection tree of each structural unit, thereby completing the connection tree of each structural unit.
[0104] The feature of the connection tree of the structural unit is used to describe the connection tree of the structural unit. For example, the connection tree of the structural unit includes at least two nodes and the connection lines between the nodes, and the different nodes are connected in sequence. After encoding the connection tree of the structural unit, the obtained feature includes the features of the nodes in the connection tree and the features corresponding to the connection relationship between the nodes, which can reflect the position of the reaction center breakage. Therefore, the connection tree can be completed at the position, thereby obtaining the connection tree of the reactant.
[0105] Optionally, the connection tree of the structural unit includes nodes and connection lines between the nodes, and the feature table includes the features of each node and the features of each connection line. The encoding of the connection tree of the structural unit to obtain the feature of the connection tree comprises: searching the feature table to determine the features of each node in the structural unit and the features of the connection lines between any two nodes, and combining the features of each node and the features of the connection lines between the nodes based on the connection relationship between the nodes, and determining the combined features as the feature of the connection tree.
[0106] Optionally, the steps 204 and 205 are executed based on an inverse synthesis model. The inverse synthesis model includes an encoding layer and a decoding layer. The encoding layer is used to encode the connection tree of each structural unit to obtain the feature of the connection tree of each structural unit. The decoding layer is used to determine the target node of each structural unit based on the feature of the connection tree of each structural unit, add the corresponding target node to the connection tree of each structural unit, and complete the connection tree of each structural unit.
[0107] Optionally, based on the feature of the reaction center, the feature of the connection tree of the structural unit, and the features of the multiple preset types, a target type is determined from the multiple preset types, and the node belonging to the target type is determined as the target node.
[0108] The reaction center in the molecular graph of the target product corresponds to an atom in the molecular graph of the structural unit, and indicates that the reaction center is originally a chemical bond connected to the atom. The characteristics of the reaction center include attribute information and structural information of the atom corresponding to the reaction center. The attribute information of the atom includes atom type, atomic number, and aromaticity parameter, and the structural information of the atom includes the connection relationship between the atom and other atoms. The characteristics of the connection tree of the structural unit include the characteristics of each node in the connection tree and the characteristics corresponding to the connection relationship between the nodes.
[0109] The electronic device stores a plurality of preset types, and each preset type represents a node type. For example, one node type represents an atom, and another node type represents three different types of atomic connections. The characteristics of the preset type include the attribute information and the structural information of the atom represented by the preset type, which can represent what kind of atom is represented by the preset type and what kind of structure is connected.
[0110] Optionally, the electronic device determines the probability of each preset type based on the characteristics of the reaction center, the characteristics of the connection tree of the at least one structural unit, and the characteristics of the plurality of preset types, selects a preset type with the highest probability from the plurality of preset types as the target type, and determines the nodes belonging to the target type as the target nodes. The probability of the preset type represents the probability of adding a node belonging to the preset type.
[0111] In one possible implementation, the electronic device determines the probability of the preset type by using the formula y(n)=softmax(U×ReLU(W1n+W2z + )).
[0112] U is a feature matrix including the characteristics of the plurality of preset types. For example, if the number of preset types is m, U is an m-dimensional matrix including the characteristics of the m preset types. For example, the number of rows of U is m, and each row represents the characteristics of a preset type, or the number of columns of U is m, and each column represents the characteristics of a preset type. n is the characteristics of the atom corresponding to the reaction center, z + is the characteristics of the connection tree of the at least one structural unit, ReLU (Rectified Linear Unit) is an activation function, W1 and W2 are weights of the reaction center and the structural unit, softmax is a likelihood function, y(n) is an identifier of the preset type with the highest probability, or y(n) is a probability feature including the probability of each preset type, and the preset type with the highest probability can be determined based on the probability feature.
[0113] Optionally, steps 204 and 205 are performed based on an inverse synthesis model, ReLU and softmax are both functions provided inside the inverse synthesis model, U, W1 and W2 are model parameters in the inverse synthesis model, the model parameters can be determined by training the inverse synthesis model, and the training process of the inverse synthesis model will be described in the following embodiments, which will not be described here.
[0114] 205. The electronic device adds a target node of each structural unit on the connection tree of each structural unit to obtain a connection tree of the reactant corresponding to each structural unit.
[0115] In the connection tree of the structural unit, adding a target node means that a target node is added to an atom corresponding to a reaction center in the connection tree of the structural unit, so that the atom is connected to the target node, and a complete connection tree is obtained, that is, the connection tree of the reactant corresponding to the structural unit.
[0116] In the embodiments of the present application, the reactant is a substance used for synthesizing the target product, and a connection tree of the reactant can be obtained for each structural unit, so that at least one connection tree of the reactant can be obtained, that is, the at least one reactant is the substance used for synthesizing the target product.
[0117] The embodiments of the present application provide a reverse synthesis scheme, which splits the molecular graph of the target product based on the reaction center to obtain a molecular graph of at least one structural unit, simulates a scenario of disconnecting the target product from the reaction center, and then no longer directly processes the molecular graph, but converts the molecular graph of the structural unit into a connection tree to represent the structural unit in the form of the connection tree, so as to process the connection tree of the structural unit to obtain the connection tree of the reactant. Since the connection tree has a relatively simple form of expression, it is more intuitive and concise than the form of expression of the molecular graph, and therefore the processing flow is relatively simple, and is still applicable even in the case that the molecular structure of the target product is relatively complex. Therefore, the processing efficiency of the embodiments of the present application is higher, and the scope of application is wider.
[0118] The method provided by the embodiments of the present application can determine the target node with the highest probability based on the characteristics of the reaction center, the characteristics of the connection tree of the at least one structural unit, and the characteristics of the plurality of preset types, and add the target node on the connection tree of the structural unit to obtain the connection tree of the reactant corresponding to the structural unit, which can comprehensively consider the characteristics of the reaction center, the characteristics of the connection tree of the at least one structural unit, and the characteristics of the plurality of preset types to determine the target node, and ensure the accuracy of the reverse synthesis.
[0119] On the basis of the above-mentioned embodiments, the electronic device can also complete the molecular graph of the structural unit to obtain the molecular graph of the reactant. That is, after obtaining the molecular graph of the at least one structural unit, the method further includes: adding the molecular structure corresponding to the target node of each structural unit in the molecular graph of each structural unit to obtain the molecular graph of each reactant.
[0120] In the molecular graph of the structural unit, adding the molecular structure corresponding to the target node means: adding the molecular structure corresponding to the target node to the atom corresponding to the reaction center on the molecular graph of the structural unit, so that the atom is connected to the molecular structure corresponding to the target node, to obtain a complete molecular graph, which is the molecular graph of the reactant corresponding to the structural unit.
[0121] In the embodiments of the present application, the reactant is a substance used for synthesizing the target product, and the molecular graph of a reactant can be obtained for each structural unit, so that the molecular graph of at least one reactant can be obtained, that is, the at least one reactant is the substance used for synthesizing the target product.
[0122] The above-mentioned process of completing the molecular graph of the structural unit can be executed in parallel or in sequence with the above-mentioned process of completing the connection tree of the structural unit in steps 204-204, so as to obtain the molecular graph and the connection tree of the reactant to represent the reactant in different forms.
[0123] It should be noted that the above-mentioned embodiments are only described in the case of completing the connection tree, and the connection tree of the structural unit does not necessarily need to be completed. The connection tree of some structural units can be the complete connection tree of the substance, and therefore does not need to be completed. Therefore, the electronic device first determines whether the connection tree of the structural unit needs to be completed.
[0124] Optionally, the method further includes: determining a second parameter based on the characteristics of the reaction center and the characteristics of the connection tree of the at least one structural unit, the second parameter being used to indicate whether the connection tree of the structural unit needs to be completed, in the case that the second parameter is greater than a preset threshold, completing the connection tree of the at least one structural unit in steps 204-205 to obtain the connection tree of the at least one reactant, and in the case that the second parameter is not greater than the preset threshold, indicating that the connection tree of the structural unit does not need to be completed, and determining the connection tree of each structural unit as the connection tree of a reactant.
[0125] The second parameter can be the probability of completing the connection tree of the structural unit.
[0126] Optionally, the electronic device uses f(n) = σ(w T ReLU(W1n+W2z +) is determined, where m is a feature of an atom corresponding to a reaction center, z + is a feature of a connection tree of the at least one structural unit, ReLU is an activation function, w is a weight of the activation function, w T is a bias of w, W1 and W2 are weights of a reaction center and a structural unit, the σ function is a Sigmoid activation function, and f(n) is the second parameter.
[0127] For example, in a case where f(n) is greater than a preset threshold 0.5, it indicates that the connection tree of the structural unit needs to be completed, and the electronic device completes the connection tree of the at least one structural unit; in a case where f(n) is not greater than the preset threshold 0.5, it indicates that the connection tree of the structural unit does not need to be completed, and the electronic device determines the connection tree of each structural unit as a connection tree of one reactant.
[0128] The method provided in the embodiments of the present application can determine whether the connection tree of the structural unit needs to be completed based on the feature of the reaction center and the feature of the connection tree of the structural unit, and can complete the connection tree of the structural unit only when it needs to be completed. In a case where the connection tree of the structural unit does not need to be completed, the connection tree of the structural unit can be directly determined as a connection tree of one reactant, and the processing efficiency can be improved.
[0129] Figure 4 is a flowchart of another reverse synthesis method provided in the embodiments of the present application, as Figure 4 indicated in the figure, the method is executed by an electronic device. The method comprises the following steps.
[0130] 401. The electronic device determines a reaction center of a target product in a molecular graph of the target product.
[0131] 402. The electronic device segments the molecular graph of the target product based on the reaction center to obtain a molecular graph of at least one structural unit.
[0132] 403. The electronic device converts the molecular graph of the at least one structural unit into a connection tree.
[0133] Afterwards, for each structural unit, the following steps 404-408 are executed.
[0134] 404. The electronic device determines a target node of the structural unit based on the connection tree of the structural unit, where the target node is a new node that needs to be added in the connection tree of the structural unit.
[0135] Steps 401-404 are the same as steps 201-204, and will not be described here again.
[0136] 405. The electronic device determines a plurality of position pairs.
[0137] In the embodiments of the present application, the target node needs to be added to the connection tree of the structural unit, and there are multiple positions in the connection tree of the structural unit, for example, the connection tree includes multiple atoms, and each of the multiple atoms can be used as a connection position, and there are also multiple positions in the target node, for example, the target node also includes multiple atoms, and each of the multiple atoms can be used as a connection position. Therefore, it is necessary to find the positions connected to each other from the connection tree of the structural unit and the target node.
[0138] Therefore, the electronic device first combines at least one position in the connection tree of the structural unit with multiple positions in the target node in pairs to obtain multiple position pairs, each position pair including a first position on the connection tree of the structural unit and a second position in the target node, the first position being a position on the connection tree of the structural unit, and the second position being a position on the target node. Moreover, the two positions in the same position pair are of the same type, for example, both positions are atoms, or both positions are chemical bonds, so that the two connection positions can be connected after the connection positions of the connection tree of the structural unit and the target node are determined.
[0139] Alternatively, based on the reaction center, after the molecular graph of the target product is segmented, a disconnected position is formed in the molecular graph of the structural unit, and after the molecular graph of the structural unit is converted into a connection tree, the disconnected position corresponds to a node in the connection tree, indicating that the node is disconnected from other nodes. Therefore, the node should be used as a connection position. Therefore, in the connection tree of the structural unit, the electronic device only considers the node and does not consider other nodes. That is, the electronic device uses the node corresponding to the reaction center in the connection tree of the structural unit as a first position, combines the first position with multiple positions in the target node in pairs to obtain multiple position pairs, and then determines a target position pair from the multiple position pairs to determine the position in the target node connected to the first position.
[0140] 406、The electronic device determines a first parameter of each position pair based on the features of the first position and the second position in each position pair.
[0141] For any position, the position can be an atom, and the features of the position can include the features of the atom, such as the atomic species, the atomic nuclear charge number, and the aromaticity parameter, and can also include the features of the chemical bond connected to the atom, such as the type or connection relationship of the chemical bond, the type of the chemical bond being a single bond, a double bond, or a triple bond, and the connection relationship of the chemical bond being the species of the atom connected to each end of the chemical bond or the number of atoms connected by the chemical bond.
[0142] The first parameter represents a probability of merging the first position and the second position in the position pair, and the higher the first parameter, the greater the probability of merging the first position and the second position, and the more likely that the connection tree obtained by merging the first position and the second position is a connection tree of a complete substance, that is, the higher the accuracy of the reactant obtained after merging the first position and the second position.
[0143] Optionally, the electronic device adopts a formula g(a)=w T The first parameter is determined by tanh(W1a+W2x+W3z), where a is a feature of the first position in the position pair, x is a feature of the second position in the position pair, z is a feature of the connection tree of the structural unit, W1, W2 and W3 are weights of the first position, the second position and the structural unit respectively, tanh is an activation function, w is a weight of the activation function, and w T is a bias of w, and g(a) is the first parameter.
[0144] 407. The electronic device selects a target position pair with the maximum first parameter from the plurality of position pairs, determines the first position in the target position pair as the first connection position, and determines the second position in the target position pair as the second connection position.
[0145] The first connection position is a position on the connection tree of the structural unit that needs to be connected to the target node, and the second connection position is a position on the target node that needs to be connected to the connection tree of the structural unit, and the first connection position and the second connection position are of the same type, for example, both the first connection position and the second connection position are atoms.
[0146] 408. The electronic device merges the first connection position on the connection tree of the structural unit and the second connection position in the target node to obtain a connection tree of the reactant.
[0147] In the embodiments of the present application, after the electronic device determines the first connection position and the second connection position, the first connection position and the second connection position are merged, that is, the two connection positions are regarded as one connection position, so as to obtain the connection tree of the reactant. For example, in the first connection position and the second connection position, only the node on one connection position is retained, and the node on the other connection position is removed, and the nodes originally connected to the node on the retained connection position are reconnected, so as to realize the merging of the first connection position and the second connection position.
[0148] The method provided in the embodiments of the present application ensures the rationality of the connection positions and improves the accuracy of the retrosynthesis by determining the first connection position and the second connection position with the highest probability based on the features of the first position in the connection tree of the structural unit and the features of the target node, and merging the first connection position on the connection tree of the structural unit with the second connection position in the target node to obtain the connection tree of the reactant.
[0149] In the above embodiments, the step of determining the connection tree of the at least one reactant based on the connection tree of the at least one structural unit is performed based on a retrosynthesis model, that is, after the molecular graph of the structural unit is converted into the connection tree, the connection tree of the structural unit is input into the retrosynthesis model, and the retrosynthesis model processes to obtain the connection tree of the reactant. In the retrosynthesis model, the detailed steps shown in the above embodiments can be performed, such as determining whether the connection tree needs to be completed, if the connection tree needs to be completed, completing the connection tree to obtain the connection tree of the reactant, and if the connection tree does not need to be completed, taking the connection tree as the connection tree of a reactant. The specific operation of the retrosynthesis model is described in detail in steps 204-205 and steps 404-408 in the above embodiments, which will not be described here.
[0150] In order to ensure the accuracy of the retrosynthesis model, the retrosynthesis model needs to be trained. Figure 5 is a flowchart of a retrosynthesis model training method provided in the embodiments of the present application, as shown in Figure 5 The method is executed by an electronic device. The method comprises the following steps.
[0151] 501. The electronic device obtains a first connection tree of a sample structural unit and a second connection tree of a sample reactant.
[0152] In the embodiments of the present application, the molecular graph of the sample product is segmented based on the reaction center to obtain the molecular graph of the sample structural unit, and the sample product is the product obtained after the sample reactant undergoes a chemical reaction. The molecular graph of the sample structural unit and the molecular graph of the sample reactant are converted into connection trees to obtain the first connection tree of the sample structural unit and the second connection tree of the sample reactant. The molecular graph of the sample product and the molecular graph of the sample reactant are stored in the electronic device, input into the electronic device by a user, or obtained by the electronic device from a database, and the database stores connection trees of multiple substances.
[0153] 502. The electronic device determines a loss value of the retrosynthesis model based on the first connection tree and the second connection tree.
[0154] The loss value can represent the accuracy of the inverse synthesis model, and the error generated by the inverse synthesis model can be understood through the loss value, so that the inverse synthesis model can be trained based on the loss value to improve the accuracy of the inverse synthesis model.
[0155] Optionally, the electronic device inputs the first connection tree into the inverse synthesis model, and the inverse synthesis model determines a third connection tree based on the first connection tree, the third connection tree being a connection tree of the sample reactant predicted by the inverse synthesis model, and a loss value of the inverse synthesis model can be determined based on the third connection tree and the second connection tree.
[0156] Optionally, in the above embodiment, the inverse synthesis model can perform three tasks: a completion prediction task, a node type prediction task, and a connection position prediction task. The completion prediction task refers to predicting whether the connection tree needs to be completed and whether a node needs to be added to the connection tree of the structural unit. The node type prediction task refers to predicting the type of node to be added, i.e., predicting the target type. The connection position prediction task refers to predicting the first connection position and the second connection position.
[0157] Correspondingly, when training the inverse synthesis model, the inverse synthesis model performs any one of the above three tasks, and each time a task is performed, a corresponding loss value is generated. Therefore, the loss value of the inverse synthesis model can include at least one of a first loss value, a second loss value, and a third loss value.
[0158] The first loss value is the loss value of the first connection position and the second connection position obtained by the inverse synthesis model after inputting the first connection tree into the inverse synthesis model. The first connection position is the position at which a new node needs to be connected in the first connection tree, and the second connection position is the position at which the first connection tree needs to be connected to the new node. For example, after inputting the first connection tree into the inverse synthesis model, the inverse synthesis model predicts the first connection position and the second connection position. The first connection position and the second connection position are labeled in the second connection tree. By comparing the predicted first connection position and the second connection position with the actual connection positions in the second connection tree, the first loss value can be obtained.
[0159] The second loss value is the loss value of the second parameter obtained by the inverse synthesis model after inputting the first connection tree into the inverse synthesis model. The second parameter represents the probability of completing the first connection tree. For example, after inputting the first connection tree into the inverse synthesis model, the inverse synthesis model predicts the second parameter. By comparing the first connection tree and the second connection tree, it can be determined whether the first connection tree needs to be completed, i.e., the true second parameter is obtained. Based on the predicted second parameter and the true second parameter, the second loss value can be determined.
[0160] The third loss value is a loss value of the target type obtained by the inverse synthesis model after inputting the first connection tree into the inverse synthesis model, and the target type is a type of a new node to be added on the first connection tree. For example, after inputting the first connection tree into the inverse synthesis model, the inverse synthesis model predicts a target type, and the second connection tree is labeled with a type label of a newly added node. Based on the target type and the type label, the third loss value can be obtained.
[0161] 503 The electronic device trains the inverse synthesis model based on the loss value, so that the loss value of the trained inverse synthesis model is reduced.
[0162] The trained inverse synthesis model is used to determine the connection tree of the reactant based on the connection tree of any structural unit.
[0163] In the embodiments of the present application, the electronic device trains the inverse synthesis model based on the loss value, that is, updates the model parameters of the inverse synthesis model based on at least one of the first loss value, the second loss value, and the third loss value, so that the loss value of the trained inverse synthesis model is reduced, and the trained inverse synthesis model can more accurately determine the connection tree of the reactant based on the connection tree of any structural unit.
[0164] Optionally, the first loss value, the second loss value, and the third loss value each have a corresponding weight, and the first loss value, the second loss value, and the third loss value are weighted, such as weighted summation or weighted average, based on the respective weights, to obtain the loss value of the inverse synthesis model, based on which the inverse synthesis model is trained.
[0165] The weight of each loss value is 1 by default, or can be adjusted according to specific needs. Adjusting the weight of the loss value can change the convergence speed when training the inverse synthesis model. If the weight is too small, the convergence speed is slow, and it takes a long time for the model to reach the optimal. If the weight is too large, the model may not converge, so a suitable weight needs to be selected during training. Moreover, the convergence speeds of each task corresponding to the first loss value, the second loss value, and the third loss value in the inverse synthesis model are different. For example, the task corresponding to the first loss value converges quickly and reaches the optimal first, while the other tasks have not reached the optimal and continue to be trained. This task may be overfitting, so the weight corresponding to the first loss value can be adjusted to control the convergence speed of the task and avoid overfitting.
[0166] The embodiments of the present application provide a method for training an inverse synthesis model, which can train the inverse synthesis model based on the first connection tree of the sample structural unit and the second connection tree of the sample reactant, improve the accuracy of the inverse synthesis model, and enable the inverse synthesis model obtained by training to determine the connection tree of the reactant based on the connection tree of any structural unit, thereby improving the efficiency of inverse synthesis.
[0167] Optionally, before training the inverse synthesis model through the above embodiments, the inverse synthesis model can also be pre-trained. The pre-training process includes: selecting at least one connection tree pair from the data set, each connection tree pair including two connection trees of structural units, and training the inverse synthesis model based on the at least one connection tree pair to enable the inverse synthesis model to distinguish different positions in the two connection trees in the same connection tree pair.
[0168] The data set includes multiple connection trees of structural units, and the two connection trees in the same connection tree pair satisfy a similarity condition, i.e., the structures of the two connection trees are similar. For example, there is only one connection position different or only one atom different in the two connection trees.
[0169] For each connection tree pair, a label is labeled at the different position in the two connection trees, the connection tree pair is input into the inverse synthesis model, the inverse synthesis model determines the different position in the two connection trees in the connection tree pair, determines a loss value based on the determined different position in the two connection trees and the labeled label, and updates the model parameters of the inverse synthesis model based on the loss value. Through one or more training, the inverse synthesis model can have the ability to distinguish the different positions in the two connection trees in the same connection tree pair. This ability facilitates the inverse synthesis model to extract features of each connection tree, and subsequently formally trains the inverse synthesis model based on the first connection tree of the sample structural unit and the second connection tree of the sample reactant, which can speed up the training speed.
[0170] During the training process, it is necessary to determine the correspondence between the structural unit and the reactant, i.e., which reactant should be obtained after the structural unit is completed. In the case of limited training sample data, only part of the corresponding reactants of the structural units can be determined. If these training samples are directly used to train the inverse synthesis model, it will result in insufficient data and affect the training effect. The above pre-training process only needs the connection trees of similar structural units, and does not need to know the correspondence between the structural unit and the reactant, i.e., even if the corresponding reactant of the structural unit is unknown, the training can still be realized. Therefore, it can well make up for the insufficient data of the training sample in the training process of the inverse synthesis model, and still ensure the training effect of the inverse synthesis model.
[0171] On the basis of the above embodiments, the step of determining the reaction center of the target product can be performed based on a reaction center recognition model, and the molecular graph of the target product is input into the reaction center recognition model, and the reaction center recognition model can determine the reaction center in the molecular graph of the target product.
[0172] In order to ensure the accuracy of the reaction center recognition model, the reaction center recognition model needs to be trained, and the training process of the reaction center recognition model will be described below.
[0173] Figure 6 is a flowchart of a reaction center recognition model training method provided by an embodiment of the present application, as shown in the figure, the method is executed by an electronic device. The method comprises: Figure 6
[0174] 601. The electronic device obtains a molecular graph and a position identifier of a sample product.
[0175] The position identifier indicates the position of the reaction center of the sample product in the molecular graph. The molecular graph and the position identifier of the sample product are stored in the electronic device, or input by a user to the electronic device, or obtained by other means, which is not limited by the embodiments of the present application.
[0176] 602. The electronic device trains a reaction center recognition model based on the molecular graph and the position identifier, so that the trained reaction center recognition model is used to determine the reaction center in the molecular graph of any product.
[0177] In the embodiments of the present application, the electronic device inputs the molecular graph of the sample product to the reaction center recognition model, the reaction center recognition model determines the position of the reaction center in the molecular graph of the sample product, obtains the predicted position identifier, the reaction center recognition model determines the loss value based on the true position identifier and the predicted position identifier of the reaction center, and updates the model parameters of the reaction center recognition model based on the loss value, so that the trained reaction center recognition model can more accurately determine the reaction center in the molecular graph of any product.
[0178] Optionally, before training the reaction center recognition model through the above-mentioned embodiments, the reaction center recognition model can also be pre-trained. The pre-training process comprises: selecting at least one connection tree pair from the data set, each connection tree pair comprising two connection trees of structural units, and training the reaction center recognition model based on the at least one connection tree pair, so that the reaction center recognition model can distinguish different positions in the two connection trees in the same connection tree pair.
[0179] For each connection tree pair, different positions in the two connection trees are labeled with labels, the connection tree pair is input to the reaction center recognition model, the reaction center recognition model determines the different positions in the two connection trees in the connection tree pair, determines the loss value based on the determined different positions in the two connection trees and the labeled labels, and updates the model parameters of the reaction center recognition model based on the loss value. Through one or more training, the reaction center recognition model can have the ability to distinguish different positions in the two connection trees in the same connection tree pair. This ability facilitates the reaction center recognition model to extract the features of each connection tree, and subsequently formally trains the reaction center recognition model based on the molecular graph and the position identifier of the sample product, which can speed up the training speed.
[0180] The position of the reaction center in the molecular graph of the sample product needs to be determined in the training process. In the case of limited amount of training sample data, the position of the reaction center in the molecular graph of only part of the sample product can be determined. If the reaction center recognition model is directly trained using these training samples, the amount of data is insufficient, which affects the training effect. The pre-training process described above only needs the connection tree of the similar structural unit, and the training can be realized without pre-knowing the position of the reaction center of the sample product. Therefore, the insufficient amount of training sample data in the training process of the reaction center recognition model can be well compensated, and the training effect of the reaction center recognition model can still be ensured.
[0181] For example, Figure 7 For example, Figure 7 is a flowchart of training an inverse synthesis model. The data set includes molecular graphs of multiple structural units. The molecular graph of each structural unit in the data set is converted into a connection tree. Then, a plurality of connection tree pairs are selected, each connection tree pair including the connection trees of two structural units. Based on the plurality of connection tree pairs, the reaction center recognition model and the inverse synthesis model are pre-trained. The pre-trained reaction center recognition model and the inverse synthesis model can distinguish different positions in the two connection trees in the same connection tree pair. After pre-training, the reaction center recognition model is first trained. After training is completed, the inverse synthesis model can be trained using the reaction center recognition model. The process of training the reaction center recognition model is the same as that of the above Figure 6 embodiment. Here, only the training process of the inverse synthesis model is exemplified.
[0182] The sample product is molecule A. The molecular graph of molecule A is input into the reaction center recognition model. The reaction center recognition model determines the reaction center in the molecular graph of molecule A. Based on the reaction center, the molecular graph of molecule A is segmented to obtain the molecular graph of at least one sample structural unit. The molecular graph of the sample structural unit is converted into a connection tree and input into the encoding layer of the inverse synthesis model. The encoding layer encodes the connection tree of the sample structural unit to obtain the features of the connection tree of the sample structural unit. The decoding layer completes the connection tree based on the features, thereby obtaining the connection tree of molecule B. The inverse synthesis model is trained based on the connection tree of molecule B and the connection tree of the sample reactant.
[0183] Figure 8 is a structural diagram of an inverse synthesis device provided by an embodiment of the present application. The device is configured in an electronic device, as Figure 8 shown, the device includes:
[0184] The reaction center determination module 801 is configured to determine the reaction center of the target product in the molecular graph of the target product.
[0185] The segmentation module 802 is configured to segment the molecular graph of the target product based on the reaction center to obtain the molecular graph of at least one structural unit.
[0186] The conversion module 803 is configured to convert the molecular graph of each structure unit into a connection tree.
[0187] The completion module 804 is configured to determine a target node of each structure unit based on the connection tree of each structure unit, the target node being a new node that needs to be added on the connection tree of the structure unit.
[0188] The completion module 804 is configured to add the target node of each structure unit on the connection tree of each structure unit to obtain a connection tree of a reactant corresponding to each structure unit, the reactant being a substance used for synthesizing the target product.
[0189] Optionally, the completion module 804 is configured to:
[0190] encode the connection tree of each structure unit to obtain a feature of the connection tree of each structure unit;
[0191] determine the target node of each structure unit based on the feature of the connection tree of each structure unit.
[0192] Optionally, the completion module 804 is configured to:
[0193] determine a target type from a plurality of preset types based on the feature of the reaction center, the feature of the connection tree of the structure unit, and the plurality of preset types, each preset type being indicative of a node type;
[0194] determine a node belonging to the target type as the target node of the structure unit.
[0195] Optionally, the completion module 804 includes:
[0196] a determination unit configured to determine a first connection position and a second connection position based on the connection tree of the structure unit and the target node of the structure unit, the first connection position being a position on the connection tree at which the target node needs to be connected, the second connection position being a position on the target node at which the connection tree needs to be connected, and the first connection position and the second connection position being of the same type;
[0197] a merging unit configured to merge the first connection position on the connection tree of the structure unit and the second connection position on the target node to obtain the connection tree of the reactant.
[0198] Optionally, the determination unit is configured to:
[0199] determine a plurality of position pairs, each position pair including a first position on the connection tree and a second position on the target node, and the two positions in the same position pair being of the same type;
[0200] determine, based on features of the first position and the second position in each position pair, a first parameter of each position pair, the first parameter representing a probability of combination of the first position and the second position in the position pair;
[0201] select, from the plurality of position pairs, a target position pair with a maximum first parameter, determine the first position in the target position pair as the first connection position, and determine the second position in the target position pair as the second connection position.
[0202] Optionally, the apparatus further includes:
[0203] a molecular graph completion module configured to add, in the molecular graph of each structural unit, a molecular structure corresponding to the target node of each structural unit, to obtain a molecular graph of each reactant.
[0204] Optionally, the apparatus further includes:
[0205] a second parameter determination module configured to determine, based on features of the reaction center and features of the connection tree of at least one structural unit, a second parameter, the second parameter representing a probability of completion of the connection tree;
[0206] a completion module 804 configured to, in a case where the second parameter is greater than a preset threshold, perform the steps of determining the target node of each structural unit and adding the target node of each structural unit on the connection tree of each structural unit;
[0207] a connection tree determination module configured to, in a case where the second parameter is not greater than the preset threshold, determine the connection tree of each structural unit as the connection tree of one reactant.
[0208] Optionally, the steps of determining the target node of each structural unit and adding the target node of each structural unit on the connection tree of each structural unit are performed based on an inverse synthesis model, and the apparatus further includes:
[0209] a sample acquisition module configured to acquire a first connection tree of a sample structural unit and a second connection tree of a sample reactant, wherein the sample structural unit is obtained by disconnecting a sample product at the reaction center, and the sample product is a product obtained after a chemical reaction of the sample reactant;
[0210] a training module configured to train the inverse synthesis model based on the first connection tree and the second connection tree, so that the trained inverse synthesis model is used to determine the connection tree of the reactant based on the connection tree of any structural unit.
[0211] Optionally, the training module is configured to:
[0212] determine, based on the first connection tree and the second connection tree, a loss value of the inverse synthesis model, the loss value including at least one of a first loss value, a second loss value, and a third loss value;
[0213] train the inverse synthesis model based on the loss value, so that the loss value of the trained inverse synthesis model is reduced;
[0214] The first loss value is the loss value of the first connection position and the second connection position obtained by the inverse synthesis model after the first connection tree is input into the inverse synthesis model, the first connection position is a position on the first connection tree where a new node needs to be connected, and the second connection position is a position on the new node where the first connection tree needs to be connected.
[0215] The second loss value is the loss value of the second parameter obtained by the inverse synthesis model after the first connection tree is input into the inverse synthesis model, and the second parameter represents the probability of completing the connection tree.
[0216] The third loss value is the loss value of the target type obtained by the inverse synthesis model after the first connection tree is input into the inverse synthesis model, and the target type is the type of the new node that needs to be added on the first connection tree.
[0217] Optionally, the device further comprises:
[0218] The connection tree pair acquisition module is configured to select at least one connection tree pair from the data set, each connection tree pair comprising two connection trees of structural units, and the two connection trees in the same connection tree pair satisfying a similarity condition, and the data set comprising connection trees of multiple structural units.
[0219] The training module is configured to train the inverse synthesis model based on the at least one connection tree pair, so that the inverse synthesis model distinguishes different positions in the two connection trees in the same connection tree pair.
[0220] Optionally, the step of determining the reaction center of the target product is performed based on a reaction center identification model, and the device further comprises:
[0221] The sample acquisition module is configured to acquire a molecular graph of a sample product and a position identifier, the position identifier representing a position of a reaction center of the sample product in the molecular graph.
[0222] The training module is configured to train the reaction center identification model based on the molecular graph and the position identifier, so that the trained reaction center identification model is used to determine the reaction center in the molecular graph of any product.
[0223] Optionally, the device further comprises:
[0224] The connection tree pair acquisition module is configured to select at least one connection tree pair from the data set, each connection tree pair comprising two connection trees of structural units, and the two connection trees in the same connection tree pair satisfying a similarity condition, and the data set comprising connection trees of multiple structural units.
[0225] The training module is configured to train the reaction center identification model based on the at least one pair of connection trees, so that the reaction center identification model distinguishes different positions in two connection trees in the same pair of connection trees.
[0226] The embodiment of the present application provides a device for reverse synthesis, which can split a molecular graph of a target product based on a reaction center to obtain a molecular graph of at least one structural unit, simulate a scenario of disconnecting the target product from the reaction center, and then no longer directly process the molecular graph, but convert the molecular graph of the structural unit into a connection tree to represent the structural unit in the form of the connection tree, so as to process the connection tree of the structural unit to obtain a connection tree of a reactant. Since the connection tree has a relatively simple form of expression, is more intuitive and concise than the molecular graph, and has a relatively simple processing flow, the device is still applicable even in the case that the molecular structure of the target product is relatively complex. Therefore, the processing efficiency of the embodiment of the present application is higher, and the application range is wider.
[0227] It should be noted that: the device for reverse synthesis provided in the above embodiment is only exemplified by the division of the above functional modules, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the electronic device is divided into different functional modules to complete all or part of the functions described above. In addition, the device for reverse synthesis and the method for reverse synthesis provided in the above embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0228] The embodiment of the present application further provides an electronic device, which includes a processor and a memory, and the memory stores at least one computer program, which is loaded and executed by the processor to realize the operations performed by the electronic device in the reverse synthesis method of the above embodiment.
[0229] Figure 9 The structure schematic diagram of the electronic device 900 provided by an example embodiment of the present application is shown.
[0230] The electronic device 900 includes a processor 901 and a memory 902.
[0231] The processor 901 can include one or more processing cores, such as a 4-core processor, an 8-core processor, and the like. The processor 901 can be implemented in the form of at least one of a DSP (Digital Signal Processing), an FPGA (Field Programmable Gate Array), a PLA (Programmable Logic Array). In some embodiments, the processor 901 can further include an AI (Artificial Intelligence) processor for processing a machine learning-related computing operation.
[0232] The memory 902 can include one or more computer-readable storage media that can be non-transitory. The memory 902 can further include a high-speed random access memory, and a nonvolatile memory such as one or more disk storage devices, flash storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 902 is used to store at least one computer program for being executed by the processor 901 to implement the inverse synthesis method provided by the method embodiments in the present application.
[0233] In some embodiments, the electronic device 900 can further optionally include a peripheral device interface 903 and at least one peripheral device. The processor 901, the memory 902, and the peripheral device interface 903 can be connected through a bus or a signal line. Each peripheral device can be connected to the peripheral device interface 903 through a bus, a signal line, or a circuit board. Optionally, the peripheral device includes at least one of a radio frequency circuit 904, a display screen 905, a camera component 906, and an audio circuit 907.
[0234] The peripheral device interface 903 can be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 901 and the memory 902. In some embodiments, the processor 901, the memory 902, and the peripheral device interface 903 are integrated on the same chip or circuit board.
[0235] The radio frequency circuit 904 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 904 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 904 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals.
[0236] Display screen 905 is used to display a user interface (UI). This UI may include graphics, text, icons, video, and any combination thereof. When display screen 905 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 901 for processing. In this case, display screen 905 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard.
[0237] The camera assembly 906 is used to acquire images or videos. Optionally, the camera assembly 906 includes a front-facing camera and a rear-facing camera. The front-facing camera is disposed on the front panel of the electronic device 900, and the rear-facing camera is disposed on the back of the electronic device 900. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions.
[0238] The audio circuit 907 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals that are input to the processor 901 for processing, or input to the radio frequency circuit 904 to realize voice communication.
[0239] In some embodiments, the electronic device 900 further includes one or more sensors 908. The one or more sensors 908 include, but are not limited to, an accelerometer 909, a gyroscope 910, a pressure sensor 911, an optical sensor 912, and a proximity sensor 913.
[0240] Those skilled in the art will understand that Figure 9 The structure shown does not constitute a limitation on the electronic device 900, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0241] This application also provides a computer-readable storage medium storing at least one computer program, which is loaded and executed by a processor to implement the operations performed by the inverse synthesis method of the above embodiments.
[0242] This application also provides a computer program product, including a computer program loaded and executed by a processor to perform the operations performed by the inverse synthesis method of the above embodiments.
[0243] In some embodiments, the computer program involved in the embodiments of the present application can be deployed on a computer device to execute, or on multiple computer devices located in one place to execute, or on multiple computer devices distributed in multiple places and interconnected through a communication network to execute. The multiple computer devices distributed in multiple places and interconnected through a communication network can constitute a blockchain system.
[0244] Those of ordinary skill in the present art can understand that all or part of the steps of the above-mentioned embodiments can be completed by hardware, or by a program instructing relevant hardware, and the program can be stored in a computer readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.
[0245] The above is only an optional embodiment of the present application, and does not limit the embodiments of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the embodiments of the present application shall be included in the protection scope of the present application.
Claims
1. A method of reverse synthesis, characterized in that, The method comprises: determining a reaction center of the target product in a molecular graph of the target product; segmenting the molecular graph of the target product based on the reaction center to obtain a molecular graph of at least one structural unit; converting the molecular graph of the at least one structural unit into a connection tree; determining a target node of each structural unit based on the connection tree of each structural unit, the target node being a new node to be added on the connection tree of the structural unit; adding the target node of each structural unit on the connection tree of each structural unit to obtain a connection tree of a reactant corresponding to each structural unit, wherein the reactant is a substance used to synthesize the target product.
2. The method of claim 1, wherein, The method further comprises: encoding the connection tree of each structural unit to obtain a feature of the connection tree of each structural unit; determining the target node of each structural unit based on the feature of the connection tree of each structural unit.
3. The method of claim 1, wherein, The method further comprises: determining a target type from a plurality of preset types based on the feature of the reaction center, the feature of the connection tree of the structural unit, and the features of the plurality of preset types, wherein each preset type refers to a node type; determining a node belonging to the target type as the target node of the structural unit.
4. The method of claim 1, wherein, The method further comprises: determining a first connection position and a second connection position based on the connection tree of the structural unit and the target node of the structural unit, the first connection position being a position on the connection tree where the target node needs to be connected, the second connection position being a position on the target node where the connection tree needs to be connected, and the first connection position and the second connection position being of the same type; merging the first connection position on the connection tree of the structural unit and the second connection position on the target node to obtain the connection tree of the reactant.
5. The method of claim 4, wherein, The method further comprises: determining a plurality of position pairs, wherein each position pair comprises a first position on the connection tree and a second position on the target node, and the two positions in the same position pair are of the same type; determining a first parameter of each position pair based on the features of the first position and the second position in the position pair, the first parameter representing the probability of merging the first position and the second position in the position pair; selecting a target position pair with the largest first parameter from the plurality of position pairs, determining the first position in the target position pair as the first connection position, and determining the second position in the target position pair as the second connection position.
6. The method of claim 1, wherein, The method further comprises: adding a molecular structure corresponding to the target node of each structure unit in a molecular graph of the each structure unit to obtain a molecular graph of each reactant.
7. The method of claim 1, wherein, The method further comprises: determining a second parameter based on the feature of the reaction center and the feature of the connection tree of the at least one structure unit, the second parameter representing a probability of completing the connection tree; in a case where the second parameter is greater than a preset threshold, performing the steps of determining the target node of each structure unit and adding the target node of each structure unit on the connection tree of the each structure unit; in a case where the second parameter is not greater than the preset threshold, determining the connection tree of each structure unit as the connection tree of one reactant.
8. The method of claim 1, wherein, The steps of determining the target node of each structure unit and adding the target node of each structure unit on the connection tree of the each structure unit are performed based on an inverse synthesis model, and the method further comprises: obtaining a first connection tree of a sample structure unit and a second connection tree of a sample reactant, wherein the sample structure unit is obtained by disconnecting a sample product at a reaction center, and the sample product is a product obtained after a chemical reaction of the sample reactant; training the inverse synthesis model based on the first connection tree and the second connection tree, so that the trained inverse synthesis model is used to determine the connection tree of a reactant based on the connection tree of any structure unit.
9. The method of claim 8, wherein, The training of the inverse synthesis model based on the first connection tree and the second connection tree comprises: determining a loss value of the inverse synthesis model based on the first connection tree and the second connection tree, the loss value comprising at least one of a first loss value, a second loss value and a third loss value; training the inverse synthesis model based on the loss value, so that the loss value of the trained inverse synthesis model is reduced; The first loss value is a loss value of first connection positions and second connection positions obtained by the inverse synthesis model after the first connection tree is input into the inverse synthesis model, the first connection positions are positions on the first connection tree where new nodes need to be connected, and the second connection positions are positions on the new nodes where the first connection tree needs to be connected. The second loss value is a loss value of a second parameter obtained by the inverse synthesis model after the first connection tree is input into the inverse synthesis model, the second parameter representing a probability of completing the connection tree. The third loss value is a loss value of a target type obtained by the inverse synthesis model after the first connection tree is input into the inverse synthesis model, the target type being a type of new nodes that need to be added on the first connection tree.
10. The method of claim 8, wherein, Before the training of the inverse synthesis model based on the first connection tree and the second connection tree, the method further comprises: selecting at least one connection tree pair from a data set, each connection tree pair comprising connection trees of two structure units, and the two connection trees in the same connection tree pair satisfying a similarity condition, the data set comprising connection trees of a plurality of structure units; The inverse synthesis model is trained based on the at least one pair of connection trees, so that the inverse synthesis model distinguishes different positions in two connection trees in the same pair of connection trees.
11. The method of claim 1, wherein, The step of determining the reaction center of the target product is performed based on a reaction center identification model, and the method further comprises: obtaining a molecular graph and a position identifier of a sample product, the position identifier representing a position of a reaction center of the sample product in the molecular graph; training the reaction center identification model based on the molecular graph and the position identifier, so that the trained reaction center identification model is used to determine a reaction center in a molecular graph of any product.
12. The method of claim 11, wherein, Before the training of the reaction center identification model based on the molecular graph and the position identifier, the method further comprises: selecting at least one pair of connection trees from a data set, each pair of connection trees comprising two connection trees of structural units, and the two connection trees in the same pair of connection trees satisfying a similarity condition, the data set comprising a plurality of connection trees of structural units; training the reaction center identification model based on the at least one pair of connection trees, so that the reaction center identification model distinguishes different positions in two connection trees in the same pair of connection trees.
13. A reverse synthesis device, characterized by The device comprises: a reaction center determination module configured to determine a reaction center of a target product in a molecular graph of the target product; a segmentation module configured to segment the molecular graph of the target product based on the reaction center, to obtain at least one molecular graph of a structural unit; a conversion module configured to convert the at least one molecular graph of a structural unit into a connection tree; a completion module configured to determine a target node of each structural unit based on the connection tree of each structural unit, the target node being a new node to be added to the connection tree of the structural unit; the completion module is configured to add the target node of each structural unit to the connection tree of each structural unit, to obtain a connection tree of a reactant corresponding to each structural unit, wherein the reactant is a substance used to synthesize the target product.
14. An electronic device, comprising: The electronic device comprises a processor and a memory, and the memory stores at least one computer program, which is loaded and executed by the processor to implement the operations performed by the inverse synthesis method according to any one of claims 1 to 12.
15. A computer-readable storage medium, characterized in that, The computer readable storage medium stores at least one computer program, which is loaded and executed by the processor to implement the operations performed by the inverse synthesis method according to any one of claims 1 to 12.
16. A computer program product comprising a computer program, characterized in that, The computer program is loaded and executed by the processor to implement the operations performed by the inverse synthesis method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Artificial intelligence-based retrosynthesis prediction method and device, equipment and storage medium
CN111524557A
Deep learning-based inverse synthesis prediction method and device, medium and equipment
CN114220496A