Training Method, Device and Electronic Device for Molecular Graph Reconstruction Model

By introducing graph matching modules and pre-training modules into the molecular graph reconstruction model, calculating the relationship matrix and supervising the reconstruction loss, the problem that the graph generation model cannot effectively supervise the output during the molecular discovery process is solved, and the applicability and efficiency of the molecular graph reconstruction model in the molecular discovery process is realized.

CN114334040BActive Publication Date: 2025-05-30TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111421790.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-26
Publication Date
2025-05-30
Estimated Expiration
2041-11-26

AI Technical Summary

Technical Problem

The existing graph generation model cannot effectively monitor whether the outputted reconstructed samples have the desired chemical properties during molecular discovery, resulting in inappropriate use of molecular discovery.

Method used

A training method for molecular graph reconstruction model is designed. By introducing a graph matching module and a pre-training module, the reconstruction loss is calculated using the relation matrix to realize the supervision of whether the reconstructed molecular graph output by the decoder is a candidate molecule of the desired chemical properties.

Benefits of technology

This method makes the molecular graph reconstruction model suitable for molecular discovery processes, reduces the matching complexity and resource consumption, improves the matching accuracy and the guidance effect of the graph matching module on the molecular graph reconstruction model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114334040B_ABST
    Figure CN114334040B_ABST
Patent Text Reader

Abstract

The present application provides a training method, device, and electronic device for a molecular graph reconstruction model. The method is applicable to the field of artificial intelligence for drugs. The molecular graph reconstruction model includes an encoder, a decoder, and a graph matching module. In the training method provided by the present application, on the one hand, the graph matching module is designed as a pre-training module, and the relationship matrix output by the graph matching module is designed to calculate the reconstruction loss between the sample molecular graph and the reconstructed molecular graph. On the other hand, the representation vector of the sample molecular graph output by the encoder is designed as the input vector of the decoder. The training method of the molecular graph reconstruction model provided by the present application can not only make the molecular graph reconstruction model applicable to the molecular discovery process, but also reduce the matching complexity and resource consumption, improve the accuracy of matching, and further improve the guiding effect of the graph matching module on the molecular graph reconstruction model. In addition, the method provided by the present application can also improve the practicality and reconstruction effect of the molecular graph reconstruction model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of artificial intelligence for drugs, and more specifically, to a training method, device, and electronic device for a molecular graph reconstruction model. Background Art

[0002] The purpose of molecular discovery is to find candidate molecules with desired chemical properties, which is a long, costly, and high-failure-rate process.

[0003] Since graph generation models can explore large latent spaces in a data-driven manner, they have shown great potential in accelerating the molecular discovery process.

[0004] As far as those skilled in the art know, the graph generation model based on the autoencoder is a model that can be used to generate reconstructed graphs. Specifically, the sample graph can be first encoded into the latent space, and then the decoder can be used to predict the node features and the topological structure of the graph in a decoding manner to realize the reconstruction of the sample graph. However, due to the invariance of graph permutation, the reconstruction loss adopted by the graph generation model cannot supervise whether the output reconstructed sample is a candidate molecule with the desired chemical properties, which in turn leads to the inapplicability of the graph generation model to the molecular discovery process.

[0005] Therefore, there is an urgent need in the art for a training method for a molecular graph reconstruction model, so as to discover candidate molecules with desired chemical properties by using the molecular graph reconstruction model. Summary of the Invention

[0006] The present application provides a training method, device, and electronic device for a molecular graph reconstruction model. The training method can not only make the molecular graph reconstruction model applicable to the molecular discovery process, but also reduce the matching complexity and resource consumption, improve the accuracy of matching, and further improve the guiding effect of the graph matching module on the molecular graph reconstruction model. In addition, the training method for the molecular reconstructed graph model provided by the present application can also improve the practicability and reconstruction effect of the molecular graph reconstruction model.

[0007] On the one hand, the present application provides a training method for a molecular graph reconstruction model, which includes an encoder, a decoder, and a graph matching module; the method includes:

[0008] Obtain a sample molecular graph;

[0009] Process the attribute values of the sample molecular graph by using the encoder to obtain a representation vector of the sample molecular graph;

[0010] Use the decoder to reconstruct the sample molecular graph based on the representation vector of the sample molecular graph to obtain a reconstructed molecular graph;

[0011] Use the graph matching module to predict the node correspondence and edge correspondence between the sample molecular graph and the reconstructed molecular graph, and obtain a relationship matrix;

[0012] Based on the relationship matrix, compare the sample molecular graph and the reconstructed molecular graph to obtain the reconstruction loss of the reconstructed molecular graph, where the reconstruction loss is used to characterize the difference between the sample molecular graph and the reconstructed molecular graph;

[0013] Based on the reconstruction loss, adjust the decoder to obtain a trained molecular graph reconstruction model.

[0014] On the other hand, the present application provides a training device for a molecular graph reconstruction model. The molecular graph reconstruction model includes an encoder, a decoder, and a graph matching module. The training device includes:

[0015] An acquisition unit for acquiring a sample molecular graph;

[0016] A processing unit for using the encoder to process the attribute values of the sample molecular graph to obtain a representation vector of the sample molecular graph;

[0017] A reconstruction unit for using the decoder to reconstruct the sample molecular graph based on the representation vector of the sample molecular graph to obtain a reconstructed molecular graph;

[0018] A prediction unit for using the graph matching module to predict the node correspondence and edge correspondence between the sample molecular graph and the reconstructed molecular graph to obtain a relationship matrix;

[0019] A calculation unit for comparing the sample molecular graph and the reconstructed molecular graph based on the relationship matrix to obtain the reconstruction loss of the reconstructed molecular graph, where the reconstruction loss is used to characterize the difference between the sample molecular graph and the reconstructed molecular graph;

[0020] An adjustment unit for adjusting the decoder based on the reconstruction loss to obtain a trained molecular graph reconstruction model.

[0021] On the other hand, the present application provides a method for constructing a molecular graph, including:

[0022] Obtain a random vector;

[0023] Using the decoder in the molecular graph reconstruction model with the random vector and a preset attribute label as inputs, construct a molecular graph with the preset attribute label; wherein, the molecular graph reconstruction model is a model trained according to the method described in the first aspect.

[0024] On the other hand, the present application provides a method for constructing a molecular graph, including:

[0025] Obtain a molecular graph to be optimized;

[0026] Taking the molecular graph to be optimized as input, using the encoder in the molecular graph reconstruction model to process the molecular graph to be optimized, and obtaining a representation vector of the molecular graph to be optimized; the molecular graph reconstruction model is a model trained according to the method described in the first aspect;

[0027] Taking the representation vector of the molecular graph to be optimized and a preset attribute label as input, and using the decoder in the molecular graph reconstruction model to construct a molecular graph obtained by optimizing the molecular graph to be optimized based on the preset attribute label.

[0028] On the other hand, the present application provides a method for predicting molecular graph attributes, including:

[0029] Obtaining a molecular graph to be predicted;

[0030] Taking the molecular graph to be predicted as input, using the encoder in the molecular graph reconstruction model to perform attribute prediction on the molecular graph to be predicted, and obtaining an attribute value of the molecular graph to be predicted; the molecular graph reconstruction model is a model trained according to the method described in the first aspect.

[0031] On the other hand, the present application provides an electronic device, including:

[0032] A processor adapted to implement computer instructions; and,

[0033] A computer-readable storage medium storing computer instructions, the computer instructions being adapted to be loaded and executed by the processor to perform the method described in the first aspect above.

[0034] On the other hand, an embodiment of the present application provides a computer-readable storage medium storing computer instructions, which, when read and executed by a processor of a computer device, cause the computer device to perform the method described in the first aspect above.

[0035] On the other hand, an embodiment of the present application provides a computer program product or computer program, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to perform the method described in the first aspect above.

[0036] Based on the above technical solutions, the present application designs the graph matching module as a pre-training module, and designs the relationship matrix output by the graph matching module to calculate the reconstruction loss between the sample molecular graph and the reconstructed molecular graph. This can not only make the molecular graph reconstruction model applicable to the molecular discovery process, but also reduce the matching complexity and resource consumption, and improve the accuracy of matching. Furthermore, it improves the guiding effect of the graph matching module on the molecular graph reconstruction model.

[0037] Specifically, the graph matching module is designed as a relationship matrix for characterizing the node correspondence and edge correspondence between the input sample molecular graph and the generated reconstructed molecular graph. Then, based on this relationship matrix, the reconstruction loss between the input sample molecular graph and the generated reconstructed molecular graph is calculated, which can make the molecular graph reconstruction model applicable to the molecular discovery process. It is beneficial to supervise whether the reconstructed molecular graph output by the decoder is a candidate molecule with the required chemical properties through this reconstruction loss, making the molecular graph reconstruction model applicable to the molecular discovery process. In addition, designing the graph matching module in the molecular graph reconstruction model as a pre-training module can not only reduce the matching complexity and resource consumption compared with the traditional matching method, but also improve the accuracy of matching. Furthermore, it can improve the guiding effect of the graph matching module on the molecular graph reconstruction model.

[0038] In addition, the present application designs the characterization vector of the sample molecular graph output by the encoder as the input vector of the decoder, which can not only improve the practicability of the molecular graph reconstruction model, but also enable the molecular graph reconstruction model to generate reconstructed molecular graphs with different property constraints, improving the reconstruction effect of the molecular graph reconstruction model.

[0039] Specifically, on the one hand, designing the characterization vector of the sample molecular graph output by the encoder as the input vector of the decoder can reconstruct the sample molecular graph to obtain the reconstructed molecular graph output by the decoder; on the other hand, during the molecular graph reconstruction process, with the preset attribute label as the input, a reconstructed molecular graph with composite property constraints can be obtained, improving the practicability of the molecular graph reconstruction model. On the other hand, the characterization vector of the molecular sample graph output by the encoder as an input vector of the decoder can be used as a property constraint of the decoder, and the preset attribute label as another input of the decoder can be used as another property constraint sample. Furthermore, a reconstructed molecular graph with multiple property constraints can be realized, improving the reconstruction effect of the molecular graph reconstruction model.

[0040] In summary, the training method of the molecular reconstruction graph model provided by the present application can not only make the molecular graph reconstruction model applicable to the molecular discovery process, but also reduce the matching complexity and resource consumption, improve the accuracy of matching, and further improve the guiding effect of the graph matching module on the molecular graph reconstruction model. In addition, the training method of the molecular reconstruction graph model provided by the present application can also improve the practicability and reconstruction effect of the molecular graph reconstruction model. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 FIG. is a schematic flowchart of a training method for a molecular graph reconstruction model provided by an embodiment of the present application.

[0042] Figure 2 FIG. is a schematic framework of a molecular graph reconstruction model provided by an embodiment of the present application.

[0043] Figure 3 FIG. is an application example of a decoder in a molecular graph reconstruction model provided by an embodiment of the present application.

[0044] Figure 4 FIG. is another application example of a decoder in a molecular graph reconstruction model provided by an embodiment of the present application.

[0045] Figure 5 FIG. is an application example of an encoder in a molecular graph reconstruction model provided by an embodiment of the present application.

[0046] Figure 6 FIG. is a schematic block diagram of a training device for a molecular graph reconstruction model provided by an embodiment of the present application.

[0047] Figure 7 FIG. is a schematic block diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0049] The solution provided by the present application may relate to the field of artificial intelligence (AI) technology.

[0050] Among them, AI is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making.

[0051] It should be understood that artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0052] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields. For example, common ones include smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, intelligent healthcare, intelligent customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0053] The embodiments of this application may be related to computer vision (CV) technology in artificial intelligence technology. Computer vision is a science that studies how to enable machines to "see". Further, it refers to using cameras and computers to replace human eyes to perform machine vision such as identifying, tracking, and measuring targets, and further performing graphic processing to make the computer-processed images more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to establish an artificial intelligence system that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.

[0054] Embodiments of the present application may also relate to Machine Learning (ML) in artificial intelligence. ML is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.

[0055] The present application also relates to the field of drugs. In the drug R & D process, after completing target identification and validation, it is necessary to screen candidate drug compounds. In the screening process, using molecular property prediction algorithms to predict the absorption, distribution, metabolism, excretion, and toxicity properties of molecules can help R & D personnel screen molecules, greatly improving the R & D efficiency and reducing the drug R & D cost.

[0056] The purpose of molecular discovery is to find candidate molecules with desired chemical properties, which is a long, high-cost, and high-failure-rate process. In the present application, molecular discovery can also be referred to as target discovery, drug discovery, or drug target discovery.

[0057] Since the graph generation model can explore a large latent space in a data-driven manner, it has shown great potential in accelerating the molecular discovery process.

[0058] As far as those skilled in the art know, the graph generation model based on the autoencoder is a model that can be used to generate reconstructed graphs. Specifically, the sample graph can be first encoded into the latent space, and then the decoder can be used to predict the node features and the topological structure of the graph in a decoding manner to achieve the reconstruction of the sample graph. However, due to the permutation invariance of the graph, the reconstruction loss adopted by the graph generation model cannot supervise whether the reconstructed sample output is a candidate molecule with the desired chemical properties, which in turn leads to the inapplicability of the graph generation model to the molecular discovery process.

[0059] Based on this, the present application provides a training method, apparatus, and electronic device for a molecular graph reconstruction model. This training method can not only make the molecular graph reconstruction model applicable to the molecular discovery process, but also reduce the matching complexity and resource consumption, improve the accuracy of matching, and thus enhance the guiding effect of the graph matching module on the molecular graph reconstruction model. In addition, the training method for the molecular reconstruction graph model provided by the present application can also improve the practicality and reconstruction effect of the molecular graph reconstruction model.

[0060] To facilitate the understanding of the solution provided by the present application, the following explains the relevant terms involved.

[0061] Refactoring: The process of converting a vector into a molecular graph.

[0062] Exemplarily, based on the characterization vector of a known molecular graph, a new molecular graph can be constructed under the guidance of preset attribute labels, or based on a random vector, a new molecular graph can be constructed under the guidance of preset attribute labels.

[0063] Characterization vector: Information of a characterization object (such as a molecular graph, node, or edge) represented in vector form.

[0064] Taking the characterization vector of a molecular graph as an example, the characterization vector of a molecular graph refers to the information of the molecular graph represented in vector form. The information of the molecular graph includes, but is not limited to, the topological structure of the molecular graph and the attribute information of the molecular graph. The topological structure includes the connection relationship between nodes and edges. In other words, the characterization vector of a molecular graph can reflect the topological structure and / or attribute information of the molecular graph. Exemplarily, the characterization vector of a molecular graph can be obtained by using deep learning. Deep learning is a branch field of machine learning, which can learn the representation of a molecular graph in vector form from the information of the molecular graph. The characterization vector of the molecular graph involved in the present application can be the representation output by any intermediate layer. For example, it can be the representation output by a hidden layer. At this time, the characterization vector of the molecular graph can also be called the characterization vector of the molecular graph in the hidden space or the hidden vector.

[0065] Relationship matrix: Used to characterize the corresponding relationship between nodes and the corresponding relationship between edges in two molecular graphs.

[0066] Exemplarily, molecule Figure 1 includes M nodes and N edges, M > 0, N > 0; molecule Figure 2 includes X nodes and Y edges, X > 0, Y > 0; molecule Figure 1 and molecule Figure 2 The relationship matrix between them can be used to characterize the matching degree between each of the M nodes and each of the X nodes, and the matching degree between each of the N edges and each of the Y edges. The value range of the matching degree can be [0, 1].

[0067] Figure 1 It is a schematic flowchart of a training method 100 for a molecular graph reconstruction model provided by an embodiment of the present application.

[0068] It should be noted that the solution provided by the embodiment of the present application can be executed by any electronic device with data processing capabilities. For example, the electronic device can be implemented as a server. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, and big data and artificial intelligence platforms. The servers can be directly or indirectly connected through wired or wireless communication methods, and the present application does not make any restrictions here.

[0069] It should be understood that the training method provided by the present application is applicable to a molecular graph reconstruction model, that is, any model including an encoder, a decoder, and a graph matching module.

[0070] As Figure 1 shown, the training method 100 may include:

[0071] S110, obtaining a sample molecular graph;

[0072] S120, using the encoder to process the attribute values of the sample molecular graph to obtain a representation vector of the sample molecular graph;

[0073] S130, using the decoder to reconstruct the sample molecular graph based on the representation vector of the sample molecular graph to obtain a reconstructed molecular graph;

[0074] S140, using the graph matching module to predict the node correspondence and edge correspondence between the sample molecular graph and the reconstructed molecular graph to obtain a relationship matrix;

[0075] S150, comparing the sample molecular graph and the reconstructed molecular graph based on the relationship matrix to obtain a reconstruction loss of the reconstructed molecular graph, where the reconstruction loss is used to characterize the difference between the sample molecular graph and the reconstructed molecular graph;

[0076] S160, adjusting the decoder based on the reconstruction loss to obtain a trained molecular graph reconstruction model.

[0077] Based on the above technical solutions, the graph matching module of this application is designed as a pre-training module, and the relationship matrix output by the graph matching module is designed to calculate the reconstruction loss between the sample molecular graph and the reconstructed molecular graph. This can not only make the molecular graph reconstruction model applicable to the molecular discovery process, but also reduce the matching complexity and resource consumption, and improve the accuracy of matching. Furthermore, it improves the guiding effect of the graph matching module on the molecular graph reconstruction model.

[0078] Specifically, the graph matching module is designed as a relationship matrix for characterizing the node correspondence and edge correspondence between the input sample molecular graph and the generated reconstructed molecular graph. Then, based on this relationship matrix, the reconstruction loss between the input sample molecular graph and the generated reconstructed molecular graph is calculated, which can make the molecular graph reconstruction model applicable to the molecular discovery process, and is conducive to supervising whether the reconstructed molecular graph output by the decoder is a candidate molecule with the required chemical properties through this reconstruction loss, making the molecular graph reconstruction model applicable to the molecular discovery process. In addition, designing the graph matching module in the molecular graph reconstruction model as a pre-training module can not only reduce the matching complexity and resource consumption compared with the traditional matching method, but also improve the accuracy of matching. Furthermore, it can improve the guiding effect of the graph matching module on the molecular graph reconstruction model.

[0079] In addition, the application designs the representation vector of the sample molecular graph output by the encoder as the input vector of the decoder, which can not only improve the practicability of the molecular graph reconstruction model, but also enable the molecular graph reconstruction model to generate reconstructed molecular graphs with different property constraints, improving the reconstruction effect of the molecular graph reconstruction model.

[0080] Specifically, on the one hand, designing the representation vector of the sample molecular graph output by the encoder as the input vector of the decoder can reconstruct the sample molecular graph to obtain the reconstructed molecular graph output by the decoder; on the other hand, in the process of molecular graph reconstruction, taking the preset attribute label as the input can obtain the reconstructed molecular graph with composite property constraints, improving the practicability of the molecular graph reconstruction model. On the other hand, the representation vector of the molecular sample graph output by the encoder as an input vector of the decoder can be used as a property constraint of the decoder, and the preset attribute label as another input of the decoder can be used as another property constraint sample of the decoder. Furthermore, it can realize the reconstructed molecular graph with multiple property constraints, improving the reconstruction effect of the molecular graph reconstruction model.

[0081] In summary, the training method of the molecular reconstruction graph model provided by this application can not only make the molecular graph reconstruction model applicable to the molecular discovery process, but also reduce the matching complexity and resource consumption, improve the accuracy of matching, and further improve the guiding effect of the graph matching module on the molecular graph reconstruction model. In addition, the training method of the molecular reconstruction graph model provided by this application can also improve the practicability and reconstruction effect of the molecular graph reconstruction model.

[0082] Next, a molecular graph reconstruction model applicable to this application will be described by way of example. Figure 2

[0083] Figure 2 FIG. 2 is a schematic framework of the molecular graph reconstruction model 200 provided by an embodiment of this application.

[0084] As Figure 2 shown, the molecular graph reconstruction model 200 may include an encoder 210, a decoder 220, and a graph matching module 230. Among them, the encoder 210 may be used to process the sample molecular graph into a representation vector of the sample molecular graph. Optionally, the encoder 210 may also be used to predict the attributes of the sample molecular graph to obtain the predicted attribute values of the sample molecular graph. The decoder 220 may be used to reconstruct the representation vector of the sample molecular graph output by the encoder into a molecular graph, that is, it may be used to output a reconstructed molecular graph. The graph matching module 230 may be used to supervise the difference between the sample molecular graph and the reconstructed molecular graph, that is, the reconstruction loss of the sample molecular graph.

[0085] In addition, the encoder 210 may be used to process the reconstructed molecular graph into a representation vector of the reconstructed molecular graph. Optionally, the encoder 210 may also be used to predict the attributes of the reconstructed molecular graph to obtain the predicted attribute values of the reconstructed molecular graph.

[0086] In some embodiments, before S140, the method 100 may further include:

[0087] Reorder the nodes and edges in the sample molecular graph using a reordering matrix to obtain a reordered molecular graph;

[0088] Use the graph matching module to predict the correspondence between the sample molecular graph and the reordered molecular graph to obtain a prediction matrix;

[0089] Adjust the parameters of the graph matching module based on the reordering matrix and the prediction matrix to obtain the trained graph matching module.

[0090] Exemplarily, sample nodes and sample edges may be randomly selected from the sample molecular graph to obtain a sample set corresponding to the sample molecular graph, which may be expressed as: where N represents the number of nodes included in the sample set, C x ​The vector dimension of the nodes in the sample set; N 2 Represents the number of edges included in the sample set. Denotes the n-dimensional Euclidean space, simply referred to as the n-dimensional space. After obtaining the sample set, the order of the nodes in the sample set can be randomly rearranged, and the edges included in the sample set can also be reordered in the same way to obtain the reordered sample set, that is, the sample set corresponding to the reordered molecular graph, which can be expressed as If the random rearrangement method is defined as the reordering matrix S target , then the node correspondence and edge correspondence between the sample molecular graph and the reordered molecular graph can be respectively expressed as:

[0091] Furthermore, obtain the sample set corresponding to the sample molecular graph and the sample set corresponding to the reordered molecular graph After that, the sample set corresponding to the sample molecular graph and the sample set corresponding to the reordered molecular graph can be respectively input into the deep graph neural network (Graph neural networks, GNN) to fuse the node features, edge features, and structural features in the sample set corresponding to the sample molecular graph , as well as fuse the node features, edge features, and structural features in the sample set corresponding to the reordered molecular graph , and then obtain the fused features: and Furthermore, the most similar nodes can be matched through the attention mechanism, that is and after normalizing it, the prediction matrix is obtained. For example, it can be normalized to S = sinkhorn(Att) through the Sinkhorn algorithm. Finally, based on the reordering matrix S target and the prediction matrix S, the graph matching module can be pre-trained to obtain the pre-trained graph matching module.

[0092] It should be noted that traditional graph matching algorithms achieve the correspondence between nodes and nodes, and edges and edges in two graphs by finding the high-order neighborhood features of nodes and edges. As the graph scale increases (the number of nodes), the computational cost of traditional graph matching algorithms will increase by the fourth power of the number of nodes, and the computational efficiency is low.

[0093] In this embodiment, the self-supervised training method can not only achieve the ability similar to traditional graph matching, but also reduce the matching complexity and improve the matching efficiency.

[0094] In addition, the present application trains the graph matching module by means of reordering, and can perform pre-training on a real dataset using samples with randomly reordered node orders without any additional annotation information, and can accurately find the node correspondence and edge correspondence between the input sample and the generated sample. Compared with the solution using traditional graph matching technology, the solution provided by the present application is more efficient and reduces a large amount of computational resource consumption. Compared with the solution using statistics, the matching degree of the solution provided by the present application is more accurate, which is beneficial to improving the guiding effect on the molecular graph reconstruction model.

[0095] Optionally, the reordering matrix is an orthogonal matrix, and the values of the elements in the orthogonal matrix include 0 and 1.

[0096] In other words, S target can be defined as an orthogonal matrix containing only [0, 1].

[0097] Optionally, calculate the first loss, the second loss, and the third loss; wherein, the first loss is used to characterize the difference between the reordering matrix and the prediction matrix; the second loss is used to characterize the difference between the first representation vector and the second representation vector of the first node in the sample molecular graph, the first representation vector is the representation vector of the first node before being transformed by the reordering matrix, and the second representation vector is the representation vector obtained by transforming the first representation vector by the reordering matrix and then by the prediction matrix; the third loss is used to characterize the difference between the third representation vector and the fourth representation vector of the first edge in the sample molecular graph, the third representation vector is the representation vector of the first edge before being transformed by the reordering matrix, and the fourth representation vector is the representation vector obtained by transforming the third representation vector by the reordering matrix and then by the prediction matrix; adjust the graph matching module based on the first loss, the second loss, and the third loss to obtain the trained graph matching module.

[0098] In this embodiment, the first loss, the second loss, and the third loss can be used as loss functions to supervise the learning of the graph matching module. In other words, the prediction effect of the graph matching module can be self-supervised through the first loss, the second loss, and the third loss. Furthermore, it can be ensured that the correspondence S predicted by the graph matching module and the set reordering relationship S target are as consistent as possible or tend to fit, improving the prediction effect of the graph matching module.

[0099] Exemplarily, the first loss can be defined as the constraint conditions of the node correspondence and edge correspondence between the sample molecular graph and the reordering molecular graph. For example, the first loss can be defined as: and / or Where I is the identity matrix, and M is used as a mask to shield the influence of empty nodes on the graph matching module. ||x|| represents the norm. For example, in a two-dimensional Euclidean geometric space, there is an Euclidean norm. Exemplarily, each vector is drawn as an arrow starting from the origin in the Cartesian coordinate system, and the Euclidean norm of each vector is the length of the arrow.

[0100] Exemplarily, the second loss can be defined as a constraint condition for the node correspondence between the sample molecular graph and the re-ranked molecular graph. For example, the second loss can be defined as:

[0101] Exemplarily, the third loss can be defined as a constraint condition for the edge correspondence between the sample molecular graph and the reconstructed molecular graph. For example, the third loss can be defined as:

[0102] In some embodiments, the S160 may include:

[0103] Calculating a first divergence and a second divergence;

[0104] Wherein, the first divergence is used to characterize the difference between the prior distribution and the distribution of the representation vectors of the sample molecular graph; the second divergence is used to characterize the difference between the prior distribution and the representation vectors of the reconstructed molecular graph;

[0105] Adjusting the parameters of the encoder based on the first divergence, and adjusting the parameters of the decoder based on the second divergence and the reconstruction loss to obtain a trained molecular graph reconstruction model.

[0106] It should be noted that, generally, a generative adversarial network is a way for two neural networks to learn by competing with each other, consisting of a generative network and a discriminator network. For example, a generative adversarial network can be implemented as a Graph Auto Encoder (GAN). The generative network needs to generate generated samples that are as similar as possible to the real samples. The input of the discriminator network is the real samples and the generated samples output by the generative network. The purpose is to distinguish the real samples as much as possible, while the generative network tries to deceive the discriminator network as much as possible, and finally makes the discriminator network unable to judge whether the output result of the generative network is real.

[0107] However, the training mode of the generative adversarial network is prone to training collapse, that is, the generative network can only generate a single sample.

[0108] In addition, variational autoencoders require the generated samples to be as close as possible to the input samples. However, the evaluation metrics for closeness do not necessarily ensure that the generated samples and the real samples have an approximate distribution in all spaces, which may lead to too low quality and accuracy of the generated samples output by the variational autoencoder and fail to meet the application requirements.

[0109] Based on this, the present application combines the adversarial training mode and the variational autoencoder in the training process of the encoder and the decoder. That is, the encoder is used as the discriminator network and the decoder is used as the generator network. Correspondingly, the input of the encoder is the sample molecular graph and the reconstructed molecular graph. On the one hand, by introducing the first divergence, it can be ensured that the sample molecular graph as the real sample conforms to the prior distribution as much as possible after being mapped to the vector space, that is, it can be ensured that the representation vector of the sample molecular graph can conform to the prior distribution as much as possible. On the other hand, by introducing the second divergence, it can be ensured that the reconstructed molecular graph as the generated sample conforms to the prior distribution as much as possible after being mapped to the vector space, that is, it can be ensured that the representation vector of the reconstructed molecular graph can conform to the prior distribution as much as possible. Thus, it can be ensured that the generated samples (reconstructed molecular graphs) and the real samples (i.e., sample molecular graphs) have an approximate distribution in all spaces, not only retaining the advantages of the generative adversarial network and the variational autoencoder but also overcoming their disadvantages. Specifically, not only the advantages of the generative adversarial network and the variational autoencoder are retained, but also the training mode of the molecular graph reconstruction model can be prevented from falling into training collapse, that is, the encoder can be applicable to generate various samples. In addition, it can make the generated samples and the real samples have an approximate distribution in all spaces. Furthermore, the quality and accuracy of the generated samples output by the encoder are improved, and the performance of the encoder is enhanced.

[0110] In other words, aiming at the problems of low accuracy of the generated samples by the variational autoencoder and the easy occurrence of mode collapse in the generative adversarial network, the present application uses the training ideas of the variational autoencoder and the generative adversarial network, and uses the proposed first divergence and second divergence to supervise the learning effect of the molecular graph reconstruction model. Specifically, on the basis that the encoder is used to process the sample molecular graph into a representation vector, it also undertakes the role of the discriminator network in the generative adversarial network. That is, by using the KL divergence between the representation vector of the reconstructed molecular graph output by the encoder and the prior distribution as the adversarial target index, the training ideas of the variational autoencoder and the generative adversarial network are unified. Not only the disadvantages of both are complemented, but also the disadvantages of both can be overcome. Specifically, not only the advantages of the generative adversarial network and the variational autoencoder are retained, but also the training mode of the molecular graph reconstruction model can be prevented from falling into training collapse, that is, the encoder can be applicable to generate various samples. In addition, it can make the generated samples and the real samples have an approximate distribution in all spaces. Furthermore, the quality and accuracy of the generated samples output by the encoder are improved, and the performance of the encoder is enhanced.

[0111] Briefly, by iteratively updating the encoder and the decoder, the present application enables the decoder to generate realistic samples while avoiding training collapse and maintaining the diversity of the generated samples.

[0112] Exemplarily, assume that the distribution of the real samples in the vector space conforms to the prior distribution, that is, after the encoder maps the real samples to the vector space, it conforms to the prior distribution. The difference between the distribution of the real samples mapped to the vector space by the encoder and the prior distribution is measured using the KL divergence. That is, the difference between the distribution of the real samples mapped to the vector space and the prior distribution can be measured using the first divergence, which can be expressed as KL real In addition, the difference between the distribution of the reconstructed molecular graph generated by the decoder after being input into the encoder and the prior distribution is also measured using the KL divergence. That is, the difference between the distribution of the reconstructed molecular graph generated by the decoder after being input into the encoder and the prior distribution can be measured using the second divergence, which can be expressed as KL fake Among them, the prior distribution is also called the prior distribution or the a priori distribution, which is a probability distribution and is opposite to the posterior distribution; the prior distribution is independent of the test results or random sampling, and the prior distribution reflects the distribution obtained based on other knowledge before the statistical test. The KL divergence can be used to measure the matching degree of two probability distributions. The greater the difference between the two distributions, the greater the KL divergence. The sample space is the set of all possible results of an experiment or random trial, and each possible result in the random trial is called a sample point.

[0113] Exemplarily, in adversarial training, the decoder tries to deceive the encoder as much as possible. Therefore, the goal of the decoder is to minimize the first divergence, that is, to minimize KL real .

[0114] Exemplarily, in adversarial training, the encoder can distinguish the real samples and the samples output by the decoder as much as possible. That is, the goal of the encoder is to maximize the second divergence, that is, to maximize KL fake .

[0115] Optionally, a first value is calculated based on the first divergence; wherein, when the first divergence is greater than or equal to a preset threshold, the gradient of the first divergence is less than the gradient of the first value; based on the first value, the parameters of the encoder are adjusted.

[0116] Exemplarily, the present application can use maximizing to replace the encoder's goal of maximizing KL fake , so as to reduce the gradient when KL fake is too large, and thus avoid the encoder from pushing the distribution of the generated samples too far from the real distribution, improving the stability of the training process.

[0117] In some embodiments, the S160 may include:

[0118] Use this encoder to predict the attribute values of the sample molecule to obtain the predicted attribute values of the sample molecular graph;

[0119] Use this encoder to predict the attribute values of the reconstructed molecule to obtain the predicted attribute values of the reconstructed molecular graph;

[0120] Calculate the fourth loss, the fifth loss, and the sixth loss;

[0121] Among them, the fourth loss is used to characterize the difference between the predicted attribute values of the sample molecular graph and the attribute labels of the sample molecular graph;

[0122] The fifth loss is used to characterize the difference between the predicted attribute values of the reconstructed molecular graph and the preset attribute labels;

[0123] The sixth loss is used to characterize the difference between the predicted attribute values of the sample molecular graph and the predicted attribute values of the reconstructed molecular graph;

[0124] Adjust the parameters of the encoder based on the fourth loss and the fifth loss, and adjust the parameters of the decoder based on the sixth loss and the reconstruction loss to obtain the trained molecular graph reconstruction model.

[0125] In this embodiment, the fourth loss, the fifth loss, and the sixth loss can be used as loss functions to supervise the learning of the molecular graph reconstruction model. In other words, through the fourth loss, the fifth loss, and the sixth loss, the reconstruction effect of the molecular graph reconstruction model can be self-supervised, and thus, the reconstruction effect of the molecular graph reconstruction model can be guaranteed.

[0126] In addition, for the prediction results of the encoder on the attribute values, as the results and supervision of the encoder's predicted properties, in some embodiments, the predicted attribute values of the sample molecular graph can be used as features in the representation vector of the sample molecular graph input to the decoder. That is, the predicted values of the encoder on the attribute values of the sample molecular graph can be used as the pre-set attribute values of the reconstructed molecular graph. Based on this, in practical applications, technicians only need to add an additional pre-set attribute label to the input of the decoder, and the decoder can generate a reconstructed molecular graph with composite set property constraints, realizing the function of a single model for generating reconstructed molecular graphs with different property constraints.

[0127] Of course, when the predicted attribute values of the sample molecular graph are used as features in the representation vector of the sample molecular graph input to the decoder, this application does not limit its specific implementation manner.

[0128] For example, the predicted attribute values of the sample molecular graph can be mapped in the vector space in the form of real numbers.

[0129] For another example, the predicted attribute value of the sample molecular graph can be in a way similar to that of a capsule network, replacing the description of the predicted attribute value from a real number with a vector, and then using the norm of the vector or the distance of the assumed distribution as the measure of the predicted attribute value. This can not only make the predicted attribute value have a more refined expression ability, but also give more adjustment space to the preset attribute labels. Furthermore, it can improve the operability of the molecular graph reconstruction model.

[0130] Exemplarily, the molecular graph reconstruction model supports the prediction of preset attribute labels and the generation guided by the preset attribute labels. For example, for the encoder in this application, it can perform encoding in combination with the prediction of the attribute value of the sample molecular graph. Specifically, the encoder can encode the predicted attribute value of the sample molecular graph to obtain the representation vector of the sample molecular graph. For example, the encoder can encode the predicted attribute value of the sample molecular graph in the vector space to obtain the representation vector of the sample molecular graph. For another example, for the decoder, it can take the encoding of the predicted attribute value of the sample molecular graph predicted by the encoder in the vector space as the input and reconstruct the sample molecular graph to obtain the reconstructed molecular graph. For yet another example, after receiving the reconstructed molecular graph output by the decoder, the encoder can also predict the attribute value of the reconstructed molecular graph to obtain the predicted attribute value of the reconstructed molecular graph.

[0131] Exemplarily, during the training process, the goal of the encoder is to minimize the difference between the predicted attribute value of the sample molecular graph and the attribute label of the sample molecular graph, that is, to minimize this fourth loss, while the goal of the decoder is to minimize the difference between the preset attribute label and the predicted attribute value of the reconstructed molecular graph, that is, to minimize this fifth loss. In addition, the goal of the encoder can also be to minimize the difference between the predicted attribute value of the sample molecular graph and the predicted attribute value of the reconstructed molecular graph, that is, to minimize this sixth loss.

[0132] In other words, during the training process, the goals of the encoder and the decoder can be as follows:

[0133] The predicted attribute value of the sample molecular graph predicted by the encoder is: c real ; the attribute label of the sample molecular graph is: l real ; the predicted attribute value of the reconstructed molecular graph predicted by the encoder is: C fake ; the preset attribute label for the reconstructed molecular graph is: l fake ; at this time, the goal of the encoder is to minimize the fourth loss and the sixth loss, that is, to minimize ||c real -l real || and ||c fake -l fake ||, while the goal of the decoder is to minimize the fifth loss, that is, to minimize ||c fake -c real ||.

[0134] In some embodiments, the property value is the value of the absorption property or the value of the Absorption, Distribution, Metabolism, Excretion, Toxicity (ADMET) property.

[0135] Exemplarily, the property value may be the physicochemical characteristics of the molecule, including but not limited to: solubility, permeability, stability, etc.; it may be biochemical characteristics, including but not limited to: metabolic process, protein binding ability, transport (absorption and excretion), etc.; it may also be toxicity characteristics, including but not limited to: clearance rate, half-life, biological activity, drug-drug interaction (DDI), etc. Exemplarily, the property value may be thermodynamic solubility and kinetic solubility; thermodynamic solubility is the solubility ability of a compound after reaching equilibrium in dissolution that we usually consider, and kinetic solubility generally refers to adding a compound dissolved in an organic solvent to an aqueous solution and then detecting the solubility. Of course, the above properties are only examples of the present application and should not be construed as a limitation to the present application.

[0136] It should be noted that the molecular graph Chinese module provided by the present application can be applied to various application scenarios, and the present application does not make specific limitations thereto. For example, it can be applied to the prediction of the property value of a molecular graph to be predicted, or it can be applied to optimize a molecular graph to be optimized, such as optimizing the molecular graph to be optimized based on a preset property label, and it can also be used to construct a molecular graph. For example, it is used to construct a molecular graph with a preset property label.

[0137] Figure 3 It is an application example of the decoder in the molecular graph reconstruction model provided by the embodiments of the present application.

[0138] As Figure 3 shown, by using the decoder 220 in the molecular graph reconstruction model, a random vector is obtained, and with this random vector and a preset property label as inputs, a molecular graph with the preset property label is constructed; wherein, the molecular graph reconstruction model is the molecular graph reconstruction model trained by the method 100 as shown in Figure 1 shown, or the molecular graph reconstruction model 200 as shown in Figure 2 shown.

[0139] In other words, the decoder in the molecular graph reconstruction model provided by the present application can be used to construct a molecular graph. For example, a molecular graph with characteristic properties is constructed.

[0140] Figure 4This is another application example of the decoder in the molecular graph reconstruction model provided by the embodiments of the present application.

[0141] As Figure 3 shown, the encoder 210 in the molecular graph reconstruction model is used to obtain the molecular graph to be optimized, and the molecular graph to be optimized is used as the input to process the molecular graph to be optimized, so as to obtain the characterization vector of the molecular graph to be optimized; wherein, the molecular graph reconstruction model is a molecular graph reconstruction model trained according to the method 100 shown in Figure 1 or the molecular graph reconstruction model 200 shown in Figure 2 ; using the characterization vector of the molecular graph to be optimized and the preset attribute label as the input, the decoder 220 in the molecular graph reconstruction model is used to construct a molecular graph obtained by optimizing the molecular graph to be optimized based on the preset attribute label.

[0142] In other words, on the basis of the existing molecular graph, the encoder in the molecular graph reconstruction model provided by the present application can be used to encode the vector space of the existing molecular graph, and then the preset attribute label is adjusted, so that the generated reconstructed molecular graph can optimize the property performance of the existing molecular graph while ensuring that the molecular configuration is substantially the same as that of the input existing molecular graph.

[0143] In addition, the predicted attribute value of the existing molecular graph can be used as a feature in the characterization vector of the sample molecular graph input to the decoder, that is, the predicted value of the attribute value of the existing molecular graph by the encoder can be used as the pre-set attribute value of the reconstructed molecular graph. Based on this, in practical applications, technicians only need to add an additional preset attribute label to the input of the decoder, and the decoder can generate a reconstructed molecular graph with composite set property constraints, realizing the function of a single model for generating reconstructed molecular graphs with different property constraints.

[0144] Figure 5 This is an application example of the encoder in the molecular graph reconstruction model provided by the embodiments of the present application.

[0145] As Figure 5 shown, the encoder 110 in the molecular graph reconstruction model is used to obtain the molecular graph to be predicted, and the molecular graph to be predicted is used as the input to predict the attributes of the molecular graph to be predicted, so as to obtain the attribute value of the molecular graph to be predicted; the molecular graph reconstruction model is a model trained according to the method in the first aspect. Wherein, the molecular graph reconstruction model is a molecular graph reconstruction model trained according to the method 100 shown in Figure 1 or the molecular graph reconstruction model 200 shown in Figure 2 .

[0146] In other words, the encoder in the molecular graph reconstruction model provided by the present application can be used to predict the attribute value of the molecular graph.

[0147] The preferred embodiments of the present application have been described in detail above in conjunction with the accompanying drawings. However, the present application is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present application, various simple modifications can be made to the technical solutions of the present application, and these simple modifications all fall within the protection scope of the present application. For example, for the various specific technical features described in the above specific embodiments, they can be combined in any appropriate manner without conflict. To avoid unnecessary repetition, the present application will not separately describe various possible combination methods. Again, for example, any combination can be made between various different embodiments of the present application, as long as it does not violate the idea of the present application, it should also be regarded as the content disclosed by the present application.

[0148] It should also be understood that in various method embodiments of the present application, the magnitudes of the serial numbers of the above processes do not mean the order of execution, and the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0149] The method provided by the embodiments of the present application has been described above. Next, the device provided by the embodiments of the present application will be described.

[0150] Figure 6 is a schematic block diagram of a training device 300 for a molecular graph reconstruction model provided by an embodiment of the present application.

[0151] As Figure 6 shown, the training device 300 for the molecular graph reconstruction model may include:

[0152] An acquisition unit 310, configured to acquire a sample molecular graph;

[0153] A processing unit 320, configured to process the attribute values of the sample molecular graph by using the encoder to obtain a characterization vector of the sample molecular graph;

[0154] A reconstruction unit 330, configured to reconstruct the sample molecular graph based on the characterization vector of the sample molecular graph by using the decoder to obtain a reconstructed molecular graph;

[0155] A prediction unit 340, configured to predict the node correspondence and edge correspondence between the sample molecular graph and the reconstructed molecular graph by using the graph matching module to obtain a relationship matrix;

[0156] A calculation unit 350, configured to compare the sample molecular graph and the reconstructed molecular graph based on the relationship matrix to obtain a reconstruction loss of the reconstructed molecular graph, and the reconstruction loss is used to characterize the difference between the sample molecular graph and the reconstructed molecular graph;

[0157] An adjustment unit 360, configured to adjust the decoder based on the reconstruction loss to obtain a trained molecular graph reconstruction model.

[0158] In some embodiments, before the prediction unit 340 uses the graph matching module to predict the node correspondence and edge correspondence between the sample molecular graph and the reconstructed molecular graph to obtain a relationship matrix, it can also be used to:

[0159] Use the reordering matrix to reorder the nodes and edges in the sample molecular graph to obtain a reordered molecular graph;

[0160] Use the graph matching module to predict the correspondence between the sample molecular graph and the reordered molecular graph to obtain a prediction matrix;

[0161] Adjust the parameters of the graph matching module based on the reordering matrix and the prediction matrix to obtain the trained graph matching module.

[0162] In some embodiments, the reordering matrix is an orthogonal matrix, and the values of the elements in the orthogonal matrix include 0 and 1.

[0163] In some embodiments, the prediction unit 340 may specifically be used to:

[0164] Calculate a first loss, a second loss, and a third loss;

[0165] Wherein, the first loss is used to characterize the difference between the reordering matrix and the prediction matrix;

[0166] The second loss is used to characterize the difference between the first representation vector and the second representation vector of the first node in the sample molecular graph. The first representation vector is the representation vector of the first node before being transformed by the reordering matrix, and the second representation vector is the representation vector obtained by transforming the first representation vector through the reordering matrix and then through the prediction matrix;

[0167] The third loss is used to characterize the difference between the third representation vector and the fourth representation vector of the first edge in the sample molecular graph. The third representation vector is the representation vector of the first edge before being transformed by the reordering matrix, and the fourth representation vector is the representation vector obtained by transforming the third representation vector through the reordering matrix and then through the prediction matrix;

[0168] Adjust the graph matching module based on the first loss, the second loss, and the third loss to obtain the trained graph matching module.

[0169] In some embodiments, the adjustment unit 360 is specifically used to:

[0170] Calculate a first divergence and a second divergence;

[0171] Wherein, the first divergence is used to characterize the difference between the prior distribution and the distribution of the representation vectors of the sample molecular graph; the second divergence is used to characterize the difference between the prior distribution and the representation vectors of the reconstructed molecular graph;

[0172] Based on the first divergence, the parameters of the encoder are adjusted, and based on the second divergence and the reconstruction loss, the parameters of the decoder are adjusted to obtain a trained molecular graph reconstruction model.

[0173] In some embodiments, the adjustment unit 360 is specifically configured to:

[0174] Calculate a first value based on the first divergence;

[0175] Wherein, when the first divergence is greater than or equal to a preset threshold, the gradient of the first divergence is less than the gradient of the first value;

[0176] Based on the first value, the parameters of the encoder are adjusted.

[0177] In some embodiments, the adjustment unit 360 is specifically configured to:

[0178] Use the encoder to predict the attribute values of the sample molecule to obtain the predicted attribute values of the sample molecular graph;

[0179] Use the encoder to predict the attribute values of the reconstructed molecule to obtain the predicted attribute values of the reconstructed molecular graph;

[0180] Calculate a fourth loss, a fifth loss, and a sixth loss;

[0181] Wherein, the fourth loss is used to characterize the difference between the predicted attribute values of the sample molecular graph and the attribute labels of the sample molecular graph;

[0182] The fifth loss is used to characterize the difference between the predicted attribute values of the reconstructed molecular graph and the preset attribute labels;

[0183] The sixth loss is used to characterize the difference between the predicted attribute values of the sample molecular graph and the predicted attribute values of the reconstructed molecular graph;

[0184] Based on the fourth loss and the fifth loss, the parameters of the encoder are adjusted, and based on the sixth loss and the reconstruction loss, the parameters of the decoder are adjusted to obtain a trained molecular graph reconstruction model.

[0185] In some embodiments, the attribute value is the value of the absorption attribute or the value of the distribution, metabolism, excretion, and toxicity (ADMET) attribute.

[0186] It should be understood that the device embodiments and the method embodiments can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, they will not be elaborated here. Specifically, the training device 300 of the molecular graph reconstruction model can correspond to the corresponding entity in the method 100 of the embodiments of the present application, and each unit in the training device 300 of the molecular graph reconstruction model respectively implements the corresponding processes in the method 100. For the sake of brevity, they will not be elaborated here.

[0187] It should also be understood that each unit in the training device 300 of the molecular graph reconstruction model involved in the embodiments of the present application can be respectively or wholly combined into one or several other units to form, or some of the units can be further split into multiple smaller units with functional division to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above units are divided based on logical functions. In practical applications, the function of one unit can be implemented by multiple units, or the functions of multiple units can be implemented by one unit. In other embodiments of the present application, the training device 300 of the molecular graph reconstruction model can also include other units. In practical applications, these functions can also be assisted by other units and can be implemented by the cooperation of multiple units. According to another embodiment of the present application, the training device 300 of the molecular graph reconstruction model involved in the embodiments of the present application can be constructed by running a computer program (including program code) capable of executing the steps involved in the corresponding method on a general computing device of a general computer including processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM), so as to respectively implement the training method of the molecular graph reconstruction model provided by the embodiments of the present application. The computer program can be recorded on a computer-readable storage medium, for example, and loaded into an electronic device through the computer-readable storage medium and run therein to implement the corresponding method of the embodiments of the present application.

[0188] In other words, the units described above can be implemented in the form of hardware, or in the form of software instructions, or in the form of a combination of software and hardware. Specifically, each step of the method embodiments in the embodiments of the present application can be completed by the integrated logic circuit in the hardware of the processor and / or software instructions. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software in the decoding processor. Optionally, the software can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, and a register. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the above method embodiments.

[0189] Figure 7It is a schematic structural diagram of the electronic device 400 provided by an embodiment of the present application.

[0190] As Figure 7 shown, the electronic device 400 at least includes a processor 410 and a computer-readable storage medium 420. Among them, the processor 410 and the computer-readable storage medium 420 can be connected through a bus or other means. The computer-readable storage medium 420 is used to store a computer program 421, and the computer program 421 includes computer instructions. The processor 410 is used to execute the computer instructions stored in the computer-readable storage medium 420. The processor 410 is the computing core and control core of the electronic device 400, and is adapted to implement one or more computer instructions, and is specifically adapted to load and execute one or more computer instructions to implement the corresponding method flow or corresponding function.

[0191] As an example, the processor 410 may also be referred to as a central processing unit (CPU). The processor 410 may include, but is not limited to: a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and so on.

[0192] As an example, the computer-readable storage medium 420 can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory; optionally, it can also be at least one computer-readable storage medium located far from the aforementioned processor 410. Specifically, the computer-readable storage medium 420 includes but is not limited to: volatile memory and / or non-volatile memory. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synch link dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).

[0193] As Figure 7 shown, the electronic device 400 may further include a transceiver 430.

[0194] Among them, the processor 410 can control the transceiver 430 to communicate with other devices. Specifically, it can send information or data to other devices, or receive information or data sent by other devices. The transceiver 430 can include a transmitter and a receiver. The transceiver 430 may further include an antenna, and the number of antennas can be one or more.

[0195] It should be understood that the various components in the communication device 400 are connected through a bus system. Among them, the bus system includes not only a data bus but also a power bus, a control bus, and a status signal bus.

[0196] In one implementation, the electronic device 400 can be any electronic device with data processing capabilities; a first computer instruction is stored in the computer-readable storage medium 420; the processor 410 loads and executes the first computer instruction stored in the computer-readable storage medium 420 to implement Figure 1 the corresponding steps in the method embodiments shown; in a specific implementation, the first computer instruction in the computer-readable storage medium 420 is loaded and executed by the processor 410 for the corresponding steps. To avoid repetition, details are not described herein again.

[0197] According to another aspect of the present application, an embodiment of the present application further provides a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in the electronic device 400 for storing programs and data. For example, the computer-readable storage medium 420. It can be understood that the computer-readable storage medium 420 here can include both the built-in storage medium in the electronic device 400 and, of course, the extended storage medium supported by the electronic device 400. The computer-readable storage medium provides a storage space that stores the operating system of the electronic device 400. And, one or more computer instructions suitable for being loaded and executed by the processor 410 are stored in this storage space. These computer instructions can be one or more computer programs 421 (including program codes).

[0198] According to another aspect of the present application, an embodiment of the present application further provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. For example, the computer program 421. At this time, the data processing device 400 can be a computer. The processor 410 reads the computer instructions from the computer-readable storage medium 420, and the processor 410 executes the computer instructions, so that the computer executes the model training method provided in the above various alternative manners.

[0199] In other words, when implemented using software, it can be implemented in the form of a computer program product in whole or in part. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes of the embodiments of the present application are run in whole or in part, or the functions of the embodiments of the present application are implemented. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website, a computer, a server, or a data center to another website, a computer, a server, or a data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.).

[0200] Those of ordinary skill in the art can realize that the units and process steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.

[0201] Finally, it should be noted that the above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A training method for a molecular graph reconstruction model, characterized in that, the molecular graph reconstruction model includes an encoder, a decoder, and a graph matching module; the method includes: Obtain a sample molecular graph; Use the encoder to process the attribute values of the sample molecular graph to obtain a representation vector of the sample molecular graph; Use the decoder to reconstruct the sample molecular graph based on the representation vector of the sample molecular graph to obtain a reconstructed molecular graph; Use the graph matching module to predict the node correspondence and edge correspondence between the sample molecular graph and the reconstructed molecular graph to obtain a relationship matrix; Based on the relationship matrix, compare the sample molecular graph and the reconstructed molecular graph to obtain a reconstruction loss of the reconstructed molecular graph, and the reconstruction loss is used to characterize the difference between the sample molecular graph and the reconstructed molecular graph; Based on the reconstruction loss, adjust the decoder to obtain a trained molecular graph reconstruction model; wherein, before using the graph matching module to predict the node correspondence and edge correspondence between the sample molecular graph and the reconstructed molecular graph to obtain a relationship matrix, the method further includes: Use a reordering matrix to reorder the nodes and edges in the sample molecular graph to obtain a reordered molecular graph; Use the graph matching module to predict the correspondence between the sample molecular graph and the reordered molecular graph to obtain a prediction matrix; Based on the reordering matrix and the prediction matrix, adjust the parameters of the graph matching module to obtain a trained graph matching module.

2. The method according to claim 1, characterized in that, the reordering matrix is an orthogonal matrix, and the values of the elements in the orthogonal matrix include 0 and 1.

3. The method according to claim 1, characterized in that, the adjusting the graph matching module based on the reordering matrix and the prediction matrix to obtain a trained graph matching module includes: Calculate a first loss, a second loss, and a third loss; wherein, the first loss is used to characterize the difference between the reordering matrix and the prediction matrix; The second loss is used to characterize the difference between the first representation vector and the second representation vector of the first node in the sample molecular graph. The first representation vector is the representation vector of the first node before being transformed by the reordering matrix, and the second representation vector is the representation vector obtained by transforming the first representation vector by the reordering matrix and then by the prediction matrix; The third loss is used to characterize the difference between the third representation vector and the fourth representation vector of the first edge in the sample molecular graph. The third representation vector is the representation vector of the first edge before being transformed by the reordering matrix, and the fourth representation vector is the representation vector obtained by transforming the third representation vector by the reordering matrix and then by the prediction matrix; Based on the first loss, the second loss, and the third loss, adjust the graph matching module to obtain a trained graph matching module.

4. The method according to any one of claims 1 to 3, It is characterized in that adjusting the decoder based on the reconstruction loss to obtain a trained molecular graph reconstruction model, including: calculating a first divergence and a second divergence; wherein, the first divergence is used to characterize the difference between the prior distribution and the distribution of the representation vectors of the sample molecular graphs; the second divergence is used to characterize the difference between the prior distribution and the representation vectors of the reconstructed molecular graphs; adjusting the parameters of the encoder based on the first divergence, and adjusting the parameters of the decoder based on the second divergence and the reconstruction loss to obtain a trained molecular graph reconstruction model.

5. The method according to claim 4, It is characterized in that the adjusting the parameters of the encoder based on the first divergence includes: calculating a first value based on the first divergence; wherein, when the first divergence is greater than or equal to a preset threshold, the gradient of the first divergence is less than the gradient of the first value; adjusting the parameters of the encoder based on the first value.

6. The method according to any one of claims 1 to 3, It is characterized in that adjusting the decoder based on the reconstruction loss to obtain a trained molecular graph reconstruction model, including: predicting the attribute values of the sample molecules by using the encoder to obtain the predicted attribute values of the sample molecular graphs; predicting the attribute values of the reconstructed molecules by using the encoder to obtain the predicted attribute values of the reconstructed molecular graphs; calculating a fourth loss, a fifth loss, and a sixth loss; wherein, the fourth loss is used to characterize the difference between the predicted attribute values of the sample molecular graphs and the attribute labels of the sample molecular graphs; the fifth loss is used to characterize the difference between the predicted attribute values of the reconstructed molecular graphs and the preset attribute labels; the sixth loss is used to characterize the difference between the predicted attribute values of the sample molecular graphs and the predicted attribute values of the reconstructed molecular graphs; adjusting the parameters of the encoder based on the fourth loss and the fifth loss, and adjusting the parameters of the decoder based on the sixth loss and the reconstruction loss to obtain a trained molecular graph reconstruction model.

7. The method according to claim 6, It is characterized in that the attribute value is the value of the absorption attribute or the value of the distribution metabolism excretion toxicity (ADMET) attribute.

8. A method for constructing a molecular graph, It is characterized in that including: obtaining a random vector; using the decoder in the molecular graph reconstruction model with the random vector and the preset attribute labels as inputs to construct a molecular graph with the preset attribute labels; wherein, the molecular graph reconstruction model is a model trained by the method according to any one of claims 1 to 7.

9. A method for constructing a molecular graph, It is characterized in that including: obtaining a molecular graph to be optimized; using the encoder in the molecular graph reconstruction model with the molecular graph to be optimized as an input to process the molecular graph to be optimized to obtain the representation vector of the molecular graph to be optimized; wherein, the molecular graph reconstruction model is a model trained by the method according to any one of claims 1 to 7; Taking the characterization vector of the molecular graph to be optimized and the preset attribute label as inputs, and using the decoder in the molecular graph reconstruction model, a molecular graph optimized based on the preset attribute label for the molecular graph to be optimized is constructed.

10. A method for predicting molecular graph attributes, characterized in that, it includes: Obtaining the molecular graph to be predicted; Taking the molecular graph to be predicted as an input, and using the encoder in the molecular graph reconstruction model to perform attribute prediction on the molecular graph to be predicted, obtaining the attribute value of the molecular graph to be predicted; wherein, the molecular graph reconstruction model is a model trained according to the method described in any one of claims 1 to 7.

11. A training device for a molecular graph reconstruction model, characterized in that, the molecular graph reconstruction model includes an encoder, a decoder and a graph matching module; the training device includes: An acquisition unit, configured to acquire a sample molecular graph; A processing unit, configured to process the attribute value of the sample molecular graph by using the encoder to obtain the characterization vector of the sample molecular graph; A reconstruction unit, configured to use the decoder to reconstruct the sample molecular graph based on the characterization vector of the sample molecular graph to obtain a reconstructed molecular graph; A prediction unit, configured to use the graph matching module to predict the node correspondence and edge correspondence between the sample molecular graph and the reconstructed molecular graph to obtain a relationship matrix; A calculation unit, configured to compare the sample molecular graph and the reconstructed molecular graph based on the relationship matrix to obtain the reconstruction loss of the reconstructed molecular graph, and the reconstruction loss is used to characterize the difference between the sample molecular graph and the reconstructed molecular graph; An adjustment unit, configured to adjust the decoder based on the reconstruction loss to obtain a trained molecular graph reconstruction model; wherein, before the prediction unit uses the graph matching module to predict the node correspondence and edge correspondence between the sample molecular graph and the reconstructed molecular graph to obtain a relationship matrix, it is further configured to: Use a reordering matrix to reorder the nodes and edges in the sample molecular graph to obtain a reordered molecular graph; Use the graph matching module to predict the correspondence between the sample molecular graph and the reordered molecular graph to obtain a prediction matrix; Adjust the parameters of the graph matching module based on the reordering matrix and the prediction matrix to obtain a trained graph matching module.

12. An electronic device, characterized in that, it includes: A processor, adapted to execute a computer program; A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by the processor, it implements the method described in any one of claims 1 to 7, the method described in claim 8, the method described in claim 9, or the method described in claim 10.

13. A computer-readable storage medium, characterized in that, it is used to store a computer program, and the computer program enables a computer to execute the method described in any one of claims 1 to 7, the method described in claim 8, the method described in claim 9, or the method described in claim 10.

14. A computer program product comprising computer programs / instructions, wherein, when the computer programs / instructions are executed by a processor, the methods described in any one of claims 1 to 7, the method described in claim 8, the method described in claim 9, or the method described in claim 10 are implemented.

Citation Information

Patent Citations

  • Drug molecule generation method based on regularization variation automatic encoder

    CN110970099A

  • Molecule generation method and device, electronic equipment and storage medium

    CN112086144A