Artificial intelligence auxiliary drug generation method and device based on molecular bond scaffold
Through an artificial intelligence-assisted drug generation method based on molecular bond scaffolds, the molecular bond scaffold is generated using a circular graph neural network and a graph neural network, and combined with a variational autoencoder and proxy model, the shortcomings in chemical space and attribute prediction in the existing technology are solved, and the efficient generation of multi-target drugs is achieved.
Patent Information
- Application Number
- CN202510356562.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
Existing molecular generation models conflict with accessible chemical space and chemical effectiveness control, fragment-based methods are limited by predefined fragment sets, while atom-based methods lack the accuracy and flexibility of attribute prediction, making it difficult to generate high-quality multi-target drugs.
Using an artificial intelligence-assisted drug generation method based on molecular bond scaffolds, molecular fragments are removed through a circular graph neural network, and bonded connections are generated using graph neural networks. Combined with variational autoencoder and proxy models, the chemical space of drug generation and molecular properties are expanded.
On the basis of ensuring chemical effectiveness and synthesis accessibility, the chemical space for drug generation has been expanded, the diversity and flexibility of the generated molecules have been improved, and the attribute prediction ability to adapt to new tasks has been adapted.
Smart Images

Figure CN120299561A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer-aided drug generation, and in particular to an artificial intelligence-assisted drug generation method and device based on a molecular bond scaffold. Background Art
[0002] Computer-aided drug design (CADD) has been widely used in all stages of drug discovery to accelerate development and reduce costs, which generally includes molecular modeling and computational chemistry. Recently, artificial intelligence (AI), especially deep learning (DL), has made significant progress in several CADD fields mainly through string-based and graph-based methods, such as structure modeling and drug screening. In the field of drug design (i.e., molecular generation), string-based methods, which adopt sequence-based molecular representations (e.g., "Simplified Molecular Input Line Entry Specification", SMILES), have been widely adopted together with deep language models. However, existing graph-based molecular generation methods have not been fully developed, and their general performance lags behind that of string-based methods. Two traditional molecular graph generation methods (atom-based and fragment-based) show conflicting trade-offs in terms of accessible chemical space and control of chemical validity or synthetic accessibility. The fragment-based method generates molecules fragment by fragment, ensuring high validity and synthetic accessibility, but novelty and diversity are limited by the predefined fragment set. In contrast, the atom-based method can generate molecules atom by atom, with high chemical space accessibility, but usually requires more effort to control chemical validity. Secondly, the accuracy of property prediction is often hindered by data scarcity, especially in terms of experimentally measured properties, which affects the quality of the generated molecules. For example, intermolecular affinity, molecular toxicity, etc. In addition, existing molecular generation models have low flexibility and are not easily adaptable to new tasks (i.e., new properties). Summary of the Invention
[0003] In view of this, the main purpose of the embodiments of the present invention is to provide an artificial intelligence-assisted drug generation method and device based on a molecular bond scaffold, in order to solve at least one of the problems in the prior art. The present invention can expand the accessible chemical space of drug generation.
[0004] To achieve the above object, on the one hand, an embodiment of the present invention provides an artificial intelligence-assisted drug generation method based on a molecular bond scaffold, the method comprising:
[0005] Obtaining a first bond scaffold of a first molecular graph;
[0006] Inputting the first molecular graph into an encoder, and removing molecular fragments of the first molecular graph according to the first bond scaffold to obtain a latent space vector;
[0007] Predicting a second bond scaffold through a decoder according to the latent space vector;
[0008] Predict the bonding connections between each of the second key scaffolds through a decoder;
[0009] Assemble each of the second key scaffolds according to the bonding connections to obtain an intermediate molecular graph;
[0010] Perform atomic prediction on the intermediate molecular graph through the decoder to obtain a target atomic type;
[0011] Obtain a target molecule according to the intermediate molecular graph and the target atomic type.
[0012] In some embodiments, obtaining the first key scaffold of the first molecular graph includes the following steps:
[0013] Fragment the first molecular graph to obtain molecular fragments;
[0014] Perform an atom removal operation on the atoms of the molecular fragments to obtain the first key scaffold.
[0015] In some embodiments, inputting the first molecular graph into an encoder, and removing the molecular fragments of the first molecular graph according to the first key scaffold to obtain a latent space vector includes the following steps:
[0016] Input the first molecular graph into the recurrent graph neural network of the encoder, and sequentially remove the molecular fragments to which each of the first key scaffolds belongs according to the breadth-first search algorithm to obtain the latent space vector.
[0017] In some embodiments, sequentially removing the molecular fragments to which each of the first key scaffolds belongs according to the breadth-first search algorithm includes the following steps:
[0018] Update the features of the nodes and edges of the first molecular graph according to the graph neural network algorithm to obtain a second molecular graph;
[0019] Remove the nodes and edges of the second molecular graph according to the breadth-first search algorithm to obtain a third molecular graph;
[0020] Use the third molecular graph as the updated first molecular graph, and input the updated first molecular graph into the next recurrent graph neural network, and return to the step of updating the features of the nodes and edges of the first molecular graph according to the graph neural network algorithm to obtain a second molecular graph until there are no nodes and edges in the first molecular graph.
[0021] In some embodiments, predicting the bonding connections between each of the second key scaffolds through a decoder includes the following steps:
[0022] Obtain the bonding probability between any two nodes of all the second key scaffolds;
[0023] When the bonding probability is greater than a first preset threshold, obtain the bonding connection.
[0024] In some embodiments, the method of obtaining the target atom type by performing atomic prediction on the intermediate molecular graph through the decoder includes the following steps:
[0025] According to the graph neural network algorithm, perform a feature update operation on the nodes of the intermediate molecular graph;
[0026] Perform an atomic decoding operation on the nodes of the intermediate molecular graph after feature update to obtain an intermediate vector;
[0027] Select the candidate element with the highest element probability in the intermediate vector as the target atom type;
[0028] Wherein, each component of the intermediate vector is the element probability corresponding to the predicted candidate element.
[0029] In some embodiments, a method for generating an artificial intelligence-assisted drug based on a molecular bond scaffold further includes the following steps: training an initial proxy model according to the latent space vector to obtain a target proxy model;
[0030] Wherein, the target proxy model is used to learn molecular properties and perform transfer learning.
[0031] To achieve the above object, on the other hand, an embodiment of the present invention proposes an artificial intelligence-assisted drug generation device based on a molecular bond scaffold, and the device includes:
[0032] A first module, configured to obtain a first key scaffold of a first molecular graph;
[0033] A second module, configured to input the first molecular graph into an encoder, and remove molecular fragments of the first molecular graph according to the first key scaffold to obtain a latent space vector;
[0034] A third module, configured to predict a second key scaffold according to the latent space vector through a decoder;
[0035] A fourth module, configured to predict a bonding connection between each of the second key scaffolds through a decoder;
[0036] A fifth module, configured to assemble each of the second key scaffolds according to the bonding connection to obtain an intermediate molecular graph;
[0037] A sixth module, configured to perform atomic prediction on the intermediate molecular graph through the decoder to obtain a target atom type;
[0038] A seventh module, configured to obtain a target molecule according to the intermediate molecular graph and the target atom type.
[0039] To achieve the above object, on the other hand, an embodiment of the present invention provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above-mentioned artificial intelligence-assisted drug generation method based on a molecular bond scaffold.
[0040] To achieve the above object, on the other hand, an embodiment of the present invention provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned artificial intelligence-assisted drug generation method based on a molecular bond scaffold.
[0041] To achieve the above object, on the other hand, an embodiment of the present invention provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of a computer device can read the computer instructions from the computer-readable storage medium, and when the processor executes the computer instructions, the computer device implements the above-mentioned artificial intelligence-assisted drug generation method based on a molecular bond scaffold.
[0042] The embodiments of the present invention at least include the following beneficial effects: The present invention provides an artificial intelligence-assisted drug generation method and device based on a molecular bond scaffold. The solution obtains a first bond scaffold of a first molecular graph; inputs the first molecular graph into an encoder, and according to the first bond scaffold, removes molecular fragments of the first molecular graph to obtain a latent space vector; predicts a second bond scaffold according to the latent space vector through a decoder; predicts the bonding connections between the second bond scaffolds through the decoder; assembles the second bond scaffolds according to the bonding connections to obtain an intermediate molecular graph; predicts atoms of the intermediate molecular graph through the decoder to obtain a target atom type; and obtains a target molecule according to the intermediate molecular graph and the target atom type, which can expand the accessible chemical space for drug generation. Description of the Drawings
[0043] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0044] Figure 1 It is a flowchart of the artificial intelligence-assisted drug generation method based on a molecular bond scaffold provided by an embodiment of the present invention;
[0045] Figure 2 It is a schematic diagram of the encoder in the drug generation model provided by an embodiment of the present invention;
[0046] Figure 3 It is a schematic diagram of the decoder in the drug generation model provided by an embodiment of the present invention;
[0047] Figure 4 It is a schematic diagram of the algorithm of the graph neural network module provided by an embodiment of the present invention;
[0048] Figure 5 It is a schematic diagram of the overall architecture of the drug generation model provided by an embodiment of the present invention;
[0049] Figure 6 It is a schematic diagram of the hardware structure of the electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0050] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the embodiments of the present invention. They are only examples of devices and methods consistent with some aspects of the embodiments of the present invention detailed in the appended claims.
[0051] It should be noted that although functional module division is performed in the system schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from the module division in the system or the order in the flowchart. The terms "first / S100" and "second / S200" in the specification, claims and the above-mentioned drawings may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "when" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0052] The terms "at least one", "a plurality", "each", "any one", etc. used in the present invention, at least one includes one, two or more than two, a plurality includes two or more than two, each refers to each of the corresponding plurality, and any one refers to any one of the plurality.
[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this invention belongs. The terms used herein are for the purpose of describing embodiments of the invention only and are not intended to limit the invention.
[0054] Before elaborating on the embodiments of the present invention in detail, some nouns and terms involved in the embodiments of the present invention are described first. The nouns and terms involved in the embodiments of the present invention are subject to the following explanations.
[0055] Machine Learning (ML) is a subfield of artificial intelligence that aims to train models through data, enabling computers to learn from experience and improve performance without explicit programming. Its core is to enable machines to automatically identify patterns from data and make predictions or decisions.
[0056] Deep Learning (DL) is a subfield of machine learning that uses multi-layer neural networks to simulate complex data patterns. It learns the features of data through multi-level abstractions and is widely applied in fields such as image recognition and natural language processing. The core of deep learning is the neural network, especially the Deep Neural Network (DNN), which is trained with a large amount of data and computing resources and can automatically extract features and make predictions.
[0057] Graph Neural Network (GNN) is a deep learning model specifically designed for processing graph-structured data. Graph-structured data consists of nodes (vertices) and edges, where nodes represent entities and edges represent the relationships between entities. GNN can capture complex relationships in the graph structure by passing and aggregating information among the nodes and edges in the graph. Its advantage lies in processing non-Euclidean data applicable to complex structures such as graphs and trees, and it can obtain global and local features through multi-level information passing.
[0058] Recurrent Neural Network (RNN) is a neural network used for processing sequential data. Its core feature is the ability to utilize the hidden state at the previous moment to influence the output at the current moment, thereby capturing the temporal dependencies in the sequence.
[0059] Variational Autoencoder (VAE) is a generative model that combines probabilistic graphical models and deep learning, mainly used to generate new samples similar to the training data.
[0060] Computer-Aided Drug Design (CADD) is a process that uses computer technology to accelerate drug discovery and optimization. It combines chemistry, biology, physics, and computer science to design new drugs by simulating and predicting molecular behavior.
[0061] Molecular generation refers to the use of computational methods to automatically design new molecular structures, usually by combining chemical rules, machine learning, or deep learning models to generate molecules that meet specific requirements.
[0062] The bond scaffold is a concept proposed in the method of the present invention and is defined as a molecular fragment that contains only chemical bonds.
[0063] Small molecule drugs play a crucial role in all aspects of public health care. The typical drug development process involves evaluating a large number of pharmacologically relevant molecules, and the estimated number of potential molecules is as high as 10 60 . In addition, in the real world, the designed molecules usually need to meet two or more properties of interest, such as binding to two different proteins simultaneously (dual-target drugs), or achieving both a high drug-likeness score and low toxicity at the same time. These multi-target drugs can exhibit more favorable pharmacological effects. For example, dual-target drugs can utilize synthetic lethality to treat cancer with high specificity. However, for decades, many multi-target drugs have been discovered accidentally rather than through rational design, which makes multi-target drug design an open challenge in the field of drug development.
[0064] Currently, Computer-Aided Drug Design (CADD) has been widely used in all stages of drug discovery to accelerate development and reduce costs, which typically includes molecular modeling and computational chemistry. Recently, artificial intelligence (AI), especially deep learning (DL), has made significant progress in several CADD fields mainly through string-based and graph-based methods, such as structure modeling and drug screening. In the field of drug design (i.e., molecular generation), string-based methods, which use sequence-based molecular representations (such as "Simplified Molecular Input Line Entry Specification", SMILES), have been widely adopted together with deep language models. However, the existing graph-based molecular generation methods have not been fully developed, and their general performance lags behind that of string-based methods.
[0065] Two existing traditional molecular graph generation methods (atom-based and fragment-based) exhibit conflicting trade-offs in terms of accessible chemical space and control of chemical validity or synthetic accessibility. The fragment-based method generates molecules fragment by fragment, ensuring high validity and synthetic accessibility, but novelty and diversity are limited by the predefined fragment set. In contrast, the atom-based method can generate molecules atom by atom, with high chemical space accessibility, but usually requires more effort to control chemical validity. Secondly, the accuracy of property prediction is often hindered by data scarcity, especially for experimentally measured properties, which affects the quality of the generated molecules. For example, intermolecular affinity, molecular toxicity, etc. In addition, existing molecular generation models have low flexibility and are not easily adaptable to new tasks (i.e., new properties).
[0066] In view of this, as Figure 1 shown, embodiments of the present invention provide an artificial intelligence-assisted drug generation method based on a molecular bond scaffold, which may include but is not limited to steps S100 to S700:
[0067] Step S100, obtaining a first bond scaffold of a first molecular graph;
[0068] Step S200, inputting the first molecular graph into an encoder, and removing molecular fragments of the first molecular graph according to the first bond scaffold to obtain a latent space vector;
[0069] Step S300, predicting a second bond scaffold through a decoder according to the latent space vector;
[0070] Step S400, predicting the bonding connections between the second bond scaffolds through the decoder;
[0071] Step S500, assembling the second bond scaffolds according to the bonding connections to obtain an intermediate molecular graph;
[0072] Step S600, predicting target atom types for the intermediate molecular graph through the decoder;
[0073] Step S700, obtaining a target molecule according to the intermediate molecular graph and the target atom types.
[0074] In steps S100 to S200 of some embodiments, as Figure 2 shown, the encoder adopts a recurrent graph neural network (RGNN) as the core module. The first molecular graph is input into the encoder, and a part of the molecular fragments of the first molecular graph are removed in the RGNN, and then the output processed molecular graph is input into the next RGNN until there are no remaining molecular fragments. Finally, the recurrent graph neural network outputs a low-dimensional latent space vector of the molecule.
[0075] In some embodiments, step S100 may include, but is not limited to, steps S101 to S102:
[0076] Step S101, fragment the first molecular graph to obtain molecular fragments;
[0077] Step S102, perform an atom removal operation on the atoms of the molecular fragments to obtain the first bond scaffold.
[0078] In steps S101 to S102 of some embodiments, the first molecular graph is fragmented into molecular fragments by classical computational chemistry methods. Optionally, the computational chemistry method may be RECAP (REtrosynthetic Combinatorial Analysis Procedure) and BRICS (Breaking of Retrosynthetically Interesting Chemical Substructures). After obtaining the molecular fragments, removing the atoms of all fragments and only retaining the chemical bonds can obtain the first bond scaffold of the molecule.
[0079] In step S200 of some embodiments, the encoder will iteratively input molecules, and after each input, remove the fragment to which a bond scaffold of the molecule belongs. Exemplarily, input the first molecular graph into the encoder, and according to the obtained first bond scaffold, sequentially remove the molecular fragments to which a first bond scaffold of the first molecular graph belongs.
[0080] Optionally, the removal order is sorted by breadth-first search (BFS). Among them, randomly select a fragment as the starting node for BFS (breadth-first search) sorting. Then, by inputting the first molecular graph into the recurrent graph neural network of the encoder, according to the breadth-first search algorithm, sequentially remove the molecular fragments to which each first bond scaffold belongs to obtain the latent space vector.
[0081] In some embodiments, in the step of sequentially removing the molecular fragments to which each first bond scaffold belongs according to the breadth-first search algorithm, it may include, but is not limited to, steps S201 to S203:
[0082] Step S201, update the features of the nodes and edges of the first molecular graph according to the graph neural network algorithm to obtain a second molecular graph;
[0083] Step S202, remove the nodes and edges of the second molecular graph according to the breadth-first search algorithm to obtain a third molecular graph;
[0084] Step S203: Use the third molecular graph as the updated first molecular graph, and input the updated first molecular graph into the next recurrent graph neural network. Return the step of updating the features of the nodes and edges of the first molecular graph according to the graph neural network algorithm to obtain a second molecular graph, until there are no nodes and edges in the first molecular graph.
[0085] In steps S201 to S203 of some embodiments, first, the features of the nodes and edges of the first molecular graph are updated by applying the graph neural network (GNN) algorithm, and then a second molecular graph is obtained after the feature update. The second molecular graph retains the structure of the first molecular graph (i.e., the connection relationship between nodes and edges), but the features of the nodes and edges have been updated to better represent the topological structure and properties of the graph. Among them, the GNN algorithm uses the structural information of the graph and the initial features of the nodes and edges, and updates the feature representation of each node and edge through an information transmission and aggregation mechanism. Then, according to the breadth-first search (BFS) algorithm, the nodes and edges of the second molecular graph are removed to obtain a third molecular graph, and this third molecular graph is used as the new first molecular graph and input into the next RGNN module, and return to step S201 until there are no nodes and edges in the first molecular graph. Exemplarily, given an initial molecular graph, which is represented by a set of nodes and edges as G = {x i , e ij}, where x i represents the i-th node and its features; e ij represents the edge connecting the i-th node and the j-th node and its features. RGNN first updates the features of the nodes and edges of the initial molecular graph according to the GNN algorithm to obtain a new graph, denoted as G″ = {x i ″, e ij ″}. Then, according to the BFS sorting, the corresponding nodes and edges in G″ are removed, and then input into the RGNN module to repeat the above steps until there are no nodes and edges in the molecular graph. Optionally, global average pooling of all nodes and edges in G″ in RGNN can obtain a vector, which can be used to update the hidden layer state of the gated recurrent unit in RGNN.
[0086] In some embodiments, the recurrent graph neural network block includes an attention-based GNN and a gated recurrent unit (GRU). It uses a memory mechanism to encode and decode molecules in a sequential manner. The RGNN block aims to learn the encoding and decoding process (i.e., node deletion or generation) of the molecular graph in a sequential manner. First, it updates the features of the nodes and edges through the GNN block, and then uses global average pooling of the graph to update the hidden state of the GRU. The GRU uses a gating mechanism to selectively remember the input information. By introducing a gating mechanism to control the flow of information, the problems of gradient disappearance and explosion can be solved.
[0087] In steps S300 to S700 of some embodiments, as Figure 3 shown, starting from the latent space vector, the decoder uses the RGNN module to predict a new bond scaffold each time. This scaffold does not contain atomic information but only contains the molecular bond information of the fragment. After the bond scaffolds are generated, the decoder predicts the bonding connections between the bond scaffolds and assembles the bond scaffolds through the bonding connections. Then, atoms are predicted and generated iteratively multiple times until a complete molecule is obtained.
[0088] In step S300 of some embodiments, the latent space vector is input into the decoder, and the bond scaffolds of the molecule are decoded sequentially through the RGNN, and then each second bond scaffold can be obtained. During training, the decoding order is the reverse of the input during encoding. Each bond scaffold is a part of the molecular subgraph. Among them, the subgraph generated in the i-th step can be expressed as
[0089] In some embodiments, step S400 may include but is not limited to steps S401 to S402:
[0090] Step S401, obtaining the bonding probability between any two nodes of all the second bond scaffolds;
[0091] Step S402, when the bonding probability is greater than the first preset threshold, obtaining the bonding connection.
[0092] In steps S401 to S402 of some embodiments, for the nodes of all the second bond scaffolds, the bonding probability between any two nodes is obtained to predict whether they are bonded. When the bonding probability is greater than the first preset threshold, the two nodes are bonded. Optionally, a shallow neural network, namely a multi-layer perceptron (MLP), is used to predict bonding. Exemplarily, each bond scaffold is a part of the molecular subgraph, and the subgraph generated in the i-th step can be expressed as For all the generated subgraphs of the nodes, pairwise prediction is made on whether they are bonded, and there is the following formula:
[0093]
[0094] where Concat represents the concatenation between vectors; x i represents the i-th node and its features; x j represents the j-th node and its features; represents the bonding probability between the i-th node and the j-th node. When actually predicting bonding, if the bonding probability is greater than 0.5, the two nodes are bonded.
[0095] In step S500 of some embodiments, after predicting the completion of bonding, according to the bonding connection, each decoded second key bracket is assembled, and then an intermediate molecular graph G containing only molecular bonds can be obtained. scaf 。
[0096] In some embodiments, step S600 may include but is not limited to steps S601 to S603:
[0097] Step S601, perform a feature update operation on the nodes of the intermediate molecular graph according to the graph neural network algorithm;
[0098] Step S602, perform an atomic decoding operation on the nodes of the intermediate molecular graph after feature update to obtain an intermediate vector;
[0099] Step S603, select the candidate element with the highest element probability in the intermediate vector as the target atomic type;
[0100] Wherein, each component of the intermediate vector is the element probability corresponding to the predicted candidate element.
[0101] In steps S601 to S603 of some embodiments, starting from the intermediate molecular graph G scaf and using the GNN algorithm to update the features of each node of the intermediate molecular graph G scaf to obtain each updated node, then the updated i-th node is represented as x i ″. Performing atomic decoding on this node, there is the following expression:
[0102]
[0103] Wherein, represents the predicted atomic classification of the i-th node, which is an intermediate vector of length 10. Each position (i.e., each component) of this intermediate vector is the element probability of the predicted candidate element. When actually predicting an atom, the candidate element with the highest element probability is taken as the finally predicted atomic type. Optionally, the intermediate vector has a length of 10. Then when actually predicting an atom, among the 10 candidate elements, the candidate element with the highest predicted element probability is taken as the finally predicted atomic type. Exemplarily, taking the intermediate vector length as 3 as an example, assuming there is an intermediate vector which represents that when actually predicting an atom, the probabilities of occurrence in 3 candidate elements (such as hydrogen, oxygen, and carbon) are 20%, 30%, and 50% respectively. Then the carbon element with an element probability of 50% is taken as the finally predicted atomic type.
[0104] In step S700 of some embodiments, starting from the intermediate molecular graph G scafStart, predict one atom each time, and then update the intermediate molecular graph G according to the predicted atom type scaf and input the updated intermediate molecular graph G scaf into the GNN to predict the next atom. The nodes for predicting atoms are random each time until all nodes have predicted atoms, and finally the target molecule is obtained.
[0105] In some embodiments, in the GNN module algorithm, the GNN block performs message passing between nodes and edges, where nodes and edges update their features by aggregating attention-weighted messages from their neighboring nodes and edges (the attention mechanism provides a weight to determine the relative "importance" of each message). Exemplarily, input a graph (a set of nodes and edges). For the i-th node x i , its adjacent nodes are x j , and the edge between the two nodes is e ij . The GNN updates the node x Figure 4 and the edge e i using the algorithm shown in ij . Here, Linear represents a linear transformation, LinearNoBias represents a linear transformation without bias, LayerNorm represents layer normalization, softmax represents the normalized exponential function, and leaky_relu represents the leaky rectified linear unit.
[0106] Refer to Figure 4 , Figure 4 for the explanation of the GNN module algorithm as follows:
[0107] 1. Input and Output
[0108] Input: Central node Neighboring nodes where j belongs to the neighbor set of node i, and the edge between node i and j is
[0109] Output: Updated central node Updated edge
[0110] 2. Algorithm Steps
[0111] 2.1 Define the GNN function:
[0112] The input parameters include the central node x i , the neighboring nodes x j , the edge e ij , and the number of heads N h = 8.
[0113] 2.2 Message Aggregation of the Attention Mechanism:
[0114] Perform layer normalization and linear transformation (without bias) on the central node x i to obtain a query vector
[0115] Perform layer normalization and linear transformation (without bias) on the neighbor node x j to obtain a key vector and a value vector
[0116] Perform layer normalization and linear transformation (without bias) on the edge e ij to obtain an edge vector
[0117] Calculate the attention coefficient by adding the edge vector to the dot product of the query vector and the key vector;
[0118] Normalize the attention coefficient using the softmax function to obtain
[0119] 2.3 Aggregate neighbor node information:
[0120] Calculate the aggregated neighbor node information and the aggregated edge information
[0121] 2.4 Feed-forward network:
[0122] Perform linear transformation, layer normalization, and activation function (leaky ReLU) operations on the aggregated central node information to obtain the updated central node
[0123] Perform linear transformation, layer normalization, and activation function (leaky ReLU) operations on the aggregated edge information to obtain the updated edge
[0124] 2.5 Return the results:
[0125] Return the updated central node and the edge
[0126] In some embodiments, an artificial intelligence-assisted drug generation method based on a molecular bond scaffold may further include: training an initial surrogate model according to the latent space vector to obtain a target surrogate model; wherein, the target surrogate model is used to learn molecular properties and perform transfer learning.
[0127] As Figure 5As shown, an encoder, a decoder, and a surrogate model are included in the overall framework of the molecular generation model of the drug. Among them, the encoder embeds the molecular graph into a low-dimensional latent space, and the decoder decodes the molecule from the latent space. Additionally, by inputting the embedding of the molecule in the latent space, the surrogate model can be trained to predict different properties of the molecule, which can be computer-simulated properties or actual in vivo / in vitro biological data. After training, new molecules that meet the desired properties can be sampled in the latent space through the surrogate model. Optionally, the surrogate model adopts lightweight classical ML algorithms, such as random forest or support vector machine, etc. Exemplarily, in the training of the surrogate model, publicly available biochemical datasets, such as BindingDB, PDBbind, TDC databases, etc., can be used as the training dataset. These databases cover the absorption, distribution, metabolism, excretion, and toxicity of drugs, etc., and can be used to train surrogate models for predicting different properties. Among them, the surrogate model uses a classical ML algorithm framework, such as random forest or support vector machine. The training of ML is very fast compared to DL. Usually, there is no need to repeatedly adjust the parameters, and the sklearn library and its default training parameters can be used. By adopting an additional lightweight surrogate model for each molecular property to learn each molecular property based on the latent space vector encoded by the variational autoencoder (VAE), it can be avoided that the main body of the model needs to be retrained when applied to a new task (i.e., learning new molecular properties).
[0128] Further, in some alternative embodiments, the encoder and the decoder are trained. The training dataset can use publicly available large and small molecule datasets, such as the ChEMBL or ZINC databases. The model was trained using the Adam optimizer. The maximum node size was set to 64, and the initial learning rate was 5e -4 Subsequently, for additional training, the maximum number of nodes was increased to 128, and the learning rate was 1e -4 50,000 randomly sampled molecules were used as the test set, and the rest were used for training. All key scaffolds were guaranteed to appear at least once in the training set.
[0129] An artificial intelligence-assisted drug generation device based on molecular key scaffolds provided by an embodiment of the present invention can implement the above-mentioned artificial intelligence-assisted drug generation method based on molecular key scaffolds. The device includes:
[0130] A first module for obtaining the first key scaffold of the first molecular graph;
[0131] A second module for inputting the first molecular graph into the encoder, removing the molecular fragments of the first molecular graph according to the first key scaffold, and obtaining a latent space vector;
[0132] A third module for predicting a second key scaffold through the decoder according to the latent space vector
[0133] A fourth module, configured to predict the bonding connections between the second key brackets through a decoder;
[0134] A fifth module, configured to assemble the second key brackets according to the bonding connections to obtain an intermediate molecular graph;
[0135] A sixth module, configured to perform atomic prediction on the intermediate molecular graph through the decoder to obtain a target atom type;
[0136] A seventh module, configured to obtain a target molecule according to the intermediate molecular graph and the target atom type.
[0137] It can be understood that the content in the above method embodiments is applicable to the device embodiments. The functions specifically implemented in the device embodiments are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0138] An embodiment of the present invention further provides an electronic device, which includes a processor and a memory. The memory stores a computer program, and when the processor executes the computer program, the above-mentioned artificial intelligence-assisted drug generation method based on molecular key brackets is implemented. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.
[0139] It can be understood that the content in the above method embodiments is applicable to the device embodiments. The functions specifically implemented in the device embodiments are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0140] Reference Figure 6 , Figure 6 illustrates the hardware structure of an electronic device according to another embodiment. The electronic device includes:
[0141] A processor 801, which can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is configured to execute relevant programs to implement the technical solutions provided by the embodiments of the present invention;
[0142] The memory 802 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 802 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 802 and are called by the processor 801 to execute the artificial intelligence-assisted drug generation method based on the molecular bond scaffold of the embodiments of the present invention;
[0143] The input / output interface 803 is used to implement information input and output;
[0144] The communication interface 804 is used to implement communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0145] The bus 805 transmits information between various components of the device (such as the processor 801, the memory 802, the input / output interface 803, and the communication interface 804);
[0146] Among them, the processor 801, the memory 802, the input / output interface 803, and the communication interface 804 are communicatively connected to each other inside the device through the bus 805.
[0147] The embodiments of the present invention also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the artificial intelligence-assisted drug generation method based on the molecular bond scaffold described above is implemented.
[0148] It can be understood that the content in the above method embodiments is applicable to the embodiments of this storage medium. The functions specifically implemented by the embodiments of this storage medium are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0149] The embodiments of the present invention also provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions. The computer instructions are stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the foregoing artificial intelligence-assisted drug generation method based on the molecular bond scaffold.
[0150] In summary, the artificial intelligence-assisted drug generation method and device based on the molecular bond scaffold in the embodiments of the present invention have the following advantages:
[0151] 1. The embodiment of the present invention adopts a drug molecule generation path based on a bond scaffold. First, a cyclic neural network is used to generate a set of molecular bond scaffolds, and then a graph neural network is used to iteratively generate each atom in turn. On the basis of ensuring high effectiveness and synthetic accessibility, the accessible chemical space can be greatly expanded. The embodiment of the present invention covers a wider chemical space than the fragment-based molecular graph generation method. The fragment-based method can only construct new molecules based on the existing fragment library, and the common fragment library is very small compared to the known chemical space. And the embodiment of the present invention has higher chemical effectiveness than the atom-based molecular graph generation method. The atom-based method usually requires additional operations to control the chemical effectiveness of the generated molecules through a predefined knowledge base. Usually, it is reflected that the molecules generated by the AI method applying the atom-based molecular graph are generally simple and the effectiveness is not as high as that of the fragment-based method.
[0152] 2. The embodiment of the present invention also uses a variational autoencoder (VAE) as the backbone. For each molecular property, an additional lightweight proxy model (ML module) is used to learn each molecular property according to the latent space vector encoded by the VAE, avoiding the need to retrain the entire model when the main body of the model is applied to a new task (i.e., learning a new molecular property), and improving the flexibility of the drug generation model.
[0153] In some alternative embodiments, the functions / operations mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the functions / operations involved, two consecutive blocks shown may actually be executed substantially simultaneously or the blocks can sometimes be executed in the reverse order. In addition, the embodiments presented and described in the flowcharts of the present invention are provided by way of example for a more comprehensive understanding of the technology. The disclosed method is not limited to the operations and logical processes presented herein. Alternative embodiments are contemplated, in which the order of various operations is changed and the sub-operations described as part of a larger operation are executed independently.
[0154] In addition, although the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the described functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present invention. Rather, given the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skills of an engineer. Thus, those skilled in the art can implement the present invention as set forth in the claims without undue experimentation. It should also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.
[0155] If the described function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0156] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definable sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0157] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or otherwise processing it as appropriate, and then storing it in a computer memory.
[0158] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0159] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0160] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the claims and their equivalents.
[0161] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present invention.
Claims
1. An artificial intelligence-assisted drug generation method based on a molecular bond scaffold, characterized in that, Including the following steps: Obtain the first bond scaffold of the first molecular graph; Input the first molecular graph into an encoder, and according to the first bond scaffold, remove the molecular fragments of the first molecular graph to obtain a latent space vector; Predict a second bond scaffold through a decoder according to the latent space vector; Predict the bond connections between the second bond scaffolds through the decoder; Assemble the second bond scaffolds according to the bond connections to obtain an intermediate molecular graph; Perform atomic prediction on the intermediate molecular graph through the decoder to obtain a target atomic type; Obtain a target molecule according to the intermediate molecular graph and the target atomic type.
2. The artificial intelligence-assisted drug generation method based on a molecular bond scaffold according to claim 1, wherein The step of obtaining the first bond scaffold of the first molecular graph includes the following steps: Fragment the first molecular graph to obtain molecular fragments; Perform an atom removal operation on the atoms of the molecular fragments to obtain the first bond scaffold.
3. The artificial intelligence-assisted drug generation method based on a molecular bond scaffold according to claim 1, characterized in that, The step of inputting the first molecular graph into an encoder, and according to the first bond scaffold, removing the molecular fragments of the first molecular graph to obtain a latent space vector includes the following steps: Input the first molecular graph into the recurrent graph neural network of the encoder, and according to the breadth-first search algorithm, sequentially remove the molecular fragments to which each first bond scaffold belongs to obtain the latent space vector.
4. An artificial intelligence-assisted drug generation method based on a molecular bond scaffold according to claim 3, characterized in that, The step of sequentially removing the molecular fragments to which each first bond scaffold belongs according to the breadth-first search algorithm includes the following steps: Update the features of the nodes and edges of the first molecular graph according to the graph neural network algorithm to obtain a second molecular graph; Remove the nodes and edges of the second molecular graph according to the breadth-first search algorithm to obtain a third molecular graph; Take the third molecular graph as the updated first molecular graph, and input the updated first molecular graph into the next recurrent graph neural network, and return to the step of updating the features of the nodes and edges of the first molecular graph according to the graph neural network algorithm to obtain a second molecular graph until there are no nodes and edges in the first molecular graph.
5. The artificial intelligence-assisted drug generation method based on a molecular bond scaffold according to claim 1, characterized in that, The step of predicting the bond connections between the second bond scaffolds through the decoder includes the following steps: Obtain the bonding probabilities between any two nodes of all the second bond scaffolds; When the bonding probability is greater than a first preset threshold, obtain the bond connection.
6. The artificial intelligence-assisted drug generation method based on a molecular bond scaffold according to claim 1, wherein The step of performing atomic prediction on the intermediate molecular graph through the decoder to obtain a target atomic type includes the following steps: Update the features of the nodes of the intermediate molecular graph according to the graph neural network algorithm; Perform an atomic decoding operation on the nodes of the intermediate molecular graph after feature update to obtain an intermediate vector; Select the candidate element with the highest element probability in the intermediate vector as the target atomic type; Wherein, each component of the intermediate vector is the element probability corresponding to the predicted candidate element.
7. A method for generating artificial intelligence-assisted drugs based on a molecular bond scaffold according to claim 1, characterized in that, It further includes the following steps: Train an initial proxy model according to the latent space vector to obtain a target proxy model; Wherein, the target proxy model is used to learn molecular properties and perform transfer learning.
8. An artificial intelligence-assisted drug generation device based on a molecular bond scaffold, characterized in that, Including: A first module for obtaining the first bond scaffold of the first molecular graph; A second module, configured to input the first molecular graph into an encoder, remove molecular fragments of the first molecular graph according to the first bond scaffold, and obtain a latent space vector; A third module, configured to predict a second bond scaffold through a decoder according to the latent space vector; A fourth module, configured to predict bond connections between the second bond scaffolds through the decoder; A fifth module, configured to assemble the second bond scaffolds according to the bond connections to obtain an intermediate molecular graph; A sixth module, configured to perform atomic prediction on the intermediate molecular graph through the decoder to obtain a target atomic type; A seventh module, configured to obtain a target molecule according to the intermediate molecular graph and the target atomic type; 9. An electronic device, characterized in that, Comprising a processor and a memory; The memory is used for storing a program; The processor executes the program to implement the method according to any one of claims 1 to 7; 10. A computer-readable storage medium, characterized in that, The storage medium stores a program, and the program is executed by a processor to implement the method according to any one of claims 1 to 7.