Drug molecule multi-attribute optimization method and system based on diffusion model
By constructing a global corpus and motif graph transformation and adjusting the molecular structure with the diffusion model, the problem of insufficient functional group structure capture and multi-attribute optimization in drug molecule optimization in the existing technology is solved, and the structural diversity and optimization effect of drug molecules in the multi-attribute optimization process is achieved.
Patent Information
- Application Number
- CN202510319713.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art is difficult to effectively capture the complex functional group structure in drug molecular optimization, ignore the complex relationship between the overall structure and attributes of the molecule, and it is difficult to optimize multiple pharmacological attributes at the same time, resulting in poor quality of optimization results.
A global corpus containing common functional groups and independent atoms is constructed, the molecular map is converted into a motif diagram, and forward diffusion and reverse generation are used to adjust the molecular structure through multi-attribute optimization conditions to ensure that the optimized molecules contain complex functional groups structures and meet the multi-attribute objectives.
It has achieved the maintenance of structural diversity and optimization effect of drug molecules in the multi-attribute optimization process, and improved the efficiency and success rate of drug molecules optimization.
Smart Images

Figure CN120260729A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of medicinal chemistry and computer technology, and particularly relates to a method and system for optimizing multiple properties of drug molecules based on a diffusion model. Background Art
[0002] The optimization of drug molecules is a key step in the drug R & D process, aiming to optimize the pharmacological properties of drugs, such as enhancing efficacy, reducing toxicity, and improving pharmacokinetic properties, by adjusting and modifying the molecular structure. This process can not only accelerate the drug R & D process, but also significantly improve the R & D success rate and reduce costs. In traditional drug molecule optimization, chemists usually rely on their own knowledge accumulation and experience, and conduct molecular improvement through laboratory synthesis and bioactivity testing. This method mainly adjusts the structure of candidate molecules step by step. Common means include functional group replacement, ring structure modification, introduction of polar groups, etc., aiming to optimize the performance of molecules in multiple pharmacological properties, such as enhancing drug efficacy, improving solubility, increasing blood-brain barrier permeability, or reducing toxicity. Although this traditional method has achieved certain success in practice, due to its reliance on a large number of experimental verifications and empirical adjustments, the R & D cycle is long, the cost is high, and the optimization efficiency is relatively limited.
[0003] With the continuous progress of artificial intelligence technology, especially deep learning methods, the way of drug molecule optimization has gradually undergone a revolutionary change. Modern deep learning methods can process and analyze the relationship between complex molecular structures and biological activities, helping chemists explore potential candidate molecules from a wider structural space. Advanced technologies such as generative adversarial networks (GANs), variational autoencoders (VAEs), and diffusion models can automatically generate new chemical structures or effectively optimize existing molecules while retaining the original pharmacological properties of the molecules. These methods learn the complex mapping relationship between molecular structures and activities through training on large-scale data sets, thereby providing more efficient and accurate optimization strategies, shortening the drug development cycle, and significantly improving the efficiency and success rate of new drug discovery.
[0004] Prior Art [1] Yu J, Xu T, Rong Y, et al. Structure-aware conditional variational auto-encoder for constrained molecule optimization [J]. Pattern Recognition, 2022, 126: 108581. This paper proposed a structure-aware conditional variational auto-encoder (SCVAE). The authors represented molecules as graph structures, modeled molecular topologies using graph neural networks, and controlled target properties through conditional encoding. The algorithm adopted joint training of reconstruction loss and property optimization objectives to ensure that the generated molecules met both the optimization objectives and the constraints. Therefore, prior art [1] SCVAE can effectively capture the overall structural information of molecules by modeling molecular topologies using graph neural networks. However, when dealing with complex functional groups or local structures, the expressive power of the graph structure may be insufficient. Especially in molecular optimization tasks, the optimized molecules generated by the decoder often may lack some key functional group structures, affecting the quality of the optimization results.
[0005] Prior Art [2] Maziarka Pocha A, Kaczmarczyk J, et al. Mol-CycleGAN: a generative model for molecular optimization [J]. Journal of Cheminformatics, 2020, 12(1): 2. This paper proposed a GAN-based molecular optimization model (Mol-CycleGAN). The model consists of two GANs with cyclic structures. The first is used to convert molecules without the target property into molecules with the property, and the second converts molecules with the target property into molecules without the property. By minimizing the difference between the original molecules and the molecules generated by the second GAN, the model realizes molecular optimization. Therefore, the optimization objective of prior art [2] Mol-CycleGAN is to minimize the difference between the original molecules and the generated molecules. However, this objective is too dependent on the contrastive loss and may not be able to fully capture the details of key pharmacological properties such as the binding of molecules to targets. In addition, during the optimization process, it is difficult for the model to optimize multiple properties simultaneously, limiting its performance on multi-property targets.
[0006] Prior Art [3] Xie J, Chen S, Lei J, et al. DiffDec: structure-aware scaffold decoration with an end-to-end diffusion model [J]. Journal of Chemical Information and Modeling, 2024, 64(7): 2554-2564. This paper proposed a molecular optimization method (DiffDec) based on the denoising diffusion probability model. By using improved diffusion techniques, three-dimensional spatial constraints were introduced to optimize molecular scaffold decoration. DiffDec can generate R groups with real geometric structures and generate multiple R groups for the same scaffold at different growth anchors. The optimized molecules showed a significant improvement in the binding affinity with the target. Therefore, in Prior Art [3], DiffDec uses the diffusion probability model to generate molecular structures, mainly relying on local modifications at different growth anchors. This local optimization strategy ignores the complex relationship between the overall structure and properties of the molecule and may lead to the inability to effectively adjust the molecule from a global level to obtain the best optimization effect.
[0007] In view of this, the present application is specifically proposed. Summary of the Invention
[0008] The present invention provides a multi-attribute optimization method and system for drug molecules based on a diffusion model. To better explore the functional group structure of molecules, the present invention constructs a global corpus containing common functional groups and single atoms, and converts the molecular graph into a motif graph through functional group matching. On this basis, the diffusion model is used to optimize the motif graph of the target molecule to ensure that the optimized molecule has a complex functional group structure. For the multi-attribute target optimization problem, the present invention defines multi-attribute optimization conditions and minimizes the multi-attribute loss during the training process, thereby guiding the molecule to be optimized towards the desired attribute target. In addition, the multi-attribute optimization diffusion model proposed by the present invention involves the addition and deletion of motif nodes during both the forward diffusion and reverse generation processes to adjust the overall structure of the molecule, ensuring that the optimized molecule maintains good diversity while meeting the multi-attribute target.
[0009] The present invention is realized through the following technical solutions:
[0010] In a first aspect, the present invention provides a multi-attribute optimization method for drug molecules based on a diffusion model, the method comprising:
[0011] Constructing a global corpus containing common functional groups and independent atoms, and based on the global corpus, converting the original molecular graph into a motif graph and initializing the motif nodes in the motif graph; the original molecular graph is the molecular graph to be optimized;
[0012] Define multi-attribute optimization conditions and construct a multi-attribute optimization diffusion model. Use motif node embedding and multi-attribute optimization conditions as the input of the multi-attribute optimization diffusion model. Based on the overall adjustment of the molecular structure by the multi-attribute optimization diffusion model, guide the molecule to optimize towards the multi-attribute optimization goal, and obtain the reconstructed motif graph;
[0013] Perform motif graph translation (i.e., translation from motif graph to molecular graph) according to the reconstructed motif graph to obtain a new molecule that meets the multi-attribute optimization goal; motif graph translation includes motif node type prediction and anchor atom prediction.
[0014] Furthermore, based on the global corpus, convert the original molecular graph into a motif graph, including:
[0015] Based on the global corpus, perform graph reduction on the original molecular graph, and represent the reduced graph as a motif graph; graph reduction includes:
[0016] Extract independent atoms and functional group structures in the original molecular graph as separate nodes, and use the separate nodes as motif nodes;
[0017] According to the connection relationship of motif nodes in the original molecular graph, establish edge connections between motif nodes; the connection relationship includes connections between independent atoms, connections between functional groups, and connections between functional groups and independent atoms;
[0018] Define the common atoms between adjacent motif nodes and the atoms connected by chemical bonds as anchor atoms, and the anchor atoms belong to the nodes within the functional group structure in the original molecular graph.
[0019] Furthermore, the connection between independent atoms includes: maintaining the original chemical bond connection;
[0020] The connection between functional groups includes:
[0021] If there are one or more common atoms between two functional groups in the original molecular graph, establish an edge connection;
[0022] If two functional groups are connected by a common chemical bond, establish an edge connection;
[0023] The connection between a functional group and an independent atom includes:
[0024] If the functional group and the independent atom are connected by a common chemical bond, establish an edge connection.
[0025] Furthermore, a global attribute condition vector c is introduced into the multi-attribute optimization conditions, and each item in the global attribute condition vector c describes the optimization goal of the molecule in a certain attribute.
[0026] Furthermore, the multi-attribute optimization diffusion model includes a forward diffusion process and a reverse generation process;
[0027] In the forward diffusion process, the node feature distribution in the motif graph is transformed into a Gaussian distribution by gradually adding noise; the entire forward diffusion process includes T time steps. In each time step, the features of all nodes are diffused once, and a node change is performed simultaneously until the preset number of time steps T is reached; among them, the initial input of the multi-attribute optimization diffusion model is obtained by concatenating the multi-attribute optimization condition c with each node feature x v concatenated;
[0028] In the reverse generation process, by learning a generation network p θ regenerates node features, predicts node operations, and predicts edge connections to reconstruct the motif graph.
[0029] Furthermore, the formula for the diffusion process is:
[0030]
[0031] where t represents the time step; x t represents the node feature at time step t; β t is the noise intensity, which is a preset hyperparameter; I is the identity matrix, and β t I is used to represent the covariance matrix of the noise; represents the Gaussian distribution, that is, it means that the diffused node feature x t obeys a Gaussian distribution with a mean of and a variance of β t I; q(x t |x t-1 ) represents the conditional probability distribution of obtaining the node feature x t-1 by adding noise to the node feature x t at time step t - 1.
[0032] Furthermore, generating node features is to generate node features by denoising based on a conditional probability model;
[0033] Predicting node operations is to predict whether each node needs to be added or deleted by introducing a classification network;
[0034] Predicting edge connections is to predict the edge connection probability between the newly added node v′ and the nodes one by one based on an edge generation network. If the prediction of the operation on node v is to add, then an edge connection relationship is directly established between the newly added node v′ and the current node v; the edge connection relationship between the newly added node and other nodes in the node set is predicted through the edge generation network.
[0035] Furthermore, the anchor atom prediction includes:
[0036] Based on the multi-class prediction model, predict the probability of each atom in the motif node represented by the functional group structure as an anchor point, and record it as the anchor point probability;
[0037] Based on the atom embeddings within the functional group and the neighbor motif node embeddings, calculate the edge connection probability between each atom within the functional group and the neighbor motif nodes;
[0038] Based on the anchor point probability and the edge connection probability, calculate the joint probability of the atoms within the functional group being connected to the neighbor motif nodes, and then select the atom with the maximum joint probability as the anchor point to connect to the neighbor motif nodes.
[0039] Furthermore, the loss function of the multi-attribute optimization objective includes node feature denoising loss, node operation loss, edge connection loss, anchor point prediction loss, and custom multi-attribute optimization condition loss.
[0040] In a second aspect, the present invention further provides a drug molecule multi-attribute optimization system based on a diffusion model, and the system includes:
[0041] A global corpus construction unit for constructing a global corpus containing common functional groups and independent atoms;
[0042] A motif graph construction unit for converting the original molecular graph into a motif graph based on the global corpus and initializing the motif nodes in the motif graph; the original molecular graph is the molecular graph to be optimized;
[0043] A motif graph reconstruction unit for defining multi-attribute optimization conditions and constructing a multi-attribute optimization diffusion model, using the motif node embeddings and the multi-attribute optimization conditions as the input of the multi-attribute optimization diffusion model, and overall adjusting the molecular structure based on the multi-attribute optimization diffusion model to guide the molecule to optimize towards the multi-attribute optimization objective, obtaining the reconstructed motif graph;
[0044] A motif graph translation unit for performing motif graph translation (i.e., translation from the motif graph to the molecular graph) according to the reconstructed motif graph to obtain a new molecule that meets the multi-attribute optimization objective; the motif graph translation includes motif node type prediction and anchor atom prediction.
[0045] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0046] A method and system for multi-attribute optimization of drug molecules based on a diffusion model. The present invention constructs a global corpus containing common functional groups and independent atoms, and converts the original molecular graph into a motif graph based on the corpus. During the molecular optimization process, the motif graph is used as the input and output of the diffusion model to ensure that the optimized molecule contains common functional group structures. The designed multi-attribute optimization diffusion model adjusts the molecular structure as a whole during the forward diffusion and reverse generation processes, guiding the molecule to optimize towards the multi-attribute target direction, thereby ensuring that the optimized molecule meets the multi-attribute target and maintains structural diversity simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, form a part of this application, and do not limit the embodiments of the present invention. In the drawings:
[0048] Figure 1 is a flowchart of a method for multi-attribute optimization of drug molecules based on a diffusion model according to the present invention;
[0049] Figure 2 is a detailed flowchart of a method for multi-attribute optimization of drug molecules based on a diffusion model according to the present invention;
[0050] Figure 3 is a schematic diagram of motif graph construction according to the present invention;
[0051] Figure 4 is a schematic diagram of the optimization diffusion process of the multi-attribute optimization diffusion model according to the present invention;
[0052] Figure 5 is a schematic diagram of motif graph translation according to the present invention;
[0053] Figure 6 is a block diagram of the structure of a system for multi-attribute optimization of drug molecules based on a diffusion model according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the embodiments and the drawings. The illustrative embodiments and descriptions of the present invention are only used to explain the present invention and do not limit the present invention.
[0055] The solution of the present invention includes steps such as the initial representation of molecules, the definition of multi-attribute optimization conditions, the forward diffusion, reverse generation, and motif graph translation of the multi-attribute optimization diffusion model. The overall process is as Figure 2 shown. The key design points of the present invention are as follows:
[0056] 1) The present invention establishes a global corpus of common functional groups and independent atoms, represents the molecule as a motif graph for optimization, and solves the problem that the optimized molecule lacks common functional group structures.
[0057] 2) Define multi-attribute optimization conditions and set a multi-attribute optimization weight vector, which can achieve the weight allocation of the optimization priorities of different attributes and realize a customized optimization goal.
[0058] 3) Design a multi-attribute optimization diffusion model, and use motif node embedding and multi-attribute optimization conditions as the inputs of the multi-attribute optimization diffusion model. In the forward diffusion process, the model adds noise to the node features and adjusts the molecular composition; in the reverse generation process, the model regenerates the node features, predicts the node operations, and predicts the edge connections, thereby reconstructing the motif graph. Finally, through motif node type prediction and anchor atom prediction, the translation from the motif graph to the molecular graph is realized, and a new molecule that meets the multi-attribute optimization goal is obtained.
[0059] Example 1
[0060] As Figure 1 shown, a method for multi-attribute optimization of drug molecules based on a diffusion model according to the present invention includes:
[0061] S1: Construct a global corpus containing common functional groups and independent atoms, and based on the global corpus, convert the original molecular graph into a motif graph and initialize the motif nodes in the motif graph; the original molecular graph is the molecular graph to be optimized;
[0062] S2: Define multi-attribute optimization conditions and construct a multi-attribute optimization diffusion model, use motif node embedding and multi-attribute optimization conditions as the inputs of the multi-attribute optimization diffusion model, and overall adjust the molecular structure based on the multi-attribute optimization diffusion model to guide the molecule to optimize towards the multi-attribute optimization goal, and obtain the reconstructed motif graph;
[0063] S3: Perform motif graph translation (i.e., translation from the motif graph to the molecular graph) according to the reconstructed motif graph to obtain a new molecule that meets the multi-attribute optimization goal; the motif graph translation includes motif node type prediction and anchor atom prediction.
[0064] In this embodiment, in step S1, constructing a global corpus containing common functional groups and independent atoms specifically includes:
[0065] The molecule is represented by an undirected graph G=(V, E), where V represents the set of atoms of the molecule and E represents the set of chemical bonds between atoms. To focus on the functional group structure in the molecule and keep the complexity of the molecular structure during the molecular optimization process, common chemical functional groups are collected and sorted to construct a functional group corpus.
[0066] In this embodiment, in step S1, converting the original molecular graph into a motif graph based on the functional group corpus includes:
[0067] Based on the functional group corpus, the original molecular graph is reduced, and the reduced graph is represented as a motif graph. The graph reduction includes: extracting independent atoms and functional group structures in the original molecular graph as separate nodes, and establishing the connection relationships between the reduced nodes. In molecular chemistry, a motif often refers to a substructure with specific functions or structural characteristics, including functional groups and other significant structural units. In the present invention, the reduced graph is called a motif graph, and the nodes in the graph are called motif nodes. A schematic diagram of the construction of the motif graph is as shown in Figure 3 as follows:
[0068] Establishment of motif nodes: RDKit is a Python toolkit for cheminformatics and can be used to match molecular fragments. In the process of extracting motif nodes from a drug molecule, the functional group corpus is traversed to read functional groups one by one, and then the pattern of the functional group is matched using RDKit. If the functional group exists in the molecule, a group of atoms involved in the functional group is extracted as a motif node. Since not all atoms in the molecule are functional groups, to ensure the integrity of the molecule, independent atoms are also regarded as motif nodes. In the present invention, the common atoms between adjacent motif nodes and the atoms connected by chemical bonds are defined as anchor atoms, and the anchor atoms belong to the nodes within the functional group structure in the original molecular graph.
[0069] Establishment of edges: Motif nodes include independent atoms and functional group structures. Edge connections are established according to their connection relationships in the original molecular graph, where the connection relationships are divided into connections between independent atoms, connections between functional groups, and connections between functional groups and independent atoms. Among them, the connections between independent atoms maintain the original chemical bond connections. There are two cases for the connections between functional groups: one is that if there are one or more common atoms between two functional groups in the original molecular graph, an edge connection is established; the other is that if two functional groups are connected by a common chemical bond (such as a double bond, a single bond, etc.), an edge connection is also established. The connection between a functional group and an independent atom depends on their common chemical bond. If they are connected by a common chemical bond, an edge connection is established.
[0070] Construction of the motif graph: Graph reduction is completed during the process of extracting motif nodes, and the reduced graph is represented as a motif graph G Motif =(V Motif , E Motif ), V Motif ={v group ∪v atom}, E Motif ={(v i , v j )|v i , v j ∈V Motif}. V Motif is the motif node set, and V Motif is composed of functional group nodes v groupIt consists of independent atomic nodes. E Motif is a set of edge connection relationships between motif nodes.
[0071] To initialize the embedding representation of a molecule, the present invention constructs a global corpus covering all possible functional groups and atomic types. Each motif node can be initially embedded and represented by a one-hot vector, and the length of the one-hot vector represents the total number of functional groups and atomic types in the corpus. There are a total of D entries in the corpus, and the length of the one-hot vector for each node is D, where only the value at the corresponding index position is 1 and the other positions are 0. Suppose the index of the functional group or atom represented by a certain node v in the corpus is j, then the one-hot vector x of this node v is shown in Formula 1.
[0072]
[0073] where k is the index in this vector, and k ∈ {0, 1,..., D - 1}.
[0074] In this embodiment, in step S2, multi-attribute optimization conditions are defined, including:
[0075] During the molecule optimization process, the present invention defines multi-attribute optimization conditions. Given a molecule M and its attribute set y = {y1, y2,..., y k}, the goal is to generate a new molecule M' that meets the requirements of multi-attribute optimization. To achieve this goal, a global attribute condition vector c is introduced, and each item in this vector describes the optimization goal of the molecule in a certain attribute. Taking the three attribute conditions of drug-likeness (QED), lipophilicity (LogP), and activity against dopamine type 2 receptor (DRD2) as examples, the definition of the condition vector is shown in Formula 2.
[0076] c = [c QED , c LogP , c DRD2 (2)
[0077] where c QED , c LogP , c DRD2 respectively represent the conditions for QED, LogP, and DRD2 activity. Drug-likeness greater than 0.8 is represented as 1, otherwise as 0; LogP between -1 and 6 is represented as 1, otherwise as 0; having activity against DRD2 is represented as 1, and having no activity is represented as 0.
[0078] In this embodiment, in step S2, a multi-attribute optimization diffusion model is constructed, including:
[0079] The diffusion model based on multi-attribute optimization is a generative deep learning framework designed to generate new molecules that meet the target design requirements under multiple attribute constraints. The multi-attribute optimization diffusion model proposed in this invention includes a forward diffusion process of gradually adding noise and a reverse generation process of denoising. The schematic diagram of the diffusion process is as shown in Figure 4 as follows.
[0080] (1) Forward diffusion process
[0081] In the forward diffusion process, the node feature distribution in the motif graph is transformed into a Gaussian distribution by gradually adding noise. The forward diffusion is divided into two parts: node feature diffusion and node change. Among them, node feature diffusion is to add noise to the features of each node in the motif graph; node change is to randomly insert or delete nodes to simulate the dynamic changes of the graph structure, so as to maintain the diversity between the optimized molecule and the original molecule. The entire forward diffusion process includes T time steps. In each time step, the features of all nodes are diffused once, and a node change is performed at the same time until the preset number of time steps T is reached to complete the entire diffusion process. The multi-attribute optimization condition c is concatenated with each node feature x v to obtain the initial input of each node in the diffusion model, as shown in formula (3).
[0082] x0 = [x v ∥c] (3)
[0083] where ∥ represents the vector concatenation operation, and x0 represents the initial input of node v in the motif graph in the diffusion model.
[0084] Node feature diffusion: Given the initial motif graph G Motif =(V Motif , E Motif ), taking the node v in the graph as an example, its diffusion process is defined as shown in formula (4).
[0085]
[0086] where t represents the time step; x t represents the node feature at time step t; β t is the noise intensity, which is a preset hyperparameter; I is the identity matrix, and β t I is used to represent the covariance matrix of the noise; represents the Gaussian distribution, that is, it means that the diffused node feature x t obeys a Gaussian distribution with a mean of and a variance of β t I; q(x t |x t-1 ) represents the node feature x t-1Above, the node feature x is obtained by adding noise. t The conditional probability distribution.
[0087] In the present invention, a linear scheduling strategy is used for β t to provide values, and β t is set to a value that increases uniformly over time:
[0088]
[0089] where β min and β max are the boundaries of the set range of the noise intensity hyperparameter schedule, and T is the total number of time steps.
[0090] Node changes: During the diffusion process, node insertion and deletion operations are random. At each time step t, let the current motif graph be G t =(V t , E t ). The probabilities of inserting or deleting a node are p add and p del respectively. If p add is greater than or equal to p del , then a new node is inserted, and its feature is initialized to a Gaussian distribution with a mean of zero and a variance of σ 2 . The newly inserted node is added to the node set V t in G t ; if p add is less than p del , then a node is randomly deleted, and the node set V t and the edge set E t and E t in G
[0091] (2) Reverse generation process
[0092] In the reverse generation process, a generative network p θ is learned to regenerate node features, predict node operations, and predict edge connections, thereby reconstructing the motif graph.
[0093] Node feature generation: Use the conditional probability model p θ (x t-1 |x t ) to denoise and generate node features, as shown in formula (6).
[0094]
[0095] where μ θ and ∑ θ are generative networks, μ θ (xt , t) and ∑ θ (x t , t) are the mean and variance calculated by the generation network respectively.
[0096] μ θ and ∑ θ The generation network of and ∑ first aggregates the information of neighboring nodes based on the graph neural network, and then generates the mean and variance of the nodes through a multi-branch network. Assume x t corresponds to the node x in the motif graph v , and the hidden representation h of the updated node is obtained after aggregating the neighboring information v , as shown in formula (7).
[0097]
[0098] where N(v) is the neighborhood of node v, l represents the l-th iteration of the graph neural network, W (l) and b (l) are both learnable parameters, and ReLU is a non-linear activation function.
[0099] To generate the mean and variance of node features, a multi-branch network is designed for prediction. The multi-branch network consists of a shared backbone network followed by two independent output branches. The backbone shared network uses a multi-layer perceptron (MLP) to perform feature transformation on the node hidden representation , as shown in formula (8).
[0100]
[0101] where MLP shared is the backbone network, composed of several fully connected layers; z v is the intermediate representation, which will be passed to the mean and variance prediction branches later.
[0102] Mean prediction branch. Taking z v as the input, it passes through a fully connected layer MLP μ to output and predict the mean μ v , as shown in formula (9).
[0103] μ v = MLP μ (z v )(9)
[0104] where MLP μ is a simple fully connected layer, and the output dimension is the same as the node feature dimension.
[0105] Variance prediction branch. Taking z v as the input, it passes through a fully connected layer MLPσ The output is the predicted log variance As shown in formula (10).
[0106]
[0107] The output log variance is to ensure the variance of the prediction is always non - negative
[0108] Node operation prediction: Introduce a classification network to predict whether each node needs to be added or deleted, as shown in formulas (11) and (12).
[0109]
[0110] where is a two - layer fully - connected layer network, W1 and W2 are learnable weight parameters, and b1 and b2 are bias parameters. The output is a probability vector of length 2, representing the probability of adding or deleting an operation for this node. If the addition probability is greater, an addition operation is performed on this node, that is, a new node is generated and directly connected to this node; otherwise, a deletion operation is performed on this node, that is, this node will be removed.
[0111] Edge connection generation: If the operation prediction for node v is addition, a direct edge connection relationship is established between the newly added node v′ and the current node v, that is, the edge (v, v′) is added to the edge set E at time step t t . In addition, there may also be connection relationships between the newly added node and other nodes in the node set, which can be predicted through the edge generation network . By traversing each node in the node set V t except node v at time step t, the edge connection probability between the newly added node v′ and these nodes is predicted one by one using the edge generation network, as shown in formula (13).
[0112]
[0113] where h v′ is the embedding vector of the newly added node, h u refers to the embedding vectors of other nodes except node v represents the neural network model for edge generation, σ is the Sigmoid function, and the output is a probability value between 0 and 1, representing the probability of establishing an edge connection between the newly added node v′ and node u.
[0114] In this embodiment, the motif graph translation in step S3 includes:
[0115] After step S2, the reconstructed motif graph is obtained. Let the reconstructed motif graph obtained by the multi-attribute optimized diffusion model be To achieve the reverse translation from the motif graph to the molecular graph, it is necessary to first determine the type of each motif node, and then predict which atom in the motif node is directly connected to the neighbor nodes. The schematic diagram of the motif graph translation is as Figure 5 shown.
[0116] During the reverse generation process of the multi-attribute optimized diffusion model, the output is the feature distribution of each node That is, the probability distribution of the node features. To recover from this distribution to the specific motif node features (such as functional groups or single atoms), it is necessary to map the predicted distribution of each node back to a definite motif node. The Softmax function can be used to map the node features to a probability vector, and the specific motif feature of this node can be determined by the maximum value in this probability vector, as shown in formula (14).
[0117]
[0118] Where, is the feature distribution of node v in the reverse-generated motif graph, with a shape of D dimensions, and D is the number of possible motif types in the established corpus; is the predicted value of node v corresponding to the i-th class; represents the exponential transformation of the node feature distribution. The purpose of the transformation is to enhance the weight of the larger feature values while reducing the influence of the smaller feature values, so as to highlight the most likely category; is the sum of the exponential values of all feature terms of the node, and the purpose is to normalize the output result so that the sum of the probabilities of all categories is 1; is the vector obtained after being mapped by the Softmax function, where each element represents the probability that node v belongs to a specific motif node (such as a functional group or an atom).
[0119] Since the motif node may be a single atom or a functional group, when the motif node is a functional group, it is impossible to directly generate a specific molecular graph only relying on the connection relationship between the motif nodes. To obtain a clear molecular structure, it is necessary to predict which atom in the functional group is used as the anchor point, and complete the construction of the molecular graph through the connection relationship between the motifs. The basic idea of the anchor point prediction is to select one or more atoms from the internal atoms of the functional group as the "anchor points" for connecting other motif nodes. The present invention first predicts the probability of each atom in the functional group as an anchor point through a multi-classification prediction model, and then calculates each atom a in the functional group i and each neighbor motif node v j The connection probability p edge,ij, finally, determine which atom is used as the anchor to connect with specific neighbor nodes by combining the probability of the anchor point and the edge connection probability. Assume that a certain functional group f is a motif node, and there are multiple atoms a1, a2, …, a inside this functional group m , and the embeddings of these atoms are represented as h1, h2, …, h m , which can be calculated by a multi-layer perceptron, as shown in formula (15).
[0120]
[0121] Among them, MLP anchor is a multi-layer perceptron, the input is the feature of the functional group node and the output is the embedding h of each atom i i .
[0122] Next, calculate the probability p that each atom is an anchor point anchor,i , as shown in formula (16).
[0123]
[0124] If there are multiple neighbor nodes corresponding to the motif node of this functional group, taking one of the neighbor nodes v j as an example, calculate the edge connection probability p between the atom a in the functional group i and v j , as shown in formula (17). edge,ij , as shown in formula (17).
[0125]
[0126] Among them, ∥ represents the vector concatenation operation, MLP edge is a multi-layer perceptron, W edge is a learnable weight parameter, b edge is a bias parameter, and σ is the Sigmoid function.
[0127] Then, calculate the joint probability that the atom a in the functional group i is connected to the neighbor node v j as the edge connection probability, as shown in formula (18).
[0128] p connect,ij = p anchor,i · p edge,ij (18)
[0129] Finally, select the atom with the largest joint probability as the anchor to connect with the neighbor node v j , as shown in formula (19).
[0130]
[0131] Among them, The function is used to obtain the atom with the maximum probability value, indicating the predicted anchor atom.
[0132] In this embodiment, the loss function of the multi-attribute optimization objective includes node feature denoising loss, node operation loss, edge connection loss, anchor prediction loss, and custom multi-attribute optimization condition loss.
[0133] Since the number of nodes changes during the reverse generation process of diffusion, the generated motif graph is inconsistent with the nodes of the initial motif graph. Therefore, when calculating the denoising loss of node features, only the denoising error corresponding to the nodes of the initial graph G0 is considered, as shown in formula (20).
[0134]
[0135] where v ∈ V t ∩V0 represents the common nodes of the initial graph and the current graph, represents the embedding vector of node v in the initial graph, is to calculate the mathematical expectation.
[0136] During the reverse generation process of diffusion, the loss of the addition or deletion operations performed on the nodes compared to the initial nodes is as shown in formula (21).
[0137]
[0138] Among them, is the predicted node operation situation, is the operation situation of this node in the original molecule.
[0139] For predicting the edge connection situation of newly added nodes with other nodes, the predicted loss is as shown in formula (22).
[0140]
[0141] Among them, is the predicted edge connection situation, is the edge connection situation in the original molecule.
[0142] For the motif node set V Motif , assuming the number of functional groups is N c , and there are m atoms in this functional group C, the cross-entropy loss of anchor prediction is as shown in formula (23).
[0143]
[0144] Among them, y anchor,iIndicates whether atom i is an anchor point of functional group C, y anchor,i = 1 indicates that atom i is an anchor point; p anchor,i is the probability that the model predicts atom i to be an anchor point.
[0145] To achieve custom optimization of specific properties, an attribute weight vector w is defined, where each item represents the weight assignment for the optimization priorities of different properties. The multi-attribute optimization loss function is shown in formula (24).
[0146]
[0147] where y j is the conditional value of the molecule on attribute j, is the normalized predicted value of the optimized molecule on attribute j. When w j = 0, attribute j is ignored; when w j > 0, attribute j is optimized.
[0148] In summary, the loss function of the multi-attribute optimization objective is:
[0149]
[0150] where is the node feature denoising loss, is the node operation loss, is the edge connection loss, is the anchor point prediction loss, is the custom multi-attribute optimization condition loss.
[0151] During the training process, the multi-attribute optimization diffusion model gradually adds noise to the motif graph through forward diffusion and denoises during the reverse generation process to gradually restore the molecular structure. During the reverse generation process, newly added nodes are marked so that the node operations and edge connection losses can be uniformly calculated after the generation is completed. When the reverse generation is completed and the optimized molecule is obtained, first calculate the node feature denoising loss to ensure that the generated motif features are close to the performance of the real molecule in motif features. Subsequently, calculate the node operation loss to supervise the addition and deletion of nodes, and calculate the edge connection loss to predict the connection mode of the newly added node to other nodes. Then, calculate the anchor point prediction loss to determine the anchor atoms of the functional group so that the functional group can be correctly connected to adjacent motif nodes. Finally, calculate the multi-attribute optimization loss to ensure that the values of the molecule on multiple attributes meet the optimization objectives. By minimizing the total loss the model can be continuously optimized during the backpropagation process so that it can generate molecules that meet the multi-attribute optimization objectives.
[0152] The present invention constructs a global corpus containing common functional groups and independent atoms, and converts the original molecular graph into a motif graph based on the corpus. During the molecular optimization process, the motif graph is used as the input and output of the diffusion model to ensure that the optimized molecule contains common functional group structures. In the designed multi-attribute optimization diffusion model, during the forward diffusion and reverse generation processes, the model adjusts the molecular structure as a whole to guide the molecule to optimize towards the multi-attribute target, thereby ensuring that the optimized molecule meets the multi-attribute target and maintains structural diversity.
[0153] Example 2
[0154] As Figure 6 shown, the difference between this example and Example 1 is that this example provides a multi-attribute optimization system for drug molecules based on a diffusion model, and this system corresponds one-to-one with a multi-attribute optimization method for drug molecules based on a diffusion model in Example 1; this system includes:
[0155] A global corpus construction unit for constructing a global corpus containing common functional groups and independent atoms;
[0156] A motif graph construction unit for converting the original molecular graph into a motif graph based on the global corpus and initializing the motif nodes in the motif graph; the original molecular graph is the molecular graph to be optimized;
[0157] A motif graph reconstruction unit for defining multi-attribute optimization conditions and constructing a multi-attribute optimization diffusion model, using motif node embedding and multi-attribute optimization conditions as the input of the multi-attribute optimization diffusion model, and adjusting the molecular structure as a whole based on the multi-attribute optimization diffusion model to guide the molecule to optimize towards the multi-attribute optimization target to obtain the reconstructed motif graph;
[0158] A motif graph translation unit for performing motif graph translation (i.e., translation from motif graph to molecular graph) according to the reconstructed motif graph to obtain a new molecule that meets the multi-attribute optimization target; motif graph translation includes motif node type prediction and anchor atom prediction.
[0159] Among them, the execution processes of each unit can be carried out according to the process steps of a multi-attribute optimization method for drug molecules based on a diffusion model in Example 1, and will not be elaborated one by one in this example.
[0160] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.
[0161] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0162] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0163] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0164] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above description is only for the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for optimizing multiple attributes of drug molecules based on a diffusion model, characterized in that, The method includes: Constructing a global corpus containing common functional groups and independent atoms, and based on the global corpus, converting the original molecular graph into a motif graph and initializing motif nodes in the motif graph; the original molecular graph is the molecular graph to be optimized; Defining multi-attribute optimization conditions and constructing a multi-attribute optimization diffusion model, using motif node embeddings and the multi-attribute optimization conditions as inputs to the multi-attribute optimization diffusion model, and globally adjusting the molecular structure based on the multi-attribute optimization diffusion model to guide the molecule to optimize towards the multi-attribute optimization goal to obtain a reconstructed motif graph; Performing motif graph translation according to the reconstructed motif graph to obtain a new molecule that meets the multi-attribute optimization goal; the motif graph translation includes motif node type prediction and anchor atom prediction.
2. The multi-attribute optimization method of drug molecules based on the diffusion model according to claim 1, wherein, Based on the global corpus, converting the original molecular graph into a motif graph includes: Based on the global corpus, performing graph reduction on the original molecular graph and representing the reduced graph as a motif graph; the graph reduction includes: Extracting independent atoms and functional group structures in the original molecular graph as separate nodes, and using the separate nodes as motif nodes; According to the connection relationships of the motif nodes in the original molecular graph, establishing edge connections between the motif nodes; the connection relationships include connections between independent atoms, connections between functional groups, and connections between functional groups and independent atoms; And defining common atoms between adjacent motif nodes and atoms connected by chemical bonds as anchor atoms, and the anchor atoms belong to nodes within the functional group structure in the original molecular graph.
3. The multi-attribute optimization method of drug molecules based on a diffusion model according to claim 2, wherein, The connection between independent atoms includes: maintaining the original chemical bond connection; The connection between functional groups includes: If there are one or more common atoms between two functional groups in the original molecular graph, establishing an edge connection; If two functional groups are connected by a common chemical bond, establishing an edge connection; The connection between a functional group and an independent atom includes: If a functional group and an independent atom are connected by a common chemical bond, establishing an edge connection.
4. A method for optimizing multiple properties of drug molecules based on a diffusion model according to claim 1, wherein, A global attribute condition vector is introduced into the multi-attribute optimization conditions, and each item in the global attribute condition vector describes the optimization goal of the molecule in a certain attribute.
5. A method for optimizing multiple properties of drug molecules based on a diffusion model according to claim 1, characterized in that, The multi-attribute optimization diffusion model includes a forward diffusion process and a reverse generation process; In the forward diffusion process, the node feature distribution in the motif graph is transformed into a Gaussian distribution by gradually adding noise; the entire forward diffusion process includes T time steps, and in each time step, the features of all nodes are diffused once, and at the same time, a node change is performed until the preset number of time steps T is reached; among them, the initial input of the multi-attribute optimization diffusion model is obtained by splicing the multi-attribute optimization conditions with each node feature; In the reverse generation process, a generation network is learned to regenerate node features, predict node operations, and predict edge connections, so as to reconstruct the motif graph.
6. The multi-attribute optimization method of drug molecules based on the diffusion model according to claim 5, characterized in that The formula for the diffusion process is: where \(t\) represents the time step; \(x\) t represents the node feature at time step \(t\); \(\beta\) t is the noise intensity, which is a pre-set hyperparameter; \(I\) is the identity matrix, and \(\beta\) t \(I\) is used to represent the covariance matrix of the noise; represents the Gaussian distribution, that is, it represents that the diffused node feature \(x\) t follows a Gaussian distribution with a mean of and a variance of \(\beta\) t \(I\); \(q(x\) t | \(x\) t-1 ) represents the conditional probability distribution of obtaining the node feature \(x\) t-1 by adding noise to the node feature \(x\) t at time step \(t - 1\).
7. A method for optimizing multiple attributes of drug molecules based on a diffusion model according to claim 5, characterized in that, The generated node features are denoised to generate node features based on a conditional probability model; The prediction of node operations is to predict whether each node needs to be added or deleted by introducing a classification network; The predicted edge connection is to predict the edge connection probability between the newly added node v and the nodes one by one based on the edge generation network. If the prediction result for the operation on node v is new addition, then the newly added node v ′ and the current node v directly establish an edge connection relationship; the edge generation network is used to predict the edge connection relationship between the newly added node and other nodes in the node set. ′ 8. A method for optimizing multiple properties of drug molecules based on a diffusion model according to claim 1, characterized in that, The anchor atom prediction includes: Based on the multi-classification prediction model, predict the probability of each atom in the motif node represented by the functional group structure as an anchor point, and record it as the anchor point probability; Based on the atom embeddings within the functional group and the neighbor motif node embeddings, calculate the edge connection probability of each atom within the functional group with the neighbor motif nodes; Based on the anchor point probability and the edge connection probability, calculate the joint probability of the atoms within the functional group being connected to the neighbor motif nodes, and then select the atom with the maximum joint probability as the anchor point to connect to the neighbor motif nodes.
9. A method for optimizing multiple properties of drug molecules based on a diffusion model according to claim 1, characterized in that The loss function of the multi-attribute optimization objective includes node feature denoising loss, node operation loss, edge connection loss, anchor point prediction loss, and custom multi-attribute optimization condition loss.
10. A drug molecule multi-attribute optimization system based on a diffusion model, characterized in that, The system includes: A global corpus construction unit for constructing a global corpus containing common functional groups and independent atoms; A motif graph construction unit for converting the original molecular graph into a motif graph based on the global corpus and initializing the motif nodes in the motif graph; the original molecular graph is the molecular graph to be optimized; A motif graph reconstruction unit for defining multi-attribute optimization conditions and constructing a multi-attribute optimization diffusion model, using the motif node embeddings and the multi-attribute optimization conditions as the input of the multi-attribute optimization diffusion model, and overall adjusting the molecular structure based on the multi-attribute optimization diffusion model to guide the molecule to optimize towards the multi-attribute optimization objective to obtain the reconstructed motif graph; A motif graph translation unit for performing motif graph translation according to the reconstructed motif graph to obtain a new molecule that meets the multi-attribute optimization objective; the motif graph translation includes motif node type prediction and anchor atom prediction.