A method and system for predicting the association between microorganisms and drugs

By constructing a microbial-drug association network and interaction network, combining multimodal attribute graphs, a regularized graph neural network model was established, which solved the problem that the existing technology could not construct interpretable node characteristics, achieved high-accurate microbial-drug association prediction, and dealt with the sparseness of the data set.

CN115472305BActive Publication Date: 2025-05-30GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210938454.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-05
Publication Date
2025-05-30
Estimated Expiration
2042-08-05

AI Technical Summary

Technical Problem

Existing methods for predicting microbial-drug associations cannot construct interpretable node characteristics of biology and drugs, and it is difficult to deal with the problem of sparsity in the microbial-drug association dataset.

Method used

By constructing a microbial-drug association and interaction network, combining multimodal attribute graphs, a regularized graph neural network model is established to generate interpretable node characteristics of biology and drugs, and to consider the sparseness of the data set.

Benefits of technology

The construction of interpretable node characteristics of biology and drugs is achieved, the accuracy of microbial-drug association prediction is improved, and the sparsity problem of data sets is effectively dealt with.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115472305B_ABST
    Figure CN115472305B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for predicting the association effect between microorganisms and drugs, comprising the following steps: S1. constructing an association network Net1 of microorganisms and drugs through a microorganism-drug association database; S2. constructing an interaction network Net2 through the microorganism-drug association database; S3. constructing a multi-modal attribute graph of microorganisms and drugs according to the comprehensive similarity of drugs, the drug network topology of drugs, the functional similarity of microorganisms, and the genomic sequences; S4. establishing a graph neural network model introducing regularization; S5. obtaining embedding representations Z1 and Z2; inputting Z1 and Z2 into the graph neural network for training to obtain a trained graph neural network; S6. obtaining a dataset to be predicted, and predicting the association effect between microorganisms and drugs in the dataset to be predicted through the trained graph neural network. The present invention solves the problem that the prior art cannot construct interpretable node features of organisms and drugs, and has the characteristic of being able to consider the sparsity problem brought by the existing microorganism-drug association datasets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of bioinformatics, and more specifically, to a method for predicting the association between microorganisms and drugs. Background Art

[0002] In recent years, the research focus in the medical field has been to explore the relationship between microbial community imbalance and drug efficacy and toxicity. However, there is still a lack of comprehensive understanding of the complex mechanisms of the interaction between the microbial community and drugs in the human body. Currently, new drug development faces two major challenges. On the one hand, the discovery of antibiotics is very limited. Most of the work focuses on optimizing or combining known compounds. It is difficult to culture the target strains under laboratory conditions, and most drugs fail during the experimental process. On the other hand, the number of drug-resistant bacteria is also increasing at an alarming rate. More and more studies have shown that there is a close interaction between microorganisms and drugs. Some connections between microorganisms and drugs have been confirmed by culture experiments, but they are not sufficient to clarify the complex interaction mechanism between human microorganisms and drugs. Therefore, there is an urgent need to develop an effective method to systematically explore the possible connections between microorganisms and drugs.

[0003] Currently, there are two types of computational methods for predicting the relationship between microorganisms and drugs.

[0004] The first type of method mainly focuses on similarity metrics. For example, the HMDAKATZ method using the KATZ metric. However, such metrics are too simple to fully reflect similarity and will lead to inaccurate association identification.

[0005] The second type of method uses graph-based learning methods, which use rich semantic information in the graph data representation and have better prediction ability compared to previous methods based on similarity metrics. Currently, there are two common graph representation learning methods: metapaths and graph convolutional networks.

[0006] The metapath algorithm mainly uses the edge information of the microorganism-drug association for prediction. The metapath algorithm combines metapath2vec with neural network suggestions to learn the low-dimensional embedding representations of microorganisms and drugs. The metapath algorithm does help to improve the prediction ability of the model, but it is too dependent on edge information. When introducing new drugs or new microorganisms, considering the lack of existing edge information, it will naturally lead to prediction failure.

[0007] Compared with the meta-path, the GCN method can not only capture edge information but also node information. Therefore, in current methods, using the GCN method to predict the microbial-drug association has attracted wide interest. Long et al. first applied the GCN encoder to the microbial-drug association method GCNMDA and introduced the conditional random field into the GCN hidden layer. There is also an existing method of node-level attention GCN, EGATMDA, to learn the embeddings of nodes (i.e., microorganisms and drugs), which can effectively retain the target neighbors of the graph and only retain relevant information. However, existing methods fail to construct node features containing biological information.

[0008] In summary, existing methods for predicting the microbial-drug association have the problem of being unable to construct rich and interpretable node features of organisms and drugs. Therefore, how to invent a method for predicting the microbial-drug association that can construct interpretable node features of organisms and drugs is a technical problem that urgently needs to be solved in this field. Summary of the Invention

[0009] In order to solve the problem that existing technologies cannot construct interpretable node features of organisms and drugs, the present invention provides a method for predicting the microbial-drug association, which has the characteristic of being able to take into account the sparsity problem brought by existing microbial-drug association data sets.

[0010] To achieve the above object of the present invention, the following technical solutions are adopted:

[0011] A method for predicting the microbial-drug association includes the following steps:

[0012] S1. Construct an association network of microorganisms and drugs through a microbial-drug association database, and call the association network Net1;

[0013] S2. Retrieve the relevant interactions of microorganisms-microorganisms through the microorganism database in the microbial-drug association database, and retrieve the relevant interactions of drugs-drugs through the drug database in the microbial-drug association database; construct an interaction network according to the relevant interactions of microorganisms-microorganisms and the relevant interactions of drugs-drugs, and call the interaction network Net2;

[0014] S3. Construct a topological attribute network of drugs through the drug database, and construct a microbial gene sequence through the microorganism database; construct a multi-modal attribute graph of microorganisms and drugs according to the comprehensive similarity attributes of drugs in the drug database, the topological attributes of the drug network, the functional similarity attributes of microorganisms in the microorganism database, and the genomic sequence attributes.

[0015] S4. Based on the multi-modal attribute graphs of Net1, Net2, and microorganism-drug, establish a graph neural network model with regularization introduced;

[0016] S5. Input Net1, Net2 combined with the multi-modal attribute graph of microorganism-drug into the graph neural network model to obtain the embedding representations Z1 and Z2; input the embedding representations Z1 and Z2 into the graph neural network for training to obtain a trained graph neural network;

[0017] S6. Obtain the dataset to be predicted, and predict the association effect between microorganisms and drugs in the dataset to be predicted through the trained graph neural network.

[0018] The present invention constructs an association network of microorganisms and drugs through a microorganism-drug association database, and further obtains the interaction networks Net1 and Net2; establishes a graph neural network model with regularization introduced, inputs Net1, Net2 combined with the multi-modal attribute graph of microorganism-drug into the graph neural network model to obtain the embedding representations Z1 and Z2, and inputs the embedding representations Z1 and Z2 into the graph neural network for training to obtain a trained graph neural network; constructs interpretable node features of organisms and drugs, and considers the sparsity problem brought by the existing microorganism-drug association datasets.

[0019] Preferably, in step S3, construct a topological attribute network of drugs through a drug database, and construct microorganism gene sequences through a microorganism database; according to the comprehensive similarity attributes of drugs in the drug database, the topological attributes of the drug network, the functional similarity attributes of microorganisms in the microorganism database, and the genomic sequence attributes, the specific steps for constructing the multi-modal attribute graph of microorganism-drug are as follows:

[0020] S301. Construct a similarity feature matrix of drugs according to the drug similarity attributes in the drug database, and construct a topological attribute network of drugs through the drug database, so as to obtain a second attribute feature matrix of drugs;

[0021] S302. Construct a similarity feature matrix of microorganisms according to the functional similarity attributes of microorganisms in the microorganism database, and construct microorganism gene sequences through the microorganism database, so as to obtain a second attribute feature matrix of microorganisms;

[0022] S303. Construct a microorganism-drug similarity feature network according to the similarity feature matrix of drugs and the similarity feature matrix of microorganisms;

[0023] S304. Construct a microorganism-drug second attribute feature network according to the second attribute feature matrix of drugs and the second attribute feature matrix of microorganisms;

[0024] S305. Combine the microbial-drug similarity feature network and the microbial-drug second attribute feature network to obtain a multi-modal attribute graph of the microbial-drug.

[0025] Further, in step S301, construct a similarity feature matrix of drugs according to the drug similarity attributes in the drug database, and construct a topological attribute network of drugs through the drug database, so as to obtain a second attribute feature matrix of drugs. The specific steps are as follows:

[0026] A1. Use the SIMCOMP2 tool to calculate the drug similarity attributes in the drug database to obtain the molecular structure similarity matrix DS struct (di, dj);

[0027] A2. Represent the drug-drug interaction spectrum in Net2 with the matrix DIP to obtain the normalized kernel bandwidth:

[0028]

[0029] where μ represents the normalized kernel bandwidth, μ′ is the original bandwidth, set to 1, DIP(d i ) represents the interaction of drug d i with other drugs, and nd represents the number of microorganisms in Net1;

[0030] A3. Represent the similarity feature matrix of drugs as S d (d i , d j ):

[0031]

[0032] A4. Construct the topological attributes of the drug network in the drug database by the random walk method with restart, perform random drift and restart on the drug network until the drug network converges, complete the construction of the drug network, so as to obtain the probability distribution vector of each drug, and construct the second attribute feature matrix F d ∈R nd×nd .

[0033] Furthermore, in step A4, the formula for random drift and restart is:

[0034]

[0035] where represents the probability that the i-th node of the drug network moves to other nodes at time t + 1, θ is the restart probability, T is the transition probability matrix, and p i (0) ∈R n×1The starting probability vector representing the i-th node of the drug network, p i (t) ∈R n×1 Represents the probability that the i-th node of the drug network moves to other nodes at time t.

[0036] Furthermore, in the step S302, the specific steps of constructing the similarity feature matrix of microorganisms according to the functional similarity attributes of microorganisms in the microorganism database and constructing the microorganism gene sequence through the microorganism database to obtain the second attribute feature matrix of microorganisms are as follows:

[0037] B1. Use the Kamneva tool to calculate the functional similarity attributes of microorganisms in the biological database to obtain the similarity feature matrix S of microorganisms m ∈R nm×nm , where nm represents the number of microorganisms in Net1; the similarity between microorganism m i and microorganism m j is represented as S m (m i , m j );

[0038] B2. Encode the original gene sequences of the microorganism data in the microorganism database to obtain the microorganism gene sequences;

[0039] B3. Pad all the encoded microorganism gene sequences with zeros so that the lengths of all the padded microorganism gene sequences are the same;

[0040] B4. Use the principal component analysis method to analyze all the padded microorganism gene sequences to obtain a k-dimensional matrix, and represent the second attribute feature matrix of microorganisms as F by the k-dimensional matrix m ∈R nm×k .

[0041] Furthermore, in the step S303, according to the similarity feature matrix of drugs and the similarity feature matrix of microorganisms, construct a microorganism-drug similarity feature network, and the specific steps are as follows:

[0042] C1. According to the similarity feature matrix of drugs and the similarity feature matrix of microorganisms, construct a microorganism-drug similarity feature network X simility :

[0043]

[0044] C2. According to the second attribute feature matrix of drugs and the second attribute feature matrix of microorganisms, construct a microorganism-drug second attribute feature network X secondary :

[0045]

[0046] C3. Combine the microbial-drug similarity feature network and the microbial-drug second-attribute feature network to obtain the multi-modal attribute graph X of microorganisms and drugs:

[0047] X = [X simility , X secondary .

[0048] Furthermore, in step S4, according to Net1, Net2, and the multi-modal attribute graph of microorganisms and drugs, a graph neural network model with regularization is established. The specific steps are as follows:

[0049] S401. Establish a feature matrix of microorganisms and drugs based on the multi-modal attribute graph of microorganisms and drugs, construct a heterogeneous matrix of microorganisms and drugs. In the heterogeneous matrix, vi represents any microorganism or drug node, and the heterogeneous matrix is expressed as:

[0050]

[0051] Among them, Y is the feature matrix of microorganisms and drugs, represents the content feature vector of node vi;

[0052] S402. Set a learnable matrix W ∈ R m ×f, and assign initial values to the elements of the learnable matrix using random numbers. Among them, f is the dimension of the node embedding representation, which is set by hyperparameters. n = nd + nm is the number of nodes, and m = 2×(nd + nm) is the feature dimension of the nodes. According to and W, generate a vector

[0053]

[0054] S403. Set a scaling constant s ∈ R, which represents the norm of the propagated hidden features, and generate a normalized feature transformation vector

[0055]

[0056] S404. Obtain the L2 regularization formula g():

[0057]

[0058] S405. Use the GNCN encoder to encode the association network of microorganisms and drugs and the multi-modal attribute graph of microorganisms and drugs:

[0059]

[0060] where \(A\in R\) nd×nm is the adjacency matrix of the association network described in step S1. If there is a known association between nodes \(i\) and \(j\) in the association network, the element \(A_{ij}\) in \(A\) is set to 1, otherwise it is 0; where \(I\) N is the \(N\)-order identity matrix, is the degree matrix of.

[0061] Furthermore, in the step S5, the multi-modal attribute graph of Net1 and Net2 combined with microorganisms and drugs is input into the graph neural network model to obtain the embedding representations \(Z1\) and \(Z2\); the specific steps of inputting the embedding representations \(Z1\) and \(Z2\) into the graph neural network for training are as follows:

[0062] S501. Propagate the normalized vector through the GNCN network to generate node embedding vectors

[0063]

[0064] where is the unit vector of \(i\) in the matrix, is the unit vector of \(j\) in the matrix, \(deg_i\) is the degree of node \(i\); \(deg_j\) is the degree of node \(j\);

[0065] S502. Generate a node embedding matrix according to the node embedding vectors to generate the latent variable \(Z\in R^{n\times f}\) of the GNCN encoder, and obtain \(Z1\) corresponding to Net1 and \(Z2\) corresponding to Net2:

[0066] \(Z_i = GNCN(X, A, s)\);

[0067] S503. Define a loss function, and the loss function is the binary cross-entropy between the multi-modal attribute graph and the reconstructed graph obtained from the trained graph neural network:

[0068]

[0069] where \(L\) is the loss function, \(N\) is the total number of all nodes, \(y\) represents the value of an element in the adjacency matrix \(A\), taking values of 0 or 1, represents the value of the corresponding element in the reconstructed adjacency matrix taking values between 0 and 1;

[0070] S504. Input \(Z1\) and \(Z2\) into the DNN classifier of the graph neural network model, set the number of training epochs \(epoch\) to \(k2\), and the training process uses stochastic gradient descent. Stop training when the loss function converges to obtain the trained graph neural network.

[0071] Furthermore, in step S5, after the graph neural network model is trained, the graph neural network model is also verified. The specific verification steps are as follows:

[0072] D1. Introduce a k-fold cross-validation framework. Under the k-fold cross-validation framework, all the known microorganism-drug association data in the existing microorganism-drug association database are randomly divided into k1 groups. Each subset of randomly sampled unknown association pairs with the same size batch in the k1 groups is selected as the test set, and the remaining known association pairs are selected as the training set;

[0073] D2. Input the test set into the trained graph neural network model to obtain the classification result;

[0074] D3. If the classification result is the positive class, it is predicted that there is an association between the microorganism and the drug. If the classification result is the negative class, it is predicted that there is no association between the microorganism and the drug;

[0075] D4. Obtain the AUC value of the trained graph neural network model according to the classification result; verify the accuracy of the graph neural network model according to the AUC value.

[0076] Furthermore, in step D4, to obtain the AUC value of the trained graph neural network model according to the classification result, the specific steps are as follows;

[0077] E1. Input the training set into the model to obtain the reconstructed graph of the training set by the current model. Take the scores of the edges between the nodes in the reconstructed graph of the training set by the current model, which is denoted as the association probability. The association probability takes values between 0 and 1;

[0078] E2. Use the association probability as the classification threshold. When other association probabilities are greater than this classification threshold, it is regarded as predicting a positive sample. When other association probabilities are less than this classification threshold, it is regarded as predicting a negative sample;

[0079] E3. According to the association relationship between microorganisms and drugs in the training set, obtain the true value of the edge label in the training set. The true value of the label takes values of 0 or 1, where 0 indicates that there is no edge, that is, there is no association relationship, which is actually a negative sample, and 1 indicates that there is an edge, that is, there is an association relationship, which is actually a positive sample;

[0080] E4. Statistically calculate the true positive rate and false positive rate under each classification threshold:

[0081]

[0082]

[0083] Among them, TPRate is the true positive rate, FPRate is the false positive rate, TP is the true positive rate, representing the number of samples that are actually negative samples but predicted as positive samples, FN is the false negative rate, representing the number of samples that are actually positive samples but predicted as negative samples, FP is the false positive rate, representing the number of samples that are actually negative samples but predicted as positive samples, and TN is the true negative rate, representing the number of samples that are actually negative samples and predicted as negative samples;

[0084] E5. Plot the ROC curve with FPRate as the horizontal axis and TPRate as the vertical axis, and use the infinitesimal method to calculate the area under the ROC curve, that is, the AUC value.

[0085] The beneficial effects of the present invention are as follows:

[0086] The present invention constructs an association network of microorganisms and drugs through a microorganism-drug association database, and further obtains interaction networks Net1 and Net2; establishes a graph neural network model introducing regularization, inputs Net1, Net2 and the multi-modal attribute graph of microorganisms and drugs into the graph neural network model to obtain embedding representations Z1 and Z2, and inputs the embedding representations Z1 and Z2 into the graph neural network for training to obtain a trained graph neural network; constructs interpretable node features of organisms and drugs, and considers the sparsity problem brought by the existing microorganism-drug association data set. BRIEF DESCRIPTION OF THE DRAWINGS

[0087] Figure 1 is a schematic flow chart of a method for predicting the association effect of microorganisms and drugs according to the present invention.

[0088] Figure 2 is a schematic flow chart of constructing a multi-modal attribute graph of a method for predicting the association effect of microorganisms and drugs according to the present invention.

[0089] Figure 3 is a schematic flow chart of obtaining the association probability of a method for predicting the association effect of microorganisms and drugs according to the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0090] The present invention will be described in detail below with reference to the drawings and specific embodiments.

[0091] Example 1

[0092] As Figure 1 shown, a method for predicting the association effect of microorganisms and drugs includes the following steps:

[0093] S1. Construct an association network of microorganisms and drugs through a microorganism-drug association database, and call the association network Net1;

[0094] S2. Retrieve the relevant interactions between microorganisms - microorganisms through the microorganism database in the microorganism - drug association database, and retrieve the relevant interactions between drugs - drugs through the drug database in the microorganism - drug association database; construct an interaction network based on the relevant interactions between microorganisms - microorganisms and drugs - drugs, and name the interaction network Net2;

[0095] S3. Construct the topological property network of drugs through the drug database, and construct the microbial gene sequences through the microorganism database; construct the multi - modal property graph of microorganisms - drugs based on the comprehensive similarity properties of drugs in the drug database, the topological properties of the drug network, the functional similarity properties of microorganisms in the microorganism database, and the genomic sequence properties;

[0096] S4. Establish a graph neural network model with regularization introduced based on Net1, Net2, and the multi - modal property graph of microorganisms - drugs;

[0097] S5. Input Net1, Net2 combined with the multi - modal property graph of microorganisms - drugs into the graph neural network model to obtain the embedding representations Z1 and Z2; input the embedding representations Z1 and Z2 into the graph neural network for training to obtain the trained graph neural network;

[0098] S6. Obtain the dataset to be predicted, and predict the associated effects between microorganisms - drugs in the dataset to be predicted through the trained graph neural network.

[0099] The present invention constructs the association network of microorganisms - drugs through the microorganism - drug association database, and further obtains the interaction networks Net1 and Net2; establishes a graph neural network model with regularization introduced, inputs Net1, Net2 combined with the multi - modal property graph of microorganisms - drugs into the graph neural network model to obtain the embedding representations Z1 and Z2, and inputs the embedding representations Z1 and Z2 into the graph neural network for training to obtain the trained graph neural network; constructs interpretable node features of organisms and drugs, and considers the sparsity problem brought by the existing microorganism - drug association datasets.

[0100] Example 2

[0101] Specifically, as Figure 2 shown, in a specific embodiment, in step S3, the specific steps of constructing the topological property network of drugs through the drug database, constructing the microbial gene sequences through the microorganism database, and constructing the multi - modal property graph of microorganisms - drugs based on the comprehensive similarity properties of drugs in the drug database, the topological properties of the drug network, the functional similarity properties of microorganisms in the microorganism database, and the genomic sequence properties are as follows:

[0102] S301. Construct a similarity feature matrix of drugs according to the drug similarity attributes in the drug database, and construct a topological attribute network of drugs through the drug database, so as to obtain the second attribute feature matrix of drugs;

[0103] S302. Construct a similarity feature matrix of microorganisms according to the functional similarity attributes of microorganisms in the microorganism database, and construct the microbial gene sequence through the microorganism database, so as to obtain the second attribute feature matrix of microorganisms;

[0104] S303. Construct a microorganism-drug similarity feature network according to the similarity feature matrix of drugs and the similarity feature matrix of microorganisms;

[0105] S304. Construct a microorganism-drug second attribute feature network according to the second attribute feature matrix of drugs and the second attribute feature matrix of microorganisms;

[0106] S305. Combine the microorganism-drug similarity feature network and the microorganism-drug second attribute feature network to obtain a multi-modal attribute graph of microorganisms-drugs.

[0107] In a specific embodiment, in step S301, constructing a similarity feature matrix of drugs according to the drug similarity attributes in the drug database, and constructing a topological attribute network of drugs through the drug database, so as to obtain the second attribute feature matrix of drugs, the specific steps are as follows:

[0108] A1. Use the SIMCOMP2 tool to calculate the drug similarity attributes in the drug database to obtain the molecular structure similarity matrix DS of drugs struct (di, dj);

[0109] A2. Use the matrix DIP to represent the drug-drug interaction spectrum in Net2 to obtain the standardized kernel bandwidth:

[0110]

[0111] where μ represents the standardized kernel bandwidth, μ′ is the original bandwidth, set to 1, DIP(d i ) represents the interaction of drug d i with other drugs, and nd represents the number of microorganisms in the Net1;

[0112] A3. Represent the similarity feature matrix of drugs as S d (d i , d j ):

[0113]

[0114] A4. Construct the topological properties of the drug network in the drug database by the random walk method with restart, perform random drift and restart on the drug network until the drug network converges, complete the construction of the drug network, thereby obtaining the probability distribution vector of each drug, and construct the second attribute feature matrix F of the drug d ∈R nd×nd 。

[0115] In a specific embodiment, in the step A4, the formula for random drift and restart is as follows:

[0116]

[0117] where, represents the probability that the i-th node of the drug network moves to other nodes at time t + 1, θ is the restart probability, T is the transition probability matrix, p i (0) ∈R n×1 represents the starting probability vector of the i-th node of the drug network, p i (t) ∈R n×1 represents the probability that the i-th node of the drug network moves to other nodes at time t.

[0118] In a specific embodiment, in the step S302, the specific steps for constructing the similarity feature matrix of microorganisms according to the functional similarity attributes of microorganisms in the microorganism database and obtaining the second attribute feature matrix of microorganisms by constructing the microorganism gene sequence through the microorganism database are as follows:

[0119] B1. Use the Kamneva tool to calculate the functional similarity attributes of microorganisms in the biological database to obtain the similarity feature matrix S of microorganisms m ∈R nm×nm , where nm represents the number of microorganisms in Net1; represent the similarity between microorganism m i and microorganism m j as S m (m i , m j );

[0120] B2. Encode the original gene sequences of microorganism data in the microorganism database to obtain microorganism gene sequences;

[0121] B3. Pad all the encoded microorganism gene sequences with zeros so that all the padded microorganism gene sequences have the same length;

[0122] B4. Use the principal component analysis method to analyze all the padded microorganism gene sequences to obtain a k-dimensional matrix, and represent the second attribute feature matrix of microorganisms as F by the k-dimensional matrixm ∈R nm×k 。

[0123] In a specific embodiment, in step S303, according to the similarity feature matrix of drugs and the similarity feature matrix of microorganisms, a microorganism-drug similarity feature network is constructed. The specific steps are as follows:

[0124] C1. Construct a microorganism-drug similarity feature network X according to the similarity feature matrix of drugs and the similarity feature matrix of microorganisms simility :

[0125]

[0126] C2. Construct a microorganism-drug second attribute feature network X according to the second attribute feature matrix of drugs and the second attribute feature matrix of microorganisms secondary :

[0127]

[0128] C3. Combine the microorganism-drug similarity feature network and the microorganism-drug second attribute feature network to obtain a multi-modal attribute graph X of microorganisms and drugs:

[0129] X = [X simility , X secondary .

[0130] In a specific embodiment, in step S4, according to Net1, Net2, and the multi-modal attribute graph of microorganisms and drugs, a graph neural network model with regularization is established. The specific steps are as follows:

[0131] S401. Establish a microorganism-drug feature matrix according to the multi-modal attribute graph of microorganisms and drugs, and construct a heterogeneous matrix of microorganisms and drugs. In the heterogeneous matrix, vi represents any microorganism or drug node, and the heterogeneous matrix is expressed as:

[0132]

[0133] where Y is the microorganism-drug feature matrix, represents the content feature vector of node vi;

[0134] S402. Set a learnable matrix W ∈ R m ×f, and assign initial values to the elements of the learnable matrix using random numbers. Where f is the dimension of the node embedding representation, which is set by hyperparameters, n = nd + nm is the number of nodes, and m = 2×(nd + nm) is the feature dimension of the nodes. According to and W, generate a vector

[0135]

[0136] S403. Set the scaling constant s ∈ R, which represents the norm of the propagated hidden features, and generate a normalized feature transformation vector by the GNCN network of the regularized graph neural network model

[0137]

[0138] S404. Obtain the L2 regularization formula g():

[0139]

[0140] S405. Use the GNCN encoder to encode the microorganism-drug association network and the microorganism-drug multimodal attribute graph

[0141]

[0142] where A ∈ R nd×nm is the adjacency matrix of the association network described in step S1. If there is a known association between nodes i and j in the association network, the element Aij in A is set to 1, otherwise 0; where I N is the N-order identity matrix, is the degree matrix of.

[0143] In a specific embodiment, as Figure 3 shown, in step S5, input Net1 and Net2 combined with the microorganism-drug multimodal attribute graph into the graph neural network model to obtain the embedding representations Z1 and Z2; the specific steps of inputting the embedding representations Z1 and Z2 into the graph neural network for training are as follows

[0144] S501. Propagate the normalized vector through the GNCN network to generate node embedding vectors

[0145]

[0146] where is the unit vector of i in the matrix, is the unit vector of j in the matrix, degi is the degree of node i; degj is the degree of node j;

[0147] S502. Generate a node embedding matrix according to the node embedding vectors to generate the latent variable Z ∈ Rn×f of the GNCN encoder, and obtain Z1 corresponding to Net1 and Z2 corresponding to Net2

[0148] Zi = GNCN(X, A, s);

[0149] S503. Define a loss function, which is the binary cross - entropy between the multi - modal attribute graph and the reconstructed graph obtained by training the graph neural network:

[0150]

[0151] where L is the loss function, N is the total number of all nodes, y represents the value of an element in the adjacency matrix A, taking values of 0 or 1, represents the value of the corresponding element in the reconstructed adjacency matrix and its value ranges from 0 to 1;

[0152] S504. Input Z1 and Z2 into the DNN classifier of the graph neural network model, set the number of training epochs to k2, and use stochastic gradient descent during the training process. Stop training when the loss function converges to obtain the trained graph neural network.

[0153] Example 3

[0154] In a specific embodiment, in step S5, after training the graph neural network model, the graph neural network model is further verified. The specific verification steps are as follows:

[0155] D1. Introduce the k - fold cross - validation framework. Under the k - fold cross - validation framework, randomly divide all the known microorganism - drug association data in the existing microorganism - drug association database into k1 groups, select each subset of randomly sampled unknown association pairs with the same batch size in the k1 groups as the test set, and select the remaining known association pairs as the training set;

[0156] D2. Input the test set into the trained graph neural network model to obtain the classification result;

[0157] D3. If the classification result is the positive class, predict that there is an association between the microorganism and the drug; if the classification result is the negative class, predict that there is no association between the microorganism and the drug;

[0158] D4. Obtain the AUC value of the trained graph neural network model according to the classification result; verify the accuracy of the graph neural network model according to the AUC value.

[0159] In a specific embodiment, in step D4, to obtain the AUC value of the trained graph neural network model according to the classification result, the specific steps are as follows;

[0160] E1. Input the training set into the model to obtain the reconstructed graph of the training set by the current model. Take the scores of the edges between the nodes in the reconstructed graph of the training set by the current model, denoted as the association probability, and the association probability ranges from 0 to 1.

[0161] E2. Use the association probability as the classification threshold. When other association probabilities are greater than this classification threshold, it is regarded as predicting a positive sample. When other association probabilities are less than this classification threshold, it is regarded as predicting a negative sample.

[0162] E3. According to the association relationship between microorganisms and drugs in the training set, obtain the true value of the label of the edge in the training set. The true value of the label takes 0 or 1. Among them, 0 means there is no edge, that is, there is no association relationship, which is actually a negative sample, and 1 means there is an edge, that is, there is an association relationship, which is actually a positive sample.

[0163] E4. Calculate the true positive rate and false positive rate under each classification threshold:

[0164]

[0165]

[0166] Among them, TPRate is the true positive rate, FPRate is the false positive rate, TP is the true positive rate, representing the number of samples that are actually negative samples but predicted as positive samples, FN is the false negative rate, representing the number of samples that are actually positive samples but predicted as negative samples, FP is the false positive rate, representing the number of samples that are actually negative samples but predicted as positive samples, and TN is the true negative rate, representing the number of samples that are actually negative samples and predicted as negative samples.

[0167] E5. Draw the ROC curve with FPRate as the horizontal axis and TPRate as the vertical axis, and use the infinitesimal method to calculate the area of the ROC curve, that is, the AUC value.

[0168] In this embodiment, in order to verify a method for predicting the association effect of microorganisms and drugs of the present invention, this embodiment uses the default parameter settings, runs this method and 5 existing methods on the MDAD dataset, and uses the AUC value as the performance evaluation index. The larger the AUC value, the higher the accuracy of the method.

[0169] In this embodiment, 5-fold cross-validation and 10-fold cross-validation are performed on all methods including the present invention. The experimentally verified drug combinations are randomly divided into 5 or 10 subsets of the same size. Each subset is used as the test set in turn, and the remaining are used to train the model. In order to eliminate the random sampling bias, this process is repeated 10 times, and the final AUC score is calculated according to the average value of the AUC values in the 10 repeated validations, and the accuracy of each method is evaluated with the final AUC score as the performance index.

[0170] The verification results are shown in the following table. In the 5-fold cross-validation, the method adopted in the present invention is denoted as G2GNAEMDA, and the final AUC score of the present invention is the highest among all methods. In the 10-fold cross-validation, the final AUC score of the present invention is also the highest among all methods. Therefore, the accuracy of the present invention is superior to that of the existing methods:

[0171]

[0172] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.

Claims

1. A method for predicting the association effect of microorganisms and drugs, characterized in that: It includes the following steps: S1. Construct an association network of microorganisms and drugs through a microorganism-drug association database, and call the association network Net1; S2. Retrieve the relevant interactions between microorganisms through the microorganism database in the microorganism-drug association database, and retrieve the relevant interactions between drugs through the drug database in the microorganism-drug association database; construct an interaction network according to the relevant interactions between microorganisms and the relevant interactions between drugs, and call the interaction network Net2; S3. Construct a topological attribute network of drugs through the drug database, and construct a microbial gene sequence through the microorganism database; construct a multi-modal attribute graph of microorganisms and drugs according to the comprehensive similarity attributes of drugs in the drug database, the topological attributes of the drug network, the functional similarity attributes of microorganisms in the microorganism database, and the genomic sequence attributes; In the step S3, the specific steps of constructing a topological attribute network of drugs through the drug database, constructing a microbial gene sequence through the microorganism database, and constructing a multi-modal attribute graph of microorganisms and drugs according to the comprehensive similarity attributes of drugs in the drug database, the topological attributes of the drug network, the functional similarity attributes of microorganisms in the microorganism database, and the genomic sequence attributes are as follows: S301. Construct a similarity feature matrix of drugs according to the drug similarity attributes in the drug database, and construct a topological attribute network of drugs through the drug database, so as to obtain a second attribute feature matrix of drugs; S302. Construct a similarity feature matrix of microorganisms according to the functional similarity attributes of microorganisms in the microorganism database, and construct a microbial gene sequence through the microorganism database, so as to obtain a second attribute feature matrix of microorganisms; S303. Construct a microorganism-drug similarity feature network according to the similarity feature matrix of drugs and the similarity feature matrix of microorganisms; S304. Construct a microorganism-drug second attribute feature network according to the second attribute feature matrix of drugs and the second attribute feature matrix of microorganisms; S305. Combine the microorganism-drug similarity feature network and the microorganism-drug second attribute feature network to obtain a multi-modal attribute graph of microorganisms and drugs S4. Establish a graph neural network model with regularization introduced according to Net1, Net2, and the multi-modal attribute graph of microorganisms and drugs; S5. Input Net1, Net2 combined with the multi-modal attribute graph of microorganisms and drugs into the graph neural network model to obtain embedding representations Z1 and Z2; input the embedding representations Z1 and Z2 into the graph neural network for training to obtain a trained graph neural network; S6. Predict the association effect of microorganisms and drugs in the microorganism-drug dataset through the trained graph neural network.

2. The method for predicting the association effect of microorganisms and drugs according to claim 1, characterized in that: In the step S301, a similarity feature matrix of drugs is constructed according to the drug similarity attributes in the drug database, and a topological attribute network of drugs is constructed through the drug database, so as to obtain a second attribute feature matrix of drugs. The specific steps are as follows: A1. Use the SIMCOMP2 tool to calculate the drug similarity attributes in the drug database and obtain the molecular structure similarity matrix DS of the drugs struct (di, dj); A2. Use the matrix DIP to represent the drug-drug interaction spectrum in Net2 to obtain the standardized kernel bandwidth: Among them, μ represents the standardized kernel bandwidth, μ′ is the original bandwidth, set to 1, DIP(d i ) represents the interaction of drug d i with other drugs, and nd represents the number of microorganisms in the said Net1; A3. Represent the similarity feature matrix of the drug as S d (d i ,d j ): A4. Construct the topological properties of the drug network in the drug database by the random walk method with restart, perform random drift and restart on the drug network until the drug network converges, complete the construction of the drug network, so as to obtain the probability distribution vector of each drug, and construct the second attribute feature matrix F of the drug d ∈R nd×nd 。 3. According to the method for predicting the microbial-drug association effect described in claim 2, It is characterized in that: In the step A4, the formula for random drift and restart is: Among them, represents the probability that the \(i\)-th node of the drug network moves to other nodes at time \(t + 1\), \(\theta\) is the restart probability, \(T\) is the transition probability matrix, \(p\) i (0) \(\in R\) n×1 represents the initial probability vector of the \(i\)-th node of the drug network, \(p\) i (t) \(\in R\) n×1 represents the probability that the \(i\)-th node of the drug network moves to other nodes at time \(t\).

4. According to the method for predicting the microbial-drug association effect described in claim 3, It is characterized in that: In the step S302, a similarity feature matrix of microorganisms is constructed according to the functional similarity attributes of microorganisms in the microorganism database, and a microorganism gene sequence is constructed through the microorganism database, so as to obtain the specific steps of the second attribute feature matrix of microorganisms are as follows: B1. Calculate the functional similarity attributes of microorganisms in the biological database using the Kamneva tool to obtain the similarity feature matrix S of microorganisms m ∈R nm×nm , where nm represents the number of microorganisms in Net1; represent the similarity between microorganism m i and microorganism m j as S m (m i , m j ); B2. Encode the original gene sequences of the microorganism data in the microorganism database to obtain microorganism gene sequences; B3. Pad all the encoded microorganism gene sequences with zeros so that the lengths of all the padded microorganism gene sequences are the same; B4. Use the principal component analysis method to analyze all the filled microbial gene sequences to obtain a k-dimensional matrix, and represent the second attribute feature matrix of the microorganism as F by the k-dimensional matrix m ∈R nm×k .

5. According to the method for predicting the microbial-drug association effect described in claim 4, It is characterized in that: In the step S303, according to the similarity feature matrix of drugs and the similarity feature matrix of microorganisms, a microbial-drug similarity feature network is constructed. The specific steps are as follows: C1. Construct a microorganism-drug similarity feature network X based on the similarity feature matrix of drugs and the similarity feature matrix of microorganisms simility : Construct a second-attribute feature network X of microorganism-drug based on the second-attribute feature matrix of the drug and the second-attribute feature matrix of the microorganism secondary : C3. Combine the microbial-drug similarity feature network and the microbial-drug second attribute feature network to obtain a multi-modal attribute graph X of microorganisms and drugs: X = [X simility , X secondary 。 6. According to the method for predicting the microbial-drug association effect described in claim 5, It is characterized in that: In the step S4, according to Net1, Net2, and the multi-modal attribute graph of microorganisms and drugs, a graph neural network model with regularization is established. The specific steps are as follows: S401. Establish a feature matrix of microorganisms and drugs according to the multi-modal attribute graph of microorganisms and drugs, and construct a heterogeneous matrix of microorganisms and drugs. In the heterogeneous matrix, vi represents any microorganism or drug of a node, and the heterogeneous matrix is expressed as: Among them, Y is the characteristic matrix of microorganism-drug, represents the content feature vector of node vi; S402. Set the learnable matrix \(W\in\mathbb{R}\) m \(^{n\times f}\), and initialize the elements of the learnable matrix with random numbers, where \(f\) is the dimension of the node embedding representation, set by hyperparameters, \(n = n_d + n_m\) is the number of nodes, \(m = 2\times(n_d + n_m)\) is the feature dimension of the nodes. According to and \(W\), generate the vector after feature transformation S403. Set a scaling constant s ∈ R, which represents the norm of the propagated hidden features and generates a normalized feature transformation vector by the GNCN network of the regularized graph neural network model S404. Obtain the formula g() of L2 regularization: S405. Use the GNCN encoder to encode the association network of microorganisms and drugs and the multi-modal attribute graph of microorganisms and drugs: where \(A\in R\) nd×nm is the adjacency matrix of the association network described in step S1. If there is a known association between nodes \(i\) and \(j\) in the association network, the element \(A_{ij}\) in \(A\) is set to 1, otherwise it is 0; where \(I\) N is the \(N\)-order identity matrix, and is the degree matrix of 7. According to the method for predicting the microbial-drug association effect described in claim 6, It is characterized in that: In the step S5, input Net1, Net2, and the multi-modal attribute graph of microorganisms and drugs into the graph neural network model to obtain the embedding representations Z1 and Z2; input the embedding representations Z1 and Z2 into the graph neural network for training to obtain the trained graph neural network. The specific steps are as follows: S501. Propagate the normalized vector through the GNCN network to generate node embedding vectors Among them, is the unit vector of i in the matrix, is the unit vector of j in the matrix, degii is the degree of node i; degj is the degree of node j; S502. Generate a node embedding matrix based on the node embedding vectors Generate a latent variable Z ∈ Rn×f of the GNCN encoder, and obtain Z1 corresponding to Net1 and Z2 corresponding to Net2: Z = GNCN(X, A, s); S503. Define a loss function, and the loss function is the binary cross entropy between the multi-modal attribute graph and the reconstructed graph obtained from the trained graph neural network: Among them, L is the loss function, N is the total number of all nodes, y represents the value of an element in the adjacency matrix A, taking values of 0 or 1, representing the value of the corresponding element in the reconstructed adjacency matrix, with values ranging from 0 to 1; S504. Input Z1 and Z2 into the DNN classifier of the graph neural network model. Set the number of training epochs to k2. The training process uses stochastic gradient descent and stops training when the loss function converges to obtain a trained graph neural network.

8. The method for predicting the microbial-drug association effect according to claim 7, wherein: in step S5, after training the graph neural network model, the accuracy of the graph neural network model is further verified. The specific verification steps are as follows: D1. Introduce a k-fold cross-validation framework. Under the k-fold cross-validation framework, randomly divide all known microbial-drug association data in the existing microbial-drug association database into k1 groups. Select each subset of randomly sampled unknown association pairs with the same batch size in the k1 groups as the test set, and select the remaining known association pairs as the training set; D2. Input the test set into the trained graph neural network model to obtain the classification result; D3. If the classification result is a positive class, it is predicted that there is an association between the microorganism and the drug. If the classification result is a negative class, it is predicted that there is no association between the microorganism and the drug; D4. Obtain the AUC value of the trained graph neural network model according to the classification result; verify the accuracy of the graph neural network model according to the AUC value.

9. The method for predicting the microbial-drug association effect according to claim 8, wherein: in step D4, the specific steps for obtaining the AUC value of the trained graph neural network model according to the classification result are as follows; E1. Input the training set into the model to obtain the reconstructed graph of the training set by the current model. Take the scores of the edges between the nodes in the reconstructed graph of the training set by the current model, which are recorded as the association probabilities. The association probabilities take values between 0 and 1; E2. Use the association probability as the classification threshold. When other association probabilities are greater than this classification threshold, it is regarded as predicting a positive sample. When other association probabilities are less than this classification threshold, it is regarded as predicting a negative sample; E3. According to the association relationship between microorganisms and drugs in the training set, obtain the true label of the edges in the training set. The true label takes values of 0 or 1, where 0 indicates that there is no edge, that is, there is no association relationship, that is, it is actually a negative sample, and 1 indicates that there is an edge, that is, there is an association relationship, that is, it is actually a positive sample; E4. Statistically calculate the true positive rate and false positive rate at each classification threshold: where TPRate is the true positive rate, FPRate is the false positive rate, TP is the true positive rate, which represents the number of samples that are actually negative samples but predicted as positive samples, FN is the false negative rate, which represents the number of samples that are actually positive samples but predicted as negative samples, FP is the false positive rate, which represents the number of samples that are actually negative samples but predicted as positive samples, and TN is the true negative rate, which represents the number of samples that are actually negative samples but predicted as negative samples; E5. Draw an ROC curve with FPRate as the horizontal axis and TPRate as the vertical axis, and use the infinitesimal method to calculate the area under the ROC curve, that is, the AUC value.

Citation Information

Patent Citations

  • Drug target prediction method based on multi-source data fusion and network structure disturbance

    CN112420126A

  • Efficient prediction method for association of miRNA and drug resistance

    CN114550814A