Drug-disease association prediction method based on graph network and auto-encoder

Through the method of combining graph network and autoencoder, the spatial and topological information of drugs and diseases are extracted, and the accuracy and robustness of drug-disease association prediction in the prior art are solved, and more efficient drug-disease association prediction is achieved.

CN120299747APending Publication Date: 2025-07-11ANHUI UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510350227.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Existing drug-disease association prediction methods have deteriorated performance when feature extraction dependence and data are missing, making it difficult to accurately and efficiently predict the association between drugs and diseases.

Method used

Using a combination of graph network and autoencoder, the characteristics of drugs and diseases in the feature space and topological space are extracted through graph attention network, and the autoencoder is used for feature fusion and deep extraction to construct a drug-disease association prediction model.

Benefits of technology

It improves the accuracy of drug-disease association prediction and model interpretability, reduces laboratory research costs, and is suitable for robust prediction of multi-source data models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299747A_ABST
    Figure CN120299747A_ABST
Patent Text Reader

Abstract

The invention discloses a drug-disease association prediction method based on a graph network and an auto-encoder, and belongs to the technical field of drug-disease association prediction, and the method comprises the following steps: S1, preprocessing drug-disease data; s2, extracting information of medicines and diseases in a topological space and a feature space; s3, deep extraction of drug and disease feature expression and fusion; s4, model training; and S5, correlation prediction. According to the method, the graph attention network and the graph encoder are adopted to extract information of drugs and diseases in a feature space and a topological space, feature expressions of the drugs and the diseases are enriched, and meanwhile, the feature expressions of the drugs and the diseases are reconstructed by adopting the auto-encoder, so that the features of the drugs and the diseases are more representative; and the accuracy of a drug-disease association prediction result is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of drug-disease association prediction, and in particular to a drug-disease association prediction method based on a graph network and an autoencoder. Background Art

[0002] The purpose of drug-disease association prediction is to discover potential drug uses and pre-evaluate possible adverse effects of drugs, etc., so as to assist drug repositioning. Researchers use multi-omics data such as genomics and proteomics to analyze the relationship between drug action targets and disease-related biomolecules to infer whether there is an association between a certain drug and a specific disease, and the nature of the association (such as whether the drug can treat the disease, whether the drug will cause the disease, etc.), or use machine learning algorithms such as support vector machines and random forest algorithms, and use known drug-disease association data, drug characteristics (such as chemical structure characteristics, pharmacological properties, etc.) and disease characteristics (such as symptom manifestations, pathological characteristics, etc.) to predict whether there is an association for unknown drug-disease combinations; or use the association network between drugs and diseases and the similarity between drugs and diseases to predict possible associations under similar circumstances.

[0003] In recent years, there have been mainly three solutions for the drug-disease association prediction task based on computational methods: machine learning methods, graph deep learning-based methods, and matrix factorization and completion methods. Specifically, most machine learning methods are data-driven. They usually generate potential features from known drug-disease interaction data, and then use various machine learning techniques to predict the potential indications of a given drug, and identify unknown drug-disease associations based on drug similarity and disease similarity. Although machine learning methods are accepted due to their high-quality prediction results, this method is too dependent on the feature extraction of drugs and diseases and the selection of samples. With the development of technology and the continuous update of databases, more and more biological information has been added to the prediction task, such as proteins, drug side effects, and genes. Therefore, graph deep learning-based methods have gradually come into the public eye. By integrating heterogeneous information networks through drug-disease, drug-protein, protein-disease associations and their biological knowledge, different learning strategies are adopted on the constructed semantic graph and functional similarity graph to obtain the feature representations of drugs and diseases, and different features are learned from topological and biological perspectives. Although graph deep learning-based methods have better interpretability than machine learning methods, their performance is sometimes unsatisfactory. Due to the flexibility of matrix factorization and completion methods in integrating prior knowledge and their good performance, they are widely popular in drug-disease association prediction. In matrix factorization methods, there is also a method based on the recommendation system. This method regards the identification of potential drug indications as a recommendation task and uses matrix factorization methods for the experimental task objectives. Although these methods are effective, they are not suitable for accurate prediction of new drugs or diseases.

[0004] In summary, drug-disease association prediction can be achieved using these three methods: machine learning, graph deep learning-based, and matrix factorization and completion. They utilize the correlations between different biological information or biological macromolecules to obtain the information required by the corresponding methods to complete the prediction task. These methods provide new ideas for drug research and development, reduce the research and development costs, accelerate the development of new drugs, and improve the success rate of new drugs. Although the above methods have good performance in drug-disease association prediction, there are still some defects. They ignore the characteristics of drugs and diseases themselves and the dependence of the topological structure between the two on the prediction task. Secondly, for multi-source data models, when data is missing, the model performance will decline rapidly, and the model is too dependent on data sources. For this reason, a drug-disease association prediction method based on graph network and autoencoder is proposed. Summary of the Invention

[0005] The technical problem to be solved by the present invention is: how to solve the key features representing drugs and diseases, and after obtaining the representative features, be able to accurately and efficiently predict the association between drugs and diseases, and achieve automatic prediction through machines, reducing the cost of laboratory research, and providing a drug-disease association prediction method based on graph network and autoencoder.

[0006] The present invention solves the above technical problem through the following technical solutions. The present invention includes the following steps:

[0007] S1: Preprocessing of drug-disease data

[0008] Calculate the similarity between drugs and between diseases in the drug-disease dataset respectively, use the similar features of drugs and diseases to construct drug and disease feature graphs, and divide them into a training set and a test set according to a set ratio;

[0009] S2: Information extraction of drugs and diseases in topological space and feature space

[0010] In the graph convolution module, use the graph attention network to extract the features of drugs and diseases in the feature space, and at the same time use the graph encoder to extract the features of drugs and diseases in the topological space;

[0011] S3: Deep extraction of drug and disease feature expressions and fusion

[0012] In the autoencoder module, use the autoencoder to further extract the features of drugs and diseases in the feature space and topological space, and fuse the extracted features;

[0013] S4: Model training

[0014] Use the training set to train the graph neural network to obtain a drug-disease association prediction model, where the graph neural network includes a graph convolution module, an autoencoder module and a prediction module;

[0015] S5: Association prediction

[0016] Send the test set data into the drug-disease association prediction model for prediction to obtain the prediction result.

[0017] Furthermore, in the step S1, calculate the similarity between drugs through the two-dimensional chemical fingerprints corresponding to the drugs, and then obtain the drug similarity matrix. Construct a K-nearest neighbor graph as the drug feature graph according to the drug similarity matrix, so as to obtain the matrix expression of drug similarity. At the same time, according to the rule that the elements connected in the K-nearest neighbor graph are set to 1 and the non-adjacent elements are set to 0, obtain the corresponding adjacency matrix.

[0018] Further, in the step S1, the similarity between diseases is obtained by calculating the contribution of each disease to another disease in the disease graph, and then the disease similarity matrix is obtained. A K-nearest neighbor graph is constructed as the disease feature graph based on the disease similarity matrix, so as to obtain the matrix expression of disease similarity. At the same time, the corresponding adjacency matrix is obtained according to the rule that the connected elements in the K-nearest neighbor graph are set to 1 and the non-adjacent elements are set to 0.

[0019] Further, in the step S2, the specific processing process is as follows:

[0020] S21: Use the graph attention network to extract the information of drugs and diseases in the feature space from the drug feature graph and the disease feature graph, that is, obtain the new drug feature representation and the new disease feature representation;

[0021] S22: Use the graph encoder to extract the information of drugs and diseases in the topological space from the drug-disease association graph, that is, obtain the drug topological structure representation and the disease topological structure representation. Among them, the drug-disease association graph is a drug-disease interaction graph, and the drug-disease interaction graph is composed of drugs and diseases as nodes and the relationships between drugs and diseases as edges.

[0022] Further, in the step S21, the specific process of obtaining the new drug feature representation is as follows:

[0023] After obtaining the drug similarity matrix and the adjacency matrix, calculate the attention weights between drugs by using the graph attention network, and perform weighted summation on the features of adjacent drugs according to the attention weights to obtain the new drug feature representation.

[0024] Further, in the step S21, the specific process of obtaining the new disease feature representation is as follows:

[0025] After obtaining the disease similarity matrix and the adjacency matrix, calculate the attention weights between diseases by using the graph attention mechanism, and perform weighted summation on the features of adjacent diseases according to the attention weights to obtain the new disease feature representation.

[0026] Further, in the step S22, the extraction process of the drug topological structure feature representation is as follows:

[0027] S2201: In the drug-disease association graph, calculate the information transmitted from disease j to drug i. The formula is as follows:

[0028]

[0029] where c ij is the symmetric normalization constant from disease j to drug i, r i is the node connected to drug i, d jis a node connected to disease j, N(r i ) and N(d j ) respectively represent the sets of drug and disease neighborhood nodes, and W is the parameter matrix of edge types;

[0030] S2202: After obtaining the information of each edge, the information passed into each node can be obtained by cumulatively summing the edges connected to each node:

[0031] h i = σ[ΣMP(μ j→i )]

[0032] where σ is the activation function;

[0033] S2203: Obtain the drug topological structure feature representation through a linear layer:

[0034] z i = W dr h i

[0035] where W dr is the weight matrix of the linear layer.

[0036] Furthermore, in the step S22, the extraction process of the disease topological structure feature representation is as follows:

[0037] S2211: In the drug-disease association graph, calculate the information transmitted from drug m to disease n, and the formula is as follows:

[0038]

[0039] where c nm is the symmetric normalization constant from drug m to disease n, r n is the node connected to disease n, d m is the node connected to drug m, N(r n ) and N(d m ) respectively represent the sets of drug and disease neighborhood nodes, and W is the parameter matrix of edge types;

[0040] S2212: After obtaining the information of each edge, the information passed into each node can be obtained by cumulatively summing the edges connected to each node:

[0041] h n = σ[ΣMP(μ m→n )]

[0042] where σ is the activation function;

[0043] S2213: Obtain the drug topological structure feature representation through a linear layer:

[0044] z n = W di h n

[0045] where W di is the weight matrix of the linear layer.

[0046] Furthermore, in the step S3, the fusion process is as follows:

[0047] S31: After obtaining the features of the drug and the disease in the feature space and the topological space, use an autoencoder to further extract the features of the drug and the disease in the feature space and the topological space to obtain the corresponding feature vectors;

[0048] S32: Concatenate the obtained feature vectors as the final feature representation of the drug and the disease.

[0049] Furthermore, in the step S4, the prediction module is a multi-layer perceptron. During prediction, the final feature representation of the drug and the disease obtained in step S32 is concatenated again and sent into the multi-layer perceptron for prediction.

[0050] The present invention has the following advantages compared with the prior art: The drug-disease association prediction method based on a graph network and an autoencoder extracts the information of drugs and diseases in the feature space and the topological space by using a graph attention network and a graph encoder, enriches the feature representation of drugs and diseases, and at the same time reconstructs the feature representation of drugs and diseases by using an autoencoder, making the drug and disease features more representative, thereby effectively improving the accuracy of the drug-disease association prediction result. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 is a schematic flowchart of the drug-disease association prediction method based on a graph network and an autoencoder in an embodiment of the present invention;

[0052] Figure 2 is a schematic diagram of the attention mechanism of the graph attention network in an embodiment of the present invention, where (a) is single-head attention and (b) is multi-head attention;

[0053] Figure 3 is a schematic structural diagram of the autoencoder in an embodiment of the present invention;

[0054] Figure 4 is a schematic diagram of the drug-disease association prediction model in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] The embodiments of the present invention will be described in detail below. The following embodiments are implemented on the premise of the technical solution of the present invention, and detailed implementation manners and specific operation processes are given. However, the protection scope of the present invention is not limited to the following embodiments.

[0056] As Figures 1 to 4 shown, this embodiment provides a technical solution: a drug-disease association prediction method based on graph network and autoencoder, including the following steps:

[0057] S1: Preprocess the drug-disease data.

[0058] In step S1, it includes the following two sub-steps:

[0059] S11: Randomly divide each drug-disease association dataset (Gdataset, Cdataset, Ldataset, LRSSL). The training set and the test set are 10 subsets of equal size and non-correlated. Each subset is regarded as the test set in turn, and the remaining nine subsets are used as the training set.

[0060] S12: Calculate the similarity between drugs through the two-dimensional chemical fingerprints corresponding to the drugs, calculate the similarity between diseases by calculating the contribution of each disease to another disease in the disease graph, and then use the similar features of drugs and diseases to construct drug and disease feature graphs. Table 1 shows the specific information of each dataset.

[0061] In step S12, calculate the similarity between drugs through the two-dimensional chemical fingerprints corresponding to the drugs, and then obtain the drug similarity matrix. Construct the K-nearest neighbor graph as the drug feature graph according to the drug similarity matrix, so as to obtain the matrix expression of drug similarity. At the same time, obtain the corresponding adjacency matrix according to the rule that the connected elements in the K-nearest neighbor graph are set to 1 and the non-adjacent elements are set to 0.

[0062] In step S12, calculate the similarity between diseases by calculating the contribution of each disease to another disease in the disease graph, and then obtain the disease similarity matrix. Construct the K-nearest neighbor graph as the disease feature graph according to the disease similarity matrix, so as to obtain the matrix expression of disease similarity. At the same time, obtain the corresponding adjacency matrix according to the rule that the connected elements in the K-nearest neighbor graph are set to 1 and the non-adjacent elements are set to 0.

[0063] S2: Extract information of drugs and diseases in the topological space and feature space;

[0064] In step S2, it includes the following sub-step:

[0065] S21: The graph attention network extracts the information of drugs and diseases in the feature space from the drug feature graph and the disease feature graph, that is, obtains the new feature representation of drugs and the new feature representation of diseases;

[0066] S22: Extract the information of drugs and diseases in the topological space from the drug-disease association graph using a graph encoder, that is, obtain the drug topological structure representation and the disease topological structure representation.

[0067] In step S22, more specifically, a graph encoder is constructed using the Graph Matrix Completion Network (GCMC) as the backbone network. The nodes in the drug-disease association graph are used as samples. By constructing different views and comparing the similarities between them, a feature representation that can capture the graph structure and semantic information is learned. On the other hand, a multi-view learning mechanism is used to fuse the features from different views to more comprehensively describe the graph data and avoid the information incompleteness or deviation that may be brought by a single view. The Graph Matrix Completion Network takes the known and unknown drug-disease association information (provided by the drug-disease association graph) in the dataset as input. Among them, the known and unknown drug-disease associations are regarded as different edge types, and a separate processing channel is assigned to each edge type. Through message passing and specific transformations, the final representations of drugs and diseases in the topological space are finally obtained, thereby obtaining the drug topological structure representation and the disease topological structure representation.

[0068] It should be noted that the drug-disease association graph is formed by the drug-disease interaction graph. The drug-disease interaction graph has drugs and diseases as nodes and the relationships between drugs and diseases as edges. However, different from the graph attention network, in the interaction graph, drugs can only be associated with diseases, and there will be no situation where drugs are associated with drugs. Diseases can only be associated with drugs, and there will be no situation where diseases are associated with diseases.

[0069] In step S21, the features of drugs and diseases are extracted through a graph attention network. The graph attention network can adaptively learn the complex information between drugs and diseases, and uses the attention mechanism to assign different weights to different neighbor nodes in the feature map, making the model decision basis clearly visible, accurately capturing semantic information, better reflecting the real relationship, and enhancing the interpretability of the model.

[0070] More specifically, the potential features of drugs and diseases in the feature space are captured by constructing the feature maps of drugs and diseases, that is, the new drug feature representation and the new disease feature representation are obtained:

[0071] Among them, the process of extracting the new drug feature representation is as follows:

[0072] First, a K-nearest neighbor graph is constructed according to the drug similarity matrix, where the drug similarity matrix is represented as follows:

[0073] X r ∈R n×n

[0074] Among them, n is the number of drugs, and the adjacency matrix of the drug feature graph can be represented by the following formula:

[0075] A r ∈R n×n

[0076] The connected elements in the adjacency matrix are set to 1, and the non - adjacent elements are set to 0. Each specific element is defined according to the following rules:

[0077]

[0078] After obtaining the drug similarity matrix and the adjacency matrix, by using the graph attention mechanism, calculate the attention weights between drugs, and perform weighted summation on the features of adjacent drugs according to the weights to obtain the new drug feature representation. The calculation formula for the attention weights between drug nodes is as follows:

[0079]

[0080] Among them, and are the feature vectors of drugs i and j respectively, W at is the shared attention weight, F is the concatenation operation, e ij represents the correlation coefficient between drugs i and j, N i is the set of adjacent nodes of drug i, α ij is the attention weight. The new drug feature representation is as follows:

[0081]

[0082] Among them, σ represents the activation function, W is the calculation weight, is the feature vector of the adjacent node j of drug node i, represents the updated feature representation of drug node i.

[0083] The process of extracting the new disease feature representation is as follows:

[0084] First, construct a K - nearest neighbor graph according to the disease similarity matrix, where the disease similarity matrix is represented as follows:

[0085] X i ∈R m×m

[0086] Among them, m is the number of drugs, then the adjacency matrix of the disease feature graph can be represented by the following formula:

[0087] A i ∈R m×m

[0088] The connected elements in the adjacency matrix are set to 1, and the non - adjacent elements are set to 0. Each specific element is defined according to the following rules:

[0089]

[0090] After obtaining the disease similarity matrix and the adjacency matrix, by using the graph attention mechanism to calculate the attention weights between diseases, and weighted summing the features of adjacent diseases according to the weights, a new feature representation of diseases is obtained. The calculation formula for the attention weights between disease nodes is as follows:

[0091]

[0092] Among them, and are the feature vectors of diseases a and b respectively, W at is the shared attention weight, F is the concatenation operation, e ab represents the correlation coefficient between diseases a and b, N a is the set of adjacent nodes of disease a, α ab is the attention weight, and the new feature representation of diseases is as follows:

[0093]

[0094] Among them, σ represents the activation function, W is the calculation weight, is the feature vector of adjacent node b of disease node a, represents the updated feature representation of disease node a.

[0095] In the step S22, the graph encoder can effectively process graph-structured data and accurately capture the associations between nodes. It can not only focus on the local features of the nodes themselves, such as the attributes of the nodes, but also extract the global features of the entire graph topology, etc. by propagating information in the graph, forming a comprehensive and in-depth feature description of the graph, providing a solid foundation for subsequent prediction tasks.

[0096] More specifically, the graph encoder extracts the features of drugs and diseases in the topological space (drug topological structure representation, disease topological structure representation) from the known drug-disease association graph. Specifically, in the drug-disease association graph, each edge connecting a drug and a disease is assigned a separate processing channel, and the graph convolutional layer calculates only considering the first-order neighborhood of the drug or the disease, and the same transformation is applied at all positions in the graph. This type of local graph convolution can be regarded as a form of message passing, where vector-valued messages are passed and transformed on the edges of the graph. In the process of implementing drug-disease association prediction, taking the extraction of drug topological structure feature representation as an example, each drug-disease associated edge is assigned a specific transformation, so as to generate edge type-specific information. The specific information transfer formula from disease j to drug i is as follows:

[0097]

[0098] Among them, c ij is the symmetric normalization constant from disease j to drug i, r i is the node connected to drug i, d j is the node connected to disease j, N(r i ) and N(d j ) respectively represent the sets of drug and disease neighborhood nodes, and W is the parameter matrix specific to the edge type;

[0099] After obtaining the information of each edge, the information passed into each node can be obtained by cumulatively summing the edges connected to each node. The specific formula is as follows:

[0100] h i = σ[∑MP(μ j→i )]

[0101] Among them, σ is the activation function. By aggregating the information of each edge to obtain the information passed into the node, the tightness and influence of the relationship between nodes can be effectively evaluated, reflecting the status and role of the node in the topological structure.

[0102] Finally, the topological structure expression of the drug is obtained through the linear layer:

[0103] z i = W dr h i

[0104] Among them, W dr is the weight matrix of the linear layer.

[0105] Similarly, the topological structure information of the disease is obtained in the same way.

[0106] In the drug-disease association graph, calculate the information transmitted from drug m to disease n. The formula is as follows:

[0107]

[0108] Among them, c nm is the symmetric normalization constant from drug m to disease n, r n is the node connected to disease n, d m is the node connected to drug m, N(r n ) and N(d m ) respectively represent the sets of drug and disease neighborhood nodes, and W is the parameter matrix of the edge type;

[0109] After obtaining the information of each edge, the information passed into each node can be obtained by cumulatively summing the edges connected to each node:

[0110] h n = σ[ΣMP(μm→n )]

[0111] Among them, σ is the activation function;

[0112] Obtain the drug topological structure feature representation through the linear layer:

[0113] z n = W di h n

[0114] Among them, W di is the weight matrix of the linear layer.

[0115] S3: Deeply extract the drug and disease feature expressions

[0116] In step S3, use the autoencoder to further extract the features of drugs and diseases in the feature space and topological space. The encoder first encodes the preliminary features of drugs and diseases extracted by the graph convolution module (obtained through the above steps S21 and S22). In this process, the input features propagate forward along with multiple layers of neurons, and the high-dimensional features are gradually mapped to the low-dimensional feature space to accurately extract more representative features of drugs. Subsequently is the decoding process. The low-dimensional feature vectors obtained by encoding start to propagate backward and are gradually restored to data with the same dimension as the original input. On the basis of ensuring that the original drug-disease features are not lost, the effective extraction of their deeper information is realized.

[0117] S4: Fusion of topological space features and feature space information

[0118] Fuse the topological space features and feature space information of drugs and diseases.

[0119] Specifically, after obtaining the features of drugs and diseases in the feature space and topological space, use the autoencoder to further extract the features of drugs and diseases in the feature space and topological space to obtain the corresponding feature vectors; concatenate the obtained feature vectors as the final feature expressions of drugs and diseases.

[0120] S5: Model network training

[0121] Use the training set to train the graph neural network to obtain a drug-disease association prediction model;

[0122] Specifically, in step S5, the specific training scheme is as follows:

[0123] S51: Set the initial learning rate to 0.01, the number of neighbors to 3, the hyperparameter of the loss function to 0.1, and train for 4000 rounds in total;

[0124] S52: Record the optimal AUPR value and AUROC value of the model during training.

[0125] More specifically, during the model training process, the binary cross-entropy loss function is used as the main loss function, and the formula is as follows:

[0126]

[0127] where i represents the drug and j represents the disease, and y ij represents the true label of the drug-disease association. Due to the common semantics between the feature space and the topological space, a consistency constraint is used here to enhance the generality and generalization ability of the model. Specifically, two normalized embedding matrices are used to capture the similarity of drug nodes in different spaces, as shown below:

[0128]

[0129] where X Fr and X Tr are matrices formed by concatenating the features of each drug in the feature space and the topological space respectively, and are the corresponding transposed matrices. Thus, the following constraints can be generated:

[0130]

[0131] In the same way, the disease feature constraint can be obtained, which is expressed as follows:

[0132]

[0133] The final loss function is obtained by weighted combination of the above three loss functions as follows:

[0134] L = L bce + εL r + εL d

[0135] where ε is a hyperparameter for balancing the above three terms.

[0136] S6: Prediction of protein-protein interactions

[0137] The test set data is fed into the drug-disease association prediction model for prediction to obtain the prediction results. The results are shown in Tables 2 and 3.

[0138] Table 2 AUROC values of different methods on different datasets

[0139] Dataset DRHGCN NIMCGCN DRWBNCF MBiRW iDrug BNNR Our Gdataset 0.948 0.821 0.923 0.896 0.905 0.937 0.946 Cdataset 0.964 0.827 0.941 0.920 0.926 0.952 0.967 Ldataset 0.851 0.777 0.824 0.765 0.838 0.866 0.875 LRSSL 0.961 0.843 0.935 0.893 0.900 0.922 0.931 Average 0.931 0.817 0.906 0.868 0.892 0.919 0.930

[0140] Table 3 AUPR values of different methods on different datasets

[0141] Dataset DRHGCN NIMCGCN DRWBNCF MBiRW iDrug BNNR Our Gdataset 0.490 0.123 0.484 0.106 0.167 0.328 0.595 Cdataset 0.580 0.174 0.559 0.161 0.250 0.431 0.680 Ldataset 0.498 0.117 0.419 0.030 0.086 0.142 0.561 LRSSL 0.384 0.087 0.349 0.032 0.070 0.226 0.485 Average 0.488 0.125 0.453 0.082 0.143 0.282 0.580

[0142] As can be seen from Table 2 and Table 3, the AUROC performance of the model proposed by the present invention on the datasets Gdataset, Cdataset, and LRSSL is slightly lower than that of the sub-optimal model. However, in terms of AUPR, the model proposed by the present invention performs far better than the baseline model on the four datasets. Although the proposed method shows extremely excellent performance in terms of the AUPR metric, there are still some minor deficiencies in terms of AUROC. However, it is worth mentioning that in the context of dealing with imbalanced datasets, the AUPR metric has more significant advantages compared to AUROC. AUPR can provide richer and more in-depth information for the evaluation of model performance. It has a more sensitive and accurate ability to depict the relationship between different category samples in the dataset, thus being more persuasive in measuring the model performance. The model proposed by the present invention has a high prediction accuracy and is expected to be deployed on edge devices in the future for production and application.

[0143] In summary, for the drug-disease association prediction method based on graph network and autoencoder in the above embodiments, the graph attention network and graph encoder are used to extract the information of drugs and diseases in the feature space and topological space, enriching the feature expressions of drugs and diseases. At the same time, the autoencoder is used to reconstruct the feature expressions of drugs and diseases, making the drug and disease features more representative, and thus effectively improving the accuracy of the drug-disease association prediction results.

[0144] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A drug-disease association prediction method based on graph network and autoencoder, characterized in that, It includes the following steps: S1: Preprocessing of drug-disease data Calculate the similarities between drugs and between diseases in the drug-disease dataset respectively, construct drug and disease feature graphs using the similar features of drugs and diseases, and divide them into a training set and a test set according to a set ratio; S2: Information extraction of drugs and diseases in topological space and feature space In the graph convolution module, use the graph attention network to extract the features of drugs and diseases in the feature space, and at the same time use the graph encoder to extract the features of drugs and diseases in the topological space; S3: Deep extraction of drug and disease feature expressions and fusion In the autoencoder module, use the autoencoder to further extract the features of drugs and diseases in the feature space and topological space, and fuse the extracted features; S4: Model training Use the training set to train the graph neural network to obtain a drug-disease association prediction model, where the graph neural network includes a graph convolution module, an autoencoder module and a prediction module; S5: Association prediction Send the test set data into the drug-disease association prediction model for prediction to obtain the prediction results.

2. The drug-disease association prediction method based on graph network and autoencoder according to claim 1, wherein In the step S1, calculate the similarity between drugs through the two-dimensional chemical fingerprints corresponding to drugs, and then obtain the drug similarity matrix. Construct a K-nearest neighbor graph as the drug feature graph according to the drug similarity matrix, so as to obtain the matrix expression of drug similarity. At the same time, obtain the corresponding adjacency matrix according to the rule that the connected elements in the K-nearest neighbor graph are set to 1 and the non-adjacent elements are set to 0.

3. The drug-disease association prediction method based on graph network and autoencoder according to claim 2, wherein In the step S1, calculate the similarity between diseases by calculating the contribution of each disease to another disease in the disease graph, and then obtain the disease similarity matrix. Construct a K-nearest neighbor graph as the disease feature graph according to the disease similarity matrix, so as to obtain the matrix expression of disease similarity. At the same time, obtain the corresponding adjacency matrix according to the rule that the connected elements in the K-nearest neighbor graph are set to 1 and the non-adjacent elements are set to 0.

4. A drug-disease association prediction method based on graph network and autoencoder according to claim 3, characterized in that In the step S2, the specific processing process is as follows: S21: Use the graph attention network to extract the information of drugs and diseases in the feature space in the drug feature graph and disease feature graph, that is, obtain the new drug feature representation and the new disease feature representation; S22: Use the graph encoder to extract the information of drugs and diseases in the topological space in the drug-disease association graph, that is, obtain the drug topological structure representation and the disease topological structure representation, where the drug-disease association graph is a drug-disease interaction graph, and the drug-disease interaction graph is composed of drugs and diseases as nodes and the relationships between drugs and diseases as edges.

5. A drug-disease association prediction method based on graph network and autoencoder according to claim 4, wherein In the step S21, the specific process of obtaining the new drug feature representation is: After obtaining the drug similarity matrix and the adjacency matrix, calculate the attention weights between drugs by using the graph attention network, and perform weighted summation on the features of adjacent drugs according to the attention weights to obtain the new drug feature representation.

6. The drug-disease association prediction method based on graph network and autoencoder according to claim 4, wherein In the step S21, the specific process of obtaining the new disease feature representation is: After obtaining the disease similarity matrix and the adjacency matrix, calculate the attention weights between diseases by using the graph attention mechanism, and perform weighted summation on the features of adjacent diseases according to the attention weights to obtain the new disease feature representation.

7. A drug-disease association prediction method based on graph network and autoencoder according to claim 4, characterized in that In the step S22, the extraction process of the drug topological structure feature representation is as follows: S2201: In the drug-disease association graph, calculate the information transmitted from disease j to drug i, and the formula is as follows: where c ij is the symmetric normalization constant from disease j to drug i, r i is the node connected to drug i, d j is the node connected to disease j, N(r i ) and N(d j ) respectively represent the sets of drug and disease neighborhood nodes, and W is the parameter matrix of edge types; S2202: After obtaining the information of each edge, accumulate and sum the edges connected to each node to obtain the information incoming at each node: h i = σ[∑MP(μ j→i )] where σ is the activation function; S2203: Obtain the drug topological structure feature representation through a linear layer: z i = W dr h i Among them, W dr is the weight matrix of the linear layer.

8. A drug-disease association prediction method based on graph network and autoencoder according to claim 4, characterized in that, In the step S22, the extraction process of the disease topological structure feature representation is as follows: S2211: In the drug-disease association graph, calculate the information transmitted from drug m to disease n, and the formula is as follows: Among them, c nm is the symmetric normalization constant from drug m to disease n, r n is the node connected to disease n, d m is the node connected to drug m, N(r n ) and N(d m ) respectively represent the sets of drug and disease neighborhood nodes, and W is the parameter matrix of edge types; S2212: After obtaining the information of each edge, accumulate and sum the edges connected to each node to obtain the information incoming at each node: h n = σ[ΣMP(μ m→n )] where σ is the activation function; S2213: Obtain the drug topological structure feature representation through a linear layer: z n = W di h n Among them, W di is the weight matrix of the linear layer.

9. A drug-disease association prediction method based on graph network and autoencoder according to claim 4, characterized in that In the step S3, the fusion process is as follows: S31: After obtaining the features of the drug and the disease in the feature space and the topological space, use an autoencoder to further extract the features of the drug and the disease in the feature space and the topological space to obtain the corresponding feature vectors; S32: Concatenate the obtained feature vectors as the final feature representation of the drug and the disease.

10. A drug-disease association prediction method based on graph network and autoencoder according to claim 9, wherein In the step S4, the prediction module is a multi-layer perceptron. During prediction, the final feature representations of the drug and the disease obtained in step S32 are concatenated again and sent into the multi-layer perceptron for prediction.