Drug redirection prediction method and device based on graph convolutional neural network
By constructing a graph convolutional neural network model of multi-source data sets, the integrated characteristics of drugs and diseases are extracted, and the problem of inaccurate prediction of traditional drug redirection is solved, and more efficient drug redirection prediction is achieved.
Patent Information
- Application Number
- CN202510344327.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-08
AI Technical Summary
The existing drug development model is time-consuming, costly and has low success rate. Traditional drug redirection methods cannot effectively utilize the deep correlation between biological entities, resulting in inaccurate drug redirection prediction.
Using a graph convolutional neural network-based method, we use graph convolutional neural networks and relationship graph convolutional networks to extract the characteristics of biological entities and integrate them through multi-layer perception machines to predict drug redirection results.
It improves the accuracy of drug redirection prediction, can extract deep correlations between diseases and drugs, enhances the utilization of correlation information between drugs and other biological entities, and improves the accuracy of prediction.
Smart Images

Figure CN120280078A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence and medical data processing, and particularly to a method and device for predicting drug repositioning based on a graph convolutional neural network. Background Art
[0002] Drug repositioning, as a means of discovering new uses of existing drugs, has received increasing attention in medical research in recent years. The traditional drug development process often consumes huge amounts, not only requiring several years and a large amount of capital investment, but also having a probability of less than 10% for a drug to successfully pass clinical trials. At the same time, due to the insufficiently understood chemical properties of new drugs, unknown side effects may occur, thereby increasing the risks and uncertainties of patient treatment. These problems indicate that there are obvious limitations in the traditional drug discovery mode. Drug repositioning can reduce the research and development time and cost through the reuse of existing drugs, providing a faster and more effective treatment plan for patients. Therefore, the research on drug repositioning has become an important breakthrough direction in the field of drug development.
[0003] In the research of drug repositioning, predicting the drug-disease relationship is the core link. Predicting the potential association between drugs and diseases can not only help discover new treatment plans, but also significantly improve the efficiency of drug repositioning. In recent years, with the development of big data and machine learning technologies, researchers have tried to integrate biological data and informatics methods to solve this problem. Early research mainly relied on matrix factorization techniques, which performed low-dimensional representation on the drug-disease similarity matrix to mine its potential association. However, although these methods can effectively utilize the known similarity information, they often cannot extract the deep associations between biological entities, resulting in limited performance in complex biological networks and unable to achieve more accurate drug repositioning predictions. Summary of the Invention
[0004] The purpose of the present application is to provide a method and device for predicting drug repositioning based on a graph convolutional neural network, which can improve the accuracy of drug repositioning prediction.
[0005] To achieve the above purpose, the present application provides the following solutions:
[0006] In the first aspect, the present application provides a method for predicting drug repositioning based on a graph convolutional neural network, including:
[0007] Obtaining information of a multi-source data set; the information of the multi-source data set includes the basic characteristics of different types of biological entities and the interaction relationships between biological entities; biological entities include drugs, genes, proteins, and diseases;
[0008] According to the information of the multi-source dataset, apply a feature extraction model based on a graph convolutional neural network for feature extraction to obtain the integrated features of drugs and the integrated features of diseases; the feature extraction model based on a graph convolutional neural network includes a graph convolutional neural network group, a relational graph convolutional network, and a multi-layer perceptron group connected in sequence; the graph convolutional neural network group is used to extract the network features of homogeneous networks and heterogeneous networks according to the multi-source dataset; the homogeneous networks include the homogeneous network composed of each drug, the homogeneous network composed of each disease, the homogeneous network composed of each gene, and the homogeneous network composed of each protein; the heterogeneous networks include the heterogeneous network composed of drugs and diseases, the heterogeneous network composed of drugs and genes, the heterogeneous network composed of drugs and proteins, the heterogeneous network composed of diseases and genes, and the heterogeneous network composed of diseases and proteins; the relational graph convolutional network is used to fuse the network features of the extracted homogeneous networks and heterogeneous networks; the multi-layer perceptron group is used to obtain the features of the integrated drugs and the features of the integrated diseases according to the features fused by the relational graph convolutional network;
[0009] Determine the redirection prediction results of each drug according to the integrated features of the drugs and the integrated features of the diseases.
[0010] In a second aspect, the present application provides a drug redirection prediction device based on a graph convolutional neural network, including:
[0011] A data acquisition module for acquiring the information of the multi-source dataset; the information of the multi-source dataset includes the basic features of different types of biological entities and the interaction relationships between biological entities; the biological entities include drugs, genes, proteins, and diseases;
[0012] A feature extraction module for applying a feature extraction model based on a graph convolutional neural network for feature extraction according to the information of the multi-source dataset to obtain the integrated features of drugs and the integrated features of diseases; the feature extraction model based on a graph convolutional neural network includes a graph convolutional neural network group, a relational graph convolutional network, and a multi-layer perceptron group connected in sequence; the graph convolutional neural network group is used to extract the network features of homogeneous networks composed of the same type of biological entities and heterogeneous networks composed of different types of biological entities according to the multi-source dataset; the relational graph convolutional network is used to fuse the network features of the extracted homogeneous networks and heterogeneous networks; the multi-layer perceptron group is used to integrate the features of drugs and the features of integrated diseases according to the features fused by the relational graph convolutional network;
[0013] A drug redirection prediction module for determining the redirection prediction results of each drug according to the integrated features of the drugs and the integrated features of the diseases.
[0014] In a third aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned drug redirection prediction method based on a graph convolutional neural network is implemented.
[0015] In a fourth aspect, the present application provides a computer program product, including a computer program which, when executed by a processor, implements the above-mentioned drug repositioning prediction method based on a graph convolutional neural network.
[0016] According to the specific embodiments provided by the present application, the following technical effects are disclosed:
[0017] The present application provides a drug repositioning prediction method and apparatus based on a graph convolutional neural network. Based on the information of a multi-source data set, a feature extraction model based on a graph convolutional neural network is applied to extract the features of drugs and diseases. Since the multi-source data set includes several drugs, several genes, several proteins, several diseases, the basic features of each drug, the basic features of each gene, the basic features of each protein, the basic features of each disease, the interaction relationships between drugs and genes, the interaction relationships between drugs and proteins, the interaction relationships between drugs and diseases, the interaction relationships between diseases and genes, and the interaction relationships between diseases and proteins, the multi-source data set not only includes several drugs and several diseases, but also includes the relevant information of genes and proteins. When extracting features, not only the direct association information between diseases and drugs can be considered, but also the indirect association information from genes and proteins is integrated, and the deep-level association between diseases and drugs can be extracted, thereby improving the accuracy of drug repositioning prediction. In addition, when performing feature extraction, a feature extraction model composed of a graph convolutional neural network and a relational graph convolutional neural network is applied, and the association information between diseases and drugs, the association information between drugs and genes, the association information between drugs and proteins, the association information between diseases and genes, and the association information between diseases and proteins can be fully mined. Finally, the integrated features of diseases and drugs extracted take into account the influence of genes and proteins, and the deep-level association between diseases and drugs can be extracted, thereby improving the accuracy of drug repositioning prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 It is an application environment diagram of a drug repositioning prediction method based on a graph convolutional neural network in an embodiment of the present application;
[0020] Figure 2 It is a flowchart of a drug repositioning prediction method based on a graph convolutional neural network provided in an embodiment of the present application;
[0021] Figure 3 Schematic diagram of the technical concept of a drug repositioning prediction method based on a graph convolutional neural network provided by an embodiment of the present application;
[0022] Figure 4 Schematic diagram of the structure of a feature extraction model based on a graph convolutional neural network provided by an embodiment of the present application;
[0023] Figure 5 Schematic diagram of the functional modules of a drug repositioning prediction device based on a graph convolutional neural network provided by an embodiment of the present application;
[0024] Figure 6 Schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0025] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0026] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0027] The drug repositioning prediction method based on a graph convolutional neural network provided by the embodiments of the present application can be applied to an application environment as shown in Figure 1 . Among them, the terminal communicates with the server through the network. The data storage system can store the data that the server needs to process. The data storage system can be set separately, integrated on the server, or placed on the cloud or other servers. The terminal can send the information of the multi-source data set to be processed to the server. The server receives the information of the multi-source data set to be processed. The server applies a feature extraction model based on a graph convolutional neural network to extract features according to the information of the multi-source data set, and obtains the integrated features of drugs and the integrated features of diseases; determines the drug repositioning prediction results according to the integrated features of drugs and the integrated features of diseases. The server can feedback the obtained drug repositioning prediction results to the terminal. In addition, in some embodiments, the drug repositioning prediction method based on a graph convolutional neural network can also be implemented by the server or the terminal alone. For example, the terminal can directly perform drug repositioning prediction on the information of the multi-source data set to be processed, or the server can obtain the information of the multi-source data set to be processed from the data storage system and perform drug repositioning prediction.
[0028] Among them, the terminal can be, but is not limited to, various desktop computers, laptop computers, smartphones, tablet computers, Internet of Things devices, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers, and can also be a cloud server.
[0029] In an exemplary embodiment, the network structure-based method provides a new solution idea. By constructing a multi-dimensional network of drugs, diseases, and other biological entities, the topological relationships among them in the biological system are analyzed. With the development of network analysis technology, more and more studies use technologies such as network propagation and label propagation to identify potential connections between drugs and diseases. Compared with traditional methods, such methods can better capture the complex associations between biological entities. However, with the application of machine learning and deep learning technologies, graph neural networks have gradually become the mainstream method in the drug-disease relationship prediction task. By modeling various biological entities such as drugs, diseases, and related proteins and genes as a heterogeneous network, graph neural networks can extract features of drugs and diseases from multiple levels and improve the prediction accuracy by learning the complex relationships between them. In future development, how to integrate more biological information from multiple perspectives to further improve the accuracy of drug-disease association prediction is still the key research direction for researchers. By introducing a more accurate negative sample selection method (by adding some incorrect samples to the training samples to make the model learn that these samples are incorrect, thus avoiding the model generating incorrect samples), reducing the noise in the prediction process, and exploring joint prediction methods under the multi-task learning framework, it is expected to promote further breakthroughs in drug repositioning research. As Figure 2 and Figure 3 shown, this application provides a drug repositioning prediction method based on a graph convolutional neural network. This method is executed by a computer device, and specifically can be executed alone by a computer device such as a terminal or a server, or jointly executed by a terminal and a server. In the embodiment of this application, taking this method applied to Figure 1 the server in
[0030] Step 101: Obtain the information of the multi-source dataset; the information of the multi-source dataset includes the basic characteristics of different types of biological entities and the interaction relationships between biological entities; biological entities include drugs, genes, proteins, and diseases. Specifically, the information of the multi-source dataset includes several drugs, several genes, several proteins, several diseases, the basic characteristics of each drug, the basic characteristics of each gene, the basic characteristics of each protein, the basic characteristics of each disease, the interaction relationship between drugs and genes, the interaction relationship between drugs and proteins, the interaction relationship between drugs and diseases, the interaction relationship between diseases and genes, the interaction relationship between diseases and proteins, the interaction relationship between genes and genes, and the interaction relationship between proteins and proteins. Examples of multi-source datasets include: DrugBank v3.0 dataset for drug data, Mesh dataset for disease data, HumanNet dataset for gene data, human protein reference dataset for protein data, Comparative Toxicogenomics dataset for drug and disease relationship data, DGIdb v5.0 dataset for drug and gene relationship dataset, etc. The interaction relationship between a drug and a disease means that the drug can be used to treat a certain type of disease. For example, drug A can treat disease B, and the connection between A and B is 1.
[0031] Step 102: According to the information of the multi-source dataset, apply a feature extraction model based on a graph convolutional neural network for feature extraction to obtain the integrated features of drugs and the integrated features of diseases. The feature extraction model based on a graph convolutional neural network includes a graph convolutional neural network group, a relational graph convolutional network, and a multi-layer perceptron group connected in sequence; the graph convolutional neural network group is used to extract the network features of homogeneous networks composed of the same type of biological entities and heterogeneous networks composed of different types of biological entities according to the multi-source dataset; the homogeneous networks include the homogeneous network composed of each drug, the homogeneous network composed of each disease, the homogeneous network composed of each gene, and the homogeneous network composed of each protein; the heterogeneous networks include the heterogeneous network composed of drugs and diseases, the heterogeneous network composed of drugs and genes, the heterogeneous network composed of drugs and proteins, the heterogeneous network composed of diseases and genes, and the heterogeneous network composed of diseases and proteins; the relational graph convolutional network is used to fuse the network features of the extracted homogeneous networks and heterogeneous networks; the multi-layer perceptron group is used to integrate the features of drugs and integrate the features of diseases according to the features fused by the relational graph convolutional network.
[0032] Step 103: Determine the redirection prediction results of each drug according to the integrated features of drugs and the integrated features of diseases.
[0033] By implementing the above-mentioned steps 101 to 103, based on the information of the multi-source dataset, a feature extraction model based on a graph convolutional neural network is applied to extract the features of drugs and diseases. Since the multi-source dataset includes several drugs, several genes, several proteins, several diseases, the basic features of each drug, the basic features of each gene, the basic features of each protein, the basic features of each disease, the interaction relationships between drugs and genes, the interaction relationships between drugs and proteins, the interaction relationships between drugs and diseases, the interaction relationships between diseases and genes, and the interaction relationships between diseases and proteins, the multi-source dataset not only includes several drugs and several diseases, but also includes the relevant information of genes and proteins. When extracting features, not only can the direct association information between diseases and drugs be considered, but also the indirect association information from genes and proteins can be integrated, and the deep association between diseases and drugs can be extracted, thereby improving the accuracy of drug repositioning prediction. In addition, when performing feature extraction, a feature extraction model composed of a graph convolutional neural network and a relational graph convolutional neural network is applied, which can fully mine the association information between diseases and drugs, the association information between drugs and genes, the association information between drugs and proteins, the association information between diseases and genes, and the association information between diseases and proteins. Finally, the integrated features of diseases and drugs extracted consider the influence of genes and proteins, and the deep association between diseases and drugs can be extracted, thereby improving the accuracy of drug repositioning prediction.
[0034] In another exemplary embodiment of the present application, as Figure 3 and Figure 4 shown, the graph convolutional neural network group includes 7 graph convolutional neural networks and 4 fully connected networks. The multi-layer perceptron group includes 2 multi-layer perceptrons.
[0035] Among them, the 7 graph convolutional neural networks are respectively denoted as the first graph convolutional neural network to the seventh graph convolutional neural network (GCN1 to GCN7) in sequence; the 4 fully connected networks are respectively denoted as the first fully connected network to the fourth fully connected network (FC1 to FC4) in sequence; the 4 multi-layer perceptrons are respectively denoted as the first multi-layer perceptron to the fourth multi-layer perceptron (MLP1 to MLP4) in sequence.
[0036] The input of the first graph convolutional neural network GCN1 is the relationship matrix between each drug and the basic features of each drug, that is, the features of the homogeneous network composed of each drug; the elements of the relationship matrix between each drug are the similarities between each drug; the inputs of the first fully connected network FC1 to the fourth fully connected network FC4 are the basic features of each drug, the basic features of each gene, the basic features of each protein, and the basic features of each disease respectively. Figure 3In it, in the drug similarity network A, the nodes are various types of drugs, the node features are the basic features of the drugs, and the edge relationships between the nodes are the similarities between the corresponding drugs. In the gene interaction network B, the nodes are various types of genes, the node features are the basic features of the genes, and the edge relationships between the nodes are the interaction relationships between the corresponding genes. In the protein interaction network C, the nodes are various types of proteins, the node features are the basic features of the proteins, and the edge relationships between the nodes are the interaction relationships between the corresponding proteins. In the disease similarity network D, the nodes are various types of diseases, the node features are the basic features of the diseases, and the edge relationships between the nodes are the similarities between the corresponding diseases.
[0037] The input of the second graph convolutional neural network GCN2 is connected to the output of the first fully connected network FC1 and the output of the second fully connected network FC2; the input of the second graph convolutional neural network GCN2 also includes the interaction relationship between drugs and genes. The input of the second graph convolutional neural network GCN2 is the basic features of each drug extracted by the first fully connected network FC1, the basic features of each gene extracted by the second fully connected network FC2, and the interaction relationship between drugs and genes, that is, the features of the heterogeneous network composed of drugs and genes.
[0038] If there is an interaction relationship between a certain drug and a certain gene, the corresponding element of the interaction relationship matrix between each drug and each gene can be set to 1, otherwise it is set to 0. Whether there is an interaction relationship between a certain drug and a certain gene is determined according to the research data obtained from the existing research on drug genes in this field.
[0039] The input of the third graph convolutional neural network GCN3 is connected to the output of the first fully connected network FC1 and the output of the third fully connected network FC3; the input of the third graph convolutional neural network GCN3 also includes the interaction relationship between drugs and proteins. The input of the third graph convolutional neural network GCN3 is the basic features of each drug extracted by the first fully connected network FC1, the basic features of each protein extracted by the third fully connected network FC3, and the interaction relationship between drugs and proteins, that is, the features of the heterogeneous network composed of drugs and proteins.
[0040] If there is an interaction relationship between a certain drug and a certain protein, the corresponding element of the interaction relationship matrix between each drug and each protein can be set to 1, otherwise it is set to 0. Whether there is an interaction relationship between a certain drug and a certain protein is determined according to the research data obtained from the existing research on drug proteins in this field.
[0041] The input of the fourth graph convolutional neural network GCN4 is connected to the output of the first fully connected network FC1 and the output of the fourth fully connected network FC4; the input of the fourth graph convolutional neural network GCN4 also includes the interaction relationship between drugs and diseases. The input of the fourth graph convolutional neural network GCN4 is the basic features of each drug extracted by the first fully connected network FC1, the basic features of each disease extracted by the fourth fully connected network FC4, and the interaction relationship between drugs and diseases, that is, the features of the heterogeneous network composed of drugs and diseases.
[0042] The input of the fifth graph convolutional neural network GCN5 is connected to the output of the fourth fully connected network FC4 and the output of the second fully connected network FC2; the input of the fifth graph convolutional neural network GCN5 also includes the interaction relationship between diseases and genes. If there is an interaction relationship between a certain disease and a certain gene, the corresponding element of the interaction relationship matrix between each disease and each gene can be set to 1, otherwise it is set to 0. Whether there is an interaction relationship between a certain disease and a certain gene is determined according to the research data obtained from the existing research on disease genes in this field. The input of the fifth graph convolutional neural network GCN5 is the basic features of each disease extracted by the fourth fully connected network FC4, the basic features of each gene extracted by the second fully connected network FC2, and the interaction relationship between diseases and genes, that is, the features of the heterogeneous network composed of diseases and genes.
[0043] The input of the sixth graph convolutional neural network GCN6 is connected to the output of the fourth fully connected network FC4 and the output of the third fully connected network FC3; the input of the sixth graph convolutional neural network GCN6 also includes the interaction relationship between diseases and proteins. The input of the sixth graph convolutional neural network GCN6 is the basic features of each disease extracted by the fourth fully connected network FC4, the basic features of each protein extracted by the third fully connected network FC3, and the interaction relationship between diseases and proteins, that is, the features of the heterogeneous network composed of diseases and proteins.
[0044] If there is an interaction relationship between a certain disease and a certain protein, the corresponding element of the interaction relationship matrix between each disease and each protein can be set to 1, otherwise it is set to 0. Whether there is an interaction relationship between a certain disease and a certain protein is determined according to the research data obtained from the existing research on disease proteins in this field.
[0045] The input of the seventh graph convolutional neural network GCN7 is the relationship matrix between each disease and the basic features of each disease, that is, the homogeneous network composed of each disease; the elements of the relationship matrix between each disease are the similarities between each disease.
[0046] The outputs of the first graph convolutional neural network GCN1 to the seventh graph convolutional neural network GCN7 are all connected to the relational graph convolutional network RGCN; the outputs of the relational graph convolutional network RGCN are respectively connected to the first multi-layer perceptron MLP1 to the fourth multi-layer perceptron MLP4.
[0047] The input of the first multi-layer perceptron MLP1 is the global node representation of the drug output by the relational graph convolutional network; the input of the second multi-layer perceptron MLP2 is the global node representation of the disease output by the relational graph convolutional network; the input of the third multi-layer perceptron MLP3 is the global node representation of the gene output by the relational graph convolutional network; the input of the fourth multi-layer perceptron MLP4 is the global node representation of the protein output by the relational graph convolutional network.
[0048] The outputs of the first multi-layer perceptron MLP1 to the fourth multi-layer perceptron MLP4 are respectively the integrated features of the drug, the integrated features of the disease, the integrated features of the gene, and the integrated features of the protein.
[0049] The feature extraction model in step 102 is a trained model, and the loss function during training is Loss = Loss dd +β(Loss drg +Loss drp +Loss dig +Loss dip ), where Loss dd represents the difference between the predicted relationship between the drug and the disease and the true relationship, β represents a hyperparameter, and Loss drg , Loss_ drp , Loss dig , Loss dip respectively represent the differences between the predicted relationships and the true relationships between the drug and the gene, the drug and the protein, the disease and the gene, and the disease and the protein.
[0050] In another exemplary embodiment of the present application, based on a multi-source dataset related to drugs and diseases, a heterogeneous information network G=(V, E) is constructed, such as the network structure shown in (c) in Figure 3 , where the node set V includes a drug node set V dr , a disease node set V di , a gene node set V g and a protein node set V p。The meta-paths included in the edge set E are of four categories: drug-disease association, drug-gene-disease association, drug-protein-disease association, and drug-gene-protein-disease association. In this application, the similarity between drug nodes is used as the initial feature vector of drugs, the similarity between disease nodes is used as the initial feature vector of diseases, the interaction between gene nodes is used as the initial feature vector of genes, and the interaction between protein nodes is used as the initial feature vector of proteins. Such a structured information network can not only effectively capture the relationship between drugs and diseases, but also provide a rich information basis for subsequent analysis and prediction through multi-level node features.
[0051] Based on the above content and Figure 4 shown in the feature extraction model structure based on the graph convolutional neural network, in step 102, according to the information of the multi-source data set, a feature extraction model based on the graph convolutional neural network is applied for feature extraction to obtain the integrated features of drugs and the integrated features of diseases, which specifically include:
[0052] (1) Calculate the similarity between drugs according to the basic features of each drug, and determine the relationship matrix between each drug according to the similarity between each drug.
[0053] (2) Calculate the similarity between diseases according to the basic features of each disease, and determine the relationship matrix between each disease according to the similarity between each disease.
[0054] (3) Input the relationship matrix between each drug and the basic features of each drug into the first graph convolutional neural network GCN1 to obtain the drug node representation; input the relationship matrix between each disease and the basic features of each disease into the seventh graph convolutional neural network GCN7 to obtain the disease node representation.
[0055] According to Figure 3 shown in (c) of the heterogeneous information network structure, the structure includes the direct association between drugs and diseases and the indirect association between drugs and diseases. First, in order to make full use of the drug similarity network and the disease similarity network to capture the potential features of drugs and diseases, the initial feature vectors of drugs and diseases and their respective basic features are respectively input into their respective graph convolutional neural networks, which are specifically represented as follows.
[0056]
[0057] Among them, l represents the number of layers of the first graph convolutional neural network and the seventh graph convolutional neural network; and respectively represent the drug node representations of the l-th layer and the (l-1)-th layer of the first graph convolutional neural network. When l is 1, Represents the relationship matrix between each drug and the basic characteristics of each drug; and respectively represent the disease node representations of the l-th layer and the (l - 1)-th layer of the seventh graph convolutional neural network. When l = 1, represents the relationship matrix between each disease and the basic characteristics of each disease; relu(·) is the activation function; and are respectively the degree matrices of the adjacency matrix of the drug node and the adjacency matrix of the disease node; and A dr and A di are respectively the adjacency matrices of the drug node and the disease node, and I dr and I di are both identity matrices; and are respectively the weight parameter matrices of the l-th layer of the first graph convolutional neural network and the seventh graph convolutional neural network; and are respectively the bias matrices of the l-th layer of the first graph convolutional neural network and the seventh graph convolutional neural network. By integrating the neighbor information in the homogeneous network (drug similarity network) corresponding to each drug and the neighbor information in the homogeneous network (disease similarity network) corresponding to each disease, more abstract and global drug features and disease features can be captured, thereby enhancing the context awareness ability of the feature extraction model.
[0058] (4) Input the basic characteristics of each drug into the first fully connected network FC1 to obtain the extracted features of each drug; input the basic characteristics of each gene into the second fully connected network FC2 to obtain the extracted features of each gene; input the basic characteristics of each protein into the third fully connected network FC3 to obtain the extracted features of each protein; input the basic characteristics of each disease into the fourth fully connected FC4 layer to obtain the extracted features of each disease.
[0059] To effectively learn the embedding representations of the four types of nodes in their respective homogeneous networks, four fully connected networks (FC) are used for feature extraction, which are specifically represented as follows.
[0060]
[0061] Among them, dr2, g1, p1, and di2 respectively represent the extracted features of each drug, the extracted features of each gene, the extracted features of each protein, and the extracted features of each disease; X dr , X g , X p and X di respectively represent the basic characteristics of drugs, genes, proteins, and diseases; is the weight parameter of the fully connected network; is the bias parameter of the fully connected network; σ(·) is the ReLU activation function.
[0062] (5) Input the drug extraction features, gene extraction features, and the interaction relationships between drugs and genes into the second graph convolutional neural network GCN2 to obtain the node representations of drugs and genes respectively; input the drug extraction features, protein extraction features, and the interaction relationships between drugs and proteins into the third graph convolutional neural network GCN3 to obtain the node representations of drugs and proteins respectively; input the drug extraction features, disease extraction features, and the interaction relationships between drugs and diseases into the fourth graph convolutional neural network GCN4 to obtain the node representations of drugs and diseases respectively; input the disease extraction features, gene extraction features, and the interaction relationships between diseases and genes into the fifth graph convolutional neural network GCN5 to obtain the node representations of diseases and genes respectively; input the disease extraction features, protein extraction features, and the interaction relationships between diseases and proteins into the sixth graph convolutional neural network GCN6 to obtain the node representations of diseases and proteins respectively.
[0063] In the heterogeneous network, different types of edges represent different semantic information. For the direct association between drugs and diseases, the present application establishes a graph convolutional neural network (corresponding to GCN4) to learn the topological structure and attributes of the nodes in the heterogeneous network composed of drugs and diseases, as follows.
[0064]
[0065] Among them, l represents the number of layers of the fourth graph convolutional neural network; represents the node representations of drugs and diseases output by the l-th layer of the fourth graph convolutional neural network; represents the node representations of drugs and diseases output by the (l - 1)-th layer of the fourth graph convolutional neural network; the input of the first layer of the fourth graph convolutional neural network is and is the weight matrix of the l-th layer of the fourth graph convolutional neural network; represents the degree matrix of the adjacency matrix of drug nodes and disease nodes; A dd is the adjacency matrix of drug nodes and disease nodes, and I dd is the identity matrix;
[0066] Meanwhile, for the indirect association, heterogeneous graph convolutional neural networks (corresponding to GCN2, GCN3, GCN5, GCN6) are established for drugs and genes, drugs and proteins, diseases and genes, and diseases and proteins respectively to explore how drugs affect diseases through specific genes or proteins. Specifically, the graph convolutional neural networks for drugs and genes, and drugs and proteins are as follows.
[0067]
[0068] The graph convolutional neural networks for diseases and genes, and diseases and proteins are as follows.
[0069]
[0070] Among them, the inputs of the first layer of GCN2, GCN3, GCN5, and GCN6 are: and and respectively represent the node representations of drugs and genes in the l-th layer and the (l-1)-th layer of the second graph convolutional neural network; and respectively represent the node representations of drugs and proteins in the l-th layer and the (l-1)-th layer of the third graph convolutional neural network; is the degree matrix of the adjacency matrix of drug nodes and gene nodes; is the degree matrix of the adjacency matrix of drug nodes and protein nodes; and A drg are respectively the adjacency matrices of drug nodes and gene nodes; A drp are respectively the adjacency matrices of drug nodes and protein nodes; I drg and I drp are both identity matrices; and respectively represent the node representations of diseases and genes in the l-th layer and the (l-1)-th layer of the fifth graph convolutional neural network; and respectively represent the node representations of diseases and proteins in the l-th layer and the (l-1)-th layer of the sixth graph convolutional neural network; is the degree matrix of the adjacency matrix of disease nodes and gene nodes; is the degree matrix of the adjacency matrix of disease nodes and protein nodes; and A dig are respectively the adjacency matrices of disease nodes and gene nodes; A dip are respectively the adjacency matrices of disease nodes and protein nodes; I dig and I dip are both identity matrices; respectively represent the weight parameters of the l-th layer of the second graph convolutional neural network, the l-th layer of the third graph convolutional neural network, the l-th layer of the fifth graph convolutional neural network, and the l-th layer of the sixth graph convolutional neural network.
[0071] (6) Input the drug node representation, the node representations of drug-gene, drug-protein, drug-disease, disease-gene, disease-protein, and disease node into the Relational Graph Convolutional Network (RGCN) to obtain the global node representations of drugs, genes, proteins, and diseases respectively.
[0072] To capture and aggregate the topological and semantic information of heterogeneous networks from a global perspective, we adopt the Relational Graph Convolutional Network (RGCN). Specifically, the new embedding representation of each node in the network is the weighted sum of the embeddings of all adjacent nodes, and these weights depend on the types of adjacent nodes and edges. Therefore, the RGCN uses multiple aggregation methods to capture the potential relationship information between nodes, thereby enhancing the learning ability and performance of the RGCN. The aggregation operation is expressed as follows:
[0073]
[0074] where R represents the type of edges in the RGCN; denotes the set of node neighbors of node i in the RGCN under edge type r; represents the weight parameter of edge type r; represents the weight parameter of node i itself; is the embedding representation of node i in the m-th layer of the RGCN. When m = 0, the input to the first layer of the RGCN is As Figure 4 shown, the output of the last layer of the RGCN model is:
[0075] h i (dr di ,dr g ,dr p ,di dr ,di g ,di p ,g dr ,g di ,p dr ,p di ).
[0076] where dr di , dr g , dr p are all drug features, representing the updated features of drug-disease, gene-protein respectively. di dr , di g , di p are all disease features, representing the updated features of disease-drug, gene-protein respectively. g dr , gdi They are all gene features, representing the features after gene and disease, and drug updates respectively. p dr 、p di They are all protein features, representing the features after protein and disease, and drug updates respectively.
[0077] In the present application, the nodes in the graph convolutional neural network and the relational graph convolutional neural network are a vector, for example, a vector with 256 dimensions. The GCN aims to learn and update the embedding representation of the nodes by aggregating the information of neighboring nodes. Each layer will aggregate the features of neighboring nodes, and finally obtain a high-dimensional vector. The node representation is dynamically learned and depends on the graph structure and node features. Different from the nodes in the knowledge graph, the nodes in the knowledge graph are a triple (head entity, relation, tail entity), such as (Alzheimer's disease, related gene, APOE), which stores the relationship between the head entity and the tail entity. The node representation of the knowledge graph aims to capture the semantic information of entities and their relationships, and is usually used in tasks such as knowledge reasoning and question answering systems. The node representation can be predefined (such as Word2Vec, TransE, etc.) or learned through a specific model (such as R-GCN), emphasizing the semantics of entities and relationships. The node representation is usually static and based on the existing structure of the knowledge graph.
[0078] (7) Input the features of the drug's global node representation that interact with the disease, the features that interact with the gene, and the features that interact with the protein into the first multi-layer perceptron MLP1 for feature integration to obtain the integrated feature of the drug (corresponding to Figure 4 dr in); Input the features of the disease's global node representation that interact with the drug, the features that interact with the gene, and the features that interact with the protein into the second multi-layer perceptron MLP2 for feature integration to obtain the integrated feature of the disease (corresponding to Figure 4 di in); Input the features of the gene's global node representation that interact with the drug and the features that interact with the disease into the third multi-layer perceptron MLP3 for feature integration to obtain the integrated feature of the gene (corresponding to Figure 4 g in); Input the features of the protein's global node representation that interact with the drug and the features that interact with the disease into the fourth multi-layer perceptron MLP4 for feature integration to obtain the integrated feature of the protein (corresponding to Figure 4 p in).
[0079] In the present application, since the research is on drug repositioning prediction, mainly mining the deep association relationship between drugs and diseases, obtaining the drug features considering the influence of diseases, genes, and proteins, and the disease features considering the influence of drugs, genes, and proteins, when performing drug repositioning prediction, only the integrated features of the drug and the integrated features of the disease need to be applied.
[0080] In another exemplary embodiment of the present application, in step (1) above, in order to better explore the potential characteristics of drugs and diseases, similarity is used as their initial characteristic. For the input isomorphic drug interaction network (in the network, the cosine similarity between drug characteristics is used as the weight of the edge between two drugs, thus forming a weighted network), a drug similarity network is constructed based on the maximum common subgraph of drugs. The similarity score between drug pairs is calculated by the Jaccard coefficient. For two drug association graphs Dr1 and Dr2, the maximum common subgraph MCS(Dr1, Dr2) can be found by searching their maximum cliques. Then, the Jaccard coefficient (a common similarity measure) is used to calculate the similarity between two drugs. JC(Dr1, Dr2) can be expressed as:
[0081]
[0082] Through the drug similarity network, the initial feature vector representations of drugs can be obtained, and these representations will be used to jointly optimize the feature vectors in the subsequent heterogeneous graph convolutional network. The initial feature vectors provide the association information between drugs, enabling them to better capture potential drug-disease relationships. This optimization process will help improve the accuracy and reliability of drug-disease association prediction.
[0083] Based on the above, calculating the similarity between drugs according to the basic characteristics of each drug specifically includes:
[0084] (1-1) Determine the drug association graph of each drug according to the basic characteristics of each drug.
[0085] (1-2) Determine the maximum common subgraph of every two drugs according to the drug association graphs of every two drugs.
[0086] (1-3) Calculate the Jaccard coefficient according to the maximum common subgraph and drug association graph of every two drugs, and obtain the similarity between every two drugs.
[0087] In another exemplary embodiment of the present application, in step (2) above, for the calculation of disease similarity, disease similarity is constructed through the semantic similarity between diseases in the MeSH database. The diseases in the MeSH database are divided into general parent diseases and child diseases according to their classification granularity. The semantic information of a disease is defined as the sum of the semantic contributions of itself and its child diseases. For a pair of diseases Di1 and Di2, the similarity between them is the semantic contribution value of the common parent disease divided by the sum of the semantic contribution values of the two diseases. The calculation formula is as follows:
[0088]
[0089] Among them, S(Di1, Di2) represents the similarity between diseases Di1 and Di2; CP(Di1, Di2) represents the semantic contribution value of the common parent disease of diseases Di1 and Di2 (existing data in the MeSH database); CP(Di1) represents the semantic contribution value of disease Di1; CP(Di2) represents the semantic contribution value of disease Di2. Compared with using disease relationship data, similarity data can capture weak indirect connections between diseases and provide more comprehensive and comprehensive data support for subsequent models.
[0090] Calculate the similarity between each disease according to the basic characteristics of each disease, including: 1) Obtain each disease node and its related information (including the basic characteristics of the disease), and establish a relationship framework between disease nodes to reflect the similarity between two diseases; 2) According to the classification hierarchy information of the disease, divide the disease into general parent diseases and child diseases; 3) Determine the semantic information of the disease, and define the semantic information of each disease as the sum of the semantic contributions of itself and its child diseases; 4) For each pair of diseases, calculate the semantic contribution value of the common parent disease, and divide it by the sum of the semantic contribution values of the two diseases to obtain the similarity of each pair of diseases.
[0091] In another exemplary embodiment of the present application, in step 103, determining each drug redirection prediction result according to the integrated characteristics of the drug and the integrated characteristics of the disease specifically includes:
[0092] (i) Calculate the cosine similarity between the integrated characteristics of each drug and the integrated characteristics of each disease.
[0093] (ii) Sort the cosine similarities to determine the redirection prediction results of each drug.
[0094] For each disease, calculate the cosine similarity of the integrated characteristics between each drug and the disease, sort the cosine similarities, and select some drugs with the highest cosine similarities as potential drugs for treating the corresponding disease.
[0095] In order to verify the effectiveness of the drug redirection prediction based on the graph convolutional neural network of the present application, a comparative experiment is given below. An heterogeneous network containing four types of nodes and nine types of edges (drug-drug, disease-disease, gene-gene, protein-protein, drug-disease, drug-gene, drug-protein, disease-gene, disease-protein) is constructed by collecting a public data set, representing diverse drug-related and disease-related information. Specifically, the multi-source data set includes 542 drugs, 394 diseases, 11,153 genes and 1,512 proteins, and the associations between them.
[0096] To evaluate the effectiveness and accuracy of the method proposed in this application, it is compared with nine baseline models: HMLKGAT, DRAGNN, MGRMF, LBMFF, DRWBNCF, REDDA, LAGCN, MGRNNM, and DTINet. To ensure the rigor of the comparative experiment, all baseline models are trained and tested on the dataset proposed in this application. The experimental results are shown in Table 1. Except that the AUC index is slightly lower than that of the DRWBNCF method, the method of this application is superior to the other nine methods in four indexes: AUPR, F1-score, Precision, and Recall. Compared with the best baseline method, the method of this application leads by 5.46%, 7.10%, 5.09%, and 9.21% respectively in these four indexes. This indicates that the method of this application has more superior prediction performance. The possible reason for this difference is that most baseline models only consider the association between drugs and diseases and do not fully utilize the complex relationships between other biological entities (genes and proteins). In contrast, the method proposed in this application not only utilizes the direct information of drugs and diseases but also integrates the indirect information from genes and proteins. This enables a more comprehensive exploration of implicit associations and obtains more complex node embedding representations. In summary, the method proposed in this application performs better in predicting drug-disease associations.
[0097] Table 1 Comparative Experiment Results
[0098]
[0099]
[0100] To evaluate the contributions of various biological entities in this application, two variants of the model (ontology) of this application are set up for ablation study.
[0101] w / o protein: In this variant, only the heterogeneous network of drug-gene-disease is used to learn the embedding representations of drug, gene, and disease nodes.
[0102] w / o gene: In this variant, only the heterogeneous network of drug-protein-disease is used to learn the embedding representations of drug, protein, and disease nodes.
[0103] Table 2 Ablation Experiment Results
[0104] AUC AUPR F1-score Precision recall w / oprotein 0.9476 0.9550 0.8767 0.8640 0.8998 w / ogene 0.9537 0.9454 0.8755 0.8485 0.9042 the method of this application 0.9584 0.9611 0.8912 0.8761 0.9072
[0105] As shown in the experimental results of Table 2, the method of this application performs better than the two variants, indicating that more comprehensive indirect associations help improve the prediction performance of drug redirection. This experimental result further verifies the conclusion of the above comparative experiment.
[0106] This application constructs a heterogeneous information network containing four node types and nine edge types by collecting public datasets from multiple sources to comprehensively characterize the diverse associations between drugs and diseases. The dataset covers 542 drugs, 394 diseases, 11,153 genes, 1,512 proteins, and various relationships between these nodes. These associations include not only direct drug-disease relationships but also indirect connections such as drug-gene, drug-protein, and gene-disease, providing a highly interconnected biological network framework. By integrating the functions of drugs, the symptoms and mechanisms of diseases, the expression patterns of genes, and the functions and interactions of proteins, the heterogeneous information network structure can reflect multi-dimensional biological background information, providing rich feature support for in-depth analysis of the complex associations between drugs and diseases and their potential interaction mechanisms. In addition, this heterogeneous network structure also provides a solid data foundation for exploring new drug targets and potential treatment methods. Comparative experiments prove that the method proposed in this application performs better than the baseline method. The ablation experiment results evaluate the contribution of each module in the method proposed in this application. In summary, the method proposed in this application is a powerful tool for drug repositioning. Here, the module can be understood as follows: the entire process of protein data processing is a module, and the entire process of gene data processing is a module. Because in the ablation experiment, one variant only uses genes and the other only uses proteins, the contribution of each module can be seen.
[0107] Drug repositioning is a key aspect of biomedical research, and predicting drug-disease associations (DDA) is an important step in drug repositioning. With the development of deep learning and neural network technologies, graph convolutional networks (GCNs) have achieved remarkable results in this research field. Although existing DDA models have made great progress, there is still room for improvement in fully utilizing and integrating information from multiple biological entities. This application proposes a new method for drug repositioning research based on convolutional neural networks to predict drug-disease associations. First, data on four biological entities are collected from multiple datasets, and a heterogeneous network of drug-gene-protein-disease containing direct and indirect associations between drugs, genes, proteins, and diseases is constructed, and meta-paths are established based on the topological information of biological entities. In addition, initial features of drugs and diseases are established by calculating similarities, and the interactions between gene nodes and between protein nodes are used as initial features of genes and proteins. Next, based on a feature extraction model containing a fully connected network (FC), a graph convolutional network (GCN), a relational graph convolutional network (RGCN), and a multi-layer perceptron (MLP), the representations of each node are learned from the similarity and association data of these entities, and the node representations of biological entities are updated. Finally, the association score (represented by cosine similarity) of the final embeddings (integrated features) of drugs and diseases is calculated to accurately predict drug-disease associations.
[0108] The present application also provides an application scenario which applies the above-mentioned drug repositioning prediction method based on graph convolutional neural network. Specifically: The drug repositioning prediction method based on graph convolutional neural network provided in this embodiment can be applied in the drug repositioning research scenario. This scenario includes a data collection link, a drug repositioning prediction link, and a drug repositioning research link; the data collection link is used to collect data sets from multiple sources, and the data sets include several drugs, several genes, several proteins, several diseases, the basic characteristics of each drug, the basic characteristics of each gene, the basic characteristics of each protein, the basic characteristics of each disease, the interaction relationships between drugs and genes, the interaction relationships between drugs and proteins, the interaction relationships between drugs and diseases, the interaction relationships between diseases and genes, and the interaction relationships between diseases and proteins; the drug repositioning prediction link is used to perform drug repositioning prediction based on the data sets collected from multiple sources; the drug repositioning research link is used to conduct relevant experimental research on the applicability of relevant drugs to a certain disease according to the prediction results of the drug repositioning prediction link.
[0109] Based on the same inventive concept, an embodiment of the present application also provides a drug repositioning prediction device based on graph convolutional neural network for implementing the above-mentioned drug repositioning prediction method based on graph convolutional neural network. The solution provided by this device for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the drug repositioning prediction device based on graph convolutional neural network provided below can refer to the limitations on the drug repositioning prediction method based on graph convolutional neural network in the above text, and will not be elaborated here.
[0110] In an exemplary embodiment, as Figure 5 shown, a drug repositioning prediction device based on graph convolutional neural network is provided, including:
[0111] A data acquisition module M1, configured to acquire information of a multi-source data set; the information of the multi-source data set includes the basic characteristics of different types of biological entities and the interaction relationships between biological entities; the biological entities include drugs, genes, proteins, and diseases. Specifically, the information of the multi-source data set includes several drugs, several genes, several proteins, several diseases, the basic characteristics of each drug, the basic characteristics of each gene, the basic characteristics of each protein, the basic characteristics of each disease, the interaction relationships between drugs and genes, the interaction relationships between drugs and proteins, the interaction relationships between drugs and diseases, the interaction relationships between diseases and genes, and the interaction relationships between diseases and proteins.
[0112] A feature extraction module M2, which is used to extract features by applying a feature extraction model based on a graph convolutional neural network according to the information of the multi-source data set, and obtain the integrated features of drugs and the integrated features of diseases. The feature extraction model based on the graph convolutional neural network includes a graph convolutional neural network group, a relational graph convolutional network, and a multi-layer perceptron group connected in sequence; the graph convolutional neural network group is used to extract the network features of homogeneous networks composed of the same type of biological entities and heterogeneous networks composed of different types of biological entities according to the multi-source data set; the homogeneous networks include homogeneous networks composed of each drug, homogeneous networks composed of each disease, homogeneous networks composed of each gene, and homogeneous networks composed of each protein; the heterogeneous networks include heterogeneous networks composed of drugs and diseases, heterogeneous networks composed of drugs and genes, heterogeneous networks composed of drugs and proteins, heterogeneous networks composed of diseases and genes, and heterogeneous networks composed of diseases and proteins; the relational graph convolutional network is used to fuse the network features of the extracted homogeneous networks and heterogeneous networks; the multi-layer perceptron group is used to integrate the features of drugs and the features of diseases according to the features fused by the relational graph convolutional network.
[0113] A drug redirection prediction module M3, which is used to determine the drug redirection prediction results according to the integrated features of drugs and the integrated features of diseases.
[0114] For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and for the related parts, please refer to the description in the method section.
[0115] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 6 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the information of the multi-source data set, the feature extraction model based on the graph convolutional neural network, and the drug redirection prediction results of each drug. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a drug redirection prediction method based on a graph convolutional neural network.
[0116] Those skilled in the art can understand,Figure 6 The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements. In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0117] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0118] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0119] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0120] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the various embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0121] The databases involved in the various embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the various embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0122] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the various technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0123] In this article, specific examples are used to elaborate on the principles and implementation methods of this application. The descriptions of the above embodiments are only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A drug repositioning prediction method based on a graph convolutional neural network, characterized in that Including: Obtaining information of a multi-source dataset; the information of the multi-source dataset includes the basic characteristics of different types of biological entities and the interaction relationships between biological entities; biological entities include drugs, genes, proteins, and diseases. According to the information of the multi-source dataset, applying a feature extraction model based on a graph convolutional neural network for feature extraction to obtain the integrated features of drugs and the integrated features of diseases; the feature extraction model based on a graph convolutional neural network includes a graph convolutional neural network group, a relational graph convolutional network, and a multi-layer perceptron group connected in sequence; the graph convolutional neural network group is used to extract the network features of a homogeneous network composed of the same type of biological entities and a heterogeneous network composed of different types of biological entities according to the multi-source dataset; the relational graph convolutional network is used to fuse the extracted network features of the homogeneous network and the heterogeneous network; the multi-layer perceptron group is used to integrate the features of drugs and the features of diseases according to the features fused by the relational graph convolutional network. Determining the redirection prediction results of each drug according to the integrated features of drugs and the integrated features of diseases.
2. The method for predicting drug repositioning based on a graph convolutional neural network according to claim 1, wherein The graph convolutional neural network group includes 7 graph convolutional neural networks and 4 fully connected networks. The 7 graph convolutional neural networks are respectively denoted as the first graph convolutional neural network to the seventh graph convolutional neural network in sequence; the 4 fully connected networks are respectively denoted as the first fully connected network to the fourth fully connected network in sequence. The input of the first graph convolutional neural network is the relationship matrix between drugs and the basic characteristics of drugs; the inputs of the first fully connected network to the fourth fully connected network are the basic characteristics of drugs, the basic characteristics of genes, the basic characteristics of proteins, and the basic characteristics of diseases respectively. The input of the second graph convolutional neural network is connected to the output of the first fully connected network and the output of the second fully connected network; the input of the second graph convolutional neural network further includes the interaction relationship between drugs and genes. The input of the third graph convolutional neural network is connected to the output of the first fully connected network and the output of the third fully connected network; the input of the third graph convolutional neural network further includes the interaction relationship between drugs and proteins. The input of the fourth graph convolutional neural network is connected to the output of the first fully connected network and the output of the fourth fully connected network; the input of the fourth graph convolutional neural network further includes the interaction relationship between drugs and diseases. The input of the fifth graph convolutional neural network is connected to the output of the fourth fully connected network and the output of the second fully connected network; the input of the fifth graph convolutional neural network further includes the interaction relationship between diseases and genes. The input of the sixth graph convolutional neural network is connected to the output of the fourth fully connected network and the output of the third fully connected network; the input of the sixth graph convolutional neural network further includes the interaction relationship between diseases and proteins. The input of the seventh graph convolutional neural network is the relationship matrix between diseases and the basic characteristics of diseases. The outputs of the first graph convolutional neural network to the seventh graph convolutional neural network are all connected to the input of the relational graph convolutional network.
3. The method for predicting drug redirection based on a graph convolutional neural network according to claim 2, wherein The multi-layer perceptron group includes 2 multi-layer perceptrons; the 2 multi-layer perceptrons are respectively denoted as the first multi-layer perceptron and the second multi-layer perceptron in sequence; the inputs of the first multi-layer perceptron and the second multi-layer perceptron are both connected to the output of the relational graph convolutional network. The input of the first multi-layer perceptron is the global node representation of the drug output by the relational graph convolutional network; The input of the second multi-layer perceptron is the global node representation of the disease output by the relational graph convolutional network; The output of the first multi-layer perceptron is the integrated feature of the drug; The output of the second multi-layer perceptron is the integrated feature of the disease.
4. The method for predicting drug redirection based on a graph convolutional neural network according to claim 3, wherein According to the information in the multi-source dataset, apply a feature extraction model based on a graph convolutional neural network for feature extraction to obtain the integrated features of drugs and diseases, specifically including: Calculate the similarity between drugs according to the basic features of each drug, and determine the relationship matrix between drugs according to the similarity between drugs; Calculate the similarity between diseases according to the basic features of each disease, and determine the relationship matrix between diseases according to the similarity between diseases; Input the relationship matrix between drugs and the basic features of each drug into the first graph convolutional neural network to obtain drug node representations; input the relationship matrix between diseases and the basic features of each disease into the seventh graph convolutional neural network to obtain disease node representations; Input the basic features of each drug into the first fully connected network to obtain the extracted features of each drug; input the basic features of each gene into the second fully connected network to obtain the extracted features of each gene; input the basic features of each protein into the third fully connected network to obtain the extracted features of each protein; input the basic features of each disease into the fourth fully connected network to obtain the extracted features of each disease; Input the extracted features of each drug, the extracted features of each gene, and the interaction relationship between drugs and genes into the second graph convolutional neural network to obtain the node representations of drugs and genes respectively; input the extracted features of each drug, the extracted features of each protein, and the interaction relationship between drugs and proteins into the third graph convolutional neural network to obtain the node representations of drugs and proteins respectively; input the extracted features of each drug, the extracted features of each disease, and the interaction relationship between drugs and diseases into the fourth graph convolutional neural network to obtain the node representations of drugs and diseases respectively; input the extracted features of each disease, the extracted features of each gene, and the interaction relationship between diseases and genes into the fifth graph convolutional neural network to obtain the node representations of diseases and genes respectively; input the extracted features of each disease, the extracted features of each protein, and the interaction relationship between diseases and proteins into the sixth graph convolutional neural network to obtain the node representations of diseases and proteins respectively; Input the drug node representations, the node representations of drugs and genes, the node representations of drugs and proteins, the node representations of drugs and diseases, the node representations of diseases and genes, the node representations of diseases and proteins, and the disease node representations into the relational graph convolutional network to obtain the global node representations of drugs, genes, proteins, and diseases respectively; Input the features of the drug that interact with diseases, genes, and proteins in the global node representation of the drug into the first multi-layer perceptron for feature integration to obtain the integrated features of the drug; input the features of the disease that interact with drugs, genes, and proteins in the global node representation of the disease into the second multi-layer perceptron for feature integration to obtain the integrated features of the disease.
5. The drug repositioning prediction method based on graph convolutional neural network according to claim 4, wherein Calculate the similarity between drugs according to the basic features of each drug, specifically including: Determine the drug association graph of each drug according to the basic features of each drug; Determine the maximum common subgraph of every two drugs according to the drug association graphs of every two drugs; Calculate the Jaccard coefficient according to the maximum common subgraph and the drug association graph of every two drugs to obtain the similarity between every two drugs.
6. The drug repositioning prediction method based on graph convolutional neural network according to claim 4, wherein Calculate the similarity between diseases according to the basic features of each disease, specifically including: For any two diseases, calculate the similarity between the two diseases according to the semantic contribution value of the common parent disease of the two diseases and the semantic contribution values of the two diseases.
7. The method for predicting drug repositioning based on a graph convolutional neural network according to claim 1, wherein Determine the redirection prediction results of each drug according to the integrated features of the drug and the integrated features of the disease, specifically including: Calculate the cosine similarity between the integrated features of each drug and the integrated features of each disease; Sort the cosine similarities to determine the redirection prediction results of each drug.
8. A drug repositioning prediction device based on a graph convolutional neural network, characterized in that, The drug redirection prediction device based on the graph convolutional neural network includes: A data acquisition module for acquiring information of a multi-source data set; the information of the multi-source data set includes the basic features of different types of biological entities and the interaction relationships between biological entities; biological entities include drugs, genes, proteins, and diseases; A feature extraction module for extracting the integrated features of drugs and the integrated features of diseases by applying a feature extraction model based on the graph convolutional neural network according to the information of the multi-source data set; the feature extraction model based on the graph convolutional neural network includes a graph convolutional neural network group, a relational graph convolutional network, and a multi-layer perceptron group connected in sequence; the graph convolutional neural network group is used to extract the network features of the homogeneous network composed of the same type of biological entities and the heterogeneous network composed of different types of biological entities according to the multi-source data set; the relational graph convolutional network is used to fuse the extracted network features of the homogeneous network and the heterogeneous network; the multi-layer perceptron group is used to integrate the features of drugs and the features of diseases according to the features fused by the relational graph convolutional network; A drug redirection prediction module for determining the redirection prediction results of each drug according to the integrated features of the drug and the integrated features of the disease.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the drug redirection prediction method based on the graph convolutional neural network according to any one of claims 1-7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the drug redirection prediction method based on the graph convolutional neural network according to any one of claims 1-7.