Biological network link prediction algorithm SSL-Bio based on self-supervised graph representation learning
By using a self-supervised graph representation learning algorithm SSL-Bio in biological network link prediction, the graph convolution neural network and Bayesian personalized ranking loss are used to solve the problems of sparseness and high noise in biological networks, which significantly improves the prediction accuracy.
Patent Information
- Application Number
- CN202310675405.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-08
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2043-06-08
AI Technical Summary
The existing biological network link prediction algorithm is affected by sparse biological networks, long-tail distribution and high noise, resulting in low prediction accuracy.
SSL-Bio, a biological network link prediction algorithm based on self-supervised graph representation learning, uses graph convolutional neural network to perform self-supervised learning, generates embedded representations of biological network nodes, and performs multi-task learning through Bayesian personalized ranking loss and self-supervised loss to optimize the model to improve prediction accuracy.
It effectively solves the problems of sparse, long-tail distribution and high noise in biological networks, and improves the accuracy of biological network link prediction.
Smart Images

Figure CN116895326B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of biological network analysis, and in particular relates to a biological network link prediction algorithm SSL-Bio based on self-supervised graph representation learning. Background Art
[0002] As biological research continues to deepen, more and more biological entity relationships are discovered, and biological networks have become a powerful tool for describing these complex biological entity relationships. Biological networks are mainly composed of two parts: biological entity nodes and edges formed by connections between biological entity nodes.
[0003] Biological network analysis is currently a very popular research direction. Mining the relationship between biological entities in biological networks is of great significance. It has many successful applications in the biomedical field [1]. Relationship mining for drug-target interaction (DTI) networks can discover some drug-protein interaction relationships that have not yet been discovered, and then provide these interaction relationships to biologists for verification. This is of great significance for drug research and development, and it can greatly shorten the drug development cycle [2]. Relationship mining for drug-disease association (DDA) networks can discover some currently unclear drug-disease relationships, which is of great significance for drug repositioning and plays a huge role in the field of disease treatment [3]. At the same time, with the development of biotechnology, protein-protein interaction relationships of various species have been discovered [4] (Protein-Protein Interaction, PPI). PPI is very important in the cell system of organisms because proteins are the basis of cell structure and function, and many important cell processes are related to proteins. At the same time, most proteins exert their functions by interacting with other proteins, so accurate prediction of PPI is crucial for cell physiology [5]. Xu et al. [6] constructed a miRNA-miRNA Association (MMA) network. Understanding the interaction between miRNA-miRNA can help us understand the synergy between different miRNAs, thereby further understanding biological processes [7]. Zhao et al. [8] collected the interactions between proteins and metabolites and constructed a protein-metabolite interaction (PMI) network. At the same time, they conducted a series of analyses based on the PMI network, which played an important role in understanding how metabolites regulate protein functions and control various cellular processes. Wishart et al. published the fifth edition of the DrugBank database [9]. The database collects the interactions between multiple drugs and builds a drug-drug interaction (DDI) network based on these interactions. The DDI network often describes adverse reactions between drugs, so understanding the relationships in the DDI network plays a certain guiding role in clinical treatment and medication
[10] . In summary, there are many successful applications of relationship discovery based on biological networks, so research on relationship discovery based on biological networks has great practical significance.
[0004] The essence of the above-mentioned biological network relationship discovery can be regarded as the problem of link prediction between nodes in the network. Link prediction
[11] was first studied on social networks. Its essence is that given a social network, the link prediction algorithm can infer what new interactions may occur between social network members in the near future. As early as 2013, Barzel et al.
[12] had developed an algorithm to perform link prediction on biological networks. At present, link prediction based on biological networks still faces a series of problems. For example, the developed biological network link prediction algorithms are often affected by the sparse and long-tail distribution of biological networks. However, the existing algorithms do not perform targeted modeling based on the characteristics of biological networks. Therefore, link prediction for biological networks still needs further exploration.
[0005] References:
[0006] [1]Peng J, Lu G, Shang XA Survey of Network Representation LearningMethods for Link Prediction in Biological Network[J]. Current pharmaceuticaldesign, 2020, 26(26): 3076-3084.
[0007] [2] Luo Y, Zhao X, Zhou J, et al. A network integration approach for drug-target interaction prediction and computational drug repositioning from heterogeneous information [J]. Nature communications, 2017, 8(1): 1-13.
[0008] [3]Guney E, Menche J, Vidal M, et a1.Network-based in silico drugefficacy screening[J].Nature communications, 2016, 7(1): 1-13.
[0009] [4]Stelzl U,Worm U,Lalowski M,et al.A human protein-proteininteraction network:a resource for annotating the proteome[J].Cell,2005,122(6):957-968.
[0010] [5]Rao V S,Srinivas K,Sujini G N,et al.Protein-protein interactiondetection:methods and analysis[J].International journal of proteomics,2014,2014.
[0011] [6]Xu J,Li C X,Li Y S,et al.MiRNA-miRNA synergistic network:construuction via coregulating functional modules and disease miRNAtopological features[J].Nucleic acids research,2011,39(3):825-836.
[0012] [7]Guo L,Zhao Y,Yang S,et al.Integrative analysis of miRNA-mRNA andmiRNA-miRNA interactions[J].BioMed research international,2014,2014.
[0013] [8]Zhao T,Liu J,Zeng X,et al.Prediction and collection of protein-metabolite interactions[J].Briefings in Bioinformatics,2021.
[0014] [9] Wishart DS, Feunang YD, Guo AC, et al. DrugBank 5.0: a major update to the DrugBank database for 2018[J]. Nucleic acids research, 2018, 46(D1): D1074-D1082.
[0015]
[10] Vilar S, Uriarte E, Santana L, et al.Similarity-based modeling inlarge-scale prediction of drug-drug interactions[J].Nature protocols, 2014, 9(9): 2147-2163.
[0016]
[11] Liben-Nowell D, Kleinberg J. The link-prediction problem for social networks[J]. Journal of the American society for information science and technology, 2007, 58(7): 1019-1031.
[0017]
[12] Barzel B, Barabási A L.Network link prediction by global silencing of indirect correlations[J]. Nature biotechnology, 2013, 31(8): 720-725. Summary of the invention
[0018] In view of the problems existing in the above-mentioned prior art, the main purpose of the present invention is to provide a biological network link prediction algorithm SSL-Bio based on self-supervised graph representation learning. SSL-Bio realizes link prediction for various biological networks through self-supervised learning based on graph convolutional neural networks, which can solve the problem of inaccurate link prediction for biological networks due to sparse biological networks, long-tail distribution, and high noise, and improve the prediction accuracy.
[0019] The purpose of the present invention is achieved through the following technical solutions:
[0020] The present invention provides a biological network link prediction algorithm SSL-Bio based on self-supervised graph representation learning, the method comprising the following steps:
[0021] Get the biological network in, represents the set of nodes in the biological network; ε represents the set of edges in the biological network;
[0022] Generate an embedding representation of a node v in a biological network using a graph convolutional neural network:
[0023]
[0024] in, represents the neighbor set of node v, Indicates the number of nodes numbered v in the node set; represents the neighbor set of node i, Indicates the number of nodes numbered i in the node set; Represents the embedding representation of node v’s neighbor node i at layer l;
[0025] The embedding representations generated using the graph convolutional neural network are summed to generate the feature representation h for node v v :
[0026]
[0027] Where L represents the total number of layers of the graph convolutional neural network in the model; α l represents the weight of node v in the feature summation at layer l, α l =1 / L; Represents the feature representation of node v after the lth convolutional neural network layer;
[0028] Use the inner product to calculate the probability score that there is an edge between node v and node i
[0029]
[0030] in, represents the transpose of the feature representation of node v; h i represents the feature representation of node i;
[0031] The Bayesian personalized ranking loss BPR is used to optimize the model, and the loss function for:
[0032]
[0033] Where N represents the number of nodes in the biological network; Represents the model prediction score of node i that has a relationship with node v; Represents the model prediction score of node j that has a relationship with node i;
[0034] The acquired biological network is enhanced by using data augmentation Generate two views;
[0035] Select nodes u and v in the two generated views and encode them into u using graph convolutional neural network i 、v i ,u i 、v i The self-supervised loss based on contrastive learning is:
[0036]
[0037] Among them, θ uses the cosine similarity function to measure the similarity between two vectors; τ represents the adjustable temperature parameter in contrastive learning; k represents the node different from i in the view; N represents the total number of nodes;
[0038] Compute the contrastive learning-based self-supervised loss for all nodes in the view
[0039]
[0040] Bayesian personalized ranking loss and self-supervised loss Perform multi-task learning, total model loss for:
[0041]
[0042] Among them, λ 1 , 2 are all adjustable parameters, λ 1 The loss range used to control the self-supervised loss, λ 2 Used to control the range of L2 regularization; W represents all trainable parameters in the model.
[0043] As a further description of the above technical solution, in the step of "using data enhancement to obtain the biological network Generate two views”, including generating two views of the acquired biological network by dropping random edges in each round of model training:
[0044]
[0045]
[0046] Among them, Mask 1 、Mask 2 Represents the mask vectors for generating the first and second data augmentation images, Mask1 ∈{0, 1} |ε| , Mask 2 ∈{0, 1} |ε| ; ⊙ represents the Hadamard product operation.
[0047] As a further description of the above technical solution, in the step of "using data enhancement to obtain the biological network Generating Two Views” also includes generating two views of the acquired biological network by adding random edges in each round of model training:
[0048]
[0049]
[0050] Among them, Add 1 、Add 2 Represent the expanded vectors of the first and second data augmentation graphs respectively; Represents an add operation.
[0051] As a further description of the above technical solution, in the step of "using data enhancement to obtain the biological network In the step of “generating two views”, the method further includes generating two views of the acquired biological network by using a random walk method, wherein the random walk method includes a random walk based on edge discarding and a random walk based on edge adding.
[0052] As a further description of the above technical solution, the process of generating two views of the acquired biological network based on the random walk with edge dropping is expressed by the following two formulas:
[0053]
[0054]
[0055] Among them, Mask 1 (l) 、Mask 2 (l) They represent the mask vectors for generating the first and second data augmentation graphs in the lth layer of the graph convolutional neural network respectively;
[0056] The process of generating two views of the acquired biological network based on edge-dropping random walk is expressed by the following two formulas:
[0057]
[0058]
[0059] Among them, Add 1(l) 、Add 2 (l) They respectively represent the expanded vectors for generating the first and second data augmentation graphs in the lth layer of the graph convolutional neural network.
[0060] In summary, the outstanding effects of the present invention are:
[0061] The biological network link prediction algorithm SSL-Bio based on self-supervised graph representation learning provided by the present invention realizes link prediction for various biological networks, can solve the problem of inaccurate biological network link prediction due to sparse biological networks, long-tail distribution, and high noise, and improves the prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without paying any creative work.
[0063] Figure 1 This is a framework diagram of the SSL-Bio model in an embodiment of the present invention;
[0064] Figure 2 A data enhancement method in a biological network in an embodiment of the present invention;
[0065] Figure 3 1 and 2 show the performance of Recall and NDCG under different temperature parameters in the embodiment of the present invention. DETAILED DESCRIPTION
[0066] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0067] In the description of the present invention, it should be noted that the directions or positional relationships indicated by the terms "upper", "middle", "lower", "inner", "outer", etc. are based on the directions or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or component referred to must have a specific direction, be constructed and operated in a specific direction, and therefore cannot be understood as limiting the present invention. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0068] The embodiment of the present invention provides a biological network link prediction algorithm SSL-bio based on self-supervised graph representation learning, the algorithm comprising the following steps:
[0069] Get a given biological network in, represents the set of nodes in the biological network; ε represents the set of edges in the biological network;
[0070] Generate an embedding representation of a node v in a biological network using a graph convolutional neural network:
[0071]
[0072] in, represents the neighbor set of node v, Indicates the number of nodes numbered v in the node set; represents the neighbor set of node i, Indicates the number of nodes numbered i in the node set; Represents the embedding representation of node v’s neighbor node i at layer l.
[0073] Through the above formula, the feature representations of all neighbors of node v can be aggregated to node v in the form of summation. The message aggregation operation here is the core operation of graph convolution.
[0074] Next, the obtained embedding representations are summed to reduce the size of the SSL-bio model, speed up the inference of the SSL-bio model, and generate the final feature representation h for node v. v :
[0075]
[0076] Where L represents the total number of layers of the graph convolutional neural network in the model; α l represents the weight of node v in the feature summation at layer l, α l =1 / L; Represents the feature representation of node v after the lth convolutional neural network layer.
[0077] Through this formula, we can get the representation of node v after the L-layer graph convolutional neural network performs neighbor feature aggregation. The above process is the first part of the SSL-bio model - the graph representation learning part (corresponding to Figure 1 -A), whose main function is to generate an embedded representation for each node in a given biological network using a graph representation learning algorithm based on the topological structure of the biological network.
[0078] After obtaining the feature representation of each node in the biological network, for any two nodes v and i in a given biological network, we use the inner product method to obtain the probability score of the existence of an edge between node v and node i.
[0079]
[0080] in, represents the transpose of the feature representation of node v; h i represents the feature representation of node i.
[0081] After obtaining the score of edges between any two nodes, this embodiment then uses the Bayesian personalized ranking loss BPR to optimize the model. The loss will make the model predict the observed edges higher than the unobserved edges. The loss is expressed as follows:
[0082]
[0083] Where N represents the number of nodes in the biological network; Represents the model prediction score of node i that has a relationship with node v; Represents the model prediction score of node j that has a relationship with node i.
[0084] The above process is the second part of the SSL-bio model, the supervision loss part (corresponding to Figure 1 -B), using known biological entity interactions to guide the model to learn embedded representations of biological entities.
[0085] Then we move on to the third part of the SSL-bio model, the self-supervised learning part (corresponding to Figure 1 -C), firstly, the acquired biological network is enhanced by using data augmentation Generate two views, and then select nodes u and v in the two generated views and encode them into u using the graph convolutional neural network i 、v i It is worth mentioning that this part completely shares the parameters of the graph neural network in the first part included in the SSL-bio model.
[0086] See also Figure 2 In this part, this embodiment adopts four data enhancement methods to solve the problem of poor node representation learning quality caused by the long-tail distribution, sparseness and noise characteristics of biological networks. The circular nodes and triangular nodes represent entity nodes in the network, and the solid lines in the figure represent the edges between entity nodes in the biological network. Figure 2 -a and Figure 2 The dotted lines in -c represent the edges that will be removed from the biological network. Figure 2-b and Figure 2 The dotted lines in -d represent the edges that will be removed in the biological network. In order to better describe the process of data enhancement, an adjacency matrix A is used here to represent the given biological network. , and assume that in the biological network link prediction problem to be solved, the given biological network is unweighted, that is, A∈{0,1}.
[0087] like Figure 2 As shown in Figure 1, in each round of model training, two views are generated for the acquired given biological network by randomly dropping edges (ED) to resist the noise of edges in the biological network, thereby generating better node representations for link prediction of the biological network. The process can be expressed using the following two formulas:
[0088]
[0089]
[0090] Among them, Mask 1 、Mask 2 Represents the mask vectors for generating the first and second data augmentation images, Mask 1 ∈{0, 1} |ε| , Mask 2 ∈{0, 1} |ε| ; ⊙ represents the delete operation.
[0091] Through the above Mask 1 and Mask 2 The operation of two masking vectors can be used to Generate two subgraphs and The adjacency matrix corresponding to the two subgraphs is A 1 and A 2 In the process of generating subgraphs, the parameter p is used to control the percentage of network edges to be deleted. If p = 0.1, it means Mask 1 and Mask 2 There are 10% 0 values in the vector and the remaining 90% are 1, which means that 10% of the edges in a given biological network are deleted.
[0092] like Figure 2 -b, the second method is to generate two views of the acquired given biological network by adding random edges (Edge Add, EA) in each round of model training, and then add the generated edges to the biological network to resist the sparsity in the biological network, thereby generating better node representations for link prediction of the biological network. The process can be expressed by the following two formulas:
[0093]
[0094]
[0095] Among them, Add 1 、Add 2 Represent the expanded vectors of the first and second data augmentation graphs respectively; Represents an add operation.
[0096] Through the above Add 1 and Add 2 The operation of two expanded vectors can be used to Generate two subgraphs and The adjacency matrix corresponding to the two subgraphs is A 1 and A 2 In the process of generating subgraphs, the parameter p is used to control the percentage of increasing the number of network edges. If p = 0.1, it means 0.1 times the number of edges in the original biological network.
[0097] The first and second enhancement methods mentioned above use the same subgraph for message aggregation for the same node, but when such a strategy is implemented, the representation learned for a given node is single because it has fixed neighboring nodes. Therefore, in this embodiment, multiple subgraphs are generated for each node, and the third and fourth enhancement methods are summarized into one, called Random Walk (RW), which is combined with the above-mentioned random edge discarding and random edge addition to generate two views of the acquired biological network. The random walk method includes the following: Figure 2 -c shows the random walk with edge drop (RWED) and Figure 2 -d shows the random walk with edge add (RWEA), where
[0098] The process of generating two views of the acquired biological network based on edge-dropping random walk is expressed by the following two formulas:
[0099]
[0100]
[0101] Among them, Mask 1 (l) 、Mask 2 (l)They represent the mask vectors for generating the first and second data augmentation graphs in the lth layer of the graph convolutional neural network respectively;
[0102] The process of generating two views of the acquired biological network based on edge-dropping random walk is expressed by the following two formulas:
[0103]
[0104]
[0105] Among them, Add 1 (l) 、Add 2 (l) They respectively represent the expanded vectors for generating the first and second data augmentation graphs in the lth layer of the graph convolutional neural network.
[0106] In the process of data enhancement in this embodiment, Figure 2 As shown in the figure, Layer1 represents the first layer of node embedding representation for biological networks using graph convolutional neural networks, and the same applies to Layer2 and Layer3. Taking the original biological network given in the figure as an example, the same biological network is used in Layer1, Layer2 and Layer3 to learn the embedding representation of nodes. Similarly, the ED and EA data enhancement methods also use the same biological network in Layer1, Layer2 and Layer3 to learn the embedding representation of nodes. However, in the process of using graph convolutional neural networks to generate embedding representations for nodes, the two data enhancement methods RWED and RWEA use different biological networks in Layer1, Layer2 and Layer3 to learn the embedding representation of nodes. By doing so, for a given node, it can aggregate richer node information, thereby enhancing the robustness of its node representation.
[0107] Finally, the fourth part of the SSL-bio model, the design of self-supervised loss and multi-task learning (corresponding to Figure 1 Specifically, in this section, the self-supervised contrastive learning loss is performed on the nodes after GCN encoding, u i 、v i The self-supervised loss based on contrastive learning is:
[0108]
[0109] Among them, θ uses the cosine similarity function to measure the similarity between two vectors; τ represents the adjustable temperature parameter in contrastive learning; k represents the nodes in the view that are different from i (i.e., negative samples); and N represents the total number of nodes.
[0110] In this formula, the denominator consists of three parts: e θ (u i , v i ) / τ describes the similarity measurement results of the same node in two views (positive sample pairs), It is the sum of the similarity metrics of all negative sample pairs in the inter-view. is the sum of the similarity metrics of all negative sample pairs in intra-view. Specifically, Figure 1 -C shows that in the two generated views, the same node is used as a positive sample. In the learning process, it is expected that the representation between positive samples is maximized. Negative samples come from two parts. One part comes from the same view (intra-view). Specifically, it is considered that in the same view, all nodes except the u node are its negative samples. The other part comes from another view (inter-view). Specifically, it is considered that in another view, all nodes except the v node are its negative samples. For all negative samples, it is hoped that the representation between negative sample pairs is minimized through contrast loss.
[0111] Next, the self-supervised loss based on contrastive learning can be calculated for all nodes in the view, and the total loss can be obtained by summing the losses of the nodes. That is, the total self-supervised loss of the model based on contrastive learning
[0112]
[0113] In multi-task learning, each task can promote each other, thereby improving the performance of the model in downstream tasks. This embodiment is based on Bayesian personalized ranking loss and self-supervised loss Perform multi-task learning, total model loss for:
[0114]
[0115] Among them, λ 1 , 2 are all adjustable parameters, λ 1 The loss range used to control the self-supervised loss, λ 2 Used to control the range of L2 regularization; W represents all trainable parameters in the model.
[0116] The above is the overall process of the SSL-Bio method, and the specific algorithm is shown in Table 1 below.
[0117] Table 1 Specific algorithms of SSL-Bio
[0118]
[0119]
[0120] Experimental results and analysis:
[0121] (I) Dataset
[0122] Specifically, in this embodiment, multiple known biological network data sets are used to evaluate the SSL-Bio model. The data sets used and the types of biological entities in the biological network included in each data set, the number of each biological entity, the total number of edges in the biological network, and the sparsity rate in the biological network are shown in Table 2 below.
[0123] Table 2 Dataset
[0124]
[0125] In the above table, the XueDTI dataset comes from the drug-protein interaction (DTI) network studied by Xue et al.; the SnapDiG dataset, SnapDrG dataset, and SnapDrS dataset come from the disease-gene association (DiG) network, drug-gene association (DrG), and drug-side effect (DrS) association network compiled by the Snap research group of Stanford University and others.
[0126] 2. Evaluation indicators
[0127] Since the link prediction task of biological networks is similar to that of recommendation systems, this embodiment uses a series of recommendation system indicators as evaluation indicators, which are: AUROC, AUPR, Recall@N, NDCG@N, MRR@N, MAP@N, and N is set to 20. More specifically, it is assumed that the biological network contains two biological entities (biological entity A and biological entity B). For a given biological entity A, through the SSL-Bio model, we calculate the probability prediction value of predicting that all biological entities B have edges with biological entity A, and then Recall@N evaluates the total number of the top N predicted values and the ratio of the known number of relationships. NDCGO@N is an evaluation indicator that takes order factors into consideration. The more top-ranked predicted values are successfully predicted, the higher the result. MRR@N (Mean Reciprocal Rank, MRR) evaluates the reciprocal of the first position of the real edge in the sorted list in the ranking of the top N predicted values, and then calculates this reciprocal value for all entities, and finally averages it. MAP@N (Mean Average Precision, MAP) is the precision value of the first N entities A in the ranking, and then the average is calculated.
[0128] Specifically, in this embodiment, in order to ensure that the distribution of training data and test data is consistent, taking the drug-protein interaction network as an example, given the interaction between any drug and all proteins, 70% of the edges are selected as training and the remaining edges are used as testing; the early stopping strategy is adopted to prevent the model from overfitting. When the Recall@20 index does not increase for 100 consecutive epochs, the model training is stopped.
[0129] (III) Experimental results
[0130] In this embodiment, the biological network link prediction algorithm SSL-Bio based on self-supervised graph representation learning is compared with the network representation learning algorithms laplacian, GRaRep, SVD, GF, and HOPE based on matrix decomposition, the network representation learning algorithms Deepwalk, node2vec, and struc2vec based on random walk, and the network representation learning algorithms LINE, SDNE, and GAE based on neural networks on four data sets XueDTI, SnapDiG, SnapDrG, and SnapDrS. The comparison results are shown in Tables 3, 4, 5, and 6, respectively.
[0131] Experimental data show that SSL-Bio performed best in all six indicators when evaluated on the XueDTI, SnapDrG, and SnapDrS datasets. Only when evaluated on the SnapDiG dataset did SSL-Bio rank second in the AUROC indicator, and it still performed best in the other indicators.
[0132] Table 3 Experimental results of XueDTI dataset
[0133]
[0134]
[0135] Table 4 Experimental results of SnapDiG dataset
[0136]
[0137] Table 5 Experimental results of SnapDrG dataset
[0138]
[0139]
[0140] Table 6 Experimental results of SnapDrS dataset
[0141]
[0142] (IV) The impact of different data enhancement methods on the model
[0143] In this embodiment, the four models, namely, the edge discarding data enhancement method ED based on biological networks, the edge adding data enhancement method EA based on biological networks, the random walk based on edge discarding RWED and the random walk based on edge adding RWEA, are named SSL-Bio-ED, SSL-Bio-EA, SSL-Bio-RWED and SSL-Bio-RWEA respectively. The experimental results of the four models on various data sets are shown in Table 7, which proves that the four enhancement methods used in this embodiment are all effective and can play a good auxiliary role in the link prediction task of biological networks.
[0144] Table 7 Evaluation index results of four data augmentation methods on different datasets
[0145]
[0146] 5. The impact of introducing self-supervised representation learning on link prediction tasks based on biological networks
[0147] The loss function of the SSL-Bio model is This assessment will Remove it from the model, that is, This verifies the impact of the introduction of self-supervised representation learning on model performance. The corresponding model is named Remove-SSL. This ablation study is performed on all datasets. The experimental results are shown in Table 8:
[0148] On the XueDTI dataset, compared with Remove-SSL, SSL-Bio's Recall, NDCG, MAP, and MRR indicators increased by 1.87%, 3.52%, 3.61%, and 6.78%, respectively. In terms of AUROC and AUPR, AUROC decreased slightly, while AUPR increased by up to 31.20%. It is worth noting that on the positive and negative sample imbalance dataset, AUPR can better reflect the quality of the model than AUROC, and we should pay more attention to the improvement of the AUPR indicator. On the SnapDiG dataset, compared with Remove-SSL, SSL-Bio's Recall, NDCG, MAP, MRR, and AUPR indicators increased by 3.33%, 1.85%, 0.92%, 0.09%, and 6.69%, respectively. On the SnapDrG dataset, compared with Remove-SSL, the Recall, NDCG, MAP, MRR, AUROC, and AUPR of SSL-Bio were improved by 1.35%, 1.01%, 0.75%, 1.23%, 0.06%, and 9.49%, respectively. On the SnapDrS dataset, compared with Remove-SSL, the Recall, NDCG, MAP, MRR, AUROC, and AUPR of SSL-Bio were improved by 3.32%, 2.21%, 2.14%, 1.41%, 0.38%, and 9.19%, respectively. This shows that the introduction of self-supervised representation learning has a significant effect on improving the link prediction performance of biological networks.
[0149] Table 8 Comparison of experimental results of Remove-SSL model and SSL-Bio model
[0150]
[0151] (VI) Effect of temperature parameter on SSL-Bio model in self-supervised representation learning
[0152] Specifically, in this embodiment, the influence of a smaller temperature parameter (0.1) and a larger temperature parameter (1.0) on the model is tested. Figure 3, Recal and NDCG achieve the best performance when the temperature parameter is equal to 0.2 on the XueDTI dataset, and then the performance decreases as the temperature parameter increases. On the SnapDiG dataset, both Recall and NDCG achieve the best performance when the temperature parameter is 0.1. On the SnapDrG dataset, both Recall and NDCG achieve the best performance when the temperature parameter is 0.5. On the SnapDrS dataset, both Recall and NDCG achieve the best performance when the temperature parameter is 0.1. Therefore, in the link prediction problem of biological networks, a smaller temperature parameter can often achieve better results.
[0153] (VII) The influence of parameter p in data enhancement process
[0154] Specifically, in this embodiment, when evaluating the impact of different data augmentation methods on the model, for the models SSL-Bio-ED and SSL-Bio-RWED, the main operation is to discard random edges from the original biological network, and the discard ratio is expressed by the parameter p 1 In order not to destroy the original biological network, the parameter p 1 ∈{0.1, 0.2, 0.3, 0.4, 0.5}. For SSL-Bio-EA and SSL-Bio-RWEA, the main operation is to add random edges to the original biological network, and the increase ratio is expressed by the parameter p 2 In order not to introduce too much noise into the original biological network, the parameter p 2 ∈{0.1, 0.2, 0.3, 0.4, 0.5}, thus, the Recall and NDCG performances of different data sets under different parameters p are as follows:
[0155] When using ED data enhancement, the Recall and NDCG indicators are both in p 1 = 0.2; on the SnapDiG dataset, the Recall index and NDCG index are both within p 1 =0.2; on the SnapDrG dataset, the Recall index and NDCG index are respectively 1 = 0.2 and p 1 = 0.5; on the SapDrS dataset, the Recall index and NDCG index are both within p 1 The best value is achieved when =0.5.
[0156] When using EA data augmentation, the Recall and NDCG indicators are both in p 2= 0.1; on the SnapDiG dataset, the Recall index and NDCG index are both within p 2 = 0.3; on the SnapDrG dataset, the Recall index and NDCG index are both within p 2 = 0.1; on the SnapDrS dataset, the Recall index and NDCG index are respectively 2 = 0.5. In most cases, when a large number of edges are deleted or added, the structure of the original biological network is affected and the learning of node representation is affected. Therefore, when using different data enhancement methods for self-supervised learning, it is possible to 1 ∈{0.1, 0.2, 0.3, 0.4, 0.5} and p 2 ∈{0.1, 0.2, 0.3, 0.4, 0.5} for parameter adjustment to achieve the best performance of the SSL-Bio model.
[0157] Finally, it should be noted that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, it is still possible for those skilled in the art to modify the technical solutions described in the aforementioned embodiments or to make equivalent substitutions for some of the technical features therein. Any changes, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A biological network link prediction algorithm SSL-Bio based on self-supervised graph representation learning, characterized in that: The method comprises the following steps: Get the biological network in, represents the set of nodes in the biological network; ε represents the set of edges in the biological network; Generate embedding representations of nodes in biological networks using graph convolutional neural networks: in, represents the neighbor set of node v, Indicates the number of nodes numbered v in the node set; represents the neighbor set of node i, Indicates the number of nodes numbered i in the node set; Represents the embedding representation of node v’s neighbor node i at layer l; The embedding representations generated using the graph convolutional neural network are summed to generate the feature representation h for node v v : Where L represents the total number of layers of the graph convolutional neural network in the model; α l represents the weight of node v in the feature summation at layer l, α l =1 / L; Represents the feature representation of node v after the lth convolutional neural network layer; Use the inner product to calculate the probability score that there is an edge between node v and node i in, represents the transpose of the feature representation of node v; h i represents the feature representation of node i; The Bayesian personalized ranking loss BPR is used to optimize the model, and the loss function for: Where N represents the number of nodes in the biological network; Represents the model prediction score of node i that has a relationship with node v; Represents the model prediction score of node j that has a relationship with node i; The acquired biological network is enhanced by using data augmentation Generate two views; Select nodes u and v in the two generated views and encode them into u using graph convolutional neural network i 、v i ,u i 、v i The self-supervised loss based on contrastive learning is: Among them, θ uses the cosine similarity function to measure the similarity between two vectors; τ represents the adjustable temperature parameter in contrastive learning; k represents the node different from i in the view; N represents the total number of nodes; Compute the contrastive learning-based self-supervised loss for all nodes in the view Bayesian personalized ranking loss and self-supervised loss Perform multi-task learning, total model loss for: Among them, λ1 and λ2 are adjustable parameters, λ1 is used to control the loss range of self-supervised loss, and λ2 is used to control the range of L2 regularization; W represents all trainable parameters in the model.
2. The biological network link prediction algorithm SSL-Bio based on self-supervised graph representation learning according to claim 1, characterized in that: In step "using data enhancement to obtain the biological network Generate two views”, including generating two views of the acquired biological network by dropping random edges in each round of model training: Among them, Mask1 and Mask2 represent the mask vectors for generating the first and second data augmentation images, respectively, Mask1∈{0,1} |ε| ,Mask2∈{0,1} |ε| ; ⊙ represents the Hadamard product operation.
3. The biological network link prediction algorithm SSL-Bio based on self-supervised learning graph representation learning according to claim 2, characterized in that: In step "using data enhancement to obtain the biological network Generating Two Views” also includes generating two views of the acquired biological network by adding random edges in each round of model training: Among them, Add1 and Add2 represent the expansion vectors of the first and second data enhancement graphs respectively; Represents an add operation.
4. The biological network link prediction algorithm SSL-Bio based on self-supervised graph representation learning according to claim 3, characterized in that: In step "using data enhancement to obtain the biological network In the step of “generating two views”, the method further includes generating two views of the acquired biological network by using a random walk method, wherein the random walk method includes a random walk based on edge discarding and a random walk based on edge adding.
5. The biological network link prediction algorithm SSL-Bio based on self-supervised graph representation learning according to claim 4, characterized in that: The process of generating two views of the acquired biological network based on edge-dropping random walk is expressed by the following two formulas: Among them, Mask1 (l) 、Mask2 (l) They represent the mask vectors for generating the first and second data augmentation graphs in the lth layer of the graph convolutional neural network respectively; The process of generating two views of the acquired biological network based on edge-dropping random walk is expressed by the following two formulas: Among them, Add1 (l) 、Add2 (l) They respectively represent the expanded vectors for generating the first and second data augmentation graphs in the lth layer of the graph convolutional neural network.
Citation Information
Patent Citations
MOOC recommendation method based on graph convolutional neural network
CN114154070A
KR20220160407A