Method for Mining Protein Functional Modules, Computer Device, and Storage Medium
Through the node-level adaptive graph convolution network model and K-means clustering algorithm, combined with the weighted adjacency matrix, the mining problem of complexes and signal pathways in the protein interaction network in the prior art is solved, and the recognition accuracy of protein functional modules is improved.
Patent Information
- Application Number
- CN202111649446.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-29
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-12-29
AI Technical Summary
The prior art is difficult to effectively mine protein complexes and signaling pathways from protein interaction networks, especially to ignore protein signaling pathways in sparse states.
The node-level adaptive graph convolutional network model is adopted to screen and expand protein complexes and signaling pathways by learning higher-order and lower-order neighbor information, combining K-means clustering algorithm and weighted adjacency matrix.
It improves the recognition accuracy of protein functional modules, and the generated protein complexes and signal pathways are more in line with the real situation, improving the understanding of protein functions and disease mechanism analysis capabilities.
Smart Images

Figure CN116417060B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of systems biology, and particularly relates to a method for mining protein function modules, a computer device, and a computer-readable storage medium. Background Art
[0002] Protein functions are often manifested through the interactions between proteins or nucleic acids. Such interactions exist in the life activities of body cells and intersect with each other to form a protein-protein interaction (PPI) network. In a PPI network, a set of proteins that complete a specific molecular process through interactions is called a protein function module, such as a protein complex or a protein signaling pathway. The mining of protein function modules can not only help understand the functional organizational structure of cells and the ways of performing physiological functions, but also contribute to people's understanding of various biological processes, revealing the occurrence mechanisms of diseases, and finding new drug targets. Therefore, mining the functional modules of protein interactions has important significance.
[0003] Based on the protein interaction network, existing technologies often predict protein functions by identifying protein complexes composed of multiple small protein molecules. A protein complex is composed of multiple small protein molecules tightly connected together and interacting pairwise, presenting a dense state in the network structure. Often, through a shallow graph neural network algorithm, relevant structural information can be well learned to identify protein complexes. However, in the protein interaction network, there are still multiple small protein molecules presenting a chain-like form, forming a protein signaling pathway. Compared with complexes, the protein signaling pathway is in a sparse state and requires higher-order information to be identified in the network structure.
[0004] The article "Protein complexes identification based on go attributed network embedding[J].2018, Bo Xu" discloses a method for identifying protein complexes. It learns protein node representations based on the accelerated attributed network embedding model (AANE), finds all maximal clique structures (with three or more protein nodes) through the protein interaction network, calculates the density of the maximal cliques based on the similarity of protein node representations, and iterates multiple times. Each time, the maximal clique structure is expanded, and according to the principle of the overall density increase after adding a new protein node, protein complexes are obtained. However, this method learns the structural information of the first-order neighbors of protein nodes based on the accelerated attributed network embedding model (AANE), and is only suitable for mining protein complexes in dense subgraphs, while ignoring the protein signaling pathways existing in the protein interaction network. Both are important bases for discovering protein functions. Summary of the Invention
[0005] In view of this, the present invention provides a method for mining protein functional modules, a computer device, and a computer-readable storage medium, so as to solve the problem of how to mine protein complexes and protein signaling pathways from a protein interaction network.
[0006] To solve the above technical problems, one aspect of the present invention is to provide a method for mining protein functional modules, which includes the steps of:
[0007] S1. Input the protein interaction network into a node-level adaptive graph convolutional network model for learning and training, so that each protein node learns high-order and low-order neighbor information, and obtain a protein node vector representation;
[0008] S2. Based on the protein node vector representation, perform clustering through the K-means clustering algorithm to obtain a soft label of the clustering result of the protein nodes. Set a loss function according to the clustering result soft label and perform backpropagation to update the network parameters of the model;
[0009] S3. Perform iterative calculations based on the above steps S1 to S2 until the model converges or reaches the maximum number of iterations, and obtain the final protein node vector representation and clustering result of the last iterative calculation;
[0010] S4. Based on the final protein node vector representation, calculate the similarity of protein nodes through a cosine similarity calculation formula, and construct a weighted adjacency matrix;
[0011] S5. Screen out the basic structure of the protein complex from the protein interaction network and expand it based on the calculation of the weighted adjacency matrix to obtain the protein complex;
[0012] S6. Screen out the basic structure of the protein signaling pathway from each clustering cluster of the clustering result and expand it based on the calculation of the weighted adjacency matrix to obtain the protein signaling pathway.
[0013] In a preferred solution, the step S1 includes:
[0014] S11. Obtain a protein interaction network, and construct a corresponding adjacency matrix A and a gene ontology attribute feature matrix X; wherein, the nodes of the protein interaction network are represented as V = {v1, v2,..., v n}, and the gene ontology attribute feature matrix X = [x1, x2,..., x n T , where n is the total number of protein nodes, the dimension of x i is d, both x and d are positive integers, and i = 1 to n;
[0015] S12. Calculate the normalized Laplacian matrix Ls based on the adjacency matrix A, and construct a low-pass filter G; the normalized Laplacian matrix Ls is Ls = I - D -1 / 2 AD -1 / 2 , and the low-pass filter G is where D is the degree matrix of the adjacency matrix A, D = diag(d1, d2,..., d n ), d i represents the number of edges of node v i , Λ is the diagonal matrix of the eigenvalues of matrix Ls, I is the diagonal matrix with all eigenvalues of matrix Ls being 1, and U is the eigenvector of matrix Ls;
[0016] S13. Set the number of iterative convolution layers k = t, and let t loop from 0 to M to perform the following steps S14 to S18. M represents the maximum value of the number of convolution layers and takes a positive integer value;
[0017] S14. Perform the k-th layer convolution operation on the gene ontology attribute feature matrix X using the low-pass graph filter G to obtain the graph representation G k X of the current convolution layer of the protein-protein interaction network. The calculation formula is as follows:
[0018]
[0019] S15. Based on the graph representation G k X of the current convolution layer, calculate the state values of each protein node in the protein-protein interaction network. The calculation formula is as follows:
[0020]
[0021] where [G k X] i represents the graph representation of node vi during the k-th layer convolution operation, represents the state value of node v i during the (k - 1)-th layer convolution operation, represents the state value of node v i during the k-th layer convolution operation, and S() is a non-linear transformation function;
[0022] S16. Calculate and evaluate the probability value for each node to stop the iterative convolution based on the state values; among them, the calculation formula for the probability value is as follows:
[0023]
[0024] where W h and b h are the learnable network parameters of the node-level adaptive graph convolution network model; σ represents the activation function; Denote the node v i The probability value during the k-th layer convolution operation;
[0025] S17. Set the probability value threshold ε. For each node v i , calculate the cumulative probability value of its first k layers of convolution and compare it with the threshold ε: If the cumulative probability value reaches or exceeds the threshold ε, then for node v i Stop the iterative convolution calculation and record its convolution layer number as N i = k'; If the convolution operation layer k iterates to M and the cumulative probability value is still less than the threshold ε, then for node v i The convolution layer number of is N i = M; The convolution layer number N at which the convolution of the node v i Stops iterating convolution is expressed as follows: i It is expressed as follows:
[0026]
[0027] S18. Calculate and obtain the probability value during the last layer convolution operation of each node v i , and the calculation formula is as follows:
[0028]
[0029] S19. For each node v i , linearly combine the graph representation of its first N i layers of convolution operations with the probability value to obtain the vector representation of the protein node:
[0030]
[0031] In a preferred solution, the step S2 includes:
[0032] S21. Set the number m of clustering clusters, where m is a positive integer;
[0033] S22. Based on the protein node vector representation, perform clustering through the K-means clustering algorithm to obtain the soft label of the clustering result of the protein nodes;
[0034] S23. According to the soft label of the clustering result, set the loss function L as:
[0035] where λ tig represents the first loss coefficient, λ sep represents the second loss coefficient, λ tig and λ sep are both constants, L tig represents the similarity between nodes within the cluster, and L sep represents the similarity between nodes between clusters;
[0036] Among them, L tig The calculation formula is as follows:
[0037]
[0038] Among them, L sep The calculation formula is as follows:
[0039]
[0040] S24. Perform backpropagation according to the loss function to update the network parameters of the node-level adaptive graph convolutional network model;
[0041] The step S3 includes: Based on a preset maximum number of iterations, repeat steps S1 to S2 for iterative calculation until the model converges or reaches the maximum number of iterations. In the last iterative calculation, obtain the final protein node vector representation in step S1, and obtain the final clustering result C = {C1, C2,... C m}, forming m clustering clusters.
[0042] In a preferred solution, the value of the number m of the clustering clusters is set in the following manner:
[0043] Based on the protein interaction network, set m = r, and perform clustering through the K-means clustering algorithm to obtain C = {C1, C2,... C r}; where r takes values from 2 to R, and R is a positive integer;
[0044] For each specific value of m, calculate the sum of squared errors SSE of the distance from each node to the cluster center through the elbow method. The calculation formula is as follows:
[0045] p is a node within the C i cluster, and center i is the center point of the C i cluster;
[0046] According to the corresponding relationship between the specific value of m and the calculated SSE value, fit a curve graph, and determine the inflection point where the decline rate of the SSE value changes from fast to slow in the fitted curve. Select the m value corresponding to the inflection point as the value of the number m of the final clustering clusters.
[0047] In a preferred solution, in the loss function L:
[0048]
[0049] In a preferred solution, the step S4 includes: based on the adjacency matrix A and the protein node vector representation, calculating the similarity between protein nodes v i and v j using the following cosine similarity calculation formula, and constructing a weighted adjacency matrix W:
[0050]
[0051] where a ij is an element of the adjacency matrix A, and w ij is an element of the weighted adjacency matrix W.
[0052] In a preferred solution, the step S5 includes:
[0053] S51. Set and initialize the sets Alternative_core, Complex_Seed_core, and Complex_set;
[0054] S52. Apply the clique mining method to screen out the maximum clique structure Clique q from the protein interaction network, and place the maximum clique structure Clique q into the set Alternativecore;
[0055] S53. Based on the weighted adjacency matrix W, calculate the density scores of all maximum cliques Clique q in the set Alternative_core, and sort them from largest to smallest according to the density scores; where the calculation formula for the density score is:
[0056]
[0057] S54. Remove the maximum clique with the largest density score from the set Alternative_core and place it into the set Complex_Seed_core as the basic structure of the protein complex;
[0058] S55. Traverse the remaining maximum clique structures in the set Alternative_core. When there are overlapping protein nodes between the remaining maximum clique and the protein nodes in the current maximum clique with the largest density score:
[0059] If the number of duplicate nodes is less than 2, delete the duplicate nodes in the remaining maximum clique, and keep the remaining part if the quantity is greater than 3; if the number of duplicate nodes is greater than or equal to 2, do not delete the duplicate nodes;
[0060] S56. Repeat the above steps S53 - S55 until the set Alternative_core is an empty set, and several basic structures of protein complexes are obtained in the set Complex_Seed_core;
[0061] S57. Based on the maximum clique Clique in the set Complex_Seed_core j , for any neighbor node p of the protein nodes in this maximum clique Clique j , calculate the correlation score between the neighbor node p i and this maximum clique Clique i based on the protein node similarity. If the correlation score is greater than the pre - set threshold θ1, then embed the neighbor protein node p j into this maximum clique Clique i ; where the calculation formula of the correlation score is as follows: j
[0062]
[0063] S58. After traversing all the neighbor protein nodes of this maximum clique Clique j , determine a protein complex, remove it from the set Complex_Seed_core and place it into the set Complex_set;
[0064] S59. Repeat the above steps S57 - S58 until the set Complex_Seed_core is an empty set, and the finally mined protein complexes are obtained in the set Complex_set.
[0065] In the preferred solution, the step S6 includes:
[0066] S61. Set and initialize the sets Pathway_Seed_core and Pathway_set;
[0067] S62. Based on the protein - protein interaction network, traverse the m clustering clusters, find the shortest paths between every two nodes within the clusters, and the length of the shortest paths does not exceed 3. Place all the selected paths into the set Pathway_Seed_core as the basic structures of protein signal pathways;
[0068] S63. Based on the weighted adjacency matrix W, calculate the density scores of all the shortest paths shortest_path q in the set Pathway_Seed_core, and sort them from large to small according to the density scores; where the calculation formula of the density score is:
[0069]
[0070] S64. Select the shortest path shortest_path with the highest density from the set Pathway_Seed_core j For the shortest path shortest_path j For any neighbor node p at the end i Calculate the correlation score between the neighbor protein node and this shortest path based on protein node similarity; if the correlation score is greater than the pre-set threshold θ2, then add the neighbor protein node p i to the end of the shortest path shortest_path j .
[0071] Among them, the calculation formula of the correlation score is as follows:
[0072]
[0073] S65. After traversing all neighbor protein nodes at the end of the shortest path shortest_path j determine a protein communication path, remove it from the set Pathwa_Seed_core and place it into the set Pathway_set;
[0074] S66. Repeat the above steps S64 - S65 until the set Pathway_Seed_core is an empty set, and obtain the finally mined protein communication path in the set Pathway_set.
[0075] To solve the above technical problems, the present invention also provides a computer device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. Among them, when the processor executes the program, it implements the steps of the above-mentioned method for mining protein function modules.
[0076] To solve the above technical problems, the present invention also provides a computer-readable storage medium. Among them, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the above-mentioned method for mining protein function modules.
[0077] The method for mining protein functional modules provided by the embodiments of the present invention is based on a node-level adaptive graph convolutional network (NASGC). Through an adaptive mechanism, each protein node learns high-order and low-order neighbor information respectively. In the vector representation information of the protein nodes obtained by learning, high-order and low-order structural information is fused on the gene ontology attribute features of the protein nodes, obtaining a more generalized protein node representation. Thus, protein complexes and protein signaling pathways can be mined from the protein interaction network, and the mined protein complexes are more in line with the actual situation, improving the accuracy of protein function recognition. Brief Description of the Drawings
[0078] Figure 1 It is a workflow diagram of the method for mining protein functional modules in the embodiments of the present invention;
[0079] Figure 2 It is a process diagram of the method for mining protein functional modules in the embodiments of the present invention;
[0080] Figure 3 It is a structural block diagram of a computer device in the embodiments of the present invention. Detailed Embodiments
[0081] To make the objectives, technical solutions, and advantages of the present invention clearer, the detailed embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Examples of these preferred embodiments are illustrated in the accompanying drawings. The embodiments of the present invention shown in the drawings and described according to the drawings are merely exemplary, and the present invention is not limited to these embodiments.
[0082] Here, it should also be noted that in order to avoid obscuring the present invention with unnecessary details, only the structures and / or processing steps closely related to the solution of the present invention are shown in the accompanying drawings, while other details less related to the present invention are omitted.
[0083] Figure 1 It is a schematic flowchart of the method for mining protein functional modules provided by the embodiments of the present invention.
[0084] The method for mining protein functional modules of the present application is applied to a terminal device. Among them, the terminal device can be a server, a mobile device, or a system in which the server and the mobile device cooperate with each other. Correspondingly, each part included in the terminal device, such as each unit, subunit, module, and submodule, can be all set in the server, all set in the mobile device, or separately set in the server and the mobile device. The terminal device is, for example, a computer device.
[0085] Further, the above server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster composed of multiple servers or as a single server. When the server is software, it can be implemented as multiple software or software modules, such as software or software modules for providing a distributed server, or as a single software or software module.
[0086] Refer to Figure 1 and Figure 2 , a method for mining protein functional modules provided by an embodiment of the present invention includes the following steps:
[0087] Step S1: Input a protein-protein interaction network (PPI network) into a node-level adaptive graph convolutional network (NASGC) model for learning and training to obtain protein node vector representations.
[0088] Specifically, in the node-level adaptive graph convolutional network model, through an adaptive mechanism, each protein node separately learns high-order and low-order neighbor information to learn the protein node vector representation. Thus, high-order and low-order structural information is fused on the gene ontology attribute features of the protein node, obtaining a more generalized protein node representation.
[0089] Step S1 of this embodiment specifically includes the following sub-steps:
[0090] Step S11: Obtain a protein-protein interaction network and construct a corresponding adjacency matrix A and a gene ontology attribute feature matrix X.
[0091] Among them, the nodes of the protein-protein interaction network are represented as V = {v1, v2,..., v n}, the gene ontology attribute feature matrix X = [x1, x2,..., x n T , n is the total number of protein nodes, the dimension of x i is d, both x and d are positive integers, and i = 1 to n.
[0092] It should be noted that the element a ij in the adjacency matrix A: when there is an interaction between node v i and node v j , a ij = 1; when there is no interaction between node v i and node v j , a ij = 0.
[0093] Step S12: Calculate the normalized Laplacian matrix Ls based on the adjacency matrix A and construct a low-pass filter G.
[0094] Specifically, the normalized Laplacian matrix Ls is Ls = I - D -1 / 2 AD -1 / 2 , and the low-pass filter G is where D is the degree matrix of the adjacency matrix A, D = diag(d1, d2,..., d n ), d i represents the number of edges of node v i , Λ is the diagonal matrix of the eigenvalues of matrix Ls, I is the diagonal matrix with all eigenvalues of matrix Ls being 1, and U is the eigenvector of matrix Ls;
[0095] Step S13: Set the number of iterative convolution layers k = t, and let t loop from 0 to M to perform the following steps S14 to S18. M represents the maximum value of the number of convolution layers and takes a positive integer value.
[0096] Step S14: Perform the k-th layer convolution operation on the gene ontology attribute feature matrix X using the low-pass graph filter G to obtain the graph representation G k X of the current convolution layer of the protein interaction network. The calculation formula is as follows:
[0097]
[0098] Step S15: Based on the graph representation G k X of the current convolution layer, calculate the state values of each protein node in the protein interaction network. The calculation formula is as follows:
[0099]
[0100] where [G k X] i represents the graph representation of node v i during the k-th layer convolution operation, represents the state value of node v i during the (k - 1)-th layer convolution operation, represents the state value of node v i during the k-th layer convolution operation, and S() is a non-linear transformation function (RNN / GRU).
[0101] Step S16: Calculate and evaluate the probability value for each node to stop iterating based on the state values. Among them, the calculation formula for the probability value is as follows:
[0102] where W h and b h are the learnable network parameters of the node-level adaptive graph convolution network model; σ represents the activation function (Sigmoid); represents node v iThe probability value during the k-th layer convolution operation.
[0103] Step S17: Set the probability value threshold ε. For each node v i , calculate the cumulative probability value of its first k layer convolution operations and compare it with the threshold ε: If the cumulative probability value reaches or exceeds the threshold ε, then for node v i stop the iterative convolution calculation and record its convolution layer number as N i = k'; If the convolution operation layer k iterates to M and the cumulative probability value is still less than the threshold ε, then the convolution layer number of node v i is N i = M; The convolution layer number N i at which the convolution of node v stops iterating i is expressed as follows:
[0104]
[0105] Step S18: Calculate the probability value during the last layer convolution operation of each node v i . The calculation formula is as follows:
[0106]
[0107] When all nodes stop iterating or reach the maximum convolution layer number M, stop the loop calculation.
[0108] Step S19: For each node v i , linearly combine the graph representation of its first N i layer convolution operations with the probability value to obtain the protein node vector representation:
[0109]
[0110] Output the vector representation of the above protein nodes and update the node information in the protein interaction network.
[0111] Through the above process, based on the node-level adaptive graph convolutional network (NASGC), each protein node learns high-order and low-order neighbor information respectively through the adaptive mechanism, and the vector representation information of the protein node is learned In which, high-order and low-order structural information is fused on the gene ontology attribute features of the protein node, so as to obtain a more generalized protein node representation.
[0112] Step S2: Based on the protein node vector representation, perform clustering through the K-means clustering algorithm to obtain the soft label of the clustering result of the protein nodes. Set the loss function according to the clustering result soft label and perform backpropagation to update the network parameters of the model.
[0113] Step S2 of this embodiment specifically includes the following sub-steps:
[0114] Step S21: Set the number m of clustering clusters, where m is a positive integer.
[0115] In this embodiment, the value of the number m of the clustering clusters is set in the following manner:
[0116] (1) Based on the gene ontology attribute feature matrix X, for the protein-protein interaction network, set m = r, and perform clustering through the K-means clustering algorithm to obtain C = {C1, C2,..., C r}; where r takes values from 2 to R, and R is a positive integer.
[0117] (2) For each specific value of m, calculate the sum of squared errors SSE of the distance from each node to the cluster center through the elbow method, and the calculation formula is as follows:
[0118] p is a node within the C i cluster, and center i is the center point of the C i cluster.
[0119] (3) According to the correspondence between the specific value of m and the calculated SSE value, fit a curve graph (such as Figure 2 shown in), determine the inflection point where the decline rate of the SSE value changes from rapid to slow in the fitted curve, and select the m value corresponding to the inflection point as the value of the final number m of clustering clusters.
[0120] Step S22: Based on the protein node vector representation, perform clustering through the K-means clustering algorithm to obtain the soft label C = {C1, C2,..., C m} of the clustering result of protein nodes.
[0121] Step S23: According to the soft label C of the clustering result, set the loss function L as:
[0122] where λ tig represents the first loss coefficient, λ sep represents the second loss coefficient, λ tig and λ sep are both constants, L tig represents the similarity between nodes within the cluster, and L sep represents the similarity between nodes between clusters.
[0123] Among them, the calculation formula of L tig is as follows:
[0124]
[0125] Among them, L sep has the following calculation formula:
[0126]
[0127] A good cluster partition should have a small intra-cluster distance, so L tig parameter representing the similarity between nodes within the cluster is introduced. A good cluster partition should also have a large inter-cluster distance, so L sep parameter representing the similarity between nodes between clusters is introduced.
[0128] For the loss coefficients λ tig and λ sep , a larger λ tig drives the nodes within the cluster closer, while λ sep drives the nodes between clusters to be well separated. λ tig and λ sep are adversarial parameters used to control the trade-off between the two metrics of compactness and separability. Among them, regarding the specific values of the loss coefficients λ tig and λ sep : The ratio of L tig and (1 / L sep ) can be observed, which can be roughly approximated by the values obtained after performing the first iteration; then the specific values of the loss coefficients λ tig and λ sep are taken to balance the two terms of λ tig L tig and in the loss function L. In a more preferred embodiment, in the loss function L: λ tig L tig : For example, it is 1:3, 1:5, 1:10, 1:15, 1:20, 1:25, 1:30, 1:35, 1:40, 1:45 or 1:50.
[0129] S24. Perform backpropagation according to the loss function to update the network parameters of the node-level adaptive graph convolutional network model.
[0130] Step S3. Based on the above steps S1 to S2, perform iterative calculations until the model converges or reaches the maximum number of iterations, and obtain the final protein node vector representation and clustering result of the last iterative calculation.
[0131] Specifically, based on a preset maximum number of iterations, steps S1 to S2 are repeated for iterative calculation until the model converges or reaches the maximum number of iterations. In the last iterative calculation, the final protein node vector representation is obtained in step S1, and the final clustering result C = {C1, C2,..., C m} is obtained in step S2, forming m clustering clusters.
[0132] Step S4: Based on the final protein node vector representation, calculate the similarity of protein nodes through the cosine similarity calculation formula, and construct a weighted adjacency matrix.
[0133] Specifically, step S3 includes: based on the adjacency matrix A and the protein node vector representation, calculate the similarity of protein nodes v i and v j through the following cosine similarity calculation formula, and construct a weighted adjacency matrix W:
[0134]
[0135] where a ij is an element of the adjacency matrix A, and w ij is an element of the weighted adjacency matrix W.
[0136] Step S5: Screen out the basic structure of protein complexes from the protein interaction network and expand it based on the calculation of the weighted adjacency matrix to obtain protein complexes.
[0137] Specifically, apply the clique mining method and combine it with the weighted adjacency matrix to screen out the basic structure of protein complexes from the protein interaction network, and embed neighbor nodes that meet predetermined conditions into the basic structure of protein complexes to obtain one of the protein functional modules: protein complexes.
[0138] Step S5 of this embodiment specifically includes the following sub-steps:
[0139] Step S51: Set and initialize the sets Alternative_core, Complex_Seed_core, and Complex_set.
[0140] Step S52: Apply the clique mining method to screen out the maximum clique structure Clique q from the protein interaction network, and place the maximum clique structure Clique q into the set Alternative_core. Where q is the number of the maximum clique structure.
[0141] Step S53: Calculate the density scores of all maximal cliques Clique in the set Alternative_core based on the weighted adjacency matrix W, and sort them in descending order according to the density scores. q Among them, the calculation formula of the density score is:
[0142] Among them, the calculation formula of the density score is:
[0143]
[0144] Step S54: Remove the maximal clique with the highest density score from the set Alternative_core and place it into the set Complex_Seed_core as the basic structure of the protein complex.
[0145] For example, in step S53, the sorting in descending order according to the density scores is Clique1, Clique2, Clique3, …; the maximal clique with the highest density score is Clique1, then the maximal clique Clique1 is removed from the set Alternative_core and placed into the set Complex_Seed_core as the basic structure of the protein complex.
[0146] Step S55: Traverse the remaining maximal clique structures in the set Alternative_core. If there are protein nodes in the remaining maximal cliques that overlap with the protein nodes in the currently maximal clique with the highest density score:
[0147] If the number of overlapping nodes is less than 2, delete the overlapping nodes in the remaining maximal clique, and keep the remaining part if the quantity is greater than 3; if the number of overlapping nodes is greater than or equal to 2, do not delete the overlapping nodes.
[0148] For example, in the first loop calculation, Clique1 is placed into the set Complex_Seed_core as the basic structure of the protein complex, then the remaining maximal clique structures in the set Alternative_core include Clique2, Clique3, Clique4, ….
[0149] Taking Clique2 as an example, if the protein nodes of Clique2 overlap with the protein nodes of the currently maximal clique Clique1 with the highest density score:
[0150] When the number of duplicate nodes is less than 2, delete the duplicate nodes in Clique2. And when the remaining nodes in Clique2 are more than 3, Clique2 is retained in the set Alternative_core; otherwise, Clique2 is removed from the set Alternative_core. When the number of duplicate nodes is greater than or equal to 2, do not delete the duplicate nodes in Clique2.
[0151] In real protein complexes, there will be multiple complexes with common internal structures, and the probability of generating protein complexes with common maximum cliques should be increased. Through the above filtering method of maximum cliques, based on the maximum clique structure as the basic framework of protein complexes, as many common protein nodes between maximum cliques as possible are retained, which can reflect the protein complexes with common maximum cliques and is more in line with the real situation.
[0152] Step S56: Repeat the above steps S53 - S55 until the set Alternative_core becomes an empty set, and several basic structures of protein complexes are obtained in the set Complex_Seed_core.
[0153] Step S57: Based on the maximum clique Clique j in the set Complex_Seed_core, for any neighbor node p j of the protein nodes in this maximum clique Clique i , calculate the correlation score between the neighbor node p i and this maximum clique Clique j based on the protein node similarity. If the correlation score is greater than the pre - set threshold θ1, then embed the protein node p i into this maximum clique Clique j . The calculation formula for the correlation score is as follows:
[0154]
[0155] Step S58: After traversing all the neighbor protein nodes of this maximum clique Clique j , complete the node expansion of the basic structure, thereby determining a protein complex, which is removed from the set Complex_Seed_core and placed into the set Complex_set.
[0156] Step S59: Repeat the above steps S57 - S58 until the set Complex_Seed_core becomes an empty set, and the finally mined protein complexes are obtained in the set Complex_set.
[0157] Step S6: Screen out the basic structure of the protein signaling pathway from each cluster in the clustering result and expand it based on the calculation of the weighted adjacency matrix to obtain the protein signaling pathway.
[0158] Specifically, screen out the shortest path within the cluster from each cluster in the clustering result C = {C1, C2,..., C m} as the basic structure of the protein signaling pathway, calculate the correlation through the weighted adjacency matrix, and embed the neighbor nodes that meet the predetermined conditions into the endpoints of the basic structure of the protein signaling pathway to obtain the second protein functional module: the protein signaling pathway.
[0159] Step S6 in this embodiment specifically includes the following sub-steps:
[0160] Step S61: Set and initialize the sets Pathway_Seed_core and Pathway_set.
[0161] Step S62: Based on the protein interaction network, traverse the m clusters, find the shortest paths between pairwise nodes within the cluster, and the length of the shortest path does not exceed 3. Place all the screened paths into the set Pathway_Seed_core as the basic structure of the protein signaling pathway.
[0162] Step S63: Based on the weighted adjacency matrix W, calculate the density scores of all the shortest paths shortest_path in the set Pathway_Seed_core q and sort them from large to small according to the density scores; where the calculation formula of the density score is:
[0163]
[0164] Step S64: Take the shortest path shortest_path with the largest density in the set Pathway_Seed_core j , for any neighbor node p j at the end of the shortest path shortest_path i , calculate the correlation score between the neighbor protein node and this shortest path based on the protein node similarity; if the correlation score is greater than the pre-set threshold θ2, then embed the neighbor protein node p i into the end of the shortest path shortest_path j ;
[0165] Among them, the calculation formula of the correlation score is as follows:
[0166]
[0167] Step S65: Traverse all neighbor protein nodes at the end of the shortest_path j to complete the node expansion of the basic structure, thereby determining a protein communication path, which is removed from the set Pathway_Seed_core and placed into the set Pathway_set.
[0168] Step S66: Repeat the above steps S64 - S65 until the set Pathway_Seed_core is an empty set, and the finally mined protein communication paths are obtained in the set Pathway_set.
[0169] The method for mining protein functional modules provided in the above embodiments is based on the node - level adaptive graph convolutional network (NASGC). Through the adaptive mechanism, each protein node separately learns high - order and low - order neighbor information. In the vector representation information of the protein nodes obtained by learning, the high - order and low - order structural information is fused on the gene ontology attribute features of the protein nodes, resulting in a more generalized protein node representation. Thus, protein complexes and protein signal pathways can be mined from the protein - protein interaction network, and the mined protein complexes are more in line with the actual situation, improving the accuracy of protein function recognition.
[0170] Based on the method for mining protein functional modules provided in the above embodiments, an embodiment of the present invention further provides a computer device, as Figure 3 shown. The computer device includes: a processor 10, a memory 20, an input device 30, and an output device 40. A GPU is set in the processor 10, and the number of processors 10 can be one or more. Figure 2 Taking one processor 10 as an example. The processor 10, memory 20, input device 30, and output device 40 in the computer device can be connected through a bus or other means.
[0171] Among them, the memory 20, as a computer - readable storage medium, can be used to store software programs, computer - executable programs, and modules. The processor 10 runs the software programs, instructions, and modules stored in the memory 20 to execute various functional applications and data processing of the device, that is, to implement the steps of the method for mining protein functional modules described in the foregoing embodiments of the present invention. The input device 30 is used to receive image data, input numerical or character information, and generate key signal inputs related to the user settings and function control of the device. The output device 40 may include a display device such as a display screen, for example, used to display images.
[0172] Based on the method for mining protein functional modules provided in the above embodiments, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for mining protein functional modules in the foregoing embodiments of the present invention are implemented. The computer storage medium may be any available medium or data storage device that can be accessed by a computer, including but not limited to magnetic memory, optical memory, and semiconductor memory, etc.
[0173] It should be noted that the above embodiments are only used to illustrate the technical concept and features of the present invention, and the purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly, and cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the spirit and essence of the present invention should be covered within the protection scope of the present invention.
Claims
1. A method for mining protein functional modules, characterized in that, Including the steps: S1. Input the protein - protein interaction network into the node - level adaptive graph convolutional network model for learning and training, enabling each protein node to learn high - order and low - order neighbor information and obtaining the protein node vector representation; S2. Based on the protein node vector representation, perform clustering through the K - means clustering algorithm to obtain the soft labels of the clustering results of protein nodes. Set the loss function according to the soft labels of the clustering results and perform backpropagation to update the network parameters of the model; S3. Perform iterative calculations based on the above steps S1 to S2 until the model converges or reaches the maximum number of iterations, obtaining the final protein node vector representation and clustering results of the last iterative calculation; S4. Based on the final protein node vector representation, calculate the similarity of protein nodes through the cosine similarity calculation formula and construct a weighted adjacency matrix; S5. Screen out the basic structure of protein complexes from the protein - protein interaction network and expand it based on the calculation of the weighted adjacency matrix to obtain protein complexes; S6. Screen out the basic structure of protein signaling pathways from each cluster of the clustering results and expand it based on the calculation of the weighted adjacency matrix to obtain protein signaling pathways.
2. The method for mining protein functional modules according to claim 1, wherein The step S1 includes: S11. Obtain a protein-protein interaction network, and construct a corresponding adjacency matrix A and a gene ontology attribute feature matrix X; wherein, the nodes of the protein-protein interaction network are represented as , and the gene ontology attribute feature matrix , where n is the total number of protein nodes, the dimension of x i is d, both x and d are positive integers, and i = 1 to n; S12. Calculate the normalized Laplacian matrix Ls based on the adjacency matrix A, and construct a low-pass filter G; the normalized Laplacian matrix Ls is , and the low-pass filter G is ; where D is the degree matrix of the adjacency matrix A , d i represents the number of edges of node v i , Λ is the diagonal matrix of the eigenvalues of matrix Ls, I is the diagonal matrix with all eigenvalues of matrix Ls being 1, and U is the eigenvector of matrix Ls; S13. Set the number of iterative convolutional layers k = t, and let t loop from 0 to M to perform the following steps S14 to S18. M represents the maximum value of the number of convolutional layers and takes a positive integer value; S14. Perform the k-th layer convolution operation using the low-pass graph filter G and the gene ontology attribute feature matrix X to obtain the graph representation G of the current convolution layer of the protein-protein interaction network, and the calculation formula is as follows: k X ; S15. Calculate the state values of each protein node in the protein interaction network based on the graph representation G k X of the current convolutional layer, and the calculation formula is as follows: , Among them, [G k X] i represents the graph representation of node v i during the k-th layer convolution operation, represents the state value of node v i during the (k - 1)-th layer convolution operation, represents the state value of node v i during the k-th layer convolution operation, and S( ) is a non-linear transformation function; S16. Calculate and evaluate the probability value for each node to stop iterative convolution based on the state value. Among them, the calculation formula of the probability value is as follows: , Among them, W h and b h are the learnable network parameters of the node-level adaptive graph convolutional network model; σ represents the activation function; represents the probability value of node v i during the k-th layer convolution operation; S17. Set a probability value threshold ε. For each node v i , calculate the cumulative probability value of its first k layers of convolution and compare it with the threshold ε: If the cumulative probability value reaches above the threshold ε, then for node v i Stop the iterative convolution calculation and record its convolution layer number as N i = k'; If the number of convolutional layers k is iterated to M and the cumulative probability value is still less than the threshold ε, then node v i has the number of convolutional layers as N i = M; The node v i The number of convolutional layers N for stopping iterative convolution i is expressed as follows: ; S18. Calculate the probability value when the last convolutional operation of each node v i is performed. The calculation formula is as follows: ; S19. For each node v i , linearly combine the graph representation of its first N i -layer convolutional operations with probability values to obtain the protein node vector representation: 。 3. The method for mining protein functional modules according to claim 1 or 2, characterized in that, The step S2 includes: S21. Set the number of clustering clusters m, where m is a positive integer; S22. Based on the protein node vector representation, perform clustering through the K - means clustering algorithm to obtain the soft labels of the clustering results of protein nodes; S23. According to the soft labels of the clustering results, set the loss function L as: ; among which, λ tig represents the first loss coefficient, λ sep represents the second loss coefficient, λ tig and λ sep are both constants, L tig represents the similarity between nodes within a cluster, L sep represents the similarity between nodes between clusters; Among them, L tig The calculation formula is as follows: ; Among them, L sep The calculation formula is as follows: ; S24. Perform backpropagation according to the loss function to update the network parameters of the node - level adaptive graph convolutional network model; The step S3 includes: based on a preset maximum number of iterations, repeating steps S1 to S2 for iterative calculation until the model converges or reaches the maximum number of iterations. In the last iterative calculation, the final protein node vector representation is obtained in step S1, and the final clustering result is obtained in step S2 , forming m clustering clusters.
4. The method for mining protein functional modules according to claim 3, wherein, The value of the number of clustering clusters m is set in the following way: Based on the protein interaction network, set m = r, and perform clustering through the K-means clustering algorithm to obtain ; where r takes values from 2 to R, and R is a positive integer; For each specific value of m, calculate the sum of squared errors SSE of the distance from each node to the cluster center through the elbow algorithm. The calculation formula is as follows: , p is a C i intra-cluster node, center i is a C i center point of the cluster; Fit a curve graph based on the corresponding relationship between the specific value of m and the calculated SSE value, determine the inflection point where the decline rate of the SSE value changes from rapid to slow in the fitted curve, and select the m value corresponding to the inflection point as the value of the final number of clustering clusters m.
5. The method for mining protein functional modules according to claim 3, wherein In the loss function L: 。 6. The method for mining protein functional modules according to claim 3, wherein The step S4 includes: calculating the similarity between protein nodes v i and v j based on the adjacency matrix A and the vector representation of protein nodes, and constructing a weighted adjacency matrix W: , where a ij is an element of the adjacency matrix A, and w ij is an element of the weighted adjacency matrix W.
7. The method for mining protein functional modules according to claim 6, characterized in that, The step S5 includes: S51. Set and initialize the sets Alternative_core, Complex_Seed_core, and Complex_set; S52. Use the clique mining method to screen out the maximum clique structure Clique from the protein interaction network q , and put the maximum clique structure Clique q into the set Alternative_core; S53. Calculate the density scores of all maximal cliques Clique in the set Alternative_core based on the weighted adjacency matrix W, and sort them in descending order according to the density scores; wherein, the calculation formula of the density score is: q and perform a descending order sorting according to the density scores; wherein, the calculation formula of the density score is: ; S54. Remove the maximum - density clique from the set Alternative_core and place it into the set Complex_Seed_core as the basic structure of the protein complex; S55. Traverse the remaining maximal clique structures in the set Alternative_core. When there are overlapping protein nodes between the remaining maximal cliques and the protein nodes in the maximal clique with the currently largest density score: If the number of overlapping nodes is less than 2, delete the overlapping nodes in the remaining maximal cliques, and retain the remaining part if its quantity is greater than 3; if the number of overlapping nodes is greater than or equal to 2, do not delete the overlapping nodes; S56. Repeat the above steps S53 - S55 until the set Alternative_core becomes an empty set, and obtain several basic structures of protein complexes in the set Complex_Seed_core; S57. Clique based on the maximum clique in the set Complex_Seed_core j , for the maximal clique j Any neighbor node p of the protein node i , calculate neighbor node p based on protein node similarity i With the maximal clique j If the correlation score is greater than the preset threshold θ1, the neighbor protein node p i Embed the maximal clique j ; The calculation formula of the correlation score is as follows: ; S58. After traversing all the neighbor protein nodes of the maximum clique Clique j a protein complex is determined, removed from the set Complex_Seed_core and placed into the set Complex_set; S59. Repeat the above steps S57 - S58 until the set Complex_Seed_core becomes an empty set, and obtain the finally mined protein complexes in the set Complex_set.
8. The method for mining protein functional modules according to claim 6, wherein The step S6 includes: S61. Set and initialize the sets Pathway_Seed_core and Pathway_set; S62. Based on the protein interaction network, traverse the m clustering clusters, find the shortest paths between pairwise nodes within the clusters, and the length of the shortest paths does not exceed 3. Place all the filtered paths into the set Pathway_Seed_core as the basic structures of protein signal pathways; S63. Calculate all shortest paths shortest_path in the set Pathway_Seed_core based on the weighted adjacency matrix W q and sort them in descending order according to the density scores; wherein, the calculation formula of the density score is: ; S64. Select the shortest path shortest_path with the highest density in the set Pathway_Seed_core j For the shortest path shortest_path j For any neighbor node p at the end i Calculate the correlation score between the neighbor protein node and this shortest path based on protein node similarity; if the correlation score is greater than the pre-set threshold θ2, then the neighbor protein node p i Embed it at the end of the shortest path shortest_path j at the end; Among them, the calculation formula of the correlation score is as follows: ; S65. After traversing all the neighbor protein nodes at the end of the shortest_path j a protein communication path is determined, removed from the set Pathway_Seed_core and placed into the set Pathway_set; S66. Repeat the above steps S64 - S65 until the set Pathway_Seed_core becomes an empty set, and obtain the finally mined protein communication paths in the set Pathway_set.
9. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method for mining protein functional modules according to any one of claims 1 - 8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the steps of the method for mining protein functional modules according to any one of claims 1 - 8.
Citation Information
Patent Citations
Protein complex recognizing method based on multi-source data fusion and multi-target optimization
CN108009403A
Protein design method and system
WO2017081687A1