Functional module identification method, device, terminal equipment and storage medium
By measuring the r-order proximity indicators between nodes in the biological network, iteratively enhance the biological network, and performing low-rank representation learning, the problem of inaccurate recognition of functional modules is solved, and higher recognition accuracy is achieved.
Patent Information
- Application Number
- CN202410223522.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-28
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-02-28
AI Technical Summary
In the prior art, the identification of functional modules is not accurate enough, and it is difficult to effectively utilize the topological information and additional information of the biological network, resulting in the improvement of the identification accuracy of functional modules in the biological network.
By measuring the r-order proximity index between nodes in the biological network, the biological network is iteratively enhanced, low-rank representation learning is performed, and the low-rank matrix is obtained, and the functional modules of each node in the biological network are divided according to the matrix.
This method can fully explore the information implicit in the biological network, explicitly express it in the enhanced network, improve the accuracy of functional module recognition, and better identify functional modules in the biological network.
Smart Images

Figure CN118072833B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the technical field of information extraction and data mining, and in particular, relates to a method, apparatus, terminal device and storage medium for identifying a functional module. Background Art
[0002] With the development of bioinformatics and systems biology research, biological networks have become a powerful tool for understanding and analyzing complex biological systems. They can not only describe the emerging network structure characteristics of biological systems, but also compare the similarities or differences in network structures of different biological systems. Functional modules in biological networks, as core components of biological networks, are crucial to understanding the functions and behaviors of biological systems. However, due to the complexity and large scale of biological networks, it is still a challenge to accurately identify the functional modules of biological networks based on the topological information of biological networks.
[0003] Most of the existing studies have only focused on the basic topological information of biological networks, resulting in inaccurate identification of functional modules in biological networks. The few methods that want to use additional information beyond the basic topological information of biological networks mainly combine artificially collected prior information or background knowledge in biological networks to supplement the topological information of biological networks. However, the large-scale and complexity of biological networks make it difficult to obtain reliable prior information to supplement the basic topological information of biological networks, and the accuracy of identifying functional modules in biological networks still needs to be improved. Summary of the invention
[0004] In view of this, embodiments of the present application provide a method, apparatus, terminal device and storage medium for identifying a functional module to solve the problem of inaccurate functional module identification in the prior art.
[0005] A first aspect of an embodiment of the present application provides a method for identifying a functional module, including:
[0006] The r-order proximity index between nodes in a biological network is measured, where r is a positive integer;
[0007] Iteratively enhancing the biological network based on the r-order proximity index to obtain an enhanced network;
[0008] Performing low-rank representation learning on the enhanced network to obtain a low-rank matrix;
[0009] The functional modules of each node in the biological network are divided according to the low-rank matrix.
[0010] In a possible implementation of the first aspect, the measuring the r-order proximity index between nodes in the biological network includes:
[0011] Identify pairs of neighboring nodes of order 1 to r in biological networks;
[0012] The inter-point mutual information index is calculated based on the 1st to rth order neighboring node pairs, and the rth order proximity index between the nodes in the biological network is measured by the inter-point mutual information index.
[0013] In a possible implementation manner of the first aspect, the calculating the inter-point mutual information index based on the 1st to rth order neighboring node pairs, and measuring the rth order proximity index between nodes by the inter-point mutual information index, includes:
[0014] Calculate the mutual information index between points based on the 1st to rth order neighboring node pairs and the order of the neighboring node pairs;
[0015] The mutual information index between the points is used as the node v i and node v j The r-order proximity index between them;
[0016] The mutual information index between points is calculated by the following formula:
[0017]
[0018] Among them, v i and v j Represents two nodes; N k (v i ) indicates that it contains node v i The number of k-order neighboring node pairs, N k (v j ) indicates that it contains node v j The number of k-order neighboring node pairs, N k (v i ,v j ) represents the k-order neighboring node pair {v i ,v j}The number of times it appears; H k represents the set of k-order neighboring node pairs, |H k | indicates H k The number of node pairs in the middle; ω k is the weight coefficient, which is negatively correlated with the order k of the neighboring node pair; P(v i ,v j ) represents node v i and node v j The mutual information index between the points; r is the order of the proximity index between the nodes to be measured.
[0019] In a possible implementation manner of the first aspect, the iteratively enhancing the biological network based on the r-order proximity index to obtain an enhanced network includes:
[0020] encoding the r-order proximity index into a biological network to obtain an enhanced network;
[0021] Measuring the r-order proximity index between nodes in the enhanced network;
[0022] Encoding the r-order proximity index into the enhanced network to obtain a new enhanced network, and using the new enhanced network as the enhanced network;
[0023] Return to the step of measuring the high-order proximity index between nodes in the enhanced network until a preset iteration condition is met.
[0024] In a possible implementation manner of the first aspect, encoding the r-order proximity indicator into the enhanced network to obtain a new enhanced network, and using the new enhanced network as the enhanced network includes:
[0025] According to the following network enhancement strategy formula, the r-order proximity index is encoded into the enhanced network to obtain a new enhanced network;
[0026] Using the new enhanced network as the enhanced network;
[0027] The network enhancement strategy formula is:
[0028]
[0029] Among them, v i and v j represents two nodes; ε is the threshold coefficient, which is a real number; P(v i ,v j ) represents node v i and node v j The mutual information index between points; a ij Represents node v i and v j The connection relationship between ij =1, indicating node v i and v j There is an edge between them; a ij =0, indicating that node v i and v j There is no edge between them; Represents node v i and v j The enhanced connection between them.
[0030] In a possible implementation manner of the first aspect, performing low-rank representation learning on the enhanced network to obtain a low-rank matrix includes:
[0031] The enhanced network is decomposed based on a capacity-expanded graph regularized symmetric non-negative matrix factorization algorithm to obtain a low-rank matrix.
[0032] In a possible implementation manner of the first aspect, the capacity-expanded graph regularized symmetric non-negative matrix decomposition algorithm decomposes the enhanced network to obtain a low-rank matrix, including:
[0033] The loss function is determined as:
[0034]
[0035] Among them, J CGF is the loss function; is the n×n matrix of the enhanced network; X, Y and U are all n×K non-negative feature matrices; and is the regularization term of the equation; the real number θ>0 is the generalized loss in the balanced loss function and The coefficient of importance; λ>0 is the graph regularization coefficient; Tr represents the trace of the computation matrix; is the Laplacian matrix, is an n×n diagonal matrix whose element values are calculated as
[0036] The optimization problem to be solved is determined as:
[0037]
[0038] The feature matrices X, Y, and U are iteratively updated according to the following learning rules until the convergence condition is reached, and the feature matrices X, Y, and U finally obtained are used as the solution to the optimization problem:
[0039]
[0040] The output feature matrix X is a low-rank matrix.
[0041] A second aspect of an embodiment of the present application provides a functional module identification device, including:
[0042] The measurement module is used to measure the r-order proximity index between nodes in the biological network, where r is a positive integer;
[0043] An iterative enhancement module, which iteratively enhances the biological network based on the r-order proximity index to obtain an enhanced network;
[0044] A low-rank representation learning module performs low-rank representation learning on the enhanced network to obtain a low-rank matrix;
[0045] A partitioning module is used to partition the functional modules of each node in the biological network according to the low-rank matrix.
[0046] A third aspect of an embodiment of the present application provides a terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method as claimed in any one of claims 1 to 7 when executing the computer program.
[0047] A fourth aspect of an embodiment of the present application provides a storage medium, which is a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method as described in any one of claims 1 to 7 are implemented.
[0048] Compared with the prior art, the beneficial effects of the embodiments of the present application are as follows: the present application measures the r-order proximity index between nodes in the biological network, and then iteratively enhances the biological network based on the r-order proximity index to obtain an enhanced network, and then performs low-rank representation learning on the enhanced network to obtain a low-rank matrix. Finally, the functional modules of each node in the biological network are divided by the low-rank matrix, so that the implicit information in the biological network can be fully mined, and then the implicit information is explicitly expressed in the enhanced network, so that a low-rank matrix that also contains comprehensive information can be obtained, and the nodes in the biological network can be divided according to the low-rank matrix, thereby improving the accuracy of functional module identification.
[0049] It can be understood that the beneficial effects of the second to fourth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0051] Figure 1 It is a schematic diagram of a biological network provided in an embodiment of the present application;
[0052] Figure 2 This is a schematic diagram of the first implementation flow of the functional module identification method provided in the embodiment of the present application;
[0053] Figure 3 This is a schematic diagram of a second implementation flow of the functional module identification method provided in an embodiment of the present application;
[0054] Figure 4This is a schematic diagram of a third implementation flow of the functional module identification method provided in an embodiment of the present application;
[0055] Figure 5 is a schematic diagram of a functional module identification device provided in an embodiment of the present application;
[0056] Figure 6 is a schematic diagram of a terminal device provided in an embodiment of the present application;
[0057] Figure 7 It is a bar chart comparing the performance of the functional module identification method provided in the embodiment of the present application with the prior art in biological network function identification. DETAILED DESCRIPTION
[0058] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0059] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.
[0060] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0061] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.
[0062] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0063] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0064] The embodiment of the present application provides a method for identifying a functional module, which can be executed by a terminal device when running a computer program with corresponding functions, and is used to realize the identification of functional modules. This method measures the r-order proximity index between nodes in a biological network, and then iteratively enhances the biological network based on the r-order proximity index to obtain an enhanced network, and then performs low-rank representation learning on the enhanced network to obtain a low-rank matrix, and finally divides the functional modules of each node in the biological network through the low-rank matrix, which can comprehensively mine the implicit information in the biological network, and then explicitly express the implicit information in the enhanced network, which can obtain a low-rank matrix that also contains comprehensive information, and divide the nodes in the biological network according to the low-rank matrix, thereby improving the accuracy of functional module identification.
[0065] In application, the functional module identification method provided in the present application can be applied to the identification of biological network functional modules, for example, protein complex prediction, identifying protein clusters or protein complexes that can perform certain biological functions from protein interaction networks; gene co-expression module identification, identifying gene clusters or gene co-expression modules with high topological overlap similarity from gene co-expression networks; micro-ribonucleic acid (miRNA)-disease association relationship prediction, accurately predicting the unknown association between miRNA and disease from the miRNA-disease heterogeneous information network.
[0066] In application, the terminal device can be a computer device such as a mobile phone, a tablet computer, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc. The embodiment of the present application does not impose any restrictions on the specific type of the terminal device.
[0067] like Figure 2 As shown, the present application provides a method for identifying a functional module, the method comprising the following steps S11 to S14:
[0068] Step S11, measuring the r-order proximity index between nodes in the biological network, where r is a positive integer.
[0069] In applications, biological networks are generally biological networks, which are abstract representations of biological systems in the form of graphs. In biological networks, the elements that make up the biological system are nodes, and the connections between elements are edges. Biological networks can be protein interaction networks, gene co-expression networks, brain neural networks, etc.
[0070] In the application, before measuring the r-order proximity index between nodes in the biological network, it also includes describing the biological network with a graph G = (V, E), where V = {v i |i∈{1,…,n} represents a set of n nodes, v i is a node in the biological network; E = {e ij |i,j∈{1,…,n}} represents a set of m edges, e ij is node v i and v j The connecting edge between.
[0071] In the application, in order to enhance the intuitiveness, simplicity and computability of the relationship between nodes in the biological network, the biological network G uses the adjacency matrix A = [a ij ] to represent and store, where a ij Represents node v i and v j For undirected and unweighted biological networks, the biological network G is usually a unipartite graph, and A is a symmetric binary matrix: if the node v i and v j If there is an edge between them, then the corresponding element a in A ij =1, if node v i and v j There is no edge between them, then the corresponding element a in A ij =0.
[0072] It is understandable that in order to facilitate the understanding of the method provided by the present application, the relevant definition of proximity is given. First Order Proximity (FOP) describes the relationship between two nodes with a direct edge. High Order Proximity (HOP) describes the relationship between two nodes connected by a multi-hop edge. If two nodes are connected by a two-hop edge, they have a second-order proximity relationship, which is described by second-order proximity; similarly, r-order proximity describes the proximity relationship between two nodes connected by r-hop edges. Figure 1In the biological network shown in Figure 2, the node pair {v a ,v b} are directly connected by edge e ab connection, with FOP, node pair {v a ,v b} is a first-order neighboring node pair; the node pair {v a ,v c} is connected by two hops {e ab ,e bc} connection, with a second-order HOP, the node pair {v a ,v c} is a second-order neighboring node pair; the node pair {v b ,v k} by four-hop edge {e bf ,e fg ,e gi ,e ik} connection, with a fourth-order HOP, node pair {v b ,v k} is a fourth-order neighboring node pair.
[0073] In the application, the r-order proximity index between nodes in the biological network includes the 1st to rth order proximity between nodes in the biological network. It can be understood that the r-order proximity index between nodes in the biological network fully includes the high-order information implicit between nodes.
[0074] like Figure 3 As shown, in one embodiment, step S11 includes the following steps S111 to S112: Step S111, identifying 1st to rth order neighboring node pairs in the biological network.
[0075] In the application, the biological network can be subjected to a depth-first search (DFS) at depths of 1 to r to obtain the 1st to rth order neighboring node pairs in the biological network. For example, the nodes in the biological network are set as the starting nodes one by one, and then the search depth of the depth-first search is set to 1, and the biological network is traversed to obtain all the first-order neighboring node pairs; the search depth of the depth-first search is set to 2, and the biological network is traversed to obtain all the second-order neighboring nodes; ...; the search depth of the depth-first search is set to r, and the biological network is traversed to obtain all the rth order neighboring node pairs.
[0076] In an application, the 1st to rth order node pairs in the identified biological network can be stored separately in the form of node pair lists.
[0077] Step S112, calculating the point-to-point mutual information index based on the 1st to rth order neighboring node pairs, and measuring the rth order proximity index between the nodes in the biological network by the point-to-point mutual information index.
[0078] In applications, the point-wise mutual information index is an extension of the point-wise mutual information. Point-wise mutual information (PMI) is used to compare the probability of observing a node pair with the probability of independently observing the two nodes in the node pair. The calculation method is:
[0079]
[0080] Where p(w1) and p(w2) are the probabilities of observing nodes w1 and w2 independently, respectively, and p(w1,w2) calculates the probability of observing w1 and w2 simultaneously, that is, the probability of observing the node pair {w1,w2}. It can be understood that in biological networks, if the joint probability p(w1,w2) is greater than the product of the independent observation probabilities p(w1)p(w2), it means that even if there is no direct edge between the two nodes w1 and w2, there is non-negligible high-order information implicit between nodes w1 and w2. Therefore, the point-to-point mutual information index is calculated based on the 1st to rth order neighboring node pairs and the point-to-point mutual information calculation method, and the rth order proximity index between nodes in the biological network is measured by the point-to-point mutual information index, which can pay more attention to the implicit information in the biological network, and is conducive to improving the accuracy and reliability of the final division of the functional modules of the nodes.
[0081] In one embodiment, step S112 includes:
[0082] Calculate the mutual information index between points based on the 1st to rth order neighboring node pairs and the order of the neighboring node pairs;
[0083] The mutual information index between the points is used as the node v i and node v j The r-order proximity index between them;
[0084] The mutual information index between points is calculated by the following formula:
[0085]
[0086] Among them, v i and v j Represents two nodes; N k (v i ) indicates that it contains node v i The number of k-order neighboring node pairs, N k (v j ) indicates that it contains node v j The number of k-order neighboring node pairs, N k (v i ,v j ) represents the k-order neighboring node pair {v i,v j}The number of times it appears; H k represents the set of k-order neighboring node pairs, |H k | indicates H k The number of node pairs in the middle; ω k is the weight coefficient, which is negatively correlated with the order k of the neighboring node pair; P(v i ,v j ) represents node v i and node v j The mutual information index between the points; r is the order of the proximity index between the nodes to be measured.
[0087] In the application, when r is 2, the calculation of the mutual information index between points is:
[0088]
[0089] Among them, P(v i ,v j ) is the node v i , and v j The mutual information index between points, v i and v j represents two nodes, H represents the first-order and second-order neighboring node pairs in the biological network obtained by DFS, |H| represents the total number of first-order and second-order neighboring node pairs in the biological network obtained by DFS, N(v i ) represents node v i The number of times it appears in H, N(v j ) represents node v j The number of times it appears in H, N(v i ,v j ) represents the node pair {v i ,v j The number of times} appears in H.
[0090] In the application, when r is greater than 2, the node v is independently observed i The probability is:
[0091]
[0092] Independently observe node v i The probability is:
[0093]
[0094] Observe that the node pair {v i ,v j The probability of} is:
[0095]
[0096] The mutual information index between points is:
[0097]
[0098] Among them, v i and v j Represents two nodes; N k (v i ) indicates that it contains node v i The number of k-order neighboring node pairs, N k (v j ) indicates that it contains node v j The number of k-order neighboring node pairs, N k (v i ,v j ) represents the k-order neighboring node pair {v i ,v j}The number of times it appears; H k represents the set of k-order neighboring node pairs, |H k | indicates H k The number of node pairs in the middle; ω k is the weight coefficient, which is negatively correlated with the order k of the neighboring node pair; P(v i ,v j ) represents node v i and node v j The mutual information index between the points; r is the order of the proximity index between the nodes to be measured.
[0099] It can be understood that as the order k of the neighboring node pair considered increases, the distance between the nodes in the network topology also increases, and the connection between the two nodes also weakens, that is, the weight coefficient ω k Negatively correlated with the order k of the neighboring node pair, for example, the reciprocal of the order k of the neighboring node pair can be used as the weight coefficient ω k , at this time, the mutual information index between points is:
[0100]
[0101] In applications, r can be set to the diameter of the biological network to ensure that the obtained r-order proximity index covers all possible node pairs, but this will cause large computational and storage overheads; in order to balance the coverage of r-order proximity and computational efficiency, according to the small world hypothesis of complex networks (also known as six degrees of separation), the average path length for any two nodes in a complex network to establish a high-order relationship is 6, so the maximum value of r can be 6, that is, r can be 2, 3, 4, 5 or 6.
[0102] Step S12, iteratively enhancing the biological network based on the r-order proximity index to obtain an enhanced network.
[0103] In the application, the r-order proximity index between nodes in the biological network comprehensively includes the high-order information implicit in the biological network, so the r-order proximity index between nodes in the biological network can be used as high-value supplementary information of the biological network. Step S12 explicitly encodes the r-order proximity index into the biological network through an iterative network topology enhancement (INE) strategy to enhance the basic network topology information of the biological network, thereby obtaining an enhanced network, so that the high-order information of the biological network can be explicitly expressed in the basic network topology information of the enhanced network.
[0104] like Figure 4 As shown, in one embodiment, step S12 includes the following steps S121 to S124:
[0105] Step S121, encoding the r-order proximity index into the biological network to obtain an enhanced network;
[0106] Step S122, measuring the r-order proximity index between nodes in the enhanced network;
[0107] Step S123, encoding the r-order proximity index into the enhanced network to obtain a new enhanced network, and using the new enhanced network as the enhanced network;
[0108] Step S124, returning to the step of measuring the high-order proximity index between nodes in the enhanced network until a preset iteration condition is met.
[0109] In the application, the preset iteration condition is that the number of executions of step S123 reaches the preset number of iterations d, and step S121 and step S123 also include outputting the adjacency matrix of the enhanced network is an n×n matrix, It can be understood that as steps S121 to S124 are executed, the results of each iteration will be accumulated, and the basic network topology information in the enhanced network will explicitly contain the high-order information implicit in the biological network, which is conducive to improving the accuracy and reliability of the subsequent division of functional modules of nodes in the biological network.
[0110] In one embodiment, the step S123 includes:
[0111] r according to the following network enhancement strategy formula, encoding the r-order proximity index into the enhanced network to obtain a new enhanced network;
[0112] Using the new enhanced network as the enhanced network;
[0113] The network enhancement strategy formula is:
[0114]
[0115] Among them, v i and v j represents two nodes; ε is the threshold coefficient, which is a real number; P(v i ,v j ) represents node v i and node v j The mutual information index between points; a ij Represents node v i and v j The connection relationship between ij =1, indicating node v i and v j There is an edge between them; a ij =0, indicating that node v i and v j There is no edge between them; Represents node v i and v j The enhanced connection between them.
[0116] In the application, the implementation of step S121 is consistent with the implementation of step S123, which will not be described in detail here; the implementation of step S122 is consistent with the implementation of step S11, which will not be described in detail here. In the application, the r-order proximity index can be explicitly encoded in the enhanced network by adding an edge between nodes. r determines whether to add a node v to the node v according to the r-order proximity index between nodes and the threshold coefficient. i and v j An edge is added between them, and the weight of the reconstructed edge is set to 1 or 0, instead of making them directly equal to the corresponding r-order proximity index. On the one hand, it avoids that each node can be connected to any other node in the network with a very low probability, making the entire network topology too smooth, making it difficult to capture more differentiated node representations; on the other hand, the adjacency matrix of the enhanced network is Some elements in Forcing it to zero retains a certain degree of network topology sparseness and can better explain the network representation results.
[0117] Step S13, performing low-rank representation learning on the enhanced network to obtain a low-rank matrix.
[0118] In one embodiment, step S13 includes:
[0119] The enhanced network is decomposed based on a capacity-expanded graph regularized symmetric non-negative matrix factorization algorithm to obtain a low-rank matrix.
[0120] In the application, the capacity-enlarged and graph-regularized factorization algorithm (CGF) based on graph regularization invented by introducing the feature capacity of the expanded model based on the symmetric non-negative matrix factorization algorithm (SNMF) not only improves the representation learning ability of low-rank matrices, but also maintains the local topological invariance of the network.
[0121] In one embodiment, the capacity-expanded graph regularized symmetric non-negative matrix factorization algorithm decomposes the enhanced network to obtain a low-rank matrix, including:
[0122] The loss function is determined as:
[0123]
[0124] Among them, J CGF is the loss function; is the n×n matrix of the enhanced network; X, Y and U are all n×K non-negative feature matrices; and is the regularization term of the equation; the real number θ>0 is the generalized loss in the balanced loss function and The coefficient of importance; λ>0 is the graph regularization coefficient; Tr represents the trace of the computation matrix; is the Laplacian matrix, is an n×n diagonal matrix whose element values are calculated as
[0125] The optimization problem to be solved is determined as:
[0126]
[0127] The feature matrices X, Y, and U are iteratively updated according to the following learning rules until the convergence condition is reached, and the feature matrices X, Y, and U finally obtained are used as the solution to the optimization problem:
[0128]
[0129] The output feature matrix X is a low-rank matrix.
[0130] In the application, the non-negative variables Y and U are used to expand the feature capacity and fit the adjacency matrix of the enhanced network. Equation regularization term and It is used to bridge and share information between feature matrices, and pass the information learned by Y and U to X. It can be understood that since Y and U expand the understanding space, X can be enhanced. The learning rule is derived by combining the Lagrangian function method and the KKT (Karush-Kuhn-Tucker) condition of the inequality constraint with the optimization problem.
[0131] In the application, the convergence condition is that the absolute value of the difference between the loss function values after two consecutive iterative training is less than 0.1 or the number of iterations reaches a preset learning number.
[0132] Step S14, dividing the functional modules of each node in the biological network according to the low-rank matrix.
[0133] In the application, the functional modules of each node in the biological network are divided according to the low-rank matrix X and the following assignment rules:
[0134]
[0135] Where V represents the biological network, v j Represents the nodes; C K is the function module k, k∈{1,2,…,K}, K is the total number of function modules; x jl is an element in the low-rank matrix X, representing node v j The probability of being divided into functional module l, x jk is the maximum value among them, indicating that node v j The probability of being divided into functional module k is the greatest. After dividing the functional modules of all nodes in the biological network, the identified functional module set can be obtained
[0136] C={C1,C2,…,C K}.
[0137] like Figure 7 As shown, taking protein complex prediction as an example, the protein complex is a functional module hidden in the protein network. Identifying such functional modules is to identify protein clusters or protein complexes that can perform certain biological functions from the protein interaction network. The functional module identification method provided in this application is significantly improved in standard mutual information, purity, accuracy, etc. compared with other methods in the prior art.
[0138] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0139] The present application also provides a functional module identification device for executing the steps in the functional module identification method embodiment. The functional module identification device can be a virtual device in a terminal device, which is run by a processor of the terminal device, or it can be the terminal device itself.
[0140] like Figure 5 As shown, the functional module identification device 5 provided in the embodiment of the present application includes modules 501 to 504:
[0141] A measurement module 501 is used to measure the r-order proximity index between nodes in the biological network, where r is a positive integer;
[0142] Iterative enhancement module 502, iteratively enhancing the biological network based on the r-order proximity index to obtain an enhanced network;
[0143] A low-rank representation learning module 503 performs low-rank representation learning on the enhanced network to obtain a low-rank matrix;
[0144] The partitioning module 504 is used to partition the functional modules of each node in the biological network according to the low-rank matrix.
[0145] In application, each module in the functional module identification device may be a software program module, or may be implemented by different logic circuits integrated in a processor, or may be implemented by multiple distributed processors.
[0146] Figure 6 This is a schematic diagram of the structure of a terminal device provided in an embodiment of the present application. Figure 6 As shown, the terminal device 6 of this embodiment includes: at least one processor 60 ( Figure 6 Only one is shown in the figure) a processor, a memory 61, and a computer program 62 stored in the memory 61 and executable on the at least one processor 60, and when the processor 60 executes the computer program 62, the steps in any of the above-mentioned functional module identification method embodiments are implemented.
[0147] The terminal device 6 may be a computer device such as a mobile phone, a tablet computer, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc. The terminal device may include, but is not limited to, a processor 60 and a memory 61. Those skilled in the art will understand that Figure 6 It is only an example of the terminal device 6 and does not constitute a limitation on the terminal device 6. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, it may also include input and output devices, network access devices, etc.
[0148] The processor 60 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0149] In some embodiments, the memory 61 may be an internal storage unit of the terminal device 6, such as a hard disk or memory of the terminal device 6. In other embodiments, the memory 61 may also be an external storage device of the terminal device 6, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the terminal device 6. Further, the memory 61 may also include both an internal storage unit of the terminal device 6 and an external storage device. The memory 61 is used to store an operating system, an application program, a boot loader (BootLoader), data, and other programs, such as the program code of the computer program, etc. The memory 61 may also be used to temporarily store data that has been output or is to be output.
[0150] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / modules are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.
[0151] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0152] An embodiment of the present application further provides a storage medium, which is a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.
[0153] An embodiment of the present application provides a computer program product. When the computer program product runs on a mobile terminal, the mobile terminal can implement the steps in the above-mentioned method embodiments when executing the computer program product.
[0154] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may at least include: any entity or device that can carry the computer program code to the device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electric carrier signal, a telecommunication signal, and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.
[0155] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0156] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0157] In the embodiments provided in the present application, it should be understood that the disclosed devices / network equipment and methods can be implemented in other ways. For example, the device / network equipment embodiments described above are merely schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0158] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0159] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for identifying a functional module, characterized in that: Applied to functional module identification in biological networks, the method comprises: The r-order proximity index between nodes in a biological network is measured, where r is a positive integer; Iteratively enhancing the biological network based on the r-order proximity index to obtain an enhanced network; Performing low-rank representation learning on the enhanced network to obtain a low-rank matrix; Divide the functional modules of each node in the biological network according to the low-rank matrix; The r-order proximity index between nodes in the biological network is measured, including: Identify pairs of neighboring nodes of order 1 to r in biological networks; Calculate the mutual information index between points based on the 1st to rth order neighboring node pairs, and measure the rth order proximity index between nodes in the biological network by the mutual information index between points; The step of calculating the mutual information index between nodes based on the 1st to rth order neighboring node pairs, and measuring the rth order proximity index between nodes by the mutual information index between nodes, comprises: Calculate the mutual information index between points based on the 1st to rth order neighboring node pairs and the order of the neighboring node pairs; The mutual information index between the points is used as the node v i and node v j The r-order proximity index between them; The mutual information index between points is calculated by the following formula: Among them, v i and v j Represents two nodes; N k (v i ) indicates that it contains node v i The number of k-order neighboring node pairs, N k (v j ) indicates that it contains node v j The number of k-order neighboring node pairs, N k (v i ,v j ) represents the k-order neighboring node pair {v i ,v j }The number of times it appears; H k represents the set of k-order neighboring node pairs, |H k | indicates H k The number of node pairs in the middle; ω k is the weight coefficient, which is negatively correlated with the order k of the neighboring node pair; P(v i ,v j ) represents node v i and node v j The mutual information index between the points; r is the order of the proximity index between the nodes to be measured.
2. The functional module identification method according to claim 1, characterized in that: Based on the r The biological network is iteratively enhanced by using the order proximity index to obtain an enhanced network, including: encoding the r-order proximity index into a biological network to obtain an enhanced network; Measuring the r-order proximity index between nodes in the enhanced network; Encoding the r-order proximity index into the enhanced network to obtain a new enhanced network, and using the new enhanced network as the enhanced network; Return to the step of measuring the high-order proximity index between nodes in the enhanced network until a preset iteration condition is met.
3. The functional module identification method according to claim 2, characterized in that: The step of encoding the r-order proximity index into the enhanced network to obtain a new enhanced network, and using the new enhanced network as the enhanced network, comprises: According to the network enhancement strategy formula, the r-order proximity index is encoded into the enhanced network to obtain a new enhanced network; Using the new enhanced network as the enhanced network; The network enhancement strategy formula is: Other situations; Among them, v i and v j represents two nodes; ε is the threshold coefficient, which is a real number; P(v i ,v j ) represents node v i and node v j The mutual information index between points; a ij Represents node v i and v j The connection relationship between ij =1, indicating node v i and v j There is an edge between them; a ij =0, indicating that node v i and v j There is no edge between them; Represents node v i and v j The enhanced connection relationship between them.
4. The functional module identification method according to claim 1, characterized in that: The step of performing low-rank representation learning on the enhanced network to obtain a low-rank matrix includes: The enhanced network is decomposed based on a capacity-expanded graph regularized symmetric non-negative matrix factorization algorithm to obtain a low-rank matrix.
5. The functional module identification method according to claim 4, characterized in that: The symmetric non-negative matrix decomposition algorithm based on graph regularization with capacity expansion decomposes the enhanced network to obtain a low-rank matrix, including: The loss function is determined as: stX≥0,Y≥0,U≥0; Among them, J CGF is the loss function; is the n×n matrix of the enhanced network; X, Y and U are all n×K non-negative feature matrices; and is the regularization term of the equation; the real number θ>0 is the generalized loss in the balanced loss function and The coefficient of importance; λ>0 is the graph regularization coefficient; Tr represents the trace of the computation matrix; is the Laplacian matrix, is an n×n diagonal matrix whose element values are calculated as The optimization problem to be solved is determined as: The feature matrices X, Y, and U are iteratively updated according to the following learning rules until the convergence condition is reached, and the feature matrices X, Y, and U finally obtained are used as the solution to the optimization problem: The output feature matrix X is a low-rank matrix.
6. A functional module identification device, characterized in that: Applied to functional module identification in biological networks, the device comprises: The measurement module is used to measure the r-order proximity index between nodes in the biological network, where r is a positive integer; An iterative enhancement module, which iteratively enhances the biological network based on the r-order proximity index to obtain an enhanced network; A low-rank representation learning module performs low-rank representation learning on the enhanced network to obtain a low-rank matrix; A partitioning module, used for partitioning the functional modules of each node in the biological network according to the low-rank matrix; The r-order proximity index between nodes in the biological network is measured, including: Identify pairs of neighboring nodes of order 1 to r in biological networks; Calculate the mutual information index between points based on the 1st to rth order neighboring node pairs, and measure the rth order proximity index between nodes in the biological network by the mutual information index between points; The step of calculating the mutual information index between nodes based on the 1st to rth order neighboring node pairs, and measuring the rth order proximity index between nodes by the mutual information index between nodes, comprises: Calculate the mutual information index between points based on the 1st to rth order neighboring node pairs and the order of the neighboring node pairs; The mutual information index between the points is used as the node v i and node v j The r-order proximity index between them; The mutual information index between points is calculated by the following formula: Among them, v i and v j Represents two nodes; N k (v i ) indicates that it contains node v i The number of k-order neighboring node pairs, N k (v j ) indicates that it contains node v j The number of k-order neighboring node pairs, N k (v i ,v j ) represents the k-order neighboring node pair {v i ,v j }The number of times it appears; H k represents the set of k-order neighboring node pairs, |H k | indicates H k The number of node pairs in the middle; ω k is the weight coefficient, which is negatively correlated with the order k of the neighboring node pair; P(v i ,v j ) represents node v i and node v j The mutual information index between the points; r is the order of the proximity index between the nodes to be measured.
7. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.
8. A storage medium, wherein the storage medium is a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Fusion network drug target relationship prediction method based on network enhancement and graph regularization
CN112270950A
Protein function module mining method, computer equipment and storage medium
CN116417060A