A multi-network joint clustering method for protein complex recognition

By employing a multi-network joint clustering method, protein complexes are identified, solving the problem of identifying dynamically changing and shared unique complexes in existing technologies, and achieving higher accuracy in protein complex identification.

CN116935972BActive Publication Date: 2026-04-03SHENZHEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-24
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing protein complex identification methods struggle to accurately identify common and unique protein complexes in different tissues or cell lines, and they are unable to capture the dynamic changes of these complexes.

Method used

A multi-network joint clustering method is adopted to obtain protein interaction networks under different states. The networks are divided into common complexes, partially common complexes and unique complexes using an objective function. The L2 regularization term and the unique complex constraint term are combined to perform iterative optimization and discretization calculation to identify protein complexes.

Benefits of technology

It enables more precise identification of protein complexes, reveals their spatial dynamics, and has higher recognition performance and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116935972B_ABST
    Figure CN116935972B_ABST
Patent Text Reader

Abstract

To address the limitations of existing technologies, this invention proposes a multi-network joint clustering method for protein complex identification. By jointly analyzing protein interaction networks under different states, the final clustering results are obtained, corresponding to the protein complexes in the networks. Protein complexes from different networks are classified into common complexes, partially common complexes, and unique complexes. Compared with existing protein complex identification algorithms, this invention can more accurately identify protein complexes and discover their spatial dynamics. Experimental results show that this invention can jointly analyze protein interaction networks under different states and identify different protein complexes, exhibiting more accurate identification performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computational biology, specifically to biological data mining, and more specifically to a multi-network joint clustering method for protein complex identification. Background Technology

[0002] Proteins are essential components of cells and tissues and are the most important material basis for biological activities. They rarely function alone, but rather perform biological functions by interacting with other proteins to form complexes. Therefore, the study of protein complexes helps to elucidate protein functions and understand biological processes, and is crucial for research in biology, pathology, and proteomics. Traditional protein complex recognition is achieved through biological experiments, such as co-immunoexpression (Co-IP) [Clancy T, Hovig E. From proteomes to complexomes in the era of systems biology[J]. Proteomics, 2014, 14(1):24-41] and RNA interference [Cullen LM, Arndt GM. Genome-wide screening for gene function using RNAi in mammalian cells[J]. Immunology and cell biology, 2005, 83(3):217-223]. However, these biological experiments are both time-consuming and difficult to scale up. With the rapid advancement of affinity purification mass spectrometry (AP-MS [Huttlin EL, Bruckner RJ, Navarrete-Perea J, et al. Dual proteome-scale networks reveal cell-specific remodeling of the human interactome[J]. Cell, 2021, 184(11):3022-3040.e28]), an increasing amount of protein-protein interaction (PPI) data has been generated, making it possible to identify protein complexes in protein-protein interaction networks based on computational methods.

[0003] Over the past decade, a series of computational model-based protein complex identification methods have emerged, which mainly utilize the topological information of protein interaction networks to identify potential protein complexes. Existing methods can be broadly categorized into subgraph density-based methods [Bader GD, Hogue CW V. An automated method for finding molecular complexes in large protein interaction networks[J]. BMCbioinformatics, 2003, 4(1): 1-27], seed expansion-based methods [Liu X, Yang Z, Sang S, et al. Detection of protein complexes from multiple protein interaction networks using graph embedding[J]. Artificial Intelligence in Medicine, 2019, 96: 107-115], and core-attachment-based methods [Leung HCM, Xiang Q, Yiu SM, et al. Predicting protein complexes from PPI data: a core-attachment approach[J]. Journal of Computational Biology, 2009, 16(2): 133-144], etc. However, recent research [Tan CSH, Go KD, Bisteau X, et al. Thermal proximity coaggregation for system-wide profiling of protein complex dynamics in cells[J]. Science, 2018, 359(6380):1170-1177] shows that protein-protein interaction networks in organisms change with spatial, temporal, and environmental variations. The methods described above primarily focus on identifying static complexes within individual protein-protein interaction networks and cannot capture the dynamic changes in these complexes.While some methods integrate multi-source data (such as tandem affinity purification and protein domains) to achieve protein complex identification, they still aim to improve the accuracy of complex identification within individual protein interaction networks. In recent years, a series of dynamic network clustering methods have been proposed to identify protein complexes in dynamic protein interaction networks. However, these methods rarely consider the commonalities and differences between protein complexes in different protein interaction networks. Therefore, a new method is urgently needed to accurately identify common and unique protein complexes in different tissues or cell lines. Summary of the Invention

[0004] To address the limitations of existing technologies, this invention proposes a multi-network joint clustering method for protein complex recognition. The technical solution adopted in this invention is as follows:

[0005] A multi-network joint clustering method for protein complex recognition includes the following steps:

[0006] S1, Obtain the protein-protein interaction network to be processed;

[0007] S2, Solve the preset objective function based on the result of step S1; The objective function is constructed based on the following division method: For complexes in protein interaction networks under different states, they are divided into common complexes, partially common complexes and unique complexes; The common complexes exist in two protein interaction networks at the same time, The partially common complexes share some protein members of the two protein interaction networks, and The unique complexes are functional modules that each protein interaction network forms only under specific states.

[0008] S3, Discretize the result of step S2 to obtain the protein-complex allocation matrix, and use the protein-complex allocation matrix as a clustering indicator.

[0009] Compared to existing technologies, this invention obtains clustering results by jointly analyzing protein interaction networks under different states, corresponding to protein complexes in the network, and classifying protein complexes from different networks into common complexes, partially common complexes, and unique complexes. Compared to existing protein complex identification algorithms, this invention can more accurately identify protein complexes and discover their spatial dynamics. Experimental results show that this invention can jointly analyze protein interaction networks under different states and identify different protein complexes, exhibiting more accurate identification performance.

[0010] As a preferred embodiment, for any two protein-protein interaction networks in different states, the objective function includes the following:

[0011]

[0012] in, This is a protein-complex relationship matrix with dimensions N×K. (t) Element value This represents the probability that protein i belongs to complex k, and its value ranges from all positive rational numbers. H (t) =[H c H p(t) H s(t) ], t=1,2; matrix as well as Let H represent the common complex, partially common complex, and unique complex in the t-th protein-protein interaction network, respectively. c element value The matrix H represents the probability that protein i belongs to complex k. p(t) element values ​​in The matrix H represents the probability that protein i belongs to complex l. s(t) The element values ​​in K represent the probability that protein i belongs to complex z; c =K p =K s =K / 5, where K is the total number of complexes in the protein-protein interaction network; This represents the adjacency matrix in the t-th protein-protein interaction network. This indicates that protein i and protein j interact in the t-th protein-protein interaction network. For a weighted network, Represents the probability that protein i and protein j interact; vector θ (t) ∈{0,1} N×1 , A represents (t) It contains information about protein i. Then it means A (t) It does not contain information about protein i. And so on.

[0013] Furthermore, the objective function includes the following constraints to cause specific complexes in the two protein interaction networks to have different protein affinity modes:

[0014]

[0015] Furthermore, the final form of the objective function is:

[0016]

[0017] Where β is the L2 regularization coefficient and λ is the constraint coefficient of the specific complex.

[0018] Furthermore, in step S2, the protein complex relationship matrix is ​​obtained. The optimal solution is obtained in step S3 for the protein complex relationship matrix. The optimal solution is discretized to obtain the protein-complex allocation matrix.

[0019] Furthermore, in step S2, during the random initialization of H... c H p(t) and H s(t) Then, the protein complex relationship matrix is ​​updated alternately. variables in Perform iterative optimization.

[0020] Furthermore, in step S2, the updates are performed alternately according to the following formula:

[0021]

[0022]

[0023]

[0024] Among them, ⊙ and These represent element-level multiplication and division, respectively, with subscripts a, b, and d indicating complexes.

[0025] This invention also includes the following:

[0026] A multi-network joint clustering system for protein complex recognition includes a sequentially connected network acquisition module, an iterative optimization module, and a discretization calculation module, wherein:

[0027] The network acquisition module is used to acquire the protein-protein interaction network to be processed;

[0028] The iterative optimization module is used to solve the preset objective function using the results of the network acquisition module. The objective function is constructed based on the following division method: for complexes in protein interaction networks under different states, they are divided into common complexes, partially common complexes, and unique complexes. The common complexes exist in two protein interaction networks at the same time. The partially common complexes share some protein members of the two protein interaction networks. The unique complexes are functional modules that are formed only in specific states in each protein interaction network.

[0029] The discretization calculation module is used to discretize the results of the iterative optimization module to obtain a protein-complex allocation matrix, which is then used as a clustering indicator.

[0030] A storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the aforementioned multi-network joint clustering method for protein complex recognition.

[0031] A computer device includes a storage medium, a processor, and a computer program stored in the storage medium and executable by the processor, wherein the computer program, when executed by the processor, implements the steps of the multi-network joint clustering method for protein complex recognition as described above. Attached Figure Description

[0032] Figure 1 This is a flowchart of the steps of the multi-network joint clustering method for protein complex recognition provided in Embodiment 1 of the present invention;

[0033] Figure 2 A schematic diagram of the components of each protein complex;

[0034] Figure 3 This is a schematic diagram illustrating the principle of the multi-network joint clustering method for protein complex recognition provided in Embodiment 1 of the present invention.

[0035] Figure 4 This is a parameter effect diagram of the protein complex identification algorithm CSNMF based on multi-network joint clustering in the evaluation experiment of Embodiment 1 of the present invention, showing the changes in parameters β and λ on the simulated dataset, including Recall, F-measure, and AUC.

[0036] Figure 5The evaluation curves of the protein complex identification algorithm CSNMF based on multi-network joint clustering in the experiment of Embodiment 1 of the present invention are shown to reflect the changes in Recall, F-measure and AUC as parameter k changes on a simulated dataset.

[0037] Figure 6 A schematic diagram of the composition of a multi-network joint clustering system for protein complex recognition provided in Embodiment 2 of the present invention. Detailed Implementation

[0038] The accompanying drawings are for illustrative purposes only and should not be construed as limiting the scope of this patent.

[0039] It should be understood that the described embodiments are merely some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of the embodiments of this application.

[0040] The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to limit the embodiments of this application. The singular forms “a,” “the,” and “the” used in the embodiments of this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0041] In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims. In the description of this application, it should be understood that the terms "first," "second," "third," etc., are used only to distinguish similar objects and are not necessarily used to describe a specific order or sequence, nor should they be construed as indicating or implying relative importance. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.

[0042] Furthermore, in the description of this application, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. The invention will be further described below with reference to the accompanying drawings and embodiments.

[0043] To address the limitations of existing technologies, this embodiment provides a technical solution. The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0044] Example 1

[0045] To identify common and unique protein complexes in protein interaction networks under two different states, thereby providing suggestions and references for understanding cell composition and activity, exploring life mechanisms, and other biological research tasks, this embodiment provides a multi-network joint clustering method for protein complex identification. This can be considered a protein complex identification algorithm based on multi-network joint clustering, CSNMF. Please refer to [link / reference]. Figure 1 This includes the following steps:

[0046] S1, Obtain the protein-protein interaction network to be processed;

[0047] S2, Solve the preset objective function using the result of step S1; the objective function is constructed based on the following partitioning method: For complexes in protein-protein interaction networks under different states, please refer to... Figure 2 The complexes are divided into common complexes, partially common complexes, and unique complexes. The common complexes exist in two protein interaction networks simultaneously. The partially common complexes share some protein members of the two protein interaction networks. The unique complexes are functional modules that each protein interaction network forms only under specific conditions.

[0048] S3, Discretize the result of step S2 to obtain the protein-complex allocation matrix, and use the protein-complex allocation matrix as a clustering indicator.

[0049] Compared to existing technologies, this invention obtains clustering results by jointly analyzing protein interaction networks under different states, corresponding to protein complexes in the network, and classifying protein complexes from different networks into common complexes, partially common complexes, and unique complexes. Compared to existing protein complex identification algorithms, this invention can more accurately identify protein complexes and discover their spatial dynamics. Experimental results show that this invention can jointly analyze protein interaction networks under different states and identify different protein complexes, exhibiting more accurate identification performance.

[0050] Protein-protein interactions change under different cellular states and conditions. Therefore, this embodiment divides the complexes in protein-protein interaction networks under two different states into three parts: shared, partially shared, and unique. Shared complexes exist in both networks simultaneously; partially shared complexes between the two networks share some protein members; and network-specific complexes are functional modules that form only under specific states. For a detailed explanation of the symbols, please refer to [link to documentation]. Figure 3 :

[0051] set up and Let represent the adjacency matrix of two protein interaction networks, where This indicates that protein i and protein j interact in network t, while in a weighted network, This represents the probability of two proteins interacting. Due to differences in gene expression and limitations in biological experimental techniques, the number of proteins covered by each network varies. Therefore, this embodiment introduces a vector θ. () ∈{0,1} N×1 ,t=1,2, where A represents () It contains information about protein i. Then it means A () The data does not contain information about protein i. The purpose of protein complex identification is to obtain the protein-protein interaction network A. () The subgraph in is denoted as Where K () Let be the number of protein complexes in network t. The goal of this embodiment is to infer a clustering indicator matrix from multiple protein-protein interaction networks. and This goal can be represented as {A} (1) A (2)}→{H (1)* H (2)*},in This indicates that protein i belongs to complex j, and vice versa.

[0052] This embodiment defines a protein-complex relationship matrix. To represent the probability that protein i belongs to complex k, a complex affinity matrix is ​​defined. The affinity fraction used to describe whether protein i and protein j may belong to the same complex is expressed as:

[0053]

[0054] This embodiment minimizes A (t) with U (t) The difference between them, from the adjacency matrix A (t) Inferring potential U(t) As an optional implementation, KL divergence is used to measure U. (t) With A (t) The difference between them is expressed as:

[0055]

[0056] Next, substitute formula (1) into formula (2) and remove the constant term. The above formula can then be modified as follows:

[0057]

[0058] By minimizing formula (3), we can obtain A (t) Estimating H (t) .

[0059] Let H (t) =[H c H p(t) H s(t) The terms ], t = 1, 2 are used to characterize shared complexes, partially shared complexes, and unique complexes, where Let K represent the common complex, partially common complex, and unique complex in the t-th network, respectively. In this embodiment, K is set... c =K p =K s =K / 5, where K is the total number of complexes in the two protein interaction networks. Substituting the above settings into formula (3), a preferred embodiment can be obtained. For any two protein interaction networks in different states, the objective function includes the following:

[0060]

[0061] in, This is a protein-complex relationship matrix with dimensions N×K. (t) Element value This represents the probability that protein i belongs to complex k, and its value ranges from all positive rational numbers. H (t) =[H c H p(t) H s(t) ], t=1,2; matrix as well as Let H represent the common complex, partially common complex, and unique complex in the t-th protein-protein interaction network, respectively. c element value The matrix H represents the probability that protein i belongs to complex k. p(t) element values ​​in The matrix H represents the probability that protein i belongs to complex l.s(t) The element values ​​in K represent the probability that protein i belongs to complex z; c =K p =K s =K / 5, which is the total number of complexes in the protein-protein interaction network; This represents the adjacency matrix in the t-th protein-protein interaction network. This indicates that protein i and protein j interact in the t-th protein-protein interaction network. For a weighted network, Represents the probability that protein i and protein j interact; vector θ (t) ∈{0,1} N×1 , A represents (t) It contains information about protein i. Then it means A (t) It does not contain information about protein i. And so on.

[0062] To prevent overfitting, this embodiment introduces an L2 regularization term. Since both networks contain their own unique complexes, this embodiment assumes that the unique complexes in the two networks have different protein affinity modes. Therefore, the objective function further includes the following constraints to ensure that the unique complexes in the two protein interaction networks have different protein affinity modes:

[0063]

[0064] Furthermore, the final form of the objective function is:

[0065]

[0066] Where β is the L2 regularization coefficient and λ is the constraint coefficient of the specific complex.

[0067] The protein complex relationship matrix is ​​defined as follows: This represents the complexes identified from different networks. Due to the optimal solution... All are continuous values; this embodiment will... Discretized into the final protein-complex allocation matrix H * In this invention, in order to obtain the overlapping complex, for each protein i, this embodiment first... Sort the i-th row in descending order using... Indicates. If and The gap between them is the largest, and at the same time but otherwise Thus, if K iIf the value is greater than 1, protein i can belong to multiple complexes. Ultimately, the protein-complex assignment matrix H corresponding to the two networks can be obtained. (1)* =[H c* H p (1)* H s(1)* ] and H (2)* =H c* H p(2)* H s(2)* ].

[0068] Therefore, further, in step S2, the protein complex relationship matrix is ​​obtained. The optimal solution is obtained in step S3 for the protein complex relationship matrix. The optimal solution is discretized to obtain the protein-complex allocation matrix.

[0069] Furthermore, in step S2, during the random initialization of H... c H p(t) and H s(t) Then, the protein complex relationship matrix is ​​updated alternately. variables in Perform iterative optimization.

[0070] Furthermore, in step S2, the updates are performed alternately according to the following formula:

[0071]

[0072]

[0073]

[0074] Among them, ⊙ and These represent element-level multiplication and division, respectively, with subscripts a, b, and d indicating complexes.

[0075] In step S2, the iterative update stops after a preset condition is met. As an optional embodiment, the algorithm stops iterative updating when the relative change of the objective function is less than 1e-6 or the number of iterations reaches a predefined maximum value (which can be set to 200).

[0076] Next, in order to further demonstrate the effectiveness of the solution of the present invention, this embodiment conducts evaluation experiments on two datasets to evaluate the method of this embodiment.

[0077] The dataset consists of a simulated dataset and a real dataset, BioPlex. The simulated dataset contains 3000 nodes, which are divided into two overlapping subsets, each containing 2350 nodes, corresponding to two interacting networks. There are 50 shared clusters C between the two networks. com And 20 pairs of partially shared clusters C pc Each network contains 30 unique clusters C sp , where |C com |=30,|C pc |=20,|C sp |=15. Each pair of partially common clusters has 10 overlapping nodes. This embodiment uses probability p. in Generate edges within a cluster using probability p. out Generate edges between clusters using probability p noise Generate edges between each pair of nodes in the network. In this dataset, this embodiment sets p... in =0.3,p out =0.005,p noise =0.01.

[0078] The BioPlex dataset contains two cell line protein interaction networks (293T and HCT116), constructed using affinity purification mass spectrometry (AP-MS) experiments. 293T has 118,163 interactions among its 13,957 proteins, while HCT116 has 70,967 interactions among its 10,115 proteins, with 9,521 shared proteins between the two networks. The dataset itself was derived from the standard human complex dataset CORUM 4.0, which extracted 1,968 complexes covering 2,458 proteins as a reference dataset.

[0079] Evaluation metrics. This embodiment uses four metrics for evaluation, including recall, F-measure, geometric precision (ACC), and the score of matched complexes (FRAC).

[0080] Let b i It is one of the predicted complexes B, q j This is a complex from the reference dataset Q. This embodiment first introduces an overlap score to measure the identified complex b. i With reference complex q j Similarities between them:

[0081]

[0082] Where |·| represents the number of proteins in the complex. If OS(b i ,q j When ω ≥ ω, the identified complex b iIt is considered to be related to the reference complex q j Matching. We set the value of ω to 0.25 in all experiments.

[0083] In this embodiment, Tp (true positive) is the number of predicted complexes that match the reference complex, FN (false negative) is the number of reference complexes that do not match the predicted complex, and FP (false positive) is the number of predicted complexes minus TP. The calculations for Recall, Precision, and F-measure are as follows:

[0084]

[0085]

[0086]

[0087] Geometric accuracy (SCC) is calculated from sensitivity and positive predictive value as follows:

[0088]

[0089]

[0090]

[0091] Sensitivity (Sn) represents the number of identified proteins covered in the reference complex, and positive predictive value (PPV) represents the probability that the identified protein complex is a true positive.

[0092] FRAC is defined by equation (16), which describes the ratio of the matched reference complex to the reference dataset:

[0093]

[0094] The evaluation metrics described above rely on a reference dataset. However, reference datasets are often incomplete, and the identified complexes may be biologically valid but not yet described. Therefore, this embodiment introduces the following P-value:

[0095]

[0096] Where |V| and |q j | These represent the number of proteins and the recognized complexes in the PPI network, respectively.

[0097] Parameter setting and effect evaluation. The method in this embodiment has four parameters: K, β, σ, and λ. K is the total number of complexes, β is the penalty coefficient for L2 regularization, σ is the tradeoff coefficient between the common part and other parts of L2 regularization, and λ controls the difference between two specific parts. In all experiments, the value of σ is set to 2. This embodiment uses a grid search to find the optimal parameters. On the synthetic dataset, λ is {2...} 0 ,2 1 , ..., 2 10 ,2 11}, {100,200,…,900,1000}; on the human protein dataset, λ is {2 -4 ,2 -3 ,…,2 3 ,2 4} and K are {5000, 10000, 15000}. We performed parameter sensitivity analysis on the three hyperparameters K, β, and λ on a simulated dataset. Figure 4 The results show that the performance of the method in this embodiment on network 1 varies with a single hyperparameter (β or λ), while other parameters remain constant. It can be observed that when λ is fixed and β is no greater than 2... 10 When β is constant, the ACC, Recall, and F-measure of this method increase with increasing β, but when β equals 2... 11 When β is constant, the method will not work. When λ is constant, ACC increases with increasing λ, indicating that specific constraint terms have a significant impact on λ, leading to improved performance. This embodiment also analyzes the impact of K on the performance of the method in this embodiment. For example... Figure 5 As shown, when K ≤ 300, ACC, Recall, and F-measure increase with the increase of the total number of complexes K. However, when K > 300, ACC, Recall, and F-measure tend to plateau.

[0098] This embodiment conducted experiments on simulated datasets and BioPlex datasets, comparing the CSNMF method of this embodiment with existing comparative methods such as PSMVC [Ou-Yang L, Zhang XF, Dai DQ, et al. Protein complex detection based on partially shared multi-view clustering[J]. BMC bioinformatics, 2016, 17(1): 1-15.], and the experimental results are shown in the following tables.

[0099] Table 1 shows the performance comparison of the protein complex identification algorithm CSNMF, based on multi-network joint clustering, with four other algorithms on a simulated dataset. The results are presented as the number of protein complexes identified, the number of proteins covered by the identification results, FRAC, Recall, F-measure, and ACC. Bold numbers in the table represent the highest scores.

[0100] As shown in Table 1, on the simulated dataset, the method of this embodiment significantly outperforms single-network clustering methods ClusterOne [Nepusz T, Yu H, Paccanaro A. Detecting overlapping protein complexes in protein-protein interaction networks[J]. Nature methods, 2012, 9(5): 471-472.], IPCA [Li M, Chen J, Wang J, et al. Modifying the DPClusalgorithm for identifying protein complexes based on new topological structures[J]. BMC bioinformatics, 2008, 9(1): 1-16.] and NCMine [Tadaka S, Kinoshita K. NCMine: Core-peripheral based functional module detection using near-cliquemining[J]. Bioinformatics, 2016, 32(22): 3454-3460] on all metrics.

[0101] Table 1

[0102]

[0103]

[0104] Table 2 shows a performance comparison between the protein complex identification algorithms CSNMF and PSMVC, which are based on multi-network joint clustering, on the BioPlex dataset. The results are presented as the number of protein complexes identified, the number of proteins covered by the identification results, FRAC, Recall, and ACC. Bold numbers in the table represent the highest scores.

[0105] Table 3 shows the gene enrichment analysis results of the protein complex identification algorithms CSNMF and PSMVC, which are based on multi-network joint clustering, on the BioPlex dataset. The bolded numbers in the table represent the highest scores.

[0106] As shown in Table 2, compared with PSMVC based on multi-view clustering, the method in this embodiment outperforms PSMVC in terms of F-measure and ACC. On the BioPlex dataset, the method in this embodiment improves FRAC and Recall by 4% and 2.8% respectively compared with PSMVC. Furthermore, as shown by the gene enrichment analysis results in Table 3, the model in this embodiment can identify more meaningful protein complexes compared with PSMVC.

[0107] Table 2

[0108] Method #complex #protein FRAC Recall ACC PSMVC 4368 11528 0.273 0.149 0.621 CSNMF 4323 11215 0.313 0.177 0.618

[0109] Table 3

[0110]

[0111]

[0112] Example 2

[0113] A multi-network joint clustering system for protein complex recognition, please refer to [link to relevant documentation]. Figure 6 It includes a network acquisition module 1, an iterative optimization module 2, and a discretization calculation module 3 connected in sequence, wherein:

[0114] The network acquisition module 1 is used to acquire the protein-protein interaction network to be processed;

[0115] The iterative optimization module 2 is used to solve the preset objective function using the results of the network acquisition module 1. The objective function is constructed based on the following division method: for complexes in protein interaction networks under different states, they are divided into common complexes, partially common complexes, and unique complexes. The common complexes exist in two protein interaction networks at the same time. The partially common complexes share some protein members of the two protein interaction networks. The unique complexes are functional modules that are formed only in specific states in each protein interaction network.

[0116] The discretization calculation module 3 is used to discretize the results of the iterative optimization module 2 to obtain the protein-complex allocation matrix, and the protein-complex allocation matrix is ​​used as a clustering indicator.

[0117] Example 3

[0118] A storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the multi-network joint clustering method for protein complex recognition as described in Example 1.

[0119] Example 4

[0120] A computer device includes a storage medium, a processor, and a computer program stored in the storage medium and executable by the processor, wherein the computer program, when executed by the processor, implements the steps of the multi-network joint clustering method for protein complex recognition as described in Example 1.

[0121] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation of the present invention. Those skilled in the art can make other variations or modifications based on the above description. It is neither necessary nor possible to exhaustively describe all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the claims of the present invention.

Claims

1. A multi-network joint clustering method for protein complex recognition, characterized in that, Includes the following steps: S1, Obtain the protein-protein interaction network to be processed; S2, Solve the preset objective function using the result of step S1; the objective function is constructed based on the following partitioning: for complexes in protein-protein interaction networks under different states, they are divided into common complexes, partially common complexes, and unique complexes; the common complexes exist simultaneously in two protein-protein interaction networks, the partially common complexes share some protein members of the two protein-protein interaction networks, and the unique complexes are functional modules formed by each protein-protein interaction network only under specific states; for any two protein-protein interaction networks under different states, the objective function includes the following: ; in, This is a protein-complex relationship matrix with dimension 1. Element value Indicates protein Belongs to complex The probability, taking values ​​from all positive rational numbers. ; ;matrix , as well as They represent the first A matrix of shared complexes, partially shared complexes, and unique complexes in a protein-protein interaction network. element value Indicates protein Belongs to complex Possibilities, matrix element values ​​in Indicates protein Belongs to complex Possibilities, matrix element values ​​in Indicates protein Belongs to complex The possibility; , This represents the total number of complexes in the protein-protein interaction network; Indicates the first Adjacency matrix in a protein-protein interaction network Indicates protein With protein In the In a protein-protein interaction network, there are interactions; for a weighted network... Indicates protein With protein The probability of interaction; vector , express Contains protein Information, Then it means It does not contain protein. Information, And so on; The objective function includes the following constraints to cause specific complexes in two protein interaction networks to have different protein affinity modes: ; The final form of the objective function is: ; in, The coefficients of the L2 regularization term. For the constraint coefficients of the specific complex; In step S2, the protein complex relationship matrix is ​​obtained. The optimal solution is obtained in step S3 for the protein complex relationship matrix. The optimal solution is discretized to obtain the protein-complex allocation matrix; In step S2, during random initialization , as well as Then, the protein complex relationship matrix is ​​updated alternately. variables in Iterative optimization is then performed. S3, Discretize the result of step S2 to obtain the protein-complex allocation matrix, and use the protein-complex allocation matrix as a clustering indicator.

2. The multi-network joint clustering method for protein complex recognition according to claim 1, characterized in that, In step S2, the updates are performed alternately according to the following formula: ; ; ; in, and These represent element-wise multiplication and division, respectively, with subscripts... , , Indicates a complex.

3. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the steps of the multi-network joint clustering method for protein complex recognition as described in claim 1 or 2.

4. A computer device, characterized in that: It includes a storage medium, a processor, and a computer program stored in the storage medium and executable by the processor, wherein the computer program, when executed by the processor, implements the steps of the multi-network joint clustering method for protein complex recognition as described in claim 1 or 2.