Method for constructing protein interaction network of Chinese mitten crab
By combining the PINA and STRING database methods, the Chinese mitten protein interaction network was constructed, which solved the problems of false positive results and method differences in the existing technology, and achieved the construction of a more accurate and comprehensive protein interaction network.
Patent Information
- Application Number
- CN202210045867.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-17
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-01-17
AI Technical Summary
The prior art has a large number of false positive results and method differences in building protein interaction networks, making it difficult to build an accurate protein interaction network.
Combined with the PINA and STRING databases, a Chinese mitten protein interactor network based on reference model organisms was constructed by comparing the protein sequences of Chinese mitten mitten mitten idyllic gene sequences with reference model organisms, and the final protein interaction network was obtained through fusion and expansion.
It improves the accuracy and coverage of the Chinese mitten protein interaction network and provides richer biological analysis information.
Smart Images

Figure CN114822687B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for constructing a protein interaction network, and in particular to a method for constructing a protein interaction network of Chinese mitten crab. Background Art
[0002] Protein interaction networks are a hot topic in systems biology. By combining computer models and experimental techniques, complex biological systems can be analyzed from a systemic perspective, providing guidance and predictions for biological experiments and systems biology research [1]. Protein interaction networks reflect the interactions between proteins in cells. The network not only contains a large amount of signal transduction information, but also includes protein relationships in gene transcription regulation, metabolism, immunity, etc. The network has scale-free and small-world characteristics. After 20 years of development, the genome-scale protein interaction networks that have been constructed cover a variety of biological types such as viruses, bacteria, eukaryotic single-cell organisms, plants, insects, vertebrates, etc., and have been applied to functional prediction, evolutionary analysis, disease research, drug development, cell engineering, etc. [2,3].
[0003] Many model organisms use high-throughput experimental methods to construct protein interaction networks. For example, the Escherichia coli protein interaction network is constructed by tandem affinity chromatography and protein in vitro binding experiments [4,5], and the Saccharomyces cerevisiae protein interaction network is constructed by yeast two-hybrid method [6]. However, there are great problems in high-throughput experimental technology. High-throughput experiments often produce a large number of false positive results, and the protein interaction networks obtained by different methods vary greatly [7,8]. Therefore, researchers began to predict protein interaction relationships and construct protein interaction networks by computer. The method based on homology alignment is the main method for constructing protein interaction networks by computer. This method compares the gene sequence or protein sequence of the target species with the protein sequence of the model organism, and refers to the known protein interaction relationships of the model organism. According to the principle that homologous proteins have similar protein interaction relationships, the protein interaction relationships of the target species are predicted, thereby constructing a protein interaction network. This method is widely applicable to various organisms.
[0004] References
[0005] 1. Snider, J.; Kotlyar, M.; Saraon, P.; Yao, Z.; Jurisica, I.; Stagljar, I., Fundamentals of protein interaction network mapping. Molecular systems biology2015, 11, (12), 848. https: / / doi.org / 10.15252 / msb.20156351
[0006] 2.Hao, T.; Peng, W.; Wang, Q.; Wang, B.; Sun, J., Reconstruction and Application of Protein-Protein Interaction Network. International journal of molecular sciences 2016, 17, (6). https: / / doi.org / 10.3390 / ijms17060907
[0007] 3. Wang Fen, Pei Huimin, Wen Di, Li Jing, Research on protein-protein interactions. Biochemical Engineering, 2020, 6, (06), 111-113.
[0008] 4. Butland, G.; Peregrin-Alvarez, JM; Li, J.; Yang, W.; Yang, coli.Nature 2005,433,(7025),531-7.https: / / doi.org / 10.1038 / nature03239
[0009] 5.Arifuzzaman,M.1Maeda,M.Itoh,A.Nishikata,K.Takita,C.Saito,R.Ara,T.Nakahigashi,K.Huang,HC,Hirai,A.Tsuzuki,K.Nakamura,S.Alt af-Ul-Amin, M. and Oshima, T. and Baba, T. and Yamamoto, N. and Kawamura, T. and Ioka-Nakamichi, T. and Kitagawa, M. and Tomita, M. and Kanaya, S. and Wada, C. and Mori, H. ,Large-scale Identification of protein-protein interactions of Escherichia coli K-12.Genome research 2006,16,(5),686-91
[0010] 6.Ito,T.andChiba,T.Ozawa,R.andYoshida,M.Hattori,M.andSakaki,Y.,A comprehensive two-hybrid analysis of the yeast protein interactome 2001,98,(8),4569-74
[0011] 7.Sprinzak,E.Sattath,S.andMargalit,H.,How reliable are experimentalprotein-protein interaction data?Journal of Molecular Biology 2003,327,(5),919-23
[0012] 8. Deane, CM; Salwinski, L.; Summary of the invention
[0013] The technical problem to be solved by the present invention is to provide a method for constructing a protein interaction network of Chinese mitten crab by combining PINA and STRING databases
[0014] The technical solution adopted by the present invention is: a method for constructing a protein interaction network of Chinese mitten crab, comprising the following steps:
[0015] 1) Data collection: including
[0016] (1.1) Obtaining the gene sequence information of Chinese mitten crab through the Chinese mitten crab genome or transcriptome;
[0017] (1.2) Using organisms in the PINA database as reference model organisms, and obtaining protein sequence information of each reference model organism from the Uniprot database;
[0018] (1.3) Obtain protein interaction information of each reference model organism from the PINA database;
[0019] (1.4) Obtain protein interaction information of each reference model organism from the STRING database;
[0020] 2) Construct a protein interaction subnetwork of Chinese mitten crab based on the reference model organism;
[0021] 3) Fusion of multiple protein interaction subnetworks of Chinese mitten crab based on reference model organisms;
[0022] 4) The fusion network of Chinese mitten crab was expanded to obtain the protein interaction network of Chinese mitten crab.
[0023] Step 2) includes:
[0024] (2.1) Using the BLASTX analysis tool, the gene sequence information of Chinese mitten crab was compared with the protein sequence information of reference model organisms obtained from the Uniprot database, and the E-value was set to 1e-005;
[0025] (2.2) The first protein sequence of the reference model organism in the comparison is used as the homologous sequence of the Chinese mitten crab gene sequence, thereby determining the reference model organism protein and Chinese mitten crab gene with homologous relationship;
[0026] (2.3) Extracting the protein interaction relationship between the reference model organism proteins homologous to the Chinese mitten crab genes and the protein interaction relationship between the proteins according to the protein interaction information of the reference model organisms obtained from the PINA database;
[0027] (2.4) The proteins and protein interactions of the reference model organisms constitute a Chinese mitten crab protein interaction subnetwork based on the reference model organism; each reference model organism forms a Chinese mitten crab protein interaction subnetwork based on the reference model organism, and several reference model organisms form several Chinese mitten crab protein interaction subnetworks based on the reference model organism.
[0028] Step 3) includes:
[0029] (3.1) Determine the fusion order of the protein interaction subnetwork based on the reference model organism Chinese mitten crab;
[0030] (3.2) The fusion order is determined by the closeness of the relationship between the reference model organism and the Chinese mitten crab. The relationship is sorted from close to distant, and the closest relationship is fused first;
[0031] (3.3) After determining the fusion order, the Chinese mitten crab protein interaction subnetworks based on the reference model organism are fused pairwise to obtain the Chinese mitten crab fusion network; each fusion must first determine the relationship between the nodes of the two fused Chinese mitten crab protein interaction subnetworks based on the reference model organism, that is, the homology relationship between the reference model organism proteins corresponding to the two fused Chinese mitten crab protein interaction subnetworks based on the reference model organism.
[0032] The pairwise fusion of the protein interaction subnetworks of the reference model organism Chinese mitten crab specifically includes:
[0033] (3.3.1) The relatively close relationship in the two protein interaction subnetworks of the reference model organism Chinese mitten crab that are first fused is set as the target network, and the relatively distant relationship is set as the query network. The target network and the query network are judged by protein sequence alignment whether they have the same protein interaction relationship. Specifically, by protein sequence alignment, it is determined that the two interacting proteins in the target network are homologous to the two interacting proteins in the query network. If the target network and the query network have the same protein interaction relationship, the same protein interaction relationship in the target network is extracted as part of the fusion network; the different protein interaction relationships in the target network and the query network are added to the fusion network to finally form a fusion network;
[0034] (3.3.2) The formed fusion network is used as the new target network, and the next Chinese mitten crab protein interaction subnetwork based on the reference model organism is used as the new query network for the second fusion. If there are N Chinese mitten crab protein interaction subnetworks based on the reference model organism, they need to be fused N-1 times to finally obtain the Chinese mitten crab fusion network.
[0035] Step 4) includes:
[0036] (4.1) Using all protein names contained in the Chinese mitten crab genome annotation information as search terms, select a model organism that is the same as the reference model organism in the STRING database one at a time, search and extract the protein interaction relationship of the model organism, until all reference model organisms are selected and the protein interaction relationship of the corresponding model organism is searched and extracted;
[0037] (4.2) The STRING database provides a confidence score for protein interactions. The confidence score ranges from 0 to 1. The closer it is to 1, the higher the confidence. After searching, only protein interactions with a confidence score of 0.9 or above are retained.
[0038] (4.3) removing duplicates from the retained protein interactions with a confidence score above 0.9 to obtain expanded data;
[0039] (4.4) The expanded data were merged with the Eriocheir sinensis fusion network, and repeated protein interaction relationships were removed to obtain the final Eriocheir sinensis protein interaction network.
[0040] The method for constructing the protein interaction network of Chinese mitten crab of the present invention has the following beneficial effects:
[0041] 1. In the method of the present invention, the gene sequence information of Chinese mitten crab can be obtained not only from the Chinese mitten crab genome, but also from the Chinese mitten crab transcriptome. The source of the gene sequence information of Chinese mitten crab is more extensive.
[0042] 2. In the method of the present invention, the protein interaction information of the reference model organisms is taken from two databases, the PINA database and the STRING database. The source of the protein interaction information of the reference model organisms is more extensive.
[0043] 3. The reference model biological protein interaction information contained in the PINA database is high-quality information that has been manually corrected, while the reference model biological protein interaction relationship searched in the STRING database only retains the part with a credibility score above 0.9. Therefore, the reference model biological protein interaction information used in the method of the present invention is of higher quality.
[0044] 4. The method for determining the fusion order of the Chinese mitten crab protein interaction subnetwork is simple and easy to operate. When the method of the present invention fuses the Chinese mitten crab protein interaction subnetwork based on the reference model organism, the order of fusion of the Chinese mitten crab protein interaction subnetwork based on the reference model organism is determined by the distance of the relationship between the reference model organism and the Chinese mitten crab. The relationship is sorted from close to far, and the close relationship is fused first. This method is simple and easy to operate.
[0045] 5. When this method fuses the protein interaction subnetworks of the Chinese mitten crab based on the reference model organism, the fused network includes both the same protein interaction relationships in the target network and the query network, as well as the different protein interaction relationships in the target network and the query network. The information of the fused network is more comprehensive.
[0046] 6. In the expansion step of the Chinese mitten crab fusion network, this method directly uses all the protein names contained in the Chinese mitten crab genome annotation information as search items to search in the STRING database. The method is simple and easy to operate.
[0047] 7. The protein interaction network of Chinese mitten crab constructed using the method of the present invention is currently the largest protein interaction network of Chinese mitten crab with the largest number of proteins and protein interaction relationships. Therefore, it provides richer information for biological analysis using this network. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 The present invention is a flow chart of the method for constructing the protein interaction network of Chinese mitten crab. DETAILED DESCRIPTION
[0049] The method for constructing the protein interaction network of Chinese mitten crab of the present invention is described in detail below with reference to the embodiments and the accompanying drawings.
[0050] like Figure 1 As shown, the method for constructing the protein interaction network of Chinese mitten crab of the present invention comprises the following steps:
[0051] 1) Data collection: including
[0052] (1.1) Obtaining the gene sequence information of Chinese mitten crab through the Chinese mitten crab genome or transcriptome;
[0053] (1.2) Using organisms in the PINA database as reference model organisms, and obtaining protein sequence information of each reference model organism from the Uniprot database;
[0054] (1.3) Obtain protein interaction information of each reference model organism from the PINA database;
[0055] (1.4) Obtain protein interaction information of each reference model organism from the STRING database.
[0056] 2) Construct a protein interaction subnetwork based on the reference model organism Chinese mitten crab; including:
[0057] (2.1) Using the BLASTX analysis tool, the gene sequence information of Chinese mitten crab was compared with the protein sequence information of reference model organisms obtained from the Uniprot database, and the E-value was set to 1e-005;
[0058] (2.2) The first protein sequence of the reference model organism in the comparison is used as the homologous sequence of the Chinese mitten crab gene sequence, thereby determining the reference model organism protein and Chinese mitten crab gene with homologous relationship;
[0059] (2.3) Extracting the protein interaction relationship between the reference model organism proteins homologous to the Chinese mitten crab genes and the protein interaction relationship between the proteins according to the protein interaction information of the reference model organisms obtained from the PINA database;
[0060] (2.4) The proteins and protein interactions of the reference model organisms constitute a Chinese mitten crab protein interaction subnetwork based on the reference model organism; each reference model organism forms a Chinese mitten crab protein interaction subnetwork based on the reference model organism, and several reference model organisms form several Chinese mitten crab protein interaction subnetworks based on the reference model organism.
[0061] 3) Fusion of multiple protein interaction subnetworks of Chinese mitten crab based on reference model organisms; including:
[0062] (3.1) Determine the fusion order of the protein interaction subnetwork based on the reference model organism Chinese mitten crab;
[0063] (3.2) The fusion order is determined by the closeness of the relationship between the reference model organism and the Chinese mitten crab. The relationship is sorted from close to distant, and the closest relationship is fused first;
[0064] (3.3) After determining the fusion order, the Chinese mitten crab protein interaction subnetworks based on the reference model organism are fused pairwise to obtain the Chinese mitten crab fusion network; each fusion must first determine the relationship between the nodes of the two fused Chinese mitten crab protein interaction subnetworks based on the reference model organism, that is, the homology relationship between the reference model organism proteins corresponding to the two fused Chinese mitten crab protein interaction subnetworks based on the reference model organism.
[0065] The pairwise fusion of the protein interaction subnetworks of the reference model organism Chinese mitten crab specifically includes:
[0066] (3.3.1) The relatively close relationship in the two protein interaction subnetworks of the reference model organism Chinese mitten crab that are first fused is set as the target network, and the relatively distant relationship is set as the query network. The target network and the query network are judged by protein sequence alignment whether they have the same protein interaction relationship. Specifically, by protein sequence alignment, it is determined that the two interacting proteins in the target network are homologous to the two interacting proteins in the query network. If the target network and the query network have the same protein interaction relationship, the same protein interaction relationship in the target network is extracted as part of the fusion network; the different protein interaction relationships in the target network and the query network are added to the fusion network to finally form a fusion network;
[0067] (3.3.2) The formed fusion network is used as the new target network, and the next Chinese mitten crab protein interaction subnetwork based on the reference model organism is used as the new query network for the second fusion. If there are N Chinese mitten crab protein interaction subnetworks based on the reference model organism, they need to be fused N-1 times to finally obtain the Chinese mitten crab fusion network.
[0068] 4) Expand the fusion network of Chinese mitten crab to obtain the protein interaction network of Chinese mitten crab;
[0069] Only part of the gene sequence information of Chinese mitten crab is included in the Chinese mitten crab fusion network because the PINA database contains less protein interaction information. Therefore, in order to obtain more complete protein interaction data, the STRING database is used to expand the Chinese mitten crab fusion network. The expansion specifically includes:
[0070] (4.1) Using all protein names contained in the Chinese mitten crab genome annotation information as search terms, select a model organism that is the same as the reference model organism in the STRING database one at a time, search and extract the protein interaction relationship of the model organism, until all reference model organisms are selected and the protein interaction relationship of the corresponding model organism is searched and extracted;
[0071] (4.2) The STRING database provides a confidence score for protein interactions. The confidence score ranges from 0 to 1. The closer it is to 1, the higher the confidence. After searching, only protein interactions with a confidence score of 0.9 or above are retained.
[0072] (4.3) removing duplicates from the retained protein interactions with a confidence score above 0.9 to obtain expanded data;
[0073] (4.4) The expanded data were merged with the Eriocheir sinensis fusion network, and repeated protein interaction relationships were removed to obtain the final Eriocheir sinensis protein interaction network.
[0074] Here are some examples:
[0075] Using the transcriptome of Chinese mitten crab as the source of gene sequences, and selecting Drosophila, nematodes, humans, house mice, brown rats, and yeast as reference model organisms, a protein interaction network of Chinese mitten crab was constructed, which contains 8225 proteins and 148524 protein interaction relationships. The specific method is as follows:
[0076] Step 1: Data Collection:
[0077] (1) The gene sequence data of Chinese mitten crab was obtained from the transcriptome data of the GEO database. The data included the transcriptome data of five tissues of Chinese mitten crab juveniles: gills, eyestalks, thoracic nerve clusters, hepatopancreas, and muscles. The data contained 246,232 gene sequences and their functional annotations.
[0078] (2) 22,006 Drosophila protein sequences, 26,618 C. elegans protein sequences, 70,236 human protein sequences, 50,694 House mouse protein sequences, 29,982 Rattus norvegicus protein sequences, and 6,721 yeast protein sequence data were obtained from the Uniprot database;
[0079] (3) From the PINA database, 52,280 Drosophila protein interactions, 14,829 C. elegans protein interactions, 166,776 human protein interactions, 13,865 house mouse protein interactions, 3,804 brown rat protein interactions, and 94,414 yeast protein interactions were obtained;
[0080] (4) From the STRING database, we obtained 7,963,454 Drosophila protein interactions, 5,778,268 C. elegans protein interactions, 11,353,056 human protein interactions, 12,614,042 house mouse protein interactions, 13,433,196 brown rat protein interactions, and 2,007,134 yeast protein interactions.
[0081] Step 2: Construct a protein interaction subnetwork based on the reference model organism Chinese mitten crab:
[0082] The BLASTX analysis tool was used to align the gene sequence of Chinese mitten crab with the protein sequence of the reference model organism. The E-value was set to 1e-005 during the alignment. The protein interaction subnetwork of Chinese mitten crab protein based on the reference model organism was constructed according to the protein interaction relationship of the reference model organism and the homologous relationship with the gene sequence of Chinese mitten crab. Each reference model organism corresponds to a subnetwork, and a total of 6 Chinese mitten crab protein interaction subnetworks based on the reference model organism were obtained. The results are shown in Table 1.
[0083] Table 1 Protein interaction subnetworks of Eriocheir sinensis based on six reference model organisms
[0084]
[0085] The third step: perform fusion of the protein interaction subnetworks of the reference model organism Eriocheir sinensis to obtain the Eriocheir sinensis fusion network:
[0086] According to the kinship, the fusion order was determined as the Chinese mitten crab protein interaction subnetwork based on fruit flies, the Chinese mitten crab protein interaction network based on nematodes, the Chinese mitten crab protein interaction subnetwork based on humans, the Chinese mitten crab protein interaction subnetwork based on house mice, the Chinese mitten crab protein interaction subnetwork based on brown rats, and the Chinese mitten crab protein interaction subnetwork based on yeast. The six Chinese mitten crab protein interaction subnetworks based on reference model organisms were fused, and after five fusions, the Chinese mitten crab fusion network was finally obtained, which contained 6015 proteins and 108034 protein interactions.
[0087] Step 4: Expand the fusion network of Chinese mitten crab to obtain the protein interaction network of Chinese mitten crab:
[0088] There are 246,232 gene sequences in the transcriptome data of Chinese mitten crab, and 25,047 gene sequences are included in the fusion network of Chinese mitten crab, accounting for only 10.2% of all gene sequences. The remaining gene sequences have not been aligned to have interactions. The annotation information of the transcriptome of Chinese mitten crab corresponds to 19,338 proteins. The protein list composed of these proteins was used to select six model organisms, namely fruit flies, nematodes, humans, house mice, brown rats, and yeast, to search for protein interactions in the STRING database. After 6 searches, protein interactions from 6 model organisms were obtained, and these interactions were merged to retain only protein interactions with a credibility of more than 0.9. The duplicates of the protein interactions with a credibility score of more than 0.9 were removed, and the final expanded data contained 5,545 proteins and 51,740 protein interactions. Finally, the expanded data was merged with the Chinese mitten crab fusion network, and the duplication was removed to complete the expansion of the Chinese mitten crab fusion network, and finally the Chinese mitten crab protein interaction network was obtained, which contained 8225 proteins and 148524 protein interaction relationships. These proteins corresponded to 31507 genes in the Chinese mitten crab transcriptome. After the expansion, the number of proteins in the protein interaction network increased by 36.7% compared with that before the expansion, and the corresponding number of Chinese mitten crab genes increased by 25.8% compared with that before the expansion.
Claims
1. A method for constructing a protein interaction network of Chinese mitten crab, It is characterized in that The steps include: 1) Data collection: including (1.1) Obtaining the gene sequence information of Chinese mitten crab through its genome or transcriptome; (1.2) Using organisms in the PINA database as reference model organisms, and obtaining protein sequence information of each reference model organism from the Uniprot database; (1.3) Obtain protein interaction information of each reference model organism from the PINA database; (1.4) Obtain protein interaction information of each reference model organism from the STRING database; 2) Construct a protein interaction subnetwork based on the reference model organism Chinese mitten crab; 3) Fusion of multiple protein interaction subnetworks of Chinese mitten crab based on reference model organisms; including: (3.1) Determine the fusion order of the protein interaction subnetwork based on the reference model organism Chinese mitten crab; (3.2) The fusion order is determined by the closeness of the relationship between the reference model organism and the Chinese mitten crab. The relationship is sorted from close to distant, and the closest relationship is fused first; (3.3) After determining the fusion order, perform pairwise fusion of the Chinese mitten crab protein interaction subnetworks based on the reference model organism to obtain the Chinese mitten crab fusion network; each fusion must first determine the relationship between the nodes of the two fused Chinese mitten crab protein interaction subnetworks based on the reference model organism, that is, the homology relationship between the reference model organism proteins corresponding to the two fused Chinese mitten crab protein interaction subnetworks based on the reference model organism; The pairwise fusion of the protein interaction subnetworks of the reference model organism Chinese mitten crab specifically includes: (3.3.1) The relatively close relationship in the two protein interaction subnetworks of the reference model organism Chinese mitten crab that are first fused is set as the target network, and the relatively distant relationship is set as the query network. The target network and the query network are judged by protein sequence alignment whether they have the same protein interaction relationship. Specifically, if the two interacting proteins in the target network and the two interacting proteins in the query network are homologous proteins through protein sequence alignment, then the target network and the query network have the same protein interaction relationship, and the same protein interaction relationship in the target network is extracted as part of the fusion network; the different protein interaction relationships in the target network and the query network are added to the fusion network to finally form a fusion network; (3.3.2) The formed fusion network is used as the new target network, and the next Chinese mitten crab protein interaction subnetwork based on the reference model organism is used as the new query network for the second fusion. If there are N Chinese mitten crab protein interaction subnetworks based on the reference model organism, it needs to be fused N-1 times to finally obtain the Chinese mitten crab fusion network; 4) The fusion network of Chinese mitten crab was expanded to obtain the protein interaction network of Chinese mitten crab.
2. The method for constructing a protein interaction network of Chinese mitten crab according to claim 1, It is characterized in that Step 2) includes: (2.1) Using the BLASTX analysis tool, the gene sequence information of Chinese mitten crab was compared with the protein sequence information of reference model organisms obtained from the Uniprot database, and the E-value was set to 1e-005; (2.2) The first protein sequence of the reference model organism in the comparison is used as the homologous sequence of the Chinese mitten crab gene sequence, thereby determining the reference model organism protein and Chinese mitten crab gene with homologous relationship; (2.3) Based on the protein interaction information of the reference model organisms obtained from the PINA database, the proteins of the reference model organisms that are homologous to the genes of the Chinese mitten crab and the protein interaction relationships formed by the proteins are extracted; (2.4) The proteins and protein interactions of the reference model organisms constitute a Chinese mitten crab protein interaction subnetwork based on the reference model organisms; each reference model organism forms a Chinese mitten crab protein interaction subnetwork based on the reference model organism, and several reference model organisms form several Chinese mitten crab protein interaction subnetworks based on the reference model organisms.
3. The method for constructing a protein interaction network of Chinese mitten crab according to claim 1, It is characterized in that Step 4) includes: (4.1) Using all protein names contained in the Chinese mitten crab genome annotation information as search terms, select a model organism that is the same as the reference model organism in the STRING database one at a time, search and extract the protein interaction relationship of the model organism, until all reference model organisms have been selected and the protein interaction relationship of the corresponding model organism has been searched and extracted; (4.2) The STRING database provides a confidence score for protein interactions. The confidence score ranges from 0 to 1. The closer the score is to 1, the higher the confidence. After searching, only protein interactions with a confidence score of 0.9 or above are retained. (4.3) Remove duplicates from the retained protein interactions with a confidence score above 0.9 to obtain expanded data; (4.4) The expanded data were merged with the aforementioned mitten crab fusion network, and repeated protein interaction relationships were removed to obtain the final mitten crab protein interaction network.
Citation Information
Patent Citations
Methods and compositions involving endopeptidases PepO2 and PepO3
US20050281914A1
Method for identifying interacting proteins
WO2001009615A2