Integration method for performing differential protein downstream focusing analysis by combining target gene set and KEGG weight
By combining target gene sets and KEGG weights and combining protein interaction network analysis, proteins that play a key role in specific biological processes are screened, solving the problem that the existing technology is difficult to focus on related key proteins, and achieving more accurate disease mechanism revelation and molecular target provision.
Patent Information
- Application Number
- CN202411683833.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-22
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-11-22
AI Technical Summary
The prior art is difficult to effectively focus on key pathways and key proteins most relevant to specific biological problems or diseases, and fails to integrate information using KEGG pathway enrichment and protein interaction network analysis.
Proteomic data mining is performed by combining the target gene set and KEGG weights, and proteins that play a key role in a specific biological process are screened using KEGG weights, degree-centricity and median-centricity scoring in the protein interaction network.
This method can more effectively reveal the mechanism of action of proteins in the development of diseases, provide more precise molecular targets, and provide support for clinical diagnosis and treatment.
Smart Images

Figure CN119964649A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of bioinformatics analysis and relates to a process method for performing key protein analysis on differential proteins or genes. Background Art
[0002] Proteomics is the science of studying the composition, structure, function and interaction of all proteins in cells. With the development of high-throughput sequencing technology, proteomics research has been able to reveal the differences in protein expression of cells in different physiological and pathological states. The discovery of these differential proteins is of great significance for understanding disease mechanisms, discovering biomarkers and developing new treatments.
[0003] There are many downstream analysis methods for differentially expressed proteins, including functional enrichment analysis, protein-protein interaction network analysis, etc. Each type of analysis contains different data sets and methods, which greatly facilitates bioinformatics analysis for key protein research. Although these methods can provide a large amount of data and information, it is often difficult to focus on key pathways and key proteins that are most relevant to specific biological problems or diseases; in addition, KEGG enrichment of pathways and protein PPI interactions are often used as parallel information mining methods and are not well integrated and utilized.
[0004] In order to solve the above problems, this application proposes a new research process that focuses on differentially expressed proteins related to the target gene set. The core of this process is to exclude most of the interfering information through the target gene set, and then convert the network relationship of genes in KEGG into kegg weights, combined with the degree centrality and betweenness centrality scores of protein interactions, to screen out proteins that play a key role in specific biological processes. This method can more effectively reveal the mechanism of action of proteins in the occurrence and development of diseases, and provide more accurate molecular targets for clinical diagnosis and treatment. Summary of the invention
[0005] The present invention first provides a proteomics data mining-functional annotation analysis method combining a target gene set and KEGG weights, the method comprising the following steps:
[0006] 1) Provide target gene sets of research interest;
[0007] 2) Provide differential protein information, including gene name (symbol) and fold change information;
[0008] 3) Based on 1) the target gene set, differentially expressed proteins were screened and the number of pathways involved in the differentially expressed proteins was counted. Meanwhile, a network diagram related to the target gene set was drawn to visualize the gene set relationship and up- and down-regulation of the differentially expressed proteins;
[0009] 4) Based on the number of differentially expressed proteins in 3), the Kegg count (kc) was calculated and the Kegg weight (kegg_weights) of the protein was assigned by the min-max method. The weight value range was 0 to 1. The Kegg weight calculation formula of the i-th differentially expressed protein was as follows:
[0010] weights(i)=kc(i)-min(kc) / max(kc)-min(kc);
[0011] Where weights(i) represents the Kegg weight of the ith protein, kc(i) represents the number of pathways of the ith protein, max(kc) and min(kc) represent the maximum and minimum number of pathways in the differentially expressed proteins, respectively.
[0012] 5) Draw a protein interaction network diagram for the differentially expressed proteins in 3), and calculate the degree centrality (degree_centrality) and betweenness centrality (betweenness_centrality) of each node in the network; degree centrality refers to the number of edges connected to a node in the network, and betweenness centrality measures the "mediation" role of a node in the network, that is, the frequency of the node as a bridge or intermediary point on the shortest path between other pairs of nodes in the network; preferably, the degree centrality (degree_centrality) and betweenness centrality values (betweenness_centrality) are calculated using the degree and betweenness commands of the R package igraph;
[0013] 6) Calculate the combined importance score of each protein node according to the following formula, where CS (combine_score) represents the final importance score, DC (degree_centrality) represents the degree centrality value, BC (betweenness_centrality) represents the betweenness centrality value, and KW (kegg_weights) represents the kegg pathway weight calculated in step 4);
[0014] CS=(DC+BC)KW;
[0015] 7) Output the first several node proteins based on CS, preferably 2 to 10;
[0016] 8) Perform subsequent functional annotation and mining based on the output node proteins.
[0017] In a specific embodiment of the present invention, the differential protein information is the differential protein information generated after stem cells intervene in healthy subjects or patients.
[0018] In a specific embodiment of the present invention, the differential protein information is obtained through the SomaScan protein sequencing platform.
[0019] In a specific embodiment of the present invention, the seven major classification gene sets of the SomaScan protein sequencing platform are used as target gene sets, and the seven major classification gene sets include proteins related to cardiovascular disease, inflammation and immune response, metabolic disease, cancer, nerves, cytokines, and respiration.
[0020] In a specific embodiment of the present invention, differential proteins are screened according to gene sets, protein interaction network analysis and KEGG weight assignment are performed, key candidate proteins are screened, and at the same time, the KEGG enrichment information of all differential proteins is combined to perform preliminary functional annotations on the candidate proteins to determine the key proteins.
[0021] In a specific embodiment of the present invention, pathway enrichment analysis of differential proteins is performed based on the KEGG database.
[0022] In a specific embodiment of the present invention, significant pathways are determined based on the p.adjust (corrected p value) of KEGG enrichment; preferably, the top 20 pathways with the smallest p.adjust values are focused on.
[0023] In a specific embodiment of the present invention, the mapping display of key proteins on the target pathway is performed by Kegg Mapper to preliminarily determine the biological significance of the target protein.
[0024] The present invention also provides a computer component storing a computer program for implementing the above method.
[0025] Beneficial technical effects
[0026] This method can reduce the noise of irrelevant proteins, more efficiently screen key target proteins and perform preliminary functional annotation. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 The top 10 protein changes after adding KEGG weights
[0028] Figure 2 Conventional downstream analysis of differentially expressed proteins
[0029] Figure 3 Network relationship diagram between seven classification function sets and differentially expressed proteins
[0030] Figure 4 Protein interaction network diagram between seven major classification-related differentially expressed proteins
[0031] Figure 5Interaction network diagram of all differentially expressed proteins
[0032] Figure 6 Hierarchical diagram of the top 20 pathways enriched by all differentially expressed proteins
[0033] Figure 7 PRKACA is mapped in the pathway. The red box is the differentially expressed protein in the pathway, and the pink box is PRKACA.
[0034] Figure 8 CASP3 is mapped in the pathway. The red box is the differentially expressed protein in the pathway, and the pink box is CASP3.
[0035] Fig. 9 Focus analysis flow chart DETAILED DESCRIPTION
[0036] This analysis process extracts and organizes differential protein information based on the functional gene set, and constructs a gene set-differential protein network diagram to show the network relationship between the gene set and the differential protein; and constructs a protein interaction network for the extracted differential proteins, and screens candidate key proteins based on the KEGG weights of the differential proteins. At the same time, pathway enrichment analysis is performed on all differential proteins, and the top 20 enriched pathways are visualized, and the pathways where the key proteins are located are observed and mapped, so as to quickly annotate the functions of potential differential proteins.
[0037] 1) Organize the target proteins. Generally, the target proteins are organized according to their own research focus. This experiment uses the seven major classification gene sets provided by the SomaScan protein sequencing platform as the research object;
[0038] 2) Seven major classification gene sets include proteins related to cardiovascular disease, inflammation and immune response, metabolic disease, cancer, nerves, cytokines, and respiration;
[0039] 3) Perform pathway enrichment analysis on the differentially expressed proteins and visualize the top 20 enrichment results, including the hierarchical information of the pathways;
[0040] 4) Mark the differentially expressed proteins according to the gene set, and construct a gene set-differentially expressed protein network diagram to preliminarily observe the relationship between the differentially expressed proteins and the gene set;
[0041] 5) Perform protein interaction mapping on the differentially expressed proteins in the gene set, calculate the KEGG weights of the differentially expressed proteins, and screen for key proteins based on the CS calculation formula results;
[0042] 6) Map the screened key proteins to the target pathway and complete the preliminary screening-functional annotation analysis.
[0043] The analysis process of the present invention is detailed in Figure 8 , the whole process can be implemented under Linux through programming.
[0044] Technical terms
[0045] SomaScan Protein Sequencing Platform
[0046] The SomaScan protein sequencing platform discovers new disease biomarkers through high-throughput protein analysis technology. Currently, it can detect more than 11,000 proteins in ultra-trace samples at one time. The detection technology covers sample types including blood, urine, body fluids, etc. The sample sources include humans and other model animals. Only 55 microliters of sample are needed to detect proteins, and the dynamic detection range is 10log, which can detect extremely low-abundance proteins.
[0047] The seven major categories described in this article come from seven panels designed by the SomaScan platform based on 1,178 literature collections targeting proteins in different directions. Specific literature can be viewed at https: / / somalogic.com / publications / .
[0048] Protein interaction network diagram
[0049] Protein-protein interaction (PPI) network is an analytical method that describes the interaction relationship between a group of proteins. Because the interaction relationship between proteins is complex, each protein is a node, and the interaction between multiple proteins is connected into lines to form a network structure, so it is called a protein interaction network.
[0050] In this paper, the protein interaction graph was implemented using stringDB, and the igraph package was used to calculate the degree centrality and betweenness centrality values of the interaction network.
[0051] Gene set-differential protein network diagram
[0052] In the network diagram of gene sets and differential proteins, the gene set classification labels and genes are used as nodes, the categories to which the genes belong are used as connecting edges, and different categories are connected through the common genes they contain. Based on this, a relationship network is drawn to display the categories with the most differential proteins and their up- and down-regulation conditions.
[0053] The determination of key differentially expressed proteins includes three steps: 1) screening based on the specified gene set; 2) PPI network construction and importance score calculation; 3) final scoring based on KEGG weights. The final most consistent score CS is the final basis for judgment. The proteins are sorted according to the CS score, and the top ranking proteins are selected for subsequent research.
[0054] Pathway enrichment analysis
[0055] Pathway enrichment analysis: Focuses on checking whether genes / metabolites in a gene set are enriched in a specific pathway. Databases such as KEGG and Reactome provide information on known biological pathways, which can be used for enrichment analysis. Specific calculation methods include:
[0056] 1) Provide differentially expressed genes n;
[0057] 2) Calculate the number k of differentially expressed genes from pathway A;
[0058] 3) Read the number of genes M in the KEGG background file;
[0059] 4) Read the number of genes N in pathway A;
[0060] 5) Calculate the enrichment fold (Fold Enrichment) or enrichment score, the calculation formula is: Enrichment Score = (k / n) / (N / M);
[0061] 6) Calculate P values to assess the significance of enrichment, such as the hypergeometric test and Fisher's exact test.
[0062] In this paper, KEGG enrichment analysis was performed using the R package clusterProfiler.
[0063] In this article, significant pathways are usually selected based on the corrected p-value (p.adjust) of the enrichment results. Generally, p.adjust < 0.05 can be considered as significant enrichment, and the enrichment results of different studies vary.
[0064] Pathway Mapping
[0065] In this article, pathway mapping refers to color-coding the differentially expressed proteins on the relevant pathways. The main purpose is to view the functional distribution of the differentially expressed proteins in the pathways and the functional positions of the candidate key proteins in the pathways of interest, as well as their upstream and downstream regulatory relationships.
[0066] Example
[0067] For a study on the effects of stem cells on healthy people, the differential proteins obtained were screened and analyzed using the method of the present invention, as follows:
[0068] Example 1 Screening protein interaction network diagram drawing and KEGG weight calculation
[0069] This experiment used the seven key categories provided by the SomaScan protein sequencing platform as the research objects. The seven classification gene sets include proteins related to cardiovascular disease, inflammation and immune response, metabolic disease, cancer, nerves, cytokines, and respiration. Different from the conventional downstream analysis of differentially expressed proteins (see Figure 2 ), first draw a network relationship diagram between the seven classification function sets and differentially expressed proteins (see Figure 3 ), where the gene set and differential protein network diagram showed that the most involved functions among the differential proteins were inflammation and immune response, and nerve-related proteins. These two types of functions are the areas where stem cells are more concentrated. The cardiovascular disease-related gene set is also relatively significant, which can be combined with subsequent observations.
[0070] Based on the network relationship diagram between the seven major classification function sets and differentially expressed proteins, a protein interaction network diagram between the seven major classification-related differentially expressed proteins was further drawn (see Figure 4 ), and the KEGG weights of the differentially expressed proteins were calculated at the same time. The differentially expressed proteins of interest were obtained by combining the degree centrality values, betweenness centrality values, and KEGG weights calculated based on the interaction network. By adding the KEGG weights, 'PRKACA' and 'CASP3' became the two most important candidate proteins (see Figure 1 ).
[0071] Example 2 Differential protein pathway enrichment analysis
[0072] Next, pathway enrichment analysis was performed on all differentially expressed proteins, and the top 20 pathways were visualized. The results are shown in Figure 4 ,Depend on Figure 4 The functional level information involved in the pathways that can be mainly enriched and the levels with the largest number are more prominent in terms of immunity-related, signal transduction-related, cancer-related, and cardiovascular disease-related levels.
[0073] Among the first 20 pathways, we will focus on mapping proteins PRKACA and CASP3 in the pathways of interest. Figure 5-7 .
[0074] Example 3 Results Analysis
[0075] For PRKACA and CASP3, which ranked high after adding KEGG, functional annotation was performed in combination with the pathways of interest. PRKACA (protein kinase A catalytic subunit α) and CASP3 (caspase 3) are two important proteins that play a key role in cell signaling, immune response and apoptosis, and play an important role in health maintenance, especially in pathways related to inflammation and cardiovascular diseases. PRKACA is the catalytic subunit of protein kinase A (PKA) and is activated when cyclic adenosine monophosphate (cAMP) levels increase, leading to the phosphorylation of multiple target proteins. PRKACA plays a role in multiple biological processes such as metabolism, gene expression, cell division and apoptosis. In the chemokine signaling pathway, PRKACA participates in downstream signaling and affects the movement and response of immune cells. The activation of PKA in immune cells can change their response to chemokines, which plays a key role in guiding immune cells to migrate to sites of infection or injury and promoting the regulation of immune responses. CASP3 is a key execution protease in the apoptotic pathway, which initiates and completes programmed cell death by cleaving various substrates. It plays an important role in clearing damaged or unwanted cells and maintaining cellular homeostasis. CASP3-mediated apoptosis of endothelial cells and vascular smooth muscle cells contributes to the progression of atherosclerosis. Apoptosis in atherosclerotic plaques can lead to plaque instability, which is crucial in the pathogenesis of cardiovascular diseases.
[0076] This analysis explains to some extent the biological function effects caused by stem cell therapy. Stem cell therapy, especially mesenchymal stem cell (MSC) therapy, regulates the functions of PRKACA and CASP3 by secreting multiple factors to affect chemokine and lipid pathways: stem cells can secrete factors that enhance the expression or activity of PKA-related signaling molecules. This regulatory effect can reduce inflammation, improve immune function, and reduce the risk of plaque formation by promoting vascular health and reducing lipid accumulation. In addition, stem cells reduce CASP3 activation in key tissues by secreting growth factors, cytokines, and anti-inflammatory molecules. This helps protect important tissues from unnecessary cell death, thereby inhibiting the aggravation of conditions such as atherosclerosis.
[0077] This indicates that stem cells may contribute to organ rejuvenation and maintenance of tissue function by targeting and regulating inflammatory responses, enhancing tissue repair, and reducing apoptosis. The comprehensive impact of stem cell therapy on signaling pathways and protein activity demonstrates its potential in delaying aging-related diseases and maintaining multi-organ health. Through the key mediation of PRKACA and CASP3, stem cell therapy can have a positive regulatory effect on the body's systemic health by reducing inflammation and supporting tissue repair. This pathway regulation shows therapeutic promise for chronic diseases and helps maintain the body's homeostasis and healthy tissue function.
[0078] Only based on the differential proteins FN1 and MMP9 screened by the protein interaction network diagram, preliminary investigation shows that matrix metalloproteinase-9 (MMP-9) is an enzyme that plays multiple roles in the human body. It is mainly involved in the degradation and remodeling of the extracellular matrix (ECM), affecting processes such as tissue repair, angiogenesis, and embryonic development. MMP-9 is also involved in the immune regulation of leukocytes, angiogenesis, neurogenesis and synaptic remodeling, as well as tumor growth and metastasis. Fibronectin 1 (FN1) is a key extracellular matrix protein that plays an important role in regulating cell adhesion, migration, proliferation, and participating in wound healing, tumor progression, tissue fibrosis, and immune regulation. FN1 affects cell signal transduction by binding to integrin receptors on the cell surface, especially in signaling pathways such as TGF-β / PI3K / Akt. Its abnormal expression is closely related to the occurrence and development of many diseases, making it an important target for drug development and disease diagnosis. Compared with PRKACA and CASP3, the impact on stem cells is not closely related.
Claims
1. A method for proteomics data mining-functional annotation analysis combining a target gene set and KEGG weights, the method comprising the following steps: 1) Provide target gene sets of research interest; 2) Provide differential protein information, including gene name (symbol) and fold change information; 3) Based on the target gene set in 1), differentially expressed proteins were screened and the number of pathways involved in the differentially expressed proteins was counted. Meanwhile, a network diagram related to the target gene set was drawn to visualize the gene set relationship and up- and down-regulation of the differentially expressed proteins; 4) Based on the number of differentially expressed proteins in 3), the Kegg count (kc) was calculated and the Kegg weight (kegg_weights) of the protein was assigned by the min-max method. The weight value range was 0 to 1. The Kegg weight calculation formula of the i-th differentially expressed protein was as follows: weights(i)=kc(i)-min(kc) / max(kc)-min(kc); 5) Draw a protein interaction network diagram for the differentially expressed proteins in 3), and calculate the degree centrality and betweenness centrality of each node in the network, respectively. Degree centrality refers to the number of edges connected to a node in the network, and betweenness centrality refers to the frequency of the node as a bridge or intermediary point on the shortest path between other pairs of nodes in the network; preferably, the degree and betweenness commands of the R package igraph are used to calculate the degree centrality and betweenness centrality; 6) The combined importance score of each protein node was calculated according to the following formula, where CS (combine_score) represents the final importance score, DC (degree_centrality) represents the degree centrality value, BC (betweenness_centrality) represents the betweenness centrality value, and KW (kegg_weights) represents the calculated KEGG pathway weight; CS=(DC+BC)KW; 7) Output several key proteins based on CS, preferably 2 to 10; 8) Carry out subsequent functional annotation and mining based on the output key proteins.
2. The method according to claim 1, wherein the differential protein information is the differential protein information generated after stem cells intervene in healthy subjects or patients.
3. The method according to claim 1 or 2, wherein the differential protein information is obtained by SomaScan protein sequencing platform.
4. The method of claim 3, wherein the target gene set is determined based on the seven major classification gene sets of the SomaScan protein sequencing platform, wherein the seven major classification gene sets include proteins related to cardiovascular disease, inflammation and immune response, metabolic disease, cancer, nerves, cytokines, and respiration.
5. According to the method of claim 4, differential proteins are screened according to the seven classification gene sets, protein interaction network analysis and KEGG weight assignment are performed, key proteins are screened, and the KEGG enrichment information of all differential proteins is combined to perform preliminary functional annotations on the key proteins.
6. The method according to claim 5, wherein pathway enrichment analysis of differentially expressed proteins is performed based on the KEGG database.
7. The method of claim 6, wherein the significant pathway is determined based on the p.adjust (corrected p value) of KEGG enrichment.
8. The method as claimed in claim 7, wherein the top 20 pathways with the smallest p.adjust values are taken as significant pathways.
9. The method according to claim 8, wherein the key protein is mapped and displayed on the target pathway by Kegg Mapper to preliminarily determine the biological significance of the key protein.
10. A computer component storing a computer program for implementing the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Protein biomarker identification method independent of database search
CN110838340A
Method for screening key genes and key protein sets related to research purpose
CN112634982A
Method for screening potential biomarker of prostate cancer and application thereof
CN114252611A
Methods and systems for annotating genomic data
US20240112752A1
Protein design method and system
WO2017081687A1