Cancer Transcriptome Data Processing Method Based on Gene Co-Expression Network Analysis
Through the method based on gene co-expression network analysis, the problem that the existing technology is difficult to effectively process and mine cancer transcriptome data is solved, the mining of gene modules and the identification of key genes is achieved, and a new direction is provided to understand the pathogenesis and clinical treatment of cancer.
Patent Information
- Application Number
- CN202210040488.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-14
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-01-14
AI Technical Summary
Existing bioinformatics analysis methods are difficult to effectively process and mine cancer transcriptome data, making it difficult to fully understand the pathogenesis and clinical treatment of cancer.
Methods based on gene co-expression network analysis are adopted, including obtaining original data sets, preprocessing data, identifying differentially expressed genes, building gene co-expression networks, mining gene modules, association analysis, enrichment analysis, identifying key genes, exploring their functions and survival analysis.
This method can effectively unearth the biologically significant gene modules and key genes, identify genes related to tumor diseases, and provide new directions to understand the pathogenesis and clinical treatment of cancer.
Smart Images

Figure CN114360642B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for processing gene data, and more particularly to a method for processing cancer transcriptome data based on gene co-expression network analysis. Background Art
[0002] In recent years, the prevalence of cancer has been increasing. However, due to the difficulty of treating this type of disease and its high recurrence rate, the research on cancer has become increasingly important. If bioinformatics methods can be used to mine functional gene modules of cancer and identify key genes among them, it will surely be able to further understand the pathogenesis of cancer and be helpful for its clinical treatment.
[0003] With the rapid development of next-generation sequencing technology, gene expression data has shown an explosive growth. How to mine hidden knowledge from a large amount of data has become one of the important tasks in the post-genomic era. At the same time, with the in-depth research, people have gradually found that in the cellular environment, various biological factors do not act alone, but cooperate with each other to complete various complex biological functions. Therefore, converting various biological data into biological networks by appropriate methods and then analyzing and mining them using the relevant knowledge of graph theory and complex network theory has become an effective method for processing massive biological data. A biological network is a network constructed using biological elements in scientific problems in the biological field. The nodes in the network represent biological elements such as proteins and genes, while the edges in the network represent the interaction relationships of biological elements in biochemical, physical, or functional aspects. A gene co-expression network is a commonly used biological network, and its emergence has opened up a new direction for the development of genomics. Summary of the Invention
[0004] In order to effectively process cancer transcriptome data, the present invention provides a method for processing cancer transcriptome data based on gene co-expression network analysis.
[0005] The technical solutions adopted by the present invention to solve the technical problems are as follows:
[0006] The method for processing cancer transcriptome data based on gene co-expression network analysis of the present invention mainly includes the following steps:
[0007] Step 1: Obtain the original data set;
[0008] Step 2: Preprocess the original data;
[0009] Step 3: Identify differentially expressed genes;
[0010] Step 4: Construct a gene co-expression network;
[0011] Step 5: Mine gene modules;
[0012] Step 6: Correlation analysis between gene modules and clinical indicators;
[0013] Step 7: Enrichment analysis of gene modules;
[0014] Step 8: Identify key genes;
[0015] Step 9: Explore the functions of key genes;
[0016] Step 10: Survival analysis of key genes.
[0017] Furthermore, in Step 1, the original dataset is from the TCGA database or the GEO database; the original dataset includes gene expression data in cancer tissue samples, gene expression data in adjacent tissue samples, and clinical data corresponding to each sample.
[0018] Furthermore, in Step 2, first filter out low-expression genes, then perform hierarchical clustering on the samples, and delete outlier samples.
[0019] Furthermore, in Step 3, use the FC-t algorithm to identify all differentially expressed genes that meet the defined conditions.
[0020] Furthermore, in Step 4, based on the gene expression data of differentially expressed genes in the samples, perform Pearson correlation analysis between pairwise genes; set defined conditions to screen all the obtained relationships, and regard two genes that meet the defined conditions as having a co-expression relationship; represent all genes with co-expression relationships and their relationships in a graph, that is, obtain the gene co-expression network.
[0021] Furthermore, in Step 5, use 4 community detection algorithms to perform network clustering on the nodes in the gene co-expression network, and obtain communities composed of genes with similar functions, namely gene modules; use "modularity" as the evaluation criterion to select the optimal module mining result.
[0022] Furthermore, in Step 6, perform principal component analysis on all gene expression data in a gene module, and define the first principal component as the module characteristic gene of the gene module; perform Pearson correlation analysis between the module characteristic genes of each gene module and different clinical indicators to obtain the correlation matrix between the gene module and clinical indicators.
[0023] Furthermore, in Step 7, perform enrichment analysis on the genes in the gene module of interest with biological processes, cellular components, and molecular functions provided by the GO database, and at the same time perform enrichment analysis on the gene with signal pathways provided by the Reactome database.
[0024] Further, in step eight, the importance of all nodes in the gene co-expression network is scored using the PageRank algorithm. The scoring criteria are based on topological principles, and then the relatively important nodes in the gene co-expression network are identified. The genes corresponding to these nodes are the key genes.
[0025] Further, in step nine, the Disgenet database is used to retrieve diseases related to the key genes to explore the functions of the key genes.
[0026] Further, in step ten, the online software onclnc is used to perform survival analysis on the key genes and draw survival curves.
[0027] The beneficial effects of the present invention are as follows:
[0028] Complex network theory plays a huge role in many disciplines. In recent years, its applications in disciplines such as computer science, physics, and sociology have been widely studied. An organism is a highly complex system, and every biological process of it requires the joint participation of many substances. It is difficult to understand the underlying molecular mechanism as a whole by studying a single gene or protein. Due to the complexity of cancer diseases, existing bioinformatics analysis methods are difficult to effectively analyze and mine their transcriptome data. Therefore, the present invention applies complex network theory to biological research and specifically applies it to the processing and analysis method of cancer transcriptome data.
[0029] The present invention proposes a method for processing cancer transcriptome data based on gene co-expression network analysis, mainly including: obtaining the original data set; preprocessing the original data; identifying differentially expressed genes; constructing a gene co-expression network; mining gene modules; performing correlation analysis between gene modules and clinical indicators; performing enrichment analysis of gene modules; identifying key genes; exploring the functions of key genes; performing survival analysis of key genes. From the GO / Reactome enrichment analysis results, it can be seen that the gene modules divided by this method have significant biological significance; from the verification results of the Disgenet database for the key genes, it can be seen that most of the key genes identified by this method are related to tumor diseases. Thus, it can be proved that the method for processing cancer transcriptome data based on gene co-expression network analysis provided by the present invention has good effects in the mining of gene modules and the identification of key genes. The method for processing cancer transcriptome data based on gene co-expression network analysis of the present invention can be used as an important tool for cancer disease transcriptome data, and the application of this method also provides a new direction for further understanding the pathogenesis of cancer diseases. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 It is a flowchart of the method for processing cancer transcriptome data based on gene co-expression network analysis of the present invention.
[0031] Figure 2 It is the flowchart of data acquisition and preprocessing in Example 1.
[0032] Figure 3 It is the hierarchical clustering tree of cancer tissue samples in Example 1.
[0033] Figure 4 It is the flowchart of differentially expressed gene identification in Example 1.
[0034] Figure 5 It is the volcano plot of differentially expressed genes in Example 1.
[0035] Figure 6 It is the flowchart of gene co-expression network construction in Example 1.
[0036] Figure 7 It is the gene co-expression network and several small networks in Example 1.
[0037] Figure 8 It is the flowchart of gene module mining in Example 1.
[0038] Figure 9 It is the module mining result of the eigenvector algorithm in Example 1.
[0039] Figure 10 It is the flowchart of the association analysis between gene modules and clinical indicators in Example 1.
[0040] Figure 11 It is the association matrix between gene modules and clinical indicators in Example 1.
[0041] Figure 12 It is the flowchart of GO / Reactome enrichment analysis of gene modules in Example 1.
[0042] Figure 13 It is the BP enrichment result of gene module m1 in Example 1.
[0043] Figure 14 It is the CC enrichment result of gene module m1 in Example 1.
[0044] Figure 15 It is the MF enrichment result of gene module m1 in Example 1.
[0045] Figure 16 It is the flowchart of key gene identification in Example 1.
[0046] Figure 17 It is the survival curve of gene NAA40 in Example 1. Detailed implementation manners
[0047] The present invention will be further described in detail below with reference to the accompanying drawings.
[0048] A method for processing cancer transcriptome data based on gene co-expression network analysis according to the present invention is as Figure 1 shown, and specifically includes the following steps:
[0049] Step 1: Obtain the original data set
[0050] Obtain the original data set from the TCGA database (https: / / cancergenome.nih.gov / ) or the GEO database (https: / / www.ncbi.nlm.nih.gov / geoprofiles). The original data set mainly includes gene expression data in cancer tissue samples, gene expression data in adjacent cancer tissue samples, and clinical data corresponding to each sample.
[0051] Step 2: Preprocess the original data
[0052] First, filter out the low-expression genes therein, that is, delete the low-expression genes with the maximum value of gene expression level (FPKM) less than 1 in cancer tissue or adjacent cancer tissue. Then, perform hierarchical clustering of the expression levels of the remaining genes in cancer tissue or adjacent cancer tissue with respect to the samples, and delete the outlier samples, so as to obtain the original data set for further mining.
[0053] Step 3: Identify differentially expressed genes
[0054] Use the FC-t algorithm to identify all differentially expressed genes that meet the specified conditions. The specified conditions are: FC >= threshold || FC <= threshold && P <= threshold, where FC represents the differential change multiple, threshold represents the differential change multiple, and P represents the statistical significance of the T-test.
[0055] Step 4: Construct a gene co-expression network
[0056] Based on the gene expression data of differentially expressed genes in the samples, perform Pearson correlation analysis between two genes to obtain their Pearson correlation coefficient (PCC) and P value; further, set the specified condition |PCC| >= threshold && P < threshold to screen all the obtained relationships, and regard two genes that meet the specified conditions as having a co-expression relationship; finally, represent all the genes with co-expression relationships and their relationships in a graph, that is, obtain the gene co-expression network.
[0057] Step 5: Mine gene modules
[0058] Four community detection algorithms (eigenvector, label-propagation, map-equation, edge-betweenness) are used to perform network clustering on the nodes (genes) in the gene co-expression network, and communities composed of genes with similar functions, namely gene modules, are obtained. Further, the "modularity" is used as an evaluation criterion to select the optimal module mining result.
[0059] The four community detection algorithms (eigenvector, label-propagation, map-equation, edge-betweenness) are implemented using the functions leading.eigenvector.community(), label.propagation.community(), infomap.community(), and edge.between ness.community() in the "igraph" package of the R software respectively.
[0060] Step 6: Association analysis between gene modules and clinical indicators
[0061] Principal component analysis is performed on the gene expression data of all genes in a gene module, and the first principal component is defined as the module eigen gene (ME) of this gene module. Further, the module eigen genes (ME) of each gene module are subjected to Pearson correlation analysis with different clinical indicators, and the absolute value of the Pearson correlation coefficient (PCC) is taken to obtain the association matrix between this gene module and the clinical indicators.
[0062] Step 7: GO / Reactome enrichment analysis of gene modules
[0063] The genes in the gene module of interest are subjected to enrichment analysis with the biological processes (BP), cellular components (CC), and molecular functions (MF) provided by the GO database (http: / / geneontology.org / ), and at the same time, the genes are subjected to enrichment analysis with the signaling pathways provided by the Reactome database (https: / / reactome.org / ).
[0064] Step 8: Identify key genes
[0065] The PageRank algorithm is used to score the importance of all nodes in the gene co-expression network. The scoring criterion is based on topological principles, and then the relatively important nodes in the gene co-expression network are identified. The genes corresponding to these nodes are the key genes. The top n genes with the highest scores can be selected as the key genes for this disease according to the situation.
[0066] Step 9. Explore the functions of key genes
[0067] Retrieve the diseases related to the key genes using the Disgenet database (http: / / www.disgenet.org / ) to explore the functions of the key genes.
[0068] Step 10. Survival analysis of key genes
[0069] Perform survival analysis on the key genes using the online software onclnc (http: / / www.oncolnc.org / ) and draw the survival curve.
[0070] Next, in combination with the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0071] Example 1 Analysis of transcriptome data of breast invasive carcinoma
[0072] Using the method for processing cancer transcriptome data based on gene co-expression network analysis of the present invention, perform data analysis on the transcriptome of "hepatocellular carcinoma", specifically including the following steps:
[0073] (1) Acquisition and preprocessing of the original data set
[0074] As Figure 2 shown, specifically including the following steps:
[0075] ① Download the expression profile data of breast invasive carcinoma (BRCA) tissues and their adjacent tissues from the TCGA database (https: / / cancergenome.nih.gov / ). There are 1,235 breast invasive carcinoma tissue samples and 149 adjacent tissue samples in total, and each sample contains 60,483 genes.
[0076] ② Delete the low-expression genes with the maximum value of gene expression level (FPKM) less than 1 in breast invasive carcinoma tissues or adjacent tissues, and a total of 14,129 genes remain.
[0077] ③ Perform hierarchical clustering of the expression levels of all the remaining genes after filtration in breast invasive carcinoma tissues or adjacent tissues with respect to the samples. The hierarchical clustering tree is as Figure 3 shown. From Figure 3It can be seen that there are a total of 3 outlier samples: TCGA-DD-AAEB, TCGA-CC-5259, and TCGA-FV-A4ZP. After deleting them, the original dataset for further analysis is obtained. Part of the preprocessed original data (gene expression levels) is shown in Table 1.
[0078] Table 1 Gene expression data of breast invasive carcinoma tissues after preprocessing
[0079] Gene ID Sample 1 Sample 2 Sample 3 Sample 4 Sample 5 ENSG00000167578 2.982962631 2.426924178 2.180554626 1.704487843 1.574196206 ENSG00000078237 1.511416409 2.962567928 3.496769794 2.590545901 0.977175405 ENSG00000146083 15.42361467 34.18583752 7.12327477 6.727115362 4.062698315 ENSG00000198242 105.7124415 207.8535728 193.8028654 113.9189313 112.3564048 ENSG00000134108 17.91677888 19.23785333 34.42522038 26.4881835 14.67476604 ENSG00000167700 25.30139043 15.73939839 12.30061718 5.993611251 58.71591007 ENSG00000060642 5.208718456 6.704560923 6.293990173 6.504866534 11.32913478 ENSG00000166391 0 0.268748468 0.194918684 0.157939948 20.1794968 ENSG00000070087 2.812958236 6.293035032 17.21057879 11.38258759 0.345172284 ENSG00000153561 8.757826693 6.304640789 15.13777725 11.0352232 15.45512232
[0080] (2) Identification of differentially expressed genes
[0081] As Figure 4 shown, it specifically includes the following steps:
[0082] ① Calculate the FC values and P values of all genes using the FC-t algorithm. Part of the calculation results are shown in Table 2.
[0083] ② Set the limiting conditions FC >= 2 || FC <= 0.5 && P <= 0.05 to identify differentially expressed genes. A total of 4130 upregulated genes and 471 downregulated genes are identified.
[0084] ③ Use the R software package "ggplot2" to draw a volcano plot of differentially expressed genes to visually display the screening results of differentially expressed genes. The volcano plot of differentially expressed genes is as Figure 5 shown.
[0085] From Table 2 and Figure 4 it can be seen that compared with normal tissues, there are a large number of genes with significant differential expression in breast invasive carcinoma tissues.
[0086] Table 2 Calculation results of the FC-t algorithm
[0087] Gene ID FC P ENSG00000146083 2.700996521 1.13E-40 ENSG00000198242 2.360743008 3.45E-33 ENSG00000167700 2.481208784 4.38E-28 ENSG00000166391 0.302473358 1.03E-17 ENSG00000127511 2.283066186 1.44E-46 ENSG00000064601 2.93550767 3.52E-58 ENSG00000227766 2.444028443 3.37E-09 ENSG00000008517 2.397033874 2.05E-13 ENSG00000070081 2.184524979 2.20E-27 ENSG00000275479 3.451166828 1.91E-23
[0088] (3) Construction of gene co-expression network
[0089] As Figure 6 shown, it specifically includes the following steps:
[0090] ① For each differentially expressed gene, calculate its Pearson correlation coefficient (PCC) and P value with other differentially expressed genes. Part of the calculation results are shown in Table 3.
[0091] ② Set the limiting condition |PCC| >= 0.65 && P < 0.05 to screen the obtained relationships, and consider two genes that meet the limiting conditions as having a co-expression relationship.
[0092] ③ Import all genes with co-expression relationships and their relationships into the Cytoscape software for visualization, asFigure 7 as shown
[0093] ④According to the visualization results, delete the genes of the small network, and the remaining large network is the gene co-expression network.
[0094] Table 3 Results of Pearson correlation analysis
[0095]
[0096]
[0097] (4) Mining of gene modules
[0098] such as Figure 8 as shown, which specifically includes the following steps:
[0099] ①Use the functions leading.eigenvector.community(), label.propagation.community(), infomap.community(), edge.betweenness.community() of the "igraph" package in R software to perform network clustering on the nodes (genes) in the gene co-expression network, and obtain communities composed of genes with similar functions, that is, gene modules.
[0100] ②Calculate the modularity of the clustering results of 4 community detection algorithms (eigenvector, label-propagation, map-equation, edge-betweenness), and select the result with the largest modularity for further research. In this embodiment, the modularity of the eigenvector algorithm is the highest, so the module mining result obtained by the eigenvector algorithm is selected for further research here.
[0101] ③Delete the communities with too few genes (in this embodiment, delete the communities with less than 50 genes). A total of 9 communities remain, corresponding to 9 gene modules. The module mining result of the eigenvector algorithm is as Figure 9 as shown.
[0102] (5) Correlation analysis between gene modules and clinical indicators
[0103] such as Figure 10 as shown, which specifically includes the following steps:
[0104] ①Perform principal component analysis on all gene expression data in each gene module to obtain the module characteristic genes (ME) of each gene module.
[0105] ②Perform Pearson correlation analysis on the module characteristic genes (ME) of each gene module and 4 clinical indicators, namely event, T, N, and M (where event represents the survival status of the patient, and T, N, and M represent tumor stages), and take the absolute value of the Pearson correlation coefficient (PCC) to obtain the association matrix between the gene module and the clinical indicators, as Figure 11 shown. As can be seen from Figure 11 , gene modules m1, m2, m3, and m7 have relatively high correlations with the clinical indicators.
[0106] (6) GO / Reactome enrichment analysis of gene modules
[0107] As Figure 12 shown, it specifically includes the following steps:
[0108] ①Perform enrichment analysis on the genes included in gene modules m1, m2, m3, and m6 respectively with the biological processes (BP), cellular components (CC), and molecular functions (MF) provided by the GO database, and select the 10 Terms with the smallest P values for research. The enrichment analysis results of gene module m1 are as Figures 13 - 15 shown.
[0109] ②Perform enrichment analysis on the genes included in gene modules m1, m2, m3, and m6 respectively with the signal pathways provided by the Reactome database, and select the 10 signal pathways with the smallest P values for research. The enrichment analysis results of gene module m1 are shown in Table 4.
[0110] From the enrichment results of each gene module, it can be seen that the enrichment results of each gene module have relatively high specificity and are mostly related to tumor diseases, which can prove the reliability of the module mining results.
[0111] Table 4 Reactome enrichment results of gene module m1
[0112] Pathway ID Pathway Name Enrichment Count P R-HSA-69278 Cell Cycle, Mitotic 64 1.11E-16 R-HSA-1640170 Cell Cycle 71 1.11E-16 R-HSA-453279 Mitotic G1 phase and G1 / S transition 24 2.57E-12 R-HSA-73886 Chromosome Maintenance 19 6.22E-12 R-HSA-69205 G1 / S-Specific Transcription 13 3.12E-11 R-HSA-69206 G1 / S Transition 20 3.50E-10 R-HSA-68886 M Phase 32 4.48E-10 R-HSA-69190 DNA strand elongation 11 1.62E-09 R-HSA-73894 DNA Repair 28 2.67E-08 R-HSA-69242 S Phase 19 3.58E-08
[0113] (7) Identification of key genes
[0114] As Figure 16 shown, it specifically includes the following steps:
[0115] ①Use the PageRank algorithm to score the importance of all genes in the gene co-expression network based on the topological structure.
[0116] ②Arrange all genes in descending order according to the scoring results.
[0117] ③ Select the top 20 genes as the key genes for hepatocellular carcinoma. The 20 key genes in this example are: FABP7, CXCL3, LOC284578, CAPN6, NRG2, HCFC1, ILF3, KANSL1, NAA40, NCOA6, PCDHB2, GRIK2, FRMD7, CCSER1, PCDHGA1, PCDHA1, LRRC37A6P, PCDHGA12, ZNF486, and PCDHGB5.
[0118] (8) Exploration of the functions of key genes
[0119] Input all the key genes into the Disgenet database (http: / / www.disgenet.org / ) in sequence for retrieving related diseases. Among them, the retrieval results of gene NAA40 are shown in Table 5.
[0120] From the retrieval results of the Disgenet database, it can be seen that most of the 20 HUB genes are related to tumor diseases, which can prove the reliability of the cancer transcriptome data processing method based on gene co-expression network analysis proposed by the present invention.
[0121] Table 5 Retrieval results of gene NAA40
[0122]
[0123] (9) Survival analysis of key genes
[0124] Perform survival analysis on all the key genes using the online software onclnc (http: / / www.oncolnc.org / ) and draw survival curves (select "BRCA" for cancer). Among them, the survival curve of gene NAA40 is as Figure 17 shown.
[0125] From the survival analysis of the key genes, it can be seen that all 20 key genes are significantly correlated with the survival of patients, which further proves that the cancer transcriptome data processing method based on gene co-expression network analysis proposed by the present invention plays a significant role in the identification of key genes.
[0126] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A method for processing cancer transcriptome data based on gene co-expression network analysis, characterized in that, it includes the following steps: Step 1, obtain the original data set; Step 2, preprocess the original data; Step 3, identify differentially expressed genes; Step 4, construct a gene co-expression network; Step 5, mine gene modules; Step 6, perform correlation analysis between gene modules and clinical indicators; Step 7, perform enrichment analysis of gene modules; Step 8, identify key genes; Step 9, explore the functions of key genes; Step 10, perform survival analysis of key genes; In Step 6, perform principal component analysis on the gene expression data of all genes in a gene module, and define the first principal component as the module characteristic gene of the gene module; Perform Pearson correlation analysis on the module characteristic genes of each gene module and different clinical indicators to obtain the correlation matrix between the gene module and the clinical indicators; In Step 4, based on the gene expression data of differentially expressed genes in the samples, perform Pearson correlation analysis on every two genes to obtain their Pearson correlation coefficient (PCC) and P value; set the limit condition |PCC| >= threshold && P < threshold to screen all the obtained relationships, and regard the two genes that meet the limit condition as having a co-expression relationship; finally, represent all the genes with co-expression relationships and their relationships in a graph, that is, obtain the gene co-expression network; In Step 5, use 4 community detection algorithms to perform network clustering on the nodes in the gene co-expression network to obtain communities composed of genes with similar functions, that is, gene modules; use "modularity" as the evaluation criterion to select the optimal module mining result; the 4 community detection algorithms are eigenvector, label-propagation, map-equation, edge-betweenness; In Step 7, perform enrichment analysis on the genes in the gene module of interest with the biological processes, cellular components and molecular functions provided by the GO database, and at the same time perform enrichment analysis on the gene with the signal pathways provided by the Reactome database; In Step 8, use the PageRank algorithm to score the importance of all nodes in the gene co-expression network. The scoring standard is based on topological principles, and then identify the relatively important nodes in the gene co-expression network. The genes corresponding to these nodes are the key genes.
2. The method for processing cancer transcriptome data based on gene co-expression network analysis according to claim 1, characterized in that, in Step 1, the original data set is from the TCGA database or the GEO database; the original data set includes gene expression data in cancer tissue samples, gene expression data in adjacent cancer tissue samples, and clinical data corresponding to each sample.
3. The method for processing cancer transcriptome data based on gene co-expression network analysis according to claim 2, characterized in that, in Step 2, first filter out low-expression genes, and then perform hierarchical clustering on the samples to delete outlier samples.
4. The method for processing cancer transcriptome data based on gene co-expression network analysis according to claim 3, characterized in that, in step three, all differentially expressed genes satisfying the defined conditions are identified by using the FC-t algorithm.
5. The method for processing cancer transcriptome data based on gene co-expression network analysis according to claim 1, characterized in that, in step nine, diseases related to the key genes are retrieved by using the Disgenet database to explore the functions of the key genes; in step ten, survival analysis of the key genes is performed by using the online software onclnc, and a survival curve is plotted.
Citation Information
Patent Citations
Method for mining radiotherapy specific genes of colorectal cancer by using weight gene co-expression network
CN109872772A
Method and device for constructing omics data analysis platform, and computer device
CN110570905A
Method for establishing pancreatic cancer miRNA prognosis model and screening targeted gene
CN112391470A