Tumor metabolism subtype classification method based on KEGG metabolic pathway GSVA score and application
Through GSVA scoring method and cluster analysis of the KEGG metabolic pathway, the metabolic subtypes of tumor cells were identified and divided, solving the problem that the role of metabolic pathways in traditional methods was not fully considered, and achieving more accurate classification of tumor metabolic subtypes and support for personalized treatment.
Patent Information
- Application Number
- CN202510152537.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-30
AI Technical Summary
Traditional tumor classification methods fail to fully consider the role of metabolic pathways in the tumor process, resulting in insufficient identification of tumor metabolic subtypes, affecting the effects of early diagnosis, prognosis evaluation and personalized treatment.
The GSVA scoring method based on the KEGG metabolic pathway was used to comprehensively analyze the metabolic pathways of tumor cells. Different metabolic subtypes were identified through K-mean clustering and consensus clustering, and differential gene expression analysis and functional enrichment analysis were performed.
The accurate division of tumor metabolic subtypes has been achieved, providing scientific basis for early diagnosis, treatment plan selection and efficacy evaluation, and improving the feasibility of personalized treatment.
Smart Images

Figure CN120072069A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biomedicine, and more specifically, to a method for classifying tumor metabolic subtypes based on the GSVA score of KEGG metabolic pathways and its application. Background Art
[0002] During the growth and proliferation of tumor cells, significant changes occur in their metabolic activities, manifested as reprogramming of metabolic pathways. A classic example is the "Warburg effect", that is, under aerobic conditions, tumor cells tend to obtain energy through the glycolysis pathway rather than the oxidative phosphorylation pathway. This phenomenon has been widely observed in most tumors. However, in addition to glycolysis, tumor cells often also re - adjust other metabolic pathways, such as lipid metabolism, amino acid metabolism, and nucleotide metabolism, etc., to support their rapid growth and proliferation.
[0003] The metabolic characteristics of tumor cells are heterogeneous among different types of tumors and are closely related to the occurrence, development, metastasis, and drug resistance of tumors. Traditional tumor classification methods are mainly based on morphology and gene expression, but they do not fully consider the role of metabolic pathways in the tumor process. Although many studies have explored the characteristics of tumor cell metabolism, these studies often focus on single metabolic pathways and lack a comprehensive analysis of the overall metabolism. Therefore, the identification of tumor metabolic subtypes is of great significance for the early diagnosis, prognosis evaluation, and formulation of precise treatment plans for cancer. Different metabolic subtypes may reflect the heterogeneity of tumor cells in metabolic activities, and cells of different subtypes may respond differently to treatment methods such as chemotherapy, targeted therapy, and immunotherapy. The classification of metabolic subtypes based on metabolic pathways can provide new ideas for the personalized treatment of tumors. Summary of the Invention
[0004] The present invention provides a method for classifying tumor metabolic subtypes based on the GSVA score of 85 KEGG metabolic pathways, which can comprehensively analyze the metabolic pathways of tumor cells, accurately classify the metabolic subtypes of tumors, and provide a basis for the early diagnosis, selection of treatment plans, and evaluation of treatment efficacy of tumors.
[0005] The technical solution of the present invention is as follows:
[0006] A method for classifying tumor metabolic subtypes based on the GSVA score of KEGG metabolic pathways, comprising the following steps:
[0007] (1) Obtain scRNA - seq or bulkRNA - seq data;
[0008] (2) The scRNA-seq data in step (1) is subjected to dimensionality reduction and clustering to obtain different cell types; the obtained epithelial cells are re-clustered by dimensionality reduction to obtain different epithelial sub-clusters;
[0009] (3) The different epithelial sub-clusters in step (2) are used as references with T cells, myeloid cells, and endothelial cells for InferCNV copy number variation analysis to preliminarily identify malignant cell sub-clusters. Subsequently, normal and tumor samples of lung cancer are downloaded from the TCGA database, and the top 50 differentially expressed genes of the normal and tumor samples are respectively selected as the gene sets of normal epithelial cells and tumor cells. Then, single-sample gene set enrichment analysis is performed to evaluate the expression of these two gene sets in the epithelial sub-population, and tumor cell sub-clusters are obtained;
[0010] (4) For the bulkRNA-seq data in step (1), the tumor samples need to be separated;
[0011] (5) Obtain the gene expression matrix of the tumor cells in step (3) or the gene expression matrix of the tumor samples in step (4);
[0012] (6) Download the gene set information of the metabolic pathways from the KEGG database; according to the information, calculate the GSVA score through the gene expression matrix in claim 5;
[0013] (7) According to the GSVA score in step (6), apply K-means clustering to identify the optimal number of metabolic subtypes;
[0014] (8) Use the ConsensusClusterPlus method to perform consensus clustering through 1,000 iterations and 80% resampling to determine this classification;
[0015] (9) Plot a heatmap to display the clustering results in step (8);
[0016] (10) Perform differential gene expression analysis on the clustered metabolic subtypes to obtain the differentially expressed genes (DEGs) of each metabolic subtype, and evaluate the differences in gene expression of each metabolic subtype;
[0017] (11) Use the DEGs obtained in step (10) for GO enrichment analysis to evaluate the functional characteristics of different metabolic subtypes.
[0018] Preferably, the gene set information in step (6) includes at least one of lipid metabolic pathways, energy metabolic pathways, amino acid metabolic pathways, and nucleotide metabolic pathways.
[0019] Preferably, the K-means clustering in step (7) includes using the K-means function in R language.
[0020] The second object of the present invention is to protect the application of the above method in the evaluation of the efficacy of tumors in non-disease treatment and diagnosis.
[0021] Advantages of the present invention
[0022] Compared with the existing metabolic subtype classification methods, the present invention has the following unique advantages and innovations: Traditional metabolic subtype classification methods usually rely on single or a few metabolic pathways (such as glycolysis, oxidative phosphorylation, etc.), while the present invention is based on the comprehensive scoring of 85 metabolic pathways in the KEGG database, covering multiple key metabolic modules such as energy metabolism, amino acid metabolism, lipid metabolism, nucleotide metabolism, carbohydrate metabolism, vitamin and cofactor metabolism, etc., and can more accurately and systematically characterize the metabolic characteristics of malignant cells; This method can perform unbiased analysis using large-scale tumor cell data, adopt a scoring system based on KEGG metabolic pathways, and combine clustering analysis (such as K-means, consensus clustering, etc.) to objectively reveal metabolic subtypes, rather than relying on empirical classification; This method can also combine single-cell transcriptome data (scRNA-seq) to analyze tumor metabolic heterogeneity at the single-cell level, thereby distinguishing the metabolic states in different microenvironments and capturing the metabolic plasticity of tumor cells during evolution and treatment. The metabolic subtype classification of the present invention can be closely associated with the clinical characteristics of patients (such as stage, treatment response, survival prognosis), and can especially be used to predict the sensitivity of metabolic targeted drugs, assist in formulating personalized treatment plans, and improve the feasibility of precision medicine. Description of the drawings
[0023] Figure 1 is the flowchart of the method of the present invention;
[0024] Figure 2 is the K-value selection diagram of Example 1 of the present invention;
[0025] Figure 3 is the consensus clustering diagram of Example 1 of the present invention;
[0026] Figure 4 is the subtype classification heat map of Example 1 of the present invention;
[0027] Figure 5 is the K-value selection diagram of Example 2 of the present invention;
[0028] Figure 6 is the consensus clustering diagram of Example 2 of the present invention;
[0029] Figure 7 is the subtype classification heat map of Example 2 of the present invention. Detailed implementation manners
[0030] To better understand the technical solution of the present invention, the technical solution of the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be noted that the embodiments described herein are based on the technical solution of the present invention, and detailed implementation manners and specific operation processes are given. However, the protection scope of the present invention is not limited to the described embodiments. On the contrary, the purpose of providing these embodiments is to make the content of the present invention more thorough and comprehensive.
[0031] Example 1
[0032] A method for classifying tumor cell metabolic subtypes based on the GSVA score of KEGG metabolic pathways. Example 1 takes the TCGA lung cancer bulk RNA-seq data as an example to illustrate the present invention.
[0033] (1) Download and obtain the TCGA lung cancer bulk RNA-seq data from the TCGA official website; log in to the TCGA database (https: / / www.cancer.gov / ccg / research / genome-sequencing / tcga) and select 600 samples of lung cancer, including 541 tumor samples and 59 adjacent cancer samples.
[0034] (2) Process the downloaded data to separate the tumor samples to obtain the tumor sample gene expression matrix;
[0035] Download 85 metabolic pathway gene sets from the KEGG official website; log in to the KEGG official website (https: / / www.genome.jp / kegg / ) and download gene sets of metabolic pathways such as lipid metabolism and energy metabolism. For example, the gene sets of nitrogen metabolism include GLUD2, GLUD1, GLUL, CPS1, CA13, CA1, CA6, CA7, CA12, CA5B, CA14, CA9, CA3, CA5A, CA8, CA2, CA4;
[0036] (3) Based on the expression matrix of the tumor samples in step (2), combined with 85 KEGG metabolic pathways, the gene set variation analysis (GSVA) method is used to score each metabolic pathway; in R, execute gsva.res <- gsva(Tumor, KEGG, method = 'gsva', kcdf = "Gaussian", min.sz = 10). "gsva" is a function in the GSVA package used to calculate the gene set variation analysis score; "Tumor" is the gene expression matrix of the tumor samples mentioned in step (2); "KEGG" is the gene set of 85 KEGG metabolic pathways mentioned in step three; "method = 'gsva'" specifies using the GSVA method for score calculation; "kcdf = \"Gaussian\"" uses the Gaussian kernel to estimate the distribution type of gene expression data; "min.sz = 10" sets the minimum gene set size to 10, which means that when performing GSVA analysis, each gene set needs to contain at least 10 genes, and the calculation results are stored in gsva.res;
[0037] (4) Based on the GSVA scores calculated in step (3), apply K - means clustering (using the K - means function in R language) to identify the optimal number of metabolic subtypes; in R, execute fviz_nbclust(data, kmeans, method = "wss"); "fviz_nbclust()" is a function in the factoextra package used to visualize the results of cluster analysis and select the optimal number of clusters k; "data" is the score calculated in step 4; "kmeans" is the selected clustering algorithm, specifying the use of k - means clustering, aiming to divide the data into k clusters; "method = \"wss\"" specifies using the "WSS" (Within - cluster sum of squares) method to evaluate the effect of clustering. WSS represents the sum of squared errors within the cluster, which indicates the distance between data points and their cluster centers. By plotting the change trend of WSS as the number of clusters k increases, it can help select a reasonable value of k. Usually, at the "elbow", where the rate of decrease of WSS begins to slow down, it is used as the basis for selecting the optimal k value; here we select the optimal K value = 3( Figure 2 );
[0038] (5) Perform consensus clustering analysis using the ConsensusClusterPlus package. Determine this classification through consensus clustering with 1,000 iterations and 80% resampling; execute results = ConsensusClusterPlus(d, maxK = 20, reps = 1000, pItem = 0.8, pFeature = 1, title = "66tiao_untitled_consensus_cluster", clusterAlg = "hc", distance = "pearson", seed = 1262118388.71279, tmyPal = NULL, writeTable = FALSE, plot = "png") in R. The "ConsensusClusterPlus()" function refers to performing consensus clustering analysis; "maxK = 20" means the maximum number of clusters is 20; "reps = 1000" means repeating 1,000 times to increase the stability of clustering; "pItem = 0.8" means randomly selecting 80% of the samples each time of clustering; "pFeature = 1" means using 100% of the metabolic pathways each time of clustering; "clusterAlg = 'hc'" means using Hierarchical Clustering; "distance = \"pearson\"" means using the Pearson correlation coefficient as the distance metric; "seed = 1262118388.71279" means setting the random number seed; we plotted a heatmap for the analysis clustering results, Figure 3 showing the clustering results of the consensus clustering.
[0039] (6) Plot a heatmap to show this clustering result, Figure 4 showing the clustering heatmap.
[0040] Through the above steps, we identified three metabolic subclusters. Specifically, cluster2 showed significant metabolic activity and was thus defined as the high - metabolic subcluster; cluster1 had moderate metabolic activity and was defined as the medium - metabolic subcluster; while cluster3 showed low metabolic activity and was classified as the low - metabolic subcluster.
[0041] Next, differential analysis can be carried out to identify significant differences in gene expression, metabolic pathways, and related biological characteristics among different metabolic subclusters, such as differential gene expression analysis, functional enrichment analysis, immune microenvironment analysis, etc. Through these differential analyses, we will more deeply understand the impact of tumor cell metabolic heterogeneity on tumor biological behavior, and further provide more valuable biomarkers or targets for personalized precision treatment.
[0042] Example 2
[0043] The present invention relates to a method for classifying tumor cell metabolic subtypes based on the GSVA score of the KEGG metabolic pathway. In Example 2, the scRNA-seq data of lung cancer in the GEO database is taken as an example to illustrate the present invention.
[0044] (1) Download the scRNA-seq data of lung cancer from the GEO database; log in to (https: / / www.ncbi.nlm.nih.gov / geo / ) and select 44 samples from 33 patients numbered GSE131907, including normal samples, adjacent cancer samples, brain metastasis samples, and lymph node metastasis samples.
[0045] (2) Process the downloaded data; the UMI (unique molecular identifier) count matrix is processed using the Seurat package in R. For single-cell data QC (quality control), each cell has at least 200 genes and at most 10,000 genes, at least 500 UMIs and at most 50,000 UMIs. The DoubletFinder package is used to identify and exclude doublets, and the merge function is used to merge 44 samples. Subsequently, the Harmony package is used to reduce batch effects, and finally, dimensionality reduction and clustering result in 150,725 single cells, of which 35,854 are epithelial cells.
[0046] (3) Identify malignant cells after re-dimensionality reduction and clustering of the 35,854 epithelial cells obtained in step (2); after re-dimensionality reduction and clustering, 34 epithelial subclusters are obtained. Subsequently, the InferCNV (copy number variation) tool is used to preliminarily identify clusters 3, 4, 6, 18, 19, and 23 as normal epithelial cells, and other subclusters as malignant cells. Then, we download normal and tumor samples of lung cancer from the TCGA database, and select the top 50 differential genes of normal and tumor samples respectively as the gene sets of normal epithelial cells and tumor cells. Then we perform single-sample gene set enrichment analysis (ssGSEA) to evaluate the expression of these two gene sets in 34 epithelial subpopulations. This analysis verifies the results of InferCNV and determines tumor cells.
[0047] (4) Isolate the tumor cells in step (3) to obtain the gene expression matrix of tumor cells.
[0048] Download 85 metabolic pathway gene sets from the KEGG official website; log in to the KEGG official website (https: / / www.genome.jp / kegg / ) and download gene sets of metabolic pathways such as lipid metabolism and energy metabolism.
[0049] (5) Based on the gene expression matrix of tumor cells in step (4), combined with 85 KEGG metabolic pathways, the gene set variation analysis (GSVA) method is used to score each metabolic pathway; execute gsva.res<-gsva(Tumor, KEGG, method='gsva', kcdf="Gaussian", min.sz=10) in R. "gsva" is a function in the GSVA package for calculating the gene set variation analysis score; "Tumor" is the gene expression matrix of tumor samples mentioned in step 2; "KEGG" is the gene set of 85 KEGG metabolic pathways mentioned in step 3; "method='gsva'" specifies using the GSVA method for score calculation; "kcdf="Gaussian"" uses the Gaussian kernel to estimate the distribution type of gene expression data; "min.sz=10" sets the minimum gene set size to 10, which means that when performing GSVA analysis, each gene set needs to contain at least 10 genes, and the calculation results are stored in gsva.res.
[0050] (6) Based on the GSVA score gsva.res calculated in step (5), apply K-means clustering (using the K-means function in R language) to identify the optimal number of metabolic subtypes; execute fviz_nbclust(gsva.res, kmeans, method="wss") in R; "fviz_nbclust()" is a function in the factoextra package for visualizing the results of cluster analysis and selecting the optimal number of clusters k; "gsva.res" is the score calculated in step 6; "kmeans" is the selected clustering algorithm, specifying the use of k-means clustering to divide the data into k clusters; "method="wss"" specifies using the "WSS" (Within-cluster sum of squares) method to evaluate the effect of clustering. WSS represents the sum of squared errors within the cluster, which indicates the distance between data points and their cluster centers. By plotting the change trend of WSS as the number of clusters k increases, it can help select a reasonable value of k. Usually, at the "elbow", that is, when the decline rate of WSS starts to slow down, it is used as the basis for selecting the optimal k value; here we select the optimal K value = 3( Figure 5 )
[0051] (7) Consensus clustering analysis was performed using the ConsensusClusterPlus package, and this classification was determined by consensus clustering through 1,000 iterations and 80% resampling; in R, execute results = ConsensusClusterPlus(d, maxK = 20, reps = 1000, pItem = 0.8, pFeature = 1, title = "66tiao_untitled_consensus_cluster", clusterAlg = "hc", distance = "pearson", seed = 1262118388.71279, tmyPal = NULL, writeTable = FALSE, plot = "png"). The "ConsensusClusterPlus()" function refers to performing consensus clustering analysis; "maxK = 20" means the maximum number of clusters is 20; "reps = 1000" means repeating 1000 times to increase the stability of clustering; "pItem = 0.8" means randomly selecting 80% of the samples each time of clustering; "pFeature = 1" means using 100% of the metabolic pathways each time of clustering; "clusterAlg = 'hc'" means using hierarchical clustering; "distance = \"pearson\"" means using the Pearson correlation coefficient as the distance metric; "seed = 1262118388.71279" means setting the random number seed; we plotted a heatmap of the analysis clustering results, Figure 6 showing the clustering results of the consensus clustering.
[0052] (8) Plot a heatmap of this clustering result to show, Figure 7 showing the clustering heatmap.
[0053] Through the above steps, we identified three metabolic subtypes in malignant cells, namely cluster1, which showed significant metabolic activity, cluster2 with moderate metabolic activity, and cluster3 with low metabolic activity.
[0054] Based on these three metabolic subtypes, we performed differential analysis; to identify significant differences in gene expression, metabolic pathways, and related biological characteristics among different metabolic subclusters, which will help us understand more deeply the impact of tumor cell metabolic heterogeneity on tumor biological behavior, and thus provide more valuable biomarkers or targets for personalized precision medicine.
[0055] The present invention provides a method for classifying tumor metabolic subtypes based on the GSVA score of KEGG metabolic pathways. By comprehensively analyzing the metabolic pathways of tumor cells, different metabolic subtypes can be accurately classified. This method comprehensively considers multiple key metabolic pathways, such as lipid metabolism, energy metabolism, amino acid metabolism, etc., rather than being limited to a single metabolic pathway, so as to more systematically and comprehensively depict the tumor metabolic characteristics. Compared with traditional methods for classifying metabolic subtypes, the present invention has higher accuracy and can provide a more scientific basis for the early diagnosis of tumors, the selection of personalized treatment plans, and the evaluation of treatment efficacy.
[0056] In addition, the present invention is not only applicable to conventional tumor samples, but can also combine single-cell transcriptome data (scRNA-seq) for the analysis of metabolic heterogeneity of tumor cells, thereby further improving the classification accuracy and revealing the metabolic status in the tumor microenvironment. Through consensus clustering analysis, combined with the GSVA score, different metabolic subtypes can be identified and differential gene expression analysis, functional enrichment analysis, etc. can be performed on them to deeply understand the roles of different metabolic subtypes in the processes of tumorigenesis, development, metastasis, etc.
[0057] The classification method of the present invention has great clinical potential, especially in the field of personalized treatment of tumors. It can predict the sensitivity of tumors to metabolic-targeted drugs and provides new ideas and tools for precision medicine. Through in-depth analysis of different metabolic subtypes, more guiding biomarkers or targets can be provided for the early screening, prognosis evaluation, and efficacy monitoring of cancer in the future, promoting the further development of tumor precision treatment.
Claims
1. A tumor metabolic subtype classification method based on KEGG metabolic pathway GSVA score, characterized in that: The following steps are involved: (1) Obtain scRNA-seq or bulk RNA-seq data; (2) performing dimensionality reduction clustering on the scRNA-seq data of step (1) to obtain different cell types; performing dimensionality reduction clustering on the obtained epithelial cells again to obtain different epithelial subclusters; (3) The different epithelial subclusters in step (2) were subjected to InferCNV copy number variation analysis with T cells, myeloid cells, and endothelial cells as references to preliminarily identify malignant cell subclusters. Then, normal and tumor samples of lung cancer were downloaded from the TCGA database, and the top 50 differentially expressed genes in normal and tumor samples were selected as gene sets for normal epithelial cells and tumor cells, respectively. Single-sample gene set enrichment analysis was then performed to evaluate the expression of these two gene sets in epithelial subclusters to obtain tumor cell subclusters. (4) For the bulk RNA-seq data in step (1), the tumor samples need to be separated; (5) obtaining the gene expression matrix of the tumor cells in step (3) or the gene expression matrix of the tumor sample in step (4); (6) Download the gene set information of metabolic pathways from the KEGG database; based on the information, calculate the GSVA score through the gene expression matrix of right 5; (7) Based on the GSVA scores in step (6), K-means clustering was applied to identify the optimal number of metabolic subtypes; (8) This classification was determined by consensus clustering using the ConsensusClusterPlus method with 1,000 iterations and 80% resampling; (9) Draw a heat map to display the clustering results of step (8); (10) Perform differential gene expression analysis on the clustered metabolic subtypes to obtain the differentially expressed genes (DEGs) of each metabolic subtype and evaluate the differences in gene expression of each metabolic subtype; (11) The DEGs obtained in step (10) were used for GO enrichment analysis to evaluate the functional characteristics of different metabolic subtypes.
2. The classification method according to claim 1, characterized in that: The gene set information of step (6) includes at least one of a lipid metabolism pathway, an energy metabolism pathway, an amino acid metabolism pathway, and a nucleotide metabolism pathway.
3. The classification method according to claim 1, characterized in that: The K-means clustering in step (7) includes using the K-means function in the R language.
4. Use of the method according to any one of claims 1 to 3 in evaluating the efficacy of tumors in non-disease treatment and diagnosis.
Citation Information
Patent Citations
Conbined penholder and blotter
CA40036A
Marker of mesenchymal subtype glioblastoma and application thereof
CN114457163A
Primary liver cancer gene classification and liver cancer tissue energy metabolism-based prognosis analysis method
CN114822688A
Method, system and equipment for classifying and predicting glioblastoma, and medium
CN116030888A
Metabolic gene-based gastric cancer molecular typing and prognosis prediction model construction method and application
CN116798632A
Cited By
Multi-omics malignant pleural effusion immunometabolism reprogramming space-time heterogeneity analysis device
CN121601145A