Liver cancer molecular subtype typing and prognosis model construction method based on glycometabolism and lactic acid metabolism related genes and application of liver cancer molecular subtype typing and prognosis model construction method

By constructing a molecular subtype classification and prognostic model for liver cancer based on genes related to glucose and lactate metabolism, the problem of inconsistent treatment efficacy in existing liver cancer treatments has been solved, enabling the formulation of individualized treatment strategies and improvement of survival outcomes for liver cancer.

CN120913640APending Publication Date: 2025-11-07CHONGQING UNIV CANCER HOSPITAL
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511031919.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Current treatments for liver cancer show significant differences in efficacy, and long-term survival rates for patients have not reached ideal levels. There is a lack of systematic research on the synergistic regulation of metabolic networks by multiple genes.

Method used

Based on the synergistic regulatory network of genes related to glucose metabolism and lactate metabolism, a molecular subtype classification and prognostic model for liver cancer was constructed. By acquiring transcriptome data from liver cancer patients, differentially expressed genes were screened, and a prognostic model was constructed using consistent clustering and random forest models. The CLM score was used for risk assessment.

Benefits of technology

It provides personalized treatment strategies for liver cancer, significantly distinguishing the overall survival, clinical characteristics, and immune cell differences among different subtypes, and improving the accuracy of survival outcome prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913640A_ABST
    Figure CN120913640A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of biomedicine, in particular to a liver cancer molecular subtype typing and prognosis model construction method based on glycometabolism and lactic acid metabolism related genes and application thereof. According to the invention, the typing and prognosis model of liver cancer subtypes is constructed for the first time on the basis of a coordinated regulation network of a glycometabolism gene and a lactic acid metabolism gene, the constructed liver cancer subtype subtypes comprise Cluster1 and Cluster2, and the total survival rates of different subtypes have significant difference; the clinical characteristics T.stage and TNM.stage of different subtypes are obviously different from each other; 12 kinds of immune cells have significant differences among different subtypes; iPS immune scores of different subtypes are significantly different; the TIDE scores of different subtypes have significant score differences; the genes with the highest mutation frequency in different subtypes comprise TP53, CTNNB1 and TTN. CLM scores of the prognosis model constructed on the basis of the collaborative regulation network of the glycometabolism genes and the lactic acid metabolism genes have no significant difference in different patient states; the 17 immune cells have significant differences between high and low risk groups. And a new way is provided for individualized treatment strategy formulation and survival outcome improvement of HCC.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of biomedical technology, and particularly relates to a liver cancer molecular subtype typing and prognosis model construction method based on sugar metabolism and lactic acid metabolism related genes and application thereof. BACKGROUND

[0002] Hepatocellular Carcinoma (HCC) is one of the most common malignant tumors and a major cause of cancer-related deaths. Its main causes include chronic hepatitis B or C virus infection, long-term alcohol abuse, rare genetic diseases and metabolic diseases. Due to the complexity of the pathogenesis of liver cancer, the existing treatment methods have significant differences in efficacy, and the long-term survival rate of patients has not yet reached the ideal level.

[0003] Metabolism is a series of orderly biochemical reactions necessary for the survival of living organisms. The occurrence and development of liver cancer are often accompanied by liver function disorders, leading to tumor metabolic reprogramming. A large amount of evidence shows that cell metabolic disorders not only drive the anabolic growth and proliferation of tumors in harsh microenvironments, but also provide a new strategy for liver cancer treatment by targeting abnormal metabolic pathways. In terms of energy acquisition pathways, liver cancer cells exhibit a typical "Warburg effect": even under rich oxygen conditions, they still prefer to produce energy through the glycolytic pathway. It is worth noting that the glycolysis product lactic acid is not only a metabolic end product, but also an energy material and precursor material that supports the growth of tumor microenvironments and distant metastatic lesions. However, current liver cancer metabolism research focuses on a single molecule, and the systematic study of the coordinated regulation of metabolic networks by multiple genes in the occurrence and development of liver cancer is still insufficient. SUMMARY

[0004] The application is based on a synergistic regulation network of sugar metabolism genes and lactic acid metabolism genes to construct a liver cancer molecular subtype typing and prognosis model. The liver cancer molecular subtype typing constructed includes Cluster1 and Cluster2. There are significant differences in overall survival rate between different subtypes. There are significant differences in clinical characteristics T.stage and TNM.stage between different subtypes. There are significant differences in 12 kinds of immune cells between different subtypes, including activated CD4 T cells, central memory CD4 T cells, eosinophils, (gamma delta T cells, memory B cells, natural killer cells, natural killer T cells, neutrophils, plasmacytoid dendritic cells, T follicular helper cells, type 1 helper T cells, type 17 helper T cells. There are significant differences in IPS immune scores between different subtypes. There are significant differences in TIDE scores between different subtypes. The genes with the highest mutation frequency in different subtypes include TP53, CTNNB1 and TTN. The CLM score in the constructed prognosis model does not have significant differences in different patient states. There are significant differences in 17 kinds of immune cells between high-risk and low-risk groups, including activated CD4 T cells, activated CD8 T cells, activated dendritic cells, CD56-high natural killer cells, central memory CD4 T cells, central memory CD8 T cells, effector memory CD4 T cells, eosinophils, gamma delta T cells, immature B cells, immature dendritic cells, myeloid-derived suppressor cells, memory B cells, plasmacytoid dendritic cells, regulatory T cells, T follicular helper cells, type 2 helper T cells. The application provides a new way for individualized treatment strategy and survival outcome improvement of HCC.

[0005] Therefore, the application aims to provide a liver cancer molecular subtype typing and prognosis model construction method based on sugar metabolism and lactic acid metabolism related genes and application thereof.

[0006] To achieve the above-mentioned purpose, the application provides the following technical scheme.

[0007] The application provides a liver cancer molecular subtype typing method based on sugar metabolism and lactic acid metabolism related genes, including the following steps:

[0008] S1. Obtain a liver cancer patient transcriptome dataset;

[0009] S2. Screen sugar metabolism related genes and lactic acid metabolism related genes; obtain differentially expressed genes through differential analysis; determine overall survival related genes through single factor Cox regression analysis;

[0010] S3. Integrate differentially expressed genes, sugar metabolism related genes, lactic acid metabolism related genes and overall survival related genes, and obtain a first key gene set by taking the intersection;

[0011] S4. Based on the first key gene set obtained in S3, the genes that are synergistically regulated by sugar metabolism and lactic acid metabolism are screened by Spearman correlation analysis;

[0012] S5. Based on the screened genes that are synergistically regulated by sugar metabolism and lactic acid metabolism, ConsensusClusterPlus is used to perform consistent clustering on patients to construct liver cancer molecular subtype typing.

[0013] Further, in S1, the liver cancer patient transcriptome data is derived from the RNA-Seq data of the hepatocellular carcinoma TCGA-LIHC cohort in the UCSC xena database, as well as the data of prognosis information, mutation information and clinical information.

[0014] Further, in S2, the sugar metabolism related genes are screened from the KEGG pathway database, the lactic acid metabolism related genes are screened from the MSigDB database and the paper Lactylation-Related Gene Signature Effectively Predicts Prognosis and Treatment Responsiveness in Hepatocellular Carcinoma, the differential expression genes are calculated by the R package DESeq2 to obtain the significance P value and the differential fold log2FC of each gene, and the genes with adjusted differential significance p.adj less than 0.05 and |log2FC|>1 are selected as the screening threshold for screening, and the total survival related genes are selected as the genes with p value less than 0.05.

[0015] Further, in S4, the screening condition for screening the genes that are synergistically regulated by sugar metabolism and lactic acid metabolism is threshold P<0.05 and |R|>0.7, and 5 genes are screened, including the sugar metabolism genes PKM, UAP1L1, FBP1 and the lactic acid metabolism genes SLC16A3, ALDOB.

[0016] Further, in S5, the liver cancer molecular subtype typing includes Cluster1 and Cluster2, and there is a significant difference in overall survival rate between different subtypes; there is a significant difference in clinical characteristics T.stage and TNM.stage between different subtypes; there are 12 kinds of immune cells that exist significantly different between different subtypes, including activated CD4 T cells, central memory CD4 T cells, eosinophils, (γδT cells, memory B cells, natural killer cells, natural killer T cells, neutrophils, plasma cell-like dendritic cells, T follicular helper cells, type 1 helper T cells, type 17 helper T cells; there is a significant difference in IPS immune score between different subtypes; there is a significant difference in TIDE score between different subtypes; the genes with the highest mutation frequency in different subtypes include TP53, CTNNB1 and TTN.

[0017] The application also discloses a liver cancer prognosis model construction method based on sugar metabolism and lactic acid metabolism related genes, comprising the S1 to S5 steps of the liver cancer molecular subtype typing method in any one of the above, and further comprising the following steps:

[0018] S6. According to the obtained subtype, the subtype grouping is performed based on the liver cancer patient transcription data set, and differential expression analysis is performed to screen subtype differential genes;

[0019] S7. The obtained subtype differential genes are intersected with total survival period related genes to obtain a second key gene set;

[0020] S8. Based on the second key gene set obtained in S7, the top 20 genes are screened according to the MeanDecreaseAccuracy average reduction precision ranking by using a random forest model, and the core genes HPX, SLC16A3 and SLC22A1 are further screened by a LASSO method;

[0021] S9. The risk score CLM score is calculated according to the expression amount of the core genes and the regression coefficient, wherein the CLM score calculation method is as follows:

[0022] CLMscore=HPX exp *-0.02828050+SLC16A3 exp *0.18513830+SLC22A1 exp *-0.02343508。

[0023] S10. The patients are divided into a high-risk group and a low-risk group by using an optimal cutoff value, and the model prediction performance is verified by KM survival analysis and a ROC curve, wherein the cutoff value is determined by using a surv_cutpoint function of an R package surviveminer.

[0024] Further, the differential expression analysis is performed by using an R software package DESeq2 in S6, and the subtype differential genes are screened according to a threshold |logFC|≥2 and P<0.05.

[0025] Further, the CLM score does not exist significant difference in different patient states; activated CD4 T cells, activated CD8 T cells, activated dendritic cells, CD56 high expression natural killer cells, central memory CD4 T cells, central memory CD8 T cells, effector memory CD4 T cells, eosinophils, γδ T cells, immature B cells, immature dendritic cells, myeloid-derived suppressor cells, memory B cells, plasma cell-like dendritic cells, regulatory T cells, T follicular helper cells and type 2 helper T cells, i.e., 17 kinds of immune cells, exist significant difference between the high-risk group and the low-risk group.

[0026] Of course, the liver cancer prognosis model constructed by any one of the above-mentioned methods for constructing a liver cancer prognosis model based on a gene related to sugar metabolism and lactic acid metabolism is also within the protection scope of the present application.

[0027] The present application also provides application of the above-mentioned liver cancer prognosis model constructed by any one of the methods for constructing a liver cancer prognosis model based on a gene related to sugar metabolism and lactic acid metabolism in auxiliary judgment of liver cancer prognosis.

[0028] The present application also provides a marker for a liver cancer prognosis model based on a gene related to sugar metabolism and lactic acid metabolism, wherein the marker is HPX, SLC16A3 or SLC22A1, and the liver cancer prognosis model based on the marker is:

[0029] CLMscore = HPX exp *-0.02828050 + SLC16A3 exp *0.18513830 + SLC22A1 exp *-0.02343508.

[0030] The present application also provides application of the above-mentioned marker for a liver cancer prognosis model based on a gene related to sugar metabolism and lactic acid metabolism in preparation of a liver cancer prognosis evaluation product. The liver cancer prognosis evaluation product includes a detection reagent, a detection kit and the like.

[0031] The application has the beneficial effects that: the application first constructs a liver cancer subtype typing and prognosis model based on a synergistic regulation network of sugar metabolism genes and lactic acid metabolism genes, the liver cancer molecular subtype typing constructed includes Cluster1 and Cluster2, there is a significant difference in overall survival rate between different subtypes; there is a significant difference in clinical characteristics T.stage and TNM.stage between different subtypes; there are 12 kinds of immune cells that exist significant differences between different subtypes, including activated CD4 T cells, central memory CD4 T cells, eosinophils, (gamma delta T cells, memory B cells, natural killer cells, natural killer T cells, neutrophils, plasmacytoid dendritic cells, T follicular helper cells, type 1 helper T cells, type 17 helper T cells; there is a significant difference in IPS immune score between different subtypes; there is a significant difference in TIDE score between different subtypes; the genes with the highest mutation frequency in different subtypes include TP53, CTNNB1 and TTN. The CLM score of the prognosis model constructed based on the synergistic regulation network of sugar metabolism genes and lactic acid metabolism genes in the application does not exist significant difference in different patient states; there are 17 kinds of immune cells that exist significant differences between high-risk and low-risk groups, including activated CD4 T cells, activated CD8 T cells, activated dendritic cells, CD56-high natural killer cells, central memory CD4 T cells, central memory CD8 T cells, effector memory CD4 T cells, eosinophils, gamma delta T cells, immature B cells, immature dendritic cells, myeloid-derived suppressor cells, memory B cells, plasmacytoid dendritic cells, regulatory T cells, T follicular helper cells, type 2 helper T cells. The application provides a new way for individualized treatment strategy and survival outcome improvement of HCC. BRIEF DESCRIPTION OF DRAWINGS

[0032] Figure 1 The figure is a volcano plot of differential genes of tumor samples vs normal samples.

[0033] Figure 2 The figure is a heat map of top10 up-regulated and down-regulated differential genes of tumor samples vs normal samples.

[0034] Figure 3 The figure is a venn diagram of the intersection of differential genes, sugar metabolism genes and OS related genes.

[0035] Figure 4 The figure is a venn diagram of the intersection of differential genes, lactic acid metabolism genes and OS related genes.

[0036] Figure 5 The figure is a correlation heat map of sugar metabolism genes and lactic acid metabolism genes.

[0037] Figure 6 The figure is a GeneMANIA diagram.

[0038] Figure 7Heatmap for subtype clustering.

[0039] Figure 8 KM curve between subtypes.

[0040] Figure 9 Heatmap for expression of genes in different subtypes and their distribution with clinical features.

[0041] Figure 10 Barplot for proportion of immune cell infiltration.

[0042] Figure 11 Boxplot for difference of immune cells between different subtypes.

[0043] Figure 12 Violin plot for difference of immune score between different subtypes.

[0044] Figure 13 Difference of immune cycle between different subtypes, **** represents p<0.0001, *** represents p<0.001, ** represents p<0.01, * represents p<0.05, ns represents p>0.05. p is the p value calculated by Wilcoxon test.

[0045] Figure 14 Difference of TIDE score between different subtypes.

[0046] Figure 15 Difference of expression level of immune checkpoints in different subtypes, **** represents p<0.0001, *** represents p<0.001, ** represents p<0.01, * represents p<0.05, ns represents p>0.05. p is the p value calculated by Wilcoxon test.

[0047] Figure 16 Waterfall plot for cluster1.

[0048] Figure 17 Waterfall plot for cluster2.

[0049] Figure 18 Volcano plot for subtype difference.

[0050] Figure 19 Heatmap for up and down top10 subtype difference.

[0051] Figure 20 GO top5 enrichment conclusion.

[0052] Figure 21 KEGG top10 enrichment bubble plot.

[0053] Figure 22 Venn diagram of intersection of subtype difference genes and OS related genes.

[0054] Figure 23 Top 20 features for random forest MeanDecreaseAccuracy.

[0055] Figure 24 LASSO cross-validation curve.

[0056] Figure 25 LASSO path plot.

[0057] Figure 26 KM survival curve between high and low risk groups in training set.

[0058] Figure 27 ROC curve at 1, 3, 5 years in training set.

[0059] Figure 28 Difference in CLM score between different subtypes in training set.

[0060] Figure 29 KM survival curve between high and low risk groups in validation set.

[0061] Figure 30 ROC curve at 1, 3, 5 years in validation set.

[0062] Figure 31 Heatmap of gene expression between high and low risk groups in validation set.

[0063] Figure 32 KM survival curve between high and low risk groups in immunotherapy cohort.

[0064] Figure 33 Heatmap of gene expression between high and low risk groups in immunotherapy cohort.

[0065] Figure 34 Proportion of different patient status in high and low risk groups.

[0066] Figure 35 Difference in risk score between different patient status.

[0067] Figure 36 Significant different immune cells in different subtypes, **** represents p<0.0001, *** represents p<0.001, ** represents p<0.01, * represents p<0.05, ns represents p>0.05. p is the p value calculated by Wilcoxon test.

[0068] Figure 37 Significant different immune cells in high and low risk groups.

[0069] Figure 38 Filtered single cell results.

[0070] Figure 39 Gene signature variance plot.

[0071] Figure 40 UMAP plot.

[0072] Figure 41 Proportion of 7 different cells in high vs low risk groups.

[0073] Figure 42 T cell subcluster annotation.

[0074] Figure 43 T cell subcluster annotation - proportion of different cells in high vs low risk groups.

[0075] Figure 44 Subpopulation annotation.

[0076] Figure 45 Subpopulation annotation - proportion of different cells in high vs low risk groups.

[0077] Figure 46 Lattice plot.

[0078] Figure 47 Calibration curve.

[0079] Figure 48 Lattice performance assessment ROC curve.

[0080] Figure 49 DCA curve.

[0081] Figure 50 HPX single gene analysis, A is KM curve, B is gene expression ridge plot.

[0082] Figure 51 SLC16A3 single gene analysis, A is KM curve, B is gene expression ridge plot.

[0083] Figure 52 SLC22A1 single gene analysis, A is KM curve, B is gene expression ridge plot. DETAILED DESCRIPTION

[0084] In order to make the technical problems solved by the present application, technical solutions and beneficial effects more clearly understood, the technical solutions of the present application will be further described below in conjunction with the drawings and specific examples. It should be understood that the following examples are only illustratively described and explained the present application, and should not be interpreted as limiting the scope of protection of the present application. Any technology implemented based on the above description of the present application is included in the scope of protection intended by the present application. It should be noted that the experimental methods used in the following examples are conventional methods, and the materials, reagents, etc. used in the following examples, if not specifically stated, can be obtained from commercial channels.

[0085] Example 1

[0086] 1. Obtain transcriptome dataset of liver cancer patients

[0087] Download RNA-Seq data of hepatocellular carcinoma (TCGA-LIHC) cohort from UCSC xena database (https: / / xena.ucsc.edu / ), as well as data of prognosis information, mutation information and clinical information. According to the data of prognosis information and clinical information, extract cancer tissue samples with prognosis information and corresponding normal tissue samples, remove samples with survival time of 0 and repeated samples, finally a total of 365 tumor tissue samples and 50 normal tissue samples are included as a training set [xena_LIHC_*.csv].

[0088] Download dataset with data number GSE14520 from NCBI GEO database (http: / / www.ncbi.nlm.nih.gov / geo / ), sequencing platform is GPL571 [HG-U133A_2] Affymetrix Human Genome U133A2.0 Array; GPL3921 [HT_HG-U133A] Affymetrix HT Human Genome U133A Array. Extract tumor tissue samples with prognosis information, a total of 221 tumor samples, as a validation set [dat.GSE14520.csv / survival.GSE14520.csv].

[0089] Download dataset with data number GSE151530 from NCBI GEO database, extract 32 Hepatocellular carcinoma single cell samples from it for analysis.

[0090] 2. Screening of sugar metabolism related genes and lactic acid metabolism related genes

[0091] Carbohydrate metabolism (CRG) related gene screening collects carbohydrate related pathway genes [CRGs_gene.csv] from KEGG pathway database (https: / / www.genome.jp / kegg / pathway.html#global) through Carbohydrate metabolism, a total of 358.

[0092] Lactate metabolism (LMRG) related gene signature The gene sets related to lactate metabolism were retrieved from MSigDB database (http: / / www.gsea-msigdb.org / gsea / downloads.jsp), including "GOBP_GLUCLOSE_CATABOLIC_PROCESS_TO_LACTATE_VIA_PYRUVATE", "GOBP_LACTATE_METABOLIC_PROCESS", "GOBP_LACTATE_TRANSMEMBRANE_TRANSPORT", "GOMF_L_LACTATE_DEHYDROGENASE_ACTIVITY", "GOMF_LACTATE_DEHYDROGENASE_ACTIVITY", "GOMF_LACTATE_TRANSMEMBRANE_TRANSPORTER_ACTIVITY", "HP_ABNORMAL_BRAIN_LACTATE_LEVEL_BY_MRS", "HP_ABNORMAL_LACTATE_DEHYDROGENASE_LEVEL", "HP_ELEVATED_LACTATE_PYRUVATE_RATIO", "HP_INCREASED_CIRCULATING_LACTATE_DEHYDROGENASE_CONCENTRATION", "HP_INCREASED_CSF_LACTATE_LEVEL", "HP_INCREASED_SERUM_LACTATE_LEVEL", and the lactylation-related genes obtained from previous study (Lactylation-Related Gene Signature Effectively Predicts Prognosis and Treatment Responsiveness in Hepatocellular Carcinoma), the gene sets were integrated, and a total of 442 genes were obtained [LMRGs_gene.csv].

[0093] Immuno-therapy dataset: IMvigor210 was obtained from R package IMvigor210CoreBiologies (version = 2.0.0; http: / / research-pub.gene.com / IMvigor210CoreBiologies).

[0094] 3. Differential analysis

[0095] Based on the TCGA transcriptome data, the significance P value and the difference fold log2FC of each gene were calculated using R package DESeq2 (version = 1.38.3; https: / / bioconductor.org / packages / release / bioc / html / DESeq2.html) for tumor samples vs normal samples. The adjusted difference significance p.adj less than 0.05 and |log2FC|>1 were selected as the screening threshold, and a total of 3977 differential genes were screened [01.DEG_sig.csv], of which 1120 were up-regulated genes and 2857 were down-regulated genes. According to the differential expression results, a volcano plot was drawn as shown in Figure 1 . At the same time, the expression of the top ten up-regulated and down-regulated genes was extracted to draw a difference heat map to observe the distribution of differential genes in different samples, as shown in Figure 2 .

[0096] 4. Single factor Cox regression analysis

[0097] Based on the TCGA transcriptome data, all protein-coding genes in the patient's RNA-seq data were subjected to single factor Cox regression analysis by R package survival (version = 3.5.7; https: / / cran.r-project.org / web / packages / survival / index.html) to determine the genes related to overall survival (OS). Only genes with p value<0.05 were selected, and 3736 overall survival-related genes were screened.

[0098] Integrating differential genes (DEG), sugar metabolism and lactic acid metabolism genes, and overall survival (OS) related genes [01.cox_0.05.csv], the intersection was taken to obtain the optimal sugar metabolism (CRGs) and lactic acid metabolism genes (LMRGs). Finally, the intersection of DEG, CRGs, and OS-related genes had 50 [02.DEG_CRGs_OS_gene.csv], and the intersection of DEG, LMRGs, and OS-related genes had 9 [02.DEG_LMRGs_OS_gene.csv]. A venn diagram was drawn, as shown in Figure 3 、 Figure 4 .

[0099] 5. Interaction between sugar metabolism and lactic acid metabolism genes

[0100] To obtain genes related to both glucose metabolism and lactate metabolism, spearman correlation was calculated between the optimal genes of glucose metabolism and lactate metabolism obtained above using R package psych (version = 2.3.9; https: / / cran.r-project.org / web / packages / psych / index.html), and the optimal genes were screened according to the threshold P < 0.05 and |R| > 0.7 [01.Correlaion.csv], and finally 5 genes were screened (glucose metabolism: PKM, UAP1L1, FBP1, lactate metabolism: SLC16A3, ALDOB) [01.sig_cor_0.7_0.05.csv], and the correlation heat map was drawn as shown in Figure 5 .

[0101] 6、GeneMANIA analysis

[0102] To further study the possible mechanism of action of glucose metabolism and lactate metabolism genes, based on the optimal genes screened above, an interaction network was constructed using GeneMANIA (https: / / genemania.org / ) database, different color lines represent different association relationships (Physical Interactions, Co-expression, Predicted Interactions, Co-localization, Genetic Interactions, Pathway, Shared Protein Domains), and different color blocks represent different metabolic processes. As shown in Figure 6 .

[0103] 7、Construction of molecular subtypes

[0104] Based on the optimal genes, the TCGA patient dataset was subjected to consensus clustering analysis by R package ConsensusClusterPlus (version = 1.62.0; https: / / bioconductor.org / packages / release / bioc / html / ConsensusClusterPlus.html) to accurately classify patients into different subtypes, and KM (Kaplan-Meier) survival analysis was performed between subtype groups using R package survival (version = 3.5.7; https: / / cran.r-project.org / web / packages / survival / index.html). The expression of genes in different subtypes and their correlation with clinical characteristics (age, gender, TNM, TNM.Stage) were analyzed.

[0105] ConsensusClusterPlus consensus clustering divided the TCGA patient dataset into two subtypes (cluster.csv), of which Cluster1 had 280 samples and Cluster2 had 85 samples, as shown in Figure 7 KM curve was drawn between subtype groups, as shown in Figure 8 It can be seen that there is a significant difference in overall survival rate between subtypes.

[0106] The expression of optimal genes in different subtypes and their distribution with clinical characteristics were analyzed, as shown in Figure 9 The heatmap directly shows the expression pattern of optimal genes in different subtypes, and the horizontal label reflects the distribution relationship between the expression of these genes and each clinical characteristic. For example, certain genes are highly expressed in a specific subtype, which may be related to certain clinical characteristics (such as M0 state or specific age group). This result suggests that the expression characteristics of different genes may form unique molecular characteristics among subtypes, thereby driving subtype-specific biological behaviors and possibly affecting the clinical outcome of patients. The difference in clinical characteristics in different subtypes is shown in Table 1, and it can be seen from the table that there is a significant difference in clinical characteristics T.stage and TNM.stage in different subtypes.

[0107] Table 1

[0108]

[0109] 8. Analysis of immune infiltration in different molecular subtypes

[0110] The proportions of 22 immune cell infiltration types in each sample were calculated using CIBERSORT in the R package IOBR (version = 0.99.0; https: / / github.com / IOBR / IOBR), and the proportions are plotted as shown in Figure 10. The differences in immune cell infiltration among different subtypes were analyzed using the rank-sum test, and 12 immune cell types showed significant differences among different subtypes [03.DE.cibersort.csv]. Figure 11 As shown.

[0111] The immune score for each sample was calculated using the IPS method in the R package IOBR [05.IPS_res.csv]. A violin plot showing the differences in immune scores between different subtypes was then plotted, as shown below. Figure 12 As shown in the figure, there are significant differences in immune scores among different subtypes (p<0.05).

[0112] The immune cycle results for each sample were obtained using TIP [TIP_res.csv], and the differences in tumor-infiltrating immune cells in the seven-step cancer immune cycle under different subtypes were plotted [04.TIP_stat.csv]. Figure 13 As shown in the image, the scores obtained through TIDE [TIDE_res.csv] were plotted as a violin graph illustrating the differences in TIDE scores across different subtypes. Figure 14 As shown. By Figure 14 As can be seen, there were significant differences in TIDE scores among different subtypes (p<0.05). Figure 13 It is evident that significant differences exist among different subtypes in the eight key steps of the immune cycle, including Step 2. Cancer antigen presentation, Step 3. Priming and activation, Step 4. Eosinophil recruitment, Step 4. MDSC recruitment, Step 4. Monocyte recruitment, Step 4. Treg cell recruitment, Step 6. Recognition of cancer cells by T cells, and Step 7. Killing of cancer cells.

[0113] 9. Analysis of Immune Checkpoint Gene Expression

[0114] Based on the TCGA patient dataset, the expression levels of immune checkpoint genes in subtypes were compared, and the expression levels of 33 immune checkpoints such as PDCD1, TIGIT, CTLA4, CD80, CD86 were screened [01.checkpoint_result.csv], and the expression levels of 29 immune checkpoints were significantly different between different subtypes [02.DE.checkpoint_result.csv], and a box plot was drawn according to the results of the significantly different immune checkpoints, as shown in Figure 15 .

[0115] 10. Somatic mutations in different molecular subtypes

[0116] Based on the TCGA patient dataset, the mutation annotation format (MAF) file of the patient was downloaded from the TCGA database using the R package TCGAmutations, and the somatic mutations of patients in different subtypes were analyzed using the R package maftools, and a waterfall plot of different subtypes was drawn, as shown in Figure 16 , 17 From the figure, it can be seen that the genes with the highest mutation frequency in different subtypes include TP53, CTNNB1, TTN, etc. The types of mutations of each gene are also divided into missense mutation (Missense Mutation), nonsense mutation (Nonsense Mutation), splice site mutation (Splice Site), frame shift insertion (Frame_Shift_Ins), frame shift deletion (Frame_Shift_Del), etc. Various categories are visually displayed in the figure by different colors. In addition, from the right bar chart, it can be seen that the mutation frequency distribution of each gene in different samples, thereby revealing the potential differences in somatic mutation patterns between different subtypes.

[0117] 11. Subtype difference analysis

[0118] According to the obtained subtypes, based on the TCGA patient dataset, the difference expression analysis was performed using the R software package DESeq2 according to the subtype grouping, and the differential genes were screened according to the threshold |logFC|≥2 and P<0.05, a total of 536 differential genes were obtained, of which 354 were up-regulated and 182 were down-regulated, and a volcano plot was drawn according to the differential expression results, as shown in Figure 18 At the same time, the expression levels of the top ten up-regulated and down-regulated genes were extracted to draw a difference heat map to observe the distribution of differential genes in different samples, as shown in Figure 19 .

[0119] 12. Subtype enrichment analysis

[0120] KEGG and GO enrichment analysis were performed using R package clusterProfiler to obtain the results of gene set enrichment (P.adj < 0.05). Among them, the GO enrichment results had a total of 356 (BP: 243; CC: 38; MF: 75), and the top 5 enrichment results of each term were selected according to p.adj from small to large to draw a column chart, as shown in Figure 20 BP: Cellular response to xenobiotic stimulus (cellular response to xenobiotic stimulus), cellular hormone metabolic process (cellular hormone metabolic process), small molecule catabolic process (small molecule catabolic process), response to xenobiotic stimulus (response to xenobiotic stimulus), retinoid metabolic process (retinoid metabolic process), MF: collagen-containing extracellular matrix (collagen-containing extracellular matrix), basal plasma membrane (basal plasma membrane), basal part of cell (basal part of cell), basolateral plasma membrane (basolateral plasma membrane), synaptic membrane (synaptic membrane), CC: steroid hydroxylase activity (steroid hydroxylase activity), tetrapyrrole binding (tetrapyrrole binding), heme binding (heme binding), receptor ligand activity (receptor ligand activity), oxidoreductase activity, acting on paired donors, with incorporation or reduction of molecular oxygen, reduced flavin or flavoprotein as one donor, and incorporation of one atom of oxygen (oxidoreductase activity, acting on paired donors, with incorporation or reduction of molecular oxygen, reduced flavin or flavoprotein as one donor, and incorporation of one atom of oxygen).

[0121] The KEGG enrichment results had a total of 262, and the top 10 enrichment results were selected according to p.adj from small to large to draw a bubble chart, as shown in Figure 21Among them are Retinol metabolism, Bile secretion, Drug metabolism-cytochrome P450, Metabolism of xenobiotics by cytochrome P450, Glycolysis / Gluconeogenesis, Neuroactive ligand-receptor interaction, Chemical carcinogenesis-DNA adducts, Glucagon signaling pathway, PPAR signaling pathway, Salivary secretion.

[0122] 13. Prognosis prediction model construction

[0123] The obtained subtype differential genes are intersected with genes related to overall survival (OS), and the obtained genes are used for subsequent analysis, such as Figure 22 As shown in the figure, a total of 164 genes are obtained.

[0124] The random forest model is used to screen genes using the R package randomForest, and the genes are ranked according to MeanDecreaseAccuracy (average decrease in accuracy). The larger the MeanDecreaseAccuracy value, the more important the feature to the prediction ability of the model. The top 20 genes [01.importance_rf.csv] are selected, as shown in Figure 23 According to the R package glmnet, the genes screened by the random forest model are further screened using the LASSO method [02.Lasso_Coefficients.csv]. According to lambda.min=0.0371391626827193, three genes (HPX, SLC16A3, SLC22A1) are finally screened, and the LASSO cross-validation curve and path diagram are drawn, as shown in Figure 24 、 25

[0125] The screened genes and regression coefficients are used to construct a prognosis prediction feature model to obtain a CLM score. The CLM score is calculated as follows:

[0126] CLMscore=HPX​exp *-0.02828050+SLC16A3 exp *0.18513830+SLC22A1 exp *-0.02343508

[0127] 14. Verify the predictive power of the CLM score.

[0128] Based on the CLM score, patients of the subtype were divided into high-risk and low-risk groups according to the optimal cutoff value using the surv_cutpoint function in the R package survminer (version = 0.4.9; https: / / cran.r-project.org / web / packages / survival / index.html). To assess the difference in survival rates between the two groups, KM survival curves were generated using the R package survival. Figure 26 In addition, the R package survivalROC (version = 1.0.3.1; https: / / github.com / cran / survivalROC) is used to generate receiver operating characteristic (ROC) curves. Figure 27 To assess the predictive ability of the prognostic model, and to analyze the differential expression of CLM scores among different subtypes. Figure 28 ).Depend on Figure 26 It is evident that there are significant differences in the KM curves between the high-risk and low-risk groups, and that... Figure 27 As can be seen, the AUC curve was greater than 0.65 at 1, 3, and 5 years, indicating that the model has high accuracy in predicting patient survival risk and strong predictive ability in clinical practice. Figure 28 It is evident that there are significant differences in CLM scores among different subtypes.

[0129] To verify the model's stability and obtain validation set data and clinical information, based on the CLM score, the data were divided into high-risk and low-risk groups using the optimal cutoff method, and KM curve survival analysis was performed. Figure 29 Finally, the ROC curve ( Figure 30 Risk heat map Figure 31 To evaluate the effectiveness of the CLM score on the validation set. Figure 29 It is evident that there are significant differences in the KM curves between the high- and low-risk groups in the validation set, and that... Figure 30 As can be seen, the AUC curve is greater than 0.65 in both the 3 and 5 years, and also exceeds 0.6 in the 1 year, very close to 0.65, indicating that the model also has strong predictive ability in the validation set. Figure 31 The gene expression was shown in the high- and low-risk groups.

[0130] Simultaneously, based on CLM score analysis, the IMvigor210 CoreBiologies immunotherapy cohort was analyzed using KM curves ( Figure 32 Risk heat map Figure 33 Assess the effectiveness of the CLM score in the validation set. Plot the proportions of patients with complete remission / partial remission / disease stability in the high- and low-risk groups. Figure 34 ).Depend on Figure 32 It is evident that there are significant differences in the KM curves between the high- and low-risk groups in the validation set. Figure 33 The expression of genes in the high- and low-risk groups was shown. It can be seen that the expression of SLC16A3 is significantly different between the high- and low-risk groups, while the expression of the other two genes is not. Figure 34 The data shows the proportions of different patient states in the high- and low-risk groups. It can be seen that the proportion of patients with disease remission (CR) / partial remission (PR) / stable disease (SD) is higher in the low-risk group than in the high-risk group, while the proportion of patients with disease progression (PD) is higher in the high-risk group than in the low-risk group. (NE indicates non-assessment). Figure 35 The figure shows the differences in CLM scores among different patient states. As can be seen from the figure, there are no significant differences in CLM scores among different patient states.

[0131] 15. ssGSEA

[0132] The enrichment scores of 28 immune cell infiltration types in each sample were calculated using the `calculate_sig_score` function from the R package `IOBR`. Differential expression of immune cells in subtype analysis and high- and low-risk groups was then analyzed. Figure 36 It was found that 12 types of immune cells (Activated CD4 T cells, Central memory CD4 T cells, Eosinophils, Gamma delta T cells, Memory B cells, Natural killer cells, Natural killer T cells, Neutrophils, Plasmacytoid dendritic cells, T follicular helper cells, Type 1 T helper cells, and Type 17 T helper cells) exhibited significant differences among different subtypes. Figure 37It can be seen that there are 17 kinds of immune cells (Activated CD4 T cell, Activated CD8 T cell, Activated dendritic cell, CD56bright natural killer cell, Central memory CD4 T cell, Central memory CD8 T cell, Effector memory CD4 T cell, Eosinophil, Gamma delta T cell, Immature B cell, Immature dendritic cell, MDSC, Memory B cell, Plasmacytoid dendritic cell, Regulatory T cell, T follicular helper cell, Type 2 T helper cell) that are significantly different between the high-risk group and the low-risk group.

[0133] 16, scRNA-seq analysis

[0134] For single-cell feature research, the scRNA-seq data set was analyzed according to the standard program of R package Seurat. According to previous reports, cells with less than 200 genes, more than 7000 total genes, and more than 20% mitochondrial genes were filtered out. The results after filtering (number of cells: 44344, number of genes: 18667) are shown in Figure 38 .

[0135] The R package harmony (version = 1.2.0; https: / / github.com / immunogenomics / harmony) can reduce the batch effect between samples. The FindVariableFeatures function is used to identify the top 2000 variable expression genes, and the gene feature variance chart is drawn as shown in Figure 39 .

[0136] After grouping cells by FindNeighbors, FindClusters(resolution=0.5), optimizing standard modules, and reducing the dimension of data by RunUMAP function, the single cells were annotated by R package SingleR, and the annotation results are shown in Figure 40 , which annotated 7 cell types in total, including B cell (B_cell), endothelial cell (Endothelial_cells), hepatocyte (Hepatocytes), macrophage (Macrophage), monocyte (Monocytes), smooth muscle cell (Smooth_muscle_cells), and T cell (T_cells).

[0137] According to the CLM score of each sample after analysis, the samples of scRNA-seq data were divided into high-risk group and low-risk group according to the median value by PesudoBulk analysis, and the proportion of different cells in high-risk group and low-risk group was compared. From Figure 41 It can be seen that the proportion of T cells in high-risk group and low-risk group has the largest difference.

[0138] Then, the T cells were sub-cluster annotated ( Figure 42 ), and the proportion difference of T cells in high-risk group and low-risk group was observed, as shown in Figure 43 , it can be seen that CD4+_effector_memory has the largest proportion difference in high-risk group and low-risk group.

[0139] The CD4+_effector_memory cells were further sub-cluster annotated, as shown in Figure 44 , and the proportion difference of CD4+_effector_memory cells in high-risk group and low-risk group was observed again ( Figure 45 ).

[0140] 17. Establishing a prediction nomogram

[0141] The R package rms (version=6.7.1; https: / / cran.r-project.org / web / packages / rms / index.html) was used to construct a nomogram based on survival time, survival status, age, gender, T stage, N stage, M stage, TNM status, and CLM score ( Figure 46 ). The Harrell consistency index (C-index), calibration curve ( Figure 47 ), and ROC curve ( Figure 48 ) were used to draw DCA decision curve by R package ggDCA (version=1.2; https: / / cran.r-project.org / web / packages / ggDCA / index.html)Figure 49 ), evaluate the performance of nomogram. It can be seen from the combination of calibration curve and ROC curve that the nomogram has good prediction effect. DCA helps to evaluate the practicability of the prediction model in clinical decision making by showing the net benefit at different time points of 1 year, 3 years and 5 years.

[0142] 18. Single gene analysis

[0143] In order to understand how the CLM score predicts the prognosis of liver cancer patients, single gene analysis, KM survival analysis were performed on the model characteristic genes of the subtypes, and the expression of the model characteristic genes in different subtypes, TCGA samples with diseases and control samples were analyzed. The KM curve of each gene and the gene expression ridge map are shown in Figures 50-52 As can be seen from the figure, according to the median value of the model characteristic genes of the subtypes, the high and low expression groups are divided, and there are significant differences in the model characteristic genes in the high and low expression groups. The ridge map shows the expression distribution of the model characteristic genes among the samples with diseases, control samples and subtypes. It can be seen that the expression distribution of all model characteristic genes in the control samples is relatively concentrated, and the expression distribution in the samples with diseases and subtypes is relatively dispersed.

[0144] Finally, it should be noted that the above examples are only used to illustrate the technical solutions of the present application and not to limit it. Although the present application has been described in detail with reference to the above examples, those skilled in the art should understand that any modification, equivalent replacement, improvement, etc. within the spirit and principles of the present application should be included in the protection scope of the present application.

Claims

1. A method for classifying a liver cancer molecular subtype based on genes related to sugar metabolism and lactic acid metabolism, characterized by, The method comprises the following steps: S1. Obtain a hepatocellular carcinoma patient transcriptome dataset; S2. Screen for sugar metabolism-related genes and lactic acid metabolism-related genes; Obtain differentially expressed genes through differential analysis; determine total survival-related genes through single-factor Cox regression analysis; S3. Integrate the differentially expressed genes, sugar metabolism-related genes, lactic acid metabolism-related genes and total survival-related genes, and obtain a first key gene set by taking the intersection; S4. Based on the first key gene set obtained in S3, screen for sugar metabolism and lactic acid metabolism synergistically regulated genes through Spearman correlation analysis; S5. Based on the screened sugar metabolism and lactic acid metabolism synergistically regulated genes, perform consistent clustering on patients by using ConsensusClusterPlus, and construct a hepatocellular carcinoma molecular subtype classification.

2. The liver cancer molecular subtyping method according to claim 1, characterized by, In S1, the hepatocellular carcinoma patient transcriptome data is derived from the RNA-Seq data of the hepatocellular carcinoma TCGA-LIHC cohort in the UCSC xena database, as well as the data of prognosis information, mutation information and clinical information; in S2, the sugar metabolism-related genes are screened from the KEGG pathway database, the lactic acid metabolism-related genes are screened from the MSigDB database and the paper Lactylation-Related Gene Signature Effectively Predicts Prognosis and Treatment Responsiveness in Hepatocellular Carcinoma, the significance P value and the difference fold log2FC of each gene are calculated by the R package DESeq2, and the screening threshold is selected as p.adj less than 0.05 and |log2FC|>1, and the total survival-related genes are selected as genes with a p value less than 0.05; in S4, the screening condition for screening the sugar metabolism and lactic acid metabolism synergistically regulated genes is threshold P<0.05 and |R|>0.7, and 5 genes are screened, including the sugar metabolism genes PKM, UAP1L1 and FBP1, and the lactic acid metabolism genes SLC16A3 and ALDOB.

3. The liver cancer molecular subtyping method according to claim 1, characterized by, In S5, the hepatocellular carcinoma molecular subtype classification comprises Cluster1 and Cluster2, and the overall survival rates of different subtypes are significantly different; the clinical characteristics T.stage and TNM.stage of different subtypes are significantly different; there are 12 immune cells that are significantly different between different subtypes, including activated CD4 T cells, central memory CD4 T cells, eosinophils, (gamma delta T cells, memory B cells, natural killer cells, natural killer T cells, neutrophils, plasma cell-like dendritic cells, T follicular helper cells, type 1 helper T cells and type 17 helper T cells; the IPS immune scores of different subtypes are significantly different; the TIDE scores of different subtypes are significantly different; the genes with the highest mutation frequency in different subtypes include TP53, CTNNB1 and TTN.

4. A method for constructing a hepatocellular carcinoma prognostic model based on genes related to glucose and lactate metabolism, characterized in that, The S1-S5 steps of the liver cancer molecular subtype typing method according to any one of claims 1-3 further comprise the following steps: S6. Based on the obtained subtype, perform differential expression analysis on the liver cancer patient transcript data set subtype grouping, and screen subtype differential genes; S7. Intersect the obtained subtype differential genes with the overall survival-related genes to obtain a second key gene set; S8. Based on the second key gene set obtained in S7, use a random forest model to rank and screen top 20 genes according to MeanDecreaseAccuracy average reduction in accuracy, and then use LASSO method to further screen to obtain core genes HPX, SLC16A3, and SLC22A1; S9. Calculate the risk score CLM score according to the expression amount of the core genes and the regression coefficient, wherein the CLM score calculation method is as follows: CLMscore = HPX exp * 0.02828050 + SLC16A3 exp * 0.18513830 + SLC22A1 exp * 0.02343508. S10. Divide the patients into a high-risk group and a low-risk group according to the optimal cutoff value, and verify the model prediction performance through KM survival analysis and ROC curve, wherein the cutoff value is determined using the surv_cutpoint function of the R package survminer.

5. The method of claim 4, wherein the liver cancer prognosis model is constructed by using the gene expression data of the liver cancer tissue and the gene expression data of the normal liver tissue. In S6, the R software package DESeq2 is used for differential expression analysis, and subtype differential genes are screened according to the threshold |logFC|≥2 and P<0.

05.

6. The method for constructing a liver cancer prognosis model based on genes related to sugar metabolism and lactic acid metabolism according to any one of claims 4 to 5, characterized in that, Further comprising: S11. Use the R package rms to construct a prediction nomogram based on survival time, survival status, age, gender, T stage, N stage, M stage, TNM status, and CLM score, and use Harrell consistency index, calibration curve, and DCA decision curve drawn using the R package ggDCA to evaluate the performance of the nomogram.

7. The liver cancer prognosis model constructed by the method according to any one of claims 4-6.

8. The application of the liver cancer prognosis model according to claim 7 in the auxiliary judgment of liver cancer prognosis.

9. A marker for a liver cancer prognosis model based on genes related to sugar metabolism and lactic acid metabolism, characterized by, The markers are HPX, SLC16A3, and SLC22A1, and the liver cancer prognosis model based on the markers is: CLMscore = HPX exp * 0.02828050 + SLC16A3 exp * 0.18513830 + SLC22A1 exp * 0.02343508.

10. The application of the markers of the liver cancer prognosis model based on sugar metabolism and lactic acid metabolism-related genes according to claim 9 in the preparation of liver cancer prognosis evaluation products.

Citation Information

Cited By

  • Metabolism-related fatty liver disease single cell data set visualization method and system

    CN121687216A