A novel system for predicting the prognosis of liver cancer patients based on the molecular characteristics of senescent T cells
By establishing a liver cancer prognostic risk assessment system based on the expression of CDCA8, MCM4, and TTK genes, the problem of inaccurate prediction for specific populations by existing models has been solved, enabling efficient prognostic risk assessment and personalized recommendations for treatment plans for liver cancer patients.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING DITAN HOSPITAL CAPITAL MEDICAL UNIVERSTY
- Filing Date
- 2026-02-06
- Publication Date
- 2026-05-26
AI Technical Summary
Existing prognostic prediction models for liver cancer patients have unclear predictive effects on certain specific populations and lack prognostic risk assessment based on T-cell aging-related genes, often resulting in missing the optimal treatment period during early diagnosis.
A predictive system based on the expression levels of three T cell senescence-related genes, CDCA8, MCM4, and TTK, was established. Patient information was obtained through a data collection module, a risk score was calculated using a prognostic risk discrimination module, and the prognostic risk level of the patient was determined by combining the scoring rules of the prognostic judgment module. A visual nomogram model was then constructed for prediction.
It can classify patients into high-risk and low-risk groups, providing more accurate prognostic assessments, helping doctors develop personalized treatment plans, and improving patient survival rates.
Smart Images

Figure CN122090924A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of healthcare informatics, specifically relating to a novel system for predicting the prognosis of liver cancer patients based on biological information features and a method for using the system to predict the prognosis of liver cancer patients. Background Technology
[0002] Hepatocellular carcinoma (HCC) accounts for over 90% of primary liver tumors and is currently the sixth most common malignant tumor worldwide, becoming the third leading cause of cancer-related deaths globally. With a deeper understanding of cancer and advancements in technology, new treatments and methods are constantly emerging, providing doctors and patients with more options. However, early-stage HCC generally presents with no specific clinical manifestations, and its progression is very insidious. Most patients are diagnosed at an advanced stage, missing the optimal treatment window, resulting in a 5-year survival rate of only 14.1%. Immunotherapy, such as immune checkpoint inhibitors, which has become a hot topic in recent years, often yields inconsistent treatment outcomes for HCC patients. Therefore, if an effective prognostic (or risk) prediction model for HCC patients could be developed to predict and classify patients, it would provide valuable reference and information for doctors and patients when choosing treatment options.
[0003] Existing technologies do include models for predicting the prognosis of liver cancer patients, such as Okuda, TNM, BCLC, Child, and CLIP. The predictive indicators of these prognostic models mainly focus on tumor burden, liver function, and living status. However, these classic models were developed relatively early, and the predictive factors they contain are relatively simple, so their predictive effects on certain specific populations are unknown (Zhou Dongdong, et al. Analysis and comparison of common liver cancer prognostic prediction models [J]. Chinese Journal of Clinical Oncology, 2020, 47(24):1281-1286.). With in-depth research on the biological characteristics related to liver cancer, prognostic models based on biomolecular characteristics closely related to the occurrence and development of liver cancer have emerged, such as the Chinese invention patent application "System for predicting the prognosis of patients with non-surgical treatment of primary liver cancer based on platelet / spleen length ratio" (publication number CN115985497A, publication date April 18, 2023).
[0004] Based on a data review of the TCGA database, the inventors discovered a strong negative correlation between high expression of T cell senescence-related genes and overall survival (OS) in HCC patients. Previous studies have shown that T cell senescence, besides exhaustion, is another key functional impairment state of T cells in the tumor microenvironment. Senescent T cells exhibit severely impaired function and secrete large amounts of inflammatory factors, exacerbating the formation of the tumor immunosuppressive microenvironment and accelerating tumor progression (Calcinotto A, et al. Cellular Senescence: Aging, Cancer, and Injury[J]. Physiological Reviews. 2019 Apr;99(2):1047–78; Muñoz-Espín D, et al. Cellular senescence: from physiology to pathology[J]. Nat Rev Mol Cell Biol. 2014 Jul;15(7):482–96.). Currently, there are no reports of prognostic models for HCC related to T cell senescence. Therefore, the inventors envision establishing a prognostic risk assessment model based on the aforementioned T-cell senescence-related genes to predict the prognosis of HCC patients, thereby providing doctors with a better understanding of the condition of HCC patients, predicting the course of the disease, and providing helpful assistance in responding to immunotherapy. Summary of the Invention
[0005] Therefore, this invention provides a novel system for predicting the prognosis of patients with primary liver cancer based on the expression levels of three T-cell senescence-related genes: CDCA8, MCM4, and TTK. This system can classify patients with primary liver cancer into a high-risk group with poor prognosis and weak response to immunotherapy, and a low-risk group with better prognosis and potentially more aggressive response to immunotherapy. This allows for prediction of patient prognosis, providing physicians with a deeper understanding of the condition of patients with primary liver cancer and offering valuable reference for clinical treatment decisions.
[0006] The present invention achieves its objective through the following technical solutions.
[0007] A system for predicting the prognosis of patients with primary liver cancer based on the molecular characteristics of senescent T cells, the system being based on the expression levels of senescence-related genes in T cells from liver cancer tissue that has been excised from the patient, comprising:
[0008] Data collection module: set to acquire data on the patient's age, sex, T stage, N stage, M stage, Grade level, Stage stage, and expression levels of genes CDCA8, MCM4, and TTK in liver cancer tissue;
[0009] The prognostic risk assessment module is configured to substitute the expression levels of the genes CDCA8, MCM4, and TTK obtained by the data collection module into formula (1) to calculate the prognostic risk score of the patient. The prognostic risk is determined based on the calculated risk score, where a prognostic risk score ≥ 0.316006455458479 is considered high risk, and a prognostic risk score < 0.316006455458479 is considered low risk.
[0010]
[0011] Among them, [CDCA8], [MCM4], and [TTK] represent the expression levels of the corresponding genes;
[0012] Prognostic module: This module is configured to use an established nomogram prognostic prediction model for primary liver cancer patients to assign scores to age, sex, T, N, M, Grade, Stage, and risk score, calculate a total score, and determine the prognostic outcome based on the total score. The scoring rules in the nomogram prognostic prediction model for primary liver cancer patients are as follows:
[0013] Age: ≤10 years old, score is 0; >10 years old, score is calculated according to formula (2):
[0014]
[0015] Where A represents the patient's age;
[0016] Gender: Females score 0, males score 1.06;
[0017] T-stage: T1 score is 0, T2 score is 52.27, T3 score is 68.27, T4 score is 71.27;
[0018] N-stage: N0 score is 0, N1 score is 12.15;
[0019] M stage: M0 score is 0, M1 score is 44.26;
[0020] Grade classification: G1 and G2 scores are 0, G3 scores are 0.53, and G4 scores are 7.11;
[0021] Stage classification: Stage I score is 0, Stage II score is 38.72, Stage III score is 45.41, and Stage IV score is 100;
[0022] Prognostic risk: Low risk score is 0, high risk score is 7.92.
[0023] Preferably, [CDCA8], [MCM4], and [TTK] are the FPKM values of the corresponding genes.
[0024] Preferably, the prognosis refers to the 1-year survival probability, 3-year survival probability, and / or 5-year survival probability after a patient is diagnosed with primary liver cancer.
[0025] Therefore, the present invention also provides a method for predicting the prognosis of patients with primary liver cancer; the method is based on the system for predicting the prognosis of patients with primary liver cancer based on the molecular characteristics of senescent T cells described in the present invention, and includes the following steps:
[0026] S1. Data Acquisition Steps
[0027] Data on the patients' age, sex, T stage, N stage, M stage, Grade level, Stage stage, and expression levels of CDCA8, MCM4, and TTK in liver cancer tissue were obtained.
[0028] S2. Data Input Steps
[0029] Input the data collected in step S1 into the data acquisition module of the above system;
[0030] S3. Prognostic Risk Assessment Steps
[0031] Using the prognostic risk discrimination module of the above system, a prognostic risk score is calculated based on the expression levels of CDCA8, MCM4 and TTK. A prognostic risk score ≥0.316006455458479 is considered high risk, and a prognostic risk score <0.316006455458479 is considered low risk.
[0032] S4. Prognostic assessment steps
[0033] Using the prognostic module of the above system, scores are assigned to the age, gender, T stage, N stage, M stage, Grade classification, Stage classification, and risk score, and the total score is calculated. The prognosis of the patient with primary liver cancer is determined based on the total score.
[0034] Preferably, the prognosis of patients with primary liver cancer refers to the 1-year survival probability, 3-year survival probability, and / or 5-year survival probability after the patient is diagnosed with primary liver cancer.
[0035] This application identifies three T-cell senescence-related genes—CDCA8, MCM4, and TTK—significantly associated with overall survival in patients with primary liver cancer. Based on the expression levels of these three genes, a prognostic risk score is calculated to determine the patient's prognostic risk (high / low). Furthermore, by combining the patient's age, sex, T stage, N stage, M stage, Grade classification, and Stage classification, a visual nomogram model predicting the prognosis (1-year, 3-year, and / or 5-year survival probability after diagnosis) of primary liver cancer patients is constructed. The prediction system and method disclosed in this application help physicians make a more in-depth and accurate prediction of the disease progression in patients with primary liver cancer, thus providing a useful reference for clinical treatment decisions. Attached Figure Description
[0036] The invention will now be further described with reference to the accompanying drawings.
[0037] Figure 1 The figure shown is the cumulative distribution function (CDF) obtained when performing consistent clustering on the 370 patients with primary liver cancer included in Study Example 1.
[0038] Figure 2 The diagram shown is the Delta plot obtained from the consistent clustering of 370 patients with primary liver cancer included in Study Example 1.
[0039] Figure 3 The diagram shown is a cluster consistency map obtained when performing consistent clustering on 370 patients with primary liver cancer included in Study Example 1.
[0040] Figure 4 The diagram shows a cluster heatmap of 370 patients with primary liver cancer included in Study Example 1, divided into two subtypes: C1 and C2.
[0041] Figure 5 The results shown are from PCA (Principal Component Analysis) dimensionality reduction clustering analysis performed on 370 patients with primary liver cancer included in Study Example 1, who were divided into two subtypes, C1 and C2.
[0042] Figure 6 The figure shows the Kaplan-Meier (KM) survival curves for the 370 patients with primary liver cancer included in Study Example 1, divided into two subtypes: C1 and C2.
[0043] Figure 7 The illustration shows the gene expression differences between the 370 patients with primary liver cancer included in Study Example 1, who were divided into two subtypes, C1 and C2.
[0044] Figure 8 The results shown are from the KEGG pathway enrichment analysis performed on differentially expressed genes (DEGs) identified in Study Example 1.
[0045] Figure 9 The results shown are from GSEA analysis of differentially expressed genes (DEGs) identified in Study Example 1.
[0046] Figure 10 The figure shows the expression levels of genes related to the senescence-associated secretory phenotype (SASP) in subtypes C1 and C2 in Study Example 1.
[0047] Figure 11 The diagram shows the network topology analysis results of differentially expressed genes in subtypes C1 and C2 using the R software package in Example 1, under different soft threshold powers. A shows a non-linear positive correlation between the scale-free fit index (y-axis) and the soft threshold power (x-axis); B shows a non-linear negative correlation between average connectivity (y-axis) and the soft threshold power (x-axis).
[0048] Figure 12 The diagram shown is a cluster diagram of 33 gene modules obtained by using the R software package to cluster differentially expressed genes of C1 and C2 subtypes in Study Example 1.
[0049] Figure 13 The diagram shown is a correlation plot between the clustering modules of 33 differentially expressed genes and the clinical feature OS in Study Example 1.
[0050] Figure 14 The results shown are the KEGG pathway enrichment analysis results of genes within the tuequoise module of differentially expressed gene clustering in Study Example 1.
[0051] Figure 15 The diagram shown is the lasso-cox regression cross-validation plot of the training set based on the selected Hub genes in Study Example 1.
[0052] Figure 16 The diagram shown is the lasso-cox regression path of the training set based on the selected Hub genes in Study Example 1.
[0053] Figure 17The example shown is from Study Example 1, where the three modeling genes CDCA8, MCM4, and TTK were analyzed in the GSE169009 dataset of healthy individuals with high aging CD8 levels. + T cells and low-aging CD8 + Results of differential expression analysis in T cells.
[0054] Figure 18 The example shown is from Study Example 1, where the three modeling genes CDCA8, MCM4, and TTK were analyzed in the GSE169009 dataset of healthy individuals with high aging CD8 levels. + Results of differential expression analysis of T cells between young and old individuals.
[0055] Figure 19 The results shown are from Study Example 1, which analyzed the differential expression of three modeling genes, CDCA8, MCM4, and TTK, in the GSE62232 dataset between young and elderly patients with primary HCC.
[0056] Figure 20 The diagram shows the differential expression of the three modeling genes, CDCA8, MCM4, and TTK, in patients with subtypes C1 and C2 in Study Example 1. A: Differential expression of CDCA8 in the two subtypes; B: Differential expression of MCM4 in the two subtypes; C: Differential expression of TTK in the two subtypes.
[0057] Figure 21 The figures shown are the KM survival curves for the high and low expression groups of the three modeling genes CDCA8, MCM4, and TTK in Study Example 1. A: KM survival curves for the high and low expression groups of CDCA8; B: KM survival curves for the high and low expression groups of MCM4; C: KM survival curves for the high and low expression groups of TTK.
[0058] Figure 22 The diagram shows the differential expression of three modeling genes, CDCA8, MCM4, and TTK, in various cancer tissues, including liver cancer, and normal tissues from the Timer database, in Example 1. A: Expression of CDCA8 in various cancer tissues and normal tissues; B: Expression of MCM4 in various cancer tissues and normal tissues; C: Expression of TTK in various cancer tissues and normal tissues.
[0059] Figure 23 The image shows the immunohistochemical results of three modeling genes, CDCA8, MCM4, and TTK, in liver cancer tumor tissues and normal tissues from the HPA database in Study Example 1. A: Expression of CDCA8 in normal tissue (left) and cancer tissue (right); B: Expression of MCM4 in normal tissue (left) and cancer tissue (right); C: Expression of TTK in normal tissue (left) and cancer tissue (right).
[0060] Figure 24 The image shows the immunohistochemical results of the three modeling genes, CDCA8, MCM4, and TTK, in liver cancer tissue and normal tissue in Study Example 1. A: Immunofluorescence images show the differences in expression of CDCA8, MCM4, and TTK between normal liver tissue and liver tumor tissue from liver cancer patients; B: Bar charts show the statistical results of the differences in expression of CDCA8, MCM4, and TTK between normal liver tissue and liver tumor tissue from liver cancer patients.
[0061] Figure 25 The diagram shown is a triptych of prognostic risk scores, clinical events (death), and the expression of three characteristic genes in the training set, which were divided into high-risk and low-risk groups based on risk scores, in Study Example 2.
[0062] Figure 26 The figure shown is the biological KM survival curve for patients in the high-risk and low-risk groups of the training set in Study Example 2.
[0063] Figure 27 The figure shown is the biological KM survival curve for patients in the high-risk and low-risk groups in Study Example 2.
[0064] Figure 28 The figure shows the ROC curves for 1-year, 3-year, and 5-year predictions from the training set patients in Study Example 2.
[0065] Figure 29 The figure shows the ROC curves for 1-year, 3-year, and 5-year predictions for patients in the validation set in Study Example 2.
[0066] Figure 30 The results shown are from a Cox univariate regression analysis of Study Example 3, which included multiple clinical features and prognostic risk scores.
[0067] Figure 31 The results shown are from a Cox multivariate regression analysis of Study Example 3, which included multiple clinical features and prognostic risk scores.
[0068] Figure 32 The image shown is a visualization of the column line graph created in Study Example 3.
[0069] Figure 33 The diagram shows the survival analysis KM curves of high- and low-risk patients selected based on the nomogram model in Study Example 3.
[0070] Figure 34 The figure shown is the time-dependent ROC curve of the nomogram model in Study Example 3.
[0071] Figure 35 The figure shown is the calibration curve of the nomogram model in Study Example 3. Detailed Implementation
[0072] The present invention will be described below with reference to research examples and specific embodiments. Those skilled in the art will understand that these research examples and embodiments are for illustrative purposes only and do not limit the scope of the invention in any way.
[0073] Unless otherwise specified, the experimental methods used in the following research examples and embodiments are conventional methods. Unless otherwise specified, the medicinal materials and reagents used in the following embodiments are commercially available products.
[0074] Study Example 1: A retrospective study of hepatocellular carcinoma patients using the TCGA-LIHC public database.
[0075] 1.1 Research Subjects
[0076] RNA-seq data and clinical information of liver cancer patients were obtained from the Liver hepatocellular carcinoma (LIHC) dataset from the TCGA public database (https: / / www.cancer.gov / ). After excluding patients with secondary liver cancer, data from 370 patients with primary liver cancer were included.
[0077] 1.2 Clustering
[0078] Unbiased and unsupervised consensus clustering was performed on 370 patients with primary liver cancer using the "ConsensusClusterPlus" R package. The sampling proportion was set to 0.8, the number of samples was set to 10, and the maximum number of clusters was set to 10. Subsequently, clustering was performed based on the cumulative distribution function (CDF), Delta area plot, and cluster consensus. Figure 3 The optimal number of clusters was determined using several parameters, and PCA (Principal Component Analysis) dimensionality reduction clustering analysis was performed using the R package stats (v.3.6.0) to further verify the discriminative power between the optimal clusters. The R package survival (v.3.5.5) was used to analyze the prognostic differences between groups and to plot Kaplan-Meier (KM) survival curves.
[0079] CDF diagram (see) Figure 1 The data shows that when k=2, i.e., when divided into two subtypes, the CDF descent slope is the smallest. (See Delta area plot). Figure 2 ) and cluster consistency graph (see Figure 3 The results also show that the optimal clustering outcome is achieved when divided into two subtypes. Therefore, the 370 samples were divided into two subtypes, C1 (N=178) and C2 (N=192). Clustering heatmaps for subtypes C1 and C2 are shown below. Figure 4As shown. PCA analysis results show good differentiation between the two subtypes C1 and C2 (see...). Figure 5 The KM curves showed that patients with the C2 subtype had shorter overall survival (OS) than those with the C1 subtype (see...). Figure 6 ).
[0080] 1.3 The underlying molecular mechanisms of OS differences between C1 and C2 subtypes
[0081] Differential expression analysis was performed on RNA-seq data from C1 and C2 subtype patients using the R package limma (v.3.4.6) to identify differentially expressed genes between the two subtypes. Genes with FDR < 0.05 and |log2FC| ≥ 1.5 were selected as differentially expressed genes (DEGs). Subsequently, KEGG (Kyoto Encyclopedia of Genes and Genomes) pathway enrichment analysis and GSEA (Gene Set Enrichment Analysis) analysis were performed on the differentially expressed genes using the R package clusterProfiler (v.3.14.3). The minimum gene set was set to 5, the maximum gene set to 5000, and gene sets with p < 0.05 were considered statistically significant.
[0082] RNA-seq differential expression analysis showed that, compared with the C2 subtype, 3061 genes were downregulated and 1742 genes were upregulated in the C1 subtype (see...). Figure 7 These genes were identified as DEGs. KEGG analysis of these 4083 DEGs showed that they were significantly enriched in signaling pathways such as metabolic pathways (hsa01100), cell cycle (hsa04110), p53 signaling pathway (hsa04115), and cellular senescence (hsa04218) (see [link to KEGG analysis]). Figure 8 Interestingly, GSEA analysis showed that cellular senescence and the p53 signaling pathway were significantly activated in the C2 isotype (see...). Figure 9 Furthermore, genes regulating the senescence-associated secretory phenotype (SASP) are expressed at higher levels in the C2 subtype (see [link to article]). Figure 10 The above results suggest that cellular senescence and overactivation of the p53 signaling pathway may be associated with poorer overall survival (OS).
[0083] 1.4 Screening differentially expressed gene modules most associated with worse OS and constructing a prognostic risk scoring model
[0084] We used the R package to perform Weighted Gene Co-Expression Network Analysis (WGCNA) to identify the DEGs most strongly associated with overall survival (OS) in patients with primary liver cancer. First, we determined the optimal soft thresholding power for network construction using scale independence analysis. Then, based on this power value, we calculated the topological overlap matrix between genes and used dynamic tree cutting to identify modules with the following parameters: mergeCutHeight = 0.25 and minModuleSize = 30. Modules with high similarity were automatically merged. To identify the modules most relevant to OS, we calculated the correlation between the eigenvector genes of each module and the clinical characteristics of patient OS. The module with the strongest negative correlation to OS was selected as the key module, and the top 10 genes with the highest connectivity within this module were considered as hub genes. Finally, we used the R package clusterProfiler (v. 3.14.3) to perform KEGG pathway enrichment analysis on all genes within the key module to reveal their potential biological functions.
[0085] Analysis showed that when the power value was 4, the scale-free fit exponent R0 was... 2 =0.9, the connectivity between genes satisfies the scale-free network degree (see Figure 11After clustering similar genes into modules and merging highly correlated modules, a total of 33 gene populations were formed, which were distinguished by different colors: dark orange, yellowgreen, saddlebrown, light cyan, paleurquoise, purple, violet, steelblue, grey60, royalblue, darkmagenta, darkgreen, and green. Yellow, salmon, dark grey, sky blue, orange, light yellow, pink, blue, cyan, turquoise, light green, white, dark red, red, dark olive green, dark turquoise, sienna3, green, black, midnight blue, and grey, where the grey module is the set of genes that cannot be assigned to any module (see...). Figure 12 Correlation analysis between gene clustering modules and clinical characteristics (OS) showed that the turquoise module was most negatively correlated with OS (correlation -0.2, p=7.3E-5, see...). Figure 13 This indicates that gene abundance within this module is significantly associated with shorter OS. KEGG pathway enrichment analysis of the 697 genes contained in the turquoise module showed that the genes were mainly enriched in the cell cycle, p53 signaling pathway, and cellular senescence pathway (see...). Figure 14 ); This is consistent with the analysis results in Section 1.3.
[0086] Weighted Gene Co-Expression Network Analysis (WGCNA) was performed on the 697 genes contained in the turquoise module in conjunction with the patient's overall survival (OS). The top 10 genes with the highest connectivity within the module were obtained, as shown in Table 1. These 10 genes were considered as hub genes and used to screen for key genes most associated with worse OS.
[0087]
[0088] Patients with a survival time of 0 were removed, and 365 patients with primary liver cancer obtained from the TCGA-LIHC database were divided into a training set (n=256) and a validation set (n=109) in a 7:3 ratio. The clinical information of the two groups of patients was balanced and comparable, as shown in Table 2.
[0089]
[0090] Integrating the data on the 10 hub genes, survival time, survival status, and gene expression selected in the previous section, lasso-cox regression was performed on the training set, and the optimal parameters were determined through 10-fold cross-validation. The results are shown in [the table below]. Figure 15 and Figure 16 .
[0091] Figure 15 The Lasso-Cox regression cross-validation plot is shown. Figure 16 The lasso-cox regression path diagram is shown. Figure 15 The parameter λmin corresponding to the dashed line on the left side of the lasso-cox regression cross-validation plot has a value of 3, representing the optimal λ value for the fit. This indicates that three genes are closely related to the overall survival of liver cancer patients. These genes are CDCA8, MCM4, and TTK. CDCA8 is an important cell cycle regulatory gene and a key regulator of mitosis and cell division. MCM4 is a key regulator of DNA replication, and its dysregulation can lead to DNA damage. TTK is associated with cell proliferation, and its related pathways include G1 phase and G1 / S transition in mitosis, as well as DNA damage. One of the important mechanisms of cellular senescence is oxidative stress and DNA damage, which leads to an abnormal cell cycle state where cells enter permanent growth arrest. Therefore, these three key genes play an important role in regulating the cell cycle, are highly correlated with DNA damage, and are strongly associated with the occurrence and development of cellular senescence.
[0092] Based on three key features, the following formula (1) was constructed to calculate the prognostic risk score for patients with primary liver cancer:
[0093]
[0094] Wherein, [CDCA8], [MCM4], and [TTK] represent the expression levels of the aforementioned genes. In this application, the expression level is the FPKM value.
[0095] 1.5 The relationship between CDCA8, MCM4, and TTK and T cell senescence
[0096] Analysis of three modeling genes (CDCA8, MCM4, and TTK) in the aging CD8+ of healthy humans based on the GSE169009 dataset. + Characteristic expression on T cells. Surprisingly, all three modeling genes are expressed on highly aging CD8 cells. + Significantly high expression on T cells (see...) Figure 17 Furthermore, this characteristic expression is independent of age (see...). Figure 18 Given the close association between T cell senescence and aging, we further analyzed the expression of these three modeling genes between young and elderly patients with primary HCC based on the GSE62232 dataset. The results still indicate that the characteristic expression of these three modeling genes in patients with primary liver cancer remains independent of age (see...). Figure 19 ).
[0097] The above research results strongly demonstrate that the high expression of the three genes CDCA8, MCM4, and TTK is closely related to T cell senescence; the risk scoring model constructed in this invention is a novel prognostic risk scoring model for hepatocellular carcinoma based on T cell senescence-related molecules.
[0098] 1.6 Validation of the relationship between CDCA8, MCM4, and TTK and the prognosis of primary liver cancer
[0099] To verify the relationship between the three modeling genes CDCA8, MCM4, and TTK and the prognosis of primary liver cancer, the expression of these three modeling genes was compared between C1 and C2 subtype patients. The results showed that CDCA8, MCM4, and TTK were significantly overexpressed in C2 subtype patients with a worse prognosis (see...). Figure 20 ).
[0100] The training and validation sets were combined into a single set. Patients were divided into high-expression and low-expression groups using the median expression level of the three modeling genes as the cut-off. KM survival curves were plotted for each high-expression and low-expression group. (See attached data.) Figure 21 . Figure 21 The results showed that patients with primary liver cancer who expressed high levels of CDCA8, MCM4, and TTK had worse overall survival (OS).
[0101] Gene differential expression analysis was performed using the Timer database. Results are shown below. Figure 22 . Figure 22 The results showed that the expression of the three modeling genes was significantly higher in tumor tissues of liver cancer patients than in normal tissues, and this phenomenon was also observed in a variety of cancers other than liver cancer, suggesting that these three genes have a certain degree of tumor specificity.
[0102] Immunohistochemical results of liver cancer tissues and normal tissues from the HPA database were compared. The results are shown in [link to results]. Figure 23 . Figure 23The results showed that the expression of the three modeling genes was significantly higher in hepatocellular carcinoma tissues than in normal liver tissues. This result is consistent with the results of RNA-seq data analysis using the Timer database.
[0103] 1.7 Experimental Validation of CDCA8, MCM4, and TTK
[0104] Multicolor immunofluorescence assays were performed according to the method reported by Liang Y et al. to compare the expression of CDCA8, MCM4, and TTK in primary hepatocellular carcinoma tissues and normal tissues (Liang Y, et al. Integrating Network Pharmacology and Experimental Validation to Decipher the Mechanism of Action of Astragalus–Atractylodes Herb Pair in Treating Hepatocellular Carcinoma. Drug Des Devel Ther. 2024 Jun 11;18:2169–87.).
[0105] Normal liver tissue and tumor tissue from patients with primary liver cancer were obtained from Beijing Ditan Hospital. Antibodies used were purchased from Wuhan Sanying Biotechnology Co., Ltd. and Wuhan Saiwei Biotechnology Co., Ltd. (CDCA8 catalog number: 12465-1-AP; MCM4 catalog number: 13043-1-AP; TTK catalog number: GB111371). After image acquisition, the positive staining rate of the acquired images was analyzed using ImageJ_v1.8.0 software (WayneRasband National Institutes of Health, USA). The results of multicolor immunofluorescence detection are shown below. Figure 24 .
[0106] like Figure 24 As shown, CDCA8, MCM4, and TTK were almost undetectable in normal liver tissue, but their expression was significantly increased in tumor tissue compared to normal liver tissue. This result is consistent with the RNA-seq data analysis results from the Timer database and the immunohistochemical results obtained from the HPA database in Section 1.6.
[0107] Therefore, the above results all suggest that the three selected modeling genes, CDCA8, MCM4 and TTK, are specific to liver cancer tissue.
[0108] Example 2: Evaluation of the prognostic risk scoring model for primary liver cancer constructed in Example 1
[0109] 1. Relationship between prognostic risk scores and clinical events, and characteristic gene expression.
[0110] The prognostic risk score for each patient in the training set was calculated. Using 0.316006455458479 as the cut-off value, the patients were divided into a high-risk group (prognostic risk score ≥ 0.316006455458479, n=111) and a low-risk group (prognostic risk score < 0.316006455458479, n=145). Tripartite plots of prognostic risk scores were plotted against clinical events (death) and the expression of three characteristic genes in both the high-risk and low-risk groups. (See below) Figure 25 . Figure 25 This indicates that mortality increases with increasing risk score and survival time decreases accordingly. The three specific key genes involved in the prognostic risk score of this invention are significant risk indicators.
[0111] 2. KM Survival Analysis Method
[0112] To evaluate the association between prognostic risk scores and patient survival time, the overall survival times of high-risk and low-risk patients in the training and validation sets were statistically analyzed, and KM curves were plotted. The results are shown below. Figure 26 and Figure 27 As shown.
[0113] Figure 26 and Figure 27 The results showed that patients in the high-risk group of both the training and validation sets had significantly lower overall survival (OS) than patients in the corresponding low-risk group (p=0.0024 in the training set and p=0.0015 in the validation set).
[0114] 3. ROC curve
[0115] Receiver operating characteristic (ROC) curves were plotted for the training and validation sets, respectively. The area under the ROC curve (AUC) was used to measure the predictive accuracy of the prognostic risk score. The results are shown in [Figure number missing]. Figure 28 and Figure 29 . Figure 28 As shown, in the training set, the AUC values for predicting 1-year, 3-year, and 5-year survival (OS) were 0.73 (95% CI 0.81–0.64), 0.72 (95% CI 0.80–0.63), and 0.77 (95% CI 0.87–0.67), respectively. Figure 29 As shown, in the validation set, the AUC values at the corresponding time points were 0.72 (95% CI 0.80-0.63), 0.70 (95% CI 0.78-0.61), and 0.75 (95% CI 0.85-0.65), respectively. The AUC values for 1 year, 3 years, and 5 years in both the training and validation sets were greater than 0.5, indicating that the prognostic risk score established in this invention has superior predictive power, and its long-term predictive power is better than its short-term predictive power.
[0116] Study Example 3: Establishment of a nomogram prognostic prediction model for patients with primary liver cancer
[0117] 1. Selection of Indicators for the Nomogram Prognostic Prediction Model
[0118] To construct a nomogram-based prognostic prediction model for patients with primary liver cancer, the prognostic risk score was combined with various clinical indicators based on the aforementioned study cases. Cox univariate and multivariate regression analyses were performed, incorporating age, gender, T stage, N stage, M stage, Grade, Stage, and prognostic risk score. Results are shown in [Figure number missing]. Figure 30 and Figure 31 .
[0119] Figure 30 and Figure 31 Studies have shown that only T stage and prognostic risk score are independent predictors of overall survival in patients with primary liver cancer (P<0.05). However, TNM stage and WHO classification are the global gold standard for tumor assessment. If the model lacks core elements such as N, M, and Grade, it will be out of touch with clinical practice and lose its usability. Secondly, the predictive model aims to accurately estimate risk, not just screen for independent factors. Retaining variables such as age, gender, NM stage, and Grade helps to fully correct for confounding effects and avoid bias in the estimation of coefficients of other significant factors (such as T stage) due to the removal of collinear variables, thereby improving the overall robustness and calibration ability of the model and maximizing predictive accuracy. Therefore, age, gender, T stage, N stage, M stage, Grade, Stage, and prognostic risk score are all included in the nomogram prognostic prediction model. 2. Construction of the nomogram prognostic prediction model
[0120] Patient age, gender, T stage, N stage, M stage, Grade, Stage, and prognostic risk score were imported into the open-source data analysis software R version 4.1.3 (http: / / www.rproject.org / ). Scores for each factor and the overall survival probabilities for 1-year, 3-year, and 5-year patients with primary liver cancer corresponding to the total score were obtained. Based on this, a nomogram visualization was created, including the scores for the eight factors, the total score axis, and the overall survival probability axes for 1-year, 3-year, and 5-year survival. (See [link to nomogram]). Figure 32 As shown.
[0121] The scoring rules for the 8 factors are as follows:
[0122] Age: ≤10 years old, score is 0; >10 years old, score is calculated according to formula (2):
[0123]
[0124] Where A represents the patient's age;
[0125] For example, according to this age-based scoring rule, a 30-year-old would score 3.45.
[0126] Gender: Females score 0, males score 1.06;
[0127] T-stage: T1 score is 0, T2 score is 52.27, T3 score is 68.27, T4 score is 71.27;
[0128] N-stage: N0 score is 0, N1 score is 12.15;
[0129] M stage: M0 score is 0, M1 score is 44.26;
[0130] Grade classification: G1 and G2 scores are 0, G3 scores are 0.53, and G4 scores are 7.11;
[0131] Stage classification: Stage I score is 0, Stage II score is 38.72, Stage III score is 45.41, and Stage IV score is 100;
[0132] Prognostic risk: Low risk score is 0, high risk score is 7.92. (For example...) Figure 32 As shown, when using the constructed visual nomogram prognostic prediction model, based on the patient's various clinicopathological characteristics, the corresponding value point is found on each variable axis in the graph, and then projected vertically upwards to the top 'score' axis to read the score obtained for each variable. Then, the scores of all variables are summed to obtain the total score. Finally, this total score is located on the bottom 'total score' axis and projected vertically downwards to read the predicted probability of the patient experiencing the target event on the corresponding 'probability' axis (corresponding to different time points such as 365 days (1 year), 1095 days (3 years), and 1825 days (5 years).
[0133] It should be noted that although the prognostic risk score is an independent prognostic factor, its contribution weight in the nomogram prognostic prediction model is lower than that of the T stage and the Stage stage. This suggests that, based on the existing clinical staging system, although this score is not yet sufficient to replace traditional core indicators such as the TNM stage, it can provide limited but stable incremental prognostic information and can be used for refined risk stratification.
[0134] 3. Evaluation of the nomogram prognostic prediction model
[0135] Study Example 1 selected 365 patients with primary liver cancer from the TCGA-LIHC database. A pre-constructed nomogram prognostic prediction model was used to predict the 5-year survival probability of these patients. Patients were divided into high-risk and low-risk groups based on their survival probability (≥50%), and survival analysis was performed to plot KM curves. The time-dependent ROC curve of the nomogram prognostic prediction model was plotted, and the AUC values at 1, 3, and 5 years were calculated. A calibration curve for the nomogram prognostic prediction model was also plotted. The plotted KM curves, ROC curves, and calibration curves are shown below. Figure 33 , Figure 34 and Figure 35 As shown.
[0136] Figure 33 As shown, the nomogram prognostic prediction model constructed using the present invention can better distinguish between high-risk and low-risk patients, and the overall survival (OS) of high-risk patients is significantly worse. Figure 34 As shown, the areas under the curves (AUCs) for predicting 1-year, 3-year, and 5-year survival rates for patients with primary liver cancer were 0.83 (95% CI 0.89-0.76), 0.77 (95% CI 0.85-0.69), and 0.78 (95% CI 0.87-0.69), respectively, all greater than 0.5. This indicates that the nomogram prognostic prediction model constructed in this invention can accurately predict the 1-year, 3-year, and 5-year survival probabilities of patients.
[0137] Figure 35 The calibration curve passes through the origin and has a slope close to 1. The closer the survival probability predicted by the model is to the actual observed survival probability, the closer the slope of the calibration curve is to 1. Therefore, the nomogram prognostic prediction model constructed in this invention can accurately predict the prognosis of patients with primary liver cancer.
[0138] In summary, this application identified three T-cell senescence-related genes—CDCA8, MCM4, and TTK—significantly associated with overall survival in patients with primary liver cancer (HCC) through a retrospective study. Based on the expression levels of these three genes, a prognostic risk score was calculated to determine the patient's prognostic risk (high / low). Furthermore, by combining the patient's age, sex, T stage, N stage, M stage, Grade classification, and Stage classification, a visual nomogram model was constructed to predict the prognosis (1-year, 3-year, and / or 5-year survival probability after diagnosis) of HCC patients. The prediction system and method disclosed in this application help physicians make a more in-depth and accurate prediction of the disease progression in HCC patients, thus providing a useful reference for clinical treatment decisions.
Claims
1. A system for predicting the prognosis of patients with primary liver cancer based on the molecular characteristics of senescent T cells, the system being based on the expression levels of senescence-related genes in T cells removed from the patient's liver cancer tissue, comprising: Data collection module: set to acquire data on the patient's age, sex, T stage, N stage, M stage, Grade level, Stage stage, and expression levels of genes CDCA8, MCM4, and TTK in liver cancer tissue; The prognostic risk assessment module is configured to substitute the expression levels of the genes CDCA8, MCM4, and TTK obtained by the data collection module into formula (1) to calculate the prognostic risk score of the patient. The prognostic risk is determined based on the calculated risk score, where a prognostic risk score ≥ 0.316006455458479 is considered high risk, and a prognostic risk score < 0.316006455458479 is considered low risk. Among them, [CDCA8], [MCM4], and [TTK] represent the expression levels of the corresponding genes; Prognostic module: This module is configured to use an established nomogram prognostic prediction model for primary liver cancer patients to assign scores to age, sex, T, N, M, Grade, Stage, and risk score, calculate a total score, and determine the prognostic outcome based on the total score. The scoring rules in the nomogram prognostic prediction model for primary liver cancer patients are as follows: Age: ≤10 years old, score is 0; >10 years old, score is calculated according to formula (2): Where A represents the patient's age; Gender: Females score 0, males score 1.06; T-stages: T1 score is 0, T2 score is 52.27, T3 score is 68.27, and T4 score is 71.27; N-stage: N0 score is 0, N1 score is 12.15; M stage: M0 score is 0, M1 score is 44.26; Grade classification: G1 and G2 scores are 0, G3 scores are 0.53, and G4 scores are 7.11; Stage classification: Stage I score is 0, Stage II score is 38.72, Stage III score is 45.41, and Stage IV score is 100; Prognostic risk: Low risk score is 0, high risk score is 7.
92.
2. The system according to claim 1, characterized in that, The [CDCA8], [MCM4], and [TTK] values are the FPKM values of the corresponding genes.
3. The system according to claim 1, characterized in that, The prognosis refers to the 1-year survival probability, 3-year survival probability, and / or 5-year survival probability after a patient is diagnosed with primary liver cancer.
4. A method for predicting the prognosis of patients with primary liver cancer; the method is based on the system for predicting the prognosis of patients with primary liver cancer based on the molecular characteristics of senescent T cells as described in any one of claims 1 to 3, and includes the following steps: S1. Data Acquisition Steps Data on the patients' age, sex, T stage, N stage, M stage, Grade level, Stage stage, and expression levels of CDCA8, MCM4, and TTK in liver cancer tissue were obtained. S2. Data Input Steps Input the data collected in step S1 into the data acquisition module of the above system; S3. Prognostic Risk Assessment Steps Using the prognostic risk discrimination module of the above system, a prognostic risk score is calculated based on the expression levels of CDCA8, MCM4 and TTK. A prognostic risk score ≥0.316006455458479 is considered high risk, and a prognostic risk score <0.316006455458479 is considered low risk. S4. Prognostic assessment steps Using the prognostic module of the above system, scores are assigned to the age, gender, T stage, N stage, M stage, Grade classification, Stage classification, and risk score, and the total score is calculated. The prognosis of the patient with primary liver cancer is determined based on the total score.
5. The method according to claim 4, characterized in that, The prognosis of patients with primary liver cancer refers to their 1-year survival probability, 3-year survival probability, and / or 5-year survival probability after being diagnosed with primary liver cancer.
Citation Information
Patent Citations
System for predicting prognosis of non-surgical treatment primary liver cancer patient mainly based on platelet / spleen length-diameter ratio
CN115985497A
Improvements in the Manufacture of Annealing Pots.
GB111371A