A liver cancer dual-pathway gene prognosis evaluation method, system and related device
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-20
- Publication Date
- 2026-08-11
AI Technical Summary
[0006]本发明的目的是提供一种肝癌双通路基因预后评估方法、系统及相关设备,旨在解决现有肝细胞癌预后评估中单一分子特征建模稳定性不足、高维基因变量相关性影响风险分层结果可靠性的技术问题
[0052] This invention maps a set of genes related to glycosylation modification and a set of genes related to ferroptosis regulation to a transcriptome expression matrix of hepatocellular carcinoma, forming a joint candidate gene expression matrix. This allows prognostic assessment to move beyond being limited to a single molecular feature source and to simultaneously incorporate information on tumor metabolic regulation, cell death regulation, and related transcriptional expression changes within the same data processing framework. This improves the matching degree between candidate feature sources and the molecular heterogeneity of hepatocellular carcinoma.
Smart Images

Figure CN122551877A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of tumor molecular diagnostics and biomedical information analysis technology, specifically to a method, system, and related equipment for prognostic assessment of liver cancer using dual-pathway genes. Background Technology
[0002] Hepatocellular carcinoma (HCC) is one of the most common malignant tumors in clinical practice. Its development and progression are influenced by a variety of factors, including genetic variations, transcriptional regulation, metabolic reprogramming, and the tumor microenvironment. Current prognostic assessment methods largely rely on pathological staging, imaging findings, serological markers, and routine clinical characteristics. While these indicators can reflect tumor burden and patient baseline status to some extent, their ability to characterize tumor heterogeneity at the molecular level, differences in disease progression, and inter-individual survival outcomes remains limited. With the development of high-throughput sequencing technology and public biomedical databases, constructing prognostic assessment models based on transcriptome data has become an important direction for assisting in tumor risk stratification.
[0003] Existing gene expression analysis methods typically focus on screening and modeling genes with single-type molecular characteristics or single-pathway related genes, making them susceptible to gene expression noise, differences in sample sources, and correlations of high-dimensional variables. When there are a large number of genes to be analyzed and there are co-expression relationships or functional associations among them, traditional statistical screening methods struggle to balance feature stability, model simplicity, and prognostic interpretability, potentially leading to insufficient repeatability of screening results and decreased model generalization ability, thereby affecting the application performance in different data cohorts.
[0004] Therefore, in the field of hepatocellular carcinoma prognostic assessment, there is still a need for a technical solution that can jointly analyze existing transcriptome expression data and clinical follow-up data. This solution would enable a continuous data processing flow between candidate molecular feature screening, key feature compression, risk score construction, and external validation, thereby improving the stability and verifiability of prognostic risk stratification results and providing auxiliary evidence for individualized prognostic assessment of hepatocellular carcinoma patients.
[0005] In view of this, the present invention proposes a method, system and related equipment for prognostic assessment of liver cancer through dual-pathway genes. Summary of the Invention
[0006] The purpose of this invention is to provide a method, system, and related equipment for prognostic assessment of liver cancer using dual-pathway genes, aiming to solve the technical problems of insufficient stability of single molecular feature modeling and the impact of high-dimensional gene variable correlation on the reliability of risk stratification results in existing prognostic assessments of hepatocellular carcinoma.
[0007] In a first aspect, the present invention provides a method for prognostic assessment of hepatocellular carcinoma based on glycosylation and ferroptosis-related genes, comprising the following steps:
[0008] Transcriptome expression matrix, tumor phenotype information and clinical follow-up data of hepatocellular carcinoma samples were obtained, and glycosylation modification-related gene set and ferroptosis regulation-related gene set were obtained. The glycosylation modification-related gene set and the ferroptosis regulation-related gene set were mapped to the transcriptome expression matrix to form a joint candidate gene expression matrix.
[0009] Based on the joint candidate gene expression matrix, a weighted gene co-expression network is constructed, and the module correlation between each gene module and the tumor phenotype is calculated. Gene modules that meet the module correlation screening criteria are identified as key co-expression modules.
[0010] Differential expression analysis was performed on genes in the key co-expression module between tumor tissue and adjacent normal tissue to obtain a candidate prognostic gene set that simultaneously meets the module correlation condition and the differential expression condition;
[0011] Based on the candidate prognostic gene set and clinical follow-up data, survival-related genes were screened, and a Cox regression model with regularization penalty term was used to perform dimensionality reduction modeling of survival-related genes, determine the core gene combination and its corresponding regression coefficient, and construct a risk scoring model.
[0012] The risk score of the hepatocellular carcinoma sample to be evaluated is calculated based on the risk scoring model, and the risk stratification results and corresponding prognostic assessment results are output according to the preset cutoff value.
[0013] As a preferred embodiment of the present invention, the formation of the joint candidate gene expression matrix includes:
[0014] Transcriptome expression data were obtained from hepatocellular carcinoma tumor tissue samples, adjacent normal tissue samples, and hepatocellular carcinoma samples with clinical follow-up data.
[0015] The transcriptome expression data were subjected to sample matching, gene identification unification, and expression level standardization to form a standardized expression matrix;
[0016] Gene identifiers from the glycosylation modification-related gene set and the ferroptosis regulation-related gene set were matched with gene identifiers in the normalized expression matrix;
[0017] Glycosylation-related genes and ferroptosis-regulating genes with expression levels recorded in the standardized expression matrix are retained to form a joint candidate gene expression matrix.
[0018] As a preferred embodiment of the present invention, the determination of the key co-expression module includes:
[0019] The sample clustering results are calculated based on the joint candidate gene expression matrix, and samples that do not meet the sample clustering quality requirements are removed based on the sample clustering results.
[0020] Iterate through multiple soft threshold parameters and calculate the scale-free topology fit index and average connectivity under different soft threshold parameters;
[0021] A weighted adjacency matrix is constructed based on a soft threshold parameter that allows the scale-free topology fitting index to meet a preset fitting condition and the average connectivity to be in a stable range.
[0022] The weighted adjacency matrix is converted into a topological overlap matrix, and gene module partitioning is performed based on the topological overlap matrix;
[0023] Calculate the module characteristic genes of each gene module, and determine the key co-expression modules based on the correlation coefficient between module characteristic genes and tumor phenotype and the results of statistical tests.
[0024] As a preferred embodiment of the present invention, obtaining the candidate prognostic gene set includes:
[0025] Extract gene expression data from the key co-expression module;
[0026] Based on the expression differences between tumor tissue samples and adjacent normal tissue samples, the fold change in expression of each gene in the key co-expression module and the statistical test value were calculated.
[0027] Genes whose expression fold change meets the preset fold change condition and whose statistical test value meets the preset significance condition are identified as differentially expressed genes;
[0028] The differentially expressed genes were used as a candidate prognostic gene set for survival-related screening and risk scoring model construction.
[0029] As a preferred embodiment of the present invention, the construction of the risk scoring model includes:
[0030] Hepatocellular carcinoma samples with complete clinical follow-up data were divided into training and testing sets.
[0031] In the training set, with total survival and survival status as survival outcome variables, univariate Cox regression analysis was performed on each gene in the candidate prognostic gene set to obtain survival-related genes.
[0032] LASSO-Cox regression analysis was performed on the survival-related genes, and the penalty parameters were determined by cross-validation.
[0033] Genes with non-zero regression coefficients under the aforementioned penalty parameters were selected as the core gene combination;
[0034] A risk scoring model is constructed based on the standardized expression levels and corresponding regression coefficients of each gene in the core gene combination.
[0035] The core gene combination includes B4GALNT1, EZH2, KIF20A, MYCN, ST6GALNAC4, CBS, G6PD, SQSTM1, B3GNTL1, and NQO1.
[0036] As a preferred embodiment of the present invention, the risk scoring model is as follows:
[0037] Risk Score = 0.312×Exp(B4GALNT1) + 0.204×Exp(EZH2) + 0.026×Exp(KIF20A)+0.030×Exp(MYCN)+0.026×Exp(ST6GALNAC4)−0.027×Exp(CBS)+0.179×Exp(G6PD)+0.097×Exp(SQSTM1) + 0.072×Exp(B3GNTL1) + 0.059×Exp(NQO1);
[0038] Where Exp(B4GALNT1), Exp(EZH2), Exp(KIF20A), Exp(MYCN), Exp(ST6GALNAC4), Exp(CBS), Exp(G6PD), Exp(SQSTM1), Exp(B3GNTL1), and Exp(NQO1) represent the normalized expression levels of the corresponding genes in the samples to be evaluated.
[0039] After calculating the risk score of the hepatocellular carcinoma sample to be evaluated based on the risk scoring model, the risk score is compared with the cutoff value determined by the training set, and the risk stratification results of high-risk group or low-risk group are output, and the corresponding prognostic assessment results are generated.
[0040] In a second aspect, the present invention provides a hepatocellular carcinoma prognostic assessment system based on glycosylation and ferroptosis-related genes, which, based on the implementation of the first aspect, includes:
[0041] The data acquisition module is used to acquire the transcriptome expression matrix, tumor phenotype information, clinical follow-up data, glycosylation modification-related gene set and ferroptosis regulation-related gene set of hepatocellular carcinoma samples, and map the glycosylation modification-related gene set and the ferroptosis regulation-related gene set to the transcriptome expression matrix to form a joint candidate gene expression matrix.
[0042] The co-expression module screening module is used to construct a weighted gene co-expression network based on the joint candidate gene expression matrix, calculate the module correlation between each gene module and the tumor phenotype, and identify key co-expression modules.
[0043] The differential expression screening module is used to perform differential expression analysis of genes in the key co-expression module between tumor tissue and adjacent normal tissue to obtain a candidate prognostic gene set.
[0044] The risk model construction module is used to screen for survival-related genes based on the candidate prognostic gene set and clinical follow-up data, and to perform dimensionality reduction modeling of survival-related genes using a Cox regression model with regularization penalty term, determine the core gene combination and its corresponding regression coefficient, and construct a risk scoring model.
[0045] The prognostic result output module is used to calculate the risk score of the hepatocellular carcinoma sample to be evaluated based on the risk scoring model, and output the risk stratification result and the corresponding prognostic assessment result according to the preset cutoff value.
[0046] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the first aspect.
[0047] Fourthly, the present invention provides an electronic device, comprising:
[0048] processor;
[0049] and a memory for storing executable instructions of the processor; wherein the processor is configured to perform the first aspect by executing the executable instructions.
[0050] Fifthly, the present invention provides the application of the first aspect in the preparation of a hepatocellular carcinoma prognostic assessment kit.
[0051] The technical effects and advantages provided by the present invention in the above technical solution are as follows:
[0052] This invention maps a set of genes related to glycosylation modification and a set of genes related to ferroptosis regulation to a transcriptome expression matrix of hepatocellular carcinoma, forming a joint candidate gene expression matrix. This allows prognostic assessment to move beyond being limited to a single molecular feature source and to simultaneously incorporate information on tumor metabolic regulation, cell death regulation, and related transcriptional expression changes within the same data processing framework. This improves the matching degree between candidate feature sources and the molecular heterogeneity of hepatocellular carcinoma.
[0053] This invention constructs a weighted gene co-expression network based on a joint candidate gene expression matrix, identifies key co-expression modules by the module correlation between gene modules and tumor phenotypes, and then performs differential expression screening on genes within key co-expression modules. This allows candidate prognostic genes to be constrained by both co-expression structural relationships and tumor tissue expression differences, reducing the accidental screening results caused by relying solely on differential expression analysis and improving the stability of the candidate prognostic gene set in different sample cohorts.
[0054] This invention screens candidate prognostic gene sets for survival relevance and uses a Cox regression model with regularization penalty to determine core gene combinations and their regression coefficients, constructing a risk scoring model. Based on the risk score, it outputs risk stratification results and prognostic assessment results, enabling high-dimensional transcriptome features to be converted into a computable and verifiable risk score form. This reduces the impact of multi-gene variable correlation on model construction and provides a consistent data processing basis for prognostic risk stratification of hepatocellular carcinoma patients. Attached Figure Description
[0055] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0056] Figure 1 This is an overall flowchart of the hepatocellular carcinoma prognostic assessment method of the present invention;
[0057] Figure 2A The graph shows the results of the correlation analysis between module characteristic genes and tumor phenotypes;
[0058] Figure 2B The cross-validation curves in the LASSO-Cox regression are shown.
[0059] Figure 2C The diagram shows the path of the regression coefficients of each candidate gene as a function of the penalty parameter.
[0060] Figure 3A Survival curves for the high-risk and low-risk groups in the training set;
[0061] Figure 3B Survival curves for the high-risk and low-risk groups in the test set;
[0062] Figure 3C Survival curves for the high-risk and low-risk groups in the GEO external validation set;
[0063] Figure 4A The ROC curve of the training set at preset time points;
[0064] Figure 4B ROC curves for the test set at preset time points;
[0065] Figure 4C The ROC curve of the GEO external validation set at the preset time points;
[0066] Figure 5 This is a schematic diagram of the nodal plot prediction model of the present invention. Detailed Implementation
[0067] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions in the embodiments of this application will be described in more detail below with reference to the accompanying drawings.
[0068] Throughout the accompanying drawings, the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions. The described embodiments are only a part of the embodiments of this application, not all of them. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application. The embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0069] Example 1
[0070] like Figure 1 As shown, this embodiment provides a prognostic assessment method for hepatocellular carcinoma based on glycosylation and ferroptosis-related genes. This method calculates prognostic risk scores and outputs risk stratification for hepatocellular carcinoma samples based on existing transcriptome expression data and clinical follow-up data. It can run on general-purpose computers, servers, or cloud computing platforms. Input data includes the transcriptome expression matrix of the hepatocellular carcinoma sample, tumor phenotype information, and clinical follow-up data. Output results include risk scores, risk stratification results, predictive efficacy evaluation results, and nomogram prediction results. The method includes the following steps:
[0071] S101. Obtain the transcriptome expression matrix, tumor phenotype information and clinical follow-up data of hepatocellular carcinoma samples, and obtain the gene set related to glycosylation modification and the gene set related to ferroptosis regulation. Map the gene set related to glycosylation modification and the gene set related to ferroptosis regulation to the transcriptome expression matrix to form a joint candidate gene expression matrix.
[0072] In this embodiment, transcriptome sequencing data from the LIHC project was downloaded from the TCGA database. This transcriptome sequencing data is RNA-seq level 3 data, including 374 hepatocellular carcinoma tumor tissue samples and 50 adjacent normal tissue samples. Simultaneously, corresponding clinical follow-up information was obtained, including overall survival, survival status, TNM stage, vascular invasion status, and other clinicopathological information. The GSE14520 dataset was downloaded from the GEO database as an external validation cohort. This GSE14520 dataset includes 247 hepatocellular carcinoma samples and their complete follow-up information. The above data was used for subsequent candidate gene screening, risk scoring model construction, and external validation. The data sources and modeling validation process have been disclosed in the technical disclosure document.
[0073] 185 glycosyltransferase genes and 656 ferroptosis-related genes were obtained from the GeneCards database. Among them, the glycosylation modification-related gene set may include glycosyltransferase-related genes, and the ferroptosis regulation-related gene set may include genes related to ferroptosis occurrence, regulation, inhibition, and redox homeostasis.
[0074] The aforementioned genes were then matched with the TCGA-LIHC transcriptome expression matrix to form a joint candidate gene expression matrix for glycosylation and ferroptosis-related genes. RNA-seq count data from 50 paired samples in the TCGA-LIHC cohort were further extracted. The expression data were normalized using the DESeq2 package with variance stabilization transformation to reduce the impact of sequencing depth differences and inter-sample expression level differences on subsequent analyses. Transcriptome expression data underwent sample matching, gene identification standardization, and expression level normalization. For RNA-seq expression data, a TPM expression matrix or a normalized expression matrix transformed with log2(TPM+1) could be used. For count data involved in differential expression analysis, DESeq2 was used for variance stabilization transformation to reduce the impact of sequencing depth differences on the expression matrix.
[0075] Subsequently, gene markers from the glycosylation modification-related gene set and the ferroptosis regulation-related gene set were matched with gene markers in the normalized expression matrix. Genes that could not be matched with expression records in the expression matrix were removed, and glycosylation modification-related genes and ferroptosis regulation-related genes with expression records were retained. The corresponding expression levels were then rearranged according to sample number to form a joint candidate gene expression matrix. This limited the original whole transcriptome data to the candidate gene range related to glycosylation modification and ferroptosis regulation, reducing interference from irrelevant features in the subsequent co-expression network construction and modeling stages.
[0076] S102. Construct a weighted gene co-expression network based on the joint candidate gene expression matrix, calculate the module correlation between each gene module and the tumor phenotype, and determine the gene modules that meet the module correlation screening conditions as key co-expression modules.
[0077] In this embodiment, the WGCNA method is used to perform weighted co-expression network analysis on the joint candidate gene expression matrix. First, cluster analysis is performed on the samples in the joint candidate gene expression matrix. Based on the sample clustering tree, samples with abnormal expression distribution are identified, and samples that do not meet the sample clustering quality requirements are removed to reduce the impact of outliers on the construction of the co-expression network.
[0078] Subsequently, multiple soft threshold parameters β are iterated, and the scale-free topology fit index and average connectivity are calculated for each soft threshold parameter. When the scale-free topology fit index reaches a preset fitting condition and the average connectivity is in a stable variation range, the corresponding soft threshold parameter is selected as the network construction parameter. In this embodiment, soft threshold parameters in the range of 1 to 20 can be iterated. When the scale-free topology fit index R² reaches 0.95 or higher and the average connectivity tends to be stable, the soft threshold parameter β is determined to be 6.
[0079] Based on the determined soft threshold parameter β, the correlation between genes is calculated and a weighted adjacency matrix is constructed. Further, the weighted adjacency matrix is converted into a topological overlap matrix. Based on the topological overlap matrix, genes are hierarchically clustered, and a dynamic tree cutting method is used to divide the gene modules. During module division, the minimum number of genes in a module can be set to 30, and modules with high similarity in their characteristic genes are merged. The merging and cutting height threshold can be set to 0.25.
[0080] After obtaining multiple gene modules, the module characteristic genes of each gene module are calculated, and these module characteristic genes are the first principal components of the gene expression matrix within that module. Then, the correlation coefficients and corresponding statistical test values between each module characteristic gene and the tumor phenotype are calculated. The tumor phenotype may include the grouping status of tumor tissue and adjacent normal tissue.
[0081] In this embodiment, the red module showed the highest correlation with the tumor phenotype, with a correlation coefficient of 0.95. The statistical test results met the preset conditions. Therefore, the red module was identified as the key co-expression module, and the genes in this key co-expression module were extracted for subsequent differential expression analysis. Instead of directly screening prognostic genes from all glycosylation and ferroptosis-related genes, the co-expression network was first used to identify the module regions related to the tumor phenotype, so that the subsequent screening was constrained by the module correlation, thereby reducing the randomness brought about by simple differential expression screening.
[0082] like Figure 2A As shown, after completing the construction of the weighted gene co-expression network, the correlation analysis results between the module characteristic genes of each gene module and the tumor phenotype can be obtained. Based on Figure 2A The results show that the red modules that meet the screening criteria for correlation with tumor phenotype are identified as key co-expression modules. The red modules have a high correlation with tumor phenotype, with a correlation coefficient of 0.95.
[0083] S103. Perform differential expression analysis on the genes in the key co-expression module between tumor tissue and adjacent normal tissue to obtain a candidate prognostic gene set that simultaneously meets the module correlation condition and the differential expression condition.
[0084] In this embodiment, gene expression data from the key co-expression modules identified in step S102 are extracted, and differential expression analysis is performed using tumor tissue samples and adjacent normal tissue samples as comparison objects. For RNA-seq counting data, DESeq2 can be used for differential expression analysis. Before analysis, the input data undergoes a sample number consistency check, and the analysis design matrix is set up according to a pairing or grouping relationship between tumor tissue samples and adjacent normal tissue samples. Subsequently, the fold change in expression of each gene in the key co-expression modules between tumor tissue and adjacent normal tissue and the statistical test value are calculated based on a statistical model.
[0085] In this embodiment, the screening criteria for differentially expressed genes can be set as follows: |log2Fold Change| corresponding to the fold change in expression is greater than 0.5, and the corrected p-value is less than 0.05. Genes that meet the above criteria are identified as differentially expressed genes. These differentially expressed genes are used as a candidate prognostic gene set for subsequent survival-related screening and risk scoring model construction.
[0086] In the specific data processing, after screening and differential expression analysis using the WGCNA module, 77 candidate prognostic genes were obtained, including both upregulated and downregulated genes. These candidate prognostic genes were simultaneously constrained by both the tumor phenotype-related module and the differential expression criteria, unlike methods that rely solely on single expression differences for candidate gene screening.
[0087] S104. Based on the candidate prognostic gene set and clinical follow-up data, survival-related genes are screened, and a Cox regression model with regularization penalty term is used to perform dimensionality reduction modeling of survival-related genes, determine the core gene combination and its corresponding regression coefficient, and construct a risk scoring model.
[0088] In this embodiment, hepatocellular carcinoma samples with complete clinical follow-up data were first screened. The clinical follow-up data included overall survival and survival status. To reduce the impact of short-term non-tumor-related events on model training, samples with an overall survival of 30 days or less were excluded. After screening, 342 hepatocellular carcinoma samples with complete survival information were obtained.
[0089] Subsequently, the 342 samples were divided into a training set and a test set in a 1:1 ratio, with each set consisting of 171 samples. The training set was used for model building, and the test set was used for internal model validation.
[0090] In the training set, using overall survival as the time variable and survival status as the outcome variable, univariate Cox regression analysis was performed on each gene in the candidate prognostic gene set obtained in step S103. Through univariate Cox regression analysis, genes associated with overall survival were screened as survival-related genes. In this embodiment, 54 survival-related genes were selected from 77 candidate prognostic genes.
[0091] Furthermore, LASSO-Cox regression analysis was performed on the 54 survival-related genes. LASSO-Cox regression introduces an L1 regularization penalty term to reduce the dimensionality of high-dimensional gene expression features. During modeling, the glmnet package can be used to set up a Cox proportional hazards model, and the penalty parameter λ is determined through 10-fold cross-validation. Specifically, the training set is divided into 10 subsets, with 9 subsets used for model training each time, and the remaining subset used for validation. This process is repeated until every subset is used as a validation set. Based on the partial likelihood bias under different λ values, the λ value corresponding to the lowest partial likelihood bias is selected as the penalty parameter.
[0092] like Figure 2B As shown, the penalty parameter of the LASSO-Cox regression model is determined through cross-validation to achieve a balance between candidate gene number compression and prediction error control. The vertical axis represents partial likelihood deviation, and the horizontal axis represents Log(λ). As λ increases, the model complexity decreases, and the number of non-zero coefficients is gradually compressed from 54 to 0. The dashed line on the left corresponds to the λ value (lambda.min) when the cross-validation error is minimized, at which point the model includes 13 non-zero coefficient variables. The dashed line on the right corresponds to the penalty parameter value within one standard error range of the minimum error, used to help determine the degree of model compression. Figure 2C The study shows the path of regression coefficient shrinkage as the penalty parameter changes for different candidate genes. Based on this, core gene combinations with non-zero regression coefficients are screened, and a risk scoring model is constructed based on the core gene combinations and their corresponding regression coefficients.
[0093] In this embodiment, the penalty parameter λ determined by 10-fold cross-validation is 0.073. Under this penalty parameter, 10 core genes with non-zero regression coefficients were selected. The core gene combination includes: B4GALNT1, EZH2, KIF20A, MYCN, ST6GALNAC4, CBS, G6PD, SQSTM1, B3GNTL1, and NQO1.
[0094] Risk Score = 0.312×Exp(B4GALNT1) + 0.204×Exp(EZH2) + 0.026×Exp(KIF20A)+0.030×Exp(MYCN)+0.026×Exp(ST6GALNAC4)−0.027×Exp(CBS)+0.179×Exp(G6PD)+0.097×Exp(SQSTM1) + 0.072×Exp(B3GNTL1) + 0.059×Exp(NQO1);
[0095] Wherein, Exp(B4GALNT1), Exp(EZH2), Exp(KIF20A), Exp(MYCN), Exp(ST6GALNAC4), Exp(CBS), Exp(G6PD), Exp(SQSTM1), Exp(B3GNTL1), and Exp(NQO1) represent the normalized expression levels of the corresponding genes in the samples to be evaluated. The normalized expression levels can be obtained by converting the expression levels using log2(TPM+1). This transforms the high-dimensional expression characteristics of glycosylation-related genes and ferroptosis-regulated genes into a risk scoring model with a fixed calculation format, facilitating subsequent consistent risk calculations for the samples to be evaluated.
[0096] S105. Calculate the risk score of the hepatocellular carcinoma sample to be evaluated based on the risk scoring model, and output the risk stratification results and corresponding prognostic assessment results according to the preset cutoff value.
[0097] In this embodiment, for the hepatocellular carcinoma sample to be evaluated, its corresponding transcriptome expression data are first obtained, and the normalized expression levels of B4GALNT1, EZH2, KIF20A, MYCN, ST6GALNAC4, CBS, G6PD, SQSTM1, B3GNTL1, and NQO1 are extracted according to the normalization method in step S101. Subsequently, the normalized expression levels of each gene are substituted into the risk scoring model constructed in step S104 to calculate the Risk Score of the sample to be evaluated.
[0098] The median risk score in the training set is used as the cutoff value, or the optimal cutoff value determined by the training set is used to divide the samples to be evaluated into high-risk or low-risk groups. When the risk score of the sample to be evaluated is greater than or equal to the cutoff value, the result for the high-risk group is output; when the risk score of the sample to be evaluated is lower than the cutoff value, the result for the low-risk group is output.
[0099] Survival curves for the high-risk and low-risk groups were generated using the Kaplan-Meier method on the training, test, and external validation sets, respectively, and the difference in overall survival between the two groups was compared using the Log-rank test. Simultaneously, the predictive efficacy of the risk scoring model at preset time points was evaluated using time-dependent receiver operating characteristic (ROC) curves. These preset time points may include 1 year, 3 years, and 5 years.
[0100] After obtaining risk scores from the training set, test set, and external validation set, the samples are divided into high-risk and low-risk groups based on the cutoff values determined in the training set, and survival curves are generated using the Kaplan-Meier method. Figure 3A As shown, the overall survival rate of patients in the high-risk group in the training set was significantly lower than that in the low-risk group, and the difference between the two groups was statistically significant (P < 0.0001). Figure 3B For the survival curves of the validation set, the prognosis of the high-risk group was still significantly worse than that of the low-risk group (P = 0.019). Figure 3C Using the external validation set GSE14520, the high-risk group also showed worse survival outcomes, with a statistically significant difference between groups (P = 0.006). These results demonstrate that the risk scoring model exhibits good prognostic discrimination ability in the training set, validation set, and independent external validation cohort, and can reliably identify high-risk individuals with poor prognosis in hepatocellular carcinoma. This indicates that the risk scoring model can stratify the prognostic risk of hepatocellular carcinoma samples based on the expression levels of core genes.
[0101] Furthermore, the predictive performance of the risk scoring model at preset time points is evaluated using time-dependent ROC curves, such as... Figure 4A The AUCs for the training set shown are 0.87, 0.81, and 0.80 for 1, 3, and 5 years, respectively. Figure 4B The AUCs for the test set at 1 year, 3 years, and 5 years were 0.81, 0.74, and 0.72, respectively. Figure 4C The 1-year, 3-year, and 5-year AUCs for the external validation set GSE14520 are 0.77, 0.67, and 0.63, respectively. These results demonstrate that the constructed risk scoring model has reproducible prognostic risk stratification capabilities in the training data, internal test data, and external validation data.
[0102] Furthermore, the risk score and clinicopathological factors were jointly input into a multivariate Cox regression model. These clinicopathological factors included at least one of TNM stage, vascular invasion status, age, sex, AFP level, and Child-Pugh classification. After multivariate Cox regression analysis, the risk score, TNM stage, and vascular invasion status could be used as input variables for the nomogram prediction model.
[0103] After validating the risk scoring model, the risk score and clinicopathological factors can be further input into a multivariate Cox regression model. These clinicopathological factors include TNM staging and vascular invasion status. Figure 5 As shown, a nomogram prediction model is constructed based on multivariate Cox regression results. This model uses Risk Score, TNM stage, and vascular invasion status as predictor variables, and the survival probability at a preset time point as the output. For the sample to be evaluated, the total score can be calculated based on the corresponding integrals of each predictor variable in the nomogram, and the prognostic probability at the corresponding time point can be obtained based on the total score.
[0104] Example 2
[0105] To further verify the technical effectiveness of the present invention in integrating glycosylation and ferroptosis-related genes and constructing a model using the LASSO-Cox algorithm compared to a single pathway model, this embodiment sets up two comparative examples.
[0106] Comparative Example 1 is a single-pathway glycosylation model. A prognostic model was constructed using only 185 glycosylation genes according to the method in Example 1. WGCNA was used to screen key modules (correlation with tumor phenotype r=0.82), DESeq2 differential analysis yielded 62 differentially expressed genes, and LASSO-Cox regression was used to screen out 8 core genes to construct a risk scoring model.
[0107] Comparative Example 2 is a single-pathway ferroptosis model. A prognostic model was constructed using only 656 ferroptosis genes in the same manner. WGCNA was used to screen key modules (correlation with tumor phenotype r=0.79), DESeq2 differential analysis yielded 71 differentially expressed genes, and LASSO-Cox regression identified 9 core genes to construct a risk scoring model.
[0108] The dual-pathway integration model of this invention uses both glycosylation-related genes and ferroptosis-related genes for candidate gene screening, as described in Example 1, integrating the glycosylation and ferroptosis pathways. WGCNA was used to screen key modules (correlation with tumor phenotype r=0.95), resulting in 77 differentially expressed genes. LASSO-Cox regression was used to screen 10 core genes, and a risk scoring model was constructed.
[0109] Table 1: Comparative experimental results evaluating the performance of each model on the TCGA test set n=171:
[0110] Comparative Example 1 0.72 0.68 0.65 0.64 HR=1.85, P=0.012 Comparative Example 2 0.74 0.70 0.66 0.66 HR=1.92, P=0.008 This invention 0.81 0.74 0.72 0.73 HR=2.85, P<0.001
[0111] The results showed that the 1-year AUC of the dual-pathway model of this invention (0.81) was significantly higher than that of Comparative Example 1 (0.72) and Comparative Example 2 (0.74), with relative improvements of 12.5% and 9.5%, respectively. In the 3-year and 5-year long-term survival prediction, the AUC of the model of this invention (0.74, 0.72) was significantly higher than that of the single-pathway model (0.68 / 0.65 and 0.70 / 0.66), demonstrating that the dual-pathway integration has a sustained advantage in long-term prognostic assessment. The risk stratification capability (HR) of the model of this invention was also demonstrated. The LASSO algorithm's HR = 2.85 significantly outperformed the single-pathway model (HR = 1.85 and 1.92), demonstrating that the 10-gene combination screened by the LASSO algorithm has a stronger prognostic discrimination ability than the single-pathway 8-9-gene combination. The model of this invention achieved an AUC of 0.77 after 1 year on the GEO external validation set (GSE14520), while the monoglycosylation model achieved 0.68 and the monoferroptosis model achieved 0.70, confirming that the dual-pathway integration model has better generalization ability and effectively avoids the overfitting risk of the single-pathway model.
[0112] Conclusion: This embodiment demonstrates that by integrating the characteristics of both glycosylation and ferroptosis pathways and combining them with the LASSO algorithm to effectively handle the problem of collinearity in high-dimensional data, the model of this invention significantly outperforms existing single-pathway models in terms of prediction accuracy, risk stratification ability, and generalization performance, achieving unexpected technical results.
[0113] Example 3
[0114] This embodiment, based on Embodiment 1, provides a hepatocellular carcinoma prognostic assessment system based on glycosylation and ferroptosis-related genes. Deployed on a general-purpose computer, server, workstation, or cloud computing platform, it executes the hepatocellular carcinoma prognostic assessment method described in Embodiment 1. The system takes the transcriptome expression matrix, tumor phenotype information, and clinical follow-up data of hepatocellular carcinoma samples as input, and outputs risk scores, risk stratification results, and corresponding prognostic assessment results. The system's module structure corresponds to the data acquisition, WGCNA module screening, differential expression analysis, LASSO-Cox modeling, and risk score output processes disclosed in the technical disclosure; it includes a data acquisition module, a co-expression module screening module, a differential expression screening module, a risk model construction module, and a prognostic result output module.
[0115] The data acquisition module is used to acquire the transcriptome expression matrix, tumor phenotype information, clinical follow-up data, glycosylation modification-related gene set, and ferroptosis regulation-related gene set of hepatocellular carcinoma samples, and map the glycosylation modification-related gene set and the ferroptosis regulation-related gene set to the transcriptome expression matrix to form a joint candidate gene expression matrix. Specifically, the data acquisition module can read RNA-seq expression data from the TCGA-LIHC cohort, phenotypic grouping information of tumor tissue and adjacent normal tissue, and clinical follow-up data such as overall survival and survival status; it can also read the expression matrix and follow-up information from the GEO external validation cohort. The data acquisition module performs sample number matching, gene identification unification, and expression level standardization on the input data, and retains the glycosylation modification-related genes and ferroptosis regulation-related genes with expression level records in the transcriptome expression matrix, thereby forming a joint candidate gene expression matrix.
[0116] The co-expression module screening module is used to construct a weighted gene co-expression network based on the joint candidate gene expression matrix, calculate the module correlation between each gene module and the tumor phenotype, and determine key co-expression modules. Specifically, the co-expression module screening module performs cluster analysis on the samples in the joint candidate gene expression matrix, removing samples that do not meet the clustering quality requirements; it iterates through multiple soft threshold parameters, calculating the scale-free topological fit index and average connectivity under different soft threshold parameters; when the scale-free topological fit index meets the preset fitting conditions and the average connectivity is in a stable variation range, the soft threshold parameters required for network construction are determined; then a weighted adjacency matrix is constructed, the weighted adjacency matrix is converted into a topological overlap matrix, and gene modules are divided based on the topological overlap matrix. The co-expression module screening module further calculates the module characteristic genes of each gene module, and determines key co-expression modules based on the correlation coefficient between the module characteristic genes and the tumor phenotype and the statistical test results.
[0117] The differential expression screening module is used to perform differential expression analysis on genes in the key co-expression module between tumor tissue and adjacent normal tissue to obtain a candidate prognostic gene set. Specifically, the differential expression screening module extracts gene expression data from the key co-expression module and calculates the fold change and statistical test value of each gene based on the expression difference between tumor tissue samples and adjacent normal tissue samples; genes whose fold change meets a preset fold change condition and whose statistical test value meets a preset significance condition are identified as differentially expressed genes, and these differentially expressed genes are used as the candidate prognostic gene set.
[0118] The risk model construction module is used to screen for survival-related genes based on the candidate prognostic gene set and clinical follow-up data, and to perform dimensionality reduction modeling of survival-related genes using a Cox regression model with a regularization penalty term to determine the core gene combination and its corresponding regression coefficients, thereby constructing a risk scoring model. Specifically, the risk model construction module divides hepatocellular carcinoma samples with complete clinical follow-up data into training and testing sets; in the training set, using overall survival and survival status as survival outcome variables, univariate Cox regression analysis is performed on each gene in the candidate prognostic gene set to obtain survival-related genes; subsequently, LASSO-Cox regression is used to perform dimensionality reduction screening of the survival-related genes, and the penalty parameter is determined through cross-validation; genes with non-zero regression coefficients under the penalty parameter are selected as core gene combinations, and a risk scoring model is constructed based on the standardized expression levels and corresponding regression coefficients of each gene in the core gene combination.
[0119] In this embodiment, the core gene combination includes B4GALNT1, EZH2, KIF20A, MYCN, ST6GALNAC4, CBS, G6PD, SQSTM1, B3GNTL1, and NQO1. The risk scoring model can be expressed as:
[0120] Risk Score = 0.312×Exp(B4GALNT1) + 0.204×Exp(EZH2) + 0.026×Exp(KIF20A)+0.030×Exp(MYCN)+0.026×Exp(ST6GALNAC4)−0.027×Exp(CBS)+0.179×Exp(G6PD)+0.097×Exp(SQSTM1) + 0.072×Exp(B3GNTL1) + 0.059×Exp(NQO1);
[0121] Wherein, Exp(B4GALNT1), Exp(EZH2), Exp(KIF20A), Exp(MYCN), Exp(ST6GALNAC4), Exp(CBS), Exp(G6PD), Exp(SQSTM1), Exp(B3GNTL1), and Exp(NQO1) represent the normalized expression levels of the corresponding genes in the samples to be evaluated.
[0122] The prognostic result output module is used to calculate the risk score of the hepatocellular carcinoma sample to be evaluated based on the risk scoring model, and output the risk stratification result and corresponding prognostic assessment result according to the preset cutoff value. Specifically, the prognostic result output module reads the standardized expression level of the core gene combination in the sample to be evaluated, substitutes it into the risk scoring model to obtain the risk score; then compares the risk score with the cutoff value determined in the training set, and outputs the risk stratification result of high-risk group or low-risk group. Further, the prognostic result output module can also output Kaplan-Meier survival analysis results, time-dependent ROC curve evaluation results, C-index evaluation results, and nomogram prediction results generated based on risk score and clinicopathological factors. This enables the hepatocellular carcinoma prognostic assessment process to be executed automatically in a computer environment and ensures that different samples obtain consistent risk scores and risk stratification results under the same calculation process.
[0123] Example 4
[0124] This embodiment provides a computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements the hepatocellular carcinoma prognostic assessment method based on glycosylation and ferroptosis-related genes as described in Specific Embodiment 1.
[0125] The computer-readable storage medium may be a read-only memory, random access memory, disk, optical disk, solid-state drive, mobile storage device, or server storage medium. The computer program may be written in R, Python, or other programming languages capable of transcriptome data processing, survival analysis, and model construction.
[0126] In one optional implementation, when the computer program is executed by the processor, it sequentially performs the following processes: acquiring the transcriptome expression matrix, tumor phenotype information, and clinical follow-up data of hepatocellular carcinoma samples; acquiring the gene set related to glycosylation modification and the gene set related to ferroptosis regulation, and mapping them to the transcriptome expression matrix to form a joint candidate gene expression matrix; constructing a weighted gene co-expression network based on the joint candidate gene expression matrix to identify key co-expression modules related to tumor phenotype; performing differential expression analysis on genes in the key co-expression modules to obtain a candidate prognostic gene set; performing survival-related screening based on the candidate prognostic gene set and clinical follow-up data, and using the LASSO-Cox regression model to determine the core gene combination and its regression coefficients; constructing a risk scoring model based on the core gene combination and its regression coefficients; calculating the risk score of the hepatocellular carcinoma sample to be evaluated according to the risk scoring model, and outputting the risk stratification results and corresponding prognostic assessment results.
[0127] In one optional implementation, the computer program is further configured to invoke a data standardization program, a WGCNA analysis program, a differential expression analysis program, a Cox regression analysis program, a LASSO-Cox regression program, a time-dependent ROC analysis program, and a nomogram generation program. The data standardization program is used to standardize gene identification and expression levels in transcriptome expression data; the WGCNA analysis program is used to perform soft-threshold screening, weighted adjacency matrix construction, topological overlap matrix transformation, and identification of key co-expression modules; the differential expression analysis program is used to calculate the fold change in expression of each gene in the key co-expression modules and the statistical test values; the LASSO-Cox regression program is used to perform dimensionality reduction screening of survival-related genes and output genes with non-zero regression coefficients; the time-dependent ROC analysis program is used to output the predictive efficacy evaluation results at preset time points; and the nomogram generation program is used to generate a prognostic probability prediction model based on risk scores and clinicopathological factors.
[0128] The hepatocellular carcinoma prognostic assessment method in Specific Implementation 1 can be solidified into an executable program through the aforementioned computer-readable storage medium, enabling the method to be repeatedly run on a general-purpose computer, server, or cloud computing platform, while maintaining consistency in risk score calculation, risk stratification output, and model validation processes.
[0129] Example 5
[0130] This embodiment provides an electronic device, which includes a processor and a memory. The memory is used to store executable instructions of the processor, and the processor is configured to execute the executable instructions to implement the hepatocellular carcinoma prognostic assessment method based on glycosylation and ferroptosis-related genes as described in Specific Embodiment 1.
[0131] The electronic device may be a desktop computer, laptop computer, server, workstation, or cloud computing platform. The processor may be a central processing unit, graphics processing unit, digital signal processor, or other processing unit capable of executing computer program instructions. The memory may include RAM, solid-state drive, hard disk drive, or other non-volatile storage media. The electronic device may also include a data input interface, a data output interface, and a display device. The data input interface is used to import the transcriptome expression matrix, tumor phenotype information, and clinical follow-up data of hepatocellular carcinoma samples. The data output interface and display device are used to output risk scores, risk stratification results, predictive efficacy evaluation results, and nomogram prediction results.
[0132] In one optional implementation, when the processor executes executable instructions in the memory, it first reads the transcriptome expression matrix, the glycosylation modification-related gene set, and the ferroptosis regulation-related gene set, and forms a joint candidate gene expression matrix; then, it performs weighted gene co-expression network analysis based on the joint candidate gene expression matrix to identify key co-expression modules; next, it performs differential expression analysis on the genes in the key co-expression modules to obtain a candidate prognostic gene set; then, it performs univariate Cox regression analysis and LASSO-Cox regression analysis based on the candidate prognostic gene set and clinical follow-up data to obtain the core gene combination and its regression coefficients; finally, it calculates the risk score based on the standardized expression level and regression coefficient of each gene in the core gene combination and outputs the corresponding risk stratification results.
[0133] In one optional embodiment, the electronic device is configured with a model building unit, a model validation unit, and a result output unit. The model building unit is used to perform candidate gene screening, core gene identification, and risk score model construction based on training set samples; the model validation unit is used to perform Kaplan-Meier survival analysis, time-dependent ROC analysis, and C-index calculation based on test set samples and external validation set samples; the result output unit is used to output the risk score of the sample to be evaluated, the results of the high-risk group or low-risk group, the predictive efficacy evaluation results at preset time points, and the nomogram prediction results generated based on the risk score and clinicopathological factors.
[0134] In one optional embodiment, the electronic device stores a risk scoring model in its memory. The risk scoring model includes the gene names, corresponding regression coefficients, and risk scoring formulas for B4GALNT1, EZH2, KIF20A, MYCN, ST6GALNAC4, CBS, G6PD, SQSTM1, B3GNTL1, and NQO1. After receiving the standardized expression levels of the sample to be evaluated, the processor calls the risk scoring formula to calculate the Risk Score and outputs the risk stratification results based on the cutoff values determined in the training set.
[0135] The aforementioned electronic devices enable automated processing of existing hepatocellular carcinoma transcriptome data and clinical follow-up data, allowing the prognostic assessment method in Implementation Method 1 to complete data import, feature screening, model construction, risk score calculation, and assessment result output in a software-based manner.
[0136] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A prognostic assessment method for hepatocellular carcinoma based on glycosylation and ferroptosis-related genes, characterized in that, Includes the following steps: S101. Obtain the transcriptome expression matrix, tumor phenotype information and clinical follow-up data of hepatocellular carcinoma samples, and obtain the gene set related to glycosylation modification and the gene set related to ferroptosis regulation. Map the gene set related to glycosylation modification and the gene set related to ferroptosis regulation to the transcriptome expression matrix to form a joint candidate gene expression matrix. S102. Construct a weighted gene co-expression network based on the joint candidate gene expression matrix, calculate the module correlation between each gene module and the tumor phenotype, and determine the gene modules that meet the module correlation screening conditions as key co-expression modules. S103. Perform differential expression analysis on the genes in the key co-expression module between tumor tissue and adjacent normal tissue to obtain a candidate prognostic gene set that simultaneously meets the module correlation condition and the differential expression condition. S104. Based on the candidate prognostic gene set and clinical follow-up data, survival-related genes are screened, and a Cox regression model with regularization penalty term is used to perform dimensionality reduction modeling of survival-related genes, determine the core gene combination and its corresponding regression coefficient, and construct a risk scoring model. S105. Calculate the risk score of the hepatocellular carcinoma sample to be evaluated based on the risk scoring model, and output the risk stratification results and corresponding prognostic assessment results according to the preset cutoff value.
2. The method for prognostic assessment of hepatocellular carcinoma based on glycosylation and ferroptosis-related genes according to claim 1, characterized in that, The formation of the joint candidate gene expression matrix includes: Transcriptome expression data were obtained from hepatocellular carcinoma tumor tissue samples, adjacent normal tissue samples, and hepatocellular carcinoma samples with clinical follow-up data. The transcriptome expression data were subjected to sample matching, gene identification unification, and expression level standardization to form a standardized expression matrix; Gene identifiers from the glycosylation modification-related gene set and the ferroptosis regulation-related gene set were matched with gene identifiers in the normalized expression matrix; Glycosylation-related genes and ferroptosis-regulating genes with expression levels recorded in the standardized expression matrix are retained to form a joint candidate gene expression matrix.
3. The method for prognostic assessment of hepatocellular carcinoma based on glycosylation and ferroptosis-related genes according to claim 1, characterized in that, The determination of the key co-expression module includes: The sample clustering results are calculated based on the joint candidate gene expression matrix, and samples that do not meet the sample clustering quality requirements are removed based on the sample clustering results. Iterate through multiple soft threshold parameters and calculate the scale-free topology fit index and average connectivity under different soft threshold parameters; A weighted adjacency matrix is constructed based on a soft threshold parameter that allows the scale-free topology fitting index to meet a preset fitting condition and the average connectivity to be in a stable range. The weighted adjacency matrix is converted into a topological overlap matrix, and gene module partitioning is performed based on the topological overlap matrix; Calculate the module characteristic genes of each gene module, and determine the key co-expression modules based on the correlation coefficient between module characteristic genes and tumor phenotype and the results of statistical tests.
4. The method for prognostic assessment of hepatocellular carcinoma based on glycosylation and ferroptosis-related genes according to claim 1, characterized in that, Obtaining the candidate prognostic gene set includes: Extract gene expression data from the key co-expression module; Based on the expression differences between tumor tissue samples and adjacent normal tissue samples, the fold change in expression of each gene in the key co-expression module and the statistical test value were calculated. Genes whose expression fold change meets the preset fold change condition and whose statistical test value meets the preset significance condition are identified as differentially expressed genes; The differentially expressed genes were used as a candidate prognostic gene set for survival-related screening and risk scoring model construction.
5. The method for prognostic assessment of hepatocellular carcinoma based on glycosylation and ferroptosis-related genes according to claim 1, characterized in that, The construction of the risk scoring model includes: Hepatocellular carcinoma samples with complete clinical follow-up data were divided into training and testing sets. In the training set, univariate Cox regression analysis was performed on each gene in the candidate prognostic gene set, with overall survival and survival status as survival outcome variables, to obtain survival-related genes. LASSO-Cox regression analysis was performed on the survival-related genes, and the penalty parameters were determined by cross-validation. Genes with non-zero regression coefficients under the penalty parameters are selected as core gene combinations; a risk scoring model is constructed based on the standardized expression levels and corresponding regression coefficients of each gene in the core gene combination; wherein, the core gene combination includes B4GALNT1, EZH2, KIF20A, MYCN, ST6GALNAC4, CBS, G6PD, SQSTM1, B3GNTL1 and NQO1.
6. The method for prognostic assessment of hepatocellular carcinoma based on glycosylation and ferroptosis-related genes according to claim 5, characterized in that, The risk scoring model is as follows: Risk Score = 0.312×Exp(B4GALNT1) + 0.204×Exp(EZH2) + 0.026×Exp(KIF20A)+0.030×Exp(MYCN)+0.026×Exp(ST6GALNAC4)−0.027×Exp(CBS)+0.179×Exp(G6PD)+0.097×Exp(SQSTM1) + 0.072×Exp(B3GNTL1) + 0.059×Exp(NQO1); Where Exp(B4GALNT1), Exp(EZH2), Exp(KIF20A), Exp(MYCN), Exp(ST6GALNAC4), Exp(CBS), Exp(G6PD), Exp(SQSTM1), Exp(B3GNTL1), and Exp(NQO1) represent the normalized expression levels of the corresponding genes in the samples to be evaluated. After calculating the risk score of the hepatocellular carcinoma sample to be evaluated based on the risk scoring model, the risk score is compared with the cutoff value determined by the training set, and the risk stratification results of high-risk group or low-risk group are output, and the corresponding prognostic assessment results are generated.
7. A prognostic assessment system for hepatocellular carcinoma based on glycosylation and ferroptosis-related genes, implemented according to any one of claims 1-6, characterized in that, include: The data acquisition module is used to acquire the transcriptome expression matrix, tumor phenotype information, clinical follow-up data, glycosylation modification-related gene set and ferroptosis regulation-related gene set of hepatocellular carcinoma samples, and map the glycosylation modification-related gene set and the ferroptosis regulation-related gene set to the transcriptome expression matrix to form a joint candidate gene expression matrix. The co-expression module screening module is used to construct a weighted gene co-expression network based on the joint candidate gene expression matrix, calculate the module correlation between each gene module and the tumor phenotype, and identify key co-expression modules. The differential expression screening module is used to perform differential expression analysis of genes in the key co-expression module between tumor tissue and adjacent normal tissue to obtain a candidate prognostic gene set. The risk model construction module is used to screen for survival-related genes based on the candidate prognostic gene set and clinical follow-up data, and to perform dimensionality reduction modeling of survival-related genes using a Cox regression model with regularization penalty term, determine the core gene combination and its corresponding regression coefficient, and construct a risk scoring model. The prognostic result output module is used to calculate the risk score of the hepatocellular carcinoma sample to be evaluated based on the risk scoring model, and output the risk stratification result and the corresponding prognostic assessment result according to the preset cutoff value.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the hepatocellular carcinoma prognostic assessment method based on glycosylation and ferroptosis-related genes as described in any one of claims 1-6.
9. An electronic device, characterized in that, include: processor; The processor also includes a memory for storing executable instructions of the processor; wherein the processor is configured to execute the hepatocellular carcinoma prognostic assessment method based on glycosylation and ferroptosis-related genes as described in any one of claims 1-7 by executing the executable instructions.
10. The application of the hepatocellular carcinoma prognostic assessment method based on glycosylation and ferroptosis-related genes as described in any one of claims 1-7 in the preparation of a hepatocellular carcinoma prognostic assessment kit.