Colorectal cancer prognosis marker, prognosis prediction model and construction method thereof

By constructing a colorectal cancer prognosis evaluation model based on ANKZF1, P4HA1 and STC2, combined with Shaboutok's small molecule inhibitor, the accuracy of prognosis evaluation in colorectal cancer patients is solved, the treatment effect and chemotherapy resistance analysis is improved for high-risk patients, and the immunotherapy improvement plan is provided.

CN120472991APending Publication Date: 2025-08-12重庆医科大学国际体外诊断研究院
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510559285.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The lack of effective biomarkers and therapeutic targets in the prognostic evaluation of colorectal cancer patients has led to limited effectiveness in most patients, especially in patients with colorectal cancer with liver metastasis, with short survival and a lack of accurate oncological strategies.

Method used

By obtaining gene data for colorectal cancer from TCGA and GEO databases, using the molecular tag database to identify genes related to hypoxia and fructose metabolism, a multigene risk model was constructed, ANKZF1, P4HA1 and STC2 were screened out as key prognostic genes, and their high expression in colorectal cancer cell lines were verified, and the high affinity of Shabbutok combined with these genes as a small molecule inhibitor was used to build a prognostic evaluation model.

Benefits of technology

Improve the accuracy of prognostic evaluation in patients with colorectal cancer, reveal treatment options in high-risk patients, enhance resistance analysis to standard chemotherapy and antiangiogenic drugs, and provide potential improvement pathways for immunotherapy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472991A_ABST
    Figure CN120472991A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of biology, in particular to a colorectal cancer prognosis marker, a prognosis prediction model and a construction method of the colorectal cancer prognosis marker. Gene expression data of a colorectal cancer patient and single-cell RNA sequencing data of a tumor immune single-cell center are obtained from a cancer genome map and a gene expression comprehensive database, and 224 hypoxia and fructose metabolism related genes are selected; a hypoxia and fructose metabolism related gene scoring system is constructed through regression analysis and various machine learning algorithms, a risk model is constructed, eight genes significantly related to poor prognosis are preliminarily selected from the risk model, and further verification is performed in combination with TCGA transcriptome data and a protein expression profile of a human protein map database. Kaplan-Meier survival curve, multi-factor regression analysis, column diagram construction and ROC curve evaluation are carried out, key prognosis genes of P4HA1, ANKZF1 and STC2 are determined, it is found that sabutok can be combined with ANKZF1, STC2 and P4HA1 in a high-affinity mode, and sabutok is revealed to serve as a promising treatment choice for high-risk patients.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biotechnology, and in particular to colorectal cancer prognostic markers, a prognostic prediction model and a construction method thereof. Background Art

[0002] Colorectal cancer is a common malignant tumor of the digestive tract, second only to gastric cancer and esophageal cancer in incidence. The most important factor affecting the prognosis of colorectal cancer patients is metastasis, with liver metastasis being the most common. According to statistics, 20%-40% of colorectal cancer patients have liver metastasis at the time of diagnosis, and the incidence of metachronous liver metastasis is even higher, reaching 50%. The average survival of patients with untreated liver metastases is 16-18 months, while the survival rate for those with extensive metastases is only 3-5 months. The vast majority of patients with liver metastases are unable to undergo radical resection due to factors such as the later development of extrahepatic metastatic lesions, liver masses involving multiple major blood vessels, or insufficient liver function.

[0003] Immunotherapy has shown efficacy in specific patient populations; however, due to the molecular heterogeneity of colorectal cancer, its clinical benefits are often limited. Therefore, the discovery of new prognostic biomarkers and therapeutic targets is crucial for the advancement of precision oncology strategies. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide colorectal cancer prognostic markers, prognostic prediction models and their construction methods, identify three key prognostic genes: P4HA1, ANKZF1 and STC2, and find that Sabutoc is a promising small molecule inhibitor that can bind to ANKZF1, STC2 and P4HA1 with high affinity.

[0005] The present invention solves the above technical problems through the following technical means:

[0006] In a first aspect, the present application discloses a colorectal cancer prognosis marker, comprising a gene combination, wherein the gene combination mainly consists of three genes: ANKZF1, P4HA1 and STC2.

[0007] In a second aspect, the present application discloses the use of the colorectal cancer prognosis marker described in the first aspect above in the preparation of a product for colorectal cancer prognosis detection.

[0008] In a third aspect, the present application discloses the use of the colorectal cancer prognostic marker described in the first aspect above in the preparation of a product for evaluating the efficacy of colorectal cancer immunotherapy response.

[0009] In a fourth aspect, the present application discloses the use of the colorectal cancer prognosis marker described in the first aspect above in the preparation of a product for predicting colorectal cancer drug sensitivity.

[0010] In a fifth aspect, the present application discloses a method for constructing a colorectal cancer prognosis evaluation model, comprising the following steps:

[0011] Colorectal cancer genetic and clinical data were obtained from the TCGA and GEO databases, and the data were preprocessed. 224 genes related to hypoxia and fructose metabolism were identified using the Molecular Signature Database (MsigDB). Univariate Cox regression analysis was performed in the TCGA dataset to screen out 38 genes significantly associated with colorectal cancer prognosis. Least absolute shrinkage and selection operator regression analysis was used on the 38 screened genes to select 16 prognostic risk genes, and a polygenic risk model was established. Based on the polygenic risk model, 8 candidate genes significantly associated with poor prognosis were screened. Differential expression analysis of the 8 candidate genes was performed in the TCGA cohort, and protein expression levels were detected using the Human Protein Atlas database. Kaplan-Meier survival analysis was performed, and receiver operating characteristic (ROC) curve analysis was further used to screen out 3 key prognostic genes, namely ANKZF1, P4HA1, and STC2. Gender, age, tumor stage, and risk score were integrated to estimate the survival probability of colorectal cancer patients and construct a colorectal cancer prognosis assessment model.

[0012] In a sixth aspect, the present application discloses a colorectal cancer prognosis assessment model obtained by the construction method described in the fifth aspect above.

[0013] The present invention uses the Kaplan-Meier method to compare the overall survival of each group of patients, and uses the "timeROC" R package to draw time-dependent receiver operating characteristic curves to quantitatively evaluate the prognostic performance of the model. Eight genes significantly associated with poor prognosis (HR>1.5, p<0.05) were initially screened from the risk model. Further validation was performed by combining transcriptome data from TCGA and protein expression profiles from the Human Protein Atlas database. Kaplan-Meier survival curves, multivariate regression analysis, nomogram construction, and receiver operating characteristic (ROC) curve evaluation were performed, ultimately identifying three key prognostic genes: P4HA1, ANKZF1, and STC2, and verifying their high expression in colorectal cancer cell lines. In addition, functional validation experiments further revealed that hypoxia significantly increases the resistance of colorectal cancer cells to standard chemotherapy drugs and anti-angiogenic drugs. Given the urgent need for targeted intervention for high-risk patients, the present application conducted molecular docking analysis and found that Sabutoc is a promising small molecule inhibitor that can bind to ANKZF1, STC2, and P4HA1 with high affinity. Molecular dynamics simulations also validated its stability, highlighting sabutol as a promising treatment option for high-risk patients. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1Univariate Cox regression analysis was used to identify genes related to hypoxia and fructose metabolism that were associated with prognosis in the TCGA-CRC cohort; Figure 2 is the risk coefficient graph of the 16 genes included in the risk model; Figure 3 This is a visualization of the distribution of risk scores in the TCGA-CRC cohort; Figure 4 Figure 3 is the Kaplan-Meier survival curve of the 1-, 3-, and 4-year overall survival rates of the high-risk and low-risk groups in the TCGA-CRC cohort, where A is the 1-year overall survival rate, B is the 3-year overall survival rate, and C is the 4-year overall survival rate. Figure 5 The Kaplan-Meier survival curves of the 1-, 3-, and 4-year overall survival rates in the GSE17536 cohort are shown, where D is the 1-year overall survival rate, E is the 3-year overall survival rate, and F is the 4-year overall survival rate. Figure 6 The Kaplan-Meier survival curves of the 1-, 3-, and 4-year overall survival rates in the GSE29621 cohort are shown, where G is the 1-year overall survival rate, H is the 3-year overall survival rate, and I is the 4-year overall survival rate. Figure 7 The ROC curves of the time-dependent 1-, 3-, and 4-year overall survival rates of the three groups of patients are shown in Figure 1. E represents the TCGA-CRC cohort, F represents the GSE17536 cohort, and G represents the GSE29621 cohort. Figure 8 The results of the gene feature enrichment analysis of the risk model are shown in Figure 5, where AC shows the gene features enriched in the GO analysis, including biological processes ( Figure 8 Middle A), cell components ( Figure 8 Middle B) and molecular function ( Figure 8 The significant pathways in C) and D) are shown in the KEGG enrichment analysis. Figure 9 These are the nine main cell types in the GSE146771 single-cell dataset; Figure 10 is the cellular composition of 10 colorectal cancer (CRC) tumor samples in the GSE146771 dataset; Figure 11 It is the expression profile of cell type-specific marker genes in the GSE146771 dataset; Figure 12 is that high expression of prognostic risk genes was observed in the monocyte / macrophage population; Figure 13 This is a visualization of risk gene expression in different cell types in the GSE146771 dataset; Figure 14 It is a cell-cell interaction analysis of the GSE146771 dataset; Figure 15 is the differential expression map of risk model genes in Mono / Macro cells stratified by TNM stage; Figure 16 The cell type annotation map and risk gene expression map in the EMTAB8107 single-cell dataset, where A is the cell type annotation map and B is the risk gene expression map; Figure 17These are the cell type annotation maps and risk gene expression maps in the GSE139555 dataset, where A is the cell type annotation map and B is the risk gene expression map; Figure 18 is the differential expression plot of eight candidate genes in the risk model of the TCGA-CRC cohort; Figure 19 is a graph of the protein levels of ANKZF1, DTNA, ENO3, P4HA1, PPFIA4, PYGM, and STC2 in colorectal cancer and normal tissues; Figure 20 Figure 2 is the overall survival analysis graph of ANKZF1, P4HA1, and STC2 in the TCGA-CRC cohort, where A is the overall survival analysis graph of STC2, B is the overall survival analysis graph of P4HA1, and C is the overall survival analysis graph of ANKZF1. Figure 21 is the ROC curve diagram of ANKZF1, P4HA1 and STC2 genes; Figure 22 is a multivariate Cox regression analysis plot combining the risk scores of ANKZF1, P4HA1, and STC2 with the clinical characteristics of the TCGA cohort; Figure 23 It is a nomogram that combines sex, age, tumor stage, and ANKZF1, P4HA1, and STC2 risk scores to predict the 2-, 3-, and 4-year overall survival rates of patients with colorectal cancer; Figure 24 This is the ROC curve of the prediction performance of the nomogram based on the ANKZF1, P4HA1, and STC2 genes, where A is the 2-year overall survival rate, B is the 3-year overall survival rate, and C is the 4-year overall survival rate; Figure 25 These are the expression level graphs of three key prognostic genes (STC2, P4HA1, and ANKZF1) in the human normal colon mucosal epithelial cell line NCM460 and the colorectal cancer cell lines LoVo, SW620, HCT116, SW480, and Caco-2. A is the expression level graph of the STC2 gene, B is the expression level graph of the P4HA1 gene, and C is the expression level graph of the ANKZF1 gene. Figure 26 This is a heat map of the correlation between the 16 risk model genes and various immune cell populations; Figure 27 This is the correlation analysis diagram of immune cells and overall immune cell infiltration with 16 risk prognostic genes; Figure 28 This is a comparison of the TIDE score, immune dysfunction score, immune rejection score, and microsatellite instability map between high-risk and low-risk groups in the TCGA-CRC cohort, where A represents the TIDE score, B represents the immune dysfunction score, C represents the immune rejection score, and D represents the microsatellite instability. Figure 29 This is a graph showing the distribution of the proportion of high-risk and low-risk patients who responded to anti-PD-1 immunotherapy in the CRC immunotherapy cohort. Figure 30 It is a comprehensive immune profile of the tumor microenvironment; Figure 31 is a heat map of metabolic features associated with the tumor microenvironment; Figure 32 Figure 2 shows significant differences in immune rejection, epithelial-mesenchymal transition (EMT), immunosuppression, and immune microenvironment between high-risk and low-risk groups. Figure 33 is the distribution of genes associated with the pan-fibroblast TGFβ signature between the high-risk and low-risk groups; Figure 34 This is a correlation analysis diagram of high-risk population and pan-fibroblast TGFβ marker-related genes; Figure 35 This is a drug sensitivity prediction map based on the GDSC database, where A is irinotecan, B is oxaliplatin, C is sorafenib, D is mitoxantrone, E is sabutol, and F is staurosporine; Figure 36 This is a schematic diagram of the in vitro experimental process for HCT116 and LoVo cell lines; Figure 37 is a morphological and quantitative comparison of HCT116 cells treated with oxaliplatin, irinotecan, or sorafenib under normoxic and hypoxic conditions; Figure 38 The graph shows the experimental results of LoVo cells treated with the same concentration of drugs under normoxic and hypoxic conditions; Figure 39 This is a graph showing the cell viability of HCT116 and LoVo cells treated with oxaliplatin, irinotecan, and sorafenib at different oxygen concentrations; Figure 40 are the conformational images of P4HA1, ANKZF1, and STC2 binding to mitoxantrone, sabutol, and staurosporine, respectively; Figure 41 This is a molecular dynamics simulation analysis diagram of the P4HA1-Sabtoker complex. DETAILED DESCRIPTION

[0015] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0016] Immune checkpoint inhibitors have introduced a new therapeutic avenue for colorectal cancer, but their effectiveness is primarily limited to patients with high tumor mutational burden (TMB-H) or high microsatellite instability (MSI-H) status. These subgroups comprise only 15% of all CRC cases and 4% of metastatic CRC cases. Consequently, the prognosis for patients with colorectal cancer remains suboptimal. The incidence and progression of CRC are driven by a combination of genetic, environmental, and dietary factors. With increasing global fructose intake, accumulating evidence suggests a strong association between excessive fructose intake and CRC progression. Hypoxia, a hallmark of the tumor microenvironment, drives metabolic reprogramming, promoting fructose utilization through glycolysis, thereby supporting tumor growth and invasion. Under hypoxic conditions, fructose serves as an adaptive energy source, and hypoxia-inducible factor-1 directly upregulates the fructose-specific transporter GLUT5 to enhance fructose uptake as an alternative metabolic fuel. This mechanism has been validated in various malignancies. For example, GLUT5 is significantly upregulated in lung cancer, accelerating tumor growth in vivo, while inhibition of GLUT5 in cholangiocarcinoma reduces fructose uptake and tumor cell invasion. Therefore, elucidating the role of the hypoxia-fructose metabolism axis in the occurrence and development of colorectal cancer is expected to improve patient prognosis.

[0017] This application obtained genetic and clinical data on colorectal cancer from the TCGA and GEO databases, preprocessed the data, and used the Molecular Signature Database (MsigDB) to identify 224 genes related to hypoxia and fructose metabolism. By integrating LASSO regression and 117 machine learning algorithms, a robust risk score model was optimized that included 16 genes related to hypoxia and fructose metabolism (including PGM2, ALDOB, CDKN1B, GLRX, NAGK, VAV2, SIAH2, ALDOC, PYGM, P4HA1, DTNA, CSRP2, STC2, ANKZF1, ENO3, and PPFIA4). This model demonstrated strong prognostic performance across different datasets, with high-risk patients exhibiting significantly poorer survival outcomes. Single-cell RNA sequencing analysis revealed that the prognostic risk genes were primarily expressed in monocytes and macrophages. Studies of immune cell infiltration confirmed this finding, revealing a significant association between model genes and M2-type tumor-associated macrophages. This suggests that genes related to hypoxia and fructose metabolism may regulate TAM polarization and alter the tumor microenvironment, ultimately impacting patient prognosis. Furthermore, single-cell RNA sequencing analysis revealed that model genes play a key role in advanced colorectal cancer and are positively correlated with TNM stage. The tumor microenvironment comprises diverse immune cell populations, among which immune checkpoint molecules play a key role in immune escape. To further elucidate the relationship between genes related to hypoxia and fructose metabolism and immunotherapy response, TIDE scores were analyzed. High-risk patients showed increased immunosuppression and immune escape scores and decreased MSI-H scores. These findings suggest that low-risk individuals may benefit more from immunotherapy, possibly due to their more favorable immune environment. Furthermore, signature analysis using the IOBR package revealed a significant upregulation of the pan-fibroblast TGF-β (Pan_F_TBRs) gene signature in high-risk patients, which may be a core barrier to immunotherapy efficacy. Previous studies have shown that the fibroblast-associated TGF-β pathway can inhibit CD8+ T cell infiltration in the tumor microenvironment, thereby weakening anti-tumor immunity and reducing the efficacy of immunotherapy. Notably, preclinical models have shown that dual blockade of TGF-β and PD-L1 can enhance T cell infiltration and enhance the efficacy of immunotherapy.

[0018] Furthermore, this application identified three key prognostic genes: ANKZF1, STC2, and P4HA1, and verified their high expression in colorectal cancer cell lines. ANKZF1 is upregulated in colorectal cancer and correlates with decreased overall survival. Silencing ANKZF1 inhibits colorectal cancer cell progression by downregulating MMP2, MMP9, N-cadherin, and HIF-α. STC2 is a hypoxia-inducible glycoprotein that is upregulated in various cancers, promoting survival, proliferation, and migration under hypoxic conditions while also promoting chemoresistance and immune evasion. P4HA1 participates in extracellular matrix remodeling during epithelial-mesenchymal transition and promotes angiogenesis through a hypoxia-inducible factor-1-mediated glycolysis pathway. Furthermore, P4HA1 regulates glycolytic metabolic reprogramming through the α-KG / TET2 / FBP1 axis, suggesting that it may be a metabolic therapeutic target for colorectal cancer. Functional validation experiments further revealed that hypoxia significantly increases the resistance of colorectal cancer cells to standard chemotherapeutic drugs and anti-angiogenic drugs.

[0019] In order to better understand the above technical solution, the above technical solution will be described in detail below with reference to the accompanying drawings and specific implementation methods.

[0020] The screening of markers, prognostic evaluation models, and their construction and verification in this application are carried out using the following methods:

[0021] (1) Data collection

[0022] Colorectal cancer RNA sequencing (RNA-seq) data and related clinical information were obtained from The Cancer Genome Atlas (TCGA) database (https: / / portal.gdc.cancer.gov), including 453 tumor samples and 41 normal samples. Two different CRC validation cohorts, GSE17536 (17) (n = 177) and GSE29621 (18) (n = 65), were obtained from the Gene Expression Omnibus (GEO) database (https: / / www.ncbi.nlm.nih.gov / geo / ). Data from the PD-1 immune checkpoint inhibitor (ICI) treatment cohort were obtained from the TIGER database (https: / / tiger.cancer-pku.cn). All datasets were preprocessed and normalized using R software to ensure data consistency and comparability. Genes associated with hypoxia and fructose metabolism were identified using the Molecular Signature Database (MSigDB, https: / / www.gsea-msigdb.org) and further refined by an extensive review of existing literature. A total of 224 genes related to hypoxia and fructose metabolism were identified.

[0023] (2) Establishing a prognostic risk scoring model

[0024] Univariate Cox regression analysis was used to screen gene signatures associated with hypoxia and fructose metabolism. Lambda was used for least absolute shrinkage and selection operator (LASSO) regression analysis to identify 16 prognostic risk genes. The "min" parameter in the R package "glmnet" was used to prevent overfitting and ensure model robustness. The "Mime" R package developed by Liu Hongwei et al. was used to randomly apply 117 machine learning algorithms. Machine learning techniques help identify complex patterns in high-dimensional data, thereby improving prognostic accuracy. Ten classic machine learning algorithms were evaluated using the "Mime" R package, including generalized boosted regression model (GBM), stepwise Cox (StepCox), CoxBoost, elastic net (Enet), partial least squares Cox regression (plsRcox), supervised principal component analysis (superpc), ridge regression, survival support vector machine (survivalsvm), LASSO, and random survival forest (RSF). The best-performing algorithm was then compared with the LASSO regression model in terms of predictive accuracy. The risk score for each colorectal cancer patient was calculated as follows:

[0025]

[0026] In the above table, GeneExpression i represents the expression level of the i-th gene in a certain sample, Coefi represents the risk coefficient of the i-th gene in the model, i represents the i-th gene, and n represents the total number of selected genes.

[0027] (3) Biological function analysis

[0028] The functional roles of 16 prognostic risk genes were analyzed using Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analysis. Genes differentially upregulated in the high-risk cohort were enriched using the "ggplot2" and "ClusterProfiler" R packages. Furthermore, GSEA and HALLMARK gene set enrichment analysis were performed using the log2-fold change distribution of differentially expressed genes to identify significantly enriched pathways and biological processes.

[0029] (4) Single-cell sequencing data analysis

[0030] Single-cell RNA sequencing datasets from colorectal cancer tissues were obtained from the Gene Expression Omnibus (GEO) database, including GSE146771, GSE139555, and the EMBL-EBI ArrayExpress dataset EMTAB8107. All collected data were obtained from the Tumor Immune Single-cell Hub (TISCH) database (http: / / tisch.comp-genomics.org). Quality control was performed using the MAESTRO analysis pipeline, and low-quality single cells and genes with low expression levels were removed to generate an initial gene expression matrix. Dimensionality reduction was performed using principal component analysis (PCA), followed by visualization using unified manifold approximation and projection (UMAP). Cluster analysis was then performed, and cell populations were annotated based on established marker genes. Specifically, the GSE146771 dataset contains 43,817 single cells from 10 CRC samples, while the GSE139555 dataset contains 10,112 single cells from two CRC samples. The EMTAB8107 dataset contains 7 CRC samples and 23,176 single cells. All datasets were generated using the 10× Genomics platform.

[0031] (5) Prognostic performance evaluation and identification of key prognostic genes

[0032] The Kaplan-Meier method was used to compare overall survival among patients in each group. Time-dependent receiver operating characteristic (ROC) curves were drawn using the "timeROC" R package to quantitatively evaluate the prognostic performance of the model. Eight genes significantly associated with poor prognosis (HR>1.5, p<0.05) were initially screened from the risk model. Further validation was performed using transcriptome data from TCGA and protein expression profiles from the Human Protein Atlas (HPA) database (https: / / www.proteinatlas.org). Kaplan-Meier survival curves, multivariate regression analysis, nomogram construction, and ROC curve evaluation were performed, ultimately identifying three key prognostic genes: P4HA1, ANKZF1, and STC2.

[0033] (6) Cell culture and real-time fluorescence quantitative PCR

[0034] Human colorectal cancer cell lines SW480, Caco2, HCT116, Lovo, SW620 and human normal mucosal epithelial cells NCM460 were derived from the Key Laboratory of Clinical Laboratory Diagnosis, School of Laboratory Medicine, Chongqing Medical University. The cells were placed in high-glucose DMEM containing 10% fetal bovine serum and 1% penicillin-streptomycin and cultured in a humidified incubator at 37°C. Total RNA was extracted using the steady-purure rapid RNA extraction kit (ACCURATE BIOLOGY, China), and then reverse transcribed using the Evo M-MLVRT Mix kit (ACCURATE BIOLOGY, China). qRT-PCR was performed using the SYBR Green PremixPRO Taq HS qPCR kit from ACCURATE BIOLOGY, China. GAPDH was used as an internal control. Primers designed and synthesized by Sangon Biotech. The qRT-PCR primer sequences used in this application are shown in Table 1:

[0035]

[0036] Table 1

[0037] (7) Cell counting kit (CCK-8) and clone formation assay

[0038] HCT116 and LoVo colorectal cancer cell lines were cultured under normoxic conditions (37°C, 5% CO2) and exposed to hypoxic conditions (1% O2) using AnaeroPack (Mitsubishi Gas Chemical). After passaging, LoVo cells were seeded in 96-well plates at a density of 3000 cells / well, and HCT116 cells were seeded in 96-well plates at a density of 8000 cells / well, and cultured under normoxic and hypoxic conditions, respectively. After adhesion was confirmed, culture medium containing 10 μM oxaliplatin, irinotecan, or sorafenib was added, and the cells were incubated for another 24 hours. Subsequently, 100 μL of culture medium containing 10% CCK-8 reagent was added to each well and incubated for 2 hours. The absorbance was measured at 450 nm using a microplate reader. In the clonogenic experiment, 1000 HCT116 or LoVo cells were seeded into 6-well plates. After cell attachment and colony formation, the medium was replaced with fresh medium containing the same concentration of oxaliplatin, irinotecan, or sorafenib. The cells were cultured for 7–14 days. Colonies were then fixed with 1% formaldehyde and stained with crystal violet.

[0039] (8) Analysis of gene expression differences between risk groups

[0040] Differential expression analysis among different HFMRGs risk groups was performed using the "Deseq2" R package. Differentially expressed genes were identified based on an adjusted p-value (Padj) less than 0.05 and an absolute log2FC greater than 1. The "ggpub" R package was used to create volcano plots for visualization of the results.

[0041] (9) Immune function analysis

[0042] The “CIBERSORT” R package preliminarily assessed the differences in immune cell infiltration by estimating the number of different immune cell types in high- and low-risk groups. Subsequently, tumor immune dysfunction and rejection (TIDE) was used to predict the response to immune checkpoint inhibitors in different HFMRGs risk groups. The “ESTIMATE” R package was used to calculate the immune score, stromal score, and comprehensive ESTIMATE score to further understand the immune microenvironment. In addition, tumor-related metabolic activity, immune evasion, epithelial-mesenchymal transition (EMT), immunosuppression, and immune microenvironment characteristics were analyzed based on predefined labels in the “IOBR” R package.

[0043] (10) Molecular docking of key prognostic proteins and potential therapeutic compounds

[0044] The Genomics of Cancer Drug Sensitivity (GDSC) database (https: / / www.cancerrxgene.org) was used to assess drug sensitivity in high- and low-risk cohorts, providing insights into the differential responses of colorectal cancer cells to different therapeutic agents. The half-maximal inhibitory concentration (IC50), a commonly used metric for evaluating therapeutic efficacy, was used to assess drug sensitivity. This analysis identified three conventional chemotherapeutic agents and three potential therapeutic candidates. To further investigate the interactions between these candidate compounds and key prognostic proteins, we retrieved the structures of P4HA1 (PDB ID: 1TJC), ANKZF1 (PDB ID: 5WHG), and STC2 (PDB ID: 8A7E) from the RCSB Protein Data Bank (https: / / www.rcsb.org). The molecular structures of the candidate compounds were obtained from the PubChem database (https: / / pubchem.ncbi.nlm.nih.gov). Molecular docking simulations were performed using AutoDockVina software to predict binding affinity and interaction patterns. The docking process involves the preparation of proteins and ligands, execution of docking simulations, and subsequent analysis of the results. Kolman charges, solvation parameters, and polar hydrogens were incorporated into the protein structure, while Gasteiger charges were assigned to the ligands. The docking grid spacing centered on the catalytic site of the heme binding domain is To ensure accurate positioning of the binding site, 100 independent docking simulations were performed for each ligand using the Lamarckian genetic algorithm to optimize the selection of docking poses. The docking results were visualized using PyMOL, enabling detailed structural analysis and interpretation of binding interactions.

[0045] (11) Molecular dynamics simulation

[0046] Molecular dynamics simulations of the P4HA1-sabutoclax complex were performed using GROMACS, while quantum chemical calculations were performed using ORCA, Multiwfn, and Sobtop. GROMACS was used as the primary simulation engine to ensure consistent system preparation and enable high-throughput molecular dynamics simulations. The simulation process consisted of four main steps: system preparation, energy minimization, equilibration, and running a production simulation. During the system preparation phase, protein topology was generated using pdb2gmx with the AMBER14SB force field and the TIP3P explicit water model. Ligand parameters and topology were exported using Sobtop to ensure compatibility with the selected force field. The complex was then positioned within a cubic simulation box, ensuring a minimum separation of 1.2 nm from the box boundary. Energy minimization used the steepest descent algorithm to resolve steric clashes and optimize the molecular geometry before equilibration. The equilibration process was performed in two distinct stages: initially, an NVT ensemble equilibration was performed over 100 picoseconds, during which the system was gradually heated from 0 to 300 Kelvin using a Berendsen thermostat. An additional 100 picoseconds of NPT ensemble equilibration was then performed to stabilize the pressure at 1 bar using a Parrinello-Rachmann barostat. During the equilibration phase, molecular bonds were constrained using the LINCS algorithm. In the subsequent production phase, Newton's equations of motion were integrated using a leapfrog algorithm. Long-range electrostatic interactions were calculated using the particle mesh Ewald (PME) method with a real-space cutoff of 1.0 nm, while van der Waals interactions were cutoff at the same distance. Periodic boundary conditions were applied in all directions, and a 50-nanosecond production mode simulation was performed with an integration time step of 2 femtoseconds.

[0047] (12) Statistical analysis

[0048] Statistical analyses were performed using open-source R software (version 4.4.2) and GraphPad Prism 10.2.1. Comparisons between two groups were performed using the t-test or Wilcoxon test. Multiple group comparisons were performed using one-way analysis of variance. P values less than 0.05 were considered statistically significant (*p < 0.05; **p < 0.01; ***p < 0.001; ****p < 0.0001).

[0049] The following results were obtained using the above experimental method:

[0050] Univariate Cox regression analysis was used to screen gene signatures related to hypoxia and fructose metabolism. Univariate Cox regression analysis in the TCGA dataset found that 38 genes were significantly associated with prognosis, such as Figure 1 Among them, 28 genes showed hazard ratios (HR) > 1, indicating an association with worse survival outcomes, while 10 genes had HRs < 1, indicating a potential protective effect.

[0051] In order to construct a reliable multi-gene prognostic model, the 38 genes were subjected to the least absolute shrinkage and selection operator (LASSO) regression analysis using λ to further refine the selection of the risk score model, thereby selecting 16 meaningful genes. The hazard ratios of these 16 genes were ranked from low to high, and the corresponding risk coefficients were displayed, as shown in Figure 2. Figure 2 Among them, 8 genes showed HR>1.5 and P<0.05, which highlighted a strong correlation with poor survival outcomes.

[0052] The risk score for each colorectal cancer patient is calculated as follows:

[0053]

[0054] Using the above formula, the risk score suitable for each patient was calculated, and the TCGA cohort was divided into high and low risk categories according to the optimal cutoff value (group.cutoff = 70), as shown in Figure 5. Figure 3 shown. Figure 3 In the study, the distribution of risk scores in the TCGA-CRC cohort was visualized, patients were ranked from low risk to high risk, survival status was classified, and gene expression patterns were presented in the form of a heat map. Patients with high HFMRGs scores had significantly lower overall survival.

[0055] The Kaplan-Meier method was used to compare the overall survival of patients in each group. The Kaplan-Meier survival curves of the 1-, 3-, and 4-year overall survival rates of the high-risk group and the low-risk group in the TCGA-CRC cohort are shown in Figure 2. Figure 4 The Kaplan-Meier survival curves of 1-, 3-, and 4-year overall survival in the GSE17536 cohort are shown in Figure 5 The Kaplan-Meier survival curves of 1-, 3-, and 4-year overall survival in the GSE29621 cohort are shown in Figure 6 As shown in Figure 2. Kaplan-Meier survival analysis confirmed that the 1-year, 3-year, and 4-year overall survival rates of high-risk patients were significantly worse. The “timeROC” R package was used to draw time-dependent receiver operating characteristic (ROC) curves to quantitatively evaluate the prognostic performance of the model, as shown in Figure 2. Figure 7 As shown. Figure 7As shown in Figure E, in the TCGA cohort, the area under the curve (AUC) values for 1-year, 3-year, and 4-year overall survival (OS) were 0.727, 0.765, and 0.800, respectively; to evaluate the general applicability of the model, two independent CRC cohorts (GSE17536 and GSE29621) were further analyzed, as shown in Figure 5E. Figure 7 As shown in Figure F, in GSE17536, the AUC values of 1-, 3-, and 4-year overall survival rates were 0.758, 0.650, and 0.632, respectively; Figure 7 As shown in G, in GSE29621, the AUC values were 0.844, 0.735, and 0.686, respectively.

[0056] To clarify the biological functions of the 16 genes selected by the selection operator regression analysis, Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway analyses were performed. The analysis results are shown in Figure 8 .like Figure 8 As shown in Figures AC, GO enrichment analysis found that these genes were related to carbohydrate metabolism, with a particular focus on glycolysis, fructose metabolism, and NADH regeneration. Figure 8 As shown in middle D, KEGG pathway analysis showed that pathways related to fructose and mannose metabolism, glycolysis / gluconeogenesis, and HIF-1 signaling pathway were significantly enriched.

[0057] To explore the expression patterns of risk marker genes in different cell types in the colorectal cancer tumor microenvironment, single-cell RNA sequencing data from the GSE146771 dataset were used for analysis. A total of 23 different cell subpopulations were identified from 10 CRC tumor samples, including 9 major immune cell and stromal cell types, such as Figure 9 The cell composition of 10 colorectal cancer (CRC) tumor samples in the GSE146771 dataset is shown in Figure 10 As shown, Figure 10 The data showed that the cell composition of different patients was different, with monocytes and macrophages constituting the main cell populations. Based on this classification, the expression of marker genes for each cell type was further analyzed, such as Figure 11 It is noteworthy that the risk signature genes are mainly enriched in clusters 2 and 9 of monocytes and macrophages, as shown in Figure 12 As shown. Further visualization, such as Figure 13 As shown in Figure 3, these risk signature genes are concentrated in myeloid cell subsets. Cell-cell interaction (CCI) analysis was performed, as shown in Figure 3. Figure 14As shown in Figure 2, cell-cell interaction analysis revealed extensive interactions between monocytes / macrophages and various T cell subsets, including consuming CD8+ T cells (CD8Tex_C8), cytotoxic CD8+ T cells (CD8T_C6, CD8T_C1), and proliferating T cells (Tprolif_C19). The differential expression of risk model genes in Mono / Macro cells stratified by TNM stage is shown in Figure 2. Figure 15 As shown, stratified analysis based on TNM stage showed that risk signature genes were significantly upregulated in monocytes and macrophages of patients with advanced disease. The upregulation of these genes was particularly associated with tumor progression in stage IIB, IIIB, and IIIC CRC cases. Two additional single-cell RNA sequencing datasets, EMTAB8107 and GSE139555, were analyzed. Figure 16 As shown in Figure A, 12 different cell types were identified in CRC tissues in the EMTAB8107 dataset, while 11 cell types were detected in the GSE139555 dataset. Figure 17 As shown in A. Figure 16 Middle B and Figure 17 As shown in middle B, risk prognostic genes were also highly expressed in monocytes / macrophages in both EMTAB8107 and GSE139555 datasets (***p<0.001).

[0058] To improve the accuracy of prognostic assessment for colorectal cancer patients, a multi-gene risk model established previously was used to screen key prognostic genes and a nomogram prediction model was established. First, eight genes significantly associated with poor prognosis (HR>1.5, p<0.05) were preliminarily screened from the risk model (see Figure 2 ) as candidate genes and then differential expression analysis was performed in the TCGA cohort, see Figure 18 , found that ANKZF1, P4HA1 and STC2 were significantly overexpressed in the tumor group. To further explore the protein expression of ANKZF1, P4HA1 and STC2 genes in CRC, the protein expression levels were detected using the Human Protein Atlas database. The protein levels of ANKZF1, DTNA, ENO3, P4HA1, PPFIA4, PYGM and STC2 in colorectal cancer and normal tissues are shown in the figure below. Figure 19 The results showed that the protein expression levels of ANKZF1, P4HA1 and STC2 were significantly increased in colorectal cancer tissues, consistent with their mRNA expression patterns. The significant consistency between transcription and translation levels suggested that these genes may play a key role in the progression of colorectal cancer. Kaplan-Meier survival analysis showed that ANKZF1 ( Figure 20 (shown in A), P4HA1 ( Figure 20(shown in B) and STC2 ( Figure 20 The results of the ROC curve analysis showed that the ANKZF1, P4HA1 and STC2 genes were significantly associated with poor overall survival, and the survival time of patients in the high expression group was significantly shorter (P = 0.0032, 0.025, 0.016). Figure 21 As shown in the results, the AUC values of P4HA1, ANKZF1 and STC2 were 0.84, 0.95 and 0.95, respectively, which were significantly better than conventional clinical features.

[0059] To evaluate the independent prognostic significance of the three key genes ANKZF1, P4HA1, and STC2, their risk scores were included in the multivariate regression analysis together with sex, age, and tumor stage, as Figure 22 As shown in the results, the analysis showed that the three-gene risk score showed a stronger ability to predict survival compared with traditional clinical features (HR = 3.03, P < 0.001). Therefore, a nomogram model was constructed to estimate the 2-year, 3-year and 4-year survival probability of colorectal cancer patients by integrating sex, age, tumor stage and risk score, as shown in the figure below. Figure 23 As shown. Figure 24 As shown in Figure 2, ROC curve analysis showed that the AUC values of the nomogram for predicting 2-year, 3-year, and 4-year survival rates were 0.809, 0.829, and 0.832, respectively. Subsequently, the expression levels of P4HA1, ANKZF1, and STC2 in different colorectal cancer cell lines were analyzed. The expression levels of three key prognostic genes (STC2, P4HA1, and ANKZF1) in the human normal colon mucosal epithelial cell line NCM460 and the colorectal cancer cell lines LoVo, SW620, HCT116, SW480, and Caco-2 were shown in Figure 2. Figure 25 As shown, compared with the normal colon epithelial cell line NCM460, STC2 was highly expressed in the SW620 cell line, ANKZF1 was upregulated in both SW620 and LoVo cell lines, and P4HA1 was significantly upregulated in LoVo, SW480, Caco2 and HCT116 cell lines.

[0060] In summary, P4HA1, ANKZF1, and STC2 were effectively identified as key prognostic genes. The nomogram model established based on the risk score has a strong predictive efficacy for the prognosis of colorectal cancer and provides a valuable tool for clinical prognostic assessment.

[0061] In this application, immune cell infiltration between different risk score groups was also evaluated, e.g. Figure 26 and 27Results showed that genes in the risk signature were significantly positively correlated with both M0 and M2 macrophages. Specifically, CDKN1B (R = 0.11, p < 0.05), CSRP2 (R = 0.3, p < 0.05), GLRX (R = 0.11, p < 0.05), NAGK (R = 0.12, p < 0.05), P4HA1 (R = 0.14, p < 0.05), PGM2 (R = 0.14, p < 0.05), PYGM (R = 0.1, p < 0.05), and SIAH2 (R = 0.16, p < 0.05) were significantly associated with M2 macrophages. Despite the significant correlation between risk signature genes and macrophages, bulk RNA sequencing immune infiltration analysis revealed that M2 macrophages were not significantly elevated in the high-risk group and showed a trend toward slightly decreased infiltration. In contrast, the infiltration of M0 macrophages was significantly increased, indicating that the microenvironment was rich in undifferentiated macrophages, which may be precursors to the immunosuppressive phenotype. In addition, the high-risk cohort showed increased levels of regulatory T cells (tregs), while the low-risk cohort showed increased infiltration of resting CD4 memory T cells, resting dendritic cells, and resting mast cells. This distribution suggests that high-risk tumors transition to an immunosuppressive microenvironment. Subsequent validation using the TIDE algorithm confirmed that the high-risk cohort showed significantly higher TIDE scores, indicating enhanced immune dysfunction and rejection, and lower microsatellite instability (MSI) scores (such as Figure 28 In addition, please refer to Figure 29 In an independent cohort, analysis of responses to anti-PD-1 immunotherapy showed that 76% of high-risk patients were classified as non-responders, suggesting a lower sensitivity to immune checkpoint blockade therapy. Figure 30 Compared with the low-risk group, the stromal score, immune score, and ESTIMATE score of the high-risk group were significantly increased. Notably, genes associated with the risk signature were mainly enriched in macrophage cluster 2, a subpopulation that may be associated with immunosuppressive function. This finding suggests that immunosuppression in high-risk individuals may not be due to a general increase in M2 macrophages, but rather to functional changes in specific macrophage subsets. These macrophages, together with enhanced regulatory T cell infiltration, form a tumor microenvironment that promotes immune escape, tumor progression, and treatment resistance.

[0062] These findings highlight a complex immunosuppressive landscape in high-risk tumors, characterized by accumulation of M0 macrophages, functional modulation of macrophage subsets, increased regulatory T-cell infiltration, and reduced sensitivity to immunotherapy, emphasizing the need to consider macrophage plasticity and T-cell dysfunction when developing immunotherapy strategies for high-risk patients.

[0063] High-risk tumors in the TCGA cohort showed a clear tendency to immune escape and decreased responsiveness to immunotherapy. To further explore the biological characteristics of TME, the “signature” module in the IOBR package was used to systematically analyze gene sets related to tumor metabolism, immune escape, epithelial-mesenchymal transition, immunosuppressive signaling, and immune cell infiltration. These signatures were quantified using principal component analysis (PCA), resulting in a comprehensive overview of tumor-related metabolic reprogramming between risk score groups, such as Figure 31 shown.

[0064] like Figure 32 As shown, subsequent analysis of immune-related signatures showed that in the high-risk cohort, immune escape, epithelial-mesenchymal transition (EMT), immunosuppressive pathways, and global remodeling of the immune microenvironment were significantly elevated, and PAN_F_TBRs were significantly enriched. Previous studies have shown that fibroblast-derived TGF-β signaling contributes to resistance to immunotherapy by promoting an immune-rejecting tumor microenvironment, thereby reducing effective anti-tumor immunity. To further elucidate this mechanism, the differential expression and correlation of genes within the Pan_F_TBRs signature were analyzed, as shown in Figure 5. Figure 33 and Figure 34 These results indicate that Figure 34 The genes shown in were significantly overexpressed in the high-risk group and were accompanied by a strong positive correlation, implying the coordinated activation of TGF-β-mediated immunosuppressive signaling pathways.

[0065] Due to tumor heterogeneity, the clinical response to chemotherapy in patients with advanced colorectal cancer varies greatly. Sensitivity to conventional drugs significantly affects the prognosis of patients. To better understand the potential treatment response, data from the GDSC2 database were used to preliminarily evaluate oxaliplatin and irinotecan (two widely used chemotherapy drugs) and sorafenib (an anti-angiogenic drug). No significant difference in the response to irinotecan was observed between the high-risk group and the low-risk group (P = 0.093, Figure 35 A) or sorafenib (P = 0.18, Figure 35 There were significant differences in the IC50 values of the three drugs (shown in C). These results suggest that patients with high HFMRG risk scores may have reduced sensitivity to conventional treatment options. Considering these observations, attempts were made to identify alternative treatment options for high-risk patients. Based on further screening, three compounds - mitoxantrone, sabutoclax and staurosporine - emerged as promising candidate drugs. The IC50 values of the three drugs in the high-risk group were significantly lower than those in the low-risk group (staurosporine: P = 0.00065; mitoxantrone: P = 0.0049; sabutoclax: P = 0.0027; Figure 35(shown in DF), indicating that these patients have higher sensitivity and treatment potential. (Pan_F_TBRs: pan-fibroblast TGFβ signaling pathway; CAF: cancer-associated fibroblasts; EMT: epithelial-mesenchymal transition; WNT_target: WNT signaling pathway target gene; TAM: tumor-associated macrophages; MDSC: myeloid-derived suppressor cells; GPAGs: glycosylation-related genes; PPAGs: phosphorylation-related genes). In order to experimentally verify whether hypoxia contributes to chemotherapy resistance, in vitro experiments were performed using two CRC cell lines (HCT116 and LoVo) cultured under normoxic and hypoxic conditions. The experimental process is as follows: Figure 36 Hypoxia-exposed cells treated with equal concentrations of oxaliplatin, irinotecan, or sorafenib showed increased viability, especially in the HCT116 cell line, accompanied by observable morphological changes, such as Figure 37 As shown in Figure 2. Clone formation experiments further confirmed that under the same drug conditions, there were higher numbers of clones in hypoxic LoVo cells compared with normoxic controls, as shown in Figure 2. Figure 38 As shown in Figure 2, it indicates that drug sensitivity is weakened in a hypoxic environment. Figure 39 As shown in Figure 3, this result was also confirmed by CCK-8 assay, where both cell lines showed higher viability under hypoxic conditions.

[0066] The interaction between important prognostic genes and potential therapeutic drugs was preliminarily verified by molecular docking analysis. The protein structures of three key prognostic genes, P4HA1, ANKZF1, and STC2, were obtained from the RCSB protein database (PDB). The molecular structures of three prospective therapeutic drugs, mitoxantrone, sabutol, and staurosporine, were obtained from the PubChem database. Molecular docking simulation was performed using AutoDockVina software, and the simulation results were visualized using PyMOL software. The results are shown in Figure 3. Figure 40 The results show that, see Figure 40P4HA1 exhibited significant binding affinity for mitoxantrone and staurosporine, with optimal docking scores of -5.4 kcal / mol and -7.0 kcal / mol, respectively. Within the P4HA1 binding pocket, mitoxantrone formed noncovalent interactions, including a hydrophobic interaction with Leu177. Staurosporine also formed noncovalent interactions with P4HA1 and hydrophobic interactions with Leu168 and Tyr158. ANKZF1 had optimal docking scores of -7.2 kcal / mol and -8.3 kcal / mol, respectively, for mitoxantrone. Mitoxantrone formed hydrogen bonds with Ser215 and Lys337, a π-π interaction with Phe225, and hydrophobic interactions with Phe211 and Trp266. Similarly, staurosporine forms noncovalent interactions with ANKZF1, including π-π stacking interactions with Phe225 and hydrophobic interactions with Asp339 and Trp266. For STC2, the optimal docking scores for mitoxantrone and staurosporine were -7.2 kcal / mol and -9.1 kcal / mol, respectively. Mitoxantrone forms hydrogen bonds with Lys1556 and Asp1605, and hydrophobic interactions with Phe1575 and Val1555. Staurosporine also interacts noncovalently with STC2, forming hydrophobic interactions with Thr870, Glu794, and Ser795.

[0067] Overall, docking scores below -4.25 kcal / mol indicate potential binding, below -5.0 kcal / mol indicate good affinity, and below -7.0 kcal / mol indicate strong affinity. The binding energy scores for P4HA1, ANKZF1, and STC2 with sabutoclax are shown in Table 2. Notably, sabutoclax exhibited strong binding affinity to all three key prognostic proteins. Sabutoclax is a broad-spectrum Bcl-2 family inhibitor that can induce tumor cell apoptosis by targeting Bcl-2, Bcl-xL, and Mcl-1. It has demonstrated promising therapeutic potential in a variety of solid tumors, highlighting its potential as a targeted therapy for key prognostic genes.

[0068]

[0069] Table 2

[0070] Since molecular docking had confirmed the binding affinity of Sarbutol to three key prognostic proteins (P4HA1, ANKZF1, and STC2), molecular dynamics (MD) simulations were performed on P4HA1. This selection was based on its highest hazard coefficient in the prognostic model, indicating its major contribution to high-risk stratification of colorectal cancer patients. Given P4HA1's potential role in tumor progression and its strong docking interaction with Sarbutol, P4HA1 represented the most relevant candidate for further stability evaluation.

[0071] To evaluate the binding stability of the Sarbutoc-P4HA1 complex, 50 ns MD simulations were performed as Figure 41 As shown, Figure 41In the figure, A represents the time-dependent fluctuations of the root mean square deviation (RMSD) of the P4HA1 protein backbone and the ligand, reflecting the stability of the complex; B represents the root mean square fluctuation (RMSF) of the protein residues, highlighting the flexible regions within the P4HA1-Sabtok complex; C represents the variation of the gyration radius (Rg) with simulation time, providing insights into the compactness of the structure; D represents the fluctuation of the solvent accessible surface area (SASA) during the simulation, quantifying the solvent exposure dynamics; E represents the binding free energy component of the complex: the van der Waals interaction energy The ligand binding free energy (RMSD) analysis revealed that the complex was stable within the final 20 ns, indicating no significant conformational drift. The ligand RMSD remained within a low range, indicating stable binding within the P4HA1 binding pocket. Root mean square fluctuation (RMSF) analysis revealed that most residues exhibited minimal fluctuations, while significant changes were observed in the C-terminal region, likely attributable to the inherent flexibility of the protein. The radius of gyration (Rg) remained unchanged, indicating that the overall protein structure remained compact without significant unfolding or collapse. Solvent-accessible surface area (SASA) analysis further confirmed the structural stability of the complex. Finally, MM / GBSA calculations revealed a negative total binding free energy (ΔTOTAL), supporting the strong and stable binding affinity of Sabtok within the P4HA1 binding pocket. Van der Waals interactions (ΔVDWAALS) were identified as the primary stabilizing force, while electrostatic interactions (ΔEEL) were initially unfavorable. However, solvation effects (ΔEGB) compensated for this, resulting in an overall stable complex. The negative solvation free energy (ΔGSOLV) highlights the critical role of the solvent environment in maintaining binding stability.

[0072] Taken together, these findings indicate that Sarbutol has a strong and stable binding affinity for P4HA1, further emphasizing its advantage as a targeted therapy for high-risk patients.

[0073] Therefore, a combination gene consisting of the three genes ANKZF1, P4HA1 and STC2 can be used as a prognostic marker for colorectal cancer, for the preparation of a colorectal cancer prognosis detection product, or for the preparation of a colorectal cancer immunotherapy response efficacy evaluation product, or for the preparation of a colorectal cancer drug sensitivity prediction product. Among them, the colorectal cancer prognosis detection product includes a product for detecting the mRNA expression level or protein expression level of the combination gene. The product for detecting the mRNA expression level or protein expression level of the combination gene includes a substance that can bind to the nucleic acid of the combination gene or can bind to the protein expressed by the combination gene. The product for detecting the mRNA expression level or protein expression level of the combination gene is selected from any one of a reagent, a kit, a test paper, a gene chip, an antibody chip, and a high-throughput sequencing platform.

[0074] The above embodiments are intended only to illustrate the technical solutions of the present invention and are not intended to limit the scope of the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art will appreciate that the technical solutions of the present invention may be modified or replaced with equivalents without departing from the spirit and scope of the technical solutions of the present invention, and such modifications or equivalents shall be encompassed by the claims of the present invention. Any techniques, shapes, and structures not described in detail herein are well known.

Claims

1. A prognostic marker for colorectal cancer, characterized in that: The invention comprises a combination gene, which is mainly composed of three genes: ANKZF1, P4HA1 and STC2.

2. Use of the colorectal cancer prognosis marker according to claim 1 in the preparation of a product for colorectal cancer prognosis detection.

3. The use according to claim 2, characterized in that The colorectal cancer prognosis detection product includes a product for detecting the mRNA expression level or protein expression level of a combination of genes.

4. The use according to claim 3, characterized in that The product for detecting the mRNA expression level or protein expression level of the combination gene includes a substance that can bind to the nucleic acid of the combination gene or can bind to the protein expressed by the combination gene.

5. The use according to claim 4, characterized in that The product for detecting the mRNA expression level or protein expression level of the combination gene is selected from any one of a reagent, a kit, a test paper, a gene chip, an antibody chip, and a high-throughput sequencing platform.

6. Use of the colorectal cancer prognostic marker according to claim 1 in the preparation of a product for evaluating the efficacy of colorectal cancer immunotherapy response.

7. Use of the colorectal cancer prognostic marker according to claim 1 in the preparation of a product for predicting colorectal cancer drug sensitivity.

8. A method for constructing a colorectal cancer prognosis evaluation model, characterized in that: The following steps are involved: Colorectal cancer genetic and clinical data were obtained from the TCGA and GEO databases. The data were preprocessed and 224 genes related to hypoxia and fructose metabolism were identified using the molecular signature database. Univariate Cox regression analysis was performed in the TCGA dataset, and 38 genes were screened out that were significantly associated with the prognosis of colorectal cancer; The 38 genes screened were analyzed using least absolute shrinkage and selection operator regression to select 16 prognostic risk genes and establish a polygenic risk model; Based on the polygenic risk model, eight candidate genes significantly associated with poor prognosis were screened; The eight candidate genes were differentially expressed in the TCGA cohort, and protein expression levels were detected using the Human Protein Atlas database. Kaplan-Meier survival analysis was performed, and receiver operating characteristic (ROC) curve analysis was further used to screen out three key prognostic genes, namely ANKZF1, P4HA1, and STC2. Gender, age, tumor stage and risk score were integrated to estimate the survival probability of colorectal cancer patients and to construct a colorectal cancer prognostic evaluation model.

9. The method for constructing a colorectal cancer prognosis evaluation model according to claim 8, characterized in that: The risk score is calculated as follows: Among them, GeneExpression i represents the expression level of the i-th gene in a certain sample, Coefi represents the risk coefficient of the i-th gene in the model, i represents the i-th gene, and n represents the total number of selected genes.

10. A colorectal cancer prognosis evaluation model obtained by the construction method according to claim 8 or 9.