Construction method of colorectal cancer prognosis optimal effect model based on integrated ML algorithm

By integrating ML algorithms and gene expression data, we construct an optimal prognosis model for colorectal cancer, which solves the problem of difficult to reflect the differences in patients' immunotherapy sensitivity and tumor microenvironment in the existing technology, and realizes the accurate evaluation of the prognosis of colorectal cancer patients and the formulation of personalized treatment strategies.

CN120032702APending Publication Date: 2025-05-23NANCHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510017398.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art is difficult to effectively reflect the sensitivity of different colorectal cancer patients to immunotherapy and the differences in tumor microenvironment, which leads to difficulty in selecting treatment strategies, and existing prognostic models are relatively rare in clinical applications.

Method used

By obtaining transcriptomic data and immune-related genes of colorectal cancer patients from the TCGA database, GEO database and IMMPORT, the gene expression differences were analyzed using the limma R package, prognosis-related genes were screened out, and a variety of ML algorithms were integrated to construct the optimal prognosis model of colorectal cancer.

Benefits of technology

The accuracy of prognosis evaluation of colorectal cancer patients has been achieved, which has the potential value of helping patients stratify and formulate personalized treatment strategies, which can better reflect the tumor microenvironment characteristics of different patients and improve the effectiveness of immunotherapy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032702A_ABST
    Figure CN120032702A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of gene technology and biomedicine, and discloses a colorectal cancer prognosis optimal effect model construction method based on an integrated ML algorithm, and the method comprises the following steps: determining immune-related prognosis genes based on transcriptome data of a colorectal cancer patient and a healthy person collected from TCGA and GEO databases and immune-related genes obtained from IMMPORT; combining an integration algorithm of different combination modes, fitting a prediction model based on a training queue TCGA-CRC, verifying the prediction model by adopting a GSE17536 training set, and calculating a consistency index C-index of each data set; and after comparison, selecting the prediction model with the highest C-index average value as the colorectal cancer prognosis optimal effect model based on the integrated ML algorithm. The invention shows great potential in showing the drug sensitivity of a patient and immune cell subset infiltration and immune cell functions in a tumor microenvironment; the colorectal cancer prognosis optimal effect model based on the integrated ML algorithm has important value in patient prognosis prediction and treatment strategy selection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of gene technology and biomedical technology, and in particular relates to a method for constructing an optimal prognosis model for colorectal cancer based on an integrated ML algorithm. Background Art

[0002] Colorectal cancer is the third most common cancer and the second leading cause of cancer death worldwide. In recent years, the incidence and mortality of colorectal cancer have continued to rise. Among patients with early-stage disease, 25%-50% will develop metastatic colorectal cancer, and their survival rate is only about 20%. There is an urgent need to develop new treatment strategies to improve survival rates.

[0003] Common colorectal cancer treatment options include surgery, chemotherapy, radiotherapy, targeted therapy, and immunotherapy. Among them, immunotherapy using ICIs (immune checkpoint inhibitors) has made significant progress. It uses the patient's own immune system to fight cancer cells. Many successful cases have been reported in recent years, especially for colorectal cancer patients with mismatch repair deficiency or dMMR / MSI-H (microsatellite high instability). However, the proportion of colorectal cancer patients showing dMMR / MSI-H is low, accounting for only 15% and 4% of colorectal cancer and metastatic colorectal cancer cases, respectively. At the same time, the AJCC (American Joint Committee on Cancer) clinical management standards, which are currently widely used in colorectal cancer treatment decisions and monitoring strategies, are difficult to reflect the differences in sensitivity to immunotherapy and TME (tumor microenvironment) of different patients, and there are large differences in clinical outcomes among different patients at the same stage. In summary, it is crucial to identify new prognostic and treatment indicators for colorectal cancer.

[0004] The occurrence and development of colorectal cancer is closely related to the immune microenvironment of the patient's tumor. The immune microenvironment of colorectal cancer patients includes stromal cells, arterial duct system, other cells that affect tumors and their functional molecular systems in various aspects of the microenvironment. Immune-related markers in the tumor immune microenvironment are key factors affecting the immunotherapy effect and prognosis of patients. The development of tumor immunotherapy has produced fruitful results so far. T cell immunosuppressive factors such as PD-1 (programmed death receptor 1) and CTLA-4 (cytotoxic T lymphocyte-associated antigen 4) have been discovered, showing great clinical potential. In recent years, immunotherapy has made significant progress in the treatment of solid tumors, and CAR-T cell therapy (chimeric antigen receptor T cell immunotherapy), immune checkpoint inhibitors and cell therapy strategies have been designed. In the study of colorectal cancer, immunotherapy has also shown great potential. PD-1 inhibitors represented by nivolumab and pembrolizumab have been approved by the FDA (U.S. Food and Drug Administration) for patients with metastatic colorectal cancer with microsatellite instability, becoming an important strategy for the treatment of colorectal cancer. However, colorectal cancer has a highly suppressive immune microenvironment, which causes the problem of immunotherapy resistance, making immunotherapy ineffective for colorectal cancer subtypes such as pMMR / MSS and dMMR / MSI-L. In recent years, although researchers have designed some ICIs combination therapies to address some limitations of colorectal cancer treatment, they still have not designed a detailed classification based on the characteristics of the colorectal cancer tumor immune microenvironment to help choose diagnosis and treatment strategies in order to address the difficulty in choosing immunotherapy methods for different colorectal cancer patients.

[0005] In recent years, the continuous development of next-generation sequencing technology has made it possible to analyze the mechanism of occurrence and development of colorectal cancer, drug sensitivity, and prognostic differences at the molecular level, which has greatly helped the study of the tumor immune microenvironment of colorectal cancer. At the same time, ML (Machine Learning) can detect complex logical relationships from large, noisy data, helping researchers analyze complex omics data. In view of the problem that the results obtained by a single machine learning algorithm may be biased, the integration of machine learning algorithms significantly improves the generalization ability of the results, making the results more clinically significant. At this stage, many studies use sequencing data to construct colorectal cancer prognosis models, but very few can be actually applied in clinical practice.

[0006] To this end, those skilled in the art provide a method for constructing an optimal prognosis model for colorectal cancer based on an integrated ML algorithm to solve the problems raised by the background technology. Summary of the invention

[0007] In order to solve the above technical problems, the present invention provides a method for constructing an optimal prognosis model for colorectal cancer based on an integrated ML algorithm to solve the problems existing in the prior art.

[0008] In one aspect, the present invention provides a method for constructing an optimal prognosis model for colorectal cancer based on an integrated ML algorithm, comprising the following steps: S1. First, obtain and download the transcriptome data of colorectal cancer patient samples and healthy human samples from the TCGA database, then obtain and download the transcriptome data of colorectal cancer patient samples from the GEO database, and then obtain immune-related genes from IMMPORT; S2. Then, the limma R package was used to analyze the expression differences of immune-related genes in samples from colorectal cancer patients and healthy people, and 1,149 genes with significant differences in expression between the two groups were screened out; S3, then univariate Cox regression was used to identify 45 prognosis-related genes from the differentially expressed immune-related genes; S4. Combine multiple ML algorithms for random combination to obtain integrated algorithms with different combinations. Fit the prediction model based on the training cohort TCGA-CRC. Use GSE17536 as the validation set to validate the prediction model and calculate the consistency index C-index in the two datasets to reflect the discriminative ability of the prediction model based on different integrated algorithms. S5. Select the prediction model with the highest C-index average value as the most effective model for colorectal cancer prognosis based on the integrated machine learning algorithm.

[0009] Preferably, in step S3, the 45 prognosis genes include: ADIPOQ, BACH2, BIRC5, BMP5, CCL11, CCL24, CCL28, CD1A, CD1B, CRABP2, CXCL1, CXCL2, CXCL3, DEFA6, EREG, F2RL1, FABP4, GLP2R, GRP, HAMP, IL13RA2, IL20RB, INHBB, LEP, LTB4R, MC1R, NGFR, NOX1, NOX4, NR3C2, NRG1, PGF, PLCG2, PLXNA3, PTH1R, RETNLB, S100P, SCG2, SEMA5B, SLC11A1, SPP1, SSTR2, TPM2, UCN, and WNT5A.

[0010] Preferably, in step S4, the combination of multiple ML algorithms is performed randomly, and the multiple ML algorithms include: 10 machine learning algorithms: random survival forest, elastic network, Lasso, Ridge, StepCox, CoxBoost, Cox's partial least squares regression, supervised principal component, generalized boosting regression model and survival support vector machine, which constitute 101 integrated algorithms through multiple combinations.

[0011] Preferably, the prediction model with the highest C-index average value is CoxBoost+Ridge, which contains 20 immune-related prognostic genes.

[0012] Preferably, the 20 immune-related prognosis genes include: CD1A, CD1B, HAMP, S100P, FABP4, CCL24, PLCG2, SEMA5B, PLXNA3, EREG, GRP, INHBB, NRG1, PGF, RETNLB, UCN, GLP2R, IL20RB, MC1R and PTH1R.

[0013] Preferably, after step S5, the method further includes: verifying the most effective model for colorectal cancer prognosis based on an integrated machine learning algorithm.

[0014] Preferably, the method for verifying the most effective model for colorectal cancer prognosis based on integrated machine learning algorithms includes survival analysis, gene pathway enrichment analysis, immune correlation analysis, immune infiltration analysis and drug sensitivity analysis.

[0015] The present invention also provides an application of the optimal model for colorectal cancer prognosis based on the integrated ML algorithm described in the aforementioned embodiment.

[0016] Compared with the prior art, the present invention has the following beneficial effects: The present invention establishes an optimal prognosis model for colorectal cancer based on an integrated ML algorithm, which can be used as a prognostic evaluation indicator and has the potential value of helping patients stratify and formulate personalized treatment strategies. In addition, the optimal prognosis model for colorectal cancer based on an integrated ML algorithm has great potential in showing the differences in immune cell subpopulation infiltration and immune function in the tumor microenvironment of different patients, which may help in the selection of treatment strategies in clinical practice and bring new hope for the development of future immunotherapy interventions. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 This is a flow chart of a method for constructing the most effective model for colorectal cancer prognosis based on integrated ML algorithm in one embodiment of the present invention; Figure 2A This is a schematic diagram of the first process of establishing the most effective model for colorectal cancer prognosis based on the integrated ML algorithm in one embodiment of the present invention; Figure 2B This is a schematic diagram of the second process of establishing the most effective model for colorectal cancer prognosis based on the integrated ML algorithm in one embodiment of the present invention; Figure 2C This is a schematic diagram of the third process of establishing the most effective model for colorectal cancer prognosis based on the integrated ML algorithm in one embodiment of the present invention; Figure 3 This is a schematic diagram of the verification and analysis of the most effective model for colorectal cancer prognosis based on the integrated ML algorithm in one embodiment of the present invention; Figure 4 This is a schematic diagram of immune cell infiltration and immune microenvironment analysis in one embodiment of the present invention; Figure 5 This is a schematic diagram of differential gene enrichment in one embodiment of the present invention. DETAILED DESCRIPTION

[0018] The following embodiments of the present invention are described in further detail in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0019] like Figure 1 As shown: Embodiment: The present invention provides a method for constructing an optimal prognosis model for colorectal cancer based on an integrated ML algorithm, comprising the following steps: S1. First, the transcriptome data of 579 colorectal cancer patient samples and 51 healthy human samples were obtained and downloaded from the TCGA database, and then the transcriptome data of 177 colorectal cancer patient samples were obtained and downloaded from the GEO database, and then 2483 immune-related genes were obtained from IMMPORT. S2. Then, the limma R package was used to analyze the expression differences of immune-related genes in samples from colorectal cancer patients and healthy people, and 1,149 genes with significant differences in expression between the two groups were screened out; S3, then univariate Cox regression was used to identify 45 prognosis-related genes from the 1149 differentially expressed immune-related genes; S4. Combine multiple ML algorithms for random combination to obtain integrated algorithms with different combinations. Fit the prediction model based on the training cohort TCGA-CRC. Use GSE17536 as the validation set to validate the prediction model and calculate the consistency index C-index in the two datasets to reflect the discriminative ability of the prediction model based on different integrated algorithms. S5. The prediction model with the highest C-index average value was selected as the most effective model for colorectal cancer prognosis based on the integrated ML algorithm. In order to further understand the related miRNAs, multivariate Cox regression analysis was performed.

[0020] When screening immune-related genes that are differentially expressed in samples from colorectal cancer patients and normal subjects, the limmaR package was used to analyze gene expression differences, and 1149 genes with significant differences in expression between the two groups were screened out from 2483 immune-related genes. A heat map of the expression differences between the two groups was drawn for the 99 genes with the most significant differences, as shown in Figure 2. Figure 2A shown.

[0021] In the univariate Cox regression analysis, 45 immune-related genes, including: ADIPOQ, BACH2, BIRC5, BMP5, CCL11, CCL24, CCL28, CD1A, CD1B, CRABP2, CXCL1, CXCL2, CXCL3, DEFA6, EREG, F2RL1, FABP4, GLP2R, GRP, HAMP, IL13RA2, IL20RB, INHBB, LEP, LTB4R, MC1R, NGFR, NOX1, NOX4, NR3C2, NRG1, PGF, PLCG2, PLXNA3, PTH1R, RETNLB, S100P, SCG2, SEMA5B, SLC11A1, SPP1, SSTR2, TPM2, UCN, WNT5A. These 45 immune-related genes were associated with prognosis, among which these three genes (SSTR2, SEMA5B, NRG1) had the strongest correlation, such as Figure 2B shown.

[0022] When constructing the integrated ML algorithm, 10 basic machine learning algorithms were used, including: random survival forest, elastic network, Lasso, Ridge, StepCox, CoxBoost, Cox's partial least squares regression, supervised principal component, generalized boosted regression model and survival support vector machine. 101 integrated ML algorithms were constructed through various combinations, among which the algorithm combination of CoxBoost and Ridge had the strongest correlation, such as Figure 2C shown.

[0023] From the above, we can see that by establishing the most effective prognostic model for colorectal cancer based on the integrated ML algorithm, it can be used as a prognostic evaluation indicator, has the potential value of helping patients stratify and formulate personalized treatment strategies, and has high accuracy.

[0024] Specifically, the validation of the most effective model for colorectal cancer prognosis based on integrated ML algorithm.

[0025] A total of 582 samples and 175 samples were downloaded from the Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) databases as validation sets, respectively. The samples were divided into high-risk and low-risk groups according to the median risk score. The Kaplan-Meier curve analysis was used to show the overall survival rate of patients in the high-risk and low-risk groups, and the survival time and survival status of the two groups of patients were analyzed. Then, the clinical data of 609 colorectal cancer patients were downloaded from the TCGA database, and the independence of the most effective prognostic model for colorectal cancer based on the integrated ML algorithm was verified. Univariate Cox regression analysis and multivariate Cox regression analysis were performed. The samples were divided into high-risk and low-risk groups according to the median risk score, and the receiver operating characteristic (ROC) curve analysis was used to evaluate the accuracy of the risk model in the validation set. Next, the C-index curve was used to compare the correlation between each clinical feature and the risk score in predicting the survival rate of colorectal cancer patients. Then, the clinical case characteristics and risk score were combined and the RMS package in R was used to draw the nomogram to predict the accuracy of the survival rate of colorectal cancer patients.

[0026] like Figure 3 As shown in the figure, the high-risk and low-risk groups were divided according to the risk score of the validation set given by the model; in the validation set, the area under the curve (AUC) showed that the 1-year, 3-year, and 5-year overall survival rates were 0.794, 0.776, and 0.760, respectively, among which the 1-year OS rate was the most accurate, indicating that the optimal prognostic model for colorectal cancer based on the integrated ML algorithm has a high accuracy in predicting the prognosis of colorectal cancer patients. Figure 3 E; Kaplan-Meier survival analysis showed that the overall survival rate of the high-risk group was significantly lower than that of the low-risk group (P<0.05). Figure 3 A and Figure 3 As shown in B, the accuracy of the optimal prognostic model for colorectal cancer based on the integrated ML algorithm is demonstrated.

[0027] In the validation set, the area under the curve (AUC) showed that the 1-year, 3-year, and 5-year overall survival rates were 0.714, 0.731, and 0.695, respectively, among which the 3-year OS rate was the most accurate, indicating that the model has a high accuracy in predicting the prognosis of colorectal cancer patients. Figure 3 C; Kaplan-Meier survival analysis showed that the overall survival rate of the high-risk group was significantly lower than that of the low-risk group (P<0.01). Figure 3 As described in D, the accuracy of the optimal prognostic model for colorectal cancer based on integrated ML algorithm was demonstrated.

[0028] The validation set was divided into high-risk group and low-risk group according to the median risk score; blue represents the low-risk group and red represents the high-risk group; the results show that the higher the patient's risk score, the shorter the patient's survival time, such as Figure 3 E and Figure 3 As shown in F.

[0029] To verify the independence of the prognostic model, multivariate Cox regression analysis was performed on age, sex, tumor stage, and risk score. Figure 3 As shown in C, the results show that age, grade and risk score can be used as independent factors affecting survival; next, by comparing the ROC-AUC method and C-index, as shown in Figure 3 D and Figure 3 As shown in Figure 1, the accuracy of evaluating patient prognosis based on age, gender, tumor stage and risk score was compared, and it was found that the risk score given by the CoxBoost+Ridge algorithm was better than the clinical characteristics, showing the superiority of the optimal prognosis model for colorectal cancer based on the integrated ML algorithm established by the present invention; based on the risk score of the patient's prognostic characteristics and other clinical indicators, a nomogram was constructed to more comprehensively predict the patient's survival rate, as shown in Figure 1. Figure 3 F, and use Figure 3 The correction curve shown in G evaluates its performance. The results show that the predicted total score of the patient is 171 points, and the accuracy of predicting the patient's 1-year, 3-year, and 5-year overall survival is 0.919, 0.815, and 0.653, respectively, indicating that the optimal prognosis model for colorectal cancer based on the integrated ML algorithm has good practicality.

[0030] As can be seen from the above, the AUC values ​​for predicting overall survival at 1 year, 3 years, and 5 years are all greater than 0.65, which makes the optimal prognosis model for colorectal cancer based on the integrated ML algorithm of the present invention have good sensitivity and specificity.

[0031] More specifically, the validation method includes immune infiltration analysis, TMB (Tumor Mutational Burden) and TIDE (Tumor Immune Dysfunction and Exclusion) analysis, gene pathway enrichment analysis and drug sensitivity analysis.

[0032] Immune infiltration analysis: The immunedeconv package in R was used to integrate XCELL, TIMER, QUANTISEQ, MCPCOUNTER, EPIC, CIBERSORT-ABS and CIBERSORT algorithms to analyze the correlation between risk score and immune cell expression, and the results were presented in the form of bubble charts. The RColorBrewer package in R was used for correlation analysis of immunophenotyping. The BiocManager package in R was used to analyze the differences in 29 immune-related functions between high-risk and low-risk groups, and the differences between high-risk and low-risk groups were presented in box plots. The Estimate package in R was used to calculate the immune score (Immune Score), stromal score (Stromal Score) and total score (Estimate Score) of the high-risk and low-risk groups, respectively, and the differences between the high-risk and low-risk groups were presented in violin plots.

[0033] To further understand immune infiltration analysis, combined with Figure 4 As shown in the figure, the patient scores were calculated according to the optimal prognostic model for colorectal cancer based on the integrated ML algorithm, and the patients were divided into high-risk and low-risk groups. The correlation between the expression of different subpopulations of tumor-infiltrating immune cells and the risk score was evaluated, and the results were visualized with bubble charts. The analysis results showed that there was an obvious negative regulatory relationship between B cells, neutrophils and CD4+T cell memory cells and risk scores, as shown in Figure 2. Figure 4 As shown in A; in the verification of immune typing, patients were divided into six immune subgroups: C1-wound healing type, C2-IFN-γ dominant type, C3-inflammatory type, C4-lymphocyte depletion type, C5-immune silent type, and C6-TGF-β dominant type. The results showed that there was no significant correlation between immune typing and risk score, as shown in Figure 4 As shown in B; a differential analysis of 29 immune-related functions was performed, and the heat map visualization results showed that most of these immune-related functions showed moderate to high significant differences between the high-risk and low-risk groups, such as Figure 4 As shown in C; in the violin plots of stromal cell score, immune cell score and ESTIMATE score, each score has significant differences in patients with high and low risk groups. The ESTIMATE score of patients in the high-risk group is significantly lower than that of patients in the low-risk group, indicating a poor prognosis. Figure 4 As shown in D.

[0034] TMB (Tumor Mutational Burden) and TIDE (Tumor Immune Dysfunction and Exclusion) analysis: The BiocManager package in R was used to analyze the differences in immune checkpoint-related gene expression between high-risk and low-risk groups, and the results were presented in the form of box plots. The website http: / / tide.dfci.harvard.edu / login / was used to analyze the differences in TIDE scores between high-risk and low-risk groups, and to evaluate the differences in immunotherapy and immune escape among patients, and the results were presented in the form of violin plots. The TCGAbiolinks package in R was used to organize the mutation data of the samples and calculate TMB. The TMB differences between high-risk and low-risk groups were presented in the form of violin plots. The correlation between risk score and TMB was evaluated by linear regression, and the ggpubr package in R was used to visualize the results.

[0035] To further understand TMB (Tumor Mutational Burden) and TIDE (Tumor Immune Dysfunction and Exclusion) analysis, combined with Figure 4 and Figure 5 As shown in the figure, the patient's score was calculated according to the constructed risk score model, and the patients were divided into high-risk and low-risk groups. The difference in the expression of 29 immune checkpoint genes in the high-risk and low-risk groups was analyzed, and the box plot showed that the expression of these immune checkpoint genes was moderately to highly correlated with the risk score as a whole, as shown in the figure. Figure 4 As shown in E; 701 colorectal cancer patients were obtained from the TCGA database for TIDE analysis. The violin plot showed that the TIDE score of patients in the high-risk group was significantly higher than that of patients in the low-risk group, indicating that the tumors in the high-risk group have a stronger immune escape ability and the immunotherapy effect is poor. Figure 4 As shown in F; the difference in TMB between high-risk and low-risk groups and its correlation with risk scores were analyzed. The violin plot showed that the TMB of patients in the high-risk group was significantly higher than that of patients in the low-risk group, indicating better antigenicity, while the linear regression plot showed that there was a significant positive correlation between TMB and risk scores, as shown in Figure 5 C and Figure 5 As shown in D.

[0036] Gene pathway enrichment analysis: KEGG (Kyoto Encyclopedia of Genes and Genomes) and GO (Gene Ontology) analysis was performed on the samples using the clusterProfiler package in R, and gene set enrichment analysis was performed using GSEA software.

[0037] To further understand gene pathway enrichment analysis, combined with Figure 5As shown, the patient's score is calculated according to the constructed risk scoring model to divide the patient into high-risk and low-risk groups; gene pathway enrichment analysis is performed to observe the enrichment of these differentially expressed genes; as shown Figure 5 GO analysis in GSEA shown in A showed that the pathways related to wound healing, blood microparticles, fibrocytes, platelet α, and platelet α granule cavity were up-regulated in the high-risk group, while the pathways related to immunoglobulin production, immune response molecular mediator production, immune complex, T cell receptor complex, and antigen binding were up-regulated in the low-risk group; Figure 5 KEGG analysis in GSEA shown in B showed that complement and coagulation cascades, cytochrome P450 in drug metabolism, PPAR signaling pathway, prion-related pathway and tyrosine-related pathway were upregulated in the high-risk group, while hematopoietic factor-related pathway, cytokine-cytokine receptor interaction-related pathway, hematopoietic cell lineage pathway, intestinal immune network-related pathway and starch and sucrose metabolism pathway were upregulated in the low-risk group.

[0038] Drug sensitivity analysis: The BiocManager package in R was used to analyze the differences in sensitivity to different drugs between high-risk and low-risk groups. p<0.05 indicated a significant difference, and the results were presented in the form of box plots.

[0039] To further understand drug sensitivity analysis, combined with Figure 5 As shown in the figure, 623 colorectal cancer patients were downloaded from the TCGA database, and the patient scores were calculated according to the constructed risk score model, so as to divide the patients into high-risk and low-risk groups; the box plot visualization results showed that the sensitivity of drugs AZD-1332, ERK2440, IGF1R_3801, Linsitinib, OSI-027, PRIMA-1MET, WIKI4 and XAV939 was higher in the low-risk group. The sensitivity of Bortezomib, AGI-5198, AZD-6482, AZD-8055, Cyclophosphamide, Erlotinib, GSK343, GSK591, GSK269962A, IAP_5620, KU-55933, MIRA-1, Ribociclib, Savolitinib, SB216763, SCH772984 and Temozolomide was higher in the high-risk group.

[0040] From the above, we can see that using the validation set to validate the model can more accurately and intuitively understand the independent prognostic value of the risk model and the results of systematic analysis of immune infiltration, TMB and TIDE, gene pathway enrichment and drug sensitivity.

[0041] In summary, by constructing an optimal prognostic model for colorectal cancer based on an integrated ML algorithm, the Kaplan-Meier curve and receiver operating characteristic (ROC) curve results confirmed that this risk feature has an accurate predictive value for the prognosis of CRC patients; and the prognostic evaluation ability of the optimal prognostic model for colorectal cancer based on an integrated ML algorithm has been verified in the test data set; the results of tumor immune-related function analysis showed that most of the 29 immune-related immune functions in the high-risk and low-risk groups were significantly different in the high-risk and low-risk groups; TMB analysis showed that TMB was significantly positively correlated with the immune score; TIDE analysis showed that there was a significant difference in TIDE between the high-risk and low-risk groups, and the TIDE of patients in the high-risk group was higher than that of patients in the low-risk group.

[0042] The embodiments of the present invention are provided for the purpose of illustration and description. Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations of the present invention. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present invention.

Claims

1. A method for constructing an optimal prognosis model for colorectal cancer based on an integrated ML algorithm, characterized in that: The following steps are involved: S1. First, obtain and download the transcriptome data of colorectal cancer patient samples and healthy human samples from the TCGA database, then obtain and download the transcriptome data of colorectal cancer patient samples from the GEO database, and then obtain immune-related genes from IMMPORT; S2. Then, the limma R package was used to analyze the expression differences of immune-related genes in samples from colorectal cancer patients and healthy people, and 1,149 genes with significant differences in expression between the two groups were screened out; S3, then univariate Cox regression was used to identify 45 prognosis-related genes from the differentially expressed immune-related genes; S4. Combine multiple ML algorithms for random combination to obtain integrated algorithms with different combinations. Fit the prediction model based on the training cohort TCGA-CRC. Use GSE17536 as the validation set to validate the prediction model and calculate the consistency index C-index in the two datasets to reflect the discriminative ability of the prediction model based on different integrated algorithms. S5. Select the prediction model with the highest C-index average value as the most effective model for colorectal cancer prognosis based on the integrated machine learning algorithm.

2. The method according to claim 1, characterized in that: In step S3, the 45 prognostic genes include: ADIPOQ, BACH2, BIRC5, BMP5, CCL11, CCL24, CCL28, CD1A, CD1B, CRABP2, CXCL1, CXCL2, CXCL3, DEFA6, EREG, F2RL1, FABP4, GLP2R, GRP, HAMP, IL13RA2, IL20RB, INHBB, LEP, LTB4R, MC1R, NGFR, NOX1, NOX4, NR3C2, NRG1, PGF, PLCG2, PLXNA3, PTH1R, RETNLB, S100P, SCG2, SEMA5B, SLC11A1, SPP1, SSTR2, TPM2, UCN, and WNT5A.

3. The method according to claim 1, characterized in that: In step S4, the combination of multiple ML algorithms is performed randomly, and the multiple ML algorithms include: 10 machine learning algorithms: random survival forest, elastic network, Lasso, Ridge, StepCox, CoxBoost, Cox's partial least squares regression, supervised principal component, generalized boosted regression model and survival support vector machine, which constitute 101 integrated algorithms through multiple combinations.

4. The method according to claim 1, characterized in that: The prediction model with the highest C-index average value is CoxBoost+Ridge, which contains 20 immune-related prognostic genes.

5. The method according to claim 4, characterized in that: The 20 immune-related prognostic genes include: CD1A, CD1B, HAMP, S100P, FABP4, CCL24, PLCG2, SEMA5B, PLXNA3, EREG, GRP, INHBB, NRG1, PGF, RETNLB, UCN, GLP2R, IL20RB, MC1R and PTH1R.

6. The method according to claim 1, characterized in that: After step S5, the method further includes: verifying the most effective model for colorectal cancer prognosis based on the integrated machine learning algorithm.

7. The method according to claim 6, characterized in that: Methods to validate the most effective prognostic model for colorectal cancer based on integrated machine learning algorithms include survival analysis, gene pathway enrichment analysis, immune correlation analysis, immune infiltration analysis, and drug sensitivity analysis.

8. An application of the most effective model for colorectal cancer prognosis based on integrated ML algorithm as described in any one of claims 1 to 7.