Biomarkers, models and their applications for predicting the prognosis risk of colorectal cancer
By constructing a prognostic risk prediction model for colorectal cancer, using gene expression analysis and mathematical modeling, the problem of prognostic risk prediction in colorectal cancer patients is solved, and the prediction of survival time and sensitivity of chemotherapy drugs is achieved, and personalized treatment is supported.
Patent Information
- Application Number
- CN202211356031.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-01
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-11-01
AI Technical Summary
The lack of effective biomarkers and models for prognostic risk prediction in patients with colorectal cancer in the prior art, resulting in limited effectiveness of personalized therapeutic interventions and difficult to predict sensitivity of chemotherapy drugs.
A colorectal cancer prognosis risk prediction model was constructed. By obtaining expression profile data from GEO and TCGA databases, differential expression analysis was performed, prognostic related genes were screened out, and gene composition model was determined using LASSO Cox regression analysis, and risk scores were calculated to evaluate the predictive performance of the model.
Effectively identify genes with prognostic value, predict survival time and chemotherapy drug sensitivity in patients with colorectal cancer, and help clinicians design personalized therapeutic interventions.
Smart Images

Figure CN116030880B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biomedical technologies, and particularly to a biomarker, a model and their applications for predicting the prognosis risk of colorectal cancer. Background Art
[0002] Colorectal cancer is one of the most common and worst-prognosis gastrointestinal malignancies at present, accounting for 10% of the annual cancer-related deaths. It is estimated that by 2030, there will be more than 2.2 million new cases of colorectal cancer, and the incidence of colorectal cancer will increase by 90% among young people. Colorectal cancer patients usually have changes in bowel habits, blood in the stool, weight loss, etc. in the early stage, but lack very obvious clinical symptoms. Therefore, colorectal cancer patients often are in the advanced stage when diagnosed and treated. Research shows that colorectal cancer is characterized by easy recurrence and metastasis, and colorectal cancer cells have strong proliferative ability. With the in-depth research, although the 5-year survival rate of colorectal cancer patients with early diagnosis reaches more than 90%, the 5-year survival rate of metastatic colorectal cancer patients is very low, only 12%.
[0003] At present, although conventional treatments such as surgical resection, chemotherapy, radiotherapy, and targeted therapy have reduced the mortality rate of colorectal cancer patients, metastasis is still the main cause of death for most colorectal cancer patients. In addition, due to individual heterogeneity, the clinical outcomes of different patients vary greatly, which limits the effectiveness of traditional treatment methods. With the rapid development of multi-omics technologies, prognosis models based on gene expression have become important biomarkers for screening cancer patients with different clinicopathological risks. Therefore, developing new multi-gene risk prediction models can provide valuable information for the identification of high-risk colorectal cancer patients, thereby helping clinicians monitor disease progression and design appropriate personalized treatment intervention measures. Summary of the Invention
[0004] To this end, the technical problem to be solved by the present invention is to overcome the problems existing in the prior art, and propose a biomarker, a model and their applications for predicting the prognosis risk of colorectal cancer, which can bring new strategies for predicting the survival time of colorectal cancer patients and the sensitivity to chemotherapeutic drugs, and help clinicians monitor disease progression and design appropriate personalized treatment intervention measures.
[0005] To solve the above technical problem, the present invention provides a prognosis risk prediction model for colorectal cancer, and the prognosis risk prediction model is constructed through the following steps:
[0006] S1: Obtain the expression profile data and clinical information data of colorectal cancer patients from the GEO and TCGA databases. These two types of data are used for the construction and verification of the prognosis risk prediction model, and the data include a training set and a validation set. At the same time, obtain the expression profile data of cancer tissue samples and normal control tissue samples from the GEO database;
[0007] S2: Perform differential expression analysis on the gene expression matrices of the cancer tissue samples and normal control tissue samples after standardization to obtain differentially expressed genes.
[0008] S3: Perform univariate Cox regression and PH test analysis on the differentially expressed genes in the training set to screen out prognosis-related genes.
[0009] S4: Perform LASSO Cox regression analysis on the screened genes, determine the optimal λ value through cross-validation, screen out the genes that make up the prognosis risk prediction model, and calculate the risk scores of each patient in each dataset based on the regression coefficients of each screened gene and its standardized expression levels in the training set and validation set.
[0010] S5: Evaluate the prediction performance of the prognosis risk prediction model according to the risk scores.
[0011] In one embodiment of the present invention, in S3, the method of performing univariate Cox regression and PH test analysis on the differentially expressed genes in the training set to screen out prognosis-related genes includes:
[0012] Delete the samples with missing overall survival time and survival outcome in the prognosis information of the training set.
[0013] In the univariate Cox regression analysis, use the standardized gene expression matrix of the differentially expressed genes in the training set, the overall survival time of the samples, and the survival outcome information as input files for univariate Cox regression analysis, and screen out the genes that are significantly related to the sample prognosis time, meet the PH regression hypothesis, and have a p-value less than the set value.
[0014] In one embodiment of the present invention, in S4, the genes include INHBA, GALNT6, ZNF239, PDE6A, PTPRR, HSD17B2, CD177, CDKN2B, MALL, LGALS2, BTNL8, SCG2, and PCOLCE2.
[0015] In one embodiment of the present invention, the regression coefficients corresponding to the genes INHBA, GALNT6, ZNF239, PDE6A, PTPRR, HSD17B2, CD177, CDKN2B, MALL, LGALS2, BTNL8, SCG2, and PCOLCE2 are 0.008048177, -0.014590776, 0.169851027, -0.141714084, 0.080175473, 0.286836794, -0.115094672, 0.070682084, 0.044238906, -0.16681854, -0.071504884, 0.281751541, and 0.070570821, respectively.
[0016] In one embodiment of the present invention, in S4, the calculation formula of the risk score is:
[0017]
[0018] where Coef represents the regression coefficient of each gene in the model, E represents the normalized expression level of each gene in the model, i represents the index of the gene in the model, and n represents the number of genes included in the model.
[0019] In one embodiment of the present invention, in S5, the method for evaluating the prediction performance of the prognostic risk prediction model according to the risk score includes:
[0020] Performing a multivariate Cox regression analysis with the risk scores and clinical indicators of the validation set and the training set as variables to determine whether the risk score based on the prognostic risk prediction model can be used as an independent prognostic factor for colorectal cancer. At the same time, perform a KM survival curve and time-dependent ROC curve analysis based on the risk score to evaluate the prediction performance of this prognostic risk prediction model for the overall survival time of colorectal cancer patients.
[0021] In one embodiment of the present invention, the method for evaluating the prediction performance of this prognostic risk prediction model for the overall survival time of colorectal cancer patients by performing a KM survival curve and time-dependent ROC curve analysis based on the risk score includes:
[0022] Arrange the samples in ascending or descending order according to the risk score, determine the median of the samples, take the samples with a risk score greater than the median as the high-risk group, and take the samples with a risk score less than the median as the low-risk group;
[0023] Create a survival object with the prognostic information of the samples in the training set and the validation set and fit the survival curve, draw a visual KM survival curve, and use the log-rank statistical test to analyze the difference in the overall survival probability between the high-risk group and the low-risk group of colorectal cancer patients;
[0024] Select the survival rates of multiple years in the sample information and prognostic information of the training set and validation set for time-dependent ROC analysis, draw the curve, calculate the area value under the curve and its confidence interval, and evaluate the prediction performance of the prognostic risk prediction model for the overall survival time of colorectal cancer patients based on the area value.
[0025] In addition, the present invention also provides a biomarker for predicting the prognostic risk of colorectal cancer, and the biomarker includes genes INHBA, GALNT6, ZNF239, PDE6A, PTPRR, HSD17B2, CD177, CDKN2B, MALL, LGALS2, BTNL8, SCG2, and PCOLCE2, which are used to construct a colorectal cancer prognostic risk prediction model as described above.
[0026] Furthermore, the present invention also provides an application of the colorectal cancer prognostic risk prediction model as described above in gene ontology analysis, immune correlation analysis, and chemotherapy drug response analysis.
[0027] The above technical solutions of the present invention have the following advantages compared with the prior art:
[0028] The biomarker, model, and its application for predicting the prognostic risk of colorectal cancer according to the present invention construct a prognostic risk prediction model for colorectal cancer based on mathematical modeling and bioinformatics methods, can effectively identify genes with prognostic value, and thus can bring new strategies for predicting the survival time of colorectal cancer patients and the sensitivity to chemotherapy drugs, helping clinicians monitor the disease progression and design appropriate personalized treatment intervention measures. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to make the content of the present invention easier to be clearly understood, the following further describes the present invention in detail according to the specific embodiments of the present invention and in combination with the drawings, wherein
[0030] Figure 1 is a flowchart of a method for constructing a prognostic risk prediction model for colorectal cancer proposed by the present invention.
[0031] Figure 2 is a diagram for identifying the common differentially expressed genes in three datasets, where A-C represent the distribution of differentially expressed genes in the three datasets; D-E represent the intersection of up-regulated genes and the intersection of down-regulated genes in the three datasets.
[0032] Figure 3 is a forest plot showing the results of univariate Cox regression analysis.
[0033] Figure 4This is a graph showing the performance evaluation results of the prognostic risk prediction model for colorectal cancer. Among them, A shows that colorectal cancer patients are divided into high-risk and low-risk groups according to the distribution of risk scores; the scatter plot in B shows the distribution of the survival status of colorectal cancer patients in different risk groups; the heat map in C shows the expression levels of model-related genes in patients in the high- and low-risk groups; D shows that the KM curve shows the survival of the high-risk and low-risk groups; E shows that the time-dependent ROC curve shows the ability of the model to predict the overall survival time of colorectal cancer patients.
[0034] Figure 5 This is a comparative analysis graph of risk scores for patients with different tumor pathological grades and cancer stages in the training set and validation set.
[0035] Figure 6 This is a KM survival curve graph showing the predictive ability of the constructed model in different clinical subgroups of the training set GSE17538.
[0036] Figure 7 This is a KM survival curve graph showing the predictive ability of the constructed model in different clinical subgroups of the validation set GSE29621.
[0037] Figure 8 This is a KM curve graph showing the difference in overall survival (OS) between the high-risk and low-risk groups in the validation set.
[0038] Figure 9 This is a KM curve graph showing the difference in disease-specific survival (DSS) between the high-risk and low-risk groups in the validation set.
[0039] Figure 10 This is a KM curve showing the difference in disease-free survival (DFS) between the high-risk and low-risk groups in the validation set.
[0040] Figure 11 This is an ROC curve graph showing the change of the prognostic performance of the model over time in the OS validation set.
[0041] Figure 12 This is an ROC curve graph showing the change of the prognostic performance of the model over time in the DSS validation set.
[0042] Figure 13 This is an ROC curve graph showing the change of the prognostic performance of the model over time in the DFS validation set.
[0043] Figure 14 This is a forest plot showing the results of the multivariate Cox regression analysis in the training set and validation set. Specific implementation method
[0044] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it. However, the exemplified embodiments are not intended to limit the present invention.
[0045] Referring to Figure 1 As shown, the embodiment of the present invention provides a prognostic risk prediction model for colorectal cancer. The prognostic risk prediction model is constructed through the following steps:
[0046] S1: Obtain the expression profile data and clinical information data of colorectal cancer patients from the GEO and TCGA databases. These two types of data are used for the construction and verification of the prognostic risk prediction model. The data includes a training set and a validation set. At the same time, obtain the expression profile data of cancer tissue samples and normal control tissue samples from the GEO database;
[0047] S2: Perform differential expression analysis on the standardized gene expression matrices of cancer tissue samples and normal control tissue samples to obtain differentially expressed genes;
[0048] S3: Perform univariate Cox regression and PH test analysis on the differentially expressed genes in the training set to screen out prognosis-related genes;
[0049] S4: Perform LASSO Cox regression analysis on the screened genes, determine the optimal λ value through cross-validation, screen out the genes that make up the prognostic risk prediction model, and calculate the risk scores of each patient in each dataset according to the regression coefficients of each screened gene and its standardized expression levels in the training set and the validation set;
[0050] S5: Evaluate the prediction performance of the prognostic risk prediction model according to the risk scores.
[0051] The above-mentioned term "prognosis" refers to the disease development status predicted by doctors based on various characteristics of patients, such as age, gender, test results, etc. In clinical practice, patients with the same disease, due to individual heterogeneity, often lead to different clinical events, such as recovery, other diseases, inflammatory reactions, pain, etc. Among them, death is the most serious clinical event. Therefore, through prognostic evaluation, patients can not only understand their own condition more clearly to cooperate with clinical treatment, but also doctors can make a more definite and personalized diagnosis of the patient's condition.
[0052] A prognostic risk prediction model for colorectal cancer according to the present invention constructs a prognostic risk prediction model for colorectal cancer based on mathematical modeling and bioinformatics methods, and can effectively identify genes with prognostic value, thereby bringing new strategies for predicting the survival time of colorectal cancer patients and the sensitivity to chemotherapeutic drugs, and helping clinicians monitor disease progression and design appropriate personalized treatment intervention measures.
[0053] Among them, in S1, the gene expression profiles and clinical information data of GSE113513, GSE89076, GSE39582, GSE17538, GSE29621 and GSE38832 can be downloaded from the GEO database. Among them, GSE17538 is used as the training set, containing sample information of 238 patients. GSE39582, GSE29621 and GSE38832 are used as validation sets, containing sample information of 579 patients, 65 patients and 122 patients respectively. At the same time, the data containing 377 colorectal cancer patients downloaded from the TCGA database is also used as a validation set. The sample information and gene expression matrix corresponding to each data are exported by using the GEO convenient converter. The data of the three groups of GSE113513, GSE89076 and GSE39582 are classified according to normal samples and tumor samples for the following analysis.
[0054] Among them, in S2, the method for performing differential expression analysis on the standardized gene expression matrix includes reading the standardized gene expression matrix, and using the "limma" R package to perform statistical analysis of differential expression on the data of the standardized expression matrix read from GSE113513, GSE89076, and GSE39582. Taking |log2(FC)|≥1.5 and adj.P.Val<0.05 as the criteria, genes that are significantly differentially expressed (differentially expressed genes, DEGs) between tumor samples and normal control samples are screened out. Further, the "limma" package in step S2 is a standard package for analyzing microarray data in R / Bioconductor, which has a complete set of processes for analyzing microarrays and can overcome the problem of small sample size. The Benjamini-Hochberg false discovery rate method, that is, the expected value of the proportion of the number of true hypotheses rejected to the number of all rejected original hypotheses, is a multiple comparison correction method. The present invention uses the Benjamini-Hochberg false discovery rate to adjust the original p-value (p value).
[0055] Among them, in S3, the method of performing univariate Cox regression and PH test analysis on the differentially expressed genes in the training set to screen out the genes significantly related to the overall survival rate of colorectal cancer patients includes deleting the samples with missing overall survival time and survival outcome in the training set prognosis information; in the univariate Cox regression analysis, the standardized gene expression matrix, the overall survival time and survival outcome information of the samples of the differentially expressed genes in the validation set are used as input files for univariate Cox regression analysis to screen out the genes significantly related to the sample prognosis time that meet the PH regression hypothesis and have a p-value less than the set value. Preferably, the coxph function and cox.zph function of the "survival" package in R language can be used to perform univariate Cox regression analysis on the common differentially expressed genes of GSE113513, GSE89076 and GSE39582 in the training set GSE17538, so as to screen out the genes significantly related to the overall survival rate of colorectal cancer patients. Specifically, in order to screen reliable prognosis-related genes, the samples with missing overall survival (OS) time and survival outcome in the training set prognosis information should be deleted. In the univariate Cox regression analysis, the expression matrix data standardized by the common differentially expressed genes of GSE113513, GSE89076 and GSE39582 in GSE17538 and the prognosis data including OS time and survival outcome are used as input files for batch univariate Cox regression analysis to screen out the genes significantly related to the sample prognosis time that meet the PH regression hypothesis and have a p-value less than 0.05.
[0056] Furthermore, the "survival" R package in step S3 was first written and applied to survival analysis in 1985, and it has various functions, including the description of survival data, hypothesis testing, the construction of cox proportional hazards regression models, and the construction of competing risk models, etc. When there are multiple independent variables, the significantly significant independent variables can be screened out through univariate Cox regression analysis.
[0057] Among them, in S4, "glmnet" is an R package for fitting generalized linear models.. LASSO Cox regression analysis is a new variable selection and shrinkage method that minimizes the logarithmic partial likelihood, shrinks the coefficients and produces some exactly zero coefficients, thus reducing the estimation variance, and at the same time can obtain a reliable final model. LASSO Cox regression analysis determines the genes that make up the model while selecting the best λ value through cross-validation, and the final regression model is obtained by extracting the regression coefficients through Coef.
[0058] Further, the genes included in the prognostic risk prediction model in step S4 are INHBA, GALNT6, ZNF239, PDE6A, PTPRR, HSD17B2, CD177, CDKN2B, MALL, LGALS2, BTNL8, SCG2, and PCOLCE2, and their corresponding regression coefficients are 0.008048177, -0.014590776, 0.169851027, -0.141714084, 0.080175473, 0.286836794, -0.115094672, 0.070682084, 0.044238906, -0.16681854, -0.071504884, 0.281751541, and 0.070570821 respectively; the risk score of each patient is calculated based on the expression level of each relevant gene participating in the model construction and its corresponding regression coefficient, and the calculation formula of the risk score is:
[0059]
[0060] where Coef represents the regression coefficient of each gene in the model, E represents the standardized expression level of each gene in the model, i represents the index of the gene in the model, and n represents the number of genes included in the model.
[0061] Among them, in S5, the method for evaluating the prediction performance of the prognostic risk prediction model according to the risk score includes performing a multivariate Cox regression analysis with the risk scores and clinical indicators of the validation set and the training set as variables to determine whether the risk score based on the prognostic risk prediction model can be used as an independent prognostic factor for colorectal cancer. At the same time, perform a KM survival curve and time-dependent ROC curve analysis according to the risk score to evaluate the prediction performance of this prognostic risk prediction model for the overall survival time of colorectal cancer patients. Preferably, the samples are sorted in ascending or descending order according to the risk score, the median of the samples is determined, the samples with a risk score greater than the median are used as the high-risk group, and the samples with a risk score less than the median are used as the low-risk group; create a survival object with the prognostic information of the samples in the training set and the validation set and fit the survival curve to draw a visual KM survival curve to analyze the difference in the overall survival probability between the high-risk group and the low-risk group of colorectal cancer patients; select the sample information and prognostic information of the training set and the validation set for time-dependent ROC analysis for multiple years and draw a curve, calculate the area value under the curve and its confidence interval, and evaluate the prediction performance of this prognostic risk prediction model for the overall survival time of colorectal cancer patients based on the area value.
[0062] As an example, in step S5, a multivariate Cox regression analysis was performed using the risk scores calculated based on the model and common clinical indicators as variables to evaluate whether the constructed model has independent prognostic value. Specifically, the R language "survival" package can be used to construct survival functions with the following data respectively, and the model genes are verified multiple times. The risk scores and clinical indicators of the validation set and the training set are used as variables for multivariate Cox regression analysis to obtain the hazard ratio, its confidence interval, and p-value of each variable, and then to determine whether the risk score based on the prognostic risk prediction model can be used as an independent prognostic factor for colorectal cancer. The R language "forestplot" package is used to create a forest plot to display the results of the multivariate Cox regression. The data used includes the OS survival index datasets (GSE29621, GSE39582, and TCGA), the DSS survival index datasets (GSE17538 and GSE38832), and the DFS survival index datasets (GSE17538 and GSE29621).
[0063] As an example, in step S5, the risk score of each colorectal cancer patient was calculated according to the standardized expression levels of the genes in the model and their corresponding regression coefficients, and the median was found by arranging the risk scores of the samples in descending order. Samples with a risk score greater than the median were used as the high-risk group, while samples with a risk score less than the median were used as the low-risk group. The R language "survival" package was used to create a survival object with the sample prognostic information and fit the survival curve, and then the R language "survminer" package was used to plot the visual Kaplan-Meier (KM) survival curve. The difference in the overall survival probability between the high- and low-risk groups of colorectal cancer patients was analyzed by log-rank statistical test (p < 0.05). Further, based on the sample information and prognostic information of the training set and the validation set, the R language "timeROC" package was used to select the survival rates at 1 year, 3 years, and 5 years for time-dependent ROC analysis and plot the curve, calculate the area under the curve (AUC) value and its confidence interval, and then the R language "ggplot2" package was used to beautify the graph, and the prediction performance of the prognostic risk prediction model for the overall survival time of colorectal cancer patients was evaluated based on the AUC value.
[0064] Further, the KM curve in step S5 presents the information from the time of the test to the prognostic event in a visual way, and it can be known the probability of a specific subject surviving after a specific time. The "survminer" package is the most commonly used package in bioinformatics to implement the drawing of survival analysis curves. It can not only generate the curve neatly, but also give parameters such as p-value and risk value.
[0065] The following describes in detail a method for constructing a prognostic risk prediction model for colorectal cancer proposed by the present invention in specific embodiments, which includes the following:
[0066] 1. Identification of differentially expressed genes in the colorectal cancer dataset: Based on the gene chip data in the GEO database, the present invention performs differential expression analysis on the expression matrices of colorectal cancer tumor tissue samples and normal control tissue samples in the GES39582, GSE89076, and GSE113513 datasets. Among them, the GSE113513 dataset contains 14 cancer tissue samples and 14 normal tissue samples, the GSE89076 dataset contains 39 cancer tissue samples and 39 normal tissue samples, and the GSE39582 dataset contains 566 cancer tissue samples and 19 normal tissue samples. Differential expression analysis is performed on the expression profile data of cancer tissue samples and normal tissue samples in the three datasets respectively, deleting non-specific genes, that is, the situation where multiple genes correspond to the same probe, and screening out 643 differentially expressed genes from the GSE39582 dataset with the criteria of |log2(FC)|≥1.5 and adj.P.Val<0.05, including 218 up-regulated genes and 425 down-regulated genes ( Figure 2 A), screening out 1535 differentially expressed genes from the GSE89076 dataset, including 545 up-regulated genes and 990 down-regulated genes ( Figure 2 B), and screening out 455 differentially expressed genes from the GSE113513 dataset, including 136 up-regulated genes and 319 down-regulated genes ( Figure 2 C). To ensure the reliability of the differentially expressed genes of colorectal cancer in the present invention, the present invention selects the intersection of the differentially expressed genes in the three datasets of GSE113513, GSE89076, and GSE39582 for subsequent analysis, that is, including 59 common up-regulated genes and 160 common down-regulated genes ( Figure 2 D and Figure 2 E), and a total of 219 common differentially expressed genes are used for subsequent analysis.
[0067] 2. Univariate Cox regression analysis of differentially expressed genes: As Figure 2As shown, first, the "limma" package in R language was used to perform statistical analysis of differential expression on the standardized expression matrix data read from GSE113513, GSE89076, and GSE39582. Among them, the normal tissue samples and cancer tissue samples in GSE89076 and GSE113513 were paired samples, so pairing information was added to perform differential expression analysis of paired samples. And with |log2(FC)|≥1.5 and adj.P.Val<0.05 as the criteria, genes with significant differential expression were screened out. Then, the genes with significant differential expression identified in the GSE113513, GSE89076, and GSE39582 datasets were respectively selected, and were divided into up-regulated genes and down-regulated genes according to their expression fold changes. The Venn diagram analysis tool was used to take the intersection of the up-regulated genes and down-regulated genes in the three datasets of GSE113513, GSE89076, and GES39582 respectively to make a Venn diagram, and the differentially expressed genes common to the three datasets were obtained for the following analysis.
[0068] 3. Construction of colorectal cancer prognosis model: LASSO Cox regression analysis was performed on the 31 differentially expressed genes obtained by univariate Cox regression analysis to construct a penalty function and further compress the coefficients of the variables in the prognosis risk prediction model to obtain the optimal λ value and the regression coefficient of each gene. Finally, 13 genes of the model were determined, and the risk scores of each patient in each dataset were calculated according to the regression coefficient of each selected gene and its standardized expression level in the training set and validation set. As Figure 3 shown, the 13 genes include INHBA, GALNT6, ZNF239, PDE6A, PTPRR, HSD17B2, CD177, CDKN2B, MALL, LGALS2, BTNL8, SCG2, and PCOLCE2.
[0069] Using the median of the risk scores as the boundary, the samples in the dataset were divided into high-risk and low-risk groups, and the survival information of the high-risk group patients and low-risk group patients in the GSE17538 dataset was observed. As Figure 4 shown in A - C, the OS survival time and survival status of the low-risk group patients were better than those of the high-risk group patients, and the expression of the 13 prognosis genes was significantly different between the high-risk and low-risk groups. At the same time, the results of the KM survival curve showed that in the GSE39582 dataset, there was a significant statistical difference in OS between the high-risk and low-risk group patients (p<0.0001)( Figure 4 D). In addition, a time-dependent ROC curve was drawn to evaluate the prediction effect of the model, and the calculated area under the curve AUC value was 0.755 at 1 year, 0.776 at 3 years, and 0.736 at 5 years( Figure 4 E).
[0070] 4. Analysis and Validation of the Prognostic Risk Prediction Model for Colorectal Cancer:
[0071] (1) Prognosis analysis of the model: To obtain the relationship between the constructed prognostic risk prediction model and various clinical manifestations, the risk scores of each group were compared under different classification conditions such as age, gender, tumor pathological grade, and cancer stage. As Figure 5 shown, in the training set GSE17538 and the validation set GSE29621, the risk scores of colorectal cancer patients in the advanced TNM stage (Stage 4) were significantly higher than those of colorectal cancer patients in the other stages, that is, the risk score increased with the progression of the patient's cancer tumor; at the same time, in the training set GSE17538 and the validation set GSE29621, the risk scores of patients with low tumor differentiation, that is, PD (Poorly differentiated, PD) grade, were significantly higher than those of patients with high tumor differentiation, that is, WD (Well differentiated, WD) grade, that is, the risk score was negatively correlated with the tumor differentiation degree. It can be seen that the constructed prognostic risk prediction model for colorectal cancer is closely related to the clinical manifestations of tumor pathological grade and cancer stage.
[0072] In addition, the present invention conducts a KM survival analysis by grouping based on the clinical indicators of tumor differentiation degree and tumor TNM stage, and further evaluates the prognostic performance of the prediction model for colorectal cancer patients. As Figure 6 and Figure 7 shown, in the training set GSE17538 and the validation set GSE29621, the survival probability of high-risk colorectal cancer patients in the MD (Moderately differentiated, MD) grade subgroup of tumor differentiation degree was significantly lower than that of low-risk colorectal cancer patients (p value less than 0.05); at the same time, in the training set GSE17538 and the validation set GSE29621, the survival probability of high-risk colorectal cancer patients in the TNM stage III / IV subgroup was significantly lower than that of low-risk colorectal cancer patients (p value less than 0.05). The survival probability of high-risk colorectal cancer patients in the remaining subgroups also showed a trend of being lower than that of low-risk patients. It can be seen that the constructed prognostic risk prediction model for colorectal cancer has good predictability for the prognosis of colorectal cancer patients in the above clinical subgroups.
[0073] (2) Model validation: To verify the credibility of the constructed prediction model, the patients were divided into high- and low-risk groups according to the median risk score of colorectal cancer patients in the validation set. As Figures 8 to 10As shown, it showed good performance in the validation results of OS data (GSE39582, GSE29621, TCGA), DSS data (GSE17538, GSE38832), and DFS data (GSE17538, GSE29621). As the survival time increased, the survival probability of patients in the high-risk group was significantly lower than that of patients in the low-risk group, and the p-values were all less than 0.05, indicating that the risk score calculated based on the prognostic genes of the model had a good effect on predicting the survival time of colorectal cancer patients.
[0074] In addition, from Figure 11 the ROC curve in, the AUC value in the ROC validation result of the OS data of the dataset GSE29621 reached 0.81 at 1 year, and was also greater than 0.7 at 3 years and 5 years. Moreover, the AUC values in the ROC validation results of the OS data of the GSE39582 and TCGA colorectal cancer datasets were all greater than 0.7 at 1 year; from Figure 12 it can be seen that the AUC values in the ROC validation results of the DSS data of the two datasets GSE17538 and GSE38832 were both around 0.8; in the results validated by DFS data, the ROC values in the GSE29621 dataset were greater than 0.9 at 1 year and 5 years, and the AUC value at 3 years also reached 0.89. Moreover, the ROC validation results of the DFS data of GSE17538 were also greater than 0.7 ( Figure 13 ). The above results further illustrate that the prognostic risk prediction model of the present invention also has a good prediction effect on colorectal cancer patients in the validation set.
[0075] Then, continue to verify whether the prognostic performance of the colorectal cancer prognostic risk prediction model is independent of other clinical characteristics. The present invention performed a multivariate Cox regression analysis on the known clinical characteristics including gender, age, and TNM stage in the validation set. It can be seen from the multivariate forest plot that the p-values of the risk score and TNM stage in all datasets were less than 0.05 ( Figure 14 ), indicating that the prognostic genes in the constructed colorectal cancer prognostic risk prediction model were verified in all 8 datasets. Therefore, it was proved that the risk score and TNM stage could be used as independent prognostic factors for colorectal cancer. At the same time, it was also found that the hazard ratios of the risk score and TNM stage in all datasets were greater than 1, indicating that they were negatively correlated with the survival time of patients.
[0076] In addition, corresponding to the above embodiments of constructing a prognostic risk prediction model for colorectal cancer, the embodiments of the present invention further provide a biomarker for prognostic risk prediction of colorectal cancer, and the biomarker includes genes INHBA, GALNT6, ZNF239, PDE6A, PTPRR, HSD17B2, CD177, CDKN2B, MALL, LGALS2, BTNL8, SCG2 and PCOLCE2, which are used to construct a prognostic risk prediction model for colorectal cancer as described above.
[0077] The above-mentioned "biomarker" is a biochemical index, and its basic definition is: "A defined characteristic that can be used as an indicator of normal biological processes and pathogenic processes." Biomarkers can also be used as indicators for prognostic treatment intervention in patients, and they can be derived from molecular, histological, radiological or physiological characteristics. Therefore, biomarkers can not only be used to describe the changes in the structure and function of cells, sub-cells, tissues, organs, etc., but also can be used for disease diagnosis and disease staging, and have a very wide range of uses, and can classify disease risks, prognoses or treatment responses. In addition, prognostic biomarkers can be used to identify the likelihood of clinical events, disease recurrence or disease progression in patients with diseases or related disease conditions, and are particularly important for predicting the risk of individual disease events or adverse outcomes.
[0078] Furthermore, corresponding to the above embodiments of constructing a prognostic risk prediction model for colorectal cancer, the embodiments of the present invention further provide an application of the prognostic risk prediction model for colorectal cancer as described above in gene ontology analysis, immune correlation analysis and chemotherapy drug response analysis, which specifically includes the following three aspects:
[0079] 1. GO and GSEA enrichment analysis: GO enrichment analysis was performed using 148 differentially expressed genes between the high- and low-risk groups of colorectal cancer patients. In terms of biological processes, it was mainly enriched in entries such as extracellular matrix, extracellular structure organization, osteogenesis, and serine / threonine transmembrane receptor protein kinase signaling pathway; in terms of cellular components, it was mainly enriched in entries such as extracellular matrix of collagen, endoplasmic reticulum lumen, and focal adhesion kinase signaling pathway; in terms of molecular functions, it was mainly enriched in entries such as extracellular matrix structural components, receptor-ligand activity, and signal receptor activator activity. In addition, the present invention performed GSEA enrichment analysis on 148 differentially expressed genes between the high- and low-risk groups of colorectal cancer patients to analyze the function of the constructed prognostic risk prediction model. The top ten tumor characteristic pathways significantly enriched in the high-risk group include apoptosis, blood coagulation process, complement signaling, epithelial-mesenchymal transition, KRAS-related signaling, protein transmembrane transport, TGF-β signaling, and TNFα / NF-kB and other pathways.
[0080] 2. Immune - related analysis: Based on the expression levels of immune cells and immune - function - related genes in the training set GSE17538, ssGSEA analysis was performed on each sample to obtain the infiltration scores of immune - related gene sets. The Wilcoxon test was used to examine the correlations between the high - and low - risk groups and immune cell types and immune cell functions. The infiltration scores of multiple immune cells were significantly different between the high - and low - risk groups, such as B cells, CD8+ T cells, dendritic cells (DCs), mast cells, natural killer cells (NK), plasmacytoid DCs (pDCs), helper T cells, and regulatory T cells (Tregs); the infiltration scores of some key immune - related functions were also significantly different between the high - and low - risk groups, such as antigen - presenting cell (APC) co - stimulation, major histocompatibility complex class I, inflammation promotion, T - cell co - inhibition, T - cell co - promotion, and IFN response. In addition, according to Pearson correlation analysis, in the training set GSE17538, the risk score was significantly correlated with the immune score and stromal score, indicating a close relationship between the risk score and the infiltration of immune cells in the tumor microenvironment. Furthermore, an analysis of the expression levels of two common immune checkpoint genes in colorectal cancer immunotherapy was performed between the high - and low - risk groups. In the GSE17538 dataset, the expression of cytotoxic T - lymphocyte - associated protein gene (CTLA - 4) in the high - risk group was significantly lower than that in the low - risk group, while the expression level of TIM - 3 in the high - risk group was significantly higher than that in the low - risk group. Pearson correlation analysis also further indicated that the expression levels of CTLA - 4 and TIM - 3 genes were significantly correlated with the risk score, indicating that the constructed prognostic risk prediction model for colorectal cancer is closely related to the expression of immune checkpoints.
[0081] 3. Analysis of chemosensitivity of colorectal cancer chemotherapy drugs: The IC50 values of 138 chemotherapy drugs for each sample in the training set and test set were calculated using the R package "pRRophetic". The calculation results showed that the IC50 values of 5 chemotherapy drugs were significantly different between the high - and low - risk groups in all datasets (p - values were all less than 0.05), and these 5 compounds were all known chemotherapy drugs for colorectal cancer, including Bexarotene, Embelin, FTI - 277, Pazopanib, and Lestaurtinib. These results indicate that these 5 drugs are more sensitive in high - risk group patients, further demonstrating that the prognostic risk prediction model constructed in the present invention can be used to predict the chemosensitivity of colorectal cancer patients to chemotherapy drugs.
[0082] Obviously, the above embodiments are merely examples given for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. And the obvious changes or modifications derived therefrom still fall within the protection scope of the present invention.
Claims
1. Application of a biomarker in constructing a prognostic risk prediction model for colorectal cancer, characterized in that: The biomarker consists of genes INHBA, GALNT6, ZNF239, PDE6A, PTPRR, HSD17B2, CD177, CDKN2B, MALL, LGALS2, BTNL8, SCG2, and PCOLCE2. Among them, the colorectal cancer prognosis risk prediction model is constructed through the following steps: S1: Obtain the expression profile data and clinical information data of colorectal cancer patients from the GEO and TCGA databases. These two types of data are used for the construction and validation of the prognosis risk prediction model. The data includes a training set and a validation set. At the same time, obtain the expression profile data of cancer tissue samples and normal control tissue samples from the GEO database; S2: Perform differential expression analysis on the standardized gene expression matrices of cancer tissue samples and normal control tissue samples to obtain differentially expressed genes; S3: Perform univariate Cox regression and PH test analysis on the differentially expressed genes in the training set to screen out prognosis-related genes; S4: Perform LASSO Cox regression analysis on the screened genes, determine the optimal λ value through cross-validation, screen out the genes that make up the prognosis risk prediction model, and calculate the risk scores of each patient in each dataset according to the regression coefficients of each screened gene and its standardized expression levels in the training set and validation set; S5: Evaluate the prediction performance of the prognosis risk prediction model according to the risk scores; Among them, the method for evaluating the prediction performance of the prognosis risk prediction model according to the risk scores includes: using the risk scores and clinical indicators of the validation set and training set as variables for multivariate Cox regression analysis to determine whether the risk scores based on the prognosis risk prediction model can be used as independent prognostic factors for colorectal cancer. At the same time, perform KM survival curve and time-dependent ROC curve analysis according to the risk scores to evaluate the prediction performance of this prognosis risk prediction model for the overall survival time of colorectal cancer patients; The method for performing KM survival curve and time-dependent ROC curve analysis according to the risk scores to evaluate the prediction performance of this prognosis risk prediction model for the overall survival time of colorectal cancer patients includes: Arrange the samples in ascending or descending order according to the risk scores, determine the median of the samples, regard the samples with risk scores greater than the median as the high-risk group, and regard the samples with risk scores less than the median as the low-risk group; create survival objects and fit survival curves with the sample prognosis information of the training set and validation set, draw a visual KM survival curve, and use the log-rank statistical test to analyze the difference in the overall survival probability between the high-risk group and low-risk group of colorectal cancer patients; select the sample information and prognosis information of the training set and validation set for time-dependent ROC analysis and draw curves for multiple years of survival rates, calculate the area value under the curve and its confidence interval, and evaluate the prediction performance of this prognosis risk prediction model for the overall survival time of colorectal cancer patients based on the area value; Multivariate Cox regression analysis was performed using the risk scores calculated based on the model and common clinical indicators as variables to evaluate whether the constructed model had independent prognostic value; the "survival" package in R language was used to construct survival functions with the following data respectively, and the model genes were verified multiple times. Multivariate Cox regression analysis was performed with the risk scores and clinical indicators of the validation set and training set as variables to obtain the hazard ratio, its confidence interval, and p-value of each variable, and then to determine whether the risk score based on the prognostic risk prediction model could be used as an independent prognostic factor for colorectal cancer. The "forestplot" package in R language was used to create a forest plot to display the results of the multivariate Cox regression. The data used included the OS survival index dataset, DSS survival index dataset, and DFS survival index dataset; The regression coefficients corresponding to the genes INHBA, GALNT6, ZNF239, PDE6A, PTPRR, HSD17B2, CD177, CDKN2B, MALL, LGALS2, BTNL8, SCG2, and PCOLCE2 were 0.008048177, -0.014590776, 0.169851027, -0.141714084, 0.080175473, 0.286836794, -0.115094672, 0.070682084, 0.044238906, -0.16681854, -0.071504884, 0.281751541, and 0.070570821 respectively; In S4, the calculation formula of the risk score is: where Coef represents the regression coefficient of each gene in the model, E represents the normalized expression level of each gene in the model, i represents the index of the gene in the model, and n represents the number of genes included in the model.
2. Use of the biomarker according to claim 1 in constructing a prognostic risk prediction model for colorectal cancer, characterized in that: In S3, the method for screening prognostic-related genes by performing univariate Cox regression and PH test analysis on the differentially expressed genes in the training set included: deleting the samples with missing overall survival time and survival outcome in the training set prognostic information; in the univariate Cox regression analysis, the normalized gene expression matrix of the differentially expressed genes in the training set, the overall survival time of the samples, and the survival outcome information were used as input files for univariate Cox regression analysis to screen out the genes that met the PH regression hypothesis and had a p-value less than the set value and were significantly related to the sample prognostic time.
Citation Information
Patent Citations
Application of model constructed based on PCD related gene combination in preparation of product for predicting prognosis of colonic adenocarcinoma
CN114540499A