Colorectal cancer prognosis model construction method based on fatty acid metabolism related genes
A fatty acid metabolism-based model for colorectal cancer prognosis identifies key genes to predict patient outcomes and assess immune therapy responsiveness, enhancing treatment strategies and overcoming drug resistance.
Patent Information
- Application Number
- CN202510366627.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-15
AI Technical Summary
In the prior art, the prognostic model of colorectal cancer fails to effectively bind fatty acid metabolism-related genes, resulting in a lack of biomarkers, affecting the effects of personalized treatment and immunotherapy.
A colorectal cancer prognosis model based on fatty acid metabolism-related genes was constructed. By obtaining data from the TCGA database, screening differentially expressed genes, performing single-factor Cox regression analysis, combining multivariate Cox regression analysis, risk scores were calculated, and tumor immunotherapy response and chemotherapy resistance were evaluated.
Effective evaluation of the prognosis of colorectal cancer patients has been achieved, the accuracy of personalized treatment has been improved, the potential of immunotherapy and chemotherapy resistance has been evaluated, and new molecular markers have been provided to improve survival outcomes.
Smart Images

Figure CN120319306A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biomedical technologies, and specifically to a method for constructing a prognostic model for colorectal cancer based on fatty acid metabolism-related genes. Background Art
[0002] Colorectal cancer (CRC) is the third most common cancer globally and the second leading cause of cancer-related deaths. Due to the inconspicuous early symptoms, it is difficult to diagnose, and the 5-year survival rate of metastatic colorectal cancer remains only about 14%. Although the benefits of existing treatment regimens have stabilized, there is still an urgent need to develop new effective treatment strategies to improve survival outcomes. With the development of clinical trials on targeted therapies, specific subgroups of colorectal cancer patients have shown significant efficacy, but the relative lack of biomarkers has delayed progress in this area.
[0003] Abnormal fatty acid metabolism is closely related to the occurrence and development of cancer, that is, the enzymes involved in fatty acid uptake, synthesis, and oxidation are abnormally regulated in various cancers, leading to reprogramming of FA metabolism and promoting the progression of the malignant phenotype of cancer.
[0004] In the prior art, tumor subtypes and prognostic models based on the fatty acid catabolism pathway of colorectal cancer have, to a certain extent, promoted the practical application of the strategy of inhibiting FA synthesis for tumor treatment, explored specific inhibitory targets and mechanisms of action, but have not deeply understood the value of prognostic models in immunotherapy analysis. For this reason, a method for constructing a prognostic model for colorectal cancer based on fatty acid metabolism-related genes is proposed. Summary of the Invention
[0005] The present invention aims to solve at least one of the technical problems existing in the prior art. For this reason, an object of the present invention is to propose a method for constructing a prognostic model for colorectal cancer based on fatty acid metabolism-related genes. A prognostic risk score model is constructed based on fatty acid metabolism-related genes to effectively predict the prognosis of colorectal cancer patients and provide new molecular markers for personalized treatment of colorectal cancer.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A method for constructing a prognostic model for colorectal cancer based on fatty acid metabolism-related genes, comprising the following steps:
[0008] S1. Obtain the transcriptome data of colorectal cancer diagnosis individuals from the TCGA database;
[0009] S2. Perform differential analysis on the expression data of fatty acid metabolism-related genes in colorectal cancer tissues and normal tissues, and screen out differentially expressed genes;
[0010] S3. Perform univariate Cox regression analysis on the differentially expressed genes to screen out the genes related to the prognosis of colorectal cancer;
[0011] S4. Using the screened prognostic genes and combining with the clinical characteristics of the patients, construct a prognostic model for colorectal cancer. Combine with the gene expression levels determined by multivariate Cox regression analysis to calculate the risk score for each patient. The formula is:
[0012] Risk score = ∑(coef * gene expression), where coef is the correlation coefficient of the corresponding gene.
[0013] As a further optimized solution of the present invention, in step S2, use the limma package and sva package to extract the expression levels of DEGs related to fatty acid metabolism, and combine the differential gene expression data with the survival data of the patients.
[0014] As a further optimized solution of the present invention, in step S3, in the TCGA cohort, use the survival and survminer packages to perform univariate cox regression analysis, select the genes related to prognosis from the DEGs related to fatty acid metabolism according to p < 0.05 and generate a forest plot. Use the R software package "maftools" to analyze the mutations and gene associations of tumor samples in the training set.
[0015] Compared with the prior art, the beneficial effects of the present invention are:
[0016] The method for constructing a prognostic model for colorectal cancer based on fatty acid metabolism-related genes proposed by the present invention improves the prior art by introducing the TIDE score to evaluate the tumor immunotherapy reactivity, thereby analyzing the immune escape potential and immunotherapy effect of tumor subtypes. And, through drug sensitivity analysis, the relationship between the fatty acid risk score and chemotherapy resistance is further evaluated, opening up a way for innovative interventions to restore or alleviate the drug resistance mechanisms exhibited by these cells. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a schematic diagram for the exploration of the prognostic characteristics and the construction of the model of the present invention;
[0018] Figure 2 It is a schematic diagram for the verification and performance analysis of the prognostic model of the present invention;
[0019] Figure 3 It is a schematic diagram for the risk score results, immune correlation and immunotherapy analysis of the present invention;
[0020] Figure 4 It is a schematic diagram for the development and verification of the nomogram of the present invention;
[0021] Figure 5Schematic diagram of drug sensitivity analysis of the present invention;
[0022] Figure 6 Schematic diagram of biological function and pathway enrichment analysis of the present invention;
[0023] Figure 7 Schematic diagram of the construction of protein - protein interaction network and identification of core genes of the present invention. Detailed implementation manners
[0024] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0025] Embodiment 1
[0026] The present invention provides a technical solution: a method for constructing a colorectal cancer prognosis model based on fatty acid metabolism - related genes, comprising the following steps:
[0027] S1. Obtain the transcriptome data of colorectal cancer - diagnosed individuals from the TCGA database;
[0028] S2. Perform differential analysis on the fatty acid metabolism - related gene expression data of colorectal cancer tissues and normal tissues, and screen out the differentially expressed genes;
[0029] Specifically, in step S2, the limma package and sva package are used to extract the expression levels of fatty acid metabolism - related degs, and the differential gene expression data is combined with the survival data of patients.
[0030] S3. Perform univariate Cox regression analysis on the differentially expressed genes to screen out the genes related to colorectal cancer prognosis;
[0031] Specifically, in step S3, in the TCGA cohort, the survival and survminer packages are used for univariate cox regression analysis. Genes related to prognosis are selected from the fatty acid metabolism - related degs according to p < 0.05 and a forest plot is generated. The R software package "maftools" is used to analyze the mutations and gene associations of tumor samples in the training set.
[0032] S4. Using the screened prognosis genes, combined with the clinical characteristics of patients, construct a colorectal cancer prognosis model, and calculate the risk score of each patient by combining the gene expression levels determined by multivariate Cox regression analysis. The formula is:
[0033] Risk score = ∑(coef * gene expression), where coef is the correlation coefficient of the corresponding gene.
[0034] Example 2
[0035] 1. Risk score calculation and grouping: Calculate the risk score of each patient using the constructed prognostic model. According to the formula in step S4 of Example 1, all patients are divided into two subgroups according to the median of the risk score: the low-risk group and the high-risk group.
[0036] 2. Principal component analysis: Perform principal component analysis on the low-risk group and the high-risk group using the R packages "limma" and "ggplot2". Observe the distribution and clustering of the two groups of patients on different principal components through the PCA plot. Evaluate the effectiveness of the grouping and judge whether there are significant differences in biological characteristics between the two groups of patients.
[0037] 3. Survival analysis: Perform survival analysis on the low-risk group and the high-risk group using the R packages "survival" and "survminer". Plot the Kaplan-Meier survival curve and compare the differences in overall survival (OS) and progression-free survival (PFS) between the two groups of patients. Use the Log-rank test to evaluate the statistical significance of the differences in survival rates between the two groups of patients.
[0038] 4. Cox regression analysis: Perform univariate and multivariate Cox regression analysis to evaluate the relationship between the risk score and the patient's survival rate, and consider the influence of other clinical characteristics on the survival rate. Evaluate the independence of the prognostic model and judge whether the risk score can be used as a prognostic prediction indicator independent of other clinical characteristics.
[0039] 5. ROC curve analysis: Use the ROC curve to evaluate the prediction accuracy of the risk model at different survival times (such as 1 year, 3 years, 5 years). Calculate the area under the ROC curve (AUC). The closer the AUC is to 1, the higher the prediction accuracy of the model. Compare with other clinical characteristics to evaluate the superiority of the risk model.
[0040] 6. Combined clinical feature ROC analysis: Combine the risk score with other clinical characteristics to construct a combined prediction model. Use the ROC curve to evaluate the prediction accuracy of the combined prediction model. Compare with the models using only the risk score or clinical characteristics to evaluate the superiority of the combined prediction model.
[0041] 7. External dataset validation: Use an external dataset to validate the constructed prognostic model and evaluate the generalization ability of the model.
[0042] Example 3
[0043] Explore the relationship between the constructed colorectal cancer prognostic model and clinical characteristics. The specific steps are as follows:
[0044] 1. Data preparation: Combine the patient's risk score with the clinical feature data.
[0045] 2. Clinical feature analysis: Use the R package "clicor" to evaluate the differences in clinical features between the high-risk group and the low-risk group.
[0046] Clinical features include: age (≤65 years vs >65 years), gender (male vs female), tumor stage (StageⅠ - StageⅣ), T status (T1 - T4), N status (N0 - N2), M status (M0 - M1).
[0047] 3. Visualization: Use the R package "ggpubr" to draw box plots to visualize the differences in clinical features between the high-risk group and the low-risk group. Box plots can intuitively show the median, interquartile range, and outliers of clinical features between the two groups.
[0048] 4. Statistical analysis: Conduct statistical analysis on each clinical feature to evaluate the statistical significance of the differences between the two groups. Based on the results of the statistical analysis, determine which clinical features have a significant correlation with the risk score.
[0049] Example 4
[0050] The specific steps for immunophenotyping analysis are as follows: Use the immuneSubtype R package for immunophenotyping to observe whether there are differences in the patient's risk scores among different immunophenotypes. And visualize the results by drawing box plots using the ggpubr package.
[0051] Example 5
[0052] The specific steps for creating a nomogram using the risk scoring system are as follows:
[0053] Nomogram construction: Use the R package "rms" to draw a nomogram. Incorporate variables such as risk score, gender, age, and stage into the nomogram and set corresponding weights. Evaluate the predictive accuracy of the nomogram through a calibration curve.
[0054] Nomogram application: Through the nomogram, predict the 1-year, 3-year, and 5-year survival rates of patients for individualized prediction.
[0055] Nomogram validation: Use the timeROC package to evaluate the accuracy level of the nomogram in predicting the patient's survival period. Use the survival package for independent prognostic analysis to evaluate whether the nomogram can be used as an independent prognostic factor independent of other clinical traits.
[0056] Example 6
[0057] The specific steps for immune cell infiltration, difference, and immune-related function analysis are as follows:
[0058] Immune cell infiltration assessment: The CIBERSORT bioinformatics algorithm was used to evaluate the infiltration of immune cells in each sample using standardized gene expression profile data.
[0059] Visualization of differential immune cell infiltration levels: The R package "limma" was used to perform differential analysis of immune cell infiltration levels.
[0060] Box plots were drawn using the "reshape2" and "ggpubr" packages to visualize the differences in immune cell infiltration between colon cancer patients in the high- and low-risk groups.
[0061] Immune-related function analysis: The R packages "GSVA" and "GSEABase" were used to quantify the expression of immune functions of each type of immune cell in colorectal cancer tissues of the high- and low-risk groups through single-sample gene set enrichment analysis algorithm.
[0062] Example Seven
[0063] The specific steps to evaluate the differences in immune characteristics and gene mutations between the high- and low-risk groups are as follows:
[0064] 1. Assessment of immune characteristic differences: The GSVA R package was used to perform differential analysis of biological processes between the high- and low-risk groups. The pathway activation levels were shown through a heat map to visualize the differences in immune characteristics between the high- and low-risk groups.
[0065] 2. Correlation analysis between gene mutations and risk scores: The mutRisk R package was used to analyze the correlation between the patient risk scores and the target gene mutations. The results were presented in the form of a box plot to show the distribution of gene mutations between different risk groups.
[0066] Example Eight
[0067] The specific steps to explore the role of the constructed prognostic model in clinical characteristics and immunotherapy prediction are as follows:
[0068] 1. Chemotherapy drug sensitivity analysis: The prophytic R software package was used to calculate the half-maximal inhibitory concentration values of various chemotherapy drugs. Twelve drugs were selected from the Cancer Drug Sensitivity Genomics dataset to compare the IC50 values between the high- and low-risk groups. The differences in drug sensitivity between the two groups were shown through a box plot, and the statistical significance level was set at P < 0.05.
[0069] 2. Tumor immune dysfunction and exclusion score: TIDE scores were performed on the high- and low-risk group samples, and the scoring data were obtained from the TIDE website. The differences in the responses to immunotherapy between the high-risk and low-risk groups were shown through a violin plot.
[0070] Example Nine
[0071] The specific steps for enrichment analysis of differentially expressed genes are as follows:
[0072] 1. Obtaining differential genes: Use the riskDiff method to obtain the differential genes between the high- and low-risk groups.
[0073] 2. GO and KEGG enrichment analysis: Use the R package ClusterProfiler to perform Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes enrichment analysis on immune-related differentially expressed genes. Set P < 0.05 as the standard for significant enrichment.
[0074] 3. Visualization of results: Use barplot and bubble plots to visualize the results of enrichment analysis, showing the enrichment levels of differential genes in GO and KEGG pathways. Display the enrichment of differential genes through bar charts, bubble charts, and circular plots.
[0075] Example Ten
[0076] The specific steps for the analysis of survival-related hub genes based on the prediction model are as follows:
[0077] 1. Construction of protein-protein interaction network: Use the STRING database to construct a protein-protein interaction network and use its visualization tool for preliminary visualization.
[0078] 2. Visualization of network relationships and node attributes: Use the Cytoscape R package to import the protein-protein interaction network into the Cytoscape software for further visualization and analysis of the network relationships and node attributes of genes.
[0079] 3. Identification of core genes: Use the CytoHubba plugin in Cytoscape to identify the core genes in the network.
[0080] Example Eleven
[0081] Explore the roles of core genes in terms of clinical characteristics and immune cell correlations. The specific steps are as follows:
[0082] 1. Clinical correlation analysis of core genes: Use the GeneSur and geneCliCor algorithms to screen out core genes with potential clinical correlations. Perform single-gene survival analysis on the screened genes and divide the samples into high-expression and low-expression groups according to the gene expression levels. Observe the differences in the expression levels of the target genes between different clinical groups and present the analysis results in the form of box plots.
[0083] 2. Core gene immune cell correlation analysis: The geneImmune package was used to study the differential expression of core genes in immune cells. The results of the immune cell correlation analysis of the genes were visualized by box plots to reveal the association between gene expression and immune cell infiltration.
[0084] Example XII
[0085] In step S2 of Example 1, data of 612 colorectal cancer patients were downloaded from the TCGA database as the research population. Based on previous studies, 308 fatty acid metabolism-related genes were selected.
[0086] Using |logFC|>0.585, BH correction, and P<0.01 as the criteria, the intersection of fatty acid-related genes and differentially expressed genes was determined, and 154 fatty acid-related genes with differential expression in tumor tissues were identified, including 65 up-regulated genes and 89 down-regulated genes.
[0087] As shown by A and B in the appendix Figure 1 showed the expression of 154 differentially expressed fatty acid-related genes in tumor tissues and normal tissues, and the volcano plot presented the differential expression levels of these genes in tumor tissues.
[0088] Example XIII
[0089] In step S3 of Example 1, among the survival data of patients from the TCGA database, after removing samples with missing information, 540 TCGA patients were left as the training set (n = 540).
[0090] Survival time, survival status and other data of the patients in the training set were extracted and merged with the expression data of fatty acid-related genes. Through univariate Cox analysis, it was found that 16 genes were differentially associated with overall survival (OS), as shown by C in the appendix Figure 1 shown.
[0091] Subsequently, a waterfall plot was used to show the mutation landscape of 16 key FAM genes, and it was found that missense mutations were the main variant type. Among them, the mutation frequency of MORC2 was the highest, as shown by D in the appendix Figure 1 shown, and G0S2 had no mutations;
[0092] Further analysis of the co-mutation relationship between genes showed that there was a significant co-mutation relationship between ACSL6 and multiple genes (such as ACADL, ACOT11, etc.), indicating their synergistic effect in certain biological processes, as shown by E in the appendix Figure 1 shown.
[0093] Eleven genes finally used to construct the prognostic model were screened by LASSO regression, as shown by the appendix Figure 1As shown by F and G in []. The regression coefficients of 11 genes are as follows:
[0094] Gene Coef
[0095] CD36 0.134940501884887
[0096] ENO3 0.609308898268095
[0097] MORC2 0.172144128161207
[0098] ELOVL3 0.492853637867911
[0099] ACOT11 -0.286999697734189
[0100] ALAD 0.471252518539128
[0101] CIDEA 0.0296481070609902
[0102] ELOVL6 -0.192502555371545
[0103] ACSL6 -0.0218849248895571
[0104] ACADL 0.457613750979409
[0105] CPT2 -0.281558037713414.
[0106] Among them, the genes CD36, ENO3, MORC2, ELOVL3, ALAD, CIDEA, and ACADL are positively correlated with the poor prognosis of CRC, while the genes ACOT11, ELOVL6, ACSL6, and CPT2 are negatively correlated with the poor prognosis of CRC.
[0107] Example 14
[0108] In step S4 of Example 1, the patient score was calculated according to the formula, and the patients were divided into a high-score group and a low-score group according to the median score. Principal component analysis was used for dimensionality reduction, as shown by A and B in the appendix Figure 2 As shown.
[0109] The Kaplan-Meier method was used to calculate the survival probabilities of patients in the high- and low-risk groups in each cohort and compare their OS differences. The analysis results showed that the survival probability of patients in the high-risk group in the training set was significantly lower than that of patients in the low-risk group, and the difference was statistically significant (P<0.01), as shown by the appendix Figure 2as shown in C in
[0110] The progression-free survival of the two groups was compared using the Kaplan-Meier survival curve and the log-rank test. The results showed that there was a significant improvement in PFS in the low-risk group compared with the high-risk group (P<0.01), as shown in Figure 2 D in
[0111] To evaluate the independent prognostic ability of the prognostic model, univariate and multivariate Cox regression analyses were performed on other clinical factors and risk scores. The results showed that in the univariate Cox analysis, the score obtained from the TCGA data cohort was significantly correlated with OS. After adjusting for other confounding factors, the multivariate Cox analysis showed that the risk Score was still an independent predictor of OS, as shown in Figure 2 E and F in , and the higher the risk score, the greater the risk of the patient (HR = 2.989, 95% CI: 2.053-4.353, P<0.001).
[0112] To evaluate the prediction efficiency of the prognostic model for 1-year, 3-year, and 5-year survival rates, the receiver operating characteristic (ROC) curve was plotted to evaluate the prognostic prediction accuracy of the constructed model for patients with different survival times. The results showed that the areas under the curve (AUC) for 1 year, 3 years, and 5 years were 0.663, 0.720, and 0.756, respectively, as shown in Figure 2 G in , indicating that the model had a good prediction effect.
[0113] The AUCs of the risk score and clinical characteristics such as age and gender were calculated respectively. The results showed that compared with the stage clinical characteristics, the risk score had significantly higher prognostic prediction accuracy, as shown in Figure 2 H in .
[0114] Example XV
[0115] In steps S5 and S6 of Example 1, the risk scores of colorectal cancer patients were analyzed for correlation with patient gender, stage, tumor stage, and immune classification. The results showed that there was no significant difference between patients of different genders, as shown in Figure 3 A in , but there were significant differences between cancer patients of different stages, tumor stages, and immune classifications: the higher the stage, the higher the score and the worse the prognosis, as shown in Figure 3 B in . The higher the degree of infiltration of the cancer, the higher the risk score, as shown in Figure 3 C-E in .
[0116] Through the clinical actual prognosis evaluation of colorectal cancer patients, a nomogram was developed based on risk scores and factors such as age, gender, and stage. By scoring each index of the patient and finally summing them up to obtain the total score, the survival probability of the patient at different survival years was obtained. For example, the probability that a patient with a comprehensive score of 255 survived for more than 1 year was 0.817, more than 2 years was 0.606, and more than 3 years was 0.36, as shown by G in Appendix Figure 3 shown.
[0117] To evaluate the predictive accuracy of this nomogram, it was calibrated. The resulting calibration curve showed that the accuracy of this model was good, as shown by H in Appendix Figure 3 shown.
[0118] ROC analysis was performed on multiple clinical characteristics to evaluate the stability of the constructed model. The results showed that the area under the ROC curve of Nomo was the largest, as shown by I in Appendix Figure 3 shown.
[0119] In the independent analysis, the results of the forest plot showed that the Nomogram results were accurate and had a large weight, as shown by J and K in Appendix Figure 3 shown.
[0120] Example XVI
[0121] In step S6 of Example 1, the immunotherapy effects of colorectal cancer patients in different risk groups were explored. The CiberSort algorithm was used to calculate the immune cells and related immune functions of the two groups of patients. The analysis results showed that among the immune cells, the proportions of T cells CD4 memory resting and Dendritic cells resting in the low-risk group were higher than those in the high-risk group, and the difference was statistically significant (P<0.05). This indicates that the relative abundance of cells in the low-risk group was higher than that in the high-risk group, as shown by A in Appendix Figure 4 shown.
[0122] In terms of immune function, the scores of functions such as HLA, type I interferon response, and type II interferon response in the low-risk group were significantly lower than those in the high-risk group (P<0.05), indicating that these immune functions were downregulated in the low-risk group, as shown by B in Appendix Figure 4 shown.
[0123] In non-highly mutated colorectal cancer, the mutation rate of APC > 80%, and the mutation rate of TP53 is about 55%-60%, which are the two most common mutations observed in colorectal cancer. As shown by C in Appendix Figure 4 shown, the risk score differences between the wild-type and mutant of the checkpoint gene were examined. Among them, the score of the wild-type of APC was higher than that of the mutant, but the difference was not statistically significant (P>0.05), as shown by Appendix Figure 4as shown in C of [reference]; while the mutant type of TP53 is significantly higher than the wild type, as shown in Figure 4 D of [reference].
[0124] Subsequently, through immunotherapy analysis, it was found that there was a significant difference in TIDE scores between the high-risk and low-risk groups of patients (P < 0.05). The results showed that the TIDE scores were higher, the potential for immune escape was relatively greater, and the effect of immunotherapy was worse, as shown in Figure 4 E of [reference].
[0125] The relationship between the fatty acid risk score and chemotherapy resistance was further evaluated through drug sensitivity analysis, and the difference in chemotherapy resistance between high FARS and low FARS was calculated using the pRophetic software package, as shown in Figure 5 [reference]. Among them, the IC50 values of axitinib, elesclomol, GNF-2, HG-6-64-1, MG-132, Navitoclax, PF-4708671, PI-103, and WH-4-023 showed stronger chemotherapy effects in the high FARS group, while LY317615 and rTRAI showed stronger chemotherapy effects in the low FARS group.
[0126] Example XVII
[0127] In step S6 of Example 1, the differentially expressed genes (DEGs) between the high-risk and low-risk groups and the expression levels of these genes were shown in region A of [reference]. After that, the differential genes were subjected to Gene Ontology (GO) enrichment analysis, and it was found that the differential genes were mainly enriched in collagen-containing extracellular tissues and participated in heparin binding, glycosaminoglycan binding, and sulfur compound binding, as shown in Figure 6 B and D of [reference]. Figure 6 [reference].
[0128] At the same time, Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analysis was performed on the differential genes, and it was found that the differential genes were mainly enriched in the P13K-Akt signaling pathway, as shown in Figure 6 C and E of [reference].
[0129] PPI network analysis showed that genes such as FN1, TAGLN, ACTG2, MYH11, LMOD1, and CNN1 had high connectivity in the network, as shown in Figure 7 A of [reference]. Performing network analysis, it can be seen that the genes with high connectivity in the network are FN1, ACTGP, and MYH1, among which FN1 has the highest connectivity, as shown in Figure 7 B of [reference]. Using cytoscape to find key genes, FN1 and MYH1 were obtained, as shown in Figure 7 C of [reference].
[0130] Example XVIII
[0131] In step S6 of Example 1, regarding the biological significance of FN1 in colorectal cancer, as shown by D in the appendix Figure 7 , it can be seen that the high expression of the FN1 gene in colorectal cancer is significantly correlated with a shorter survival time.
[0132] The clinicopathological characteristics of patients with high and low FN1 expression were analyzed, including age, gender, clinical stage, T stage, N stage, M stage, etc. As shown by E in the appendix Figure 7 , there was no significant statistical difference in gender between the high and low FN1 expression groups, while there was a certain statistical significance in age (with 65 years old as the boundary), as shown by F in the appendix Figure 7 .
[0133] As the clinical stage increased, the expression level of the FN1 gene also gradually increased, and there was a statistical difference in the FN1 expression level between stage I tumors and other stages, as shown by G - J in the appendix Figure 7 .
[0134] The analysis results showed that based on clinicopathological characteristics, high FN1 expression represented a higher degree of malignancy in colorectal cancer. Among immune cells, the proportions of Plasma cells, T cells CD4 memory resting, and Dendritic cells resting in the low FN1 expression group were higher than those in the high FN1 expression group, and the differences were statistically significant (P < 0.05), indicating that the relative abundances of these cells were higher in the low FN1 expression group compared to the high FN1 expression group. While the relative abundance of Macrophages M0 was higher in the high FN1 expression group, as shown by K in the appendix Figure 7 .
[0135] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with that embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0136] In the drawings of the disclosed embodiments of the present invention, only the structures related to the disclosed embodiments are involved, and other structures can refer to the general design. Without conflict, the same embodiment and different embodiments of the present invention can be combined with each other.
[0137] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for constructing a prognostic model for colorectal cancer based on fatty acid metabolism-related genes, characterized in that Including the following steps: S1. Obtain the transcriptome data of colorectal cancer diagnosed individuals from the TCGA database; S2. Perform differential analysis on the gene expression data related to fatty acid metabolism in colorectal cancer tissues and normal tissues, and screen out the differentially expressed genes; S3. Perform univariate Cox regression analysis on the differentially expressed genes, and screen out the genes related to the prognosis of colorectal cancer; S4. Use the screened prognostic genes, combined with the clinical characteristics of the patients, to construct a colorectal cancer prognosis model, and calculate the risk score of each patient by combining the gene expression levels determined by multivariate Cox regression analysis. The formula is: Risk score = ∑(coef * gene expression), where coef is the correlation coefficient of the corresponding gene.
2. The method for constructing a colorectal cancer prognosis model based on fatty acid metabolism-related genes according to claim 1, wherein: In step S2, the limma package and sva package are used to extract the deg expression levels related to fatty acid metabolism, and the differential gene expression data is merged with the survival data of the patients.
3. The method for constructing a colorectal cancer prognosis model based on fatty acid metabolism-related genes according to claim 1, wherein: In step S3, in the TCGA cohort, the survival and survminer packages are used for univariate cox regression analysis. Genes related to prognosis are selected from the deg related to fatty acid metabolism according to p < 0.05 and a forest plot is generated. The R software package "maftools" is used to analyze the mutations and gene associations of the tumor samples in the training set.
Citation Information
Cited By
Prognosis evaluation method and system for II-III stage colorectal cancer patient
CN120674084A
Machine learning-based bladder cancer subtype classification system and molecular typing method
CN121565270A