Lung adenocarcinoma prognosis risk prediction model and construction method and application thereof

By constructing a prognostic model of genes related to CAF infiltration, using multivariate Cox regression analysis and Nomogram tools, the prognosis problem of lung adenocarcinoma was solved, efficient prognostic evaluation and personalized treatment guidance were achieved, and the survival rate and treatment effect of lung adenocarcinoma patients were improved.

CN120280151APending Publication Date: 2025-07-08SUZHOU UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510425425.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art is difficult to effectively predict the prognosis of lung adenocarcinoma, the lack of specific biomarkers leads to difficulty in early diagnosis, affects treatment choices and reduces survival rates, and the complexity of the tumor microenvironment increases the difficulty of diagnosis and treatment.

Method used

A prognostic model based on genes related to tumor-associated fibroblasts (CAF) infiltration was constructed, and the regression coefficients of genes such as COX6A1, ENOX1, FERMT2, NID1, LOX, SNAI2, GLI2, ZNF154, COX7A1, NXPH3, FRMD4A, SYT11 and ENTPD1 were determined through multivariate Cox regression analysis, and a Nomogram tool was constructed for personalized prediction.

Benefits of technology

It has achieved efficient detection of lung adenocarcinoma prognosis, with a 5-year survival AUC reaching 0.839, providing independent prognostic factors, guiding personalized treatment, and improving the accuracy of prognostic evaluation and the accuracy of treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120280151A_ABST
    Figure CN120280151A_ABST
Patent Text Reader

Abstract

The invention discloses a lung adenocarcinoma prognosis risk prediction model and a construction method and application thereof, and belongs to the technical field of biomedical detection. The lung adenocarcinoma (LUAD) prognosis model is successfully constructed by systematically analyzing the immune infiltration related gene (CAFRG) of tumor-associated fibroblasts (CAF), and the application value of the lung adenocarcinoma (LUAD) prognosis model in prediction of patient prognosis is verified. The CAFRG risk score is negatively related to the immune microenvironment score and is remarkably related to immune checkpoint gene expression, which suggests that a high-risk patient may have the characteristic of immune escape. Furthermore, based on TIDE tool analysis, the patients in the low-risk group have more active T cell immune response. The risk score is also closely related to the sensitivity of antitumor drugs, especially Doramapimod. Therefore, a new theoretical basis is provided for personalized treatment of lung adenocarcinoma, and a new direction is provided for future clinical research.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biomedical detection technologies, and particularly to a prognostic risk prediction model for lung adenocarcinoma, a construction method thereof, and an application thereof. Background Art

[0002] Lung adenocarcinoma (LUAD) is the most common subtype of non-small cell lung cancer (NSCLC), accounting for approximately 40% of all lung cancer cases. Despite significant progress in early detection and treatment strategies, lung cancer, especially the major subtype LUAD, remains the leading cause of cancer-related deaths globally. Different from other non-small cell lung cancer subtypes (such as lung squamous cell carcinoma LUSC) that may present symptoms (such as cough, hemoptysis) earlier, lung adenocarcinoma is usually asymptomatic in the early stage and lacks effective prognostic biomarkers. Many patients are diagnosed at an advanced stage, which limits treatment options and affects survival outcomes. In addition, there are more difficulties in the diagnosis of LUAD. For example, LUAD has higher heterogeneity compared to other subtypes, such as genetic heterogeneity, transcriptional heterogeneity, metabolic heterogeneity, etc. In particular, the degree of differentiation affects the judgment, increasing the difficulty of diagnosis. The metabolic heterogeneity of lung adenocarcinoma also constitutes a diagnostic obstacle - even within the same tumor, different regions may show significant differences in the activities of pathways such as glycolysis and glutamine metabolism, resulting in uneven distribution of SUV values in FDG-PET imaging and being easily misjudged as necrosis or inflammation. In addition, the degree of differentiation is closely related to abnormal epigenetic regulation: DNA methylation profile disorders often occur in poorly differentiated adenocarcinomas, but the clinical detection standardization of such epigenetic markers is still insufficient and it is difficult to be used as a routine diagnostic basis. Therefore, there is an urgent need for new prognostic biomarkers to enhance early detection ability, predict patient outcomes, and guide treatment strategies.

[0003] The tumor microenvironment (TME) plays a key role in the development and metastasis of cancer, and cancer-associated fibroblasts (CAFs) are a central component of this microenvironment. CAFs are derived from normal fibroblasts or other stromal cells, are activated under the stimulation of tumor signals, and play roles in all aspects of cancer biology, including tumor growth, angiogenesis, immune evasion, and resistance to treatment. CAFs secrete various growth factors, cytokines, and extracellular matrix proteins, promoting the migration, invasion, and metastasis of tumor cells.

[0004] The interaction between CAFs and tumor cells is mediated through a complex network of signaling molecules and altered gene expression. Many studies have identified specific CAF-related genes that contribute to tumor invasiveness and metastasis. These genes include those involved in extracellular matrix remodeling, inflammatory responses, and immune cell recruitment, all of which play crucial roles in tumor progression. Recent studies have shown that CAF-related genes can serve as powerful prognostic biomarkers. In the context of LUAD, the expression levels of specific CAF-related genes have been shown to be associated with poor survival rates, making them promising candidates for inclusion in prognostic models. However, there is a lack of comprehensive models integrating CAF-related gene signatures to predict clinical outcomes in LUAD patients. Summary of the Invention

[0005] To address the above problems, the present invention develops a prognostic model based on genes related to the tumor-associated fibroblast (CAF) infiltration score in patients with lung adenocarcinoma (LUAD) and validates its accuracy and robustness. The key objectives include identifying CAF-related genes (CAFRGs), constructing a prognostic model, validating the effectiveness of the model in an independent cohort, evaluating the clinical significance of the risk score, and performing in vitro functional validation of the key genes. These findings aim to improve the precision of prognosis and treatment decision-making for LUAD patients.

[0006] The first objective of the present invention is to provide a prognostic risk prediction model for lung adenocarcinoma, and the prediction model evaluates the predictive performance of the prognosis of lung adenocarcinoma according to a risk score, and the risk score conforms to the following formula:

[0007] Among them, Exp is the expression value of each CAFRG in the model (which can be the normalized expression level value), coef is the regression coefficient of each CAFRG in the model, i represents the index of the CAFRG in the model, n represents the number of CAFRGs included in the model, and the CAFRGs include the COX6A1 gene (Cytochrome c oxidase subunit 6A1), the ENOX1 gene (ecto-NOX disulfide-thiol exchanger 1), the FERMT2 gene (FERM domain containing kindlin 2), the NID1 gene (Nidogen 1), the LOX gene (Lysyl Oxidase), the SNAI2 gene (snail family transcriptional repressor 2), the GLI2 gene (GLI family zinc finger 2), the ZNF154 gene (zinc finger protein 154), the COX7A1 gene (cytochrome c oxidase subunit 7A1), the NXPH3 gene (Neurexophilin 3), the FRMD4A gene (FERM domain containing 4A), the SYT11 gene (Synaptotagmin 11), and the ENTPD1 gene (Ectonucleoside triphosphate diphosphohydrolase 1).

[0008] Furthermore, the regression coefficients are obtained through multivariate Cox regression analysis. Specifically, all CAFRGs are incorporated into the multivariate Cox regression to evaluate the relationship between each CAFRG and the survival risk of lung adenocarcinoma patients, and the regression coefficients of each CAFRG are obtained.

[0009] Further, the regression coefficients corresponding to the genes COX6A1, ENOX1, FERMT2, NID1, LOX, SNAI2, GLI2, ZNF154, COX7A1, NXPH3, FRMD4A, SYT11, and ENTPD1 are 0.491, 0.409, 0.319, 0.257, 0.223, 0.148, 0.137, -0.136, -0.179, -0.182, -0.258, -0.391, and -0.402, respectively.

[0010] The second object of the present invention is to provide a method for constructing the prediction model, including the following steps:

[0011] S1. Obtain tissue samples of lung adenocarcinoma patients and control tissue samples, analyze the correlation between the gene expression levels and the infiltration degree of immune cells including CAFs therein, and select those with a high correlation with CAF infiltration as candidate CAFRGs;

[0012] S2. Use univariate Cox regression analysis to evaluate the relationship between the expression level of each CAFRG in the candidate CAFRGs and the survival period, and perform a secondary screening of the candidate CAFRGs according to the risk values calculated for each CAFRG;

[0013] S3. Use LASSO regression analysis to further screen the results of the secondary screening in S2 to obtain CAFRGs with a higher correlation with the survival period, which are the prognosis-related genes (i.e., the genes COX6A1, ENOX1, FERMT2, NID1, LOX, SNAI2, GLI2, ZNF154, COX7A1, NXPH3, FRMD4A, SYT11, and ENTPD1);

[0014] S4. Perform multivariate Cox regression analysis on the prognosis-related genes, and calculate the risk scores of each patient according to the regression coefficients and expression levels of each CAFRG;

[0015] S5. Evaluate the prediction performance of the prognosis risk prediction model according to the risk scores.

[0016] Further, in step S1, analyzing the correlation between the gene expression levels and the infiltration degree of cellular immunity including tumor-associated fibroblasts therein includes the following steps: obtaining the infiltration score of tumor-associated fibroblasts in tumor samples using the xCell method, and calculating the correlation coefficient between the gene expression levels and the infiltration score using the Spearman method.

[0017] Further, in step S5, a multivariate Cox regression analysis is performed with the risk score and clinical indicators as variables. If the risk score and clinical indicators based on the prognostic risk prediction model can be used as independent prognostic factors for lung adenocarcinoma respectively, a Nomogram is constructed by combining the risk score and clinical indicators to evaluate the predictive performance of the prognostic risk prediction model for the overall survival time of lung adenocarcinoma patients.

[0018] Further, in step S5, the correlation between the risk score and the anti-tumor drug sensitivity calculated based on the OncoPredict algorithm is used to evaluate the predictive performance of the prognostic risk prediction model for anti-tumor drug treatment of lung adenocarcinoma patients.

[0019] Further, according to the expression levels of prognosis-related genes in lung adenocarcinoma patient samples, the OncoPredict algorithm is used to calculate the anti-tumor drug sensitivity of each lung adenocarcinoma patient.

[0020] The third object of the present invention is to provide a prognostic risk prediction system for lung adenocarcinoma, which contains the prognostic risk prediction model for lung adenocarcinoma and a reagent for detecting the expression level of CAFRG.

[0021] The CAFRG includes COX6A1 gene, ENOX1 gene, FERMT2 gene, NID1 gene, LOX gene, SNAI2 gene, GLI2 gene, ZNF154 gene, COX7A1 gene, NXPH3 gene, FRMD4A gene, SYT11 gene and ENTPD1 gene.

[0022] The fourth object of the present invention is to provide the application of the prognostic risk prediction model for lung adenocarcinoma or the prognostic risk prediction system for lung adenocarcinoma in the preparation of a prognostic detection product for lung adenocarcinoma.

[0023] The beneficial effects of the present invention:

[0024] The present invention constructs a prognostic model based on CAF immune infiltration-related genes. Only 13 CAF-related genes specific to the prognostic risk of a specific lung cancer subtype - lung adenocarcinoma are involved in this prognostic model, enabling efficient detection of its prognosis, and the AUC for 5-year survival is as high as 0.839. This model provides an effective tool for the prognostic evaluation of LUAD patients and opens up a new direction for the personalized treatment of lung adenocarcinoma. Description of the Drawings

[0025] Figure 1Identification of CAF infiltration-related genes. The scatter plot shows the correlation between CAF infiltration scores and the expression of all genes in the TCGA LUAD dataset, calculated by the MCPCOUNTER (A) and XCELL (B) algorithms respectively. The Venn diagram shows the overlap of genes positively (C) and negatively (D) correlated with CAF infiltration scores based on the two algorithms. TCGA: The Cancer Genome Atlas; LUAD: Lung adenocarcinoma; CAF: Cancer-associated fibroblasts.

[0026] Figure 2 Construction and performance analysis of the CAFRGs prognostic model. (A) Univariate COX analysis identified CAFRGs related to prognosis in the training set; (B) Multivariate COX analysis constructed the prognostic model; (C) The TimeROC curve shows the ROC curves and AUC values of patients in the training set at 1-5 years; (D) Survival curves of patients in the high- and low-risk groups in the training set; Distribution of risk scores (E) and survival outcomes and time (F) of samples in the high- and low-risk groups of the training set; (G) Heatmap showing the expression differences of CAFRGs in the model in samples of the high- and low-risk groups of the training set.

[0027] Figure 3 To verify the robustness of the model in the TCGA LUAD and GSE31210 datasets. (A) The TimeROC curve shows the ROC curves and their AUC values at 1 to 5 years for patients in the TCGA_LUAD dataset; (B) Comparison of survival curves of patients in the high-risk and low-risk groups in the TCGA_LUAD dataset; (C) Distribution of risk scores and survival outcomes of samples in different risk groups in the TCGA_LUAD dataset; (D) Heatmap showing the expression differences of genes related to fibroblasts (CAFRGs) in the model in samples of the high- and low-risk groups of the TCGA_LUAD dataset; (E) The TimeROC curve shows the ROC curves and their AUC values at 1 to 5 years for patients in the GSE31210 dataset; (F) Comparison of survival curves of patients in different risk groups in the GSE31210 dataset; (G) Distribution of risk scores and survival outcomes of patients in the high- and low-risk groups in the GSE31210 dataset; (H) Heatmap presenting the expression patterns of CAFRGs in the model in samples of the high- and low-risk groups of the GSE31210 dataset. TCGA, LUAD, ROC, CAFRG, AUC.

[0028] Figure 4For the validation of the model robustness in the GSE13213 dataset. (A) The ROC curve shows the ROC and AUC values of the 1-year survival in the GSE13213 dataset; (B) Comparison of the survival curves of high-risk and low-risk patients in the GSE13213 dataset; (C) Risk score distribution and survival outcomes of high-risk and low-risk samples in the GSE13213 dataset. ROC: Receiver Operating Characteristic curve; AUC: Area Under the Curve.

[0029] Figure 5 The CAFRG risk score is an independent prognostic factor. (A) The forest plot shows the results of the univariate COX analysis of the CAFRG risk score and major clinicopathological features for patient prognosis in TCGA LUAD; (B) The forest plot shows the results of the multivariate COX analysis of the significant factors in the univariate COX analysis results in TCGA LUAD for patient prognosis; (C) The forest plot shows the results of the univariate COX analysis of the CAFRG risk score and major clinicopathological features for patient prognosis in GSE31210; (D) The forest plot shows the results of the multivariate COX analysis of the significant factors in the univariate COX analysis results in GSE31210 for patient prognosis. CAFRG, COX, TCGA, LUAD.

[0030] Figure 6 To construct and evaluate a clinical prediction nomogram. (A) A clinical prediction nomogram was constructed based on the independent prognostic factors obtained from the multivariate COX analysis in the TCGA LUAD dataset; the calibration plot (B), ROC curve (C), and DCA curve (D) were used to verify the accuracy of the nomogram constructed based on the TCGA LUAD dataset. (E) A clinical prediction nomogram was constructed based on the independent prognostic factors obtained from the multivariate COX analysis in the TCGA LUAD dataset; the calibration plot (F), Kaplan-Meier survival curve (G), and DCA curve (H) were used to verify the accuracy of the nomogram constructed based on the GSE31210 dataset. TCGA, LUAD, COX, DCA, Decision Curve Analysis.

[0031] Figure 7 To construct a Nomogram-based clinical prediction tool. (A) A visual prediction online tool was designed for the clinical prediction nomogram constructed based on the TCGA LUAD and GSE31210 datasets; the predicted survival curve (B) and the survival probabilities at 1, 2, 3, and 5 years (C) were set with parameters of T4, N0, and a risk score of 5; the predicted survival curve (D) and the survival probabilities at 1, 2, 3, and 5 years (E) were set with parameters of T1, N0, and a risk score of 7. TCGA, LUAD.

[0032] Figure 8The risk score is related to immune cell infiltration and immunotherapy in the TCGA LUAD dataset.

[0033] (A) The lollipop plot shows the results of the correlation analysis between the risk score and the immune cell infiltration score obtained based on the xCell algorithm; the scatter plot shows the correlation between the risk score and the immune microenvironment score (B) and the immune score (C); (D) The lollipop plot shows the results of the correlation analysis between the risk score and the expression of immune checkpoint genes; the scatter plot shows the correlation between the risk score and the expression of BTLA (E) and VSIR (F); the scatter plot and box plot show the correlation between the risk score and the T cell Dysfunction (G), Exclusion (H), MDSC infiltration score (I), and CAF infiltration score (J) calculated based on the TIDE algorithm. TCGA, LUAD, TIDE, CAF, MDSC.

[0034] Figure 9 The CAFRG risk score is related to the progression of lung adenocarcinoma. The heatmap shows the correlation between the risk score and oncogenes (A) and the anti-tumor drug sensitivity (B) calculated based on OncoPredict analyzed from the TCGA LUAD and GSE31210 datasets; the Lollipop plots show the GSEA enrichment analysis results of the CAFRG risk score based on the TCGA LUAD dataset, biological functions (C), and signaling pathways (E); the GSEA plots show the enrichment of the CAFRG risk score in several important biological functions (D) and signaling pathways (F). TCGA, LUAD, CAFRG, GSEA. Detailed implementation manners

[0035] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the embodiments cited do not limit the present invention.

[0036] The solutions involved in the present invention are as follows:

[0037] The present invention successfully constructs and validates a prognostic model for lung adenocarcinoma (LUAD) by systematically analyzing tumor-associated fibroblast (CAF) immune infiltration-related genes, reveals the important role of CAF in the tumor immune microenvironment, and provides new insights and theoretical basis for precision medicine. We found that this model can not only effectively predict the prognosis of LUAD patients, but also be closely related to the tumor immune microenvironment, drug sensitivity, and tumor biological characteristics (such as proliferation, migration, and senescence). These findings provide important clues for further understanding the mechanism of action of CAF in lung adenocarcinoma and developing new treatment strategies.

[0038] One of the core findings of this invention is that the CAFRG risk score is significantly correlated with the immune microenvironment score and immune cell infiltration in LUAD, especially closely related to the expression of immune checkpoint genes in LUAD. Changes in the immune microenvironment play a crucial role in the process of tumor immune escape. As an important part of the tumor microenvironment, CAFs may affect the tumor immune escape mechanism by regulating the infiltration of immune cells and the activation of immune checkpoints. We observed that in LUAD, the CAFRG risk score is correlated with the infiltration of multiple immune cells and negatively correlated with the immune microenvironment score. These results indicate that the immune microenvironment and immune activity of high-risk patients are lower, which may predict the characteristics of immune escape. Based on the TIDE tool, the risk score was found to be negatively correlated with T cell dysfunction and positively correlated with exclusion, indicating that patients in the low-risk group may have a more active T cell immune response, and the immune system can effectively fight against tumors, while the high-risk group may face the challenge of tumor immune escape, manifested as impaired T cell function and T cell exclusion, resulting in weakened immune surveillance. Immune checkpoint genes (such as PD-1, CTLA-4, etc.) are usually closely related to the immune escape mechanism. Tumor cells evade host immune surveillance by expressing immune checkpoint molecules to inhibit the attack of the immune system. We found that the risk score is negatively correlated with the expression of these immune checkpoint genes, and these results further imply the prediction of the CAFRG risk score for lung adenocarcinoma immune escape and immunotherapy response. In addition, we also explored the relationship between the risk score and tumor biological functions and found that it is closely related to the related signaling pathways and biological functions such as proliferation, migration, and immune system activation in lung adenocarcinoma. Tumor drug resistance has become a major challenge in tumor clinical treatment. Drug resistance not only significantly reduces the effectiveness of treatment, but also leads to a worse prognosis for patients, increasing the complexity and cost of treatment. The risk score was also found to be closely related to the sensitivity of anti-tumor drugs, especially significantly negatively correlated with the sensitivity to Doramapimod (a MAPK / ERK inhibitor).

[0039] Example 1

[0040] The establishment of the LUAD prognosis model includes the following steps:

[0041] 1. Dataset

[0042] This study used data from multiple public databases on lung adenocarcinoma (LUAD) for analysis. First, gene expression data and clinical information of LUAD patients were obtained from the Cancer Genome Atlas (TCGA) database. The TCGA dataset contains gene expression data of 572 patients, including 513 tumor tissues and 59 normal adjacent tissue samples, as well as corresponding clinical characteristics (such as age, gender, tumor stage, etc.). To verify the effectiveness of the model, the GSE31210 dataset containing gene expression and prognosis information of 226 primary LUAD samples was additionally extracted from the Gene Expression Omnibus (GEO) database.

[0043] 2. Immune infiltration analysis

[0044] To evaluate the immune microenvironment in LUAD samples, especially the infiltration level of cancer-associated fibroblasts (CAFs), the present invention used the xCell and TIMER algorithms for immune infiltration analysis. We obtained the infiltration scores of CAFs and other immune cells and stromal cells in tumor samples through the xCell algorithm. The infiltration degrees of multiple immune cells (including T cells, B cells, macrophages, etc.) were quantified by the TIMER algorithm, and the T cell dysfunction and exclusion scores of each sample were quantified.

[0045] 3. Prognostic model construction and validation

[0046] All genes in the TCGA LUAD dataset that were related to the CAF score (|correlation coefficient| > 0.25, P < 0.05) were analyzed using the Spearman method and defined as CAF infiltration-related genes (CAFRGs). The CAFRG prognostic model was constructed and validated according to the following steps:

[0047] (a) Dataset splitting: The TCGA LUAD dataset was split into a training set and an internal test set at a ratio of 7:3.

[0048] (b) Clinical feature analysis: The chi-square test was used to analyze the distribution differences of clinical characteristics (such as gender, age, tumor stage, lymph node metastasis, etc.) in the training set and the validation set.

[0049] (c) Univariate Cox regression analysis: The relationship between the expression of each CAFRG and the overall survival (OS) was evaluated, and the hazard ratio (HR) and corresponding p-value of each CAFRG were calculated. CAFRGs with p-values less than 0.05 were selected as candidate genes.

[0050] (d) LASSO regression analysis: To reduce model complexity and select the best prognostic factors, LASSO (Least Absolute Shrinkage and Selection Operator) regression analysis was used to screen out the CAFRGs most relevant to OS from univariate Cox regression.

[0051] (e) Multivariate Cox regression analysis: Based on the LASSO regression analysis, multivariate Cox step - wise regression analysis was further used to establish a prognostic model of CAF - related genes. According to the Cox regression results, the risk score of each patient was calculated, and the risk score formula was: where, Exp i is the expression value of the i - th CAFRG, and coef i is the Cox regression coefficient of this gene.

[0052] (f) Model validation: The predictive performance of the prognostic model in the training set, internal test set, and validation set was evaluated by Kaplan - Meier survival analysis and time - dependent ROC curve (using the "TimeROC" package).

[0053] The results of this example are as follows:

[0054] (1) Construction and validation of a prognostic model of CAF - related genes

[0055] Two LUAD cohorts from the TCGA and GEO databases, along with corresponding clinical data, were used in this invention. Table 1 summarizes the demographic and clinical characteristics of the training set, internal test, and independent validation set. After excluding samples with missing clinical information in the TCGA - LUAD dataset, a total of 504 LUAD patients were included, among which 183 were alive and 321 were dead (median follow - up time: 1.789 years). This dataset was randomly divided into a training set (n = 353) and an internal test set (n = 151) at a ratio of 7:3. As expected, there were no significant differences in the main clinicopathological characteristics among the training set, test, and the entire TCGA - LUAD cohort (Table 1). In addition, we included the GEO dataset GSE31210, which included 226 LUAD patients, and the mortality rate at the end of follow - up was 37.81% (median follow - up time: 4.720 years).

[0056] Table 1

[0057]

[0058] (2) Construction and validation of a prognostic model of CAF - related genes

[0059] Based on the TCGA LUAD dataset, the CAF immune infiltration scores of each sample were obtained using the MCPCOUNTER and XCELL algorithms respectively. Correlation analysis was used to calculate the genes whose expression levels were correlated with the CAF infiltration scores (CAFRGs, Figure 1 A - B). Further, the intersection genes of the analysis results of the two algorithms were obtained by taking the intersection. A total of 1154 genes positively correlated with CAF infiltration and 17 negatively correlated genes were identified ( Figure 1 C - D). Based on the training set data, univariate COX analysis was used, and a total of 174 prognosis - related CAFRGs were identified ( Figure 2 A). Gene screening was performed based on LASSO regression to obtain 30 characteristic genes. Further, the characteristic genes were incorporated into stepwise multivariate COX regression. As shown in Figure 2 B for the multivariate COX results. Finally, a risk model for CAFRGs was obtained: Risk score = COX6A1 Exp×(0.491)+ENOX1 Exp×(0.409)+FERMT2 Exp×(0.319)+NID1 Exp×(0.257)+LOX Exp×(0.223)+SNAI2 Exp×(0.148)+GLI2 Exp×(0.137)+ZNF154 Exp×(-0.136)+COX7A1 Exp×(-0.179)+NXPH3 Exp×(-0.182)+FRMD4A Exp×(-0.258)+SYT11Exp×(-0.391)+ENTPD1 Exp×(-0.402). The ROC curve showed that the risk score had good predictive performance for patient prognosis, and the AUC values at 1, 4, and 5 years were 0.790, 0.819, and 0.839 respectively ( Figure 2 C), and the prognosis of high - risk patients was significantly worse than that of low - risk patients ( Figure 2 D). In the training set, the overall survival time of high - risk patients was shorter and the number of deaths was larger compared to low - risk patients ( Figure 2 E - F). The heatmap showed the expression distribution of the genes in the model in high - and low - risk samples ( Figure 2 G).

[0060] (3) Validation of the CAF risk model

[0061] We validated the robustness of the model in the entire TCGA LUAD dataset and another independent dataset GSE31210. Based on the TCGA LUAD dataset, the ROC curve of the risk score for patient prognosis was plotted, and the AUC value at 1 year was 0.786 ( Figure 3 A). After dividing the risk score into high - and low - risk groups according to the median, survival curve analysis showed that the prognosis of low - risk group patients was significantly better than that of high - risk groupFigure 3 B). The distribution of risk scores and survival outcomes showed that patients in the low-risk group had a lower mortality rate, while the mortality rate in the high-risk group was significantly higher ( Figure 3 C). The heatmap demonstrated the expression patterns of CAFRGs in high- and low-risk group samples of the TCGA LUAD dataset ( Figure 3 D). Similarly, in the GSE31210 dataset, the ROC curves showed the ROC curves and AUC values at 1- to 5-year time points, further validating the robustness of the model in an independent dataset ( Figure 3 E). The survival curves indicated that patients in the high-risk group had a poor prognosis ( Figure 3 F). Additionally, the overall survival time of high-risk patients was shorter compared to that of low-risk patients, and the number of deaths was higher ( Figure 3 G). The heatmap revealed the expression differences of CAFRGs in high- and low-risk group samples of the GSE31210 dataset ( Figure 3 H). In another independent cohort GSE13213, the AUC of the 1-year ROC curve was 0.882 ( Figure 4 A), and the survival curves also showed a poor prognosis in the high-risk group ( Figure 4 B), with the overall survival time in the high-risk group being lower than that in the low-risk group and more death cases ( Figure 4 C). These results confirmed the reliability and robustness of the CAFRG prognostic model.

[0062] (4) The CAFRG risk score is an independent prognostic factor

[0063] In the TCGA LUAD dataset, univariate COX regression analysis showed that the CAFRG risk score was significantly correlated with major clinicopathological features (such as distant metastasis, lymph node metastasis, invasion depth, and stage) and the risk score, and all were risk factors (HR > 0, P < 0.05, Figure 5 A). Further multivariate COX regression analysis indicated that the CAFRG risk score was still an independent prognostic factor (HR = 2.37, 95% CI: 2.16–4.04, P < 0.001, Figure 5 B). In the GSE31210 dataset, univariate COX analysis identified that the CAFRG risk score and stage had a significant impact on patient prognosis (HR > 0, P < 0.05, Figure 5C). Further multivariate COX regression analysis verified the status of the CAFRG risk score as an independent prognostic factor (HR = 2.82, 95% CI: 1.23–6.046, P = 0.014). These results consistently indicate that the CAFRG risk score serves as an independent and reliable prognostic marker in different datasets and has important clinical application value.

[0064] Example 2 Construction of a clinical prediction tool

[0065] 1. Construction and evaluation of the Nomogram

[0066] According to the results of multivariate Cox regression analysis, a Nomogram was constructed to provide an individualized survival prediction tool for patients. The Nomogram combines the CAFRG risk score with the independent prognostic factors obtained from multivariate COX analysis to generate a prediction of the patient's survival probability.

[0067] Construction of the Nomogram: The rms package was used to combine the coefficients of multivariate Cox regression to construct a comprehensive model, visually demonstrating the impact of risk scores and clinical characteristics on patient survival.

[0068] Model evaluation: The predictive performance of the Nomogram was evaluated by calculating the C-index (concordance index). A calibration curve was used to evaluate the applicability of the model in actual clinical practice. Decision curve analysis (DCA) was used to evaluate the clinical decision-making value of the Nomogram.

[0069] 2. Construction of a clinical prediction tool

[0070] Based on the prognostic model and the Nomogram, a clinical prediction tool was further developed to assist clinicians in predicting the survival risk of patients based on their clinical characteristics and CAFRG risk score. This tool can draw a survival curve based on the patient's clinical information and risk score and predict the survival probability at 1 year, 3 years, and 5 years. This tool is implemented through the R Shiny package, enabling doctors to conveniently use this tool to help formulate more personalized treatment plans. Through this tool, the accuracy of prognostic assessment of cancer patients can be improved, and clinical treatment decisions can be guided.

[0071] The results of this example are as follows:

[0072] (1) Construction and evaluation of a clinical prediction nomogram

[0073] In the TCGA LUAD dataset, a clinical prediction nomogram was constructed using the independent prognostic factors screened by multivariate COX regression analysis ( Figure 6A). The accuracy of the nomogram was verified as follows: (1) The calibration plot showed that the predicted results of 1-year, 3-year, and 5-year survival rates were highly consistent with the actual survival rates, demonstrating good predictive performance. Figure 6 B); (2) The ROC curve showed that the AUCs of the nomogram model for predicting 1-, 3-, and 5-year survival probabilities were 0.811, 0.754, and 0.787, respectively. Figure 6 C); (3) The results of decision curve analysis (DCA) indicated that the nomogram provided a relatively high net benefit at different risk thresholds, supporting its effectiveness in clinical decision-making. Figure 6 D). For the GSE31210 dataset, a corresponding clinical prediction nomogram was constructed based on independent prognostic factors. Figure 6 E). The calibration plot Figure 6 F), the ROC curve Figure 6 G), and the DCA curve Figure 6 H) demonstrated the potential application value of the nomogram in clinical practice. In summary, the nomograms constructed based on the TCGA LUAD and GSE31210 datasets both showed excellent predictive ability and clinical practicability, and could provide effective support for the individualized treatment of LUAD patients.

[0074] (2) Construction of a Nomogram-based clinical prediction tool

[0075] We used R language to construct a visualization online tool for the clinical prediction nomogram based on the TCGA LUAD and GSE31210 datasets. This tool allows users to perform individualized survival prediction according to different clinical characteristics and risk scores. By setting different clinical parameters, users can intuitively obtain the predicted survival curve and survival probability of individual patients. Under the current parameters, that is, T4, N0, and a risk score of 5, the survival probabilities of patients at 1 year, 2 years, 3 years, and 5 years are 84%, 66%, 49%, and 18%, respectively. Figure 7 A). Figure 7 B - C show the predicted survival curves when the parameters are set to T4, N0, and a risk score of 5, presenting lower survival probabilities. Further, when the parameters are set to T1, N0, and a risk score of 7, the corresponding survival probabilities at 1 year, 2 years, 3 years, and 5 years are 47%, 17%, 5%, and 1%, respectively, indicating a poor long-term survival prognosis. Figure 7 D - E).

[0076] Example 3 Risk score is related to immune cell infiltration and immunotherapy

[0077] Based on the TCGA LUAD dataset, we analyzed the correlation between the risk score and the immune cell infiltration score calculated by the xCell algorithm. The results showed that the risk score was significantly correlated with the infiltration levels of multiple immune cell subsets, especially classical dendritic cells (cDC), M2 macrophages, HSCs (hematopoietic stem cells), and mast cells, etc. Figure 8 A). In addition, the risk score showed a strong correlation with the immune microenvironment score and the immune score (correlation coefficients were -0.445 and -0.435, respectively, Figure 8 B-C). As shown in Figure 8 D, further correlation analysis showed that the risk score was significantly correlated with the expression levels of multiple immune checkpoint genes, especially BTLA (r = -0.330, Figure 8 E) and VSIR (r = -0.311, Figure 8 F). Further, based on the TIDE algorithm, we further obtained the CAF and MDSC immune cell infiltration scores and the T cell dysfunction and exclusion scores. Further correlation analysis showed that the risk score was correlated with T cell dysfunction ( Figure 8 G) and exclusion ( Figure 8 H), as well as the MDSC ( Figure 8 I) and CAF ( Figure 8 J) immune infiltration scores. Similarly, based on the GSE31210 dataset, the results of the correlation analysis were also related to multiple immune cell infiltration scores and the T cell dysfunction and exclusion calculated by the TIDE algorithm, etc. These results suggest that the risk score can indicate the response of LUAD patients to immunotherapy.

[0078] Example 4 The risk score is related to the progression of lung adenocarcinoma

[0079] 1. Drug sensitivity analysis

[0080] Drug sensitivity analysis aims to evaluate the sensitivity of different tumor samples to various anti-tumor drugs, in order to provide a basis for clinical treatment. The drug sensitivity data were obtained from the GDSC database (https: / / www.cancerrxgene.org / ), which provides cell line sensitivity data for various drugs including chemotherapy drugs and targeted drugs. First, according to the expression of CAFRG in patient samples, the OncoPredict package was used to predict the drug sensitivity of each sample.

[0081] 2. GSEA enrichment analysis

[0082] Gene Set Enrichment Analysis (GSEA) was used to analyze the signaling pathways and biological functions related to the risk score. First, all genes related to the risk score were obtained based on correlation analysis and sorted according to the correlation coefficient. GSEA was conducted using the "ClusterProfiler" R package based on pre-defined C2 (curated gene sets) and C5 (GO gene sets). Pathways significantly related to the risk score were screened according to the normalized enrichment score (NES) and FDR (false discovery rate) values.

[0083] 3. Cell culture

[0084] The cell lines used in this experiment included human lung adenocarcinoma (LUAD) cell lines A549 and H1299. The cells were all purchased from ATCC (American Type Culture Collection) and cultured according to standard cell culture conditions. A549 cells and H1299 cells were cultured in RPMI-1640 medium containing 10% fetal bovine serum (FBS) respectively. The cells were cultured in a constant temperature incubator at 37 °C and 5% CO2, and passaged after reaching 70-80% confluence. All cells were examined by routine morphological inspection and used after confirming no cell contamination (such as mycoplasma contamination).

[0085] The results of this example are as follows:

[0086] As Figure 9 A, The heatmap shows the correlation between the CAFRG risk score and oncogenes based on the analysis of TCGA LUAD and GSE31210 datasets. The results show that the CAFRG risk score is significantly positively correlated with multiple known oncogenes (such as PLK1, CDK1, FOXM1, etc.). Further correlation analysis results show that there is a significant correlation between the CAFRG risk score and the anti-tumor drug sensitivity calculated based on the OncoPredict algorithm, such as Doramapimod, Axitinib, Uprosertib, and Niraparib, etc. Figure 9 B). In addition, based on the GSEA algorithm, we analyzed the biological functions and signaling pathways related to the CAFRG risk score using the TCGA LUAD dataset. In the biological function analysis, the risk score was related to functions such as DNA replication, DNA double-strand break, cell adhesion regulation, and immune response activation. Figure 9 C-D). In the signaling pathway analysis, the risk score was related to multiple cancer-related signaling pathways, such as the cell cycle, mismatch repair, DNA replication, JAK STAT signaling pathway, cell adhesion molecules, etc. Figure 9E-F).

[0087] Obviously, the above embodiments are merely examples given for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or alterations can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. And the obvious changes or alterations derived therefrom still fall within the protection scope of this invention.

Claims

1. A prognostic risk prediction model for lung adenocarcinoma, characterized in that, The prediction model evaluates the prediction performance of the prognosis of lung adenocarcinoma based on the risk score, and the risk score conforms to the following formula: where Exp is the expression value of each CAFRG in the model, coef is the regression coefficient of each CAFRG in the model, i represents the index of CAFRG in the model, and n represents the number of CAFRGs included in the model. The CAFRGs include the COX6A1 gene, ENOX1 gene, FERMT2 gene, NID1 gene, LOX gene, SNAI2 gene, GLI2 gene, ZNF154 gene, COX7A1 gene, NXPH3 gene, FRMD4A gene, SYT11 gene, and ENTPD1 gene.

2. The prognostic risk prediction model for lung adenocarcinoma according to claim 1, wherein The regression coefficient is obtained through multivariate Cox regression analysis, specifically including: incorporating all CAFRGs into multivariate Cox regression, evaluating the relationship between each CAFRG and the survival risk of lung adenocarcinoma patients, and obtaining the regression coefficient of each CAFRG.

3. The lung adenocarcinoma prognosis risk prediction model according to claim 2, wherein The regression coefficients corresponding to the genes COX6A1, ENOX1, FERMT2, NID1, LOX, SNAI2, GLI2, ZNF154, COX7A1, NXPH3, FRMD4A, SYT11, and ENTPD1 are 0.491, 0.409, 0.319, 0.257, 0.223, 0.148, 0.137, -0.136, -0.179, -0.182, -0.258, -0.391, and -0.402, respectively.

4. The method for constructing the prognostic risk prediction model for lung adenocarcinoma according to any one of claims 1 to 3, characterized in that, It includes the following steps: S1. Obtain tissue samples of lung adenocarcinoma patients and control tissue samples, analyze the correlation between the gene expression level and the infiltration degree of immune cells including cancer-associated fibroblasts (CAFs) therein, and screen out candidate CAFRGs; S2. Use univariate Cox regression analysis to evaluate the relationship between the expression level of each CAFRG in the candidate CAFRGs and the survival period, and perform secondary screening on the candidate CAFRGs according to the risk values calculated for each CAFRG; S3. Use LASSO regression analysis to perform re-screening on the results of the secondary screening in S2 to obtain CAFRGs with a higher correlation with the survival period, and obtain prognosis-related genes; S4. Perform multivariate Cox regression analysis on the prognosis-related genes, and calculate the risk scores of each patient according to the regression coefficient and expression level of each CAFRG; S5. Evaluate the prediction performance of the prognosis risk prediction model according to the risk score.

5. The construction method according to claim 4, wherein, In step S1, analyzing the correlation between the gene expression level and the infiltration degree of cellular immunity including cancer-associated fibroblasts therein includes the following steps: obtaining the infiltration score of cancer-associated fibroblasts in tumor samples using the xCell method, and calculating the correlation coefficient between the gene expression level and the infiltration score using the Spearman method.

6. The construction method according to claim 4, characterized in that In step S5, a multivariate Cox regression analysis was performed with the risk score and clinical indicators as variables. If the risk score and clinical indicators based on the prognostic risk prediction model can be used as independent prognostic factors for lung adenocarcinoma respectively, a Nomogram was constructed by combining the risk score and clinical indicators to evaluate the predictive performance of the prognostic risk prediction model for the overall survival time of lung adenocarcinoma patients.

7. The construction method according to claim 4, wherein In step S5, the correlation between the risk score and the anti-tumor drug sensitivity calculated based on the OncoPredict algorithm was used to evaluate the predictive performance of the prognostic risk prediction model for anti-tumor drug treatment of lung adenocarcinoma patients.

8. The construction method according to claim 7, characterized in that According to the expression levels of prognosis-related genes in lung adenocarcinoma patient samples, the OncoPredict algorithm was used to calculate the anti-tumor drug sensitivity of each lung adenocarcinoma patient.

9. A prognostic risk prediction system for lung adenocarcinoma, characterized in that, The prediction system contains the lung adenocarcinoma prognostic risk prediction model according to any one of claims 1-3, and a reagent for detecting the expression level of CAFRG. The CAFRG includes COX6A1 gene, ENOX1 gene, FERMT2 gene, NID1 gene, LOX gene, SNAI2 gene, GLI2 gene, ZNF154 gene, COX7A1 gene, NXPH3 gene, FRMD4A gene, SYT11 gene and ENTPD1 gene.

10. Use of the lung adenocarcinoma prognostic risk prediction model according to any one of claims 1-3 or the lung adenocarcinoma prognostic risk prediction system according to claim 9 in the preparation of a lung adenocarcinoma prognosis detection product.

Citation Information

Cited By

  • Hepatocellular carcinoma data processing method and system

    CN120600124A

  • A method and system for processing hepatocellular carcinoma data

    CN120600124B

  • Clinical prediction model construction method and system for treating oligometastatic non-small cell lung cancer through radioactive particle implantation

    CN121260462A