Method for constructing pancreatic cancer prognosis model based on endoplasmic reticulum stress (ERS) related lncRNA
By constructing a pancreatic cancer prognosis model based on ERS-related lncRNA, the existing models have solved the problems of complex data integration and limited prediction accuracy, achieving more accurate prognostic evaluation and the provision of personalized treatment plans.
Patent Information
- Application Number
- CN202510147589.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-30
AI Technical Summary
The existing pancreatic cancer prognosis model has complex data integration and limited prediction accuracy, making it difficult to fully capture multi-level biological characteristics, especially in the prediction of tumor proliferation, invasion and metastasis.
A pancreatic cancer prognosis model based on endoplasmic reticulum stress (ERS)-related lncRNA was constructed. By obtaining data from the TCGA database, standardized processing, differential analysis and co-expression analysis, 9 key lncRNAs were screened out, a multi-factor cox proportional hazard model was constructed, the patients' risk scores were calculated, and the patients were divided into high-risk group and low-risk group.
It significantly improves the accuracy of the evaluation of prognosis of pancreatic cancer patients, effectively distinguishes between high-risk groups and low-risk groups, and provides scientific basis for stratified management and personalized treatment strategies to assist in judging the patient's immune status and its potential response to immunotherapy.
Smart Images

Figure CN120072067A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of constructing a prognostic model for pancreatic cancer, and specifically to a method for constructing a prognostic model for pancreatic cancer based on endoplasmic reticulum stress (ERS)-related lncRNAs. Background Art
[0002] In the prognostic assessment of pancreatic cancer, bioinformatics models often identify survival-related molecular markers by integrating high-throughput transcriptome data and patients' clinical information. These methods mainly rely on a single type of marker (such as mRNA or lncRNA), or use combined analysis of multiple markers to construct risk prediction models. However, existing methods generally face problems of complex data integration and limited prediction accuracy, and it is difficult to comprehensively capture the multi-level biological characteristics of pancreatic cancer, especially in predicting the processes of tumor proliferation, invasion, and metastasis.
[0003] Although multi-marker models can increase the information coverage, data heterogeneity, sampling noise, and biological background differences often lead to limited model stability and reduced prediction ability. In addition, single-marker methods fail to fully reflect the complex biological processes in cancer progression, limiting the clinical application value and generalization ability of the models. Therefore, there is an urgent need for a model that can more accurately assess the prognostic risk to improve the accuracy and reliability of pancreatic cancer survival prediction. So we propose a method for constructing a prognostic model for pancreatic cancer based on endoplasmic reticulum stress (ERS)-related lncRNAs to solve the problems raised above.
[0004] The above information disclosed in this background art is only used to increase the understanding of the background art of the present invention. Therefore, it may include prior art that is not known to those of ordinary skill in the art. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for constructing a prognostic model for pancreatic cancer based on endoplasmic reticulum stress (ERS)-related lncRNAs to solve the problems raised in the above background art.
[0006] To achieve the above purpose, the present invention provides the following technical solution: A method for constructing a prognostic model for pancreatic cancer based on endoplasmic reticulum stress (ERS)-related lncRNAs, comprising the following steps:
[0007] Step 1: Obtain the transcriptome data and clinical information of pancreatic cancer patients from the Cancer Genome Atlas (TCGA) database;
[0008] Step 2: Standardize and preprocess the data obtained in Step 1 to form a sample dataset. Select 293 ERS genes, and use the limma R package and pheatmap R package to perform differential analysis and co-expression analysis on tumor samples and normal samples, identify 224 lncRNAs related to ERS, and visualize the co-expression relationship through volcano plots and heatmaps;
[0009] Step 3: Through LASSO regression and univariate Cox regression analysis, perform feature selection on the screened ERS-related lncRNAs, screen out 9 key lncRNAs for the final construction of a prognostic risk model for pancreatic cancer, obtain a scoring formula through a multivariate cox proportional hazards model to calculate the risk score of patients, and divide the patients into high-risk and low-risk groups according to the median of the scores;
[0010] Step 4: Use the Kaplan-Meier survival curve test to evaluate the survival differences between the high- and low-risk groups, draw risk score distribution plots, heatmaps, and forest plots to visually display the relationship between the expression of key lncRNAs and patient survival, and develop a nomogram based on the risk score and age and gender factors;
[0011] Step 5: Evaluate the relationship between the expression of ERS-related lncRNAs and the infiltration of immune cells in the pancreatic cancer microenvironment, analyze the differences in immune cell composition and immune function scores between the high- and low-risk groups, and at the same time, through tumor microenvironment (TME) and immune correlation analysis and immune checkpoint analysis, reveal the potential differences in immune therapy responsiveness between the high- and low-risk groups;
[0012] Step 6: Predict the sensitivity of high- and low-risk group patients to anticancer drugs through the oncoPredict package, combine the drug sensitivity data for analysis, and evaluate the drug response differences between the high- and low-risk groups.
[0013] Preferably, in the above Step 1, the transcriptome data and clinical information include but are not limited to follow-up time, survival status, age, gender, tumor grade, stage, and TNM stage.
[0014] Preferably, in the above Step 2, the dataset is divided into a 1:1 training set and test set for model construction and independent verification.
[0015] Preferably, in the above Step 3, the feature selection uses the Lasso regression algorithm.
[0016] Preferably, in the above Step 3, the feature selection threshold for ERS-related lncRNAs is the optimal range of Log lambda.
[0017] Compared with the prior art, the beneficial effects of the present invention are:
[0018] (1) The method of the present invention is simple, has high prediction accuracy and wide applicability. By screening and modeling ERS-related lncRNAs, it can effectively distinguish high-risk and low-risk groups of pancreatic cancer patients, significantly improve the accuracy of assessing the prognosis of patients, and verify the independent prognostic ability of the model.
[0019] (2) The model of the present invention can not only provide a scientific basis for the stratified management and personalized treatment strategies of patients, but also assist in judging the immune status of patients and their potential response to immunotherapy; there are significant differences in immune cell infiltration and drug sensitivity between high- and low-risk group patients, which helps to develop personalized treatment plans, improve treatment effects and reduce unnecessary medical expenses, and is suitable for large-scale clinical applications and precision medicine research.
[0020] The above summary is only for the purpose of the specification and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments and features described above, further aspects, embodiments and features of the present invention will be readily apparent by reference to the drawings and the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 :
[0022] A: Heatmap of the expression of 224 differentially expressed ERS-lncRNAs in tumor tissues and normal tissues;
[0023] B: Volcano plot of the differential expression levels of 224 differentially expressed ERS-lncRNAs in tumor tissues;
[0024] C: Forest plot of 57 ERS-related lncRNAs significantly associated with the prognosis of PAAD screened by univariate Cox analysis;
[0025] D: Heatmap of the expression levels of 57 lncRNAs in normal tissues and tumor tissues;
[0026] E: Relationship diagram between Log lambda and partial likelihood deviance;
[0027] F: Schematic diagram of the selection process of the optimal value of λ;
[0028] Figure 2 :
[0029] A: K-M survival curve of the survival status of high- and low-risk group patients in the total cohort;
[0030] B: K-M survival curve of the survival status of high- and low-risk group patients in the test set;
[0031] C: K-M survival curve of the survival status of high- and low-risk group patients in the training set;
[0032] D: In the total cohort, the risk scores of patients are arranged in descending order;
[0033] E: In the test set, the risk scores of patients are arranged in descending order;
[0034] F: In the training set, the risk scores of patients are arranged in descending order;
[0035] G: Scatter plot between the survival time (futime) and survival status (fustat) of patients in the total cohort;
[0036] H: Scatter plot between the survival time (futime) and survival status (fustat) of patients in the test set;
[0037] I: Scatter plot between the survival time (futime) and survival status (fustat) of patients in the training set;
[0038] J: Heatmap of the expression of 9 ERS-related lncRNAs in each PAAD patient in the total cohort;
[0039] K: Heatmap of the expression of 9 ERS-related lncRNAs in each PAAD patient in the test set;
[0040] L: Heatmap of the expression of 9 ERS-related lncRNAs in each PAAD patient in the training set;
[0041] Figure 3 :
[0042] A, B: Visualization graphs of the results of univariate and multivariate Cox regression analyses;
[0043] C: Receiver operating characteristic (ROC) curve graph;
[0044] D: ROC curves of the model for each clinical feature at 1 year, 3 years, and 5 years;
[0045] E: Kaplan-Meier survival curve graph of the survival status of patients in the high- and low-risk groups when age <= 65;
[0046] F: Kaplan-Meier survival curve graph of the survival status of patients in the high- and low-risk groups when age > 65;
[0047] G: Nomogram of the survival probability of patients in different survival years;
[0048] H: Calibration curve graph of the risk model drawn by fitting the model with clinical case characteristics and risk scores;
[0049] Figure 4 :
[0050] A: Heatmap of differentially expressed genes between high- and low-risk groups;
[0051] B: Volcano plot of gene expression levels;
[0052] C: Result graph of GO analysis;
[0053] D: Result graph of KEGG analysis;
[0054] E: Result graph of GSEA enrichment analysis. The pathways enriched in the low-risk group are mainly chemokine signaling pathway, neuroactive ligand-receptor interaction, primary immunodeficiency, steroid hormone biosynthesis, etc.;
[0055] F: Result graph of GSEA enrichment analysis. The pathways enriched in the high-risk group are mainly cell cycle, p53 signaling pathway, proteasome, etc.;
[0056] Figure 5 :
[0057] A, B, C: The ESTIMATE algorithm was used to evaluate the immune cell infiltration degree, stromal cell abundance and tumor purity of each tumor sample, and the corresponding immune score, stromal score and estimate score result graphs were obtained;
[0058] D: Result graph of the correlation analysis between immune cells and risk score;
[0059] E, F: The ssGSEA analysis was used to evaluate the immune cells and related immune functions of the two groups of patients;
[0060] G: Difference graph of the expression levels of immune checkpoint-related genes between high- and low-risk groups;
[0061] Figure 6 : Difference graph of the survival status of patients in the high-expression group and the low-expression group;
[0062] Figure 7 : Drug sensitivity analysis graph;
[0063] Figure 8 : All patients were subjected to consensus clustering analysis according to the expression levels of 9 ERS-related lncRNAs involved in model construction, and the clustering effect was optimal when k = 3;
[0064] Figure 9 :
[0065] A: Survival analysis was performed on patients belonging to three clusters to compare the OS differences after clustering;
[0066] B: Sankey diagram of the correspondence between sample subgroups and risk scores;
[0067] C, D, E, F: Perform principal component analysis (PCA) and t-distributed stochastic neighbor embedding (t-SNE) on the samples. The analysis results show that there are differences in different dimensions among the patients in the three clusters, while the dimensional differences between the patients in the high- and low-risk groups are not significant enough;
[0068] Figure 10 :
[0069] A, B, C: TME score difference diagrams between every two of the 3 clusters;
[0070] D: Heatmap of immune infiltration analysis of the clustered samples;
[0071] E: Diagram of differential expression analysis of immune checkpoint-related genes in the three clusters;
[0072] Figure 11 : Schematic diagrams of drugs with different sensitivities;
[0073] Figure 12 : Schematic diagram of the process of the present invention. Detailed implementation manners
[0074] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0075] Please refer to Figure 1 , a method for constructing a prognostic model for pancreatic cancer based on endoplasmic reticulum stress (ERS)-related lncRNA, comprising the following steps:
[0076] Step 1: Obtain the transcriptome data and clinical information of 182 pancreatic cancer patients from the Cancer Genome Atlas (TCGA) database; the transcriptome data and clinical information include, but are not limited to, follow-up time, survival status, age, gender, tumor grade, stage, and TNM stage;
[0077] Step 2: Standardize and preprocess the data obtained in Step 1 to form a sample data set. Finally, a total of 4 normal samples and 182 tumor sample data are included. Select 293 ERS genes, and use the limma R package and the pheatmap R package to perform differential analysis and co-expression analysis on the tumor samples and normal samples, identify 224 lncRNAs related to ERS, and visualize the co-expression relationship through volcano plots and heatmaps;
[0078] The dataset was divided into a training set and a test set at a ratio of 1:1 for model construction and independent verification;
[0079] Step 3: Through LASSO regression and univariate Cox regression analysis, feature selection was performed on the selected ERS-related lncRNAs, and 9 key lncRNAs were screened for the final construction of a pancreatic cancer prognosis risk model. A scoring formula was obtained through a multivariate cox proportional hazards model to calculate the risk score of patients, and patients were divided into a high-risk group and a low-risk group according to the median of the scores;
[0080] Step 4: The Kaplan-Meier survival curve test was used to evaluate the survival differences between the high- and low-risk groups. Risk score distribution plots, heatmaps, and forest plots were drawn to visually display the relationship between the expression of key lncRNAs and patient survival. A nomogram was developed based on the risk score and age and gender factors;
[0081] Step 5: The relationship between the expression of ERS-related lncRNAs and immune cell infiltration in the pancreatic cancer microenvironment was evaluated. The differences in immune cell composition and immune function scores between the high- and low-risk groups were analyzed. At the same time, through tumor microenvironment (TME) and immune correlation analysis and immune checkpoint analysis, the potential differences in immune therapy responsiveness between the high- and low-risk groups were revealed;
[0082] Step 6: The oncoPredict package was used to predict the sensitivity of high- and low-risk group patients to anti-cancer drugs. Combined with drug sensitivity data analysis, the drug response differences between the high- and low-risk groups were evaluated.
[0083] Screening of differentially expressed ERS-related lncRNAs in PAAD
[0084] Transcriptome data and corresponding clinical information of 182 PAAD (pancreatic cancer) patients were downloaded from the TCGA database. Based on existing research, 293 ERS genes were selected, and 1551 ERS-related lncRNAs were identified through co-expression analysis with these genes. Further analysis of the expression differences of these ERS-related lncRNAs between normal samples and tumor tissues showed that a total of 224 ERS-related lncRNAs with significant expression differences (|logFC|≥1 and FDR<0.05) were screened out in PAAD tumors and normal tissues. Among them, 112 lncRNAs were down-regulated in tumor tissues and 112 lncRNAs were up-regulated. The heatmap presented the expression of 224 differentially expressed ERS-lncRNAs in tumor tissues and normal tissues, and the volcano plot presented the differential expression levels of these lncRNAs in tumor tissues ( Figure 1 -A,B).
[0085] Construction of a prognostic model based on ERS-related lncRNAs associated with PAAD prognosis
[0086] Next, we sorted out the survival data of PAAD patients obtained previously. After excluding samples with missing information, we finally retained 171 patients and randomly divided these samples into a training set (n = 86) and a test set (n = 85) for subsequent analysis. We extracted data such as the survival time and survival status of the patients and merged them with the lncRNA expression data. Through univariate Cox analysis, 57 ERS-related lncRNAs significantly associated with PAAD prognosis were screened out ( Figure 1 -C), and the heatmap shows the expression levels of these 57 lncRNAs in normal and tumor tissues ( Figure 1 -D). We further screened out 9 lncRNAs finally used for constructing the prognostic model through LASSO regression (model parameters are shown in Figure 1 -E,F) and multivariate Cox regression analysis: LINC00852, AL117382.1, AC019186.1, AC002401.4, AC092171.3, SENCR, GUSBP11, AC001226.1, AP005233.2. According to the prognostic model we constructed, the risk score is calculated according to the following formula: risk score = LINC00852*(-2.54822542090913)+AL117382.1*(-0.845372252572961)+AC019186.1*(2.80999868600992)+AC002401.4*(0.764501281953383)+AC092171.3*(-1.23253927055708)+SENCR*(-2.36655618495886)+GUSBP11*(-4.03192901171483)+AC001226.1*(-3.47015206599368)+AP005233.2*(0.37518317427093).
[0087] Validation and efficacy evaluation of the prognostic model
[0088] To verify the accuracy of the constructed prognostic prediction model, we first performed survival analysis. The patients in each cohort were divided into a high-risk group and a low-risk group according to the 50th percentile of the risk score. The Kaplan-Meier method was used to calculate the survival probabilities of the high- and low-risk group patients in each cohort and compare the differences in their overall survival (OS). The analysis results showed that in each cohort, compared with the high-risk group, the low-risk group patients had significantly longer OS( Figure 2-A, B, C), and the difference was statistically significant (P < 0.05). At the same time, we also used a scatter plot to present the risk score distribution of patients in different survival states. It can be intuitively seen that in the test set, patients with the end event of death were mostly concentrated in the high-risk group, and this trend could also be observed in the other two cohorts ( Figure 2 -D, E, F, G, H, I), which indicates that the mortality rate of patients increases with the increase of risk value. In addition, the expression of 9 ERS-related lncRNAs in each PAAD patient was presented in a heat map ( Figure 2 -J, K, L). We observed that each lncRNA had the same expression trend in different cohorts.
[0089] Next, in order to evaluate the independent prognostic ability of the prognostic model, we performed univariate and multivariate Cox regression analyses on other clinical factors and risk scores. The results showed that compared with clinical factors such as gender and stage, both risk score and age were independent influencing factors for patient prognosis ( Figure 3 -A, B), and the higher the risk score, the greater the risk of the patient (HR = 1.038, 95% CI: 1.021 - 1.056, P < 0.001). Next, we plotted the receiver operating characteristic (ROC) curve to evaluate the prognostic prediction accuracy of the constructed model for patients with different survival times. The results showed that the area under the curve (AUC) at 1 year, 3 years, and 5 years were 0.747, 0.834, and 0.954 respectively ( Figure 3 -C), which indicates that our model has good precision, and the prediction accuracy of the risk score (AUC = 0.834) was significantly higher than that of clinical factors such as age (AUC = 0.662), gender (AUC = 0.458), Grade (AUC = 0.0604), and Stage (AUC = 0.588) ( Figure 3 -D). In addition, we also analyzed the survival status of high- and low-risk patients in different age groups with age as the stratification factor. The results showed that in patients 65 years old and below, the OS of high-risk group patients was significantly shorter than that of the low-risk group (P < 0.001), and the same result was also observed in patients over 65 years old (P < 0.05), which indicates that the risk score is applicable to patients of different ages. These analysis results indicate that the model we constructed has good robustness.
[0090] Development and evaluation of a nomogram for PAAD clinical prognosis
[0091] To achieve the clinical actual prognosis assessment of PAAD patients, we developed a nomogram based on risk scores and factors such as age and gender. By scoring each index of the patient and finally summing them up to obtain the total score, the survival probability of the patient at different survival years can be obtained. For example, the probability that a patient with a comprehensive score of 741 survives for more than 1 year is 0.96, the probability of surviving for more than 2 years is 0.854, and the probability of surviving for more than 3 years is 0.82( Figure 3 -G). And to evaluate the prediction accuracy of this nomogram, we corrected it, and the C-index of the obtained calibration curve is 0.717 (95% CI: 0.672 - 0.762), indicating that the accuracy of this model is good( Figure 3 -H).
[0092] Signal pathway and biological function enrichment analysis
[0093] We used heat maps( Figure 4 -A) and volcano plots( Figure 4 -B) to show the differentially expressed genes (DEGs) between the high- and low-risk groups and the expression levels of these genes. Next, gene ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analyses were performed on the above DEGs to explore potential biological functions and signal pathways. The results of GO analysis showed that these genes were mainly enriched in pathways such as regulation of chemical synaptic transmission, regulation of immune response signaling pathway, activation signal of immune response, etc. at the biological process level, mainly enriched in the outer side of the plasma membrane, plasma membrane signal receptor complex, immunoglobulin complex, etc. at the cellular component level, and mainly enriched in activities of passive transmembrane transporter proteins, single atom ion-gated channel activities, etc. at the molecular function level( Figure 4 -C). The results of KEGG analysis showed that these genes were mainly enriched in pathways such as cell adhesion molecules, MAPK signaling pathway, T cell receptor signaling pathway, pancreatic secretion, etc( Figure 4 -D). To further explore the differential enrichment of biological functions and signal pathways in the high- and low-risk groups, we further performed GSEA enrichment analysis. The analysis results showed that the pathways mainly enriched in the low-risk group were chemokine signaling pathway, neuroactive ligand-receptor interaction, primary immunodeficiency, steroid hormone biosynthesis, etc( Figure 4 -E), while the pathways mainly enriched in the high-risk group were cell cycle, P53 signaling pathway, proteasome, etc( Figure 4 -F).
[0094] Tumor microenvironment analysis
[0095] The tumor microenvironment (TME) plays an important role in tumorigenesis, development, and disease progression. To explore the differences in the tumor microenvironment between high- and low-risk groups of pancreatic cancer patients, we used the ESTIMATE algorithm to evaluate the immune cell infiltration, stromal cell abundance, and tumor purity of each tumor sample, and accordingly obtained three scores: immune score, stromal score, and estimate score ( Figure 5 -A, B, C). The analysis results showed that the above three scores of the low-risk group patients were higher than those of the high-risk group patients, but only the difference in the immune score was statistically significant (P < 0.05), indicating that the low-risk group patients had higher immune cell infiltration.
[0096] Immune correlation analysis and immune analysis
[0097] To further investigate the relationship between immune cell infiltration and the risk score of pancreatic cancer patients, we performed a correlation analysis between immune cells and the risk score using 7 algorithms ( Figure 5 -D). The research results showed that most infiltrating immune cells such as T cell CD4+, T cell CD8+, Macrophage, B cell Monocyte, NK cell, etc. were all significantly negatively correlated with the risk score (cor < 0 and P < 0.05), while the infiltration of cells such as Macrophage M0, Macrophage M2, Neutrophil, etc. was significantly positively correlated with the risk score (cor > 0 and p < 0.05). Next, we divided the samples into high-expression and low-expression groups based on the expression levels of infiltrating immune cells, and compared the differences in the survival status of patients in the two groups through survival analysis ( Figure 6 ). We found that the overall survival time of patients in the high-expression group of cells such as B cell plasma, Granulocyte-monocyte progenitor, Macrophage M0, Macrophage M1, T cell CD4+ Th2, uncharacterized cell, etc. was shorter than that in the low-expression group, while Class-switched memory B cell, Macrophage M2, Monocyte, T cell CD4+ central memory, T cell CD4+ etc. were such that the OS of the high-expression group was longer than that of the low-expression group, and the above differences were statistically significant (P < 0.05).
[0098] To further explore the immunotherapy effects of pancreatic cancer patients in different risk groups, we used ssGSEA analysis to evaluate the immune cells and related immune functions of the two groups of patientsFigure 5 -E, F). The analysis results showed that in immune cells, the ssGSEA scores of B cells, CD8+ T cells, Mast cells, and tumor-infiltrating lymphocytes (TIL) in the low-risk group were higher than those in the high-risk group, and the differences were statistically significant (P<0.05), indicating that the relative abundances of these cells were higher in the low-risk group than in the high-risk group. In terms of immune functions, the scores of functions such as Cytolytic activity, T cell co-stimulation, and type II interferon response were significantly higher in the low-risk group than in the high-risk group (P<0.05), indicating that these immune functions were upregulated in the low-risk group. However, the ssGSEA score of type I interferon response was significantly higher in the low-risk group than in the high-risk group (P<0.05), suggesting that this immune function was upregulated in the high-risk group. Next, we analyzed the differences in the expression levels of immune checkpoint-related genes between the high- and low-risk groups ( Figure 5 -G). From the analysis results, we observed that there were 25 immune checkpoint-related genes with statistically significant differences in expression between the high- and low-risk groups (P<0.05). Among them, the expressions of genes such as BTLA, CD200R1, CD40LG, CD27, CD160, and LAG3 were significantly upregulated in the low-risk group, while the expressions of genes such as HHLA2, CD70, TNFSF4, TNFSF9, and CD276 were significantly upregulated in the high-risk group. These findings of ours may provide potential new targets for the immunotherapy of pancreatic cancer patients.
[0099] Drug sensitivity analysis based on the ERS-related lncRNA prognostic model
[0100] To promote the individualized treatment of pancreatic cancer patients and guide clinical drug use, we performed drug sensitivity analysis on the samples. By calculating the IC50 of different drugs in the high- and low-risk groups, we compared and obtained drugs with significantly different sensitivities ( Figure 7)。The analysis results showed that compared with the high-risk group, patients in the low-risk group were more sensitive to drugs such as Afuresertib (AKT kinase inhibitor), AZ6102 (TNKS1 / 2 inhibitor), Rapamycin (mTOR inhibitor), AZD8055 (mTOR inhibitor), BMS-754807 (insulin receptor inhibitor), Cediranib, Gefitinib, Mitoxantrone, Navitoclax (Bcl-2 inhibitor), Oxaliplatin, Vincristine, etc. (P<0.05). However, the IC50 of drugs such as Trametinib (MEK1 / 2 inhibitor), SCH772984 (ERK1 / 2 inhibitor), ERK6604 (MAPK signaling pathway inhibitor) for the high-risk group was significantly lower than that of the low-risk group (P<0.05), indicating that patients in the high-risk group were more sensitive to these drugs than those in the low-risk group. The above findings may help clinicians develop personalized treatment plans.
[0101] Consensus clustering analysis based on ERS-related lncRNAs
[0102] To further classify and compare pancreatic cancer patients at the molecular level, we performed consensus clustering analysis on all patients according to the expression levels of the 9 ERS-related lncRNAs involved in model construction mentioned above. After comparing the clustering effects of different numbers of clusters (k), we found that the clustering effect was optimal when k = 3 ( Figure 8 ). Then we clustered all PAAD patient samples into three clusters: cluster 1 (n = 67), cluster 2 (n = 50), cluster 3 (n = 54), and performed survival analysis on the patients belonging to the three clusters to compare the OS differences after clustering ( Figure 9 -A), and the results showed that there were significant differences in OS among the patients in the three clusters (P<0.05), and the OS of C3 patients was the longest and that of C2 was the shortest. We also used a Sankey diagram to show the correspondence between sample subgroups and risk scores ( Figure 9 -B), and we found that most samples in C3 were low-risk patients, while most samples in C2 were high-risk patients, and the samples in C1 were similarly distributed in the two risk groups. To further explore the discrimination of samples under different clustering bases, we performed principal component (PCA) and t-distributed stochastic neighbor embedding (t-SNE) analyses on the samples. From the analysis results, we observed that there were differences in different dimensions among the patients in the three clusters, but the dimensional differences between the patients in the high- and low-risk groups were not significant enough ( Figure 9 -C, D, E, F).
[0103] Next, we also performed TME analysis on the clustered samples and compared the TME score differences between any two of the three clusters ( Figure 10-A, B, C), the results of the stromal score showed that there were significant differences in the scores between Cluster 1 and Cluster 2, with Cluster 1 having a higher score than Cluster 2, and the score of Cluster 3 was also significantly higher than that of Cluster 2 (P < 0.05), while there was no significant difference in the stromal score between Cluster 3 and Cluster 2 (P > 0.05); there were no significant differences in the immune scores between any two of the three clusters; and the ESTIMATE scores of Cluster 1 and Cluster 3 were significantly higher than that of Cluster 2 (P < 0.05), and there was no significant difference in the ESTIMATE score between Cluster 3 and Cluster 2 (P > 0.05).
[0104] In addition, we also performed immune infiltration analysis on the clustered samples. From the heatmap, we observed a higher degree of immune cell infiltration in Cluster 3 ( Figure 10 -D). Subsequently, we performed differential expression analysis of immune checkpoint-related genes in the three clusters ( Figure 10 -E), and the results showed that genes such as TNFRSF14, TBFRSF25, CD70, CD40, LGALS9 were significantly upregulated in Cluster 2 (P < 0.05), while the expressions of PDCD1LG2, CD200, and NRP1 were significantly downregulated in Cluster 2. Finally, we performed drug sensitivity analysis on the samples of the three subgroups. By comparing the IC50 values between any two of the three groups, 70 drugs with different sensitivities among the three clusters were screened out (
[0105] ). We found that among these three clusters, Cluster 3 was more sensitive to most drugs, including AZD4547 (FGFRs inhibitor), AZD5363 (AKT kinase inhibitor), AZD8055 (mTOR inhibitor), BMS-754807 (insulin receptor inhibitor), cediranib, axitinib, etc., while drugs such as MK-1775 (Wee1 inhibitor), PD0325901 (MEK inhibitor), gemcitabine, SCH772984 (ERK12 inhibitor), trametinib, etc. were more sensitive to Cluster 2. These results suggest that patients classified as Cluster 3 may have a better prognosis and may also obtain more favorable treatment effects compared to patients in Cluster 2. Figure 11 )
[0106] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for constructing a pancreatic cancer prognosis model based on endoplasmic reticulum stress (ERS)-related lncRNA, characterized in that: The steps include: Step 1: Obtain transcriptome data and clinical information of pancreatic cancer patients from the Cancer Genome Atlas TCGA database; Step 2: The data obtained in step 1 were standardized and preprocessed to form a sample data set. 293 ERS genes were selected, and differential analysis and co-expression analysis were performed on tumor samples and normal samples using the limma R package and pheatmap R package. 224 ERS-related lncRNAs were identified, and the co-expression relationship was visualized using volcano plots and heat maps; Step 3: LASSO regression and univariate Cox regression analysis were used to perform feature selection on the ERS-related lncRNAs screened out, and 9 key lncRNAs were screened out for the final construction of the pancreatic cancer prognostic risk model. The scoring formula was obtained through the multivariate Cox proportional hazard model to calculate the patient's risk score, and the patients were divided into high-risk group and low-risk group according to the median score; Step 4: Use the Kaplan-Meier survival curve test to evaluate the survival difference between the high-risk and low-risk groups, draw risk score distribution maps, heat maps, and forest maps to visualize the relationship between the expression of key lncRNAs and patient survival, and develop a nomogram based on risk scores and age and gender factors; Step 5: Evaluate the relationship between ERS-related lncRNA expression and immune cell infiltration in the pancreatic cancer microenvironment, analyze the differences in immune cell composition and immune function scores between high-risk and low-risk groups, and reveal potential differences in immunotherapy responsiveness between high-risk and low-risk groups through tumor microenvironment (TME) and immune correlation analysis and immune checkpoint analysis; Step 6: Use the oncoPredict package to predict the sensitivity of patients in the high-risk and low-risk groups to anticancer drugs, and analyze the drug sensitivity data to evaluate the differences in drug response between the high-risk and low-risk groups.
2. The method for constructing a pancreatic cancer prognosis model based on endoplasmic reticulum stress (ERS)-related lncRNA according to claim 1, characterized in that: In step 1, the transcriptome data and clinical information include but are not limited to follow-up time, survival status, age, gender, tumor grade, stage and TNM stage.
3. The method for constructing a pancreatic cancer prognosis model based on endoplasmic reticulum stress (ERS)-related lncRNA according to claim 1, characterized in that: In step 2, the data set is divided into a training set and a test set in a 1:1 ratio for model construction and independent verification.
4. The method for constructing a pancreatic cancer prognosis model based on endoplasmic reticulum stress (ERS)-related lncRNA according to claim 1, characterized in that: In step 3, the feature selection uses the Lasso regression algorithm.
5. The method for constructing a pancreatic cancer prognosis model based on endoplasmic reticulum stress (ERS)-related lncRNA according to claim 1, characterized in that: In step 3, the ERS-related lncRNA feature selection threshold is the optimal range of Log lambda.
Citation Information
Cited By
Hepatocellular carcinoma data processing method and system
CN120600124A