Construction method of pancreatic cancer prognosis model based on transporter-related lncRNA

By constructing a pancreatic cancer prognosis model based on migratory-related lncRNA, the problems of complex data integration and limited prediction accuracy in the existing technology are solved, and high-precision prognosis evaluation and personalized treatment strategies for pancreatic cancer patients are realized, which improves the biological relevance and clinical application value of the model.

CN120299533APending Publication Date: 2025-07-11NANCHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510374011.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The data integration of the prior art in the prognosis assessment of pancreatic cancer is complex, with limited prediction accuracy, and it is difficult to fully capture multi-level biological characteristics, especially in the prediction of tumor migration and invasion processes. The single marker method fails to fully reflect the complex biological processes in cancer progression, limiting the clinical application value and generalization ability of the model.

Method used

A pancreatic cancer prognosis model based on migrant-related lncRNA was constructed. By obtaining data from the TCGA database, standardizing and pre-processing, 180 migrant-related lncRNAs were screened out, and feature selection was used using Lasso regression model, risk scoring model was constructed, and the patient's survival risk was evaluated in combination with immune cell infiltration and drug sensitivity analysis.

Benefits of technology

It significantly improves the accuracy of evaluation of prognosis in pancreatic cancer patients, improves the accuracy of survival risk prediction, provides personalized treatment strategies and scientific basis for immunotherapy, enhances the biological relevance and clinical practicality of the model, and is suitable for large-scale clinical applications and precision medical research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299533A_ABST
    Figure CN120299533A_ABST
Patent Text Reader

Abstract

The invention discloses a construction method of a pancreatic cancer prognosis model based on transference body related lncRNA, and relates to the technical field of pancreatic cancer prognosis model construction. According to the invention, through screening and modeling of the transference body-related lncRNA, a high-risk group and a low-risk group of pancreatic cancer patients are effectively distinguished, the evaluation precision of patient prognosis is significantly improved, and the independent prognosis ability of the model is verified; by focusing on migration and invasion characteristics of pancreatic cancer, the accuracy of survival risk prediction is remarkably improved. According to the method, the blank in existing prognosis evaluation is filled, a scientific basis can be provided for hierarchical management and personalized treatment strategies of pancreatic cancer patients, the immune state of the patients and the potential response of the patients to immunotherapy can be judged in an auxiliary manner, and the biological correlation and clinical practicability of the model are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of constructing a prognostic model for pancreatic cancer, and particularly to a method for constructing a prognostic model for pancreatic cancer based on migrasome-related lncRNA. Background Art

[0002] In the prognostic assessment of pancreatic cancer, bioinformatics models often identify survival-related molecular markers by integrating high-throughput transcriptome data and patients' clinical information. These methods mainly rely on a single type of marker (such as mRNA or lncRNA), or use a combined analysis of multiple markers to construct a risk prediction model. However, existing methods generally face problems such as complex data integration and limited prediction accuracy, and it is difficult to comprehensively capture the multi-level biological characteristics of pancreatic cancer, especially in the prediction of tumor migration and invasion processes.

[0003] Although multi-marker models can increase the information coverage, the heterogeneity of data, sampling noise, and biological background differences often lead to limited model stability and decreased prediction ability. In addition, the method of single marker fails to fully reflect the complex biological processes in cancer progression, limiting the clinical application value and generalization ability of the model. Therefore, there is an urgent need for a model that can more accurately evaluate the prognostic risk to improve the accuracy and reliability of pancreatic cancer survival prediction. So we propose a method for constructing a prognostic model for pancreatic cancer based on migrasome-related lncRNA to solve the problems mentioned above.

[0004] The above information disclosed in this background art is only used to increase the understanding of the background art of the present invention. Therefore, it may include prior art that is not known to those of ordinary skill in the art. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for constructing a prognostic model for pancreatic cancer based on migrasome-related lncRNA to solve the problems in the above background art that existing methods generally face complex data integration and limited prediction accuracy, it is difficult to comprehensively capture the multi-level biological characteristics of pancreatic cancer, especially in the prediction of tumor migration and invasion processes, and the method of single marker fails to fully reflect the complex biological processes in cancer progression, limiting the clinical application value and generalization ability of the model.

[0006] To achieve the above purpose, the present invention provides the following technical solution: A method for constructing a prognostic model for pancreatic cancer based on migrasome-related lncRNA, comprising the following steps:

[0007] Step 1: Obtain the transcriptome data and clinical information of pancreatic cancer patients from the Cancer Genome Atlas (TCGA) database;

[0008] Step 2: Standardize and preprocess the data obtained in Step 1 to form a sample data set, extract the expression levels of 8 key genes related to migrasomes, perform differential analysis and co-expression analysis on tumor samples and normal samples using the limma and edgeR R packages, screen out 180 lncRNAs related to migrasomes, and use the ggalluvial package to generate a Sankey diagram to visualize the co-expression relationship;

[0009] Step 3: Use the Lasso regression model in the glmnet R package to perform feature selection on the migrasome-related lncRNAs screened in Step 2, screen out 4 key lncRNAs and construct a prognostic risk model for pancreatic cancer, obtain a scoring formula through a multifactorial cox proportional hazards model to calculate the risk score of patients, and divide the patients into high-risk and low-risk groups according to the median of the scores;

[0010] Step 4: Use the Kaplan-Meier survival curve and Log-rank test to evaluate the survival differences between the high- and low-risk groups, draw survival status diagrams, risk score distribution diagrams, and heatmaps to visually display the relationship between the expression of key lncRNAs and patient survival, and develop a nomogram based on the risk score and age and gender factors;

[0011] Step 5: Evaluate the relationship between the expression of migrasome-related lncRNAs and the infiltration of immune cells in the pancreatic cancer microenvironment, analyze the differences in the composition of immune cells and immune function scores between the high- and low-risk groups, and at the same time, through tumor mutation burden (TMB) and immune escape analysis, reveal the potential differences in immune therapy responsiveness between the high- and low-risk groups;

[0012] Step 6: Predict the sensitivity of high- and low-risk group patients to anti-cancer drugs through the oncoPredict package, analyze in combination with drug sensitivity data, and evaluate the drug response differences between the high- and low-risk groups.

[0013] Preferably, in the above Step 1, the transcriptome data and clinical information include but are not limited to mRNA expression profiles, lncRNA expression profiles, mutation information, and the age, gender, TNM stage, and survival data of patients.

[0014] Preferably, in the above Step 2, the data set is divided into a training set and a test set at a ratio of 1:1 for model construction and independent verification.

[0015] Preferably, in the above Step 3, the feature selection uses the Lasso regression algorithm.

[0016] Preferably, in the above Step 3, the feature selection threshold for migrasome-related lncRNAs is the optimal range of Log lambda.

[0017] Compared with the prior art, the beneficial effects of the present invention are:

[0018] (1) Through the screening and modeling of migrasome-related lncRNAs, the present invention effectively distinguishes high-risk and low-risk groups of pancreatic cancer patients, significantly improves the accuracy of patient prognosis assessment, and verifies the independent prognostic ability of the model; by focusing on the migratory and invasive characteristics of pancreatic cancer, the accuracy of survival risk prediction is significantly improved.

[0019] (2) The present invention fills the gap in existing prognosis assessment. It can not only provide a scientific basis for the stratified management and personalized treatment strategies of pancreatic cancer patients, but also assist in judging the immune status of patients and their potential response to immunotherapy, enhancing the biological relevance and clinical practicability of the model.

[0020] (3) It is known from the present invention that there are significant differences in immune cell infiltration, tumor mutational burden (TMB), and drug sensitivity between high- and low-risk group patients, which helps to formulate personalized treatment plans, improve treatment effects and reduce unnecessary medical expenses, and is suitable for large-scale clinical applications and precision medicine research.

[0021] The above summary is only for the purpose of the specification and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features of the present invention will become apparent by reference to the drawings and the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 : Schematic diagram of the process of the present invention;

[0023] Figure 2 :

[0024] A: Sankey diagram visualizing the flow relationship between migrasome-related genes and prognosis-related lncRNAs (migrasome-related genes are on the right and above, and migrasome prognosis-related lncRNAs are below);

[0025] B: Correlation heatmap visualizing the co-expression analysis between migrasome-related genes and prognosis-related lncRNAs (10 migrasome-related genes are on the left, and prognosis-related lncRNAs are below; red color indicates statistical significance, and the more * in the box, the higher the correlation between the two);

[0026] C: Forest plot of 45 migrasome prognosis-related IncRNAs significantly associated with the prognosis of pancreatic cancer selected by univariate Cox analysis (green represents protective factors, and red represents risk factors);

[0027] D: Relationship diagram between Log lambda and partial likelihood deviation. By determining the relationship between partial likelihood deviation and Log lambda, the optimal penalty parameter range is determined, and then the optimal variables are selected and the results are visualized;

[0028] E: Schematic diagram of the selection process of the optimal value of λ;

[0029] Figure 3 : Survival assessment of pancreatic cancer was performed in the total cohort, test set, and training set. Patients were divided into high-risk and low-risk groups according to the median level of the risk score (blue represents the low-risk group, and red represents the high-risk group). Among them,

[0030] A: Heatmap showing the expression of 4 migratory prognostic-related lncRNAs in each pancreatic cancer patient in the total cohort;

[0031] B: Heatmap showing the expression of 4 migratory prognostic-related lncRNAs in each pancreatic cancer patient in the test set;

[0032] C: Heatmap showing the expression of 4 migratory prognostic-related lncRNAs in each pancreatic cancer patient in the training set;

[0033] D: In the total cohort, the risk scores of patients are arranged in descending order;

[0034] E: In the test set, the risk scores of patients are arranged in descending order;

[0035] F: In the training set, the risk scores of patients are arranged in descending order;

[0036] G: Scatter plot between the survival time (futime) and survival status (fustat) of patients in the total cohort;

[0037] H: Scatter plot between the survival time (futime) and survival status (fustat) of patients in the test set;

[0038] I: Scatter plot between the survival time and survival status of the training set;

[0039] Figure 4 : Patients with pancreatic cancer with different clinical characteristics were divided into high- and low-risk groups, and Kaplan-Meier survival curves were plotted (P < 0.05 was statistically significant; blue was the low-risk group, and red was the high-risk group). Among them,

[0040] A: K-M survival curve of the survival status of high- and low-risk group patients in the total cohort (P < 0.001 was statistically significant);

[0041] B: Kaplan-Meier survival curves of the survival status of patients in the high- and low-risk groups in the test set (statistically significant difference with P < 0.05);

[0042] C: Kaplan-Meier survival curves of the survival status of patients in the high- and low-risk groups in the training set (statistically significant difference with P < 0.001);

[0043] D: Kaplan-Meier survival curves of the progression-free survival (PFS) between the high- and low-risk groups of patients (statistically significant difference with P < 0.001);

[0044] E: Kaplan-Meier survival curves of the survival status of high- and low-risk group patients when age > 65 (statistically significant difference with P < 0.05);

[0045] F: Kaplan-Meier survival curves of the survival status of high- and low-risk group patients when age <= 65 (statistically significant difference with P < 0.001);

[0046] G: Kaplan-Meier survival curves of the survival status of high- and low-risk group female patients (statistically significant difference with P < 0.001);

[0047] H: Kaplan-Meier survival curves of the survival status of high- and low-risk group male patients (statistically significant difference with P < 0.05);

[0048] I: Kaplan-Meier survival curves of the survival status of high- and low-risk group patients with tumor grade 1-2 (statistically significant difference with P < 0.001);

[0049] J: Kaplan-Meier survival curves of the survival status of high- and low-risk group patients with tumor grade 3-4 (statistically significant difference with P < 0.05);

[0050] K: Kaplan-Meier survival curves of the survival status of high- and low-risk group patients with tumor stage I-II (statistically significant difference with P < 0.001);

[0051] L: Kaplan-Meier survival curves of the survival status of high- and low-risk group patients with tumor stage III-IV (P > 0.05 indicates no significant statistical difference);

[0052] Figure 5 : To validate the constructed prognostic model, where

[0053] A: Receiver operating characteristic (ROC) curves of each clinical feature (including gender, age, tumor grade, tumor stage) and risk score (an area under the curve (AUC value) greater than 0.5 indicates high model prediction accuracy);

[0054] B: The ROC curves of the model for each clinical feature at 1 year, 3 years, and 5 years (an area under the curve (AUC value) greater than 0.5 indicates high prediction accuracy of the model). The AUC values for the three time periods are 0.670, 0.816, and 0.928 respectively, indicating that the model has relatively high accuracy in predicting the prognosis of colorectal cancer patients;

[0055] C, D: Univariate analysis and multivariate analysis including pancreatic cancer and clinical factors, and visualizing the results, indicating that the risk score, age risk score, and age are independent factors affecting the prognosis of patients. (HR: hazard ratio, CI: confidence interval; HR = 1.207, 95% CI: 1.073 - 1.357, P < 0.001);

[0056] E: C-index curve graph of gender, age (65 years old), tumor grade (G1 - 2 / G3 - 4), tumor stage (stage1 - 2 / 3 - 4) and survival time (a C-index value greater than 0.5 represents good prediction performance of the model);

[0057] F: Calibration curves of the risk model drawn by fitting the model with clinical case characteristics and risk scores at 1 year, 3 years, and 5 years;

[0058] G: Scoring each index of the patient, adding up the scores of each index to obtain a total score to evaluate the survival rate of the patient, shown as a nomogram of the survival probability of patients in different survival years (the left side shows each clinical index, the upper side shows the corresponding scores, and the lower side shows the survival rate of patients corresponding to the total score);

[0059] Figure 6 :

[0060] A, B: Using Gene Ontology (GO) enrichment analysis to evaluate the potential functions of lncRNA-corresponding genes in biological processes (BP), cellular components (CC), and molecular functions (MF), and visualizing them using bar charts and bubble charts respectively (p-value and q-value both less than 0.05 indicate statistical significance);

[0061] C, D: Using Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analysis to explore the functional enrichment of lncRNA-corresponding genes in related signaling pathways, and visualizing them using bar charts and bubble charts respectively (p-value and q-value both less than 0.05 indicate statistical significance);

[0062] E: Three-dimensional principal component analysis (PCA) graph of high- and low-risk groups of genes related to the migrasome in pancreatic cancer;

[0063] F: Three-dimensional principal component analysis (PCA) plot of high- and low-risk groups of lncRNAs related to the prognosis of pancreatic cancer migrasomes;

[0064] G: Three-dimensional principal component analysis (PCA) plot of high- and low-risk groups of lncRNAs related to the risk of pancreatic cancer migrasomes;

[0065] H: Gene set enrichment analysis (GSEA) curve plot of the top 5 significantly expressed biological pathways in the low-risk group based on the KEGG gene set;

[0066] I: Gene set enrichment analysis (GSEA) curve plot of the top 5 significantly expressed biological pathways in the high-risk group based on the GO gene library;

[0067] J: Gene set enrichment analysis (GSEA) curve plot of the top 5 significantly expressed biological pathways in the low-risk group based on the GO gene library;

[0068] Figure 7 : Using the ESTIMATE algorithm and the CIBERSORT algorithm to evaluate the tumor microenvironment (TME) score and immune cell infiltration respectively, and analyze the gene mutation characteristics (blue for the low-risk group, red for the high-risk group). Among them,

[0069] A: Schematic diagram of immune cell infiltration in the high-risk and low-risk groups (the right side shows the immune cell types);

[0070] B: Box plot showing the expression differences of each immune cell type in the high-risk and low-risk groups (*The more, the more significant the statistical difference);

[0071] C: Box plot showing the differences between the high- and low-risk groups in different immune function scores (*The more, the more significant the statistical difference);

[0072] D: Violin plot showing the differences between the high-risk and low-risk groups in different TME scores (immune score, stromal score, and ESTIMATE composite score) (P > 0.05 indicates no significant statistical difference);

[0073] E: Tumor mutation overview map of high-frequency mutated genes in the high-risk group (the left side shows the mutated genes, and different colors below represent different mutation types);

[0074] F: Tumor mutation overview map of high-frequency mutated genes in the low-risk group (the left side shows the mutated genes, and different colors below represent different mutation types);

[0075] Figure 8 :

[0076] A: Violin plot of tumor mutation burden (TMB) between the high- and low-risk groups (P > 0.05 indicates no significant statistical difference);

[0077] B: Violin plot of Tumor Immune Dysfunction and Exclusion score (TIDE) between high- and low-risk groups for predicting tumor immune escape ability (P>0.05 indicates no significant statistical difference);

[0078] C: Kaplan-Meier survival curve between high- and low-tumor mutation burden (TMB) groups (P<0.05 indicates significant statistical difference);

[0079] D: Kaplan-Meier survival curves among high-risk group with high TMB, high-risk group with low TMB, low-risk group with high TMB, and low-risk group with low TMB (P<0.001 indicates significant statistical difference);

[0080] Figure 9 : Sensitivity analysis of 49 drugs for patients in high- and low-risk groups of pancreatic cancer (P<0.05 indicates significant statistical difference, that is, there are differences in the efficacy of drugs between high- and low-risk groups). Detailed implementation manners

[0081] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0082] Embodiment 1

[0083] Please refer to Figure 1 , a method for constructing a prognostic model of pancreatic cancer based on migrasome-related lncRNA, comprising the following steps:

[0084] Step 1: Obtain the transcriptome data and clinical information of 185 pancreatic cancer patients from the Cancer Genome Atlas (TCGA) database;

[0085] The transcriptome data and clinical information include but are not limited to mRNA expression profiles, lncRNA expression profiles, mutation information, and the age, gender, TNM stage, and survival data of the patients;

[0086] Step 2: Standardize and preprocess the data obtained in Step 1 to form a sample data set, extract the expression levels of 8 key genes related to migrasomes, perform differential analysis and co-expression analysis on tumor samples and normal samples using limma and edgeR R packages, screen out 180 lncRNAs related to migrasomes, and use the ggalluvial package to generate a Sankey diagram to visualize the co-expression relationship;

[0087] The dataset was divided into a 1:1 training set and a test set for model construction and independent verification;

[0088] Step 3: Use the Lasso regression model in the glmnet R package to perform feature selection on the migrasome-related lncRNAs screened in Step 2, screen out 4 key lncRNAs and construct a prognostic risk model for pancreatic cancer. Obtain a scoring formula through the multivariate cox proportional hazards model to calculate the risk score of patients, and divide the patients into high-risk and low-risk groups according to the median of the scores;

[0089] Feature selection uses the Lasso regression algorithm;

[0090] The feature selection threshold for migrasome-related lncRNAs is the optimal range of Log lambda;

[0091] Step 4: Use the Kaplan-Meier survival curve and Log-rank test to evaluate the survival differences between the high- and low-risk groups, draw survival status plots, risk score distribution plots, and heatmaps to visually display the relationship between the expression of key lncRNAs and patient survival, and develop a nomogram based on the risk score and age and gender factors;

[0092] Step 5: Evaluate the relationship between the expression of migrasome-related lncRNAs and immune cell infiltration in the pancreatic cancer microenvironment, analyze the differences in immune cell composition and immune function scores between the high- and low-risk groups, and at the same time, through tumor mutational burden (TMB) and immune escape analysis, reveal the potential differences in immune therapy responsiveness between the high- and low-risk groups;

[0093] Step 6: Predict the sensitivity of high- and low-risk group patients to anti-cancer drugs through the oncoPredict package, combine the drug sensitivity data for analysis, and evaluate the drug response differences between the high- and low-risk groups.

[0094] Identification of migrasome-related lncRNAs

[0095] Through the co-expression analysis of migrasome-related genes and lncRNAs, lncRNAs that meet the criteria of |r|>0.3 and P<0.05 were screened out, and the relationship between these lncRNAs and migrasome-related genes was shown through a Sankey diagram ( Figure 2 -A). The results showed ( Figure 2 -B) that the migrasome-related lncRNAs identified by co-expression analysis were significantly different between the high- and low-risk groups, identifying potential key lncRNAs, which were further used for the construction of the prognostic model.

[0096] Construction and validation of the prognostic model

[0097] Feature selection: Lasso regression was used to screen lncRNAs related to migrasomes. By plotting the relationship between Log lambda and the partial likelihood deviance ( Figure 2 -D), the optimal penalty parameter (lambda) was determined, and the features were gradually shrunk. Eventually, a small number of lncRNAs with non-zero coefficients were retained. This step ensures that the model is concise and effectively avoids overfitting. Figure 2 -E shows the variation of the number of features with the penalty parameter. By gradually increasing the lambda value and observing the changes in the lncRNA regression coefficients, a suitable lambda value was selected, and this process was visualized (when λ = 4, the rationality of the model can be ensured, that is, 4 lncRNAs related to the prognosis of migrasomes were selected).

[0098] Establishment of the risk score model and grouping: Based on the lncRNAs screened by Lasso regression and their regression coefficients, a prognostic risk score model for patients was constructed. As Figure 3 shown, survival assessment of pancreatic cancer was performed in the total cohort, test set, and training set. According to the median risk score, patients were divided into high-risk and low-risk groups. Survival analysis showed that patients in the high-risk group had significantly worse prognoses, while patients in the low-risk group had higher survival rates, with significant differences. The performance of this risk score in the test set verified its stability and generalization ability.

[0099] Survival analysis and PFS verification

[0100] As Figure 4 shown, in order to further verify the predictive effect of the risk score model, the survival differences between high- and low-risk group patients were analyzed by Kaplan-Meier survival curves. There were significant survival differences between the high- and low-risk groups in the overall, training set, and test set. The results of the Log-rank test illustrated the significant survival differences between the high- and low-risk groups, indicating the effectiveness of the risk score model in different patient datasets. In addition, according to clinical features such as gender, age, tumor grade, and stage, the patients were divided into different subgroups in the present invention. The Kaplan-Meier curves of the subgroup analysis also showed significant survival differences between the high- and low-risk groups in each subgroup. In the further progression-free survival (PFS) analysis, the results also showed significant differences in PFS between the high- and low-risk groups. Patients in the high-risk group had lower PFS and faster disease progression, while patients in the low-risk group had a significantly prolonged PFS. This analysis further demonstrated the applicability of the risk score model in predicting disease progression.

[0101] Risk score distribution and independent prognostic ability

[0102] The risk score distribution map and the survival status map further reveal the relationship between the risk score and the patient's survival time. After arranging the risk scores in ascending order, it shows that the survival time of the high-risk group is shorter and the mortality rate is higher. The survival status map visually demonstrates the differences in the survival status of patients in the high- and low-risk groups. In addition, the heat map ( Figure 3 -A, B, C) shows the expression of the characteristic lncRNAs in the high- and low-risk groups, presenting obvious expression differences, further supporting the importance of these lncRNAs in prognostic prediction and verifying the prediction effect of the model.

[0103] To evaluate the independent prognostic role of the risk score, univariate and multivariate Cox proportional hazards regression analyses were performed. After controlling for other clinical variables, the risk score remained a significant independent prognostic factor. Through visual analysis using the forest plot ( Figure 2 -C), the risk score had a relatively high hazard ratio among all variables, intuitively verifying the reliability of the risk score.

[0104] Model prediction performance evaluation

[0105] The time-dependent ROC curve was used to evaluate the prediction effect of the model at 1 year, 3 years, and 5 years. The results showed ( Figure 5 -A, B) that the AUC values of the risk score model at 1 year, 3 years, and 5 years were at relatively high levels, indicating that the model had strong ability to distinguish the survival risks of patients at different time points. Especially, the AUC values at 1 year and 3 years were close to 0.8, indicating excellent performance of the model in short-term prognostic prediction, and the AUC at 5 years also reached a relatively high level, demonstrating the stability of the model in long-term prognostic prediction.

[0106] To further verify the prediction ability of the model, the C-index (concordance index) of the model was calculated. The C-index is used to measure the discrimination ability of the model, that is, the accuracy of the model in predicting the survival ranking of patients. The results showed ( Figure 5 -E) that the C-index value of the risk score model was significantly higher than other clinical characteristics (such as age, tumor stage, etc.), indicating that the risk score model had high accuracy and robustness in distinguishing patients with different prognoses. This result further supported the effectiveness of the risk score model as an independent prognostic indicator.

[0107] To comprehensively apply the risk score and other clinical characteristics to individualized survival prediction, a nomogram of the multivariate Cox model was constructed ( Figure 5 -G). The nomogram combines clinical variables such as the risk score, age, gender, and tumor stage, assigns corresponding risk scores to each variable, and thus generates predicted values of 1-year, 3-year, and 5-year survival probabilities for individual patients. To verify the prediction accuracy of the nomogram, calibration curves at 1 year, 3 years, and 5 years were plotted ( Figure 5-F), the consistency between the survival probabilities predicted by the model and the actual survival rates was compared. The calibration curve showed that the prediction results of the nomogram at each time point were highly consistent with the actually observed survival rates, indicating that the nomogram had good accuracy and consistency in predictions at different time points, reducing the errors caused by model bias. Through the calibration process of 1,000 times of Bootstrap sampling, the overfitting risk of the model was further reduced, demonstrating the reliability of the nomogram.

[0108] Principal component analysis (PCA)

[0109] To further verify the effectiveness of the risk score model in patient classification, principal component analysis (PCA) was used to visualize the differences between the high-risk group and the low-risk group in gene expression data. In the PCA analysis, we selected the first three principal components (PC1, PC2, and PC3) and projected the high- and low-risk groups into this three-dimensional space. The results showed ( Figure 6 -E, F, G) that the distributions of the low-risk group and the high-risk group in the PCA space were significantly separated. The samples of the low-risk group mostly concentrated in one area of the space, while the samples of the high-risk group were distributed on the other side. This spatial separation indicated that there were significant differences in the gene expression patterns of patients in different risk groups. It supported the classification ability of the risk score model in distinguishing patients with different survival prognoses.

[0110] Gene function enrichment analysis

[0111] GO and KEGG enrichment analyses were performed on the selected differentially expressed genes. The results of the GO analysis showed ( Figure 6 -A, B) that these genes were significantly enriched in multiple migration-related functions in biological processes (BP), cellular components (CC), and molecular functions (MF). In terms of biological processes, these genes were mainly enriched in processes such as neural signal transmission, trans-synaptic signal regulation, hormone secretion, and cell-cell communication, indicating that neural and endocrine signals might play important roles in the progression of pancreatic cancer. The cellular component analysis showed that the enriched genes were concentrated in neuron-related structures (such as neuron cell bodies, axon terminals) and vesicle membranes, suggesting that these structures might be involved in the material transport and signal transmission between cancer cells and support the growth and spread of tumors. The molecular function analysis revealed significant enrichment of these genes in ion channels (such as potassium ion channels) and enzyme activities (such as serine peptidase activity). The changes in such ion channels and enzyme activities might affect the signal conduction, membrane potential regulation, and tissue invasiveness of cancer cells.

[0112] KEGG pathway analysis further revealed that these genes were enriched in multiple signal pathways related to cancer progression. The KEGG pathway enrichment analysis showed ( Figure 6-C, D), the differentially expressed genes were significantly enriched in multiple pathways such as pancreatic metabolism, neural signal transduction, cell signal transduction, and endocrine regulation. These genes play important roles in metabolic and endocrine functions such as pancreatic secretion, protein digestion, and insulin secretion, and may affect the basic physiological functions of the pancreas. The enrichment of neural signaling pathways such as neuroactive ligand-receptor interaction, dopaminergic and GABAergic synapses suggests that these genes may be involved in signal transduction in cancer cells and shaping the tumor microenvironment. Meanwhile, the enrichment of key signaling pathways such as MAPK and cAMP indicates that these genes may affect tumor progression by regulating cell growth and differentiation. The enrichment of the endocrine regulation pathway implies that there are differences in gene expression between the high- and low-risk groups in terms of hormone synthesis and regulation, and these mechanisms jointly affect the malignant behavior of pancreatic cancer and the survival prognosis of patients.

[0113] In addition, single-sample gene set enrichment analysis (ssGSEA) further revealed ( Figure 6 -H, I, J), the differences in immune function between different risk groups. By calculating the enrichment scores of each sample on different immune function gene sets, it showed that the low-risk group was enriched in pathways such as calcium signaling pathway, complement and coagulation cascades, neuroactive ligand-receptor interaction, and steroid hormone biosynthesis, reflecting its activity in metabolism, immune regulation, and endocrine function, which may provide a more stable microenvironment for the low-risk group and inhibit tumor progression. On the contrary, the high-risk group was enriched in pathways related to epidermal cell differentiation, keratinization, and chromatin structure, indicating higher cell differentiation activity, increased genomic instability, and prone to form highly invasive cell populations, resulting in a worse prognosis. These findings further explain the differences in biological characteristics between the high- and low-risk groups and support the role of the risk score model in prognosis prediction.

[0114] Tumor Microenvironment (TME) and Immune Analysis

[0115] The tumor microenvironment scores were compared among different risk groups, as Figure 7 shown, using the ESTIMATE algorithm and the CIBERSORT algorithm to evaluate the tumor microenvironment (TME) score and immune cell infiltration respectively, and analyze the gene mutation characteristics. The ESTIMATE algorithm was used to evaluate the immune score, stromal score, and comprehensive score of patients in the high-risk and low-risk groups. The results showed that the TME score of patients in the high-risk group was significantly lower than that in the low-risk group, especially with significant differences in the immune score and stromal score. This result indicates that the tumor microenvironment of patients in the high-risk group may have lower immune and stromal component contents, indicating higher malignancy and more invasive tumor characteristics. The violin plot visually demonstrated the distribution differences of the TME score between different risk groups, and this significant difference was confirmed by the Wilcoxon rank-sum test.

[0116] Subsequently, the CIBERSORT algorithm was used to analyze the immune cell infiltration in each patient sample, including multiple immune cell components. The results showed that there were significant differences in the infiltration proportions of immune cells between the high-risk group and the low-risk group, especially in anti-tumor immune cell types such as T cells and natural killer (NK) cells, where the infiltration proportion in the low-risk group was significantly higher. This infiltration feature indicates that patients in the low-risk group have a more active anti-tumor immune response ability, which may contribute to their better survival prognosis.

[0117] Analysis of tumor mutations and immune escape ability

[0118] To explore the differences in tumor mutation characteristics between the high- and low-risk groups, the maftools was used to analyze the gene mutation data of the patients. The results Figure 8 showed that there were significant differences between the high-risk group and the low-risk group in some common mutated genes (such as TP53, KRAS, etc.). The mutation frequency in the high-risk group was higher, indicating that the genomes of these patients had stronger instability, which might lead to more malignant and rapidly developing tumor characteristics.

[0119] Subsequently, survival analysis was performed by combining TMB with the risk score, and the patients were further divided into high-TMB and low-TMB subgroups, forming four combinations: low-TMB + low-risk, low-TMB + high-risk, high-TMB + low-risk, and high-TMB + high-risk. The survival differences among the four groups of patients were analyzed by the Kaplan-Meier survival curve. The results showed that the survival time of patients in the high-TMB + high-risk group was significantly shorter, while the survival situation of the low-TMB + low-risk group was the best. This result indicates that combining TMB with the risk score can more effectively predict the survival prognosis of patients.

[0120] In addition, to evaluate the potential response of patients in different risk groups to immunotherapy, the TIDE (Tumor Immune Dysfunction and Exclusion) algorithm was used to analyze the immune escape ability of the patients. The TIDE score is mainly used to evaluate the ability of tumors to escape immune detection and anti-tumor immunity, so as to predict the response degree of patients to immunotherapy. The results showed that the TIDE score of the high-risk group was significantly higher than that of the low-risk group, indicating that patients in the high-risk group were more likely to have immune escape, thus reducing their response to immunotherapy. This finding proves that the tumors of patients in the high-risk group may escape the host's immune surveillance through more effective immune suppression mechanisms in the immune microenvironment.

[0121] Drug sensitivity analysis

[0122] By analyzing the drug sensitivities of patients in the high- and low-risk groups, as Figure 9As shown, the figure demonstrates the response differences of different drugs between the two groups. The oncoPredict package was used to predict the sensitivity of high- and low-risk groups to various anti-cancer drugs and visualize them in the form of box plots.

[0123] The results showed that there were significant differences in Alisertib and AMG-319 between the high- and low-risk groups (p < 0.05), with the high-risk group being more sensitive to these two drugs, suggesting that they may be more effective in high-risk patients. Similarly, significant sensitivity differences were also shown in BMS-345541, Carmustine, and CDK9_5038. In particular, the sensitivity of CDK9_5038 was significantly higher in the high-risk group, indicating its potential therapeutic effect on high-risk patients.

[0124] In addition, the sensitivities of Eg5_9814 and GSK1904529A were significantly higher in the high-risk group than in the low-risk group (p < 0.05), showing that they may have greater therapeutic advantages in high-risk patients. Significant differences were also shown in JAK1_8709 and KRAS(G12C) Inhibitor-12. In particular, KRAS(G12C) Inhibitor-12 showed stronger sensitivity in the high-risk group, indicating its greater therapeutic potential in high-risk patients. The sensitivity analysis of Mitoxantrone, MN-64, and OSI-027 also showed significant differences, suggesting that these drugs may have more advantages in the treatment of the high-risk group.

[0125] In addition, significant sensitivity differences were shown in PAK_5339, PCI-34051, and PF-4708671 in the high-risk group (p < 0.05). In particular, PCI-34051 showed higher sensitivity to high-risk patients. Finally, there were statistically significant differences in RO-3306 and Sinularin between the high- and low-risk groups (p < 0.05), with the high-risk group being more sensitive to these drugs, further indicating that these drugs may have greater application potential in high-risk patients.

[0126] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0127] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for constructing a prognostic model of pancreatic cancer based on migrasome-related lncRNA, characterized in that, The steps are as follows: Step 1: Obtain the transcriptome data and clinical information of pancreatic cancer patients from the Cancer Genome Atlas (TCGA) database; Step 2: Standardize and preprocess the data obtained in Step 1 to form a sample dataset, extract the expression levels of 8 key genes related to migrasomes, perform differential analysis and co-expression analysis on tumor samples and normal samples using the limma and edgeR R packages, screen out 180 lncRNAs related to migrasomes, and generate a Sankey diagram using the ggalluvial package to visualize the co-expression relationship; Step 3: Use the Lasso regression model in the glmnet R package to perform feature selection on the migrasome-related lncRNAs screened in Step 2, screen out 4 key lncRNAs and construct a prognostic risk model for pancreatic cancer, obtain a scoring formula through a multifactorial cox proportional hazards model to calculate the risk score of patients, and divide the patients into high-risk and low-risk groups according to the median of the scores; Step 4: Use the Kaplan-Meier survival curve and Log-rank test to evaluate the survival differences between the high- and low-risk groups, draw survival status diagrams, risk score distribution diagrams, and heatmaps to visually display the relationship between the expression of key lncRNAs and patient survival, and develop a nomogram based on the risk score and age and gender factors; Step 5: Evaluate the relationship between the expression of migrasome-related lncRNAs and immune cell infiltration in the pancreatic cancer microenvironment, analyze the differences in immune cell composition and immune function scores between the high- and low-risk groups, and at the same time, through tumor mutational burden (TMB) and immune escape analysis, reveal the potential differences in immune therapy responsiveness between the high- and low-risk groups; Step 6: Predict the sensitivity of high- and low-risk group patients to anti-cancer drugs through the oncoPredict package, combine the drug sensitivity data for analysis, and evaluate the drug response differences between the high- and low-risk groups.

2. The construction method of a pancreatic cancer prognosis model based on migrasome-related lncRNA according to claim 1, characterized in that: In Step 1, the transcriptome data and clinical information include, but are not limited to, mRNA expression profiles, lncRNA expression profiles, mutation information, and the age, gender, TNM stage, and survival data of patients.

3. The construction method of a pancreatic cancer prognosis model based on migrasome-related lncRNA according to claim 1, characterized in that: In Step 2, the dataset is divided into a 1:1 training set and test set for model construction and independent verification.

4. The construction method of a pancreatic cancer prognosis model based on migrasome-related lncRNA according to claim 1, characterized in that: In Step 3, the feature selection uses the Lasso regression algorithm.

5. The construction method of a pancreatic cancer prognosis model based on migrasome-related lncRNA according to claim 1, characterized in that: In Step 3, the feature selection threshold for migrasome-related lncRNAs is the optimal range of Log lambda.