Pancreatic cancer prognosis prediction model based on intercellular effect related genes, construction method and application
By constructing a pancreatic cancer prognosis prediction model based on genes related to cell burial effects, using genes such as ASPH, EMP1, FNDC3B and LGALS3, combined with machine learning methods, the dynamics and population heterogeneity of pancreatic cancer prognosis assessment are solved, and the possibility of early warning and individualized treatment is achieved.
Patent Information
- Application Number
- CN202510617651.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-09-02
AI Technical Summary
The prognostic assessment tools for pancreatic cancer patients in the prior art lack dynamic and population heterogeneity analysis, and insufficient prediction methods for genes related to cell burial effects, which makes it difficult to achieve early warning and individualized treatment strategies.
A pancreatic cancer prognosis prediction model based on genes related to cell burial action was constructed. By screening genes such as ASPH, EMP1, FNDC3B and LGALS3, using TCGA database and machine learning methods, a risk scoring model was established to divide high-risk and low-risk groups.
It provides a more accurate prediction tool for pancreatic cancer prognosis, which can identify high-risk patients early, provide reference for individualized treatment plans, and improve clinical management effectiveness.
Smart Images

Figure CN120581071A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biomedicine, and in particular to a pancreatic cancer prognosis prediction model based on efferocytosis-related genes using machine learning technology, as well as a construction method and application thereof. Background Art
[0002] Pancreatic ductal adenocarcinoma (PDAC) is the most common and aggressive histological subtype of pancreatic cancer and one of the leading causes of cancer-related deaths worldwide, with a five-year survival rate of less than 10%. Due to the insidious nature of the disease and nonspecific symptoms, PDAC patients are often diagnosed at an advanced or metastatic stage, significantly limiting the potential for treatment benefit. Therefore, the discovery of innovative biomarkers and molecular profiles is urgently needed to enable earlier disease warning, more precise prognostic stratification, and the development of personalized treatment strategies, thereby improving clinical management and curbing the impact of the highly malignant phenotype.
[0003] Efferocytosis (ER), as a mechanism for the phagocytic clearance of apoptotic cells, maintains the stability of the tumor microenvironment by preventing secondary necrosis, which can trigger an inflammatory response and exacerbate tissue damage. Although evidence suggests that ER is particularly important in the context of tumor biology and plays a key role in tumor development and progression by reshaping the tumor microenvironment, there are still problems such as insufficient analysis of population heterogeneity, a lack of dynamic prognostic assessment tools, and limitations in efferocytosis-based PDAC prognostic gene screening methods. Therefore, current research on efferocytosis to predict the prognosis of pancreatic cancer patients is still insufficient. Summary of the Invention
[0004] To address the above-mentioned technical issues, the present invention provides a pancreatic cancer prognosis prediction model based on genes associated with efferocytosis, as well as a method for its construction and application. This application not only addresses the existing issues of insufficient analysis of population heterogeneity and the lack of dynamic prognostic assessment tools, but also provides new theoretical basis and technical support for pancreatic cancer prognosis prediction, with significant clinical application value.
[0005] In a first aspect, the present invention provides a use of a molecular marker related to efferocytosis, which is achieved through the following technical solutions.
[0006] An application of an efferocytosis-related molecular marker in the preparation of a product for predicting the prognosis of pancreatic cancer, wherein the efferocytosis-related molecular marker consists of the following genes: ASPH, EMP1, FNDC3B and LGALS3.
[0007] In a second aspect, the present invention provides a pancreatic cancer prognosis prediction model based on efferocytosis-related genes, which is achieved through the following technical solutions.
[0008] A pancreatic cancer prognosis prediction model based on efferocytosis-related genes was constructed based on the above markers, and the calculation formula was: Risk Score = ASPH*(0.0360)+EMP1*(0.1314)+FNDC3B*(0.0372)+
[0009] LGALS3*(0.2998), with the quartile of risk score as the critical value, was divided into two prognostic subgroups: high-risk group and low-risk group.
[0010] In a third aspect, the present invention provides a method for constructing a pancreatic cancer prognosis prediction model based on genes related to efferocytosis, which is achieved through the following technical solutions.
[0011] A method for constructing the above model includes the following steps:
[0012] Genes related to efferocytosis were obtained by screening the molecular feature database;
[0013] Obtain transcriptome and clinical data of pancreatic cancer in the TCGA database;
[0014] Using overall survival time and survival status as outcome variables and efferocytosis-related gene expression values as predictor variables, univariate Cox proportional hazards regression analysis was performed on efferocytosis-related genes. Using p < 0.05 as the significance threshold, genes with no statistically significant difference were eliminated to screen out efferocytosis-related genes that were significantly associated with pancreatic cancer prognosis.
[0015] Using behavioral genes, gene expression matrices with columns as samples, and patient survival data including survival time and survival status as input data, multiple iterations of LASSO regularized Cox proportional hazard model regression analysis were performed;
[0016] The frequency of each gene selected in the iterative process of the LASSO regularized Cox proportional hazard model regression iteration results was counted, and genes with a frequency threshold > 50% were identified as high-frequency stable genes;
[0017] The expression levels of high-frequency stable genes and the average values of their regression coefficients in multiple iterations were weighted and summed to obtain a pancreatic cancer prognosis prediction model based on efferocytosis-related genes.
[0018] Furthermore, the specific method for obtaining efferocytosis-related genes through screening of the molecular feature database is as follows: all efferocytosis-related genes are extracted from the molecular feature database, genes related to the overall survival time and disease-free survival of pancreatic cancer patients are screened out, and representative survival-negative-related genes are screened out. After that, genes closely related to the expression of representative survival-negative-related genes are screened out in pancreatic cancer tissues, and genes expressed in normal tissues are eliminated, and finally efferocytosis-related genes are obtained.
[0019] Furthermore, the specific method for obtaining the transcriptome and clinical data of pancreatic cancer in the TCGA database is as follows: obtaining the mRNA TPM expression profile and supporting clinical information data, performing log2 normalization conversion on the original TPM expression matrix, and retaining only samples with sample numbers ending in "01A".
[0020] Furthermore, LASSO regularized Cox proportional hazards model regression analysis was performed for 100 iterations with the following parameters: alpha = 1, family = "cox", type.measure = "deviance", nfolds = 10, and standardize = TRUE.
[0021] In a fourth aspect, the present invention provides a computer-readable storage medium, which is implemented through the following technical solutions.
[0022] A computer-readable storage medium stores a computer program, which controls the device where the computer-readable storage medium is located to execute the above-mentioned model when the computer program is running.
[0023] In a fifth aspect, the present invention provides a method for screening high-frequency stable genes of efferocytosis that are significantly correlated with the prognosis of pancreatic cancer, which is achieved through the following technical solutions.
[0024] A method for screening high-frequency stable genes of efferocytosis significantly associated with the prognosis of pancreatic cancer comprises the following steps:
[0025] Genes related to efferocytosis were obtained by screening the molecular feature database;
[0026] Obtain transcriptome and clinical data of pancreatic cancer in the TCGA database;
[0027] Using overall survival time and survival status as outcome variables and efferocytosis-related gene expression values as predictor variables, univariate Cox proportional hazards regression analysis was performed on efferocytosis-related genes. Using p < 0.05 as the significance threshold, genes with no statistically significant difference were eliminated to screen out efferocytosis-related genes that were significantly associated with pancreatic cancer prognosis.
[0028] Using behavioral genes, gene expression matrices with columns as samples, and patient survival data including survival time and survival status as input data, multiple iterations of LASSO regularized Cox proportional hazard model regression analysis were performed;
[0029] The frequency of each gene selected in the iterative process of the LASSO regularized Cox proportional hazard model regression iteration results was counted, and genes with a frequency threshold >50% were screened to obtain high-frequency stable genes.
[0030] This application has the following beneficial effects.
[0031] This application is based on the perspective of multidimensional interactions between tumor cells, matrix components and tumor-associated immune cells. By integrating the TCGA (Cancer Genome Atlas), GTEx (Genotype-Tissue Expression) database and multiple GEO (Gene Expression Omnibus) data sets, cell burial-related gene expression profile analysis was performed in pancreatic cancer patient tissues. Secondly, the expression profile analysis results were used to screen out a group of cell burial-related genes that may be related to the survival of pancreatic cancer patients as potential pancreatic cancer prognostic genes. Finally, machine deep learning was used to obtain 4 core genes through the LASSO algorithm, and the ROC curve was used to confirm that the combination of the 4 core genes has good predictive efficacy. The above-mentioned gene combination can not only predict the prognosis of PDAC patients, but also provide a reference for exploring new treatment options. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0033] Figure 1 The present invention is a screening of efferocytosis-related genes (wherein, A: significant prognostic value of ERGs for OS and DFS in pancreatic ductal adenocarcinoma (PDAC); BC: box plots showing upregulated genes (B) and downregulated genes (C) in OS and DFS, respectively; D: box plot of RNA expression of 5 genes associated with poor survival in PDAC tumor tissues and normal tissues; E: KM survival curve showing that these 5 genes are associated with poor survival in PDAC patients; F: top 11 genes with the highest correlation with these 5 genes in PDAC based on PCC analysis; G: Venn diagram showing the distribution of the above 11 genes in PDAC tumors and normal tissues; H: univariate Cox proportional hazards regression analysis results of 15 efferocytosis-related genes;
[0034] Figure 2The present invention develops and validates prognostic indicators associated with four high-frequency stable genes (A: Coefficient selection in LASSO regression analysis based on four high-frequency stable genes, with vertical lines indicating positions determined by optimal λ values; B: Cross-validation process for parameter adjustment by LASSO regression; C: Feature selection frequencies of 14 pre-selected genes screened by 100 deep learning steps in the TCGA-PDAC cohort; D: Survival curves of high-risk and low-risk groups generated by the Kaplan-Meier method; E: Time-dependent ROC curve analysis of the TCGA-PDAC cohort after deep learning; F: Validation of the four high-frequency stable genes in the database;
[0035] Figure 3 It is a flow chart of the method of the present invention. DETAILED DESCRIPTION
[0036] Hereinafter, preferred embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. Prior to the description, it should be understood that the terms used in the specification and the appended claims are not to be construed as limited to their general and dictionary meanings, but rather should be interpreted based on the meanings and concepts corresponding to the technical aspects of the present invention, based on the principle that allows the inventor to appropriately define the terms for the best interpretation. Therefore, the description herein is merely a preferred example for illustrative purposes and is not intended to limit the scope of the present invention. It should be understood that other equivalent implementations and modifications may be made without departing from the spirit and scope of the present invention.
[0037] like Figure 3 As shown, the present invention systematically constructs and verifies a pancreatic cancer (PDAC) prognosis prediction model based on efferocytosis gene characteristics by integrating multi-omics data and machine learning methods. First, efferocytosis-related genes are screened through a molecular feature database (such as MsigDB), and the transcriptome and clinical data of the TCGA-PAAD cohort are used, combined with univariate Cox regression (p<0.05), LASSO regression and multivariate Cox analysis to screen out genes significantly associated with prognosis. Through multiple iterative modeling, high-frequency stable genes and their corresponding regression coefficient means are selected to construct an efferocytosis prognostic risk scoring model, and patients are divided into high / low risk groups according to the quartiles of all patient scores. The prognostic prediction efficacy is verified by survival analysis and time-dependent ROC curves. The present invention provides new related molecular markers and calculation methods based on efferocytosis for PDAC prognosis prediction.
[0038] The invention is further described below with reference to the accompanying drawings and examples. Unless otherwise specified, the experimental methods used in the present invention are conventional methods, and the experimental equipment, materials, reagents, etc. used can be purchased from relevant material sales companies.
[0039] This application used R version 4.3.3, downloaded from https: / / cran.r-project.org / bin / windows / base / old / 4.3.3 / . The easyTCGA package was downloaded from https: / / github.com / , and the survival, glmnet, survmine, timeROC, and ggplot2 packages were downloaded from https: / / cran.r-project.org / .
[0040] Data sources for this invention include: (1) 178 PDAC patient samples were obtained from the TCGA database (phs000178), covering clinical information, transcriptome expression data, and copy number variation (CNV) data; (2) 171 normal pancreatic tissue samples were included as controls using the GTEx project on the GEPIA / GEPIA2 platform (PUCH); (3) the GEPIA / GEPIA2 platform (PUCH) provides pan-cancer cohort gene expression profiles and cancer-normal tissue interaction analysis functions; (4) six PDAC datasets (GSE111672, 141017, 148673, 154778, 162708, and 165399) and one CRA database (CRA001160) were screened from the GEO database for confirmatory analysis.
[0041] 1. The screening process for efferocytosis-related genes (ER) involved in the present invention is as follows: 167 efferocytosis-related genes (ERGs) were extracted from the MSigDB database (USCD). The 167 efferocytosis-related genes were from the following literature, the link is as follows: https: / / cancerci.biomedcentral.com / articles / 10.1186 / s12935-024-03571-3, and the 167 efferocytosis-related genes are shown in Table 1.
[0042] Table 1. 167 genes related to efferocytosis
[0043]
[0044] Among the 167 efferocytosis-related genes, 39 genes were associated with the overall survival (OS) and disease-free survival (DFS) of pancreatic cancer patients, including 21 OS-related genes and 18 DFS-related genes (see Figure 1 A, Figure 1 In A, red boxes represent negatively correlated genes, and blue boxes represent positively correlated genes). Five representative genes negatively correlated with survival were screened from 39 genes ( Figure 1 B), and further screened out 11 genes closely related to the expression of the above five genes in pancreatic cancer tissues ( Figure 1 F) Among the above 16 genes, 1 gene expressed in normal tissues was eliminated, resulting in 15 genes related to efferocytosis.
[0045] 2. The above-mentioned downloaded TCGA-PAAD and clinical data and CNV data sources: Based on R version 4.3.3 (released on February 29, 2024), the R language software package easyTCGA (version number 0.0.4.2000) was used to obtain the mRNA TPM expression spectrum and supporting clinical information data (download date: January 1, 2025). Through the bioinformatics preprocessing process, the original TPM expression matrix was first log2 normalized (log2 (TPM+1)) to eliminate the skewed distribution of the data and improve the reliability of the analysis. To ensure the homogeneity of the research subjects, the present invention only retains samples with sample numbers ending in "01A". These samples are all primary tumor samples, which effectively excludes the interference of metastatic tumors, blood samples and normal tissue samples, and the follow-up time (time) is less than or equal to 2000 days. Survival events were defined as status (1 = Dead, 0 = Alive). For patients who had no survival events (status = 0), only the longest follow-up time within 2000 days was retained. For patients who had an event (status = 1), only the last event was retained. Ultimately, 144 qualified samples were screened from 178 PDAC patient samples for the next analysis.
[0046] 3. The prognostic analysis method of efferocytosis-related genes of the present invention specifically includes the following steps: First, based on the TCGA-PAAD dataset, with 2000 days as the observation endpoint, the overall survival time (OS) and survival status (1=Dead, 0=Alive) of the patient were used as outcome variables, and the expression values of 15 efferocytosis-related genes were used as predictor variables. The 15 efferocytosis-related genes were subjected to univariate Cox proportional hazard regression analysis using the R language survival software package (version 3.5-8), and the hazard ratio (HR), regression coefficient (Coef), statistical significance level (p value) and 95% confidence interval (95% CI) of each gene were calculated. Using p < 0.05 as the significance threshold, genes with no statistical difference (RTN4) were eliminated, and finally 14 efferocytosis-related genes significantly associated with the prognosis of pancreatic cancer were screened out ( Figure 1 H and Table 2).
[0047] Table 2. 15 efferocytosis-related genes associated with pancreatic cancer prognosis
[0048]
[0049]
[0050] 4. The present invention provides a method for constructing a prognostic model of epiblast-related genes based on LASSO regularized Cox proportional hazard model regression analysis. The specific implementation steps are as follows: using the cv.glmnet function in the R language glmnet software package (version 4.1-8), taking the gene expression matrix (behavioral genes, columns as samples) and patient survival data (including survival time and survival status) as input data, setting the random number seed to set.seed (123) to ensure the reproducibility of the results. The LASSO regularized Cox proportional hazard model regression analysis was repeated 100 times, and each analysis used a 10-fold cross-validation technique. The key parameters were set as follows: alpha = 1 (pure LASSO penalty), family = "cox" (Cox proportional hazard model), type.measure = "deviance" (based on partial likelihood deviation evaluation criteria), nfolds = 10 (10-fold cross-validation), and standardize = TRUE (expression value standardization). Each iteration determines the optimal penalty parameter lambda.min through cross-validation, and records the regression coefficient of each gene (the coefficient of the unselected gene is recorded as 0), as shown in Table 3, Figure 2 As shown in A and 2B.
[0051]
[0052] 5. The results of 100 LASSO regularized Cox proportional hazard model regression iterations were integrated and analyzed, and the frequency of each gene selected during the iteration was counted to form a gene selection frequency table (Table 4). Using strict screening criteria, genes with a frequency threshold of >50% (threshold = 0.5) were set as high-frequency stable genes. Four high-frequency stable genes ("LEAF") were screened out from the 14 candidate genes ( Figure 2 C). The average regression coefficients of the four high-frequency stable genes across 100 iterations (accurate to four decimal places) were further calculated to form a gene coefficient table (Table 5). This screening method ensures the stability of the results through multiple iterations. The four genes obtained will be used to construct a prognostic risk score model for pancreatic cancer.
[0053] Table 4. Selection frequency of 14 efferocytosis-related genes in 100 iterations based on LASSO regularized Cox proportional hazards model regression analysis
[0054]
[0055] Table 5. Average regression coefficients of four high-frequency stable genes related to efferocytosis in 100 iterations based on LASSO regularized Cox proportional hazards model regression analysis
[0056]
[0057] 6. The expression levels of high-frequency stable genes and their average coefficients were weighted and summed to obtain the risk score: Risk Score = ASPH*(0.0360)+EMP1*(0.1314)+FNDC3B*(0.0372)+LGALS3*(0.2998).
[0058] 7. The Risk Score calculation formula was applied to quantify the risk score of each patient in the TCGA-PAAD cohort. The quartiles of the risk score of all patients were used as critical values (critical value 1 (lowest 25%) = 3.83, critical value 2 (highest 25%) = 4.18), and the patients were clearly divided into two prognostic subgroups: high-risk group ("High") and low-risk group ("Low"). The Kaplan-Meier survival model was constructed using the survfit function of the survival package (version 3.5-8) in the R language. The input parameters included: 1) survival time variable time; 2) survival status variable status (1 = Dead, 0 = Alive); 3) risk group variable (High / Low). The ggsurvplot function of the survminer package (version 0.5.0) was used to generate a survival curve with statistical significance marks to intuitively display the survival differences between high-risk and low-risk groups ( Figure 2 D).
[0059] 8. The predictive performance of the Risk Score was evaluated based on time-dependent receiver operating characteristic (ROC) analysis. Clinical follow-up time was first converted to standard days (1 year = 365 days, 3 years = 1095 days, and 5 years = 1825 days). Modeling and analysis were performed using the timeROC function in the R language timeROC package (version 0.4). Input parameters included: 1) survival time variable T (number of days); 2) survival status variable delta (1 = dead, 0 = alive); 3) risk score marker = Risk Score; 4) event type parameter cause = 1 (dead event); 5) weighting method weighting = "marginal" (marginal weight); and 6) iid = TRUE (calculation of confidence intervals). Statistical power was ensured by systematically checking the number of events before each time point. The false positive rate (FPR), true positive rate (TPR), and area under the curve (AUC) values (including 95% confidence intervals) were extracted to construct the evaluation data framework. Finally, the ggplot2 package (version 3.5.1) was used for visualization, including: 1) using geom_smooth (method = "loess") to draw a smooth ROC curve; 2) differentiating the color according to the predicted time point (1 year = red, 3 years = green, 5 years = blue); 3) adding a gray dotted reference line (AUC = 0.5); 4) accurately marking the AUC value of each time point in the lower right corner of the chart (keeping two decimal places) ( Figure 2 E).
[0060] 9. To expand the applicability of the results, this application retrieved 6 GSE and one CRA single-cell sequencing datasets related to pancreatic cancer patients from the GEO database, and verified the four high-frequency stable genes screened out one by one. The enrichment correlation of each gene in the single-cell dataset with the tumor cell, fibroblast, and immune cell populations was calculated respectively. Finally, the degree of association of each gene with the malignant tumor cell population and tumor-related stromal cell population in the single-cell sequencing was evaluated. The results are as follows: Figure 2 As shown in F.
[0061] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. Use of a molecular marker related to efferocytosis in the preparation of a product for predicting the prognosis of pancreatic cancer, characterized in that: Evoked cell function-related molecular markers consist of the following genes: ASPH, EMP1, FNDC3B and LGALS3.
2. A pancreatic cancer prognosis prediction model based on efferocytosis-related genes, characterized by: The model is constructed based on the markers described in claim 1, and the calculation formula is: Risk Score = ASPH*(0.0360)+EMP1*(0.1314)+FNDC3B*(0.0372)+LGALS3*(0.2998). The quartile of the risk score is used as the critical value to divide the patients into two prognostic subgroups: high-risk group and low-risk group.
3. A method for constructing the model according to claim 2, characterized in that: The following steps are involved: Genes related to efferocytosis were obtained by screening the molecular feature database; Obtain transcriptome and clinical data of pancreatic cancer in the TCGA database; Using overall survival time and survival status as outcome variables and efferocytosis-related gene expression values as predictor variables, univariate Cox proportional hazards regression analysis was performed on efferocytosis-related genes. Using p < 0.05 as the significance threshold, genes with no statistically significant difference were eliminated to screen out efferocytosis-related genes that were significantly associated with pancreatic cancer prognosis. Using behavioral genes, gene expression matrices with columns as samples, and patient survival data including survival time and survival status as input data, multiple iterations of LASSO regularized Cox proportional hazard model regression analysis were performed; The frequency of each gene selected in the iterative process of the LASSO regularized Cox proportional hazard model regression iteration results was counted, and genes with a frequency threshold >50% were identified as high-frequency stable genes; The expression levels of high-frequency stable genes and the average values of their regression coefficients in multiple iterations were weighted and summed to obtain a pancreatic cancer prognosis prediction model based on efferocytosis-related genes.
4. The construction method according to claim 3, characterized in that: The specific method for obtaining efferocytosis-related genes through screening of a molecular feature database is as follows: extract all efferocytosis-related genes from the molecular feature database, screen out genes related to the overall survival time and disease-free survival of pancreatic cancer patients, and then screen out representative survival-negative-related genes. Subsequently, genes closely related to the expression of representative survival-negative-related genes are screened in pancreatic cancer tissues, and genes expressed in normal tissues are eliminated, ultimately obtaining efferocytosis-related genes.
5. The construction method according to claim 3, characterized in that: The specific method for obtaining the transcriptome and clinical data of pancreatic cancer in the TCGA database is as follows: obtain the mRNA TPM expression profile and supporting clinical information data, perform log2 normalization transformation on the original TPM expression matrix, and retain only samples with sample numbers ending in "01A".
6. The construction method according to claim 3, characterized in that: The LASSO regularized Cox proportional hazards model regression analysis was performed for 100 iterations with the following parameters: alpha = 1, family = "cox", type.measure = "deviance", nfolds = 10, and standardize = TRUE.
7. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the model described in claim 2.
8. A method for screening high-frequency stable genes involved in efferocytosis that are significantly associated with the prognosis of pancreatic cancer, characterized by: The following steps are involved: Genes related to efferocytosis were obtained by screening the molecular feature database; Obtain transcriptome and clinical data of pancreatic cancer in the TCGA database; Using overall survival time and survival status as outcome variables and efferocytosis-related gene expression values as predictor variables, univariate Cox proportional hazards regression analysis was performed on efferocytosis-related genes. Using p < 0.05 as the significance threshold, genes with no statistically significant difference were eliminated to screen out efferocytosis-related genes that were significantly associated with pancreatic cancer prognosis. Using behavioral genes, gene expression matrices with columns as samples, and patient survival data including survival time and survival status as input data, multiple iterations of LASSO regularized Cox proportional hazard model regression analysis were performed; The frequency of each gene selected in the iterative process of the LASSO regularized Cox proportional hazard model regression iteration results was counted, and genes with a frequency threshold >50% were screened to obtain high-frequency stable genes.