Pancreatic cancer dependency gene scoring prognosis model construction method based on multi-algorithm optimization
By integrating multiple machine learning algorithms and pancreatic cancer-dependent gene screening, a high-precision prognosis model is constructed, which solves the problem of difficult to accurately predict the immunotherapy response and prognosis of pancreatic cancer patients, and achieves personalized treatment and survival prolongation.
Patent Information
- Application Number
- CN202510435584.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-25
AI Technical Summary
The prior art is difficult to accurately predict the response and prognosis of pancreatic cancer patients to immunotherapy, resulting in poor treatment effect and low patient survival rate.
By integrating multiple machine learning algorithms, combining pancreatic cancer-dependent gene screening and personalized immunotherapy evaluation, a high-precision and stable performance prognosis model is built, biomarkers related to pancreatic cancer prognosis are identified, and personalized treatment plans are formulated.
Accurate prediction of the risk and prognosis of immunotherapy in pancreatic cancer patients has been achieved, the treatment effect has been improved, the patient's survival has been extended, new therapeutic targets have been discovered, and the development of pancreatic cancer treatment has been promoted.
Smart Images

Figure CN120375918A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biomedicine, in particular to the innovation of prognostic evaluation and personalized immunotherapy for pancreatic cancer. By integrating multiple machine learning algorithms, a method for constructing a prognostic model based on multi-algorithm optimization of pancreatic cancer-dependent gene scoring is constructed to accurately predict the risk and prognosis of patients receiving immunotherapy, thereby providing a scientific basis for clinical decision-making. Background Art
[0002] Pancreatic cancer is a highly malignant digestive system tumor, ranking fourth in malignant tumor-related mortality, with a 5-year survival rate of no more than 7%. Due to the insidious onset of pancreatic cancer and the lack of obvious early symptoms, most patients are already in the middle or late stages when diagnosed, missing the best time for treatment. In addition, pancreatic cancer is highly invasive and metastatic, and responds poorly to traditional treatments such as surgery, radiotherapy, and chemotherapy, resulting in a generally poor prognosis for patients, with a 5-year survival rate of 1% to 5% after surgery, posing a serious threat to human life and health.
[0003] Tumor dependency genes refer to the use of RNAi, CRISPR and other technologies to screen hundreds of cancer cell lines in heterogeneous cancer cell models on a large scale to identify adaptive genes required for tumor cell survival. The existence of tumor dependency genes makes the treatment of cancers such as pancreatic cancer extremely complex and challenging. The "cold tumor" and highly heterogeneous characteristics of pancreatic cancer indicate that there are likely to be a large number of pancreatic cancer dependency gene biomarkers in pancreatic cancer that are related to clinical diagnosis, survival and prognosis, but there are currently no reports on the study of pancreatic cancer dependency gene maps.
[0004] With the rapid development of biotechnology and immunology, immunotherapy, as an emerging treatment method, has shown significant efficacy and potential in the treatment of various tumors. Immunotherapy activates the patient's own immune system to identify and attack tumor cells, with high specificity and durability. However, although immunotherapy has achieved remarkable results in some tumors, its application effect in pancreatic cancer is not ideal.
[0005] On the one hand, the immune microenvironment of pancreatic cancer is complex and changeable, and tumor cells evade the surveillance and attack of the immune system through various mechanisms, which limits the effectiveness of immunotherapy. On the other hand, there are significant individual differences between different patients, including gene expression characteristics, immune status, clinical characteristics, etc. These factors will affect the efficacy and prognosis of immunotherapy. Therefore, how to accurately predict the response and prognosis of pancreatic cancer patients to immunotherapy has become a key issue that needs to be solved urgently.
[0006] At present, although some studies have attempted to predict the prognosis and immunotherapy response of pancreatic cancer through means such as gene expression profiles and immune-related markers. For example, Chen Liang et al. used DEGs at the transcriptome level to screen highly expressed genes in pancreatic cancer and construct a prognostic model. The AUC under the ROC curve of the model maintained around 0.75 in the test dataset, and the AUC under the ROC curve of the model in the external validation dataset was between 0.7 and 0.85. The prognostic models constructed by this method have problems such as low accuracy, poor stability, and insufficient specificity, and it is difficult to meet the needs of clinical decision-making. Therefore, developing a new method that can accurately predict the immunotherapy risk and prognosis of pancreatic cancer patients is of great significance for improving the treatment effect and prolonging the patient's survival period. Summary of the Invention
[0007] The object of the present invention is to provide a method for constructing a prognostic model for optimizing pancreatic cancer dependence gene scores based on multiple algorithms, which can use machine learning algorithms to identify biomarkers related to the prognosis and survival of pancreatic cancer, construct a prognostic model with high precision and stable performance, and combine the individual differences of patients, including gene expression characteristics, clinical characteristics, etc., to realize the formulation of personalized immunotherapy plans and risk assessment.
[0008] By integrating a variety of advanced machine learning algorithms, combined with the screening of pancreatic cancer dependence genes and personalized immunotherapy evaluation, it aims to construct a prognostic model with high precision and stable performance, provide a scientific basis for the clinical decision-making of pancreatic cancer patients, and thus improve the treatment effect and the patient's survival period.
[0009] The specific steps are as follows:
[0010] Step 1: Identification of pancreatic cancer malignant cells and pancreatic cancer dependence gene activity;
[0011] Pancreatic cancer cells originate from ductal cells. By analyzing the copy number variations (CNVs) of ductal cells in the gene expression data of the single-cell sequencing dataset (GSE212966), all malignant cells and non-malignant cells are distinguished; re-clustering analysis is performed on ductal cells, and the clusters derived from normal populations are selected as reference cells. Unsupervised clustering technology is used to distinguish high-CNVs and low-CNVs cells, and the clusters with significantly higher CNV scores than other cells are defined as malignant ductal cells;
[0012] Subsequently, the activity scores of pancreatic cancer dependence genes (PADGs) of each malignant ductal cell are calculated, and the natural value is taken as the threshold to distinguish the malignant cells into two cell populations: PADG highly active malignant cells and PADG low active malignant cells;
[0013] Step 2: Consensus clustering to identify PADG subgroups;
[0014] Download CRISPR gene effect data from 22 pancreatic cancer cell lines and 295 cell lines of other cancer types on the DepMap portal website; define genes with gene effects in pancreatic cancer cell lines < those in other cancer cell lines as pancreatic cancer-dependent genes; determine different subtypes of pancreatic ductal adenocarcinoma by unsupervised clustering analysis based on the expression of PADG genes, and analyze the prognostic differences among these subtypes through survival analysis;
[0015] Step 3: Construction of the prognostic model;
[0016] Combine the differentially expressed genes among PADG subtypes, the differentially expressed genes between tumors and control groups, and the differential genes between highly active malignant PADG cells and all other cells. The intersection of the three groups of differential genes was summarized to obtain a common differential gene; further, through univariate COX regression analysis, genes significantly related to prognosis were screened out;
[0017] To construct an efficient PADG-related prognostic model, ten advanced machine learning algorithms were adopted and tested on the GSE183795 cohort and two external validation datasets, including GSE62452 and GSE28735, to select the model with the highest average C-index value in the validation cohort; based on the RSF+StepCox[forward] algorithm, the most valuable PADG-related feature genes were identified, and the prognostic model with the best performance was constructed;
[0018] Step 4: Analysis of the correlation of prognostic genes;
[0019] Analyze the impact of the expression patterns of the above prognostic genes on the survival and prognosis of PDAC patients, and compare the expression differences of these genes among PADG subtypes; define genes that are obtained by the above methods and have significant differences as PADG biomarker genes;
[0020] Step 5: Evaluation of immunotherapy for the machine learning prognostic model based on the pancreatic cancer-dependent gene score;
[0021] Perform a correlation analysis on the obtained PADG biomarker genes and 68 existing immune checkpoints to evaluate their immune response potential and immune activity; when evaluating the immunotherapy reactivity, evaluate the reactivity of patients to immunotherapy based on the risk stratification of biomarker gene expression by calculating the between-group differences in the dysfunction score, exclusion score, and comprehensive TIDE score of PADG subtypes.
[0022] In the first step, consistency clustering was used to identify the PADG subgroups. Re-clustering analysis was performed on ductal cells, resulting in a total of eight clusters: cluster0, 1, 2, 3, 4, 5, 6, and 7. Cluster 5, which originated from normal individuals, was selected as the reference cell. Unsupervised clustering technology was used to distinguish between high-CNVs and low-CNVs cells. It was found that the CNV scores of Cluster 0, 1, 2, 3, 4, and 6 were significantly higher than those of other clusters. Cluster 0, 1, 2, 3, 4, and 6 were defined as malignant ductal cells. Subsequently, the pancreatic cancer dependency gene (PADG) activity scores of each malignant ductal cell were calculated, and using a threshold of 0.15, the malignant cells were divided into two cell populations: PADG-highly active malignant cells and PADG-lowly active malignant cells. By comparing the PADG-highly active malignant cells and PADG-lowly active malignant cells, a total of 442 differentially expressed genes (DEGs) were identified. The differences in these genes between the two groups were statistically significant, with an adjusted p < 0.05 and |Log2 fold change| > 0.5.
[0023] In the second step, consistency clustering was used to identify the PADG subgroups. Based on the expression of the PADG gene, two different pancreatic ductal adenocarcinoma subtypes were determined through unsupervised clustering analysis. The optimal clustering with k = 2 indicated that the data were reliable and stably differentiated into two PADG activity clusters with different activities. Moreover, Kaplan-Meier survival analysis showed significant prognostic differences between these subtypes. Specifically, the survival outcomes of patients in PADG Cluster1 were significantly better than those of patients in Cluster 2. By comparing different PADG subtypes, a total of 634 differentially expressed genes (DEGs) were identified. The differences in these genes between the two groups were statistically significant, with |log2Fold Change| > 1 and an FDR-adjusted P value < 0.05.
[0024] In Step 3, for the construction of the prognostic model, by comparing tumors and controls in the transcriptome data of GSE28735, GSE183795, and GSE62452, a total of 471 differentially expressed genes (DEGs) were identified. The differences in these genes between the two groups were statistically significant, with |log2Fold Change| > 1 and p < 0.05 after FDR correction. All DEGs were visually displayed through volcano plots. By comparing highly active malignant cells of PADG with all other cells, a total of 857 differentially expressed genes (DEGs) were identified. The differences in these genes between the two groups were statistically significant, with corrected p < 0.05 and |Log2 foldchange| > 0.5. By combining the differentially expressed genes between PADG subtypes, the differentially expressed genes between tumors and control groups, and the differential genes between highly active malignant cells of PADG and all other cells, the intersection of the three groups of differential genes was summarized to obtain a list containing 39 genes. Further, through univariate COX regression analysis, 22 genes significantly related to prognosis were screened out, with pvalue < 0.05.
[0025] To construct an efficient PADG-related prognostic model, ten advanced machine learning algorithms were adopted and tested on the GSE183795 cohort and two external validation datasets to select the model with the highest average C-index value in the validation cohort. Finally, based on the RSF+StepCox[forward] algorithm, 18 most valuable PADG-related characteristic genes were identified, and the best-performing prognostic model was constructed.
[0026] Survival analysis was performed on the GSE62452, GSE183795, GSE28735, and TCGA datasets. The results showed that a high risk score was associated with a reduction in survival time. The AUC values of the GSE183795 dataset at 1 year, 3 years, and 5 years were 0.952, 0.976, and 0.971, respectively. The AUC values of the GSE62452 dataset were 0.850, 0.890, and 0.850, respectively. The AUC values of the GSE28735 dataset were 0.842, 0.846, and NA.
[0027] Step 4 described above, prognostic gene correlation analysis, analyzed the impact of the expression patterns of these 18 prognostic genes on the prognosis of PDAC patients. The results showed that 9 of these genes had a significant impact on the survival of PAAD. Through comparison, it was found that there were significant differences in the expression of 8 out of these 9 genes among PADG subtypes, p < 0.05; there were significant differences in the expression of 8 genes between the disease and control groups, p < 0.05. These 8 genes were defined as PADG biomarker genes, including COL17A1, ERBB3, ESRP1, GPRC5A, KLF5, MAL2, PCDH1, and SERPINB5. The KM survival curve showed that the high expression of the 8 PADG biomarker genes was associated with a poor prognosis in PADC patients, p < 0.05.
[0028] Step 5 described above, immunotherapy evaluation of the machine learning prognostic model based on the pancreatic cancer dependency gene score, conducted a detailed correlation analysis of 8 PADG biomarker genes and 68 immune checkpoints. The results showed that the biomarker genes COL17A1, ERBB3, ESRP1, GPRC5A, KLF5, MAL2, PCDH1, and SERPINB5 showed significant positive correlations with immune-related molecules involved in cell adhesion, apoptosis regulation, and immune regulation, indicating their potential to enhance the immune response. In addition, these 8 PADG biomarker genes were significantly negatively correlated with the advantages of immune checkpoints, suggesting their inhibitory ability on immune activity. In addition, significant differential expression patterns of immune checkpoints in PADG subtypes were noted. When evaluating the immunotherapy reactivity, the dysfunction score, exclusion score, and comprehensive TIDE score of the PADG subtypes were calculated, all of which showed significant inter-group differences. The results indicated that the machine learning prognostic model based on the pancreatic cancer dependency gene score was expected to predict the reactivity of patients to immunotherapy.
[0029] Advantages of the present invention:
[0030] The method for constructing a prognostic model based on multi-algorithm optimized pancreatic cancer dependency gene score described in the present invention constructs a prognostic model with high precision and stable performance by integrating multiple machine learning algorithms, and can accurately predict the risk and prognosis of pancreatic cancer patients receiving immunotherapy. Achieving personalized treatment: By combining the expression of 8 PADG biomarker genes in the individual prognostic model of PDAC patients, a personalized immunotherapy plan can be formulated for pancreatic cancer patients, improving the treatment effect and prolonging the patient's survival period. Discovering new treatment targets: By screening pancreatic cancer dependency genes and PADG biomarker genes, new ideas and methods for the treatment of pancreatic cancer are provided, which helps to discover new treatment targets and promote the development of the field of pancreatic cancer treatment. Description of the Drawings
[0031] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments:
[0032] Figure 1 For the analysis of copy number variation of ductal cells and the identification of pancreatic cancer dependence activity;
[0033] Figure 2 For the differential analysis between different PADG subtypes in pancreatic ductal carcinoma patients;
[0034] Figure 3 For constructing a prognostic model based on an integrated method of machine learning;
[0035] Figure 4 For the survival analysis and expression of prognostic genes;
[0036] Figure 5 For the immune-related analysis of pancreatic cancer patients. Specific embodiments
[0037] Example 1
[0038] The method for constructing a prognostic model for pancreatic cancer dependence gene score based on multi-algorithm optimization of the present invention aims to construct a prognostic model with high precision and stable performance by integrating a variety of advanced machine learning algorithms, combined with the screening of pancreatic cancer dependence genes and the evaluation of personalized immunotherapy, so as to provide a scientific basis for the clinical decision-making of pancreatic cancer patients, thereby improving the treatment effect and the survival period of patients.
[0039] Step 1, identification of pancreatic cancer malignant cells and pancreatic cancer dependence gene activity;
[0040] Perform reclustering analysis on ductal cells to obtain a total of eight clusters, namely cluster0, 1, 2, 3, 4, 5, 6, and 7. See Appendix Figure 1 A. Select cluster5 from the normal population as the reference cells, and use unsupervised clustering technology to distinguish high-CNVs and low-CNVs cells. It is found that the CNV scores of Cluster 0, 1, 2, 3, 4, and 6 are significantly higher than those of other clusters. See Appendix Figure 1 B, 1C. Therefore, Cluster 0, 1, 2, 3, 4, and 6 are defined as malignant ductal cells. See Appendix Figure 1 D. Subsequently, calculate the pancreatic cancer dependence gene (PADG) activity scores of each malignant ductal cell, and use 0.15 as the threshold to divide the malignant cells into two cell populations: PADG highly active malignant cells and PADG low active malignant cells. See Appendix Figure 1E. By comparing PADG-highly active malignant cells and PADG-lowly active malignant cells, a total of 442 differentially expressed genes (DEGs) were identified, and the differences in these genes between the two groups were statistically significant, with an adjusted p < 0.05 and |Log2 fold change| > 0.5.
[0041] Appendix Figure 1 Analysis of copy number variation in ductal cells and identification of pancreatic cancer dependencies. (A) shows the re-clustering results of ductal cells among different groups. (B) The inferCNV heatmap shows the cell CNV scores. (C) The differential bar chart shows the CNV scores between different cells. (D) The t-SNE plot shows the annotation results of different ductal cells. (E) The box plot shows the distribution of PADG-related activity scores in 6 clusters of malignant ductal cells.
[0042] Step 2, consensus clustering to identify PADG subgroups;
[0043] According to the expression of PADG genes, 2 different subtypes of pancreatic ductal adenocarcinoma were identified by unsupervised clustering analysis, as shown in Appendix Figure 2 A). The optimal clustering with k = 2 indicates that the data is reliable and stably differentiated into 2 PADG activity Clusters with different activities, as shown in Appendix Figure 2 B. Moreover, Kaplan-Meier survival analysis showed significant prognostic differences between these subtypes. Specifically, the survival outcomes of patients in PADG Cluster1 were significantly better than those of patients in Cluster 2, as shown in Appendix Figure 2 C. By comparing different PADG subtypes including Cluster1 and Cluster2, a total of 634 differentially expressed genes (DEGs) were identified, and the differences in these genes between the two groups were statistically significant, with |log2Fold Change| > 1 and the P value after FDR correction < 0.05.
[0044] Appendix Figure 2 Differential analysis between different PADG subtypes in patients with pancreatic ductal carcinoma. (A) Consensus clustering according to K = 2. (B) The cumulative distribution function (CDF) plot shows the consensus distribution of each cluster (k) of patients with pancreatic ductal carcinoma; (C) Survival analysis shows the estimated survival probabilities of different PADG subtypes in patients with pancreatic ductal carcinoma.
[0045] Step 3, construction of a prognostic model;
[0046] By comparing tumors and controls in the transcriptome data of GSE28735, GSE183795, and GSE62452, a total of 471 differentially expressed genes (DEGs) were identified. The differences in these genes between the two groups were statistically significant, with |log2FoldChange| > 1 and p < 0.05 after FDR correction. All DEGs were visualized using a volcano plot, as shown in Appendix Figure 3 A). By comparing highly active malignant cells of PADG with all other cells, a total of 857 differentially expressed genes (DEGs) were identified. The differences in these genes between the two groups were statistically significant, with corrected p < 0.05 and |Log2 fold change| > 0.5. Combining the differentially expressed genes between PADG subtypes, between tumor and control groups, and between highly active malignant cells of PADG and all other cells, the intersection of the three groups of differentially expressed genes was summarized to obtain a list of 39 genes, as shown in Appendix Figure 3 B. Further, through univariate COX regression analysis, 22 genes significantly associated with prognosis were screened out, with p value < 0.05, as shown in Appendix Figure 3 C.
[0047] To construct an efficient PADG-related prognostic model, ten advanced machine learning algorithms were used and tested on the GSE183795 cohort and two external validation datasets, including GSE62452 and GSE28735, to select the model with the highest average C-index value in the validation cohort. Finally, 18 most valuable PADG-related feature genes were identified based on the RSF+StepCox[forward] algorithm, and the best-performing prognostic model was constructed, as shown in Appendix Figure 3 D).
[0048] Survival analysis was performed on the GSE62452, GSE183795, GSE28735, and TCGA datasets. The results showed that a high risk score was associated with a reduced survival time, as shown in Appendix Figure 3 E-H). The AUC values of the GSE183795 dataset at 1 year, 3 years, and 5 years were 0.952, 0.976, and 0.971, respectively; the AUC values of the GSE62452 dataset were 0.850, 0.890, and 0.850, respectively; the AUC values of the GSE28735 dataset were 0.842, 0.846, and NA, respectively. The results emphasized the prognostic significance of the PADG-related model.
[0049] Appendix Figure 3Construct a prognostic model based on a machine learning-based integrated method. (A) A volcano plot depicts the distribution of differentially expressed genes between pancreatic cancer and control samples. Orange, blue, and gray dots represent gene expression levels associated with upregulation, downregulation, and no significant expression, respectively. (B) A Venn diagram shows the final key genes obtained from different differentially expressed genes. (C) Univariate Cox analysis reveals the correlation between key genes and prognosis. (D) A heatmap of C-index values of 101 machine learning algorithm combinations in different datasets, and the right bar chart shows the average C-index values of different algorithm combinations in the validation cohort. (E) KM survival curves demonstrate the association between risk scores and overall survival in the GSE183795 dataset. (F) KM survival curves demonstrate the association between risk scores and overall survival in the GSE28735 dataset. (G) KM survival curves demonstrate the association between risk scores and overall survival in the GSE62452 dataset. (H) KM survival curves demonstrate the association between risk scores and overall survival in the TCGA-PAAD dataset.
[0050] Step 4, Prognostic gene correlation analysis;
[0051] Analyze the impact of the expression patterns of these 18 prognostic genes on the prognosis of PDAC patients. The results show that 9 of these genes have a significant impact on the survival of PAAD. Through comparison, it is found that there are significant differences in the expression of 8 of these 9 genes among PADG subtypes, p < 0.05, see Appendix Figure 4 A. There are significant differences in the expression of 8 genes between the disease and control groups, p < 0.05, Appendix Figure 4 B. Define these 8 genes as PADG biomarker genes, including COL17A1, ERBB3, ESRP1, GPRC5A, KLF5, MAL2, PCDH1, and SERPINB5. KM survival curves show that high expression of the 8 PADG biomarker genes is associated with poor prognosis in PADC patients, p < 0.05, see Appendix Figure 4 C.
[0052] Appendix Figure 4 Survival analysis and expression of prognostic genes. (A) Box plots show the expression distribution of prognostic genes in PADG subtypes. (B) Box plots show the expression distribution of prognostic genes between pancreatic cancer and control groups. (C) Survival curves of high and low expression groups of prognostic genes. Asterisks indicate p-values: ****p < 0.0001, ***p < 0.001, **p < 0.01, *p < 0.05.
[0053] Step 5, Immunotherapy evaluation of a machine learning prognostic model based on pancreatic cancer-dependent gene scores;
[0054] An exhaustive correlation analysis was performed on 8 PADG biomarker genes and 68 immune checkpoints. The results showed that the biomarker genes COL17A1, ERBB3, ESRP1, GPRC5A, KLF5, MAL2, PCDH1, and SERPINB5 showed significant positive correlations with immune-related molecules involved in cell adhesion, apoptosis regulation, and immune regulation, thus indicating their potential to enhance the immune response. In addition, these 8 PADG biomarker genes were significantly negatively correlated with the predominance of immune checkpoints, suggesting their inhibitory ability on immune activity. See Appendix Figure 5 A. In addition, significant differential expression patterns of immune checkpoints, including BTN2A2, BTNL9, and CEACAM1, were noted in the PADG subtypes. See Appendix Figure 5 B. When evaluating immunotherapy responsiveness, dysfunction scores, exclusion scores, and comprehensive TIDE scores of the PADG subtypes were calculated, all of which showed significant inter-group differences. See Appendix Figure 5 C. The results indicated that a machine learning prognostic model based on the pancreatic cancer dependency gene score was promising for predicting patients' responsiveness to immunotherapy.
[0055] To gain in-depth understanding of the immune status of high-risk and low-risk patient cohorts, the changes in TIDE scores and risk scores were carefully examined, and most patients were concentrated in the high-risk group. See Appendix Figure 5 D. These observations highlight that high-risk patients in risk stratification may benefit more from immunotherapy.
[0056] Appendix Figure 5 Immune-related analysis of pancreatic cancer patients. (A) Bubble plot showing the correlation between 8 biomarker genes and 68 immune checkpoints. (B) Box plot showing the expression distribution of immune checkpoints in PADG subtypes. (C) Box plot showing the distribution of Dysfunction scores, Exclusion scores, and comprehensive TIDE scores among PADG subtypes. (D) Sankey diagram of risk scores and immune responses. Asterisks indicate p-values: ****p < 0.0001, ***p < 0.001, **p < 0.01, *p < 0.05.
Claims
1. A method for constructing a prognostic model based on a pancreatic cancer dependency gene score optimized by multiple algorithms, characterized in that: Use machine learning algorithms to identify biomarkers related to pancreatic cancer prognosis and survival, and construct a prognostic model with high precision and stable performance; The specific steps are as follows: Step 1: Identification of pancreatic cancer malignant cells and pancreatic cancer-dependent gene activities; Pancreatic cancer cells originate from ductal cells. All malignant and non-malignant cells are distinguished by analyzing the copy number variations (CNVs) of ductal cells in the gene expression data of the single-cell sequencing dataset (GSE212966); re-clustering analysis is performed on ductal cells, and the clusters derived from normal populations are selected as reference cells. Unsupervised clustering technology is used to distinguish high-CNVs and low-CNVs cells, and the clusters with significantly higher CNV scores than other cells are defined as malignant ductal cells; Subsequently, the pancreatic cancer-dependent gene (PADG) activity scores of each malignant ductal cell are calculated, and the natural value is taken as the threshold to divide the malignant cells into two cell populations: PADG highly active malignant cells and PADG low-active malignant cells; Step 2: Consensus clustering to identify PADG subgroups; Download CRISPR gene effect data from 22 pancreatic cancer cell lines and 295 cell lines of other cancer types on the DepMap portal website; define the gene effects of pancreatic cancer cell lines < the gene effects of other cancer cell lines as pancreatic cancer-dependent genes; according to the expression of PADG genes, different pancreatic ductal adenocarcinoma subtypes are determined by unsupervised clustering analysis, and the prognostic differences between these subtypes are analyzed by survival analysis; Step 3: Construction of a prognostic model; Combining the differentially expressed genes between PADG subtypes, the differentially expressed genes between tumors and control groups, and the differential genes between PADG highly active malignant cells and all other cells, the intersection of the three groups of differential genes is summarized to obtain a co-differential gene; further, through univariate COX regression analysis, genes significantly related to prognosis are screened out; To construct an efficient PADG-related prognostic model, ten advanced machine learning algorithms are used and tested on the GSE183795 cohort and two external validation datasets to select the model with the highest average C-index value in the validation cohort; based on the RSF+StepCox[forward] algorithm, the most valuable PADG-related feature genes are identified, and the best-performing prognostic model is constructed; Step 4: Analysis of the correlation of prognostic genes; Analyze the impact of the expression patterns of the above prognostic genes on the survival and prognosis of PDAC patients, and compare the expression differences of these genes between PADG subtypes; define the genes obtained by the above methods and all having significant differences as PADG biomarker genes; Step 5: Immunotherapy evaluation of the machine learning prognostic model based on the pancreatic cancer-dependent gene score; Perform a correlation analysis on the obtained PADG biomarker genes and the existing 68 immune checkpoints to evaluate their immune response potential and immune activity; when evaluating the responsiveness to immunotherapy, evaluate the responsiveness of patients to immunotherapy based on the risk stratification of biomarker gene expression by calculating the between-group differences in the dysfunction score, exclusion score, and comprehensive TIDE score of the PADG subtypes.
2. The method for constructing a prognostic model for pancreatic cancer-dependent gene scoring optimized based on multiple algorithms according to claim 1, characterized in that: In step one described above, consensus clustering is used to identify PADG subgroups, and reclustering analysis is performed on ductal cells to obtain a total of eight clusters: cluster0, 1, 2, 3, 4, 5, 6, and 7; cluster5 derived from normal populations is selected as the reference cell, and unsupervised clustering technology is used to distinguish high-CNVs and low-CNVs cells. It is found that the CNV scores of Cluster0, 1, 2, 3, 4, and 6 are significantly higher than those of other clusters, and Cluster 0, 1, 2, 3, 4, and 6 are defined as malignant ductal cells; subsequently, the pancreatic cancer dependence gene (PADG) activity scores of each malignant ductal cell are calculated, and using 0.15 as the threshold, the malignant cells are divided into two cell populations: PADG highly active malignant cells and PADG low-active malignant cells; by comparing PADG highly active malignant cells and PADG low-active malignant cells, a total of 442 differentially expressed genes (DEGs) are identified, and the differences between the two groups are statistically significant, with an adjusted p < 0.05 and |Log2 fold change| > 0.
5.
3. The method for constructing a prognostic model for pancreatic cancer dependence gene score optimized based on multiple algorithms according to claim 1, wherein: In step two described above, consensus clustering is used to identify PADG subgroups, and according to the expression of PADG genes, two different pancreatic ductal adenocarcinoma subtypes are determined by unsupervised clustering analysis; the optimal clustering with k = 2 indicates that the data are reliably and stably differentiated into two PADG activity Clusters with different activities; moreover, Kaplan-Meier survival analysis shows significant prognostic differences between these subtypes; specifically, the survival outcomes of patients in PADG Cluster1 are significantly better than those of patients in Cluster 2; by comparing different PADG subtypes, a total of 634 differentially expressed genes are identified, and the differences between the two groups are statistically significant, with |log2Fold Change| > 1 and an FDR-adjusted P value < 0.
05.
4. The method for constructing a prognostic model for pancreatic cancer dependence gene score optimized based on multiple algorithms according to claim 1, wherein: In Step 3, for the construction of the prognostic model, by comparing tumors and controls in the transcriptome data of GSE28735, GSE183795, and GSE62452, a total of 471 differentially expressed genes were identified. The differences in these genes between the two groups were statistically significant, with |log2Fold Change| > 1 and p < 0.05 after FDR correction. All DEGs were visually displayed through volcano plots. By comparing highly active malignant cells of PADG and all other cells, a total of 857 differentially expressed genes were identified. The differences in these genes between the two groups were statistically significant, with corrected p < 0.05 and |Log2 foldchange| > 0.
5. Combining the differentially expressed genes between PADG subtypes, between tumors and control groups, and between highly active malignant cells of PADG and all other cells, the intersection of the three groups of differentially expressed genes was summarized to obtain a list containing 39 genes. Further, through univariate COX regression analysis, 22 genes significantly related to prognosis were screened out, with pvalue < 0.
05. To construct an efficient PADG-related prognostic model, ten advanced machine learning algorithms were used and tested on the GSE183795 cohort and two external validation datasets to select the model with the highest average C-index value in the validation cohort. Finally, based on the RSF+StepCox[forward] algorithm, 18 most valuable PADG-related feature genes were identified, and the prognostic model with the best performance was constructed. Survival analysis was performed on the GSE62452, GSE183795, GSE28735, and TCGA datasets. The results showed that a high risk score was associated with a reduced survival time. The AUC values of the GSE183795 dataset at 1 year, 3 years, and 5 years were 0.952, 0.976, and 0.971, respectively. The AUC values of the GSE62452 dataset were 0.850, 0.890, and 0.850, respectively. The AUC values of the GSE28735 dataset were 0.842, 0.846, and NA, respectively.
5. The method for constructing a prognostic model for pancreatic cancer-dependent gene scores optimized based on multiple algorithms according to claim 1, wherein: In Step 4, for the prognostic gene correlation analysis, the effect of the expression patterns of these 18 prognostic genes on the prognosis of PDAC patients was analyzed. The results showed that 9 of these genes had a significant impact on the survival of PAAD. Through comparison, it was found that there were significant differences in the expression of 8 of these 9 genes between PADG subtypes, with p < 0.
05. There were significant differences in the expression of 8 genes between the disease and control groups, with p < 0.
05. These 8 genes were defined as PADG biomarker genes, including COL17A1, ERBB3, ESRP1, GPRC5A, KLF5, MAL2, PCDH1, and SERPINB5. The KM survival curve showed that the high expression of the 8 PADG biomarker genes was associated with a poor prognosis in PADC patients, with p < 0.
05.
6. The method for constructing a prognostic model for pancreatic cancer dependency gene score optimized based on multiple algorithms according to claim 1, wherein: Step five described above: Step five, immunotherapy evaluation of the machine learning prognosis model based on pancreatic cancer dependency gene score. An exhaustive correlation analysis was performed on 8 PADG biomarker genes and 68 immune checkpoints. The results showed that the biomarker genes COL17A1, ERBB3, ESRP1, GPRC5A, KLF5, MAL2, PCDH1, and SERPINB5 showed significant positive correlations with immune-related molecules involved in cell adhesion, apoptosis regulation, and immune regulation, thus indicating their potential to enhance the immune response; In addition, these 8 PADG biomarker genes were significantly negatively correlated with the advantages of immune checkpoints, suggesting their inhibitory ability on immune activity; in addition, significant differential expression patterns of immune checkpoints in PADG subtypes were noted; when evaluating the immunotherapy reactivity, the dysfunction score, exclusion score, and comprehensive TIDE score of PADG subtypes were calculated, all of which produced significant inter-group differences; the results showed that the machine learning prognosis model based on pancreatic cancer dependency gene score is expected to predict the reactivity of patients to immunotherapy.