Prognostic prediction model for PD-L1 inhibitor targeted therapy and construction method
By constructing a prognostic prediction model for targeted therapy of PD-L1 inhibitors based on the expression of immune gene RNA, screening the immune gene pairs that are most related to clinical responses, the problem of insufficient prediction accuracy of PD-L1 inhibitor treatment in urothelial carcinoma is solved, and the personalized treatment plan is improved, especially for patients who cannot receive platinum drugs treatment.
Patent Information
- Application Number
- CN202410645815.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-23
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2044-05-23
AI Technical Summary
The existing prognostic prediction model for PD-L1 inhibitors in the treatment of urothelial carcinoma is insufficient in accuracy and cannot effectively guide personalized treatment plans, especially for patients who cannot receive platinum drugs.
A prediction model for targeted therapy of PD-L1 inhibitors based on the expression of immune gene RNA was constructed. The immune gene pairs that are most related to clinical responses were screened through the immune gene expression sequence algorithm, and the immune gene expression sequence matrix was constructed, and the prediction model was trained to predict the response of tumor patients to PD-L1 inhibitors.
It improves the prognostic prediction accuracy and efficiency of PD-L1 inhibitor treatment response, helps to achieve personalized treatment plans, improves treatment success rate, and demonstrates its application potential in other tumor types.
Smart Images

Figure CN118692700B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of tumor-targeted drugs, and particularly to the prediction of the therapeutic effect of tumor PD-L1 inhibitors. Background Art
[0002] Urothelial carcinoma is one of the most common malignant tumors of the urinary system. Urothelial carcinoma is the most common histological type of bladder cancer and can occur throughout the urinary tract, including the ureters, renal pelvis, and urethra, but most commonly in the bladder.
[0003] The current standard treatment regimen for urothelial carcinoma is chemotherapy based on gemcitabine combined with cisplatin. For patients who are unable to receive cisplatin treatment due to renal insufficiency, severe illness, or other major complications, a treatment regimen with carboplatin is used. In addition, many patients with urothelial carcinoma are diagnosed at an advanced age, and platinum-based chemotherapy may not be applicable given their age or the presence of other serious complications, thus limiting the treatment options for these patients and resulting in poor efficacy.
[0004] In recent years, major breakthroughs in the field of immunology have revolutionized the treatment strategy for urothelial carcinoma. In particular, the in-depth understanding of how tumors evade immune surveillance has promoted the widespread use of immune checkpoint inhibitors (ICIs). These drugs have shown a high objective response rate and durable efficacy in various tumor types and have been approved for the treatment of multiple cancers, including malignant melanoma, non-small cell lung cancer, and urothelial carcinoma. Immune checkpoint inhibitors significantly enhance the anti-tumor activity of many patients with advanced urothelial carcinoma who are resistant to platinum-based drug treatment by enhancing the body's immune response to tumors; in addition, these drugs have also shown potential as a first-line treatment regimen in clinical trials of patients who are intolerant to cisplatin.
[0005] In urothelial carcinoma and other solid malignancies, the most clinically valuable immune checkpoint inhibitor is programmed death receptor-1 (PD-1). PD-1 is a protein expressed on the surface of T cells, which interacts with multiple ligands, including programmed death 1 ligand 1 (PD-L1) and programmed death ligand-2 (PD-L2). Programmed cell death ligand-1 (PD-L1) is a key factor in the tumor immune escape mechanism. PD-L1 is usually expressed in immune cells other than T cells, including antigen-presenting cells and tumor cells. The interaction after PD-L1 binds to its ligand leads to the inhibition of T cell activity and immune response, thus helping tumor cells escape the attack of the immune system. Therefore, PD-L1 has become an important biomarker for evaluating tumor treatment response and prognosis. Immune checkpoint inhibitors targeting PD-L1 have become an important means of tumor treatment. By developing inhibitors targeting PD-L1, the inhibitory effect on T cells can be offset in the short term, thereby enhancing the anti-tumor immune response.
[0006] In recent years, with the successive application of relevant PD-L1 inhibitors in clinical practice, the prognostic study after PD-L1 inhibitor treatment has become a new research hotspot. There have been many reports on PD-L1 inhibitors and prognosis. For example, in 2020, the Dana-Farber Cancer Institute analyzed the survival period of patients treated with PD-L1 inhibitors through five factors: ECOG-PS, liver metastasis, platelet count, neutrophil-lymphocyte ratio, and lactate dehydrogenase in 405 patients with urothelial carcinoma, and constructed a risk assessment model to predict the survival probability of patients. In 2023, a research team from the Icahn School of Medicine at Mount Sinai analyzed the disease-free survival of 709 patients with urothelial carcinoma based on PD-L1 expression in tumor and immune cells to predict the prognosis of patients. Summary of the Invention
[0007] One of the technical problems to be solved by the present invention is to provide a method for constructing a prognostic prediction model for PD-L1 inhibitor targeted therapy. The model constructed by this method can accurately predict the drug response of PD-L1 inhibitors in tumors, especially urothelial carcinoma.
[0008] To solve the above technical problems, the construction method of the PD-L1 inhibitor targeted therapy prognosis prediction model of the present invention constructs an expression data matrix of immune genes based on the RNA expression levels of immune genes, applies an immune gene expression sequence algorithm to screen immune gene pairs most relevant to the clinical response of treating tumors with PD-L1 inhibitors and constructs an immune gene expression sequence matrix, and obtains a prognosis prediction model for PD-L1 inhibitor targeted therapy of tumors through training.
[0009] The immune gene expression sequence algorithm preferably includes the following steps:
[0010] 1) Perform rank transformation on the expression values of immune genes in the expression data matrix of immune genes to obtain an immune gene rank transformation matrix;
[0011] 2) Calculate the gene rank difference for each sample in the immune gene rank transformation matrix to obtain an immune gene rank difference matrix;
[0012] 3) Transpose the immune gene rank difference matrix and add the clinical information of each sample to obtain an immune gene rank difference matrix with clinical information;
[0013] 4) Perform feature selection on the immune gene rank difference matrix with clinical information to screen immune genes related to the clinical response;
[0014] 5) Extract the immune gene rank difference matrix according to the immune genes screened in step 4) to obtain a high-feature immune gene rank difference matrix; calculate the rank difference of each pair of immune genes through nested loops to obtain an immune gene pair rank difference matrix, and add the clinical information of each sample to the immune gene pair rank difference matrix to obtain an immune gene pair rank difference matrix with clinical information;
[0015] 6) Perform feature selection on the immune gene pair rank difference matrix with clinical information obtained in step 5) to screen immune gene pairs most relevant to the clinical response and obtain an immune gene expression sequence matrix.
[0016] The clinical information described above includes recurrence-free survival time, recurrence-free survival status, and PD-L1 inhibitor treatment response data.
[0017] The second technical problem to be solved by the present invention is to provide a PD-L1 inhibitor targeted therapy prognosis prediction model constructed by the above method.
[0018] The third technical problem to be solved by the present invention is to provide the use of the above PD-L1 inhibitor targeted therapy prognosis prediction model, which can be used for prognosis prediction after PD-L1 inhibitor targeted therapy for tumor patients (especially urothelial cancer patients).
[0019] The present invention constructs a prognostic prediction model for PD-L1 inhibitor targeted therapy based on the RNA expression levels of paired immune genes. By deeply analyzing the expression levels of relevant immune genes in tumor samples after PD-L1 inhibitor treatment, it predicts the possible responses of tumor patients, especially urothelial cancer patients, to PD-L1 inhibitor treatment, greatly improving the prognostic prediction accuracy and efficiency of PD-L1 inhibitor treatment responses. This helps to achieve personalized treatment plans and improve treatment success rates. Verification through pan-cancer data shows that this model also has the ability and potential to be applied to other tumor types. Description of the Drawings
[0020] Figure 1 It is the survival curve of the high-risk group and the low-risk group in the training set of Example 1 of the present invention; the treatment response of PD-L1 inhibitor is significantly correlated with the overall survival (OS).
[0021] Figure 2 It is the ROC curve of the PD-L1 inhibitor targeted therapy prognostic prediction model constructed in Example 1 of the present invention in the training set.
[0022] Figure 3 It is the survival curve of the high-risk group and the low-risk group in the validation set of Example 2 of the present invention.
[0023] Figure 4 It is the ROC curve of the PD-L1 inhibitor targeted therapy prognostic prediction model in the validation set of Example 2 of the present invention.
[0024] Figure 5 It is the survival curve of the high-risk group and the low-risk group in the extended validation set of Example 3 of the present invention.
[0025] Figure 6 It is the ROC curve of the PD-L1 inhibitor targeted therapy prognostic prediction model in the extended validation set of Example 3 of the present invention. Detailed Embodiments
[0026] For a more specific understanding of the technical content, features and effects of the present invention, the technical solutions of the present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0027] Example 1 Construction of the PD-L1 Inhibitor Targeted Therapy Prognostic Prediction Model
[0028] I. Collecting Gene Expression Profile Data and Clinical Data of Tumor Patients
[0029] Collect 476 tumor patient samples, including 378 urothelial cancer samples and 98 pan-cancer samples.
[0030] Among 378 urothelial carcinoma samples, 298 samples were from the gene expression data in the IMvigor210 Core Biologies package (Mariathasan, S., Turley, S., Nickles, D. et al. TGFβ attenuates tumour response to PD-L1 blockade by contributing to exclusion of T cells. Nature 554, 544–548 (2018).), and 80 samples were from the applicant's own sequencing data. At the same time, the corresponding clinical data of all samples were obtained (including recurrence-free survival time, recurrence-free survival status, and PD-L1 inhibitor treatment response data).
[0031] The gene expression data of 98 pan-cancer samples were collected from the literature Jiao, X., Wei, X., Li, S. et al. A genomic mutation signature predicts the clinical outcomes of immunotherapy and characterizes immunophenotypes in gastrointestinal cancer. npj Precis. Oncol. 5, 36 (2021). The information of pan-cancer samples is shown in Table 1.
[0032] Table 1 Information of 98 pan-cancer samples
[0033]
[0034]
[0035] II. Transcriptome data preprocessing
[0036] The Trimmomatic software was used to perform quality control on the transcriptome sequencing data of the above 80 urothelial carcinoma patient samples owned by the applicant, removing adapter sequences and filtering out low-quality fragments. The quality-controlled data were mapped to the human reference genome hg19 using the STAR software to obtain the aligned BAM file. The BAM file was used for gene expression quantification with the RSEM tool and FPKM normalization to obtain the gene expression values after FPKM normalization. The FPKM normalization conversion formula is:
[0037]
[0038] where i is the sample number, j is the gene number, and R ij is the reads count value of gene j in sample i, and F ijis the FPKM value of gene j in sample i, L j is the coding region length of gene j, T i is the number of sequencing reads of sample i.
[0039] Calculate the median expression value of each gene, and delete the genes with a median value of 0 to filter out low-expressed genes.
[0040] Merge the expression data of 298 urothelial carcinoma samples in the IMvigor210CoreBiologies package dataset with the expression data of 80 urothelial carcinoma samples owned by the applicant, and use removeBatchEffect of limma to remove batch effects.
[0041] III. Screening of immune-related genes
[0042] According to relevant literature and databases (1. Vogelstein B, Papadopoulos N, Velculescu VE, Zhou S, Diaz LA Jr, Kinzler KW. Cancer genome landscapes. Science. 2013 Mar
[0043] 29; 339(6127):1546 - 58.; 2, Thorsson V, Gibbs DL, et al. The Immune Landscape of Cancer. Immunity. 2018 Apr 17;48(4):812 - 830.e14.; 3, Kanehisa M, Goto S. KEGG: kyoto encyclopedia of genes and genomes. Nucleic Acids Res. 2000 Jan 1;28(1):27 - 30.), collect and organize immune - related genes. Immune genes can be classified into the following major categories: immune regulatory genes, cell signal transduction genes, cell surface receptor genes, cytokine genes, growth factor genes, cytokine receptor genes. Specifically, they can be divided into immune checkpoint and regulatory genes, immune activation genes, antigen - presentation - related genes, cytokines and their receptors, T - cell, B - cell, natural killer cell marker genes, inflammation - related genes, cell adhesion and cell signal transduction genes, heat - shock proteins and stress - response genes, immune signaling pathway genes, etc. Blast the screened genes against all human genes to find all homologous genes, and then take the intersection with the expression data obtained in the second step above. Finally, the following 1039 immune genes are screened out: AZGP1, B2M, CALR, CANX, CD1A, CD1B, CD1C, CD1D, CD1E, CD4, CD8A, CD8B, CD74, CREB1, CTSB, CTSE, CTSL1, CTSS, FCER1G, FCGRT, PDIA3, HFE, HLA - A, HLA - B, HLA - C, HLA - DMA, HLA - DMB, HLA - DOA, HLA - DOB, HLA - DPA1, HLA - DPB1, HLA - DQA1, HLA - DQA2, HLA - DQB1, HLA - DRA, HLA - DRB1, HLA - DRB3, HLA - DRB4, HLA - DRB5, HLA - E, HLA - F, HLA - G, HLA - H, MR1, HSPA1B, HSPA1L, HSPA2, HSPA4, HSPA5, HSPA6, HSPA8, HSP90AA1, HSP90AB1, ICAM1, IFNA2, IFNA4, IFNA5, IFNA6, IFNA7, IFNA8, IFNA10, IFNA13, IFNA14, IFNA16, IFNA17, IFNA21, IFNG, KIR2DL1, KIR2DL2, KIR2DL3, KIR2DL4, KIR2DS1, KIR2DS3, KIR2DS4, KIR2DS5, KIR3DL1, KIR3DL2, KLRC1,KLRC2, KLRC3, KLRD1, LTA, CIITA, MICA, MICB, NFYA, NFYB, NFYC, LGMN, PSMB8, PSMC1, PSMC2, PSMC3, PSMC4, PSMC5, PSMC6, PSMD1, PSMD2, PSMD3, PSMD4, PSMD5, PSMD7, PSMD8, PSMD10, PSMD11, PSMD13, PSME1, PSME2, RELB, RFX5, RFXAP, SLC10A2, TAP1, TAP2, TAPBP, THBS1, SHFM1, KLRC4, AP3B1, RFXANK, PSMD6, PSME3, PSMD14, CLEC4M, IFI30, PROCR, ADRM1, KIAA0368, TRPC4AP, CD209, UBXN1, ERAP1, TAPBPL, KIR2DL5A, ERAP2, ULBP3, ULBP2, ULBP1, KIR3DL3, RAET1E, RAET1L, UBR1, RAET1G, PDIA2, CD79A, CD79B, LYN, SYK, BTK, BLNK, VAV3, VAV1, VAV2, RAC1, RAC2, RAC3, PPP3CA, PPP3CB, PPP3CC, CHP, PPP3R1, PPP3R2, CHP2, NFAT5, NFATC1, NFATC2, NFATC3, NFATC4, HRAS, KRAS, NRAS, FOS, JUN, CARD11, BCL10, MALT1, CHUK, IKBKB, IKBKG, NFKB1, RELA, NFKBIA, NFKBIB, NFKBIE, CD81, CD19, CR2, PIK3R5, PIK3R1, PIK3R2, PIK3R3, PIK3CA, PIK3CB, PIK3CD, PIK3CG, AKT3, AKT1, AKT2, GSK3B, INPP5D, CD22, CD72, PTPN6, LILRB3, FCGR2B, RASGRP3, PLCG2, PRKCB, IFITM1, IGHA1, IGHA2, IGHD, IGHE, IGHG1, IGHG2, IGHG3, IGHG4, IGHM, IGHV2-70, IGHV3-23, IGKC, IGKV1-5, IGKV2-40, IGKV3-20, IGKV3D-11, IGKV3D-20, IGKV4-1, IGLC3, IGLV7-43, C3, C5, CAMP, CCL1, CCL11, CCL13, CCL14, CCL15, CCL16, CCL17, CCL18, CCL19, CCL2CCL20, CCL21, CCL22, CCL23, CCL24, CCL25, CCL26, CCL27, CCL28, CCL3, CCL3L3, CCL4, CCL4L1, CCL5, CCL7, CCL8, CKLF, CMA1, CTSG, CX3CL1, CXCL1, CXCL10, CXCL11, CXCL12, CXCL13, CXCL14, CXCL16, CXCL17, CXCL2, CXCL3, CXCL5, CXCL6, CXCL9, CYR61, DEFA1, DEFA3, DEFA5, DEFB1, DEFB104A, DEFB4, EDN1, EDN2, EDN3, FGF10, FGF2, HTN3, IL8, LECT2, PF4, PF4V1, PLAU, PPBP, PROK2, RNASE2, SAA1, SBDS, SEMA3A, SEMA3B, SEMA3C, SEMA3D, SEMA3E, SEMA3F, SEMA3G, SEMA4A, SEMA4B, SEMA4C, SEMA4D, SEMA4F, SEMA4G, SEMA5A, SEMA5B, SEMA6A, SEMA6B, SEMA6C, SEMA6D, SEMA7A, SLIT1, SLIT2, TNC, TYMP, XCL1, XCL2, C5AR1, CCBP2, CCR1, CCR10, CCR3, CCR4, CCR5, CCR6, CCR7, CCR8, CCR9, CCRL1, CCRL2, CMKLR1, CX3CR1, CXCR3, CXCR4, CXCR5, CXCR6, CXCR7, CYSLTR1, CYSLTR2, DARC, EDNRA, EDNRB, FPR1, FPR2, GPR17, GPR32, GPR33, GPR44, GPR77, IL8RA, IL8RB, LTB4R, LTB4R2, PLAUR, PLXNA1, PLXNA2, PLXNA3, PLXNA4, PLXNB1, PLXNB2, PLXNB3, PLXNC1, PLXND1, PTAFR, ROBO1, ROBO2, ROBO3, RXFP3, XCR1, ADIPOQ, ADM, ADM2, AGRP, AGT, AMBN, AMELX, AMH, ANGPTL5, ANGPTL7, APLN, AREG, ARMET, ARMETL1, ARTN, AVP, AZU1, BDNF, BMP1, BMP10, BMP15, BMP2, BMP3, BMP4, BMP5, BMP6, BMP7, BMP8A, BMP8B, BTC, C19orf10, CALCA, CALCBCAT, CCK, CD320, CD40LG, CD70, CECR1, CER1, CGA, CGB, CGB2, CHGA, CHGB, CLCF1, CLEC11A, CMTM1, CMTM2, CMTM3, CMTM4, CMTM5, CMTM6, CMTM7, CMTM8, CNTF, CORT, CRH, CSF1, CSF2, CSF3, CSH1, CSHL1, CSPG5, CTF1, CTGF, DKK1, EBI3, EGF, EPGN, EPO, EREG, ESM1, FAM3B, FAM3C, FAM3D, FASLG, FGF1, FGF11, FGF12, FGF13, FGF14, FGF16, FGF17, FGF18, FGF19, FGF20, FGF21, FGF22, FGF23, FGF3, FGF4, FGF5, FGF6, FGF7, FGF8, FGF9, FIGF, FIGNL2, FLT3LG, FSHB, GAL, GALP, GAST, GCG, GDF1, GDF10, GDF11, GDF15, GDF2, GDF3, GDF5, GDF6, GDF7, GDF9, GDNF, GH1, GH2, GHRH, GHRL, GIP, GKN1, GMFB, GMFG, GNRH1, GNRH2, GPHA2, GPHB5, GPI, GREM1, GREM2, GRN, GRP, GUCA2A, HAMP, HBEGF, HDGF, HDGFRP3, HGF, IAPP, IFNB1, IFNE, IFN K, IFN W1, IGF1, IGF2, IL10, IL11, IL12A, IL12B, IL13, IL15, IL16, IL17A, IL17B, IL17C, IL17D, IL17F, IL18, IL19, IL1A, IL1B, IL1F10, IL1F5, IL1F6, IL1F7, IL1F8, IL1F9, IL1RN, IL2, IL20, IL21, IL22, IL23A, IL24, IL25, IL26, IL27, IL28A, IL28B, IL29, IL3, IL31, IL32, IL33, IL34, IL4, IL5, IL6, IL6ST, IL7, IL9, INHA, INHBA, INHBB, INHBC, INHBE, INS, INSL3, INSL4, INSL5, INSL6, JAG1, JAG2, KGFLP1, KGFLP2, KITLG, KL, LACRT, LEFTY1, LEFTY2, LEP, LHB, LIF, LRSAM1, LTB, LTBP1, LTBP2, LTBP3LTBP4, MDK, MIA, MIF, MLN, MSTN, NAMPT, NDP, NENF, NGF, NMB, NODAL, NOV, NPFF, NPPA, NPPB, NPPC, NPY, NRG1, NRG2, NRG3, NRG4, NRTN, NTF3, NTF4, NTS, NUDT6, OGN, OSGIN1, OSM, OSTN, OXT, P11, PDGFA, PDGFB, PDGFC, PDGFD, PDGFRA, PDGFRB, PDGFRL, PDYN, PENK, PGF, PMCH, PNOC, POMC, PPY, PRL, PRLH, PROK1, PSPN, PTH, PTH2, PTHLH, PTN, PYY, QRFP, RABEP1, RABEP2, REG1A, RETN, RETNLB, RLN1, RLN2, RLN3, S100A6, SCG2, SCGB3A1, SCT, SCYE1, SECTM1, SLURP1, SPP1, SST, STC1, STC2, TAC1, TDGF1, TDGF3, TG, TGFA, TGFB1, TGFB2, TGFB3, THPO, TNF, TNFRSF11B, TNFSF10, TNFSF11, TNFSF12, TNFSF13, TNFSF13B, TNFSF14, TNFSF15, TNFSF18, TNFSF4, TNFSF8, TNFSF9, TOR2A, TRH, TSHB, TSLP, TXLNA, UCN, UCN2, UCN3, UTS2, UTS2D, VEGFA, VEGFB, VEGFC, VGF, VIP, ACVR1B, ACVR1C, ACVR2A, ACVR2B, ACVRL1, ADCYAP1R1, ADIPOR1, ADIPOR2, ADRB1, ADRB2, AGTR1, AGTR2, AMHR2, ANGPT1, ANGPT4, ANGPTL1, ANGPTL2, ANGPTL3, ANGPTL4, ANGPTL6, APLNR, AR, AVPR1A, AVPR1B, AVPR2, BMPR1A, BMPR1B, BMPR2, BRD8, C3AR1, CALCR, CALCRL, CD40, CNTFR, CRHR1, CRHR2, CRIM1, CRLF1, CRLF2, CRLF3, CSF1R, CSF2RA, CSF2RB, CSF3R, EGFR, ENG, EPOR, ESR1, ESR2, ESRRA, ESRRB, ESRRG, FGFR1, FGFR2, FGFR3, FGFR4, FGFRL1, FLT1, FLT3, FLT4FSHR, GALR2, GALR3, GCGR, GHR, GHRHR, GHSR, GIPR, GLP1R, GLP2R, GNRHR, GPER, HNF4A, HNF4G, HTR3A, HTR3B, HTR3C, HTR3D, HTR3E, IFNAR1, IFNAR2, IFNGR1, IFNGR2, IGF1R, IGF2R, IL10RA, IL10RB, IL11RA, IL12RB1, IL12RB2, IL13RA1, IL13RA2, IL15RA, IL17RA, IL17RB, IL17RC, IL17RD, IL17RE, IL18R1, IL18RAP, IL1R1, IL1R2, IL1RAP, IL1RL1, IL1RL2, IL20RA, IL20RB, IL21R, IL22RA1, IL22RA2, IL23R, IL27RA, IL28RA, IL2RA, IL2RB, IL2RG, IL31RA, IL3RA, IL4R, IL5RA, IL6R, IL7R, IL9R, INSR, KDR, LEPR, LGR4, LGR5, LGR6, LHCGR, LIFR, LTBR, MC1R, MC2R, MC3R, MC4R, MCHR1, MCHR2, MET, MLNR, MPL, MTNR1A, MTNR1B, NGFR, NMBR, NPR1, NPR3, NR0B1, NR0B2, NR1D1, NR1D2, NR1H2, NR1H3, NR1H4, NR1 I2, NR1 I3, NR2C1, NR2C2, NR2E1, NR2E3, NR2F1, NR2F2, NR2F6, NR3C1, NR3C2, NR4A1, NR4A2, NR4A3, NR5A1, NR5A2, NR6A1, NRP1, NRP2, OGFR, OPRD1, OPRK1, OPRL1, OPRM1, OSMR, OXTR, PGR, PGRMC2, PPARA, PPARD, PPARG, PRLHR, PRLR, PTGDR, PTGDS, PTGER1, PTGER2, PTGER3, PTGER4, PTGFR, PTH1R, PTH2R, RARA, RARB, RARG, RORA, RORB, RORC, RXFP1, RXFP2, RXRA, RXRB, RXRG, S1PR1, S1PR2, SCTR, SDC1, SDC2, SDC3, SDC4, SORT1, SSTR1, SSTR2, SSTR5, TACR1, TEK, TGFBR1, TGFBR2, TGFBR3, THRA, THRB, TIE1, TNFRSF10ATNFRSF10B, TNFRSF10C, TNFRSF10D, TNFRSF11A, TNFRSF12A, TNFRSF13B, TNFRSF13C, TNFRSF14, TNFRSF17, TNFRSF18, TNFRSF19, TNFRSF1A, TNFRSF1B, TNFRSF21, TNFRSF25, TNFRSF4, TNFRSF6B, TNFRSF8, TNFRSF9, TRHR, TSHR, TUBB3, VDR, VIPR1, VIPR2, PTPN11, ICAM2, ITGAL, ITGB2, PTK2B, PAK1, MAP2K1, MAP2K2, MAPK1, MAPK3, NCR2, TYROBP, LCK, FCGR3A, FCGR3B, NCR1, NCR3, CD247, ZAP70, LCP2, LAT, PLCG1, SH3BP2, FYN, SHC2, SHC4, SHC3, SHC1, GRB2, SOS1, SOS2, ARAF, BRAF, RAF1, KLRK1, HCST, CD48, CD244, PRKCA, PRKCG, SH2D1B, SH2D1A, FAS, GZMB, PRF1, CASP3, BID, CD3D, CD3E, CD3G, PTPRC, ITK, TEC, NCK1, NCK2, GRAP2, PAK2, PAK3, PAK4, PAK6, PAK7, RHOA, CDC42, CD28, ICOS, MAP3K8, MAP3K14, PDCD1, CTLA4, CBLC, CBL, CBLB, CDK4, RASGRP1, PDK1, PRKCQ, TRBC1, TRBV12-3, TRGV3.
[0044] IV. Screening of Features Most Relevant to Clinical Response
[0045] Select 298 samples from urothelial carcinoma samples as the training set. Integrate the gene expression data of the 298 samples with the above 1039 immune genes screened out, take the intersection, and retain the immune genes existing in the expression data and the corresponding expression data to obtain the expression data matrix of immune genes. The columns of the matrix are samples, the rows are genes, and the data in the matrix are immune gene expression values.
[0046] According to the expression data matrix of immune genes, use the following immune gene expression sequence algorithm to screen the features most relevant to clinical response:
[0047] 1) Rank transformation of immune gene expression values: Multiply the immune gene expression value of each sample by -1, and use the rank function in R language to rank the immune gene expression values of each sample. In this way, the immune gene expression values in each sample will be assigned a new rank value according to the relative level of their expression. The rank transformation formula is:
[0048] Rank(X a )=floor(rank(-X a ))
[0049] where X a is the expression value of the immune gene in the a-th column, and Rank(X a ) is the rank of the immune gene in the a-th column.
[0050] After rank transformation, the values in each sample will change based on their ranking in that sample. The row and column of the transformed expression matrix are still genes and samples, and the values in the matrix change from immune gene expression values to rank values of immune genes. The transformed matrix is called the immune gene rank transformation matrix.
[0051] 2) Calculate the Rank Difference
[0052] For each sample in the immune gene rank transformation matrix, subtract the median rank of that row from each element (the rank of the gene) to calculate the gene rank difference. That is, the gene rank difference calculation formula is:
[0053] Difference(Y ij )=Y ij -median(Y i )
[0054] where Y ij is the rank value of the i-th row and j-th column, median(Y i ) is the median rank of the i-th row, and Difference(Y ij ) is the difference between the rank of the immune gene in the i-th row and j-th column and the median rank of the i-th row.
[0055] By calculating the gene rank difference, the expression level of each gene will be adjusted relative to its own median to identify the genes that truly show significant changes. The transformed matrix is called the immune gene rank difference matrix.
[0056] 3) Transpose the immune gene rank difference matrix. At this time, the rows of the matrix are samples and the columns are genes. Add the clinical data of each sample (including recurrence-free survival time, recurrence-free survival status, PD-L1 inhibitor treatment response data) to the transposed rank value matrix to obtain the immune gene rank difference matrix with clinical information.
[0057] 4) First feature selection
[0058] The random forest algorithm is used to perform the first feature selection on the rank difference matrix of immune genes with clinical information. The sbfControl and rfeControl in the caret package are used to implement the Sequential Backward Selection (SBS) function and the Recursive Feature Elimination (RFE) function respectively, and the features related to clinical response are selected by the method of cross-validation (cv). The features at this time are immune genes related to clinical response.
[0059] 5) Construct the rank difference matrix of immune gene pairs
[0060] Extract the rank difference matrix of immune genes in step 3) according to the immune genes screened in step 4) above to obtain the rank difference matrix of high-feature immune genes. The rows of the matrix are samples, and the columns are the immune genes screened in step 4).
[0061] The rank differences of each pair of immune genes are calculated by nested loops, ensuring that each gene pair is calculated only once and without repetition. If it is a positive difference (that is, the rank difference of one gene is greater than or equal to the rank difference of the other gene), then this position is marked as 1; if it is a negative difference (that is, the rank difference of one gene is less than the rank difference of the other gene), then this position is marked as -1. The names of the gene pairs are connected with the ">" symbol, such as "Gene A > Gene B". Create a new matrix, where the rows represent samples and the columns represent gene pairs, and the positive and negative difference values calculated above are stored in the columns. This new matrix is called the rank difference matrix of immune gene pairs. By constructing the rank difference matrix of immune gene pairs, not only can a gene regulatory network be constructed to facilitate the identification of key gene pairs, but also feature extraction can be facilitated.
[0062] Add the clinical data of each sample (including recurrence-free survival time, recurrence-free survival status, PD-L1 inhibitor treatment response data) to the rank difference matrix of immune gene pairs to obtain the rank difference matrix of immune gene pairs with clinical information.
[0063] 6) Second feature selection
[0064] Again, the random forest algorithm is adopted, and the SBS and RFE methods in the caret package are used to perform the second feature selection on the rank difference matrix of immune genes with clinical information in step 5) above, screening out the immune gene pairs most relevant to the clinical response, as follows: "CASP3>CCL24", "CASP3>FAS", "CASP3>RAC2", "CASP3>TNFRSF10A", "CASP3>TNFRSF10B", "CASP3>TNFRSF10D", "CCL24>GRB2", "CCL24>KIR2DL4",
[0065] "CCL24>KIR3DL1", "CCL24>KIR3DL2", "CCL24>KLRC3", "CCL24>KLRC4", "CCL24>NFKBIB", "CCL24>PSMD3", "CCL24>PSMD4", "CCL24>PSMD7", "CCL24>PSMD8", "CCL24>RAC1", "CCL24>TRH", "CD8A>FAS", "FAS>IFNG", "FAS>KIR2DL4", "FAS>KLRC2", "FAS>KLRC3", "FAS>TRH", "GRB2>PSMD3", "IFNG>KIR3DL2", "IFNG>NFKB1", "IFNG>NFKBIB", "IFNG>PSMD8", "IFNG>RAC2", "IFNG>TNFRSF14", "KIR2DL4>KLRK1", "KIR2DL4>NFKB1", "KIR2DL4>TNFRSF10B", "KIR2DL4>TNFRSF11B", "KIR2DL4>TNFRSF14", "KIR3DL1>PSMD3", "KIR3DL1>PSMD4", "KIR3DL1>PSMD8", "KIR3DL1>TNFRSF10A", "KLRC1>TNFRSF10D", "KLRC2>KLRK1", "KLRC2>NFKB1", "KLRC2>NFKBIB", "KLRC2>PSMD4", "KLRC2>PSMD7", "KLRC2>PSMD8", "KLRC2>TNFRSF10A", "KLRC2>TNFRSF10B", "KLRC2>TNFRSF10D", "KLRC2>TNFRSF14", "KLRC2>TNFRSF1A", "KLRC3>KLRK1", "KLRC3>NFKB1", "KLRC3>TNFRSF14", "KLRC3>TNFRSF1A", "KLRC4>NFKB1", "KLRC4>PIK3R1", "KLRC4>RAC2", "KLRC4>TNFRSF10A", "KLRC4>TNFRSF10D", "KLRC4>TNFRSF11B", "KLRC4>TNFRSF14", "KLRK1>TNFRSF10A", "KLRK1>TNFRSF11B", "NFKB1>PSMD3", "NFKB1>TNFRSF10D", "NFKBIB>TNFRSF10B", "NFKBIB>TNFRSF1A", "PIK3R1>RAC1", "PIK3R1>TRH", "PSMD3>TNFRSF10D""PSMD3>TNFRSF1A", "PSMD4>TNFRSF1A", "PSMD7>TNFRSF10B", "PSMD7>TNFRSF10D", "PSMD7>TNFRSF1A", "PSMD8>TNFRSF10D", "PSMD8>TNFRSF1A", "RAC1>TNFRSF10A", "RAC1>TNFRSF10B", "RAC1>TNFRSF10D", "RAC1>TNFRSF11B", "RAC1>VAV2", "RAC3>TNFRSF10B", "RAC3>TNFRSF10D", "TNFRSF10D>TRH".
[0066] The rank difference matrix of immune gene pairs screened in this step is called the immune gene expression order matrix. The rows of the matrix are samples, and the columns are immune gene pairs.
[0067] V. Constructing a prognostic prediction model for PD-L1 inhibitor targeted therapy
[0068] Use the caret package in R language to train the model with the screened immune gene expression order matrix. The specific method is as follows:
[0069] 1. Set model training parameters
[0070] a) Use the trainControl function to configure model training parameters. Select the bootstrap method as the resampling method and set the number of resampling times to 200 times.
[0071] b) Turn on class probability calculation (classProbs = TRUE) so that the model can predict the probability that the observed value belongs to each class.
[0072] c) Use the twoClassSummary function (binary classification summary function) as the performance evaluation index, and pay attention to sensitivity, specificity, and the area under the ROC curve (AUC).
[0073] 2. Iterative training
[0074] a) Conduct 40 loop iterations. In each iteration, use the createDataPartition function to add PD-L1 inhibitor treatment response data to the training.
[0075] b) Apply the train function to train the random forest model, and use the predict function with the prediction type of probability (type = "prob") to predict the recurrence risk probability of urothelial carcinoma.
[0076] 3. Evaluate model performance
[0077] The roc function of the pROC package was used to calculate the ROC curve (Receiver Operating Characteristic curve) and AUC value for each iteration to evaluate the model performance. When the AUC value of the model exceeded 0.9, the loop was terminated to obtain the optimal PD-L1 inhibitor targeted therapy prognosis prediction model.
[0078] 4. Define the risk
[0079] The recurrence risk probability values predicted by the predict function were defined. When the risk probability was greater than the median of all risk values, the sample was defined as high risk ("High"); otherwise, it was defined as low risk ("Low"). According to the risk definition results of each sample, 298 samples were divided into a high-risk group and a low-risk group.
[0080] VI. Survival analysis
[0081] The survfit function of the survival package in R language was used to perform survival analysis on the risk prediction results, and the univariate Kaplan-Meier survival analysis method was adopted to perform survival analysis on the prediction results of the optimal PD-L1 inhibitor targeted therapy prognosis prediction model to evaluate the prognosis of patients treated with PD-L1 inhibitors predicted by the model.
[0082] The results were as Figure 1 shown. The p-value of the log rank test was 0.0014, indicating a significant difference in the survival curves between the high- and low-risk groups, suggesting a significant difference in the recurrence-free survival rate between patients in different risk groups. As Figure 2 shown, in further analysis, the ROC curve analysis showed that the AUC value was 0.93, indicating that using the PD-L1 expression level as a predictive marker for the treatment effect of PD-L1 inhibitors in urothelial carcinoma had high accuracy and could effectively distinguish patients with good responses to PD-L1 inhibitor treatment.
[0083] Example 2 Validation of the PD-L1 inhibitor targeted therapy prognosis prediction model
[0084] Another 80 samples from the above 378 urothelial carcinoma samples were used as the validation set to validate the PD-L1 inhibitor targeted therapy prognosis prediction model constructed in Example 1 to ensure the generalization ability and accuracy of the model on an independent dataset.
[0085] 1. Data integration and immune gene feature screening
[0086] Integrate the gene expression data of 80 samples with the 1039 immune genes screened in Example 1, take the intersection, and retain the immune genes present in the expression data and the corresponding expression data to obtain the expression data matrix of the immune genes of the validation set samples.
[0087] According to the expression data matrix of the immune genes of the validation set samples, follow the immune gene expression order algorithm in Example 1 to screen the immune gene pair features most relevant to the clinical response in the validation set samples and construct an immune gene expression order matrix.
[0088] 2. Model Prediction
[0089] Use the predict function in the caret package and the PD-L1 inhibitor targeted therapy prognosis prediction model trained in Example 1 to predict the validation set, set the prediction type to prob, and obtain the urothelial cancer recurrence risk probability prediction values for each sample in the validation set.
[0090] 3. Model Performance Evaluation on the Validation Set
[0091] Use the roc function in the pROC package to calculate the ROC curve and the area under the curve (AUC) on the validation set to evaluate the classification performance of the model on the independent dataset. By calculating the AUC value, we can intuitively understand the ability of the model to distinguish patients with good and poor responses to PD-L1 inhibitor treatment.
[0092] 4. Define Risk
[0093] Define the risk probability values predicted by the predict function. When the risk probability is greater than the median of all risk values, define the sample as high risk ("High"), otherwise, define it as low risk ("Low"). According to the risk definition results of each sample, divide the 80 samples into a high-risk group and a low-risk group.
[0094] 5. Survival Analysis
[0095] Apply the survfit function in the survival package of R language to perform survival analysis on the validation set based on the model prediction results and evaluate the model prediction results.
[0096] The results are as Figure 3 shown. The p-value of the log rank test is 0.012, indicating that there is a significant difference in the survival curves between the high- and low-risk groups, suggesting that there is a significant difference in the recurrence-free survival rate between patients in different risk groups. As Figure 4As shown, in further analysis, ROC curve analysis showed that the AUC value was 0.819, indicating that the PD-L1 inhibitor-targeted therapy prognosis prediction model constructed in Example 1 had good performance and could effectively predict the response of urothelial cancer patients to PD-L1 inhibitor therapy. Through the calculation of the ROC curve and AUC value, the model demonstrated a high degree of discrimination ability, proving its potential value in actual clinical applications.
[0097] Expansion Verification of the PD-L1 Inhibitor-Targeted Therapy Prognosis Prediction Model in Example 3
[0098] Using 98 pan-cancer samples as an expansion validation set, the generalization ability of the PD-L1 inhibitor-targeted therapy prognosis prediction model was further verified to evaluate the application effect of the model in a wider range of cancer types.
[0099] 1. Data Integration and Construction of the Immune Gene Expression Order Model
[0100] Integrate the gene expression data of 98 pan-cancer samples with the 1039 immune genes screened in Example 1, take the intersection, and retain the immune genes present in the expression data and the corresponding expression data to obtain the expression data matrix of the immune genes of the pan-cancer samples.
[0101] According to the expression data matrix of the immune genes of the pan-cancer samples, follow the immune gene expression order algorithm in Example 1 to screen the immune gene pair features most relevant to the clinical response in the expansion validation set samples and construct an immune gene expression order matrix.
[0102] 2. Model Prediction
[0103] Use the predict function in the caret package and the optimal PD-L1 inhibitor-targeted therapy prognosis prediction model trained in Example 1 to predict the expansion validation set, select prob as the prediction type, and obtain the cancer recurrence risk probability prediction values for each sample in the expansion validation set.
[0104] 3. Model Performance Evaluation on the Expansion Validation Set
[0105] Use the roc function of the pROC package to calculate the ROC curve and the area under the curve (AUC) to evaluate the classification performance of the model on the expansion validation set. The AUC value can measure the ability of the model to predict the response to PD-L1 inhibitor therapy in different cancer samples.
[0106] 4. Define Risk
[0107] Define the risk probability values predicted by the predict function. When the risk probability is greater than the median of all risk values, define the sample as high risk ("High"); otherwise, define it as low risk ("Low"). According to the risk definition results of each sample, divide 98 samples into a high-risk group and a low-risk group.
[0108] 5. Survival analysis
[0109] Apply the survfit function in the survival package of R language to perform survival analysis on the extended validation set based on the model prediction results and evaluate the model prediction results.
[0110] The results are as Figure 5 shown. The p-value of the log rank test is 0.0067, indicating that there is a significant difference in the survival curves between the high- and low-risk groups, which means that there is a significant difference in the recurrence-free survival rate between patients in different risk groups. As Figure 6 shown, in further analysis, the ROC curve analysis shows that the AUC value is 0.848, indicating that the PD-L1 inhibitor targeted therapy prognosis prediction model constructed in Example 1 has good performance and can effectively predict the response of pan-cancer samples to PD-L1 inhibitor therapy. Through the calculation of the ROC curve and the AUC value, the model shows a high discrimination ability, demonstrating its potential value in actual clinical applications.
[0111] The above embodiments are only feasible or preferred embodiments of the present invention and are used to illustrate the present invention, not to limit the scope of the patent application of the present invention. Therefore, all equivalent changes and modifications made according to the scope of the patent application of the present invention should fall within the scope covered by the patent of the present invention.
Claims
1. A method for constructing a prognostic prediction model for PD-L1 inhibitor targeted therapy, characterized in that, Construct an expression data matrix of immune genes based on the RNA expression levels of immune genes, apply the immune gene expression sequence algorithm to screen the immune gene pairs most relevant to the clinical response of treating tumors with PD-L1 inhibitors and construct an immune gene expression sequence matrix, and obtain a prognostic prediction model for PD-L1 inhibitor targeted tumor treatment through training; the immune gene expression sequence algorithm includes the following steps: 1) Perform rank transformation on the expression values of immune genes in the expression data matrix of immune genes to obtain an immune gene rank transformation matrix; 2) Calculate the gene rank difference for each sample in the immune gene rank transformation matrix to obtain an immune gene rank difference matrix; 3) Transpose the immune gene rank difference matrix and add the clinical information of each sample to obtain an immune gene rank difference matrix with clinical information; 4) Use the random forest algorithm to perform feature selection on the immune gene rank difference matrix with clinical information to screen immune genes related to the clinical response; 5) Extract the immune gene rank difference matrix according to the immune genes screened in step 4) to obtain a high-feature immune gene rank difference matrix; calculate the rank difference of each pair of immune genes through nested loops to obtain an immune gene pair rank difference matrix, and add the clinical information of each sample to the immune gene pair rank difference matrix to obtain an immune gene pair rank difference matrix with clinical information; 6) Use the random forest algorithm to perform feature selection on the immune gene pair rank difference matrix with clinical information obtained in step 5) to screen the immune gene pairs most relevant to the clinical response and obtain an immune gene expression sequence matrix.
2. The method according to claim 1, characterized in that, The immune genes are screened according to the following method: Collect and organize immune-related genes according to the literature and databases, perform blast alignment with human genes to find all homologous genes, and then take the intersection with the gene expression data of tumor samples to screen out immune genes.
3. The method according to claim 1 or 2, characterized in that, The immune genes include the following 1039 genes: AZGP1, B2M, CALR, CANX, CD1A, CD1B, CD1C, CD1D, CD1E, CD4, CD8A, CD8B, CD74, CREB1, CTSB, CTSE, CTSL1, CTSS, FCER1G, FCGRT, PDIA3, HFE, HLA-A, HLA-B, HLA-C, HLA-DMA, HLA-DMB, HLA-DOA, HLA-DOB, HLA-DPA1, HLA-DPB1, HLA-DQA1, HLA-DQA2, HLA-DQB1, HLA-DRA, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5, HLA-E, HLA-F, HLA-G, HLA-H, MR1, HSPA1B, HSPA1L, HSPA2, HSPA4, HSPA5, HSPA6, HSPA8, HSP90AA1, HSP90AB1, ICAM1, IFNA2, IFNA4, IFNA5, IFNA6, IFNA7, IFNA8, IFNA10, IFNA13, IFNA14, IFNA16, IFNA17, IFNA21, IFNG, KIR2DL1, KIR2DL2, KIR2DL3, KIR2DL4, KIR2DS1, KIR2DS3, KIR2DS4, KIR2DS5, KIR3DL1, KIR3DL2, KLRC1, KLRC2, KLRC3, KLRD1, LTA, CIITA, MICA, MICB, NFYA, NFYB, NFYC, LGMN, PSMB8, PSMC1, PSMC2, PSMC3, PSMC4, PSMC5, PSMC6, PSMD1, PSMD2, PSMD3, PSMD4, PSMD5, PSMD7, PSMD8, PSMD10, PSMD11, PSMD13, PSME1, PSME2, RELB, RFX5, RFXAP, SLC10A2, TAP1, TAP2, TAPBP, THBS1, SHFM1, KLRC4, AP3B1, RFXANK, PSMD6, PSME3, PSMD14, CLEC4M, IFI30, PROCR, ADRM1, KIAA0368, TRPC4AP, CD209, UBXN1, ERAP1, TAPBPL, KIR2DL5A, ERAP2, ULBP3, ULBP2, ULBP1, KIR3DL3, RAET1E, RAET1L, UBR1, RAET1G, PDIA2, CD79A, CD79B, LYN, SYK, BTK, BLNK, VAV3, VAV1, VAV2, RAC1, RAC2, RAC3PPP3CA, PPP3CB, PPP3CC, CHP, PPP3R1, PPP3R2, CHP2, NFAT5, NFATC1, NFATC2, NFATC3, NFATC4, HRAS, KRAS, NRAS, FOS, JUN, CARD11, BCL10, MALT1, CHUK, IKBKB, IKBKG, NFKB1, RELA, NFKBIA, NFKBIB, NFKBIE, CD81, CD19, CR2, PIK3R5, PIK3R1, PIK3R2, PIK3R3, PIK3CA, PIK3CB, PIK3CD, PIK3CG, AKT3, AKT1, AKT2, GSK3B, INPP5D, CD22, CD72, PTPN6, LILRB3, FCGR2B, RASGRP3, PLCG2, PRKCB, IFITM1, IGHA1, IGHA2, IGHD, IGHE, IGHG1, IGHG2, IGHG3, IGHG4, IGHM, IGHV2-70, IGHV3-23, IGKC, IGKV1-5, IGKV2-40, IGKV3-20, IGKV3D-11, IGKV3D-20, IGKV4-1, IGLC3, IGLV7-43, C3, C5, CAMP, CCL1, CCL11, CCL13, CCL14, CCL15, CCL16, CCL17, CCL18, CCL19, CCL2, CCL20, CCL21, CCL22, CCL23, CCL24, CCL25, CCL26, CCL27, CCL28, CCL3, CCL3L3, CCL4, CCL4L1, CCL5, CCL7, CCL8, CKLF, CMA1, CTSG, CX3CL1, CXCL1, CXCL10, CXCL11, CXCL12, CXCL13, CXCL14, CXCL16, CXCL17, CXCL2, CXCL3, CXCL5, CXCL6, CXCL9, CYR61, DEFA1, DEFA3, DEFA5, DEFB1, DEFB104A, DEFB4, EDN1, EDN2, EDN3, FGF10, FGF2, HTN3, IL8, LECT2, PF4, PF4V1, PLAU, PPBP, PROK2, RNASE2, SAA1, SBDS, SEMA3A, SEMA3B, SEMA3C, SEMA3D, SEMA3E, SEMA3F, SEMA3G, SEMA4A, SEMA4B, SEMA4C, SEMA4D, SEMA4F, SEMA4G, SEMA5A, SEMA5B, SEMA6A, SEMA6B, SEMA6C, SEMA6DSEMA7A, SLIT1, SLIT2, TNC, TYMP, XCL1, XCL2, C5AR1, CCBP2, CCR1, CCR10, CCR3, CCR4, CCR5, CCR6, CCR7, CCR8, CCR9, CCRL1, CCRL2, CMKLR1, CX3CR1, CXCR3, CXCR4, CXCR5, CXCR6, CXCR7, CYSLTR1, CYSLTR2, DARC, EDNRA, EDNRB, FPR1, FPR2, GPR17, GPR32, GPR33, GPR44, GPR77, IL8RA, IL8RB, LTB4R, LTB4R2, PLAUR, PLXNA1, PLXNA2, PLXNA3, PLXNA4, PLXNB1, PLXNB2, PLXNB3, PLXNC1, PLXND1, PTAFR, ROBO1, ROBO2, ROBO3, RXFP3, XCR1, ADIPOQ, ADM, ADM2, AGRP, AGT, AMBN, AMELX, AMH, ANGPTL5, ANGPTL7, APLN, AREG, ARMET, ARMETL1, ARTN, AVP, AZU1, BDNF, BMP1, BMP10, BMP15, BMP2, BMP3, BMP4, BMP5, BMP6, BMP7, BMP8A, BMP8B, BTC, C19orf10, CALCA, CALCB, CAT, CCK, CD320, CD40LG, CD70, CECR1, CER1, CGA, CGB, CGB2, CHGA, CHGB, CLCF1, CLEC11A, CMTM1, CMTM2, CMTM3, CMTM4, CMTM5, CMTM6, CMTM7, CMTM8, CNTF, CORT, CRH, CSF1, CSF2, CSF3, CSH1, CSHL1, CSPG5, CTF1, CTGF, DKK1, EBI3, EGF, EPGN, EPO, EREG, ESM1, FAM3B, FAM3C, FAM3D, FASLG, FGF1, FGF11, FGF12, FGF13, FGF14, FGF16, FGF17, FGF18, FGF19, FGF20, FGF21, FGF22, FGF23, FGF3, FGF4, FGF5, FGF6, FGF7, FGF8, FGF9, FIGF, FIGNL2, FLT3LG, FSHB, GAL, GALP, GAST, GCG, GDF1, GDF10, GDF11, GDF15, GDF2, GDF3, GDF5, GDF6, GDF7, GDF9, GDNF, GH1, GH2, GHRHGHRL, GIP, GKN1, GMFB, GMFG, GNRH1, GNRH2, GPHA2, GPHB5, GPI, GREM1, GREM2, GRN, GRP, GUCA2A, HAMP, HBEGF, HDGF, HDGFRP3, HGF, IAPP, IFNB1, IFNE, IFNK, IFNW1, IGF1, IGF2, IL10, IL11, IL12A, IL12B, IL13, IL15, IL16, IL17A, IL17B, IL17C, IL17D, IL17F, IL18, IL19, IL1A, IL1B, IL1F10, IL1F5, IL1F6, IL1F7, IL1F8, IL1F9, IL1RN, IL2, IL20, IL21, IL22, IL23A, IL24, IL25, IL26, IL27, IL28A, IL28B, IL29, IL3, IL31, IL32, IL33, IL34, IL4, IL5, IL6, IL6ST, IL7, IL9, INHA, INHBA, INHBB, INHBC, INHBE, INS, INSL3, INSL4, INSL5, INSL6, JAG1, JAG2, KGFLP1, KGFLP2, KITLG, KL, LACRT, LEFTY1, LEFTY2, LEP, LHB, LIF, LRSAM1, LTB, LTBP1, LTBP2, LTBP3, LTBP4, MDK, MIA, MIF, MLN, MSTN, NAMPT, NDP, NENF, NGF, NMB, NODAL, NOV, NPFF, NPPA, NPPB, NPPC, NPY, NRG1, NRG2, NRG3, NRG4, NRTN, NTF3, NTF4, NTS, NUDT6, OGN, OSGIN1, OSM, OSTN, OXT, P11, PDGFA, PDGFB, PDGFC, PDGFD, PDGFRA, PDGFRB, PDGFRL, PDYN, PENK, PGF, PMCH, PNOC, POMC, PPY, PRL, PRLH, PROK1, PSPN, PTH, PTH2, PTHLH, PTN, PYY, QRFP, RABEP1, RABEP2, REG1A, RETN, RETNLB, RLN1, RLN2, RLN3, S100A6, SCG2, SCGB3A1, SCT, SCYE1, SECTM1, SLURP1, SPP1, SST, STC1, STC2, TAC1, TDGF1, TDGF3, TG, TGFA, TGFB1, TGFB2, TGFB3, THPO, TNF, TNFRSF11B, TNFSF10TNFSF11, TNFSF12, TNFSF13, TNFSF13B, TNFSF14, TNFSF15, TNFSF18, TNFSF4, TNFSF8, TNFSF9, TOR2A, TRH, TSHB, TSLP, TXLNA, UCN, UCN2, UCN3, UTS2, UTS2D, VEGFA, VEGFB, VEGFC, VGF, VIP, ACVR1B, ACVR1C, ACVR2A, ACVR2B, ACVRL1, ADCYAP1R1, ADIPOR1, ADIPOR2, ADRB1, ADRB2, AGTR1, AGTR2, AMHR2, ANGPT1, ANGPT4, ANGPTL1, ANGPTL2, ANGPTL3, ANGPTL4, ANGPTL6, APLNR, AR, AVPR1A, AVPR1B, AVPR2, BMPR1A, BMPR1B, BMPR2, BRD8, C3AR1, CALCR, CALCRL, CD40, CNTFR, CRHR1, CRHR2, CRIM1, CRLF1, CRLF2, CRLF3, CSF1R, CSF2RA, CSF2RB, CSF3R, EGFR, ENG, EPOR, ESR1, ESR2, ESRRA, ESRRB, ESRRG, FGFR1, FGFR2, FGFR3, FGFR4, FGFRL1, FLT1, FLT3, FLT4, FSHR, GALR2, GALR3, GCGR, GHR, GHRHR, GHSR, GIPR, GLP1R, GLP2R, GNRHR, GPER, HNF4A, HNF4G, HTR3A, HTR3B, HTR3C, HTR3D, HTR3E, IFNAR1, IFNAR2, IFNGR1, IFNGR2, IGF1R, IGF2R, IL10RA, IL10RB, IL11RA, IL12RB1, IL12RB2, IL13RA1, IL13RA2, IL15RA, IL17RA, IL17RB, IL17RC, IL17RD, IL17RE, IL18R1, IL18RAP, IL1R1, IL1R2, IL1RAP, IL1RL1, IL1RL2, IL20RA, IL20RB, IL21R, IL22RA1, IL22RA2, IL23R, IL27RA, IL28RA, IL2RA, IL2RB, IL2RG, IL31RA, IL3RA, IL4R, IL5RA, IL6R, IL7R, IL9R, INSR, KDR, LEPR, LGR4, LGR5, LGR6, LHCGR, LIFR, LTBR, MC1R, MC2R, MC3RMC4R, MCHR1, MCHR2, MET, MLNR, MPL, MTNR1A, MTNR1B, NGFR, NMBR, NPR1, NPR3, NR0B1, NR0B2, NR1D1, NR1D2, NR1H2, NR1H3, NR1H4, NR1I2, NR1I3, NR2C1, NR2C2, NR2E1, NR2E3, NR2F1, NR2F2, NR2F6, NR3C1, NR3C2, NR4A1, NR4A2, NR4A3, NR5A1, NR5A2, NR6A1, NRP1, NRP2, OGFR, OPRD1, OPRK1, OPRL1, OPRM1, OSMR, OXTR, PGR, PGRMC2, PPARA, PPARD, PPARG, PRLHR, PRLR, PTGDR, PTGDS, PTGER1, PTGER2, PTGER3, PTGER4, PTGFR, PTH1R, PTH2R, RARA, RARB, RARG, RORA, RORB, RORC, RXFP1, RXFP2, RXRA, RXRB, RXRG, S1PR1, S1PR2, SCTR, SDC1, SDC2, SDC3, SDC4, SORT1, SSTR1, SSTR2, SSTR5, TACR1, TEK, TGFBR1, TGFBR2, TGFBR3, THRA, THRB, TIE1, TNFRSF10A, TNFRSF10B, TNFRSF10C, TNFRSF10D, TNFRSF11A, TNFRSF12A, TNFRSF13B, TNFRSF13C, TNFRSF14, TNFRSF17, TNFRSF18, TNFRSF19, TNFRSF1A, TNFRSF1B, TNFRSF21, TNFRSF25, TNFRSF4, TNFRSF6B, TNFRSF8, TNFRSF9, TRHR, TSHR, TUBB3, VDR, VIPR1, VIPR2, PTPN11, ICAM2, ITGAL, ITGB2, PTK2B, PAK1, MAP2K1, MAP2K2, MAPK1, MAPK3, NCR2, TYROBP, LCK, FCGR3A, FCGR3B, NCR1, NCR3, CD247, ZAP70, LCP2, LAT, PLCG1, SH3BP2, FYN, SHC2, SHC4, SHC3, SHC1, GRB2, SOS1, SOS2, ARAF, BRAF, RAF1, KLRK1, HCST, CD48, CD244, PRKCA, PRKCG, SH2D1B, SH2D1A, FAS, GZMB, PRF1, CASP3, BIDCD3D, CD3E, CD3G, PTPRC, ITK, TEC, NCK1, NCK2, GRAP2, PAK2, PAK3, PAK4, PAK6, PAK7, RHOA, CDC42, CD28, ICOS, MAP3K8, MAP3K14, PDCD1, CTLA4, CBLC, CBL, CBLB, CDK4, RASGRP1, PDK1, PRKCQ, TRBC1, TRBV12-3, TRGV3.
4. The method according to claim 1, characterized in that The construction method of the expression data matrix of the immune genes includes the following steps: Integrate the gene expression data of the training set samples with the immune genes, take the intersection, retain the immune genes existing in the gene expression data and the corresponding expression data, and generate an expression data matrix of the immune genes. The columns of the matrix are the training set samples, the rows are the immune genes, and the data in the matrix are the expression values of the immune genes.
5. The method according to claim 1, wherein In step 1), the rank transformation formula is: Rank(X a ) = floor(rank(-X a )) Among them, X a is the expression value of the immune genes in the ath column, and Rank(X a ) is the rank of the immune genes in the ath column.
6. The method according to claim 1, characterized in that, In step 2), the gene rank difference calculation formula is: Difference(Y ij ) = Y ij - median(Y i ) Among them, Y ij is the rank of the i-th row and j-th column, median(Y i ) is the median rank of the i-th row, Difference(Y ij ) is the difference between the rank of the immune gene in the i-th row and j-th column and the median rank of the i-th row.
7. The method according to claim 1, characterized in that, In step 5), the immune gene pair rank difference matrix is in binary form. When the rank difference of one gene is greater than or equal to the rank difference of another gene, it is a positive difference, and the corresponding position is marked as 1; otherwise, it is a negative difference, and the corresponding position is marked as -1.
8. The method according to claim 1, characterized in that The clinical information includes recurrence-free survival time, recurrence-free survival status, and PD-L1 inhibitor treatment response data.
9. The method according to claim 1, characterized in that The immune gene pairs most relevant to clinical response include: "CASP3>CCL24", "CASP3>FAS", "CASP3>RAC2", "CASP3>TNFRSF10A", "CASP3>TNFRSF10B", "CASP3>TNFRSF10D", "CCL24>GRB2", "CCL24>KIR2DL4", "CCL24>KIR3DL1", "CCL24>KIR3DL2", "CCL24>KLRC3", "CCL24>KLRC4", "CCL24>NFKBIB", "CCL24>PSMD3", "CCL24>PSMD4", "CCL24>PSMD7", "CCL24>PSMD8", "CCL24>RAC1", "CCL24>TRH", "CD8A>FAS", "FAS>IFNG", "FAS>KIR2DL4", "FAS>KLRC2", "FAS>KLRC3", "FAS>TRH", "GRB2>PSMD3", "IFNG>KIR3DL2", "IFNG>NFKB1", "IFNG>NFKBIB", "IFNG>PSMD8", "IFNG>RAC2", "IFNG>TNFRSF14", "KIR2DL4>KLRK1", "KIR2DL4>NFKB1", "KIR2DL4>TNFRSF10B", "KIR2DL4>TNFRSF11B", "KIR2DL4>TNFRSF14", "KIR3DL1>PSMD3", "KIR3DL1>PSMD4", "KIR3DL1>PSMD8", "KIR3DL1>TNFRSF10A", "KLRC1>TNFRSF10D", "KLRC2>KLRK1", "KLRC2>NFKB1", "KLRC2>NFKBIB", "KLRC2>PSMD4", "KLRC2>PSMD7", "KLRC2>PSMD8", "KLRC2>TNFRSF10A", "KLRC2>TNFRSF10B", "KLRC2>TNFRSF10D", "KLRC2>TNFRSF14", "KLRC2>TNFRSF1A", "KLRC3>KLRK1", "KLRC3>NFKB1", "KLRC3>TNFRSF14", "KLRC3>TNFRSF1A", "KLRC4>NFKB1", "KLRC4>PIK3R1", "KLRC4>RAC2", "KLRC4>TNFRSF10A", "KLRC4>TNFRSF10D", "KLRC4>TNFRSF11B", "KLRC4>TNFRSF14", "KLRK1>TNFRSF10A", "KLRK1>TNFRSF11B", "NFKB1>PSMD3", "NFKB1>TNFRSF10D", "NFKBIB>TNFRSF10B", "NFKBIB>TNFRSF1A", "PIK3R1>RAC1", "PIK3R1>TRH", "PSMD3>TNFRSF10D", "PSMD3>TNFRSF1A", "PSMD4>TNFRSF1A", "PSMD7>TNFRSF10B", "PSMD7>TNFRSF10D", "PSMD7>TNFRSF1A", "PSMD8>TNFRSF10D", "PSMD8>TNFRSF1A", "RAC1>TNFRSF10A", "RAC1>TNFRSF10B", "RAC1>TNFRSF10D", "RAC1>TNFRSF11B", "RAC1>VAV2", "RAC3>TNFRSF10B", "RAC3>TNFRSF10D", and "TNFRSF10D>TRH".
10. The method according to claim 1, characterized in that, The training of the prognostic prediction model for PD-L1 inhibitor targeted therapy uses the random forest algorithm, and the steps include: 1) Use the trainControl function to configure the model training parameters, select the bootstrap method as the resampling method, enable the calculation of class probabilities, and use the binary classification summary function as the performance evaluation index; 2) Perform loop iteration. In each iteration, use the createDataPartition function to add the PD-L1 inhibitor treatment response data to the training; apply the train function to train the random forest model, and use the predict function with the prediction type of probability to predict the probability of tumor recurrence risk; 3) Calculate the ROC curve and AUC value for each iteration. When the AUC value exceeds the target value, terminate the loop to obtain the optimal prognostic prediction model for PD-L1 inhibitor targeted therapy.
11. A PD-L1 inhibitor targeted therapy prognostic prediction model product constructed by the method according to any one of claims 1-10.
12. The application method of the model according to claim 11 in the prognostic prediction of tumor PD-L1 inhibitor targeted therapy.
Citation Information
Patent Citations
Biomarker system screening method based on multi-omics
CN108830045A
Gene pair screening method and device, computer equipment and storage medium
CN111383716A