Methods for predicting and selecting neoantigens
By performing whole-exome sequencing and RNA sequencing on tumor samples from cancer patients, combined with a neoantigen classifier, the immunotherapy responders in cancer patients can be accurately identified. This addresses the shortcomings of the existing TMB method and improves the predictive accuracy and efficacy of immunotherapy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2026-04-10
AI Technical Summary
Existing tumor mutational burden (TMB)-based methods are not precise enough in identifying immunotherapy responders in cancer patients and cannot effectively assess neoantigen presentation, resulting in poor immunotherapy efficacy.
By performing whole-exome sequencing and RNA sequencing on tumor samples from patients, the quantity and abundance of neoantigens, unique T and B cell receptors, and enrichment of immune cell populations are quantified. Combined with MHC allele typing and neoantigen classifiers, cancer patients can be identified as immunotherapy responders.
It improves the accuracy of identifying immunotherapy responders in cancer patients, with positive and negative predictive values at least 10%-50% higher than traditional methods, thus enhancing the specificity and effectiveness of immunotherapy.
Smart Images

Figure CN121844384A_ABST
Abstract
Description
[0001] This patent application claims the benefit of U.S. Provisional Patent Application Serial No. 63 / 661,737, filed June 19, 2024, which is incorporated in its entirety by reference for all purposes. Background Technology
[0002] Tumor mutational burden (TMB), reflecting the number of cancer mutations, has become a predictive biomarker for immunotherapy. While a high TMB status leads to increased neoantigen presentation and T-cell recognition, not all mutations generate neoantigens or induce an immune response, making TMB an imperfect biomarker. Therefore, there is a need for an improved method to identify subgroups of cancer patients who are immunotherapy responders using a more effective method for assessing neoantigen presentation. Such a method would also enable the creation of more effective therapies based on the presented neoantigens. Summary of the Invention
[0003] In one aspect, this disclosure relates to a method for identifying cancer patients as immunotherapy responders, the method comprising: performing whole-exome sequencing and whole-genome sequencing on a tumor sample from the patient to quantify the number of neoantigens in the tumor sample; performing RNA sequencing (such as whole-transcriptome RNA sequencing or targeted T and B cell receptor sequencing using extracted DNA) on the tumor sample from the patient to quantify the number of unique T and B cell receptors and the enrichment of immune cell populations in the tumor sample; and identifying the cancer patient as an immunotherapy responder using the number of neoantigens in the tumor sample, the number of unique T and B cell receptors, the abundance of each unique T and B cell receptor, and the enrichment of immune cell populations in the tumor sample. In some instances, a custom kit of approximately 100 to 500 cancer genes is used.
[0004] In some embodiments, the method further includes whole-exome sequencing of germline and / or tumor samples from the patient to genotype the patient’s MHC I and MHC II alleles.
[0005] In some embodiments, quantifying neoantigens in tumor samples includes (i) genotyping the patient’s MHC I and MHC II alleles by germline and / or tumor whole-exome sequencing; (ii) identifying somatic mutations in the patient’s tumor sample that cause changes in protein sequence and filtering out somatic mutations from genes that are never expressed based on RNA sequencing of the tumor sample; and (iii) pairing each of the patient’s MHC I and MHC II alleles obtained in (if) with each peptide of length 8-12 (MHC I) or 10-30 (MHC II) amino acids included in at least one somatic mutation obtained in (ii), and identifying one or more neoantigens based on MHC-peptide binding and T-cell activation.
[0006] In some embodiments, somatic mutations include single nucleotide variants (SNVs), multiple nucleotide variants (MNVs), copy number variants (CNVs), insertions or deletions, gene fusions, structural variants, or combinations thereof.
[0007] In some embodiments, one or more neoantigen classifiers are used to identify neoantigens.
[0008] In some embodiments, quantifying the number and abundance of unique T and B cell receptors includes: (i) deconvolving the proportion of immune cells in a tumor sample based on RNA sequencing data, and (ii) assembling B and T cell receptors to quantify the number of unique T and B cell receptors.
[0009] In some embodiments, the enrichment of immune cell populations in tumor samples is determined by tumor gene expression based on RNA sequencing data.
[0010] In some embodiments, the patient's tumor sample is derived from a solid tumor.
[0011] In some embodiments, the method identifies cancer patients as immunotherapy responders with a positive predictive value (PPV) at least 10%, at least 20%, at least 30%, at least 40%, or at least 50% higher than using TMB alone. In some embodiments, the method identifies cancer patients as immunotherapy non-responders with a negative predictive value (NPV) at least 10%, at least 20%, at least 30%, at least 40%, or at least 50% higher than using TMB alone.
[0012] In some embodiments, the method identifies cancer patients as immunotherapy responders with a sensitivity at least 10%, at least 20%, at least 30%, at least 40%, or at least 50% higher than using TMB alone. In some embodiments, the method identifies cancer patients as immunotherapy non-responders with a specificity at least 10%, at least 20%, at least 30%, at least 40%, or at least 50% higher than using TMB alone.
[0013] On the other hand, this disclosure relates to a method for treating cancer, the method comprising administering treatment to a cancer patient who has been identified as an immunotherapy responder by the method described herein.
[0014] In some embodiments, treatment includes checkpoint inhibitors, CAR-T therapy, TCR-T therapy, NK cell therapy, cancer vaccines, oncolytic viruses, cytokines, monoclonal antibodies, or combinations thereof.
[0015] In some embodiments, immunotherapy includes PD1 inhibitors, PD-L1 inhibitors, CTLA-4 inhibitors, or combinations thereof.
[0016] In some embodiments, immunotherapy includes pembrolizumab, nivolumab, cimiplimab, dostarlimab, atezolizumab, avelumab, durvalumab, ipilimumab, or tremelimumab. In some embodiments, immunotherapy includes vopratelimab, spartalizumab, camrelizumab, sintilimab, tislelizumab, toripalimab, INCMGA00012, AMP-224, AMP-514, KN035, cosibelimab, AUNP12, CA-170, or BMS-986189.
[0017] In some embodiments, the cancer patient has been treated with surgery, chemotherapy, or radiation therapy, either or simultaneously.
[0018] In some embodiments, the cancer is breast cancer, colorectal cancer, gastrointestinal cancer, kidney cancer, lung cancer, bladder cancer, ovarian cancer, or pancreatic cancer.
[0019] In some embodiments, cancer is a cancer or tumor of the abdomen or abdominal wall, adrenal gland, anus, appendix, bladder, bone, brain, breast, cervix, chest wall, colon, diaphragm, duodenum, ear, endometrium, esophagus, fallopian tube, gallbladder, gastroesophageal junction, head and neck, kidney, larynx, liver, lung, lymph nodes, malignant effusion, mediastinum, nasal cavity, omentum, ovary, pancreas, pancreaticobiliary duct, parotid gland, pelvis, penis, pericardium, peritoneum, pleura, prostate, rectum, salivary gland, skin, small intestine, soft tissue, spleen, stomach, thyroid gland, tongue, trachea, ureter, uterus, vagina, vulva, or Whipple's resection site.
[0020] In another aspect, this disclosure relates to a method for identifying cancer patients as responders to PD-1, PD-L1, or CTLA-4 immunotherapy, the method comprising: performing whole-exome sequencing on a tumor sample from the patient to quantify the number of neoantigens in the tumor sample; performing RNA sequencing on the tumor sample from the patient to quantify the number of unique T and B cell receptors, enrichment of immune cell populations, and expression of PD-1, PD-L1, or CTLA-4 in the tumor sample; and identifying the cancer patient as a responder to PD-1, PD-L1, or CTLA-4 immunotherapy using the number of neoantigens in the tumor sample, the number and abundance of unique T and B cell receptors in the tumor sample, the enrichment of immune cell populations, and the expression of PD-1, PD-L1, or CTLA-4.
[0021] In another aspect, this disclosure relates to a method for treating cancer, the method comprising administering immunotherapy to a cancer patient who has been identified as a PD-1 or PD-L1 or CTLA-4 immunotherapy responder by means of the methods described herein, wherein the immunotherapy comprises pembrolizumab, nivolumab, simiprelimab, dostalimab, atezolizumab, avelumab, durvalumab, ipilimumab, or trimemumab.
[0022] In other embodiments, a method is provided for generating a library of approximately 10-100 personalized cancer vaccines. The personalized cancer vaccines are matched to patients. A predictive model is developed using clinical genomics data. Patient clusters with a certain type of cancer are identified. In some instances, these clusters tend to share one or more common neoantigens. In some instances, the predictive model utilizes standard clustering algorithms. In other instances, the phylogenetic or clonal evolution of tumors is used to identify patient clusters with a certain type of cancer.
[0023] The vaccine is constructed for each patient cluster. In some instances, the vaccine covers the most immunogenic neoantigen at the centroid of each cluster. Patients (such as new patients with a certain type of cancer) can then be matched with vaccines that have been constructed by identifying which cluster the patient belongs to and / or predicting the patient's immune response score to each available vaccine. At least one computing device includes at least one processor configured to: In some embodiments, systems, methods, and non-transitory computer-readable media are provided for outputting a catalog of one or more somatic mutations. Data from subjects is fed into a predictive model, and neoantigens appearing in a subset of data from the subjects are scored against one or more parameters. Based on the scores against one or more parameters, one or more neoantigens that meet an immunostimulation threshold appearing in the subset of subjects are identified. Somatic mutations that meet the immunostimulation threshold for one or more neoantigens appearing in the subset of subjects are identified. A predictive catalog of one or more of the immunostimulation threshold-meeting somatic mutations appearing in the subset of subjects is output, which predicts somatic mutations.
[0024] In some respects, the data comes from subjects who have previously received treatment for a certain type of cancer, and the predictive model is a machine learning model. In other respects, the predictive model is better at predicting immune stimulation for a subset of subjects than tumor mutational burden (TMB) status. Parameters may further include one or more of the following: immune checkpoint inhibitor (ICI) response, ctDNA outcome, age, sex, and ECOG score.
[0025] In some embodiments, systems, methods, and non-transitory computer-readable media are provided for outputting one or more vaccines to be administered to a subject. Data from the subject is fed into a trained model. Neoantigens appearing in the data for the subject are scored against one or more parameters. Based on the scores against one or more parameters, one or more neoantigens that meet an immune stimulation threshold for the subject are identified. Somatic mutations of one or more neoantigens that meet an immune stimulation threshold for the subject are identified. One or more vaccines to be administered to the subject are output for the somatic mutations of one or more neoantigens that meet an immune stimulation threshold. The one or more vaccines to be administered to the subject are selected from a predicted catalog of somatic mutations appearing in a subset of subjects who have previously been treated for a certain type of cancer.
[0026] In each respect, one or more vaccines are selected from a library of pre-prepared vaccines to be administered. In other respects, the vaccine is one or more of the following: a peptide-based synthetic vaccine, a messenger RNA (mRNA) vaccine, or a conventional vaccine.
[0027] In other embodiments, a vaccine composition is provided. This vaccine composition is prepared by a method comprising the steps of: feeding data from a subject suffering from a certain type of cancer into a predictive model; scoring neoantigens appearing in the data for the subject against one or more parameters; determining one or more neoantigens appearing in the subject that satisfy an immune stimulation threshold based on the scores against the one or more parameters; determining somatic mutations of the one or more neoantigens appearing in the subject that satisfy the immune stimulation threshold; and preparing one or more vaccine compositions to be administered to the subject against one or more somatic mutations of the one or more neoantigens that satisfy the immune stimulation threshold, wherein the one or more vaccines to be administered to the subject are selected from a predictive catalog of somatic mutations appearing in a subset of subjects previously treated for that type of cancer.
[0028] In various aspects, the one or more parameters include a combination of one or more of the following: peptide processing and presentation, RNA expression, MHC binding fold change, T cell activation, and dissimilarity to a reference human proteome. In other aspects, one or more vaccine compositions are selected from a library of pre-prepared vaccine compositions to be administered. One or more vaccine compositions can be one or more of the following: peptide-based synthetic vaccines, messenger RNA (mRNA) vaccines, or conventional vaccines. In other aspects, one or more neoantigens are combined with liquid nanoparticles (LNPs) to prepare one or more vaccine compositions. Attached Figure Description
[0029] The currently disclosed embodiments will be further explained with reference to the accompanying drawings, in which similar numbers in several views denote similar structures. The drawings shown are not necessarily to scale, but rather focus on illustrating the principles of the currently disclosed embodiments.
[0030] Figure 1A This is a flowchart illustrating the extraction of clinical data based on the examples described in this article.
[0031] Figure 1B This is a block diagram of an exemplary immune response prediction model based on the examples described in this article.
[0032] Figures 2A to 2D This is a graphical representation of the immune response of an individual with melanoma, based on the examples described in this article.
[0033] Figures 3A to 3D This is a graphical representation of the immune response in an individual with NSCLC based on the examples described in this article.
[0034] Figures 4A to 4D This is a graphical representation of the immune response in an individual with MSI-H CRC based on the examples described in this article.
[0035] Figures 5A to 5D This is a graphical representation of the immune response of an individual with a solid tumor, based on the examples described in this article.
[0036] Figure 6 This is a block diagram illustrating a computing system based on the examples described in this article.
[0037] Figure 7 This is a block diagram illustrating the prediction engine based on the example described in this article.
[0038] Figure 8 This is a block diagram of a bioinformatics pipeline based on the examples described in this article.
[0039] Figure 9 This is a flowchart illustrating a method for training an immune stimulation prediction model based on the examples described in this paper.
[0040] Figure 10 This is a flowchart illustrating a method for generating immune stimulation and predicting vaccines, based on the examples described in this article.
[0041] Figure 11 This is a block diagram of an exemplary bioinformatics pipeline based on the examples described in this article.
[0042] Figure 12 This is a graphical representation of the neoantigen priority ranking performance based on the examples described in this article.
[0043] Figure 13 This is a graphical representation of the prediction of true positive neoantigens based on the examples described in this article.
[0044] Figure 14 This is a graphical representation of the neoantigen prediction and prioritization method based on the examples described in this paper, compared to other methods.
[0045] Figure 15 This is a graphical representation of RNAseq data used with neoantigen prediction and prioritization methods, based on the examples described in this article.
[0046] Figure 16 This is a graphical representation of the neoantigen prediction and prioritization method described in this paper compared to TMB.
[0047] Figures 17A to 17H This is a graphical representation of the ICI response and neoantigen prediction and prioritization methods based on the examples described in this article. Detailed Implementation
[0048] General Overview The methods and compositions provided herein improve immunotherapy for treating cancer. In one aspect, this disclosure relates to a method for identifying cancer patients as immunotherapy responders, the method comprising: performing whole-exome sequencing, whole-genome sequencing, or a custom kit on a tumor sample from the patient to quantify the number of neoantigens in the tumor sample; performing RNA sequencing on the tumor sample from the patient to quantify the number of unique T and B cell receptors and the enrichment of immune cell populations in the tumor sample; and identifying the cancer patient as an immunotherapy responder using the number of neoantigens in the tumor sample, the number and abundance of unique T and B cell receptors, and the enrichment of immune cell populations in the tumor sample.
[0049] On the other hand, this disclosure relates to a method for treating cancer, the method comprising administering immunotherapy to a cancer patient who has been identified as an immunotherapy responder by the method described herein.
[0050] In another aspect, this disclosure relates to a method for identifying cancer patients as responders to PD-1, PD-L1, or CTLA-4 immunotherapy, the method comprising: performing whole-exome sequencing, whole-genome sequencing, or custom kits of cancer genes on a tumor sample from the patient to quantify the number of neoantigens in the tumor sample; performing RNA sequencing on the tumor sample from the patient to quantify the number of unique T and B cell receptors, enrichment of immune cell populations, and expression of PD-1, PD-L1, or CTLA-4 in the tumor sample; and using the number of neoantigens in the tumor sample, the number and abundance of unique T and B cell receptors in the tumor sample, the enrichment of immune cell populations, and the expression of PD-1, PD-L1, or CTLA-4 to identify the cancer patient as a responder to PD-1, PD-L1, or CTLA-4 immunotherapy.
[0051] In another aspect, this disclosure relates to a method for treating cancer, the method comprising administering immunotherapy to a cancer patient who has been identified as a PD-1 or PD-L1 or CTLA-4 immunotherapy responder by means of the methods described herein, wherein the immunotherapy comprises pembrolizumab, nivolumab, simiprelimab, dostalimab, atezolizumab, avelumab, durvalumab, ipilimumab, or trimemumab.
[0052] Sample collection The methods disclosed herein improve immunotherapy for treating a variety of cancers in patients. Those skilled in the art will understand that different types of cancer will require the collection of different types of samples as described herein.
[0053] In some embodiments, the cancer is a solid tumor, and the biological sample is a tumor biopsy sample. Performing a biopsy typically involves using a sharp instrument to remove a small amount of tissue from a patient suspected of containing diseased cells or tissue, such as a tumor. There are many different types of biopsies, such as needle biopsy, CT-guided biopsy, ultrasound-guided biopsy, bone biopsy, bone marrow biopsy, liver biopsy, kidney biopsy, aspiration biopsy, prostate biopsy, skin biopsy, and surgical biopsy (such as laparoscopic biopsy). In some embodiments, the biological sample is obtained via liquid biopsy. In some embodiments, the biological sample is a blood, serum, plasma, or urine sample. Furthermore, biological liquid samples can be extracted from a variety of animal fluids containing cell-free DNA, including but not limited to blood, serum, plasma, bone marrow, urine vitreous humor, sputum, tears, sweat, saliva, semen, mucosal excretions, mucus, cerebrospinal fluid, amniotic fluid, lymph, etc. Cell-free DNA can be of fetal origin (via fluid taken from a pregnant subject) or can be derived from the subject's own tissue.
[0054] In some embodiments, the cancer is a leukemia, and the biological sample is a liquid sample. In some embodiments, the cancer is a leukemia, and the biological sample is a blood, serum, plasma, or bone marrow sample. In some embodiments, both the cancer DNA and the matching normal DNA are obtained from the blood sample by separating and dissecting the plasma and erythrocyte sedimentation rate (ESR) tannin layer. The DNA obtained from the ESR tannin layer can serve as normal DNA that matches the circulating tumor DNA obtained from the plasma fraction.
[0055] In some embodiments, the method of this disclosure further includes longitudinally collecting multiple liquid biopsy samples from a patient. In some embodiments, the liquid biopsy samples are obtained from the patient after the patient has received cancer-specific treatment. In some embodiments, the liquid biopsy samples are blood, serum, plasma, or urine samples.
[0056] Methods for identifying immunotherapy responders This disclosure relates to a method for identifying cancer patients as immunotherapy responders, the method comprising: performing whole-exome sequencing, whole-genome sequencing, or custom kit sequencing on a tumor sample from the patient to quantify the number of neoantigens in the tumor sample; performing RNA sequencing on the tumor sample from the patient to quantify the number and abundance of unique T and B cell receptors and the enrichment of immune cell populations in the tumor sample; and identifying the cancer patient as an immunotherapy responder using the number of neoantigens in the tumor sample, the number of unique T and B cell receptors, and the enrichment of immune cell populations in the tumor sample.
[0057] In some embodiments, the method further includes whole-exome sequencing of a germline sample from the patient to genotype the patient's MHC I and / or MHC II alleles. In some embodiments, whole-exome sequencing, whole-genome sequencing, or custom kit sequencing is performed on cellular DNA obtained from solid tumors and from matched normal tissues (such as erythrocyte sedimentation rate saturates). By comparing sequencing data of DNA obtained from tumor samples with DNA obtained from matched normal tissues, neoantigens can be identified and used to determine whether a cancer patient is a responder to immunotherapy.
[0058] In some embodiments, the term “whole exome sequencing” refers to sequencing all protein-coding regions (also known as the exome) of genes in the genome. Whole exome sequencing of tumor biopsy samples is described, for example, in WO2015 / 164432 and WO2019 / 200228, which are incorporated herein by reference in their entirety.
[0059] In some embodiments, whole-exome sequencing may first involve the step of isolating a subset of DNA encoding proteins (called exons) prior to sequencing. This first step can be performed using capture techniques for the isolated exons, namely array-based capture or in-solution capture as described elsewhere herein. Target enrichment methods allow for the selective capture of genomic regions of interest from a DNA sample prior to sequencing using enrichment methods such as hybridization capture or targeted amplification. The genomic regions of interest may include all exon regions of the genome to prepare a sample for whole-exome sequencing (WES).
[0060] In some embodiments, quantifying neoantigens in tumor samples includes (i) genotyping the patient’s MHC I and MHC II alleles by germline whole-exome sequencing (or other types of seq) or sequencing of the tumor sample; (ii) identifying somatic mutations in the patient’s tumor sample that cause changes in protein sequence and filtering out somatic mutations from genes that are never expressed based on RNA sequencing of the tumor sample; and (iii) pairing each of the patient’s MHC I and MHC II alleles obtained in (i) with each peptide of length 8-12 (MHC I) or 10-30 (MHC II) amino acids included in at least one somatic mutation obtained in (ii) and identifying one or more neoantigens based on MHC-peptide binding and T-cell activation.
[0061] In some embodiments, somatic mutations identified in tumor samples include single nucleotide variants (SNVs), multinucleotide variants (MNVs), copy number variants (CNVs), insertions / deletions, gene fusions, structural variants, aberrant splicing variants, or combinations thereof. The term "insertion / deletion" refers to both the insertion and deletion of nucleic acids in the genome. The term "gene fusion" refers to any genomic alteration caused by the insertion and / or deletion of DNA in the genome, resulting in the fusion of two different genomic loci. The term "structural variant" refers to genomic alterations involving DNA segments larger than 1 kilobase (kb), such as deletions or insertions, and can be microscopic or submicroscopic.
[0062] In some embodiments, the somatic mutations identified in the tumor sample are protein-coding mutations. In some embodiments, the protein-coding mutations include one or more oncogenes, tumor suppressor genes, genes that enhance or inhibit cell proliferation, invasion or metastasis, genes that promote or inhibit apoptosis, and pro-angiogenic or anti-angiogenic genes. In some embodiments, the protein-coding mutations include AKT1 (14q32.33, ALK (2p23.2-23.1), APC (5q22.2), AR (Xq12), ARAF (Xp11.3), ARID1A (1p36.11), ATM (11q22.3), BRAF (7q34), BRCA1 (17q21.31), BRCA2 (13q13.1), CCND1 (11q13.3), CCND2 (12p13.32), CCNE1 (19q12), CDH1 (16q22.1), CDK4 (12q14.1), CDK6 (7q21.2), CDKN2A (9p21.3), CTNNB1 (3p22.1), DDR2 (1q23.3), EGFR (7p11.2), ERBB2 (17q12), and ESR1. (6q25.1-25.2), EZH2 (7q36.1), FBXW7 (4q31.3), FGFR1 (8p11.23), FGFR2 (10q26.13), FGFR3 (4p16.3), GATA3 (10p14), GNA11 (19p13.3), GNAQ (9q21.2), GNAS (20q13.32), HNF1A (12q24.31), HRAS (11p15.5), IDH1 (2q34), IDH2 (15q26.1), JAK2 (9p24.1), JAK3 (19p13.11), KIT (4q12), KRAS (12p12.1), MAP2K1 (15q22.31), MAP2K2 (19p13.3), MAPK1 (22q11.22), MAPK3 (16p11.2), MET (7q31.2), MLH1 (3p22.2), MPL (1p34.2), MTOR (1p36.22), MYC(8q24.21), NF1 (17q11.2), NFE2L2 (2q31.2), NOTCH1 (9q34.3), NPM1 (5q35.1), NRAS(1p13.2), NTRK1 (1q23.1), NTRK3 (15q25.3), PDGFRA (4q12), PIK3CA (3q26.32), PTEN(10q23.31), PTPN11(12q24.13), RAF1 (3p25.2), RB1 (13q14.2), RET (10q11.21), RHEB (7q36.1), RHOA (3p21.31), RIT1 (1q22), ROS1 (6q22.1), SMAD4 (18q21.2), SMO (7q32.1), STK11 (19p13.3), TERT (5p15.33), TP53 (17p13.1), TSC1 (9q34.13) and / or VHL (3p25.3). In some embodiments, the protein-coding mutation is located in the exon regions of one or more of the following genes: ABL1, ACVR1B, AKT1, AKT2, AKT3, ALK, ALOX12B, AMER1 (FAM123B), APC, AR, ARAF, ARFRP1, ARID1A, ASXL1, ATM, ATR, ATRX, AURKA, AURKB, AXIN1, AXL, BAP1, BARD1, BCL2, BCL2L1, BCL2L2, BCL6, BCOR, BCORL1, BRAF, BRCA1, BRCA2, BRD4, BRIP1, BTG1, BTG2, BTK, C11orf30 (EMSY), CALR, CARD11, CASP8, CBFB, CBL, CCND1, CCND2, CCND3, CCNE1, CD22, CD274 (PD-L1), CD70, CD79A, CD79B, CDC73, CDH1, CDK12, CDK4, CDK6, CDK8, CDKN1A, CDKN1B, CDKN2A. CDKN2BCDKN2C CEBPA CHEK1 CHEK2 CIC CREBBP CRKL CSF1R CSF3R CTCF CTNNA1 CTNNB1 CUL3CUL4A CXCR4 CYP17A1 DAXX DDR1 DDR2 DIS3 DNMT3A DOT1L EED EGFR EP300 EPHA3EPHB1 EPHB4 ERBB2 ERBB3 ERBB4 ERCC4 ERG ERRFI1 ESR1 EZH2 FAM46C FANCA FANCCFANCG FANCL FAS FBXW7 FGF10 FGF12 FGF14 FGF19 FGF23 FGF3 FGF4 FGF6 FGFR1FGFR2 FGFR3 FGFR4 FH FLCN FLT1 FLT3 FOXL2 FUBP1 GABRA6 GATA3 GATA4 GATA6 GID4(C17orf39)GNA11 GNA13 GNAQ GNAS GRM3 GSK3B H3F3A HDAC1 HGF HNF1A HRAS HSD3B1ID3 IDH1 IDH2 IGF1R IKBKE IKZF1 INPP4B IRF2 IRF4 IRS2 JAK1 JAK2 JAK3 JUNKDM5A KDM5C KDM6A KDR KEAP1 KEL KIT KLHL6 KMT2A (MLL) KMT2D (MLL2) KRAS LTKLYN MAF MAP2K1 (MEK1) MAP2K2 (MEK2) MAP2K4 MAP3K1 MAP3K13 MAPK1 MCL1 MDM2MDM4 MED12 MEF2B MEN1 MERTK MET MITF MKNK1 MLH1 MPL MRE11A MSH2 MSH3 MSH6MST1R MTAP MTOR MUTYH MYC MYCL (MYCL1) MYCN MYD88 NBN NF1 NF2 NFE2L2 NFKBIANKX2-1 NOTCH1 NOTCH2 NOTCH3 NPM1 NRAS NT5C2 NTRK1 NTRK2 NTRK3 P2RY8 PALB2PARK2 PARP1 PARP2 PARP3 PAX5 PBRM1 PDCD1 (PD-1) PDCD1LG2 (PD-L2) PDGFRAPDGFRB PDK1 PIK3C2B PIK3C2G PIK3CA PIK3CB PIK3R1 PIM1 PMS2 POLD1 POLE PPARGPPP2R1A PPP2R2A PRDM1 PRKAR1A PRKCI PTCH1 PTEN PTPN11 PTPRO QKI RAC1 RAD21RAD51 RAD51B RAD51C RAD51D RAD52 RAD54L RAF1 RARA RB1 RBM10 REL RET RICTORRNF43 ROS1 RPTOR SDHA SDHB SDHC SDHD SETD2 SF3B1 SGK1 SMAD2 SMAD4 SMARCA4SMARCB1 SMO SNCAIP SOCS1 SOX2 SOX9 SPEN SPOP SRC STAG2 STAT3 STK11 SUFU SYKTBX3 TEK TET2 TGFBR2 TIPARP TNFAIP3 TNFRSF14 TP53TSC1 TSC2 TYRO3 U2AF1 VEGFAVHL WHSC1 (MMSET) WHSC1L1 WT1 XPO1 XRCC2 ZNF217 ZNF703.
[0063] In some embodiments, proteins encoding somatic mutations identified from each patient using WES are selected, and the most severe consequence of each mutation is retained. Sliding windows of lengths of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, and 30 amino acids are arranged on each mutation, and all windows containing mutations are retained. Each of these mutated amino acid sequences is then paired with each of the MHC-I and MHC-II alleles in all combinations of the patient's sequences, and analyzed using one or more neoantigens and / or neoantigen classifiers.
[0064] In some embodiments, a neoantigen classifier is used to identify neoantigens that predict not only whether the mutation will be displayed as a novel epitope on the cell surface, but also immunogenicity (i.e., whether T cells are likely to respond to the novel epitope). Neoantigen classifiers predict immunogenicity by examining the physicochemical properties of mutant peptides that bind to MHC. They are machine learning models trained on experimental datasets consisting of novel epitopes that stimulate T cell activation and other novel epitopes that do not stimulate immune activation. The additional prediction of immunogenicity by the neoantigen classifier improves the prediction of immunotherapy responses.
[0065] In some embodiments, a variety of different neoantigens and / or neoantigen classifiers (such as at least 2, at least 3, at least 4, or at least 5 different neoantigens and / or neoantigen classifiers) are used to identify neoantigens.
[0066] In some embodiments, the neoantigen and / or neoantigen classifier can assess whether a patient's MHC alleles can bind to the mutated amino acid sequence and / or whether the mutated amino acid sequence can stimulate a T cell response (immunogenicity). The T cell response can be a CD8+ cell response and / or a CD4+ cell response. Mutated amino acid sequences that strongly bind to the patient's MHC I allele and / or are immunogenic (e.g., IC50) are retained. 50 <600 nM, or IC 50 <550 nM, or IC 50 <500nM, or IC 50 <450 nM, or IC 50<400 nM; and / or percentile rank <0.6%, or percentile rank <0.55%, or percentile rank <0.5%, or percentile rank <0.45%, or percentile rank <0.4%. Sum the number of neoantigens and / or neoantigens for each of the neoantigens and / or neoantigen classifiers.
[0067] In some embodiments, RNAseq data are used to assemble T and B cell receptor sequences present within tumor biopsies. In some embodiments, the number of unique T and B cell receptors (α-diversity) within the biopsy RNAseq data is quantified. In some embodiments, tumor gene expression data are used to determine the enrichment of immune cell populations present within the tumor. In some embodiments, mutated amino acid sequences originating from unexpressed genes are removed.
[0068] In some embodiments, quantifying the number of unique T and B cell receptors includes: (i) deconvolving the proportion of immune cells in a tumor sample based on RNA sequencing data, and (ii) assembling B and T cell receptors to quantify the number of unique T and B cell receptors.
[0069] In some embodiments, tumor gene expression data are used to determine the enrichment of the following immune cell populations present in the tumor: classical monocytes (monocytes.C), plasmacytoid dendritic cells (pDC), or both. In some embodiments, tumor gene expression data are used to determine the enrichment of the following immune cell populations present in the tumor: mucosa-associated invariant T cells (MAIT), myeloid dendritic cells (mDC), low-density neutrophils (neutrophils.LD), CD4+, and so on. + Memory T cells (T.CD4 memory), CD4 + Primary T cells (T.CD4 primary), IFNG single gene expression, CD274 single gene expression, or combinations thereof. In some embodiments, tumor gene expression data are used to determine the enrichment of the following immune cell populations present in the tumor: memory B cells (B. memory), primary B cells (B. primary), low-density basophils (basophil.LD), non-classical intermediate monocytes (monocyte.NC.I), natural killer cells (NK), plasmablasts, CD8+. + Memory T cells (T.CD8 memory), CD8 + Primary T cells (T.CD8.primary), non-VD2γδT cells (T.gd.non.Vd2), VD2γδT cells (T.gd.Vd2), CD8A single gene expression, CD8B single gene expression, CD4 single gene expression, CD8 / CD4 expression ratio, TCGA subtypes or combinations thereof.
[0070] In some embodiments, RNAseq data are used to determine the expression of one or more immune checkpoint molecules in tumor samples. In some embodiments, immune checkpoint molecules include CD137, CD134, PD-1, KIR, LAG-3, PD-L1, PDL2, CTLA-4, B7.1, B7.2, B7-DC, B7-H1, B7-H2, B7-H3, B7-H4, B7-H5, B7-H6, B7-H7, BTLA, LIGHT, HVEM, GAL9, TIM-3, TIGHT, VISTA, 2B4, CGEN-15049, CHK1, CHK2, A2aR, TGF-β, PI3Kγ, GITR, ICOS, IDO, TLR, IL-2R, IL-10, PVRIG, CCRY, OX-40, CD160, CD20, CD52, CD47, CD73, CD27-CD70, CD40, or combinations thereof.
[0071] In some embodiments, RNA sequencing is batch RNA sequencing. In some embodiments, RNA sequencing is single-cell RNA sequencing or sequencing targeting T and B cell receptors.
[0072] In some embodiments, the patient's tumor sample is from a solid tumor. In some embodiments, the patient's tumor sample is from a solid tumor. In some embodiments, the patient's tumor sample is from the abdomen or abdominal wall, adrenal gland, anus, appendix, bladder, bone, brain, breast, cervix, chest wall, colon, diaphragm, duodenum, ear, endometrium, esophagus, fallopian tube, gallbladder, gastroesophageal junction, head and neck, kidney, larynx, liver, lung, lymph nodes, malignant effusion, mediastinum, nasal cavity, omentum, ovary, pancreas, pancreaticobiliary duct, parotid gland, pelvis, penis, pericardium, peritoneum, pleura, prostate, rectum, salivary glands, skin, small intestine, soft tissue, spleen, stomach, thyroid gland, tongue, trachea, ureter, uterus, vagina, vulva, or Whipple's resection site. In some embodiments, the patient has breast cancer, colorectal cancer, gastrointestinal cancer, kidney cancer, lung cancer, multiple myeloma, ovarian cancer, or pancreatic cancer.
[0073] In some embodiments, cancer patients are identified as immunotherapy responders using analytical models based on the number of neoantigens in tumor samples, the number and abundance of unique T and B cell receptors, and the enrichment of immune cell populations in tumor samples. In some embodiments, the analytical models further consider pre-immunotherapy ctDNA results and clinical characteristics (e.g., age, cancer stage, prior therapy, metastasis). Detection of ctDNA in blood samples from cancer patients is described, for example, in WO2015 / 164432 and WO2019 / 200228, which are incorporated herein by reference in their entirety.
[0074] In some embodiments, a machine learning model, such as a random forest model, is trained using the number of neoantigens in the tumor sample, the number of unique T and B cell receptors, enrichment of immune cell populations in the tumor sample, pre-immunotherapy ctDNA results, and clinical characteristics (e.g., age, cancer stage, prior therapy, metastasis). The random forest model is trained using 80% of the patient dataset. The trained model is then tested using the remaining 20% of the patient dataset. In other instances, the machine learning model can be a linear regression, logistic regression, a random forest classifier, a cutoff value learned from historical data, a support vector machine, a neural network, or a deep neural network.
[0075] In some embodiments, cancer patients are identified as immunotherapy responders by using the number of neoantigens in the tumor sample, the number of unique T and B cell receptors, enrichment of classical monocytes (monocyte C) and plasmacytoid dendritic cell (pDC) populations in the tumor sample, and optionally, the pre-immunotherapy ctDNA positivity status in the patient's blood sample.
[0076] In some embodiments, based on the number of unique T and B cell receptors in the tumor sample, classical monocytes (monocytes C), plasmacytoid dendritic cells (pDC), mucosa-associated invariant T cells (MAIT) (medullary dendritic cells (mDC)), low-density neutrophils (neutrophils.LD), CD4 + Memory T cells (T.CD4 memory), CD4 + Enrichment of the primary T cell (T.CD4 primary) population, IFNG single gene expression and CD274 single gene expression, and, optionally, pre-immunotherapy ctDNA positivity in the patient's blood sample, the number of neoantigens in the tumor sample were used to identify cancer patients as immunotherapy responders.
[0077] In some embodiments, based on the number of unique T and B cell receptors in the tumor sample, classical monocytes (monocytes C), plasmacytoid dendritic cells (pDC), mucosa-associated invariant T cells (MAIT) (medullary dendritic cells (mDC)), low-density neutrophils (neutrophils.LD), CD4 + Memory T cells (T.CD4 memory), CD4 + Enrichment of primordial T cell (T.CD4 primordial) population, IFNG single gene expression and CD274 single gene expression, node status, baseline ECOG, and, optionally, pre-immunotherapy ctDNA positivity in blood samples from patients, were used to identify cancer patients as immunotherapy responders using the amount of neoantigens in tumor samples.
[0078] In some embodiments, the methods described herein identify cancer patients as immunotherapy responders without requiring TMB analysis. In some embodiments, the methods described herein do not include any step of determining TMB.
[0079] In some embodiments, the method identifies cancer patients as immunotherapy responders with a positive predictive value (PPV) at least 10%, at least 20%, at least 30%, at least 40%, or at least 50% higher than using TMB alone. In some embodiments, the method identifies cancer patients as immunotherapy non-responders with a negative predictive value (NPV) at least 10%, at least 20%, at least 30%, at least 40%, or at least 50% higher than using TMB alone.
[0080] In some embodiments, the method identifies cancer patients as immunotherapy responders with a sensitivity at least 10%, at least 20%, at least 30%, at least 40%, or at least 50% higher than using TMB alone. In some embodiments, the method identifies cancer patients as immunotherapy non-responders with a specificity at least 10%, at least 20%, at least 30%, at least 40%, or at least 50% higher than using TMB alone.
[0081] In another aspect, this disclosure relates to a method for identifying cancer patients as responders to PD-1, PD-L1, or CTLA-4 immunotherapy, the method comprising: performing whole-exome sequencing on a tumor sample from the patient to quantify the number of neoantigens in the tumor sample; performing RNA sequencing on the tumor sample from the patient to quantify the number and abundance of unique T and B cell receptors, enrichment of immune cell populations, and expression of PD-1, PD-L1, or CTLA-4 in the tumor sample; and identifying the cancer patient as a responder to PD-1, PD-L1, or CTLA-4 immunotherapy using the number of neoantigens in the tumor sample, the number of unique T and B cell receptors in the tumor sample, the enrichment of immune cell populations, and the expression of PD-1, PD-L1, or CTLA-4.
[0082] In some embodiments, cancer patients are identified as responders to PD-1 or PD-L1 or CTLA-4 immunotherapy by using the number of neoantigens in the tumor sample, the number and abundance of unique T and B cell receptors, the enrichment of classical monocytes (monocytes C) and plasmacytoid dendritic cell (pDC) populations in the tumor sample, and optionally the pre-immunotherapy ctDNA positivity status and PD-1 or PD-L1 expression in the tumor sample.
[0083] In some embodiments, cancer patients are identified as responders to PD-1, PD-L1, or CTLA-4 immunotherapy based on the number of unique T and B cell receptors in the tumor sample, enrichment of classical monocytes (monocyte C), plasmacytoid dendritic cells (pDC), mucosa-associated invariant T cells (MAIT), myeloid dendritic cells (mDC), low-density neutrophils (neutrophils.LD), CD4+ memory T cells (T.CD4 memory), CD4+ primordial T cells (T.CD4 primordial), IFNG single gene expression, CD274 single gene expression, and optionally, pre-immunotherapy ctDNA positivity and PD-1 or PD-L1 expression in the tumor sample.
[0084] In some embodiments, the number of neoantigens in the tumor sample, the number and abundance of unique T and B cell receptors in the tumor sample, classical monocytes (monocyte C), plasmacytoid dendritic cells (pDC), mucosa-associated invariant T cells (MAIT) (medullary dendritic cells), low-density neutrophils (neutrophils.LD), and CD4 are used. + Memory T cells (T.CD4 memory), CD4 + Enrichment of the primary T cell (T.CD4 primary) population, IFNG single gene expression and CD274 single gene expression, node status, baseline ECOG, and, optionally, pre-immunotherapy ctDNA positivity and PD-1 or PD-L1 or CTLA-4 expression in tumor samples, identified cancer patients as responders to PD-1 or PD-L1 or CTLA-4 immunotherapy.
[0085] In some embodiments, the methods described herein identify cancer patients as responders to PD-1, PD-L1, or CTLA-4 immunotherapy without TMB analysis. In some embodiments, the methods described herein do not include any step of determining TMB.
[0086] In some embodiments, the term “cancer” refers to or describes a physiological condition in animals or humans that is typically characterized by unregulated cell growth.
[0087] In some embodiments, a "tumor" comprises one or more cancerous cells. There are several main types of cancer. A carcinoma is a cancer that begins in the skin or in tissues that line or cover internal organs. A sarcoma is a cancer that begins in bone, cartilage, fat, muscle, blood vessels, or other connecting or supporting tissues. Leukemia is a cancer that begins in blood-forming tissues, such as bone marrow, and causes a large number of abnormal blood cells to be produced and enter the bloodstream. Lymphoma and multiple myeloma are cancers that begin in the cells of the immune system. Central nervous system cancers are cancers that begin in the tissues of the brain and spinal cord. In some embodiments, the cancer has metastasized. In some embodiments, the cancer has not metastasized.
[0088] Methods for treating cancer using immunotherapy This disclosure further relates to a method for treating cancer, the method comprising identifying a cancer patient as an immunotherapy responder using the method described herein, and administering an effective amount of immunotherapy to the cancer patient.
[0089] In some embodiments, immunotherapy includes checkpoint inhibitors, CAR-T therapy, TCR-T therapy, NK cell therapy, neoantigen vaccines, oncolytic viruses, cytokines, monoclonal antibodies, or combinations thereof.
[0090] In some embodiments, immunotherapy includes immune checkpoint inhibitors. Immune checkpoint molecules may be CD137, CD134, PD-1, KIR, LAG-3, PD-L1, PDL2, CTLA-4, B7.1, B7.2, B7-DC, B7-H1, B7-H2, B7-H3, B7-H4, B7-H5, B7-H6, B7-H7, BTLA, LIGHT, HVEM, GAL9, TIM-3, TIGHT, VISTA, 2B4, CGEN-15049, CHK1, CHK2, A2aR, TGF-β, PI3Kγ, GITR, ICOS, IDO, TLR, IL-2R, IL-10, PVRIG, CCRY, OX-40, CD160, CD20, CD52, CD47, CD73, CD27-CD70, CD40, or combinations thereof. In some embodiments, immunotherapy includes immune checkpoint inhibitors selected from PD-1 inhibitors, PD-L1 inhibitors, CLTA-4 inhibitors, or combinations thereof.
[0091] In some embodiments, immunotherapy includes pembrolizumab, nivolumab, simipremab, dostalimumab, atezolizumab, avelumab, durvalumab, ipilimumab, or trimemumab. In some embodiments, immunotherapy includes voraprazumab, spartazolizumab, camrelizumab, syndilimumab, tislelizumab, toripalimab, INCMGA00012, AMP-224, AMP-514, KN035, cochilimumab, AUNP12, CA-170, or BMS-986189.
[0092] In some embodiments, immunotherapy includes a therapeutic cell composition. Exemplary therapeutic cell compositions include, but are not limited to, T cells, natural killer (NK) cells and dendritic cells, chimeric antigen receptor T (CAR-T) cells, T cell receptor-engineered T (TCR-T) cells, and chimeric antigen receptor-natural killer (CAR-NK) cells. In some embodiments, immunotherapy includes a combination of a therapeutic cell composition and a cancer vaccine.
[0093] This disclosure further relates to a method for treating cancer, the method comprising identifying a cancer patient as a responder to PD-1 or PD-L1 or CTLA-4 immunotherapy using the method described herein, and administering an effective amount of immunotherapy to the cancer patient, wherein the PD-1 or PD-L1 or CTLA-4 immunotherapy includes pembrolizumab, nivolumab, simiprelimab, dostalimab, atezolizumab, avelumab, durvalumab, ipilimumab, or trimemumab.
[0094] In some embodiments, the cancer patient has been treated with surgery, chemotherapy, or radiation therapy, either or simultaneously.
[0095] In some embodiments, the cancer is breast cancer, colorectal cancer, gastrointestinal cancer, kidney cancer, lung cancer, bladder cancer, ovarian cancer, or pancreatic cancer.
[0096] In some embodiments, cancer is a cancer or tumor of the abdomen or abdominal wall, adrenal gland, anus, appendix, bladder, bone, brain, breast, cervix, chest wall, colon, diaphragm, duodenum, ear, endometrium, esophagus, fallopian tube, gallbladder, gastroesophageal junction, head and neck, kidney, larynx, liver, lung, lymph nodes, malignant effusion, mediastinum, nasal cavity, omentum, ovary, pancreas, pancreaticobiliary duct, parotid gland, pelvis, penis, pericardium, peritoneum, pleura, prostate, rectum, salivary gland, skin, small intestine, soft tissue, spleen, stomach, thyroid gland, tongue, trachea, ureter, uterus, vagina, vulva, or Whipple's resection site.
[0097] Exemplary cancers used in any of the methods described herein include solid tumors, carcinomas, sarcomas, lymphomas, leukemias, germ cell tumors, or blastomas. In some embodiments, the cancer is acute lymphoblastic leukemia, acute myeloid leukemia, adrenocortical carcinoma, anal cancer, appendiceal cancer, astrocytoma (such as cerebellar or cerebral astrocytoma in children), basal cell carcinoma, bile duct cancer (such as extrahepatic bile duct cancer), bladder cancer, bone tumor (such as osteosarcoma or malignant fibrous histiocytoma), brainstem glioma, brain cancer (such as cerebellar astrocytoma, cerebral astrocytoma / malignant glioma, ependymoma, neuroblastoma, supratentorial primitive neuroectodermal tumor, or optic pathway and hypothalamic glioma), glioblastoma, breast cancer, bronchial adenoma or carcinoid, Burkitt's lymphoma, carcinoid tumor (such as pediatric or gastrointestinal carcinoid tumors). Cancers including: central nervous system lymphoma, cerebellar astrocytoma or malignant glioma (such as pediatric cerebellar astrocytoma or malignant glioma), cervical cancer, childhood cancer, chronic lymphocytic leukemia, chronic myeloid leukemia, chronic myeloproliferative disorder, colon cancer, cutaneous T-cell lymphoma, desmoplastic small round cell tumor, endometrial cancer, ependymoma, esophageal cancer, Ewing's sarcoma, tumors in the Ewing's tumor family, extracranial germ cell tumors (such as pediatric extracranial germ cell tumors), gonadal extragerminal tumors, eye cancer (such as intraocular melanoma or retinoblastoma), gallbladder cancer, gastric cancer, gastrointestinal carcinoid tumors, gastrointestinal stromal tumors, and germ cell tumors. (Such as extracranial, extragonadal, or ovarian germ cell tumors), gestational trophoblastic tumors, gliomas (such as brainstem, pediatric astrocytoma, or pediatric visual pathway and hypothalamic gliomas), gastric carcinoid tumors, piloblastic leukemia, head and neck cancer, heart cancer, hepatocellular carcinoma, Hodgkin's lymphoma, hypopharyngeal cancer, hypothalamic and visual pathway gliomas (such as pediatric visual pathway gliomas), islet cell carcinomas (such as endocrine or islet cell carcinomas), Kaposi's sarcoma, kidney cancer, laryngeal cancer, leukemias (such as acute lymphoblastic, acute myeloid, chronic lymphocytic, chronic myeloid, or piloblastic leukemia), lip or oral cancers, liposarcomas, liver cancers (such as non-small cell or small cell hepatocellular carcinomas). Cellular cancers), lung cancer, lymphomas (such as AIDS-related, Burkitt's, cutaneous T-cell, Hodgkin's, non-Hodgkin's, or central nervous system lymphomas), macroglobulinemia (such as Waldenström's macroglobulinemia, malignant fibrous histiocytoma or osteosarcoma of the skeleton, neuroblastoma (such as neuroblastoma in children), melanoma, Merkel cell carcinoma, mesothelioma (such as mesothelioma in adults or children), occult metastatic squamous neck cancer, oral cancer, multiple endocrine adenoma syndrome (such as multiple endocrine adenoma syndrome in children), multiple myeloma or plasmacytoma, mycosis fungoides, myelodyplasia syndrome, myelodyplasia or myeloproliferative disorders,Myeloid leukemia (such as chronic myeloid leukemia), myeloid leukemia (such as acute myeloid leukemia in adults or children), myeloproliferative disorders (such as chronic myeloproliferative disorders), nasal cavity or paranasal sinus cancer, nasopharyngeal carcinoma, neuroblastoma, oral cancer, oropharyngeal cancer, osteosarcoma or malignant fibrous histiocytoma of bone, ovarian cancer, ovarian epithelial cancer, ovarian germ cell tumors, low-grade potential ovarian tumors, pancreatic cancer (such as islet cell pancreatic cancer), paranasal sinus or nasal cavity cancer, parathyroid cancer, penile cancer, pharyngeal cancer, pheochromocytoma, pineal astrocytoma, pineal germ cell tumor, pineal gland... Cytomas or supratentorial primitive neuroectodermal tumors (such as pineal cell carcinoma or supratentorial primitive neuroectodermal tumors in children), pituitary adenomas, plasma cell vegetations, pleural pulmonary blastomas, primary central nervous system lymphomas, cancers, rectal cancer, renal cell carcinoma, renal pelvis or ureter cancers (such as transitional cell carcinoma of the renal pelvis or ureter), retinoblastomas, rhabdomyosarcomas (such as rhabdomyosarcoma in children), salivary gland cancers, sarcomas (such as sarcomas in the Ewing's tumor family, Kaburg's tumor, soft tissue or uterine sarcoma), Seychelles syndrome, skin cancers (such as non-melanoma, melanoma, or Merkel cell skin cancer). Cell skin cancer), small bowel cancer, squamous cell carcinoma, supratentorial primitive neuroectodermal tumors (such as supratentorial primitive neuroectodermal tumors in children), T-cell lymphomas (such as cutaneous T-cell lymphomas), testicular cancer, laryngeal cancer, thymoma (such as thymoma in children), thymoma or thymic carcinoma, thyroid cancer (such as thyroid cancer in children), trophoblastic tumors (such as trophoblastic tumors of pregnancy), cancers of unknown primary site (such as cancers of unknown primary site in adults or children), urethral cancers (such as endometrial uterine cancer), uterine sarcoma, vaginal cancer, optic pathway or hypothalamic gliomas (such as optic pathway or hypothalamic gliomas in children), vulvar cancer, Waldenström macroglobulinemia or Wilms' tumor (such as Wilms' tumor in children).
[0098] A machine learning tool is provided that predicts individual responses to immune checkpoint inhibitor (ICI) therapy based on a large number of individuals with complete clinical information. By utilizing whole-exome sequencing data, ctDNA timepoint analysis, and clinical data, a deep learning-based immune response prediction model successfully predicted which individuals would respond to and benefit from ICI therapy across multiple cancer types.
[0099] The examples in this paper utilize ctDNA status and / or neoantigen load for risk stratification and treatment response prediction, combining them with biomarkers / clinicopathological features to significantly improve treatment strategies. The development of such multifactorial models requires a large cohort of individuals with detailed clinical information. The immune response prediction model was developed using a comprehensive and extensive dataset, and machine learning algorithms evaluated and identified the optimal set of features that could accurately predict an individual's response to ICI therapy. This immune response prediction model was tested and validated in individuals with melanoma, NSCLC, and CRC, and validated in additional cancer types where ICI therapy is a potential line of treatment. Because this immune response prediction model is based on available genomic and clinicopathological data collected as part of routine development of tumor informed consent, bespoke ctDNA assays, and evaluation of tumor driver variants via comprehensive genomic profiling, it can be used at the start of treatment and ctDNA assay development for individuals with tumors who may benefit from ICI therapy.
[0100] Example 1 Predictive biomarkers can help individuals most likely to benefit from immune checkpoint inhibitors (ICIs), but their predictive accuracy and specificity are limited. An immune response prediction model is provided that incorporates neoantigen load and clinicopathological variables to improve the prediction of ICI responses compared to tumor mutational burden (TMB).
[0101] The immune response prediction model was based on tumor and matched normal whole-exome sequencing data from individuals with solid tumors (melanoma n=238, non-small cell lung cancer (NSCLC) n=133, and high microsatellite instability colorectal cancer (MSI-high CRC) n=172) who underwent ICI and were referred for personalized, tumor-informed mPCR-NGS circulating tumor (ct) DNA assays. The model was then trained using predicted neoantigens based on several open-source classifiers and clinicopathological characteristics of 80% of the individuals. Progression-free survival (PFS) after ICI was used to evaluate the model's performance relative to TMB status alone in the remaining 20% of individuals and in an independent individual validation set enrolled in the BESPOKE IO clinical trial (NCT04761783).
[0102] The parameters selected for the immune response prediction model included neoantigen predictors, nodal status, baseline ECOG score, and pre-ICI ctDNA. In melanoma individuals, both the model and TMB predicted the response, even though the model (HR=12.68, P<0.001, area under the curve (AUC)=0.94) was more accurate than TMB (HR=3.23, P=0.021, AUC=0.82). In NSCLC individuals, the model outperformed TMB in predicting PFS (model: HR=6.73, P=0.006, AUC=0.88; TMB: HR=1.31, P=0.69, AUC=0.64), and a similar trend was observed in individuals with high MSI and CRC (model: HR=6.10, P=0.014, AUC=0.85; TMB: HR=3.52, P=0.076, AUC=0.5). These results were further validated on the BESPOKE IO dataset (model: HR=3.062, P<0.001, AUC=0.75; TMB: HR=1.76, P=0.094, AUC=0.70). Compared with TMB, the inclusion of neoantigen load and other clinical variables significantly improved the prediction of individual ICI responses. This model can be used with a variety of different cancer types.
[0103] Study design and clinical data extraction To develop an immune response prediction model, a retrospective analysis was performed on whole-exome sequencing (WES) data from clinical genomic databases (N=86,635), such as... Figure 1A As shown. Figure 1A This is a graph of the individuals included in the sub-analysis. WES was performed as part of the ctDNA test. The analysis included individuals with malignant melanoma, non-small cell lung cancer (NSCLC), and colorectal cancer (CRC), totaling 48,228 individuals. Clinical outcomes of individuals who had undergone clinically indicated testing were collected as part of ongoing quality assurance procedures. Individuals who received ICI without other concurrent therapies and who were ctDNA positive at baseline (before ICI) or at any time after ICI initiation and whose complete clinical outcome information was available in the clinical testing database were included in the real-world data analysis (melanoma: 238, NSCLC: 133, MSI-high CRC: 172), as shown. Figure 1A As shown. Deidentified personal information (including age, sex, ECOG score, cancer type, and stage) was extracted and included in these analyses according to the IRB-approved quality assurance protocol (Salus Protocol # 20099-01). The study was conducted in accordance with the Declaration of Helsinki.
[0104] In addition, the analysis included a blinded validation cohort of 99 individuals with lung disease, melanoma, and MSI-high CRC who were enrolled in a prospective, longitudinal, multicenter observational study (BESPOKE IO, https: / / clinicaltrials.gov / study / NCT04761783, approved by the Ethics and Independent Review Service Protocol # Natera-20-043-NCP BESPOKEctDNA-guided Immunotherapy Study). The cohort included here represents a temporary group of individuals with fully annotated data and a median clinical follow-up of 10 months.
[0105] WES DNA analysis Figure 1B This is a schematic diagram of an exemplary immune response prediction model. Figure 1B This is a representation of the process involved in developing a machine learning-based immune response prediction model. For each individual, WES data were processed from FFPE tumor tissue and matched with normal blood to obtain somatic and germline genomic information (green boxes). This information was used for phase variants and as input to the variant effect predictor, and for genotyping of MHC-I and MHC-II alleles. These data were used to evaluate MHC peptide binding prediction and T cell activation prediction. In addition to clinical variables and ctDNA results, these data were also included in the immune response prediction model.
[0106] Formalin-fixed and paraffin-embedded tumor tissue and matched normal blood samples underwent WES and bioinformatics processing. Tumor and germline common single nucleotide variants (SNVs) and insertions / deletions (indels) were identified using a call algorithm via aligned tumor and normal BAM files. Somatic SNVs and insertions / deletions were phased with germline variants to determine mutant germline allele haplotypes and then annotated (VEP v109) to determine the highest impact consequence for each mutation. MHC class I (MHC-I) alleles were genotyped using OptiType v1.3.3. MHC class II (MHC-II) alleles were genotyped using HISAT genotyping v1.3.3.
[0107] Peptide-MHC binding prediction Peptide-MHC binding prediction was performed in two steps. First, pVAC-Seq v4.0.5 was used to summarize peptide-MHC binding metrics using all detected amino acid sequence alteration mutations and MHC-I and MHC-II alleles, and MHCflurry v2.1 and MHCflurryEL v2.1 were used to predict subsequent MHC-I mutant peptide treatment and presentation scores. Next, the predicted T cell immunogenicity of each mutant peptide-MHC-I combination was evaluated, such as... Figure 1B As shown.
[0108] Personalized, tumor-informed ctDNA testing In summary, to develop personalized assays, 16 individual-specific somatic single nucleotide variants (SNVs) were selected from WES tumors and matched against normal data to design individual-specific primers. Plasma samples isolated from whole blood were then analyzed for ctDNA detection. Plasma samples with at least two of the 16 variants detected above a predefined algorithm confidence threshold were defined as ctDNA positive and their concentrations were measured and reported as mean tumor molecules (MTM) / mL plasma.
[0109] Tumor mutation burden TMB is calculated as the number of nonsynonymous somatic mutations identified by WES divided by the WES hybridization capture manifold size (in megabases). A TMB value ≥ 10 mutations / Mb is considered "high", while a TMB value < 10 mutations / Mb is considered "low".
[0110] Univariate analysis of clinicopathological characteristics Using R v4.2.2 library survival v3.4, the effects of individual age (above / below 65 years), sex, ECOG score, cancer stage, pre-ICI ctDNA results (positive / negative), and TMB (above / below 10 nonsynonymous mutations / Mb) on progression-free survival (PFS) were individually tested using Cox proportional hazards. PFS was measured from the date of ICI initiation to the first recorded radiographic sign of progression, with data reviewed at the last follow-up or death. Forest plots were generated using R library survival analysis v0.3.0.
[0111] Model training and testing The immunogenicity of neoantigens varies due to a range of properties / characteristics and cannot be captured by a single score. Here, the neoantigen set for each individual is scored cumulatively across all combinations of five properties: processing (0-0.025, 0.025-0.05, 0.05-0.1, 0.1-0.5, 0.5-0.5-1), presentation (0-0.1, 0.1-0.2, 0.2-0.4, 0.4-0.6, 0.6-1), MHC binding fold change as measured by log2 (WT IC50 / MT IC50) (0-1, 1-2, 2-3, 3-4, 4-5, >5), predicted immunogenicity based on percentage ranking (0-0.25, 0.25-0.5, 0.5-1, 1-5, 5-10, >10), and dissimilarity to the reference human proteome based on (<0.75, ≥0.75).
[0112] The above data, pre-ICI ctDNA results, and clinical characteristics (e.g., age, sex, ECOG score, cancer stage) were then used to train a random survival forest model to predict cancer progression. 80% of the individuals were used for training, and the remaining 20% were used for testing. Training was performed using 5-fold cross-validation, and the best model was selected based on the integrated Brier score. The training and testing datasets were balanced to include an equal proportion of individuals with progression. Analysis was performed using the R v4.2.2, tidymodelsv1.0.0, and partykit 1.2-20 libraries. KM curves were plotted using the R survminer v0.4.9 library; AUC was calculated using the pROC v1.18.5 library.
[0113] result The training and testing set included a total of 543 individuals with selective cancer indications, such as... Figure 1A As shown. These individuals included 238 (43.8%) diagnosed with melanoma, 133 (24.5%) diagnosed with NSCLC, and 172 (31.7%) diagnosed with MSI-high CRC. An additional 99 patients with solid tumors were included in the validation set. The most common ICI regimens varied based on cancer type: ipilimumab and nivolumab for melanoma (33.9%), pembrolizumab for NSCLC (47.9%) and CRC (70.0%), and ipilimumab and nivolumab for the Bespoke IO dataset (37.4%). The median PFS was 161 days for melanoma, 253 days for NSCLC, 259 days for MSI-high CRC, and 148 days for the BESPOKE IO validation dataset. Complete demographic information is included in Table 1 below.
[0114] Table 1. Individual Demographic Data and Clinical Characteristics
[0115] Melanoma Figure 2AThis is a graphical representation of Kaplan-Meir progression-free survival estimates for 49 patients with melanoma stratified by TMB status. The median follow-up period for melanoma individuals (N=49) was 6 months (range 1–62 months). When assessing the predictive value of TMB, 22.6% (7 out of 31) of individuals with TMB > 10 mut / Mb reported progression, compared to 61.1% (11 out of 18) of individuals with TMB < 10 mut / Mb (HR: 3.23, 95% CI 1.20–8.76, P=0.021). Figure 2A As shown.
[0116] Figure 2B This is a graphical representation of the Kaplan-Mayer progression-free survival estimates for 49 individuals with melanoma stratified based on a response prediction model. When assessing the model's predictive value, 87.5% (14 out of 16) of non-responders (predicted progression) reported progression, compared to 12.1% (4 out of 33) of responders (individuals not predicted to progress) (HR: 12.68, 95% CI 3.57–45.02, P < 0.001). Figure 2B As shown.
[0117] Figure 2C This is a graphical representation of the area under the curve (AUC) of the ICI response used to predict TMB and the prediction model. The prediction accuracy of both TMB and the model was calculated, and the AUC for TMB alone was 0.82, while the AUC for the model was 0.94. Figure 2C As shown. Figure 2D This is a graphical representation of a univariate analysis used to evaluate the predictive value of the immune response prediction model and other clinicopathological factors. Among other clinicopathological features, in the univariate analysis, the immune response prediction model was the most important factor predicting an increased risk of progression (p<0.001), followed by baseline positivity (p=0.034) and TMB status (p=0.021), such as... Figure 2D As shown, survival analysis was performed using the Kaplan-Mayer estimator and the Cox method.
[0118] NSCLC Figure 3AThis is a graphical representation of Kaplan-Mayer progression-free survival estimates for 24 NSCLC patients stratified by TMB status. The median follow-up period for NSCLC individuals was 12 months (range 1–82 months). When assessing the predictive value of TMB, 42.9% (3 out of 7) of individuals with TMB > 10 mut / Mb reported progression, compared to 58.8% (10 out of 17) of individuals with TMB < 10 mut / Mb (HR: 1.31, 95% CI 0.35–4.85, P = 0.69). Figure 3A As shown.
[0119] Figure 3B This is a graphical representation of the Kaplan-Mayer progression-free survival estimates for 24 individuals with NSCLC stratified based on a response prediction model. When assessing the model's predictive value, 100% (9 out of 9) of the non-responders reported progression, compared to 26.7% (4 out of 15) of the non-responders (HR: 6.73, 95% CI 1.74–26.07, P = 0.006). Figure 3B As shown.
[0120] Figure 3C This is a graphical representation of the area under the curve (AUC) of the ICI response used to predict TMB and the prediction model. The prediction accuracy of both TMB and the model was calculated, and the AUC for TMB alone was 0.64, while the AUC for the model was 0.88. Figure 3C As shown. Figure 3D This is a graphical representation of a univariate analysis used to evaluate the predictive value of the immune response prediction model and other clinicopathological factors. Among other clinicopathological features, in the univariate analysis, the immune response prediction model was the only significant factor predicting an increased risk of progression (p=0.006), such as... Figure 3D As shown, survival analysis was performed using the Kaplan-Mayer estimator and the Cox method.
[0121] CRC Since ICI treatment is only applicable to CRC individuals with MSI-high tumors, individuals meeting this criterion are included in this analysis. The median follow-up period for melanoma individuals was 14 months (range 2–37 months).
[0122] Figure 4AThis is a graphical representation of progression-free survival in 36 patients with MSI-H CRC stratified by TMB status. When assessing the predictive value of TMB, 18.5% (5 out of 27) of individuals with TMB > 10 mut / Mb reported progression, compared to 44.4% (4 out of 9) of individuals with TMB < 10 mut / Mb (HR: 3.52, 95% CI 0.88–14.11, P = 0.076). Figure 4A As shown.
[0123] Figure 4B This is a graphical representation of the Kaplan-Mayer progression-free survival estimates for 36 individuals with MSI-H CRC stratified by progression based on a response prediction model. When assessing the model's predictive value, 75.0% (3 out of 4) of non-responders reported progression, compared to 18.8% (6 out of 32) of responders (HR: 6.10, 95% CI 1.45–25.69, P = 0.014). Figure 4B As shown.
[0124] Figure 4C This is a graphical representation of the area under the curve (AUC) of the ICI response used to predict TMB and the prediction model. The prediction accuracy of both TMB and the model was calculated, and the AUC for TMB alone was 0.50, while the AUC for the model was 0.85. Figure 4C As shown.
[0125] Figure 4D This is a graphical representation of a univariate analysis used to evaluate the predictive value of the immune response prediction model and other clinicopathological factors. The univariate analysis shows that the immune response prediction model was the only significant factor for predicting an increased risk of progression (p = 0.014), as... Figure 4D As shown, survival analysis was performed using the Kaplan-Mayer estimator and the Cox method.
[0126] BESPOKE IO (Blind Validation Group) Figure 5A This is a graphical representation of Kaplan-Mayer progression-free survival estimates for 99 patients with solid tumors stratified by TMB status. The median follow-up period for individuals in the blinded validation cohort was 10 months (range 2–35 months). When assessing the predictive value of TMB, 34.2% (13 out of 38) of individuals with TMB > 10 mut / Mb reported progression, compared to 45.9% (28 out of 61) of individuals with TMB < 10 mut / Mb (HR: 1.76, 95% CI 0.91–3.42, P = 0.094). Figure 5A As shown.
[0127] Figure 5B This is a graphical representation of the Kaplan-Mayer progression-free survival estimates for 99 individuals with solid tumors stratified by progression based on a response prediction model. When assessing the model's predictive value, 65.2% (13 out of 23) of the responders reported progression compared to 34.2% (26 out of 76) of the non-responders (HR: 3.06, 95% CI 1.60–5.86, P = 0.0007). Figure 5B As shown.
[0128] Figure 5C This is a graphical representation of the area under the curve (AUC) of the ICI response used to predict TMB and the prediction model. The prediction accuracy of both TMB and the model was calculated, and the AUC for TMB alone was 0.70, while the AUC for the model was 0.75. Figure 5C As shown.
[0129] Figure 5D This is a graphical representation of a univariate analysis used to evaluate the predictive value of the immune response prediction model and other clinicopathological factors. In the univariate analysis of individual clinicopathological characteristics of PFS, the immune response prediction model was the most important factor predicting an increased risk of progression (p<0.001), followed by stage (p=0.018) and baseline ctDNA positivity (p=0.023). Figure 5D As shown, other traditionally used clinicopathological risk factors were not significant. Survival analysis was performed using the Kaplan-Mayer estimator and the Cox method.
[0130] discuss Immune response prediction models assess multiple biomarkers representing different elements known to be important for responses to stimulating ICI therapy. The benefit of this approach is access to comprehensive genomic, clinicopathological, and individual treatment outcomes through clinical genomic databases and standard tools for assessing neoantigen load and T-cell response. In summary, based on a threshold of approximately 10 mutations per Mb, immune response prediction models are better able to identify responders and non-responders to ICI therapy compared to TMB alone.
[0131] The immune response prediction model considers several different neoantigen properties that influence immunogenicity and bins them across a range of scores. In this way, the results better define the neoantigens and ultimately improve predictions of which individuals are most likely to respond to ICI therapy.
[0132] The described immune response prediction model incorporates both results from baseline ctDNA testing and clinicopathological characteristics (i.e., node status, baseline ECOG score). The response prediction model considers both intrinsic tumor features and other peripheral features, which collectively improve predictive performance compared to using only a single data source.
[0133] RNAseq data can also be used. RNAseq data can improve the prediction of therapy response by retaining only neoantigen variants expressed at high levels of heterologity, such as gene-gene fusions and aberrant splicing events.
[0134] The predictive model is developed using one or more machine learning algorithms that evaluate a comprehensive dataset and identify the optimal set of features that can accurately predict neoantigen candidates, utilize immune stimulation in response to ICI therapy, and rank the neoantigen candidates. The predictive model can identify the optimal set of neoantigens, which can be incorporated into personalized cancer vaccines for individuals most likely to benefit from ICI therapy.
[0135] Predicting and selecting neoantigens Turn now Figures 6 to 10 Various apparatuses, systems and methods according to various aspects of this disclosure will be described.
[0136] Figures 6 to 10 The system, methods, and instructions for identifying and ranking the most common somatic mutations of a new antigen (candidate new antigen) for a certain type of cancer and for establishing a catalog of somatic mutations and vaccines are described. Figures 6 to 10 The system, method, and description are further described for predicting the immune response of an individual for treatment of a certain type of cancer.
[0137] In some instances, the identified common neoantigens represent 1-50% of the addressable histology of a particular type of cancer. Neoantigens are peptides generated by somatic mutations and are recognized as distinct from themselves and presented by antigen-presenting cells. A catalogue of somatic mutations and vaccines against a particular type of cancer can then be used for new subjects with that type of cancer to compare DNA and RNA sequencing and identify candidate vaccines for the new subjects.
[0138] Using historical data from a population of subjects treated for a specific type of cancer, somatic mutations in candidate neoantigens for that cancer type were identified through a population-level study of somatic mutations against neoantigens. Based on the identification of common somatic mutations in a specific cancer type, a vaccine library targeting that cancer type was developed using neoantigen candidate prediction and ranking. The most clinically significant somatic mutations against neoantigens for different cancer types were identified. In one example, the cancer types were melanoma, lung cancer, and colorectal cancer. Using data from hundreds of subjects treated for a specific cancer type with known outcomes, a predictive model identified somatic mutations in vaccine candidates. The predictive model included parameters (characteristics) of the subjects and somatic mutations in the neoantigen. The predictive engine identified subjects who had responded to therapy and had tumor samples with the same or similar somatic mutations.
[0139] In one instance, a machine learning model is used to rank neoantigen candidates based on methods for predicting neoantigen candidates, immune stimulation, and vaccine neoantigen candidates. Parameters are selected and used in the machine learning model to determine how subjects stratify over time following immunosuppressant therapy. Protein sequences of neoantigens based on DNA and RNA sequencing can be identified as unique to each subject (or multiple subjects). Survival benefits for subjects are determined based on the somatic mutations and characteristics of the resulting proteins in the subject's sample.
[0140] It should be understood that any type and combination of AI / machine learning algorithms, publicly available tools, open-source tools, and commercial products can be used for parameter selection and determination and for neoantigen candidate selection models, immune stimulation models, and / or neoantigen candidate ranking models (“predictive models”), and will be expected to be used in a similar manner. Those skilled in the art will understand that various commercial models are suitable for incorporation, such as Google… TM ALPHA FOLD, IntFOLD, RaptorX, HHpred, Phyre, Phyre2, and I-TASSER. Exemplary machine learning algorithms that can be used may include, but are not limited to, linear regression, logistic regression, decision trees, SVM algorithms, Naive Bayes algorithms, KNN algorithms, K-means, random forest algorithms, dimensionality reduction algorithms, gradient boosting algorithms, AdaBoosting algorithms, deep learning, and neural networks.
[0141] In some instances, machine learning models utilize several parameters to determine which somatic mutations are vaccine candidates. These parameters include the subject's response to treatment (such as checkpoint inhibitors), the fold change difference between native proteins in the WES germline and somatic mutant proteins from tumor samples, the protein expression level of candidate neoantigen proteins (in some cases, TCGA25), MHC-I: mutant peptide processing and presentation score (IQ), as well as T cell activation, dissimilarity to a reference human proteome, and MHC binding fold change. Additional parameters may include response to immune checkpoint inhibitors (ICIs), ctDNA results, age, sex, and ECOG score. It should be understood that any combination of these parameters can be used with the neoantigen candidate models, immune stimulation models, and neoantigen candidate ranking models described below.
[0142] One parameter is the expression level of the neoantigen from the somatic mutation. For example, the number of subjects with the same or similar somatic mutations, and whether *a* expresses a protein from the somatic mutation. This can be determined based on historical data from DNA and RNA sequencing of subjects with a certain type of cancer. Furthermore, RNA expression levels are used to determine whether the somatic mutation is expressed as RNA and therefore may be able to produce a protein.
[0143] Another parameter is the IC50 value for somatic mutations in HLA proteins. In this example, a low IC50 indicates that the somatic mutation will bind strongly to the subject's HLA. If the somatically mutated protein binds to HLA better than the germline sequence protein, then the somatically mutated protein is more likely to shuttle to the cell surface, making it detectable by the immune system.
[0144] Another parameter is the difference in fold change between the somatic mutant sequence and the germline sequence for the subject. Vaccine candidates may possess a somatic mutant sequence that is found to bind better than the germline DNA sequence for the subject. The folding difference between the somatic mutant protein and the protein in the germline DNA sequence distinguishes the somatic mutant from the "self" protein displayed on the cell surface. The immune system is trained to ignore "self" proteins in order to prevent autoimmune diseases. Therefore, when used as a vaccine, a larger folding difference between the somatic mutant protein and the protein in the germline DNA sequence for the subject will elicit more immune stimulation. In some instances, the greater the difference between the somatic mutant protein and its self-allelic counterpart, the greater the likelihood that the immune system can detect a protein that is more different from its own protein.
[0145] Another parameter is the difference in protein sequence between somatic mutations and the general proteome. The folding differences between somatic mutation protein sequences and proteins in germline DNA sequences distinguish somatic mutations from other “self” proteins displayed on the cell surface. Therefore, greater differences between somatic mutation proteins and other proteins in germline DNA can improve the immune stimulation of somatic mutations. Another parameter is the degree to which T cells respond to somatic mutations.
[0146] The machine learning model described in this article can generate a score for each somatic mutation based on a set of parameters (such as those described above). These parameters may have a stronger weight in the machine learning model than other parameters for the subject.
[0147] In one instance, a machine learning model identifies the most prevalent and highest-scoring somatic mutations in a neoantigen for a particular type of cancer. For example, the most prevalent and highest-scoring somatic mutations for a particular type of cancer can be identified by analyzing data from 1,000 subjects treated using a machine learning model.
[0148] For somatic mutations present in multiple subjects targeting different cancer types, a machine learning model assigns scores to somatic mutations based on a set of parameters and ranks them to determine which somatic mutations are most likely to infer a survival benefit and / or function as a vaccine for a particular type of cancer. A catalog of somatic mutations identified as vaccine candidates is compiled, and a pre-made vaccine library for these candidates is generated.
[0149] The resulting vaccines can include peptide-based vaccines, mRNA vaccines, or other types of vaccines. The somatic mutations used to generate the vaccine are somatic mutations already identified based on historical subject data, used to stimulate the immune system using machine learning models. Additional adjuvants can be added to peptide-based or mRNA-based vaccines to improve the immune system's response to the vaccine, including those against metals, bacteria, and derived toxic peptides.
[0150] Once a machine learning model is developed, it can be used to predict vaccines for new subjects with a certain type of cancer. In one instance, a new subject might be eligible for one or more pre-made vaccines based on their somatic mutation. The new subject's somatic mutation is identified and compared to a catalog of somatic mutations against a specific type of cancer, compiled by the machine learning model. In some instances, the new subject's HLA genotype is identified, and it is determined whether the HLA genotype is compatible with pre-made vaccines in a library targeting somatic mutations. Vaccines are then identified for the new subject and administered. In some instances, a certain percentage of new subjects might be eligible for pre-made vaccines in a vaccine library, while others might not be.
[0151] Vaccine bank Figure 6 The system 10 showcases a vaccine library 20, a vaccine management application 100, a data storage device 200, a prediction engine 700, and new subject data 500.
[0152] Vaccine repository 20 includes pre-prepared vaccines 35 stored in a housing facility that includes a suitable temperature control system to maintain the vaccines under feasible conditions. Exemplary pre-prepared vaccines 35 can be various types of vaccines, including peptide-based synthetic vaccines (epitope vaccines), messenger RNA (mRNA) vaccines, and conventional vaccines. Additional vaccine types may include gene therapy and editing, such as gene editing using adeno-associated virus (AAV) and adenomatous polyposis (APC).
[0153] Vaccine repository 20 includes vaccine catalog 30. Vaccine catalog 30 includes vaccines that have been cataloged according to predetermined characteristics against somatic mutations (neoantigens). A somatic mutation is a variation in the DNA sequence of a subject between a germline sequence and a sequence from a tumor sample of the subject. Predetermined characteristics against somatic mutations may include major histocompatibility complex (MHC), genotype information (e.g., single nucleus polymorphisms, copy number variants (CNVs), insertions and deletions, gene fusions, structural variants or combinations thereof, and SNPs of specific nucleic acid sequences associated with genes or genomic or mitochondrial DNA). Cataloging may constitute the creation of a centralized record or database of characteristics obtained for each somatic mutation.
[0154] The vaccine repository facilitates the selection of pre-prepared vaccines suitable for administration to subjects suffering from a particular type of cancer from multiple samples. The vaccine repository catalog can be stored on a computer or server located at or away from the vaccine repository 20. The vaccine catalog 30 can be transmitted using any desired communication protocol and via any type of networking technology with any desired network topology.
[0155] The vaccine management application 100 communicates with the catalog 30, the data storage device 200, and the prediction engine 700. The vaccine management application 100 may include a browser-based graphical user interface (GUI) or a command-line interface (CLI) to query the catalog 30.
[0156] The prediction engine 700, described in more detail below, includes a method and system for identifying a library or catalog of 10-100 somatic mutations for neoantigens used in personalized cancer vaccines using one or more models. The pre-prepared vaccine 35 is generated against the identified somatic mutations of the neoantigen and can be for various types of vaccines, including peptide-based synthetic vaccines (epitope vaccines), messenger RNA (mRNA) vaccines, and conventional vaccines. In some instances, for the mRNA pre-prepared vaccine 35 generated against the identified somatic mutations, the neoantigen can be stored in a library using liquid nanoparticles (LNPS). In other instances, the pre-prepared vaccine and the LNP mRNA can be stored separately in a library and combined when new subjects are identified for administration of the pre-prepared vaccine.
[0157] In other instances, pre-prepared vaccines can be stored in a library as polymer nanostructures of peptide, mRNA, or gene-edited DNA vaccines. Somatic mutations of neoantigens from the genome can be stored separately from somatic mutations of neoantigens from the exome. In some instances, neoantigens and LNPs are stored separately and combined when new subjects are identified for administration of the pre-prepared vaccine.
[0158] In one instance, prediction engine 700 identifies somatic mutations (neoantigens) by identifying clusters of subjects who tend to have one or more common neoantigens. Prediction engine 700 identifies clusters of subjects with a certain type of cancer that tends to have one or more common neoantigens, for example, by looking at parameters for the subjects (such as the clonal evolution of the tumor) and / or using clustering algorithms. Other parameters used for prediction may include copy number variation (CNV) and subdomains to detect the clonal nature of somatic mutations in the subjects.
[0159] In one instance, based on the identification of somatic mutations in neoantigens by the prediction engine 700, a catalog 30 and a pre-made vaccine library 35 are generated, which cover the most immunogenic somatic mutations of neoantigens in the centroid of each cluster of subjects with a certain type of cancer.
[0160] For new subjects with a certain type of cancer, matching subjects with pre-prepared vaccines can be achieved by identifying the cluster to which the subject belongs. In one instance, each new subject's immune response score to available pre-prepared vaccines is determined by identifying which cluster the patient likely belongs to and / or predicting the patient's immune response score to each available pre-prepared vaccine, thus matching the new subject with the pre-prepared vaccine. Pre-prepared vaccines can also be cataloged based on genomic or exome DNA sequencing of somatic mutations.
[0161] New participant data 500 can be remotely received by the vaccine administration application 100. New participant data 500 may include germline (natural DNA) and DNA and RNA sequencing data from tumor samples of the new participants. In some instances, the DNA and RNA sequencing data may be whole-exome sequencing (WES) germline sequencing data for the participants, WES for the tumor samples of the new participants, and RNA sequencing for the tumor samples of the new participants. In some instances, the new participant data is not part of the historical participant data used to train the predictive model, but is used in conjunction with the trained predictive model to predict the vaccine to be administered to the new participants.
[0162] It can monitor vaccine administration to subjects from a vaccine bank and collect additional data. This data may include the vaccine's efficacy in stimulating the immune system, treatment outcomes for subjects, additional DNA and RNA data, including identification of subclonal variants, copy number and variant allele frequency (VAR), and information on minimum disease resistance (MRD).
[0163] The data storage device 200 can be maintained by the vaccine management application 100 or can be maintained independently. Various other data can also be stored in the data storage device 200.
[0164] New subject data 500 can be received by vaccine administration application 100 and used by prediction engine 700 with prediction model to determine or predict what vaccine to administer to new subjects.
[0165] Prediction Engine Figure 7 A more detailed example of Prediction Engine 700 is shown. For example... Figure 7 As shown, the prediction engine 700 includes processing resources 701 (e.g., a processor, CPU, GPU, SoC, or other processing resources) and storage medium 705 (e.g., a non-transitory computer-readable medium). Storage medium 705 includes model training instructions 710 and model prediction instructions 720. Model training instructions 710 are executable to train a neoantigen candidate model, an immune stimulation model, and a neoantigen candidate ranking model, while model prediction instructions 720 are executable to predict neoantigen candidates to be used as a vaccine to be administered to subjects with a certain type of cancer.
[0166] The prediction engine 700 uses historical datasets from the data storage device 200 to train neoantigen candidate models, immune stimulation models, and neoantigen candidate ranking models. Model training instructions 410 include instructions 711 for loading historical subject data from the data storage device 200. Model training instructions 710 also include instructions 712 for determining the parameters of the machine learning model.
[0167] The data used to train the machine learning model (in this case, data from data storage device 200) is prepared and loaded for model development. The data can be added directly to the same directory as the machine learning model's code. The data can be uploaded as part of an experimental directory. The data can also be obtained using a distributed file system to allow machine clusters to access shared datasets. The data can also be provided as an object storage device for managing the data.
[0168] These parameters may include somatic mutations in subjects with a certain type of cancer, subject response to treatment (such as checkpoint inhibitors), fold change differences between native proteins in the WES lineage and somatic mutant proteins from tumor samples, protein expression levels of candidate neoantigens (in some cases, TCGA25), MHC-I: mutant peptide processing and presentation score (IQ), T cell activation, and dissimilarity to the reference human proteome.
[0169] Other historical data for the subjects include the treatments used on the subjects, the checkpoint inhibitors used for the treatment, and subject outcome data in response to the treatment. Immune checkpoint inhibitors used for the treatment may include PD-1 inhibitors, PD-L1 inhibitors, CLTA-4 inhibitors, or combinations thereof.
[0170] In this example, the prediction engine 700 compares germline WES sequencing data of the subject with WES sequencing data from a tumor sample from the same subject to identify changes in DNA sequencing, such as somatic mutations. Somatic mutations in neoantigens can include single nucleotide variants (SNVs), multinucleotide variants (MNVs), copy number variants (CNVs), insertions / deletions, gene fusions, structural variants, aberrant splicing variants, or combinations thereof. DNA sequencing data obtained from the tumor sample is compared with DNA obtained from normal subject tissue to identify somatic mutations that are present in the tumor sample but not in the subject's germline. Somatic mutations in a patient's tumor sample cause changes in protein sequences.
[0171] In some instances, WES sequencing data from tumor samples are compared to a reference genome to identify somatic mutations and germline variants that differ from a normal reference genome. An individual may possess both somatic mutations and germline variants (which are normal but differ from the reference genome). Immune response prediction models can predict amino acid sequences containing both types of variants. Somatic mutations are phased together with germline variants to determine the mutation-germline allele haplotype. When creating personalized cancer vaccines for individuals, both normal germline variants (such as normal germline variants that differ from a normal reference genome) and somatic mutations are considered.
[0172] In some instances, RNA sequencing (RNA-seq) is performed to determine the presence and quantity of RNA in a biological sample, representing a aggregated snapshot of the cellular dynamic RNA library (also known as the transcriptome). In these instances, available RNA sequencing methods include whole transcriptome sequencing, ribosomal RNA depletion sequencing, and sequencing of targeted gene expression within RNA. In some instances, RNA sequencing can be inferred when RNA sequencing data is unavailable.
[0173] Predictive models can be trained in a machine learning development environment. An exemplary machine learning development environment that can be used to train a predictive model has one or more professional agents and pilots. The host, agents, and pilots can reside on a single computer server or can be distributed across computer servers in a cloud computing environment.
[0174] The host is the central component and is responsible for storing experimental, trial, and historical subject data from data storage device 200. The host schedules and dispatches tasks for agents. The host can also manage and deconfigure agents within the cloud environment. The host can advance the state machine of experiments, trials, and workloads over time.
[0175] An agent manages multiple slots, which are computing devices, typically central processing units (CPUs). Agents are stateless and communicate with the host. Each agent is responsible for discovering local computing devices (slots) and sending data about them to the host. Agents run workloads requested by the host. Agents monitor containers and send information about them to the host.
[0176] Pilots run experiments in a containerized environment. Pilots are expected to have access to the data that will be used in training. An agent is responsible for reporting the pilot's status to the host. The machine learning development environment is prepared by determining the CPUs used to train the model. The container can be the default container or a custom one.
[0177] The machine learning environment described above is one example of a machine learning environment that can train the neoantigen candidate model, immune stimulation model, and neoantigen candidate ranking model of the prediction engine 700, but it should be understood that any other desired machine learning environment can be used.
[0178] The model training code can be converted to utilize APIs within a machine learning development environment. Hyperparameters can be fine-tuned to select data, features, model architecture, and learning algorithms to produce a predictive model.
[0179] Although for ease of description in Figure 7The diagram shows model training instructions 710 and model prediction instructions 720 as part of the same storage medium 705, but in some instances these may be provided separately. For example, prediction engine 700 may include multiple computing systems, including one or more systems for training neoantigen candidate models, immune stimulation models, and one or more systems for neoantigen candidate models, immune stimulation models, and neoantigen candidate ranking models.
[0180] In some instances, the system used for training may have model training instructions 710 but not model prediction instructions 720, and conversely, the system used for prediction may have model prediction instructions 720 but not model training instructions 710. In other instances, the prediction engine 700 may include one or more computing systems that include both model training instructions 710 and model prediction instructions 720.
[0181] Neoantigen candidate selection model Figure 8 A bioinformatics pipeline for identifying and prioritizing candidate immunogenic neoantigens using a combination of standard and newly developed algorithms and models is described. At point 810, candidate neoantigens are identified. Tumor and matched normal WES samples are aligned to a reference genome and phased to identify coding sequence alterations (including both germline and somatic variants). WES data are also used for HLA-I and HLA-II haplotype analysis. Variants and HLA haplotypes are evaluated together to determine MHC-I and MHC-II binding and presentation. Finally, T cell activation is predicted. Although parameters for MHC binding prediction, T cell activation, and HLA 1 haplotype analysis are used to predict neoantigen candidates, it should be understood that these parameters can be used in any combination, and additional parameters can be used with Prediction Engine 700 to predict neoantigen candidates.
[0182] In order to identify candidate neoantigens, Figure 7 The prediction engine 700 compares WES sequencing data of a subject's lineage with WES sequencing data of tumor samples from the same subject to identify changes in DNA sequencing, such as somatic mutations, in a group of subjects diagnosed with a certain type of cancer. In an embodiment, the prediction engine 700 performs these steps for each subject in a group diagnosed with a certain type of cancer from a sample set. This includes the number and abundance of unique T and B cell receptors from the subject's RNA sequencing data, and the enrichment of immune cell populations in the tumor samples.
[0183] Prediction Engine 700 determines the number of base pairs in somatic mutations of neoantigens that differ from the normal reference sequence. In this example, Prediction Engine 700 uses aligned tumor and normal BAM files via a invoked algorithm to identify tumor and germline common single nucleotide variants (SNVs) and insertions / deletions (indels). Somatic SNVs and insertions / deletions are phased with germline variants to determine mutant germline allele haplotypes and are then annotated (VEP v109) to determine the highest impact consequence for each mutation. Genotyping of MHC class I (MHC-I) alleles is performed using OptiType v1.3.3. Genotyping of MHC class II (MHC-II) alleles is performed using HISAT genotyping v1.3.3.
[0184] In some instances, Prediction Engine 700 utilizes RNA sequencing data from tumor samples to filter out somatic mutations in unexpressed genes. Prediction Engine 700 identifies the RNA sequence of the somatic mutation. In one instance, Prediction Engine 700 uses a computer-simulated translation tool. The translation tool translates DNA sequence changes into protein sequences and identifies the number of RNA sequences associated with the somatic mutation. Peptide-MHC binding prediction can then be performed in two steps to determine whether the somatic mutation will bind to the individual's major histocompatibility complex (MHC). First, all mutations altering the amino acid sequence, except for the MHC-I and MHC-II alleles, are fed into pVAC-Seq v4.0.5. pVAC-Seq is a pipeline that generates mutant and germline peptide sequences within known MHC binding lengths, produces all MHC-peptide combinations, executes many published peptide-MHC binding prediction tools, and compiles peptide-MHC binding prediction metrics. Prediction Engine 700 then determines whether T cells will bind to the MHC peptide complex.
[0185] In this example, artificial intelligence (AI) can be used to assess the genetic regions surrounding somatic mutations in neoantigens to predict neoantigen expression levels and binding. It should be understood that any type and combination of AI / machine learning algorithms, publicly available tools, open-source tools, and commercial products can be used for parameter selection in neoantigen candidate selection models, immune stimulation models, and / or neoantigen candidate ranking models, and is expected to be used in a similar manner. Those skilled in the art will understand that various commercial models are suitable for incorporation into the application, such as Google... TM ALPHA FOLD, IntFOLD, RaptorX, HHpred, Phyre, Phyre2 and I-TASSER.
[0186] Parameter selection refer to Figure 8Once a neoantigen candidate is selected, at position 820, the prediction engine at position 700 determines additional parameters (or classifiers) for the neoantigen candidate. These parameters include TCGA25, MT IC50 (MHC-1: mutant peptide processing and presentation score (IQ)), T cell activation prediction score, MHC-I binding fold change (MHC-I binding fold change log, (WT IC, 0)), and dissimilarity to a reference proteome. Parameters described as having demonstrated effects on different aspects of immunogenicity are applied and assigned scores. Once scores are assigned, candidate neoantigens with matching properties are grouped together and summed.
[0187] Any suitable combination of these parameters can be used to train a machine learning model. This pipeline incorporates combinations of parameters that can be used for predictions using neoantigen candidate models, immune stimulation models, and neoantigen candidate ranking models. The prediction engine 700 described herein incorporates a variety of relevant parameters, including any combination of the parameters described above, as well as additional parameters for predicting MHC / peptide binding to identify neoantigen candidates, immune stimuli, and / or candidate neoantigens used as vaccine candidates. Other parameters may include clonal evolution from tumor DNA sequencing and / or parameters using clustering algorithms.
[0188] The prediction engine 700 uses historical datasets of a group of subjects with a certain type of cancer to determine the characteristics or parameters of the subjects used to train the machine learning model. In some embodiments, the group of subjects with a certain type of cancer has received cancer treatment.
[0189] When RNA sequencing data is unavailable, prediction engine 700 can infer RNA expression. RNA expression databases (such as the Cancer Genome Atlas (TCGA)) can be used to infer an individual's gene expression. Prediction engine 700 determines the percentage of individuals in the RNA expression database who are expressing proteins of genes that are candidates for neoantigens. For example, if an individual from the RNA expression database expresses a specific gene at or above the 25th percentile, then RNA expression can be inferred. Prediction engine 700 can also use RNA expression databases to rank neoantigen characteristics. For example, it can determine the proportion of genes (such as somatic mutations) expressed by individuals with cancer. Exemplary parameters for determining whether RNA is expressed can include the expression level of somatic mutations. This can be determined based on DNA and RNA sequencing data from a subject with a certain type of cancer. Furthermore, RNA expression levels are used to determine whether somatic mutations are expressed as RNA and therefore may be able to produce proteins. In this example, using DNA and genetic motif data, commercially available AI tools can be used to predict RNA expression levels.
[0190] Prediction Engine 700 also determines the MT IC50 (MHC-1: mutant peptide processing and presentation score (IQ)) of neoantigen candidates to predict those neoantigen candidates that strongly bind to an individual's MHC allele and / or are immunogenic (e.g., IC50). 50 <600 nM, or IC 50 <550 nM, or IC 50 <500nM, or IC 50 <450 nM, or IC 50 <400 nM; and / or percentile rank <0.6%, or percentile rank <0.55%, or percentile rank <0.5%, or percentile rank <0.45%, or percentile rank <0.4%.
[0191] The IC50 value of a subject's HLA protein against a somatic mutant sequence can be used. In this example, a low IC50 indicates that the somatic mutation will bind strongly to the subject's HLA. If the somatic mutant protein binds to HLA better than the germline sequence protein, then the somatic mutant protein is more likely to shuttle to the cell surface, making it detectable by the immune system.
[0192] Prediction Engine 700 determines the binding ratio of mutant peptides to wild-type peptides (Log2 WT IC50 / MT IC50). Compared to wild-type peptide binding, Prediction Engine 700 favors mutant peptides with stronger MHC binding.
[0193] Prediction Engine 700 determines the T-cell activation score of neoantigen candidates. The T-cell activation score can be determined by quantifying the number of unique T and B cell receptors, including: (i) deconvolution of the proportion of immune cells in the tumor sample based on RNA sequencing data, and (ii) assembly of B and T cell receptors to quantify the number of unique T and B cell receptors. The T-cell activation score can be ranked by percentage and RNA expression level.
[0194] Another parameter is the difference in fold change between the somatic mutant sequence and the germline sequence for the subject. Vaccine candidates may possess a somatic mutant sequence that is found to bind better than the germline DNA sequence for the subject. The folding difference between the somatic mutant protein and the protein in the germline DNA sequence distinguishes the somatic mutant from the "self" protein displayed on the cell surface. The immune system is trained to ignore "self" proteins in order to prevent autoimmune diseases. Therefore, when used as a vaccine, a larger folding difference between the somatic mutant protein and the protein in the germline DNA sequence for the subject will elicit more immune stimulation. In some instances, the greater the difference between the somatic mutant protein and its self-allelic counterpart, the greater the likelihood that the immune system can detect a protein that is more different from its own protein.
[0195] Another parameter is the difference in protein sequence between somatic mutants and the general proteome. The folding differences between somatic mutant protein sequences and proteins in germline DNA sequences distinguish somatic mutants from other “self” proteins displayed on the cell surface. The immune system is trained to ignore “self” proteins to prevent autoimmune diseases. Therefore, greater differences between somatic mutant proteins and other proteins in germline DNA can improve the immune stimulation of somatic mutants.
[0196] Random forest algorithms can be used for parameter selection and model training. While random forest algorithms are used in this example, it should be understood that any type and combination of AI / machine learning algorithms, publicly available tools, open-source tools, and commercial products can be used for parameter selection in neoantigen candidate selection models, immune stimulation models, and / or neoantigen candidate ranking models, and will be expected to be done in a similar manner. Those skilled in the art will understand that various commercial models are suitable for incorporation into the application, such as Google... TM ALPHA FOLD, IntFOLD, RaptorX, HHpred, Phyre, Phyre2 and I-TASSER.
[0197] Immune response model refer to Figure 8 At point 830, the Prediction Engine 700 developed a model to predict which patients are more likely to progress or not progress after ICI therapy. This model incorporates various clinical parameters, including baseline ctDNA results, as well as candidate neoantigens based on mutant IC50, changes in MHC-I binding folds, dissimilarity of candidate neoantigens to the reference human proteome, and T-cell activation prediction scores.
[0198] A machine learning tool is provided that predicts individual responses to immune checkpoint inhibitor (ICI) therapy based on a large number of individuals with complete clinical information. By utilizing whole-exome sequencing data, ctDNA timepoint analysis, and clinical data, a deep learning-based immune response prediction model successfully predicted which individuals would respond to and benefit from ICI therapy across multiple cancer types. Individuals responding to ICI therapy may have multiple or strong neoantigens.
[0199] exist Figure 7 In the example shown, the immune stimulation model incorporates the random survival forest algorithm from machine learning. The random forest algorithm is a non-parametric machine learning strategy that can be used to construct risk prediction models in survival analysis. In a survival setting, the predictors are a set formed by combining the results of many survival trees. While the random forest algorithm is used in this example, it should be understood that any suitable type and combination of machine learning algorithms can be used for the immune stimulation model.
[0200] Although the random forest algorithm was used in the example, it should be understood that any type and combination of AI / machine learning algorithms, publicly available tools, open-source tools, and commercial products can be used for parameter selection in neoantigen candidate selection models, immune stimulation models, and / or neoantigen candidate ranking models, and will be expected to be done in a similar manner.
[0201] Neoantigen sequencing model refer to Figure 8 At step 840, candidate neoantigens are ranked using a machine learning model based on ICI response, TCGA25, mutant IC50, and T cell activation prediction scores. ICI responses in the training set are evaluated using a set of matched scores for these three characteristics in the context of the neoantigen. In some embodiments, the ICI response can be determined using the immune stimulation model described above.
[0202] Candidate neoantigens are assigned weighted scores based on the association between the neoantigen and the ICI response, and these scores are then used to rank the neoantigens. As described above, the prediction engine 700 uses a machine learning model to determine which somatic mutations are most likely to infer survival benefits and / or be used as vaccines against a certain type of cancer, thereby assigning scores to the somatic mutations of neoantigens (candidate neoantigens) and ranking them based on a set of parameters. A catalog of somatic mutations identified as vaccine candidates is compiled, and a pre-made vaccine library of vaccine candidates is generated.
[0203] One or more parameters can be used to predict candidate neoantigens for use as vaccine candidates. Any combination of the parameters described herein can be used to predict and rank candidate neoantigens.
[0204] Model training instructions 710 include instructions 713 for determining scores to predict the probability of certain categories (neoantigen candidates, immune stimulation, and neoantigen ranking for vaccine candidates) based on a combination of parameters. In this example, a random survival forest can determine the scores for neoantigens that qualify as vaccine candidates. Somatic mutations that most likely qualify as immune stimulation and vaccine candidates are 714. While a random forest algorithm is used in this example, it should be understood that any type and combination of AI / machine learning algorithms, publicly available tools, open-source tools, and commercial products can be used for parameter selection in neoantigen candidate selection models, immune stimulation models, and / or neoantigen candidate ranking models, and will be expected to be done in a similar manner. Those skilled in the art will understand that various commercial models are suitable for incorporation into the application, such as Google. TM ALPHA FOLD, IntFOLD, RaptorX, HHpred, Phyre, Phyre2 and I-TASSER.
[0205] Vaccine Catalog Model training instructions 710 include instructions 715 for outputting a catalog of somatic mutations identified as immune stimuli and vaccine candidates. For example, somatic mutations that appear in multiple subjects with a certain type of cancer and are ranked highly by a machine learning model are categorized as immune stimuli and vaccine candidates. This catalog may include DNA and RNA sequencing data, parameters for selecting somatic mutations, and scoring and ranking information.
[0206] In some instances, some of the data from data storage device 200 can be used for training, while other data can be used for testing and validation. In one instance, 80% of the dataset from data storage device 200 can be used to train a gradient boosting regression algorithm, while the remaining 20% can be used for immune stimulation and vaccine candidate prediction from the prediction model.
[0207] In some instances, users can determine when a model is adequately trained based on test errors or biases, for example, using a test set bias graph. In other instances, the prediction engine 700 can be configured with logic to determine when a model is adequately trained. For example, the prediction engine 700 can identify a model as adequately trained when the test error rate is equal to or below a threshold (or consistently equal to or below a threshold for a defined number of test runs to account for variance in the errors).
[0208] In some instances, users can specify which features (parameters) should be removed, for example, using the aforementioned feature importance map. In other instances, the prediction engine 700 can be configured with logic to select which features to remove. For example, all features with a correlation below a defined correlation threshold can be omitted.
[0209] Vaccine prediction for new subjects Refer again Figure 7 When new subject data for a certain type of cancer is received, prediction engine 700 is used to predict immune stimulation and vaccine candidates.
[0210] In one instance, the vaccine administration application 100 makes a Representational State Transfer (REST) application programming interface (API) call to the prediction engine 700. The payload of the REST API call includes data on immune stimulation and vaccine prediction for each new subject.
[0211] In this example, in response to an API call, model prediction instruction 720 is executed. Model prediction instruction 720 includes instruction 721 for loading new subject data. New subject data 500 may include germline (natural DNA) and DNA and RNA sequencing data of the new subject's tumor sample. In some examples, the DNA and RNA sequencing data are whole-exome sequencing (WES) germline sequencing data for the subject, WES for the new subject's tumor sample, and RNA sequencing for the new subject's tumor sample. In this example, the new subject data is not part of the historical subject data used to train the prediction model, but is used in conjunction with the trained prediction model to predict immune stimulation and candidate vaccines to be administered to the new subject.
[0212] The model prediction instruction 720 further includes instruction 722 for identifying somatic mutations in a new subject by comparing data from the WES germline with data from the new subject's WES tumor sample. The model prediction instruction 720 further includes instruction 723 for selecting somatic mutations for the new subject that are homologous to somatic mutations cataloged by the machine learning model. In some instances, the new subject's somatic mutations are 70%, 80%, 90%, or 99% homologous to one or more somatic mutations cataloged by the machine learning model.
[0213] In one instance, based on the identification of somatic mutations (neoantigens) by the prediction engine 700, a catalog 30 and a pre-made vaccine library 35 are generated, which cover the most immunogenic somatic mutations (neoantigens) in the centroid of each cluster of subjects with a certain type of cancer.
[0214] For a new subject with a certain type of cancer, the subject can be matched with pre-prepared vaccines 35 by identifying the cluster to which the subject belongs. In one instance, each new subject's immune response score to available pre-prepared vaccines is determined by identifying which cluster the patient might belong to and / or predicting the patient's immune response score to each available pre-prepared vaccine, in order to match the new subject with a pre-prepared vaccine. Instruction 424 outputs predicted immune stimuli and vaccine candidates for one or more somatic mutations in the new subject. The new subject's somatic mutations, DNA sequencing, and RNA sequencing are fed into a trained model to determine the predicted immune stimuli and vaccine candidates for the new subject. In other words, new subject data from the REST API payload is loaded into a trained predictive model to generate predicted immune stimuli and vaccine candidates.
[0215] method Figure 9 This is a flowchart illustrating a method for training an immune stimulation prediction model according to the examples described herein. This method can be executed by any suitable processor or other hardware discussed herein (e.g., the processor or hardware included in prediction engine 700). Specifically, in some instances, prediction engine 700 or components thereof are instantiated by one or more processors executing machine-readable instructions, which at least partially include... Figure 9 The corresponding instructions for the operation of the method.
[0216] The method begins at step 920, where subject data from data storage device 200 is loaded or fed for use by prediction engine 700. This data includes historical data of subjects undergoing treatment for a certain type of cancer.
[0217] At step 930, prediction engine 700 defines the parameters of the machine learning model. In some instances, feature importance is plotted by prediction engine 700. Prediction engine 700 determines which subject parameters (features) meet a relevance threshold and are predictive features for neoantigen ranking used for neoantigen candidate selection, immune stimulation, and vaccine candidate eligibility based on the determined mathematical relationships. Prediction engine 700 determines the mathematical relationships between parameters and neoantigens by feeding historical subject data as input into one or more machine learning models.
[0218] At step 940, a score for each of the neoantigens in at least some of the subjects is determined using data from subjects who have previously been treated for a certain type of cancer. The scores for each of the immune stimulation parameters in at least some of the subjects are aggregated, and an aggregated score for candidate neoantigens appearing in a subset of the subjects is generated.
[0219] At step 950, each of the neoantigens appearing in a subset of subjects is ranked according to its aggregate score. Somatic mutations are ranked based on predicted probabilities. Somatic mutations and vaccine candidates with the highest probability of immune stimulation are ranked highest.
[0220] At step 960, the predicted somatic mutations of the neoantigen are determined. Somatic mutations of the neoantigen can be determined based on data from subjects treating a certain type of cancer, processed by comparing tumor DNA sequencing of the subjects with germline DNA sequencing to identify somatic mutations in the subjects. Subject data can also be preprocessed by formatting the prediction engine, checking for completeness and bias, checking and estimating missing values, and smoothing and binning the data.
[0221] At step 970, a catalogue of somatic mutations against neoantigens that are both immune-stimulating and vaccine candidates is output. For example, somatic mutations of neoantigens that appear in multiple subjects with a certain type of cancer and are ranked highly by a machine learning model are categorized as both immune-stimulating and vaccine candidates. This catalogue may include DNA and RNA sequencing data, parameters for selecting somatic mutations for neoantigens, and scoring and ranking information.
[0222] use Figure 9 The process involves generating a trained model. This trained model can then be used to predict immune stimulation in new subjects and vaccine candidates, such as... Figure 10 As described in [the text].
[0223] Figure 10 This is a flowchart illustrating a method for predicting immune stimulation and vaccine candidates for new subjects, based on examples described herein. This method can be executed by any suitable processor or other hardware discussed herein (e.g., the processor or hardware included in prediction engine 700). Specifically, in some instances, prediction engine 700 or components thereof are instantiated by one or more processors executing machine-readable instructions, which at least partially include... Figure 10 The corresponding instructions for the operation of the method.
[0224] The method begins at step 1010, where the prediction engine 700 receives a request from the vaccine administration application. At step 1020, data from new subjects is received. New subject data 500 may include germline (natural DNA) and DNA and RNA sequencing data from tumor samples of the new subjects. In this example, the DNA and RNA sequencing data may be whole-exome sequencing (WES) germline sequencing data for the subject, whole-genome sequencing, cancer genome sequencing, WES of tumor samples from the new subject, and RNA sequencing of tumor samples from the new subject.
[0225] In step 1030, prediction engine 700 identifies predicted immune stimulation and vaccine candidates for one or more somatic mutations in the new subject. New subject data is fed as input into the immune stimulation prediction model generated by prediction engine 700.
[0226] In step 1040, the prediction engine 700 identifies from the catalog somatic mutations that are homologous to one or more somatic mutations for a new subject in a subject undergoing treatment for a certain type of cancer.
[0227] In step 1050, prediction engine 700 generates and formats predicted immune stimulation and vaccine candidates for one or more somatic mutations in a new subject. In one instance, the predicted immune stimulation and vaccine candidates are returned as payload data. Vaccine candidates can then be selected from a vaccine library containing pre-made vaccine libraries targeting somatic mutations in different types of cancer.
[0228] Example 2 In another instance, patients who responded well to current immunotherapies were rich in immunogenic neoantigens, and their response to ICIs provided better neoantigen prediction and prioritization. Furthermore, these neoantigen characteristics helped predict patient responses to ICIs.
[0229] like Figure 11 As shown, candidate neoantigen selection and prioritization for PCV design, as well as ICI response prediction and prioritization, were performed. Tumor and matched normal WES were used as data for identifying candidate neoantigens. Candidate neoantigens were binned according to MHC-I binding, predicted T-cell immunogenicity, and gene expression (TCGA-based). ICI response was measured by progression-free survival (from a real-world clinical cohort). A strong correlation between predicted neoantigen immunogenicity and ICI outcomes was observed using the Cox proportional hazards ratio.
[0230] like Figure 12 As shown, the performance of neoantigen prioritization using neoantigen prediction methods was evaluated using publicly available benchmark datasets. The NCI1, TESLA2, and HiTIDE3 datasets were used with known immunogenic neoantigens, as evaluated via in vitro experiments. When prioritizing the top 20 neoantigens (the common number of neoantigen targets included in the current PCV), approximately half of the true positive neoantigens were identified using neoantigen prediction and prioritization methods. Figure 13 The box plot for each patient depicts the proportion of true positive neoantigens prioritized in the top 20 by neoantigen prediction and prioritization methods.
[0231] like Figure 14As shown, the neoantigen prediction and prioritization method described in this paper performs better than other prediction models. The results of the neoantigen prediction and prioritization method described in this paper are compared with other neoantigen prioritization methods in the TESLA cohort, which includes five patients (three with melanoma and two with NSCLC).
[0232] Next reference Figure 15 In this example, the inclusion of patient-matched RNAseq data had almost no impact on neoantigen prioritization performance. Figure 15 This depicts the number of true positive neoantigens selected using neoantigen prediction and prioritization models trained with and without patient-matched RNA. In some cases, population-based data from TCGA are sufficient for gene expression inference using neoantigen prediction and prioritization models. Although in other instances, RNA data can enhance neoantigen and prioritization methods.
[0233] like Figure 16 As shown, the neoantigen prediction and prioritization method provides better ICI response prediction than using TMB alone. In this example, neoantigen load is the most predictive factor in classifying patients as responders and non-responders. Kaplan-Mayer estimates of PFS indicate that the neoantigen prediction and prioritization method classifies patients as responders and non-responders more accurately than using TMB.
[0234] like Figures 17A to 17H As shown, the neoantigen prediction and prioritization method is informed by the in vivo ICI response. This method effectively identifies immunogenic targets designed for PCV and can be used to avoid ICI treatment in patients who are unlikely to benefit.
[0235] The following provides illustrative examples of the systems, methods, compositions, and computer products described herein. Embodiments of the systems, methods, and / or computer products described herein may include any one or more of the aspects / embodiments / features described below in any order and / or any combination thereof: 1. A system comprising: at least one computing device, the at least one computing device including at least one processor, the at least one processor being configured to: feed data from subjects into a prediction model; score a neoantigen appearing in data from a subset of the subjects for one or more parameters; determine one or more neoantigens appearing in the subset of subjects that satisfy an immune stimulation threshold based on the scores for the one or more parameters; determine somatic mutations for the one or more neoantigens appearing in the subset of subjects that satisfy the immune stimulation threshold; and output a prediction catalog of one or more of the somatic mutations appearing in the subset of subjects that satisfy the immune stimulation threshold, the prediction catalog predicting the somatic mutations.
[0236] 2. The system of claim 1, wherein the data comes from a subject who has previously been treated for a certain type of cancer.
[0237] 3. The system of claim 1, wherein the prediction model is more predictive of immune stimulation for a subset of the subjects than tumor mutational burden (TMB) status.
[0238] 4. The system according to claim 2, wherein the prediction model is a machine learning model.
[0239] 5. The system of claim 4, wherein the one or more parameters include one or more of the following: peptide processing and presentation, RNA expression, MHC binding fold change, T cell activation, and dissimilarity to a reference human proteome.
[0240] 6. The system of claim 5, wherein the parameters further include one or more of the following: immune checkpoint inhibitor (ICI) response, ctDNA result, age, sex, and ECOG score.
[0241] 7. A system comprising: at least one computing device, the at least one computing device including at least one processor, the at least one processor being configured to: feed data for a newly diagnosed subject with a certain type of cancer into a predictive model; score neoantigens appearing in the data for the subject against one or more parameters; determine one or more neoantigens appearing for the subject that satisfy an immune stimulation threshold based on the scores against the one or more parameters; determine somatic mutations of the one or more neoantigens appearing for the subject that satisfy the immune stimulation threshold; and output one or more vaccines to be administered to the subject in response to the somatic mutations of the one or more neoantigens that satisfy the immune stimulation threshold, wherein the one or more vaccines to be administered to the subject are selected from a predictive catalog of somatic mutations appearing in a subset of subjects previously treated for the type of cancer.
[0242] 8. The system of claim 7, wherein one or more vaccines are selected from a pre-prepared vaccine library to be administered.
[0243] 9. The system of claim 8, wherein the vaccine is one or more of the following: a peptide-based synthetic vaccine, a messenger RNA (mRNA) vaccine, or a conventional vaccine.
[0244] 10. A non-transitory computer-readable medium storing instructions executable by a processor to cause the processor to: at least one computing device, the at least one computing device including at least one processor, the at least one processor being configured to: feed data from a subject into a prediction model; score neoantigens appearing in a subset of data from the subject against one or more parameters; determine one or more neoantigens appearing in the subset of the subject that satisfy an immune stimulation threshold based on the scores against the one or more parameters; determine somatic mutations for the one or more neoantigens appearing in the subset of the subject that satisfy the immune stimulation threshold; and output a prediction catalog of one or more of the somatic mutations appearing in the subset of the subject that satisfy the immune stimulation threshold, the prediction catalog predicting the somatic mutations.
[0245] 11. The non-transitory computer-readable medium of claim 10, wherein the data originates from a subject who has previously been treated for a certain type of cancer.
[0246] 12. The non-transitory computer-readable medium of claim 10, wherein the prediction model is more predictive of immune stimulation for a subset of the subject than tumor mutational burden (TMB) status.
[0247] 13. The non-transitory computer-readable medium of claim 11, wherein the prediction model is a machine learning model.
[0248] 14. The non-transitory computer-readable medium of claim 13, wherein the parameters include one or more of the following: peptide processing and presentation, RNA expression, MHC binding fold change, T cell activation, and dissimilarity to a reference human proteome.
[0249] 15. The non-transitory computer-readable medium of claim 14, wherein the one or more parameters further include one or more of the following: immune checkpoint inhibitor (ICI) response, ctDNA result, age, sex, and ECOG score.
[0250] 16. A vaccine composition prepared by a method comprising the steps of: feeding data of a subject suffering from a certain type of cancer into a predictive model; scoring a neoantigen appearing in the data of the subject against one or more parameters; determining one or more neoantigens appearing in the subject that satisfy an immune stimulation threshold based on the scores against the one or more parameters; determining somatic mutations of the one or more neoantigens appearing in the subject that satisfy the immune stimulation threshold; and preparing one or more vaccine compositions to be administered to the subject against one or more somatic mutations of the one or more neoantigens that satisfy the immune stimulation threshold, wherein the one or more vaccines to be administered to the subject are selected from a predictive catalog of somatic mutations appearing in a subset of subjects previously treated for that type of cancer.
[0251] 17. The vaccine composition of claim 16, wherein the one or more parameters include one or more of the following: peptide processing and presentation, RNA expression, MHC binding fold change, T cell activation, and dissimilarity to a reference human proteome.
[0252] 18. The vaccine composition of claim 16, wherein the one or more vaccine compositions are selected from a library of pre-prepared vaccine compositions to be administered.
[0253] 19. The vaccine composition of claim 16, wherein the one or more of the vaccine compositions are one or more of the following: a peptide-based synthetic vaccine, a messenger RNA (mRNA) vaccine, or a conventional vaccine.
[0254] 20. The vaccine composition of claim 19, wherein the method further comprises: combining the one or more neoantigens with liquid nanoparticles (LNPs) to prepare the one or more vaccine compositions.
[0255] The methods, systems, apparatus, and equipment described herein may be implemented, comprise, or performed by one or more computer systems. The methods described above may also be stored on a non-transitory computer-readable medium. Many elements may be, include, or comprise computer systems.
[0256] It should be understood that both the general description and the detailed description provide examples that are essentially illustrative and are intended to provide an understanding of this disclosure without limiting its scope. Various mechanical, compositional, structural, electronic, and operational changes may be made without departing from the scope of this specification and the claims. In some cases, well-known circuits, structures, and techniques have not been shown or described in detail so as not to obscure the examples. Similar numbers in two or more figures represent the same or similar elements.
[0257] Additionally, unless the context otherwise indicates, the singular forms “a / an” and “the” are intended to include the plural forms as well. Furthermore, the terms “comprises,” “comprising,” “includes,” etc., specify the presence of stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups. Components described as coupled may be directly coupled electronically or mechanically, or they may be indirectly coupled via one or more intermediate components, unless otherwise specifically indicated. Mathematical and geometric terms are not necessarily intended to be used according to their strict definitions unless the context of the specification otherwise indicates, as those skilled in the art will understand that, for example, substantially similar elements acting in substantially similar ways may readily fall within the scope of descriptive terms, even if those terms also have strict definitions.
[0258] Where feasible, an element and its associated aspects described in detail with reference to one instance may be included in other instances where they are not specifically shown or described. For example, if an element is described in detail with reference to one instance but not with reference to a second instance, the element may still be claimed as included in the second instance.
[0259] Given the disclosure herein, further modifications and alternative examples will be apparent to those skilled in the art. For example, apparatus and methods may include additional components or steps omitted from the figures and description for clarity of operation. Therefore, this description is to be construed as merely illustrative and intended to teach those skilled in the art the general manner of implementing the invention. It should be understood that the various examples shown and described herein should be considered exemplary. Elements and materials, and the arrangement of such elements and materials, may be substituted for those shown and described herein, components and processes may be reversed, and certain features of this teaching may be utilized independently, all of which will be apparent to those skilled in the art upon benefiting from the description herein. Changes may be made to the elements described herein without departing from the scope of this teaching and the following claims.
[0260] It should be understood that the specific examples described herein are non-limiting, and modifications can be made to the structure, dimensions, materials, and methods without departing from the scope of this teaching.
[0261] Other instances of this disclosure will be apparent to those skilled in the art upon consideration of the specification and practice of its contents. The specification and examples are intended to be illustrative only, and the appended claims shall enjoy their fullest breadth, including equivalents, in accordance with applicable law.
Claims
1. A system comprising: At least one computing device, the at least one computing device including at least one processor, the at least one processor being configured to: Data from the subjects is fed into the predictive model; A score is given for a neoantigen appearing in a subset of data from the subjects, based on one or more parameters. One or more neoantigens that meet the immune stimulation threshold and appear in a subset of the subjects are determined based on scores for one or more of the parameters. Identify somatic mutations of one or more neoantigens that satisfy the immune stimulation threshold and appear in the subset of the subjects; as well as Output a prediction catalog of one or more of the somatic mutations that satisfy the immune stimulation threshold and appear in the subset of the subjects, the prediction catalog predicting the somatic mutations.
2. The system of claim 1, wherein the data comes from a subject who has previously been treated for a certain type of cancer.
3. The system of claim 1, wherein the prediction model is more predictive of immune stimulation for a subset of the subject than tumor mutational burden (TMB) status.
4. The system according to claim 2, wherein the prediction model is a machine learning model.
5. The system of claim 4, wherein the one or more parameters include one or a combination of the following: peptide processing and presentation, RNA expression, MHC binding fold change, T cell activation, and dissimilarity to a reference human proteome.
6. The system of claim 5, wherein the parameters further include one or more of the following: immune checkpoint inhibitor (ICI) response, ctDNA result, age, sex, and ECOG score.
7. A system comprising: At least one computing device, the at least one computing device including at least one processor, the at least one processor being configured to: Data from newly diagnosed subjects with a certain type of cancer is fed into the predictive model; The new antigens appearing in the data for the subjects are scored based on one or more parameters; One or more neoantigens that meet the immune stimulation threshold for the subject are determined based on scores for one or more of the parameters. Identify somatic mutations of one or more neoantigens that satisfy the immune stimulation threshold in the subject; as well as For the somatic mutations of the one or more neoantigens that satisfy the immune stimulation threshold, output one or more vaccines to be administered to the subject, wherein the one or more vaccines to be administered to the subject are selected from a predicted catalog of somatic mutations appearing in a subset of subjects previously treated for the type of cancer.
8. The system of claim 7, wherein one or more vaccines are selected from a pre-prepared vaccine library to be administered.
9. The system of claim 8, wherein the vaccine is one or more of the following: a peptide-based synthetic vaccine, a messenger RNA (mRNA) vaccine, or a conventional vaccine.
10. A non-transitory computer-readable medium storing instructions executable by a processor to cause the processor to: At least one computing device, the at least one computing device including at least one processor, the at least one processor being configured to feed data from subjects into a prediction model; The new antigens appearing in a subset of data from the subjects are scored based on one or more parameters; One or more neoantigens that meet the immune stimulation threshold and appear in a subset of the subjects are determined based on scores for one or more of the parameters. Identify somatic mutations of one or more neoantigens that satisfy the immune stimulation threshold and appear in the subset of the subjects; as well as Output a prediction catalog of one or more of the somatic mutations that satisfy the immune stimulation threshold and appear in the subset of the subjects, the prediction catalog predicting the somatic mutations.
11. The non-transitory computer-readable medium of claim 10, wherein the data originates from a subject who has previously been treated for a certain type of cancer.
12. The non-transitory computer-readable medium of claim 10, wherein the prediction model is more predictive of immune stimulation for a subset of the subject than tumor mutational burden (TMB) status.
13. The non-transitory computer-readable medium of claim 11, wherein the prediction model is a machine learning model.
14. The non-transitory computer-readable medium of claim 13, wherein the parameters include one or more of the following: peptide processing and presentation, RNA expression, MHC binding fold change, T cell activation, and dissimilarity to a reference human proteome.
15. The non-transitory computer-readable medium of claim 14, wherein the one or more parameters further include one or more of the following: immune checkpoint inhibitor (ICI) response, ctDNA result, age, sex, and ECOG score.
16. A vaccine composition prepared by a method comprising the following steps: Data from subjects with a certain type of cancer is fed into the predictive model; The new antigens appearing in the data for the subjects are scored based on one or more parameters; One or more neoantigens that meet the immune stimulation threshold for the subject are determined based on scores for one or more of the parameters. Identify somatic mutations of one or more neoantigens that satisfy the immune stimulation threshold in the subject; as well as For one or more somatic mutations of one or more neoantigens that meet an immune stimulation threshold, one or more vaccine compositions are prepared to be administered to the subject, wherein the one or more vaccines to be administered to the subject are selected from a predicted catalog of somatic mutations occurring in a subset of subjects previously treated for the type of cancer.
17. The vaccine composition of claim 16, wherein the one or more parameters include one or a combination of the following: peptide processing and presentation, RNA expression, MHC binding fold change, T cell activation, and dissimilarity to a reference human proteome.
18. The vaccine composition of claim 16, wherein the one or more vaccine compositions are selected from a library of pre-prepared vaccine compositions to be administered.
19. The vaccine composition of claim 16, wherein the one or more of the vaccine compositions are one or more of the following: a peptide-based synthetic vaccine, a messenger RNA (mRNA) vaccine, or a conventional vaccine.
20. The vaccine composition according to claim 19, wherein the method further comprises: The one or more neoantigens are combined with liquid nanoparticles (LNPs) to prepare the one or more vaccine compositions.
Citation Information
Patent Citations
Detecting mutations and ploidy in chromosomal segments
WO2015164432A1
Methods for cancer detection and monitoring by means of personalized detection of circulating tumor DNA
WO2019200228A1