Neoantigen analysis

By combining multivariate analysis with genomic and protein information, neoantigens with high potential clinical efficacy were screened, solving the problem of inaccurate identification of candidate neoantigens in existing technologies and achieving efficient screening for cancer immunotherapy.

CN121963858APending Publication Date: 2026-05-01PERSONAL GENOME DIAGNOSTICS INC
View PDF 23 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PERSONAL GENOME DIAGNOSTICS INC
Filing Date
2016-07-14
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies lack sensitivity and specificity, failing to effectively identify candidate neoantigens, leading to costly validation procedures and unsuccessful cancer treatments.

Method used

By combining multivariate analysis with genomic and protein information, and utilizing sequencing and HLA typing, the priority of neoantigens can be predicted, and candidate neoantigens with high potential clinical efficacy can be screened out, reducing the need for additional experimental validation.

Benefits of technology

It improves the sensitivity and positive predictive value of neoantigen identification, saves time and money, and prioritizes the most promising potential antigens for cancer immunotherapy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

Cancer immunology provides a desirable new approach for cancer treatment, but verifies that the potential neoantigens to be directed against targets are cost-effective and expensive. Analysis of MHC binding affinity, antigen processing, similarity to known antigens, predicted expression levels (as mRNA or proteins), self-similarity, and mutant allele frequency provides screening methods for identifying and prioritizing candidate neoantigens using sequencing data. The methods of the invention save time and money by identifying preferential candidate neoantigens for further experimental verification.
Need to check novelty before this filing date? Find Prior Art

Description

Neoantigen analysis

[0001] This application is a divisional application of Chinese Invention Patent Application No. 201680050906.1, filed on July 14, 2016, entitled "Analysis of Neoantigens".

[0002] Related applications

[0003] This application claims the benefit and priority of U.S. Provisional Application No. 62 / 192,373, filed July 14, 2015, which is incorporated herein by reference in its entirety. Technical Field

[0004] This invention relates to the field of neoantigen analysis. Specifically, it relates to the identification and prioritization of mutation-derived neoantigens for the development of cancer vaccines and T-cell therapies. Background Technology

[0005] Cancer is characterized by the proliferation of abnormal cells. The success of conventional treatments depends on the type of cancer and the stage at which it is detected. Many treatments involve expensive and painful surgery and chemotherapy, and are often unsuccessful or only moderately prolong the patient's life. Promising treatments under development include tumor vaccines or T-cell therapies that target tumor antigens, enabling the patient's immune system to distinguish between tumor and healthy cells and triggering an immune response. See Chen et al., Oncology Meets Immunology: The Cancer-Immunity Cycle, Immunity 39, July 25, 2013, the contents of which are incorporated herein by reference in their entirety for all purposes.

[0006] Neoantigens are a class of immunogens associated with tumor-specific mutations specific to a patient's cancer. Neoantigens have shown promise as targets for antitumor immunotherapies, including adaptive T-cell transfer with tumor-infiltrating lymphocytes (TILs), cancer vaccines, and checkpoint inhibitors. See Hacohen et al., Getting Personal with Neoantigen-Based Therapeutic Cancer Vaccines, Cancer Immunol Res, July 1, 11, 2013; Robbins et al., Mining exomic sequencing data to identify mutated antigens recognized by adoptively transferred tumor-reactive T cells, Nature Medicine 19, 747-752 (2013); each of these contents is incorporated herein by reference in its full text for all purposes.

[0007] While strategies exist for using sequenced tumor DNA and HLA typing to identify and prioritize candidate neoantigens, conventional techniques lack the sensitivity and specificity to identify some candidate neoantigens and provide non-priority results that still require expensive validation procedures. Snyder et al., Genetic Basis for Clinical Response to CTLA-4 Blockade in Melanoma, N Engl J Med, 2014; 371:2189-2199; Segal et al., Epitope landscape in breast and colorectal cancer, Cancer Res., 2008 Feb 1; 68(3):889-92; Fritsch et al., 2014, HLA-Binding Properties of Tumor Neoepitopes in Humans, Cancer Immunol Res., 2(6); 1-8; Each of these contents is incorporated herein by reference in its full text for all purposes. Summary of the Invention

[0008] This invention relates to a screening method for identifying and prioritizing candidate neoantigens. The invention recognizes key factors in co-operating to prioritize neoantigens for effective treatment. As a result of this recognition, the invention provides multivariate operations using genomic and protein-based information to prioritize neoantigens for highly personalized efficacy in cancer immunotherapy. Based on the application of the claimed method and the potential for clinical efficacy in patients from whom samples are obtained, peptide sequences are prioritized as candidate neoantigens.

[0009] In some embodiments, the method of the present invention utilizes sequencing and matched normal controls to achieve high levels of sensitivity and positive predictive value in identifying mutations or variants, even below low mutant allele frequencies in tumors. Once the mutated sequence and the corresponding candidate neoantigen peptide sequence are identified in tumor tissue, a neoantigen priority score is generated for each candidate neoantigen peptide sequence using the individual's HLA type and two or more of the following: similarity of the peptide sequence to known antigens; self-similarity of the peptide sequence; mutant allele frequency of the peptide sequence; predicted binding affinity of the peptide sequence to one or more HLA alleles of the individual; predicted antigen processing of the peptide sequence; and mRNA or protein expression analysis of the peptide sequence. Predicted antigen processing may include peptide cleavage prediction or antigen processing-associated transporter (TAP) affinity prediction. Various inputs used to calculate neoantigen priority may be weighted in some embodiments. Based on sequencing data, priority scores are used to identify and prioritize candidate neoantigens with high potential clinical utility, thereby focusing further investigation only on the most promising potential antigens. Therefore, the method of the present invention helps researchers increase their chances of successfully identifying neoantigens by providing priority reports, while reducing additional experiments, providing a screening method that saves time and money on costly experimental validation.

[0010] In some aspects, the present invention provides methods for predicting and prioritizing potential neoantigens. An exemplary method includes obtaining tumor nucleic acid sequences and normal nucleic acid sequences of an individual. The tumor nucleic acid sequences are compared with the normal nucleic acid sequences to identify a plurality of translational peptide sequences that may have tumor-specific mutations. The individual's HLA type is then determined, wherein the HLA type comprises one or more HLA alleles. The method further includes predicting the major histocompatibility complex (MHC) binding affinity between each of the plurality of peptide sequences and the HLA alleles, and predicting the antigenic peptide processing fraction for each of the plurality of peptide sequences. Mutant allele frequencies are determined for each of the plurality of peptide sequences, and each of the plurality of peptide sequences is compared with a known antigen to determine a similarity score with a known antigen. The method of the present invention further includes determining the self-similarity score of each of the plurality of peptide sequences derived from the normal nucleic acid sequence, and determining the mRNA expression level or protein expression level of each of the plurality of peptide sequences. For each of the plurality of peptide sequences, a multivariate operation is performed using terms including MHC binding affinity, antigen peptide processing fraction, similarity to known antigen fraction, self-similarity fraction, and mRNA expression level or protein expression level to generate a neoantigen priority score for each of the plurality of peptide sequences for each of these terms. A report containing the neoantigen priority score for each of the plurality of peptide sequences is then prepared.

[0011] In some embodiments, the method of the present invention may include determining the tumor nucleic acid sequence by whole-exome sequencing of tumor nucleic acids extracted from an individual's tumor tissue. Whole-exome sequencing may include next-generation sequencing or Sanger sequencing, or both. In some embodiments, a normal reference nucleic acid sequence is obtained from a database of shared sequences. In alternative embodiments, the normal nucleic acid sequence may be derived from non-tumor tissue of the same individual from which the sample was collected. The method of the present invention may include determining the normal nucleic acid sequence by whole-exome sequencing of normal nucleic acids obtained from an individual's non-tumor tissue.

[0012] In various embodiments, the antigen peptide processing fraction may include peptide cleavage prediction or antigen processing-associated transporter (TAP) affinity prediction. HLA type can be determined from tumor nucleic acid sequences or normal nucleic acid sequences, or by serotyping or cytological assays. In some embodiments, one or more steps of the method may be performed using a computer including a processor coupled to tangible nontransitory memory and input / output devices. The method of the invention may further include sending a report to an output device. In various methods of the invention, each of the plurality of peptide sequences may have a predicted MHC binding affinity of less than 500 nM, expressed as IC50. In some embodiments, the known antigen sequence may be obtained from a database of known antigen sequences. Attached Figure Description

[0013] Figure 1 illustrates a method for identifying and prioritizing candidate neoantigens.

[0014] Figure 2 illustrates another method for identifying and then prioritizing candidate neoantigens.

[0015] Figure 3 illustrates the germline and somatic changes detected in a series of cases and the importance of using matched normal controls to identify tumor-specific mutations.

[0016] Figure 4 shows a sample report of the present invention. Detailed Implementation

[0017] This invention provides a method for identifying and prioritizing candidate neoantigens for cancer immunotherapy. The method utilizes multivariate analysis to provide a priority score for determining which candidate neoantigens are most likely to be successfully used in the development of cancer immunotherapies. This method is particularly useful for determining individualized neoantigen prioritization to maximize therapeutic efficacy for specific tumors in specific patients. As a result of this method, clinicians may have a better understanding of which neoantigen therapies to bring into or advance in clinical trials to produce effective immunomodulatory therapeutics.

[0018] Data obtained from tumor nucleic acids, along with HLA typing, peptide similarity analysis, and other markers as described herein, generate scores reflecting the potential therapeutic efficacy of candidate neoantigens. Key inputs in the claimed multivariate analysis are provided herein. These inputs are pooled to prioritize candidate neoantigens for further development. In some embodiments, the method of the present invention relies on whole-exome sequencing and matched normal control sequencing to achieve high levels of sensitivity and positive predictive value in identifying mutations or variants, even at low mutant allele frequencies. Once a mutated sequence and a corresponding candidate neoantigen peptide sequence are identified for an individual, a weighted multivariate operation uses the individual's HLA type and two or more of the following to generate a neoantigen priority score for each candidate neoantigen peptide sequence: similarity of the peptide sequence to known antigens; self-similarity of the peptide sequence; mutant allele frequency of the peptide sequence; predicted major histocompatibility complex (MHC) binding affinity between the peptide sequence and one or more HLA alleles of the individual; predicted antigen processing of the peptide sequence; and mRNA or protein expression analysis of the peptide sequence. Predicted antigen processing may include peptide cleavage prediction or antigen processing-associated transporter (TAP) affinity prediction. Therefore, the method of this invention helps researchers increase their chances of successfully identifying neoantigens by providing priority reports, while reducing additional experiments, providing a time- and cost-effective initial screening for expensive experimental validation. Candidate neoantigens used herein can be provided as peptide sequences.

[0019] Figures 1 and 2 illustrate exemplary methods of the present invention, which include obtaining tumor nucleic acid sequencing data and normal nucleic acid sequencing data of an individual tumor. Mutations are identified together with wild-type and somatic peptide pairs in FASTA format, typically between 8 and 11 amino acids in length. In different embodiments, the length of the wild-type and somatic peptide pairs can be between 11 and 20 amino acids. FASTQ format nucleic acid sequencing data is used for HLA-genotyping via computer analysis, or alternatively, conventional, experimentally validated HLA-genotyping information (e.g., HLA-A01:01 HLA-A26:01) can be manually provided to the methods.

[0020] HLA alleles of identified individuals were used in conjunction with FASTA format data of wild-type and somatic peptide pairs to predict MHC binding affinity of the peptides and each HLA allele, for example, using NetMHCpan. Antigen processing, such as peptide cleavage and TAP transporter affinity, was predicted using tumor nucleic acid sequencing data. Candidate neoantigens or peptide sequences were then selected using antigen processing and MHC binding affinity predictions, where, for example, the predicted MHC binding affinity (in IC50) was less than 500 nM. Self-similarity was then assessed using tumor and normal nucleic acid sequences. Candidate neoantigens were also compared to known antigens to determine similarity and to obtain mRNA or protein expression of the gene carrying the peptide. A neoantigen priority score was then generated for each candidate neoantigen using a weighted multivariate operation, incorporating mutant allele frequency, expression, similarity to known antigens, self-similarity, and predicted MHC binding affinity. This score can be used to prioritize neoantigens for subsequent experiments.

[0021] Sample preparation, sequencing and mutation identification

[0022] The method of the present invention may include identifying and prioritizing candidate neoantigens or peptide sequences from provided nucleic acid sequences, or in some embodiments, may include sample preparation and sequencing techniques for generating nucleic acid sequences. In some embodiments, samples from an individual or patient may be obtained in the form of frozen tissue, FFPE blocks or slides, pleural effusion, cells, DNA, cell lines, blood, saliva, or xenografts. Samples may be obtained from tumor tissue, and in some embodiments, may also be obtained from normal tissue to provide a source of normal or matched normal nucleic acids. Normal nucleic acids may be obtained from any non-tumor tissue or from sources such as saliva or whole blood. Tumor nucleic acids and normal nucleic acids may be extracted from the sample using known methods. In a preferred embodiment, at least 50 ng of DNA should be obtained for sequencing.

[0023] Nucleic acids can comprise deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). Nucleic acid samples can be sequenced using any known method. Nucleic acid samples can be sequenced using a classic dideoxy sequencing reaction (Sanger sequencing) with labeled terminators or primers and gel separation in a plate or capillary. Other techniques that can be used with the methods of this invention include sequencing by synthesis using reversibly terminated labeled nucleotides, pyrosequencing, 454 sequencing, Illumina / Solexa sequencing, allele-specific hybridization with a labeled oligonucleotide probe library, synthetic sequencing by ligation following allele-specific hybridization with a labeled clonal library, real-time monitoring of labeled nucleotide incorporation during the polymerization step, Polony sequencing, translocation via nanopores or nanochannels, digestion or polymerization of DNA with nucleotide binding detected in nanopores or nanochannels, optical detection of nucleotides in the strand localized by nanopores or nanochannels, and SOLiD sequencing. Separated molecules can be sequenced by sequential or single growth reactions using polymerases or ligases and by single or sequential differential hybridization with probe libraries.

[0024] In some embodiments, sequencing technologies (e.g., next-generation sequencing technologies) are used to sequence a portion of one or more captured targets (e.g., or their amplicon), and these sequences are used to count the number of distinct barcodes present. Therefore, in some embodiments, aspects of the invention relate to highly multiplexed qPCR reactions.

[0025] Sequencing technologies that can be used include, for example, Illumina sequencing. Illumina sequencing is based on DNA amplification on a solid surface using foldback PCR and anchored primers. The DNA is fragmented, and adaptors are added to the 5' and 3' ends of the fragments. The DNA fragments attached to the surface of the flow cell channels are extended and bridged amplified. The fragments become double-stranded, and the double-stranded molecules are denatured. Multiple cycles of solid-phase amplification followed by denaturation can generate millions of clusters of approximately 1,000 copies of single-stranded DNA molecules with the same template in each channel of the flow cell. Sequencing is performed using primers, DNA polymerase, and reversible termination nucleotides labeled with four fluorophores. After nucleotide incorporation, the fluorophores are excited using a laser, and an image is captured and the identity of the first base is recorded. The 3' terminator and fluorophore from each incorporated base are removed, and the incorporation, detection, and identification steps are repeated. Sequencing based on this technology is described in U.S. Patent Nos. 7,960,120; 7,835,871; 7,232,656; 7,598,035; 6,911,345; 6,833,246; 6,828,100; 6,306,597; 6,210,891; U.S. Publication No. 2011 / 0009278; U.S. Publication No. 2007 / 0114362; U.S. Publication No. 2006 / 0292611; and U.S. Publication No. 2006 / 0024681, each of which is incorporated herein by reference in its entirety.

[0026] Sequencing produces multiple reads. Reads typically include data on nucleotide sequences less than about 150 bases or less than about 90 bases in length. In some embodiments, reads are between about 80 and 90 bases in length, for example, about 85 bases. In some embodiments, these are very short reads, i.e., less than about 50 or about 30 bases in length.

[0027] Sequencing techniques that can be used in the methods of the provided invention include, for example, 454 sequencing (454 LifeSciences, Roche, Branford, Connecticut) (Margulies, M et al., Nature, 437:376-380 (2005); U.S. Patent No. 5,583,024; U.S. Patent No. 5,674,713; and U.S. Patent No. 5,700,673). 454 sequencing involves two steps. In the first step, DNA is cut into fragments of approximately 300-800 base pairs and the fragments are blunted. Oligonucleotide adaptors are then attached to the ends of the fragments. The adaptors act as primers for amplification and sequencing of the fragments. The fragments can be attached to DNA-capturing beads, such as streptoacidin-coated beads, using, for example, adaptor B containing a 5'-biotin tag. The fragments attached to the beads are then subjected to PCR amplification in an oil-water emulsion droplet. The result is multiple copies of the amplified DNA fragment cloned on each bead. In the second step, the beads are captured in wells (microliter size). Pyrosequencing is performed in parallel on each DNA fragment. The addition of one or more nucleotides generates a light signal, which is recorded by a CCD camera in the sequencing instrument. The signal intensity is proportional to the number of nucleotides incorporated. Pyrosequencing utilizes the pyrophosphate (PPi) released upon nucleotide addition. PPi is converted to ATP by ATP sulfatase in the presence of adenosine 5'-phosphorylsulfate. Luteinase uses ATP to convert luciferin to oxidized luciferin, and this reaction produces the light that is detected and analyzed.

[0028] Another example of a DNA sequencing technology that can be used in the methods of the provided invention is the SOLiD technology from AppliedBiosystems, Life Technologies Corporation (Carlsbad, CA). In SOLiD sequencing, DNA is cut into fragments, and adaptors are attached to the 5' and 3' ends of the fragments to generate a fragment library. Alternatively, internal adaptors can be introduced to generate a paired library by ligating adaptors to the 5' and 3' ends of the fragments, circularizing the fragments, digesting the circularized fragments to generate internal adaptors, and attaching the adaptors to the 5' and 3' ends of the resulting fragments. Next, a population of clonal beads is prepared in a microreactor containing beads, primers, template, and PCR components. After PCR, the template is denatured and the beads are enriched to isolate beads with extended template. The template on the selected beads is subjected to 3' modification, which allows it to adhere to a glass slide. The sequence can be determined by sequentially hybridizing and ligating partially random oligonucleotides with a centrally defined base (or base pair) identified by a specific fluorophore. After the color is recorded, the linked oligonucleotides are cleaved and removed, and then the process is repeated.

[0029] Another example of a DNA sequencing technology that can be used in the methods of the provided invention is Ion Torrent sequencing, such as that described in U.S. Publications 2009 / 0026082, 2009 / 0127589, 2010 / 0035252, 2010 / 0137143, 2010 / 0188073, 2010 / 0197507, 2010 / 0282617, 2010 / 0300559, 2010 / 0300895, 2010 / 0301398, and 2010 / 0304982, the contents of which are incorporated herein by reference in their entirety. In Ion Torrent sequencing, DNA is cut into fragments of approximately 300-800 base pairs and the fragments are blunted. Oligonucleotide adaptors are then ligated to the ends of the fragments. The adaptors act as primers for the amplification and sequencing of the fragments. These fragments can be attached to a surface at a single resolution, making them individually resolvable. The addition of one or more nucleotides releases protons (H). + The signal is detected and recorded in the sequencing instrument. The signal intensity is proportional to the number of incorporated nucleotides.

[0030] Another example of a sequencing technology that can be used in the methods of the provided invention is Illumina sequencing. Illumina sequencing is based on DNA amplification on a solid surface using foldback PCR and anchored primers. The DNA is fragmented, and adaptors are added to the 5' and 3' ends of the fragments. The DNA fragments attached to the surface of the flow cell channels are extended and bridged amplified. The fragments become double-stranded, and the double-stranded molecules are denatured. Multiple cycles of solid-phase amplification followed by denaturation can generate millions of clusters of approximately 1,000 copies of single-stranded DNA molecules with the same template in each channel of the flow cell. Sequence sequencing is performed using primers, DNA polymerase, and reversible termination nucleotides labeled with four fluorophores. After nucleotide incorporation, the fluorophores are excited using a laser, and an image is captured and the identity of the first base is recorded. The 3' terminator and fluorophore from each incorporated base are removed, and the incorporation, detection, and identification steps are repeated. Sequencing based on this technology is described in the following documents: U.S. Publication No. 2011 / 0009278, U.S. Publication No. 2007 / 0114362, U.S. Publication No. 2006 / 0024681, U.S. Publication No. 2006 / 0292611, U.S. Patent No. 7,960,120, U.S. Patent No. 7,835,871, U.S. Patent No. 7,232,656, U.S. Patent No. 7,598,035, U.S. Patent No. 6,306,597, U.S. Patent No. 6,210,891, U.S. Patent No. 6,828,100, U.S. Patent No. 6,833,246, and U.S. Patent No. 6,911,345, each of which is incorporated herein by reference in its entirety.

[0031] Another example of sequencing technology that can be used in the methods of the provided invention includes the Single Molecule Real-Time (SMRT) technology of Pacific Biosciences (Menlo Park, California). In SMRT, each of the four DNA bases is attached to one of four different fluorescent dyes. These dyes are phosphate-linked. A single DNA polymerase is immobilized at the bottom of a zero-mode waveguide (ZMW) with a single-molecule template of single-stranded DNA. The ZMW is a confinement structure that allows observation of single nucleotides incorporated by the DNA polymerase against a background of rapidly diffused fluorescent nucleotides in and out of the ZMW (in microseconds). The incorporation of nucleotides into the growing strand takes several milliseconds. During this time, the fluorescent tag is excited and produces a fluorescent signal, and the fluorescent tag is cleaved. Detection of the corresponding fluorescence of the dye indicates which base has been incorporated. The process is repeated.

[0032] Another example of a sequencing technology that can be used in the methods of the provided invention is nanopore sequencing (Soni, GV and Meller, A., Clin Chem [Clinical Chemistry] 53: 1996-2001 (2007)). A nanopore is a tiny pore on the order of 1 nanometer in diameter. Immersing a nanopore in a conductive fluid and applying a potential to it results in a small current flowing through it due to the conduction of ions through the nanopore. The amount of current flowing is sensitive to the size of the nanopore. As a DNA molecule passes through the nanopore, each nucleotide on the DNA molecule impedes the nanopore to varying degrees. Therefore, the change in the current flowing through the nanopore as a DNA molecule passes through it represents a reading of the DNA sequence.

[0033] Another example of sequencing technology that can be used in the methods of the provided invention involves sequencing DNA using a chemically sensitive field-effect transistor (chemFET) array (e.g., as described in U.S. Publication No. 2009 / 0026082). In one example of this technology, a DNA molecule can be placed in a reaction chamber, and a template molecule can be hybridized with sequencing primers that bind a polymerase. The incorporation of one or more triphosphates at the 3' end of the sequencing primers into a new nucleic acid strand can be detected by a change in current at the chemFET. An array can have multiple chemFET sensors. In another example, a single nucleic acid can be attached to a bead, and the nucleic acid can be amplified on the bead. Individual beads can be transferred to individual reaction chambers on a chemFET array, each chamber having a chemFET sensor, and the nucleic acid can be sequenced.

[0034] Another example of sequencing technology that can be used in the methods of the provided invention involves the use of an electron microscope (Moudrianakis EN and Beer M., PNAS [Proceedings of the National Academy of Sciences], 53:564-71 (1965)). In one example of the technique, individual DNA molecules are labeled with metal tags that are distinguishable by an electron microscope. These molecules are then stretched on a flat surface and imaged using an electron microscope to measure the sequence.

[0035] Another example of sequencing technology that can be used in the methods of the provided invention involves the Rapid Aneuploidy Screening Test-Sequencing System (FAST-SeqS) as described in PCT application PCT / US2013 / 033451, which is linked by reference. See also Kinde et al., “FAST-SeqS: A Simple and Efficient Method for the Detection of Aneuploidy by Massively Parallel Sequencing” DOI: 10.1371 / journal.pone.0041162, which is linked by reference. FAST-SeqS uses specific primers, specifically a pair of primers, which anneal to a subgroup of sequences scattered throughout the genome. These regions are selected due to similarity, allowing them to be amplified with a single primer pair, but unique enough to allow differentiation of most amplified loci. Unlike traditional whole-genome amplification libraries (where each tag must be aligned independently), FAST-SeqS generates sequence alignments with a smaller number of locations.

[0036] Sequence assembly can be performed using methods known in the art, including reference-based assembly, de novo assembly, alignment-by-alignment assembly, or a combination thereof. In some embodiments, sequence assembly uses the Low Coverage Sequence Assembly Software (LOCAS) tool, described by Klein et al., LOCAS-A low coverage sequence assembly tool for re-sequencing projects, PLoS One 6(8) article 23455 (2011), the contents of which are hereby incorporated by reference in their entirety. Sequence assembly is described in U.S. Patent Nos. 8,165,821; 7,809,509; 6,223,128; 2011 / 0257889; and U.S. Publication No. 2009 / 0318310, each of which is hereby incorporated by reference in its entirety.

[0037] Once the tumor nucleic acid sequence is obtained, it can be compared with a normal nucleic acid sequence to identify mutations in the tumor nucleic acid sequence. In some embodiments, the normal nucleic acid can be a reference genome, such as HG18 or HG19, or any human reference sequence compiled by the International Human Genome Sequencing Consortium or the 1000 Genomes Project. In a preferred embodiment, the normal nucleic acid sequence is a matching normal nucleic acid that can be obtained from an individual's non-tumor tissue or from an associated individual. Using a matching normal tissue as a reference sequence for determining variants or mutations can help identify germline mutations present in an individual's tumor and non-tumor cells and can allow for the elimination of false positives and more accurate identification of tumor-specific variants or mutations. Figure 2 shows a bar graph illustrating the germline and somatic changes detected in a series of cases, demonstrating the importance of using a matching normal control to identify tumor-specific mutations. Mutations, as used herein, can include, for example, modifications, chromosomal alterations, substitutions, insertions or deletions, single nucleotide polymorphisms, translocations, inversions, duplications, and copy number variations.

[0038] In one exemplary embodiment, commercially available technologies such as CANCERXOME, available from Personal Genome Diagnostics, Inc. (Baltimore, Maryland), are used to identify tumor-specific mutations.

[0039] HLA typing

[0040] HLA typing of an individual or patient can be performed using a variety of known methods, including cytology, serological typing, genotyping, or analysis of sequence data via computer.

[0041] In a preferred embodiment, HLA typing is performed via computer analysis using one or more techniques such as OptiType, which run on a computing device. See Szolek et al., OptiType: precision HLA typing from next-generation sequencing data, Bioinformatics. Dec 1, 2014; 30(23), incorporated herein in its entirety for all purposes. Various other computer analysis techniques may also be used. See Major et al., HLA typing from 1000 genomes whole genome and whole exome Illumina data, PLoS One. 6 November 2013; 8(11):e78410; Wittig et al., Development of a high-resolution NGS-based HLA-typing and analysis pipeline, Nucl. Acids Res. (2015), first published online on 9 March 2015, doi:10.1093 / nar / gkv184.

[0042] In some embodiments, HLA alleles can be determined by other means (e.g., HLA-A01:01, HLA-A26:01), and the results can also be used by the method, thereby avoiding the need for prediction via computer analysis.

[0043] Identification of candidate neoantigens and prioritization of candidate neoantigens

[0044] Using an individual's HLA typing information and peptide sequences with identified mutations, various computer analysis techniques and programs can be used to predict MHC binding affinity for each peptide sequence (e.g., NetMHCpan version 2.8 available at http: / / www.cbs.dtu.dk / services / NetMHCpan / , MHC-I antigen peptide processing prediction (MAPPP) available at http: / / www.mpiib-berlin.mpg.de / MAPPP / , Bioinformatics and Molecular Analysis (BIMAS) HLA peptide binding prediction available at http: / / www.bimas.cit.nih.gov / molbio / hla_bind / , Rankpep MHC peptide binding prediction available at http: / / imed.med.ucm.es / Tools / rankpep.html, or the SYFPEITHI epitope predictor available at http: / / www.syfpeithi.de / bin / MHCServer.dll / EpitopePrediction.htm).

[0045] Proper peptide processing, including peptide cleavage and antigen processing-associated transporter (TAP), is necessary before MHC presentation and binding. According to the method of the invention, candidate neoantigens can be identified in part using antigen peptide processing predictions that may include an antigen peptide processing fraction. In some embodiments, the antigen peptide processing fraction may include peptide cleavage prediction and TAP binding affinity prediction. Peptide cleavage can be predicted from the peptide sequence using computer analysis techniques or computer programs (e.g., the MAPPP proteasome cleavage predictor available at http: / / www.mpiib-berlin.mpg.de / MAPPP / cleavage.html or the Rankpep cleavage predictor available at http: / / imed.med.ucm.es / Tools / rankpep.html). Similarly, known methods can be used to predict TAP binding affinity from peptide sequences, such as those described in Doytchinova et al., Transporter associated with antigen processing preselection of peptides binding to the MHC: a bioinformatic evaluation, J Immunol. Dec 1, 2004; 173(11); Tenzer et al., Modeling the MHC class I pathway by combining predictions of proteasomal cleavage, TAP transport and MHC class I binding, Cell Mol Life Sci. May 2005; 62(9): 1025-37; Zhang et al., PREDTAP: a system for prediction of peptide binding to the human transporter associated with antigen processing [PREDTAP: A system for predicting peptides that bind to human transporters associated with antigen processing], Immunome Research, May 2006, 2:3; its contents are incorporated herein by reference in their entirety and for all purposes.

[0046] Based on antigenic peptide processing prediction or scores, candidate neoantigens can be classified into epitope (E) or non-antigen (NA) antigenic peptide processing categories, with E classification taking precedence over NA classification.

[0047] Using MHC binding affinity cutoff values, such as IC50 values ​​less than 100 nM, 200 nM, 300 nM, 400 nM, 500 nM, 600 nM, 700 nM, 800 nM, 900 nM, 1000 nM, etc., candidate neoantigen or peptide sequences with predicted MHC binding affinity higher than the cutoff value can be excluded from further analysis or consideration.

[0048] Candidate neoantigens can be further characterized by analyzing their similarity to known antigens, predicted expression levels (as mRNA or protein), self-similarity measurements, and mutant allele frequencies. Mutant allele frequencies can be determined by analyzing tumor nucleic acid sequencing data to ascertain the frequency of a subject mutant allele in the sequenced nucleic acid compared to other alleles of the stated nucleic acid or gene. Mutant allele frequencies can be determined, for example, as average expression in tumor nucleic acids. Generally, an increased mutant allele frequency will indicate an increased likelihood of clinical utility for the peptide sequence or candidate neoantigen.

[0049] Self-similarity can be determined by comparing the mutant peptide sequence to an equivalent normal peptide sequence to establish a similarity score. In some embodiments, self-similarity can be determined amino acid-by-amino along the peptide sequence. Self-similarity can be determined as a percentage value. Generally, a lower level of self-similarity will indicate an increased likelihood of clinical utility for the peptide sequence or candidate neoantigen. In a preferred embodiment, an amino acid substitution PMBEC matrix is ​​used to calculate the similarity score, with a score less than 0.05 reflecting a loss of similarity in the mutant peptide to the parental wild-type peptide. (See http: / / www.biomedcentral.com / 1471-2105 / 10 / 394).

[0050] Known antigen similarity can be determined by comparing a peptide sequence to the peptide sequence of a known antigen. In some embodiments, similarity to a known antigen can be determined amino acid-by-amino along the peptide sequence. Known antigens can be obtained from databases, such as the IEDB (Immunotope Database and Analysis Resource), available at http: / / www.iedb.org / home_v3.php. A score can be assigned to a peptide sequence or candidate neoantigen, which may include, for example, a percentage similarity value, which may be the highest value determined from a series of comparisons with multiple known antigens. Generally, a higher level of similarity to a known antigen will indicate an increased likelihood of clinical utility for the peptide sequence or candidate neoantigen. Known antigen similarity can be determined, for example, by searching for sequence homology with known antigens through a sequence similarity search of neoantigen candidates against IEDB, or by searching databases of other bacterial proteins that may reflect neoantigens.

[0051] Protein or mRNA expression levels of peptide sequences can be predicted by measuring the expression of relevant genes in tumor samples, for example, using RNAseq analysis or microarrays, or by referencing databases of known expression data associated with specific tumor types (e.g., cancer genome maps).

[0052] In various embodiments, multivariate operations can be performed on two or more of the following terms to generate a neoantigen priority score for each of a plurality of peptide sequences: MHC binding affinity, antigenic peptide processing score, similarity to known antigen score, self-similarity score, and mRNA expression level or protein expression level. In various embodiments, weight values ​​can be used to weight one or more terms to increase or decrease their influence on the neoantigen priority score relative to other terms.

[0053] In one exemplary embodiment, neoantigen prioritization can be determined by applying rules or a set of rules to one or more features identified or characterized for each of a plurality of candidate neoantigens. Rules may include exclusion clauses and multifactor classification parameters that prioritize neoantigen candidate features such as MHC binding affinity, antigenic peptide processing fraction or classification, similarity to known antigens, self-similarity fraction, and mRNA or protein expression levels. Examples of such embodiments can be found below.

[0054] In some embodiments, neoantigen prioritization may be included in a report prepared according to the method of the invention. A sample report is shown in Figure 3. The report may consist of any combination of identified candidate neoantigens or peptide sequences, associated neoantigen priority scores, and values ​​of any combination of determinant terms. In some embodiments, candidate neoantigens may be prioritized, for example, from highest to lowest priority. The report may be physical, printed or written on paper using the output devices described below, or it may be electronic, prepared and stored on a computing device. The report may be sent to interested parties, such as the individual or patient being tested, the ordering physician or laboratory, or other physicians or laboratories, or other entities. The report may be delivered in physical form or sent electronically, for example, by email.

[0055] computing devices

[0056] As those skilled in the art will recognize, methods necessary or best suited to carry out the invention may include one or more computing devices, computing systems or computers, including one or more processors (e.g., central processing unit (CPU), graphics processing unit (GPU), etc.), computer-readable storage devices (e.g., main memory, static memory, etc.) or combinations thereof communicating with each other via a bus.

[0057] The processor may include any suitable processor known in the art, such as the processor sold by Intel under the trademark XEON E7 (Santa Clara, California) or the processor sold by AMD under the trademark OPTERON 6200 (Sunnyvale, California).

[0058] The memory preferably includes at least one tangible, non-transitory medium capable of storing: one or more sets of instructions executable to cause the system to perform the functions described herein (e.g., software embodied in any methods or functions found herein or the aforementioned computer programs); data (e.g., images of drug data sources, personal data, or drug databases); or both. While a computer-readable storage device may be a single medium in exemplary embodiments, the term "computer-readable storage device" should be understood to include a single or multiple media (e.g., centralized or distributed databases and / or associated caches and servers) that store instructions or data. Therefore, the term "computer-readable storage device" should be understood to include, but is not limited to: solid-state storage (e.g., a Subscriber Identity Module (SIM) card, a Secure Digital Card (SD card), a MicroSD card, or a Solid State Drive (SSD)), optical and magnetic media, and any other tangible storage media.

[0059] Any suitable service can be used for storage, such as Amazon Web Services, memory for computing systems, cloud storage, servers, or other computer-readable storage.

[0060] The input / output device according to the present invention may include one or more of the following: a display unit (e.g., a liquid crystal display (LCD) or a cathode ray tube (CRT) monitor), an alphanumeric input device (e.g., a keyboard), an indicator control device (e.g., a mouse or a touchpad), a disk drive unit, a printer, a signal generating device (e.g., a speaker), a touch screen, buttons, an accelerometer, a microphone, a cellular radio frequency antenna, and a network interface device (which may be, for example, a network interface card (NIC), a Wi-Fi card, or a cellular modem, or any combination thereof).

[0061] Those skilled in the art will recognize that the methods described herein can be implemented using any suitable development environment or programming language. For example, the methods herein can be implemented using Perl, Python, C++, C#, Java, JavaScript, Visual Basic, Ruby on Rails, Groovy, and Grails, or any other suitable tool. For mobile devices, native Xcode or Android Java is preferred.

[0062] Example

[0063] Example 1

[0064] In one exemplary embodiment, given a set of candidate neoantigen peptides, which can be determined by their association with somatic mutations identified by sequence analysis, the following set of rules can be applied. MHC binding affinity can be determined as described above to determine the predicted IC50 affinity. All candidate neoantigens with predicted IC50 affinity greater than, for example, 500 nM can then be excluded from further examination.

[0065] The remaining candidate neoantigens can then be classified according to a multifactorial classification starting with MHC binding affinity. As mentioned above, candidate neoantigens can be classified as SB or WB (strong and weak binders) and categorized such that SB peptides are given higher priority than WB peptides.

[0066] The antigenic peptide processing of candidate neoantigens can then be determined or predicted as described above. Candidate neoantigens can then be classified as E or NA, and further classified within their MHC binding affinity classification order, such that SB peptides classified as E are preferred over SB peptides classified as NA, which in turn are preferred over WB peptides classified as E.

[0067] The reference gene expression level of the candidate neoantigen can then be determined using the method described above. The candidate neoantigen can then be further classified based on its current MHC binding affinity and antigen processing priority using the reference gene expression level, with higher levels having higher priority than lower levels.

[0068] The resulting ranking list of candidate neoantigens may include priority groups of candidate neoantigens, with the highest priority candidate neoantigens at the top of the list. This list may be presented and delivered to the requesting individual or entity in the form of a report as described elsewhere.

[0069] In some embodiments, the method may include treating a patient with a vaccine or T-cell therapy targeting a neoantigen based on the ranking of candidate neoantigens after grading them. The method may include experimentally validating candidate neoantigens based on the ranking of neoantigens. The method may also include inducing treatment in a patient with a vaccine or T-cell therapy targeting a neoantigen based on the ranking of neoantigens.

[0070] Example 2

[0071] In a second exemplary embodiment, given a set of candidate neoantigen peptides, which can be determined by their association with somatic mutations identified by sequence analysis, the following set of rules can be applied. MHC binding affinity can be determined as described above to determine the predicted IC50 affinity. All candidate neoantigens having a predicted IC50 affinity greater than, for example, 1000 nM can then be excluded from further examination.

[0072] RNAseq expression values ​​can be determined for the remaining candidate neoantigens, and peptides with RNAseq expression values ​​of related genes below the threshold of approximately 10 reads / kilobases / megamapped reads (RPKM) can be excluded from further examination.

[0073] The remaining candidate neoantigens can then be classified according to a multifactorial classification starting with MHC binding affinity. As mentioned above, candidate neoantigens can be classified as SB or WB and further classified such that SB peptides are given higher priority than WB peptides.

[0074] The antigenic peptide processing of candidate neoantigens can then be determined or predicted as described above. Candidate neoantigens can then be classified as E or NA, and further classified within their MHC binding affinity classification order, such that SB peptides classified as E are preferred over SB peptides classified as NA, which in turn are preferred over WB peptides classified as E.

[0075] The self-similarity of candidate neoantigens can then be determined using the PMBEC comparison described above. Candidate neoantigens can then be secondary-classified by self-similarity scores within their current MHC binding affinity and antigen processing priority, with lower scores having higher priority than higher scores.

[0076] The resulting ranking list of candidate neoantigens may include priority groups of candidate neoantigens, with the highest priority candidate neoantigens at the top of the list. This list may be presented and delivered to the requesting individual or entity in the form of a report as described elsewhere.

[0077] Example 3

[0078] In a third exemplary embodiment, given a set of candidate neoantigen peptides, which can be determined by their association with somatic mutations identified by sequence analysis, the following set of rules can be applied. MHC binding affinity can be determined as described above to determine the predicted IC50 affinity. All candidate neoantigens having a predicted IC50 affinity greater than, for example, 750 nM can then be excluded from further examination.

[0079] RNAseq expression values ​​can be determined for the remaining candidate neoantigens, and peptides with RNAseq expression values ​​of related genes below the threshold of approximately 25 reads / kilobases / megamapped reads (RPKM) can be excluded from further examination.

[0080] The remaining candidate neoantigens can then be classified according to a multifactorial classification starting with MHC binding affinity. As mentioned above, candidate neoantigens can be classified as SB or WB and further classified such that SB peptides are given higher priority than WB peptides.

[0081] The antigenic peptide processing of candidate neoantigens can then be determined or predicted as described above. Candidate neoantigens can then be classified as E or NA, and further classified within their MHC binding affinity classification order, such that SB peptides classified as E are preferred over SB peptides classified as NA, which in turn are preferred over WB peptides classified as E.

[0082] The similarity of candidate neoantigens to known antigens can then be determined using methods such as those described above. Sub-classification of candidate neoantigens can then be based on amino acid matches with 100% identity to known antigens (longer perfect matches reflect higher priority).

[0083] The resulting list of candidate neoantigens may include a priority group of candidate neoantigens, with the highest priority candidate neoantigens at the top of the list. This list may be presented and delivered to requesting individuals or entities in the form of a report as described elsewhere. The list may be used to experimentally validate or select, administer, or result in the administration of a vaccine or T-cell therapy containing a target neoantigen from the priority candidate neoantigens in the list.

[0084] Example 4

[0085] Using sequencing data; known neoantigens from Fritsch et al., Cancer Immunol Res [Cancer Immunology Research] 2014; experimentally validated neoantigens from Robbins et al., Nat Med [Nature Medicine] 2013; and predictive biomarkers of checkpoint inhibitors identified using the technique of Snyder et al., NEJM [New England Journal of Medicine] 2014; the method of the present invention is applied to the sequencing data, and the priority peptide sequences or candidate neoantigens are compared with validated neoantigens using the application of the rule set described in Example 1.

[0086] The number of preferred candidate neoantigens generated by the operation and the ranking of experimentally validated neoantigens from Robbins et al., Nat Med [Nature Medicine] 2013 are shown in Table 1 below:

[0087] Table 1

[0088] The operation sorts experimentally validated neoantigens into the top 20% of all candidate neoantigens.

[0089] The comparison of the operation involved identifying candidate neoantigens with known neoantigens from Fritsch et al., Cancer Immunol Res [Cancer Immunology Research] 2014, which revealed that the operation identified 18 of the 19 known neoantigens, as shown in Table 2, with a sensitivity greater than 90%.

[0090] Table 2

[0091] Example 5

[0092] Using cancer genomics databases (e.g., TCGA) and cancer mutation databases (e.g., COSMIC), we identified over 1000 frequent mutations occurring at least 1% of any tumor type. We then predicted the protein regions flanked by the mutations and the neoORF due to frameshift mutations. Using dbMHC from NCBI, we compiled the most prevalent HLA class I alleles in the North American population, resulting in 90 unique 4-elements, each with a population frequency of ≥ 0.15%.

[0093] The method of this invention was applied to over 1000 somatic mutation-related peptides and the 90 HLA alleles to predict and prioritize candidate neoantigens. Table 3 provides a partial list of frequent somatic mutations in which mutation-related peptides were predicted as candidate neoantigens for at least one of the compiled HLA alleles. The reported frequencies of HLA alleles in the North American population allow for assessment of the likelihood that a patient with a specific neoantigen-associated somatic mutation will possess at least one HLA allele that recognizes that neoantigen.

[0094] The frequent somatic mutations identified here may generate neoantigens that could be promising targets for effective vaccines and T-cell therapies. Because these antigens are expressed only on tumor cells and not on other cells, vaccines and T-cell therapies targeting them would produce a concentrated immune response with reduced cytotoxicity. Furthermore, since these mutations occur in multiple patients, most of them may confer a growth advantage on the tumor, thus halting tumor growth upon eradication. Additionally, for many mutations, the resulting candidate neoantigens may bind to HLA alleles in a large number of subgroups of the North American population, suggesting that vaccines or T-cell therapies targeting these neoantigens could potentially benefit many patients. Unsurprisingly, many of the mutations identified in the analyses listed above have been shown to induce anti-tumor immunity, including IDH1-R132H (Schumacher et al., Nature 2014), KRAS-G12 mutation (Chaft et al., Clin Lung Cancer 2014) and EGFR-VIII deletion (Taylor et al., Curr Cancer Drug Targets 2012).

[0095] Table 3

[0096] This invention relates to the following embodiments:

[0097] 1. A method for prioritizing candidate neoantigens for a patient, the method comprising the steps of:

[0098] Multiple candidate neoantigens were obtained;

[0099] Determine the self-similarity of the members of the plurality of candidate neoantigens;

[0100] Determine the similarity between members of the plurality of candidate neoantigens and known antigens;

[0101] Determine the expression levels of members of the plurality of candidate neoantigens;

[0102] Identify the frequencies of mutant alleles in the exons encoding members of the plurality of candidate neoantigens; and

[0103] The rules are applied to the results of the determination and identification steps to rank the members of the plurality of candidate neoantigens according to the likelihood of clinical significance.

[0104] 2. The method of embodiment 1, the method further comprising preparing a report containing members of the sorted plurality of candidate neoantigens.

[0105] 3. The method as described in embodiment 1, wherein the plurality of candidate neoantigens are derived from patient tumor samples.

[0106] 4. The method of embodiment 1, wherein the plurality of candidate neoantigens are obtained by determining the HLA genotype and MHC binding affinity of candidate peptides obtained from tumor samples.

[0107] 5. The method as described in embodiment 4, wherein the HLA genotype and MHC binding affinity of the candidate peptide are determined from peptide sequence data via computer analysis.

[0108] 6. The method as described in embodiment 4, wherein the HLA genotype and MHC binding affinity of the candidate peptide are determined by assay.

[0109] 7. The method of embodiment 4, wherein the candidate peptide is obtained by comparing a peptide from the tumor sample with a corresponding peptide from a normal sample, wherein the candidate peptide contains a mutation relative to the corresponding peptide.

[0110] 8. The method of embodiment 4, wherein the application step includes excluding candidate neoantigens having an MHC binding affinity greater than 1000 nM from the plurality of candidate neoantigens.

[0111] 9. The method of embodiment 8, wherein the application step includes excluding candidate neoantigens having an MHC binding affinity greater than 750 nM from the plurality of candidate neoantigens.

[0112] 10. The method of embodiment 9, wherein the application step includes excluding candidate neoantigens having an MHC binding affinity greater than 500 nM from the plurality of candidate neoantigens.

[0113] 11. The method of embodiment 1, wherein each of the plurality of candidate neoantigens is assigned to a strongly binding (SB) or weakly binding (WB) MHC classifier, and the application step includes sorting the plurality of neoantigens such that SB candidate neoantigens are ranked higher than WB candidate neoantigens.

[0114] 12. The method as described in embodiment 1, wherein the method further comprises:

[0115] Determine the antigenic peptide processing classification of the members of the plurality of neoantigens.

[0116] 13. The method of embodiment 12, wherein the plurality of neoantigens are assigned to a classification of epitopes (E) or non-antigens (NA), and the application step includes sorting the plurality of neoantigens such that candidate E neoantigens are ranked higher than candidate NA neoantigens.

[0117] 14. The method of embodiment 1, wherein the application step includes sorting the plurality of neoantigens such that candidate neoantigens with lower self-similarity are ranked higher than neoantigens with higher self-similarity.

[0118] 15. The method of embodiment 1, wherein the expression level comprises RNAseq expression values, and the application step comprises excluding candidate neoantigens from the plurality of candidate neoantigens having expression values ​​of less than 10 reads / kilobase / megamapped reads (RPKM).

[0119] 16. The method of embodiment 15, wherein the application step includes excluding candidate neoantigens from the plurality of candidate neoantigens that have an expression value of less than 25 readings / kilobases / megamapped readings (RPKM).

[0120] 17. The method of embodiment 1, wherein the application step includes ranking the plurality of neoantigens based on 100% amino acid identity with a portion of a known antigen, such that candidate neoantigens having amino acid identity with a longer portion of a known antigen are ranked higher than candidate neoantigens having amino acid identity with a shorter portion of a known antigen.

[0121] 18. The method of embodiment 12, wherein the antigen peptide processing classification is determined using peptide cleavage prediction or antigen processing-associated transporter (TAP) affinity prediction.

[0122] 19. The method of embodiment 2, wherein one or more of the acquisition, determination, identification, or application steps are performed using a computer comprising a processor coupled to tangible nontransitory memory and input / output devices.

[0123] 20. The method of embodiment 10, the method further comprising sending the report to the output device.

[0124] 21. A method for identifying shared neoantigens, the method comprising the following steps:

[0125] Selected from multiple frequent mutations occurring in more than one tumor type;

[0126] Based on the predicted peptide sequence, determine which of the multiple frequent mutations is a potential neoantigen; and

[0127] Based on the prevalence of HLA class I alleles and the prevalence of frequent mutations across multiple tumor types, the potential neoantigens were identified as common neoantigens.

[0128] By incorporating via reference

[0129] In the entirety of this disclosure, references and citations have been made to other sources such as patents, patent applications, patent publications, journals, books, papers, and web content. All of these sources are incorporated herein by reference in their entirety for all purposes.

[0130] equivalent

[0131] In addition to those shown and described herein, various modifications of the invention and many other embodiments thereof will become apparent to those skilled in the art from the entire contents of this document (including references to the scientific and patent literature cited herein). The subject matter herein contains important information, examples, and guidance applicable to the practice of the invention in its various embodiments and equivalents.

Claims

1. A method for prioritizing candidate neoantigens for a patient, the method comprising the steps of: Obtain multiple candidate neoantigens; determine the self-similarity of the members of the multiple candidate neoantigens; Determine the similarity between members of the plurality of candidate neoantigens and known antigens; The expression levels of members of the plurality of candidate neoantigens are determined; the frequencies of mutant alleles in the exons encoding members of the plurality of candidate neoantigens are identified; and rules are applied to the results of the determination and identification steps to rank the members of the plurality of candidate neoantigens according to the likelihood of clinical significance.

2. The method of claim 1, the method further comprising preparing a report containing members of the sorted plurality of candidate neoantigens.

3. The method of claim 1, wherein the plurality of candidate neoantigens are derived from patient tumor samples.

4. The method of claim 1, wherein the plurality of candidate neoantigens are obtained by determining the HLA genotype and MHC binding affinity of candidate peptides obtained from tumor samples.

5. The method of claim 4, wherein the HLA genotype and MHC binding affinity of the candidate peptide are determined from peptide sequence data via computer analysis.

6. The method of claim 4, wherein the HLA genotype and MHC binding affinity of the candidate peptide are determined by assay.

7. The method of claim 4, wherein the candidate peptide is obtained by comparing a peptide from the tumor sample with a corresponding peptide from a normal sample, wherein the candidate peptide contains a mutation relative to the corresponding peptide.

8. The method of claim 4, wherein the application step comprises excluding candidate neoantigens having an MHC binding affinity greater than 1000 nM from the plurality of candidate neoantigens.

9. The method of claim 8, wherein the application step comprises excluding candidate neoantigens having an MHC binding affinity greater than 750 nM from the plurality of candidate neoantigens.

10. A method for identifying shared neoantigens, the method comprising the following steps: Select multiple frequent mutations occurring in more than one tumor type; determine which of the multiple frequent mutations is a potential neoantigen based on the predicted peptide sequence; and identify the potential neoantigen as a common neoantigen based on the prevalence of HLA class I alleles and the prevalence of frequent mutations across multiple tumor types.

Citation Information

Patent Citations

  • Methods for producing a paired tag from a nucleic acid sequence and methods of use thereof

    US20060024681A1

  • Paired end sequencing

    US20060292611A1

  • Confocal imaging methods and apparatus

    US20070114362A1

  • Methods and apparatus for measuring analytes using large scale FET arrays

    US20090026082A1

  • Methods and apparatus for measuring analytes using large scale FET arrays

    US20090127589A1