Allele-specific expression

The method uses RNA-seq data and a Bayesian framework to account for tumor purity and sequencing errors, enhancing the sensitivity and accuracy of tumor-specific mutation detection, particularly for identifying neoantigens.

JP2025534267APending Publication Date: 2025-10-15IOVANCE BIOTHERAPEUTICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025517289
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-23
Filing Date
2023-09-06
Publication Date
2025-10-15

AI Technical Summary

Technical Problem

Existing methods for determining tumor-specific mutant allele expression are insufficiently sensitive and inaccurate due to the mixing of tumor and normal cells in samples, leading to false negatives and exclusion of potentially expressed variants.

Method used

A method using RNA-seq data to calculate the probability of tumor-specific mutation expression by considering tumor purity, sequencing error rates, and gene expression levels, employing a Bayesian framework to determine the likelihood of mutation expression and power of detection.

Benefits of technology

Accurately identifies tumor-specific mutations likely to be expressed, even at low levels, reducing false negatives and providing a robust, reproducible method for identifying neoantigens for cancer treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025534267000001_ABST
    Figure 2025534267000001_ABST
Patent Text Reader

Abstract

A method for determining whether a tumor-specific mutation is likely to be expressed in a subject is described. The method includes obtaining RNA-seq data from one or more samples from the subject containing tumor genetic material, the RNA-seq data including, for each of the one or more samples, a number of RNA reads in the sample that exhibit the tumor-specific mutation (b) and a total number of RNA reads at the location of the tumor-specific mutation (d). The method further includes determining the likelihood of the sequence data if the tumor-specific mutation (i) is expressed and (ii) is not expressed. The method is used to detect tumor-specific mutations in the RNA-seq data, identify neoantigens, or provide therapies targeting such neoantigens. Related methods, systems, and products are also described.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to methods for determining whether tumor-specific mutant alleles are likely to be expressed and for identifying tumor-specific mutant alleles expressed in tumors. The present disclosure also relates to methods and compositions for the treatment of cancer using or targeting neoantigens. [Background technology]

[0002] It is generally accepted that cancer antigens arising from cancer-specific variants (also called "tumor-specific mutations") represent promising therapeutic targets for immunotherapy, provided they are expressed by cancer cells (see, e.g., Heemskerk, Kvitsborg, & Schumacher, 2013). However, approaches to determining whether these antigens (also called "neoantigens") are expressed have been relatively crude. For example, it has been proposed to use whole-genome (WGS) or whole-exome (WES) sequencing to identify mutations in genomic DNA, and then use RNA sequencing to identify genes that are expressed or differentially expressed in tumor material and focus all efforts on them (see, e.g., Heemskerk, Kvitsborg, & Schumacher, 2013). Similarly, it has been proposed that RNA sequencing data could be used instead of WGS / WES to identify tumor-specific mutations, with the advantage of not being limited to known genes but being able to identify, for example, intragenic fusions, novel transcripts, etc. However, this is complicated by factors such as differences in mRNA expression levels of any given transcript in tumor and normal tissues (making comparisons difficult), and it has been recognized that this approach is not suitable for identifying variants of RNA species present at low levels (see, for example, Heemskerk, Kvitsborg & Schumacher, 2013). Improved methods for determining tumor-specific variant expression are therefore needed. Summary of the Invention

[0003] The present inventors recognized that a simplistic gene-centric approach to RNA expression is insufficient in the context of cancer-specific mutations. This is because this approach combines information from gene variants and normal transcripts (both of which may be present in tumor cells) to provide a collective signal from healthy and tumor cells (because tumor samples are typically mixed samples containing both types of cells, the latter potentially being genetically heterogeneous), and as a whole, provides highly uncertain information about whether a variant is actually expressed. The present inventors discovered that approaches that simply examine gene expression levels within tumor samples as a filter for candidate neoantigens unnecessarily exclude many variants that may actually be expressed. Therefore, the present inventors recognized the need for a more sensitive approach that can take all of the above factors into account to more confidently identify the presence of variants in RNA expression data. In particular, the inventors have devised an approach for using RNA-seq data from one or more samples containing tumor cells or genetic material derived therefrom to determine the probability that a tumor-specific mutation is expressed in a tumor, and further, to determine whether the available RNA-seq data provides sufficient information to determine whether a tumor-specific mutation is expressed in a sample. This method has been found to be particularly useful for identifying neoantigens, for example, for purposes of cancer treatment or prognosis.

[0004] In particular, when evaluating whether T cell responses can be observed in vitro for multiple tumor-specific mutations identified, for example, from genome sequence data, the inventors identified many mutations that drive responses but are not identified as expressed in the RNA-seq data. In other words, the inventors identified a subset of mutations whose expression was not detected but that were found to be immunogenic. There are several possible explanations for this. For example, these mutations may be immunogenic but not expressed by immunoediting. In other words, the mutations may not be expressed at all. Another possibility is that the variant is expressed, but the RNA-seq data has low power to detect expression at this locus. To be able to distinguish between these two situations, the inventors devised a method to calculate the power of mutation detection in expression data. This method is particularly useful for identifying whether known mutations that may be expressed at low levels (when the mutation and / or locus are expressed at low levels and / or the tumor purity of the sample is low) are likely to be completely absent or false negatives (mutations that may be expressed but are not detected in the available data). This method can also be used to identify mutations directly from RNA expression data. This may be particularly useful for identifying splicing variants (e.g., retained introns and kipped exons) or any variants that cannot be directly identified using genomic data due to mappable issues such as fusions. We evaluated this method using a cell line titration series. In this series, high-confidence variants (e.g., single-nucleotide variants) could be identified in pure samples, their expression confirmed (one or more variant reads in the pure sample), and their expression assessed in dilution series containing 50%, 20%, or 10% variant cells. This demonstrated that false-negative mutations (i.e., mutations expressed in the cell line but not in the diluted samples) have low predictive power for allele-specific expression determined according to the new method.The method uses a rigorous statistical framework to classify individual variants as expressed and provide probabilities that reflect the confidence of the assignment. The method is fast, flexible, robust, and reproducible, is based on interpretable assumptions, and allows for the flexible incorporation of genotype and purity data in providing predictions.

[0005] Thus, according to a first aspect, there is provided a method of determining whether a tumor-specific mutation is likely to be expressed in a subject, the method comprising: providing or obtaining RNA-seq data from one or more samples from the subject comprising tumor genetic material, the RNA-seq data including, for each of the one or more samples, the number of RNA reads in the sample that exhibit the tumor-specific mutation (b) and the total number of RNA reads at the location of the tumor-specific mutation (d); and determining the likelihood of the sequence data if the tumor-specific mutation (i) is expressed (P(b,d|α,t,E=1)) and (ii) is not expressed (P(b,d|α,t,E=0)). The likelihood of the sequence data may also be referred to as the probability of observing the sequence data if the mutation is expressed and the probability of observing the sequence data if the mutation is not expressed.

[0006] The method of this aspect may have one or more of the following features.

[0007] The likelihood may be the probability of the sequence data given (conditional on) the tumor fraction of each of one or more samples and the proportion of the total number of reads at the tumor-specific mutation position that originate from a cell population (e.g., normal cells) within one or more samples that does not contain the tumor-specific mutation, where the samples further include a cell population (e.g., tumor cells) containing the tumor-specific mutation, representing a sample proportion equal to the tumor fraction. The likelihood varies depending on the probability of sampling a sequence read containing the tumor-specific mutation from each sample when the tumor-specific mutation is expressed or not expressed, and may depend on the sequencing error rate, the genotype of the tumor and normal cell populations, the tumor fraction of the sample, and the proportion of the total read count of the gene containing the tumor-specific mutation that originates from the normal cell population. The likelihood of the sequence data may be the probability of observing the sequence data given (conditional on) the tumor fraction of each of one or more samples (t) and the proportion of the total number of reads at the tumor-specific mutation position that originate from a cell population (α) within one or more samples that does not contain the tumor-specific mutation. Thus, a sample may contain a cell population that does not contain tumor-specific mutations and a tumor cell population that contains tumor-specific mutations, the latter representing a proportion of the sample equal to the tumor fraction. The likelihood may be determined using Equation (5) or Equation (5').

[0008] The method may further comprise comparing the resulting likelihoods to determine whether the tumor-specific mutation is likely to occur in the subject. Comparing the likelihoods determines the posterior probability (P(E=1|b, d, α, t)) of the tumor-specific mutation occurring, relative to the prior probability (μ ρ) and the likelihood of the sequence data when the tumor-specific mutation is (i) expressed (P(b,d|α,t,E=1)) and (ii) not expressed (P(b,d|α,t,E=0)). The posterior probability may be determined using either Equation (13) (particularly any of the different formulations of P(E=1|b,d,α,t) provided in Equation (13)), Equation (14) or Equation (14') (particularly any of the different formulations of P(E=1|b,d,α,t) provided in Equation (14) or Equation (14'). Comparing likelihoods determines the power to detect whether a tumor-specific mutation is expressed at a given false positive rate. The power to detect whether a tumor-specific mutation is expressed is determined by the number of RNA reads (b) in a sample that exhibit the tumor-specific mutation (P(b,d|α,t,E=1),P(b|E=1),P(b|M1)) at a threshold number of reads (b c ) is the area under the curve of the likelihood of the number of reads exhibiting tumor-specific mutations exceeding the threshold number of reads (b c ) may include determining the number of reads such that the area under the curve of the likelihood of the number of reads exhibiting the tumor-specific mutation above a threshold number of reads when the tumor-specific mutation is not expressed as a function of the number of RNA reads in the sample exhibiting the tumor-specific mutation (b) (P(b,d|α,t,E=0),P(b|E=0),P(b|M0)) is equal to a predetermined false positive rate. The power to detect the mutation being expressed may be determined using equation (8), where P(b|M1) is the likelihood of the sequence data when the mutation is expressed, and may be defined by equation (5) or (5'). The selection of the false positive rate may depend on the circumstances and user preference. For example, a false positive rate of 0.05 or 0.01 may be selected. The method may further include determining the area under the curve of the likelihood of the number of RNA reads in the sample exhibiting the tumor-specific mutation (b) above a threshold number of reads (b). c ) or more.

[0009] The inventors have devised a Bayesian framework for determining the probability that a tumor-specific mutation will be expressed, and a statistical hypothesis testing framework for determining whether a tumor-specific mutation will be expressed and quantifying the power of this test. Both frameworks are based on quantifying the likelihood (in terms of the total number of RNA reads at the position of the tumor-specific mutation) of observing the number of RNA reads containing the mutation when the tumor-specific mutation is expressed, and the likelihood (in terms of the total number of RNA reads at the position of the tumor-specific mutation) of observing the number of RNA reads containing the mutation when the tumor-specific mutation is not expressed. These are then compared to obtain a posterior probability (dependent on the likelihood ratio), or to quantify the true positive rate associated with the threshold number of reads containing the mutation that must be declared to be expressed as a tumor-specific mutation. Thus, both frameworks allow for rigorous determination of whether a tumor-specific mutation will be expressed, taking into account data on the expression at the mutation's locus, using quantities and parameters that are directly interpretable and linked to biological phenomena. The Bayesian framework can further incorporate informative priors (e.g., knowledge of whether genes containing tumor-specific mutations are expected to be expressed, or whether mutations in this disease type are typically expressed).

[0010] If the posterior probability of tumor-specific mutations being expressed is above a predetermined threshold, the tumor-specific mutation may be considered to be likely to be expressed. If the posterior probability of tumor-specific mutations being expressed is below a predetermined threshold, the tumor-specific mutation may be considered to be unlikely to be expressed. The predetermined threshold may be selected as a threshold that maximizes the true positive rate while keeping the false positive rate below a predetermined value (e.g., 0.05) in a set of tumor-specific mutations with known expression status. The set of tumor-specific mutations with known expression status may be obtained using one or more pure tumor samples (e.g., samples with tumor purity = 1). Such samples may be obtained by purification of one or more tumor cell lines. The predetermined threshold may be about 0.5, between 0.4 and 0.7, or between 0.5 and 0.6. The power to detect whether a tumor-specific mutation is expressed is above a predetermined threshold, and the number of RNA reads in a sample (b) that exhibit a tumor-specific mutation is less than a threshold number of reads (b c ) is greater than a threshold number of reads (b), a tumor-specific mutation may be considered likely to be expressed. The power to detect whether a tumor-specific mutation is expressed is greater than a predetermined threshold, and the number of RNA reads in a sample that exhibit a tumor-specific mutation (b) is greater than a threshold number of reads (b c ), the likelihood of tumor-specific mutations being expressed may be considered low. If the power to detect whether a tumor-specific mutation is expressed is below a predetermined threshold and the number of RNA reads in a sample that exhibit a tumor-specific mutation (b) is below a threshold number of reads (b c ), the tumor-specific mutation may be considered likely to be expressed. In such cases, the RNA-seq data alone may be insufficient to confidently determine whether the tumor-specific mutation is truly expressed or not, and therefore the tumor-specific mutation may be considered likely to be expressed by default. This contrasts with previous approaches that simply ignore tumor-specific mutations detected at low levels in RNA-seq data, regardless of whether this low level truly indicates a lack of expression or instead arises from limitations of the data itself.

[0011] Obtaining RNA-seq data including the number of RNA reads in the sample that exhibit the tumor-specific mutation (b), the number of RNA reads in the sample that exhibit the corresponding germline allele, and the total number of RNA reads at the position of the tumor-specific mutation (d) may alternatively include obtaining RNA-seq data including at least two of the number of RNA reads in the sample that exhibit the tumor-specific mutation (b), the number of RNA reads in the sample that exhibit the corresponding germline allele(s) (or that do not exhibit the tumor-specific mutation), and the total number of RNA reads at the position of the tumor-specific mutation (d). The RNA-seq data may include RNA-seq data from multiple samples from the subject. The multiple samples may be tumor samples. The posterior probability that the tumor-specific mutation is expressed may be the posterior probability that the tumor-specific mutation is ubiquitously expressed in the subject's tumor. Such a posterior probability may be determined using Equation (14') (particularly any of the formulations of Equation (14')).

[0012] Each sample was genotyped to contain at least one copy of the tumor-specific mutation (G v ) and a first cell population with a genotype (G Nand a second cell population, each having a genotype (t), where the first cell population represents a proportion equal to the tumor fraction (t) for each sample. The likelihood may be obtained as a sum of multiple joint genotypes, including the genotype for the first cell population and the genotype for the second cell population. The sum may be a weighted sum, where the likelihood given each respective joint genotype is weighted by the probability associated with the joint genotype. The probabilities associated with each joint genotype may sum to 1. The second cell population may be assumed to have a homozygous diploid reference allele genotype (AA). The joint genotype may be determined using estimates of major and minor copy numbers of loci derived from DNA sequence data of the sample or related samples. A joint genotype may be determined from the estimates of the major and minor copy numbers for a locus, assuming that the possible genotypes of the first cell population include any genotypes with a copy number equal to the sum of the major and minor copy numbers and a number of copies of the variant allele between the minor and major copy numbers. The possible genotypes of the first cell population may include genotypes in which a mutation to a variant allele is assumed to have occurred before or prior to any copy number event (a loss or gain of copy number at the locus compared to the diploid genotype). Each of the possible genotypes considered may be associated with equal probability. The tumor-specific mutation may be assumed to be ubiquitous in one or more samples. The tumor-specific mutation may be a clonal mutation or a mutation assumed to be clonal in the subject's tumor. The method may include determining whether the tumor-specific mutation is likely to be clonal in the subject. Accordingly, a method for determining whether a clonal tumor-specific mutation is likely to be expressed in a subject is also described herein.

[0013] The likelihood of sequence data when a tumor-specific mutation is (i) expressed and (ii) not expressed may depend on the proportion (α) of total expression of the gene containing the tumor-specific mutation (reads per million or transcripts per million, TPM) originating from cell populations within one or more samples that do not contain the tumor-specific mutation. Reference to total expression of a gene containing a tumor-specific mutation may refer to the total expression of one or more transcripts derived from the gene. The parameter α may be set to a predetermined default value, or α may be estimated from data including RNA-seq data of multiple samples with different tumor fractions. For example, α may be set to a predetermined default value of α=0.5, which is equivalent to assuming that tumor cells and normal cells contribute equally to the total expression at the locus. Alternatively, α may be set to a value corresponding to the expression level of the gene in one or more tumor and normal samples. The expression level of the gene in one or more tumor and normal samples may be obtained from a database. For example, the expected expression of the gene in samples from the same tumor type and normal samples, or the differential expression between one or more samples from the same tumor type and normal samples, may be used. If the genotype of a cell population in one or more samples has a genotype containing at least one copy of a tumor-specific mutation but does not have any non-mutated alleles at the locus of the tumor-specific mutation, α may be set to 1 when determining the likelihood of the sequence data when the tumor-specific mutation is not expressed. The parameter α may be estimated from data including RNA-seq data for multiple samples with various tumor fractions, and α is derived from the slope (g) of a regression model (e.g., a linear regression model) fitted to the total expression of genes containing tumor-specific mutations in multiple samples as a function of their tumor purity. The value of the parameter α may be set to

[0014]

number

[0015] The parameters may be estimated separately for each gene i containing a tumor-specific mutation. Thus, the above formula can also be written as:

[0016]

number

[0017] The probability of tumor-specific mutations occurring is calculated by the mean μ of the beta distribution (p(ρ)) for the parameter ρ of the Bernoulli distribution for the variable (E) that reflects whether tumor-specific mutations occur. ρ The parameter μ ρ may be set to a predetermined value. If the total number of RNA reads at the position of the tumor-specific mutation is above or equal to a predetermined threshold, the parameter μ ρ may be set to a first predetermined value. The predetermined threshold may be 0 (d>0). μ is determined depending on whether the gene containing the tumor-specific mutation is expected to be expressed when the total number of RNA reads at the position of the tumor-specific mutation is below the predetermined threshold. ρmay be set to a first predetermined value or a second predetermined value. A gene containing a tumor-specific mutation may be considered to be predicted to be expressed if the total expression of the gene in the sample is above a predetermined threshold, and may be considered to be predicted to be unexpressed otherwise. ρ The first predetermined value of μ may be 0.5. ρThe second predetermined value of may be less than 0.5 (e.g., 0.2, 0.1, 0.05, or 0.01). The predetermined threshold for the total expression of genes containing tumor-specific mutations may be 1 TPM (transcripts per million). The predetermined threshold for the total number of RNA reads at the position of a tumor-specific mutation may be 0. The predetermined threshold for the total expression of genes containing tumor-specific mutations may be set according to the expected level of expression of the genes in one or more tumor samples, for example, based on the expression level of the genes in one or more additional tumor samples (e.g., tumor samples from tumors of the same type as the subject's tumor). The expression level of the genes in the one or more additional tumor samples may be obtained, for example, from a database. The value of the prior probability that the mutation will be expressed may depend on the subject, the tumor, the mutation, or a combination thereof. For example, the value may be determined using previously obtained data on a related patient cohort, such as patients suffering from the same type or subtype of cancer. Alternatively, the value may be arbitrarily set based on prior knowledge of the cancer type or mutation. For example, a specific mutation discovered across multiple cancer samples and identified as frequently expressed in these samples may be assigned a probability greater than 0.5. Therefore, we adapted the naive Bayesian framework with a single prior probability to instead include a prior probability with multiple possible values ​​in cases where there is insufficient evidence from the data (i.e., when the total number of reads at the tumor-specific mutation locus is below a predetermined threshold). The multiple possible values ​​include at least a first value when there is evidence that the gene is expressed and a second value when there is insufficient evidence that the gene is expressed. For example, if TPM > 1 = 1, the gene may be assumed to be expressed, and the prior probability may be set to 0.5. This reflects the assumption that the lack of read detection at a locus (whether a variant allele or a reference allele) is not indicative of whether the tumor-specific mutation is expressed, but is likely the result of technical issues. If TPM = 0 at a locus, the gene may be assumed not to be expressed, and the prior probability may be set to a value less than 0.5.This reflects the assumption that if a gene is not expressed, it is unlikely that a tumor-specific mutation will also be expressed.

[0018] The likelihood may be conditional on the tumor purity of one or more samples. The tumor purity of one or more samples may be estimated using genomic sequence data from one or more samples. For example, the tumor purity may be estimated using methods known in the art, such as ASCAT or Sequenza. The tumor purity of a sample may be the most reliable purity (such as a maximum a posteriori estimate) obtained from a probabilistic method for estimating tumor purity from DNA sequence data. The genotypes (also referred to as "joint genotypes") of the first and second cell populations may be set to a predetermined genotype, for example, the first cell population may be considered to be a heterozygous diploid containing a copy of a tumor-specific mutation and a copy of a non-mutated allele, and the second cell population may be considered to be a diploid with two non-mutated alleles. Alternatively, the genotypes of the first and second cell populations may be estimated using methods known in the art, such as ASCAT or Sequenza. The genotype of the sample may be the most reliable genotype (e.g., a maximum a posteriori joint genotype) obtained from a probabilistic method for estimating the genotype of a mixed population from DNA sequence data. The probability of observing the sequence data may depend on the genotypes of the first and second cell populations, the total number of reads at the tumor-specific mutation locus, the tumor fraction (t), the proportion of total expression of genes containing the tumor-specific mutation originating from the second cell population (α), and the ratio of expression of allele(s) not containing the tumor-specific mutation to total expression at the tumor-specific mutation locus (θ). The ratio of expression of allele(s) not containing the tumor-specific mutation to total expression at the tumor-specific mutation locus (θ) may be assumed to be a random variable having a first distribution when the tumor-specific mutation is expressed (i.e., for the purpose of estimating the likelihood of the data when the tumor-specific mutation is expressed) and a second distribution when the tumor-specific mutation is not expressed (i.e., for the purpose of estimating the likelihood of the data when the tumor-specific mutation is not expressed). The second distribution may be a beta distribution with parameters α0 and β0, and the first distribution may be a beta distribution with parameters α1 and β1.Any of the following parameters may be used alone or in combination: α0>1 (e.g., α0=9999), β0=1, α1=1, and β1=1.

[0019] The likelihood may be a marginal likelihood obtained from the probability of observing the sequence data by integrating the ratio (θ) of expression of alleles that do not contain tumor-specific mutations to total expression at the tumor-specific mutation locus. Thus, determining the likelihood may involve numerical integration, e.g., by a processor, of the probability of observing the sequence data over all possible values ​​(i.e., 0 to 1) of the ratio (θ) of expression of alleles that do not contain tumor-specific mutations to total expression at the tumor-specific mutation locus. In other words, the likelihood may be the integral, over all possible values ​​θ, of the probability of observing the sequence data multiplied by the distribution of the ratio (θ) of expression of alleles that do not contain tumor-specific mutations to total expression at the tumor-specific mutation locus. The likelihood may be obtained as a sum of integrals (e.g., sum of marginal likelihoods) for multiple possible genotypes of the first and second cell populations. The sum may be weighted by the probability of each of the multiple possible genotypes.

[0020] Genotype G of the first and second cell populations v ,G N The probability of observing sequence data conditional on the total number of reads at the tumor-specific mutation locus, the tumor fraction (t), the proportion of total expression of genes containing the tumor-specific mutation originating from the second cell population (α), and the ratio of expression of alleles not containing the tumor-specific mutation to total expression at the tumor-specific mutation locus (θ) may be assumed to follow a binomial distribution with parameters ξ(G,θ,α,t) representing the probability of sampling a read containing a tumor-specific mutation from the tumor sample, where ξ(G,θ,α,t) is given by any of Equation (2), Equation (2'), or Equation (2"), and optionally

[0021]

number

[0022]

number

[0023] The method may be implemented by a computer. Thus, the step of acquiring RNA-seq data may be performed by a processor, and the step of determining the likelihood may be performed by the processor. The step of acquiring RNA-seq data may include receiving RNA-seq data including sequence reads from one or more samples from the subject, and determining from the sequence reads at least two of the number of RNA reads in the sample that exhibit the tumor-specific mutation (b), the number of RNA reads in the sample that exhibit the corresponding germline allele(s) (or do not exhibit the tumor-specific mutation), and the total number of reads at the position of the tumor-specific mutation (d). The step of determining the posterior probability of the tumor-specific mutation being expressed and / or the power of detecting the tumor-specific mutation being expressed may be performed by a computer. The step of determining the posterior probability of the tumor-specific mutation being expressed and / or the power of detecting the tumor-specific mutation being expressed may include a step of numerical integration to obtain the likelihood. In particular, this step may involve determining the posterior probability of a mutation being expressed given the prior probability of the mutation being expressed, and the probability of observing sequence data when the tumor-specific mutation (i) is expressed and (ii) is not expressed, by solving multiple one-dimensional integrals (e.g., a pair of integrals for each sample representing the assumption that the mutation is expressed and the assumption that the mutation is not expressed) that integrate the probability of observed sequence data over all possible values ​​of the proportion of RNA reads from tumor cells in the sample that contain the tumor-specific mutation (θ). These numerical integrals may be solved independently (e.g., in parallel) for each sample, each mutation, and each model (i.e., assuming that the tumor-specific mutation is (i) expressed and (ii) not expressed). Similarly, determining the power of detecting a tumor-specific mutation being expressed may involve solving multiple one-dimensional integrals (e.g., a pair of integrals for each sample representing the assumption that the mutation is expressed and the assumption that the mutation is not expressed) that integrate the probability of observed sequence data over all possible values ​​of the proportion of RNA reads from tumor cells in the sample that contain the tumor-specific mutation (θ). The providing step may include one or more steps, all or part of which are computer-implemented.

[0024] The method may further include obtaining or providing at least one estimate of tumor fraction for each sample, and optionally at least one corresponding set of one or more candidate joint genotypes, where the joint genotypes include the genotypes of a cell population not containing a tumor-specific mutation (e.g., a normal cell population) and the genotypes of a cell population containing a tumor-specific mutation (e.g., a tumor population). The estimate of tumor fraction may be obtained using a method for determining an allele-specific copy number profile in a sample containing a mixture of tumor cells and normal cells. Methods for doing this using genome sequencing or array data are known in the art, for example, expressing allele-specific data as a function of parameters (including allele-specific copy number, tumor aneuploidy, and tumor cell fraction) and identifying the values ​​of these parameters that best fit all the data. Examples of such methods include, among others, ASCAT (Van Loo et al., 2010). Alternatively, the estimate of tumor fraction may be determined experimentally. Thus, the method may further include obtaining an estimate of tumor fraction for each of one or more samples. In particular, the method may include obtaining, by a processor, for each sample at least one estimate of tumor fraction, where the processor uses genomic sequence data from one or more samples to determine estimates of tumor fraction and allele-specific copy number, and determining, by the processor, a set of one or more candidate joint genotypes associated with the allele-specific copy numbers. For tumor cells in the mixed sample, the allele-specific copy numbers, or variables derived therefrom (or conversely, variables from which such allele-specific copy numbers can be derived, examples of which are B allele fraction and logR), may be used to obtain a set of one or more candidate genotypes.The allele-specific copy number of tumor cells in a mixed sample may be obtained using a method for determining the allele-specific copy number profile in a sample containing a mixture of tumor cells and normal cells, examples of which include, among others, ASCAT (Van Loo et al., 2010), Sequenza (Favero et al., 2015), or ascatNgs (Raine et al., 2016). Thus, the method may further include, for each of one or more samples, obtaining estimates for at least two of the copy number of the major allele in tumor cells in the sample, the copy number of the minor allele in tumor cells in the sample, and the total copy number at the position of the tumor-specific mutation in tumor cells in the sample. The estimate of the copy number of tumor cells in the sample may represent a summarized (e.g., average) estimate for the entire population of tumor cells in the sample.

[0025] The method may further include repeating the method for multiple tumor-specific mutations identified in the subject. The method may further include ranking or otherwise prioritizing the multiple tumor-specific mutations based at least in part on the probability that they are determined to be expressed in the subject. The method may further include identifying one or more tumor-specific mutations in the subject. Identifying one or more tumor-specific mutations in the subject may be performed using genome or transcriptome sequence data from one or more samples from the subject containing tumor genetic material and sequence data from one or more germline samples from the subject, for example, by comparing the sequence data. Identifying one or more tumor-specific mutations in the subject may include aligning sequence data from at least one sample containing tumor genetic material to a reference sequence and identifying positions where the sequence of the sample differs from the reference sequence. The method may further include aligning sequence data from at least one germline sample to the reference sequence and identifying positions where the sequence of the sample containing tumor genetic material differs from the germline sample. The reference sequence may be a reference genome or a reference transcriptome.

[0026] Providing sequence data from one or more samples from the subject may include or consist of receiving sequence data from a user (e.g., via a user interface), from one or more computing device(s), or from one or more data stores or databases. Providing sequence data (whether genomic sequence data, RNA sequence data, or transcriptome sequence data) may further include sequencing one or more samples from the subject containing tumor genetic material (or otherwise determining the sequence composition of the genetic material present in the samples). The method may further include sequencing one or more germline samples from the subject (or otherwise determining the sequence composition of the genomic material present in the samples). The method may further include obtaining one or more samples containing tumor genetic material and, optionally, one or more germline samples from the subject. Genetic material, as used herein, includes RNA molecules (e.g., mRNA transcripts) and, optionally, DNA molecules (e.g., genomic DNA). The method may further include providing, e.g., via a user interface, the determined probability that the tumor-specific mutation will be expressed and / or a value derived therefrom or associated therewith, examples of which include the likelihood of the sequence data assuming that the tumor-specific mutation will or will not be expressed, the relative likelihood of the sequence data assuming that the tumor-specific mutation will or will not be expressed, and / or the power of detecting an expressed tumor-specific mutation given a distribution of the likelihood of the sequence data assuming that the tumor-specific mutation will or will not be expressed at one or more false positive rates. For example, the method may include providing an "expression status" flag or value based on the determined probability that the tumor-specific mutation will be expressed and / or the likelihood and power value of the tumor-specific mutation being expressed. As another example, the method may include providing information identifying the mutation (e.g., the sequence of the mutation and its genomic location, etc.). According to a further aspect, a method is provided for identifying one or more neoantigens in a subject.The method includes identifying a plurality of tumor-specific mutations in a subject (e.g., using genomic and / or transcriptomic data from one or more tumor samples from the subject), and using the method of any embodiment of the preceding aspect to determine whether one or more of the tumor-specific mutations are likely to be expressed in the subject's tumor, and determining whether one or more of the tumor-specific mutations are likely to give rise to neo-antigens, where the neo-antigens are tumor-specific mutations that meet one or more predetermined criteria for whether they are likely to be expressed in the tumor, and optionally one or more further criteria for whether they are likely to give rise to neo-antigens. Also according to this aspect, a method of identifying one or more neo-antigens in a subject is described. The method includes: identifying, by a processor, a plurality of tumor-specific mutations in a subject using sequence data from one or more samples from the subject; determining, by the processor, whether one or more of the tumor-specific mutations are likely to occur in the subject using a method as described in any preceding claim; and selecting, by the processor, one or more of the tumor-specific mutations as candidate neo-antigens, wherein the candidate neo-antigens are tumor-specific mutations that satisfy at least one or more predetermined criteria as to whether the tumor-specific mutation is likely to occur and, optionally, one or more criteria as to whether the tumor-specific mutation is likely to give rise to the neo-antigen.

[0027] The method of this aspect may have any one or more of the following features.

[0028] The neo-antigen may be a tumor-specific mutation that meets at least a criterion selected from: having a probability of expression above a predetermined threshold; having a probability of expression above a threshold adaptively set to select a predetermined number of tumor-specific mutations with the highest probability of expression from among the tumor-specific mutations whose probabilities have been determined; having a probability of expression above a threshold adaptively set to select a predetermined top percentile of tumor-specific mutations from among the tumor-specific mutations whose probabilities have been determined; having a power to detect expressed mutations above a predetermined threshold and a number of RNA reads indicating tumor-specific mutations above a threshold number associated with the power to detect expressed mutations; and having a power to detect expressed mutations below a predetermined threshold and a number of RNA reads indicating tumor-specific mutations below a threshold number associated with the power to detect expressed mutations.

[0029] The method may further include determining whether one or more of the tumor-specific mutations are likely to be clonal in the tumor of interest and identifying whether one or more of the tumor-specific mutations are likely to give rise to clonal neo-antigens. Clonal neo-antigens may be tumor-specific mutations that meet at least a criterion selected from: having a probability of being clonal above a predetermined threshold; having a probability of being clonal above a threshold adaptively set to select a predetermined number of tumor-specific mutations with the highest probability of being clonal from among the tumor-specific mutations whose probabilities have been determined; or having a probability of being clonal above a threshold adaptively set to select a predetermined top percentile of tumor-specific mutations from among the tumor-specific mutations whose probabilities have been determined. Thus, one or more predetermined criteria for whether a tumor-specific mutation is likely to be clonal may be selected from: mutations that have a likelihood of being clonal above a predetermined threshold; and mutations that have a likelihood of being clonal above a threshold that is adaptively set to select a predetermined number of tumor-specific mutations whose likelihoods have been determined that have the highest likelihood of being clonal, and that also have a likelihood of being clonal above a threshold that is adaptively set to select a predetermined top percentile of tumor-specific mutations whose likelihoods have been determined.

[0030] A neoantigen may be a tumor-specific mutation that meets at least one of the following criteria: predicted to result in a protein or peptide that is not expressed in the subject's normal cells; predicted to result in at least one peptide that is likely to be presented by an MHC molecule; predicted to result in at least one peptide that is likely to be presented by an MHC allele known to be present in the subject; or predicted to result in a protein or peptide that is immunogenic. For example, a neoantigen may be a tumor-specific mutation that meets the criteria of predicted to result in a change in the protein sequence (e.g., because it encodes, because it affects a splice site, because it results in a truncated peptide, etc.), thus resulting in a protein or peptide that is not necessarily expressed in the subject's normal cells. Whether this is the case or not may be further confirmed, for example, by comparison with the subject's predicted normal proteome. Thus, one or more criteria for whether a tumor-specific mutation is likely to give rise to a neoantigen may be selected from: a mutation associated with an expression product expressed in tumor cells; a mutation predicted to give rise to a protein or peptide not expressed in normal cells of the subject; a mutation predicted to give rise to at least one peptide likely to be presented by an MHC molecule; a mutation predicted to give rise to at least one peptide likely to be presented by an MHC allele known to be present in the subject; and a mutation predicted to give rise to a protein or peptide that is immunogenic.

[0031] The method may further include identifying one or more peptides associated with one or more neo-antigens (i.e., one or more peptide sequences predicted to be present in tumor cells as a result of the presence of tumor-specific mutations, where the tumor-specific mutations meet one or more criteria (e.g., criteria related to probability of expression, likelihood of giving rise to neo-antigens, and / or likelihood of giving rise to clonal neo-antigens), as described above.

[0032] As one skilled in the art will appreciate, the complexity of the operations as described herein (due at least to the complexity of obtaining posterior probabilities, which requires numerical integration as described herein, and the amount of data typically generated by sequencing genetic material) is beyond the scope of mental activity. Thus, unless the context indicates otherwise (e.g., when sample preparation or acquisition steps are described), all steps of the methods described herein are computer-implemented.

[0033] According to a third aspect, there is provided a method of determining whether a tumor-specific mutation is likely to give rise to a neoantigen, the method comprising: determining whether a tumor-specific mutation is likely to be expressed in a tumor of a subject using the method of any embodiment of the first aspect; and determining whether the tumor-specific mutation satisfies one or more predetermined criteria, and optionally one or more further criteria, that are applied to the result of the step of determining whether the tumor-specific mutation is likely to be expressed.

[0034] The method of this embodiment may have any one or more of the features of any of the previous embodiments.

[0035] According to a fourth aspect, there is provided a method of providing immunotherapy to a subject diagnosed with cancer, the method comprising identifying one or more neo-antigens that are likely to be expressed in the subject's tumor using a method as described herein (such as a method according to any embodiment of the first or second aspect (e.g., by identifying one or more tumor-specific mutations in the subject)), determining whether one or more of the tumor-specific mutations are likely to be expressed in the subject's tumor using a method of any embodiment of the first aspect, selecting one or more tumor-specific mutations from the identified tumor-specific mutations based on the result of the determining and optionally one or more further criteria (examples of which are described, for example, in relation to the second aspect), and designing an immunotherapy that targets one or more of the identified neo-antigens.

[0036] According to a further aspect, a method of treating a subject diagnosed with cancer is provided. The method includes identifying one or more neo-antigens in the subject by identifying multiple tumor-specific mutations, determining whether one or more of the tumor-specific mutations are likely to be expressed in the subject, selecting one or more of the tumor-specific mutations as candidate neo-antigens, where the candidate neo-antigens are tumor-specific mutations that meet at least one or more predetermined criteria for whether the tumor-specific mutation is likely to be expressed, and treating the subject with an immunotherapy targeting one or more of the selected candidate neo-antigens. According to this aspect, determining whether the subject is likely to be expressed the tumor-specific mutations includes obtaining, by a processor, RNA-seq data from one or more samples from the subject containing tumor genetic material, the RNA-seq data including, for each of the one or more samples, a number of reads in the sample that indicate the tumor-specific mutation (b) and a total number of reads at the position of the tumor-specific mutation (d), and determining, by the processor, a probability of observing the RNA-seq data when the tumor-specific mutation (i) is expressed and (ii) a probability of observing the RNA-seq data when it is not expressed. The method may further include determining a posterior probability that the tumor-specific mutation will be expressed as a function of the prior probability that the mutation will be expressed and the probability of observing the RNA-seq data when the tumor-specific mutation is (i) expressed and (ii) not expressed. The method may have any of the features described in connection with any of the previous embodiments.

[0037] The method may have any one or more of the following features.

[0038] The present disclosure also relates to immunotherapies targeting one or more neoantigens associated with tumor-specific mutations determined to be expressed in a tumor using the methods described herein, and methods for designing and / or providing such immunotherapies. According to any embodiment as described herein, the immunotherapy may be an immunogenic composition, a composition comprising immune cells, or a therapeutic antibody. The immunogenic composition may include one or more identified neoantigens (e.g., neoantigen peptides or proteins, or cells that display the neoantigens, etc.), or a substance sufficient for expression of the identified one or more neoantigens (e.g., DNA or RNA molecules encoding the neoantigen(s)). The composition comprising immune cells may include T cells, B cells, and / or dendritic cells. The composition comprising therapeutic antibodies may include one or more antibodies that recognize at least one of the identified neoantigens. The antibodies may be monoclonal antibodies.

[0039] In any embodiment of any aspect, the cancer may be selected from bladder cancer, gastric cancer, esophageal cancer, breast cancer, colon cancer, cervical cancer, ovarian cancer, endometrial cancer, kidney cancer (renal cell), lung cancer (small cell, non-small cell, and mesothelioma), brain cancer (glioma, astrocytoma, glioblastoma), melanoma, lymphoma, small intestine cancer (duodenal and jejunal), leukemia, pancreatic cancer, hepatobiliary tumors, germ cell cancer, prostate cancer, head and neck cancer, thyroid cancer, and sarcoma. The cancer may be lung cancer. The cancer may be melanoma. The cancer may be bladder cancer. The cancer may be head and neck cancer. In any embodiment of any aspect, the subject may be human. Designing an immunotherapy targeting the identified neo-antigen(s) may include designing one or more candidate peptides for each of the one or more targeted neo-antigens, each peptide comprising at least a portion of the targeted neo-antigen. Designing or providing an immunotherapy may include obtaining one or more candidate peptides. The method may further include assaying the one or more candidate peptides for one or more properties. The assaying may be performed in vitro or in silico. For example, the one or more peptides may be assayed for immunogenicity, propensity to be displayed by MHC molecules (optionally by specific MHC molecule alleles, where the alleles may be selected according to the MHC alleles expressed by the subject), ability to induce proliferation of an immune cell population, etc. The method may further include delivering the immunotherapy. The method may further include obtaining a dendritic cell population pulsed with one or more of the candidate peptides. The immunotherapy may be a composition comprising T cells that recognize at least one of the one or more identified neo-antigens. The composition may be enriched for T cells that target at least one of the one or more identified neo-antigens. The method may include obtaining a T cell population and expanding the T cell population to increase the number or relative proportion of T cells that target at least one of the one or more identified neo-antigens. The method may further include obtaining the T cell population.A T cell population may be isolated from a subject, for example, from one or more tumor samples obtained from the subject, or from a peripheral blood sample or other tissue sample from the subject. The T cell population may include tumor-infiltrating lymphocytes. T cells may be isolated using methods well known in the art. For example, T cells may be purified from a single-cell suspension generated from a sample based on expression of CD3, CD4, or CD8. T cells may be enriched from a sample by passage through a Ficoll-Opaque gradient. The method may further include expanding the T cell population. For example, T cells may be expanded by ex vivo culture under conditions known to provide mitogenic stimulation to T cells. Illustratively, T cells may be cultured with cytokines such as IL-2 or mitogenic antibodies such as anti-CD3 and / or CD28. T cells may also be co-cultured with antigen-presenting cells (APCs), which may be irradiated. The APCs may be dendritic cells or B cells. Dendritic cells may be pulsed with peptides containing one or more identified neoantigens, either as single stimulators or as a pool of stimulatory neoantigen peptides. T cell expansion may be performed using methods known in the art, including, for example, the use of artificial antigen-presenting cells (aAPCs) to provide additional costimulatory signals, and the use of autologous PBMCs presenting appropriate peptides. Autologous PBMCs may be pulsed with neoantigen-containing peptides, as discussed herein, either as a single stimulatory agent or, alternatively, as a pool of stimulatory neoantigens.

[0040] According to a further aspect, a method is provided for expanding a T cell population for use in treating cancer in a subject. The method includes identifying one or more neo-antigens using a method as described herein (e.g., a method according to any embodiment of the second aspect), obtaining a T cell population comprising T cells capable of specifically recognizing one of the identified neo-antigens, and co-culturing the T cell population with a composition comprising the identified neo-antigen. The method may have one or more of the following features: The obtained T cell population may comprise T cells capable of specifically recognizing one of the identified neo-antigens. The method preferably includes identifying multiple neo-antigens. The neo-antigens may be clonal neo-antigens. The T cell population may comprise multiple T cells, each of which specifically recognizes one of the multiple identified neo-antigens, and the T cell population may be co-cultured with a composition comprising the multiple identified neo-antigens. The co-culture may result in expansion of the T cell population specifically recognizing one or more neo-antigens. The expansion may be carried out by co-culturing the neo-antigen-containing T cells with antigen-presenting cells. The antigen-presenting cells may be dendritic cells. Thus, the expansion may be selective expansion of T cells specific for the neo-antigen. The expansion may further include one or more non-selective expansion steps.

[0041] According to a further aspect, there is provided a composition comprising a T cell population obtained or obtainable by a method according to any embodiment of the preceding aspects. According to a further aspect, there is provided a composition comprising a neoantigen, a neoantigen-specific immune cell, or an antibody recognizing a neoantigen, for use in treating or preventing cancer in a subject, wherein the neoantigen has been identified as a neoantigen using the methods described herein (e.g., identified as derived from a tumor-specific mutation expressed in a tumor of the subject). According to a further aspect, there is provided a composition comprising a neoantigen, a neoantigen-specific immune cell, or an antibody recognizing a neoantigen, wherein the neoantigen has been identified as a neoantigen using the methods described herein (e.g., identified as derived from a tumor-specific mutation expressed in a tumor of the subject). According to a further aspect, there is provided a cell or population of cells expressing a neoantigen on its surface, wherein the neoantigen has been identified as a neoantigen using the methods described herein (e.g., identified as derived from a tumor-specific mutation expressed in a tumor of the subject). According to a further aspect, there is provided a neoantigen, a neoantigen-recognizing immune cell, or an antibody recognizing a neoantigen, for use in treating or preventing cancer in a subject. wherein the neoantigen has been identified as a neoantigen using the methods described herein (e.g., identified as resulting from a tumor-specific mutation expressed in the subject's tumor). According to a further aspect, there is provided use of a neoantigen, an immune cell that recognizes a neoantigen, or an antibody that recognizes a neoantigen, in the manufacture of a medicament for use in treating or preventing cancer in a subject, wherein the neoantigen has been identified as a neoantigen using the methods described herein (e.g., identified as resulting from a tumor-specific mutation expressed in the subject's tumor). According to a further aspect, there is provided a method of treating a subject diagnosed with cancer, the method comprising administering an immunotherapy provided using the methods described herein, or a composition as described herein. According to a further aspect, a system is provided.The system includes a processor and a computer-readable medium comprising instructions that, when executed by the processor, cause the processor to perform the steps of any method described herein, such as a method according to any embodiment of the first, second, third or fourth aspect above.

[0042] According to a further aspect, there is provided one or more non-transitory computer-readable media comprising instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of any method described herein, such as a method according to any embodiment of the first, second, third or fourth aspect above.

[0043] According to a further aspect, there is provided a computer program comprising code which, when executed on a computer, causes the computer to perform the steps of any method described herein, such as a method according to any embodiment of the first, second, third or fourth aspect above. [Brief explanation of the drawings]

[0044] [Figure 1A] Schematic of the problem of determining allele-specific expression in a homogenous diploid cell population. [Figure 1B] Schematically illustrates the problem of determining allele-specific expression in a mixed sample containing variant cells (e.g., tumor cells) and reference cells (e.g., healthy cells). [Figure 1C]Schematic depiction of a model for determining the power of detection of mutant allele expression in a mixed sample. The model's variables for detecting a mutant allele (here, denoted "C") in a mixed sample include a variant cell population (e.g., tumor cells, shown here as having a heterozygous genotype Gv for the mutant allele) that has at least one copy of the allele, and a reference cell population (e.g., healthy cells, shown here as containing only the reference allele, denoted "T"; in the illustrated example, the healthy cells have a homozygous genotype GN for T) that does not have any copies of the allele. The sample contains a proportion t of variant cells (also referred to as "tumor purity"). Expression data consists of a total number d of reads covering a locus, of which b reads represent the variant allele. The set of genotypes in the population is denoted G = {Gv,GN}. Variant cells express both the reference allele and the variant allele in a proportion θ (the proportion of expression in tumor cells due to the reference allele), and healthy cells express a proportion α of the total number of reads observed at the locus. These variables are used to determine the likelihood of observing a given number b of reads containing the variant allele, based on a null model (M0) in which the variant allele is not expressed, and based on an alternative model (M1) in which the variant allele is expressed. In the example shown, this is estimated as the binomial probability of a variant read in sample b, given that the total number of reads is d and the probability of sampling a read containing a mutation from a mixed sample is given by ξ. [Figure 1D]Schematic representation of model characteristics for determining the power of variant allele expression in mixed samples. Null and Alternative Model Likelihood. The likelihood of the null model (no expression, also denoted P(b|M0) or P(b,d|α,t,E=0)) decreases as the number of variant reads increases. The likelihood of the alternative model (variant expression, also denoted P(b|M1) or P(b,d|α,t,E=1)) has a maximum that depends on the model parameters. Therefore, using the likelihood of M0, it is possible to identify a critical value bc such that the probability of rejecting the null model when it is correct (P(false positive)), which occurs when b > bc, is equal to a given value (the area under the likelihood curve for M0 to the right of bc). The probability of correctly identifying the actual expressed variant using this bc is given by the area under the likelihood curve for M1 to the right of bc. This is the power of the statistical test to accept or reject M0 based on the likelihood of the lead data under M0 and M1. [Figure 1E] C and D show how the probability that a variant will be expressed (P(E=1|b,d,α,t)) can be determined using the models described. The two plots in the upper left corner show the likelihood of the null model (M0,E=0) and the alternative model (M1,E=1), as shown in D. The plot in the upper right corner shows the prior probabilities for the parameter θ of the null and variant models, the distribution of the variable E (E=1, with the variant expressed, or E=0, without), and the hyperprior probabilities for the parameter ρ of the distribution of the variable E, according to an example of the described method. [Figure 2A] FIG. 1 is a flow chart illustrating a method for determining whether a tumor-specific mutation is likely to be clonal, and its use in identifying clonal neoantigens. [Figure 2B-1] 1 is a flow chart that schematically illustrates a method of providing immunotherapy. [Figure 2B-2] Continued from Figure 2B-1. [Figure 3]1 illustrates an embodiment of a system for determining whether a tumor-specific mutation is likely to be clonal, and / or for identifying clonal neoantigens, and / or for providing immunotherapy. [Figure 4] 1 shows a schematic representation of the model used to evaluate sequence data according to the methods disclosed herein. [Figure 5] (A) Schematic representation of a model used to determine the probability of a variant being expressed according to the methods disclosed herein based on data from a single tumor sample. (B) Schematic representation of a model used to determine the probability of a variant being expressed according to the methods disclosed herein based on data from multiple tumor samples. [Figure 6] The plot shows the behavior of the posterior probability of a variant being expressed (y-axis) as a function of the relative likelihood (r, x-axis) of a model assuming the variant is expressed versus a model assuming the variant is not expressed. The plot shows curves for each of several parameters μρ (the mean of the beta distribution for ρ, used as the hyperprior for ρ, where ρ is a parameter of the Bernoulli variable prior for E, and E is a binary variable reflecting whether the variant is expressed or not). The curves shown are for increasing values ​​of μρ, from the bottom curve (μρ=0.1) to the top curve (μρ=0.9). [Figure 7] Schematic showing how the proportion of total expression at a locus due to the normal cell population (α) can be estimated by fitting a line to the total expression data at that locus for samples of varying purity (t). The plot shows the relationship of total expression at a locus (total TPM) as a function of purity (t) for various values ​​of α (from 0 to 1 clockwise from the upper right quadrant), where TPMN+TPMT=5. [Figure 8]Shown are the rates of true negatives and false positives (first bar in each subplot, variants known not to be expressed), and false negatives and true positives (second bar in each subplot, variants known to be expressed). These were obtained using the methods described herein to identify whether a variant is expressed in a dataset of known dilutions of cell lines containing a known variant with a known expression state. Each column represents a different dilution, and each row represents the results using a different source for the genotype and purity used in the model. The number of mutations in each category is superimposed on the bars. In each case, the smallest portion of the bar is the misassigned mutation (FP or FN mutation). [Figure 9A] Boxplots of mutation power estimated as described herein for the data in Figure 8 are shown, separated by known expression status (top = mutations known not to be expressed, bottom = mutations known to be expressed). Each column represents a different dilution, and each subplot shows the distribution of mutations not detected as expressed (left, i.e., TN in top row, FN in bottom row) and mutations detected as not expressed (right). Ground truth genotype and purity. [Figure 9B] Boxplots of mutation power estimated as described herein for the data in Figure 8, broken down by known expression state (top = mutations known not to be expressed, bottom = mutations known to be expressed). Each column represents a different dilution, and each subplot shows the distribution of mutations not detected as expressed (left, i.e., TN in top row, FN in bottom row) and mutations detected as not expressed (right). Fixed genotype, ground truth purity. [Figure 9C] Boxplots of mutation power estimated as described herein for the data in Figure 8, broken down by known expression state (top = mutations known not to be expressed, bottom = mutations known to be expressed). Each column represents a different dilution, and each subplot shows the distribution of mutations not detected as expressed (left, i.e., TN in top row, FN in bottom row) and mutations detected as not expressed (right). Estimated genotype, ground truth purity. [Figure 9D] Boxplots of mutation power estimated as described herein for the data in Figure 8, broken down by known expression state (top = mutations known not to be expressed, bottom = mutations known to be expressed). Each column represents a different dilution, and each subplot shows the distribution of mutations not detected as expressed (left, i.e., TN in top row, FN in bottom row) and mutations detected as not expressed (right). Ground truth genotype, estimated purity. [Figure 9E] Boxplots of mutation power estimated as described herein for the data in Figure 8, broken down by known expression state (top = mutations known not to be expressed, bottom = mutations known to be expressed). Each column represents a different dilution, and each subplot shows the distribution of mutations not detected as expressed (left, i.e., TN in top row, FN in bottom row) and mutations detected as not expressed (right). Fixed genotype, estimated purity. [Figure 10] The true positive rate is shown as a function of power for expression in the data in Figures 8 and 9, with the variants sorted by power. Each plot is for a model using a different source of genotype and purity. [Figure 11] The rates of true negatives and false positives (first bar in each subplot, variants known not to be expressed) and false negatives and true positives (second bar in each subplot, variants known to be expressed) are shown. These were obtained using the method described herein, which uses estimates of cellular expression ratio values ​​inferred from the data to identify whether a variant is expressed in a dataset of known dilutions of cell lines containing known variants with known expression states. Each column represents a different dilution, and each row represents the results using a different source for the genotype and purity used in the model. The number of mutations in each category is superimposed on the bars. In each case, the smallest portion of the bar is an incorrectly assigned mutation (FP or FN mutation). The data is the same as that used in Figure 8B. [Figure 12A]Figure 13 shows the distribution of estimated power for expression variants obtained using the methods described herein for the data in Figure 11. Each subplot shows the results using a different source for the genotype and purity used in the model. Ground truth genotype and purity. [Figure 12B] Figure 13 shows the distribution of estimated power for expression variants obtained using the methods described herein for the data in Figure 11. Each subplot shows the results using a different source for the genotype and purity used in the model: Fixed genotype, ground truth purity. [Figure 12C] Figure 12 shows the distribution of estimated power for expression variants obtained using the methods described herein for the data in Figure 11. Each subplot shows the results using a different source for the genotype and purity used in the model: Estimated genotype, ground truth purity. [Figure 12D] Figure 13 shows the distribution of estimated power for expression variants obtained using the methods described herein for the data in Figure 11. Each subplot shows the results using a different source for the genotype and purity used in the model. Ground truth genotype, estimated purity. [Figure 12E] Figure 13 shows the distribution of estimated power for expression variants obtained using the methods described herein for the data in Figure 11. Each subplot shows the results using a different source for the genotype and purity used in the model. Fixed genotype, estimated purity. [Figure 13] The calculated (ground truth) true positive rate is shown as a function of estimated power for the data in Figures 11 and 12. Each row shows the results using a different source for the genotypes and purity used in the model. [Figure 14]For the data in Figures 11-13, the number of variants associated with a value between 0 and 1 for the probability of the mutation being expressed as determined using the methods described herein is shown, color-coded by read count (where read count = 0 or not zero), and split into never-expressed mutations (top row) and truly-expressed mutations (bottom row). Each column shows results for a different purity (5%, 10%, or 30%). [Figure 15] For the data in Figures 11-14, the number of variants associated with a value between 0 and 1 for the probability that the variant is expressed, as determined using the variant allele fraction (VAF) of the variant in the read data, is shown. Only data points with non-zero read counts are shown, as VAF cannot be calculated for variants with a read count of 0. The data are split into never-expressed variants (top row) and truly expressed variants (bottom row). Each column shows results for a different purity (5%, 10%, or 30%). [Figure 16] (A) Receiver operating characteristic (ROC) curves for the data in Figures 11-15 are shown for identifying whether a variant is expressed or not, using the probability that the variant is expressed as determined using the methods described herein. Data from a model using ground truth genotypes and purities, and alpha values ​​estimated from data of various purities using a regression model, are displayed. (B) Receiver operating characteristic (ROC) curves for the data in Figures 11-15 are shown for identifying whether a variant is expressed or not, using the probability that the variant is expressed as determined using the variant allele fraction of the variant in the read data. The ROC curves show the true positive rate as a function of the false positive rate for various values ​​of the probability cutoff used to assess whether a variant is expressed or not (shown as the color of the curve, along with a scale on the right hand side of the plot). This data only considers variants with a total number of reads >0. The plot also shows the area under the curve (AUC), the probability threshold (cutoff) that yields the highest true positive rate while maintaining a false positive rate <0.05, and the corresponding false positive rate (FPR) and true positive rate (TPR). [Figure 17] (A) shows a calibration curve for a model using the data from Figures 11-16 to determine the probability that a mutation will be expressed as determined using the methods described herein. Data for the model using ground truth genotypes and purities, and alpha values ​​estimated from data of various purities using a regression model, are displayed. (B) shows a calibration curve for a model using the data from Figures 11-16 to determine the probability that a mutation will be expressed as determined using the variant allele fraction of the mutation in the read data. The calibration curve shows the proportion of true positives (expressed variants) as a function of the probability bin (10% probability bin) obtained for each approach in A and B. This data only considers variants with a total number of reads >0. [Figure 18] (A) Receiver operating characteristic (ROC) curves for the data in Figures 11-17 are shown for identifying whether a variant is expressed or not, using the probability that the variant is expressed as determined using the methods described herein. The data are displayed for models using ground truth genotypes and purities, and alpha values ​​estimated from data of various purities using a regression model. (B) Receiver operating characteristic (ROC) curves for the data in Figures 11-17 are shown for identifying whether a variant is expressed or not, using the probability that the variant is expressed as determined using the variant allele fraction of the variant in the read data. The ROC curves show the true positive rate as a function of the false positive rate for various values ​​of the probability cutoff used to assess whether a variant is expressed or not (shown as the color of the curve, along with a scale on the right-hand side of the plot). This figure shows the same information as Figure 16, but is determined using data including cases where the total number of reads = 0. The plot also shows the area under the curve (AUC) and the probability threshold (cutoff) that yields the highest true positive rate while maintaining a false positive rate <0.05, along with the corresponding false positive rate (FPR) and true positive rate (TPR). [Figure 19](A) shows a calibration curve for a model using the data from Figures 11-18 to determine the probability that a mutation will be expressed as determined using the methods described herein. Data from the model using ground truth genotypes and purities, and alpha values ​​estimated from data of various purities using a regression model, are displayed. (B) shows a calibration curve for a model using the data from Figures 11-18 to determine the probability that a mutation will be expressed as determined using the variant allele fraction of the mutation in the read data. The calibration curve shows the proportion of true positives (expressed variants) as a function of the probability bin (10% probability bin) obtained for each approach in A and B. The data only considers variants with total reads > 0 and even variants with total reads = 0. DETAILED DESCRIPTION OF THE INVENTION

[0045] It is generally accepted that cancer antigens arising from cancer-specific variants (also called "tumor-specific mutations") represent promising therapeutic targets, provided they are expressed by cancer cells (see, e.g., Heemskerk, Kvitsborg, & Schumacher, 2013). However, approaches to determining whether these antigens are expressed have been relatively crude. For example, it has been proposed to use whole-genome or whole-exome sequencing to identify mutations in genomic DNA, and then use RNA sequencing to identify genes expressed in the tumor material and focus all efforts on them (see, e.g., Heemskerk, Kvitsborg, & Schumacher, 2013). Similarly, it has been proposed that RNA sequencing data may be used to identify tumor-specific mutations, with the advantage that they are not limited to known genes but can identify, for example, intragenic fusions, novel transcripts, etc. However, this is complicated by factors such as differences in mRNA expression levels of any given transcript between tumor and normal tissues (making comparisons difficult), making it unlikely that such approaches are suitable for identifying variants in RNA species (see, e.g., Heemskerk, Kvitsborg, & Schumacher, 2013). Gene-centric approaches to RNA expression combine information from gene variants and normal transcripts (both of which may be present in tumor cells). Furthermore, because tumor samples are typically mixed samples containing tumor and normal cells, they also provide collective signals from healthy and tumor cells, collectively providing highly uncertain information about whether a variant is actually expressed. Therefore, more sensitive approaches that can take all these factors into account are needed to more confidently identify the presence of variants in RNA expression data. In terms of determining allelic imbalance, approaches to determining allele-specific expression have been developed. For example, Castel et al. (2015) proposed an approach that combines quality control measures with a binomial test to determine whether the ratio of two germline alleles is significantly different from the expected 0.5.However, such approaches only determine whether two germline alleles expected to exist in a pure diploid heterozygous population are differentially expressed. They do not apply to the much more complex problem of detecting allele-specific expression of tumor-specific mutations. The present disclosure provides a method to solve this problem.

[0046] In this disclosure, the following terms are employed and are intended to be defined as indicated below.

[0047] The term "allele-specific expression" (also referred to as "allele expression") refers to the amount of mRNA transcribed from a particular allele or whether a particular allele is expressed (i.e., whether any mRNA is transcribed from a particular allele). In the context of this disclosure, determining allele-specific expression refers to determining whether an allele is expressed, unless otherwise indicated. Allele-specific expression is typically a property of a particular cell, cell population, tissue, sample, or individual. In the context of this disclosure, an allele of interest is an allele that is present in a tumor and absent in a germline population. Thus, an allele of interest (i.e., one whose expression is detected and / or quantified) may be referred to interchangeably as a "mutation" or "mutant allele" or "variant allele." A mutant allele may be a tumor-specific mutation, and a reference allele may be the corresponding germline allele. A germline allele may also be referred to as a "normal," "healthy," or "reference" allele. A pair of alleles within a population may be referred to as a "major allele" and a "minor allele." Here, the major allele is the allele with the highest VAF for that locus, and the minor allele is the allele with the lowest VAF for that locus (in a population with only two alleles, the VAF マイナー +VAF メジャー = 1). Depending on the sample (e.g., the proportion of tumor cells in the sample and the cancer cell fraction of the mutation), the germline allele or the mutant allele may be the major allele.

[0048] Determining the allele-specific expression of a variant allele in a mixed sample containing variant and germline cells is a significantly more complex problem than determining allele imbalance. Allelic imbalance, sometimes referred to as "allelic expression" in the prior art, quantifies the expression variation between two haplotypes in a diploid individual with heterozygous sites. This is conceptually distinct from the problem at hand (determining whether a variant is expressed) because a statistical test (usually a binomial test) can be used to assess deviations from the expected 0.5 allele expression balance. Indeed, the methods described herein aim to determine whether a variant allele is present in the expression data, rather than whether the variant and reference alleles are expressed to different degrees. Figure 1A illustrates the problem of determining allele-specific expression in a homogeneous diploid cell population. Furthermore, allelic imbalance refers to the germline context, where a single, pure population of heterozygous cells is considered. Figure 1B illustrates the problem of determining allele-specific expression in a mixed sample containing variant and reference cells. This is a significantly more complex situation than when considering a homogenous diploid cell population (such as in determining allele expression / allelic imbalance in the germline context). As shown in Figure 1A, at the simplest level, assuming a pure diploid heterozygous population (i.e., a population in which all cells have the same diploid heterozygous genotype), the genomic VAF is expected to be equal to 0.5, and genomic DNA reads should show equal representation of each allele. However, even within the same population, two alleles may not have the same expression level, so the number of RNA reads representing each allele can be anywhere between 0 and 1. In other words, any transcriptome VAF between 0 and 1 may be reasonable. Every sequencing process samples a population of molecules present in a sample. The sampling process results in a population of reads that are more or less likely to accurately represent the original population, depending on the sampling depth, the amount of molecules at a particular locus actually present in the sample (i.e., the expression level of the locus), and the sequencing error rate.Gene expression is known to have an extremely wide dynamic range, with some genes being abundantly expressed and others being expressed only at very low levels. At high expression levels of a genomic locus, the number of reads for each allele should be sufficient to infer VAF expression with relatively good confidence, especially if neither allele has extremely low expression. However, at low expression levels of a locus, the minor allele may not be sampled at all (mere chance) and / or the number of reads that actually sample the minor allele may be so low that it is impossible to distinguish between the presence of a variant and a sequencing error, given the error rate of the sequencing platform used. As shown in Figure 1B, this becomes even more complicated when examining mixed populations (e.g., samples containing tumor and non-tumor cells). Indeed, in such cases, the number of genomic copies that can express a variant allele (tumor-specific allele) in a tumor sample depends, at least, on the tumor purity (the proportion of cells in the sample that are tumor cells) and the copy number of the variant allele in the tumor population. Furthermore, in addition to the different expression levels of the variant and reference alleles in each cell, the expression level of a locus may differ between tumor and non-tumor cells.

[0049] The present inventors have developed a method for determining the power of detecting the expression of a mutant allele in a mixed sample containing a variant cell population (e.g., tumor cells) with at least one copy of the allele and a reference cell population (e.g., normal cells) without any copies of the allele. The method evaluates the likelihood of observing a given number b of reads containing the variant allele based on a null model in which the variant allele is not expressed and an alternative model in which the variant allele is expressed. Statistical tests for determining whether a variant is expressed have four outcomes: false positive (FP, where a variant is identified as expressed when in fact it is not), false negative (FN, where a variant is identified as not expressed when in fact it is), true positive (TP, where a variant that is actually expressed is identified as expressed), and true negative (TN, where a variant that is not actually expressed is identified as not expressed). There is a trade-off between false positives and false negatives; a more permissive test will reflect more true positives, but also more false positives. Therefore, the FP rate can be fixed at an acceptable level (e.g., 5%), and the power (probability of avoiding false negatives) can be calculated based on the FN rate at that selected FP rate. As shown in Figure 1D, the two likelihood values ​​(based on the null model M0 at the top and the alternative model M1 at the bottom) assume a binomial distribution for the number of variant reads observed under each hypothesis and behave differently as a function of the number of variant reads b. The critical number of variant reads b c It is possible to identify a critical number above which the null model associated with any chosen false positive (FP) rate (FPR, the probability of a false positive, i.e., the probability that a variant is declared expressed when it is not actually expressed) is rejected (when the number of variant reads is b ≥ b_ c For example, a suitable FPR selection may be FPR=0.05. Therefore, b c is based on M0, P(b≧b c )=0.05. At this threshold, b ≥ b cAny variant where b is determined to be expressed is determined to be expressed, and the probability of detecting a truly expressed variant (P(b ≥ b c )) is the true positive (TP) rate (TPR), also called the power of the test. As shown in Figure 1C and further described below, the likelihoods M1 and M0 are estimated based on the probability of observing sequence data, which includes the total number of reads covering the locus (d) and the number of reads covering the locus and carrying the mutation (b), as well as the tumor purity (t) and the genotypes of the variant and reference cells (G, G). v and G N The model takes into account the proportion of expression at a locus due to the reference allele (θ, reference read count / total read count), and the proportion of total expression (both variant and reference) due to the reference cells (α). The model detailed below assumes that the variant and reference cell populations are genetically homogeneous at the locus (i.e., one genotype G v and one genotype G N However, the same principles can be applied to cases where the variant population contains populations that carry the variant and populations that do not carry the variant (but may not have the same genotype at the same loci as the reference population). Thus, the default implementation assumes that the mutation is ubiquitous in the sample(s) under analysis. The variable ξ is the product of the genotype G v Variant cells with genotype G Nis the probability of sampling a read containing a mutation from a mixed sample containing a reference cell with ξ. Here, variant cells represent the population proportion t. It is proportional to the copy number of the variant allele in the mixed population and its relative expression within the population. In the example shown below, the probability of observing the number of variant reads b is assumed to follow a binomial distribution conditional on the total number of local reads and the value of the variable ξ. The likelihoods of the null and alternative models are obtained using Bayes' rule and various priors for the null and variant model parameters θ (P(θ)). For example, in the null model, the prior for θ may place a higher density around θ = 1 (where most expression is expected to come from the reference allele) than around any other value (shown in the upper right of Figure 1E). On the other hand, in the alternative model, the prior for θ gives a wide range of possible values ​​for θ between 0 and 1 (shown as a flat prior in the upper right of Figure 22). As shown in Figure 1E, the likelihoods of M0 and M1 (and in particular their ratio r) can be used to estimate the posterior probability that a variant will be expressed given the expression data (P(E=1|b,d,α,t)). The probabilities p(θ|E=1) and p(θ|E=0), shown in the upper right corner of Figure 1E, are priors for θ (the reference ratio, which is the balance between reference allele expression and total expression at the locus), conditional on E (whether the variant is expressed or not). In the example below, a prior distribution with a high probability mass placed at 1 is used for E=0 (no expression; all expression is expected to come from the reference allele), and a relatively flat prior (similar mass for all values ​​of θ) is used for E=1 (the variant allele may be expressed at any level). If we assume that E is a Bernoulli variable with parameter ρ, where ρ represents prior knowledge of expression, then the probability p(ρ) becomes a hyperprior with respect to ρ. Using estimates of the data b, d, and t (and optionally α), it is possible to calculate the joint posterior distribution for E and ρ, from which we obtain the likelihood ratio r between the two models, and the mean (hyper-prior) μ of the beta distribution for ρ. ρBy integrating with respect to ρ, we can calculate the marginal distribution over E, as shown in Figure 1E.

[0050] The output of this method may be used in multiple ways. For example, if the power of detection of expression of a particular variant is determined to be low and / or the probability of expression of the variant is determined to be low, the variant will not necessarily be removed from the list of potentially immunogenic variants, despite low evidence of expression of the variant. This information may be combined with one or more additional criteria, such as the likelihood of a variant-derived peptide binding to one or more MHC alleles, the likelihood of presentation of a variant-derived peptide by one or more MHC alleles, the likelihood of processing a variant-derived peptide, the likelihood of a variant-derived peptide being immunogenic, or the likelihood of differential binding affinity and / or presentation between a variant-derived peptide and the corresponding germline peptide. Conversely, a variant identified as being associated with high power of detection of expression and / or a low probability of expression may be removed from the list of potentially immunogenic variants if evidence of variant expression is low. Furthermore, the methods described herein can be applied to multiple samples, individually or jointly, to determine whether a variant is likely to be present / expressed in all of multiple tumor samples from the same subject. This provides an indication of whether the variant is likely to be present (and expressed) in all cells of the cancer—sometimes called a ubiquitously expressed variant—which can be assumed to be clonality (although strictly speaking, true clonality cannot be reliably established, as single-cell sequence data would be required for all, or at least a representative proportion, of the cancer cells).

[0051] As used herein, a "sample" may be a cell or tissue sample, biological fluid, or extract (e.g., a DNA extract obtained from a subject) from which genomic and / or transcriptomic material can be obtained for genomic and / or transcriptomic analysis, examples of which include genome sequencing (e.g., whole genome sequencing, whole exome sequencing) or RNA sequencing (also referred to as "RNAseq" or "RNA-seq"). A sample may also be a sample of cells, tissue, or biological fluid (e.g., a biopsy) taken from a subject. Such a sample may be referred to as a "subject sample." In particular, a sample may be a blood sample or a tumor sample, or a sample derived therefrom. For purposes of obtaining RNA sequence data, a "sample" as used herein is a cell or tissue sample, or an extract (e.g., an RNA extract obtained from a subject) from which transcriptomic material can be obtained. As used herein, for purposes of obtaining DNA / genomic sequence data, a "sample" may be a cell or tissue sample, biological fluid, or extract (e.g., a DNA extract obtained from a subject) from which genomic material can be obtained for genomic analysis, examples of which include genome sequencing (e.g., whole genome sequencing, whole exome sequencing). A sample may be freshly taken from a subject or may be processed and / or stored (e.g., frozen, restored, or subjected to one or more purification, enrichment, or extraction steps) prior to genome / transcriptome analysis. A sample may also be a cell or tissue culture sample. Thus, a sample as described herein may refer to any type of sample containing cells, or genomic and / or transcriptomic material derived therefrom, whether derived from a biological sample obtained from a subject or from a sample obtained, for example, from a cell line. In embodiments, the sample is a sample obtained from a subject, such as a human subject.The sample is preferably from a mammal (e.g., a mammalian cell sample, or a sample from a mammalian subject, examples of which include a cat, dog, horse, donkey, sheep, pig, goat, cow, mouse, rat, rabbit, or guinea pig), and is preferably from a human (e.g., a human cell sample, or a sample from a human subject). Furthermore, the sample may be transported and / or stored, collection may occur at a location remote from the location of sequence data acquisition (e.g., sequencing), and / or the steps of any computer-implemented method described herein may occur at a location remote from the location of sample collection and / or the location of sequence data acquisition (e.g., sequencing) (e.g., the steps of the computer-implemented method may be performed by means of networked computers, an example of which is by means of a "cloud" provider).

[0052] A "mixed sample" refers to a sample that is expected to contain multiple cell types or genetic material derived from multiple cell types. In the context of the present disclosure, a mixed sample typically contains tumor cells or is expected (or anticipated) to contain tumor cells or genetic material derived from tumor cells. The genetic material can include genomic material (e.g., DNA) or transcriptomic material (e.g., RNA). A sample collected from a subject (such as a tumor sample) is typically a mixed sample (unless subjected to one or more purification and / or separation steps). Typically, the sample contains tumor cells and at least one other cell type (and / or genetic material derived therefrom). For example, a mixed sample may be a tumor sample. A "tumor sample" refers to a sample derived from or obtained from a tumor. Such a sample may include tumor cells and normal (non-tumor) cells. Normal cells may include immune cells (e.g., lymphocytes) and / or other normal (non-tumor) cells. Lymphocytes in such a mixed sample may be referred to as "tumor-infiltrating lymphocytes" (TILs). The tumor may be a solid tumor, a non-solid tumor, or a hematological tumor. The tumor sample may be a primary tumor sample, a tumor-associated lymph node sample, or a sample from a metastatic site in a subject. The sample containing tumor cells or genetic material derived from tumor cells may be a bodily fluid sample. Thus, the genetic material derived from tumor cells may be circulating tumor DNA or tumor DNA in exosomes. Alternatively, or in addition, the sample may contain circulating tumor cells. The mixed sample may be a cell, tissue, or bodily fluid sample that has been processed to extract genetic material. Methods for extracting genetic material from biological samples are known in the art. The mixed sample may have undergone one or more processing steps that may alter the ratio of multiple cell types or genetic material derived from multiple cell types within the sample. For example, a mixed sample containing tumor cells may have been processed to enrich the sample in tumor cells.Thus, even if the sample is assumed to be pure for a particular purpose (i.e., the tumor fraction is 1 or 100%), a sample of purified tumor cells may be referred to as a "mixed sample" based on the possible presence of small amounts of other cell types.

[0053] The term "tumor fraction" (sometimes referred to as "tumor purity" or simply "purity," or aberrant cell fraction (ACF)) refers to the proportion of DNA-containing cells (which are tumor cells) within a mixed sample, or the equivalent proportion at which a particular mixture of genetic material from tumor and non-tumor cells within a sample is assumed to occur. Methods for determining the tumor fraction within a sample are known in the art. For example, in terms of a cell or tissue sample, the tumor fraction may be estimated by analyzing pathology slides (e.g., hematoxylin and eosin (H&E)-stained slides, or other histochemistry or immunohistochemistry slides, by counting tumor cells in one or more representative areas of the sample) or by using high-throughput assays such as flow cytometry. In terms of samples containing genomic material, tumor fraction may be estimated using sequence analysis processes that attempt to deconvolute the tumor and germline genomes, examples of which include, for example, ASCAT (Van Loo et al., 2010), ABSOLUTE (Carter et al., 2012), or ichorCNA (Adalsteinsson et al., 2017).

[0054] A "normal sample," "healthy sample," or "germline cell sample" refers to a sample that is assumed to be free of tumor cells or genetic material derived from tumor cells. A germline cell sample may be a blood sample, tissue sample, or purified sample, an example of which is a peripheral blood mononuclear cell sample from a subject. Similarly, the terms "normal," "germline," or "wild-type," when referring to a sequence or genotype, refer to the sequence / genotype of cells other than tumor cells. A germline cell sample may contain a small proportion of tumor cells or genetic material derived therefrom, but may still be assumed to be free of such cells or genetic material for practical purposes. In other words, all cells or genetic material may be assumed to be normal, and / or sequence data that does not fit that assumption may be disregarded.

[0055] The term "sequence data" refers to information indicating the presence, and preferably the amount, of genetic material in a sample having a particular sequence. Such information may be obtained using sequencing technology (e.g., next-generation sequencing (NGS), such as whole exome sequencing (WES), whole genome sequencing (WGS), RNA sequencing, or sequencing of reflected genomic loci (targeted sequencing or panel sequencing)), or using array technology (e.g., copy number variation array, expression array, or other molecular counting assay). When using NGS technology, the sequence data may include a count of the number of sequencing reads having a particular sequence. When using non-digital technology, such as array technology, the sequence data may include a signal (e.g., intensity value) indicating the number of sequences in a sample having a particular sequence, for example, by comparison with an appropriate control. The sequence data may be mapped to a reference sequence, such as a reference genome or transcriptome, using methods known in the art (e.g., Bowtie (Langmead et al., 2009)). Thus, a sequencing read count or equivalent non-digital signal may be associated with a particular genomic location ("genomic location" refers to a location in a reference genome to which sequence data is mapped). Furthermore, a genomic location may contain mutations, in which case a sequencing read count or equivalent non-digital signal may be associated with each possible variant (also called an "allele") at a particular genomic location. The process of identifying the presence of a mutation at a particular location in a sample is called "variant calling," and can be performed using methods known in the art (e.g., GATK HaplotypeCaller, gatk.broadinstitute.org / hc / en-us / articles / 360037225632-HaplotypeCaller, etc.).For example, the sequence data may include a count of the number of reads (or equivalent non-digital signals) that match a germline (sometimes called a "reference") allele at a particular genomic location, and a count of the number of reads (or equivalent non-digital signals) that match a mutated (sometimes called an "alternate") allele at a genomic location.

[0056] Furthermore, genome sequence data may be used to infer copy number profiles along the genome using methods known in the art. A copy number profile refers to the number of genome copies of a genomic region. The copy number profile may be allele-specific. In the context of the present disclosure, the copy number profile is preferably allele-specific and tumor / normal sample-specific. In other words, the copy number profile used in the present disclosure is preferably obtained by analyzing a sample containing a mixture of tumor cells and normal cells and using a method designed to generate allele-specific copy number profiles of the tumor cells and normal cells within the sample. The allele-specific copy number profile of the mixed sample may be obtained from the sequence data (e.g., using the read counts described above), for example, using ASCAT (Van Loo et al., 2010). Other methods are known and equally suitable. Preferably, within the context of the present disclosure, the method used to obtain the allele-specific copy number profile reports multiple possible copy number solutions and associated quality / reliability metrics. For example, ASCAT outputs a goodness-of-fit metric for each combination of ploidy (not segment-specific, but overall tumor sample ploidy) and purity values ​​at which the corresponding allele-specific copy number profile was evaluated. It should be noted that tumor-specific copy number profiles generated by such methods represent an average or summary of the entire tumor cell population.

[0057] The term "total copy number" refers to the total number of copies of a genomic region in a sample. The term "major copy number" refers to the number of genomic copies of the most prevalent allele in a sample. Conversely, the term "minor copy number" refers to the number of genomic copies of an allele other than the most prevalent allele in a sample. Unless otherwise indicated, these terms refer to the estimated major copy number and major copy number (and total copy number) of an estimated tumor copy number profile. The term "normal copy number" or "normal total copy number" refers to the number of copies of a genomic region in normal cells in a sample. Normal cells typically have two copies of each chromosome (unless the cell is genetically male and the chromosome is a sex chromosome). Therefore, in embodiments, the normal copy number may be assumed to be equal to 2 (unless the genomic region is on the X or Y chromosome and the sample under analysis is from a male subject, in which case the normal copy number may be assumed to be equal to 1). Alternatively, the normal copy number of a particular genomic region may be determined using a normal sample.

[0058] The term "logR value" (sometimes referred to as "logR," "logRR," or "LLR") refers to a measure of normalized total signal intensity that quantifies the total genome copy number at a genomic locus. In the context of the present disclosure, this term typically refers to the logR value of a sample containing tumor genetic material, and normalization is typically performed by reference to a normal sample (preferably a matched normal sample, but which may also be a process-matched normal sample or other suitable normal reference sample). For example, when NGS is used, logR may be obtained as a normalized logarithmic transformation of read depth (log(read depth tumor / read depth normal)). The term "mean B allele frequency" (MBAF, sometimes referred to as "B allele frequency" (BAF)) is a measure of the normalized allele intensity ratio at a genomic location. In the context of the present disclosure, this term typically refers to the BAF value of a sample containing tumor genetic material, and normalization is typically performed by reference to a normal sample (preferably a matched normal sample, but which may also be a process-matched normal sample or other suitable normal reference sample). For example, BAF may be obtained as the ratio of the allele frequencies of tumor alleles to normal alleles. Copy number profiles typically include copy number estimates across genomic regions called "segments." Thus, the BAF and logR associated with a genomic location may refer to the BAF and logR of the segment that overlaps with a specific genomic location (e.g., the genomic location of a mutation). Furthermore, BAF and logR can be used to obtain the corresponding major and minor copy numbers. In embodiments, copy number index values ​​may be provided for both tumor and normal copy number profile estimates, even if only tumor copy number profile values ​​are used.

[0059] The terms "tumor-specific mutation," "tumor-specific variant," "somatic mutation," or simply "mutation" or "variant" are used interchangeably and refer to differences in the nucleotide sequence (e.g., DNA or RNA) of tumor cells compared to healthy cells from the same subject. Differences in nucleotide sequence can result in the expression of a protein that is not expressed in healthy cells from the same subject. For example, a mutation can be a single nucleotide variant (SNV), a multiple nucleotide variant (MNV), a deletion mutation, an insertion mutation, a translocation, a missense mutation, a translocation, a fusion, a splice site mutation, or any other change in the genetic material of tumor cells. A mutation can result in the expression of a protein or peptide that is not present in healthy cells from the same subject. Mutations can be identified by exome sequencing, RNA sequencing, whole genome sequencing, and / or targeted gene panel sequencing, and / or conventional Sanger sequencing of a single gene, followed by sequence alignment to compare the DNA and / or RNA sequence from a tumor sample with the DNA and / or RNA of a reference sample or sequence (e.g., germline DNA and / or RNA sequence, or a reference sequence from a database). Suitable methods are known in the art.

[0060] An "indel mutation" refers to the insertion and / or deletion of bases in the nucleotide sequence (e.g., DNA or RNA) of an organism. Typically, indel mutations occur in the DNA of an organism, preferably genomic DNA. In embodiments, the indel may be 1 to 150 bases, e.g., 1 to 90, 1 to 50, 1 to 23, or 1 to 10 bases. The indel mutation may be a frameshift indel mutation. A frameshift indel mutation is a change in the read frame of a nucleotide sequence caused by the insertion or deletion of one or more nucleotides. Such frameshift indel mutations may generate a novel open reading frame that is significantly different from the polypeptide encoded by the non-mutated DNA / RNA in the corresponding healthy cells of a subject.

[0061] "Neoantigens" (or "neo-antigens") are antigens that arise as a result of mutations in cancer cells. Thus, neoantigens are not expressed (or are expressed at significantly lower levels) in normal (i.e., non-tumor) cells. Neoantigens may be engineered to generate individual peptides that can be recognized by T cells when presented in the context of MHC molecules. As described herein, neoantigens may be used as the basis for cancer immunotherapy. References herein to "neoantigens" are intended to include peptides derived from neoantigens. As used herein, the term "neoantigen" is intended to encompass any portion of a neoantigen that is immunogenic. An "antigen" molecule, as referred to herein, is a molecule that, itself or a portion thereof, can stimulate an immune response when presented to the immune system or immune cells in an appropriate manner. Binding of neoantigens to specific MHC molecules (encoded by specific HLA alleles) may be predicted using methods known in the art. Examples of methods for predicting MHC binding include those described by Lundegaard et al., O'Donnel et al., and Bullik-Sullivan et al. For example, MHC binding of neoantigens may be predicted using the netMHC-3 (Lundegaard et al.) and netMHCpan4 (Jurtz et al.) algorithms. Neoantigens predicted to bind to a particular MHC molecule are thereby predicted to be presented by that MHC molecule on the cell surface.

[0062] A "clonal neo-antigen" (sometimes referred to as a "truncal neo-antigen") is a neo-antigen that results from a mutation that is present in essentially all tumor cells in one or more samples from a subject (or that can be assumed to be present in essentially all tumor cells from which the tumor genetic material in the sample(s) was derived). Similarly, a "clonal mutation" (sometimes referred to as a "truncal mutation") is a mutation that is present in essentially all tumor cells in one or more samples from a subject (or that can be assumed to be present in essentially all tumor cells from which the tumor genetic material in the sample(s) was derived). Thus, a clonal mutation may be a mutation that is present in all tumor cells in one or more samples from a subject. A "subclonal" neo-antigen is a neo-antigen that results from a mutation that is present in a subset or portion of cells in one or more tumor samples from a subject (or that can be assumed to be present in a subset of tumor cells from which the tumor genetic material in the sample(s) was derived). Similarly, a "subclonal" mutation is a mutation that is present in a subset or portion of cells in one or more tumor samples from a subject (or that can be assumed to be present in a subset of tumor cells from which the tumor genetic material in the sample(s) was derived). A neo-antigen or mutation may be clonal in terms of one or more samples from a subject, but not truly clonal in terms of the entire tumor cell population that may be present in the subject (e.g., including all areas of the primary tumor and metastases). Thus, a clonal mutation may be "truly clonal," in the sense that it is a mutation that is present in essentially all tumor cells (i.e., all tumor cells) in a subject. This is because one or more samples may not be representative of each and every subset of cells present in a subject. Thus, in terms of the present disclosure, a "clonal neo-antigen" or "clonal mutation" may also be referred to as a "ubiquitous neo-antigen" or "ubiquitous mutation," indicating that the neo-antigen is present in essentially all tumor cells analyzed, but not necessarily in all tumor cells that may be present in a subject.The terms "clonal" and "ubiquitous" are used interchangeably unless the context indicates that a reference to "true clonality" is intended. The phrase "essentially all tumor cells" in relation to one or more samples or subjects may refer to at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the tumor cells in one or more samples or subjects.

[0063] Nevertheless, neoantigens / mutations identified as likely to be clonal (or "ubiquitous") may be considered to be more likely to be truly clonal, or at least more likely to be truly clonal than neoantigens / mutations identified as unlikely to be clonal. Furthermore, the confidence in the probability that a clonal neoantigen / mutation identified in a subject is truly clonal increases if the sample(s) used to identify the clonal neoantigen / mutation reflect a more complete picture of the tumor's genetic diversity (e.g., by including multiple samples from the subject (e.g., samples from different regions of the tumor) and / or by including samples that inherently reflect the diversity of tumor cells (e.g., ctDNA samples)). Conversely, neoantigens / mutations identified as unlikely to be clonal are less likely to be truly clonal, because identifying a neoantigen / mutation as unlikely to be clonal indicates that there is evidence that the neoantigen / mutation is not present in all tumor cells, even within the limited view afforded by the sampling process. Therefore, the process of identifying clonal neoantigens / mutations may be viewed as prioritizing which candidate neoantigens / mutations are most likely to be clonal based on the limited view of the clonal structure of the subject's tumor available from one or more samples.

[0064] The term "cancer cell fraction" (or "CCF") refers to the proportion of tumor cells that contain mutations, such as mutations resulting in a particular neo-antigen. Cancer cell fractions may be estimated based on one or more samples and, therefore, may not necessarily equal the true cancer cell fraction in a subject (as explained above). Nevertheless, a cancer cell fraction estimated based on one or more samples may provide a useful indication of the likely true cancer cell fraction. Furthermore, as explained above, the accuracy of such an estimate may be increased if the sample(s) used to estimate the cancer cell fraction reflect a more complete picture of the tumor's genetic diversity. Additional noise sources and confounding factors within genomic data mean that cancer cell fractions determined from one or more samples represent an estimate. Thus, while a truly clonal mutation / neo-antigen should have a CCF = 1, mutations / neo-antigens that are actually more likely to be clonal are expected to be associated with higher CCF estimates (not necessarily equal to 1), whereas mutations that are less likely to be clonal are expected to be associated with lower CCF estimates.

[0065] For example, as described by Landau et al. (2013), an estimate of the cancer cell fraction may be obtained by integrating variant allele frequencies with copy number and purity estimates. Such CCF estimates can also be used to identify mutations that are likely to be clonal. For example, a clonal mutation may be defined as a mutation having an estimated cancer cell fraction (CCF) of 0.75 or greater (e.g., CCF ≧ 0.80, 0.85, 0.90, 0.95 or 1.0). A subclonal mutation may be defined as a mutation having a CCF of less than 0.95, 0.90, 0.85, 0.80 or 0.75. Further, the CCF estimate may be associated with (e.g., derived from) a distribution that associates a probability with each of a plurality of possible values of CCF from 0 to 1, and a statistically reliable estimate may be obtained therefrom. For example, if the 95% CCF confidence interval is >= 0.75, i.e., the upper limit of the 95% confidence interval of the estimated CCF is 0.75 or greater, the mutation may be defined as likely to be a clonal mutation. In other words, if there is an interval of CCF having a lower limit L and an upper limit H such that H >= 0.75 and P(L < CCF < H) = 95%, the mutation may be defined as likely to be a clonal mutation. Alternatively, if the probability that the cancer cell fraction (CCF) reaches or exceeds the required value (e.g., 0.75 or 0.95) defined above is greater than 50% (examples thereof include probabilities of 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95% or greater), the mutation may be identified as a clone. In other words, if P(CCF > 0.75) >= 0.5, the mutation may be identified as a clone. For example, the mutation may be classified as a likely clone or subclone, respectively, based on whether the posterior probability that the CCF exceeds 0.95 (or 0.75, or any other selected threshold) is greater or less than 0.5.

[0066] In this context, as described further below, a mutation may be identified as likely to be clonal if P(CCF=1) exceeds a threshold. The threshold may be fixed. For example, a mutation may be identified as likely to be clonal if P(CCF=1)>0.05. Alternatively, the threshold may be determined for the particular set of mutations being investigated. In embodiments, the threshold may be set based on a benchmark dataset with known clonal / non-clonal status to reach a predetermined precision and / or recall. The benchmark dataset may be obtained using synthetic data and / or using a dataset obtained from a population with known clonal structure (e.g., cell line mixture data). For example, a mutation may be identified as likely to be clonal if P(CCF=1)>t, where t is the maximum value such that 95% (or any other value, e.g., 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%) of the true clonal mutations in the benchmark dataset are identified (i.e., a false negative rate of at most 5%). As another example, a mutation may be identified as likely to be clonal if P(CCF=1)>t, where t is the maximum value such that at least 50% (or any other value, e.g., 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%) of the mutations above a threshold in the benchmark dataset are true clonal. The threshold is the minimum value that results in a high probability of clonal mutation (i.e., a true positive rate of at least 50%). Alternatively, the threshold may be set so that any mutation (or a certain percentage of mutations) associated with a putative CCF with a confidence interval that meets the above criteria (e.g., the upper limit of the 95% confidence interval for the putative CCF is 0.75 or greater) is selected as likely to be clonal. Alternatively, the threshold may be set so that any mutation (or a certain percentage of mutations) associated with a putative CCF with a posterior probability distribution that meets the above criteria (e.g., the posterior probability that the CCF exceeds 0.95 (or 0.75, or any other selected threshold) is greater than 0.5) is selected as likely to be clonal.

[0067] Cancer immunotherapy (or simply "immunotherapy") refers to a therapeutic approach that involves administering to a subject an immunogenic composition (e.g., a vaccine), a composition comprising immune cells, or an immunoactive agent, such as a therapeutic antibody. The term "immunotherapy" can also refer to the therapeutic composition itself. In the context of the present disclosure, immunotherapy typically targets a neoantigen. For example, an immunogenic composition or vaccine may include a neoantigen, neoantigen-presenting cells, or substances necessary for the expression of a neoantigen. As another example, a composition comprising immune cells may include T cells and / or B cells that recognize a neoantigen. Immune cells may be isolated from a tumor or other tissue (including, but not limited to, lymph nodes, blood, or ascites), expanded ex vivo or in vitro, and readministered to a subject (a process known as "adoptive cell therapy"). Alternatively, or in addition, T cells may be isolated from a subject, engineered to target a neoantigen (e.g., by inserting a chimeric antigen receptor that binds to the neoantigen), and readministered to the subject. As another example, a therapeutic antibody may be an antibody that recognizes a neoantigen. Those skilled in the art will understand that if the neoantigen is a cell surface antigen, the antibody referred to herein recognizes the neoantigen. If the neoantigen is an intracellular antigen, the antibody recognizes a neoantigen peptide-MHC complex. As referred to herein, an antibody that "recognizes" a neoantigen encompasses both of these possibilities. Furthermore, immunotherapy may target multiple neoantigens. For example, an immunogenic composition may include multiple neoantigens, cells that present multiple neoantigens, or materials necessary for the expression of multiple neoantigens. As another example, a composition may include immune cells that recognize multiple neoantigens. Similarly, a composition may include multiple immune cells that recognize the same neoantigen. As another example, a composition may include multiple therapeutic antibodies that recognize multiple neoantigens. Similarly, a composition may include multiple therapeutic antibodies that recognize the same neoantigen.

[0068] The composition as described herein may be a pharmaceutical composition further comprising a pharmaceutically acceptable carrier, diluent, or excipient.The pharmaceutical composition may optionally comprise one or more additional pharmaceutically active polypeptides and / or compounds.Such a formulation may be, for example, in a form suitable for intravenous infusion.

[0069] Reference to "immune cells" is intended to encompass cells of the immune system (e.g., T cells, NK cells, NKT cells, B cells, and dendritic cells). In a preferred embodiment, the immune cells are T cells. The neo-antigen-recognizing immune cells may be engineered T cells. The neo-antigen-specific T cells may express a chimeric antigen receptor (CAR) or T cell receptor (TCR) that specifically binds the neo-antigen or neo-antigen peptide, or an affinity-enhanced T cell receptor (TCR) that specifically binds the neo-antigen or neo-antigen peptide (discussed further below). For example, the T cells may express a chimeric antigen receptor (CAR) or T cell receptor (TCR) that specifically binds the neo-antigen or neo-antigen peptide (e.g., an affinity-enhanced T cell receptor (TCR) that specifically binds the neo-antigen or neo-antigen peptide). Alternatively, the neo-antigen-recognizing immune cell population may be a T cell population isolated from a tumor-bearing subject. For example, a T cell population may be generated from T cells in a sample isolated from a subject (e.g., a tumor sample, a peripheral blood sample, or a sample from another tissue of the subject). A T cell population may be generated from a sample from a tumor in which a neoantigen has been identified. In other words, a T cell population may be isolated from a sample derived from a tumor of a patient to be treated, in which case a neoantigen is also identified from a sample from the tumor. A T cell population may include tumor-infiltrating lymphocytes (TILs).

[0070] The term "antibody" (Ab) includes monoclonal antibodies, polyclonal antibodies, multispecific antibodies (e.g., bispecific antibodies), and antibody fragments that exhibit the desired biological activity. The term "immunoglobulin" (Ig) may be used interchangeably with "antibody." Once a suitable neo-antigen has been identified, for example, by the methods of the present disclosure, antibodies can be generated using methods known in the art.

[0071] An "immunogenic composition" is a composition capable of inducing an immune response in a subject. This term is used interchangeably with the term "vaccine." The immunogenic compositions or vaccines described herein may induce the generation of an immune response in a subject. The "immune response" that may be generated may be humoral and / or cellular, examples of which include the stimulation of antibody production or the stimulation of cytotoxic or killer cells, which may recognize and destroy (or otherwise eliminate) cells expressing an antigen corresponding to the antigen in the vaccine on their surface. The immunogenic composition may include one or more neo-antigens or substances necessary for the expression of one or more neo-antigens. Additionally, neo-antigens may be delivered in the form of cells such as antigen-presenting cells, e.g., dendritic cells. Antigen-presenting cells, such as dendritic cells, may be pulsed or loaded with neo-antigens or neo-antigen peptides, or genetically modified (by DNA or RNA transfer) to express one, two, or more neo-antigens or neo-antigen peptides (e.g., 2, 3, 4, 5, 6, 7, 8, 9, or 10 neo-antigens or neo-antigen peptides). Methods for preparing dendritic cell immunogenic compositions or vaccines are known in the art.

[0072] Neo-antigenic peptides may be synthesized using methods known in the art. The term "peptide" is used in its standard sense to refer to a series of residues (usually L-amino acids) joined together by peptide bonds, typically between the amino and carboxyl groups of adjacent amino acids. This term includes modified peptides and synthetic peptide analogs. Neo-antigenic peptides may contain cancer cell-specific mutations (e.g., non-silent amino acid substitutions encoded by single nucleotide variants (SNVs)) at any residue position within the peptide. By way of example, peptides capable of binding to MHC class I molecules are typically 7 to 13 amino acids in length. Thus, amino acid substitutions may occur at positions 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, or 13 of a 13-amino acid peptide. In embodiments, longer peptides (e.g., 21- to 31-mers) may be used, and mutations may occur at any position, including the center of the peptide, e.g., positions 10, 11, 12, 13, 14, 15, or 16. Such peptides can also be used to stimulate both CD4 and CD8 cells to recognize neoantigens.

[0073] As used herein, "treatment" refers to reducing, alleviating, or eliminating one or more symptoms of the disease being treated compared to the symptoms before treatment. "Prevention" (or prophylaxis) refers to delaying or preventing the onset of symptoms of the disease. Prevention may be absolute (the disease does not occur) or may be effective only in some individuals or for a limited time.

[0074] As used herein, the term "computer system" includes hardware, software, and data storage devices for embodying a system or performing a method according to the above embodiments. For example, a computer system may include a central processing unit (CPU), input means, output means, and data storage, which may be embodied as one or more connected computing devices. Preferably, a computer system has a display or includes a computing device with a display for providing a visual output display (e.g., a business process design). Data storage may include RAM, a disk drive, or other non-transitory computer-readable medium. A computer system may include multiple computing devices connected by a network and capable of communicating with each other over the network. It is expressly contemplated that a computer system may consist of or include a cloud computer.

[0075] As used herein, the term "computer-readable medium" includes, but is not limited to, any non-transitory medium(s) that can be read and accessed directly by a computer or computer system. The media can include, but are not limited to, magnetic storage media (such as floppy disks, hard disk storage media, and magnetic tape), optical storage media (such as optical disks or CD-ROMs), electrical storage media (such as memory, including RAM, ROM, and flash memory), and hybrids and combinations of the above (such as magnetic / optical storage media).

[0076] Identifying the mutations that are expressed The present disclosure provides a method for determining whether tumor-specific mutations are likely to be expressed using sequence data from one or more samples containing tumor cells or genetic material derived therefrom. The present disclosure also provides a method for identifying neoantigens, which includes determining whether one or more tumor-specific mutations are likely to be expressed. An exemplary method is described with reference to FIG. 2A. The method may include an optional step 10 of obtaining one or more samples containing tumor genetic material (e.g., one or more tumor samples). The sample(s) may be a mixed sample containing genomic material from multiple cell types, including tumor cells and non-tumor cells (also referred to as "reference cells," "healthy cells," "normal cells," or "germline cells"). One or more germline samples may also be obtained that do not contain tumor genetic material. The germline samples may be matched germline samples obtained from the same subject from which the one or more tumor samples were obtained. Using matched germline samples improves the accuracy of somatic (tumor-specific) mutation calling by allowing variant positions identified in the tumor sample to be compared with any variant positions in the matched germline sample to exclude germline variants. The same matched germline sample may be used to analyze multiple tumor samples from a subject. Furthermore, the matched germline sample and one or more tumor samples may be obtained at different times. For example, a first tumor sample and a matched germline sample may be obtained at the time of tumor diagnosis or resection, and additional tumor samples may be obtained at a later time point and analyzed together with the initial matched germline sample. If a matched germline sample is not available, a reference sample or genome containing common germline variants may be used. Alternatively, a process-matched normal sample may be used, but it may not be obtained from the same subject or may be obtained from a pool of subjects.

[0077] Optionally, the sample may be sequenced in step 12 to obtain at least RNA sequence data and, optionally, DNA sequence data. The RNA sequence data may be obtained by RNA sequencing. The DNA sequence data may be obtained using either whole-exome sequencing or whole-genome sequencing. For example, although alternative methods (e.g., allele-specific copy number arrays or expression arrays) may be used, sequencing methods are preferred because they generate a digital output representative of the number of each specific sequence in the sample. Alternatively, the sequence data may have been previously obtained and received from a user interface, computing device, or database. In optional step 14, the sequence data may be analyzed to identify one or more mutations that are present in tumor cells but likely absent in non-cancerous cells. These represent tumor-specific mutations. These may be used as candidate neoantigens, as further described below. Step 14 may include, in step 14A, aligning sequences from one or more samples (i.e., the mixed sample(s) and, if available, the germline sample(s)) to a reference (e.g., a reference genome or transcriptome) and identifying genomic locations (also called "positions" or "locuses") where the tumor sequence differs from the germline sequence or can be assumed to differ from the germline sequence (e.g., if the subject's germline sequence is not available). This analysis may be performed using RNA-seq data, DNA-seq data, or both. For example, DNA-seq data may be used to identify tumor-specific mutations that are single-nucleotide variants, multi-nucleotide variants, or indels, and RNA-seq data may be used to identify tumor-specific mutations that are gene fusions or splicing variants. Thus, the tumor-specific mutations analyzed may be somatic mutations present in the tumor of the subject from whom the sample was taken.One or more of any identified (or otherwise selected by a user via, for example, a user interface or retrieved from a computing device or database) tumor-specific mutations may then be analyzed to determine whether they are likely to be expressed. In step 16, sequence data may be obtained for each sample containing tumor genetic material, including the number of RNA reads in the sample that exhibit the tumor-specific mutation (b) (these may also be referred to as reads that "contain" the tumor-specific mutation or reads that "support" the tumor-specific mutation) and the total number of reads at the position of the tumor-specific mutation (d). Alternatively, the sequence data may include any two of the number of RNA reads that exhibit the tumor-specific mutation, the number of RNA reads that exhibit the corresponding germline allele, and the total number of RNA reads at the position of the tumor-specific mutation (as all three numbers can be obtained from any two of these).

[0078] In optional step 18, information regarding at least one copy number solution that fits each sample containing tumor genetic material and information regarding the tumor fraction (also referred to as "purity") within the sample may be obtained. This information may include allele-specific copy number metrics for the tumor fraction of the sample selected from the major copy number, minor copy number, total copy number, average B allele frequency, log R value, and tumor ploidy, as well as normal copy number, or information derived from these metrics (e.g., a set of candidate joint genotypes that fit these allele-specific copy number metrics). Not all such allele-specific copy number metrics are necessary, as some contain redundant information and / or may be associated with suitable default values. For example, the normal copy number can be associated with a suitable default value (e.g., 2, assuming normal cells are diploid). Furthermore, only two of the major copy number, total copy number, and minor copy number are required to infer the third copy number. Similarly, these three values ​​can be inferred from the MBAF and log R values ​​(or vice versa). Optionally, the copy number solutions may be associated with a corresponding reliability metric. If such a metric is unavailable, each copy number solution may be assumed to be equally likely. Each candidate joint genotype includes the genotype at the position of the tumor-specific mutation in the normal population and the genotype of the tumor cell population containing the tumor-specific mutation.

[0079] In step 20, whether a tumor-specific mutation is likely to be expressed is determined by determining the likelihood of the sequence data (the number of reads containing the tumor-specific mutation and the total number of reads at the tumor-specific mutation locus) when the tumor-specific mutation is expressed and when it is not expressed. These likelihoods vary depending on the probability of sampling a sequence read containing the tumor-specific mutation from each sample when the tumor-specific mutation is expressed or not expressed, and may depend on the genotypes of the tumor and normal cell populations, the sequencing error rate, the tumor fraction of the sample, and the proportion of the total read count of the gene containing the tumor-specific mutation that comes from the normal cell population. The likelihoods may be compared to determine whether a tumor-specific mutation is likely to be expressed. This is done by comparing the posterior probability of a tumor-specific mutation being expressed (P(E=1|b,d,α,t)) with the prior probability of the mutation being expressed (μ ρ ) and the likelihood of the sequence data when the tumor-specific mutation is (i) expressed (P(b,d|α,t,E=1)) and (ii) not expressed (P(b,d|α,t,E=0)). Alternatively, comparing the likelihoods may include determining whether the tumor-specific mutation is expressed and determining the power to detect whether the tumor-specific mutation is expressed at a predetermined false positive rate. This involves determining a threshold number of reads (b) as the number of reads such that the area under the curve of the likelihood of the number of reads exhibiting the tumor-specific mutation when the tumor-specific mutation is not expressed (P(b,d|α,t,E=0),P(b|E=0),P(b|M0)) is below a predetermined false positive rate as a function of the number of RNA reads in the sample exhibiting the tumor-specific mutation (b). c ) to determine whether a tumor-specific mutation is likely to be expressed. If the number of reads showing the mutation exceeds this threshold number of reads, it is considered that the tumor-specific mutation is likely to be expressed. The power to detect whether a tumor-specific mutation is expressed is calculated by the following formula: if the tumor-specific mutation is expressed as a function of the number of RNA reads (b) in the sample showing the tumor-specific mutation (P(b,d|α,t,E=1),P(b|E=1),P(b|M1)), then the threshold number of reads (b c) may be the area under the curve of the likelihood of the number of reads exhibiting the tumor-specific mutation exceeding a predetermined threshold. Step 20 may include determining that the tumor-specific mutation is likely to occur if the posterior probability is above a predetermined threshold. Step 20 may include determining that the tumor-specific mutation is likely to occur if the number of reads exhibiting the mutation is above a threshold number of reads. Step 20 may include determining that the tumor-specific mutation is likely to occur if the number of reads exhibiting the mutation is below a threshold number of reads, but the power to detect whether the tumor-specific mutation is exhibited is below a predetermined threshold.

[0080] Step 20 may further include estimating a value for the fraction (α) of the total read counts of the gene containing the tumor-specific mutation that is attributable to the normal cell population before determining the likelihood. This may include obtaining the total expression of the gene containing the tumor-specific mutation in a plurality of samples having different tumor purities, fitting a regression model to the total expression values ​​as a function of purity, and determining the value of α from the fitted regression model. As described further below, this may be performed using equation (31). Step 20 may further include determining whether the total expression of the gene containing the tumor-specific mutation is above a predetermined threshold. Step 20 may further include obtaining a prior probability that the tumor-specific mutation will be expressed. The prior probability may be obtained from a user, a computing device, or a database. The prior probability may be selected from a plurality of values ​​depending on (i) whether the total expression of the gene containing the tumor-specific mutation is above a predetermined threshold, and (ii) whether the total number of reads at the position of the tumor-specific mutation is above a predetermined threshold. For example, if the total number of reads at the position of the tumor-specific mutation is low (e.g., 0), but the total expression of the gene is not low (e.g., TPM>1), the tumor-specific mutation may be considered likely to be expressed, and the prior probability may be set to 0.5. As another example, if the total number of reads at the position of the tumor-specific mutation is low (e.g., 0), and the total expression of the gene is low (e.g., TPM≦1), the tumor-specific mutation may be considered unlikely to be expressed (because the entire gene is not expressed), and the prior probability may be set to a value less than 0.5 (e.g., 0.05). The total expression of the gene may be the total expression of the gene in one or more samples, or one or more other samples (e.g., samples from the same or similar tumor type), or expression data derived therefrom (e.g., expression data from one or more databases (e.g., The Cancer Genome Atlas, TCGA - www.cancer.gov / about-nci / organization / ccg / research / structural-genomics / tcga)).

[0081] In optional step 22, it is determined whether the tumor-specific mutation is likely to give rise to a neoantigen. For example, it may be determined whether the mutation is likely to result in a peptide or protein that is not expressed by germline cells (whose genomes do not contain the mutation). As another example, it may be determined whether the mutation is likely to be cloned in the tumor. This step may be performed at any time after step 14, but does not necessarily have to be performed specifically after steps 16-20. For example, candidate tumor-specific mutations may be filtered according to whether they are likely to give rise to a neoantigen before determining whether they are likely to be expressed. In step 24, tumor-specific mutations that satisfy one or more criteria applied to the results of step 20 and one or more criteria applied to the results of step 22 may be identified. These may be considered to represent candidate neoantigens, optionally candidate clonal neoantigens. In optional step 26, the results of any of the preceding steps (particularly any of steps 20-24) may be provided to a user, for example, via a user interface. These results may be used, for example, to provide immunotherapy or a prognosis for the subject, as further described below.

[0082] Applicable The methods described herein may be used to determine whether a mutation is indeed likely to be expressed, and whether a mutation is identified as unlikely to be expressed due to low power to detect its expression in the sequence data at hand. Additionally, the methods described herein may be used to detect mutations from RNA data (e.g., mutations that are detectable solely from RNA-seq data, or that are advantageously detectable from RNA-seq data, examples of which include splice variants such as gene fusions and retained introns). Accordingly, methods for detecting candidate tumor-specific mutations are also described herein. These methods involve using the methods described herein to determine whether a candidate tumor-specific mutation is likely to be expressed in a sample. The methods described herein may also be used to determine the RNA sequencing depth required to enable detection of a particular mutation with a predetermined minimum power. For example, this approach may be used to determine the power to detect that a mutation will be expressed, given multiple candidate sequencing depths (resulting in corresponding b and d values). For use in sequencing samples suspected of having or expressing a mutation, a significance level P(b ≧ b c |M0) and b>b cA sequencing depth that satisfies the above condition and has a power greater than a predetermined value may be selected. Accordingly, a method for determining a sequencing depth to use for sequencing RNA in a tumor sample is also described herein. The method includes: (i) simulating RNA-seq data, including the number of RNA reads (b) in the sample that exhibit the tumor-specific mutation and the total number of RNA reads (d) at the position of the tumor-specific mutation, corresponding to one or more sequencing depths, if the tumor-specific mutation is truly expressed in the sample; using the RNA-seq data to determine the likelihood of the sequence data when the tumor-specific mutation (i) is expressed (P(b,d|α,t,E=1)) and (ii) is not expressed (P(b,d|α,t,E=0)); thereby determining whether the tumor-specific mutation is likely to be expressed in the sample when RNA-seq data is obtained from the sample at one or more candidate sequencing depths; and (ii) selecting a sequencing depth sufficient to determine that the tumor-specific mutation is likely to be expressed in the sample. In other words, also described herein is a method for determining the sequencing depth used for sequencing the RNA in tumor samples.This method includes: Use the RNA sequence data at one or more candidate sequencing depths to carry out the method for determining whether tumor-specific variants are likely to be expressed; And use the result of determination to select candidate sequencing depths, so that the tumor-specific variants that truly express are determined to be likely to be expressed according to the method described herein.

[0083] The above-mentioned method finds application in the aspects of cancer diagnosis, prognosis, and treatment approaches. In particular, the above-mentioned method may be used to provide immunotherapy targeting neoantigens. Accordingly, a method for providing immunotherapy to a subject is also described herein. The method includes identifying one or more neoantigens from one or more samples from a subject.

[0084] 2B schematically illustrates an exemplary method for providing immunotherapy. In optional step 210, one or more samples containing tumor genetic material and, optionally, one or more germline samples are obtained from a subject. The subject may be a subject diagnosed with cancer, and may be (but need not be) the same subject to whom immunotherapy is provided. In step 212, a list of candidate neoantigens is obtained. This may include step 212' of obtaining a list of candidate neo-antigens from genomic sequence data and / or step 212" of obtaining a list of candidate neo-antigens from RNA-seq data. In optional step 212', a list of candidate neo-antigens may be obtained from genomic sequence data from the sample(s) using methods known in the art, for example, as described in WO 2016 / 16174085, Landau et al. (2013), Lu et al. (2018), Leko et al. (2019), Hundal et al. (2019), and others. The list of candidate neo-antigens may include a single neo-antigen, or multiple neo-antigens. Preferably, the list includes multiple neo-antigens. The neo-antigens may be clonal neo-antigens. Methods for identifying clonal neo-antigens are known in the art and include methods described in WO 2016 / 16174085, Landau et al. (2013), Roth et al. (2014), McGranahan et al. al. (2016), and methods described in co-pending application PCT / EP2022 / 058793. Alternatively, or in addition to step 212', step 212" may identify one or more candidate neo-antigens from the RNA-seq data from the sample(s), for example, by identifying one or more RNA-seq reads that contain variants. Variants may be identified by comparing them to expected healthy sequences, such as a reference genome or transcriptome or RNA / DNA sequences from a normal sample of the subject.Thus, step 212" may include optional step 212a", in which the RNA-seq content of one or more mixed samples, and optionally the RNA-seq content of matched samples, may be determined, for example, by sequencing the RNA (or mRNA) in the samples using RNA sequencing. While alternative methods, such as expression arrays, may be used, sequencing methods are preferred because they generate a digital output representative of the number of each specific sequence in the sample. Step 212" may further include optional step 212b", analyzing the RNA-seq data to identify one or more mutations that are likely present in tumor cells but absent in non-cancerous cells. These represent tumor-specific mutations and may be used as candidate neo-antigens. This may include aligning the RNA sequences from one or more samples (i.e., the mixed sample(s) and, if available, the germline sample(s). This may further include identifying positions where the tumor's RNA sequence differs from the germline sequence, or can be assumed to differ from the germline sequence (e.g., if the germline sequence of the subject is not available), for example, based on a reference genome or transcriptome. Alternatively, or in addition, step 212b" may include aligning RNA sequences from one or more samples to a reference genome or transcriptome and identifying sequences (e.g., novel transcripts, splice variants, fusions, single nucleotide variants, indels) that are not expected to be present in such reference. For example, fusions or splice variants may be identified if one or more reads are aligned to non-contiguous sections of a reference transcriptome or genome, or if they are not aligned to a reference genome or transcriptome.

[0085] Step 212' may include optional step 212a', in which the sequence content of one or more mixed samples, and optionally the sequence content of matched samples, may be determined by sequencing the genomic material within the samples, for example, using either whole-exome sequencing or whole-genome sequencing. For example, while alternative methods (e.g., allele-specific copy number arrays) may be used, sequencing methods are preferred because they generate a digital output representative of the number of each specific sequence within the sample. Optional step 212b' may analyze the sequence data to identify one or more mutations that are likely present in tumor cells but absent in non-cancerous cells. These represent tumor-specific mutations and may be used as candidate neoantigens. This may include aligning sequences from one or more samples (i.e., the mixed sample(s) and, if available, the germline sample(s)) and identifying genomic locations where the tumor sequence differs from the germline sequence or can be assumed to differ from the germline sequence (e.g., if the subject's germline sequence is not available). In optional step 212c', genomic sequence data of the mixed sample at the genomic location of the candidate tumor-specific mutation is obtained. This data includes counts of reads supporting the mutant allele (also called "non-reference allele"), counts of reads supporting the germline allele(s) at the genomic location (A, collectively called "germline alleles" when the locus is heterozygous in the germline population, also called "reference," "wildtype," or "normal" alleles), and / or total counts of reads at the genomic location of the candidate tumor-specific mutation. Because any two of these metrics can infer a third, only two of them need to be obtained. The sequence data may alternatively or additionally include read data or intensity data from which counts can be obtained. In optional step 212d', information regarding at least one copy number solution matching each sample containing tumor genetic material may be obtained.This information may include allele-specific copy number metrics for the tumor fraction of the sample selected from the major copy number, minor copy number, total copy number, mean B allele frequency (MBAF), log R value, and tumor ploidy, as well as normal copy number, or information derived from these metrics (e.g., a set of candidate joint genotypes that match these allele-specific copy number metrics). Not all such allele-specific copy number metrics are necessary because some contain redundant information and / or may be associated with suitable default values. For example, the normal copy number may be associated with a suitable default value as described above. Furthermore, to infer the third copy number, only two of the major copy number, total copy number, and minor copy number are required. Similarly, these three values ​​can be inferred from the MBAF and log R values ​​(or vice versa). Optionally, the copy number solution may be associated with a corresponding reliability metric. If such metrics are unavailable, each copy number solution may be assumed to be equally likely. Each candidate joint genotype includes a genotype at the position of the tumor-specific mutation in a normal population, a reference tumor population that does not contain the tumor-specific mutation, and a variant tumor cell population that contains the tumor-specific mutation. In optional step 212e', the posterior probability that the tumor-specific mutation is clonal is determined based on the prior probability that the mutation is clonal, taking into account the tumor fraction for each of one or more samples and one or more candidate joint genotypes, and the probability of observing sequence data when the tumor-specific mutation is (i) clonal and (ii) non-clonal. Methods for obtaining such posterior probabilities are further described below. A prior probability is a probability that represents a belief about a quantity before any evidence is taken into account. For example, the prior probability that a mutation is clonal may represent the probability that a mutation is clonal in a tumor based on prior knowledge or assumptions and without considering sequence data from mixed samples.

[0086] In step 214, it is determined whether the candidate neo-antigen is likely to be expressed in the patient's tumor using the method described in connection with FIG. 2A. In step 216, tumor-specific mutations that satisfy one or more criteria applied to the results of step 214 and, optionally, one or more other criteria, are selected. For example, tumor-specific mutations associated with a probability of expression above a predetermined threshold may be selected. Alternatively, tumor-specific mutations associated with a probability of expression below a predetermined threshold may be excluded. As another example, tumor-specific mutations associated with the power of expression above a predetermined threshold and the likelihood of a model assuming that mutations above the predetermined threshold are not expressed may be excluded. As another example, tumor-specific mutations associated with the likelihood of a model assuming that mutations below the predetermined threshold are expressed may be excluded unless they are also associated with the power of expression below the predetermined threshold. One or more additional criteria may be applied in conjunction with the determination of step 214 to provide information regarding whether the candidate neo-antigen is likely to represent a true neo-antigen. For example, it may be determined whether the mutation is likely to result in a peptide or protein that is not expressed by germline cells (whose genomes do not contain the mutation). This step may be performed at any time after candidate neo-antigens have been identified. One or more additional criteria may be applied that provide information about whether the candidate neo-antigen is likely to be cloned in the tumor. For example, tumor-specific mutations associated with a probability of clonality (e.g., as determined in step 212e') above a predetermined threshold may be selected. These criteria may be applied in any order. For example, candidate tumor-specific mutations may be filtered according to whether they are likely to give rise to neo-antigens before determining whether the tumor-specific mutation is likely to be clonal, or vice versa.

[0087] In step 218, an immunotherapy targeting at least one (and optionally multiple) of the selected candidate neo-antigens is designed. Designing such an immunotherapy may include, in step 218A, identifying one or more candidate peptides for each of the candidate clonal neo-antigens. For example, for at least one of the candidate clonal neo-antigens, multiple peptides may be designed that differ in length and / or location of sequence variations characterizing the neo-antigen compared to the corresponding germline peptide. In step 218B, the identified peptide or peptides may be assayed in vitro and / or in silico to evaluate one or more properties (e.g., immunogenicity, likelihood of being displayed by an MHC molecule, etc.). In optional step 218C, one or more of the peptides may be selected, e.g., based on the results of step 218B. In step 220, the selected peptides may be obtained. Peptides having the selected sequences may be obtained using any method known in the art (e.g., using an expression system or by direct synthesis). In step 222, the one or more candidate peptides may be used to achieve immunotherapy. The immunotherapy may include one or more candidate peptides or substances sufficient for their expression (e.g., in the case of an immunogenic composition or vaccine), or may include molecules or cells obtained using the candidate peptides (e.g., in the case of a therapeutic antibody that selectively binds the candidate peptides, or immune cells that specifically recognize the candidate peptides). In optional step 224, the immunotherapy may be administered to a subject, preferably the subject from whom the sample used to identify the neoantigens was obtained. An example is described in which immunotherapy is achieved including a T cell population selectively enriched for T cells that recognize one or more (optionally clonal) neoantigens. In step 222A, a T cell population may be obtained. The T cells may be harvested from the subject to be treated, but this is not required. The T cells may be obtained from a tumor sample, a blood sample, or any other tissue sample. In step 222B, a dendritic cell population may be obtained.For example, the dendritic cell population may be derived from mononuclear cells (e.g., peripheral blood mononuclear cells, PBMCs) from the subject to be treated. In step 222C, the dendritic cell population may be pulsed with a candidate peptide. In step 222D, the pulsed dendritic cell population may be used to selectively expand a T cell population. Additional growth factors (e.g., cytokines or stimulatory antibodies) may also be used.

[0088] Thus, the present disclosure also provides a T cell composition comprising a T cell population selectively enriched for T cells that recognize one or more neo-antigens likely to be expressed in tumors, where the one or more neo-antigens have been identified using any of the methods described herein. The likely expressed neo-antigens may be neo-antigens selected using the results of a method as described herein (e.g., as described with reference to step 216 of FIG. 2B).

[0089] In the T cell compositions described herein, the expanded population of neoantigen-reactive T cells may have higher activity than the non-expanded T cell population, as measured by the response of the T cell population to restimulation with neoantigen peptides. Activity may be measured by cytokine production, with higher activity being a 5-10 fold or greater increase in activity.

[0090] Reference to multiple neo-antigens refers to multiple peptides or proteins, each of which may contain a different tumor-specific mutation that generates the neo-antigen. The plurality may be 2-250, 3-200, 4-150, or 5-100 tumor-specific mutations (e.g., 5-75 or 10-50 tumor-specific mutations). Each tumor-specific mutation may be represented by one or more neo-antigen peptides. In other words, the multiple neo-antigens may include multiple different peptides, some of which contain sequences that include the same tumor-specific mutation (e.g., at various positions within the peptide sequence or within peptides of various lengths). The tumor-specific mutations selected according to the methods described herein may be mutations determined using the methods described herein to be likely to be expressed in the tumor of the patient to be treated. The tumor-specific mutations selected according to the methods described herein may be mutations determined to be likely to be clonal in the tumor of the patient to be treated.

[0091] T cell populations produced in accordance with the present disclosure will have an increased number or proportion of T cells targeting one or more neo-antigens predicted to be expressed (and optionally clonal). That is, the composition of the T cell population will differ from that of a "natural" T cell population (i.e., a population that has not undergone the expansion process discussed herein) in that there will be an increased proportion or proportion of T cells targeting neo-antigens predicted to be expressed (and optionally clonal). A T cell population according to the present disclosure may have at least about 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95 or 100% of the T cells target the neo-antigen predicted to be expressed.

[0092] The immunotherapeutics described herein may be used to treat cancer. Accordingly, the present disclosure also provides a method of treating cancer in a subject, comprising administering to the subject an immunotherapeutic composition as described herein.

[0093] Suitably, in any embodiment of any aspect described herein, the cancer may be ovarian cancer, breast cancer, endometrial cancer, kidney cancer (renal cell), lung cancer (small cell, non-small cell, and mesothelioma), brain cancer (glioma, astrocytoma, glioblastoma), melanoma, Merkel cell carcinoma, clear cell renal cell carcinoma (ccRCC), lymphoma, small intestine cancer (duodenal and jejunal), leukemia, pancreatic cancer, hepatobiliary tumor, germ cell cancer, prostate cancer, head and neck cancer, thyroid cancer, and sarcoma. For example, the cancer may be lung cancer (such as lung adenocarcinoma or lung squamous cell carcinoma). As another example, the cancer may be melanoma. In embodiments, the cancer may be selected from melanoma, Merkel cell carcinoma, kidney cancer, non-small cell lung cancer (NSCLC), urothelial carcinoma of the bladder (BLAC), head and neck squamous cell carcinoma (HNSC), and microsatellite instability (MSI)-high cancer. In some embodiments, the cancer is non-small cell lung cancer (NSCLC). In other embodiments, the cancer is melanoma.

[0094] Treatment using the compositions and methods of the present disclosure may also include targeting circulating tumor cells and / or tumor-derived metastases. The methods and uses for treating cancer described herein may be practiced in combination with additional cancer therapies. In particular, the T cell compositions described herein may be administered in combination with immune checkpoint intervention, costimulatory antibodies, chemotherapy and / or radiation therapy, targeted therapy, or monoclonal antibody therapy. "In combination" may refer to administering an additional treatment before, simultaneously with, or after administration of a T cell composition as described herein.

[0095] The present disclosure also provides methods for producing immunotherapeutic compositions, which include identifying neoantigens that are likely to be expressed and producing immunotherapeutic compositions that target the neoantigens.

[0096] Also described herein is a method of treating a subject diagnosed with cancer, the method comprising: identifying a plurality of tumor-specific mutations in the subject; determining whether one or more of the tumor-specific mutations are likely to be expressed in the subject; optionally, determining whether one or more of the tumor-specific mutations are likely to be clonal in the subject; selecting one or more of the tumor-specific mutations as candidate neo-antigens, where the candidate neo-antigens are tumor-specific mutations that meet at least one or more predetermined criteria for whether the tumor-specific mutations are likely to be expressed; and treating the subject with an immunotherapy targeting one or more of the selected candidate neo-antigens, wherein determining whether the tumor-specific mutations are likely to be expressed in the subject is performed using the methods described herein.

[0097] In particular, determining whether a tumor-specific mutation is likely to be expressed in a subject may include: obtaining, by a processor, RNA-seq data from one or more samples from the subject containing tumor genetic material, the sequence data including, for each of the one or more samples, at least two of: (b) the number of RNA-seq reads in the sample that represent the tumor-specific mutation; (d) the number of RNA-seq reads in the sample that represent the corresponding germline allele; and (d) the total number of RNA-seq reads at the location of the tumor-specific mutation; and determining, by the processor, a posterior probability that the tumor-specific mutation will be expressed as a function of the average of the prior probabilities that the mutation will be expressed and a likelihood of the sequence data if the tumor-specific mutation is (i) expressed and (ii) not expressed, wherein the likelihood of the sequence data is conditional on the tumor fraction for each of the one or more samples and the proportion of total expression at the locus of the tumor-specific mutation that is assumed to be derived from a normal cell population that does not contain the tumor-specific mutation.

[0098] The method includes determining whether a tumor-specific mutation is likely to be clonal in the subject by obtaining, by a processor, genomic sequence data from one or more samples from the subject containing tumor genetic material, the sequence data comprising, for each of the one or more samples, a number of reads within the sample that exhibit the tumor-specific mutation (d b ), the number of reads in the sample representing the corresponding germline allele, and the total number of reads at the position of the tumor-specific mutation (d); and determining, by the processor, the posterior probability that the tumor-specific mutation is clonal, given the tumor fraction for each of the one or more samples and one or more candidate joint genotypes, as a function of the prior probability that the mutation is clonal and the probability of observing the sequence data when the tumor-specific mutation is (i) clonal and (ii) non-clonal, wherein each candidate joint genotype includes the genotype at the position of the tumor-specific mutation of a normal population, a reference tumor population that does not contain the tumor-specific mutation, and a variant tumor cell population that contains the tumor-specific mutation.

[0099] Candidate neo-antigens may be selected as tumor-specific mutations that further satisfy at least one or more predetermined criteria related to whether the tumor-specific mutation is likely to be clonal and / or likely to give rise to a neo-antigen. Selecting one or more of the tumor-specific mutations as candidate neo-antigens by the processor may include determining whether the one or more tumor-specific mutations satisfy one or more criteria related to whether the tumor-specific mutation is likely to give rise to a neo-antigen, selected from a mutation predicted to give rise to a protein or peptide that is not expressed in normal cells of the subject, a mutation predicted to give rise to at least one peptide likely to be presented by an MHC molecule, a mutation predicted to give rise to at least one peptide likely to be presented by an MHC allele known to be present in the subject, a mutation likely to be clonal, and a mutation predicted to give rise to a protein or peptide that is immunogenic. Selecting, by the processor, one or more of the tumor-specific mutations as candidate neo-antigens may include determining, by the processor, whether the one or more tumor-specific mutations satisfy one or more predetermined criteria related to whether the tumor-specific mutations are likely to be clonal, selected from: mutations that have a likelihood of being clonal that exceeds a predetermined threshold; and mutations that have a likelihood of being clonal that exceeds a threshold adaptively set to select a predetermined number of tumor-specific mutations with the highest likelihood of being clonal among the tumor-specific mutations whose likelihoods have been determined, and that exceed a threshold adaptively set to select a predetermined top percentile of the tumor-specific mutations whose likelihoods have been determined.

[0100] The immunotherapy targeting one or more of the selected neo-antigens may be an immunogenic composition, a composition comprising immune cells, or a therapeutic antibody. The immunotherapy may be a composition comprising T cells that recognize at least one of the one or more identified selected neo-antigens. The composition may be enriched for T cells that target at least one of the one or more identified selected neo-antigens. The method may include obtaining a T cell population and expanding the T cell population to increase the number or relative proportion of T cells that target at least one of the one or more identified selected neo-antigens.

[0101] Determining the posterior probability that a candidate tumor-specific mutation is clonal, given the tumor fraction for each of one or more samples and one or more candidate joint genotypes, as a function of the prior probability that the mutation is expressed and the probability of observing sequence data when the tumor-specific mutation is (i) expressed and (ii) not expressed, may be performed using the approach described in the following section, where each candidate joint genotype includes the genotype at the position of the tumor-specific mutation in a normal population, a reference tumor population that does not contain the tumor-specific mutation, and a variant tumor cell population that contains the tumor-specific mutation.

[0102] Identification of clonal mutations Embodiments of the methods described herein may include determining whether a tumor-specific mutation is likely to be clonal. One possible method for this determination is described in further detail herein. Identifying candidate tumor-specific mutations that are likely to be clonal may use data including allele counts of N mutations (n=1,...,N) from S samples (s=1,...,S). For simplicity, and because the method can analyze a single sample and mutation, the mutation index n and sample index s are not explicitly included in the notation used in this section. The model described in this section assumes that each mutation divides the set of sequenced cells into three subpopulations: (i) a normal cell population composed of cells with a healthy germline genome (likely diploid at the mutated region), (ii) a reference cell population composed of cancer cells without the mutation in question (which may be aneuploid at the mutated region), and (iii) a variant cell population composed of cancer cells with the mutation in question (which may be aneuploid at the mutated region and may not have the same copy number at that region as the reference population). The term "mutation" is intended herein in its broadest sense to refer to any genetic change detectable in sequence data, particularly genomic sequence data, including, among others, single nucleotide variants (SNVs), multiple nucleotide variants (MNVs), indels, and the like.

[0103] Let G = (A, B, AA, AB, AAA, AABB, ...) be the set of all genotypes, where A and B represent the reference allele and the variant allele, respectively. For example, AB represents a heterozygous variant (containing one reference / normal allele A and one variant allele B) with a total copy number of 2. Based on this notation, in Figure 4, the normal population has genotype AA (both As can be the same or different, i.e., the normal population can be homozygous or heterozygous, but both alleles are normal), the reference population has genotype AAA (the A allele is selected from the A alleles of the normal population), and the variant population has genotype AABB (the A allele is selected from the A alleles of the normal population, and the B allele is any non-reference allele). We assume that the genotype of all cells in each subpopulation is constant (i.e., with reference to Figure 4, all cells in the normal population have genotype AA, all cells in the reference population have genotype AAA, and all cells in the variant population have genotype AABB). H ;G R ;G V )∈G 3 Let be a vector, where each entry is a genotype for the normal (healthy), reference, and variant populations, respectively (each of these individual genotypes will be referred to collectively below as "G"). Let t be the proportion of cancer cells in the sample. This is often referred to as the tumor content, tumor purity, or cellularity of the sample. Let φ be the proportion of cancer cells in the sample that have the mutation, i.e., the relative proportion of cancer cells in the variant population. This is often referred to as the cancer cell fraction (CCF) or cellular prevalence of the mutation. Let ε be the expected sequencing error rate. The following function is defined, i.e., a(G):G→N is a function that maps genotype to the number of A alleles (e.g., if G is AA, then a(G)=2) b(G):G→N is a function that maps genotype to the number of B alleles (e.g., if G is AA, then b(G)=0) c(G):G→N is a function that maps genotype to total copy number at a locus (i.e., c(G)=a(G)+b(G); e.g., if G is AA, then c(G)=2). μ(G):G→N is a function that maps genotypes to the value μ(G)=min{max{(b(G) / c(G)), ε}, (1-ε)}, which can be interpreted as the probability of sampling a read containing a mutation from a population with genotype G.

[0104] Let ξ(G, φ, t) be the probability of sampling a read containing a variant allele. Assuming the initial cell population sampled during sequencing is infinite, the probability of sampling a read containing a variant allele is roughly proportional to the copy number of the variant allele in the input pool of DNA. More formally, taking into account sequencing error, the probability of sampling a variant allele (given a set of genotypes G, tumor content t, and cancer cell fraction φ) is given by the following formula (Equation (1)):

[0105]

number

[0106] The variable d is the total number of reads covering the mutation in the sample, of which d b contains the mutant allele. Therefore, the number of these reads, d, d b The probability of observing (P(d, d b |G, φ, t)) is the parameter d b and ξ(G, φ, t) (Equation (3)). This means that the sum of m Bernoulli random variables with parameter p is given by the parameters m, ρ 2For example, if the data have more variance than can be explained by the Binomial model, a BetaBinomial model with mean ξ(G, φ, t) and precision (the inverse of the variance) Y (Equation (4)) can be used instead. P(d, d b |G, φ, t) = Binomial(d b |d,ξ(G,φ,t)) (103) P(d, d b |G, φ, t, Y)=BetaBinomial(d b |d, ξ(G, φ, t), Y) (104) In the examples below, the parameter Y is set to 200, but other values ​​are possible. So far, we have assumed that the genotypes of the subpopulations are known. In general, this may be true for healthy populations (e.g., matched germline samples), but not for the reference and variant populations. Instead, it is common to observe allele-specific copy number estimates for regions that overlap with the variant. Using this information, we can derive prior probabilities for a set of plausible genotypes. We explain how in the next section. We now assume that we have a vector of prior probabilities, π, where π i is the i of the population th The second plausible joint genotype G i The probability that the observed data are marginalized across all plausible genotypes can be expressed as follows (Equations (103a) and (104a)): P(d, d b |π, φ, t)=Σ i π i Binomial(d b |d, ξ(G i , φ, t)) (103a) P(d, d b |π,φ,t,Y)=Σ i π i BetaBinomial(d b |d, ξ(G i , φ, t), Y) (104a) In the following sections, Pr(d, db We will refer to the expressions in Equation (3a) and Equation (4a) similarly using the notation |π, φ, t). Since φ and t are associated with individual samples, the above notation translates to φ and t, respectively. s and t s It should be noted that this is an abbreviation of

[0107] Derivation of variant genotype prior probabilities The above model uses either known joint genotypes, or a prior probability π, where π i is the i of the population th The second plausible joint genotype G i is the prior probability of (i.e., G i (where σ is one possible combination of genotypes in the healthy, variant, and reference populations). Various methods can be used to set potential genotype priors. The same principles apply to determining whether a tumor-specific mutation is likely to be expressed, with the difference that the joint genotype G refers to the genotypes of the tumor (variant) and normal (no variant / germline) populations (i.e., there is no reference tumor population).

[0108] For example, one possible method is called the "major copy number" method. major and c minor Let σ represent the major and minor allele copy numbers for the region overlapping with the mutation in the tumor sample. The "major copy number method" considers the following two cases: (a) In the first case, the mutation occurs before the copy number event. In this case, the genotype of the reference population matches that of the normal population. The maximum c containing the variant major For a variant population with chromosomes, all possible variant genotypes are considered. (b) In the second case, the mutation occurs after a copy number event. In this case, the reference population is major +c minorThe variant population has one variant allele and c major +c minor -Has one reference allele.

[0109] Prior weights were set to be equal for all possible variant genotypes. For example, c major =2, c minor = 1 and the normal copy number is 2. The possible genotypes are as follows: G1=(AA, AA, AAB) G2=(AA, AA, ABB) G3=(AA, AAA, AAB) The prior probability for each is 1 / 3. If allele-specific copy numbers are not available, major Set to the total number of copies and c minor Note that we can set σ to zero. This approach assumes that a mutation occurs only once, so if there is more than one copy of the mutant allele in a variant population, this must have occurred because the mutation preceded a copy number change at the locus and subsequently amplified. This approach strikes a good balance between accounting for uncertainty in the genotypes of a population and, on the other hand, not accounting for too many states.

[0110] In terms of determining whether a tumor-specific variant is expressed, in the above example the following possible genotypes are considered: G1 = (AA, AAB), G2 = (AA, ABB), G3 = (AA, AAB). References below to using maximum a posteriori (MAP) estimates of genotype refer to the use of MAP estimates of major and minor copy number and purity from DNA sequence data for a sample(s) using the major copy number method described above.

[0111] Alternative approaches may be used to set the prior probability of variant genotypes. Another possible approach is to simply assume that each variant is diploid and heterozygous (i.e., variants in the variant population occur on only one of the two chromosomes), and in terms of determining whether a tumor-specific mutation is expressed, G = (G H =AA, G R =AA, G V =AB), or G=(G H =AA, G V =AB). This is sometimes called the "AB prior." Yet another simple approach is to assume that each mutation is both diploid and homozygous (i.e., variants in the variant population occur on both chromosomes, and in terms of determining whether a tumor-specific mutation is expressed, G = (G H =AA, G R =AA, G V =BB), or G=(G H =AA, G V =BB). This is sometimes called the "BB prior." Yet another possible simple approach is to assume that the genotypes in the variant population have the expected total copy number in the region of the mutation and have exactly one variant allele (i.e., assuming a total copy number of 3, then in terms of determining whether a tumor-specific mutation is expressed, G = (G H =AA, G R =AA, G V =AAB), or G=(G H =AA, G V = AAB), i.e., the "major copy number" method above results in only G1 being considered. This is sometimes called "no zygosity prior." These approaches may be overly simplistic in many cases, as they essentially consider a single possible genotype.

[0112] Another possible approach is to assume that the genotype of the variant population has the expected total copy number at the region of the mutation and has at least one variant allele, and that the reference population is either AA or has a genotype with a copy number equal to the expected total copy number and no variant allele (equally likely). This is sometimes called the "total copy number prior," and intuitively means that the genotype of the variant population at the locus can have the expected total copy number and any number (>0) of copies of the variant allele (i.e., assuming a total copy number of 3, in terms of determining whether the tumor-specific variant is expressed, the possible genotypes are, with equal probability, G = (G H =AA, G R =AA, G V =AAB),G2=(G H =AA, G R =AA, G V =ABB),G3=(G H =AA, G R =AA, G V =BBB),G4=(G H =AA, G R =AAA,G V =AAB) or G R(That is, it essentially ignores the major and minor copy number values ​​and considers all possible genotypes with n copies. This leads to additional genotypes being considered compared to the "major copy number" method described above.) Yet another approach that can be used is to "trust" the number of major and minor alleles predicted by the copy number caller, so that only genotypes with multiple mutant alleles corresponding to either the major or minor copy number are considered. This is sometimes referred to as the "parental" mode. For example, if the major copy number = 3 and the minor copy number = 1, this approach considers the following possible genotypes with equal probability in terms of determining whether the tumor-specific variant is expressed: G1 = (AA, AA, AAAB), G2 = (AA, AA, ABBB), G3 = (AA, AAAA, AAAB) (i.e., there are either one or three mutant alleles in the variant population), or G R In contrast, the "major copy number" approach "trusts" the range of possible major copies by considering all values ​​between 1 and the expected major copy number, but does not trust its absolute value. In the example above with major copy number = 3 and minor copy number = 1, this leads to one more genotype being considered compared to the "parental" mode. That is, in terms of determining whether a tumor-specific variant is expressed, G1 = (AA, AA, AAAB), G2 = (AA, AA, AABB), G3 = (AA, AA, ABBB), G4 = (AA, AAAA, AAAB), or G R Therefore, the "major copy number" approach strikes a good balance between considering additional uncertainty from copy number calling (compared to the "parental" approach) and not having to consider too much uncertainty (compared to the "total copy number" approach).

[0113] Clonality estimation model Based on the above, a hierarchical Bayesian model may be used to identify ubiquitous mutations. Let Z be a Bernoulli variable, equal to 1 if the mutation is ubiquitous (assumed to be clonal) and 0 otherwise. Let ρ be the prior probability that the mutation is ubiquitous. In the example below, this is set to 0.5. As mentioned above, φ is the proportion of cancer cells in the sample that carry the mutation. Therefore, the model can be expressed as: Z|ρ≒Bernoulli(Z|ρ) (105) φ|Z≈Beta(φ|α=1,β=1), where Z=0; Beta(φ|α,β=1), where Z=1 (106) d b , d|π,φ,t≒Pr(d,d b |π, φ, t) (107) where α is a parameter greater than 1 in the distribution with φ|Z=1. In the following example, it is set to α=99. A beta distribution with parameters α=99 and β=1 is skewed toward 1, reflecting the assumption that the higher the cancer cell fraction φ, the more abundant clonal mutations should be. While other values ​​for the parameter α are possible, a value that reflects this assumption is preferred. As noted above, the probability in equation (107) is given by equation (103) / (103a) or equation (104) / (104a).

[0114] The joint distribution can be expressed by the following equation (Equation (108)). p(d b , d, φ, Z=z|π, t, ρ)=p(Z=z|ρ)Pr(d b , d|π, φ, t)p(φ|Z=z) (108) However, if there is one sample, or if there are multiple samples,

[0115]

number

[0116]

number

[0117]

number

[0118] This amount

[0119]

number

[0120]

number

[0121] Finally, the quantity we wish to estimate is the observed lead (d b , d), the genotype prior probability (π), the tumor fraction estimate (t), and the prior probability that the mutation is clonal (ρ) are the probability that the mutation is clonal (i.e., P(Z=1|d b, d, π, t, ρ) we want to estimate. Taking the above into consideration, we can express it as follows:

[0122]

number

[0123]

number

[0124] Thus, estimating equation (111) for the case z=1 (i.e., equation (111a)) gives the probability that the mutation is ubiquitous (i.e., assumed to be clonal given the available sample or samples). This requires evaluating S one-dimensional integrals (one for each sample in equations (109), (110)), which can be done efficiently using known numerical integration. Any numerical integration algorithm known in the art may be used for this purpose. For example, a grid approximation may be used. This is advantageously simple and sufficient, considering that there is only one parameter (φ) to integrate.

[0125] This allows estimates of the probability that a mutation is clonal to be calculated efficiently given the available data, be easily interpretable (given a rigorous statistical model that uses explicit and clear assumptions), be obtainable without manual input for any mutation, be independent of any other mutations analyzed, can strictly include prior knowledge about the mutation, and can be used to objectively and automatically prioritize lists of mutations (including their associated probabilities) for testing and / or use.

[0126] Accounting for uncertainty in copy number prediction While the above models already present numerous advantages, they can be further enhanced by considering the uncertainty in the prediction of the copy number estimates used in the models. Indeed, the above models assume that copy numbers (e.g., the major / minor / total / copy numbers used to derive genotype prior probabilities) are accurately predicted. In reality, these values ​​may have some uncertainty. Indeed, the problem of allele-specific copy number analysis of tumors is complex, and many solutions have been proposed to perform this task. One commonly used approach is ASCAT (Allele-Specific Copy Number Analysis of Tumors, Van Loo et al., 2010). It considers both tumor cell aneuploidy and non-aberrant cell infiltration when interpreting bulk copy number profiles, and outputs estimated allele-specific copy number profiles and associated tumor purity estimates. In other words, ASCAT evaluates multiple possible combinations of tumor ploidy and tumor fraction based on the assumption that the associated allele-specific copy number calls should be as close as possible to the non-negative natural number of germline heterozygous single nucleotide polymorphisms (SNPs). The solution that is judged to be optimal (estimates of tumor ploidy, tumor purity, and allele-specific copy number calls for the tumor and normal portions of the sample) is then reported along with its goodness of fit (based on the assumptions above).

[0127] The model provided above can be adjusted to account for multiple copy number solutions and their uncertainties. To do so, π is modified to include genotype entries from each predicted copy number state (e.g., each proposed solution including the major and minor copy number states), weighted by the probability associated with this state. Furthermore, since tumor purity estimates may be estimated together with these copy number states (e.g., as is the case when approaches such as ASCAT are used), the associated tumor purity estimates can also be taken into account. Note that this is not necessarily the case, for example, if tumor purity is estimated or measured separately and is not inherently associated with the copy number state estimate. Nevertheless, for generality, a set of C possible copy number / tumor content states (e.g., c major , c minor , and C possible sets of estimates of t). c Let be a vector, where each entry represents the probability of each such possible set of estimates. For each state C, we define a vector of possible genotypes, π, as explained above. CG It is possible to calculate π CG We can obtain the final genotype vector by multiplying πc by the entry for state C. This produces the slightly modified formula: P(d, d b |π, φ) = Σ i π i Binomial(d b |d, ξ(G i ,φ,t i )) (103b) P(d, d b |π, φ, Y)=Σ i π i BetaBinomial(d b |d, ξ(G i ,φ,t i ), Y) (104b) where the tumor content t i may depend on the specific state (and π i is π CG to πc (These new densities can be substituted into the relevant equations above. In particular, the problem to be solved may then be expressed as solving equation (111a), where Pr(d b , d|π, φ, t) is given by equation (103b) or equation (104b). i , C major , C minor (Hence, the model used is therefore fitted to CG ), and π c The value of is provided as an output of many methods for performing allele-specific copy number analysis of tumors, as described above, including, but not limited to, ASCAT. For the avoidance of any doubt, any approach that generates an allele-specific copy number state estimate (typically associated with a tumor purity estimate) using a reliability or other evaluation metric that can be used to relatively weight multiple solutions may be used for this purpose.

[0128] implementation The method described throughout this specification can be implemented using any programming language known in the art.For example, for the purpose of identifying clonal mutation, the Python script that implements the above method can be used, and can take as input for each mutation: mutation identifier, sample identifier, the count of the number of reads that match with the reference allele at mutation position, the count of the number of reads that match with the alternative allele at mutation position; and as input for each of one or more copy number solutions, can take: the major copy number (for tumor) that overlaps with mutation for specified copy number solution; the minor copy number (for tumor) that overlaps with mutation for specified copy number solution; the copy number for normal cells in mutation (default=2 for autosomes, or can be set to 1 for sex chromosomes in male subjects); and the tumor purity value for specified copy number solution (this can be obtained as the output of ASCAT for example, or can be obtained separately). The major and minor copy numbers overlapping with the mutation in the tumor population can be obtained directly from ASCAT for a given copy number solution (e.g., using ascatNgs, Raine et al., 2016) or can be derived from the output of ASCAT, for example, using the mean B allele frequency of the copy number segment overlapping with the mutation, the log R value of the copy number segment overlapping with the mutation, and the ploidy of the solution. For example, the allele-specific copy number estimate for the tumor at position i

[0129]

number

[0130]

number

[0131] As another example, a Python script implementing the above method may be used to identify expressed mutations, taking as inputs for each mutation a mutation identifier, a sample identifier, a count of the number of RNA reads matching the reference allele at the mutation position, a count of the number of RNA reads matching the mutant allele at the mutation position, and the genotypes of normal and tumor cells (e.g., obtained using ASCAT and / or any of the approaches described herein for inferring joint genotypes of tumor and normal populations from mixed samples), and a tumor purity value (optionally associated with a specific joint genotype, an example of which is obtained as an output from ASCAT, etc.). The script may generate as output the mutation identifier and the posterior probability that the mutation will be expressed.

[0132] system FIG. 3 illustrates an embodiment of a system for determining whether a tumor-specific mutation is likely to occur, and / or for identifying neoantigens, and / or for providing immunotherapy based at least in part on the identified neoantigens, according to the present disclosure. The system includes a computing device 1, which includes a processor 101 and a computer-readable memory 102. In the illustrated embodiment, the computing device 1 also includes a user interface 103, which is shown as a screen but may include any other means of communicating information to a user, such as via an audible or visual signal. The computing device 1 is communicatively connected, for example, via a network 6, to a sequence data acquisition means 3 (such as a sequencing instrument) and / or one or more databases 2 that store sequence data. The one or more databases may additionally store other types of information that can be used by the computing device 1, such as reference sequences, parameters, and the like. The computing device may be a smartphone, tablet, personal computer, or other computing device. The computing device is configured to perform a method for determining whether a tumor-specific mutation is likely to occur, as described herein. In an alternative embodiment, computing device 1 is configured to communicate with a remote computing device (not shown), which itself is configured to perform a method for determining whether a tumor-specific mutation is likely to occur, as described herein. In such a case, the remote computing device may also be configured to transmit results of the method to the computing device. Communication between computing device 1 and the remote computing device may be via a wired or wireless connection, and may be over a local network or a public network, examples of which include, for example, the public Internet or WiFi.

[0133] The sequence data acquisition means 3 may be wired to the computing device 1 or may communicate via a wireless connection, such as via a network 6, as shown. The connection between the computing device 1 and the sequence data acquisition means 3 may be direct or indirect (e.g., via a remote computer). The sequence data acquisition means 3 is configured to acquire sequence data from nucleic acid samples (e.g., RNA samples, and optionally genomic DNA samples extracted from cell and / or tissue samples). In some embodiments, the sample may have undergone one or more pre-processing steps, examples of which include RNA purification, fragmentation, library preparation, target sequence reflection (e.g., exon reflection and / or panel sequence reflection, etc.). Preferably, the sample has not undergone amplification, or if it has undergone amplification, it has undergone amplification in the presence of amplification bias control means, examples of which include, for example, the use of unique molecular identifiers. Any sample preparation process suitable for use in determining genome copy number profiles (genome-wide or sequence-specific) may be used within the scope of the present disclosure. The sequence data acquisition means is preferably a next-generation sequencer. The sequence data acquisition means 3 may be in direct or indirect connection with one or more databases 2 in which sequence data (raw data or partially processed data) may be stored.

[0134] The following are offered by way of example and should not be construed as limiting the scope of the claims. [Example]

[0135] These examples describe methods for identifying clonal mutations in accordance with the present disclosure and demonstrate their use using simulated data and multiple types of experimental data.

[0136] preface Allelic expression (AE, also referred to herein as "allele-specific expression," ASE) has been reported as a predictor of immunogenicity (Gartner et al. 2021). Confirming the presence of variant reads in RNA-seq data may determine whether a mutant allele is expressed, but the power to detect its expression depends on factors such as tumor purity / gene expression level, and it cannot distinguish between the absence of ASE and the power to detect ASE.

[0137] Furthermore, immunogenic mutations without ASE may be due to immunoediting (tumor cells silencing their expression to avoid immune detection). This is important because a power measure can separate 1) potential false negatives (low power) from 2) potential immunoedited mutations (high power).

[0138] The most basic approach is to check whether one or more variant reads within the sequencing reads cover the mutation locus and determine whether (i) there is no AE if zero variant reads are observed, or (ii) there is an AE if one or more variant reads are observed. However, the likelihood of observing one or more variant reads when a mutation is truly expressed varies depending on the tumor purity, genotype, and gene expression level of the sample. If the purity is 100% and gene expression is high, there is a high likelihood of observing one or more variant reads if the mutation is truly expressed. If the purity is 5% and gene expression is low, there is a low likelihood of observing one or more variant reads even if the mutation is truly expressed.

[0139] In this study, we present a set of statistical methods called ALExA (Achilles Likelihood of Expressed Alleles) for assessing allele-specific expression in the context of tumors. In particular, we devised a method to calculate the power to detect variant expression, taking into account tumor purity, genotype, and gene expression (Example 1). We further present a statistical method to estimate the probability of variant expression based on RNA-seq read count data (Example 2), using the concepts developed in Example 1.

[0140] Example 1 - Power of detection of allele-specific expression in tumors In this section, we present a model for the number of variant reads we expect to observe if the variant is expressed or not expressed (shown in Figure 4). We use this to derive the likelihood of observing a given data set if the variant is expressed or not expressed, and then frame this likelihood in terms of statistical hypotheses to identify the power of detecting the variant being expressed at a chosen false positive rate.

[0141] method Modeling the number of variant reads This model assumes that each mutation divides the population of sequenced cells into two subpopulations. A normal cell population consisting of cells with a healthy germline genome (genotype G N (has) A variant cell population (genotype G) consisting of cancer cells carrying the mutation in question v (has) This model assumes that the mutation is clonal (within the sample under investigation) and that the cancer cell fraction φ within the tumor is 1 (i.e., all cancer cells have the variant). In connection with determining whether a mutation is clonal, as explained above, it is possible to set a prior probability for φ (e.g., a beta distribution) and extend it to subclonal mutations by including a reference tumor population in the genotype. The prior probability for φ may be obtained, for example, from the analysis of DNA sequence data, as explained above. It should be noted that a variant cell population consisting of cancer cells carrying the mutation in question may be aneuploid in the region of the mutation in question and may not have the same copy number in that region as the normal population. The term "mutation," as used herein, is intended in its broadest sense to refer to any genetic change detectable in sequence data, particularly genomic sequence data. This includes, among other things, single nucleotide variants (SNVs), multi-nucleotide variants (MNVs), indels, etc. The present method relates to the detection of variants present in a tumor population; therefore, the mutation is a somatic mutation.

[0142] Let G = (A, B, AA, AB, AAA, AABB, ...) be the set of all genotypes, where A and B represent the reference allele and the variant allele, respectively. For example, AB represents a heterozygous variant (containing one reference / normal allele A and one variant allele B) with a total copy number of 2. Thus, a normal population has the genotype AA (both As can be the same or different, i.e., the normal population can be homozygous or heterozygous, but both alleles are normal), and a tumor population may have, for example, the genotype AABB (the A allele is selected from the A alleles of the normal population, and the B allele is any non-reference allele). The following notation is used: · d is the total number of reads covering the variants in the sample. b is the number of reads containing the mutation in the sample. · ε is the expected sequencing error rate. · a(G):G→N is a function that maps genotype to the number of A (reference) alleles (e.g., if G is AA, a(G)=2). · b(G):G→N is a function that maps genotype to the number of B (variant) alleles (e.g., if G is AA, then b(G)=0). · c(G):G→N is a function that maps genotype to total copy number at a locus (i.e., c(G)=a(G)+b(G). For example, if G is AA, then c(G)=2). t is the tumor purity (proportion of cancer cells in the sample). The value of this parameter can be assumed, known, or estimated, e.g., from DNA sequence data associated with the sample(s) under investigation. -G={G N ;G v}∈G 2 is a vector whose entries are the genotypes of the normal (healthy) and variant populations, respectively. The value of this parameter can be assumed, known, or estimated, e.g., from DNA sequence data associated with the sample(s) under investigation. Θ is the reference ratio (reference read counts / total read counts), which is the proportion of total read counts attributable to the reference allele. This parameter is integrated; therefore, its exact value cannot be determined. α is the fraction of total expression of genes containing tumor-specific mutations that is attributable to the normal cell population (e.g., TPM 正常 / (TPM 正常 +TPM 腫瘍 ), where TPM is transcripts per million and is a common expression metric in RNA sequencing. The value of this parameter can be assumed, known, or estimated, for example, from RNA-seq data from multiple samples of varying purity, as described further below. μ(G):G→N is the probability of sampling a read containing a mutation from a population with genotype G, defined as μ(G, Θ, ε) = p(m=1|G, Θ), i.e., the probability that a read selected from a sequenced cell population with genotype G carries the mutation of interest (m=1).

[0143] Probability of sampling a read containing a mutation from a population with genotype G The model presented herein models the process of sampling RNA-seq reads containing a mutation of interest from a tumor sample (the tumor sample is assumed to contain a mixed population that may include cells with at least some copies of the mutation (i.e., tumor cells) and cells without any copies of the mutation (i.e., normal cells)). This allows for differential allelic expression of the mutant allele and the reference allele. The probability of sampling a read containing the mutation of interest from a sequenced population with genotype G is given by p(m=1|G, Θ) (i.e., the probability that a read selected from a sequenced cell population with genotype G carries the mutation of interest (m=1)), as provided in Equation (1).

[0144]

number

[0145]

number

[0146] It is important to note that this model assumes that sequencing errors are independent of alleles. If this is not the case, separate error rates may be used depending on the type of mutation. For example, in the case of substitutions, different error rates may be provided for each class of substitution (transposition / transversion, or individual mutations, e.g., A->C, A->G, A->T). Furthermore, aligning and calling variant alleles may be more error-prone compared to the reference allele. For example, in some situations, such as some indel variants, reads aligned to the variant allele may be undercounted. This can be addressed by adding a bias term to the model, e.g., correcting the number of observed variant reads according to the expected proportion of reads that may not align to the variant.

[0147] Probability of sampling a read containing a mutation from a tumor sample (mixed population) Assuming an infinite initial cell population sampled during sequencing, such that the probability of sampling a read containing a variant allele is proportional to the copy number of the variant allele (b(G) for each genotype) and its relative expression (θ compared to the reference allele), the probability of sampling a read containing a mutation from a tumor sample (given the set of genotypes G, tumor content t, and normal expression fraction α) is given by equation (2)-(2'').

[0148]

number

[0149] Model likelihood If we observe d total reads covering the mutation in a sample, of which b contain the mutant allele, the likelihood of observing these number of reads is given by equation (3). P(b, d|G, θ, α, t)=Binomial(b|d, ξ(G, θ, α, t)) (3)

[0150] Let E be a binary variable representing whether the mutant allele at a site is expressed (E=1). We use the following prior probability for the reference ratio θ: p(θ|E=1)=Beta(θ, α1, β1) p(θ|E=0)=Beta(θ, α0, β0) (4) Here, the index "0" refers to the null model (E=0), denoted by M0 (the mutant allele is not expressed), and the index "1" refers to the alternative model (E=1), denoted by M1 (the mutant allele is truly expressed). The parameters of these models are chosen to give a wide prior probability when E=1, and a probability mass peaking around 1 when E=0. In particular, in these examples, α0=9999, β0=1, α1=1, and β1=1.

[0151] For the M0 and M1 models, the results are as follows:

[0152]

number

[0153] P(b|d, G g, θ, α, t) are provided by Equation (3), and the values ​​of p(θ|E=0) and p(θ|E=1) are provided by Equation (4). In this example, the genotype G was estimated by obtaining the major and minor copy numbers from the sample's DNA sequence data and using the major copy number approach described above to obtain a genotype estimate. In particular, this approach assumes that the normal population has genotype AA, and the variant population has any genotype with a number of variant alleles (B) between 1 and the major copy number. When there is a copy number gain or loss (when the sum of the major and minor copy numbers does not equal 2), this approach considers two cases: one in which the mutation occurred before the copy number event (e.g., minor copy number = 1, major copy number = 2: Gv = AAB, ABB), and one in which the mutation occurred after the copy number event (Gv = AAB). As a result, the genotype distribution of the variant population is Gv=AAB with a probability of 2 / 3 and Gv=ABB with a probability of 1 / 3.

[0154] Furthermore, the sequencing error rate was set separately for the null model (MO) and the alternative model (M1). For the null model, the sequencing error rate was set to a value of ε0 = 0.001, which assumes that there may be an extremely small probability of variant reads due solely to sequencing errors. For the alternative model, the sequencing error rate was set to ε0 = 0, which assumes that all variant reads are due to true variants and not sequencing errors. Alternatively, a single value may be used for both models.

[0155] It should be noted that although b and d are both observed variables, this approach only models b, since d is treated as a parameter in the model. Thus, a term such as P(b|M0) is equivalent to P(b, d|M0) (similarly, P(b|M1) = P(b, d|M1), P(b|d, G g , θ, α, t)=P(b, d|G g, θ, α, t), P(b, d|α, t, E=1)=P(b|d, α, t, E=1), etc.).

[0156] Edge cases of variant populations with no reference allele This model further defines a(G V )=0, i.e., the genotype G where the reference allele is not present V This is because, for this genotype, (model M i The sequencing noise parameter ε i (using

[0157]

number

[0158]

number

[0159] Power to detect expressed mutations When evaluating RNAseq data to determine whether a mutant allele is expressed in a sample, there are four different scenarios: True positive: The mutant allele is expressed and expression is detected. False positive: The mutant allele is not expressed, but expression is detected (type I error) True negative: the mutant allele is not expressed and no expression is detected False negative: The mutant allele is expressed, but no expression is detected (type II error).

[0160] We want to test each mutation site and determine whether the mutant allele is expressed or not. Given the null model M0, we can calculate the rejection region for the hypothesis test, for example: H0: α0=999, β0=1 H1: α0≠999, β0≠1 That is, b ≥ b c The critical value b that rejects H0 when c (the critical value of the number of reads containing the mutation) can be found. c is calculated by using equation (5) to calculate the distribution of P(b|M0) for the number of variant reads and obtaining the cumulative distribution to obtain P(b≧b c )<0.05. The test statistic b can be used as follows: ·Significance level P(b≧b c |M0) and b ≥ b cIf , reject H0. ·b c Accept H0 if Now assume that the variant alleles are expressed and that the data are actually derived from an alternative model with α1=1 and β1=1. In that case, ·b c If , we accept the null hypothesis. However, if we assume that H1 is actually true, this is a false negative (a type II error). The probability of a false negative is P(b c |M1). b≧b c If , then reject the null hypothesis. Assuming that H1 is in fact true, this is a true positive.

[0161] The probability of a true positive (and the probability of avoiding a false negative) is given by: Power = P(b ≧ b c |M1) (8) Here, power is calculated by calculating the distribution of P(b|M1) for the number of variant reads using equation (5) and obtaining the cumulative distribution to obtain P(b ≥ b c |M1) and determine the significance level P(b ≧ b c |M0) and test b ≥ b c Therefore, using model M0, we can detect b c (P(b≧b c b ≥ b at the selected significance level given by |M0) c (value such that H0 is rejected when c The value of is used to calculate the power to detect a mutation that is expressed when it actually is expressed using model M1.

[0162] ​​​This approach may be used to determine whether a mutation is indeed likely to be expressed (and, if so, whether the low probability of expression is due to low power to detect expression in the available sequence data), to detect mutations from RNA data (e.g., mutations that are detectable only or advantageously from RNA-seq data, examples of which include gene fusions and splice variants such as retained introns, among others), and to determine the RNA-seq depth required to enable detection of a particular mutation with a given minimum power. For example, this approach may be used to determine the power to detect that a mutation is expressed, given multiple possible sequencing depths (resulting in corresponding coverage values ​​d and possible b values ​​depending on the expected ratio of expression of normal and variant alleles, as well as tumor purity and genotype), and to determine the significance level P(b≧b c |M0) and b>b c Depths that satisfy both and have a power above a predetermined value may be selected for use in sequencing samples in which the mutation is suspected to be present or expressed.

[0163] Extension to multi-region sequencing The above approach may be repeated for each of multiple samples from the same subject to determine the power to detect tumor-specific mutations expressed within each sample.

[0164] result The above approach was exemplified using data from a cell line (HCC1395, a human breast cancer cell line) mixed with a reference cell line (a matched normal cell line, a B-lymphocyte-derived cell line, HCC1395L) using known ratios (10, 20, 50, and 100%) to model tumor samples with tumor purities of 10, 20, 50, and 100%. DNA and RNA sequence data were obtained for all samples. The genotype and purity of each sample were estimated using Sequenza (Favero et al., 2015) and set to maximum a posteriori (MAP) estimates. Using the baseline model described above, α = 5, ε = 0.001, and p(θ|M0)~Beta(9999,1) p(θ|M1)~Beta(1,1) where α is the cellular expression ratio, ε is the sequencing error rate, and M0 and M1 are the null and alternative models with different prior probabilities for the reference ratio θ. V and G N ) and tumor purity t were set to their MAP estimates.

[0165] The 100% sample can be used to assess the ground truth number of mutations and whether they are expressed. Two replicate experiments were analyzed in an identical manner.

[0166] For replicate 1: A total of 336 highly reliable mutations in the data were identified in 100% of samples and had at least one variant read in all samples at the DNA level. Of these, 222 (66.07%) had more than one variant RNA-seq read in 100% of samples and are assumed to be truly expressed mutations (true positives). Copy number and tumor purity could be estimated for 333 of the 336 mutations at all dilutions (10, 20, 50, 100). Of these, 221 were expressed (true positives for which copy number and purity could be estimated and used as the true positive set in these analyses) and 112 were not expressed.

[0167] For replicate 2: A total of 333 highly reliable mutations in the data passed AVID in 100% of samples and had at least one variant read in all samples at the DNA level. Of these, 220 (66.07%) had more than one variant RNA-seq read in 100% of samples and are assumed to be truly expressed mutations (true positives). For 332 of the 333 mutations, copy number and tumor purity could be estimated at all dilutions (5, 10, 30, 100). Of these, 220 were expressed and 112 were non-expressing.

[0168] To quantify the influence of tumor purity and genotype in the model when calculating power, the model was run on five different data sets (in each replicate). Fixed genotype, ground truth purity: variant genotype is assumed to be major copy number = 1, minor copy number = 1, and tumor purity is the ground truth (i.e., known % of tumor cells used). Ground truth genotypes and purity: Variant genotypes are "ground truth" using Sequenza to estimate major and minor copy numbers for each variant at 100% purity. Purity is the ground truth value. · Estimated genotype, ground truth purity: The genotype of the variant is derived from the major and minor copy number estimates for each mutation as described above, and the purity is ground truth. Ground truth genotypes, estimated purity: Variant genotypes are "ground truth" using a modified Sequenza to estimate major and minor copy numbers for each variant at 100% purity. Purity is estimated. Fixed genotype, estimated purity: Variant genotype is assumed to be major copy number = 1, minor copy number = 1, and purity is estimated.

[0169] Purity and genotype were jointly estimated using Sequenza (Sequenza is used to determine the probability of a tumor having a particular purity / ploidy value along a grid of possible values ​​and to assign a genotype for each variant given the associated LRR, BAF, and ploidy value). As part of the goal for the various datasets described above was to explore the effect of changing one variable at a time, models were not run with estimated genotype and purity; however, given the results obtained, this is not expected to produce significantly different results.

[0170] In this example, α was set to 0.5 (assuming that tumor and normal cells contribute equal amounts to the total expression at the locus).

[0171] For each mutation for which copy number and purity estimates were available in each series (5, 10, 30, 100), the following were calculated: ·b c : Critical value for testing mutation expression. The number of variant reads at that mutation must be at least b c (b≧b c ), the mutation is considered detected (i.e., expressed) and the null model of no expression is rejected. Alpha: the relevant significance value of the test (probability of a false positive - probability of detecting a mutation when it is not expressed), where alpha < 0.05. As mentioned above, bc is chosen so that alpha < 0.05, but it should be noted that since bc is an integer, the actual value of alpha that meets the criterion may vary (in particular, it is unlikely to be = 0.05). · Power: The power to detect the expressed variant (the probability of rejecting the null hypothesis when the alternative is true).

[0172] For each variant, the number of variant reads is defined as b c Compared to b, whether the variant is detected at this particular alpha (b ≥ b c ) was determined.

[0173] For each sample, we also calculated the false positive (FP) and false negative (FN) rates using ground truth data (i.e., 221 mutations known to be expressed in 100% of samples and 112 mutations known not to be expressed in 100% of samples). When evaluating RNAseq data from a sample to determine whether a variant is expressed, there are four different scenarios: True positive: The mutant allele is expressed and expression is detected. False positive: The mutant allele is not expressed, but expression is detected (type I error) True negative: the mutant allele is not expressed and no expression is detected False negative: The mutant allele is expressed, but no expression is detected (type II error).

[0174] These results are shown in Figure 8 (for replicate 2, data from replicate 1 is similar but not shown for brevity). This figure shows that the FN rate (and to a lesser extent FP) increases as tumor purity decreases. Error rates for all datasets (various genotype and purity sources) were relatively similar, although the estimated genotype + ground truth purity had a slightly higher TP rate.

[0175] The average accuracy (true negative rate - TN, true positive rate - TP) across the 10, 20 and 50 purity series was, for replicate 1, as follows: · Estimated genotype, ground truth purity: TN: 0.9792; TP: 0.7617 Fixed genotype, estimated purity: TN: 0.9792; TP: 0.7602 Fixed genotype, ground truth purity: TN: 0.9792; TP: 0.7647 Ground truth genotype and purity: TN: 0.9792; TP: 0.7647 Ground truth genotype, estimated purity: TN: 0.9792; TP: 0.7632 For iteration 2, the results were as follows: · Estimated genotype, ground truth purity: TN: 0.9583; TP: 0.7561 Fixed genotype, estimated purity: TN: 0.9583; TP: 0.7530 Fixed genotype, ground truth purity: TN: 0.9583; TP: 0.7545 Ground truth genotype and purity: TN: 0.9583; TP: 0.7545 Ground truth genotype, estimated purity: TN: 0.9583, TP: 0.7545.

[0176] TN was very similar across all cases, due to the use of the same false positive rate (<0.05) in all cases. The distribution of estimated power, grouped by different outcomes (TP, FP, FN, TP) for each assessed tumor purity, is shown in Figures 9A-9E for each data set (repeat 2). Data from replicate 1 are similar and are not shown for brevity. This shows that for truly expressed mutations (FN+TP), the power to detect false negatives (FN) is lower than the power to detect true positives (TP). This indicates that the method can identify cases where the data provide low power for mutation detection (hence, lack of detection does not necessarily mean lack of expression).

[0177] Mutations were then grouped by power value within 10% bins of the power scale, and the true positive and false negative rates were calculated for each bin. The results are shown in Figure 10 (replicate 2). Data from replicate 1 are similar and are not shown for brevity. This shows that, consistent with expectations, the TP rate correlates with the power estimated by the model. The solid line is a straight line y = x (as expected for the model). This indicates that in some cases, the model may underestimate the true positive rate (since the true TP rate is actually slightly above this line, at least for power bins below 60%). The mean squared error (MSE) and mean error (ME) of the power estimates are shown in Table 1.

[0178] [Table 1]

[0179] Example 2 - Estimation of the Alpha Parameter In this section, we will use the parameter α (the proportion of total expression at a locus that is due to a normal cell population (e.g., TPM)) used in Example 1. 正常 / (TPM 正常 +TPM 腫瘍This example introduces an approach to estimate the value of α. As mentioned in Example 1, the value of the parameter α is typically unknown for a particular sample. In the simplest implementation, it can be set to a suitable default value, such as 0.5. This assumes that the normal and tumor cell populations contribute equally to the total expression signal (e.g., number of reads) at a locus. However, intuitively, there should be a relationship between this parameter and tumor purity. Furthermore, some loci may be overexpressed in tumor cells compared to normal cells. In this example, we introduce a framework for estimating the value of α in situations where multiple samples with varying purities are available. Other approaches can also be used, for example, when differential expression between normal and tumor cells is known to occur and this estimate can be obtained (e.g., from previous data or a database).

[0180] method Relationship between TPM, purity t and α The TPM value of gene i is given as:

[0181]

number

[0182]

number

[0183] The superscripts T and N are used to refer to reads from tumor cells and reads from normal cells, respectively. Assume the following: The sample contains n cells (having (1-t)n normal cells and tn tumor cells), ·Lead on Gene i

[0184]

number

[0185]

number

[0186]

number

[0187]

number

[0188]

number

[0189] The total TPM in a sample with purity t is given by:

[0190]

number

[0191]

number

[0192]

number

[0193] Define the following:

[0194]

number

[0195]

number

[0196]

number

[0197]

number

[0198]

number

[0199]

number

[0200] Therefore, TPM at multiple purities t i By fitting a regression line to the values ​​of α i These can be obtained, for example, from multiple tumor samples from the same patient (e.g., if multi-region sequencing is available), or from multiple samples from multiple patients (such as multiple patients with a particular cancer type that are recorded in a database, an example of which is TCGA (The Cancer Genome Atlas dataset, https: / / portal.gdc.cancer.gov / ).

[0201] In fact, the following formula is obtained:

[0202]

number

[0203]

number

[0204]

number

[0205] result This example uses the same data as in Example 1 (Repeat 2) to demonstrate the use of the method described herein, including the estimation of the α parameter (cellular expression ratio) for each mutation. The baseline model described above was used, assuming a sequencing error rate ε = 0.001 (note that this can generally be set depending on the expected error rate of the sequencing platform used), p(θ|M0) ≒ Beta(9999,1) and p(θ|M1) ≒ Beta(1,1) (prior probabilities for the reference ratio under the null and alternative models). Cellular expression ratios were estimated for each mutation by a linear regression model using purity and gene expression values ​​(transcripts per million, TPM) across the entire titration series, as described above.

[0206] To quantify the influence of tumor purity and genotype on the model when calculating expression power, we ran the model on five different datasets as in Example 1 (i.e., "Fixed genotype, ground truth purity," "Ground truth genotype, ground truth purity," "Estimated genotype, ground truth purity," "Ground truth genotype, estimated purity," and "Fixed genotype, estimated purity"). Only mutations with copy number and purity estimates in each sample of the titration series were used (301 mutations, of which 200 were more than one variant RNA-seq read expressed in 100% of samples).

[0207] For each mutation, the following values ​​were calculated: ·b c: Critical value for testing mutation expression. The number of variant reads at that mutation must be at least b c (b≧b c ), the mutation is considered detected (i.e., expressed) and the null model of no expression is rejected. Alpha: the significance value associated with the test (probability of a false positive - probability of detecting a variant when it is not expressed), where alpha is less than 0.05 as explained above. · Power: The power to detect the expressed variant (the probability of rejecting the null hypothesis when the alternative is true).

[0208] For each mutation, the number of variant reads is c Compared to b, whether the variant is detected at this particular alpha (b ≥ b c ) was determined.

[0209] For each sample, the false positive rate (FP) and false negative rate (FN) were also calculated. When evaluating RNAseq data from a sample to determine whether a variant is expressed, there are four different scenarios: True positive: The mutant allele is expressed and expression is detected. False positive: The mutant allele is not expressed, but expression is detected (type I error) True negative: the mutant allele is not expressed and no expression is detected False negative: The mutant allele is expressed, but expression is not detected (type II error) The results are shown in Figure 11 and Table 2. As expected, both types of errors increased as purity decreased. Error rates were similar across all datasets, with slightly higher TP rates for ground truth genotypes and estimated purity. Mutation power provides an indication of whether we can confidently predict whether a mutant allele will be expressed. The distribution of power estimated from the model, grouped by different outcomes (TP, FN, FP, TP) and separated by tumor purity, is shown in Figures 12A-12E (for each of the five datasets). This indicates that for truly expressed mutations, the power to detect false negatives is lower than the power to detect true positives.

[0210] [Table 2]

[0211] The mutations were then grouped by power value (using 10% bins). For each bin, the true positive and true negative rates were calculated. The results are shown in Figure 13, which shows that the TP rate correlates with the power estimated by the model, as expected. The mean squared error was calculated based on the TP rate and the average estimated predictive power per bin. · Estimated genotype, ground truth purity: mse=0.0293, me=0.1278 Fixed genotype, estimated purity: mse=0.0646, me=0.2158 Fixed genotype, ground truth purity: mse=0.0334, me=0.1476 Ground truth genotype and purity: mse=0.0098, me=0.0754 Ground truth genotype, estimated purity: mse=0.0396, me=0.1611 This data shows that estimating the α parameter improves the correlation between the estimated power and the TP rate. Therefore, when samples of varying purity are available, estimating α using linear regression can improve the accuracy of the model for detecting allele-specific expression.

[0212] Example 3 - Probability of Mutations Occurring In this section, we introduce a unified model for estimating the detection power of expressed mutations and determining whether a mutation is expressed or not, using the statistical hypothesis framework introduced above. Example 1 describes a variant expression model and a non-variant expression model (illustrated in Figure 4) using a binomial likelihood for the number of variant reads at a site with coverage d (Equation (3)). Example 2 illustrates the estimation process of the α parameter (cellular expression ratio) in the model of Example 1. Based on the framework of Example 1, this example presents models for determining the probability that a mutation is expressed or not for a single sample (e.g., a single tumor region) and multiple samples (e.g., multiple tumor regions), respectively, as shown in Figures 5A and 5B. This approach advantageously combines the information in the model of Example 1 with prior information about expressed mutations within a Bayesian framework. Furthermore, this approach provides a single, interpretable probability that a tumor-specific mutation will be expressed. This can be used, for example, to rank or otherwise prioritize tumor-specific mutations for further analysis, for use as therapeutic targets, for use as diagnostic markers, etc. Because this is within the context of a Bayesian framework, the probabilities advantageously combine information from available data (RNA-seq data) in the form of likelihoods of the data under various models (variant expression / non-expression), as well as information about prior beliefs about whether the mutation will be expressed (hence the reference to a "unified" model that combines the likelihood of expression with prior probability).

[0213] method Probability measure of whether a variant will be expressed Denoting the expression of the variant allele by E (E = 1 means the variant is truly expressed, E = 0 means the variant is not expressed), the prior probability on θ (reference ratio = reference read count / total read count, the balance between reference allele expression and total allele expression) is E (Equation (4)), and the likelihood of the model is:

[0214]

number

[0215] Let E be a Bernoulli variable with parameter ρ, where p(E=1)=ρ and ρ(E=0)=(1-ρ). This is the prior probability of E, i.e., the prior probability that the mutation will truly manifest (E=1) or not manifest at all (E=0). Also, by setting a hyperprior probability for ρ, we obtain the following equation:

[0216]

number

[0217] Based on the sequence data from the sample, it is possible to calculate the joint posterior distribution for E and ρ.

[0218]

number

[0219]

number

[0220]

number

[0221]

number

[0222] To further enhance the intuitive interpretation of the model, equation (13) can be rewritten as equation (14), such that the posterior distribution for E=1 is given by a logistic function whose arguments are the log-likelihood ratios of the two outcomes of E and the log-prior odds (as shown in Figure 6).

[0223]

number

[0224] therefore, In cases where the power to detect the variant allele is high and the likelihood supports the absence of the mutant allele, c and there is strong evidence in favor of M0, and r becomes very small. As r approaches 0 (r → 0), P(E=1|b, d, a, t) = 0. In cases where the power to detect the variant allele is high and the likelihood supports the expression of the variant allele, b ≥ b c and there is strong evidence in favor of M1, and r becomes very large. As r approaches infinity (r→∞), P(E=1|b, d, a, t)=1. In cases where the power to detect the variant allele is low and the likelihood favors the absence of the mutant allele, c and since the data are equally likely under both M0 and M1, r approaches 1 (r=1), and P(E=1|b, d, a, t)=μ ρ becomes. In cases where the power to detect the variant allele is low and the likelihood supports the expression of the variant allele, b ≥ b c ​​and there can exist values ​​of b where r is very high, and as r approaches (r → ∞), P(E=1|b, d, a, t) = 1. This is the desired behavior because, when there is strong evidence that a variant will be expressed, we want the probability to reflect high confidence that the variant allele will be expressed (E=1), even if the power to detect the variant is low.

[0225] If there is no coverage for a mutation, i.e., d≈0, the following may be the cause: The biological cause (e.g., the gene carrying the mutation) is downregulated, making the mutation unlikely to be expressed. This effectively creates a scenario of missing data and therefore missing likelihood. In Bayesian models, we revert to prior probabilities (i.e., the posterior probability equals the prior probability). This does not make intuitive sense in the current context, because without formal parameters to model the gene expression process, one would expect low / absent gene expression to mean that the variant allele is unlikely to be expressed (i.e., d approaching 0 (i.e., missing data) is informative as to whether the variant allele (and reference allele) is likely to be expressed). Technical issues such that there is no data and it is not possible to infer whether a mutation has or has not been expressed.

[0226] To resolve this, prior knowledge of whether a gene is expressed or not (reference data, e.g., normal tissue, tumor tissue, tumors of the same type, etc.) can be used to determine whether the absence of reads for a locus is a technical issue. If a gene is known to be highly expressed in a sample and there is no read coverage at a particular position for a variant, then the absence of reads is likely a technical issue, and μ p = 0.5. If a gene is known to be lowly expressed and there is no read coverage, μ ρIt could potentially be set to <0.5 (because if the gene is not supposed to be expressed, the mutation is unlikely to be expressed either).

[0227] A simple solution to this problem is to set a gene expression threshold (e.g., TPM>1) for genes with variants. TPM is the number of reads mapped to a gene divided by a scaling factor equal to the total number of reads in the sample divided by 1,000,000, normalized by the length of the gene. If a gene does not meet this condition in the analyzed samples (i.e., TPM=0 or 1), μ ρ It will be <0.5.

[0228] If the gene exceeds this, μ ρ can be set to 0.5 or greater. The threshold used can be set to any predetermined value, but is preferably set to a low value to reflect genes with low expression. Alternatively, using reference expression data (e.g., expression data from a cohort of samples), the threshold can be set to a value at which the model performs well in detecting genes with low expression, or to a value that reflects the TPM value below which genes are unlikely to be expressed. When multiple samples are used, the condition can be tested for each sample individually, or the results for the multiple samples can be combined to determine whether a gene satisfies the condition. For example, it may be necessary to test the condition for each sample or for a majority of samples.

[0229] In the results shown below, the prior probability is μ ρ = 0.5. In such cases, the prior probability is set to a lower value of μ ρ =0.05.

[0230] It should be noted that this is a significant adaptation of the "traditional" Bayesian approach to the specific biological problem under investigation. Indeed, in a traditional Bayesian framework in the absence of data, the posterior probability should equal the prior probability (there is no likelihood to inform a decision that deviates from the prior probability). Therefore, in the absence of data, it is typically the desired and expected behavior of a Bayesian framework that the probability of a gene being expressed should revert to, say, a probability of p=0.5 (i.e., a 50-50 chance of expression vs. no expression). However, the inventors identified that in the current situation, the lack of data (no reads at the locus under investigation) may itself be informative, since it may be due to a lack of gene expression at the locus. In the absence of gene expression at a locus, the likelihood that the variant is not expressed is rather high. Therefore, the inventors adapted the naive Bayesian framework with a single prior probability to instead include a prior probability with multiple possible values. This multiple possible values ​​includes at least a first value when there is evidence that the locus is expressed and a second value when there is insufficient evidence that the locus is expressed.

[0231] Extension to multi-region sequencing For multi-region sequencing, when there is sequencing data from multiple regions of the tumor, the model can be extended to use multiple information sources. There are S regions, denoted by the subscript s. th Let s denote the th region. For a locus, for each of the S regions, the likelihood of expression / non-expression is given by:

[0232]

number

[0233]

number

[0234]

number

[0235] result The model described in the Methods section can be used with the default value for α, or with an estimated value. The latter is shown here. The same data as in Example 2 were used, including the same parameters for the sequencing error rate ε and the prior probability of θ under the null and alternative models. The posterior probability of each mutation being expressed was calculated as described in the Methods section above for the data described in Example 1 (repeat 2). The distributions of these probabilities, split by ground truth purity (5, 10, 30%) and ground truth representation state (represented at bottom or not represented at top), are shown in Figure 14. These were compared to simpler models in which the probability of a mutation being expressed is given by its variant allele frequency (i.e., the proportion of reads containing the variant, b / d). The distributions of probabilities for these simpler VAF models are shown in Figure 15.

[0236] Figure 14 shows that, as expected, our method defaults variants with total read count d=0 to the prior probability (in this case, μ if the gene is expected to be expressed (i.e., TPM>1)). ρ = 0.5, if the gene is expected to be lowly expressed (TPM ≤ 1), μ ρ= 0.05, where the expected expression value was determined based on the calculated TPM value of each data set in the purity series. For mutations that did not have a total read count d = 0, the prior probability was set to 0.5 for all mutations. Mutations that were never expressed (top row) were assigned very low probabilities (see the peak near 0), truly expressed mutations (bottom row) were mostly assigned high probabilities (see the peak near 1), and there were very few false negative mutations with read counts above 0 and probabilities near the lower end of the scale.

[0237] In contrast, Figure 15 shows that the "probabilities" derived from the VAF are spread over a much wider range for mutations that actually express. Because many mutations have very low VAFs, there is no natural threshold for identifying whether a mutation is expressed or not. Therefore, using this model will result in either very high false positives or very high false negatives. In other words, comparing Figures 14 and 15 shows that the methods described herein can integrate data within a biologically based model to boost confidence in expression vs. non-expression from a priori belief to a confident call for expression vs. non-expression, unless the data honestly does not provide enough information to make such a call.

[0238] The probability of expression of each mutation at each of three different purities (5, 10, and 30%) was calculated and used to quantify several performance metrics. ROC curve: TP and FP rates using different thresholds to classify variants as expressing / non-expressing: this shows the impact of the decision threshold on the TP and FP rates and makes it possible to calculate the AUC (area under the ROC curve) and the threshold that gives the best TP rate with an FP rate below 0.05. AUC: This represents the probability that a random positive example (a truly expressed mutation) will be ranked higher (more likely to be expressed) than a random negative example (a mutation that does not express at all). Calibration curve: AUC and ROC indicate the ability of the model to distinguish between positive and negative cases, but we want to assess whether the probabilities used to predict whether a variant will be expressed are correctly calibrated in that they reflect the true likelihood, i.e., whether 50% of all variants assigned a 50% probability of being expressed will be expressed. Model calibration was tested by binning the data by probability (10% bins) and calculating, for each bin, the average probability of all variants in the bin, as well as the proportion of positives. Brier score: Brier score is used to evaluate the accuracy of probability prediction, and is calculated as mean(true label l (0 or 1) - probability 2 ) is given by Prediction mean squared error: Using the calibration curve, we calculated the mean squared error across bins, i.e., mean((proportion of TPs - mean probability) 2 ).

[0239] These metrics were calculated for the model described herein and for a simpler VAF model that uses VAF as the probability that a mutation is expressed. First, we calculated them by removing edge cases where the total read depth was 0. The ROC curves for the current VAF model and the simpler VAF model are shown in Figures 16A and 16B, respectively. The cutoff values ​​were chosen to maximize the TPR while keeping the FPR below 0.05. These results show that the simpler VAF model has a slightly higher AUC, but on the left side of the curve, the model outperforms the VAF model. On the left side of the curve, the TPR of the simpler VAF model drops off as the FPR drops below approximately 9%, but the model's performance remains good throughout. This is due to the misclassification of expressed low-frequency mutations. The model achieved a TPR of 75% and an FPR of 5% on this dataset, indicating that the model provides highly informative predictions (any AUC above 50% is better than random guessing). As discussed further below, this model can likely be improved by further calibration, particularly by adapting preselection. While models using the VAF (proportion of reads containing a variant) may appear informative in the specific cases presented, such models have serious limitations. For example, they provide estimates that cannot be compared across samples, making it impossible to make reliable, reproducible, and verifiable decisions for the same patient. Indeed, a VAF of 10% in a 5%-purity sample strongly suggests that the variant will be expressed, whereas a VAF of 10% in a 90%-purity sample is much more likely to be a false positive due to, for example, sequencing error. Additional factors, such as genotype, may also affect whether the same VAF is considered reliable or unreliable. In other words, the use of VAF as an indicator of the likelihood of expression does not take into account the effects of genotype and tumor purity, resulting in an unreliable indicator of whether the variant will be expressed.

[0240] Figures 17A and 17B show the calibration curves for our model and the simpler VAF model, respectively. They show that both models are poorly calibrated, resulting in lower predicted probability estimates and underestimating the true probabilities. However, the Brier scores for our model are relatively good. Additional model calibration may result in further improvement.

[0241] Because the VAF is not defined for the absence of reads, including edge cases (total read depth = 0), the simpler VAF model cannot be used. Therefore, for any mutation with d = 0, we fit our model by using a gene expression threshold (TPM > 1) and setting p = 0.05 below this threshold and p = 0.5 otherwise. ROC curves are shown for our model and the simpler VAF model in Figures 18A and 18B, respectively, and calibration curves are shown for our model and the VAF model in Figures 19A and 19B, respectively. As can be seen from Figures 18B and 19B, the VAF is not a properly calibrated probability, and therefore cannot integrate prior assumptions (used in edge cases to estimate expression probabilities), resulting in poor performance for the simpler VAF model.

[0242] In contrast, with edge cases, the calibration curve of our model is slightly better than without edge cases (see Figure 18A), with a Brier score of 18.8%. The prediction at 50% overestimates the true probability of expression (see Figure 19A). The majority of mutations assigned to 50% are This suggests that further calibration of this dataset may improve the probability of expression (for edge cases).

[0243] This example introduces and demonstrates the use of an approach to estimate the probability of a variant being expressed based on RNA-seq data. It uses a probabilistic framework incorporating estimates of tumor purity, genotype, and ploidy to model the likelihood of detecting a variant read, taking into account variable expression at both the allele level (allele-specific expression) and the cell level (cell-specific expression). Cell-specific expression α was identified to play a key role in the accuracy of the power estimation output by the model, and a method for estimating this, when data are available to do so, was also introduced and demonstrated. Finally, the performance of the approach was evaluated using cell line data, demonstrating good classification performance (accuracy in classifying expressed vs. non-expressed mutations).

[0244] References Gartner, Jared J., et al. 2021. “A Machine Learning Model for Ranking Candidate HLA Class I Neoantigens Based on Known Neoepitopes from Multiple Human Tumor Types.” Nature Cancer 2(5): 563-74

[0245] Carter SL, Cibulskis K, Helman E, McKenna A, Shen H, Zack T, Laird PW, Onofrio RC, Winckler W, Weir BA, Beroukhim R, Pellman D, Levine DA, Lander ES, Meyerson M, Getz G. Absolute quantification of somatic DNA alterations in human cancer. Nat Biotechnol. 2012 May;30(5):413-21.

[0246] Vanessa Jurtz、Sinu Paul、Massimo Andreatta、Paolo Marcatili、Bjoern Peters and Morten Nielsen. NetMHCpan-4.0: Improved Peptide-MHC Class I Interaction Predictions Integrating Eluted Ligand and Peptide Binding Affinity Data. J Immunol November 1、2017、199(9) 3360-3368.

[0247] Langmead、B.、Trapnell、C.、Pop、M. et al. Ultrafast and memory-efficient alignment of short DNA sequences to the human genome. Genome Biol 10、R25(2009).

[0248] Lundegaard C、Lamberth K、Harndahl M、Buus S、Lund O、Nielsen M. NetMHC-3.0: accurate web accessible predictions of human、mouse and monkey MHC class I affinities for peptides of length 8-11. Nucleic Acids Res. 2008 Jul 1;36(Web Server issue):W509-12.

[0249] McGranahan, N. Furness, AJ, Rosenthal, R. Ramskov, S. Lyngaa, R. Saini, SK, Jamal-Hanjani, M. Wilson, GA, and Birkbak NJ, Hiley, CT, Watkins, TB, Shafi, S., Murugaesu, N., Mitter, R., Akarca, AU, Linares, J., Marafioti, T., Henry, JY, Van Allen, EM, Miao, D.,... Swanton, C. (2016). Clonal neoantigens elicit T cell immunoreactivity and sensitivity to immune checkpoint blockade. Science(New York, NY), 351(6280), 1463-1469.

[0250] Landau DA, Carter SL, Stojanov P, McKenna A, Stevenson K, Lawrence MS, Sougnez C, Stewart C, Sivachenko A, Wang L, Wan Y, Zhang W, Shukla SA, Vartanov A, Fernandes SM, Saksena, G, Cibulskis, K, Tesar, B, Gabriel, S, Hacohen, N, Meyerson, M, Lander, ES, Neuberg, D, Brown, JR, Getz, G, Wu, CJ. Evolution and impact of subclonal mutations in chronic lymphocytic leukemia. Cell. 2013 Feb 14;152(4):714-26. doi: 10.1016 / j. cell.2013.01.019.

[0251] Raine KM、Van Loo P、Wedge DC、Jones D、Menzies A、Butler AP、Teague JW、Tarpey P、Nik-Zainal S、Campbell PJ. ascatNgs: Identifying Somatically Acquired Copy-Number Alterations from Whole-Genome Sequencing Data. Curr Protoc Bioinformatics. 2016 Dec 8;56:15.9.1-15.9.17. doi: 10.1002 / cpbi. 17.

[0252] Heemskerk B、Kvistborg P、Schumacher TNM. The cancer antigenome. The EMBO Journal Vol. 32、No. 2、2013.

[0253] Castel、SE、Levy-Moonshine A、Mohammadi P、Banks E、Lappalainen T. Tools and best practice for data processing in allelic expression analysis. Genome Biology(2015) 16:195.

[0254] Favero F、Joshi T、Marquard AM、Birkbak NJ、Krzystanek M、Li Q、Szallasi Z、Eklund AC. Sequenza: allele-specific copy number and mutation profiles from tumor sequencing data. Ann Oncol. 2015 Jan;26(1):64-70.

[0255] All references cited herein are incorporated by reference in their entirety for all purposes to the same extent as if each individual publication or patent or patent application was specifically and individually indicated to be incorporated by reference in its entirety.

[0256] The specific embodiments described herein are provided by way of example and not by way of limitation. Various modifications and variations of the described compositions, methods, and applications of the present technology will be apparent to those skilled in the art without departing from the scope and spirit of the present technology as described. Any subheadings are provided herein for convenience only and should not be construed as limiting the present disclosure in any way. Unless otherwise indicated by context, the descriptions and definitions of features presented above are not limited to any particular aspect or embodiment of the present invention, but apply equally to all aspects and embodiments described.

[0257] The methods of any embodiment described herein may be provided as a computer program or as a computer program product or computer readable medium carrying a computer program configured to perform the method(s) when run on a computer.

[0258] Throughout this specification and claims, the following terms take the meanings expressly associated therewith, unless the context clearly dictates otherwise. As used herein, the phrase "in one embodiment" does not necessarily refer to the same embodiment, although it may. Additionally, as used herein, the phrase "in another embodiment" does not necessarily refer to different embodiments, although it may. Thus, as described below, various embodiments of the invention may be readily combined without departing from the scope or spirit of the invention.

[0259] It should be noted that as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Ranges may be expressed herein as from "about" one particular value and / or to "about" another particular value. When such a range is expressed, another embodiment includes from the one particular value and / or to the other particular value. Similarly, by use of the antecedent "about," when values ​​are expressed as approximations, it will be understood that the particular value forms another embodiment. The term "about" with respect to numerical values ​​is optional and means, for example, ±10%.

[0260] Throughout this specification, including the claims which follow, unless the context requires otherwise, the words "include" and "comprise", and variations such as "comprises", "comprising" and "including", will be understood to mean the inclusion of a stated integer or step or group of integers or steps, but not the exclusion of any other integer or step or group of integers or steps.

[0261] Other aspects and embodiments of the present invention provide those aspects and embodiments described above, with the term "comprising" replaced by the term "consisting of" or "consisting essentially of," unless the context dictates otherwise.

[0262] As used herein, "and / or" should be construed as a specific disclosure of each of the two specifically stated features or components, with or without the other. For example, "A and / or B" should be construed as a specific disclosure of (i) A, (ii) B, and (iii) each of A and B, as if each were individually set forth herein.

[0263] The features disclosed in the above description, or the following claims, or the accompanying drawings are suitably expressed in their specific form, or in terms of means for performing a disclosed function, or in terms of a method or process for obtaining a disclosed result, and such features may be utilized individually or in any combination to realize the invention in various of its forms.

Claims

1. 1. A method for determining whether a subject is likely to develop a tumor-specific mutation, comprising: obtaining RNA-seq data from one or more samples from the subject containing tumor genetic material, the RNA-seq data including, for each of the one or more samples, the number of RNA reads in the sample that exhibit the tumor-specific mutation (b) and the total number of RNA reads at the location of the tumor-specific mutation (d); determining the likelihood of the sequence data if the tumor-specific mutation (i) is expressed and (ii) is not expressed; and comparing the obtained likelihoods to thereby determine whether the tumor-specific mutation is likely to be expressed in the subject.

2. Comparing the likelihoods includes: (a) determining the posterior probability of the tumor-specific mutation being expressed by The prior probability that the mutation will occur (μ ρ )and, the likelihood of the sequence data if the tumor-specific mutation (i) is expressed and (ii) is not expressed; and and / or (b) determining the power to detect whether the tumor-specific mutation is expressed at a predetermined false positive rate, The power to detect whether the tumor-specific mutation is expressed is determined by a threshold number of reads (b) where the tumor-specific mutation is expressed as a function of the number of RNA reads (b) in the sample that exhibit the tumor-specific mutation. c ) is the area under the curve of the likelihood of the number of reads exhibiting the tumor-specific mutation exceeding The threshold number of reads (b c ) is the number of reads such that the area under the curve of the likelihood of the number of reads exhibiting the tumor-specific mutation exceeding the threshold number of reads, when the tumor-specific mutation is not expressed as a function of the number of RNA reads in the sample exhibiting the tumor-specific mutation (b), is equal to the predetermined false positive rate; The method of claim 1 , comprising the determining step.

3. 10. The method of any preceding claim, wherein the likelihood of the sequence data is the probability of the sequence data given the tumor fraction (t) of each of the one or more samples and the fraction (α) of the total number of reads at the position of the tumor-specific mutation that originate from a cell population within the one or more samples that does not contain the tumor-specific mutation.

4. Each sample was genotyped (G v ) and a first cell population having a genotype (G N 10. The method of claim 9, wherein the first cell population represents a proportion equal to the tumor fraction (t) for each of the samples.

5. 10. The method of any preceding claim, wherein the tumor-specific mutation is assumed to be ubiquitous within the one or more samples and / or the tumor-specific mutation is a clonal mutation or a mutation assumed to be clonal in the tumor of the subject.

6. 10. The method of any preceding claim, wherein the RNA-seq data comprises RNA-seq data from a plurality of samples from the subject, optionally from a plurality of tumor samples from the subject, and wherein the posterior probability that the tumor-specific mutation is expressed is a posterior probability that the tumor-specific mutation is ubiquitously expressed in the subject's tumor.

7. 10. The method of any preceding claim, wherein the likelihood of the sequence data that the tumor-specific mutation is (i) expressed and (ii) not expressed depends on the fraction (α) of total expression of genes containing the tumor-specific mutation that originates from cell populations within the one or more samples that do not contain the tumor-specific mutation, where α is set to a predetermined default value or α is estimated from data comprising RNA-seq data of multiple samples with different tumor fractions.

8. 8. The method of claim 7, wherein if the genotype of a cell population within the one or more samples having a genotype comprising at least one copy of the tumor-specific mutation does not have any non-mutated alleles at the locus of the tumor-specific mutation, then α is set to 1 when determining the likelihood of the sequence data that the tumor-specific mutation is not expressed.

9. a is estimated from data comprising RNA-seq data for a plurality of samples with different tumor fractions, where a is derived from the slope (g) of a linear regression model fitted to the total expression of genes comprising said tumor-specific mutations in a plurality of samples as a function of their tumor purity, and optionally a is provided by: [Equation 1] and TPM T and TPM N are the total expression values ​​of the gene for the tumor sample and the normal sample, respectively, and optionally, TPM T and TPM N 9. The method of claim 7 or claim 8, wherein t is obtained using a regression model fitted to the total expression values ​​(TPM) of genes of multiple tumor samples having at least two different tumor fractions (t) as a function of tumor fraction.

10. The prior probability that the tumor-specific mutation will be expressed is determined by the mean (μ) of the beta distribution (p(ρ)) for the parameter ρ of the Bernoulli distribution for the variable (E) that reflects whether the tumor-specific mutation will be expressed. ρ ), and μ ρ The method according to any one of claims 2 to 9, wherein is set to a predetermined value.

11. μ when the total number of RNA reads at the position of the tumor-specific mutation is equal to or exceeds a predetermined threshold. ρ and setting μ to a first predetermined value, and optionally, the predetermined threshold is 0 (d>0), and / or μ is set depending on whether the gene containing the tumor-specific mutation is expected to be expressed when the total number of RNA reads at the position of the tumor-specific mutation is below a predetermined threshold. ρ 11. The method of claim 10, wherein ρ is set to a first predetermined value or a second predetermined value, and optionally the gene comprising the tumor-specific mutation is considered to be predicted to be expressed if the total expression of genes in the sample is above a predetermined threshold, and is considered to be predicted to be not expressed otherwise, and / or optionally the first predetermined value is 0.5 and / or the second predetermined value is less than 0.

5.

12. 10. The method of any preceding claim, wherein the likelihood is conditional on the tumor purity of the one or more samples, and optionally the tumor purity of the one or more samples is estimated using genomic sequence data from the one or more samples.

13. 13. The method of any of claims 4 to 12, wherein the probability of observing the sequence data depends on the genotypes of the first cell population and the second cell population, the total number of reads at the locus of the tumor-specific mutation, the tumor fraction (t), the proportion of total expression for the gene containing the tumor-specific mutation originating from the second cell population (α), and the ratio of expression of the allele not containing the tumor-specific mutation to total expression at the locus of the tumor-specific mutation (θ).

14. 14. The method of claim 13, wherein the ratio of expression of the allele that does not contain the tumor-specific mutation to the total expression at the locus of the tumor-specific mutation (θ) is assumed to be a random variable having a first distribution when the tumor-specific mutation is expressed and a second distribution when the tumor-specific mutation is not expressed, and / or the likelihood is a marginal likelihood obtained from the probability of observing the sequence data by integrating the ratio of expression of the allele that does not contain the tumor-specific mutation to the total expression at the locus of the tumor-specific mutation (θ).

15. The second distribution has a parameter α 0 , β 0 and the first distribution is a beta distribution having a parameter α 1 , β 1 and optionally, α 0 >1 (e.g., α 0 = 9999), β 0 = 1, α 1 = 1, and β 1 15. The method of claim 14, wherein: =1.

16. The genotypes (G v , G N 16. The method of any of claims 13 to 15, wherein the probability of observing the sequence data conditional on G, the total number of reads at the locus of the tumor-specific mutation, the tumor fraction (t), the proportion of total expression of the gene containing the tumor-specific mutation originating from the second cell population (α), and the ratio of expression of the allele not containing the tumor-specific mutation to total expression at the locus of the tumor-specific mutation (θ) is assumed to follow a binomial distribution with parameters ξ(G, θ, α, t) representing the probability of sampling a read containing the tumor-specific mutation from a tumor sample.

17. ξ(G, θ, α, t) is given by one of equations (2), (2′), and (2″), and optionally: [Equation 2] and c(G v ) and c(G N ) are the total number of copies of the locus containing the tumor-specific mutation in the first cell population and the second cell population, respectively; ε is the sequencing error rate; and μ(G N , θ, ε) is the probability of sampling a read containing the tumor-specific mutation from the second cell population, and μ(G V 17. The method of claim 16, wherein θ, ε) is the probability of sampling a read containing the tumor-specific mutation from the first cell population.

18. μ (G N , θ, ε) and μ(G V , θ, ε) is given by equation (1), and G = G N and G = G v 18. The method of claim 17, wherein a(G) is the genotype of the second cell population and the first cell population, respectively, and b(G) is the number of copies of the allele containing the tumor-specific mutation in cells of each of the cell populations.

19. The posterior probability depends on the ratio (r) of the likelihood of the sequence data when the tumor-specific mutation (i) is expressed (P(b, d|α, t, E=1)) and (ii) is not expressed (P(b, d|α, t, E=0)), and / or the posterior probability is given by Equation (13), Equation (14), or Equation (14′), where μ ρ The method of any one of claims 2 to 18, wherein is the prior probability that the mutation will occur.

20. moreover, repeating the method for a plurality of tumor-specific mutations identified in the subject, and optionally ranking or otherwise prioritizing the plurality of tumor-specific mutations based at least in part on the probability that they are determined to be expressed in the subject; and / or identifying one or more tumor-specific mutations in said subject; and / or 10. The method of any preceding claim, comprising sequencing one or more samples from the subject containing tumor genetic material to obtain RNA sequence reads and optionally DNA sequence reads.

21. 1. A method for identifying one or more neoantigens in a subject, comprising: identifying a plurality of tumor-specific mutations in the subject; and Determining whether one or more of the tumor-specific mutations are likely to be expressed in the subject's tumor using the method of any one of claims 1 to 21; determining whether one or more of the tumor-specific mutations are likely to give rise to neo-antigens, wherein the neo-antigens are tumor-specific mutations that meet one or more predetermined criteria as to whether the tumor-specific mutations are likely to be expressed in the tumor.

22. 1. A method of providing immunotherapy to a subject diagnosed with cancer, comprising: identifying one or more tumor-specific mutations in the subject; and Determining whether one or more of the tumor-specific mutations are likely to be expressed in the subject's tumor using the method of any one of claims 1 to 21; selecting one or more tumor-specific mutations from the identified tumor-specific mutations based on the results of said determining and optionally one or more further criteria; and designing an immunotherapy that targets one or more neo-antigens derived from the selected one or more tumor-specific mutations.

23. 1. A method of treating a subject diagnosed with cancer, comprising: identifying a plurality of tumor-specific mutations in the subject; and determining whether the subject is likely to express one or more of the tumor-specific mutations; selecting one or more of the tumor-specific mutations as candidate neo-antigens, wherein the candidate neo-antigens are tumor-specific mutations that meet at least one or more predetermined criteria regarding whether the tumor-specific mutation is likely to be expressed; treating the subject with an immunotherapy targeting one or more of the selected candidate neo-antigens; identifying one or more neo-antigens by Determining whether a subject is likely to develop a tumor-specific mutation includes: obtaining, by a processor, RNA-seq data from one or more samples from the subject containing tumor genetic material, the RNA-seq data including, for each of the one or more samples, a number of reads in the sample that exhibit the tumor-specific mutation (b) and a total number of reads at the position of the tumor-specific mutation (d); determining, by the processor, the probability of observing the RNA-seq data when the tumor-specific mutation (i) is expressed and (ii) is not expressed.

24. A system comprising a processor and a computer readable medium, the computer readable medium comprising instructions that, when executed by the processor, cause the processor to perform the steps of any method described herein, such as the method of any of claims 1 to 22.

25. One or more non-transitory computer-readable media comprising instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of any method described herein, such as the method of any of claims 1-22.