Technologies for neoantigen prioritization based on tumor modeling and clonality states

By classifying tumor mutations into discrete clone types and prioritizing biologically clonal and prevalent mutations, the method addresses the ineffectiveness of current cancer therapies by targeting high-quality neoantigens for personalized cancer vaccines and T-cell therapies.

WO2026159260A2PCT designated stage Publication Date: 2026-07-30BIONTECH SE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BIONTECH SE
Filing Date
2026-01-23
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Current cancer therapies are often ineffective due to the molecular heterogeneity of tumors, failing to target high-quality neoantigen epitopes that could be effective in eliminating the entire tumor mass.

Method used

A method for identifying and prioritizing biologically clonal and highly prevalent tumor cell mutations for use in personalized cancer vaccines and T-cell therapies by classifying mutations into discrete clone types based on sequencing data, excluding low-prevalence subclones.

Benefits of technology

This approach allows for the selection of effective neoepitopes for immunotherapy treatments, ensuring that biologically clonal and prevalent mutations are targeted while avoiding less effective subclones, thereby enhancing treatment efficacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000008_0001
    Figure IMGF000008_0001
  • Figure IMGF000022_0001
    Figure IMGF000022_0001
  • Figure IMGF000094_0001
    Figure IMGF000094_0001
Patent Text Reader

Abstract

Presented herein are methods and systems that allow for identification and selection of tumor cell mutations that represent / correspond to high quality neoantigen epitope targets for cancer immunotherapy, such as for inclusion in personalized cancer vaccines (PC Vs), use in T-cell therapies, T-cell receptor (TCR)-based therapies, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No. 2013237-1502TECHNOLOGIES FOR NEOANTIGEN PRIORITIZATION BASED ON TUMOR MODELING AND CLONALITY STATES CROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to and benefit of U. S. Provisional Application No. 63 / 749,339, filed on January 24, 2025, the content of which is hereby incorporated by reference herein in its entirety.BACKGROUND

[0002] Cancer is a primary cause of mortality, accounting for 1 in 4 of all deaths. Despite recent advances in the field of cancer immunotherapy there remains no single, broadly applicable treatment. Molecular heterogeneity of tumors renders many therapies ineffective for cancer patients.SUMMARY

[0003] Presented herein are methods and systems that allow for identification and selection of tumor cell mutations that represent / correspond to high quality - i.e., likely to be effective - neoantigen epitope targets for cancer immunotherapy, such as for inclusion in personalized cancer vaccines (PCVs), use in T-cell therapies, T-cell receptor (TCR)-based therapies, etc. Among other things, the present disclosure provides clone type classification technologies that operate on sequencing data (e.g., exome sequencing data; e.g., whole genome sequencing data) obtained from a tumor sample, detect candidate mutations and classify them according to clone types that group mutations according to, for example, their likelihood of being biologically clonal (e.g., based on statistical hypothesis testing, probability calculations, etc.) and / or estimated prevalence in a sample. In certain embodiments, additionally or alternatively, the present disclosure provides approaches for using identified clone types to improve target selection e.g., via prioritizing certain mutations) for immunotherapies. Also described herein are technologies for leveraging tumor modeling to screen rare subclones, which are unlikely to provide valuable targets.

[0004] Among other things, the accurate clone type assignments that are made possible via the methods and systems described herein allow for neoepitopes based on particular cancer mutations to be selected for inclusion in immunotherapy treatments in a way that allows biologically clonal and / or highly prevalent mutation targets to be identified and prioritized, while avoiding low prevalence subclones that are only present in a small fraction of cancer cells (e.g., less than 50%; e.g., less than 40%) and, accordingly, are unlikely to be - 1 - 13241940vlAttorney Docket No. 2013237-1502effective in eliminating an entire tumor mass. In certain embodiments, rare subclones appearing in, e.g., less than 15% e.g., less than 10%) of the tumor cells may be screened via methods described herein for any sample, irrespective of the ability to assign clone types.

[0005] In some aspects, the present disclosure provides methods [e.g., for classifying clonality states of tumor cell mutations (e.g., and selecting neo-epitopes for use in immunotherapy (e.g., personalized cancer vaccines, T cell receptor (TCR) therapy, etc.) based thereon)], the method comprising: (a) receiving, by a processor of a computing device, sequencing data [e.g., whole genome sequencing (WGS) data; e.g., whole exome sequencing (WES) data; e.g., single nucleotide polymorphism (SNP) array data; e.g., RNA sequencing data; e.g., wherein the sequencing data comprises a plurality of replicates (e.g., obtained from multiple samples of extracted gDNA from tumor and normal samples; e.g., obtained for multiple library preparations (e.g., from a single sample); e.g., e.g., obtained via multiple sequencing runs on a single library (e.g., technical replicates))] from a tumor sample obtained from a subject; (b) detecting, by the processor, based on the sequencing data, a plurality of candidate mutations, each representing a mutation occurring within a population of tumor cells of the tumor sample [e.g., a single nucleotide variation (SNV), an insertion and / or deletion (indel), a structural variant (SV), etc.]; (c) for each particular candidate mutation of at least a portion of the plurality of candidate mutations, determining, by the processor, a corresponding clone type selected from a set of discrete clone types; and (d) selecting, by the processor, a subset of the candidate mutations {e.g., for inclusion in and / or targeting via a construct [e.g., a therapeutic agent (e.g., a personalized therapeutic agent, such as a personalized cancer vaccine, a T cell receptor (TCR) therapy, etc.)]} based at least in part on the determined clone type classifications.

[0006] In certain embodiments, each member of the set of discrete clone types corresponds to one or more (e.g., distinct) subpopulations of cancer cells, each subpopulation having a characteristic set (e.g., combination) of mutations. In certain embodiments, the one or more subpopulations of cancer cells correspond to variant subpopulations evolved from an initial, clonal, population arising from a progenitor cell.

[0007] In certain embodiments, each member of the set of discrete clone types represents a (e.g., relative) prevalence of the corresponding subpopulation(s) of cancer cells.- 2 - 13241940vlAttorney Docket No. 2013237-1502

[0008] In certain embodiments, the set of discrete clone types comprises a first clone type that represents a major subpopulation comprising a characteristic set of mutations present in a majority of cancer cells (e.g., at the time of tumor sample extraction).

[0009] In certain embodiments, the set of discrete clone types comprises a second clone type that represents a minor subpopulation comprising a characteristic set of mutations present in a minority of cancer cells (e.g., at the time of tumor sample extraction).

[0010] In certain embodiments, each discrete clone type is characterized by a (e.g., estimated) nominal cellularity based on cellularities of mutations assigned to that clone type (for example, taking an average cellularity of all mutations assigned to a given clone type).

[0011] In certain embodiments, provided methods comprise determining, for at least a portion of the mutations, the assigned clone type based on a clustering of cellularity values.

[0012] In certain embodiments, provided methods comprise determining an (e.g., relative) evolutionary time point associated with each subpopulation and / or mutation thereof.

[0013] In certain embodiments, provided methods comprise associating a particular discrete clone type with a driver gene (e.g., thereby identifying the driver gene as responsible for driving growth of one or more subpopulations represented by the discrete clone type).

[0014] In certain embodiments, each member of the set of discrete clone types represents a quantized biological cellularity and corresponds to a different estimated cellularity range for the particular candidate mutation.

[0015] In certain embodiments, a set of discrete clone types comprises a plurality of states, each associated with a particular (e.g., quantized) range of biological cellularity and to which mutations are assigned based at least in part on their estimated cellularity values (e.g., having been determined to be within the particular associated range of biological cellularity for a given state).

[0016] In certain embodiments, a set of discrete clone types comprises more than two discrete clone types (e.g., is not binary, merely distinguishing between clonal and non-clonal or clonal vs sub clonal).

[0017] In certain embodiments, a set of discrete clone types comprise one or more of the following: a clone type representing mutations with a biological cellularity of 1 (e.g., a putative clonal state); a clone type representing mutations with a biological cellularity of approximately 1 [e.g., greater than or equal to about 0.8 (e.g., greater than or equal to about - 3 - 13241940vlAttorney Docket No. 2013237-15020.85; e.g., greater than or equal to about 0.9; e.g., greater than or equal to 0.95)] (e.g., a nearly clonal state); and a plurality of clone types representing multiple populations of biologically subclonal mutations, each associated with a particular (e.g., quantized) range of biological cellularity.

[0018] In certain embodiments, a plurality of clone types representing multiple populations of biologically subclonal mutations comprises at least a major subclone state and a minor subclone state, the major subclone state representing a first range of biological cellularity and the minor subclone state representing a second range of biological cellularity, the first range greater than the second (e.g., the lower bound of the first range greater than or equal to the upper bound of the second range).

[0019] In certain embodiments, a major subclone state is subdivided into a plurality of sub-states, each associated with and representing a different range of biological cellularity.

[0020] In certain embodiments, a minor subclone state represents mutations present in a minority of tumor cells (e.g., optionally wherein the minor subclone state is subdivided into a plurality of sub-states, each associated with a cellularity range (e.g., below 50%)).

[0021] In certain embodiments, a plurality of clone types representing multiple populations of biologically subclonal mutations comprises one or more (e.g., a plurality of) prevalent subclone states representing mutations present in a majority of tumor cells (e.g., highly prevalent in the population tumor cells, but not necessarily present in all tumor cells) (e.g., a plurality of prevalent subclone states)[e.g., said plurality comprising a first prevalent state and a second prevalent state, the first prevalent state representing mutations more common that the second prevalent state (e.g., the first prevalent state representing mutations having a biological cellularity within a first range and the second prevalent state representing mutations having a biological cellularity within a second range, the first range greater than the second)].

[0022] In certain embodiments, each member of the set of discrete clone types represents a distinct clone type indicative of (i) a likelihood that the particular candidate mutation is a biologically clonal mutation (e.g., based on a P- value determined from a statistical hypothesis test) and / or (ii) its prevalence within the tumor sample [e.g., a quantized prevalence, such a particular range of prevalence values; e.g., a biological and / or estimated cellularity (e.g., a quantized cellularity, such as a particular range of cellularity values)]- 4 - 13241940vlAttorney Docket No. 2013237-1502within the population tumor cells, thereby determining clone type classifications for the portion of the candidate mutations.

[0023] In certain embodiments, provided methods (e.g., further) comprise receiving, by the processor, sequencing data obtained from a normal sample, and using the normal sample sequencing data together with the tumor sample sequencing data to perform one or more (e.g., all) of steps (b), (c), and (d) [e.g., at step (b), detecting, by the processor, the plurality of candidate mutations based on the tumor sample sequencing data and the normal sample sequencing data; e.g., at step (c), determining the corresponding clone type for each of at least a portion of the plurality of candidate mutations using the tumor sample sequencing data and the normal sample sequencing data; e.g., at step (d), selecting the subset of candidate mutations based further in part on the tumor sample sequencing data and the normal sample sequencing data].

[0024] In certain embodiments, provided methods comprise, following step (d), producing a pharmaceutical composition comprising a polyribonucleotide encoding one or more neoepitopes, wherein at least a portion of the one or more neoepitopes are encoded by nucleotide sequence comprising a member of the selected subset of candidate mutations.

[0025] In certain embodiments, provided methods comprise [e.g., following step (d)], using the selected subset of candidate mutations to produce an enriched population of T-cells for the subject, said enriched population of T-cells capable of specifically binding to a plurality of complexes, each complex of the plurality comprising at least a portion of a neoepitope encoded by a member of the selected subset of candidate mutations.

[0026] In certain embodiments, a plurality further comprises a major histocompatibility complex (MHC) protein expressed by cancer cells (e.g., of the subject) and / or antigen presenting cells (APCs) (e.g., of the subject).

[0027] In certain embodiments, provided methods comprise [e.g., following step (d)], using the selected subset of candidate mutations to produce an enriched population of tumor infiltrated lymphocytes (TILs) for the subject, said enriched population of TILs capable of specifically binding to a plurality of complexes, each complex of the plurality comprising at least a portion of a neoepitope encoded by a member of the selected subset of candidate mutations.- 5 - 13241940vlAttorney Docket No. 2013237-1502

[0028] In certain embodiments, each complex of the plurality further comprises a major histocompatibility complex (MHC) protein expressed by cancer cells (e.g., of the subject) and / or antigen presenting cells (APCs) (e.g., of the subject).

[0029] In certain embodiments, provided methods comprise [e.g., following step (d)], using the selected subset of candidate mutations to produce a T-cell receptor (e.g., from a T-cell from blood and / or from a TIL) for the subject, said T-cell receptor capable of specifically binding to a complex comprising at least a portion of a neoepitope encoded by a member of the selected subset of candidate mutations.

[0030] In certain embodiments, a complex further comprises a major histocompatibility complex (MHC) protein expressed by cancer cells (e.g., of the subject) and / or antigen presenting cells (APCs) (e.g., of the subject).

[0031] In certain embodiments, provided methods comprise [e.g., following step (d)], using the selected subset of candidate mutations to produce a chimeric antigen receptor (CAR) for the subject, said CAR capable of specifically binding to a complex comprising at least a portion of a neoepitope encoded by a member of the selected subset of candidate mutations.

[0032] In certain embodiments, a complex further comprises a major histocompatibility complex (MHC) protein expressed by cancer cells (e.g., of the subject) and / or antigen presenting cells (APCs) (e.g., of the subject).

[0033] In certain embodiments, provided methods comprise: determining, by the processor, based on the sequencing data, an estimated tumor sample purity (e.g., an overall fraction and / or percentage of cells in the tumor sample that are tumor, and not normal, cells) and / or purity bound (e.g., an upper bound; e.g., a lower bound); and prior to proceeding to step (b) and / or step (c), determining, by the processor, the estimated tumor sample purity and / or the purity bound to be greater than or equal to a purity threshold (e.g., 20% purity; e.g.; 40% purity; e.g., 50% purity).

[0034] In certain embodiments, a tumor sample comprises tumor cells and normal (e.g., non-cancerous) cells [e.g., and wherein (e.g., the sequencing data is WES data and) at least 40% or more (e.g., at least 50% or more) of cells of the tumor sample are tumor cells; e.g., and wherein (e.g., the sequencing data is WGS data and) at least 10% or more (e.g., at least 20% or more) of cells of the tumor sample are tumor cells].- 6 - 13241940vlAttorney Docket No. 2013237-1502

[0035] In certain embodiments, provided methods comprise performing, by the processor, tumor deconvolution (e.g., to determine an estimated tumor sample purity and / or purity bound in as described herein).

[0036] In certain embodiments, provided methods comprise evaluating one or more quality control metrics with respect to the performed tumor deconvolution and accepting result(s) (e.g., a solution corresponding to the estimated purity; e.g., a particular tumor model; e.g., a primary copy number) of the tumor deconvolution based thereon (e.g., based on a determination of passing the one or more quality control metrics) (e.g., or, if the tumor deconvolution fails one or more of the quality control metrics, determining the purity bound).

[0037] In certain embodiments, provided methods comprise performing the tumor deconvolution using a plurality of tumor models [e.g., wherein each tumor model is associated with (e.g., a function of) a different absolute copy number scale (e.g., a prospective, hypothesized, absolute copy number of primary balanced heterozygous segments) and yields a (e.g., different) corresponding purity estimate].

[0038] In certain embodiments, provided methods comprise (e.g., responsive to a determination that the tumor deconvolution is unsuccessful, e.g., based on failure of one or more of the quality control metrics), for each particular candidate mutation, determining, by the processor, a plurality of prospective clone types, each prospective clone type determined using a distinct one of the plurality of tumor models and using the plurality of prospective clone types to determine the (e.g., final) corresponding clone type for the particular candidate mutation (e.g., selecting the corresponding clone type as a most common one of the prospective clone types, e.g., if a minimum fraction (up to all) of the prospective clone types are the same, otherwise selecting the corresponding clone type as undetermined).

[0039] In certain embodiments, provided methods comprise (e.g., responsive to a determination that the tumor deconvolution is unsuccessful, e.g., based on failure of one or more of the quality control metrics), for each particular candidate mutation determining, by the processor, a plurality of prospective clone types, each prospective clone type determined using a distinct one of the plurality of tumor models, identifying candidate mutations for which all of the corresponding prospective clone types are a same clone type and assigning each of those mutations a scale invariant clone type.

[0040] In certain embodiments, provided methods comprise, prior to step (c), identifying, by the processor, a low confidence subset of the candidate mutations, said low - 7 - 13241940vlAttorney Docket No. 2013237-1502confidence subset comprising candidate mutations determined to be likely false positives and / or rare subclones; and excluding the low confidence subset of the candidate mutations from (i) the portion of the candidate mutations for which clone type classifications are determined at step (c) and / or (ii) the subset selected for inclusion in the construct at step (d).

[0041] In certain embodiments, identifying a low confidence subset comprises: determining, for each candidate mutation, a mutation confidence score that measures that measures a likelihood that a site associated with (e.g., detected as having) a given mutation encodes that given mutation (e.g., Hl: a ZZ XX / XZ event with possible copy number variation and normal contamination) in comparison with a likelihood that it does not (encode that mutation) (e.g., HO:ZZ); and determining whether a particular candidate mutation is a likely false positive and / or rare subclone based on the mutation confidence score.

[0042] In certain embodiments, a set of discrete clone types comprises one or more clonal states indicative of a high likelihood that a candidate mutation is a biologically clonal mutation and / or highly prevalent (e.g., so as to be nearly clonal) {e.g., determined or estimated to be present in about 70% or more [e.g., about 75% or more (e.g., about 80% or more)] tumor cells}.

[0043] In certain embodiments, a set of discrete clone types comprises one or more clonal states indicative of a high likelihood that a candidate mutation is a biologically clonal mutation.

[0044] In certain embodiments, a set of discrete clone types comprises at least two (e.g., distinct) clonal states [e.g., a putative clonal state representing mutations with a biological cellularity of 1 and a nearly clonal state, representing mutations with a biological cellularity of nearly 1 {e.g., greater than or equal to about 0.8 [e.g., greater than or equal to about 0.85 (e.g., greater than or equal to 0.9 (e.g., greater than or equal to 0.95)]}, each of the at least two clonal states associated with a distinct set of classification criteria and representing differing likelihoods that a candidate mutation is a biologically clonal mutation (e.g., as determined via differing levels of rigor based statistical significance, thresholds, or other metrics).

[0045] In certain embodiments, provided methods comprise, for at least one particular candidate mutation: determining, by the processor, a cellularity confidence interval (CI) lower bound (e.g., a 95% CI lower bound or a 68% CI lower bound) for the particular mutation; and classifying, by the processor, the particular candidate mutation as belonging to - 8 - 13241940vlAttorney Docket No. 2013237-1502a particular one of the one or more clonal states based at least in part on the cellularity CI lower bound {e.g., based at least in part on the cellularity CI lower bound having been determined to exceed a CI threshold value [e.g., wherein the CI threshold value is a puritydependent threshold value, having two or more values, each associated with a different range of purities (e.g., as described herein)]}.

[0046] In certain embodiments, provided methods comprise, for at least one particular candidate mutation: determining, by the processor, for the particular candidate mutation, a first P- value representing a probability of observing a cellularity of the particular candidate mutation below one (e.g., given a null hypothesis that the particular mutation is biologically clonal); and classifying, by the processor, the particular candidate mutation as belonging to a particular one of the one or more clonal states based at least in part on the first P- value (e.g., based at least in part on the first P- value having been determined to exceed a particular P-value threshold).

[0047] In certain embodiments, provided methods comprise: determining, by the processor, an estimated cellularity for the particular candidate mutation; and classifying, by the processor, the particular candidate mutation as belonging to the particular one of the one or more clonal states based (e.g., further, at least in part) on the estimated cellularity (e.g., based on the estimated cellularity having been determined to be less than one).

[0048] In certain embodiments, provided methods comprise, for at least one particular candidate mutation: determining, by the processor, for the particular candidate mutation, a second P-value representing a probability of observing a cellularity of the particular candidate mutation at or above one (e.g., given a null hypothesis that the particular mutation is biologically clonal); and classifying, by the processor, the particular candidate mutation as belonging to a particular one of the one or more clonal states based at least in part on the second P- value (e.g., based at least in part on the second P- value having been determined to exceed a particular P- value threshold).

[0049] In certain embodiments, provided methods comprise: determining, by the processor, an estimated cellularity for the particular candidate mutation; and classifying, by the processor, the particular candidate mutation as belonging to the particular one of the one or more clonal states based (e.g., further, at least in part) on the estimated cellularity (e.g., based on the estimated cellularity having been determined to be greater than or equal to one).- 9 - 13241940vlAttorney Docket No. 2013237-1502

[0050] In certain embodiments, one or more clonal states comprise a plurality of clonal states (e.g., two clonal states) and provided methods comprise, for at least one particular candidate mutation: determining, by the processor, one or both of (i) and (ii) as follows: (i) a first P- value representing a probability of observing a cellularity of the particular candidate mutation below one (e.g., given a null hypothesis that the particular mutation is biologically clonal); and (ii) a second P- value representing a probability of observing a cellularity of the particular candidate mutation at or above one (e.g., given a null hypothesis that the particular mutation is biologically clonal); and classifying, by the processor, the particular candidate mutation as belonging to a first clonal state (e.g., a putative clonal state) of the plurality of clonal states if the first P-value and / or the second P-value is / are above a first P- value threshold, or classifying, by the processor, the first candidate mutation as belonging to a second clonal state (e.g., a nearly clonal state) of the plurality of clonal states if the first P-value and / or second P-value is / are below the first P-value threshold (e.g., and above a second P-value threshold, lower than the first P-value threshold).

[0051] In certain embodiments, provided methods comprise, for at least one particular candidate mutation: determining, by the processor, an estimated cellularity (ρ) for the particular candidate mutation; and classifying, by the processor, the particular candidate mutation as belonging to a particular one of the one or more clonal states based at least in part on the estimated cellularity (e.g., based on the estimated cellularity having been determined to exceed a particular cellularity threshold).

[0052] In certain embodiments, a set of discrete clone types comprises a minor subclone state indicative of a high likelihood that a candidate mutation is (i) a biologically subclonal mutation and (ii) present in tumor cells as a fraction within a range that is lower than at least one other subclonal state (e.g., a major subclone state; e.g., one or more prevalent subclone states).

[0053] In certain embodiments, a minor subclone state represents candidate mutations present in a minority {e.g., about 50% or fewer [e.g., less than about 45% (e.g., less than about 40%)]} of tumor cells (e.g., wherein the minor subclone state is indicative of both the high likelihood that a candidate mutation is biologically subclonal and present in the minority of tumor cells).- 10 - 13241940vlAttorney Docket No. 2013237-1502

[0054] In certain embodiments, provided methods comprise, for at least one particular candidate mutation: determining, by the processor, an estimated cellularity (ρ) for the particular candidate mutation; and classifying, by the processor, the particular candidate mutation as belonging to the minor subclone state based at least in part on the estimated cellularity {e.g., based on the estimated cellularity having been determined to be less than a cellularity threshold value [e.g., less than or equal to about 50% (e.g., less than about 45% (e.g., less than about 40%))]}.

[0055] In certain embodiments, provided methods comprise, for at least one particular candidate mutation: determining, by the processor, for the particular candidate mutation, a first P- value representing a probability of observing a cellularity of the particular candidate mutation below one (e.g., given a null hypothesis that the particular candidate mutation is biologically clonal); and classifying, by the processor, the particular candidate mutation as belonging to the minor subclone state based at least in part on the first P-value.

[0056] In certain embodiments, classifying a particular candidate mutation comprises determining the first P- value to be below a particular P- value threshold(s) {e.g., wherein the particular P- value threshold is a purity-dependent threshold, having two or more values, each associated with a different range of purities (e.g., as recited in claim 75)} {e.g., wherein the value of the particular P- value threshold is dependent on an estimated cellularity for the particular candidate mutation}.

[0057] In certain embodiments, provided methods comprise, for at least one particular candidate mutation: determining, by the processor, a cellularity confidence interval (CI) (e.g., a 95% confidence interval or a 68% confidence interval) upper bound for the particular candidate mutation; and classifying, by the processor, the particular candidate mutation as belonging to the minor subclone state based at least in part on the cellularity CI upper bound {e.g., based at least in part on the cellularity CI upper bound having been determined to be below a CI threshold value [e.g., wherein the CI threshold value is a purity-dependent threshold value, having two or more values, each associated with a different range of purities (e.g., as described herein)]}.

[0058] In certain embodiments, a set of discrete clone types comprises an undetermined state (e.g., N / A) indicative of mutations for which clone type assignments are determined and / or predicted to be unreliable [e.g., indicative of any of (e.g., significant) noise- 11 - 13241940vlAttorney Docket No. 2013237-1502(such as stochastic noise, biological noise, experimental noise), local bias, alignment errors, potential errors in estimated parameter values, model assumptions].

[0059] In certain embodiments, provided methods comprise, for at least one particular candidate mutation: determining, by the processor, an estimated copy number (CN) for the particular candidate mutation [e.g., determining an estimated CN (e.g., an absolute CN) of a segment (e.g., a gene, locus, site, etc.) to which the particular candidate mutation maps]; and classifying, by the processor, the particular candidate mutation as belonging to the undetermined state based at least in part on the estimated copy number being about zero (0).

[0060] In certain embodiments, provided methods comprise, for at least one particular candidate mutation: determining, by the processor, an estimated cellularity (ρ) for the particular candidate mutation; and classifying, by the processor, the particular candidate mutation as belonging to the undetermined state based at least in part on the estimated cellularity {e.g., exceeding a particular cellularity threshold value [e.g., said particular cellularity threshold being greater than 1 (e.g., greater than 1.2, e.g., greater than 1.25, e.g., greater than 1.3, etc.) (e.g., thereby reflecting presence of a stochastic error and / or errors in one or more measured and / or estimated parameters, such as cellularity)] }.

[0061] In certain embodiments, provided methods comprise, for at least one particular candidate mutation: determining, by the processor, for the particular candidate mutation, a P-value representing a probability of observing a cellularity above one for the particular candidate mutation (e.g., given a null hypothesis that the particular mutation is biologically clonal); and classifying, by the processor, the particular candidate mutation as belonging to the undetermined state based at least in part on the P- value falling below corresponding P-value threshold (e.g., thereby reflecting that the null hypotheses may be rejected and / or a measured cellularity value exceeding one is not merely due to noise fluctuations).

[0062] In certain embodiments, a set of discrete clone types comprises one or more prevalent subclone states [e.g., representing and / or reserved for mutations that are not determined to be biologically clonal, but are present at a sufficiently high rate in cancer cells (e.g., to be a viable target for an immunotherapy; e.g., present in a majority (e.g., greater than half (e.g., 50% or more)) of tumor cells (e.g., present in 60% or more of the tumor cells))].

[0063] In certain embodiments, a set of discrete clone types comprises a plurality of prevalent subclone states, each representing mutations present in a majority of tumor cells (e.g., highly prevalent in the population tumor cells, but not necessarily present in all tumor - 12 - 13241940vlAttorney Docket No. 2013237-1502cells, said plurality comprising at least a first prevalent subclone state (e.g., prevalent-high) and a second prevalent subclone state (e.g., prevalent- low), the first prevalent state representing mutations more common that the second prevalent state (e.g., the first prevalent state representing mutations having a biological cellularity within a first range and the second prevalent state representing mutations having a biological cellularity within a second range, the first range greater than the second)].

[0064] In certain embodiments, provided methods comprise, for at least one particular candidate mutation: determining, by the processor, an estimated cellularity (ρ) for the particular candidate mutation; and classifying, by the processor, the particular candidate mutation as belonging to a particular one of the one or more prevalent subclone states based at least in part on the estimated cellularity [e.g., based on the estimated cellularity having been determined to be greater than a first cellularity threshold; e.g., based on the estimated cellularity having been determined to be between two cellularity threshold values (e.g., falling within a particular cellularity range)].

[0065] In certain embodiments, provided methods comprise, for at least one particular candidate mutation: determining, by the processor, a cellularity confidence interval (CI) lower bound (e.g., a 95% CI lower bound or a 68% CI lower bound) for the particular mutation; and classifying, by the processor, the particular candidate mutation as belonging to a particular one of the one or more prevalent subclone states based at least in part on the cellularity CI lower bound {e.g., based at least in part on the cellularity CI lower bound having been determined to exceed a CI threshold value [e.g., wherein the first CI threshold value is a purity-dependent threshold value, having two or more values, each associated with a different range of purities (e.g., as described herein)] }.

[0066] In certain embodiments, a set of discrete clone types comprises two or more states.

[0067] In certain embodiments, a set of discrete clone types comprises five or more states [e.g., a putative clonal state, a nearly clonal state, a prevalent (e.g., high and / or low) subclone state, a minor subclone state, and an undetermined (e.g., “N / A”) state].

[0068] In certain embodiments, a set of discrete clone types comprises six states, said six states comprising two clonal states, two prevalent subclone state (e.g., a prevalent-high and a prevalent- low state), a minor subclone state, and an undetermined (e.g., N / A) state.- 13 - 13241940vlAttorney Docket No. 2013237-1502

[0069] In certain embodiments, a set of discrete clone types comprises five states, said five states comprising a clonal state, two prevalent subclone state (e.g., a prevalent-high and a prevalent- low state), a minor subclone state, and an undetermined (e.g., N / A) state.

[0070] In certain embodiments, a set of discrete clone types comprises four states, said four states comprising a clonal state, a prevalent subclone state, a minor subclone state, and an undetermined state (e.g., NA).

[0071] In certain embodiments, a set of discrete clone types comprises three states, said three states comprising a clonal state, a subclonal state, and an undetermined state (e.g., N / A).

[0072] In certain embodiments, a set of discrete clone types comprises at least a clonal state and / or a subclonal state (e.g., wherein the set of discrete clone types is a binary classification, e.g., between clonal and not-clonal, clonal and subclonal, not-subclonal and subclonal, etc.).

[0073] In certain embodiments, step (c) comprises, for at least one particular candidate mutation: determining, via a statistical hypothesis test, a likelihood (e.g., a based on a P- value) that the particular candidate mutation is biologically clonal; and determining the corresponding clone type for the particular candidate mutation based at least in part on the determined likelihood.

[0074] In certain embodiments, step (c) comprises, for at least one particular candidate mutation: determining, using the sequencing data, an estimated cellularity for the particular candidate mutation (e.g., wherein the estimated cellularity measures an estimated fraction of tumor cells within the sample that harbor the particular candidate mutation); and determining the corresponding clone type for the particular candidate cell mutation based at least in part on the cellularity of the particular candidate mutation.

[0075] In certain embodiments, step (c) comprises, for at least one particular candidate mutation: determining, using the sequencing data, a likelihood (e.g., probability) of observing a cellularity of the particular candidate mutation below one (e.g., based on a null hypothesis of the particular candidate mutation being biologically clonal) (e.g., a P-value); and determining the corresponding clone type for the particular candidate mutation based at least in part on the determined likelihood (e.g., the determined P-value).- 14 - 13241940vlAttorney Docket No. 2013237-1502

[0076] In certain embodiments, step (c) comprises, for at least one particular candidate mutation: determining, using the sequencing data, a likelihood (e.g., a probability) of observing a cellularity of the particular candidate mutation above one (e.g., based on a null hypothesis of the particular candidate mutation being biologically clonal) (e.g., a P-value); and determining the corresponding clone type for the particular candidate mutation based at least in part on the determined likelihood (e.g., the determined P-value).

[0077] In certain embodiments, step (c) comprises, for at least one particular candidate mutation: determining a cellularity confidence interval (CI) lower and / or upper bound for the particular candidate mutation; and determining the corresponding clone type for the particular candidate mutation based at least in part on the determined cellularity CI lower and / or upper bound(s).

[0078] In certain embodiments, step (c) comprises, for at least one particular candidate mutation: evaluating, by the processor, a first set of criteria with respect to the particular candidate mutation and determining the first set of criteria to be satisfied; classifying, by the processor, the particular candidate mutation as belonging to a first clone type of the set of discrete clone types based on the determined satisfaction of the first set of criteria; evaluating, by the processor, a second set of criteria with respect to the particular candidate mutation and determining the second set of criteria to be satisfied [e.g., wherein the second set of criteria includes a previous clonality state assignment]; classifying, by the processor, the particular candidate mutation as belonging to a second clone type of the set of discrete clone types, different from the first, based on the determined satisfaction of the second set of criteria, thereby re-assigning the clone type classification for the particular candidate mutation.

[0079] In certain embodiments, evaluating the second set of criteria comprises determining a purity of the tumor sample to be below a particular purity threshold {e.g., such that the re-assignment occurs when sample purity is determined to be low [e.g., reassigning a mutation initially classified as belonging to a prevalent subclone state (e.g., a prevalent-high subclone state, representing mutations having greater prevalence than those in other prevalent (e.g., prevalent-low) states) to a clonal state (e.g., a nearly clonal state; a putative clonal state) when the estimated purity is below the particular purity threshold (e.g., 0.35) (e.g., optionally, together with other criteria, such as a cellularity below one and a P-value representing a probability of observing a cellularity of the mutation below one P-value being greater than a P- value threshold, e.g., 0.05)]}.- 15 - 13241940vlAttorney Docket No. 2013237-1502

[0080] In certain embodiments, step (c) comprises determining a corresponding clone type for each (e.g., all) of the plurality of candidate mutations.

[0081] In certain embodiments, step (c) comprises, for at least one particular candidate mutation of the plurality of candidate mutations: evaluating a plurality of sets of criteria with respect to the particular candidate mutation; determining none of the sets of criteria to be satisfied for the particular candidate mutation; and classifying the particular candidate mutation as belonging to the undetermined state (e.g., based on the determination of none of the sets of criteria to be satisfied).

[0082] In certain embodiments, step (c) comprises, for at least one particular candidate mutation, determining values of one or more mutation characterization metrics for the particular candidate mutation.

[0083] In certain embodiments, one or more mutation characterization metrics comprises a cellularity confidence interval (CI) lower bound (e.g., a 95% CI lower bound; e.g., a 90% CI lower bound; e.g., a 80% CI lower bound, e.g. a 70% CI lower bound, e.g., a 65% lower bound; etc.).

[0084] In certain embodiments, one or more mutation characterization metrics comprises a cellularity CI upper bound (e.g., a 95% CI upper bound; e.g., a 90% CI upper bound; etc.).

[0085] In certain embodiments, one or more mutation characterization metrics comprises an estimated cellularity (p).

[0086] In certain embodiments, one or more mutation characterization metrics comprises a P-value representing a probability of observing a cellularity of the particular candidate mutation below one (e.g., given a null hypothesis that the particular mutation is biologically clonal).

[0087] In certain embodiments, one or more mutation characterization metrics comprises a P-value representing a probability of observing a cellularity of the particular candidate mutation at or above one (e.g., given a null hypothesis that the particular mutation is biologically clonal).

[0088] In certain embodiments, step (c) comprises comparing one or more of the mutation characterization metrics to one or more thresholds, at least a portion of which are purity dependent thresholds, having two or more values, each associated with a different - 16 - 13241940vlAttorney Docket No. 2013237-1502range of purities (e.g., estimated purities, e.g., upper or lower bounds of an estimated purity, e.g., determined purities) of the tumor sample (e.g., wherein the thresholds are more stringent at lower sample purities).

[0089] In certain embodiments, a set of discrete clonality states comprises a minor subclone state {e.g., indicative of a high likelihood that a candidate is a biologically subclonal mutation and present in a minority [e.g., less than half (e.g., less than 40%)] of tumor cells} and the method comprises: at step (c), determining the minor subclone state as the corresponding clone type for one or more of the candidate mutations; and at step (d), excluding the candidate mutations classified as belonging to the minor subclone state from the selected subset (e.g., for inclusion in and / or targeting via the construct).

[0090] In certain embodiments, provided methods comprise classifying one or more of the candidate mutations as clonal (e.g., passing a rigorous hypothesis test of clonality) and, at step (d), selecting at least a portion of the candidate mutations classified as clonal for inclusion in the selected subset (e.g., for inclusion in and / or targeting via the construct).

[0091] In certain embodiments, determined clone type classifications include a plurality of distinct clone types and wherein step (d) comprises prioritizing mutations for inclusion in the selected subset according to their determined clone types.

[0092] In certain embodiments, provided methods comprise ranking the plurality of discrete clone types such that mutations assigned to higher ranked clone types are prioritized for inclusion in the selected subset over mutations assigned to lower ranked clone types.

[0093] In certain embodiments, a plurality of discrete clone types comprises an undetermined state and a minor subclone state, and wherein the method comprises ranking the undetermined state above (e.g., as a higher priority) the minor subclone state (e.g., such that mutations assigned to the undetermined state are prioritized over mutations assigned to the minor subclone state.

[0094] In certain embodiments, a plurality of discrete clone types comprises: (A) one or more clonal states; (B) one or more (e.g., a plurality of) major subclone states (e.g., prevalent subclone states); and (C) one or more minor subclone state(s).

[0095] In certain embodiments, provided methods comprise ranking the one or more clonal states above the one or more prevalent subclone states and ranking the one or more major (e.g., prevalent) subclone states above the minor subclone state [e.g., optionally,- 17 - 13241940vlAttorney Docket No. 2013237-1502wherein the method comprises grouping the one or more clonal states together and / or with at least one (or more) of the prevalent subclone states and ranking those groups above the minor subclone state].

[0096] In certain embodiments, one or more clonal states comprise a putative clonal state and a nearly clonal state and the method comprises ranking the putative clonal state above the nearly clonal state.

[0097] In certain embodiments, a plurality of discrete clone types comprises two or more prevalent subclone states, wherein said two or more prevalent subclone states comprising a first prevalent state (e.g., prevalent-high) and a second prevalent state (e.g., prevalent- low), the first prevalent state representing mutations more common that the second prevalent state (e.g., the first prevalent state representing mutations having a biological cellularity within a first range and the second prevalent state representing mutations having a biological cellularity within a second range, the first range greater than the second), and wherein the method comprises ranking the first prevalent state above the second prevalent state.

[0098] In certain embodiments, a ranking is determined at least in part based on an identification of a driver phenotype (e.g., a fractional zygosity of 1) associated with one or more of the mutations.

[0099] In certain embodiments, provided methods comprise identifying at least a portion of the mutations as driver mutations (e.g., associated with a driver phenotype) and ranking the plurality of discrete clone types such that mutations that are (i) assigned to a particular clone type and (ii) are identified as driver mutations are ranked above mutations that are (i) also assigned to the particular clone type, but (ii) not identified as driver mutations.

[0100] In certain embodiments, a ranking is determined based at least in part on a purity (e.g., an estimated sample purity; e.g., an estimated purity bound; e.g., a measured purity) of the sample [e.g., based on the purity falling below a certain threshold (e.g., 45%)] [e.g., merging two or more clone types (e.g., clonal and prevalent-high)].

[0101] In certain embodiments, a ranking is determined based at least in part on a tumor mutational burden (TMB).- 18 - 13241940vlAttorney Docket No. 2013237-1502

[0102] In certain embodiments, a ranking is determined based at least in part on a number of non-synonymous (NS) mutations.

[0103] In certain embodiments, a ranking is determined based at least in part on one or more parameters that impact one or more members (e.g., biological properties) selected from the group consisting of expression, presentation, and immunogenicity (e.g., recognition).

[0104] In certain embodiments, a ranking is determined based at least in part on a cancer type.

[0105] In certain embodiments, a ranking is determined based at least in part on a number of mutations assigned to the clonal state.

[0106] In certain embodiments, provided methods comprise, at step (d), selecting the subset based at least in part on the classified clonality states in combination with an immunogenicity (e.g., MHC -binding) prediction score.

[0107] In certain embodiments, provided methods comprise, at step (d), selecting the subset based at least in part on the clone type classifications in combination with an expression score (e.g., quantifying an RNA expression).

[0108] In certain embodiments, provided methods comprise, at step (d), selecting the subset based at least in part on the clone type classifications in combination with one or more members selected from the group consisting of T cell receptor (TCR) recognition, copy number, zygosity, fractional zygosity, and essential gene (e.g., and / or other biological factors that can impact immunogenicity and / or tumor control).

[0109] In certain embodiments, step (c) comprises, for each particular candidate mutation, evaluating one or more sets of criteria in a stepwise fashion (e.g., one after another), wherein each set of criteria is associated with, and used to assign a given mutation to, a particular member of the set of discrete clone types.

[0110] In certain embodiments, provided methods comprise, for each particular mutation, evaluating at least a portion of the one or more sets of criteria by, at each particular step of one or more steps: evaluating a corresponding current set of the one or more sets of criteria with respect to the given mutation, the current set associated with, and used to assign a given mutation to, a corresponding one of the one or more discrete clone types and, either: (i) determining the criteria of the current set to be satisfied, and, based on said determination - 19 - 13241940vlAttorney Docket No. 2013237-1502of the criteria of the current set being satisfied, assigning the particular mutation to the corresponding clone type, or (ii) determining the criteria of the current set not to be satisfied and, based on said determination of the criteria of the current set not being satisfied, proceeding to subsequent step.

[0111] In certain embodiments, a set of discrete clone types comprises an undetermined classification state and wherein a first step of the one or more steps is associated with the undetermined state and the method comprises, at the first step, determining a corresponding first criteria to be satisfied and, based on the determined satisfaction of the first criteria, assigning the particular mutation to the undetermined state.

[0112] In certain embodiments, a set of discrete clone types comprises an undetermined state and wherein, for at least one particular candidate mutation, the method comprises determining, for each of the one or more steps, the corresponding criteria not to be satisfied for the particular candidate mutation and (e.g., based on a determination that none of the criteria are satisfied) assigning the particular candidate mutation to the undetermined state.

[0113] In certain embodiments, two or more of the one or more steps are associated with a same clone type but comprise different sets of criteria.

[0114] In certain embodiments, provided methods comprise, for at least one particular candidate mutation: evaluating, by the processor, a first set of criteria with respect to the particular candidate mutation and determining the first set of criteria to be satisfied; classifying, by the processor, the particular candidate mutation as belonging to a first clone type of the set of discrete clone types based on the determined satisfaction of the first set of criteria; evaluating, by the processor, a second set of criteria with respect to the particular candidate mutation and determining the second set of criteria to be satisfied [e.g., wherein the second set of criteria includes a previous clonality state assignment]; and classifying, by the processor, the particular candidate mutation as belonging to a second clone type of the set of discrete clone types, different from the first, based on the determined satisfaction of the second set of criteria, thereby re-assigning the clone type classification for the particular candidate mutation.

[0115] In certain embodiments, evaluating the second set of criteria comprises determining a purity of the tumor sample to be below a particular purity threshold {e.g., such that the re-assignment occurs when sample purity is determined to be low [e.g., reassigning a -20 - 13241940vlAttorney Docket No. 2013237-1502mutation initially classified as belonging to a prevalent subclone state (e.g., a prevalent-high subclone state, representing mutations having greater prevalence than those in other prevalent (e.g., prevalent-low) states) to a clonal state (e.g., a nearly clonal state; a putative clonal state) when the estimated purity is below the particular purity threshold (e.g., 0.35) (e.g., optionally, together with other criteria, such as a cellularity below one and a P-value representing a probability of observing a cellularity of the mutation below one P-value being greater than a P- value threshold, e.g., 0.05)]}.

[0116] The method of any one of the preceding claims, comprising, causing, by the processor, display of a graphical clonality state spectrum view in which mutations assigned a particular clone type are grouped according their absolute copy numbers and zygosities (e.g., from left to right, along a horizontal, x-, axis) and a graphical representation (e.g., a line plot, a scatter plot, etc.) of VAFs and / or cellularity values of the mutations assigned the particular clone type is displayed.

[0117] In certain embodiments, provided methods comprise, for each particular clone type of one or more clone types: grouping corresponding mutations (e.g., determined to have the particular clone type) into one or more subsets according to their absolute copy numbers and zygosities, such that each of the one or more subsets comprises mutations having a same absolute copy number and zygosity; and causing display of, within the graphical clonality state spectrum view, a graphical representation of VAFS and / or cellularity values for each subset of mutations [e.g., wherein the subsets of corresponding mutations are distributed along a horizontal in order of increasing absolute copy number then increasing zygosity (e.g., such that the graphical representation of VAFs and / or cellularity values for each subset of mutations comprises a series of lines)].

[0118] In certain embodiments, sequencing data comprises WGS and WES data, and wherein step (b) comprises using the WES to detect the plurality of candidate mutations and wherein step (c) comprises determining absolute copy numbers for segments comprising the plurality of candidate mutations using the WGS data (e.g., by performing tumor deconvolution; e.g., as recited in claims 23-25) and using the determined absolute copy numbers to determine the corresponding clone type for each particular candidate mutation.

[0119] In certain embodiments, sequencing data comprises RNAseq data and step (b) comprises filtering false positives using the RNAseq data (e.g., to identify false positives due to PCR errors).- 21 - 13241940vlAttorney Docket No. 2013237-1502

[0120] In some aspects, the present disclosure provides methods for detecting and filtering out rare sub-clones within sets of tumor cell mutations [e.g., for use in immunotherapy (e.g., as neoantigen targets for personalized cancer vaccines, T cell receptor (TCR) therapy, etc.)], the method comprising: (a) receiving, by a processor of a computing device, sequencing data from a tumor sample obtained from a subject; (b) detecting, by the processor, based on the sequencing data, an initial set of potential tumor cell mutations, the initial set comprising a plurality of detected mutations, each representing a potential mutation present in one or more tumor cells of the tumor sample; (c) for each particular detected mutation within at least a portion of the initial set of potential tumor cell mutations, determining, by the processor, a mutation confidence score that measures a likelihood that a site associated with the particular detected mutation encodes the particular detected mutation (e.g., Hl: a ZZ XX / XZ event with possible copy number variation and normal contamination) in comparison a likelihood that the site does not encode the particular detected mutation (e.g.,HO: ZZ), thereby determining a plurality of mutation confidence scores; (d) filtering, by the processor, the initial set of potential tumor cell mutations to exclude one or more detected mutations identified as excessively rare or unlikely based at least in part on the plurality of mutation confidence scores, to generate a high confidence set of tumor cell mutations; and (e) selecting, by the processor, from the high confidence set, one or more target mutations for inclusion in and / or targeting by a construct [e.g., a therapeutic agent (e.g., a personalized therapeutic agent, such as a personalized cancer vaccine, a T cell receptor (TCR) therapy, etc.)].

[0121] In certain embodiments, provided methods comprise: (f) producing a pharmaceutical composition comprising a polyribonucleotide encoding one or more neoantigen epitopes, wherein at least a portion of the one or more neoepitopes are encoded by nucleotide sequence comprising one or more of the one or more target mutations.

[0122] In certain embodiments, provided methods comprise, [e.g., following step (e)], using the one or more target mutations to produce an enriched population of T-cells for the subject, said enriched population of T-cells capable of specifically binding to a plurality of complexes, each complex of the plurality comprising at least a portion of a neoepitope encoded by at least one of the one or more target mutations.

[0123] In certain embodiments, each complex of the plurality further comprises a major histocompatibility complex (MHC) protein expressed by cancer cells (e.g., of the subject) and / or antigen presenting cells (APCs) (e.g., of the subject).- 22 - 13241940vlAttorney Docket No. 2013237-1502

[0124] In certain embodiments, provided methods comprise, [e.g., following step (d)], using the selected subset of candidate mutations to produce an enriched population of tumor infiltrated lymphocytes (TILs) for the subject, said enriched population of TILs capable of specifically binding to a plurality of complexes, each complex of the plurality comprising at least a portion of a neoepitope encoded by a member of the selected subset of candidate mutations.

[0125] In certain embodiments, each complex of the plurality further comprises a major histocompatibility complex (MHC) protein expressed by cancer cells (e.g., of the subject) and / or antigen presenting cells (APCs) (e.g., of the subject).

[0126] In certain embodiments, provided methods comprise, [e.g., following step (d)], using the one or more target mutations to produce a T-cell receptor (e.g., based on T-cells from blood and / or TILs) for the subject, said T-cell receptor capable of specifically binding to a complex comprising at least a portion of a neoepitope encoded by at least one of the one or more target mutations.

[0127] In certain embodiments, a complex further comprises a major histocompatibility complex (MHC) protein expressed by cancer cells (e.g., of the subject) and / or antigen presenting cells (APCs) (e.g., of the subject).

[0128] In certain embodiments, provided methods comprise, [e.g., following step (e)], using the one or more target mutations to produce a chimeric antigen receptor (CAR) for the subject, said CAR capable of specifically binding to a complex comprising at least a portion of a neoepitope encoded by at least one of the one or more target mutations.

[0129] In certain embodiments, a complex further comprises a major histocompatibility complex (MHC) protein expressed by cancer cells (e.g., of the subject) and / or antigen presenting cells (APCs) (e.g., of the subject).

[0130] In certain embodiments, a mutation confidence score comprises determining the likelihoods (e.g., testing them) over ranges of one or more of: sample purities, absolute copy number scale(s) (e.g., absolute CNs of PBHS’s), e.g., optionally, limiting variant allele frequencies (VAFs).

[0131] In certain embodiments, provided methods comprise determining (e.g., using a tumor model) an estimated purity (e.g., value or bound) and / or an estimated absolute copy number of the site encoding the detected mutation, and using the estimated purity and / or- 23 - 13241940vlAttorney Docket No. 2013237-1502estimated absolute copy number to determine the mutation confidence score, thereby determining, and filtering the detected mutations based on, a refined mutation confidence score (RMCS).

[0132] In certain embodiments, an estimated purity is an estimated (e.g., upper and / or lower) bound on a purity of the tumor sample.

[0133] In certain embodiments, provided methods comprise determining the estimated purity using a plurality of tumor models [e.g., wherein each tumor model is associated with (e.g., a function of) a different absolute copy number and yields a (e.g., different) corresponding purity estimate].

[0134] In certain embodiments, provided methods comprise: evaluating one or more quality control metrics with respect to the plurality of tumor models; determining, based on the one or more quality control metrics, a failure of tumor deconvolution; and responsive to the tumor deconvolution failure, determining, based on the plurality of tumor models, the estimated bound on the purity of the tumor sample [e.g., wherein each tumor model is associated with (e.g., a function of) a different absolute copy number scale (e.g., a prospective, hypothesized, absolute copy number of primary balanced heterozygous segments)and yields a (e.g., different) corresponding purity estimate and determining, as the estimated bound, a maximum of the purity estimates from the plurality of tumor models)].

[0135] In certain embodiments, provided methods comprise determining (e.g., using a tumor model) an estimated absolute copy number of an alternate allele (e.g., zygosity) for each detected mutation and using the estimated absolute copy number of an alternate allele to determine the mutation confidence score, thereby determining, and filtering the detected mutations based on, a refined mutation confidence score (RMCS).

[0136] In certain embodiments, filtering comprises comparing the mutation confidence score (e.g., MCS or RMCS) to a purity dependent threshold (e.g., when the purity is low then the threshold is reduced for the confidence - e.g., allowing for less confident mutations through - e.g., if the threshold isn’t reduced then it may block everything; e.g., adjusting position on the sensitivity vs. specificity scale).

[0137] In some aspects, the present disclosure provides methods [e.g., for classifying clonality states of tumor cell mutations (e.g., and selecting neo-epitopes for use in personalized cancer vaccines and / or TCR therapies based thereon)], the method comprising: (a) receiving, by a processor of a computing device, sequencing data [e.g., whole genome - 24 - 13241940vlAttorney Docket No. 2013237-1502sequencing (WGS) data; e.g., whole exome sequencing (WES) data; e.g., single nucleotide polymorphism (SNP) array data] from a tumor sample obtained from a subject; (b) detecting, by the processor, based on the sequencing data, a plurality of candidate mutations, each representing a mutation occurring within a population of tumor cells of the tumor sample [e.g., a single nucleotide variation (SNV), an insertion and / or deletion (indel), a structural variant (SV), etc.]; and (c) assigning, by the processor, the plurality of candidate mutations to one or more classes selected from a set of discrete clone types, where the set of discrete clone types includes an undetermined state indicative of mutations for which clone type assignments are determined and / or predicted to be unreliable [e.g., indicative of any of (e.g., significant) stochastic noise, local bias, alignment errors, potential errors in estimated parameter values, model assumptions], and wherein a first subset of the plurality of candidate mutations to the undetermined state (e.g., based at least in part on a determination of a presence of one or more stochastic errors and / or errors in measurement and / or estimation of one or more parameters measuring a prevalence of a given mutation; e.g., based on criteria as recited in claims above).

[0138] In certain embodiments, provided methods comprise: assigning, by the processor, a second subset of the plurality of candidate mutations to (i) a clonal state indicative of a given mutation having a high likelihood of being biologically clonal and / or (ii) a prevalent classification state indicative of a given mutation being highly prevalent in tumor cells, though not necessarily clonal; ranking and / or scoring (e.g., by the processor) the plurality of candidate mutations based at least in part on their assigned clone types, wherein candidate mutations of the second subset are ranked and / or scored higher than those of the first subset; and selecting (e.g., by the processor) a final subset of the candidate mutations for inclusion in a construct [e.g., a therapeutic agent (e.g., a personalized therapeutic agent, such as a personalized cancer vaccine, a T cell receptor (TCR) therapy, etc.)] based at least in part on the ranking and / or scoring of the candidate mutations.

[0139] In certain embodiments, provided methods comprise: assigning, by the processor, a third subset of the plurality of candidate mutations to a minor subclone state indicative of a given mutation having (i) a high likelihood of being biologically subclonal and / or (ii) a low level of prevalence in tumor cells (e.g., prevalence in a minority of tumor cells); ranking and / or scoring (e.g., by the processor) the plurality of candidate mutations based at least in part on their assigned clone types, wherein candidate mutations of the first subset are ranked and / or scored higher than those of the third subset; and selecting (e.g., by - 25 - 13241940vlAttorney Docket No. 2013237-1502the processor) a final subset of the candidate mutations for inclusion in a construct [e.g., a therapeutic agent (e.g., a personalized therapeutic agent, such as a personalized cancer vaccine, a T cell receptor (TCR) therapy, etc.)] based at least in part on the ranking and / or scoring of the candidate mutations.

[0140] In some aspects, the present disclosure provides systems comprising a processor of a computing device and memory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to various aspects and embodiments described herein.

[0141] In some aspects, the present disclosure provides methods of producing an immunotherapy construct (e.g., a cancer vaccine) for a subject, said provided methods comprising: detecting a plurality of candidate mutations (e.g., somatic mutations; e.g., non-synonymous somatic mutations) in tumor cells of from the subject and selecting, using a method or system of an aspect or embodiment described herein, a subset of the plurality of candidate mutations for inclusion in the immunotherapy construct; and synthesizing, as the immunotherapy construct, a polyribonucleotide encoding a plurality of neoepitopes, each encoded neoepitope corresponding to a candidate mutation of the selected subset [e.g., such that at least a portion (e.g., two or more, three or more, ten, 15, 20) (e.g., not necessarily all) of the selected subset have a corresponding candidate mutation].

[0142] In certain embodiments, provided methods comprise screening candidate mutations of the selected subset against T-cells and / or T-cell receptors (TCRs) derived from the subject.

[0143] In some aspects, the present disclosure provides methods comprising: determining a subset of T-cells within a population of T-cells that are capable of specifically binding a plurality of complexes, wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancerspecific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, and wherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer-specific mutations within the selected subset of candidate mutations determined by a method or system of an aspect or embodiment described herein; and enriching for (e.g., expanding) the subset of T-cells that are capable of specifically binding the plurality of complexes.- 26 - 13241940vlAttorney Docket No. 2013237-1502

[0144] In some aspects, the present disclosure provides methods comprising: administering to a subject a population of T-cells, wherein the population of T cells are capable of specifically binding a plurality of complexes, wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, wherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancerspecific mutations within the selected subset of candidate mutations determined by a method or system of an aspect or embodiment described herein, and wherein the genome of at least some of the subject’s cells comprises a subset (e.g., all) of the cancer-specific mutations.

[0145] In some aspects, the present disclosure provides methods comprising: determining a subset of TILs within a population of TILs that are capable of specifically binding a plurality of complexes, wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancerspecific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, and wherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer-specific mutations within the selected subset of candidate mutations determined by a method or system of an aspect or embodiment described herein; and enriching for (e.g., expanding) the subset of TILs that are capable of specifically binding the plurality of complexes.

[0146] In some aspects, the present disclosure provides methods comprising: administering to a subject a population of TILs, wherein the population of TILs are capable of specifically binding a plurality of complexes, wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer-specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, wherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer-specific mutations within the selected subset of candidate mutations determined by a method or system of an aspect or embodiment described herein, and wherein the genome of at least some of the subject’s cells comprises a subset (e.g., all) of the cancer- specific mutations.

[0147] In some embodiments, provided methods comprise obtaining a tumor sample from the subject and using the tumor sample to detect the plurality of candidate mutations (e.g., by sequencing the tumor sample).- 27 - 13241940vlAttorney Docket No. 2013237-1502

[0148] In some embodiments, provided methods comprise (e.g., further) comprising obtaining a normal sample from the subject and using the normal sample (e.g., together with the tumor sample) to detect the plurality of cancer mutations (e.g., by sequencing the normal sample).

[0149] In some embodiments, provided methods comprise sequencing the tumor and / or normal sample (e.g., in replicates).

[0150] In some aspects, the present disclosure provides pharmaceutical compositions comprising a polyribonucleotide encoding (e.g., a polyepitopic vaccine construct comprising) a plurality of neoepitopes, wherein at least a portion of the plurality of neoepitopes are individualized (e.g., patient- specific) neoepitopes corresponding to (e.g., encoded by a nucleotide sequence comprising) a somatic mutation from a population of candidate mutations detected in a tumor sample of a subject and selected for inclusion in the pharmaceutical composition using the method or system of any one of the preceding claims (e.g., each particular neoepitope encoded by a nucleotide sequence comprising a member of the selected subset of candidate mutations determined via a method or system of an aspect or embodiment described herein).

[0151] In some aspects, the present disclosure provides individualized pharmaceutical compositions (e.g., associated with and / or intended for administration to a particular subject) comprising a polyribonucleotide encoding (e.g., a polyepitopic vaccine construct comprising) a plurality of neoepitopes for a particular subject, wherein at least a portion of the neoepitopes are individualized (e.g., patient-specific) neoepitopes corresponding to (e.g., encoded by a nucleotide sequence comprising) a somatic mutation from a population of candidate mutations detected in a tumor sample from the particular subject and selected for inclusion in the pharmaceutical composition using a method or system of an aspect or embodiment described herein.

[0152] In some aspects, the present disclosure provides a population of T-cells, wherein the population of T-cells is capable of specifically binding to a plurality of complexes, wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer-specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, and wherein the plurality of complexes comprises a plurality of neoepitope- 28 - 13241940vlAttorney Docket No. 2013237-1502portions encoded by cancer-specific mutations within the selected subset of candidate mutations determined by a method or system of an aspect or embodiment described herein.

[0153] In some aspects, the present disclosure provides a T-cell capable of specifically binding to a complex, wherein the complex comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer-specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, wherein the one or more cancer-specific mutations are within the selected subset of candidate mutations determined by a method or system of an aspect or embodiment described herein.

[0154] In some aspects, the present disclosure provides a population of T-cells, wherein the population of T-cells is capable of specifically binding to a plurality of complexes, wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer-specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, and wherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer-specific mutations within the selected subset of candidate mutations determined by a method or system of an aspect or embodiment described herein.

[0155] In some aspects, the present disclosure provides a T-cell capable of specifically binding to a complex, wherein the complex comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer-specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, wherein the one or more cancer-specific mutations are within the selected subset of candidate mutations determined by a method or system of an aspect or embodiment described herein.

[0156] In some aspects, the present disclosure provides a T-cell receptor capable of specifically binding to a complex, wherein the complex comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer-specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, wherein the one or more cancer-specific mutations are within the selected subset of candidate mutations determined by a method or system of an aspect or embodiment described herein.- 29 - 13241940vlAttorney Docket No. 2013237-1502

[0157] In some aspects, the present disclosure provides a chimeric antigen receptor (CAR) capable of specifically binding to a complex, wherein the complex comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancerspecific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, wherein the one or more cancer-specific mutations are within the selected subset of candidate mutations determined by a method or system of an aspect or embodiment described herein.

[0158] In some embodiments, at least a portion of the individualized (e.g., patientspecific) neoepitopes correspond to biologically clonal mutations.

[0159] In some embodiments, at least a portion of the individualized (e.g., patientspecific) neoepitopes correspond to mutations determined to have a high likelihood of being biologically clonal and / or highly prevalent (e.g., having been classified as belonging to a particular one of one or more clonal states, e.g., as described herein).

[0160] In some embodiments, individualized (e.g., patient- specific) neoepitopes do not correspond to mutations that are biologically subclonal and present in a minority of (e.g., less than about 50% of) tumor cells (e.g., as measured based on the tumor sample and / or a representative tumor sample taken from the subject) (e.g., having been classified as belonging to a minor subclone state as described herein).

[0161] In some embodiments, none of the individualized (e.g., patient-specific) neoepitopes correspond to mutations that are biologically subclonal and present in a minority of tumor cells (e.g., as measured based on the tumor sample and / or a representative tumor sample taken from the subject) (e.g., having been classified as belonging to a minor subclone state as described herein).

[0162] In some embodiments, at least a portion of the individualized (e.g., patientspecific) neoepitopes correspond to mutations having been classified as belonging to an undetermined state indicative of mutations for which clone type assignments are determined and / or predicted to be unreliable [e.g., indicative of any of (e.g., significant) stochastic noise, local bias, alignment errors, potential errors in estimated parameter values, model assumptions] (e.g., as described herein).

[0163] In some embodiments, at least a portion of the individualized (e.g., patientspecific) neoepitopes correspond mutations that are biologically subclonal, but present in a majority (e.g., greater than 50%) of tumor cells of the subject (e.g., as measured based on the - 30 - 13241940vlAttorney Docket No. 2013237-1502tumor sample and / or a representative tumor sample taken from the subject) (e.g., having been classified as belonging to one of one or more prevalent subclone states, e.g., as described herein).

[0164] In some embodiments, all of the portion of the individualized (e.g., patientspecific) neoepitopes correspond to mutations that are present in a majority of tumor cells of the subject.

[0165] In some embodiments, a population of candidate mutations comprises at least two sub-populations of mutations that are biologically subclonal, but present in a majority of (e.g., greater than 50% of) tumor cells, a first of the at least two sub-populations comprises mutations having cellularity values within a first range and a second of the at least two-subpopulations comprises mutations having cellularity values within a second range, said first range above the second range, and the individualized (e.g., patient-specific) neoepitopes comprise more neoepitopes corresponding to mutations from the first sub-population than the second (e.g., wherein the encoded neoepitopes comprise neoepitopes corresponding to mutations from the first sub-population, but not the second).

[0166] In some embodiments, a population of candidate mutations comprises at least four sub-populations, including: (A) a first sub-population comprising mutations that are biologically clonal and / or present in all or nearly all (e.g., 90% or more; e.g., about 80% or more) tumor cells; (B) a second sub-population comprising mutations that are biologically sub-clonal, but present in a majority of tumor cells at a rate within a first range (e.g., from about 60% to about 80% of tumor cells); (C) a third sub-population comprising mutations that are biologically sub-clonal, but present in a majority of tumor cells at a rate within a second range, the second range below the first range (e.g., from about 50% to 60% of tumor cells); and (D) a fourth sub-population, comprising mutations that are biologically sub-clonal and present in a minority (e.g., less than 50%) of tumor cells.

[0167] In some embodiments, the first sub-population is over-represented in the mutations to which the individualized (e.g., patient- specific) neoepitopes correspond (e.g., in comparison with the other sub-populations) (e.g., wherein the composition includes / encodes more individualized (e.g., patient-specific) neoepitopes corresponding to mutations from the first sub-population than the others; e.g., wherein a fraction of neoepitopes corresponding to mutations from the first sub-population is greater than an overall fraction of candidate mutations within the first sub-population).- 31 - 13241940vlAttorney Docket No. 2013237-1502

[0168] In some embodiments, the second sub-population is over-represented in the mutations to which the individualized (e.g., patient- specific) neoepitopes correspond (e.g., in comparison with the third and fourth sub-populations).

[0169] In some embodiments, substantially all (e.g., each and every) of the individualized (e.g., patient- specific) neoepitopes correspond to mutations from the first subpopulation.

[0170] In some embodiments, substantially all (e.g., each and every) of the individualized (e.g., patient- specific) neoepitopes correspond to mutations from the first and sub-populations.

[0171] In some aspects, the present disclosure provides methods comprising: obtaining a list of candidate TCR sequences for a subject; detecting a plurality of candidate mutations in tumor cells of from the subject and selecting, using a method or system of any one of various aspects and embodiments described herein, e.g., in paragraphs above, a subset of the plurality of candidate mutations; identifying, using the list of candidate TCR sequences and the selected subset of candidate mutations, one or more TCR-mutation pairings, each candidate TCR-mutation pairing comprising a particular TCR sequence of the list of candidate TCR sequences and a particular candidate mutation of the selected subset, wherein the particular candidate TCR sequence is identified as capable of specifically binding to a complex comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising the particular candidate mutation, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells; and using the one or more candidate TCR-mutation pairings to produce a personalized cancer immunotherapy.

[0172] Features of embodiments described with respect to one aspect of the invention may be applied with respect to another aspect of the invention.BRIEF DESCRIPTION OF THE DRAWING

[0173] The foregoing and other objects, aspects, features, and advantages of the present disclosure will become more apparent and better understood by referring to the following description taken in conjunction with the accompanying drawings, in which:- 32 - 13241940vlAttorney Docket No. 2013237-1502

[0174] FIG. 1A is a schematic illustrating correspondence between a collection of cancer cell subpopulations in a bulk tumor and clone type classification of bulk tumor sequence data according to an illustrative embodiment.

[0175] FIG. IB is a schematic illustrating dynamical evolution of cancer cell subpopulations (in, e.g., a bulk tumor), including emergence of certain cancer subpopulations (e.g., a major subpopulation arising due to a driver mutation).

[0176] FIG. 1C is a schematic illustrating another example approach for selecting a subset of candidate mutations for inclusion in a therapeutic construct based on clone type classifications and biological features of the candidate mutations, according to an illustrative embodiment.

[0177] FIG. ID is a block flow diagram showing an example process for obtaining sequencing data from tumor and / or normal sample(s), according to an illustrative embodiment.

[0178] FIG. IE is a schematic illustrating tumor and normal genome characteristics, according to an illustrative embodiment.

[0179] FIG.2A is a block flow diagram showing an example tumor deconvolution process, according to an illustrative embodiment.

[0180] FIG.2B is a block flow diagram showing an example process for using multiple tumor models for obtaining purity estimates and / or estimated purity bounds, according to an illustrative embodiment.

[0181] FIG.2C is a schematic illustrating certain categories of segments within a genome e.g., a tumor genome and / or a normal genome) and approaches for their identification, according to an illustrative embodiment.

[0182] FIG.2D is a schematic illustrating a visualization of segments and their corresponding potential copy number solutions to be determined from sequencing data, according to an illustrative embodiment.

[0183] FIG.2E is a graph and schematic illustrating an approach for computing global cluster density, according to an illustrative embodiment.

[0184] FIG.3 is a block flow diagram showing an example process for detecting candidate mutations from sequencing data, according to an illustrative embodiment.- 33 - 13241940vlAttorney Docket No. 2013237-1502

[0185] FIG.4A is a block flow diagram showing an example process for classifying candidate mutations according to discrete clone types and selecting mutations according to their determined clone type classifications for inclusion in a construct, such as a PCV, TCR, etc., according to an illustrative embodiment.

[0186] FIG.4B is a block flow diagram showing an example process for classifying candidate mutations according to discrete clone types, according to an illustrative embodiment.

[0187] FIG.4C is a block flow diagram illustrating certain steps and criteria used in certain embodiments of clone type classification processes described herein.

[0188] FIG.4D is a block flow diagram illustrating certain steps and criteria arrangements used in certain embodiments of clone type classification processes described herein.

[0189] FIG.4E is a block flow diagram illustrating certain steps and criteria arrangements used in certain embodiments of clone type classification processes described herein.

[0190] FIG.5 is a block flow diagram showing an example process for filtering candidate mutations, according to an illustrative embodiment.

[0191] FIG.6A is a diagram illustrating interrelation between parameters and / or measurements characterizing tumor sample properties, according to an illustrative embodiment.

[0192] FIG.6B is a diagram showing possible purity values for different copy numbers and zygosities, determined for a VAF of 0.2. Highlighted rows in the diagram correspond to balanced SNVs.

[0193] FIG.7 is a block flow diagram of an example tumor deconvolution process for determining a tumor sample purity estimate, according to an illustrative embodiment.

[0194] FIG.8 is a block flow diagram of an example process for identifying and / or selecting particular subsets of segments, according to an illustrative embodiment.

[0195] FIG.9 is a schematic illustrating certain categories of segments within a genome (e.g., a tumor genome and / or a normal genome) and approaches for their identification, according to an illustrative embodiment.- 34 - 13241940vlAttorney Docket No. 2013237-1502

[0196] FIG. 10 is a schematic illustrating certain categories of segments within a genome (e.g., a tumor genome and / or a normal genome) and approaches for their identification, according to an illustrative embodiment.

[0197] FIG. 11 is a schematic with an arrangement of example plots illustrating an example tumor deconvolution workflow, including how estimates of tumor sample properties such as purity, and absolute copy number and allele specific copy number assignments can be used to detect and characterize cancer-specific mutations, such as SNV events, according to an illustrative embodiment.

[0198] FIG. 12 is a block flow diagram of an example process for using multiple tumor models for obtaining purity estimates and / or estimated purity bounds, according to an illustrative embodiment.

[0199] FIG. 13 is a block flow diagram showing an exemplary SNV-based purity estimation process including quality control steps, according to an illustrative embodiment.

[0200] FIG. 14 is a schematic illustrating an example construct encoding selected neoantigens according to an illustrative embodiment.

[0201] FIG. 15 is a block diagram of an exemplary cloud computing environment, used in certain embodiments.

[0202] FIG. 16 is a block diagram of an example computing device and an example mobile computing device used in certain embodiments.

[0203] FIG. 17A is a schematic illustrating tumor and normal genome segments and approaches for identifying certain categories of segments and estimating initial values for auxiliary parameters used in certain embodiments of tumor modelling systems and methods, according to an illustrative embodiment.

[0204] FIG. 17B is a schematic illustrating certain steps in an example tumor deconvolution process, according to an illustrative embodiment.

[0205] FIG. 18A is a bar graph comparing ploidy estimates determined via an example tumor deconvolution approach with experimental measurements obtained from fluorescence activated cell sorting (FACS) and spectral karyotyping (SKY) for various cell lines.- 35 - 13241940vlAttorney Docket No. 2013237-1502

[0206] FIG. 18B is a scatter plot graph comparing average ploidy estimates determined via an example tumor deconvolution approach with values obtained from FACS and SKY measurements.

[0207] FIG. 19 is a heatmap graph showing accuracy of ploidy estimates determined via an example tumor deconvolution approach at varying sample purity and across a plurality of cell lines.

[0208] FIG.20A is a graph plotting accuracy of ploidy estimates determined via a tumor deconvolution approach as a function of purity, showing a series of box and whisker plots at different admixture purities for the n = 27 cell lines.

[0209] FIG.20B is a graph plotting normalized ploidy estimates determined via an example tumor deconvolution approach as compared with a reference.

[0210] FIG.20C is a scatter plot graph showing estimated purities and purity upper bounds determined via an example tumor deconvolution approach.

[0211] FIG.21 A is a heatmap graph showing accuracy of copy number predictions determined via an example tumor deconvolution approach, over various admixture purities and cell lines under test.

[0212] FIG.21B is a graph plotting accuracy of copy number estimates determined via an exemplary tumor deconvolution approach as admixture purity is varied.

[0213] FIG.22 is a block flow diagram illustrating an exemplary clone type classification procedure, according to an illustrative embodiment.

[0214] FIG.23A is a diagram illustrating possible states (e.g., base calls) that may occur in sequencing data, and associated probabilities, without an underlying physical mutation occurring, according to an illustrative embodiment.

[0215] FIG.23B is a diagram illustrating possible states (e.g., base calls) that may occur in sequencing data, and associated probabilities, of a normal allele in a mutated tumor gene, according to an illustrative embodiment.

[0216] FIG.23C is a diagram illustrating possible states (e.g., base calls) that may occur in sequencing data, and associated probabilities, of a mutated tumor allele of a tumor genome, according to an illustrative embodiment.- 36 - 13241940vlAttorney Docket No. 2013237-1502

[0217] FIG.24A is a set of two bar charts comparing mutations detected via a mutation caller with (right) and without (left) filtering to reject rare subclones, according to an illustrative embodiment.

[0218] FIG.24B is a graph showing a histogram of mutations detected using various approaches, according to an illustrative embodiment.

[0219] FIG.24C is a graph showing mutations detected at different variant allele frequencies (VAFs).

[0220] FIG.25A is a screenshot of an example report generated via an example mutation caller and clone type classification system, according to an illustrative embodiment.

[0221] FIG.25B is an example chart generated via an example mutation caller and clone type classification system, according to an illustrative embodiment.

[0222] FIG.26A is a bar chart plotting number of mutations detected and clone type assignments at varying cellularity values.

[0223] FIG.26B is a bar chart plotting percentages of single nucleotide variants (SNVs) classified in different clone types.

[0224] FIG.26C is an example screenshot of summary statistics for mutations detected and classified via an example approach in accordance with various embodiments described herein.

[0225] FIG.27 is a block flow diagram showing an example experimental approach for generating samples with biologically clonal and subclonal mutations at controlled cellularities.

[0226] FIG.28 is a Venn diagram showing mutations associated with two different cell lines.

[0227] FIG.29A is a graph plotting estimated purity against theoretical purity for an engineered sample with controlled mutation properties.

[0228] FIG.29B is a graph plotting estimated ploidy against theoretical purity for an engineered sample.

[0229] FIG.30A is a block flow diagram showing an example process for generating experimental samples with engineered clonal mutations.- 37 - 13241940vlAttorney Docket No. 2013237-1502

[0230] FIG.30B is a block flow diagram showing an example process for generating experimental samples with engineered subclonal mutations.

[0231] FIG.31A is a bar graph plotting absolute copy numbers of experimentally generated defined clonal mutations.

[0232] FIG.31B is a bar graph plotting absolute copy numbers of experimentally generated engineered subclonal mutations.

[0233] FIG.32A is a graph showing distributions of cellularities for a pure (i.e., 100% purity) sample engineered to have controlled mutation properties for experimental validation, according to an illustrative embodiment.

[0234] FIG.32B is a graph showing distributions of cellularities for a 53.2 percent purity sample engineered to have controlled mutation properties for experimental validation, according to an illustrative embodiment.

[0235] FIG.33A is a graph plotting classification sensitivity for an example clone type assignment corresponding to a clonal state (e.g., “CLONAL”), according to an illustrative embodiment.

[0236] FIG.33B is a graph plotting classification sensitivity showing percentage of mutations captured by two example clone types corresponding to a clonal state and one of two possible prevalent subclone states (e.g., “CLONAL” and “PREVALENT-HIGH”), according to an illustrative embodiment.

[0237] FIG.33C is a graph plotting classification sensitivity show percentage of mutations captured by the two example clone types of FIG.25B plus the second of the two possible prevalent subclone states (e.g., “CLONAL”, “PREVALENT-HIGH,” and “PREVALENT-LOW”), according to an illustrative embodiment.

[0238] FIG.33D is a graph plotting classification sensitivity for an example clone type assignment corresponding to a minor subclone state (e.g., “SUBCLONAL”), according to an illustrative embodiment.

[0239] FIG.34A is a graph plotting percentages of experimentally created biologically subclonal (and low prevalence) mutations that were rejected by a rare subclone filter at varying sample purities, according to an illustrative embodiment.- 38 - 13241940vlAttorney Docket No. 2013237-1502

[0240] FIG.34B is a graph plotting percentages of experimentally created biologically subclonal (and low prevalence) mutations that were assigned to a minor subclone state at varying sample purities, according to an illustrative embodiment.

[0241] FIG.35A is a graph plotting percentages of experimentally created biologically subclonal (and low prevalence) mutations that were detected via an example mutation caller, according to an illustrative embodiment.

[0242] FIG.35B is a graph plotting percentages of experimentally created biologically subclonal (and low prevalence) mutations that were detected via an example mutation caller, but rejected by subsequent filtering, according to an illustrative embodiment.

[0243] FIG.36A is a graph plotting percentages of experimentally created biologically subclonal (and low prevalence) mutations that were detected and rejected via a rare subclone filter, according to an illustrative embodiment.

[0244] FIG.36B is a graph plotting total numbers of experimentally created biologically subclonal (and low prevalence) mutations that were detected and classified as low-confidence mutations via one or more filters, according to an illustrative embodiment.

[0245] FIG.37 is a graph and a heatmap both showing cellularities and identifying clone type assignments for an experimental admixture series of a bulk tumor cell line at varying sample purities, according to an illustrative embodiment.

[0246] FIG.38 is a block flow diagram of an example approach for generating experimental samples with controlled mutation properties for validation of mutation detection and clone type classification technologies in accordance with embodiments described herein.

[0247] FIG.39A is a plot showing mutation lineages and corresponding clone type classifications determined via an illustrative embodiment of clone type classification technologies described herein.

[0248] FIG.39B is a schematic showing lineages of various mutations used to evaluate an example clone type classification method described herein.

[0249] FIG.40A is a schematic illustrating an example approach for selecting a subset of candidate mutations for inclusion in a therapeutic construct based on clone type classifications, according to an illustrative embodiment.- 39 - 13241940vlAttorney Docket No. 2013237-1502

[0250] FIG.40B is a schematic illustrating another example approach for selecting a subset of candidate mutations for inclusion in a therapeutic construct based on clone type classifications, according to an illustrative embodiment.

[0251] FIG.40C is a schematic illustrating another example approach for selecting a subset of candidate mutations for inclusion in a therapeutic construct based on clone type classifications, according to an illustrative embodiment.

[0252] FIG.40D is a schematic illustrating another example approach for selecting a subset of candidate mutations for inclusion in a therapeutic construct based on clone type classifications, according to an illustrative embodiment.

[0253] FIG.40E is a schematic illustrating another example approach for selecting a subset of candidate mutations for inclusion in a therapeutic construct based on clone type classifications, according to an illustrative embodiment.

[0254] FIG.40F is a schematic illustrating another example approach for selecting a subset of candidate mutations for inclusion in a therapeutic construct based on clone type classifications, according to an illustrative embodiment.

[0255] FIG.40G is a schematic illustrating another example approach for selecting a subset of candidate mutations for inclusion in a therapeutic construct based on clone type classifications, according to an illustrative embodiment.

[0256] FIGs.41A and 41B show VAF values and clone type classifications for a low ploidy tumor sample.

[0257] FIGs.42A and 42B show VAF values and clone type classifications for a high ploidy tumor sample.

[0258] FIG.42C is a schematic illustrating how two mutations with the same VAF may have different cellularities.

[0259] FIG.43A is a Venn diagram showing reproducibility of clone type classification for a 41% pure fresh frozen tumor sample after tumor enrichment to 83% purity, according to an illustrative embodiment.

[0260] FIG.43B is a probability distribution function of tumor / normal segment count ratios for SNVs in balanced heterozygous segments for the enriched tumor sample (83% pure), according to an illustrative embodiment.-40 - 13241940vlAttorney Docket No. 2013237-1502

[0261] FIG.43C is a CNV cluster diagram, plotting tumor / normal segment count ratios for heterozygous SNPs based on their allele frequency, according to an illustrative embodiment.

[0262] FIG.43D is a histogram showing clone type classification percentages across clone types for SNVs in an enriched tumor sample, according to an illustrative embodiment, and a table summarizing the percentages.

[0263] FIG.43E is a histogram showing estimated cellularities for SNVs with different clone type classifications, according to an illustrative embodiment.

[0264] FIG.43F is a table comparing clone type classifications according to an illustrative embodiment for SNVs in an original tumor sample and an enriched tumor sample.

[0265] FIG.43G is plot depicting a VAF-based clone type classification spectrum, wherein measured VAFs are plotted according to clone type and grouped by absolute copy number (CNmut) and zygosity ( ) (i.e., grouped by mutation node), according to an illustrative embodiment.

[0266] FIG.43H is a plot depicting a cellularity-based clone type classification spectrum, wherein estimated cellularities are plotted according to clone type and grouped by absolute copy number (CNmut) and zygosity ( ) (i.e., grouped by mutation node), according to an illustrative embodiment.

[0267] FIG.44A is a Venn diagram showing reproducibility of clone type classification for a low ploidy 28.2% pure fresh frozen tumor sample after tumor enrichment to 98% purity, according to an illustrative embodiment.

[0268] FIG.44B is a probability distribution function of tumor / normal segment count ratios for balanced heterozygous segments for an enriched tumor sample (98% pure), according to an illustrative embodiment.

[0269] FIG.44C is a CNV cluster diagram, plotting tumor / normal segment count ratios for heterozygous SNPs based on their allele frequency, according to an illustrative embodiment.

[0270] FIG.44D is a histogram showing clone type classification percentages across clone types for SNVs in an enriched tumor sample (98% pure), according to an illustrative embodiment, and a table summarizing the percentages.- 41 - 13241940vlAttorney Docket No. 2013237-1502

[0271] FIG.44E is a histogram showing estimated cellularities for SNVs in an enriched tumor sample (98% pure) with different clone type classifications, according to an illustrative embodiment.

[0272] FIG.44F is a histogram showing estimated cellularities for SNVs in an enriched tumor sample (98% pure) classified as SUBCLONAL, according to an illustrative embodiment.

[0273] FIG.44G is plot depicting a VAF-based clone type classification spectrum for SNVs in an enriched tumor sample (98% pure), wherein measured VAFs are plotted according to clone type and grouped by absolute copy number (CNmut) and zygosity (Cx) (i.e., grouped by mutation node), according to an illustrative embodiment.

[0274] FIG.44H is a plot depicting a cellularity-based clone type classification spectrum for SNVs in an enriched tumor sample (98% pure), wherein estimated cellularities are plotted according clone type and grouped by absolute copy number (CNmut) and zygosity ( ) (i.e., grouped by mutation node), according to an illustrative embodiment.

[0275] FIG.45A is a CNV cluster diagram, plotting tumor / normal segment count ratios for heterozygous SNPs based on their allele frequency for an FFPE non-small cell lung cancer (NSCLC) tumor sample, according to an illustrative embodiment.

[0276] FIG.45B is a histogram showing estimated cellularities for SNVs in an FFPE NSCLC tumor sample with different clone type classifications, according to an illustrative embodiment.

[0277] FIG.45C is a histogram showing clone type classification percentages across clone types for SNVs in an FFPE NSCLC tumor sample according to an illustrative embodiment, and a table summarizing the percentages.

[0278] FIG.45D is a series of plots showing measured versus expected VAF for SNVs in an FFPE NSCLC tumor sample, split by clone type classification, according to an illustrative embodiment.

[0279] FIG.45E is plot depicting a VAF-based clone type classification spectrum for SNVs in an FFPE NSCLC tumor sample, wherein measured VAFs are plotted according to clone type and grouped by absolute copy number (CNmut and zygosity ( ) (i.e., grouped by mutation node), according to an illustrative embodiment.- 42 - 13241940vlAttorney Docket No. 2013237-1502

[0280] FIG.45F is a plot depicting a cellularity-based clone type classification spectrum for SNVs in an FFPE NSCLC tumor sample, wherein estimated cellularities are plotted according to clone type and grouped by absolute copy number (CNmut) and zygosity ( ) (i.e., grouped by mutation node), according to an illustrative embodiment.

[0281] FIG.46A is a CNV cluster diagram, plotting tumor / normal segment count ratios for heterozygous SNPs based on their allele frequency for an FFPE soft tissue tumor sample, according to an illustrative embodiment.

[0282] FIG.46B is a histogram showing estimated cellularities for SNVs an FFPE soft tissue tumor sample with different clone type classifications, according to an illustrative embodiment.

[0283] FIG.46C is a histogram showing clone type classification percentages across clone types for SNVs in an FFPE soft tissue tumor sample, according to an illustrative embodiment, and a table summarizing the percentages.

[0284] FIG.46D is a series of plots showing measured versus expected VAF for SNVs in an FFPE soft tissue tumor sample, split by clone type classification, according to an illustrative embodiment.

[0285] FIG.46E is plot depicting a VAF-based clone type classification spectrum for SNVs in an FFPE soft tissue tumor sample, wherein measured VAFs are plotted according to clone type and grouped by absolute copy number (CNmut) and zygosity (Cx) (i.e., grouped by mutation node), according to an illustrative embodiment.

[0286] FIG.46F is a plot depicting a cellularity-based clone type classification spectrum for SNVs in an FFPE soft tissue tumor sample, wherein estimated cellularities are plotted according to clone type and grouped by absolute copy number (CNmut and zygosity (Cx) (i.e., grouped by mutation node), according to an illustrative embodiment.

[0287] FIG.47A is a CNV cluster diagram, plotting tumor / normal segment count ratios for heterozygous SNPs based on their allele frequency for an FFPE soft tissue sarcoma sample, according to an illustrative embodiment.

[0288] FIG.47B is a histogram showing estimated cellularities for SNVs in an FFPE soft tissue sarcoma sample with different clone type classifications, according to an illustrative embodiment.- 43 - 13241940vlAttorney Docket No. 2013237-1502

[0289] FIG.47C is a histogram showing clone type classification percentages across clone types for SNVs in an FFPE soft tissue sarcoma sample, according to an illustrative embodiment, and a table summarizing the percentages.

[0290] FIG.47D is a series of plots showing measured versus expected VAF for SNVs in an FFPE soft tissue sarcoma sample, split by clone type classification, according to an illustrative embodiment.

[0291] FIG.47E is plot depicting a VAF-based clone type classification spectrum for SNVs in an FFPE soft tissue sarcoma sample, wherein measured VAFs are plotted according to clone type and grouped by absolute copy number (CNmut) and zygosity (Cx) (i.e., grouped by mutation node), according to an illustrative embodiment.

[0292] FIG.47F is a plot depicting a cellularity-based clone type classification spectrum for SNVs in an FFPE soft tissue sarcoma sample, wherein estimated cellularities are plotted according to clone type and grouped by absolute copy number (CNmut) and zygosity ( ) (i.e., grouped by mutation node), according to an illustrative embodiment.

[0293] FIG.48A is a CNV cluster diagram, plotting tumor / normal segment count ratios for heterozygous SNPs based on their allele frequency for an FFPE breast cancer sample, according to an illustrative embodiment.

[0294] FIG.48B is a histogram showing estimated cellularities for SNVs in an FFPE breast cancer sample with different clone type classifications, according to an illustrative embodiment.

[0295] FIG.48C is a histogram showing clone type classification percentages across clone types for SNVs in an FFPE breast cancer sample, according to an illustrative embodiment, and a table summarizing the percentages.

[0296] FIG.48D is a series of plots showing measured versus expected VAF for SNVs in an FFPE breast cancer sample, split by clone type classification, according to an illustrative embodiment.

[0297] FIG.48E is plot depicting a VAF-based clone type classification spectrum for SNVs in an FFPE breast cancer sample, wherein measured VAFs are plotted according to clone type and grouped by absolute copy number (CNmut and zygosity ( ) (i.e., grouped by mutation node), according to an illustrative embodiment.- 44 - 13241940vlAttorney Docket No. 2013237-1502

[0298] FIG.48F is a plot depicting a cellularity-based clone type classification spectrum for SNVs in an FFPE breast cancer sample, wherein estimated cellularities are plotted according to clone type and grouped by absolute copy number (CNmut) and zygosity ( ) (i.e., grouped by mutation node), according to an illustrative embodiment.

[0299] FIG.49A is a table providing parameters estimated by tumor deconvolution and clone type classification methods for four fresh frozen tumor samples, according to an illustrative embodiment.

[0300] FIG.49B is a table providing T cell recognition screening experiment data for four fresh frozen tumor samples, according to an illustrative embodiment.

[0301] FIG.49C is a table providing T cell recognition enrichment data for four fresh frozen tumor samples, according to an illustrative embodiment.

[0302] FIG.49D is a plot depicting distribution of cellularity by clone type classification for a fresh frozen tumor sample, according to an illustrative embodiment.

[0303] FIG.50A is a block flow diagram showing an example process for improving sensitivity and specificity of a neoantigen prediction method (e.g., a method that takes DNA or RNA data from a sample as input) by leveraging methods as disclosed herein to improve the performance of the neoantigen prediction method.

[0304] FIG.50B is a schematic depicting an exemplary method for slicing a tumor tissue sample (e.g., an FFPE or fresh frozen sample) whereby tissue is sliced in an interleaved manner such that alternating slices are used for either DNA or RNA extraction.

[0305] FIG.51 is a block flow diagram depicting a method of providing personalized cancer immunotherapy, according to an illustrative embodiment.

[0306] FIG.52 is a schematic depicting the advantages of methods disclosed herein for identifying and / or selecting neoepitopes that pair with T cell receptors (TCRs) in the same patient, according to an illustrative embodiment, relative to existing generic mutation callers.

[0307] FIG.53 is a block flow diagram depicting a method of providing personalized immunotherapy and / or personalized T cell / TCR immunotherapy, according to an illustrative embodiment.

[0308] FIG.54A is a scatter plot showing estimated purity (y-axis) as a function of simulated purity (x-axis) for cell line sample admixtures (100bp).- 45 - 13241940vlAttorney Docket No. 2013237-1502

[0309] FIG.54B is a scatter plot showing estimated purity (y-axis) as a function of simulated purity (x-axis) for FFPE sample admixtures (100bp).

[0310] FIG.55A is a scatter plot showing estimated ploidy as a function of simulated histological tumor content (HTC) for cell line sample admixtures (100bp), according to an illustrative embodiment.

[0311] FIG.55B is a scatter plot showing estimated ploidy as a function of simulated histological tumor content (HTC) for FFPE sample admixtures (100bp), according to an illustrative embodiment.

[0312] FIG.56 is a bar plot showing estimated ploidy according to an illustrative embodiment compared to ploidy estimated with FACS for raw cell line samples (100bp).

[0313] FIG.57 is a bar plot showing estimated ploidy according to an illustrative embodiment, compared to ploidy estimated with FACs and SKY for raw cell line samples (50bp).

[0314] FIG.58A is a box and whisker plot showing sensitivity (assessed by relative number of mutations present in reference) as a function of simulated HTC (sHTC) for cell line sample admixtures (100bp), according to an illustrative embodiment.

[0315] FIG.58B is a box and whisker plot showing sensitivity (as assessed by relative number of mutations present in reference) as a function of simulated HTC (sHTC) for FFPE sample admixtures (100bp), according to an illustrative embodiment.

[0316] FIG.58C is a box and whisker plot showing uniqueness (as assessed by relative number of mutations absent in reference) as a function of sHTC for cell line sample admixtures (100bp), according to an illustrative embodiment.

[0317] FIG.58D is a box and whisker plot showing uniqueness (as assessed by relative number of filtered mutations absent in reference) as a function of sHTC for FFPE sample admixtures (100bp), according to an illustrative embodiment.

[0318] FIG.59A is a box and whisker plot showing sensitivity (as assessed by relative number of mutations present in reference) as a function of sHTC for cell line sample admixtures (50bp), according to an illustrative embodiment.

[0319] FIG.59B is a box and whisker plot showing uniqueness (as assessed by relative number of mutations absent in reference) as a function of sHTC for cell line sample admixtures (50bp), according to an illustrative embodiment.- 46 - 13241940vlAttorney Docket No. 2013237-1502

[0320] FIG.60A is a plot showing sensitivity for all mutations (as assessed by relative number of mutations present in reference) as a function of simulated purity for FFPE sample admixtures (100bp), according to an illustrative embodiment.

[0321] FIG.60B is a plot showing sensitivity for mutations in exon coding regions (as assessed by relative number of mutations present in reference) as a function of simulated purity for FFPE sample admixtures (100bp), according to an illustrative embodiment.

[0322] FIG.60C is a plot showing sensitivity for mutations in exon coding regions determined to have a CLONAL clone type (as assessed by relative number of mutations present in reference) as a function of simulated purity for FFPE sample admixtures (100bp), according to an illustrative embodiment.

[0323] FIG.61A is a bar plot showing number of detected mutations in raw FFPE samples (100bp) by clone type classification, with numbers on top of bars indicating total number of detected mutations per sample, and percentages indicating proportion of mutations classified as CLONAL, according to an illustrative embodiment.

[0324] FIG.61B is a bar plot showing number of detected mutations in raw cell line samples (100bp) by clone type classification, with numbers on top of bars indicating total number of detected mutations per sample, and percentages overlaid on bars indicating proportion of mutations classified as CLONAL, according to an illustrative embodiment.

[0325] FIG.62 is a bar plot showing number of detected mutations in raw cell line samples (50bp) by clone type classification, with numbers on top of bars indicating total number of detected mutations per sample, and percentages overlaid on bars indicating proportion of mutations classified as CLONAL, according to an illustrative embodiment.

[0326] FIG.63 is a series of heatmaps showing clone type classification (“Clonality State”) of mutations detected in certain FFPE raw samples and admixtures (all 100bp), wherein each panel corresponds to data from a patient sample, according to an illustrative embodiment.

[0327] FIG.64 is a series of heatmaps showing clone type classification (“Clonality State”) of mutations detected in certain raw cell line samples and admixtures, wherein each panel corresponds to data from a cell line, according to an illustrative embodiment.

[0328] FIG.65 is a series of heatmaps showing clone type classification (“Clonality State”) of mutations detected in certain FFPE raw samples and admixtures (all 50bp),- 47 - 13241940vlAttorney Docket No. 2013237-1502wherein each panel corresponds to data from a patient sample, according to an illustrative embodiment.

[0329] FIG.66A is a scatter plot showing CLONAL scores as a function of simulated purity for FFPE sample admixtures (100bp), according to an illustrative embodiment.

[0330] FIG.66B is a scatter plot showing Clonality Scores as a function of simulated purity for FFPE sample admixtures (100bp), according to an illustrative embodiment.

[0331] FIG.66C is a scatter plot showing CLONAL scores as a function of simulated purity for cell line sample admixtures (100bp), according to an illustrative embodiment.

[0332] FIG.66D is a scatter plot showing Clonality Scores as a function of simulated purity for cell line sample admixtures (100bp), according to an illustrative embodiment.

[0333] FIG.67A is a scatter plot showing CLONAL scores as a function of simulated purity for cell line sample admixtures (50bp), according to an illustrative embodiment.

[0334] FIG.67B is a box and whisker plot showing CLONAL scores as a function of simulated purity for cell line sample admixtures (100bp), according to an illustrative embodiment.

[0335] FIG.67C is a scatter plot showing Clonality Scores as a function of simulated purity for cell line sample admixtures (50bp), according to an illustrative embodiment.

[0336] FIG.67D is a box and whisker plot showing Clonality Scores as a function of simulated purity for cell line sample admixtures (50bp), according to an illustrative embodiment.

[0337] FIG.68A is a bar plot showing SNVs confirmed to have a de novo response across patients, according to an illustrative embodiment.

[0338] FIG.68B is a bar plot showing SNVs confirmed to have a de novo response across patients, according to an illustrative embodiment.

[0339] FIG.68C is a bar plot showing SNVs confirmed to have a de novo response across patients, according to an illustrative embodiment.

[0340] FIG.69A is a bar plot showing SNVs confirmed to have a de novo response across patients, according to an illustrative embodiment.

[0341] FIG.69B is a bar plot showing SNVs confirmed to have a de novo response across patients, according to an illustrative embodiment.- 48 - 13241940vlAttorney Docket No. 2013237-1502

[0342] FIG.70A is a bar plot showing SNVs for which a TCR response was raised during an ex vivo priming process, according to an illustrative embodiment.

[0343] FIG.70B is a bar plot showing SNVs for which a TCR response was raised during an ex vivo priming process, according to an illustrative embodiment.

[0344] FIG.70C is a bar plot showing SNVs for which a TCR response was not raised during an ex vivo priming process, according to an illustrative embodiment.

[0345] FIG.70D is a bar plot showing SNVs for which a TCR response was not raised during an ex vivo priming process, according to an illustrative embodiment.

[0346] FIG.71 is a schematic with an arrangement of example plots illustrating an example tumor deconvolution workflow, including how estimates of tumor sample properties such as purity, and absolute copy number and allele specific copy number assignments can be used to detect and characterize cancer-specific mutations, such as SNV events, according to an illustrative embodiment.

[0347] FIG.72 is an example of a graphical representation of tumor genome segments with heterozygous SNP balanced state calculation results overlaid thereon, according to an illustrative embodiment.

[0348] FIG.73 is a set of graphs providing results of various steps in an example tumor deconvolution process for two melanoma samples and an ovarian cancer sample.

[0349] FIG.74 is a set of graphs demonstrating use of a variance stabilizing transformation to produce a normal distribution of residual errors, according to an illustrative embodiment.

[0350] FIG.75 a graph showing an example CNV cluster plot.

[0351] FIG.76A is a graph showing an example CNV cluster plot determined for an ovarian cancer sample, according to an illustrative embodiment.

[0352] FIG.76B is a graph showing an example CNV cluster plot determined for an ovarian cancer sample, according to an illustrative embodiment.

[0353] FIG.76C is a graph showing an example CNV cluster plot determined for an ovarian cancer sample, according to an illustrative embodiment.

[0354] FIG.76D is a graph showing an example CNV cluster plot determined for an ovarian cancer sample, according to an illustrative embodiment.- 49 - 13241940vlAttorney Docket No. 2013237-1502

[0355] FIG.76E is a graph showing an example CNV cluster plot determined for an ovarian cancer sample, according to an illustrative embodiment.

[0356] FIG.76F is a graph showing an example CNV cluster plot determined for an ovarian cancer sample, according to an illustrative embodiment.

[0357] FIG.77A is a graph showing an example CNV cluster plot determined for a melanoma sample, according to an illustrative embodiment.

[0358] FIG.77B is a graph showing an example CNV cluster plot determined for a melanoma sample, according to an illustrative embodiment.

[0359] FIG.77C is a graph showing an example CNV cluster plot determined for a glioblastoma sample, according to an illustrative embodiment.

[0360] FIG.77D is a graph showing an example CNV cluster plot determined for glioblastoma sample, according to an illustrative embodiment.

[0361] FIG.77E is a graph showing an example CNV cluster plot determined for a breast cancer sample, according to an illustrative embodiment.

[0362] FIG.77F is a graph showing an example CNV cluster plot determined for a breast cancer sample, according to an illustrative embodiment.

[0363] FIG.78A is a graph showing an example CNV cluster plot determined for a melanoma sample, according to an illustrative embodiment.

[0364] FIG.78B is a graph showing an example fan plot determined for the Melanoma sample shown in FIG.78A according to an illustrative embodiment.

[0365] FIG.79A is a graph showing an example CNV cluster plot with rectangular decision boundaries, overlaid according to an illustrative embodiment.

[0366] FIG.79B is a graph showing an example CNV cluster plot with rectangular decision boundaries overlaid, according to an illustrative embodiment.

[0367] FIG.80 is a set of graphs showing results at various steps in an example CNV-based purity estimation for a melanoma sample, according to an illustrative embodiment.

[0368] FIG.81 is a set of graphs showing results at various stages in an example CNV-based purity estimation for a melanoma sample, according to an illustrative embodiment.- 50 - 13241940vlAttorney Docket No. 2013237-1502

[0369] FIG.82 a set of graphs showing example probability distribution functions for allele frequencies and residual errors used to determine a '-slalislic value, according to an illustrative embodiment.

[0370] FIG.83 is a graph showing a CNV cluster plot with outliers identified for exclusion from a '-slalislic calculation marked, according to an illustrative embodiment.

[0371] FIG.84 is a schematic illustrating a cluster density metric, according to an illustrative embodiment.

[0372] FIG.85 is a graph showing a decision plane used in a multivariate classifier, according to an illustrative embodiment.

[0373] FIG.86A is a graphical display showing results of an error correction procedure that identifies and corrects certain errors in copy number assignments, according to an illustrative embodiment.

[0374] FIG.86B is a graphical display showing results of an error correction procedure that identifies and corrects certain errors in copy number assignments, according to an illustrative embodiment.

[0375] FIG.87 is a set of graphs demonstrating performance of an error correction procedure, used in certain embodiments.

[0376] FIG.88A is a graph plotting parity error rates as a function of primary copy number (CApc) for a melanoma sample, according to an illustrative embodiment.

[0377] FIG.88B is a graph plotting parity error rates as a function of primary copy number (CApc) for a melanoma sample, according to an illustrative embodiment.

[0378] FIG.89 is a block flow diagram showing a decision tree for selecting between various purity estimation results, according to an illustrative embodiment.

[0379] FIG.90 is a block flow diagram showing an exemplary process for determining a putative mutation (e.g., an initial event) based on extreme statistics.

[0380] FIG.91A is a line graph showing true positive rate as a function of coverage depth for an exemplary extreme event detector (EED) classifier and an exemplary binomial classifier for simulated heterozygous SNPs.

[0381] FIG.91B is a line graph showing a false discovery rate (number of wrong genotype detection events over total detection events) as a function of coverage depth for an - 51 - 13241940vlAttorney Docket No. 2013237-1502exemplary EED classifier and an exemplary binomial classifier for simulated heterozygous SNPs.

[0382] FIG.91C is a line graph showing false discovery rate (false positives over total detection events) as a function of coverage for an exemplary EED classifier and an exemplary binomial classifier for simulated wild-type sites.

[0383] FIG.92A is a graph showing estimation of noise based on normal versus normal analysis of singletons according to an illustrative embodiment, Mutect2, or Strelka.

[0384] FIG.92B is a graph showing estimation of noise based on normal versus normal analysis of replicates or merged singletons according to an illustrative embodiment, Mutect2, or Strelka.

[0385] FIG.93A is a graph showing predicted PPV as a function of TMB based on normal versus normal analysis of singletons according to an illustrative embodiment, Mutect2, or Strelka (including confidence intervals).

[0386] FIG.93B is a graph showing predicted PPV as a function of TMB based on normal versus normal analysis of replicates or merged singletons according to an illustrative embodiment, Mutect2, or Strelka (including confidence intervals).

[0387] FIG.94A is a bar graph showing number of SNVs detected across permutations according to an illustrative embodiment or Mutect2.

[0388] FIG.94B is a bar graph showing number of SNVs (shared and unique variants) detected across permutations according to an illustrative embodiment or Mutect2.

[0389] FIG.94C is a histogram showing cellularity distributions for SNVs detected by Mutect2 only (“Mutect2 unique”), an illustrative embodiment only (“Embodiment unique”), or both (“Shared”).

[0390] FIG.95A is a scatter plot with unity line plotting estimated purity as a function of simulated purity across patient samples, according to an illustrative embodiment.

[0391] FIG.95B is a scatter plot with unity line plotting estimated purity as a function of simulated HTC across patient samples, according to an illustrative embodiment.

[0392] FIG.95C is a scatter plot with unity line plotting simulated HTC as a function of simulated purity across patient samples, according to an illustrative embodiment.- 52 - 13241940vlAttorney Docket No. 2013237-1502

[0393] FIG.96A is a box and whisker plot showing relative sensitivity percentage as a function of simulated HTC across patient samples according to an illustrative embodiment in singleton mode.

[0394] FIG.96B is a box and whisker plot showing relative sensitivity percentage as a function of simulated HTC across patient samples according to Mutect2 in singleton mode.

[0395] FIG.96C is a box and whisker plot showing relative sensitivity percentage as a function of simulated HTC across patient samples according to Strelka2 in singleton mode.

[0396] FIG.97A is a plot showing relative sensitivity percentage as a function of simulated purity across patient samples according to an illustrative embodiment in singleton mode.

[0397] FIG.97B is a plot showing relative sensitivity percentage as a function of simulated purity across patient samples according to Mutect2 in singleton mode.

[0398] FIG.97C is a plot showing relative sensitivity percentage as a function of simulated purity across patient samples according to Strelka2 in singleton mode.

[0399] FIG.98A is a box and whisker plot showing relative sensitivity percentage as a function of simulated HTC across patient samples according to an illustrative embodiment in replicate mode.

[0400] FIG.98B is a box and whisker plot showing relative sensitivity percentage as a function of simulated HTC across patient samples according to Mutect2 in replicate mode.

[0401] FIG.98C is a box and whisker plot showing relative sensitivity percentage as a function of simulated HTC across patient samples according to Strelka2 in singleton merged mode.

[0402] FIG.99A is a plot showing relative sensitivity percentage as a function of simulated purity across patient samples according to an illustrative embodiment in replicate mode.

[0403] FIG.99B is a plot showing relative sensitivity percentage as a function of simulated purity across patient samples according to Mutect2 in replicate mode.

[0404] FIG.99C is a plot showing relative sensitivity percentage as a function of simulated purity across patient samples according to Strelka2 in singleton merged mode.- 53 - 13241940vlAttorney Docket No. 2013237-1502

[0405] FIG. 100A is a graph showing relative sensitivity percentage (mean + / -standard deviation) as a function of simulated HTC according to an illustrative embodiment in singleton mode, Mutect2 in singleton mode, or Strelka2 in singleton mode.

[0406] FIG. 100B is a graph showing relative sensitivity percentage (mean + / -standard deviation) as a function of simulated HTC according to an illustrative embodiment in replicate mode, Mutect2 in replicate mode, or Strelka2 in singleton merged mode.

[0407] FIG. 101A is a graph showing relative sensitivity percentage (mean + / -standard deviation) as a function of simulated purity according to an illustrative embodiment in singleton mode, Mutect2 in singleton mode, or Strelka2 in singleton mode.

[0408] FIG. 101B is a graph showing relative sensitivity percentage (mean + / -standard deviation) as a function of simulated purity according to an illustrative embodiment in replicate mode, Mutect2 in replicate mode, or Strelka2 in singleton merged mode.

[0409] FIG. 102A is a box and whisker plot showing number of mutations unique to admixture as a function of simulated HTC across patient samples according to an illustrative embodiment in singleton mode.

[0410] FIG. 102B is a box and whisker plot showing number of mutations unique to admixture as a function of simulated HTC across patient samples according to Mutect2 in singleton mode.

[0411] FIG. 102C is a box and whisker plot showing number of mutations unique to admixture as a function of simulated HTC across patient samples according to Strelka2 in singleton mode.

[0412] FIG. 103A is a box and whisker plot showing number of mutations unique to admixture as a function of simulated HTC across patient samples according to an illustrative embodiment in replicate mode.

[0413] FIG. 103B is a box and whisker plot showing number of mutations unique to admixture as a function of simulated HTC across patient samples according to Mutect2 in replicate mode.

[0414] FIG. 103C is a box and whisker plot showing number of mutations unique to admixture as a function of simulated HTC across patient samples according to Strelka2 in singleton merged mode.- 54 - 13241940vlAttorney Docket No. 2013237-1502

[0415] FIG. 104A is a graph showing number of mutations unique to admixture (mean + / - standard deviation) as a function of simulated HTC according to an illustrative embodiment in singleton mode, Mutect2 in singleton mode, or Strelka2 in singleton mode.

[0416] FIG. 104B is a graph showing number of mutations unique to admixture (mean + / - standard deviation) as a function of simulated HTC according to an illustrative embodiment in replicate mode, Mutect2 in replicate mode, or Strelka2 in singleton merged mode.

[0417] FIG. 105A is a box and whisker plot showing a Jaccard index as a function of simulated HTC across patient samples according to an illustrative embodiment in singleton mode.

[0418] FIG. 105B is a box and whisker plot showing a Jaccard index as a function of simulated HTC across patient samples according to Mutect2 in singleton mode.

[0419] FIG. 105C is a box and whisker plot showing a Jaccard index as a function of simulated HTC across patient samples according to Strelka2 in singleton mode.

[0420] FIG. 106A is a plot showing a Jaccard index as a function of simulated purity across patient samples according to an illustrative embodiment in singleton mode.

[0421] FIG. 106B is a plot showing a Jaccard index as a function of simulated purity across patient samples according to Mutect2 in singleton mode.

[0422] FIG. 106C is a plot showing a Jaccard index as a function of simulated purity across patient samples according to Strelka2 in singleton mode.

[0423] FIG. 107A is a box and whisker plot showing number of mutations unique to one set as a function of simulated HTC across patient samples according to an illustrative embodiment in singleton mode.

[0424] FIG. 107B is a box and whisker plot showing number of mutations unique to one set as a function of simulated HTC across patient samples according to Mutect2 in singleton mode.

[0425] FIG. 107C is a box and whisker plot showing number of mutations unique to one set as a function of simulated HTC across patient samples according to Strelka2 in singleton mode.- 55 - 13241940vlAttorney Docket No. 2013237-1502

[0426] FIG. 108A is a plot showing number of mutations unique to one set as a function of simulated HTC across patient samples according to an illustrative embodiment in singleton mode.

[0427] FIG. 108B is plot showing number of mutations unique to one set as a function of simulated HTC across patient samples according to Mutect2 in singleton mode.

[0428] FIG. 108C is a plot showing number of mutations unique to one set as a function of simulated HTC across patient samples according to Strelka2 in singleton mode.

[0429] FIG. 109 is a bar plot showing number of mutations unique to one set of cross replicates (average and standard deviation) according to an illustrative embodiment, Mutect2, and Strelka2, based on cross replicate analysis for nine patients.

[0430] FIG. 110A is a box and whisker plot showing relative sensitivity percentage as a function of simulated HTC across patient samples according to an illustrative embodiment in replicate mode.

[0431] FIG. HOB is a box and whisker plot showing relative sensitivity percentage as a function of simulated HTC across patient samples according to Mutect2 in replicate mode filtered with VAF >=0.05.

[0432] FIG. HOC is a box and whisker plot showing relative sensitivity percentage as a function of simulated HTC across patient samples according to Strelka2 in singleton merged mode filtered with VAF >=0.05.

[0433] FIG. 111A is a plot showing relative sensitivity percentage as a function of simulated purity across patient samples according to an illustrative embodiment in replicate mode.

[0434] FIG. H1B is a plot showing relative sensitivity percentage as a function of simulated purity across patient samples according to Mutect2 in replicate mode filtered with VAF >=0.05.

[0435] FIG. H1C is a plot showing relative sensitivity percentage as a function of simulated purity across patient samples according to Strelka2 in singleton merged mode filtered with VAF >=0.05.

[0436] FIG. 112 is a graph showing relative sensitivity percentage (mean + / - standard deviation) as a function of simulated purity according to an illustrative embodiment in- 56 - 13241940vlAttorney Docket No. 2013237-1502replicate mode, Mutect2 in replicate mode filtered with VAF>=0.05, or Strelka2 in singleton merged mode filtered with VAF>=0.05.

[0437] The features and advantages of the present disclosure will become more apparent from the detailed description set forth below when taken in conjunction with the drawings, in which like reference characters identify corresponding elements throughout. In the drawings, like reference numbers generally indicate identical, functionally similar, and / or structurally similar elements.CERTAIN DEFINITIONS

[0438] About-. The term “about”, when used herein in reference to a value, refers to a value that is similar, in context to the referenced value. In general, those skilled in the art, familiar with the context, will appreciate the relevant degree of variance encompassed by “about” in that context. For example, in some embodiments, the term “about” may encompass a range of values that within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or less of the referred value.

[0439] Absolute Copy Number'. The term “absolute copy number” as used herein refers to a number of physical copies of a particular segment within a cell comprising a particular genome. For example, an absolute copy number of a segment in the normal genome can be defined as the number of physical copies of the given segment in a healthy cell. For example, an absolute copy number of a segment in the tumor genome can be defined as the number of physical copies of the given segment in a tumor cell. In certain embodiments, if only a part of a segment is amplified or deleted in a genome, then such a partial copy of the segment can either be counted as a copy of the segment or not counted as a copy of the segment. In certain embodiments, copies of the segment spanning less than 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10% or 5% of the segment length can be ignored.

[0440] Agent-. As used herein, the term “agent,” may refer to a physical entity. In some embodiments, an agent may be characterized by a particular feature and / or effect. For example, as used herein, the term “therapeutic agent” refers to a physical entity has a therapeutic effect and / or elicits a desired biological and / or pharmacological effect. In some embodiments, an agent may be a compound, molecule, or entity of any chemical class including, for example, a small molecule, polypeptide, nucleic acid, saccharide, lipid, metal, or a combination or complex thereof. In some embodiments, part or all of an agent may be - 57 - 13241940vlAttorney Docket No. 2013237-1502depicted herein as a chemical structure, or may be described using chemical nomenclature and / or with reference to general principles of organic chemistry, e.g., in accordance with the Periodic Table of Elements, CAS version, Handbook of Chemistry and Physics, 75th Ed; “Organic Chemistry”, Thomas Sorrell, University Science Books, Sausalito: 1999, and / or “March’s Advanced Organic Chemistry”, 5th Ed., Ed.: Smith, M. B. and March, J., John Wiley & Sons, New York: 2001, the entire contents of which are hereby incorporated by reference. Unless otherwise stated or clear from context, chemical structures depicted herein may be considered to reference or include one or more, or all, stereoisomeric (e.g., enantiomeric or diastereomeric) forms of the structure, and / or one or more, or all, geometric or conformational isomeric forms of the structure. For example, unless otherwise indicated or clear, both R and S configurations of a stereocenter may be contemplated in embodiments of the disclosure. In some embodiments, a compound may be described and / or utilized as a particular single stereochemical isomer; alternatively or additionally, in some embodiments, such a compound may be described and / or utilized as a combination (e.g., a mixture) of one or more enantiomeric (e.g., diastereomeric) forms (e.g., as a racemic preparation).Analogously, in some embodiments, a single geometric isomer may be described and / or utilized; in some embodiments, a combination (e.g., a mixture) of geometric (or conformational) isomers may be described and / or utilized. Unless otherwise stated or clear from context, all tautomeric forms of provided compounds are within the scope of the disclosure. Still further, unless otherwise indicated or clear from context, in some embodiments, a particular chemical compound (e.g., as may be represented by a depicted chemical structure) may be described and / or utilized in an alternative isotopic form - i.e., in a form in which one or more atoms is isotopically altered (e.g., so that a hydrogen is replaced by deuterium or tritium, and / or a carbon is replaced by 13C- or 14C-. Thus, in some embodiments, a particular compound may be described and / or utilized as or in an isotopically enriched preparation.

[0441] Allele-Specific Copy Number: As used herein, the term “allele-specific copy number” refers to a number of physical copies of a particular allele within a cell comprising a particular genome. In certain embodiments, for example, a heterozygous segment comprises a SNP. A cell comprising a genome with the heterozygous segment may comprise zero, one, or more copies of a first (e.g., maternal) allele and zero, one, or more copies of a second (e.g., paternal) allele. For example, a balanced heterozygous segment in a normal diploid genome may comprise distinguishable maternal and paternal alleles, each having an allele- specific - 58 - 13241940vlAttorney Docket No. 2013237-1502copy number of one. In certain embodiments, a corresponding heterozygous segment within a tumor genome may not have zero, one, or more copies of the paternal and maternal alleles, e.g., as a result of copy number variation (CNV) events that may occur in cancer cells. For example, a loss of heterozygosity (LOH) may result in a deletion of copies of one allele (e.g., a maternal or paternal allele), such that an allele-specific copy number for the deleted allele is zero and, if a single copy of the other allele remains, its allele- specific copy number is one. In certain embodiments, duplication events may produce other allele-specific copy numbers, greater than one for one or both alleles.

[0442] Amino acid: In its broadest sense, as used herein, the term “amino acid” refers to a compound and / or substance that can be, is, or has been incorporated into a polypeptide chain, e.g., through formation of one or more peptide bonds. In some embodiments, an amino acid has the general structure H2N-C(H)(R)-COOH. In some embodiments, an amino acid is a naturally-occurring amino acid. In some embodiments, an amino acid is a non-natural amino acid; in some embodiments, an amino acid is a D-amino acid; in some embodiments, an amino acid is an L-amino acid. “Standard amino acid” refers to any of the twenty standard L-amino acids commonly found in naturally occurring peptides. “Nonstandard amino acid” refers to any amino acid, other than the standard amino acids, regardless of whether it is prepared synthetically or obtained from a natural source. In some embodiments, an amino acid, including a carboxy- and / or amino-terminal amino acid in a polypeptide, can contain a structural modification as compared with the general structure above. For example, in some embodiments, an amino acid may be modified by methylation, amidation, acetylation, pegylation, glycosylation, phosphorylation, and / or substitution (e.g., of the amino group, the carboxylic acid group, one or more protons, and / or the hydroxyl group) as compared with the general structure. In some embodiments, such modification may, for example, alter the circulating half-life of a polypeptide containing the modified amino acid as compared with one containing an otherwise identical unmodified amino acid. In some embodiments, such modification does not significantly alter a relevant activity of a polypeptide containing the modified amino acid, as compared with one containing an otherwise identical unmodified amino acid. As will be clear from context, in some embodiments, the term “amino acid” may be used to refer to a free amino acid; in some embodiments it may be used to refer to an amino acid residue of a polypeptide.

[0443] Antigen: term “antigen”, as used herein, refers to (i) an agent that elicits an immune response; and / or (ii) an agent that binds to a T cell receptor (e.g., when presented by - 59 - 13241940vlAttorney Docket No. 2013237-1502an MHC molecule) or to an antibody. In some embodiments, an antigen elicits a humoral response (e.g., including production of antigen-specific antibodies); in some embodiments, an elicits a cellular response (e.g., involving T-cells whose receptors specifically interact with the antigen). In some embodiments, an antigen binds to an antibody and may or may not induce a particular physiological response in an organism. In general, an antigen may be or include any chemical entity such as, for example, a small molecule, a nucleic acid, a polypeptide, a carbohydrate, a lipid, a polymer [in some embodiments other than a biologic polymer (e.g., other than a nucleic acid or amino acid polymer)] etc. In some embodiments, an antigen is or comprises a polypeptide. In some embodiments, an antigen is or comprises a glycan. Those of ordinary skill in the art will appreciate that, in general, an antigen may be provided in isolated or pure form, or alternatively may be provided in crude form (e.g., together with other materials, for example in an extract such as a cellular extract or other relatively crude preparation of an antigen-containing source). In some embodiments, antigens utilized in accordance with the present invention are provided in a crude form. In some embodiments, an antigen is a recombinant antigen.

[0444] Biologically clonal mutation(s) and biologically subclonal mutation(s): As used herein, the terms “biologically clonal” and “biologically subclonal” when used in reference to mutations, such as cancer mutations, are used to specify whether a particular mutations or group of mutations are physically clonal or subclonal. In certain embodiments, a biologically clonal mutation is a mutation that is present in all tumor cells of a tumor sample or biopsy. In certain embodiments, a biologically subclonal mutation is a mutation that is not present in all tumor cells of a tumor sample or biopsy. The use of the adjective “biologically” is used to make clear that the terms “biologically clonal” and “biologically subclonal” refer to the actual physical character of a given mutation, which may or may not be known. The terms “biologically clonal” and “biologically subclonal” thus contrast with the terms “clone type”, “clonal state”, “prevalent subclone state”, and “minor subclone state”, described below, which refer to clonality classification states that are, e.g., labels, determined for (e.g., assigned to) a given mutation.

[0445] Cancer-. The term “cancer” is used herein to generally refer to a disease or condition in which cells of a tissue of interest exhibit relatively abnormal, uncontrolled, and / or autonomous growth, so that they exhibit an aberrant growth phenotype characterized by a significant loss of control of cell proliferation. In some embodiments, cancer may comprise cells that are precancerous e.g., benign), malignant, pre-metastatic, metastatic, - 60 - 13241940vlAttorney Docket No. 2013237-1502and / or non- metastatic. In some embodiments, cancer may be characterized by a solid tumor. In some embodiments, cancer may be characterized by a hematologic tumor. In general, examples of different types of cancers known in the art include, for example, triple negative breast cancer (TNBC), hematopoietic cancers including leukemias, lymphomas (Hodgkin’s and non-Hodgkin’s), myelomas and myeloproliferative disorders; sarcomas, melanomas, adenomas, carcinomas of solid tissue, squamous cell carcinomas of the mouth, throat, larynx, and lung, liver cancer, genitourinary cancers such as prostate, cervical, bladder, uterine, and endometrial cancer and renal cell carcinomas, bone cancer, pancreatic cancer, skin cancer, cutaneous or intraocular melanoma, cancer of the endocrine system, cancer of the thyroid gland, cancer of the parathyroid gland, head and neck cancers, ovarian cancer, breast cancer, glioblastomas, colorectal cancer, gastro-intestinal cancers and nervous system cancers, benign lesions such as papillomas, and the like.

[0446] Clone type: As used herein the term “clone type” is used to refer to certain discrete states or classes that may be determined for and assigned to a given mutation, for example, based on evaluation of various criteria. Different clone types may aim to capture, for example, a likelihood that a given mutation is biologically clonal or biologically subclonal, as well as, for example, a prevalence of the given mutation. In certain embodiments, a likelihood that a given mutation is biologically clonal may be assessed based on various metrics or tests, including, without limitation, statistical hypothesis tests, computed probabilities (e.g., of clonality or sub-clonality, such as a posteriori probability of clonality), etc. For example, in certain embodiments, a likelihood of whether a given mutation is biologically clonal may be assessed using metrics, such as P- values, determined using hypotheses tests, that quantify probabilities of whether for a given data (such as sequencing data) a hypothesis of clonality can be rejected. For example, in certain embodiments, sequencing data may be used to determine an observed cellularity of a given mutation (e.g., a fraction of tumor cells harboring the given mutation). Although, nominally, biologically clonal mutations have a cellularity of 1 (i.e., they are present in all tumor cells), factors, such as measurement error, sample purity, stochastic noise, etc., may lead to observed cellularity values below one, even for biologically clonal mutations. Accordingly, in certain embodiments, a P- value representing a probability of observing a cellularity below one for the given mutation (e.g., given a null hypothesis that the mutation is biologically clonal) may be determined. In this way, for example, a low P- value, e.g., below a threshold, may indicate that it is unlikely that the null hypothesis (of clonality) is true, and it can be - 61 - 13241940vlAttorney Docket No. 2013237-1502rejected (hence unlikely that a mutation is clonal), whereas a higher P- value may indicate that there is a good chance that the observed cellularity, below 1, is due to practical factors, such as measurement error, sample purity, stochastic noise, and the like. In certain embodiments, prevalence may be measured by parameters, such as cellularity, e.g., a determined fraction of tumor cells that harbor a given mutation. Whereas the terms biologically clonal and biologically subclonal are used to refer to underlying physical characteristics of a given mutation, clone types are labels, which aim to classify mutations according to measured physical properties, determined, e.g., based on sequencing data. In certain embodiments, clone types include clonal states, which aim to capture or label mutations that are highly likely to be biologically clonal and / or highly prevalent. In certain embodiments, clone types include multiple clonal classification states, for example, reflecting differing levels of certainty and / or likelihoods that a given mutation is biologically clonal and / or prevalence. In certain embodiments, clone types include one or more prevalent subclone states, capturing mutations that, while biologically subclonal, are prevalent at high rates (e.g., above 50%) in tumor cells. In certain embodiments, clone types include a minor subclonal state, which may be assigned to mutations that are determined to be likely subclonal and rare (e.g., occurring in 50% or less, 45% or less, %40 or less, %35 or less, %25 or less cancer cells).

[0447] Comparable: As used herein, the term “comparable” refers to two or more agents, entities, situations, sets of conditions, etc., that may not be identical to one another but that are sufficiently similar to permit comparison therebetween so that one skilled in the art will appreciate that conclusions may reasonably be drawn based on differences or similarities observed. In some embodiments, comparable sets of conditions, circumstances, individuals, or populations are characterized by a plurality of substantially identical features and one or a small number of varied features. Those of ordinary skill in the art will understand, in context, what degree of identity is required in any given circumstance for two or more such agents, entities, situations, sets of conditions, etc., to be considered comparable. For example, those of ordinary skill in the art will appreciate that sets of circumstances, individuals, or populations are comparable to one another when characterized by a sufficient number and type of substantially identical features to warrant a reasonable conclusion that differences in results obtained or phenomena observed under or with different sets of circumstances, individuals, or populations are caused by or indicative of the variation in those features that are varied.- 62 - 13241940vlAttorney Docket No. 2013237-1502

[0448] Corresponding to: As used herein, the term “corresponding to” refers to a relationship between two or more entities. For example, the term “corresponding to” may be used to designate the position / identity of a structural element in a compound or composition relative to another compound or composition (e.g., to an appropriate reference compound or composition). For example, in some embodiments, a monomeric residue in a polymer (e.g., an amino acid residue in a polypeptide or a nucleic acid residue in a polynucleotide) may be identified as “corresponding to” a residue in an appropriate reference polymer. For example, those of ordinary skill will appreciate that, for purposes of simplicity, residues in a polypeptide are often designated using a canonical numbering system based on a reference related polypeptide, so that an amino acid “corresponding to” a residue at position 190, for example, need not actually be the 190th amino acid in a particular amino acid chain but rather corresponds to the residue found at 190 in the reference polypeptide; those of ordinary skill in the art readily appreciate how to identify “corresponding” amino acids. For example, those skilled in the art will be aware of various sequence alignment strategies, including software programs such as, for example, BLAST, CS-BLAST, CUSASW++, DIAMOND, FASTA, GGSEARCH / GLSEARCH, Genoogle, HMMER, HHpred / HHsearch, IDF, Infernal, KLAST, USEARCH, parasail, PSI-BLAST, PSI-Search, ScalaBLAST, Sequilab, SAM, SSEARCH, SWAPHI, SWAPHLLS, SWIMM, or SWIPE that can be utilized, for example, to identify “corresponding” residues in polypeptides and / or nucleic acids in accordance with the present disclosure. Those of skill in the art will also appreciate that, in some instances, the term “corresponding to” may be used to describe an event or entity that shares a relevant similarity with another event or entity (e.g., an appropriate reference event or entity). To give but one example, a gene or protein in one organism may be described as “corresponding to” a gene or protein from another organism in order to indicate, in some embodiments, that it plays an analogous role or performs an analogous function and / or that it shows a particular degree of sequence identity or homology, or shares a particular characteristic sequence element.

[0449] Encode: As used herein, the term “encode” or “encoding” refers to sequence information of a first molecule that guides production of a second molecule having a defined sequence of nucleotides (e.g., a polyribonucleotide) or a defined sequence of amino acids. For example, a DNA molecule can encode an RNA molecule (e.g., by a transcription process that includes a DNA-dependent RNA polymerase enzyme). An RNA molecule can encode a polypeptide e.g., by a translation process). Thus, a gene, a cDNA, or an RNA molecule encodes a polypeptide if transcription and translation of RNA corresponding to that gene - 63 - 13241940vlAttorney Docket No. 2013237-1502produces the polypeptide in a cell or other biological system. In some embodiments, a coding region of a polyribonucleotide encoding a target antigen refers to a coding strand, the nucleotide sequence of which is identical to the polyribonucleotide sequence of such a target antigen. In some embodiments, a coding region of a polyribonucleotide encoding a target antigen refers to a non-coding strand of such a target antigen, which may be used as a template for transcription of a gene or cDNA.

[0450] Epitope: As used herein, the term “epitope” refers to a moiety that is specifically recognized by an immune system (e.g., an immune system component) of a subject. For example, in some embodiments, an epitope may be a moiety that is specifically recognized by a T cell, a B cell, an immunoglobulin (e.g., antibody or receptor), immunoglobulin (e.g., antibody or receptor), binding component or an aptamer. In some embodiments, an epitope is comprised of a plurality of chemical atoms or groups on an antigen. In some embodiments, such chemical atoms or groups are surface-exposed when the antigen adopts a relevant three-dimensional conformation. In some embodiments, such chemical atoms or groups are physically near to each other in space when the antigen adopts such a conformation. In some embodiments, at least some such chemical atoms or groups are physically separated from one another when the antigen adopts an alternative conformation (e.g., is linearized).

[0451] Estimated tumor sample purity. As used herein, the term “estimated tumor sample purity” is used to refer to an estimate of tumor content of a tumor sample, such as a fraction, percentage etc. of cancer cells within a tumor sample and / or, equivalently, an estimate of normal contamination, such as a fraction, percentage, etc. of normal (e.g., healthy) cells within a tumor sample. It should be understood that, in certain embodiments, tumor samples are assumed to be comprised of tumor cells and normal cells, such that a fraction of tumor cells in a tumor sample is equal to 1 minus a fraction of normal cells (e.g., 1 - p), estimates of tumor content and / or normal contamination equivalently measure an estimated tumor sample purity.

[0452] Expression: As used herein, the term “expression” of a nucleic acid sequence refers to the generation of a gene product from the nucleic acid sequence. In some embodiments, a gene product can be a transcript, e.g., a polyribonucleotide as provided herein. In some embodiments, a gene product can be a polypeptide. In some embodiments, expression of a nucleic acid sequence involves one or more of the following: (1) production of an RNA template from a DNA sequence (e.g., by transcription); (2) processing of an RNA - 64 - 13241940vlAttorney Docket No. 2013237-1502transcript (e.g., by splicing, editing, etc.); (3) translation of an RNA into a polypeptide or protein; and / or (4) post-translational modification of a polypeptide or protein.

[0453] Heterozygous Segment-. The term “heterozygous segment” as used here refers to a segment of a normal genome that comprises at least one heterozygous SNP, as well as any corresponding segments of a tumor genome or reference genome. In other words, a particular segment of a particular genome is defined as heterozygous or not according to whether the corresponding segment of a normal genome comprises a heterozygous SNP or not. For example, a reference genome may be partitioned into a plurality of segments, as described herein, in order to identify and define corresponding segments in a normal and tumor genome. Accordingly, if, for a given segment, the corresponding segment in the normal genome is determined to comprise a heterozygous SNP, then that segment is defined as a heterozygous segment.. For purposes of determining heterozygous segments, a normal genome may be a reference genome obtained from a database, a normal reference determined and / or compiled based on one or more subject (e.g., a panel), determined by sequencing a particular subject (e.g., the same subject whose tumor is being sequenced).

[0454] Homology -. As used herein, the term “homology” or “homolog” refers to the overall relatedness between polynucleotide molecules (e.g., DNA molecules and / or RNA molecules) and / or between polypeptide molecules. In some embodiments, polynucleotide molecules (e.g., DNA molecules and / or RNA molecules) and / or polypeptide molecules are considered to be “homologous” to one another if their sequences are at least 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical. In some embodiments, polynucleotide molecules (e.g., DNA molecules and / or RNA molecules) and / or polypeptide molecules are considered to be “homologous” to one another if their sequences are at least 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% similar (e.g., containing residues with related chemical properties at corresponding positions). For example, as is well known by those of ordinary skill in the art, certain amino acids are typically classified as similar to one another as “hydrophobic” or “hydrophilic” amino acids, and / or as having “polar” or “non-polar” side chains. Substitution of one amino acid for another of the same type may often be considered a “homologous” substitution.

[0455] Identity. As used herein, the term “identity” refers to the overall relatedness between polynucleotide molecules (e.g., DNA molecules and / or RNA molecules) and / or between polypeptide molecules. In some embodiments, polynucleotide molecules (e.g., - 65 - 13241940vlAttorney Docket No. 2013237-1502DNA molecules and / or RNA molecules) and / or between polypeptide molecules are considered to be “substantially identical” to one another if their sequences are at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical. Calculation of the percent identity of two nucleic acid or polypeptide sequences, for example, can be performed by aligning the two sequences for optimal comparison purposes (e.g., gaps can be introduced in one or both of a first and a second sequence for optimal alignment and non-identical sequences can be disregarded for comparison purposes). In certain embodiments, the length of a sequence aligned for comparison purposes is at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or substantially 100% of the length of a reference sequence. The nucleotides at corresponding positions are then compared. When a position in the first sequence is occupied by the same residue (e.g., nucleotide or amino acid) as the corresponding position in the second sequence, then the molecules are identical at that position. The percent identity between the two sequences is a function of the number of identical positions shared by the sequences, taking into account the number of gaps, and the length of each gap, which needs to be introduced for optimal alignment of the two sequences. The comparison of sequences and determination of percent identity between two sequences can be accomplished using a mathematical algorithm. For example, the percent identity between two nucleotide sequences can be determined using the algorithm of Meyers and Miller, 1989, which has been incorporated into the ALIGN program (version 2.0). In some exemplary embodiments, nucleic acid sequence comparisons made with the ALIGN program use a PAM120 weight residue table, a gap length penalty of 12 and a gap penalty of 4. The percent identity between two nucleotide sequences can, alternatively, be determined using the GAP program in the GCG software package using an NWSgapdna. CMP matrix.

[0456] Increased, Induced, or Reduced'. As used herein, these terms or grammatically comparable comparative terms, indicate values that are relative to a comparable reference measurement. For example, in some embodiments, an assessed value achieved with a provided composition (e.g., a pharmaceutical composition) may be “increased” relative to that obtained with a comparable reference composition. Alternatively or additionally, in some embodiments, an assessed value achieved in a subject may be “increased” relative to that obtained in the same subject under different conditions (e.g., prior to or after an event; or presence or absence of an event such as administration of a composition (e.g., a pharmaceutical composition) as described herein, or in a different,- 66 - 13241940vlAttorney Docket No. 2013237-1502comparable subject (e.g., in a comparable subject that differs from the subject of interest in prior exposure to a condition, e.g., absence of administration of a composition (e.g., a pharmaceutical composition) as described herein.). In some embodiments, comparative terms refer to statistically relevant differences (e.g., that are of a prevalence and / or magnitude sufficient to achieve statistical relevance). Those skilled in the art will be aware, or will readily be able to determine, in a given context, a degree and / or prevalence of difference that is required or sufficient to achieve such statistical significance. In some embodiments, the term “reduced” or equivalent terms refers to a reduction in the level of an assessed value by at least 5%, at least 10%, at least 20%, at least 50%, at least 75% or higher, as compared to a comparable reference. In some embodiments, the term “reduced” or equivalent terms refers to a complete or essentially complete inhibition, i.e., a reduction to zero or essentially to zero. In some embodiments, the term “increased” or “induced” refers to an increase in the level of an assessed value by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 80%, at least 100%, at least 200%, at least 500%, or higher, as compared to a comparable reference.

[0457] Initial event'. As used herein, the term “initial event” refers to a portion of a tumor genome which may encode a mutation. In certain embodiments, initial events are detected in tumor sequencing data of a tumor sample obtained from a subject. In some embodiments, an initial event is any portion of (e.g., individual sites in, contiguous subsequences in) a tumor genome (e.g., such that a set or list of initial events may include all sites in a tumor genome). In some embodiments, an initial event is a portion of a tumor genome associated with an observed statistical deviation from an expected read count in tumor sequencing data (e.g., wherein an expected read count is based on normal sequencing data which may be matched normal sequencing data). Given experimental and / or bioinformatic noise that may be present in sequenced samples, it is expected that among a plurality of initial events, any number of initial events may be determined to be false positives (e.g., may be determined to not reflect underlying mutations). In certain embodiments, an initial event refers to a variant portion of a tumor genome (e.g., identifying one or more sites within the tumor genome) determined, based on the sequencing data, to (e.g., potentially) have a different copy number and / or a different nucleotide sequence relative to a corresponding (e.g., aligned with the corresponding variant portion) normal reference. In certain embodiments, the different nucleotide sequence is or comprises an alternate allele (e.g., an individual nucleotide), different from an allele of the corresponding - 67 - 13241940vlAttorney Docket No. 2013237-1502normal reference. In certain embodiments, the different nucleotide sequence is or comprises an insertion and / or a deletion (e.g., of one or more nucleotides) relative to the corresponding normal reference (e.g., an indel). In certain embodiments, the variant portion is or comprises a structural variation relative to the normal reference. In certain embodiments, an initial event is a potential point mutation at a particular site. In certain embodiments, an initial event is a potential indel. In certain embodiments, an initial event is a potential structural variation. In certain embodiments, an initial event is a potential copy number variation. In certain embodiments, a normal reference is obtained from a database (e.g., a h 19 reference genome). In certain embodiments, a normal reference is determined based on normal sequencing data obtained for a subject (e.g., by sequencing a normal sample, e.g., of healthy tissue and / or cells)].

[0458] In order-. As used herein with reference to a polynucleotide or polyribonucleotide, “in order” refers to the order of features from 5' to 3' along the polynucleotide or polyribonucleotide. As used herein with reference to a polypeptide, “in order” refers to the order of features moving from the N-terminal-most of the features to the C-terminal-most of the features along the polypeptide. “In order” does not mean that no additional features can be present among the listed features. For example, if Features A, B, and C of a polynucleotide are described herein as being “in order, Feature A, Feature B, and Feature C,” this description does not exclude, e.g., Feature D being located between Features A and B.

[0459] Linker-. As used herein, the term “linker” refers to a portion of a polypeptide that connects different regions, portions, or antigens to one another.

[0460] Lipid: As used herein, the terms “lipid” and “lipid- like material” are broadly defined as molecules which comprise one or more hydrophobic moieties or groups and optionally also one or more hydrophilic moieties or groups. Molecules comprising hydrophobic moieties and hydrophilic moieties are also typically denoted as amphiphiles.

[0461] Neoantigen: As used herein, the term “neoantigen” refers to an antigen that is not present in a reference, such as a normal non-cancerous or germline cell, but is present in a cancer cell. In some embodiments, a neoantigen includes one or more mutations relative to a corresponding antigen present in a normal non-cancerous or germline cell.- 68 - 13241940vlAttorney Docket No. 2013237-1502

[0462] Neoantigen epitope: As used herein, the term “neoantigen epitope” refers to an epitope that is not present in a reference, such as a normal non-cancerous or germline cell, but is present in a cancer cell.

[0463] Nucleic acid! Polynucleotide: As used herein, the term “nucleic acid” refers to a polymer of at least 10 nucleotides or more. In some embodiments, a nucleic acid is or comprises DNA. In some embodiments, a nucleic acid is or comprises RNA. In some embodiments, a nucleic acid is or comprises peptide nucleic acid (PNA). In some embodiments, a nucleic acid is or comprises a single stranded nucleic acid. In some embodiments, a nucleic acid is or comprises a double- stranded nucleic acid. In some embodiments, a nucleic acid comprises both single and double- stranded portions. In some embodiments, a nucleic acid comprises a backbone that comprises one or more phosphodiester linkages. In some embodiments, a nucleic acid comprises a backbone that comprises both phosphodiester and non-phosphodiester linkages. For example, in some embodiments, a nucleic acid may comprise a backbone that comprises one or more phosphorothioate or 5'-N-phosphoramidite linkages and / or one or more peptide bonds, e.g., as in a “peptide nucleic acid”. In some embodiments, a nucleic acid comprises one or more, or all, natural residues (e.g., adenine, cytosine, deoxyadenosine, deoxycytidine,deoxy guanosine, deoxy thymidine, guanine, thymine, uracil). In some embodiments, a nucleic acid comprises on or more, or all, non-natural residues. In some embodiments, a non-natural residue comprises a nucleoside analog (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3 -methyl adenosine, 5 -methylcytidine, C-5 propynyl-cytidine, C-5 propynyl-uridine, 2-aminoadenosine, C5-bromouridine, C5 -fluorouridine, C5-iodouridine, C5-propynyl-uridine, C5 -propynyl-cytidine, C5 -methylcytidine, 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, 6-O-methylguanine, 2-thiocytidine, methylated bases, intercalated bases, and combinations thereof). In some embodiments, a non-natural residue comprises one or more modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose) as compared to those in natural residues. In some embodiments, a nucleic acid has a nucleotide sequence that encodes a functional gene product such as an RNA or polypeptide. In some embodiments, a nucleic acid has a nucleotide sequence that comprises one or more introns. In some embodiments, a nucleic acid may be prepared by isolation from a natural source, enzymatic synthesis (e.g., by polymerization based on a complementary template, e.g., in vivo or in vitro), reproduction in a recombinant cell or system, or chemical synthesis. In - 69 - 13241940vlAttorney Docket No. 2013237-1502some embodiments, a nucleic acid is at least 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 1 10, 120, 130, 140, 150, 160, 170, 180, 190, 20, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500, 10,000, 10,500, 11,000, 11,500, 12,000, 12,500, 13,000, 13,500, 14,000, 14,500, 15,000, 15,500, 16,000, 16,500, 17,000, 17,500, 18,000, 18,500, 19,000, 19,500, or 20,000 or more residues or nucleotides long.

[0464] Ploidy. As used herein, the term “ploidy,” for example of a tumor genome, is used to refer to an average of absolute copy numbers of all segments (e.g., across an entire region of a tumor genome), weighted by the length of each segment. A ploidy of a region of a tumor genome can be defined as the average of absolute copy numbers of all segments in the region, weighted by the length of each segment.

[0465] Polypeptide'. As used herein, the term “polypeptide” refers to a polymeric chain of amino acids. In some embodiments, a polypeptide has an amino acid sequence that occurs in nature. In some embodiments, a polypeptide has an amino acid sequence that does not occur in nature. In some embodiments, a polypeptide has an amino acid sequence that is engineered in that it is designed and / or produced through action of the hand of man. In some embodiments, a polypeptide may comprise or consist of natural amino acids, non-natural amino acids, or both. In some embodiments, a polypeptide may comprise or consist of only natural amino acids or only non-natural amino acids. In some embodiments, a polypeptide may comprise D-amino acids, L-amino acids, or both. In some embodiments, a polypeptide may comprise only D-amino acids. In some embodiments, a polypeptide may comprise only L-amino acids. In some embodiments, a polypeptide may include one or more pendant groups or other modifications, e.g., modifying or attached to one or more amino acid side chains, at the polypeptide’s N-terminus, at the polypeptide’s C-terminus, or any combination thereof. In some embodiments, such pendant groups or modifications comprise acetylation, amidation, lipidation, methylation, pegylation, etc., including combinations thereof. In some embodiments, a polypeptide may be cyclic, and / or may comprise a cyclic portion. In some embodiments, a polypeptide is not cyclic and / or does not comprise any cyclic portion. In some embodiments, a polypeptide is linear. In some embodiments, a polypeptide may be or comprise a stapled polypeptide. In some embodiments, the term “polypeptide” may be appended to a name of a reference polypeptide, activity, or structure; in such instances it is used herein to refer to polypeptides that share the relevant activity or structure and thus can - 70 - 13241940vlAttorney Docket No. 2013237-1502be considered to be members of the same class or family of polypeptides. For each such class, the present specification provides and / or those skilled in the art will be aware of exemplary polypeptides within the class whose amino acid sequences and / or functions are known; in some embodiments, such exemplary polypeptides are reference polypeptides for the polypeptide class or family. In some embodiments, a member of a polypeptide class or family shows significant sequence homology or identity with, shares a common sequence motif (e.g., a characteristic sequence element) with, and / or shares a common activity (in some embodiments at a comparable level or within a designated range) with a reference polypeptide of the class; in some embodiments with all polypeptides within the class). For example, in some embodiments, a member polypeptide shows an overall degree of sequence homology or identity with a reference polypeptide that is at least about 30-40%, and is often greater than about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more and / or includes at least one region (e.g., a conserved region that may in some embodiments be or comprise a characteristic sequence element) that shows very high sequence identity, often greater than 90% or even 95%, 96%, 97%, 98%, or 99%. Such a conserved region usually encompasses at least 3-4 and often up to 35 or more amino acids; in some embodiments, a conserved region encompasses at least one stretch of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35 or more contiguous amino acids. In some embodiments, a relevant polypeptide may comprise or consist of a fragment of a parent polypeptide.

[0466] Read count-. As used herein, the term “read count” refers to a number of sequencing data reads that map to a particular segment, portion thereof, or individual location (such as a SNP) within a genome. For example, the phrases “read count of a particular segment” and “segment read count” as used herein refer to a number of reads that map to the particular segment. For example, the phrases “read count of a particular SNP” and “SNP read count” as used herein refer to a number of reads that map to the particular SNP. The term “read count” may be preceded by an indication of a particular set of sequencing data and / or sequenced sample. For example, when a tumor sample is sequenced to produce tumor sequencing data comprising a plurality of tumor sequencing reads, the phrase “tumor read count” is used to refer to the number of tumor sequencing reads that map to a particular segment, portion thereof, or individual location (e.g., SNP) within a genome. Likewise, when a normal sample is sequenced to produce normal sequencing data comprising a plurality of normal sequencing reads, the phrase “normal read count” is used to refer to the number of - 71 - 13241940vlAttorney Docket No. 2013237-1502normal sequencing reads that map to a particular segment, portion thereof, or individual location (e.g., SNP) within a genome.

[0467] Reference'. As used herein, the term “reference” describes a standard or control relative to which a comparison is performed. For example, in some embodiments, an agent, animal, individual, population, sample, sequence or value of interest is compared with a reference or control agent, animal, individual, population, sample, sequence or value. In some embodiments, a reference or control is tested and / or determined substantially simultaneously with the testing or determination of interest. In some embodiments, a reference or control is a historical reference or control, optionally embodied in a tangible medium. Typically, as would be understood by those skilled in the art, a reference or control is determined or characterized under comparable conditions or circumstances to those under assessment. Those skilled in the art will appreciate when sufficient similarities are present to justify reliance on and / or comparison to a particular possible reference or control.

[0468] Ribonucleic acid (RNA) or Polyribonucleotide-. As used herein, the term “ribonucleic acid,” “RNA,” or “polyribonucleotide” refers to a polymer of ribonucleotides. In some embodiments, an RNA is single stranded. In some embodiments, an RNA is double stranded. In some embodiments, an RNA comprises both single and double stranded portions. In some embodiments, an RNA can comprise a backbone structure as described in the definition of “Nucleic acid / Polynucleotide” above. An RNA can be a regulatory RNA (e.g., siRNA, microRNA, etc.), or a messenger RNA (mRNA). In some embodiments, an RNA is a mRNA. In some embodiments, where an RNA is an mRNA, an RNA typically comprises at its 3' end a poly(A) region. In some embodiments, where an RNA is a mRNA, an RNA typically comprises at its 5' end an art-recognized cap structure, e.g., for recognizing and attachment of a mRNA to a ribosome to initiate translation. In some embodiments, an RNA is a synthetic RNA. Synthetic RNAs include RNAs that are synthesized in vitro (e.g., by enzymatic synthesis methods and / or by chemical synthesis methods).

[0469] Ribonucleotide-. As used herein, the term “ribonucleotide” encompasses unmodified ribonucleotides and modified ribonucleotides. For example, unmodified ribonucleotides include the purine bases adenine (A) and guanine (G), and the pyrimidine bases cytosine (C) and uracil (U). Modified ribonucleotides may include one or more modifications including, but not limited to, for example, (a) end modifications, e.g., 5' end modifications (e.g., phosphorylation, dephosphorylation, conjugation, inverted linkages, etc.), 3' end modifications (e.g., conjugation, inverted linkages, etc.), (b) base modifications, e.g.,- 72 - 13241940vlAttorney Docket No. 2013237-1502replacement with modified bases, stabilizing bases, destabilizing bases, or bases that base pair with an expanded repertoire of partners, or conjugated bases, (c) sugar modifications (e.g., at the 2' position or 4' position) or replacement of the sugar, and (d) internucleoside linkage modifications, including modification or replacement of the phosphodiester linkages. The term “ribonucleotide” also encompasses ribonucleotide triphosphates including modified and non-modified ribonucleotide triphosphates.

[0470] Secretory signal: As used herein, the term “secretory signal” refers to an amino acid sequence motif that targets associated polypeptides for translocation to a secretory pathway.

[0471] Segment, Segments: As used herein, the terms “segment” or “segments” (e.g., when used in regard to genetic material) refer to specific pre-defined regions of one or more genomes. For example, a particular reference genome may be subdivided into a plurality of segments, each a specific subsequence of consecutive nucleotides in the reference genome. In certain embodiments, a particular genome may be subdivided into its constituent genes, such that each segment corresponds to a particular, different, gene of the particular genome. In certain embodiments, exons of a particular genome are identified and retained, such that each segment corresponds to a particular, different, exon. In certain embodiments, each segment corresponds to a locus (e.g., a particular location on a chromosome where a particular gene, genetic marker, or allele is located). In certain embodiments, a particular genome may be subdivided into segments of a same or substantially same size (e.g., number of bases). As will be understood by one of skill in the art, a set sequencing data obtained from a particular sample, such as reads obtained via next generation sequencing (NGS) data obtained by sequencing a particular sample, may be aligned to a reference genome. In this way, reference genome may be subdivided into a plurality of segments and corresponding segments identified within a genome characteristic of the particular sample. Reads from the set of sequencing data can, accordingly, be identified as mapping to various particular segments within the genome characteristic of the particular sample and used to characterize them. In certain embodiments, multiple sets of sequencing data may be obtained for different samples (e.g., tumor sequencing data from a tumor sample, normal sequencing data from a normal sample) and aligned to a common reference genome. In this way, corresponding segments that comprises the same or substantially same (e.g., all save for variations due to e.g., single nucleotide polymorphisms (SNPs), single nucleotide variations (SNVs), insertions, deletions, etc.) base positions as from genomes characteristic of different samples - 73 - 13241940vlAttorney Docket No. 2013237-1502can be identified. That is, given a particular segment from one genome, associated with one sample, a corresponding segment of another genome, associated with another sample, may be identified. Corresponding segments may have a same and / or substantially same length (e.g., accounting for insertions, deletions, etc.). A particular segment is referred to herein as encoding or comprising a particular SNP and / or SNV if that particular SNP and / or SNV is within the particular segment.

[0472] Single Nucleotide Polymorphism (SNP): As used herein, the term “single nucleotide polymorphism” or “SNP” refers to a particular site (e.g., base position) in a genome where alternative bases are known and / or determined to distinguish one allele from another.

[0473] Single Nucleotide Variation (SNV): As used herein, the term “single nucleotide variation” is used to refer to a difference in the nucleic acid sequence (substitution of one base for another) at a particular site (allele) when comparing a genome from a diseased cell, such as a tumor cell, and a genome of a normal, non-diseased cell or a reference genome. In some embodiments, detecting mutations may refer to detecting nucleotide substitution mutations. In certain embodiments, a SNV is a somatic point mutation that occurs only in diseased (e.g., cancer) cells.

[0474] Subject: As used herein, the term “subject” refers to an organism to be administered with a composition described herein, e.g., for experimental, diagnostic, prophylactic, and / or therapeutic purposes. Typical subjects include animals (e.g., mammals such as mice, rats, rabbits, non-human primates, domestic pets, etc.) and humans. In some embodiments, a subject is a human subject. In some embodiments, a subject is suffering from a disease, disorder, or condition (e.g., cancer and / or a cancer-associated condition). In some embodiments, a subject is susceptible to a disease, disorder, or condition (e.g., cancer and / or a cancer-associated condition). In some embodiments, a subject displays one or more symptoms or characteristics of a disease, disorder, or condition (e.g., cancer and / or a cancer-associated condition). In some embodiments, a subject displays one or more non-specific symptoms of a disease, disorder, or condition (e.g., cancer and / or a cancer-associated condition). In some embodiments, a subject does not display any symptom or characteristic of a disease, disorder, or condition (e.g., cancer and / or a cancer-associated condition). In some embodiments, a subject is someone with one or more features characteristic of susceptibility to or risk of a disease, disorder, or condition (e.g., cancer and / or a cancer-- 74 - 13241940vlAttorney Docket No. 2013237-1502associated condition). In some embodiments, a subject is a patient. In some embodiments, a subject is an individual to whom diagnosis and / or therapy is and / or has been administered.

[0475] Therapy. The term “therapy” refers to an administration or delivery of an agent or intervention that has a therapeutic effect and / or elicits a desired biological and / or pharmacological effect (e.g., has been demonstrated to be statistically likely to have such effect when administered to a relevant population). In some embodiments, a therapeutic agent or therapy is any substance that can be used to alleviate, ameliorate, relieve, inhibit, prevent, delay onset of, reduce severity of, and / or reduce incidence of one or more symptoms or features of a disease, disorder, and / or condition (e.g., cancer and / or a cancer-associated condition). In some embodiments, a therapeutic agent or therapy is a medical intervention that can be performed to alleviate, relieve, inhibit, present, delay onset of, reduce severity of, and / or reduce incidence of one or more symptoms or features of a disease, disorder, and / or condition.

[0476] Treat'. As used herein, the term “treat,” “treatment,” or “treating” refers to any method used to partially or completely alleviate, ameliorate, relieve, inhibit, prevent, delay onset of, reduce severity of, and / or reduce incidence of one or more symptoms or features of a disease, disorder, and / or condition (e.g., cancer and / or a cancer-associated condition). Treatment may be administered to a subject who does not exhibit signs of a disease, disorder, and / or condition (e.g., cancer and / or a cancer-associated condition). In some embodiments, treatment may be administered to a subject who exhibits only early signs of the disease, disorder, and / or condition (e.g., cancer and / or a cancer-associated condition), for example for the purpose of decreasing the risk of developing pathology associated with the disease, disorder, and / or condition. In some embodiments, treatment may be administered to a subject at a later-stage of disease, disorder, and / or condition (e.g., cancer and / or a cancer-associated condition).

[0477] Wild-Type. As used herein, the term “wild-type” refers to an entity having a structure and / or activity as found in nature in a “normal” (as contrasted with mutant, diseased, altered, etc.) state or context. Those of ordinary skill in the art will appreciate that wild-type genes and polypeptides often exist in multiple different forms (e.g., alleles). For example, for a subject with cancer, wild-type genes, segments, SNPs, or properties thereof may be those present in a genome of that subject’s normal, non-cancerous, cells, as opposed to altered versions of those genes, segments, SNPs, or properties thereof that appear in genomes of cancer cells within the subject.- 75 - 13241940vlAttorney Docket No. 2013237-1502DETAILED DESCRIPTION

[0478] It is contemplated that systems, architectures, devices, methods, and processes of the claimed invention encompass variations and adaptations developed using information from the embodiments described herein. Adaptation and / or modification of the systems, architectures, devices, methods, and processes described herein may be performed, as contemplated by this description.

[0479] Throughout the description, where articles, devices, systems, and architectures are described as having, including, or comprising specific components, or where processes and methods are described as having, including, or comprising specific steps, it is contemplated that, additionally, there are articles, devices, systems, and architectures of the present invention that consist essentially of, or consist of, the recited components, and that there are processes and methods according to the present invention that consist essentially of, or consist of, the recited processing steps.

[0480] It should be understood that the order of steps or order for performing certain action is immaterial so long as the invention remains operable. Moreover, two or more steps or actions may be conducted simultaneously.

[0481] The mention herein of any publication, for example, in the Background section, is not an admission that the publication serves as prior art with respect to any of the claims presented herein. The Background section is presented for purposes of clarity and is not meant as a description of prior art with respect to any claim.

[0482] Documents are incorporated herein by reference as noted. Where there is any discrepancy in the meaning of a particular term, the meaning provided in the Definition section above is controlling.

[0483] Headers are provided for the convenience of the reader - the presence and / or placement of a header is not intended to limit the scope of the subject matter described herein.

[0484] Among other things, the present disclosure is based on an insight that subpopulations of cancer cells present in a tumor and / or biological features thereof can be accurately identified and / or estimated using certain methods for classifying and / or estimating- 76 - 13241940vlAttorney Docket No. 2013237-1502clonality classification state assignments of tumor cell mutations based on sequencing data derived from a tumor sample.

[0485] Accordingly, presented herein are technologies for determining clonality classification states of somatic mutations present in cancer cells based on sequencing data. Among other things, the clonality classification state assignments, that are made possible via the methods and systems described herein, accurately reflect underlying, true, physical characteristics of cancer cell mutations, including, without limitation, whether mutations are biologically clonal and / or subclonal, as well as their prevalence, and thereby allow for neoepitopes based on particular cancer mutations to be selected for inclusion in immunotherapy treatments, such as personalized cancer vaccines, in a way that allows biologically clonal and highly prevalent mutation targets to be identified and prioritized, while avoiding rare and / or biologically subclonal targets that are only present in a small fraction of cancer cells and, accordingly, will not result in effective treatments.Biological Significance of Clone Types

[0486] As tumors evolve, cancer cell subpopulations emerge, an evolution which may occur spontaneously, by virtue of the mutations generated when a cancer cell divides and / or in response to selection pressures from, for example, the environment, immune system, and / or any currently or previously administered therapy. A tumor (e.g., as represented by a bulk tumor sample) may therefore be viewed as an ensemble of tumor genomes, wherein certain lineages or subpopulations of tumor cells each have identical or nearly identical genomes. Accordingly, a tumor may be described as a collection of distinct subpopulations of tumor cells, each defined by a unique combination of mutations present in the subpopulation genome. Such subpopulations each consist of single cancer cells with identical or nearly identical genomes, whereas each subpopulation is distinguishable from other subpopulations in that its set of mutations is unique, although certain mutations may be shared across subpopulations.

[0487] An ideal observer would have access to a full and accurate genetic sequence of each individual tumor cell and could therefore identify the unique set of mutations of each tumor cell to determine the genetic lineage of each mutation. However, individual tumor cells from a bulk sample are generally mixed for sequencing analysis, along with cells that- 77 - 13241940vlAttorney Docket No. 2013237-1502are not tumor cells, and estimation is required at multiple steps to assess the presence, prevalence, and identity of mutations.

[0488] Turning to FIG. 1A, a bulk tumor sample may be viewed, from a biological standpoint, as comprising a collection of distinct subpopulations (1A130), each subpopulation comprising a unique set of mutations (e.g., “A” - “F”), whereby mutations have different prevalences in the bulk tumor (1A140) (e.g., have different biological cellularities).However, a bulk tumor sample may equally be viewed, from the standpoint of an illustrative embodiment of methods and systems described herein (1A110), as comprising a collection of mutations, where each one is classified into a clone type based on analysis of sequencing data obtained from the tumor and from a matched normal sample (1A120). The illustrative embodiment may detect and classify a set of mutations {A} present in the tumor sample assigned a CLONAL clone type and / or a set of mutations {B } assigned a PREVALENT clone type, and / or sets of mutations {C, E, F} assigned a SUBCLONAL clone type and / or eliminates a set of mutations { D } determined by the embodiment to be rare subclones.

[0489] In FIG. 1A, clone type classifications shown in 1A120 correspond to prevalences shown in 1A140. For example, mutation A, which is a biologically clonal mutation, is shared by all subclonal populations in 1A130 and is classified as CLONAL in 1A120. Mutation B is present in subpopulation 1 and in subpopulation 2, which is nested within subpopulation 1. Since subpopulation 1 is prevalent in the bulk tumor, mutation B is assigned as having a PREVALENT clone type in 1A120. Mutations C, E, and F are less prevalent biologically subclonal subpopulations (2, 4 and 3, respectively) and therefore are classified as having a SUBCLONAL clone type. Mutation D is present just in subpopulation 5 and is therefore a rare biological subclone.

[0490] Accordingly, these two views of the tumor sample converge in that each relates to a cellularity for each mutation (1A140). Across biological subpopulations, the collection of tumor genomes defines the biological cellularity of each mutation, whereas methods provided by the present disclosure classify mutations into clone types according to estimates of cellularity (p). Both biological and estimated cellularity are continuous measures due, in full or in part, to statistical variability. However, biological cellularity is inherently quantized, as it reflects the prevalence of a mutation across one or more subpopulations of tumor cells, as shown in FIG. 1A (1A130). An accurate estimate of cellularity for a given mutation, e.g., as estimated with methods and systems of the present disclosure, is a proxy for biological cellularity and is therefore expected to be similarly - 78 - 13241940vlAttorney Docket No. 2013237-1502quantized. Examples 2- 7 herein provide confirmatory evidence that clone type classification methods of the present disclosure are consistent with these biological features of tumor cell mutations.

[0491] Accordingly, methods for accurate estimation and / or classification of clone types of tumor cell mutations provided by the present disclosure provide insight into biological features of a tumor sample. As discussed in more detail below, this insight into the biological features of a sample is useful in a number of applications, including in identifying, prioritizing and / or selecting mutations to be targeted using cancer immunotherapies (e.g., personalized cancer immunotherapies such as personalized cancer vaccines or adoptive T cell therapies).

[0492] An exemplary dynamical process of cancer cell subpopulation emergence is schematized in FIG. IB, which illustrates evolution of a tumor and emergence of cancer cell subpopulations over time, up to the point of tumor sampling (1B130). In this example, evolution begins with a progenitor cancer cell that begins to divide unchecked, such that a tumor initially contains a population of tumor cells with identical or nearly identical genomes (1B100). At a later timepoint, a new variant may emerge, e.g. due to a driver mutation D (1B140), that provides that variant with a fitness advantage over the existing population. Accordingly, the subpopulation carrying this driver mutation expands, overtaking the original population which comprises the first set of mutation(s) and giving rise to a major subpopulation, wherein the driver mutation D is present in nearly all of the tumor cells sampled, alternatively or in addition to any mutation(s) present in the genome of the progenitor cell. Further subpopulations may emerge at different timepoints and with different fitness advantages (e.g., IB 120). For example, at the time of tumor extraction such additional subpopulations IB 120 present in a minority of tumor cells may be considered minor subpopulations.

[0493] The present disclosure is based in part on a recognition that, given the complexity of tumor evolution, a view of mutations as either clonal or subclonal is overly limiting and does not capture the richness of the underlying biology. Further, exemplary methods of tumor deconvolution, mutation detection, and clone type classification disclosed herein as applied to tumor sample data provide an insight that cellularities of putative biologically subclonal mutations cluster into discrete, well-separated distributions or discrete spectrums (see, e.g., Example 3) which are captured by clone type classification and which more directly correspond to biological features of mutations present in tumor cells.- 79 - 13241940vlAttorney Docket No. 2013237-1502Quantized and population-based clone type classification approaches

[0494] In some embodiments, clone type classification is performed in a quantizationbased manner, whereby a putative mutation is assigned to a clone type based on its estimated cellularity as assessed relative to nominal ranges of cellularity associated with a given set of clone types. In some embodiments, quantization-based clone type classification quantizes a cellularity scale into regions, allowing discrimination between putative biologically clonal mutations, major subclones and minor subclones.

[0495] In some embodiments, quantization-based clone type classification is advantageous in that for many applications (e.g., personalized cancer immunotherapy applications, TCR enrichment applications), it is particularly useful to separate mutations into categories based on their absolute abundance in the tumor in order to effectively prioritize mutations.

[0496] The present disclosure provides quantization-based clone type classification methods that are precise and quantitative. For example, as shown in FIG.39A and FIG.39B, performance of an exemplary clone type classification method disclosed herein was validated using a bulk tumor cell line mixed to generate subclones at nominal cellularities, such that performance of the exemplary method recapitulated a nominal (true) cellularity.

[0497] Without wishing to be bound by any particular theory, quantization-based clone type classification approaches as described herein may be a particularly useful approach due to their generality and lack of assumptions regarding the nature of cancer cell subpopulations or the cellularities of mutations, as such approaches are therefore independent of the number of subpopulations that exist within a given patient’s tumor or their prevalence within the tumor.

[0498] In some embodiments, clone type classification may be performed in a population-based manner, whereby mutations specific to an n-th subpopulation are assigned an n-th clone type, wherein each clone type is characterized by an estimated nominal cellularity corresponding to the cellularities of mutations assigned to that clone type. For example, in some embodiments, one or more mutations specific to a biological minor subpopulation may be assigned a first clone type, and mutations specific to a biological major subpopulation may be assigned a second clone type, and an estimated nominal cellularity of the first clone type of the minor subpopulation is lower than an estimated nominal cellularity - 80 - 13241940vlAttorney Docket No. 2013237-1502of the second clone type of the major subpopulation. In some embodiments, an estimated nominal cellularity is calculated by taking an average of cellularities of all mutations assigned to a given clonality state.

[0499] Population-based clone type classification approaches provided herein leverage a recognition that, in some embodiments, clone type classification links estimated cellularities to tumor cell subpopulations which exist in a tumor. Accordingly, in some embodiments, clone type reflects the nominal cellularity of mutations specific to a given tumor cell subpopulation, because all mutations specific to a given subpopulation should all have the same nominal cellularity. For example, in some embodiments, where a tumor is comprised of a small number of cancer cell subpopulations, with each cancer cell subpopulation having a defined nominal cellularity, population-based clone type classification methods cluster cellularities and assign these cellularities into different groups, where each group would ideally be a proxy for a given cancer cell subpopulation.

[0500] In some embodiments, population-based clone type classification comprises clustering by including genomic information (e.g., the mutation node), or timing information (e.g., using CNV events to time mutations), thereby providing additional information for clustering constructed to capture naturally occurring subpopulations within the tumor. In some embodiments, clustering information may be used to rank the subpopulations accurately in terms of their abundance and / or to predict the degree of aggressiveness of different subpopulations by considering the function of the mutations specific to that subpopulation / variant.

[0501] In some embodiments, a population-based clone type classification approach may be particularly useful when a nominal (true) cellularity of a subpopulation is on the border between two different cellularity classifications (e.g., on the border between being a major subpopulation and a minor subpopulation, or on the border between PREVALENT and SUBCLONAL), as a population-based clone type classification approach may more accurately resolve classification of such mutations.

[0502] Another disadvantage is that a clustering approach can be error prone when estimated cellularities are noisy. As a result, a clustering approach may inaccurately group mutations that belong to subpopulations with different nominal cellularities together. If mutations belonging to minor and major subpopulations are grouped together and labeled with the same Clonality State, this can lead to errors in the prioritization procedure.- 81 - 13241940vlAttorney Docket No. 2013237-1502Alternatively, if the clustering method attempts to minimize such errors by excluding mutations with large, expected errors, then this may increase the number of mutations that have undetermined Clonality States.

[0503] In some embodiments, clone type classification may be performed in a hybrid quantization / population-based manner. Without wishing to be bound by any particular theory, a hybrid method of clone type classification may address limitations of quantization-based or population-based clone type classification approaches. For example, in some embodiments, quantization-based clone type classification approach is first applied to identify clonal mutations, and after removing clonal mutations, a population-based clone type classification approach is used to group mutations based on cellularities. In some embodiments, a hybrid method of clone type classification incorporates information such as genomic information (location of mutations), mutation nodes and / or timing information derived from CNV events.

[0504] In some embodiments, a hybrid method of clone type classification entails first applying a quantization-based clone type classification approach and subsequently applying a clustering approach to resolve ambiguities and / or errors in classification. For example, in some embodiments, estimated or nominal cellularities close to the boundary between two different clone types (e.g., CLONAL / PREVALENT-HIGH, PREVALENT-HIGH / PREVALENT-LOW, and / or PREVALENT-LOW / SUBCLONAL) may be incorrectly classified with a pure quantitative-based Clonality State approach. In some embodiments, mutations are initially classified by a quantization-based clone type classification approach, and clone type assignments are refined by swapping mutations between adjacent clone types by leveraging clustering information, e.g. informed by mutation node information and / or timing information. In some embodiments, timing information is from tumor evolution models that leverage CNV events to deduce timing of mutations.Clinical relevance of subclonal populations and prioritization thereof

[0505] Cancer cell subpopulations are clinically relevant because they can potentially drive tumor growth. This can be achieved by several mechanisms, such as: (1) adapting to environmental changes and exploiting new growth opportunities; (2) promoting competition and cooperation between different subpopulations, for example by competing for resources such as nutrients and space, driving aggressive growth, or, alternatively by cooperating and sharing resources supporting each other's survival and proliferation; (3) evolving mechanisms - 82 - 13241940vlAttorney Docket No. 2013237-1502to evade the immune system, allowing the tumor to grow unchecked by immune surveillance; (4) providing metastatic potential by enhancing abilities to invade surrounding tissues and spread to other parts of the body, contributing to metastasis and overall tumor progression; (5) promoting resistance to therapy by harboring mutations that confer resistance, enabling them to survive treatment and continue to drive tumor growth.

[0506] A cancer cell subpopulation emerges when a new driver mutation appears, conferring that variant with a fitness advantage that enables it to expand into a subpopulation with a significant cellularity. Given that such a subpopulation could be driving tumor growth and tumor progression, targeting cancer cell subpopulations with a personalized cancer immunotherapy may improve the clinical outcome of the patient. In some embodiments, clone type classification methods of the present disclosure identify driver mutations within these subpopulations. Without wishing to be bound by any particular theory, therapeutic approaches which target driver mutations and / or other mutations that confer a particular fitness advantage to a tumor cell population may be particularly useful, since the subpopulation cannot evolve to escape the therapy without giving up the fitness advantage. Accordingly, in some embodiments, clone type classification methods of the present disclosure prioritize driver mutations over passenger mutations within the same clone type.

[0507] In some embodiments, a fitness advantage is or comprises a metabolic adaptation (e.g., an enhanced ability to survive or function when nutrients are limited). In some embodiments, a fitness advantage is or comprises a stress response adaptation (e.g., an improved ability to manage oxidative stress or DNA damage). In some embodiments, a fitness advantage is or comprises a vascularization adaptation (e.g., an enhanced formation of blood vessels). In some embodiments, a fitness advantage is or comprises one or more altered signaling pathways (e.g., in pathways that provide a proliferative advantage). In some embodiments, a fitness advantage is or comprises enhanced ability to invade, survive and / or proliferate in a distant tissue. In some embodiments, a fitness advantage is or comprises an enhanced ability to evade immune surveillance.

[0508] Without wishing to be bound by any particular theory, targeting mutations with non-CLONAL clone types (e.g., PREVALENT-HIGH, PREVALENT-LOW, and / or SUBCLONAL) that are present within cancer cell subpopulations may help arrest tumor growth and / or prevent aggressive lineages from emerging. A subset of mutations with a non-CLONAL clone types may be the new driver mutations that emerged in the given subpopulation and that are responsible for conferring the fitness advantage to a given - 83 - 13241940vlAttorney Docket No. 2013237-1502subpopulation. Targeting a subclonal driver mutation has the benefit that a cancer cell cannot silence that driver mutation without losing its fitness advantage for the subpopulation carrying that driver. In contrast, if a subclonal passenger mutation is targeted, a cancer cell may simply silence that mutation (e.g., via a loss of heterozygosity event) and thereby escape the therapy without losing fitness.

[0509] Accordingly, in some embodiments, mutations with a non-CLONAL clone type and a driver phenotype can be prioritized over mutations with a non-CLONAL clone type without a driver phenotype, as the former may be driving the fitness advantage of an expanding subpopulation. Without wishing to be bound by any particular theory, this approach may lead to targeting more aggressive subclonal populations with higher clinical relevance.

[0510] Accordingly, FIG. 1C illustrates an approach whereby clonality state predictions based on technologies described herein may also utilize biological features to prioritize mutations for inclusion in an immunotherapy. For example, in some embodiments, a method for developing a personalized cancer immunotherapy prioritizes mutations according to the following priority order (highest to lowest) giving precedence to mutations that are drivers over mutations that are not drivers:1. Clonality State is CLONAL and, in addition, the mutation has either a driver phenotype and / or a fully mutated essential gene phenotype2. Clonality State is CLONAL3. Clonality State is PREVALENT and the mutation has a driver phenotype4. Clonality State PREVALENT5. Clonality State N / A and the mutation has a driver phenotype6. Clonality State N / A7. Clonality State is SUBCLONAL and the mutation has a driver phenotype8. Clonality State is SUBCLONAL

[0511] In some embodiments, a method for developing a personalized cancer immunotherapy prioritizes mutations according to the following priority order (highest to lowest), wherein PREVALENT-HIGH and PREVALENT-LOW clone types are distinguished:1. Clonality State is CLONAL and, in addition, the mutation has either a driver phenotype and / or a fully mutated essential gene phenotype2. Clonality State is CLONAL3. Clonality State is PREVALENT-HIGH and the mutation has a driver phenotype 4. Clonality State PREVALENT -HIGH- 84 - 13241940vlAttorney Docket No. 2013237-15025. Clonality State is PREVALENT - LOW and the mutation has a driver phenotype 6. Clonality State PREVALENT -LOW7. Clonality State N / A and the mutation has a driver phenotype8. Clonality State N / A9. Clonality State is SUBCLONAL and the mutation has a driver phenotype10. Clonality State is SUBCLONALA. Cancer Genomics, Mutations, and Immunotherapy

[0512] Certain cancer mutations are unique to a patient’ s cancer and, when expressed, produce proteins and / or peptides that are distinct from those produced by normal cells. These distinct proteins and / or peptides can, accordingly, be specifically targeted via immunotherapy approaches that leverage the patient’ s own immune system to clear cancer cells while avoiding damage to normal cells. Technologies of the present disclosure, among other things, leverage and analyze sequencing data to identify potential cancer-specific mutations within genome(s) of a patient’ s tumor (tumor genome) and, moreover, characterize them in a manner that allows those mutations that will be the most effective targets of immunotherapies to be identified and prioritized, for example for as targets for personalized cancer vaccines, T-cell receptor (TCR) therapies, and the like.A.i Samples and Sequencing Data

[0513] Among other things, technologies of the present disclosure may be used in connection with sequencing and / or genomic analysis methods and systems that generate sequencing data for a subject and allow for detection of mutations that can be targeted, for example, for disease treatment via immunotherapy. For example, as shown in FIG. ID, in certain embodiments, sequencing data 128 may be generated for a subject having and / or suspected of having cancer, by obtaining a tumor sample 102 from the subject. The tumor sample 102 may be processed 104, for example, to extract and prepare nucleic acid material for sequencing and sequenced 106 to generate tumor sequencing data - i.e., sequencing data representing and obtained from nucleic acid material 104 from a tumor sample 102. In certain embodiments, as illustrated in FIG. ID, sequencing data 128 may be generated in and / or comprise replicates. For example, in certain embodiments, multiple tumor samples may be extracted and sequenced independently; in certain embodiments, a single tumor sample may be extracted and used to prepare multiple libraries (e.g., such that processing - 85 - 13241940vlAttorney Docket No. 2013237-1502steps of extracting, fragmenting, and amplifying nucleic from the sample are performed repeatedly and independently), which are then sequenced; in certain embodiments, library preparation may be performed repeatedly on a single pool of extracted nucleic acid, and the multiple libraries sequenced; in certain embodiments, a single library is sequenced multiple times (e.g., as in a technical replicate). In certain embodiments, a normal sample 112 may also be obtained, processed 114, and sequenced 116, to generate normal sequencing data 118 - i.e., sequencing data representing, and obtained from, nucleic acid material from a normal sample 112. In certain embodiments, as with tumor sample sequencing data, e.g., as illustrated in FIG. ID, normal sample sequencing data may also comprise a plurality of replicates.

[0514] A tumor sample 102 may be any sample derived from a particular subject and comprising, and / or expected to comprise, cancer cells (e.g., of the particular subject). In certain embodiments, a tumor sample is or comprises a liquid sample, such as serum, plasma, blood, urine, etc. For example, a liquid sample, such as blood, may comprise, or be suspected of comprising, cancer cells, such as circulating tumor cells (CTCs). In certain embodiments, a tumor sample is or comprises a tissue sample, for example, obtained from a subject via biopsy. A tumor sample may be representative of a subject’s primary tumor and / or one or more metastases. For example, a primary tumor sample may be obtained via biopsy of a region of a subject known or expected to harbor a primary tumor. A metastasis sample or metastases samples may be obtained via biopsy of one or more region(s) of a subject known or expected to harbor metastases. Tumor samples may be obtained and / or preserved in a variety of formats, such as, for example, isolated cells (e.g., CTCs) from blood samples, fresh tissue, flash frozen tissue, formalin-fixed paraffin embedded tumor tissue (FFPE samples), and the like.

[0515] As illustrated in FIG. ID, a tumor sample may comprise cancer cells 102b as well as, in certain cases, normal (i.e., non-cancerous) cells 102a. Accordingly, in certain embodiments, a tumor sample purity may measure relative fraction of cancer cells within a tumor sample. Tumor sample purity may, for example, be computed as (1 — / r) = rT / t]T+?7W), where « is a contamination fraction, representing a relative fraction of normal cells infiltrating a tumor sample, given by / r = TN / TT+ tN) and / T and / / N are a number of tumor and normal cells in a tumor sample, respectively. Tumor sample purity may be expressed as a decimal value, percentage, etc. As described in further detail herein, typically, purity does not need to be measured directly (e.g., via direct measuring / counting amounts of tumor and - 86 - 13241940vlAttorney Docket No. 2013237-1502normal cells in a sample), but, rather, can be determined and / or estimated using sequencing data, for example, via tumor modelling approaches.

[0516] A normal sample 112 may be any sample derived from a particular subject and comprising, and / or expected to comprise, the subject’s normal cells, but not cancer cells (e.g., in certain embodiments, a normal sample 112 does not contain any cancer cells). In certain embodiments, a normal sample is or comprises a liquid sample, such as serum, plasma, blood, urine, saliva, etc. In certain embodiments, a normal sample is or comprises a tissue sample, for example obtained from a subject via biopsy. Normal samples may be obtained and / or preserved in a variety of formats, such as, for example, isolated cells [e.g., peripheral blood mononuclear cells (PBMCs)] from blood samples, fresh tissue, flash frozen tissue, formalin-fixed paraffin embedded tumor tissue (FFPE samples), and the like.

[0517] As illustrated in FIG. ID, a normal sample 112 nominally contains only normal patient cells 112a. In certain embodiments, a normal sample may be obtained from blood of a subject. In certain embodiments, a normal sample may still comprise a small (e.g., negligible) number of cancer cells. For example, in certain embodiments, a normal cell may be obtained from a region of a subject near a tumor e.g., in an effort to obtain a normal sample from a same or similar underlying tissue type). In certain embodiments, a normal sample comprises less than 1%, e.g., less than 0.1%, e.g., less than 0.01%, e.g., less than 0.001% tumor cells.

[0518] Samples, such as tumor samples and / or normal samples, may be processed to obtain, and / or prepare, nucleic acid material therefrom for sequencing. For example, nucleic acid material, such as DNA and / or RNA, may be extracted and prepared for sequencing (e.g., via amplification, fragmentation, labeling, etc. ) as appropriate, depending on a particular desired sequencing method and / or data format. For example, in certain embodiments, sequencing data may be whole genome sequencing (WGS) data; in certain embodiments, sequencing data may be whole exome sequencing (WES) data. Various commercially available kits and instruments may be used to prepare samples for and obtain WGS and / or WES data, including, but not limited to, those provided by Illumina, Inc., PacBio, Oxford Nanopore Technologies, Thermo Fisher Scientific’s Ion Torrent™, etc.

[0519] In certain embodiments, technologies described herein may utilize various combinations of sequencing data. For example, in certain embodiments, WGS data may be used (e.g., alone). In certain embodiments, WGS may be used in combination with WES - 87 - 13241940vlAttorney Docket No. 2013237-1502data. In particular, WGS may be used to determine copy numbers and purity and WES used for mutation detection. In this way, discrete clone types as described herein may be determined based on mutations detected in WES, but using copy number predictions (e.g., via tumor deconvolution techniques, such as CNV-based approaches described herein) based on WGS data. In certain embodiments, WES and / or WGS data may be used in combination with RNAseq data. In certain embodiments, RNAseq can be used to confirm predictions for mutations that are likely to be false positives, in particular false positives due to PCR errors.

[0520] In some embodiments, methods and systems of the present disclosure utilize whole exome sequencing (WES) or whole genome sequencing (WGS) data in singletons (e.g., one sequenced tumor sample and one sequenced normal sample). In some embodiments, methods and systems of the present disclosure utilize whole exome sequencing (WES) or whole genome sequencing (WGS) data in replicates (e.g., two or more sequenced tumor samples and two or more sequenced normal samples). In some embodiments, methods and systems of the present disclosure utilize whole exome sequencing (WES) or whole genome sequencing (WGS) data in conjunction with RNA-seq data. In some embodiments, RNA-seq data is used in a method or system of the present disclosure to filter errors such as PCR errors, FFPE, sequencing artifacts, and / or other types of errors. For example, a mutation detected in DNA sequencing data may only be accepted if it is also detected using RNA-seq reads. In some embodiments, RNA seq reads are high quality RNA-seq reads. In some embodiments, if only singletons are used, coverage of singletons is required to be comparable to coverage of merged replicates (e.g., in order to ensure comparable sensitivity).

[0521] In certain embodiments, to generate sequencing data, libraries are created from the normal gDNA and tumor gDNA. In certain embodiments, from each gDNA sample, two or more libraries can be generated. In certain embodiments, from each gDNA sample, two libraries can be generated. For example, in certain embodiments, as illustrated in FIG.2, sequencing data 228 may be generated in and / or comprise replicates created by preparing multiple (e.g., two or more) libraries associated with each (e.g., gDNA) sample. The libraries can be created for whole exome sequencing and / or whole genome sequencing and / or RNA sequencing (RNAseq). Samples are then sequenced using high throughput sequencing such as NGS.

[0522] For example, in certain embodiments, sequencing data 228 may be generated in and / or comprise replicates. For example, multiple tumor samples may be extracted and sequenced independently; in certain embodiments, a single tumor sample may be extracted - 88 - 13241940vlAttorney Docket No. 2013237-1502and used to prepare multiple libraries (e.g., such that processing steps of extracting, fragmenting, and amplifying nucleic from the sample are performed repeatedly and independently), which are then sequenced; in certain embodiments, library preparation may be performed repeatedly on a single pool of extracted nucleic acid, and the multiple libraries sequenced; in certain embodiments, a single library is sequenced multiple times (e.g., as in a technical replicate). In certain embodiments, as with tumor sample sequencing data, e.g., as illustrated in FIG.2, normal sample sequencing data may also comprise a plurality of replicates.

[0523] Sequencing data 128, 108, 118, may be stored and / or presented in a variety of formats, such as FASTQ, SAM, BAM, etc. For example, sequencing data for a particular sample may comprise a plurality of reads, each read representing a nucleotide sequence of a polynucleotide fragment corresponding to (e.g., that maps to) a portion of a subject’s tumor genome and / or exome, and / or portion of the subject’s normal genome and / or exome. In certain embodiments, sequencing data typically comprises multiple overlapping reads, which may be aligned to a reference genome, such as an hl9 or h38 reference genome (see, e.g., ncbi.nlm.nih.gov / datasets / genome / GCF_000001405.13 / and ncbi.nlm.nih.gov / datasets / genome / GCF_000001405.26 / , respectively) for human patients, to map each read to a particular region of the overall genome. Aligned reads may be stored and / or provided in file formats, such as sequence alignment map (SAM) or the binary compressed version thereof (BAM). In certain embodiments, sequencing data duplicate reads may be marked and / or removed from the sequencing data. In certain embodiments, sequencing adapters may be removed from the sequencing data. In certain embodiments, e.g., in the case of case of short read sequencing, sequencing can be paired end or single end, wherein different read lengths can be used (e.g., 50bp, 100bp, 150bp, etc.)A.ii Tumor Genomics

[0524] Turning to FIG. IE, sequencing data 128 may be used to piece together and / or infer properties relating to a tumor genome and / or normal genome of a subject. In certain embodiments, a genome, such as a tumor genome and / or a normal genome, may be subdivided into a plurality of segments (black bars in the normal genome and tumor genome schematics), such as segments 152 and 154. In certain embodiments, each segment corresponds to a particular, different, gene. In certain embodiments, each segment- 89 - 13241940vlAttorney Docket No. 2013237-1502corresponds to exonic portions of a particular, different, gene (e.g., just the exons), for example when WES data is obtained. In certain embodiments, each segment corresponds to a locus (e.g., a particular location on a chromosome where a particular gene, genetic marker, or allele is located).

[0525] As illustrated in FIG. IE, a normal (e.g., human) genome is diploid, comprising two copies or versions of each gene or segment e.g., corresponding to the subject’s maternal and paternal DNA). Segments may be heterozygous, in that the copies differ, corresponding to different alleles, or may be homozygous, comprising identical alleles.

[0526] Segments in a tumor genome are not necessarily diploid and are not necessarily balanced (though certain segments of a tumor genome may be diploid and / or balanced). For example, as shown in FIG. IE, a number of copies of each allele in a segment may vary from segment to segment, and may be less than two, equal to two, or greater than two. Segments in a tumor genome need not be balanced - i.e., tumor genome heterozygous segments do not necessarily comprise a same number of copies of each allele, but, in certain embodiments, may comprise a greater number of copies of one allele (a “major allele”) than the other (a “minor allele”).

[0527] Accordingly, in certain embodiments, various parameters are used to characterize and represent physical properties of a particular segment, along with, in certain embodiments, mutations identified therein. For example, in certain embodiments, a segment j may be characterized by an absolute copy number, CN, computed as a number of copies of a particular segment (e.g., within a single tumor cell). For example, in FIG. IE, absolute copy numbers for each segment are listed above the tumor genome schematic. Shown in further detail, segment 154 is a heterozygous segment [e.g., a segment comprising one or more heterozygous single nucleotide polymorphism(s) (SNPs)], with a single nucleotide polymorphism (SNP) 156 in a first allele. As illustrated in the figure, segment 154 has an absolute copy number 162 of 5. For heterozygous segments, such as segment 154, an allele specific copy number 164 may be determined as a maximum absolute number of copies of a major allele - i.e., the allele having a number of copies greater than or equal to that of the other, minor, allele - i.e., CNx > CNx, where X and Y denote the major and minor alleles, respectively. For example, for the particular segment 154 shown in FIG. IB, there are three copies of the major allele, such that the allele specific copy number is three (CNx = 3).-90 - 13241940vlAttorney Docket No. 2013237-1502

[0528] Parameters representing properties and characteristics of mutations may also be determined. For example, if a certain gene is mutated, a number of physical copies of the mutated gene in a given tumor cell, referred to herein as the zygosity of the mutation, may be determined. In certain embodiments, a fractional zygosity of a mutation (Q may be computed as a ratio of a zygosity of the mutation and the absolute copy number of the segment harboring the mutation in the tumor genome. For example, in FIG. IE, segment 154 harbors a mutation 158 having a zygosity of two and a fractional zygosity,, of 2 / 5 (0.4).

[0529] Mutation 158 shows an example where two of the three alleles corresponding to the same parent contain a mutation. Another possible, potentially more likely evolutionary scenario is that either (i) the mutation occurred after the CNV event (a late mutation) in which case only one allele would be mutated and the zygosity would be 1, or (ii) the mutation occurred before the CNV event (an early mutation), in which case all alleles belonging to the given parent would be mutated. In this case, the zygosity would either be equal to the (major) allele specific copy number, or to the minor allele specific copy number, which is given by the absolute copy number minus the (major) allele specific copy number. Given evolutionary considerations, more likely scenarios in the example shown would (i) either all alleles of the parent indicated in 164 are mutated, or (ii) the two alleles of the other parent would be mutated, or (iii) one allele of any parent would be mutated. In certain embodiments the zygosity can be limited only to likely evolutionary zygosities, namely 1, the major allele specific copy number, or the minor allele specific copy.

[0530] In certain embodiments, a cellularity of a mutation, denoted by p, may be determined. Cellularity as used herein refers to the fraction of tumor cells that harbor a given mutation. A mutation is said to be biologically clonal if all cancer cells in a tumor sample harbor the given mutation. Biologically clonal mutations are characterized by having a nominal cellularity of 1 (p = 1). In certain embodiments, approaches described herein determine cellularity estimates for mutations. In certain embodiments, a cellularity estimate is an estimated mean cellularity (e.g., indicating, if multiple tumor samples were obtained and sequence, a given mutation is estimated to be present, on average, in a fraction of tumor cells given by the mean cellularity). In certain embodiments, confidence intervals for cellularity estimates may be determined, with lower and upper bounds denoted μγ(ρ) and νγ(ρ), where y is the confidence level of the estimate (e.g., also referred to as degree of confidence or confidence coefficient). A cellularity confidence interval (CI) [μγ(ρ), νγ(ρ)], may, for example, indicate that, if multiple tumor samples were obtained, the estimated cellularity - 91 - 13241940vlAttorney Docket No. 2013237-1502would be on the interval [e.g., at or between μγ(ρ) and νγ(ρ)] y percent of the time. For example, in certain embodiments, a 95% CI lower and upper bound are determined. In certain embodiments, a 90% CI lower and upper bound are determined. In certain embodiments, a 68% CI lower and upper bound are determined.

[0531] In certain embodiments, a cellularity confidence interval (CI) may be defined as p ± ασp. For example, for a 95% CI, α = 1.96, for a 75% CI, α = 1.15, for a 50% CI, α = 0.67. In some embodiments, a standard deviation of cellularity opis calculated as σp= √(p(1 — p) / N), where N is total coverage at a given site or a number of reads mapping to a wildtype allele plus the number of reads mapping to alternate allele. In some embodiments, N is computed after selecting only reads with a quality score above a predetermined threshold. In some embodiments, a predetermined threshold is between about 25 and about 35. In some embodiments, N is computed after filtering poor quality reads. In some embodiments, where ρ = VAF / Px, a standard deviation of cellularity σpis given by σVAF / Px, where σVAFis a standard deviation of an observed variant allele frequency, given by √(VAF(1 - VAF) / N), and Pxis an expected allele frequency of a given variant allele assuming the mutation is clonal. For example, in certain embodiments, a 90% CI lower and upper bound are determined. In certain embodiments, a 95% CI lower and upper bound are determined. In certain embodiments, a 68% CI lower and upper bound are determined.

[0532] As illustrated in FIGs. ID and IE, and described herein, while tumor samples comprise, or are expected to comprise, cancer cells, they often also comprise - e.g., are contaminated by - normal cells. Accordingly, a tumor sample may also be characterized by a sample purity (1 - p) and / or a contamination fraction, p. Sample purity (1 - p) and / or contamination fraction may be determined, for example, directly and / or via modelling approaches based on sequencing data, as described in further detail herein. Additionally, or alternatively, upper and / or lower bounds for sample purity and / or contamination fraction may be estimated. For example, as described in further detail herein, in certain embodiments, at low physical sample purities, accuracy of estimation methods, such as tumor deconvolution, may be reduced such that purity estimates based on sequencing data are expected to be of insufficient accuracy to be used in and of themselves. In such cases, however, an upper bound (e.g., a maximum purity) and / or a lower bound (e.g., a minimum purity) may still be estimates and used in certain processing steps.- 92 - 13241940vlAttorney Docket No. 2013237-1502B. Tumor Modeling and Mutation Detection

[0533] In certain embodiments, determining properties of a tumor sample and detecting and characterizing mutations utilize tumor deconvolution (also referred to herein as tumor modelling) and mutation calling methods.B.i Tumor Deconvolution

[0534] As illustrated in FIG. IE, underlying physical properties of a tumor genome and segments thereof, such as absolute copy numbers, allele specific copy numbers, presence of mutations, etc., as well as sample properties, such as purity, may be reflected in sequencing data. However, due to statistical properties of sequencing data, various sources of error and noise, and scaling factors that may be introduced from run to run and sample to sample, determining genomic and sample properties from sequencing data is non-trivial. Accordingly, in certain embodiments, tumor modelling (also referred to as deconvolution) approaches use sequence data for tumor samples to determine characteristic features of a tumor genome and / or segments thereof, such as absolute copy numbers, allele specific copy numbers, as well as tumor sample properties, such as purity and average ploidy (e.g., wherein ploidy refers to an average of copy numbers across genes in a sample, e.g., weighted by gene length).

[0535] Values of parameters determined via tumor modelling approaches may include any of those described herein, such as sample parameters like tumor sample purity, tumor average ploidy, as well as local, segment-specific parameters, such as absolute copy numbers, allele- specific copy numbers, and the like, for various segments of the tumor genome. In certain embodiments, for example, once one or more tumor genome mutations are detected, tumor modelling may, additionally or alternatively, involve determining values of parameters that characterize properties of one or more particular mutations, such as zygosity, fractional zygosity, cellularity (p), and the like. As described in further detail herein, tumor modeling approaches, such as those of the present disclosure, may be leveraged for mutation detection, for example improving accuracy. For example, mutation detection may utilize purity estimates determined from tumor modeling. In certain embodiments, in turn, tumor modeling approaches may utilize results from mutation detection, for example to refine estimates of purity. Accordingly, in certain embodiments, tumor modelling and mutation detection may- 93 - 13241940vlAttorney Docket No. 2013237-1502be performed in tandem, for example, with outputs of each being repeatedly used to refine solutions of the other.

[0536] Turning to FIG.2A, in certain embodiments, a tumor modelling approach may, among other things, construct and fit one or more tumor models that accurately explain sequencing data - such as aligned reads for tumor and, optionally, normal genome for a patient. As shown in FIG.2A, in an example tumor modelling process 200, sequencing data may be obtained 202. As described herein, sequencing data 202 may comprise tumor sequencing data and / or normal sequencing data. In certain embodiments, sequencing data comprises tumor sequencing data and normal sequencing data (e.g., which may comprise replicates).

[0537] A tumor modelling process 200 may, at various steps, identify particular segments and / or particular classes or subsets of segments of a tumor genome 204. Among other things, certain subsets - e.g., types, classes - of segments have desired properties that makes them useful / appropriate for certain tumor models and determining values of particular parameters.

[0538] For example, as shown in FIG.2B, in certain embodiments, heterozygous segments (HS) may be identified. In certain embodiments, balanced segments may be identified. In certain embodiments, balanced heterozygous segments may be identified. In certain embodiments, primary balanced heterozygous segments may be identified. In certain embodiments, primary segments may be identified.

[0539] Balanced heterozygous segments may be identified using sequencing data, for example, by comparing a number of reads corresponding to each of two alleles, A and B, of a particular segment. For example, using tumor sequencing data, the number of reads corresponding to the A and B alleles of a given,segment can be determined and metrics that are indicative of whether the jthsegment is balanced, such as a ratio (e.g., NjA / NjB), relative fraction [e.g., NjA / (NjA + NjB), NjB / (NjA + NjB), etc.], difference (e.g., |NjA-NjB|), etc.), statistics, such as a --statistic [e.g., z = (NjA — p × Nj) / √(Nj × p × (1 — p))] are determined. In certain embodiments, a discrete metric, such as a binary balance state value - e.g., having a value of 0 or 1 depending on whether a particular segment is balanced or not - may be determined, for example, by comparing values of any continuous metrics such as any of the aforementioned metrics (e.g., ratio, relative fraction, difference, z-statistic, etc.) to a threshold. In certain embodiments, values of metrics indicating whether a particular segment - 94 - 13241940vlAttorney Docket No. 2013237-1502is balanced of multiple adjacent segments can be used in conjunction to filter and / or correct for noise, for example, via a contour filtering approach as illustrated in FIG.2C. In certain embodiments, constraints, such as ensuring SNPs be balanced in both normal and tumor data, may be imposed.

[0540] In certain embodiments, primary balanced heterozygous segments may be identified and used to determine initial parameters for tumor modelling. For example, as shown in FIG.2C, a number of reads for each segment may be determined and plotted against a tumor over normal segment read count ratio. The corresponding distribution may be fit, for example, to a one-dimensional Gaussian mixture model (1D-GMM) (e.g., a sum of Gaussians) and a primary component identified. In certain embodiments, primary segments may be identified, for example, as those segments whose tumor to read count ratio falls within the primary component [e.g., within a particular number (e.g., 1, 1.5) of standard deviations of from a mean of the primary component].

[0541] Turning again to FIG. 2A, in certain embodiments, a tumor modelling approach may use sequencing data corresponding to at least a portion of the tumor genome segments (and, optionally, normal genome) to compute observed values of one or more various parameters 206. Observed values may be, or be functions of, particular genomic and / or sample properties for which estimates are desired. For example, parameter values, such as absolute copy number and purity values, may be determined 212 from sequencing data by fitting 210 a tumor model 208 to various observed values computed from sequencing data corresponding to at least a portion (e.g., particular subsets) of the segments.

[0542] In particular, in certain embodiments, values of certain observed parameters and / or statistical distributions thereof, may be determined. For example, as illustrated in FIG. 2D, sequencing data can be used to compute, for a given heterozygous segment, an observed frequency of a major allele, denoted Px, and a residual error coordinate, denoted cr, which is a function of the number of reads for the tumor and normal segments. As will be clear from context, Px may be used, in reference to a heterozygous segment or SNP, refer to a SNP allele frequence, as it does in connection with FIG. 2D. When used in reference to a variant allele harboring a mutation, Px is used to refer to the allele frequency of the variant allele (e.g., 1 - Px(k ). The graph in FIG.2D plots, for each / th segment of a plurality of segments e.g., heterozygous segments, e.g., primary heterozygous segments), Px against a ratio of tumor to normal read counts (e.g., NT, J / ANJ) for a plurality of heterozygous segments (HS) from sequencing data. Both the major allele frequency, Px, and tumor to normal read - 95 - 13241940vlAttorney Docket No. 2013237-1502count ratio can be determined directly from sequencing reads, the former is determined via Eq. (B1), below, where Nx,j is the observed read counts from tumor sequencing data that correspond to the major allele of thejthsegment and the Nj are the total read counts corresponding to thejthsegment.Eq. (B1) PX= Nxj / Nj

[0543] For an idealized pure sample, without noise, Px ~ CNx I CNj and NT,j / NN,j ~ CNj / 2. Accordingly, for a given absolute copy number CAj, possible values for CAx and Px vary as shown in Table 1, below.Table 1. Absolute Copy Numbers and idealized corresponding tumor-to-normal read count ratios, possible allele-specific copy numbers, and allele frequencies.Absolute Copy Tumor to Normal Possible Cx Possible Px Number (CN) Read Counts Values Values1 0.5 1 0, 12 1 1, 2 0.5, 13 1.5 2, 3 0.66, 14 2 2, 3, 4 0.5, 0.75, 1 5 2.5 3, 4, 5 0.6, 0.8, 1

[0544] The grid and black dots in the graph shown in FIG.2D represent nodes -locations in the space of possible solutions for various absolute copy numbers (CAj’s), computed using a tumor model that takes into account sample purity and a residual error. In the approach illustrated in FIG.2D, in the tumor model, the statistical spread of values about an ideal node coordinate, shown as the black dots, is modelled as a two-dimensional Gaussian along the major allele frequency, Px, and a residual error coordinate, calculated based on tumor and normal segment read counts. In this way, each node corresponds to a (cr, Px) value pair. In the particular example approach illustrated in FIG.2D and further in Example 1, the following expression was used: er= tobs— pth· nobs, where tobsis the observed number of reads in the tumor sample mapping to the given segment, nobsis the observed number of reads in the normal sample mapping to the given segment, pthis the predicted boundary between absolute copy numbers. Example 1 further provides an example formula for determining er. By fitting a tumor model to the distribution of Px and ercomputed from - 96 - 13241940vlAttorney Docket No. 2013237-1502tumor and normal sequence data, estimated purity values and copy numbers that maximize the quality of the fit can be determined. Model fit quality may be evaluated using a variety of fit quality metrics, including, without limitation, maximum likelihood (ML) scores, K-statistics, and specific cluster density.

[0545] Fitting a tumor model to sequence data in this manner may be performed via a variety of approaches and optimization techniques, including, without limitation, maximum likelihood estimate (MLE) as described in Example 1.

[0546] In this manner, values of parameters such as sample purity, absolute copy number (C V), and allele specific copy number (CTVx), may be determined. Fitting procedures may, for example, be used to determine values of sample purity, absolute copy number, and allele specific copy number that best fit the observed distribution of Px and er. That is, in certain embodiments, tumor models 208 may have variable parameters, such as CN, CNx, and μ, and be used to compute predicted values of observable quantities such as Px, er, as well as, in certain embodiments, functions and or statistical distributions thereof. Tumor deconvolution may, accordingly, determine a tumor model and / or values of variable parameters thereof that best fit observable data, e.g., based on values of fit quality metrics as described above.

[0547] In certain embodiments, various parameter values may be varied and optimized together, in a same manner (e.g., by maximizing or minimizing a value of an objective function based on one or more fit quality metrics). In certain embodiments, parameter values, such as purity and absolute copy number, may be determined in a multi-step fashion. For example, as illustrated in FIG.2B, in certain embodiments, a plurality of tumor models 220 may be fit, e.g., independently, to observed data, such as statistical distributions of Px and erHEIGHT="12" WIDTH="11" SRC="imgf000097_0001.tif" / > Each tumor model of the plurality 220 may correspond to and use a different potential absolute copy number solution and be separately fit, e.g., via a MLE-based approach, to observed data, to determine, for each model, a corresponding MLE-based purity estimate, thereby determining a plurality of MLE-based purity estimates 222. In certain embodiments, for each model, a ML score may be determined, for example as an overall likelihood at the MLE-based purity, for example as described in Example 1. In this manner, an overall purity estimate determined as the MLE-based purity estimate for the best performing model 224 and an absolute copy number estimate determined as the absolute copy number corresponding to, and used by, that best performing tumor model. In certain embodiments, a best performing model 224 may be identified and evaluated, for example, via - 97 - 13241940vlAttorney Docket No. 2013237-1502a quality control step 226, for example, based on additional fit quality scores such as ML score, K statistic, specific cluster density, and combinations thereof.

[0548] In certain embodiments, for example, since absolute copy number (CN) is a segment-specific property, the fitting approaches may be performed using a subset of the segments, such as those with a same absolute copy number, such as primary heterozygous segments or primary balanced heterozygous segments.

[0549] In certain embodiments, an initial tumor model selection may pass 228a additional quality control 226, and absolute copy number, purity, etc. values determined via the initial tumor model confirmed 230a and used for further processing (e.g., mutation identification, property calculation, clonality type assignment, etc.).

[0550] In certain embodiments, an initial tumor model may be rejected 228b at quality control step 226. In certain embodiments, where initial tumor model is rejected 228b, MLE-based purity estimates 222 from plurality of tumor models 220 may be used to compute, in lieu of an estimated sample purity, an upper and / or lower bound on sample purity. For example, in certain embodiments, based on scores determined in quality control step 226, it may be determined that data is insufficient and / or sources of uncertainty, such as noise etc., are too high to allow for selecting one of plurality of tumor models 220 over others with sufficient confidence. Accordingly, while a single estimate of purity and absolute copy number - obtained via selection of a particular, single, tumor model out of the plurality 220 - may not be obtainable, an upper and / or lower bound of a purity estimate can still be obtained, for example by taking a maximum and / or minimum of the MLE-based purity estimates 222.

[0551] In certain embodiments, a quality estimation step 232 may be performed to evaluate a quality of a determined solution. Quality estimation 232 may be performed based on a variety of metrics. For example, in certain embodiments, assessing quality control for a tumor deconvolution model comprises determining a persistence distance. Persistence distance may be determined as the mean run length of a constant copy number in a given chromosome, averaged over all chromosomes. In certain embodiments, an error corrected persistence distance is used and computed as the average run length of a given error corrected (EC) absolute copy number across the exome. Units of this parameter may be given in genes. Without wishing to be bound to any particular theory, a short persistence distance is believed to indicate a noisy deconvolution solution and / or a potential false solution. In certain- 98 - 13241940vlAttorney Docket No. 2013237-1502embodiments, a global cluster density metric is determined and used for tumor deconvolution quality control. Global cluster density may be used to determine / evaluate a degree to which genes in a given cluster have the same cluster assignment, for example as illustrated in FIG.2E. In certain embodiments, a low global cluster density may be indicative of either poor quality deconvolution or a false solution. In certain embodiments, a high-quality deconvolution will have a high global cluster density score.

[0552] For example, as shown in FIG.2E, to compute a global cluster density metric, each heterozygous segment (HS) in a given cluster is mapped back to the genome and the local density of gene g is determined as the percent of genes in the local neighborhood of gene g that have the same cluster assignment as g along the given chromosome, as shown in the figure. Next, a median of the local density per cluster per chromosome is calculated for all genes belonging to a given cluster on the corresponding chromosome. Finally, a local density per cluster is calculated by taking the median local density for each cluster across all chromosomes. The global cluster density is calculated as the weighted average of local cluster densities.

[0553] In certain embodiments, global cluster density can be mathematically(cdetermined as follows: Let cluster c of the 2d-GMM model contain Ng genes. The GMM model assigns each gene g belonging to cluster c a given node, denoted n(^). Each gene g can be mapped to a certain position on a chromosome, chr. The local density of a given gene Diocai g)onthe genome can be determined as a percentage of genes within a small window spanning, e.g., +L = 20 genes, from g on chromosome chr, which are also assigned to node n(^). Nearby genes have a minimal coverage: nobs> 1500. Thus, Dlocal(g) is the percent of genes nearby to g that have the same cluster assignment as g. This percent score, Dlocal(g ∈ c),canbe calculated for every gene g in cluster c. To avoid spurious results, only chromosomes with at least 20 genes ∈ c may be considered, and for these chromosomes, the median local density per cluster is calculated: mDlocal(chr, c) = median [Dlocal(g ∈c & g E c r)]. Finally, the median across all chromosomes is calculated as such: mDlocal(c) = median [mDlocal(chr, c)]. The global density is determined as the mean local chrdensity across all clusters weighted by the size of each cluster, wc: DDglobal(CNpc) =wcmDlocal(c), where wc= Ngc / ΣNg. To improve the robustness of this metric,- 99 - 13241940vlAttorney Docket No. 2013237-1502clusters selected for this calculation are determined to satisfy the following conditions: (1) CNmut< 2CNpc(2) nobs> nth, and (3) > 75.

[0554] Accordingly, in certain embodiments, tumor deconvolution processes used in in connection with clonality classification technologies of the present disclosure may, for example, in certain instances, determine an estimated upper bound purity, additionally or alternatively to a purity and / or ploidy estimate. For example, estimates of ploidy and / or purity may become less accurate for low purity tumor samples, when a sequencing bias is present, and / or when the tumor sample deviates in a significant way from tumor model assumptions. Accordingly, in certain embodiments, tumor deconvolution approaches of the present disclosure may account for loss in accuracy by generating an upper bound estimate of tumor purity that may, in turn, be compared to a threshold purity value. Where the estimated purity upper bound falls below the threshold purity value, a sample and / or estimates of purity and / or ploidy may be identified as potentially inaccurate, for example, via an error flag, replacing purity and / or ploidy estimates with an indeterminate value, etc. In certain embodiments, sample purity may be used in a quality control check.

[0555] In certain embodiments, if tumor deconvolution fails one or more quality control tests 226, then a bound (e.g., an upper bound) on purity is determined, and clone types will not be assigned to mutations. In further embodiments, even if all quality control tests 226 pass, if the tumor model fails additional quality estimation steps 232 (e.g., low persistence distance and / or low global cluster density), then clone types might be ignored for the purpose of prioritizing neoantigens. In further embodiments, even if all quality control tests 226 pass, if the estimated purity is below a certain threshold, then then clone types might be ignored for the purpose of prioritizing neoantigens. In certain embodiments, if quality estimation and / or purity threshold tests fail, clone types may still be determined and used for prioritizing neoantigens. For example, in certain embodiments, e.g., responsive to a determination that the tumor deconvolution is unsuccessful (e.g., based on failure of one or more of the quality control metrics and / or quality estimation steps), for each particular candidate mutation, approaches described herein may be used to determine a plurality of prospective clone types, each prospective clone type determined using a distinct one of the plurality of tumor models. The plurality of prospective clone types to determine a (e.g., final) corresponding clone type for the particular candidate mutation, for example selecting the corresponding clone type as a most common one of the prospective clone types, e.g., if a- 100 - 13241940vlAttorney Docket No. 2013237-1502minimum fraction (up to all) of the prospective clone types are the same (otherwise selecting the corresponding clone type as undetermined).

[0556] In this manner, clonality classification technologies described herein may ensure that determinations and predictions are sufficiently accurate. Maintaining a high level of accuracy and confidence in predictions is particularly important for immunotherapy applications (e.g., as opposed to research or diagnostic tools), wherein selecting or rejecting a particular neoepitope candidate for inclusion in a therapeutic may dramatically impact efficacy.

[0557] Further methods and systems for performing tumor deconvolution are provided in Section D herein (“Tumor Deconvolution and Mutation Detection”).B.ii Mutation Detection

[0558] Turning to FIG. 3 clone type classification technologies of the present disclosure may leverage approaches for detecting mutations 300 within a tumor genome under test. Mutation detection - also referred to as “mutation calling” - may involve identifying mismatches between (e.g., aligned) reads of tumor sequencing data and normal sequencing data 302. Beyond identifying mismatches between tumor and normal sequencing data, mutation callers 304 may employ techniques, such as statistical sorting, assembly, etc., to distinguish true or likely physical mutations from artifacts that may result from, for example, noise and / or errors introduced during sample preparation, sequencing, and data preprocessing steps, such as alignment.

[0559] Example mutation callers include, without limitation, MuTect2, described in Cibulskis et al., “Sensitive detection of somatic point mutations in impure and heterogeneous cancer samples,” Nat. Biotechnol. 31 (3):213-219 (2013), Strelka, described in Saunders et al., “Strelka: accurate somatic small-variant calling from sequenced tumor-normal sample pairs,” Bioinformatics, 28(14) 1811-1817 (2012), VarScan2, described in Kobolt et al., “VarScan 2: somatic mutation and copy number alteration discovery in cancer by exome sequencing,” Genome Res. 22(3) 568-76 (2012), SomaticSniper, described in Larson et al., “SomaticSniper: identification of somatic point mutations in whole genome sequencing data,” Bioinformatics, 28, 311-317 (2012), JointSNVMix, described in further detail in Roth et al., “JointSNVMix: a probabilistic model for accurate detection of somatic mutations in normal / tumour paired next generation sequencing data,” Bioinformatics 28, 907-913, as well - 101 - 13241940vlAttorney Docket No. 2013237-1502as Illumina’s Dynamic Read Analysis for GENomics (DRAGAN), see e.g., illumina.com / products / by-type / informatics-products / basespace-sequence-hub / apps / edico-genome-inc-dragen-somatic-pipeline.html.

[0560] As described in further detail herein, while these mutation callers detect potential mutations, they do not include or contemplate the approaches for determining and assigning mutations to discrete clone types that are introduced in the present disclosure and are developed based on key design considerations relevant for selecting effective targets for inclusion in clinically relevant cancer immunotherapies. Moreover, these mutation callers also do not rely on a tumor model to improve the performance (i.e., accuracy and sensitivity) of mutation detection. Additionally, or alternatively, the present disclosure also includes a filtering step, whereby potential mutations (e.g., which may be identified via mutation callers such as those above) may be filtered to reject potential mutations that are either likely false positives or excessively rare subclones.

[0561] Accordingly, as illustrated in FIG.3, in certain embodiments, one or more mutation callers 304 may be used to determine an initial list of potential mutation(s) 306 from sequencing data 302. In certain embodiments, one mutation caller may be used. In certain embodiments, a plurality of mutation callers may be used. For example, multiple mutation callers may be utilized, with each generating their own corresponding list of potential mutations. These may then be merged, for example, by retaining only those mutations detected by each, or at least n e.g., 2, 3, 4, etc.), mutation callers, to create initial list of potential mutation(s) 306. Additional filtering approaches, such as mutation confidence score (MCS)-based filtering and / or RMCS filtering, described in further detail in Section C.iv below, may be used to reject and filter out rare subclones 308, resulting in a final list of candidate mutations 310 (e.g., high confidence mutations) for clone type classification.C. Clonality Classification and Screening Subclonal Mutations

[0562] Among other things, without wishing to be bound to any particular theory, immunotherapies, such as personalized cancer vaccines (PCV), personalized T-cell therapies, TCR therapies, etc., targeting biologically clonal mutations are expected to be more effective than therapies targeting biologically subclonal mutations since, among other things, biologically clonal mutations are, in principle, present in all tumor cells. Accordingly, absent immune escape, therapies targeting biologically clonal mutations offer the potential to- 102 - 13241940vlAttorney Docket No. 2013237-1502eradicate an entire tumor mass. Since biologically subclonal mutations are not present in all tumor cells, targeting a particular biologically subclonal mutation are not, in general, expected to eradicate an entire tumor mass, even if each cell harboring that particular biologically subclonal mutation is eradicated. However, targeting subclonal mutations may be clinically useful if, for example, different subclones are present in different populations of tumor cells and / or killing of tumor cells leads to epitope spreading, i.e. developing an immune response against neoepitopes that were not original targets of the immunotherapy. Therefore, under certain circumstances, subclonal mutations could have certain potential clinical benefit.

[0563] Accordingly, among other things, the ability to accurately identify and differentiate between biologically clonal and biologically subclonal neoepitopes is particularly valuable for selecting neoepitopes for inclusion in a PCV, since it allows one to prioritize biologically clonal neoepitopes and / or down-prioritize biologically subclonal neoepitopes.

[0564] Additionally, or alternatively, as described in further detail herein, clonality state classification technologies of the present disclosure include the recognition that for certain cancer types, such as those with low tumor mutation burden (TMB), it may be challenging or infeasible to identify and use a sufficient number of biologically clonal mutations. Accordingly, in certain embodiments, clonality classification approaches described herein go beyond distinguishing between biologically clonal and biologically subclonal mutations, by allowing mutations that are biologically subclonal to be sub-divided into multiple discrete categories, based on, for example, predicted or estimated prevenances in tumor cells. In this manner, mutations that are biologically subclonal, but nonetheless prevalent, can be identified and included as e.g., neoepitope targets, in situations where, for example, few biologically clonal mutations can be identified (or those that are found to be unsuitable for other reasons, such as poor predicted MHC binding, low expression rates, etc.). Among other things, the present disclosure includes an insight that, rather than discarding biologically subclonal mutations, including mutations that while being biologically subclonal are, nonetheless, prevalent within tumor cells can still allow for effective tumor eradications, by virtue of including multiple neoepitope targets as well as effects, such as epitope spreading.- 103 - 13241940vlAttorney Docket No. 2013237-1502C.i Assigning Mutations to Discrete Clone Types

[0565] In certain embodiments, mutation clonality analysis technologies of the present disclosure use a classification problem framework to determine and assign clonality classification state labels to mutations. The approach and framework followed by the methods and systems of the present disclosure, accordingly, contrasts with other techniques that attempt to rely on estimation frameworks, such as the ABSOLUTE method described in Carter et al. “Absolute quantification of somatic DNA alterations in human cancer.” Nat. Biotechnol. 30, 413-421 (2012) and / or clustering VAF values, for example, as performed by the PyClone method, described in Roth et al., “PyClone: statistical inference of clonal population structure in cancer,” Nat. Meth. 11, 396-398 (2014). Among other things, a classification framework, as opposed to estimation and / or clustering frameworks, ensures that each mutation is assigned a concrete, particular clone type, thereby providing actionable information that can be leveraged to prioritize and select mutations as neoepitope targets.

[0566] FIG. 4A shows an example process 400 whereby clone types can be determined for tumor mutations and used to select neoepitope targets. In certain embodiments, as described herein, sequencing data 402 may be generated, for example, via next generation sequencing technologies, SNP arrays, etc. and used as input for tumor modeling and mutation detection 404 to determine a set of candidate mutations 406. In certain embodiments, for example as illustrated in FIG. 4A, tumor modeling and mutation detection are coupled, e.g., and performed together (e.g., in an iterative fashion). In certain embodiments, at least a portion of candidate mutations 406 may be evaluated via a clonality classification 408 step in order to assign to each evaluated candidate mutation a corresponding clonality classification state, e.g., selected from a set of discrete (e.g., predefined potential) clone types. In certain embodiments, for example, by virtue of an undetermined clone type, described in further detail below, clonality classification process 400 may allow for each and every detected candidate mutation of set 406 to be assigned a particular, discrete, clone type.

[0567] For example, additionally or alternatively to providing cellularity estimates for detected mutations, methods and systems of the present disclosure classify mutations into a discrete set of clonality classification states (also referred to as “clone types”), wherein mutations of a given clone type are predicted to have similar biological clonality characteristics and / or prevalence within cancer cells of a subject’s tumor. For example, in certain embodiments, clone types used in connection with technologies of the present- 104 - 13241940vlAttorney Docket No. 2013237-1502disclosure include classes or labels, such as one or more “clonal state(s)”, one or more “prevalent subclone state(s)”, a “minor subclone state”, and an “undetermined state”, described in further detail herein.

[0568] In certain embodiments, by assigning each mutation one of a discrete set of clone types 410, mutations can be prioritized and selected 412, for example, for inclusion as a construct, such as a therapeutic agent 414, based on their assigned clone type(s). For example, mutations classified as belonging to a clonal state can be prioritized as one group, mutations classified as belonging to a prevalent subclone state [e.g., a state which captures biologically non-clonal mutations with a cellularity greater than about fifty percent (50%), as described in further detail herein] can be prioritized as a second group, and mutations that are classified as having a minor subclone state [e.g., having a cellularity less than about fifty percent (50%)] can be prioritized as a third group. Other approaches for using discrete clone types to differentiate between and / or prioritize mutations may be used. For example, a binary schemes, such as a scheme that prioritizes mutations classified as belonging to a clonal state over mutations assigned other clone types (e.g., biologically subclonal mutations, such as prevalent subclones, minor subclones, as well as, in certain embodiments, mutations classified as having an undetermined clone type) or a scheme that prioritizes mutations assigned clone types other than a minor subclone state over those that are classified as belonging to the minor subclone state, are also contemplated. As described in further detail herein, clonality state classification and prioritization schemes may be tailored, for example, depending on cancer type and expected mutation burden, number of identified targets, etc. In certain embodiments, for example, multiple clone types spanning a full range of cellularity values may be desired. For example, as described and demonstrated, e.g., in Example 1, herein, mutations classified as prevalent subclones - biologically subclonal mutations that, while not present in all tumor cells, are found in a large fraction of cells - can serve as large reservoirs of targets, which can be important in low mutation burden samples.

[0569] Clone type assignments used in technologies of the present disclosure may be performed based on various criteria, such as values of various estimated parameters of candidate mutations and / or statistical distributions and / or tests based thereon. These include, without limitation, parameters, such as estimated cellularity, cellularity confidence interval upper and / or lower bounds, P- values computed based on various statistical tests, etc.- 105 - 13241940vlAttorney Docket No. 2013237-1502C.i.l Hypotheses Tests and P -Values

[0570] In certain embodiments, clone type assignments performed via technologies of the present disclosure utilize statistical tests. For example, in certain embodiments, a statistical hypothesis test in which a null hypothesis that a given mutation is biologically clonal is evaluated. For example, in certain embodiments, a P- value for a null hypothesis that a particular mutation is clonal may be determined according to Eq. (B2):Eq. (B2) Prob{p < 1} = Proh{P£bs< P^11} = F(NX\NX+ Nz, P^h)where p is an observed cellularity, Pxbsis an observed allele frequency, and F is a Binomial cumulative distribution function. Accordingly, a P-value determined, for example, via equation 2, may represent a probability that an observed cellularity for a particular given mutation may be less than one, although that particular mutation is biologically clonal. That is, biologically clonal mutations have nominal cellularities of 1, however the statistical nature of sequencing data, sources of noise, such as instrument error, and the like, may cause observed cellularities for mutations that are biologically subclonal to deviate from 1. The P-value computed via equation 2 gives a likelihood that, for a given mutation, an observed cellularity below 1 is due to statistical factors and noise, as opposed to the mutation itself being biologically subclonal.

[0571] Additionally or alternatively, a P-value may be computed for cellularity values exceeding one. These may be computed via a binomial distribution, e.g., if Problp < 1) = binocdf(Ax, n_std, Px) and Prob(p > 1) = 1 -binocdf(Ax- 1, n_std, Px(binocdf referring to F, the Bionomial cumulative distribution function), where Px is the allele frequency of the variant allele. Noise, instrument error, etc. can cause cellularity values to exceed one (but nominal cellularity cannot be greater than one, physically) - or errors in parameter estimates, model assumptions and the like, may cause abnormally high cellularity estimates. P-values for cellularities above one may, accordingly, be used to test whether parameter estimates and / or model assumptions on which they are based are inaccurate for particular mutations, rendering clone type assignments inaccurate.

[0572] In certain embodiments, using a null hypothesis of biological clonality (i.e., that a mutation is biologically clonal, as opposed to biologically subclonal) is advantageous since it allows a maximum likelihood estimation of zygosity and for rigorous and optimal- 106 - 13241940vlAttorney Docket No. 2013237-1502testing the statistical likelihood of the null hypothesis. Among other things, this hypothesis testing framework provides actionable and concrete information regarding clonality.

[0573] In certain embodiments, testing a null hypothesis of biological clonality is superior to testing an alternative hypothesis of biological subclonality. Among other things, determining a probability of biological subclonality does not provide the same kind of actionable information that the test of the biological clonality hypothesis does. Moreover, testing a hypothesis of biological subclonality requires invoking a model of subclonality, which is essentially unknown and complex unfounded biological model for subclonality containing multiple unknown parameters, greatly limiting its power. In certain embodiments, under a null hypothesis of subclonality, the zygosity of the mutation is unknown, requiring all possible zygosities to be considered, whereas under a null hypothesis of clonality, zygosity can be directly estimated, for example, under a maximum likelihood framework. Finally, determining a probability that a mutation is clonal or subclonal does not have concrete thresholds that can be assigned, whereas in a hypothesis testing framework significance levels can be invoked. For example, as described in further detail herein, methods such as ABSOLUTE utilize complex models for biological subclonality that involve multiple unknown parameters.

[0574] In certain embodiments, a P- value for a particular mutation, such as a P value determined via a hypothesis test as described above, may be compared with a threshold value for rejecting the null hypothesis (e.g., a P- value threshold). The particular mutation may then be assigned a clone type (e.g., label) based on whether the null hypothesis can be rejected. For example, with regard to equation 2, above, if a determined P value for a particular mutation is low, then it is unlikely that a measured cellularity can be explained by noise fluctuations, assuming that mutation is biologically clonal and, accordingly, the null hypothesis - that the mutation is biologically clonal - can be rejected. For example, P value threshold values such as IxlO’7(IxlO’5%), IxlO’6(IxlO’4%), IxlO’5(0.001 %), IxlO’4(0.01 %), IxlO’3(0.1 %), 0.01 (1%), 0.05 (5%), 0.1 (10%) may be used.

[0575] In certain embodiments, a P- value threshold to which a P value determined for a particular mutation is compared depends on a [e.g., known and / or determined (e.g., estimated)] sample purity. That is, in certain embodiments, a P- value threshold may be a function of sample purity, or, additionally or alternatively, multiple different P- value thresholds, each associated with a particular sample purity range or cutoff, may be used. For example, in certain embodiments, a P-value threshold is decreased with purity, which, in - 107 - 13241940vlAttorney Docket No. 2013237-1502certain embodiments, effectively increases the stringency of a statistical test. Without wishing to be bound to any particular theory, it is believed that lower sample purity increases overall uncertainty by increasing noise. Decreasing a P- value threshold for lower sample purities, accordingly, may compensate for increased noise and avoid misclassifications.C.i.2 Cellularity Estimates

[0576] In certain embodiments, criteria used to classify a particular mutation as belonging to a particular clone type may be based on cellularity estimates. For example, in certain embodiments, approaches of the present disclosure may determine, for a particular mutation, a corresponding cellularity value, p. Accordingly, in certain embodiments, a cellularity value determined for a particular mutation may be compared with one or more cellularity threshold values or ranges, and the mutation assigned a particular clone type based at least in part on whether its corresponding cellularity value lies above or below one or more particular threshold values and / or within a particular range. Example cellularity threshold values range from 0 to 1, such as 0.3, 0.4, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, 0.95, and 1.

[0577] In certain embodiments, as described in further detail herein, mutations that, for example, for which a P value determined via equation 2 above, exceeds a desired threshold may be assigned clonality classification states that reflect their biologically subclonal nature, but account for their prevalence in terms of cellularity. For example, as shown in the detailed experimental data presented in Example 1, below, biologically subclonal mutations may nonetheless be present at high cellularities and, accordingly, be prevalent and potentially viable targets for inclusion in immunotherapies, particularly where tumor mutational burden is low, and a limited number of strictly biologically clonal targets are available.C.i.3 Cellularity Confidence Intervals (Cis)

[0578] In certain embodiments, additionally or alternatively, confidence intervals may be determined for a cellularity, for example, based on an assumed statistical distribution e.g., Binomial) and used, e.g., analogously, to assign a mutation to a particular clone type.- 108 - 13241940vlAttorney Docket No. 2013237-1502

[0579] In certain embodiments, in addition to or instead of comparing a determined cellularity value for a given mutation to a cellularity threshold, one or more confidence interval bounds may be compared to various corresponding cellularity thresholds. For example, in certain embodiments, upper and / or lower bounds of a 95% confidence interval for an estimated cellularity may be determined. In certain embodiments, upper and / or lower bounds for a 90% confidence interval for an estimated cellularity may be determined. In certain embodiments, upper and / or lower bounds for a 68% confidence interval for an estimated cellularity may be determined. In general, upper and lower cellularity confidence interval bounds, w7( ) and v7(p) respectively, may be determined, where y denotes a particular percentage confidence level, such as 99.7%, 95%, 90%, 85%, 80%, 75%, 68% etc.

[0580] In certain embodiments, an estimated cellularity, p, and / or upper and / or lower confidence interval bounds, Wy(p) and v7(p), may be compared with one or more cellularity threshold values. In certain embodiments, cellularity threshold values may vary depending on sample purity, for example, so as to account for increasing noise (and, e.g., accordingly, increased likelihood of mis-classifications) with decreasing sample purity.

[0581] In certain embodiments, for example, to ensure accuracy of clone type assignments, clonality analysis technologies of the present disclosure include an “undetermined” category, which is designed to capture all cases where a confidence in the classification is low due to statistical considerations and / or due to a low confidence in inference of any of the parameters that lead to the classification. This includes identification of potentially incorrect absolute copy number estimations in either the normal sample or the tumor sample.

[0582] Examples of how hypothesis test values, determined cellularity values, and corresponding cellularity confidence intervals can be used to assign mutations to various clone types are included below in Section C.iii and Example 1.C.ii Certain Advantages of Discrete Clone Types

[0583] Among other things, classifying mutations according to discrete clone types in accordance with approaches described herein provides several advantages over alternative approaches, such as sorting mutations according to values of continuous variables, such as variant allele frequency (VAF) and cellularity e.g., alone). Certain advantages of discrete clone type assignments over approaches based on sorting and prioritization on the basis of - 109 - 13241940vlAttorney Docket No. 2013237-1502continuous variables alone are particularly relevant in the context of selecting neoepitope targets for cancer immunotherapies.

[0584] For example, sorting mutations according to VAF and / or cellularity, e.g., for prioritization in the context of immunotherapy is suboptimal since VAF and cellularity need not be correlated when copy number variations are present. Accordingly, mutations with higher VAFs can have lower cellularities in the presence of copy number variations.Additionally, VAF is a continuous parameter and, therefore, it cannot be used to separate mutations into discrete groups of functionally equivalent mutations in terms of their biological clonality characteristics. For example, all biologically clonal mutations are functionally equivalent in terms of clonality, but if sorted by their VAF will be ranked by a random order, adding noise to a prioritization algorithm based on VAF. Third, VAF is susceptible to stochastic errors, which are then carried over to sorting, thereby adding further noise and errors to prioritization. Stochastic errors can be high when, for example, the absolute coverage at a given site is low or when tumor sample purity is low and a mutation is covered by a small number of reads.

[0585] Although cellularity is a better predictor of the biologically clonal nature of a mutation than the VAF is, sorting mutations by their cellularity alone for prioritization is a suboptimal solution because of the following two reasons. First, cellularity, like VAF, is a continuous parameter and, therefore, does not separate mutations into discrete groups in which mutations are functionally equivalent. For example, experimental data shows that even when purity is 100%, observed cellularity values of biologically clonal mutations can be as low as 80%, and when the purity is approximately 50%, observed cellularities can be as low as 50%. Accordingly, when sorting by cellularity, mutations that are functionally equivalent would obtain, by definition, different rankings, essentially adding noise to the process of prioritization.

[0586] Second, cellularity is derived from the VAF, and, therefore, like the VAF, cellularity suffers from stochastic noise. In the example above, when purity is low (-50%), the width of the distribution of cellularities of biologically clonal mutations is much wider and can reach 50%, thus overlapping with cellularity distributions of biologically non-clonal mutations, including mutations with very low nominal cellularities. Therefore, simply sorting mutations by their cellularity can lead to ranking errors during prioritization.- 110 - 13241940vlAttorney Docket No. 2013237-1502

[0587] Examples 1-13 include data that demonstrates these shortcomings and, accordingly, advantages of the clonality state classification approach described herein.C.iii Example Clonality State Classification Procedures

[0588] As described herein, clone type classification technologies may be used to assign mutations to a plurality of different clone types, including, but not limited to, on the one hand, clonal states capturing mutations that are determined to be and / or have a sufficiently high likelihood of being biologically clonal, and, on the other, subclonal classification states for mutations that are determined to be and / or have a sufficiently high likelihood of being biologically subclonal.C.iii.1 Criteria and Conditions for Clone Type Assignments

[0589] In certain embodiments, a plurality of potential clonality classification states to which a mutation may be assigned to includes one or more clonal classification states that aim to capture (e.g., are designed to encompass) mutations that have a sufficiently high likelihood of being biologically clonal mutations and / or, in certain embodiments, are so prevalent as to be very nearly biologically clonal. In certain embodiments, a mutation may be classified as belonging to a particular clone type and / or belonging to one of a set of predefined clonal classification states based on criteria, including conditions on a given mutation’s estimated and / or observed cellularity, cellularity confidence intervals, and P values determined via hypotheses tests as described, e.g., with regard to equation 2, above.

[0590] Clonal States. In certain embodiments, a set of discrete clone types comprises one or more clonal states, to which a mutation may be assigned. In certain embodiments, assignment of a candidate mutation to a clonal state is indicative of a high likelihood that the candidate mutation is biologically clonal and / or highly prevalent. For example, in certain embodiments, mutations classified as belonging to a clonal state may be determined and / or estimated to be present in about 70% or more [e.g., 75% or more (e.g., about 80% or more)] tumor cells.

[0591] For example, a given mutation may be classified as belonging to a clonal state based on criteria regarding estimated cellular values and / or confidence interval bound(s) determined for the given mutation. For example, criteria for classifying a mutation as - Ill - 13241940vlAttorney Docket No. 2013237-1502belonging to a particular clonal state may include a requirement that a cellularity confidence interval bound, such as a lower bound, be greater than or equal to a particular cellularity CI threshold value. In certain embodiments, the cellularity CI threshold is a purity dependent threshold, for example, having a first value when the purity is above a particular purity threshold and a second value when the purity is below the particular purity threshold. The second value may be greater than the first, thereby increasing the stringency of the criteria at lower purities.

[0592] In certain embodiments, for example, where multiple sub-classes of clonal classification states are used - such as PUTATIVE CLONAL and NEARLY CLONAL classes described in Example 1 - different criteria associated with each sub-class may include different cellularity threshold values. In certain embodiments, cellularity threshold values may also vary with sample purity. For example, cellularity threshold values may increase as sample purity decreases, thereby imposing increasingly stringent conditions that mutations have overall higher cellularity values and / or increased confidence interval lower bounds in order to be classified as belonging to clonal states.

[0593] Additionally, or alternatively, e.g., alone or in combination with cellularity conditions, criteria associated with a clonal classification state may include a condition that a P-value associated with a particular hypothesis test (e.g., a null hypothesis that a mutation is biologically clonal) should be above a particular P- value threshold. For example, where a P value is associated with a null hypothesis that a mutation is biologically clonal and reflects a probability that the given mutation’s observed cellularity is less than one, criteria for assigning a mutation to a clonal classification state may include a condition that the P value is greater than a particular P value threshold.

[0594] Minor Subclone States. In certain embodiments, a plurality of potential clone types to which a mutation may be assigned includes one or more minor subclonal states. As described herein, minor subclonal states aim to capture e.g., are designed to encompass) mutations that have a sufficiently high likelihood of being biologically subclonal mutations and are sufficiently rare - that is appear in a sufficiently low fraction of tumor cells - so as to make low priority and / or poor candidates for targets of, e.g., cancer immunotherapies. Minor subclonal states may be assigned based on criteria including conditions associated with a given mutation’s estimated and / or observed cellularity, cellularity upper and / or lower confidence interval bounds, and hypothesis test P values. For example, in certain embodiments, criteria associated with a minor subclonal classification state includes- 112 - 13241940vlAttorney Docket No. 2013237-1502conditions that a cellularity value and / or confidence interval bound determined for a given mutation should be below a particular cellularity threshold.

[0595] Undetermined States. In certain embodiments, a plurality of potential clone types to which a mutation may be assigned includes an undetermined (e.g., referred to in certain examples as “N / A” for short) state, capturing mutations for which evidence to make a reliable classification is determined to be insufficient. A mutation may be classified as undetermined based on criteria including conditions on a given mutation’ s copy number estimates, estimated and / or observed cellularity, cellularity upper and / or lower confidence interval bounds, and hypothesis test P values. For example, in certain embodiments, a particular mutation may be assigned to the undetermined state based at least in part on a copy number estimate for the particular mutation being equal to zero. In certain embodiments, a particular mutation may be assigned to the undetermined category based at least in part on a cellularity value being above a particular threshold value. In certain embodiments, a particular mutation may be assigned to the undetermined category based at least in part on a cellularity value being above a cellularity threshold value that exceeds one (e.g., a cellularity threshold value of 1.1, 1.2, 1.25, 1.3, 1.4, 1.5), in combination with a P value for cellularity being greater than 1 being sufficiently low, e.g., below a particular P value threshold (e.g., 0.001).

[0596] Prevalent Subclone States. In certain embodiments, a plurality of potential clone types to which a mutation may be assigned includes one or more prevalent subclone states. For example, prevalent subclone states may capture mutations that, while not biologically clonal, are present in a sufficiently high fraction of cancer cells so as to be viable targets for immunotherapy, having sufficient value to include e.g., in a multiepitope composition, for example, if an insufficient number mutations determined to fall within a clonal state (e.g., having a high likelihood of being biologically clonal) are available.

[0597] Accordingly, including prevalent subclone categories of mutations, as opposed to, for example, simply categorizing mutations as clonal or not or clonal or subclonal, recognizes and is informed by the insight that in certain clinical applications, a tiered hierarchy of mutations can be especially valuable. For example, for low mutational burden cancers, insufficient numbers of mutations that are biologically clonal and / or determined to be likely to be biologically clonal may be available for immunotherapy, and, accordingly, mutations that are present at high rates, albeit not necessarily biologically clonal or likely to be biologically clonal, can serve as valuable reservoirs of candidates for inclusion in, for - 113 - 13241940vlAttorney Docket No. 2013237-1502example, a personalized cancer vaccine. Moreover, without wishing to be bound to any particular theory, biophysical phenomena, such as epitope spreading, may increase the breadth of a host immune response to an initially presented epitope (e.g., neoepitope) based on, e.g., a prevalent subclone mutation, thereby rendering prevalent, though not necessarily clonal, mutations and / or combinations of multiple prevalent mutations effective options on which to base immunotherapies.

[0598] In certain embodiments, prevalent subclone mutations may be identified and classified as such based on, for example, cellularity values determined to fall within particular ranges.C.iii.2 Cascading Clonality State Assignment Processes

[0599] Turning to FIGs.4B-4D, in certain embodiments, clone types are assigned to mutations using a cascading procedure. For example, in one example cascading clone type assignment process 450, multiple sets of criteria are evaluated for a given mutation, one after another. Accordingly, at each step, a mutation is evaluated against a particular set of criteria and, either assigned to a corresponding clone type e.g., if it meets the particular set of criteria) or proceeds on to a next set of criteria. For example, as illustrated in FIG.4B, the first set of criteria 452a, associated with the first clone type 454a may be evaluated. If, for a given mutation, the first criteria 452a are met, then the given mutation is assigned to the first clone type 454a. If the first criteria 452a are not satisfied, then the cascading process 450 moves along to evaluate the given mutation according to the second set of criteria 452b associated with the second clone type 454b. If the second set of criteria 452a is met, then the given mutation is assigned to the second clone type 454b, or, if not, then the cascading process 450 moves along to a next set of criteria, associated with a next clone type, and so on, until the A1' set of criteria 452n are evaluated. Based on the final set of criteria 452n, the given mutation may be assigned to the associated clone type 454n. The clone type 454n may be a new class, different from any of the preceding classes, or may be one of the preceding classes, revisited on this last step).

[0600] In certain embodiments, if no criteria of a cascading process 450 are satisfied for a given mutation, the mutation may be assigned to an undetermined class 456, as described in further detail, below.- 114 - 13241940vlAttorney Docket No. 2013237-1502

[0601] In this manner, clone type classification may be performed via a cascading approach, comprising a plurality of steps, each step associated with a particular set of criteria and each particular set of criteria, in turn, associated with a particular corresponding clone type. At each step, the associated criteria is evaluated, and a mutation being assigned may be assigned to the corresponding clone type e.g., if the associated criteria are met) or the process proceeds to a subsequent, next step, for evaluation against a subsequent set of criteria.

[0602] In certain embodiments, each step may be associated with a distinct set of criteria and a distinct clone type. In certain embodiments, as illustrated in FIG. 4C, each step does not need to be associated with a distinct clone type. In certain embodiments, two or more steps may be associated with and used to determine whether to assign a given mutation to a same particular clone type. In this way, two different steps may evaluate two different sets of criteria, such as criteria 462a and 462c illustrated in FIG. 4C, associated with a same particular clone type 464x. In between the two steps, there may be one or more other steps with other sets of criteria e.g., 462b illustrated in FIG. 4C), associated with different clonality classification states 464y.

[0603] Turning to FIG. 4D, in certain embodiments, a given mutation, once assigned to a first clone type 474a, for example, based on first criteria 472a, may, subsequently be reassigned, to a second, different, clone type 474b, for example, based on second criteria 472b.For example, as illustrated in FIG. 4D, while a given mutation may be assigned to a first clone type 474a initially (e.g., early on) in a cascading clone type classification procedure, further on in the procedure, a second set of criteria 472b may be evaluated. In certain embodiments, if second set of criteria 472b is satisfied, a given mutation, having been first assigned to first clone type 474a, may be re-classified as belonging to a second clone type 474b. In certain embodiments, if second set of criteria 472b is not satisfied, a given mutation retains its previously assigned, first, clone type 474a. In certain embodiments, re-assignment at a given step takes into account prior clone type assignments, such that, for example, only mutations assigned previously to a particular, first, clone type are evaluated for reassignment, while others, e.g., having been previously assigned to other (i.e., not the first) clone types, are not.

[0604] Turning again to FIG. 4B, in certain embodiments, mutations that, on the basis of all sets of criteria, are not assigned to one of the associated clone types are, at the end of cascading process 450, assigned to an undetermined state 456, designed, e.g., to capture - 115 - 13241940vlAttorney Docket No. 2013237-1502cases where a confidence in the classification is low due to statistical considerations and / or due to a low confidence in estimation of the various parameters used to perform classification, as described herein.

[0605] Additionally, or alternatively, in certain embodiments, an initial check, for example, based on an initial or first set of criteria 482a may be performed as a first step of cascading process, as illustrated in FIG.4E. Initial criteria 482a may be used to determine, e.g., at an outset, whether a particular mutation should be classified as undetermined 486, prior to testing for assignment to other clone types 484b...484n, for example, based on subsequently evaluated criteria 482b...482n. Accordingly, in certain embodiments, a first set of criteria 482a is associated with and tests whether a given mutation is an inconsistent and / or low confidence case. In certain embodiments, a given mutation may be assigned to an undetermined class 486 based on (z) an initial set of criteria 482a and / or (z'z) failure to be assigned to a particular (e.g., distinct) clone type during cascading process 450, as illustrated in FIGs.4B and 4E. In this manner, inconsistent and / or low confidence cases are either identified at an outset and / or may be captured at a final, e.g., last, stage of a cascading process. In certain embodiments, by virtue of including an undetermined class as an option, each and every candidate mutation may be assigned to a particular one of a set of discrete clone types - for example, either e.g., confidently) a particular clone type that captures a particular likelihood of biological clonality and / or prevalence within a tumor sample or (e.g., if noise, low purity, or other sources of uncertainty do not permit a confident assignment based on biological clonality and / or prevalence) to an undetermined class. In certain embodiments, use of an undetermined class may be particularly useful in the context of neoepitope selection for therapeutic design. In particular, as described herein, mutation...

Claims

1. Attorney Docket No. 2013237-1502CLAIMSWhat is claimed is:

1. A method, the method comprising:(a) receiving, by a processor of a computing device, sequencing data from a tumor sample obtained from a subject;(b) detecting, by the processor, based on the sequencing data, a plurality of candidate mutations, each representing a mutation occurring within a population of tumor cells of the tumor sample;(c) for each particular candidate mutation of at least a portion of the plurality of candidate mutations, determining, by the processor, a corresponding clone type selected from a set of discrete clone types; and(d) selecting, by the processor, a subset of the candidate mutations based at least in part on the determined clone type classifications.

2. The method of claim 1, wherein each member of the set of discrete clone types corresponds to one or more subpopulations of cancer cells, each subpopulation having a characteristic set of mutations and / or wherein the one or more subpopulations of cancer cells correspond to variant subpopulations evolved from an initial, clonal, population arising from a progenitor cell.

3. The method of any one of the preceding claims, wherein each member of the set of discrete clone types represents a prevalence of the corresponding subpopulation(s) of cancer cells.

4. The method of any one of the preceding claims, wherein the set of discrete clone types comprises a first clone type that represents a major subpopulation comprising a characteristic set of mutations present in a majority of cancer cells.- 332 - 13241940vlAttorney Docket No. 2013237-15025. The method of any one of the preceding claims, wherein the set of discrete clone types comprises a second clone type that represents a minor subpopulation comprising a characteristic set of mutations present in a minority of cancer cells.

6. The method of any one of the preceding claims, wherein each discrete clone type is characterized by a nominal cellularity based on cellularities of mutations assigned to that clone type (for example, taking an average cellularity of all mutations assigned to a given clone type).

7. The method of claim 6, comprising determining, for at least a portion of the mutations, the assigned clone type based on a clustering of cellularity values.

8. The method of the preceding claims, comprising determining an evolutionary time point associated with each subpopulation and / or mutation thereof.

9. The method of any one of the preceding claims, comprising associating a particular discrete clone type with a driver gene.

10. The method of any one of the preceding claims, wherein each member of the set of discrete clone types represents a quantized biological cellularity and corresponds to a different estimated cellularity range for the particular candidate mutation.

11. The method of any one of the preceding claims, wherein the set of discrete clone types comprises a plurality of states, each associated with a particular range of biological cellularity and to which mutations are assigned based at least in part on their estimated cellularity values.

12. The method of any one of the preceding claims, wherein the set of discrete clone types comprises more than two discrete clone types.- 333 - 13241940vlAttorney Docket No. 2013237-150213. The method of any one of the preceding claims, wherein the set of discrete clone types comprise one or more of the following:a clone type representing mutations with a biological cellularity of 1;a clone type representing mutations with a biological cellularity of approximately 1; anda plurality of clone types representing multiple populations of biologically subclonal mutations, each associated with a particular range of biological cellularity.

14. The method of claim 13, wherein the plurality of clone types representing multiple populations of biologically subclonal mutations comprises at least a major subclone state and a minor subclone state, the major subclone state representing a first range of biological cellularity and the minor subclone state representing a second range of biological cellularity, the first range greater than the second.

15. The method of claim 14, wherein the major subclone state is subdivided into a plurality of sub-states, each associated with and representing a different range of biological cellularity.

16. The method of claim 14 or 15, wherein the minor subclone state represents mutations present in a minority of tumor cells.

17. The method of any one of claims 13 to 16, wherein the plurality of clone types representing multiple populations of biologically subclonal mutations comprises one or more prevalent subclone states representing mutations present in a majority of tumor cells.

18. The method of any one of the preceding claims, wherein each member of the set of discrete clone types represents a distinct clone type indicative of (i) a likelihood that the particular candidate mutation is a biologically clonal mutation and / or (ii) its prevalence- 334 - 13241940vlAttorney Docket No. 2013237-1502within the tumor sample within the population tumor cells, thereby determining clone type classifications for the portion of the candidate mutations.

19. The method of any one of the preceding claims, comprising receiving, by the processor, sequencing data obtained from a normal sample, and using the normal sample sequencing data together with the tumor sample sequencing data to perform one or more of steps (b), (c), and (d).

20. The method of any one of the preceding claims, comprising, following step (d), producing a pharmaceutical composition comprising a polyribonucleotide encoding one or more neoepitopes, wherein at least a portion of the one or more neoepitopes are encoded by nucleotide sequence comprising a member of the selected subset of candidate mutations.

21. The method of any one of the preceding claims, comprising, using the selected subset of candidate mutations to produce an enriched population of T-cells for the subject, said enriched population of T-cells capable of specifically binding to a plurality of complexes, each complex of the plurality comprising at least a portion of a neoepitope encoded by a member of the selected subset of candidate mutations.

22. The method of claim 21, wherein each complex of the plurality further comprises a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells (APCs).

23. The method of any one of the preceding claims, comprising, using the selected subset of candidate mutations to produce an enriched population of tumor infiltrated lymphocytes (TILs) for the subject, said enriched population of TILs capable of specifically binding to a plurality of complexes, each complex of the plurality comprising at least a portion of a neoepitope encoded by a member of the selected subset of candidate mutations.- 335 - 13241940vlAttorney Docket No. 2013237-150224. The method of claim 23, wherein each complex of the plurality further comprises a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells (APCs).

25. The method of any one of the preceding claims, comprising, using the selected subset of candidate mutations to produce a T-cell receptor for the subject, said T-cell receptor capable of specifically binding to a complex comprising at least a portion of a neoepitope encoded by a member of the selected subset of candidate mutations.

26. The method of claim 25, wherein the complex further comprises a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells (APCs).

27. The method of any one of the preceding claims, comprising, using the selected subset of candidate mutations to produce a chimeric antigen receptor (CAR) for the subject, said CAR capable of specifically binding to a complex comprising at least a portion of a neoepitope encoded by a member of the selected subset of candidate mutations.

28. The method of claim 27, wherein the complex further comprises a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells (APCs).

29. The method of any one of the preceding claims, comprising:determining, by the processor, based on the sequencing data, an estimated tumor sample purity and / or purity bound; andprior to proceeding to step (b) and / or step (c), determining, by the processor, the estimated tumor sample purity and / or the purity bound to be greater than or equal to a purity threshold.- 336 - 13241940vlAttorney Docket No. 2013237-150230. The method of any one of the preceding claims, wherein the tumor sample comprises tumor cells and normal cells.

31. The method of any one of the preceding claims, comprising performing, by the processor, tumor deconvolution.

32. The method of claim 31, comprising evaluating one or more quality control metrics with respect to the performed tumor deconvolution and accepting result(s) of the tumor deconvolution based thereon.

33. The method of claim 31 or 32, comprising performing the tumor deconvolution using a plurality of tumor models.

34. The method of claim 33, comprising, for each particular candidate mutation, determining, by the processor, a plurality of prospective clone types, each prospective clone type determined using a distinct one of the plurality of tumor models and using the plurality of prospective clone types to determine the corresponding clone type for the particular candidate mutation (e.g., selecting the corresponding clone type as a most common one of the prospective clone types, e.g., if a minimum fraction (up to all) of the prospective clone types are the same, otherwise selecting the corresponding clone type as undetermined).35 The method of claim 33 or 34, comprising,for each particular candidate mutation determining, by the processor, a plurality of prospective clone types, each prospective clone type determined using a distinct one of the plurality of tumor models,identifying candidate mutations for which all of the corresponding prospective clone types are a same clone type and assigning each of those mutations a scale invariant clone type-- 337 - 13241940v1Attorney Docket No. 2013237-150236. The method of any one of the preceding claims, comprising,prior to step (c), identifying, by the processor, a low confidence subset of the candidate mutations, said low confidence subset comprising candidate mutations determined to be likely false positives and / or rare subclones; andexcluding the low confidence subset of the candidate mutations from (i) the portion of the candidate mutations for which clone type classifications are determined at step (c) and / or (ii) the subset selected for inclusion in the construct at step (d).

37. The method of claim 36, wherein identifying the low confidence subset comprises:determining, for each candidate mutation, a mutation confidence score that measures that measures a likelihood that a site associated with a given mutation encodes that given mutation in comparison with a likelihood that it does not (encode that mutation); and determining whether a particular candidate mutation is a likely false positive and / or rare subclone based on the mutation confidence score.

38. The method of any one of the preceding claims, wherein the set of discrete clone types comprises one or more clonal states indicative of a high likelihood that a candidate mutation is a biologically clonal mutation and / or highly prevalent.

39. The method of any one of the preceding claims, wherein the set of discrete clone types comprises one or more clonal states indicative of a high likelihood that a candidate mutation is a biologically clonal mutation.

40. The method of claim 38 or 39, wherein the set of discrete clone types comprises at least two clonal states [e.g., a putative clonal state representing mutations with a biological cellularity of 1 and a nearly clonal state, representing mutations with a biological cellularity of nearly 1, each of the at least two clonal states associated with a distinct set of classification criteria and representing differing likelihoods that a candidate mutation is a biologically clonal mutation.- 338 - 13241940vlAttorney Docket No. 2013237-150241. The method of any one of claims 38 to 40, wherein the method comprises, for at least one particular candidate mutation:determining, by the processor, a cellularity confidence interval (CI) lower bound for the particular mutation; andclassifying, by the processor, the particular candidate mutation as belonging to a particular one of the one or more clonal states based at least in part on the cellularity CI lower bound.

42. The method of any one of claims 38 to 41, wherein the method comprises, for at least one particular candidate mutation:determining, by the processor, for the particular candidate mutation, a first P- value representing a probability of observing a cellularity of the particular candidate mutation below one; andclassifying, by the processor, the particular candidate mutation as belonging to a particular one of the one or more clonal states based at least in part on the first P- value.

43. The method of claim 42, comprising:determining, by the processor, an estimated cellularity for the particular candidate mutation; andclassifying, by the processor, the particular candidate mutation as belonging to the particular one of the one or more clonal states based on the estimated cellularity.

44. The method of any one of claims 38 to 43, wherein the method comprises, for at least one particular candidate mutation:determining, by the processor, for the particular candidate mutation, a second P-value representing a probability of observing a cellularity of the particular candidate mutation at or above one; andclassifying, by the processor, the particular candidate mutation as belonging to a particular one of the one or more clonal states based at least in part on the second P-value.- 339 - 13241940vlAttorney Docket No. 2013237-150245. The method of claim 44, comprising:determining, by the processor, an estimated cellularity for the particular candidate mutation; andclassifying, by the processor, the particular candidate mutation as belonging to the particular one of the one or more clonal states based on the estimated cellularity.

46. The method of any one of claims 38 to 45, wherein the one or more clonal states comprise a plurality of clonal states and wherein the method comprises, for at least one particular candidate mutation:determining, by the processor, one or both of (i) and (ii) as follows:(i) a first P-value representing a probability of observing a cellularity of the particular candidate mutation below one; and(ii) a second P-value representing a probability of observing a cellularity of the particular candidate mutation at or above one; andclassifying, by the processor, the particular candidate mutation as belonging to a first clonal state of the plurality of clonal states if the first P-value and / or the second P-value is / are above a first P-value threshold, or classifying, by the processor, the first candidate mutation as belonging to a second clonal state of the plurality of clonal states if the first P-value and / or second P-value is / are below the first P-value threshold.

47. The method of any one of claims 38 to 46, wherein the method comprises, for at least one particular candidate mutation:determining, by the processor, an estimated cellularity (ρ) for the particular candidate mutation; andclassifying, by the processor, the particular candidate mutation as belonging to a particular one of the one or more clonal states based at least in part on the estimated cellularity.- 340 - 13241940vlAttorney Docket No. 2013237-150248. The method of any one of the preceding claims, wherein the set of discrete clone types comprises a minor subclone state indicative of a high likelihood that a candidate mutation is (i) a biologically subclonal mutation and (ii) present in tumor cells as a fraction within a range that is lower than at least one other subclonal state.

49. The method of claim 48, wherein the minor subclone state represents candidate mutations present in a minority of tumor cells.

50. The method of claim 48 or 49, wherein the method comprises, for at least one particular candidate mutation:determining, by the processor, an estimated cellularity (ρ) for the particular candidate mutation; andclassifying, by the processor, the particular candidate mutation as belonging to the minor subclone state based at least in part on the estimated cellularity.

51. The method of any one of claims 48 to 50, wherein the method comprises, for at least one particular candidate mutation:determining, by the processor, for the particular candidate mutation, a first P- value representing a probability of observing a cellularity of the particular candidate mutation below one; andclassifying, by the processor, the particular candidate mutation as belonging to the minor subclone state based at least in part on the first P-value.

52. The method of claim51 wherein classifying the particular candidate mutation comprises determining the first P-value to be below a particular P-value threshold(s).

53. The method of either one of any one of claims 39 to 43, wherein the method comprises, for at least one particular candidate mutation:- 341 - 13241940vlAttorney Docket No. 2013237-1502determining, by the processor, a cellularity confidence interval (CI) upper bound for the particular candidate mutation; andclassifying, by the processor, the particular candidate mutation as belonging to the minor subclone state based at least in part on the cellularity CI upper bound.

54. The method of any one of the preceding claims, wherein the set of discrete clone types comprises an undetermined state indicative of mutations for which clone type assignments are determined and / or predicted to be unreliable.

55. The method claim 45, wherein the method comprises, for at least one particular candidate mutation:determining, by the processor, an estimated copy number (CN) for the particular candidate mutation; andclassifying, by the processor, the particular candidate mutation as belonging to the undetermined state based at least in part on the estimated copy number being about zero (0).

56. The method of any one of claims 45 to 46, wherein the method comprises, for at least one particular candidate mutation:determining, by the processor, an estimated cellularity (ρ) for the particular candidate mutation; andclassifying, by the processor, the particular candidate mutation as belonging to the undetermined state based at least in part on the estimated cellularity.

57. The method of any one of claims 54 to 56, wherein the method comprises, for at least one particular candidate mutation:determining, by the processor, for the particular candidate mutation, a P- value representing a probability of observing a cellularity above one for the particular candidate mutation; and- 342 - 13241940vlAttorney Docket No. 2013237-1502classifying, by the processor, the particular candidate mutation as belonging to the undetermined state based at least in part on the P-value falling below corresponding P-value threshold.

58. The method of any one of the preceding claims, wherein the set of discrete clone types comprises one or more prevalent subclone states.

59. The method of claim 58, wherein the set of discrete clone types comprises a plurality of prevalent subclone states, each representing mutations present in a majority of tumor cells (e.g., highly prevalent in the population tumor cells, but not necessarily present in all tumor cells, said plurality comprising at least a first prevalent subclone state and a second prevalent subclone state, the first prevalent state representing mutations more common that the second prevalent state].

60. The method of claim 58 or 59, wherein the method comprises, for at least one particular candidate mutation:determining, by the processor, an estimated cellularity (ρ) for the particular candidate mutation; andclassifying, by the processor, the particular candidate mutation as belonging to a particular one of the one or more prevalent subclone states based at least in part on the estimated cellularity.

61. The method of claims 58 to 60, wherein the method comprises, for at least one particular candidate mutation:determining, by the processor, a cellularity confidence interval (CI) lower bound for the particular mutation; andclassifying, by the processor, the particular candidate mutation as belonging to a particular one of the one or more prevalent subclone states based at least in part on the cellularity CI lower bound.- 343 - 13241940vlAttorney Docket No. 2013237-150262. The method of any one of the preceding claims, wherein the set of discrete clone types comprises two or more states.

63. The method of any one of the preceding claims, wherein the set of discrete clone types comprises five or more states.

64. The method of any one of the preceding claims, wherein the set of discrete clone types comprises six states, said six states comprising two clonal states, two prevalent subclone state, a minor subclone state, and an undetermined state.

65. The method of any one of the preceding claims, wherein the set of discrete clone types comprises five states, said five states comprising a clonal state, two prevalent subclone state, a minor subclone state, and an undetermined state.

66. The method of any one of the preceding claims, wherein the set of discrete clone types comprises four states, said four states comprising a clonal state, a prevalent subclone state, a minor subclone state, and an undetermined state.

67. The method of any one of the preceding claims, wherein the set of discrete clone types comprises three states, said three states comprising a clonal state, a subclonal state, and an undetermined state.

68. The method of any one of the preceding claims, wherein the set of discrete clone types comprises at least a clonal state and / or a subclonal state.

69. The method of any one of the preceding claims, wherein step (c) comprises, for at least one particular candidate mutation:determining, via a statistical hypothesis test, a likelihood that the particular candidate mutation is biologically clonal; and- 344 - 13241940vlAttorney Docket No. 2013237-1502determining the corresponding clone type for the particular candidate mutation based at least in part on the determined likelihood.

70. The method of any one of the preceding claims, wherein step (c) comprises, for at least one particular candidate mutation:determining, using the sequencing data, an estimated cellularity for the particular candidate mutation; anddetermining the corresponding clone type for the particular candidate cell mutation based at least in part on the cellularity of the particular candidate mutation.

71. The method of any one of the preceding claims, wherein step (c) comprises, for at least one particular candidate mutation:determining, using the sequencing data, a likelihood of observing a cellularity of the particular candidate mutation below one; anddetermining the corresponding clone type for the particular candidate mutation based at least in part on the determined likelihood.

72. The method of any one of the preceding claims, wherein step (c) comprises, for at least one particular candidate mutation:determining, using the sequencing data, a likelihood of observing a cellularity of the particular candidate mutation above one; anddetermining the corresponding clone type for the particular candidate mutation based at least in part on the determined likelihood.

73. The method of any one of the preceding claims, wherein step (c) comprises, for at least one particular candidate mutation:determining a cellularity confidence interval (CI) lower and / or upper bound for the particular candidate mutation; and- 345 - 13241940vlAttorney Docket No. 2013237-1502determining the corresponding clone type for the particular candidate mutation based at least in part on the determined cellularity CI lower and / or upper bound(s).

74. The method of any one of the preceding claims, wherein step (c) comprises, for at least one particular candidate mutation:evaluating, by the processor, a first set of criteria with respect to the particular candidate mutation and determining the first set of criteria to be satisfied;classifying, by the processor, the particular candidate mutation as belonging to a first clone type of the set of discrete clone types based on the determined satisfaction of the first set of criteria;evaluating, by the processor, a second set of criteria with respect to the particular candidate mutation and determining the second set of criteria to be satisfied;classifying, by the processor, the particular candidate mutation as belonging to a second clone type of the set of discrete clone types, different from the first, based on the determined satisfaction of the second set of criteria, thereby re-assigning the clone type classification for the particular candidate mutation.

75. The method of claim 76, wherein evaluating the second set of criteria comprises determining a purity of the tumor sample to be below a particular purity threshold.

76. The method of any one of the preceding claims, wherein step (c) comprises determining a corresponding clone type for each of the plurality of candidate mutations.

77. The method of any one of the preceding claims, wherein step (c) comprises, for at least one particular candidate mutation of the plurality of candidate mutations:evaluating a plurality of sets of criteria with respect to the particular candidate mutation;determining none of the sets of criteria to be satisfied for the particular candidate mutation; and- 346 - 13241940vlAttorney Docket No. 2013237-1502classifying the particular candidate mutation as belonging to the undetermined state.

78. The method of any one of the preceding claims, wherein step (c) comprises, for at least one particular candidate mutation, determining values of one or more mutation characterization metrics for the particular candidate mutation.

79. The method of claim 78, wherein the one or more mutation characterization metrics comprises a cellularity confidence interval (CI) lower bound.

80. The method of claim 78 or 79, wherein the one or more mutation characterization metrics comprises a cellularity CI upper bound.

81. The method of any one of claims 78 to 80, wherein the one or more mutation characterization metrics comprises an estimated cellularity (p).

82. The method of any one of claims 78 to 81, wherein the one or more mutation characterization metrics comprises a P-value representing a probability of observing a cellularity of the particular candidate mutation below one.

83. The method of any one of claims 78 to 82 wherein the one or more mutation characterization metrics comprises a P-value representing a probability of observing a cellularity of the particular candidate mutation at or above one.

84. The method of any one of claims 78 to 83, wherein step (c) comprises comparing one or more of the mutation characterization metrics to one or more thresholds, at least a portion of which are purity dependent thresholds, having two or more values, each associated with a different range of purities of the tumor sample.- 347 - 13241940vlAttorney Docket No. 2013237-150285. The method of any one of the preceding claims, wherein the set of discrete clonality states comprises a minor subclone state and the method comprises:at step (c), determining the minor subclone state as the corresponding clone type for one or more of the candidate mutations; andat step (d), excluding the candidate mutations classified as belonging to the minor subclone state from the selected subset.

86. The method of any one of the preceding claims, comprising classifying one or more of the candidate mutations as clonal and, at step (d), selecting at least a portion of the candidate mutations classified as clonal for inclusion in the selected subset.

87. The method of any one of the preceding claims, wherein the determined clone type classifications include a plurality of distinct clone types and wherein step (d) comprises prioritizing mutations for inclusion in the selected subset according to their determined clone types.

88. The method of claim 87, wherein the method comprises ranking the plurality of discrete clone types such that mutations assigned to higher ranked clone types are prioritized for inclusion in the selected subset over mutations assigned to lower ranked clone types.

89. The method of claim 88, wherein the plurality of discrete clone types comprises an undetermined state and a minor subclone state, and wherein the method comprises ranking the undetermined state above the minor subclone state (e.g., such that mutations assigned to the undetermined state are prioritized over mutations assigned to the minor subclone state.

90. The method of claim 87 or 88, wherein the plurality of discrete clone types comprises:(A) one or more clonal states;(B) one or more major subclone states; and(C) one or more minor subclone state(s).- 348 - 13241940vlAttorney Docket No. 2013237-150291. The method of claim 90, wherein the method comprises ranking the one or more clonal states above the one or more prevalent subclone states and ranking the one or more major subclone states above the minor subclone state.

92. The method of claim 90 or 91, wherein the one or more clonal states comprise a putative clonal state and a nearly clonal state and the method comprises ranking the putative clonal state above the nearly clonal state.

93. The method of any one of claims 90 to 92, wherein the plurality of discrete clone types comprises two or more prevalent subclone states, wherein said two or more prevalent subclone states comprising a first prevalent state and a second prevalent state, the first prevalent state representing mutations more common that the second prevalent state, and wherein the method comprises ranking the first prevalent state above the second prevalent state.94 The method of any one of claims 88-93, wherein the ranking is determined at least in part based on an identification of a driver phenotype associated with one or more of the mutations.

95. The method of any one of claims 88-94, comprising identifying at least a portion of the mutations as driver mutations and ranking the plurality of discrete clone types such that mutations that are (i) assigned to a particular clone type and (ii) are identified as driver mutations are ranked above mutations that are (i) also assigned to the particular clone type, but (ii) not identified as driver mutations.

96. The method of any one of claims 88 to 95, wherein the ranking is determined based at least in part on a purity of the sample.- 349 - 13241940vlAttorney Docket No. 2013237-150297. The method of any one of claims 88 to 96, wherein the ranking is determined based at least in part on a tumor mutational burden (TMB).

98. The method of any one of claims 88 to 97, wherein the ranking is determined based at least in part on a number of non-synonymous (NS) mutations.

99. The method of any one of claims 88 to 98, wherein the ranking is determined based at least in part on one or more parameters that impact one or more members selected from the group consisting of expression, presentation, and immunogenicity.

100. The method of any one of claims 88 to 99, wherein the ranking is determined based at least in part on a cancer type.

101. The method of any one of claims 88 to 100 wherein the ranking is determined based at least in part on a number of mutations assigned to the clonal state.

102. The method of any one of the preceding claims, comprising, at step (d), selecting the subset based at least in part on the classified clonality states in combination with an immunogenicity prediction score.

103. The method of any one of the preceding claims, comprising, at step (d), selecting the subset based at least in part on the clone type classifications in combination with an expression score.

104. The method of any one of the preceding claims, comprising, at step (d), selecting the subset based at least in part on the clone type classifications in combination with one or more members selected from the group consisting of T cell receptor (TCR) recognition, copy number, zygosity, fractional zygosity, and essential gene.- 350 - 13241940vlAttorney Docket No. 2013237-1502105. The method of any one of the preceding claims, wherein step (c) comprises, for each particular candidate mutation, evaluating one or more sets of criteria in a stepwise fashion, wherein each set of criteria is associated with, and used to assign a given mutation to, a particular member of the set of discrete clone types.

106. The method of claim 94, comprising for each particular mutation, evaluating at least a portion of the one or more sets of criteria by, at each particular step of one or more steps: evaluating a corresponding current set of the one or more sets of criteria with respect to the given mutation, the current set associated with, and used to assign a given mutation to, a corresponding one of the one or more discrete clone types and, either:(i) determining the criteria of the current set to be satisfied, and, based on said determination of the criteria of the current set being satisfied, assigning the particular mutation to the corresponding clone type, or(ii) determining the criteria of the current set not to be satisfied and, based on said determination of the criteria of the current set not being satisfied, proceeding to subsequent step.

107. The method of claim 106, wherein the set of discrete clone types comprises an undetermined classification state and wherein a first step of the one or more steps is associated with the undetermined state and the method comprises, at the first step, determining a corresponding first criteria to be satisfied and, based on the determined satisfaction of the first criteria, assigning the particular mutation to the undetermined state.

108. The method of claim 106 or 107, wherein the set of discrete clone types comprises an undetermined state and wherein, for at least one particular candidate mutation, the method comprises determining, for each of the one or more steps, the corresponding criteria not to be satisfied for the particular candidate mutation and assigning the particular candidate mutation to the undetermined state.- 351 - 13241940vlAttorney Docket No. 2013237-1502109. The method of any one of claims 106 to 108 wherein two or more of the one or more steps are associated with a same clone type but comprise different sets of criteria.

110. The method of any one of claims 106 to 109, comprising for at least one particular candidate mutation:evaluating, by the processor, a first set of criteria with respect to the particular candidate mutation and determining the first set of criteria to be satisfied;classifying, by the processor, the particular candidate mutation as belonging to a first clone type of the set of discrete clone types based on the determined satisfaction of the first set of criteria;evaluating, by the processor, a second set of criteria with respect to the particular candidate mutation and determining the second set of criteria to be satisfied; and classifying, by the processor, the particular candidate mutation as belonging to a second clone type of the set of discrete clone types, different from the first, based on the determined satisfaction of the second set of criteria, thereby re-assigning the clone type classification for the particular candidate mutation.

111. The method of claim 110 wherein evaluating the second set of criteria comprises determining a purity of the tumor sample to be below a particular purity threshold.

112. The method of any one of the preceding claims, comprising, causing, by the processor, display of a graphical clonality state spectrum view in which mutations assigned a particular clone type are grouped according their absolute copy numbers and zygosities and a graphical representation of VAFs and / or cellularity values of the mutations assigned the particular clone type is displayed.

113. The method of claim 112, comprising, for each particular clone type of one or more clone types:- 352 - 13241940vlAttorney Docket No. 2013237-1502grouping corresponding mutations into one or more subsets according to their absolute copy numbers and zygosities, such that each of the one or more subsets comprises mutations having a same absolute copy number and zygosity; andcausing display of, within the graphical clonality state spectrum view, a graphical representation of VAFS and / or cellularity values for each subset of mutations.

114. The method of any one of the preceding claims, wherein the sequencing data comprises WGS and WES data, and wherein step (b) comprises using the WES to detect the plurality of candidate mutations and wherein step (c) comprises determining absolute copy numbers for segments comprising the plurality of candidate mutations using the WGS data and using the determined absolute copy numbers to determine the corresponding clone type for each particular candidate mutation.115 The method of any one of the preceding claims, wherein the sequencing data comprises RNAseq data and step (b) comprises filtering false positives using the RNAseq data.

116. A method for detecting and filtering out rare sub-clones within sets of tumor cell mutations, the method comprising:(a) receiving, by a processor of a computing device, sequencing data from a tumor sample obtained from a subject;(b) detecting, by the processor, based on the sequencing data, an initial set of potential tumor cell mutations, the initial set comprising a plurality of detected mutations, each representing a potential mutation present in one or more tumor cells of the tumor sample;(c) for each particular detected mutation within at least a portion of the initial set of potential tumor cell mutations, determining, by the processor, a mutation confidence score that measures a likelihood that a site associated with the particular detected mutation encodes the particular detected mutation in comparison a likelihood that the site does not encode the particular detected mutation, thereby determining a plurality of mutation confidence scores;- 353 - 13241940vlAttorney Docket No. 2013237-1502(d) filtering, by the processor, the initial set of potential tumor cell mutations to exclude one or more detected mutations identified as excessively rare or unlikely based at least in part on the plurality of mutation confidence scores, to generate a high confidence set of tumor cell mutations; and(e) selecting, by the processor, from the high confidence set, one or more target mutations for inclusion in and / or targeting by a construct.

117. The method of claim 116, comprising:(f) producing a pharmaceutical composition comprising a polyribonucleotide encoding one or more neoantigen epitopes, wherein at least a portion of the one or more neoepitopes are encoded by nucleotide sequence comprising one or more of the one or more target mutations.

118. The method of any one of claims 116 to 117, comprising, using the one or more target mutations to produce an enriched population of T-cells for the subject, said enriched population of T-cells capable of specifically binding to a plurality of complexes, each complex of the plurality comprising at least a portion of a neoepitope encoded by at least one of the one or more target mutations.

119. The method of claim 118, wherein each complex of the plurality further comprises a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells (APCs).

120. The method of any one of claims 116 to 119, comprising, using the selected subset of candidate mutations to produce an enriched population of tumor infiltrated lymphocytes (TILs) for the subject, said enriched population of TILs capable of specifically binding to a plurality of complexes, each complex of the plurality comprising at least a portion of a neoepitope encoded by a member of the selected subset of candidate mutations.- 354 - 13241940vlAttorney Docket No. 2013237-1502121. The method of claim 120, wherein each complex of the plurality further comprises a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells (APCs).

122. The method of any one of claims 116 to 121, comprising, using the one or more target mutations to produce a T-cell receptor for the subject, said T-cell receptor capable of specifically binding to a complex comprising at least a portion of a neoepitope encoded by at least one of the one or more target mutations.

123. The method of claim 122, wherein the complex further comprises a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells (APCs).

124. The method of any one of claims 116 to 123, comprising, using the one or more target mutations to produce a chimeric antigen receptor (CAR) for the subject, said CAR capable of specifically binding to a complex comprising at least a portion of a neoepitope encoded by at least one of the one or more target mutations.

125. The method of claim 124, wherein the complex further comprises a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells (APCs).

126. The method of any one of claims 116 to 125, wherein the mutation confidence score comprises determining the likelihoods over ranges of one or more of: sample purities, absolute copy number scale(s), e.g., optionally, limiting variant allele frequencies (VAFs).

127. The method of any one of claims 116 to 126, comprising determining an estimated purity and / or an estimated absolute copy number of the site encoding the detected mutation, and using the estimated purity and / or estimated absolute copy number to determine the- 355 - 13241940vlAttorney Docket No. 2013237-1502mutation confidence score, thereby determining, and filtering the detected mutations based on, a refined mutation confidence score (RMCS).

128. The method of claim 127, wherein the estimated purity is an estimated bound on a purity of the tumor sample.

129. The method of claim 128, comprising determining the estimated purity using a plurality of tumor models.

130. The method of claim 129, comprising:evaluating one or more quality control metrics with respect to the plurality of tumor models;determining, based on the one or more quality control metrics, a failure of tumor deconvolution; andresponsive to the tumor deconvolution failure, determining, based on the plurality of tumor models, the estimated bound on the purity of the tumor sample.

131. The method of any one of claims 116 to 130, comprising determining an estimated absolute copy number of an alternate allele for each detected mutation and using the estimated absolute copy number of an alternate allele to determine the mutation confidence score, thereby determining, and filtering the detected mutations based on, a refined mutation confidence score (RMCS).

132. The method of any one of claims 116 to 131, wherein the filtering comprises comparing the mutation confidence score to a purity dependent threshold.

133. A method, the method comprising:(a) receiving, by a processor of a computing device, sequencing data from a tumor sample obtained from a subject;- 356 - 13241940vlAttorney Docket No. 2013237-1502(b) detecting, by the processor, based on the sequencing data, a plurality of candidate mutations, each representing a mutation occurring within a population of tumor cells of the tumor sample; and(c) assigning, by the processor, the plurality of candidate mutations to one or more classes selected from a set of discrete clone types, where the set of discrete clone types includes an undetermined state indicative of mutations for which clone type assignments are determined and / or predicted to be unreliable, and wherein a first subset of the plurality of candidate mutations to the undetermined state.

134. The method of claim 133, comprising:assigning, by the processor, a second subset of the plurality of candidate mutations to (i) a clonal state indicative of a given mutation having a high likelihood of being biologically clonal and / or (ii) a prevalent classification state indicative of a given mutation being highly prevalent in tumor cells, though not necessarily clonal;ranking and / or scoring the plurality of candidate mutations based at least in part on their assigned clone types, wherein candidate mutations of the second subset are ranked and / or scored higher than those of the first subset; andselecting a final subset of the candidate mutations for inclusion in a construct based at least in part on the ranking and / or scoring of the candidate mutations.

135. The method of claim 133 or 134, comprising:assigning, by the processor, a third subset of the plurality of candidate mutations to a minor subclone state indicative of a given mutation having (i) a high likelihood of being biologically subclonal and / or (ii) a low level of prevalence in tumor cells;ranking and / or scoring the plurality of candidate mutations based at least in part on their assigned clone types, wherein candidate mutations of the first subset are ranked and / or scored higher than those of the third subset; andselecting a final subset of the candidate mutations for inclusion in a construct based at least in part on the ranking and / or scoring of the candidate mutations.- 357 - 13241940vlAttorney Docket No. 2013237-1502136. A system comprising a processor of a computing device and memory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of the preceding claims.

137. A method of producing an immunotherapy construct for a subject, the method comprising:detecting a plurality of candidate mutations in tumor cells of from the subject and selecting, using a method or system of any one of claims 1-136, a subset of the plurality of candidate mutations for inclusion in the immunotherapy construct; andsynthesizing, as the immunotherapy construct, a polyribonucleotide encoding a plurality of neoepitopes, wherein each encoded neoepitope corresponds to a candidate mutation of the selected subset.138 The method of claim 137, comprising screening candidate mutations of the selected subset against T-cells and / or T-cell receptors (TCRs) derived from the subject.

139. A method comprising:determining a subset of T-cells within a population of T-cells that are capable of specifically binding a plurality of complexes,wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer-specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, andwherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer-specific mutations within the selected subset of candidate mutations determined by the method or system of any one of claims 1-136; andenriching for the subset of T-cells that are capable of specifically binding the plurality of complexes.- 358 - 13241940vlAttorney Docket No. 2013237-1502140. A method comprising:administering to a subject a population of T-cells, wherein the population of T cells are capable of specifically binding a plurality of complexes,wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer-specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells,wherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer-specific mutations within the selected subset of candidate mutations determined by the method or system of any one of claims 1-136, andwherein the genome of at least some of the subject’s cells comprises a subset of the cancer- specific mutations.

141. A method comprising:determining a subset of TILs within a population of TILs that are capable of specifically binding a plurality of complexes,wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer-specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, andwherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer-specific mutations within the selected subset of candidate mutations determined by the method or system of any one of claims 1-136; andenriching for the subset of TILs that are capable of specifically binding the plurality of complexes.

142. A method comprising:administering to a subject a population of TILs, wherein the population of TILs are capable of specifically binding a plurality of complexes,- 359 - 13241940vlAttorney Docket No. 2013237-1502wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer-specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells,wherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer-specific mutations within the selected subset of candidate mutations determined by the method or system of any one of claims 1-136, andwherein the genome of at least some of the subject’s cells comprises a subset of the cancer- specific mutations.

143. The method of any one of claims 124 to 126, comprising obtaining a tumor sample from the subject and using the tumor sample to detect the plurality of candidate mutations.

144. The method of claim 127, comprising obtaining a normal sample From the subject and using the normal sample to detect the plurality of cancer mutations.

145. The method of claim 127 or 128, comprising sequencing the tumor and / or normal sample.

146. A pharmaceutical composition comprising a polyribonucleotide encoding a plurality of neoepitopes, wherein at least a portion of the plurality of neoepitopes are individualized neoepitopes corresponding to a somatic mutation from a population of candidate mutations detected in a tumor sample of a subject and selected for inclusion in the pharmaceutical composition using the method or system of any one of the preceding claims.

147. An individualized pharmaceutical composition comprising a polyribonucleotide encoding a plurality of neoepitopes for a particular subject, wherein at least a portion of the neoepitopes are individualized neoepitopes corresponding to a somatic mutation from a population of candidate mutations detected in a tumor sample from the particular subject and- 360 - 13241940vlAttorney Docket No. 2013237-1502selected for inclusion in the pharmaceutical composition using the method of any one of the claims 1-135 or the system of claim 136.

148. A population of T-cells,wherein the population of T-cells is capable of specifically binding to a plurality of complexes,wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer-specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, andwherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer-specific mutations within the selected subset of candidate mutations determined by the method or system of any one of claims 1-136.

149. A T-cell capable of specifically binding to a complex, wherein the complex comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells,wherein the one or more cancer- specific mutations are within the selected subset of candidate mutations determined by the method or system of any one of claims 1-136.

150. A population of T-cells,wherein the population of T-cells is capable of specifically binding to a plurality of complexes,wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer-specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, and- 361 - 13241940vlAttorney Docket No. 2013237-1502wherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer-specific mutations within the selected subset of candidate mutations determined by the method or system of any one of claims 1-136.

151. A T-cell capable of specifically binding to a complex, wherein the complex comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells,wherein the one or more cancer- specific mutations are within the selected subset of candidate mutations determined by the method or system of any one of claims 1-136.

152. A T-cell receptor capable of specifically binding to a complex, wherein the complex comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer-specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells,wherein the one or more cancer- specific mutations are within the selected subset of candidate mutations determined by the method or system of any one of claims 1-136.

153. A chimeric antigen receptor (CAR) capable of specifically binding to a complex, wherein the complex comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer-specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells,wherein the one or more cancer- specific mutations are within the selected subset of candidate mutations determined by the method or system of any one of claims 1-136.

154. The pharmaceutical composition, the population of T-cells, the T-cell, the TIL population, the T-cell receptor, or the CAR of any one of claims 146-153, wherein at least a portion of the individualized neoepitopes correspond to biologically clonal mutations.- 362 - 13241940vlAttorney Docket No. 2013237-1502155. The pharmaceutical composition, the population of T-cells, the T-cell, the TIL population, the T-cell receptor, or the CAR of any one of claims 146-154, wherein at least a portion of the individualized neoepitopes correspond to mutations determined to have a high likelihood of being biologically clonal and / or highly prevalent.

156. The pharmaceutical composition, the population of T-cells, the T-cell, the TIL population, the T-cell receptor, or the CAR of any one of claims 146-155, wherein the individualized neoepitopes do not correspond to mutations that are biologically subclonal and present in a minority of tumor cells.

157. The pharmaceutical composition, the population of T-cells, the T-cell, the TIL population, the T-cell receptor, or the CAR of any one of claims 146-156, wherein none of the individualized neoepitopes correspond to mutations that are biologically subclonal and present in a minority of tumor cells.

158. The pharmaceutical composition, the population of T-cells, the T-cell, the TIL population, the T-cell receptor, or the CAR of any one of claims 146-157, wherein at least a portion of the individualized neoepitopes correspond to mutations having been classified as belonging to an undetermined state indicative of mutations for which clone type assignments are determined and / or predicted to be unreliable.

159. The pharmaceutical composition, the population of T-cells, the T-cell, the TIL population, the T-cell receptor, or the CAR of any one of claims 146-158, wherein at least a portion of the individualized neoepitopes correspond mutations that are biologically subclonal, but present in a majority of tumor cells of the subject.

160. The pharmaceutical composition, population of T-cells, T-cell, TIL population, T-cell receptor, or CAR of claim 159, wherein all of the portion of the individualized neoepitopes correspond to mutations that are present in a majority of tumor cells of the subject.- 363 - 13241940vlAttorney Docket No. 2013237-1502161. The pharmaceutical composition, population of T-cells, T-cell, TIL population, T-cell receptor, or CAR of claim 159 or 160, wherein:the population of candidate mutations comprises at least two sub-populations of mutations that are biologically subclonal, but present in a majority of tumor cells,a first of the at least two sub-populations comprises mutations having cellularity values within a first range and a second of the at least two- subpopulations comprises mutations having cellularity values within a second range, said first range above the second range, and the individualized neoepitopes comprise more neoepitopes corresponding to mutations from the first sub-population than the second.

162. The pharmaceutical composition, the population of T-cells, the T-cell, the TIL population, the T-cell receptor, or the CAR of any one of claims 159 to 161, wherein the population of candidate mutations comprises at least four sub-populations, including:(A) a first sub-population comprising mutations that are biologically clonal and / or present in all or nearly all tumor cells;(B) a second sub-population comprising mutations that are biologically sub-clonal, but present in a majority of tumor cells at a rate within a first range;(C) a third sub-population comprising mutations that are biologically sub-clonal, but present in a majority of tumor cells at a rate within a second range, the second range below the first range; and(D) a fourth sub-population, comprising mutations that are biologically sub-clonal and present in a minority of tumor cells.

163. The pharmaceutical composition, the population of T-cells, the T-cell, the TIL population, the T-cell receptor, or the CAR of claim 162, wherein the first sub-population is over-represented in the mutations to which the individualized neoepitopes correspond.

164. The pharmaceutical composition, the population of T-cells, the T-cell, the TIL population, the T-cell receptor, or the CAR of any one of claims 162or 163, wherein the- 364 - 13241940vlAttorney Docket No. 2013237-1502second sub-population is over-represented in the mutations to which the individualized neoepitopes correspond.

165. The pharmaceutical composition, the population of T-cells, the T-cell, the TIL population, the T-cell receptor, or the CAR of any one of claims 162to 164, wherein substantially all of the individualized neoepitopes correspond to mutations from the first subpopulation.

166. The pharmaceutical composition, the population of T-cells, the T-cell, the TIL population, the T-cell receptor, or the CAR of any one of claims 162to 165, wherein substantially all of the individualized neoepitopes correspond to mutations from the first and sub-populations.

167. A method comprising:obtaining a list of candidate TCR sequences for a subject;detecting a plurality of candidate mutations in tumor cells of from the subject and selecting, using a method or system of any one of claims 1-136, a subset of the plurality of candidate mutations;identifying, using the list of candidate TCR sequences and the selected subset of candidate mutations, one or more TCR-mutation pairings, each candidate TCR-mutation pairing comprising a particular TCR sequence of the list of candidate TCR sequences and a particular candidate mutation of the selected subset, wherein the particular candidate TCR sequence is identified as capable of specifically binding to a complex comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising the particular candidate mutation, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells; andusing the one or more candidate TCR-mutation pairings to produce a personalized cancer immunotherapy.- 365 - 13241940vl