Technologies for determining tumor sample purity and / orgenomic characteristics using single nucleotide variations

SNV-based analysis identifies clonal mutations in tumor genomes to accurately estimate purity, addressing the challenge of molecular heterogeneity in tumors and improving cancer therapy precision.

WO2026159256A1PCT designated stage Publication Date: 2026-07-30BIONTECH SE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BIONTECH SE
Filing Date
2026-01-23
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Current cancer therapies are often ineffective due to the molecular heterogeneity of tumors, and existing methods struggle to accurately determine tumor sample purity, especially at low tumor content levels, limiting the detection and characterization of somatic mutations.

Method used

A method using single nucleotide variation (SNV) analysis to identify biologically clonal mutations within balanced diploid regions of the tumor genome, allowing for accurate estimation of tumor sample purity without requiring exact copy number determination, and leveraging variant allele frequencies to distinguish between different subpopulations.

Benefits of technology

Enables precise determination of tumor sample purity and somatic mutations, enhancing cancer diagnosis, treatment planning, and personalized therapy development, particularly at low tumor content levels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2026051711_30072026_PF_FP_ABST
    Figure EP2026051711_30072026_PF_FP_ABST
Patent Text Reader

Abstract

Presented herein are methods and systems that allow biophysical and / or genomic properties of tumor samples to be determined via analysis of sequencing data, obtained, for example, via next-generation sequencing of tumor samples from a subject. Among other things, tumor deconvolution technologies of the present disclosure identify and analyze mutations, such as single nucleotide variations (SNVs) that occur in cancer cells, to characterize tumor samples, for example by estimating tumor sample purity. Among other things, obtaining accurate estimates of tumor sample purity allows putative mutations to be detected and / or confirmed as actual, underlying physical, somatic mutations present in cancer cells at high levels of accuracy. In turn, an accurate list of somatic mutations characteristic of a subject's cancer can be used to diagnose the patient, determine treatment - for example via selection of targeted therapies, and / or develop personalized treatments, such as individualized cancer vaccines.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No. 2013237-1500TECHNOLOGIES FOR DETERMINING TUMOR SAMPLE PURITY AND / OR GENOMIC CHARACTERISTICS USING SINGLE NUCLEOTIDE VARIATIONSCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to and benefit of U. S. Provisional Application No.63 / 749,339, filed on January 24, 2025, the content of which is hereby incorporated by reference herein in its entirety.BACKGROUND

[0002] Cancer is a primary cause of mortality, accounting for 1 in 4 of all deaths.Despite recent advances in the field of cancer immunotherapy there remains no single, broadly applicable treatment. Molecular heterogeneity of tumors renders many therapies ineffective for cancer patients.SUMMARY

[0003] Presented herein are methods and systems that allow biophysical and / or genomic properties of tumor samples to be determined via analysis of sequencing data, obtained, for example, via next-generation sequencing of tumor samples from a subject. Among other things, tumor deconvolution technologies of the present disclosure identify and analyze mutations, such as single nucleotide variations (SNVs) that occur in cancer cells, to characterize tumor samples, for example by estimating tumor sample purity. Among other things, obtaining accurate estimates of tumor sample purity allows putative mutations to be detected and / or confirmed as actual, underlying physical, somatic mutations present in cancer cells at high levels of accuracy. In turn, an accurate list of somatic mutations (e.g., biologically clonal mutations) characteristic of a subject’s cancer can be used (e.g., as a mutation signature) to diagnose the patient, determine treatment - for example via selection of targeted therapies, and / or develop personalized treatments, such as individualized cancer vaccines.

[0004] In certain embodiments, using SNV mutation signatures for estimating tumor sample purity in accordance with technologies described herein provides advantages at low tumor sample purities in comparison with other techniques. For example, at low tumor sample purities, other signatures of cancer cells, such as copy number variations (CNVs) may be difficult to resolve in sequencing data (e.g., may be overwhelmed by noise) and, accordingly, are - 1 - 1324629 IvlAttomey Docket No. 2013237-1500limited in their ability to serve as a basis for purity estimation techniques. Among other things, as described and demonstrated in further detail herein, in certain embodiments, SNV-based purity estimation technologies of the present disclosure allow for tumor sample purity to be accurately determined from whole exome sequencing (WES) data even when tumor content ranges from as little as ten to fifteen percent of a sample. Among other things, accurate estimation of sample purity at such low levels is believed to be particularly impactful on sensitivity and specificity of mutation detection.

[0005] In certain embodiments, SNV-based purity estimation technologies of the present disclosure identify and leverage a particular subset of SNVs to determine tumor sample purity. In particular, in certain embodiments, SNV events that are likely to represent underlying mutations that are biologically clonal and lie within balanced diploid regions of a tumor genome are identified and used for estimating tumor sample purity. Among other things, the present disclosure includes the insight that isolating and basing subsequent analysis on these particular SNV events (i.e., biologically clonal mutations within balanced diploid regions of the tumor genome) allows tumor deconvolution techniques to avoid issues that can impact accuracy. In particular, the present disclosure includes recognition of, and techniques for addressing, challenges associated with degeneracies arising, for example, from SNVs located in regions that have copy numbers above two, e.g., due to CNV events that often occur in and complicate analysis of tumor genomes.

[0006] Moreover, in certain embodiments, technologies described herein allow particular desired subsets of SNV events to be identified and isolated without necessarily requiring an exact absolute copy number to be determined for and / or assigned to segments in which SNV events are located, let alone a detailed characterization of CNVs across a tumor genome. Instead, in certain embodiments, methods and systems described herein include approaches for identifying balanced heterozygous segments within a tumor genome and separating them into distinguishable subpopulations according to their absolute copy number. That is, without receiving (e.g., a priori) and / or determining specific values of absolute copy numbers for balanced heterozygous segments within a tumor genome, technologies described herein may nonetheless use sequencing data to identify and distinguish between different subpopulations of segments that have different copy numbers. Among other things, the ability to distinguish- 2 - 1324629 IvlAttomey Docket No. 2013237-1500between and select a desired subset of SNVs located within balance diploid regions of a tumor genome without explicitly determining absolute copy numbers for segments throughout the tumor genome is significant because at low tumor sample purities, the ability to resolve CNV events in sequencing data may fail.

[0007] Accordingly, in certain embodiments, methods and systems described herein may identify a subpopulation of segments that have a same, minimal, absolute copy number that is less than other even numbered absolute copy numbers present in the tumor genome. Since the genome of normal cells is diploid, the subpopulation of segments identified as having the minimal even numbered absolute copy number is likely to be a subpopulation of balanced diploid segments. Additionally or alternatively, the present disclosure further recognizes that certain events occurring in cancer cells, such as whole genome duplications (WGDs) may produce minimal copy numbers above two (e.g., four). Accordingly, in certain embodiments, methods and systems described herein include approaches for (e.g., independently) verifying a likelihood that the minimal copy number is indeed two and / or determining if it is e.g., four or greater. Among other things, these approaches leverage the insight that variant allele frequencies (VAFs) of SNVs that are observed in sequencing data reflect, among other things, a discrete set of possible zygosities that SNVs may take on, depending on the underlying absolute copy number of the regions in the tumor genome where they are located. Accordingly, by first identifying and isolating a particular subset of SNVs that are located within tumor genome segments that have a same, albeit unknown, absolute copy number, their VAF distribution can be analyzed to infer the absolute copy number of the particular subset of SNVs.

[0008] In this way, tumor deconvolution and VAF analysis technologies described herein allow for accurate estimation of sample purities, as well as for verifying and / or determining an absolute copy number for certain subsets of SNVs. As described herein, tumor sample purity estimates determined in accordance with various technologies described herein can be useful in and of themselves, e.g., for characterizing and / or assessing quality of tumor samples, and / or may be used to enhance accuracy with which somatic mutations and detected and / or characterized. Among other things, these capabilities allow for improved cancer diagnosis, treatment planning and / or selection, and, in certain embodiments, creation of personalized therapies, such as individualized cancer vaccines.- 3 - 1324629 IvlAttomey Docket No. 2013237-1500

[0009] In some aspects, the present disclosure provides a method (e.g., a computer-implemented method) (e.g., for determining an estimated purity of a tumor sample based on sequencing data) comprising: (a) obtaining, by a processor of a computing device, tumor sequencing data for a tumor sample obtained from a subject [e.g., the tumor sequencing data comprising a plurality of tumor reads, each tumor read representing a partial nucleotide sequence determined by sequencing the tumor sample]; (b) obtaining, by the processor, a list of putative single nucleotide variant (SNV) events detected based on the tumor sequencing data, each putative SNV event representing one or more potential underlying mutations occurring in a tumor genome of the subject; (c) identifying, of the list of putative SNV events, a subset of minimal balanced SNV events using the tumor sequencing data, the subset of minimal balanced SNV events comprising SNV events determined to represent mutations that: (i) are determined to be balanced [e.g., are determined to be located within locally balanced regions of the tumor genome; e.g., are determined to have an equal number of copies of corresponding maternal and paternal alleles (e.g., via a balanced test and / or contour filtering refinement); e.g., are determined to be located within balanced heterozygous segments] within the tumor genome; and (ii) are located within segments determined to have an absolute copy number equal to a same particular minimal copy number whose exact value is unknown but predicted to be lower than copy numbers of other balanced segments (e.g., the exact value of the minimal copy number is initially unknown, but determined to be smaller relative to other absolute copy numbers of balanced segments); (d) determining, by the processor, an estimated tumor sample purity and / or bound thereon based at least in part on the tumor sequencing data and the subset of minimal balanced SNV events; and (e) storing and / or providing, by the processor, the estimated tumor sample purity and / or bound thereon for display and / or further processing.

[0010] In some embodiments, provided methods comprise using (e.g., by the processor) the estimated tumor sample purity and / or bound thereon to detect and / or prioritize a plurality of somatic mutations.

[0011] In some embodiments, provided methods comprise using at least a portion of the detected and / or prioritized somatic mutations in a personalized immunotherapy [e.g., creating a personalized cancer vaccine encoding one or more neoepitopes corresponding to at least a portion of the somatic mutations].- 4 - 1324629 IvlAttomey Docket No. 2013237-1500

[0012] In some embodiments, provided methods comprise using the plurality of somatic mutations and / or prioritization thereof to determine a cancer status (e.g., a particular type of cancer; e.g., a particular stage of cancer) for the subject.

[0013] In some embodiments, provided methods comprise using the plurality of somatic mutations and / or prioritization thereof to select a therapy for the subject.

[0014] In some embodiments, provided methods comprise using the estimated tumor sample purity and / or bound thereon to determine a cancer status (e.g., a particular type of cancer; e.g., a particular stage of cancer) for the subject.

[0015] In some embodiments, provided methods comprise using the plurality of somatic mutations and / or prioritization thereof to select a therapy for the subject.

[0016] In some embodiments, provided methods comprise using the estimated tumor sample purity and / or bound thereon as a quality control [e.g., identifying the tumor sequencing data as insufficient quality based on the estimated tumor sample purity and / or bound thereon (e.g., having been determined to be below a threshold value); e.g., identifying the tumor sample as insufficient quality based on the estimated tumor sample purity and / or bound thereon].

[0017] In some embodiments, provided methods comprise identifying, by the processor, a set of balanced heterozygous segments (BHSs) within the tumor genome of the subject (e.g., and, optionally, also within a normal genome associated with a sample), each BHS of the set having been determined to comprise one or more heterozygous SNPs, at least a portion of which are balanced [e.g., are determined to have an equal number of copies of corresponding maternal and paternal alleles (e.g., via a balanced test and / or contour filtering refinement)] within the tumor genome; and selecting, by the processor, a subset of the set of BHSs as minimal balanced heterozygous segments (MBHSs) having the same particular minimal copy number (e.g., a minimal copy number), thereby identifying the plurality of MBHSs.

[0018] In some embodiments, provided methods comprise determining, by the processor, for each of at least a portion of the set of BHSs, a measure of tumor read counts based at least in part on the tumor sequencing data, thereby determining a distribution of the measure of tumor read counts for set of BHSs; identifying, by the processor, one or more components of the distribution, each of the one or more subcomponents representing (e.g., representing values of- 5 - 1324629 IvlAttomey Docket No. 2013237-1500measure of tumor read counts of) a subpopulation of BHSs having a particular absolute copy number, distinct from that of other subpopulations represented by other components of the distribution; and selecting, by the processor, based on the identified components of the distribution, the subset of MBHSs (e.g., having the minimal copy number which is less than that of other subpopulations).

[0019] In some embodiments, provided methods comprise obtaining, by the processor, normal sequencing data for a normal sample obtained from the subject [e.g., the normal sequencing data comprising a plurality of normal reads, each normal read representing a partial nucleotide sequence determined by sequencing the normal sample], and wherein the measure of tumor read counts is a tumor- to-normal read count ratio determined, for a given BHS, based on a number of tumor reads mapping to the given BHS and a number of normal reads mapping to the given BHS (e.g., as a ratio).

[0020] In some embodiments, provided methods comprise determining, by the processor, for each of at least a portion of the set of BHSs, a tumor- to-normal read count ratio based on the tumor sequencing data and the normal sequencing data, thereby determining a tumor-to-normal read count ratio distribution for the set of BHSs; identifying, by the processor, one or more components of the tumor-to-normal read count ratio distribution, each of the one or more subcomponents representing (e.g., representing tumor-to-normal read count ratios of) a subpopulation of BHSs having a particular absolute copy number, distinct from that of other subpopulations represented by other components of the tumor-to-normal read count ratio distribution; and selecting, by the processor, based on the identified components of the tumor-to-normal read count ratio distribution, the subset of MBHSs (e.g., having the minimal copy number which is less than that of other subpopulations).

[0021] In some embodiments, provided methods comprise identifying, by the processor, of the one or more components, a minimal component having a mean value of the measure of tumor read counts (e.g., a mean tumor-to-normal read count ratio value) lower than that of other components [e.g., determining, for each of the one or more components, a corresponding mean tumor-to-normal read count ratio value (e.g., representative of an average tumor-to-normal read count ratio of the subpopulation of BHSs represented by the component) and selecting, by the processor, the particular component having the lowest mean tumor-to-normal read count ratio - 6 - 1324629 IvlAttomey Docket No. 2013237-1500value as the minimal component]; and selecting, by the processor, BHSs determined to be members of the subpopulation represented by the minimal component for inclusion in the subset of MBHSs {e.g., by: for each segment of at least a portion of the BHSs, determining one or more posterior probabilities associated with the segment and corresponding to a particular one of the one or more components of the distribution (e.g., the tumor-to-normal read count ratio distribution) for the set of BHSs and measuring a likelihood of the segment being a member of the subpopulation represented by the particular component (e.g., based on a tumor-to-normal read count ratio for the segment); and identifying the members of the subpopulation represented by the minimal component based on the determined posterior probabilities for each of the portion of BHSs [e.g., identifying segments for which the associated posterior probability corresponding to the minimal component is larger than the posterior probabilities associated with the segment] }.

[0022] In some embodiments, provided methods comprise identifying the minimal component, comprising fitting, by the processor, a mixture model (e.g., a ID Gaussian Mixture Model) to the (e.g., tumor-to-normal read count ratio) distribution, thereby determining, based on the fit, a plurality of parameter values (e.g., for each component a corresponding mean, amplitude, and standard deviation), including, for each of the one or more components, a corresponding mean; and identifying the minimal component as having a lowest corresponding mean.

[0023] In some embodiments, provided methods comprise selecting, by the processor, a subpopulation of minimal segments identified as also having the same particular minimal copy number (e.g., comprising both balanced and unbalanced segments and / or both heterozygous and homozygous segments) [e.g., by determining a sample distribution of one or more sequencing data observations (e.g., a tumor-to-normal read count ratio, e.g., a residual error distribution) based on the set of MBHSs and using the sample distribution to select the subpopulation of minimal segments ].

[0024] In some embodiments, provided methods comprise refining, by the processor, the set of MBHSs using the subpopulation of minimal segments (e.g., by removing segments from the population of MBHSs that are not in the subpopulation of minimal segments, thereby refining the population of MBHSs).- 7 - 1324629 IvlAttomey Docket No. 2013237-1500

[0025] In some embodiments, provided methods comprise selecting, by the processor, from the list of putative SNV events, a subset of minimal SNV events comprising putative SNV events determined to map to at least one of the one or more minimal segments.

[0026] In some embodiments, provided methods comprise identifying the subset of minimal balanced SNV events from (e.g., by selecting a subset of) the subset of minimal SNV events.

[0027] In some embodiments, provided methods comprise filtering, by the processor, the subset of minimal SNV events to retain SNV events determined to be located in locally balanced regions of the tumor genome, thereby identifying the subset of minimal balanced SNV events.

[0028] In some embodiments, provided methods comprise selecting, by the processor, for inclusion in the set of minimal balanced SNV events, one or more of the subset of minimal SNV events determined to map to one or more balanced heterozygous segments.

[0029] In some embodiments, provided methods comprise selecting, by the processor, from the list of putative SNV events, SNV events determined to map to at least one of the one or more minimal balanced heterozygous segments (MBHSs) for inclusion in the set of minimal balanced SNV events.

[0030] In some embodiments, provided methods comprise determining, by the processor, for each SNV event of the subset of minimal balanced SNV events, a variant allele frequency (VAF), thereby determining a VAF distribution for the subset of minimal balanced SNV events; identifying, by the processor, within the VAF distribution for the subset of minimal balanced SNV events, a plurality of (e.g., discrete and distinguishable) VAF subpopulations, each corresponding to and representing VAFs of a distinct SNV subpopulation (e.g., two VAF subpopulations corresponding to two SNV subpopulations; e.g., three VAF subpopulations corresponding to three SNV subpopulations); identifying, by the processor, from the plurality of components, a main VAF subpopulation corresponding to a desired subpopulation of clonal heterozygous SNVs located in balanced diploid regions of the tumor genome; filtering, by the processor, the subset of minimal balanced SNVs based on the VAF distribution to retain and / or enrich for SNVs corresponding to the main VAF subpopulation [e.g., while excluding (e.g., removing) other subpopulations of SNVs corresponding to other components], thereby determining a filtered subset of (e.g., clonal) minimal balanced SNV events; and using, by the - 8 - 1324629 IvlAttomey Docket No. 2013237-1500processor, the filtered subset of minimal balanced SNV events to determine the estimated tumor sample purity at step (e).

[0031] In some embodiments, a plurality of VAF subpopulations comprises a low VAF subpopulation corresponding to a subpopulation of subclonal SNVs having VAFs below those of the main VAF subpopulation, and filtering the subset of minimal balanced SNVs comprises identifying and excluding SNVs of the subclonal subpopulation (e.g., based on their VAFs in relation to one or more parameters (e.g., a mean, median, standard deviation) of the low VAF subpopulation and / or the main VAF subpopulation].

[0032] In some embodiments, provided methods comprise confirming the main VAF subpopulation and the low VAF subpopulation are not harmonics [e.g., confirming that a relative difference and / or ratio between VAF(s) (e.g., as measured by a representative VAF, such as a mean, median, etc.) of the main subpopulation (e.g., the VAF subpopulation having the most SNV events) and the low VAF subpopulation (e.g., as measured by a representative VAF, such as a mean, median, etc.) is not indicative of two SNV subpopulations with a same even-numbered absolute copy number (e.g., greater than 2, e.g., 4, e.g., 6, e.g., 8) but different zygosities (e.g., wherein the different zygosities are 1 and 2, 1 and 3, or 1 and 4)].

[0033] In some embodiments, the plurality of VAF subpopulations comprises a high VAF component corresponding to a subpopulation of unbalanced SNVs, and filtering the subset of minimal balanced SNVs comprises identifying and excluding SNVs of the unbalanced subpopulation (e.g., based on their VAFs in relation to one or more parameters (e.g., a mean, median, standard deviation) of the high VAF subpopulation and / or the main VAF subpopulation].

[0034] In some embodiments, the plurality of VAF subpopulations comprises a low VAF subpopulation corresponding to a subpopulation of high copy number and low zygosity SNVs, and filtering the subset of minimal balanced SNVs comprises identifying and excluding SNVs of the high copy number and low zygosity subpopulation (e.g., based on their VAFs in relation to one or more parameters (e.g., a mean, median, standard deviation) of the low VAF subpopulation and / or the main VAF subpopulation].

[0035] In some embodiments, the plurality of VAF subpopulations comprises a high VAF subpopulation corresponding to a subpopulation of high copy number (and high zygosity) SNVs, and filtering the subset of minimal balanced SNVs comprises identifying and excluding SNVs of - 9 - 1324629 IvlAttomey Docket No. 2013237-1500the high copy number subpopulation (e.g., based on their VAFs in relation to one or more parameters (e.g., a mean, median, standard deviation) of the high VAF subpopulation and / or the main VAF subpopulation].

[0036] In some embodiments, the one or more sets of VAF outliers comprise a set of high VAF outliers comprising minimal balanced SNVs having VAFs significantly greater than a first representative (e.g., average, median, etc.) VAF of the subset of minimal balanced SNV events and / or representative VAF of a subpopulation thereof (e.g., a main subpopulation).

[0037] In some embodiments, the one or more sets of VAF outliers comprise a set of high VAF outliers comprising minimal balanced SNVs having VAFs significantly greater than a first representative (e.g., average, median, etc.) VAF of the subset of minimal balanced SNV events and / or a subpopulation thereof (e.g., a main subpopulation).

[0038] In some embodiments, provided methods comprise determining, by the processor, a second representative VAF for the set of minimal balanced SNVs (e.g., a measure of central tendency, such as a median, a mean, etc.) based on the VAF distribution for the subset of minimal balanced SNV events; identifying, by the processor, of the set of minimal balanced SNVs, a main (e.g., putative clonal) subpopulation comprising a subset of minimal balanced SNVs having VAFs similar [e.g., within predetermined factor times an expected standard deviation for a given SNV; e.g., statistically close to (e.g., within a standard statistical error range, e.g., calculated based on coverage for the SNV)] to the second representative VAF determined for the set of minimal balanced SNVs; determining, by the processor, a first representative VAF (e.g., a measure of central tendency, such as a median, a mean, etc.) for the main population; and identifying, by the processor, the one or more sets of VAF outliers based at least in part on the first representative VAF determined for the main population.

[0039] In some embodiments, provided methods comprise identifying, by the processor, a refined main subpopulation comprising a subset of minimal balanced SNVs having VAFs similar [e.g., within predetermined factor times a an expected standard deviation for a given SNV; e.g., statistically close to (e.g., within a standard statistical error range, e.g., calculated based on coverage for the SNV)] the first representative VAF determined for the main subpopulation; determining, by the processor, a representative VAF for the refined main- 10 - 1324629 IvlAttorney Docket No. 2013237-1500subpopulation; and identifying, by the processor, the one or more sets of VAF outliers based at least in part on the first representative VAF determined for the refined main population.

[0040] In some embodiments, the one or more sets of VAF outliers comprise a set of low VAF outliers and provided methods comprise identifying, by the processor, an initial set of low VAF outliers comprising minimal balanced SNVs having VAFs below (e.g., a statistically significant standardized distance below) the representative VAF for the refined main subpopulation; determining, by the processor, an initial filtered subset of (e.g., clonal) minimal balanced SNVs excluding the initial set of low VAF outliers; and iteratively updating, by the processor, the set of low VAF outliers and the initial filtered subset of minimal balanced SNVs by, beginning with a highest VAF member of the set of low VAF outliers, iteratively: moving a SNV from the set of low VAF outliers to the initial filtered subset if its VAF is within a predetermined threshold value of a minimum VAF of the initial filtered subset; and progressing to evaluate a next lower VAF SNV of the set of low VAF outliers at a subsequent iteration.

[0041] In some embodiments, the one or more sets of VAF outliers comprise a set of high VAF outliers and provided method comprises identifying, by the processor, an initial set of high VAF outliers comprising minimal balanced SNVs having VAFs above (e.g., a statistically significant standardized distance above) the representative VAF for the refined main subpopulation; determining, by the processor, an initial filtered subset of (e.g., clonal) minimal balanced SNVs excluding the initial set of high VAF outliers; and iteratively updating, by the processor, the set of high VAF outliers and the initial filtered subset of minimal balanced SNVs by, beginning with a lowest VAF member of the set of low VAF outliers, iteratively: moving a SNV from the set of high VAF outliers to the initial filtered subset if its VAF is within a predetermined threshold value of a maximum VAF of the initial filtered subset; and progressing to evaluate a next higher VAF SNV of the set of high VAF outliers at a subsequent iteration.

[0042] In some embodiments, provided methods comprise, at step (e), determining, by the processor, one or more sequencing data observations for the subset of minimal balanced SNV events (e.g., allele- specific read counts, variant allele frequencies, and / or distributions thereof) based on the tumor sequencing data, thereby determining one or more measured sequencing data observations; determining, by the processor, for each of one or more SNV models, one or more corresponding initial purity estimates, wherein: each SNV model generates one or more- 11 - 1324629 IvlAttomey Docket No. 2013237-1500predicted sequencing data observations as a function of a variable tumor sample purity parameter and determining the initial purity estimate for a corresponding SNV model comprises determining a value of the variable tumor sample purity that optimizes a fit between the predicted sequencing data observations of the SNV model and corresponding measured sequencing data observations; and determining, by the processor, the estimated tumor sample purity based at least in part on the one or more initial purity estimates.

[0043] In some embodiments, one or more SNV models do not fix (e.g., predetermine) a SNV genotype and determining the initial purity estimate for a corresponding SNV model (e.g., further) comprises determining individual SNV genotypes for the set of minimal balanced SNV events and / or a value of one or more variable parameters that measure a quantity of SNV events having a particular genotype.

[0044] In some embodiments, provided methods comprise, at step (e), determining, by the processor, for each particular SNV of at least a portion of the set of minimal balanced SNV events, a corresponding set of allele- specific read counts, wherein, for the particular SNV, the corresponding set of allele- specific read counts comprises a number of reads mapping to each of four possible alleles (e.g., corresponding to each of four possible nucleotide bases at a location of the particular SNV, e.g., as illustrated in FIG. 3 IB); determining, by the processor, a first purity value (e.g., a maximum likelihood estimate) based on a plurality of SNV likelihood functions, each associated with a particular SNV event of the portion of minimal balanced SNV events and measuring a likelihood of measuring the corresponding set of allele- specific read counts for the particular SNV as a function of tumor sample purity, wherein the first purity value is a value of tumor sample purity determined to maximize an aggregate (e.g., log-likelihood sum) of the plurality of SNV likelihood functions; determining, by the processor, the estimated tumor sample purity based on the first purity value.

[0045] In some embodiments, provided methods comprise, at step (e), determining, by the processor, an observed distribution of observed VAFs for at least a portion of the subset of minimal balanced SNV events; determining, by the processor, a predicted distribution of VAFs for the subset of minimal balanced SNV events (e.g., given the determined value of the minimal copy number and an assumed zygosity equal to the determined value of the minimal copy number divided by 2); determining, by the processor, a second purity value (e.g., a value equal to - 12 - 1324629 IvlAttorney Docket No. 2013237-15001 - a normal contamination,u) that optimizes one or more test statistics measuring similarity between the observed and predicted distributions; and determining, by the processor, at step (e), the estimated tumor sample purity based at least in part on the second purity value.

[0046] In some embodiments, the one or more test statistics comprises a ^-statistic that measures (e.g., is inversely proportional to) a difference between the observed distribution and the predicted distribution (e.g., a Kullback-Leibler (KL) divergence).

[0047] In some embodiments, provided methods comprise, at step (e), determining the estimated tumor sample purity based on both an initial purity estimate and a second initial purity estimate.

[0048] In some embodiments, provided methods comprise using the tumor sequencing data (e.g., and normal sequencing data) together with the estimated tumor sample purity to detect a second plurality of SNVs within the tumor genome of the subject and / or to refine the list of putative single nucleotide variant (SNV) events.

[0049] In some embodiments, provided methods comprise obtaining [e.g., receiving and / or accessing; e.g., detecting, by the processor, based on the tumor sequencing data (e.g., and the normal sequencing data; and the estimated tumor sample purity)], by the processor, a list of candidate mutations (e.g., SNVs) identifying a plurality of mutations determined to be present in the tumor genome of the subject (e.g., wherein the list of candidate mutations is the list of putative SNV events and / or a filtered version thereof; e.g., wherein the list of candidate mutations comprises at least a portion of SNV events of the list of putative SNV events; e.g., wherein the list of candidate mutations is different from the list of putative SNV events); identifying, by the processor, for each particular mutation of the list of candidate mutations, a corresponding segment within the tumor genome comprising the particular mutation; determining, by the processor, for each particular mutation of the list of candidate mutations, an estimated cellularity and / or variability thereof (e.g., standard deviation; e.g., a confidence interval) based at least in part on the estimated tumor sample purity, thereby determining estimated cellularity values and / or variabilities thereof for the list of candidate mutations; and selecting, by the processor, a subset of the list of candidate mutations {e.g., for inclusion in and / or targeting via a construct [e.g., a therapeutic agent (e.g., a personalized therapeutic agent, such as a personalized cancer vaccine, a T cell receptor (TCR) therapy, etc.)]} based at least in - 13 - 1324629 IvlAttomey Docket No. 2013237-1500part on the estimated cellularity values and / or variabilities thereof for the list of candidate mutations.

[0050] In some embodiments, provided methods comprise obtaining [e.g., receiving and / or accessing; e.g., detecting, by the processor, based on the tumor sequencing data (e.g., and the normal sequencing data; and the estimated tumor sample purity)], by the processor, a list of candidate mutations (e.g., SNVs) identifying a plurality of mutations determined to be present in the tumor genome of the subject (e.g., wherein the list of candidate mutations is the list of putative SNV events and / or a filtered version thereof; e.g., wherein the list of candidate mutations comprises at least a portion of SNV events of the list of putative SNV events; e.g., wherein the list of candidate mutations is different from the list of putative SNV events); identifying, by the processor, for each particular mutation of the list of candidate mutations, a corresponding segment within the tumor genome comprising the particular mutation; determining, by the processor, for each particular mutation of the list of candidate mutations, an identification of the particular mutation as clonal or subclonal based at least in part on the estimated sample purity, thereby identifying clonal and / or subclonal mutations within the list of candidate mutations; and selecting, by the processor, a subset of the list of candidate mutations {e.g., for inclusion in and / or targeting via a construct [e.g., a therapeutic agent (e.g., a personalized therapeutic agent, such as a personalized cancer vaccine, a T cell receptor (TCR) therapy, etc.)]} based at least in part on the identified clonal and / or subclonal mutations within the list of candidate mutations.

[0051] In some embodiments, tumor sequencing data and / or normal sequencing data are whole genome sequencing data (WGS), whole exome sequencing data (WES), or single nucleotide polymorphism (SNP) array data.

[0052] In some embodiments, tumor sequencing data and / or normal sequencing data comprise a plurality of replicates.

[0053] In some embodiments, provided methods comprise using the plurality of replicates (e.g., to correct errors and reduce false positives) to detect at least a portion of the putative SNVs of the list [e.g., wherein the plurality of replicates are obtained from multiple samples of extracted gDNA from tumor and normal samples, obtained for multiple library- 14 - 1324629 IvlAttorney Docket No. 2013237-1500preparations (e.g., from a single sample), obtained via multiple sequencing runs on a single library (e.g., technical replicates)].

[0054] In some embodiments, the tumor sample is a formalin-fixed paraffin embedded (FFPE) sample.

[0055] In some embodiments, the tumor sample is a fresh frozen sample.

[0056] In some embodiments, provided methods comprise obtaining, by the processor, normal sequencing data for a normal sample obtained from the subject [e.g., the normal sequencing data comprising a plurality of normal reads, each normal read representing a partial nucleotide sequence determined by sequencing the normal sample]; and at step (d), identifying the subset of minimal balanced SNV events based at least in part on the tumor sequencing data and (e.g., in combination with) the normal sequencing data.

[0057] In some embodiments, provided methods comprise obtaining, by the processor, normal sequencing data for a normal sample obtained from the subject [e.g., the normal sequencing data comprising a plurality of normal reads, each normal read representing a partial nucleotide sequence determined by sequencing the normal sample]; and at step (e), determining the estimated tumor sample purity based on the normal sequencing data together with the tumor sequencing data and the subset of minimal balanced SNV events.

[0058] In some embodiments, provided methods comprise, at step (b), detecting, by the processor, at least a portion of the putative SNV events of the list based at least in part on the tumor sequencing data.

[0059] In some embodiments, provided methods comprise detecting the portion of the putative SNV events using, the tumor sequencing data together with normal sequencing data obtained for a normal sample obtained from the subject.

[0060] In some embodiments, the normal sample is a buffy coat sample.

[0061] In some embodiments, provided methods comprise detecting the portion of the putative SNV events using, the tumor sequencing data together with a normal reference genome.

[0062] In some embodiments, the tumor sample is obtained from a subject diagnosed as having or at risk of having cancer.- 15 - 1324629 IvlAttomey Docket No. 2013237-1500

[0063] In some embodiments, the tumor sample is obtained from a subject diagnosed as having or at risk of having a cancer associated with low tumor mutation burden.

[0064] In some embodiments, the tumor sample is obtained from a subject diagnosed as having or at risk of having breast cancer, prostate cancer, pancreatic cancer, pediatric cancer (e.g., pediatric acute lymphoblastic leukemia), neuroblastoma, ovarian cancer, renal cell carcinoma, Merkel cell carcinoma, hematologic cancer, colorectal cancer, melanoma, head and neck squamous cell carcinoma, or non-small cell lung cancer.

[0065] In some embodiments, the estimated tumor sample purity is less than about 0.3 (e.g., less than 0.3, less than 0.25, less than 0.20, less than 0.19, less than 0.18, less than 0.17, less than 0.16, less than 0.15, less than 0.14, less than 0.13, less than 0.12, or less than 0.11).

[0066] In some embodiments, the estimated tumor sample purity is at least about 0.10 (e.g., at least 0.10, at least 0.11, at least 0.12, at least 0.13, at least 0.14, at least 0.15, at least 0.16, at least 0.17, at least 0.18, at least 0.19).

[0067] In some aspects, the present disclosure provides a method (e.g., a computer implemented method) for determining a minimal copy number, or a lower bound thereon, of a set of balanced heterozygous segments in a tumor genome, the method comprising: (a) obtaining, by a processor of a computing device, tumor sequencing data for a tumor sample obtained from a subject [e.g., the tumor sequencing data comprising a plurality of tumor reads, each tumor read representing a partial nucleotide sequence determined by sequencing the tumor sample]; (b) obtaining, by the processor, a list of putative single nucleotide variant (SNV) events detected based on the tumor sequencing data, each putative SNV event representing one or more potential underlying mutations occurring in a tumor genome of the subject; (c) identifying, of the list of putative SNV events, a subset of minimal balanced SNV events comprising SNV events determined to represent mutations that: (i) are determined to be balanced [e.g., are determined to have an equal number of copies of corresponding maternal and paternal alleles (e.g., via a balanced test and / or contour filtering refinement); e.g., are determined to be located within balanced heterozygous segments] within the tumor genome; and (ii) are located within segments determined to have an absolute copy number equal to a same particular minimal copy number whose exact value is unknown but predicted to be lower than copy numbers of other balanced segments (e.g., the exact value of the minimal copy number is initially unknown, but determined - 16 - 1324629 IvlAttorney Docket No. 2013237-1500to be smaller relative to other absolute copy numbers of balanced segments); (d) determining, by the processor, for at least a portion of the subset of minimal balanced SNV events, a corresponding variant allele frequency (VAF) based on the tumor sequencing data, thereby determining a VAF distribution for the subset of minimal balanced SNV events; (e) identifying, by the processor, one or more (e.g., a plurality) of VAF subpopulations within the VAF distribution for the subset of minimal balanced SNV events, each VAF subpopulation corresponding to and representing VAFs of a distinct subpopulation of SNV events; (f) determining, by the processor, an estimated value of the minimal copy number and / or a lower bound thereon (e.g., four or higher), based on the one or more VAF subpopulations; and (g) storing and / or providing, by the processor, for further processing, (i) the estimated value of the minimal copy number and / or the lower bound, for display and / or further processing.

[0068] In some embodiments, the identified VAF subpopulations comprise at least two subpopulations, including a first VAF subpopulation corresponding to a subpopulation of (e.g., clonal) SNVs having a first absolute copy number and a first zygosity and a second VAF subpopulation corresponding to a subpopulation of (e.g., clonal) SNVs (e.g., also) having the first absolute copy number and a second zygosity, different from the first zygosity.

[0069] In some embodiments, provided methods comprise determining the estimated value of the minimal copy number and / or lower bound thereon based at least in part on the identification of the first and second VAF subpopulations.

[0070] In some embodiments, the estimated value of the minimal copy number and / or lower bound thereon is four (4).

[0071] In some embodiments, provided methods comprise identifying the first and second VAF subpopulations as harmonics [e.g., corresponding to subpopulations of SNVs for which the first and second zygosities having a ratio of 1 to half the first absolute copy number (e.g., the first zygosity is one and the second zygosity is half the first absolute copy number)] (e.g., based on a difference between representative VAF measures of the first and second VAF subpopulations).

[0072] In some embodiments, provided methods comprise identifying the first VAF subpopulation to correspond to subpopulation of late SNV events, occurring after a duplication event (e.g., a whole genome duplication event) and identifying the second VAF subpopulation to - 17 - 1324629 IvlAttorney Docket No. 2013237-1500correspond to a subpopulation of early SNV events, occurring prior to a duplication event (e.g., the whole genome duplication event).

[0073] In some embodiments, provided methods comprise determining a first representative VAF measure (e.g., a mean, median, etc.) of the first VAF subpopulation and a second representative VAF measure (e.g., a mean, median, etc.) of the second VAF subpopulation; and determining the minimal copy number and / or lower bound thereon based at least in part on the first representative VAF measure and the second representative VAF measure.

[0074] In some embodiments, provided methods comprise identifying the first and second VAF subpopulations as harmonics based at least in part on the first and second representative VAF measures (e.g., based on a ratio of the first and second representative VAF values).

[0075] In some embodiments, provided methods comprise determining the minimal copy number to be four based at least in part on a ratio between the first and second representative VAF measures being approximately equal to 1:2.

[0076] In some embodiments, provided methods comprise detecting presence of a whole genome duplication event based on the one or more VAF subpopulations.

[0077] In some embodiments, provided methods comprise determining a number of SNVs in each of the first and second VAF subpopulations and estimating the value of the minimal copy number and / or bound thereon based at least in part on the number of SNVs in the first and second VAF subpopulations.

[0078] In some embodiments, provided methods comprise aborting, by the processor, a SNV-based purity estimation procedure based at least in part on the estimated minimal copy number and / or bound thereon.

[0079] In some embodiments, provided methods comprise aborting, by the processor, a CNV-based purity estimation procedure based at least in part on the estimated minimal copy number and / or bound thereon {e.g., determining, by the processor, a CNV-based estimate of the minimal copy number based on the CNV-based purity estimation procedure; determining, by the processor, the CNV-based estimate of the minimal copy number to differ from the inferred value of the minimal copy number; and responsive to the determining the CNV-based estimate of the - 18 - 1324629 IvlAttomey Docket No. 2013237-1500minimal copy number to differ from the inferred value of the minimal copy number, aborting, by the processor, the CNV-based purity estimation procedure.}.

[0080] In some embodiments, provided methods comprise using the estimated minimal copy number and / or to determine a CNV-based estimate of tumor sample purity and / or a bound thereon.

[0081] In some embodiments, one or more VAF subpopulations comprises a main VAF subpopulation corresponding to a subpopulation of balanced clonal SNVs having a single zygosity.

[0082] In some embodiments, provided methods comprise determining the estimated minimal copy number and / or lower bound thereon to be two (2) based at least in part on the identified main VAF subpopulation.

[0083] In some embodiments, one or more VAF subpopulations comprises a low VAF subpopulation corresponding to a subpopulation of subclonal SNVs.

[0084] In some embodiments, provided methods comprise filtering the set of minimal balanced SNVs to remove the subpopulation of subclonal SNVs corresponding to the low VAF subpopulation.

[0085] In some embodiments, provided methods comprise aborting an SNV-based purity estimation procedure based at least in part on the low VAF subopulation (e.g., a number of SNVs within the subpopulation of subclonal SNVs corresponding to the low VAF subpopulation).

[0086] In some aspects, the present disclosure provides a method (e.g., a computer implemented method) for determining a primary copy number, or a lower bound thereon, for a set of balanced heterozygous segments in a tumor genome, the method comprising: (a) obtaining, by a processor of a computing device, tumor sequencing data for a tumor sample obtained from a subject [e.g., the tumor sequencing data comprising a plurality of tumor reads, each tumor read representing a partial nucleotide sequence determined by sequencing the tumor sample]; (b) obtaining, by the processor, a list of putative single nucleotide variant (SNV) events detected based on the tumor sequencing data, each putative SNV event representing one or more potential underlying mutations occurring in a tumor genome of the subject; (c) identifying, of the list of putative SNV events, a subset of primary balanced SNV events comprising SNV events- 19 - 1324629 IvlAttomey Docket No. 2013237-1500determined to represent mutations that: (i) are determined to be balanced [e.g., are determined to have an equal number of copies of corresponding maternal and paternal alleles (e.g., via a balanced test and / or contour filtering refinement); e.g., are determined to be located within balanced heterozygous segments] within the tumor genome; and (ii) are located within segments determined to have an absolute copy number equal to a same particular primary copy number whose exact value is unknown but predicted to be a most frequently occurring even absolute copy number [(e.g., the exact value of the primary copy number is initially unknown, but determined to be even and to be more frequently occurring than other even copy numbers (e.g., determined to be the most commonly occurring absolute copy number within a set of balanced heterozygous segments)]; (d) determining, by the processor, for at least a portion of the subset of primary balanced SNV events, a corresponding variant allele frequency (VAF) based on the tumor sequencing data, thereby determining a VAF distribution for the subset of primary balanced SNV events; (e) identifying, by the processor, one or more (e.g., a plurality) of VAF subpopulations within the VAF distribution for the subset of primary balanced SNV events, each VAF subpopulation corresponding to and representing VAFs of a distinct subpopulation of SNV events; (f determining, by the processor, an estimated value of the primary copy number and / or a lower bound thereon based on the one or more VAF subpopulations, (g) storing and / or providing, by the processor, for further processing, the estimated value of the primary copy number and / or the lower bound thereon for display and / or further processing.

[0087] In some aspects, the present disclosure provides a method (e.g., a computer implemented method) for determining a minimal copy number, or a lower bound thereon, of a set of balanced heterozygous segments in a tumor genome, the method comprising: (a) obtaining, by a processor of a computing device, tumor sequencing data for a tumor sample obtained from a subject [e.g., the tumor sequencing data comprising a plurality of tumor reads, each tumor read representing a partial nucleotide sequence determined by sequencing the tumor sample]; (b) obtaining, by the processor, a list of putative single nucleotide variant (SNV) events detected based on the tumor sequencing data, each putative SNV event representing one or more potential underlying mutations occurring in a tumor genome of the subject; (c) identifying, of the list of putative SNV events, a subset of minimal balanced SNV events comprising SNV events determined to represent mutations that: (i) are determined to be balanced [e.g., are determined to have an equal number of copies of corresponding maternal and paternal alleles (e.g., via a - 20 - 1324629 IvlAttorney Docket No. 2013237-1500balanced test and / or contour filtering refinement); e.g., are determined to be located within balanced heterozygous segments] within the tumor genome; and (ii) are located within segments determined to have an absolute copy number equal to a same particular minimal copy number whose exact value is unknown but predicted to be lower than copy numbers of other balanced segments (e.g., the exact value of the minimal copy number is initially unknown, but determined to be smaller relative to other absolute copy numbers of balanced segments); (d) determining, by the processor, for at least a portion of the subset of minimal balanced SNV events, a corresponding variant allele frequency (VAF) based on the tumor sequencing data, thereby determining a VAF distribution for the subset of minimal balanced SNV events; (e) identifying, by the processor, a plurality of VAF subpopulations within the VAF distribution for the subset of minimal balanced SNV events, each VAF subpopulation corresponding to and representing VAFs of a distinct subpopulation of SNV events; (f) detecting, by the processor, presence of a whole genome duplication event based at least in part on the plurality of VAF subpopulations; and (g) based at least in part on the detected presence of the WGD event, (i) determining a cancer status (e.g., a particular type of cancer; e.g., a particular stage of cancer) for the subject and / or (ii) selecting a therapy for the subject.

[0088] In some aspects, the present disclosure provides a method (e.g., a computer implemented method) for determining a primary copy number, or a lower bound thereon, for a set of balanced heterozygous segments in a tumor genome, the method comprising: (a) obtaining, by a processor of a computing device, tumor sequencing data for a tumor sample obtained from a subject [e.g., the tumor sequencing data comprising a plurality of tumor reads, each tumor read representing a partial nucleotide sequence determined by sequencing the tumor sample]; (b) obtaining, by the processor, a list of putative single nucleotide variant (SNV) events detected based on the tumor sequencing data, each putative SNV event representing one or more potential underlying mutations occurring in a tumor genome of the subject; (c) identifying, of the list of putative SNV events, a subset of primary balanced SNV events comprising SNV events determined to represent mutations that: (i) are determined to be balanced [e.g., are determined to have an equal number of copies of corresponding maternal and paternal alleles (e.g., via a balanced test and / or contour filtering refinement); e.g., are determined to be located within balanced heterozygous segments] within the tumor genome; and (ii) are located within segments determined to have an absolute copy number equal to a same particular primary copy - 21 - 1324629 IvlAttomey Docket No. 2013237-1500number whose exact value is unknown but predicted to be a most frequently occurring even absolute copy number [(e.g., the exact value of the primary copy number is initially unknown, but determined to be even and to be more frequently occurring than other even copy numbers (e.g., determined to be the most commonly occurring absolute copy number within a set of balanced heterozygous segments)]; (d) determining, by the processor, for at least a portion of the subset of primary balanced SNV events, a corresponding variant allele frequency (VAF) based on the tumor sequencing data, thereby determining a VAF distribution for the subset of primary balanced SNV events; (e) identifying, by the processor, one or more (e.g., a plurality) of VAF subpopulations within the VAF distribution for the subset of primary balanced SNV events, each VAF subpopulation corresponding to and representing VAFs of a distinct subpopulation of SNV events; (f) detecting, by the processor, presence of a whole genome duplication event based at least in part on the plurality of VAF subpopulations; and (g) based at least in part on the detected presence of the WGD event, (i) determining a cancer status (e.g., a particular type of cancer; e.g., a particular stage of cancer) for the subject and / or (ii) selecting a therapy for the subject.

[0089] In some aspects, the present disclosure provides a system (e.g., for determining an estimated purity of a tumor sample based on sequencing data) comprising: a processor of a computing device; and memory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to: (a) obtain tumor sequencing data for a tumor sample obtained from a subject [e.g., the tumor sequencing data comprising a plurality of tumor reads, each tumor read representing a partial nucleotide sequence determined by sequencing the tumor sample]; (b) obtain a list of putative single nucleotide variant (SNV) events detected based on the tumor sequencing data, each putative SNV event representing one or more potential underlying mutations occurring in a tumor genome of the subject; (c) identify, from the list of putative SNV events, a subset of minimal balanced SNV events using the tumor sequencing data, the subset of minimal balanced SNV events comprising SNV events determined to represent mutations that: (i) are determined to be balanced [e.g., are determined to be located within locally balanced regions of the tumor genome; e.g., are determined to have an equal number of copies of corresponding maternal and paternal alleles (e.g., via a balanced test and / or contour filtering refinement); e.g., are determined to be located within balanced heterozygous segments] within the tumor genome; and (ii) are located within segments determined to have an absolute copy number equal to a same particular minimal copy number whose exact value is - 22 - 1324629 IvlAttorney Docket No. 2013237-1500unknown but predicted to be lower than copy numbers of other balanced segments (e.g., the exact value of the minimal copy number is initially unknown, but determined to be smaller relative to other absolute copy numbers of balanced segments); (d) determine an estimated tumor sample purity and / or bound thereon based at least in part on the tumor sequencing data and the subset of minimal balanced SNV events; and (e) store and / or provide the estimated tumor sample purity and / or bound thereon for display and / or further processing.

[0090] In some aspects, the present disclosure provides a system (e.g., for determining a minimal copy number, or a lower bound thereon, of a set of balanced heterozygous segments in a tumor genome) comprising: a processor of a computing device; and memory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to: (a) obtain tumor sequencing data for a tumor sample obtained from a subject [e.g., the tumor sequencing data comprising a plurality of tumor reads, each tumor read representing a partial nucleotide sequence determined by sequencing the tumor sample]; (b) obtain a list of putative single nucleotide variant (SNV) events detected based on the tumor sequencing data, each putative SNV event representing one or more potential underlying mutations occurring in a tumor genome of the subject; (c) identify, from the list of putative SNV events, a subset of minimal balanced SNV events comprising SNV events determined to represent mutations that: (i) are determined to be balanced [e.g., are determined to have an equal number of copies of corresponding maternal and paternal alleles (e.g., via a balanced test and / or contour filtering refinement); e.g., are determined to be located within balanced heterozygous segments] within the tumor genome; and (ii) are located within segments determined to have an absolute copy number equal to a same particular minimal copy number whose exact value is unknown but predicted to be lower than copy numbers of other balanced segments (e.g., the exact value of the minimal copy number is initially unknown, but determined to be smaller relative to other absolute copy numbers of balanced segments); (d) determine, for at least a portion of the subset of minimal balanced SNV events, a corresponding variant allele frequency (VAF) based on the tumor sequencing data, thereby determining a VAF distribution for the subset of minimal balanced SNV events; (e) identify one or more (e.g., a plurality) of VAF subpopulations within the VAF distribution for the subset of minimal balanced SNV events, each VAF subpopulation corresponding to and representing VAFs of a distinct subpopulation of SNV events; (f) determine an estimated value of the minimal copy number and / or a lower bound thereon (e.g., four or - 23 - 1324629 IvlAttomey Docket No. 2013237-1500higher), based on the one or more VAF subpopulations; and (g) store and / or provide the estimated value of the minimal copy number and / or the lower bound, for display and / or further processing.

[0091] In some aspects, the present disclosure provides a system (e.g., for determining a primary copy number, or a lower bound thereon, for a set of balanced heterozygous segments in a tumor genome) comprising: a processor of a computing device; and memory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to: (a) obtain tumor sequencing data for a tumor sample obtained from a subject [e.g., the tumor sequencing data comprising a plurality of tumor reads, each tumor read representing a partial nucleotide sequence determined by sequencing the tumor sample]; (b) obtain a list of putative single nucleotide variant (SNV) events detected based on the tumor sequencing data, each putative SNV event representing one or more potential underlying mutations occurring in a tumor genome of the subject; (c) identify, from the list of putative SNV events, a subset of primary balanced SNV events comprising SNV events determined to represent mutations that: (i) are determined to be balanced [e.g., are determined to have an equal number of copies of corresponding maternal and paternal alleles (e.g., via a balanced test and / or contour filtering refinement); e.g., are determined to be located within balanced heterozygous segments] within the tumor genome; and (ii) are located within segments determined to have an absolute copy number equal to a same particular primary copy number whose exact value is unknown but predicted to be a most frequently occurring even absolute copy number [(e.g., the exact value of the primary copy number is initially unknown, but determined to be even and to be more frequently occurring than other even copy numbers (e.g., determined to be the most commonly occurring absolute copy number within a set of balanced heterozygous segments)]; (d) determine, for at least a portion of the subset of primary balanced SNV events, a corresponding variant allele frequency (VAF) based on the tumor sequencing data, thereby determining a VAF distribution for the subset of primary balanced SNV events; (e) identify one or more (e.g., a plurality) of VAF subpopulations within the VAF distribution for the subset of primary balanced SNV events, each VAF subpopulation corresponding to and representing VAFs of a distinct subpopulation of SNV events; (f) determine an estimated value of the primary copy number and / or a lower bound thereon based on the one or more VAF subpopulations, (g) store and / or- 24 - 1324629 IvlAttorney Docket No. 2013237-1500provide the estimated value of the primary copy number and / or the lower bound thereon for display and / or further processing.

[0092] In some aspects, the present disclosure provides a system (e.g., for detecting presence of a whole genome duplication event) comprising: a processor of a computing device; and memory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to: (a) obtain tumor sequencing data for a tumor sample obtained from a subject [e.g., the tumor sequencing data comprising a plurality of tumor reads, each tumor read representing a partial nucleotide sequence determined by sequencing the tumor sample]; (b) obtain a list of putative single nucleotide variant (SNV) events detected based on the tumor sequencing data, each putative SNV event representing one or more potential underlying mutations occurring in a tumor genome of the subject; (c) identify, from the list of putative SNV events, a subset of minimal balanced SNV events comprising SNV events determined to represent mutations that: (i) are determined to be balanced [e.g., are determined to have an equal number of copies of corresponding maternal and paternal alleles (e.g., via a balanced test and / or contour fdtering refinement); e.g., are determined to be located within balanced heterozygous segments] within the tumor genome; and (ii) are located within segments determined to have an absolute copy number equal to a same particular minimal copy number whose exact value is unknown but predicted to be lower than copy numbers of other balanced segments (e.g., the exact value of the minimal copy number is initially unknown, but determined to be smaller relative to other absolute copy numbers of balanced segments); (d) determine, for at least a portion of the subset of minimal balanced SNV events, a corresponding variant allele frequency (VAF) based on the tumor sequencing data, thereby determining a VAF distribution for the subset of minimal balanced SNV events; (e) identify a plurality of VAF subpopulations within the VAF distribution for the subset of minimal balanced SNV events, each VAF subpopulation corresponding to and representing VAFs of a distinct subpopulation of SNV events; (f) detect presence of a whole genome duplication event based at least in part on the plurality of VAF subpopulations; and (i) based at least in part on the detected presence of the WGD event, (i) determine a cancer status (e.g., a particular type of cancer; e.g., a particular stage of cancer) for the subject and / or (ii) select a therapy for the subject.- 25 - 1324629 IvlAttomey Docket No. 2013237-1500

[0093] In some aspects, the present disclosure provides a system (e.g., for detecting presence of a whole genome duplication event) comprising: a processor of a computing device; and memory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to: (a) obtain tumor sequencing data for a tumor sample obtained from a subject [e.g., the tumor sequencing data comprising a plurality of tumor reads, each tumor read representing a partial nucleotide sequence determined by sequencing the tumor sample]; (b) obtain a list of putative single nucleotide variant (SNV) events detected based on the tumor sequencing data, each putative SNV event representing one or more potential underlying mutations occurring in a tumor genome of the subject; (c) identify, from the list of putative SNV events, a subset of primary balanced SNV events comprising SNV events determined to represent mutations that: (i) are determined to be balanced [e.g., are determined to have an equal number of copies of corresponding maternal and paternal alleles (e.g., via a balanced test and / or contour filtering refinement); e.g., are determined to be located within balanced heterozygous segments] within the tumor genome; and (ii) are located within segments determined to have an absolute copy number equal to a same particular primary copy number whose exact value is unknown but predicted to be a most frequently occurring even absolute copy number [(e.g., the exact value of the primary copy number is initially unknown, but determined to be even and to be more frequently occurring than other even copy numbers (e.g., determined to be the most commonly occurring absolute copy number within a set of balanced heterozygous segments)]; (d) determine, for at least a portion of the subset of primary balanced SNV events, a corresponding variant allele frequency (VAF) based on the tumor sequencing data, thereby determining a VAF distribution for the subset of primary balanced SNV events; (e) identify one or more (e.g., a plurality) of VAF subpopulations within the VAF distribution for the subset of primary balanced SNV events, each VAF subpopulation corresponding to and representing VAFs of a distinct subpopulation of SNV events; (f) detect presence of a whole genome duplication event based at least in part on the plurality of VAF subpopulations; and (g) based at least in part on the detected presence of the WGD event, (i) determine a cancer status (e.g., a particular type of cancer; e.g., a particular stage of cancer) for the subject and / or (ii) select a therapy for the subject.

[0094] In some aspects, the present disclosure provides a method of producing an immunotherapy construct (e.g., a cancer vaccine) for a subject, the method comprising: detecting - 26 - 1324629 IvlAttorney Docket No. 2013237-1500a plurality of candidate mutations (e.g., somatic mutations; e.g., non-synonymous somatic mutations) in tumor cells of from the subject using a method or system disclosed herein; and synthesizing, as the immunotherapy construct, a polyribonucleotide encoding a plurality of neoepitopes, each encoded neoepitope corresponding to a candidate mutation of the plurality of candidate mutations.

[0095] In some aspects, the present disclosure provides a method comprising determining a subset of T-cells within a population of T-cells that are capable of specifically binding a plurality of complexes, wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, and wherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer-specific mutations detected using the method or system disclosed herein; and enriching for (e.g., expanding) the subset of T-cells that are capable of specifically binding the plurality of complexes.

[0096] In some aspects, the present disclosure provides a method comprising administering to a subject a population of T-cells, wherein the population of T cells are capable of specifically binding a plurality of complexes, wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, wherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer-specific mutations detected using a method or system disclosed herein, and wherein the genome of at least some of the subject’s cells comprises a subset (e.g., all) of the cancer-specific mutations.

[0097] In some aspects, the present disclosure provides a method comprising determining a subset of TILs within a population of TILs that are capable of specifically binding a plurality of complexes, wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, and wherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer- specific mutations detected using a method or system disclosed - 27 - 1324629 IvlAttorney Docket No. 2013237-1500herein; and enriching for (e.g., expanding) the subset of TILs that are capable of specifically binding the plurality of complexes.

[0098] In some aspects, the present disclosure provides a method comprising administering to a subject a population of TILs, wherein the population of TILs are capable of specifically binding a plurality of complexes, wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancerspecific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, wherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer- specific mutations detected using a method or system disclosed herein, and wherein the genome of at least some of the subject’s cells comprises a subset (e.g., all) of the cancer-specific mutations.

[0099] In some aspects, the present disclosures provides a method comprising obtaining a tumor sample from the subject and using the tumor sample to detect the plurality of candidate mutations (e.g., by sequencing the tumor sample).

[0100] In some embodiments, provided methods (e.g., further) comprise obtaining a normal sample from the subject and using the normal sample (e.g., together with the tumor sample) to detect the plurality of cancer mutations (e.g., by sequencing the normal sample).

[0101] In some embodiments, provided methods comprise sequencing the tumor and / or normal sample (e.g., in replicates).

[0102] In some aspects, the present disclosures provide a pharmaceutical composition comprising a polyribonucleotide encoding (e.g., a polyepitopic vaccine construct comprising) a plurality of neoepitopes, wherein at least a portion of the plurality of neoepitopes are individualized (e.g., patient- specific) neoepitopes corresponding to (e.g., encoded by a nucleotide sequence comprising) a somatic mutation from a population of candidate mutations detected in a tumor sample of a subject and selected for inclusion in the pharmaceutical composition using a method or system disclosed herein (e.g., each particular neoepitope encoded by a nucleotide sequence comprising one or more candidate mutation detected using a method or system disclosed herein).- 28 - 1324629 IvlAttomey Docket No. 2013237-1500

[0103] In some aspects, the present disclosures provide an individualized pharmaceutical composition (e.g., associated with and / or intended for administration to a particular subject) comprising a polyribonucleotide encoding (e.g., a polyepitopic vaccine construct comprising) a plurality of neoepitopes for a particular subject, wherein at least a portion of the neoepitopes are individualized (e.g., patient- specific) neoepitopes corresponding to (e.g., encoded by a nucleotide sequence comprising) a somatic mutation from a population of candidate mutations detected in a tumor sample from the particular subject and selected for inclusion in the pharmaceutical composition using a method or system disclosed herein.

[0104] In some aspects, the present disclosures provide a population of T-cells, wherein the population of T-cells is capable of specifically binding to a plurality of complexes, wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, and wherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancerspecific mutations using a method or system disclosed herein.

[0105] In some aspects, the present disclosures provide a T-cell capable of specifically binding to a complex, wherein the complex comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, wherein the one or more cancer-specific mutations are detected using a method or system disclosed herein.

[0106] In some aspects, the present disclosures provide a population of T-cells, wherein the population of T-cells is capable of specifically binding to a plurality of complexes, wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, and wherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancerspecific mutations detected using the method or system disclosed herein.

[0107] In some aspects, the present disclosures provide a T-cell capable of specifically binding to a complex, wherein the complex comprises (i) a portion of a neoepitope encoded by a - 29 - 1324629 IvlAttorney Docket No. 2013237-1500nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, wherein the one or more cancer-specific mutations are within the selected subset of candidate mutations detected using a method or system disclosed herein.

[0108] In some aspects, the present disclosures provide a T-cell receptor capable of specifically binding to a complex, wherein the complex comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, wherein the one or more cancer-specific mutations are within the selected subset of candidate mutations detected using a method or system disclosed herein.

[0109] In some aspects, the present disclosures provide a chimeric antigen receptor (CAR) capable of specifically binding to a complex, wherein the complex comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer-specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, wherein the one or more cancer- specific mutations are within the selected subset of candidate mutations detected using a method or system disclosed herein.

[0110] Features of embodiments described with respect to one aspect of the invention may be applied with respect to another aspect of the invention.BRIEF DESCRIPTION OF THE DRAWING

[0111] The foregoing and other objects, aspects, features, and advantages of the present disclosure will become more apparent and better understood by referring to the following description taken in conjunction with the accompanying drawings, in which:

[0112] FIG. 1 is a block flow diagram showing an example process for creating a personalized cancer immunotherapy, according to an illustrative embodiment.

[0113] FIG. 2 is a block flow diagram showing an example process for obtaining sequencing data from tumor and / or normal sample(s), according to an illustrative embodiment.- 30 - 1324629 IvlAttorney Docket No. 2013237-1500

[0114] FIG. 3 is a schematic showing a model of a normal genome and a tumor genome of a subject, according to an illustrative embodiment.

[0115] FIG. 4 is a schematic of a tumor sample, according to an illustrative embodiment.

[0116] FIG. 5A is a diagram illustrating interrelation between parameters and / or measurements characterizing tumor sample properties, according to an illustrative embodiment.

[0117] FIG. 5B is a diagram showing possible purity values for different copy numbers and zygosities, determined for a VAF of 0.2. Highlighted rows in the diagram correspond to balanced SNVs.

[0118] FIG. 6 is a block flow diagram of an example tumor deconvolution process for determining a tumor sample purity estimate, according to an illustrative embodiment.

[0119] FIG. 7 is a block flow diagram of an example process for identifying and / or selecting particular subsets of segments, according to an illustrative embodiment.

[0120] FIG. 8 is a schematic illustrating certain categories of segments within a genome (e.g., a tumor genome and / or a normal genome) and approaches for their identification, according to an illustrative embodiment.

[0121] FIG. 9 is a schematic illustrating certain categories of segments within a genome (e.g., a tumor genome and / or a normal genome) and approaches for their identification, according to an illustrative embodiment.

[0122] FIG. 10 is a schematic with an arrangement of example plots illustrating an example tumor deconvolution workflow, including how estimates of tumor sample properties, such as purity, absolute copy number, and allele specific copy number assignments, can be used to detect and characterize cancer-specific mutations, such as SNV events, according to an illustrative embodiment.

[0123] FIG. 11 is a block flow diagram of an example process for using multiple tumor models for obtaining purity estimates and / or estimated purity bounds, according to an illustrative embodiment.

[0124] FIG. 12 is a block flow diagram showing an exemplary SNV-based purity estimation process including quality control steps, according to an illustrative embodiment.- 31 - 1324629 IvlAttomey Docket No. 2013237-1500

[0125] FIG. 13A is a graph showing an exemplary scenario for a distribution of tumor over normal segment read count ratios for balanced heterozygous segments (e.g., as distributed a 1D GMM model as described herein) having two distinct components (a minimal component which is distinct from a primary component), according to an illustrative embodiment.

[0126] FIG. 13B is a graph showing an exemplary scenario for a distribution of tumor over normal segment read count ratios for balanced heterozygous segments (e.g., as distributed according to a ID GMM model as described herein) having a single component wherein the minimal component is the primary component, according to an illustrative embodiment.

[0127] FIG. 13C is a graph showing an exemplary scenario for a distribution of tumor over normal segment read count ratios for balanced heterozygous segments (e.g., as distributed according to a ID GMM model as described herein) having two distinct components wherein the component with the lower tumor / normal segment read count ratio (minimal component) is also the primary component, according to an illustrative embodiment.

[0128] FIG. 13D is a graph showing an exemplary scenario for a distribution of tumor over normal segment read count ratios for balanced heterozygous segments (e.g., as distributed according to a ID GMM model as described herein) having three distinct components wherein the minimal component is distinct from the other components (including the primary component), according to an illustrative embodiment.

[0129] FIG. 14 is a block flow diagram showing an exemplary process for determining putative minimal balanced SNV events and putative primary balanced SNV events, according to an illustrative embodiment.

[0130] FIG. 15 is a block flow diagram showing an exemplary process for estimating an optimal purity for mutation detection comprising SNV-based purity estimation and CNV-based purity estimation, according to an illustrative embodiment.

[0131] FIG. 16 is a block flow diagram showing a further exemplary process for estimating an optimal purity for mutation detection comprising SNV-based purity estimation and CNV-based purity estimation, according to an illustrative embodiment.

[0132] FIG. 17 is a block flow diagram showing a decision tree for selecting between various purity estimation results, according to an illustrative embodiment.- 32 - 1324629 IvlAttorney Docket No. 2013237-1500

[0133] FIG. 18A is a block flow diagram of an example process for analyzing VAF distributions of SNV subsets, according to an illustrative embodiment.

[0134] FIG. 18B is a block flow diagram of an example process for analyzing VAF distributions of SNV subsets, according to an illustrative embodiment.

[0135] FIG. 19A is a graph showing an exemplary scenario for a distribution of VAFs for putative minimal balanced SNV events across segments wherein a minimal copy number is equal to 2, according to an illustrative embodiment.

[0136] FIG. 19B is a graph showing an exemplary scenario for a distribution of VAFs for putative minimal balanced SNV events across segments wherein a minimal copy number is equal to 4, according to an illustrative embodiment.

[0137] FIG. 19C is a graph showing an exemplary scenario for a distribution of VAFs for putative minimal balanced SNV events across segments wherein a minimal copy number is equal to 6, according to an illustrative embodiment.

[0138] FIG. 19D is a graph showing an exemplary scenario for a distribution of VAFs for putative minimal balanced SNV events across segments wherein a minimal copy number is equal to 8, according to an illustrative embodiment.

[0139] FIG. 20 is a graph showing expected values of VAF and purity for SNV events with varying copy numbers and zygosities, according to an illustrative embodiment.

[0140] FIG. 21A is a graph showing an exemplary scenario for a distribution of VAFs for putative minimal balanced SNV events with a minimal copy number of 2, comprising a main subpopulation of minimal balanced SNV events, according to an illustrative embodiment.

[0141] FIG. 21B is a graph showing an exemplary scenario for a distribution of VAFs for putative minimal balanced SNV events with a minimal copy number of 2, comprising a main and a subclonal subpopulation of minimal balanced SNV events, according to an illustrative embodiment.

[0142] FIG. 21C is a pair of graphs showing exemplary scenarios for a distribution of VAFs for putative minimal balanced SNV events with a minimal copy number of 2, comprising- 33 - 1324629 IvlAttomey Docket No. 2013237-1500an outlier and a main subpopulation of minimal balanced SNV events, according to an illustrative embodiment.

[0143] FIG. 21D is a graph showing an exemplary scenario for a distribution of VAFs for putative minimal balanced SNV events with a minimal copy number of 2, comprising an outlier, a main, and a subclonal subpopulation of minimal balanced SNV events, according to an illustrative embodiment.

[0144] FIG. 21E is a graph showing an exemplary scenario for a distribution of VAFs for putative minimal balanced SNV events with a minimal copy number of 4, comprising harmonic outlier and main subpopulations of minimal balanced SNV events, according to an illustrative embodiment.

[0145] FIG. 21F is a graph showing an exemplary scenario for a distribution of VAFs for putative minimal balanced SNV events with a minimal copy number of 4, comprising harmonic main and subclonal subpopulations of minimal balanced SNV events, according to an illustrative embodiment.

[0146] FIG. 21G is a graph showing an exemplary scenario for a distribution of VAFs for putative minimal balanced SNV events with a minimal copy number of 4, comprising harmonic outlier and main subpopulations and a subclonal subpopulation of minimal balanced SNV events, according to an illustrative embodiment.

[0147] FIG. 21H is a graph showing an exemplary scenario for a distribution of VAFs for putative minimal balanced SNV events with a minimal copy number of 4, comprising harmonic main and subclonal subpopulations of, as well as additional subclonal, minimal balanced SNV events, according to an illustrative embodiment.

[0148] FIG. 21I is a graph showing an exemplary scenario for a distribution of VAFs for putative minimal balanced SNV events with a minimal copy number of 4, comprising harmonic outlier and clonal subpopulations as well as a main subpopulation of minimal balanced SNV events, according to an illustrative embodiment.

[0149] FIG. 21J is a graph showing an exemplary scenario for a distribution of VAFs for putative minimal balanced SNV events with a minimal copy number of 4, comprising a main subpopulation of minimal balanced SNV events, according to an illustrative embodiment.- 34 - 1324629 IvlAttorney Docket No. 2013237-1500

[0150] FIG. 21K is a graph showing an exemplary scenario for a distribution of VAFs for putative minimal balanced SNV events with a minimal copy number of 4, comprising a main and a subclonal subpopulation of minimal balanced SNV events, according to an illustrative embodiment.

[0151] FIG. 22 is a block flow diagram showing an exemplary process for VAF analysis and filtering of putative minimal balanced SNV events for improving the performance of SNV-based purity estimation and CNV-based purity estimation, according to an illustrative embodiment

[0152] FIG. 23 is a block flow diagram showing an exemplary process for VAF analysis and filtering of putative minimal balanced SNV events for improving the performance of CNV-based purity estimation, according to an illustrative embodiment.

[0153] FIG. 24 is a block flow diagram showing an exemplary process for determining subpopulations of putative minimal balanced SNV events and / or putative primary balanced SNV events, according to an illustrative embodiment.

[0154] FIG. 25 is a block flow diagram showing an exemplary process for determining population features of putative minimal balanced SNV events and / or putative primary balanced SNV events, according to an illustrative embodiment.

[0155] FIG. 26 is a block flow diagram showing an exemplary recursive process for refining a population of putative minimal balanced SNV events comprising filtering of low VAF outliers, according to an illustrative embodiment.

[0156] FIG. 27 is a block flow diagram showing an exemplary recursive process for refining a population of putative minimal balanced SNV events comprising filtering of high VAF outliers, according to an illustrative embodiment.

[0157] FIG. 28 is a schematic illustrating an example construct encoding selected neoantigens, according to an illustrative embodiment.

[0158] FIG. 29 is a block diagram of an exemplary cloud computing environment, used in certain embodiments.- 35 - 1324629 IvlAttomey Docket No. 2013237-1500

[0159] FIG. 30 is a block diagram of an example computing device and an example mobile computing device used in certain embodiments.

[0160] FIG. 31A is a state diagram for sequencing an allele for a homozygous genotype, according to an illustrative embodiment.

[0161] FIG. 31B is a state diagram for selecting and sequencing a major allele associated with a particular SNV, according to an illustrative embodiment.

[0162] FIG. 31C is a state diagram for selecting and sequencing a minor allele associated with a particular heterozygous SNP, according to an illustrative embodiment.

[0163] FIG. 32A is a graph plotting ML functions (e.g., empirical SNV likelihood and predicted analytical SNV likelihood) across values of p for a melanoma tumor sample, wherein the round marker corresponds to the global maximum of L, and the filled and empty red triangles are the predicted global and local maxima of the analytical log likelihood function, according to an illustrative embodiment.

[0164] FIG. 32B is a graph plotting generalized likelihood ratio statistics for a KL divergence metric (GLRT, also referred to as DKL herein) across values of p, wherein the round marker corresponds to the global maximum, for a melanoma tumor sample, according to an illustrative embodiment.

[0165] FIG. 32C is a graph plotting observed and predicted probability distribution function estimations for putative minimal balanced SNV events with and without VAF filtering of outliers, for a melanoma tumor sample, according to an illustrative embodiment.

[0166] FIG. 32D is a graph plotting log likelihood of putative minimal balanced SNV events being minimal balanced SNV events for a melanoma tumor sample, according to an illustrative embodiment.

[0167] FIG. 32E is a graph plotting predicted fractions of homozygous genotypes Pxx based on maximum likelihood (ML) estimation of genotype as a function of estimated normal contamination p for a melanoma tumor sample, according to an illustrative embodiment.- 36 - 1324629 IvlAttorney Docket No. 2013237-1500

[0168] FIG. 33 is a graph showing lower and upper boundary regions as a function of a true purity (x-axis) delineating solution regions for an analytical SNV-likelihood function (y-axis), according to an illustrative embodiment.

[0169] FIG. 34 is a graph showing a landscape of predicted maxima for an analytical SNV-likelihood function as a function of true purity (x-axis) and true fraction of homozygous SNVs (y-axis), according to an illustrative embodiment.

[0170] FIG. 35A is a graph showing log likelihood results from Monte Carlo simulations comparing an analytical likelihood function (blue) to an empirical SNV-likelihood function (red) for Region A of FIG.34, according to an illustrative embodiment.

[0171] FIG. 35B is a graph showing log likelihood results from Monte Carlo simulations comparing an analytical likelihood function (blue) to an empirical SNV-likelihood function (red) for Region B of FIG.34, according to an illustrative embodiment.

[0172] FIG. 35C is a graph showing log likelihood results from Monte Carlo simulations comparing an analytical likelihood function (blue) to an empirical SNV-likelihood function (red) for Region C of FIG.34, according to an illustrative embodiment.

[0173] FIG. 35D is a graph showing log likelihood results from Monte Carlo simulations comparing an analytical likelihood function (blue) to an empirical SNV-likelihood function (red) for Region D of FIG.34, according to an illustrative embodiment.

[0174] FIG. 35E is a graph showing log likelihood results from Monte Carlo simulations comparing an analytical likelihood function (blue) to an empirical SNV-likelihood function (red) for Region E of FIG.34, according to an illustrative embodiment.

[0175] FIG. 36 is a graph showing differentiation between a global maximum of an analytical SNV-likelihood function and a potential secondary maximum across a ( / / o, pxx) parameter space, according to an illustrative embodiment.

[0176] FIG. 37A is a graph plotting ML functions (empirical SNV log likelihood - blue and predicted analytical SNV log likelihood - red) having multiple peaks across values of for a tumor sampled at 5000x, according to an illustrative embodiment.- 37 - 1324629 IvlAttorney Docket No. 2013237-1500

[0177] FIG. 37B is a graph plotting KL divergence metrics (with outlier removal - blue and without outlier removal - red) having a sharp global maximum across values of p for a tumor sampled at 5000x, according to an illustrative embodiment.

[0178] FIG. 37C is a graph plotting ML functions (empirical SNV log likelihood - blue and predicted analytical SNV log likelihood - red) having multiple peaks across values of p for a tumor sampled at lOOx, according to an illustrative embodiment.

[0179] FIG. 37D is a graph plotting KL divergence metrics (with outlier removal - blue and without outlier removal - red) having a sharp global maximum across values of p for a tumor sampled at lOOx, according to an illustrative embodiment.

[0180] FIG. 38A is a graph showing results of a maximum likelihood SNV-based purity estimator for simulated tumor samples with various levels of normal contamination, p, and various fractions of homozygous SNVs, xx for number of SNVs L = 50 and depth of coverage C = 400x, according to an illustrative embodiment.

[0181] FIG. 38B is a graph showing results of a KL divergence SNV-based purity estimator for simulated tumor samples with various levels of normal contamination, p, and various fractions of homozygous SNVs, xx for number of SNVs L = 50 and depth of coverage C = 400x, according to an illustrative embodiment.

[0182] FIG. 38C is a graph showing results of a combined maximum likelihood and KL divergence SNV-based purity estimator for simulated tumor samples with various levels of normal contamination, p, and various fractions of homozygous SNVs, xx for number of SNVs L = 50 and depth of coverage C = 400x, according to an illustrative embodiment.

[0183] FIG. 38D is a graph showing results of a maximum likelihood SNV-based purity estimator for simulated tumor samples with various levels of normal contamination, p, and various fractions of homozygous SNVs, xx for number of SNVs L = 50 and depth of coverage C = 200x, according to an illustrative embodiment.

[0184] FIG. 38E is a graph showing results of a KL divergence SNV-based purity estimator for simulated tumor samples with various levels of normal contamination, p, and various fractions of homozygous SNVs, xx for number of SNVs L = 50 and depth of coverage C = 200x, according to an illustrative embodiment.- 38 - 1324629 IvlAttomey Docket No. 2013237-1500

[0185] FIG. 38F is a graph showing results of a combined maximum likelihood and KL divergence SNV-based purity estimator for simulated tumor samples with various levels of normal contamination, p, and various fractions of homozygous SNVs, xx for number of SNVs L = 50 and depth of coverage C = 200x, according to an illustrative embodiment.

[0186] FIG. 38G is a graph showing results of a maximum likelihood SNV-based purity estimator for simulated tumor samples with various levels of normal contamination, p, and various fractions of homozygous SNVs, xx for number of SNVs L = 20 and depth of coverage C = 200x, according to an illustrative embodiment.

[0187] FIG. 38H is a graph showing results of a KL divergence SNV-based purity estimator for simulated tumor samples with various levels of normal contamination, p, and various fractions of homozygous SNVs, xx for number of SNVs L = 20 and depth of coverage C = 200x, according to an illustrative embodiment.

[0188] FIG. 38I is a graph showing results of a combined maximum likelihood and KL divergence SNV-based purity estimator for simulated tumor samples with various levels of normal contamination, p, and various fractions of homozygous SNVs, xx for number of SNVs L = 20 and depth of coverage C = 200x, according to an illustrative embodiment.

[0189] FIG. 39A is a series of plots showing results for SNV-based purity estimation for a melanoma sample, according to an illustrative embodiment.

[0190] FIG. 39B is a CNV cluster plot showing results for CNV-based purity estimation for a melanoma sample, according to an illustrative embodiment.

[0191] FIG. 40A is a histogram showing estimated cellularities p for balanced diploid regions of a genome of a melanoma sample, according to an illustrative embodiment.

[0192] FIG. 40B is a histogram showing estimated cellularities p for balanced diploid regions of a genome of an ovarian tumor sample, according to an illustrative embodiment.

[0193] FIG. 40C is a block flow diagram of an example process for detecting putative mutations, according to an illustrative embodiment.

[0194] FIG. 40D is a graph providing results comparing an example mutation detection process with a binomial classifier, according to an illustrative embodiment.- 39 - 1324629 IvlAttomey Docket No. 2013237-1500

[0195] FIG. 40E is a graph providing results comparing an example mutation detection process with a binomial classifier, according to an illustrative embodiment.

[0196] FIG. 40F is a graph providing results comparing an example mutation detection process with a binomial classifier, according to an illustrative embodiment.

[0001] FIG. 41A is a graph plotting a VAF distribution for putative minimal balanced SNV events associated with different genomic sites for a first exemplary ovarian tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0002] FIG. 41B is a graph plotting observed and predicted probability distribution function estimations for putative minimal balanced SNV events with and without VAF filtering of outliers, for a first exemplary ovarian tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0003] FIG. 41C is a graph plotting log likelihood of putative minimal balanced SNV events being minimal balanced SNV events for a first exemplary ovarian tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0004] FIG. 41D is a graph plotting predicted fractions of homozygous genotypes Pxx based on maximum likelihood (ML) estimation of genotype as a function of estimated normal contamination,zz, for a first exemplary ovarian tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0005] FIG. 41E is a graph plotting generalized likelihood ratio statistics for a KL divergence metric (GLRT, also referred to as DKL herein) across values of,zz. wherein the round marker corresponds to the global maximum, for a first exemplary ovarian tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0006] FIG. 41F is a graph plotting ML functions (empirical SNV likelihood and predicted analytical SNV likelihood) across values of,zz. for a first exemplary ovarian tumor sample with a minimal copy number of 2, wherein the round marker corresponds to the global maximum of L, and the filled and empty red triangles are the predicted global and local maxima of the analytical log likelihood function, according to an illustrative embodiment.- 40 - 1324629 IvlAttorney Docket No. 2013237-1500

[0007] FIG. 41G is a CNV cluster diagram, plotting tumor / normal segment count ratios for heterozygous SNPs based on their allele frequencies for a first exemplary ovarian tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0008] FIG. 41H is a graph plotting a 1D-GMM of balanced heterozygous segments for a first exemplary ovarian tumor sample with a minimal (and primary) copy number of 2, according to an illustrative embodiment.

[0009] FIG. 411 is a graph plotting likelihood of,zz for all four scaling solutions with an SNV-based estimate of,zz superimposed (triangle) for a first exemplary ovarian tumor sample, according to an illustrative embodiment.

[0010] FIG. 41J shows an output report for a first exemplary ovarian tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0011] FIG. 41K shows a further output report for a first exemplary ovarian tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0012] FIG. 42A is a graph plotting a VAF distribution for putative minimal balanced SNV events associated with different genomic sites for a second exemplary ovarian tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0013] FIG. 42B is a graph plotting observed and predicted probability distribution function estimations for putative minimal balanced SNV events with and without VAF filtering of outliers for a second exemplary ovarian tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0014] FIG. 42C is a graph plotting log likelihood of putative minimal balanced SNV events being minimal balanced SNV events for a second exemplary ovarian tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0015] FIG. 42D is a graph plotting predicted fractions of homozygous genotypes Pxx based on maximum likelihood (ML) estimation of genotype as a function of estimated normal contamination,zz, for a second exemplary ovarian tumor sample with a minimal copy number of 2, according to an illustrative embodiment.- 41 - 1324629 IvlAttomey Docket No. 2013237-1500

[0016] FIG. 42E is a graph plotting generalized likelihood ratio statistics for a KL divergence metric (GLRT, also referred to as DKL herein) across values of,zz. wherein the round marker corresponds to the global maximum, for a second exemplary ovarian tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0017] FIG. 42F is a graph plotting ML functions (e.g., empirical SNV likelihood and predicted analytical SNV likelihood) across values of,zz. for a second exemplary ovarian tumor sample with a minimal copy number of 2, wherein the round marker corresponds to the global maximum of L, and the filled and empty red triangles are the predicted global and local maxima of the analytical log likelihood function, according to an illustrative embodiment.

[0018] FIG. 42G is a CNV cluster diagram, plotting tumor / normal segment count ratios for heterozygous SNPs based on their allele frequencies for a second exemplary ovarian tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0019] FIG. 42H is a graph plotting a 1D-GMM of balanced heterozygous segments for a second exemplary ovarian tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0020] FIG. 421 is a graph plotting likelihood of,zz for all four scaling solutions with an SNV-based estimate of,zz superimposed (triangle) for a second exemplary ovarian tumor sample, according to an illustrative embodiment.

[0021] FIG. 42J shows an output report for a second exemplary ovarian tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0022] FIG. 42K shows a further output report for a second exemplary ovarian tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0023] FIG. 43A is a graph plotting a VAF distribution for putative minimal balanced SNV events associated with different genomic sites for a third exemplary PBMC tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0024] FIG. 43B is a graph plotting observed and predicted probability distribution function estimations for putative minimal balanced SNV events with and without VAF filtering of outliers for a third exemplary PBMC tumor sample with a minimal copy number of 2, according to an illustrative embodiment.- 42 - 1324629 IvlAttorney Docket No. 2013237-1500

[0025] FIG. 43C is a graph plotting log likelihood of putative minimal balanced SNV events being minimal balanced SNV events for a third exemplary PBMC tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0026] FIG. 43D is a graph plotting predicted fractions of homozygous genotypes Pxx based on maximum likelihood (ML) estimation of genotype as a function of estimated normal contamination,zz, for a third exemplary PBMC tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0027] FIG. 43E is a graph plotting generalized likelihood ratio statistics for a KL divergence metric (GLRT. also referred to as DKL herein) across values of,zz. wherein the round marker corresponds to the global maximum, for a third exemplary PBMC tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0028] FIG. 43F is a graph plotting ML functions (empirical SNV likelihood and predicted analytical SNV likelihood) across values of,zz. for a third exemplary PBMC tumor sample with a minimal copy number of 2, wherein the round marker corresponds to the global maximum of L, and the filled and empty red triangles are the predicted global and local maxima of the analytical log likelihood function, according to an illustrative embodiment.

[0029] FIG. 43G is a graph plotting a 1D-GMM of balanced heterozygous segments for a third exemplary PBMC tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0030] FIG. 43H is a graph plotting likelihood of,zz for all four scaling solutions with an SNV-based estimate of,zz superimposed (triangle) for a third exemplary PBMC tumor sample, according to an illustrative embodiment.

[0031] FIG. 431 shows an output report for a third exemplary PBMC tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0032] FIG. 43J shows a further output report for a third exemplary PBMC tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0033] FIG. 44A is a graph plotting a VAF distribution for putative minimal balanced SNV events associated with different genomic sites for a fourth exemplary tumor sample with a minimal copy number of 2, according to an illustrative embodiment.- 43 - 1324629 IvlAttomey Docket No. 2013237-1500

[0034] FIG. 44B is a graph plotting observed and predicted probability distribution function estimations for putative minimal balanced SNV events with and without VAF filtering of outliers for a fourth exemplary tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0035] FIG. 44C is a graph plotting log likelihood of putative minimal balanced SNV events being minimal balanced SNV events for a fourth exemplary tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0036] FIG. 44D is a graph plotting predicted fractions of homozygous genotypes Pxx based on maximum likelihood (ML) estimation of genotype as a function of estimated normal contamination,zz, for a fourth exemplary tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0037] FIG. 44E is a graph plotting generalized likelihood ratio statistics for a KL divergence metric (GLRT, also referred to as DKL herein) across values of,zz. wherein the round marker corresponds to the global maximum, for a fourth exemplary tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0038] FIG. 44F is a graph plotting ML functions (empirical SNV likelihood and predicted analytical SNV likelihood) across values of,zz. for a fourth exemplary tumor sample with a minimal copy number of 2, wherein the round marker corresponds to the global maximum of L, and the filled and empty red triangles are the predicted global and local maxima of the analytical log likelihood function, according to an illustrative embodiment.

[0039] FIG. 44G is a graph plotting a 1D-GMM of balanced heterozygous segments for a fourth exemplary tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0040] FIG. 44H is a graph plotting likelihood of,zz for all four scaling solutions with an SNV-based estimate of,zz superimposed (triangle) for a fourth exemplary tumor sample, according to an illustrative embodiment.

[0041] FIG. 441 shows an output report for a fourth exemplary tumor sample with a minimal copy number of 2, according to an illustrative embodiment.- 44 - 1324629 IvlAttorney Docket No. 2013237-1500

[0042] FIG. 44J shows a further output report for a fourth exemplary tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0043] FIG. 45A is a graph plotting a VAF distribution for putative minimal balanced SNV events associated with different genomic sites for a fifth exemplary glioblastoma (GBM) tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0044] FIG. 45B is a graph plotting observed and predicted probability distribution function estimations for putative minimal balanced SNV events with and without VAF filtering of outliers for a fifth exemplary GBM tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0045] FIG. 45C is a graph plotting log likelihood of putative minimal balanced SNV events being minimal balanced SNV events for a fifth exemplary GBM tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0046] FIG. 45D is a graph plotting predicted fractions of homozygous genotypes Pxx based on maximum likelihood (ML) estimation of genotype as a function of estimated normal contamination,zz, for a fifth exemplary GBM tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0047] FIG. 45E is a graph plotting generalized likelihood ratio statistics for a KL divergence metric (GLRT. also referred to as DKL herein) across values of,zz. wherein the round marker corresponds to the global maximum, for a fifth exemplary GBM tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0048] FIG. 45F is a graph plotting ML functions (empirical SNV likelihood and predicted analytical SNV likelihood) across values of,zz. for a fifth exemplary GBM tumor sample with a minimal copy number of 2, wherein the round marker corresponds to the global maximum of L, and the filled and empty red triangles are the predicted global and local maxima of the analytical log likelihood function, according to an illustrative embodiment.

[0049] FIG. 45G is a CNV cluster diagram, plotting tumor / normal segment count ratios for heterozygous SNPs based on their allele frequencies for a fifth exemplary GBM tumor sample with a minimal copy number of 2, according to an illustrative embodiment.- 45 - 1324629 IvlAttomey Docket No. 2013237-1500

[0050] FIG. 45H is a graph plotting a 1D-GMM of balanced heterozygous segments for a fifth exemplary GBM tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0051] FIG. 451 is a graph plotting likelihood of,zz for all four scaling solutions with an SNV-based estimate of,zz superimposed (triangle) for a fifth exemplary GBM tumor sample, according to an illustrative embodiment.

[0052] FIG. 45J shows an output report for a fifth exemplary GBM tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0053] FIG. 45K shows a further output report for a fifth exemplary GBM tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0054] FIG. 46A is a graph plotting a VAF distribution for putative minimal balanced SNV events associated with different genomic sites for a sixth exemplary ovarian tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0055] FIG. 46B is a graph plotting a VAF distribution for putative minimal balanced SNV events associated with different genomic sites filtering outliers for a sixth exemplary ovarian tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0056] FIG. 46C is a graph plotting observed and predicted probability distribution function estimations for putative minimal balanced SNV events with and without VAF filtering of outliers for a sixth exemplary ovarian tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0057] FIG. 46D is a graph plotting log likelihood of putative minimal balanced SNV events being minimal balanced SNV events for a sixth exemplary ovarian tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0058] FIG. 46E is a graph plotting predicted fractions of homozygous genotypes Pxxbased on maximum likelihood (ML) estimation of genotype as a function of estimated normal contamination,zz, for a sixth exemplary ovarian tumor sample with a minimal copy number of 2, according to an illustrative embodiment.- 46 - 1324629 IvlAttorney Docket No. 2013237-1500

[0059] FIG. 46F is a graph plotting generalized likelihood ratio statistics for a KL divergence metric (GLRT, also referred to as DKL herein) across values of,zz. wherein the round marker corresponds to the global maximum, for a sixth exemplary ovarian tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0060] FIG. 46G is a graph plotting ML functions (e.g., empirical SNV likelihood and predicted analytical SNV likelihood) across values of,zz. for a sixth exemplary ovarian tumor sample with a minimal copy number of 2, wherein the round marker corresponds to the global maximum of L, and the filled and empty red triangles are the predicted global and local maxima of the analytical log likelihood function, according to an illustrative embodiment.

[0061] FIG. 46H is a CNV cluster diagram, plotting tumor / normal segment count ratios for heterozygous SNPs based on their allele frequencies for a sixth exemplary ovarian tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0062] FIG. 461 is a graph plotting a 1D-GMM of balanced heterozygous segments for a sixth exemplary ovarian tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0063] FIG. 46J is a graph plotting likelihood of,zz for all four scaling solutions with an SNV-based estimate of,zz superimposed (triangle) for a sixth exemplary ovarian tumor sample, according to an illustrative embodiment.

[0064] FIG. 46K shows an output report for a sixth exemplary ovarian tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0065] FIG. 46L shows a further output report for a sixth exemplary ovarian tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0066] FIG. 47A is a graph plotting a VAF distribution for putative minimal balanced SNV events associated with different genomic sites for a seventh exemplary GBM tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0067] FIG. 47B is a CNV cluster diagram, plotting tumor / normal segment count ratios for heterozygous SNPs based on their allele frequencies for a seventh exemplary GBM tumor sample with a minimal copy number of 2, according to an illustrative embodiment.- 47 - 1324629 IvlAttorney Docket No. 2013237-1500

[0068] FIG. 47C shows an output report for a seventh exemplary GBM tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0069] FIG. 47D shows a further output report for a seventh exemplary GBM tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0070] FIG. 47E is a graph plotting a VAF distribution for putative minimal balanced SNV events associated with different genomic sites for an eighth exemplary GBM tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0071] FIG. 47F is a graph plotting a VAF distribution for putative minimal balanced SNV events associated with different genomic sites filtering outliers for an eighth exemplary GBM tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0072] FIG. 47G is a CNV cluster diagram, plotting tumor / normal segment count ratios for heterozygous SNPs based on their allele frequencies for an eighth exemplary GBM tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0073] FIG. 47H shows an output report for an eighth exemplary GBM tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0074] FIG. 471 shows a further output report for an eighth exemplary GBM tumor sample with a minimal copy number of 2, according to an illustrative embodiment.

[0075] FIG. 48A is a graph plotting a VAF distribution for putative minimal balanced SNV events associated with different genomic sites for a ninth exemplary melanoma sample with a minimal copy number not equal to 2, according to an illustrative embodiment.

[0076] FIG. 48B is a graph plotting observed and predicted probability distribution function estimations for putative minimal balanced SNV events with and without VAF filtering of outliers for a ninth exemplary melanoma sample with a minimal copy number not equal to 2, according to an illustrative embodiment.

[0077] FIG. 48C is a graph plotting log likelihood of putative minimal balanced SNV events being minimal balanced SNV events for a ninth exemplary melanoma sample with a minimal copy number not equal to 2, according to an illustrative embodiment.- 48 - 1324629 IvlAttomey Docket No. 2013237-1500

[0078] FIG. 48D is a graph plotting predicted fractions of homozygous genotypes Px* based on maximum likelihood (ML) estimation of genotype as a function of estimated normal contamination,zz, for a ninth exemplary melanoma sample a minimal copy number not equal to 2, according to an illustrative embodiment.

[0079] FIG. 48E is a graph plotting generalized likelihood ratio statistics for a KL divergence metric (GLRT. also referred to as DKL herein) across values of,zz. wherein the round marker corresponds to the global maximum, for a ninth exemplary melanoma sample with a minimal copy number not equal to 2, according to an illustrative embodiment.

[0080] FIG. 48F is a graph plotting ML functions (e.g., empirical SNV likelihood and predicted analytical SNV likelihood) across values of,zz. for a ninth exemplary melanoma sample a minimal copy number not equal to 2, wherein the round marker corresponds to the global maximum of L, and the filled and empty red triangles are the predicted global and local maxima of the analytical log likelihood function, according to an illustrative embodiment.

[0081] FIG. 48G is a CNV cluster diagram, plotting tumor / normal segment count ratios for heterozygous SNPs based on their allele frequencies for a ninth exemplary melanoma sample with a minimal copy number not equal to 2, according to an illustrative embodiment.

[0082] FIG. 48H is a graph plotting a 1D-GMM of balanced heterozygous segments for a ninth exemplary melanoma sample with a minimal copy number not equal to 2, according to an illustrative embodiment.

[0083] FIG. 481 shows an output report for a ninth exemplary melanoma sample with a minimal copy number not equal to 2, according to an illustrative embodiment.

[0084] FIG. 48J shows a further output report for a ninth exemplary melanoma sample with a minimal copy number not equal to 2, according to an illustrative embodiment.

[0085] FIG. 49A is a graph plotting a VAF distribution for putative minimal balanced SNV events associated with different genomic sites for a tenth exemplary ovarian tumor sample with a minimal copy number not equal to 2, according to an illustrative embodiment.

[0086] FIG. 49B is a graph plotting observed and predicted probability distribution function estimations for putative minimal balanced SNV events with and without VAF filtering- 49 - 1324629 IvlAttomey Docket No. 2013237-1500of outliers for a tenth exemplary ovarian tumor sample with a minimal copy number not equal to 2, according to an illustrative embodiment.

[0087] FIG. 49C is a graph plotting log likelihood of putative minimal balanced SNV events being minimal balanced SNV events for a tenth exemplary ovarian tumor sample with a minimal copy number not equal to 2, according to an illustrative embodiment.

[0088] FIG. 49D is a graph plotting predicted fractions of homozygous genotypes Pxx based on maximum likelihood (ML) estimation of genotype as a function of estimated normal contamination,zz, for a tenth exemplary ovarian tumor sample a minimal copy number not equal to 2, according to an illustrative embodiment.

[0089] FIG. 49E is a graph plotting generalized likelihood ratio statistics for a KL divergence metric (GLRT, also referred to as DKL herein) across values of,zz. wherein the round marker corresponds to the global maximum, for a tenth exemplary ovarian tumor sample with a minimal copy number not equal to 2, according to an illustrative embodiment.

[0090] FIG. 49F is a graph plotting ML functions (e.g., empirical SNV likelihood and predicted analytical SNV likelihood) across values of,zz. for a tenth exemplary ovarian tumor sample with a minimal copy number not equal to 2, wherein the round marker corresponds to the global maximum of L, and the filled and empty red triangles are the predicted global and local maxima of the analytical log likelihood function, according to an illustrative embodiment.

[0091] FIG. 49G is a CNV cluster diagram, plotting tumor / normal segment count ratios for heterozygous SNPs based on their allele frequencies for a tenth exemplary ovarian tumor sample with a minimal copy number not equal to 2, according to an illustrative embodiment.

[0092] FIG. 49H is a graph plotting a 1D-GMM of balanced heterozygous segments for a tenth exemplary ovarian tumor sample with a minimal copy number not equal to 2, according to an illustrative embodiment.

[0093] FIG. 491 shows an output report for a tenth exemplary ovarian tumor sample with a minimal copy number not equal to 2, according to an illustrative embodiment.

[0094] FIG. 49J shows a further output report for a tenth exemplary ovarian tumor sample with a minimal copy number not equal to 2, according to an illustrative embodiment.- 50 - 1324629 IvlAttorney Docket No. 2013237-1500

[0095] FIG. 50A is a graph plotting a VAF distribution for putative minimal balanced SNV events associated with different genomic sites for an exemplary hepatocellular carcinoma cell line with a minimal copy number not equal to 2, according to an illustrative embodiment.

[0096] FIG. 50B is a graph plotting a 1D-GMM of balanced heterozygous segments for an exemplary hepatocellular carcinoma tumor cell line with a minimal copy number not equal to 2, according to an illustrative embodiment.

[0097] FIG. 50C is a CNV cluster diagram, plotting tumor / normal segment count ratios for heterozygous SNPs based on their allele frequencies for an exemplary hepatocellular carcinoma tumor cell line with a minimal copy number not equal to 2, according to an illustrative embodiment.

[0098] FIG. 50D shows an output report for an exemplary hepatocellular carcinoma cell tumor line with a minimal copy number not equal to 2, according to an illustrative embodiment.

[0099] FIG. 50E shows a further output report for an exemplary hepatocellular carcinoma tumor cell line with a minimal copy number not equal to 2, according to an illustrative embodiment.

[0100] FIG. 51A is a graph plotting a VAF distribution for putative minimal balanced SNV events associated with different genomic sites for a twelfth exemplary ovarian tumor sample with a minimal copy number not equal to 2, according to an illustrative embodiment.

[0101] FIG. 51B is a graph plotting a 1D-GMM of balanced heterozygous segments for a twelfth exemplary ovarian tumor sample with a minimal copy number not equal to 2, according to an illustrative embodiment.

[0102] FIG. 51C is a CNV cluster diagram, plotting tumor / normal segment count ratios for heterozygous SNPs based on their allele frequencies for a twelfth exemplary ovarian tumor sample with a minimal copy number not equal to 2, according to an illustrative embodiment.

[0103] FIG. 51D shows an output report for a twelfth exemplary ovarian tumor sample with a minimal copy number not equal to 2, according to an illustrative embodiment.

[0104] FIG. 51E shows a further output report for a twelfth exemplary ovarian tumor sample with a minimal copy number not equal to 2, according to an illustrative embodiment.- 51 - 1324629 IvlAttomey Docket No. 2013237-1500

[0105] FIG. 52A is a graph plotting a VAF distribution for putative minimal balanced SNV events associated with different genomic sites for a thirteenth exemplary ovarian tumor sample with a minimal copy number not equal to 2, according to an illustrative embodiment.

[0106] FIG. 52B is a graph plotting a 1D-GMM of balanced heterozygous segments for a thirteenth exemplary ovarian tumor sample with a minimal copy number not equal to 2, according to an illustrative embodiment.

[0107] FIG. 52C is a CNV cluster diagram, plotting tumor / normal segment count ratios for heterozygous SNPs based on their allele frequencies for a fourteenth exemplary ovarian tumor sample with a minimal copy number not equal to 2, according to an illustrative embodiment.

[0108] FIG. 52D is a CNV cluster diagram, plotting tumor / normal segment count ratios for heterozygous SNPs based on their allele frequencies for a fourteenth exemplary ovarian tumor sample with a minimal copy number not equal to 2, according to an illustrative embodiment.

[0109] FIG. 52E shows an output report for a fourteenth exemplary ovarian tumor sample with a minimal copy number not equal to 2, according to an illustrative embodiment.

[0110] FIG. 52F shows a further output report for a fourteenth exemplary ovarian tumor sample with a minimal copy number not equal to 2, according to an illustrative embodiment.

[0111] FIG. 53A is a graph plotting a VAF distribution for putative minimal balanced SNVs for a further exemplary tumor sample with a minimal copy number of 4, according to an illustrative embodiment.

[0112] FIG. 53B is a graph plotting a VAF distribution for putative minimal balanced SNVs for an additional exemplary tumor sample with a minimal copy number of 4 and harmonic subclonal population, according to an illustrative embodiment.

[0113] FIG. 53C is a graph plotting a VAF distribution for putative minimal balanced SNVs for a further exemplary tumor sample with a minimal copy number of 4 and harmonic subclonal population, according to an illustrative embodiment.- 52 - 1324629 IvlAttorney Docket No. 2013237-1500

[0114] FIG. 53D is a graph plotting a VAF distribution for putative minimal balanced SNVs for a further exemplary tumor sample with a minimal copy number of 4 and harmonic subclonal population, according to an illustrative embodiment.

[0115] The features and advantages of the present disclosure will become more apparent from the detailed description set forth below when taken in conjunction with the drawings, in which like reference characters identify corresponding elements throughout. In the drawings, like reference numbers generally indicate identical, functionally similar, and / or structurally similar elements.CERTAIN DEFINITIONS

[0116] About'. The term “about”, when used herein in reference to a value, refers to a value that is similar, in context to the referenced value. In general, those skilled in the art, familiar with the context, will appreciate the relevant degree of variance encompassed by “about” in that context. For example, in some embodiments, the term “about” may encompass a range of values that within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or less of the referred value.

[0117] Absolute Copy Number'. The term “absolute copy number” as used herein refers to a number of physical copies of a particular segment within a cell comprising a particular genome. For example, an absolute copy number of a segment in the normal genome can be defined as the number of physical copies of the given segment in a healthy cell. For example, an absolute copy number of a segment in the tumor genome can be defined as the number of physical copies of the given segment in a tumor cell. In certain embodiments, if only a part of a segment is amplified or deleted in a genome, then such a partial copy of the segment can either be counted as a copy of the segment or not counted as a copy of the segment. In certain embodiments, copies of the segment spanning less than 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10% or 5% of the segment length can be ignored.

[0118] Agent'. As used herein, the term “agent,” may refer to a physical entity. In some embodiments, an agent may be characterized by a particular feature and / or effect. For example, as used herein, the term “therapeutic agent” refers to a physical entity has a therapeutic effect- 53 - 1324629 IvlAttomey Docket No. 2013237-1500and / or elicits a desired biological and / or pharmacological effect. In some embodiments, an agent may be a compound, molecule, or entity of any chemical class including, for example, a small molecule, polypeptide, nucleic acid, saccharide, lipid, metal, or a combination or complex thereof. In some embodiments, part or all of an agent may be depicted herein as a chemical structure, or may be described using chemical nomenclature and / or with reference to general principles of organic chemistry, e.g., in accordance with the Periodic Table of Elements, CAS version, Handbook of Chemistry and Physics, 75th Ed; “Organic Chemistry”, Thomas Sorrell, University Science Books, Sausalito: 1999, and / or “March’s Advanced Organic Chemistry”, 5th Ed., Ed.: Smith, M. B. and March, J., John Wiley & Sons, New York: 2001, the entire contents of which are hereby incorporated by reference. Unless otherwise stated or clear from context, chemical structures depicted herein may be considered to reference or include one or more, or all, stereoisomeric (e.g., enantiomeric or diastereomeric) forms of the structure, and / or one or more, or all, geometric or conformational isomeric forms of the structure. For example, unless otherwise indicated or clear, both R and S configurations of a stereocenter may be contemplated in embodiments of the disclosure. In some embodiments, a compound may be described and / or utilized as a particular single stereochemical isomer; alternatively or additionally, in some embodiments, such a compound may be described and / or utilized as a combination e.g., a mixture) of one or more enantiomeric (e.g., diastereomeric) forms (e.g., as a racemic preparation). Analogously, in some embodiments, a single geometric isomer may be described and / or utilized; in some embodiments, a combination (e.g., a mixture) of geometric (or conformational) isomers may be described and / or utilized. Unless otherwise stated or clear from context, all tautomeric forms of provided compounds are within the scope of the disclosure. Still further, unless otherwise indicated or clear from context, in some embodiments, a particular chemical compound (e.g., as may be represented by a depicted chemical structure) may be described and / or utilized in an alternative isotopic form - i.e., in a form in which one or more atoms is isotopically altered (e.g., so that a hydrogen is replaced by deuterium or tritium, and / or a carbon is replaced by 13C- or 14C-. Thus, in some embodiments, a particular compound may be described and / or utilized as or in an isotopically enriched preparation.

[0119] Allele-Specific Copy Number: As used herein, the term “allele- specific copy number” refers to a number of physical copies of a particular allele within a cell comprising a particular genome. In certain embodiments, for example, a heterozygous segment comprises a - 54 - 1324629 IvlAttorney Docket No. 2013237-1500SNP. A cell comprising a genome with the heterozygous segment may comprise zero, one, or more copies of a first (e.g., maternal) allele and zero, one, or more copies of a second (e.g., paternal) allele. For example, a balanced heterozygous segment in a normal diploid genome may comprise distinguishable maternal and paternal alleles, each having an allele-specific copy number of one. In certain embodiments, a corresponding heterozygous segment within a tumor genome may not have zero, one, or more copies of the paternal and maternal alleles, e.g., as a result of copy number variation (CNV) events that may occur in cancer cells. For example, a loss of heterozygosity (LOH) may result in a deletion of copies of one allele (e.g., a maternal or paternal allele), such that an allele- specific copy number for the deleted allele is zero and, if a single copy of the other allele remains, its allele- specific copy number is one. In certain embodiments, duplication events may produce other allele-specific copy numbers, greater than one for one or both alleles.

[0120] Amino acid'. In its broadest sense, as used herein, the term “amino acid” refers to a compound and / or substance that can be, is, or has been incorporated into a polypeptide chain, e.g., through formation of one or more peptide bonds. In some embodiments, an amino acid has the general structure H2N-C(H)(R)-COOH. In some embodiments, an amino acid is a naturally-occurring amino acid. In some embodiments, an amino acid is a non-natural amino acid; in some embodiments, an amino acid is a D-amino acid; in some embodiments, an amino acid is an L-amino acid. “Standard amino acid” refers to any of the twenty standard L-amino acids commonly found in naturally occurring peptides. “Nonstandard amino acid” refers to any amino acid, other than the standard amino acids, regardless of whether it is prepared synthetically or obtained from a natural source. In some embodiments, an amino acid, including a carboxy- and / or amino-terminal amino acid in a polypeptide, can contain a structural modification as compared with the general structure above. For example, in some embodiments, an amino acid may be modified by methylation, amidation, acetylation, pegylation, glycosylation, phosphorylation, and / or substitution e.g., of the amino group, the carboxylic acid group, one or more protons, and / or the hydroxyl group) as compared with the general structure. In some embodiments, such modification may, for example, alter the circulating half-life of a polypeptide containing the modified amino acid as compared with one containing an otherwise identical unmodified amino acid. In some embodiments, such modification does not significantly alter a relevant activity of a polypeptide containing the modified amino acid, as - 55 - 1324629 IvlAttomey Docket No. 2013237-1500compared with one containing an otherwise identical unmodified amino acid. As will be clear from context, in some embodiments, the term “amino acid” may be used to refer to a free amino acid; in some embodiments it may be used to refer to an amino acid residue of a polypeptide.

[0121] Antigen -, term “antigen”, as used herein, refers to (i) an agent that elicits an immune response; and / or (ii) an agent that binds to a T cell receptor (e.g., when presented by an MHC molecule) or to an antibody. In some embodiments, an antigen elicits a humoral response (e.g., including production of antigen-specific antibodies); in some embodiments, an elicits a cellular response e.g., involving T-cells whose receptors specifically interact with the antigen). In some embodiments, and antigen binds to an antibody and may or may not induce a particular physiological response in an organism. In general, an antigen may be or include any chemical entity such as, for example, a small molecule, a nucleic acid, a polypeptide, a carbohydrate, a lipid, a polymer [in some embodiments other than a biologic polymer (e.g., other than a nucleic acid or amino acid polymer)] etc. In some embodiments, an antigen is or comprises a polypeptide. In some embodiments, an antigen is or comprises a glycan. Those of ordinary skill in the art will appreciate that, in general, an antigen may be provided in isolated or pure form, or alternatively may be provided in crude form (e.g., together with other materials, for example in an extract such as a cellular extract or other relatively crude preparation of an antigen-containing source). In some embodiments, antigens utilized in accordance with the present invention are provided in a crude form. In some embodiments, an antigen is a recombinant antigen.

[0122] Biologically clonal mutation(s) and biologically subclonal mutation(s): As used herein, the terms “biologically clonal” and “biologically subclonal” when used in reference to mutations, such as cancer mutations, are used to specify whether a particular mutations or group of mutations are physically clonal or subclonal. In certain embodiments, a biologically clonal mutation is a mutation that is present in all tumor cells of a tumor sample or biopsy. In certain embodiments, a biologically subclonal mutation is a mutation that is not present in all tumor cells of a tumor sample or biopsy. The use of the adjective “biologically” is used to make clear that the terms “biologically clonal” and “biologically subclonal” refer to the actual physical character of a given mutation, which may or may not be known. The terms “biologically clonal” and “biologically subclonal” thus contrast with the terms “clone type”, “clonal state”, “prevalent- 56 - 1324629 IvlAttorney Docket No. 2013237-1500subclone state”, and “minor subclone state”, described below, which refer to clonality classification states that are, e.g., labels, determined for (e.g., assigned to) a given mutation.

[0123] Cancer. The term “cancer” is used herein to generally refer to a disease or condition in which cells of a tissue of interest exhibit relatively abnormal, uncontrolled, and / or autonomous growth, so that they exhibit an aberrant growth phenotype characterized by a significant loss of control of cell proliferation. In some embodiments, cancer may comprise cells that are precancerous e.g., benign), malignant, pre-metastatic, metastatic, and / or non-metastatic. In some embodiments, cancer may be characterized by a solid tumor. In some embodiments, cancer may be characterized by a hematologic tumor. In general, examples of different types of cancers known in the art include, for example, triple negative breast cancer (TNBC), hematopoietic cancers including leukemias, lymphomas (Hodgkin’s and non-Hodgkin’s), myelomas and myeloproliferative disorders; sarcomas, melanomas, adenomas, carcinomas of solid tissue, squamous cell carcinomas of the mouth, throat, larynx, and lung, liver cancer, genitourinary cancers such as prostate, cervical, bladder, uterine, and endometrial cancer and renal cell carcinomas, bone cancer, pancreatic cancer, skin cancer, cutaneous or intraocular melanoma, cancer of the endocrine system, cancer of the thyroid gland, cancer of the parathyroid gland, head and neck cancers, ovarian cancer, breast cancer, glioblastomas, colorectal cancer, gastro-intestinal cancers and nervous system cancers, benign lesions such as papillomas, and the like.

[0124] Comparable'. As used herein, the term “comparable” refers to two or more agents, entities, situations, sets of conditions, etc., that may not be identical to one another but that are sufficiently similar to permit comparison therebetween so that one skilled in the art will appreciate that conclusions may reasonably be drawn based on differences or similarities observed. In some embodiments, comparable sets of conditions, circumstances, individuals, or populations are characterized by a plurality of substantially identical features and one or a small number of varied features. Those of ordinary skill in the art will understand, in context, what degree of identity is required in any given circumstance for two or more such agents, entities, situations, sets of conditions, etc., to be considered comparable. For example, those of ordinary skill in the art will appreciate that sets of circumstances, individuals, or populations are comparable to one another when characterized by a sufficient number and type of substantially- 57 - 1324629 IvlAttorney Docket No. 2013237-1500identical features to warrant a reasonable conclusion that differences in results obtained or phenomena observed under or with different sets of circumstances, individuals, or populations are caused by or indicative of the variation in those features that are varied.

[0125] Corresponding to: As used herein, the term “corresponding to” refers to a relationship between two or more entities. For example, the term “corresponding to” may be used to designate the position / identity of a structural element in a compound or composition relative to another compound or composition (e.g., to an appropriate reference compound or composition). For example, in some embodiments, a monomeric residue in a polymer (e.g., an amino acid residue in a polypeptide or a nucleic acid residue in a polynucleotide) may be identified as “corresponding to” a residue in an appropriate reference polymer. For example, those of ordinary skill will appreciate that, for purposes of simplicity, residues in a polypeptide are often designated using a canonical numbering system based on a reference related polypeptide, so that an amino acid “corresponding to” a residue at position 190, for example, need not actually be the 190th amino acid in a particular amino acid chain but rather corresponds to the residue found at 190 in the reference polypeptide; those of ordinary skill in the art readily appreciate how to identify “corresponding” amino acids. For example, those skilled in the art will be aware of various sequence alignment strategies, including software programs such as, for example, BLAST, CS-BLAST, CUSASW++, DIAMOND, FASTA, GGSEARCH / GLSEARCH, Genoogle, HMMER, HHpred / HHsearch, IDF, Infernal, KLAST, USEARCH, parasail, PSL BLAST, PSI-Search, ScalaBLAST, Sequilab, SAM, SSEARCH, SWAPHI, SWAPHLLS, SWIMM, or SWIPE that can be utilized, for example, to identify “corresponding” residues in polypeptides and / or nucleic acids in accordance with the present disclosure. Those of skill in the art will also appreciate that, in some instances, the term “corresponding to” may be used to describe an event or entity that shares a relevant similarity with another event or entity (e.g., an appropriate reference event or entity). To give but one example, a gene or protein in one organism may be described as “corresponding to” a gene or protein from another organism in order to indicate, in some embodiments, that it plays an analogous role or performs an analogous function and / or that it shows a particular degree of sequence identity or homology, or shares a particular characteristic sequence element.- 58 - 1324629 IvlAttomey Docket No. 2013237-1500

[0126] Encode. As used herein, the term “encode” or “encoding” refers to sequence information of a first molecule that guides production of a second molecule having a defined sequence of nucleotides (e.g., a polyribonucleotide) or a defined sequence of amino acids. For example, a DNA molecule can encode an RNA molecule (e.g., by a transcription process that includes a DNA-dependent RNA polymerase enzyme). An RNA molecule can encode a polypeptide e.g., by a translation process). Thus, a gene, a cDNA, or an RNA molecule encodes a polypeptide if transcription and translation of RNA corresponding to that gene produces the polypeptide in a cell or other biological system. In some embodiments, a coding region of a polyribonucleotide encoding a target antigen refers to a coding strand, the nucleotide sequence of which is identical to the polyribonucleotide sequence of such a target antigen. In some embodiments, a coding region of a polyribonucleotide encoding a target antigen refers to a noncoding strand of such a target antigen, which may be used as a template for transcription of a gene or cDNA.

[0127] Epitope'. As used herein, the term “epitope” refers to a moiety that is specifically recognized by an immune system (e.g., an immune system component) of a subject. For example, in some embodiments, an epitope may be a moiety that is specifically recognized by a T cell, a B cell, an immunoglobulin (e.g., antibody or receptor), immunoglobulin (e.g., antibody or receptor), binding component or an aptamer. In some embodiments, an epitope is comprised of a plurality of chemical atoms or groups on an antigen. In some embodiments, such chemical atoms or groups are surface-exposed when the antigen adopts a relevant three-dimensional conformation. In some embodiments, such chemical atoms or groups are physically near to each other in space when the antigen adopts such a conformation. In some embodiments, at least some such chemical atoms or groups are physically separated from one another when the antigen adopts an alternative conformation (e.g., is linearized).

[0128] Estimated tumor sample purity. As used herein, the term “estimated tumor sample purity” is used to refer to an estimate of tumor content of a tumor sample, such as a fraction, percentage etc. of cancer cells within a tumor sample and / or, equivalently, an estimate of normal contamination, such as a fraction, percentage, etc. of normal (e.g., healthy) cells within a tumor sample. It should be understood that, in certain embodiments, tumor samples are assumed to be comprised of tumor cells and normal cells, such that a fraction of tumor cells in a- 59 - 1324629 IvlAttomey Docket No. 2013237-1500tumor sample is equal to 1 minus a fraction of normal cells (e.g., 1 - p), estimates of tumor content and / or normal contamination equivalently measure an estimated tumor sample purity.

[0129] Expression-. As used herein, the term “expression” of a nucleic acid sequence refers to the generation of a gene product from the nucleic acid sequence. In some embodiments, a gene product can be a transcript, e.g., a polyribonucleotide as provided herein. In some embodiments, a gene product can be a polypeptide. In some embodiments, expression of a nucleic acid sequence involves one or more of the following: (1) production of an RNA template from a DNA sequence (e.g., by transcription); (2) processing of an RNA transcript (e.g., by splicing, editing, etc.); (3) translation of an RNA into a polypeptide or protein; and / or (4) post-translational modification of a polypeptide or protein.

[0130] Heterozygous Segment-. The term “heterozygous segment” as used here refers to a segment of a normal genome that comprises at least one heterozygous SNP, as well as any corresponding segments of a tumor genome or reference genome. In other words, a particular segment of a particular genome is defined as heterozygous or not according to whether the corresponding segment of a normal genome comprises a heterozygous SNP or not. For example, a reference genome may be partitioned into a plurality of segments, as described herein, in order to identify and define corresponding segments in a normal and tumor genome. Accordingly, if, for a given segment, the corresponding segment in the normal genome is determined to comprise a heterozygous SNP, then that segment is defined as a heterozygous segment. For purposes of determining heterozygous segments, a normal genome may be a reference genome obtained from a database, a normal reference determined and / or compiled based on one or more subject (e.g., a panel), determined by sequencing a particular subject (e.g., the same subject whose tumor is being sequenced).

[0131] Homology. As used herein, the term “homology” or “homolog” refers to the overall relatedness between polynucleotide molecules (e.g., DNA molecules and / or RNA molecules) and / or between polypeptide molecules. In some embodiments, polynucleotide molecules (e.g., DNA molecules and / or RNA molecules) and / or polypeptide molecules are considered to be “homologous” to one another if their sequences are at least 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical. In some embodiments, polynucleotide molecules (e.g., DNA molecules and / or RNA - 60 - 1324629 IvlAttorney Docket No. 2013237-1500molecules) and / or polypeptide molecules are considered to be “homologous” to one another if their sequences are at least 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% similar (e.g., containing residues with related chemical properties at corresponding positions). For example, as is well known by those of ordinary skill in the art, certain amino acids are typically classified as similar to one another as “hydrophobic” or “hydrophilic” amino acids, and / or as having “polar” or “non-polar” side chains. Substitution of one amino acid for another of the same type may often be considered a “homologous” substitution.

[0132] Identity. As used herein, the term “identity” refers to the overall relatedness between polynucleotide molecules (e.g., DNA molecules and / or RNA molecules) and / or between polypeptide molecules. In some embodiments, polynucleotide molecules (e.g., DNA molecules and / or RNA molecules) and / or between polypeptide molecules are considered to be “substantially identical” to one another if their sequences are at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical. Calculation of the percent identity of two nucleic acid or polypeptide sequences, for example, can be performed by aligning the two sequences for optimal comparison purposes (e.g., gaps can be introduced in one or both of a first and a second sequence for optimal alignment and non-identical sequences can be disregarded for comparison purposes). In certain embodiments, the length of a sequence aligned for comparison purposes is at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or substantially 100% of the length of a reference sequence. The nucleotides at corresponding positions are then compared. When a position in the first sequence is occupied by the same residue (e.g., nucleotide or amino acid) as the corresponding position in the second sequence, then the molecules are identical at that position. The percent identity between the two sequences is a function of the number of identical positions shared by the sequences, taking into account the number of gaps, and the length of each gap, which needs to be introduced for optimal alignment of the two sequences. The comparison of sequences and determination of percent identity between two sequences can be accomplished using a mathematical algorithm. For example, the percent identity between two nucleotide sequences can be determined using the algorithm of Meyers and Miller, 1989, which has been incorporated into the ALIGN program (version 2.0). In some exemplary embodiments, nucleic acid sequence comparisons made with the ALIGN - 61 - 1324629 IvlAttomey Docket No. 2013237-1500program use a PAM 120 weight residue table, a gap length penalty of 12 and a gap penalty of 4. The percent identity between two nucleotide sequences can, alternatively, be determined using the GAP program in the GCG software package using an NWSgapdna. CMP matrix.

[0133] Increased, Induced, or Reduced'. As used herein, these terms or grammatically comparable comparative terms, indicate values that are relative to a comparable reference measurement. For example, in some embodiments, an assessed value achieved with a provided composition (e.g., a pharmaceutical composition) may be “increased” relative to that obtained with a comparable reference composition. Alternatively or additionally, in some embodiments, an assessed value achieved in a subject may be “increased” relative to that obtained in the same subject under different conditions (e.g., prior to or after an event; or presence or absence of an event such as administration of a composition (e.g., a pharmaceutical composition) as described herein, or in a different, comparable subject (e.g., in a comparable subject that differs from the subject of interest in prior exposure to a condition, e.g., absence of administration of a composition (e.g., a pharmaceutical composition) as described herein.). In some embodiments, comparative terms refer to statistically relevant differences (e.g., that are of a prevalence and / or magnitude sufficient to achieve statistical relevance). Those skilled in the art will be aware, or will readily be able to determine, in a given context, a degree and / or prevalence of difference that is required or sufficient to achieve such statistical significance. In some embodiments, the term “reduced” or equivalent terms refers to a reduction in the level of an assessed value by at least 5%, at least 10%, at least 20%, at least 50%, at least 75% or higher, as compared to a comparable reference. In some embodiments, the term “reduced” or equivalent terms refers to a complete or essentially complete inhibition, i.e., a reduction to zero or essentially to zero. In some embodiments, the term “increased” or “induced” refers to an increase in the level of an assessed value by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 80%, at least 100%, at least 200%, at least 500%, or higher, as compared to a comparable reference.

[0134] Initial event'. As used herein, the term “initial event” refers to a portion of a tumor genome which may encode a mutation. In certain embodiments, initial events are detected in tumor sequencing data of a tumor sample obtained from a subject. In some embodiments, an initial event is any portion of (e.g., individual sites in, contiguous subsequences in) a tumor genome (e.g., such that a set or list of initial events may include all sites in a tumor genome). In- 62 - 1324629 IvlAttorney Docket No. 2013237-1500some embodiments, an initial event is a portion of a tumor genome associated with an observed statistical deviation from an expected read count in tumor sequencing data (e.g., wherein an expected read count is based on normal sequencing data which may be matched normal sequencing data). Given experimental and / or bioinformatic noise that may be present in sequenced samples, it is expected that among a plurality of initial events, any number of initial events may be determined to be false positives (e.g., may be determined to not reflect underlying mutations). In certain embodiments, an initial event refers to a variant portion of a tumor genome (e.g., identifying one or more sites within the tumor genome) determined, based on the sequencing data, to (e.g., potentially) have a different copy number and / or a different nucleotide sequence relative to a corresponding (e.g., aligned with the corresponding variant portion) normal reference. In certain embodiments, the different nucleotide sequence is or comprises an alternate allele (e.g., an individual nucleotide), different from an allele of the corresponding normal reference. In certain embodiments, the different nucleotide sequence is or comprises an insertion and / or a deletion (e.g., of one or more nucleotides) relative to the corresponding normal reference (e.g., an indel). In certain embodiments, the variant portion is or comprises a structural variation relative to the normal reference. In certain embodiments, an initial event is a potential point mutation at a particular site. In certain embodiments, an initial event is a potential indel. In certain embodiments, an initial event is a potential structural variation. In certain embodiments, an initial event is a potential copy number variation. In certain embodiments, a normal reference is obtained from a database (e.g., a hl9 reference genome). In certain embodiments, a normal reference is determined based on normal sequencing data obtained for a subject (e.g., by sequencing a normal sample, e.g., of healthy tissue and / or cells)].

[0135] In order. As used herein with reference to a polynucleotide or polyribonucleotide, “in order” refers to the order of features from 5' to 3' along the polynucleotide or polyribonucleotide. As used herein with reference to a polypeptide, “in order” refers to the order of features moving from the N-terminal-most of the features to the C-terminal-most of the features along the polypeptide. “In order” does not mean that no additional features can be present among the listed features. For example, if Features A, B, and C of a polynucleotide are described herein as being “in order, Feature A, Feature B, and Feature C,” this description does not exclude, e.g., Feature D being located between Features A and B.- 63 - 1324629 IvlAttomey Docket No. 2013237-1500

[0136] Linker. As used herein, the term “linker” refers to a portion of a polypeptide that connects different regions, portions, or antigens to one another.

[0137] Lipid'. As used herein, the terms “lipid” and “lipid-like material” are broadly defined as molecules which comprise one or more hydrophobic moieties or groups and optionally also one or more hydrophilic moieties or groups. Molecules comprising hydrophobic moieties and hydrophilic moieties are also typically denoted as amphiphiles.

[0138] Neoantigen'. As used herein, the term “neoantigen” refers to an antigen that is not present in a reference, such as a normal non-cancerous or germline cell, but is present in a cancer cell. In some embodiments, a neoantigen includes one or more mutations relative to a corresponding antigen present in a normal non-cancerous or germline cell.

[0139] Neoantigen epitope. As used herein, the term “neoantigen epitope” refers to an epitope that is not present in a reference, such as a normal non-cancerous or germline cell, but is present in a cancer cell.

[0140] Nucleic acid / Polynucleotide'. As used herein, the term “nucleic acid” refers to a polymer of at least 10 nucleotides or more. In some embodiments, a nucleic acid is or comprises DNA. In some embodiments, a nucleic acid is or comprises RNA. In some embodiments, a nucleic acid is or comprises peptide nucleic acid (PNA). In some embodiments, a nucleic acid is or comprises a single stranded nucleic acid. In some embodiments, a nucleic acid is or comprises a double- stranded nucleic acid. In some embodiments, a nucleic acid comprises both single and double- stranded portions. In some embodiments, a nucleic acid comprises a backbone that comprises one or more phosphodiester linkages. In some embodiments, a nucleic acid comprises a backbone that comprises both phosphodiester and non-phosphodiester linkages. For example, in some embodiments, a nucleic acid may comprise a backbone that comprises one or more phosphorothioate or 5'-N-phosphoramidite linkages and / or one or more peptide bonds, e.g., as in a “peptide nucleic acid”. In some embodiments, a nucleic acid comprises one or more, or all, natural residues (e.g., adenine, cytosine, deoxy adenosine, deoxycytidine, deoxy guanosine, deoxythymidine, guanine, thymine, uracil). In some embodiments, a nucleic acid comprises on or more, or all, non-natural residues. In some embodiments, a non-natural residue comprises a nucleoside analog (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3 -methyl adenosine, 5-methylcytidine, C-5 propynyl-cytidine, C-5 propynyl-uridine, 2- - 64 - 1324629 IvlAttorney Docket No. 2013237-1500aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyl-uridine, C5 -propynyl-cytidine, C5-methylcytidine, 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, 6-O-methylguanine, 2-thiocytidine, methylated bases, intercalated bases, and combinations thereof). In some embodiments, a non-natural residue comprises one or more modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose) as compared to those in natural residues. In some embodiments, a nucleic acid has a nucleotide sequence that encodes a functional gene product such as an RNA or polypeptide. In some embodiments, a nucleic acid has a nucleotide sequence that comprises one or more introns. In some embodiments, a nucleic acid may be prepared by isolation from a natural source, enzymatic synthesis (e.g., by polymerization based on a complementary template, e.g., in vivo or in vitro), reproduction in a recombinant cell or system, or chemical synthesis. In some embodiments, a nucleic acid is at least 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 1 10, 120, 130, 140, 150, 160, 170, 180, 190, 20, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500, 10,000, 10,500, 11,000, 11,500, 12,000, 12,500, 13,000, 13,500, 14,000, 14,500, 15,000, 15,500, 16,000, 16,500, 17,000, 17,500, 18,000, 18,500, 19,000, 19,500, or 20,000 or more residues or nucleotides long.

[0141] Ploidy. As used herein, the term “ploidy,” for example of a tumor genome, is used to refer to an average of absolute copy numbers of all segments (e.g., across an entire region of a tumor genome), weighted by the length of each segment. A ploidy of a region of a tumor genome can be defined as the average of absolute copy numbers of all segments in the region, weighted by the length of each segment.

[0142] Polypeptide. As used herein, the term “polypeptide” refers to a polymeric chain of amino acids. In some embodiments, a polypeptide has an amino acid sequence that occurs in nature. In some embodiments, a polypeptide has an amino acid sequence that does not occur in nature. In some embodiments, a polypeptide has an amino acid sequence that is engineered in that it is designed and / or produced through action of the hand of man. In some embodiments, a polypeptide may comprise or consist of natural amino acids, non-natural amino acids, or both. In some embodiments, a polypeptide may comprise or consist of only natural amino acids or only non-natural amino acids. In some embodiments, a polypeptide may comprise D-amino acids, L-- 65 - 1324629 IvlAttomey Docket No. 2013237-1500amino acids, or both. In some embodiments, a polypeptide may comprise only D-amino acids. In some embodiments, a polypeptide may comprise only L- amino acids. In some embodiments, a polypeptide may include one or more pendant groups or other modifications, e.g., modifying or attached to one or more amino acid side chains, at the polypeptide’s N-terminus, at the polypeptide’s C-terminus, or any combination thereof. In some embodiments, such pendant groups or modifications comprise acetylation, amidation, lipidation, methylation, pegylation, etc., including combinations thereof. In some embodiments, a polypeptide may be cyclic, and / or may comprise a cyclic portion. In some embodiments, a polypeptide is not cyclic and / or does not comprise any cyclic portion. In some embodiments, a polypeptide is linear. In some embodiments, a polypeptide may be or comprise a stapled polypeptide. In some embodiments, the term “polypeptide” may be appended to a name of a reference polypeptide, activity, or structure; in such instances it is used herein to refer to polypeptides that share the relevant activity or structure and thus can be considered to be members of the same class or family of polypeptides. For each such class, the present specification provides and / or those skilled in the art will be aware of exemplary polypeptides within the class whose amino acid sequences and / or functions are known; in some embodiments, such exemplary polypeptides are reference polypeptides for the polypeptide class or family. In some embodiments, a member of a polypeptide class or family shows significant sequence homology or identity with, shares a common sequence motif (e.g., a characteristic sequence element) with, and / or shares a common activity (in some embodiments at a comparable level or within a designated range) with a reference polypeptide of the class; in some embodiments with all polypeptides within the class). For example, in some embodiments, a member polypeptide shows an overall degree of sequence homology or identity with a reference polypeptide that is at least about 30-40%, and is often greater than about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more and / or includes at least one region (e.g., a conserved region that may in some embodiments be or comprise a characteristic sequence element) that shows very high sequence identity, often greater than 90% or even 95%, 96%, 97%, 98%, or 99%. Such a conserved region usually encompasses at least 3-4 and often up to 35 or more amino acids; in some embodiments, a conserved region encompasses at least one stretch of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35 or more- 66 - 1324629 IvlAttorney Docket No. 2013237-1500contiguous amino acids. In some embodiments, a relevant polypeptide may comprise or consist of a fragment of a parent polypeptide.

[0143] Read count'. As used herein, the term “read count” refers to a number of sequencing data reads that map to a particular segment, portion thereof, or individual location (such as a SNP) within a genome. For example, the phrases “read count of a particular segment” and “segment read count” as used herein refer to a number of reads that map to the particular segment. For example, the phrases “read count of a particular SNP” and “SNP read count” as used herein refer to a number of reads that map to the particular SNP. The term “read count” may be preceded by an indication of a particular set of sequencing data and / or sequenced sample. For example, when a tumor sample is sequenced to produce tumor sequencing data comprising a plurality of tumor sequencing reads, the phrase “tumor read count” is used to refer to the number of tumor sequencing reads that map to a particular segment, portion thereof, or individual location (e.g., SNP) within a genome. Likewise, when a normal sample is sequenced to produce normal sequencing data comprising a plurality of normal sequencing reads, the phrase “normal read count” is used to refer to the number of normal sequencing reads that map to a particular segment, portion thereof, or individual location (e.g., SNP) within a genome.

[0144] Reference'. As used herein, the term “reference” describes a standard or control relative to which a comparison is performed. For example, in some embodiments, an agent, animal, individual, population, sample, sequence or value of interest is compared with a reference or control agent, animal, individual, population, sample, sequence or value. In some embodiments, a reference or control is tested and / or determined substantially simultaneously with the testing or determination of interest. In some embodiments, a reference or control is a historical reference or control, optionally embodied in a tangible medium. Typically, as would be understood by those skilled in the art, a reference or control is determined or characterized under comparable conditions or circumstances to those under assessment. Those skilled in the art will appreciate when sufficient similarities are present to justify reliance on and / or comparison to a particular possible reference or control.

[0145] Ribonucleic acid (RNA) or Polyribonucleotide'. As used herein, the term “ribonucleic acid,” “RNA,” or “polyribonucleotide” refers to a polymer of ribonucleotides. In some embodiments, an RNA is single stranded. In some embodiments, an RNA is double - 67 - 1324629 IvlAttorney Docket No. 2013237-1500stranded. In some embodiments, an RNA comprises both single and double stranded portions. In some embodiments, an RNA can comprise a backbone structure as described in the definition of “Nucleic acid / Polynucleotide” above. An RNA can be a regulatory RNA (e.g., siRNA, microRNA, etc.), or a messenger RNA (mRNA). In some embodiments, an RNA is a mRNA. In some embodiments, where an RNA is a mRNA, an RNA typically comprises at its 3' end a poly(A) region. In some embodiments, where an RNA is a mRNA, an RNA typically comprises at its 5' end an art-recognized cap structure, e.g., for recognizing and attachment of a mRNA to a ribosome to initiate translation. In some embodiments, an RNA is a synthetic RNA. Synthetic RNAs include RNAs that are synthesized in vitro (e.g., by enzymatic synthesis methods and / or by chemical synthesis methods).

[0146] Ribonucleotide'. As used herein, the term “ribonucleotide” encompasses unmodified ribonucleotides and modified ribonucleotides. For example, unmodified ribonucleotides include the purine bases adenine (A) and guanine (G), and the pyrimidine bases cytosine (C) and uracil (U). Modified ribonucleotides may include one or more modifications including, but not limited to, for example, (a) end modifications, e.g., 5' end modifications (e.g., phosphorylation, dephosphorylation, conjugation, inverted linkages, etc.), 3' end modifications (e.g., conjugation, inverted linkages, etc.), (b) base modifications, e.g., replacement with modified bases, stabilizing bases, destabilizing bases, or bases that base pair with an expanded repertoire of partners, or conjugated bases, (c) sugar modifications (e.g., at the 2' position or 4' position) or replacement of the sugar, and (d) internucleoside linkage modifications, including modification or replacement of the phosphodiester linkages. The term “ribonucleotide” also encompasses ribonucleotide triphosphates including modified and non-modified ribonucleotide triphosphates.

[0147] Secretory signal'. As used herein, the term “secretory signal” refers to an amino acid sequence motif that targets associated polypeptides for translocation to a secretory pathway.

[0148] Segment, Segments: As used herein, the terms “segment” or “segments” (e.g., when used in regard to genetic material) refer to specific pre-defined regions of one or more genomes. For example, a particular reference genome may be subdivided into a plurality of segments, each a specific subsequence of consecutive nucleotides in the reference genome. In certain embodiments, a particular genome may be subdivided into its constituent genes, such that - 68 - 1324629 IvlAttomey Docket No. 2013237-1500each segment corresponds to a particular, different, gene of the particular genome. In certain embodiments, exons of a particular genome are identified and retained, such that each segment corresponds to a particular, different, exon. In certain embodiments, each segment corresponds to a locus (e.g., a particular location on a chromosome where a particular gene, genetic marker, or allele is located). In certain embodiments, a particular genome may be subdivided into segments of a same or substantially same size (e.g., number of bases). As will be understood by one of skill in the art, a set sequencing data obtained from a particular sample, such as reads obtained via next generation sequencing (NGS) data obtained by sequencing a particular sample, may be aligned to a reference genome. In this way, reference genome may be subdivided into a plurality of segments and corresponding segments identified within a genome characteristic of the particular sample. Reads from the set of sequencing data can, accordingly, be identified as mapping to various particular segments within the genome characteristic of the particular sample and used to characterize them. In certain embodiments, multiple sets of sequencing data may be obtained for different samples (e.g., tumor sequencing data from a tumor sample, normal sequencing data from a normal sample) and aligned to a common reference genome. In this way, corresponding segments that comprises the same or substantially same (e.g., all save for variations due to e.g., single nucleotide polymorphisms (SNPs), single nucleotide variations (SNVs), insertions, deletions, etc.) base positions as from genomes characteristic of different samples can be identified. That is, given a particular segment from one genome, associated with one sample, a corresponding segment of another genome, associated with another sample, may be identified. Corresponding segments may have a same and / or substantially same length (e.g., accounting for insertions, deletions, etc.). A particular segment is referred to herein as encoding or comprising a particular SNP and / or SNV if that particular SNP and / or SNV is within the particular segment.

[0149] Single Nucleotide Polymorphism (SNP): As used herein, the term “single nucleotide polymorphism” or “SNP” refers to a particular site (e.g., base position) in a genome where alternative bases are known and / or determined to distinguish one allele from another.

[0150] Single Nucleotide Variation (SNV): As used herein, the term “single nucleotide variation” is used to refer to a difference in the nucleic acid sequence (substitution of one base for another) at a particular site (allele) when comparing a genome from a diseased cell, such as a- 69 - 1324629 IvlAttomey Docket No. 2013237-1500tumor cell, and a genome of a normal, non-diseased cell or a reference genome. In some embodiments, detecting mutations may refer to detecting nucleotide substitution mutations. In certain embodiments, a SNV is a somatic point mutation that occurs only in diseased (e.g., cancer) cells.

[0151] Subject-. As used herein, the term “subject” refers to an organism to be administered with a composition described herein, e.g., for experimental, diagnostic, prophylactic, and / or therapeutic purposes. Typical subjects include animals (e.g., mammals such as mice, rats, rabbits, non-human primates, domestic pets, etc.) and humans. In some embodiments, a subject is a human subject. In some embodiments, a subject is suffering from a disease, disorder, or condition (e.g., cancer and / or a cancer-associated condition). In some embodiments, a subject is susceptible to a disease, disorder, or condition (e.g., cancer and / or a cancer-associated condition). In some embodiments, a subject displays one or more symptoms or characteristics of a disease, disorder, or condition (e.g., cancer and / or a cancer-associated condition). In some embodiments, a subject displays one or more non-specific symptoms of a disease, disorder, or condition (e.g., cancer and / or a cancer-associated condition). In some embodiments, a subject does not display any symptom or characteristic of a disease, disorder, or condition (e.g., cancer and / or a cancer-associated condition). In some embodiments, a subject is someone with one or more features characteristic of susceptibility to or risk of a disease, disorder, or condition (e.g., cancer and / or a cancer-associated condition). In some embodiments, a subject is a patient. In some embodiments, a subject is an individual to whom diagnosis and / or therapy is and / or has been administered.

[0152] Therapy. The term “therapy” refers to an administration or delivery of an agent or intervention that has a therapeutic effect and / or elicits a desired biological and / or pharmacological effect (e.g., has been demonstrated to be statistically likely to have such effect when administered to a relevant population). In some embodiments, a therapeutic agent or therapy is any substance that can be used to alleviate, ameliorate, relieve, inhibit, prevent, delay onset of, reduce severity of, and / or reduce incidence of one or more symptoms or features of a disease, disorder, and / or condition (e.g., cancer and / or a cancer-associated condition). In some embodiments, a therapeutic agent or therapy is a medical intervention that can be performed to- 70 - 1324629 IvlAttorney Docket No. 2013237-1500alleviate, relieve, inhibit, present, delay onset of, reduce severity of, and / or reduce incidence of one or more symptoms or features of a disease, disorder, and / or condition.

[0153] Treat'. As used herein, the term “treat,” “treatment,” or “treating” refers to any method used to partially or completely alleviate, ameliorate, relieve, inhibit, prevent, delay onset of, reduce severity of, and / or reduce incidence of one or more symptoms or features of a disease, disorder, and / or condition (e.g., cancer and / or a cancer-associated condition). Treatment may be administered to a subject who does not exhibit signs of a disease, disorder, and / or condition (e.g., cancer and / or a cancer-associated condition). In some embodiments, treatment may be administered to a subject who exhibits only early signs of the disease, disorder, and / or condition (e.g., cancer and / or a cancer-associated condition), for example for the purpose of decreasing the risk of developing pathology associated with the disease, disorder, and / or condition. In some embodiments, treatment may be administered to a subject at a later-stage of disease, disorder, and / or condition (e.g., cancer and / or a cancer-associated condition).

[0154] Wild-Type'. As used herein, the term “wild-type” refers to an entity having a structure and / or activity as found in nature in a “normal” (as contrasted with mutant, diseased, altered, etc.) state or context. Those of ordinary skill in the art will appreciate that wild-type genes and polypeptides often exist in multiple different forms (e.g., alleles). For example, for a subject with cancer, wild-type genes, segments, SNPs, or properties thereof may be those present in a genome of that subject’s normal, non-cancerous, cells, as opposed to altered versions of those genes, segments, SNPs, or properties thereof that appear in genomes of cancer cells within the subject.DETAILED DESCRIPTION

[0155] It is contemplated that systems, architectures, devices, methods, and processes of the claimed invention encompass variations and adaptations developed using information from the embodiments described herein. Adaptation and / or modification of the systems, architectures, devices, methods, and processes described herein may be performed, as contemplated by this description.- 71 - 1324629 IvlAttomey Docket No. 2013237-1500

[0156] Throughout the description, where articles, devices, systems, and architectures are described as having, including, or comprising specific components, or where processes and methods are described as having, including, or comprising specific steps, it is contemplated that, additionally, there are articles, devices, systems, and architectures of the present invention that consist essentially of, or consist of, the recited components, and that there are processes and methods according to the present invention that consist essentially of, or consist of, the recited processing steps.

[0157] It should be understood that the order of steps or order for performing certain action is immaterial so long as the invention remains operable. Moreover, two or more steps or actions may be conducted simultaneously.

[0158] The mention herein of any publication, for example, in the Background section, is not an admission that the publication serves as prior art with respect to any of the claims presented herein. The Background section is presented for purposes of clarity and is not meant as a description of prior art with respect to any claim.

[0159] Documents are incorporated herein by reference as noted. Where there is any discrepancy in the meaning of a particular term, the meaning provided in the Definition section above is controlling.

[0160] Headers are provided for the convenience of the reader - the presence and / or placement of a header is not intended to limit the scope of the subject matter described herein.

[0161] Presented herein are methods and systems that allow biophysical and / or genomic properties of tumor samples to be determined via analysis of sequencing data, obtained, for example, via next-generation sequencing of tumor samples obtained from a subject. Among other things, the present disclose allows for accurate estimation of tumor sample purity at a variety of sample conditions that are experienced in clinical practice. The ability to determine tumor sample purity is particularly relevant for cancer diagnostics and treatment approaches that rely on mutation signatures of patients’ tumors. For example, personalized cancer immunotherapies rely on accurate and sensitive detection of mutations in the patient’s tumor. In order to achieve the most sensitive mutation detection without errors across all tumor content ranges, a precise estimation of the purity of the tumor sample should be obtained.- 72 - 1324629 IvlAttorney Docket No. 2013237-1500

[0162] Purity can be estimated by two primary means, either by leveraging somatic copy number variations (CNV s) in the tumor genome in conjunction with heterozygous SNPs by means of a tumor deconvolution approach (CNV-based purity estimation) or by leveraging the mutations in the tumor genome (SNV-based purity estimation). SNV-based purity estimation is an important pathway for purity estimation, in particular in cases when the tumor content is low or when CNV-based purity estimation is not successful and, in some embodiments, only yields an upper bound on purity. Achieving accurate purity estimation, in particular when the tumor content is low, is critical for achieving highly sensitive mutation detection, which is required for designing highly efficacious personalized cancer immunotherapies. Among other things, the present disclosure describes methods by which the success and accuracy of both SNV-based purity estimation and CNV-based purity estimation can be enhanced.

[0163] Among other things, the present disclosure provides methods and systems which provide accurate and sensitive measures of tumor sample purity, enabling improved mutation detection and tumor modeling for a large fraction of patients across all purity and tumor mutation burden (TMB) ranges. The methods of the present disclosure provide certain advantages for indications and / or samples wherein tumor content is low or when the TMB is low (and especially in the most challenging scenarios when both cases co-occur), because it is in these types of tumor samples that the performance of mutation detection is most critical. Low TMB tumors may be hallmarks of many indications but can also occur in any indication, and low tumor content samples are often encountered in the clinic. Taken together, methods and systems of the present disclosure provide an important contribution for personalized cancer immunotherapies (e.g., in typical treatment settings).

[0164] CNV-based purity estimation encompasses methods for determining purity of a sequenced tumor sample. A purity can be defined as the fraction of tumor cells present in the sequenced tumor sample, e.g., out of all tumor and normal cells that were sequenced. CNV-based purity estimation may also involve determining the absolute copy numbers of all genes and / or segments within the genome or exome, and, when possible, determining also the allele specific copy numbers of genes and / or segments, a process also referred to herein as tumor deconvolution.- 73 - 1324629 IvlAttomey Docket No. 2013237-1500

[0165] An advantage of CNV-based purity estimation is that this estimation relies on thousands of heterozygous segments (HSs) [e.g., segments of the genome containing at least one heterozygous SNP] resulting in a robust estimation of purity. In some embodiments, a CNV-based purity estimation method as described herein has a further advantage that if a “scaling” of copy numbers cannot be determined with confidence, an upper bound on purity can always be determined by taking the maximum purity across all possible scalings. In this context, “scaling” can mean the value assumed for the primary copy number, which must be an even non-negative integer. In some embodiments, a CNV-based purity estimation method as described herein determines a correct “scaling” for copy numbers, which in the context of proposed method translates to determining the value of the primary copy number. When it is not possible to determine a scaling of copy numbers with confidence because there is no clear scaling that solves the problem satisfactorily, a CNV-based purity estimation method will fail one or more internal CNV-based purity estimation QC metrics. In these cases, an upper bound on purity can always be determined, e.g., by taking the maximum purity across all possible scaling values, i.e., taking the maximum estimated purity when assigning a primary copy number any potential value (e.g., 2, 4, 6 and 8).

[0166] Although methods for CNV-based purity estimation as described herein are highly robust methods for purity estimation, their performance generally diminishes with tumor content. Due to finite coverage and finite number of heterozygous available for estimation, variables extracted from a tumor genome, such as allele frequencies and read counts, have stochastic errors, and such errors begin to dominate as purity diminishes. Below a certain tumor content threshold, it may no longer be possible to resolve CNV clusters and only a bound on purity may be determined. In practice, this means that, in certain embodiments, ability to determine a precise estimation of purity using CNV-based purity estimation is limited (e.g., to a certain minimal purity). For example, when performing CNV-based purity estimation based on whole exome sequencing (WES) data, a minimal purity that can be estimated using this approach is in the range of approximately 20% to approximately 30%, (e.g., about 25%).

[0167] SNV-based purity estimation methods of the present disclosure can provide accurate purity estimation at lower purities compared to CNV-based purity estimation. For example, when performing SNV-based purity estimation based on WES data, a minimal purity- 74 - 1324629 IvlAttorney Docket No. 2013237-1500may be estimated in the range of 10% to 15%. Ranges of purity for which a given method may be accurately performed can depend on various factors, such as, for example, stochastic errors due to finite coverage, experimental or sequencing-related errors, bioinformatic or alignment related errors, genomic complexity, presence of intratumor heterogeneity (ITH). Accordingly, a particular advantage of SNV-based purity estimation methods and systems of the present disclosure is that such methods and systems are capable of estimating purity down to significantly lower values compared to CNV-based purity estimation. In some embodiments, accurate purity estimation at low tumor contents may be critical for achieving high sensitivity of mutation detection without increasing a rate of false positives, since the lower the purity, the more impact accurate purity estimation has on sensitivity and specificity. For example, a 1 % difference in the estimated purity when the purity is 10% is likely to have a much greater impact on mutation detection performance compared to a 1% difference in the estimated purity when the purity is 40%. Therefore, in some embodiments, if an accurate purity estimation is not provided to a mutation detection method when tumor content is low, performance of mutation detection in terms of sensitivity and / or specificity may be significantly diminished. Hence, in such embodiments, ability to provide an accurate purity estimation for low tumor content samples may be particularly useful for achieving optimal or close to optimal mutation detection performance in terms of sensitivity and specificity.

[0168] In some embodiments, if CNV-based purity estimation is successful, this estimation will be preferred over SNV-based purity estimation because CNV-based purity estimation is a robust approach that relies on thousands of heterozygous segments (HSs) across the entire tumor genome. Since virtually all tumor samples contain CNVs, this method can be applied to every patient. However, in some embodiments, CNV-based purity estimation is not always possible. For example, when the tumor content is low, genomic complexity is particularly high coupled with high ploidy, or ITH is significant, CNV-based purity estimation may not be feasible. Accordingly, in some embodiments, either the estimated purity will lie on the lower bound of a purity grid search range, e.g., which is set to a value in the range of 20 to 30%, or, if a CNV-based purity estimation QC metric fails, CNV-based purity estimation will yield an upper bound on purity as described above. In such cases, an estimated purity may effectively be an upper bound on a true purity. Taking an upper bound can reduce loss of sensitivity without risking decreasing specificity. However, in some embodiments, an upper bound on purity may - 75 - 1324629 IvlAttomey Docket No. 2013237-1500reduce the sensitivity of mutation detection if a true purity is lower, especially if tumor content is low. In such cases, SNV-based purity estimation becomes a better method to obtain a purity estimation.

[0169] The safety and efficacy of a personalized cancer immunotherapy may depend, among other factors, on the targets (e.g., neoantigen epitopes) selected to be included in it. For example, when a personalized cancer immunotherapy has a limited number of slots available for targets and those targets include false positives, efficacy of the personalized cancer immunotherapy may in theory be reduced for some patients. This is particularly true if the number of targets included in the personalized cancer immunotherapy is low, because each target has a finite probability for generating an immune response. This effect can be exacerbated if antigen competition plays a role in the immune response: if a strong immune response is directed against a false positive, this immune response may come at the expense of immune responses against the other true positives included in the therapy.

[0170] Another factor that can exacerbate clinical impact of errors is a TMB of a tumor sample. Given a fixed rate of false positives, the lower the TMB, the more likely a false positive may be selected to be included in the personalized cancer immunotherapy. This relation can be quantified by the positive predictive value (PPV) of mutation detection, given by PPV=TP / (TP+FP), where TP is a number of true positives detected in the tumor sample, and FP is a number of false positives detected in the tumor sample. According to this equation, the same FP value will have a greater impact on the PPV when TP is lower. Accordingly, in general, lower TMBs are associated with more significant error terms (FPs) f, leading to a greater drop in the PPV for an individual patient. Therefore, lower TMBs, are associated with larger impact of errors on performance of a personized cancer immunotherapy.

[0171] Including a false positive as a target in the personalized cancer immunotherapy also has a theoretical potential to lead to side effects. For example, if a false positive happens to be a SNP present in the patient’s normal genome, there is a potential theoretical risk for breaking tolerance and generating an autoimmune response against that target. Therefore, to maximize safety and efficacy of personalized cancer immunotherapies it is preferable to utilize a mutation detection method that does not predict false positives.- 76 - 1324629 IvlAttorney Docket No. 2013237-1500

[0172] On the other hand, in order to select the targets that will lead to the most effective clinical response, it is important that a pool of potential targets be as large as possible. This is because targets are prioritized based on selection criteria such as expression, likelihood of eliciting an immune response, ability to confirm tumor control, and so on. Therefore, having the largest possible pool of mutations to select from may increase the overall clinical performance of the personalized cancer immunotherapy.

[0173] Achieving high sensitivity is particularly important when tumor content is low because sensitivity tends to diminish below a certain tumor content as mutations become increasingly more difficult to detect in sequencing data. Achieving high sensitivity is also important when the TMB is low because in this case every mutation that can be detected can be significant for target selection. Achieving high sensitivity is most critical when both TMB and tumor content are low. In extreme cases, and in particular in low TMB / low tumor content scenarios, there is a risk that if sensitivity of mutation detection is not sufficiently high, not enough targets will be detected to design a therapy and a patient will be excluded. Therefore, achieving high sensitivity for mutation detection is critical for the clinical success of personalized cancer immunotherapies.

[0174] Accordingly, in some embodiments, e.g., in the context of a personalized cancer immunotherapy, particularly useful mutation detection methods maximize sensitivity while controlling for errors.

[0175] The present disclosure provides an insight that maximizing sensitivity of mutation detection while controlling for errors across a full range of possible tumor contents should involve incorporation of purity in the process of mutation detection. However, if estimated purity is lower or higher compared to the true purity, this can negatively impact performance of mutation detection and therefore, by extension, performance of a personalized cancer immunotherapy.

[0176] In the case of SNV-based purity estimation, a risk is that if the SNVs included in the SNV-based purity estimation method violate assumptions underlying this method, this could lead to an error or bias in purity estimation, and in turn, lead to either lower sensitivity and / or lower specificity of mutation detection. If an estimated purity is lower than the true purity, this could lead to over- sensitivity to detect mutations, which can increase the rate of false positives.- 77 - 1324629 IvlAttorney Docket No. 2013237-1500For example, in some embodiments, a mutation confidence score can be used to assess the likelihood that a given site encodes a mutation. Such a mutation confidence score can be devised to take into account the estimated purity of the tumor sample. In some embodiments, where a purity is estimated to be lower than a true purity, a mutation confidence score can assign putative mutations (including false positives) a higher confidence score, causing putative mutations to be erroneously accepted thereby leading to a higher rate of false positives. In some embodiments, one or more false positive (FP) filters (filters designed to discard false positives arising from specific sources of noise) can be potentiated by an estimated purity of the tumor sample and / or by a mutation confidence score. In some embodiments, if an estimated purity is lower than a true purity, this may lower sensitivity of filters to remove false positives and therefore increase overall rate of false positives. Impact of an underestimated purity may be exacerbated when tumor content is low because in such a regime, even a small underestimation of the purity can lead to oversensitivity to detect mutations, in which case false positives may also be detected.

[0177] Conversely, in some embodiments, if an estimated purity is higher than a true purity, this may lead to a lower sensitivity of mutation detection. For example, if a mutation confidence score receives as input a purity estimation that is higher than a true purity then the sensitivity of this metric to detect mutations will diminish, leading to an overall lower mutation detection sensitivity. Likewise, if a FP filter potentiated by the purity receives as input a purity that is higher than a true purity, such a filter may be overly sensitive to filter putative mutations, resulting in an overall lower sensitivity to detect mutations.

[0178] In some embodiments, if SNVs included in SNV-based purity estimation method violate model assumptions, this can cause SNV-based purity estimation to fail a SNV-based purity estimation QC metric, in which case purity estimation will rely on CNV-based purity estimation. When the tumor content is low, CNV-based purity estimation will generally yield an upper bound purity, which will, in turn, limit the sensitivity of mutation detection and hence can also potentially diminish clinical performance of the personalized cancer immunotherapy.

[0179] In some embodiments, when TMB is low and / or when purity is low, an effect of an error or bias in the estimated purity may be more impactful. An underestimation of purity can lead to an increase in the rate of false positives, which has a greater impact when TMB is low or when a number of detected targets is low due to low purity. An over-estimation of purity can lead - 78 - 1324629 IvlAttomey Docket No. 2013237-1500to reduced sensitivity, which has a greater potential to impact clinical performance of a personalized cancer immunotherapy when the TMB is low or when a number of detected targets is low due to low purity. In some embodiments, when both TMB and tumor content are low, an error or bias in the estimated purity may be particularly impactful. The present disclosure provides, among other things, is based on an insight that enabling accurate purity estimation may be particularly useful for success of personalized immunotherapies, in particular in low tumor content and / or low TMB tumors.

[0180] To conclude, in some embodiments, including SNVs that violate assumptions underlying SNV-based purity estimation may lead to suboptimal mutation detection and consequently potentially suboptimal clinical performance of a personalized cancer immunotherapy. In some embodiments, in which TMB and / or tumor content are low, possibility of inclusion of SNVs that violate assumptions underlying SNV-based purity estimation and / or possibility of suboptimal mutation detection may be higher. In some embodiments, suboptimal mutation detection may lead to exclusion of the patient due to lack of a sufficient number of targets to design a personalized cancer immunotherapy. Thus, while in some embodiments, SNV-based purity estimation can provide accurate purity estimations (e.g., when the CNV-based purity estimation is not possible), which may be important for achieving high sensitivity of mutation detection (e.g., when the tumor content is low), there is a risk that using SNVs that violate model assumptions will lead to incorrect SNV-based purity estimation, and consequently suboptimal selection of targets for a personalized cancer immunotherapy. Accordingly, the present disclosure provides methods and systems to increase success and quality control of both SNV-based purity estimation and CNV-based purity estimation.

[0181] The present disclosure provides, among other things, an insight that in order for SNV-based purity estimation to have a unique solution it must be applied to a special subset of SNVs, namely, clonal SNV events in balanced regions of the tumor with a copy number of 2. In some embodiments (e.g., if SNV events are not clonal SNV events and / or are not in balanced regions of a tumor genome and / or are not in regions of a tumor genome with a copy number of 2) then purity estimation may either be biased or not have a unique solution due to mathematical degeneracy that arises when SNV events with other copy numbers are present. In other words, possible VAFs observed in a tumor sample for a given SNV event depend on the absolute copy- 79 - 1324629 IvlAttomey Docket No. 2013237-1500number of the gene or segment containing the SNV event. An insight of the present disclosure is that such complexity may be circumvented by restricting estimation of purity to minimal balanced SNV events (as disclosed herein) and assuming that the minimal copy number is 2 (as disclosed here).

[0182] An advantage provided by methods and systems of the present disclosure for SNV-based purity estimation is that estimation of purity becomes possible when CNV-based purity estimation is not feasible (e.g., when tumor content is low). In such embodiments, absolute copy numbers across a tumor genome will not be known, and therefore it is not possible to use copy number estimates to identify balanced diploid regions in the tumor genome. Further, in such embodiments, it will not be possible to leverage tumor deconvolution to determine a copy number of SNV events or to determine if SNV events are clonal.

[0183] A further insight of the present disclosure is that in such embodiments, other means may be used to enhance a likelihood that SNV-based purity estimation is successful and control for its quality. In some embodiments, an SNV-based purity estimation method of the present disclosure also improves the performance of CNV-based purity estimation.

[0184] In some embodiments of mutation detection methods, mutation detection performance is iteratively improved as tumor modeling is iteratively improved. In some embodiments, iterative improvement is performed using a feedback mechanism. For example, as purity (and possibly copy number estimations) are iteratively refined, false positive filters are potentiated by increasingly more accurate purity (and possibly copy number) estimations, leading to a higher sensitivity to detect mutations while controlling for false positives. Such feedback procedures are particularly useful when tumor content is low, because the lower the tumor content, the more critical it is to have accurate purity estimation in order to obtain maximal sensitivity of mutation detection without increasing a rate of false positives. As noted, when purity is low, small differences in purity estimation can have outsized impact on mutation detection sensitivity and / or specificity. Therefore, in some embodiments, achieving accurate purity estimation is particularly important in order for a mutation detector to achieve optimal performance.

[0185] In some embodiments, since a mechanism to obtain accurate purity estimation in a low tumor content range is SNV-based purity estimation, in order for a feedback tumor - 80 - 1324629 IvlAttorney Docket No. 2013237-1500modeling approach to be maximally effective, it may be important to (1) to maximize a likelihood that SNV-based purity estimation can be applied, (2) ensure that an estimated purity is accurate and unbiased, (3) perform quality control to abort or avert SNV-based purity estimation if appropriate SNV events cannot be identified. In some embodiments, where it can be ascertained that SNV-based purity estimation fails, then CNV-based purity estimation may be preferred. In some embodiments, the present disclosure provides methods and systems to increase a likelihood that SNV-based purity estimation will be successful and will provide an accurate and unbiased estimation of the purity, and in cases of failure, provide, under certain circumstances, feedback that can improve accuracy of a CNV-based purity estimation.

[0186] A certain advantage of methods and systems of the present disclosure for SNV-based purity estimation is that they do not rely on knowledge of copy numbers. In certain embodiments, e.g., where knowledge of copy numbers is not available when methods and systems SNV-based purity estimation may be particularly useful.A. Creating Personalized Cancer Immunotherapies Based on Tumor Genome Analysis

[0187] Certain cancer mutations are unique to a patient’s cancer and, when expressed, produce proteins and / or peptides that are distinct from those produced by normal cells. These distinct proteins and / or peptides can, accordingly, be specifically targeted via immunotherapy approaches that leverage the patient’s own immune system to clear cancer cells while avoiding damage to normal cells. Technologies of the present disclosure, among other things, leverage and analyze sequencing data to identify potential cancer-specific mutations within genome(s) of a patient’s tumor (tumor genome) and, moreover, characterize them in a manner that allows those mutations that will be the most effective targets of immunotherapies to be identified and prioritized, for example for as targets for personalized cancer vaccines, T-cell receptor (TCR) therapies, and the like.

[0188] As illustrated in FIG. 1, biological samples obtained from a patient can be sequenced and the resultant sequencing data analyzed to detect and characterize mutations (SNV events) unique to the patient’s cancer cells. Detected mutations can be prioritized according to various metrics that reflect, for example, their prevalence as well as propensity to induce an immune response. Non- synonymous mutations that are both highly prevalent (e.g., present in a - 81 - 1324629 IvlAttomey Docket No. 2013237-1500substantial fraction, up to all, of the patient’s cancer cells) and likely to be effective in priming the patient’s immune system to mount a strong response can, accordingly, be selected for inclusion in a personalized immunotherapy for the patient. A personalized therapeutic can thus be designed, manufactured, and administered to the patient as treatment.

[0189] For example, as illustrated in FIG. 1, in certain embodiments, normal 102a and tumor tissue 102b samples are obtained from the patient. Normal genomic DNA (gDNA) is extracted from a normal tissue sample 102a and tumor gDNA is extracted from a tumor tissue sample 102b. Sequencing (104) may then be performed using the extracted normal and tumor gDNA to generate sequencing data.

[0190] Sequencing (104) may be performed, for example as described in further detail herein, using next generation sequencing (NGS) techniques. Accordingly, in certain embodiments, various pre-processing steps (106) are performed, for example to align reads of sequencing data to a reference genome.

[0191] In certain embodiments, sequencing data [e.g., having been pre-processed (e.g., to align reads to a reference genome)] is used (e.g., as input) for tumor deconvolution and / or mutation detection (108) techniques of the present disclosure. For example, as described in further detail herein, tumor deconvolution and / or mutation detection techniques of the present disclosure operate on sequencing data to determine one or more biophysical and / or genomic features 110 of a patient’s cancer. These determined biophysical and / or genomic features characterize genomic properties of the patient’s cancer (e.g., to the extent represented in the tumor tissue sample), physical properties of the tumor sample 102b. For example, in certain embodiments, one or more tumor genomic features 110a are determined. Tumor genomic features may include, without limitation, copy numbers (e.g., absolute copy numbers and / or allele- specific copy numbers) of one or more segments within a tumor genome (e.g., all segments; e.g., a particular subset of segments, such as heterozygous segments) and / or mutations (SNVs) therein, as well as characteristics of detected mutations (SNVs), such as their zygosity and / or clonality state. In certain embodiments, one or more sample features 110b are determined. Sample features 110b may include, without limitation, a sample purity and / or contamination fraction, which reflect the potential for and amount of tumor samples to include a non-trivial and, at times, substantial, fraction of normal, non-cancerous, cells. In certain - 82 - 1324629 IvlAttomey Docket No. 2013237-1500embodiments, mutations (e.g., SNVs) are detected 110c. Detected mutations may, for example, be determined by tumor deconvolution and / or mutation detection technologies and provided, for example as a standardized file such as a variant call format (.vcf) file.

[0192] In certain embodiments, detected mutations 110c are filtered to identify non-synonymous mutations (112).

[0193] In certain embodiments, non-synonymous mutations are prioritized (114) to select a subset of mutations for inclusion in a personalized cancer immunotherapy 120. For example, in certain embodiments, mutations (e.g., non-synonymous mutations) are evaluated to determine which particular mutations are, or are predicted to be, present in a substantial fraction (e.g., above a certain threshold, up to all) of a patient’s cancer cells, so that an immune response that targets and eliminates cells expressing one or more particular mutations is likely to eliminate a substantial fraction of the patient’s cancer cells. In certain embodiments, additionally or alternatively, mutations that are determined to, or predicted likely to, elicit a strong immune response within the patient are prioritized and selected for.

[0194] For example, in certain embodiments, one or more scoring metrics 116 are determined for each of at least a portion of detected mutations 100c [e.g., a subset identified as non-synonymous (112), or a portion thereof]. Scoring metrics 116 may be used to prioritize and select a subset of mutations for inclusion in a personalized cancer immunotherapy 120. Scoring metrics 116 may include metrics characterizing potential for a particular mutation to elicit an immune response and may include, without limitation, major histocompatibility complex (MHC) binding predictions, T-cell receptor (TCR) recognition predictions, expression level predictions, and the like. These immune response metrics may be determined via various approaches, including, for example, machine learning and other techniques. Additionally or alternatively, scoring metrics may include metrics determined via tumor deconvolution and mutation detection technologies described herein, such as, without limitation, absolute copy number values, clonality state classifications, values and / or classifications indicating a function of a gene harboring a given mutation, values and / or classifications indicating whether a given mutation is truncal and / or early, and zygosity and / or fractional zygosity values.

[0195] In certain embodiments, a prioritized subset of detected mutations may be used for a personalized cancer immunotherapy 120. For example, in certain embodiments,- 83 - 1324629 IvlAttomey Docket No. 2013237-1500compositions encoding one or more of a prioritized subset of mutations may be manufactured and administered to a patient, for example as a personalized cancer vaccine.A.i Patient Samples and Sequencing Data

[0196] Turning to FIG. 2, as described herein, tumor deconvolution and / or mutation detection technologies of the present disclosure may be used in connection with (e.g., to analyze) sequencing data for a subject to determine genomic and / or biophysical properties of, and / or detect mutations characteristic of, a patient’s cancer. Among other things, technologies of the present disclosure may be used in connection with sequencing and / or genomic analysis methods and systems that generate sequencing data for a subject and allow for detection of mutations that can be targeted, for example, for disease treatment via immunotherapy. Among other things, technologies of the present disclosure may be used in connection with sequencing and / or genomic analysis methods and systems that generate sequencing data for a subject and allow for detection of mutations that can be targeted, for example, for disease treatment via immunotherapy.

[0197] For example, as shown in FIG. 2, in certain embodiments, sequencing data 228 may be generated for a subject having and / or suspected of having cancer, by obtaining a tumor sample 202 from the subject. The tumor sample 202 may be processed 204, for example, to extract and prepare nucleic acid material for sequencing and sequenced 206 to generate tumor sequencing data 208 - e.g., sequencing data representing and obtained from nucleic acid material 204 from a tumor sample 202. Likewise, in certain embodiments, sequencing data 228 may (e.g., also) include normal sequencing data 218 from normal sample 212 having been obtained, processed 214, and sequenced 216, to generate normal sequencing data 218 - e.g., sequencing data representing, and obtained from, nucleic acid material from a normal sample 212.

[0198] A tumor sample 102 may be any sample derived from a particular subject and comprising, and / or expected to comprise, cancer cells e.g., of the particular subject). In certain embodiments, a tumor sample is or comprises a liquid sample, such as serum, plasma, blood, urine, etc. For example, a liquid sample, such as blood, may comprise, or be suspected of comprising, cancer cells, such as circulating tumor cells (CTCs). In certain embodiments, a tumor sample is or comprises a tissue sample, for example, obtained from a subject via biopsy. A - 84 - 1324629 IvlAttomey Docket No. 2013237-1500tumor sample may be representative of a subject’s primary tumor and / or one or more metastases. For example, a primary tumor sample may be obtained via biopsy of a region of a subject known or expected to harbor a primary tumor. A metastasis sample or metastases samples may be obtained via biopsy of one or more region(s) of a subject known or expected to harbor metastases. Tumor samples may be obtained and / or preserved in a variety of formats, such as, for example, isolated cells (e.g., CTCs) from blood samples, fresh tissue, flash frozen tissue, formalin-fixed paraffin embedded tumor tissue (FFPE samples), and the like.

[0199] As illustrated in FIG.2, a tumor sample may comprise cancer cells 202b as well as, in certain cases, normal (z.e., non-cancerous) cells 202a. Accordingly, in certain embodiments, a tumor sample purity may measure relative fraction of cancer cells within a tumor sample. Tumor sample purity may, for example, be computed as (1 — ) =+?]w), where LI is a contamination fraction, representing a relative fraction of normal cells infiltrating a tumor sample, given by j = rN / (TT+ 7]w) and Z T and Z N are a number of tumor and normal cells in a tumor sample, respectively. Tumor sample purity may be expressed as a decimal value, percentage, etc. As described in further detail herein, typically, purity does not need to be measured directly (e.g., via direct measuring / counting amounts of tumor and normal cells in a sample), but, rather, can be determined and / or estimated using sequencing data, for example, via tumor modelling approaches such as those described herein.

[0200] As described in further detail herein, in certain embodiments, additionally or alternatively, bounds, such as upper and / or lower bounds for sample purity (e.g., and / or contamination fraction), may be estimated, e.g., via tumor modeling approaches of the present disclosure. For example, as described in further detail herein, in certain embodiments, at low physical sample purities, accuracy of estimation methods, such as tumor deconvolution, may be reduced such that purity estimates based on sequencing data are expected to be of insufficient accuracy to be used in and of themselves. In such cases, however, an upper bound (e.g., a maximum purity) and / or a lower bound (e.g., a minimum purity) may still be estimates and used in certain processing steps.

[0201] A normal sample 212 may be any sample derived from a particular subject and comprising, and / or expected to comprise, the subject’s normal cells, but not cancer cells (e.g., in certain embodiments, a normal sample 212 does not contain any cancer cells). In certain - 85 - 1324629 IvlAttomey Docket No. 2013237-1500embodiments, a normal sample is or comprises a liquid sample, such as serum, plasma, blood, urine, saliva, etc. In certain embodiments, a normal sample is or comprises a tissue sample, for example obtained from a subject via biopsy. Normal samples may be obtained and / or preserved in a variety of formats, such as, for example, isolated cells [e.g., peripheral blood mononuclear cells (PBMCs)] from blood samples, fresh tissue, flash frozen tissue, formalin-fixed paraffin embedded tumor tissue (FFPE samples), and the like.

[0202] As illustrated in FIG.2, a normal sample 212 nominally contains only normal patient cells 212a. In certain embodiments, a normal sample may be obtained from blood of a subject. In certain embodiments, a normal sample may still comprise a small (e.g., negligible) number of cancer cells. For example, in certain embodiments, a normal cell may be obtained from a region of a subject near a tumor e.g., in an effort to obtain a normal sample from a same or similar underlying tissue type). In certain embodiments, a normal sample comprises less than 1%, e.g., less than 0.1%, e.g., less than 0.01%, e.g., less than 0.001% tumor cells.

[0203] Samples, such as tumor samples and / or normal samples, may be processed to obtain, and / or prepare, nucleic acid material therefrom for sequencing. For example, nucleic acid material, such as DNA and / or RNA, may be extracted and prepared for sequencing (e.g., via amplification, fragmentation, labeling, etc.) as appropriate, depending on a particular desired sequencing method and / or data format. In certain embodiments, tumor sample 202 is processed 204 and / or normal sample 212 is processed 214 to extract nucleic acid material, such as gDNA. In certain embodiments, normal gDNA is extracted from normal tissue sample 212 and tumor gDNA is extracted from tumor tissue sample 202. For example, in certain embodiments, sequencing data may be whole genome sequencing (WGS) data; in certain embodiments, sequencing data may be whole exome sequencing (WES) data. Various commercially available kits and instruments may be used to prepare samples for and obtain WGS and / or WES data, including, but not limited to, those provided by Illumina, Inc., PacBio, Oxford Nanopore Technologies, Thermo Fisher Scientific’s Ion Torrent™, etc. In some embodiments, methods and systems of the present disclosure utilize whole exome sequencing (WES) or whole genome sequencing (WGS) data in singletons (e.g., one sequenced tumor sample and one sequenced normal sample). In some embodiments, methods and systems of the present disclosure utilize whole exome sequencing (WES) or whole genome sequencing (WGS) data in replicates (e.g.,- 86 - 1324629 IvlAttomey Docket No. 2013237-1500two or more sequenced tumor samples and two or more sequenced normal samples). In some embodiments, methods and systems of the present disclosure utilize whole exome sequencing (WES) or whole genome sequencing (WGS) data in conjunction with RNA-seq data. In some embodiments, RNA-seq data is used in a method or system of the present disclosure to filter errors such as PCR errors, FFPE, sequencing artifacts, and / or other types of errors. For example, a mutation detected in DNA sequencing data may only be accepted if it is also detected using RNA-seq reads. In some embodiments, RNA seq reads are high quality RNA-seq reads. In some embodiments, if only singletons are used, coverage of singletons is required to be comparable to coverage of merged replicates (e.g., in order to ensure comparable sensitivity).

[0204] In certain embodiments, to generate sequencing data, libraries are created from the normal gDNA and tumor gDNA. In certain embodiments, from each gDNA sample, two or more libraries can be generated. In certain embodiments, from each gDNA sample, two libraries can be generated. For example, in certain embodiments, as illustrated in FIG.2, sequencing data 228 may be generated in and / or comprise replicates created by preparing multiple (e.g., two or more) libraries associated with each (e.g., gDNA) sample. The libraries can be created for whole exome sequencing and / or whole genome sequencing and / or RNA sequencing (RNAseq). Samples are then sequenced using high throughput sequencing such as NGS.

[0205] For example, in certain embodiments, such that, for example, sequencing data 228 may be generated in and / or comprise replicates. For example, multiple tumor samples may be extracted and sequenced independently; in certain embodiments, a single tumor sample may be extracted and used to prepare multiple libraries (e.g., such that processing steps of extracting, fragmenting, and amplifying nucleic from the sample are performed repeatedly and independently), which are then sequenced; in certain embodiments, library preparation may be performed repeatedly on a single pool of extracted nucleic acid, and the multiple libraries sequenced; in certain embodiments, a single library is sequenced multiple times (e.g., as in a technical replicate). In certain embodiments, as with tumor sample sequencing data, e.g., as illustrated in FIG.2, normal sample sequencing data may also comprise a plurality of replicates.

[0206] Sequencing data 228, 208, 218, may be stored and / or presented in a variety of formats, such as FASTQ, SAM, BAM, etc. For example, sequencing data for a particular sample may comprise a plurality of reads, each read representing a nucleotide sequence of a- 87 - 1324629 IvlAttomey Docket No. 2013237-1500polynucleotide fragment corresponding to (e.g., that maps to) a portion of a subject’s tumor genome and / or exome, and / or portion of the subject’s normal genome and / or exome. In certain embodiments, sequencing data typically comprises multiple overlapping reads, which may be aligned to a reference genome, such as an hl9 or h38 reference genome (see, e.g., Ncbi.nlm.nih.gov / datasets / genome / GCF_000001405.13 / and Ncbi.nlm.nih.gov / datasets / genome / GCF_000001405.26 / , respectively; see also Karolchik, D. et al. Nucleic Acids Res. 32, D493-D496, 2004 and Kent WJ et al., Genome Res.l2(6):996-1006, 2002) for human patients, to map each read to a particular region of the overall genome. A reference genome may, accordingly, be a publicly available reference genome. In certain embodiments, a reference genome may be generated based on a normal genome of the subject.

[0207] For example, in certain embodiments, NGS sequencing is performed and generates FASTQ files as output. In certain embodiments, NGS output reads, stored, for example, in FASTQ files, are aligned. Aligned reads may be stored and / or provided in file formats, such as sequence alignment map (SAM) or the binary compressed version thereof (BAM). In certain embodiments, sequencing data duplicate reads may be marked and / or removed from the sequencing data. In certain embodiments, sequencing adapters may be removed from the sequencing data. In certain embodiments, e.g., in the case of short read sequencing, sequencing can be paired end or single end, wherein different read lengths can be used (e.g., 50bp, lOObp, 150bp, etc.).

[0208] In certain embodiments, aligned reads, e.g., may be stored in one or more BAM file(s), are used as input for tumor deconvolution and / or mutation detection technologies described here. For example, in certain embodiments, samples are sequenced using high throughput sequencing such as next generation sequencing, generating, for example, FASTQ files as output. Next, FASTQ files are aligned and, in certain embodiments, duplicate reads are marked. An alignment step may be performed, and generate, as output, a BAM file for a normal sample and a BAM file for a tumor sample. In certain embodiments, output of an alignment step is two normal BAM files and two tumor BAM files.

[0209] As described in further detail in the following, based on input sequencing data (e.g., comprising tumor sequencing data and normal sequencing data), tumor deconvolution and mutation detection technologies of the present disclosure may determine various [biophysical - 88 - 1324629 IvlAttomey Docket No. 2013237-1500and genomic] properties of a tumor sample and detect mutations (e.g., SNVs) occurring within the tumor genome.A.ii Tumor Genomics

[0210] Turning to FIG.3, sequencing data may be used to piece together and / or infer properties relating to a tumor genome and / or normal genome of a subject.A.ii.a Segments

[0211] As illustrated in FIG.3, a genome may be subdivided into a plurality of segments, each segment representing a particular sub-region (e.g., a subsequence) of the genome. In certain embodiments, a reference genome may be subdivided into a plurality of segments. In certain embodiments, a tumor genome and / or a normal genome may be subdivided into a plurality of segments (black bars in the normal genome and tumor genome schematics).

[0212] Among other things, in certain embodiments, as illustrated in FIG.3, this approach provides a common coordinate system or set of subregions between reference genome 302, normal genome 322, and tumor genome 342. For example, (e.g., sequencing data, such as reads, corresponding to) normal genome 322 and (e.g., sequencing data, such as reads, corresponding to) tumor genome 342 may be aligned to reference genome 302. In this way reads can be mapped to particular regions of reference genome 302 and a coordinate system for [reads of] normal genome 322 and tumor genome 342 can be provided. For example, reference genome 302 can be used to determine a chromosome number, a nucleotide position (in the chromosome), as well as a directionality of a given read.

[0213] Thereafter, by partitioning reference genome 302 into a plurality of segments, regions of normal genome 322 and tumor genome 342 that correspond to a particular segment can be identified and, for example, analyzed and compared.

[0214] For example, in the illustrative schematic shown in FIG.3, nine segments 304a, 304b, 304c...304i, and 304j (collectively 304) of reference genome 302 are shown, with- 89 - 1324629 IvlAttomey Docket No. 2013237-1500corresponding segments of normal genome 322 and tumor genome 342 indicated via the vertical dashed lines.A.ii.b Copy Number Variations

[0215] As illustrated in FIG.3, the majority of normal genome 322 is typically diploid, with two copies of each particular segment (e.g., except for segments located on male sex chromosomes for which normal genome 322 contains either a single maternal chromosome or a single paternal chromosome), with one copy associated with (e.g., originating from) a maternal allele 326a and another associated with (e.g., originating from) a paternal allele 326b. In certain embodiments, methods and systems described herein may exclude male sex chromosomes and / or segments that map to locations on male sex chromosomes from analysis since male sex chromosomes do not contain heterozygous SNPs, which, as explained in further detail herein, can be used to facilitate determining tumor genomic properties and / or sample purities.

[0216] As illustrated in FIG.3, for segments in a normal e.g., human) genome that, absent germline CNV events, is diploid, segments may be heterozygous, in that the copies differ, corresponding to different alleles, or may be homozygous, comprising identical alleles.

[0217] In contrast, segments in a tumor genome 342 are not necessarily diploid. Nor are segments in a tumor genome necessarily balanced. Certain segments of a tumor genome may, however, be diploid and / or balanced. For example, as shown in FIG.3, a number of copies of each allele in a segment may vary from segment to segment, and may be less than two, equal to two, or greater than two. Segments in a tumor genome need not be balanced - i.e., tumor genome heterozygous segments do not necessarily comprise a same number of copies of each allele, but, in certain embodiments, may comprise a greater number of copies of one allele (a “major allele”) than the other (a “minor allele”).

[0218] Accordingly, in certain embodiments, various parameters are used to characterize and represent physical properties of a particular segment (e.g., within a normal and / or tumor genome), along with, in certain embodiments, mutations (such as SNVs) identified therein.- 90 - 1324629 IvlAttorney Docket No. 2013237-1500

[0219] For example, in certain embodiments, a segment j may be characterized by an absolute copy number, CNj, which is computed as a number of copies of a particular segment (e.g., within a single tumor cell).

[0220] For example, in the schematic shown in FIG.3, while all segments in normal genome 322 have one copy of a maternal allele and one copy of a paternal allele, copy numbers of maternal and paternal alleles in tumor genome 342 may vary from segment to segment.Accordingly, while all segments in normal genome 322 that are shown in FIG.3 have an absolute copy number of two, in tumor genome 342 segments 344a, 344b, 344c, 344d, 344e, 344f, 344g, 344h, 344i, 344j have absolute copy numbers of 2, 4, 3, 1, 3, 5, 0 (a deletion), 2, 4, and 1, respectively. Moreover, as illustrated in FIG.3, in tumor genome 342, the number of copies of maternal and paternal alleles also may vary from segment to segment such that, for example segment 344a has one copy of each a maternal and paternal allele, segment 344b has two copies of each, and segment 344f has three copies of the maternal allele and two copies of the paternal allele.

[0221] Normal segments can also have other copy number configurations due to germline CNV events, however, without limiting the generality of the method, such events are not described in the figure.A.ii.c Single Nucleotide Polymorphisms and Heterozygous Segments

[0222] In certain embodiments, maternal and paternal alleles of a particular segment in a normal genome 322 may harbor different variants of a single nucleotide polymorphism (SNP). In this case, the particular segment is referred to as comprising a heterozygous SNP. The particular segment in the normal genome 322, along with corresponding segments in reference and tumor genomes, are referred to as heterozygous segments.

[0223] For example, as illustrated in FIG.3, reference genome 302 is subdivided into a plurality of segments. For segment 304a of reference genome 302, corresponding segment 324a in normal genome 322 comprises at least one heterozygous SNP, such that segment 304a of reference genome 302 and corresponding segments 324a and 344a of normal and tumor genome, respectively, are referred to as heterozygous segments. In contrast, segment 3041 of reference- 91 - 1324629 IvlAttomey Docket No. 2013237-1500genome corresponds to segment 324i of normal genome, which does not comprise any heterozygous SNPs. Accordingly, segment 304i and corresponding segments 324i and 344i in normal and tumor genomes are not heterozygous segments.

[0224] Overall, FIG.3 depicts seven (7) heterozygous segments and three (3) homogeneous segments. As illustrated in the figure, for each of the heterozygous segments, copies in the normal genome 322 harbor at least one heterozygous SNP (illustrated schematically via different colored orange and green dots in the maternal and paternal alleles), whereas normal genome 322 copies of the homogeneous segments do not contain any heterozygous SNPs.Different normal genomes (e.g., of different individual subjects) will generally have different heterozygous segments since different normal genomes encode different heterozygous SNPs.

[0225] As explained herein, while normal genome 322 typically has two copies of each segment (apart from those located on male sex chromosomes) the number of copies of particular segments in a tumor genome may differ from two and can vary from segment to segment.Additionally or alternatively, as illustrated in FIG.3, tumor genomes do not necessarily have equal numbers of maternal and paternal alleles for each heterozygous segment and / or, in certain cases, may entirely lack a maternal or paternal copy. For example, segment 304a is a heterozygous segment for which corresponding tumor genome segment 344a has two copies: one maternal and one paternal. Second segment 304b is another heterozygous segment. In this case, however, corresponding segment 344b in tumor genome 342 has four copies: two maternal copies of the segment in the tumor genome and two paternal copies of the segment in the tumor genome. Therefore, the segment has an absolute copy number of 4 in the tumor genome.Reference number 344c shows three copies of a heterozygous segment: one maternal copy of the segment in the tumor genome and two paternal copies of the segment in the tumor genome, therefore the segment has an absolute copy number of 3 in the tumor genome.

[0226] In certain embodiments, a heterozygous segment may comprise only paternal or only maternal alleles, referred to herein as a loss of heterozygosity (LOH) event. For example, segments 344d and 344e in tumor genome each correspond to a normal diploid segment that comprises a heterozygous SNP - i.e., a heterozygous segment. However, as illustrated in FIG.3, segments 344d and 344e in tumor genome lack any maternal allele copies - they have only (one- 92 - 1324629 IvlAttorney Docket No. 2013237-1500and three, respectively) copies of the paternal alleles. Accordingly, these are segments that have undergone a LOH event.

[0227] In certain embodiments, for a given segment (e.g., in reference or normal genome), the corresponding segment may be entirely absent from tumor genome 344g, indicating that, for example, all copies of the segment (both maternal and paternal) were deleted (complete deletion). Segment 344g in tumor genome, accordingly, has an absolute copy number of O.A.ii.d Tumor Genome Mutations and Characteristics

[0228] FIG. 3 illustrates an expanded view of segments 322f and 344f. As illustrated in FIG. 3, a normal cell 352 comprises normal genome 322, including maternal 354a and paternal 354b alleles of segment 322f, with maternal allele comprising an alternative variant of SNP 356.Tumor cell 372 comprises tumor genome 342, including segment 344f, which corresponds to normal segment 322f. Unlike normal cell, tumor cell 372 comprises multiple copies of each allele - namely, three copies of maternal allele and two copies of paternal allele.

[0229] As described herein, corresponding segments 322f and 344f in normal 322 and tumor 324 genomes, respectively, can be characterized by values of parameters such as an absolute copy number and an allele specific copy number. For example, as shown in FIG.3, normal segment 322f has an absolute copy number of two - (CAwt= 2) and corresponding tumor genome segment 344f has an absolute copy number 373 of five (e.g., CAj = 5). As illustrated in FIG. 3, although a normal segment will typically have one copy of each parental allele, CNV events may cause a tumor genome segment to have one or multiple (e.g., two or more) copies of each parental allele. Moreover, the number of maternal and paternal alleles need not be equal in a tumor genome. Accordingly, heterozygous tumor genome segments may also be characterized by an allele- specific copy number. For heterozygous segments, such as segment 354, an allele specific copy number 364 may be determined as a maximum absolute number of copies of a major allele - i.e., the allele having a number of copies greater than or equal to that of the other, minor, allele - i.e., CNx > CNx, where X and Y denote the major and minor alleles, respectively. For example, for the particular segment 354 shown in FIG.3, there are three copies of the major allele, such that the allele specific copy number is three (CNx = 3).- 93 - 1324629 IvlAttomey Docket No. 2013237-1500

[0230] As shown in the bottom portion of FIG. 3, tumor genome segments may harbor mutations, such as an SNV 382. In the notation used herein, a copy number of a segment harboring a mutation may be denoted CAmut (e.g., for segment 344f in tumor cell 372, CAmut = 5). As shown in FIG. 3, mutations may occur in particular alleles, but are not necessarily present in each copy of a particular allele. For example, while SNV 382 occurs in maternal allele 374, it is not present in all three copies of maternal allele 374 - it is present in only two copies.Accordingly, additional parameters may be determined to characterize genomic properties of mutations like SNVs.

[0231] For example, if a certain segment is mutated, a number of physical copies of the mutated segment in a given tumor cell, referred to herein as the zygosity of the mutation, may be determined. In certain embodiments, a fractional zygosity of a mutation (Q may be computed as a ratio of the zygosity of a mutation (Cx) and the absolute copy number (CAmut) of the segment harboring the mutation in the tumor genome. For example, in FIG. 3, segment 344f harbors a mutation 382 having a zygosity of two and a fractional zygosity, of 2 / 5 (0.4).

[0232] Turning to FIG. 4, as described herein, a sample of tumor tissue, may comprise normal cells 400 and tumor cells 410. As shown in FIG. 4 and described herein, segments in normal cells 415 are assumed to have an absolute copy number of two, except for segments occurring on male sex chromosomes for which an absolute copy number is one. As shown in the figure, a segment may be amplified in the tumor cells. In the example illustrated in FIG. 4, the tumor cell segment has an absolute copy number of five 420. As shown in the figure, example mutation 430 (black star) has three copies, and, accordingly, its zygosity is three (440). In the illustrative example shown in FIG. 4, mutation 430 (black star) is present in all tumor cells and is therefore a clonal mutation, whereas a second mutation 450 (red star) is present in a subset of tumor cells (just one of the three tumor cells in the diagram in FIG. 4) and is therefore a subclonal mutation. Clonal mutations with a zygosity greater than 1 can arise, for example, when the mutation occurred before the copy number amplification event (hence are considered “early” mutations), whereas subclonal mutations with a zygosity of 1 can arise, for example, when the mutation occurred after the copy number amplification event (hence are considered “late” mutations). In this model all CNV events are considered to be clonal, however, in a more general model, CNV events can also be present in just a subset of tumor cells.- 94 - 1324629 IvlAttorney Docket No. 2013237-1500

[0233] In certain embodiments, a cellularity of a mutation, denoted by p, may be determined. Cellularity as used herein refers to the fraction of tumor cells that harbor a given mutation. A mutation is said to be biologically clonal if all cancer cells in a tumor sample harbor the given mutation. Biologically clonal mutations are characterized by having a nominal cellularity of 1 (p = 1). For example, clonal mutation 430 is present in all three tumor cells and, accordingly, has a nominal cellularity of 1, whereas subclonal mutation 450 is present in 1 out of 3 tumor cells, and, accordingly, has a nominal cellularity p = 1 / 3.

[0234] In certain embodiments, approaches described herein determine cellularity estimates for mutations. In certain embodiments, a cellularity estimate is an estimated mean cellularity (e.g., indicating, if multiple tumor samples were obtained and sequence, a given mutation is estimated to be present, on average, in a fraction of tumor cells given by the mean cellularity). In certain embodiments, confidence intervals for cellularity estimates may be determined, with lower and upper bounds denoted wT( ) and vY( ), where y is the confidence level of the estimate (e.g., also referred to as degree of confidence or confidence coefficient). A cellularity confidence interval (CI) [wT( ), vY( )], may, for example, indicate that, if multiple tumor samples were obtained, the estimated cellularity would be on the interval [e.g., at or between wY( ) and vY( )] y percent of the time. For example, in certain embodiments, a 90% CI lower and upper bound are determined. In certain embodiments, a 95% CI lower and upper bound are determined. In certain embodiments, a cellularity confidence interval may be determined as p±ao_p. For example, for a 95% C. I. a=1.96, for a 75% C. I a=1.15, for a 50% C. I a=0.67. The standard deviation of the cellularity can be calculated as (p(l-p) / N) where N can be the total coverage at the given site or the number reads mapping to the wildtype allele plus the number reads mapping to alternate allele, for example, after selecting only reads with a quality score above a predetermined threshold, for example, in the range of 25 to 35, and potentially after filtering certain poor quality reads. In certain embodiments, given that p=VAF / P_X, the standard deviation of the cellularity can be given by c_VAF / P_X, where c_VAF is the standard deviation of the observed variant allele frequency, given by (VAF(1-VAF) / N), and P_X is the expected allele frequency of the given variant allele assuming the mutation is clonal.B. Tumor Deconvolution and Mutation Detection- 95 - 1324629 IvlAttomey Docket No. 2013237-1500

[0235] Among other things, this application describes technologies for determining genetic and biophysical properties of tumor samples based on sequencing data (a procedure referred to, in certain cases, as “tumor deconvolution”).

[0236] As described in further detail herein, tumor deconvolution technologies of the present disclosure allow for complex characteristics of tumor genomes to be determined based on sequencing data. For example, cancer cells undergo extensive mutations, including copy number variation (CNV) events and single nucleotide variations (SNVs). As a result, unlike genomes extracted from normal cells, tumor genomes are not reliably diploid. Instead, the number maternal and paternal copies of genetic material varies from segment to segment, depending on the CNV events that took place over the lifetime of a given population of cancer cells. Additionally, cancer cells harbor collections of mutations at individual sites - SNV events - such as substitutions, insertions, deletions, etc.

[0237] Among other things, mutations, such as SNV events, that are found in tumor cells make valuable targets for therapeutics, including immunotherapies such as individualized cancer therapies. Accordingly, the ability to accurately and rapidly detect and characterize mutations is an important step in treating cancer patients.B.i Tumor Modeling and Mutation Detection Challenges

[0238] Accurate detection and characterization of tumor mutations, however, is a highly complex process. Among other things, as illustrated in FIG.5A, genetic and biophysical features of tumor samples may both impact particular tumor read counts and quantities, such as allele frequencies, that are derived therefrom and observed based on sequencing data.

[0239] For example, as explained above, tumor samples often include a fraction of normal, healthy cells. Read counts for particular mutations and / or portions of a tumor genome that are observed in sequencing data may be impacted by tumor sample purity. Accordingly, the certainty with which an event (e.g., a sequencing data event), such as a collection of reads with unique (e.g., abnormal) base calls at a particular site, can be determined to indicate true underlying cancer cell mutations depends on tumor sample purity. Additionally or alternatively, CNV events, which may increase or decrease relative amounts of genetic material - and thus- 96 - 1324629 IvlAttorney Docket No. 2013237-1500read counts - for particular portions (e.g., genes) of the tumor genome, may also impact how underlying, true, mutations manifest in observable sequencing data.

[0240] Accordingly, among other things, tumor deconvolution technologies of the present disclosure allow estimation of tumor sample purity and, in certain embodiments, characterization of CNV events across a tumor genome. As described in further detail herein, in certain embodiments, determining biophysical and genomic properties of a tumor sample in this manner can be used to improve accuracy with which cancer mutations are detected.

[0241] Additionally or alternatively, in certain embodiments, mutations may be prioritized as targets for immunotherapy, for example to select high value targets for inclusion in personalized cancer vaccines, T-cell therapies, and the like, according to genomic characteristics, such as whether they are determined to be biologically clonal or subclonal, their zygosity, etc. Accordingly, ability to characterize genomic properties of mutations themselves and / or portions of a tumor genome where they are located (e.g., copy numbers of segments harboring mutations) can be highly valuable in the context of therapeutic approaches.

[0242] Determining biophysical properties of tumor samples, such as their purity, and characterizing genomic properties, such as copy number variations across a tumor genome, in non-trivial. Among other things, both tumor sample purity and CNV events affect how observed sequencing reads (e.g., such as impacting relative read counts of various segments in a tumor genome).

[0243] Additionally or alternatively, the present disclosure and techniques described herein appreciate the fact that tumor genomes vary widely from patient to patient and from indication to indication in terms of the frequency of mutations, the frequency and extent CNV events, and the degree of subclonality of these genetic features. For example, the number of SNVs detected in an exome can vary across patients by 3 orders of magnitude, and across indications by at least 4 orders of magnitude from -1 (in certain pediatric cancers) up to -104 (Alexandrov, L. B. et al., 2013, Nature 500, 415-421, Lawrence, M. S. et al., 2013, Nature 499, 214-218). Similar diversity is observed for CNVs. For example, breast cancers can contain anywhere from thousands of CNV events and other structural variation events to nearly none. Likewise, lung squamous cell tumors can contain anywhere from hundreds of CNV events and other structural variation events to none. On the other hand, indications like kidney renal clear - 97 - 1324629 IvlAttorney Docket No. 2013237-1500cell carcinoma and medulloblastoma appear to contain significantly fewer structural variations (Yang, L. et al., 2013, Cell 153, 919-929, Network, C. G. A. R., 2013, Nature 499, 43-49, Parsons, D. W. et al., 2011, Science 331, 435-439). Adding to these difficulties is the fact that tumor sample purity is often, in practice, not high, and frequently even very low, leading to a reduction in signal to noise ratio of the genetic features sought to be estimated.

[0244] Accordingly, among other things, tumor deconvolution technologies of the present disclosure allow estimation of tumor sample purity from sequencing data, accurately accounting for the manner in which CNV events impact sequencing data observations such as read counts, variant allele frequencies (VAFs), and the like. Notably, tumor deconvolution methods and systems described herein leverage signal from SNV events themselves and allow for tumor sample purities to be determined without necessarily requiring direct characterization of CNV events. Accordingly, as described in further detail herein, even when characterization of CNV events is not feasible. As described in further detail herein, in certain embodiments, determining biophysical and genomic properties of a tumor sample in this manner can be used to improve accuracy with which cancer mutations are detected.B.i.a Degeneracy in SNV-based purity estimation

[0245] Turning to FIG.5B, in general, purity estimation is important for achieving optimal mutation detection. Purity estimation can be achieved, for example, by leveraging mutations present in a tumor genome or by leveraging CNV events present in a tumor genome, as these are markers that may be unique to a tumor sample and absent in a normal genome. Since tumor genomes are diverse, in some embodiments, a given tumor genome will contain sufficient SNVs or CNVs for purity estimation, in some embodiments, even if sufficient CNVs are present, a method for purity estimation based on these events may not be successful due to the complexity involved in solving the problem. Moreover, in some embodiments, (e.g., at low purities), it is no longer feasible to estimate purity based on CNVs; in such embodiments, methods of purity estimation based on mutations (e.g., SNV-based purity estimation methods) are a particularly useful alternative. Accordingly, the present disclosure provides methods and systems of SNV-based purity estimation which are in certain embodiments, useful alternatives to CNV-based purity estimation methods and systems (e.g., when CNV-based purity estimation - 98 - 1324629 IvlAttomey Docket No. 2013237-1500fails). However, the present disclosure identifies a problem that, without knowledge of copy numbers and zygosities of SNV s, an estimation of purity is mathematically not unique when using SNVs (e.g., SNV events) as markers.

[0246] To illustrate this point, Eq. 1 shows a purity of a tumor sample for a diploid site in a normal genome in the absence of noise given an observed variant allele frequency (VAF) of a SNV in the tumor sample:2 ■ VAFEq. (1) pur^y -2_yAF + CNmu^K_VAF)

[0247] In Eq. (1) CNmutis a copy number of a mutation in the tumor genome and K is a fractional zygosity of a mutation given by K = Cx / CNmut, wherein Cxis a zygosity of the mutation (number of mutated copies of a gene to which the mutations maps). To illustrate the problem of degeneracy identified by the present disclosure, consider a SNV measured to have a VAF of 0.2. FIG.5B shows possible purity values for various copy numbers and zygosities for an SNV, illustrating the point that without restriction of SNVs to a particular copy number and zygosity, purity is not unique.

[0248] Further, the present disclosure provides an insight that restricting SNVs to just diploid SNVs (SNVs with a copy number of 2) would not resolve the identified problem because the zygosity can still impact the estimated purity, as shown in the second and third rows of FIG.5B; restricting SNVs just to SNVs in balanced regions of a tumor genome would also not resolve the problem because, as shown in FIG.5B, balanced SNVs (rows highlighted in green in Table 1) also have degenerate purities; and restricting SNVs to balanced SNVs with a copy number of 4, corresponding to minimal balanced SNVs when a minimal copy number is equal to 4 (CNmut= 2, Cx= 1 or 2) also does not resolve the problem because here too different zygosities can yield different purity estimations (seventh and eighth row in in FIG.5B). Further, however, the present disclosure provides an insight that restricting SNVs to balanced SNVs in diploid regions (i.e., CNmut= 2, Cx= 1) would resolve the problem and yield a unique purity prediction since there is only one combination of copy numbers and zygosities for this subset of SNVs.- 99 - 1324629 IvlAttomey Docket No. 2013237-1500

[0249] Accordingly, methods and systems of the present disclosure are directed to resolving the identified problem by using SNVs in balanced diploid regions of a tumor genome (e.g., to perform SNV-based purity estimation), as under such conditions, a unique purity estimation can be obtained (e.g., under asymptotic sampling conditions). Accordingly, the present disclosure also provides methods and systems for identifying SNV events occurring in balanced diploid regions of a tumor genome without knowledge of copy numbers, and resolving situations wherein balanced diploid SNV events do not occur in a tumor genome. In certain embodiments, the present disclosure provides methods and systems of SNV-based purity estimation which utilize CNV-based purity estimation, (e.g., an CNV-based purity estimation which provides uncertain information), and which may improve CNV-based purity estimation.B.ii Identifying and Using Subpopulations of Mutations for Tumor Modeling

[0250] Turning to FIG.6, in certain embodiments, a tumor deconvolution process 600 may, among other things, construct and fit one or more tumor models that accurately explain sequencing data - such as aligned reads for tumor and, optionally, normal genome for a patient. As shown in FIG.6, in an example tumor modelling process 600, sequencing data may be obtained 602. As described herein, sequencing data 602 may comprise tumor sequencing data and / or normal sequencing data. In certain embodiments, sequencing data comprises tumor sequencing data and normal sequencing data (e.g., which may comprise replicates).

[0251] In certain embodiments, tumor modelling process 600 obtains a list of putative SNV events 604, with each putative SNV event representing one or more potential underlying mutations present in a tumor genome representative of the tumor sample. That is, as described herein, a SNV event refers to an event or anomaly in sequencing data, such as a collection of reads that reflect abnormal base calls (e.g., different from the normal base observed e.g., in normal sequencing data, e.g., obtained from the subject, or determined from a panel, or determined from a reference genome, etc.) at particular sites in a tumor genome, and thereby indicate the presence of a potential, underlying physical mutation within a cancer cells. A list of putative SNV events may include SNV events detected based on tumor sequencing data, for example via various mutation calling approaches whereby, for example, sequencing data is analyzed and sites determined to be likely to encode an alternate allele (e.g., nucleotide base),- 100 - 1324629 IvlAttorney Docket No. 2013237-1500different from a wild-type allele (e.g., nucleotide base) characteristic of a normal genome, identified as putative SNV events.

[0252] Tumor modeling process 600 may, in certain embodiments, identify and / or select a subset of SNV events 606 to use for purity estimation. Among other things, as described herein, quantities observed in sequencing data, such as a distribution of tumor read counts for various alleles, variant allele frequencies (VAFs), and the like, for SNV events may be impacted by factors such as the biological clonality or subclonality of a given SNV event as well as the copy number and zygosity of the SNV (e.g., the region in which the SNV is located).Accordingly, in certain embodiments, rather than attempt to model a complex constellation of possible underlying biological scenarios and how they present in sequencing data, tumor deconvolution approaches described herein identify and filter SNVs to obtain a selected subset determined to have, or to be likely to have, a particular, intentionally limited, underlying biology. In this way, a single, simplified tumor model that aims to capture a particular limited range of biological scenarios used and fit to a particular subset of sequencing data for which its assumptions are appropriate.

[0253] For example, in certain embodiments tumor modeling process 600 identifies and utilizes a subset of SNVs events that are determined to represent mutations that are balanced within a tumor genome and are located in segments that have a particular absolute copy number. In certain embodiments, the particular absolute copy number is a minimal copy number, whose exact value is unknown but predicted to be lower than copy numbers of other balanced segments (e.g., the exact value of the minimal copy number is initially unknown, but determined to be smaller relative to other absolute copy numbers of balanced segments). In certain embodiments, the particular absolute copy number is a primary copy number, whose exact value is unknown, but which is determined to be a most frequently occurring even-numbered absolute copy number (e.g., the exact value of the primary copy number is initially unknown, but is determined to occur more frequently than other even-numbered copy numbers within the tumor genome).B.ii.a Identifying Balanced Regions of a Tumor Genome

[0254] FIGs. 7-9 show example processes by which particular subsets of segments and / or SNVs may be identified and selected based on sequencing data.- 101 - 1324629 IvlAttomey Docket No. 2013237-1500

[0255] FIG. 7 shows an example process 700 by which particular segments within a tumor genome, such as balanced heterozygous segments, those having a minimal copy number, and / or a primary copy number, may be identified. FIGs.8 and 9 illustrate example approaches for identifying and selecting particular subpopulations of segments within a tumor genome. FIG. 8 is a schematic illustrating an example process for identifying and selecting a subset of primary balanced heterozygous segments and using the selected subset of primary balanced heterozygous segments to identify and select primary segments. As described in further detail below, primary balanced heterozygous segments are balanced heterozygous segments that are determined to have a same, particular, primary copy number that occurs more frequently than other even-numbered absolute copy numbers. Primary segments are identified as also having the primary copy number, but are not necessarily balanced and / or heterozygous within the tumor genome. FIG.9 is a schematic illustrating an example process for identifying and selecting a subset of minimal balanced heterozygous segments and using the selected subset of minimal balanced heterozygous segments to identify and select minimal segments. As described herein, e.g., in further detail below, minimal balanced heterozygous segments are balanced heterozygous segments that are determined to have a same, particular, minimal copy number that is determined to be less than other even-numbed copy numbers. Minimal segments are identified as also having the minimal copy number, but are not necessarily balanced and / or heterozygous within the tumor genome.

[0256] In certain embodiments, process 700 identifies and / or receives heterozygous SNPs (e.g., present in a subject’s normal genome) 702.

[0257] SNPs that are present in a normal genome of a subject (also referred to herein as “wild-type”) may be identified using normal sequencing data, for example via a commercial SNP or variant caller (e.g., Illumina DRAGEN, Qiagen Genomics, Torrent Variant Caller, etc.) based on sequence alignment to a reference and statistical modeling. Additionally or alternatively, an initial list of SNPs may be determined via orthogonal methods, such as SNP arrays or TaqMan assays. Additionally or alternatively, an initial list of SNPs may be obtained from a database, for example of potential SNPs present in an average individual of a particular population of which the subject is a member.- 102 - 1324629 IvlAttomey Docket No. 2013237-1500

[0258] In certain embodiments, normal sequencing data may be used to evaluate and filter an initial list of SNPs and / or heterozygous SNPs for consistency with an expected underlying model of a heterozygous SNP, e.g., a normal diploid genome with, for a given SNP, two alleles present in equal copy number. Accordingly, in certain embodiments, a balanced test may be used to determine whether sequencing data supports, with sufficiently high confidence, an underlying hypothesis that two alleles of a SNP are balanced - e.g., present in equal copy numbers. For example, statistical tests, such as a Z-statistic test, may be used to classify a particular SNP as balanced or unbalanced. SNPs or putative heterozygous SNPs in an initial list that do not meet a balanced test may be filtered and removed from the initial list of heterozygous SNPs.

[0259] In certain embodiments, heterozygous SNPs of the initial list may be assessed to determine whether sequencing data meets particular quality criteria. For example, each heterozygous SNP may be required to have a particular minimum coverage level in both normal and tumor sequencing data and / or a quality score above a particular minimum value.

[0260] In certain embodiments, heterozygous SNPs may be filtered according to the particular chromosome where they are located. For example, SNPs located on particular chromosomes may be excluded from analysis and removed from the list of heterozygous SNPs. In certain embodiments, only heterozygous SNPs located on particular chromosomes, such as chromosomes 1 through 22, are used.

[0261] In certain embodiments, a final list of heterozygous SNPs 704 produced in this manner may then be evaluated to identify SNPs that are balanced in the tumor genome and, in turn, a set of balanced heterozygous segments 706.

[0262] For example, in certain embodiments a balanced test as described above is performed for each SNP using tumor sequencing data, e.g., to determine whether a particular SNP is balanced in the tumor genome.

[0263] Accordingly, in certain embodiments, approaches described herein determine a balanced state of one or more heterozygous SNPs, for example in a tumor genome. In certain embodiments, a balanced state of a given heterozygous SNP is a value that indicates whether or not the given heterozygous SNP is balanced. For example, a balanced state may be a- 103 - 1324629 IvlAttomey Docket No. 2013237-1500classification, such as “balanced” or “unbalanced”, a Boolean value (e.g., True or False), a numeric value (e.g., 1 or 0), etc.

[0264] Turning to FIGs.8 and 9, example processes for determining primary balanced heterozygous segments and minimal balanced heterozygous segments, respectively, both begin by receiving a list of SNPs and determined balanced states identifying whether particular SNPs are balanced or not within the tumor genome. FIGs.8 and 9 show illustrative plots 830, 930 of balanced states for SNPs located the segments of the example tumor genome illustrated schematically in each figure. The balanced states shown in FIGs.8 and 9 are numeric values - 0 for unbalanced and 1 for balanced SNPs. For sake of simplicity, the figures only show one SNP per segment, but it should be understood that, as described herein, a particular segment may comprise one, or more than one, heterozygous SNP. As shown in FIG.8, in certain embodiments, balanced states are determined for heterozygous SNPs including SNPs within segments 819, 820, where a loss of heterozygosity (LOH) event has occurred (since it may not be known a priori whether a LOH event has occurred). FIG.9 also illustrates SNPs within segments 919, 920 where a LOH event has occurred, and for which balanced states are also determined. In certain embodiments, a balanced state is not determined for segments that do not comprise any heterozygous SNPs.

[0265] As illustrated in FIGs.8 and 9, statistical tests and / or other approaches for computing balanced states for SNPs may correctly reflect the true state of the tumor genome at a high level of accuracy, but errors may still occur. Plots 830, 930 show a balanced state for each heterozygous SNP as a red dot, such as 832, 932, which are examples of a heterozygous SNP determined to be balanced. In these examples, the balanced state shown in the figures reflects the correct balanced state for all heterozygous SNPs except for SNPs 834, 934, which are examples of errors that may occur when determining balanced states for individual SNPs.Balance state 834 is an example of an erroneous determination of the balanced property of heterozygous SNP 820. Balanced state 934 is an example of an erroneous determination of the balanced property of heterozygous SNP 920. Such errors can occur when using a statistical test, in particular when the purity of the tumor is low so that the distinction between the state of being balanced and the state of being unbalanced diminishes, and more so when the coverage of the given site is low.- 104 - 1324629 IvlAttomey Docket No. 2013237-1500

[0266] In certain embodiments, balanced states of one or more heterozygous SNPs may be contour filtered. For example, as shown in FIG. 8, a contour filter to a profile of balance states 836 may be determined. Contour filtering can correct errors in the determination of the balance state of a heterozygous SNP by identifying regions in the genome that have a high density of balanced and / or unbalanced heterozygous SNPs. For example, contour filtering based on the balance state of neighboring heterozygous SNPs determined that the heterozygous SNP 834 is in a locally unbalanced region, and therefore the balance state of heterozygous SNP 834 can be corrected to a contour-filtered balance state of 1 (unbalanced), wherein heterozygous SNPs are, for example, regarded as balanced if they are balance both before and after contour filtering. As a result, the heterozygous segment corresponding to 820 may be classified as unbalanced and hence discarded from further analysis. Contour filtering is also illustrated in FIG. 9, with contour 936 being used to determine that SNP 934 is in a locally unbalanced region, such that an initially determined balance state indicating that SNP 934 is balanced (e.g., a value of 0) may be updated to a value indicative of SNP 934 being unbalanced (e.g., a value of 1). Accordingly, in certain embodiments, contour filtering can be used to discard HSs wrongly classified as BHSs, in particular when the purity is low and more such errors are expected.

[0267] Turning again to FIG.7, in certain embodiments, balanced states of individual heterozygous SNPs can be used (e.g., aggregated) to determine segment-level balanced states, and thereby identify a set of balanced heterozygous segments 708. A particular segment may comprise one or more heterozygous SNPs. Accordingly, in certain embodiments, a balanced state of a particular segment may be determined based on balanced states of the one or more heterozygous SNPs that it comprises. For example, a balanced state for a particular segment may be determined as the mean (average), median, mode, etc. of the individual balanced states of the one or more heterozygous SNPs located within the particular segment. In certain embodiments, segment-level balanced states are (e.g., additionally or alternatively to individual SNP balanced states) contour filtered.

[0268] Accordingly, in certain embodiments, a set of balanced heterozygous segments may be identified by selecting segments with balanced states that satisfy certain criteria, such as having balanced states that are at and / or above a particular threshold value (e.g., greater than or equal to 0.5, greater than or equal to 0.6, greater than or equal to 0.7, greater than or equal to 0.8,- 105 - 1324629 IvlAttomey Docket No. 2013237-1500greater than or equal to 0.9; e.g., equal to 1; e.g., greater than 0.5, greater than 0.6, greater than 0.7, greater than 0.8, greater than 0.9).B.ii.b Decomposing and Identifying Segments According to Absolute Copy Number

[0269] In certain embodiments, a particular subset of BHS’s are identified. For example, as described herein within a tumor genome, various subpopulations of segments can be identified, for example based on whether or not they are balanced and / or their absolute copy number. Accordingly, tumor modeling approaches may identify and isolate a particular subpopulation of segments that (i) are balanced (heterozygous segments) and (ii) have a same particular copy number.

[0270] For example, in certain embodiments, once a set of balanced heterozygous segments are identified 708, sequencing data corresponding to (e.g., tumor and / or normal reads that map to) members of the set of balanced heterozygous segments can be used to identify one or more subpopulations thereof (e.g., of the set of balanced heterozygous segments) 710, each subpopulation of BHSs comprising (e.g., only) segments having a same absolute copy number. Since subpopulations are identified within the set of balanced heterozygous segments, each subpopulation is expected to have an even-numbered absolute copy number (e.g., by virtue of the segments all being balanced). Subsets of balanced heterozygous segments corresponding to one or more particular, desired subpopulations may then be selected 712, for example to select segments associated with a most frequently occurring (e.g., even numbered) absolute copy number (e.g., a primary copy number), such as a subset of primary balanced heterozygous segments 714a and / or a minimal (e.g., even numbered) copy number, less than other (e.g., even numbered) absolute copy numbers of other segment subpopulations, such as a set of minimal balanced heterozygous segments 714b.

[0271] For example, in certain embodiments, tumor and normal read counts for each segment can be used to determine a probability density function (pdf) for values of a tumor-to-normal read count ratios 720. For example, a tumor-to-normal read count ratio can be determined for a particular segment as a ratio of tumor read counts to normal read counts for the particular segment (e.g., tumor read count for the particular segment divided by the normal read - 106 - 1324629 IvlAttomey Docket No. 2013237-1500count for the particular segment). Tumor-to-normal read count ratios may be determined for a plurality of segments, such as BHS’s, to determine a distribution of tumor-to-normal read count ratios for the plurality of segments (e.g., set of BHSs) 720. A distribution may be represented via a pdf (e.g., via creation of a histogram, or other approaches) 722. As illustrated in FIGs. 8 and 9, distribution 722 may comprise multiple components, each corresponding to a particular subset of segments having a different absolute copy number.

[0272] For example, FIG.8 shows a graph of an illustrative pdf 840 constructed based on the tumor-to-normal read count ratios for a set of balanced heterozygous segments. As illustrated in FIG.8, pdf 840 comprises multiple components (e.g., normal-like components), appearing as distinct peaks within the plot. Each component reflects a subpopulation of segments having a particular, different, absolute copy number.

[0273] Without wishing to be bound to any particular theory, FIG.8 illustrates a connection between underlying variations in copy numbers from segment to segment and components of pdf 840. For example, BHS’s 813, 814, 815, 816 each have an absolute copy number of four (4) and correspond to component 844 (e.g., centered around a higher mean tumor-to-normal read count ratio), while BHS’s 810, 811, 812 each have an absolute copy number of two (2) and correspond to component 846 (e.g., centered around a lower mean tumor-to-normal read count ratio).

[0274] FIG. 9 illustrates a connection between a measured tumor-to-normal read count ratio pdf 940 and underlying variations in copy number from segment to segment as well. In FIG. 9, pdf 940 has two peaks, or components, with a lower mean (left-most) subcomponent corresponding 944 to segments 913, 914, 915, and 916, which are balanced diploid segments, and a higher mean component corresponding 946 segments 910, 911, and 912 each have a copy number of four (4) and, accordingly, correspond to a higher mean peak in pdf 940.

[0275] Accordingly, in certain embodiments, an empirical pdf providing a distribution of tumor-to-normal read count ratios for balanced heterozygous segments may be decomposed into one or more subcomponents 724. Individual components may, accordingly, be distinguished from each other and particular desired subcomponents selected 726. Subsets of segments that are associated with a particular subcomponent may, accordingly, be identified and selected 712.- 107 - 1324629 IvlAttorney Docket No. 2013237-1500

[0276] For example, turning again to FIGs. 8 and 9, in certain embodiments, pdfs such as pdf 840 and pdf 940 can be decomposed into different normal-like components 850 and 950, respectively via a decomposition algorithm, such as a fit to a model pdf. For example, in certain embodiments a one-dimensional Gaussian mixture model (1D-GMM) can be used to approximate and be fit to a pdf such as pdfs 840, 940. Parameters determined via the fit - e.g., an amplitude, mean, and standard deviation of each constituent Gaussian - may, accordingly, be used to select a particular component, such as a primary component that corresponds to a largest subpopulation of underlying segments (e.g., accordingly, a most frequently occurring absolute copy number) and / or a minimal component that corresponds to a subpopulation of segments with a minimal absolute copy number that is smaller than other even-numbered copy numbers, and, accordingly, likely to be two.Selecting Primary Balanced Heterozygous Segments and / or Primary Segments

[0277] Turning to FIG. 8, in certain embodiments, an amplitude of each component of a pdf, such as pdf 840, may be determined and the component having the largest amplitude selected as the primary component 854. In this way, in certain embodiments, the component identified as the primary component corresponds to a largest subpopulation of segments, e.g., having a most frequently occurring even-numbered absolute copy number.

[0278] In certain embodiments, a mean tumor- to-normal read count ratio of each component may be determined and used to select the primary component, e.g., alone or in conjunction with the component amplitudes. For example, in certain embodiments, a component having the lowest mean tumor-to-normal read count ratio may be identified, and its amplitude compared with the maximum amplitude (across all components). In certain embodiments, if an amplitude of the component having the lowest mean tumor-to-normal read count ratio is at least (e.g., at or above) a particular minimum faction of the maximum amplitude, then the lowest mean ( ) component is selected as the primary component.

[0279] Among other things, selecting a particular subset of segments - e.g., balanced heterozygous segments - from which to construct tumor-to-normal read count pdf facilitates its decomposition into multiple (e.g., normal-like) components. Among other things, limiting the pdf to balanced heterozygous segments eliminates segments with odd copy numbers, so that the - 108 - 1324629 IvlAttomey Docket No. 2013237-1500individual components (e.g., peaks) of the distribution are well separated. Accordingly, pdf is amendable to robust decomposition even for low purity tumor samples. Moreover, by selecting a particular subset of balanced heterozygous segments, the range of possible copy numbers of the selected subset (e.g., primary balanced heterozygous segments) is restricted to even integers, thereby reducing the possible candidate primary copy numbers by a factor of two.

[0280] In certain embodiments, one or more auxiliary parameters may be determined from the primary component 854 of pdf 840. In certain embodiments, a mean of the primary component, referred to herein as a “primary slope,” (denoted rpc) may be determined and used as an auxiliary parameter. In certain embodiments, a standard deviation of the primary component, referred to herein as the “primary standard deviation,” (denoted opc) may be determined. In certain embodiments, an amplitude of the primary component, denoted Ape, may be determined. In certain embodiments, as described in further detail herein, values of one or more auxiliary parameters, including, but not limited to, a primary slope (rpc) and / or primary standard deviation (opc) may be used to determine a purity estimate (e.g., a CNV-based purity estimation).

[0281] In certain embodiments, a set of primary balanced heterozygous segments are selected 860 from the set balanced heterozygous segments. Parameters 852 characterizing each component of the balanced heterozygous segment tumor-to-normal read count pdf 840, for example individual component amplitudes (Ai), means ( ), and standard deviations (oi), may be used to select a subset of primary balanced heterozygous segments. For example, in certain embodiments, a subset of primary balanced heterozygous segments may be selected from a set of balanced heterozygous segments utilizing, for example, the principle of maximum probability based on the parameters 852 obtained from the decomposition of the pdf. For example, in certain embodiments a maximum a-posteriori probability (MAP) approach can be used to determine a likelihood of a particular heterozygous segment belonging to a particular one of the one or more components of pdf 840. Segments having a highest likelihood for belonging to the primary component may, accordingly, be selected for inclusion in the set of primary balanced heterozygous segments.

[0282] In certain embodiments, after primary balanced heterozygous segments are identified, the set of identified primary balanced heterozygous segments 824 may be used estimate additional auxiliary parameters and / or refine values of those determined (e.g., initially)- 109 - 1324629 IvlAttorney Docket No. 2013237-1500based on decomposition and identification of primary component 826. For example, in certain embodiments, the set of identified primary balanced heterozygous segments 824 may be used to determine a primary allele frequency standard deviation 870. In certain embodiments, the set of identified primary balanced heterozygous is used to determine a primary residual standard deviation 872.

[0283] For example, in certain embodiments, a primary allele frequency standard deviation may be determined as a standard deviation of observed allele frequencies across the set of primary balanced heterozygous segments 860.

[0284] In certain embodiments, a residual error may be computed for a particular segment to compare (e.g., measure a measure of a difference between) (z) an observed tumor read count for the segment and (zz) an expected tumor read count for the segment predicted based on its normal read count. For example, for a particular, j-th segment, an observed tumor read count zj may be compared with a predicted tumor read count, tj, that is computed according to Eq. (2), below - e.g., as a linear prediction based on that segment’s observed normal read count.Eq. (2) tj = a • nj

[0285] The parameter a in Eq. (2) unknown but expected to depend on an absolute copy number of a given segment.

[0286] In certain embodiments, for the set of primary balanced heterozygous segments, an initial estimate of the value of a can be approximated as the mean of the primary component -e.g., the primary slope, rpc (e.g., since / 'pc is the average tumor-to-normal read count ratio for the subpopulation of segments represented by the primary component).

[0287] In certain embodiments, a particular functional form of a residual error may be based on a variance stabilizing transform that is known or expected to produce a particular distribution of values. For example, in certain embodiments, a residual error may be determined based on observed and expected (e.g., a linear prediction of) tumor read counts accordingly to a variance stabilizing transform such that residual errors are normally distributed when computed - 110 - 1324629 IvlAttorney Docket No. 2013237-1500across the set of primary balanced heterozygous segments. For example, as demonstrated in Example 1, in certain embodiments, a difference of square root functions may be used as a variance stabilizing transform, with residual error computed as:Eq. (3) Residual ErrorEq. (3)

[0288] Other variance stabilizing transformations may be used, including, without limitation a logarithmic transformation (e.g., e = log (t) — log (an)), an arc-sine square root transformation (e.g., e = arcsin( t) — arcsin (Van)), a reciprocal transformation (e.g., e = t-1— (an)-1), an exponential transformation (e.g., e = exp(t) — exp(an)), a box-Cox Transformation (e.g.,T*1)- for0; and log(T), for X = 0, where parameter Z. is chosen to best stabilize variance and approximate normality), an Anscombe transform, as well as other possible transformations.

[0289] In certain embodiments, residual errors are determined for each segment of the set of primary balanced heterozygous segments, thereby determining a distribution of primary residual errors 874. A standard deviation may be determined, for example, by taking advantage of the expectation that the primary residual errors are normally distributed, by virtue of the variance stabilizing transformation. For example, in certain embodiments, a Gaussian fit to the distribution of primary residual errors may be used with, for example, the standard deviation extracted from a best fit (e.g., by minimizing a least squares error between an empirical probability distribution function of the primary residual error, q, and a Normal distribution with an unknown standard deviation).

[0290] In certain embodiments, a standard deviation of primary residual errors (e) 872 may be used, in turn, to identify, from all segments 880 (e.g., not just balanced heterozygous segments) a primary segment subpopulation that comprises those segments (e.g., whether heterozygous or not, balanced or not) having an absolute copy number equal to the primary copy number. For example, in certain embodiments, this is performed by selecting segments for - Ill - 1324629 IvlAttorney Docket No. 2013237-1500which the absolute of the primary residual standard deviation 880 is small in proportion to the primary residual standard deviation 872. In certain embodiments, since the primary residual errors 880 are normally distributed, selection of primary segments benefits from rigorous statistical tolerance.

[0291] Primary segments may, in turn, be used to refine the set of primary balanced heterozygous segments 860, which may be used for CNV-based purity estimation approaches such as those described in Example 1. For example, once primary segments have been precisely estimated based on a rigorous statistical distribution, primary segments can be used to further refine the primary balanced heterozygous segments by checking that each primary balanced heterozygous segment of the initially determined set (e.g., based on an MAP approach and parameters of a 1D-GMM decomposition) also appears in the set of primary segments 890. A refined subset of primary balanced heterozygous segments can be used to refine the primary slope 892 and / or the primary allele frequency standard deviation 894.

[0292] In certain embodiments, as described in further detail herein, the set of primary balanced heterozygous segments and / or auxiliary parameter values determined therefrom may be used to determine various parameters used in tumor models, which are, in turn, fit to observed sequencing data. For example, in certain embodiments, auxiliary parameters estimated from primary balanced heterozygous segments may be used to approximate (e.g., fixed) values of certain parameters in tumor models used for (e.g., CNV-based) purity estimation, e.g., thereby reducing a number of variable parameters in tumor models to be fit to observable distributions from sequencing data. In certain embodiments, auxiliary parameters estimated from primary balanced heterozygous segments may be used as initial estimates of values of certain parameters in tumor models used for (e.g., CNV-based) purity estimation, e.g., so as to provide an accurate starting point for a fitting procedure. In certain embodiments, fitting procedures may refine a parameter value, updating it from its initial value taken from an auxiliary parameter. In certain embodiments, as described in further detail herein, tumor model parameters may be established relative to the primary copy number that characterizes the set of primary balanced heterozygous segments, such as rather than estimate and / or fit multiple copy number values, a single, primary, copy number is determined via tumor mode fitting. Among other things, in this way, identifying primary balanced heterozygous segments and determining auxiliary parameters therefrom- 112 - 1324629 IvlAttorney Docket No. 2013237-1500simplifies and / or makes tractable complex tumor modeling procedures and allows for accurate purity estimates and absolute copy number determinations. As described in further detail herein, CNV-based purity estimation approaches that utilize a subset of primary balanced heterozygous segments may be complementary to, and used additionally or alternatively to SNV-based purity estimation techniques described herein. In certain embodiments, additionally or alternatively, VAF analysis methods described herein may be used to determine a minimal and / or primary copy number value and / or a lower bound thereon, which, in turn, can be used to inform and improve CNV-based purity estimation techniques which assign absolute copy number values to segments throughout the tumor genome.Selecting Minimal Balanced Heterozygous Segments and / or Minimal Segments

[0293] Turning to FIG. 9, as with procedures for selecting and refining a subset of primary balanced heterozygous segments and / or primary segments described above in regard to FIG. 8, a subset of minimal balanced heterozygous segments and / or minimal segments can be identified and selected using a decomposition of an empirical pdf representing a distribution of tumor read counts (e.g., tumor read counts normalized to corresponding normal read counts, as in a tumor-to-normal read count distribution) 720.

[0294] As shown in FIG. 9, in certain embodiments, a mean (n) of each component of a pdf, such as pdf 940, may be determined and the component having the lowest mean selected as the minimal component 954. In this way, in certain embodiments, the component identified as the minimal component corresponds to a subpopulation of segments having a lowest even numbered absolute copy number - a minimal copy number. As described herein, since normal cells are typically diploid, this minimal copy number is typically two.

[0295] As described herein, among other things, selecting a particular subset of segments - e.g., balanced heterozygous segments - from which to construct tumor-to-normal read count pdf facilitates its decomposition into multiple (e.g., normal-like) components by eliminating segments with odd copy numbers, such that the individual components (e.g., peaks) of the distribution are well separated.- 113 - 1324629 IvlAttorney Docket No. 2013237-1500

[0296] In certain embodiments, one or more auxiliary parameters may be determined from the minimal component 954 of pdf 940. In certain embodiments, a mean of the minimal component, referred to herein as a “minimal slope,” (denoted rmin) may be determined and used as an auxiliary parameter. In certain embodiments, a standard deviation of the primary component, referred to herein as the “minimal component standard deviation,” (denoted crmin) may be determined. In certain embodiments, an amplitude of the minimal component, denoted Amin, may be determined.

[0297] In certain embodiments, a set of minimal balanced heterozygous segments are selected 960 from the set balanced heterozygous segments. As with selection of primary balanced heterozygous segments, parameters 952 characterizing each component of the balanced heterozygous segment tumor-to-normal read count pdf 940, for example individual component amplitudes (Ai), means ( ), and standard deviations (cq), may be used to select a subset of minimal balanced heterozygous segments. For example, in certain embodiments, a subset of minimal balanced heterozygous segments may be selected from a set of balanced heterozygous segments utilizing, for example, the principle of maximum probability based on the parameters 952 obtained from the decomposition of the pdf. For example, in certain embodiments a maximum a-posteriori probability (MAP) approach can be used to determine a likelihood of a particular heterozygous segment belonging to a particular one of the one or more components of pdf 940. Segments having a highest likelihood for belonging to the minimal component may, accordingly, be selected for inclusion in the set of minimal balanced heterozygous segments.

[0298] In certain embodiments, after minimal balanced heterozygous segments are identified, the set of identified minimal balanced heterozygous segments 924 may be used estimate additional auxiliary parameters and / or refine values of those determined (e.g., initially) based on decomposition and identification of minimal component 926.

[0299] In certain embodiments, the set of identified minimal balanced heterozygous is used to determine a minimal residual standard deviation 972. In particular, as described above with regard to primary balanced heterozygous segments, a residual error may be computed for a particular segment to compare (e.g., measure a measure of a difference between) (z) an observed tumor read count for the segment and (zz) an expected tumor read count for the segment predicted- 114 - 1324629 IvlAttorney Docket No. 2013237-1500based on its normal read count. As described above, this may be accomplished using a linear predictor, for example as in Eq. (2).

[0300] In certain embodiments, for the set of minimal balanced heterozygous segments, an initial estimate of the value of a in Eq. (2) can be approximated as the mean of the minimal component - e.g., the minimal slope, rmin(e.g., since rminis the average tumor-to-normal read count ratio for the subpopulation of segments represented by the minimal component). As described above in regard to Eq. (3), in certain embodiments, residual errors may be computed through the use of a variance stabilizing transform that is known or expected to produce a particular distribution of values.

[0301] In certain embodiments, residual errors are determined for each segment of the set of minimal balanced heterozygous segments, thereby determining a distribution of minimal (component) residual errors 974. A standard deviation may be determined, for example, by taking advantage of the expectation that the minimal residual errors are normally distributed, by virtue of the variance stabilizing transformation. For example, in certain embodiments, a Gaussian fit to the distribution of primary residual errors may be used with, for example, the standard deviation extracted from a best fit (e.g., by minimizing a least squares error between an empirical probability distribution function of the primary residual error, q, and a Normal distribution with an unknown standard deviation).

[0302] In certain embodiments, a standard deviation of minimal residual errors ( < ) 972 may be used, in turn, to identify, from all segments 980 (e.g., not just balanced heterozygous segments) a minimal segment subpopulation that comprises those segments (e.g., whether heterozygous or not, balanced or not) having an absolute copy number equal to the minimal copy number. For example, in certain embodiments, this is performed by selecting segments for which the absolute of the minimal residual standard deviation 980 is small in proportion to the minimal residual standard deviation 972. In certain embodiments, since the minimal residual errors 980 are normally distributed, selection of minimal segments benefits from rigorous statistical tolerance.

[0303] Minimal segments may, in turn, be used to refine the set of minimal balanced heterozygous segments 960, which may be used for SNV-based purity estimation approaches such as those described in Example 1. For example, once minimal segments have been- 115 - 1324629 IvlAttorney Docket No. 2013237-1500precisely estimated based on a rigorous statistical distribution, minimal segments can be used to further refine the minimal balanced heterozygous segments by checking that each primary balanced heterozygous segment of the initially determined set (e.g., based on an MAP approach and parameters of a 1D-GMM decomposition) also appears in the set of minimal segments 990.B.ii.c Identifying Minimal Balanced SNVs

[0304] Turning again to FIG.6, in certain embodiments, subsets of minimal balanced heterozygous segments and / or minimal segments may be used to identify a subset of minimal balanced SNV events that can be used for a SNV-based purity estimation procedures, such as process 600.

[0305] In certain embodiments, a subset of minimal balanced SNV events may be identified from SNV list 604 using a subset of minimal balanced heterozygous segments. For example, in certain embodiments, SNV events of SNV list 604 that are located within a minimal balanced heterozygous segment can be identified as minimal balanced SNV events.

[0306] In certain embodiments, a subset of minimal balanced SNV events may be identified from SNV list using a subset of minimal segments followed by one or more filtering steps to identify SNV events that are located within balanced regions of a tumor genome. For example, in certain embodiments, SNV events of SNV list 604 that are located within a minimal segment can be identified. SNV events determined to be located within a minimal segment may then be evaluated to identify those that are within balanced regions and, accordingly, are minimal balanced SNV events.

[0307] SNVs may be determined to be balanced using various balanced tests and / or filtering (e.g., contour filtering) approaches described herein, for example in Section B.ii.a. In certain embodiments, a SNV may be identified as balanced based on one or more nearby SNPs. For example, in certain embodiments, a balanced state may be determined for one or more nearby SNPs using statistical tests and / or filtering, such as contour filtering, approaches. SNVs located within a certain distance of one or more balanced SNPs, in between two balanced SNPs, etc. may accordingly be identified as balanced and selected for inclusion in a set of minimal balanced SNV events. In certain embodiments, SNV’s selected for inclusion in a set of minimal- 116 - 1324629 IvlAttorney Docket No. 2013237-1500balanced SNV events may also be determined to be balanced in a normal genome, which may be a reference genome obtained e.g., from a database and / or may be determined via sequencing data, such as normal sequencing data obtained from a normal sample obtained from the subject. Other filtering approaches, such as coverage and / or quality score filters may also be used.Additionally or alternatively, once an initial subset of minimal balanced SNV events is selected, outliers associated with unbalanced SNV events (e.g., that were not removed by filtering approaches used to obtain the initial subset of minimal balanced SNV events) may be identified and removed, e.g., via VAF filtering approaches described herein, for example in Section B.iv.B.iii SNV-Based Tumor Modeling Using Minimal Balanced Heterozygous Segments

[0308] Turning again to FIG.6, in certain embodiments, tumor deconvolution technologies described herein use a particular subset of SNV events 606, such as a subset of minimal balanced SNV events, to estimate a purity of a tumor sample.

[0309] For example, in certain embodiments, one or more sequencing data observations are determined for the particular subset of SNV events 608. For example, quantities such as allele- specific read counts, variant allele frequencies, etc. and / or distributions thereof may be determined for a subset of minimal balanced SNV events. For example, allele- specific read counts may be determined for each of one or more alleles of a given site at which a putative SNV event occurs. For example, a SNV event may represent an underlying single nucleotide substitution at a particular site within a tumor genome, wherein a normal, wild-type, allele has a first nucleotide at the particular site and an alternate, mutated, allele has a second, different, nucleotide at the particular site. Allele specific read counts may be determined for the normal and alternate alleles as the number of reads having base calls corresponding to the first and second nucleotides, respectively, at the particular site. In certain embodiments, allele- specific read counts are also determined for two remaining noise alleles. Additionally or alternatively, a VAF may be determined for a particular allele as a fraction of reads corresponding to the alternate allele (e.g., relatively to all reads mapping to the particular site). In certain embodiments, sequencing data observations, such as allele- specific read counts and / or VAFs may be determined using a filtered portion of sequencing data reads, for example, to exclude those with coverage levels and / or quality scores below particular threshold values.- 117 - 1324629 IvlAttorney Docket No. 2013237-1500

[0310] Measured sequencing data observations 608 may be used in connection with a tumor model 612 to determine 610 one or more purity estimates. For example, a tumor model 612 may be used to determine predicted values and / or distributions of the sequencing data observations as a function of one or more variable parameters. Values and / or distributions of tumor model predictions can then be compared, and fit to, those of the measured sequencing data observations. The one or more variable parameters may include a tumor sample purity variable parameter, such that tumor model predictions of sequencing data observations can be fit to measured sequencing data observations by adjusting a value of the tumor sample purity parameter. In this way, an estimate of tumor sample purity may be obtained by determining the value of the variable tumor sample purity that optimizes the fit between the predicted sequencing data observations of the tumor model and the corresponding measured sequencing data observations.

[0311] For example, a tumor model may be a SNV model 612 that predicts frequencies with which different alleles - e.g., an alternate (e.g., mutated) allele and a normal (e.g., wildtype) allele, as well as, in certain embodiments, noise alleles - are expected to be observed in sequencing measurements. An SNV model 612 may, for example, assume mutations are biologically clonal and occur in balanced diploid regions of a tumor genome and predict, as a function of tumor sample purity and, optionally, a sequencing error rate, frequencies with which normal, alternate, and noise alleles are expected to be observed in sequencing data. Agreement between predictions made by an SNV model 612 and measured sequencing data observations 608 can be quantified using various metric functions. For example, in certain embodiments, a likelihood function (e.g., using a maximum likelihood-based approach) 616a that measures, for each SNV event, a likelihood of measuring a corresponding set of allele- specific read counts as a function of tumor sample purity. An overall likelihood function may then be determined based on an aggregation (e.g., a sum, a product) of all individual SNV likelihood functions for the subset of minimal balanced SNV events. In certain embodiments, additionally or alternatively, SNV model may be used to generate, as a function of tumor sample purity, a predicted distribution of VAFs for the set of minimal balanced SNVs, which can be compared with the measured VAF distribution 616b. Various metrics, including, but not limited to, a Kullback-Leibler (KL) divergence function, may be used to quantify agreement between predicted and measured VAF distributions. Accordingly, values of the tumor sample purity parameter can be - 118 - 1324629 IvlAttorney Docket No. 2013237-1500tested and a best-fit value that maximizes or minimizes one or more metric functions determined 610. Various procedures may be used to test and search for best-fit values of the tumor sample purity parameter using a given metric function, including, without limitation, grid searching, gradient descent, and the like.

[0312] Example 1 demonstrates use of maximum likelihood estimation (MLE) and KL-divergence-based approaches for determining tumor sample purities.

[0313] In certain embodiments, SNV models may allow for multiple SNV genotypes, rather than e.g., restrict SNV events to a single genotype. For example, in a diploid...

Claims

1. Attomey Docket No. 2013237-1500CLAIMSWhat is claimed is:

1. A method, the method comprising:(a) obtaining, by a processor of a computing device, tumor sequencing data for a tumor sample obtained from a subject;(b) obtaining, by the processor, a list of putative single nucleotide variant (SNV) events detected based on the tumor sequencing data, each putative SNV event representing one or more potential underlying mutations occurring in a tumor genome of the subject;(c) identifying, of the list of putative SNV events, a subset of minimal balanced SNV events using the tumor sequencing data, the subset of minimal balanced SNV events comprising SNV events determined to represent mutations that:(i) are determined to be balanced within the tumor genome; and (ii) are located within segments determined to have an absolute copy number equal to a same particular minimal copy number whose exact value is unknown but predicted to be lower than copy numbers of other balanced segments;(d) determining, by the processor, an estimated tumor sample purity and / or bound thereon based at least in part on the tumor sequencing data and the subset of minimal balanced SNV events; and(e) storing and / or providing, by the processor, the estimated tumor sample purity and / or bound thereon for display and / or further processing.

2. The method of claim 1, comprising using the estimated tumor sample purity and / or bound thereon to detect and / or prioritize a plurality of somatic mutations.- 295 - 1324629 IvlAttomey Docket No. 2013237-15003. The method of claim 1, comprising using at least a portion of the detected and / or prioritized somatic mutations in a personalized immunotherapy and / or using the plurality of somatic mutations and / or prioritization thereof to determine a cancer status for the subject.

4. The method of claim 1 to 3 comprising using the plurality of somatic mutations and / or prioritization thereof to select a therapy for the subject.

5. The method of any one of the preceding claims, comprising using the estimated tumor sample purity and / or bound thereon to determine a cancer status for the subject.

6. The method of claim 1 to 5 comprising using the plurality of somatic mutations and / or prioritization thereof to select a therapy for the subject.

7. The method of any one of the preceding claims, comprising using the estimated tumor sample purity and / or bound thereon as a quality control.

8. The method of any one of the preceding claims, comprising:identifying, by the processor, a set of balanced heterozygous segments (BHSs) within the tumor genome of the subject, each BHS of the set having been determined to comprise one or more heterozygous SNPs, at least a portion of which are balanced within the tumor genome; and selecting, by the processor, a subset of the set of BHSs as minimal balanced heterozygous segments (MBHSs) having the same particular minimal copy number, thereby identifying the plurality of MBHSs.

9. The method of claim 8, comprising,- 296 - 1324629 IvlAttorney Docket No. 2013237-1500determining, by the processor, for each of at least a portion of the set of BHSs, a measure of tumor read counts based at least in part on the tumor sequencing data, thereby determining a distribution of the measure of tumor read counts for set of BHSs;identifying, by the processor, one or more components of the distribution, each of the one or more subcomponents representing a subpopulation of BHSs having a particular absolute copy number, distinct from that of other subpopulations represented by other components of the distribution; andselecting, by the processor, based on the identified components of the distribution, the subset of MBHSs.

10. The method of claim 9, wherein the method comprises obtaining, by the processor, normal sequencing data for a normal sample obtained from the subject, and wherein the measure of tumor read counts is a tumor- to-normal read count ratio determined, for a given BHS, based on a number of tumor reads mapping to the given BHS and a number of normal reads mapping to the given BHS.

11. The method of any one of claims 8 - 10, comprising:determining, by the processor, for each of at least a portion of the set of BHSs, a tumor-to-normal read count ratio based on the tumor sequencing data and the normal sequencing data, thereby determining a tumor- to-normal read count ratio distribution for the set of BHSs;identifying, by the processor, one or more components of the tumor- to-normal read count ratio distribution, each of the one or more subcomponents representing a subpopulation of BHSs having a particular absolute copy number, distinct from that of other subpopulations represented by other components of the tumor-to-normal read count ratio distribution; andselecting, by the processor, based on the identified components of the tumor-to-normal read count ratio distribution, the subset of MBHSs.

12. The method of any one of claims 8 - 11, comprising:- 297 - 1324629 IvlAttomey Docket No. 2013237-1500identifying, by the processor, of the one or more components, a minimal component having a mean value of the measure of tumor read counts lower than that of other components; andselecting, by the processor, BHSs determined to be members of the subpopulation represented by the minimal component for inclusion in the subset of MBHSs.

13. The method of claim 12, wherein identifying the minimal component comprises:fitting, by the processor, a mixture model to the distribution, thereby determining, based on the fit, a plurality of parameter values, including, for each of the one or more components, a corresponding mean; andidentifying the minimal component as having a lowest corresponding mean.

14. The method of any one of the preceding claims, comprising selecting, by the processor, a subpopulation of minimal segments identified as also having the same particular minimal copy number.

15. The method of claim 14, comprising refining, by the processor, the set of MBHSs using the subpopulation of minimal segments.

16. The method of claim 14 or 15, wherein step (d) comprises selecting, by the processor, from the list of putative SNV events, a subset of minimal SNV events comprising putative SNV events determined to map to at least one of the one or more minimal segments.

17. The method of claim 16, comprising identifying the subset of minimal balanced SNV events from the subset of minimal SNV events.- 298 - 1324629 IvlAttomey Docket No. 2013237-150018. The method of claim 17, comprising filtering, by the processor, the subset of minimal SNV events to retain SNV events determined to be located in locally balanced regions of the tumor genome, thereby identifying the subset of minimal balanced SNV events.

19. The method of claim 17 or claim 18, comprising selecting, by the processor, for inclusion in the set of minimal balanced SNV events, one or more of the subset of minimal SNV events determined to map to one or more balanced heterozygous segments.

20. The method of any one of claims 8 to 19, wherein step (d) comprises selecting, by the processor, from the list of putative SNV events, SNV events determined to map to at least one of the one or more minimal balanced heterozygous segments (MBHSs) for inclusion in the set of minimal balanced SNV events.

21. The method of any one of the preceding claims, comprising:determining, by the processor, for each SNV event of the subset of minimal balanced SNV events, a variant allele frequency (VAF), thereby determining a VAF distribution for the subset of minimal balanced SNV events;identifying, by the processor, within the VAF distribution for the subset of minimal balanced SNV events, a plurality of VAF subpopulations, each corresponding to and representing VAFs of a distinct SNV subpopulation;identifying, by the processor, from the plurality of components, a main VAF subpopulation corresponding to a desired subpopulation of clonal heterozygous SNVs located in balanced diploid regions of the tumor genome;filtering, by the processor, the subset of minimal balanced SNVs based on the VAF distribution to retain and / or enrich for SNVs corresponding to the main VAF subpopulation, thereby determining a fdtered subset of minimal balanced SNV events; andusing, by the processor, the filtered subset of minimal balanced SNV events to determine the estimated tumor sample purity at step (e).- 299 - 1324629 IvlAttomey Docket No. 2013237-150022. The method of claim 21, wherein:the plurality of VAF subpopulations comprises a low VAF subpopulation corresponding to a subpopulation of subclonal SNVs having VAFs below those of the main VAF subpopulation, andfiltering the subset of minimal balanced SNVs comprises identifying and excluding SNVs of the subclonal subpopulation (e.g., based on their VAFs in relation to one or more parameters of the low VAF subpopulation and / or the main VAF subpopulation].

23. The method of claim 22, comprising confirming the main VAF subpopulation and the low VAF subpopulation are not harmonics.

24. The method of any one of claims 21 to 23, wherein:the plurality of VAF subpopulations comprises a high VAF component corresponding to a subpopulation of unbalanced SNVs, andfiltering the subset of minimal balanced SNVs comprises identifying and excluding SNVs of the unbalanced subpopulation (e.g., based on their VAFs in relation to one or more parameters of the high VAF subpopulation and / or the main VAF subpopulation].

25. The method of any one of claims 21 to 24, wherein:the plurality of VAF subpopulations comprises a low VAF subpopulation corresponding to a subpopulation of high copy number and low zygosity SNVs, andfiltering the subset of minimal balanced SNVs comprises identifying and excluding SNVs of the high copy number and low zygosity subpopulation (e.g., based on their VAFs in relation to one or more parameters of the low VAF subpopulation and / or the main VAF subpopulation].- 300 - 1324629 IvlAttorney Docket No. 2013237-150026. The method of any one of claims 21 to 25, wherein:the plurality of VAF subpopulations comprises a high VAF subpopulation corresponding to a subpopulation of high copy number (and high zygosity) SNVs, andfiltering the subset of minimal balanced SNVs comprises identifying and excluding SNVs of the high copy number subpopulation (e.g., based on their VAFs in relation to one or more parameters of the high VAF subpopulation and / or the main VAF subpopulation].

27. The method of any one of claims 21 to 26, wherein the one or more sets of VAF subpopulations comprise a set of high VAF outliers comprising minimal balanced SNVs having VAFs significantly greater than a first representative VAF of the subset of minimal balanced SNV events and / or representative VAF of a subpopulation thereof.

28. The method of any one of claims 21 to 27, wherein the one or more sets of VAF subpopulations comprise a set of high VAF outliers comprising minimal balanced SNVs having VAFs significantly greater than a first representative VAF of the subset of minimal balanced SNV events and / or a subpopulation thereof.

29. The method of any one of claims 21-28, comprising:determining, by the processor, a second representative VAF for the set of minimal balanced SNVs based on the VAF distribution for the subset of minimal balanced SNV events; andidentifying, by the processor, of the set of minimal balanced SNVs, a main subpopulation comprising a subset of minimal balanced SNVs having VAFs similar to the second representative VAF determined for the set of minimal balanced SNVs;determining, by the processor, a first representative VAF for the main population; and identifying, by the processor, the one or more sets of VAF outliers based at least in part on the first representative VAF determined for the main population.- 301 - 1324629 IvlAttomey Docket No. 2013237-150030. The method of claim 29, comprising:identifying, by the processor, a refined main subpopulation comprising a subset of minimal balanced SNVs having VAFs similar the first representative VAF determined for the main subpopulation;determining, by the processor, a representative VAF for the refined main subpopulation; andidentifying, by the processor, the one or more sets of VAF outliers based at least in part on the first representative VAF determined for the refined main population.

31. The method of claim 30, wherein the one or more sets of VAF outliers comprise a set of low VAF outliers and the method comprises:identifying, by the processor, an initial set of low VAF outliers comprising minimal balanced SNVs having VAFs below the representative VAF for the refined main subpopulation;determining, by the processor, an initial filtered subset of minimal balanced SNVs excluding the initial set of low VAF outliers; anditeratively updating, by the processor, the set of low VAF outliers and the initial filtered subset of minimal balanced SNVs by, beginning with a highest VAF member of the set of low VAF outliers, iteratively:moving a SNV from the set of low VAF outliers to the initial filtered subset if its VAF is within a predetermined threshold value of a minimum VAF of the initial fdtered subset; and progressing to evaluate a next lower VAF SNV of the set of low VAF outliers at a subsequent iteration.

32. The method of claim 30 or 31, wherein the one or more sets of VAF outliers comprise a set of high VAF outliers and the method comprises:- 302 - 1324629 IvlAttomey Docket No. 2013237-1500identifying, by the processor, an initial set of high VAF outliers comprising minimal balanced SNVs having VAFs above the representative VAF for the refined main subpopulation;determining, by the processor, an initial filtered subset of minimal balanced SNVs excluding the initial set of high VAF outliers; anditeratively updating, by the processor, the set of high VAF outliers and the initial filtered subset of minimal balanced SNVs by, beginning with a lowest VAF member of the set of low VAF outliers, iteratively:moving a SNV from the set of high VAF outliers to the initial filtered subset if its VAF is within a predetermined threshold value of a maximum VAF of the initial filtered subset; and progressing to evaluate a next higher VAF SNV of the set of high VAF outliers at a subsequent iteration.

33. The method of any one of the preceding claims, wherein step (e) comprises:determining, by the processor, one or more sequencing data observations for the subset of minimal balanced SNV events based on the tumor sequencing data, thereby determining one or more measured sequencing data observations;determining, by the processor, for each of one or more SNV models, one or more corresponding initial purity estimates, wherein:each SNV model generates one or more predicted sequencing data observations as a function of a variable tumor sample purity parameter anddetermining the initial purity estimate for a corresponding SNV model comprises determining a value of the variable tumor sample purity that optimizes a fit between the predicted sequencing data observations of the SNV model and corresponding measured sequencing data observations; anddetermining, by the processor, the estimated tumor sample purity based at least in part on the one or more initial purity estimates.- 303 - 1324629 IvlAttomey Docket No. 2013237-150034. The method of claim 33, wherein the one or more SNV models do not fix a SNV genotype and determining the initial purity estimate for a corresponding SNV model comprises determining individual SNV genotypes for the set of minimal balanced SNV events and / or a value of one or more variable parameters that measure a quantity of SNV events having a particular genotype.

35. The method of any one of the preceding claims, wherein step (e) comprises:determining, by the processor, for each particular SNV of at least a portion of the set of minimal balanced SNV events, a corresponding set of allele- specific read counts, wherein, for the particular SNV, the corresponding set of allele-specific read counts comprises a number of reads mapping to each of four possible alleles;determining, by the processor, a first purity value based on a plurality of SNV likelihood functions, each associated with a particular SNV event of the portion of minimal balanced SNV events and measuring a likelihood of measuring the corresponding set of allele- specific read counts for the particular SNV as a function of tumor sample purity, wherein the first purity value is a value of tumor sample purity determined to maximize an aggregate of the plurality of SNV likelihood functions;determining, by the processor, the estimated tumor sample purity based on the fist purity value.

36. The method of any one of the preceding claims, wherein step (e) comprises:determining, by the processor, an observed distribution of observed VAFs for at least a portion of the subset of minimal balanced SNV events;determining, by the processor, a predicted distribution of VAFs for the subset of minimal balanced SNV events;determining, by the processor, a second purity value that optimizes one or more test statistics measuring similarity between the observed and predicted distributions; and- 304 - 1324629 IvlAttomey Docket No. 2013237-1500determining, by the processor, at step (e), the estimated tumor sample purity based at least in part on the second purity value.

37. The method of claim 36, wherein the one or more test statistics comprises a K-statistic that measures a difference between the observed distribution and the predicted distribution.

38. The method of any one of claims 33 to 37, comprising, at step (e), determining the estimated tumor sample purity based on both an initial purity estimate and the second initial purity estimate.

39. The method of any one of the preceding claims, comprising using the tumor sequencing data together with the estimated tumor sample purity to detect a second plurality of SNVs within the tumor genome of the subject and / or to refine the list of putative single nucleotide variant (SNV) events.

40. The method of any one of the preceding claims, comprising:obtaining, by the processor, a list of candidate mutations identifying a plurality of mutations determined to be present in the tumor genome of the subject;identifying, by the processor, for each particular mutation of the list of candidate mutations, a corresponding segment within the tumor genome comprising the particular mutation;determining, by the processor, for each particular mutation of the list of candidate mutations, an estimated cellularity and / or variability thereof based at least in part on the estimated tumor sample purity, thereby determining estimated cellularity values and / or variabilities thereof for the list of candidate mutations; andselecting, by the processor, a subset of the list of candidate mutations based at least in part on the estimated cellularity values and / or variabilities thereof for the list of candidate mutations.- 305 - 1324629 IvlAttomey Docket No. 2013237-150041. The method of any one of the preceding claims, comprising:obtaining, by the processor, a list of candidate mutations identifying a plurality of mutations determined to be present in the tumor genome of the subject;identifying, by the processor, for each particular mutation of the list of candidate mutations, a corresponding segment within the tumor genome comprising the particular mutation;determining, by the processor, for each particular mutation of the list of candidate mutations, an identification of the particular mutation as clonal or subclonal based at least in part on the estimated sample purity, thereby identifying clonal and / or subclonal mutations within the list of candidate mutations; andselecting, by the processor, a subset of the list of candidate mutations based at least in part on the identified clonal and / or subclonal mutations within the list of candidate mutations.

42. The method of any one of the preceding claims, wherein the tumor sequencing data and / or normal sequencing data are whole genome sequencing data (WGS), whole exome sequencing data (WES), or single nucleotide polymorphism (SNP) array data.

43. The method of any one of the preceding claims, wherein the tumor sequencing data and / or normal sequencing data comprise a plurality of replicates.44 The method of claim 43, comprising using the plurality of replicates to detect at least a portion of the putative SNVs of the list.

45. The method of any one of the preceding claims, wherein the tumor sample is a formalin-fixed paraffin embedded (FFPE) sample.- 306 - 1324629 IvlAttomey Docket No. 2013237-150046. The method of any one of the preceding claims, wherein the tumor sample is a fresh frozen sample.

47. The method of any one of the preceding claims, comprising:obtaining, by the processor, normal sequencing data for a normal sample obtained from the subject; andat step (d), identifying the subset of minimal balanced SNV events based at least in part on the tumor sequencing data and the normal sequencing data.

48. The method of any one of the preceding claims, comprising:obtaining, by the processor, normal sequencing data for a normal sample obtained from the subject; andat step (e), determining the estimated tumor sample purity based on the normal sequencing data together with the tumor sequencing data and the subset of minimal balanced SNV events.

49. The method of any one of the preceding claims, wherein step (b) comprises detecting, by the processor, at least a portion of the putative SNV events of the list based at least in part on the tumor sequencing data.

50. The method of claim 49, comprising detecting the portion of the putative SNV events using the tumor sequencing data together with normal sequencing data obtained for a normal sample obtained from the subject.

51. The method of any one of claims 48 to 50, wherein the normal sample is a buffy coat sample.- 307 - 1324629 IvlAttomey Docket No. 2013237-150052. The method of claim 51, comprising detecting the portion of the putative SNV events using, the tumor sequencing data together with a normal reference genome.

53. The method of any one of the preceding claims, wherein the tumor sample is obtained from a subject diagnosed as having or at risk of having cancer.

54. The method of any one of the preceding claims, wherein the tumor sample is obtained from a subject diagnosed as having or at risk of having a cancer associated with low tumor mutation burden.

55. The method of any one of the preceding claims, wherein the tumor sample is obtained from a subject diagnosed as having or at risk of having breast cancer, prostate cancer, pancreatic cancer, pediatric cancer, neuroblastoma, ovarian cancer, renal cell carcinoma, Merkel cell carcinoma, hematologic cancer, colorectal cancer, melanoma, head and neck squamous cell carcinoma, or non-small cell lung cancer.

56. The method of any one of the preceding claims, wherein the estimated tumor sample purity is less than about 0.3.

57. The method of any one of the preceding claims, wherein the estimated tumor sample purity is at least about 0.10.

58. A method for determining a minimal copy number, or a lower bound thereon, of a set of balanced heterozygous segments in a tumor genome, the method comprising:(a) obtaining, by a processor of a computing device, tumor sequencing data for a tumor sample obtained from a subject;- 308 - 1324629 IvlAttomey Docket No. 2013237-1500(b) obtaining, by the processor, a list of putative single nucleotide variant (SNV) events detected based on the tumor sequencing data, each putative SNV event representing one or more potential underlying mutations occurring in a tumor genome of the subject;(c) identifying, of the list of putative SNV events, a subset of minimal balanced SNV events comprising SNV events determined to represent mutations that:(i) are determined to be balanced within the tumor genome; and (ii) are located within segments determined to have an absolute copy number equal to a same particular minimal copy number whose exact value is unknown but predicted to be lower than copy numbers of other balanced segments;(d) determining, by the processor, for at least a portion of the subset of minimal balanced SNV events, a corresponding variant allele frequency (VAF) based on the tumor sequencing data, thereby determining a VAF distribution for the subset of minimal balanced SNV events;(e) identifying, by the processor, one or more of VAF subpopulations within the VAF distribution for the subset of minimal balanced SNV events, each VAF subpopulation corresponding to and representing VAFs of a distinct subpopulation of SNV events;(f) determining, by the processor, an estimated value of the minimal copy number and / or a lower bound thereon (e.g., four or higher), based on the one or more VAF subpopulations; and(g) storing and / or providing, by the processor, for further processing, (i) the estimated value of the minimal copy number and / or the lower bound, for display and / or further processing.

59. The method of claim 58, wherein the identified VAF subpopulations comprise at least two subpopulations, including a first VAF subpopulation corresponding to a subpopulation of SNVs having a first absolute copy number and a first zygosity and a second VAF subpopulation corresponding to a subpopulation of SNVs having the first absolute copy number and a second zygosity, different from the first zygosity.- 309 - 1324629 IvlAttomey Docket No. 2013237-150060. The method of claim 59, comprising determining the estimated value of the minimal copy number and / or lower bound thereon based at least in part on the identification of the first and second VAF subpopulations.

61. The method of claim 59 or claim 60, wherein the estimated value of the minimal copy number and / or lower bound thereon is four (4).

62. The method of any one of claims 59 to 61, comprising identifying the first and second VAF subpopulations as harmonics.

63. The method of any one of claims 59 to 62, comprising identifying the first VAF subpopulation to correspond to subpopulation of late SNV events, occurring after a duplication event and identifying the second VAF subpopulation to correspond to a subpopulation of early SNV events, occurring prior to a duplication event64. The method of any one of claims 59 to 63, comprising:determining a first representative VAF measure of the first VAF subpopulation and a second representative VAF measure of the second VAF subpopulation; anddetermining the minimal copy number and / or lower bound thereon based at least in part on the first representative VAF measure and the second representative VAF measure.

65. The method of claim 64, comprising identifying the first and second VAF subpopulations as harmonics based at least in part on the first and second representative VAF measures66. The method of claim 65, comprising determining the minimal copy number to be four based at least in part on a ratio between the first and second representative VAF measures being approximately equal to 1:2.- 310 - 1324629 IvlAttorney Docket No. 2013237-150067. The method of any one of claims 58 to 66, comprising detecting presence of a whole genome duplication event based on the one or more VAF subpopulations.

68. The method of any one of claims 58 to 67 comprising determining a number of SNVs in each of the first and second VAF subpopulations and estimating the value of the minimal copy number and / or bound thereon based at least in part on the number of SNVs in the first and second VAF subpopulations.

69. The method of any one of claims 58 to 68, comprising aborting, by the processor, a SNV-based purity estimation procedure based at least in part on the estimated minimal copy number and / or bound thereon.

70. The method of any one of claims 58 to claim 69, comprising aborting, by the processor, a CNV-based purity estimation procedure based at least in part on the estimated minimal copy number and / or bound thereon.

71. The method of any one of claims 58 to 70, comprising using the estimated minimal copy number and / or to determine a CNV-based estimate of tumor sample purity and / or a bound thereon.

72. The method of any one of claims 58 to 71, wherein the one or more VAF subpopulations comprises a main VAF subpopulation corresponding to a subpopulation of balanced clonal SNVs having a single zygosity.- 311 - 1324629 IvlAttomey Docket No. 2013237-150073. The method of claim 72, comprising determining the estimated minimal copy number and / or lower bound thereon to be two (2) based at least in part on the identified main VAF subpopulation.

74. The method of any one of claims 72 to 73, wherein the one or more VAF subpopulations comprises a low VAF subpopulation corresponding to a subpopulation of subclonal SNVs.

75. The method of claim 73, comprising filtering the set of minimal balanced SNVs to remove the subpopulation of subclonal SNVs corresponding to the low VAF subpopulation.

76. The method of claim 74, comprising aborting an SNV-based purity estimation procedure based at least in part on the low VAF subopulation.

77. A method for determining a primary copy number, or a lower bound thereon, for a set of balanced heterozygous segments in a tumor genome, the method comprising:(a) obtaining, by a processor of a computing device, tumor sequencing data for a tumor sample obtained from a subject;(b) obtaining, by the processor, a list of putative single nucleotide variant (SNV) events detected based on the tumor sequencing data, each putative SNV event representing one or more potential underlying mutations occurring in a tumor genome of the subject;(c) identifying, of the list of putative SNV events, a subset of primary balanced SNV events comprising SNV events determined to represent mutations that:(i) are determined to be balanced within the tumor genome; and (ii) are located within segments determined to have an absolute copy number equal to a same particular primary copy number whose exact value is unknown but predicted to be a most frequently occurring even absolute copy number;- 312 - 1324629 IvlAttomey Docket No. 2013237-1500(d) determining, by the processor, for at least a portion of the subset of primary balanced SNV events, a corresponding variant allele frequency (VAF) based on the tumor sequencing data, thereby determining a VAF distribution for the subset of primary balanced SNV events;(e) identifying, by the processor, one or more of VAF subpopulations within the VAF distribution for the subset of primary balanced SNV events, each VAF subpopulation corresponding to and representing VAFs of a distinct subpopulation of SNV events;(f) determining, by the processor, an estimated value of the primary copy number and / or a lower bound thereon based on the one or more VAF subpopulations,(g) storing and / or providing, by the processor, for further processing, the estimated value of the primary copy number and / or the lower bound thereon for display and / or further processing.

78. A method for determining a minimal copy number, or a lower bound thereon, of a set of balanced heterozygous segments in a tumor genome, the method comprising:(a) obtaining, by a processor of a computing device, tumor sequencing data for a tumor sample obtained from a subject;(b) obtaining, by the processor, a list of putative single nucleotide variant (SNV) events detected based on the tumor sequencing data, each putative SNV event representing one or more potential underlying mutations occurring in a tumor genome of the subject;(c) identifying, of the list of putative SNV events, a subset of minimal balanced SNV events comprising SNV events determined to represent mutations that:(i) are determined to be balanced within the tumor genome; and (ii) are located within segments determined to have an absolute copy number equal to a same particular minimal copy number whose exact value is unknown but predicted to be lower than copy numbers of other balanced segments;(d) determining, by the processor, for at least a portion of the subset of minimal balanced SNV events, a corresponding variant allele frequency (VAF) based on the tumor - 313 - 1324629 IvlAttomey Docket No. 2013237-1500sequencing data, thereby determining a VAF distribution for the subset of minimal balanced SNV events;(e) identifying, by the processor, a plurality of VAF subpopulations within the VAF distribution for the subset of minimal balanced SNV events, each VAF subpopulation corresponding to and representing VAFs of a distinct subpopulation of SNV events;(f) detecting, by the processor, presence of a whole genome duplication event based at least in part on the plurality of VAF subpopulations; and(g) based at least in part on the detected presence of the WGD event, (i) determining a cancer status for the subject and / or (ii) selecting a therapy for the subject.

79. A method for determining a primary copy number, or a lower bound thereon, for a set of balanced heterozygous segments in a tumor genome, the method comprising:(a) obtaining, by a processor of a computing device, tumor sequencing data for a tumor sample obtained from a subject;(b) obtaining, by the processor, a list of putative single nucleotide variant (SNV) events detected based on the tumor sequencing data, each putative SNV event representing one or more potential underlying mutations occurring in a tumor genome of the subject;(c) identifying, of the list of putative SNV events, a subset of primary balanced SNV events comprising SNV events determined to represent mutations that:(i) are determined to be balanced within the tumor genome; and (ii) are located within segments determined to have an absolute copy number equal to a same particular primary copy number whose exact value is unknown but predicted to be a most frequently occurring even absolute copy number;(d) determining, by the processor, for at least a portion of the subset of primary balanced SNV events, a corresponding variant allele frequency (VAF) based on the tumor sequencing data, thereby determining a VAF distribution for the subset of primary balanced SNV events;- 314 - 1324629 IvlAttomey Docket No. 2013237-1500(e) identifying, by the processor, one or more of VAF subpopulations within the VAF distribution for the subset of primary balanced SNV events, each VAF subpopulation corresponding to and representing VAFs of a distinct subpopulation of SNV events;(f) detecting, by the processor, presence of a whole genome duplication event based at least in part on the plurality of VAF subpopulations; and(g) based at least in part on the detected presence of the WGD event, (i) determining a cancer status for the subject and / or (ii) selecting a therapy for the subject.

80. A system comprising:a processor of a computing device; andmemory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to:(a) obtain tumor sequencing data for a tumor sample obtained from a subject; (b) obtain a list of putative single nucleotide variant (SNV) events detected based on the tumor sequencing data, each putative SNV event representing one or more potential underlying mutations occurring in a tumor genome of the subject;(c) identify, from the list of putative SNV events, a subset of minimal balanced SNV events using the tumor sequencing data, the subset of minimal balanced SNV events comprising SNV events determined to represent mutations that:(i) are determined to be balanced within the tumor genome; and (ii) are located within segments determined to have an absolute copy number equal to a same particular minimal copy number whose exact value is unknown but predicted to be lower than copy numbers of other balanced segments;(d) determine an estimated tumor sample purity and / or bound thereon based at least in part on the tumor sequencing data and the subset of minimal balanced SNV events; and- 315 - 1324629 IvlAttomey Docket No. 2013237-1500(e) store and / or provide the estimated tumor sample purity and / or bound thereon for display and / or further processing.

81. A system comprising:a processor of a computing device; andmemory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to:(a) obtain tumor sequencing data for a tumor sample obtained from a subject;(b) obtain a list of putative single nucleotide variant (SNV) events detected based on the tumor sequencing data, each putative SNV event representing one or more potential underlying mutations occurring in a tumor genome of the subject;(c) identify, from the list of putative SNV events, a subset of minimal balanced SNV events comprising SNV events determined to represent mutations that:(i) are determined to be balanced within the tumor genome; and (ii) are located within segments determined to have an absolute copy number equal to a same particular minimal copy number whose exact value is unknown but predicted to be lower than copy numbers of other balanced segments;(d) determine, for at least a portion of the subset of minimal balanced SNV events, a corresponding variant allele frequency (VAF) based on the tumor sequencing data, thereby determining a VAF distribution for the subset of minimal balanced SNV events;(e) identify one or more of VAF subpopulations within the VAF distribution for the subset of minimal balanced SNV events, each VAF subpopulation corresponding to and representing VAFs of a distinct subpopulation of SNV events;(f) determine an estimated value of the minimal copy number and / or a lower bound thereon, based on the one or more VAF subpopulations; and(g) store and / or provide the estimated value of the minimal copy number and / or the lower bound, for display and / or further processing.- 316 - 1324629 IvlAttomey Docket No. 2013237-150082. A system comprising:a processor of a computing device; andmemory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to:(a) obtain tumor sequencing data for a tumor sample obtained from a subject;(b) obtain a list of putative single nucleotide variant (SNV) events detected based on the tumor sequencing data, each putative SNV event representing one or more potential underlying mutations occurring in a tumor genome of the subject;(c) identify, from the list of putative SNV events, a subset of primary balanced SNV events comprising SNV events determined to represent mutations that:(i) are determined to be balanced within the tumor genome; and (ii) are located within segments determined to have an absolute copy number equal to a same particular primary copy number whose exact value is unknown but predicted to be a most frequently occurring even absolute copy number;(d) determine, for at least a portion of the subset of primary balanced SNV events, a corresponding variant allele frequency (VAF) based on the tumor sequencing data, thereby determining a VAF distribution for the subset of primary balanced SNV events;(e) identify one or more of VAF subpopulations within the VAF distribution for the subset of primary balanced SNV events, each VAF subpopulation corresponding to and representing VAFs of a distinct subpopulation of SNV events;(f) determine an estimated value of the primary copy number and / or a lower bound thereon based on the one or more VAF subpopulations,(g) store and / or provide the estimated value of the primary copy number and / or the lower bound thereon for display and / or further processing.

83. A system comprising:- 317 - 1324629 IvlAttomey Docket No. 2013237-1500a processor of a computing device; andmemory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to:(a) obtain tumor sequencing data for a tumor sample obtained from a subject;(b) obtain a list of putative single nucleotide variant (SNV) events detected based on the tumor sequencing data, each putative SNV event representing one or more potential underlying mutations occurring in a tumor genome of the subject;(c) identify, from the list of putative SNV events, a subset of minimal balanced SNV events comprising SNV events determined to represent mutations that:(i) are determined to be balanced within the tumor genome; and (ii) are located within segments determined to have an absolute copy number equal to a same particular minimal copy number whose exact value is unknown but predicted to be lower than copy numbers of other balanced segments;(d) determine, for at least a portion of the subset of minimal balanced SNV events, a corresponding variant allele frequency (VAF) based on the tumor sequencing data, thereby determining a VAF distribution for the subset of minimal balanced SNV events;(e) identify a plurality of VAF subpopulations within the VAF distribution for the subset of minimal balanced SNV events, each VAF subpopulation corresponding to and representing VAFs of a distinct subpopulation of SNV events;(f) detect presence of a whole genome duplication event based at least in part on the plurality of VAF subpopulations; and(i) based at least in part on the detected presence of the WGD event, (i) determine a cancer status for the subject and / or (ii) select a therapy for the subject.

84. A system comprising:a processor of a computing device; and- 318 - 1324629 IvlAttomey Docket No. 2013237-1500memory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to:(a) obtain tumor sequencing data for a tumor sample obtained from a subject;(b) obtain a list of putative single nucleotide variant (SNV) events detected based on the tumor sequencing data, each putative SNV event representing one or more potential underlying mutations occurring in a tumor genome of the subject;(c) identify, from the list of putative SNV events, a subset of primary balanced SNV events comprising SNV events determined to represent mutations that:(i) are determined to be balanced within the tumor genome; and (ii) are located within segments determined to have an absolute copy number equal to a same particular primary copy number whose exact value is unknown but predicted to be a most frequently occurring even absolute copy number;(d) determine, for at least a portion of the subset of primary balanced SNV events, a corresponding variant allele frequency (VAF) based on the tumor sequencing data, thereby determining a VAF distribution for the subset of primary balanced SNV events;(e) identify one or more of VAF subpopulations within the VAF distribution for the subset of primary balanced SNV events, each VAF subpopulation corresponding to and representing VAFs of a distinct subpopulation of SNV events;(f) detect presence of a whole genome duplication event based at least in part on the plurality of VAF subpopulations; and(g) based at least in part on the detected presence of the WGD event, (i) determine a cancer status for the subject and / or (ii) select a therapy for the subject.

85. A method of producing an immunotherapy construct for a subject, the method comprising:detecting a plurality of candidate mutations in tumor cells of from the subject using a method or system of any one of claims 1-79; and- 319 - 1324629 IvlAttomey Docket No. 2013237-1500synthesizing, as the immunotherapy construct, a polyribonucleotide encoding a plurality of neoepitopes, each encoded neoepitope corresponding to a candidate mutation of the plurality of candidate mutations.

86. A method comprising:determining a subset of T-cells within a population of T-cells that are capable of specifically binding a plurality of complexes,wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, andwherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer- specific mutations detected using the method or system of any one of claims 1-79; and enriching for the subset of T-cells that are capable of specifically binding the plurality of complexes.

87. A method comprising:administering to a subject a population of T-cells, wherein the population of T cells are capable of specifically binding a plurality of complexes,wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells,wherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer- specific mutations detected using the method or system of any one of claims 1-79, and wherein the genome of at least some of the subject’s cells comprises a subset of the cancer-specific mutations.- 320 - 1324629 IvlAttorney Docket No. 2013237-150088. A method comprising:determining a subset of TILs within a population of TILs that are capable of specifically binding a plurality of complexes,wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, andwherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer- specific mutations detected using the method or system of any one of claims 1-79; and enriching for the subset of TILs that are capable of specifically binding the plurality of complexes.

89. A method comprising:administering to a subject a population of TILs, wherein the population of TILs are capable of specifically binding a plurality of complexes,wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells,wherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer- specific mutations detected using the method or system of any one of claims 1-79, and wherein the genome of at least some of the subject’s cells comprises a subset of the cancer-specific mutations.

90. The method of any one of claims 88 to 89, comprising obtaining a tumor sample from the subject and using the tumor sample to detect the plurality of candidate mutations.- 321 - 1324629 IvlAttomey Docket No. 2013237-150091. The method of claim 90, comprising obtaining a normal sample from the subject and using the normal sample to detect the plurality of cancer mutations.

92. The method of claim 90 or 91, comprising sequencing the tumor and / or normal sample.

93. A pharmaceutical composition comprising a polyribonucleotide encoding a plurality of neoepitopes, wherein at least a portion of the plurality of neoepitopes are individualized neoepitopes corresponding to a somatic mutation from a population of candidate mutations detected in a tumor sample of a subject and selected for inclusion in the pharmaceutical composition using the method or system of any one of the preceding claims.

94. An individualized pharmaceutical composition comprising a polyribonucleotide encoding a plurality of neoepitopes for a particular subject, wherein at least a portion of the neoepitopes are individualized neoepitopes corresponding to a somatic mutation from a population of candidate mutations detected in a tumor sample from the particular subject and selected for inclusion in the pharmaceutical composition using the method or system of any one of the claims 1-79.

95. A population of T-cells,wherein the population of T-cells is capable of specifically binding to a plurality of complexes,wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, andwherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer- specific mutations using the method or system of any one of claims 1-79.- 322 - 1324629 IvlAttomey Docket No. 2013237-150096. A T-cell capable of specifically binding to a complex, wherein the complex comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancerspecific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells,wherein the one or more cancer-specific mutations are detected using the method or system of any one of claims 1-79.

97. A population of T-cells,wherein the population of T-cells is capable of specifically binding to a plurality of complexes,wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, andwherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer- specific mutations detected using the method or system of any one of claims 1-79.

98. A T-cell capable of specifically binding to a complex, wherein the complex comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancerspecific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells,wherein the one or more cancer-specific mutations are within the selected subset of candidate mutations detected using the method or system of any one of claims 1-79.

99. A T-cell receptor capable of specifically binding to a complex, wherein the complex comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or- 323 - 1324629 IvlAttomey Docket No. 2013237-1500more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells,wherein the one or more cancer-specific mutations are within the selected subset of candidate mutations detected using the method or system of any one of claims 1-79.

100. A chimeric antigen receptor (CAR) capable of specifically binding to a complex, wherein the complex comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells,wherein the one or more cancer-specific mutations are within the selected subset of candidate mutations detected using the method or system of any one of claims 1-79.- 324 - 1324629 Ivl