Technologies for determining tumor sample purity and / or genomic characteristics

By identifying a subset of balanced heterozygous segments, the method optimizes tumor sample purity estimation, addressing inefficiencies and inaccuracies in existing techniques, providing robust and accurate genomic insights for personalized cancer therapies.

WO2026159254A1PCT designated stage Publication Date: 2026-07-30BIONTECH SE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BIONTECH SE
Filing Date
2026-01-23
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Existing methods for determining tumor sample purity and genomic characteristics are inefficient and prone to local optima, requiring complex parameter optimization and assumptions that may not hold for diverse patient populations, leading to inaccurate and computationally intensive results.

Method used

A method that identifies a subset of balanced heterozygous segments within a tumor genome, using these segments to estimate tumor sample purity by optimizing a reduced parameter space, avoiding local optima and reducing computational intensity, while providing unbiased and robust estimates.

Benefits of technology

The method achieves accurate and robust estimation of tumor sample purity and genomic properties, suitable for clinical applications, enabling precise detection of cancer-specific mutations and personalized therapies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2026051707_30072026_PF_FP_ABST
    Figure EP2026051707_30072026_PF_FP_ABST
Patent Text Reader

Abstract

Presented herein are methods and systems that allow biophysical and / or genomic properties of tumor samples to be determined via analysis of sequencing data, obtained, for example, via next-generation sequencing of tumor samples obtained from a subject. Among other things, tumor deconvolution technologies of the present disclosure leverage signals in sequencing data that originate from copy number variation (CNV) events that occur in cancer cells to estimate tumor sample purity and / or determine genomic properties, such as absolute copy numbers of segments within a tumor genome. Among other things, accurate assessments of tumor sample purity and / or genomic properties made possible via techniques described herein can be utilized to accurately detect and characterize cancer-specific mutations, such as single nucleotide variations (SNVs) that may serve as targets for personalized cancer therapies.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No.: 2013237-0970TECHNOLOGIES FOR DETERMINING TUMOR SAMPLE PURITY AND / OR GENOMIC CHARACTERISTICSCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to and benefit of U. S. Provisional Application No.63 / 749,339, filed on January 24, 2025, the content of which is hereby incorporated by reference herein in its entirety.BACKGROUND

[0002] Cancer is a primary cause of mortality, accounting for 1 in 4 of all deaths. Despite recent advances in the field of cancer immunotherapy there remains no single, broadly applicable treatment. Molecular heterogeneity of tumors renders many therapies ineffective for cancer patients.SUMMARY

[0003] Presented herein are methods and systems that allow biophysical and / or genomic properties of tumor samples to be determined via analysis of sequencing data, obtained, for example, via next-generation sequencing of tumor samples obtained from a subject. Among other things, tumor deconvolution technologies of the present disclosure leverage signals in sequencing data that originate from copy number variation (CNV) events that occur in cancer cells to estimate tumor sample purity and / or determine genomic properties, such as absolute copy numbers of segments within a tumor genome. Among other things, accurate assessments of tumor sample purity and / or genomic properties made possible via techniques described herein can be utilized to accurately detect and characterize cancer- specific mutations, such as single nucleotide variations (SNVs) that may serve as targets for personalized cancer therapies.

[0004] Among other things, in certain embodiments, the present disclosure addresses challenges associated with estimation of purity from CNV events, which typically involves optimizing multiple parameters of a large space of parameters. For example, estimation of purity from CNV events relies on the absolute copy number of segments, which are generally not known in advance. Accordingly, in certain embodiments, approaches described herein provide an advantage by first identifying a particular subset of segments within a tumor genome -namely a subset of segments that are heterozygous, balanced (e.g., have equal number ofAttomey Docket No.: 2013237-0970maternal and paternal alleles) and have a same, but as yet unknown, absolute copy number.Among other things, as described herein, having first identified PBHSs, certain auxiliary parameters used for estimation of the tumor sample purity can be directly estimated from the PBHSs, thereby drastically reducing the dimensionality of the problem of CNV-based tumor deconvolution. For example, instead of having to optimize multiple parameters of a large space of parameters, only the purity (e.g., one continuous parameter) and the value of the absolute copy number of the selected subset of balanced heterozygous segments - referred to herein as a primary copy number (one discrete parameter) need to be optimized. Moreover, owing to judicious selection of the particular subset of balanced heterozygous segments used to initiate tumor deconvolution as described herein, the primary copy number can be restricted to a small number of possible even integers. In certain embodiments, for a vast majority of samples across many indications the primary copy number will be either 2 or 4.

[0005] In certain embodiments, methods described herein have the further advantage that the process of finding an optimal purity based on CNV events (1) is highly robust because local (false) optima, which occur when searching for optima among a large parameter space, are avoided, (2) is unbiased, e.g., no assumptions need to be made with regards to the tumor genome in order to reduce the parameter space (e.g., by assuming certain recurrent cancer karyotypes), and (3) is substantially less computationally intensive. Such assumptions will likely be violated for a fair subset of patients when applying the method to many patients in the general population.

[0006] Without wishing to be bound to any particular theory, it is believed that if the search space was larger and included additional unknown parameters then the optimization of the purity would be prone to local (false) optima, and the true solution would possibly not be found. In this way, tumor deconvolution methods and systems described herein may offer levels of accuracy and robustness suitable for clinical applications and decision making in ways that previous techniques do not.

[0007] In some aspects, the present disclosure provides a method (e.g., a computer-implemented method) (e.g., for determining genomic and biophysical properties of a tumor sample from sequencing data) comprising: (a) obtaining, by a processor of a computing device, tumor sequencing data for a tumor sample obtained from a subject [e.g., the tumor sequencing data comprising a plurality of tumor reads, each tumor read representing a partial nucleotide - 2 - 13247547v 1Attomey Docket No.: 2013237-0970sequence determined by sequencing the tumor sample]; (b) identifying, by the processor, using the tumor sequencing data, a plurality of primary balanced heterozygous segments within a tumor genome for the subject, each primary balanced heterozygous segment having been determined to: (i) comprise one or more heterozygous SNPs, at least a portion of which are determined to be balanced [e.g., are determined to have an equal number of copies of corresponding maternal and paternal alleles (e.g., via a balanced test and / or contour filtering refinement)] within the tumor genome; and (ii) have an absolute copy number equal to a particular, unknown-but-to-be-determined, primary copy number (e.g., the exact value of the copy number is initially unknown, but all the primary balanced heterozygous segments are determined to have same absolute copy number, e.g., and known to be even, e.g., by virtue of the PBHSs being balanced); (c) determining, by the processor, based on the tumor sequencing data and the primary balanced heterozygous segments, an estimated tumor sample purity and / or bound thereon for the tumor sample; and (d) storing and / or providing, by the processor, the estimated tumor sample purity and / or bound thereon for display and / or further processing.

[0008] In some embodiments, provided methods comprise using (e.g., by the processor) the estimated tumor sample purity and / or bound thereon to detect and / or prioritize a plurality of somatic mutations.

[0009] In some embodiments, provided methods comprise using at least a portion of the detected and / or prioritized somatic mutations in a personalized immunotherapy [e.g., creating a personalized cancer vaccine encoding one or more neoepitopes corresponding to at least a portion of the somatic mutations].

[0010] In some embodiments, provided methods comprise using the plurality of somatic mutations and / or prioritization thereof to determine a cancer status (e.g., a particular type of cancer; e.g., a particular stage of cancer) for the subject.

[0011] In some embodiments, provided methods comprise using the plurality of somatic mutations and / or prioritization thereof to select a therapy for the subject.

[0012] In some embodiments, provided methods comprise using the estimated tumor sample purity and / or bound thereon to determine a cancer status (e.g., a particular type of cancer; e.g., a particular stage of cancer) for the subject.- 3 - 13247547v 1Attomey Docket No.: 2013237-0970

[0013] In some embodiments, provided methods comprise using the estimated tumor sample purity and / or bound thereon to select a therapy for the subject.

[0014] In some embodiments, provided methods comprise using the estimated tumor sample purity and / or bound thereon as a quality control [e.g., identifying the tumor sequencing data as insufficient quality based on the estimated tumor sample purity and / or bound thereon (e.g., having been determined to be below a threshold value); e.g., identifying the tumor sample as insufficient quality based on the estimated tumor sample purity and / or bound thereon].

[0015] In some embodiments, provided methods comprise repeating sequencing of the tumor sample based on the estimated tumor sample purity and / or bound thereon.

[0016] In some embodiments, provided methods comprise: identifying, by the processor, a set of balanced heterozygous segments within the tumor genome for the subject (e.g., and, optionally, also within a normal genome associated with a sample), each balanced heterozygous segment of the set having been determined to comprise one or more heterozygous SNPs, at least a portion of which are balanced [e.g., are determined to have an equal number of copies of corresponding maternal and paternal alleles (e.g., via a balanced test and / or contour filtering refinement)] within the tumor genome; and selecting, by the processor, a subset of the set of balanced heterozygous segments as primary balanced heterozygous segments having a same particular absolute copy number (e.g., a primary copy number), thereby identifying the plurality of primary balanced heterozygous segments.

[0017] In some embodiments, provided methods comprise: determining, by the processor, for each of at least a portion of the set of BHSs, a tumor-to-normal read count ratio based on the tumor sequencing data and the normal sequencing data, thereby determining a tumor-to-normal read count ratio distribution for the set of BHSs; identifying, by the processor, one or more components of the tumor-to-normal read count ratio distribution, each of the one or more subcomponents representing (e.g., tumor-to-normal read count ratios of) a subpopulation of BHSs having a particular (e.g., even, e.g., by virtue of being a subpopulation of BHSs) absolute copy number, distinct from that of other subpopulations represented by other components of the tumor-to-normal read count ratio distribution; and selecting, by the processor, based on the identified components of the tumor-to-normal read count ratio distribution, the subset of primary balanced heterozygous segments (PBHSs).- 4 - 13247547v 1Attomey Docket No.: 2013237-0970

[0018] In some embodiments, provided methods comprise obtaining, by the processor, normal sequencing data for a normal sample obtained from the subject [e.g., the normal sequencing data comprising a plurality of normal reads, each normal read representing a partial nucleotide sequence determined by sequencing the normal sample], and wherein the measure of tumor read counts is a tumor- to-normal read count ratio determined, for a given BHS, based on a number of tumor reads mapping to the given BHS and a number of normal reads mapping to the given BHS (e.g., as a ratio).

[0019] In some embodiments, provided methods comprise identifying, by the processor, a particular one of the one or more components as a primary component representing a largest subpopulation of the BHSs (e.g., having a greatest number of balanced heterozygous segments therein) (e.g., having a most frequently occurring absolute copy number and / or tumor-to-normal read count ratio) (e.g., determining, for each of the one or more components, a corresponding amplitude representing a number and / or relative frequency of tumor-to-normal read count ratio values associated with the component and selecting, by the processor, the particular component having the largest corresponding amplitude); and selecting, by the processor, BHSs determined to be members of the largest subpopulation represented by the primary component as for inclusion in the subset of PBHSs {e.g., by: for each segment of at least a portion of the BHSs, determining one or more posterior probabilities associated with the segment, each corresponding to a particular one of the one or more components of the tumor-to-normal read count ratio distribution for the set of BHSs and measuring a likelihood of the segment being a member of the subpopulation represented by the particular component (e.g., based on a tumor-to-normal read count ratio for the segment); and identifying the members of the subpopulation represented by the primary component based on the determined posterior probabilities for each of the portion of BHSs [e.g., identifying segments for which the associated posterior probability corresponding to the primary component is larger than other posterior probabilities associated with the segment (e.g., corresponding to other components)].}

[0020] In some embodiments, identifying the primary component comprises: fitting, by the processor, a mixture model (e.g., a ID Gaussian Mixture Model) to the tumor-to-normal read count ratio distribution, thereby determining, based on the fit, a plurality of parameter values (e.g., for each component a corresponding mean, amplitude, and standard deviation), including,- 5 - 13247547v 1Attomey Docket No.: 2013237-0970for each of the one or more components, a corresponding amplitude; and identifying, by the processor, the primary component as having a largest corresponding amplitude.

[0021] In some embodiments, provided methods comprise determining, by the processor, based on the plurality of primary balanced heterozygous segments, values of one or more auxiliary parameters.

[0022] In some embodiments, the one or more auxiliary parameters comprises a primary slope that measures a (e.g., average) tumor-to-normal read count ratio for the plurality of primary balanced heterozygous segments.

[0023] In some embodiments, the one or more auxiliary parameters comprises a primary residual standard deviation that measures a standard deviation of residual errors for the plurality of primary balanced heterozygous segments.

[0024] In some embodiments, the one or more auxiliary parameters comprises a primary allele frequency standard deviation that measures a standard deviation of allele frequency values determined for the plurality of primary balanced heterozygous segments.

[0025] In some embodiments, provided methods comprise, at step (d), using the determined values of the one or more auxiliary parameters to determine the estimated sample purity.

[0026] In some embodiments, provided methods comprise, at step (d), determining, by the processor, for each of at least a portion of (e.g., heterozygous) segments within the tumor genome, values for each of one or more sequencing data observations (e.g., an allele frequency, a tumor-to-normal read count ratio, a differential allele frequency, a residual error) based on the tumor sequencing data and / or normal sequencing data, thereby determining a set of measured distributions comprising, for each of the one or more sequencing data observations, a corresponding distribution of values across the portion of tumor genome segments; determining, by the processor, one or more tumor model fits based on the distributions corresponding to the one or more sequencing data observations and a tumor model, wherein determining each one or more tumor model fits comprises: generating, using the tumor model, a set of predicted distributions for the one or more sequencing data observations as a function of one or more parameters (e.g., a primary copy number, a tumor sample purity, a coupling constant) of the- 6 - 13247547v 1Attorney Docket No.: 2013237-0970tumor model; and determining best-fit values for at least a portion of the one or more model parameters that optimize a model accuracy function that measures a degree to which the set of predicted distributions match the set of measured distributions, such that each tumor model fit is associated with (i) a set of best- fit values of at least a portion of the one or more model parameters and (ii) a model accuracy score corresponding to a value of the model accuracy function when the set of best fit values are used as values for the respective model parameters.

[0027] In some embodiments, one or more parameters of the tumor model comprise a tumor sample purity, which is included in the portion of the one or more model parameters for which the best-fit values are determined by optimizing the model accuracy, such determining the one or more tumor model fits comprises determining one or more best-fit values of the tumor sample purity, each associated with a corresponding tumor model fit.

[0028] In some embodiments, provided methods, at step (d), comprise determining the estimated tumor sample purity based on (e.g., as a selected one of) the one or more best-fit values of the tumor sample purity.

[0029] In some embodiments, one or more parameters of the tumor model comprises the primary copy number and wherein determining the one or more tumor model fits comprises, for each candidate primary copy number value of a set of candidate values, determining a corresponding tumor model fit by using the candidate primary copy number value for the primary copy number model parameter and determining best-fit values that optimize the model accuracy function for a remainder of the one or more model parameters, such that each tumor model fit is associated with (i) the corresponding candidate primary copy number value, (ii) the set of best- fit values for the remainder of the one or more model parameters and (iii) the model accuracy score corresponding to the value of the model accuracy function when the corresponding candidate primary copy number value and the set of best fit values for the remainder of the model parameters are used.

[0030] In some embodiments, one or more model parameters comprise a coupling constant that measures how a change in a value of the absolute copy number of the plurality of PBHSs impacts a change in observed tumor-to-normal read count ratios for the plurality of PBHSs.13247547v 1Attomey Docket No.: 2013237-0970

[0031] In some embodiments, provided methods comprise using a primary slope auxiliary parameter (e.g., the primary slope auxiliary parameter of various embodiments described herein) as an initial value for the coupling constant.

[0032] In some embodiments, one or more sequencing data observations comprise an observed allele frequency [e.g., determined, for a given heterozygous segment, as a fraction of total tumor read counts for the given heterozygous segment that map to a particular allele (e.g., a major allele) of the given heterozygous] or a function thereof.

[0033] In some embodiments, one or more sequencing data observations comprises a differential allele frequency that measures, for a given heterozygous segment, a difference between (i) the observed allele frequency and (ii) a predicted allele frequency for the given heterozygous segment.

[0034] In some embodiments, at least one of the one or more sequencing data observations is determined, for a given segment, as a function of tumor read counts for the given segment and / or normal read counts for the given segment.

[0035] In some embodiments, one or more sequencing data observations comprises a tumor-to-normal read count ratio, determined, for a given segment, as a ratio of tumor read counts to normal read counts for the given segment.

[0036] In some embodiments, one or more sequencing data observations comprises a residual error determined based on (i) tumor read counts for the given segment and (ii) normal read counts for the given segment (e.g., based on a difference between an observed tumor read count for the given segment and a value predicted via a linear predictor based on normal read counts of a corresponding segment).

[0037] In some embodiments, residual error is determined based on a variance stabilizing transformation.

[0038] In some embodiments, provided methods comprise determining the one or more tumor model fits comprises, for at least a portion of the tumor model parameters, determining best-fit values that maximize a value of a model likelihood function that measures (e.g., quantifies) an accuracy with which the tumor model, when particular values of the model- 8 - 13247547v 1Attomey Docket No.: 2013237-0970parameters are used (e.g., when the model parameters are set to the particular values), predicts or explains the distributions of the one or more sequencing data observations.

[0039] In some embodiments, a tumor model predicts a set of grid points, each grid point representing a predicted (e.g., average) allele frequency and a predicted tumor-to-normal read count ratio corresponding to a particular node comprising a specific combination of an absolute copy number value and an allele specific copy number (e.g., that a heterozygous segment could have).

[0040] In some embodiments, a predicted set of grid points is a function of the one or more model parameters (e.g., the tumor sample purity, the primary copy number, and the coupling constant).

[0041] In some embodiments, provided methods comprise determining, by the processor, a value of the primary copy number.

[0042] In some embodiments, provided methods comprise using a multivariate classifier to determine the value of the primary copy number.

[0043] In some embodiments, provided methods comprise, for each candidate primary copy number value of a set of potential primary copy number values [e.g., a set of even integers, ranging from 2 (e.g., CNpc = 2) to a (e.g., predefined) maximum copy number (e.g., 8, 10, 12, 20, 30, 40, 50, 66, 100, etc.)], determining, by the processor, corresponding values for each of one or more test statistics (e.g., each test statistic a metric that quantifies a degree of accuracy or plausibility of a given primary copy number value); and selecting, by the processor, a particular one of the candidate primary copy number value as the value of the primary copy number, based on (i) the values of the one or more test statistics corresponding to each of the candidate primary copy number and (ii) the multivariate classifier (e.g., using the values of the one or more test statistics as input to the multivariate classifier).

[0044] In some embodiments, one or more test statistics comprises a maximum value of a model likelihood function that measures (e.g., quantifies) an accuracy with which the tumor model predicts or explains the distributions of the one or more sequencing data observations for a given candidate primary copy number value.

[0045] In some embodiments, one or more test statistics comprises a ^-statistic that measures (e.g., is inversely proportional to) a difference between, for one or more of the- 9 - 13247547v 1Attomey Docket No.: 2013237-0970sequencing data observations, the distribution determined based on the tumor and normal sequencing data and a predicted distribution (e.g., a Kullback-Leibler (KL) divergence) [e.g., wherein the one or more of the sequencing data observations comprise a differential allele frequency (e.g., an / coordinate) and a residual error (e.g., an e coordinate)].

[0046] In some embodiments, one or more test statistics comprises a cluster density {e.g., a cluster density (e.g., that measures the degree of genomic continuity of contiguous regions on the genome to which a given cluster of heterozygous segments assigned to a same particular node map (e.g., averaged over one or more clusters); e.g., a global cluster density, determined as an average the cluster density over all clusters; e.g., a specific cluster density, determined as an average cluster density over all clusters except those with corresponding duplicate nodes)}.

[0047] In some embodiments, provided methods comprise assigning, by the processor, based at least in part on the estimated sample purity, absolute copy number values to each of at least a portion (e.g., substantially all) of segments within the tumor genome.

[0048] In some embodiments, provided methods comprise, for each to-be-assigned segment of the portion of segments within the tumor genome: determining, by the processor, one or more node assignment likelihoods for the to-be-assigned segment, each node likelihood assignment likelihood measuring a likelihood of the particular (e.g., heterozygous) segment having a specific absolute copy number (C / Vmut) and allele specific copy number (Cx) based at least in part on: (i) observed tumor reads counts for the particular heterozygous segment and / or a particular allele thereof (e.g., based on the tumor sequencing data) and observed normal read counts for the particular heterozygous segment and / or the particular allele thereof (e.g., based on the tumor sequencing data); and (ii) the estimated tumor sample purity (e.g., and the primary copy number; e.g., and a best-fit value of the coupling constant).

[0049] In some embodiments, provided methods comprise identifying, by the processor, a first heterozygous segment having been assigned an odd absolute copy number; determining, by the processor, the first heterozygous segment to be balanced in the tumor genome (e.g., based on a previously determined balanced state and / or refined balanced state; e.g., by evaluating a balanced state test for the first heterozygous segment), thereby determining the first heterozygous segment to be balanced in the tumor genome yet having been assigned an odd absolute copy number; and responsive to the determining the first heterozygous segment to be - 10 - 13247547v 1Attorney Docket No.: 2013237-0970balanced in the tumor genome yet having been assigned an odd absolute copy number, updating, by the processor, the assigned odd absolute copy number to an even, error-corrected value.

[0050] In some embodiments, an error corrected value to which the absolute copy number of the first heterozygous segment is updated is an absolute copy number represented by a nearest even node relative to the first heterozygous segment (e.g., an even node for which a node assignment likelihood for the first heterozygous segment is higher than node assignment likelihoods for all other even nodes) [e.g., and wherein the method comprises: (i) receiving and / or accessing a nearest even node assignment likelihood; (ii) receiving and / or accessing a current node assignment likelihood; and (iii) determining the nearest even node assignment likelihood to be at least a minimum threshold fraction of the current node assignment likelihood and, responsive thereto, updating the assigned absolute copy number to the absolute copy number represented by the nearest even node].

[0051] In some embodiments, provided methods comprise identifying one or more correctly assigned neighboring segments of the first heterozygous segment (e.g., segments within no more than a particular threshold distance of the first heterozygous segment and determined to not require error correction) and determining the error-corrected value based on absolute copy numbers of the one or more correctly assigned neighboring segments.

[0052] In some embodiments, provided methods comprise identifying, by the processor, a local genetic neighborhood of a first segment; and updating, by the processor, the assigned absolute copy number of the first segment based on absolute copy numbers of one or more segments within the local genetic neighborhood.

[0053] In some embodiments, provided methods comprise evaluating one or more (e.g., a plurality of) quality assurance criteria {e.g., each criterion based in whole or in part on one or more members selected from the group consisting of: the estimated tumor sample purity; a determined value of the primary copy number; a result of a multivariate classifier (e.g., used to select a primary copy number value); a cluster density; a node divergence score; a fraction of CNV outliers; a measure of cluster variance; and a degree of smearing of a primary component cluster.}; determining a failure in purity estimation based on the one or more quality assurance criteria; and responsive to the determined failure in purity estimation, determining (i) a bound- 11 - 13247547v 1Attomey Docket No.: 2013237-0970(e.g., an upper bound) on tumor sample purity and / or (ii) selecting an alternative estimate of tumor sample purity or bound therein, determined via an alternate method.

[0054] In some embodiments, provided methods comprise using tumor sequencing data (e.g., and normal sequencing data) together with the estimated tumor sample purity to detect a plurality of SNVs within the tumor genome of the subject.

[0055] In some embodiments, provided methods comprise using obtaining [e.g., receiving and / or accessing; e.g., detecting, by the processor, based on the tumor sequencing data (e.g., and the normal sequencing data; and the estimated tumor sample purity)], a list of candidate mutations (e.g., SNVs) identifying a plurality of mutations determined to be present in the tumor genome of the subject; identifying, by the processor, for each particular mutation of the list of candidate mutations, a corresponding segment within the tumor genome comprising the particular mutation; determining, by the processor, for each particular mutation of the list of candidate mutations, an absolute copy number of the corresponding segment comprising the particular mutation based at least in part on the estimated sample purity (e.g., together with a determined value of the primary copy number), thereby determining absolute copy numbers for the list of candidate mutations; and selecting, by the processor, a subset of the list of candidate mutations {e.g., for inclusion in and / or targeting via a construct [e.g., a therapeutic agent (e.g., a personalized therapeutic agent, such as a personalized cancer vaccine, a T cell receptor (TCR) therapy, etc.)]} based at least in part on the determined absolute copy numbers for the list of candidate mutations.

[0056] In some embodiments, provided methods comprise obtaining [e.g., receiving and / or accessing; e.g., detecting, by the processor, based on the tumor sequencing data (e.g., and the normal sequencing data; and the estimated tumor sample purity)], by the processor, a list of candidate mutations (e.g., SNVs) identifying a plurality of mutations determined to be present in the tumor genome of the subject; identifying, by the processor, of the list of candidate mutations, one or more mutations as mapping to (e.g., locations within) heterozygous segments (e.g., of the normal genome of the subject) and, for each particular mutation of the one or more mutations identified as mapping to heterozygous segments, identifying, a corresponding segment within the tumor genome comprising the particular mutation; determining, by the processor, for each particular mutation of the one or more mutations identified as mapping to heterozygous- 12 - 13247547v 1Attorney Docket No.: 2013237-0970segments, an allele specific copy number and / or a fractional zygosity of the corresponding segment comprising the particular mutation based at least in part on the estimated sample purity (e.g., together with a determined value of the primary copy number), thereby determining allele specific copy numbers and / or fractional zygosities for the one or more mutations mapping to heterozygous segments; and selecting, by the processor, a subset of the list of candidate mutations {e.g., for inclusion in and / or targeting via a construct [e.g., a therapeutic agent (e.g., a personalized therapeutic agent, such as a personalized cancer vaccine, a T cell receptor (TCR) therapy, etc.)]} based at least in part on the determined allele specific copy numbers and / or fractional zygosities for the one or more mutations mapping to heterozygous segments.

[0057] In some embodiments, provided methods comprise obtaining [e.g., receiving and / or accessing; e.g., detecting, by the processor, based on the tumor sequencing data (e.g., and the normal sequencing data; and the estimated tumor sample purity)], by the processor, a list of candidate mutations (e.g., SNVs) identifying a plurality of mutations determined to be present in the tumor genome of the subject; identifying, by the processor, for each particular mutation of the list of candidate mutations, a corresponding segment within the tumor genome comprising the particular mutation; determining, by the processor, for each particular mutation of the list of candidate mutations, an estimated cellularity and / or variability thereof (e.g., standard deviation; e.g., a confidence interval) based at least in part on the estimated sample purity (e.g., together with a determined value of the primary copy number), thereby determining estimated cellularity values and / or variabilities thereof for the list of candidate mutations; and selecting, by the processor, a subset of the list of candidate mutations {e.g., for inclusion in and / or targeting via a construct [e.g., a therapeutic agent (e.g., a personalized therapeutic agent, such as a personalized cancer vaccine, a T cell receptor (TCR) therapy, etc.)] } based at least in part on the estimated cellularity values and / or variabilities thereof for the list of candidate mutations.

[0058] In some embodiments, provided methods comprise obtaining [e.g., receiving and / or accessing; e.g., detecting, by the processor, based on the tumor sequencing data (e.g., and the normal sequencing data; and the estimated tumor sample purity)], by the processor, a list of candidate mutations (e.g., SNVs) identifying a plurality of mutations determined to be present in the tumor genome of the subject; identifying, by the processor, for each particular mutation of the list of candidate mutations, a corresponding segment within the tumor genome comprising- 13 - 13247547v 1Attomey Docket No.: 2013237-0970the particular mutation; determining, by the processor, for each particular mutation of the list of candidate mutations, an identification of the particular mutation as clonal or subclonal based at least in part on the estimated sample purity (e.g., together with a determined value of the primary copy number), thereby identifying clonal and / or subclonal mutations within the list of candidate mutations; and selecting, by the processor, a subset of the list of candidate mutations {e.g., for inclusion in and / or targeting via a construct [e.g., a therapeutic agent (e.g., a personalized therapeutic agent, such as a personalized cancer vaccine, a T cell receptor (TCR) therapy, etc.)] } based at least in part on the identified clonal and / or subclonal mutations within the list of candidate mutations.

[0059] In some embodiments, provided methods comprise causing, by the processor, display and / or rendering of a graphical user interface (GUI) and / or graphical report comprising a graphical representation of at least a portion (e.g., a particular chromosome; e.g., a user-selected subregion) of the tumor genome of the subject.

[0060] In some embodiments, a graphical representation of the portion of the tumor genome of the subject comprises one or more segment display tracks, each visually representing a plurality of segments within the portion of the tumor genome {e.g., wherein the one or more segment display tracks visually represent the plurality of segments along with determined genomic properties [e.g., identifications of particular subsets of segments (e.g., heterozygous segments, balanced heterozygous segments, primary balanced heterozygous segments); e.g., absolute copy number values (e.g., before and / or after error correction)}.

[0061] In some embodiments, one or more segment display tracks visually represent each segment via a corresponding graphical element [e.g., a line, a rectangular patch, a marker (e.g., a square, rectangular, circle, triangle, asterisk, etc.)], wherein positions and / or orientations of the graphical elements visually represent (e.g., relative) locations of the corresponding segments within the tumor genome and / or the graphical elements are color-coded to visually represent a particular genomic property of the corresponding segment.

[0062] In some embodiments, a graphical representation of the portion of the tumor genome comprises one or more mutation display tracks, each visually representing a plurality of detected mutations within the portion of the tumor genome (e.g., overlaid on the one or more segment display tracks).- 14 - 13247547v 1Attorney Docket No.: 2013237-0970

[0063] In some embodiments, tumor sequencing data and / or normal sequencing data are whole genome sequencing data (WGS), whole exome sequencing data (WES), or single nucleotide polymorphism (SNP) array data.

[0064] In some embodiments, tumor sequencing data and / or normal sequencing data comprise a plurality of replicates.

[0065] In some embodiments, a plurality of replicates are obtained from multiple samples of extracted gDNA from tumor and normal samples, obtained for multiple library preparations (e.g., from a single sample), obtained via multiple sequencing runs on a single library (e.g., technical replicates).

[0066] In some embodiments, a tumor sample is a formalin-fixed paraffin embedded (FFPE) sample.

[0067] In some embodiments, a tumor sample is a fresh frozen sample.

[0068] In some embodiments, a tumor sample is obtained from a subject diagnosed as having or at risk of having cancer.

[0069] In some embodiments, a tumor sample is obtained from a subject diagnosed as having or at risk of having a cancer associated with low tumor mutation burden.

[0070] In some embodiments, a tumor sample is obtained from a subject diagnosed as having or at risk of having breast cancer, prostate cancer, pancreatic cancer, pediatric cancer (e.g., pediatric acute lymphoblastic leukemia), neuroblastoma, ovarian cancer, renal cell carcinoma, Merkel cell carcinoma, hematologic cancer, colorectal cancer, melanoma, head and neck squamous cell carcinoma, or non-small cell lung cancer.

[0071] In some embodiments, an estimated tumor sample purity is less than about 0.3 (e.g., less than 0.3, less than 0.25, less than 0.20, less than 0.19, less than 0.18, less than 0.17, less than 0.16, less than 0.15, less than 0.14, less than 0.13, less than 0.12, or less than 0.11).

[0072] In some embodiments, an estimated tumor sample purity is at least about 0.10 (e.g., at least 0.10, at least 0.11, at least 0.12, at least 0.13, at least 0.14, at least 0.15, at least 0.16, at least 0.17, at least 0.18, at least 0.19).- 15 - 13247547v 1Attomey Docket No.: 2013237-0970

[0073] In some embodiments, provided methods comprise determining, by the processor, a value of the primary copy number (e.g., via a method disclosed herein) and detecting, based on the value of the primary copy number, presence of a whole genome duplication (WGD) event in the tumor sample.

[0074] In some embodiments, provided methods comprise detecting presence of the whole genome duplication event based at least in part of the value of the primary copy number having been determined to be at least four.

[0075] In some aspects, the present disclosure provides a method (e.g., for correcting errors in copy numbers assigned to segments of a tumor genome) comprising: (a) obtaining [e.g., receiving and / or accessing (e.g., based on a database; e.g., detecting, by a processor of a computing device (e.g., based on the normal sequencing data)], by a processor of a computing device, a list of nominally heterozygous segments within a tumor genome of a subject, and each segment of the list corresponding to a heterozygous segment within a normal genome of the subject; (b) for each nominally heterozygous segment of the list: obtaining [e.g., receiving and / or accessing; e.g., determining], by the processor, a currently assigned absolute copy number representing a number of copies of the segment having been determined to be present in the tumor genome; and obtaining [e.g., receiving and / or accessing; e.g., determining], by the processor, a balanced state indicating whether the segment is balanced, thereby determining, for the list of nominally heterozygous segments, balanced states of corresponding tumor genome segments; (c) identifying, by the processor, based on the balanced states of the list of nominally heterozygous segments, a subset of balanced heterozygous segments; (d) identifying, by the processor, within the subset of balanced heterozygous segments, one or more segments with oddvalued corresponding absolute copy numbers as potential copy number assignments errors; (e) updating, by the processor, absolute copy numbers of the potential copy number assignment errors to an error-corrected value, wherein each error-corrected value is an even number (e.g., integer); and (f) storing, by the processor, for display and / or further processing, the error corrected values.

[0076] In some embodiments, provided methods comprise, at step (e), for at least one particular segment identified as a potential copy number assignment error: receiving and / or determining, by the processor, for the particular segment, a first node assignment likelihood - 16 - 13247547v 1Attorney Docket No.: 2013237-0970corresponding to the currently assigned absolute copy number of the particular segment and a second node assignment likelihood corresponding to a second absolute copy number, wherein the second absolute copy number is even-valued and wherein the first and second node assignment likelihoods represent a likelihood of the particular segment having the currently assigned and second absolute copy number, respectively; and based on the first and second node assignment likelihoods, updating, by the processor, the currently assigned absolute copy number by replacing it with the second absolute copy number.

[0077] In some embodiments, provided methods comprise, at step (e), for at least one particular segment identified as a potential copy number assignment error: identifying, by the processor, within the list of nominally heterozygous segments, one or more genomic neighbor segments with correctly assigned copy numbers, the one or more genomic neighbors comprising (i) a nearest upstream segment (e.g., in a 5’ direction) and (ii) a nearest downstream segment (e.g., in a 3’ direction); and updating, by the processor, the currently assigned absolute copy number of the particular segment based at least in part on absolute copy numbers of the nearest upstream segment and the nearest downstream segment (e.g., determining the absolute copy numbers of the nearest upstream segment and the nearest downstream segment to be equal and, based thereon, replacing the currently assigned absolute copy number of the particular segment with the absolute copy number of the nearest upstream and downstream segments).

[0078] In some aspects, the present disclosure provides a method (e.g., for correcting errors in copy numbers assigned to segments of a tumor genome) comprising: (a) obtaining [e.g., receiving and / or accessing (e.g., based on a database; e.g., detecting, by a processor of a computing device (e.g., based on the normal sequencing data)], by a processor of a computing device, a list of tumor genome segments and, for each segment of the list, a currently assigned absolute copy number; (b) for each particular segment of at least a portion of the list: identifying, by the processor, a plurality of genomic neighbor segments, each genomic neighbor segment a segment of the list having been determined to be located within a threshold number of segments and / or a threshold bases [e.g., upstream (e.g., in a 5’ direction) and / or downstream (e.g., in a 3’ direction)] of the particular segment, and including the plurality of genomic neighbors together with the particular segment in a local genomic neighborhood; determining, by the processor, a representative copy number of the local genomic neighborhood based on absolute copy numbers- 17 - 13247547v 1Attorney Docket No.: 2013237-0970of segments within the local genomic neighborhood (e.g., as a mode, an average, a median, an interquartile range, etc.); and updating, by the processor, the currently assigned absolute copy number of the particular segment to a value determined based on the representative copy number of the local genomic neighborhood, thereby determining error-corrected copy numbers for at least the portion of the list; and (c) storing and / or providing, by the processor, for display and / or further processing, the error corrected copy numbers for the portion of the list of tumor genome segments.

[0079] In some aspects, the present disclosure provides a system (e.g., for determining genomic and biophysical properties of a tumor sample from sequencing data) comprising: a processor of a computing device; memory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to: (a) obtain tumor sequencing data for a tumor sample obtained from a subject [e.g., the tumor sequencing data comprising a plurality of tumor reads, each tumor read representing a partial nucleotide sequence determined by sequencing the tumor sample]; (b) identify, using the tumor sequencing data, a plurality of primary balanced heterozygous segments within a tumor genome for the subject, each primary balanced heterozygous segment having been determined to: (i) comprise one or more heterozygous SNPs, at least a portion of which are determined to be balanced [e.g., are determined to have an equal number of copies of corresponding maternal and paternal alleles (e.g., via a balanced test and / or contour filtering refinement)] within the tumor genome; and (ii) have an absolute copy number equal to a particular, unknown-but-to-be-determined, primary copy number (e.g., the exact value of the copy number is initially unknown, but all the primary balanced heterozygous segments are determined to have same absolute copy number, e.g., and known to be even, e.g., by virtue of the PBHSs being balanced); (c) determine, based on the tumor sequencing data and the primary balanced heterozygous segments, an estimated tumor sample purity and / or bound thereon for the tumor sample; and (d) store and / or provide the estimated tumor sample purity and / or bound thereon for display and / or further processing.

[0080] In some aspects, the present disclosure provides a system (e.g., for correcting errors in copy numbers assigned to segments of a tumor genome) comprising: a processor of a computing device; memory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to: (a) obtain [e.g., receiving and / or accessing- 18 - 13247547v 1Attomey Docket No.: 2013237-0970(e.g., based on a database; e.g., detecting (e.g., based on the normal sequencing data)], a list of nominally heterozygous segments within a tumor genome of a subject, and each segment of the list corresponding to a heterozygous segment within a normal genome of the subject; (b) for each nominally heterozygous segment of the list: obtain [e.g., receiving and / or accessing; e.g., determining] a currently assigned absolute copy number representing a number of copies of the segment having been determined to be present in the tumor genome; and obtain [e.g., receiving and / or accessing; e.g., determining] a balanced state indicating whether the segment is balanced, thereby determining, for the list of nominally heterozygous segments, balanced states of corresponding tumor genome segments; (c) identify, based on the balanced states of the list of nominally heterozygous segments, a subset of balanced heterozygous segments; (d) identify, within the subset of balanced heterozygous segments, one or more segments with odd- valued corresponding absolute copy numbers as potential copy number assignments errors; (e) update absolute copy numbers of the potential copy number assignment errors to an error-corrected value, wherein each error-corrected value is an even number (e.g., integer); and (f) store and / or provide the error corrected values for display and / or further processing.

[0081] In some embodiments, at step (e), the instructions cause the processor to, for at least one particular segment identified as a potential copy number assignment error: receive and / or determine, for the particular segment, a first node assignment likelihood corresponding to the currently assigned absolute copy number of the particular segment and a second node assignment likelihood corresponding to a second absolute copy number, wherein the second absolute copy number is even-valued and wherein the first and second node assignment likelihoods represent a likelihood of the particular segment having the currently assigned and second absolute copy number, respectively; and based on the first and second node assignment likelihoods, update the currently assigned absolute copy number by replacing it with the second absolute copy number.

[0082] In some embodiments, at step (e), the instructions cause the processor to, for at least one particular segment identified as a potential copy number assignment error: identify, within the list of nominally heterozygous segments, one or more genomic neighbor segments with correctly assigned copy numbers, the one or more genomic neighbors comprising (i) a nearest upstream segment (e.g., in a 5’ direction) and (ii) a nearest downstream segment (e.g., in- 19 - 13247547v 1Attomey Docket No.: 2013237-0970a 3’ direction); and update the currently assigned absolute copy number of the particular segment based at least in part on absolute copy numbers of the nearest upstream segment and the nearest downstream segment (e.g., determining the absolute copy numbers of the nearest upstream segment and the nearest downstream segment to be equal and, based thereon, replacing the currently assigned absolute copy number of the particular segment with the absolute copy number of the nearest upstream and downstream segments).

[0083] In some aspects, the present disclosure provides a system (e.g., for correcting errors in copy numbers assigned to segments of a tumor genome) comprising: a processor of a computing device; memory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to: (a) obtain [e.g., receive and / or access (e.g., based on a database; e.g., detect (e.g., based on the normal sequencing data)], a list of tumor genome segments and, for each segment of the list, a currently assigned absolute copy number; (b) for each particular segment of at least a portion of the list: identify a plurality of genomic neighbor segments, each genomic neighbor segment a segment of the list having been determined to be located within a threshold number of segments and / or a threshold bases [e.g., upstream (e.g., in a 5’ direction) and / or downstream (e.g., in a 3’ direction)] of the particular segment, and including the plurality of genomic neighbors together with the particular segment in a local genomic neighborhood; determine a representative copy number of the local genomic neighborhood based on absolute copy numbers of segments within the local genomic neighborhood (e.g., as a mode, an average, a median, an interquartile range, etc.); and update the currently assigned absolute copy number of the particular segment to a value determined based on the representative copy number of the local genomic neighborhood, thereby determining error-corrected copy numbers for at least the portion of the list; and (c) store and / or provide the error corrected copy numbers for the portion of the list of tumor genome segments for display and / or further processing.

[0084] In some aspects, the present disclosure provides a method of producing an immunotherapy construct (e.g., a cancer vaccine) for a subject, the method comprising: detecting a plurality of candidate mutations (e.g., somatic mutations; e.g., non-synonymous somatic mutations) in tumor cells from the subject using a method or system disclosed herein; and synthesizing, as the immunotherapy construct, a polyribonucleotide encoding a plurality of- 20 - 13247547v 1Attorney Docket No.: 2013237-0970neoepitopes, each encoded neoepitope corresponding to a candidate mutation of the plurality of candidate mutations.

[0085] In some aspects, the present disclosure provides a method comprising determining a subset of T-cells within a population of T-cells that are capable of specifically binding a plurality of complexes, wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, and wherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer-specific mutations detected using a method or system disclosed herein; and enriching for (e.g., expanding) the subset of T-cells that are capable of specifically binding the plurality of complexes.

[0086] In some aspects, the present disclosure provides a method comprising administering to a subject a population of T-cells, wherein the population of T cells are capable of specifically binding a plurality of complexes, wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, wherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer-specific mutations detected using a method or system disclosed herein, and wherein the genome of at least some of the subject’s cells comprises a subset (e.g., all) of the cancer-specific mutations.

[0087] In some aspects, the present disclosure provides a method comprising determining a subset of TILs within a population of TILs that are capable of specifically binding a plurality of complexes, wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, and wherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer- specific mutations detected using a method or system disclosed herein; and enriching for (e.g., expanding) the subset of TILs that are capable of specifically binding the plurality of complexes.- 21 - 13247547v 1Attomey Docket No.: 2013237-0970

[0088] In some aspects, the present disclosure provides a method comprising administering to a subject a population of TILs, wherein the population of TILs are capable of specifically binding a plurality of complexes, wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancerspecific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, wherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer- specific mutations detected using a method or system disclosed herein, and wherein the genome of at least some of the subject’s cells comprises a subset (e.g., all) of the cancer-specific mutations.

[0089] In some embodiments, provided methods comprise obtaining a tumor sample from the subject and using the tumor sample to detect the plurality of candidate mutations (e.g., by sequencing the tumor sample).

[0090] In some embodiments, provided methods (e.g., further) comprise obtaining a normal sample from the subject and using the normal sample (e.g., together with the tumor sample) to detect the plurality of cancer mutations (e.g., by sequencing the normal sample).

[0091] In some embodiments, provided methods comprise sequencing the tumor and / or normal sample (e.g., in replicates).

[0092] In some aspects, the present disclosure provides a pharmaceutical composition comprising a polyribonucleotide encoding (e.g., a polyepitopic vaccine construct comprising) a plurality of neoepitopes, wherein at least a portion of the plurality of neoepitopes are individualized (e.g., patient- specific) neoepitopes corresponding to (e.g., encoded by a nucleotide sequence comprising) a somatic mutation from a population of candidate mutations detected in a tumor sample of a subject and selected for inclusion in the pharmaceutical composition using a method or system described herein (e.g., each particular neoepitope encoded by a nucleotide sequence comprising one or more candidate mutation detected using a method or system disclosed herein).

[0093] In some aspects, the present disclosure provides an individualized pharmaceutical composition (e.g., associated with and / or intended for administration to a particular subject) comprising a polyribonucleotide encoding (e.g., a polyepitopic vaccine construct comprising) a plurality of neoepitopes for a particular subject, wherein at least a portion of the neoepitopes are - 22 - 13247547v 1Attorney Docket No.: 2013237-0970individualized (e.g., patient- specific) neoepitopes corresponding to (e.g., encoded by a nucleotide sequence comprising) a somatic mutation from a population of candidate mutations detected in a tumor sample from the particular subject and selected for inclusion in the pharmaceutical composition using a method or system disclosed herein.

[0094] In some aspects, the present disclosure provides a population of T-cells, wherein the population of T-cells is capable of specifically binding to a plurality of complexes, wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, and wherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancerspecific mutations using a method or system disclosed herein.

[0095] In some aspects, the present disclosure provides a T-cell capable of specifically binding to a complex, wherein the complex comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, wherein the one or more cancer-specific mutations are detected using a method or system disclosed herein.

[0096] In some aspects, the present disclosure provides a population of T-cells, wherein the population of T-cells is capable of specifically binding to a plurality of complexes, wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, and wherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancerspecific mutations detected using a method or system disclosed herein.

[0097] In some aspects, the present disclosure provides a T-cell capable of specifically binding to a complex, wherein the complex comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, wherein the one or more cancer-specific mutations are within the selected subset of candidate mutations detected using a method or system disclosed herein.- 23 - 13247547v 1Attomey Docket No.: 2013237-0970

[0098] In some aspects, the present disclosure provides a T-cell receptor capable of specifically binding to a complex, wherein the complex comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, wherein the one or more cancer-specific mutations are within the selected subset of candidate mutations detected using a method or system disclosed herein.

[0099] In some aspects, the present disclosure provides a chimeric antigen receptor (CAR) capable of specifically binding to a complex, wherein the complex comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer-specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, wherein the one or more cancer- specific mutations are within the selected subset of candidate mutations detected using a method or system disclosed herein.

[0100] In some aspects, the present disclosure provides a method (e.g., for analyzing sequencing data to detect one or more potential mutations), the method comprising: (a) obtaining, by the processor, sequencing data for a biological sample obtained from a subject, the sequencing data comprising a plurality of reads, each read represents a partial nucleotide sequence of a sequenced genome associated with the sample (e.g., and each read aligned to a reference genome); (b) selecting, by the processor, a particular site within the sequenced genome and determining, for the particular site: (i) a corresponding coverage (e.g., a total number of reads mapping to the site), (ii) a corresponding number of reads mapping to a particular alternate allele, different from a wild-type allele, and (iii) an error rate; (c) determining, by the processor, for the particular selected site, a false positive probability representing a probability of observing at least the corresponding number of reads mapping to the particular alternate allele without a true underlying mutation being present in the sequenced genome, given the corresponding coverage and the error rate; (d) determining, by the processor, for the particular selected site, a missed-detection probability representing a probability of no more than the corresponding number of reads mapping to the particular alternate allele with a true underlying mutation being present in the sequenced genome, given the corresponding coverage [e.g., and one or more of: an assumed absolute copy number of the particular selected site (e.g., of two), an assumed fractional zygosity (e.g., of 0.5), an assumed tumor content (e.g., of 0.75), and an assumed error rate (e.g.,- 24 - 13247547v 1Attorney Docket No.: 2013237-0970of 0)]; and (e) determining, by the processor, a false discovery rate based on the false positive probability, the missed-detection probability, and representative overall mutation likelihood for the sample; (d) determining, by the processor, the particular selected site to be a potential mutation based on the false positive probability, the missed-detection probability and the false discovery rate; and (e) responsive to determining the selected site to be a potential mutation, storing, by the processor, the selected site in a list of detected mutations and / or providing, by the processor, the selected site for further processing.

[0101] In certain embodiments, a false positive probability is determined so as to account for (i) the corresponding number of reads mapping to the particular alternate allele and (ii) a distribution of the maximum of the three noise alleles (e.g., based on extreme event statistics).

[0102] In certain embodiments, step (b) comprises, for each of a number of noise reads ranging from (i) the corresponding number for reads mapping to the particular alternate allele to (ii) the corresponding coverage, determining, by the processor, a plurality of possible configurations for distributing the number of noise reads across three possible non-wild-type alleles and, for each possible configuration, a corresponding probability.

[0103] In certain embodiments, step (d) comprises comparing the false positive probability to a first threshold value [e.g., a predetermined threshold value (e.g., a value determined to reflect a per-exome error rate (e.g., one (e.g., 2x10-8), two, five, ten errors per exome)].

[0104] In certain embodiments, provided methods comprise: determining, by the processor, the false positive probability rate to be less than or equal to the first threshold value and, responsive to determining the false positive probability rate to be less than or equal to the first threshold value, determining, by the processor, the particular selected site to be a potential mutation.

[0105] In certain embodiments, step (d) comprises comparing the missed-detection probability rate to a second threshold value (e.g., 0.01, 0.02, 0.05, 0.1, etc.).

[0106] In certain embodiments, provided methods comprise: determining, by the processor, the missed-detection probability rate to be greater than the second threshold value; and based at least in part on the determining the missed-detection probability rate to be greater- 25 - 13247547v 1Attomey Docket No.: 2013237-0970than the second threshold value, determining, by the processor, the particular selected site to be a potential mutation.

[0107] In certain embodiments, provided methods comprise: responsive to the determining the missed-detection probability rate to be greater than the second threshold value, comparing the false discovery rate to a third threshold value (e.g., 1%, 5%, 10%, etc.).

[0108] In certain embodiments, provided methods comprise: determining, by the processor, the false discovery rate to be less than or equal to (e.g., less than) the third threshold value; and responsive to the determining the false discovery rate less than or equal to (e.g., less than) the third threshold value, determining, by the processor, the particular selected site to be a potential mutation.

[0109] In certain embodiments, a sample is a tumor sample and the representative overall mutation likelihood is selected based on a particular cancer type for the tumor sample.

[0110] In certain embodiments, a representative overall mutation likelihood is a value determined to represent a particular number of mutations per exome [e.g., 1,000 SNVs per exome, 2,000 SNVs per exome, 5,000 SNVs per exome (e.g., 10-4), 10,000 SNVs per exome, etc.].

[0111] In some aspects, the present disclosure provides systems comprising a processor of a computing device; and memory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of various aspects and embodiments described herein.

[0112] Features of embodiments described with respect to one aspect of the invention may be applied with respect to another aspect of the invention.BRIEF DESCRIPTION OF THE DRAWING

[0113] The foregoing and other objects, aspects, features, and advantages of the present disclosure will become more apparent and better understood by referring to the following description taken in conjunction with the accompanying drawings, in which:- 26 - 13247547v 1Attorney Docket No.: 2013237-0970

[0114] FIG. 1 is a block flow diagram showing an example process for creating a personalized cancer immunotherapy, according to an illustrative embodiment.

[0115] FIG. 2 is a block flow diagram showing an example process for obtaining sequencing data from tumor and / or normal sample(s), according to an illustrative embodiment.

[0116] FIG. 3 is a schematic showing a model of a normal genome and a tumor genome of a subject, according to an illustrative embodiment.

[0117] FIG. 4 is a schematic of a tumor sample, according to an illustrative embodiment.

[0118] FIG. 5 is a diagram illustrating interrelation between parameters and / or measurements characterizing tumor sample properties, according to an illustrative embodiment.

[0119] FIG. 6A is a block flow diagram of an example tumor deconvolution process, according to an illustrative embodiment.

[0120] FIG. 6B is a block flow diagram of an example process for using a particular subpopulation of segments (e.g., primary balanced heterozygous segments) to perform tumor deconvolution and / or facilitate mutation detection and / or characterization, according to an illustrative embodiment.

[0121] FIG. 7 is a schematic illustrating certain categories of segments within a genome (e.g., a tumor genome and / or a normal genome) and approaches for their identification, according to an illustrative embodiment.

[0122] FIG. 8 is a block flow diagram of an example process for using sequencing data to identify a particular subpopulation of tumor genome segments (e.g., primary balanced heterozygous segments), according to an illustrative embodiment.

[0123] FIG. 9 is a block flow diagram of an example process for using multiple tumor models for obtaining purity estimates and / or estimated purity bounds, according to an illustrative embodiment.

[0124] FIG. 10 is a block flow diagram of an example process for performing tumor deconvolution, according to an illustrative embodiment.- 27 - 13247547v 1Attorney Docket No.: 2013237-0970

[0125] FIG. 11 is a block flow diagram of an example process for incorporating purity estimates and / or tumor genomic properties into a mutation detection process, according to an illustrative embodiment.

[0126] FIG. 12A is an example plot illustrating a step in an example tumor deconvolution workflow, including how estimates of tumor sample properties such as purity, absolute copy number, and allele specific copy number assignments, can be used to detect and characterize cancer- specific mutations, such as SNV events, according to an illustrative embodiment.

[0127] FIG. 12B is an example plot illustrating a step in an example tumor deconvolution workflow, including how estimates of tumor sample properties such as purity, absolute copy number, and allele specific copy number assignments, can be used to detect and characterize cancer- specific mutations, such as SNV events, according to an illustrative embodiment.

[0128] FIG. 12C is an example plot illustrating a step in an example tumor deconvolution workflow, including how estimates of tumor sample properties such as purity, absolute copy number, and allele specific copy number assignments, can be used to detect and characterize cancer- specific mutations, such as SNV events, according to an illustrative embodiment.

[0129] FIG. 12D is an example plot illustrating a step in an example tumor deconvolution workflow, including how estimates of tumor sample properties such as purity, absolute copy number, and allele specific copy number assignments, can be used to detect and characterize cancer- specific mutations, such as SNV events, according to an illustrative embodiment.

[0130] FIG. 12E is an example plot illustrating a step in an example tumor deconvolution workflow, including how estimates of tumor sample properties such as purity, absolute copy number, and allele specific copy number assignments, can be used to detect and characterize cancer- specific mutations, such as SNV events, according to an illustrative embodiment.- 28 - 13247547v 1Attomey Docket No.: 2013237-0970

[0131] FIG. 12F is an example plot illustrating a step in an example tumor deconvolution workflow, including how estimates of tumor sample properties such as purity, absolute copy number, and allele specific copy number assignments, can be used to detect and characterize cancer- specific mutations, such as SNV events, according to an illustrative embodiment.

[0132] FIG. 12G is an example plot illustrating a step in an example tumor deconvolution workflow, including how estimates of tumor sample properties such as purity, absolute copy number, and allele specific copy number assignments, can be used to detect and characterize cancer- specific mutations, such as SNV events, according to an illustrative embodiment.

[0133] FIG. 13 is a block flow diagram of an example process for estimating tumor sample characteristics (e.g., genomic and / or biophysical properties) and using estimated tumor sample characteristics to detect and characterize cancer mutations, according to an illustrative embodiment.

[0134] FIG. 14A is a block flow diagram of an example process for correcting copy number assignment errors, according to an illustrative embodiment.

[0135] FIG. 14B is a block flow diagram of an example process for correcting copy number assignment errors, according to an illustrative embodiment.

[0136] FIG. 15 is a schematic illustrating an example construct encoding selected neoantigens, according to an illustrative embodiment.

[0137] FIG. 16 is a block diagram of an exemplary cloud computing environment, used in certain embodiments.

[0138] FIG. 17 is a block diagram of an example computing device and an example mobile computing device used in certain embodiments.

[0139] FIG. 18A is an example plot illustrating a step in an example tumor deconvolution workflow, including how estimates of tumor sample properties such as purity, absolute copy number, and allele specific copy number assignments, can be used to detect and characterize cancer- specific mutations, such as SNV events, according to an illustrative embodiment.- 29 - 13247547v 1Attomey Docket No.: 2013237-0970

[0140] FIG. 18B is an example plot illustrating a step in an example tumor deconvolution workflow, including how estimates of tumor sample properties such as purity, absolute copy number, and allele specific copy number assignments, can be used to detect and characterize cancer- specific mutations, such as SNV events, according to an illustrative embodiment.

[0141] FIG. 18C is an example plot illustrating a step in an example tumor deconvolution workflow, including how estimates of tumor sample properties such as purity, absolute copy number, and allele specific copy number assignments, can be used to detect and characterize cancer- specific mutations, such as SNV events, according to an illustrative embodiment.

[0142] FIG. 18D is an example plot illustrating a step in an example tumor deconvolution workflow, including how estimates of tumor sample properties such as purity, absolute copy number, and allele specific copy number assignments, can be used to detect and characterize cancer- specific mutations, such as SNV events, according to an illustrative embodiment.

[0143] FIG. 18E is an example plot illustrating a step in an example tumor deconvolution workflow, including how estimates of tumor sample properties such as purity, absolute copy number, and allele specific copy number assignments, can be used to detect and characterize cancer- specific mutations, such as SNV events, according to an illustrative embodiment.

[0144] FIG. 18F is an example plot illustrating a step in an example tumor deconvolution workflow, including how estimates of tumor sample properties such as purity, absolute copy number, and allele specific copy number assignments, can be used to detect and characterize cancer- specific mutations, such as SNV events, according to an illustrative embodiment.

[0145] FIG. 18G is an example plot illustrating a step in an example tumor deconvolution workflow, including how estimates of tumor sample properties such as purity, absolute copy number, and allele specific copy number assignments, can be used to detect and characterize cancer- specific mutations, such as SNV events, according to an illustrative embodiment.- 30 - 13247547v 1Attorney Docket No.: 2013237-0970

[0146] FIG. 19 is an example of a graphical representation of tumor genome segments with heterozygous SNP balanced state calculation results overlaid thereon, according to an illustrative embodiment.

[0147] FIG. 20A is a graph providing results of a steps in an example tumor deconvolution process for two melanoma samples and an ovarian cancer sample, according to an illustrative embodiment.

[0148] FIG. 20B is a graph providing results of a steps in an example tumor deconvolution process for two melanoma samples and an ovarian cancer sample, according to an illustrative embodiment.

[0149] FIG. 20C is a graph providing results of a steps in an example tumor deconvolution process for two melanoma samples and an ovarian cancer sample, according to an illustrative embodiment.

[0150] FIG. 20D is a graph providing results of a steps in an example tumor deconvolution process for two melanoma samples and an ovarian cancer sample, according to an illustrative embodiment.

[0151] FIG. 20E is a graph providing results of a steps in an example tumor deconvolution process for two melanoma samples and an ovarian cancer sample, according to an illustrative embodiment.

[0152] FIG. 20F is a graph providing results of a steps in an example tumor deconvolution process for two melanoma samples and an ovarian cancer sample, according to an illustrative embodiment.

[0153] FIG. 20G is a graph providing results of a steps in an example tumor deconvolution process for two melanoma samples and an ovarian cancer sample, according to an illustrative embodiment.

[0154] FIG. 20H is a graph providing results of a steps in an example tumor deconvolution process for two melanoma samples and an ovarian cancer sample, according to an illustrative embodiment.- 31 - 13247547v 1Attomey Docket No.: 2013237-0970

[0155] FIG. 201 is a graph providing results of a steps in an example tumor deconvolution process for two melanoma samples and an ovarian cancer sample, according to an illustrative embodiment.

[0156] FIG. 20J is a graph providing results of a steps in an example tumor deconvolution process for two melanoma samples and an ovarian cancer sample, according to an illustrative embodiment.

[0157] FIG. 20K is a graph providing results of a steps in an example tumor deconvolution process for two melanoma samples and an ovarian cancer sample, according to an illustrative embodiment.

[0158] FIG. 20L is a graph providing results of a steps in an example tumor deconvolution process for two melanoma samples and an ovarian cancer sample, according to an illustrative embodiment.

[0159] FIG. 21A is a state diagram for sequencing a single non-mutated site, according to an illustrative embodiment.

[0160] FIG. 21B is a state diagram for selecting and sequencing a major allele associated with a particular heterozygous SNP, according to an illustrative embodiment.

[0161] FIG. 21C is a state diagram for selecting and sequencing a minor allele associated with a particular heterozygous SNP, according to an illustrative embodiment.

[0162] FIG. 22A is a graph demonstrating use of a variance stabilizing transformation to produce a normal distribution of residual errors, according to an illustrative embodiment.

[0163] FIG. 22B is a graph demonstrating use of a variance stabilizing transformation to produce a normal distribution of residual errors, according to an illustrative embodiment.

[0164] FIG. 22C is a graph demonstrating use of a variance stabilizing transformation to produce a normal distribution of residual errors, according to an illustrative embodiment.

[0165] FIG. 23 is a graph showing an example CNV cluster plot, according to an illustrative embodiment.

[0166] FIG. 24A is a graph showing an example CNV cluster plot determined for an ovarian cancer sample, according to an illustrative embodiment.- 32 - 13247547v 1Attorney Docket No.: 2013237-0970

[0167] FIG. 24B is a graph showing an example CNV cluster plot determined for an ovarian cancer sample, according to an illustrative embodiment.

[0168] FIG. 24C is a graph showing an example CNV cluster plot determined for an ovarian cancer sample, according to an illustrative embodiment.

[0169] FIG. 24D is a graph showing an example CNV cluster plot determined for an ovarian cancer sample, according to an illustrative embodiment.

[0170] FIG. 24E is a graph showing an example CNV cluster plot determined for an ovarian cancer sample, according to an illustrative embodiment.

[0171] FIG. 24F is a graph showing an example CNV cluster plot determined for an ovarian cancer sample, according to an illustrative embodiment.

[0172] FIG. 25A is a graph showing an example CNV cluster plot determined for a melanoma sample, according to an illustrative embodiment.

[0173] FIG. 25B is a graph showing an example CNV cluster plot determined for a melanoma sample, according to an illustrative embodiment.

[0174] FIG. 25C is a graph showing an example CNV cluster plot determined for a glioblastoma sample, according to an illustrative embodiment.

[0175] FIG. 25D is a graph showing an example CNV cluster plot determined for glioblastoma sample, according to an illustrative embodiment.

[0176] FIG. 25E is a graph showing an example CNV cluster plot determined for a breast cancer sample, according to an illustrative embodiment.

[0177] FIG. 25F is a graph showing an example CNV cluster plot determined for a breast cancer sample, according to an illustrative embodiment.

[0178] FIG. 26A is a graph showing an example CNV cluster plot determined for a melanoma sample, according to an illustrative embodiment.

[0179] FIG. 26B is a graph showing an example fan plot determined for the melanoma sample shown in FIG.26A, according to an illustrative embodiment.- 33 - 13247547v 1Attomey Docket No.: 2013237-0970

[0180] FIG. 27A is a graph showing an example CNV cluster plot with rectangular decision boundaries overlaid, according to an illustrative embodiment.

[0181] FIG. 27B is a graph showing an example CNV cluster plot with rectangular decision boundaries overlaid, according to an illustrative embodiment.

[0182] FIG. 28A is a graph showing results at a steps in an example CNV-based purity estimation for a melanoma sample, according to an illustrative embodiment.

[0183] FIG. 28B is a graph showing results at a steps in an example CNV-based purity estimation for a melanoma sample, according to an illustrative embodiment..

[0184] FIG. 28C is a graph showing results at a steps in an example CNV-based purity estimation for a melanoma sample, according to an illustrative embodiment.

[0185] FIG. 28D is a graph showing results at a steps in an example CNV-based purity estimation for a melanoma sample, according to an illustrative embodiment.

[0186] FIG. 28E is a graph showing results at a steps in an example CNV-based purity estimation for a melanoma sample, according to an illustrative embodiment.

[0187] FIG. 28F is a graph showing results at a steps in an example CNV-based purity estimation for a melanoma sample, according to an illustrative embodiment.

[0188] FIG. 28G is a graph showing results at a steps in an example CNV-based purity estimation for a melanoma sample, according to an illustrative embodiment.

[0189] FIG. 28H is a graph showing results at a steps in an example CNV-based purity estimation for a melanoma sample, according to an illustrative embodiment.

[0190] FIG. 29A is a graph showing results at a stages in an example CNV-based purity estimation for a melanoma sample, according to an illustrative embodiment.

[0191] FIG. 29B is a graph showing results at a stages in an example CNV-based purity estimation for a melanoma sample, according to an illustrative embodiment.

[0192] FIG. 29C is a graph showing results at a stages in an example CNV-based purity estimation for a melanoma sample, according to an illustrative embodiment.- 34 - 13247547v 1Attorney Docket No.: 2013237-0970

[0193] FIG. 29D is a graph showing results at a stages in an example CNV-based purity estimation for a melanoma sample, according to an illustrative embodiment.

[0194] FIG. 29E is a graph showing results at a stages in an example CNV-based purity estimation for a melanoma sample, according to an illustrative embodiment.

[0195] FIG. 29F is a graph showing results at a stages in an example CNV-based purity estimation for a melanoma sample, according to an illustrative embodiment.

[0196] FIG. 29G is a graph showing results at a stages in an example CNV-based purity estimation for a melanoma sample, according to an illustrative embodiment.

[0197] FIG. 29H is a graph showing results at a stages in an example CNV-based purity estimation for a melanoma sample, according to an illustrative embodiment.

[0198] FIG. 30A a graphs of a cluster plot of tumor genome segments, according to an illustrative embodiment.

[0199] FIG. 30B a set of graphs showing example probability distribution functions for allele frequencies and residual errors used to determine a ^-statistic value, according to an illustrative embodiment.

[0200] FIG. 31 is a graph showing a CNV cluster plot with outliers identified for exclusion from a ^-statistic calculation marked, according to an illustrative embodiment.

[0201] FIG. 32 is a schematic illustrating a cluster density metric, according to an illustrative embodiment.

[0202] FIG. 33 is a graph showing a decision plane used in a multivariate classifier, according to an illustrative embodiment.

[0203] FIG. 34A is a graphical display showing results of an error correction procedure that identifies and corrects certain errors in copy number assignments, according to an illustrative embodiment.

[0204] FIG. 34B is a graphical display showing results of an error correction procedure that identifies and corrects certain errors in copy number assignments, according to an illustrative embodiment.- 35 - 13247547v 1Attomey Docket No.: 2013237-0970

[0205] FIG. 35A is a set of graphs demonstrating performance of an error correction procedure, used in certain embodiments.

[0206] FIG. 35B is a set of graphs demonstrating performance of an error correction procedure, used in certain embodiments.

[0207] FIG. 35C is a set of graphs demonstrating performance of an error correction procedure, used in certain embodiments.

[0208] FIG. 35D is a set of graphs demonstrating performance of an error correction procedure, used in certain embodiments.

[0209] FIG. 36A is a graph plotting parity error rates as a function of primary copy number (CApc) for a melanoma sample, according to an illustrative embodiment.

[0210] FIG. 36B is a graph plotting parity error rates as a function of primary copy number (CApc) for a melanoma sample, according to an illustrative embodiment.

[0211] FIG. 37 is a block flow diagram showing a decision tree for selecting between various purity estimation results, according to an illustrative embodiment.

[0212] FIG. 38A is a graph plotting a histogram of cellularity values determined for SNV mutations detected from a melanoma sample, according to an illustrative embodiment.

[0213] FIG. 38B is a graph plotting a histogram of cellularity values determined for SNV mutations detected from an ovarian cancer sample, according to an illustrative embodiment.

[0214] FIG. 39 is a block flow diagram of an example process for detecting putative mutations, according to an illustrative embodiment.

[0215] FIG. 40A is a graph providing results comparing an example mutation detection process with a binomial classifier, according to an illustrative embodiment.

[0216] FIG. 40B is a graph providing results comparing an example mutation detection process with a binomial classifier, according to an illustrative embodiment.

[0217] FIG. 40C is a graph providing results comparing an example mutation detection process with a binomial classifier, according to an illustrative embodiment.- 36 - 13247547v 1Attorney Docket No.: 2013237-0970

[0218] FIG. 41A is a graph showing results of a gender estimation procedure for a sample from a male subject, used in certain embodiments.

[0219] FIG. 41B is a graph showing results of a gender estimation procedure for a sample from a female subject, used in certain embodiments.

[0220] FIG. 42 is a snapshot of a cancer genome graphical display generated in accordance with certain embodiments.

[0221] FIG. 43A is a set of cancer genome graphical displays representing determined properties of a tumor genome for chromosomes 1 through 15 of a small cell lung cancer cell line, according to an illustrative embodiment.

[0222] FIG. 43B is a set of cancer genome graphical displays representing determined properties of a tumor genome for chromosomes 2 through 22, X, and Y of the small cell lung cancer cell line of FIG.43A, according to an illustrative embodiment.

[0223] FIG. 43C is a plot of spectral karyotyping analysis (SKY) results for the small cell lung cancer cell line of FIGs.43A and 43B, according to an illustrative embodiment.

[0224] FIG. 44A is a set of cancer genome graphical displays representing determined properties of a tumor genome for chromosomes 1 through 15 of a melanoma sample, according to an illustrative embodiment.

[0225] FIG. 44B is a set of cancer genome graphical displays representing determined properties of a tumor genome for chromosomes 2 through 22, X, and Y of the melanoma sample of FIG. 44A, according to an illustrative embodiment.

[0226] FIG. 45 is a set of graphs showing results for a normal versus normal analysis for a melanoma cohort to determine the background noise of the embodiment compared to published mutation callers, according to an illustrative embodiment.

[0227] FIG. 46 is a set of graphs showing results for a cross replicate analysis of a melanoma cohort compared to published mutation callers, according to an illustrative embodiment.- 37 - 13247547v 1Attorney Docket No.: 2013237-0970

[0228] FIG. 47A is a set of graph showing results for sensitivity and experimentally determined accuracy of an embodiment of technologies described herein for a melanoma cohort compared to published mutation callers, according to an illustrative embodiment.

[0229] FIG. 47B is a set of graph showing results for sensitivity and experimentally determined accuracy of an embodiment of technologies described herein for a melanoma cohort compared to published mutation callers, according to an illustrative embodiment.

[0230] FIG. 47C is a set of graph showing results for sensitivity and experimentally determined accuracy of an embodiment of technologies described herein for a melanoma cohort compared to published mutation callers, according to an illustrative embodiment.

[0231] FIG. 47D is a set of graph showing results for sensitivity and experimentally determined accuracy of an embodiment of technologies described herein for a melanoma cohort compared to published mutation callers, according to an illustrative embodiment.

[0232] FIG. 47E is a set of graph showing results for sensitivity and experimentally determined accuracy of an embodiment of technologies described herein for a melanoma cohort compared to published mutation callers, according to an illustrative embodiment.

[0233] FIG. 47F is a set of graph showing results for sensitivity and experimentally determined accuracy of an embodiment of technologies described herein for a melanoma cohort compared to published mutation callers, according to an illustrative embodiment.

[0234] FIG. 48A is a graph showing results demonstrating sensitivity determined for a cohort of ovarian tumor samples where mutations were confirmed by an orthogonal technology, according to an illustrative embodiment.

[0235] FIG. 48A is a graph showing results demonstrating sensitivity determined for a cohort of ovarian tumor samples where mutations were confirmed by an orthogonal technology, according to an illustrative embodiment.

[0236] FIG. 49A is a set of graph showing VAF distributions of true positives versus false positives for a melanoma cohort comparing an embodiment of technologies described herein with published mutation callers, according to an illustrative embodiment.- 38 - 13247547v 1Attomey Docket No.: 2013237-0970

[0237] FIG. 49B is a set of graph showing VAF distributions of true positives versus false positives for a melanoma cohort comparing an embodiment of technologies described herein with published mutation callers, according to an illustrative embodiment.

[0238] FIG. 49C is a set of graph showing VAF distributions of true positives versus false positives for a melanoma cohort comparing an embodiment of technologies described herein with published mutation callers, according to an illustrative embodiment.

[0239] FIG. 49D is a set of graph showing VAF distributions of true positives versus false positives for a melanoma cohort comparing an embodiment of technologies described herein with published mutation callers, according to an illustrative embodiment.

[0240] FIG. 49E is a set of diagrams showing intersection between predicted mutations for a melanoma cohort compared to published mutation callers, according to an illustrative embodiment.

[0241] FIG. 49F is a Venn diagram showing correspondence between MiSeq VAF, a nominal admixture VAF,and NGS VAF for the experimental admixture series used to determine the limit of detection of the MiSeq target confirmation assay, according to an illustrative embodiment.

[0242] FIG. 49G is a Venn diagram showing correspondence between MiSeq VAF, a nominal admixture VAF,and NGS VAF for the experimental admixture series used to determine the limit of detection of the MiSeq target confirmation assay, according to an illustrative embodiment.

[0243] FIG. 50A is a graph showing results of a tumor modeling solution of P001 resulting in a CNV-based purity estimation, according to an illustrative embodiment.

[0244] FIG. 50B is a graph showing results of a tumor modeling solution of P001 resulting in a CNV-based purity estimation, according to an illustrative embodiment.

[0245] FIG. 50C is a graph showing results of a tumor modeling solution of P001 resulting in a CNV-based purity estimation, according to an illustrative embodiment.

[0246] FIG. 50D is a graph showing results of a tumor modeling solution of P001 resulting in a CNV-based purity estimation, according to an illustrative embodiment.- 39 - 13247547v 1Attomey Docket No.: 2013237-0970

[0247] FIG. 50E is a graph showing results of a tumor modeling solution of P001 resulting in a CNV-based purity estimation, according to an illustrative embodiment.

[0248] FIG. 50F is a graph showing results of a tumor modeling solution of P001 resulting in a CNV-based purity estimation, according to an illustrative embodiment.

[0249] FIG. 50G is a graph showing results of a tumor modeling solution of P001 resulting in a CNV-based purity estimation, according to an illustrative embodiment.

[0250] FIG. 50H is a screenshot of a genome view display for P001, according to an illustrative embodiment.

[0251] FIG. 51A is a graphs showing results of a tumor modeling solution of P004 resulting in a CNV-based purity estimation, according to an illustrative embodiment.

[0252] FIG. 51B is a graphs showing results of a tumor modeling solution of P004 resulting in a CNV-based purity estimation, according to an illustrative embodiment.

[0253] FIG. 51C is a graphs showing results of a tumor modeling solution of P004 resulting in a CNV-based purity estimation, according to an illustrative embodiment.

[0254] FIG. 51D is a graphs showing results of a tumor modeling solution of P004 resulting in a CNV-based purity estimation, according to an illustrative embodiment.

[0255] FIG. 51E is a graphs showing results of a tumor modeling solution of P004 resulting in a CNV-based purity estimation, according to an illustrative embodiment.

[0256] FIG. 51F is a graphs showing results of a tumor modeling solution of P004 resulting in a CNV-based purity estimation, according to an illustrative embodiment.

[0257] FIG. 51G is a graphs showing results of a tumor modeling solution of P004 resulting in a CNV-based purity estimation, according to an illustrative embodiment.

[0258] FIG. 51H is a screenshot of the genome view display for P004, according to an illustrative embodiment.

[0259] FIG. 52A is a graph showing results at a step of a tumor modeling solution of P003 resulting in resulting in a CNV-based upper bound on purity, according to an illustrative embodiment.-40 - 13247547v 1Attorney Docket No.: 2013237-0970

[0260] FIG. 52B is a graph showing results of a step of a tumor modeling solution of P003 resulting in a CNV-based purity estimation, according to an illustrative embodiment.

[0261] FIG. 52C is a graph showing results of a step of a tumor modeling solution of P003 resulting in a CNV-based purity estimation, according to an illustrative embodiment.

[0262] FIG. 52D is a set of graphs showing results of a step of a tumor modeling solution of P003 resulting in a CNV-based purity estimation, according to an illustrative embodiment.

[0263] FIG. 52E is a set of graphs showing results of a step of a tumor modeling solution of P003 resulting in a CNV-based purity estimation, according to an illustrative embodiment.

[0264] FIG. 53A is a graph showing results of evaluation of accuracy of purity and ploidy predictions of ovarian and glioblastoma cohorts, according to an illustrative embodiment.

[0265] FIG. 53B is a graph showing results of evaluation of accuracy of purity and ploidy predictions of ovarian and glioblastoma cohorts, according to an illustrative embodiment.

[0266] FIG. 53C is a graph showing results of evaluation of accuracy of purity and ploidy predictions of ovarian and glioblastoma cohorts, according to an illustrative embodiment.

[0267] FIG. 54 is a graph showing efficacy of error correction tested across 241 ovarian, glioblastoma, and breast cancer samples, according to an illustrative embodiment.

[0268] FIG. 55 is a graph showing accuracy of gender estimation tested against annotation from multiple cohorts across 344 patients, according to an illustrative embodiment.

[0269] FIG. 56A is a schematic illustrating how tumor control of personalized cancer immunotherapies can be enhanced by prioritizing mutations based on their genomic features and / or functionality of the genes encoding the mutations, according to certain embodiments.

[0270] FIG. 56B is a schematic illustrating how tumor control of personalized cancer immunotherapies can be enhanced by prioritizing mutations based on their genomic features and / or functionality of the genes encoding the mutations, according to certain embodiments.

[0271] FIG. 56C is a schematic illustrating how tumor control of personalized cancer immunotherapies can be enhanced by prioritizing mutations based on their genomic features and / or functionality of the genes encoding the mutations, according to certain embodiments.- 41 - 13247547v 1Attomey Docket No.: 2013237-0970

[0272] FIG. 56D is a schematic illustrating how tumor control of personalized cancer immunotherapies can be enhanced by prioritizing mutations based on their genomic features and / or functionality of the genes encoding the mutations, according to certain embodiments.

[0273] The features and advantages of the present disclosure will become more apparent from the detailed description set forth below when taken in conjunction with the drawings, in which like reference characters identify corresponding elements throughout. In the drawings, like reference numbers generally indicate identical, functionally similar, and / or structurally similar elements.CERTAIN DEFINITIONS

[0274] About-. The term “about”, when used herein in reference to a value, refers to a value that is similar, in context to the referenced value. In general, those skilled in the art, familiar with the context, will appreciate the relevant degree of variance encompassed by “about” in that context. For example, in some embodiments, the term “about” may encompass a range of values that within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or less of the referred value.

[0275] Absolute Copy Number-. The term “absolute copy number” as used herein refers to a number of physical copies of a particular segment within a cell comprising a particular genome. For example, an absolute copy number of a segment in the normal genome can be defined as the number of physical copies of the given segment in a healthy cell. For example, an absolute copy number of a segment in the tumor genome can be defined as the number of physical copies of the given segment in a tumor cell. In certain embodiments, if only a part of a segment is amplified or deleted in a genome, then such a partial copy of the segment can either be counted as a copy of the segment or not counted as a copy of the segment. In certain embodiments, copies of the segment spanning less than 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10% or 5% of the segment length can be ignored.

[0276] Agent-. As used herein, the term “agent,” may refer to a physical entity. In some embodiments, an agent may be characterized by a particular feature and / or effect. For example, as used herein, the term “therapeutic agent” refers to a physical entity has a therapeutic effect and / or elicits a desired biological and / or pharmacological effect. In some embodiments, an agent may be a compound, molecule, or entity of any chemical class including, for example, a small - 42 - 13247547v 1Attorney Docket No.: 2013237-0970molecule, polypeptide, nucleic acid, saccharide, lipid, metal, or a combination or complex thereof. In some embodiments, part or all of an agent may be depicted herein as a chemical structure, or may be described using chemical nomenclature and / or with reference to general principles of organic chemistry, e.g., in accordance with the Periodic Table of Elements, CAS version, Handbook of Chemistry and Physics, 75th Ed; “Organic Chemistry”, Thomas Sorrell, University Science Books, Sausalito: 1999, and / or “March’s Advanced Organic Chemistry”, 5th Ed., Ed.: Smith, M. B. and March, J., John Wiley & Sons, New York: 2001, the entire contents of which are hereby incorporated by reference. Unless otherwise stated or clear from context, chemical structures depicted herein may be considered to reference or include one or more, or all, stereoisomeric (e.g., enantiomeric or diastereomeric) forms of the structure, and / or one or more, or all, geometric or conformational isomeric forms of the structure. For example, unless otherwise indicated or clear, both R and S configurations of a stereocenter may be contemplated in embodiments of the disclosure. In some embodiments, a compound may be described and / or utilized as a particular single stereochemical isomer; alternatively or additionally, in some embodiments, such a compound may be described and / or utilized as a combination e.g., a mixture) of one or more enantiomeric (e.g., diastereomeric) forms (e.g., as a racemic preparation). Analogously, in some embodiments, a single geometric isomer may be described and / or utilized; in some embodiments, a combination (e.g., a mixture) of geometric (or conformational) isomers may be described and / or utilized. Unless otherwise stated or clear from context, all tautomeric forms of provided compounds are within the scope of the disclosure. Still further, unless otherwise indicated or clear from context, in some embodiments, a particular chemical compound (e.g., as may be represented by a depicted chemical structure) may be described and / or utilized in an alternative isotopic form - i.e., in a form in which one or more atoms is isotopically altered (e.g., so that a hydrogen is replaced by deuterium or tritium, and / or a carbon is replaced by 13C- or 14C-. Thus, in some embodiments, a particular compound may be described and / or utilized as or in an isotopically enriched preparation.

[0277] Allele-Specific Copy Number: As used herein, the term “allele- specific copy number” refers to a number of physical copies of a particular allele within a cell comprising a particular genome. In certain embodiments, for example, a heterozygous segment comprises a SNP. A cell comprising a genome with the heterozygous segment may comprise zero, one, or more copies of a first (e.g., maternal) allele and zero, one, or more copies of a second (e.g., - 43 - 13247547v 1Attomey Docket No.: 2013237-0970paternal) allele. For example, a balanced heterozygous segment in a normal diploid genome may comprise distinguishable maternal and paternal alleles, each having an allele-specific copy number of one. In certain embodiments, a corresponding heterozygous segment within a tumor genome may not have zero, one, or more copies of the paternal and maternal alleles, e.g., as a result of copy number variation (CNV) events that may occur in cancer cells. For example, a loss of heterozygosity (LOH) may result in a deletion of copies of one allele (e.g., a maternal or paternal allele), such that an allele- specific copy number for the deleted allele is zero and, if a single copy of the other allele remains, its allele- specific copy number is one. In certain embodiments, duplication events may produce other allele-specific copy numbers, greater than one for one or both alleles.

[0278] Amino acid'. In its broadest sense, as used herein, the term “amino acid” refers to a compound and / or substance that can be, is, or has been incorporated into a polypeptide chain, e.g., through formation of one or more peptide bonds. In some embodiments, an amino acid has the general structure H2N-C(H)(R)-COOH. In some embodiments, an amino acid is a naturally-occurring amino acid. In some embodiments, an amino acid is a non-natural amino acid; in some embodiments, an amino acid is a D-amino acid; in some embodiments, an amino acid is an L-amino acid. “Standard amino acid” refers to any of the twenty standard L-amino acids commonly found in naturally occurring peptides. “Nonstandard amino acid” refers to any amino acid, other than the standard amino acids, regardless of whether it is prepared synthetically or obtained from a natural source. In some embodiments, an amino acid, including a carboxy- and / or amino-terminal amino acid in a polypeptide, can contain a structural modification as compared with the general structure above. For example, in some embodiments, an amino acid may be modified by methylation, amidation, acetylation, pegylation, glycosylation, phosphorylation, and / or substitution e.g., of the amino group, the carboxylic acid group, one or more protons, and / or the hydroxyl group) as compared with the general structure. In some embodiments, such modification may, for example, alter the circulating half-life of a polypeptide containing the modified amino acid as compared with one containing an otherwise identical unmodified amino acid. In some embodiments, such modification does not significantly alter a relevant activity of a polypeptide containing the modified amino acid, as compared with one containing an otherwise identical unmodified amino acid. As will be clear- 44 - 13247547v 1Attorney Docket No.: 2013237-0970from context, in some embodiments, the term “amino acid” may be used to refer to a free amino acid; in some embodiments it may be used to refer to an amino acid residue of a polypeptide.

[0279] Antigen: term “antigen”, as used herein, refers to (i) an agent that elicits an immune response; and / or (ii) an agent that binds to a T cell receptor (e.g., when presented by an MHC molecule) or to an antibody. In some embodiments, an antigen elicits a humoral response (e.g., including production of antigen-specific antibodies); in some embodiments, an elicits a cellular response e.g., involving T-cells whose receptors specifically interact with the antigen). In some embodiments, and antigen binds to an antibody and may or may not induce a particular physiological response in an organism. In general, an antigen may be or include any chemical entity such as, for example, a small molecule, a nucleic acid, a polypeptide, a carbohydrate, a lipid, a polymer [in some embodiments other than a biologic polymer (e.g., other than a nucleic acid or amino acid polymer)] etc. In some embodiments, an antigen is or comprises a polypeptide. In some embodiments, an antigen is or comprises a glycan. Those of ordinary skill in the art will appreciate that, in general, an antigen may be provided in isolated or pure form, or alternatively may be provided in crude form (e.g., together with other materials, for example in an extract such as a cellular extract or other relatively crude preparation of an antigen-containing source). In some embodiments, antigens utilized in accordance with the present invention are provided in a crude form. In some embodiments, an antigen is a recombinant antigen.

[0280] Biologically clonal mutation(s) and biologically subclonal mutation(s): As used herein, the terms “biologically clonal” and “biologically subclonal” when used in reference to mutations, such as cancer mutations, are used to specify whether a particular mutations or group of mutations are physically clonal or subclonal. In certain embodiments, a biologically clonal mutation is a mutation that is present in all tumor cells of a tumor sample or biopsy. In certain embodiments, a biologically subclonal mutation is a mutation that is not present in all tumor cells of a tumor sample or biopsy. The use of the adjective “biologically” is used to make clear that the terms “biologically clonal” and “biologically subclonal” refer to the actual physical character of a given mutation, which may or may not be known. The terms “biologically clonal” and “biologically subclonal” thus contrast with the terms “clone type”, “clonal state”, “prevalent subclone state”, and “minor subclone state”, described below, which refer to clonality classification states that are, e.g., labels, determined for (e.g., assigned to) a given mutation (see,- 45 - 13247547v 1Attomey Docket No.: 2013237-0970e.g., US Provisional Appln. No. 63 / 749,339 for methods and systems for determining clone type).

[0281] Cancer. The term “cancer” is used herein to generally refer to a disease or condition in which cells of a tissue of interest exhibit relatively abnormal, uncontrolled, and / or autonomous growth, so that they exhibit an aberrant growth phenotype characterized by a significant loss of control of cell proliferation. In some embodiments, cancer may comprise cells that are precancerous e.g., benign), malignant, pre-metastatic, metastatic, and / or non-metastatic. In some embodiments, cancer may be characterized by a solid tumor. In some embodiments, cancer may be characterized by a hematologic tumor. In general, examples of different types of cancers known in the art include, for example, triple negative breast cancer (TNBC), hematopoietic cancers including leukemias, lymphomas (Hodgkin’s and non-Hodgkin’s), myelomas and myeloproliferative disorders; sarcomas, melanomas, adenomas, carcinomas of solid tissue, squamous cell carcinomas of the mouth, throat, larynx, and lung, liver cancer, genitourinary cancers such as prostate, cervical, bladder, uterine, and endometrial cancer and renal cell carcinomas, bone cancer, pancreatic cancer, skin cancer, cutaneous or intraocular melanoma, cancer of the endocrine system, cancer of the thyroid gland, cancer of the parathyroid gland, head and neck cancers, ovarian cancer, breast cancer, glioblastomas, colorectal cancer, gastro-intestinal cancers and nervous system cancers, benign lesions such as papillomas, and the like.

[0282] Clone type: As used herein the term “clone type” is used to refer to certain discrete states or classes that may be determined for and assigned to a given mutation, for example, based on evaluation of various criteria. Different clone types may aim to capture, for example, a likelihood that a given mutation is biologically clonal or biologically subclonal, as well as, for example, a prevalence of the given mutation. In certain embodiments, a likelihood that a given mutation is biologically clonal may be assessed based on various metrics or tests, including, without limitation, statistical hypothesis tests, computed probabilities (e.g., of clonality or sub-clonality, such as a posteriori probability of clonality), etc. For example, in certain embodiments, a likelihood of whether a given mutation is biologically clonal may be assessed using metrics, such as P-values, determined using hypotheses tests, that quantify probabilities of whether for a given data (such as sequencing data) a hypothesis of clonality can- 46 - 13247547v 1Attorney Docket No.: 2013237-0970be rejected. For example, in certain embodiments, sequencing data may be used to determine an observed cellularity of a given mutation (e.g., a fraction of tumor cells harboring the given mutation). Although, nominally, biologically clonal mutations have a cellularity of 1 (i.e., they are present in all tumor cells), factors, such as measurement error, sample purity, stochastic noise, etc., may lead to observed cellularity values below one, even for biologically clonal mutations. Accordingly, in certain embodiments, a P- value representing a probability of observing a cellularity below one for the given mutation (e.g., given a null hypothesis that the mutation is biologically clonal) may be determined. In this way, for example, a low P-value, e.g., below a threshold, may indicate that it is unlikely that the null hypothesis (of clonality) is true, and it can be rejected (hence unlikely that a mutation is clonal), whereas a higher P-value may indicate that there is a good chance that the observed cellularity, below 1, is due to practical factors, such as measurement error, sample purity, stochastic noise, and the like. In certain embodiments, prevalence may be measured by parameters, such as cellularity, e.g., a determined fraction of tumor cells that harbor a given mutation. Whereas the terms biologically clonal and biologically subclonal are used to refer to underlying physical characteristics of a given mutation, clone types are labels, which aim to classify mutations according to measured physical properties, determined, e.g., based on sequencing data. In certain embodiments, clone types include clonal states, which aim to capture or label mutations that are highly likely to be biologically clonal and / or highly prevalent. In certain embodiments, clone types include multiple clonal classification states, for example, reflecting differing levels of certainty and / or likelihoods that a given mutation is biologically clonal and / or prevalence. In certain embodiments, clone types include one or more prevalent subclone states, capturing mutations that, while biologically subclonal, are prevalent at high rates (e.g., above 50%) in tumor cells. In certain embodiments, clone types include a minor subclonal state, which may be assigned to mutations that are determined to be likely subclonal and rare e.g., occurring in 50% or less, 45% or less, 40% or less, 35% or less, or 25% or less cancer cells).

[0283] Comparable'. As used herein, the term “comparable” refers to two or more agents, entities, situations, sets of conditions, etc., that may not be identical to one another but that are sufficiently similar to permit comparison therebetween so that one skilled in the art will appreciate that conclusions may reasonably be drawn based on differences or similarities observed. In some embodiments, comparable sets of conditions, circumstances, individuals, or - 47 - 13247547v 1Attorney Docket No.: 2013237-0970populations are characterized by a plurality of substantially identical features and one or a small number of varied features. Those of ordinary skill in the art will understand, in context, what degree of identity is required in any given circumstance for two or more such agents, entities, situations, sets of conditions, etc., to be considered comparable. For example, those of ordinary skill in the art will appreciate that sets of circumstances, individuals, or populations are comparable to one another when characterized by a sufficient number and type of substantially identical features to warrant a reasonable conclusion that differences in results obtained or phenomena observed under or with different sets of circumstances, individuals, or populations are caused by or indicative of the variation in those features that are varied.

[0284] Corresponding to: As used herein, the term “corresponding to” refers to a relationship between two or more entities. For example, the term “corresponding to” may be used to designate the position / identity of a structural element in a compound or composition relative to another compound or composition (e.g., to an appropriate reference compound or composition). For example, in some embodiments, a monomeric residue in a polymer (e.g., an amino acid residue in a polypeptide or a nucleic acid residue in a polynucleotide) may be identified as “corresponding to” a residue in an appropriate reference polymer. For example, those of ordinary skill will appreciate that, for purposes of simplicity, residues in a polypeptide are often designated using a canonical numbering system based on a reference related polypeptide, so that an amino acid “corresponding to” a residue at position 190, for example, need not actually be the 190th amino acid in a particular amino acid chain but rather corresponds to the residue found at 190 in the reference polypeptide; those of ordinary skill in the art readily appreciate how to identify “corresponding” amino acids. For example, those skilled in the art will be aware of various sequence alignment strategies, including software programs such as, for example, BLAST, CS-BLAST, CUSASW++, DIAMOND, FASTA, GGSEARCH / GLSEARCH, Genoogle, HMMER, HHpred / HHsearch, IDF, Infernal, KLAST, USEARCH, parasail, PSL BLAST, PSI-Search, ScalaBLAST, Sequilab, SAM, SSEARCH, SWAPHI, SWAPHLLS, SWIMM, or SWIPE that can be utilized, for example, to identify “corresponding” residues in polypeptides and / or nucleic acids in accordance with the present disclosure. Those of skill in the art will also appreciate that, in some instances, the term “corresponding to” may be used to describe an event or entity that shares a relevant similarity with another event or entity (e.g., an appropriate reference event or entity). To give but one example, a gene or protein in one - 48 - 13247547v 1Attomey Docket No.: 2013237-0970organism may be described as “corresponding to” a gene or protein from another organism in order to indicate, in some embodiments, that it plays an analogous role or performs an analogous function and / or that it shows a particular degree of sequence identity or homology, or shares a particular characteristic sequence element.

[0285] Encode-. As used herein, the term “encode” or “encoding” refers to sequence information of a first molecule that guides production of a second molecule having a defined sequence of nucleotides (e.g., a polyribonucleotide) or a defined sequence of amino acids. For example, a DNA molecule can encode an RNA molecule (e.g., by a transcription process that includes a DNA-dependent RNA polymerase enzyme). An RNA molecule can encode a polypeptide e.g., by a translation process). Thus, a gene, a cDNA, or an RNA molecule encodes a polypeptide if transcription and translation of RNA corresponding to that gene produces the polypeptide in a cell or other biological system. In some embodiments, a coding region of a polyribonucleotide encoding a target antigen refers to a coding strand, the nucleotide sequence of which is identical to the polyribonucleotide sequence of such a target antigen. In some embodiments, a coding region of a polyribonucleotide encoding a target antigen refers to a noncoding strand of such a target antigen, which may be used as a template for transcription of a gene or cDNA.

[0286] Epitope-. As used herein, the term “epitope” refers to a moiety that is specifically recognized by an immune system (e.g., an immune system component) of a subject. For example, in some embodiments, an epitope may be a moiety that is specifically recognized by a T cell, a B cell, an immunoglobulin (e.g., antibody or receptor), immunoglobulin (e.g., antibody or receptor), binding component or an aptamer. In some embodiments, an epitope is comprised of a plurality of chemical atoms or groups on an antigen. In some embodiments, such chemical atoms or groups are surface-exposed when the antigen adopts a relevant three-dimensional conformation. In some embodiments, such chemical atoms or groups are physically near to each other in space when the antigen adopts such a conformation. In some embodiments, at least some such chemical atoms are groups are physically separated from one another when the antigen adopts an alternative conformation (e.g., is linearized).

[0287] Estimated tumor sample purity. As used herein, the term “estimated tumor sample purity” is used to refer to an estimate of tumor content of a tumor sample, such as a - 49 - 13247547v 1Attomey Docket No.: 2013237-0970fraction, percentage etc. of cancer cells within a tumor sample and / or, equivalently, an estimate of normal contamination, such as a fraction, percentage, etc. of normal (e.g., healthy) cells within a tumor sample. It should be understood that, in certain embodiments, tumor samples are assumed to be comprised of tumor cells and normal cells, such that a fraction of tumor cells in a tumor sample is equal to 1 minus a fraction of normal cells (e.g., 1 - p), estimates of tumor content and / or normal contamination equivalently measure an estimated tumor sample purity.

[0288] Expression-. As used herein, the term “expression” of a nucleic acid sequence refers to the generation of a gene product from the nucleic acid sequence. In some embodiments, a gene product can be a transcript, e.g., a polyribonucleotide as provided herein. In some embodiments, a gene product can be a polypeptide. In some embodiments, expression of a nucleic acid sequence involves one or more of the following: (1) production of an RNA template from a DNA sequence (e.g., by transcription); (2) processing of an RNA transcript (e.g., by splicing, editing, etc.); (3) translation of an RNA into a polypeptide or protein; and / or (4) post-translational modification of a polypeptide or protein.

[0289] Heterozygous Segment-. The term “heterozygous segment” as used here refers to a segment of a normal genome that comprises at least one heterozygous SNP, as well as any corresponding segments of a tumor genome or reference genome. In other words, a particular segment of a particular genome is defined as heterozygous or not according to whether the corresponding segment of a normal genome comprises a heterozygous SNP or not. For example, a reference genome may be partitioned into a plurality of segments, as described herein, in order to identify and define corresponding segments in a normal and tumor genome. Accordingly, if, for a given segment, the corresponding segment in the normal genome is determined to comprise a heterozygous SNP, then that segment is defined as a heterozygous segment. For purposes of determining heterozygous segments, a normal genome may be a reference genome obtained from a database, a normal reference determined and / or compiled based on one or more subject (e.g., a panel), determined by sequencing a particular subject (e.g., the same subject whose tumor is being sequenced).

[0290] Homology. As used herein, the term “homology” or “homolog” refers to the overall relatedness between polynucleotide molecules (e.g., DNA molecules and / or RNA molecules) and / or between polypeptide molecules. In some embodiments, polynucleotide - 50 - 13247547v 1Attorney Docket No.: 2013237-0970molecules (e.g., DNA molecules and / or RNA molecules) and / or polypeptide molecules are considered to be “homologous” to one another if their sequences are at least 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical. In some embodiments, polynucleotide molecules (e.g., DNA molecules and / or RNA molecules) and / or polypeptide molecules are considered to be “homologous” to one another if their sequences are at least 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% similar (e.g., containing residues with related chemical properties at corresponding positions). For example, as is well known by those of ordinary skill in the art, certain amino acids are typically classified as similar to one another as “hydrophobic” or “hydrophilic” amino acids, and / or as having “polar” or “non-polar” side chains. Substitution of one amino acid for another of the same type may often be considered a “homologous” substitution.

[0291] Identity. As used herein, the term “identity” refers to the overall relatedness between polynucleotide molecules (e.g., DNA molecules and / or RNA molecules) and / or between polypeptide molecules. In some embodiments, polynucleotide molecules (e.g., DNA molecules and / or RNA molecules) and / or between polypeptide molecules are considered to be “substantially identical” to one another if their sequences are at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical. Calculation of the percent identity of two nucleic acid or polypeptide sequences, for example, can be performed by aligning the two sequences for optimal comparison purposes (e.g., gaps can be introduced in one or both of a first and a second sequence for optimal alignment and non-identical sequences can be disregarded for comparison purposes). In certain embodiments, the length of a sequence aligned for comparison purposes is at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or substantially 100% of the length of a reference sequence. The nucleotides at corresponding positions are then compared. When a position in the first sequence is occupied by the same residue (e.g., nucleotide or amino acid) as the corresponding position in the second sequence, then the molecules are identical at that position. The percent identity between the two sequences is a function of the number of identical positions shared by the sequences, taking into account the number of gaps, and the length of each gap, which needs to be introduced for optimal alignment of the two sequences. The comparison of sequences and determination of percent identity - 51 - 13247547v 1Attomey Docket No.: 2013237-0970between two sequences can be accomplished using a mathematical algorithm. For example, the percent identity between two nucleotide sequences can be determined using the algorithm of Meyers and Miller, 1989, which has been incorporated into the ALIGN program (version 2.0). In some exemplary embodiments, nucleic acid sequence comparisons made with the ALIGN program use a PAM 120 weight residue table, a gap length penalty of 12 and a gap penalty of 4. The percent identity between two nucleotide sequences can, alternatively, be determined using the GAP program in the GCG software package using an NWSgapdna. CMP matrix.

[0292] Increased, Induced, or Reduced'. As used herein, these terms or grammatically comparable comparative terms, indicate values that are relative to a comparable reference measurement. For example, in some embodiments, an assessed value achieved with a provided composition (e.g., a pharmaceutical composition) may be “increased” relative to that obtained with a comparable reference composition. Alternatively or additionally, in some embodiments, an assessed value achieved in a subject may be “increased” relative to that obtained in the same subject under different conditions (e.g., prior to or after an event; or presence or absence of an event such as administration of a composition (e.g., a pharmaceutical composition) as described herein, or in a different, comparable subject (e.g., in a comparable subject that differs from the subject of interest in prior exposure to a condition, e.g., absence of administration of a composition (e.g., a pharmaceutical composition) as described herein.). In some embodiments, comparative terms refer to statistically relevant differences (e.g., that are of a prevalence and / or magnitude sufficient to achieve statistical relevance). Those skilled in the art will be aware, or will readily be able to determine, in a given context, a degree and / or prevalence of difference that is required or sufficient to achieve such statistical significance. In some embodiments, the term “reduced” or equivalent terms refers to a reduction in the level of an assessed value by at least 5%, at least 10%, at least 20%, at least 50%, at least 75% or higher, as compared to a comparable reference. In some embodiments, the term “reduced” or equivalent terms refers to a complete or essentially complete inhibition, i.e., a reduction to zero or essentially to zero. In some embodiments, the term “increased” or “induced” refers to an increase in the level of an assessed value by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 80%, at least 100%, at least 200%, at least 500%, or higher, as compared to a comparable reference.- 52 - 13247547v 1Attorney Docket No.: 2013237-0970

[0293] In order. As used herein with reference to a polynucleotide or polyribonucleotide, “in order” refers to the order of features from 5' to 3' along the polynucleotide or polyribonucleotide. As used herein with reference to a polypeptide, “in order” refers to the order of features moving from the N-terminal-most of the features to the C-terminal-most of the features along the polypeptide. “In order” does not mean that no additional features can be present among the listed features. For example, if Features A, B, and C of a polynucleotide are described herein as being “in order, Feature A, Feature B, and Feature C,” this description does not exclude, e.g., Feature D being located between Features A and B.

[0294] Linker. As used herein, the term “linker” refers to a portion of a polypeptide that connects different regions, portions, or antigens to one another.

[0295] Lipid'. As used herein, the terms “lipid” and “lipid-like material” are broadly defined as molecules which comprise one or more hydrophobic moieties or groups and optionally also one or more hydrophilic moieties or groups. Molecules comprising hydrophobic moieties and hydrophilic moieties are also typically denoted as amphiphiles.

[0296] Neoantigen'. As used herein, the term “neoantigen” refers to an antigen that is not present in a reference, such as a normal non-cancerous or germline cell, but is present in a cancer cell. In some embodiments, a neoantigen includes one or more mutations relative to a corresponding antigen present in a normal non-cancerous or germline cell.

[0297] Neoantigen epitope. As used herein, the term “neoantigen epitope” refers to an epitope that is not present in a reference, such as a normal non-cancerous or germline cell, but is present in a cancer cell.

[0298] Nucleic acid / Polynucleotide'. As used herein, the term “nucleic acid” refers to a polymer of at least 10 nucleotides or more. In some embodiments, a nucleic acid is or comprises DNA. In some embodiments, a nucleic acid is or comprises RNA. In some embodiments, a nucleic acid is or comprises peptide nucleic acid (PNA). In some embodiments, a nucleic acid is or comprises a single stranded nucleic acid. In some embodiments, a nucleic acid is or comprises a double- stranded nucleic acid. In some embodiments, a nucleic acid comprises both single and double- stranded portions. In some embodiments, a nucleic acid comprises a backbone that comprises one or more phosphodiester linkages. In some embodiments, a nucleic acid comprises a backbone that comprises both phosphodiester and non-phosphodiester linkages. For - 53 - 13247547v 1Attomey Docket No.: 2013237-0970example, in some embodiments, a nucleic acid may comprise a backbone that comprises one or more phosphorothioate or 5'-N-phosphoramidite linkages and / or one or more peptide bonds, e.g., as in a “peptide nucleic acid”. In some embodiments, a nucleic acid comprises one or more, or all, natural residues (e.g., adenine, cytosine, deoxy adenosine, deoxycytidine, deoxy guanosine, deoxythymidine, guanine, thymine, uracil). In some embodiments, a nucleic acid comprises on or more, or all, non-natural residues. In some embodiments, a non-natural residue comprises a nucleoside analog (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3 -methyl adenosine, 5-methylcytidine, C-5 propynyl-cytidine, C-5 propynyl-uridine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyl-uridine, C5 -propynyl-cytidine, C5-methylcytidine, 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, 6-O-methylguanine, 2-thiocytidine, methylated bases, intercalated bases, and combinations thereof). In some embodiments, a non-natural residue comprises one or more modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose) as compared to those in natural residues. In some embodiments, a nucleic acid has a nucleotide sequence that encodes a functional gene product such as an RNA or polypeptide. In some embodiments, a nucleic acid has a nucleotide sequence that comprises one or more introns. In some embodiments, a nucleic acid may be prepared by isolation from a natural source, enzymatic synthesis (e.g., by polymerization based on a complementary template, e.g., in vivo or in vitro), reproduction in a recombinant cell or system, or chemical synthesis. In some embodiments, a nucleic acid is at least 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 1 10, 120, 130, 140, 150, 160, 170, 180, 190, 20, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500, 10,000, 10,500, 11,000, 11,500, 12,000, 12,500, 13,000, 13,500, 14,000, 14,500, 15,000, 15,500, 16,000, 16,500, 17,000, 17,500, 18,000, 18,500, 19,000, 19,500, or 20,000 or more residues or nucleotides long.

[0299] Ploidy. As used herein, the term “ploidy,” for example of a tumor genome, is used to refer to an average of absolute copy numbers of all segments (e.g., across an entire region of a tumor genome), weighted by the length of each segment. A ploidy of a region of a tumor genome can be defined as the average of absolute copy numbers of all segments in the region, weighted by the length of each segment.- 54 - 13247547v 1Attorney Docket No.: 2013237-0970

[0300] Polypeptide'. As used herein, the term “polypeptide” refers to a polymeric chain of amino acids. In some embodiments, a polypeptide has an amino acid sequence that occurs in nature. In some embodiments, a polypeptide has an amino acid sequence that does not occur in nature. In some embodiments, a polypeptide has an amino acid sequence that is engineered in that it is designed and / or produced through action of the hand of man. In some embodiments, a polypeptide may comprise or consist of natural amino acids, non-natural amino acids, or both. In some embodiments, a polypeptide may comprise or consist of only natural amino acids or only non-natural amino acids. In some embodiments, a polypeptide may comprise D-amino acids, L-amino acids, or both. In some embodiments, a polypeptide may comprise only D-amino acids. In some embodiments, a polypeptide may comprise only L- amino acids. In some embodiments, a polypeptide may include one or more pendant groups or other modifications, e.g., modifying or attached to one or more amino acid side chains, at the polypeptide’s N-terminus, at the polypeptide’s C-terminus, or any combination thereof. In some embodiments, such pendant groups or modifications comprise acetylation, amidation, lipidation, methylation, pegylation, etc., including combinations thereof. In some embodiments, a polypeptide may be cyclic, and / or may comprise a cyclic portion. In some embodiments, a polypeptide is not cyclic and / or does not comprise any cyclic portion. In some embodiments, a polypeptide is linear. In some embodiments, a polypeptide may be or comprise a stapled polypeptide. In some embodiments, the term “polypeptide” may be appended to a name of a reference polypeptide, activity, or structure; in such instances it is used herein to refer to polypeptides that share the relevant activity or structure and thus can be considered to be members of the same class or family of polypeptides. For each such class, the present specification provides and / or those skilled in the art will be aware of exemplary polypeptides within the class whose amino acid sequences and / or functions are known; in some embodiments, such exemplary polypeptides are reference polypeptides for the polypeptide class or family. In some embodiments, a member of a polypeptide class or family shows significant sequence homology or identity with, shares a common sequence motif (e.g., a characteristic sequence element) with, and / or shares a common activity (in some embodiments at a comparable level or within a designated range) with a reference polypeptide of the class; in some embodiments with all polypeptides within the class). For example, in some embodiments, a member polypeptide shows an overall degree of sequence homology or identity with a reference polypeptide that is at least about 30-40%, and is often - 55 - 13247547v 1Attomey Docket No.: 2013237-0970greater than about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more and / or includes at least one region (e.g., a conserved region that may in some embodiments be or comprise a characteristic sequence element) that shows very high sequence identity, often greater than 90% or even 95%, 96%, 97%, 98%, or 99%. Such a conserved region usually encompasses at least 3-4 and often up to 35 or more amino acids; in some embodiments, a conserved region encompasses at least one stretch of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35 or more contiguous amino acids. In some embodiments, a relevant polypeptide may comprise or consist of a fragment of a parent polypeptide.

[0301] Read count'. As used herein, the term “read count” refers to a number of sequencing data reads that map to a particular segment, portion thereof, or individual location (such as a SNP) within a genome. For example, the phrases “read count of a particular segment” and “segment read count” as used herein refer to a number of reads that map to the particular segment. For example, the phrases “read count of a particular SNP” and “SNP read count” as used herein refer to a number of reads that map to the particular SNP. The term “read count” may be preceded by an indication of a particular set of sequencing data and / or sequenced sample. For example, when a tumor sample is sequenced to produce tumor sequencing data comprising a plurality of tumor sequencing reads, the phrase “tumor read count” is used to refer to the number of tumor sequencing reads that map to a particular segment, portion thereof, or individual location (e.g., SNP) within a genome. Likewise, when a normal sample is sequenced to produce normal sequencing data comprising a plurality of normal sequencing reads, the phrase “normal read count” is used to refer to the number of normal sequencing reads that map to a particular segment, portion thereof, or individual location (e.g., SNP) within a genome.

[0302] Reference'. As used herein, the term “reference” describes a standard or control relative to which a comparison is performed. For example, in some embodiments, an agent, animal, individual, population, sample, sequence or value of interest is compared with a reference or control agent, animal, individual, population, sample, sequence or value. In some embodiments, a reference or control is tested and / or determined substantially simultaneously with the testing or determination of interest. In some embodiments, a reference or control is a historical reference or control, optionally embodied in a tangible medium. Typically, as would- 56 - 13247547v 1Attorney Docket No.: 2013237-0970be understood by those skilled in the art, a reference or control is determined or characterized under comparable conditions or circumstances to those under assessment. Those skilled in the art will appreciate when sufficient similarities are present to justify reliance on and / or comparison to a particular possible reference or control.

[0303] Ribonucleic acid (RNA) or Polyribonucleotide'. As used herein, the term “ribonucleic acid,” “RNA,” or “polyribonucleotide” refers to a polymer of ribonucleotides. In some embodiments, an RNA is single stranded. In some embodiments, an RNA is double stranded. In some embodiments, an RNA comprises both single and double stranded portions. In some embodiments, an RNA can comprise a backbone structure as described in the definition of “Nucleic acid / Polynucleotide” above. An RNA can be a regulatory RNA (e.g., siRNA, microRNA, etc.), or a messenger RNA (mRNA). In some embodiments, an RNA is a mRNA. In some embodiments, where an RNA is a mRNA, an RNA typically comprises at its 3' end a poly(A) region. In some embodiments, where an RNA is a mRNA, an RNA typically comprises at its 5' end an art-recognized cap structure, e.g., for recognizing and attachment of a mRNA to a ribosome to initiate translation. In some embodiments, an RNA is a synthetic RNA. Synthetic RNAs include RNAs that are synthesized in vitro (e.g., by enzymatic synthesis methods and / or by chemical synthesis methods).

[0304] Ribonucleotide'. As used herein, the term “ribonucleotide” encompasses unmodified ribonucleotides and modified ribonucleotides. For example, unmodified ribonucleotides include the purine bases adenine (A) and guanine (G), and the pyrimidine bases cytosine (C) and uracil (U). Modified ribonucleotides may include one or more modifications including, but not limited to, for example, (a) end modifications, e.g., 5' end modifications (e.g., phosphorylation, dephosphorylation, conjugation, inverted linkages, etc.), 3' end modifications (e.g., conjugation, inverted linkages, etc.), (b) base modifications, e.g., replacement with modified bases, stabilizing bases, destabilizing bases, or bases that base pair with an expanded repertoire of partners, or conjugated bases, (c) sugar modifications (e.g., at the 2' position or 4' position) or replacement of the sugar, and (d) internucleoside linkage modifications, including modification or replacement of the phosphodiester linkages. The term “ribonucleotide” also encompasses ribonucleotide triphosphates including modified and non-modified ribonucleotide triphosphates.- 57 - 13247547v 1Attorney Docket No.: 2013237-0970

[0305] Secretory signal: As used herein, the term “secretory signal” refers to an amino acid sequence motif that targets associated polypeptides for translocation to a secretory pathway.

[0306] Segment, Segments: As used herein, the terms “segment” or “segments” (e.g., when used in regard to genetic material) refer to specific pre-defined regions of one or more genomes. For example, a particular reference genome may be subdivided into a plurality of segments, each a specific subsequence of consecutive nucleotides in the reference genome. In certain embodiments, a particular genome may be subdivided into its constituent genes, such that each segment corresponds to a particular, different, gene of the particular genome. In certain embodiments, exons of a particular genome are identified and retained, such that each segment corresponds to a particular, different, exon. In certain embodiments, each segment corresponds to a locus (e.g., a particular location on a chromosome where a particular gene, genetic marker, or allele is located). In certain embodiments, a particular genome may be subdivided into segments of a same or substantially same size (e.g., number of bases). As will be understood by one of skill in the art, a set sequencing data obtained from a particular sample, such as reads obtained via next generation sequencing (NGS) data obtained by sequencing a particular sample, may be aligned to a reference genome. In this way, reference genome may be subdivided into a plurality of segments and corresponding segments identified within a genome characteristic of the particular sample. Reads from the set of sequencing data can, accordingly, be identified as mapping to various particular segments within the genome characteristic of the particular sample and used to characterize them. In certain embodiments, multiple sets of sequencing data may be obtained for different samples (e.g., tumor sequencing data from a tumor sample, normal sequencing data from a normal sample) and aligned to a common reference genome. In this way, corresponding segments that comprises the same or substantially same (e.g., all save for variations due to e.g., single nucleotide polymorphisms (SNPs), single nucleotide variations (SNVs), insertions, deletions, etc.) base positions as from genomes characteristic of different samples can be identified. That is, given a particular segment from one genome, associated with one sample, a corresponding segment of another genome, associated with another sample, may be identified. Corresponding segments may have a same and / or substantially same length (e.g., accounting for insertions, deletions, etc.). A particular segment is referred to herein as encoding or comprising a particular SNP and / or SNV if that particular SNP and / or SNV is within the particular segment.- 58 - 13247547v 1Attomey Docket No.: 2013237-0970

[0307] Single Nucleotide Polymorphism (SNP): As used herein, the term “single nucleotide polymorphism” or “SNP” refers to a particular site (e.g., base position) in a genome where alternative bases are known and / or determined to distinguish one allele from another.

[0308] Single Nucleotide Variation (SNV): As used herein, the term “single nucleotide variation” is used to refer to a difference in the nucleic acid sequence (substitution of one base for another) at a particular site (allele) when comparing a genome from a diseased cell, such as a tumor cell, and a genome of a normal, non-diseased cell or a reference genome. In some embodiments, detecting mutations may refer to detecting nucleotide substitution mutations. In certain embodiments, a SNV is a somatic point mutation that occurs only in diseased (e.g., cancer) cells.

[0309] Subject-. As used herein, the term “subject” refers to an organism to be administered with a composition described herein, e.g., for experimental, diagnostic, prophylactic, and / or therapeutic purposes. Typical subjects include animals (e.g., mammals such as mice, rats, rabbits, non-human primates, domestic pets, etc.) and humans. In some embodiments, a subject is a human subject. In some embodiments, a subject is suffering from a disease, disorder, or condition (e.g., cancer and / or a cancer-associated condition). In some embodiments, a subject is susceptible to a disease, disorder, or condition (e.g., cancer and / or a cancer-associated condition). In some embodiments, a subject displays one or more symptoms or characteristics of a disease, disorder, or condition (e.g., cancer and / or a cancer-associated condition). In some embodiments, a subject displays one or more non-specific symptoms of a disease, disorder, or condition (e.g., cancer and / or a cancer-associated condition). In some embodiments, a subject does not display any symptom or characteristic of a disease, disorder, or condition (e.g., cancer and / or a cancer-associated condition). In some embodiments, a subject is someone with one or more features characteristic of susceptibility to or risk of a disease, disorder, or condition (e.g., cancer and / or a cancer-associated condition). In some embodiments, a subject is a patient. In some embodiments, a subject is an individual to whom diagnosis and / or therapy is and / or has been administered.

[0310] Therapy. The term “therapy” refers to an administration or delivery of an agent or intervention that has a therapeutic effect and / or elicits a desired biological and / or pharmacological effect (e.g., has been demonstrated to be statistically likely to have such effect - 59 - 13247547v 1Attomey Docket No.: 2013237-0970when administered to a relevant population). In some embodiments, a therapeutic agent or therapy is any substance that can be used to alleviate, ameliorate, relieve, inhibit, prevent, delay onset of, reduce severity of, and / or reduce incidence of one or more symptoms or features of a disease, disorder, and / or condition (e.g., cancer and / or a cancer-associated condition). In some embodiments, a therapeutic agent or therapy is a medical intervention that can be performed to alleviate, relieve, inhibit, present, delay onset of, reduce severity of, and / or reduce incidence of one or more symptoms or features of a disease, disorder, and / or condition.

[0311] Treat'. As used herein, the term “treat,” “treatment,” or “treating” refers to any method used to partially or completely alleviate, ameliorate, relieve, inhibit, prevent, delay onset of, reduce severity of, and / or reduce incidence of one or more symptoms or features of a disease, disorder, and / or condition (e.g., cancer and / or a cancer-associated condition). Treatment may be administered to a subject who does not exhibit signs of a disease, disorder, and / or condition (e.g., cancer and / or a cancer-associated condition). In some embodiments, treatment may be administered to a subject who exhibits only early signs of the disease, disorder, and / or condition (e.g., cancer and / or a cancer-associated condition), for example for the purpose of decreasing the risk of developing pathology associated with the disease, disorder, and / or condition. In some embodiments, treatment may be administered to a subject at a later-stage of disease, disorder, and / or condition (e.g., cancer and / or a cancer-associated condition).

[0312] Wild-Type'. As used herein, the term “wild-type” refers to an entity having a structure and / or activity as found in nature in a “normal” (as contrasted with mutant, diseased, altered, etc.) state or context. Those of ordinary skill in the art will appreciate that wild-type genes and polypeptides often exist in multiple different forms (e.g., alleles). For example, for a subject with cancer, wild-type genes, segments, SNPs, or properties thereof may be those present in a genome of that subject’s normal, non-cancerous, cells, as opposed to altered versions of those genes, segments, SNPs, or properties thereof that appear in genomes of cancer cells within the subject.DETAILED DESCRIPTION

[0313] It is contemplated that systems, architectures, devices, methods, and processes of the claimed invention encompass variations and adaptations developed using information from - 60 - 13247547v 1Attorney Docket No.: 2013237-0970the embodiments described herein. Adaptation and / or modification of the systems, architectures, devices, methods, and processes described herein may be performed, as contemplated by this description.

[0314] Throughout the description, where articles, devices, systems, and architectures are described as having, including, or comprising specific components, or where processes and methods are described as having, including, or comprising specific steps, it is contemplated that, additionally, there are articles, devices, systems, and architectures of the present invention that consist essentially of, or consist of, the recited components, and that there are processes and methods according to the present invention that consist essentially of, or consist of, the recited processing steps.

[0315] It should be understood that the order of steps or order for performing certain action is immaterial so long as the invention remains operable. Moreover, two or more steps or actions may be conducted simultaneously.

[0316] The mention herein of any publication, for example, in the Background section, is not an admission that the publication serves as prior art with respect to any of the claims presented herein. The Background section is presented for purposes of clarity and is not meant as a description of prior art with respect to any claim.

[0317] Documents are incorporated herein by reference as noted. Where there is any discrepancy in the meaning of a particular term, the meaning provided in the Definition section above is controlling.

[0318] Headers are provided for the convenience of the reader - the presence and / or placement of a header is not intended to limit the scope of the subject matter described herein.

[0319] Presented herein are methods and systems that allow biophysical and / or genomic properties of tumor samples to be determined via analysis of sequencing data, obtained, for example, via next-generation sequencing of tumor samples obtained from a subject. Among other things, tumor deconvolution technologies of the present disclosure leverage signals in sequencing data that originate from copy number variation (CNV) events that occur in cancer cells to estimate tumor sample purity and / or determine genomic properties, such as absolute copy numbers of segments within a tumor genome. Among other things, accurate assessments of- 61 - 13247547v 1Attomey Docket No.: 2013237-0970tumor sample purity and / or genomic properties made possible via techniques described herein can be utilized to accurately detect and characterize cancer- specific mutations, such as single nucleotide variations (SNVs) that may serve as targets for personalized cancer therapies.

[0320] In particular, in certain embodiments, tumor deconvolution technologies of the present disclosure utilize a particular subset of balanced heterozygous segments, referred to herein as primary balanced heterozygous segments. As described in further detail herein, primary balanced heterozygous segments may be identified as balanced segments - e.g., segments that are first determined to have equal numbers of copies of material and paternal alleles within a tumor genome - and have a same absolute copy number. The absolute copy number of primary balanced heterozygous segments is, in certain embodiments, a most frequently occurring even-numbered absolute copy number within the tumor genome.

[0321] In this way, in certain embodiments, a subset of primary balanced heterozygous segments (PBHSs) provides a substantially uniform and / or sufficiently large population of heterozygous segments that can be leveraged to rigorously and / or robustly extract various parameters used to determine purity estimates and / or absolute copy numbers of segment throughout a tumor genome.

[0322] For example, as described in further detail herein, in certain embodiments, using PBHSs to estimate sample purity from CNV events allows for a robust estimation process. For example, given a set of PBHSs, certain auxiliary parameters (e.g., a primary slope, a primary standard deviation, a primary allele frequency standard deviation, and a primary residual standard deviation) can be initially estimated without making any assumptions on a tumor sample or tumor genome. Then, a candidate primary copy number can be assumed from a set of possible candidate primary copy numbers (for example, 2, 4, 6, and 8), and for each candidate primary copy number, a candidate purity that best explains the HSs is estimated, wherein the estimation of the purity makes use of one or more auxiliary parameters. Finally, a most probable primary copy number can be selected from the set of candidate primary copy numbers using, for example, a multivariate classier. Purity can then be determined as the candidate purity corresponding to the most probable primary copy number. Once purity has been estimated from CNV events it is possible to estimate all parameters. If purity estimation fails, it is still possible to provide an upper bound on a purity of the tumor sample.- 62 - 13247547v 1Attorney Docket No.: 2013237-0970

[0323] As described herein, CNV-based deconvolution techniques (e.g., purity estimation) can provide reliable estimates down to tumor sample purities as low as between 20-30% (0.2-0.3) (e.g., at purities ranging from 0.2 or 0.3 to 1) when whole exome sequencing data is used.

[0324] For example, once purity has been estimated from CNV events, which entails also determination of a most probable (optimal) primary copy number, it is possible to estimate: an absolute copy number of a segment in the tumor genome, an allele-specific copy number of a heterozygous segment, a ploidy, which can be the ploidy of the tumor genome or any part thereof, e.g., a ploidy of a chromosome and / or a ploidy of a region of chromosome.

[0325] An estimated ploidy may be used as an indication of how much replication has taken place within the tumor cell. This information can be used by other methods that require, for example, knowing an absolute amount of genetic matter (DNA) in a tumor cell. This method can complement or replace experimental methods for determining ploidy, such as flow cytometry, Spectral Karyotyping (SKY) analysis, digital PCR, quantitative PCR, and so on. Moreover, ploidy can be used to normalize absolute copy numbers of putative SNVs when determining suitability of a SNV as a neoepitope for a cancer vaccine.

[0326] In certain embodiments, estimation of an absolute copy number and an allelespecific copy number of HSs can be performed by assigning the HSs to the most likely nodes using a two-dimensional Gaussian mixture model (2d-GMM).

[0327] Additionally, or alternatively, in certain embodiments, if putative SNVs are provided to the method, once the purity has been estimated from CNV events it is possible to also estimate an absolute copy number of a putative SNV and / or an absolute copy number of an alternate allele (zygosity) of a putative SNV.

[0328] Additionally, or alternatively, the present disclosure includes the insight that under certain circumstances it may not be possible to reliably estimate purity of a tumor sample either from CNV events or from SNVs. Accordingly, e.g., instead of aborting purity estimation in such cases, the present disclosure recognizes that it is possible to obtain an upper bound on purity. Such an upper bound is useful, for example, in order to set more sensitive thresholds for mutation detection, which allows in turn to improve sensitivity and specificity of mutation detection. Although optimal thresholds can be obtained only when purity is precisely known, - 63 - 13247547v 1Attomey Docket No.: 2013237-0970even having an upper bound on purity can significantly improve the sensitivity of mutation detection.

[0329] The ability to estimate an upper bound on purity when purity cannot be estimated is particularly desirable when the method is intended for clinical and / or diagnostic applications. In such applications there is no option of dropping out a sample from analysis, as if often done in cohort analyses for research purposes (e.g., wherein methods are used for studying cohorts of tumor samples), because every patient sample needs to be analyzed in the best possible way.

[0330] In certain embodiments, accuracy of CNV-based purity estimation may be reduced, for example, if tumor sample purity is very low, significant CNV subclonality is present, there are very few or no CNV events in a tumor sample, sequencing coverage is not sufficiently high, or due to other confounding genomic and / or experimental factors. In certain embodiments, accuracy of SNV-based purity estimation may be reduced, for example, if tumor sample purity is very low, if there are too few or no minimal balanced SNVs, if there is significant SNV subclonality, if there are a significant number of false positive SNVs, or if there are a significant number misclassified minimal balanced SNVs. Accordingly, in certain embodiments, depending on the purity of the tumor sample, the genomic complexity of the tumor genome, and the quality of the sequencing and the quality of the mutation detection, purity estimation may be insufficiently accurate.

[0331] Accordingly, in certain embodiments, an upper bound on purity of a tumor sample may be calculated either based on CNV events, if a number of CNV events is not negligible, or based on SNV events, for example, when a number of CNV events is negligible.

[0332] Accordingly, certain embodiments described herein allow for estimation of an upper bound on purity based on CNV events (referred to as the CNV-based upper bound on purity) because the problem of purity estimation is separated into two parts: selecting a primary copy number from a small set of even integers and estimating an optimal purity for a given hypothesized primary copy number. Since the latter process is a highly robust optimization process, an upper bound is virtually always guaranteed. Since in a typical approach purity would be estimated by brute force over a large space of continuous variables, it is not feasible by such an approach to derive an upper bound on purity in a robust, accurate, and computationally reasonable manner.- 64 - 13247547v 1Attorney Docket No.: 2013237-0970

[0333] Estimation of an upper bound on purity from SNV events (referred to as the SNV-based upper bound on purity) can be achieved by assuming that all SNVs have a heterozygous genotype.

[0334] The decision which form of upper bound to provide as a result - the SNV-based upper bound on purity or the CNV-based upper bound on purity can be based in certain embodiments on the PBHSs: in certain embodiments, if the fraction of HSs that are MBHSs is above a predetermined threshold (e.g., above 80%, above 95%) then the SNV-based upper bound on purity is preferred since all putative SNVs are likely to have a minimal copy number of 2, and there are not enough CNV events on which to base a CNV-based calculation. Otherwise, in these embodiments, CNV-based upper bound on purity would be preferred.

[0335] Accordingly, tumor deconvolution systems and methods described herein may determine one or more parameters of a tumor sample. In certain embodiments, one or more parameters of the tumor sample comprise one or more of: a purity of the tumor sample based on CNVs, a purity of the tumor sample based on SNVs, an upper bound on the purity of the tumor sample based on CNVs, an upper bound on the purity of the tumor sample based on SNVs, a lower bound on the purity of the tumor sample based on SNVs, a quality score for the SNV-based purity estimation in the form of a p- value associated with the SNV-based purity, a quality score for the SNV-based purity estimation indicating a fraction of homozygous mutations in PBHSs, an absolute copy number of a segment comprising a heterozygous SNP, an absolute copy number of a segment not comprising a heterozygous SNP, an allele- specific copy number of a segment comprising a heterozygous SNP, an absolute copy number of one or more mutated alleles for one or more provided putative SNVs, an absolute copy number of a segment that contains a SNV, a fraction of diseased cells containing a putative mutated allele, a decision whether a SNV is present in all cells of the tumor sample, a decision whether an additional possible evolutionary scenario exists that could explain an expected SNV subclonality, a consistency score, a global density score, a specific density score, a quality score indicating whether the CNV-based purity estimation was successful or aborted, a quality score indicating whether a SNV deconvolution was successful or aborted, an estimated error rate of absolute copy number assignment based on parity correction, an amount of CNV subclonality, an amount of SNV subclonality, a genomic persistence distance, a ploidy of the tumor genome, and a gender of- 65 - 13247547v 1Attomey Docket No.: 2013237-0970the patient, a Clonality State (clone type) determined for each mutation. Approaches for determining clonality types are described, for example, in U. S. Provisional application no.63 / 749,339.A. Creating Personalized Cancer Immunotherapies Based on Tumor Genome Analysis

[0336] Certain cancer mutations are unique to a patient’s cancer and, when expressed, produce proteins and / or peptides that are distinct from those produced by normal cells. These distinct proteins and / or peptides can, accordingly, be specifically targeted via immunotherapy approaches that leverage the patient’s own immune system to clear cancer cells while avoiding damage to normal cells. Technologies of the present disclosure, among other things, leverage and analyze sequencing data to identify potential cancer-specific mutations within genome(s) of a patient’s tumor (tumor genome) and, moreover, characterize them in a manner that allows those mutations that will be the most effective targets of immunotherapies to be identified and prioritized, for example as targets for personalized cancer vaccines, T-cell receptor (TCR) therapies, and the like.

[0337] As illustrated in FIG. 1, biological samples obtained from a patient can be sequenced and the resultant sequencing data can be analyzed to detect and characterize mutations (SNV events) unique to the patient’s cancer cells. Detected mutations can be prioritized according to various metrics that reflect, for example, their prevalence as well as propensity to induce an immune response. Non-synonymous mutations that are both highly prevalent (e.g., present in a substantial fraction, up to all, of the patient’s cancer cells) and likely to be effective in priming a patient’s immune system to mount a strong response can, accordingly, be selected for inclusion in a personalized immunotherapy for the patient. A personalized therapeutic can thus be designed, manufactured, and administered to the patient as treatment.

[0338] For example, as illustrated in FIG. 1, in certain embodiments, normal 102a and tumor tissue 102b samples are obtained from the patient. Normal genomic DNA (gDNA) is extracted from a normal tissue sample 102a and tumor gDNA is extracted from a tumor tissue sample 102b. Sequencing (104) may then be performed using the extracted normal and tumor gDNA to generate sequencing data.- 66 - 13247547v 1Attorney Docket No.: 2013237-0970

[0339] Sequencing (104) may be performed, for example as described in further detail herein, using next generation sequencing (NGS) techniques. Accordingly, in certain embodiments, various pre-processing steps (106) are performed, for example to align reads of sequencing data to a reference genome.

[0340] In certain embodiments, sequencing data [e.g., having been pre-processed (e.g., to align reads to a reference genome)] is used (e.g., as input) for tumor deconvolution and / or mutation detection (108) techniques of the present disclosure. For example, as described in further detail herein, tumor deconvolution and / or mutation detection techniques of the present disclosure operate on sequencing data to determine one or more [biophysical and genomic] features 110 of a patient’s cancer. These determined [biophysical and genomic] features may characterize genomic properties of the patient’s cancer (e.g., to the extent represented in the tumor tissue sample) and / or physical properties of the tumor sample 102b. For example, in certain embodiments, one or more tumor genomic features 110a are determined. Tumor genomic features may include, without limitation, copy numbers (e.g., absolute copy numbers and / or allele- specific copy numbers) of one or more segments within a tumor genome (e.g., all segments; e.g., a particular subset of segments, such as heterozygous segments) and / or mutations (SNVs) therein, as well as characteristics of detected mutations (SNVs), such as their zygosity and / or clone type. In certain embodiments, one or more sample features 110b are determined. Sample features 110b may include, without limitation a sample purity and / or contamination fraction, which reflect the potential for and amount of tumor samples to include a non-trivial and, at times, substantial, fraction of normal, non-cancerous, cells. In certain embodiments, mutations (e.g., SNVs) are detected 100c. Detected mutations may, for example, be determined by tumor deconvolution and / or mutation detection technologies and provided, for example as a standardized file such as a variant call format (.vcf) file.

[0341] In certain embodiments, detected mutations 110c are filtered to identify non-synonymous mutations (112).

[0342] In certain embodiments, non-synonymous mutations are prioritized (114) to select a subset of mutations for inclusion in a personalized cancer immunotherapy 120. For example, in certain embodiments, mutations (e.g., non-synonymous mutations) are evaluated to determine which particular mutations are, or are predicted to be, present in a substantial fraction (e.g.,- 67 - 13247547v 1Attorney Docket No.: 2013237-0970above a certain threshold, up to all) of a patient’s cancer cells, so that an immune response that targets and eliminates cells expressing one or more particular mutations is likely to eliminate a substantial fraction of the patient’s cancer cells. In certain embodiments, additionally or alternatively, mutations that are determined to, or predicted likely to, elicit a strong immune response within the patient are prioritized and selected for.

[0343] For example, in certain embodiments, one or more scoring metrics 116 are determined for each of at least a portion of detected mutations 100c [e.g., a subset identified as non-synonymous (112), or a portion thereof]. Scoring metrics 116 may be used to prioritize and select a subset of mutations for inclusion in a personalized cancer immunotherapy 120. Scoring metrics 116 may include metrics characterizing potential for a particular mutation to elicit an immune response and may include, without limitation, major histocompatibility complex (MHC) binding predictions, T-cell receptor (TCR) recognition predictions, expression level predictions, and the like. These immune response metrics may be determined via various approaches, including, for example, machine learning and other techniques. Additionally, or alternatively, scoring metrics may include metrics determined via tumor deconvolution and mutation detection technologies described herein, such as, without limitation, absolute copy number values, clone type classifications, values and / or classifications indicating a function of a gene harboring a given mutation, values and / or classifications indicating whether a given mutation is truncal and / or early, and zygosity and / or fractional zygosity values.

[0344] In certain embodiments, a prioritized subset of detected mutations may be used for a personalized cancer immunotherapy 120. For example, in certain embodiments, compositions encoding one or more of a prioritized subset of mutations may be manufactured and administered to a patient, for example as a personalized cancer vaccine.A.i Patient Samples and Sequencing Data

[0345] Turning to FIG.2, as described herein, tumor deconvolution and / or mutation detection technologies of the present disclosure may be used in connection with (e.g., to analyze) sequencing data for a subject to determine genomic and biophysical properties of, and detect mutations characteristic of, a patient’s cancer. Among other things, technologies of the present disclosure may be used in connection with sequencing and / or genomic analysis methods and - 68 - 13247547v 1Attomey Docket No.: 2013237-0970systems that generate sequencing data for a subject and allow for detection of mutations that can be targeted, for example, for disease treatment via immunotherapy. Among other things, technologies of the present disclosure may be used in connection with sequencing and / or genomic analysis methods and systems that generate sequencing data for a subject and allow for detection of mutations that can be targeted, for example, for disease treatment via immunotherapy.

[0346] For example, as shown in FIG. 2, in certain embodiments, sequencing data 228 may be generated for a subject having and / or suspected of having cancer, by obtaining a tumor sample 202 from the subject. The tumor sample 202 may be processed 204, for example, to extract and prepare nucleic acid material for sequencing and sequenced 206 to generate tumor sequencing data 208 - e.g., sequencing data representing and obtained from nucleic acid material 204 from a tumor sample 202. Likewise, in certain embodiments, sequencing data 228 may (e.g., also) include normal sequencing data 218 from normal sample 212 having been obtained, processed 214, and sequenced 216, to generate normal sequencing data 218 - e.g., sequencing data representing, and obtained from, nucleic acid material from a normal sample 212.

[0347] A tumor sample 102 may be any sample derived from a particular subject and comprising, and / or expected to comprise, cancer cells (e.g., of the particular subject). In certain embodiments, a tumor sample is or comprises a liquid sample, such as serum, plasma, blood, urine, etc. For example, a liquid sample, such as blood, may comprise, or be suspected of comprising, cancer cells, such as circulating tumor cells (CTCs). In certain embodiments, a tumor sample is or comprises a tissue sample, for example, obtained from a subject via biopsy. A tumor sample may be representative of a subject’s primary tumor and / or one or more metastases. For example, a primary tumor sample may be obtained via biopsy of a region of a subject known or expected to harbor a primary tumor. A metastasis sample or metastases samples may be obtained via biopsy of one or more region(s) of a subject known or expected to harbor metastases. Tumor samples may be obtained and / or preserved in a variety of formats, such as, for example, isolated cells (e.g., CTCs) from blood samples, fresh tissue, flash frozen tissue, formalin-fixed paraffin embedded tumor tissue (FFPE samples), and the like.

[0348] As illustrated in FIG. 2, a tumor sample may comprise cancer cells 202b as well as, in certain cases, normal (i.e., non-cancerous) cells 202a. A tumor sample may contain a - 69 - 13247547v 1Attomey Docket No.: 2013237-0970normal genome in addition to the tumor genome. For example, since tumor samples excised from patients, they are typically impure, containing, for example, normal cells from the stroma, normal tissue, and / or tumor infiltrating lymphocytes, and so on, which lead to contamination of the tumor sample with the normal genome of the subject. Tumor cell lines, on the other hand, may be considered to be pure tumor samples with no normal genome contamination (normal contamination). Accordingly, in certain embodiments, a tumor sample purity may measure relative fraction of cancer cells within a tumor sample. Tumor sample purity may, for example, be computed as (1 — j ) =+ ’IN), where LI is a contamination fraction, representing a relative fraction of normal cells infiltrating a tumor sample, given by j = rN / (TT+?]w) and T T and T N are a number of tumor and normal cells in a tumor sample, respectively. Tumor sample purity may be expressed as a decimal value, percentage, etc. As described in further detail herein, typically, purity does not need to be measured directly (e.g., via direct measuring / counting amounts of tumor and normal cells in a sample), but, rather, can be determined and / or estimated using sequencing data, for example, via tumor modelling approaches such as those described herein.

[0349] As described in further detail herein, in certain embodiments, additionally or alternatively, bounds, such as upper and / or lower bounds for sample purity (e.g., and / or contamination fraction) may be estimated, e.g., via tumor modeling approaches of the present disclosure. For example, as described in further detail herein, in certain embodiments, at low physical sample purities, accuracy of estimation methods, such as tumor deconvolution, may be reduced such that purity estimates based on sequencing data are expected to be of insufficient accuracy to be used in and of themselves. In such cases, however, an upper bound (e.g., a maximum purity) and / or a lower bound e.g., a minimum purity) may still be estimates and used in certain processing steps.

[0350] A normal sample 212 may be any sample derived from a particular subject and comprising, and / or expected to comprise, the subject’s normal cells, but not cancer cells (e.g., in certain embodiments, a normal sample 212 does not contain any cancer cells). In certain embodiments, a normal sample is or comprises a liquid sample, such as serum, plasma, blood, urine, saliva, etc. In certain embodiments, a normal sample is or comprises a tissue sample, for example obtained from a subject via biopsy. Normal samples may be obtained and / or preserved- 70 - 13247547v 1Attorney Docket No.: 2013237-0970in a variety of formats, such as, for example, isolated cells [e.g., peripheral blood mononuclear cells (PBMCs)] from blood samples, fresh tissue, flash frozen tissue, formalin-fixed paraffin embedded tumor tissue (FFPE samples), and the like.

[0351] As illustrated in FIG.2, a normal sample 212 nominally contains only normal patient cells 212a. In certain embodiments, a normal sample may be obtained from blood of a subject. In certain embodiments, a normal sample may still comprise a small (e.g., negligible) number of cancer cells. For example, in certain embodiments, a normal cell may be obtained from a region of a subject near a tumor e.g., in an effort to obtain a normal sample from a same or similar underlying tissue type). In certain embodiments, a normal sample comprises less than 1%, e.g., less than 0.1%, e.g., less than 0.01%, e.g., less than 0.001% tumor cells.

[0352] Samples, such as tumor samples and / or normal samples, may be processed to obtain, and / or prepare, nucleic acid material therefrom for sequencing. For example, nucleic acid material, such as DNA and / or RNA, may be extracted and prepared for sequencing (e.g., via amplification, fragmentation, labeling, etc.) as appropriate, depending on a particular desired sequencing method and / or data format. In certain embodiments, a tumor sample 202 is processed 204 and / or normal sample 212 is processed 214 to extract nucleic acid material, such as gDNA. In certain embodiments, normal gDNA is extracted from normal tissue sample 212 and tumor gDNA is extracted from tumor tissue sample 202. For example, in certain embodiments, sequencing data may be whole genome sequencing (WGS) data; in certain embodiments, sequencing data may be whole exome sequencing (WES) data. Various commercially available kits and instruments may be used to prepare samples for and obtain WGS and / or WES data, including, but not limited to, those provided by Illumina, Inc., PacBio, Oxford Nanopore Technologies, Thermo Fisher Scientific’s Ion Torrent™, etc. In some embodiments, methods and systems of the present disclosure utilize whole exome sequencing (WES) or whole genome sequencing (WGS) data in singletons (e.g., one sequenced tumor sample and one sequenced normal sample). In some embodiments, methods and systems of the present disclosure utilize whole exome sequencing (WES) or whole genome sequencing (WGS) data in replicates (e.g., two or more sequenced tumor samples and two or more sequenced normal samples). In some embodiments, methods and systems of the present disclosure utilize whole exome sequencing (WES) or whole genome sequencing (WGS) data in conjunction with RNA-seq data.- 71 - 13247547v 1Attomey Docket No.: 2013237-0970In some embodiments, RNA-seq data is used in a method or system of the present disclosure to filter errors such as PCR errors, FFPE, sequencing artifacts, and / or other types of errors. For example, a mutation detected in DNA sequencing data may only be accepted if it is also detected using RNA-seq reads. In some embodiments, RNA seq reads are high quality RNA-seq reads. In some embodiments, if only singletons are used, coverage of singletons is required to be comparable to coverage of merged replicates (e.g., in order to ensure comparable sensitivity).

[0353] In certain embodiments, to generate sequencing data, libraries are created from the normal gDNA and tumor gDNA. In certain embodiments, from each gDNA sample, two or more libraries can be generated. In certain embodiments, from each gDNA sample, two libraries can be generated. For example, in certain embodiments, as illustrated in FIG.2, sequencing data 228 may be generated in and / or comprise replicates created by preparing multiple (e.g., two or more) libraries associated with each (e.g., gDNA) sample. The libraries can be created for whole exome sequencing and / or whole genome sequencing and / or RNA sequencing (RNAseq). Samples are then sequenced using high throughput sequencing such as NGS.

[0354] For example, in certain embodiments, sequencing data 228 may be generated in and / or comprise replicates. For example, multiple tumor samples may be extracted and sequenced independently; in certain embodiments, a single tumor sample may be extracted and used to prepare multiple libraries (e.g., such that processing steps of extracting, fragmenting, and amplifying nucleic from the sample are performed repeatedly and independently), which are then sequenced; in certain embodiments, library preparation may be performed repeatedly on a single pool of extracted nucleic acid, and the multiple libraries sequenced; in certain embodiments, a single library is sequenced multiple times (e.g., as in a technical replicate). In certain embodiments, as with tumor sample sequencing data, e.g., as illustrated in FIG.2, normal sample sequencing data may also comprise a plurality of replicates.

[0355] Sequencing data 228, 208, 218, may be stored and / or presented in a variety of formats, such as FASTQ, SAM, BAM, etc. For example, sequencing data for a particular sample may comprise a plurality of reads, each read representing a nucleotide sequence of a polynucleotide fragment corresponding to e.g., that maps to) a portion of a subject’s tumor genome and / or exome, and / or portion of the subject’s normal genome and / or exome. In certain embodiments, sequencing data typically comprises multiple overlapping reads, which may be - 72 - 13247547v 1Attorney Docket No.: 2013237-0970aligned to a reference genome, such as an hl9 or h38 reference genome (see, e.g., Ncbi.nlm.nih.gov / datasets / genome / GCF_000001405.13 / and Ncbi.nlm.nih.gov / datasets / genome / GCF_000001405.26 / , respectively; see also Karolchik, D. et al. Nucleic Acids Res. 32, D493-D496, 2004 and Kent WJ et al., Genome Res.12(6):996-1006, 2002) for human patients, to map each read to a particular region of the overall genome. A reference genome may, accordingly, be a publicly available reference genome. In certain embodiments, a reference genome may be generated based on a normal genome of the subject. In certain embodiments, a reference genome may be used to provide a coordinate system for the normal genome and the diseased genome, e.g. the tumor genome. A reference genome can be used for mapping reads and providing a coordinate system for the normal genome and the tumor genome, wherein the coordinate system allows for the provision of the chromosome number, a nucleotide position in the chromosome, as well as directionality of the read. A reference genome can be based on the genome of one or more members from the same species as the subject providing the diseased sample, or can be based on the normal genome of the subject (a matched genome).

[0356] For example, in certain embodiments, NGS sequencing is performed and generates FASTQ files as output. In certain embodiments, NGS output reads, stored, for example, in FASTQ files, are aligned. Aligned reads may be stored and / or provided in file formats, such as sequence alignment map (SAM) or the binary compressed version thereof (BAM). In certain embodiments, sequencing data duplicate reads may be marked and / or removed from the sequencing data. In certain embodiments, sequencing adapters may be removed from the sequencing data. In certain embodiments, e.g., in the case of short read sequencing, sequencing can be paired end or single end, wherein different read lengths can be used (e.g., 50bp, lOObp, 150bp, etc.).

[0357] In certain embodiments, aligned reads, e.g., as stored in one or more BAM file(s), are used as input for tumor deconvolution and / or mutation detection technologies described here. For example, in certain embodiments, samples are sequenced using high throughput sequencing such as next generation sequencing, generating, for example, FASTQ files as output. Next, FASTQ files are aligned and, in certain embodiments, duplicate reads are marked. An alignment step may be performed, and generate, as output, a BAM file for a normal sample and a BAM file- 73 - 13247547v 1Attomey Docket No.: 2013237-0970for a tumor sample. In certain embodiments, output of an alignment step is two normal BAM files and two tumor BAM files.

[0358] As described in further detail in the following, based on input sequencing data (e.g., comprising tumor sequencing data and normal sequencing data), tumor deconvolution and mutation detection technologies of the present disclosure may determine various properties of a tumor sample and detect mutations (e.g., SNVs) occurring within the tumor genome.A.i.a Sequencing Data Reads, Alignment, and Quality Scores

[0359] In certain embodiments, sequencing information is provided in the form of reads after alignment to a reference genome.

[0360] In certain embodiments, normal reads are reads sequenced from the normal sample, and tumor reads are reads sequenced from the tumor sample.

[0361] Normal reads sequenced from the normal sample can span the entire genome, the exome, the coding regions, or any part, segment, or region of the genome.

[0362] Tumor reads sequenced from the tumor sample can span the entire genome, the exome, the coding regions, or any part, segment, or region of the genome.

[0363] In certain embodiments, normal reads and tumor reads span the same regions of the genome.

[0364] In certain embodiments, a normal read / tumor read can comprise all kinds of information obtainable from next generation sequencing or other sequencing methods. For example, a normal read / tumor read can comprise a nucleotide sequence, a paired read if the sequencing is paired end, and / or a quality score per base.

[0365] A read alignment can comprise all kinds of information obtainable from next generation sequencing or other sequencing methods after the mapping of a read to the reference genome. For example, a read alignment can comprise an normal or tumor read (pair), an indication if the read is a forward read or a reverse read, wherein a forward read is a read mapped to the forward strand, and a reverse read is a read mapped to the reverse strand, a mapping quality score, and / or a chromosome and a position to which the read maps in a reference- 74 - 13247547v 1Attorney Docket No.: 2013237-0970genome, wherein it is understood that forward tumor reads are forward reads of the tumor sample, and reverse tumor reads are reverse reads of the tumor sample.

[0366] After mapping each read to the site in the reference genome defined by a chromosome and a position in the chromosome, the following information may be obtainable from next generation sequencing for each site per given replicate: a number of reads mapping to each of the four possible nucleotides, where reads are possibly filtered to retain only reads exceeding a predetermined quality score threshold, a number of forward reads mapping to each of the four possible nucleotides, where reads are possibly filtered to retain only reads exceeding a predetermined quality score threshold, a number of reverse reads mapping to each of the four possible nucleotides, where reads are possibly filtered to retain only reads exceeding a predetermined quality score threshold, a coverage of all reads, a coverage of all forward reads, a coverage of all reverse reads, a list of measured nucleotides and their corresponding quality scores, and / or a reference base.

[0367] A quality score can reflect the accuracy of base calling at a particular site, and can indicate the probability that a given base is called incorrectly. In certain embodiments, the quality score can indicate the probability that a given base is called incorrectly by the sequencing instrument. In certain embodiments, the quality score of a given base in a given read can reflect the probability that the given base in the given read is called incorrectly by the sequencing instrument.

[0368] Quality scores may be measured in different ways and based on different scales. In certain embodiments, a quality score can mean a property that is logarithmically related to a probability of the base error to occur, such that the quality score, Q, can be given by -10 multiplied by the log (base 10) of the probability of the base error to occur, P, i.e., Q =— 101og10P. For example, according to the above definition, a quality score of 30 to a base can mean that the probability of an incorrect base call is 1 in 1000 times, which implies that the probability for a correct base call is 99.9%. Likewise, a quality score of 20 to a base can mean that the probability of an incorrect base call is 1 in 100 times, which implies that the probability for a correct base call is 99%, and so on and so forth. In certain embodiments, it is assumed that quality scores are inversely proportional to a probability of the base error to occur and / or the error rate per base, and without limiting the generality of the method it is assumed that P =- 75 - 13247547v 1Attomey Docket No.: 2013237-0970— 101og10P. It is understood that if a different relationship between P and Q is invoked, then related definitions, such as quality score filtering, as well as thresholds are changed accordingly, without limiting the generality of the method.

[0369] A coverage of a given site can refer to a total number of reads mapping to a given site. A coverage can be calculated given all measured reads, or given a subset of reads. For example, a coverage can be calculated based on all reads with a quality score equal to or above a given predetermined threshold. A coverage can be calculated based on the forward reads or a subset thereof, or based on the reverse reads or a subset thereof, with or without retaining reads with a quality score equal to or above a given predetermined threshold.

[0370] A read can be said “to map to” or “to contain” an allele at a given site in the genome if at the position in the read corresponding to the given site the read encodes the said allele.

[0371] In some embodiments, a read can be a specific base in the read mapping to a given site in the genome. Although for every base in a read there may be a corresponding different quality score, it is understood that when “a read is said to have a quality score”, this can refer to the quality score of the base in the read corresponding to the given site. Thus, when a read is said to have a quality score, this can refer to the quality score of the base in the read corresponding to a given site (as opposed to all other quality scores provided for all other nucleotides in the said read), wherein the given site can include, for example, a site in the genome, a putative mutation (that occurs at a given site in the genome), or a particular allele (that corresponds to a given site in the genome). Likewise, “the error rate per base of a given read” may be interpreted as the error rate per base of the specific base in the read corresponding to the given site, which is determined by the quality score of the base in the read corresponding to the given site.

[0372] For example, “a quality score of read mapping to a given allele” (at a presumed site in the genome) can refer to the quality score of the nucleotide corresponding to the given site in the read, wherein the said read encodes the given allele at the position in the read corresponding to the given site.

[0373] In certain embodiments, a read can be mapped to a reference genome by finding the position in the reference genome to which the read most likely originated from, wherein the - 76 - 13247547v 1Attorney Docket No.: 2013237-0970accuracy of mapping can be given by a mapping quality score. In certain embodiments, the higher the mapping quality score the more likely the read originated from the mapped position. In other embodiments, a read can comprise one or more measured bases of the sequenced genome and the position of the measured base with respect to the reference genome and / or sequenced genome can be obtained by other means.

[0374] In certain embodiments, a read can comprise one or more nucleotides that can be mapped to a known position of the genome with high fidelity. In certain embodiments, a read can span the entire sequenced genome. In certain embodiments, the abundance of segments in the sample being sequenced can be provided by another measured parameter that may be discrete or continuous.A.i.b Sequencing Data Error Rates

[0375] In certain embodiments, tumor deconvolution methods and systems described herein account for error rates, e.g., in calling particular bases, that may occur during sequencing, such as NGS, and which may be present in reads of sequencing data. An error rate per base may be estimated in different ways. In certain embodiments, the error rate per base may be equal to the average error rate per base reported by the manufacturer of the sequencing device.Additionally or alternatively, in certain embodiments, the error rate per base may be estimated based on quality scores provided by, for example, the sequencing instrument, or by a software analyzing the results produced by the sequencing instrument. The error rate per base may also be estimated based on an analysis of the reads and related information including, for example, the quality scores provided by the sequencing instrument and / or the quality of mapping and / or alignment of the reads.

[0376] The error rate per base may be estimated for each base measured by the sequenced instrument. The error rate per base may also be estimated for all bases measured by the sequencing instrument mapping to a particular site. The error rate per base may also be estimated for all bases measured by the sequencing instrument for all sites in the sequenced genome.

[0377] The error rate per base may also be estimated based on the quality score of the individual base in the individual read mapping to the given site. The error rate per base may also - 77 - 13247547v 1Attorney Docket No.: 2013237-0970be estimated based on the quality scores of one or more bases encoded in one or more reads mapping to a given site in the genome.

[0378] The error rate per base can also be an upper bound on the probability of the base error to occur. For example, the error rate per base may also be estimated based on the minimal quality score of one or more bases encoded in one or more reads mapping to a given site in the genome, wherein it is also possible to filter the reads such that only reads equal to or exceeding a predetermined quality thresholds are retained, where quality scores are assumed, without limiting the generally of the method, to be inversely proportional to the probability of the base error to occur.

[0379] The error rate per base may also be an upper bound of any of the aforementioned error rates per base. The error rate per base can be calculated separately for the normal reads and for the tumor reads. In certain embodiments, the error rate per base is derived from the quality scores of the reads used for the analysis.

[0380] In certain embodiments, the method comprises a step wherein reads are quality score filtered by retaining only reads with quality scores above or equal to a predetermined quality score threshold and assuming that the error rate per base for all retained reads can be given by an upper bound provided by a constrained error rate per base, wherein the constrained error rate per base can be given by the error rate corresponding to the predetermined quality score threshold. In certain embodiments, the predetermined quality score threshold is in the range of 20 to 40. In certain embodiments, the predetermined quality score threshold is in the range of 25 to 35.

[0381] Quality score filtering can enhance the accuracy of statistical tests that require knowledge of an error rate per base compared to approaches wherein an average error rate per read is assumed.

[0382] A false positive may be understood to occur if a prediction that a site encodes a mutation is accepted, when in fact no mutation exists at the given site. A potential false positive may be understood to occur if a prediction that a site encodes a mutation is accepted, when there is a potential that no mutation exists at the given site. Filtering false positives may be understood to be the rejection of sites that are determined by a statistical or other test to potentially lead to the false positive unless rejected. Filtering false positives may also be understood to be the - 78 - 13247547v 1Attomey Docket No.: 2013237-0970rejection of potential false positive. Filtering false positives may reduce the rate of false positive predictions of the method. The objective of a false positive filter may be to filter false positives.A.i.c Quality Score Filtering

[0383] In certain embodiments, reads can be quality score filtered, which means that only reads with quality scores above or equal to a predetermined quality score threshold of Q are retained for analysis, so that a probability of a base error to occur is less than or equal a constrained error rate per base given by 10-Q / 10, whereby, without limiting the generally of the method, the aforementioned definition of the quality score is invoked. Quality score filtering may also imply that the error rate per base is set to the constrained error rate per base, which in turn allows to use the error rate per base in rigorous statistical tests because the error rate per base serves as an upper bound on the true probability of the base error to occur in any of the retained reads. In certain embodiments, a predetermined quality score threshold Q is a number in the range of 20 to 40. In certain embodiments, a predetermined quality score threshold Q is a number in the range of 25 and 35. In certain embodiments, a predetermined quality score threshold Q is a number in the range of 27 and 33. A predetermined quality score threshold may increase as the sequencing technology and reagents improve with time.A.ii Tumor Genomics

[0384] Turning to FIG.3, sequencing data may be used to piece together and / or infer properties relating to a tumor genome and / or normal genome of a subject.A.ii.a Segments

[0385] As illustrated in FIG.3, a genome may be subdivided into a plurality of segments, each segment representing a particular sub-region (e.g., a subsequence) of the genome. In certain embodiments, a reference genome may be subdivided into a plurality of segments. In certain embodiments, a tumor genome and / or a normal genome may be subdivided into a plurality of segments (black bars in the normal genome and tumor genome schematics).- 79 - 13247547v 1Attomey Docket No.: 2013237-0970

[0386] In certain embodiments, segments are non-overlapping segments. In certain embodiments, a segment can span a gene, e.g. as defined in a reference genome that the reads are aligned to. In certain embodiments, a segment can be a fragment of a gene, an exon, a union of exons, or the union of exons associated within a given gene. In certain embodiments, various other sets of predetermined regions in a reference genome (with or without introns), or another set of predetermined regions in a reference genome based on the normal genome, may be used as a basis for subdividing genomes into a plurality of segments. In specific embodiments, a segment can be a region of a reference genome with a given constant copy number and / or a given allele- specific copy number in the diseased genome or alternatively a fragment of a gene with a given constant copy number and / or allele- specific copy number in the diseased genome. A segment can be defined to include or to exclude introns.

[0387] Among other things, in certain embodiments, as illustrated in FIG.3, this approach provides a common coordinate system or set of subregions between reference genome 302, normal genome 322, and tumor genome 342. For example, (e.g., sequencing data, such as reads, corresponding to) normal genome 322 and (e.g., sequencing data, such as reads, corresponding to) tumor genome 342 may be aligned to reference genome 302. In this way reads can be mapped to particular regions of reference genome 302 and a coordinate system for reads of a normal genome 322 and tumor genome 342 can be provided. For example, reference genome can be used to determine a chromosome number, a nucleotide position (in the chromosome), and / or a directionality of a given read.

[0388] In certain embodiments, when a segment spans more than a predetermined length threshold it may be discarded from analysis for use in tumor deconvolution approaches described herein. In certain embodiments, a predetermined length threshold is in the range of 0.1 to 20 megabases (Mb). In certain embodiments, a predetermined length threshold is in a range of 0.1 to 3 Mb. Embodiments of methods and systems described herein may assume that segments spanning lengths below the predetermined length threshold do not have breakpoints (a change in a copy number), or that the number of segments spanning lengths below the predetermined length threshold that contain breakpoints is very small compared to the total number of segments considered.- 80 - 13247547v 1Attorney Docket No.: 2013237-0970

[0389] Thereafter, by partitioning reference genome 302 into a plurality of segments, regions of normal genome 322 and tumor genome 342 that correspond to a particular segment can be identified and, for example, analyzed and compared.

[0390] For example, in the illustrative schematic shown in FIG.3, nine segments 304a, 304b, 304c...304i, and 304j (collectively 304) of reference genome 302 are shown, with corresponding segments of normal genome 322 and tumor genome 342 indicated via the vertical dashed lines.A.ii.b Copy Number Variations

[0391] As illustrated in FIG.3, the majority of normal genome 322 is typically diploid, with two copies of each particular segment (e.g., except for segments located on male sex chromosomes for which normal genome 322 contains either a single maternal chromosome or a single paternal chromosome), with one copy associated with (e.g., originating from) a maternal allele 326a and another associated with (e.g., originating from) a paternal allele 326b. In certain embodiments, methods and systems described herein may exclude male sex chromosomes and / or segments that map to locations on male sex chromosomes from analysis since male sex chromosomes do not contain heterozygous SNPs, which, as explained in further detail herein, can be used to facilitate determining tumor genomic properties and / or sample purities.

[0392] As illustrated in FIG.3, for segments in a normal e.g., human) genome that, absent germline CNV events, is diploid, segments may be heterozygous, in that the copies differ, corresponding to different alleles, or may be homozygous, comprising identical alleles.

[0393] In contrast, segments in a tumor genome 342 are not necessarily diploid. Nor are segments in a tumor genome necessarily balanced. Certain segments of a tumor genome may, however, be diploid and / or balanced. For example, as shown in FIG.3, a number of copies of each allele in a segment may vary from segment to segment, and may be less than two, equal to two, or greater than two. Segments in a tumor genome need not be balanced - i.e., tumor genome heterozygous segments do not necessarily comprise a same number of copies of each allele, but, in certain embodiments, may comprise a greater number of copies of one allele (a “major allele”) than the other (a “minor allele”).- 81 - 13247547v 1Attomey Docket No.: 2013237-0970

[0394] Accordingly, in certain embodiments, various parameters are used to characterize and represent physical properties of a particular segment (e.g., within a normal and / or tumor genome), along with, in certain embodiments, mutations (such as SNVs) identified therein.

[0395] For example, in certain embodiments, a segment j may be characterized by an absolute copy number, CAj, which is computed as a number of copies of a particular segment (e.g., within a single tumor cell).

[0396] For example, in the schematic shown in FIG.3, while all segments in normal genome 322 have one copy of a maternal allele and one copy of a paternal allele, copy numbers of maternal and paternal alleles in tumor genome 342 may vary from segment to segment. Accordingly, while all segments in normal genome 322 that are shown in FIG.3 have an absolute copy number of two, in tumor genome 342 segments 344a, 344b, 344c, 344d, 344e, 344f, 344g, 344h, 344i, 344j have absolute copy numbers of 2, 4, 3, 1, 3, 5, 0 (a deletion), 2, 4, and 1, respectively. Moreover, as illustrated in FIG.3, in tumor genome 342, the number of copies of maternal and paternal alleles also may vary from segment to segment such that, for example segment 344a has one copy of each a maternal and paternal allele, segment 344b has two copies of each, and segment 344f has three copies of the maternal allele and two copies of the paternal allele.

[0397] Normal segments can also have other copy number configurations due to germline CNV events, however, without limiting the generality of the method, such events are not described in the figure.A.ii.c Single Nucleotide Polymorphisms and Heterozygous Segments

[0398] In certain embodiments, maternal and paternal alleles of a particular segment in a normal genome 322 may harbor different variants of a single nucleotide polymorphism (SNP). In this case, the particular segment is referred to as comprising a heterozygous SNP. The particular segment in the normal genome 322, along with corresponding segments in reference and tumor genomes, are referred to as heterozygous segments.

[0399] For example, as illustrated in FIG.3, reference genome 302 is subdivided into a plurality of segments. For segment 304a of reference genome 302, corresponding segment 324a- 82 - 13247547v 1Attomey Docket No.: 2013237-0970in normal genome 322 comprises at least one heterozygous SNP, such that segment 304a of reference genome 302 and corresponding segments 324a and 344a of normal and tumor genome, respectively, are referred to as heterozygous segments. In contrast, segment 304i of reference genome corresponds to segment 324i of normal genome, which does not comprise any heterozygous SNPs. Accordingly, segment 304i and corresponding segments 324i and 344i in normal and tumor genomes are not heterozygous segments.

[0400] Overall, FIG.3 depicts seven (7) heterozygous segments and three (3) homogeneous segments. As illustrated in the figure, for each of the heterozygous segments, copies in the normal genome 322 harbor at least one heterozygous SNP (illustrated schematically via different colored orange and green dots in the maternal and paternal alleles), whereas normal genome 322 copies of the homogeneous segments do not contain any heterozygous SNPs.Different normal genomes (e.g., of different individual subjects) will generally have different heterozygous segments since different normal genomes encode different heterozygous SNPs.

[0401] As explained herein, while normal genome 322 typically has two copies of each segment (apart from those located on male sex chromosomes) the number of copies of particular segments in a tumor genome may differ from two and can vary from segment to segment.Additionally or alternatively, as illustrated in FIG.3, tumor genomes do not necessarily have equal numbers of maternal and paternal alleles for each heterozygous segment and / or, in certain cases, may entirely lack a maternal or paternal copy. For example, segment 304a is a heterozygous segment for which corresponding tumor genome segment 344a has two copies: one maternal and one paternal. Second segment 304b is another heterozygous segment. In this case, however, corresponding segment 344b in tumor genome 342 has four copies: two maternal copies of the segment in the tumor genome and two paternal copies of the segment in the tumor genome. Therefore, the segment has an absolute copy number of 4 in the tumor genome.Reference number 344c shows three copies of a heterozygous segment: one maternal copy of the segment in the tumor genome and two paternal copies of the segment in the tumor genome, therefore the segment has an absolute copy number of 3 in the tumor genome.

[0402] In certain embodiments, a heterozygous segment may comprise only paternal or only maternal alleles, referred to herein as a loss of heterozygosity (LOH) event. For example, segments 344d and 344e in tumor genome each correspond to a normal diploid segment that - 83 - 13247547v 1Attomey Docket No.: 2013237-0970comprises a heterozygous SNP - i.e., a heterozygous segment. However, as illustrated in FIG.3, segments 344d and 344e in tumor genome lack any maternal allele copies - they have only (one and three, respectively) copies of the paternal alleles. Accordingly, while these are segments that have undergone a LOH event.

[0403] In certain embodiments, for a given segment (e.g., in reference or normal genome), the corresponding segment may be entirely absent from tumor genome 344g, indicating that, for example, all copies of the segment (both maternal and paternal) were deleted (complete deletion). Segment 344g in tumor genome, accordingly, has an absolute copy number of 0.A.ii.d Tumor Genome Mutations and Characteristics

[0404] FIG. 3 illustrates an expanded view of segments 322f and 344f. As illustrated in FIG. 3, a normal cell 352 comprises normal genome 322, including maternal 354a and paternal 354b alleles of segment 322f, with maternal allele comprising an alternative variant of SNP 356.Tumor cell 372 comprises tumor genome 342, including segment 344f, which corresponds to normal segment 322f. Unlike normal cell, tumor cell 372 comprises multiple copies of each allele - namely, three copies of maternal allele and two copies of paternal allele.

[0405] As described herein, corresponding segments 322f and 344f in normal 322 and tumor 324 genomes, respectively, can be characterized by values of parameters such as an absolute copy number and an allele specific copy number. For example, as shown in FIG.3, normal segment 322f has an absolute copy number of two - (CAwt = 2) and corresponding tumor genome segment 344f has an absolute copy number 373 of five (e.g., CNj = 5). As illustrated in FIG. 3, although a normal segment will typically have one copy of each parental allele, CNV events may cause a tumor genome segment to have one or multiple (e.g., two or more) copies of each parental allele. Moreover, the number of maternal and paternal alleles need not be equal in a tumor genome. Accordingly, heterozygous tumor genome segments may also be characterized by an allele- specific copy number. For heterozygous segments, such as segment 354, an allele specific copy number 364 may be determined as a maximum absolute number of copies of a major allele - i.e., the allele having a number of copies greater than or equal to that of the other, minor, allele - i.e., CNx > CNx, where X and Y denote the major and minor alleles, respectively.- 84 - 13247547v 1Attomey Docket No.: 2013237-0970For example, for the particular segment 354 shown in FIG. 3, there are three copies of the major allele, such that the allele specific copy number is three (CNx = 3).

[0406] As shown in the bottom portion of FIG. 3, tumor genome segments may harbor mutations, such as a SNV 382. In the notation used herein, a copy number of a segment harboring a mutation may be denoted CAmut (e.g., for segment 344f in tumor cell 372, CAmut = 5). As shown in FIG. 3, mutations may occur in particular alleles, but are not necessarily present in each copy of a particular allele. For example, while SNV 382 occurs in maternal allele 374, it is not present in all three copies of maternal allele 374 - it is present in only two copies.Accordingly, additional parameters may be determined to characterize genomic properties of mutations like SNVs.

[0407] For example, if a certain segment is mutated, a number of physical copies of the mutated segment in a given tumor cell, referred to herein as the zygosity of the mutation, may be determined. In certain embodiments, a fractional zygosity of a mutation (Q may be computed as a ratio of the zygosity of a mutation (Cx) and the absolute copy number (CAmut) of the segment harboring the mutation in the tumor genome. For example, in FIG. 3, segment 344f harbors a mutation 382 having a zygosity of two and a fractional zygosity, of 2 / 5 (0.4).

[0408] Turning to FIG. 4, as described herein, a sample of tumor tissue, may comprise normal cells 400 and tumor cells 430. As shown in FIG. 4 and described herein, segments in normal cells 415 are assumed to have an absolute copy number of two, except for segments occurring on male sex chromosomes for which an absolute copy number is one. As shown in the figure, a segment may be amplified in the tumor cells. In the example illustrated in FIG. 4, the tumor cell segment has an absolute copy number of five 420. As shown in the figure, example mutation 430 (black star) has three copies, and, accordingly, its zygosity is three (440). In the illustrative example shown in FIG. 4, mutation 430 (black star) is present in all tumor cells and is therefore a clonal mutation, whereas a second mutation 450 (red star) is present in a subset of tumor cells (just one of the three tumor cells in the diagram in FIG. 4) and is therefore a subclonal mutation. Clonal mutations with a zygosity greater than 1 can arise, for example, when the mutation occurred before the copy number amplification event (hence are considered “early” mutations), whereas subclonal mutations with a zygosity of 1 can arise, for example, when the mutation occurred after the copy number amplification event (hence are considered - 85 - 13247547v 1Attomey Docket No.: 2013237-0970“late” mutations). In this model all CNV events are considered to be clonal, however, in a more general model, CNV events can also be present in just a subset of tumor cells.

[0409] In certain embodiments, a cellularity of a mutation, denoted by p, may be determined. Cellularity as used herein refers to the fraction of tumor cells that harbor a given mutation. A mutation is said to be biologically clonal if all cancer cells in a tumor sample harbor the given mutation. Biologically clonal mutations are characterized by having a nominal cellularity of 1 (p = 1). For example, clonal mutation 430 is present in all three tumor cells and, accordingly, has a nominal cellularity of 1, whereas subclonal mutation 450 is present in 1 out of 3 tumor cells, and, accordingly, has a nominal cellularity p = 1 / 3.

[0410] In certain embodiments, approaches described herein determine cellularity estimates for mutations. In certain embodiments, a cellularity estimate is an estimated mean cellularity (e.g., indicating, if multiple tumor samples were obtained and sequence, a given mutation is estimated to be present, on average, in a fraction of tumor cells given by the mean cellularity). In certain embodiments, confidence intervals for cellularity estimates may be determined, with lower and upper bounds denoted wY( ) and vY( ), where y is the confidence level of the estimate (e.g., also referred to as degree of confidence or confidence coefficient). A cellularity confidence interval (CI) [wY( ), vY( )], may, for example, indicate that, if multiple tumor samples were obtained, the estimated cellularity would be on the interval [e.g., at or between wY( ) and vY( )] y percent of the time. For example, in certain embodiments, a 90% CI lower and upper bound are determined. In certain embodiments, a 95% CI lower and upper bound are determined. In certain embodiments, a 68% CI lower and upper bound are determined. In certain embodiments, a cellularity confidence interval may be determined as p±ao_p. For example, for a 95% C. I. a = 1.96, for a 75% C. I a = 1.15, for a 50% C. I a = 0.67. The standard deviation of the cellularity can be calculated as √(1 — p) / N where N can be the total coverage at the given site or the number reads mapping to the wildtype allele plus the number reads mapping to alternate allele, e.g., after selecting only reads with a quality score above a predetermined threshold, e.g., in the range of 25 to 35, and potentially after filtering certain poor quality reads. In more preferred embodiments, given that p = VAF / Px, the standard deviation of the cellularity can be given by σVAF / Px, where σVAFis the standard deviation of the- 86 - 13247547v 1Attomey Docket No.: 2013237-0970observed variant allele frequency, given by √VAF(1 — VAF) / N, and Pxis the expected allele frequency of the given variant allele assuming the mutation is clonal.B. Tumor Deconvolution and Mutation Detection

[0411] Among other things, this application describes technologies for determining genetic and biophysical properties of tumor samples based on sequencing data (a procedure referred to, in certain cases, as “tumor deconvolution”).

[0412] As described in further detail herein, tumor deconvolution technologies of the present disclosure allow for complex characteristics of tumor genomes to be determined based on sequencing data. For example, cancer cells undergo extensive mutations, including copy number variation (CNV) events and single nucleotide variations (SNVs). As a result, unlike genomes extracted from normal cells, tumor genomes are not reliably diploid. Instead, the number maternal and paternal copies of genetic material varies from segment to segment, depending on the CNV events that took place over the lifetime of a given population of cancer cells. Additionally, cancer cells harbor collections of mutations at individual sites - SNV events - such as substitutions, insertions, deletions, etc.

[0413] Because mutations - in particular SNVs - that are found in tumor cells make prime targets for individualized cancer therapies (e.g., immunotherapies), the ability to accurate and rapidly detect and characterize them is key.B.i Tumor Modeling and Mutation Detection Challenges

[0414] Accurate detection and characterization of tumor mutations, however, is a highly complex process. Among other things, as illustrated in FIG.5, genetic and biophysical features of tumor samples may both impact particular tumor read counts and quantities, such as allele frequencies, that are derived therefrom and observed based on sequencing data.

[0415] For example, as explained above, tumor samples often include a fraction of normal, healthy cells. Read counts for particular mutations and / or portions of a tumor genome that are observed in sequencing data may be impacted by tumor sample purity. Accordingly, the certainty with which an event, such as a collection of reads with unique (e.g., abnormal) base - 87 - 13247547v 1Attomey Docket No.: 2013237-0970calls at a particular site, can be determined to indicate true underlying cancer cell mutations depends on tumor sample purity. Additionally or alternatively, CNV events, which may increase or decrease relative amounts of genetic material - and thus read counts - for particular portions (e.g., genes) of the tumor genome, may also impact how underlying, true, mutations manifest in observable sequencing data.

[0416] Accordingly, among other things, tumor deconvolution technologies of the present disclosure allow estimation of tumor sample purity and, in certain embodiments, characterization of CNV events across a tumor genome. As described in further detail herein, in certain embodiments, determining biophysical and genomic properties of a tumor sample in this manner can be used to improve accuracy with which cancer mutations are detected.

[0417] Additionally, or alternatively, in certain embodiments, mutations may be prioritized as targets for immunotherapy, for example to select high value targets for inclusion in personalized cancer vaccines, T-cell therapies, and the like, according to genomic characteristics, such as whether they are determined to be biologically clonal or subclonal, their zygosity, etc. Accordingly, ability to characterize genomic properties of mutations themselves and / or portions of a tumor genome where they are located (e.g., copy numbers of segments harboring mutations) can be highly valuable in the context of therapeutic approaches.

[0418] Determining biophysical properties of tumor samples, such as their purity, and characterizing genomic properties, such as copy number variations across a tumor genome, in non-trivial. Among other things, both tumor sample purity and CNV events affect how observed sequencing reads (e.g., such as impacting relative read counts of various segments in a tumor genome). This disclosure includes, accordingly, among other things, insight that estimating tumor sample purity and determining underlying genomic properties from sequencing data are, accordingly, coupled problems. Tumor deconvolution techniques described herein, accordingly, address these two problems in parallel, modeling and solving for tumor sample purity and CNV properties of a tumor genome, as well as their observed signatures in sequencing data, together, in certain embodiments in an iterative fashion.

[0419] Additionally, or alternatively, the present disclosure and techniques described herein appreciate the fact that tumor genomes vary widely from patient to patient and from indication to indication in terms of the frequency of mutations, the frequency and extent CNV - 88 - 13247547v 1Attomey Docket No.: 2013237-0970events, and the degree of subclonality of these genetic features. For example, the number of SNVs detected in an exome can vary across patients by 3 orders of magnitude, and across indications by at least 4 orders of magnitude from ~1 (in certain pediatric cancers) up to ~104(Alexandrov, L. B. et al., 2013, Nature 500, 415-421, Lawrence, M. S. et al., 2013, Nature 499, 214-218). Similar diversity is observed for CNVs. For example, breast cancers can contain anywhere from thousands of CNV events and other structural variation events to nearly none. Likewise, lung squamous cell tumors can contain anywhere from hundreds of CNV events and other structural variation events to none. On the other hand, indications like kidney renal clear cell carcinoma and medulloblastoma appear to contain significantly fewer structural variations (Yang, L. et al., 2013, Cell 153, 919-929, Network, C. G. A. R., 2013, Nature 499, 43-49, Parsons, D. W. et al., 2011, Science 331, 435-439). Adding to these difficulties is the fact that tumor sample purity is often, in practice, not high, and frequently even very low, leading to a reduction in signal to noise ratio of the genetic features sought to be estimated.

[0420] Tumor deconvolution and mutation characterization techniques of the present disclosure, accordingly, include approaches that allow for robust and accurate determination of sample purity and / or absolute copy numbers, for example via leveraging primary balanced heterozygous segments as described herein.

[0421] Additionally, or alternatively, coupling between the various tumor sample biophysical and genomic parameters, may, in certain embodiments, result in errors propagating from estimation of one parameter to another. That is, inaccuracies in estimating a value one particular parameter (e.g., tumor sample purity, certain absolute copy numbers, coupling constants) may impact more than the accuracy with which that particular parameter is estimated; they may impact accuracies with which values other parameters are estimated. Accordingly, among other things, tumor deconvolution technologies of the present disclosure include various multi-step and iterative refinement procedures, as well as quality control and error correction procedures, which serve to limit and / or correct inevitable inaccuracies in values determined via individual steps. Improved accuracy provided by approaches described herein is especially significant since small differences in an absolute copy number of a gene or mutation can have clinical implications.- 89 - 13247547v 1Attomey Docket No.: 2013237-0970

[0422] Among other things, the problem of purity estimation from somatic CNV events is generally considered a complicated problem because purity estimation is strongly coupled with the problem of estimation of absolute copy numbers of segments. Since typically the solution to the problem of purity estimation requires a joint estimation / optimization of many parameters, searching for an optimal solution of the many a priori unknown parameters required for CNV-based purity estimation means that the optimization is performed over a large parameter space. Such an optimization approach to solving the problem is not only computationally intensive and slow but can also be prone to false solutions due to local optima. A falsely estimated purity can have negative clinical consequences for patients by, for example, leading to detection of false mutations, leading to wrong predictions of absolute copy numbers of clinically actionable genes, and so on.

[0423] The large parameter space can be reduced by making certain assumptions regarding the tumor genome (e.g., assuming certain common precomputed tumor models, assuming certain recurrent cancer karyotypes). Such an approach could possibly be applicable for a method intended for cohort analysis for research purposes, wherein samples that violate such assumptions can be discarded, or wherein errors in sample analysis due to violation of these assumptions has no consequences. However, when the method is intended for clinical applications for the general population where it is nearly certain that in some percentage of cases the assumptions will be violated, a calculation method based on assumptions runs the risk of producing false predictions that can then have clinical implications and therefore should be avoided.

[0424] For example, an erroneous purity estimation may have negative clinical implications, since a wrong purity estimation can lead to wrong estimations of absolute copy numbers, on the basis of which clinical decisions can be made, or result in predicting wrong mutations (e.g., if the sensitivity of the mutation detection is increased too much due to an underestimated purity). One way of coping with a large search space is to assume certain common tumor models frequently observed for a given indication, for example assuming certain recurrent cancer karyotypes. Such an approach is highly problematic when the method is intended for clinical applications because when applying the method to the general population certain individuals are bound to violate assumptions, leading to false predictions for these- 90 - 13247547v 1Attorney Docket No.: 2013237-0970individuals, predictions that may have clinical implications if the treatment is based or relies on these predictions. Clinical implications in the case of treatment may include, for example, side effects and / or reduced efficacy of the treatment.

[0425] By using PBHSs to solve for the purity of the tumor sample based on CNV events such problems can be solved. By using PBHSs to solve for the purity of the tumor sample based on CNV events the dimensionality and complexity of the problem of estimating purity based on CNV events (CNV-based purity estimation) can be drastically reduced. This can be achieved by using PBHSs to estimate one or more auxiliary parameters (which is performed independently of solving the problem of purity estimation) that, once estimated, greatly simplifies the problem of CNV-based purity estimation. Auxiliary parameters can be parameters that may possibly not be of interest in and of themselves (i.e., they may not have biological and / or clinical significance) but can be used to estimate biologically meaningful and / or clinically significant parameters, such as the purity, absolute copy numbers of segments or SNVs, and so on.

[0426] Thus, accordingly, certain embodiments described herein can have the advantage that by basing the solution on PBHSs (which are isolated beforehand), the solution is decoupled to an estimation of a single continuous parameter (the purity) and determining one additional discrete parameter (the primary copy number). The complexity of the problem can therefore be dramatically reduced to a simple one-dimensional search for an optimal purity in combination with determination of the most plausible primary copy number from a discrete set of copy numbers.

[0427] Furthermore, because by construction, the primary copy number of PBHSs must be an even integer (because absolute copy numbers of BHSs must be even integers) the range of possible values for the primary copy number is halved (only even integers rather than all integers), wherein for the vast majority of tumors primary copy numbers will be either 2 or 4, and generally are not expected to exceed a value of 10. This drastic reduction in complexity can be made possible due to the fact that when PBHSs are isolated, certain auxiliary parameters that help to solve the problem can be directly estimated from the PBHSs.

[0428] PBHSs can essentially serve as an “anchor point”, wherein, although initially no parameter of interest required to solve the problem is known, having isolated this special group of segments allows to estimate certain auxiliary parameters that once estimated, reduces the - 91 - 13247547v 1Attomey Docket No.: 2013237-0970problem of purity estimation from CNV events to a one-dimensional search on a purity and determination of a primary copy number that can take on a small set of even integers.

[0429] In the context of purity estimation, accordingly, certain embodiments described herein can have the further advantage that by using PBHSs to solve for the purity of the tumor sample based on CNV events, the purity can be much more accurately and robustly solved across many tumor types of varying genomic complexity, degrees of subclonality (intratumor heterogeneity), tumor sampling conditions, and so on, including challenging tumor samples due to, for example, low purity, low coverage, high or low genetic complexity of the tumor genome, subclonality of CNV events, and so on. This is because isolation of the PBHSs is a very robust process in and of itself, which is in particular robust to the effects of low purity, low coverage, high or low genetic complexity of the tumor genome, subclonality of CNV events, and so on, factors that normally make the problem of purity estimation from CNV events highly challenging and error prone.

[0430] Moreover, the nature and extent of CNV events can greatly vary between cancer indications and among different individuals suffering from the same cancer type (see, for example, Yang et al., 2013, Cell 153:919-929.). Certain embodiments described herein have the further advantage that by using PBHSs to solve for the purity of the tumor sample based on CNV events the purity can be much more accurately and robustly solved across different indications, wherein each indication can have a different preponderance to the type, complexity, frequency, extent, spectrum of absolute copy numbers, etc., of CNV events, and hence the method may be applicable for high throughput automated personalized pipelines for analyzing tumor samples.

[0431] Certain embodiments described herein can have the further advantage that by using PBHSs to solve for the purity of the tumor sample based on CNV events the purity can be much more accurately and robustly solved for different individuals that suffer from the same cancer type, wherein each individual can have a different preponderance to the type, complexity, frequency, extent, spectrum of absolute copy numbers, etc., of CNV events. Hence, these embodiments can be suited for high throughput automated personalized pipelines for analyzing tumor samples.

[0432] Certain embodiments described herein can have the further advantage that by reducing the dimensionality of the problem to an estimation one continuous parameter (the - 92 - 13247547v 1Attorney Docket No.: 2013237-0970purity) and one discrete parameter (the primary copy number) with a small set of possible values, the execution time of the method is reduced, which is consequential if the method is intended for clinical applications, and in particular if it intended to be applied to masses of samples.

[0433] Certain embodiments described herein can reduce the dimensionality of the problem to an estimation of one continuous parameter (the purity) and one discrete parameter (the primary copy number) with a small set of possible values there are fewer solutions that need to be considered, and hence less chance for a false solution.

[0434] Since the dimensionality of the problem of purity estimation from CNV events is dramatically reduced by utilizing PBHSs, embodiments may not require making any assumptions about the tumor genome, assumptions that are likely to be violated when applying the method to the general population.

[0435] By being more accurate, more robust, computationally less intensive and faster, other genetic characteristics of the tumor genome that rely on purity estimation and determination of absolute copy numbers (such as estimation of mutations in the tumor genome and their full genetic characterization) can also be performed in a more accurate, more robust, computationally less intensive and faster way.B.ii Identifying and Using Subpopulations of Segments for Tumor Modeling

[0436] Turning to FIG.6A, in certain embodiments, a tumor modelling approach may, among other things, construct and fit one or more tumor models that accurately explain sequencing data - such as aligned reads for tumor and, optionally, normal genome for a patient. As shown in FIG.6A, in an example tumor modelling process 600, sequencing data may be obtained 602. As described herein, sequencing data 602 may comprise tumor sequencing data and / or normal sequencing data. In certain embodiments, sequencing data comprises tumor sequencing data and normal sequencing data (e.g., which may comprise replicates).

[0437] A tumor modelling process 600 may, at various steps, identify particular segments and / or particular classes or subsets of segments of a tumor genome 604. Among other things, certain subsets - e.g., types, classes - of segments have desired properties that makes them useful / appropriate for certain tumor models and determining values of particular parameters.- 93 - 13247547v 1Attomey Docket No.: 2013237-0970

[0438] For example, as shown in FIG.6B, in certain embodiments, heterozygous segments (HS) may be identified. In certain embodiments, balanced segments may be identified. In certain embodiments, balanced heterozygous segments may be identified. In certain embodiments, primary balanced heterozygous segments may be identified. In certain embodiments, primary segments may be identified.

[0439] As shown in FIG.6B and FIG.7, in certain embodiments tumor modeling technologies of the present disclosure include steps whereby a particular subset of segments are identified and used to facilitate downstream sample purity and / or copy number analysis. For example, certain identified subsets of segments, such as certain subpopulations of balanced heterozygous segments, can be used to compute approximate values of and / or initial estimates of certain model parameters that may, in turn, improve accuracy, efficiency, or make tractable downstream purity and / or copy number calculations.

[0440] In particular, in certain embodiments tumor modeling technologies described herein identify a particular subpopulation of balanced heterozygous segments - referred to herein as primary balanced heterozygous segments (PBHSs) and use them for subsequent tumor modeling and CNV-based purity estimations. Among other things, as described in further detail herein, PBHSs can be used to estimate auxiliary parameters which, in turn, may be used to facilitate tumor modeling, including estimation of copy numbers and / or sample purities.

[0441] Turning to FIG. 6B, in an example process 650, PBHSs are identified 652. Once the PBHSs are identified 652, they may be used to determine values of auxiliary parameters and / or initial estimates of certain parameters 654 that can be used to estimate sample purity and, in certain embodiments, a copy number of the subpopulation of PBHSs, referred to herein as a primary copy number 656.

[0442] For example, as described in further detail herein, in certain embodiments, methods and systems of the present disclosure search for a tumor model comprising an estimated sample purity and copy number assignments for each of at least a portion of segments (e.g., substantially all segments; e.g., a portion of segments meeting particular coverage and other sequencing data quality metrics) of the tumor genome that best fits observed sequencing data -e.g., a collection of tumor and normal reads and particular segments each read maps to. In certain embodiments, an estimated sample purity is varied (e.g., thereby varying a tumor model - 94 - 13247547v 1Attorney Docket No.: 2013237-0970based thereon) and tumor model quality assessed (e.g., quantified) to identify a best fit, for example via a maximum likelihood approach. Among other things, a set of identified primary balanced heterozygous segments and / or auxiliary parameters determined therefrom can be used to determine values of other tumor model parameters and / or initial estimates thereof, simplifying and / or making tractable a complex estimation procedure.

[0443] In certain embodiments, estimated sample purity 658a and / or primary copy number 658b are determined, and assessed via various quality control procedures 660. As described in further detail herein, quality control procedure may be used, for example, to determine if purity and / or copy number estimates are sufficiently high quality and, in certain embodiments, substitute bounds (e.g., an upper bound; e.g., a lower bound) for particular values.

[0444] In certain embodiments, sample purity estimates 658a and / or primary copy number 658b can be used to determine copy number values for other segments of the tumor genome, for example assigning copy numbers to substantially all segments (e.g., not just primary balanced heterozygous segments) 662 and providing a tumor model that comprises a substantially complete picture of the tumor’s genomic structure. In certain embodiments, sample purity and / or primary copy number values may be used for mutation detection and characterization 664.B.iii Selecting Primary Balanced Heterozygous Segments

[0445] FIGs. 7 and 8 illustrates example procedures for identifying primary balanced heterozygous segments (PBHSs) and using them to estimate auxiliary parameters which may, in turn, be used for tumor deconvolution (e.g., including purity estimation, such as CNV-based purity estimation).

[0446] As shown in FIG.7, in certain embodiments, a tumor genome (e.g., or model thereof) may be decomposed into balanced heterozygous segments (BHSs) and primary balanced heterozygous segments (PBHSs). FIG.7 illustrates 12 segments 700 in normal genome 702, and tumor genome 704. Above each representation of a segment in the normal genome is a number indicating the absolute number of copies of the segment in the normal genome 706. Likewise,- 95 - 13247547v 1Attomey Docket No.: 2013237-0970above each representation of a segment in the tumor genome is a number indicating the absolute number of copies of the segment in the tumor genome 708.

[0447] In the illustrative schematic shown in FIG. 7, three HSs, whose copies in tumor genome are indicated by 710, 711, 712, of the 11 HSs represented in the tumor genome 710, 711, 712, 713, 714, 715, 716, 717, 719, 720, 721 have an absolute copy number of two in the tumor genome, with one paternal copy of the segment and one maternal copy of the segment. The next four HSs represented in the tumor genome 713, 714, 715, 716 have an absolute copy number of four in the tumor genome, with each segment having two paternal copies and two maternal copies. The next heterozygous segment represented in the tumor genome 717 also has an absolute copy number of four in the tumor genome, but in this case the segment has one maternal copy and three paternal copies. The segment represented in the tumor genome 718 also has an absolute copy number of four in the tumor genome, however, in this case the segment does not contain heterozygous SNPs. The segment represented in the tumor genome 719 has two paternal copies. The segment represented in the tumor genome 720 has one paternal copy. The segment represented in the tumor genome 721 has an absolute copy number of three, comprising one maternal copy of the segment and three paternal copies of the segment.

[0448] The 12 segments depicted in the figure can be divided into three groups: primary segments 724, balanced heterozygous segments (BHSs) 726, and primary balanced heterozygous segments (PBHSs) 728. Primary segments, highlighted by the green box 724, are segments that the number of copies of which in the tumor genome is given by the primary copy number. In this example, the primary copy number is four. BHSs are highlighted by the pink box 726, wherein BHSs are HSs for which the number of copies of the maternal allele of the segment in the tumor genome equals the number of copies of the paternal allele of the segment in the tumor genome. Finally, PBHSs, highlighted by the blue box 728, are the subset of BHSs which have an absolute copy number in the tumor genome equal to the primary copy number.

[0449] In the illustrative schematic shown in FIG. 7, heterozygous SNPs corresponding to the HSs represented in the tumor genome 710, 711, 712, 713, 714, 715, 716 are balanced in the tumor genome. The remaining heterozygous SNPs 717, 719, 720, 721 are unbalanced in the tumor genome.- 96 - 13247547v 1Attorney Docket No.: 2013237-0970

[0450] The particular copy number variations from segment to segment and resultant subpopulations of segments that make up a tumor genome are not a priori known. Instead, they manifest characteristics of observed sequencing data, such as how many tumor reads map to each particular segment. Accordingly, tumor modeling technologies described herein aim to determine, based on sequencing data the underlying genomic properties of the cancer cells. Accordingly, among other things, the present disclosure includes the insight that a tumor genome to be determined can be decomposed into particular subpopulations of segments, for example based on segment heterozygosity and copy number symmetries, as shown in the schematic of FIG. 7. Moreover, this disclosure includes the further insight that decomposing a tumor genome into particular subpopulations in this way could prove useful for facilitating tumor modeling, e.g., to determine estimated sample purities and copy numbers of various segments and / or particular alleles within a tumor genome.B.iii.a Identifying Balanced Segments

[0451] Turning to FIG.8, tumor modeling technologies described herein may identify a set of PBHSs via process such as a step-wise process 800.

[0452] In certain embodiments, process 800 identifies and / or receives heterozygous SNPs (e.g., present in a subject’s normal genome) 802.

[0453] SNPs that are present in a normal genome of a subject (also referred to herein as “wild-type”) may be identified using normal sequencing data, for example via a commercial SNP or variant caller (e.g., Illumina DRAGEN, Qiagen Genomics, Torrent Variant Caller, etc.) based on sequence alignment to a reference and statistical modeling. Additionally, or alternatively, an initial list of SNPs may be determined via orthogonal methods, such as SNP arrays or TaqMan assays. Additionally, or alternatively, an initial list of SNPs may be obtained from a database, for example of potential SNPs present in an average individual of a particular population of which the subject is a member.

[0454] In certain embodiments, normal sequencing data may be used to evaluate and filter an initial list of SNPs and / or heterozygous SNPs for consistency with an expected underlying model of a heterozygous SNP, e.g., a normal diploid genome with, for a given SNP,- 97 - 13247547v 1Attorney Docket No.: 2013237-0970two alleles present in equal copy number. Accordingly, in certain embodiments, a balanced test may be used to determine whether sequencing data supports, with sufficiently high confidence, an underlying hypothesis that two alleles of a SNP are balanced - e.g., present in equal copy numbers. For example, statistical tests, such as a Z-statistic test, may be used to classify a particular SNP as balanced or unbalanced. SNPs or putative heterozygous SNPs in an initial list that do not meet a balanced test may be filtered and removed from the initial list of heterozygous SNPs.

[0455] In certain embodiments, heterozygous SNPs of the initial list may be assessed to determine whether sequencing data meets particular quality criteria. For example, each heterozygous SNP may be required to have a particular minimum coverage level in both normal and tumor sequencing data and / or a quality score above a particular minimum value.

[0456] In certain embodiments, heterozygous SNPs may be filtered according to the particular chromosome where they are located. For example, SNPs located on particular chromosomes may be excluded from analysis and removed from the list of heterozygous SNPs. In certain embodiments, only heterozygous SNPs located on particular chromosomes, such as chromosomes 1 through 22, are used.

[0457] In certain embodiments, a final list of heterozygous SNPs 804 produced in this manner may then be evaluated to identify SNPs that are balanced in the tumor genome and, in turn, a set of balanced heterozygous segments 806.

[0458] For example, in certain embodiments a balanced test as described above is performed for each SNP using tumor sequencing data, e.g., to determine whether a particular SNP is balanced in the tumor genome.

[0459] Accordingly, in certain embodiments, approaches described herein determine a balanced state of one or more heterozygous SNPs, for example in a tumor genome. In certain embodiments, a balanced state of a given heterozygous SNP is a value that indicates whether or not the given heterozygous SNP is balanced. For example, a balanced state may be a classification, such as “balanced” or “unbalanced”, a Boolean value (e.g., True or False), a numeric value (e.g., 1 or 0), etc.- 98 - 13247547v 1Attomey Docket No.: 2013237-0970

[0460] FIG. 7 shows an illustrative plot 730 of balanced states for SNPs located the 12 segments of the example tumor genome illustrated schematically in the figure. The balanced states shown in FIG. 7 are numeric values - 0 for unbalanced and 1 for balanced SNPs. For sake of simplicity, the figure only shows one SNP per segment, but it should be understood that, as described herein, a particular segment may have one, or more than one, heterozygous SNP. As shown in FIG. 7, in certain embodiments, balanced states are determined for heterozygous SNPs including SNPs within segments 719, 720, where a loss of heterozygosity event has occurred (since it may not be known a priori whether a LOH event has occurred). In certain embodiments, a balanced state is not determined for segments that do not comprise any heterozygous SNPs.

[0461] As illustrated in FIG. 7, statistical tests and / or other approaches for computing balanced states for SNPs may correctly reflect the true state of the tumor genome at a high level of accuracy, but errors may still occur. Plot 730 shows the balance state for each heterozygous SNP as a red dot, such as 732, which is an example of a heterozygous SNP deduced to be balanced. In this example, the balance state shown in the figure reflects the correct balanced state for all heterozygous SNPs except for the SNP 734, which shows an example of an error. The balance state indicated 734 is an example of an erroneous determination of the balanced property of the heterozygous SNP 720. Such errors can occur when using a statistical test, in particular when the purity of the tumor is low so that the distinction between the state of being balanced and the state of being unbalanced diminishes, and more so when the coverage of the given site is low.

[0462] In certain embodiments, balanced states of one or more heterozygous SNPs may be contour filtered. For example, as shown in FIG. 7, a contour filter to a profile of balance states 736 may be determined. Contour filtering can correct errors in the determination of the balance state of a heterozygous SNP by identifying regions in the genome that have a high density of balanced and / or unbalanced heterozygous SNPs. For example, contour filtering based on the balance state of neighboring heterozygous SNPs determined that the heterozygous SNP 734 is in a locally unbalanced region, and therefore the balance state of heterozygous SNP 734 can be corrected to a contour-filtered balance state of 1 (unbalanced), wherein heterozygous SNPs may be regarded as balanced if they are balanced both before and after contour filtering.- 99 - 13247547v 1Attomey Docket No.: 2013237-0970As a result, the heterozygous segment corresponding to 720 may be classified as unbalanced and hence discarded from further analysis. Accordingly, in certain embodiments, contour filtering can be used to discard HSs wrongly classified as BHSs, in particular when the purity is low and more such errors are expected.

[0463] In certain embodiments, balanced states of individual heterozygous SNPs can be used (e.g., aggregated) to determine segment-level balanced states, and thereby identify a set of balanced heterozygous segments 808. A particular segment may comprise one or more heterozygous SNPs. Accordingly, in certain embodiments, a balanced state of a particular segment may be determined based on balanced states of the one or more heterozygous SNPs that it comprises. For example, a balanced state for a particular segment may be determined as the mean (average), median, mode, etc. of the individual balanced states of the one or more heterozygous SNPs located within the particular segment. In certain embodiments, segmentlevel balanced states are (e.g., additionally or alternatively to individual SNP balanced states) contour filtered.

[0464] Accordingly, in certain embodiments, a set of balanced heterozygous segments may be identified by selecting segments with balanced states that satisfy certain criteria, such as having balanced states that are at and / or above a particular threshold value (e.g., greater than or equal to 0.5, greater than or equal to 0.6, greater than or equal to 0.7, greater than or equal to 0.8, greater than or equal to 0.9; e.g., equal to 1; e.g., greater than 0.5, greater than 0.6, greater than 0.7, greater than 0.8, greater than 0.9).B.iii.b Identifying Primary Balanced Heterozygous Segments

[0465] In certain embodiments, a particular subset of BHSs are identified. For example, as described herein within a tumor genome, various subpopulations of segments can be identified, for example based on whether or not they are balanced and / or their absolute copy number. Accordingly, tumor modeling approaches may identify and isolate a particular subpopulation of segments that (i) are balanced (heterozygous segments) and (ii) all have a same particular copy number.

[0466] For example, in certain embodiments, once a set of balanced heterozygous segments are identified 808, sequencing data corresponding to (e.g., tumor and / or normal reads - 100 - 13247547v 1Attorney Docket No.: 2013237-0970that map to) members of the set of balanced heterozygous segments can be used to identify one or more subpopulations thereof (e.g., of the set of balanced heterozygous segments), each subpopulation of BHSs comprising (e.g., only) segments having a same absolute copy number.

[0467] For example, as illustrated in FIG.7, in certain embodiments, tumor and normal read counts for each segment can be used to determine a probability density function (pdf) for values of tumor-to-normal read count ratios. For example, a tumor-to-normal read count ratio can be determined for a particular segment as a ratio of tumor read counts to normal read counts for the particular segment (e.g., tumor read count for the particular segment divided by the normal read count for the particular segment). Tumor-to-normal read count ratios may be determined for a plurality of segments, such as BHSs, to determine a distribution of tumor-to-normal read count ratios for the plurality of segments (e.g., set of BHSs) 810. A distribution may be represented via a pdf (e.g., via creation of a histogram, or other approaches) 812.

[0468] FIG. 7 shows a graph of an illustrative pdf 740 constructed based on the tumor-to-normal read count ratios for the BHSs. As illustrated in FIG.7, pdf 740 comprises multiple components (e.g., normal-like components), appearing as distinct peaks within the plot. Each component reflects a subpopulation of segments having a particular, different, absolute copy number.

[0469] Without wishing to be bound to any particular theory, FIG.7 illustrates the connection between the underlying CNVs from segment to segment and components of pdf 740.For example, BHSs 713, 714, 715, 716 each have an absolute copy number of four (4) and correspond to component 744 (e.g., centered around a higher mean tumor-to-normal read count ratio), while BHSs 710, 711, 712 each have an absolute copy number of two (2) and correspond to component 746 (e.g., centered around a lower mean tumor-to-normal read count ratio).

[0470] Accordingly, in certain embodiments, components of tumor-to-normal read count ratio distribution may be identified 814. For example, a decomposition 816 of empirical pdf 812 may be created, e.g., in which individual components are distinguishable and selectable. In certain embodiments, a particular one of the one or more components may be selected as a primary component 818 and used to determine a set of PBHSs and / or determine various auxiliary parameters as described herein.- 101 - 13247547v 1Attomey Docket No.: 2013237-0970

[0471] In certain embodiments, pdf 740 can be decomposed into different normal-like components 750 via a decomposition algorithm, such as a fit to a model pdf. For example, in certain embodiments a one-dimensional Gaussian mixture model (1D-GMM) can be used to approximate and be fit to pdf 740. Parameters determined via the fit - e.g., an amplitude, mean, and standard deviation of each constituent Gaussian - may, accordingly, be used to select the primary component.

[0472] In certain embodiments, an amplitude of each component of pdf 740 may be determined and the component having the largest amplitude selected as the primary component. In certain embodiments, a mean tumor- to-normal read count ratio of each component may be determined and used to select the primary component, e.g., alone or in conjunction with the component amplitudes. For example, in certain embodiments, a component having the lowest mean tumor-to-normal read count ratio may be identified, and its amplitude compared with the maximum amplitude (across all components). In certain embodiments, if an amplitude of the component having the lowest mean tumor-to-normal read count ratio is at least (e.g., at or above) a particular minimum faction of the maximum amplitude, then the lowest mean (n) component is selected as the primary component.

[0473] Among other things, selecting a particular subset of segments - e.g., BHSs - from which to construct tumor-to-normal read count pdf facilitates its decomposition into multiple (e.g., normal-like) components. Among other things, limiting the pdf to BHSs eliminates segments with odd copy numbers, so that the individual components (e.g., peaks) of the distribution are well separated. Accordingly, pdf is amendable to robust decomposition even for low purity tumor samples. Moreover, by selecting a particular subset of BHSs, the range of possible copy numbers of the selected subset (e.g., PBHSs) is restricted to even integers, thereby reducing the possible candidate primary copy numbers by a factor of two.

[0474] In certain embodiments, one or more auxiliary parameters may be determined 820a from the primary component 754 of pdf 740. In certain embodiments, a mean of the primary component, referred to herein as a “primary slope,” (denoted rpc) may be determined and used as an auxiliary parameter. In certain embodiments, a standard deviation of the primary component, referred to herein as the “primary standard deviation,” (denoted opc) may be determined. In certain embodiments, an amplitude of the primary component, denoted Ape, may - 102 - 13247547v 1Attomey Docket No.: 2013237-0970be determined. In certain embodiments, as described in further detail herein, values of one or more auxiliary parameters, including, but not limited to, a primary slope (rpc) and / or primary standard deviation (opc) may be used to determine a purity estimate (e.g., a SNV-based purity estimation and / or a CNV-based purity estimation).

[0475] Auxiliary parameters may include, as described in further detail herein, a primary allele frequency standard deviation, and a primary residual standard deviation.

[0476] A primary slope can be estimated as the mean the tumor over normal segment read count ratio of PBHSs, a primary standard deviation can be estimated as a standard deviation of the tumor over normal segment read count ratio of PBHSs, a primary allele frequency standard deviation can be estimated as a standard deviation of observed allele frequencies of the PBHSs, and primary residual standard deviation can be estimated as the standard deviation of residual errors of the PBHSs. The tumor over normal segment read count ratio for a given segment may be given by the ratio of a number of tumor reads that map to the given segment over a number of normal reads that map to the given segment.

[0477] In certain embodiments, one or more auxiliary parameters are used to define a two-dimensional Gaussian mixture model (2d-GMM) used for a maximum likelihood estimation of the purity based on CNV events.

[0478] In certain embodiments, a set of PBHSs are selected 760, 820b from the set BHSs. Parameters 752 characterizing each component of the BHS tumor- to-normal read count pdf 740, for example individual component amplitudes (Ai), means ( ), and standard deviations (<7i), may be used to select PBHS. For example, in certain embodiments, PBHSs may be selected from BHSs utilizing, for example, the principle of maximum probability based on the parameters 752 obtained from the decomposition of the pdf. For example, Example 1 describes an approach whereby a maximum a-posteriori probability (MAP) approach can be used to determine a likelihood of a particular heterozygous segment belonging to a particular one of the one or more components of pdf 740. Segments having a highest likelihood for belonging to the primary component may, accordingly, be selected for inclusion in the set of PBHSs.

[0479] In certain embodiments, after PBHSs are identified, the set of identified PBHSs 824 may be used to estimate additional auxiliary parameters and / or refine values of those determined (e.g., initially) based on decomposition and identification of primary component 826.- 103 - 13247547v 1Attomey Docket No.: 2013237-0970For example, in certain embodiments, the set of identified PBHSs 824 may be used to determine a primary allele frequency standard deviation 770. In certain embodiments, the set of identified PBHSs is used to determine a primary residual standard deviation 772.

[0480] For example, in certain embodiments, a primary allele frequency standard deviation may be determined as a standard deviation of observed allele frequencies across the set of PBHSs 760.

[0481] In certain embodiments, a residual error may be computed for a particular segment to compare (e.g., measure a measure of a difference between) (i) an observed tumor read count for the segment and (ii) an expected tumor read count for the segment predicted based on its normal read count. For example, for a particular, j-th segment, an observed tumor read count tj may be compared with a predicted tumor read count, tj, that is computed according to Eq. (1), below - e.g., as a linear prediction based on that segment’s observed normal read count.Eq. (1) tj = a • nj

[0482] The parameter a in Eq. (1) unknown but expected to depend on an absolute copy number of a given segment.

[0483] In certain embodiments, for the set of PBHSs, an initial estimate of the value of a can be approximated as the mean of the primary component - e.g., the primary slope, rpc (e.g., since rpc is the average tumor-to-normal read count ratio for the subpopulation of segments represented by the primary component).

[0484] In certain embodiments, a particular functional form of a residual error may be based on a variance stabilizing transform that is known or expected to produce a particular distribution of values. For example, in certain embodiments, a residual error may be determined based on observed and expected (e.g., a linear prediction of) tumor read counts accordingly to a variance stabilizing transform such that residual errors are normally distributed when computed across the set of PBHSs. For example, as demonstrated in Example 1, in certain embodiments, a difference of square root functions may be used as a variance stabilizing transform, with residual error computed as:- 104 - 13247547v 1Attomey Docket No.: 2013237-0970Eq. (2)

[0485] Other variance stabilizing transformations may be used, including, without limitation a logarithmic transformation (e.g., e = log (t) — log (an)), an arc-sine square root transformation (e.g., e = arcsin( t) — arcsin (Van) ), a reciprocal transformation (e.g., e = t-1— (an)-1), an exponential transformation (e.g., e = exp(t) — exp(an)), a box-Cox Transformation^. g.,, K ( ) = (KA / l — 1) / , for A 0; and log(Y), for = 0, where parameter X is chosen to best stabilize variance and approximate normality), an Anscombe transform, as well as other possible transformations.

[0486] In certain embodiments, residual errors are determined for each segment of the set of PBHSs, thereby determining a distribution of primary residual errors. A standard deviation may be determined, for example, by taking advantage of the expectation that the primary residual errors are normally distributed, by virtue of the variance stabilizing transformation. For example, in certain embodiments, a Gaussian fit to the distribution of primary residual errors may be used with, for example, the standard deviation extracted from a best fit (e.g., by minimizing a least squares error between an empirical probability distribution function of the primary residual error, q, and a Normal distribution with an unknown standard deviation).

[0487] In certain embodiments, a standard deviation of primary residual errors (e) 772 may be used, in turn, to identify, from all segments 780 (e.g., not just balanced heterozygous segments) a primary segment subpopulation 782 that comprises those segments (e.g., whether heterozygous or not, balanced or not) having an absolute copy number equal to the primary copy number. For example, in certain embodiments, this is performed by selecting segments for which the absolute of the primary residual standard deviation 780 is small in proportion to the primary residual standard deviation 772. In certain embodiments, since the primary residual errors 780 are normally distributed, selection of primary segments benefits from rigorous statistical tolerance.

[0488] Primary segments may be used to refine the set of PBHSs 760, which are further required for CNV-based purity estimation. For example, once primary segments have been- 105 - 13247547v 1Attomey Docket No.: 2013237-0970precisely estimated based on a rigorous statistical distribution, primary segments can be used to further refine the PBHSs by checking that each PBHS of the initially determined set (e.g., based on an MAP approach and parameters of a 1D-GMM decomposition) also appears in the set of primary segments 790. Refined PBHSs can be used to refine the primary slope 792 and / or the primary allele frequency standard deviation 794.

[0489] In certain embodiments, as described in further detail herein, the set of PBHSs and / or auxiliary parameter values determined therefrom may be used to determine various parameters used in tumor models, which are, in turn, fit to observed sequencing data 828. For example, in certain embodiments, auxiliary parameters estimated from PBHSs may be used to approximate (e.g., fixed) values of certain parameters in tumor models used for (e.g., CNV-based) purity estimation, e.g., thereby reducing a number of variable parameters in tumor models to be fit to observable distributions from sequencing data. In certain embodiments, auxiliary parameters estimated from PBHSs may be used as initial estimates of values of certain parameters in tumor models used for (e.g., CNV-based) purity estimation, e.g., so as to provide an accurate starting point for a fitting procedure. In certain embodiments, fitting procedures may refine a parameter value, updating it from its initial value taken from an auxiliary parameter. In certain embodiments, as described in further detail herein, tumor model parameters may be established relative to the primary copy number that characterizes the set of PBHSs, such as rather than estimate and / or fit multiple copy number values, a single, primary, copy number is determined via tumor mode fitting. Among other things, in this way, identifying PBHSs and determining auxiliary parameters therefrom simplifies and / or makes tractable complex tumor modeling procedures and allows for accurate purity estimates and absolute copy number determinations.B.iv Estimation of Sample Purity and Absolute Copy Number

[0490] Turning to FIG.9, in certain embodiments, tumor modeling technologies of the present disclosure leverage a set of identified PBHSs and / or various auxiliary parameters determined therefrom to estimate a purity of a tumor sample. In certain embodiments, additionally or alternatively, tumor modeling technologies described herein may determine a value of a primary copy number - e.g., the absolute copy number of primary segments. As described in further detail herein, once determined, the primary copy number may be used as a - 106 - 13247547v 1Attomey Docket No.: 2013237-0970point of reference from which absolute copy numbers of other segments (e.g., not just the primary segments) can be determined.

[0491] As illustrated in the schematic of FIG.9, procedures described herein for estimating sample purity may include a maximum likelihood procedure whereby one or more initial estimates of sample purity are determined 922, each corresponding to a candidate primary copy number 920. In certain embodiments, thereafter, a classification step 924 is included to select a particular one of the one or more candidate primary copy numbers and a final determined primary copy number for the tumor sample. A final estimated sample purity may then be determined as the initial purity estimate corresponding to the selected primary copy number 930a. In certain embodiments, as explained herein, quality control procedures 926 may be used to determine whether the final estimated sample purity may be retained 928a, and used for further analysis (e.g., mutation detection and / or genomically characterizing mutations), or if a bound (such as an upper bound) or other value (e.g., determined from an alternative pathway, such as a SNV-based purity estimate) should be used instead 930b.

[0492] As in further detail herein, results of tumor modeling may be used to (1) improve the sensitivity and accuracy of mutation detection and (2) genomically characterize mutations, including performing a clonality analysis.

[0493] FIG. 10 shows a schematic of an example procedure for tumor modeling, e.g., for improving mutation detection as well as characterizing mutations genomically. In the example schematic of FIG. 10, tumor modeling and its use is illustrated in a modular fashion, with CNV-based purity estimation, purity estimation logic, absolute copy number estimation, refinement of a list of putative SNVs, zygosity estimation and subclonality estimation each being performed by various sub steps. Among other things, FIG. 10 shows how an example tumor modeling and purity estimation procedure may include steps such as (i) maximum likelihood estimation of purity given a set of candidate primary copy numbers, (ii) selecting an optimal primary copy number with associated quality control, (iii) selection of optimal method for purity estimation, (iv) estimation of genome-wide absolute copy numbers with error control, (v) refinement of mutation detection using tumor modeling, (vi) maximum likelihood zygosity estimation; and (vii) clonality analysis.- 107 - 13247547v 1Attorney Docket No.: 2013237-0970B.iv.a Purity Estimation

[0494] In certain embodiments, tumor modelling approaches herein estimate a sample purity. For example, as shown in FIG. 10, in certain embodiments, a CNV-based purity estimation process may determine one or more tumor model fits based on a set of candidate primary copy numbers. Each candidate primary copy number of the set may be an even integer. Candidate primary copy numbers of the set may range from a minimum value, e.g., 2, up to a maximum primary copy number, e.g., 8. Maximum primary copy number may be a predetermined or preset value. For example, a set of candidate primary copy numbers may be {2, 4, 6, 8}.

[0495] In certain embodiments, a set or range of candidate purity values may also be determined. For example, candidate purity values may range from 0 to 1.

[0496] Next, likelihood values 1001 may be determined for various combinations of candidate primary copy numbers and candidate purity values based on sequencing data. In particular, in certain embodiments, tumor sequencing data reads and normal sequencing data reads mapping to heterozygous segments may be used to determine a likelihood of each candidate primary copy number and candidate purity combination.

[0497] In certain embodiments, a maximum likelihood estimation of purity is calculated 1004.

[0498] For example, in certain embodiments a node assignment likelihood function may be used to determine, for a particular individual segment, a likelihood of that segment having a particular absolute copy number and a particular allele- specific copy number - a combination referred to herein as a node. Node assignment likelihood function may be described in terms of a two-dimensional Gaussian mixture model (2d-GMM) where one dimension is the residual error coordinate 1006, and one coordinate is the differential allele frequency coordinate 1008.

[0499] In certain embodiments, to compute an overall model likelihood of a given purity, both coordinates are calculated for each of at least a portion of identified heterozygous segments 1002 given all assumed possible combinations of an absolute copy number of a given heterozygous segment and the allele- specific copy number of the heterozygous segment (wherein- 108 - 13247547v 1Attomey Docket No.: 2013237-0970each unique combination of absolute copy number and allele- specific copy number is referred to as a node). Both these coordinates have a normal or Normal-like distribution and therefore the 2d-GMM is appropriate. In certain embodiments, other M coordinates (e.g., M > 1) can be selected, as long as they have a normal or Normal-like distribution (e.g., by considering other variance stabilizing transformations for the tumor and normal reads).

[0500] Each node assignment likelihood function and, accordingly, overall model likelihood function, depends on the undetermined primary copy number, and a set of additional parameters 1010. In certain embodiments, a set of additional parameters can comprise, for example, previously estimated auxiliary parameters such as a primary slope, rpc, a primary residual standard deviation, < T(. and a primary allele frequency standard deviation, opobs. In certain embodiments, a set of additional parameters comprises an error rate per base, s. In certain embodiments, auxiliary parameters can be used to calculate residual errors (primary slope) and / or provide estimates for the standard deviations of a 2D-GMM used to model distributions of observable quantities that can be determined from sequencing data, such as a primary residual standard deviation and / or a primary allele frequency standard deviation. In certain embodiments, an error rate per base, s, is determined by quality score filtering of the tumor reads, wherein only tumor reads having a quality score equal to or above a predetermined quality score threshold are retained and setting the error rate per base to the error rate corresponding to the predetermined quality score threshold.

[0501] In certain embodiments, likelihood function can depend on additional parameters that can be refined by maximum likelihood. For example, likelihood function may depend on a coupling constant (0), whose value may be initially approximated as the primary slope, rpc, determined from a 1D-GMM decomposition as described herein, and then refined via a fitting procedure (e.g., that searches for values of 0 in the vicinity of the initial, rpc, guess that improve likelihood function). In certain embodiments, the likelihood function can be weighted, giving certain nodes a greater weight to facilitate more robust convergence.

[0502] Graph 1012 shows an example of how an overall model likelihood may vary as a function of candidate purities for various candidate primary copy numbers. The graph shows that in the given example an optimal purity can be determined for each candidate primary copy number 1022.- 109 - 13247547v 1Attorney Docket No.: 2013237-0970

[0503] In certain embodiments, ML-based purity estimation procure shown in FIG. 10 may be augmented by including and / or considering purity estimates obtained from various, e.g., alternative or additional methods. For example, as described in further detail in the PCT application entitled “Technologies for Determining Tumor Sample Purity and / or Genomic Characteristics Using Single Nucleotide Variations” filed January 23, 2026, the content of which is incorporated by reference herein in its entirety, at certain primary copy numbers (e.g., for the case where the primary copy number is assumed to be equal to 2) a purity estimation that leverages signal in sequencing data from SNV events may be used, e.g., and selected in addition or instead of CNV-based purity estimation, e.g., based on certain optimality criteria whether to select the CNV-based purity estimation or the SNV-based purity estimation, indicated by reference number 1020.B.iv.b Determining a Primary Copy Number and Node Assignments

[0504] In certain embodiments, (e.g., in a next step) tumor modelling technologies of the present disclosure determine a most suitable primary copy number based on tumor and normal reads. In certain embodiments, a most suitable primary copy number is determined using a multivariate classifier 1000. In certain embodiments, input to multivariate classifier 1000 comprises values of various test statistics, including, e.g., (1) a maximum likelihood score (ML score) 1032, (2) a K statistic 1034, which is a statistic based on the KL-divergence measured for both the residual error coordinate, and the allele frequency coordinate, and (3) a cluster density 1036, indicating a degree to which neighboring HSs on the tumor genome are assigned to the same node (e.g., in certain embodiments, a cluster density is a specific cluster density where nodes that are duplications of primary copy number 2 nodes are ignored).

[0505] In certain embodiments, based values of ML score 1032, K statistic 1034, and cluster density 1036, multivariate classifier may determine 1038 to reject or pass (e.g., not reject) a given candidate primary number. One or more candidate primary copy numbers may be passed and, accordingly, retained by multivariate classifier. Multivariate classifier may then select a particular one of the one or more retained candidate primary copy numbers as a final, optimal, primary copy number. In certain embodiments, multivariate classifier may determine a particular one of the one or more retained candidate primary copy numbers as a most plausible of the- 110 - 13247547v 1Attorney Docket No.: 2013237-0970retained primary copy numbers and select it as the optimal primary copy number 1038. In certain embodiments, multivariate classifier may abort the CNV-based purity estimation if, for example, the input parameters indicate that a plausible self-consistent solution cannot be determined.

[0506] Diagram 1042 illustrates an example multivariate classifier, showing input ML score and K statistic projected on a two-dimensional plane partitioned into different regions corresponding to possible outcomes of the classifier. Regions are labeled: pass (P), reject (R), or fail (F). In regions labeled with multiple possible outcomes a final classification is based on a value of specific cluster density input. The black dot in diagram 1042 represents a possible combination of the ML score and K statistic for a given candidate primary copy number.

[0507] In certain embodiments, multivariant classifier can be augmented and / or replaced by a neural network or other machine learning approach, which is trained on one or more of the following features: the maximum CNV-likelihood score (normalized or otherwise), the K statistic (normalized or otherwise), the specific cluster density, the f coordinate for each HS, the r coordinate for each HS, the node assignment for each HS, the genomic location of each HS

[0508] In certain embodiments, once a final primary copy number value is selected, corresponding MLE-determined tumor sample purity and coupling constant may be selected as well, such that a final set of tumor model parameters - namely, a primary copy number, tumor sample purity, and coupling constant - are determined.

[0509] In certain embodiments, final set of tumor model parameters may be used to assign each of at least a portion (e.g., substantially all) of heterozygous segments in a node - that is, a particular combination of an absolute copy number and an allele specific copy number. For example, for each of the portion of the heterozygous segments, a node assignment likelihood function may be evaluated using final set of tumor model parameters to determine a node assignment likelihood score for each of one or more potential nodes, and a given heterozygous segment assigned to the highest scoring node.

[0510] In certain embodiments, various quality assurance tests 1044 may be evaluated based on final tumor model parameters and / or node assignments determined therefrom to evaluate an overall quality of tumor deconvolution. In certain embodiments, based on overall- Ill - 13247547v 1Attorney Docket No.: 2013237-0970tumor deconvolution quality, quality assurance steps may be passed, or tumor deconvolution determined to be aborted.B.iv.c Purity Estimation Logic

[0511] In certain embodiments, tumor modeling technologies described herein include purity estimation logic to determine whether to retain CNV-based purity estimate determined as described above or to output a bound, such as a lower or upper bound on purity. In certain embodiments, purity estimation logic may evaluate whether to utilize a purity estimate obtained via another, additional or alternative method, such as a SNV-based purity estimate.

[0512] For example, when estimation of purity fails, it can, nevertheless, be desirable to estimate an upper bound on purity. An upper bound on purity can be used, for example, to improve the sensitivity and specificity of mutation detection, in particular when the tumor sample is highly impure, by taking into account a conservative estimate (upper bound) of the sample purity, when determining if a site encodes a SNV. For example, a likelihood ratio test comparing a model of a mutation to a model where no mutation exists can take such an upper bound on the purity into account. Estimation of the tumor sample purity based on CNV events using PBHSs can allow estimating an upper bound on purity in a highly robust and accurate fashion.

[0513] Estimation of the purity of the tumor sample based on CNV events using PBHSs can decouple the problem of purity estimation (which is a single continuous variable) from determination of the primary copy number (a parameter exclusively related to the PBHSs). The value of the primary copy number can be limited to a small discrete set of even integers that serve as candidate primary copy numbers (typically just 2 or 4, and generally not greater than 10). For each candidate primary copy number an optimal purity can be estimated. The upper bound on purity can preferably be taken to be the maximum of all estimated optimal purities. Therefore, by basing the solution on PBHSs the method can allow for a robust estimation of an upper bound on purity.

[0514] Because estimation of an optimal purity for any assumed primary copy number is a very robust process, the upper bound can be estimated robustly. Estimation of an upper bound- 112 - 13247547v 1Attorney Docket No.: 2013237-0970on purity in this manner also works well for difficult to analyze samples, wherein it is difficult to determine the correct primary copy number.

[0515] Because the primary copy number is limited to a small set of values, the method has the advantage that the upper bound can be a tight upper bound (close to the true purity).

[0516] Certain embodiments described herein can be used for high throughput automated personalized pipelines for analyzing tumor samples, wherein difficult to analyze samples are anticipated for a subset of subjects, and for these subjects estimation of a CNV-based upper bound on purity could be attainable, allowing more sensitive and specific detection of mutations for these subjects.

[0517] For example, as illustrated in FIG. 10, purity estimation logic may include a first decision branch point. In certain embodiments, possible results of first decision branch point may be a pass or initial abort decision. In certain embodiments, if a result of first decision branch point is a pass, purity estimation logic may evaluate certain criteria and, based thereon, select the CNV-based purity 1052 or another purity estimate, such as an SNV-based purity estimate 1054.In certain embodiments, otherwise, if purity estimation is aborted, it is determined whether to provide a SNV-based upper bound on purity 1056, a CNV-based upper bound on purity 1058, or to abort purity estimation 1059.

[0518] In certain embodiments, a CNV-based upper bound on purity 1058 can be determined by taking a maximum purity of the purities estimated for all candidate primary copy numbers. For example, in 1012 the CNV-based upper bound on purity would be 1 —B.v Characterizing Tumor Genomic Properties and Improving Mutation DetectionB.v.a Genome- Wide Absolute Copy Number Estimation

[0519] In certain embodiments, once (e.g., if) a purity estimate (1052 or 1054) is obtained, substantially all segments in a tumor genome can be assigned an absolute copy number using optimal purity and optimal primary copy number (e.g., and possibly additional optimized parameters) 1060.- 113 - 13247547v 1Attorney Docket No.: 2013237-0970

[0520] In certain embodiments, absolute copy number assignments may utilize a gender of the subject. In certain embodiments the subject’s gender is provided. In certain embodiments, the subject’s gender is estimated based on the number of reads mapping to genes in the X and Y chromosomes 1070.

[0521] Once absolute copy numbers have been assigned, the absolute copy number can be error corrected, e.g., using parity error correction or other methods. For example, potential errors in absolute copy number assignments may be detected and corrected (e.g., replaced with a most likely corrected estimate). In certain embodiments, if the segment is a heterozygous segment, then the allele- specific copy number can also be corrected.B.v.b SNV Refinement

[0522] In certain embodiments, tumor deconvolution processes and results thereof described herein may be used to improve mutation detection accuracy. For example, in certain embodiments, putative SNVs may be detected based on sequencing data. In certain embodiments, a list of putative SNVs 1082 may be filed using a mutation confidence score as described herein, for example in Example 1 (e.g., wherein only putative SNVs for which the mutation confidence score is above a predetermined threshold are retained for further analysis). As described in Example 1, mutation confidence score-based filtering may (e.g., initially) be performed without knowledge of tumor model parameters.

[0523] In certain embodiments, a refined mutation confidence score that does take into account values of tumor model parameters and / or upper and / or lower bounds thereof may be determined for each SNV, and used to (e.g., further) filter a list of putative SNVs 1082, thereby generating a refined list of SNVs 1084.

[0524] For example, in certain embodiments, a refined mutation confidence score may be determined for a particular (e.g., given, input) SNV based on an estimated purity (1052 or 1054) and / or an estimated upper bound on purity (1056 or 1058). In certain embodiments, if a particular tumor sample purity estimate is determined (e.g., as opposed to a bound thereon), then an absolute copy number of each site encoding a SNV may be determined based on the (e.g., error corrected) absolute copy number of the segment to which the given SNV is mapped. In certain embodiments, an absolute copy number of a SNV (e.g., which may be error corrected) - 114 - 13247547v 1Attorney Docket No.: 2013237-0970can be used to estimate a zygosity (e.g., an absolute copy number of the alternate allele) of a SNV 1086. In certain embodiments, absolute copy numbers and zygosities of SNV may be used to improve accuracy of refined mutation confidence scores 1088.

[0525] In certain embodiments, e.g., if a particular value of a tumor sample purity was estimated, subclonality (e.g., whether a given SNV is subclonal or clonal) of SNVs may be determined, e.g., using an estimated tumor sample purity along with an absolute copy number and zygosity of the SNV 1089.

[0526] Finally, a refined list of SNVs is selected for output via intersection with 1090.

[0527] FIG. 11 shows an example process whereby mutation detection may take into account tumor sample purity (or an upper bound of the tumor purity), as well as, in certain embodiments, an absolute copy number of the putative mutation, and / or an absolute copy number of the alternate allele of the putative mutation. In certain embodiments, the normal and tumor BAM files 1100 are provided to a front-end detector 1110. BAM files 1100 may, for example, be a result of exome sequencing. Front-end detector 1110 produces a list of putative mutations that may, in turn, be initially filtered via mutation confidence score filter 1120.Mutation confidence score filter may, for example, determine and filter mutations based on a mutation confidence score that does not account for a tumor sample purity and / or absolute copy numbers of various segments in the tumor genome. The resulting putative mutations may be provided to a purity / CNV estimation module 1130. Purity / CNV estimation module 1130 may also be provided the normal read counts per segment and the tumor read counts per segment 1140 extracted from the normal and tumor BAM files 1100, as well as a list of putative heterozygous SNPs 1150. Purity / CNV estimation module 1130 may then estimate tumor sample purity (e.g., or an upper bound on a tumor sample purity) 1160 either by CNV-based purity estimation or by SNV-based purity estimation. As described herein, CNV-based purity estimation may also be used to estimate absolute copy numbers of all segments in the sequenced tumor genome, which in turn may allow for estimation of the absolute copy number of the putative mutations 1160, and the absolute copy number of the alternate allele of the putative mutations 1160. One or more of these parameters 1160 can be used to calculate a refined mutation confidence score, which can then replace the mutation confidence score 1170 in the mutation confidence score filter 1120 to produce a refined list of putative mutations.- 115 - 13247547v 1Attorney Docket No.: 2013237-0970C. Example CNV-Based Tumor Deconvolution Procedures

[0528] FIGs. 12A-G provides an overview of the steps involved in a CNV-based purity estimation process, used in certain embodiments. In certain embodiments, as described herein, a tumor genome is first subdivided into balanced regions (FIG. 12A). Subdividing a tumor genome into balanced and unbalanced regions may make use of statistical tests that determine balanced states of individual heterozygous SNPs in a tumor genome and / or contour filtering procedures that evaluate or refine balanced states of SNPs in view of balanced states of neighboring SNPs, e.g., in a local vicinity. In this way, locally balanced and locally unbalanced regions can be identified. Based on this subdivision of the tumor genome, a set of balanced heterozygous segments (BHSs) may be identified, e.g., as those segments located within locally balanced regions of the tumor genome. For example, heterozygous segments (HS) may be identified in a normal genome, e.g., based on SNPs detected in normal sequencing data and / or identified in a database as described herein. Based on identified normal HSs, corresponding, nominally heterozygous, segments in a tumor genome may be identified. For each HS, allele frequencies of the corresponding tumor genome segments may be determined, along with a tumor- to-normal read count ratio. FIG. 12B plots these determined (tumor genome) allele frequencies (y-axis) and the tumor-to-normal read count ratios (x-axis) for each HS identified in a normal genome. BHSs are identified via green coloring.

[0529] In certain embodiments, (e.g., next) the distribution of tumor-to-normal read count ratios of BHSs (FIG. 12C) is decomposed into one or more (e.g., Normal-like) components. In certain embodiments, a primary (e.g., Normal-like) component (dashed curve), which is associated with a primary copy number and represents tumor-to-normal read count ratios of primary balanced heterozygous segments (PBHSs), may be identified. In FIG. 12A, PBHSs are shown as red dots. Among other things, without wishing to be bound to any particular theory, the process of isolating PBHSs is highly robust even at low tumor sample purities. Among other things, the procedure for subdividing the tumor genome into balanced regions is robust due to contour filtering approaches described herein. Additionally, or alternatively, by focusing only on balanced regions of the tumor genome, odd copy numbers are rejected, which leads the distribution of the tumor-to-normal segment read count ratio to contain- 116 - 13247547v 1Attorney Docket No.: 2013237-0970only a small number of spaced out (e.g., Normal-like) components corresponding to even copy numbers (FIG. 12C).

[0530] In certain embodiments once identified, PBHSs are then used to determine various parameters such as a primary slope (e.g., mean of the primary component), a primary standard deviation (e.g., standard deviation of the primary component), a primary residual standard deviation, and a primary allele frequency standard deviation. These parameters may, in turn, be used to first refine identification of PBHSs, such that a refined set of PBHS can be determined and subsequently used to further refine estimated values of these parameters (e.g., improving precision).

[0531] In certain embodiments, these refined parameters may be used to determine a two-dimensional likelihood function based on a residual error coordinate and an allele frequency coordinate. This likelihood function may be parametrized by a hypothesized primary copy number, a normal contamination ( / / ) and a tumor / normal coupling constant parameter (0). An example of this likelihood function for a specific primary number is shown in FIG. 12D. FIG.12E shows the likelihood function for each candidate primary copy number as a function of tumor sample purity. As described herein, in certain embodiments, likelihood function may be optimized parentheses (e.g., determining optimal values of sample purity and coupling constant) for each of a set of candidate primary copy numbers in this way multiple model fits are determined, one for each candidate primary copy number, with each model fit associated with an optimal value of p and 9.

[0532] In certain embodiments, a multivariate classifier is used to determine which of the candidate primary copy numbers is a correct solution (FIG. 12F). In certain embodiments, if no solution is plausible (1240), or the selected primary copy number is rejected (1250) by any of a series of quality control metrics (1230), then the estimated purity will be an upper bound taken over all, or a subset of, candidate primary copy numbers (1200), otherwise the method determines a tumor model, comprising an estimated purity, absolute copy numbers of segments, and allele specific copy number of (heterozygous) segments.

[0533] FIG. 12G shows an example of a cluster plot representing a tumor model solution. The tumor model associates each HS with a specific node, where each node represents a specific absolute copy number (x-axis) and allele specific copy number (y-axis) combination.- 117 - 13247547v 1Attorney Docket No.: 2013237-0970

[0534] In certain embodiments, a tumor deconvolution solution that passes all tumor deconvolution QC metrics (1230) provides a tumor model comprising of the purity of the tumor sample, absolute copy numbers of segments, and allele specific copy number of segments (1210). In certain embodiments, if the multivariate classifier failed to determine the correct primary copy number (1240), or the tumor deconvolution solution failed one or more tumor deconvolution QC metrics (1250), then the estimated purity will be an upper bound taken over all, or a subset of, candidate primary copy numbers (1200).

[0535] Subsequently the zygosity, cellularity and clone types of SNVs can be determined (1220). Further details regarding determining clone types for SNVs are described in U. S.Provisional Patent Application No. 63 / 749,339, entitled “Technologies for Neoantigen Prioritization Based on Tumor Modeling and Clonality States,” and the PCT application having the same title filed January 23, 2026, the content of each of which is incorporated by reference herein in its entirety.

[0536] FIG. 13 provides further detail regarding tumor deconvolution techniques and how they may be used in the context of mutation detection and characterization. In particular, in the example system shown in FIG. 13, input 1300 provided to tumor deconvolution techniques of the present disclosure may include a list of putative SNVs 1302, a list of putative heterozygous SNPs 1304, tumor and normal read counts per segment 1306 (e.g., number of reads mapping to each allele, e.g., along with quality scores per base per read and mapping quality scores per read), and potentially a gender of the patient 1307.

[0537] List of putative SNVs 1302 may be passed through a mutation confidence score filter 1303. Mutation confidence score filter may determine and filter mutations based on a mutation confidence score, which may be calculated without prior knowledge of tumor-related parameters such as the purity of the tumor sample or absolute copy numbers of segments. MCS filter may retain only putative SNVs having a MCS above a predetermined threshold.

[0538] Putative heterozygous SNPs 1304 may, in certain embodiments, be crossed against a SNP database 1308 such as dbSNP to remove false positives and / or filtered by selecting only heterozygous SNPs that are balanced in the normal genome 1308.

[0539] The input is then further processed by the purity / CNV module comprised of 1310, 1311, and 1312, wherein in 1310 PBHSs are selected, in 1312 the purity of the tumor sample is - 118 - 13247547v 1Attorney Docket No.: 2013237-0970estimated either using SNV-based purity estimation or CNV-based purity estimation using auxiliary parameters extracted from PBHSs, and in 1314 absolute copy numbers are assigned to segments and SNVs and additional genomic features of the SNVs are estimated.

[0540] The purity / CNV module identifies segments comprising at least one heterozygous SNP as heterozygous segments 1309. A balanced subset of HSs, BHSs, is selected 1316 by identifying those HSs whose heterozygous SNPs are balanced in the tumor genome. This step is assisted by a contour filter that identifies locally balanced regions in the tumor genome.

[0541] Selecting, and initially focusing on, BHSs is helpful since their absolute copy numbers must be even integers. This is manifest in observed empirical probability distribution functions (pdf) of the tumor (t) over normal (n) segment read count ratios, which are expected to contains only Normal-like components corresponding to even absolute copy numbers. In certain embodiments, a number of even components detected in the empirical pdf is estimated, for example, using a constrained one-dimensional Gaussian mixture model (1D-GMM) with a variable number of components. In certain embodiments, a primary component from the even components is selected (with an unknown even copy number, called the primary copy number CNPC= 2k, with k being an unknown integer). PBHSs are then selected to be the subset of BHSs for which the probability of being assigned to the primary component is greater than that for all other components (if present) 1318.

[0542] In certain embodiments, (e.g., next), two independent methods for purity estimation may be implemented in the purity / CNV module. In certain embodiments, one submodule 1320 may perform purity estimation based on SNVs (SNV-based purity estimation), and another submodule 1322 may perform purity estimation based on CNV events (CNV-based purity estimation). In certain embodiments, an SNV-based purity estimation approach 1320 can improve, under certain conditions, purity estimated by CNV-based purity estimation submodule 1322. Additionally, or alternatively, if there are few or no CNV events in the tumor genome, SNV-based purity estimation submodule 1320 can compensate for failure of CNV-based purity estimation submodule 1322.

[0543] In certain embodiments, SNV-based purity estimation submodule 1320 initially uses a primary residual standard deviation auxiliary parameter, ae(derived from PBHSs) to (e.g., rigorously) identify primary segments, and select from the SNVs, a subset of primary SNVs 724- 119 - 13247547v 1Attomey Docket No.: 2013237-0970that map to such segments. Further details regarding SNV-based purity estimation are described in the PCT application entitled “Technologies for Determining Tumor Sample Purity and / or Genomic Characteristics Using Single Nucleotide Variations” filed January 23, 2026, the content of which is hereby incorporated by reference herein in its entirety. In certain embodiments, various filters may be applied to select primary SNVs in balanced regions of the tumor genome 1326, resulting in a subset of minimal balanced SNVs. In certain embodiments, e.g., to avoid false mathematical solutions, SNV-based purity estimation may be applied to minimal balanced SNVs for hypotheses for which the primary copy number, CNPC, is 2.

[0544] In certain embodiments, SNV-based purity estimation uses a combined maximum likelihood approach, and a2distributed goodness-of-fit metric based on the Kullback-Leibler (KL) divergence to find an optimal pu...

Claims

Attomey Docket No.: 2013237-0970CLAIMSWhat is claimed is:

1. A method, the method comprising:(a) obtaining, by a processor of a computing device, tumor sequencing data for a tumor sample obtained from a subject;(b) identifying, by the processor, using the tumor sequencing data, a plurality of primary balanced heterozygous segments within a tumor genome for the subject, each primary balanced heterozygous segment having been determined to:(i) comprise one or more heterozygous SNPs, at least a portion of which are determined to be balanced within the tumor genome; and(ii) have an absolute copy number equal to a particular, unknown-but-to-be- determined, primary copy number;(c) determining, by the processor, based on the tumor sequencing data and the primary balanced heterozygous segments, an estimated tumor sample purity and / or bound thereon for the tumor sample; and(d) storing and / or providing, by the processor, the estimated tumor sample purity and / or bound thereon for display and / or further processing.

2. The method of claim 1, comprising using the estimated tumor sample purity and / or bound thereon to detect and / or prioritize a plurality of somatic mutations.

3. The method of claim 2, comprising using at least a portion of the detected and / or prioritized somatic mutations in a personalized immunotherapy4. The method of any one of claims 2-3, comprising using the plurality of somatic mutations and / or prioritization thereof to determine a cancer status for the subject.- 360 - 13247547v 1Attorney Docket No.: 2013237-09705. The method of any one of claims 2-4 comprising using the plurality of somatic mutations and / or prioritization thereof to select a therapy for the subject.

6. The method of any one of the preceding claims, comprising using the estimated tumor sample purity and / or bound thereon to determine a cancer status for the subject.

7. The method of any one of claims 2-4 comprising using the estimated tumor sample purity and / or bound thereon to select a therapy for the subject.

8. The method of any one of the preceding claims, comprising using the estimated tumor sample purity and / or bound thereon as a quality control.

9. The method of claim 8, comprising repeating sequencing of the tumor sample based on the estimated tumor sample purity and / or bound thereon.

10. The method of any one of the preceding claims, comprising:identifying, by the processor, a set of balanced heterozygous segments within the tumor genome for the subject, each balanced heterozygous segment of the set having been determined to comprise one or more heterozygous SNPs, at least a portion of which are balanced within the tumor genome; andselecting, by the processor, a subset of the set of balanced heterozygous segments as primary balanced heterozygous segments having a same particular absolute copy number, thereby identifying the plurality of primary balanced heterozygous segments.

11. The method of claim 10, comprising:- 361 - 13247547v 1Attomey Docket No.: 2013237-0970determining, by the processor, for each of at least a portion of the set of BHSs, a tumor-to-normal read count ratio based on the tumor sequencing data and the normal sequencing data, thereby determining a tumor-to-normal read count ratio distribution for the set of BHSs;identifying, by the processor, one or more components of the tumor-to-normal read count ratio distribution, each of the one or more subcomponents representing a subpopulation of BHSs having a particular absolute copy number, distinct from that of other subpopulations represented by other components of the tumor-to-normal read count ratio distribution; andselecting, by the processor, based on the identified components of the tumor-to-normal read count ratio distribution, the subset of primary balanced heterozygous segments (PBHSs).

12. The method of claim 11, wherein the method comprises obtaining, by the processor, normal sequencing data for a normal sample obtained from the subject, and wherein the measure of tumor read counts is a tumor-to-normal read count ratio determined, for a given BHS, based on a number of tumor reads mapping to the given BHS and a number of normal reads mapping to the given BHS.

13. The method of any one of claims 11-12, comprising:identifying, by the processor, a particular one of the one or more components as a primary component representing a largest subpopulation of the BHSs; andselecting, by the processor, BHSs determined to be members of the largest subpopulation represented by the primary component as for inclusion in the subset of PBHSs.

14. The method of claim 13, wherein identifying the primary component comprises: fitting, by the processor, a mixture model to the tumor-to-normal read count ratio distribution, thereby determining, based on the fit, a plurality of parameter values including, for each of the one or more components, a corresponding amplitude; andidentifying, by the processor, the primary component as having a largest corresponding amplitude.- 362 - 13247547v 1Attomey Docket No.: 2013237-097015. The method of any one of the preceding claims, comprising determining, by the processor, based on the plurality of primary balanced heterozygous segments, values of one or more auxiliary parameters.

16. The method of claim 15, wherein the one or more auxiliary parameters comprises a primary slope that measures a tumor-to-normal read count ratio for the plurality of primary balanced heterozygous segments.

17. The method of claim 15 or 16, wherein the one or more auxiliary parameters comprises a primary residual standard deviation that measures a standard deviation of residual errors for the plurality of primary balanced heterozygous segments.

18. The method of any one of claims 15-17, wherein the one or more auxiliary parameters comprises a primary allele frequency standard deviation that measures a standard deviation of allele frequency values determined for the plurality of primary balanced heterozygous segments.

19. The method of any one of claims 15-18, wherein step (d) comprises using the determined values of the one or more auxiliary parameters to determine the estimated sample purity.

20. The method of any one of the preceding claims, wherein step (d) comprises:determining, by the processor, for each of at least a portion of segments within the tumor genome, values for each of one or more sequencing data observations based on the tumor sequencing data and / or normal sequencing data, thereby determining a set of measured distributions comprising, for each of the one or more sequencing data observations, a corresponding distribution of values across the portion of tumor genome segments;- 363 - 13247547v 1Attomey Docket No.: 2013237-0970determining, by the processor, one or more tumor model fits based on the distributions corresponding to the one or more sequencing data observations and a tumor model, wherein determining each one or more tumor model fits comprises:generating, using the tumor model, a set of predicted distributions for the one or more sequencing data observations as a function of one or more parameters of the tumor model; and determining best- fit values for at least a portion of the one or more model parameters that optimize a model accuracy function that measures a degree to which the set of predicted distributions match the set of measured distributions,such that each tumor model fit is associated with (i) a set of best- fit values of at least a portion of the one or more model parameters and (ii) a model accuracy score corresponding to a value of the model accuracy function when the set of best fit values are used as values for the respective model parameters.

21. The method of claim 20, wherein the one or more parameters of the tumor model comprise a tumor sample purity, which is included in the portion of the one or more model parameters for which the best- fit values are determined by optimizing the model accuracy, such determining the one or more tumor model fits comprises determining one or more best-fit values of the tumor sample purity, each associated with a corresponding tumor model fit.

22. The method of claim 21, wherein step (d) comprises determining the estimated tumor sample purity based on the one or more best-fit values of the tumor sample purity.

23. The method of any one of claims 20-22,wherein the one or more parameters of the tumor model comprises the primary copy number andwherein determining the one or more tumor model fits comprises, for each candidate primary copy number value of a set of candidate values,- 364 - 13247547v 1Attomey Docket No.: 2013237-0970determining a corresponding tumor model fit by using the candidate primary copy number value for the primary copy number model parameter; anddetermining best-fit values that optimize the model accuracy function for a remainder of the one or more model parameters,such that each tumor model fit is associated with:(i) the corresponding candidate primary copy number value,(ii) the set of best-fit values for the remainder of the one or more model parameters, and(iii) the model accuracy score corresponding to the value of the model accuracy function when the corresponding candidate primary copy number value and the set of best- fit values for the remainder of the model parameters are used in the tumor model.

24. The method of any one of claims 20-23, wherein the one or more model parameters comprise a coupling constant that measures how a change in a value of the absolute copy number of the plurality of PBHSs impacts a change in observed tumor-to-normal read count ratios for the plurality of PBHSs.

25. The method of claim 24, comprising using a primary slope auxiliary parameter as an initial value for the coupling constant.

26. The method of any one of claims 20-25, wherein the one or more sequencing data observations comprise an observed allele frequency or a function thereof.

27. The method of claim 26, wherein the one or more sequencing data observations comprises a differential allele frequency that measures, for a given heterozygous segment, a- 365 - 13247547v 1Attomey Docket No.: 2013237-0970difference between (i) the observed allele frequency and (ii) a predicted allele frequency for the given heterozygous segment.

28. The method of any one of claims 20-27, wherein at least one of the one or more sequencing data observations is determined, for a given segment, as a function of tumor read counts for the given segment and / or normal read counts for the given segment.

29. The method of claim 28, wherein the one or more sequencing data observations comprises a tumor-to-normal read count ratio, determined, for a given segment, as a ratio of tumor read counts to normal read counts for the given segment.

30. The method of claim 28 or 29, wherein the one or more sequencing data observations comprises a residual error determined based on (i) tumor read counts for the given segment and (ii) normal read counts for the given segment.

31. The method of claim 30, wherein the residual error is determined based on a variance stabilizing transformation.

32. The method of any one of claims 30-31, wherein determining the one or more tumor model fits comprises, for at least a portion of the tumor model parameters, determining best-fit values that maximize a value of a model likelihood function that measures an accuracy with which the tumor model, when particular values of the model parameters are used, predicts or explains the distributions of the one or more sequencing data observations.

33. The method of any one of claims 30-32, wherein the tumor model predicts a set of grid points, each grid point representing a predicted allele frequency and a predicted tumor-to-read count ratio corresponding to a particular node comprising a specific combination of an absolute copy number value and an allele specific copy number.- 366 - 13247547v 1Attomey Docket No.: 2013237-097034. The method of claim 33, wherein the predicted set of grid points is a function of the one or more model parameters.

35. The method of any one of the preceding claims comprising determining, by the processor, a value of the primary copy number.

36. The method of claim 35, comprising using a multivariate classifier to determine the value of the primary copy number.

37. The method of claim 35 or 36 comprising:for each candidate primary copy number value of a set of potential primary copy number values, determining, by the processor, corresponding values for each of one or more test statistics; andselecting, by the processor, a particular one of the candidate primary copy number value as the value of the primary copy number, based on (i) the values of the one or more test statistics corresponding to each of the candidate primary copy number and (ii) the multivariate classifier.

38. The method of any one of claims 32-37, wherein the one or more test statistics comprises a maximum value of a model likelihood function that measures an accuracy with which the tumor model predicts or explains the distributions of the one or more sequencing data observations for a given candidate primary copy number value.

39. The method of claim 37 or 38, wherein in the one or more test statistics comprises a K-statistic that measures a difference between, for one or more of the sequencing data observations, the distribution determined based on the tumor and normal sequencing data and a predicted distribution- 367 - 13247547v 1Attomey Docket No.: 2013237-097040. The method of any one of claims 37 to 39, wherein the one or more test statistics comprises a cluster density.

41. The method of any one of the preceding claims, comprising assigning, by the processor, based at least in part on the estimated sample purity, absolute copy number values to each of at least a portion of segments within the tumor genome.

42. The method of claim 41, comprising, for each to-be-assigned segment of the portion of segments within the tumor genome:determining, by the processor, one or more node assignment likelihoods for the to-be-assigned segment, each node likelihood assignment likelihood measuring a likelihood of the particular segment having a specific absolute copy number (C / Vmut) and allele specific copy number (Cx) based at least in part on:(i) observed tumor reads counts for the particular heterozygous segment and / or a particular allele thereof and observed normal read counts for the particular heterozygous segment and / or the particular allele thereof; and(ii) the estimated tumor sample purity.

43. The method of claim 41 or 42, comprising:identifying, by the processor, a first heterozygous segment having been assigned an odd absolute copy number;determining, by the processor, the first heterozygous segment to be balanced in the tumor genome, thereby determining the first heterozygous segment to be balanced in the tumor genome yet having been assigned an odd absolute copy number; andresponsive to the determining the first heterozygous segment to be balanced in the tumor genome yet having been assigned an odd absolute copy number, updating, by the processor, the assigned odd absolute copy number to an even, error-corrected value.- 368 - 13247547v 1Attomey Docket No.: 2013237-097044. The method of claim 43, wherein the error corrected value to which the absolute copy number of the first heterozygous segment is updated is an absolute copy number represented by a nearest even node relative to the first heterozygous segment.

45. The method of either claim 43 or 44, comprising identifying one or more correctly assigned neighboring segments of the first heterozygous segment and determining the error-corrected value based on absolute copy numbers of the one or more correctly assigned neighboring segments.

46. The method of any one of claims 41-45, comprising:identifying, by the processor, a local genetic neighborhood of a first segment; and updating, by the processor, the assigned absolute copy number of the first segment based on absolute copy numbers of one or more segments within the local genetic neighborhood.

47. The method of any one of the preceding claims, comprising:evaluating one or more quality assurance criteria;determining a failure in purity estimation based on the one or more quality assurance criteria; andresponsive to the determined failure in purity estimation, determining (i) a bound on tumor sample purity and / or (ii) selecting an alternative estimate of tumor sample purity or bound therein, determined via an alternate method.

48. The method of any one of the preceding claims, comprising using the tumor sequencing data together with the estimated tumor sample purity to detect a plurality of SNVs within the tumor genome of the subject.- 369 - 13247547v 1Attomey Docket No.: 2013237-097049. The method of any one of the preceding claims, comprising:obtaining, by the processor, a list of candidate mutations identifying a plurality of mutations determined to be present in the tumor genome of the subject;identifying, by the processor, for each particular mutation of the list of candidate mutations, a corresponding segment within the tumor genome comprising the particular mutation;determining, by the processor, for each particular mutation of the list of candidate mutations, an absolute copy number of the corresponding segment comprising the particular mutation based at least in part on the estimated sample purity, thereby determining absolute copy numbers for the list of candidate mutations; andselecting, by the processor, a subset of the list of candidate mutations based at least in part on the determined absolute copy numbers for the list of candidate mutations.

50. The method of any one of the preceding claims, comprising:obtaining, by the processor, a list of candidate mutations identifying a plurality of mutations determined to be present in the tumor genome of the subject;identifying, by the processor, of the list of candidate mutations, one or more mutations as mapping to heterozygous segments and, for each particular mutation of the one or more mutations identified as mapping to heterozygous segments, identifying, a corresponding segment within the tumor genome comprising the particular mutation;determining, by the processor, for each particular mutation of the one or more mutations identified as mapping to heterozygous segments, an allele specific copy number and / or a fractional zygosity of the corresponding segment comprising the particular mutation based at least in part on the estimated sample purity, thereby determining allele specific copy numbers and / or fractional zygosities for the one or more mutations mapping to heterozygous segments; and- 370 - 13247547v 1Attorney Docket No.: 2013237-0970selecting, by the processor, a subset of the list of candidate mutations based at least in part on the determined allele specific copy numbers and / or fractional zygosities for the one or more mutations mapping to heterozygous segments.

51. The method of any one of the preceding claims, comprising:obtaining, by the processor a list of candidate mutations identifying a plurality of mutations determined to be present in the tumor genome of the subject;identifying, by the processor, for each particular mutation of the list of candidate mutations, a corresponding segment within the tumor genome comprising the particular mutation;determining, by the processor, for each particular mutation of the list of candidate mutations, an estimated cellularity and / or variability thereof based at least in part on the estimated sample purity, thereby determining estimated cellularity values and / or variabilities thereof for the list of candidate mutations; andselecting, by the processor, a subset of the list of candidate mutations based at least in part on the estimated cellularity values and / or variabilities thereof for the list of candidate mutations.

52. The method of any one of the preceding claims, comprising:obtaining, by the processor, a list of candidate mutations identifying a plurality of mutations determined to be present in the tumor genome of the subject;identifying, by the processor, for each particular mutation of the list of candidate mutations, a corresponding segment within the tumor genome comprising the particular mutation;determining, by the processor, for each particular mutation of the list of candidate mutations, an identification of the particular mutation as clonal or subclonal based at least in part on the estimated sample purity, thereby identifying clonal and / or subclonal mutations within the list of candidate mutations; and- 371 - 13247547v 1Attomey Docket No.: 2013237-0970selecting, by the processor, a subset of the list of candidate mutations based at least in part on the identified clonal and / or subclonal mutations within the list of candidate mutations.

53. The method of any one of the preceding claims, comprising causing, by the processor, display and / or rendering of a graphical user interface (GUI) and / or graphical report comprising a graphical representation of at least a portion of the tumor genome of the subject.

54. The method of claim 53, wherein the graphical representation of the portion of the tumor genome of the subject comprises one or more segment display tracks, each visually representing a plurality of segments within the portion of the tumor genome.

55. The method of claim 54, wherein the one or more segment display tracks visually represent each segment via a corresponding graphical element, wherein positions and / or orientations of the graphical elements visually represent locations of the corresponding segments within the tumor genome and / or the graphical elements are color-coded to visually represent a particular genomic property of the corresponding segment.

56. The method of any one of claims 52-55, wherein the graphical representation of the portion of the tumor genome comprises one or more mutation display tracks, each visually representing a plurality of detected mutations within the portion of the tumor genome.

57. The method of any one of the preceding claims, wherein the tumor sequencing data and / or normal sequencing data are whole genome sequencing data (WGS), whole exome sequencing data (WES), or single nucleotide polymorphism (SNP) array data.

58. The method of any one of the preceding claims, wherein the tumor sequencing data and / or normal sequencing data comprise a plurality of replicates.- 372 - 13247547v 1Attomey Docket No.: 2013237-097059. The method of claim 58, wherein the plurality of replicates are obtained from multiple samples of extracted gDNA from tumor and normal samples, obtained for multiple library preparations, obtained via multiple sequencing runs on a single library.

60. The method of any one of the preceding claims, wherein the tumor sample is a formalin-fixed paraffin embedded (FFPE) sample.

61. The method of any one of the preceding claims, wherein the tumor sample is a fresh frozen sample.

62. The method of any one of the preceding claims, wherein the tumor sample is obtained from a subject diagnosed as having or at risk of having cancer.

63. The method of any one of the preceding claims, wherein the tumor sample is obtained from a subject diagnosed as having or at risk of having a cancer associated with low tumor mutation burden.

64. The method of any one of the preceding claims, wherein the tumor sample is obtained from a subject diagnosed as having or at risk of having breast cancer, prostate cancer, pancreatic cancer, pediatric cancer, neuroblastoma, ovarian cancer, renal cell carcinoma, Merkel cell carcinoma, hematologic cancer, colorectal cancer, melanoma, head and neck squamous cell carcinoma, or non-small cell lung cancer.

65. The method of any one of the preceding claims, wherein the estimated tumor sample purity is less than about 0.3.- 373 - 13247547v 1Attomey Docket No.: 2013237-097066. The method of any one of the preceding claims, wherein the estimated tumor sample purity is at least about 0.10.

67. The method of any one of the preceding claims, comprising determining, by the processor, a value of the primary copy number and detecting, based on the value of the primary copy number, presence of a whole genome duplication (WGD) event in the tumor sample.

68. The method of claim 67 comprising detecting presence of the whole genome duplication event based at least in part of the value of the primary copy number having been determined to be at least four.

69. A method comprising:(a) obtaining, by a processor of a computing device, a list of nominally heterozygous segments within a tumor genome of a subject, and each segment of the list corresponding to a heterozygous segment within a normal genome of the subject;(b) for each nominally heterozygous segment of the list:obtaining, by the processor, a currently assigned absolute copy number representing a number of copies of the segment having been determined to be present in the tumor genome; andobtaining, by the processor, a balanced state indicating whether the segment is balanced, thereby determining, for the list of nominally heterozygous segments, balanced states of corresponding tumor genome segments;(c) identifying, by the processor, based on the balanced states of the list of nominally heterozygous segments, a subset of balanced heterozygous segments;(d) identifying, by the processor, within the subset of balanced heterozygous segments, one or more segments with odd- valued corresponding absolute copy numbers as potential copy number assignments errors;- 374 - 13247547v 1Attomey Docket No.: 2013237-0970(e) updating, by the processor, absolute copy numbers of the potential copy number assignment errors to an error-corrected value, wherein each error-corrected value is an even number; and(f) storing, by the processor, for display and / or further processing, the error corrected values.

70. The method of claim 69, wherein step (e) comprises, for at least one particular segment identified as a potential copy number assignment error:receiving and / or determining, by the processor, for the particular segment, a first node assignment likelihood corresponding to the currently assigned absolute copy number of the particular segment and a second node assignment likelihood corresponding to a second absolute copy number, wherein the second absolute copy number is even-valued and wherein the first and second node assignment likelihoods represent a likelihood of the particular segment having the currently assigned and second absolute copy number, respectively; andbased on the first and second node assignment likelihoods, updating, by the processor, the currently assigned absolute copy number by replacing it with the second absolute copy number.

71. The method of claim 69 or 70, wherein step (e) comprises, for at least one particular segment identified as a potential copy number assignment error:identifying, by the processor, within the list of nominally heterozygous segments, one or more genomic neighbor segments with correctly assigned copy numbers, the one or more genomic neighbors comprising (i) a nearest upstream segment and (ii) a nearest downstream segment; andupdating, by the processor, the currently assigned absolute copy number of the particular segment based at least in part on absolute copy numbers of the nearest upstream segment and the nearest downstream segment.

72. A method comprising:- 375 - 13247547v 1Attomey Docket No.: 2013237-0970(a) obtaining, by a processor of a computing device, a list of tumor genome segments and, for each segment of the list, a currently assigned absolute copy number;(b) for each particular segment of at least a portion of the list:identifying, by the processor, a plurality of genomic neighbor segments, each genomic neighbor segment a segment of the list having been determined to be located within a threshold number of segments and / or a threshold bases of the particular segment, and including the plurality of genomic neighbors together with the particular segment in a local genomic neighborhood;determining, by the processor, a representative copy number of the local genomic neighborhood based on absolute copy numbers of segments within the local genomic neighborhood; andupdating, by the processor, the currently assigned absolute copy number of the particular segment to a value determined based on the representative copy number of the local genomic neighborhood,thereby determining error-corrected copy numbers for at least the portion of the list; and(c) storing and / or providing, by the processor, for display and / or further processing, the error corrected copy numbers for the portion of the list of tumor genome segments.

73. A system comprising:a processor of a computing device;memory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to:(a) obtain tumor sequencing data for a tumor sample obtained from a subject; (b) identify, using the tumor sequencing data, a plurality of primary balanced heterozygous segments within a tumor genome for the subject, each primary balanced heterozygous segment having been determined to:- 376 - 13247547v 1Attorney Docket No.: 2013237-0970(i) comprise one or more heterozygous SNPs, at least a portion of which are determined to be balanced within the tumor genome; and(ii) have an absolute copy number equal to a particular, unknown-but- to-be-determined, primary copy number(c) determine, based on the tumor sequencing data and the primary balanced heterozygous segments, an estimated tumor sample purity and / or bound thereon for the tumor sample; and(d) store and / or provide the estimated tumor sample purity and / or bound thereon for display and / or further processing.

74. A system comprising:a processor of a computing device;memory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to:(a) obtain, a list of nominally heterozygous segments within a tumor genome of a subject, and each segment of the list corresponding to a heterozygous segment within a normal genome of the subject;(b) for each nominally heterozygous segment of the list:obtain a currently assigned absolute copy number representing a number of copies of the segment having been determined to be present in the tumor genome; andobtain a balanced state indicating whether the segment is balanced, thereby determining, for the list of nominally heterozygous segments, balanced states of corresponding tumor genome segments;(c) identify, based on the balanced states of the list of nominally heterozygous segments, a subset of balanced heterozygous segments;- 377 - 13247547v 1Attomey Docket No.: 2013237-0970(d) identify, within the subset of balanced heterozygous segments, one or more segments with odd- valued corresponding absolute copy numbers as potential copy number assignments errors;(e) update absolute copy numbers of the potential copy number assignment errors to an error-corrected value, wherein each error-corrected value is an even number; and(f) store and / or provide the error corrected values for display and / or further processing.

75. The system of claim 73, wherein at step (e), the instructions cause the processor to, for at least one particular segment identified as a potential copy number assignment error:receive and / or determine, for the particular segment, a first node assignment likelihood corresponding to the currently assigned absolute copy number of the particular segment and a second node assignment likelihood corresponding to a second absolute copy number, wherein the second absolute copy number is even-valued and wherein the first and second node assignment likelihoods represent a likelihood of the particular segment having the currently assigned and second absolute copy number, respectively; andbased on the first and second node assignment likelihoods, update the currently assigned absolute copy number by replacing it with the second absolute copy number.

76. The system of claim 74 or 75, wherein at step (e) the instructions cause the processor to, for at least one particular segment identified as a potential copy number assignment error:identify, within the list of nominally heterozygous segments, one or more genomic neighbor segments with correctly assigned copy numbers, the one or more genomic neighbors comprising (i) a nearest upstream segment and (ii) a nearest downstream segment; and update the currently assigned absolute copy number of the particular segment based at least in part on absolute copy numbers of the nearest upstream segment and the nearest downstream segment.- 378 - 13247547v 1Attomey Docket No.: 2013237-097076. A system comprising:a processor of a computing device;memory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to:(a) obtain a list of tumor genome segments and, for each segment of the list, a currently assigned absolute copy number;(b) for each particular segment of at least a portion of the list:identify a plurality of genomic neighbor segments, each genomic neighbor segment a segment of the list having been determined to be located within a threshold number of segments and / or a threshold bases of the particular segment, and including the plurality of genomic neighbors together with the particular segment in a local genomic neighborhood;determine a representative copy number of the local genomic neighborhood based on absolute copy numbers of segments within the local genomic neighborhood; andupdate the currently assigned absolute copy number of the particular segment to a value determined based on the representative copy number of the local genomic neighborhood,thereby determining error-corrected copy numbers for at least the portion of the list; and(c) store and / or provide the error corrected copy numbers for the portion of the list of tumor genome segments for display and / or further processing.

77. A method of producing an immunotherapy construct for a subject, the method comprising:detecting a plurality of candidate mutations in tumor cells from the subject using a method or system of any one of claims 1-76; and- 379 - 13247547v 1Attomey Docket No.: 2013237-0970synthesizing, as the immunotherapy construct, a polyribonucleotide encoding a plurality of neoepitopes, each encoded neoepitope corresponding to a candidate mutation of the plurality of candidate mutations.

78. A method comprising:determining a subset of T-cells within a population of T-cells that are capable of specifically binding a plurality of complexes,wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, andwherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer- specific mutations detected using the method or system of any one of claims 1-76; andenriching for the subset of T-cells that are capable of specifically binding the plurality of complexes.

79. A method comprising:administering to a subject a population of T-cells, wherein the population of T cells are capable of specifically binding a plurality of complexes,wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells,wherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer- specific mutations detected using the method or system of any one of claims 1-76, and- 380 - 13247547v 1Attorney Docket No.: 2013237-0970wherein the genome of at least some of the subject’s cells comprises a subset of the cancer-specific mutations.

80. A method comprising:determining a subset of TILs within a population of TILs that are capable of specifically binding a plurality of complexes,wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, andwherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer- specific mutations detected using the method or system of any one of claims 1-76; andenriching for the subset of TILs that are capable of specifically binding the plurality of complexes.

81. A method comprising:administering to a subject a population of TILs, wherein the population of TILs are capable of specifically binding a plurality of complexes,wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells,wherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer- specific mutations detected using the method or system of any one of claims 1-76, andwherein the genome of at least some of the subject’s cells comprises a subset of the cancer- specific mutations.- 381 - 13247547v 1Attomey Docket No.: 2013237-097082. The method of any one of claims 80 to 81, comprising obtaining a tumor sample from the subject and using the tumor sample to detect the plurality of candidate mutations.

83. The method of claim 82, comprising obtaining a normal sample from the subject and using the normal sample to detect the plurality of cancer mutations.

84. The method of claim 82 or 83, comprising sequencing the tumor and / or normal sample.

85. A pharmaceutical composition comprising a polyribonucleotide encoding a plurality of neoepitopes, wherein at least a portion of the plurality of neoepitopes are individualized neoepitopes corresponding to a somatic mutation from a population of candidate mutations detected in a tumor sample of a subject and selected for inclusion in the pharmaceutical composition using the method or system of any one of the preceding claims.

86. An individualized pharmaceutical composition comprising a polyribonucleotide encoding a plurality of neoepitopes for a particular subject, wherein at least a portion of the neoepitopes are individualized neoepitopes corresponding to a somatic mutation from a population of candidate mutations detected in a tumor sample from the particular subject and selected for inclusion in the pharmaceutical composition using the method or system of any one of the claims 1-76.

87. A population of T-cells,wherein the population of T-cells is capable of specifically binding to a plurality of complexes,wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major- 382 - 13247547v 1Attomey Docket No.: 2013237-0970histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, andwherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer- specific mutations using the method or system of any one of claims 1-76.

88. A T-cell capable of specifically binding to a complex, wherein the complex comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancerspecific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells,wherein the one or more cancer-specific mutations are detected using the method or system of any one of claims 1-76.

89. A population of T-cells,wherein the population of T-cells is capable of specifically binding to a plurality of complexes,wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, andwherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer- specific mutations detected using the method or system of any one of claims 1-76.

90. A T-cell capable of specifically binding to a complex, wherein the complex comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancerspecific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells,wherein the one or more cancer-specific mutations are within the selected subset of candidate mutations detected using the method or system of any one of claims 1-76.- 383 - 13247547v 1Attomey Docket No.: 2013237-097091. A T-cell receptor capable of specifically binding to a complex, wherein the complex comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells,wherein the one or more cancer-specific mutations are within the selected subset of candidate mutations detected using the method or system of any one of claims 1-76.

92. A chimeric antigen receptor (CAR) capable of specifically binding to a complex, wherein the complex comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells,wherein the one or more cancer-specific mutations are within the selected subset of candidate mutations detected using the method or system of any one of claims 1-76.

93. A method (e.g., for analyzing sequencing data to detect one or more potential mutations), the method comprising:(a) obtaining, by the processor, sequencing data for a biological sample obtained from a subject, the sequencing data comprising a plurality of reads, each read represents a partial nucleotide sequence of a sequenced genome associated with the sample (e.g., and each read aligned to a reference genome);(b) selecting, by the processor, a particular site within the sequenced genome and determining, for the particular site: (i) a corresponding coverage (e.g., a total number of reads mapping to the site), (ii) a corresponding number of reads mapping to a particular alternate allele, different from a wild-type allele, and (iii) an error rate;(c) determining, by the processor, for the particular selected site, a false positive probability representing a probability of observing at least the corresponding number of reads mapping to the particular alternate allele without a true underlying mutation being present in the sequenced genome, given the corresponding coverage and the error rate;- 384 - 13247547v 1Attomey Docket No.: 2013237-0970(d) determining, by the processor, for the particular selected site, a missed-detection probability representing a probability of no more than the corresponding number of reads mapping to the particular alternate allele with a true underlying mutation being present in the sequenced genome, given the corresponding coverage [e.g., and one or more of: an assumed absolute copy number of the particular selected site (e.g., of two), an assumed fractional zygosity (e.g., of 0.5), an assumed tumor content (e.g., of 0.75), and an assumed error rate (e.g., of 0)]; and(e) determining, by the processor, a false discovery rate based on the false positive probability, the missed-detection probability, and representative overall mutation likelihood for the sample;(d) determining, by the processor, the particular selected site to be a potential mutation based on the false positive probability, the missed-detection probability and the false discovery rate; and(e) responsive to determining the selected site to be a potential mutation, storing, by the processor, the selected site in a list of detected mutations and / or providing, by the processor, the selected site for further processing.

94. The method of claim 93, wherein the false positive probability is determined so as to account for (i) the corresponding number of reads mapping to the particular alternate allele and (ii) a distribution of the maximum of the three noise alleles (e.g., based on extreme event statistics).

95. The method of claim 93 or claim 94, wherein step (b) comprises, for each of a number of noise reads ranging from (i) the corresponding number for reads mapping to the particular alternate allele to (ii) the corresponding coverage, determining, by the processor, a plurality of possible configurations for distributing the number of noise reads across three possible non-wild-type alleles and, for each possible configuration, a corresponding probability.- 385 - 13247547v 1Attomey Docket No.: 2013237-097096. The method of any one of claims 93-95, wherein step (d) comprises comparing the false positive probability to a first threshold value [e.g., a predetermined threshold value (e.g., a value determined to reflect a per-exome error rate (e.g., one (e.g., 2x10-8), two, five, ten errors per exome)].

97. The method of claim 96, comprising:determining, by the processor, the false positive probability rate to be less than or equal to the first threshold value and,responsive to determining the false positive probability rate to be less than or equal to the first threshold value, determining, by the processor, the particular selected site to be a potential mutation.

98. The method of any one of claims 93-97, wherein step (d) comprises comparing the missed-detection probability rate to a second threshold value (e.g., 0.01, 0.02, 0.05, 0.1, etc.).

99. The method of claim 98, comprising:determining, by the processor, the missed-detection probability rate to be greater than the second threshold value;based at least in part on the determining the missed-detection probability rate to be greater than the second threshold value, determining, by the processor, the particular selected site to be a potential mutation.

100. The method of claim 99, comprising:responsive to the determining the missed-detection probability rate to be greater than the second threshold value, comparing the false discovery rate to a third threshold value (e.g., 1%, 5%, 10%, etc.).- 386 - 13247547v 1Attomey Docket No.: 2013237-0970101. The method of claim 100, comprising:determining, by the processor, the false discovery rate to be less than or equal to (e.g., less than) the third threshold value; andresponsive to the determining the false discovery rate less than or equal to (e.g., less than) the third threshold value, determining, by the processor, the particular selected site to be a potential mutation.

102. The method of any one of the preceding claims, wherein the sample is a tumor sample and the representative overall mutation likelihood is selected based on a particular cancer type for the tumor sample.

103. The method of any one of the preceding claims, wherein the representative overall mutation likelihood is a value determined to represent a particular number of mutations per exome [e.g., 1,000 SNVs per exome, 2,000 SNVs per exome, 5,000 SNVs per exome (e.g., 10-4), 10,000 SNVs per exome, etc.].

104. A system comprising a processor of a computing device; and memory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of claims 93-103.- 387 - 13247547v 1