Technologies for mutation detection using tumor modeling and feedback
The method addresses the ineffectiveness of current cancer therapies by refining mutation detection through tumor sequencing data analysis and purity estimation, enabling precise identification of cancer-specific mutations for personalized treatments.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- BIONTECH SE
- Filing Date
- 2026-01-23
- Publication Date
- 2026-07-30
AI Technical Summary
Current cancer therapies are often ineffective due to the molecular heterogeneity of tumors, lacking a single, broadly applicable treatment for cancer patients.
A method for detecting cancer-specific mutations using tumor sequencing data, involving the estimation of tumor sample purity to filter initial events and determine putative mutations, utilizing a processor to refine mutation confidence scores based on tumor model parameters and sequencing data.
Accurately identifies and characterizes cancer-specific mutations, such as SNVs, for personalized cancer therapies by filtering false positives and refining mutation detection, enhancing the precision of mutation detection technologies.
Smart Images

Figure EP2026051709_30072026_PF_FP_ABST
Abstract
Description
Attorney Docket. No.: 2013237-0970TECHNOLOGIES FOR MUTATION DETECTION USING TUMOR MODELING AND FEEDBACKCROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to and benefit of U.S. Provisional Application No.63 / 749,339, filed on January 24, 2025, the content of which is hereby incorporated by reference herein in its entirety.BACKGROUND
[0002] Cancer is a primary cause of mortality, accounting for 1 in 4 of all deaths.Despite recent advances in the field of cancer immunotherapy there remains no single, broadly applicable treatment. Molecular heterogeneity of tumors renders many therapies ineffective for cancer patients.SUMMARY
[0003] Presented herein are methods and systems that allow biophysical and / or genomic properties of tumor samples to be determined via analysis of sequencing data, obtained, for example, via next-generation sequencing of tumor samples obtained from a subject. Among other things, mutation detection technologies of the present disclosure leverage estimated tumor sample purity to filter a set of initial events for putative mutations (e.g., in an iterative fashion). Among other things, techniques described herein can be utilized to accurately detect and characterize cancer- specific mutations, such as single nucleotide variations (SNVs) that may serve as targets for personalized cancer therapies.
[0004] In some aspects, the present disclosure provides a method (e.g., a computer-implemented method) (e.g., for detecting one or more cancer mutations for a subject) comprising: (a) obtaining, by a processor of a computing device, tumor sequencing data for a tumor sample obtained from a subject, the tumor sequencing data representing a tumor genome associated with the tumor sample [e.g., the tumor sequencing data comprising a plurality of tumor reads, each tumor read representing a partial nucleotide sequence of the tumor genome as determined by sequencing the tumor sample]; (b) detecting, by the processor, using on the tumor sequencing data, a plurality of initial events, each initial event corresponding to a variant portion of the tumor genome (e.g., identifying one or more sites within the tumor genome), wherein, for - 1 - 13241894vlAttorney Docket. No.: 2013237-0970each initial event, the corresponding variant portion is determined, based on the sequencing data, to (e.g., potentially) have a different copy number and / or a different nucleotide sequence relative to a corresponding (e.g., aligned with the corresponding variant portion) normal reference [e.g., wherein the normal reference is obtained from a database (e.g., a hl9 reference genome); e.g., wherein the normal reference is determined based on normal sequencing data obtained for the subject (e.g., by sequencing a normal sample, e.g., of healthy tissue and / or cells)]; (c) determining, by the processor, an estimated purity for the tumor sample and / or a bound (e.g., an upper bound; e.g., a lower bound) thereon (e.g., based on the tumor sequencing data); (d) filtering, by the processor, the plurality of initial events to exclude initial events identified as false-positives based on (i) the tumor sequencing data and (ii) the estimated purity for the tumor sample and / or bound thereon and retaining other initial events (e.g., not identified as false-positives and excluded) as putative mutations for inclusion in a mutation list; and (e) storing and / or providing, for display and / or further processing, the mutation list.
[0005] In some embodiments, for each of at least a portion of the plurality of initial events, the corresponding variant portion is determined, based on the sequencing data, to (e.g., potentially) have a different nucleotide sequence relative to the corresponding normal reference.
[0006] In some embodiments, the different nucleotide sequence is or comprises an alternate allele (e.g., an individual nucleotide), different from an allele of the corresponding normal reference (e.g., a wild-type allele).
[0007] In some embodiments, the different nucleotide sequence is or comprises an insertion and / or a deletion (e.g., of one or more nucleotides) relative to the corresponding normal reference (e.g., an indel).
[0008] In some embodiments, the variant portion is or comprises a structural variation relative to the normal reference.
[0009] In some embodiments, for each of at least a portion of the plurality of initial events, the corresponding variant portion is determined, based on the sequencing data, to (e.g., potentially) have a different copy number than the corresponding normal reference.- 2 - 13241894vlAttorney Docket. No.: 2013237-0970
[0010] In some embodiments, for each of at least a portion of the plurality of initial events, the corresponding variant portion is determined, based on the sequencing data, to (e.g., potentially) have a different nucleotide sequence relative to the corresponding normal reference.
[0011] In some embodiments, the different nucleotide sequence is or comprises an alternate allele (e.g., an individual nucleotide), different from a normal allele of the corresponding normal reference.
[0012] In some embodiments, step (d) comprises, for each initial event of at least a portion of the initial events: determining, by the processor, a mutation confidence score that quantifies (i) a likelihood of the initial event encoding the alternate allele and thereby representing a true underlying mutation relative to (ii) a likelihood of the initial event encoding the normal allele, wherein the mutation confidence score for a given initial event is determined based at least in part on (i) reads of the tumor sequencing data that map to the given initial event and (ii) a set of tumor model parameters; and identifying, by the processor, the initial event as (i) a false positive to be excluded from the mutation list or (ii) a putative mutation for inclusion in the mutation list based on the mutation confidence score determined for the initial event in comparison with a (e.g., predetermined) mutation confidence score threshold value.
[0013] In some embodiments, the set of tumor model parameters comprises a measured tumor content of the tumor sample (e.g., determined via histology, such as histological tumor content).
[0014] In some embodiments, the set of tumor model parameters comprises one or more predetermined purity values (e.g., a predetermined purity value obtained from a database; e.g., a predetermined range of purity values) and / or one or more predetermined copy numbers.
[0015] In some embodiments, the set of tumor model parameters comprises the estimated purity and / or bound thereon for the tumor sample, such that the mutation confidence score is a refined mutation confidence score (e.g., determined based on an estimated purity specific to tumor sample).
[0016] In some embodiments, the mutation confidence score is a ratio of the likelihood of the initial event encoding the alternate allele to the likelihood of the initial event encoding the normal allele.- 3 - 13241894vlAttorney Docket. No.: 2013237-0970
[0017] In some embodiments, the refined mutation confidence score is a difference of a log-likelihood of the initial event encoding the alternate allele and a log-likelihood of the initial event encoding the normal allele.
[0018] In some embodiments, provided methods comprise prior to step (c), filtering the plurality of initial events by, for each initial event of at least a portion of the initial events: determining, by the processor, a mutation confidence score that quantifies (i) a likelihood of the initial event encoding the alternate allele and thereby representing a true underlying mutation relative to (ii) a likelihood of the initial event encoding the normal allele, wherein the mutation confidence score is determined, for a given initial event, based at least in part on (i) reads of the tumor sequencing data that map to the given initial event and (ii) a set of tumor model parameters, and wherein the set of tumor model parameters comprises a predetermined set of one or more purity values and / or range of purity values; and identifying, by the processor, the initial event as a false positive to be excluded from a preliminary mutation list or identifying, by the processor, the initial event as a putative mutation for inclusion in the preliminary mutation list based on the mutation confidence score determined for the initial event in comparison with a (e.g., predetermined) mutation confidence score threshold value, thereby determining the preliminary mutation list, comprising a set of initial events retained following filtering according to their mutation confidence scores; and at step (c), using the preliminary mutation list to determine the estimated purity for the tumor sample and / or the bound thereon.
[0019] In some embodiments, step (d) comprises, following step (c): updating the set of tumor model parameters according to the estimated tumor sample purity and / or bound thereon (e.g., to replace, in the set of tumor model parameters, the predetermined set of one or more purity values and / or range of purity values with the estimated tumor sample purity and / or bound thereon), thereby determining a refined set of tumor model parameters; and for each initial event identified as a putative mutation and included in the preliminary mutation list: determining, by the processor, a refined mutation confidence score that quantifies (i) a likelihood of the initial event encoding the alternate allele and thereby representing a true underlying mutation relative to (ii) a likelihood of the initial event encoding the wild-type allele, wherein the refined mutation confidence score is determined, for a given initial event, based at least in part on (i) reads of the tumor sequencing data that map to the given initial event and (ii) the refined set of tumor model- 4 - 13241894vlAttorney Docket. No.: 2013237-0970parameters; and identifying, by the processor, the initial event as a false positive to be excluded from the mutation list or retaining, by the processor, the identification of the initial event as a putative mutation for inclusion in the mutation list based on the refined mutation confidence score determined for the initial event in comparison with a (e.g., predetermined) refined mutation confidence score threshold value.
[0020] In some embodiments, the refined mutation confidence score threshold value is lower than the mutation confidence score threshold value.
[0021] In some embodiments, step (d) comprises, for each initial event of at least a portion of the initial events: determining, by the processor, a value of a test statistic for the initial event based at least in part on the tumor sequencing data; determining, by the processor, a mutation confidence score that quantifies (i) a likelihood of the initial event encoding the alternate allele and thereby representing a true underlying mutation relative to (ii) a likelihood of the initial event encoding the wild-type allele, wherein the mutation confidence score is determined, for a given initial event, based at least in part on (i) reads of the tumor sequencing data that map to the given initial event and (ii) a set of tumor model parameters; selecting, by the processor, a value of a discrimination threshold (e.g., P-value significance threshold, thresholds on computed or measured quantities, etc.) based at least in part on the mutation confidence score determined for the initial event; and identifying, by the processor, the initial event as (i) a false positive to be excluded from the mutation list or (ii) a putative mutation for inclusion in the mutation list based on the value of the test statistic determined for the initial event and the selected value of the discrimination threshold.
[0022] In some embodiments, the set of tumor model parameters comprises a measured tumor content of the tumor sample (e.g., determined via histology, such as histological tumor content).
[0023] In some embodiments, the set of tumor model parameters comprises one or more predetermined purity values (e.g., a predetermined purity value obtained from a database; e.g., a predetermined range of purity values) and / or one or more predetermined copy numbers.
[0024] In some embodiments, the set of tumor model parameters comprises the estimated purity and / or bound thereon for the tumor sample, such that the mutation confidence score is a- 5 - 13241894vlAttorney Docket. No.: 2013237-0970refined mutation confidence score (e.g., determined based on an estimated purity specific to tumor sample).
[0025] In some embodiments, selecting the value of the discrimination threshold comprises selecting one of a plurality of threshold value settings, each threshold value setting associated with a range of refined mutation confidence scores, and lower threshold value settings are associated with higher- valued ranges of refined mutation confidence scores.
[0026] In some embodiments, selecting the value of the discrimination threshold comprises adjusting a threshold value setting to reduce the value of the discrimination threshold as the mutation confidence score increases or to increase the value of the discrimination threshold as the refined mutation confidence score decreases.
[0027] In some embodiments, step (d) comprises: selecting (e.g., adjusting), by the processor, value(s) of one or more discrimination thresholds based on the estimated purity and / or bound thereon; and using, by the processor, the selected value(s) of the one or more discrimination thresholds to filter the plurality of initial events.
[0028] In some embodiments, step (d) comprises, for each initial event of at least a portion of the initial events: determining, by the processor, a value of a test statistic for the initial event based at least in part on the tumor sequencing data; selecting (e.g., adjusting), by the processor, a value of a discrimination threshold based at least in part on the estimated purity for the tumor sample and / or the bound thereon; and identifying, by the processor, the initial event as (i) a false positive to be excluded from the mutation list or (ii) a putative mutation for inclusion in the mutation list based on the value of the test statistic determined for the initial event and the selected value of the discrimination threshold.
[0029] In some embodiments, provided methods comprise selecting the value of the discrimination threshold from one of a plurality of threshold settings, wherein each of the plurality of threshold settings is associated with a range of tumor sample purity values and a corresponding value for the discrimination threshold, and wherein threshold settings associated with ranges of lower tumor sample purities have lower corresponding values for the discrimination threshold than threshold settings associated with ranged of higher tumor sample purities (e.g., such that the discrimination threshold is relaxed at lower purities).- 6 - 13241894vlAttorney Docket. No.: 2013237-0970
[0030] In some embodiments, provided methods comprise selecting the value of the discrimination threshold comprises reducing the value of the discrimination threshold as the estimated tumor sample purity decreases or increasing the value of the discrimination threshold as the estimated tumor sample purity increases (e.g., such that the discrimination threshold is relaxed at lower purities).
[0031] In some embodiments, the test statistic is a mutation confidence score and / or a refined mutation confidence score.
[0032] In some embodiments, the test statistic is determined based on (e.g., is a function of) the estimated tumor sample purity and / or bound thereon.
[0033] In some embodiments, the test statistic is determined based on (e.g., is a function of) a ploidy of the tumor sample.
[0034] In some embodiments, the test statistic is determined based on (e.g., is a function of) a measured histological tumor content of the tumor sample.
[0035] In some embodiments, the value of the discrimination threshold is selected based (e.g., further) at least in part on a ploidy of the tumor sample.
[0036] In some embodiments, the value of the discrimination threshold is selected based (e.g., further) at least in part on a value of a secondary test statistic, said secondary test statistic determined as a function of a tumor-model derived parameter.
[0037] In some embodiments, step (d) comprises: selecting, by the processor, a particular false positive filter of a plurality of false positive filters based at least in part on the estimated tumor sample purity; and identifying, by the processor, the initial event as (i) a false positive to be excluded from the mutation list or (ii) a putative mutation for inclusion in the mutation list, using the particular selected false positive filter.
[0038] In some embodiments, each false positive filter of the plurality of false positive filters is associated with a range of tumor sample purities.
[0039] In some embodiments, the plurality of false positive filters comprises a series of filters associated with ranges corresponding progressively higher tumor sample purity vales and wherein a stringency of filters of the series increases with increasing tumor sample purity.- 7 - 13241894vlAttorney Docket. No.: 2013237-0970
[0040] In some embodiments, provided methods comprise determining, by the processor, the estimated tumor sample purity to be within the range of tumor sample purities associated with the selected false positive filter, thereby selecting the false positive filter associated with the range of tumor sample purities that includes the estimated tumor sample purity.
[0041] In some embodiments, provided methods comprise determining, by the processor, a value of a tumor sample purity-derived parameter (e.g., a parameter determined as function of tumor sample purity) based on the estimated tumor sample purity; and selecting, by the processor, the particular false positive filter of the plurality of false positive filters based on the value of the tumor- sample purity-derived parameter.
[0042] In some embodiments, each false positive filter of the plurality of false positive filters is associated with a range of values for the tumor sample purity-derived parameter.
[0043] In some embodiments, provided methods comprise determining, by the processor, the tumor sample purity-derived parameter to be within the range of tumor sample purities associated with the particular selected false positive filter, thereby selecting the false positive filter associated with the range of tumor sample purities that includes the tumor sample purity-derived parameter.
[0044] In some embodiments, the tumor sample purity-derived parameter is a mutation confidence score (e.g., a refined mutation confidence score) that quantifies (i) a likelihood of the initial event encoding the alternate allele and thereby representing a true underlying mutation relative to (ii) a likelihood of the initial event encoding the wild-type allele, wherein the mutation confidence score (e.g., refined mutation confidence score) for a given initial event is determined based at least in part on (i) reads of the tumor sequencing data that map to the given initial event and (ii) a set of one or more tumor model parameters.
[0045] In some embodiments, the plurality of false positive filters comprises a series of filters of differing levels of stringency, and wherein a stringency of filters of the series increases with decreasing values of the mutation confidence score.
[0046] In some embodiments, step (d) comprises using a strand bias filter that measures a likelihood of strand bias for a given initial event based on a number of (e.g., unfiltered) reads- 8 - 13241894vlAttorney Docket. No.: 2013237-0970measured in a forward direction and a reverse direction for each of the alternate allele and the wild- type allele.
[0047] In some embodiments, provided methods comprise determining, for a given initial event, a forward direction mutation confidence score based on forward reads mapping to the given initial event and a reverse direction mutation confidence score based on reverse reads mapping to the given initial event; and filtering or retaining the given initial event based at least in part on (i) the likelihood of strand bias, (ii) the forward direction mutation confidence score, and (iii) the reverse direction mutation confidence score.
[0048] In some embodiments, provided methods comprise determining, for a given initial event: a forward direction refined mutation confidence score based on forward reads mapping to the given initial event and the estimated tumor sample purity; and a reverse direction refined mutation confidence score based on reverse reads mapping to the given initial event and the estimated tumor sample purity; and filtering or retaining the given initial event based at least in part on (i) the likelihood of strand bias, (ii) the forward direction refined mutation confidence score, and (iii) the reverse direction refined mutation confidence score.
[0049] In some embodiments, provided methods comprise obtaining, by the processor, normal sequencing data and step (d) comprises using a normal coverage filter that excludes initial events if coverage for corresponding locations in the normal sequencing data is below a (e.g., predetermined) coverage level threshold.
[0050] In some embodiments, provided methods comprise, at step (d), determining, by the processor, a refined mutation confidence score based on tumor reads mapping to the initial event and retaining initial events whose refined mutation confidence score is above a predetermined threshold level (e.g., even if their coverage in the normal sequencing data is below the coverage level threshold).
[0051] In some embodiments, step (c) comprises using a measured histological tumor content of the tumor sample as the estimated purity for the tumor sample.
[0052] In some embodiments, step (c) comprises determining the estimated purity for the tumor sample and / or bound thereon based at least in part on the tumor sequencing data.- 9 - 13241894vlAttorney Docket. No.: 2013237-0970
[0053] In some embodiments, provided methods comprise, determining, by the processor, the estimated purity for the tumor sample and / or bound thereon using a copy number variation (CNV)-based purity estimation procedure [e.g., e.g., that analyzes a read distribution of CNV events within the tumor genome (e.g., that fits one or more predicted copy number grids to a distribution of read counts and major allele frequencies determined for a plurality of heterozygous segments within a tumor genome)] based on the tumor sequencing data.
[0054] In some embodiments, provided methods comprise determining absolute copy numbers and / or allele specific copy numbers for one or more segments (e.g., a primary copy number for a subset of primary balanced segments) within the tumor genome using the CNV-based purity estimation procedure.
[0055] In some embodiments, step (d) comprises filtering the plurality of initial events based at least in part on the absolute copy numbers and / or allele specific copy numbers [e.g., by assigning absolute copy numbers to tumor genome comprising the plurality of initial events and, for each initial event, comparing a measured VAF with a predicted VAF determined based on the absolute copy number of the tumor genome segment comprising the initial event and the estimated tumor sample purity].
[0056] In some embodiments, provided methods comprise determining the estimated purity for the tumor sample and / or bound thereon using a single nucleotide variant (SNV)-based purity estimation procedure that models expected variant allele frequencies (VAF) and / or an expected distribution of VAFs for at least a portion of the plurality of initial events.
[0057] In some embodiments, provided methods comprise modeling the initial events as biologically clonal (e.g., occurring in each cancer cell of the tumor sample).
[0058] In some embodiments, a measured purity of the tumor sample is less than about 0.4 (e.g., less than 0.3, less than 0.25, less than 0.20, less than 0.19, less than 0.18, less than 0.17, less than 0.16, less than 0.15, less than 0.14, less than 0.13, less than 0.12, or less than 0.11) [e.g., wherein the measured purity of the tumor sample is a histological tumor content (e.g., determined via histology)] and / or wherein the estimated purity of the tumor sample is less than about 0.3.- 10 - 13241894vlAttorney Docket. No.: 2013237-0970
[0059] In some embodiments, provided methods comprise obtaining, by the processor, normal sequencing data for a normal sample obtained from the subject and using the normal sequencing data and the tumor sequencing data to determine, and filter initial events based on, a likelihood that a given initial event arises from noise.
[0060] In some embodiments, tumor sequencing data comprises RNA sequencing data (RNAseq) and the method comprises using the RNAseq data to filter errors [e.g., using RNAseq in combination with DNA sequencing data; e.g., prior to detecting initial events at step (b)].
[0061] In some aspects, the present disclosure provides a method (e.g., for iterative mutation detection and refinement of a tumor model) comprising: (a) obtaining, by a processor of a computing device, tumor sequencing data for a tumor sample obtained from a subject, the tumor sequencing data representing a tumor genome associated with the tumor sample [e.g., the tumor sequencing data comprising a plurality of tumor reads, each tumor read representing a partial nucleotide sequence of the tumor genome as determined by sequencing the tumor sample]; (b) detecting, by the processor, using the tumor sequencing data, a plurality of initial events; (c) iteratively: filtering, by the processor, the plurality of initial events based on a set of one or more tumor model parameters, to generate a filtered list of putative mutations and (e.g., then, e.g., within a given iteration) using, by the processor, the filtered list of putative mutations to refine the set of tumor model parameters, wherein the refined set of tumor model parameters determined at one iteration is used to filter the plurality of initial events in a subsequent iteration, and, at a final iteration, the plurality of initial events is filtered based on a final refined set of tumor model parameters determined at a previous (e.g., next to last) iteration, thereby generating a final list of putative mutations; and (d) storing and / or providing, by the processor, the final list of putative mutations for display and / or further processing.
[0062] In some embodiments, the filtering (e.g., at step (c)) is performed by the method of any of the aspects and embodiments described herein, for example in paragraphs above.
[0063] In some embodiments, at a first iteration, the set of tumor model parameters is an initial set of tumor model parameters, having been determined and / or obtained without using the plurality of initial events and / or the filtered list of putative mutations.
[0064] In some embodiments, step (c) comprises, at a first iteration, filtering the plurality of initial events based on a generic (e.g., stored, predetermined) set of tumor model parameters - 11 - 13241894vlAttorney Docket. No.: 2013237-0970comprising, for each of the one or more tumor model parameters of the set, (i) a (e.g., predetermined) range of expected values and / or (ii) an (e.g., predetermined) expected value.
[0065] In some embodiments, filtering the plurality of initial events based on the generic set of tumor model parameters comprises, for each of at least a portion of the initial events: determining, by the processor, a mutation confidence score that quantifies (i) a likelihood of the initial event encoding an alternate allele and thereby representing a true underlying mutation relative to (ii) a likelihood of the initial event encoding a normal allele, wherein the mutation confidence score is determined, for a given initial event, based at least in part (i) reads of the tumor sequencing data that map to the given initial event and (ii) the generic set of tumor model parameters; and identifying, by the processor, the initial event as a false positive to be excluded from a preliminary mutation list or identifying, by the processor, the initial event as a putative mutation for inclusion in the preliminary mutation list based on the mutation confidence score determined for the initial event in comparison with a (e.g., predetermined) mutation confidence score threshold value.
[0066] In some embodiments, step (c) comprises, at a first iteration: determining, by the processor, an initial set of tumor model parameters using a copy number variation (CNV)-based purity estimation procedure [e.g., that analyzes a read distribution of CNV events within the tumor genome (e.g., that fits one or more predicted copy number grids to a distribution of read counts and major allele frequencies determined for a plurality of heterozygous segments within the tumor genome)] based on the tumor sequencing data; and filtering the plurality of initial events based on the initial set of tumor model parameters.
[0067] In some embodiments, provided methods comprise determining absolute copy numbers and / or allele specific copy numbers for one or more segments (e.g., a primary copy number for a subset of primary balanced segments) within the tumor genome using the CNV-based purity estimation procedure.
[0068] In some embodiments, step (d) comprises filtering the plurality of initial events based at least in part on the absolute copy numbers and / or allele specific copy numbers [e.g., by assigning absolute copy numbers to tumor genome comprising the plurality of initial events and, for each initial event, comparing a measured VAF with a predicted VAF determined based on the- 12 - 13241894vlAttorney Docket. No.: 2013237-0970absolute copy number of the tumor genome segment comprising the initial event and the estimated tumor sample purity].
[0069] In some embodiments, provided methods comprise, at a first iteration: obtaining a measured tumor content of the tumor sample (e.g., determined via histology, such as histological tumor content) and including the measured tumor content in an initial set of tumor model parameters; and filtering the plurality of initial events based on the initial set of tumor model parameters.
[0070] In some embodiments, provided methods comprise, at a first iteration: obtaining one or more predetermined purity values (e.g., a predetermined purity value obtained from a database; e.g., a predetermined range of purity values) and / or one or more predetermined copy numbers and including the measured tumor content and / or one or more predetermined copy numbers in an initial set of tumor model parameters; and filtering the plurality of initial events based on the initial set of tumor model parameters
[0071] In some embodiments, at step (c), for one or more iterations, refining the set of tumor model parameters comprises determining, by the processor, an estimated tumor sample purity and / or bound thereon based on (i) the filtered list of putative mutations and (ii) the tumor sequencing data.
[0072] In some embodiments, determining the estimated tumor sample purity and / or bound thereon comprises using a single nucleotide variant (SNV)-based purity estimation procedure that models expected variant allele frequencies (VAF) and / or an expected distribution of VAFs for at least a portion of the filtered list of putative mutations.
[0073] In some embodiments, determining the estimated tumor sample purity and / or bound thereon comprises using a copy number variation (CNV)-based purity estimation procedure [e.g., e.g., that analyzes a read distribution of CNV events within the tumor genome (e.g., that fits one or more predicted copy number grids to a distribution of read counts and major allele frequencies determined for a plurality of heterozygous segments within a tumor genome)] based on the tumor sequencing data.
[0074] In some embodiments, provided methods comprise determining absolute copy numbers and / or allele specific copy numbers for one or more segments (e.g., a primary copy- 13 - 13241894vlAttorney Docket. No.: 2013237-0970number for a subset of primary balanced segments) within the tumor genome using the CNV-based purity estimation procedure.
[0075] In some embodiments, provided methods comprise determining a VAF distribution based on the filter list of putative mutations.
[0076] In some embodiments, determining the estimated tumor sample purity and / or bound thereon comprises: determining a first estimated tumor sample purity and / or a first tumor sample purity bound using a single nucleotide variant (SNV)-based purity estimation procedure; determining a second estimated tumor sample purity and / or a second tumor sample purity bound using a copy number variation (CNV)-based purity estimation procedure that fits a predicted a copy number grid to a distribution of read counts and major allele frequencies determined for a plurality of heterozygous segments within a tumor genome; and selecting, by the processor, for use as the estimated tumor sample purity and / or bound thereon, one of: the first estimated tumor sample purity, the first tumor sample purity bound, the second estimated tumor sample purity, and the second tumor sample purity bound.
[0077] In some embodiments, each initial event corresponds to a variant portion of the tumor genome (e.g., identifying one or more sites within the tumor genome), wherein, for each initial event, the corresponding variant portion is determined, based on the sequencing data, to (e.g., potentially) have a different copy number and / or a different nucleotide sequence relative to a corresponding (e.g., aligned with the corresponding variant portion) normal reference [e.g., wherein the normal reference is obtained from a database (e.g., a hl9 reference genome); e.g., wherein the normal reference is determined based on normal sequencing data obtained for the subject (e.g., by sequencing a normal sample, e.g., of healthy tissue and / or cells)].
[0078] In some embodiments, for each of at least a portion of the plurality of initial events, the corresponding variant portion is determined, based on the sequencing data, to (e.g., potentially) have a different nucleotide sequence relative to the corresponding normal reference.
[0079] In some embodiments, the different nucleotide sequence is or comprises an alternate allele (e.g., an individual nucleotide), different from an allele of the corresponding normal reference (e.g., a wild-type allele).- 14 - 13241894vlAttorney Docket. No.: 2013237-0970
[0080] In some embodiments, the different nucleotide sequence is or comprises an insertion and / or a deletion (e.g., of one or more nucleotides) relative to the corresponding normal reference (e.g., an indel).
[0081] In some embodiments, the variant portion is or comprises a structural variation relative to the normal reference.
[0082] In some embodiments, for each of at least a portion of the plurality of initial events, the corresponding variant portion is determined, based on the sequencing data, to (e.g., potentially) have a different copy number than the corresponding normal reference.
[0083] In some embodiments, for each of at least a portion of the plurality of initial events, the corresponding variant portion is determined, based on the sequencing data, to (e.g., potentially) have a different nucleotide sequence relative to the corresponding normal reference.
[0084] In some embodiments, the different nucleotide sequence is or comprises an alternate allele (e.g., an individual nucleotide), different from a normal allele of the corresponding normal reference.
[0085] In some embodiments, step (d) comprises, for each initial event of at least a portion of the initial events: determining, by the processor, a mutation confidence score that quantifies (i) a likelihood of the initial event encoding the alternate allele and thereby representing a true underlying mutation relative to (ii) a likelihood of the initial event encoding the normal allele, wherein the mutation confidence score for a given initial event is determined based at least in part on (i) reads of the tumor sequencing data that map to the given initial event and (ii) a set of tumor model parameters; and identifying, by the processor, the initial event as (i) a false positive to be excluded from the mutation list or (ii) a putative mutation for inclusion in the mutation list based on the mutation confidence score determined for the initial event in comparison with a (e.g., predetermined) mutation confidence score threshold value.
[0086] In some embodiments, the set of tumor model parameters comprises the estimated purity and / or bound thereon for the tumor sample, such that the mutation confidence score is a refined mutation confidence score (e.g., determined based on an estimated purity specific to tumor sample).- 15 - 13241894vlAttorney Docket. No.: 2013237-0970
[0087] In some embodiments, step (d) comprises, for each initial event of at least a portion of the initial events: determining, by the processor, a value of a test statistic for the initial event based at least in part on the tumor sequencing data; determining, by the processor, a mutation confidence score that quantifies (i) a likelihood of the initial event encoding the alternate allele and thereby representing a true underlying mutation relative to (ii) a likelihood of the initial event encoding the wild-type allele, wherein the mutation confidence score is determined, for a given initial event, based at least in part on (i) reads of the tumor sequencing data that map to the given initial event and (ii) a set of tumor model parameters; selecting, by the processor, a value of a discrimination threshold (e.g., P-value significance threshold, thresholds on computed or measured quantities, etc.) based at least in part on the mutation confidence score determined for the initial event; and identifying, by the processor, the initial event as (i) a false positive to be excluded from the mutation list or (ii) a putative mutation for inclusion in the mutation list based on the value of the test statistic determined for the initial event and the selected value of the discrimination threshold.
[0088] In some embodiments, the set of tumor model parameters comprises the estimated purity and / or bound thereon for the tumor sample, such that the mutation confidence score is a refined mutation confidence score (e.g., determined based on an estimated purity specific to tumor sample).
[0089] In some embodiments, step (d) comprises: selecting (e.g., adjusting), by the processor, value(s) of one or more discrimination thresholds based on the estimated purity and / or bound thereon; and using, by the processor, the selected value(s) of the one or more discrimination thresholds to filter the plurality of initial events.
[0090] In some embodiments, step (d) comprises, for each initial event of at least a portion of the initial events: determining, by the processor, a value of a test statistic for the initial event based at least in part on the tumor sequencing data; selecting (e.g., adjusting), by the processor, a value of a discrimination threshold based at least in part on the estimated purity for the tumor sample and / or the bound thereon; and identifying, by the processor, the initial event as (i) a false positive to be excluded from the mutation list or (ii) a putative mutation for inclusion in the mutation list based on the value of the test statistic determined for the initial event and the selected value of the discrimination threshold.- 16 - 13241894vlAttorney Docket. No.: 2013237-0970
[0091] In some embodiments, step (d) comprises: selecting, by the processor, a particular false positive filter of a plurality of false positive filters based at least in part on the estimated tumor sample purity; and identifying, by the processor, the initial event as (i) a false positive to be excluded from the mutation list or (ii) a putative mutation for inclusion in the mutation list, using the particular selected false positive filter.
[0092] In some embodiments, a measured purity of the tumor sample is less than about 0.4 (e.g., less than 0.3, less than 0.25, less than 0.20, less than 0.19, less than 0.18, less than 0.17, less than 0.16, less than 0.15, less than 0.14, less than 0.13, less than 0.12, or less than 0.11) (e.g., wherein the measured purity of the tumor sample is a histological tumor content (e.g., determined via histology)] and / or wherein the estimated purity of the tumor sample is less than about 0.3.
[0093] In some embodiments, provided methods comprise obtaining, by the processor, normal sequencing data for a normal sample obtained from the subject and using the normal sequencing data and the tumor sequencing data to determine, and filter initial events based on, a likelihood that a given initial event arises from noise.
[0094] In some embodiments, the tumor sequencing data comprises RNA sequencing data (RNAseq) and the method comprises using the RNAseq data to filter errors [e.g., using RNAseq in combination with DNA sequencing data; e.g., prior to detecting initial events at step (b)].
[0095] In some aspects, the present disclosure provides a method (e.g., a computer-implemented method) (e.g., for detecting one or more cancer mutations for a subject) comprising: (a) obtaining, by a processor of a computing device, tumor sequencing data for a tumor sample obtained from a subject, the tumor sequencing data representing a tumor genome associated with the tumor sample [e.g., the tumor sequencing data comprising a plurality of tumor reads, each tumor read representing a partial nucleotide sequence determined by sequencing the tumor sample]; (b) detecting, by the processor, using the tumor sequencing data, a plurality of initial events, each initial event corresponding to variant portion of the tumor genome (e.g., identifying one or more sites within a tumor genome), wherein, for each initial event, the corresponding variant portion is determined, based on the sequencing data, to (e.g., potentially) have a different copy number and / or a different nucleotide sequence relative to a - 17 - 13241894vlAttorney Docket. No.: 2013237-0970corresponding (e.g., aligned with the corresponding variant portion) normal reference [e.g., wherein the normal reference is obtained from a database; e.g., wherein the normal reference is determined based on normal sequencing data obtained for the subject (e.g., by sequencing a normal sample, e.g., of healthy tissue and / or cells)]; (c) filtering, by the processor, the plurality of initial events to exclude initial events identified as false-positives based on the tumor sequencing data, and retaining other initial events (e.g., not identified as false-positives and excluded) of the plurality of initial events as putative mutations for inclusion in an initial mutation list; (d) determining, by the processor, an initial tumor sample purity estimate and / or a bound (e.g., an upper bound; e.g., a lower bound) thereon (e.g., based on the tumor sequencing data and initial mutation list); (e) filtering, by the processor, the plurality of initial of initial events to exclude putative mutations identified as false-positives based on (i) the tumor sequencing data and (ii) the initial tumor sample purity estimate and / or bound thereon, and retaining other initial events (e.g., not identified as false-positives and excluded) of the plurality of initial events as putative mutations for inclusion in an refined mutation list; and (g) storing and / or providing, for display and / or further processing, the refined mutation list.
[0096] In some embodiments, provided methods comprise determining, by the processor, a refined tumor sample purity estimate and / or a bound thereon based on the tumor sequencing data and the refined mutation list; filtering, by the processor, the plurality of initial events to exclude putative mutations identified as false-positives based on (i) the tumor sequencing data and (ii) the refined tumor sample purity estimate and / or bound thereon, and retaining other putative mutations (e.g., not identified as false positives and thus excluded) as putative mutations for inclusion in an updated version of the refined mutation list; and storing and / or providing, for display and / or further processing, the updated version of the refined mutation list.
[0097] In some embodiments, provided methods comprise, iteratively: determining, by the processor, a refined tumor sample purity estimate and / or a bound thereon based on the tumor sequencing data and the refined mutation list; filtering, by the processor, the plurality of initial events to exclude putative mutations identified as false-positives based on (i) the tumor sequencing data and (ii) the refined tumor sample purity estimate and / or bound thereon, and retaining other putative mutations (e.g., not identified as false positives and thus excluded) as putative mutations for inclusion in an updated version of the refined mutation list; and iteratively- 18 - 13241894vlAttorney Docket. No.: 2013237-0970(i) updating the refined tumor sample purity estimate and / or bound thereon using the updated version of the refined mutation list and (ii) filtering the plurality of initial events using the updated tumor sample purity estimate and / or bound thereon to further update the refined mutation list.
[0098] In some aspects, the present disclosure provides a method (e.g., a computer-implemented method) (e.g., for detecting one or more cancer mutations for a subject) comprising: (a) obtaining, by a processor of a computing device, tumor sequencing data for a tumor sample obtained from a subject, the tumor sequencing data representing a tumor genome associated with the tumor sample [e.g., the tumor sequencing data comprising a plurality of tumor reads, each tumor read representing a partial nucleotide sequence determined by sequencing the tumor sample]; (b) detecting, by the processor, using on the tumor sequencing data, a plurality of initial events, each initial event corresponding to variant portion of the tumor genome (e.g., identifying one or more sites within a tumor genome), wherein, for each initial event, the corresponding variant portion is determined, based on the sequencing data, to (e.g., potentially) have a different copy number and / or a different nucleotide sequence relative to a corresponding (e.g., aligned with the corresponding variant portion) normal reference [e.g., wherein the normal reference is obtained from a database; e.g., wherein the normal reference is determined based on normal sequencing data obtained for the subject (e.g., by sequencing a normal sample, e.g., of healthy tissue and / or cells)]; (c) filtering, by the processor, the plurality of initial events to exclude initial events identified as false-positives based on the tumor sequencing data, and retaining other initial events (e.g., not identified as false-positives and excluded) as putative mutations for inclusion in an initial mutation list; (d) determining, by the processor, an initial set of tumor model parameters based on the tumor sequencing data (e.g., and the initial mutation list); (e) filtering, by the processor, the plurality of initial of initial events to exclude putative mutations identified as false-positives based on (i) the tumor sequencing data and (ii) the initial set of tumor model parameters, and retaining other initial events (e.g., not identified as false-positives and excluded) as putative mutations for inclusion in an refined mutation list; (g) storing and / or providing, for display and / or further processing, the refined mutation list.- 19 - 13241894vlAttorney Docket. No.: 2013237-0970
[0099] In some embodiments, provided methods comprise determining, by the processor, a refined set of tumor model parameters based on the tumor sequencing data and the refined mutation list; filtering, by the processor, the plurality of initial events to exclude putative mutations identified as false-positives based on (i) the tumor sequencing data and (ii) the refined set of tumor model parameters, and retaining other putative mutations (e.g., not identified as false positives and thus excluded) as putative mutations for inclusion in an updated version of the refined mutation list; and storing and / or providing, for display and / or further processing, the updated version of the refined mutation list.
[0100] In some embodiments, provided methods comprise, iteratively: determining, by the processor, a refined set of tumor model parameters based on the tumor sequencing data and the refined mutation list; and filtering, by the processor, the plurality of initial events to exclude putative mutations identified as false-positives based on (i) the tumor sequencing data and (ii) the refined set of tumor model parameters, and retaining other putative mutations (e.g., not identified as false positives and thus excluded) as putative mutations for inclusion in an updated version of the refined mutation list; and iteratively (i) updating the refined set of tumor model parameters using the updated version of the refined mutation list and (ii) filtering the plurality of initial events using the updated version of the refined set of tumor model parameters to further update the refined mutation list.
[0101] In some aspects, the present disclosure provides a system (e.g., for detecting one or more cancer mutations for a subject) comprising: a processor of a computing device; and memory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to: (a) obtain tumor sequencing data for a tumor sample obtained from a subject, the tumor sequencing data representing a tumor genome associated with the tumor sample [e.g., the tumor sequencing data comprising a plurality of tumor reads, each tumor read representing a partial nucleotide sequence of the tumor genome as determined by sequencing the tumor sample]; (b) detect, using on the tumor sequencing data, a plurality of initial events, each initial event corresponding to a variant portion of the tumor genome (e.g., identifying one or more sites within the tumor genome), wherein, for each initial event, the corresponding variant portion is determined, based on the sequencing data, to (e.g., potentially) have a different copy number and / or a different nucleotide sequence relative to a corresponding- 20 - 13241894vlAttorney Docket. No.: 2013237-0970(e.g., aligned with the corresponding variant portion) normal reference [e.g., wherein the normal reference is obtained from a database (e.g., a hl9 reference genome); e.g., wherein the normal reference is determined based on normal sequencing data obtained for the subject (e.g., by sequencing a normal sample, e.g., of healthy tissue and / or cells)]; (c) determine an estimated purity for the tumor sample and / or a bound (e.g., an upper bound; e.g., a lower bound) thereon (e.g., based on the tumor sequencing data); (d) filter the plurality of initial events to exclude initial events identified as false-positives based on (i) the tumor sequencing data and (ii) the estimated purity for the tumor sample and / or bound thereon and retaining other initial events (e.g., not identified as false-positives and excluded) as putative mutations for inclusion in a mutation list; and (e) store and / or provide, for display and / or further processing, the mutation list.
[0102] In some aspects, the present disclosure provides a system (e.g., for detecting one or more cancer mutations for a subject) comprising: a processor of a computing device; and memory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to: (a) obtain tumor sequencing data for a tumor sample obtained from a subject, the tumor sequencing data representing a tumor genome associated with the tumor sample [e.g., the tumor sequencing data comprising a plurality of tumor reads, each tumor read representing a partial nucleotide sequence of the tumor genome as determined by sequencing the tumor sample]; (b) detecting, by the processor, using the tumor sequencing data, a plurality of initial events; (c) iteratively: filter the plurality of initial events based on a set of one or more tumor model parameters, to generate a filtered list of putative mutations and (e.g., then, e.g., within a given iteration) use the filtered list of putative mutations to refine the set of tumor model parameters, wherein the refined set of tumor model parameters determined at one iteration is used to filter the plurality of initial events in a subsequent iteration, and, at a final iteration, the plurality of initial events is filtered based on a final refined set of tumor model parameters determined at a previous (e.g., next to last) iteration, thereby generating a final list of putative mutations; and (d) store and / or provide, the final list of putative mutations for display and / or further processing.
[0103] In some aspects, the present disclosure provides a system (e.g., for detecting one or more cancer mutations for a subject) comprising: a processor of a computing device; and- 21 - 13241894vlAttorney Docket. No.: 2013237-0970memory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to: (a) obtain tumor sequencing data for a tumor sample obtained from a subject, the tumor sequencing data representing a tumor genome associated with the tumor sample [e.g., the tumor sequencing data comprising a plurality of tumor reads, each tumor read representing a partial nucleotide sequence determined by sequencing the tumor sample]; (b) detect, using the tumor sequencing data, a plurality of initial events, each initial event corresponding to variant portion of the tumor genome (e.g., identifying one or more sites within a tumor genome), wherein, for each initial event, the corresponding variant portion is determined, based on the sequencing data, to (e.g., potentially) have a different copy number and / or a different nucleotide sequence relative to a corresponding (e.g., aligned with the corresponding variant portion) normal reference [e.g., wherein the normal reference is obtained from a database; e.g., wherein the normal reference is determined based on normal sequencing data obtained for the subject (e.g., by sequencing a normal sample, e.g., of healthy tissue and / or cells)]; (c) filter the plurality of initial events to exclude initial events identified as false-positives based on the tumor sequencing data, and retaining other initial events (e.g., not identified as false-positives and excluded) of the plurality of initial events as putative mutations for inclusion in an initial mutation list; (d) determine an initial tumor sample purity estimate and / or a bound (e.g., an upper bound; e.g., a lower bound) thereon (e.g., based on the tumor sequencing data and initial mutation list); (e) filter the plurality of initial of initial events to exclude putative mutations identified as false-positives based on (i) the tumor sequencing data and (ii) the initial tumor sample purity estimate and / or bound thereon, and retaining other initial events (e.g., not identified as false-positives and excluded) of the plurality of initial events as putative mutations for inclusion in an refined mutation list; and (g) store and / or provide, for display and / or further processing, the refined mutation list.
[0104] In some aspects, the present disclosure provides a system (e.g., for detecting one or more cancer mutations for a subject) comprising: a processor of a computing device; and memory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to: (a) obtain tumor sequencing data for a tumor sample obtained from a subject, the tumor sequencing data representing a tumor genome associated with the tumor sample [e.g., the tumor sequencing data comprising a plurality of tumor reads, each tumor read representing a partial nucleotide sequence determined by sequencing the tumor sample]; (b)- 22 - 13241894vlAttorney Docket. No.: 2013237-0970detect, using on the tumor sequencing data, a plurality of initial events, each initial event corresponding to variant portion of the tumor genome (e.g., identifying one or more sites within a tumor genome), wherein, for each initial event, the corresponding variant portion is determined, based on the sequencing data, to (e.g., potentially) have a different copy number and / or a different nucleotide sequence relative to a corresponding (e.g., aligned with the corresponding variant portion) normal reference [e.g., wherein the normal reference is obtained from a database; e.g., wherein the normal reference is determined based on normal sequencing data obtained for the subject (e.g., by sequencing a normal sample, e.g., of healthy tissue and / or cells)]; (c) filter the plurality of initial events to exclude initial events identified as false-positives based on the tumor sequencing data, and retaining other initial events (e.g., not identified as false-positives and excluded) as putative mutations for inclusion in an initial mutation list; (d) determine an initial set of tumor model parameters based on the tumor sequencing data (e.g., and the initial mutation list); (e) filter the plurality of initial of initial events to exclude putative mutations identified as false-positives based on (i) the tumor sequencing data and (ii) the initial set of tumor model parameters, and retaining other initial events (e.g., not identified as false-positives and excluded) as putative mutations for inclusion in an refined mutation list; (g) store and / or provide, for display and / or further processing, the refined mutation list.
[0105] In some aspects, the present disclosure provides a method of producing an immunotherapy construct (e.g., a cancer vaccine) for a subject, the method comprising detecting a plurality of candidate mutations (e.g., somatic mutations; e.g., non-synonymous somatic mutations) in tumor cells of from the subject using a method or system disclosed herein; and synthesizing, as the immunotherapy construct, a polyribonucleotide encoding a plurality of neoepitopes, each encoded neoepitope corresponding to a candidate mutation of the plurality of candidate mutations.
[0106] In some aspects, the present disclosure provides a method comprising: determining a subset of T-cells within a population of T-cells that are capable of specifically binding a plurality of complexes, wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer-specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, and wherein the plurality of complexes comprises a plurality of- 23 - 13241894vlAttorney Docket. No.: 2013237-0970neoepitope portions encoded by cancer-specific mutations detected using a method or system disclosed herein; and enriching for (e.g., expanding) the subset of T-cells that are capable of specifically binding the plurality of complexes.
[0107] In some aspects, the present disclosure provides a method comprising: administering to a subject a population of T-cells, wherein the population of T cells are capable of specifically binding a plurality of complexes, wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, wherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer-specific mutations detected using a method or system disclosed herein, and wherein the genome of at least some of the subject’s cells comprises a subset (e.g., all) of the cancer-specific mutations.
[0108] In some aspects, the present disclosure provides a method comprising: determining a subset of TILs within a population of TILs that are capable of specifically binding a plurality of complexes, wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, and wherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer-specific mutations detected using a method or system disclosed herein; and enriching for (e.g., expanding) the subset of TILs that are capable of specifically binding the plurality of complexes.
[0109] In some aspects, the present disclosure provides a method comprising: administering to a subject a population of TILs, wherein the population of TILs are capable of specifically binding a plurality of complexes, wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancerspecific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, wherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer- specific mutations detected using a method or system disclosed herein, and wherein the genome of at least some of the subject’s cells comprises a subset (e.g., all) of the cancer-specific mutations.- 24 - 13241894vlAttorney Docket. No.: 2013237-0970
[0110] In some embodiments, provided methods comprise obtaining a tumor sample from the subject and using the tumor sample to detect the plurality of candidate mutations (e.g., by sequencing the tumor sample).
[0111] In some embodiments, provided methods (e.g., further) comprise obtaining a normal sample from the subject and using the normal sample (e.g., together with the tumor sample) to detect the plurality of cancer mutations (e.g., by sequencing the normal sample).
[0112] In some embodiments, provided methods comprise sequencing the tumor and / or normal sample (e.g., in replicates).
[0113] In some aspects, the present disclosure provides a pharmaceutical composition comprising a polyribonucleotide encoding (e.g., a polyepitopic vaccine construct comprising) a plurality of neoepitopes, wherein at least a portion of the plurality of neoepitopes are individualized (e.g., patient- specific) neoepitopes corresponding to (e.g., encoded by a nucleotide sequence comprising) a somatic mutation from a population of candidate mutations detected in a tumor sample of a subject and selected for inclusion in the pharmaceutical composition using a method or system disclosed herein (e.g., each particular neoepitope encoded by a nucleotide sequence comprising one or more candidate mutation detected using a method or system disclosed herein).
[0114] In some aspects, the present disclosure provides an individualized pharmaceutical composition (e.g., associated with and / or intended for administration to a particular subject) comprising a polyribonucleotide encoding (e.g., a polyepitopic vaccine construct comprising) a plurality of neoepitopes for a particular subject, wherein at least a portion of the neoepitopes are individualized (e.g., patient- specific) neoepitopes corresponding to (e.g., encoded by a nucleotide sequence comprising) a somatic mutation from a population of candidate mutations detected in a tumor sample from the particular subject and selected for inclusion in the pharmaceutical composition using a method or system disclosed herein.
[0115] In some aspects, the present disclosure provides a population of T-cells, wherein the population of T-cells is capable of specifically binding to a plurality of complexes, wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, and wherein - 25 - 13241894vlAttorney Docket. No.: 2013237-0970the plurality of complexes comprises a plurality of neoepitope portions encoded by cancerspecific mutations using a method or system disclosed herein.
[0116] In some aspects, the present disclosure provides a T-cell capable of specifically binding to a complex, wherein the complex comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, wherein the one or more cancer-specific mutations are detected using a method or system disclosed herein.
[0117] In some aspects, the present disclosure provides a population of T-cells, wherein the population of T-cells is capable of specifically binding to a plurality of complexes, wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, and wherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancerspecific mutations detected using a method or system disclosed herein.
[0118] In some aspects, the present disclosure provides a T-cell capable of specifically binding to a complex, wherein the complex comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, wherein the one or more cancer-specific mutations are within the selected subset of candidate mutations detected using a method or system disclosed herein.
[0119] In some aspects, the present disclosure provides a T-cell receptor capable of specifically binding to a complex, wherein the complex comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, wherein the one or more cancer-specific mutations are within the selected subset of candidate mutations detected using a method or system disclosed herein.
[0120] In some aspects, the present disclosure provides a chimeric antigen receptor (CAR) capable of specifically binding to a complex, wherein the complex comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer-specific - 26 - 13241894vlAttorney Docket. No.: 2013237-0970mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, wherein the one or more cancer- specific mutations are within the selected subset of candidate mutations detected using a method or system disclosed herein.
[0121] Features of embodiments described with respect to one aspect of the invention may be applied with respect to another aspect of the invention.BRIEF DESCRIPTION OF THE DRAWING
[0122] The foregoing and other objects, aspects, features, and advantages of the present disclosure will become more apparent and better understood by referring to the following description taken in conjunction with the accompanying drawings, in which:
[0123] FIG. 1 is a block flow diagram showing an example process for creating a personalized cancer immunotherapy, according to an illustrative embodiment.
[0124] FIG. 2 is a block flow diagram showing an example process for obtaining sequencing data from tumor and / or normal sample(s), according to an illustrative embodiment.
[0125] FIG. 3 is a schematic showing a model of a normal genome and a tumor genome of a subject, according to an illustrative embodiment.
[0126] FIG. 4 is a schematic of a tumor sample, according to an illustrative embodiment.
[0127] FIG. 5A is a diagram illustrating interrelation between parameters and / or measurements characterizing tumor sample properties, according to an illustrative embodiment.
[0128] FIG. 5B is a diagram showing possible purity values for different copy numbers and zygosities, determined for a VAF of 0.2, wherein highlighted rows in the diagram correspond to balanced SNVs.
[0129] FIG. 6 is a block flow diagram of an example tumor deconvolution process for determining a tumor sample purity estimate, according to an illustrative embodiment.
[0130] FIG. 7 is a block flow diagram of an example process for identifying and / or selecting particular subsets of segments, according to an illustrative embodiment.- 27 - 13241894vlAttorney Docket. No.: 2013237-0970
[0131] FIG. 8 is a schematic illustrating certain categories of segments within a genome (e.g., a tumor genome and / or a normal genome) and approaches for their identification, according to an illustrative embodiment.
[0132] FIG. 9 is a schematic illustrating certain categories of segments within a genome (e.g., a tumor genome and / or a normal genome) and approaches for their identification, according to an illustrative embodiment.
[0133] FIG. 10 is a schematic with an arrangement of example plots illustrating an example tumor deconvolution workflow, including how estimates of tumor sample properties such as purity, and absolute copy number and allele specific copy number assignments can be used to detect and characterize cancer- specific mutations, such as SNV events, according to an illustrative embodiment.
[0134] FIG. 11 is a block flow diagram of an example process for using multiple tumor models for obtaining purity estimates and / or estimated purity bounds, according to an illustrative embodiment.
[0135] FIG. 12 is a block flow diagram showing an exemplary SNV-based purity estimation process including quality control steps, according to an illustrative embodiment.
[0136] FIG. 13 is a block flow diagram showing a decision tree for selecting between various purity estimation results, according to an illustrative embodiment.
[0137] FIG. 14 is a block flow diagram of an example process for analyzing VAF distributions of SNV subsets, according to an illustrative embodiment.
[0138] FIG. 15 is a block flow diagram of an exemplary process for performing filtering and tumor modeling with feedback, according to an illustrative embodiment.
[0139] FIG. 16 is a block flow diagram of an exemplary process for performing filtering and tumor modeling with feedback, according to an illustrative embodiment.
[0140] FIG. 17 is a block flow diagram of an exemplary process for performing filtering and tumor modeling with feedback, according to an illustrative embodiment.- 28 - 13241894vlAttorney Docket. No.: 2013237-0970
[0141] FIG. 18 is a block flow diagram of an exemplary process for detection and selection of mutations to be targeted with a personalized cancer immunotherapy which utilizes filtering and tumor modeling with feedback, according to an illustrative embodiment.
[0142] FIG. 19 is a block flow diagram of an exemplary process for performing filtering and tumor modeling with feedback, according to an illustrative embodiment.
[0143] FIG. 20 is a block flow diagram of an exemplary process for performing filtering and tumor modeling with feedback, according to an illustrative embodiment.
[0144] FIG. 21 is a schematic depicting modulation of sensitivity and / or specificity of a discrimination threshold of a false positive filter based on a mutation confidence score, as disclosed herein.
[0145] FIG. 22 is a schematic depicting improvement of a mutation confidence filter informed by a tumor model, as disclosed herein.
[0146] FIG. 23 is a schematic showing improvement in mutation detection using a refined mutation confidence score, as disclosed herein.
[0147] FIG. 24 is a block flow diagram of an exemplary pre-false positive (pre-FP) filter module, according to an illustrative embodiment.
[0148] FIG. 25 is a block flow diagram of an exemplary pre-FP filter module, according to an illustrative embodiment.
[0149] FIG. 26 is a block flow diagram of an exemplary pre-FP filter module, according to an illustrative embodiment.
[0150] FIG. 27 is a block flow diagram of an exemplary pre-FP filter module, according to an illustrative embodiment.
[0151] FIG. 28 is a block flow diagram of an exemplary post-FP filter module, according to an illustrative embodiment.
[0152] FIG. 29A is a diagram illustrating possible states (e.g., base calls) that may occur in sequencing data, and associated probabilities, without an underlying physical mutation occurring, according to an illustrative embodiment.- 29 - 13241894vlAttorney Docket. No.: 2013237-0970
[0153] FIG. 29B is a diagram illustrating possible states (e.g., base calls) that may occur in sequencing data, and associated probabilities, of a normal allele in a mutated tumor gene, according to an illustrative embodiment.
[0154] FIG. 29C is a diagram illustrating possible states (e.g., base calls) that may occur in sequencing data, and associated probabilities, of a mutated tumor allele of a tumor genome, according to an illustrative embodiment.
[0155] FIG. 30 is a schematic illustrating an example construct encoding selected neoantigens according to an illustrative embodiment.
[0156] FIG. 31 is a block diagram of an exemplary cloud computing environment, used in certain embodiments.
[0157] FIG. 32 is a block diagram of an example computing device and an example mobile computing device used in certain embodiments.
[0158] FIG. 33A is a state diagram for sequencing an allele for a homozygous phenotype, according to an illustrative embodiment.
[0159] FIG. 33B is a state diagram for selecting and sequencing a major allele associated with a particular SNV, according to an illustrative embodiment.
[0160] FIG. 33C is a state diagram for selecting and sequencing a minor allele associated with a particular heterozygous SNP, according to an illustrative embodiment.
[0161] FIG. 34A is a graph plotting ML functions (empirical SNV likelihood and predicted analytical SNV likelihood) across values of μ for a melanoma tumor sample, wherein the round marker corresponds to the global maximum of L, and the filled and empty red triangles are the predicted global and local maxima of the analytical log likelihood function, according to an illustrative embodiment.
[0162] FIG. 34B is a graph plotting generalized likelihood ratio statistics for a KL divergence metric (GLRT. also referred to as DKLherein) across values of μ, wherein the round marker corresponds to the global maximum, for a melanoma tumor sample, according to an illustrative embodiment.- 30 - 13241894vlAttorney Docket. No.: 2013237-0970
[0163] FIG. 34C is a graph plotting observed and predicted probability distribution function estimations for putative minimal balanced SNV events with and without VAF filtering of outliers, for a melanoma tumor sample, according to an illustrative embodiment.
[0164] FIG. 34D is a graph plotting log likelihood of putative minimal balanced SNV events being minimal balanced SNV events for melanoma tumor sample, according to an illustrative embodiment.
[0165] FIG. 34E is a graph plotting predicted fractions of homozygous genotypes Pxx based on maximum likelihood (ML) estimation of genotype as a function of estimated normal contamination p for a melanoma tumor sample, according to an illustrative embodiment.
[0166] FIG. 35 is a graph showing lower and upper boundary regions as a function of a true purity (x-axis) delineating solution regions for an analytical SNV-likelihood function (y-axis).
[0167] FIG. 36 is a graph showing a landscape of predicted maxima for an analytical SNV-likelihood function as a function of true purity (x-axis) and true fraction of homozygous SNVs (y-axis).
[0168] FIG. 37A is a graph comparing ML functions (empirical SNV log likelihood -blue and analytical SNV log likelihood - red) from Monte Carlo simulations of Region A of FIG.36, according to an illustrative embodiment.
[0169] FIG. 37B is a graph comparing ML functions (empirical SNV log likelihood -blue and analytical SNV log likelihood - red) from Monte Carlo simulations of Region B of FIG.36, according to an illustrative embodiment.
[0170] FIG. 37C is a graph comparing ML functions (empirical SNV log likelihood -blue and analytical SNV log likelihood - red) from Monte Carlo simulations of Region C of FIG.36, according to an illustrative embodiment.
[0171] FIG. 37D is a graph comparing ML functions (empirical SNV log likelihood -blue and analytical SNV log likelihood - red) from Monte Carlo simulations of Region D of FIG. 36, according to an illustrative embodiment.- 31 - 13241894vlAttorney Docket. No.: 2013237-0970
[0172] FIG. 37E is a graph comparing ML functions (empirical SNV log likelihood -blue and analytical SNV log likelihood - red) from Monte Carlo simulations of Region E of FIG.36, according to an illustrative embodiment.
[0173] FIG. 38 is a graph showing differentiation between a global maximum of an analytical SNV-likelihood function and a potential secondary maximum across a tzo, pxx) parameter space, according to an illustrative embodiment.
[0174] FIG. 39A is a graph showing ML functions (empirical SNV log likelihood - blue and analytical SNV log likelihood - red) having multiple peaks across values of p for a tumor sampled at 5000x, according to an illustrative embodiment.
[0175] FIG. 39B is a graph plotting KL divergence metrics (with outlier removal - blue and without outlier removal - red) having a sharp global maximum across values of p for a tumor sampled at 5000x, according to an illustrative embodiment.
[0176] FIG. 39C is a graph plotting ML functions (empirical SNV log likelihood - blue and predicted analytical SNV log likelihood - red) having multiple peaks across values of p for a tumor sampled at lOOx, according to an illustrative embodiment.
[0177] FIG. 39D is a graph plotting KL divergence metrics (with outlier removal - blue and without outlier removal - red) having a sharp global maximum across values of p for a tumor sampled at lOOx, according to an illustrative embodiment.
[0178] FIG. 40A is a graph showing results of a maximum likelihood SNV-based purity estimator for simulated tumor samples with various levels of normal contamination, p, and various fractions of homozygous SNVs, Pxx for number of SNVs L = 50 and depth of coverage C = 400x, according to an illustrative embodiment.
[0179] FIG. 40B is a graph showing results of a KL divergence SNV-based purity estimator for simulated tumor samples with various levels of normal contamination, p, and various fractions of homozygous SNVs, Pxx for number of SNVs L = 50 and depth of coverage C = 400x, according to an illustrative embodiment.
[0180] FIG. 40C is a graph showing results of a combined maximum likelihood and KL divergence SNV-based purity estimator for simulated tumor samples with various levels of- 32 - 13241894vlAttorney Docket. No.: 2013237-0970normal contamination, p, and various fractions of homozygous SNVs, Pxx for number of SNVs L = 50 and depth of coverage C = 400x, according to an illustrative embodiment.
[0181] FIG. 40D is a graph showing results of a maximum likelihood SNV-based purity estimator for simulated tumor samples with various levels of normal contamination, p, and various fractions of homozygous SNVs, Pxx for number of SNVs L = 50 and depth of coverage C = 200x, according to an illustrative embodiment.
[0182] FIG. 40E is a graph showing results of a KL divergence SNV-based purity estimator for simulated tumor samples with various levels of normal contamination, p, and various fractions of homozygous SNVs, Pxx for number of SNVs L = 50 and depth of coverage C = 200x, according to an illustrative embodiment.
[0183] FIG. 40F is a graph showing results of a combined maximum likelihood and KL divergence SNV-based purity estimator for simulated tumor samples with various levels of normal contamination, p, and various fractions of homozygous SNVs, Pxx for number of SNVs L = 50 and depth of coverage C = 200x, according to an illustrative embodiment.
[0184] FIG. 40G is a graph showing results of a maximum likelihood SNV-based purity estimator for simulated tumor samples with various levels of normal contamination, p, and various fractions of homozygous SNVs, Pxx for number of SNVs L = 20 and depth of coverage C = 200x, according to an illustrative embodiment.
[0185] FIG. 40H is a graph showing results of a KL divergence SNV-based purity estimator for simulated tumor samples with various levels of normal contamination, p, and various fractions of homozygous SNVs, pxx for number of SNVs L = 20 and depth of coverage C = 200x, according to an illustrative embodiment.
[0186] FIG. 401 is a graph showing results of a combined maximum likelihood and KL divergence SNV-based purity estimator for simulated tumor samples with various levels of normal contamination, p, and various fractions of homozygous SNVs, pxx for number of SNVs L = 20 and depth of coverage C = 200x, according to an illustrative embodiment.
[0187] FIG. 41A is a series of plots showing results for SNV-based purity estimation for a melanoma sample, according to an illustrative embodiment.- 33 - 13241894vlAttorney Docket. No.: 2013237-0970
[0188] FIG. 41B is a CNV cluster plot showing results for CNV-based purity estimation for the melanoma sample, according to an illustrative embodiment.
[0189] FIG. 42 is a block flow diagram showing an exemplary process for determining a putative mutation (e.g., an initial event) based on extreme statistics.
[0190] FIG. 43A is a line graph showing true positive rate as a function of coverage depth for an exemplary extreme event detector (EED) classifier and an exemplary binomial classifier for simulated heterozygous SNPs.
[0191] FIG. 43B is a line graph showing a false discovery rate (number of wrong genotype detection events over total detection events) as a function of coverage depth for an exemplary EED classifier and an exemplary binomial classifier for simulated heterozygous SNPs.
[0192] FIG. 43C is a line graph showing false discovery rate (false positives over total detection events) as a function of coverage for an exemplary EED classifier and an exemplary binomial classifier for simulated wild-type sites.
[0193] FIG. 44 is a histogram of percentages of histological tumor content (HTC) of a cohort of 28 patient tumor samples used for a benchmark analysis.
[0194] FIG. 45 is a bar graph showing estimated purity according to certain illustrative embodiments (an initial tumor model and a final tumor model) across a cohort of 28 patient tumor samples.
[0195] FIG. 46A is a bar graph showing number of mutations estimated in patient tumor samples 1 - 16 (of 28) according to certain illustrative embodiments.
[0196] FIG. 46B is a bar graph showing number of mutations estimated in patient tumor samples 17-27 (of 28), according to certain illustrative embodiments.
[0197] FIG. 46C is a bar graph showing number of mutations estimated in patient tumor sample 28 (of 28), according to certain illustrative embodiments.
[0198] FIG. 47A is a bar graph showing fold change in sensitivity for an illustrative embodiment (Case 1) relative to another illustrative embodiment (Baseline) estimated for a cohort of 28 patient tumor samples.- 34 - 13241894vlAttorney Docket. No.: 2013237-0970
[0199] FIG. 47B is a bar graph showing percentage of sensitivity boost (i.e., increase) for an illustrative embodiment (Case 2) relative to another illustrative embodiment (Case 1) estimated for a cohort of 28 patient tumor samples.
[0200] FIG. 47C is a bar graph showing percentage of sensitivity boost for an illustrative embodiment (Case 3) relative to another illustrative embodiment (Case 2) estimated for a cohort of 28 patient tumor samples.
[0201] FIG. 47D is a bar graph showing percentage of sensitivity boost for an illustrative embodiment (Case 4) relative to another illustrative embodiment (Case 3) estimated for a cohort of 28 patient tumor samples.
[0202] FIG. 48 is a graph showing percentage of sensitivity boost for an illustrative embodiment (Case 4) relative to another illustrative embodiment (Case 1) estimated for a cohort of 28 patient tumor samples.
[0203] FIG. 49A is a box and whisker plot showing fold change in sensitivity for an illustrative embodiment (Case 1) relative to another illustrative embodiment (Baseline) estimated for a cohort of 28 patient tumor samples.
[0204] FIG. 49B is a box and whisker plot showing percentages of sensitivity boost for certain illustrative embodiments relative to other illustrative embodiments estimated for a cohort of 28 patient tumor samples.
[0205] FIG. 49C is a box and whisker plot showing percentage of sensitivity boost for an illustrative embodiment (Case 4) relative to another illustrative embodiment (Case 1) estimated for a cohort of 28 patient tumor samples.
[0206] FIG. 49D is a box and whisker plot showing fold change in sensitivity for an illustrative embodiment (Case 4) relative to another illustrative embodiment (Baseline) estimated for a cohort of 28 patient tumor samples.
[0207] FIG. 50A is a scatter plot showing number of false positives identified with illustrative embodiments of the present disclosure and published mutation callers.
[0208] FIG. 50B is a line graph showing PPV as a function of TMB given a number of false positives.- 35 - 13241894vlAttorney Docket. No.: 2013237-0970
[0209] FIG. 50C is a set of bar plots showing number of false positives per patient sample identified with an illustrative embodiment (left) or with previously published mutation callers MuTect (middle) or Strelka (right), using singletons.
[0210] FIG. 51A is a bar plot showing false discovery rate (FDR[%]) for four patient samples according to an illustrative embodiment, MuTect, or Strelka, using singletons.
[0211] FIG. 51B is a bar plot showing number of true positives (TPs) (left y axis) and total predicted SNVs (right y axis) for four patient samples according to an illustrative embodiment, MuTect, or Strelka, using singletons.
[0212] FIG. 51C is a bar plot showing false discover rate (FDR[%]) for four patient samples according to an illustrative embodiment using replicates, an illustrative embodiment using merged replicates, MuTect using merged replicates, or Strelka using merged replicates.
[0213] FIG. 51D is a bar plot showing number of true positives (TPs) (left y axis) and total predicted SNVs (right y axis) for four patient samples according to an illustrative embodiment using replicates, an illustrative embodiment using merged replicates, Mutect using merged replicates, or Strelka using merged replicates.
[0214] FIG. 51E is a bar plot showing number of false positives identified in patient sample P006 according to an illustrative embodiment, Mutect, or Strelka, using singletons.
[0215] FIG. 51F is a bar plot showing number of false positives identified in patient sample P006 according to an illustrative embodiment using replicates, an illustrative embodiment using merged replicates, Mutect using merged replicates, or Strelka using merged replicates.
[0216] FIG. 52A is a table showing results from evaluation of PPV of seven FFPE tumor samples using a 50bp-HiSeq4000 pipeline, according to an illustrative embodiment.
[0217] FIG. 52B is a table showing results from evaluation of PPV of seven FFPE tumor samples using a 100bp-NovaSeq6000 pipeline, according to an illustrative embodiment.
[0218] FIG. 53A is a scatter plot showing estimated purity (y-axis) as a function of simulated purity (x-axis) for cell line sample admixtures (lOObp).- 36 - 13241894vlAttorney Docket. No.: 2013237-0970
[0219] FIG. 53B is a scatter plot showing estimated purity (y-axis) as a function of simulated purity (x-axis) for FFPE sample admixtures (lOObp).
[0220] FIG. 54A is a scatter plot showing estimated ploidy as a function of simulated histological tumor content (HTC) for cell line sample admixtures (lOObp), according to an illustrative embodiment.
[0221] FIG. 54B is a scatter plot showing estimated ploidy as a function of simulated histological tumor content (HTC) for FFPE sample admixtures (lOObp), according to an illustrative embodiment.
[0222] FIG. 55 is a bar plot showing estimated ploidy according to an illustrative embodiment compared to ploidy estimated with FACS for raw cell line samples (lOObp).
[0223] FIG. 56 is a bar plot showing estimated ploidy according to an illustrative embodiment, compared to ploidy estimated with FACs and SKY for raw cell line samples (50bp).
[0224] FIG. 57A is a box and whisker plot showing sensitivity (assessed by relative number of mutations present in reference) as a function of simulated HTC (sHTC) for cell line sample admixtures (lOObp), according to an illustrative embodiment.
[0225] FIG. 57B is a box and whisker plot showing sensitivity (as assessed by relative number of mutations present in reference) as a function of simulated HTC (sHTC) for FFPE sample admixtures (lOObp), according to an illustrative embodiment.
[0226] FIG. 57C is a box and whisker plot showing uniqueness (as assessed by relative number of mutations absent in reference) as a function of sHTC for cell line sample admixtures (lOObp), according to an illustrative embodiment.
[0227] FIG. 57D is a box and whisker plot showing uniqueness (as assessed by relative number of filtered mutations absent in reference) as a function of sHTC for FFPE sample admixtures (lOObp), according to an illustrative embodiment.
[0228] FIG. 58A is a box and whisker plot showing sensitivity (as assessed by relative number of mutations present in reference) as a function of sHTC for cell line sample admixtures (50bp), according to an illustrative embodiment.- 37 - 13241894vlAttorney Docket. No.: 2013237-0970
[0229] FIG. 58B is a box and whisker plot showing uniqueness (as assessed by relative number of mutations absent in reference) as a function of sHTC for cell line sample admixtures (50bp), according to an illustrative embodiment.
[0230] FIG. 59A is a plot showing sensitivity for all mutations (as assessed by relative number of mutations present in reference) as a function of simulated purity for FFPE sample admixtures (lOObp), according to an illustrative embodiment.
[0231] FIG. 59B is a plot showing sensitivity for mutations in exon coding regions (as assessed by relative number of mutations present in reference) as a function of simulated purity for FFPE sample admixtures (lOObp), according to an illustrative embodiment.
[0232] FIG. 59C is a plot showing sensitivity for mutations in exon coding regions determined to have a CLONAL clone type (as assessed by relative number of mutations present in reference) as a function of simulated purity for FFPE sample admixtures (lOObp), according to an illustrative embodiment.
[0233] FIG. 60A is a bar plot showing number of detected mutations in raw FFPE samples (lOObp) by clone type classification, with numbers on top of bars indicating total number of detected mutations per sample, and percentages indicating proportion of mutations classified as CLONAL, according to an illustrative embodiment.
[0234] FIG. 60B is a bar plot showing number of detected mutations in raw cell line samples (lOObp) by clone type classification, with numbers on top of bars indicating total number of detected mutations per sample, and percentages overlaid on bars indicating proportion of mutations classified as CLONAL, according to an illustrative embodiment.
[0235] FIG. 61 is a bar plot showing number of detected mutations in raw cell line samples (50bp) by clone type classification, with numbers on top of bars indicating total number of detected mutations per sample, and percentages overlaid on bars indicating proportion of mutations classified as CLONAL, according to an illustrative embodiment.
[0236] FIG. 62 is a series of heatmaps showing clone type classification (“Clonality State”) of mutations detected in certain FFPE raw samples and admixtures (all lOObp), wherein each panel corresponds to data from a patient sample, according to an illustrative embodiment.- 38 - 13241894vlAttorney Docket. No.: 2013237-0970
[0237] FIG. 63 is a series of heatmaps showing clone type classification (“Clonality State”) of mutations detected in certain raw cell line samples and admixtures, wherein each panel corresponds to results from a cell line, according to an illustrative embodiment.
[0238] FIGs. 64A-E is a series of heatmaps showing clone type classification (“Clonality State”) of mutations detected in certain FFPE raw samples and admixtures (all 50bp), wherein each panel corresponds to results from a patient sample, according to an illustrative embodiment.
[0239] FIG. 65A is a scatter plot showing CLONAL scores as a function of simulated purity for FFPE sample admixtures (lOObp), according to an illustrative embodiment.
[0240] FIG. 65B is a scatter plot showing Clonality Scores as a function of simulated purity for FFPE sample admixtures (lOObp), according to an illustrative embodiment.
[0241] FIG. 65C is a scatter plot showing CLONAL scores as a function of simulated purity for cell line sample admixtures (lOObp), according to an illustrative embodiment.
[0242] FIG. 65D is a scatter plot showing Clonality Scores as a function of simulated purity for cell line sample admixtures (lOObp), according to an illustrative embodiment.
[0243] FIG. 66A is a scatter plot showing CLONAL scores as a function of simulated purity for cell line sample admixtures (50bp), according to an illustrative embodiment.
[0244] FIG. 66B is a box and whisker plot showing CLONAL scores as a function of simulated purity for cell line sample admixtures (lOObp), according to an illustrative embodiment.
[0245] FIG. 66C is a scatter plot showing Clonality Scores as a function of simulated purity for cell line sample admixtures (50bp), according to an illustrative embodiment.
[0246] FIG. 66D is a box and whisker plot showing Clonality Scores as a function of simulated purity for cell line sample admixtures (50bp), according to an illustrative embodiment.
[0247] FIG. 67A is a bar plot showing events predicted by MuTect2 across four permutations (bars) and split (colors) according to whether events were reproduced between permutations.- 39 - 13241894vlAttorney Docket. No.: 2013237-0970
[0248] FIG. 67B is a bar plot showing events predicted by an illustrative embodiment across four permutations (bars) and split (colors) according to whether events were reproduced between permutations.
[0249] FIG. 67C is a histogram showing a distribution of VAFs for events detected by both the illustrative embodiment and MuTect2 (“Shared”), detected by only MuTect2, and detected only by the illustrative embodiment.
[0250] FIG. 68 is a graph and a heatmap both showing cellularities and identifying clone type assignments for an experimental admixture series of a bulk tumor cell line at varying sample purities, according to an illustrative embodiment.
[0251] FIG. 69 is a block flow diagram of an example approach for generating experimental samples with controlled mutation properties for validation of mutation detection technologies in accordance with embodiments described herein.
[0252] FIG. 70 is a plot showing mutation lineages and corresponding clone type classifications determined via an illustrative embodiment of technologies described herein.
[0253] FIG. 71A is a graph plotting percentages of experimentally created biologically subclonal (and low prevalence) mutations that were rejected by a rare subclone filter at varying sample purities, according to an illustrative embodiment.
[0254] FIG. 71B is a graph plotting percentages of experimentally created biologically subclonal (and low prevalence) mutations that were assigned to a minor subclone state at varying sample purities, according to an illustrative embodiment.
[0255] FIG. 72A is a graph plotting percentages of experimentally created biologically subclonal (and low prevalence) mutations that were detected via an example mutation caller, according to an illustrative embodiment.
[0256] FIG. 72B is a graph plotting percentages of experimentally created biologically subclonal (and low prevalence) mutations that were detected via an example mutation caller, but rejected by subsequent filtering, according to an illustrative embodiment- 40 - 13241894vlAttorney Docket. No.: 2013237-0970
[0257] FIG. 73A is a graph plotting percentages of experimentally created biologically subclonal (and low prevalence) mutations that were detected and rejected via a rare subclone filter, according to an illustrative embodiment.
[0258] FIG. 73B is a graph plotting total numbers of experimentally created biologically subclonal (and low prevalence) mutations that were detected and classified as low-confidence mutations via one or more filters, according to an illustrative embodiment.
[0259] FIG. 74A is a graph showing estimation of noise based on normal versus normal analysis of singletons according to an illustrative embodiment, Mutect2, or Strelka.
[0260] FIG. 74B is a graph showing estimation of noise based on normal versus normal analysis of replicates or merged singletons according to an illustrative embodiment, Mutect2, or Strelka.
[0261] FIG. 75A is a graph showing predicted PPV as a function of TMB based on normal versus normal analysis of singletons according to an illustrative embodiment, Mutect2, or Strelka (including confidence intervals).
[0262] FIG. 75B is a graph showing predicted PPV as a function of TMB based on normal versus normal analysis of replicates or merged singletons according to an illustrative embodiment, Mutect2, or Strelka (including confidence intervals).
[0263] FIG. 76A is a bar graph showing number of SNVs detected across permutations according to an illustrative embodiment or Mutect2.
[0264] FIG. 76B is a bar graph showing number of SNVs (shared and unique variants) detected across permutations according to an illustrative embodiment or Mutect2.
[0265] FIG. 76C is a histogram showing cellularity distributions for SNVs detected by Mutect2 only (“Mutect2 unique”), an illustrative embodiment only (“Embodiment unique”), or both (“Shared”).
[0266] FIG. 77A is a scatter plot with unity line plotting estimated purity as a function of simulated purity across patient samples, according to an illustrative embodiment.
[0267] FIG. 77B is a scatter plot with unity line plotting estimated purity as a function of simulated HTC across patient samples, according to an illustrative embodiment.- 41 - 13241894vlAttorney Docket. No.: 2013237-0970
[0268] FIG. 77C is a scatter plot with unity line plotting simulated HTC as a function of simulated purity across patient samples, according to an illustrative embodiment.
[0269] FIG. 78A is a box and whisker plot showing relative sensitivity percentage as a function of simulated HTC across patient samples according to an illustrative embodiment in singleton mode.
[0270] FIG. 78B is a box and whisker plot showing relative sensitivity percentage as a function of simulated HTC across patient samples according to Mutect2 in singleton mode.
[0271] FIG. 78C is a box and whisker plot showing relative sensitivity percentage as a function of simulated HTC across patient samples according to Strelka2 in singleton mode.
[0272] FIG. 79A is a plot showing relative sensitivity percentage as a function of simulated purity across patient samples according to an illustrative embodiment in singleton mode.
[0273] FIG. 79B is a plot showing relative sensitivity percentage as a function of simulated purity across patient samples according to Mutect2 in singleton mode.
[0274] FIG. 79C is a plot showing relative sensitivity percentage as a function of simulated purity across patient samples according to Strelka2 in singleton mode.
[0275] FIG. 80A is a box and whisker plot showing relative sensitivity percentage as a function of simulated HTC across patient samples according to an illustrative embodiment in replicate mode.
[0276] FIG. 80B is a box and whisker plot showing relative sensitivity percentage as a function of simulated HTC across patient samples according to Mutect2 in replicate mode.
[0277] FIG. 80C is a box and whisker plot showing relative sensitivity percentage as a function of simulated HTC across patient samples according to Strelka2 in singleton merged mode.
[0278] FIG. 81A is a plot showing relative sensitivity percentage as a function of simulated purity across patient samples according to an illustrative embodiment in replicate mode.- 42 - 13241894vlAttorney Docket. No.: 2013237-0970
[0279] FIG. 81B is a plot showing relative sensitivity percentage as a function of simulated purity across patient samples according to Mutect2 in replicate mode.
[0280] FIG. 81C is a plot showing relative sensitivity percentage as a function of simulated purity across patient samples according to Strelka2 in singleton merged mode.
[0281] FIG. 82A is a graph showing relative sensitivity percentage (mean + / - standard deviation) as a function of simulated HTC according to an illustrative embodiment in singleton mode, Mutect2 in singleton mode, or Strelka2 in singleton mode.
[0282] FIG. 82B is a graph showing relative sensitivity percentage (mean + / - standard deviation) as a function of simulated HTC according to an illustrative embodiment in replicate mode, Mutect2 in replicate mode, or Strelka2 in singleton merged mode.
[0283] FIG. 83A is a graph showing relative sensitivity percentage (mean + / - standard deviation) as a function of simulated purity according to an illustrative embodiment in singleton mode, Mutect2 in singleton mode, or Strelka2 in singleton mode.
[0284] FIG. 83B is a graph showing relative sensitivity percentage (mean + / - standard deviation) as a function of simulated purity according to an illustrative embodiment in replicate mode, Mutect2 in replicate mode, or Strelka2 in singleton merged mode.
[0285] FIG. 84A is a box and whisker plot showing number of mutations unique to admixture as a function of simulated HTC across patient samples according to an illustrative embodiment in singleton mode.
[0286] FIG. 84B is a box and whisker plot showing number of mutations unique to admixture as a function of simulated HTC across patient samples according to Mutect2 in singleton mode.
[0287] FIG. 84C is a box and whisker plot showing number of mutations unique to admixture as a function of simulated HTC across patient samples according to Strelka2 in singleton mode.
[0288] FIG. 85A is a box and whisker plot showing number of mutations unique to admixture as a function of simulated HTC across patient samples according to an illustrative embodiment in replicate mode.- 43 - 13241894vlAttorney Docket. No.: 2013237-0970
[0289] FIG. 85B is a box and whisker plot showing number of mutations unique to admixture as a function of simulated HTC across patient samples according to Mutect2 in replicate mode.
[0290] FIG. 85C is a box and whisker plot showing number of mutations unique to admixture as a function of simulated HTC across patient samples according to Strelka2 in singleton merged mode.
[0291] FIG. 86A is a graph showing number of mutations unique to admixture (mean + / -standard deviation) as a function of simulated HTC according to an illustrative embodiment in singleton mode, Mutect2 in singleton mode, or Strelka2 in singleton mode.
[0292] FIG. 86B is a graph showing number of mutations unique to admixture (mean + / -standard deviation) as a function of simulated HTC according to an illustrative embodiment in replicate mode, Mutect2 in replicate mode, or Strelka2 in singleton merged mode.
[0293] FIG. 87A is a box and whisker plot showing a Jaccard index as a function of simulated HTC across patient samples according to an illustrative embodiment in singleton mode.
[0294] FIG. 87B is a box and whisker plot showing a Jaccard index as a function of simulated HTC across patient samples according to Mutect2 in singleton mode.
[0295] FIG. 87C is a box and whisker plot showing a Jaccard index as a function of simulated HTC across patient samples according to Strelka2 in singleton mode.
[0296] FIG. 88A is a plot showing a Jaccard index as a function of simulated purity across patient samples according to an illustrative embodiment in singleton mode.
[0297] FIG. 88B is a plot showing a Jaccard index as a function of simulated purity across patient samples according to Mutect2 in singleton mode.
[0298] FIG. 88C is a plot showing a Jaccard index as a function of simulated purity across patient samples according to Strelka2 in singleton mode.
[0299] FIG. 89A is a box and whisker plot showing number of mutations unique to one set as a function of simulated HTC across patient samples according to an illustrative embodiment in singleton mode.- 44 - 13241894vlAttorney Docket. No.: 2013237-0970
[0300] FIG. 89B is a box and whisker plot showing number of mutations unique to one set as a function of simulated HTC across patient samples according to Mutect2 in singleton mode.
[0301] FIG. 89C is a box and whisker plot showing number of mutations unique to one set as a function of simulated HTC across patient samples according to Strelka2 in singleton mode.
[0302] FIG. 90A is a plot showing number of mutations unique to one set as a function of simulated HTC across patient samples according to an illustrative embodiment in singleton mode.
[0303] FIG. 90B is plot showing number of mutations unique to one set as a function of simulated HTC across patient samples according to Mutect2 in singleton mode.
[0304] FIG. 90C is a plot showing number of mutations unique to one set as a function of simulated HTC across patient samples according to Strelka2 in singleton mode.
[0305] FIG. 91 is a bar plot showing number of mutations unique to one set of cross replicates (average and standard deviation) according to an illustrative embodiment, Mutect2, and Strelka2, based on cross replicate analysis for nine patients.
[0306] FIG. 92A is a box and whisker plot showing relative sensitivity percentage as a function of simulated HTC across patient samples according to an illustrative embodiment in replicate mode.
[0307] FIG. 92B is a box and whisker plot showing relative sensitivity percentage as a function of simulated HTC across patient samples according to Mutect2 in replicate mode filtered with VAF >=0.05.
[0308] FIG. 92C is a box and whisker plot showing relative sensitivity percentage as a function of simulated HTC across patient samples according to Strelka2 in singleton merged mode filtered with VAF >=0.05.
[0309] FIG. 93A is a plot showing relative sensitivity percentage as a function of simulated purity across patient samples according to an illustrative embodiment in replicate mode.- 45 - 13241894vlAttorney Docket. No.: 2013237-0970
[0310] FIG. 93B is a plot showing relative sensitivity percentage as a function of simulated purity across patient samples according to Mutect2 in replicate mode filtered with VAF >=0.05.
[0311] FIG. 93C is a plot showing relative sensitivity percentage as a function of simulated purity across patient samples according to Strelka2 in singleton merged mode filtered with VAF >=0.05.
[0312] FIG. 94 is a graph showing relative sensitivity percentage (mean + / - standard deviation) as a function of simulated purity according to an illustrative embodiment in replicate mode, Mutect2 in replicate mode filtered with VAF>=0.05, or Strelka2 in singleton merged mode filtered with VAF>=0.05.
[0313] The features and advantages of the present disclosure will become more apparent from the detailed description set forth below when taken in conjunction with the drawings, in which like reference characters identify corresponding elements throughout. In the drawings, like reference numbers generally indicate identical, functionally similar, and / or structurally similar elements.CERTAIN DEFINITIONS
[0314] About'. The term “about”, when used herein in reference to a value, refers to a value that is similar, in context to the referenced value. In general, those skilled in the art, familiar with the context, will appreciate the relevant degree of variance encompassed by “about” in that context. For example, in some embodiments, the term “about” may encompass a range of values that within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or less of the referred value.
[0315] Absolute Copy Number'. The term “absolute copy number” as used herein refers to a number of physical copies of a particular segment within a cell comprising a particular genome. For example, an absolute copy number of a segment in the normal genome can be defined as the number of physical copies of the given segment in a healthy cell. For example, an absolute copy number of a segment in the tumor genome can be defined as the number of physical copies of the given segment in a tumor cell. In certain embodiments, if only a part of a - 46 - 13241894vlAttorney Docket. No.: 2013237-0970segment is amplified or deleted in a genome, then such a partial copy of the segment can either be counted as a copy of the segment or not counted as a copy of the segment. In certain embodiments, copies of the segment spanning less than 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10% or 5% of the segment length can be ignored.
[0316] Agent-. As used herein, the term “agent,” may refer to a physical entity. In some embodiments, an agent may be characterized by a particular feature and / or effect. For example, as used herein, the term “therapeutic agent” refers to a physical entity has a therapeutic effect and / or elicits a desired biological and / or pharmacological effect. In some embodiments, an agent may be a compound, molecule, or entity of any chemical class including, for example, a small molecule, polypeptide, nucleic acid, saccharide, lipid, metal, or a combination or complex thereof. In some embodiments, part or all of an agent may be depicted herein as a chemical structure, or may be described using chemical nomenclature and / or with reference to general principles of organic chemistry, e.g., in accordance with the Periodic Table of Elements, CAS version, Handbook of Chemistry and Physics, 75th Ed; “Organic Chemistry”, Thomas Sorrell, University Science Books, Sausalito: 1999, and / or “March’s Advanced Organic Chemistry”, 5th Ed., Ed.: Smith, M. B. and March, J., John Wiley & Sons, New York: 2001, the entire contents of which are hereby incorporated by reference. Unless otherwise stated or clear from context, chemical structures depicted herein may be considered to reference or include one or more, or all, stereoisomeric (e.g., enantiomeric or diastereomeric) forms of the structure, and / or one or more, or all, geometric or conformational isomeric forms of the structure. For example, unless otherwise indicated or clear, both R and S configurations of a stereocenter may be contemplated in embodiments of the disclosure. In some embodiments, a compound may be described and / or utilized as a particular single stereochemical isomer; alternatively or additionally, in some embodiments, such a compound may be described and / or utilized as a combination e.g., a mixture) of one or more enantiomeric (e.g., diastereomeric) forms (e.g., as a racemic preparation). Analogously, in some embodiments, a single geometric isomer may be described and / or utilized; in some embodiments, a combination (e.g., a mixture) of geometric (or conformational) isomers may be described and / or utilized. Unless otherwise stated or clear from context, all tautomeric forms of provided compounds are within the scope of the disclosure. Still further, unless otherwise indicated or clear from context, in some embodiments, a particular chemical compound (e.g., as may be represented by a depicted chemical structure) may be - 47 - 13241894vlAttorney Docket. No.: 2013237-0970described and / or utilized in an alternative isotopic form - i.e., in a form in which one or more atoms is isotopically altered (e.g., so that a hydrogen is replaced by deuterium or tritium, and / or a carbon is replaced by 13C- or 14C-. Thus, in some embodiments, a particular compound may be described and / or utilized as or in an isotopically enriched preparation.
[0317] Allele-Specific Copy Number: As used herein, the term “allele- specific copy number” refers to a number of physical copies of a particular allele within a cell comprising a particular genome. In certain embodiments, for example, a heterozygous segment comprises a SNP. A cell comprising a genome with the heterozygous segment may comprise zero, one, or more copies of a first (e.g., maternal) allele and zero, one, or more copies of a second (e.g., paternal) allele. For example, a balanced heterozygous segment in a normal diploid genome may comprise distinguishable maternal and paternal alleles, each having an allele-specific copy number of one. In certain embodiments, a corresponding heterozygous segment within a tumor genome may not have zero, one, or more copies of the paternal and maternal alleles, e.g., as a result of copy number variation (CNV) events that may occur in cancer cells. For example, a loss of heterozygosity (LOH) may result in a deletion of copies of one allele (e.g., a maternal or paternal allele), such that an allele- specific copy number for the deleted allele is zero and, if a single copy of the other allele remains, its allele- specific copy number is one. In certain embodiments, duplication events may produce other allele-specific copy numbers, greater than one for one or both alleles.
[0318] Amino acid'. In its broadest sense, as used herein, the term “amino acid” refers to a compound and / or substance that can be, is, or has been incorporated into a polypeptide chain, e.g., through formation of one or more peptide bonds. In some embodiments, an amino acid has the general structure H2N–C(H)(R)–COOH. In some embodiments, an amino acid is a naturally-occurring amino acid. In some embodiments, an amino acid is a non-natural amino acid; in some embodiments, an amino acid is a D-amino acid; in some embodiments, an amino acid is an L-amino acid. “Standard amino acid” refers to any of the twenty standard L-amino acids commonly found in naturally occurring peptides. “Nonstandard amino acid” refers to any amino acid, other than the standard amino acids, regardless of whether it is prepared synthetically or obtained from a natural source. In some embodiments, an amino acid, including a carboxy- and / or amino-terminal amino acid in a polypeptide, can contain a structural- 48 - 13241894vlAttorney Docket. No.: 2013237-0970modification as compared with the general structure above. For example, in some embodiments, an amino acid may be modified by methylation, amidation, acetylation, pegylation, glycosylation, phosphorylation, and / or substitution (e.g., of the amino group, the carboxylic acid group, one or more protons, and / or the hydroxyl group) as compared with the general structure. In some embodiments, such modification may, for example, alter the circulating half-life of a polypeptide containing the modified amino acid as compared with one containing an otherwise identical unmodified amino acid. In some embodiments, such modification does not significantly alter a relevant activity of a polypeptide containing the modified amino acid, as compared with one containing an otherwise identical unmodified amino acid. As will be clear from context, in some embodiments, the term “amino acid” may be used to refer to a free amino acid; in some embodiments it may be used to refer to an amino acid residue of a polypeptide.
[0319] Antigen -, term “antigen”, as used herein, refers to (i) an agent that elicits an immune response; and / or (ii) an agent that binds to a T cell receptor (e.g., when presented by an MHC molecule) or to an antibody. In some embodiments, an antigen elicits a humoral response (e.g., including production of antigen-specific antibodies); in some embodiments, an elicits a cellular response e.g., involving T-cells whose receptors specifically interact with the antigen). In some embodiments, and antigen binds to an antibody and may or may not induce a particular physiological response in an organism. In general, an antigen may be or include any chemical entity such as, for example, a small molecule, a nucleic acid, a polypeptide, a carbohydrate, a lipid, a polymer [in some embodiments other than a biologic polymer (e.g., other than a nucleic acid or amino acid polymer)] etc. In some embodiments, an antigen is or comprises a polypeptide. In some embodiments, an antigen is or comprises a glycan. Those of ordinary skill in the art will appreciate that, in general, an antigen may be provided in isolated or pure form, or alternatively may be provided in crude form (e.g., together with other materials, for example in an extract such as a cellular extract or other relatively crude preparation of an antigen-containing source). In some embodiments, antigens utilized in accordance with the present invention are provided in a crude form. In some embodiments, an antigen is a recombinant antigen.
[0320] Biologically clonal mutation(s) and biologically subclonal mutation(s): As used herein, the terms “biologically clonal” and “biologically subclonal” when used in reference to mutations, such as cancer mutations, are used to specify whether a particular mutations or group- 49 - 13241894vlAttorney Docket. No.: 2013237-0970of mutations are physically clonal or subclonal. In certain embodiments, a biologically clonal mutation is a mutation that is present in all tumor cells of a tumor sample or biopsy. In certain embodiments, a biologically subclonal mutation is a mutation that is not present in all tumor cells of a tumor sample or biopsy. The use of the adjective “biologically” is used to make clear that the terms “biologically clonal” and “biologically subclonal” refer to the actual physical character of a given mutation, which may or may not be known. The terms “biologically clonal” and “biologically subclonal” thus contrast with the terms “clone type”, “clonal state”, “prevalent subclone state”, and “minor subclone state”, described below, which refer to clonality classification states that are, e.g., labels, determined for (e.g., assigned to) a given mutation.
[0321] Cancer. The term “cancer” is used herein to generally refer to a disease or condition in which cells of a tissue of interest exhibit relatively abnormal, uncontrolled, and / or autonomous growth, so that they exhibit an aberrant growth phenotype characterized by a significant loss of control of cell proliferation. In some embodiments, cancer may comprise cells that are precancerous e.g., benign), malignant, pre-metastatic, metastatic, and / or non-metastatic. In some embodiments, cancer may be characterized by a solid tumor. In some embodiments, cancer may be characterized by a hematologic tumor. In general, examples of different types of cancers known in the art include, for example, triple negative breast cancer (TNBC), hematopoietic cancers including leukemias, lymphomas (Hodgkin’s and non-Hodgkin’s), myelomas and myeloproliferative disorders; sarcomas, melanomas, adenomas, carcinomas of solid tissue, squamous cell carcinomas of the mouth, throat, larynx, and lung, liver cancer, genitourinary cancers such as prostate, cervical, bladder, uterine, and endometrial cancer and renal cell carcinomas, bone cancer, pancreatic cancer, skin cancer, cutaneous or intraocular melanoma, cancer of the endocrine system, cancer of the thyroid gland, cancer of the parathyroid gland, head and neck cancers, ovarian cancer, breast cancer, glioblastomas, colorectal cancer, gastro-intestinal cancers and nervous system cancers, benign lesions such as papillomas, and the like.
[0322] Clone type: As used herein the term “clone type” is used to refer to certain discrete states or classes that may be determined for and assigned to a given mutation, for example, based on evaluation of various criteria. Different clone types may aim to capture, for example, a likelihood that a given mutation is biologically clonal or biologically subclonal, as- 50 - 13241894vlAttorney Docket. No.: 2013237-0970well as, for example, a prevalence of the given mutation. In certain embodiments, a likelihood that a given mutation is biologically clonal may be assessed based on various metrics or tests, including, without limitation, statistical hypothesis tests, computed probabilities (e.g., of clonality or sub-clonality, such as a posteriori probability of clonality), etc. For example, in certain embodiments, a likelihood of whether a given mutation is biologically clonal may be assessed using metrics, such as P-values, determined using hypotheses tests, that quantify probabilities of whether for a given data (such as sequencing data) a hypothesis of clonality can be rejected. For example, in certain embodiments, sequencing data may be used to determine an observed cellularity of a given mutation (e.g., a fraction of tumor cells harboring the given mutation). Although, nominally, biologically clonal mutations have a cellularity of 1 (i.e., they are present in all tumor cells), factors, such as measurement error, sample purity, stochastic noise, etc., may lead to observed cellularity values below one, even for biologically clonal mutations. Accordingly, in certain embodiments, a P-value representing a probability of observing a cellularity below one for the given mutation (e.g., given a null hypothesis that the mutation is biologically clonal) may be determined. In this way, for example, a low P-value, e.g., below a threshold, may indicate that it is unlikely that the null hypothesis (of clonality) is true, and it can be rejected (hence unlikely that a mutation is clonal), whereas a higher P-value may indicate that there is a good chance that the observed cellularity, below 1, is due to practical factors, such as measurement error, sample purity, stochastic noise, and the like. In certain embodiments, prevalence may be measured by parameters, such as cellularity, e.g., a determined fraction of tumor cells that harbor a given mutation. Whereas the terms biologically clonal and biologically subclonal are used to refer to underlying physical characteristics of a given mutation, clone types are labels, which aim to classify mutations according to measured physical properties, determined, e.g., based on sequencing data. In certain embodiments, clone types include clonal states, which aim to capture or label mutations that are highly likely to be biologically clonal and / or highly prevalent. In certain embodiments, clone types include multiple clonal classification states, for example, reflecting differing levels of certainty and / or likelihoods that a given mutation is biologically clonal and / or prevalence. In certain embodiments, clone types include one or more prevalent subclone states, capturing mutations that, while biologically subclonal, are prevalent at high rates e.g., above 50%) in tumor cells. In certain embodiments, clone types include a minor subclonal state, which may be assigned to - 51 - 13241894vlAttorney Docket. No.: 2013237-0970mutations that are determined to be likely subclonal and rare e.g., occurring in 50% or less, 45% or less, 40% or less, 35% or less, or 25% or less cancer cells). In some embodiments, clone type may be determined as described in U. S. Provisional Patent Application No. 63 / 749,339, entitled “Technologies for Neoantigen Prioritization Based on Tumor Modeling and Clonality States,” and the PCT application having the same title filed January 23, 2026, the content of each of which is incorporated by reference herein in its entirety.
[0323] Comparable'. As used herein, the term “comparable” refers to two or more agents, entities, situations, sets of conditions, etc., that may not be identical to one another but that are sufficiently similar to permit comparison therebetween so that one skilled in the art will appreciate that conclusions may reasonably be drawn based on differences or similarities observed. In some embodiments, comparable sets of conditions, circumstances, individuals, or populations are characterized by a plurality of substantially identical features and one or a small number of varied features. Those of ordinary skill in the art will understand, in context, what degree of identity is required in any given circumstance for two or more such agents, entities, situations, sets of conditions, etc., to be considered comparable. For example, those of ordinary skill in the art will appreciate that sets of circumstances, individuals, or populations are comparable to one another when characterized by a sufficient number and type of substantially identical features to warrant a reasonable conclusion that differences in results obtained or phenomena observed under or with different sets of circumstances, individuals, or populations are caused by or indicative of the variation in those features that are varied.
[0324] Corresponding to: As used herein, the term “corresponding to” refers to a relationship between two or more entities. For example, the term “corresponding to” may be used to designate the position / identity of a structural element in a compound or composition relative to another compound or composition (e.g., to an appropriate reference compound or composition). For example, in some embodiments, a monomeric residue in a polymer (e.g., an amino acid residue in a polypeptide or a nucleic acid residue in a polynucleotide) may be identified as “corresponding to” a residue in an appropriate reference polymer. For example, those of ordinary skill will appreciate that, for purposes of simplicity, residues in a polypeptide are often designated using a canonical numbering system based on a reference related polypeptide, so that an amino acid “corresponding to” a residue at position 190, for example,- 52 - 13241894vlAttorney Docket. No.: 2013237-0970need not actually be the 190th amino acid in a particular amino acid chain but rather corresponds to the residue found at 190 in the reference polypeptide; those of ordinary skill in the art readily appreciate how to identify “corresponding” amino acids. For example, those skilled in the art will be aware of various sequence alignment strategies, including software programs such as, for example, BLAST, CS-BLAST, CUSASW++, DIAMOND, FASTA, GGSEARCH / GLSEARCH, Genoogle, HMMER, HHpred / HHsearch, IDF, Infernal, KLAST, USEARCH, parasail, PSI-BLAST, PSI-Search, ScalaBLAST, Sequilab, SAM, SSEARCH, SWAPHI, SWAPHLLS, SWIMM, or SWIPE that can be utilized, for example, to identify “corresponding” residues in polypeptides and / or nucleic acids in accordance with the present disclosure. Those of skill in the art will also appreciate that, in some instances, the term “corresponding to” may be used to describe an event or entity that shares a relevant similarity with another event or entity (e.g., an appropriate reference event or entity). To give but one example, a gene or protein in one organism may be described as “corresponding to” a gene or protein from another organism in order to indicate, in some embodiments, that it plays an analogous role or performs an analogous function and / or that it shows a particular degree of sequence identity or homology, or shares a particular characteristic sequence element.
[0325] Encode-. As used herein, the term “encode” or “encoding” refers to sequence information of a first molecule that guides production of a second molecule having a defined sequence of nucleotides (e.g., a polyribonucleotide) or a defined sequence of amino acids. For example, a DNA molecule can encode an RNA molecule (e.g., by a transcription process that includes a DNA-dependent RNA polymerase enzyme). An RNA molecule can encode a polypeptide e.g., by a translation process). Thus, a gene, a cDNA, or an RNA molecule encodes a polypeptide if transcription and translation of RNA corresponding to that gene produces the polypeptide in a cell or other biological system. In some embodiments, a coding region of a polyribonucleotide encoding a target antigen refers to a coding strand, the nucleotide sequence of which is identical to the polyribonucleotide sequence of such a target antigen. In some embodiments, a coding region of a polyribonucleotide encoding a target antigen refers to a noncoding strand of such a target antigen, which may be used as a template for transcription of a gene or cDNA.- 53 - 13241894vlAttorney Docket. No.: 2013237-0970
[0326] Epitope'. As used herein, the term “epitope” refers to a moiety that is specifically recognized by an immune system (e.g., an immune system component) of a subject. For example, in some embodiments, an epitope may be a moiety that is specifically recognized by a T cell, a B cell, an immunoglobulin (e.g., antibody or receptor), immunoglobulin (e.g., antibody or receptor), binding component or an aptamer. In some embodiments, an epitope is comprised of a plurality of chemical atoms or groups on an antigen. In some embodiments, such chemical atoms or groups are surface-exposed when the antigen adopts a relevant three-dimensional conformation. In some embodiments, such chemical atoms or groups are physically near to each other in space when the antigen adopts such a conformation. In some embodiments, at least some such chemical atoms are groups are physically separated from one another when the antigen adopts an alternative conformation (e.g., is linearized).
[0327] Estimated tumor sample purity. As used herein, the term “estimated tumor sample purity” is used to refer to an estimate of tumor content of a tumor sample, such as a fraction, percentage etc. of cancer cells within a tumor sample and / or, equivalently, an estimate of normal contamination, such as a fraction, percentage, etc. of normal (e.g., healthy) cells within a tumor sample. It should be understood that, in certain embodiments, tumor samples are assumed to be comprised of tumor cells and normal cells, such that a fraction of tumor cells in a tumor sample is equal to 1 minus a fraction of normal cells (e.g., 1 – μ), estimates of tumor content and / or normal contamination equivalently measure an estimated tumor sample purity.
[0328] Expression'. As used herein, the term “expression” of a nucleic acid sequence refers to the generation of a gene product from the nucleic acid sequence. In some embodiments, a gene product can be a transcript, e.g., a polyribonucleotide as provided herein. In some embodiments, a gene product can be a polypeptide. In some embodiments, expression of a nucleic acid sequence involves one or more of the following: (1) production of an RNA template from a DNA sequence (e.g., by transcription); (2) processing of an RNA transcript (e.g., by splicing, editing, etc.); (3) translation of an RNA into a polypeptide or protein; and / or (4) post-translational modification of a polypeptide or protein.
[0329] Heterozygous Segment'. The term “heterozygous segment” as used here refers to a segment of a normal genome that comprises at least one heterozygous SNP, as well as any corresponding segments of a tumor genome or reference genome. In other words, a particular - 54 - 13241894vlAttorney Docket. No.: 2013237-0970segment of a particular genome is defined as heterozygous or not according to whether the corresponding segment of a normal genome comprises a heterozygous SNP or not. For example, a reference genome may be partitioned into a plurality of segments, as described herein, in order to identify and define corresponding segments in a normal and tumor genome. Accordingly, if, for a given segment, the corresponding segment in the normal genome is determined to comprise a heterozygous SNP, then that segment is defined as a heterozygous segment. For purposes of determining heterozygous segments, a normal genome may be a reference genome obtained from a database, a normal reference determined and / or compiled based on one or more subject (e.g., a panel), determined by sequencing a particular subject (e.g., the same subject whose tumor is being sequenced).
[0330] Homology. As used herein, the term “homology” or “homolog” refers to the overall relatedness between polynucleotide molecules (e.g., DNA molecules and / or RNA molecules) and / or between polypeptide molecules. In some embodiments, polynucleotide molecules (e.g., DNA molecules and / or RNA molecules) and / or polypeptide molecules are considered to be “homologous” to one another if their sequences are at least 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical. In some embodiments, polynucleotide molecules (e.g., DNA molecules and / or RNA molecules) and / or polypeptide molecules are considered to be “homologous” to one another if their sequences are at least 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% similar (e.g., containing residues with related chemical properties at corresponding positions). For example, as is well known by those of ordinary skill in the art, certain amino acids are typically classified as similar to one another as “hydrophobic” or “hydrophilic” amino acids, and / or as having “polar” or “non-polar” side chains. Substitution of one amino acid for another of the same type may often be considered a “homologous” substitution.
[0331] Identity. As used herein, the term “identity” refers to the overall relatedness between polynucleotide molecules (e.g., DNA molecules and / or RNA molecules) and / or between polypeptide molecules. In some embodiments, polynucleotide molecules (e.g., DNA molecules and / or RNA molecules) and / or between polypeptide molecules are considered to be “substantially identical” to one another if their sequences are at least 80%, 85%, 90%, 95%,- 55 - 13241894vlAttorney Docket. No.: 2013237-097096%, 97%, 98%, or 99% identical. Calculation of the percent identity of two nucleic acid or polypeptide sequences, for example, can be performed by aligning the two sequences for optimal comparison purposes (e.g., gaps can be introduced in one or both of a first and a second sequence for optimal alignment and non-identical sequences can be disregarded for comparison purposes). In certain embodiments, the length of a sequence aligned for comparison purposes is at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or substantially 100% of the length of a reference sequence. The nucleotides at corresponding positions are then compared. When a position in the first sequence is occupied by the same residue (e.g., nucleotide or amino acid) as the corresponding position in the second sequence, then the molecules are identical at that position. The percent identity between the two sequences is a function of the number of identical positions shared by the sequences, taking into account the number of gaps, and the length of each gap, which needs to be introduced for optimal alignment of the two sequences. The comparison of sequences and determination of percent identity between two sequences can be accomplished using a mathematical algorithm. For example, the percent identity between two nucleotide sequences can be determined using the algorithm of Meyers and Miller, 1989, which has been incorporated into the ALIGN program (version 2.0). In some exemplary embodiments, nucleic acid sequence comparisons made with the ALIGN program use a PAM 120 weight residue table, a gap length penalty of 12 and a gap penalty of 4. The percent identity between two nucleotide sequences can, alternatively, be determined using the GAP program in the GCG software package using an NWSgapdna. CMP matrix.
[0332] Increased, Induced, or Reduced'. As used herein, these terms or grammatically comparable comparative terms, indicate values that are relative to a comparable reference measurement. For example, in some embodiments, an assessed value achieved with a provided composition (e.g., a pharmaceutical composition) may be “increased” relative to that obtained with a comparable reference composition. Alternatively or additionally, in some embodiments, an assessed value achieved in a subject may be “increased” relative to that obtained in the same subject under different conditions (e.g., prior to or after an event; or presence or absence of an event such as administration of a composition (e.g., a pharmaceutical composition) as described herein, or in a different, comparable subject (e.g., in a comparable subject that differs from the subject of interest in prior exposure to a condition, e.g., absence of administration of a- 56 - 13241894vlAttorney Docket. No.: 2013237-0970composition (e.g., a pharmaceutical composition) as described herein.). In some embodiments, comparative terms refer to statistically relevant differences (e.g., that are of a prevalence and / or magnitude sufficient to achieve statistical relevance). Those skilled in the art will be aware, or will readily be able to determine, in a given context, a degree and / or prevalence of difference that is required or sufficient to achieve such statistical significance. In some embodiments, the term “reduced” or equivalent terms refers to a reduction in the level of an assessed value by at least 5%, at least 10%, at least 20%, at least 50%, at least 75% or higher, as compared to a comparable reference. In some embodiments, the term “reduced” or equivalent terms refers to a complete or essentially complete inhibition, i.e., a reduction to zero or essentially to zero. In some embodiments, the term “increased” or “induced” refers to an increase in the level of an assessed value by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 80%, at least 100%, at least 200%, at least 500%, or higher, as compared to a comparable reference.
[0333] Initial event-. As used herein, the term “initial event” refers to a portion of a tumor genome which may encode a mutation. In certain embodiments, initial events are detected in tumor sequencing data of a tumor sample obtained from a subject. In some embodiments, an initial event is any portion of (e.g., individual sites in, contiguous subsequences in) a tumor genome (e.g., such that a set of initial events may include all sites in a tumor genome). In some embodiments, an initial event is a portion of a tumor genome associated with an observed statistical deviation from an expected read count in tumor sequencing data (e.g., wherein an expected read count is based on normal sequencing data which may be matched normal sequencing data). Given experimental and / or bioinformatic noise that may be present in sequenced samples, it is expected that among a plurality of initial events, any number of initial events may be determined to be false positives (e.g., may be determined to not reflect underlying mutations). In certain embodiments, an initial event refers to a variant portion of a tumor genome (e.g., identifying one or more sites within the tumor genome) determined, based on the sequencing data, to (e.g., potentially) have a different copy number and / or a different nucleotide sequence relative to a corresponding (e.g., aligned with the corresponding variant portion) normal reference. In certain embodiments, the different nucleotide sequence is or comprises an alternate allele (e.g., an individual nucleotide), different from an allele of the corresponding normal reference. In certain embodiments, the different nucleotide sequence is or comprises an insertion and / or a deletion (e.g., of one or more nucleotides) relative to the corresponding normal - 57 - 13241894vlAttorney Docket. No.: 2013237-0970reference (e.g., an indel). In certain embodiments, the variant portion is or comprises a structural variation relative to the normal reference. In certain embodiments, an initial event is a potential point mutation at a particular site. In certain embodiments, an initial event is a potential indel. In certain embodiments, an initial event is a potential structural variation. In certain embodiments, an initial event is a potential copy number variation. In certain embodiments, a normal reference is obtained from a database (e.g., a hl9 reference genome). In certain embodiments, a normal reference is determined based on normal sequencing data obtained for a subject (e.g., by sequencing a normal sample, e.g., of healthy tissue and / or cells)].
[0334] In order. As used herein with reference to a polynucleotide or polyribonucleotide, “in order” refers to the order of features from 5' to 3' along the polynucleotide or polyribonucleotide. As used herein with reference to a polypeptide, “in order” refers to the order of features moving from the N-terminal-most of the features to the C-terminal-most of the features along the polypeptide. “In order” does not mean that no additional features can be present among the listed features. For example, if Features A, B, and C of a polynucleotide are described herein as being “in order, Feature A, Feature B, and Feature C,” this description does not exclude, e.g., Feature D being located between Features A and B.
[0335] Linker. As used herein, the term “linker” refers to a portion of a polypeptide that connects different regions, portions, or antigens to one another.
[0336] Lipid'. As used herein, the terms “lipid” and “lipid-like material” are broadly defined as molecules which comprise one or more hydrophobic moieties or groups and optionally also one or more hydrophilic moieties or groups. Molecules comprising hydrophobic moieties and hydrophilic moieties are also typically denoted as amphiphiles.
[0337] Neoantigen'. As used herein, the term “neoantigen” refers to an antigen that is not present in a reference, such as a normal non-cancerous or germline cell, but is present in a cancer cell. In some embodiments, a neoantigen includes one or more mutations relative to a corresponding antigen present in a normal non-cancerous or germline cell.
[0338] Neoantigen epitope. As used herein, the term “neoantigen epitope” refers to an epitope that is not present in a reference, such as a normal non-cancerous or germline cell, but is present in a cancer cell.- 58 - 13241894vlAttorney Docket. No.: 2013237-0970
[0339] Nucleic acid / Polynucleotide'. As used herein, the term “nucleic acid” refers to a polymer of at least 10 nucleotides or more. In some embodiments, a nucleic acid is or comprises DNA. In some embodiments, a nucleic acid is or comprises RNA. In some embodiments, a nucleic acid is or comprises peptide nucleic acid (PNA). In some embodiments, a nucleic acid is or comprises a single stranded nucleic acid. In some embodiments, a nucleic acid is or comprises a double- stranded nucleic acid. In some embodiments, a nucleic acid comprises both single and double- stranded portions. In some embodiments, a nucleic acid comprises a backbone that comprises one or more phosphodiester linkages. In some embodiments, a nucleic acid comprises a backbone that comprises both phosphodiester and non-phosphodiester linkages. For example, in some embodiments, a nucleic acid may comprise a backbone that comprises one or more phosphorothioate or 5'-N-phosphoramidite linkages and / or one or more peptide bonds, e.g., as in a “peptide nucleic acid”. In some embodiments, a nucleic acid comprises one or more, or all, natural residues (e.g., adenine, cytosine, deoxy adenosine, deoxycytidine, deoxy guanosine, deoxythymidine, guanine, thymine, uracil). In some embodiments, a nucleic acid comprises on or more, or all, non-natural residues. In some embodiments, a non-natural residue comprises a nucleoside analog (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyl adenosine, 5-methylcytidine, C-5 propynyl-cytidine, C-5 propynyl-uridine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyl-uridine, C5-propynyl-cytidine, C5-methylcytidine, 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, 6-O-methylguanine, 2-thiocytidine, methylated bases, intercalated bases, and combinations thereof). In some embodiments, a non-natural residue comprises one or more modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose) as compared to those in natural residues. In some embodiments, a nucleic acid has a nucleotide sequence that encodes a functional gene product such as an RNA or polypeptide. In some embodiments, a nucleic acid has a nucleotide sequence that comprises one or more introns. In some embodiments, a nucleic acid may be prepared by isolation from a natural source, enzymatic synthesis (e.g., by polymerization based on a complementary template, e.g., in vivo or in vitro), reproduction in a recombinant cell or system, or chemical synthesis. In some embodiments, a nucleic acid is at least 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 20, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000,- 59 - 13241894vlAttorney Docket. No.: 2013237-09703500, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500, 10,000, 10,500, 11,000, 11,500, 12,000, 12,500, 13,000, 13,500, 14,000, 14,500, 15,000, 15,500, 16,000, 16,500, 17,000, 17,500, 18,000, 18,500, 19,000, 19,500, or 20,000 or more residues or nucleotides long.
[0340] Ploidy. As used herein, the term “ploidy,” for example of a tumor genome, is used to refer to an average of absolute copy numbers of all segments (e.g., across an entire region of a tumor genome), weighted by the length of each segment. A ploidy of a region of a tumor genome can be defined as the average of absolute copy numbers of all segments in the region, weighted by the length of each segment.
[0341] Polypeptide. As used herein, the term “polypeptide” refers to a polymeric chain of amino acids. In some embodiments, a polypeptide has an amino acid sequence that occurs in nature. In some embodiments, a polypeptide has an amino acid sequence that does not occur in nature. In some embodiments, a polypeptide has an amino acid sequence that is engineered in that it is designed and / or produced through action of the hand of man. In some embodiments, a polypeptide may comprise or consist of natural amino acids, non-natural amino acids, or both. In some embodiments, a polypeptide may comprise or consist of only natural amino acids or only non-natural amino acids. In some embodiments, a polypeptide may comprise D-amino acids, L-amino acids, or both. In some embodiments, a polypeptide may comprise only D-amino acids. In some embodiments, a polypeptide may comprise only L- amino acids. In some embodiments, a polypeptide may include one or more pendant groups or other modifications, e.g., modifying or attached to one or more amino acid side chains, at the polypeptide’s N-terminus, at the polypeptide’s C-terminus, or any combination thereof. In some embodiments, such pendant groups or modifications comprise acetylation, amidation, lipidation, methylation, pegylation, etc., including combinations thereof. In some embodiments, a polypeptide may be cyclic, and / or may comprise a cyclic portion. In some embodiments, a polypeptide is not cyclic and / or does not comprise any cyclic portion. In some embodiments, a polypeptide is linear. In some embodiments, a polypeptide may be or comprise a stapled polypeptide. In some embodiments, the term “polypeptide” may be appended to a name of a reference polypeptide, activity, or structure; in such instances it is used herein to refer to polypeptides that share the relevant activity or structure and thus can be considered to be members of the same class or family of polypeptides. For each such class, the present specification provides and / or those skilled in the- 60 - 13241894vlAttorney Docket. No.: 2013237-0970art will be aware of exemplary polypeptides within the class whose amino acid sequences and / or functions are known; in some embodiments, such exemplary polypeptides are reference polypeptides for the polypeptide class or family. In some embodiments, a member of a polypeptide class or family shows significant sequence homology or identity with, shares a common sequence motif (e.g., a characteristic sequence element) with, and / or shares a common activity (in some embodiments at a comparable level or within a designated range) with a reference polypeptide of the class; in some embodiments with all polypeptides within the class). For example, in some embodiments, a member polypeptide shows an overall degree of sequence homology or identity with a reference polypeptide that is at least about 30-40%, and is often greater than about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more and / or includes at least one region (e.g., a conserved region that may in some embodiments be or comprise a characteristic sequence element) that shows very high sequence identity, often greater than 90% or even 95%, 96%, 97%, 98%, or 99%. Such a conserved region usually encompasses at least 3-4 and often up to 35 or more amino acids; in some embodiments, a conserved region encompasses at least one stretch of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35 or more contiguous amino acids. In some embodiments, a relevant polypeptide may comprise or consist of a fragment of a parent polypeptide.
[0342] Read count'. As used herein, the term “read count” refers to a number of sequencing data reads that map to a particular segment, portion thereof, or individual location (such as a SNP) within a genome. For example, the phrases “read count of a particular segment” and “segment read count” as used herein refer to a number of reads that map to the particular segment. For example, the phrases “read count of a particular SNP” and “SNP read count” as used herein refer to a number of reads that map to the particular SNP. The term “read count” may be preceded by an indication of a particular set of sequencing data and / or sequenced sample. For example, when a tumor sample is sequenced to produce tumor sequencing data comprising a plurality of tumor sequencing reads, the phrase “tumor read count” is used to refer to the number of tumor sequencing reads that map to a particular segment, portion thereof, or individual location (e.g., SNP) within a genome. Likewise, when a normal sample is sequenced to produce normal sequencing data comprising a plurality of normal sequencing reads, the phrase “normal- 61 - 13241894vlAttorney Docket. No.: 2013237-0970read count” is used to refer to the number of normal sequencing reads that map to a particular segment, portion thereof, or individual location (e.g., SNP) within a genome.
[0343] Reference-. As used herein, the term “reference” describes a standard or control relative to which a comparison is performed. For example, in some embodiments, an agent, animal, individual, population, sample, sequence or value of interest is compared with a reference or control agent, animal, individual, population, sample, sequence or value. In some embodiments, a reference or control is tested and / or determined substantially simultaneously with the testing or determination of interest. In some embodiments, a reference or control is a historical reference or control, optionally embodied in a tangible medium. Typically, as would be understood by those skilled in the art, a reference or control is determined or characterized under comparable conditions or circumstances to those under assessment. Those skilled in the art will appreciate when sufficient similarities are present to justify reliance on and / or comparison to a particular possible reference or control.
[0344] Ribonucleic acid (RNA) or Polyribonucleotide-. As used herein, the term “ribonucleic acid,” “RNA,” or “polyribonucleotide” refers to a polymer of ribonucleotides. In some embodiments, an RNA is single stranded. In some embodiments, an RNA is double stranded. In some embodiments, an RNA comprises both single and double stranded portions. In some embodiments, an RNA can comprise a backbone structure as described in the definition of “Nucleic acid / Polynucleotide” above. An RNA can be a regulatory RNA (e.g., siRNA, microRNA, etc.), or a messenger RNA (mRNA). In some embodiments, an RNA is a mRNA. In some embodiments, where an RNA is a mRNA, an RNA typically comprises at its 3' end a poly(A) region. In some embodiments, where an RNA is a mRNA, an RNA typically comprises at its 5' end an art-recognized cap structure, e.g., for recognizing and attachment of a mRNA to a ribosome to initiate translation. In some embodiments, an RNA is a synthetic RNA. Synthetic RNAs include RNAs that are synthesized in vitro (e.g., by enzymatic synthesis methods and / or by chemical synthesis methods).
[0345] Ribonucleotide-. As used herein, the term “ribonucleotide” encompasses unmodified ribonucleotides and modified ribonucleotides. For example, unmodified ribonucleotides include the purine bases adenine (A) and guanine (G), and the pyrimidine bases cytosine (C) and uracil (U). Modified ribonucleotides may include one or more modifications - 62 - 13241894vlAttorney Docket. No.: 2013237-0970including, but not limited to, for example, (a) end modifications, e.g., 5' end modifications (e.g., phosphorylation, dephosphorylation, conjugation, inverted linkages, etc.), 3' end modifications (e.g., conjugation, inverted linkages, etc.), (b) base modifications, e.g., replacement with modified bases, stabilizing bases, destabilizing bases, or bases that base pair with an expanded repertoire of partners, or conjugated bases, (c) sugar modifications (e.g., at the 2' position or 4' position) or replacement of the sugar, and (d) internucleoside linkage modifications, including modification or replacement of the phosphodiester linkages. The term “ribonucleotide” also encompasses ribonucleotide triphosphates including modified and non-modified ribonucleotide triphosphates.
[0346] Secretory signal'. As used herein, the term “secretory signal” refers to an amino acid sequence motif that targets associated polypeptides for translocation to a secretory pathway.
[0347] Segment, Segments: As used herein, the terms “segment” or “segments” (e.g., when used in regard to genetic material) refer to specific pre-defined regions of one or more genomes. For example, a particular reference genome may be subdivided into a plurality of segments, each a specific subsequence of consecutive nucleotides in the reference genome. In certain embodiments, a particular genome may be subdivided into its constituent genes, such that each segment corresponds to a particular, different, gene of the particular genome. In certain embodiments, exons of a particular genome are identified and retained, such that each segment corresponds to a particular, different, exon. In certain embodiments, each segment corresponds to a locus (e.g., a particular location on a chromosome where a particular gene, genetic marker, or allele is located). In certain embodiments, a particular genome may be subdivided into segments of a same or substantially same size (e.g., number of bases). As will be understood by one of skill in the art, a set sequencing data obtained from a particular sample, such as reads obtained via next generation sequencing (NGS) data obtained by sequencing a particular sample, may be aligned to a reference genome. In this way, reference genome may be subdivided into a plurality of segments and corresponding segments identified within a genome characteristic of the particular sample. Reads from the set of sequencing data can, accordingly, be identified as mapping to various particular segments within the genome characteristic of the particular sample and used to characterize them. In certain embodiments, multiple sets of sequencing data may be obtained for different samples (e.g., tumor sequencing data from a tumor sample, normal- 63 - 13241894vlAttorney Docket. No.: 2013237-0970sequencing data from a normal sample) and aligned to a common reference genome. In this way, corresponding segments that comprises the same or substantially same (e.g., all save for variations due to e.g., single nucleotide polymorphisms (SNPs), single nucleotide variations (SNVs), insertions, deletions, etc.) base positions as from genomes characteristic of different samples can be identified. That is, given a particular segment from one genome, associated with one sample, a corresponding segment of another genome, associated with another sample, may be identified. Corresponding segments may have a same and / or substantially same length (e.g., accounting for insertions, deletions, etc.). A particular segment is referred to herein as encoding or comprising a particular SNP and / or SNV if that particular SNP and / or SNV is within the particular segment.
[0348] Single Nucleotide Polymorphism (SNP): As used herein, the term “single nucleotide polymorphism” or “SNP” refers to a particular site (e.g., base position) in a genome where alternative bases are known and / or determined to distinguish one allele from another.
[0349] Single Nucleotide Variation (SNV): As used herein, the term “single nucleotide variation” is used to refer to a difference in the nucleic acid sequence (substitution of one base for another) at a particular site (allele) when comparing a genome from a diseased cell, such as a tumor cell, and a genome of a normal, non-diseased cell or a reference genome. In some embodiments, detecting mutations may refer to detecting nucleotide substitution mutations. In certain embodiments, a SNV is a somatic point mutation that occurs only in diseased (e.g., cancer) cells.
[0350] Subject: As used herein, the term “subject” refers to an organism to be administered with a composition described herein, e.g., for experimental, diagnostic, prophylactic, and / or therapeutic purposes. Typical subjects include animals (e.g., mammals such as mice, rats, rabbits, non-human primates, domestic pets, etc.) and humans. In some embodiments, a subject is a human subject. In some embodiments, a subject is suffering from a disease, disorder, or condition (e.g., cancer and / or a cancer-associated condition). In some embodiments, a subject is susceptible to a disease, disorder, or condition (e.g., cancer and / or a cancer-associated condition). In some embodiments, a subject displays one or more symptoms or characteristics of a disease, disorder, or condition (e.g., cancer and / or a cancer-associated condition). In some embodiments, a subject displays one or more non-specific symptoms of a - 64 - 13241894vlAttorney Docket. No.: 2013237-0970disease, disorder, or condition (e.g., cancer and / or a cancer-associated condition). In some embodiments, a subject does not display any symptom or characteristic of a disease, disorder, or condition (e.g., cancer and / or a cancer-associated condition). In some embodiments, a subject is someone with one or more features characteristic of susceptibility to or risk of a disease, disorder, or condition (e.g., cancer and / or a cancer-associated condition). In some embodiments, a subject is a patient. In some embodiments, a subject is an individual to whom diagnosis and / or therapy is and / or has been administered.
[0351] Therapy. The term “therapy” refers to an administration or delivery of an agent or intervention that has a therapeutic effect and / or elicits a desired biological and / or pharmacological effect (e.g., has been demonstrated to be statistically likely to have such effect when administered to a relevant population). In some embodiments, a therapeutic agent or therapy is any substance that can be used to alleviate, ameliorate, relieve, inhibit, prevent, delay onset of, reduce severity of, and / or reduce incidence of one or more symptoms or features of a disease, disorder, and / or condition (e.g., cancer and / or a cancer-associated condition). In some embodiments, a therapeutic agent or therapy is a medical intervention that can be performed to alleviate, relieve, inhibit, present, delay onset of, reduce severity of, and / or reduce incidence of one or more symptoms or features of a disease, disorder, and / or condition.
[0352] Treat-. As used herein, the term “treat,” “treatment,” or “treating” refers to any method used to partially or completely alleviate, ameliorate, relieve, inhibit, prevent, delay onset of, reduce severity of, and / or reduce incidence of one or more symptoms or features of a disease, disorder, and / or condition (e.g., cancer and / or a cancer-associated condition). Treatment may be administered to a subject who does not exhibit signs of a disease, disorder, and / or condition (e.g., cancer and / or a cancer-associated condition). In some embodiments, treatment may be administered to a subject who exhibits only early signs of the disease, disorder, and / or condition (e.g., cancer and / or a cancer-associated condition), for example for the purpose of decreasing the risk of developing pathology associated with the disease, disorder, and / or condition. In some embodiments, treatment may be administered to a subject at a later-stage of disease, disorder, and / or condition (e.g., cancer and / or a cancer-associated condition).
[0353] Wild-Type'. As used herein, the term “wild-type” refers to an entity having a structure and / or activity as found in nature in a “normal” (as contrasted with mutant, diseased, - 65 - 13241894vlAttorney Docket. No.: 2013237-0970altered, etc.) state or context. Those of ordinary skill in the art will appreciate that wild-type genes and polypeptides often exist in multiple different forms (e.g., alleles). For example, for a subject with cancer, wild-type genes, segments, SNPs, or properties thereof may be those present in a genome of that subject’s normal, non-cancerous, cells, as opposed to altered versions of those genes, segments, SNPs, or properties thereof that appear in genomes of cancer cells within the subject.DETAILED DESCRIPTION
[0354] It is contemplated that systems, architectures, devices, methods, and processes of the claimed invention encompass variations and adaptations developed using information from the embodiments described herein. Adaptation and / or modification of the systems, architectures, devices, methods, and processes described herein may be performed, as contemplated by this description.
[0355] Throughout the description, where articles, devices, systems, and architectures are described as having, including, or comprising specific components, or where processes and methods are described as having, including, or comprising specific steps, it is contemplated that, additionally, there are articles, devices, systems, and architectures of the present invention that consist essentially of, or consist of, the recited components, and that there are processes and methods according to the present invention that consist essentially of, or consist of, the recited processing steps.
[0356] It should be understood that the order of steps or order for performing certain action is immaterial so long as the invention remains operable. Moreover, two or more steps or actions may be conducted simultaneously.
[0357] The mention herein of any publication, for example, in the Background section, is not an admission that the publication serves as prior art with respect to any of the claims presented herein. The Background section is presented for purposes of clarity and is not meant as a description of prior art with respect to any claim.- 66 - 13241894vlAttorney Docket. No.: 2013237-0970
[0358] Documents are incorporated herein by reference as noted. Where there is any discrepancy in the meaning of a particular term, the meaning provided in the Definition section above is controlling.
[0359] Headers are provided for the convenience of the reader - the presence and / or placement of a header is not intended to limit the scope of the subject matter described herein.
[0360] Presented herein are methods and systems that allow mutations present in a tumor genome to be identified (e.g., detected) with high sensitivity and accuracy via analysis of sequencing data obtained, for example, via next-generation sequencing of tumor samples obtained from a subject. Among other things, mutation detection technologies of the present disclosure are informed by tumor model parameters, and through sequential iteration steps, predictions of mutations and predictions of a tumor model are sequentially improved. Among other things, sensitive and accurate detection and characterization of cancer- specific mutations made possible via techniques described herein can be utilized to detect and / or select targets for personalized cancer therapies.A. Creating Personalized Cancer Immunotherapies Based on Tumor Genome Analysis
[0361] Certain cancer mutations are unique to a patient’s cancer and, when expressed, produce proteins and / or peptides that are distinct from those produced by normal cells. These distinct proteins and / or peptides can, accordingly, be specifically targeted via immunotherapy approaches that leverage the patient’s own immune system to clear cancer cells while avoiding damage to normal cells. Technologies of the present disclosure, among other things, leverage and analyze sequencing data to identify potential cancer-specific mutations within genome(s) of a patient’s tumor (tumor genome) and, moreover, characterize them in a manner that allows those mutations that will be the most effective targets of immunotherapies to be identified and prioritized, for example for as targets for personalized cancer vaccines, T-cell receptor (TCR) therapies, and the like.
[0362] As illustrated in FIG. 1, biological samples obtained from a patient can be sequenced and the resultant sequencing data analyzed to detect and characterize mutations (SNV events) unique to the patient’s cancer cells. Detected mutations can be prioritized according to- 67 - 13241894vlAttorney Docket. No.: 2013237-0970various metrics that reflect, for example, their prevalence as well as propensity to induce an immune response. Non- synonymous mutations that are both highly prevalent (e.g., present in a substantial fraction, up to all, of the patient’s cancer cells) and likely to be effective in priming the patient’s immune system to mount a strong response can, accordingly, be selected for inclusion in a personalized immunotherapy for the patient. A personalized therapeutic can thus be designed, manufactured, and administered to the patient as treatment.
[0363] For example, as illustrated in FIG. 1, in certain embodiments, normal 102a and tumor tissue 102b samples are obtained from the patient. Normal genomic DNA (gDNA) is extracted from a normal tissue sample 102a and tumor gDNA is extracted from a tumor tissue sample 102b. Sequencing (104) may then be performed using the extracted normal and tumor gDNA to generate sequencing data.
[0364] Sequencing (104) may be performed, for example as described in further detail herein, using next generation sequencing (NGS) techniques. Accordingly, in certain embodiments, various pre-processing steps (106) are performed, for example to align reads of sequencing data to a reference genome.
[0365] In certain embodiments, sequencing data [e.g., having been pre-processed (e.g., to align reads to a reference genome)] is used (e.g., as input) for tumor deconvolution and / or mutation detection (108) techniques of the present disclosure. For example, as described in further detail herein, tumor deconvolution and / or mutation detection techniques of the present disclosure operate on sequencing data to determine one or more biophysical and / or genomic features 110 of a patient’s cancer. These determined biophysical and / or genomic features characterize genomic properties of the patient’s cancer (e.g., to the extent represented in the tumor tissue sample), physical properties of the tumor sample 102b. For example, in certain embodiments, one or more tumor genomic features 110a are determined. Tumor genomic features may include, without limitation, copy numbers (e.g., absolute copy numbers and / or allele- specific copy numbers) of one or more segments within a tumor genome (e.g., all segments; e.g., a particular subset of segments, such as heterozygous segments) and / or mutations (SNVs) therein, as well as characteristics of detected mutations (SNVs), such as their zygosity and / or clone type. In certain embodiments, one or more sample features 110b are determined. Sample features 110b may include, without limitation a sample purity and / or contamination - 68 - 13241894vlAttorney Docket. No.: 2013237-0970fraction, which reflect the potential for and amount of tumor samples to include a non-trivial and, at times, substantial, fraction of normal, non-cancerous, cells. In certain embodiments, mutations (e.g., SNVs) are detected 100c. Detected mutations may, for example, be determined by tumor deconvolution and / or mutation detection technologies and provided, for example as a standardized file such as a variant call format (.vcf) file.
[0366] In certain embodiments, detected mutations 110c are filtered to identify non-synonymous mutations (112).
[0367] In certain embodiments, non-synonymous mutations are prioritized (114) to select a subset of mutations for inclusion in a personalized cancer immunotherapy 120. For example, in certain embodiments, mutations (e.g., non-synonymous mutations) are evaluated to determine which particular mutations are, or are predicted to be, present in a substantial fraction (e.g., above a certain threshold, up to all) of a patient’s cancer cells, so that an immune response that targets and eliminates cells expressing one or more particular mutations is likely to eliminate a substantial fraction of the patient’s cancer cells. In certain embodiments, additionally or alternatively, mutations that are determined to, or predicted likely to, elicit a strong immune response within the patient are prioritized and selected for.
[0368] For example, in certain embodiments, one or more scoring metrics 116 are determined for each of at least a portion of detected mutations 100c [e.g., a subset identified as non-synonymous (112), or a portion thereof]. Scoring metrics 116 may be used to prioritize and select a subset of mutations for inclusion in a personalized cancer immunotherapy 120. Scoring metrics 116 may include metrics characterizing potential for a particular mutation to elicit an immune response and may include, without limitation, major histocompatibility complex (MHC) binding predictions, T-cell receptor (TCR) recognition predictions, expression level predictions, and the like. These immune response metrics may be determined via various approaches, including, for example, machine learning and other techniques. Additionally, or alternatively, scoring metrics may include metrics determined via tumor deconvolution and mutation detection technologies described herein, such as, without limitation, absolute copy number values, clone type classifications, values and / or classifications indicating a function of a gene harboring a given mutation, values and / or classifications indicating whether a given mutation is truncal and / or early, and zygosity and / or fractional zygosity values.- 69 - 13241894vlAttorney Docket. No.: 2013237-0970
[0369] In certain embodiments, a prioritized subset of detected mutations may be used for a personalized cancer immunotherapy 120. For example, in certain embodiments, compositions encoding one or more of a prioritized subset of mutations may be manufactured and administered to a patient, for example as a personalized cancer vaccine.A.i Patient Samples and Sequencing Data
[0370] Turning to FIG. 2, as described herein, tumor deconvolution and / or mutation detection technologies of the present disclosure may be used in connection with (e.g., to analyze) sequencing data for a subject to determine [genomic and biophysical properties of, and detect mutations characteristic of, a patient’s cancer. Among other things, technologies of the present disclosure may be used in connection with sequencing and / or genomic analysis methods and systems that generate sequencing data for a subject and allow for detection of mutations that can be targeted, for example, for disease treatment via immunotherapy. Among other things, technologies of the present disclosure may be used in connection with sequencing and / or genomic analysis methods and systems that generate sequencing data for a subject and allow for detection of mutations that can be targeted, for example, for disease treatment via immunotherapy.
[0371] For example, as shown in FIG. 2, in certain embodiments, sequencing data 228 may be generated for a subject having and / or suspected of having cancer, by obtaining a tumor sample 202 from the subject. The tumor sample 202 may be processed 204, for example, to extract and prepare nucleic acid material for sequencing and sequenced 206 to generate tumor sequencing data 208 - i.e., sequencing data representing and obtained from nucleic acid material 204 from a tumor sample 202. Likewise, in certain embodiments, sequencing data 228 may (e.g., also) include normal sequencing data 218 from normal sample 212 having been obtained, processed 214, and sequenced 216, to generate normal sequencing data 218 - i.e., sequencing data representing, and obtained from, nucleic acid material from a normal sample 212.
[0372] A tumor sample 102 may be any sample derived from a particular subject and comprising, and / or expected to comprise, cancer cells e.g., of the particular subject). In certain embodiments, a tumor sample is or comprises a liquid sample, such as serum, plasma, blood, urine, etc. For example, a liquid sample, such as blood, may comprise, or be suspected of - 70 - 13241894vlAttorney Docket. No.: 2013237-0970comprising, cancer cells, such as circulating tumor cells (CTCs). In certain embodiments, a tumor sample is or comprises a tissue sample, for example, obtained from a subject via biopsy. A tumor sample may be representative of a subject’s primary tumor and / or one or more metastases. For example, a primary tumor sample may be obtained via biopsy of a region of a subject known or expected to harbor a primary tumor. A metastasis sample or metastases samples may be obtained via biopsy of one or more region(s) of a subject known or expected to harbor metastases. Tumor samples may be obtained and / or preserved in a variety of formats, such as, for example, isolated cells (e.g., CTCs) from blood samples, fresh tissue, flash frozen tissue, formalin-fixed paraffin embedded tumor tissue (FFPE samples), and the like.
[0373] As illustrated in FIG.2, a tumor sample may comprise cancer cells 202b as well as, in certain cases, normal (z.e., non-cancerous) cells 202a. Accordingly, in certain embodiments, a tumor sample purity may measure relative fraction of cancer cells within a tumor sample. Tumor sample purity may, for example, be computed as (1 − μ) = ηT / (ηT+ ηN), where μ is a contamination fraction, representing a relative fraction of normal cells infiltrating a tumor sample, given by μ = ηN / (ηT+ ηN) and ηTand ηNare a number of tumor and normal cells in a tumor sample, respectively. Tumor sample purity may be expressed as a decimal value, percentage, etc. As described in further detail herein, typically, purity does not need to be measured directly (e.g., via direct measuring / counting amounts of tumor and normal cells in a sample), but, rather, can be determined and / or estimated using sequencing data, for example, via tumor modelling approaches such as those described herein.
[0374] As described in further detail herein, in certain embodiments, additionally or alternatively, bounds, such as upper and / or lower bounds for sample purity (e.g., and / or contamination fraction) may be estimated, e.g., via tumor modeling approaches of the present disclosure. For example, as described in further detail herein, in certain embodiments, at low physical sample purities, accuracy of estimation methods, such as tumor deconvolution, may be reduced such that purity estimates based on sequencing data are expected to be of insufficient accuracy to be used in and of themselves. In such cases, however, an upper bound (e.g., a maximum purity) and / or a lower bound (e.g., a minimum purity) may still be estimates and used in certain processing steps.- 71 - 13241894vlAttorney Docket. No.: 2013237-0970
[0375] A normal sample 212 may be any sample derived from a particular subject and comprising, and / or expected to comprise, the subject’s normal cells, but not cancer cells (e.g., in certain embodiments, a normal sample 212 does not contain any cancer cells). In certain embodiments, a normal sample is or comprises a liquid sample, such as serum, plasma, blood, urine, saliva, etc. In certain embodiments, a normal sample is or comprises a tissue sample, for example obtained from a subject via biopsy. Normal samples may be obtained and / or preserved in a variety of formats, such as, for example, isolated cells [e.g., peripheral blood mononuclear cells (PBMCs)] from blood samples, fresh tissue, flash frozen tissue, formalin-fixed paraffin embedded tumor tissue (FFPE samples), and the like.
[0376] As illustrated in FIG.2, a normal sample 212 nominally contains only normal patient cells 212a. In certain embodiments, a normal sample may be obtained from blood of a subject. In certain embodiments, a normal sample may still comprise a small (e.g., negligible) number of cancer cells. For example, in certain embodiments, a normal cell may be obtained from a region of a subject near a tumor e.g., in an effort to obtain a normal sample from a same or similar underlying tissue type). In certain embodiments, a normal sample comprises less than 1%, e.g., less than 0.1%, e.g., less than 0.01%, e.g., less than 0.001% tumor cells.
[0377] Samples, such as tumor samples and / or normal samples, may be processed to obtain, and / or prepare, nucleic acid material therefrom for sequencing. For example, nucleic acid material, such as DNA and / or RNA, may be extracted and prepared for sequencing (e.g., via amplification, fragmentation, labeling, etc.) as appropriate, depending on a particular desired sequencing method and / or data format. In certain embodiments, tumor sample 202 is processed 204 and / or normal sample 212 is processed 214 to extract nucleic acid material, such as gDNA. In certain embodiments, normal gDNA is extracted from normal tissue sample 212 and tumor gDNA is extracted from tumor tissue sample 202. For example, in certain embodiments, sequencing data may be whole genome sequencing (WGS) data; in certain embodiments, sequencing data may be whole exome sequencing (WES) data. Various commercially available kits and instruments may be used to prepare samples for and obtain WGS and / or WES data, including, but not limited to, those provided by Illumina, Inc., PacBio, Oxford Nanopore Technologies, Thermo Fisher Scientific’s Ion Torrent™, etc. In some embodiments, methods and systems of the present disclosure utilize whole exome sequencing (WES) or whole genome- 72 - 13241894vlAttorney Docket. No.: 2013237-0970sequencing (WGS) data in singletons (e.g., one sequenced tumor sample and one sequenced normal sample). In some embodiments, methods and systems of the present disclosure utilize whole exome sequencing (WES) or whole genome sequencing (WGS) data in replicates (e.g., two or more sequenced tumor samples and two or more sequenced normal samples). In some embodiments, methods and systems of the present disclosure utilize whole exome sequencing (WES) or whole genome sequencing (WGS) data in conjunction with RNA-seq data. In some embodiments, RNA-seq data is used in a method or system of the present disclosure to filter errors such as PCR errors, FFPE, sequencing artifacts, and / or other types of errors. For example, a mutation detected in DNA sequencing data may only be accepted if it is also detected using RNA-seq reads. In some embodiments, RNA seq reads are high quality RNA-seq reads. In some embodiments, if only singletons are used, coverage of singletons is required to be comparable to coverage of merged replicates (e.g., in order to ensure comparable sensitivity).
[0378] In certain embodiments, to generate sequencing data, libraries are created from the normal gDNA and tumor gDNA. In certain embodiments, from each gDNA sample, two or more libraries can be generated. In certain embodiments, from each gDNA sample, two libraries can be generated. For example, in certain embodiments, as illustrated in FIG.2, sequencing data 228 may be generated in and / or comprise replicates created by preparing multiple (e.g., two or more) libraries associated with each (e.g., gDNA) sample. The libraries can be created for whole exome sequencing and / or whole genome sequencing and / or RNA sequencing (RNAseq). Samples are then sequenced using high throughput sequencing such as NGS.
[0379] For example, in certain embodiments, such that, for example, sequencing data 228 may be generated in and / or comprise replicates. For example, multiple tumor samples may be extracted and sequenced independently; in certain embodiments, a single tumor sample may be extracted and used to prepare multiple libraries (e.g., such that processing steps of extracting, fragmenting, and amplifying nucleic from the sample are performed repeatedly and independently), which are then sequenced; in certain embodiments, library preparation may be performed repeatedly on a single pool of extracted nucleic acid, and the multiple libraries sequenced; in certain embodiments, a single library is sequenced multiple times (e.g., as in a technical replicate). In certain embodiments, as with tumor sample sequencing data, e.g., as illustrated in FIG.2, normal sample sequencing data may also comprise a plurality of replicates.- 73 - 13241894vlAttorney Docket. No.: 2013237-0970
[0380] Sequencing data 228, 208, 218, may be stored and / or presented in a variety of formats, such as FASTQ, SAM, BAM, etc. For example, sequencing data for a particular sample may comprise a plurality of reads, each read representing a nucleotide sequence of a polynucleotide fragment corresponding to (e.g., that maps to) a portion of a subject’s tumor genome and / or exome, and / or portion of the subject’s normal genome and / or exome. In certain embodiments, sequencing data typically comprises multiple overlapping reads, which may be aligned to a reference genome, such as an hl9 or h38 reference genome (see, e.g., Ncbi.nlm.nih.gov / datasets / genome / GCF_000001405.13 / and Ncbi.nlm.nih.gov / datasets / genome / GCF_000001405.26 / , respectively; see also Karolchik, D. et al. Nucleic Acids Res. 32, D493-D496, 2004 and Kent WJ et al., Genome Res.12(6):996-1006, 2002) for human patients, to map each read to a particular region of the overall genome. A reference genome may, accordingly, be a publicly available reference genome. In certain embodiments, a reference genome may be generated based on a normal genome of the subject.
[0381] For example, in certain embodiments, NGS sequencing is performed and generates FASTQ files as output. In certain embodiments, NGS output reads, stored, for example, in FASTQ files, are aligned. Aligned reads may be stored and / or provided in file formats, such as sequence alignment map (SAM) or the binary compressed version thereof (BAM). In certain embodiments, sequencing data duplicate reads may be marked and / or removed from the sequencing data. In certain embodiments, sequencing adapters may be removed from the sequencing data. In certain embodiments, e.g., in the case of short read sequencing, sequencing can be paired end or single end, wherein different read lengths can be used (e.g., 50bp, 100bp, 150bp, etc.).
[0382] In certain embodiments, aligned reads, e.g., as stored in one or more BAM file(s), are used as input for tumor deconvolution and / or mutation detection technologies described here. For example, in certain embodiments, samples are sequenced using high throughput sequencing such as next generation sequencing, generating, for example, FASTQ files as output. Next, FASTQ files are aligned, and, in certain embodiments, duplicate reads are marked. An alignment step may be performed, and generate, as output, a BAM file for a normal sample and a BAM file for a tumor sample. In certain embodiments, output of an alignment step is two normal BAM files and two tumor BAM files.- 74 - 13241894vlAttorney Docket. No.: 2013237-0970
[0383] As described in further detail in the following, based on input sequencing data (e.g., comprising tumor sequencing data and normal sequencing data), tumor deconvolution and mutation detection technologies of the present disclosure may determine various [biophysical and genomic] properties of a tumor sample and detect mutations (e.g., SNVs) occurring within the tumor genome.A.i.a Histological Tumor Content (HTC)
[0384] In some embodiments, a histological tumor content is an input to a method of the present disclosure. For example, in some embodiments, purity is used as an input, and the purity used as input is HTC.
[0385] In certain embodiments, an additional and / or alternative purity measurement, e.g., obtained via another (e.g., external, not based on sequencing data) may be used as input. For example, an additional purity measurement may be determined / obtained via an additional and / or alternative measurement approach such as a histology (e.g., using H& E staining), to obtain, e.g., a histological tumor content (HTC) as an additional purity measurement. In certain embodiments, imaging (e.g., fluorescence) or spectroscopic (e.g., infrared, Raman, etc.) technique may be used to obtain an additional purity measurement. In certain embodiments, an additional purity measurement may be a value retrieved from a database (e.g., of stored characteristic sample purities), for example based on cancer type, stage, tumor location, sampling method, etc.
[0386] In some embodiments, an HTC is provided as an input to a tumor modeling procedure (e.g., to provide an initial tumor content estimation before the tumor model is estimated based on sequencing data). In some embodiments, an HTC is provided as an input to a generic tumor model. In some embodiments, an HTC is used to determine if iterative tumor modeling is activated. In some embodiments, an HTC is used to improve performance of an FP filter.
[0387] In general, without wishing to be bound by any particular theory, a purity estimate based on tumor sequencing data may be preferred over HTC and / or is used in addition to HTC in methods and systems of the present disclosure because HTC may be error prone. For example, HTC may not reflect true tumor content within a sequenced tumor sample because interpretation of tumor tissue is complex (e.g., for human pathologists as well as Al methods); different - 75 - 13241894vlAttorney Docket. No.: 2013237-0970pathologists may provide different estimations for the same tumor tissue, and even the same pathologist may provide different estimations of tumor content for the same sample on different occasions. Similarly, Al or computer-assisted methods may make significantly different predictions depending on the algorithms and training procedures used. Further, in some cases, even if HTC is accurately estimated, such an HTC estimation is determined based on a tissue section and may not accurately reflect tumor content within an entire tissue block. If an HTC does not reflect true tumor content of a sequenced tumor sample, either because of measurement bias or because it does not correspond to a true tumor content of a sequenced tumor sample, this may result in hyper-sensitization of mutation detection which can inflate errors, or undersensitization of mutation detection, in which case ability to identify low VAF mutations (including clonal mutations in low tumor content samples) can be adversely impacted.Accordingly, in many cases, tumor modeling based on sequencing data may be a more accurate method to determine purity of a sequenced tumor.A.ii Tumor Genomics
[0388] Turning to FIG.3, sequencing data may be used to piece together and / or infer properties relating to a tumor genome and / or normal genome of a subject.A.ii.a Segments
[0389] As illustrated in FIG.3, a genome may be subdivided into a plurality of segments, each segment representing a particular sub-region (e.g., a subsequence) of the genome. In certain embodiments, a reference genome may be subdivided into a plurality of segments. In certain embodiments, a tumor genome and / or a normal genome may be subdivided into a plurality of segments (black bars in the normal genome and tumor genome schematics).
[0390] Among other things, in certain embodiments, as illustrated in FIG.3, this approach provides a common coordinate system or set of subregions between reference genome 302, normal genome 322, and tumor genome 342. For example, (e.g., sequencing data, such as reads, corresponding to) normal genome 322 and (e.g., sequencing data, such as reads, corresponding to) tumor genome 322 may be aligned to reference genome 302. In this way reads - 76 - 13241894vlAttorney Docket. No.: 2013237-0970can be mapped to particular regions of reference genome 302 and a coordinate system for [reads of] normal genome 322 and tumor genome 342 can be provided. For example, reference genome can be used to determine a chromosome number, a nucleotide position (in the chromosome), as well as a directionality of a given read.
[0391] Thereafter, by partitioning reference genome 302 into a plurality of segments, regions of normal genome 322 and tumor genome 342 that correspond to a particular segment can be identified and, for example, analyzed and compared.
[0392] For example, in the illustrative schematic shown in FIG.3, nine segments 304a, 304b, 304c...304i, and 304j (collectively 304) of reference genome 302 are shown, with corresponding segments of normal genome 322 and tumor genome 342 indicated via the vertical dashed lines.A.ii.b Copy Number Variations
[0393] As illustrated in FIG.3, the majority of normal genome 322 is typically diploid, with two copies of each particular segment (e.g., except for segments located on male sex chromosomes for which normal genome 322 contains either a single maternal chromosome or a single paternal chromosome), with one copy associated with (e.g., originating from) a maternal allele 326a and another associated with (e.g., originating from) a paternal allele 326b. In certain embodiments, methods and systems described herein may exclude male sex chromosomes and / or segments that map to locations on male sex chromosomes from analysis since male sex chromosomes do not contain heterozygous SNPs, which, as explained in further detail herein, can be used to facilitate determining tumor genomic properties and / or sample purities.
[0394] As illustrated in FIG.3, for segments in a normal e.g., human) genome that, absent germline CNV events, is diploid, segments may be heterozygous, in that the copies differ, corresponding to different alleles, or may be homozygous, comprising identical alleles.
[0395] In contrast, segments in a tumor genome 342 are not necessarily diploid. Nor are segments in a tumor genome necessarily balanced. Certain segments of a tumor genome may, however, be diploid and / or balanced. For example, as shown in FIG.3, a number of copies of each allele in a segment may vary from segment to segment, and may be less than two, equal to - 77 - 13241894vlAttorney Docket. No.: 2013237-0970two, or greater than two. Segments in a tumor genome need not be balanced - i.e., tumor genome heterozygous segments do not necessarily comprise a same number of copies of each allele, but, in certain embodiments, may comprise a greater number of copies of one allele (a “major allele”) than the other (a “minor allele”).
[0396] Accordingly, in certain embodiments, various parameters are used to characterize and represent physical properties of a particular segment (e.g., within a normal and / or tumor genome), along with, in certain embodiments, mutations (such as SNVs) identified therein.
[0397] For example, in certain embodiments, a segment j may be characterized by an absolute copy number, CNj, which is computed as a number of copies of a particular segment (e.g., within a single tumor cell).
[0398] For example, in the schematic shown in FIG.3, while all segments in normal genome 322 have one copy of a maternal allele and one copy of a paternal allele, copy numbers of maternal and paternal alleles in tumor genome 342 may vary from segment to segment. Accordingly, while all segments in normal genome 322 that are shown in FIG.3 have an absolute copy number of two, in tumor genome 342 segments 344a, 344b, 344c, 344d, 344e, 344f, 344g, 344h, 344i, 344j have absolute copy numbers of 2, 4, 3, 1, 3, 5, 0 (a deletion), 2, 4, and 1, respectively. Moreover, as illustrated in FIG.3, in tumor genome 342, the number of copies of maternal and paternal alleles also may vary from segment to segment such that, for example segment 344a has one copy of each a maternal and paternal allele, segment 344b has two copies of each, and segment 344f has three copies of the maternal allele and two copies of the paternal allele.
[0399] Normal segments can also have other copy number configurations due to germline CNV events, however, without limiting the generality of the method, such events are not described in the figure.A.ii.c Single Nucleotide Polymorphisms and Heterozygous Segments
[0400] In certain embodiments, maternal and paternal alleles of a particular segment in a normal genome 322 may harbor different variants of a single nucleotide polymorphism (SNP). In this case, the particular segment is referred to as comprising a heterozygous SNP. The - 78 - 13241894vlAttorney Docket. No.: 2013237-0970particular segment in the normal genome 322, along with corresponding segments in reference and tumor genomes, are referred to as heterozygous segments.
[0401] For example, as illustrated in FIG.3, reference genome 302 is subdivided into a plurality of segments. For segment 304a of reference genome 302, corresponding segment 324a in normal genome 322 comprises at least one heterozygous SNP, such that segment 304a of reference genome 302 and corresponding segments 324a and 344a of normal and tumor genome, respectively, are referred to as heterozygous segments. In contrast, segment 304i of reference genome corresponds to segment 324i of normal genome, which does not comprise any heterozygous SNPs. Accordingly, segment 304i and corresponding segments 324i and 344i in normal and tumor genomes are not heterozygous segments.
[0402] Overall, FIG.3 depicts seven (7) heterozygous segments and three (3) homogeneous segments. As illustrated in the figure, for each of the heterozygous segments, copies in the normal genome 322 harbor at least one heterozygous SNP (illustrated schematically via different colored orange and green dots in the maternal and paternal alleles), whereas normal genome 322 copies of the homogeneous segments do not contain any heterozygous SNPs.Different normal genomes (e.g., of different individual subjects) will generally have different heterozygous segments since different normal genomes encode different heterozygous SNPs.
[0403] As explained herein, while normal genome 322 typically has two copies of each segment (apart from those located on male sex chromosomes) the number of copies of particular segments in a tumor genome may differ from two and can vary from segment to segment.Additionally, or alternatively, as illustrated in FIG.3, tumor genomes do not necessarily have equal numbers of maternal and paternal alleles for each heterozygous segment and / or, in certain cases, may entirely lack a maternal or paternal copy. For example, segment 304a is a heterozygous segment for which corresponding tumor genome segment 344a has two copies: one maternal and one paternal. Second segment 304b is another heterozygous segment. In this case, however, corresponding segment 344b in tumor genome 342 has four copies: two maternal copies of the segment in the tumor genome and two paternal copies of the segment in the tumor genome. Therefore, the segment has an absolute copy number of 4 in the tumor genome.Reference number 344c shows three copies of a heterozygous segment: one maternal copy of the- 79 - 13241894vlAttorney Docket. No.: 2013237-0970segment in the tumor genome and two paternal copies of the segment in the tumor genome, therefore the segment has an absolute copy number of 3 in the tumor genome.
[0404] In certain embodiments, a heterozygous segment may comprise only paternal or only maternal alleles, referred to herein as a loss of heterozygosity (LOH) event. For example, segments 344d and 344e in tumor genome each correspond to a normal diploid segment that comprises a heterozygous SNP - i.e., a heterozygous segment. However, as illustrated in FIG.3, segments 344d and 344e in tumor genome lack any maternal allele copies - they have only (one and three, respectively) copies of the paternal alleles. Accordingly, while these are segments that have undergone a LOH event.
[0405] In certain embodiments, for a given segment (e.g., in reference or normal genome), the corresponding segment may be entirely absent from tumor genome 344g, indicating that, for example, all copies of the segment (both maternal and paternal) were deleted (complete deletion). Segment 344g in tumor genome, accordingly, has an absolute copy number of O.A.ii.d Tumor Genome Mutations and Characteristics
[0406] FIG. 3 illustrates an expanded view of segments 322f and 344f. As illustrated in FIG. 3, a normal cell 352 comprises normal genome 322, including maternal 354a and paternal 354b alleles of segment 322f, with maternal allele comprising an alternative variant of SNP 356.Tumor cell 372 comprises tumor genome 342, including segment 344f, which corresponds to normal segment 322f. Unlike normal cell, tumor cell 372 comprises multiple copies of each allele - namely, three copies of maternal allele and two copies of paternal allele.
[0407] As described herein, corresponding segments 322f and 344f in normal 322 and tumor 324 genomes, respectively, can be characterized by values of parameters such as an absolute copy number and an allele specific copy number. For example, as shown in FIG.3, normal segment 322f has an absolute copy number of two - (CAwt = 2) and corresponding tumor genome segment 344f has an absolute copy number 373 of five (e.g., CNj= 5). As illustrated in FIG. 3, although a normal segment will typically have one copy of each parental allele, CNV events may cause a tumor genome segment to have one or multiple (e.g., two or more) copies of- 80 - 13241894vlAttorney Docket. No.: 2013237-0970each parental allele. Moreover, the number of maternal and paternal alleles need not be equal in a tumor genome. Accordingly, heterozygous tumor genome segments may also be characterized by an allele- specific copy number. For heterozygous segments, such as segment 354, an allele specific copy number 364 may be determined as a maximum absolute number of copies of a major allele - i.e., the allele having a number of copies greater than or equal to that of the other, minor, allele - i.e., CNX≥ CNY, where X and Y denote the major and minor alleles, respectively. For example, for the particular segment 354 shown in FIG. 3, there are three copies of the major allele, such that the allele specific copy number is three (CNx = 3).
[0408] As shown in the bottom portion of FIG. 3, tumor genome segments may harbor mutations, such as a SNV 382. In the notation used herein, a copy number of a segment harboring a mutation may be denoted CAmut (e.g., for segment 344f in tumor cell 372, CAmut = 5). As shown in FIG. 3, mutations may occur in particular alleles but are not necessarily present in each copy of a particular allele. For example, while SNV 382 occurs in maternal allele 374, it is not present in all three copies of maternal allele 374 - it is present in only two copies.Accordingly, additional parameters may be determined to characterize genomic properties of mutations like SNVs.
[0409] For example, if a certain segment is mutated, a number of physical copies of the mutated segment in a given tumor cell, referred to herein as the zygosity of the mutation, may be determined. In certain embodiments, a fractional zygosity of a mutation (Q may be computed as a ratio of the zygosity of a mutation (Cx) and the absolute copy number (CAmut) of the segment harboring the mutation in the tumor genome. For example, in FIG. 3, segment 344f harbors a mutation 382 having a zygosity of two and a fractional zygosity, of 2 / 5 (0.4).
[0410] Turning to FIG. 4, as described herein, a sample of tumor tissue, may comprise normal cells 400 and tumor cells 430. As shown in FIG. 4 and described herein, segments in normal cells 415 are assumed to have an absolute copy number of two, except for segments occurring on male sex chromosomes for which an absolute copy number is one. As shown in the figure, a segment may be amplified in the tumor cells. In the example illustrated in FIG. 4, the tumor cell segment has an absolute copy number of five 420. As shown in the figure, example mutation 430 (black star) has three copies, and, accordingly, its zygosity is three (440). In the illustrative example shown in FIG. 4, mutation 430 (black star) is present in all tumor cells and - 81 - 13241894vlAttorney Docket. No.: 2013237-0970is therefore a clonal mutation, whereas a second mutation 450 (red star) is present in a subset of tumor cells (just one of the three tumor cells in the diagram in FIG.4) and is therefore a subclonal mutation. Clonal mutations with a zygosity greater than 1 can arise, for example, when the mutation occurred before the copy number amplification event (hence are considered “early” mutations), whereas subclonal mutations with a zygosity of 1 can arise, for example, when the mutation occurred after the copy number amplification event (hence are considered “late” mutations). In this model all CNV events are considered to be clonal, however, in a more general model, CNV events can also be present in just a subset of tumor cells.
[0411] In certain embodiments, a cellularity of a mutation, denoted by p, may be determined. Cellularity as used herein refers to the fraction of tumor cells that harbor a given mutation. A mutation is said to be biologically clonal if all cancer cells in a tumor sample harbor the given mutation. Biologically clonal mutations are characterized by having a nominal cellularity of 1 (p = 1). For example, clonal mutation 430 is present in all three tumor cells and, accordingly, has a nominal cellularity of 1, whereas subclonal mutation 450 is present in 1 out of 3 tumor cells, and, accordingly, has a nominal cellularity p = 1 / 3.
[0412] In certain embodiments, approaches described herein determine cellularity estimates for mutations. In certain embodiments, a cellularity estimate is an estimated mean cellularity (e.g., indicating, if multiple tumor samples were obtained and sequence, a given mutation is estimated to be present, on average, in a fraction of tumor cells given by the mean cellularity). In certain embodiments, confidence intervals for cellularity estimates may be determined, with lower and upper bounds denoted uγ(p) and vγ(p), where y is the confidence level of the estimate (e.g., also referred to as degree of confidence or confidence coefficient). A cellularity confidence interval (CI) [uγ(p), vγ(p)], may, for example, indicate that, if multiple tumor samples were obtained, the estimated cellularity would be on the interval [e.g., at or between uγ(p) and vγ(p)] y percent of the time. For example, in certain embodiments, a 90% CI lower and upper bound are determined. In certain embodiments, a 95% CI lower and upper bound are determined. In certain embodiments, a 68% CI lower and upper bound are determined.- 82 - 13241894vlAttorney Docket. No.: 2013237-0970B. Tumor Deconvolution and Mutation Detection
[0413] Among other things, this application describes technologies for determining genetic and biophysical properties of tumor samples based on sequencing data (a procedure referred to, in certain cases, as “tumor deconvolution”). Tumor deconvolution procedures described herein may comprise, among other things, one or more tumor models that accurately explain tumor sequencing data.
[0414] As described in further detail herein, tumor deconvolution technologies of the present disclosure allow for complex characteristics of tumor genomes to be determined based on sequencing data. For example, cancer cells undergo extensive mutations, including copy number variation (CNV) events and single nucleotide variations (SNVs). As a result, unlike genomes extracted from normal cells, tumor genomes are not reliably diploid. Instead, the number maternal and paternal copies of genetic material varies from segment to segment, depending on the CNV events that took place over the lifetime of a given population of cancer cells. Additionally, cancer cells harbor collections of mutations at individual sites - SNV events - such as substitutions, insertions, deletions, etc.
[0415] Among other things, mutations, such as SNV events, that are found in tumor cells make valuable targets for therapeutics, including immunotherapies such as individualized cancer therapies. Accordingly, the ability to accurately and rapidly detect and characterize mutations is an important step in treating cancer patients.B.i Challenges and Insights: Tumor Models and Mutation Detection
[0416] Accurate detection and characterization of tumor mutations, however, is a highly complex process. Among other things, as illustrated in FIG.5A, genetic and biophysical features of tumor samples may both impact particular tumor read counts and quantities, such as allele frequencies, that are derived therefrom and observed based on sequencing data.
[0417] For example, as explained above, tumor samples often include a fraction of normal, healthy cells. Read counts for particular mutations and / or portions of a tumor genome that are observed in sequencing data may be impacted by tumor sample purity. Accordingly, the certainty with which an event, such as a collection of reads with unique (e.g., abnormal) base - 83 - 13241894vlAttorney Docket. No.: 2013237-0970calls at a particular site, can be determined to indicate true underlying cancer cell mutations depends on tumor sample purity. Additionally, or alternatively, CNV events, which may increase or decrease relative amounts of genetic material - and thus read counts - for particular portions (e.g., genes) of the tumor genome, may also impact how underlying, true, mutations manifest in observable sequencing data.
[0418] Accordingly, among other things, tumor deconvolution technologies of the present disclosure allow estimation of tumor sample purity and, in certain embodiments, characterization of CNV events across a tumor genome. As described in further detail herein, in certain embodiments, determining biophysical and genomic properties of a tumor sample in this manner can be used to improve accuracy with which cancer mutations are detected.
[0419] Additionally, or alternatively, in certain embodiments, mutations may be prioritized as targets for immunotherapy, for example to select high value targets for inclusion in personalized cancer vaccines, T-cell therapies, and the like, according to genomic characteristics, such as whether they are determined to be biologically clonal or subclonal, their zygosity, etc. Accordingly, ability to characterize genomic properties of mutations themselves and / or portions of a tumor genome where they are located (e.g., copy numbers of segments harboring mutations) can be highly valuable in the context of therapeutic approaches.
[0420] Determining biophysical properties of tumor samples, such as their purity, and characterizing genomic properties, such as copy number variations across a tumor genome, in non-trivial. Among other things, both tumor sample purity and CNV events affect how observed sequencing reads (e.g., such as impacting relative read counts of various segments in a tumor genome).
[0421] Additionally, or alternatively, the present disclosure and techniques described herein appreciate the fact that tumor genomes vary widely from patient to patient and from indication to indication in terms of the frequency of mutations, the frequency and extent CNV events, and the degree of subclonality of these genetic features. For example, the number of SNVs detected in an exome can vary across patients by 3 orders of magnitude, and across indications by at least 4 orders of magnitude from -1 (in certain pediatric cancers) up to -104 (Alexandrov, L. B. et al., 2013, Nature 500, 415-421, Lawrence, M. S. et al., 2013, Nature 499, 214-218). Similar diversity is observed for CNVs. For example, breast cancers can contain - 84 - 13241894vlAttorney Docket. No.: 2013237-0970anywhere from thousands of CNV events and other structural variation events to nearly none. Likewise, lung squamous cell tumors can contain anywhere from hundreds of CNV events and other structural variation events to none. On the other hand, indications like kidney renal clear cell carcinoma and medulloblastoma appear to contain significantly fewer structural variations (Yang, L. et al., 2013, Cell 153, 919-929, Network, C. G. A. R., 2013, Nature 499, 43-49, Parsons, D. W. et al., 2011, Science 331, 435-439). Adding to these difficulties is the fact that tumor sample purity is often, in practice, not high, and frequently even very low, leading to a reduction in signal to noise ratio of the genetic features sought to be estimated.
[0422] Accordingly, among other things, tumor deconvolution technologies of the present disclosure allow estimation of tumor sample purity from sequencing data, accurately accounting for the manner in which CNV events impact sequencing data observations such as read counts, variant allele frequencies (VAFs), and the like. Notably, tumor deconvolution methods and systems described herein leverage signal from SNV events themselves and allow for tumor sample purities to be determined without necessarily requiring direct characterization of CNV events. Accordingly, as described in further detail herein, even when characterization of CNV events is not feasible. As described in further detail herein, in certain embodiments, determining biophysical and genomic properties of a tumor sample in this manner can be used to improve accuracy with which cancer mutations are detected.B.i.a Achieving high sensitivity and high accuracy is challenging, particularly when tumor content and / or tumor mutation burden is low
[0423] Personalized cancer immunotherapies are designed to train the immune system to target certain neoepitopes specific to a tumor. A step in design of personalized cancer immunotherapies is identification of neoepitope targets. Since neoepitopes originate from mutations present in a tumor genome, developing personalized cancer immunotherapies requires detection of mutations present in a genome of a tumor sample. Accuracy and sensitivity of mutation detection will directly impact clinical efficacy of personalized cancer immunotherapies. Accordingly, in order to maximize therapeutic efficacy of personalized cancer immunotherapies, mutations should be identified using a method with high sensitivity (i.e., identifying as many useful mutations as possible) and accuracy (i.e., minimizing detection of false positives) and - 85 - 13241894vlAttorney Docket. No.: 2013237-0970which works consistently for every patient sample. As discussed further below, achieving high sensitivity and accuracy of mutation detection can be particularly useful, and particularly challenging, in samples in which tumor content and / or tumor mutation burden (TMB) are low, and such conditions are often encountered in the clinic and in the same tumor sample.
[0424] Provided herein are methods and systems for identifying mutations with high sensitivity and accuracy, including in samples with low tumor content. Among other things, the present disclosure is based on an insight that performing accurate and sensitive mutation detection and performing accurate tumor modeling are coupled problems: identifying mutations with high accuracy and high sensitivity is challenging because, in order to maximize sensitivity and accuracy of mutation detection for a tumor sample, an accurate model of the tumor sample (including, e.g., an estimate of tumor sample purity) should be determined, while, on the other hand, in order to maximize accuracy of a model of a tumor sample, mutations should be identified and leveraged (e.g., as in a tumor modeling or deconvolution approach as described herein). Accordingly, among other things, the present disclosure provides methods and systems for mutation detection which employ an iterative feedback-driven estimation procedure based on a model of a tumor sample (e.g., comprising determining an initial list of putative mutations (e.g., based on a set of initial events)) in a tumor sample, estimating a model of the tumor sample based on the initial list of putative mutations, iteratively refining the list of putative mutations based on the model of the tumor sample, and refining the model of the tumor sample used to refine the list of putative mutations).B.i.b Clinical efficacy of personalized cancer immunotherapies is related to accuracy and sensitivity of mutation detection
[0425] Personalized cancer immunotherapies are designed de novo and tailored to each patient. Since personalized cancer immunotherapies are designed to target mutations that are specific to each patient’s tumor, design of personalized cancer immunotherapies involves detecting mutations encoded in a patient’s tumor genome. To maximize therapeutic potential and efficacy of personalized cancer immunotherapies across patients, mutation detection should ideally be robust to variation in quality of sequencing, quality of tumor sample, tumor content of tumor sample and TMB of a tumor genome.- 86 - 13241894vlAttorney Docket. No.: 2013237-0970B.i.c Factors that may impact mutation detection performanceQuality of a tumor sample and quality of sequencing
[0426] Quality of genomic material present in a tumor sample can vary in terms of its integrity, its fragment size, and its fidelity. For example, older tumor samples are more likely to contain more degraded genomic DNA (gDNA) than more recently collected tumor samples. In many cases, tumor samples are preserved using a Formalin-Fixed Paraffin-Embedded (FFPE) process than can directly quality of gDNA. Further, quality of RNA samples can also be highly variable from patient to patient, depending on how tumor tissue was initially excised and how tumor tissue was stored, fixed, transported, etc.
[0427] In addition, quality of sequencing can also vary. Quality of tumor sample and quality of sequencing can also be related. For example, poorer quality tumor samples containing degraded or fragmented gDNA can yield poorer sequencing quality. In general, it is desirable that mutation detection be robust to variation occurring due to sample and / or sequencing quality.Tumor content
[0428] Tumor content of a tumor sample is another parameter that can impact mutation detection performance. In general, it is expected that as tumor content decreases, signal (e.g., associated with genomic features of interest such as mutations) decreases, and it may become more difficult to distinguish signal from noise. Given that noise from various sources is generally present in tumor samples and / or tumor sequencing data, statistically significant events detected in tumor sequencing data (e.g. initial events as described herein) may not be true positives.Accordingly, in general, it is expected to be more difficult to identify mutations with a similar degree of sensitivity and specificity in lower tumor content samples compared to higher tumor content samples. However, ideally, mutation detection should be minimally impacted by tumor content of a tumor sample. Accordingly, a mutation detector (e.g., which comprises one or more false positives filters) should ideally detect true positives and reject false positives with high sensitivity and high specificity even when tumor content is low.- 87 - 13241894vlAttorney Docket. No.: 2013237-0970
[0429] In certain embodiments, e.g., in clinical settings, tumor samples have low tumor content (e.g., below about 40%, below about 30%, or below about 20%). In certain embodiments (e.g., certain indications), a majority of tumor samples can have a low tumor content (e.g., below about 40%, below about 30%, or below about 20%).Tumor mutation burden (TMB)
[0430] TMB is another parameter that may impact accuracy of mutation detection. TMB is generally defined as a total number of mutations present in a patient’s tumor. Different indications typically have different ranges of characteristic TMBs (see, e.g., Alexandrov et al., 2013, doi.org / 10.1038 / naturel2477). For example, indications such as medulloblastoma, pancreatic cancer and prostate cancer are associated with low average TMB, whereas indications such as bladder cancer, lung cancer, and melanoma are associated with high average TMB. However, across patients within each indication there can be a wide range of TMBs, and in some indications, despite a higher average TMB, there may still be a significant percentage of patients with low TMB.
[0431] Accuracy of mutation detection can be measured by a positive predictive value (PPV) of mutation detection, given by ratio of a number of predicted mutations that are true positives (TP) and a total number of predicted mutations (given by a sum of TP and a number of predicted mutations that are false positives (FP)):Eq. (1) TPPPV = - TP+FP
[0432] Another way to interpret Eq. 1 is to set TP to a TMB of a given patient, and interpret FP to be a random variable (RV) corresponding to a number of FPs predicted by a given mutation detector per patient:Eq. (2) TMBPPV = - TMB+FP- 88 - 13241894vlAttorney Docket. No.: 2013237-0970
[0433] Since mutation callers are generally not tuned to specific indications and do not employ a tumor model (e.g., in contrast to methods of detecting mutations as described herein), and because a decision to call a mutation occurs at each site in a genome independently (e.g., in contrast to methods of detecting mutations as described herein), a distribution of FP is expected to be independent of TMB and independent of tumor content of a tumor sample and therefore represents an additive noise term. A distribution of FPs can therefore represent a characteristic accuracy of a mutation detector. For each tumor sample analyzed by a mutation detector, a number of false positives that the detector predicts (FP) is drawn from a distribution that characterizes accuracy of that mutation detector. Eq.2 suggests that lower TMBs are associated with a greater impact of a noise term. For example, having 10 false positives when TP=100 will result in a PPV of 90.9%.; however, since FP is generally independent of TMB and is a characteristic of a given mutation detector, if TMB is 10 and TP=10, the same mutation detector would still predict on average 10 false positives, leading to a PPV of 50%.
[0434] Accordingly, the present disclosure is based in part on an insight that, while TMB may not directly impact a mutation caller, TMB has an important indirect effect on mutation detection in that PPV is expected to decrease with decreasing TMB unless FP is much smaller than TMB (FP«TMB).
[0435] As a result, personalized cancer immunotherapies designed for patients who have tumors with a low TMB may be likely to be more susceptible to effects of errors (FP) predicted by a mutation detector. For this reason, in order to achieve a PPV close to 1 for all patients, including those with low TMB tumors, a mutation detector should ideally make less than one error per exome (FP<1) and, ideally, significantly less than one error per exome (FP<<1)
[0436] The present disclosure provides methods and systems which increase positive predictive value (PPV) of mutation detection (e.g., by reducing number of false positives and / or reducing a false discover rate FDR = 1 — PPV) relative to existing methods and systems. For example, as described in Example 3 in analysis of four fresh frozen (FF) tumor samples with varying TMB, methods of the present disclosure can achieve FDRs of zero or near zero with comparable sensitivity to previously published mutation callers (e.g., MuTect and Strelka), while previously published mutation callers have FDRs that can reach -25% (see, e.g., FIGs.51 - - 89 - 13241894vlAttorney Docket. No.: 2013237-097051F). Further, results of Example 3 show that for these published mutation callers, FDR is inversely proportional to TMB (as predicted by Eq.2; see, e.g., FIG.50B). For example, as shown in Example 3, a noise distribution (FP) for existing published mutation callers was determined using normal-versus-normal analysis (see, e.g., FIGs.51A-51F, FIGs.50A - 50C) and showed that estimated noise was consistent with measured FDR of clinical samples as predicted by Eq.2. Without wishing to be bound by any particular theory, these results support an idea that, for mutation callers that do not utilize a tumor model (e.g., that do not utilize a tumor model based on tumor sequencing data, as disclosed herein) or parameters estimated therefrom, noise can be regarded as an additive component, and a normal-versus-normal analysis can be used to predict a distribution of a random additive noise component FP. Estimated noise can, in turn, be used to predict a potential range of PPV for a given TMB.
[0437] In contrast, noise distributions were also determined using a normal-versus-normal analysis of an illustrative embodiment of the present disclosure and shown to be close to zero for singleton mode of the illustrative embodiment (see, e.g., FIGs.50A - 50C). Without wishing to be bound by any particular theory, in replicate mode, a noise component according to methods and systems of the present disclosure is expected to be even lower. Further, results presented in Example 3 demonstrate that methods and systems of the present disclosure can achieve a PPV greater than 99% for FFPE samples, regardless of TMB and using various sequencing pipelines (see, e.g., FIGs.51A - 51F, FIGs.52A - 52B).Combination of low TMB and low tumor content
[0438] In certain scenarios, mutation detection is particularly challenging, e.g., when tumor samples have both low TMB and low tumor content. In order to achieve high sensitivity at low tumor content, mutations should be identified for which evidence for the mutation (i.e., “signal”) is close to a noise floor. However, when TMB is low, PPV will be highly sensitive to errors, as discussed above (see, e.g., Eq.2). Thus, in general, when tumor content is low there is likely to be a greater risk that increased sensitivity will come at a cost of reduced accuracy and lead to reduced PPVs (e.g., when evaluated at the level of individual patients).
[0439] In clinical settings, low tumor content samples are often encountered, and for certain cohorts, low tumor samples may be typically encountered. At the same time, certain - 90 - 13241894vlAttorney Docket. No.: 2013237-0970indications such as, e.g., PDAC, prostate, ovarian, certain breast cancer subtypes, and pediatric cancers, are characterized by low TMB. Even for patients that have indications typically characterized by a high TMB, low TMB tumors are often encountered. Therefore, a combination of low tumor content and low TMB samples can often be encountered in clinical settings.
[0440] Accordingly, the present disclosure provides method of detecting mutations with both high sensitivity and high accuracy such that FP is minimized (e.g. approaches <<1) (e.g., for use in identifying mutations present in a tumor genome associated with at low tumor content). Such methods are based on an insight that, by iteratively estimating a tumor model, thereby increasing accuracy of the tumor model, while potentiating FP and mutation confidence score filters using a tumor model or parameters derived therefrom, it is possible achieve high sensitivity, while controlling error rate (i.e., maintaining or increasing accuracy). The present disclosure further provides methods and systems capable of performing highly sensitive mutation detection (e.g., such that PPV approaches 1), independent of TMB and when tumor content is low. Since tumor samples with low TMB and low tumor content are often encountered in the clinic, the present disclosure provides methods and systems for designing personalized cancer immunotherapies that can provide each patient with a therapy capable of achieving maximal clinical efficacy.B.i.d Mutation detection sensitivity can impact efficacy of personalized cancer immunotherapyOnly a subset of mutations may be used as targets for a personalized cancer immunotherapy
[0441] A pool of mutations that can serve as targets in a personalized immunotherapy may not be large and depends on TMB: the lower the TMB, the smaller the pool of potential targets available for a personalized cancer immunotherapy. Moreover, while TMB determines a total number of target candidates, in many cases, only a subset of mutations detected in a tumor serve as targets for a personalized cancer immunotherapy. Mutations useful as targets for a personalized cancer immunotherapy have certain features, for example: (i) they encode a non-synonymous mutation or otherwise lead to transcription of a non- wildtype sequence; (ii) a mutated allele is transcribed and translated; (iii) in MHC class I presentation, a peptide encoding a mutated sequence is recognized by a Transporter Associated with Antigen Processing (TAP) - 91 - 13241894vlAttorney Docket. No.: 2013237-0970protein complex so that it can subsequently be transported from a cell's cytoplasm into an endoplasmic reticulum (ER) where it is paired with an MHC class I molecule; (iv) a peptide fragment encoding a mutation is presented on an MHC molecule (either class I or class II) (e.g., form a pMHC complex); and / or (v) a pMHC complex is recognized by a T cell and generates an immune response specific to the mutated sequence (i.e., a response not triggered by a corresponding wild-type sequence).
[0442] In addition, mutations may have other useful features: for example, useful targets may be biologically clonal in a tumor sample, may be less likely to lead to tumor escape, etc.
[0443] In many cases, when designing a personalized cancer immunotherapy, algorithms are applied to predict which mutations have features such as those described above so that only mutations with a therapeutic potential are included in an immunotherapy. Therefore, in some cases, an effective pool of mutations that can be used as targets for a personalized cancer immunotherapy may be significantly smaller than that predicted by TMB.Lower mutation detection sensitivity can limit clinical efficacy of personalized cancer immunotherapy
[0444] Sensitivity of mutation detection is an important factor impacting clinical efficacy of personalized cancer immunotherapy because only a fraction of mutations detected in the tumor sample are generally used as targets. As discussed above, useful mutations are nonsynonymous and expressed in the tumor, presented on MHC molecules and recognized by T cells. Mutations also vary in their capacity to confer tumor control, prevent tumor escape, and so on. Therefore, having a larger pool of identified mutations from which to select targets for a personalized immunotherapy can increase quality of the final selected targets, leading to a potentially more efficacious immunotherapy (e.g., a therapy leading to a stronger or more potent immune response, a therapy that confers stronger or more effective tumor control, a therapy that is more lasting and / or less likely to undergo tumor escape over time, as assessed relative to an immunotherapy which utilizes targets selected from a smaller pool of identified mutations).
[0445] Sensitivity of mutation detection can therefore be a factor in determining cohort level efficacy of a personalized cancer immunotherapy. If mutation detection sensitivity is low,- 92 - 13241894vlAttorney Docket. No.: 2013237-0970this may result in a smaller pool of mutations from which neoepitopes can be selected, which may in turn lead to potentially design of a personalized immunotherapies that for some patients may not reach its full potential therapeutic efficacy.Lower TMB can be associated with greater impact sensitivity on performance of personalized cancer immunotherapies
[0446] For samples with lower TMBs, mutation detection sensitivity may have larger effects on identification and selection of targets for a personalized cancer immunotherapy and thereby may have a larger effect on the efficacy of such an immunotherapy. For example, when a pool of neoantigens is smaller, as in a low TMB tumor, additional mutations discovered with improved sensitivity and targets derived therefrom may result in a higher potential efficacy of an immunotherapy.
[0447] Furthermore, personalized cancer immunotherapy can have a cutoff (e.g., a minimum number of neoepitopes) below which a treatment cannot be designed. If sensitivity is not sufficiently high, then in cases of low TMB there may simply not be enough targets to design a personalized cancer immunotherapy, and in such cases, a patient may not be eligible to receive a personalized cancer immunotherapy. Since multiple cancer indications are considered to have low TMBs, and across all indications there can be a significant fraction of patients with low TMBs, obtaining high sensitivity is important in order to be able to realize personalized immunotherapeutic treatment for as many patients as possible.Low TMB tumor samples with low tumor content can be particularly challenging
[0448] Tumor samples with low TMB and low tumor content are often encountered in the clinic. Generally speaking, lower tumor content tends to be associated with a higher difficulty of detecting mutations without making false predictions, because lower tumor content means there is a weaker signal for the genomic features of interest compared to noise (e.g., wherein noise can be of experimental, bioinformatic, biological or other origin). Therefore, as tumor content diminishes, increasing sensitivity without impacting accuracy is expected to be more - 93 - 13241894vlAttorney Docket. No.: 2013237-0970challenging. If accuracy (e.g., as assessed by a rate of false positives) is fixed (as discussed further below), then as tumor content diminishes and signal associated with mutations becomes more difficult to distinguish from noise, attaining a similar level of sensitivity without increasing a rate of false positives becomes increasingly challenging. When TMB is low, as tumor content diminishes, any decrease in sensitivity can directly impact quality of targets that are selected for therapy. In extreme cases, poor sensitivity at low tumor contents could result in failure to design a therapy and exclusion of a patient. Therefore, the ability to detect mutations with high sensitivity is particularly useful when tumor content is low.
[0449] The present disclosure provides methods and systems for increasing sensitivity in mutation detection, while also achieving a high level of accuracy, that are particularly useful for samples with low tumor content and / or low TMB. In some embodiments, methods and systems of the present disclosure achieve high sensitivity and high accuracy by estimating mutations and a tumor model with increasing accuracy and sensitivity in an iterative manner using a feedback mechanism, while applying FP filters and mutation confidence score filters dependent on tumor model parameters.B.i.e Mutation detection accuracy can impact efficacy of personalized cancer immunotherapies
[0450] Accuracy of mutation detection can also impact clinical performance of personalized cancer immunotherapies. In some embodiments, personalized cancer immunotherapies target a predetermined number of neoepitopes or have a predetermined capacity (e.g., to include a certain number of neoantigens but allowing flexibility in the number depending on neoantigen length). Generally speaking, a lower PPV of mutation detection can lead to fewer neoantigen epitopes associated with true mutations being included in an immunotherapy, which in turn can potentially result in a less efficacious therapy. Other factors, such as epitope spreading may contribute to an effect of a personalized cancer immunotherapy, however, ultimately, efficacy of an immunotherapy generally depends on targets included in the immunotherapy. For example, not all targets elicit a similar degree of T cell response, and different targets may have different capacity to confer tumor control or prevent tumor escape. Therefore, if fewer neoepitopes associated with true positive mutations are included in an - 94 - 13241894vlAttorney Docket. No.: 2013237-0970immunotherapy, clinical efficacy of an immunotherapy may be diminished, with impact dependent on number and / or quality of true positives associated with neoepitope targets for an immunotherapy.False positives can come at expense of true positives
[0451] Since, as discussed above, personalized cancer immunotherapies generally have a finite capacity for targets, every false positive included in an immunotherapy takes up a space which would preferably be used for a true positive, thereby potentially reducing maximal potential efficacy of the immunotherapy. For example, if a personalized cancer immunotherapy targets Nneoneoepitopes, a number of targets encoding true positives will be given on average by:Eq. (3)J / vVTTDP = p i p i v v ■J / v’neowhere PPV is a positive predictive value of mutation detection. Thus, a lower PPV could lead to fewer true positives being included in a personalized cancer immunotherapy, resulting in fewer relevant T cells responses and in some embodiments, reduced clinical efficacy.
[0452] Although a personalized cancer immunotherapy may include NTP neoepitopes targeting true positive mutations, it is possible that only a fraction of those neoepitopes will elicit an immune response. A number of neoepitopes eliciting an immune response (e.g., a clinically relevant immune response) may be referred to as a number effective targets. Impact of PPV on potential clinical efficacy can depend on a number of effective targets. A number of effective targets can be determined by different factors.
[0453] In some embodiments, (e.g., since algorithms for predicting immunogenicity have limited predictive power), a subset of targets included in an immunotherapy (i.e., not all targets included) may lead to an immune response. Therefore, in some embodiments, a number of “effective targets” included in a personalized immunotherapy (i.e., neoepitopes included in a personalized immunotherapy leading to an immune response) may be lower than a number of targets included in a design of an immunotherapy.- 95 - 13241894vlAttorney Docket. No.: 2013237-0970
[0454] In some embodiments (e.g., for low TMB tumors and samples derived therefrom), a pool of mutations that can elicit an immune response is also lower. Accordingly, in some embodiments, a number of neoepitopes capable of generating an immune response may overall be lower due to an lower TMB associated with the tumor derived therefrom (e.g., as compared to a tumor associated with a higher TMB).
[0455] In some embodiments (e.g., in personalized cancer vaccines), when multiple neoepitopes are delivered together, immune response to certain neoepitopes may be diminished or biased by presence of other neoepitopes (e.g., due to antigen competition). Without wishing to be bound by any particular theory, there are several mechanisms that may lead to antigen competition, including, for example: competition between multiple peptides in antigen presenting cells for a same HLA binding site (e.g., such that high-affinity peptides can displace peptides with weaker affinities, limiting their presentation); T cells with higher-affinity TCRs and / or higher precursor frequency can engage dendritic cells more efficiently and for longer, leading to increased costimulation associated with certain peptides; dominance of certain clones can constrain expansion of weaker clones; and / or inhibitory receptors can preferentially suppress of lower-avidity responses by, e.g., regulatory T cells.
[0456] In some embodiments, only a subset of TCR clonotypes robustly expand in response to a personalized cancer immunotherapy. In some embodiments (e.g., if antigen competition is present), an effective number of targets in an immunotherapy may be reduced.Impact of PPV may be higher when a number of effective targets is low
[0457] When a number of effective targets is low, potential impact of PPV on clinical efficacy of an immunotherapy can be more pronounced in that an immunotherapy depends on a smaller number neoepitopes to achieve a desired clinical response.
[0458] For example, if a PPV is 50% and there are 10 effective targets (10 neoepitopes generating an immune response), that leaves, on average, 5 neoepitopes targeting true positives to drive a clinical response of an immunotherapy which selects those targets. The probability that all neoepitopes target false positives and that accordingly, such an immunotherapy will not be effective is small: (l-0.5)A10=10A(-3). In this example, the immune response could be robust to- 96 - 13241894vlAttorney Docket. No.: 2013237-0970presence of some false positives. However, if there are only 2 effective targets, that leaves on average only one neoepitope targeting a true positive to drive the clinical response of the immunotherapy, and the probability that all neoepitopes target false positives would be significant: (l-0.5)A2=0.25. Achieving a PPV of 50% is a realistic scenario for existing (e.g., previously published) mutation detectors when TMB is low. For example, if a mutation detector has 91% accuracy when TMB is 100, this suggests that FP-10. Therefore, when TMB drops to 10, PPV will be 10 / ( 10+10)=0.5.
[0459] In an extreme example of just a single effective target, clinical efficacy would be directly proportional to PPV of mutation detection. For example, if PPV is 75%, there is a 25% chance that an immunotherapy will target a false positive and therefore have no clinical efficacy beyond any adjuvant effect of the immunotherapy itself that is not related to the selected targets. In a case where PPV of mutation detection is 75%, for example, the probability that at least one of four patients would not show a clinical response would be very high: 1 -( l-0.75)A4=0.68. Thus, in this case, there would be a 68% chance that one or more of these four patients would not respond to immunotherapy.
[0460] In some embodiments, a number of effective targets can be reduced, even to one, due to: (1) fewer adequate targets as a result of low TMB; (2) poor performance of an immunogenicity prediction algorithm; and / or (3) presence of antigen competition.
[0461] A low PPV may have a potential clinical impact even when all detected targets are used by an immunotherapy, e.g., due to antigen competition, such that a neoepitope encoding a false positive mutation becomes immunodominant, in which case such an immunotherapy may have reduced or even no clinical efficacy. Therefore, if a number of effective targets is low, clinical response may be less robust to presence of false positives.
[0462] These examples show that achieving a low rate of errors (FP<~1 error per exome, ideally <<~1) is particularly useful to ensure efficacy of personalized cancer immunotherapies for tumor samples for which a number of effective targets is low.- 97 - 13241894vlAttorney Docket. No.: 2013237-0970Challenge of achieving high accuracy and high sensitivity when TMB is low and tumor content is low
[0463] Particularly for low TMB tumors, in order to maximize the chance of finding the best targets out of a smaller pool of underlying mutations, sensitivity of mutation detection is ideally as high as possible. In more extreme cases where TMB is very low, high sensitivity is required to find a sufficient number of targets to include in an immunotherapy. However, when TMB is low, PPV will be more sensitive to an additive noise component, FP (see Eq.2). Thus, if a mutation detector strives to be highly sensitive, unless a noise component, FP, is extremely low, this can lead to a drop in PPV. This is particularly relevant when tumor content is low, since in such cases a mutation detector is ideally highly sensitive to detect true positives near a noise floor, but this can lead, in turn, to an increased additive noise component (FP), leading to an even sharper drop in PPV. A goal of the present disclosure was therefore to design a method capable of achieving high sensitivity, ideally close to 1, while keeping additive noise component low, e.g., «1.
[0464] In line with such a goal, the present disclosure provides methods and systems for achieving a low noise floor (FP<<1) when detecting mutations with high sensitivity in low tumor content samples. In some embodiments, methods and systems of the present disclosure perform an iterative feedback-driven estimation process of mutations and a tumor model, where false positive filters (FP filters) are potentiated with an increasingly more accurate tumor model. Such methods enable FP filters to tune filtering to a noise floor, while at the same time tuning sensitivity to detect low VAF mutations given an increasingly accurate tumor model.B.i.f Tumor modeling parameters can be used to enrich for clonal mutations
[0465] Clonal mutations are generally considered to be good targets for a personalized cancer immunotherapy because clonal mutations (and, by extension, corresponding neoepitopes) are present in all tumor cells. However, in order to determine with a high degree of confidence whether a mutation is biologically clonal, it is useful to have an accurate estimate of purity of a tumor sample and absolute copy number of the segments to which the mutation maps (e.g., both segments which contain the mutation, and segments not containing the mutation). However, even in absolute copy numbers are unknown or cannot be reliably estimated, knowledge of estimated - 98 - 13241894vlAttorney Docket. No.: 2013237-0970purity of a tumor sample can allow for enrichment for clonal mutations. The present disclosure provides methods and systems for integrating tumor model parameters such as estimated purity in a mutation detection procedure (e.g., using an RMCS as described herein).B.i.g Degeneracy in SNV-based purity estimation
[0466] Turning to FIG.5B, in general, purity estimation is important for achieving highly sensitive and accurate mutation detection. Purity estimation can be achieved, for example, by leveraging mutations present in a tumor genome or by leveraging CNV events present in a tumor genome, as these are markers that may be unique to a tumor sample and absent in a normal genome. Since tumor genomes are diverse, in some embodiments, a given tumor genome will contain sufficient SNVs or CNVs for purity estimation, in some embodiments, even if sufficient CNVs are present, a method for purity estimation based on these events may not be successful due to the complexity involved in solving the problem. Moreover, in some embodiments, (e.g., at low purities), it is no longer feasible to estimate purity based on CNVs; in such embodiments, methods of purity estimation based on mutations (e.g., SNV-based purity estimation methods) are a particularly useful alternative. Accordingly, the present disclosure provides methods and systems of SNV-based purity estimation which are in certain embodiments, useful alternatives to CNV-based purity estimation methods and systems (e.g., when CNV-based purity estimation fails). However, the present disclosure identifies a problem that, without knowledge of copy numbers and zygosities of SNV s, an estimation of purity is mathematically not unique when using SNVs (e.g., SNV events) as markers.
[0467] To illustrate this point, Eq.4 shows a purity of a tumor sample for a diploid site in a normal genome in the absence of noise given an observed variant allele frequency (VAF) of a SNV in the tumor sample:2 ■ VAFEq. (4) purity = 2 · VAF / (2 · VAF + CNmut(κ − VAF))- 99 - 13241894vlAttorney Docket. No.: 2013237-0970wherein CNmutis a copy number of a mutation in the tumor genome and K is a fractional zygosity of a mutation given by K = Cx / CNmut, wherein Cxis a zygosity of the mutation (number of mutated copies of a gene to which the mutations maps). To illustrate the problem of degeneracy identified by the present disclosure, consider a SNV measured to have a VAF of 0.2.FIG. 5B shows possible purity values for various copy numbers and zygosities for an SNV, illustrating the point that without restriction of SNVs to a particular copy number and zygosity, purity is not unique.
[0468] Further, the present disclosure provides an insight that restricting SNVs to just diploid SNVs (SNVs with a copy number of 2) would not resolve the identified problem because the zygosity can still impact the estimated purity, as shown in the second and third rows of FIG.5B; restricting SNVs just to SNVs in balanced regions of a tumor genome would also not resolve the problem because, as shown in FIG.5B, balanced SNVs (rows highlighted in FIG.5B) also have degenerate purities; and restricting SNVs to balanced SNVs with a copy number of 4, corresponding to minimal balanced SNVs when a minimal copy number is equal to 4 (CNmut= 2, Cx= 1 or 2) also does not resolve the problem because here too different zygosities can yield different purity estimations (seventh and eighth row in in FIG.5B). Further, however, the present disclosure provides an insight that restricting SNVs to balanced SNVs in diploid regions (i.e., CNmut= 2, CX= 1) would resolve the problem and yield a unique purity prediction since there is only one combination of copy numbers and zygosities for this subset of SNVs.
[0469] Accordingly, methods and systems of the present disclosure are directed to resolving the identified problem by using SNVs in balanced diploid regions of a tumor genome (e.g., to perform SNV-based purity estimation), as under such conditions, a unique purity estimation can be obtained (e.g., under asymptotic sampling conditions). Accordingly, the present disclosure also provides methods and systems for identifying SNV events occurring in balanced diploid regions of a tumor genome without knowledge of copy numbers and resolving situations wherein balanced diploid SNV events do not occur in a tumor genome. In certain embodiments, the present disclosure provides methods and systems of SNV-based purity estimation which utilize CNV-based purity estimation, (e.g., a CNV-based purity estimation which provides uncertain information), and which may improve CNV-based purity estimation.- 100 - 13241894vlAttorney Docket. No.: 2013237-0970B.ii Identifying and Using Subpopulations of Mutations for Tumor Modeling
[0470] Turning to FIG.6, in certain embodiments, a tumor deconvolution process 600 may, among other things, construct and fit one or more tumor models that accurately explain sequencing data - such as aligned reads for tumor and, optionally, normal genome for a patient. As shown in FIG.6, in an example tumor modelling process 600, sequencing data may be obtained 602. As described herein, sequencing data 602 may comprise tumor sequencing data and / or normal sequencing data. In certain embodiments, sequencing data comprises tumor sequencing data and normal sequencing data (e.g., which may comprise replicates).
[0471] In certain embodiments, tumor modelling process 600 obtains a list of putative SNV events 604, with each putative SNV event representing one or more potential underlying mutations present in a tumor genome representative of the tumor sample. A list of putative SNV events may include SNV events detected based on tumor sequencing data, for example via various mutation calling approaches whereby, for example, sequencing data is analyzed and sites determined to be likely to encode an alternate allele (e.g., nucleotide base), different from a wildtype allele (e.g., nucleotide base) characteristic of a normal genome, identified as putative SNV events.
[0472] Tumor modeling process 600 may, in certain embodiments, identify and / or select a subset of SNV events 606 to use for purity estimation. Among other things, as described herein, quantities observed in sequencing data, such as a distribution of tumor read counts for various alleles, variant allele frequencies (VAFs), and the like, for SNV events may be impacted by factors such as the biological clonality or subclonality of a given SNV event as well as the copy number and zygosity of the SNV (e.g., the region in which the SNV is located).Accordingly, in certain embodiments, rather than attempt to model a complex constellation of possible underlying biological scenarios and how they present in sequencing data, tumor deconvolution approaches described herein identify and filter SNVs to obtain a selected subset determined to have, or to be likely to have, a particular, intentionally limited, underlying biology. In this way, a single, simplified tumor model that aims to capture a particular limited range of biological scenarios used and fit to a particular subset of sequencing data for which its assumptions are appropriate.- 101 - 13241894vlAttorney Docket. No.: 2013237-0970
[0473] For example, in certain embodiments tumor modeling process 600 identifies and utilizes a subset of SNVs events that are determined to represent mutations that are balanced within a tumor genome and are located in segments that have a particular absolute copy number. In certain embodiments, the particular absolute copy number is a minimal copy number, whose exact value is unknown but predicted to be lower than copy numbers of other balanced segments (e.g., the exact value of the minimal copy number is initially unknown but determined to be smaller relative to other absolute copy numbers of balanced segments). In certain embodiments, the particular absolute copy number is a primary copy number, whose exact value is unknown, but which is determined to be a most frequently occurring even-numbered absolute copy number (e.g., the exact value of the primary copy number is initially unknown but is determined to occur more frequently than other even-numbered copy numbers within the tumor genome).B.ii.a Identifying Balanced Regions of a Tumor Genome
[0474] FIGs. 7-9 show example processes by which particular subsets of segments and / or SNVs may be identified and selected based on sequencing data.
[0475] FIG. 7 shows an example process 700 by which particular segments within a tumor genome, such as balanced heterozygous segments, those having a minimal copy number, and / or a primary copy number, may be identified. FIGs.8 and 9 illustrate example approaches for identifying and selecting particular subpopulations of segments within a tumor genome. FIG. 8 is a schematic illustrating an example process for identifying and selecting a subset of primary balanced heterozygous segments and using the selected subset of primary balanced heterozygous segments to identify and select primary segments. As described in further detail below, primary balanced heterozygous segments are balanced heterozygous segments that are determined to have a same, particular, primary copy number that occurs more frequently than other even-numbered absolute copy numbers. Primary segments are identified as also having the primary copy number but are not necessarily balanced and / or heterozygous within the tumor genome. FIG.9 is a schematic illustrating an example process for identifying and selecting a subset of minimal balanced heterozygous segments and using the selected subset of minimal balanced heterozygous segments to identify and select minimal segments. As described herein, e.g., in further detail below, minimal balanced heterozygous segments are balanced heterozygous - 102 - 13241894vlAttorney Docket. No.: 2013237-0970segments that are determined to have a same, particular, minimal copy number that is determined to be less than other even-numbered copy numbers. Minimal segments are identified as also having the minimal copy number but are not necessarily balanced and / or heterozygous within the tumor genome.
[0476] In certain embodiments, process 700 identifies and / or receives heterozygous SNPs (e.g., present in a subject’s normal genome) 702.
[0477] SNPs that are present in a normal genome of a subject (also referred to herein as “wild-type”) may be identified using normal sequencing data, for example via a commercial SNP or variant caller (e.g., Illumina DRAGEN, Qiagen Genomics, Torrent Variant Caller, etc.) based on sequence alignment to a reference and statistical modeling. Additionally, or alternatively, an initial list of SNPs may be determined via orthogonal methods, such as SNP arrays or TaqMan assays. Additionally, or alternatively, an initial list of SNPs may be obtained from a database, for example of potential SNPs present in an average individual of a particular population of which the subject is a member.
[0478] In certain embodiments, normal sequencing data may be used to evaluate and filter an initial list of SNPs and / or heterozygous SNPs for consistency with an expected underlying model of a heterozygous SNP, e.g., a normal diploid genome with, for a given SNP, two alleles present in equal copy number. Accordingly, in certain embodiments, a balanced test may be used to determine whether sequencing data supports, with sufficiently high confidence, an underlying hypothesis that two alleles of a SNP are balanced - e.g., present in equal copy numbers. For example, statistical tests, such as a Z-statistic test, may be used to classify a particular SNP as balanced or unbalanced. SNPs or putative heterozygous SNPs in an initial list that do not meet a balanced test may be filtered and removed from the initial list of heterozygous SNPs.
[0479] In certain embodiments, heterozygous SNPs of the initial list may be assessed to determine whether sequencing data meets particular quality criteria. For example, each heterozygous SNP may be required to have a particular minimum coverage level in both normal and tumor sequencing data and / or a quality score above a particular minimum value.
[0480] In certain embodiments, heterozygous SNPs may be filtered according to the particular chromosome where they are located. For example, SNPs located on particular- 103 - 13241894vlAttorney Docket. No.: 2013237-0970chromosomes may be excluded from analysis and removed from the list of heterozygous SNPs. In certain embodiments, only heterozygous SNPs located on particular chromosomes, such as chromosomes 1 through 22, are used.
[0481] In certain embodiments, a final list of heterozygous SNPs 704 produced in this manner may then be evaluated to identify SNPs that are balanced in the tumor genome and, in turn, a set of balanced heterozygous segments 706.
[0482] For example, in certain embodiments a balanced test as described above is performed for each SNP using tumor sequencing data, e.g., to determine whether a particular SNP is balanced in the tumor genome.
[0483] Accordingly, in certain embodiments, approaches described herein determine a balanced state of one or more heterozygous SNPs, for example in a tumor genome. In certain embodiments, a balanced state of a given heterozygous SNP is a value that indicates whether or not the given heterozygous SNP is balanced. For example, a balanced state may be a classification, such as “balanced” or “unbalanced”, a Boolean value (e.g., True or False), a numeric value (e.g., 1 or 0), etc.
[0484] Turning to FIGs.8 and 9, example processes for determining primary balanced heterozygous segments and minimal balanced heterozygous segments, respectively, both begin by receiving a list of SNPs and determined balanced states identifying whether particular SNPs are balanced or not within the tumor genome. FIGs.8 and 9 show illustrative plots 830, 930 of balanced states for SNPs located the segments of the example tumor genome illustrated schematically in each figure. The balanced states shown in FIGs.8 and 9 are numeric values - 0 for unbalanced and 1 for balanced SNPs. For sake of simplicity, the figures only show one SNP per segment, but it should be understood that, as described herein, a particular segment may comprise one, or more than one, heterozygous SNP. As shown in FIG.8, in certain embodiments, balanced states are determined for heterozygous SNPs including SNPs within segments 819, 820, where a loss of heterozygosity (LOH) event has occurred (since it may not be known a priori whether a LOH event has occurred). FIG.9 also illustrates SNPs within segments 919, 920 where a LOH event has occurred, and for which balanced states are also determined. In certain embodiments, a balanced state is not determined for segments that do not comprise any heterozygous SNPs.- 104 - 13241894vlAttorney Docket. No.: 2013237-0970
[0485] As illustrated in FIGs.8 and 9, statistical tests and / or other approaches for computing balanced states for SNPs may correctly reflect the true state of the tumor genome at a high level of accuracy, but errors may still occur. Plots 830, 930 show a balanced state for each heterozygous SNP as a red dot, such as 832, 932, which are examples of a heterozygous SNP determined to be balanced. In these examples, the balanced state shown in the figures reflects the correct balanced state for all heterozygous SNPs except for SNPs 834, 934, which are examples of errors that may occur when determining balanced states for individual SNPs.Balance state 834 is an example of an erroneous determination of the balanced property of heterozygous SNP 820. Balanced state 934 is an example of an erroneous determination of the balanced property of heterozygous SNP 920. Such errors can occur when using a statistical test, in particular when the purity of the tumor is low so that the distinction between the state of being balanced and the state of being unbalanced diminishes, and more so when the coverage of the given site is low.
[0486] In certain embodiments, balanced states of one or more heterozygous SNPs may be contour filtered. For example, as shown in FIG. 8, a contour filter to a profile of balance states 836 may be determined. Contour filtering can correct errors in the determination of the balance state of a heterozygous SNP by identifying regions in the genome that have a high density of balanced and / or unbalanced heterozygous SNPs. For example, contour filtering based on the balance state of neighboring heterozygous SNPs determined that the heterozygous SNP 834 is in a locally unbalanced region, and therefore the balance state of heterozygous SNP 834 can be corrected to a contour-filtered balance state of 1 (unbalanced), wherein heterozygous SNPs may be regarded as balanced if they are balanced both before and after contour filtering. As a result, the heterozygous segment corresponding to 820 may be classified as unbalanced and hence discarded from further analysis. Contour filtering is also illustrated in FIG.9, with contour 936 being used to determine that SNP 934 is in a locally unbalanced region, such that an initially determined balance state indicating that SNP 934 is balanced (e.g., a value of 0) may be updated to a value indicative of SNP 934 being unbalanced (e.g., a value of 1). Accordingly, in certain embodiments, contour filtering can be used to discard HSs wrongly classified as BHSs, in particular when the purity is low and more such errors are expected.- 105 - 13241894vlAttorney Docket. No.: 2013237-0970
[0487] Turning again to FIG.7, in certain embodiments, balanced states of individual heterozygous SNPs can be used (e.g., aggregated) to determine segment-level balanced states, and thereby identify a set of balanced heterozygous segments 708. A particular segment may comprise one or more heterozygous SNPs. Accordingly, in certain embodiments, a balanced state of a particular segment may be determined based on balanced states of the one or more heterozygous SNPs that it comprises. For example, a balanced state for a particular segment may be determined as the mean (average), median, mode, etc. of the individual balanced states of the one or more heterozygous SNPs located within the particular segment. In certain embodiments, segment-level balanced states are (e.g., additionally or alternatively to individual SNP balanced states) contour filtered.
[0488] Accordingly, in certain embodiments, a set of balanced heterozygous segments may be identified by selecting segments with balanced states that satisfy certain criteria, such as having balanced states that are at and / or above a particular threshold value (e.g., greater than or equal to 0.5, greater than or equal to 0.6, greater than or equal to 0.7, greater than or equal to 0.8, greater than or equal to 0.9; e.g., equal to 1; e.g., greater than 0.5, greater than 0.6, greater than 0.7, greater than 0.8, greater than 0.9).B.ii.b Decomposing and Identifying Segments According to Absolute Copy Number
[0489] In certain embodiments, a particular subset of BHS’s are identified. For example, as described herein within a tumor genome, various subpopulations of segments can be identified, for example based on whether or not they are balanced and / or their absolute copy number. Accordingly, tumor modeling approaches may identify and isolate a particular subpopulation of segments that (i) are balanced (heterozygous segments) and (ii) have a same particular copy number.
[0490] For example, in certain embodiments, once a set of balanced heterozygous segments are identified 708, sequencing data corresponding to (e.g., tumor and / or normal reads that map to) members of the set of balanced heterozygous segments can be used to identify one or more subpopulations thereof (e.g., of the set of balanced heterozygous segments) 710, each subpopulation of BHSs comprising (e.g., only) segments having a same absolute copy number.- 106 - 13241894vlAttorney Docket. No.: 2013237-0970Since subpopulations are identified within the set of balanced heterozygous segments, each subpopulation is expected to have an even-numbered absolute copy number (e.g., by virtue of the segments all being balanced). Subsets of balanced heterozygous segments corresponding to one or more particular, desired subpopulations may then be selected 712, for example to select segments associated with a most frequently occurring (e.g., even numbered) absolute copy number (e.g., a primary copy number), such as a subset of primary balanced heterozygous segments 714a and / or a minimal (e.g., even numbered) copy number, less than other (e.g., even numbered) absolute copy numbers of other segment subpopulations, such as a set of minimal balanced heterozygous segments 714b.
[0491] For example, in certain embodiments, tumor and normal read counts for each segment can be used to determine a probability density function (pdf) for values of a tumor-to-normal read count ratios 720. For example, a tumor-to-normal read count ratio can be determined for a particular segment as a ratio of tumor read counts to normal read counts for the particular segment (e.g., tumor read count for the particular segment divided by the normal read count for the particular segment). Tumor-to-normal read count ratios may be determined for a plurality of segments, such as BHS’s, to determine a distribution of tumor-to-normal read count ratios for the plurality of segments (e.g., set of BHSs) 720. A distribution may be represented via a pdf (e.g., via creation of a histogram, or other approaches) 722. As illustrated in FIGs.8 and 9, distribution 722 may comprise multiple components, each corresponding to a particular subset of segments having a different absolute copy number.
[0492] For example, FIG.8 shows a graph of an illustrative pdf 840 constructed based on the tumor-to-normal read count ratios for a set of balanced heterozygous segments. As illustrated in FIG.8, pdf 840 comprises multiple components (e.g., normal-like components), appearing as distinct peaks within the plot. Each component reflects a subpopulation of segments having a particular, different, absolute copy number.
[0493] Without wishing to be bound to any particular theory, FIG.8 illustrates a connection between underlying variations in copy numbers from segment to segment and components of pdf 840. For example, BHS’s 813, 814, 815, 816 each have an absolute copy number of four (4) and correspond to component 844 (e.g., centered around a higher mean tumor-to-normal read count ratio), while BHS’s 810, 811, 812 each have an absolute copy- 107 - 13241894vlAttorney Docket. No.: 2013237-0970number of two (2) and correspond to component 846 (e.g., centered around a lower mean tumor-to-normal read count ratio).
[0494] FIG. 9 illustrates a connection between a measured tumor-to-normal read count ratio pdf 940 and underlying variations in copy number from segment to segment as well. In FIG. 9, pdf 940 has two peaks, or components, with a lower mean (left-most) subcomponent corresponding 944 to segments 913, 914, 915, and 916, which are balanced diploid segments, and a higher mean component corresponding 946 segments 910, 911, and 912 each have a copy number of four (4) and, accordingly, correspond to a higher mean peak in pdf 940.
[0495] Accordingly, in certain embodiments, an empirical pdf providing a distribution of tumor-to-normal read count ratios for balanced heterozygous segments may be decomposed into one or more subcomponents 724. Individual components may, accordingly, be distinguished from each other and particular desired subcomponents selected 726. Subsets of segments that are associated with a particular subcomponent may, accordingly, be identified and selected 712.
[0496] For example, turning again to FIGs.8 and 9, in certain embodiments, pdfs such as pdf 840 and pdf 940 can be decomposed into different normal-like components 850 and 950, respectively via a decomposition algorithm, such as a fit to a model pdf. For example, in certain embodiments a one-dimensional Gaussian mixture model (1D-GMM) can be used to approximate and be fit to a pdf such as pdfs 840, 940. Parameters determined via the fit - e.g., an amplitude, mean, and standard deviation of each constituent Gaussian - may, accordingly, be used to select a particular component, such as a primary component that corresponds to a largest subpopulation of underlying segments (e.g., accordingly, a most frequently occurring absolute copy number) and / or a minimal component that corresponds to a subpopulation of segments with a minimal absolute copy number that is smaller than other even-numbered copy numbers, and, accordingly, likely to be two.Selecting Primary Balanced Heterozygous Segments and / or Primary Segments
[0497] Turning to FIG. 8, in certain embodiments, an amplitude of each component of a pdf, such as pdf 840, may be determined and the component having the largest amplitude- 108 - 13241894vlAttorney Docket. No.: 2013237-0970selected as the primary component 854. In this way, in certain embodiments, the component identified as the primary component corresponds to a largest subpopulation of segments, e.g., having a most frequently occurring even-numbered absolute copy number.
[0498] In certain embodiments, a mean tumor- to-normal read count ratio of each component may be determined and used to select the primary component, e.g., alone or in conjunction with the component amplitudes. For example, in certain embodiments, a component having the lowest mean tumor-to-normal read count ratio may be identified, and its amplitude compared with the maximum amplitude (across all components). In certain embodiments, if an amplitude of the component having the lowest mean tumor-to-normal read count ratio is at least (e.g., at or above) a particular minimum faction of the maximum amplitude, then the lowest mean ( ) component is selected as the primary component.
[0499] Among other things, selecting a particular subset of segments - e.g., balanced heterozygous segments - from which to construct tumor-to-normal read count pdf facilitates its decomposition into multiple (e.g., normal-like) components. Among other things, limiting the pdf to balanced heterozygous segments eliminates segments with odd copy numbers, so that the individual components (e.g., peaks) of the distribution are well separated. Accordingly, pdf is amendable to robust decomposition even for low purity tumor samples. Moreover, by selecting a particular subset of balanced heterozygous segments, the range of possible copy numbers of the selected subset (e.g., primary balanced heterozygous segments) is restricted to even integers, thereby reducing the possible candidate primary copy numbers by a factor of two.
[0500] In certain embodiments, one or more auxiliary parameters may be determined from the primary component 854 of pdf 840. In certain embodiments, a mean of the primary component, referred to herein as a “primary slope,” (denoted rpc) may be determined and used as an auxiliary parameter. In certain embodiments, a standard deviation of the primary component, referred to herein as the “primary standard deviation,” (denoted opc) may be determined. In certain embodiments, an amplitude of the primary component, denoted Ape, may be determined. In certain embodiments, as described in further detail herein, values of one or more auxiliary parameters, including, but not limited to, a primary slope (rpc) and / or primary standard deviation (opc) may be used to determine a purity estimate (e.g., a CNV-based purity estimation).- 109 - 13241894vlAttorney Docket. No.: 2013237-0970
[0501] In certain embodiments, a set of primary balanced heterozygous segments are selected 860 from the set balanced heterozygous segments. Parameters 852 characterizing each component of the balanced heterozygous segment tumor-to-normal read count pdf 840, for example individual component amplitudes (Ai), means ( ), and standard deviations (oi), may be used to select a subset of primary balanced heterozygous segments. For example, in certain embodiments, a subset of primary balanced heterozygous segments may be selected from a set of balanced heterozygous segments utilizing, for example, the principle of maximum probability based on the parameters 852 obtained from the decomposition of the pdf. For example, in certain embodiments a maximum a-posteriori probability (MAP) approach can be used to determine a likelihood of a particular heterozygous segment belonging to a particular one of the one or more components of pdf 840. Segments having a highest likelihood for belonging to the primary component may, accordingly, be selected for inclusion in the set of primary balanced heterozygous segments.
[0502] In certain embodiments, after primary balanced heterozygous segments are identified, the set of identified primary balanced heterozygous segments 824 may be used estimate additional auxiliary parameters and / or refine values of those determined (e.g., initially) based on decomposition and identification of primary component 826. For example, in certain embodiments, the set of identified primary balanced heterozygous segments 824 may be used to determine a primary allele frequency standard deviation 870. In certain embodiments, the set of identified primary balanced heterozygous is used to determine a primary residual standard deviation 872.
[0503] For example, in certain embodiments, a primary allele frequency standard deviation may be determined as a standard deviation of observed allele frequencies across the set of primary balanced heterozygous segments 860.
[0504] In certain embodiments, a residual error may be computed for a particular segment to compare (e.g., measure a measure of a difference between) (i) an observed tumor read count for the segment and (ii) an expected tumor read count for the segment predicted based on its normal read count. For example, for a particular, j-th segment, an observed tumor read count tj may be compared with a predicted tumor read count, tj, that is computed according to Eq. (5), below - e.g., as a linear prediction based on that segment’s observed normal read count.- 110 - 13241894vlAttorney Docket. No.: 2013237-0970Eq. (2) tj = a • nj
[0505] The parameter a in Eq. (5) unknown but expected to depend on an absolute copy number of a given segment.
[0506] In certain embodiments, for the set of primary balanced heterozygous segments, an initial estimate of the value of a can be approximated as the mean of the primary component -e.g., the primary slope, rpc (e.g., since rpc is the average tumor-to-normal read count ratio for the subpopulation of segments represented by the primary component).
[0507] In certain embodiments, a particular functional form of a residual error may be based on a variance stabilizing transform that is known or expected to produce a particular distribution of values. For example, in certain embodiments, a residual error may be determined based on observed and expected (e.g., a linear prediction of) tumor read counts accordingly to a variance stabilizing transform such that residual errors are normally distributed when computed across the set of primary balanced heterozygous segments. For example, as demonstrated in Example 1, in certain embodiments, a difference of square root functions may be used as a variance stabilizing transform, with residual error computed as:Eq. (6)
[0508] Other variance stabilizing transformations may be used, including, without limitation a logarithmic transformation (e.g., e = log (t) — log (an)), an arc-sine square root transformation (e.g., e = arcsin( t) — arcsin (Von)), a reciprocal transformation (e.g., e = t-1— (an)-1), an exponential transformation (e.g., e = exp(t) — exp(an)), a box-Cox Transformation (e.g., 1^ = (Y1'-- l) / 2, for A 0; and log(T), for A = 0, where parameter is chosen to best stabilize variance and approximate normality), an Anscombe transform, as well as other possible transformations.- Ill - 13241894vlAttorney Docket. No.: 2013237-0970
[0509] In certain embodiments, residual errors are determined for each segment of the set of primary balanced heterozygous segments, thereby determining a distribution of primary residual errors 874. A standard deviation may be determined, for example, by taking advantage of the expectation that the primary residual errors are normally distributed, by virtue of the variance stabilizing transformation. For example, in certain embodiments, a Gaussian fit to the distribution of primary residual errors may be used with, for example, the standard deviation extracted from a best fit (e.g., by minimizing a least squares error between an empirical probability distribution function of the primary residual error, q, and a Normal distribution with an unknown standard deviation).
[0510] In certain embodiments, a standard deviation of primary residual errors (cre) 872 may be used, in turn, to identify, from all segments 880 (e.g., not just balanced heterozygous segments) a primary segment subpopulation that comprises those segments (e.g., whether heterozygous or not, balanced or not) having an absolute copy number equal to the primary copy number. For example, in certain embodiments, this is performed by selecting segments for which the absolute of the primary residual standard deviation 880 is small in proportion to the primary residual standard deviation 872. In certain embodiments, since the primary residual errors 880 are normally distributed, selection of primary segments benefits from rigorous statistical tolerance.
[0511] Primary segments may, in turn, be used to refine the set of primary balanced heterozygous segments 860, which may be used for CNV-based purity estimation approaches such as those described in Example 1. For example, once primary segments have been precisely estimated based on a rigorous statistical distribution, primary segments can be used to further refine the primary balanced heterozygous segments by checking that each primary balanced heterozygous segment of the initially determined set (e.g., based on an MAP approach and parameters of a 1D-GMM decomposition) also appears in the set of primary segments 890. A refined subset of primary balanced heterozygous segments can be used to refine the primary slope 892 and / or the primary allele frequency standard deviation 894.
[0512] In certain embodiments, as described in further detail herein, the set of primary balanced heterozygous segments and / or auxiliary parameter values determined therefrom may be used to determine various parameters used in tumor models, which are, in turn, fit to observed - 112 - 13241894vlAttorney Docket. No.: 2013237-0970sequencing data. For example, in certain embodiments, auxiliary parameters estimated from primary balanced heterozygous segments may be used to approximate (e.g., fixed) values of certain parameters in tumor models used for (e.g., CNV-based) purity estimation, e.g., thereby reducing a number of variable parameters in tumor models to be fit to observable distributions from sequencing data. In certain embodiments, auxiliary parameters estimated from primary balanced heterozygous segments may be used as initial estimates of values of certain parameters in tumor models used for (e.g., CNV-based) purity estimation, e.g., so as to provide an accurate starting point for a fitting procedure. In certain embodiments, fitting procedures may refine a parameter value, updating it from its initial value taken from an auxiliary parameter. In certain embodiments, as described in further detail herein, tumor model parameters may be established relative to the primary copy number that characterizes the set of primary balanced heterozygous segments, such as rather than estimate and / or fit multiple copy number values, a single, primary, copy number is determined via tumor mode fitting. Among other things, in this way, identifying primary balanced heterozygous segments and determining auxiliary parameters therefrom simplifies and / or makes tractable complex tumor modeling procedures and allows for accurate purity estimates and absolute copy number determinations. As described in further detail herein, CNV-based purity estimation approaches that utilize a subset of primary balanced heterozygous segments may be complementary to and used additionally or alternatively to SNV-based purity estimation techniques described herein. In certain embodiments, additionally or alternatively, VAF analysis methods described herein may be used to determine a minimal and / or primary copy number value and / or a lower bound thereon, which, in turn, can be used to inform and improve CNV-based purity estimation techniques which assign absolute copy number values to segments throughout the tumor genome.Selecting Minimal Balanced Heterozygous Segments and / or Minimal Segments
[0513] Turning to FIG.9, as with procedures for selecting and refining a subset of primary balanced heterozygous segments and / or primary segments described above in regard to FIG. 8, a subset of minimal balanced heterozygous segments and / or minimal segments can be identified and selected using a decomposition of an empirical pdf representing a distribution of - 113 - 13241894vlAttorney Docket. No.: 2013237-0970tumor read counts (e.g., tumor read counts normalized to corresponding normal read counts, as in a tumor-to-normal read count distribution) 720.
[0514] As shown in FIG. 9, in certain embodiments, an mean (ri) of each component of a pdf, such as pdf 940, may be determined and the component having the lowest mean selected as the minimal component 954. In this way, in certain embodiments, the component identified as the minimal component corresponds to a subpopulation of segments having a lowest even numbered absolute copy number - a minimal copy number. As described herein, since normal cells are typically diploid, this minimal copy number is typically two.
[0515] As described herein, among other things, selecting a particular subset of segments - e.g., balanced heterozygous segments - from which to construct tumor-to-normal read count pdf facilitates its decomposition into multiple (e.g., normal-like) components by eliminating segments with odd copy numbers, such that the individual components (e.g., peaks) of the distribution are well separated.
[0516] In certain embodiments, one or more auxiliary parameters may be determined from the minimal component 954 of pdf 940. In certain embodiments, a mean of the minimal component, referred to herein as a “minimal slope,” (denoted rmin) may be determined and used as an auxiliary parameter. In certain embodiments, a standard deviation of the primary component, referred to herein as the “minimal component standard deviation,” (denoted crmin) may be determined. In certain embodiments, an amplitude of the minimal component, denoted Amin, may be determined.
[0517] In certain embodiments, a set of minimal balanced heterozygous segments are selected 960 from the set balanced heterozygous segments. As with selection of primary balanced heterozygous segments, parameters 952 characterizing each component of the balanced heterozygous segment tumor-to-normal read count pdf 940, for example individual component amplitudes (Ai), means ( ), and standard deviations (cq), may be used to select a subset of minimal balanced heterozygous segments. For example, in certain embodiments, a subset of minimal balanced heterozygous segments may be selected from a set of balanced heterozygous segments utilizing, for example, the principle of maximum probability based on the parameters 952 obtained from the decomposition of the pdf. For example, in certain embodiments a maximum a-posteriori probability (MAP) approach can be used to determine a likelihood of a - 114 - 13241894vlAttorney Docket. No.: 2013237-0970particular heterozygous segment belonging to a particular one of the one or more components of pdf 940. Segments having a highest likelihood for belonging to the minimal component may, accordingly, be selected for inclusion in the set of minimal balanced heterozygous segments.
[0518] In certain embodiments, after minimal balanced heterozygous segments are identified, the set of identified minimal balanced heterozygous segments 924 may be used estimate additional auxiliary parameters and / or refine values of those determined (e.g., initially) based on decomposition and identification of minimal component 926.
[0519] In certain embodiments, the set of identified minimal balanced heterozygous is used to determine a minimal residual standard deviation 972. In particular, as described above with regard to primary balanced heterozygous segments, a residual error may be computed for a particular segment to compare (e.g., measure a measure of a difference between) (i) an observed tumor read count for the segment and (ii) an expected tumor read count for the segment predicted based on its normal read count. As described above, this may be accomplished using a linear predictor, for example as in Eq. (5).
[0520] In certain embodiments, for the set of minimal balanced heterozygous segments, an initial estimate of the value of a in Eq. (5) can be approximated as the mean of the minmal component - e.g., the minimal slope, rmin (e.g., since rmin is the average tumor-to-normal read count ratio for the subpopulation of segments represented by the minimal component). As described above in regard to Eq. (6), in certain embodiments, residual errors may be computed through the use of a variance stabilizing transform that is known or expected to produce a particular distribution of values.
[0521] In certain embodiments, residual errors are determined for each segment of the set of minimal balanced heterozygous segments, thereby determining a distribution of minimal (component) residual errors 974. A standard deviation may be determined, for example, by taking advantage of the expectation that the minimal residual errors are normally distributed, by virtue of the variance stabilizing transformation. For example, in certain embodiments, a Gaussian fit to the distribution of primary residual errors may be used with, for example, the standard deviation extracted from a best fit (e.g., by minimizing a least squares error between an empirical probability distribution function of the primary residual error, ej, and a Normal distribution with an unknown standard deviation).- 115 - 13241894vlAttorney Docket. No.: 2013237-0970
[0522] In certain embodiments, a standard deviation of minimal residual errors (σc) 972 may be used, in turn, to identify, from all segments 980 (e.g., not just balanced heterozygous segments) a minimal segment subpopulation that comprises those segments (e.g., whether heterozygous or not, balanced or not) having an absolute copy number equal to the minimal copy number. For example, in certain embodiments, this is performed by selecting segments for which the absolute of the minimal residual standard deviation 980 is small in proportion to the minimal residual standard deviation 972. In certain embodiments, since the minimal residual errors 980 are normally distributed, selection of minimal segments benefits from rigorous statistical tolerance.
[0523] Minimal segments may, in turn, be used to refine the set of minimal balanced heterozygous segments 960, which may be used for SNV-based purity estimation approaches such as those described in Example 1. For example, once minimal segments have been precisely estimated based on a rigorous statistical distribution, minimal segments can be used to further refine the minimal balanced heterozygous segments by checking that each primary balanced heterozygous segment of the initially determined set (e.g., based on an MAP approach and parameters of a 1D-GMM decomposition) also appears in the set of minimal segments 890.B.ii.c Identifying Minimal Balanced SNVs
[0524] Turning again to FIG.6, in certain embodiments, subsets of minimal balanced heterozygous segments and / or minimal segments may be used to identify a subset of minimal balanced SNV events that can be used for a SNV-based purity estimation procedures, such as process 600.
[0525] In certain embodiments, a subset of minimal balanced SNV events may be identified from SNV list 604 using a subset of minimal balanced heterozygous segments. For example, in certain embodiments, SNV events of SNV list 604 that are located within a minimal balanced heterozygous segment can be identified as minimal balanced SNV events.
[0526] In certain embodiments, a subset of minimal balanced SNV events may be identified from SNV list using a subset of minimal segments followed by one or more filtering steps to identify SNV events that are located within balanced regions of a tumor genome. For- 116 - 13241894vlAttorney Docket. No.: 2013237-0970example, in certain embodiments, SNV events of SNV list 604 that are located within a minimal segment can be identified. SNV events determined to be located within a minimal segment may then be evaluated to identify those that are within balanced regions and, accordingly, are minimal balanced SNV events.
[0527] SNVs may be determined to be balanced using various balanced tests and / or filtering (e.g., contour filtering) approaches described herein. In certain embodiments, a SNV may be identified as balanced based on one or more nearby SNPs. For example, in certain embodiments, a balanced state may be determined for one or more nearby SNPs using statistical tests and / or filtering, such as contour filtering, approaches. SNVs located within a certain distance of one or more balanced SNPs, in between two balanced SNPs, etc. may accordingly be identified as balanced and selected for inclusion in a set of minimal balanced SNV events. In certain embodiments, SNVs selected for inclusion in a set of minimal balanced SNV events may also be determined to be balanced in a normal genome, which may be a reference genome obtained e.g., from a database and / or may be determined via sequencing data, such as normal sequencing data obtained from a normal sample obtained from the subject. Other filtering approaches, such as coverage and / or quality score filters may also be used. Additionally, or alternatively, once an initial subset of minimal balanced SNV events is selected, outliers associated with unbalanced SNV events (e.g., that were not removed by filtering approaches used to obtain the initial subset of minimal balanced SNV events) may be identified and removed, e.g., via VAF filtering approaches described herein.B.iii SNV-Based Tumor Modeling Using Minimal Balanced Heterozygous Segments
[0528] Turning again to FIG.6, in certain embodiments, tumor deconvolution technologies described herein use a particular subset of SNV events 606, such as a subset of minimal balanced SNV events, to estimate a purity of a tumor sample.
[0529] For example, in certain embodiments, one or more sequencing data observations are determined for the particular subset of SNV events 608. For example, quantities such as allele- specific read counts, variant allele frequencies, etc. and / or distributions thereof may be determined for a subset of minimal balanced SNV events. For example, allele- specific read counts may be determined for each of one or more alleles of a given site at which a putative SNV - 117 - 13241894vlAttorney Docket. No.: 2013237-0970event occurs. For example, a SNV event may represent an underlying single nucleotide substitution at a particular site within a tumor genome, wherein a normal, wild-type, allele has a first nucleotide at the particular site and an alternate, mutated, allele has a second, different, nucleotide at the particular site. Allele specific read counts may be determined for the normal and alternate alleles as the number of reads having base calls corresponding to the first and second nucleotides, respectively, at the particular site. In certain embodiments, allele- specific read counts are also determined for two remaining noise alleles. Additionally, or alternatively, a VAF may be determined for a particular allele as a fraction of reads corresponding to the alternate allele (e.g., relatively to all reads mapping to the particular site). In certain embodiments, sequencing data observations, such as allele- specific read counts and / or VAFs may be determined using a filtered portion of sequencing data reads, for example, to exclude those with coverage levels and / or quality scores below particular threshold values.
[0530] Measured sequencing data observations 608 may be used in connection with a tumor model 612 to determine 610 one or more purity estimates. For example, a tumor model 612 may be used to determine predicted values and / or distributions of the sequencing data observations as a function of one or more variable parameters. Values and / or distributions of tumor model predictions can then be compared, and fit to, those of the measured sequencing data observations. The one or more variable parameters may include a tumor sample purity variable parameter, such that tumor model predictions of sequencing data observations can be fit to measured sequencing data observations by adjusting a value of the tumor sample purity parameter. In this way, an estimate of tumor sample purity may be obtained by determining the value of the variable tumor sample purity that optimizes the fit between the predicted sequencing data observations of the tumor model and the corresponding measured sequencing data observations.
[0531] For example, a tumor model may be a SNV model 612 that predicts frequencies with which different alleles - e.g., an alternate (e.g., mutated) allele and a normal (e.g., wildtype) allele, as well as, in certain embodiments, noise alleles - are expected to be observed in sequencing measurements. An SNV model 612 may, for example, assume mutations are biologically clonal and occur in balanced diploid regions of a tumor genome and predict, as a function of tumor sample purity and, optionally, a sequencing error rate, frequencies with which- 118 - 13241894vlAttorney Docket. No.: 2013237-0970normal, alternate, and noise alleles are expected to be observed in sequencing data. Agreement between predictions made by an SNV model 612 and measured sequencing data observations 608 can be quantified using various metric functions. For example, in certain embodiments, a likelihood function (e.g., using a maximum likelihood-based approach) 616a that measures, for each SNV event, a likelihood of measuring a corresponding set of allele- specific read counts as a function of tumor sample purity. An overall likelihood function may then be determined based on an aggregation (e.g., a sum, a product) of all individual SNV likelihood functions for the subset of minimal balanced SNV events. In certain embodiments, additionally or alternatively, SNV model may be used to generate, as a function of tumor sample purity, a predicted distribution of VAFs for the set of minimal balanced SNVs, which can be compared with the measured VAF distribution 616b. Various metrics, including, but not limited to, a Kullback-Leibler (KL) divergence function, may be used to quantify agreement between predicted and measured VAF distributions. Accordingly, values of the tumor sample purity parameter can be tested and a best-fit value that maximizes or minimizes one or more metric functions determined 610. Various procedures may be used to test and search for best-fit values of the tumor sample purity parameter using a given metric function, including, without limitation, grid searching, gradient descent, and the like.
[0532] Example 1 demonstrates use of maximum likelihood estimation (MLE) and KL-divergence-based approaches for determining tumor sample purities.
[0533] In certain embodiments, SNV models may allow for multiple SNV genotypes, rather than e.g., restrict SNV events to a single genotype. For example, in a diploid region, in principle, a SNV event may give rise to a mutation on one or both alleles. Without wishing to be bound to any particular theory, for SNVs occurring in diploid, balanced, heterozygous regions of the tumor genome, SNV events are expected to be limited to mutations on a single allele, since the likelihood of two identical mutations (e.g., a same substitution at a same site) independently occurring on two different alleles is expected to be small. In certain embodiments, however, rather than assume SNV events are heterozygous, a SNV model and / or fitting procedure may allow for the possibility of homozygous SNVs, for example by computing expected read counts or VAFs for both genotypes and selecting the maximum likelihood genotype and / or by modeling SNV events as a mixture of heterogenous and homozygous genotypes with the fraction of- 119 - 13241894vlAttorney Docket. No.: 2013237-0970homozygous genotypes allowed to vary as a free parameter that can be optimized along with tumor sample purity. In this way, a number and / or fraction of homozygous SNVs can be determined while estimating purity 610. Since homozygous SNVs are expected to be rare, the number and / or fraction of homozygous SNVs can be used as a quality control check indicative of a reliability of purity estimation procedure 600 and / or a particular model approach used to determine purity 610.
[0534] In certain embodiments, once initial purity estimates are obtained, they may be used to determine a (e.g., final) purity estimate 614. For example, in certain embodiments, a single initial purity estimate is obtained, for example via a MLE 616a or KL divergence 616b approach and used as purity estimate 614. In certain embodiments, multiple initial purity estimates, such as (e.g., both) a MLE 616a and KL divergence 616b based purity estimate are obtained and a particular one of the multiple initial purity estimates selected as a final purity estimate 614, for example based on one or more selection criteria.B.iv CNV-Based Tumor Modeling Using Primary Balanced Heterozygous Segments
[0535] Turning to FIGs. 10 and 11, in certain embodiments, SNV-based tumor modeling technologies of the present disclosure may be used in connection with CNV-based approaches that leverage a set of identified PBHSs and / or various auxiliary parameters determined therefrom to estimate a purity of a tumor sample. In certain embodiments, additionally or alternatively, tumor modeling technologies described herein may determine a value of a primary copy number - i.e., the absolute copy number of primary segments. As described in further detail herein, once determined, the primary copy number may be used as a point of reference from which absolute copy numbers of other segments (e.g., not just the primary segments) can be determined.
[0536] As illustrated in the schematic of FIG. 11, procedures described herein for estimating sample purity may include a maximum likelihood procedure whereby one or more initial estimates of sample purity are determined 1122, each corresponding to a candidate primary copy number 1120. In certain embodiments, thereafter, a classification step 1124 is included to select a particular one of the one or more candidate primary copy numbers and a final determined primary copy number for the tumor sample. A final estimated sample purity may then be determined as the initial purity estimate corresponding to the selected primary copy - 120 - 13241894vlAttorney Docket. No.: 2013237-0970number 1130a. In certain embodiments, as explained herein, quality control procedures 1126 may be used to determine whether the final estimated sample purity may be retained 1128a, and used for further analysis (e.g., mutation detection and / or genomically characterizing mutations), or if a bound (such as an upper bound) or other value (e.g., determined from an alternative pathway, such as a SNV-based purity estimate) should be used instead 1130b.
[0537] As in further detail herein, results of tumor modeling may be used to (1) improve the sensitivity and accuracy of mutation detection and (2) genomically characterize mutations, including performing a clonality analysis.
[0538] FIG. 10 shows a schematic of an example procedure for tumor modeling, e.g., for the improving mutation detection as well as characterizing mutations genomically. In the example schematic of FIG. 10, tumor modeling and its use is illustrated in a modular fashion, with CNV-based purity estimation, purity estimation logic, absolute copy number estimation, refinement of a list of putative SNVs, zygosity estimation and subclonality estimation each being performed by various sub steps. Among other things, FIG....
Claims
Attorney Docket. No.: 2013237-0970CLAIMSWhat is claimed is:
1. A method comprising:(a) obtaining, by a processor of a computing device, tumor sequencing data for a tumor sample obtained from a subject, the tumor sequencing data representing a tumor genome associated with the tumor sample;(b) detecting, by the processor, using on the tumor sequencing data, a plurality of initial events, each initial event corresponding to a variant portion of the tumor genome, wherein, for each initial event, the corresponding variant portion is determined, based on the sequencing data, to have a different copy number and / or a different nucleotide sequence relative to a corresponding normal reference;(c) determining, by the processor, an estimated purity for the tumor sample and / or a bound thereon;(d) filtering, by the processor, the plurality of initial events to exclude initial events identified as false-positives based on (i) the tumor sequencing data and (ii) the estimated purity for the tumor sample and / or bound thereon and retaining other initial events as putative mutations for inclusion in a mutation list; and(e) storing and / or providing, for display and / or further processing, the mutation list.
2. The method of claim 1, wherein, for each of at least a portion of the plurality of initial events, the corresponding variant portion is determined, based on the sequencing data, to have a different nucleotide sequence relative to the corresponding normal reference.
3. The method of claim 2, wherein the different nucleotide sequence is or comprises an alternate allele, different from an allele of the corresponding normal reference.- 326 - 13241894vlAttorney Docket. No.: 2013237-09704. The method of claim 2 or claim 3, wherein the different nucleotide sequence is or comprises an insertion and / or a deletion relative to the corresponding normal reference.
5. The method of any one of claims 2-4, wherein the variant portion is or comprises a structural variation relative to the normal reference.
6. The method of any one of claims 1 to 5, wherein, for each of at least a portion of the plurality of initial events, the corresponding variant portion is determined, based on the sequencing data, to have a different copy number than the corresponding normal reference.
7. The method of any one of claims 1-6, wherein, for each of at least a portion of the plurality of initial events, the corresponding variant portion is determined, based on the sequencing data, to have a different nucleotide sequence relative to the corresponding normal reference.
8. The method of claim 7, wherein the different nucleotide sequence is or comprises an alternate allele, different from a normal allele of the corresponding normal reference.
9. The method of claim 8, wherein step (d) comprises, for each initial event of at least a portion of the initial events:determining, by the processor, a mutation confidence score that quantifies (i) a likelihood of the initial event encoding the alternate allele and thereby representing a true underlying mutation relative to (ii) a likelihood of the initial event encoding the normal allele, wherein the mutation confidence score for a given initial event is determined based at least in part on (i) reads of the tumor sequencing data that map to the given initial event and (ii) a set of tumor model parameters; andidentifying, by the processor, the initial event as (i) a false positive to be excluded from the mutation list or (ii) a putative mutation for inclusion in the mutation list based on the- 327 - 13241894vlAttorney Docket. No.: 2013237-0970mutation confidence score determined for the initial event in comparison with a mutation confidence score threshold value.
10. The method of claim 9, wherein the set of tumor model parameters comprises a measured tumor content of the tumor sample.
11. The method of claim 9 or claim 10, wherein the set of tumor model parameters comprises one or more predetermined purity values and / or one or more predetermined copy numbers.
12. The method of any one of claims 9-11, wherein the set of tumor model parameters comprises the estimated purity and / or bound thereon for the tumor sample, such that the mutation confidence score is a refined mutation confidence score.
13. The method of any one of claims 9-12, wherein the mutation confidence score is a ratio of the likelihood of the initial event encoding the alternate allele to the likelihood of the initial event encoding the normal allele.
14. The method of any one of claims 9-13, wherein the mutation confidence score is a difference of a log-likelihood of the initial event encoding the alternate allele and a loglikelihood of the initial event encoding the normal allele.
15. The method of any one of claims 8-14, comprising:prior to step (c), filtering the plurality of initial events by, for each initial event of at least a portion of the initial events:determining, by the processor, a mutation confidence score that quantifies (i) a likelihood of the initial event encoding the alternate allele and thereby representing a true underlying mutation relative to (ii) a likelihood of the initial event encoding the normal - 328 - 13241894vlAttorney Docket. No.: 2013237-0970allele, wherein the mutation confidence score is determined, for a given initial event, based at least in part on (i) reads of the tumor sequencing data that map to the given initial event and (ii) a set of tumor model parameters, and wherein the set of tumor model parameters comprises a predetermined set of one or more purity values and / or range of purity values; andidentifying, by the processor, the initial event as a false positive to be excluded from a preliminary mutation list or identifying, by the processor, the initial event as a putative mutation for inclusion in the preliminary mutation list based on the mutation confidence score determined for the initial event in comparison with a mutation confidence score threshold value,thereby determining the preliminary mutation list, comprising a set of initial events retained following filtering according to their mutation confidence scores; and at step (c), using the preliminary mutation list to determine the estimated purity for the tumor sample and / or the bound thereon.
16. The method of claim 15, wherein step (d) comprises, following step (c):updating the set of tumor model parameters according to the estimated tumor sample purity and / or bound thereon, thereby determining a refined set of tumor model parameters; and for each initial event identified as a putative mutation and included in the preliminary mutation list:determining, by the processor, a refined mutation confidence score that quantifies (i) a likelihood of the initial event encoding the alternate allele and thereby representing a true underlying mutation relative to (ii) a likelihood of the initial event encoding the wildtype allele, wherein the refined mutation confidence score is determined, for a given initial event, based at least in part on (i) reads of the tumor sequencing data that map to the given initial event and (ii) the refined set of tumor model parameters; and identifying, by the processor, the initial event as a false positive to be excluded from the mutation list or retaining, by the processor, the identification of the initial event- 329 - 13241894vlAttorney Docket. No.: 2013237-0970as a putative mutation for inclusion in the mutation list based on the refined mutation confidence score determined for the initial event in comparison with a refined mutation confidence score threshold value.
17. The method of claim 16, wherein the refined mutation confidence score threshold value is lower than the mutation confidence score threshold value.
18. The method of any one of the preceding claims, wherein step (d) comprises, for each initial event of at least a portion of the initial events:determining, by the processor, a value of a test statistic for the initial event based at least in part on the tumor sequencing data;determining, by the processor, a mutation confidence score that quantifies (i) a likelihood of the initial event encoding the alternate allele and thereby representing a true underlying mutation relative to (ii) a likelihood of the initial event encoding the wild-type allele, wherein the mutation confidence score is determined, for a given initial event, based at least in part on (i) reads of the tumor sequencing data that map to the given initial event and (ii) a set of tumor model parameters;selecting, by the processor, a value of a discrimination threshold based at least in part on the mutation confidence score determined for the initial event; andidentifying, by the processor, the initial event as (i) a false positive to be excluded from the mutation list or (ii) a putative mutation for inclusion in the mutation list based on the value of the test statistic determined for the initial event and the selected value of the discrimination threshold.
19. The method of claim 18, wherein the set of tumor model parameters comprises a measured tumor content of the tumor sample.- 330 - 13241894vlAttorney Docket. No.: 2013237-097020. The method of claim 18 or claim 19, wherein the set of tumor model parameters comprises one or more predetermined purity values and / or one or more predetermined copy numbers.
21. The method of any one of claims 18-20, wherein the set of tumor model parameters comprises the estimated purity and / or bound thereon for the tumor sample, such that the mutation confidence score is a refined mutation confidence score.
22. The method of any one of claims 18-21,wherein selecting the value of the discrimination threshold comprises selecting one of a plurality of threshold value settings, each threshold value setting associated with a range of refined mutation confidence scores, andwherein lower threshold value settings are associated with higher-valued ranges of refined mutation confidence scores.
23. The method of claim 22, wherein selecting the value of the discrimination threshold comprises adjusting a threshold value setting to reduce the value of the discrimination threshold as the mutation confidence score increases or to increase the value of the discrimination threshold as the refined mutation confidence score decreases.
24. The method of any one of the preceding claims, wherein step (d) comprises:selecting, by the processor, value(s) of one or more discrimination thresholds based on the estimated purity and / or bound thereon; andusing, by the processor, the selected value(s) of the one or more discrimination thresholds to filter the plurality of initial events.- 331 - 13241894vlAttorney Docket. No.: 2013237-097025. The method of any one of the preceding claims, wherein step (d) comprises, for each initial event of at least a portion of the initial events:determining, by the processor, a value of a test statistic for the initial event based at least in part on the tumor sequencing data;selecting, by the processor, a value of a discrimination threshold based at least in part on the estimated purity for the tumor sample and / or the bound thereon; andidentifying, by the processor, the initial event as (i) a false positive to be excluded from the mutation list or (ii) a putative mutation for inclusion in the mutation list based on the value of the test statistic determined for the initial event and the selected value of the discrimination threshold.
26. The method of claim 25, comprising selecting the value of the discrimination threshold from one of a plurality of threshold settings,wherein each of the plurality of threshold settings is associated with a range of tumor sample purity values and a corresponding value for the discrimination threshold, and wherein threshold settings associated with ranges of lower tumor sample purities have lower corresponding values for the discrimination threshold than threshold settings associated with ranged of higher tumor sample purities.
27. The method of claim 25 or claim 26, wherein selecting the value of the discrimination threshold comprises reducing the value of the discrimination threshold as the estimated tumor sample purity decreases or increasing the value of the discrimination threshold as the estimated tumor sample purity increases.
28. The method of any one of claims 25 to 27, wherein the test statistic is a mutation confidence score and / or a refined mutation confidence score.- 332 - 13241894vlAttorney Docket. No.: 2013237-097029. The method of any one of claims 25 to 28, wherein the test statistic is determined based on the estimated tumor sample purity and / or bound thereon.
30. The method of any one of claims 25 to 29, wherein the test statistic is determined based on a ploidy of the tumor sample.
31. The method of any one of claims 25 to 30, the test statistic is determined based on a measured histological tumor content of the tumor sample.
32. The method of any one of claims 25 to 31, wherein the value of the discrimination threshold is selected based at least in part on a ploidy of the tumor sample.
33. The method of any one of claims 25 to 32, wherein the value of the discrimination threshold is selected based at least in part on a value of a secondary test statistic, said secondary test statistic determined as a function of a tumor-model derived parameter.
34. The method of any one of the preceding claims, wherein step (d) comprises:selecting, by the processor, a particular false positive filter of a plurality of false positive filters based at least in part on the estimated tumor sample purity; andidentifying, by the processor, the initial event as (i) a false positive to be excluded from the mutation list or (ii) a putative mutation for inclusion in the mutation list, using the particular selected false positive filter.
35. The method of claim 34, wherein each false positive filter of the plurality of false positive filters is associated with a range of tumor sample purities.- 333 - 13241894vlAttorney Docket. No.: 2013237-097036. The method of claim 35, wherein the plurality of false positive filters comprises a series of filters associated with ranges corresponding progressively higher tumor sample purity vales and wherein a stringency of filters of the series increases with increasing tumor sample purity.
37. The method of claim 35 or 36, comprising determining, by the processor, the estimated tumor sample purity to be within the range of tumor sample purities associated with the selected false positive filter, thereby selecting the false positive filter associated with the range of tumor sample purities that includes the estimated tumor sample purity.
38. The method of any one of claims 34-37, comprising:determining, by the processor, a value of a tumor sample purity-derived parameter based on the estimated tumor sample purity; andselecting, by the processor, the particular false positive filter of the plurality of false positive filters based on the value of the tumor-sample purity-derived parameter.
39. The method of claim 38, wherein each false positive filter of the plurality of false positive filters is associated with a range of values for the tumor sample purity-derived parameter.
40. The method of claim 39, comprising determining, by the processor, the tumor sample purity-derived parameter to be within the range of tumor sample purities associated with the particular selected false positive filter, thereby selecting the false positive filter associated with the range of tumor sample purities that includes the tumor sample purity-derived parameter.
41. The method of any one of claims 38-40, wherein the tumor sample purity-derived parameter is a mutation confidence score that quantifies (i) a likelihood of the initial event encoding the alternate allele and thereby representing a true underlying mutation relative to (ii) a likelihood of the initial event encoding the wild-type allele, wherein the mutation confidence score for a given initial event is determined based at least in part on (i) reads of the tumor- 334 - 13241894vlAttorney Docket. No.: 2013237-0970sequencing data that map to the given initial event and (ii) a set of one or more tumor model parameters.
42. The method of claim 41, wherein the plurality of false positive filters comprises a series of filters of differing levels of stringency, and wherein a stringency of filters of the series increases with decreasing values of the mutation confidence score.
43. The method of any one of the preceding claims, wherein step (d) comprises using a strand bias filter that measures a likelihood of strand bias for a given initial event based on a number of reads measured in a forward direction and a reverse direction for each of the alternate allele and the wild-type allele.
44. The method of claim 43, comprising:determining, for a given initial event, a forward direction mutation confidence score based on forward reads mapping to the given initial event and a reverse direction mutation confidence score based on reverse reads mapping to the given initial event; andfiltering or retaining the given initial event based at least in part on (i) the likelihood of strand bias, (ii) the forward direction mutation confidence score and (iii) the reverse direction mutation confidence score.
45. The method of claim 44, comprising:determining, for a given initial event:a forward direction refined mutation confidence score based on forward reads mapping to the given initial event and the estimated tumor sample purity; anda reverse direction refined mutation confidence score based on reverse reads mapping to the given initial event and the estimated tumor sample purity; and- 335 - 13241894vlAttorney Docket. No.: 2013237-0970filtering or retaining the given initial event based at least in part on (i) the likelihood of strand bias, (ii) the forward direction refined mutation confidence score and (iii) the reverse direction refined mutation confidence score.
46. The method of any one of the preceding claims, comprising obtaining, by the processor, normal sequencing data and wherein step (d) comprises using a normal coverage filter that excludes initial events if coverage for corresponding locations in the normal sequencing data is below a coverage level threshold.
47. The method of claim 46, comprising, at step (d), determining, by the processer, a refined mutation confidence score based on tumor reads mapping to the initial event and retaining initial events whose refined mutation confidence score is above a predetermined threshold level.
48. The method of any one of the preceding claims, wherein step (c) comprises using a measured histological tumor content of the tumor sample as the estimated purity for the tumor sample.
49. The method of any one of the preceding claims, wherein step (c) comprises determining the estimated purity for the tumor sample and / or bound thereon based at least in part on the tumor sequencing data.
50. The method of claim 49, comprising determining, by the processor, the estimated purity for the tumor sample and / or bound thereon using a copy number variation (CNV)-based purity estimation procedure based on the tumor sequencing data.
51. The method of claim 50, comprising determining absolute copy numbers and / or allele specific copy numbers for one or more segments within the tumor genome using the CNV-based purity estimation procedure.- 336 - 13241894vlAttorney Docket. No.: 2013237-097052. The method of claim 50 or 51, wherein step (d) comprises filtering the plurality of initial events based at least in part on the absolute copy numbers and / or allele specific copy numbers.
53. The method of any one of claims 49- 52, comprising determining the estimated purity for the tumor sample and / or bound thereon using a single nucleotide variant (SNV)-based purity estimation procedure that models expected variant allele frequencies (VAF) and / or an expected distribution of VAFs for at least a portion of the plurality of initial events.
54. The method of claim 53, comprising modeling the initial events as biologically clonal.
55. The method of any one of the preceding claims, wherein a measured purity of the tumor sample is less than about 0.4 and / or wherein the estimated purity of the tumor sample is less than about 0.3.
56. The method of any one of the preceding claims, comprising obtaining, by the processor, normal sequencing data for a normal sample obtained from the subject and using the normal sequencing data and the tumor sequencing data to determine, and filter initial events based on, a likelihood that a given initial event arises from noise.
57. The method of any one of the preceding claims, wherein the tumor sequencing data comprises RNA sequencing data (RNAseq) and the method comprises using the RNAseq data to filter errors.
58. A method, the method comprising:- 337 - 13241894vlAttorney Docket. No.: 2013237-0970(a) obtaining, by a processor of a computing device, tumor sequencing data for a tumor sample obtained from a subject, the tumor sequencing data representing a tumor genome associated with the tumor sample;(b) detecting, by the processor, using the tumor sequencing data, a plurality of initial events;(c) iteratively: filtering, by the processor, the plurality of initial events based on a set of one or more tumor model parameters, to generate a filtered list of putative mutations and using, by the processor, the filtered list of putative mutations to refine the set of tumor model parameters,wherein the refined set of tumor model parameters determined at one iteration is used to filter the plurality of initial events in a subsequent iteration, and,at a final iteration, the plurality of initial events is filtered based on a final refined set of tumor model parameters determined at a previous iteration, thereby generating a final list of putative mutations; and(d) storing and / or providing, by the processor, the final list of putative mutations for display and / or further processing.
59. The method of claim 58, wherein the filtering is performed by the method of any one of claims 1-47.
59. The method of claim 58, wherein at a first iteration, the set of tumor model parameters is an initial set of tumor model parameters, having been determined and / or obtained without using the plurality of initial events and / or the filtered list of putative mutations and, optionally, wherein step (c) comprises, at a first iteration, filtering the plurality of initial events based on a generic set of tumor model parameters comprising, for each of the one or more tumor model parameters of the set, (i) a range of expected values and / or (ii) an expected value.- 338 - 13241894vlAttorney Docket. No.: 2013237-097061. The method of claim 60, wherein filtering the plurality of initial events based on the generic set of tumor model parameters comprises, for each of at least a portion of the initial events:determining, by the processor, a mutation confidence score that quantifies (i) a likelihood of the initial event encoding an alternate allele and thereby representing a true underlying mutation relative to (ii) a likelihood of the initial event encoding a normal allele, wherein the mutation confidence score is determined, for a given initial event, based at least in part (i) reads of the tumor sequencing data that map to the given initial event and (ii) the generic set of tumor model parameters; andidentifying, by the processor, the initial event as a false positive to be excluded from a preliminary mutation list or identifying, by the processor, the initial event as a putative mutation for inclusion in the preliminary mutation list based on the mutation confidence score determined for the initial event in comparison with a mutation confidence score threshold value.
62. The method of claim 2, wherein step (c) comprises, at a first iteration:determining, by the processor, an initial set of tumor model parameters using a copy number variation (CNV)-based purity estimation procedure based on the tumor sequencing data; andfiltering the plurality of initial events based on the initial set of tumor model parameters.
63. The method of claim 62, comprising determining absolute copy numbers and / or allele specific copy numbers for one or more segments within the tumor genome using the CNV-based purity estimation procedure.
64. The method of claim 63, wherein step (d) comprises filtering the plurality of initial events based at least in part on the absolute copy numbers and / or allele specific copy numbers.
65. The method of any one of claims 58-64, comprising, at a first iteration:- 339 - 13241894vlAttorney Docket. No.: 2013237-0970obtaining a measured tumor content of the tumor sample and including the measured tumor content in an initial set of tumor model parameters; andfiltering the plurality of initial events based on the initial set of tumor model parameters.
66. The method of any one of 58-65, comprising, at a first iteration:obtaining one or more predetermined purity values and / or one or more predetermined copy numbers and including the measured tumor content and / or one or more predetermined copy numbers in an initial set of tumor model parameters; andfiltering the plurality of initial events based on the initial set of tumor model parameters67. The method of any one of the preceding claims, wherein at step (c), for one or more iterations, refining the set of tumor model parameters comprises determining, by the processor, an estimated tumor sample purity and / or bound thereon based on (i) the filtered list of putative mutations and (ii) the tumor sequencing data.
68. The method of claim 67, wherein determining the estimated tumor sample purity and / or bound thereon comprises using a single nucleotide variant (SNV)-based purity estimation procedure that models expected variant allele frequencies (VAF) and / or an expected distribution of VAFs for at least a portion of the filtered list of putative mutations.
69. The method of claim 67 or 68, wherein determining the estimated tumor sample purity and / or bound thereon comprises using a copy number variation (CNV)-based purity estimation procedure based on the tumor sequencing data.
70. The method of claim 69, comprising determining absolute copy numbers and / or allele specific copy numbers for one or more segments within the tumor genome using the CNV-based purity estimation procedure.- 340 - 13241894vlAttorney Docket. No.: 2013237-097071. The method of any one of claims 67-70, comprising determining a VAF distribution based on the filter list of putative mutations.
72. The method of any one of claims 67 to 71, wherein determining the estimated tumor sample purity and / or bound thereon comprises:determining a first estimated tumor sample purity and / or a first tumor sample purity bound using a single nucleotide variant (SNV)-based purity estimation procedure;determining a second estimated tumor sample purity and / or a second tumor sample purity bound using a copy number variation (CNV)-based purity estimation procedure that fits a predicted a copy number grid to a distribution of read counts and major allele frequencies determined for a plurality of heterozygous segments within a tumor genome; andselecting, by the processor, for use as the estimated tumor sample purity and / or bound thereon, one of: the first estimated tumor sample purity, the first tumor sample purity bound, the second estimated tumor sample purity, and the second tumor sample purity bound.
73. The method of any one of claims 58-72, wherein each initial event corresponds to a variant portion of the tumor genome, wherein, for each initial event, the corresponding variant portion is determined, based on the sequencing data, to have a different copy number and / or a different nucleotide sequence relative to a corresponding normal reference.
74. The method of claim 73, wherein, for each of at least a portion of the plurality of initial events, the corresponding variant portion is determined, based on the sequencing data, to have a different nucleotide sequence relative to the corresponding normal reference.
75. The method of any one of claims 73-74, wherein the different nucleotide sequence is or comprises an alternate allele, different from an allele of the corresponding normal reference.- 341 - 13241894vlAttorney Docket. No.: 2013237-097076. The method of any one of claims 73-75, wherein the different nucleotide sequence is or comprises an insertion and / or a deletion relative to the corresponding normal reference.
77. The method of any one of claims 73-76, wherein the variant portion is or comprises a structural variation relative to the normal reference.
78. The method of any one of claims 73-77, wherein, for each of at least a portion of the plurality of initial events, the corresponding variant portion is determined, based on the sequencing data, to have a different copy number than the corresponding normal reference.
79. The method of any one of claims 73-78, wherein, for each of at least a portion of the plurality of initial events, the corresponding variant portion is determined, based on the sequencing data, to have a different nucleotide sequence relative to the corresponding normal reference.
80. The method of claim 79, wherein the different nucleotide sequence is or comprises an alternate allele, different from a normal allele of the corresponding normal reference.
81. The method of any one of claims 58-80, wherein step (d) comprises, for each initial event of at least a portion of the initial events:determining, by the processor, a mutation confidence score that quantifies (i) a likelihood of the initial event encoding the alternate allele and thereby representing a true underlying mutation relative to (ii) a likelihood of the initial event encoding the normal allele, wherein the mutation confidence score for a given initial event is determined based at least in part on (i) reads of the tumor sequencing data that map to the given initial event and (ii) a set of tumor model parameters; and- 342 - 13241894vlAttorney Docket. No.: 2013237-0970identifying, by the processor, the initial event as (i) a false positive to be excluded from the mutation list or (ii) a putative mutation for inclusion in the mutation list based on the mutation confidence score determined for the initial event in comparison with a mutation confidence score threshold value.
82. The method of any one of claim 81, wherein the set of tumor model parameters comprises the estimated purity and / or bound thereon for the tumor sample, such that the mutation confidence score is a refined mutation confidence score.
83. The method of any one of claims 58-82, wherein step (d) comprises, for each initial event of at least a portion of the initial events:determining, by the processor, a value of a test statistic for the initial event based at least in part on the tumor sequencing data;determining, by the processor, a mutation confidence score that quantifies (i) a likelihood of the initial event encoding the alternate allele and thereby representing a true underlying mutation relative to (ii) a likelihood of the initial event encoding the wild-type allele, wherein the mutation confidence score is determined, for a given initial event, based at least in part on (i) reads of the tumor sequencing data that map to the given initial event and (ii) a set of tumor model parameters;selecting, by the processor, a value of a discrimination threshold based at least in part on the mutation confidence score determined for the initial event; andidentifying, by the processor, the initial event as (i) a false positive to be excluded from the mutation list or (ii) a putative mutation for inclusion in the mutation list based on the value of the test statistic determined for the initial event and the selected value of the discrimination threshold.- 343 - 13241894vlAttorney Docket. No.: 2013237-097084. The method of claim 83, wherein the set of tumor model parameters comprises the estimated purity and / or bound thereon for the tumor sample, such that the mutation confidence score is a refined mutation confidence score.
85. The method of any one of claims 58-84, wherein step (d) comprises:selecting, by the processor, value(s) of one or more discrimination thresholds based on the estimated purity and / or bound thereon; andusing, by the processor, the selected value(s) of the one or more discrimination thresholds to filter the plurality of initial events.
86. The method of any one of claims 58-85, wherein step (d) comprises, for each initial event of at least a portion of the initial events:determining, by the processor, a value of a test statistic for the initial event based at least in part on the tumor sequencing data;selecting, by the processor, a value of a discrimination threshold based at least in part on the estimated purity for the tumor sample and / or the bound thereon; andidentifying, by the processor, the initial event as (i) a false positive to be excluded from the mutation list or (ii) a putative mutation for inclusion in the mutation list based on the value of the test statistic determined for the initial event and the selected value of the discrimination threshold.
87. The method of any one of claim 86, wherein step (d) comprises:selecting, by the processor, a particular false positive filter of a plurality of false positive filters based at least in part on the estimated tumor sample purity; andidentifying, by the processor, the initial event as (i) a false positive to be excluded from the mutation list or (ii) a putative mutation for inclusion in the mutation list, using the particular selected false positive filter.- 344 - 13241894vlAttorney Docket. No.: 2013237-097088. The method of any one of claims 58-87, wherein a measured purity of the tumor sample is less than about 0.4 (e.g., wherein the measured purity of the tumor sample is a histological tumor content] and / or wherein the estimated purity of the tumor sample is less than about 0.3.
89. The method of any one of claims 58-88, comprising obtaining, by the processor, normal sequencing data for a normal sample obtained from the subject and using the normal sequencing data and the tumor sequencing data to determine, and filter initial events based on, a likelihood that a given initial event arises from noise.
90. The method of any one of claims 58-89, wherein the tumor sequencing data comprises RNA sequencing data (RNAseq) and the method comprises using the RNAseq data to filter errors.
91. A method comprising:(a) obtaining, by a processor of a computing device, tumor sequencing data for a tumor sample obtained from a subject, the tumor sequencing data representing a tumor genome associated with the tumor sample;(b) detecting, by the processor, using the tumor sequencing data, a plurality of initial events, each initial event corresponding to variant portion of the tumor genome, wherein, for each initial event, the corresponding variant portion is determined, based on the sequencing data, to have a different copy number and / or a different nucleotide sequence relative to a corresponding normal reference;(c) filtering, by the processor, the plurality of initial events to exclude initial events identified as false-positives based on the tumor sequencing data, and retaining other initial events of the plurality of initial events as putative mutations for inclusion in an initial mutation list;(d) determining, by the processor, an initial tumor sample purity estimate and / or a bound thereon;- 345 - 13241894vlAttorney Docket. No.: 2013237-0970(e) filtering, by the processor, the plurality of initial of initial events to exclude putative mutations identified as false-positives based on (i) the tumor sequencing data and (ii) the initial tumor sample purity estimate and / or bound thereon, and retaining other initial events of the plurality of initial events as putative mutations for inclusion in an refined mutation list; and (g) storing and / or providing, for display and / or further processing, the refined mutation list.
92. The method of claim 91, comprising:determining, by the processor, a refined tumor sample purity estimate and / or a bound thereon based on the tumor sequencing data and the refined mutation list;filtering, by the processor, the plurality of initial events to exclude putative mutations identified as false-positives based on (i) the tumor sequencing data and (ii) the refined tumor sample purity estimate and / or bound thereon, and retaining other putative mutations as putative mutations for inclusion in an updated version of the refined mutation list; andstoring and / or providing, for display and / or further processing, the updated version of the refined mutation list.
93. The method of claim 91 or 92, comprising, iteratively:determining, by the processor, a refined tumor sample purity estimate and / or a bound thereon based on the tumor sequencing data and the refined mutation list; andfiltering, by the processor, the plurality of initial events to exclude putative mutations identified as false-positives based on (i) the tumor sequencing data and (ii) the refined tumor sample purity estimate and / or bound thereon, and retaining other putative mutations as putative mutations for inclusion in an updated version of the refined mutation list; anditeratively (i) updating the refined tumor sample purity estimate and / or bound thereon using the updated version of the refined mutation list and (ii) filtering the plurality of initial events using the updated tumor sample purity estimate and / or bound thereon to further update the refined mutation list.- 346 - 13241894vlAttorney Docket. No.: 2013237-097094. A method comprising:(a) obtaining, by a processor of a computing device, tumor sequencing data for a tumor sample obtained from a subject, the tumor sequencing data representing a tumor genome associated with the tumor sample;(b) detecting, by the processor, using on the tumor sequencing data, a plurality of initial events, each initial event corresponding to variant portion of the tumor genome, wherein, for each initial event, the corresponding variant portion is determined, based on the sequencing data, to have a different copy number and / or a different nucleotide sequence relative to a corresponding normal reference;(c) filtering, by the processor, the plurality of initial events to exclude initial events identified as false-positives based on the tumor sequencing data, and retaining other initial events as putative mutations for inclusion in an initial mutation list;(d) determining, by the processor, an initial set of tumor model parameters based on the tumor sequencing data;(e) filtering, by the processor, the plurality of initial of initial events to exclude putative mutations identified as false-positives based on (i) the tumor sequencing data and (ii) the initial set of tumor model parameters, and retaining other initial events as putative mutations for inclusion in an refined mutation list;(g) storing and / or providing, for display and / or further processing, the refined mutation list.
95. The method of claim 94, comprising:determining, by the processor, a refined set of tumor model parameters based on the tumor sequencing data and the refined mutation list;filtering, by the processor, the plurality of initial events to exclude putative mutations identified as false-positives based on (i) the tumor sequencing data and (ii) the refined set of- 347 - 13241894vlAttorney Docket. No.: 2013237-0970tumor model parameters, and retaining other putative mutations as putative mutations for inclusion in an updated version of the refined mutation list; andstoring and / or providing, for display and / or further processing, the updated version of the refined mutation list.
96. The method of claim 94 or 95, comprising, iteratively:determining, by the processor, a refined set of tumor model parameters based on the tumor sequencing data and the refined mutation list; andfiltering, by the processor, the plurality of initial events to exclude putative mutations identified as false-positives based on (i) the tumor sequencing data and (ii) the refined set of tumor model parameters, and retaining other putative mutations as putative mutations for inclusion in an updated version of the refined mutation list; anditeratively (i) updating the refined set of tumor model parameters using the updated version of the refined mutation list and (ii) filtering the plurality of initial events using the updated version of the refined set of tumor model parameters to further update the refined mutation list.
97. A system comprising:a processor of a computing device; andmemory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to:(a) obtain tumor sequencing data for a tumor sample obtained from a subject, the tumor sequencing data representing a tumor genome associated with the tumor sample;(b) detect, using on the tumor sequencing data, a plurality of initial events, each initial event corresponding to a variant portion of the tumor genome, wherein, for each initial event, the corresponding variant portion is determined, based on the- 348 - 13241894vlAttorney Docket. No.: 2013237-0970sequencing data, to have a different copy number and / or a different nucleotide sequence relative to a corresponding normal reference;(c) determine an estimated purity for the tumor sample and / or a bound thereon;(d) filter the plurality of initial events to exclude initial events identified as false-positives based on (i) the tumor sequencing data and (ii) the estimated purity for the tumor sample and / or bound thereon and retaining other initial events as putative mutations for inclusion in a mutation list; and(e) store and / or provide, for display and / or further processing, the mutation list.
98. A system comprising:a processor of a computing device; andmemory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to:(a) obtain tumor sequencing data for a tumor sample obtained from a subject, the tumor sequencing data representing a tumor genome associated with the tumor sample;(b) detecting, by the processor, using the tumor sequencing data, a plurality of initial events;(c) iteratively: filter the plurality of initial events based on a set of one or more tumor model parameters, to generate a filtered list of putative mutations and use the filtered list of putative mutations to refine the set of tumor model parameters,wherein the refined set of tumor model parameters determined at one iteration is used to filter the plurality of initial events in a subsequent iteration, and,- 349 - 13241894vlAttorney Docket. No.: 2013237-0970at a final iteration, the plurality of initial events is filtered based on a final refined set of tumor model parameters determined at a previous iteration, thereby generating a final list of putative mutations; and(d) store and / or provide, the final list of putative mutations for display and / or further processing.
99. A system comprising:a processor of a computing device; andmemory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to:(a) obtain tumor sequencing data for a tumor sample obtained from a subject, the tumor sequencing data representing a tumor genome associated with the tumor sample;(b) detect, using the tumor sequencing data, a plurality of initial events, each initial event corresponding to variant portion of the tumor genome, wherein, for each initial event, the corresponding variant portion is determined, based on the sequencing data, to have a different copy number and / or a different nucleotide sequence relative to a corresponding normal reference;(c) filter the plurality of initial events to exclude initial events identified as false-positives based on the tumor sequencing data, and retaining other initial events of the plurality of initial events as putative mutations for inclusion in an initial mutation list;(d) determine an initial tumor sample purity estimate and / or a bound thereon; (e) filter the plurality of initial of initial events to exclude putative mutations identified as false-positives based on (i) the tumor sequencing data and (ii) the initial tumor sample purity estimate and / or bound thereon, and retaining other initial events of the plurality of initial events as putative mutations for inclusion in an refined mutation list; and- 350 - 13241894vlAttorney Docket. No.: 2013237-0970(g) store and / or provide, for display and / or further processing, the refined mutation list.
100. A system comprising:a processor of a computing device; andmemory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to:(a) obtain tumor sequencing data for a tumor sample obtained from a subject, the tumor sequencing data representing a tumor genome associated with the tumor sample;(b) detect, using on the tumor sequencing data, a plurality of initial events, each initial event corresponding to variant portion of the tumor genome, wherein, for each initial event, the corresponding variant portion is determined, based on the sequencing data, to have a different copy number and / or a different nucleotide sequence relative to a corresponding normal reference;(c) filter the plurality of initial events to exclude initial events identified as false-positives based on the tumor sequencing data, and retaining other initial events as putative mutations for inclusion in an initial mutation list;(d) determine an initial set of tumor model parameters based on the tumor sequencing data;(e) filter the plurality of initial of initial events to exclude putative mutations identified as false-positives based on (i) the tumor sequencing data and (ii) the initial set of tumor model parameters, and retaining other initial events as putative mutations for inclusion in an refined mutation list;(g) store and / or provide, for display and / or further processing, the refined mutation list.- 351 - 13241894vlAttorney Docket. No.: 2013237-0970101. A method of producing an immunotherapy construct for a subject, the method comprising:detecting a plurality of candidate mutations in tumor cells of from the subject using a method or system of any one of claims 1-100; andsynthesizing, as the immunotherapy construct, a polyribonucleotide encoding a plurality of neoepitopes, each encoded neoepitope corresponding to a candidate mutation of the plurality of candidate mutations.
102. A method comprising:determining a subset of T-cells within a population of T-cells that are capable of specifically binding a plurality of complexes,wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, andwherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer- specific mutations detected using the method or system of any one of claims 1-100; andenriching for the subset of T-cells that are capable of specifically binding the plurality of complexes.
103. A method comprising:administering to a subject a population of T-cells, wherein the population of T cells are capable of specifically binding a plurality of complexes,wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells,- 352 - 13241894vlAttorney Docket. No.: 2013237-0970wherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer- specific mutations detected using the method or system of any one of claims 1-100, andwherein the genome of at least some of the subject’s cells comprises a subset of the cancer-specific mutations.
104. A method comprising:determining a subset of TILs within a population of TILs that are capable of specifically binding a plurality of complexes,wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, andwherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer- specific mutations detected using the method or system of any one of claims 1-100; andenriching for the subset of TILs that are capable of specifically binding the plurality of complexes.
105. A method comprising:administering to a subject a population of TILs, wherein the population of TILs are capable of specifically binding a plurality of complexes,wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells,- 353 - 13241894vlAttorney Docket. No.: 2013237-0970wherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer- specific mutations detected using the method or system of any one of claims 1-100, andwherein the genome of at least some of the subject’s cells comprises a subset of the cancer-specific mutations.
106. The method of any one of claims 104 to 105, comprising obtaining a tumor sample from the subject and using the tumor sample to detect the plurality of candidate mutations.
107. The method of claim 106, comprising obtaining a normal sample from the subject and using the normal sample to detect the plurality of cancer mutations.
108. The method of claim 106 or 107, comprising sequencing the tumor and / or normal sample.
109. A pharmaceutical composition comprising a polyribonucleotide encoding a plurality of neoepitopes, wherein at least a portion of the plurality of neoepitopes are individualized neoepitopes corresponding to a somatic mutation from a population of candidate mutations detected in a tumor sample of a subject and selected for inclusion in the pharmaceutical composition using the method or system of any one of the preceding claims.
110. An individualized pharmaceutical composition comprising a polyribonucleotide encoding a plurality of neoepitopes for a particular subject, wherein at least a portion of the neoepitopes are individualized neoepitopes corresponding to a somatic mutation from a population of candidate mutations detected in a tumor sample from the particular subject and selected for inclusion in the pharmaceutical composition using the method or system of any one of the claims 1-100.- 354 - 13241894vlAttorney Docket. No.: 2013237-0970111. A population of T-cells,wherein the population of T-cells is capable of specifically binding to a plurality of complexes,wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, andwherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer- specific mutations using the method or system of any one of claims 1-100.
112. A T-cell capable of specifically binding to a complex, wherein the complex comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancerspecific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells,wherein the one or more cancer-specific mutations are detected using the method or system of any one of claims 1-100.
113. A population of T-cells,wherein the population of T-cells is capable of specifically binding to a plurality of complexes,wherein each complex of the plurality comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells, andwherein the plurality of complexes comprises a plurality of neoepitope portions encoded by cancer- specific mutations detected using the method or system of any one of claims 1-100.- 355 - 13241894vlAttorney Docket. No.: 2013237-0970114. A T-cell capable of specifically binding to a complex, wherein the complex comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancerspecific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells,wherein the one or more cancer-specific mutations are within the selected subset of candidate mutations detected using the method or system of any one of claims 1-100.
115. A T-cell receptor capable of specifically binding to a complex, wherein the complex comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells,wherein the one or more cancer-specific mutations are within the selected subset of candidate mutations detected using the method or system of any one of claims 1-100.
116. A chimeric antigen receptor (CAR) capable of specifically binding to a complex, wherein the complex comprises (i) a portion of a neoepitope encoded by a nucleotide sequence comprising one or more cancer- specific mutations, and (ii) a major histocompatibility complex (MHC) protein expressed by cancer cells and / or antigen presenting cells,wherein the one or more cancer-specific mutations are within the selected subset of candidate mutations detected using the method or system of any one of claims 1-100.- 356 - 13241894vl