System and method for evaluating tumor fraction

The described methods allow for precise and efficient determination of tumor fraction levels in samples by analyzing allele fractions across sub-genomic intervals, addressing the limitations of current genomic research translation into clinical practice for cancer analysis.

JP7702360B2Active Publication Date: 2025-07-03FOUNDATION MEDICINE INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2021568292
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-05-20
Filing Date
2020-05-20
Publication Date
2025-07-03
Estimated Expiration
2040-05-20

AI Technical Summary

Technical Problem

Current methods for translating genomic research into clinical practice for cancer analysis are costly, time-consuming, and technically difficult, necessitating the need for novel approaches to accurately assess tumor fraction levels in samples.

Method used

The methods and systems described enable the evaluation of tumor fraction levels by obtaining an accuracy metric from a sample, comparing it to a reference, and utilizing a predetermined relationship between accuracy metrics and tumor fractions, allowing for precise determination of tumor content in samples through allele fraction analysis across multiple sub-genomic intervals.

Benefits of technology

This approach provides rapid and accurate determination of tumor fractions, even at low levels, enabling effective tumor treatment and monitoring, reducing the need for costly and time-consuming traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007702360000100
    Figure 0007702360000100
  • Figure 0007702360000101
    Figure 0007702360000101
  • Figure 0007702360000102
    Figure 0007702360000102
Patent Text Reader

Abstract

Disclosed at least in part herein is a method for identifying a tumor fraction of a sample from a subject. The method may include, for example, obtaining a value for a target variable associated with a subgenomic interval in the sample, identifying a probability index from the target variable, accessing a determined relationship between the stored probability index and the stored tumor fraction, and identifying the tumor fraction of the sample by referring to the probability index and the determined relationship.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit of priority of U.S. Provisional Patent Application No. 62 / 850,474, filed on May 20, 2019, the content of which is incorporated herein by reference in its entirety.

Background Art

[0002] Cancer cells accumulate mutations during the development and progression of cancer. These mutations can be the result of DNA repair, inherent dysfunctions in copying or modification, or exposure to external mutagens. Specific mutations confer a growth advantage to cancer cells and are actively selected in the microenvironment of the tissue where cancer develops. However, translating genomic research into routine clinical practice remains costly, time - consuming, and technically difficult.

[0003] Accordingly, there remains a need for novel approaches, including genomic profiling, for analyzing cancer - related samples.

Summary of the Invention

[0004] The methods and systems described herein enable the evaluation of tumor fraction levels in a sample, a biopsy, or a subject. Typically, the tumor fraction is expressed or measured as the level or ratio of tumor - derived DNA in a sample relative to a reference in the sample, such as non - tumor DNA or total DNA. In the methods described herein, a value of an accuracy metric of the sample is obtained, and the value can be evaluated with respect to a reference, for example, by comparison to the reference. The accuracy metric can itself be a function of a target variable that reflects the level of alleles in a sub - genomic interval. The target variable can include a variable that is a function of the allele fraction, as well as a variable that is a function of the reads of a sub - genomic interval.

[0005] In some embodiments, the value of the target variable is obtained from, e.g., directly obtained from, the sample. Typically, the reference against which the accuracy index of the sample is compared is a relevant accuracy index value (or multiple accuracy index values) that correlates with, e.g., the level of the tumor fraction. The accuracy index value incorporated by reference can be based on, e.g., an entity or relationship within the sample (e.g., 0.5 for alleles in a heterozygous subgenomic interval) or outside the sample (e.g., a standard curve generated from one or more other subjects).

[0006] In some examples, the target variable can be the allele fraction in one or more subgenomic intervals. Other examples of target variables include variables such as log2 ratios, which are a function of the number of reads in one or more subgenomic intervals. Typically, multiple subgenomic intervals (e.g., 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, or more subgenomic intervals) are analyzed to identify the tumor fraction. The multiple subgenomic intervals can be present on the same chromosome or on different chromosomes (e.g., distributed across 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, or more chromosomes). In one embodiment, at least a portion of the multiple subgenomic intervals is heterozygous (with respect to the alleles in the subgenomic interval).

[0007] In one embodiment, the accuracy index for a sample from a subject is compared to a curve related to the accuracy index for the tumor fraction, and a value for the sample tumor fraction is obtained.

[0008] In one embodiment, the accuracy index is a function of the target variable, e.g., the allele fraction. As an example, the accuracy index can be related to the degree to which the observed allele fraction deviates from a reference, e.g., an expected allele fraction or log2 ratio, and can be compared to a reference related to the level of the tumor fraction. In other examples, the accuracy index can measure the relative accuracy of the target variable, e.g., the entropy index described herein.

[0009] Thus, the methods described herein include methods for evaluating, e.g., estimating, the tumor fraction of a sample. Such methods include, for example, obtaining a value of a target variable of the sample, obtaining a reference value, e.g., a measure of certainty as a function of the target variable, comparing the value of the sample to a reference value to obtain a value of the tumor fraction of the sample.

[0010] In some embodiments, a method for determining the tumor fraction of a sample from a subject comprises obtaining a plurality of values, each value indicative of an allele fraction at a corresponding locus within a subgenomic interval in the sample, identifying a measure of certainty indicative of the variance of the plurality of values, accessing a predetermined relationship between one or more stored measures of certainty and one or more stored tumor fractions, and determining the tumor fraction of the sample from the measure of certainty and the predetermined relationship.

[0011] In some embodiments, each value among the plurality of values is an allele fraction. In some embodiments, each value among the plurality of values comprises a ratio of the abundance difference between the maternal and paternal alleles to the abundance of either the maternal or paternal allele at the corresponding locus. In some embodiments, the measure of certainty indicates the deviation of each of the plurality of values from an expected value. In some embodiments, the expected value is a locus-specific expected value.

[0012] In some embodiments, the measure of certainty is the root mean square deviation from an expected value. In some embodiments, the expected value is the expected allele frequency for non-tumor. In some embodiments, each value among the plurality of values is an allele fraction and the expected value is 0.5.

[0013] In some embodiments, each value among a plurality of values is a ratio of the difference in abundance between a maternal allele and a paternal allele to the abundance of the maternal allele or the paternal allele at a corresponding locus, the expected value includes the expected ratio of the difference in abundance between the maternal allele and the paternal allele to the abundance of the maternal allele or the paternal allele, and the expected value is the expected ratio for a non-tumor sample. In some embodiments, the expected value is 0.

[0014] In some embodiments, the plurality of values includes a plurality of allele coverages.

[0015] In some embodiments, the method further includes identifying a probability distribution function of the plurality of values, and the confidence metric is identified using the probability distribution function. In some embodiments, the confidence metric is the entropy of the probability distribution function.

[0016] In some embodiments, the corresponding locus includes one or more loci having different maternal and paternal alleles. In some embodiments, the corresponding locus consists of loci having different maternal and paternal alleles. In some embodiments, the corresponding locus includes one or more loci having the same maternal and paternal alleles.

[0017] In some aspects, a method for identifying the tumor fraction of a sample from a subject includes obtaining a plurality of values, each value indicating a difference between the allele coverage of a locus in a tumor sample and the allele coverage of the same locus in a non-tumor sample at a plurality of loci within a subgenomic interval; identifying a confidence metric indicating the variance of the plurality of values; accessing a predetermined relationship between one or more stored confidence metrics and one or more stored tumor fractions; and identifying the tumor fraction of the sample from the confidence metric and the predetermined relationship.

[0018] In some embodiments, each value among the plurality of values comprises a ratio of the allelic coverage of a locus in a tumor sample compared to the allelic coverage of the same locus in a non-tumor sample.

[0019] In some embodiments, each value among the plurality of values comprises a log ratio of the allelic coverage of a locus in a tumor sample compared to the allelic coverage of the same locus in a non-tumor sample. In some embodiments, the log ratio is a log2 ratio.

[0020] In some embodiments, each value among the plurality of values comprises a ratio of the difference between the allelic coverage of a locus in a tumor sample and the allelic coverage of the same locus in a non-tumor sample to the allelic coverage of the same locus in the non-tumor sample.

[0021] In some embodiments, the confidence metric indicates the deviation of each value among the plurality of values from an expected value across the corresponding locus, where the expected value is the value that would be expected if the tumor sample were a non-tumor sample.

[0022] In some embodiments, each value comprises a ratio of the allelic coverage of a locus in a tumor sample compared to the allelic coverage of the same locus in a non-tumor sample and the expected value is 1, or each value comprises a log ratio of the allelic coverage of a locus in a tumor sample compared to the allelic coverage of the same locus in a non-tumor sample and the expected value is 0, or each value comprises a ratio of the difference between the allelic coverage of a locus in a tumor sample and the allelic coverage of the same locus in a non-tumor sample to the allelic coverage of the same locus in the non-tumor sample and the expected value is 0.

[0023] In some embodiments, the confidence metric is the root mean square deviation from the expected value.

[0024] In some embodiments, the method further comprises identifying a probability distribution function of the plurality of values, and the confidence metric is determined using the probability distribution function. In some embodiments, the confidence metric is the entropy of the probability distribution function.

[0025] In some embodiments, the allelic coverage includes the allelic coverage of maternal alleles and paternal alleles.

[0026] In some embodiments, the allelic coverage consists of the allelic coverage of maternal alleles and paternal alleles.

[0027] In some embodiments of the above method, the plurality of loci includes at least one nucleotide associated with a single nucleotide polymorphism (SNP). In some aspects, the plurality of loci each include two or more nucleotides each associated with a single nucleotide polymorphism (SNP). In some embodiments, the SNP is associated with cancer.

[0028] In some embodiments of the above method, at least a portion of the plurality of loci is associated with a copy number variation (CNV). In some embodiments, the CNV is associated with cancer.

[0029] In some embodiments of the above method, the method further includes sequencing a sample to identify the abundance or coverage of alleles at each locus.

[0030] In some embodiments of the above method, the method further includes performing array hybridization on the sample to identify the abundance or coverage of alleles at each locus.

[0031] In some embodiments of the above method, the method further includes accessing a training dataset that includes a plurality of relationships between a plurality of training confidence metrics and a training tumor fraction, and applying a machine learning process to the training dataset to identify a predetermined relationship between a training accuracy metric and the training tumor fraction.

[0032] In some embodiments of the above method, the method further includes generating a report that includes the subject and information identifying the specified tumor fraction. In some embodiments, the method further includes providing the report to the subject or a healthcare provider. In some embodiments, the method further includes formatting the report for an electronic health record.

[0033] In some aspects, a method of treating a subject's tumor includes administering an effective amount of tumor therapy to the subject in response to the specified tumor fraction, the tumor fraction being specified according to any one of the above methods. In some aspects, the method includes identifying the presence of a tumor in a patient based on the specified tumor fraction. In some aspects, the tumor therapy includes chemotherapy, radiation therapy, or surgery.

[0034] In some aspects, a method of monitoring tumor progression or recurrence in a subject includes: (a) identifying a first tumor fraction of a first sample obtained from the subject at a first time point according to any one of the above methods; (b) identifying a second tumor fraction of a second sample obtained from the subject at a second time point; and (c) comparing the first tumor fraction with the second tumor fraction to thereby monitor tumor progression.

[0035] In some embodiments of the method of monitoring tumor progression or recurrence, identifying the second tumor fraction comprises obtaining a plurality of second values, each value indicating an allele fraction at a corresponding locus within a subgenomic interval in the second tumor sample, the subgenomic interval in the second sample being the same as or different from the subgenomic interval in the first sample; identifying a second accuracy metric indicative of the variance of the plurality of second values; accessing a predetermined relationship between one or more stored accuracy metrics and one or more stored tumor fractions; and identifying the second tumor fraction of the second sample from the second accuracy metric and the predetermined relationship.

[0036] In some embodiments of methods for monitoring tumor progression or recurrence, identifying a second tumor fraction comprises obtaining a second plurality of values, each value indicative of the difference between the allele coverage of loci in a second tumor sample within a subgenomic interval in a sample and the allele coverage of the same loci in a non-tumor sample, wherein the subgenomic interval used to identify the second tumor fraction is the same as or different from the subgenomic interval used to identify the first tumor fraction, obtaining, identifying a second accuracy metric indicative of the variance of the second plurality of values, accessing a predetermined relationship between one or more stored accuracy metrics and one or more stored tumor fractions, and identifying a second tumor fraction of the second tumor sample from the second accuracy metric and the predetermined relationship.

[0037] In some aspects of methods for monitoring tumor progression or recurrence, the method further comprises adjusting tumor therapy in response to tumor progression. In some aspects, the method comprises adjusting the dosage of a tumor therapy or selecting a different tumor therapy in response to tumor progression. In some aspects, the method comprises administering the adjusted tumor therapy to the subject.

[0038] In some aspects of methods for monitoring tumor progression or recurrence, the method comprises the first time point being before the subject is administered tumor therapy and the second time point being after the subject is administered tumor therapy.

[0039] In some embodiments of any of the methods described above, the subject has cancer, is at risk of having cancer, or is suspected of having cancer. In some embodiments, the cancer is a solid tumor. In some embodiments, the cancer is a blood cancer.

[0040] In some embodiments of any of the methods described above, the sample is a liquid sample.

[0041] In some embodiments of any of the methods described above, the sample is a solid sample.

[0042] In some embodiments of any of the above methods, the sample comprises cell-free DNA (cfDNA) or circulating tumor DNA (ctDNA).

[0043] In some embodiments of any of the above methods, one or more stored accuracy metrics comprise a plurality of stored accuracy metrics, and one or more stored tumor fractions comprise a plurality of stored tumor fractions.

[0044] Disclosed herein is a computer system comprising a processor and a memory communicatively coupled to the processor and configured to store a predetermined relationship between one or more stored accuracy metrics and one or more associated stored tumor fractions, the memory storing instructions that, when executed by the processor, cause the processor to: (a) (i) obtain a plurality of values indicative of allele fractions at corresponding loci within a subgenomic interval in a sample, or (ii) obtain a plurality of values indicative of the difference between allele coverage of loci in a tumor sample and allele coverage of the same loci in a non-tumor sample at a plurality of loci within a subgenomic interval; (b) identify an accuracy metric indicative of the variance of the plurality of values; (c) access the stored predetermined relationship; (d) identify the tumor fraction of the sample from the accuracy metric and the predetermined relationship.

[0045] In some embodiments of the computer system, the memory further stores instructions that, when executed by the processor, cause the processor to access a training dataset comprising a plurality of relationships between a plurality of training accuracy metrics and associated training tumor fractions and apply a machine learning process to the training dataset to identify a predetermined relationship between the training accuracy metrics and the training tumor fractions.

[0046] In some embodiments of the computer system, the instructions, when executed by the processor, cause the processor to execute any one of the above methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Various aspects of at least one example are described below with reference to the accompanying drawings which are not drawn to scale. The drawings are included to provide an illustration and further understanding of the various aspects and examples and are incorporated herein and constitute a part thereof, but are not intended as a definition of the limits of a particular example. The drawings, together with the remainder of the specification, serve to explain the principles and operation of the described and claimed aspects and examples. In the figures, each component that is the same or substantially the same as shown in the various figures is represented by like reference numerals. For clarity, not all components are labeled in all of the figures.

[0048]

Figure 1

[0049]

Figure 2

[0050]

Figure 3

[0051]

Figure 4

[0052] Methods and systems for determining the tumor fraction of a sample from a subject are described herein. Also described are methods of treating a subject's tumor in response to the determined tumor fraction, and methods and systems for monitoring the progression or recurrence of a subject's tumor, including determining the tumor fraction in samples obtained from the subject at two or more time points. Rapid and accurate determination of the tumor fraction can substantially enhance tumor therapy by ensuring that a subject receives effective treatment during the initial stages of the tumor or during tumor recurrence, particularly at low tumor fraction levels. Other uses of the tumor fraction are also contemplated and further discussed herein. For example, the tumor fraction can be used in some embodiments to analyze tumor biopsies. In some aspects, the tumor fraction is used to characterize variants (e.g., as somatic or germline, or as homozygous, heterozygous or subclonal) using, for example, a somatic germline zygosity (SGZ) algorithm. The methods and systems described herein provide accurate tumor fraction determination even at low tumor fraction levels.

[0053] As further described herein, the tumor fraction is closely related to the variance of allele fractions across multiple analyzed loci. The variance can be referred to as an "accuracy metric". Using the relationship between one or more accuracy metrics and one or more corresponding tumor fractions, the tumor fraction of a sample can be determined from the determined accuracy metric of the sample from the subject. The relationship receives the determined accuracy metric as input and outputs the tumor fraction of the sample. This relationship can be applied to determine the tumor fraction of a sample from a subject, thereby enabling effective tumor treatment, monitoring of the subject for tumor progression or recurrence, and / or analysis of tumor samples.

[0054] In some embodiments, the tumor fraction of a sample is determined for a tumor sample using a tumor sample and a non-tumor sample (e.g., a healthy tissue sample). The tumor sample and the non-tumor sample can be obtained from the same individual (i.e., a matched normal control) or from different individuals. The accuracy metric can be the variance of a plurality of values, each of which indicates the difference between the coverage of a locus in the tumor sample at a plurality of loci and the coverage of the same locus in the non-tumor sample. As described above, the relationship between the accuracy metric and the tumor fraction can be used to determine the tumor fraction of a sample from the determined accuracy metric of the sample from a subject. The relationship receives the determined accuracy metric as an input and outputs the tumor fraction of the sample. This relationship can be applied to determine the tumor fraction of a sample from a subject, thereby enabling effective tumor treatment, monitoring of the subject for tumor progression or recurrence, and / or analysis of the tumor sample.

[0055] Tumor fraction determination An important metric in cancer monitoring, diagnosis, and treatment is the tumor fraction. In some embodiments, the tumor fraction is a measure of the tumor genomic content in a sample (e.g., a biopsy), proportional to the total genomic content regardless of cell origin. Generally, it is advantageous to identify (e.g., estimate) tumor content or changes in tumor content from a sample, as this can serve both to report changes and to provide information regarding the presence or progression of a disease. For example, a liquid biopsy that typically utilizes a blood sample from a cancer patient can be useful when a solid biopsy is not possible or not recommended. The methods described herein can be used to determine the tumor fraction in various types of samples, such as solid and liquid samples. In some embodiments, the methods described herein are used for solid samples, for example, as an alternative to or in combination with visual screening methods. In other embodiments, the methods described herein are used for liquid samples, for example, when visual screening methods are ineffective or not available.

[0056] In some embodiments, the tumor fraction in a cell-free sample includes a measure of tumor DNA released from a primary tumor into the vasculature or lymphatics, carried around the body in the bloodstream, relative to the amount of total DNA (e.g., tumor and normal) released into the bloodstream. The tumor fraction can be used to monitor patients at risk of cancer (regardless of the presence or absence of a current diagnosis). As a factor used in the diagnosis of cancer; or to determine whether a current treatment regimen is effective, e.g., has a beneficial effect.

[0057] Traditional approaches for measuring tumor fraction typically require that modeled parameters of both purity and ploidy be inferred from either or both of log ratios and allele frequencies, or from a pathological review. In some embodiments, the tumor fraction can be considered a modeled parameter of fragments of cancer cells in a heterogeneous tumor sample, and tumor purity or other measurements can be taken into account. In some embodiments, tumor cell ploidy can refer to the average weighted copy number of all chromosomes (or portions thereof). The ploidy observed in a sample can be affected by various degrees of aneuploidy of the tumor cells, the heterogeneity of the sample (e.g., different ratios of tumor cells to normal cells), or both.

[0058] Traditional approaches for predicting tumor fraction can be highly unreliable for low tumor content due to poorly fitting models. In some embodiments, the methods described herein can overcome certain drawbacks of conventional efforts by, for example, identifying the tumor fraction (and associated confidence levels) based on the effects of tumor cell aneuploidy, e.g., as measured by allele coverage or allele fraction in one or more sub-genomic intervals in a sample. In some embodiments, the sub-genomic intervals include heterozygous single nucleotide polymorphism (SNP) sites. In other embodiments, the sub-genomic intervals include two or more nucleotide positions.

[0059] As used herein, the term "allele coverage" or simply "coverage" or "Cvg" refers to the number of reads (e.g., unique reads) generated from DNA sequence identification of a subgenomic interval in a sample. As used herein, the term "allele intensity" or simply "intensity" refers to the number of signals (e.g., unique signals) generated from genomic hybridization in a subgenomic interval in a sample. "Read" or "signal" is intended to encompass situations where there may be duplicates of the same "unique read" or "unique signal" (i.e., duplicates are not removed prior to performing the methods described herein), but since duplicates are represented in both the numerator and denominator, it will be understood that any ratio calculated using the described methods will result in a value very similar to the "unique" read or signal ratio.

[0060] As used herein, the term "allele fraction" refers to the relative level (e.g., abundance) of an allele in a subgenomic interval in a sample. The allele fraction can be expressed as a ratio or percentage. For example, the allele fraction can be expressed as the ratio of the number of one particular allele (e.g., A, T, C, or G) in a subgenomic interval to the number of all different alleles in that subgenomic interval. In some embodiments, the allele fraction is measured by calculating the ratio of the coverage or intensity from one particular allele (e.g., A, T, C, or G) in a given subgenomic interval to the total coverage or intensity from all different alleles. Sometimes, the terms "allele fraction" and "allele frequency" are used interchangeably herein. As used herein, the log ratio is typically measured by log2(T / R), where T is the level (e.g., abundance) of one or more alleles associated with a subgenomic interval in a sample, and R is the level (e.g., abundance) of one or more alleles associated with the subgenomic interval in a reference sample. As used herein, the term "allele" refers to one of two or more alternative forms (e.g., genes or any portion thereof) of a genomic sequence. For example, if a "C" to "T" SNP is associated with a subgenomic interval, the subgenomic interval can be described as being associated with alleles "C" and "T" with respect to the SNP.

[0061] In some embodiments, there are two or more different alleles associated with a subgenomic interval. If two or more different alleles are present in a sample, the subgenomic interval is considered to be heterozygous for the sample. If the subgenomic interval is not heterozygous for the sample, in some embodiments, it can be homozygous, semi-zygous, or hemizygous.

[0062] As used herein, the term "abundance" refers to the amount, number, or quantity of an object. For example, the abundance of an allele associated with a subgenomic interval can mean the amount, number, or quantity of the allele associated with the subgenomic interval in a sample, as determined, for example, by sequence-specific or array-based comparative genomic hybridization (aCGH). For example, if there are two alleles "A" and "G" associated with a particular subgenomic interval and there are 10 copies of allele "A" and 20 copies of allele "G" in the sample, the abundance of allele "A" can be considered 10 and the abundance of allele "G" can be considered 20. In some embodiments, the abundance of an allele is measured by allele coverage or allele intensity. For example, the number of unique reads for allele "A" or "G" reflects how many copies of allele "A" or "G" are present in the sample.

[0063] As used herein, the term "accuracy metric" refers to a metric derived from a measure or value of a target variable. In some embodiments, the target variable can represent the abundance of a subgenomic interval or an allele associated with a subgenomic interval in a sample. In some examples, the accuracy metric can be the deviation of the allele fraction from the expected allele fraction. In other examples, the accuracy metric can be a measure of allele intensity. These examples are illustrative and other accuracy metrics may be used.

[0064] As an example, in the case of a heterozygous SNP, an allele fraction value of 0.50 can indicate a typical diploid sub-genomic interval. An allele fraction that deviates from the expected value of 0.50 indicates aneuploidy at that site. In these examples, this deviation in allele coverage can be correlated with the tumor fraction within the training set in order to build a model that identifies (e.g., predicts or estimates) the tumor fraction based on allele coverage. In some embodiments, the methods described herein correlate the deviation of the allele fraction or log ratio with the tumor fraction, thereby eliminating the need to model the purity and ploidy of the tumor. In some embodiments, the methods described herein enable more accurate identification of low levels, e.g., tumor fractions less than 30%. In one embodiment, the allele fraction or log ratio is identified by a method that includes sequence identification, e.g., next-generation sequencing (NGS). It will be understood that the method for identifying the allele fraction or log ratio is not limited to sequence identification. For example, any method for measuring the coverage or relative level (e.g., abundance) of SNPs, as well as any method for measuring coverage from a larger genomic region, can be used. In one embodiment, the allele fraction or log ratio is identified by a method other than sequence identification, e.g., by array-based comprehensive genomic hybridization (aCGH). In one embodiment, the tumor fraction is, is expected to be, or is predicted to be 0.25 or less, 0.2 or less, 0.15 or less, or 0.1 or less, e.g., between 0.1 and 0.3, between 0.1 and 0.2, between 0.2 and 0.3, or between 0.15 and 0.25.

[0065] In some embodiments, the methods described herein use the allele fraction or log ratio to indicate the expected proportion of coverage, but it will be understood that the present disclosure is generally intended to describe the correlation of the tumor fraction with the deviation of the expected coverage without being limited to the allele fraction, log ratio, or any other specific metric.

[0066] As used herein, "single nucleotide polymorphism" or SNP refers to a change in a single nucleotide that occurs at a specific location in the genome. In some embodiments, such a change exists to a recognizable extent within a population (e.g., >1%). Typically, SNPs are germline changes and not somatic single nucleotide variants (SNVs).

[0067] In one embodiment, the tumor fraction is a numerical representation (e.g., a ratio or percentage) indicating the amount of DNA from tumor cells relative to the total amount of DNA (e.g., tumor and non-tumor DNA) in a sample. In one embodiment, the sample is a liquid biopsy material. In one embodiment, the sample is a solid tissue sample. In one embodiment, the tumor is a solid tumor. In one embodiment, the tumor is a blood cancer. In one embodiment, the tumor fraction in a liquid biopsy indicates the presence or level of a detectable tumor in the body.

[0068] An exemplary method of determining the tumor fraction of a sample from a subject is to obtain a plurality of values, each value indicating an allele fraction at a corresponding locus within a subgenomic interval in the sample, identify an accuracy metric indicating the variance of the plurality of values, access a predetermined relationship between the stored accuracy metric and the stored tumor fraction, and determine the tumor fraction of the sample from the accuracy metric and the predetermined relationship.

[0069] The value indicating the allele fraction can be determined for each corresponding locus. A locus can include one or more nucleotides. In some embodiments, the corresponding locus includes one or more loci having different maternal and paternal alleles. In some embodiments, the corresponding locus consists of loci having different maternal and paternal alleles. In some embodiments, the corresponding locus includes one or more loci having the same maternal and paternal alleles.

[0070] In some embodiments, the plurality of values indicating the allele fractions at a plurality of corresponding loci in a sample are the plurality of allele fractions at the plurality of corresponding loci in the sample. The allele fraction at each of the corresponding loci can be determined, for example, by sequencing nucleic acid molecules in a tumor sample and assigning allele coverage for each allele at each locus. For example, the allele fraction at locus JPEG0007702360000001.jpg9170 can be determined as follows. JPEG0007702360000002.jpg16170Where JPEG0007702360000003.jpg8170 is the coverage of allele a at locus i, JPEG0007702360000004.jpg8170 is the coverage of allele b at locus i. In some embodiments, allele a and allele b are JPEG0007702360000005.jpg11170 are assigned as such.

[0071] In some embodiments, the expected allele fraction is the allele fraction expected in a healthy individual or healthy sample (i.e., a non-tumor sample). For example, the allele fraction at a heterozygous locus (i.e., having different maternal and paternal alleles) is expected to be 0.5, and the allele fraction at a homozygous locus (i.e., the maternal and paternal alleles are the same) is expected to be 1.0.

[0072] The allele fraction is an exemplary value for specifying the tumor fraction according to the methods described herein. However, in some embodiments, other values indicating the allele fraction may be used. In some embodiments, the value indicating the allele fraction is the relative difference in allele frequency. For example, the value indicating the allele fraction can be the ratio of the difference in abundance (e.g., coverage or sequence-specific depth) between the maternal allele and the paternal allele to the abundance of the maternal allele or the paternal allele. That is, in some embodiments, the value is the relative difference as follows. JPEG0007702360000006.jpg16170

[0073] wherein JPEG0007702360000007.jpg10170 is the coverage of allele a at locus i, JPEG0007702360000008.jpg8170 is the coverage of allele b at locus i. In a healthy individual or healthy sample, the difference and relative difference in allele frequency are expected to be 0. In some embodiments, a probability distribution function is specified for a plurality of values indicating the allele fraction. For example, in some embodiments, the probability distribution function is specified for a plurality of allele fractions at a plurality of corresponding loci in a sample. In some embodiments, the probability distribution function of the plurality of allele fractions is defined as follows. JPEG0007702360000009.jpg17170 wherein JPEG0007702360000010.jpg9170 is the coverage of allele a at locus i, JPEG0007702360000011.jpg8170 is the coverage of allele b at locus i.

[0074] The dispersion (or certainty index) can be, for example, a deviation from the expected allele fraction (or a value indicating the expected allele fraction) across multiple loci. In some embodiments, the certainty index is the root mean square deviation from the expected allele fraction (or a value indicating it). For example, in some embodiments, the certainty index is the root mean square deviation (RMSD) defined as follows. JPEG0007702360000012.jpg20170 Wherein, JPEG0007702360000013.jpg8170 is the allele frequency (or a value indicating the allele frequency such as the relative difference ratio) at locus i, JPEG0007702360000014.jpg9170 is the expected allele frequency at locus i, and N is the number of loci at a plurality of corresponding loci. For example, for some loci, JPEG0007702360000015.jpg9170 can be 0.5, and at other loci, JPEG0007702360000016.jpg9170 can be 1. In some embodiments, the loci include only loci having different maternal and paternal alleles. Thus, JPEG0007702360000017.jpg8170 can be defined as 0.5 across all loci, and RMSD can be defined as follows. JPEG0007702360000018.jpg21170

[0075] In some embodiments, the value indicating the allele fraction can be the ratio of the abundance difference (e.g., coverage or sequence-specific depth) between the maternal and paternal alleles to the abundance of the maternal or paternal allele, JPEG0007702360000019.jpg8170 can be defined as 0. Thus, RMSD can be defined as follows. JPEG0007702360000020.jpg21170 Wherein, JPEG0007702360000021.jpg8170 is the coverage of allele a at locus i, JPEG0007702360000022.jpg8170 is the coverage of allele b at locus i.

[0076] In some embodiments, a probability distribution (e.g., a probability distribution function) can be specified for allele fractions across multiple loci. A confidence metric (e.g., a variance) can be an indicator of the probability distribution, such as the entropy of the probability distribution. For example, in some embodiments, the allele fraction probability distribution function The entropy of JPEG0007702360000023.jpg9170 can be defined as follows. JPEG0007702360000024.jpg17170 Where JPEG0007702360000025.jpg9170 is the allele fraction probability distribution function and n is the base of the logarithm. In some embodiments, the base of the logarithm is 2 (i.e., log2). Thus, in some embodiments, the entropy of the allele fraction probability distribution function JPEG0007702360000026.jpg9170 can be defined as follows. JPEG0007702360000027.jpg16170

[0077] In some embodiments, a method for determining the tumor fraction of a sample from a subject comprises obtaining a plurality of values, each value indicating a difference between the allelic coverage of a locus in a tumor sample and the allelic coverage of the same locus in a non-tumor sample at a plurality of loci within a sub-genomic interval; identifying a confidence metric indicative of the variance of the plurality of values; accessing a predetermined relationship between the stored confidence metric and the stored tumor fraction; and determining the tumor fraction of the sample from the confidence metric and the predetermined relationship. In some embodiments, the tumor sample and the non-tumor sample are obtained from the same individual (i.e., a matched normal control). In some embodiments, the tumor sample and the non-tumor sample are obtained from different individuals. The coverage may be raw coverage (e.g., the raw number of sequence-specific reads), normalized coverage (e.g., normalized to the average sequence-specific depth or the median of the sequence-specific depths), and / or other bias-corrected coverage (e.g., GC-bias-corrected coverage depth). In some embodiments, the allelic coverage comprises the coverage of the maternal allele and the coverage of the paternal allele (e.g., the sum of the coverage of the maternal allele and the coverage of the paternal allele). In some embodiments, the allelic coverage consists of the coverage of the maternal allele and the coverage of the paternal allele (such as the sum of the coverage of the maternal allele and the coverage of the paternal allele).

[0078] In some embodiments, each value indicating the difference between the allele coverage of a locus in a tumor sample and the allele coverage of the same locus in a non-tumor sample includes the ratio of the allele coverage of the locus in the tumor sample compared to the allele coverage of the same locus in the non-tumor sample. In some embodiments, the allele coverage includes the coverage of the maternal allele and the coverage of the paternal allele (e.g., the sum of the coverage of the maternal allele and the coverage of the paternal allele). In some embodiments, the allele coverage consists of the coverage of the maternal allele and the coverage of the paternal allele (such as the sum of the coverage of the maternal allele and the coverage of the paternal allele). For example, in some embodiments, the ratio may be defined as follows. JPEG0007702360000028.jpg21170Wherein, JPEG0007702360000029.jpg11170 is the coverage of the maternal allele at locus i in the tumor sample, JPEG0007702360000030.jpg11170 is the coverage of the maternal allele at locus i in the tumor sample, JPEG0007702360000031.jpg11170 is the coverage of the maternal allele at locus i in the non-tumor sample, JPEG0007702360000032.jpg11170 is the coverage of the maternal allele at locus i in the non-tumor sample.

[0079] In some embodiments, each value indicating the difference between the allele coverage of a locus in a tumor sample and the allele coverage of the same locus in a non-tumor sample is the log ratio (such as log2 ratio) of the allele coverage of the locus in the tumor sample compared to the allele coverage of the same locus in the non-tumor sample. In some embodiments, the allele coverage includes the coverage of the maternal allele and the coverage of the paternal allele (e.g., the sum of the coverage of the maternal allele and the coverage of the paternal allele). In some embodiments, the allele coverage consists of the coverage of the maternal allele and the coverage of the paternal allele (such as the sum of the coverage of the maternal allele and the coverage of the paternal allele). For example, the log ratio can be defined as follows in some embodiments. JPEG0007702360000033.jpg21170Where, log n is the logarithm to the base n, JPEG0007702360000034.jpg10170is the coverage of the maternal allele at locus i in the tumor sample, JPEG0007702360000035.jpg10170is the coverage of the maternal allele at locus i in the tumor sample, JPEG0007702360000036.jpg10170is the coverage of the maternal allele at locus i in the non-tumor sample, JPEG0007702360000037.jpg12170is the coverage of the maternal allele at locus i in the non-tumor sample. For example, the log ratio may be a log2 ratio. In some embodiments, the log ratio is defined as follows. JPEG0007702360000038.jpg21170Where, JPEG0007702360000039.jpg12170is the coverage of the maternal allele at locus i in the tumor sample, JPEG0007702360000040.jpg10170is the coverage of the maternal allele at locus i in the tumor sample, JPEG0007702360000041.jpg11170 is the coverage of the maternal allele at locus i in the non-tumor sample, JPEG0007702360000042.jpg11170 is the coverage of the maternal allele at locus i in the non-tumor sample, is the coverage of the maternal allele at locus i in the non-tumor sample.

[0080] In some embodiments, each value indicating the difference between the allele coverage of a locus in a tumor sample and the allele coverage of the same locus in a non-tumor sample includes the ratio of the difference in the allele coverage of the locus in the tumor sample compared to the allele coverage of the same locus in the non-tumor sample, relative to the allele coverage of the same locus in the non-tumor sample. In some embodiments, the allele coverage includes the coverage of the maternal allele and the coverage of the paternal allele (e.g., the sum of the coverage of the maternal allele and the coverage of the paternal allele). In some embodiments, the allele coverage consists of the coverage of the maternal allele and the coverage of the paternal allele (such as the sum of the coverage of the maternal allele and the coverage of the paternal allele). For example, in some embodiments, the ratio is defined as follows. JPEG0007702360000043.jpg21170 Wherein, JPEG0007702360000044.jpg11170 is the coverage of the maternal allele at locus i in the tumor sample, JPEG0007702360000045.jpg11170 is the coverage of the maternal allele at locus i in the tumor sample, JPEG0007702360000046.jpg10170 is the coverage of the maternal allele at locus i in the non-tumor sample, JPEG0007702360000047.jpg11170 is the coverage of the maternal allele at locus i in the non-tumor sample.

[0081] In some embodiments, the probability distribution function is specified for a plurality of values representing the difference between the allele coverage of a locus in a tumor sample and the allele coverage of the same locus in a non-tumor sample. In some embodiments, the allele coverage includes the coverage of the maternal allele and the coverage of the paternal allele (e.g., the sum of the coverage of the maternal allele and the coverage of the paternal allele). In some embodiments, the allele coverage consists of the coverage of the maternal allele and the coverage of the paternal allele (such as the sum of the coverage of the maternal allele and the coverage of the paternal allele). For example, in some embodiments, the probability distribution function is specified for a plurality of ratios of the allele coverage of a locus in a tumor sample compared to the allele coverage of the same locus in a non-tumor sample (e.g., a log ratio, such as a log2 ratio). In some embodiments, the probability distribution function of the plurality of allele fractions is defined as follows. JPEG0007702360000048.jpg23170Where log n is the logarithm to base n, JPEG0007702360000049.jpg10170is the coverage of the maternal allele at locus i in the tumor sample, JPEG0007702360000050.jpg10170is the coverage of the maternal allele at locus i in the tumor sample, JPEG0007702360000051.jpg10170is the coverage of the maternal allele at locus i in the non-tumor sample, JPEG0007702360000052.jpg10170is the coverage of the maternal allele at locus i in the non-tumor sample. In some embodiments, the log ratio is a log2 ratio. For example, in some embodiments, the probability distribution function of the plurality of allele fractions is defined as follows. JPEG0007702360000053.jpg23170Where JPEG0007702360000054.jpg11170 is the coverage of the maternal allele at locus i in the tumor sample, JPEG0007702360000055.jpg12170 is the coverage of the maternal allele at locus i in the tumor sample, JPEG0007702360000056.jpg11170 is the coverage of the maternal allele at locus i in the non-tumor sample, JPEG0007702360000057.jpg11170 is the coverage of the maternal allele at locus i in the non-tumor sample.

[0082] The variance (or accuracy metric) can be, for example, the deviation of each value among a plurality of values from the expected value across the corresponding locus. The expected value is the value that would be expected if the tumor sample were a non-tumor (e.g., healthy) sample. In some embodiments, the accuracy metric is the root mean square deviation from the expected value. For example, in some embodiments, the accuracy metric is the root mean square deviation (RMSD) defined as follows. JPEG0007702360000058.jpg27170

[0083] In some embodiments, the value indicating the allele fraction is the ratio of the difference in the allele coverage of the locus in the tumor sample compared to the allele coverage of the same locus in the non-tumor sample, to the allele coverage of the same locus in the non-tumor sample. Thus, the RMSD can be defined as follows. JPEG0007702360000059.jpg21170

[0084] In some embodiments, a probability distribution (e.g., a probability distribution function) can be specified for a plurality of values indicating the difference between the allele coverage of the locus in the tumor sample and the allele coverage of the same locus in the non-tumor sample. The accuracy metric (e.g., the variance) can be an indicator of the probability distribution such as the entropy of the probability distribution. For example, in some embodiments, the allele fraction probability distribution function The entropy of JPEG0007702360000060.jpg9170 can be defined as follows. JPEG0007702360000061.jpg16170 where JPEG0007702360000062.jpg25170 where log n is the logarithm with base n, and JPEG0007702360000063.jpg11170 is the coverage of the maternal allele at locus i in the tumor sample, JPEG0007702360000064.jpg12170 is the coverage of the maternal allele at locus i in the tumor sample, JPEG0007702360000065.jpg11170 is the coverage of the maternal allele at locus i in the non - tumor sample, JPEG0007702360000066.jpg11170 is the coverage of the maternal allele at locus i in the non - tumor sample. In some embodiments, the base of the logarithm is 2 (i.e., log2). Thus, in some embodiments, the allele fraction probability distribution function JPEG0007702360000067.jpg8170's entropy can be defined as follows. JPEG0007702360000068.jpg15170 where JPEG0007702360000069.jpg21170 where JPEG0007702360000070.jpg10170 is the coverage of the maternal allele at locus i in the tumor sample, JPEG0007702360000071.jpg11170 is the coverage of the maternal allele at locus i in the tumor sample, JPEG0007702360000072.jpg12170 is the coverage of the maternal allele at locus i in the non - tumor sample, JPEG0007702360000073.jpg11170 is the coverage of the maternal allele at locus i within the non-tumor sample.

[0085] Using the relationship between one or more stored accuracy metrics and one or more stored tumor fractions, a tumor fraction can be identified based on the identified accuracy metric. In some embodiments, the model is trained to use a training dataset that includes training accuracy metrics and associated tumor fractions to identify the relationship between the accuracy metric and the tumor fraction. The training dataset can be identified, for example, using a plurality of clinical samples having known (i.e., training) tumor fractions (e.g., as identified by maximum somatic allele frequency (MSAF), where the tumor fraction filters germline variant calls from all calls in the tumor sample and compares the remaining variants (i.e., maximum somatic variants) to the total variants (maximum somatic variants + germline variants) to identify the maximum somatic allele frequency). Nucleic acid molecules in the clinical samples can be sequenced to identify allele frequencies (or values indicative of allele frequencies) across a plurality of loci, as well as associated training accuracy metrics. The training accuracy metrics can be correlated with the training tumor fractions to identify the relationship between the accuracy metric and the tumor fraction. In another approach, serial dilutions can be performed from one or more clinical samples to obtain a plurality of different tumor fractions, which can be correlated with the accuracy metrics of the serially diluted samples to identify the relationship.

[0086] In some aspects, a training sub-process is first performed to identify (e.g., estimate) the tumor fraction. The dataset can be constructed from clinical specimens. Using the training set and in-silico dilutions of the training set, the tumor fraction can be correlated with the variation in allele fraction or log ratio corresponding to the aneuploidy typically observed in tumors. In other examples, cell line / clinical sample dilutions can be performed.

[0087] In some embodiments, the confidence metric can be a function of coverage in a particular SNP bin for a particular allele and / or allele frequency (e.g., in the range from 0 to 0.5). In some examples, the training data uses a deviation metric (e.g., allelic fraction deviation or log ratio deviation) as an input and returns an estimated tumor fraction along with a lower and upper bound. Values that deviate from 0 and 1 (i.e., fall in between) and are not 0.5 (exclusive) can be considered "noise", and the averaged noise can be correlated with the expected or estimated tumor fraction. In other examples, the training data provides as an input a log ratio deviation metric, or generally any metric that quantifies the coverage deviation from an expected value. In either case, the allelic coverage deviation metric or the log ratio deviation metric can be a measure of the tumor fraction.

[0088] These correlations derived during training can be utilized to estimate or evaluate the patient's tumor fraction with upper and lower bounds. Coverage metrics such as SNP allelic coverage variation metrics can be used in generating the correlations.

[0089] The methods described herein can, for example, improve the ability to identify whether a tumor is present in a biological sample and provide a tumor fraction determination (e.g., an estimate) with known estimation limits. Provide a systematic and orthogonal approach for evaluating somatic variants; Provide a framework for a new, inexpensive tumor tracking / identification assay.

[0090] In some embodiments, the methods described herein also provide advantages in certain cases of liquid biopsy (although the present disclosure is not limited to liquid biopsy). Solid tumors have multiple different means for estimating tumor content, including pathological review, somatic allele frequency (MSAF), and analytical copy number alteration (CNA) modeling. However, liquid biopsy is typically not suitable for these methods or requires significant rework. Since cell-free DNA floats freely in the blood, its presence is microscopic and thus cannot be examined by a pathologist. Furthermore, the amount of DNA that a tumor tends to release into the bloodstream can be small compared to normal DNA. Thus, analytical CNA modeling can fail due to low tumor content.

[0091] The methods described herein typically do not require a pathological review. They are sensitive enough so that analytical CNA modeling is not required to identify the presence or content of a tumor; there is no analytical equation; they are independent of short variant calling and provide orthogonal assessment of short variants. And they are improved (e.g., not confounded) in the presence of CNA events.

[0092] The methods described herein enable the development of new, inexpensive tumor tracking (e.g., monitoring) assays. For example, if a patient presents tumor content in an assay (e.g., a comprehensive assay) that covers a sufficient number of sub-genomic intervals (e.g., sub-genomic intervals that include one or more SNPs), this method can be based only on SNP mutations, so that tumor progression can be tracked over time with a second assay and at a fairly low cost. In some embodiments, the first assay encompasses more sub-genomic intervals than the second assay. In other embodiments, the first assay covers fewer sub-genomic intervals than the second assay. In certain specific embodiments, the first assay and the second assay cover essentially the same number of sub-genomic intervals.

[0093] The gene panels included in the first and second assays can have the same or different sizes. For example, an assay including a panel of at least about 100, 150, 200, 250, 300, 350, 400, 450, 500 or more genes can be considered a large panel, and an assay including about 100, 90, 80, 70, 60, 50, 40, 30, 20 or less than 10 genes can be considered a small panel. The "large" and "small" panel sizes are typically specified by the purpose of the assay and should not be limited to the exemplary sizes above. In some embodiments, the first assay includes a large panel and the second assay includes the same or a different large panel. In other embodiments, the first assay includes a small panel and the second assay includes the same or a different small panel. In certain embodiments, the first assay includes a large panel and the second assay includes a small panel, or vice versa. The first and second assays need not be of the same assay type. For example, the first assay can be based on sequence identification (e.g., NGS), the second assay can be based on genomic hybridization, or vice versa.

[0094] In some embodiments, the sub-genomic intervals covered by the second assay can be a subset of the sub-genomic intervals covered by the first assay. In some embodiments, the sub-genomic intervals covered by the first assay can be a subset of the sub-genomic intervals covered by the second assay. In other embodiments, the sub-genomic intervals covered by the second assay overlap, but are not the same as, the sub-genomic intervals covered by the first assay. In certain embodiments, the first assay covers one or more sub-genomic intervals not covered by the second assay. In certain embodiments, the second assay covers one or more sub-genomic intervals not covered by the first assay.

[0095] In some embodiments, even though the estimated tumor fraction may have a wide error range across patients, any within-patient comparison provides a small error range and results in the ability to track the progression of tumors initially identified in a comprehensive assay (e.g., FoundationOne, FoundationOne CDx or FoundationOne Liquid assay). Since the second assay can be much less expensive than the comprehensive assay, it can be used as a standard screening technique for subsets of patients, such as at least at-risk patients, to answer the question of whether a patient has cancer.

[0096] FIG. 1 shows a method 100 for estimating a tumor fraction from a sample. Method 100 begins at step 102. At step 104, values for target variables associated with subgenomic intervals are obtained directly from, for example, a sample from a subject. The target variable may be, for example, an allele fraction. The sample can be, for example, a liquid sample or a solid sample.

[0097] In some examples, the patient allele fraction for at least one heterozygous single nucleotide polymorphism (SNP) site is identified from a biopsy taken from the patient. In one example, the biopsy can be a liquid biopsy, i.e., a non-solid biological tissue, such as a sample of blood. However, the present disclosure is not so limited and is intended to encompass any solid or liquid assay or biopsy without limitation. In one embodiment, the liquid biopsy includes a blood sample. In one embodiment, the liquid biopsy includes cell-free DNA (cfDNA). In one embodiment, the liquid biopsy includes circulating tumor DNA (ctDNA). In one embodiment, the liquid biopsy includes DNA shedding from a tumor. In one embodiment, the liquid biopsy includes nucleic acids other than DNA, such as RNA. In one embodiment, the liquid biopsy includes circulating tumor cells (CTCs). Other types of liquid biopsies are described, for example, in Crowley et al., Nat Rev Clin Oncol. 2013;10(8):472-484, the entire contents of which are incorporated by reference.

[0098] In step 106, a confidence metric can be specified from a target variable, and in step 108, the specified relationship is accessed between the stored confidence metric and the stored tumor fraction. The specified relationship can include historical sample data (collected from a patient or other test subject) that associates a confidence metric (e.g., a sampled allele fraction deviation) for at least one heterozygous SNP site with a corresponding sampled tumor fraction. In some examples, the sampled allele coverage deviation is a "noise" metric that reflects the degree to which the allele fraction varies from an expected value. In some examples, the number of data points correlating the tumor fraction with the noise metric calculated from the allele portion can exceed one hundred (100), one thousand (1,000), ten thousand (10,000), or more.

[0099] In one example, the specified relationship may be derived from an in-silico process and the analysis may be performed by a machine learning process. This process may start with a particular tumor fraction and perform sample dilution (e.g., using a matched norm) to correlate one or more coverage deviation metrics (e.g., allele fraction values) across one or more sub-genomic intervals (e.g., SNPs, SNP bins, and / or chromosomes). The metric can be a measure of the frequency and degree to which the tumor fraction falls between the values of 0 or 1. The average "noise" metric between 0 and 1 (exclusive) can correlate with the expected or estimated tumor fraction.

[0100] The number of elements related to the sub-genomic intervals contributing to the calculation of the confidence metric value that correlates with the tumor fraction can be on the order of ten (10), one hundred (100), one thousand (1,000), ten thousand (10,000), or more.

[0101] Due to the large number of elements related to the sub-genomic intervals that contribute to the calculation of the confidence metric in the correlation, the elements can be "binned" or aggregated in some examples by sub-genomic interval position or other characteristics. Binning can prevent a single (or small set of) elements from disproportionately weighting the correlation of the confidence metric and adversely affecting the estimated tumor fraction. For example, if one element of a single sub-genomic interval represents 5,000 copies of a copy variant, it can lead to an inaccurately high estimated tumor fraction. Thus, in some examples, the elements contributing to the confidence metric are averaged or aggregated by chromosome, for example, for each of the 22 relevant chromosomes. These 22 aggregated chromosome values can then be used to calculate a confidence metric that is next correlated with the tumor fraction, ensuring that a single sub-genomic interval (e.g., an SNP site) does not disproportionately affect the correlation. Other methods can be utilized to limit the impact of extreme copy number events, such as but not limited to preventing outlier elements from entering the confidence metric calculation.

[0102] In some examples, the correlation can be a mean (i.e., average) correlation, and upper and lower bound correlations are also calculated. In this way, the mean correlation is bounded by a 95% confidence interval.

[0103] A sub-genomic interval may contain one or several sub-genomic intervals and, in some examples, may be at least one heterozygous SNP site. The sub-genomic intervals can be selected based on various criteria. For example, the sub-genomic intervals can be selected based on how polymorphic the sub-genomic intervals are in a general healthy population and healthy sub-populations (including different genders, ages or ethnic backgrounds). It may be advantageous for the sub-genomic intervals to be quite different in the healthy population. The sequence-specific characteristics of the sub-genomic intervals can also be selected based on being "well-behaved", i.e., close to the expected allele frequencies such as 0, 0.5 and 1.0. Further, the region can be selected based on being "sufficiently covered", i.e., having typical coverage across the population of that site. Sub-genomic intervals can be excluded if they occur in simple repeats of gene families or any generally repeating sequences of DNA, as this feature can challenge the alignment methodology. In one embodiment, the sub-genomic intervals can be located in genomic regions that do not contain, or essentially do not contain, high homology, simple repeats or gene families.

[0104] In one embodiment, the sub-genomic interval contains a minor allele. As used herein, a "minor allele" is an allele other than the most common allele associated with a particular sub-genomic interval in a given population (e.g., the second most common allele or the least common allele). In one embodiment, at least 10, 20, 50, 100, 150, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1200, 1400, 1600, 1800, 2000 or 10000 heterozygous sub-genomic intervals are selected. In one example, 10, 20, 50, 100, 150, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1200, 1400, 1600, 1800, 2000 or 10000 or fewer heterozygous SNP sites are selected.

[0105] In one example, the selected sub-genomic intervals and / or correlations can be universal, i.e., across all disease ontologies, to provide a broad screening technique. In other examples, the sub-genomic intervals are selected based on a disease ontology (e.g., tumor type), and the correlations can be adjusted.

[0106] One or more accuracy metrics can be used to correlate a target variable (e.g., allele coverage deviation and / or allele fraction variation) with the tumor fraction. For example, a metric regarding the allele fraction can be applied. In one example, an allele frequency entropy metric or a root mean square deviation (RMSD) metric can be used. Allele frequency entropy: JPEG0007702360000074.jpg17170Root mean square deviation: JPEG0007702360000075.jpg19170Where i is the SNP bin and af is the allele frequency in the range from 0 to 0.5. The folded SNP allele frequencies are used here by convention (e.g., as described in Nielsen. Hum Genomics. 2004;1(3):218 - 224 and Marth et al., Genetics. 2004;6(1):351 - 372), but the methodology holds if the full range from 0 to 1 is utilized. Other metrics such as metrics based on log2 ratios can also be used. Any of these metrics can incorporate factors such as coverage at a particular SNP bin, and a "bin" can be defined as one or more base pairs. In some embodiments, the accuracy metric may be written as a function of coverage such that accuracy metric = f(Cvg). Further, any mathematical transformation or operation acting on the accuracy metric can also be considered an accuracy metric.

[0107] In some examples, the confidence metric can be a deviation from an expected log2 ratio for at least one sub-genomic interval. In other examples, the confidence metric can be a deviation from an expected allele fraction in a healthy population for at least one sub-genomic interval (e.g., SNP) known to be heterozygous. In other examples, the confidence metric can be a deviation from an expected allele coverage in a healthy population for at least one sub-genomic interval (e.g., SNP) known to be heterozygous.

[0108] Table 1 shows exemplary confidence metrics that can be used, including any p-moment or combinations thereof.

Table 1

[0109] In step 110, the tumor fraction of the sample is identified (e.g., estimated) with reference to the confidence metric and the identified relationship. In some examples, the coefficients of the identified relationship are applied to the confidence metric identified from the patient sample, and the product reaches the tumor fraction that has been evaluated (e.g., estimated) in total. It will be understood that other functions can be performed to obtain the final estimated tumor fraction. For example, the estimated tumor ratio can be scaled, normalized, or adjusted in other ways from the measured value of the initial or raw estimated tumor ratio.

[0110] In step 112, method 100 ends.

[0111] The estimated tumor fraction can be used by medical practitioners in several ways. For example, the estimated tumor fraction can be used to monitor patients at risk for one or more types of cancer. The estimated tumor fraction can also be used to diagnose cancer or to determine whether cancer treatment is successfully affecting the tumor.

[0112] The estimated tumor fraction can also be used in connection with other screening techniques to confirm or validate test results. For example, CNA screening can yield multiple possible combinations of purity and ploidy for a patient, particularly a patient with a low tumor fraction (e.g., less than 30%). This technique can be used to clarify such results.

[0113] In some embodiments, a report including the estimated tumor fraction can be generated. In one embodiment, the report further includes treatment options based on the estimated tumor fraction. In one embodiment, the report further includes a prognosis based on the estimated tumor fraction.

[0114] Methods of Treating and Monitoring Tumors Also disclosed is a method of treating a subject's disease. The method is to administer an effective amount of therapy to the subject in response to the identification (e.g., estimation) of a tumor fraction (e.g., identified according to the methods described herein), thereby treating the disease, wherein the estimation of the tumor fraction includes obtaining a value of a target variable associated with a subgenomic interval in a sample, identifying an accuracy metric from the target variable, accessing a specified relationship between the stored accuracy metric and the stored tumor fraction, and identifying the tumor fraction of the sample with reference to the accuracy metric and the specified relationship.

[0115] In one embodiment, the method further includes administering a second therapy to the subject. In one embodiment, the method further includes discontinuing a second therapy for the subject. In one embodiment, the method further includes identifying the presence of somatic changes (e.g., somatic changes associated with the disease) in the subject.

[0116] In one embodiment, the allele fraction is determined by a method that includes sequence identification, such as next-generation sequence identification (NGS). In one embodiment, the allele fraction is determined by a method that further includes target selection, such as solution hybridization. In other embodiments, other methodologies used to detect DNA (e.g., cfDNA, ctDNA, etc.), such as microarrays, can be used.

[0117] A method for evaluating a disease in a subject, wherein determining (e.g., estimating) a tumor fraction (e.g., determined according to the methods described herein) includes obtaining a value for a target variable associated with a subgenomic interval in a sample, identifying a confidence metric from the target variable, accessing a specified relationship between the stored confidence metric and the stored tumor fraction, and referring to the confidence metric and the specified relationship to identify the tumor fraction of the sample, thereby evaluating the disease, is also described. In one embodiment, the allele fraction is determined by a method that includes sequence identification, such as NGS. In one embodiment, the allele fraction is determined by a method that further includes target selection, such as solution hybridization. In other embodiments, other methodologies used to detect DNA (e.g., cfDNA, ctDNA, etc.), such as microarrays, can be used. In one embodiment, the method further includes selecting a treatment method for the disease. In one embodiment, the method further includes discontinuing treatment of the subject. In one embodiment, the method further includes selecting a subject for a clinical trial. In one embodiment, the method further includes determining a disease state, such as remission, stability, recurrence, etc. In one embodiment, the disease is evaluated periodically, e.g., monthly, every two months, every three months, every six months, or annually. In one embodiment, the method further includes identifying the presence of somatic changes (e.g., somatic changes associated with the disease) in the subject.

[0118] A method for evaluating a subject, wherein the determination (e.g., estimation) of the tumor fraction (e.g., as determined according to the method described herein) includes obtaining a value for a target variable related to a subgenomic interval in a sample, identifying a confidence metric from the target variable, accessing a specified relationship between the stored confidence metric and the stored tumor fraction, and referring to the confidence metric and the specified relationship to identify the tumor fraction of the sample, thereby evaluating the subject. In one embodiment, the allele fraction is determined by a method including sequence identification, e.g., by NGS. In one embodiment, the allele fraction is determined by a method further including target selection, e.g., by solution hybridization. In other embodiments, other methodologies for detecting DNA (e.g., cfDNA, ctDNA, etc.), such as using microarrays, can be used.

[0119] In one embodiment, the method further includes selecting a subject for treatment. In one embodiment, the method further includes discontinuing treatment of a subject. In one embodiment, the method further includes selecting a subject for a clinical trial.

[0120] In one embodiment, the subject is evaluated periodically, e.g., monthly, every two months, every three months, every six months, or annually.

[0121] In one embodiment, the method further includes identifying the presence of somatic changes (e.g., somatic changes associated with a disease) in the subject.

[0122] In one embodiment, the target variable (e.g., allele fraction) is determined by a method including sequence identification, e.g., by NGS. In one embodiment, the allele fraction is determined by a method further including target selection, e.g., by solution hybridization. In other embodiments, other methodologies for detecting DNA (e.g., cfDNA, ctDNA, etc.), such as using microarrays, can be used.

[0123] A method for evaluating a treatment, wherein specifying (e.g., estimating) a tumor fraction (e.g., specified according to the method described herein) includes obtaining a value for a target variable related to a sub-genomic interval in a sample, specifying a confidence metric from the target variable, accessing a specified relationship between the stored confidence metric and the stored tumor fraction, and specifying the tumor fraction of the sample with reference to the confidence metric and the specified relationship, thereby evaluating the treatment, is also described.

[0124] In one embodiment, the target variable (e.g., allele fraction) is specified by a method including sequence identification, e.g., by NGS. In one embodiment, the allele fraction is specified by a method further including target selection, e.g., by solution hybridization. In other embodiments, other methodologies used to detect DNA (e.g., cfDNA, ctDNA, etc.), such as using microarrays, can be used.

[0125] In one embodiment, the method further includes selecting a treatment for the subject.

[0126] In one embodiment, the treatment is evaluated periodically, e.g., monthly, every two months, every three months, every six months, or annually.

[0127] A method for providing a report (e.g., for reporting a tumor fraction specified according to the method described herein) is described. The method includes obtaining a value for a target variable related to a sub-genomic interval in a sample, specifying a confidence metric from the target variable, accessing a specified relationship between the stored confidence metric and the stored tumor fraction, specifying the tumor fraction of the sample with reference to the confidence metric and the specified relationship, and recording the estimated tumor fraction in a report, thereby providing the report.

[0128] In one embodiment, the allele fraction is determined by a method that includes sequence identification, such as by NGS. In one embodiment, the allele fraction is determined by a method that further includes target selection, such as by solution hybridization. In other embodiments, other methodologies used to detect DNA (e.g., cfDNA, ctDNA, etc.), such as microarrays, can be used.

[0129] In one embodiment, the method further includes transmitting a report to the subject or a third party. In one embodiment, the report further includes treatment options based on the estimated tumor fraction.

[0130] In one embodiment, reporting further includes the subject's genomic profile (e.g., a genomic profile related to a disease).

[0131] A method for evaluating a biopsy from a subject (e.g., including determining a tumor fraction according to the methods described herein) is described. The method includes obtaining a value of a target variable related to a sub-genomic interval in a sample from the biopsy, determining an accuracy metric from the target variable, accessing a specified relationship between the stored accuracy metric and the stored tumor fraction, and specifying the tumor fraction of the sample with reference to the accuracy metric and the specified relationship, thereby evaluating the biopsy.

[0132] In one embodiment, an estimated tumor fraction exceeding a threshold indicates that the biopsy is suitable for genomic profiling.

[0133] Exemplary computer implementation forms The above process is merely an exemplary embodiment of a system that can be used to estimate a tumor fraction. Such exemplary embodiments are not intended to limit the scope of the present disclosure. Neither the embodiments described herein nor the claims are intended to be limited to any particular embodiment unless such claims explicitly enumerate a limitation to a particular embodiment.

[0134] Various embodiments, their operations, and processes and methods related to various embodiments and variations of these methods and operations can be defined, individually or in combination, by a computer-readable signal tangibly embodied on a computer-readable medium, such as a non-volatile recording medium, an integrated circuit memory element, or a combination thereof. According to one embodiment, the computer-readable medium can be non-transitory in that computer-executable instructions can be stored on the medium permanently or semi-permanently. Such a signal can define instructions, for example, as part of one or more programs that, as a result of being executed by a computer, instruct the computer to perform one or more of the methods or operations described herein and / or their various embodiments, variations, and combinations. Such instructions can be written in any of a plurality of programming languages, such as Java, Visual Basic, C, C#, or C++, Fortran, Pascal, Eiffel, Basic, COBOL, etc., or any of their various combinations. The computer-readable medium on which such instructions are stored can be present in one or more of the components of the general-purpose computer described above and can be distributed across one or more of such components.

[0135] The computer-readable medium can be transportable such that the instructions stored thereon can be loaded into any computer system resource for implementing aspects of the present disclosure described herein. Further, it should be understood that the instructions stored on the computer-readable medium described above are not limited to instructions embodied as part of an application program executed on a host computer. Rather, the instructions can be embodied as any type of computer code (e.g., software or microcode) that can be used to program a processor to implement the above-described aspects of the present disclosure.

[0136] Various embodiments according to the present disclosure can be implemented on one or more computer systems. These computer systems can be, for example, general-purpose computers based on an Intel PENTIUM-type processor, Motorola PowerPC, Sun UltraSPARC, Hewlett-Packard PA-RISC processor, ARM Cortex processor, Qualcomm Scorpion processor, or any other type of processor. It should be understood that one or more of any type of computer system can be used to partially or fully automate the extension of offers to users and the repayment of offers according to various embodiments of the present disclosure. Further, the software design system may be located on a single computer or may be distributed among multiple computers connected by a communication network.

[0137] The computer system can include specially programmed dedicated hardware, such as an application-specific integrated circuit (ASIC). Aspects of the present disclosure can be implemented in software, hardware, firmware, or any combination thereof. Further, such methods, operations, systems, system elements, and their components may be implemented as part of the computer systems described above or as independent components.

[0138] The computer system may be a general-purpose computer system programmable using a high-level computer programming language. The computer system may also be implemented using specially programmed dedicated hardware. The computer system may typically include a processor, which is a commercially available processor such as a well-known Pentium-class processor available from Intel Corporation. Many other processors are available. Such processors typically execute an operating system, which may be, for example, the Windows NT, Windows 2000 (Windows ME), Windows XP, Windows Vista or Windows 7 operating system available from Microsoft Corporation, the MAC OS X Snow Leopard, MAC OS X Lion operating system available from Apple Computer, the Solaris operating system available from Oracle Corporation, the iOS, Blackberry OS, Windows 7 Mobile or Android OS operating system, or UNIX available from various sources. Many other operating systems may be used.

[0139] Some aspects of the present disclosure may be implemented as distributed application components that can be executed on several different types of systems coupled via a computer network. Some components may be located and executed on mobile devices, servers, tablets, or other system types. Other components of the distributed system, such as databases or other component types, may also be used.

[0140] Both the processor and the operating system define a computer platform on which application programs in a high-level programming language are written. It should be understood that the present disclosure is not limited to a particular computer system platform, processor, operating system, set of computational algorithms, code, or network. Further, it should be understood that in a distributed computer system implementing various aspects of the present disclosure, multiple computer platform types can be used. Also, it will be apparent to those skilled in the art that the present disclosure is not limited to a particular programming language, set of computational algorithms, code, or computer system. Further, it should be understood that other suitable programming languages and other suitable computer systems can also be used.

[0141] One or more portions of a computer system may be distributed across one or more computer systems coupled to a communication network. These computer systems may also be general-purpose computer systems. For example, various aspects of the present disclosure may be distributed among one or more computer systems configured to provide services (e.g., servers) to one or more client computers or to perform overall tasks as part of a distributed system. For example, various aspects of the present disclosure may be executed on a client-server system that includes distributed components among one or more server systems that perform various functions according to various embodiments of the present disclosure. These components may be executable code, intermediate code (e.g., IL), or interpreted code (e.g., Java) that communicates via a communication network (e.g., the Internet) using a communication protocol (e.g., TCP / IP). Particular aspects of the present disclosure may also be implemented on a cloud-based computer system (e.g., an EC2 cloud-based computing platform provided by Amazon.com), a distributed computer network including clients and servers, or any combination of systems.

[0142] It should be understood that the present disclosure is not limited to being executed on any particular system or group of systems. It should also be understood that the present disclosure is not limited to any particular distributed architecture, network, or communication protocol.

[0143] Various embodiments of the present disclosure can be programmed using object-oriented programming languages such as SmallTalk, Java, C++, Ada, or C# (C-Sharp). Other object-oriented programming languages may also be used. Alternatively, functional, scripting, and / or logic programming languages may be used. Various aspects of the present disclosure may be implemented in an environment that is not programmed (e.g., a document created in HTML, XML, or other format that renders aspects of a graphical user interface (GUI) or performs other functions when displayed in a window of a browser program). Various aspects of the present disclosure may be implemented as programmed or unprogrammed elements, or any combination thereof.

[0144] Furthermore, in each of one or more computer systems including one or more components of a device, each component can be present at one or more locations on the system. For example, different portions of a component of a device may be present in different regions of memory (e.g., RAM, ROM, disk, etc.) on one or more computer systems. Each of such one or more computer systems can include a plurality of known components such as, among other components, one or more processors, a memory system, a disk storage system, one or more network interfaces, and one or more buses or other internal communication links interconnecting the various components.

[0145] The present disclosure can be implemented on a computer system, which will be described later in connection with FIGS. 2 and 3. In particular, FIG. 2 shows an exemplary computer system 200 used to implement various aspects. FIG. 3 shows an exemplary storage system that can be used.

[0146] System 200 is merely an exemplary embodiment of a computer system suitable for implementing various aspects of the present disclosure. Such exemplary embodiments are not intended to limit the scope. For example, any number of other implementations of the system are possible and are intended to fall within the scope of the present disclosure. For example, a virtual computing platform can be used. None of the claims described below are intended to be limited to any particular implementation of the system, unless such claims explicitly enumerate specific embodiments.

[0147] Various embodiments according to the present disclosure can be implemented on one or more computer systems. These computer systems may be general-purpose computers, such as those based on an Intel PENTIUM-type processor, Motorola PowerPC, Sun UltraSPARC, Hewlett-Packard PA-RISC processor, or any other type of processor. It should be understood that one or more of any type of computer system can be used to partially or fully automate the integration of security services with other systems and services according to various embodiments of the present disclosure. Further, the software design system may be located on a single computer or distributed among multiple computers connected by a communication network.

[0148] For example, various aspects of the present disclosure may be implemented as specialized software executed on a general-purpose computer system 200 as shown in FIG. 2. The computer system 200 can include a processor 203 connected to one or more memory devices 204 such as a disk drive, memory, or other device for storing data. The memory 204 is typically used to store programs and data during operation of the computer system 200. The components of the computer system 200 can be coupled by an interconnection mechanism 205, which can include one or more buses (e.g., between components integrated within the same machine) and / or networks (e.g., between components present in separate individual machines). The interconnection mechanism 205 enables the exchange of communication (e.g., data, instructions) between the system components of the system 200. The computer system 200 also includes one or more input devices 202 such as, for example, a keyboard, mouse, trackball, microphone, touch screen, etc., and one or more output devices 201 such as, for example, a printing device, display screen, and / or speaker. Further, the computer system 200 can include one or more interfaces (not shown) that connect the computer system 200 to a communication network (in addition to, or instead of, the interconnection mechanism 205).

[0149] The storage system 206 is shown in more detail in FIG. 3 and typically includes a computer-readable and writable non-volatile recording medium that stores signals defining a program to be executed by a processor or information stored on or within a medium 301 to be processed by the program. The medium can be, for example, a disk or flash memory. Typically, during operation, the processor causes the non-volatile recording medium 301 to read data into another memory 302 that enables faster access to the information by the processor than the medium 301. This memory 302 is typically a volatile random access memory such as dynamic random access memory (DRAM) or static random access memory (SRAM).

[0150] As shown in the figure, the data may be arranged within the storage system 206 or within the memory system 204. The processor 203 generally operates on the data in the integrated circuit memories 204, 202 and then copies the data to the medium 301 after the processing is completed. Various mechanisms are known for managing the data movement between the medium 301 and the integrated circuit memory element 302, and the present disclosure is not limited thereto. The present disclosure is not limited to a particular memory system 204 or storage system 206.

[0151] The computer system can include specially programmed dedicated hardware, such as an application specific integrated circuit (ASIC). Aspects of the present disclosure can be implemented in software, hardware, or firmware, or any combination thereof. Further, such methods, operations, systems, system elements, and their components may be implemented as part of the computer system described above or as independent components.

[0152] The computer system 200 is shown as an example of one type of computer system that can implement various aspects of the present disclosure, but it should be understood that aspects of the present disclosure are not limited to being implemented on a computer system as shown in FIG. 2. Various aspects of the present disclosure can be implemented on one or more computers having an architecture or components different from those shown in FIG. 2.

[0153] The computer system 200 may be a general-purpose computer system programmable using a high-level computer programming language. The computer system 300 may also be implemented using specially-programmed dedicated hardware. In the computer system 200, the processor 203 is typically a commercially available processor such as a well-known Pentium, Core, Core Vpro, Xeon, or Itanium class processor available from Intel Corporation. Many other processors are available. Such processors typically execute an operating system such as Linux, Windows NT, Windows 2000 (Windows ME), Windows XP, Windows Vista, Windows 7, or Windows 10 available from Microsoft Corporation, MAC OS Snow Leopard, MAC OS X Lion operating system available from Apple Computer, Solaris operating system available from Sun Microsystems, iOS, Blackberry OS, Windows 7 Mobile or Android OS operating system, or a UNIX which may be available from various sources. Many other operating systems may be used.

[0154] Together, the processor and the operating system define a computer platform on which application programs in a high-level programming language are written. It should be understood that the present disclosure is not limited to a particular computer system platform, processor, operating system, or network. It will also be apparent to those skilled in the art that the present disclosure is not limited to a particular programming language or computer system. Further, it should be understood that other suitable programming languages and other suitable computer systems may be used.

[0155] One or more portions of a computer system may be distributed across one or more computer systems (not shown) coupled to a communication network. These computer systems may also be general-purpose computer systems. For example, various aspects of the present disclosure may be distributed among one or more computer systems configured to provide services (e.g., servers) to one or more client computers or to perform overall tasks as part of a distributed system. For example, various aspects of the present disclosure may be executed on a client-server system that includes components distributed among one or more server systems that perform various functions according to various embodiments of the present disclosure. These components may be executable code, intermediate code (e.g., IL), or interpreted code (e.g., Java) that communicates via a communication network (e.g., the Internet) using a communication protocol (e.g., TCP / IP).

[0156] It should be understood that the present disclosure is not limited to being executed on any particular system or group of systems. It should also be understood that the present disclosure is not limited to any particular distributed architecture, network, or communication protocol.

[0157] Various embodiments of the present disclosure can be programmed using object - oriented programming languages such as SmallTalk, Java, C++, Ada, or C# (C - Sharp). Other object - oriented programming languages may also be used. Alternatively, functional, scripting, and / or logic programming languages may be used. Various aspects of the present disclosure may be implemented in an unprogrammed environment (e.g., a document created in HTML, XML, or other format that renders aspects of a graphical user interface (GUI) or performs other functions when displayed in a window of a browser program). Various aspects of the present disclosure can be implemented using various Internet technologies such as, for example, well - known Common Gateway Interface (CGI) scripts, PHP Hypertext Preprocessor (PHP), Active Server Pages (ASP), Hypertext Markup Language (HTML), Extensible Markup Language (XML), Java, JavaScript, Asynchronous JavaScript and XML (AJAX), Flash, and other programming methods. Further, various aspects of the present disclosure can be implemented on cloud - based computing platforms such as, among others, the well - known EC2 platform commercially available from Amazon.com (Seattle, WA). Various aspects of the present disclosure may be implemented as programmed or unprogrammed elements, or any combination thereof.

[0158] Definitions Certain terms are defined. Additional terms are defined throughout this specification.

[0159] As used herein, the articles “a” and “an” refer to one or more than one (e.g., up to at least one) of the grammatical object of the article.

[0160] "About" and "approximately" generally mean an acceptable degree of error of the measured quantity, taking into account the nature or accuracy of the measurement. Exemplary degrees of error are within 20% (%), typically within 10%, more typically within 5% of a given value or range of values.

[0161] As used herein, "obtain" or "obtaining" refers to obtaining ownership of a physical entity or value, such as a numerical value, either "directly obtain" or "indirectly obtain" a physical entity or value. "Directly obtain" means performing a process for obtaining a physical entity or value (e.g., performing a synthesis or analysis method). "Indirectly obtain" refers to receiving a physical entity or value from another entity or source (e.g., a third-party research institute that directly obtains a physical entity or value). Directly obtaining a physical entity includes performing a process that includes a physical change of a physical substance, such as a starting material. Exemplary changes include making a physical entity from two or more starting materials, shearing or fragmenting a substance, separating or purifying a substance, combining two or more separate entities into a mixture, and performing a chemical reaction that includes breaking or forming a covalent or non-covalent bond. Directly obtaining a value includes performing a process that includes a physical change of a substance, such as a sample, analyte, or reagent (sometimes referred to herein as "physical analysis"), performing a process that includes a physical change of a sample or another substance, separating or purifying an analyte, or a fragment or other derivative thereof, from another substance, combining an analyte, or a fragment or other derivative thereof, with another substance, such as a buffer, solvent, or reactant, or changing the structure of an analyte, or a fragment or other derivative thereof, for example, by breaking or forming a covalent or non-covalent bond between a first atom and a second atom of the analyte, or changing the structure of a reagent, or a fragment or other derivative thereof, for example, by breaking or forming a covalent or non-covalent bond between a first atom and a second atom of the reagent, and performing a method that includes one or more of the foregoing.

[0162] "Obtaining an array" or "obtaining a read", as the term is used herein, refers to obtaining ownership of a nucleotide or amino acid sequence by "directly obtaining" or "indirectly obtaining" the array or read. "Directly obtaining" an array or read means performing a process for obtaining the array, such as performing a sequence identification method (e.g., next-generation sequencing (NGS) method), i.e., performing a synthesis or analysis method. "Indirectly obtaining" an array or read means receiving information or knowledge of the array from another entity or source (e.g., a third-party laboratory that directly obtained the sequence), or receiving the array. The obtained array or read need not be a complete sequence, e.g., a sequence identification of at least one nucleotide, or obtaining information or knowledge that identifies one or more of the variations disclosed herein as being present in a sample, biopsy, or subject constitutes obtaining the sequence.

[0163] Directly obtaining an array or a read involves performing a process that includes physical changes to a physical substance, such as a starting material like a sample described herein. Exemplary changes include creating a physical entity from two or more starting materials, shearing or fragmenting a substance such as genomic DNA fragments, separating or purifying a substance (e.g., isolating a nucleic acid sample from a tissue), combining two or more separate entities into a mixture, performing a chemical reaction that includes breaking or forming covalent or non-covalent bonds. Directly obtaining a value involves performing a process that includes physical changes to a sample or another substance as described above. The size of the fragment (e.g., the average size of the fragment) can be 2500 bp or less, 2000 bp or less, 1500 bp or less, 1000 bp or less, 800 bp or less, 600 bp or less, 400 bp or less, or 200 bp or less. In some embodiments, the size of the fragment (e.g., cfDNA) is from about 150 bp to about 200 bp (e.g., from about 160 bp to about 170 bp). In some embodiments, the size of the fragment (e.g., DNA fragments from FFPE samples) is from about 150 bp to about 250 bp. In some embodiments, the size of the fragment (e.g., cDNA fragments obtained from RNA in FFPE samples) is from about 100 bp to about 150 bp.

[0164] "Obtaining a sample", as the term is used herein, refers to obtaining ownership of a sample, e.g., a sample described herein, either "directly obtaining" or "indirectly obtaining" the sample. "Directly obtaining a sample" means performing a process for obtaining the sample (e.g., performing a physical method such as surgery or extraction). "Indirectly obtaining a sample" refers to receiving a sample from another entity or source (e.g., a third-party laboratory that directly obtained the sample). Directly obtaining a sample includes performing a process that involves physical changes to physical substances, e.g., starting materials, e.g., tissues, e.g., tissues of a human patient or tissues previously isolated from a patient. Exemplary alterations include creating a physical entity from a starting material, slicing or rubbing a tissue, separating or purifying a substance (e.g., a tissue sample or a nucleic acid sample); combining two or more separate entities into a mixture; performing a chemical reaction that involves breaking or forming a covalent or non-covalent bond. Directly obtaining a sample includes, for example as described above, performing a process that involves physical changes to the sample or another substance.

[0165] As used herein, a "change" or "altered structure" of a gene or gene product (e.g., a marker gene or gene product) refers to a mutation within the gene or gene product, such as the presence of a mutation, that affects the integrity, sequence, structure, amount, or activity of the gene or gene product as compared to a normal or wild-type gene. A change can be the amount, structure, and / or activity in cancerous tissue or cancer cells as compared to its amount, structure, and / or activity in normal or healthy tissue or cells (e.g., a control), and is associated with a disease state such as cancer. For example, a change associated with cancer or predictive of responsiveness to anti-cancer therapy can be an altered nucleotide sequence (e.g., a mutation), amino acid sequence, chromosomal translocation, intrachromosomal inversion, copy number, expression level, protein level, protein activity, epigenetic modification (e.g., methylation or acetylation state, or post-translational modification) in cancerous tissue or cancer cells as compared to normal healthy tissue or cells. Exemplary mutations include, but are not limited to, point mutations (e.g., silent, missense, or nonsense), deletions, insertions, inversions, duplications, amplifications, translocations, interchromosomal and intrachromosomal rearrangements. Mutations can be present in the coding or non-coding regions of a gene. In certain embodiments, a change is detected as a genomic rearrangement, including a rearrangement, e.g., one or more introns or fragments thereof (e.g., one or more rearrangements in the 5'-UTR and / or 3'-UTR). In certain aspects, a change is (or is not) associated with a phenotype, such as a cancerous phenotype (e.g., one or more of cancer risk, cancer progression, cancer treatment, or resistance to cancer treatment). In one embodiment, a change (or tumor mutation burden) is associated with one or more of a genetic risk factor for cancer, a positive treatment response predictor, a negative treatment response predictor, a positive prognostic factor, a negative prognostic factor, or a diagnostic factor.

[0166] As used herein, the term "indel" refers to an insertion, deletion, or both of one or more nucleotides in the nucleic acid of a cell. In certain embodiments, an indel includes both an insertion and a deletion of one or more nucleotides, where both the insertion and the deletion are in close proximity on the nucleic acid. In certain embodiments, an indel results in a net change in the total number of nucleotides. In certain embodiments, an indel results in a net change of from about 1 to about 50 nucleotides.

[0167] "Clonal profile", as the term is used herein, refers to the occurrence, identity, variability, distribution, expression (occurrence or level of transcriptional copies of subgenomic signatures) or abundance, e.g., relative abundance, of one or more sequences, e.g., alleles or signatures, of a target interval (or of a cell comprising the same). In one embodiment, a clonal profile is a value for the relative abundance of one sequence, allele or signature with respect to a target interval when a plurality of sequences, alleles or signatures with respect to the target interval (or of a cell comprising the same) are present in a sample. For example, in one embodiment, a clonal profile includes values for the relative abundance of one or more of a plurality of VDJ or VJ combinations with respect to a target interval. In one embodiment, a clonal profile includes a value for the relative abundance of a selected V segment with respect to a target interval. In one embodiment, a clonal profile includes values for diversity, e.g., arising from somatic hypermutation within the sequence of a target interval. In one embodiment, a clonal profile includes values for the occurrence or level of expression of a sequence, allele or signature, as evidenced, e.g., by the occurrence or level of expression of an expressed subgenomic interval comprising the sequence, allele or signature.

[0168] "Expression sub-genomic interval", as the term is used herein, refers to the transcribed sequences of a sub-genomic interval. In one embodiment, the sequence of the expression sub-genomic interval may be different from the sub-genomic interval that is transcribed, for example, because some sequences may not be transcribed.

[0169] "Minor allele frequency" (MAF), as the term is used herein, refers to the relative frequency of a minor allele at a particular locus, e.g., in a sample. In some embodiments, the minor allele frequency is expressed as a ratio or percentage.

[0170] "Signature", as the term is used herein, refers to the sequence of a target interval. A signature can diagnose the occurrence of one of a plurality of possibilities in a target interval. For example, a signature can diagnose: the occurrence of a selected V segment in a rearranged heavy chain variable region gene or light chain variable region gene, the presence of a selected VJ junction, e.g., the presence of a selected V and a selected J segment in a rearranged heavy chain variable region gene. In one embodiment, a signature comprises a plurality of specific nucleic acid sequences. Thus, a signature is not limited to a specific nucleic acid sequence, but rather is sufficiently unique to distinguish between a first group of sequences or possibilities in a target interval and a second group of possibilities in the target interval, e.g., to distinguish between a first V segment and a second V segment, e.g., to enable the evaluation of the use of various V segments. The term "signature" includes the term "specific signature" which is a specific nucleic acid sequence. In one embodiment, a signature indicates or is the product of a particular event, e.g., a rearrangement event.

[0171] "Sub-genomic interval", as the term is used herein, refers to a portion of a genomic sequence. In one embodiment, the sub-genomic interval can be a single nucleotide position, for example, a variant at that position is (positively or negatively) associated with a tumor phenotype. In one embodiment, the sub-genomic interval comprises two or more nucleotide positions. Such embodiments include sequences that are at least 2, 5, 10, 50, 100, 150 or 250 nucleotides in length. The sub-genomic interval can include an entire gene or a portion thereof, such as a coding region (or a portion thereof), an intron (or a portion thereof) or an exon (or a portion thereof). The sub-genomic interval can include naturally occurring, for example, all or part of a fragment of genomic DNA, nucleic acid. For example, the sub-genomic interval can correspond to a fragment of genomic DNA that is subjected to a sequence-specific reaction. In one embodiment, the sub-genomic interval is a continuous sequence from a genomic source. In one embodiment, the sub-genomic interval includes sequences that are not contiguous in the genome, for example, a sub-genomic interval in cDNA can include exon-exon junctions formed as a result of splicing. In one embodiment, the sub-genomic interval includes a tumor nucleic acid molecule. In one embodiment, the sub-genomic interval includes a non-tumor nucleic acid molecule.

[0172] In one embodiment, the sub-genomic interval corresponds to a rearranged sequence, for example, a sequence in a B or T cell that results from the ligation of a V segment and a D segment, a D segment and a J segment, a V segment and a J segment, or a J segment and a class segment.

[0173] In one embodiment, the sub-genomic interval is represented by one sequence. In one embodiment, the sub-genomic interval is represented by two or more sequences, for example, a sub-genomic interval covering a VD sequence can be represented by two or more signatures.

[0174] In one embodiment, the sub-genomic interval is an intragenic region or an intergenic region; an exon or an intron, or a fragment thereof, typically an exon sequence or a fragment thereof; a coding region or a non-coding region, such as a promoter, an enhancer, a 5'untranslated region (5'UTR), or a 3'untranslated region (3'UTR), or a fragment thereof; a cDNA or a fragment thereof; an SNP; a somatic mutation, a germline mutation, or both. Changes, such as a point or a single mutation; a deletion mutation (e.g., an in-frame deletion, an intragenic deletion, a complete gene deletion); an insertion mutation (e.g., an intragenic insertion); an inversion mutation (e.g., an intrachromosomal inversion); an inverted duplication mutation; a tandem duplication (e.g., an intrachromosomal tandem duplication); a translocation (e.g., a chromosomal translocation, a non-reciprocal translocation); a rearrangement (e.g., a genomic rearrangement (e.g., a rearrangement of one or more introns, a rearrangement of one or more exons, or a combination and / or fragment thereof; the rearranged intron can include 5' and / or 3'-UTR)); a change in gene copy number; a change in gene expression; a change in RNA level; or a combination thereof, are included or consist of. "Gene copy number" refers to the number of DNA sequences in a cell that encode a particular gene product. Generally, for a given gene, a mammal has two copies of each gene. The copy number can be increased, for example, by gene amplification or duplication, or decreased by deletion.

[0175] "Target interval", when the term is used in this specification, refers to a subgenomic interval or an expressed subgenomic interval. In one embodiment, the subgenomic interval and the expressed subgenomic interval correspond, meaning that the expressed subgenomic interval contains sequences expressed from the corresponding subgenomic interval. In one embodiment, the subgenomic interval and the expressed subgenomic interval do not correspond, meaning that the expressed subgenomic interval does not contain sequences expressed from the non-corresponding subgenomic interval, but rather corresponds to a different subgenomic interval. In one embodiment, the subgenomic interval and the expressed subgenomic interval partially correspond, meaning that the expressed subgenomic interval contains sequences expressed from the corresponding subgenomic interval and sequences expressed from different corresponding subgenomic intervals.

[0176] As used herein, the term "library" refers to a collection of nucleic acid molecules. In one embodiment, the library comprises a collection of nucleic acid molecules, such as a collection of whole genomes, sub-genomic fragments, cDNA, cDNA fragments, RNA, such as mRNA, RNA fragments, or combinations thereof. Typically, the nucleic acid molecules are DNA molecules, such as genomic DNA or cDNA. The nucleic acid molecules can be fragmented, such as sheared or enzymatically prepared genomic DNA. The nucleic acid molecules contain sequences derived from a subject and can also contain sequences not derived from the subject, such as adapter sequences, primer sequences, or other sequences that enable identification, such as "barcode" sequences. In one embodiment, some or all of the library nucleic acid molecules contain adapter sequences. The adapter sequences can be located at one or both ends. The adapter sequences can be useful, for example, in sequence identification methods (such as NGS methods), amplification, reverse transcription, or cloning into vectors. The library can comprise a collection of nucleic acid molecules, such as target nucleic acid molecules (such as tumor nucleic acid molecules, reference nucleic acid molecules, or combinations thereof). The nucleic acid molecules of the library can be derived from a single individual. In an embodiment, the library can contain nucleic acid molecules from two or more subjects (such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30 or more subjects), for example, two or more libraries from different subjects can be combined to form a library containing nucleic acid molecules from two or more subjects. In one embodiment, the subject is a human having or at risk of having cancer or a tumor.

[0177] "Library catch" refers to a subset of the library, such as a subset in which a target interval is enriched, such as a product captured by hybridization with a target capture reagent.

[0178] As used herein, "target capture reagent" refers to a molecule capable of capturing a target. A target capture reagent (e.g., a bait or a target capture oligonucleotide) can include a nucleic acid molecule, such as a DNA or RNA molecule, that can hybridize (e.g.) and thereby enable capture of the target nucleic acid. In one embodiment, the target capture reagent includes a DNA molecule (e.g., a naturally occurring or modified DNA molecule), an RNA molecule (e.g., a naturally occurring or modified RNA molecule), or a combination thereof. In one embodiment, the target capture reagent is suitable for solution-phase hybridization.

[0179] "Complementary" refers to sequence complementarity between regions of two nucleic acid strands or between two regions of the same nucleic acid strand. It is known that an adenine residue in a first nucleic acid region can form a specific hydrogen bond ("base pairing") with a residue in a second nucleic acid region that is antiparallel to the first region when the residue is thymine or uracil. Similarly, it is known that a cytosine residue in a first nucleic acid strand can base pair with a residue in a second nucleic acid strand that is antiparallel to the first strand when the residue is guanine. When two regions are arranged antiparallel, a first region of a nucleic acid is complementary to a second region of the same or a different nucleic acid if at least one nucleotide residue of the first region can base pair with a residue of the second region. In certain embodiments, the first region includes a first portion and the second region includes a second portion, and when the first and second portions are arranged antiparallel, at least about 50%, at least about 75%, at least about 90%, or at least about 95% of the nucleotide residues of the first portion can base pair with the nucleotide residues of the second portion. In other embodiments, all of the nucleotide residues of the first portion can base pair with the nucleotide residues of the second portion.

[0180] The terms "cancer" and "tumor" are used interchangeably herein. These terms refer to the presence of cells having characteristics typical of cells that cause cancer, such as uncontrolled growth, immortality, metastatic ability, rapid growth and proliferation rates, and certain characteristic morphological features. Cancer cells are often in the form of tumors, but such cells can exist alone within an animal or can be non-tumorigenic cancer cells such as leukemia cells. These terms include solid tumors, soft tissue tumors, or metastatic lesions. As used herein, the term "cancer" includes pre-cancerous as well as malignant cancers.

[0181] As used herein, "likely" or "likelihood" refers to the probability that an article, object, thing or person will occur. Thus, in one example, a subject likely to respond to treatment has a higher probability of responding to treatment compared to a reference subject or group of subjects.

[0182] "Unlikely" refers to a decrease in the probability that an event, item, object, thing or person will occur relative to a reference. Thus, a subject unlikely to respond to treatment has a lower probability of responding to treatment compared to a reference subject or group of subjects.

[0183] "Control nucleic acid molecule" refers to a nucleic acid molecule having a sequence derived from a non-tumor cell.

[0184] As used herein, "next generation sequencing" or "NGS" or "NG sequencing" refers to the nucleotide sequence of individual nucleic acid molecules (e.g., in single molecule sequencing) or a clonally amplified proxy of individual nucleic acid molecules (e.g., 10 3 、10 4 、10 5Refers to any sequencing method that identifies (or simultaneously identifies) molecules above a certain number of molecules in a high-throughput manner. In one embodiment, the relative abundance of nucleic acid species in a library can be estimated by counting the relative number of occurrences of their homologous sequences in the data generated by the identification experiment. Next-generation sequencing methods are known in the art and are described, for example, in Metzker, M. (2010) Nature Biotechnology Reviews 11:31-46, which is incorporated herein by reference. Next-generation sequencing can detect variants present in less than 5% or less than 1% of the nucleic acids in a sample.

[0185] As used herein, "nucleotide value" refers to the identity of the nucleotide occupying or assigned to a nucleotide position. Typical nucleotide values include deletions (e.g., removed); insertions of one or more nucleotides (e.g., whose identity may or may not be included); or presence (occupation); A; T; C; or G. Other values may be, for example, Y as A, T, G, or C, not Y, A or X (where X is one or two of T, G or C), T or X (where X is one or two of A, G or C), G or X (where X is one or two of T, A or C), C or X (where X is one or two of T, G or A), pyrimidine nucleotides, or purine nucleotides. The nucleotide value can be the frequency of one or more, e.g., 2, 3 or 4 bases (or other values described herein, e.g., missing or additional) at a nucleotide position. For example, the nucleotide value can include the frequency of A and the frequency of G at a nucleotide position.

[0186] "Or" is used herein to mean the term "and / or" and is used interchangeably therewith, unless the context clearly indicates otherwise. The use of the term "and / or" in several places in this specification does not mean that the use of the term "or" is not interchangeable with the term "and / or", unless the context clearly indicates otherwise.

[0187] "Primary control" refers to non-tumor tissue other than normal adjacent tissue (NAT) in a sample. A typical primary control is blood.

[0188] As used herein, "sample" refers to a biological sample obtained from or derived from a source of interest, as described herein. In some embodiments, the source of interest includes an organism such as an animal or a human. The source of the sample can be fresh, frozen, and / or preserved organs, tissue samples, biopsies, excisions, smears, or solid tissue from aspirates, blood or any blood component; body fluids such as cerebrospinal fluid, amniotic fluid, peritoneal fluid, interstitial fluid; or cells from any point in the pregnancy or development of the subject. In some aspects, the source of the sample is blood or a blood component.

[0189] In some embodiments, the sample is, or comprises, a biological tissue or body fluid. The sample can include compounds that are not naturally mixed with tissue in nature, such as preservatives, anticoagulants, buffers, fixatives, nutrients, antibiotics, etc. In one embodiment, the sample is stored as a frozen sample or as a formaldehyde or paraformaldehyde-fixed paraffin-embedded (FFPE) tissue preparation. For example, the sample can be embedded in a matrix, such as an FFPE block or a frozen sample. In another embodiment, the sample is a blood or blood component sample. In yet another embodiment, the sample is a bone marrow aspiration sample. In another embodiment, the sample contains cell-free DNA (cfDNA). In some embodiments, the cfDNA is DNA from cells undergoing apoptosis or necrotic cells. Typically, cfDNA is bound by proteins (e.g., histones) and protected by nucleases. CfDNA can be used as a biomarker for non-invasive prenatal testing (NIPT), organ transplantation, cardiomyopathy, microbiota, and cancer. In another embodiment, the sample contains circulating tumor DNA (ctDNA). In some embodiments, the ctDNA is cfDNA with genetic or epigenetic changes (e.g., somatic changes or methylation signatures) that can distinguish it from that derived from tumor and non-tumor cells. In another embodiment, the sample contains circulating tumor cells (CTC). In some aspects, CTCs are cells shed into the circulation from primary or metastatic tumors. In some aspects, CTC apoptosis is a source of ctDNA in the blood / lymph.

[0190] In some embodiments, the biological sample can be, or can include, bone marrow; blood; blood cells; ascites; tissue or fine needle biopsy sample; cell-containing body fluid; cell-free floating nucleic acid; sputum; saliva; urine; cerebrospinal fluid, peritoneal fluid; pleural fluid; feces; lymph; gynecological fluid; skin swab; vaginal swab; oral swab; nasal swab; lavage or lavage fluid such as bronchial alveolar lavage fluid; aspirate; scraping; bone marrow specimen; tissue biopsy specimen; surgical sample; feces, other body fluids, secretions and / or excretions; and / or cells therefrom. In some embodiments, the biological sample is or comprises cells obtained from an individual. In some aspects, the obtained cells are or comprise cells from the individual from whom the sample was obtained.

[0191] In some embodiments, the sample is a "primary sample" obtained directly from the source of interest by any suitable means. For example, in some embodiments, the primary biological sample is obtained by a method selected from biopsies (e.g., fine needle aspiration or tissue biopsy), surgery, collection of body fluids (e.g., blood, lymph or feces), etc. In some embodiments, as will be apparent from the context, the term "sample" refers to a preparation obtained by processing the primary sample (e.g., by removing one or more components and / or by adding one or more agents), e.g., by filtering through a semipermeable membrane. Such a "processed sample" can include, for example, nucleic acids or proteins extracted from the sample or obtained by subjecting the primary sample to techniques such as amplification or reverse transcription of mRNA, isolation and / or purification of specific components.

[0192] In one embodiment, the sample is a cell associated with a tumor, such as a tumor cell or a tumor infiltrating lymphocyte (TIL). In one embodiment, the sample comprises one or more pre-malignant or malignant cells. In one embodiment, the sample is obtained from a hematologic malignancy (or pre-malignant tumor), such as a hematologic malignancy (or pre-malignant tumor) described herein. In certain aspects, the sample is obtained from a solid tumor, a soft tissue tumor, or a metastatic lesion. In other embodiments, the sample comprises tissue or cells from a surgical margin. In another embodiment, the sample comprises one or more circulating tumor cells (CTC) (e.g., CTC obtained from a blood sample). In one embodiment, the sample is a cell not associated with a tumor, such as a non-tumor cell or a peripheral blood lymphocyte.

[0193] As used herein, "sensitivity" is a measure of the ability of a method to detect sequence variants in a heterogeneous population of sequences. The method has an ST% sensitivity for F% variants if, given a sample in which the sequence variant is present as at least F% of the sequences in the sample, the method can detect the sequence at a confidence of C% at that time 、 For example, considering a sample in which the variant sequence is present as at least 5% of the sequences in the sample, if the method can detect the sequence with 99% confidence 9 out of 10 times (F = 5%; C = 99%; ST = 90%), the method has 90% sensitivity for 5% variants. Exemplary sensitivities include ST = 95%, 99%, 99.9% sensitivities for F = 1%, 5%, 10%, 20%, 50%, 100% sequence variants at confidence levels of C = 90%, 90%, 95%, and 99%.

[0194] As used herein, "specificity" is a measure of the ability of a method to distinguish true sequence variants from sequence-specific artifacts or other closely related sequences. It is the ability to avoid false positive detections. False positive detections can arise from errors introduced into the sequence of interest during sample preparation, sequence-specific errors, or inadvertent sequence-specificity of closely related sequences such as pseudogenes or nucleic acid molecules of a gene family. X True The sequence is a true variant, and X Not trueN that is not a true variant Total When applied to a sample set of sequences, the method has a specificity of X% if it selects at least X% of the non-true variants as not being variants. For example, when applied to a sample set of 1,000 sequences where 500 are true variants and 500 are not true variants, the method has a specificity of 90% and selects 90% of the 500 sequences that are not true variants as not being variants. Exemplary specificities include 90, 95, 98, and 99%.

[0195] As used herein, "control nucleic acid" or "reference nucleic acid" refers to a nucleic acid molecule from a control or reference sample. Typically, it is DNA that does not contain changes or mutations in a gene or gene product, e.g., genomic DNA, or cDNA derived from RNA. In certain embodiments, the reference or control nucleic acid sample is a wild-type or non-mutated sequence. In certain embodiments, the reference nucleic acid sample is purified or isolated (e.g., it is removed from its natural state). In other embodiments, the reference nucleic acid sample is derived from a blood control, normal adjacent tissue (NAT), or any other non-cancerous sample from the same or a different subject. In some embodiments, the reference nucleic acid sample contains a normal DNA mixture. In some embodiments, the normal DNA mixture is a process compatibility control. In some embodiments, the reference nucleic acid sample has germline variants. In some embodiments, the reference nucleic acid sample has no somatic changes and functions, for example, as a negative control.

[0196] "Sequence identification" of a nucleic acid molecule requires identifying the identity of at least one nucleotide within the molecule (e.g., a DNA molecule, an RNA molecule, or a cDNA molecule derived from an RNA molecule). In embodiments, the identity of less than all of the nucleotides in the molecule is identified. In other embodiments, the identity of most or all of the nucleotides in the molecule is identified.

[0197] As used herein, a "threshold value" is a value that is a function of the number of reads that need to be present to assign a nucleotide value to a target interval (e.g., a sub-genomic interval or an expressed sub-genomic interval). For example, this is a function of the number of reads having a particular nucleotide value, e.g., "A", at a nucleotide position necessary to assign that nucleotide value to that nucleotide position within the sub-genomic interval. The threshold value can be expressed, for example, as a number of reads, e.g., as an integer (or as a function thereof), or as a percentage of the reads having that value. As an example, if the threshold value is X and there are X + 1 reads having a nucleotide value of "A", the value of "A" is assigned to the position within the target interval (e.g., a sub-genomic interval or an expressed sub-genomic interval). The threshold value can also be expressed as a function of the expected value of a mutation or variant, the mutation frequency, or the Bayesian prior value. In one embodiment, the mutation frequency will require the number or percentage of reads having a nucleotide value, e.g., A or G, at a position in order to call that nucleotide value. In embodiments, the threshold value can be a function of the mutation prediction, e.g., the mutation frequency, and the tumor type. For example, a variant at a nucleotide position can have a first threshold value if the patient has a first tumor type and a second threshold value if the patient has a second tumor type.

[0198] As used herein, a "target nucleic acid molecule" refers to a nucleic acid molecule that one wishes to isolate from a nucleic acid library. In one embodiment, the target nucleic acid molecule can be a tumor nucleic acid molecule, a reference nucleic acid molecule, or a control nucleic acid molecule, as described herein.

[0199] As used herein, the term "tumor nucleic acid molecule" or other similar terms (e.g., "tumor or cancer-related nucleic acid molecule") refers to a nucleic acid molecule having a sequence derived from a tumor cell. The terms "tumor nucleic acid molecule" and "tumor nucleic acid" may be used interchangeably herein. In one embodiment, the tumor nucleic acid molecule includes a target region having a sequence (e.g., a nucleotide sequence) with a change (e.g., a mutation) associated with a cancerous phenotype. In other embodiments, the tumor nucleic acid molecule includes a target region having a wild-type sequence (e.g., a wild-type nucleotide sequence). For example, a target region from a heterozygous or homozygous wild-type allele present in a cancer cell. The tumor nucleic acid molecule can include a reference nucleic acid molecule. Typically, it is DNA derived from a sample, e.g., genomic DNA, or cDNA derived from RNA. In certain embodiments, the sample is purified or isolated (e.g., it is removed from its natural state). In some embodiments, the tumor nucleic acid molecule is cfDNA. In some embodiments, the tumor nucleic acid molecule is ctDNA. In some embodiments, the tumor nucleic acid molecule is DNA derived from CTC.

[0200] As used herein, the term "reference nucleic acid molecule" or other similar terms (e.g., "control nucleic acid molecule") refers to a nucleic acid molecule that includes a target region having a sequence (e.g., a nucleotide sequence) not associated with a cancerous phenotype. In one embodiment, the reference nucleic acid molecule includes the wild-type or non-mutated nucleotide sequence of a gene or gene product that is associated with a cancerous phenotype when mutated. The reference nucleic acid molecule can be present in cancer cells or non-cancer cells.

[0201] As used herein, the term "variant" refers to a structure that can be present in a sub-genomic region that can have two or more structures, e.g., alleles of a polymorphic locus.

[0202] An "isolated" nucleic acid molecule is one that has been separated from other nucleic acid molecules that are present in the natural source of the nucleic acid molecule. In certain embodiments, an "isolated" nucleic acid molecule does not contain sequences (such as protein-coding sequences) that are naturally adjacent to the nucleic acid in the genomic DNA of the organism from which the nucleic acid is derived (i.e., the sequences located at the 5' and 3' termini of the nucleic acid). For example, in various embodiments, an isolated nucleic acid molecule may contain less than about 5 kB, less than about 4 kB, less than about 3 kB, less than about 2 kB, less than about 1 kB, less than about 0.5 kB, or less than about 0.1 kB of nucleotide sequences that are naturally adjacent to the nucleic acid molecule in the genomic DNA of the cell from which the nucleic acid is derived. Further, an "isolated" nucleic acid molecule, such as an RNA molecule or a cDNA molecule, may, for example, when produced by recombinant techniques, be substantially free of other cellular material or culture medium, or, for example, when chemically synthesized, be substantially free of chemical precursors or other chemicals.

[0203] The term "substantially free of other cellular material or culture medium" includes preparations of nucleic acid molecules in which the molecule is separated from the cellular components of the cell in which it is isolated or recombinantly produced. Thus, a nucleic acid molecule that is substantially free of cellular material includes preparations of nucleic acid molecules having less than about 30%, less than about 20%, less than about 10%, or less than about 5% (by dry weight) of other cellular material or culture medium.

[0204] As used herein, "X is a function of Y" means, for example, that one variable X is associated with another variable Y. The relationship between X and Y can be direct or indirect. In one embodiment, when X is a function of Y, a causal relationship between X and Y may be implied, but not necessarily present.

[0205] Titles, such as (a), (b), (i), etc., are presented simply to make the specification and claims easier to read. The use of headings in the specification or claims does not require that steps or elements be performed in alphabetical or numerical order, or in the order in which they are presented. The use of headings in the specification or claims also does not require all executions of steps or elements.

[0206] Multigene analysis The methods described herein can be used in combination with, or as part of, a method for evaluating a set of target intervals from, for example, a set of genes or gene products described herein.

[0207] In certain embodiments, the set of genes includes genes in mutant form that are related to effects on cell division, proliferation, or survival, or genes related to cancer, such as the cancers described herein.

[0208] In certain embodiments, the set of genes includes at least about 50 or more, about 100 or more, about 150 or more, about 200 or more, about 250 or more, about 300 or more, about 350 or more, about 400 or more, about 450 or more, about 500 or more, about 550 or more, about 600 or more, about 650 or more, about 700 or more, about 750 or more, or about 800 or more genes, as described herein. In some embodiments, the set of genes includes at least about 50 or more, about 100 or more, about 150 or more, about 200 or more, about 250 or more, about 300 or more, or all of the selected genes described in Tables 2A - 5B.

[0209] In certain embodiments, the method includes obtaining a library comprising a plurality of tumor nucleic acid molecules from a sample. In certain embodiments, the method further includes contacting the library with a target capture reagent to provide selected tumor nucleic acid molecules, wherein the target capture reagent hybridizes with tumor nucleic acid molecules from the library, thereby providing a library catch. In certain embodiments, the method further includes obtaining a read for a target region, such as by next-generation sequencing, that includes a change (e.g., a somatic change) from a tumor nucleic acid molecule from the library or the library catch. In certain particular embodiments, the method further includes aligning the read for the target region by an alignment method, such as an alignment method described herein. In certain embodiments, the method further includes assigning a nucleotide value for a nucleotide position from the read for the target region, such as by a mutation calling method described herein.

[0210] In certain embodiments, the method includes one, two, three, four, or all of the following: (a) obtaining a library comprising a plurality of tumor nucleic acid molecules from a sample; (b) contacting the library with a plurality of target capture reagents to provide selected tumor nucleic acid molecules, wherein the plurality of target capture reagents hybridize with the tumor nucleic acid molecules, thereby providing a library catch; (c) obtaining a read for a target region that includes a change (e.g., a somatic change) from a tumor nucleic acid molecule from the library catch, such as by next-generation sequencing, to obtain the read for the target region; (d) aligning the read by an alignment method, such as an alignment method described herein; or (e) assigning a nucleotide value for a nucleotide position from the read, such as by a mutation calling method described herein.

[0211] In certain embodiments, obtaining reads for a target interval comprises sequencing the target interval from at least about 50 or more, about 100 or more, about 150 or more, about 200 or more, about 250 or more, about 300 or more, about 350 or more, about 400 or more, about 450 or more, about 500 or more, about 550 or more, about 600 or more, about 650 or more, about 700 or more, about 750 or more, or about 800 or more genes. In certain embodiments, obtaining reads for a target interval comprises sequencing the target interval from at least about 50 or greater, about 100 or greater, about 150 or greater, about 200 or greater, about 250 or greater, about 300 or greater, or all of the genes listed in Tables 2A - 5B.

[0212] In certain embodiments, obtaining a read for a target interval includes sequencing at an average depth of 100X or more. In certain embodiments, obtaining a read for a target interval includes sequencing at an average depth of about 250X or more. In other embodiments, obtaining a read for a target interval includes sequencing at an average depth of about 500X or more. In certain embodiments, obtaining a read for a target interval includes sequencing at an average depth of about 800X or more. In other embodiments, obtaining a read for a target interval includes sequencing at an average depth of about 1,000X or more. In other embodiments, obtaining a read for a target interval includes sequencing at an average depth of about 1,500X or more. In other embodiments, obtaining a read for a target interval includes sequencing at an average depth of about 2,000X or more. In other embodiments, obtaining a read for a target interval includes sequencing at an average depth of about 2,500X or more. In certain embodiments, obtaining a read for a target interval includes sequencing at an average depth of about 3,000X or more. In certain embodiments, obtaining a read for a target interval includes sequencing at an average depth of about 3,500X or more. In certain embodiments, obtaining a read for a target interval includes sequencing at an average depth of about 4,000X or more. In certain embodiments, obtaining a read for a target interval includes sequencing at an average depth of about 4,500X or more. In certain embodiments, obtaining a read for a target interval includes sequencing at an average depth of about 5,000X or more. In certain embodiments, obtaining a read for a target interval includes sequencing at an average depth of about 5,500X or more. In certain embodiments, obtaining a read for a target interval includes sequencing at an average depth of about 6,000X or more.

[0213] In certain embodiments, obtaining reads for a target interval involves sequencing at an average depth of about 100X or more over about 99% of the sequenced gene (e.g., exon). In certain embodiments, obtaining reads for a target interval involves sequencing at an average depth of about 250X or more over about 99% of the sequenced gene (e.g., exon). In other embodiments, obtaining reads for a target interval involves sequencing at an average depth of about 500X or more over more than about 95% of the sequenced gene (e.g., exon). In other embodiments, obtaining reads for a target interval involves sequencing at an average depth of about 800X or more over more than about 95% of the sequenced gene (e.g., exon). In other embodiments, obtaining reads for a target interval involves sequencing at an average depth greater than about 1,000X over more than about 90% of the sequenced gene (e.g., exon). In other embodiments, obtaining reads for a target interval involves sequencing at an average depth of about 2,000X or more over more than about 90% of the sequenced gene (e.g., exon). In other embodiments, obtaining reads for a target interval involves sequencing at an average depth of about 3,000X or more over more than about 90% of the sequenced gene (e.g., exon). In other embodiments, obtaining reads for a target interval involves sequencing at an average depth of about 3,500X or more over more than about 90% of the sequenced gene (e.g., exon). In other embodiments, obtaining reads for a target interval involves sequencing at an average depth of about 4,000X or more over more than about 90% of the sequenced gene (e.g., exon). In other embodiments, obtaining reads for a target interval involves sequencing at an average depth of about 4,500X or more over more than about 90% of the sequenced genes (e.g., exons).In other embodiments, obtaining reads for a target interval includes sequencing at an average depth of about 5,000X or more across more than about 90% of the arrayed genes (e.g., exons). In other embodiments, obtaining reads for a target interval includes sequencing at an average depth of about 5,500X or more across more than about 90% of the genes (e.g., exons) that are arrayed. In other embodiments, obtaining reads for a target interval includes sequencing at an average depth of about 6,000X or more across more than about 90% of the genes (e.g., exons) that are arrayed. In certain embodiments, obtaining reads for a target interval includes sequencing at an average depth of greater than about 100X, greater than about 250X, greater than about 500X, greater than about 1,000X, greater than about 1,500X, greater than about 2,000X, greater than about 2,500X, greater than about 3,000X, greater than about 3,500X, greater than about 4,000X, greater than about 4,500X, greater than about 5,000X, greater than about 5,500X or greater than about 6,000X across more than about 99% of the arrayed genes (e.g., exons).

[0214] In certain embodiments, the sequences of the sets of target intervals described herein (e.g., encoding the target intervals), e.g., nucleotide sequences, are provided by the methods described herein. In certain embodiments, the sequences are provided without using methods that include a matching normal control (e.g., wild-type control), a matching tumor control (e.g., primary to metastatic) or both.

[0215] Gene Selection Groups or sets of target intervals for analysis are described herein, e.g., sub-genomic intervals, expression sub-genomic intervals or both, e.g., sets or groups of genes and other regions, sub-genomic intervals of sets or groups.

[0216] In some embodiments, the method includes, for example, identifying target regions from at least 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, or more genes or gene products from a nucleic acid sample obtained, for example, by next-generation sequencing methods, wherein the genes are selected from Tables 2A to 5B.

[0217] In some aspects, the method includes, for example, identifying target regions from at least 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, or more genes or gene products from a sample obtained, for example, by next-generation sequencing methods, wherein the genes are selected from Tables 2A to 5B.

[0218] In another embodiment, one target region of the following sets or groups is analyzed. For example, target regions associated with tumor or cancer genes or gene products and reference (e.g., wild-type) genes or gene products can provide a group or set of subgenomic regions from a sample.

[0219] In one embodiment, the method obtains reads, e.g., sequences, a set of target regions from a sample, and the target regions are selected from at least 1, 2, 3, 4, 5, 6, 7 or all of the following. A) At least 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, or more target regions, e.g., subgenomic regions from mutant or wild-type genes according to Tables 2A to 5B, or expression subgenomic regions, or both; B) At least 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, or more target intervals from genes or gene products related to tumors or cancers (e.g., which are positive or negative treatment response predictors, positive or negative prognostic factors, or enable differential diagnosis of tumors or cancers, such as genes according to Tables 2A - 5B); C) At least 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, or more target intervals from mutant or wild - type genes or gene products (e.g., single - nucleotide polymorphisms (SNPs)) of sub - genomic intervals present in genes selected from Tables 2A - 5B; D) At least 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, or more target intervals from mutant or wild - type genes (e.g., single - nucleotide polymorphisms (SNPs)) of target intervals present in genes selected from Tables 2A - 5B, wherein the target intervals are associated with one or more of: (i) better survival rates of cancer patients treated with a drug (e.g., better survival rates of breast cancer patients treated with paclitaxel); (ii) paclitaxel metabolism; (iii) toxicity to a drug; or (iv) side effects to a drug; E) Multiple translocation changes involving at least 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, or more genes or gene products according to Tables 2A - 5B; F) At least 5 genes selected from Tables 2A - 5B, wherein, for example, an allelic mutation at a certain position is related to the type of tumor and the allelic mutation is present in less than 5% of the cells of the said tumor type; G) At least 5 genes selected from Tables 2A - 5B embedded in GC - rich regions; or At least five genes that indicate genetic (e.g., germline risk) factors for developing cancer (e.g., where the gene or gene product is selected from Tables 2A - 5B).

[0220] In yet another embodiment, the method obtains reads, e.g., sequences, for a set of target intervals from a sample, where the target intervals are selected from 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, or all of the genes described in Tables 2A - 2C.

[0221] In yet another embodiment, the method obtains reads, e.g., sequences, for a set of target intervals from a sample, where the target intervals are selected from 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, or all of the genes described in Tables 3A - 3B.

[0222] In yet another embodiment, the method obtains reads, e.g., sequences, for a set of target intervals from a sample, where the target intervals are selected from 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, or all of the genes described in Tables 4A - 4C.

[0223] In yet another embodiment, the method obtains reads, e.g., sequences, for a set of target intervals from a sample, where the target intervals are selected from 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, or all of the genes described in Tables 5A - 5B.

[0224] The selected gene or gene product (also referred to herein as the "target gene or gene product") can include a target interval that includes an intronic region or an intergenic region. For example, the target interval can include an exon or intron or a fragment thereof, typically an exon sequence or a fragment thereof. The target interval can include a coding region or a non-coding region, such as a promoter, enhancer, 5' untranslated region (5'UTR) or 3' untranslated region (3'UTR), or a fragment thereof. In other embodiments, the target interval includes cDNA or a fragment thereof. In other embodiments, the target interval includes SNPs, as described herein for example.

[0225] In other embodiments, the target interval includes substantially all exons in the genome, such as one or more of the target intervals described herein (e.g., exons from a selected gene or gene product of interest (e.g., a gene or gene product associated with the cancerous phenotype described herein)). In one embodiment, the target interval includes somatic mutations, germline mutations, or both. In one embodiment, the target interval includes changes, such as point or single mutations, deletion mutations (e.g., in-frame deletions, intragenic deletions, complete gene deletions), insertion mutations (e.g., intragenic insertions), inversion mutations (e.g., intrachromosomal inversions), ligation mutations, ligation insertion mutations, inversion duplication mutations, tandem duplications (e.g., intrachromosomal tandem duplications), translocations (e.g., chromosomal translocations, non-reciprocal translocations), rearrangements, changes in gene copy number, or combinations thereof. In certain embodiments, the target interval constitutes less than 5%, 1%, 0.5%, 0.1%, 0.05%, 0.01%, 0.005%, or 0.001% of the coding region of the genome of tumor cells in the sample. In other aspects, the target interval is not involved in a disease, e.g., is not associated with the cancerous phenotype described herein.

[0226] In one embodiment, the target gene or gene product is a biomarker. As used herein, "biomarker" or "marker" is a gene, mRNA or protein that can be altered, and said alteration is associated with cancer. The alteration can be the amount, structure and / or activity in cancer tissue or cancer cells as compared to its amount, structure and / or activity in normal or healthy tissue or cells (e.g., control), and is associated with a disease state such as cancer. For example, a marker associated with cancer or predictive of responsiveness to anti-cancer treatment may have an altered nucleotide sequence, amino acid sequence, chromosomal translocation, intrachromosomal inversion, copy number, expression level, protein level, protein activity, epigenetic modification (e.g., methylation or acetylation state, or post-translational modification) in cancer tissue or cancer cells as compared to normal healthy tissue or cells. Further, a "marker" includes a molecule that is different from the wild-type sequence at the nucleotide or amino acid level, for example by substitution, deletion or insertion, when present in a tissue or cell whose structure has changed, e.g., is mutated (including mutations), and is associated with a disease state such as cancer.

[0227] In one embodiment, the target gene or gene product comprises a single nucleotide polymorphism (SNP). In another embodiment, the gene or gene product has a small deletion, such as a small intragenic deletion (e.g., in-frame or frameshift deletion). In yet another embodiment, the target sequence results from a deletion of an entire gene. In yet another embodiment, the target sequence has a small insertion, such as a small intragenic insertion. In one embodiment, the target sequence results from an inversion, such as an intrachromosomal inversion. In another embodiment, the target sequence results from an interchromosomal translocation. In yet another embodiment, the target sequence has a tandem duplication. In one embodiment, the target sequence has an undesirable feature (e.g., high GC content or repetitive element). In another embodiment, the target sequence has a portion of a nucleotide sequence that cannot itself be readily targeted, e.g., due to its repetitive nature. In one embodiment, the target sequence results from alternative splicing. In another embodiment, the target sequence is selected from a gene or gene product or a fragment thereof according to Tables 2A to 5B.

[0228] In one embodiment, the target gene or gene product or a fragment thereof is an antibody gene or gene product, an immunoglobulin superfamily receptor (e.g., B cell receptor (BCR) or T cell receptor (TCR)) gene or gene product, or a fragment thereof.

[0229] Human antibody molecules (and B cell receptors) are composed of heavy and light chains that have both constant (C) and variable (V) regions encoded by genes on at least the following three loci. 1. The immunoglobulin heavy chain locus (IGH@) on chromosome 14, which contains gene segments for immunoglobulin heavy chains; 2. The immunoglobulin kappa (κ) locus (IGK@) on chromosome 2, which contains gene segments for immunoglobulin light chains; 3. The immunoglobulin lambda (λ) locus (IGL@) on chromosome 22, which contains gene segments for immunoglobulin light chains.

[0230] Each heavy and light chain gene contains multiple copies of three different types of gene segments for the variable region of the antibody protein. For example, the immunoglobulin heavy chain region can contain one of five different classes, γ, δ, α, μ, and ε, 44 variable (V) gene segments, 27 diversity (D) gene segments, and 6 joining (J) gene segments. The light chain can also have a number of V and J gene segments, but does not have D gene segments. The lambda light chain has seven possible C regions and the kappa light chain has one.

[0231] The immunoglobulin heavy chain locus (IGH@) is a region on human chromosome 14 that contains genes for the heavy chains of human antibodies (or immunoglobulins). For example, the IGH locus contains IGHV (variable), IGHD (diversity), IGHJ (joining), and IGHC (constant) genes. Exemplary genes encoding immunoglobulin heavy chains include IGHV1-2, IGHV1-3, IGHV1-8, IGHV1-12, IGHV1-14, IGHV1-17, IGHV1-18, IGHV1-24, IGHV1-45, IGHV1-46, IGHV1-58, IGHV1-67, IGHV1-68, IGHV1-69, IGHV1-38-4, IGHV1-69-2, IGHV2-5, IGHV2-10, IGHV2-26, IGHV2-70, IGHV3-6, IGHV3-7, IGHV3-9, IGHV3-11, IGHV3-13, IGHV3-15, IGHV3-16, IGHV3-19, IGHV3-20, IGHV3-21, IGHV3-22, IGHV3-23, IGHV3-25, IGHV3-29, IGHV3-30, IGHV3-30-2, IGHV3-30-3, IGHV3-30-5, IGHV3-32, IGHV3-33, IGHV3-33-2, IGHV3-35, IGHV3-36, IGHV3-37, IGHV3-38, IGHV3-41, IGHV3-42, IGHV3-43, IGHV3-47, IGHV3-48, IGHV3-49, IGHV3-50, IGHV3-52, IGHV3-53, IGHV3-54, IGHV3-57, IGHV3-60, IGHV3-62, IGHV3-63, IGHV3-64, IGHV3-65, IGHV3-66, IGHV3-71, IGHV3-72, IGHV3-73, IGHV3-74, IGHV3-75, IGHV3-76, IGHV3-79, IGHV3-38-3, IGHV3-69-1, IGHV4-4, IGHV4-28, IGHV4-30-1, IGHV4-30-2, IGHV4-30-4, IGHV4-31, IGHV4-34, IGHV4-39, IGHV4-55, IGHV4-59, IGHV4-61, IGHV4-80, IGHV4-38-2, IGHV5-51, IGHV5-78, IGHV5-10-1, IGHV6-1, IGHV7-4-1, IGHV7-27, IGHV7-34-1,IGHV7-40, IGHV7-56, IGHV7-81, IGHVII-1-1, IGHVII-15-1, IGHVII-20-1, IGHVII-22-1, IGHVII-26-2, IGHVII-28-1, IGHVII-30-1, IGHVII-31-1, IGHVII-33-1, IGHVII-40-1, IGHVII-43-1, IGHVII-44-2, IGHVII-46-1, IGHVII-49-1, IGHVII-51-2, IGHVII-53-1, IGHVII-60-1, IGHVII-62-1, IGHVII-65-1, IGHVII-67-1, IGHVII-74-1, IGHVII-78-1, IGHVIII-2-1, IGHVIII-5-1, IGHVIII-5-2, IGHVIII-11-1, IGHVIII-13-1, IGHVIII-16-1, IGHVIII-22-2, IGHVIII-25-1, IGHVIII-26-1, IGHVIII-38-1, IGHVIII-44, IGHVIII-47-1, IGHVIII-51-1, IGHVIII-67-2, IGHVIII-67-3, IGHVIII-67-4, IGHVIII-76-1, IGHVIII-82, IGHVIV-44-1, IGHD1-1, IGHD1-7, IGHD1-14, IGHD1-20, IGHD1-26, IGHD2-2, IGHD2-8, IGHD2-15, IGHD2-21 IGHD3-3, IGHD3-9, IGHD3-10, IGHD3-16, IGHD3-22, IGHD4-4, IGHD4-11, IGHD4-17, IGHD4-23, IGHD5-5, IGHD5-12, IGHD5-18, IGHD5-24, IGHD6-6, IGHD6-13, IGHD6-19, IGHD6-25, IGHD7-27, IGHJ1, IGHJ1P, IGHJ2, IGHJ2P, IGHJ3, IGHJ3P, IGHJ4, IGHJ5, IGHJ6, IGHA1, IGHA2, IGHG1, IGHG2, IGHG3, IGHG4, IGHGP, IGHD, IGHE, IGHEP1, IGHM, and IGHV1-69D, are included.

[0232] The immunoglobulin kappa locus (IGK@) is a region on human chromosome 2 that contains genes for the kappa (κ) light chain of antibodies (or immunoglobulins). For example, the IGK locus contains IGKV (variable), IGKJ (joining), and IGKC (constant) genes. Exemplary genes encoding immunoglobulin kappa light chains include, but are not limited to, IGKV1-5, IGKV1-6, IGKV1-8, IGKV1-9, IGKV1-12, IGKV1-13, IGKV1-16, IGKV1-17, IGKV1-22, IGKV1-27, IGKV1-32, IGKV1-33, IGKV1-35, IGKV1-37, IGKV1-39, IGKV1D-8, IGKV1D-12, IGKV1D-13, IGKV1D-16, IGKV1D-17, IGKV1D-22, IGKV1D-27, IGKV1D-32, IGKV1D-33, IGKV1D-35, IGKV1D-37, IGKV1D-39, IGKV1D-42, IGKV1D-43, IGKV2-4, IGKV2-10, IGKV2-14, IGKV2-18, IGKV2-19, IGKV2-23, IGKV2-24, IGKV2-26, IGKV2-28, IGKV2-29, IGKV2-30, IGKV2-36, IGKV2-38, IGKV2-40, IGKV2D-10, IGKV2D-14, IGKV2D-18, IGKV2D-19, IGKV2D-23, IGKV2D-24, IGKV2D-26, IGKV2D-28, IGKV2D-29, IGKV2D-30, IGKV2D-36, IGKV2D-38, IGKV2D-40, IGKV3-7, IGKV3-11, IGKV3-15, IGKV3-20, IGKV3-25, IGKV3-31, IGKV3-34, IGKV3D-7, IGKV3D-11, IGKV3D-15, IGKV3D-20, IGKV3D-25, IGKV3D-31, IGKV3D-34, IGKV4-1, IGKV5-2, IGKV6-21, IGKV6D-21, IGKV6D-41, IGKV7-3, IGKJ1, IGKJ2, IGKJ3, IGKJ4, IGKJ5, and IGKC. The immunoglobulin lambda locus (IGL@) is a region on human chromosome 22 that contains genes for the lambda light chain of an antibody (or immunoglobulin). For example, the IGL locus contains IGLV (variable), IGLJ (joining), and IGLC (constant) genes. Exemplary genes encoding immunoglobulin lambda light chains include, but are not limited to, IGLV1-36, IGLV1-40, IGLV1-41, IGLV1-44, IGLV1-47, IGLV1-50, IGLV1-51, IGLV1-62, IGLV2-5, IGLV2-8, IGLV2-11, IGLV2-14, IGLV2-18, IGLV2-23, IGLV2-28, IGLV2-33, IGLV2-34, IGLV3-1, IGLV3-2, IGLV3-4, IGLV3-6, IGLV3-7, IGLV3-9, IGLV3-10, IGLV3-12, IGLV3-13, IGLV3-15, IGLV3-16, IGLV3-17, IGLV3-19, IGLV3-21, IGLV3-22, IGLV3-24, IGLV3-25, IGLV3-26, IGLV3-27, IGLV3-29, IGLV3-30, IGLV3-31, IGLV3-32, IGLV4-3, IGLV4-60, IGLV4-69, IGLV5-37, IGLV5-39, IGLV5-45, IGLV5-48, IGLV5-52, IGLV6-57, IGLV7-35, IGLV7-43, IGLV7-46, IGLV8-61, IGLV9-49, IGLV10-54, IGLV10-67, IGLV11-55, IGLVI-20, IGLVI-38, IGLVI-42, IGLVI-56, IGLVI-63, IGLVI-68, IGLVI-70, IGLIV-53, IGLVIV-59, IGLVIV-64, IGLVIV-65, IGLVIV-66-1, IGLVV-58, IGLVV-66, IGLVVI-22-1, IGLVVI-25-1, IGLVVII25-1, IGLVVII-41-1, IGLJ1, IGLJ2, IGLJ3, IGLJ4, IGLJ5, IGLJ6, IGLJ7, IGLC1, IGLC2, IGLC3, IGLC4, IGLC5, IGLC6, IGLC7.

[0233] The B cell receptor (BCR) consists of two parts: i) a membrane-bound immunoglobulin molecule of one isotype (e.g., IgD or IgM). Except for the presence of the endogenous membrane domain, these are in their secreted form and ii) a signaling moiety linked together by disulfide bridges: it can be identical to a heterodimer called Ig-α / Ig-β (CD79). Each nucleic acid molecule of the dimer spans the plasma membrane and has a cytoplasmic tail with an immunoreceptor activation tyrosine motif (ITAM).

[0234] The T cell receptor (TCR) consists of two different protein chains (i.e., a heterodimer). In 95% of T cells, this consists of an alpha (α) chain and a beta (β) chain, while in 5% of T cells, this consists of a gamma (γ) chain and a delta (δ) chain. This ratio can vary during ontogeny and in disease states. The T cell receptor genes are similar to immunoglobulin genes in that they also contain multiple V, D, and J gene segments that are rearranged during lymphocyte development to provide each cell with a unique antigen receptor in their beta and delta chains (and the V and J gene segments of their alpha and gamma chains).

[0235] The T cell receptor alpha locus (TRA) is a region on human chromosome 14 that contains genes for the TCR alpha chain. For example, the TRA locus includes, for example, TRAV (variable), TRAJ (joining), and TRAC (constant) genes. Exemplary genes encoding the T cell receptor alpha chain include, but are not limited to, TRAV1-1, TRAV1-2, TRAV2, TRAV3, TRAV4, TRAV5, TRAV6, TRAV7, TRAV8-1, TRAV8-2, TRAV8-3, TRAV8-4, TRAV8-5, TRAV8-6, TRAV8-7, TRAV9-1, TRAV9-2, TRAV10, TRAV11, TRAV12-1, TRAV12-2, TRAV12-3, TRAV13-1, TRAV13-2, TRAV14DV4, TRAV15, TRAV16, TRAV17, TRAV18, TRAV19, TRAV20, TRAV21, TRAV22, TRAV23DV6, TRAV24, TRAV25, TRAV26-1, TRAV26-2, TRAV27, TRAV28, TRAV29DV5, TRAV30, TRAV31, TRAV32, TRAV33, TRAV34, TRAV35, TRAV36DV7, TRAV37, TRAV38-1, TRAV38-2DV8, TRAV39, TRAV40, TRAV41, TRAJ1, TRAJ2, TRAJ3, TRAJ4, TRAJ5, TRAJ6, TRAJ7, TRAJ8, TRAJ9, TRAJ10, TRAJ11, TRAJ12, TRAJ13, TRAJ14, TRAJ15, TRAJ16, TRAJ17, TRAJ18, TRAJ19, TRAJ20, TRAJ21, TRAJ22, TRAJ23, TRAJ24, TRAJ25, TRAJ26, TRAJ27, TRAJ28, TRAJ29, TRAJ30, TRAJ31, TRAJ32, TRAJ33, TRAJ34, TRAJ35, TRAJ36, TRAJ37, TRAJ38, TRAJ39, TRAJ40, TRAJ41, TRAJ42, TRAJ43, TRAJ44, TRAJ45, TRAJ46, TRAJ47, TRAJ48, TRAJ49, TRAJ50, TRAJ51, TRAJ52, TRAJ53, TRAJ54, TRAJ55, TRAJ56, TRAJ57, TRAJ58, TRAJ59, TRAJ60, TRAJ61, and TRAC.

[0236] The T cell receptor beta locus (TRB) is a region on human chromosome 7 that contains genes for the TCR beta chain. For example, the TRB locus includes, for example, TRBV (variable), TRBD (diversity), TRBJ (joining), and TRBC (constant) genes. Exemplary genes encoding the T cell receptor beta chain include, but are not limited to, TRBV1, TRBV2, TRBV3-1, TRBV3-2, TRBV4-1, TRBV4-2, TRBV4-3, TRBV5-1, TRBV5-2, TRBV5-3, TRBV5-4, TRBV5-5, TRBV5-6, TRBV5-7, TRBV6-2, TRBV6-3, TRBV6-4, TRBV6-5, TRBV6-6, TRBV6-7, TRBV6-8, TRBV6-9, TRBV7-1, TRBV7-2, TRBV7-3, TRBV7-4, TRBV7-5, TRBV7-6, TRBV7-7, TRBV7-8, TRBV7-9, TRBV8-1, TRBV8-2, TRBV9, TRBV10-1, TRBV10-2, TRBV10-3, TRBV11-1, TRBV11-2, TRBV11-3, TRBV12-1, TRBV12-2, TRBV12-3, TRBV12-4, TRBV12-5, TRBV13, TRBV14, TRBV15, TRBV16, TRBV17, TRBV18, TRBV19, TRBV20-1, TRBV21-1, TRBV22-1, TRBV23-1, TRBV24-1, TRBV25-1, TRBV26, TRBV27, TRBV28, TRBV29-1, TRBV30, TRBVA, TRBVB, TRBVB5-8, TRBV6-1, TRBD1, TRBD2, TRBJ1-1, TRBJ1-2, TRBJ1-3, TRBJ1-4, TRBJ1-5, TRBJ1-6, TRBJ2-1, TRBJ2-2, TRBJ2-2P, TRBJ2-3, TRBJ2-4, TRBJ2-5, TRBJ2-6, TRBJ2-7, TRBC1, TRBC2.

[0237] The T cell receptor delta locus (TRD) is a region on human chromosome 14 that contains genes for the TCR delta chain. For example, the TRD locus includes, for example, TRDV (variable), TRDJ (joining), and TRDC (constant) genes. Exemplary genes encoding the T cell receptor delta chain include, but are not limited to, TRDV1, TRDV2, TRDV3, TRDD1, TRDD2, TRDD3, TRDJ1, TRDJ2, TRDJ3, TRDJ4, and TRDC.

[0238] The T cell receptor gamma locus (TRG) is a region on human chromosome 7 that contains genes for the TCR gamma chain. For example, the TRG locus includes, for example, TRGV (variable), TRGJ (joining), and TRGC (constant) genes. Exemplary genes encoding the T cell receptor gamma chain include, but are not limited to, TRGV1, TRGV2, TRGV3, TRGV4, TRGV5, TRGV5 P, TRGV6, TRGV7, TRGV8, TRGV9, TRGV10, TRGV11, TRGVA, TRGVB, TRGJ1, TRGJ2, TRGJP, TRGJP1, TRGJP2, TRGC1, and TRGC2.

[0239] In one embodiment, the target gene or gene product or fragment thereof is selected from any of the genes or gene products described in Tables 2A - 5B. [Table 2A] JPEG0007702360000081.jpg240170JPEG0007702360000082.jpg241170 [Table 2B] [Table 2C] JPEG0007702360000085.jpg240170 [Table 3A] JPEG0007702360000087.jpg236170JPEG0007702360000088.jpg237170

Table 3B

Table 4A

Table 4B

Table 4C

Table 5A

Table 5B

[0240] Further exemplary genes are described, for example, in Tables 1-11 of International Application Publication No. WO2012 / 092426, the contents of which are incorporated by reference in their entirety.

[0241] The uses of the foregoing methods include, but are not limited to, the use of a library of oligonucleotides comprising all known sequence variants (or subsets thereof) of one or more specific genes for sequence identification in a medical specimen.

[0242] Type of change The methods described herein can be used in combination with, or as part of, methods for assessing genomic changes described herein.

[0243] Various types of changes (e.g., somatic changes) can be evaluated and used for the analysis of genomic changes. For example, genomic changes associated with cancer and / or tumor mutation burden can be analyzed. In some embodiments, the methods described herein are useful for analyzing samples with low tumor content and / or low amounts of tumor nucleic acids.

[0244] Somatic changes In certain embodiments, the changes evaluated according to the methods described herein are somatic changes.

[0245] In certain embodiments, the modification (e.g., somatic change) is a short coding variant, such as a base substitution or indel (insertion or deletion). In certain embodiments, the change (e.g., somatic change) is a point mutation. In other embodiments, the change (e.g., somatic change) is other than a rearrangement, such as other than a translocation. In certain embodiments, the change (e.g., somatic change) is a splice variant.

[0246] In certain embodiments, the change (e.g., somatic change) is a silent mutation, such as a synonymous change. In other embodiments, the change (e.g., somatic change) is a non-synonymous single nucleotide variant (SNV). In other embodiments, the modification (e.g., somatic change) is a passenger mutation, e.g., a modification having no detectable effect on the fitness of a clone of cells. In certain embodiments, the change (e.g., somatic change) is a variant of unknown significance (VUS), e.g., a change for which pathogenicity cannot be confirmed or excluded. In certain embodiments, the change (e.g., somatic change) is not identified as being associated with a cancer phenotype.

[0247] In certain embodiments, the change (e.g., somatic change) is not associated with, or is not known to be associated with, an effect on cell division, growth or survival. In other embodiments, the change (e.g., somatic change) is associated with an effect on cell division, growth or survival.

[0248] In certain embodiments, an increase in the level of somatic changes is an increase in the level of one or more classes or types of somatic changes (e.g., rearrangements, point mutations, indels, or any combination thereof). In certain embodiments, an increase in the level of somatic changes is an increase in the level of one class or type of somatic change (e.g., rearrangements only, point mutations only, or indels only). In certain embodiments, an increase in the level of somatic changes is an increase in the level of somatic changes at a position (e.g., nucleotide position, e.g., one or more nucleotide positions) or in a region (e.g., in a nucleotide region, e.g., in one or more nucleotide regions). In certain embodiments, an increase in the level of somatic changes is an increase in the level of somatic changes (e.g., somatic changes as described herein).

[0249] Functional alteration In certain embodiments, the change (e.g., somatic change) is a functional change in a subgenomic interval. In other embodiments, the change (e.g., somatic change) is not a known functional change in a subgenomic interval. For example, when assessing tumor mutational burden, the number of changes (e.g., somatic changes) can exclude one or more functional changes.

[0250] In some embodiments, the functional change is a change that affects cell division, growth or survival, such as a change that promotes cell division, growth or survival, as compared to a reference sequence, such as a wild-type or non-mutated sequence. In certain embodiments, the functional change is so identified by inclusion in a database of functional changes, such as the COSMIC database (cancer.sanger.ac.uk / cosmic; Forbes et al., Nucl. Acids Res. 2015;43(D1):D805-D811). In other embodiments, the functional change is a change that has a known functional state, such as a change that occurs as a known somatic change in the COSMIC database. In certain embodiments, the functional change is a change that has a potential functional state, such as the disruption of a tumor suppressor gene. In certain embodiments, the functional change is a driver mutation, such as a change that confers a selective advantage to clones in its microenvironment, for example by increasing the survival or proliferation of cells. In other embodiments, the functional change is a change that can cause clonal expansion. In certain embodiments, the functional change is a change that can cause one, two, three, four, five, or all of the following: (a) self-sufficiency in growth signals; (b) reduced responsiveness to growth inhibitory signals, e.g., insensitivity; (c) decreased apoptosis; (d) increased copy potential; (e) sustained angiogenesis; or (f) tissue invasion or metastasis.

[0251] In certain embodiments, the functional change is not a passenger mutation, e.g., a change that has no detectable effect on the fitness of a clone of cells. In certain embodiments, the functional change is not a variant of unknown significance (VUS), e.g., a change for which pathogenicity cannot be confirmed or excluded.

[0252] In certain embodiments, multiple (e.g., about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or more) functional changes in the genes described in Tables 2A - 5B are excluded. In certain embodiments, all functional changes in the genes described in Tables 2A - 5B are excluded. In certain embodiments, multiple functional changes in multiple genes described in Tables 2A - 5B are excluded. In certain embodiments, all functional changes in all genes described in Tables 2A - 5B are excluded.

[0253] Germline changes In certain embodiments, the modification is a germline modification. In other embodiments, the modification is not a germline modification. In certain embodiments, the modification is not the same as or similar to, e.g., distinguishable from, a germline modification. For example, when evaluating tumor mutational burden, the number of changes can exclude the number of germline changes.

[0254] In certain embodiments, germline changes are single nucleotide polymorphisms (SNPs), base substitutions, indels (e.g., insertions or deletions), or silent changes (e.g., synonymous changes).

[0255] In certain embodiments, germline variations are identified by use of methods that do not use comparison to a matched normal sequence. In other embodiments, germline variations are identified by methods that include use of the SGZ algorithm. In certain embodiments, germline variations are so identified by inclusion in a database of germline variations, such as the dbSNP database (www.ncbi.nlm.nih.gov / SNP / index.html; Sherry et al., Nucleic Acids Res. 2001;29(1):308-311). In other embodiments, germline variations are so identified by inclusion in two or more counts in the ExAC database (exac.broadinstitute.org; Exome Aggregation Consortium et al., “Analysis of protein-coding genetic variation in 60,706 humans,” bioRxiv preprint. October 30, 2015). In some embodiments, germline variations are identified by inclusion in the 1000 Genomes Project database (www.1000genomes.org; McVean et al., Nature. 2012;491, 56-65). In some embodiments, germline variations are identified by inclusion in the ESP database (Exome Variant Server, NHLBI GO Exome Sequencing Project (ESP), Seattle, WA (evs.gs.washington.edu / EVS / )).

[0256] Sample The methods described herein can be used to evaluate tumor fractions in various types of samples from several different sources.

[0257] In some embodiments, the sample comprises nucleic acids, such as DNA, RNA, or both. In certain embodiments, the sample comprises one or more nucleic acids derived from a tumor. In certain aspects, the sample further comprises one or more non-nucleic acid components derived from a tumor, such as cells, proteins, carbohydrates, or lipids. In certain embodiments, the sample further comprises one or more nucleic acids from non-tumor cells or tissues.

[0258] In certain aspects, the sample is obtained from a liquid biopsy. In certain aspects, the sample is not obtained from a tissue biopsy. In certain embodiments, the sample is a liquid sample. In certain embodiments, the sample does not contain or essentially does not contain solids.

[0259] In certain embodiments, the sample is obtained from a subject having a solid tumor, a blood cancer, or a metastatic form thereof. In certain embodiments, the sample is obtained from a subject having cancer or at risk of having cancer. In certain embodiments, the sample is obtained from a subject who has not received, is receiving, or has received treatment for treating cancer as described herein.

[0260] In some aspects, the sample comprises one or more nucleic acids, such as DNA, RNA, or both, from pre-malignant or malignant cells, cells from solid tumors, soft tissue tumors or metastatic lesions, cells from blood cancers, histologically normal cells, circulating tumor cells (CTCs), or combinations thereof. In some aspects, the sample comprises one or more cells selected from pre-malignant or malignant cells, solid tumors, soft tissue tumors or metastatic lesions, blood cancers, histologically normal cells, circulating tumor cells (CTCs), or combinations thereof.

[0261] In certain embodiments, the sample comprises cell-free DNA (cfDNA). In certain embodiments, the sample comprises circulating tumor DNA (ctDNA). In certain particular embodiments, the sample comprises blood, serum, or plasma. In certain embodiments, the sample comprises cerebrospinal fluid (CSF). In certain embodiments, the sample comprises pleural effusion. In certain embodiments, the sample comprises ascites. In certain embodiments, the sample comprises urine. In certain aspects, the sample comprises an excision, a needle biopsy, a fine needle aspirate, or a cytology smear. In certain embodiments, the sample is a formalin-fixed paraffin-embedded (FFPE) sample.

[0262] A variety of tissues can be sources of the samples used in the methods. Genomic or sub-genomic nucleic acids (e.g., DNA or RNA) can be isolated from a sample of interest (e.g., a sample containing tumor cells, a blood sample, a blood component sample, a sample containing cell-free DNA (cfDNA), a sample containing circulating tumor DNA (ctDNA), a sample containing circulating tumor cells (CTC), or any normal control (e.g., normal adjacent tissue (NAT))).

[0263] In some embodiments, the sample comprises nucleic acids, e.g., nucleic acids derived from a tumor, e.g., DNA, RNA, or both. The nucleic acid can be DNA or RNA. In certain aspects, the sample further comprises non-nucleic acid components derived from a tumor, e.g., cells, proteins, carbohydrates, or lipids. In certain embodiments, the sample further comprises nucleic acids from normal cells or tissues.

[0264] In certain embodiments, the sample is stored as a frozen sample or as a formaldehyde or paraformaldehyde-fixed paraffin-embedded (FFPE) tissue preparation. For example, the sample can be embedded in a matrix, such as an FFPE block or a frozen sample. In certain embodiments, the sample is a blood sample. In certain embodiments, the tissue sample is a blood component sample. In certain embodiments, the sample is a cfDNA sample. In certain embodiments, the sample is a ctDNA sample. In certain embodiments, the sample is a CTC sample. In other embodiments, the tissue sample is a bone marrow aspirate (BMA) sample. The isolation step can include flow sorting of individual chromosomes. And / or includes microdissecting the sample of the subject (e.g., the samples described herein).

[0265] In other embodiments, the sample contains one or more pre-malignant or malignant cells. In certain aspects, the sample is obtained from a solid tumor, a soft tissue tumor, or a metastatic lesion. In certain embodiments, the sample is obtained from a hematologic malignancy or pre-malignant tumor. In other embodiments, the sample contains tissue or cells from a surgical margin. In certain embodiments, the sample contains tumor-infiltrating lymphocytes. The sample can be histologically normal tissue. In one embodiment, the sample contains one or more non-malignant cells.

[0266] In certain embodiments, the FFPE sample has one, two, or all of the following characteristics. (a) A surface area of about 10 mm 2 or more, about 25 mm 2 or more, or about 50 mm 2 or more; (b) about 0.1 mm 3 or more, about 0.2 mm 3 or more, about 0.3 mm 3 or more, about 0.4 mm 3 or more, about 0.5 mm 3 or more, about 0.6 mm 3 or more, about 0.7 mm 3 or more, about 0.8 mm 3 or more, about 0.9 mm 3 or more, about 1 mm 3 or more, about 2 mm 3 or more, about 3 mm3 Greater than or equal to about 4 mm 3 Greater than or equal to about 5 mm 3 Having a sample volume of greater than or equal to; (c) having a cellularity of greater than or equal to about 50%, greater than or equal to about 60%, greater than or equal to about 70%, greater than or equal to about 80%, or greater than or equal to about 90%; and / or (d) having a number of nucleated cells of greater than or equal to about 10,000 cells, greater than or equal to about 20,000 cells, greater than or equal to about 30,000 cells, greater than or equal to about 40,000 cells, or greater than or equal to about 50,000 cells.

[0267] In one embodiment, the method further comprises obtaining a sample, such as a sample described herein. The sample can be obtained directly or indirectly. In one embodiment, the sample is obtained, for example, by isolation or purification from a sample containing cfDNA. In one embodiment, the sample is obtained, for example, by isolation or purification from a sample containing ctDNA. In one embodiment, the sample is obtained, for example, by isolation or purification from a sample containing both malignant and non-malignant cells (e.g., tumor infiltrating lymphocytes). In one embodiment, the sample is obtained, for example, by isolation or purification from a sample containing CTCs.

[0268] In other embodiments, the method comprises evaluating a sample, such as a histologically normal sample, from a surgical margin, for example, using the methods described herein. In some embodiments, a sample obtained from histologically normal tissue (e.g., a histologically normal tissue margin otherwise) may still have the changes described herein. Thus, the method may further comprise reclassifying the sample based on the presence of the detected changes. In one embodiment, for example, multiple samples from different subjects are processed simultaneously.

[0269] In one embodiment, the method comprises isolating nucleic acid from the sample to provide an isolated nucleic acid sample. In one embodiment, the method comprises isolating nucleic acid from a control to provide an isolated control nucleic acid sample. In one embodiment, the method further comprises rejecting samples that do not contain detectable nucleic acids.

[0270] In one embodiment, the method further includes determining whether a primary control is available and, if available, isolating a control nucleic acid (e.g., DNA) from the primary control. In one embodiment, the method further includes determining whether a NAT is present in the sample (e.g., when no primary control sample is available). In one embodiment, the method further includes obtaining a sub-sample enriched in non-tumor cells by macrodissecting non-tumor tissue from the NAT in a sample without a primary control, for example. In one embodiment, the method further includes determining that no primary control and NAT are available and marking the sample for analysis without a matched control.

[0271] In one embodiment, the method further includes obtaining a value of the nucleic acid yield in the sample and comparing the obtained value with a reference standard. For example, if the obtained value is less than the reference standard, the method further includes amplifying the nucleic acid before library construction. In one embodiment, the method further includes obtaining a value of the size of nucleic acid fragments in the sample and comparing the obtained value with a reference standard, for example, a size of at least 300, 600, or 900 bps, such as an average size. The parameters described herein can be adjusted or selected according to this particularity.

[0272] In certain embodiments, the method includes isolating nucleic acids from an aged sample, such as an aged FFPE sample. The aged sample can be, for example, 1 year old, 2 years old, 3 years old, 4 years old, 5 years old, 10 years old, 15 years old, 20 years old, 25 years old, 50 years old, 75 years old, or 100 years old or older.

[0273] Nucleic acids can be obtained from samples of various sizes. For example, nucleic acids can be isolated from samples of 5 to 200 μm or more. For example, the sample can be measured at 5 μm, 10 μm, 20 μm, 30 μm, 40 μm, 50 μm, 70 μm, 100 μm, 110 μm, 120 μm, 150 μm, or 200 μm or more.

[0274] Protocols for DNA isolation from samples are known in the art, for example, as provided in Example 1 of International Patent Application Publication No. WO2012 / 092426. Further methods for isolating nucleic acids (e.g., DNA) from formaldehyde- or paraformaldehyde-fixed, paraffin-embedded (FFPE) tissues are disclosed, for example, in Cronin M. et al., (2004) Am J Pathol. 164(1):35-42; Masuda N. et al., (1999) Nucleic Acids Res. 27(22):4436-4443; Specht K. et al., (2001) Am J Pathol. 158(2):419-429, Ambion RecoverAll™ Total Nucleic Acid Isolation Protocol (Ambion, Catalog No. AM1975, September 2008), Maxwell® 16 FFPE Plus LEV DNA Purification Kit Technical Manual (Promega Literature#TM349, February 2011), E.Z.N.A.® FFPE DNA Kit Handbook (OMEGA bio-tek, Norcross, GA, Product Nos. D3399-00, D3399-01, and D3399-02; June 2009) and QIAamp® DNA FFPE Tissue Handbook (Qiagen, Catalog No. 37625, October 2007). The RecoverAll™ Total Nucleic Acid Isolation Kit solubilizes paraffin-embedded samples using xylene at high temperature and captures nucleic acids by passing them over a glass fiber filter. The Maxwell® 16 FFPE Plus LEV DNA Purification Kit is used with the Maxwell® 16 Instrument to purify genomic DNA from 1- to 10-μm sections of FFPE tissue. DNA is purified using silica-coated paramagnetic particles (PMP) and eluted in a low elution volume. The E.Z.N.A.® FFPE DNA Kit uses spin columns and buffer systems for the isolation of genomic DNA.The QIAamp® DNA FFPE Tissue Kit uses QIAamp® DNA Micro technology for the purification of genomic and mitochondrial DNA. Protocols for DNA isolation from blood are disclosed, for example, in the Maxwell® 16 LEV Blood DNA Kit and Maxwell 16 Buccal Swab LEV DNA Purification Kit Technical Manual (Promega Literature #TM333, January 1, 2011).

[0275] Protocols for RNA isolation are disclosed, for example, in the Maxwell® 16 Total RNA Purification Kit Technical Bulletin (Promega Literature #TB351, August 2009).

[0276] The isolated nucleic acid (e.g., genomic DNA) can be fragmented or sheared by performing routine techniques. For example, genomic DNA can be fragmented by physical shearing methods, enzymatic cleavage methods, chemical cleavage methods, and other methods well known to those skilled in the art. The nucleic acid library can contain all or substantially all of the genomic complexity. The term "substantially all" in this context actually refers to the possibility that there may be some undesirable loss of genomic complexity during the initial steps of the procedure. The methods described herein are also useful when the nucleic acid library is a part of the genome, for example, when the genomic complexity is reduced by design. In some embodiments, any selected portion of the genome can be used with the methods described herein. In certain embodiments, the entire exome or a subset thereof is isolated.

[0277] In certain embodiments, the method further comprises isolating nucleic acids from a sample to provide a library (e.g., a nucleic acid library as described herein). In certain embodiments, the sample comprises a whole genome, sub-genomic fragments, or both. The isolated nucleic acids can be used to prepare a nucleic acid library. Protocols for isolating and preparing libraries from whole genomes or sub-genomic fragments are known in the art (e.g., Illumina's genomic DNA sample preparation kits). In certain embodiments, genomic or sub-genomic DNA fragments are isolated from a subject sample (e.g., a sample as described herein). In one embodiment, the sample is a preserved sample, such as a matrix, such as a FFPE block or a sample embedded in a frozen sample. In certain embodiments, the isolation step includes flow sorting individual chromosomes and / or microdissecting the sample. In certain embodiments, the amount of nucleic acid used to generate the nucleic acid library is less than 5 micrograms, less than 1 microgram, or less than 500 ng, less than 200 ng, less than 100 ng, less than 50 ng, less than 10 ng, less than 5 ng, or less than 1 ng.

[0278] In still other embodiments, the nucleic acid used to generate the library comprises RNA or cDNA derived from RNA. In some aspects, the RNA comprises total cellular RNA. In other embodiments, certain abundant RNA sequences (e.g., ribosomal RNA) are depleted. In some embodiments, poly(A)-tailed mRNA fragments in the total RNA preparation are enriched. In some embodiments, the cDNA is made by random-prime cDNA synthesis. In other embodiments, cDNA synthesis is initiated at the poly(A) tail of mature mRNA by priming with an oligo(dT)-containing oligonucleotide. Methods for depletion, poly(A) enrichment, and cDNA synthesis are well known to those of skill in the art.

[0279] In other embodiments, the nucleic acids are fragmented or sheared by physical or enzymatic methods, optionally ligated to synthetic adapters, size selected (e.g., by preparative gel electrophoresis), and amplified (e.g., by PCR). For example, alternative methods for DNA shearing are known in the art, as described in Example 4 of International Patent Application Publication No. 2012 / 092426. For example, alternative DNA shearing methods may be more automatable and / or more efficient (e.g., for degraded FFPE samples). Alternative methods to DNA shearing methods can also be used to avoid the ligation step during library preparation.

[0280] In other embodiments, the isolated DNA (e.g., genomic DNA) is fragmented or sheared. In some embodiments, the library comprises less than 50% genomic DNA, e.g., a fractional amount of genomic DNA that is a reduced representation or defined portion of the genome that has been fractionated by other means. In other embodiments, the library comprises all or substantially all of the genomic DNA.

[0281] In other embodiments, the fragmented and adapter-ligated nucleic acids are used without explicit size selection or amplification prior to hybrid selection. In some embodiments, the nucleic acids are amplified by specific or non-specific nucleic acid amplification methods well known to those of skill in the art. In some embodiments, the nucleic acids are amplified by whole genome amplification methods such as, for example, random primer strand displacement amplification.

[0282] The methods described herein can be carried out using small amounts of nucleic acid, for example, when the amount of source DNA or RNA is limiting (e.g., even after whole genome amplification). In one embodiment, the nucleic acid comprises a nucleic acid sample of about 5 μg, 4 μg, 3 μg, 2 μg, 1 μg, 0.8 μg, 0.7 μg, 0.6 μg, 0.5 μg or 400 ng, 300 ng, 200 ng, 100 ng, 50 ng, 10 ng, 5 ng, 1 ng or less. For example, one can typically start with 50-100 ng of genomic DNA. However, if genomic DNA (e.g., using PCR) is amplified prior to the hybridization step, for example solution hybridization, one can start with even less. Thus, it is possible, but not essential, to amplify genomic DNA prior to hybridization, for example solution hybridization.

[0283] In one embodiment, the sample comprises DNA, RNA (or cDNA derived from RNA), or both, from non-cancerous or non-malignant cells, such as tumor infiltrating lymphocytes. In one embodiment, the sample comprises DNA, RNA (or cDNA derived from RNA), or both, from non-cancerous or non-malignant cells, such as tumor infiltrating lymphocytes, and does not contain, or essentially does not contain, DNA, RNA (or cDNA derived from RNA), or both, from cancerous or malignant cells.

[0284] In one embodiment, the sample comprises DNA, RNA (or cDNA derived from RNA) from cancerous or malignant cells. In one embodiment, the sample comprises DNA, RNA (or cDNA derived from RNA) from cancerous or malignant cells and does not contain, or essentially does not contain, DNA, RNA (or cDNA derived from RNA), or both, from non-cancerous or non-malignant cells, such as tumor infiltrating lymphocytes.

[0285] In one embodiment, the sample includes DNA, RNA (or cDNA derived from RNA), or both from non-cancerous or non-malignant cells, such as tumor infiltrating lymphocytes, and DNA, RNA (or cDNA derived from RNA), or both from cancerous or malignant cells.

[0286] In certain embodiments, the sample is obtained from a subject having cancer. Exemplary cancers include, but are not limited to, B cell cancers such as multiple myeloma, melanoma, breast cancer, lung cancer (such as non-small cell lung cancer or NSCLC), bronchial cancer, colorectal cancer, prostate cancer, pancreatic cancer, gastric cancer, ovarian cancer, bladder cancer, brain or central nervous system cancer, peripheral nervous system cancer, esophageal cancer, cervical cancer, uterine or endometrial cancer, oral or pharyngeal cancer, liver cancer, kidney cancer, testicular cancer, biliary tract cancer, small intestine or accessory organ cancer, salivary gland cancer, thyroid cancer, adrenal gland cancer, osteosarcoma, chondrosarcoma, cancers of the blood tissue, adenocarcinoma, inflammatory myofibroblastic tumor, gastrointestinal stromal tumor (GIST), colon cancer, multiple myeloma (MM), myelodysplastic syndrome (MDS), myeloproliferative disorder (MPD), acute lymphoblastic leukemia (ALL), acute myeloblastic leukemia (AML), chronic myeloblastic leukemia (CML), chronic lymphocytic leukemia (CLL), polycythemia vera, Hodgkin lymphoma, non-Hodgkin lymphoma (NHL), soft tissue sarcoma, fibrosarcoma, myxosarcoma, liposarcoma, osteogenic sarcoma, chordoma, angiosarcoma, endotheliosarcoma synovioma, mesothelioma, Ewing tumor, leiomyosarcoma, rhabdomyosarcoma, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinoma, medullary carcinoma, bronchiogenic carcinoma, renal cell carcinoma, hepatoma, cholangiocarcinoma, choriocarcinoma, seminoma, embryonal carcinoma, Wilms tumor, bladder cancer, epithelial carcinoma, glioma, astrocytoma, medulloblastoma, craniopharyngioma, ependymoma, pinealoma, hemangioblastoma, acoustic neuroma, oligodendroglioma, meningioma, neuroblastoma, retinoblastoma, cell lymphoma, mantle cell lymphoma, hepatocellular carcinoma a, thyroid cancer, gastric cancer, head and neck cancer, small cell carcinoma, essential thrombocythemia, agnogenic myeloid metaplasia, hypereosinophilic syndrome, systemic mastocytosis, the familiar hypereosinophilia, chronic eosinophilic leukemia, neuroendocrine cancer, carcinoid tumor, and the like.

[0287] In one embodiment, the cancer is a hematologic malignancy (or pre-malignancy). As used herein, a hematologic malignancy refers to a tumor of hematopoietic or lymphoid tissue, such as a tumor affecting the blood, bone marrow, or lymph nodes. Exemplary hematologic malignancies include leukemia (e.g., acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic leukemia (CLL), chronic myeloid leukemia (CML), hairy cell leukemia, acute monocytic leukemia (AMoL), chronic myelomonocytic leukemia (CMML), juvenile myelomonocytic leukemia (JMML), or large granular lymphocytic leukemia), lymphoma (e.g., AIDS-related lymphoma, cutaneous T-cell lymphoma, Hodgkin lymphoma (e.g., classical Hodgkin lymphoma or nodular lymphocyte-predominant Hodgkin lymphoma), mycosis fungoides, non-Hodgkin lymphoma (e.g., B-cell non-Hodgkin lymphoma (e.g., Burkitt lymphoma, small lymphocytic lymphoma (CLL / SLL), diffuse large B-cell lymphoma, follicular lymphoma, immunoblastic large cell lymphoma, precursor B-lymphoblastic lymphoma, or mantle cell lymphoma) or T-cell non-Hodgkin lymphoma (mycosis fungoides, anaplastic large cell lymphoma, or precursor T-lymphoblastic lymphoma)), primary central nervous system, including but not limited to these. As used herein, pre-malignant refers to tissue that is not yet malignant but is ready to become malignant.

[0288] In some embodiments, the sample described herein is also referred to as a specimen. In some aspects, the sample is a tissue sample, a blood sample, or a bone marrow sample.

[0289] In some embodiments, the blood sample contains cell-free DNA (cfDNA). In some embodiments, the cfDNA contains DNA from healthy tissue, such as non-diseased cells, or tumor tissue, such as tumor cells. In some embodiments, the cfDNA from tumor tissue contains circulating tumor DNA (ctDNA). In some embodiments, the ctDNA sample is obtained, e.g., collected, from a patient having a solid tumor, such as lung cancer, breast cancer, or colon cancer.

[0290] In some embodiments, the sample, such as a specimen, is a formalin-fixed paraffin-embedded (FFPE) specimen. In some aspects, the FPPE specimen includes, but is not limited to, specimens selected from core needle biopsies, fine needle aspirates, or exfoliative cytologies. In some aspects, the sample includes an FPPE block and one original hematoxylin and eosin (H&E) stained slide. In some embodiments, the sample includes unstained slides (e.g., positively charged unburned slides with a thickness of 4-5 microns; e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more such slides) and one or more H&E stained slides.

[0291] In some embodiments, the sample includes an FPPE block or unstained slides, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 or more unstained slides and one or more H&E slides. In some embodiments, the sample includes tissue that has been formalin-fixed and embedded in a paraffin block, for example, using standard fixation methods as described herein.

[0292] In some embodiments, the sample has a surface area of at least 1-30 mm 2 , for example, about 5-25 mm 2 . In some embodiments, the sample has a surface area of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 mm 2 , for example, 5 mm 2 . In some embodiments, the sample has a surface area of at least 5 mm 2 . In some embodiments, the sample has a surface area of about 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30 mm 2 , for example, 25 mm 2 . In some embodiments, the sample has a surface area of 25 mm 2 .

[0293] In some embodiments, the sample has a thickness of at least 1-5 mm 3 , for example, about 2 mm3 includes the surface volume thereof. In some embodiments, about 2 mm 3 of the surface volume is about 80 microns, for example at least or exceeding 80 microns in depth and about 25 mm 2 of the surface area of the sample.

[0294] In some embodiments, the sample includes tumor contents including, for example, tumor nuclei. In some embodiments, the sample includes a tumor content having at least 5-50%, 10-40%, 15-25%, or 20-30% tumor nuclei. In some embodiments, the sample includes a tumor content of at least 20% tumor nuclei. In some embodiments, the sample includes a tumor content of about 30% tumor nuclei. In some aspects, the percentage of tumor nuclei is determined, for example calculated, by dividing the number of tumor cells by the total number of all cells having nuclei. In some embodiments, a higher tumor content may be required if the sample is, for example, a liver sample including hepatocytes. In some embodiments, hepatocytes have nuclei that are, for example, twice the DNA content of other, for example non-hepatocyte somatic nuclei, for example twice as many nuclei. In some aspects, the sensitivity of detection of a change (e.g., a change as described herein) depends on the tumor content of the sample, for example, a lower tumor content may result in a lower detection sensitivity.

[0295] In some embodiments, DNA is extracted from nucleated cells from the sample. In some embodiments, the sample has low nucleated cellity, for example, if the sample consists mainly of red blood cells, diseased cells containing excessive cytoplasm, or tissue having fibrosis. In some embodiments, a sample with low nucleated cellity may require more, for example a larger tissue volume, for example 2 mm 3 of tissue volume or more for DNA extraction.

[0296] In some embodiments, the FFPE sample, such as a specimen, is prepared using standard fixation methods to preserve the integrity of the nucleic acids. In some embodiments, the standard fixation methods include using 10% neutral buffered formalin for, for example, 6 to 72 hours. In some embodiments, the methods do not include fixatives such as Dutch Bouin, B5, AZF, etc. In some embodiments, the method does not include decalcification. In some embodiments, the method includes decalcification. In some embodiments, decalcification is performed using EDTA. In some embodiments, strong acids, such as hydrochloric acid, sulfuric acid, or picric acid, are not used for decalcification.

[0297] In some aspects, the sample includes FFPE blocks or unstained slides, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 or more unstained slides and 1 or more H&E slides. In some embodiments, the sample includes tissue that has been formalin-fixed and embedded in paraffin blocks, for example, using standard fixation methods as described herein.

[0298] In some embodiments, the sample includes peripheral whole blood or bone marrow aspirate. In some aspects, the sample, such as a diseased tissue, includes at least 20% nucleated elements. In some aspects, the peripheral whole blood sample or bone marrow aspirate sample is collected in a volume of about 2.5 ml. In some embodiments, the blood sample is shipped on the same day as collection, for example, at ambient temperature, such as 43 - 99°F or 6 - 37°C. In some embodiments, the blood sample is not frozen or refrigerated.

[0299] In some embodiments, the sample includes isolated, for example, extracted nucleic acids, such as DNA or RNA. In some embodiments, the isolated nucleic acid includes DNA or RNA, for example, in nuclease-free water.

[0300] In some embodiments, the sample includes a blood sample, such as a peripheral whole blood sample. In some embodiments, the peripheral whole blood sample is collected, for example, into two tubes using, for example, about 8.5 ml of blood per tube. In some embodiments, the peripheral whole blood sample is collected by venipuncture, for example, according to CLSI H3-A6. In some embodiments, the blood is immediately mixed, for example, about 8-10 times by gentle inversion, for example, by a complete, for example, 180° rotation of the wrist. In some embodiments, the blood sample is shipped on the same day as collection, for example, at ambient temperature, for example, 43-99°F or 6-37°C. In some embodiments, the blood sample is not frozen or refrigerated. In some embodiments, the collected blood sample is maintained, for example, stored, at 43-99°F or 6-37°C.

[0301] Subject In some embodiments, the sample is obtained, for example, collected from a subject, such as a patient, having a condition or disease, such as a proliferative disease (e.g., as described herein) or a non-cancer indication. In some embodiments, the disease is a proliferative disease. In some embodiments, the proliferative disease is cancer, such as a solid tumor or a blood cancer. In some embodiments, the cancer is a solid tumor. In some embodiments, the cancer is a blood cancer, such as leukemia or lymphoma.

[0302] In some embodiments, the subject has cancer. In some embodiments, the subject has been treated or is being treated for cancer. In some aspects, the subject needs to be monitored for cancer progression or regression, for example, after being treated with a cancer therapy. In some aspects, the subject needs to be monitored for cancer recurrence. In some embodiments, the subject is at risk of having cancer. In some embodiments, the subject has not been treated with a cancer therapy. In some embodiments, the subject has a genetic predisposition to cancer (e.g., having a mutation that increases the baseline risk of developing cancer). In some embodiments, the subject has been exposed to an environment (e.g., radiation or a chemical substance) that increases the risk of developing cancer. In some embodiments, the subject needs to be monitored for cancer development.

[0303] In some aspects, the patient has been previously treated with a targeted therapy, e.g., one or more targeted therapies. In some aspects, for a patient who has been previously treated with a targeted therapy, a sample, e.g., a specimen, is obtained, e.g., collected, after the targeted therapy. In some aspects, the post-targeted therapy sample is a sample that was obtained, e.g., collected, after completion of the targeted therapy.

[0304] In some aspects, the patient has not been previously treated with a targeted therapy. In some aspects, for a patient who has not been previously treated with a targeted therapy, the sample includes an excision, e.g., an original excision, or a recurrence, e.g., a post-treatment disease recurrence, e.g., including a non-targeted therapy. In some aspects, the sample is or a part of a primary tumor or a metastasis, e.g., a metastatic biopsy. In some aspects, the sample is obtained from a site that has the highest percentage of tumor, e.g., tumor cells, compared to an adjacent site, e.g., an adjacent site having tumor cells, e.g., a tumor site. In some aspects, the sample is obtained from a site that has the largest tumor focus compared to an adjacent site, e.g., an adjacent site having tumor cells, e.g., a tumor site.

[0305] In some embodiments, the disease is selected from non-small cell lung cancer (NSCLC), melanoma, breast cancer, colorectal cancer (CRC), or ovarian cancer. In some embodiments, the NSCLC described herein includes, for example, NSCLC having an EGFR change (e.g., exon 19 deletion or exon 21 L858R change), an ALK rearrangement, or BRAF V600E. In some embodiments, the melanoma described herein includes melanoma having a BRAF change, such as V600E and / or V600K. In some embodiments, the breast cancer described herein includes breast cancer having an ERBB2 (HER2) amplification. In some embodiments, the colorectal cancer described herein includes colorectal cancer having wild-type KRAS, e.g., no mutation at codons 12 and / or 13, or no mutation at codons 2, 3, and / or 4. In some embodiments, the colorectal cancer described herein includes colorectal cancer having wild-type NRAS, e.g., no mutation at codons 2, 3, and / or 4. In some embodiments, the colorectal cancer described herein includes colorectal cancer having, for example, wild-type KRAS as described herein, and, for example, wild-type NRAS as described herein. In some embodiments, the ovarian cancer described herein includes ovarian cancer having a change in BRCA1 and / or BRCA2.

[0306] Target capture reagent The methods described herein provide optimized sequencing of a plurality of genes and gene products from a sample from one or more subjects, e.g., from a cancer as described herein, by appropriate selection of a target capture reagent for selecting a sequence-specified target nucleic acid molecule, e.g., a target capture reagent for use in solution hybridization.

[0307] Any combination of two, three, four, five, or more than one target capture reagent, for example, a first and a second plurality of target capture reagents; a first and a third plurality of target capture reagents; a first and a fourth plurality of target capture reagents; a first and a fifth plurality of target capture reagents; a second and a third plurality of target capture reagents; a second and a fourth plurality of target capture reagents; a second and a fifth plurality of target capture reagents; a third and a fourth plurality of target capture reagents; a third and a fifth plurality of target capture reagents; a fourth and a fifth plurality of target capture reagents; a first, a second, and a third plurality of target capture reagents; a first, a second, and a fourth plurality of target capture reagents; a first, a second, and a fifth plurality of target capture reagents; a first, a second, a third, and a fourth plurality of target capture reagents; a first, a second, a third, a fourth, and a fifth plurality of target capture reagents, etc. can be used.

[0308] In some embodiments, the method comprises (a) obtaining a library comprising a plurality of nucleic acid molecules (e.g., target nucleic acid molecules) from a sample, e.g., a sample as described herein, e.g., a plurality of tumor nucleic acid molecules from a sample; (b) contacting the library with two, three, or more than one target capture reagent to provide selected nucleic acid molecules (e.g., library catch); (c) obtaining reads for a target region from nucleic acid molecules, e.g., tumor nucleic acid molecules from the library or library catch, by a method comprising sequencing, e.g., using a next-generation sequencing method; (d) aligning the reads by an alignment method, e.g., the alignment method described herein; (e) assigning nucleotide values (e.g., calling mutations, e.g., using a Bayesian method or a method described herein) from the reads for nucleotide positions.

[0309] In some embodiments, as used herein, the level of sequence-specific depth (e.g., X-fold level of sequence-specific depth) indicates the number of reads (e.g., unique reads) after detection and removal of duplicate reads, such as PCR duplicate reads. In other embodiments, duplicate reads are evaluated, for example, to assist in the detection of copy number alterations (CNA).

[0310] In one embodiment, the target capture reagent selects a target region that includes one or more rearrangements, such as an intron that includes a genomic rearrangement. In such embodiments, the target capture reagent is designed such that repetitive sequences are masked to increase selection efficiency. In embodiments where the rearrangement has a known junction sequence, complementary target capture reagents can be designed to the junction sequence to increase selection efficiency.

[0311] In some aspects, the method includes the use of a target capture reagent designed to capture two or more different target categories, each category having a different design strategy. In some embodiments, the methods and compositions disclosed herein (e.g., hybridization capture methods) capture a subset of target sequences (e.g., target nucleic acid molecules) and provide uniform coverage of the target sequences while minimizing coverage outside of the subset. In one embodiment, the target sequences include the entire exome or a selected subset thereof from genomic DNA. In another embodiment, the target sequences include a large chromosomal region, such as an entire chromosomal arm. The methods and compositions disclosed herein provide different target capture reagents for achieving different sequence-specific depth and coverage patterns for complex target nucleic acid sequences (e.g., nucleic acid libraries).

[0312] In one embodiment, the method includes providing selected nucleic acid molecules of one or more nucleic acid libraries (e.g., library catch). For example, the method Providing one or more libraries (e.g., one or more nucleic acid libraries) comprising a plurality of nucleic acid molecules, such as target nucleic acid molecules (e.g., comprising a plurality of tumor nucleic acid molecules and / or reference nucleic acid molecules), Contacting one or more libraries with two, three, or more target capture reagents (e.g., oligonucleotide target capture reagents), for example, in a solution-based reaction, to form a hybridization mixture comprising a plurality of target capture reagent / nucleic acid molecule hybrids; Separating the plurality of target capture reagent / nucleic acid molecule hybrids from the hybridization mixture, for example, by contacting the hybridization mixture with a binding entity that permits separation of the plurality of target capture reagent / nucleic acid molecule hybrids from the hybridization mixture; Thereby providing a library catch (e.g., a selected or enriched subset of nucleic acid molecules from one or more libraries).

[0313] In one embodiment, each of the first, second, or third plurality of target capture reagents has a unique recovery efficiency. In some embodiments, at least two or three of the plurality of target capture reagents have different recovery efficiency values.

[0314] In certain embodiments, the value of the recovery efficiency is modified by one or more of differential representation of different target capture reagents, differential overlap of target capture reagent subsets, differential target capture reagent parameters, mixing of different target capture reagents, and / or use of different types of target capture reagents. For example, variations in recovery efficiency (e.g., relative sequence coverage of each target capture reagent / target category) can be, for example, within and / or between a plurality of target capture reagents, (i) Differential representation of different target capture reagents - The design of a target capture reagent for capturing a given target (e.g., a target nucleic acid molecule) can include more / fewer copies to enhance / reduce relative target sequence specification depth, (ii) Differential overlap of target capture reagent subsets - The design of target capture reagents for capturing a given target (e.g., a target nucleic acid molecule) can include longer or shorter overlaps between adjacent target capture reagents to enhance / reduce the relative target sequence determination depth. (iii) Differential target capture reagent parameters - The design of target capture reagents for capturing a given target (e.g., a target nucleic acid molecule) can include sequence modifications / shorter lengths to reduce the capture efficiency and reduce the relative target sequence determination depth. (iv) Mixing of different target capture reagents - Target capture reagents designed to capture different target sets can be mixed in different molar ratios to enhance / reduce the relative target sequence determination depth. (v) Use of different types of oligonucleotide target capture reagents - In certain embodiments, the target capture reagent can be one or more of the following: (a) One or more chemically (e.g., non-enzymatically) synthesized (e.g., individually synthesized) target capture reagents; (b) One or more target capture reagents synthesized on an array; and (c) One or more enzymatically prepared (e.g., in vitro transcribed) target capture reagents; (d) Any combination of (a), (b), and / or (c); (e) One or more DNA oligonucleotides (e.g., natural or non-natural DNA oligonucleotides); (f) One or more RNA oligonucleotides (e.g., natural or non-natural RNA oligonucleotides); (g) A combination of (e) and (f); or (h) Any combination of the above.

[0315] Combinations of different oligonucleotides can be mixed at different ratios, for example, ratios selected from 1:1, 1:2, 1:3, 1:4, 1:5, 1:10, 1:20, 1:50, 1:100, 1:1000, etc. In one embodiment, the ratio of the chemically synthesized target capture reagent to the array generation target capture reagent is selected from 1:5, 1:10, or 1:20. The DNA or RNA oligonucleotides can be natural or non-natural. In certain embodiments, the target capture reagent contains one or more non-natural nucleotides, for example, to raise the melting temperature. Exemplary non-natural oligonucleotides include modified DNA or RNA nucleotides. Exemplary modified nucleotides (e.g., modified RNA or DNA nucleotides) include, but are not limited to, locked nucleic acid (LNA), where the ribose portion of the LNA nucleotide is modified with an extra bridge connecting the 2'-oxygen and the 4'-carbon. Peptide nucleic acid (PNA), e.g., PNA composed of repeating N-(2-aminoethyl)-glycine units linked by peptide bonds; DNA or RNA oligonucleotides modified to capture low GC regions; bicyclic nucleic acids (BNA); bridged oligonucleotides; modified 5-methyldeoxycytidine; and 2,6-diaminopurine. Other modified DNA and RNA nucleotides are known in the art.

[0316] In certain embodiments, substantially uniform or uniform coverage of the target sequence (e.g., target nucleic acid molecule) is obtained. For example, within each target capture reagent / target category, the coverage uniformity can be optimized by changing the target capture reagent parameters, for example, by one or more of the following. (i) The presentation or overlap of the target capture reagent can be increased or decreased to enhance / reduce the coverage of targets (e.g., target nucleic acid molecules) that are under / over covered relative to other targets within the same category. (ii) When the coverage is low and the target sequence is difficult to capture (e.g., high GC content sequence), the region targeted by the target capture reagent can be expanded to cover, for example, adjacent sequences (e.g., adjacent sequences with less GC richness). (iii) By using the modification of the target capture reagent array, the secondary structure of the target capture reagent can be reduced and its recovery efficiency can be enhanced. (iv) Changing the length of the target capture reagent can be used to equalize the melting hybridization rates of different target capture reagents within the same category. The length of the target capture reagent can be changed directly (by generating target capture reagents of various lengths) or indirectly (by generating target capture reagents of a certain length and replacing the ends of the target capture reagents with any sequence). (v) Modifying target capture reagents with different orientations for the same target region (i.e., the sense and antisense strands) can have different binding efficiencies. A target capture reagent with any orientation that provides optimal coverage for each target can be selected. (vi) Changing the amount of binding entity present on each target capture reagent, such as a capture tag (e.g., biotin), can affect its binding efficiency. Increasing / decreasing the tag level of a target capture reagent targeting a specific target can be used to enhance / reduce the relative target coverage. (vii) By using the change in the type of nucleotide used for different target capture reagents, the binding affinity to the target can be affected and the relative target coverage can be enhanced / reduced. (viii) By using modified oligonucleotide target capture reagents, for example, those with more stable base pairing, the melting hybridization rates between regions with low or normal GC content compared to regions with high GC content can be equalized.

[0317] In one embodiment, the method includes the use of a plurality of target capture reagents, including a target capture reagent that selects a tumor nucleic acid molecule, e.g., a nucleic acid molecule containing a target region from a tumor cell. The tumor nucleic acid molecule can be any nucleotide sequence present in a tumor cell, e.g., a mutation, wild type, reference, or intron nucleotide sequence described herein that is present in a tumor or cancer cell. In one embodiment, the tumor nucleic acid molecule includes changes that occur at a low frequency (e.g., one or more mutations), e.g., less than about 5% of the cells from the sample have changes in their genomes. In other embodiments, the tumor nucleic acid molecule includes changes (e.g., one or more mutations) that occur at a frequency of about 10% of the cells from the sample. In other embodiments, the tumor nucleic acid molecule includes an intron sequence, e.g., a sub-genomic region from an intron sequence described herein, a reference sequence present in a tumor cell.

[0318] In other embodiments, the method includes amplifying the library catch (e.g., by PCR). In other embodiments, the library catch is not amplified.

[0319] In another aspect, the invention features the target capture reagents described herein and combinations of the individual plurality of target capture reagents described herein. The target capture reagent can be part of a kit that can include, optionally, instructions, standards, buffers, or enzymes or other reagents.

[0320] Design and Construction of Target Capture Reagents In some embodiments, the target capture reagent is a molecule that can bind to a target molecule, thereby enabling capture of the target molecule. For example, the target capture reagent can be a bait, such as a nucleic acid molecule, such as a DNA or RNA molecule, that can hybridize (e.g., complementarily) and thereby enable capture of the target nucleic acid. In some embodiments, the target capture reagent, such as the bait, is a capture oligonucleotide. In certain embodiments, the target nucleic acid is a genomic DNA molecule. In other embodiments, the target nucleic acid is an RNA molecule or a cDNA molecule derived from an RNA molecule. In one embodiment, the target capture reagent is a DNA molecule. In one embodiment, the target capture reagent is an RNA molecule. In one embodiment, the target capture reagent is suitable for solution-phase hybridization. In one embodiment, the target capture reagent is suitable for solid-phase hybridization. In one embodiment, the target capture reagent is suitable for both solution-phase and solid-phase hybridization.

[0321] Typically, DNA molecules are used as the target capture reagent sequences, although RNA molecules can also be used. In some embodiments, the DNA molecule target capture reagent can be single-stranded DNA (ssDNA) or double-stranded DNA (dsDNA).

[0322] In some embodiments, the RNA-DNA duplex is more stable than the DNA-DNA duplex and thus potentially provides better nucleic acid capture. RNA target capture reagents can be made as described elsewhere herein using methods known in the art including, but not limited to, de novo chemical synthesis and transcription of DNA molecules using DNA-dependent RNA polymerases. In one embodiment, the target capture reagent sequence is generated using known nucleic acid amplification methods such as PCR using, for example, human DNA or a pooled human DNA sample as a template. The oligonucleotide can then be converted to an RNA target capture reagent. In one embodiment, in vitro transcription is used based on, for example, adding an RNA polymerase promoter sequence to one end of the oligonucleotide. In one embodiment, the RNA polymerase promoter sequence is added to the end of the target capture reagent by amplifying or re-amplifying the target capture reagent sequence, for example, by tailing one of the primers of each target-specific primer pair with the RNA promoter sequence using, for example, PCR or another nucleic acid amplification method. In one embodiment, the RNA polymerase is T7 polymerase, SP6 polymerase, or T3 polymerase. In one embodiment, the RNA target capture reagent is labeled with a tag, such as an affinity tag. In one embodiment, the RNA target capture reagent is made by in vitro transcription using, for example, biotinylated UTP. In another embodiment, the RNA target capture reagent is made without biotin and then biotin is cross-linked to the RNA molecule using methods well known in the art such as psoralen cross-linking. In one embodiment, the RNA target capture reagent is an RNase-resistant RNA molecule that can be made by generating an RNase-resistant RNA molecule using, for example, modified nucleotides during transcription. In one embodiment, the RNA target capture reagent corresponds to only one strand of a double-stranded DNA target. Typically, such RNA target capture reagents are not self-complementary and are more effective as hybridization drivers.

[0323] The target capture reagent can be designed from a reference sequence such that the target capture reagent is optimal for selecting a target of the reference sequence. In some embodiments, the target capture reagent sequence is designed using degenerate bases (e.g., wobble). For example, degenerate bases can be included in the target capture reagent sequence at positions of common SNPs or mutations to optimize the target capture reagent sequence to capture both alleles (e.g., SNPs and non-SNPs; variants and non-variants). In some embodiments, all known sequence variations (or a subset thereof) can be targeted with a plurality of oligonucleotide target capture reagents rather than using degenerate oligonucleotides.

[0324] In certain embodiments, the target capture reagent comprises an oligonucleotide (or oligonucleotides) that is about 100 nucleotides to 300 nucleotides in length. Typically, the target capture reagent comprises an oligonucleotide (or oligonucleotides) that is about 130 nucleotides to 230 nucleotides, or about 150 nucleotides to 200 nucleotides in length. In other embodiments, the target capture reagent comprises an oligonucleotide (or oligonucleotides) that is about 300 nucleotides to 1000 nucleotides in length.

[0325] In some embodiments, the target nucleic acid molecule specific sequence in the oligonucleotide is about 40 to 1000 nucleotides, about 70 to 300 nucleotides, about 100 to 200 nucleotides in length, typically about 120 to 170 nucleotides in length.

[0326] In some aspects, the target capture reagent comprises a binding entity. The binding entity can be an affinity tag. In some embodiments, the affinity tag is a biotin molecule or a hapten. In certain embodiments, the binding entity enables separation of the target capture reagent / nucleic acid molecule hybrid from the hybridization mixture by binding to a partner such as an avidin molecule or an antibody that binds to the hapten or an antigen-binding fragment thereof.

[0327] In other embodiments, the oligonucleotide in the target capture reagent includes a forward complementary sequence and a reverse complementary sequence to the same target nucleic acid molecule sequence, whereby the oligonucleotide having a reverse complementary nucleic acid molecule specific sequence also has a reverse complementary universal tail. This can result in RNA transcripts that are the same strand, i.e., not complementary to each other.

[0328] In other embodiments, the target capture reagent includes an oligonucleotide containing degenerate or mixed bases at one or more positions. In yet other embodiments, the target capture reagent includes a plurality or substantially all known sequence variants present in a population of a single species or community of organisms. In one embodiment, the target capture reagent includes a plurality or substantially all known sequence variants present in a human population.

[0329] In other embodiments, the target capture reagent includes or is derived from a cDNA sequence. In other embodiments, the target capture reagent includes an amplification product (e.g., a PCR product) amplified from genomic DNA, cDNA, or cloned DNA.

[0330] In other embodiments, the target capture reagent includes an RNA molecule. In some embodiments, the set includes RNA molecules that are chemically, enzymatically modified, or in vitro transcribed (including but not limited to those that are more stable and resistant to RNase).

[0331] In yet other embodiments, the target capture reagent is described in US Patent Application Publication No. 2010 / 0029498 and Gnirke, A. et al. (2009) Nat Biotechnol. 27(2):182-189. For example, a biotinylated RNA target capture reagent can be produced by obtaining a pool of synthetic long oligonucleotides first synthesized on a microarray and amplifying the oligonucleotides to generate the target capture reagent sequence. In some embodiments, the target capture reagent is generated by adding an RNA polymerase promoter sequence to one end of the target capture reagent sequence and synthesizing an RNA sequence using RNA polymerase. In one embodiment, a library of synthetic oligodeoxynucleotides can be obtained from a commercial supplier such as Agilent Technologies, Inc. and amplified using known nucleic acid amplification methods.

[0332] Accordingly, a method for making the aforementioned target capture reagent is provided. The method includes, for example, selecting one or more target capture reagents, such as target-specific bait oligonucleotide sequences (e.g., one or more of the mutation capture, reference or control oligonucleotide sequences described herein), and obtaining a pool of target capture reagents, such as a pool of target-specific bait oligonucleotide sequences (e.g., synthesizing a pool of target-specific bait oligonucleotide sequences, such as by microarray synthesis), and optionally amplifying the target capture reagent, such as the target-specific bait oligonucleotide sequence.

[0333] In other embodiments, the method further comprises amplifying the oligonucleotide (e.g., by PCR) using one or more biotinylated primers. In some embodiments, the oligonucleotide comprises a universal sequence at the end of each oligonucleotide bound to the microarray. The method can further comprise removing the universal sequence from the oligonucleotide. Such methods can also include removing the complementary strand of the oligonucleotide, annealing the oligonucleotide, and extending the oligonucleotide. In some of these embodiments, the method for amplifying the oligonucleotide (e.g., by PCR) uses one or more biotinylated primers. In some embodiments, the method further comprises size selecting the amplified oligonucleotide.

[0334] In one embodiment, an RNA target capture reagent is produced. The method includes producing a set of target capture reagent sequences according to the methods described herein, adding an RNA polymerase promoter sequence to one end of the target capture reagent sequences, and synthesizing an RNA sequence using RNA polymerase. The RNA polymerase can be selected from T7 RNA polymerase, SP6 RNA polymerase, or T3 RNA polymerase. In other embodiments, the RNA polymerase promoter sequence is added to the end of the target capture reagent sequence by amplifying the target capture reagent sequence (e.g., by PCR). In embodiments where the target capture reagent sequence is amplified by PCR using a specific primer pair from genomic DNA or cDNA, a PCR product that can be transcribed into an RNA target capture reagent using standard methods can be obtained by adding an RNA promoter sequence to the 5' end of one of the two specific primers of each pair.

[0335] In other embodiments, human DNA or a pooled human DNA sample can be used as a template to generate target capture reagents. In such embodiments, the oligonucleotides are amplified by polymerase chain reaction (PCR). In other embodiments, the amplified oligonucleotides are re-amplified by rolling circle amplification or hyper-branched rolling circle amplification. Using the same method, target capture reagent sequences can also be generated using human DNA or a pooled human DNA sample as a template. Using the same method, genomic sub-fragments obtained by other methods including, but not limited to, restriction digestion, pulsed field gel electrophoresis, flow sorting, CsCl density gradient centrifugation, selective kinetic reassociation, microdissection of chromosome preparations, and other fractionation methods known to those of skill in the art can also be used to generate target capture reagent sequences.

[0336] In certain embodiments, the number of target capture reagents (e.g., baits) among the plurality of target capture reagents is less than 1,000. In other embodiments, the number of target capture reagents (e.g., baits) among the plurality of target capture reagents is greater than 1,000, greater than 5,000, greater than 10,000, greater than 20,000, greater than 50,000, greater than 100,000, or greater than 500,000.

[0337] The length of the target capture reagent sequence can be from about 70 nucleotides to 1,000 nucleotides. In one embodiment, the length of the target capture reagent is from about 100 to 300 nucleotides, 110 to 200 nucleotides, or 120 to 170 nucleotides in length. In addition to the above, intermediate oligonucleotide lengths of about 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 300, 400, 500, 600, 700, 800 and 900 nucleotides in length can be used in the methods described herein. In some embodiments, oligonucleotides of about 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220 or 230 bases can be used.

[0338] Each target capture reagent array can include a target-specific (e.g., nucleic acid molecule-specific) target capture reagent array and universal tails at one or both ends. As used herein, the term "target capture reagent sequence" can refer to a target-specific target capture reagent sequence, or the entire oligonucleotide including the target-specific "target capture reagent sequence" and other nucleotides of the oligonucleotide. The target-specific sequence in the target capture reagent is about 40 nucleotides to 1000 nucleotides in length. In one embodiment, the target-specific sequence is about 70 nucleotides to 300 nucleotides in length. In another embodiment, the target-specific sequence is about 100 nucleotides to 200 nucleotides in length. In yet another embodiment, the target-specific sequence is about 120 nucleotides to 170 nucleotides in length, typically 120 nucleotides in length. In addition to the above, intermediate lengths, such as target-specific sequences of about 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 300, 400, 500, 600, 700, 800 and 900 nucleotides in length, as well as target-specific sequences of lengths between the above lengths, can also be used in the methods described herein.

[0339] In one embodiment, the target capture reagent is an oligomer (e.g., composed of RNA oligomers, DNA oligomers, or combinations thereof) having a length of about 50 to 200 nucleotides (e.g., about 50, 60, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 190, or 200 nucleotides in length). In one embodiment, each target capture reagent oligomer contains about 120 to 170, or typically about 120, nucleotides that are target-specific target capture reagent sequences. The target capture reagent can include additional non-target-specific nucleotide sequences at one or both ends. The additional nucleotide sequences can be used, for example, for PCR amplification or as target capture reagent identifiers. In certain embodiments, the target capture reagent further includes a binding entity (e.g., an affinity tag such as a biotin molecule) described herein. The binding entity, such as a biotin molecule, can be attached to the target capture reagent, for example, at the 5' end, 3' end, or internally (e.g., by incorporating a biotinylated nucleotide) of the target capture reagent. In one embodiment, the biotin molecule is attached to the 5' end of the target capture reagent.

[0340] In an exemplary embodiment, the target capture reagent is an oligonucleotide about 150 nucleotides in length, of which 120 nucleotides are the target-specific "target capture reagent sequence". The other 30 nucleotides (e.g., 15 nucleotides at each end) are arbitrary tails that are universal for PCR amplification. The tails can be any sequence selected by the user. For example, a pool of synthetic oligonucleotides can include oligonucleotides having the sequence 5'-ATCGCACCAGCGTGTN 120 CACTGCGGCTCCTCA-3' (SEQ ID NO: 1), where N 120 represents the target-specific target capture reagent sequence.

[0341] The target capture reagent sequences described herein can be used for the selection of exons and short target sequences. In one embodiment, the target capture reagent is about 100 nucleotides to 300 nucleotides in length. In another embodiment, the target capture reagent is about 130 nucleotides to 230 nucleotides in length. In yet another embodiment, the target capture reagent is about 150 nucleotides to 200 nucleotides in length. For example, the target-specific sequences in the target capture reagent for the selection of exons and short target sequences are about 40 nucleotides to 1000 nucleotides in length. In one embodiment, the target-specific sequence is about 70 nucleotides to 300 nucleotides in length. In another embodiment, the target-specific sequence is about 100 nucleotides to 200 nucleotides in length. In yet another embodiment, the target-specific sequence is about 120 nucleotides to 170 nucleotides in length.

[0342] In some embodiments, the long oligonucleotides can minimize the number of oligonucleotides necessary to capture the target sequence. For example, one oligonucleotide can be used per exon. It is known in the art that the average and median lengths of protein-coding exons in the human genome are approximately 164 and 120 base pairs, respectively. Longer target capture reagent sequences are more specific than shorter ones and can capture better. As a result, the success rate per oligonucleotide target capture reagent sequence is higher than that of short oligonucleotides. In one embodiment, the minimum target capture reagent coverage sequence is, for example, the size of one target capture reagent (e.g., 120-170 bases) for capturing an exon-sized target. When specifying the length of the target capture reagent sequence, it can also be taken into account that an unnecessarily long target capture reagent captures more unwanted DNA directly adjacent to the target. Longer oligonucleotide target capture reagents can also be more resistant to polymorphisms in the target region in the DNA sample than shorter ones. Typically, the target capture reagent sequence is derived from a reference genome sequence. If the target sequence in the actual DNA sample deviates from the reference sequence, for example, if it contains a single nucleotide polymorphism (SNP), it may not hybridize efficiently to the target capture reagent, and thus may not be represented or may be completely absent in the sequence hybridized to the target capture reagent sequence. For example, a single mismatch of 120-170 bases may have less impact on hybridization stability than a single mismatch of 20 or 70 bases, which are typical target capture reagent or primer lengths in multiplex amplification and microarray capture, respectively. Therefore, allelic dropout due to SNPs may be less likely to occur with longer synthetic target capture reagent molecules.

[0343] To select targets that are long compared to the length of the capture reagent for capture targets such as genomic regions, the length of the target capture reagent array is typically in the same size range as the target capture reagent for the short targets above, except that there is no need to limit the maximum size of the target capture reagent array for the sole purpose of minimizing the targeting of adjacent sequences. Alternatively, oligonucleotides can be tiled over a much wider window (typically 600 bases). This method can be used to capture DNA fragments that are much larger (e.g., about 500 bases) than typical exons. As a result, many more unwanted adjacent non-target sequences are selected.

[0344] Synthesis of the target capture reagent The target capture reagent can be, for example, any type of oligonucleotide, such as DNA or RNA. DNA or RNA target capture reagents ("oligo target capture reagents") can be synthesized individually or in an array as DNA or RNA target capture reagents (e.g., "array baits"). Oligo target capture reagents are typically single-stranded, whether provided in array format or as isolated oligos. The target capture reagent can further include a binding entity (e.g., an affinity tag such as a biotin molecule) described herein. A binding entity, such as a biotin molecule, can be attached to the target capture reagent, for example, at the 5' or 3' end of the target capture reagent, typically at the 5' end of the target capture reagent. The target capture reagent can be synthesized by methods described in the art, such as those described in International Patent Application Publication No. 2012 / 092426 or International Patent Application Publication No. 2015 / 021080, the entire contents of which are incorporated herein by reference.

[0345] Hybridization conditions The method featured in the present invention involves contacting a library (e.g., a nucleic acid library) with a plurality of target capture reagents to provide a selected library catch. The contacting step can be performed by solution hybridization. In certain embodiments, the method includes repeating the hybridization step by one or more additional solution hybridizations. In some embodiments, the method further includes subjecting the library catch to one or more additional solution hybridizations with the same or a different set of target capture reagents. Hybridization methods that can be adapted for use in the methods herein are described in the art, for example, as described in International Patent Application Publication No. 2012 / 092426.

[0346] Further embodiments or features of the present invention are as follows.

[0347] In certain embodiments, the method includes determining the presence or absence of a change, e.g., a positive or negative change, associated with a cancerous phenotype in a sample (e.g., at least 10, 20, 30, 50 or more changes in the genes or gene products described herein). In other embodiments, the method includes identifying a genomic signature, e.g., a continuous / composite biomarker (e.g., the level of tumor mutational burden). In other embodiments, the method includes determining the presence or absence of one or more genomic signatures, e.g., continuous / composite biomarkers, e.g., the level of microsatellite instability, or loss of heterozygosity (LOH). The method includes contacting nucleic acids in a sample by a solution-based reaction with any of the methods and target capture reagents described herein to obtain a library catch, and determining the presence or absence of a change in the genes or gene products described herein by sequencing all or a subset of the library catch (e.g., by next-generation sequencing).

[0348] In certain embodiments, the target capture reagent comprises an oligonucleotide (or oligonucleotides) that is about 100 nucleotides to 300 nucleotides in length. Typically, the target capture reagent comprises an oligonucleotide (or oligonucleotides) that is about 130 nucleotides to 230 nucleotides, or about 150 nucleotides to 200 nucleotides in length. In other embodiments, the target capture reagent comprises an oligonucleotide (or oligonucleotides) that is about 300 nucleotides to 1000 nucleotides in length.

[0349] In other embodiments, the target capture reagent comprises or is derived from a cDNA sequence. In one embodiment, the cDNA is prepared from an RNA sequence, such as RNA derived from a tumor or cancer cell, such as RNA obtained from a tumor-FFPE sample, a blood sample or a bone marrow aspirate sample. In other embodiments, the target capture reagent comprises an amplification product (e.g., a PCR product) amplified from genomic DNA, cDNA or cloned DNA.

[0350] In certain embodiments, a library (e.g., a nucleic acid library) comprises a collection of nucleic acid molecules. As described herein, the nucleic acid molecules of the library can comprise target nucleic acid molecules (e.g., tumor nucleic acid molecules, reference nucleic acid molecules and / or control nucleic acid molecules; also referred to herein as first, second and / or third nucleic acid molecules, respectively). The nucleic acid molecules of the library can be derived from a single individual. In some embodiments, the library can comprise nucleic acid molecules from two or more subjects (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30 or more subjects), e.g., two or more libraries from different subjects can be combined to form a library having nucleic acid molecules from two or more subjects. In one embodiment, the subject is a human having or at risk of having cancer or a tumor.

[0351] In some embodiments, the method comprises contacting one or more libraries (e.g., one or more nucleic acid libraries) with a plurality of target capture reagents to provide a selected subgroup of nucleic acids, e.g., a library catch. In one embodiment, the contacting step is performed on a solid support, e.g., in an array. Solid supports suitable for hybridization are described, for example, in Albert, T.J. et al. (2007) Nat. Methods 4(11):903-5; Hodges, E. et al. (2007) Nat. Genet. 39(12):1522-7; and Okou, D.T. et al. (2007) Nat. Methods 4(11):907-9, the contents of which are incorporated herein by reference. In other embodiments, the contacting step is performed by solution hybridization. In certain embodiments, the method comprises repeating the hybridization step by one or more additional hybridizations. In some embodiments, the method further comprises subjecting the library catch to one or more additional hybridizations with the same or a different set of target capture reagents.

[0352] In still other embodiments, the method further comprises subjecting the library catch to genotyping to identify the genotype of the selected nucleic acids.

[0353] In certain embodiments, the method comprises i) fingerprinting the sample; ii) quantifying the abundance of a gene or gene product (e.g., a gene or gene product as described herein) in the sample (e.g., quantifying the relative abundance of transcripts in the sample); iii) identifying the sample as belonging to a particular subject (e.g., a normal control or a cancer patient); iv) identifying a genetic trait in the sample (e.g., the genetic makeup of one or more subjects (e.g., ethnicity, race, familial traits)); v) determining the ploidy in a nucleic acid sample; determining the loss of heterozygosity in the sample; (vi) determining the presence or absence of changes described in this specification, such as nucleotide substitutions, copy number changes, indels or rearrangements in a sample; (vii) identifying the level of tumor mutation burden and / or microsatellite instability (and / or other complex biomarkers) in a sample; (viii) identifying the level of tumor / normal cell mixture in a sample, and comprising.

[0354] Combinations of different oligonucleotides can be mixed at different ratios, for example, ratios selected from 1:1, 1:2, 1:3, 1:4, 1:5, 1:10, 1:20, 1:50, 1:100, 1:1000, etc. In one embodiment, the ratio of the chemically synthesized target capture reagent (e.g., bait) to the array-generated target capture reagent (e.g., bait) is selected from 1:5, 1:10, or 1:20. The DNA or RNA oligonucleotides can be natural or non-natural. In certain embodiments, the target capture reagent (e.g., bait) contains one or more non-naturally occurring nucleotides, for example, to increase the melting temperature. Exemplary non-natural oligonucleotides include modified DNA or RNA nucleotides. An exemplary modified RNA nucleotide is locked nucleic acid (LNA), and the ribose portion of the LNA nucleotide is modified by an extra bridge connecting the 2'-oxygen and 4'-carbon (Kaur, H; Arora, A; Vengel, J; Maiti, S. (2006). "Thermodynamic, counterion, and hydration effects for incorporation of locked nucleic acid nucleotides into DNA duplexes" Biochemistry 45(23):7347-55). Other exemplary modified DNA and RNA nucleotides include peptide bonds (Egholm, M. et al. (1993) Nature 365(6446):566-8), DNA or RNA oligonucleotides modified to capture low GC regions; bicyclic nucleic acids (BNA) or bridged oligonucleotides; modified 5-methyldeoxycytidine; and peptide nucleic acids (PNA) composed of repeated N-(2-aminoethyl)-glycine units linked by 2,6-diaminopurine, but are not limited thereto. Other modified DNA and RNA nucleotides are known in the art.

[0355] In one embodiment, the method further includes obtaining a library, wherein the size of the nucleic acid fragments in the library is below a reference value, and the library is prepared without a fragmentation step between DNA isolation and library preparation.

[0356] In one embodiment, the method further comprises obtaining a nucleic acid fragment, wherein the size of the nucleic acid fragment is equal to or greater than a reference value, the nucleic acid fragment is fragmented, and then such nucleic acid fragments are made into a library.

[0357] In one embodiment, the method further comprises labeling each of the plurality of library nucleic acid molecules, for example, by adding a distinguishable separate nucleic acid sequence (barcode) to each of the plurality of nucleic acid molecules.

[0358] In one embodiment, the method further comprises attaching a primer to each of the plurality of library nucleic acid molecules.

[0359] In one embodiment, the method further comprises providing a plurality of target capture reagents and selecting a plurality of target capture reagents, wherein the selection is based on: 1) patient characteristics, such as age, tumor stage, previous treatment, or resistance; 2) tumor type; 3) sample characteristics; 4) characteristics of a control sample; 5) presence or type of a control; 6) characteristics of an isolated tumor (or control) nucleic acid sample; 7) library characteristics; 8) mutations known to be associated with the type of tumor in the sample; 9) mutations not known to be associated with the type of tumor in the sample; 10) the ability to identify sequences to be sequenced (or hybridized or recovered) or difficulties associated with sequences having mutations, such as high GC regions or rearrangements; or 11) the gene to be sequenced.

[0360] In one embodiment, the method further comprises selecting a target capture reagent or a plurality of target capture reagents, for example, in response to the identification of a small number of tumor cells in the sample, to provide relatively highly efficient capture of nucleic acid molecules of a first gene as compared to nucleic acid molecules of a second gene, wherein, for example, a mutation in the first gene is associated with the tumor phenotype of the tumor type of the sample and optionally a mutation in the second gene is not associated with the tumor phenotype of the tumor type of the sample.

[0361] In one embodiment, the method further includes obtaining a library capture property, such as a value of nucleic acid concentration, and comparing the obtained value to a reference standard of the property.

[0362] In one embodiment, the method further includes selecting a library having a value of a library property that meets a reference standard for library quantification.

[0363] Sequence identification The methods and systems described herein can be used in combination with, or as part of, a method or system for nucleic acid sequence identification.

[0364] In some embodiments, nucleic acid molecules from a library are isolated, e.g., using solution hybridization, thereby providing a library capture. The library capture or a subgroup thereof can be sequence identified. Accordingly, the methods described herein can further include analyzing the library capture. In some embodiments, the library capture is analyzed by a sequence identification method, such as the next generation sequence identification methods described herein. In some embodiments, the method includes isolating a library capture by solution hybridization and subjecting the library capture to nucleic acid sequence identification. In certain embodiments, the library capture is re-sequence identified.

[0365] Any sequence identification method known in the art can be used. For example, sequence identification of nucleic acids isolated by solution hybridization is typically performed using next generation sequence identification (NGS). Sequence identification methods suitable for use herein are described in the art, e.g., as described in International Patent Application Publication No. WO 2012 / 092426.

[0366] In one embodiment, at least 10, 20, 30, 40, 50, 60, 70, 80, or 90% of the reads obtained or analyzed are for target intervals from genes described herein, such as genes from Tables 2A - 5B. In one embodiment, at least 0.01, 0.02, 0.03, 0.04, 0.05, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 2.0, 5.0, 10, 15, or 30 megabases, such as genomic bases, are sequenced. In one embodiment, the method includes obtaining nucleotide sequence reads from a sample described herein. In one embodiment, the reads are provided by an NGS sequencing method.

[0367] The methods disclosed herein can be used to detect changes present in a subject's genome, whole exome, or transcriptome and are applicable to DNA and RNA sequencing, such as targeted DNA and / or RNA sequencing. In some embodiments, transcripts of the genes described herein are sequenced. In other embodiments, the method includes detecting a change in the level (e.g., increase or decrease) of a gene or gene product, such as a change in the expression of a gene or gene product described herein. The method can optionally include the step of enriching the sample for the target RNA. In other embodiments, the method includes depleting a sample of a specific highly abundant RNA, such as ribosomal RNA or globin RNA. The RNA sequencing method can be used alone or in combination with the DNA sequencing methods described herein. In one embodiment, the method includes performing a DNA sequencing step and an RNA sequencing step. The methods can be performed in any order. For example, the method can include confirming the expression of the changes described herein by RNA sequencing, such as confirming the expression of a mutation or fusion detected by the DNA sequencing method of the invention. In other embodiments, the method includes performing an RNA sequencing step followed by a DNA sequencing step.

[0368] Alignment The methods disclosed herein can integrate the use of multiple individually tailored alignment methods or algorithms to optimize performance in array identification methods, particularly methods that rely on large-scale parallel array identification of multiple diverse gene events in a number of diverse genes, such as those described herein, for example, in methods of analyzing cancer-derived samples.

[0369] In some embodiments, the alignment methods used to analyze reads are not individually customized or adjusted for each of the numerous variants in different genes. In some embodiments, a multiplex alignment method that is individually customized or adjusted for at least a subset of the numerous variants in different genes is used to analyze reads. In some embodiments, a multiplex alignment method that is individually customized or adjusted for each of the numerous variants in different genes is used to analyze reads. In some aspects, the adjustment can be a function of the gene (or other target interval) being arrayed, the tumor type in the sample, the variant being arrayed, or one or more of the characteristics of the sample or subject. The selection or use of alignment conditions that are individually adjusted for several target intervals being arrayed enables optimization of speed, sensitivity, and specificity. This method is particularly effective when the alignment of reads to a relatively large number of diverse target intervals is optimized.

[0370] In some embodiments, reads from each of X distinct target intervals are aligned with a distinct alignment method, where distinct target intervals (e.g., target intervals or expressed target intervals) means different from the other X - 1 target intervals, and distinct alignment methods means different from the other X - 1 alignment methods, and X is at least 2.

[0371] In one embodiment, target intervals from at least X genes, such as at least X genes from Tables 2A - 5B, are aligned by a unique alignment method, where X is 2, 3, 4, 5, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500 or more.

[0372] In one embodiment, the method includes selecting or using an alignment method for analyzing, e.g., aligning, reads, and the alignment method is (i) a tumor type, such as the tumor type in the sample; (ii) a gene in which the target interval (e.g., target interval or expressed target interval) being sequenced is located, or a gene type, such as a gene or gene type characterized by a mutation, e.g., a variant or mutant form, or a frequency of mutation; (iii) an analysis site (e.g., nucleotide position); (iv) the type of variant within the target interval being evaluated (e.g., target interval or expressed target interval), such as a substitution; (v) the type of sample, such as the samples described herein; and (vi) the sequence within or near the target interval being evaluated, such as the expected tendency for misalignment for the target interval (e.g., target interval or expressed target interval), e.g., the presence of repetitive sequences within or near the target interval (e.g., target interval or expressed target interval), is one or more or all functions of, or is selected in response to, or is optimized with respect to,

[0373] As mentioned elsewhere in this specification, in some embodiments, the method is particularly effective when the alignment of reads for a relatively large number of target intervals is optimized. Thus, in one embodiment, at least X unique alignment methods are used to analyze reads for at least X unique target intervals, where the unique means is different from the other X - 1, and X is 2, 3, 4, 5, 10, 15, 20, 30, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000 or more.

[0374] In one embodiment, target intervals from at least X genes from Tables 2A - 5B are analyzed, where X is 2, 3, 4, 5, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500 or more.

[0375] In one embodiment, a unique alignment method is applied to target intervals in each of at least 3, 5, 10, 20, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400 or 500 different genes.

[0376] In one embodiment, nucleotide values are assigned to nucleotide positions of at least 20, 40, 60, 80, 100, 120, 140, 160 or 180, 200, 300, 400, or 500 genes, such as the genes in Tables 2A - 5B. In one embodiment, a unique alignment method for the target interval is applied to each of at least 10, 20, 30, 40, or 50% of the analyzed genes.

[0377] The methods disclosed herein enable the rapid and efficient alignment of troublesome reads, e.g., reads having rearrangements. Thus, in embodiments where the reads for the target interval (e.g., the target interval or the represented target interval) include nucleotide positions involving rearrangements, e.g., translocations, the method may be appropriately adjusted and may include using an alignment method that includes the following. Selecting a rearrangement reference array for alignment with the reads, wherein the rearrangement reference array aligns with a rearrangement (in some embodiments, the reference array is not the same as a genomic rearrangement); and Comparing the reads with the rearrangement reference array, e.g., aligning them.

[0378] In some embodiments, different methods, e.g., alternative methods, are used to align difficult reads. These methods are particularly effective when the alignment of reads to a relatively large number of diverse target intervals is optimized. By way of example, a method of analyzing a sample is Performing a comparison, e.g., an alignment comparison, of the reads under a first set of parameters (e.g., a first mapping algorithm or a first reference array); and Determining whether the reads meet a first alignment criterion (e.g., the reads can be aligned with the first reference array, e.g., with a low number of mismatches); and If the reads do not meet the first alignment criterion, performing a second alignment comparison under a second set of parameters (e.g., a second mapping algorithm or a second reference array); and Optionally, determining whether the reads meet the second criterion (e.g., the reads can be aligned with the second reference array with less than a predetermined number of mismatches), and includes The second set of parameters includes a set of parameters that are more likely to result in an alignment of the reads with a variant, e.g., a rearrangement, e.g., an insertion, deletion or translocation, compared to the first set of parameters, e.g., the use of the second reference array.

[0379] In an embodiment, the alignment method from the section entitled "Alignment" in this specification is combined with a mutation calling method from the section entitled "Mutation Calling" in this specification and / or a target capture reagent from the section entitled "Target Capture Reagent" in this specification and / or a target capture reagent from the section entitled "Design and Construction of Target Capture Reagents" in this specification. This method can be applied to a set of target intervals from the section entitled "Gene Selection" in this specification and / or a sample from the section entitled "Sample" in this specification from a subject from the section entitled "Subject" in this specification.

[0380] Alignment is typically the process of matching reads to a location, e.g., a genomic location. Misalignment (e.g., the placement of base pairs from a short read at an incorrect location within the genome). For example, reads of alternative alleles can be shifted from the main pile-up of alternative allele reads, so misalignments due to the sequence context of reads around an actual cancer mutation (e.g., the presence of repetitive sequences) can result in a decrease in the sensitivity of mutation detection. If a problematic sequence situation occurs when there is no actual mutation, misalignment can introduce artifact reads of "mutated" alleles by placing the actual reads of the reference genome bases at the wrong location. Since mutation calling algorithms for multiplexed multi-gene analysis must be sensitive even to low-abundance mutations, these misalignments can increase the false positive discovery rate / decrease the specificity.

[0381] As discussed herein, a decrease in sensitivity to actual mutations can be addressed by evaluating (either manually or in an automated fashion) the quality of the alignment around the predicted mutation sites of the gene being analyzed. The sites to be evaluated can be obtained from a database of cancer mutations (e.g., COSMIC). Regions identified as problematic can be repaired using an algorithm selected to provide better performance in the relevant sequence context by alignment optimization (or realignment) using a slower but more accurate alignment algorithm such as the Smith-Waterman alignment. If the general alignment algorithm cannot improve the problem, a customized alignment approach can be created, for example, by adjusting the maximum difference mismatch penalty parameter for genes likely to contain substitutions, by adjusting specific mismatch penalty parameters based on specific mutation types common to a particular tumor type (e.g., C→T); or by adjusting specific mismatch penalty parameters based on specific mutation types common in a particular sample type (e.g., substitutions common in FFPE).

[0382] A decrease in the specificity of the evaluated gene region (increase in false positive rate) due to misalignment can be evaluated by manual or automated inspection of all mutation calls in the sequenced sample. Regions found to be prone to false mutation calls due to misalignment can receive the same alignment rescue as described above. If algorithmic improvements are not possible, "mutations" from the problem regions can be classified or screened out of the test panel.

[0383] The methods disclosed herein enable the use of multiple individually tailored alignment methods or algorithms to optimize performance in methods that rely on rearrangement, such as sequencing of target regions related to indels, particularly large-scale parallel sequencing of multiple diverse genetic events in a large number of diverse genes, such as a large number of diverse genes from a sample. In some embodiments, a multiple alignment method that is individually customized or adjusted for each of a number of rearrangements in different genes is used to analyze reads. In some embodiments, the adjustment can be a function of the target region(s) being sequenced (e.g., one or more genes), the tumor type associated with the sample, the variant being sequenced, or one or more characteristics of the sample or subject. This selection or use of alignment conditions that are finely tuned to the multiple target regions being sequenced enables optimization of speed, sensitivity, and specificity. This method is particularly effective when the alignment of reads to a relatively large number of diverse target regions is optimized. In embodiments, the method includes the use of an alignment method optimized for rearrangement and other alignment methods optimized for target regions not related to rearrangement.

[0384] In some embodiments, an alignment selector is used. As used herein, an "alignment selector" refers to a parameter that enables or directs the selection of an alignment method, such as an alignment algorithm or parameters, that can optimize the sequencing of a target region. An alignment selector can be specific to, or selected as a function of, for example, one or more of the following. 1. The sequence context of the target interval (e.g., the nucleotide position being evaluated) related to the tendency of misalignment of reads for the target interval, e.g., the sequence context. For example, the presence of sequence elements within or near the target interval being evaluated that are repeated elsewhere in the genome can cause misalignment, thereby potentially degrading performance. By selecting an algorithm or algorithm parameters that minimize misalignment, performance can be improved. In this case, the value of the alignment selector can be a function of the sequence situation, e.g., the presence or absence of sequences of a length that are repeated at least several times in the genome (or in the portion of the genome being analyzed). 2. The tumor type being analyzed. For example, a particular tumor type can be characterized by an increased deletion rate. Thus, performance can be improved by selecting an algorithm or algorithm parameters that are more sensitive to indels. In this case, the value of the alignment selector can be a function of the tumor type, e.g., it can be an identifier of the tumor type. In one embodiment, the value is the identity of the tumor type, e.g., a solid tumor or a hematological malignancy (or a pre-malignant tumor). 3. The gene or type of gene being analyzed, e.g., the gene or type of gene can be analyzed. By way of example, oncogenes are often characterized by substitutions or in-frame indels. Thus, performance is particularly sensitive to these variations, and can be improved by selecting an algorithm or algorithm parameters that are specific to these over others. Tumor suppressor genes are often characterized by frameshift indels. Thus, performance can be improved by selecting an algorithm or algorithm parameters that are particularly sensitive to these variations. Thus, performance can be improved by selecting an algorithm or algorithm parameters that match the target interval. In this case, the value of the alignment selector can be a function of the gene or genotype, e.g., it can be an identifier of the gene or genotype. In one embodiment, the value is the identity of the gene. 4. The site being analyzed (e.g., nucleotide position). In this case, the value of the alignment selector can be a function of the site or type of site, e.g., an identifier for the site or site type. In one embodiment, the value is the identity of the site. (For example, when the gene containing the site is highly homologous to another gene, a normal / high-speed short-read alignment algorithm (e.g., BWA) may have difficulty distinguishing the two genes and may even require a more powerful alignment method (Smith-Waterman) or assembly (ARACHNE). Similarly, when the gene sequence contains a low-complexity region (e.g., AAAAAA), a more focused alignment method may be required.) 5. The variant or type of variant associated with the target interval being evaluated. For example, it includes substitutions, insertions, deletions, translocations, or other rearrangements. Thus, performance can be improved by selecting an algorithm or algorithm parameters that are more sensitive to a particular variant type. In this case, the value of the alignment selector can be a function of the type of variant, e.g., an identifier for the type of variant. In one embodiment, the value is the identity of the type of variant, e.g., a substitution. 6. The type of sample, e.g., the samples described herein. The sample type / quality can affect the error (false observation of non-reference sequences) rate. Thus, performance can be improved by selecting an algorithm or algorithm parameters that accurately model the true error rate of the sample. In this case, the value of the alignment selector can be a function of the type of sample, e.g., an identifier for the sample type. In one embodiment, the value is the identification information of the sample type.

[0385] Generally, accurate detection of indel mutations is a matter of alignment because the false indel rate on sequencing platforms that are invalidated herein is relatively low (thus, even a few observations of correctly aligned indels can be strong evidence of a mutation). However, accurate alignment in the presence of indels can be difficult (especially as indel length increases). In addition to general problems associated with alignment, such as substitutions, the indels themselves can cause alignment problems. (For example, a deletion of a 2bp dinucleotide repeat cannot be easily and definitively placed.) False placement of shorter (<15bp) apparent indel-containing reads can reduce both sensitivity and specificity. Larger indels (approaching the length of an individual read, e.g., a 36bp read) may be unable to align the read at all, making indel detection impossible in a standard set of aligned reads.

[0386] A cancer mutation database can be used to address these problems and improve performance. To reduce false positive indel discoveries (improve specificity), regions around generally expected indels can be examined for alignment problems due to sequence context and addressed similarly to the substitutions above. To improve the sensitivity of indel detection, several different approaches can be used that utilize information about indels expected in cancer. For example, short reads containing the expected indels can be simulated and alignment attempted. The alignment can be examined, and problematic indel regions can have alignment parameters adjusted, for example, by reducing gap open / extension penalties or by aligning partial reads (e.g., the first or second half of the read).

[0387] Alternatively, the initial alignment can also be attempted with alternative versions of the genome that include not only the normal reference genome but also each of the known or likely cancer indel mutations. In this approach, reads of indels that initially failed to align or were misaligned are successfully placed in the alternative (mutated) version of the genome.

[0388] In this way, indel alignment (and thus calling) can be optimized for the expected cancer genes / sites. As used herein, a sequence alignment algorithm embodies a computational method or approach used to identify where a read sequence (e.g., a short sequence, e.g., a short sequence from next-generation sequencing identification) is most likely to originate from by evaluating the similarity between the read sequence and the reference sequence. Various algorithms can be applied to the sequence alignment problem. Some algorithms are relatively slow but allow for relatively high specificity. These include, for example, dynamic programming-based algorithms. Dynamic programming is a way of solving complex problems by breaking them down into simpler steps. Other approaches are relatively efficient but typically not as thorough. These include, for example, heuristic algorithms and probabilistic methods designed for large database searches.

[0389] Alignment parameters are used in an alignment algorithm to adjust the performance of the algorithm, for example, to yield an optimal global or local alignment between the read sequence and the reference sequence. Alignment parameters can assign weights to matches, mismatches, and indels. For example, lower weights allow for alignments with more mismatches and indels.

[0390] The situation of the array, such as the presence of repetitive arrays (e.g., tandem repeats, interspersed repeats), low-complexity regions, indels, pseudogenes or paralogs, can affect alignment specificity (e.g., causing misalignment). As used herein, a misalignment refers to the arrangement of base pairs from short reads at incorrect positions within the genome.

[0391] When an alignment algorithm is selected or alignment parameters are adjusted based on a tumor type, e.g., a tumor type having a tendency to have a specific mutation or mutation type, the sensitivity of the alignment can be enhanced.

[0392] Selecting an alignment algorithm or adjusting alignment parameters based on a specific genotype (e.g., oncogene, tumor suppressor gene) can enhance the sensitivity of the alignment. Mutations in different types of cancer-related genes can have different effects on the cancer phenotype. For example, mutant oncogene alleles are typically dominant. Mutant tumor suppressor alleles are typically recessive, which means that in most cases both alleles of the tumor suppressor gene must be affected before an effect is manifested.

[0393] The sensitivity of the alignment can be adjusted (e.g., increased) when an alignment algorithm is selected or when alignment parameters are adjusted based on a mutant type (e.g., single nucleotide polymorphism, indel (insertion or deletion), inversion, translocation, tandem repeat).

[0394] When an alignment algorithm is selected or alignment parameters are adjusted based on a mutation site (e.g., mutation hot spot), the sensitivity of the alignment can be adjusted (e.g., increased). A mutation hot spot refers to a site within the genome where mutations occur at a frequency up to 100 times the normal mutation rate.

[0395] When an alignment algorithm is selected or when alignment parameters are adjusted based on a sample type (e.g., cfDNA sample, ctDNA sample, FFPE sample, or CTC sample), the sensitivity / specificity of the alignment can be adjusted (e.g., increased).

[0396] In some embodiments, NGS reads can be aligned to a known reference sequence or assembled de novo. For example, NGS reads can be aligned to a reference sequence (e.g., a wild-type sequence). Methods for sequence alignment for NGS are described, for example, in Trapnell C. and Salzberg S.L. Nature Biotech, 2009, 27:455-457. Examples of de novo assembly are described, for example, in Warren R. et al., Bioinformatics, 2007, 23:500-501, Butler J. et al., Genome Res., 2008, 18:810-820; and Zerbino D.R. and Birney E., Genome Res., 2008, 18:821-829. Sequence alignment or assembly can be performed using read data from one or more NGS platforms, such as by mixing Roche / 454 and Illumina / Solexa read data.

[0397] Optimization of alignment is described in the art, for example, as described in International Patent Application Publication No. 2012 / 092426.

[0398] Mutation calling The methods disclosed herein can incorporate the use of customized or adjusted mutation calling parameters to optimize performance in sequence identification methods, particularly methods that rely on large-scale parallel sequence identification of a large number of diverse genes, such as those derived from cancer as described herein, e.g., a large number of diverse gene events from a sample.

[0399] In some embodiments, the mutation calls for each of a number of target intervals are not individually customized or fine-tuned. In some embodiments, the mutation calls for at least a subset of some of the target intervals are individually customized or fine-tuned. In some embodiments, the mutation calls for each of some of the target intervals are individually customized or fine-tuned. The customization or adjustment can be based on factors described herein, such as the type of cancer in the sample, the gene in which the targeted interval to be sequenced is located, or one or more of the identified variants. This selection or use of alignment conditions fine-tuned to a number of targeted intervals to be sequenced enables optimization of speed, sensitivity, and specificity. This method is particularly effective when the alignment of reads to a relatively large number of diverse target intervals is optimized.

[0400] In some embodiments, nucleotide values are assigned for nucleotide positions in each of X distinct target intervals, where distinct target intervals (meaning different from the other X - 1 target intervals, e.g., sub-genomic intervals, expression sub-genomic intervals, or both), distinct calling methods (meaning different from the other X - 1 calling methods), and X is at least 2. The calling methods are different and can thereby be unique, for example, by depending on different Bayesian prior values.

[0401] In one embodiment, assigning the nucleotide value is a function of an expected value or a value representing an expected value (e.g., from the literature) of a variant, e.g., a lead indicating a mutation, at the nucleotide position in the type of tumor.

[0402] In one embodiment, the method includes assigning nucleotide values (e.g., mutation calls) for at least 10, 20, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1,000 nucleotide positions, where each assignment is a function of an expected value (e.g., from the literature) or a unique value (as contrasted with the values of other assignments) that represents whether a variant, e.g., a mutation, is present at the nucleotide position in the type of tumor before observing a read indicating the variant.

[0403] In one embodiment, assigning the nucleotide values is a function of a set of values representing the probability of observing a read indicating the variant at the nucleotide position when the variant is present in the sample at a frequency (e.g., 1%, 5%, 10%, etc.) and / or when the variant is not present (e.g., observed in a read due to base calling error only).

[0404] In one embodiment, the mutation calling method described herein includes the following steps, namely, for each nucleotide position in each of the X target intervals, (i) a first value that is an expected value (e.g., from the literature) or represents an expected value of observing a read indicating a variant, e.g., a mutation, at the nucleotide position within the X type of tumor; and (ii) a second set of values representing the probability of observing a read indicating the variant at the nucleotide position when the variant is present in the sample at a frequency (e.g., 1%, 5%, 10%, etc.) and / or when the variant is not present (e.g., observed in a read due to base calling error only); obtaining responding to the values, e.g., by weighing a comparison between the values in the second set using the first value (e.g., calculating the posterior probability of the presence of a mutation) by the Bayesian method described herein, and analyzing the sample by assigning a nucleotide value (e.g., a mutation call) from the read for each of the nucleotide positions.

[0405] In one embodiment, the method comprises: (i) assigning nucleotide values (e.g., mutation calls) to at least 10, 20, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1,000 nucleotide positions, each assignment being based on a unique (as opposed to other assignments) first and / or second value; (ii) assigning the method of (i), at least 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, or 500 of the assignments being made using a first value that is a function of the probability of a variant present in less than 5, 10, or 20% of the cells in the tumor type, for example; (iii) assigning nucleotide values (e.g., mutation calls) to at least X nucleotide positions, each of which is associated with a variant having a unique (as opposed to the other X - 1 assignments) probability of being present in a tumor of the type of the sample, e.g., tumor type, and optionally, each of the X assignments being based on a unique (as opposed to the other X - 1 assignments) first and / or second value (where X = 2, 3, 5, 10, 20, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, or 500); (iv) assigning nucleotide values (e.g., mutation calls) to a first and a second nucleotide position, wherein the probability of a first variant at the first nucleotide position being present in a tumor of a type (e.g., the tumor type of the sample) is at least 2, 5, 10, 20, 30, or 40 times greater than the probability of a second variant at the second nucleotide position being present, and optionally, each assignment being based on a unique (as opposed to other assignments) first and / or second value; (v) Assigning nucleotide values to a plurality of nucleotide positions (e.g., calling mutations), wherein the plurality includes one or more, e.g., at least 3, 4, 5, 6, 7, or all of the following probability percentage ranges: 0.01 or less; greater than 0.01 and 0.02 or less, greater than 0.02 and 0.03 or less, greater than 0.03 and 0.04 or less, greater than 0.04 and 0.05 or less, greater than 0.05 and 0.1 or less, greater than 0.1 and 0.2 or less, greater than 0.2 and 0.5 or less, greater than 0.5 and 1.0 or less, greater than 1.0 and 2.0 or less, greater than 2.0 and 5.0 or less, greater than 5.0 and 10.0 or less, greater than 10.0 and 20.0 or less, greater than 20.0 and 50.0 or less, greater than 50 and 100.0% or less, including the assignment of mutants classified in all cases, The probability range is the range of the probability that a mutant at a nucleotide position exists in a tumor type (e.g., the tumor type of the sample), or the probability that a mutant at a nucleotide position exists in the sample, a library from the sample, or the enumerated percentage (%) of cells in a library catch from that library, for a preselected type (e.g., the tumor type of the sample), Optionally, each assignment is based on unique first and / or second values (e.g., unique compared to other assignments within the enumerated probability range, or unique compared to one or more or all of the first and / or second values of other enumerated probability ranges), assigning; (vi) Independently for each of at least 1, 2, 3, 5, 40, 25, 15, 10, 5, 4, 3, 2, 1, 0.5, 0.4, 0.3, 1,000, or 0.2 nucleotide positions having mutants present in 50, 10, 20, 40, 20, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or less than 0.1% of the DNA in the sample, assigning nucleotide values (e.g., calling mutations), optionally, each assignment is based on unique first and / or second values (compared to other assignments); (vii) Assigning nucleotide values (e.g., mutation calls) to the first and second nucleotide positions, wherein the likelihood of a variant at the first position in the DNA of the sample is at least 2, 5, 10, 20, 30, or 40 times greater than the likelihood of a variant at the second nucleotide position in the DNA of the sample, and optionally, each assignment is based on unique first and / or second values (as opposed to other assignments); (viii) Assigning nucleotide values (e.g., mutation calls) to one or more or all of the following: (1) At least 1, 2, 3, 4, or 5 nucleotide positions having variants present in less than 1% of the cells in the sample, nucleic acids in a library from the sample, or nucleic acids in a library catch from that library; (2) At least 1, 2, 3, 4, or 5 nucleotide positions having variants present in 1-2% of the cells in the sample, nucleic acids in a library from the sample, or nucleic acids in a library catch from that library; (3) At least 1, 2, 3, 4, or 5 nucleotide positions having variants present in more than 2% and up to 3% of the cells in the sample, nucleic acids in a library from the sample, or nucleic acids in a library catch from that library (4) At least 1, 2, 3, 4, or 5 nucleotide positions having variants present in more than 3% and up to 4% of the cells in the sample, nucleic acids in a library from the sample, or nucleic acids in a library catch from that library; (5) At least 1, 2, 3, 4, or 5 nucleotide positions having variants present in more than 4% and up to 5% of the cells in the sample, nucleic acids in a library from the sample, or nucleic acids in a library catch from that library; (6) At least 1, 2, 3, 4, or 5 nucleotide positions having variants present in more than 5% and up to 10% of the cells in the sample, nucleic acids in a library from the sample, or nucleic acids in a library catch from that library; (7) At least 1, 2, 3, 4, or 5 nucleotide positions having variants that are present at more than 10% and less than or equal to 20% of the cells in the sample, nucleic acids in the library from the sample, or nucleic acids in the library catch from that library; (8) At least 1, 2, 3, 4, or 5 nucleotide positions having variants that are present at more than 20% and less than or equal to 40% of the cells in the sample, nucleic acids in the library from the sample, or nucleic acids in the library catch from that library; (9) At least 1, 2, 3, 4, or 5 nucleotide positions having variants that are present at more than 40% and less than or equal to 50% of the cells in the sample, nucleic acids in the library from the sample, or nucleic acids in the library catch from that library; or (10) At least 1, 2, 3, 4, or 5 nucleotide positions having variants that are present at more than 50% and less than or equal to 100% of the cells in the sample, nucleic acids in the library from the sample, or nucleic acids in the library catch from that library; Optionally, each assignment is made based on a unique first and / or second value (e.g., unique compared to other assignments within the enumerated range (e.g., a range of less than 1% in (1)) or unique compared to the first and / or second values for determination in one or more or all of the other enumerated ranges), the act of assigning; (ix) The act of assigning a nucleotide value (e.g., a mutation call) to each of X nucleotide positions, where each nucleotide position independently has a likelihood (of a variant present in the DNA of the sample) that is unique compared to the likelihood of variants at the other X - 1 nucleotide positions, X is 1, 2, 3, 5, 10, 20, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1,000 or more, and each assignment is made based on a unique first and / or second value (compared to other assignments), including one or more or all of the above.

[0406] In some embodiments, a "threshold value" is used, for example, for calling mutations at specific positions in a gene, to evaluate reads and select values for nucleotide positions from those reads. In some embodiments, the threshold value for each of a number of target intervals is customized or fine-tuned. The customization or adjustment can be based on one or more of the factors described herein, such as the type of cancer in the sample, the gene in which the sequenced target interval (sub-genomic interval or expression sub-genomic interval) is located, or the sequenced variant. This provides finely tuned calls for each of the number of target intervals to be sequenced. In some embodiments, the method is particularly effective when a relatively large number of diverse sub-genomic intervals are analyzed.

[0407] Thus, in another embodiment, the method includes the following mutation calling method: For each of the X target intervals, obtaining a threshold value, each of the obtained X threshold values being unique compared to the other X - 1 threshold values, thereby providing X unique threshold values; For each of the X target intervals, comparing an observed value, which is a function of the number of reads having a nucleotide value at a nucleotide position, to its unique threshold value, thereby applying its unique threshold value to each of the X target intervals; Optionally, in response to the result of the comparison, assigning a nucleotide value to the nucleotide position, where X is 2 or more, and the assigning, a method.

[0408] In one embodiment, the method includes assigning a nucleotide value to at least 2, 3, 5, 10, 20, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900 or 1,000 nucleotide positions, each independently having a first value that is a function of a probability less than 0.5, 0.4, 0.25, 0.15, 0.10, 0.05, 0.04, 0.03, 0.02 or 0.01.

[0409] In one embodiment, the method comprises assigning a nucleotide value to each of at least X nucleotide positions, each having, independently of the other X-1 first values, a unique first value, each of the X first values being a function of a probability of less than 0.5, 0.4, 0.25, 0.15, 0.10, 0.05, 0.04, 0.03, 0.02, or 0.01, where X is 1, 2, 3, 5, 10, 20, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1,000 or more.

[0410] In one embodiment, nucleotide values are assigned to nucleotide positions of at least 20, 40, 60, 80, 100, 120, 140, 160 or 180, 200, 300, 400, or 500 genes, such as the genes of Tables 2A to 5B. In one embodiment, unique first and / or second values are applied to target intervals in at least 10, 20, 30, 40 or 50% of each of the genes analyzed.

[0411] Embodiments of the method can be applied, for example, when the thresholds for a relatively large number of target intervals are optimized, as can be seen from the following embodiments.

[0412] In one embodiment, a unique threshold is applied to a target interval, such as a sub-genomic interval or an expression sub-genomic interval, in each of at least 3, 5, 10, 20, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900 or 1,000 different genes.

[0413] In one embodiment, nucleotide values are assigned to the nucleotide positions of at least 20, 40, 60, 80, 100, 120, 140, 160, or 180, 200, 300, 400, or 500 genes, such as the genes in Tables 2A - 5B. In one embodiment, unique thresholds are applied to sub-genomic intervals in at least 10, 20, 30, 40, or 50% of each of the analyzed genes.

[0414] In one embodiment, nucleotide values are assigned to the nucleotide positions of at least 5, 10, 20, 30, or 40 genes in Tables 2A - 5B. In one embodiment, unique thresholds for a target interval (e.g., a sub-genomic interval or an expression sub-genomic interval) are applied in at least 10, 20, 30, 40, or 50% of each of the analyzed genes.

[0415] The elements of that module can be included in a method for analyzing a tumor. In embodiments, alignment methods from the section entitled "Mutation Calling" can be combined with alignment methods from the section entitled "Alignment" and / or target capture reagents from the section entitled "Target Capture Reagents" and / or the sections entitled "Design and Construction of Target Capture Reagents" and "Competition of Target Capture Reagents" herein. The method can be applied to a set of target intervals from the section entitled "Gene Selection" herein and / or a sample from the section entitled "Sample" from a subject from the section entitled "Subject" herein.

[0416] Base calling refers to the raw output of a sequencing device. Mutation calling refers to the process of selecting a nucleotide value, e.g., A, G, T, or C, for a nucleotide position for which the sequence has been determined. Typically, the sequence-determined reads (or base calls) for a position provide more than one value; e.g., some reads give T and some reads give G. Mutation calling is the process of assigning a nucleotide value, e.g., one of those values, to the sequence. Although called "mutation" calling, it can be applied to assign nucleotide values to any nucleotide position, e.g., positions corresponding to mutant alleles, wild-type alleles, alleles not characterized as mutant or wild-type, or positions not characterized by variability. Methods for mutation calling can include one or more of the following: making independent calls based on information at each position within a reference sequence (e.g., examining sequence reads; examining base calls and quality scores; calculating the probability of the observed bases and quality scores given a potential genotype; and assigning a genotype (e.g., using Bayes' rule)); removing false positives (e.g., using a depth threshold to reject SNPs with read depths much lower or higher than expected; local realignment to remove false positives due to small indels); performing linkage disequilibrium (LD) / imputation-based analysis to improve the calls.

[0417] Equations for calculating genotype likelihoods related to specific genotypes and positions are described, for example, in Li H. and Durbin R., Bioinformatics, 2010;26(5):589-95. Prior expectations for specific mutations in a specific cancer type can be used when evaluating samples from that cancer type. Such possibilities can be obtained from public databases of cancer mutations, such as Catalogue of Somatic Mutation in Cancer (COSMIC), HGMD (Human Gene Mutation Database), The SNP Consortium, Breast Cancer Mutation Data Base (BIC), and Breast Cancer Gene Database (BCGD).

[0418] Examples of LD / imputation-based analysis are described, for example, in Browning B.L. and Yu Z., Am. J. Hum. Genet. 2009, 85(6):847-61. See also Li Y. et al., Annu. Rev. Genomics Hum. Genet. 2009, 10:387-406 for examples of low-coverage SNP calling methods.

[0419] After alignment, substitution detection can be performed using a calling method, such as a Bayesian mutation calling method. This is applied to each base of each target interval, for example, the exons of the gene being evaluated, and the presence of alternative alleles is observed. This method compares the probability of observing read data in the presence of a mutation to the probability of observing read data in the presence of basecall errors only. If this comparison strongly supports the presence of a mutation, the mutation can be called.

[0420] Methods have been developed to address limited deviations from 50% or 100% frequencies for the analysis of cancer DNA. (For example, SNVMix Bioinformatics. March 15, 2010; 26(6):730-736) However, the methods disclosed herein allow for consideration of the possibility that mutant alleles exist anywhere between 1% and 100% of the sample DNA, particularly at levels below 50%. This approach is particularly important for the detection of mutations in low-purity FFPE samples of native (multiclonal) tumor DNA.

[0421] The advantage of the Bayesian mutation detection approach is that the comparison of only the probability of the presence of a mutation and the probability of a base calling error can be weighted by the prior expectation of the presence of a mutation at that site. If several reads of alternative alleles are observed at sites that are frequently mutated for a given cancer type, the presence of a mutation can be reliably called even if the amount of evidence for the mutation does not meet the normal threshold. This flexibility can then be used to increase the detection sensitivity for rarer mutations / lower purity samples or to make the test more robust to a decrease in read coverage. The probability that a random base pair in the genome is mutated in cancer is approximately 1e-6. The probability of a specific mutation at many sites in a typical polygenic cancer genome panel can be orders of magnitude higher. These likelihoods can be obtained from public databases of cancer mutations (e.g., COSMIC). Indel calling is the process of finding bases in sequence-specific data that differ from the reference sequence by an insertion or deletion, typically including a relevant confidence score or statistical evidence metric.

[0422] The method of indel calling may include the steps of identifying candidate indels, calculating genotype likelihoods by local realignment, and performing LD-based genotype inference and calling. Typically, the Bayesian method is used to obtain potential indel candidates, which are then tested with the reference sequence within the Bayesian framework.

[0423] Algorithms for generating candidate indels can be found, for example, in McKenna A. et al., Genome Res. 2010;20(9):1297-303; Ye K. et al., Bioinformatics, 2009;25(21):2865-71; Lunter G. and Goodson M., Genome Res., 2011;21(6):936-9; and Li H. et al., Bioinformatics 2009, Bioinformatics 25(16):2078-9.

[0424] Examples of methods for generating indel calls and individual-level genotype likelihoods include, for example, the Dindel algorithm (Albers C.A. et al., Genome Res. 2011;21(6):961-73). For example, using the Bayesian EM algorithm, reads can be analyzed to make an initial indel call, genotype likelihoods can be generated for each candidate indel, and then genotypes can be complemented, for example, using QCALL (Le S.Q. and Durbin R., Genome Res. 2011;21(6):952-60). Parameters such as prior expectations of observing indels can be adjusted based on the size or position of the indel (e.g., increased or decreased).

[0425] In one embodiment, at least 10, 20, 30, 40, 50, 60, 70, 80, or 90% of the mutation calls made by this method are for target intervals from the genes or gene products described herein, such as the genes or gene products in Tables 2A - 5B. In one embodiment, at least 10, 20, 30, 40, 50, 60, 70, 80, or 90% of the unique thresholds described herein are for target intervals from the genes or gene products described herein, such as the genes or gene products in Tables 2A - 5B. In one embodiment, at least 10, 20, 30, 40, 50, 60, 70, 80, or 90% of the annotated or third-party-reported mutation calls are for target intervals from the genes or gene products described herein, such as the genes or gene products in Tables 2A - 5B.

[0426] In one embodiment, the assigned value for the nucleotide position is optionally transmitted to a third party with explanatory annotations. In one embodiment, the assigned value for the nucleotide position is not transmitted to a third party. In one embodiment, the assigned values for a plurality of nucleotide positions are optionally transmitted to a third party with explanatory annotations, and the assigned values for a second plurality of nucleotide positions are not transmitted to a third party.

[0427] In one embodiment, the method includes assigning one or more reads, for example, by barcode deconvolution.

[0428] In one embodiment, the method includes assigning one or more reads as tumor reads or control reads, for example, by barcode backconvolution. In one embodiment, the method includes mapping each of the one or more reads, for example, by alignment with a reference sequence. In one embodiment, the method includes storing the called mutations.

[0429] In one embodiment, the method includes annotating so-called mutations, for example, annotating so-called mutations with mutation structures, for example, missense mutations, or functions, for example, disease phenotypes. In one embodiment, the method includes obtaining nucleotide sequence reads for tumor nucleic acids and control nucleic acids. In one embodiment, the method includes calling nucleotide values, for example, variants, for example, mutations, for each of the target intervals (e.g., subgenomic intervals, expression subgenomic intervals, or both), for example, using a Bayesian calling method or a non-Bayesian calling method. In one embodiment, the method includes evaluating a plurality of reads containing at least one SNP. In one embodiment, the method includes identifying SNP allele ratios in the sample and / or control reads.

[0430] In some embodiments, the method further includes constructing a database of alignment artifacts for a target subgenomic region. In one embodiment, the database can be used to exclude false mutation calls and improve specificity. In one embodiment, the database is constructed by sequencing unrelated samples or cell lines and recording non-reference allele events that occur more frequently than expected due to random sequencing errors in one or more of these normal samples. This approach can classify germline mutations as artifacts, but is acceptable in methods for somatic mutations. This incorrect classification of germline mutations as artifacts can be improved, if necessary, by filtering this database for known germline mutations (removing common variants) and for artifacts that appear in only one individual (removing rarer variants).

[0431] Optimization of mutation calling is described in the art, for example, as described in International Patent Application Publication No. WO 2012 / 092426.

[0432] SGZ algorithm Various types of changes, such as somatic changes and germline mutations, can be detected by the methods described herein (e.g., sequencing, alignment, or mutation calling methods). In certain embodiments, germline mutations are further identified by methods using the SGZ (Somatic-Germline-Zygosity) algorithm. See, for example, U.S. Patent No. 9,792,403 and Sun et al., A computational approach to distinguish somatic vs. germline origin genomic alteration from a deep sequencing of cancer specimens without matched normal, PLOS Computational Biology (February 2018).

[0433] In clinical diagnosis and treatment, matched normal controls are generally not available. In some embodiments, well-characterized genomic changes do not require normal tissue for interpret...

Claims

1. A method for determining the tumor fraction of a sample from a subject, comprising: obtaining a plurality of values, each value indicating an allele fraction at a corresponding locus within a subgenomic interval in the sample; determining a confidence metric indicative of the variance of the plurality of values; accessing a predetermined relationship between one or more stored confidence metrics and one or more stored tumor fraction values stored in a memory; and determining the tumor fraction of the sample based on the determined confidence metric and the corresponding tumor fraction in the memory corresponding to the predetermined confidence metric. A method comprising the above.

2. The method of claim 1, wherein (a) each value among the plurality of values is an allele fraction, or (b) each value among the plurality of values comprises a ratio of the abundance difference between a maternal allele and a paternal allele to the abundance of either the maternal allele or the paternal allele at the corresponding locus.

3. The method of claim 1 or 2, wherein the confidence metric indicates a deviation from an expected value for each of the plurality of values.

4. The method of claim 3, wherein (a) the expected value is a locus-specific expected value, (b) the confidence metric is a root mean square deviation from the expected value, (c) the expected value is an expected allele frequency for a non-tumor sample, or (d) each value among the plurality of values is an allele fraction and the expected value is 0.

5.

5. The method of claim 3 or 4, wherein each value among the plurality of values is a ratio of the abundance difference between a maternal allele and a paternal allele to the abundance of either the maternal allele or the paternal allele at the corresponding locus, the expected value comprises the expected ratio of the abundance difference between the maternal allele and the paternal allele to the abundance of either the maternal allele or the paternal allele, and the expected value is the expected ratio for a non-tumor sample.

6. The method of claim 5, wherein the expected value is 0.

7. The method according to any one of claims 1 to 6, wherein the plurality of values includes a plurality of allele coverages.

8. The method of claim 1, further comprising determining a probability distribution function of the plurality of values, wherein the confidence metric is determined using the probability distribution function.

9. The method of claim 8, wherein the confidence metric is the entropy of the probability distribution function.

10. (a) wherein the corresponding locus comprises one or more loci having different maternal and paternal alleles; (b) wherein the corresponding locus consists of loci having different maternal and paternal alleles, or (c) wherein the corresponding locus comprises one or more loci having the same maternal and paternal alleles, the method according to any one of claims 1 to 9.

11. A method for determining the tumor fraction of a sample from a subject, comprising: obtaining a plurality of values, each value indicating a difference between the allele coverage of a locus in a tumor sample and the allele coverage of the same locus in a non-tumor sample at a plurality of loci within a sub-genomic interval; determining an accuracy metric indicative of the variance of the plurality of values; accessing a predetermined relationship between one or more stored accuracy metrics and one or more stored tumor fraction values stored in a memory; and determining the tumor fraction of the sample based on the determined accuracy metric corresponding to the predetermined accuracy metric and the corresponding tumor fraction in the memory. A method comprising:

12. (a) each value among the plurality of values comprises a ratio of the allele coverage of a locus in the tumor sample compared to the allele coverage of the same locus in the non-tumor sample; (b) each value among the plurality of values comprises a log ratio of the allele coverage of a locus in the tumor sample compared to the allele coverage of the same locus in the non-tumor sample; (c) each value among the plurality of values includes the log ratio of the allelic coverage of the locus in the tumor sample compared to the allelic coverage of the same locus in the non-tumor sample, and the log ratio is log 2 ratio (d) each value among the plurality of values comprises a ratio of the difference in allele coverage of the locus in the tumor sample and the same locus in the non-tumor sample to the allele coverage of the same locus in the non-tumor sample; (e) the accuracy metric indicates the deviation of each value among the plurality of values from an expected value across the corresponding locus, the expected value being the value expected if the tumor sample were a non-tumor sample; (f) the accuracy metric indicates the deviation of each value among the plurality of values from an expected value across the corresponding locus, the expected value being the value expected if the tumor sample were a non-tumor sample, each value comprises a ratio of the allele coverage of a locus in the tumor sample compared to the allele coverage of the same locus in the non-tumor sample, and the expected value is 1. (g) The accuracy index indicates the deviation of each value among the plurality of values from the expected value across the corresponding locus, the expected value being the value expected if the tumor sample were a non-tumor sample, each value including the log ratio of the allelic coverage of the locus in the tumor sample compared to the allelic coverage of the same locus in the non-tumor sample, and the expected value being 0. (h) The accuracy index indicates the deviation of each value among the plurality of values from the expected value across the corresponding locus, the expected value being the value expected if the tumor sample were a non-tumor sample, each value including the ratio of the difference in allelic coverage of the locus in the tumor sample and the allelic coverage of the same locus in the non-tumor sample to the allelic coverage of the same locus in the non-tumor sample, and the expected value being 0. (i) The accuracy index is the root mean square deviation from the expected value. (j) Further comprising specifying a probability distribution function of the plurality of values, the accuracy index being specified using the probability distribution function. (k) Further comprising specifying a probability distribution function of the plurality of values, the accuracy index being specified using the probability distribution function, and the accuracy index being the entropy of the probability distribution function. (l) The allelic coverage includes the allelic coverage of maternal alleles and paternal alleles, or (m) The allelic coverage consists of the allelic coverage of maternal alleles and paternal alleles. The method according to claim 11.

13. The method according to claim 11 or 12, wherein the plurality of loci includes at least one nucleotide associated with a single nucleotide polymorphism (SNP).

14. (a) The plurality of loci each includes two or more nucleotides each associated with a single nucleotide polymorphism (SNP), or (b) The SNP is associated with cancer. The method according to claim 13.

15. The method according to any one of claims 11 to 14, wherein at least a part of the plurality of loci is associated with a copy number variation (CNV).

16. The method according to claim 15, wherein the CNV is associated with cancer.

17. (a) Further comprising sequencing the sample to identify the abundance or coverage of alleles at each locus. (b) further comprising performing array hybridization on the sample to identify the abundance or coverage of alleles at each locus (c) accessing a training dataset that includes a plurality of relationships between a plurality of training accuracy metrics and associated training tumor fractions applying a machine learning process to the training dataset to identify the predetermined relationship between the training accuracy metric and the training tumor fraction further comprising (d) further comprising generating a report that includes information identifying the subject and the identified tumor fraction (e) further comprising generating a report that includes information identifying the subject and the identified tumor fraction, and providing the report to the subject or a healthcare provider, or (f) further comprising generating a report that includes information identifying the subject and the identified tumor fraction, and formatting the report for an electronic health record The method according to any one of claims 1 to 16.

18. A method for assisting in monitoring tumor progression or recurrence in a subject, comprising: (a) identifying a first tumor fraction of a first sample obtained from the subject at a first time point according to the method according to any one of claims 1 to 17; (b) identifying a second tumor fraction of a second sample obtained from the subject at a second time point; (c) comparing the first tumor fraction with the second tumor fraction to thereby monitor the tumor progression A method comprising

19. (a) wherein identifying the second tumor fraction is obtaining a second plurality of values, each value indicating an allele fraction at a corresponding locus within a subgenomic interval in a second tumor sample, wherein the subgenomic interval in the second sample is the same as or different from the subgenomic interval in the first sample, obtaining; identifying a second accuracy metric indicating the variance of the second plurality of values; accessing a predetermined relationship between one or more stored accuracy metrics and one or more stored tumor fractions; identifying the second tumor fraction of the second sample from the second accuracy metric and the predetermined relationship comprising, or (b) wherein identifying the second tumor fraction is Obtaining a second plurality of values, each value indicating a difference between the allelic coverage of loci in a second tumor sample at a plurality of loci within a sub-genomic interval in the sample and the allelic coverage of the same loci in a non-tumor sample, and the sub-genomic interval used to identify the second tumor fraction being the same as or different from the sub-genomic interval used to identify the first tumor fraction, and obtaining; Identifying a second accuracy metric indicative of the variance of the second plurality of values; Accessing a predetermined relationship between one or more stored accuracy metrics and one or more stored tumor fractions; Identifying the second tumor fraction of the second tumor sample from the second accuracy metric and the predetermined relationship The method according to claim 18, comprising.

20. The method according to claim 18 or 19, wherein the first time point is before the subject undergoes tumor therapy, and the second time point is after the subject has undergone the tumor therapy.

21. The method according to any one of claims 1 to 20, wherein the subject has cancer, is at risk of having cancer, or is suspected of having cancer.

22. (a) the cancer is a solid tumor, or (b) the cancer is a blood cancer, the method according to claim 21.

23. (a) the sample is a liquid sample, (b) the sample is a solid sample, or (c) the sample contains cell-free DNA (cfDNA) or circulating tumor DNA (ctDNA), the method according to any one of claims 1 to 22.

24. The method according to any one of claims 1 to 23, wherein the one or more stored accuracy metrics include a plurality of stored accuracy metrics, and the one or more stored tumor fractions include a plurality of stored tumor fractions.

25. A computer system, a processor; a memory communicatively coupled to the processor, a predetermined relationship between one or more stored accuracy metrics and one or more associated stored tumor fraction values; and when executed by the processor, cause the processor to, (a) (i) obtaining a plurality of values, each value indicating an allele fraction at a corresponding locus within a sub-genomic interval in the sample, or (ii) obtaining a plurality of values, each value indicating a difference between the allele coverage of a locus in a tumor sample and the allele coverage of the same locus in a non-tumor sample at a plurality of loci within a sub-genomic interval; (b) specifying a confidence metric indicative of the variance of the plurality of values; (c) accessing the stored predetermined relationship; and (d) specifying the tumor fraction of the sample based on the specified confidence metric corresponding to a predetermined confidence metric and a corresponding tumor fraction in memory, a memory storing instructions to cause the execution of: comprising, and (a) when executed by the processor, the memory causes the processor to access a training dataset including a plurality of relationships between a plurality of training confidence metrics and associated training tumor fractions; and apply a machine learning process to the training dataset to specify a predetermined relationship between the training confidence metric and the training tumor fraction, further comprising instructions to cause the execution of, and / or (b) when executed by the processor, the instructions cause the processor to execute the method according to any one of claims 1 to 24, a computer system.

26. (a) when executed by the processor, the memory causes the processor to access a training dataset including a plurality of relationships between a plurality of training confidence metrics and associated training tumor fractions; and apply a machine learning process to the training dataset to specify a predetermined relationship between the training confidence metric and the training tumor fraction, further comprising instructions to cause the execution of, and (b) when executed by the processor, the instructions cause the processor to execute the method according to any one of claims 1 to 24, the computer system according to claim 25.

Citation Information

Patent Citations

  • Detection of mutations and ploidy of chromosome segments

    JP2017519488A

  • Molecular quality assurance methods for use in sequencing

    JP2018536430A

  • Methods and systems for assessing tumor mutation burden

    JP2019512218A

  • Therapeutic and diagnostic methods for cancer

    WO2019018757A1