Targeted depletion sequencing for use in minimum residual disease assays
By depleting low-utility DNA regions and enriching for tumor-specific mutations, the method improves the sensitivity and specificity of MRD assays, enabling effective monitoring of cancer recurrence and treatment response.
Patent Information
- Application Number
- PCT/US2025/037594
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-16
- Filing Date
- 2025-07-14
- Publication Date
- 2026-01-22
AI Technical Summary
Existing methods for detecting minimal residual disease (MRD) in cancer patients face challenges due to low levels of cell-free DNA from tumor cells, leading to hindered detection of disease-associated mutations.
The method involves obtaining DNA samples from tumor and non-tumor sources, depleting regions of low utility, and using oligonucleotide probes to enrich for tumor-specific somatic mutations, thereby generating a patient-specific panel for improved MRD assays.
This approach enhances the sensitivity and specificity of MRD assays by efficiently identifying tumor-specific mutations, allowing for better monitoring of cancer recurrence and treatment response.
Smart Images

Figure US2025037594_22012026_PF_FP_ABST
Abstract
Description
TARGETED DEPLETION SEQUENCING FOR USE IN MINIMUM RESIDUAL DISEASE ASSAYSCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims the benefit of U.S. provisional application No. 63 / 672,178, filed on July 16, 2024, the entire disclosure of which is incorporated by reference herein.TECHNICAL FIELD
[0002] Described herein are methods of improving the sensitivity and specificity of tumor-informed minimal residual disease (MRD) assays and panels. More specifically, the present disclosure provides methods for more efficiently and effectively preparing samples for patient-specific panels MRD assays and related analysis.BACKGROUND
[0003] The following description of the background of the present technology is provided simply as an aid in understanding the present technology and is not admitted to describe or constitute prior art to the present technology.
[0004] The discovery of cell-free deoxyribonucleic acid (cfDNA) has promoted the non-invasive detection of alterations in genomic sequences that occur in various disease states and patient populations. However, in some instances, e.g., cancer, the ability to determine the presence of a mutation associated with a disease phenotype has been hindered by the extremely low levels of cell-free DNA from the target cell population, e.g., tumor cells. Thus, methods that allow for enhanced and accurate detection of disease-associated mutations in DNA samples comprising the low levels of the cfDNA and other sources of DNA remain desirable.SUMMARY
[0005] The present disclosure provides methods of reducing error and improving performance of minimal residual disease (MRD) assays through the targeted deletion of sequences of low utility for the assays or the removal of excess DNA from patient samples.
[0006] In a first aspect, the present disclosure provides methods of preparing a patient-specific panel of tumor-specific somatic mutations, comprising: (a) obtainingfrom a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample; (b) depleting DNA fragments that comprise a region of low utility from the first DNA sample and the second DNA sample;(c) sequencing the first depleted DNA sample and the second depleted DNA sample; and(d) generating a patient-specific panel of tumor-specific somatic mutations that are present in the first depleted DNA sample and absent in the second depleted DNA sample. In some embodiments, these methods may further comprise preparing a plurality of oligonucleotide probes, wherein each probe in the plurality of oligonucleotide probes hybridizes to a tumor-specific somatic mutation in the patient-specific signature panel.
[0007] In another aspect, the present disclosure provides methods of preparing a patient-specific panel of tumor-specific somatic mutations, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample; (b) sequencing the first DNA sample and the second DNA sample; and (c) generating a patient-specific panel of tumorspecific mutations that are present in the first DNA sample and absent in the second DNA sample. In some embodiments, these methods may further comprise preparing a plurality of oligonucleotide probes, wherein each prove in the plurality of oligonucleotide probes hybridizes to a tumor-specific somatic mutation in the patient-specific signature panel.
[0008] In another aspect, the present disclosure provides methods of preparing a circulating tumor DNA (ctDNA)-enriched fraction from a biological sample, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample; (b) depleting DNA fragments that comprise a region of low utility from the first DNA sample and the second DNA sample; (c) sequencing the first depleted DNA sample and the second depleted DNA sample; (d) obtaining a further non-tumor sample from the subject; (e) extracting cell free DNA (cfDNA) from the further non-tumor sample; and (f) enriching the cfDNA from the second non-tumor sample with a plurality of oligonucleotide probes, wherein each probe in the plurality of oligonucleotide probes is capable of hybridizing to a DNA fragment comprising a tumor-specific somatic mutation that is present in the first depleted DNA sample and absent in the second depleted DNA sample, thereby obtaining a DNA fraction that is enriched for ctDNA.
[0009] In another aspect, the present disclosure provides methods of preparing a circulating tumor DNA (ctDNA)-enriched fraction from a biological sample, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample; (b) depleting DNA fragments that comprise a region of low utility from the first DNA sample and the second DNA sample; (c) sequencing the first depleted DNA sample and the second depleted DNA sample; (d) obtaining a further non-tumor sample from the subject; (e) extracting cell-free DNA (cfDNA) from the further non-tumor sample; and (f) depleting DNA fragments that comprise a region of low utility from the cfDNA, wherein the remaining DNA fragments comprise the sequences having tumor-specific somatic mutations present in the first depleted DNA sample and absent in the second depleted DNA sample, thereby obtaining a DNA fraction that is enriched for ctDNA.
[0010] In another aspect, the present disclosure provides methods of preparing a circulating tumor DNA (ctDNA)-enriched fraction from a biological sample, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample; (b) sequencing the first DNA sample and the second DNA sample; (c) obtaining a further non-tumor sample from the subject; (e) extracting cell-free DNA (cfDNA) from the further non-tumor sample; and (e) depleting DNA fragments that comprise a region of low utility from the cfDNA, wherein the remaining DNA fragments comprise the sequences having tumorspecific somatic mutations present in the first DNA sample and absent in the second DNA sample, thereby obtaining a DNA fraction that is enriched for ctDNA.
[0011] In another aspect, the present disclosure provides methods of detecting circulating tumor DNA (ctDNA) in a sample, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample; (b) depleting DNA fragments that comprise a region of low utility from the first DNA sample and the second DNA sample; (c) sequencing the first depleted DNA sample and the second depleted DNA sample; (d) obtaining a further non-tumor sample from the subject; (e) extracting cell free DNA (cfDNA) from the further non-tumor sample; (f) enriching the cfDNA from the further non-tumor sample with a plurality of oligonucleotide probes, wherein each probe in the plurality of oligonucleotide probes is capable of hybridizing to a DNA fragmentcomprising a tumor-specific somatic mutation that is present in the first depleted DNA sample and absent in the second depleted DNA sample, thereby obtaining an enriched DNA fraction; (g) sequencing the enriched DNA fraction, thereby obtaining a plurality of sequence reads; and (h) detecting the presence or absence of ctDNA, wherein the presence of a sequence read in the plurality of sequence reads comprising a nucleic acid sequence comprising one or more of the tumor-specific somatic mutations indicates the presence of ctDNA. In some embodiments, (d)-(h) are repeated at one or more times during a cancer treatment, following completion of a cancer treatment, while the subject is in remission, or coinciding with or prior to surgery.
[0012] In another aspect, the present disclosure provides methods of detecting circulating tumor DNA (ctDNA) in a sample, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample; (b) depleting DNA fragments that comprise a region of low utility from the first DNA sample and the second DNA sample; (c) sequencing the first depleted DNA sample and the second depleted DNA sample; (d) obtaining a further non-tumor sample from the subject; (e) extracting cell-free DNA (cfDNA) from the further non-tumor sample; (f) depleting DNA fragments that comprise a region of low utility from the cfDNA, wherein the remaining DNA fragments comprise the sequences having tumor-specific somatic mutations present in the first depleted DNA sample and absent in the second depleted DNA sample, thereby obtaining a DNA fraction that is enriched for ctDNA; (g) sequencing the enriched DNA fraction, thereby obtaining a plurality of sequence reads; and (h) detecting the presence or absence of ctDNA, wherein the presence of a sequence read in the plurality of sequence reads comprising a nucleic acid sequence comprising one or more of the tumor-specific somatic mutations indicates the presence of ctDNA. In some embodiments, (d)-(h) are repeated at one or more times during a cancer treatment, following completion of a cancer treatment, while the subject is in remission, or coinciding with or prior to surgery.
[0013] In another aspect, the present disclosure provides methods of detecting circulating tumor DNA (ctDNA) in a sample, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample; (b) sequencing the first DNA sample and the second DNA sample; (c) obtaining a further non-tumor sample from the subject; (d)extracting cell-free DNA (cfDNA) from the further non-tumor sample; (e) depleting DNA fragments that comprise a region of low utility from the cfDNA, wherein the remaining DNA fragments comprise the sequences having tumor-specific somatic mutations present in the first DNA sample and absent in the second DNA sample, thereby obtaining a DNA fraction that is enriched for ctDNA; (f) sequencing the enriched DNA fraction, thereby obtaining a plurality of sequence reads; and (g) detecting the presence or absence of ctDNA, wherein the presence of a sequence read in the plurality of sequence reads comprising a nucleic acid sequence comprising one or more of the tumor-specific somatic mutations indicates the presence of ctDNA. In some embodiments, (c)-(g) are repeated at one or more times during a cancer treatment, while the subject is in remission, or coinciding with or prior to surgery.
[0014] In another aspect, the present disclosure provides methods of monitoring cancer recurrence, response to treatment, or disease progression in a subject previously treated for cancer, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a nontumor sample; (b) depleting DNA fragments that comprise a region of low utility from the first DNA sample and the second DNA sample; (c) sequencing the first depleted DNA sample and the second depleted DNA sample; (d) obtaining a further non-tumor sample from the subject; (e) extracting cell free DNA (cfDNA) from the further non-tumor sample; (f) enriching the cfDNA from the second non-tumor sample with a plurality of oligonucleotide probes, wherein each probe in the plurality of oligonucleotide probes is capable of hybridizing to a DNA fragment comprising a tumor-specific somatic mutation that is present in the first depleted DNA sample and absent in the second depleted DNA sample, thereby obtaining an enriched DNA fraction; (g) sequencing the enriched DNA fraction, thereby obtaining a plurality of sequence reads; and (h) detecting the presence or absence of any sequence reads in the plurality of sequence reads, wherein the presence of a sequence read comprising a nucleic acid sequence comprising one or more of the tumorspecific somatic mutations indicates (i) cancer recurrence, (ii) a lack of responsiveness to treatment, (iii) disease progression, or (iv) any combination thereof. In some embodiments, (d)-(h) are repeated at one or more times during a cancer treatment, following completion of a cancer treatment, while the subject is in remission, or coinciding with or prior to surgery.
[0015] In another aspect, the present disclosure provides methods of monitoring cancer recurrence, response to treatment, or disease progression in a subject previously treated for cancer, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a nontumor sample; (b) depleting DNA fragments that comprise a region of low utility from the first DNA sample and the second DNA sample; (c) sequencing the first depleted DNA sample and the second depleted DNA sample; (d) obtaining a further non-tumor sample from the subject; (e) extracting cell-free DNA (cfDNA) from the further non-tumor sample; (f) depleting DNA fragments that comprise a region of low utility from the cfDNA, wherein the remaining DNA fragments comprise the sequences having tumorspecific somatic mutations present in the first depleted DNA sample and absent in the second depleted DNA sample, thereby obtaining a DNA fraction that is enriched for ctDNA; (g) sequencing the enriched DNA fraction, thereby obtaining a plurality of sequence reads; and (h) detecting the presence or absence of any sequence reads in the plurality of sequence reads, wherein the presence of a sequence read comprising a nucleic acid sequence comprising one or more of the tumor-specific somatic mutations indicates (i) cancer recurrence, (ii) a lack of responsiveness to treatment, (iii) disease progression, or (iv) any combination thereof. In some embodiments, (d)-(h) are repeated at one or more times during a cancer treatment, following completion of a cancer treatment, while the subject is in remission, or coinciding with or prior to surgery.
[0016] In another aspect, the present disclosure provides methods of monitoring cancer recurrence, response to treatment, or disease progression in a subject previously treated for cancer, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non- tumor sample; (b) sequencing the first DNA sample and the second DNA sample; (c) obtaining a further non-tumor sample from the subject; (d) extracting cell-free DNA (cfDNA) from the further non-tumor sample; (e) depleting DNA fragments that comprise a region of low utility from the cfDNA, wherein the remaining DNA fragments comprise the sequences having tumor-specific somatic mutations present in the first DNA sample and absent in the second DNA sample, thereby obtaining a DNA fraction that is enriched for ctDNA; (f) sequencing the enriched DNA fraction, thereby obtaining a plurality of sequence reads; and (g) detecting the presence or absence of ctDNA, wherein the presence of a sequence read comprising a nucleic acid sequence comprising one or moreof the tumor-specific somatic mutations indicates (i) cancer recurrence, (ii) a lack of responsiveness to treatment, (iii) disease progression, or (iv) any combination thereof. In some embodiments, (c)-(g) are repeated at one or more times during a cancer treatment, following completion of a cancer treatment, while the subject is in remission, or coinciding with or prior to surgery.
[0017] In some embodiments, the region of low utility comprises repetitive elements, low GC content, high GC content, or any combination thereof.
[0018] In some embodiments, the region of low utility does not contain or is unlikely to contain somatic mutations.
[0019] In some embodiments, depleting from the first DNA sample and the second DNA sample any DNA fragments that comprise a region of low utility comprises nuclease-based depletion, restriction enzyme-based depletion, subtractive hybridization, or kinetic-based depletion. In some embodiments, depleting from the cfDNA sample any DNA fragments that comprise a region of low utility comprises nuclease-based depletion, restriction enzyme-based depletion, subtractive hybridization, or kinetic-based depletion. In some embodiments, the depletion method used to deplete regions of low utility from the first DNA sample and the second DNA sample and from the cfDNA sample may be the same. In some embodiments, the depletion method used to deplete regions of low utility from the first DNA sample and the second DNA sample and from the cfDNA sample may be different.
[0020] In some embodiments, nuclease-based depletion comprises targeting depletion of sequences with a DNA-guided nuclease, an RNA-guided nuclease, or a structure- guided nuclease. For example, the DNA-guided endonuclease can be Thermus thermophilus Argonaute (TtAgo) or a transcription activator-like effector nuclease (TALEN), or the RNA-guided endonuclease can be a Cas endonuclease. In some embodiments, endonuclease-based depletion comprises Depletion of Abundant Sequences by Hybridization (DASH).
[0021] In some embodiments, kinetic-based depletion comprises depleting DNA fragments based on denaturation and annealing kinetics.
[0022] In some embodiments, restriction enzyme-based depletion comprises reduced- representation sequencing.
[0023] In another aspect, the present disclosure provides methods of preparing a patient-specific panel of tumor-specific somatic mutations, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample; (b) reducing the size of the first DNA sample and the second DNA sample by digesting the first DNA sample and the second DNA sample with at least one restriction enzyme; (c) sequencing the first reduced DNA sample and the second reduced DNA sample; and (d) generating a patient-specific panel of tumor-specific somatic mutations that are present in the first DNA sample and absent in the second DNA sample. In some embodiments, the method may further comprise preparing a plurality of oligonucleotide probes, wherein each probe in the plurality of oligonucleotide probes hybridizes to a tumor-specific somatic mutation in the patient-specific signature panel.
[0024] In another aspect, the present disclosure provides methods of preparing a circulating tumor DNA (ctDNA)-enriched fraction from a biological sample, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample; (b) reducing the size of the first DNA sample and the second DNA sample by digesting the first DNA sample and the second DNA sample with at least one restriction enzyme; (c) sequencing the first reduced DNA sample and the second reduced DNA sample; (d) obtaining a further non- tumor sample from the subject; (e) extracting cell free DNA (cfDNA) from the further non-tumor sample; and (f) enriching the cfDNA from the second non-tumor sample with a plurality of oligonucleotide probes, wherein each probe in the plurality of oligonucleotide probes is capable of hybridizing to a DNA fragment comprising a tumorspecific somatic mutation that is present in the first reduced DNA sample and absent in the second reduced DNA sample, thereby obtaining a DNA fraction that is enriched for ctDNA.
[0025] In another aspect, the present disclosure provides methods of preparing a circulating tumor DNA (ctDNA)-enriched fraction from a biological sample, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample; (b) reducing the sizeof the first DNA sample and the second DNA sample by digesting the first DNA sample and the second DNA sample with at least one restriction enzyme; (c) sequencing the first reduced DNA sample and the second reduced DNA sample; (d) obtaining a further nontumor sample from the subject; (e) extracting cell-free DNA (cfDNA) from the further non-tumor sample; and (f) reducing the size of the cfDNA sample by digesting the cfDNA sample with at least one restriction enzyme, wherein the remaining DNA fragments comprise the sequences having tumor-specific somatic mutations present in the first reduced DNA sample and absent in the second reduced DNA sample, thereby obtaining a DNA fraction that is enriched for ctDNA. In some embodiments, the restriction enzyme used to digest the first DNA sample and the second DNA sample is the same restriction enzyme used to digest the cfDNA sample.
[0026] In another aspect, the present disclosure provides methods of preparing a circulating tumor DNA (ctDNA)-enriched fraction from a biological sample, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample; (b) sequencing the first DNA sample and the second DNA sample; (c) obtaining a further non-tumor sample from the subject; (d) extracting cell-free DNA (cfDNA) from the further non-tumor sample; and (e) reducing the size of the cfDNA sample by digesting the cfDNA sample with at least one restriction enzyme, wherein the remaining DNA fragments comprise the sequences having tumor-specific somatic mutations present in the first DNA sample and absent in the second DNA sample, thereby obtaining a DNA fraction that is enriched for ctDNA.
[0027] In another aspect, the present disclosure provides methods of detecting circulating tumor DNA (ctDNA) in a sample, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample; (b) reducing the size of the first DNA sample and the second DNA sample by digesting the first DNA sample and the second DNA sample with at least one restriction enzyme; (c) sequencing the first reduced DNA sample and the second reduced DNA sample; (d) obtaining a further non-tumor sample from the subject; (e) extracting cell free DNA (cfDNA) from the further non-tumor sample; (f) enriching the cfDNA from the further non-tumor sample with a plurality of oligonucleotide probes, wherein each probe in the plurality of oligonucleotide probes is capable of hybridizing toa DNA fragment comprising a tumor-specific somatic mutation that is present in the first reduced DNA sample and absent in the second reduced DNA sample, thereby obtaining an enriched DNA fraction; (g) sequencing the enriched DNA fraction, thereby obtaining a plurality of sequence reads; and (h) detecting the presence or absence of ctDNA, wherein the presence of a sequence read in the plurality of sequence reads comprising a nucleic acid sequence comprising one or more of the tumor-specific somatic mutations indicates the presence of ctDNA. In some embodiments, (d)-(h) are repeated at one or more times during a cancer treatment, following completion of a cancer treatment, while the subject is in remission, or coinciding with or prior to surgery.
[0028] In another aspect, the present disclosure provides methods of detecting circulating tumor DNA (ctDNA) in a sample, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample; (b) reducing the size of the first DNA sample and the second DNA sample by digesting the first DNA sample and the second DNA sample with at least one restriction enzyme; (c) sequencing the first reduced DNA sample and the second reduced DNA sample; (d) obtaining a further non-tumor sample from the subject; (e) extracting cell-free DNA (cfDNA) from the further non-tumor sample; (f) reducing the size of the cfDNA sample by digesting the cfDNA sample with at least one restriction enzyme, wherein the remaining DNA fragments comprise the sequences having tumorspecific somatic mutations present in the first reduced DNA sample and absent in the second reduced DNA sample, thereby obtaining a DNA fraction that is enriched for ctDNA; (g) sequencing the enriched DNA fraction, thereby obtaining a plurality of sequence reads; and (h) detecting the presence or absence of ctDNA, wherein the presence of a sequence read in the plurality of sequence reads comprising a nucleic acid sequence comprising one or more of the tumor-specific somatic mutations indicates the presence of ctDNA. In some embodiments, (d)-(h) are repeated at one or more times during a cancer treatment, following completion of a cancer treatment, while the subject is in remission, or coinciding with or prior to surgery. In some embodiments, the restriction enzyme used to digest the first DNA sample and the second DNA sample is the same restriction enzyme used to digest the cfDNA sample.
[0029] In another aspect, the present disclosure provides methods of detecting circulating tumor DNA (ctDNA) in a sample, comprising: (a) obtaining from a subjectwith a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample; (b) sequencing the first DNA sample and the second DNA sample ; (c) obtaining a further non-tumor sample from the subject; (d) extracting cell-free DNA (cfDNA) from the further non-tumor sample; (e) reducing the size of the cfDNA sample by digesting the cfDNA sample with at least one restriction enzyme, wherein the remaining DNA fragments comprise the sequences having tumorspecific somatic mutations present in the first DNA sample and absent in the second DNA sample, thereby obtaining a DNA fraction that is enriched for ctDNA; (f) sequencing the enriched DNA fraction, thereby obtaining a plurality of sequence reads; and (g) detecting the presence or absence of ctDNA, wherein the presence of a sequence read in the plurality of sequence reads comprising a nucleic acid sequence comprising one or more of the tumor-specific somatic mutations indicates the presence of ctDNA. In some embodiments, (c)-(g) are repeated at one or more times during a cancer treatment, following completion of a cancer treatment, while the subject is in remission, or coinciding with or prior to surgery.
[0030] In another aspect, the present disclosure provides methods of monitoring cancer recurrence, response to treatment, or disease progression in a subject previously treated for cancer, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non- tumor sample; (b) reducing the size of the first DNA sample and the second DNA sample by digesting the first DNA sample and the second DNA sample with at least one restriction enzyme; (c) sequencing the first reduced DNA sample and the second reduced DNA sample; (d) obtaining a further non-tumor sample from the subject; (e) extracting cell free DNA (cfDNA) from the further non-tumor sample; (f) enriching the cfDNA from the second non-tumor sample with a plurality of oligonucleotide probes, wherein each probe in the plurality of oligonucleotide probes is capable of hybridizing to a DNA fragment comprising a tumor-specific somatic mutation that is present in the first reduced DNA sample and absent in the second reduced DNA sample, thereby obtaining an enriched DNA fraction; (g) sequencing the enriched DNA fraction, thereby obtaining a plurality of sequence reads; and (h) detecting the presence or absence of any sequence reads in the plurality of sequence reads, wherein the presence of a sequence read comprising a nucleic acid sequence comprising one or more of the tumor-specific somatic mutations indicates (i) cancer recurrence, (ii) a lack of responsiveness to treatment, (iii)disease progression, or (iv) any combination thereof. In some embodiments, (d)-(h) are repeated at one or more times during a cancer treatment, following completion of a cancer treatment, while the subject is in remission, or coinciding with or prior to surgery.
[0031] In another aspect, the present disclosure provides methods of monitoring cancer recurrence, response to treatment, or disease progression in a subject previously treated for cancer, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a nontumor sample; (b) reducing the size of the first DNA sample and the second DNA sample by digesting the first DNA sample and the second DNA sample with at least one restriction enzyme; (c) sequencing the first reduced DNA sample and the second reduced DNA sample; (d) obtaining a further non-tumor sample from the subject; (e) extracting cell-free DNA (cfDNA) from the further non-tumor sample; (f) reducing the size of the cfDNA sample by digesting the cfDNA sample with at least one restriction enzyme, wherein the remaining DNA fragments comprise the sequences having tumor-specific somatic mutations present in the first reduced DNA sample and absent in the second reduced DNA sample, thereby obtaining a DNA fraction that is enriched for ctDNA; (g) sequencing the enriched DNA fraction, thereby obtaining a plurality of sequence reads; and (h) detecting the presence or absence of any sequence reads in the plurality of sequence reads, wherein the presence of a sequence read comprising a nucleic acid sequence comprising one or more of the tumor-specific somatic mutations indicates (i) cancer recurrence, (ii) a lack of responsiveness to treatment, (iii) disease progression, or (iv) any combination thereof. In some embodiments, (d)-(h) are repeated at one or more times during a cancer treatment, following completion of a cancer treatment, while the subject is in remission, or coinciding with or prior to surgery. In some embodiments, the restriction enzyme used to digest the first DNA sample and the second DNA sample is the same restriction enzyme used to digest the cfDNA sample.
[0032] In another aspect, the present disclosure provides methods of monitoring cancer recurrence, response to treatment, or disease progression in a subject previously treated for cancer, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non- tumor sample; (b) sequencing the first DNA sample and the second DNA sample; (c) obtaining a further non-tumor sample from the subject; (d) extracting cell-free DNA(cfDNA) from the further non-tumor sample; (e) reducing the size of the cfDNA sample by digesting the cfDNA sample with at least one restriction enzyme, wherein the remaining DNA fragments comprise the sequences having tumor-specific somatic mutations present in the first DNA sample and absent in the second DNA sample, thereby obtaining a DNA fraction that is enriched for ctDNA; (f) sequencing the enriched DNA fraction, thereby obtaining a plurality of sequence reads; and (g) detecting the presence or absence of any sequence reads in the plurality of sequence reads, wherein the presence of a sequence read comprising a nucleic acid sequence comprising one or more of the tumorspecific somatic mutations indicates (i) cancer recurrence, (ii) a lack of responsiveness to treatment, (iii) disease progression, or (iv) any combination thereof. In some embodiments, (c)-(g) are repeated at one or more times during a cancer treatment, following completion of a cancer treatment, while the subject is in remission, or coinciding with or prior to surgery.
[0033] In some embodiments of any of the foregoing aspects, the tumor-specific somatic mutation is obtained by aligning sequences of the first DNA sample to a reference human genome that is not from the subject, aligning sequences of the second DNA sample to a reference human genome that is not from the subject, and selecting mutations that are present in the first DNA sample and absent in the second DNA sample.
[0034] In some embodiments of any of the foregoing aspects, the tumor-specific somatic mutation is obtained by aligning sequences of the first DNA sample to sequences of the second DNA sample.
[0035] In some embodiments of any of the foregoing aspects, the tumor-specific somatic mutation(s) comprises one or more somatic mutations selected from SNVs, insertions, deletions, and translocations.
[0036] In some embodiments of any of the foregoing aspects, first non-tumor sample comprises a tissue sample matched to a tissue of origin of the tumor sample. In some embodiments of any of the foregoing aspects, first non-tumor sample comprises a fluid sample selected from a buffy coat sample, blood, blood plasma, blood serum, urine, saliva, and cerebral spinal fluid (CSF).
[0037] In some embodiments, the plurality of oligonucleotide probes is capable of detecting at least 10, at least 50, at least 100, at least 150, at least 200, at least 250, or at least 500 tumor-specific somatic mutations.
[0038] In some embodiments, enriching the cfDNA from the second non-tumor sample comprises (i) hybrid capture-based enrichment, (ii) PCR-target enrichment, or (iii) on-sequencer enrichment.
[0039] In some embodiments, the further non-tumor sample comprises a fluid sample selected from a buffy coat sample, blood, blood plasma, blood serum, urine, saliva, and cerebral spinal fluid (CSF).
[0040] In some embodiments, sequencing the enriched DNA fraction comprises next generation sequencing.
[0041] In some embodiments, the enriched DNA fraction comprises sequencing of introns, exons, intergenic regions, or a combination thereof.
[0042] In some embodiments of any of the foregoing aspects, the cancer is selected from adrenal cancer, anal cancer, bile duct cancer, bladder cancer, bone cancer, a brain / CNS tumor, breast cancer, Castleman disease, cervical cancer, colon or rectum cancer, endometrial cancer, esophagus cancer, a Ewing tumor, eye cancer, gallbladder cancer, a gastrointestinal carcinoid tumor, a gastrointestinal stromal tumor (GIST), gestational trophoblastic disease, Hodgkin disease, Kaposi sarcoma, kidney cancer, laryngeal and hypopharyngeal cancer, leukemia, liver cancer, lung cancer, lymphoma, malignant mesothelioma, multiple myeloma, myelodysplastic Syndrome, nasal cavity or paranasal sinus cancer, nasopharyngeal cancer, neuroblastoma, oral cavity or oropharyngeal cancer, osteosarcoma, ovarian cancer, pancreatic cancer, penile cancer, a pituitary tumor, prostate cancer, retinoblastoma, rhabdomyosarcoma, salivary gland cancer, skin cancer, small intestine cancer, stomach cancer, testicular cancer, thymus cancer, thyroid cancer, uterine sarcoma, vaginal cancer, vulvar cancer, Waldenstrom macroglobulinemia, and Wilms tumor.
[0043] In another aspect, the present disclosure provides enriched ctDNA fractions prepared according to the methods disclosed herein. In some embodiments, the enriched ctDNA fraction comprises cfDNA fragments that collectively comprise at least 10, atleast 50, at least 100, at least 150, at least 200, at least 250, or at least 500 tumor-specific somatic mutations. In some embodiments, the DNA fraction is enriched for cfDNA fragments averaging less than about 200 base pairs in length.BRIEF DESCRIPTION OF THE DRAWINGS
[0044] FIGs. 1A, IB, and 1C depict exemplary methods of integrating targeted depletion and / or reduction of genome size with MRD assays. FIG. 1 A shows an exemplary method wherein (a) Phase I includes (i) targeted depletion or reduction in genome size of a DNA sample from a tumor sample and a DNA sample from a non-tumor sample coupled with sequencing of the depleted or reduced DNA sample, (ii) utilizing the data from the sequencing to identify somatic variants, and (iii) utilizing the identified somatic variants to develop an MRD panel of probes and wherein (b) Phase II includes using cfDNA and the MRD panel of probes for targeted sequencing utilizing hybrid capture methods to enrich for ctDNA in the cfDNA sample and finally detect the presence of somatic variants in the ctDNA. FIG. IB depicts a second exemplary method wherein (a) Phase I includes (i) targeted depletion or reduction in genome size of a DNA sample from a tumor sample and a DNA sample from a non-tumor sample coupled with sequencing of the depleted or reduced DNA sample and (ii) using the data from the sequencing to identify somatic variants and wherein (b) Phase II includes (i) targeted depletion or reduction in genome size of a cfDNA sample coupled with sequencing of the depleted or reduced DNA sample and (ii) utilizing the data from the sequencing to detect the presence of the identified somatic variants in the ctDNA. FIG. 1C depicts another exemplary method wherein (a) Phase I includes (i) whole genome sequencing or whole exome sequencing of a DNA sample from a tumor sample and a DNA sample from a non-tumor sample and (ii) using the data from the whole genome sequencing or whole exome sequencing to identify somatic variants and wherein (b) Phase II includes (i) targeted depletion or reduction in genome size of a cfDNA sample coupled with sequencing of the depleted or reduced DNA sample and (ii) utilizing the data from the sequencing to detect the presence of the identified somatic variants in the ctDNA.
[0045] FIG. 2 depicts the quantification of various DNA fragments generated from shearing three buffy coat DNA samples on Covaris R230 under FFPE protocol (20 iterations total), FFPE protocol + 10 iterations (30 iterations total), and FFPE protocol +20 iterations (40 iterations total). All three samples generated fragmented DNA in the desired 300 to 450 bp range.
[0046] FIG. 3 shows quantification of the DNA yield from the hybridization supernatant for four different input concentrations (0 ng, 50 ng, 100 ng, and 500 ng) with the probe (left bar for each input) and without the probe (right bar for each input).
[0047] FIGs. 4A, 4B, 4C, and 4D show the qPCR results following hybridization. FIG. 4A shows the Ct Values for the samples, from left to right, of 0 ng input with probe, 50 ng input with probe, 100 ng input with probe, 500 ng input with probe, 50 ng input without probe, 100 ng input without probe, 500 ng input without probe, the input control, and the NTC control. The first bar for each sample represents the qPCR results using primer set 1, identified on FIG. 4A as ALU 101-206. The second bar for each sample represents the qPCR results using primer set 2, identified on FIG. 4A as ALU 98-188. The third bar for each sample represents the qPCR results using primer set 3, identified on FIG. 4A as ALU. The fourth bar for each sample represents the qPCR result using the housekeeping gene GAPDH primer set. FIG. 4B shows the fold change versus input using primer set 1, identified on FIG. 4B as ALU 101-206, for inputs of 50 ng, 100 ng, and 500 ng. The left bar for each sample represents samples without probe added, while the right bar for each sample represents samples with probe added. FIG. 4C shows the fold change versus input using primer set 2, identified on FIG. 4C as ALU 98-188, for inputs of 50 ng, 100 ng, and 500 ng. The left bar for each sample represents samples without probe added, while the right bar for each sample represents samples with probe added. FIG. 4D shows the fold change versus input using primer set 3, identified on FIG. 4D as ALU, for inputs of 50 ng, 100 ng, and 500 ng. The left bar for each sample represents samples without probe added, while the right bar for each sample represents samples with probe added.
[0048] FIG. 5 depicts the TapeStation quantifications results following amplification via PCR using primer set 2, identified on FIG. 5 as ALU 98-188, to test the amount of probe captured on the first, second, and third streptavidin pulldown captures using various input amounts.
[0049] FIGs. 6A and 6B show the results of the tested clean-up methods for the two restriction enzyme digestions. FIG. 6A shows the results for the HpyCH4V digestion samples. FIG. 6B shows the results for the Alul digestion samples.
[0050] FIG. 7 shows the quantification of the DNA concentration for cleaned up indexed library samples using TapeStation (D1000).
[0051] FIG. 8 shows the quantification for the DNA concentration of the unpurified PCR reaction productions from the hybridization supernatant samples.
[0052] FIG. 9 shows the fold change of targets measured by qPCR following repetitive element depletion.
[0053] FIGs. 10A and 10B show the fold change of targets measured by qPCR following restriction enzyme digestion. FIG. 10A shows samples digested with HpyCH4V. FIG. 10B shows samples digested with Alul.
[0054] FIG. 11 shows analysis of clusters for sampled depleted using either hybridization capture or restriction enzyme digestion.
[0055] FIGs. 12 A, 12B, 12C, 12D, and 12E show the mean coverage for the targeted depleted sequence libraries. FIG. 12A shows the summary of the analysis of the samples including mean coverage, CV coverage, and median insert size. FIG. 12B shows the quantification of the total number of reads for each sample. FIG. 12C shows the mean coverage for each sample. FIG. 12D shows the insert size for each sample. FIG. 12E shows the GC content for each sample.
[0056] FIGs. 13A, 13B, and 13C show normalized depth ratios between test and control samples for each target. FIG. 13 A shows the average depth for each repetitive element. FIG. 13B shows the average depth normalized to account for variability in raw sequencing depth for each repetitive element. FIG. 13C shows the normalized repetitive element site depths divided by the average of the normalized depth of the corresponding sites from the whole genome sequencing replicates for each repetitive element. The y-axis is restricted to better represent the majority of the data, but some outliers are not shown.
[0057] FIG. 14 shows the median normalized depth ratio compared to the probe-to- target ratio.
[0058] FIGs. 15A and 15B show the depletion at each target site correlated to BLAST Hits. FIG. 15A shows the normalized depth at each target site correlated to BLAST Hits. FIG. 15B shows the normalized depth ratio at each target site correlated to BLAST Hits.
[0059] FIGs. 16A and 16B show the depth ratio between test and control samples for each repetitive element following restriction enzyme digestion. FIG. 16A shows the average depth for each repetitive element site. FIG. 16B shows the repetitive element site depths divided by the average depth of the corresponding sites from the whole genome sequencing replicates to get the depth ratio for each repetitive element site.DETAILED DESCRIPTION
[0060] The present disclosure provides methods for detecting and quantifying circulating tumor DNA in a sample from a subject and methods of preparing samples for such detection. The disclosed methods include processes for generating a patient-specific panel of somatic mutations that are present in a tumor sample and absent from a nontumor sample. Such panels cane be used, for example, to track or monitor minimal residual disease (MRD) or disease recurrence in a subject that was previously diagnosed with cancer.
[0061] Conventionally, tumor-informed MRD assays begin with sequencing DNA from a tumor sample and a non-tumor sample from a subject. This has conventionally been done either via whole genome sequencing (WGS) or whole exome sequencing (WES). The present disclosure provides improvements to existing MRD methodologies utilizing WGS by depleting unnecessary DNA from the tumor and / or non-tumor sample prior to sequencing, and in so doing provides processes that are faster, more efficient, more sensitive, and less expensive, as the disclosed methods do not require as much sequencing to establish a patient-specific panel of mutations. The present disclosure also provides improvements to existing MRD methodologies utilizing WES by increasing the number of somatic mutations that can be identified as compared to WES methods and thereby increasing the sensitivity of the tests, as the number of somatic mutations identified is proportional to the amount of the genome that is sequenced. Particular methodologies for depleting unnecessary DNA or “regions or low utility” are described herein, along with further discussion of the improvements to the technology provided by the disclosed methods.
[0062] It is to be appreciated that certain aspects, modes, embodiments, variations and features of the present methods are described below in various levels of detail in order to provide a substantial understanding of the present technology. It is to be understood that the present disclosure is not limited to particular uses, methods, reagents, compounds, compositions or biological systems, which can, of course, vary. It is also to be understood that the terminology used herein for the purpose of describing particular embodiments only and is not intended to be limiting.I. Definitions
[0063] Unless defined otherwise, all technical and scientific terms used herein generally have the same meaning as commonly understood by one of ordinary skill in the art to which this technology belongs. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise. For example, reference to a “a cell” includes a combination of two or more cells, and the like. Generally, the nomenclature used herein and the laboratory procedures in cell culture, molecular genetics, organic chemistry, analytical chemistry and nucleic acid chemistry and hybridization described below are those well- known and commonly employed in the art.
[0064] As used herein, the term “about” in reference to a number is generally taken to include numbers that fall within a range of 1%, 5%, or 10% in either direction (greater than or less than) of the number unless otherwise stated or otherwise evident from the context (except where such number would be less than 0% or exceed 100% of a possible value).
[0065] It is understood that aspects and variation of the invention described herein include “consisting” and / or “consisting essentially of’ aspects and variations. Throughout the description, where compositions are described as having, including, or comprising specific components, or where processes and methods are described as having, including, or comprising specific steps, it is contemplated that, additionally, there are compositions of the present disclosure that consist essentially of, or consist of, the recited components, and that there are processes and methods according to the present disclosure that consist essentially of, or consist of, the recited processing steps.
[0066] A “set” of reads refers to all sequence reads with a common parent nucleic acid strand, which may or may not have had errors introduced during sequencing or amplification of the parent nucleic acid strand.
[0067] Numeric ranges are inclusive of the numbers defining the range.
[0068] The term “mutation” herein refers to a change introduced into a reference sequence, including, but not limited to, substitutions, insertions, and deletions (including truncations) relative to the reference sequence. Mutations can involve large sections of DNA (e.g., a copy number variant). Mutations can involve whole chromosomes (e.g., aneuploidy). Mutations can involve small sections of DNA. Examples of mutations involving small sections of DNA, including e.g., point mutations or single nucleotide polymorphisms (SNPs), single nucleotide variants (SNVs), multiple nucleotide polymorphisms, insertions (e.g., insertion of one or more nucleotides at a locus but less than the entire locus), multiple nucleotide changes, deletions (e.g., deletion of one or more nucleotides at a locus), inversions (e.g., reversal of a sequence of one or more nucleotides), and genomic rearrangements (e.g., deletions, duplications, inversions, and translocations). In some embodiments, the reference sequence is a parental sequence. In some embodiments, the reference sequence is a reference human genome, e.g., hl9. In some embodiments, the reference sequence is derived from a non-cancer (or non-tumor) sequence. In some embodiments, the mutation is inherited. In some embodiments, the mutation is spontaneous or de novo. In some embodiments, the mutation is a “somatic” mutation or variant.
[0069] As used herein, the term “somatic variant” or “somatic mutation” refers to a variant arising after conception in non-germline DNA of an individual. Somatic variants may include single-nucleotide variants (SNVs), multi-nucleotide variants, insertions and deletions e.g., indel variants), and genomic rearrangements, for example. The terms “somatic variant” and “somatic mutation” are used interchangeably herein. In some embodiments, the terms “somatic variant” and “somatic mutation” refer to a collection of somatic variants that are specific to a patient.
[0070] The term “patient-specific panel” or “set of somatic variants” herein refers to a collection of sequences comprising somatic mutations that are specific to a patient, ormarkers that distinguish between two or more individuals. A panel may distinguish one sample from another.
[0071] The term “subset of somatic variants” or “subset panel” herein refers to a subset of somatic variants of the patient-specific panel or set of somatic variants.
[0072] The term “tumor burden” herein refers to the total amount of tumor material present in a patient, which can be reflected by the tumor fraction as determined according to the methods provided herein.
[0073] The term “tumor fraction” herein refers to the proportion of circulating cell- free tumor DNA (ctDNA) relative to the total amount of cell-free DNA (cfDNA). Tumor fraction may be indicative of the size of the tumor.
[0074] As used herein, the term “genomic DNA” refers to DNA of a cellular genome. The genomic DNA can be cellular, i.e., contained within a cell, or it can be cell-free.
[0075] The term “sample” herein refers to any substance containing or presumed to contain nucleic acid. The sample can be a biological sample obtained from a subject or patient. The nucleic acids can be RNA or DNA (e.g., genomic DNA). In some embodiments, the biological sample is a biological fluid sample. The fluid sample can be whole blood, plasma, serum, ascites, cerebrospinal fluid, sweat, urine, tears, saliva, buccal sample, cavity rinse, or organ rinse. The fluid sample can be an essentially cell-free liquid sample (e.g., plasma, serum, sweat, urine, tears, etc.). In some embodiments, the biological sample is a solid biological sample, e.g., feces or tissue biopsy, such as a tumor biopsy. In some embodiments, the sample is a tumor sample. In some embodiments, the sample is a non-tumor sample. A “sample” may include, but is not limited to, tissue, blood, plasma, saliva, urine, semen, amniotic fluid, oocytes, skin, hair, feces, cheek swabs, or pap smear lysate from an individual. In some embodiments, the sample is blood, plasma, or serum.
[0076] The term “target detection sequence” herein refers to a selected target detection polynucleotide, e.g., a sequence present in a cfDNA molecule, whose presence, amount, and / or nucleotide sequence, or changes in these, are desired to be determined. Target detection sequences are interrogated for the presence or absence of a somaticvariant. The target detection polynucleotide can be a region of a gene associated with a disease. In some embodiments, the region is an exon. The disease can be cancer.
[0077] As used herein, the term “target detection polynucleotide” refers to a nucleic acid molecule or polynucleotide in a population of nucleic acid molecules having a target detection sequence to which one or more oligonucleotides are designed to hybridize. “Target detection polynucleotide” may be used to refer to a double-stranded nucleic acid molecule that includes a target detection sequence on one or both strands, or a singlestranded nucleic acid molecule including a target detection sequence, and may be derived from any source of or process for isolating or generating nucleic acid molecules. A target detection polynucleotide may include one or more (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more) target sequences, which may be the same or different. In general, different target detection polynucleotides include different sequences, such as one or more different nucleotides or one or more different target detection sequences.
[0078] The terms “anneal,” “hybridize,” or “bind” can refer to two polynucleotide sequences, segments, or strands, and can be used interchangeably and have the usual meaning in the art. Two complementary sequences (e.g., DNA and / or RNA) can anneal or hybridize by forming hydrogen bonds with complementary bases to produce a doublestranded polynucleotide or a double-stranded region of a polynucleotide.
[0079] The term “marker” or “segregating marker” refers to a moiety that is used to discriminate between two or more samples, e.g., two or more individuals or tissues. A marker may be a nucleic acid (e.g., a gene), small molecule, peptide, fatty acid, metabolite, protein, lipid, etc. A marker may be a mutation. A marker may be a synthetic nucleic acid. A marker or set of markers may define a genetic signature of an entity, e.g., an individual, relative to a second nucleic acid, e.g., a reference nucleic acid sequence.
[0080] The terms “treat,” “treatment,” and “treating” refer to the reduction or amelioration of the progression, severity, and / or duration of a proliferative disorder, e.g., cancer, or the amelioration of a proliferative disorder resulting from the administration of one or more therapies.
[0081] As used herein, the term “barcode” (also termed “single molecule identifier” or “SMI”) refers to a known nucleic acid sequence that allows some feature of a polynucleotide with which the barcode is associated to be identified. In someembodiments, the feature of the polynucleotide to be identified is the sample from which the polynucleotide is derived. In some embodiments, barcodes are about or at least about 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more nucleotides in length. In some embodiments, barcodes are shorter than 10, 9, 8, 7, 6, 5, or 4 nucleotides in length. In some embodiments, barcodes associated with some polynucleotides are of different lengths than barcodes associated with other polynucleotides. In general, barcodes are of sufficient length and include sequences that are sufficiently different to allow the identification of samples based on barcodes with which they are associated. In some embodiments, a barcode, and the sample source with which it is associated, can be identified accurately after the mutation, insertion, or deletion of one or more nucleotides in the barcode sequence, such as the mutation, insertion, or deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides. In some embodiments, each barcode in a plurality of barcodes differ from every other barcode in the plurality at least three nucleotide positions, such as at least 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotide positions. A plurality of barcodes may be represented in a pool of samples, each sample including polynucleotides comprising one or more barcodes that differ from the barcodes contained in the polynucleotides derived from the other samples in the pool. Samples of polynucleotides including one or more barcodes can be pooled based on the barcode sequences to which they are joined, such that all four of the nucleotide bases A, G, C, and T are approximately evenly represented at one or more positions along each barcode in the pool (such as at 1, 2, 3, 4, 5, 6, 7, 8, or more positions, or all positions of the barcode).
[0082] The term “copy number variant,” “CNV,” or “copy number” refers to any duplication or deletion of a genomic segment. In some embodiments, the copy number is the copy number of each somatic variant in the set of somatic variants.
[0083] The term “derived from” encompasses the terms “originated from,” “obtained from,” “obtainable from,” “isolated from,” and “created from,” and generally indicates that one specified material (e.g., a biological sample) finds its origin in another specified material or individual or has features that can be described with reference to another specified material.
[0084] The term “library” or “sequencing library” herein refers to a collection or plurality of template molecules, i.e., target DNA duplexes, which share common sequences at their 5’ ends and common sequences at their 3’ ends. Use of the term“library” to refer to a collection or plurality of template molecules should not be taken to imply that the templates making up the library are derived from a particular source, or that the “library” has a particular composition. By way of example, use of the term “library” should not be taken to imply that the individual templates within the library must be of different nucleotide sequence or that the templates must be related in terms of sequence and / or source. In general, the term “sequencing library” herein refers to DNA that is processed for sequencing, e.g., using massively parallel methods, e.g., Next Generation Sequencing. The DNA may optionally be amplified to obtain a population of multiple copies of processed DNA, which can be sequenced by Next Generation Sequencing.
[0085] The term “Next Generation Sequencing” or “NGS” refers to sequencing methods that allow for massively parallel sequencing of clonally amplified molecules and of single nucleic acid molecules during which a plurality (e.g., millions) of nucleic acid fragments from a single sample or from multiple different samples are sequenced in unison. Non-limiting examples of NGS include sequencing-by-synthesis, sequencing-by- ligation, real-time sequencing, and nanopore sequencing.
[0086] The term “sequence read” or simply “read” herein refers to sequence information of a nucleic acid fragment obtained through a sequencing assay, such as a Next Generation Sequencing (NGS) assay. In some embodiments, a sequence read refers to data representing a sequence of nucleotide bases that were measured using a clonal sequencing method. Clonal sequencing may produce sequence data representing single, or clones, or clusters of one original DNA molecule. A sequence read may also have an associated quality score at each base position of the sequence, indicating the probability that the nucleotide has been called correctly.
[0087] The term “mapping a sequence read” herein refers to the process of determining a sequence read’s location of origin in the genome sequence of a particular organism. The location of origin of sequence reads is based on the similarity of the nucleotide sequence of the read and the genome sequence.
[0088] The term “preferential enrichment” of DNA that corresponds to a locus, or preferential enrichment of DNA at a locus, refers to any method that results in the percentage of molecules of DNA in a post-enrichment DNA mixture that corresponds tothe locus being higher than the percentage of molecules of DNA in the pre-enrichment DNA mixture that correspond to the locus. The method may involve selective amplification of DNA molecules that correspond to a locus. The method may involve removing DNA molecules that do not correspond to a locus. The method may involve a combination of methods. The degree of enrichment is defined as the percentage of molecules of DNA in the post-enrichment mixture that corresponds to the locus divided by the percentage of molecules of DNA in the pre-enrichment mixture that correspond to the locus. Preferential enrichment may be carried out at a plurality of loci. In some embodiments of the present disclosure, the degree of enrichment is greater than 20. In some embodiments of the present disclosure, the degree of enrichment is greater than 200. In some embodiments of the present disclosure, the degree of enrichment is greater than 2,000. When preferential enrichment is carried out at a plurality of loci, the degree of enrichment may refer to the average degree of enrichment of all of the loci in the set of loci.
[0089] The term “amplification,” with respect to nucleic acid sequences, herein refers to methods that increase the representation of a population of nucleic acid sequences in a sample. Copies of a particular target nucleic acid generated in vitro in an amplification reaction are called “amplicons” or “amplification products.” Amplification may be exponential or linear. A target nucleic acid may be DNA (such as, for example, genomic DNA and cDNA) or RNA. While the exemplary method described hereinafter relate to amplification using polymerase chain reaction (PCR), numerous other methods, such as isothermal methods, rolling circle methods, etc., are available to the skilled artisan. The skilled artisan will understand that these other methods may be used either in place of, or together with, PCR methods. See, e.g., Saiki, “Amplification of Genomic DNA" in PCR Protocols, Innis et al., Eds., Academic Press, San Diego, CA 1990, pp. 13-20; Wharam et al., Nucleic Acids Res. 29(11):E54-E54 (2001).
[0090] The term “selective amplification” herein refers to a method that increases the number of copies of a particular molecule of DNA, or molecules of DNA, that correspond to a particular region of DNA. It may also refer to a method that increases the number of copies of a particular targeted molecule of DNA or targeted region of DNA more than it increases non-targeted molecules or regions of DNA. Selective amplification may be a method of preferential enrichment.
[0091] The term “direct amplification” herein refers to a nucleic acid amplification reaction in which the target nucleic acid is amplified from the sample without prior purification, extraction, or concentration.
[0092] The term “amplification mixture” herein refers to a mixture of reagents that are used in a nucleic acid amplification reaction but does not contain primers or sample. An amplification mixture comprises a buffer, dNTPs, and a DNA polymerase. An amplification mixture may further comprise at least one of MgCh, KC1, nonionic and ionic detergents (including cationic detergents). In general, amplification methods disclosed herein include an amplification mixture. The term “amplification master mix” refers to an amplification mixture, primers, and / or probes for amplifying one or more target nucleic acids but does not contain the sample to be amplified. The term “reactionsample mixture” herein refers to a mixture containing amplification master mix and a sample.
[0093] The term “multiplex PCR” herein refers to the simultaneous generation of two or more PCR products or amplicons within the same reaction vessel. Similarly, a “2-plex PCR” refers to the simultaneous generation of two PCR products or amplicons within the same reaction vessel. Each PCR product is primed using a distinct primer pair. A multiplex reaction may further include specific probes for each product that are labeled with different detectable moieties.
[0094] The term “universal priming sequence” refers to a DNA sequence that may be appended to a population of target DNA molecules, for example by ligation, PCR, or ligation-mediated PCR. Once added to the population of target molecules, primers specific to the universal priming sequences can be used to amplify the target population using a single pair of amplification primers. Universal priming sequences are typically not related to the target sequences.
[0095] The term “universal adapters” or “ligation adapters” or “library tags” are DNA molecules containing a universal priming sequence that can be covalently linked to the 5’ and 3’ end of a population of target double-stranded DNA molecules. The addition of the adapters provides universal priming sequences to the 5’ and 3’ end of the target population from which PCR amplification can take place, amplifying all molecules from the target population, using a single pair of amplification primers.
[0096] The term “targeting” herein refers to a method used to selectively amplify or otherwise preferentially enrich those molecules of DNA the correspond to a set of loci in a mixture of DNA.
[0097] The term “primer” herein refers to an oligonucleotide, whether occurring naturally or produced synthetically, which is capable of acting as a point of initiation of nucleic acid synthesis when placed under conditions in which synthesis of a primer extension product which is complementary to a nucleic acid strand is induced, e.g., in the presence of four different nucleotide triphosphates and a polymerase enzyme, e.g., a thermostable enzyme, in an appropriate buffer (“buffer” includes pH, ionic strength, cofactors, etc.) and at a suitable temperature. The primer is preferably single-stranded for maximum efficiency in amplification but may alternatively be double-stranded. If doublestranded, the primer is first treated to separate its strands before being used to prepare extension products. Preferably, the primer is an oligodeoxyribonucleotide. The primer must be sufficiently long to prime the synthesis of extension products in the presence of the polymerase, e.g., thermostable polymerase enzyme. The exact lengths of a primer will depend on many factors, including temperature, source of primer, and use of the method. For example, depending on the complexity of the target sequence, the oligonucleotide primer typically contains 15-25 nucleotides, although it may contain more or fewer nucleotides. Short primer molecules generally require colder temperatures to form sufficiently stable hybrid complexes with the template.
[0098] A “hybrid capture probe” herein refers to any nucleic acid sequence, possibly modified, that is generated by various methods, such as PCR or direct synthesis, and intended to be complementary to one strand of a specific target DNA sequence in a sample. The exogenous hybrid capture probes may be added to a prepared sample and hybridized through a denature-reannealing process to form duplexes of exogenous- endogenous fragments. These duplexes may then be physically separated from the sample by various means.
[0099] A “spacer” may consist of a repeated single nucleotide (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more of the same nucleotide in a row), or a sequence of 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides repeated 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more times. A spacer may comprise or consist of a specific sequence, such as a sequence that does not hybridize toany target sequence in a sample. A spacer may comprise or consist of a sequence of randomly selected nucleotides.
[0100] The terms “substantially similar” and “substantially identical” in the context of at least two nucleic acids typically means that a polynucleotide includes a sequence that has at least about 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 99.5% sequence identical, in comparison with a reference (e.g., wild-type) polynucleotide or polypeptide. Sequence identity may be determined using known programs such as BLAST, ALIGN, and CLUSTAL using standard parameters. See, e.g., Altshul et al. (1990) J. Mol. Biol. 215:403-410; Henikoff et al. (1989) Proc. Natl. Acad. Sci. 89: 10915; Karin et al. (1993) Proc. Natl. Acad. Sci. 90:5873; and Higgins et al. (1988) Gene 73:237. Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information. Also, databases may be searched using FASTA (Person et al. (1988) Proc. Natl. Acad. Sci. 85:2444-2448). In some embodiments, substantially identical nucleic acid molecules hybridize to each other under stringent conditions (e.g., within a range of medium to high stringency).
[0101] The term “tag” refers to a detectable moiety that may be one or more atom(s) or molecule(s), or a collection of atoms and molecules. A tag may provide an optical, fluorescent, electrochemical, magnetic, or electrostatic (e.g., inductive, capacitive) signature.
[0102] The term “tagged nucleotide” herein refers to a nucleotide that includes a tag (or tag species) that is coupled to any location of the nucleotide including, but not limited to a phosphate (e.g., terminal phosphate), sugar or nitrogenous base moiety of the nucleotide. Tags may be one or more atom(s) or molecule(s), or a collection of atoms or molecules. A tag may provide an optical, electrochemical, magnetic, or electrostatic (e.g., inductive, capacitive) signature.
[0103] The term “template DNA molecule” herein refers to a strand of a nucleic acid from which a complementary nucleic acid strand is synthesized by a DNA polymerase, for example, in a primer extension reaction.
[0104] A “portion adjacent to a region of interest” refers to a sequence that is immediately proximal to a region of interest. Reference to a “portion of or adjacent to aregion of interest” refers to a sequence that 1) is entirely within the region of interest, 2) is entirely outside by immediately proximal to the region of interest, or 3) includes a contiguous sequence from within and immediately proximal to the region of interest. Reference to a “sequence that is substantially complementary to a portion of or adjacent to a region of interest” refers to 1) a sequence that is substantially complementary to a sequence entirely within the region of interest, 2) a sequence substantially complementary to a sequence entirely outside but immediately proximal to the region of interest, or 3) a sequence that is substantially complementary to a contiguous sequence from within and immediately proximal to the region of interest.
[0105] “Noisy genetic data” herein refers to genetic data with any of the following: allele dropouts, uncertain base pair measurements, incorrect base pair measurements, missing base pair measurements, uncertain measurements of insertions or deletions, uncertain measurements of chromosome segment copy numbers, spurious signals, missing measurements, other errors, or combinations thereof.
[0106] A “region of low utility” refers to a sequence or portion of genomic DNA that does is not necessary for preparing a patient-specific panel of somatic mutation or that does not have information relevant to the detection assays disclosed herein. Regions of low utility include, but are not limited to, repetitive elements that have a high proportion of off-target reads that will be enriched and sequenced in the detection assay. Regions of low utility can also include regions of the genome that contain few somatic mutations or that do not contain regions of interest for disease detection. Regions of low utility comprise target depletion sequences. Additionally, regions of low utility can include, but are not limited to, regions of DNA that contain high GC content, low GC content, or are otherwise prone to causing sequencing errors during subsequent sequencing analysis of the DNA. Regions having low GC content may be regions having less than or equal to 30% GC content, less than or equal to 20% GC content, or less than or equal to 10% GC content. Regions having high GC content may be regions having greater than or equal to 70% GC content, greater than or equal to 80% GC content, or greater than or equal to 90% GC content.
[0107] The term “target depletion sequence” refers to a sequence within a region of low utility that is used to identify or target the region of low utility for depletion. The target depletion sequence may be a nucleotide sequence complementary to the DNA orRNA guide for a targeted endonuclease. Alternatively, the target depletion sequence may be a sequence that has high or low GC content, or a sequence that contains a restriction enzyme recognition site.
[0108] Confidence” herein refers to the statistical likelihood that the called SNP, SNV, variant, copy number, etc. correctly represents the real genetic state of the individual.II. Minimal Residual Disease Detection
[0109] The goal of a minimum residual disease (MRD) assay is to detect and / or quantify circulating tumor (ctDNA) so researchers and clinicians can detect recurrence early and monitor the progress of the disease throughout or following treatment or during remission. In general, an MRD assay will rely on patient-specific and tumor-specific panels (i.e., a “panel” or “signature panel” or “a set of somatic variants”) for assessing the presence of ctDNA in a patient sample. The panel can be prepared with the general steps of : (1) profiling a tumor or cancer sample from a patient and (2) identifying a subset of somatic mutations to target, and, at one or more later time points, (3) taking a subsequent sample from the patient, (4) enriching cell-free DNA (cfDNA) for the target somatic mutation sites, and (5) determining or estimating the ctDNA content of the cfDNA given the tumor profile and sequencing data.
[0110] The disclosed methods allow for generating more sensitive patient-specific panels while reducing cost and processing time. As described in further detail below, the method includes targeting regions of low utility to specifically deplete these regions prior to the sequencing in order to remove noisy genetic data. Specific aspects of this targeted depletion steps are discussed in more detail below.
[0111] Depleting regions of low utility (e.g., via targeted depletion) or unnecessary DNA (e.g., via restriction enzyme digestion) can improve MRD methodologies that rely on either whole genome sequencing (WGS) or whole exome sequencing (WES). WGS- based MRD methods require sequencing of the patient’s entire genome to generate the subject-specific panels and to possibly also track the presence of said mutations in the sequencing libraries derived from the subject’s tumor and non-tumor samples. These methods are highly costly and time-intensive, but they are also highly sensitive and able to identify a high number of somatic variants of interest based on the high level ofsequencing coverage. WES-based MRD methods depend upon sequencing only the coding regions of genomic DNA, which make up around 1% of the whole human genome. These methods are much cheaper and faster to perform; however, there is a dramatically reduced capacity to identify somatic variants of interest as there is much lower sequencing coverage. MRD assays have higher performance when there are more high confidence somatic variants that are identified and tracked; thus, while WES-based MRD assays may be cheaper and quicker than WGS-based MRD assays, WES-based MRD assays will have greatly reduced performance when compared to WGS-based MRD assays.
[0112] Removing regions of low utility (e.g., via targeted depletion) and / or unnecessary DNA (e.g., via restriction enzyme digestion) from the sequencing libraries prepared from a subject’s samples means that less sequencing is required to be performed as compared to WGS-based MRD methods. This results in improved cost-efficiency and processing time over conventional whole genome sequencing assays. Additionally, those practicing the disclosed methods can perform sequencing on depleted DNA samples at a greater depth while maintaining or decreasing cost compared to WGS. Increasing the depth of sequencing can improve detection of unique somatic variants, which may ultimately improve variant calling during subsequent MRD analysis. Depleting or removing regions of DNA that are prone to sequencing errors, such as regions with high or low GC content, can also reduce the burden (both time and cost) of post-sequencing analysis and improve calling by reducing the amount of sequencing errors that may otherwise be erroneously counted as tumor-specific somatic mutations. These benefits are significant given the context of using the disclosed methods to track potential cancer recurrence or responsiveness to treatment because timely results can make a difference in patient survival.
[0113] Targeted depletion and / or reducing the size of the DNA sample may be done (a) during the initial profiling of the tumor or cancer sample from the patient in order to generate the tumor-specific panel of somatic mutations, (b) during the enriching of the cfDNA from the subsequent non-tumor sample from the sample, or (c) during both of these parts of the MRD assay (FIGs. 1A-1C).
[0114] Removing regions of low utility (e.g., via targeted depletion) and / or unnecessary DNA (e.g., via restriction enzyme digestion) can be controlled in order toensure that enough DNA is being sequenced in order to identify enough high confidence somatic variants to ensure the MRD assay has a high level of performance. This results in an improvement over WES-based MRD methods as WES relies on a very limited portion of the genome to identify variants, which reduces the possible level of performance.
[0115] Taken together, DNA depletion allows for balancing between cost- and timeconsiderations and performance requirements. Both of these benefits are significant given the context of using the disclosed methods to track potential cancer recurrence or responsiveness to treatment because timely results generated using a high-performing assay can enhance patient survival by more accurately and more rapidly detecting disease presence.
[0116] In sum, the disclosed methods provide tangible and meaningful improvements over conventional MRD analysis by improving variant calling, decreasing cost and time requirements, and generally improving the efficiency in preparing patient-specific panels for MRD analysis.
[0117] Once a patient-specific and tumor-specific panel (i.e., a “panel”) has been established, such a panel can be used to enrich ctDNA in subsequent samples taken from the cancer patient. The subsequent samples may be taken from a patient at various time points during the course of treatment or during a period of remission. For example, after a surgical removal of a tumor, the tumor may be profiled as described herein to determine tumor-specific somatic mutations, and at one or more subsequent time points, a subsequent sample may be taken from the subject to search for the presence of any ctDNA comprising any one of the identified tumor-specific somatic mutations. The detection or presence of ctDNA comprising a tumor-specific somatic mutation may be indicative of cancer recurrence. Additionally or alternatively, similar assessment can be performed throughout the course of a patient’s treatment (e.g., with chemotherapy, radiation, immunotherapy, cell therapy, etc.) to detect or quantify ctDNA and determine whether the amount of ctDNA is increasing or decreasing, as this may be indicative of responsiveness to the therapy. Accordingly, assessment of a subsequent sample may be repeated 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more times throughout the course of a patient’s remission or treatment. The assessment of a subsequent sample may be repeated monthly, every other month, once every three months, once every four months, once every fivemonths, once every six months, once every seven months, once every eight months, once every nine months, once every ten months, once every eleven months, or annually.
[0118] The type of sample used for the one or more subsequent samples is generally a blood sample, a plasma sample, or a serum sample, but any biological sample that contains cfDNA or potentially contains ctDNA would be acceptable. In some embodiments, the one or more subsequent samples may be cell-free samples.
[0119] Enrichment of ctDNA in the one or more subsequent samples can be performed by methods including, but not limited to, hybrid capture-based enrichment, PCR-target enrichment, or on-sequencer enrichment. Briefly, enrichment may comprise extracting cfDNA from a subsequent sample taken from the cancer patient and contacting the extracted cfDNA with a plurality of oligonucleotides (i.e., oligonucleotide probes), wherein each oligonucleotide in the plurality of oligonucleotides comprises a nucleic acid sequence that is capable of hybridizing to a cfDNA fragment comprising one of the tumor-specific mutation sequences identified by comparing the sequences of the patient’s tumor DNA and non-tumor DNA. Thus, enrichment may utilize a set of oligonucleotide probes to selectively enrich ctDNA that may be in the subsequent sample by binding to previously identified tumor-specific somatic mutation sequences.
[0120] Enrichment of ctDNA in the one or more subsequent samples can be also performed by methods including, but not limited to, depleting DNA fragments that comprise low utility from the cfDNA derived from the one or more subsequent samples or reducing the size of the cfDNA sample from the one or more subsequent samples by digesting the cfDNA sample with at least one restriction enzyme. Briefly, enrichment may comprise extracting cfDNA from a subsequent sample taken from the cancer patient and processing the sample via the targeted depletion or genome size reduction methods described below.
[0121] A panel may comprise 10-5,000 tumor-specific somatic mutations. For example, a panel may comprise 10-4,000, 10-3,000, 10-2,500, 10-2,000, 10-1,500, 10- 1,000, 10-950, 10-900, 10-850, 10-800, 10-750, 10-700, 10-650, 10-600, 10-550, 10-500, 50-5,000, 50-4,000, 50-3,000, 50-2,500, 50-2,000, 50-1,500, 50-1,000, 50-950, 50-900, 50-850, 50-800, 50-750, 50-700, 50-650, 50-600, 50-550, 50-500, 100-5,000, 100-4,000, 100-3,000, 100-2,500, 10-2,000, 100-1,500, 100-1,000, 100-950, 100-900, 100-850, 100-800, 100-750, 100-700, 100-650, 100-600, 100-550, 100-500, 200-5,000, 200-4,000, 200- 3,000, 200-2,500, 200-2,000, 200-1,500, 200-1,000, 200-950, 200-900, 200-850, 200- 800, 200-750, 200-700, 200-650, 200-600, 200-550, 200-500, 300-5,000, 300-4,000, 300- 3,000, 300-2,500, 300-2,000, 300-1,500, 300-1,000, 300-950, 300-900, 300-850, SOO- SOO, 300-750, 300-700, 300-650, 300-600, 300-550, 300-500, 400-5,000, 400-4,000, 400- 3,000, 400-2,500, 400-2,000, 400-1,500, 400-1,000, 400-950, 400-900, 400-850, 400- 800, 400-750, 400-700, 400-650, 400-600, 400-550, 400-500, 500-5,000, 500-4,000, 500- 3,000, 500-2,500, 500-2,000, 500-1,500, 500-1,000, 500-950, 500-900, 500-850, 500- 800, 500-750, 500-700, 500-650, 500-600, or 500-550 tumor-specific somatic mutations. In some embodiments, a panel may comprise or consist of about 10, about 20, about 30, about 40, about 50, about 75, about 100, about 150, about 200, about 250, about 300, about 350, about 400, about 450, about 500, about 550, about 600, about 650, about 700, about 750, about 800, about 850, about 900, about 950, about 1,000, about 1,100, about 1,150, about 1,200, about 1,250, about 1,300, about 1,350, about 1,400, about 1,450, about 1,500, about 1,550, about 1,600, about 1,650, about 1,700, about 1,750, about 1,800, about 1,850, about 1,900, about 1,950, about 2,000 or more tumor-specific somatic mutations. In some embodiments, a panel may comprise at least 10, at least 20, at least 30, at least 40, at least 50, at least 75, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1,000, at least 1,050, at least 1,100, at least 1,150, at least 1,200, at least 1,250, at least 1,300, at least 1,350, at least 1,400, at least 1,450, at least 1,500, at least 1,550, at least 1,600, at least 1,650, at least 1,700, at least 1,750, at least 1,800, at least 1,850, at least 1,900, at least 1,950, or at least 2,000 tumor-specific somatic mutations. The tumorspecific somatic mutations may be in introns, exons, or a combination thereof. In some embodiments, the tumor-specific mutations may be one or more somatic mutations selected from SNVs, SNPs, insertions, deletions, and translocations.
[0122] After enrichment or concurrently with enrichment, the enriched DNA is sequenced. This sequencing may be performed by, for example, Next Generation Sequencing (NGS). Deep sequencing may allow for more sensitive detection, and so the depth of the sequencing may be at least 50X, at least 100X, at least 150X, at least 200X, at least 250X, at least 300X, at least 350X, at least 400X, at least 450X, at least 500X, at least 550X, at least 600X, at least 650X, at least 700X, at least 750X, at least 800X, atleast 850X, at least 900X, at least 950X, or at least l,000X. In other words, the depth of the sequencing may be about 50X, about 100X, about 150X, about 200X, about 250X, about 300X, about 350X, about 400X, about 450X, about 500X, about 550X, about 600X, about 650X, about 700X, about 750X, about 800X, about 850X, about 900X, about 950X, or about l,000X. The detection sensitivity of about 20 to about 50 ctDNA fragments comprising one or more of the set of somatic mutations in the fluid sample per a total background of about 500,000 cfDNA fragments.
[0123] The disclosed methods may be used for tracking and assessing recurrence in any cancer patient. For example, the cancer patient may have a cancer selected from, but not limited to, adrenal cancer, anal cancer, bile duct cancer, bladder cancer, bone cancer, a brain / CNS tumor, breast cancer, Castelman disease, cervical cancer, colon or rectum cancer, endometrial cancer, esophagus cancer, a Ewing tumor, eye cancer, gallbladder cancer, a gastrointestinal carcinoid tumor, a gastrointestinal stromal tumor (GIST), gestational trophoblastic disease, Hodgkin disease, Kaposi sarcoma, kidney cancer, laryngeal and hypopharyngeal cancer, leukemia, liver cancer, lung cancer, lymphoma, malignant mesothelioma, multiple myeloma, myelodysplastic syndrome, nasal cavity or paranasal sinus cancer, nasopharyngeal cancer, neuroblastoma, oral cavity or oropharyngeal cancer, osteosarcoma, ovarian cancer, pancreatic cancer, penile cancer, a pituitary tumor, prostate cancer, retinoblastoma, rhabdomyosarcoma, salivary gland cancer, skin cancer, small intestine cancer, stomach cancer, testicular cancer, thymus cancer, thyroid cancer, uterine sarcoma, vaginal cancer, vulvar cancer, Waldenstrom macroglobulinemia, and Wilms tumor. In some embodiments, the cancer may be a bloodborne or hematological cancer, such as leukemia or lymphoma.
[0124] The disclosed MRD assay, specifically the obtaining and testing of subsequent samples from a cancer patient, may be repeated one or more times following completion of a cancer treatment; one or more times while the cancer patient is in remission; one or more times coinciding with or prior to surgery; following, during, or prior to administration of chemotherapy; following, during, or prior to radiation therapy; following, during, or prior to immunotherapy; or following, during, or prior to cell therapy. The disclosed MRD assay may also be repeated at times prior to, coinciding with, and / or following an imaging test, such as a PET scan, a PET / CT scan, an MRI, or an X-ray.
[0125] Provided herein are methods of preparing a patient-specific panel of tumorspecific somatic mutations, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample; (b) depleting from the first DNA sample and the second DNA sample any DNA fragments that comprise a region of low utility; (c) sequencing the first depleted DNA sample and the second depleted DNA sample; and (d) generating a patientspecific panel of tumor-specific somatic mutations that are present in the first DNA depleted sample and absent in the second depleted DNA sample. The method may further comprise preparing a plurality of oligonucleotide probes, wherein each probe in the plurality of oligonucleotide probes hybridizes to a tumor-specific somatic mutation in the patient-specific signature panel.
[0126] Provided herein are methods of preparing a patient-specific panel of tumorspecific somatic mutations, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample; (b) reducing the size of the first DNA sample and the second DNA sample by digesting the first DNA sample and the second DNA sample with at least one restriction enzyme ; (c) sequencing the first reduced DNA sample and the second reduced DNA sample; and (d) generating a patient-specific panel of tumor specific somatic mutations that are present in the first reduced DNA sample and absent in the second reduced DNA sample. The method may further comprise preparing a plurality of oligonucleotide probes, wherein each probe in the plurality of oligonucleotide probes hybridizes to a tumor-specific somatic mutation in the patient-specific signature panel.
[0127] Further provided herein are methods of preparing a patient-specific panel of tumor-specific somatic mutations, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample; (b) sequencing the first DNA sample and the second DNA sample; and (c) generating a patient-specific panel of tumor-specific somatic mutations that are present in the first DNA sample and absent in the second DNA sample. The method may further comprise preparing a plurality of oligonucleotide probes, wherein each probe in the plurality of oligonucleotide probes hybridizes to a tumor-specific somatic mutation in the patient-specific signature panel.
[0128] Also provided herein are methods of preparing a circulating tumor DNA (ctDNA)-enriched fraction from a biological sample, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample; (b) depleting from the first DNA sample and the second DNA sample any DNA fragments that comprise a region of low utility; (c) sequencing the first depleted DNA sample and the second depleted DNA sample; (d) obtaining a second non-tumor sample from the subject; (e) extracting cell-free DNA (cfDNA) from the second non-tumor sample; and (f) enriching the cfDNA from the second non-tumor sample with a plurality of oligonucleotide probes, wherein each probe in the plurality of oligonucleotide probes is capable of hybridizing to a DNA fragment comprising a tumor-specific somatic mutation that is present in the first depleted DNA sample and absent in the second depleted DNA sample, thereby obtaining a DNA fraction that is enriched for ctDNA. Additionally provided herein are methods of preparing a circulating tumor DNA (ctDNA)-enriched fraction from a biological sample, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample; (b) reducing the size of the first DNA sample and the second DNA sample by digesting the first DNA sample and the second DNA sample with at least one restriction enzyme; (c) sequencing the first reduced DNA sample and the second reduced DNA sample; (d) obtaining a second non- tumor sample from the subject; (e) extracting cell-free DNA (cfDNA) from the second non-tumor sample; and (f) enriching the cfDNA from the second non-tumor sample with a plurality of oligonucleotide probes, wherein each probe in the plurality of oligonucleotide probes is capable of hybridizing to a DNA fragment comprising a tumorspecific somatic mutation that is present in the first reduced DNA sample and absent in the second reduced DNA sample, thereby obtaining a DNA fraction that is enriched for ctDNA. Also provided herein is the enriched ctDNA fraction prepared according to the methods disclosed herein. In some embodiments, the plurality of oligonucleotide probes used to prepare the enriched ctDNA fraction may be capable of detecting at least 10, at least 50, at least 100, at least 150, at least 200, at least 250, or at least 500 tumor-specific somatic mutations. In some embodiments, the enriched ctDNA fraction may include cfDNA fragments that collectively include at least 10, at least 50, at least 100, at least 150, at least 200, at least 250, or at least 500 tumor-specific somatic mutations. In some embodiments, the enriched ctDNA fraction may be enriched for cfDNA fragments averaging less than about 200 base pairs in length.
[0129] Provided herein are methods of preparing a circulating tumor DNA (ctDNA)- enriched fraction from a biological sample, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample; (b) depleting from the first DNA sample and the second DNA sample any DNA fragments that comprise a region of low utility; (c) sequencing the first depleted DNA sample and the second depleted DNA sample; (c) sequencing the first depleted DNA sample and the second depleted DNA sample; (d) obtaining a second non-tumor sample from the subject; (e) extracting cell-free DNA (cfDNA) from the second non-tumor sample; and (f) depleting DNA fragments that comprise a region of low utility from the cfDNA, wherein the remaining DNA fragments comprise the sequences having tumor-specific somatic mutations present in the first depleted DNA sample and absent in the second depleted DNA sample, thereby obtaining a DNA fraction that is enriched for ctDNA. Additionally provided herein are methods of preparing a circulating tumor DNA (ctDNA)-enriched fraction from a biological sample, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample; (b) reducing the size of the first DNA sample and the second DNA sample by digesting the first DNA sample and the second DNA sample with at least one restriction enzyme; (c) sequencing the first reduced DNA sample and the second reduced DNA sample; (d) obtaining a second non-tumor sample from the subject; (e) extracting cell-free DNA (cfDNA) from the second non-tumor sample; and (f) reducing the size of the cfDNA sample by digesting the cfDNA sample with at least one restriction enzyme, wherein the remaining DNA fragments comprise the sequences having tumor-specific somatic mutations present in the first reduced DNA sample and absent in the second reduced DNA sample, thereby obtaining a DNA fraction that is enriched for ctDNA. Also provided herein is the enriched ctDNA fraction prepared according to the methods disclosed herein. In some embodiments, the enriched ctDNA fraction may include cfDNA fragments that collectively include at least 10, at least 50, at least 100, at least 150, at least 200, at least 250, or at least 500 tumor-specific somatic mutations. In some embodiments, the enriched ctDNA fraction may be enriched for cfDNA fragments averaging less than about 200 base pairs in length.
[0130] Provided herein are methods of preparing a circulating tumor DNA (ctDNA)- enriched fraction from a biological sample, comprising: (a) obtaining from a subject witha prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample; (b) sequencing the first DNA sample and the second DNA sample; (c) obtaining a second non-tumor sample from the subject; (d) extracting cell-free DNA (cfDNA) from the second non-tumor sample; and (e) depleting DNA fragments that comprise a region of low utility from the cfDNA, wherein the remaining DNA fragments comprise the sequences having tumor-specific somatic mutations present in the first DNA sample and absent in the second DNA sample, thereby obtaining a DNA fraction that is enriched for ctDNA. Additionally provided herein are methods of preparing a circulating tumor DNA (ctDNA)-enriched fraction from a biological sample, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample; (b) sequencing the first DNA sample and the second DNA sample; (c) obtaining a second non-tumor sample from the subject; (d) extracting cell-free DNA (cfDNA) from the second non-tumor sample; and (e) reducing the size of the cfDNA sample by digesting the cfDNA sample with at least one restriction enzyme, wherein the remaining DNA fragments comprise the sequences having tumor-specific somatic mutations present in the first DNA sample and absent in the second DNA sample, thereby obtaining a DNA fraction that is enriched for ctDNA. Also provided herein is the enriched ctDNA fraction prepared according to the methods disclosed herein. In some embodiments, the enriched ctDNA fraction may include cfDNA fragments that collectively include at least 10, at least 50, at least 100, at least 150, at least 200, at least 250, or at least 500 tumor-specific somatic mutations. In some embodiments, the enriched ctDNA fraction may be enriched for cfDNA fragments averaging less than about 200 base pairs in length.
[0131] Provided herein are methods of detecting circulating tumor DNA (ctDNA) in a sample, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample; (b) depleting from the first DNA sample and the second DNA sample any DNA fragments that comprise a region of low utility; (c) sequencing the first DNA sample and the second DNA sample; (d) obtaining a further non-tumor sample from the subject; (e) extracting cell-free DNA (cfDNA) from the second non-tumor sample; (f) enriching the cfDNA from the second non-tumor sample with a plurality of oligonucleotide probes, wherein each probe in a plurality of oligonucleotide probes is capable of hybridizing to a DNA fragment comprising a tumor-specific somatic mutation that is present in the firstDNA sample and absent in the second DNA sample, thereby obtaining an enriched DNA fraction; (g) sequencing the enriched DNA fraction, thereby obtaining a plurality of sequence reads; and (h) detecting the presence or absence of ctDNA, wherein the presence of a sequence read in the plurality of sequence reads comprising a nucleic acid sequence comprising one or more of the tumor-specific somatic mutations indicates the presence of ctDNA. Additionally provided herein are methods of monitoring cancer recurrence, response to treatment, or disease progression in a subject previously treated for cancer, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample; (b) depleting from the first DNA sample and the second DNA sample any DNA fragments that comprise a region of low utility; (c) sequencing the first DNA sample and the second DNA sample; (d) obtaining a further non-tumor sample from the subject; (e) extracting cell-free DNA (cfDNA) from the second non-tumor sample; (f) enriching the cfDNA from the second non-tumor sample with a plurality of oligonucleotide probes, wherein each probe in the plurality of oligonucleotide probes is capable of hybridizing to a DNA fragment comprising a tumor-specific somatic mutation that is present in the first DNA sample and absent in the second DNA sample, thereby obtaining an enriched DNA fraction; (g) sequencing the presence or absence of any sequence reads in the plurality of sequences reads; and (h) detecting the presence or absence of any sequence reads in the plurality of sequence reads, wherein the presence of a sequence read comprising a nucleic acid sequence comprising one or more of the tumor-specific somatic mutations indicates (i) cancer recurrence, (ii) a lack of responsiveness to treatment, (iii) disease progression, or (iv) any combination thereof. In some embodiments of these two methods, steps (d) through (g) are repeated at one or more times during a cancer treatment, following completion of a cancer treatment, while the subject is in remission, or coinciding with or prior to surgery.
[0132] Provided herein are methods of detecting circulating tumor DNA (ctDNA) in a sample, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample;(b) depleting from the first DNA sample and the second DNA sample any DNA fragments that comprise a region of low utility; (c) sequencing the first DNA sample and the second DNA sample; (d) obtaining a further non-tumor sample from the subject; (e) extracting cell-free DNA (cfDNA) from the second non-tumor sample; (f) depleting DNAfragments that comprise a region of low utility from the cfDNA, wherein the remaining DNA fragments comprise the sequences having tumor-specific somatic mutations present in the first depleted DNA sample and absent in the second depleted DNA sample, thereby obtaining a DNA fraction that is enriched for ctDNA; (g) sequencing the enriched DNA fraction, thereby obtaining a plurality of sequence reads; and (h) detecting the presence or absence of ctDNA, wherein the presence of a sequence read in the plurality of sequence reads comprising a nucleic acid sequence comprising one or more of the tumor-specific somatic mutations indicates the presence of ctDNA. Additionally provided herein are methods of monitoring cancer recurrence, response to treatment, or disease progression in a subject previously treated for cancer, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample; (b) depleting from the first DNA sample and the second DNA sample any DNA fragments that comprise a region of low utility; (c) sequencing the first DNA sample and the second DNA sample; (d) obtaining a further non-tumor sample from the subject; (e) extracting cell-free DNA (cfDNA) from the second non-tumor sample; (f) depleting DNA fragments that comprise a region of low utility from the cfDNA, wherein the remaining DNA fragments comprise the sequences having tumor-specific somatic mutations present in the first depleted DNA sample and absent in the second depleted DNA sample, thereby obtaining a DNA fraction that is enriched for ctDNA; (g) sequencing the enriched DNA fraction, thereby obtaining a plurality of sequence reads; and (h) detecting the presence or absence of any sequence reads in the plurality of sequence reads, wherein the presence of a sequence read comprising a nucleic acid sequence comprising one or more of the tumor-specific somatic mutations indicates (i) cancer recurrence, (ii) a lack of responsiveness to treatment, (iii) disease progression, or (iv) any combination thereof. In some embodiments of these two methods, steps (d) through (g) are repeated at one or more times during a cancer treatment, following completion of a cancer treatment, while the subject is in remission, or coinciding with or prior to surgery.
[0133] Provided herein are methods of detecting circulating tumor DNA (ctDNA) in a sample, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample; (b) sequencing the first DNA sample and the second DNA sample; (c) obtaining a further non-tumor sample from the subject; (d) extracting cell-free DNA (cfDNA) from thesecond non-tumor sample; (e) depleting DNA fragments that comprise a region of low utility from the cfDNA, wherein the remaining DNA fragments comprise the sequences having tumor-specific somatic mutations present in the first depleted DNA sample and absent in the second depleted DNA sample, thereby obtaining a DNA fraction that is enriched for ctDNA; (f) sequencing the enriched DNA fraction, thereby obtaining a plurality of sequence reads; and (g) detecting the presence or absence of ctDNA, wherein the presence of a sequence read in the plurality of sequence reads comprising a nucleic acid sequence comprising one or more of the tumor-specific somatic mutations indicates the presence of ctDNA. Additionally provided herein are methods of monitoring cancer recurrence, response to treatment, or disease progression in a subject previously treated for cancer, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample; (b) sequencing the first DNA sample and the second DNA sample; (c) obtaining a further non-tumor sample from the subject; (d) extracting cell-free DNA (cfDNA) from the second non-tumor sample; (e) depleting DNA fragments that comprise a region of low utility from the cfDNA, wherein the remaining DNA fragments comprise the sequences having tumor-specific somatic mutations present in the first depleted DNA sample and absent in the second depleted DNA sample, thereby obtaining a DNA fraction that is enriched for ctDNA; (f) sequencing the presence or absence of any sequence reads in the plurality of sequences reads; and (g) detecting the presence or absence of any sequence reads in the plurality of sequence reads, wherein the presence of a sequence read comprising a nucleic acid sequence comprising one or more of the tumor-specific somatic mutations indicates (i) cancer recurrence, (ii) a lack of responsiveness to treatment, (iii) disease progression, or (iv) any combination thereof. In some embodiments of these two methods, steps (c) through (g) are repeated at one or more times during a cancer treatment, following completion of a cancer treatment, while the subject is in remission, or coinciding with or prior to surgery.
[0134] Provided herein are methods of detecting circulating tumor DNA (ctDNA) in a sample, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample;(b) reducing the size of the first DNA sample and the second DNA sample by digesting the first DNA sample and the second DNA sample with at least one restriction enzyme; ;(c) sequencing the first DNA sample and the second DNA sample; (d) obtaining a furthernon-tumor sample from the subject; (e) extracting cell-free DNA (cfDNA) from the second non-tumor sample; (f) enriching the cfDNA from the second non-tumor sample with a plurality of oligonucleotide probes, wherein each probe in a plurality of oligonucleotide probes is capable of hybridizing to a DNA fragment comprising a tumorspecific somatic mutation that is present in the first DNA sample and absent in the second DNA sample, thereby obtaining an enriched DNA fraction; (g) sequencing the enriched DNA fraction, thereby obtaining a plurality of sequence reads; and (h) detecting the presence or absence of ctDNA, wherein the presence of a sequence read in the plurality of sequence reads comprising a nucleic acid sequence comprising one or more of the tumorspecific somatic mutations indicates the presence of ctDNA. Additionally provided herein are methods of monitoring cancer recurrence, response to treatment, or disease progression in a subject previously treated for cancer, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample; (b) reducing the size of the first DNA sample and the second DNA sample by digesting the first DNA sample and the second DNA sample with at least one restriction enzyme; (c) sequencing the first DNA sample and the second DNA sample; (d) obtaining a further non-tumor sample from the subject; (e) extracting cell-free DNA (cfDNA) from the second non-tumor sample; (f) enriching the cfDNA from the second non-tumor sample with a plurality of oligonucleotide probes, wherein each probe in the plurality of oligonucleotide probes is capable of hybridizing to a DNA fragment comprising a tumor-specific somatic mutation that is present in the first DNA sample and absent in the second DNA sample, thereby obtaining an enriched DNA fraction; (g) sequencing the presence or absence of any sequence reads in the plurality of sequences reads; and (h) detecting the presence or absence of any sequence reads in the plurality of sequence reads, wherein the presence of a sequence read comprising a nucleic acid sequence comprising one or more of the tumor-specific somatic mutations indicates (i) cancer recurrence, (ii) a lack of responsiveness to treatment, (iii) disease progression, or (iv) any combination thereof. In some embodiments of these two methods, steps (d) through (h) are repeated at one or more times during a cancer treatment, following completion of a cancer treatment, while the subject is in remission, or coinciding with or prior to surgery.
[0135] Provided herein are methods of detecting circulating tumor DNA (ctDNA) in a sample, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a firstDNA sample from a tumor sample and a second DNA sample from a non-tumor sample;(b) reducing the size of the first DNA sample and the second DNA sample by digesting the first DNA sample and the second DNA sample with at least one restriction enzyme;(c) sequencing the first reduced DNA sample and the second reduced DNA sample; (d) obtaining a further non-tumor sample from the subject; (e) extracting cell-free DNA (cfDNA) from the second non-tumor sample; (f) reducing the size of the cfDNA sample by digesting the cfDNA sample with at least one restriction enzyme, wherein the remaining DNA fragments comprise the sequences having tumor-specific somatic mutations present in the first reduced DNA sample and absent in the second reduced DNA sample, thereby obtaining a DNA fraction that is enriched for ctDNA; (g) sequencing the enriched DNA fraction, thereby obtaining a plurality of sequence reads; and (h) detecting the presence or absence of ctDNA, wherein the presence of a sequence read in the plurality of sequence reads comprising a nucleic acid sequence comprising one or more of the tumor-specific somatic mutations indicates the presence of ctDNA. Additionally provided herein are methods of monitoring cancer recurrence, response to treatment, or disease progression in a subject previously treated for cancer, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample; (b) reducing the size of the first DNA sample and the second DNA sample by digesting the first DNA sample and the second DNA sample with at least one restriction enzyme; (c) sequencing the first reduced DNA sample and the second reduced DNA sample; (d) obtaining a further non- tumor sample from the subject; (e) extracting cell-free DNA (cfDNA) from the second non-tumor sample; (f) reducing the size of the cfDNA sample by digesting the cfDNA sample with at least one restriction enzyme, wherein the remaining DNA fragments comprise the sequences having tumor-specific somatic mutations present in the first reduced DNA sample and absent in the second reduced DNA sample, thereby obtaining a DNA fraction that is enriched for ctDNA; (g) sequencing the presence or absence of any sequence reads in the plurality of sequences reads; and (h) detecting the presence or absence of any sequence reads in the plurality of sequence reads, wherein the presence of a sequence read comprising a nucleic acid sequence comprising one or more of the tumorspecific somatic mutations indicates (i) cancer recurrence, (ii) a lack of responsiveness to treatment, (iii) disease progression, or (iv) any combination thereof. In some embodiments of these two methods, steps (d) through (h) are repeated at one or moretimes during a cancer treatment, following completion of a cancer treatment, while the subject is in remission, or coinciding with or prior to surgery.
[0136] Provided herein are methods of detecting circulating tumor DNA (ctDNA) in a sample, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample; (b) sequencing the first DNA sample and the second DNA sample; (c) obtaining a further non-tumor sample from the subject; (d) extracting cell-free DNA (cfDNA) from the second non-tumor sample; (e) reducing the size of the cfDNA sample by digesting the cfDNA sample with at least one restriction enzyme, wherein the remaining DNA fragments comprise the sequences having tumor-specific somatic mutations present in the first DNA sample and absent in the second DNA sample, thereby obtaining a DNA fraction that is enriched for ctDNA; (f) sequencing the enriched DNA fraction, thereby obtaining a plurality of sequence reads; and (g) detecting the presence or absence of ctDNA, wherein the presence of a sequence read in the plurality of sequence reads comprising a nucleic acid sequence comprising one or more of the tumor-specific somatic mutations indicates the presence of ctDNA. Additionally provided herein are methods of monitoring cancer recurrence, response to treatment, or disease progression in a subject previously treated for cancer, comprising: (a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample; (b) sequencing the first DNA sample and the second DNA sample; (c) obtaining a further non-tumor sample from the subject; (d) extracting cell-free DNA (cfDNA) from the second non-tumor sample; (e) reducing the size of the cfDNA sample by digesting the cfDNA sample with at least one restriction enzyme, wherein the remaining DNA fragments comprise the sequences having tumor-specific somatic mutations present in the first DNA sample and absent in the second DNA sample, thereby obtaining a DNA fraction that is enriched for ctDNA; (f) sequencing the presence or absence of any sequence reads in the plurality of sequences reads; and (g) detecting the presence or absence of any sequence reads in the plurality of sequence reads, wherein the presence of a sequence read comprising a nucleic acid sequence comprising one or more of the tumor-specific somatic mutations indicates (i) cancer recurrence, (ii) a lack of responsiveness to treatment, (iii) disease progression, or (iv) any combination thereof. In some embodiments of these two methods, steps (c) through (g) are repeated at one ormore times during a cancer treatment, following completion of a cancer treatment, while the subject is in remission, or coinciding with or prior to surgery.
[0137] In some embodiments of any of the methods disclosed herein, the region of low utility may be repetitive elements, low GC content, high GC content, regions lacking somatic mutations of interest, or any combination thereof. In some embodiments of any of the methods disclosed herein, depleting from the first DNA sample and the second DNA sample any DNA fragments that include a region of low utility may include nuclease-based depletion, restriction enzyme-based depletion, subtractive hybridization, or kinetic-based depletion. In some further embodiments of any of the methods disclosed herein, nuclease-based depletion includes targeting depletion of sequences with a DNA- or RNA-guided endonuclease. In some further embodiments of any of the methods disclosed herein, the DNA-guided endonuclease may be Thermus thermophilus Argonaute (TtAgo) or a transcription activator-like effector nuclease (TALEN). In some further embodiments of the methods disclosed herein, the RNA-guided endonuclease may be a Cas endonuclease. In some further embodiments, nuclease-based depletion may be Depletion of Abundant Sequences by Hybridization (DASH). In some further embodiments of the methods disclosed herein, kinetic-based depletion includes depleting DNA fragments based on denaturation and annealing kinetics. In some further embodiments of the methods disclosed herein, restriction enzyme-based depletion includes reduced representation sequencing.
[0138] In some embodiments of any of the methods disclosed herein, reducing the size of a DNA sample may include restriction enzyme-based depletion.
[0139] In some embodiments of any of the methods disclosed herein, the tumorspecific somatic mutations may be obtained by aligning sequences of the first DNA sample to a reference human genome that is not from the subject, aligning sequences of the second DNA sample to a reference human genome that is not from the subject, and selecting mutations that are present in the first DNA sample and absent in the second DNA sample. In some embodiments of any of the methods disclosed herein, the tumorspecific somatic mutation may be obtained by aligning sequences of the first DNA sample to sequences of the second DNA sample. In some embodiments of any of the methods disclosed herein, the tumor-specific somatic mutation(s) may include one or more somatic mutations selected from SNVs, insertions, deletions, and translocations. Insome embodiments of any of the methods disclosed herein, the first non-tumor sample may be a tissue sample matched to a tissue of origin of the tumor sample. In some embodiments of any of the methods disclosed herein, the first non-tumor sample may be a fluid sample selected from a buffy coat sample, blood, blood plasma, blood serum, urine, saliva, and cerebral spinal fluid (CSF). In some embodiments of any of the methods disclosed herein, the plurality of oligonucleotide probes may be capable of detecting at least 10, at least 50, at least 100, at least 150, at least 200, at least 250, or at least 500 tumor-specific somatic mutations. In some embodiments of any of the methods disclosed herein, enriching the cfDNA from the second non-tumor sample may include (i) hybrid capture-based enrichment, (ii) PCR-target enrichment, or (iii) on-sequencer enrichment.
[0140] In some embodiments of any of the methods disclosed herein, the further non- tumor sample may include a fluid sample selected from a buffy coat sample, blood, blood plasma, blood serum, urine, saliva, and cerebral spinal fluid (CSF). In some embodiments of any of the methods disclosed herein, sequencing the enriched DNA fraction may include whole genome sequencing or targeted sequencing. In some embodiments of any of the methods disclosed herein, the targeted sequencing may include sequencing of introns, exons, intergenic regions, or a combination thereof.
[0141] In some embodiments of any of the methods disclosed herein, the cancer may be selected from adrenal cancer, anal cancer, bile duct cancer, bladder cancer, bone cancer, a brain / CNS tumor, breast cancer, Castleman disease, cervical cancer, colon or rectum cancer, endometrial cancer, esophagus cancer, a Ewing tumor, eye cancer, gallbladder cancer, a gastrointestinal carcinoid tumor, a gastrointestinal stromal tumor (GIST), gestational trophoblastic disease, Hodgkin disease, Kaposi sarcoma, kidney cancer, laryngeal and hypopharyngeal cancer, leukemia, liver cancer, lung cancer, lymphoma, malignant mesothelioma, multiple myeloma, myelodysplastic Syndrome, nasal cavity or paranasal sinus cancer, nasopharyngeal cancer, neuroblastoma, oral cavity or oropharyngeal cancer, osteosarcoma, ovarian cancer, pancreatic cancer, penile cancer, a pituitary tumor, prostate cancer, retinoblastoma, rhabdomyosarcoma, salivary gland cancer, skin cancer, small intestine cancer, stomach cancer, testicular cancer, thymus cancer, thyroid cancer, uterine sarcoma, vaginal cancer, vulvar cancer, Waldenstrom macroglobulinemia, and Wilms tumor.
[0142] The disclosed methods allow for detecting ctDNA or determining the tumor fraction from a biological sample from a patient that has, previously had, or is suspected of having cancer. As described in further detail below, the methods can be represented by two phases. In a first phase, or enrollment phase, a set of somatic mutations or variants that are specific to a patient are identified, and then filtered to generate a subset of somatic variants or mutations that include only specific types of somatic mutations or variants or show a preference for specific types of somatic mutations or variants. For the purposes of the disclosed methods, the subset of somatic mutations or variants will comprise or consist of single nucleotide variants, single nucleotide polymorphisms, multinucleotide variants, small indels, and genomic rearrangements for the reasons described herein. A panel of capture probes is then generated that are specific to the subset panel of somatic mutations or variants, which can be used to enrich a sample before sequencing. In a second phase, tumors may be detected or monitored by analyzing cfDNA using the set of somatic mutations or variants that are specific to the patient and the panel of capture probes.
[0143] Specific aspects of the MRD processes are discussed in more detail below.III. Phase I - Panel of Markers / Mutations and Capture Probes a. DNA Library Preparation
[0144] In some embodiments of the method disclosed herein, a DNA library is obtained or prepared from cfDNA obtained from a patient, e.g., a cancer patient. In some embodiments, a DNA library is obtained or prepared from the genome of the patient. In some embodiments, the DNA has been previously sequenced and mutations or variants identified.
[0145] When producing a DNA library from genomic DNA, the genomic DNA can be fragmented, for example, by using a hydrodynamic shear or other mechanical force, or fragmented by chemical or enzymatic digestion, such as restriction digesting. This fragmentation process allows the DNA molecules present in the genome to be sufficiently short for analysis, such as sequencing or digital PCR. cfDNA, however, is generally sufficiently short such that no fragmentation is necessary. cfDNA originates from genomic DNA. A portion of the cfDNA obtained from a plasma sample of a cancerpatient may originate from cancer cells (i.e., circulating tumor DNA or ctDNA), and a portion of the cfDNA may originate from non-cancer cells.
[0146] In some embodiments, the DNA molecules are subjected to targeted depletion prior to library preparation and enrichment or simultaneously with library preparation and enrichment. Targeted depletion may be done to deplete sample of DNA molecules of regions of low utility, including, but not limited to, repetitive regions. Targeted depletion may be accomplished through methods including, but not limited to, targeted endonuclease degradation of regions of low utility, reduced representation sequencing (e.g., RAD Seq), affinity -based hybrid capture, and DNA denaturation and reannealing kinetic-dependent methods. Targeted depletion removes regions of low utility from the sample in order to promote efficiency and accuracy and to reduce off-target amplification and sequencing. These methods are discussed in more detail below.
[0147] In some embodiments, the DNA sample is reduced in size prior to library preparation and enrichment or simultaneously with library preparation and enrichment. Reduction in size of the DNA sample may be done to deplete the sample of an amount of DNA in order to generate a reduced DNA sample that is smaller in comparison to the original DNA sample. This reduction may be done without reference to whether the DNA lost contains information relevant to disease detection. Reduction in sample size may be accomplished through methods including, but not limited to, restriction enzyme-based depletion. In some embodiments, the DNA sample is contacted with a restriction enzyme, and fragments equal to or larger than a threshold are retained while smaller fragments are removed from the sample. These methods are discussed in more detail below.
[0148] In some embodiments, the DNA molecules are not subjected to targeted depletion prior to library preparation and enrichment or simultaneously with library preparation and enrichment.
[0149] In some embodiments of any of the method disclosed herein, the DNA molecules are subjected to additional modification, resulting in the attachment of oligonucleotides to the DNA molecules, to facilitate library preparation and enrichment. The oligonucleotides can comprise an adapter sequence or a molecular barcode or both. In some embodiments, the adapter sequence is common to all oligonucleotides in a plurality of oligonucleotides that are used to form the DNA library. In someembodiments, the molecular barcodes are unique or have low redundancy. By way of example, the oligonucleotide can be attached to the DNA molecules by ligation. Direct attachment of the oligonucleotides to the DNA molecules in the DNA library can be used, for example, when enrichment occurs in a downstream process. In some embodiments, a DNA library is prepared by direct attachment of an oligonucleotide comprising a molecular barcode and an adapter sequence, followed by enrichment (for example, by hybridization) of DNA molecules comprising a region of interest or a portion of a region of interest.
[0150] In some embodiments of any of the methods disclosed herein, library preparation and enrichment occur simultaneously. For example, in some embodiments, DNA molecules comprising a targeted detection region of interest or a portion thereof are preferentially amplified. This can be done, for example, by combining the depleted cfDNA (or genomic DNA) with oligonucleotides comprising a target-specific sequence, an adapter sequence, and a molecular barcode, and amplifying the DNA molecules. As before, in some embodiments, the adapter sequence is common to all oligonucleotides in a plurality of oligonucleotides, and the molecular barcode is unique or of low redundancy. The detection target-specific sequence is unique to the targeted detection region of interest or portion thereof. Thus, PCR amplification selectively amplifies the DNA molecules comprising the targeted detection region of interest or portion thereof.
[0151] When methods include the use of tags or molecular barcodes, the tag or molecular barcode may also be ligated to the fragments or included within the oligonucleotide. The independent attachment of the tag or molecular barcode, as opposed to incorporating the tag or molecular barcode, may vary with the enrichment method. For example, when using hybrid capture-based target enrichment, the adapter can include the molecular barcode, while when using PCR targeted enrichment, detection target-specific primer pairs and overhangs are used that will incorporate the sequencing adapters and sample-specific molecular barcodes, and when using on-sequencer enrichment, the adapter may be separately ligated from the tag or molecular barcode. b. Panel of Mutations / Markers
[0152] In some embodiments, sequencing of the nucleic acids from the sample is performed using Whole Genome Sequencing (WGS). In some embodiments, targetedsequencing is performed and may be either DNA or RNA sequencing. In some embodiments, the targeted sequencing may be to a subset of the whole genome. In some embodiments, targeted depletion is performed prior to sequencing. In some embodiments, targeted sequencing and targeted depletion are performed. In other embodiments, targeted Whole Exome Sequencing (WES) of the DNA from the sample is performed. The DNA is sequenced using a Next Generation Sequencing (NGS) platform, which is massively parallel sequencing. NGS technologies provide high throughput sequencing information and provide digital quantitative information, in that each sequence read that aligns to the sequence of interest is countable. In certain embodiments, clonally amplified DNA templates or single DNA molecules are sequenced in a massively parallel fashion within a flow cell. In addition to high-throughput sequence information, NGS provides quantitative information, in that each sequence read is countable and represents an individual clonal DNA template or single DNA molecule. The sequencing technologies of NGS include pyrosequencing, sequencing-by-synthesis with reversible dye terminators, and sequencing by oligonucleotide probe ligation and ion semiconductor sequencing. DNA from individual samples can be sequenced individually (i.e., singleplex sequencing) or DNA from multiple samples can be pooled and sequenced as indexed genomic molecules (i.e., multiplex sequencing) on a single sequencing run, to generate up to several hundred million reads of DNA sequences. Commercially available platforms include, e.g., platforms for sequencing-by-synthesis, ion semiconductor sequencing, pyrosequencing, reversible dye terminator sequencing, sequencing by ligation, singlemolecule sequencing, sequencing by hybridization, and nanopore sequencing. Platforms for sequencing by synthesis are available from, e.g., Illumina, 454 Life Sciences, Helicos Biosciences, and Qiagen. Illumina platforms can include, e.g., Illumina’s Solexa platform and Illumina’s Genome Analyzer. 454 Life Sciences platforms include, e.g., the GS Flex and GS Junior and are describe in U.S. Patent No. 7,323,305. Platforms from Helicos Biosciences include the True Single Molecule Sequencing platform. Ion Torrent, an alternative NGS system, is available from ThermoScientific and is a semiconductor-based technology that detects hydrogen ions that are released during polymerization of nucleic acids. Any detection method that allows for the detection of segregating markers may be used with the assay provided for herein.
[0153] In some embodiments, Whole Genome Sequencing (WGS) of tumor and normal DNA is performed.
[0154] In other embodiments, Whole Exome Sequencing (WES) of tumor and normal DNA is performed. WES comprises selecting DNA sequences that encode proteins and sequencing that DNA using any high throughput DNA sequencing technology. Methods that can be used to target exome DNA include the use of polymerase chain reaction (PCR), molecular inversion probes (MIP), hybrid capture, and in-solution capture. The utility of targeted genome approaches is well established, and commercially available methods of WES include the Roche NimbleGen Capture Array (Roche NimbleGen Inc., Madison, WI), Agilent SureSelect (Agilent Technologies, Santa Clara, CA), and RainDance Technologies emulsion PCR (RainDance Technologies, Lexington, MA), IDT xGEN® Exome Research Panel, and others.
[0155] Sequence reads may comprise about 20 bp, about 25 bp, about 30 bp, about 35 bp, about 40 bp, about 45 bp, about 50 bp, about 55 bp, about 60 bp, about 65 bp, about 70 bp, about 75 bp, about 80 bp, about 85 bp, about 90 bp, about 95 bp, about 100 bp, about 110 bp, about 120 bp, about 130 bp, about 140 bp, about 150 bp, about 200 bp, about 250 bp, about 300 bp, about 350 bp, about 400 bp, about 450 bp, about 500 bp, or more than 500 bp.
[0156] In some embodiments of the methods described herein, the somatic mutations identified will be analyzed and filtered to generate a subset panel of markers. For example, the subset panel of markers may comprise one or more types of somatic mutations, including but not limited to single-nucleotide variants (SNVs), multinucleotide variants (MNVs), insertions and deletions (e.g., indel variants), and genomic rearrangements. In some embodiments, the subset panel will only include somatic mutations that comprise multiple changes compared to the normal sample, i.e., the subset panel will not include any SNVs. In some embodiments, the subset panel of somatic mutations can include greater than 50, up to 100, up to 200, up to 300, up to 400, up to 500, up to 600, up to 700, up to 800, up to 900, up to 1,000, up to 1,500, up to 2,000, up to 2,500, up to 3,000, up to 4,000, up to 5,000, up to 6,000, up to 7,000, up to 8,000, up to 9,000, up to 10,000, up to 11,000, up to 12,000, up to 13,000, up to 14,000, up to 15,000 or more than 15,000 mutations, which may comprise MNVs, small indels, genomic rearrangements, or combinations thereof. In other embodiments, the subset panel includes between 50 and 15,000 mutations, between 100 and 15,000 mutations, between 500 and13,000 mutations, between 1,000 and 10,000 mutations, between 2,000 and 8,000 mutations, or between 4,000 and 6,000 mutations. c. Probes
[0157] The somatic variants or subset panel may be represented by a set of oligonucleotide probes (e.g., capture probes) each designed to at least partially hybridize to a target sequence that has been identified to comprise a mutation or variant identified in the tumor sample from the patient or in the parental sequence.
[0158] For the purposes of the disclosed methods, probes that are used to identify fragments of interest (i.e., fragments comprising a somatic mutation or variant) can include, but are not limited to capture probes, primers (for the purposes of PCR-based enrichment), or any other suitable type of nucleic acid probe.
[0159] In some embodiments, the panel comprises capture probes comprising the somatic variants identified in the patient’s tumor. In some embodiments, each capture probe is designed to selectively hybridize to a target sequence. The capture probe can be at least 70%, 75%, 80%, 85%, 90%, 95%, or more than 95% complementary to a target sequence. In some embodiments, the capture probe is 100% complementary to a target sequence. In some embodiments, the capture probes are DNA probes. In other embodiments, the capture probes are RNA.
[0160] The capture probe generally is sufficiently long to encompass the sequence of a somatic mutation or corresponding normal sequence comprised in the genomic sequence targeted by the capture probe. The length and composition of a capture probe can depend on many factors including temperature of the annealing reaction, source and base composition of the oligonucleotide, and the estimated ratio of probe to genomic target sequence. Additionally, the length of the capture probe is dependent on the length of the target sequence it is designed to capture. The method provided utilizes cfDNA including circulating tumor DNA (ctDNA) as the source of the target sequences that are to be captured. Accordingly, as cfDNA is highly fragmented to an average of about 170 bp, the capture robe can be, for example, between 100 and 300 bp, between 150 and 250 bp, or between 175 and 200 bp. Probes are typically longer than 120 bases. In some embodiments, if the allele is one or a few bases, then the capture probes may be less than about 110 bases, less than about 100 bases, less than about 90 bases, less than about 80bases, less than about 70 bases, less than about 60 bases, less than about 50 bases, less than about 40 bases, less than about 30 bases, and less than about 25 bases, and this is sufficient to ensure equal enrichment from all alleles.
[0161] When the mixture of DNA that is to be enriched using the hybrid capture technology is a mixture comprising cfDNA isolated from blood, the average length of DNA is quite short, typically less than 200 bases. The use of shorter probes results in a greater chance that the hybrid capture probes will capture desired DNA fragments. Larger variations may require longer probes. In some embodiments, the variations of interest are more than one base in length. In some embodiments, the variations of interest are one base in length, e.g., single nucleotide polymorphisms (SNPs). In some embodiments, targeted regions in the genome can be preferentially enriched using hybrid capture probes wherein the hybrid capture probes are shorter than 90 bases, and can be less than 80 bases, less than 70 bases, less than 60 bases, less than 50 bases, less than 40 bases, less than 30 bases, or less than 25 bases. In some embodiments, to increase the chance that the desired allele is sequenced, the length of the probe that is designed to hybridize to the regions flanking the polymorphic allele location can be decreased from above 90 bases to about 80 bases, or to about 70 bases, or to about 60 bases, or to about 50 bases, or to about 40 bases, or to about 30 bases, or to about 25 bases.
[0162] Hybrid capture probes can be designed such that the region of the capture probe with DNA that is complementary to the DNA found in regions flanking the polymorphic allele is not immediately adjacent to the polymorphic site. Instead, the capture probe can be designed such that the region of the capture probe that is designed to hybridize to the DNA flanking the polymorphic site of the target is separated from the portion of the capture probe that will be in van der Waals contact with the polymorphic site by a small distance that is equivalent in length to one or a small number of bases. In some embodiments, the hybrid capture probe is designed to hybridize to a region that is flanking the polymorphic allele but does not cross it; this may be termed a “flanking capture probe.” The length of the flanking capture probe may be less than about 120 bases, less than about 110 bases, less than about 100 bases, less than about 90 bases, less than about 80 bases, less than about 70 bases, less than about 60 bases, less than about 50 bases, less than about 40 bases, less than about 30 bases, or less than about 25 bases. Theregion of the genome that is targeted by the flanking capture probe may be separated by the polymorphic locus by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11-20, or more than 20 base pairs.
[0163] For small insertions or deletions, one or more probes that overlap the mutation may be sufficient to capture and sequence fragments comprising the mutation. Hybridization may be less efficient between the probe-limiting capture efficiency, typically designed to the reference genome sequence. To ensure capture of fragments comprising the mutation, one could design two probes, one matching the normal allele and one matching the mutant allele. A longer probe may enhance hybridization. Multiple overlapping probes may enhance capture. Finally, placing a probe immediately adjacent to, but not overlapping, the mutation may permit relatively similar capture efficiency of the normal and mutant alleles.
[0164] For Short Tandem Repeats (STRs), a probe overlapping these highly variable sites is unlikely to capture the fragment well. To enhance capture, a probe could be placed adjacent to, but not overlapping, the variable site. The fragment could then be sequenced as normal to reveal the length and composition of the STR.
[0165] For large deletions, a series of overlapping probes, a common approach currently used in exon capture systems may not work. However, with this approach, it may be difficult to determine whether or not an individual is heterozygous. According to the method provided, custom probes are designed to ensure capture of the unique set of somatic mutations identified in the patient’s tumor.
[0166] Capture probes can be modified to comprise purification moieties that serve to isolate the capture duplex from the unhybridized, untargeted cfDNA sequences by binding to a purification moiety binding partner. Suitable binding pairs for use in the invention include, but are not limited to, antigens / antibodies (for example, digoxigenin / antidigoxigenin, dinitrophenyl (DNP) / anti-DNP, dansyl-X / antidansyl, fluorescein / anti-fluorescein, lucifer yellow / anti-lucifer yellow, and rhodamine / anti- rhodamine); biotin / avidin (or biotin / streptavidin); calmodulin binding protein (CBP) / calmodulin; hormone / hormone receptor; lectin / carbohydrate; peptide / cell membrane receptor; protein A / antibody; hapten / anti-hapten; enzyme / co-factor; and enzyme / substrate. Other suitable binding pairs include polypeptides, such as the FLAG- peptide (Hopp et al., BioTechnology, 6: 1204-1210 (1988)); the KT3 epitope peptide(Martin et aL, Science, 255: 192-194 (1992)); tubulin epitope peptide (Skinner et aL, J. Biol. Chem., 266: 15163-15166 (1991)); and the T7 gene 10 protein peptide tag (Lutz- Freyermuth et al., Proc. Natl. Acad. Sci. USA, 87:6393-6397 (1990)) and the antibodies each thereto.
[0167] Further non-limiting examples of binding partners include agonists and antagonists for cell membrane receptors, toxins and venoms, viral epitopes, hormones such as steroids, hormone receptors, peptides, enzymes and other catalytic polypeptides, enzyme substrates, cofactors, drugs including small organic molecule drugs, opiates, opiate receptors, lectins, sugars, saccharides including polysaccharides, proteins, and antibodies including monoclonal antibodies and synthetic antibody fragments, cells, cell membranes and moieties therein including cell membrane receptors and organelles. In some embodiments, the first binding partner is a reactive moiety, and the second binding partner is a reactive surface that reacts with the reactive moiety, such as described herein with respect to other aspects of the invention. In some embodiments, the oligonucleotide primers are attached to the solid surface prior to initiating the extension reaction. Methods for the addition of binding partners to capture oligonucleotide probes are known in the art and include addition during (such as by using a modified nucleotide comprising the binding partner) or after synthesis. Additionally, the capture probes can be tethered to a solid surface, e.g., a magnetic bead, which facilitates the isolation of captured sequences.IV. Phase II - Detection and Monitoring Tumors by Analyzing cfDNA a. Targeted Enrichment of a Region of Interest
[0168] The disclosed methods generally comprise enriching a detection target sequence in a region of interest. Examples of enrichment techniques include, but are not limited to, hybrid capture, selective circularization (also referred to as molecular inversion probes (MIPs)), and PCR amplification of targeted detection regions of interest. Hybrid capture methods are based on the selective hybridization of the detection target genomic regions to user-designed oligonucleotides. The hybridization can be to oligonucleotides immobilized on high- or low-density microarrays (on-array capture), or solution-phase hybridization to oligonucleotides modified with a ligand (e.g., biotin), which can subsequently be immobilized to a solid surface, such as a bead (in solution capture). Molecular inversion probe (MIP) -based methods rely on construction ofnumerous single-stranded linear oligonucleotide probes, consisting of a common linker flanked by detection target-specific sequences. Upon annealing to a detection target sequence, the probe gap region is filled via polymerization and ligation, resulting in a circularized probe. The circularized probes are then released and amplified using primers directed to the common linker region. PCR-based methods employ highly parallel PCR amplification, where each detection target sequence in the sample has a corresponding pair of unique, sequence-specific primers. In some embodiments, enrichment of a detection target sequence occurs at the time of sequencing.
[0169] Additionally or alternatively, enriching the detection target sequence of interest may be done via targeted depletion and / or reduction of cfDNA sample size. These methods are discussed in more detail below. These methods of enrichment do not rely on oligonucleotide probes to enrich the cfDNA sample for ctDNA.
[0170] In this second phase of the method, samples that are used for determining the tumor fraction of the patient include samples that contain nucleic acids that are cell-free. Cell-free nucleic acids, including cfDNA, can be obtained by various methods from biological samples, including, but not limited to, plasma, serum, and urine. Other biological fluid samples include, but are not limited to, blood, sweat, tears, sputum, ear flow, lymph, saliva, cerebrospinal fluid, ravages, bone marrow suspension, vaginal flow, transcervical lavage, brain fluid, ascites, milk, secretions of the respiratory, intestinal, and genitourinary tracts, amniotic fluids, and leukophoresis samples. In some embodiments, the sample is a sample that is easily obtainable by non-invasive procedures, e.g., blood, plasma, serum, sweat, tears, sputum, urine, ear flow, saliva, or feces. In certain embodiments, the sample is a peripheral blood sample, or the plasma and / or serum fractions of a peripheral blood sample. In other embodiments, the biological sample is a swab or smear, a biopsy specimen, or a cell culture. In another embodiment, the sample is a mixture of two or more biological samples, e.g., a biological sample can comprise two or more of a biological fluid sample, a tissue sample, and a cell culture sample.
[0171] In some embodiments, the cfDNA sample is subjected to targeted depletion prior to sequencing. Targeted depletion may be done to deplete sample of DNA molecules of regions of low utility, including, but not limited to, repetitive regions. Targeted depletion may be accomplished through methods including, but not limited to, targeted endonuclease degradation of regions of low utility, reduced representationsequencing (e.g., RAD Seq), affinity-based hybrid capture, and DNA denaturation and reannealing kinetic-dependent methods. Targeted depletion removes regions of low utility from the sample in order to promote efficiency and accuracy and to reduce off-target amplification and sequencing. These methods are discussed in more detail below.
[0172] In some embodiments, the cfDNA sample is reduced in size prior sequencing. Reduction in size of the cfDNA sample may be done to deplete the sample of an amount of DNA in order to generate a reduced cfDNA sample that is smaller in comparison to the original DNA sample and enriched for ctDNA. This reduction may be done without reference to whether the DNA lost contains information relevant to disease detection. Reduction in sample size may be accomplished through methods including, but not limited to, restriction enzyme-based depletion. In some embodiments, the cfDNA sample is contacted with a restriction enzyme, and fragments equal to or larger than a threshold are retained while smaller fragments are removed from the sample. These methods are discussed in more detail below.
[0173] In various embodiments, the cfDNA present in the sample can be enriched specifically or non-specifically prior to use (e.g., prior to capture and sequencing). Nonspecific enrichment of sample DNA refers to the amplification of the genomic DNA fragments of the sample that can be used to increase the level of the sample DNA prior to capture and sequencing. Non-specific enrichment can be the selective enrichment of exomes. Suitable methods for whole genome amplification include, but are not limited to, degenerate oligonucleotide-primed PCR (DOP), primer extension PCR technique (PEP), and multiple displacement amplification (MDA). In some embodiments, the sample is unenriched for cfDNA.
[0174] As is described elsewhere herein, cfDNA is present as fragments averaging about 170 bp. Accordingly, further fragmentation of cfDNA is not needed. In some embodiments, sufficient cfDNA is obtained from a 10 mL blood sample to confidently determine the presence or absence of cancer in a patient. The blood samples used in the method provided can be of about 5 mL, about 10 mL, about 15 mL, about 20 mL, about 25 mL, or more than 25 mL. Typically, 20 mL of blood plasma contains between 5,000 and 10,000 genome equivalents and provides more than sufficient cfDNA for determining tumor fraction according to the method provided. In some embodiments, sufficient cfDNA is obtained from 10 mL to 20 mL of blood to determine tumor fraction.
[0175] To separate cfDNA from cells in a sample, various methods including, but not limited to fractionation, centrifugation (e.g., density gradient centrifugation), DNA- specific precipitation, or high-throughput cell sorting and / or other separation methods can be used. Commercially available kits for manual and automated separation of cfDNA are available (Roche Diagnostics, Indianapolis, IN; Qiagen, Germantown, MD).
[0176] cfDNA can be end-repaired and, optionally, dA-tailed, and double-stranded adapters comprising sequences complementary to amplification and sequencing primers are ligated to the ends of the cfDNA molecules to enable NGS, e.g., using an Illumina platform. Additionally, each of the double-stranded adapters further comprises a nonrandom barcode sequence, which serves to differentiate individual cfDNA molecules. In some embodiments, the barcode sequences are random sequences. In other embodiments, the barcode sequences are non-random barcode sequence. Non-random barcode sequences provide a significant advantage over random barcode sequences because nonrandom barcode sequences enable unambiguous identification of the sequence reads described below. The non-random barcode sequences are designed specifically to be base-balanced, both within and across all barcodes. Additionally, in some embodiments, the non-random barcodes can comprise a T nucleotide at the 3’ end, which is complementary to the A nucleotide of dA-tailed cfDNA molecules. In embodiments utilizing a T nucleotide overhang at the 3 ’ end of the barcode, barcodes of three different lengths can be designed to avoid a single base flashing across the entire flow-cell of the sequencer. Non-random barcode sequences can be present in adapters as sequences of 13, 14, and 15 bp; 10, 11, and 12 bp; 11, 12, and 13 bp; 14, 15, and 16 bp; 15, 16, and 17 bp; and the like. In some embodiments, the shortest barcode sequence can be 8 bp, and the longest barcode sequence can be 100 bp.
[0177] Each sequence of the subpanel that is present in the cfDNA sample is targeted by one or more capture probes described elsewhere herein and is isolated for further analysis. b. Sequencing and Analysis
[0178] The disclosed methods generally comprise sequencing one or more samples. Sequencing methods include, but are not limited to, Maxam Gilbert sequencing-based techniques, chain-termination-based techniques, shotgun sequencing, bridge PCRsequencing, single-molecule real-time sequencing, ion semiconductor sequencing (Ion Torrent sequencing), nanopore sequencing, pyrosequencing, sequencing by synthesis, sequencing by ligation (SOLiD sequencing), sequencing by electron microscopy, dideoxy sequencing reactions (Sanger method), massively parallel sequencing, polony sequencing, duplex sequencing, and DNA nanoball sequencing. In some embodiments, sequencing involves hybridizing a primer to the template to form a template / primer duplex, contacting the duplex with a polymerase enzyme in the presence of detectably-labeled nucleotides under conditions that permit the polymerase to add nucleotides to the primer in a template-dependent manner, detecting a signal from the incorporated labeled nucleotide, and sequentially repeating the contacting and detecting steps at least once, wherein sequential detection of incorporated labeled nucleotides determines the sequence of the nucleic acid. In some embodiments, the sequencing comprises obtaining paired end reads. The accuracy or average accuracy of the sequence information may be greater than 80%, 90%, 95%, 99%, or 99.98%. In some embodiments, the sequence information obtained in more than 50 bp, 100 bp, or 200 bp. The sequence information may be obtained in less than 1 month, 2 weeks, 1 week, 1 day, 3 hours, 1 hour, 30 minutes, 10 minutes, or 5 minutes. The sequence accuracy or average accuracy may be greater than 95% or 99%. The sequencing coverage may be greater than 20-fold or less than 500-fold. Examples of detectable labels include radiolabels, fluorescent labels, enzymatic labels, etc. In some embodiments, the detectable label may be an optically detectable label, such as a fluorescent label. Examples in fluorescent labels include cyanine, rhodamine, fluorescein, coumarin, BODIPY, alexa, or conjugated multi-dyes. In some embodiments, the nucleotide is flagged if one or more of its sequence segments are substantially similar to one or more sequence segments of another nucleotide within the same partition.
[0179] Some methods of sequencing may require or involve a prior target enrichment step. For example, use of on-sequencer enrichment, such as with a nanopore sequencer, allows for the simultaneous enrichment and sequencing of the sequence library by realtime rejection of molecules that are not from the region of interest. Alternatively, sequences can be selectively and preferentially sequenced from the region of interest.
[0180] Captured sequences can be analyzed using the sequencing-by-synthesis technology of Illumina, which uses fluorescent reversible terminator deoxyribonucleotides. The reads generated by the sequencing process are aligned to areference sequence and associated with a sequence of the somatic sequence panel specific for the patient. Mapping of the sequence reads can be achieved by comparing the sequence of the reads with the sequence of the reference genome to determine the specific genetic information, and optionally, the chromosomal origin of the sequenced nucleic acid (e.g., cfDNA) molecule. A number of computer algorithms are available for aligning sequences, including without limitation BLAST (Altschul etal., 199), BLITZ (MPsrch) (Sturrock & Collins, 1993), FASTA (Person & Lipman, 1988), BOWTIE (Langmead et al., Genome Biology 10:R25.1-R25-10 (2009)), or ELAND (Illumina, Inc., San Diego, CA, USA). In one embodiment, the sequencing data is processed by bioinformatic alignment analysis for the Illumina Genome Analyzer, which uses the Efficient Large- Scale Alignment of Nucleotide Databases (ELAND) software. Additional software includes SAMtools (SAMtools, Bioinformatics, 2009, 25(16):2078-2079), and the Burroughs-Wheeler block sorting compression procedure, which involves block sorting or pre-processing to make compression more efficient.
[0181] The barcoded cfDNA fragments isolated from the patient’s fluid sample, e.g., blood sample, can be amplified, e.g., by PCR, and captured using the hybrid probes. Capturing of the barcoded fragments comprises obtaining single strands of barcoded cfDNA and hybridizing the barcoded cfDNA with different hybrid probes. Each of the different hybrid probes hybridizes to a single-stranded barcoded cfDNA target sequence to form a target-hybrid probe duplex. The duplex is isolated from unhybridized cfDNA by binding the purification binding moiety comprised in the hybrid probe to the corresponding purification moiety binding partner. As described elsewhere herein, the corresponding purification moiety binding partner can be immobilized on a solid surface, e.g., a magnetic bead, which facilitates the separation of the capture duplex from unhybridized cfDNA molecules in solution. The barcoded cfDNA of the duplex is released and is subjected to sequencing using an NGS instrument.
[0182] The error rate in sequencing using NGS methods is of approximately 1 in 500 bases, which results in many sequencing errors. The high error rate becomes problematic, especially when attempting to identify somatic mutations in mixtures of DNA sequences comprising only a small fraction of mutated species or sequences comprising single nucleotide variants. Additionally, NGS methods typically utilize single-stranded DNA as the primary source of sequencing material. Any error included during the amplificationstep of the DNA molecule prior to sequencing is perpetuated and becomes indistinguishable as an extraneous technology-dependent mistake. Chemical errors occur at a frequency of approximately 1 in 1,000 bases. This combination of sequencing and chemical errors obscures the limit of detection (LOD).
[0183] Accordingly, in some embodiments, double-stranded sequencing of the cfDNA is performed. As described elsewhere herein cfDNA can be end-repaired, and optionally dA-tailed, and double-stranded adaptors comprising sequences complementary to amplification and sequencing primers are ligated to the ends of the cfDNA molecules to enable NGS sequencing, e.g., using an Illumina platform.
[0184] The tumor fraction can then be calculated as the proportion of different cfDNA sequences, each comprising at least one somatic mutation, i.e., ctDNA sequences, relative to the total number of different cfDNA, i.e., ctDNA and corresponding normal sequences. Unlike the single-stranded approach, the current method corrects for random sequencing errors.
[0185] Additionally, in some embodiments, the methods described herein avoid such errors by analyzing target sequences that comprise somatic mutations having multiple changes relative to a reference sequence. In some embodiments, the method includes eliminating false positives by including in the patient-specific panel only variants that include multiple changes relative to a reference sequence. It would be unlikely that the variant would be observed only due to an assay error, sequencing error, or single background mutations in the genome. This method may eliminate certain true disease- related mutations from inclusion in the panel, preventing the effective detection of tumor cells. In some embodiments, the methods further include obtaining a second non-tumor sample from the subject; extracting genomic DNA from the second non-tumor sample; enriching the genomic DNA from the second non-tumor sample for sequences corresponding to a preliminary patient-specific panel, thereby obtaining an enriched DNA fraction from the second non-tumor sample; sequencing the enriched DNA fraction from the second non-tumor sample to detect the presence or absence of any of the plurality of the preliminary tumor-specific somatic mutations; and removing from the preliminary patient-specific panel any of the plurality of preliminary tumor-specific somatic mutations that are present in the enriched DNA fraction from the second non-tumor sample, thereby obtaining a patient-specific panel. This method provides additional benefits, includingacting as comprehensive quality control for the bespoke enrichment library before it is applied to a ctDNA sample and enabling identification of alterations that are not true somatic changes in the tumor but may instead be germ-line changes, clonal hematopoiesis of indeterminate potential (CHIP) mutations, sequencing errors, or other artifacts. Removing such artifacts from the patient-specific panel would lead to a reduction in false positive results. c. Molecular Barcodes
[0186] In some embodiments, an identified sequence, i.e., a molecular barcode, may be used to identify unique DNA molecules or target sequences in a DNA library. Molecular barcodes aid in reconstruction of a contiguous DNA sequence or assist in copy number variation determinations. Exemplary markers include nucleic acid binding proteins, optical labels, nucleotide analogs, nucleic acid sequences, and others known in the art.
[0187] In some embodiments, the molecular barcode is a nanostructure barcode. In some embodiments, the molecular barcode comprises a nucleic acid sequence that, when joined to a detection target polynucleotide, serves as an identifier of the sample or sequence from which the detection target polynucleotide was derived. In some embodiments, molecular barcodes are at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more nucleotides in length. In some embodiments, molecular barcodes are shorter than 10, 9, 8, 7, 6, 5, or 4 nucleotides in length. In some embodiments, each molecular barcode in a plurality of molecular barcodes differ from every other molecular barcode in the plurality in at least three nucleotide positions, such as at least 3, 4, 5, 6, 7, 8, 9, 10, or more positions. In some embodiments, molecular barcodes associated with some polynucleotides are of different length than molecular barcodes associated with other polynucleotides. In general, molecular barcodes are of sufficient length and comprise sequences that are sufficiently different to allow the identification of samples based on molecular barcodes with which they are associated. In some embodiments, both the forward and reverse adapter comprise at least one of a plurality of molecular barcode sequences. In some embodiments, each reverse adapter comprises at least one of a plurality of molecular barcode sequences, wherein each molecular barcode sequence of the plurality of molecular barcode sequences differs from every other molecular barcode sequence in the plurality of molecular barcode sequences.
[0188] In some embodiments, every molecular barcode in a set is unique, that is, any two molecular barcodes chosen out of a given set will differ in at least one nucleotide position. Furthermore, it is contemplated that molecular barcodes have certain biochemical properties that are selected based on how the set will be used. For instance, certain sets of molecular barcodes that are used in an RT-PCR reaction should not have complementary sequences to any sequence in the genome of a certain organism or set of organisms. A requirement for non-complementarity helps to ensure that the use of a particular molecular barcode sequence will not result in mis-priming during molecular biological manipulations requiring primers, such as reverse transcription or PCR. Certain sets satisfy other biochemical properties imposed by the requirements associated with the processing of the sequence molecules into which the barcodes are incorporated.
[0189] Examples of sequencing technologies for sequencing molecular barcodes, as well as any generated nucleotide-based sequence, include, but are not limited to, Maxam- Gilbert sequencing-based techniques, chain-termination-based techniques, shotgun sequencing, bridge PCR sequencing, single-molecule real-time sequencing, ion semiconductor sequencing (Ion Torrent sequencing), nanopore sequencing, pyrosequencing, sequencing by synthesis, sequencing by ligation (SOLiD sequencing), sequencing by electron microscopy, dideoxy sequencing reactions (Sanger sequencing), massively parallel sequencing, polony sequencing, and DNA nanoball sequencing.
[0190] In some embodiments, molecular barcodes are used to improve the power of copy-number calling algorithms by reducing non-independence from PCR duplication. In another embodiment, molecular barcodes can be used to improve test specificity be reducing sequence error generated during amplification.V. Methods of Generating Patient-Specific Panels using Targeted Depletion or Reduction in Sample Size
[0191] Detecting somatic mutations that are present in a low frequency through Next Generation Sequencing (NGS) presents a promising method of detecting disease presence, including cancer reoccurrence, early and through less-invasive procedures. Deep sequencing is an NGS approach that is highly sensitive and can detect rare mutations or polymorphisms from a sample. Deep sequencing relies upon a high depth of coverage, meaning the process will generate a high number of reads of a given nucleotidesequence in a single experiment. By increasing the depth of coverage by increasing the number of reads of the given sequence, it will be possible to detect rare mutations or polymorphisms in a sample containing many copies of the same nucleotide sequence by increasing the likelihood that the mutated sequence will be amplified and detected. Thus, tools utilizing NGS are much more sensitive and able to detect even very rare mutations in a given sample.
[0192] Deep sequencing of whole genomes present issues. Genomes, including the human genome, include regions of low utility in the detection assays, including repetitive elements that make up approximately 50% of the human genome. Whole genome sequencing will involve sequencing of these regions as well, which increases the amount of DNA to be sequenced. These regions may also hybridize to the probes designed for other detection target sequences of interest, resulting in reduced numbers of reads for the sequences to be detected and false identification of regions of low utility as regions of interest. Thus, the need for a high number of reads for deep sequencing may result in increased costs, processing time, and risk of error.
[0193] One method of addressing this issue is through a process termed targeted depletion. This method generally functions by targeting the regions of low utility in the DNA sample and depleting them from the sample thereby reducing the total amount of DNA that can be amplified and sequenced in order to prevent off-target sequencing and reduce the total number of reads necessary to reach the requisite depth to detect the regions of interest in the detection assay.
[0194] Targeted depletion may be accomplished through various strategies. In one embodiment, regions of low utility may be depleted from the DNA sample using endonucleases. The endonucleases may be DNA-guided endonucleases that utilize DNA guide oligonucleotides to target depletion sequences for cleavage. The endonucleases may be RNA-guided endonucleases that utilize RNA guide oligonucleotides to target depletion sequences for cleavage. The DNA or RNA guide oligonucleotides may be complementary to depletion sequences and target portions of the DNA sample having the depletion sequence for cleavage by the endonuclease. The DNA-guided endonucleases include, but are not limited to, Thermus thermophilus Argonaute (TtAgo). When the DNA-guided endonuclease is TtAgo, the DNA guide oligonucleotide can be a DNA oligonucleotide having about 16 to 18 nucleotides that is phosphorylated at the 5’ end ofthe oligonucleotide. The RNA-guided endonucleases include, but are not limited to, CRISPR-associated endonucleases, e.g., Cas9 endonuclease. When the RNA-guided endonuclease is a Cas9 endonuclease, the RNA guide oligonucleotide can be an RNA oligonucleotide having about 15 to 24 nucleotides. The endonucleases may be structure- guided endonucleases (SGNs) that do not rely on a specific target sequence but instead targets a specific target structure for cleavage. The SGNs include, but are not limited to, a FEN-1-Fnl fusion protein as described in Xu et al., Genome Biol. 17: 186 (2016).
[0195] Targeted depletion using the endonucleases may comprise contacting the DNA sample, either before fragmentation or after fragmentation, with the endonucleases and a plurality of guide oligonucleotides. The plurality of guide oligonucleotides may comprise one or more unique guide oligonucleotides, wherein each unique guide oligonucleotide is complementary to and hybridizes to a specific depletion sequence. Contacting the DNA sample with the endonuclease and plurality of guide oligonucleotides results in the targeted cleavage of the depletion sequences. Cleavage of the depletion sequences results in the removal of the depletion sequences from the sample and prevents subsequent amplification and sequencing of the depletion sequences in subsequent steps of the detection assay. The remainder of the DNA sample may be further processed and sequenced as discussed above.
[0196] In another embodiment, regions of low utility may be depleted from the DNA sample using restriction enzymes in what is termed “reduced-representation sequencing.” Reduced-representation sequencing includes, but is not limited to, RAD-Seq, 2bRAD- Seq, ddRAD-Seq, and GBS-Seq. This method may involve selection of restriction enzymes based on their ability to digest the DNA sample in regions of low utility; therefore, in some embodiments, restriction enzymes that preferentially digest the DNA sample at regions of low utility may be selected for use in this targeted depletion method. In some embodiments, reduced-representation sequencing may be restriction enzyme- mediated reduced representation sequencing. The DNA sample may be contacted with a restriction enzyme, which cleaves the DNA at various recognition sites throughout the genome. Optionally, the cleavage products may next be ligated with a first adapter sequence. These ligated cleavage products may be randomly sheared before being ligated with a second adapter sequence. Only ligated cleavage products that have both adapter sequences will be amplified, thus ensuring that sequences in proximity to the restrictionsites will be amplified for further processing and sequencing, while sequences distant from the restriction sites and generated from shearing only will not be amplified and processed. Additionally or alternatively, the cleavage products resulting from contacting the DNA sample with the restriction enzyme may be subjected to a size selection step in order to eliminate fragments that fall outside the desired fragment length range. The cleavage products that have a length within the desired fragment length range may then be ligated with adapter sequences in order to amplify the cleavage products of interest. One or more restriction enzyme may be utilized to digest the DNA sample, and one or more first adapter sequences may be used to ligate to the cleavage sites of the DNA sample based on which restriction enzyme generated the specific cleavage. The remainder of the DNA sample may be further processed and sequenced as discussed above.
[0197] In another embodiment, regions of low utility may be depleted from the DNA sample using affinity-based hybrid capture methods. In some embodiments, the DNA sample may be randomly sheared to generate short DNA segments that are between 200 and 600 nucleotides in length (e.g., 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, or 60 bp in length, or any lengths in between). In some embodiments, the DNA sample may be randomly sheared to generate short DNA segments that are 200-500, 200-400, 200-300, 250-500, 250-450, 250-400, 250-250, 300-500, 300-450, or 300-400 nucleotides in length. These short DNA segments may be contacted with a plurality of capture oligonucleotides or capture probes following denaturation of the double-stranded structure of the short DNA segments to generate single-stranded short DNA segments. The plurality of capture oligonucleotides or capture probes may be one or more oligonucleotides that contain sequences that are substantially or completely complementary to a sequence in the shorter DNA segments comprising the regions of low utility. The plurality of capture oligonucleotides or capture probes may include an affinity tag, including, but not limited to, biotin and dA tails.When the single-stranded short DNA segments are contacted with the plurality of capture oligonucleotides or capture probes, the single-stranded short DNA segments with regions of low utility that comprises sequences complementary to the sequence of one or more of the plurality of capture oligonucleotides or capture probes hybridize to the one or more of the plurality of capture oligonucleotides or capture probes to form a double-stranded DNA-capture oligonucleotide structure or double-stranded DNA-capture probe structure.These double-stranded DNA-capture oligonucleotide structures or double-stranded DNA- capture probe structures may be pulled down from the DNA sample by contacting the sample with reagents that bind the affinity tag of the capture oligonucleotides or capture probes. These reagents include, but are not limited to, magnetic beads comprising streptavidin when the affinity tag is biotin, magnetic beads comprising avidin when the affinity tag is biotin, glass beads comprising streptavidin when the affinity tag is biotin, glass beads comprising avidin when the affinity tag is biotin, membranes comprising streptavidin when the affinity tag is biotin, membranes comprising avidin when the affinity tag is biotin, magnetic beads comprising a polyT oligonucleotide when the affinity tag is dA tails, and glass beads comprising a polyT oligonucleotide when the affinity tag is dA tails. Pull-down of the double-stranded DNA-capture oligonucleotides or double-stranded DNA-capture probes depletes the DNA sample of the short DNA segments that comprise the regions of low utility. The remainder of the DNA sample may be further processed and sequenced as discussed above.
[0198] In another embodiment, regions of low utility may be depleted from the DNA sample using methods dependent upon DNA denaturation and reannealing kinetics. In some embodiments, the methods dependent upon DNA denaturation and reannealing kinetics include, but are not limited to, Cot analysis of the DNA sample. In some embodiments, the DNA sample may be sheared to generate short DNA segments that are between 200 and 600 nucleotides in length (e.g., 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, or 60 bp in length, or any lengths in between). These short DNA segments may be divided into two or more samples. The samples of the short DNA segments may be denatured to produce singlestranded short DNA segments. The samples of the single-stranded short DNA segments may then be allowed to reassociate to different Cot values, wherein the Cot value is the product of nucleotide concentration in moles per liter and reassociation time in seconds and, if applicable, a factor based upon cation concentration of the buffer. For each sample, renatured DNA may be separated from single-stranded DNA. The renatured DNA may be separated from the single-stranded DNA using methods including, but not limited to, hydroxyapatite (HAP) chromatography. Following separation of the renatured DNA, the percentage of the sample that has not reassociated (%ssDNA) may be determined. The logarithm of the sample’s Cot value is plotted against the sample’s%ssDNA to generate a Cot point, and a graph of Cot points ranging from little or no reassociation until reassociation approaches completion is used to generate a Cot curve. The Cot curve may include three regions, including a highly repetitive (HR) component, moderately repetitive (MR) component, and a single / low-copy (SL) component, which are, respectively, characterized by fast, intermediate, and slow reassociation. It is believed that the slowest reassociating component of a Cot curve generally represents single-copy DNA sequences. Assuming that the SL component has a repetition frequency of 1, the average repetition frequency of the DNA in the HR and the MR components may be estimated by dividing their reassociation rates by the reassociation value of the SL component. Thus, the various components may be studied to identify regions of low utility, such as repetitive elements, that may be desirable targets for depletion. In some embodiments, the HR component samples may include the renatured DNA isolated at a low Cot value that has been separated from the remaining sample. In some embodiments, the MR component samples may include the renatured DNA isolated at an intermediate Cot value that has been separated from the remaining sample. In some embodiments, the SL component samples may include the renatured DNA isolated at a high Cot value and any additional single-stranded DNA that has not renatured. The Cot values for the three components may be determined by statistical analysis of the generated Cot curve for the total DNA sample. In some embodiments, only the SL component samples may be used for further processing and sequencing as discussed above. In some embodiments, the SL component samples and the MR component samples may be used for further processing and sequencing as discussed above. In some embodiments the SL component samples, the MR component samples, and a portion of the HR component samples that are predicted to have a lower repetition frequency may be used for further processing and sequencing as discussed above. In some embodiments, a random portion of one or more of each of the component samples may be use for further processing and sequencing as discussed above. In some embodiments, a random portion of one or more of the component samples may be used for further processing and sequencing as discussed above.
[0199] In another embodiment, the size of a DNA sample may be reduced using restriction enzymes in what is termed “reduced-representation sequencing.” Reduced- representation sequencing includes, but is not limited to, RAD-Seq, 2bRAD-Seq, ddRAD-Seq, and GBS-Seq. This method does not involve selection of restriction enzymes based on their ability to digest the DNA sample in regions of low utility, butinstead allow for the use of any restriction enzymes regardless of where they are predicted to digest the sample. In some embodiments, reduced-representation sequencing may be restriction enzyme-mediated reduced representation sequencing. The DNA sample may be contacted with a restriction enzyme, which cleaves the DNA at various recognition sites throughout the genome. Optionally, the cleavage products may next be ligated with a first adapter sequence. These ligated cleavage products may be randomly sheared before being ligated with a second adapter sequence. Only ligated cleavage products that have both adapter sequences will be amplified, thus ensuring that sequences in proximity to the restriction sites will be amplified for further processing and sequencing, while sequences distant from the restriction sites and generated from shearing only will not be amplified and processed. Additionally or alternatively, the cleavage products resulting from contacting the DNA sample with the restriction enzyme may be subjected to a size selection step in order to eliminate fragments that fall outside the desired fragment length range. The cleavage products that have a length within the desired fragment length range may then be ligated with adapter sequences in order to amplify the cleavage products of interest. One or more restriction enzyme may be utilized to digest the DNA sample, and one or more first adapter sequences may be used to ligate to the cleavage sites of the DNA sample based on which restriction enzyme generated the specific cleavage. The remainder of the DNA sample may be further processed and sequenced as discussed above.VI. Methods of Enriching ctDNA from cfDNA Samples using Targeted Depletion or Reduction in Sample Size
[0200] Similar to the reasons provided the above section, it may be beneficial to apply the targeted depletion or reduction in sample size methods in enriching the cfDNA for ctDNA. By increasing the depth of coverage by increasing the number of reads of the given sequence by reducing the overall amount of DNA being sequenced, it will be possible to detect rare mutations or polymorphisms in a sample containing many copies of the same nucleotide sequence by increasing the likelihood that the mutated sequence will be amplified and detected. Thus, tools utilizing NGS are much more sensitive and able to detect even very rare mutations in a given sample.
[0201] Targeted depletion may be accomplished through various strategies. In one embodiment, regions of low utility may be depleted from the cfDNA sample usingendonucleases. The endonucleases may be DNA-guided endonucleases that utilize DNA guide oligonucleotides to target depletion sequences for cleavage. The endonucleases may be RNA-guided endonucleases that utilize RNA guide oligonucleotides to target depletion sequences for cleavage. The DNA or RNA guide oligonucleotides may be complementary to depletion sequences and target portions of the cfDNA sample having the depletion sequence for cleavage by the endonuclease. The DNA-guided endonucleases include, but are not limited to, Thermus thermophilus Argonaute (TtAgo). When the DNA-guided endonuclease is TtAgo, the DNA guide oligonucleotide can be a DNA oligonucleotide having about 16 to 18 nucleotides that is phosphorylated at the 5’ end of the oligonucleotide. The RNA-guided endonucleases include, but are not limited to, CRISPR-associated endonucleases, e.g., Cas9 endonuclease. When the RNA-guided endonuclease is a Cas9 endonuclease, the RNA guide oligonucleotide can be an RNA oligonucleotide having about 15 to 24 nucleotides. The endonucleases may be structure- guided endonucleases (SGNs) that do not rely on a specific target sequence but instead targets a specific target structure for cleavage. The SGNs include, but are not limited to, a FEN-1-Fnl fusion protein as described in Xu et al., Genome Biol. 17: 186 (2016).
[0202] Targeted depletion using the endonucleases may comprise contacting the cfDNA sample, either before or after optional fragmentation, with the endonucleases and a plurality of guide oligonucleotides. The plurality of guide oligonucleotides may comprise one or more unique guide oligonucleotides, wherein each unique guide oligonucleotide is complementary to and hybridizes to a specific depletion sequence. Contacting the cfDNA sample with the endonuclease and plurality of guide oligonucleotides results in the targeted cleavage of the depletion sequences and enrichment of ctDNA. Cleavage of the depletion sequences results in the removal of the depletion sequences from the sample and prevents subsequent amplification and sequencing of the depletion sequences in subsequent steps of the detection assay. The remainder of the cfDNA sample which is enriched for the ctDNA may be further processed and sequenced as discussed above.
[0203] In another embodiment, regions of low utility may be depleted from the cfDNA sample using restriction enzymes in what is termed “reduced-representation sequencing.” Reduced-representation sequencing includes, but is not limited to, RAD- Seq, 2bRAD-Seq, ddRAD-Seq, and GBS-Seq. This method may involve selection ofrestriction enzymes based on their ability to digest the cfDNA sample in regions of low utility; therefore, in some embodiments, restriction enzymes that preferentially digest the cfDNA sample at regions of low utility may be selected for use in this targeted depletion method. In some embodiments, reduced-representation sequencing may be restriction enzyme-mediated reduced representation sequencing. The cfDNA sample may be contacted with one or more restriction enzymes, which cleaves the cfDNA at various recognition sites throughout the genome. Optionally, the cleavage products may next be ligated with a first adapter sequence. These ligated cleavage products may be randomly sheared before being ligated with a second adapter sequence. Only ligated cleavage products that have both adapter sequences will be amplified, thus ensuring that sequences in proximity to the restriction sites will be amplified for further processing and sequencing, while sequences distant from the restriction sites and generated from shearing only will not be amplified and processed. Additionally or alternatively, the cleavage products resulting from contacting the cfDNA sample with the one or more restriction enzyme may be subjected to a size selection step in order to eliminate fragments that fall outside the desired fragment length range. The cleavage products that have a length within the desired fragment length range may then be ligated with adapter sequences in order to amplify the cleavage products of interest. Additionally or alternatively, the cfDNA sample may be ligated with partial adapters prior to being contacted with the one or more restriction enzymes. Following contact with the one or more restriction enzymes, the uncut molecules may be purified with the sample and amplified using polymerase chain reaction (PCR) utilizing primers that may hybridize to the partial adapters in order to amplify the sequences that lack the restriction enzyme recognition site. One or more restriction enzyme may be utilized to digest the cfDNA sample, and one or more first adapter sequences or partial adapters may be used to ligate to the cleavage sites of the cfDNA sample based on which restriction enzyme generated the specific cleavage. The remainder of the cfDNA sample which is enriched for the ctDNA may be further processed and sequenced as discussed above.
[0204] In another embodiment, regions of low utility may be depleted from the cfDNA sample using affinity -based hybrid capture methods. These short DNA segments may be contacted with a plurality of capture oligonucleotides or capture probes following denaturation of the double-stranded structure of the short DNA segments to generate single-stranded short DNA segments. The plurality of capture oligonucleotides or captureprobes may be one or more oligonucleotides that contain sequences that are substantially or completely complementary to a sequence in the shorter DNA segments comprising the regions of low utility. The plurality of capture oligonucleotides or capture probes may include an affinity tag, including, but not limited to, biotin and dA tails. When the singlestranded short DNA segments are contacted with the plurality of capture oligonucleotides or capture probes, the single-stranded short DNA segments with regions of low utility that comprises sequences complementary to the sequence of one or more of the plurality of capture oligonucleotides or capture probes hybridize to the one or more of the plurality of capture oligonucleotides or capture probes to form a double-stranded DNA-capture oligonucleotide structure or double-stranded DNA-capture probe structure. These doublestranded DNA-capture oligonucleotide structures or double-stranded DNA-capture probe structures may be pulled down from the cfDNA sample by contacting the sample with reagents that bind the affinity tag of the capture oligonucleotides or capture probes. These reagents include, but are not limited to, magnetic beads comprising streptavidin when the affinity tag is biotin, magnetic beads comprising avidin when the affinity tag is biotin, glass beads comprising streptavidin when the affinity tag is biotin, glass beads comprising avidin when the affinity tag is biotin, membranes comprising streptavidin when the affinity tag is biotin, membranes comprising avidin when the affinity tag is biotin, magnetic beads comprising a polyT oligonucleotide when the affinity tag is dA tails, and glass beads comprising a polyT oligonucleotide when the affinity tag is dA tails. Pulldown of the double-stranded DNA-capture oligonucleotides or double-stranded DNA- capture probes depletes the DNA sample of the short DNA segments that comprise the regions of low utility. The remainder of the cfDNA sample which is enriched for the ctDNA may be further processed and sequenced as discussed above.
[0205] In another embodiment, regions of low utility may be depleted from the cfDNA sample using methods dependent upon cfDNA denaturation and reannealing kinetics. In some embodiments, the methods dependent upon cfDNA denaturation and reannealing kinetics include, but are not limited to, Cot analysis of the cfDNA sample, as discussed above. In some embodiments, the HR component samples may include the renatured cfDNA isolated at a low Cot value that has been separated from the remaining sample. In some embodiments, the MR component samples may include the renatured cfDNA isolated at an intermediate Cot value that has been separated from the remaining sample. In some embodiments, the SL component samples may include the renaturedcfDNA isolated at a high Cot value and any additional single-stranded cfDNA that has not renatured. The Cot values for the three components may be determined by statistical analysis of the generated Cot curve for the total cfDNA sample. In some embodiments, only the SL component samples may be used for further processing and sequencing as discussed above. In some embodiments, the SL component samples and the MR component samples may be used for further processing and sequencing as discussed above. In some embodiments the SL component samples, the MR component samples, and a portion of the HR component samples that are predicted to have a lower repetition frequency may be used for further processing and sequencing as discussed above. In some embodiments, a random portion of one or more of each of the component samples may be use for further processing and sequencing as discussed above. In some embodiments, a random portion of one or more of the component samples may be used for further processing and sequencing as discussed above. The final mixture of one or more component samples enriched for ctDNA sample may be further processed and sequenced as discussed above.
[0206] In another embodiment, the size of a DNA sample may be reduced using restriction enzymes in what is termed “reduced-representation sequencing.” Reduced- representation sequencing includes, but is not limited to, RAD-Seq, 2bRAD-Seq, ddRAD-Seq, and GBS-Seq. This method does not involve selection of restriction enzymes based on their ability to digest the cfDNA sample in regions of low utility, but instead allow for the use of any restriction enzymes regardless of where they are predicted to digest the sample. In some embodiments, reduced-representation sequencing may be restriction enzyme-mediated reduced representation sequencing. The cfDNA sample may be contacted with a restriction enzyme, which cleaves the cfDNA at various recognition sites throughout the genome. Optionally, the cleavage products may next be ligated with a first adapter sequence. These ligated cleavage products may be randomly sheared before being ligated with a second adapter sequence. Only ligated cleavage products that have both adapter sequences will be amplified, thus ensuring that sequences in proximity to the restriction sites will be amplified for further processing and sequencing, while sequences distant from the restriction sites and generated from shearing only will not be amplified and processed. Additionally or alternatively, the cleavage products resulting from contacting the cfDNA sample with the restriction enzyme may be subjected to a size selection step in order to eliminate fragments that fall outside the desired fragment lengthrange. The cleavage products that have a length within the desired fragment length range may then be ligated with adapter sequences in order to amplify the cleavage products of interest. One or more restriction enzyme may be utilized to digest the cfDNA sample, and one or more first adapter sequences may be used to ligate to the cleavage sites of the cfDNA sample based on which restriction enzyme generated the specific cleavage. The remainder of the cfDNA which is enriched for ctDNA sample may be further processed and sequenced as discussed above.
[0207] All publications and patents cited in this specification are herein incorporated by reference as if each individual publication or patent were specifically and individually indicated to be incorporated by reference and are incorporated herein by reference to disclose and describe the methods and / or materials in connection with which the publications are cited. The citation of any publication is for its disclosure prior to the filing date and should not be construed as an admission that the present invention is not entitled to antedate such publication by virtue of prior invention. Further, the dates of publication provided may be different from the actual publication dates that may need to be independently confirmed.EXAMPLES
[0208] The present technology is further illustrated by the following Examples, which should not be construed as limiting in any way.Example 1: Alu Y Hybridization Capture Depletion
[0209] Alu elements are highly repetitive DNA sequences that can be classified as short interspersed elements (SINEs), which are a type of retrotransposon. Alu repeats are the most abundant family of repeats in the human genome and make up an estimated 10% of the genome.
[0210] To test one possible method of subtractive hybridization for targeted depletion, Alu elements, specifically AluY elements, are targeted for depletion.
[0211] Libraries of mock cfDNA were prepared as the test DNA sample for targeted depletion. These mock cfDNA test samples were prepared by obtaining buffy coat DNA. These buffy coat DNA samples were then sheared on Covaris R230. Specifically, 2 pg ofthe buffy coat DNA were sheared in 50 pL to produce mock cfDNA test samples with DNA segments of between around 350 nucleotides and around 400 nucleotides in length (FIG. 2, Table 1). The mock cfDNA samples were dried down for around 30 minutes at 60 °C and then resuspended in hybridization reaction mix.
[0212] Table 1 : Quantification of Sheared Buffy Coat DNA Samples and Biotinylated AluY Probes
[0213] Probes designed to hybridize to AluY sequences were prepared biotinylated using biotinylated PCR primers to produce probes that are between around 120 to around 145 nucleotides in length that are complementary to the AluY sequence and have a biotin tag on the 5’ end.
[0214] The mock cfDNA test samples were denatured to produce single-stranded DNA segments. These single-stranded DNA segments were then contacted with the biotinylated AluY probes, and the probes are allowed to hybridize to any AluY sequences in the single-stranded DNA segments of the mock cfDNA sample. Specifically, a series of samples were utilized using varied amounts of DNA input (0 ng, 50 ng, 100 ng, and 500ng) with either a constant probe amount of 100 ng or water with no probe for nonspecific depletion controls. The samples were incubated at 95 °C for two minutes to denature the test samples and at 60 °C for at least 16 hours.
[0215] Table 2: Hybridization Reaction Sample Set-Up
[0216] Following hybridization, the samples were contacted with streptavidin beads, which are capable of binding to the biotin tags on the biotinylated AluY probes (Table 3). 0.5 mg of streptavidin beads were washed in lx Bead Wash buffer and resuspended in hybridization reaction mix. A first aliquot of resuspended beads was added to the samples, mixed via eversion five times every ten minutes, and incubated for 20 minutes. The first aliquot of beads was pulled down from the sample, and the supernatant was transferred to a second aliquot of beads for another round of mixing and incubation as described above. After incubation with the second aliquot of beads, the second aliquot of beads was pulled down from the sample, and the supernatant was transferred to a final third aliquot of beads for another round of mixing and incubation as described above. After incubation with the third aliquot of beads, the third aliquot of beads was pulled down from the sample, and the supernatant was transferred to an empty container. 20 pL of water was added to each aliquot of beads.
[0217] Table 3: Biotinylated Primers
[0218] The remaining supernatant for each sample was diluted to a total volume of 150 pL with elution buffer. The diluted supernatant was cleaned up using the 1.3X AMPure bead cleanup protocol. The supernatant was allowed to bind for 10 minutes, shaking at 1,800 RPM for several minutes before being washed twice with 500 pL of 80% ethanol. The supernatant samples were allowed to dry for five minutes. The dried supernatant samples were then eluted in 52 pL of elution buffer for five minutes, shaking at 1,800 RPM for a minute. 50 pL of the final purified DNA was recovered from each supernatant sample.
[0219] The final purified DNA was quantified on Qubit (FIG. 3). There was no obvious difference in the amount of DNA quantified between samples that were subjected to hybridization with the probe or without the probe added.
[0220] The efficiency of the subtractive hybridization method was then tested using quantitative real-time PCR (qPCR) of the final purified DNA from each sample to measure the amount of AluY remaining in the depleted mock cfDNA test sample as compared to a control mock cfDNA test sample that was not contacted with the AluY probes. To assess any change in Alu abundance between the various hybridization conditions, delta-delta Ct analysis was performed for each of the Alu primer sets relative to the housekeeping primer set (GAPDH), and using the input sample as the reference sample, a fold change in Alu abundance was calculated via using the delta-delta Ct approach for each Alu primer set. Three sets of Alu qPCR primers were used to detect the remaining amount of AluY in the depleted mock cfDNA test samples along with a control GAPDH primer set (Table 4). The Ct values were much lower for all three of the Alu primer sets than the GAPDH control (Table 5, FIG. 4A). In generally, the Ct values decreased with increasing sample input as expected, but surprisingly, there was significant amplification of all three Alu primer sets in the probe-only control, indicating that a significant amount of probe is not removed from the supernatant. It was found that two of the Alu primer sets (probe set 1 and probe set 2) readily amplified from the biotinylated probe as well, indicating that that levels of Alu detected in the mock cfDNA samples may not be accurate using those primer sets.
[0221] Table 4: Alu Primer Sets
[0222] Table 5 : qPCR Results Following Hybridization
[0223] First, there was a significant increase in the abundance of Alu under almost all hybrid capture conditions relative to input, regardless of whether probes were present or not. This suggests that the hybridization capture incubation on its own somehow enriches for repetitive elements. In general, the fold increase in Alu was inversely correlated with sample input. The addition of biotinylated capture probe had a mixed impact on the fold change depending on the Alu primer set used in the qPCR detection step (FIGS. 4B-4D). This analysis is confounded by the fact that it appears not all of the probe was removed from the capture reaction and that two primer sets are capable of amplifying the probe itself. In all, there was one data point that suggested that there was depletion of Alu relative to the input material using the one primer set that did not have detect the capture probe (500 ng unput + probe). It was found that there was around a two-fold reduction in the abundance of Alu levels in this sample, thus indicating that the method is effective in targeted depletion.
[0224] To determine how much probe was captured with the first, second, and third streptavidin pulldown, the DNA bound to the bead was used as input for a PCR reaction using one of the primer sets known to amplify directly from the probe. The post-PCR material was run on TapeStation to assess amplification (FIG. 5).
[0225] Amplification was readily detected from the DNA from the first aliquot of beads. Limited amplification was observed for the DNA from the second aliquot of beads, primarily in the probe-only condition, and no amplification was observed using the DNAfrom the third aliquot of beads. This suggests that the vast majority of biotinylated probe was captured with the first pulldown. This raises the probability that a significant number of molecules do not contain the biotin modification, possibly due to incomplete biotinylation from IDR (combined with standard desalting) or due to primer-independent priming during the PCR reaction to generate the probe. Unbiotinylated probe confounds analysis by qPCR if the primers amplify the probe as well.Example 2: Hybrid Capture and Reduced Representation via Restriction Enzyme Digestion (RED) Depletion
[0226] Samples for this example consisted of buffy coat isolated from duplicate Streck tubes from Prequel. Samples were research-allowed and de-identified. Genomic DNA was extracted from around 100 pL using the MagMAX DNA Multi-Sample Ultra 2.0 kit and the ThermoFisher Flex. Extracted DNA was quantified via Qubit (Table 6).
[0227] Table 6: Concentrations of Genomic DNA Samples
[0228] DNA fragmentation
[0229] Acoustic Shearing: Covaris R230 was used to shear DNA to around 350 nucleotides, which took three runs of the program or around 60 cycles to achieve this size. The sheared DNA samples were diluted to have around 4 pg in a total volume of 100 pL (40 ng / pL) and split across two wells of a TPX strip.
[0230] Restriction Enzyme Digestion: Restriction digest of intact buffy genomic DNA was performed using two blunt cutters with 4 base recognition sites: HpyCH4V (TG|CA) and Alul (AG|CT). 1 pg of DNA was digested in a total reaction volume of 50 pL for 15 minutes at 37 °C per manufacturer’s instructions. Two replicates, for a total of 4 pg, per sample per restriction enzyme were included.
[0231] Following the digestion incubation, the two replicates were pooled (100 pL; 2 pg) and then split between two purification strategies: (a) 2.5X AMPure cleanup (70 pL) and (b) Blue Pippin size-selection (30 pL; 2% agarose cassette, 200-550 base pairs).Results were similar between the two biological samples (Al and Bl) and for both restriction enzymes (FIGs. 6A and 6B). Purification via Blue Pippin had a lower recovery rate (around 12% versus around 30%) but resulted in a tighter distribution of fragments as compared to the 2.5X AMPure cleanup purification (Table 7).
[0232] Table 7: Concentration and Percent Recovery Following Purification
[0233] Library Preparation
[0234] Libraries were prepared from 100 ng fragmented DNA (sample Al) using the IDT xGen cfDNA and FFPE library prep kit. Replicates were included for the Covaris- sheared DNA to provide unique indexes so hybridization capture conditions could be tested and sequenced together. Libraries were prepared per IDT’s instructions, with cleanup performed with Counsyl AMPure beads.
[0235] Following adapter ligation, indexing PCR (50 pL) was performed using the included IDT 2X HiFi PCR mix and 10 base pair UDI primers. PCR was performed for 8 cycles using an Eppendorf MasterCycler 50i (8 cycles to ensure sufficient DNA for multiple capture runs). A 1.3X AMPure cleanup was performed to purify the indexed libraries. Libraries were quantified on the Qubit (HS) (Table 8) and TapeStation (D1000) (FIG. 7). All samples had sufficient yield for multiple hybridizations.
[0236] Table 8: Concentrations and Total Library Yields
[0237] Probes
[0238] A candidate list of repetitive elements (n = 595) was selected, and probes (n = 1,775) were designed using the MRD Oligo Designer, merging the BLAST hit coordinates for all probes and calculating the ROI size returns bases, which is around 10% of the human genome. Analysis of the probes indicated that only 538 of the MRD probes overlapped with the repetitive regions, indicating that over 98% of the somatic sites would still be identifiable after depletion.
[0239] Subtractive Hybridization Capture
[0240] A typical hybridization capture reaction uses 500 ng of library and 4 pL of biotinylated probes, resulting in a probe-to-target ratio of around 1,000 to 1, assuming that the target is unique in the human genome.
[0241] For targeting repetitive elements that are highly abundant throughout the genome, the probe-to-target ratio will be much lower when standard conditions are used. Here, the probe-to-target ratio was varied by increasing the amount of probe and / or decreasing the amount of library input (Table 9).
[0242] Table 9: Samples Utilized for Hybridization Capture Depletion
[0243] Library, probes, and 10 base pair blockers were combined and dried down via vacufuge for 30 minutes at 60 °C. The DNA (probes, blockers, library) was thoroughly resuspended in hybridization reaction mix (hybridization buffer, enhancer, and water), denatured at 95 °C for two minutes, and hybridized at 60 °C overnight for at least 16 hours.
[0244] Streptavidin beads were washed according to IDT’s protocol and were resuspended in hybridization reaction mix. The beads were preheated to 60 °C for 5 minutes before being added to the overnight hybridization reactions. The beads were incubated with the hybridization reaction samples for 30 minutes and were mixed every 10 minutes to ensure the beads remained in solution. After 30 minutes, the beads were collected using a magnet, and the supernatant was added to another aliquot of pre-washed beads for a second round of streptavidin : biotin capture. After the second 30-minute bead capture incubation, the final depleted supernatant as transferred to a new well.
[0245] The final hybridization supernatant or depleted supernatant was diluted around 4x with elution buffer to reduce the formamide concentration and to produce a final volume of 150 pL. The depleted supernatant was then purified via a 1.3X AMPure bead (Counsyl) cleanup with a 10-minute binding incubation to maximize recovery. The recovered DNA was eluted in 50 pL of elution buffer and quantified via Qubit (lx HS) (Table 10). The DNA yield was consistently around 20% to around 30% of input. It was observed that lower library input exhibited a higher percent recovery.
[0246] Table 10: Total DNA and Percent Recovery in Depleted Supernatant Samples
[0247] DNA from the hybridization supernatant was then PCR-amplified with p5 / p7 primers to ensure DNA was double stranded for more accurate quantification prior to sequencing. Libraries were amplified with Kapa HiFi reagents for six cycles. Two input amounts were used: 2 pL and 4 pL. Following PCR, 1 pL of the unpurified PCR reaction was run on the TapeStation (D1000) (FIG. 8).
[0248] The remaining material was purified via a 1.3X AMPure bead (Counsyl) cleanup, and the eluted DNA (50 pL) was quantified via Qubit (lx HS). Library sizes were within the expected range, and quantifications were within reason.
[0249] qPCR
[0250] Primers were designed to target two repetitive elements from the RE depletion panel, representing target sites / probes with a high and low number of BLAST hits. qPCR was performed for each sample, and the fold change was calculated via delta-delta CT analysis.
[0251] As expected, the target site with a single BLAST hit exhibited better depletion than the target site with many BLAST hits. Varying the library-to-probe ratio had mixed effects as measured by different amplicons. For example, 1888_set3 and 3832_set4 exhibited little change, but 3832_set3 suggests increased absolute probe concentration has a more pronounced effect than the ratio of library and probe (FIG. 9). These results suggest repetitive element depletion was generally successful, but the effectiveness wasdependent on the number of target sites. In the best case, site 3832_set3 was reduced to less than 10% of input.
[0252] The same primers were used to measure the impact of reduced representation via restriction enzyme digest (RED) (FIGs. 10A and 10B). Again, fold change for the targets varied with restriction enzyme and whether or not post-digestion size selections was performed (200-550 bp). In some cases, sites were reduced to less than 1% of input.
[0253] Sequencing
[0254] All of the samples were pooled, along with two control samples and sequenced on the NovaSeq 6000. All of the libraries were normalized to 1.5 nM and pooled at equal volume, except the restriction enzyme digested libraries, which were pooled at 1 / 3 the volume due to the greater reduction in effective genome size (Table 11).
[0255] Table 11 : Pooling Ratios for Sequencing Trials
[0256] The final 1.5 nM library pool was denatured and loaded onto an S2 flow cell for 2x150 paired-end sequencing, targeting around 30x Whole Genome Sequencing coverage.
[0257] On average the TDS / WGS libraries receiver around 497 M clusters PF and the RED libraries received -160 M clusters PF, as expected (FIG. 11).
[0258] The sequencing data and a sample sheet were transferred to DNAnexus for alignment and analysis. Alignment was performed using the current MRD alignment workflow.
[0259] Results
[0260] Mean coverage was around 30x for most of the TDS libraries, except for 250ng_2x which only had 22x coverage, and around 25x for the control whole genome sequenced libraries (FIGs. 12A-12E). Apparent coverage for the RED was very low (around 2x), but this is due to inappropriate de-duplication based on start-stop positions. Additionally, 8 base pairs was trimmed off each read prior to or during alignment due to incorrect UMVbase-mask settings applied during read analysis and alignment.
[0261] Depletion Analysis - Hybridization Capture
[0262] Depletion was calculated by normalized depth ratio between test and control samples for each repetitive element target. First, the average depth for each repetitive element site was calculating using Bedtools Coverage (FIG. 13 A). The average depths for each sample were normalized to account for variability in raw sequencing depth (reads / sample) (FIG. 13B). For each sample, the normalized repetitive element site depths were then divided by the average of the normalized depth of the corresponding sites from the whole genome sequencing replicates to get the normalized depth ratio (FIG. 13C). In the plots, the y-axis is restricted to better represent the majority of the data, but some outliers are not shown.
[0263] Comparison of Experimental Conditions: Overall, depletion was relatively consistent across the experimental conditions, but increasing the probe-to-target ratio from 1,000: 1 to 4,000:1 slightly improved depletion (Table 12). In the experiment, the probe-to-target ratio was modulated by increasing probe concentration and / or decreasing library concentration. Increasing probe concentration (2x vs lx) appeared to have a bigger impact than reducing the library concentration (500 ng vs. 250 ng input), consistent with hybridization reaction kinetics / efficiency being concentration-dependent (FIG. 14).
[0264] Table 12: Normalized Depth Ratios for Probe to Target Ratio
[0265] The depletion at each target site was also roughly correlated with the number of BLAST hits for the associated probes (FIG. 15A = Normalized Depth; FIG. 15B = Norm Depth Ratio); however, the probes were heavily skewed toward fewer BLAST hits. Generally ,the more BLAST hits a probe had, the less depletion observed at the intended target, suggesting the probe to target ratio is too low for effective depletion.
[0266] Examples of Maximum Depletion (Table 13): The lowest depth ratio, corresponding to maximum depletion, was observed for the TDS_250ng_2xProbe condition. The minimum depth ratio was 0.078, corresponding to an around 13-fold depletion of reads at that target.
[0267] Table 13 : Depth Ratio for Probes with Different BLAST Hits
[0268] Subtractive hybridization resulted in around a two-fold depletion (maximum of 13 -fold depletion) of repetitive elements with less than 10,000 BLAST hits. Depletion varied widely but was roughly correlated with the number of BLAST hits. Increasing probe concentration helped more than reducing DNA input.
[0269] Depletion Analysis - Restrict Enzyme Digestion
[0270] Although the RED approach does not specifically target repetitive elements like the hybridization capture approach described above, a similar analysis was performed just to evaluate the effectiveness of depletion across the same sites. Depletion was calculated as the depth ratio between test and control samples for each repetitive element target. First, the average depth for each RE site was calculate using Bedtools Coverage (FIG. 16A). The normalization step was skipped because the RED samples should require significantly fewer reads than the whole genome sequencing controls. For each sample, the repetitive element site depths were then divided by the average depth of the corresponding sites from two whole genome sequencing replicates to get the depth ratio (FIG. 16B). In the plots, the y-axis is restricted to better represent the majority of the data, but some outliers are not shown.
[0271] Several sites had depth ratios equal to zero, corresponding to complete lack of coverage (Table 14).
[0272] Table 14: Samples with No Coverage
[0273] Impact on ExplO Somatic Sites: A .csv file that contains all of the ExplO target sites and probes used was generated. The file represents around 1,000 somatic sites times around 30 individuals to produce around 30,000 sites. Bedtools Coverage was used to estimate the number of ExplO sites retained. It was found that digestion using the HpyCH4V enzyme resulted in around 33.5% of the ExplO sites having the 120 base pair target sequence retained, while around 51.8% of the ExplO sites had greater than or equal to 50% of the 120 base pair target sequence retained. Digestion using the Alul enzyme resulted in around 32.5% of the ExplO sites having the 120 base pair target sequenceretained, while around 50.3% of the Exp 10 sites had greater than or equal to 50% of the 120 base pair target sequence retained.
[0274] Restriction enzyme digestion resulted in the complete depletion of certain regions of the genome. Digestion with HpyCH4V or Alul and selection for fragments that are between 200 base pairs and 550 base pairs resulted in around a 70% reduction in genome size and around a 50% reduction in coverage of Exp 10 somatic sites.Example 3 — Target Depletion Sequencing (TDS) compatibility with somatic calling
[0275] DNA from a matched Tumor-Normal cell line pair (2343_2363) was used to prepare sequencing libraries and 4 replicate libraries were prepared per cell line (i.e., 4x Tumor, 4x Normal). Tumor and Normal replicates were divided between 4 experimental conditions (see Table 15) to compare standard Whole Genome Sequencing (WGS) to Target Depletion Sequencing (TDS).
[0276] Two panels were used in the TDS arm of this experiment: a set of bioiotynylated DNA oligos (Twist) targeting specific Repetitive Elements (RE). For the RE Probe set, ROI = -1,700 probes covering -153 Kb (not including “off-target” base pairs) or -49.7 Mb (factoring in 324 median number of blast hits per target region), representing -0.005 - 1.5% of the human genome. The second panel contained biotinylated Cot-1 DNA covering 10-50% of the human genome (up to 1,500 Mb).
[0277] Libraries were sequenced to a target depth of coverage of ~30x on a NovaSeq X (25B flow cell). Samples were processed to perform demux, alignment, and somatic calling. Additional analyses were performed on a DNAnexus workstation. Average depths were computed across various regions of interest (ROIs) using mosdepth software, which is compatible with the new CRAM outputs of the Dragen alignment and depths were normalized based on the total number of sequencing reads for each library.
[0278] Sequencing Depth
[0279] Depth, normalized depth, and the normalized depth ratio (relative to the WGS control) were calculated for each target site (-1,700) from the RE Probes panel for each of the Tumor and Normal libraries.
[0280] TDS performed with the RE Probe Panel resulted in the lowest median normalized depth ratio (i.e., largest reduction in coverage) for the associated ROI. TDS performed with the Cot-1 Probe Panel also resulted in a slightly reduced median normalized depth ratio, indicating there is some overlap between the two ROIs.
[0281] Somatic Calling and Somatic Variant Read Counts
[0282] Approximately 12,000 high-confidence (Tier 1) somatic variants were identified in each sample. The median normalized tumor read counts (both alt and total) across the identified variants was higher (+2-4%) for samples that underwent TDS with either the RE Probe Panel or the Cot-1 Probe Pabel. This indicates that targeted depletion of repetitive elements can increase coverage of somatic variants. Increased coverage at somatic sites improves confidence in somatic calling results.
[0283] Table 15: Effect of Targeted Depletion of Repetitive Elements on Coverage ofSomatic Variants
[0284] This experiment indicates that TDS is generally compatible with somatic calling and can increase read counts for somatic variants. Increased coverage increases confidence in somatic calls.
[0285] This experiment evaluated TDS of tumor-normal libraries using two panels. The first panel (RE Probes) had higher depletion efficiency but a smaller ROI than the second panel (Cot-1 probes), which had a lower depletion efficiency but a larger ROI. Both panels resulted in an increase in coverage of somatic sites, indicating breadth and depletion efficiency are both key factors for maximizing TDS performance.
[0286] Assuming repetitive elements represent -50% of the human genome and 100% depletion efficiency, TDS can increase coverage for the same number ofsequencing reads without TDS, increasing somatic calling performance and resulting in reduced sequencing cost.EQUIVALENTS
[0287] The present technology is not to be limited in terms of the particular embodiments described in this application, which are intended as single illustrations of individual aspects of the present technology. Many modifications and variations of this present technology can be made without departing from its spirit and scope, as will be apparent to those skilled in the art. Functionally equivalent methods and apparatuses within the scope of the present technology, in addition to those enumerated herein, will be apparent to those skilled in the art from the foregoing descriptions. Such modifications and variations are intended to fall within the scope of the present technology. It is to be understood that this present technology is not limited to particular methods, reagents, compounds, compositions or biological systems, which can, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting.
[0288] All patents, patent applications, provisional applications, and publications referred to or cited herein are incorporated by reference in their entirety, including all figures and tables, to the extent that are not inconsistent with the explicit teachings of this specification.
Claims
CLAIMSWhat is claimed is:
1. A method of preparing a patient-specific panel of tumor-specific somatic mutations, comprising:(a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample;(b) depleting DNA fragments that comprise a region of low utility from the first DNA sample and the second DNA sample;(c) sequencing the first depleted DNA sample and the second depleted DNA sample; and(d) generating a patient-specific panel of tumor-specific somatic mutations that are present in the first depleted DNA sample and absent in the second depleted DNA sample.
2. The method of claim 1 further comprising preparing a plurality of oligonucleotide probes, wherein each probe in the plurality of oligonucleotide probes hybridizes to a tumor-specific somatic mutation in the patient-specific signature panel.
3. A method of preparing a circulating tumor DNA (ctDNA)-enriched fraction from a biological sample, comprising:(a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample;(b) depleting DNA fragments that comprise a region of low utility from the first DNA sample and the second DNA sample;(c) sequencing the first depleted DNA sample and the second depleted DNA sample;(d) obtaining a further non-tumor sample from the subject;(e) extracting cell free DNA (cfDNA) from the further non-tumor sample; and(f) enriching the cfDNA from the further non-tumor sample, wherein enriching comprises:(i) contacting the cfDNA from the further non-tumor sample with a plurality of oligonucleotide probes, wherein each probe in the plurality of oligonucleotide probes is capable of hybridizing to a DNA fragment comprising a tumor-specific somatic mutation that is present in the first depleted DNA sample and absent in thesecond depleted DNA sample, thereby obtaining a DNA fraction that is enriched for ctDNA,(ii) depleting DNA fragments that comprise a region of low utility from the cfDNA from the further non-tumor sample, thereby obtaining a DNA fraction that is enriched for ctDNA, or(iii) reducing the size of the cfDNA sample by digesting the cfDNA sample with at least one restriction enzyme, thereby obtaining a DNA fraction that is enriched for ctDNA.
4. A method of detecting circulating tumor DNA (ctDNA) in a sample, comprising:(a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample;(b) depleting DNA fragments that comprise a region of low utility from the first DNA sample and the second DNA sample;(c) sequencing the first depleted DNA sample and the second depleted DNA sample;(d) obtaining a further non-tumor sample from the subject;(e) extracting cell free DNA (cfDNA) from the second non-tumor sample;(f) sequencing the cfDNA, thereby obtaining a plurality of sequence reads; and(g) detecting the presence or absence of ctDNA, wherein the presence of a sequence read in the plurality of sequence reads comprising a nucleic acid sequence comprising one or more of the tumor-specific somatic mutations indicates the presence of ctDNA.
5. A method of monitoring cancer recurrence, response to treatment, or disease progression in a subject previously treated for cancer, comprising:(a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample;(b) depleting DNA fragments that comprise a region of low utility from the first DNA sample and the second DNA sample;(c) sequencing the first depleted DNA sample and the second depleted DNA sample;(d) obtaining a further non-tumor sample from the subject;(e) extracting cell free DNA (cfDNA) from the second non-tumor sample;(f) sequencing the cfDNA, thereby obtaining a plurality of sequence reads; and(g) detecting the presence or absence of any sequence reads in the plurality of sequence reads, wherein the presence of a sequence read comprising a nucleic acidsequence comprising one or more of the tumor-specific somatic mutations indicates (i) cancer recurrence, (ii) a lack of responsiveness to treatment, (iii) disease progression, or (iv) any combination thereof.
6. The method of claim 4 or 5, wherein (d)-(g) are repeated at one or more times during a cancer treatment, following completion of a cancer treatment, while the subject is in remission, or coinciding with or prior to surgery.
7. The method of any one of claims 4-6, wherein prior to sequencing the cfDNA, the cfDNA is enriched by (i) contacting the cfDNA from the further non-tumor sample with a plurality of oligonucleotide probes, wherein each probe in the plurality of oligonucleotide probes is capable of hybridizing to a DNA fragment comprising a tumor-specific somatic mutation that is present in the first depleted DNA sample and absent in the second depleted DNA sample, thereby obtaining a DNA fraction that is enriched for ctDNA, (ii) depleting DNA fragments that comprise a region of low utility from the cfDNA from the further non-tumor sample, thereby obtaining a DNA fraction that is enriched for ctDNA, or (iii) reducing the size of the cfDNA sample by digesting the cfDNA sample with at least one restriction enzyme, thereby obtaining a DNA fraction that is enriched for ctDNA.
8. The method of any one of claims 1-7, wherein (i) the region of low utility from the first DNA sample and the second DNA sample and (ii) the region of low utility from the cfDNA each independently comprise repetitive elements, low GC content, high GC content, or any combination thereof.
9. The method of any one of claims 1-8, wherein (i) the region of low utility from the first DNA sample and the second DNA sample and (ii) the region of low utility from the cfDNA each independently do not contain or are unlikely to contain somatic mutations.
10. The method of any one of claims 1-9, wherein depleting DNA fragments that comprise a region of low utility from the first DNA sample and the second DNA sample comprises nuclease-based depletion, restriction enzyme-based depletion, subtractive hybridization, or kinetic-based depletion.
11. The method of any one of claims 7-10, wherein depleting DNA fragments that comprise a region of low utility from the cfDNA from the further non-tumor sample comprises nuclease-based depletion, restriction enzyme-based depletion, subtractive hybridization, or kinetic-based depletion.
12. The method of claim 10 or 11, wherein nuclease-based depletion comprises targeting depletion of sequences with a DNA-guided nuclease, an RNA-guided nuclease, or a structure-guided nuclease.
13. The method of claim 12, wherein the DNA-guided endonuclease is Thermus thermophilus Argonaute (TtAgo) or a transcription activator-like effector nuclease (TALEN).
14. The method of claim 12, wherein the RNA-guided endonuclease is a Cas endonuclease15. The method of claim 10 or 11, wherein endonuclease-based depletion comprises Depletion of Abundant Sequences by Hybridization (DASH).
16. The method of claim 10 or 11, wherein kinetic-based depletion comprises depleting DNA fragments based on denaturation and annealing kinetics.
17. The method of claim 10 or 11, wherein restriction enzyme-based depletion comprises reduced-repre sentati on sequencing .
18. A method of preparing a patient-specific panel of tumor-specific somatic mutations, comprising:(a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample;(b) reducing the size of the first DNA sample and the second DNA sample by digesting the first DNA sample and the second DNA sample with at least one restriction enzyme;(c) sequencing the first reduced DNA sample and the second reduced DNA sample; and(d) generating a patient-specific panel of tumor-specific somatic mutations that are present in the first DNA sample and absent in the second DNA sample.
19. The method of claim 18 further comprising preparing a plurality of oligonucleotide probes, wherein each probe in the plurality of oligonucleotide probes hybridizes to a tumor-specific somatic mutation in the patient-specific signature panel.
20. A method of preparing a circulating tumor DNA (ctDNA)-enriched fraction from a biological sample, comprising:(a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample;(b) reducing the size of the first DNA sample and the second DNA sample by digesting the first DNA sample and the second DNA sample with at least one restriction enzyme;(c) sequencing the first reduced DNA sample and the second reduced DNA sample;(d) obtaining a further non-tumor sample from the subject;(e) extracting cell free DNA (cfDNA) from the further non-tumor sample; and(f) enriching the cfDNA from the further non-tumor sample, wherein enriching comprises:(i) contacting the cfDNA from the further non-tumor sample with a plurality of oligonucleotide probes, wherein each probe in the plurality of oligonucleotide probes is capable of hybridizing to a DNA fragment comprising a tumor-specific somatic mutation that is present in the first reduced DNA sample and absent in the second reduced DNA sample, thereby obtaining a DNA fraction that is enriched for ctDNA,(ii) depleting DNA fragments that comprise a region of low utility from the cfDNA from the further non-tumor sample, thereby obtaining a DNA fraction that is enriched for ctDNA, or(iii) reducing the size of the cfDNA sample by digesting the cfDNA sample with at least one restriction enzyme, thereby obtaining a DNA fraction that is enriched for ctDNA21. A method of detecting circulating tumor DNA (ctDNA) in a sample, comprising:(a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample;(b) reducing the size of the first DNA sample and the second DNA sample by digesting the first DNA sample and the second DNA sample with at least onerestriction enzyme;(c) sequencing the first reduced DNA sample and the second reduced DNA sample;(d) obtaining a further non-tumor sample from the subject;(e) extracting cell free DNA (cfDNA) from the second non-tumor sample;(f) sequencing the enriched DNA fraction, thereby obtaining a plurality of sequence reads; and(g) detecting the presence or absence of ctDNA, wherein the presence of a sequence read in the plurality of sequence reads comprising a nucleic acid sequence comprising one or more of the tumor-specific somatic mutations indicates the presence of ctDNA.
22. A method of monitoring cancer recurrence, response to treatment, or disease progression in a subject previously treated for cancer, comprising:(a) obtaining from a subject with a prior diagnosis of a cancer a first DNA sample from a tumor sample and a second DNA sample from a non-tumor sample;(b) reducing the size of the first DNA sample and the second DNA sample by digesting the first DNA sample and the second DNA sample with at least one restriction enzyme;(c) sequencing the first reduced DNA sample and the second reduced DNA sample;(d) obtaining a further non-tumor sample from the subject;(e) extracting cell free DNA (cfDNA) from the second non-tumor sample;(f) sequencing the enriched DNA fraction, thereby obtaining a plurality of sequence reads; and(g) detecting the presence or absence of any sequence reads in the plurality of sequence reads, wherein the presence of a sequence read comprising a nucleic acid sequence comprising one or more of the tumor-specific somatic mutations indicates (i) cancer recurrence, (ii) a lack of responsiveness to treatment, (iii) disease progression, or (iv) any combination thereof.
23. The method of claim 21 or 22, wherein (d)-(g) are repeated at one or more times during a cancer treatment, following completion of a cancer treatment, while the subject is in remission, or coinciding with or prior to surgery.
24. The method of any one of claims 21-23, wherein prior to sequencing the cfDNA, the cfDNA is enriched by (i) contacting the cfDNA from the further non-tumor sample with a plurality of oligonucleotide probes, wherein each probe in the plurality ofoligonucleotide probes is capable of hybridizing to a DNA fragment comprising a tumor-specific somatic mutation that is present in the first depleted DNA sample and absent in the second depleted DNA sample, thereby obtaining a DNA fraction that is enriched for ctDNA, (ii) depleting DNA fragments that comprise a region of low utility from the cfDNA from the further non-tumor sample, thereby obtaining a DNA fraction that is enriched for ctDNA, or (iii) reducing the size of the cfDNA sample by digesting the cfDNA sample with at least one restriction enzyme, thereby obtaining a DNA fraction that is enriched for ctDNA.
25. The method of any one of claims 1-24, wherein the tumor-specific somatic mutation is obtained by aligning sequences of the first DNA sample to a reference human genome that is not from the subject, aligning sequences of the second DNA sample to a reference human genome that is not from the subject, and selecting mutations that are present in the first DNA sample and absent in the second DNA sample.
26. The method of any one of claims 1-24, wherein the tumor-specific somatic mutation is obtained by aligning sequences of the first DNA sample to sequences of the second DNA sample.
27. The method of any one of claims 1-26, wherein the tumor-specific somatic mutation(s) comprises one or more somatic mutations selected from SNVs, insertions, deletions, and translocations.
28. The method of any one of claims 1-27, wherein first non-tumor sample comprises a tissue sample matched to a tissue of origin of the tumor sample.
29. The method of any one of claims 1-28, wherein first non-tumor sample comprises a fluid sample selected from a buffy coat sample, blood, blood plasma, blood serum, urine, saliva, and cerebral spinal fluid (CSF).
30. The method of any one of claims 2-17 or 19-29, wherein the plurality of oligonucleotide probes is capable of detecting at least 10, at least 50, at least 100, at least 150, at least 200, at least 250, or at least 500 tumor-specific somatic mutations.
31. The method of any one of claims 3-17 or 20-30, wherein enriching the cfDNA from the second non-tumor sample comprises (i) hybrid capture-based enrichment, (ii) PCR-target enrichment, or (iii) on-sequencer enrichment.
32. The method of any one of claims 3-17 or 20-31, wherein the further non-tumor sample comprises a fluid sample selected from a buffy coat sample, blood, blood plasma, blood serum, urine, saliva, and cerebral spinal fluid (CSF).
33. The method of any one of claims 4-17 or 21-32, wherein sequencing the enriched DNA fraction comprises next generation sequencing.
34. The method of claim 7-17 or 24-33, wherein the enriched DNA fraction comprises sequencing of introns, exons, intergenic regions, or a combination thereof35. The method of any one of claims 1-34, wherein the cancer is selected from adrenal cancer, anal cancer, bile duct cancer, bladder cancer, bone cancer, a brain / CNS tumor, breast cancer, Castleman disease, cervical cancer, colon or rectum cancer, endometrial cancer, esophagus cancer, a Ewing tumor, eye cancer, gallbladder cancer, a gastrointestinal carcinoid tumor, a gastrointestinal stromal tumor (GIST), gestational trophoblastic disease, Hodgkin disease, Kaposi sarcoma, kidney cancer, laryngeal and hypopharyngeal cancer, leukemia, liver cancer, lung cancer, lymphoma, malignant mesothelioma, multiple myeloma, myelodysplastic Syndrome, nasal cavity or paranasal sinus cancer, nasopharyngeal cancer, neuroblastoma, oral cavity or oropharyngeal cancer, osteosarcoma, ovarian cancer, pancreatic cancer, penile cancer, a pituitary tumor, prostate cancer, retinoblastoma, rhabdomyosarcoma, salivary gland cancer, skin cancer, small intestine cancer, stomach cancer, testicular cancer, thymus cancer, thyroid cancer, uterine sarcoma, vaginal cancer, vulvar cancer, Waldenstrom macroglobulinemia, and Wilms tumor.
36. An enriched ctDNA fraction prepared according to the method of claim 3 or 20.
37. The enriched ctDNA fraction of claim 36, wherein the plurality of oligonucleotide probes is capable of detecting at least 10, at least 50, at least 100, at least 150, at least 200, at least 250, or at least 500 tumor-specific somatic mutations.
38. The enriched ctDNA fraction of claim 36 or 37, wherein the enriched ctDNA fraction comprises cfDNA fragments that collectively comprise at least 10, at least 50, at least 100, at least 150, at least 200, at least 250, or at least 500 tumor-specific somatic mutations.
39. The enriched ctDNA fraction of any one of claims 36-38, wherein the DNA fraction is enriched for cfDNA fragments averaging less than about 200 base pairs in length.
Citation Information
Patent Citations
Methods of amplifying and sequencing nucleic acids
US7323305B2
Improvements in variant detection
CA3119078A1
Methods and systems for analyzing nucleic acid molecules
US11634779B2
Methods of enriching nucleic acids
WO2023148235A1