Genetic and epigenetic detection in a single workflow
Patent Information
- Application Number
- JP2024539661
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-12-29
- Filing Date
- 2022-12-28
- Publication Date
- 2025-12-16
AI Technical Summary
The prior art is difficult to efficiently detect gene and epigenetic information simultaneously in a single workflow, resulting in high detection costs and increased complexity. Especially in the case of limited DNA amounts in liquid biopsy, multiple detections cannot be effectively realized assays.
By using a resistant nucleic acid polymerase and a specific nucleotide mixture, a round of primer extension is performed to generate the first and second strands of gene and epigenetic information, and the resistant nucleotide bodies introduced during primer extension are used to protect the gene information from being mistransformed, while simultaneously generating detectable epigenetic information.
It realizes efficient and simplified genetic and epigenetic information detection in a single workflow, improves the sensitivity and information output of liquid biopsy, and reduces detection costs.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 294,640, filed December 29, 2021, which is incorporated by reference herein in its entirety.
[0002] Electronic Sequence Listing Reference The contents of the electronic sequence listing (197102007240seqlist.xml, size: 4,557 bytes, and creation date: December 22, 2022) are incorporated by reference herein in their entirety.
[0003] Provided herein are methods relating to detecting genetic and epigenetic information in a single workflow, as well as related diagnostic, prognostic, monitoring, screening, and therapeutic methods, systems, and computer-readable storage media. [Background technology]
[0004] Genetic mutations are the hallmark of cancer. On the other hand, aberrant DNA methylation is a widespread form of epigenetic alteration that occurs early in carcinogenesis. Rather than using genetic alteration analysis alone, new approaches in precision medicine attempt to characterize and detect signals from multiple tumors to provide a holistic picture of a patient's disease state. Taking this multi-omic approach, a high-resolution tumor molecular phenotype, composed of the tumor's genetic, transcriptional, and epigenetic profile, can be determined.
[0005] Obtaining multi-omic signals typically requires performing multiple assays on the same input material. However, limited input material and the high cost and complexity of performing multiple assays can make multi-omic approaches prohibitive. In the case of liquid biopsy, the limited amount of cfDNA can limit the number of assays performed. Simultaneous identification of genetic and epigenetic signals can increase the sensitivity of cancer detection, for example, in liquid biopsy or other samples. In diagnostic settings, simplified workflows are needed to maximize information output.
[0006] Traditional approaches used to detect cytosine methylation leave genomic variant calling at a disadvantage. For example, in standard methylation detection methods, chemical or enzymatic treatment of DNA converts most of the cytosine bases to uracil bases, making cytosine variant calling error-prone. In this process, only methylated cytosine bases are protected from change. Thus, researchers must decide whether to evaluate patients' gene sequence variants or methylation variants as primary biomarkers.
[0007] Due to the importance of analyzing cancer-associated alterations in both epigenetic and genetic alterations, there remains a need for improved methods and systems that provide integrated and efficient analysis of both genetic (e.g., sequence alterations) and epigenetic (e.g., methylation) information in a single workflow.
[0008] All references cited herein, including patent applications and publications, are incorporated by reference in their entirety. Summary of the Invention
[0009] The present disclosure provides, inter alia, methods for detecting genetic and epigenetic sequence information in a single workflow, which are based at least in part on the data disclosed herein demonstrating methods providing simultaneous base-level detection of methylation and genetic variants in a single workflow, which may be used, for example, in detecting sequence variants and / or methylation variants, and in the detection, monitoring, screening, diagnosis, and / or prognosis of cancer, or response to cancer therapy.
[0010] This disclosure describes various techniques that allow for the simultaneous detection of both genetic and epigenetic changes in a single workflow. In standard methylation methods, chemical or enzymatic treatment of DNA converts the majority of cytosines to uracil, making cytosine variant calling error-prone. Only methylated cytosines are protected from conversion. In the disclosed methods, a conversion-resistant copy of the original molecule is created to preserve the genetic information. This is achieved, for example, by using primer extension to copy the original DNA molecule using 5-methylcytosine (5mC) or another cytosine analogue that is resistant to the particular conversion chemistry used. The original DNA molecule maintains the methylation information, and the protected strand maintains the genetic information and preserves the possibility to call genetic variants. Other capabilities include the possibility to use NGS index sequences to identify not only the sample but also the strand type (methyl or gene). In addition, methyl and genetic information can be paired to understand multi-omic signatures at the single molecule level.
[0011] In one aspect, a method for detecting genetic information and epigenetic information in a single workflow is provided herein. In some embodiments, the method includes: providing a plurality of first single-stranded DNA fragments; subjecting the plurality of first single-stranded DNA fragments to a round of primer extension in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragments, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides that contain cytosine analogs that are resistant to cytosine conversion, thereby generating a plurality of first strands corresponding to the first single-stranded DNA fragments and a plurality of second strands that are complementary to the first strands, the second strands containing cytosine analogs that are resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosine is converted if present in the first strand; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands. In some embodiments, a method includes providing a plurality of first single-stranded DNA fragments; subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragments, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides comprising a cytosine analog that is resistant to TET-assisted pyridine borane treatment, thereby generating a plurality of first strands corresponding to the first single-stranded DNA fragments and a plurality of second strands complementary to the first strands, wherein the second strands comprise a cytosine analog that is resistant to TET-assisted pyridine borane treatment; subjecting the plurality of first and second strands to TET-assisted pyridine borane treatment under conditions such that any methylated cytosines present in the first strand undergo cytosine conversion; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.
[0012] In some embodiments, the method further comprises subjecting the plurality of first strands and / or the plurality of second strands to amplification prior to detecting, and detecting comprises detecting at least a portion of the plurality of first and second strands and their amplification products, if present. In some embodiments, the method comprises subjecting the plurality of second strands to amplification in the presence of one or more primers that anneal to at least a portion of the plurality of second strands but not to the plurality of first strands. In some embodiments, the method includes subjecting the plurality of first strands to amplification in the presence of one or more primers that anneal to at least a portion of the plurality of first strands but not to the plurality of second strands. In some embodiments, one or more primers anneal to at least a portion of the plurality of first strands after the first strands undergo cytosine conversion, but not before the first strands undergo cytosine conversion.In some embodiments, one or more primers anneal to at least a portion of the plurality of first strands only if the cytosine of the first strand does not undergo cytosine conversion.In some embodiments, amplification occurs after cytosine conversion.
[0013] In some embodiments, the method further comprises enriching the plurality of second strands or their amplification products prior to detection. In some embodiments, the method further comprises enriching the plurality of first strands or their amplification products prior to detection. In some embodiments, the method comprises separating one or more first strands or their amplification products from one or more second strands or their amplification products prior to detection. In some embodiments, the separating comprises (a) combining one or more bait molecules with the plurality of first and second strands, where the one or more bait molecules preferentially hybridize to one or more of the first strands or their amplification products, or to one or more of the second strands or their amplification products, thereby producing a nucleic acid hybrid, and (b) isolating the nucleic acid hybrid. In some embodiments, the one or more bait molecules hybridize to one or more of the first strands or their amplification products but not to the plurality of second strands or their amplification products. In some embodiments, the separation occurs after cytosine conversion treatment, and the one or more bait molecules hybridize to one or more of the first strands or their amplification products if one or more cytosines of the first strand undergo cytosine conversion. In some embodiments, the one or more bait molecules hybridize to one or more of the second strands or their amplification products, but not to the plurality of first strands or their amplification products. In some embodiments, the separating comprises: (a) combining one or more first bait molecules with a plurality of first and second strands, where the one or more first bait molecules preferentially hybridize to one or more of the first strands or their amplification products, thereby producing a first nucleic acid hybrid; isolating the first nucleic acid hybrid; combining one or more second bait molecules with the plurality of first and second strands, where the one or more second bait molecules preferentially hybridize to one or more of the second strands or their amplification products, thereby producing a second nucleic acid hybrid; and isolating the second nucleic acid hybrid. In some embodiments, the plurality of first and second strands or their amplification products are detected together.In some embodiments, the plurality of first and second strands or their amplification products are detected separately. In some embodiments, the method comprises subjecting the plurality of first single-stranded DNA fragments to two or more rounds of primer extension prior to cytosine conversion in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragments, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides that contain cytosine analogs that are resistant to cytosine conversion. In some embodiments, the method further comprises attaching one or more nucleic acid adaptors to one or more of the first single-stranded DNA fragments. In some embodiments, the method further comprises attaching one or more nucleic acid adaptors to one or more of the second strands. In some embodiments, the one or more nucleic acid adaptors are attached to the first or second strand by ligation, translocation, tailing, or template switching.
[0014] In one aspect, provided herein is a method for detecting genetic and epigenetic information in a single workflow, the method comprising providing a plurality of first single-stranded DNA fragments; attaching a first adaptor nucleic acid to at least a portion of the plurality of first single-stranded DNA fragments; and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragment downstream of or at the first adaptor nucleic acid, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides that include a cytosine analog that is resistant to cytosine conversion, thereby detecting the first adaptor nucleic acid. generating a plurality of first strands comprising a first single-stranded DNA fragment having a target nucleic acid, and a second single-stranded DNA fragment complementary to the first single-stranded DNA fragment, and a plurality of second strands comprising a second adaptor nucleic acid complementary to the first adaptor nucleic acid, wherein the second strand comprises a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosines present in the first strand undergo cytosine conversion, wherein after cytosine conversion, the sequences of the first and second adaptor nucleic acids are no longer complementary; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments; attaching a first adapter nucleic acid to at least a portion of the plurality of first single-stranded DNA fragments; and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragment downstream of or at the first adapter nucleic acid, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides that include a cytosine analog that is resistant to TET-assisted pyridine borane treatment, thereby generating a plurality of first single-stranded DNA fragments comprising the first single-stranded DNA fragments having the first adapter nucleic acid. generating a plurality of second strands comprising a first strand, a second single-stranded DNA fragment complementary to the first single-stranded DNA fragment, and a second adapter nucleic acid complementary to the first adapter nucleic acid, wherein the second strand comprises a cytosine analog that is resistant to TET-assisted pyridine borane treatment; subjecting the plurality of first and second strands to TET-assisted pyridine borane treatment under conditions such that any methylated cytosines present in the first strand undergo cytosine conversion, where after cytosine conversion the sequences of the first and second adapter nucleic acids are no longer complementary; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.
[0015] In some embodiments, the first adaptor nucleic acid is attached to the 3' end of at least a portion of the plurality of first single-stranded DNA fragments. In some embodiments, the first adaptor nucleic acid is attached to the 5' end of at least a portion of the plurality of first single-stranded DNA fragments. In some embodiments, the first adaptor nucleic acid comprises one or more unmethylated cytosines that are converted during cytosine conversion. In some embodiments, the first adaptor nucleic acid comprises one or more methylated cytosines and the primer comprises one or more unmethylated cytosines that are converted during cytosine conversion.
[0016] In one aspect, provided herein is a method for detecting genetic and epigenetic information in a single workflow, the method comprising: providing a plurality of first single-stranded DNA fragments; and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer, the primer comprising: (i) a second adaptor nucleic acid portion that does not anneal to the first single-stranded DNA fragments; and (ii) a portion that anneals to at least a portion of the first single-stranded DNA fragments; (b) a nucleic acid polymerase; and (c) a mixture of nucleotides comprising a cytosine analog that is resistant to cytosine conversion, thereby generating a first adaptor nucleic acid complementary to the second adaptor nucleic acid. generating a plurality of first strands comprising a first single-stranded DNA fragment having a nucleic acid, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragment, and a second adaptor nucleic acid non-complementary to the first adaptor nucleic acid, the second strand comprising a cytosine analog resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosines present in the first strand undergo cytosine conversion, wherein after cytosine conversion, the sequences of the first and second adaptor nucleic acids are no longer complementary; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer, the primer comprising: (i) a second adapter nucleic acid portion that does not anneal to the first single-stranded DNA fragment; and (ii) a portion that anneals to at least a portion of the first single-stranded DNA fragment, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides comprising a cytosine analog that is resistant to TET-assisted pyridine borane treatment, thereby generating a plurality of first single-stranded DNA fragments comprising first single-stranded DNA fragments having a first adapter nucleic acid complementary to the second adapter nucleic acid. generating a plurality of second strands comprising a first single-stranded DNA fragment, a second single-stranded DNA fragment complementary to the first single-stranded DNA fragment, and a second adapter nucleic acid non-complementary to the first adapter nucleic acid, wherein the second strand comprises a cytosine analog that is resistant to TET-assisted pyridine borane treatment; subjecting the plurality of first and second strands to TET-assisted pyridine borane treatment under conditions such that any methylated cytosines present in the first strand undergo cytosine conversion, where after cytosine conversion the sequences of the first and second adapter nucleic acids are no longer complementary; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.
[0017] In some embodiments, the second adapter nucleic acid portion of the primer is 5' to the portion of the primer that anneals to the first single-stranded DNA fragment. In some embodiments, the primer contains one or more unmethylated cytosines that become part of the second adapter nucleic acid after primer extension and are converted during cytosine conversion.
[0018] In one aspect, provided herein is a method for detecting genetic and epigenetic information in a single workflow, the method comprising: providing a plurality of first single-stranded DNA fragments; and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer, the primer comprising: (i) a second adapter nucleic acid portion comprising a non-complementary 5' overhang that does not anneal to the first single-stranded DNA fragment; and (ii) a portion that anneals to at least a portion of the first single-stranded DNA fragment; (b) a nucleic acid polymerase; and (c) a mixture of nucleotides comprising a cytosine analog that is resistant to cytosine conversion; thereby generating a plurality of first strands comprising a first single-stranded DNA fragment having a first adapter nucleic acid complementary to a portion of a second adapter nucleic acid portion annealed to the first single-stranded DNA fragment, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragment and a second adapter nucleic acid, the second strand comprising a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosine undergoes cytosine conversion if present in the first strand; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.In one aspect, provided herein is a method for detecting genetic and epigenetic information in a single workflow, the method comprising: providing a plurality of first single-stranded DNA fragments; and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer, the primer comprising: (i) a second adapter nucleic acid portion comprising a non-complementary 5′ overhang that does not anneal to the first single-stranded DNA fragment; and (ii) a portion that anneals to at least a portion of the first single-stranded DNA fragment; (b) a nucleic acid polymerase; and (c) a mixture of nucleotides comprising a cytosine analog that is resistant to TET-assisted pyridine borane treatment, thereby detecting the plurality of first single-stranded DNA fragments. to generate a plurality of first strands comprising first single-stranded DNA fragments having a first adapter nucleic acid complementary to a portion of a second adapter nucleic acid portion annealed to the first single-stranded DNA fragment, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragment and a second adapter nucleic acid, the second strands comprising a cytosine analog that is resistant to TET-assisted pyridine borane treatment; subjecting the plurality of first and second strands to TET-assisted pyridine borane treatment under conditions such that any methylated cytosines present in the first strand undergo cytosine conversion; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.
[0019] In some embodiments, the method further comprises demultiplexing the sequence information from the first and second strands based on the first and / or second adaptor nucleic acid. In some embodiments, the method further comprises subjecting the plurality of first strands and / or the plurality of second strands to amplification prior to detecting, and detecting comprises detecting at least a portion of the plurality of first and second strands and their amplification products, if present. In some embodiments, the method comprises subjecting the plurality of second strands to amplification in the presence of one or more primers that anneal to at least a portion of the plurality of second strands but not to the plurality of first strands. In some embodiments, the one or more primers that anneal to at least a portion of the plurality of second strands anneal to at least a portion of the second adaptor nucleic acid. In some embodiments, the method comprises subjecting the plurality of first strands to amplification in the presence of one or more primers that anneal to at least a portion of the plurality of first strands but not to the plurality of second strands. In some embodiments, the one or more primers that anneal to at least a portion of the plurality of first strands anneal to at least a portion of the first adaptor nucleic acid. In some embodiments, the one or more primers anneal to at least a portion of the plurality of first strands after the first strands undergo cytosine conversion, but not before the first strands undergo cytosine conversion. In some embodiments, the one or more primers anneal to at least a portion of the plurality of first strands only if the cytosines of the first strands do not undergo cytosine conversion. In some embodiments, the method further comprises enriching the plurality of second strands or their amplification products prior to detection. In some embodiments, the method further comprises enriching the plurality of first strands or their amplification products prior to detection. In some embodiments, the method comprises separating the one or more first strands or their amplification products from the one or more second strands or their amplification products prior to detection.In some embodiments, the separation includes (a) combining one or more bait molecules with a plurality of first and second strands, where the one or more bait molecules preferentially hybridize to one or more of the first strands or their amplification products, or to one or more of the second strands or their amplification products, thereby producing a nucleic acid hybrid, and (b) isolating the nucleic acid hybrid. In some embodiments, the one or more bait molecules hybridize to one or more of the first strands or their amplification products, but not to the plurality of second strands or their amplification products. In some embodiments, the one or more bait molecules hybridize to at least a portion of the first adaptor nucleic acid. In some embodiments, the separation occurs after cytosine conversion treatment, where the one or more bait molecules hybridize to one or more of the first strands or their amplification products if one or more cytosines of the first strand undergo cytosine conversion. In some embodiments, the one or more bait molecules hybridize to one or more of the second strands or their amplification products, but not to the plurality of first strands or their amplification products. In some embodiments, the one or more bait molecules hybridize to at least a portion of the second adaptor nucleic acid. In some embodiments, the separating comprises combining one or more first bait molecules with the plurality of first and second strands, where the one or more first bait molecules preferentially hybridize to one or more of the first strands or their amplification products, thereby producing a first nucleic acid hybrid, isolating the first nucleic acid hybrid, combining one or more second bait molecules with the plurality of first and second strands, where the one or more second bait molecules preferentially hybridize to one or more of the second strands or their amplification products, thereby producing a second nucleic acid hybrid, and isolating the second nucleic acid hybrid. In some embodiments, one or more first bait molecules hybridize to at least a portion of a first adaptor nucleic acid, and / or one or more second bait molecules hybridize to at least a portion of a second adaptor nucleic acid.In some embodiments, the plurality of first and second strands or their amplification products are detected together. In some embodiments, the plurality of first and second strands or their amplification products are detected separately. In some embodiments, the method includes subjecting the plurality of first single-stranded DNA fragments to two or more rounds of primer extension prior to cytosine conversion in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragments, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides that include a cytosine analog that is resistant to cytosine conversion.
[0020] In some embodiments according to any of the embodiments described herein, the detecting is by sequencing (e.g., NGS), microarray, quantitative PCR (qPCR), digital droplet PCR (ddPCR), or molecular inversion probe. In some embodiments, the plurality of first strands are sequenced at a different sequencing depth than the plurality of second strands. In some embodiments, the plurality of second strands are sequenced at a sequencing depth that is at least 2 times, at least 4 times, at least 8 times, at least 10 times, at least 50 times, at least 100 times, or at least 1000 times higher than the sequencing depth of the plurality of first strands. In some embodiments, the plurality of first strands are sequenced at a sequencing depth that is at least 2 times, at least 4 times, at least 8 times, at least 10 times, at least 50 times, at least 100 times, or at least 1000 times higher than the sequencing depth of the plurality of second strands. In some embodiments, after primer extension, the plurality of first strands are hybridized with the plurality of second strands in the plurality of double-stranded nucleic acids. In some embodiments, the method further comprises denaturing the plurality of double-stranded DNA fragments to provide a plurality of single-stranded DNA fragments prior to the round of primer extension. In some embodiments, the method further comprises obtaining a plurality of single-stranded or double-stranded DNA fragments from the sample. In some embodiments, the method further comprises obtaining a sample from an individual. In some embodiments, the individual has cancer, is suspected of having cancer, or is undergoing treatment for cancer. In some embodiments, the individual is being screened for cancer or recurrence of cancer. In some embodiments, the sample comprises tissue, cells, and / or nucleic acid from cancer. In some embodiments, the sample comprises tissue, cells, and / or nucleic acid from normal tissue. In some embodiments, the sample comprises a tissue biopsy sample, a liquid biopsy sample, or a normal control. In some embodiments, the sample is from a tumor biopsy, a tumor specimen, or circulating tumor cells. In some embodiments, the sample is a liquid biopsy sample and comprises blood, plasma, serum, cerebrospinal fluid, sputum, stool, urine, or saliva.In some embodiments, cytosine analogs that are resistant to cytosine conversion include 5-methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), 5-carboxylcytosine (5caC), 5-formylcytosine, 5-(beta-D-glucosylmethyl)cytosine (5gmC), 5-ethyl dCTP, 5-methyl dCTP, 5-fluoro dCTP, 5-bromo dCTP, 5-iodo dCTP, 5-chloro dCTP, 5-trifluoromethyl dCTP, or 5-aza dCTP. In some embodiments, cytosine conversion is by bisulfite treatment, TET-assisted bisulfite treatment, oxidized bisulfite treatment, APOBEC, or TET / beta-glucosyltransferase-assisted APOBEC treatment. In some embodiments, if unmethylated cytosines are present in the first strand, at least 80%, at least 85%, at least 90%, at least 95%, or 100% undergo cytosine conversion as a result of the cytosine conversion treatment. In some embodiments, if unmethylated cytosines are present in the first strand, about 80% to about 97% undergo cytosine conversion as a result of the cytosine conversion treatment. In some embodiments, up to 20%, up to 15%, up to 10%, up to 5%, up to 2%, up to 1%, or up to 0.5% of the cytosine analogs in the second strand undergo cytosine conversion as a result of the cytosine conversion treatment. In some embodiments, about 0.5% to about 5% of the cytosine analogs in the second strand undergo cytosine conversion as a result of the cytosine conversion treatment. In some embodiments, the nucleic acid polymerase is capable of incorporating the cytosine analogs into the nucleic acid. In some embodiments, the method further comprises subjecting the plurality of first single-stranded DNA fragments to end repair prior to primer extension. In some embodiments, the method further comprises, after sequencing the plurality of first and second strands or their amplification products, comparing the plurality of second strand sequences to the plurality of first strand sequences. In some embodiments, the method further comprises, after sequencing the plurality of first and second strands or their amplification products, comparing the plurality of first and / or second strand sequences to a reference genome sequence.
[0021] In one aspect, provided herein is a method of detecting cancer in an individual, the method comprising detecting methylation levels and / or somatic mutations in a sample comprising a plurality of nucleic acids obtained from the individual according to the method of any one of the above embodiments, wherein the methylation levels and / or somatic mutations detected in the sample identify the individual as having cancer.
[0022] In one aspect, provided herein is a method of detecting minimal residual disease in an individual who has been treated for or is being treated for cancer, the method comprising detecting methylation levels and / or somatic mutations in a sample comprising a plurality of nucleic acids obtained from the individual according to the method of any one of the above embodiments, wherein the methylation levels and / or somatic mutations detected in the sample identify the individual as having or lack thereof minimal residual disease.
[0023] In one aspect, provided herein is a method of screening an individual suspected of having cancer, the method comprising detecting methylation levels and / or somatic mutations in a sample comprising a plurality of nucleic acids obtained from the individual according to the method of any one of the above embodiments, wherein the methylation levels and / or somatic mutations detected in the sample identify the individual as likely to have cancer.
[0024] In one aspect, provided herein is a method of determining a prognosis of an individual having cancer, the method comprising detecting methylation levels and / or somatic mutations in a sample comprising a plurality of nucleic acids obtained from the individual according to the method of any one of the above embodiments, wherein the methylation levels and / or somatic mutations detected in the sample at least partially determine the prognosis of the individual.
[0025] In one aspect, provided herein is a method of predicting survival of an individual having cancer, the method comprising detecting methylation levels and / or somatic mutations in a sample comprising a plurality of nucleic acids obtained from the individual according to the method of any one of the above embodiments, wherein the methylation levels and / or somatic mutations detected in the sample are at least partially predictive of survival of the individual.
[0026] In one aspect, provided herein is a method of predicting or detecting tumor burden in an individual having cancer, the method comprising detecting methylation levels and / or somatic mutations in a sample comprising a plurality of nucleic acids obtained from the individual according to the method of any one of the above embodiments, wherein the methylation levels and / or somatic mutations detected in the sample at least partially predict or detect the tumor burden of the individual.
[0027] In one aspect, provided herein is a method of predicting responsiveness of an individual having cancer to a treatment, the method comprising detecting methylation levels and / or somatic mutations in a sample comprising a plurality of nucleic acids obtained from the individual according to the method of any one of the above embodiments, wherein the methylation levels and / or somatic mutations detected in the sample are used at least in part to predict the individual's responsiveness to the treatment.
[0028] In one aspect, provided herein is a method of monitoring the response of an individual undergoing treatment for cancer, the method comprising administering the treatment to an individual having cancer and detecting methylation levels and / or somatic mutations in a sample obtained from the individual comprising a plurality of nucleic acids according to the method of any one of the above embodiments, wherein the methylation levels and / or somatic mutations detected in the sample are used at least in part to monitor the response to the treatment.
[0029] In one aspect, provided herein is a method of monitoring cancer in an individual, the method comprising: detecting methylation levels and / or somatic mutations in a first sample comprising a plurality of nucleic acids obtained from the individual according to the method of any one of the above embodiments; detecting methylation levels and / or somatic mutations in a second sample comprising a plurality of nucleic acids obtained from the individual according to the method of any one of the above embodiments; and determining a difference in methylation levels and / or somatic mutations between the first and second samples, thereby monitoring cancer in the individual.
[0030] In one aspect, provided herein is a system comprising one or more processors and a memory configured to store one or more computer program instructions, the one or more computer program instructions being configured to, when executed by the one or more processors, perform a method according to any one of the above embodiments. In one aspect, provided herein is a system comprising one or more processors and a memory configured to store one or more computer program instructions, which when executed by the one or more processors are configured to obtain a first plurality of sequence reads of one or more first nucleic acid molecules or their amplification products, where the first nucleic acid molecules have been subjected to a cytosine conversion treatment under conditions such that unmethylated cytosines, if present in the first nucleic acid molecule, are subjected to cytosine conversion; obtain a second plurality of sequence reads of one or more second nucleic acid molecules or their amplification products, where the second nucleic acid molecules are complementary to the first nucleic acid molecule prior to cytosine conversion and comprise a cytosine analog that is resistant to cytosine conversion; analyze the first plurality of sequence reads for the presence or absence of methylation at one or more cytosines inferred based on the cytosine conversion or lack thereof; and analyze the second plurality of sequence reads for sequence information.
[0031] In some embodiments, the one or more first nucleic acid molecules or their amplification products further comprise a first adaptor nucleic acid sequence, and the one or more computer program instructions, when executed by the one or more processors, are further configured to demultiplex the first and second plurality of sequence reads based at least in part on the detection of the first adaptor nucleic acid sequence in the first plurality of sequence reads. In some embodiments, the first adaptor nucleic acid is attached to a 3' end of at least a portion of the plurality of first nucleic acid molecules. In some embodiments, the first adaptor nucleic acid is attached to a 5' end of at least a portion of the plurality of first nucleic acid molecules. In some embodiments, the first adaptor nucleic acid comprises one or more unmethylated cytosines that are converted during cytosine conversion. In some embodiments, the first adaptor nucleic acid comprises one or more methylated cytosines and the primer comprises one or more unmethylated cytosines that are converted during cytosine conversion. In some embodiments, the first adaptor nucleic acid is between the 5' end and the 3' end of at least a portion of the plurality of first nucleic acid molecules. In some embodiments, the one or more second nucleic acid molecules or their amplification products further comprise a second adaptor nucleic acid sequence, and the one or more computer program instructions, when executed by the one or more processors, are further configured to demultiplex the first and second plurality of sequence reads based at least in part on detecting the second adaptor nucleic acid sequence in the second plurality of sequence reads. In some embodiments, the second adaptor nucleic acid is attached to a 3' end of at least a portion of the plurality of second nucleic acid molecules. In some embodiments, the second adaptor nucleic acid is attached to a 5' end of at least a portion of the plurality of second nucleic acid molecules. In some embodiments, the second adaptor nucleic acid is between the 5' end and the 3' end of at least a portion of the plurality of second nucleic acid molecules. In some embodiments, the first plurality of sequence reads are at least 2-fold, at least 4-fold, at least 8-fold, at least 10-fold, at least 50-fold, at least 100-fold, or at least 1000-fold more abundant than the second plurality of sequence reads.In some embodiments, the second plurality of sequence reads are at least 2-fold, at least 4-fold, at least 8-fold, at least 10-fold, at least 50-fold, at least 100-fold, or at least 1000-fold more abundant than the first plurality of sequence reads. In some embodiments, the first nucleic acid molecule is obtained from a sample prior to cytosine conversion. In some embodiments, the sample is from an individual having or suspected of having cancer. In some embodiments, the sample comprises tissue, cells, and / or nucleic acid from cancer. In some embodiments, the sample comprises tissue, cells, and / or nucleic acid from normal tissue. In some embodiments, the first and / or second plurality of sequence reads are obtained by sequencing, optionally, the sequencing comprises the use of massively parallel sequencing (MPS) technology, whole genome sequencing (WGS), whole exome sequencing, targeted sequencing, direct sequencing, or Sanger sequencing technology, and optionally, the massively parallel sequencing technology comprises next generation sequencing (NGS). In some embodiments, the one or more program instructions, when executed by the one or more processors, are further configured to generate a molecular profile for the sample based at least in part on the analysis. In some embodiments, the individual is administered a treatment based at least in part on the molecular profile. In some embodiments, the molecular profile further comprises results from a comprehensive genomic profiling (CGP) test, a gene expression profiling test, a cancer hotspot panel test, a DNA methylation test, a DNA fragmentation test, an RNA fragmentation test, or any combination thereof. In some embodiments, the molecular profile further comprises results from a nucleic acid sequencing-based test. In some embodiments, the one or more computer program instructions, when executed by the one or more processors, are further configured to compare a sequence read of the second plurality of sequence reads to a sequence read of the first plurality of sequence reads. In some embodiments, the one or more computer program instructions, when executed by the one or more processors, are further configured to compare a sequence read of the first and / or second plurality of sequence reads to a reference genome sequence.
[0032] In one aspect, provided herein is a non-transitory computer-readable storage medium comprising one or more programs executable by one or more computer processors for performing a method according to any one of the above embodiments. In one aspect, provided herein is a non-transitory computer readable storage medium comprising one or more programs executable by one or more computer processors for performing a method, the method comprising: obtaining a first plurality of sequence reads of one or more first nucleic acid molecules or amplification products thereof, where the first nucleic acid molecules have been subjected to a cytosine conversion treatment under conditions such that unmethylated cytosines, if present in the first nucleic acid molecule, undergo cytosine conversion; obtaining a second plurality of sequence reads of one or more second nucleic acid molecules or amplification products thereof, where the second nucleic acid molecules are complementary to the first nucleic acid molecule prior to cytosine conversion and comprise a cytosine analog that is resistant to cytosine conversion; analyzing the first plurality of sequence reads for the presence or absence of methylation at one or more cytosines inferred based on the cytosine conversion or lack thereof; and analyzing the second plurality of sequence reads for sequence information.
[0033] In some embodiments, the one or more first nucleic acid molecules or their amplification products further comprise a first adaptor nucleic acid sequence, and the method further comprises demultiplexing the first and second plurality of sequence reads based at least in part on detection of the first adaptor nucleic acid sequence in the first plurality of sequence reads. In some embodiments, the first adaptor nucleic acid is attached to a 3' end of at least a portion of the plurality of first nucleic acid molecules. In some embodiments, the first adaptor nucleic acid is attached to a 5' end of at least a portion of the plurality of first nucleic acid molecules. In some embodiments, the first adaptor nucleic acid comprises one or more unmethylated cytosines that are converted during cytosine conversion. In some embodiments, the first adaptor nucleic acid comprises one or more methylated cytosines and the primer comprises one or more unmethylated cytosines that are converted during cytosine conversion. In some embodiments, the one or more second nucleic acid molecules or their amplification products further comprise a second adaptor nucleic acid sequence, and the method further comprises demultiplexing the first and second plurality of sequence reads based at least in part on detection of the second adaptor nucleic acid sequence in the second plurality of sequence reads. In some embodiments, the second adaptor nucleic acid is attached to a 3' end of at least a portion of the plurality of second nucleic acid molecules. In some embodiments, the second adaptor nucleic acid is attached to a 5' end of at least a portion of the plurality of second nucleic acid molecules. In some embodiments, the first adaptor nucleic acid is between the 5' end and the 3' end of at least a portion of the plurality of first nucleic acid molecules. In some embodiments, the first plurality of sequence reads is at least 2-fold, at least 4-fold, at least 8-fold, at least 10-fold, at least 50-fold, at least 100-fold, or at least 1000-fold more abundant than the second plurality of sequence reads. In some embodiments, the second plurality of sequence reads are at least 2-fold, at least 4-fold, at least 8-fold, at least 10-fold, at least 50-fold, at least 100-fold, or at least 1000-fold more abundant than the first plurality of sequence reads. In some embodiments, the first nucleic acid molecule is obtained from a sample prior to cytosine conversion. In some embodiments, the sample is from an individual having or suspected of having cancer.In some embodiments, the sample comprises tissue, cells, and / or nucleic acid from a cancer. In some embodiments, the sample comprises tissue, cells, and / or nucleic acid from a normal tissue. In some embodiments, the first and / or second plurality of sequence reads are obtained by sequencing, optionally, the sequencing comprises using a massively parallel sequencing (MPS) technology, whole genome sequencing (WGS), whole exome sequencing, targeted sequencing, direct sequencing, or Sanger sequencing technology, and optionally, the massively parallel sequencing technology comprises next generation sequencing (NGS). In some embodiments, the method further comprises generating a molecular profile for the sample based at least in part on the analysis. In some embodiments, the individual is administered a treatment based at least in part on the molecular profile. In some embodiments, the molecular profile further comprises results from a comprehensive genomic profiling (CGP) test, a gene expression profiling test, a cancer hotspot panel test, a DNA methylation test, a DNA fragmentation test, an RNA fragmentation test, or any combination thereof. In some embodiments, the molecular profile further comprises results from a nucleic acid sequencing-based test. In some embodiments, the method further comprises comparing a sequence read of the second plurality of sequence reads to a sequence read of the first plurality of sequence reads, In some embodiments, the method further comprises comparing a sequence read of the first and / or second plurality of sequence reads to a reference genome sequence.
[0034] It should be understood that one, some, or all of the features of the various embodiments described herein may be combined to form other embodiments of the present invention. These and other aspects of the present invention will become apparent to those skilled in the art. These and other embodiments of the present invention are further described in the following detailed description. [Brief description of the drawings]
[0035] [Figure 1] 1 illustrates an exemplary single workflow assay for generating genomic and methylation strands, according to some embodiments. [Diagram 2]1 illustrates an exemplary strategy for labeling and demultiplexing genomic and methylated strands in a single workflow assay according to some embodiments. The sequences shown are CTGATCGTGGTT (SEQ ID NO: 3, top) and CTGATCGTGGCC (SEQ ID NO: 4, bottom). [Figure 3A] 1 illustrates an exemplary strategy of workflow tags for labeling and amplifying genomic and methylated strands in a single workflow assay, according to some embodiments. [Figure 3B] An exemplary strategy is illustrated in which the workstream tags contain unique sequences that distinguish the genomic strand from the methylated strand without requiring cytosine conversion. [Figure 3C] 1 illustrates an exemplary strategy for strand-specific amplification, according to some embodiments. [Figure 4A-4F]4A provides a flow chart of steps of an exemplary single workflow assay according to some embodiments. FIG. 4A illustrates an exemplary assay including end repair, adapter ligation, primer extension, cytosine conversion treatment, PCR amplification / library generation, and parallel sequencing of genomic and methylated strands. FIG. 4B illustrates an exemplary assay including end repair, adapter ligation, primer extension, cytosine conversion treatment, PCR amplification / library generation, hybrid capture to enrich genomic and methylated strands, and sequencing of genomic and methylated strands. FIG. 4C illustrates an exemplary assay including end repair, adapter ligation, primer extension, cytosine conversion treatment, and PCR amplification / library generation, followed by enrichment and sequencing of separate libraries of genomic and methylated strands. FIG. 4D illustrates an exemplary assay including end repair, adapter ligation, primer extension, cytosine conversion treatment, and PCR amplification / library generation, followed by enrichment of the genomic strand library, and sequencing of the genomic strand library and the methylated strand library. Figure 4E illustrates an exemplary assay including end repair, adapter ligation, primer extension, cytosine conversion treatment, and PCR amplification / library generation, followed by enrichment of the methylated strand library and sequencing of the genomic strand library and the methylated strand library. Figure 4F illustrates an exemplary assay including end repair, adapter ligation, primer extension, cytosine conversion treatment, PCR amplification / library generation, hybrid capture to enrich the genomic strand and the methylated strand, and generation and sequencing of separate targeted libraries based on the genomic strand and the methylated strand. [Figure 5A-5C] We show a proof of concept of analyzing genetic (e.g., sequence) and epigenetic (e.g., methylation) information in a single workflow. An overview of the workflow used is provided in Figure 5A. Figure 5B shows the percentage of sequence reads identified from the genomic and methylated strands and their amplification products. Figure 5C shows the results of cytosine conversion, where the genomic strand was largely preserved after cytosine conversion, while the methylated strand was efficiently converted. [Figure 6] Using a single workflow method, we show the observed protection efficiency (from cytosine conversion) in the genomic and methylated strands. [Figure 7] The abundance of reads identified on the genomic and methylated strands after primer extension using a single workflow method is shown. [Figure 8] Correlation between cancer methylation scores (assessing consensus methylation sites from individual DNA) obtained using a single workflow and standard whole genome (WG) enzymatic methylation sequencing methods. [Figure 9A-9E] Correlation of methylation levels observed using a single workflow or standard WG methylation methodology is shown. Methylation levels from smaller bins and functional regions were analyzed, including 1 kb bins (Figure 9A), 10 kb bins (Figure 9B), 100 kb bins (Figure 9C), CpG islands (Figure 9D), and CpG shores (Figure 9E). [Figure 10] Figure 1 shows the Pearson correlation coefficient of the average methylation fraction (AMF) in a single workflow compared to standard enzymatic methylation sequencing. The average (r) obtained from each bin / functional region is shown. [Figure 11] Correlation of allele frequencies of major genetic variants observed using a single workflow or standard whole genome sequencing (WGS) methodology. [Figure 12] The percentage of reads identified on the genomic and methylated strands is shown with and without preferential amplification of the genomic strand using strand-specific primers. [Figure 13] The percentage of reads identified on the genomic and methylated strands is shown with and without preferential amplification of the genomic and / or methylated strand using strand-specific primers. [Figure 14A]Average specific coverage after multi-omic hybrid capture is shown. Biotinylated capture baits targeting genomic and methylation biomarkers were combined for simultaneous capture of a single workflow library. High specific coverage was observed for genomic and methylation targets after hybrid capture. [Figure 14B] We show high coverage uniformity across both methylated and genomic capture regions as measured by Fold-80 base penalty. There is high on-target rate for both genomic and methylated regions in combined hybrid capture, indicating efficient multi-omic strand enrichment. [Figure 14C] 1 depicts a block diagram of an exemplary process for detecting methylation sequence information and genomic sequence information in a single workflow, according to some embodiments. [Figure 15] 1 depicts an exemplary system according to some embodiments. [Figure 16] 1 depicts an exemplary device according to some embodiments. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0036] The present disclosure relates generally to detecting sequence and methylation information from nucleic acids in a single, integrated workflow.
[0037] The present disclosure shows a method that provides efficient, comprehensive, and integrated analysis of sequence variants and methylation variants in a single workflow. The present disclosure shows that both types of information can be obtained efficiently from a single sample using small amounts of input DNA, and thus can be useful for a variety of applications, including the analysis of cancer-related cfDNA. These methods generate genomic strands that preserve sequence information and are resistant to cytosine conversion, and methylated strands that preserve methylation information and are susceptible to cytosine conversion. Both methylation levels and sequence variant detection from the single workflow method were found to correlate with results obtained using dedicated whole genome analysis of either sequence or methylation levels using existing methods. Advantageously, the single workflow method can be adapted to amplify and / or enrich genomic strands and / or methylated strands depending on the preference or focus of the assay.
[0038] I. General Techniques The techniques and procedures described or referenced herein may generally be found in, for example, Sambrook et al., Molecular Cloning: A Laboratory Manual 3d edition (2001) Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, Current Protocols in Molecular Biology (FMA Usubel, et al. eds., (2003)), the series Methods in Enzymology (Academic Press, Inc.): PCR 2: A Practical Approach (MJ MacPherson, B.D. Hames and G.R. Taylor eds. (1995)), Harlow and Lane, eds. (1988) Antibodies, A Laboratory Manual, and Animal Cell Culture (RI Freshney, ed. (1987)), Oligonucleotide Synthesis (MJ Gait, ed., 1984), Methods in Molecular Biology, Humana Press, Cell Biology: A Laboratory Manual, ... Notebook (JECellis, ed., 1998) Academic Press, Animal Cell Culture (RIFreshney), ed., 1987), Introduction to Cell and Tissue Culture (JP Mather and PE Roberts, 1998) Plenum Press, Cell and Tissue Culture: Laboratory Procedures (A. Doyle, JBGriffiths, and DG Newell, eds., 1993-8) J. Wiley and Sons, Handbook of Experimental Immunology (DMWeir and CCBlackwell, eds.), Gene Transfer Vectors for Mammalian Cells (JMMiller and MPCalos, eds., 1987), PCR: The Polymerase Chain Reaction, (Mullis et al., eds., 1994), Current Protocols in Immunology (JEColigan et al., eds., 1991), Short Protocols in Molecular Biology (Wiley and Sons, 1999), Immunobiology (CA Janeway and P. Travers, 1997), Antibodies (P. Finch, 1997), Antibodies: A Practical Approach (D. Catty., ed., IRL Press, 1988-1989), Monoclonal Antibodies: A Practical Approach (P. Shepherd and C. Dean, eds., Oxford University Press, 2000), Using Antibodies: A Laboratory Manual (E. Harlow and D. Lane (Cold Spring Harbor Laboratory Press, 1999), The These methods are well understood and commonly used by those skilled in the art using routine methodology, such as the widely used methodology described in Therapeutic Antibodies (M. Zanetti and JD Capra, eds., Harwood Academic Publishers, 1995), and Cancer: Principles and Practice of Oncology (VT DeVita et al., eds., J.B. Lippincott Company, 1993).
[0039] II. Definition As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. Thus, for example, reference to "a molecule" optionally includes a combination of two or more such molecules, and so forth.
[0040] The term "about" as used herein refers to a normal error range for the respective value, which is readily known to a person skilled in the art. Reference herein to "about" a value or parameter includes (and describes) embodiments that are directed to the value or parameter itself.
[0041] It is understood that aspects and embodiments of the invention described herein include "comprising," "consisting of," and / or "consisting essentially of" aspects and embodiments.
[0042] The terms "cancer" and "cancerous" refer to or describe the physiological condition in mammals that is typically characterized by unregulated cell growth. This definition includes benign and malignant cancers.
[0043] The term "tumor," as used herein, refers to all neoplastic cell growth and proliferation, whether malignant or benign, and all pre-cancerous and cancerous cells and tissues. The terms "cancer," "cancerous," and "tumor" are not mutually exclusive when referred to herein.
[0044] "Polynucleotide" or "nucleic acid", as used interchangeably herein, refers to a polymer of nucleotides of any length, including DNA and RNA. The nucleotides can be deoxyribonucleotides, ribonucleotides, modified nucleotides or bases, and / or their analogs, or any substrate that can be incorporated into a polymer by a DNA or RNA polymerase or by a synthetic reaction. Thus, for example, polynucleotides as defined herein include, but are not limited to, single-stranded and double-stranded DNA, DNA including single-stranded and double-stranded regions, single-stranded and double-stranded RNA, RNA including single-stranded and double-stranded regions, hybrid molecules including DNA and RNA that may be single-stranded or may typically be double-stranded or include single-stranded and double-stranded regions. In addition, the term "polynucleotide" as used herein refers to triple-stranded regions that include RNA or DNA, or both RNA and DNA. The strands in such regions may be from the same molecule or from different molecules. A region may include all of one or more of the molecules, but more typically involves only some regions of the molecule. One of the molecules of the triple helix region is often an oligonucleotide. The term "polynucleotide" specifically includes cDNA.
[0045] A polynucleotide may comprise modified nucleotides, such as methylated nucleotides and their analogs. Modifications to the nucleotide structure, if present, may be imparted before or after assembly of the polymer. The sequence of nucleotides may be interrupted by non-nucleotide components. A polynucleotide may be further modified after synthesis, such as by conjugation with a label. Other types of modifications include, for example, substitution of one or more of the naturally occurring nucleotides with "caps", analogs, internucleotide modifications, such as those with uncharged linkages (e.g., methylphosphonates, phosphotriesters, phosphoamidates, carbamates, etc.) and those with charged linkages (e.g., phosphorothioates, phosphorodithioates, etc.), those with pendant moieties, such as proteins (e.g., nucleases, toxins, antibodies, signal peptides, poly-L-lysine, etc.), those with intercalators (e.g., acridine, psoralen, etc.), those containing chelators (e.g., metals, radioactive metals, boron, metal oxides, etc.), those containing alkylators, those with modified linkages (e.g., alpha anomeric nucleic acids), as well as unmodified forms of polynucleotides. Additionally, any of the hydroxyl groups normally present in the sugar may be replaced, for example, by phosphonate groups, phosphate groups, protected by standard protecting groups, or activated to prepare additional linkages to additional nucleotides, or conjugated to solid or semi-solid supports. The 5' and 3' terminal OH may be phosphorylated or substituted with amines or organic capping group moieties of 1-20 carbon atoms. Other hydroxyls may also be derivatized to standard protecting groups. Polynucleotides may also contain analogous forms of ribose or deoxyribose sugars that are commonly known in the art, including, for example, 2'-0-methyl-, 2'-0-allyl-, 2'-fluoro-, or 2'-azido-ribose, carbocyclic sugar analogs, a-anomeric sugars, epimeric sugars such as arabinose, xylose or lyxose, pyranose sugars, furanose sugars, sedoheptulose, acyclic analogs, and abasic nucleoside analogs such as methyl riboside.One or more phosphodiester linkages may be replaced by alternative linking groups. These alternative linking groups include, but are not limited to, embodiments in which phosphate is replaced by P(0)S ("thioate"), P(S)S ("dithioate"), "(O)NR2 ("amidate"), P(0)R, P(0)OR', CO or CH2 ("formacetal"), where each R or R' is independently H or a substituted or unsubstituted alkyl (1-20C) optionally containing an ether (-0-) linkage, aryl, alkenyl, cycloalkyl, cycloalkenyl, or araldyl. Not all linkages in a polynucleotide need be identical. A polynucleotide may contain one or more different types of modifications and / or multiple modifications of the same type described herein. The preceding description applies to all polynucleotides referred to herein, including RNA and DNA.
[0046] "Oligonucleotide", as used herein, generally refers to a short, single-stranded polynucleotide, not necessarily less than about 250 nucleotides in length. Oligonucleotides may be synthetic. The terms "oligonucleotide" and "polynucleotide" are not mutually exclusive. The above description of polynucleotides is equally and fully applicable to oligonucleotides.
[0047] The term "detection" includes any means of detection, including direct detection and indirect detection.
[0048] "Amplification," as used herein, generally refers to the process of generating multiple copies of a desired sequence. "Multiple copies" means at least two copies. "Copy" does not necessarily mean perfect sequence complementarity or identity to the template sequence. For example, copies may contain nucleotide analogs, such as cytosine analogs that are resistant to cytosine conversion, deliberate sequence changes (e.g., sequence changes introduced by primers that contain sequences that are hybridizable to the template but are not complementary to the template), and / or sequence errors that occur during amplification.
[0049] The technique of "polymerase chain reaction" or "PCR" as used herein generally refers to a procedure in which minute quantities of specific pieces of nucleic acid, RNA, and / or DNA are amplified, for example, as described in U.S. Pat. No. 4,683,195. Generally, sequence information from or beyond the ends of the region of interest must be available so that oligonucleotide primers can be designed that are identical or similar in sequence to opposite strands of the template to be amplified. The 5' terminal nucleotides of the two primers may coincide with the ends of the amplified material. PCR can be used to amplify specific RNA sequences, specific DNA sequences from total genomic DNA, cDNA transcribed from total cellular RNA, bacteriophage, or plasmid sequences, etc. See generally Mullis et al., Cold Spring Harbor Symp. Quant. Biol. 51:263 (1987), and Erlich, ed., PCR Technology (Stockton Press, NY, 1989). As used herein, PCR is considered to be one example, but not the only example, of a nucleic acid polymerase reaction method for amplifying a nucleic acid test sample that involves the use of known nucleic acids (DNA or RNA) as primers and a nucleic acid polymerase to amplify or generate a specific piece of nucleic acid, or a specific piece of nucleic acid that is complementary to a specific nucleic acid.
[0050] The term "diagnosis" is used herein to refer to the identification or classification of a molecular or pathological state, disease, or condition (e.g., cancer). For example, "diagnosis" can refer to the identification of a particular type of cancer. "Diagnosis" can also refer to the classification of a particular subtype of cancer, for example, by histopathological criteria or by molecular features (e.g., subtypes characterized by expression of one or a combination of biomarkers (e.g., particular genes or proteins encoded by such genes), or aberrant DNA methylation levels and / or patterns).
[0051] The term "aiding in diagnosis" is used herein to refer to a method of aiding in making a clinical decision regarding the presence of, or the nature of, a particular type of symptom or condition of a disease or disorder (e.g., cancer). For example, a method of aiding in the diagnosis of a disease or condition (e.g., cancer) may include measuring specific somatic mutations or DNA methylation levels and / or patterns in a biological sample from an individual.
[0052] The term "sample" as used herein refers to a composition obtained or derived from a subject and / or individual of interest that contains cells and / or other molecular entities to be characterized and / or identified, for example, based on physical, biochemical, chemical, and / or physiological characteristics. For example, the phrase "disease sample" and variations thereof refer to any sample obtained from a subject of interest that is expected to contain or is known to contain the cells and / or molecular entities to be characterized. Samples include, but are not limited to, tissue samples, primary or cultured cells or cell lines, cell supernatants, cell lysates, platelets, serum, plasma, vitreous fluid, lymphatic fluid, synovial fluid, follicular fluid, semen, amniotic fluid, milk, whole blood, plasma, serum, cells from blood, urine, cerebrospinal fluid, saliva, sputum, tears, sweat, mucus, tumor lysates, and tissue culture media, tissue extracts, e.g., homogenized tissue, tumor tissue, cell extracts, and combinations thereof. In some cases, the sample is a whole blood sample, a plasma sample, a serum sample, or a combination thereof. In some embodiments, the sample is from a tumor (e.g., a "tumor sample"), e.g., from a biopsy. In some embodiments, the sample is a formalin-fixed paraffin-embedded (FFPE) sample.
[0053] As used herein, "tumor cell" refers to any tumor cell present in a tumor or a sample thereof. Tumor cells can be distinguished from other cells that may be present in a tumor sample (e.g., stromal cells and tumor-infiltrating immune cells) using methods known in the art and / or described herein.
[0054] A "reference sample," "reference cell," "reference tissue," "control sample," "control cell," or "control tissue," as used herein, refers to a sample, cell, tissue, standard, or level used for comparison purposes.
[0055] "Correlate" or "correlating" refers, in any case, to comparing the performance and / or results of a first analysis or protocol to the performance and / or results of a second analysis or protocol. For example, the results of the first analysis or protocol may be used in carrying out the second protocol and / or to determine whether the second analysis or protocol should be performed. With respect to embodiments of a polypeptide analysis or protocol, the results of a polypeptide expression analysis or protocol may be used to determine whether a particular therapeutic regimen should be performed. With respect to embodiments of a polynucleotide analysis or protocol, the results of a polynucleotide expression analysis or protocol may be used to determine whether a particular therapeutic regimen should be performed.
[0056] "Individual response" or "response" may be assessed using any endpoint that indicates benefit to an individual, including, but not limited to, (1) inhibition to some extent of disease progression (e.g., cancer progression), including slowing or completely halting; (2) reduction in tumor size; (3) inhibition (i.e., reducing, slowing, or completely halting) of cancer cell invasion into adjacent surrounding organs and / or tissues; (4) inhibition (i.e., reducing, slowing, or completely halting) of metastasis; (5) alleviation to some extent of one or more symptoms associated with a disease or disorder (e.g., cancer); (6) increased or prolonged length of survival, including overall survival and progression-free survival; and / or (7) reduction in mortality at a given time point following treatment.
[0057] An "effective response" of a patient or "responsiveness" of a patient to treatment with a pharmaceutical agent, and similar phrases, refers to a clinical or therapeutic benefit conferred on a patient at risk for or suffering from a disease or disorder, such as cancer. In one embodiment, such benefit includes extending survival (including overall survival and / or progression-free survival), obtaining an objective response (including a complete or partial response), or ameliorating the signs or symptoms of cancer.
[0058] "Effective amount" refers to the amount of a therapeutic agent to treat or prevent a disease or disorder in a mammal. In the case of cancer, a therapeutically effective amount of a therapeutic agent may reduce the number of cancer cells, reduce the size of a primary tumor, inhibit (i.e., slow to some extent, and in some embodiments, stop) cancer cell invasion into surrounding organs, inhibit (i.e., slow to some extent, and in some embodiments, stop) tumor metastasis, inhibit tumor growth to some extent, and / or alleviate to some extent one or more of the symptoms associated with the disorder. To the extent that the drug may prevent the growth and / or kill existing cancer cells, the drug may be cytostatic and / or cytotoxic. In the case of cancer therapy, in vivo efficacy may be measured, for example, by assessing the duration of survival, the time to disease progression (TTP), response rate (e.g., CR or PR), duration of response, and / or quality of life.
[0059] The term "pharmaceutical formulation" refers to a preparation that is in a form that allows the biological activity of the active ingredients contained therein to be effective and does not contain additional components that are unacceptably toxic to the subject to which the formulation is administered.
[0060] "Pharmaceutically acceptable carrier" refers to an ingredient in a pharmaceutical formulation, other than an active ingredient, that is non-toxic to a subject. Pharmaceutically acceptable carriers include, but are not limited to, buffers, excipients, stabilizers, or preservatives.
[0061] As used herein, "treatment" (and grammatical variations thereof, such as "treat" or "treating") refers to a clinical intervention that seeks to alter the natural course of the individual being treated, and may be performed either prophylactically or during the course of clinical pathology. Desirable effects of treatment include, but are not limited to, preventing the onset or recurrence of disease, alleviating symptoms, reducing any direct or indirect pathological consequences of the disease, preventing metastasis, reducing the rate of disease progression, ameliorating or alleviating the disease state, and remission or improved prognosis.
[0062] As used herein, the terms "individual," "patient," or "subject" are used interchangeably and refer to any single animal for which treatment is desired, e.g., mammals (including, e.g., dogs, cats, horses, rabbits, zoo animals, cows, pigs, sheep, and non-human animals such as non-human primates). In certain embodiments, a patient herein is a human.
[0063] As used herein, "administering" refers to a method of giving a dosage of a compound (e.g., an antagonist) or a pharmaceutical composition (e.g., a pharmaceutical composition including an antagonist) to a subject (e.g., a patient). Administration may be by any suitable means, including parenteral, intrapulmonary, and intranasal, and, if desired, for localized treatment, intralesional administration. Parenteral injections include, for example, intramuscular, intravenous, intraarterial, intraperitoneal, or subcutaneous administration. Administration may be by any suitable route, for example, by injection (e.g., intravenous or subcutaneous injection), depending in part on whether administration is temporary or chronic. Various dosing schedules are contemplated herein, including, but not limited to, single or multiple administrations over various time points, bolus administration, and pulse infusion.
[0064] The term "concurrently" is used herein to refer to the administration of two or more therapeutic agents, where at least a portion of the administration overlaps in time. Thus, concurrent administration includes dosing regimens where administration of one or more agents continues after administration of one or more other agents is discontinued.
[0065] The term "package insert" is used to refer to instructions typically included in the commercial packaging of a therapeutic product that contain information about the indications, use, dosage, administration, concomitant therapy, contraindications, and / or warnings regarding the use of such therapeutic product.
[0066] An "article of manufacture" is any product (e.g., package or container), or kit that contains at least one reagent, e.g., a pharmaceutical agent for the treatment of a disease or disorder described herein (e.g., cancer), or a probe for specifically detecting a biomarker (e.g., DNA methylation). In certain embodiments, the product or kit is promoted, distributed, or sold as a unit for performing a method described herein.
[0067] The term "methylation" is used herein (unless the context indicates otherwise) to refer to the presence of a methyl group at the C5 position of a cytosine nucleotide in a DNA nucleic acid. The term includes cytosine nucleotides in which the methyl group has been further modified, such as 5-methylcytosine (5mC), as well as 5-hydroxymethylcytosine (5hmC). The term also includes DNA nucleic acids that have been subjected to a chemical or enzymatic conversion of the nucleotide, such as a conversion that deaminates unmodified cytosine to uracil.
[0068] The term "aberrant methylation" is used herein to refer to a pattern of methylation that is not typically present in normal tissue. For example, the term can refer to increased methylation at a site that is not normally methylated in normal tissue, or decreased methylation at a site that is normally methylated in normal tissue. In some embodiments, nucleic acids from cancer cells (e.g., cancer nucleic acids) are characterized by aberrant methylation when their pattern and / or amount of methylation at one or more genomic loci differs from that normally present at the corresponding loci in a particular type of tissue.
[0069] The term "CpG dinucleotide" is used herein to refer to a region of two or more DNA bases in which a cytosine nucleotide is followed in a 5'→3' direction by a guanine nucleotide, e.g., 5'-C-phosphate-G-3'. In many genomes, CpG dinucleotides can often be found in "clusters" or regions of DNA containing multiple CpG dinucleotides (also called CpG islands). Much or most of the DNA methylation in many genomes occurs at CpG dinucleotides (wherein the cytosines are methylated or hydroxymethylated).
[0070] III. METHODS, SYSTEMS, AND DEVICES Certain aspects of the present disclosure relate to a method for detecting genetic and epigenetic information in a single workflow. In some embodiments, the method includes: providing a plurality of first single-stranded DNA fragments; subjecting the plurality of first single-stranded DNA fragments to a round of primer extension in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragments; (b) a nucleic acid polymerase; and (c) a mixture of nucleotides that contain a cytosine analogue that is resistant to cytosine conversion, thereby generating a plurality of first strands corresponding to the first single-stranded DNA fragments and a plurality of second strands that are complementary to the first strands, the second strands containing a cytosine analogue that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosine is converted if present in the first strand; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands. In some embodiments, detection is by microarray, quantitative PCR (qPCR), digital droplet PCR (ddPCR), molecular inversion probes, and / or sequencing. Exemplary steps for primer extension and cytosine conversion are illustrated in FIG.
[0071] As discussed herein, in some embodiments, the first strand of the present disclosure (i.e., containing methylation information such as one or more methylated and / or unmethylated cytosines) may be referred to as the "methylated strand". The methylated strand will contain a methylation level / pattern / mark based on the input DNA (e.g., the first single-stranded DNA fragment). As discussed herein, in some embodiments, the second strand of the present disclosure (i.e., containing the cytosine analogs of the present disclosure, such as those introduced via primer extension based on the first strand of the present disclosure) and / or its amplification product may be referred to as the "genomic strand". The genomic strand will be more likely to preserve sequence information, for example, after cytosine conversion as disclosed herein. In some embodiments, reference to the "methylated strand" includes the original methylated strand and its amplification product. In some embodiments, reference to the "genomic strand" includes the original genomic strand and its amplification product.
[0072] In some embodiments, the method of the present disclosure further comprises enriching one or both of the genomic strand and / or the methylated strand, or their amplification products. Enrichment can be achieved, for example, using hybrid capture, PCR (e.g., using strand-specific primers), multiple rounds of primer extension (e.g., before cytosine conversion treatment), etc. For example, in some embodiments, the method of the present disclosure further comprises enriching a plurality of second strands or their amplification products, e.g., before detection. In some embodiments, the method of the present disclosure further comprises enriching a plurality of first strands or their amplification products, e.g., before detection. In some embodiments, the method of the present disclosure further comprises enriching a plurality of first strands or their amplification products, e.g., before detection, and enriching a plurality of second strands or their amplification products, e.g., before detection.
[0073] In some embodiments, the method includes separating one or more first strands or their amplification products from one or more second strands or their amplification products prior to detection. In some embodiments, the separation is achieved by hybrid capture, for example, using a bait molecule that specifically binds / hybridizes / anneals to the genomic strand or the methylated strand or their amplification products. In some embodiments, the bait molecule specifically binds / hybridizes / anneals to, for example, a nucleic acid adapter sequence attached to the first and / or second strand.
[0074] In some embodiments, the separation includes (a) combining one or more bait molecules with the plurality of first and second strands, where the one or more bait molecules preferentially hybridize to one or more of the first strands or their amplification products, or to one or more of the second strands or their amplification products, thereby producing a nucleic acid hybrid, and (b) isolating the nucleic acid hybrid. In some embodiments, the one or more bait molecules hybridize to one or more of the first strands or their amplification products, but not to the plurality of second strands or their amplification products. In some embodiments, the separation occurs after the cytosine conversion treatment, where the one or more bait molecules hybridize to one or more of the first strands or their amplification products if one or more cytosines of the first strand undergo cytosine conversion. In some embodiments, the one or more bait molecules hybridize to one or more of the second strands or their amplification products, but not to the plurality of first strands or their amplification products. In some embodiments, the separating comprises combining one or more first bait molecules with the plurality of first and second strands, where one or more of the first bait molecules preferentially hybridize to one or more of the first strands or their amplification products, thereby producing a first nucleic acid hybrid; combining; isolating the first nucleic acid hybrid; combining one or more second bait molecules with the plurality of first and second strands, where one or more of the second bait molecules preferentially hybridize to one or more of the second strands or their amplification products, thereby producing a second nucleic acid hybrid; combining; isolating the second nucleic acid hybrid.
[0075] In some embodiments, the method further comprises subjecting the plurality of first strands and / or the plurality of second strands to amplification. In some embodiments, the amplification is prior to detecting (e.g., sequencing). In some embodiments, the detecting comprises detecting at least a portion of the plurality of first and second strands and their amplification products, if present. In some embodiments, the amplification is performed after a cytosine conversion treatment. See, e.g., Figures 4A-4F.
[0076] In some embodiments, the method includes subjecting the plurality of second strands to amplification in the presence of one or more primers that anneal to at least a portion of the plurality of second strands but not to the plurality of first strands. For example, the primers can be specific for a genomic strand.
[0077] In some embodiments, the method comprises subjecting the plurality of first strands to amplification in the presence of one or more primers that anneal with at least a portion of the plurality of first strands but not with the plurality of second strands. For example, the primers can be specific to methylated strands. In some embodiments, the one or more primers anneal with at least a portion of the plurality of first strands after the first strands undergo cytosine conversion but not before the first strands undergo cytosine conversion. In some embodiments, the one or more primers anneal with at least a portion of the plurality of first strands only if the cytosines of the first strands do not undergo cytosine conversion.
[0078] In some embodiments, the method includes subjecting a plurality of first single-stranded DNA fragments to two or more rounds of primer extension prior to cytosine conversion in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragments, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides that contain a cytosine analog that is resistant to cytosine conversion. For example, multiple rounds of primer extension (instead of one) can be performed prior to cytosine conversion (using a cytosine analog) to enrich for genomic strands.
[0079] In some embodiments, multiple first and second strands or their amplification products are detected together (eg, simultaneously) or separately.
[0080] In some embodiments, the plurality of first and second strands or their amplification products are sequenced at the same sequence depth. In some embodiments, the plurality of first and second strands or their amplification products are sequenced at different sequence depths. In some embodiments, the plurality of second strands are sequenced at a sequencing depth that is at least 2 times, at least 4 times, at least 8 times, at least 10 times, at least 50 times, at least 100 times, or at least 1000 times higher than the sequencing depth of the plurality of first strands. In some embodiments, the genomic strand or its amplification products are sequenced at a sequencing depth that is at least 2 times, at least 4 times, at least 8 times, at least 10 times, at least 50 times, at least 100 times, or at least 1000 times higher than the sequencing depth of the methylated strand. In some embodiments, the plurality of first strands are sequenced at a sequencing depth that is at least 2 times, at least 4 times, at least 8 times, at least 10 times, at least 50 times, at least 100 times, or at least 1000 times higher than the sequencing depth of the plurality of second strands. In some embodiments, the methylated strand or its amplification product is sequenced at a sequencing depth that is at least 2 times, at least 4 times, at least 8 times, at least 10 times, at least 50 times, at least 100 times, or at least 1000 times higher than the sequencing depth of the genomic strand. Using different sequencing depths can be advantageous depending on the application. For example, lower confidence calls such as single nucleotide polymorphisms can be subjected to higher sequencing depths to increase confidence.
[0081] In some embodiments, after primer extension, the plurality of first strands are hybridized to the plurality of second strands in the plurality of double-stranded nucleic acids.
[0082] Certain aspects of the present disclosure relate to adaptor nucleic acid sequences. In some embodiments, adaptor nucleic acid sequences may be used to distinguish genomic strands from methylated strands or their amplification products. In some embodiments, the methods of the present disclosure include demultiplexing sequence information from the first and / or second strands. In some embodiments, the methods of the present disclosure include demultiplexing sequence information to distinguish sequence reads from the first strand or methylated strand from sequence reads from the second strand or genomic strand. In some embodiments, the demultiplexing is based at least in part on the first and / or second adaptor nucleic acids of the present disclosure. In some embodiments, the adaptor nucleic acid sequence may further include sequences for encoding other types of information, including but not limited to sequences for identifying the sample, laboratory, physician, date, sequencing run, instrument, replicates, etc. In some embodiments, after cytosine conversion, the adaptor nucleic acid attached to the genomic strand is no longer complementary to the adaptor nucleic acid attached to the methylated strand, see, e.g., FIG. 2. In some embodiments, the adapter nucleic acid (e.g., attached to the methylated strand or the genomic strand) comprises a cytosine that is methylated or unmethylated, see, e.g., FIG. 3A. In some embodiments, the adapter nucleic acid comprises a non-complementary 5' overhang such that after primer extension, the adapter nucleic acid attached to the genomic strand is not complementary to the adapter nucleic acid attached to the methylated strand even without cytosine conversion. For example, FIG. 3B shows how the adapter nucleic acid sequence can comprise a region of non-complementarity sandwiched between two complementary portions of single-stranded DNA (see FIG. 3B, A), or an adapter sequence (see FIG. 3B, C), or a non-complementary 5' overhang (see FIG. 3B, B), such that after primer extension, a unique sequence in the adapter nucleic acid that distinguishes the two strands is generated independently of cytosine conversion.
[0083] In some embodiments, the disclosed method includes selective amplification of the methylated strand and / or the genomic strand. It is contemplated herein that strand-specific primers may be used to selectively amplify the methylated strand and / or the genomic strand, for example, using an adaptor nucleic acid sequence as a primer anchor, as shown in FIG. 3C. In the exemplary method illustrated in FIG. 3C, one or more rounds of PCR amplification are performed after cytosine conversion (e.g., using unbiased primers), after which the methylated strand and the genomic strand and their amplification products are amplified separately using strand-specific primers. Without wishing to be bound by theory, it is believed that preferential amplification performed on the products of the first and second strands, for example, after universal amplification, can prevent loss of genomic strand content when amplifying the methylated strand, and vice versa. If preferential amplification were instead performed on the original strand, one strand would be diluted and information content would be lost when specifically amplifying the other strand.
[0084] In some embodiments, the methods of the disclosure include attaching one or more nucleic acids to one or more of the first and / or second single-stranded DNA fragments. In some embodiments, the adaptors are attached via ligation, transposition, tailing, template switching, etc.
[0085] In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer, the primer comprising: (i) a second adaptor nucleic acid portion that does not anneal to the first single-stranded DNA fragments; and (ii) a portion that anneals to at least a portion of the first single-stranded DNA fragments; (b) a nucleic acid polymerase; and (c) a mixture of nucleotides that comprise a cytosine analog that is resistant to cytosine conversion, thereby generating a first single-stranded DNA fragment having a first adaptor nucleic acid complementary to the second adaptor nucleic acid. generating a plurality of first strands comprising the A fragment, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragment and a second adaptor nucleic acid, the second strand comprising a cytosine analog resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosines present in the first strand undergo cytosine conversion, where after cytosine conversion, the sequences of the first and second adaptor nucleic acids are no longer complementary; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands. In some embodiments, the detection is by microarray, quantitative PCR (qPCR), digital droplet PCR (ddPCR), molecular inversion probes, and / or sequencing.
[0086] In some embodiments, the second adaptor nucleic acid portion of the primer is between the two portions of the primer that anneal to the first single-stranded DNA fragment. In some embodiments, the second adaptor nucleic acid portion of the primer is 5' to the portion of the primer that anneals to the first single-stranded DNA fragment. For example, in some embodiments, the second adaptor nucleic acid portion does not anneal to the first single-stranded DNA fragment. Thus, the second adaptor nucleic acid portion can be introduced via the primer, and the complementary sequence can be introduced via backfilling from the 5' end of the first single-stranded DNA fragment (see, for example, FIG. 3, middle row). In some embodiments, the primer contains one or more unmethylated cytosines that become part of the second adaptor nucleic acid after primer extension and are converted during cytosine conversion.
[0087] In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments; attaching a first adaptor nucleic acid to at least a portion of the plurality of first single-stranded DNA fragments; and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragment downstream of or at the first adaptor nucleic acid, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides that include a cytosine analog that is resistant to cytosine conversion, thereby generating a plurality of first single-stranded DNA fragments comprising the first single-stranded DNA fragments having the first adaptor nucleic acid. generating a plurality of first strands, a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragment, and a second adaptor nucleic acid complementary to the first adaptor nucleic acid, the second strand comprising a cytosine analog resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosines present in the first strand undergo cytosine conversion, where after cytosine conversion, the sequences of the first and second adaptor nucleic acids are no longer complementary; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands. In some embodiments, the detection is by microarray, quantitative PCR (qPCR), digital droplet PCR (ddPCR), molecular inversion probes, and / or sequencing.
[0088] In some embodiments, the first adaptor nucleic acid is attached to the 3' end of at least a portion of the plurality of first single-stranded DNA fragments. In some embodiments, the first adaptor nucleic acid is attached to the 5' end of at least a portion of the plurality of first single-stranded DNA fragments. In some embodiments, the first adaptor nucleic acid comprises one or more unmethylated cytosines that are converted during cytosine conversion. In some embodiments, the first adaptor nucleic acid comprises one or more methylated cytosines and the primer comprises one or more unmethylated cytosines that are converted during cytosine conversion.
[0089] In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer, the primer comprising: (i) a second adaptor nucleic acid portion comprising a non-complementary 5' overhang that does not anneal to the first single-stranded DNA fragment; and (ii) a portion that anneals to at least a portion of the first single-stranded DNA fragment; (b) a nucleic acid polymerase; and (c) a mixture of nucleotides comprising a cytosine analog that is resistant to cytosine conversion, thereby generating a second adaptor nucleic acid portion that anneals to the first single-stranded DNA fragments. generating a plurality of first strands comprising first single-stranded DNA fragments having a first adaptor nucleic acid complementary to a portion of the adaptor nucleic acid portion, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragments and a second adaptor nucleic acid non-complementary to the first adaptor nucleic acid, the second strands comprising a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosines undergo cytosine conversion if present in the first strand; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.
[0090] In some embodiments, the methods of the present disclosure include end repair, adaptor ligation, primer extension, cytosine conversion treatment, PCR amplification / library generation, and parallel sequencing of the genomic and methylated strands as described herein. See, e.g., FIG. 4A.
[0091] In some embodiments, the methods of the present disclosure include end repair, adaptor ligation, primer extension, cytosine conversion treatment, PCR amplification / library generation, hybrid capture to enrich genomic and methylated strands, and sequencing of genomic and methylated strands as described herein. See, e.g., FIG. 4B.
[0092] In some embodiments, the method of the present disclosure includes end repair, adapter ligation, primer extension, cytosine conversion treatment, and PCR amplification / library generation as described herein. In some embodiments, separate libraries are constructed for genomic strands and methylated strands. In some embodiments, each library can be enriched (e.g., via hybrid capture) and sequenced. See, for example, FIG. 4C.
[0093] In some embodiments, the method of the present disclosure includes end repair, adaptor ligation, primer extension, cytosine conversion treatment, and PCR amplification / library generation as described herein. In some embodiments, separate libraries are constructed for genomic strand and methylated strand. In some embodiments, only the genomic strand library is enriched (e.g., via hybrid capture) and both libraries are sequenced. See, for example, FIG. 4D.
[0094] In some embodiments, the method of the present disclosure includes end repair, adaptor ligation, primer extension, cytosine conversion treatment, and PCR amplification / library generation as described herein. In some embodiments, separate libraries are constructed for genomic strand and methylated strand. In some embodiments, only the methylated strand library is enriched (e.g., via hybrid capture) and both libraries are sequenced. See, for example, FIG. 4E.
[0095] In some embodiments, the method of the present disclosure includes end repair, adaptor ligation, primer extension, cytosine conversion treatment, PCR amplification / library generation, hybrid capture to enrich genomic strand and methylated strand, and generate separate libraries based on genomic strand and methylated strand as described herein.Then, both libraries can be sequenced.See, for example, Figure 4F.
[0096] In some embodiments, the single-stranded DNA fragments of the present disclosure are obtained from double-stranded DNA, for example, from a sample of the present disclosure. In some embodiments, the method includes, for example, before primer extension, denaturing a plurality of double-stranded DNA fragments to provide a plurality of first single-stranded DNA fragments.
[0097] In some embodiments, the methods of the present disclosure are used to detect methylation of one of many CpG sites, islands, or shores. A CpG dinucleotide or site typically refers to a region of DNA where a cytosine nucleotide is located immediately adjacent to a guanine nucleotide in a linear sequence. "CpG" refers to a cytosine and a guanine separated by a phosphate (i.e., --C--phosphate --G--). Regions of DNA with a higher frequency or concentration of CpG sites are known as "CpG islands." Many genes in mammalian genomes have CpGs associated with the transcription start site (including promoters) of the gene, which plays a vital role in controlling gene expression. See, for example, U.S. Patent Publication No. US2014 / 0357497. Aberrant methylation patterns are observed in many types of cancer. For example, in normal tissues, CpG islands are often unmethylated, but a subset of islands are methylated during tumorigenesis, cell development, and various disease states. Hypermethylation (i.e., increased methylation levels) of CpG sites within the promoters of genes can lead to their silencing (e.g., silencing of tumor suppressor genes), a feature found, for example, in many human cancers.
[0098] In some embodiments, the disclosed methods include subjecting a nucleic acid (e.g., a plurality of first and second strands) to a cytosine conversion treatment, e.g., under conditions such that if an unmethylated cytosine is present in the first strand, it undergoes cytosine conversion (e.g., to uracil, which is converted to thymine during PCR amplification / primer extension). In some embodiments, the cytosine conversion is by bisulfite treatment, TET-assisted bisulfite treatment, oxidized bisulfite treatment, APOBEC, or TET / beta-glucosyltransferase-assisted APOBEC treatment.
[0099] In some embodiments, the cytosine conversion treatment converts unmethylated cytosines present in the first strand with a particular efficiency, for example, in some embodiments, at least 80%, at least 85%, at least 90%, at least 95%, or 100% of unmethylated cytosines present in the first strand undergo cytosine conversion as a result of the cytosine conversion treatment. In some embodiments, if unmethylated cytosines are present in the first strand, a percentage undergoes cytosine conversion as a result of the cytosine conversion treatment having an upper limit of 100%, 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, or 81% and an independently selected lower limit of 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, or 98%, where the upper limit is higher than the lower limit. In some embodiments, when unmethylated cytosines are present in the first strand, about 80% to about 97%, about 85% to about 97%, about 90% to about 97%, about 95% to about 97%, about 80% to about 95%, about 85% to about 95%, or about 90% to about 95% are converted to cytosine as a result of the cytosine conversion treatment. The present disclosure shows below that when unmethylated cytosines are present in the first strand, they are converted to cytosine with very high efficiency.
[0100] In some embodiments, the cytosine analogs of the present disclosure in the second strand are protected from the cytosine conversion process with a certain efficiency. For example, in some embodiments, up to 20%, up to 15%, up to 10%, up to 5%, up to 2%, up to 1%, up to 0.5%, or up to 0.1% of the cytosine analogs in the second strand undergo cytosine conversion as a result of the cytosine conversion process. In some embodiments, the cytosine analogs of the second strand undergo cytosine conversion as a result of the cytosine conversion process at a percentage having an upper limit of 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1% and a lower limit independently selected from 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, or 18%, where the upper limit is higher than the lower limit. In some embodiments, about 0.5% to about 5%, about 0.5% to about 3%, about 0.5% to about 1%, about 1% to about 5%, about 1% to about 3%, or about 1% to about 2% of the cytosine analogs in the second strand undergo cytosine conversion as a result of the cytosine conversion treatment. The present disclosure below shows that cytosine analogs in the second strand undergo cytosine conversion with very low efficiency.
[0101] Commonly used methods for determining methylation levels and / or patterns of DNA require methylation state-dependent conversion of cytosine to distinguish between methylated and unmethylated CpG dinucleotide sequences. For example, methylation of CpG dinucleotide sequences can be measured by using cytosine conversion-based techniques, which rely on methylation state-dependent chemical modification of CpG sequences in isolated genomic DNA or fragments thereof, followed by DNA sequence analysis. Chemical reagents capable of distinguishing between methylated and unmethylated CpG dinucleotide sequences include hydrazine, which cleaves nucleic acids, and bisulfite treatment. Bisulfite treatment followed by alkaline hydrolysis can be used to cleave nucleic acids, as described by Olek A., Nucleic Acids Res. 24:5064-6, 1996 or Frommer et al., Proc. Natl. Acad. Sci. USA 89:1827-1831 (1992). It specifically converts unmethylated cytosine to uracil and leaves 5-methylcytosine unmodified. Bisulfite-treated DNA can then be analyzed by conventional molecular techniques such as PCR amplification, sequencing, and detection including oligonucleotide hybridization. See, for example, U.S. Patent No. 10,174,372.
[0102] Various methodologies for cytosine conversion are known in the art. In some embodiments, the nucleic acids or nucleic acid fragments of the present disclosure are subjected to cytosine conversion, for example, by bisulfite treatment, TET-assisted bisulfite treatment, TET-assisted pyridine borane treatment, oxidized bisulfite treatment, or APOBEC treatment, prior to detection.
[0103] Thus, in some embodiments, the method of the present disclosure includes treating a plurality of nucleic acids or nucleic acid fragments of the present disclosure with bisulfite. Bisulfite sequencing is a method commonly used in the art for generating methylation data at single base resolution. Bisulfite conversion or treatment refers to a biochemical process for converting unmethylated cytosine residues to uracil or thymine residues (e.g., deamination to uracil followed by amplification as thymine during PCR), thereby preserving methylated cytosine residues (e.g., 5-methylcytosine, 5mC, or 5-hydroxymethylcytosine, 5hmC). Reagents for converting cytosine to uracil are known to those skilled in the art and include bisulfite reagents such as sodium bisulfite, potassium bisulfite, ammonium bisulfite, magnesium bisulfite, sodium pyrosulfite, potassium pyrosulfite, ammonium pyrosulfite, and magnesium pyrosulfite.
[0104] In some embodiments, the method of the present disclosure includes treating a plurality of nucleic acids or nucleic acid fragments of the present disclosure with enzymatic digestion and bisulfite treatment. The principle of the method is that DNA fragmentation is not achieved by sonication, but by combined enzymatic digestion with a plurality of endonucleases (MseI, Tsp509I, NlaIII, and Hpy CH4V), the restriction enzyme cleavage sites of MseI, Tsp509I, NlaIII, and Hpy CH4V being TTAA, AATT, CATG, and TGCA, respectively. See, for example, Smiraglia DJ, et al. Oncogene 2002;21:5414-5426. This is followed by, for example, bisulfite treatment as described herein.
[0105] Enzymatic methods for cytosine conversion, such as enzymatic methyl sequencing, are also known. Such approaches can be advantageous because they use enzymes instead of bisulfite, which can damage and fragment DNA, leading to DNA loss and potentially biased sequencing. For example, TET2 (Ten-eleven translocation (Tet) family 2 methylcytosine dioxygenase) and T4-BGT (T4 phage beta-glucosyltransferase) can be used to convert 5mC and 5hmC into products that cannot be deaminated by APOBEC3A (apolipoprotein B mRNA editing enzyme, catalytic polypeptide-like 3A), which is then used to deaminate unmodified cytosines by converting them to uracil. See, for example, Vaisvila, R. et al. (2021) Genome Res. 31:1-10.
[0106] In some embodiments, the method of the present disclosure includes treating a plurality of nucleic acids or nucleic acid fragments of the present disclosure with TET-assisted bisulfite (e.g., TAB-seq). In the TAB-seq approach, beta-glucosyltransferase (βGT) is used to convert 5hmC to β-glucosyl-5-hydroxymethylcytosine (5gmC), and a Tet enzyme (e.g., mTet1) is used to oxidize 5mC to 5-carboxylcytosine (5caC). The nucleic acid is then treated with bisulfite. See, e.g., Yu, M. et al. (2018) Methods Mol. Biol. 1708:645-663.
[0107] In some embodiments, the method of the present disclosure includes treating a plurality of nucleic acids or nucleic acid fragments of the present disclosure with TET-assisted pyridine borane (e.g., TAPS). In the TAPS approach, TET methylcytosine dioxygenase is used to oxidize 5mC and 5hmC to 5caC, which is then reduced to dihydrouracil (DHU) via pyridine borane. DHU is converted to thymine during subsequent PCR. See, for example, Liu, Y. et al. (2019) Nat. Biotechnol. 37: 424-429.
[0108] In some embodiments, the method of the present disclosure includes treating a plurality of nucleic acids or nucleic acid fragments of the present disclosure with oxidized bisulfite (e.g., oxBS). In the oxBS approach, 5hmC is oxidized to 5-formylcytosine (5fC), which can be converted to uracil under bisulfite. Sequencing results from bisulfite vs. oxidized bisulfite treatment can then be used to estimate 5hmC levels from 5mC. See, e.g., Booth, MJ et al. (2013) Nat. Protocols 8:1841-1851. This approach can be scaled at a genome-wide level in oxBS-seq, see, e.g., Kirschner, K. et al. (2018) Methods Mol. Biol. 1708:665-678.
[0109] In some embodiments, the method of the present disclosure comprises treating a plurality of nucleic acids or nucleic acid fragments of the present disclosure with APOBEC. Enzyme reagents for converting cytosine to uracil, i.e., cytosine deaminases, include those of the APOBEC family, such as APOBEC-seq or APOBEC3A. APOBEC family members are cytidine deaminases that convert cytosine to uracil while preserving 5-methylcytosine, i.e., without changing 5-methylcytosine. Such enzymes are described in US2013 / 0244237 and WO2018 / 165366 and are commercially available (see, e.g., NEBNext® Enzymatic Methyl-seq Kit, New England Biolabs). Non-limiting examples of APOBEC family proteins include APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D, APOBEC3F, APOBEC3G, APOBEC3H, APOBEC4, and activation-induced (cytidine) deaminase.
[0110] In some embodiments, the method of the present disclosure includes one or more rounds of primer extension using a mixture of nucleotides that contain cytosine analogs that are resistant to cytosine conversion. In some embodiments, the primer extension generates a plurality of first strands corresponding to the first single-stranded DNA fragments, and a plurality of second strands that are complementary to the first strands, the second strands containing cytosine analogs that are resistant to cytosine conversion. In some embodiments, the genomic strands and optionally their amplification products contain cytosine analogs. In some embodiments, the mixture of nucleotides includes adenine, guanine, thymine / uracil, and cytosine analogs.
[0111] A variety of cytosine analogs are contemplated for use in the methods described herein. The type of cytosine analog may depend on the type of cytosine conversion treatment, i.e., to ensure that the cytosine analog is resistant to the treatment / conversion. In some embodiments, cytosine analogs that are resistant to cytosine conversion include 5-methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), 5-carboxylcytosine (5caC), 5-formylcytosine, 5-(beta-D-glucosylmethyl)cytosine (5gmC), 5-ethyl dCTP, 5-methyl dCTP, 5-fluoro dCTP, 5-bromo dCTP, 5-iodo dCTP, 5-chloro dCTP, 5-trifluoromethyl dCTP, or 5-aza dCTP. In some embodiments, cytosine analogs that are resistant to cytosine conversion include, for example, cytosine for TET-assisted pyridine borane treatment.
[0112] In some embodiments, the polymerase used for primer extension is capable of incorporating cytosine analogs into nucleic acids.
[0113] In some embodiments, the method further comprises subjecting the plurality of first single-stranded DNA fragments to end repair, e.g., prior to primer extension.
[0114] In some embodiments, the method includes the use of TET-assisted pyridine borane sequencing (TAPS) processing. As known in the art, TET-assisted pyridine borane processing converts 5mC and 5hmC to uracil (which can be converted to thymine during PCR amplification). Thus, methylated cytosines are converted. Therefore, it is contemplated that the methods disclosed herein can be adapted for TAPS use by simply reversing which cytosines are converted (i.e., methylated vs. unmethylated).
[0115] In some embodiments, a method includes providing a plurality of first single-stranded DNA fragments; subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragments, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides comprising a cytosine analog that is resistant to TET-assisted pyridine borane treatment, thereby generating a plurality of first strands corresponding to the first single-stranded DNA fragments and a plurality of second strands complementary to the first strands, wherein the second strands comprise a cytosine analog that is resistant to TET-assisted pyridine borane treatment; subjecting the plurality of first and second strands to TET-assisted pyridine borane treatment under conditions such that any methylated cytosines present in the first strand undergo cytosine conversion; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands. In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments; attaching a first adapter nucleic acid to at least a portion of the plurality of first single-stranded DNA fragments; and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragment downstream of or at the first adapter nucleic acid, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides that include a cytosine analog that is resistant to TET-assisted pyridine borane treatment, thereby generating a plurality of first single-stranded DNA fragments comprising the first single-stranded DNA fragments having the first adapter nucleic acid. generating a plurality of second strands comprising a first strand, a second single-stranded DNA fragment complementary to the first single-stranded DNA fragment, and a second adapter nucleic acid complementary to the first adapter nucleic acid, wherein the second strand comprises a cytosine analog that is resistant to TET-assisted pyridine borane treatment; subjecting the plurality of first and second strands to TET-assisted pyridine borane treatment under conditions such that any methylated cytosines present in the first strand undergo cytosine conversion, where after cytosine conversion the sequences of the first and second adapter nucleic acids are no longer complementary; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer, the primer comprising: (i) a second adapter nucleic acid portion that does not anneal to the first single-stranded DNA fragment; and (ii) a portion that anneals to at least a portion of the first single-stranded DNA fragment, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides comprising a cytosine analog that is resistant to TET-assisted pyridine borane treatment, thereby generating a plurality of first single-stranded DNA fragments comprising first single-stranded DNA fragments having a first adapter nucleic acid complementary to the second adapter nucleic acid. generating a plurality of second strands comprising a first single-stranded DNA fragment, a second single-stranded DNA fragment complementary to the first single-stranded DNA fragment, and a second adapter nucleic acid non-complementary to the first adapter nucleic acid, wherein the second strand comprises a cytosine analog that is resistant to TET-assisted pyridine borane treatment; subjecting the plurality of first and second strands to TET-assisted pyridine borane treatment under conditions such that any methylated cytosines present in the first strand undergo cytosine conversion, where after cytosine conversion the sequences of the first and second adapter nucleic acids are no longer complementary; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.
[0116] In some aspects, the disclosure provides methods of detecting cancer in an individual. In some embodiments, the methods are performed on a plurality of nucleic acids obtained from the individual (e.g., from samples obtained from the individual), and methylation levels and / or somatic mutations or lack thereof detected in the samples identify the individual as having or not having cancer. In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments; subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragments, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides that include a cytosine analog that is resistant to cytosine conversion, thereby generating a plurality of first strands corresponding to the first single-stranded DNA fragments and a plurality of second strands that are complementary to the first strands, the second strands including a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosine undergoes cytosine conversion if present in the first strand; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer, the primer comprising: (i) a second adaptor nucleic acid portion that does not anneal to the first single-stranded DNA fragments; and (ii) a portion that anneals to at least a portion of the first single-stranded DNA fragments; (b) a nucleic acid polymerase; and (c) a mixture of nucleotides that comprise a cytosine analog that is resistant to cytosine conversion, thereby generating a first single-stranded DNA fragment having a first adaptor nucleic acid complementary to the second adaptor nucleic acid. generating a plurality of first strands comprising the A fragment, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragment, and a second adaptor nucleic acid, the second strand comprising a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosines present in the first strand undergo cytosine conversion, where after cytosine conversion, the sequences of the first and second adaptor nucleic acids are no longer complementary; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments; attaching a first adaptor nucleic acid to at least a portion of the plurality of first single-stranded DNA fragments; and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragment downstream of or at the first adaptor nucleic acid, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides that include a cytosine analog that is resistant to cytosine conversion, thereby generating a plurality of first single-stranded DNA fragments comprising the first single-stranded DNA fragments having the first adaptor nucleic acid. generating a plurality of first strands, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragment, and a second adaptor nucleic acid complementary to the first adaptor nucleic acid, wherein the second strand comprises a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosines present in the first strand undergo cytosine conversion, wherein after cytosine conversion, the sequences of the first and second adaptor nucleic acids are no longer complementary; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer, the primer comprising: (i) a second adaptor nucleic acid portion comprising a non-complementary 5' overhang that does not anneal to the first single-stranded DNA fragment; and (ii) a portion that anneals to at least a portion of the first single-stranded DNA fragment; (b) a nucleic acid polymerase; and (c) a mixture of nucleotides comprising a cytosine analog that is resistant to cytosine conversion, thereby generating a second adaptor nucleic acid portion that anneals to the first single-stranded DNA fragments. generating a plurality of first strands comprising first single-stranded DNA fragments having a first adaptor nucleic acid complementary to a portion of the adaptor nucleic acid portion, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragments and a second adaptor nucleic acid non-complementary to the first adaptor nucleic acid, the second strands comprising a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosines undergo cytosine conversion if present in the first strand; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.
[0117] In some aspects, the disclosure provides methods of detecting minimal residual disease in an individual who has been or is being treated for cancer. In some embodiments, the method is performed on a plurality of nucleic acids obtained from the individual (e.g., from samples obtained from the individual), and methylation levels and / or somatic mutations or lack thereof detected in the samples identify the individual as having minimal residual disease or lack thereof. In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments; subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragments, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides that include a cytosine analog that is resistant to cytosine conversion, thereby generating a plurality of first strands corresponding to the first single-stranded DNA fragments and a plurality of second strands that are complementary to the first strands, the second strands including a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosine undergoes cytosine conversion if present in the first strand; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer, the primer comprising: (i) a second adaptor nucleic acid portion that does not anneal to the first single-stranded DNA fragments; and (ii) a portion that anneals to at least a portion of the first single-stranded DNA fragments; (b) a nucleic acid polymerase; and (c) a mixture of nucleotides that comprise a cytosine analog that is resistant to cytosine conversion, thereby generating a first single-stranded DNA fragment having a first adaptor nucleic acid complementary to the second adaptor nucleic acid. generating a plurality of first strands comprising the A fragment, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragment, and a second adaptor nucleic acid, the second strand comprising a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosines present in the first strand undergo cytosine conversion, where after cytosine conversion, the sequences of the first and second adaptor nucleic acids are no longer complementary; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments; attaching a first adaptor nucleic acid to at least a portion of the plurality of first single-stranded DNA fragments; and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragment downstream of or at the first adaptor nucleic acid, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides that include a cytosine analog that is resistant to cytosine conversion, thereby generating a plurality of first single-stranded DNA fragments comprising the first single-stranded DNA fragments having the first adaptor nucleic acid. generating a plurality of first strands, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragment, and a second adaptor nucleic acid complementary to the first adaptor nucleic acid, wherein the second strand comprises a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosines present in the first strand undergo cytosine conversion, wherein after cytosine conversion, the sequences of the first and second adaptor nucleic acids are no longer complementary; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer, the primer comprising: (i) a second adaptor nucleic acid portion comprising a non-complementary 5' overhang that does not anneal to the first single-stranded DNA fragment; and (ii) a portion that anneals to at least a portion of the first single-stranded DNA fragment; (b) a nucleic acid polymerase; and (c) a mixture of nucleotides comprising a cytosine analog that is resistant to cytosine conversion, thereby generating a second adaptor nucleic acid portion that anneals to the first single-stranded DNA fragments. generating a plurality of first strands comprising first single-stranded DNA fragments having a first adaptor nucleic acid complementary to a portion of the adaptor nucleic acid portion, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragments and a second adaptor nucleic acid non-complementary to the first adaptor nucleic acid, the second strands comprising a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosines undergo cytosine conversion if present in the first strand; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.
[0118] In some aspects, the disclosure provides methods of screening an individual suspected of having cancer. In some embodiments, the methods are performed on a plurality of nucleic acids obtained from the individual (e.g., from samples obtained from the individual), and methylation levels and / or somatic mutations or lack thereof detected in the samples identify the individual as likely to have cancer or not likely to have cancer. In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments; subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragments, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides that include a cytosine analog that is resistant to cytosine conversion, thereby generating a plurality of first strands corresponding to the first single-stranded DNA fragments and a plurality of second strands that are complementary to the first strands, the second strands including a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosine undergoes cytosine conversion if present in the first strand; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer, the primer comprising: (i) a second adaptor nucleic acid portion that does not anneal to the first single-stranded DNA fragments; and (ii) a portion that anneals to at least a portion of the first single-stranded DNA fragments; (b) a nucleic acid polymerase; and (c) a mixture of nucleotides that comprise a cytosine analog that is resistant to cytosine conversion, thereby generating a first single-stranded DNA fragment having a first adaptor nucleic acid complementary to the second adaptor nucleic acid. generating a plurality of first strands comprising the A fragment, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragment, and a second adaptor nucleic acid, the second strand comprising a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosines present in the first strand undergo cytosine conversion, where after cytosine conversion, the sequences of the first and second adaptor nucleic acids are no longer complementary; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments; attaching a first adaptor nucleic acid to at least a portion of the plurality of first single-stranded DNA fragments; and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragment downstream of or at the first adaptor nucleic acid, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides that include a cytosine analog that is resistant to cytosine conversion, thereby generating a plurality of first single-stranded DNA fragments comprising the first single-stranded DNA fragments having the first adaptor nucleic acid. generating a plurality of first strands, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragment, and a second adaptor nucleic acid complementary to the first adaptor nucleic acid, wherein the second strand comprises a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosines present in the first strand undergo cytosine conversion, wherein after cytosine conversion, the sequences of the first and second adaptor nucleic acids are no longer complementary; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer, the primer comprising: (i) a second adaptor nucleic acid portion comprising a non-complementary 5' overhang that does not anneal to the first single-stranded DNA fragment; and (ii) a portion that anneals to at least a portion of the first single-stranded DNA fragment; (b) a nucleic acid polymerase; and (c) a mixture of nucleotides comprising a cytosine analog that is resistant to cytosine conversion, thereby generating a second adaptor nucleic acid portion that anneals to the first single-stranded DNA fragments. generating a plurality of first strands comprising first single-stranded DNA fragments having a first adaptor nucleic acid complementary to a portion of the adaptor nucleic acid portion, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragments and a second adaptor nucleic acid non-complementary to the first adaptor nucleic acid, the second strands comprising a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosines undergo cytosine conversion if present in the first strand; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.
[0119] In some aspects, the disclosure provides methods of determining the prognosis of an individual with cancer. In some embodiments, the methods are performed on a plurality of nucleic acids obtained from the individual (e.g., from samples obtained from the individual), and the methylation levels and / or somatic mutations or lack thereof detected in the samples at least partially determine the prognosis of the individual. In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments; subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragments, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides that include a cytosine analog that is resistant to cytosine conversion, thereby generating a plurality of first strands corresponding to the first single-stranded DNA fragments and a plurality of second strands that are complementary to the first strands, the second strands including a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosine undergoes cytosine conversion if present in the first strand; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer, the primer comprising: (i) a second adaptor nucleic acid portion that does not anneal to the first single-stranded DNA fragments; and (ii) a portion that anneals to at least a portion of the first single-stranded DNA fragments; (b) a nucleic acid polymerase; and (c) a mixture of nucleotides that comprise a cytosine analog that is resistant to cytosine conversion, thereby generating a first single-stranded DNA fragment having a first adaptor nucleic acid complementary to the second adaptor nucleic acid. generating a plurality of first strands comprising the A fragment, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragment, and a second adaptor nucleic acid, the second strand comprising a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosines present in the first strand undergo cytosine conversion, where after cytosine conversion, the sequences of the first and second adaptor nucleic acids are no longer complementary; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments; attaching a first adaptor nucleic acid to at least a portion of the plurality of first single-stranded DNA fragments; and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragment downstream of or at the first adaptor nucleic acid, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides that include a cytosine analog that is resistant to cytosine conversion, thereby generating a plurality of first single-stranded DNA fragments comprising the first single-stranded DNA fragments having the first adaptor nucleic acid. generating a plurality of first strands, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragment, and a second adaptor nucleic acid complementary to the first adaptor nucleic acid, wherein the second strand comprises a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosines present in the first strand undergo cytosine conversion, wherein after cytosine conversion, the sequences of the first and second adaptor nucleic acids are no longer complementary; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer, the primer comprising: (i) a second adaptor nucleic acid portion comprising a non-complementary 5' overhang that does not anneal to the first single-stranded DNA fragment; and (ii) a portion that anneals to at least a portion of the first single-stranded DNA fragment; (b) a nucleic acid polymerase; and (c) a mixture of nucleotides comprising a cytosine analog that is resistant to cytosine conversion, thereby generating a second adaptor nucleic acid portion that anneals to the first single-stranded DNA fragments. generating a plurality of first strands comprising first single-stranded DNA fragments having a first adaptor nucleic acid complementary to a portion of the adaptor nucleic acid portion, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragments and a second adaptor nucleic acid non-complementary to the first adaptor nucleic acid, the second strands comprising a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosines undergo cytosine conversion if present in the first strand; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.
[0120] In some aspects, the disclosure provides methods of predicting survival of an individual with cancer. In some embodiments, the methods are performed on a plurality of nucleic acids obtained from an individual (e.g., from samples obtained from the individual), and methylation levels and / or somatic mutations or lack thereof detected in the samples are at least partially predictive of survival of the individual. In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments; subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragments, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides that include a cytosine analog that is resistant to cytosine conversion, thereby generating a plurality of first strands corresponding to the first single-stranded DNA fragments and a plurality of second strands that are complementary to the first strands, the second strands including a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosine undergoes cytosine conversion if present in the first strand; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer, the primer comprising: (i) a second adaptor nucleic acid portion that does not anneal to the first single-stranded DNA fragments; and (ii) a portion that anneals to at least a portion of the first single-stranded DNA fragments; (b) a nucleic acid polymerase; and (c) a mixture of nucleotides that comprise a cytosine analog that is resistant to cytosine conversion, thereby generating a first single-stranded DNA fragment having a first adaptor nucleic acid complementary to the second adaptor nucleic acid. generating a plurality of first strands comprising the A fragment, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragment, and a second adaptor nucleic acid, the second strand comprising a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosines present in the first strand undergo cytosine conversion, where after cytosine conversion, the sequences of the first and second adaptor nucleic acids are no longer complementary; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments; attaching a first adaptor nucleic acid to at least a portion of the plurality of first single-stranded DNA fragments; and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragment downstream of or at the first adaptor nucleic acid, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides that include a cytosine analog that is resistant to cytosine conversion, thereby generating a plurality of first single-stranded DNA fragments comprising the first single-stranded DNA fragments having the first adaptor nucleic acid. generating a plurality of first strands, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragment, and a second adaptor nucleic acid complementary to the first adaptor nucleic acid, wherein the second strand comprises a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosines present in the first strand undergo cytosine conversion, wherein after cytosine conversion, the sequences of the first and second adaptor nucleic acids are no longer complementary; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer, the primer comprising: (i) a second adaptor nucleic acid portion comprising a non-complementary 5' overhang that does not anneal to the first single-stranded DNA fragment; and (ii) a portion that anneals to at least a portion of the first single-stranded DNA fragment; (b) a nucleic acid polymerase; and (c) a mixture of nucleotides comprising a cytosine analog that is resistant to cytosine conversion, thereby generating a second adaptor nucleic acid portion that anneals to the first single-stranded DNA fragments. generating a plurality of first strands comprising first single-stranded DNA fragments having a first adaptor nucleic acid complementary to a portion of the adaptor nucleic acid portion, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragments and a second adaptor nucleic acid non-complementary to the first adaptor nucleic acid, the second strands comprising a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosines undergo cytosine conversion if present in the first strand; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.
[0121] In some aspects, the disclosure provides methods of predicting or detecting tumor burden in an individual with cancer. In some embodiments, the methods are performed on a plurality of nucleic acids obtained from an individual (e.g., from samples obtained from the individual), and methylation levels and / or somatic mutations or lack thereof detected in the samples at least partially predict or detect the tumor burden of the individual. In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments; subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragments, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides that include a cytosine analog that is resistant to cytosine conversion, thereby generating a plurality of first strands corresponding to the first single-stranded DNA fragments and a plurality of second strands that are complementary to the first strands, the second strands including a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosine undergoes cytosine conversion if present in the first strand; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer, the primer comprising: (i) a second adaptor nucleic acid portion that does not anneal to the first single-stranded DNA fragments; and (ii) a portion that anneals to at least a portion of the first single-stranded DNA fragments; (b) a nucleic acid polymerase; and (c) a mixture of nucleotides that comprise a cytosine analog that is resistant to cytosine conversion, thereby generating a first single-stranded DNA fragment having a first adaptor nucleic acid complementary to the second adaptor nucleic acid. generating a plurality of first strands comprising the A fragment, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragment, and a second adaptor nucleic acid, the second strand comprising a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosines present in the first strand undergo cytosine conversion, where after cytosine conversion, the sequences of the first and second adaptor nucleic acids are no longer complementary; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments; attaching a first adaptor nucleic acid to at least a portion of the plurality of first single-stranded DNA fragments; and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragment downstream of or at the first adaptor nucleic acid, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides that include a cytosine analog that is resistant to cytosine conversion, thereby generating a plurality of first single-stranded DNA fragments comprising the first single-stranded DNA fragments having the first adaptor nucleic acid. generating a plurality of first strands, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragment, and a second adaptor nucleic acid complementary to the first adaptor nucleic acid, wherein the second strand comprises a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosines present in the first strand undergo cytosine conversion, wherein after cytosine conversion, the sequences of the first and second adaptor nucleic acids are no longer complementary; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer, the primer comprising: (i) a second adaptor nucleic acid portion comprising a non-complementary 5' overhang that does not anneal to the first single-stranded DNA fragment; and (ii) a portion that anneals to at least a portion of the first single-stranded DNA fragment; (b) a nucleic acid polymerase; and (c) a mixture of nucleotides comprising a cytosine analog that is resistant to cytosine conversion, thereby generating a second adaptor nucleic acid portion that anneals to the first single-stranded DNA fragments. generating a plurality of first strands comprising first single-stranded DNA fragments having a first adaptor nucleic acid complementary to a portion of the adaptor nucleic acid portion, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragments and a second adaptor nucleic acid non-complementary to the first adaptor nucleic acid, the second strands comprising a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosines undergo cytosine conversion if present in the first strand; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.
[0122] In some aspects, the disclosure provides methods of predicting a response to a treatment of an individual having cancer. In some embodiments, the methods are performed on a plurality of nucleic acids obtained from an individual (e.g., from a sample obtained from the individual), and methylation levels and / or somatic mutations or lack thereof detected in the samples are used, at least in part, to predict the individual's responsiveness to the treatment. In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments; subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragments, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides that include a cytosine analog that is resistant to cytosine conversion, thereby generating a plurality of first strands corresponding to the first single-stranded DNA fragments and a plurality of second strands that are complementary to the first strands, the second strands including a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosine undergoes cytosine conversion if present in the first strand; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer, the primer comprising: (i) a second adaptor nucleic acid portion that does not anneal to the first single-stranded DNA fragments; and (ii) a portion that anneals to at least a portion of the first single-stranded DNA fragments; (b) a nucleic acid polymerase; and (c) a mixture of nucleotides that comprise a cytosine analog that is resistant to cytosine conversion, thereby generating a first single-stranded DNA fragment having a first adaptor nucleic acid complementary to the second adaptor nucleic acid. generating a plurality of first strands comprising the A fragment, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragment, and a second adaptor nucleic acid, the second strand comprising a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosines present in the first strand undergo cytosine conversion, where after cytosine conversion, the sequences of the first and second adaptor nucleic acids are no longer complementary; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments; attaching a first adaptor nucleic acid to at least a portion of the plurality of first single-stranded DNA fragments; and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragment downstream of or at the first adaptor nucleic acid, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides that include a cytosine analog that is resistant to cytosine conversion, thereby generating a plurality of first single-stranded DNA fragments comprising the first single-stranded DNA fragments having the first adaptor nucleic acid. generating a plurality of first strands, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragment, and a second adaptor nucleic acid complementary to the first adaptor nucleic acid, wherein the second strand comprises a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosines present in the first strand undergo cytosine conversion, wherein after cytosine conversion, the sequences of the first and second adaptor nucleic acids are no longer complementary; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer, the primer comprising: (i) a second adaptor nucleic acid portion comprising a non-complementary 5' overhang that does not anneal to the first single-stranded DNA fragment; and (ii) a portion that anneals to at least a portion of the first single-stranded DNA fragment; (b) a nucleic acid polymerase; and (c) a mixture of nucleotides comprising a cytosine analog that is resistant to cytosine conversion, thereby generating a second adaptor nucleic acid portion that anneals to the first single-stranded DNA fragments. generating a plurality of first strands comprising first single-stranded DNA fragments having a first adaptor nucleic acid complementary to a portion of the adaptor nucleic acid portion, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragments and a second adaptor nucleic acid non-complementary to the first adaptor nucleic acid, the second strands comprising a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosines undergo cytosine conversion if present in the first strand; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.
[0123] In some aspects, the present disclosure provides a method for monitoring cancer in an individual. In some embodiments, the method is performed on a plurality of nucleic acids obtained from the individual (e.g., from a sample obtained from the individual). In some embodiments, the method is performed on a first plurality of nucleic acids obtained from the individual (e.g., from a first sample obtained from the individual) and a second plurality of nucleic acids obtained from the individual (e.g., from a second sample obtained from the individual), and determining the difference in methylation levels and / or somatic mutations between the first sample and the second sample is used to monitor cancer in the individual. In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments; subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragments, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides that include a cytosine analog that is resistant to cytosine conversion, thereby generating a plurality of first strands corresponding to the first single-stranded DNA fragments and a plurality of second strands that are complementary to the first strands, the second strands including a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosine undergoes cytosine conversion if present in the first strand; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer, the primer comprising: (i) a second adaptor nucleic acid portion that does not anneal to the first single-stranded DNA fragments; and (ii) a portion that anneals to at least a portion of the first single-stranded DNA fragments; (b) a nucleic acid polymerase; and (c) a mixture of nucleotides that comprise a cytosine analog that is resistant to cytosine conversion, thereby generating a first single-stranded DNA fragment having a first adaptor nucleic acid complementary to the second adaptor nucleic acid. generating a plurality of first strands comprising the A fragment, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragment, and a second adaptor nucleic acid, the second strand comprising a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosines present in the first strand undergo cytosine conversion, where after cytosine conversion, the sequences of the first and second adaptor nucleic acids are no longer complementary; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments; attaching a first adaptor nucleic acid to at least a portion of the plurality of first single-stranded DNA fragments; and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragment downstream of or at the first adaptor nucleic acid, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides that include a cytosine analog that is resistant to cytosine conversion, thereby generating a plurality of first single-stranded DNA fragments comprising the first single-stranded DNA fragments having the first adaptor nucleic acid. generating a plurality of first strands, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragment, and a second adaptor nucleic acid complementary to the first adaptor nucleic acid, wherein the second strand comprises a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosines present in the first strand undergo cytosine conversion, wherein after cytosine conversion, the sequences of the first and second adaptor nucleic acids are no longer complementary; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer, the primer comprising: (i) a second adaptor nucleic acid portion comprising a non-complementary 5' overhang that does not anneal to the first single-stranded DNA fragment; and (ii) a portion that anneals to at least a portion of the first single-stranded DNA fragment; (b) a nucleic acid polymerase; and (c) a mixture of nucleotides comprising a cytosine analog that is resistant to cytosine conversion, thereby generating a second adaptor nucleic acid portion that anneals to the first single-stranded DNA fragments. generating a plurality of first strands comprising first single-stranded DNA fragments having a first adaptor nucleic acid complementary to a portion of the adaptor nucleic acid portion, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragments and a second adaptor nucleic acid non-complementary to the first adaptor nucleic acid, the second strands comprising a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosines undergo cytosine conversion if present in the first strand; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.
[0124] In some aspects, the disclosure provides methods of monitoring the response of an individual undergoing treatment for cancer. In some embodiments, the method is performed on a plurality of nucleic acids obtained from the individual (e.g., from a sample obtained from the individual), e.g., after treatment. In some embodiments, the method includes administering the treatment to the individual and detecting methylation levels and / or somatic mutations or lack thereof, where the methylation levels and / or somatic mutations or lack thereof detected in the sample are used at least in part to monitor the response to the treatment. In some embodiments, the method is performed on a first plurality of nucleic acids obtained from the individual (e.g., from a first sample obtained from the individual) and a second plurality of nucleic acids obtained from the individual (e.g., from a second sample obtained from the individual), where determining the difference in methylation levels and / or somatic mutations between the first sample and the second sample is used to monitor cancer in the individual. In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments; subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragments, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides that include a cytosine analog that is resistant to cytosine conversion, thereby generating a plurality of first strands corresponding to the first single-stranded DNA fragments and a plurality of second strands that are complementary to the first strands, the second strands including a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosine undergoes cytosine conversion if present in the first strand; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer, the primer comprising: (i) a second adaptor nucleic acid portion that does not anneal to the first single-stranded DNA fragments; and (ii) a portion that anneals to at least a portion of the first single-stranded DNA fragments; (b) a nucleic acid polymerase; and (c) a mixture of nucleotides that comprise a cytosine analog that is resistant to cytosine conversion, thereby generating a first single-stranded DNA fragment having a first adaptor nucleic acid complementary to the second adaptor nucleic acid. generating a plurality of first strands comprising the A fragment, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragment, and a second adaptor nucleic acid, the second strand comprising a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosines present in the first strand undergo cytosine conversion, where after cytosine conversion, the sequences of the first and second adaptor nucleic acids are no longer complementary; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments; attaching a first adaptor nucleic acid to at least a portion of the plurality of first single-stranded DNA fragments; and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragment downstream of or at the first adaptor nucleic acid, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides that include a cytosine analog that is resistant to cytosine conversion, thereby generating a plurality of first single-stranded DNA fragments comprising the first single-stranded DNA fragments having the first adaptor nucleic acid. generating a plurality of first strands, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragment, and a second adaptor nucleic acid complementary to the first adaptor nucleic acid, wherein the second strand comprises a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosines present in the first strand undergo cytosine conversion, wherein after cytosine conversion, the sequences of the first and second adaptor nucleic acids are no longer complementary; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.In some embodiments, the method includes providing a plurality of first single-stranded DNA fragments and subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer, the primer comprising: (i) a second adaptor nucleic acid portion comprising a non-complementary 5' overhang that does not anneal to the first single-stranded DNA fragment; and (ii) a portion that anneals to at least a portion of the first single-stranded DNA fragment; (b) a nucleic acid polymerase; and (c) a mixture of nucleotides comprising a cytosine analog that is resistant to cytosine conversion, thereby generating a second adaptor nucleic acid portion that anneals to the first single-stranded DNA fragments. generating a plurality of first strands comprising first single-stranded DNA fragments having a first adaptor nucleic acid complementary to a portion of the adaptor nucleic acid portion, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragments and a second adaptor nucleic acid non-complementary to the first adaptor nucleic acid, the second strands comprising a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosines undergo cytosine conversion if present in the first strand; and detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.
[0125] detection In some embodiments, detecting (e.g., at least a portion of the plurality of first strands and / or at least a portion of the plurality of second strands of the present disclosure) is by microarray, quantitative PCR (qPCR), digital droplet PCR (ddPCR), molecular inversion probes, and / or sequencing.
[0126] In some embodiments, at least a portion of the plurality of first strands and at least a portion of the plurality of second strands of the present disclosure are detected by sequencing. In some embodiments, a plurality of sequence reads are obtained from at least a portion of the plurality of first strands of the present disclosure. In some embodiments, a plurality of sequence reads are obtained from at least a portion of the plurality of second strands of the present disclosure. In some embodiments, the sequencing is whole genome methyl sequencing (WGMS) or next generation sequencing (NGS). For example, in some embodiments, at least a portion of the plurality of first strands of the present disclosure (e.g., methylated strands) are detected via WGMS (e.g., enzymatic methylation sequencing), and / or at least a portion of the plurality of second strands of the present disclosure and / or their amplification products (e.g., genomic strands) are detected via WGS (e.g., NGS).
[0127] Various methods for WGMS are known in the art. Generally, these methods combine cytosine conversion (e.g., using the above-mentioned method) with whole genome sequencing technology. For example, in some embodiments, WGMS includes bisulfite sequencing, whole genome bisulfite sequencing (WGBS), APOBEC-seq, methyl-CpG binding domain (MBD) protein capture, methyl-DNA immunoprecipitation (MeDIP-seq), methylation-sensitive restriction enzyme sequencing (MSRE / MRE-Seq or Methyl-Seq), enzymatic methylation sequencing, oxidized bisulfite sequencing (oxBS-Seq), reduced representation bisulfite sequencing (RRBS), or Tet-assisted bisulfite sequencing (TAB-Seq).
[0128] Some WGMS methods rely on library construction and adapter ligation followed by standard bisulfite conversion and sequencing (e.g., WGBS). Alternatively, bisulfite treatment can be performed before adapter ligation (see, e.g., Miura, F. et al. (2012) Nucleic Acids Res. 40: e136). More recent techniques use other cytosine conversion methods, such as enzymatic approaches, to reduce DNA damage caused by bisulfite, for example in the commercially available NEBNext® Enzymatic Methyl-seq Kit (New England Biolabs). Library amplification, quantification, and sequencing steps generally follow bisulfite conversion. In some embodiments, nucleic acids are extracted from the sample prior to WGMS. In some embodiments, nucleic acids are subjected to fragmentation, repair, and adapter ligation prior to WGMS. As already mentioned, cytosine conversion can be performed before or after adapter ligation. In some embodiments, DNA repair is performed after cytosine conversion. PCR amplification (typically at least two cycles) is performed after cytosine conversion to convert uracil (previously generated by unmethylated cytosine) to thymine, accomplished using a polymerase capable of reading uracil (excluding polymerases with proofreading and repair activity). In some embodiments, prior to sequencing, fragments are enriched for the desired length. In some embodiments, prior to sequencing, nucleic acids are enriched for methylated sequences, such as by immunoprecipitation using an antibody specific for 5mC, as in the MeDIP approach (see, e.g., Pomraning, KR et al. (2009) Methods 47:142-150).
[0129] NGS methods are known in the art and are described, for example, in Metzker, M. (2010) Nature Biotechnology Reviews 11:31-46. Platforms for next-generation sequencing include, for example, Genome Sequencer (GS) FLX System from Roche / 454, Genome Analyzer (GA) from Illumina / Solexa, HiSeq 2500, HiSeq 3000, HiSeq 4000, and NovaSeq 6000 Sequencing System from Illumina, Support Oligonucleotide Ligation Detection (SOLiD) System from Life / APG, G.007 System from Polonator, HeliScope Gene Sequencing System from Helicos BioSciences, and PacBio RS System from Pacific Biosciences. NGS technology can include one or more steps, for example, template preparation, sequencing and imaging, and data analysis. Methods for template preparation may include steps such as randomly degrading nucleic acids (e.g., genomic DNA) into smaller sizes to generate sequencing templates (e.g., fragment templates or mate pair templates). Spatially separated templates may be attached or immobilized to a solid surface or support, allowing a large number of sequencing reactions to be performed simultaneously. Template types that may be used for NGS reactions include, for example, clonal amplified templates derived from a single DNA molecule, and single DNA molecule templates. Exemplary sequencing and imaging steps for NGS include, for example, circular reversible termination (CRT), sequencing by ligation (SBL), single molecule addition (pyrosequencing), and real-time sequencing. After NGS reads are generated, they may be aligned to a known reference sequence or assembled de novo. For example, identifying genetic alterations such as single nucleotide polymorphisms and structural variants in a sample (e.g., a tumor sample) may be achieved by aligning NGS reads to a reference sequence (e.g., a wild-type sequence).Methods for sequence alignment of NGS are described, for example, in Trapnell C. and Salzberg SL Nature Biotech., 2009, 27:455-457. Examples of de novo assembly are described, for example, in Warren R. et al., Bioinformatics, 2007, 23:500-501, Butler J. et al., Genome Res., 2008, 18:810-820, and Zerbino DR and Birney E., Genome Res., 2008, 18:821-829. Sequence alignment or assembly can be performed using read data from one or more NGS platforms, for example, mixing Roche / 454 and Illumina / Solexa read data. In some embodiments, NGS is performed according to the methods described, for example, in Frampton, GM et al. (2013) Nat. Biotech. 31:1023-1031, and / or Montesion, M., et al., Cancer Discovery (2021) 11(2):282-92.
[0130] In some embodiments, the method further comprises subjecting the plurality of nucleic acids to fragmentation prior to sequencing the plurality of polynucleotides or providing the plurality of sequence reads. Various DNA fragmentation techniques are used in the art prior to NGS or WGMS approaches. In some embodiments, the nucleic acid is fragmented by nebulization, where compressed gas is used to mechanically shear the nucleic acid through a small opening. In some embodiments, the nucleic acid is fragmented by sonication, where ultrasound is used to shear the nucleic acid. In some embodiments, the nucleic acid is enzymatically fragmented, for example, using one or more enzymes to digest the nucleic acid into fragments. For example, a mixture of two enzymes, one that randomly generates dsDNA nicks and one that recognizes the nick site, cleaves the opposite strand, and generates dsDNA breaks, see NEBNext® dsDNA Fragmentase.
[0131] In some embodiments, the method further comprises selectively enriching the plurality of nucleic acids or nucleic acid fragments, e.g., methylated strands and / or genomic strands, as described above, prior to sequencing the plurality of polynucleotides or providing the plurality of sequence reads. For example, one or more baits or probes can be used to hybridize with a genomic locus of interest or a fragment thereof, e.g., comprising a cluster of two or more CpG dinucleotides or comprising a genetic variant / mutation of interest. See, e.g., Graham, BI et al. Twist Fast Hybridization targeted methylation sequencing: a tunable target enrichment solution for methylation detection [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2021, 2021 Apr 10-15 and May 17-21. Philadelphia (PA): AACR, Cancer Res 2021, 81(13_Suppl): Abstract nr 2098. In some embodiments, two or more baits or probes are used, one set of baits or probes for selectively enriching the methylated strand and one set of baits or probes for selectively enriching the genomic strand. In some embodiments, two or more baits or probes are used, one set of baits or probes for selectively enriching the library generated using the methylated strand and one set of baits or probes for selectively enriching the library generated using the genomic strand. Thus, a combined hybrid capture for both methylation data and genomic data can be achieved, resulting in deep coverage and information for methylation information and genomic information in a single workflow, for example, as illustrated in Figure 14B.
[0132] In some embodiments, the method further comprises amplifying the plurality of nucleic acids or nucleic acid fragments by polymerase chain reaction (PCR) before sequencing the plurality of polynucleotides or providing the plurality of sequence reads. Various PCR techniques suitable for WGMS and NGS are known in the art. As mentioned above, in some embodiments, the plurality of nucleic acids or nucleic acid fragments are amplified by PCR after cytosine conversion. In some embodiments, PCR is achieved using the cytosine analogs of the present disclosure.
[0133] In some embodiments, the method further comprises, prior to sequencing the plurality of polynucleotides or prior to providing a plurality of sequence reads, contacting the mixture of polynucleotides with a bait molecule under conditions suitable for hybridization, wherein the mixture comprises a plurality of polynucleotides capable of hybridizing to the bait molecule, and isolating the plurality of polynucleotides hybridized to the bait molecule, wherein the isolated plurality of polynucleotides hybridized to the bait molecule are sequenced by NGS.
[0134] In some embodiments, the plurality of sequence reads are obtained by performing sequencing on the nucleic acid captured by hybridization with the bait molecule. In some embodiments, the plurality of sequence reads are obtained by performing whole exome sequencing on the nucleic acid captured by hybridization with the bait molecule. In some embodiments, the plurality of sequence reads are obtained by performing next generation sequencing (NGS), whole exome sequencing, or methylation sequencing (e.g., WGMS) on the nucleic acid captured by hybridization with the bait molecule.
[0135] In some embodiments, a hybrid capture approach is used. Further details on this and other hybrid capture processes can be found in U.S. Patent No. 9,340,830, Frampton, GM et al. (2013) Nat. Biotech. 31: 1023-1031, and Montesion, M., et al., Cancer Discovery (2021) 11(2): 282-92. In some embodiments, the method further comprises obtaining a sample from the individual prior to contacting the mixture of polynucleotides with the bait molecule, the sample comprising tumor cells and / or tumor nucleic acid, and extracting the mixture of polynucleotides from the sample, the mixture of polynucleotides being derived from tumor cells and / or tumor nucleic acid. In some embodiments, the sample further comprises non-tumor cells.
[0136] In some embodiments, the multiple sequence reads of the present disclosure include paired-end sequence reads. In general, paired-end sequencing methodology is described in, for example, WO2007 / 010252, WO2007 / 091077, and WO03 / 74734. This approach utilizes pairwise sequencing of double-stranded polynucleotide templates, which results in sequential determination of nucleotide sequences in two different, separate regions of the polynucleotide template. Paired-end methodology allows obtaining two combined or paired reads of sequence information from each double-stranded template on a cluster array, rather than only a single sequencing read as can be obtained by other methods. Paired-end sequencing techniques may specially use cluster arrays, which are typically formed by solid-phase amplification, as described, for example, in WO03 / 74734. Adapter-attached target polynucleotide duplexes are immobilized to a solid support at the 5'-end of each strand of each duplex, for example, via bridge amplification as described above, to form a dense cluster of double-stranded DNA. As both strands are immobilized at their 5'-ends, a sequencing primer is then hybridized to the free 3'-end and sequencing by synthesis is performed. Adapter sequences may be inserted between the target sequences, as described in WO2007 / 091077, allowing up to four reads from each duplex. In a further adaptation of this methodology, a particular strand may be cleaved in a controlled manner, as described in WO2007 / 010252. As a result, the timing of the sequencing reads from each strand may be controlled, allowing sequential determination of the nucleotide sequence in two different, separate regions on the complementary strand of a double-stranded template. See, for example, U.S. Pat. No. 10,174,372.
[0137] In some embodiments, the plurality of sequence reads comprises unpaired sequence reads.
[0138] In some embodiments, the methods of the present disclosure include demultiplexing sequence information from the first and / or second strand. In some embodiments, the methods of the present disclosure include demultiplexing sequence information to distinguish sequence reads from the first strand or methylated strand from sequence reads from the second strand or genomic strand. In some embodiments, the demultiplexing is based at least in part on the first and / or second adaptor nucleic acids of the present disclosure. Other capabilities include the possibility of using NGS index sequences to identify not only the sample but also the strand type (methyl or gene).
[0139] In some embodiments, the method further comprises comparing a plurality of second strand sequences to a plurality of first strand sequences, e.g., a sequence from a methylated strand can be compared to a corresponding sequence from a genomic strand, e.g., to detect conversions or lack thereof at one or more cytosines.
[0140] In some embodiments, the method further comprises comparing sequences from the plurality of first and / or second strands to a reference genome sequence. For example, a sequence from the methylated strand can be compared to a corresponding sequence from the reference genome sequence, e.g., to detect conversions or lack thereof at one or more cytosines, or to a corresponding sequence from the reference genome sequence, e.g., to detect sequence variants / mutations.
[0141] In some embodiments, the method of the present disclosure further comprises performing an alignment of sequence reads from the plurality to a reference genome, e.g., a human reference genome, prior to determining the consensus methylation pattern and CCF. In some embodiments, the alignment is a three-letter alignment to the human reference genome. In some embodiments, the method of the present disclosure further comprises filtering out sequencing reads from the plurality that could not undergo cytosine conversion prior to determining the consensus methylation pattern and CCF. In some embodiments, the method of the present disclosure further comprises filtering out sequence reads having a base other than cytosine or thymine at the first position of at least one of the CpG dinucleotides prior to determining the consensus methylation pattern and CCF. For example, these may be due to sequencing errors or mutations (somatic or germline). In some embodiments, the method of the present disclosure further comprises filtering out sequence reads having a base quality below a threshold base quality prior to determining the consensus methylation pattern and CCF. In some embodiments, a base call at a cytosine within a CpG dinucleotide is determined using two overlapping paired-end reads.
[0142] In some embodiments, detecting (e.g., at least a portion of the plurality of first strands and / or at least a portion of the plurality of second strands of the present disclosure) is by microarray. Suitable microarray technologies for detecting genetic variants (e.g., based on genomic strands) are known in the art, and in some embodiments, the microarray comprises probes specific for one or more genetic variants / mutations of interest. Suitable microarray technologies for detecting methylation (e.g., based on methylation strands) are known in the art, see, e.g., Deathage, DE et al. (2009) Methods Mol. Biol. 556:117-139.
[0143] In some embodiments, detecting (e.g., at least a portion of the plurality of first strands and / or at least a portion of the plurality of second strands of the present disclosure) is by quantitative PCR (qPCR). Suitable qPCR techniques for detecting genetic variants (e.g., based on genomic strands) are known in the art, and in some embodiments, the qPCR uses primers and / or probes specific for one or more genetic variants / mutations of interest. Suitable qPCR techniques for detecting methylation (e.g., based on genomic strands) are known in the art, and in some embodiments, the qPCR uses primers and / or probes specific for the methylation state at particular methylated / unmethylated cytosines (see, e.g., Dugast-Darzacq, C. and Grange, T. (2009) Methods Mol. Biol. 507:281-303).
[0144] In some embodiments, detecting (e.g., at least a portion of the plurality of first strands and / or at least a portion of the plurality of second strands of the present disclosure) is by digital droplet PCR (ddPCR). Suitable ddPCR techniques for detecting genetic variants (e.g., based on genomic strands) are known in the art, and in some embodiments, the ddPCR uses primers and / or probes specific for one or more genetic variants / mutations of interest. Suitable ddPCR techniques for detecting methylation (e.g., based on methylation strands) are known in the art, see, e.g., Yu, M. et al. (2018) Methods Mol. Biol. 1768: 363-383.
[0145] In some embodiments, detecting (e.g., at least a portion of the plurality of first strands and / or at least a portion of the plurality of second strands of the present disclosure) is by molecular inversion probe (MIP). MIP techniques suitable for detecting genetic variants (e.g., based on genomic strands) are known in the art, see, for example, Absalan, F. and Ronaghi, M. (2007) Methods Mol Biol. 396: 315-330. MIP techniques suitable for detecting methylation (e.g., based on methylation strands) are also known in the art, see, for example, Carrascosa, LGet al. (2014) Chem Commun (Camb) 50: 3585-3588.
[0146] Samples and cancer In some embodiments, the single-stranded and / or double-stranded DNA fragments of the present disclosure are obtained from a sample.
[0147] In some embodiments, the method of the present disclosure further comprises isolating a plurality of nucleic acids from the sample. In some embodiments, the nucleic acids are obtained from a sample that includes, for example, tumor cells and / or tumor nucleic acids. For example, the sample can include tumor cells, circulating tumor cells, tumor nucleic acids (e.g., tumor circulating tumor DNA, cfDNA, or cfRNA), part or all of a tumor biopsy, bodily fluids, cells, tissues, mRNA, cDNA, DNA, RNA, cell-free DNA, and / or cell-free RNA. In some embodiments, the sample is from a tumor biopsy or tumor specimen. In some embodiments, the sample further includes non-tumor cells and / or non-tumor nucleic acids. In some embodiments, the bodily fluids include blood, serum, plasma, saliva, semen, cerebrospinal fluid, amniotic fluid, peritoneal fluid, interstitial fluid, etc. In some embodiments, the sample further includes non-tumor cells and / or non-tumor nucleic acids.
[0148] In some embodiments, the sample comprises tissue, cells, and / or nucleic acid from cancer and / or tissue, cells, and / or nucleic acid from normal tissue. In some embodiments, the sample comprises a tissue biopsy sample, a liquid biopsy sample, or a normal control. In some embodiments, the sample is from a tumor biopsy, a tumor specimen, or circulating tumor cells. In some embodiments, the sample is a liquid biopsy sample and comprises blood, plasma, serum, cerebrospinal fluid, sputum, stool, urine, or saliva.
[0149] In some embodiments, the sample comprises a percentage of tumor nucleic acid that is less than 1% of the total nucleic acid, less than 0.5% of the total nucleic acid, less than 0.1% of the total nucleic acid, or less than 0.05% of the total nucleic acid, in some embodiments, the sample comprises a percentage of tumor nucleic acid that is at least 0.01%, at least 0.05%, or at least 0.1% of the total nucleic acid. In some embodiments, the sample is eluted with an upper limit of 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.9%, 0.8%, 0.7%, 0.6%, 0.5%, 0.4%, 0.3%, 0.2%, 0.1%, 0.09%, 0.08%, 0.07%, 0.06%, 0.05%, 0.04%, 0.03%, or 0.02% of the total nucleic acid and 0.0001%, 0.0002%, 0.0003%, 0.0004%, 0.0005%, 0.0006%, 0.0007%, 0.0008%, 0.0009%, 0.0010%, 0.0012%, 0.0014%, 0.0016%, 0.0018%, 0.0018%, 0.0019%, 0.0011%, 0.0012%, 0.0013%, 0.0014%, 0.0015%, 0.0016%, 0.0017%, 0.0018%, 0.0019%, 0.0019%, 0.0016%, 0.0019 ... 8%, 0.0009%, 0.001%, 0.002%, 0.003%, 0.004%, 0.005%, 0.006%, 0.007%, 0.008%, 0.009%, 0.01%, 0.02%, 0.03%, 0.04%, 0.05%, 0.06%, 0.07%, 0.08%, 0.09%, 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, or 1%, independently selected lower limit, wherein the upper limit is higher than the lower limit.
[0150] In some embodiments, the sample is or comprises a biological tissue or biological fluid. The sample may comprise compounds that are not naturally mixed with tissue in nature, such as preservatives, anticoagulants, buffers, fixatives, nutrients, antibiotics, etc. In one embodiment, the sample is preserved as a frozen sample or as a formaldehyde or paraformaldehyde fixed paraffin embedded (FFPE) tissue preparation. For example, the sample may be embedded in a matrix, such as an FFPE block or a frozen sample. In another embodiment, the sample is a blood or blood component sample. In yet another embodiment, the sample is a bone marrow aspirate sample. In another embodiment, the sample comprises cell-free DNA (cfDNA) or circulating cell-free DNA (ccfDNA), such as tumor cfDNA or tumor ccfDNA. Without wishing to be bound by theory, in some embodiments, the cfDNA is believed to be DNA from apoptotic or necrotic cells. Typically, cfDNA is bound by proteins (e.g., histones) and protected by nucleases. cfDNA can be used as a biomarker for, for example, non-invasive prenatal testing (NIPT), organ transplantation, cardiomyopathy, microbiome, and cancer. In another embodiment, the sample comprises circulating tumor DNA (ctDNA). Without wishing to be bound by theory, ctDNA is cfDNA that has genetic or epigenetic changes (e.g., somatic changes or methylation signatures) that can distinguish between tumor cells and those derived from non-tumor cells. In another embodiment, the sample comprises circulating tumor cells (CTCs). Without wishing to be bound by theory, in some embodiments, CTCs are believed to be cells released into circulation from primary or metastatic tumors. In some embodiments, CTC apoptosis is the source of ctDNA in blood / lymph.
[0151] In some embodiments of any of the methods provided herein, the cancer is carcinoma, sarcoma, lymphoma, leukemia, myeloma, germ cell cancer, or blastoma. In some embodiments, the cancer is a solid tumor. In some embodiments, the cancer is a hematological malignancy. In some embodiments, the cancer is a B-cell cancer, melanoma, breast cancer, lung cancer, bronchial cancer, colorectal cancer, prostate cancer, pancreatic cancer, gastric cancer, ovarian cancer, bladder cancer, brain cancer, central nervous system cancer, peripheral nervous system cancer, esophageal cancer, cervical cancer, endometrial cancer, oral cancer, pharyngeal cancer, liver cancer, kidney cancer, testicular cancer, biliary tract cancer, small intestine cancer, appendix cancer, salivary gland cancer, thyroid cancer, adrenal cancer, osteosarcoma, chondrosarcoma, cancer of the blood tissue, Adenocarcinoma, Inflammatory myofibroblastoma, Gastrointestinal stromal tumor (GIST), Colon cancer, Multiple myeloma (MM), Myelodysplastic syndrome (MDS), Myeloproliferative disorder (MPD), Acute lymphocytic leukemia (ALL), Acute myeloid leukemia (AML), Chronic myelocytic leukemia (CML), Chronic lymphocytic leukemia (CLL), Polycythemia vera, Hodgkin's lymphoma, Non-Hodgkin's lymphoma (NHL), Soft tissue sarcoma, Fibrosarcoma, Myxosarcoma, Liposarcoma , osteosarcoma, chordoma, angiosarcoma, endothelial sarcoma, lymphangiosarcoma, lymphangioendothelial sarcoma, synovium, mesothelioma, Ewing's tumor, leiomyosarcoma, rhabdomyosarcoma, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinoma, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, hepatocellular carcinoma, bile duct carcinoma, choriocarcinoma, seminoma, embryonal carcinoma, Wilms' tumor, bladder cancer, epithelial carcinoma, glioma, astrocytoma, medulloblastoma, craniopharyngioma , ependymoma, pineal cell tumor, glioblastoma, acoustic neuroblastoma, oligodendroglioma, meningioma, neuroblastoma, retinoblastoma, follicular lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, hepatocellular carcinoma, thyroid cancer, gastric cancer, head and neck cancer, small cell carcinoma, essential thrombocythemia, aplastic myeloid metaplasia, hypereosinophilic syndrome, systemic mastocytosis, familial eosinophilia, chronic eosinophilic leukemia, neuroendocrine carcinoma, or carcinoid tumor.
[0152] In some embodiments, the cancer is appendix adenocarcinoma, bladder adenocarcinoma, urothelial (transitional cell) carcinoma, breast cancer not otherwise specified (NOS), breast carcinoma NOS, invasive ductal carcinoma (IDC), invasive lobular carcinoma (ILC), cervical squamous cell carcinoma (SCC), colon adenocarcinoma (CRC), esophageal adenocarcinoma, esophageal carcinoma NOS, esophageal squamous cell carcinoma (SCC), intraocular melanoma of the eye, gallbladder adenocarcinoma, gastroesophageal junction adenocarcinoma, intrahepatic cholangiocarcinoma, kidney carcinoma NOS, hepatocellular carcinoma of the liver (HCC), lung cancer NOS, lung adenocarcinoma, lung large cell carcinoma, non-small cell lung carcinoma (NSCLC) of the lung, small cell undifferentiated carcinoma of the lung, squamous cell carcinoma of the lung (SCC), ovarian cancer NOS, pancreatic cancer NOS, pancreatic ductal adenocarcinoma, pancreaticobiliary carcinoma, prostate cancer NOS, prostate acinar adenocarcinoma, prostate ductal adenocarcinoma, rectal adenocarcinoma (CRC), skin melanoma, small intestine adenocarcinoma, soft tissue sarcoma NOS, gastric adenocarcinoma NOS, carcinoma of unknown primary source NOS, adenocarcinoma of unknown primary source, carcinoma of unknown primary source (CUP) NOS, neuroendocrine tumor of unknown primary source, squamous cell carcinoma of unknown primary source (SCC), or endometrial adenocarcinoma of the uterus NOS.
[0153] In some embodiments, a sample of the present disclosure is obtained from an individual. In some embodiments, the individual has cancer. In some embodiments, the individual is suspected of having cancer. In some embodiments, the individual has been screened for cancer or its recurrence or remission. In some embodiments, the individual is undergoing or has undergone treatment for, for example, cancer.
[0154] Software, Systems, and Devices In another aspect, provided herein is a system comprising one or more processors and a memory configured to store one or more computer program instructions, which when executed by the one or more processors are configured to obtain a first plurality of sequence reads of one or more first nucleic acid molecules or their amplification products, where the first nucleic acid molecules have been subjected to a cytosine conversion treatment under conditions such that unmethylated cytosines, if present in the first nucleic acid molecule, are subjected to cytosine conversion; obtain a second plurality of sequence reads of one or more second nucleic acid molecules or their amplification products, where the second nucleic acid molecules are complementary to the first nucleic acid molecule prior to cytosine conversion and comprise a cytosine analog that is resistant to cytosine conversion; analyze the first plurality of sequence reads for the presence or absence of methylation at one or more cytosines inferred based on the cytosine conversion or lack thereof; and analyze the second plurality of sequence reads for sequence information. In some embodiments, the one or more first nucleic acid molecules or their amplification products further comprise a first adaptor nucleic acid sequence, and the one or more computer program instructions, when executed by the one or more processors, are further configured to demultiplex the first and second plurality of sequence reads based at least in part on the detection of the first adaptor nucleic acid sequence in the first plurality of sequence reads. In some embodiments, the one or more second nucleic acid molecules or their amplification products further comprise a second adaptor nucleic acid sequence, and the one or more computer program instructions, when executed by the one or more processors, are further configured to demultiplex the first and second plurality of sequence reads based at least in part on the detection of the second adaptor nucleic acid sequence in the second plurality of sequence reads. In some embodiments, the first plurality of sequence reads are at least 2-fold, at least 4-fold, at least 8-fold, at least 10-fold, at least 50-fold, at least 100-fold, or at least 1000-fold more abundant than the second plurality of sequence reads. In some embodiments, the second plurality of sequence reads are at least 2-fold, at least 4-fold, at least 8-fold, at least 10-fold, at least 50-fold, at least 100-fold, or at least 1000-fold more abundant than the first plurality of sequence reads.In some embodiments, the first and / or second plurality of sequence reads are obtained by sequencing, optionally, the sequencing includes using a massively parallel sequencing (MPS) technology, a whole genome sequencing (WGS), a whole exome sequencing, a targeted sequencing, a direct sequencing, or a Sanger sequencing technology, optionally, the massively parallel sequencing technology includes next generation sequencing (NGS). In some embodiments, the one or more program instructions, when executed by the one or more processors, are further configured to generate a molecular profile for the sample based at least in part on the analysis. In some embodiments, the molecular profile further includes results from a comprehensive genomic profiling (CGP) test, a gene expression profiling test, a cancer hotspot panel test, a DNA methylation test, a DNA fragmentation test, an RNA fragmentation test, or any combination thereof. In some embodiments, the individual is administered a treatment based at least in part on the molecular profile. In some embodiments, the molecular profile further includes results from a nucleic acid sequencing-based test. In some embodiments, the one or more computer program instructions, when executed by the one or more processors, are further configured to compare a sequence read of the second plurality of sequence reads to a sequence read of the first plurality of sequence reads. In some embodiments, the one or more computer program instructions, when executed by the one or more processors, are further configured to compare one or more sequence reads of the first and / or second plurality of sequence reads to a reference genome sequence.
[0155] In another aspect, provided herein is a non-transitory computer readable storage medium comprising one or more programs executable by one or more computer processors to perform a method, the method comprising: obtaining a first plurality of sequence reads of one or more first nucleic acid molecules or amplification products thereof, where the first nucleic acid molecules have been subjected to a cytosine conversion treatment under conditions such that unmethylated cytosines, if present in the first nucleic acid molecule, are subjected to cytosine conversion; obtaining a second plurality of sequence reads of one or more second nucleic acid molecules or amplification products thereof, where the second nucleic acid molecules are complementary to the first nucleic acid molecule prior to cytosine conversion and comprise a cytosine analog that is resistant to cytosine conversion; analyzing the first plurality of sequence reads for presence or absence of methylation at one or more cytosines inferred based on the cytosine conversion or lack thereof; and analyzing the second plurality of sequence reads for sequence information. In some embodiments, the one or more first nucleic acid molecules or their amplification products further comprise a first adaptor nucleic acid sequence, and the method further comprises demultiplexing the first and second plurality of sequence reads based at least in part on detection of the first adaptor nucleic acid sequence in the first plurality of sequence reads. In some embodiments, the one or more second nucleic acid molecules or their amplification products further comprise a second adaptor nucleic acid sequence, and the method further comprises demultiplexing the first and second plurality of sequence reads based at least in part on detection of the second adaptor nucleic acid sequence in the second plurality of sequence reads. In some embodiments, the first plurality of sequence reads are at least 2-fold, at least 4-fold, at least 8-fold, at least 10-fold, at least 50-fold, at least 100-fold, or at least 1000-fold more abundant than the second plurality of sequence reads. In some embodiments, the second plurality of sequence reads are at least 2-fold, at least 4-fold, at least 8-fold, at least 10-fold, at least 50-fold, at least 100-fold, or at least 1000-fold more abundant than the first plurality of sequence reads.In some embodiments, the first and / or second plurality of sequence reads are obtained by sequencing, optionally, the sequencing includes using a massively parallel sequencing (MPS) technology, a whole genome sequencing (WGS), a whole exome sequencing, a targeted sequencing, a direct sequencing, or a Sanger sequencing technology, optionally, the massively parallel sequencing technology includes next generation sequencing (NGS). In some embodiments, the method further includes generating a molecular profile for the sample based at least in part on the analysis. In some embodiments, the molecular profile further includes results from a comprehensive genomic profiling (CGP) test, a gene expression profiling test, a cancer hotspot panel test, a DNA methylation test, a DNA fragmentation test, an RNA fragmentation test, or any combination thereof. In some embodiments, the individual is administered a treatment based at least in part on the molecular profile. In some embodiments, the molecular profile further includes results from a nucleic acid sequencing-based test. In some embodiments, the method further includes comparing a sequence read of the second plurality of sequence reads to a sequence read of the first plurality of sequence reads. In some embodiments, the method further comprises comparing a sequence read of the first and / or second plurality of sequence reads to a reference genome sequence.
[0156] In some embodiments, a molecular profile or report of the present disclosure includes results from sequencing the methylated strand and / or the genomic strand, for example, using the methods of the present disclosure.
[0157] FIG. 16 illustrates an example of a computing device according to one embodiment. The device 1100 may be a host computer connected to a network. The device 1100 may be a client computer or a server. As shown in FIG. 16, the device 1100 may be any suitable type of microprocessor-based device, such as a personal computer, a workstation, a server, or a handheld computing device (portable electronic device) such as a phone or tablet. The device may include, for example, a processor 1110, input devices 1120, output devices 1130, storage 1140, communication devices 1160, a power supply 1170, an operation system 1180, and a system bus 1190. The input devices 1120 and output devices 1130 may generally correspond to those described herein and may be connectable to or integrated with the computer.
[0158] The input device 1120 may be any suitable device that provides input, such as a touch screen, a keyboard or keypad, a mouse, or a voice recognition device. The output device 1130 may be any suitable device that provides output, such as a touch screen, a tactile device, or a speaker.
[0159] Storage 1140 can be any suitable device that provides storage (e.g., electrical, magnetic, or optical memory, including RAM (volatile and non-volatile), cache, hard drive, or removable storage disk). Communications device 1160 can include any suitable device that can send and receive signals over a network, such as a network interface chip or device. The components of a computer can be connected in any suitable manner, for example, via wired media (e.g., a physical bus, Ethernet, or any other wired transmission technology) or wirelessly (e.g., Bluetooth, Wi-Fi, or any other wireless technology). For example, in FIG. 16, the components are connected by a system bus 1190.
[0160] The detection module 1150 can be stored as executable instructions in the storage 1140 and executed by the processor 1110 and can include, for example, processes embodying functions of the present disclosure (e.g., embodied in the devices described above).
[0161] The detection module 1150 may also be stored and / or transferred in any non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, or device (e.g., those described herein) and may fetch instructions associated with the software from and execute the instructions. In the context of the present disclosure, a computer-readable storage medium may be any medium, such as the storage 1140, that may include or store processes for use by or in connection with an instruction execution system, apparatus, or device. Examples of computer-readable storage media may include memory units, such as hard drives, flash drives, and distribution modules that operate as a single functional unit. Also, the various processes described herein may be embodied as modules configured to operate according to the above embodiments and techniques. Furthermore, while the processes may be shown and / or described separately, one skilled in the art will understand that the above processes may be routines or modules within other processes.
[0162] The detection module 1150 may also propagate in any transmission medium for use by or in connection with an instruction execution system, apparatus, or device (such as those mentioned above) and may fetch instructions associated with the software from and execute the instructions. In the context of this disclosure, a transmission medium may be any medium that may communicate, propagate, or transmit transmission programming for use by or in connection with an instruction execution system, apparatus, or device. A transmission-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, or infrared wired or wireless propagation media.
[0163] The device 1100 may be connected to a network (e.g., network 1004 shown in FIG. 15 and / or described below), which may be any suitable type of interconnected communications system. The network may implement any suitable communications protocol and may be protected by any suitable security protocol. The network may include any suitable arrangement of network links that may implement transmission and reception of network signals, such as wireless network connections (T1 or T3 lines), cable networks, DSL, or telephone lines.
[0164] The device 1100 may implement any operating system suitable for operating on a network (e.g., operating system 1180). The detection module 1150 may be written in any suitable programming language, such as C, C++, Java, or Python. In various embodiments, application software embodying functionality of the present disclosure may be deployed in different configurations (e.g., in a client / server arrangement, or via a web browser as a web-based application or web service). In some embodiments, the operating system 1180 is executed by one or more processors, such as the processor 1110.
[0165] The device 1100 may further include a power source 1170, which may be any suitable power source.
[0166] In some embodiments, detection module 1150 is a module for detecting LOH and / or tumor mutation burden of one or more HLA-I genes and includes a process that embodies the functionality of the present disclosure (e.g., as embodied in the devices described herein).
[0167] 15 illustrates an example of a computing system according to one embodiment. In system 1000, device 1100 (e.g., as described above and illustrated in FIG. 16) is connected to network 1004, which is also connected to device 1006. In some embodiments, device 1006 is a sequencer. Exemplary sequencers may include, but are not limited to, Roche / 454's Genome Sequencer (GS) FLX System, Illumina / Solexa's Genome Analyzer (GA), Illumina's HiSeq 2500, HiSeq 3000, HiSeq 4000, and NovaSeq 6000 sequencing systems, Life / APG's Support Oligonucleotide Ligation Detection (SOLiD) system, Polonator's G.007 system, HeliScope Gene sequencing system from Helicos BioSciences, or Pacific Biosciences' PacBio RS system. The devices 1100 and 1006 can communicate using a suitable communication interface over a network 1004, such as, for example, a local area network (LAN), a virtual private network (VPN), or the Internet. In some embodiments, the network 1004 can be, for example, the Internet, an intranet, a virtual private network, a cloud network, a wired network, or a wireless network. The devices 1100 and 1006 can communicate partially or entirely over wireless or wired communications, such as Ethernet, IEEE 802.11b wireless, and the like. In addition, the devices 1100 and 1006 can communicate over a second network, such as, for example, a mobile / cellular network, using a suitable communication interface. The communication between the devices 1100 and 1006 can further include or communicate with various servers, such as a mail server, a mobile server, a media server, a telephone server, and the like.In some embodiments, devices 1100 and 1006 may communicate directly (instead of or in addition to communicating via network 1004), e.g., via wireless or wired communication, such as Ethernet, IEEE 802.11b wireless, etc. In some embodiments, devices 1100 and 1006 communicate via communication 1008, which may be a direct connection or may occur over a network (e.g., network 1004).
[0168] One or all of the devices 1100 and 1006 generally include logic (e.g., http web server logic) or are programmed to format data accessed from local or remote databases or other sources of data and content to provide and / or receive information over the network 1004 in accordance with various embodiments described herein.
[0169] FIG. 14C illustrates an exemplary process 1400 for detecting genetic and epigenetic information in a single workflow, according to some embodiments of the present disclosure. The process 1400 is performed, for example, using one or more electronic devices implementing a software program. In some examples, the process 1400 is performed using a client-server system, and the blocks of the process 1400 are divided in any manner between a server and a client device. In other examples, the blocks of the process 1400 are divided between a server and multiple client devices. Thus, while portions of the process 1400 are described herein as being performed by a particular device of a client-server system, it will be understood that the process 1400 is not so limited. In some embodiments, the steps performed may be performed across many systems, for example, in a cloud environment. In other examples, the process 1400 is performed using only a client device or only multiple client devices. In the process 1400, some blocks are optionally combined, the order of some blocks is optionally changed, and some blocks are optionally omitted. In some examples, additional steps may be performed in combination with process 1400. Accordingly, the operations illustrated (and described in more detail below) are exemplary in nature and, therefore, should not be considered as limiting.
[0170] In block 1402, a first plurality of sequence reads of one or more first nucleic acid molecules or their amplification products are obtained, and the first nucleic acid molecules are subjected to a cytosine conversion treatment under conditions (e.g., as described herein) such that unmethylated cytosines are subjected to cytosine conversion if present in the first nucleic acid molecule.
[0171] In block 1404, a second plurality of sequence reads are obtained for one or more second nucleic acid molecules or their amplification products, where the second nucleic acid molecules are complementary to the first nucleic acid molecules prior to cytosine conversion and contain cytosine analogs that are resistant to cytosine conversion (e.g., as described herein).
[0172] In block 1406, an exemplary system (e.g., one or more electronic devices) analyzes the first plurality of sequence reads for the presence or absence of methylation at one or more cytosines inferred based on cytosine conversion or lack thereof.
[0173] At block 1408, an exemplary system (e.g., one or more electronic devices) analyzes the second plurality of sequence reads for sequence information.
[0174] Optionally, the method further includes demultiplexing the first and second plurality of sequence reads based at least in part on the detection of the first adaptor nucleic acid sequence in the first plurality of sequence reads.
[0175] In some embodiments, the first and / or second plurality of sequence reads are obtained using a sequencer described herein or otherwise known in the art, such as one for performing WGMS and / or WGS.
[0176] In some embodiments, prior to block 1402 and / or 1402, the first and / or second plurality of sequence reads are generated by any of the methods of the present disclosure, e.g., based on the methylation strand and / or the genomic strand as described herein.
[0177] report In some embodiments, the methods provided herein include generating a report and / or providing the report to a party. In some embodiments, the report includes one or more treatment options identified for the individual based at least in part on the methylation and / or somatic mutations, or lack thereof, detected in a sample from the individual, e.g., as described herein. In some embodiments, the one or more treatment options are based at least in part on the methylation and / or somatic mutations, or lack thereof, detected, in a sample from the individual.
[0178] A report according to the present disclosure may be in electronic, web-based, or paper form. The report may be provided to an individual or patient (e.g., an individual or patient with cancer) or to an individual or entity other than the individual or patient (e.g., other than the individual or patient with cancer), such as one or more of a caregiver, a physician, an oncologist, a hospital, a clinic, a third party payer, an insurance company, or a government agency. In some embodiments, the report is provided or delivered to the individual or entity within about 1 day or more, about 7 days or more, about 14 days or more, about 21 days or more, about 30 days or more, about 45 days or more, or about 60 days or more after obtaining a sample from an individual (e.g., an individual with cancer). In some embodiments, the report is provided or delivered to the individual or entity within about 1 day or more, about 7 days or more, about 14 days or more, about 21 days or more, about 30 days or more, about 45 days or more, or about 60 days or more after detecting methylation and / or somatic mutations in a sample obtained from an individual (e.g., an individual with cancer).
[0179] IV. Illustrative Embodiments The following exemplary embodiments are representative of certain aspects of the present invention. Embodiment 1. A method for detecting genetic and epigenetic information in a single workflow, comprising: providing a plurality of first single-stranded DNA fragments; subjecting a plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragments, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides that contain cytosine analogs that are resistant to cytosine conversion, thereby generating a plurality of first strands corresponding to the first single-stranded DNA fragments and a plurality of second strands that are complementary to the first strands, wherein the second strands contain cytosine analogs that are resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosines present in the first strand undergo cytosine conversion; detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands. Embodiment 2. The method of embodiment 1, wherein the detecting is by sequencing. Embodiment 3. The method of embodiment 2, wherein the sequencing is next generation sequencing (NGS). Embodiment 4. The method of embodiment 1, wherein the detecting is by microarray, quantitative PCR (qPCR), digital droplet PCR (ddPCR), or molecular inversion probes. Embodiment 5. The method of any one of embodiments 1-4, wherein the method further comprises subjecting the plurality of first strands and / or the plurality of second strands to amplification prior to detecting, and wherein detecting comprises detecting at least a portion of the plurality of first and second strands and their amplification products, if present. Embodiment 6. The method of embodiment 5, comprising subjecting the plurality of second strands to amplification in the presence of one or more primers that anneal to at least a portion of the plurality of second strands but not to the plurality of first strands. Embodiment 7. The method of embodiment 5 or embodiment 6, comprising subjecting the plurality of first strands to amplification in the presence of one or more primers that anneal to at least a portion of the plurality of first strands but not to the plurality of second strands. Embodiment 8. The method of embodiment 7, wherein the one or more primers anneal to at least a portion of the plurality of first strands after the first strands have undergone cytosine conversion but not before the first strands have undergone cytosine conversion. Embodiment 9. The method of embodiment 7, wherein the one or more primers anneal to at least a portion of the plurality of first strands only if a cytosine of the first strand does not undergo cytosine conversion. Embodiment 10. The method according to any one of embodiments 5 to 9, wherein the amplification occurs after cytosine conversion. Embodiment 11. The method according to any one of embodiments 1 to 10, wherein the method further comprises enriching the plurality of second strands or their amplification products prior to detection. Embodiment 12. The method according to any one of embodiments 1 to 11, wherein the method further comprises enriching the plurality of first strands or their amplification products prior to detection. Embodiment 13. The method of embodiment 11 or embodiment 12, wherein the method comprises separating one or more first strands or amplification products thereof from one or more second strands or amplification products thereof prior to detection. Embodiment 14. The method of embodiment 13, wherein the separating comprises: (a) combining one or more bait molecules with a plurality of first and second strands, where one or more bait molecules preferentially hybridize to one or more of the first strands or their amplification products, or to one or more of the second strands or their amplification products, thereby producing nucleic acid hybrids; and (b) isolating the nucleic acid hybrids. Embodiment 15. The method of embodiment 14, wherein the one or more bait molecules hybridize to one or more of the first strands or their amplification products but not to a plurality of the second strands or their amplification products. Embodiment 16. The method of embodiment 15, wherein the separation occurs after a cytosine conversion treatment, and the one or more bait molecules hybridize to one or more of the first strands or their amplification products if one or more cytosines of the first strand undergo cytosine conversion. Embodiment 17. The method of embodiment 14, wherein the one or more bait molecules hybridize to one or more of the second strands or their amplification products but not to a plurality of the first strands or their amplification products. Embodiment 18. The separation is (a) combining one or more first bait molecules with a plurality of first and second strands, where the one or more first bait molecules preferentially hybridize to one or more of the first strands or their amplification products, thereby producing a first nucleic acid hybrid; (b) isolating the first nucleic acid hybrid; and (c) combining one or more second bait molecules with the plurality of first and second strands, where the one or more second bait molecules preferentially hybridize to one or more of the second strands or their amplification products, thereby producing second nucleic acid hybrids; (d) isolating the second nucleic acid hybrid. Embodiment 19. The method according to any one of embodiments 1 to 18, wherein a plurality of first and second strands or their amplification products are detected together. Embodiment 20. The method according to any one of embodiments 1 to 18, wherein the first and second strands or their amplification products are detected separately. Embodiment 21. The method of any one of embodiments 1 to 20, wherein the method comprises subjecting a plurality of first single-stranded DNA fragments to two or more rounds of primer extension prior to cytosine conversion in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragments, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides that includes a cytosine analog that is resistant to cytosine conversion. Embodiment 22. The method of any one of embodiments 1 to 21, further comprising attaching one or more nucleic acid adaptors to one or more of the first single-stranded DNA fragments. Embodiment 23. The method of any one of embodiments 1 to 22, further comprising attaching one or more nucleic acid adapters to one or more of the second strands. Embodiment 24. The method of embodiment 22 or embodiment 23, wherein one or more nucleic acid adapters are attached to the first strand or the second strand by ligation, transposition, tailing, or template switching. Embodiment 25. A method for detecting genetic and epigenetic information in a single workflow, comprising: providing a plurality of first single-stranded DNA fragments; attaching a first adaptor nucleic acid to at least a portion of a plurality of first single-stranded DNA fragments; subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragment downstream of or at the first adaptor nucleic acid, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides that include a cytosine analog that is resistant to cytosine conversion, thereby generating a plurality of first strands that include the first single-stranded DNA fragments with the first adaptor nucleic acid, and a plurality of second strands that include a second single-stranded DNA fragment complementary to the first single-stranded DNA fragments, and a second adaptor nucleic acid complementary to the first adaptor nucleic acid, wherein the second strands include a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that if an unmethylated cytosine is present in the first strand, it undergoes cytosine conversion, wherein after cytosine conversion, the sequences of the first and second adaptor nucleic acids are no longer complementary; detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands. Embodiment 26 The method of embodiment 25, wherein a first adaptor nucleic acid is attached to the 3'-end of at least a portion of the plurality of first single-stranded DNA fragments. Embodiment 27 The method of embodiment 25, wherein the first adaptor nucleic acid is attached to the 5' ends of at least a portion of the plurality of first single-stranded DNA fragments. Embodiment 28. The method of any one of embodiments 25 to 27, wherein the first adapter nucleic acid comprises one or more unmethylated cytosines that are converted during cytosine conversion. Embodiment 29. The method of any one of embodiments 25 to 27, wherein the first adapter nucleic acid comprises one or more methylated cytosines and the primer comprises one or more unmethylated cytosines that are converted during cytosine conversion. Embodiment 30. A method for detecting genetic and epigenetic information in a single workflow, comprising: providing a plurality of first single-stranded DNA fragments; subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer, the primer comprising (i) a second adapter nucleic acid portion that does not anneal to the first single-stranded DNA fragment and (ii) a portion that anneals to at least a portion of the first single-stranded DNA fragment, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides that comprise a cytosine analog that is resistant to cytosine conversion, thereby generating a plurality of first strands comprising first single-stranded DNA fragments having a first adapter nucleic acid complementary to the second adapter nucleic acid, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragments and a second adapter nucleic acid non-complementary to the first adapter nucleic acid, the second strands comprising a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that if an unmethylated cytosine is present in the first strand, it undergoes cytosine conversion, wherein after cytosine conversion, the sequences of the first and second adaptor nucleic acids are no longer complementary; detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands. Embodiment 31 The method of embodiment 30, wherein the second adapter nucleic acid portion of the primer is 5' to the portion of the primer that anneals to the first single-stranded DNA fragment. Embodiment 32. The method of embodiment 31, wherein the primer comprises one or more unmethylated cytosines that become part of the second adapter nucleic acid after primer extension and are converted during cytosine conversion. Embodiment 33. A method for detecting genetic and epigenetic information in a single workflow, comprising: providing a plurality of first single-stranded DNA fragments; subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer, the primer comprising: (i) a second adapter nucleic acid portion comprising a non-complementary 5' overhang that does not anneal to the first single-stranded DNA fragment; and (ii) a portion that anneals to at least a portion of the first single-stranded DNA fragment; (b) a nucleic acid polymerase; and (c) a mixture of nucleotides comprising a cytosine analog that is resistant to cytosine conversion, thereby generating a plurality of first strands comprising first single-stranded DNA fragments having a first adapter nucleic acid complementary to a portion of the second adapter nucleic acid portion that anneals to the first single-stranded DNA fragment, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragment and a second adapter nucleic acid, the second strands comprising a cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that if an unmethylated cytosine is present in the first strand, it undergoes cytosine conversion, wherein after cytosine conversion, the sequences of the first and second adaptor nucleic acids are no longer complementary; detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands. Embodiment 34. The method of any one of embodiments 25 to 33, wherein detecting is by sequencing. Embodiment 35. The method of embodiment 34, wherein the sequencing is next generation sequencing (NGS). Embodiment 36. The method of any one of embodiments 25 to 33, wherein the detecting is by microarray, quantitative PCR (qPCR), digital droplet PCR (ddPCR), or molecular inversion probes. Embodiment 37. The method of any one of embodiments 25 to 36, further comprising demultiplexing sequence information from the first and second strands based on the first and / or second adaptor nucleic acid. Embodiment 38. The method of any one of embodiments 25 to 37, wherein the method further comprises subjecting the plurality of first strands and / or the plurality of second strands to amplification prior to detecting, and wherein detecting comprises detecting at least a portion of the plurality of first and second strands and their amplification products, if present. Embodiment 39. The method of embodiment 38, comprising subjecting the plurality of second strands to amplification in the presence of one or more primers that anneal to at least a portion of the plurality of second strands but not to the plurality of first strands. Embodiment 40 The method of embodiment 39, wherein one or more primers that anneal to at least a portion of the plurality of second strands anneal to at least a portion of a second adaptor nucleic acid. Embodiment 41. The method of any one of embodiments 38 to 40, comprising subjecting the plurality of first strands to amplification in the presence of one or more primers that anneal to at least a portion of the plurality of first strands but not to the plurality of second strands. Embodiment 42 The method of embodiment 41, wherein one or more primers that anneal to at least a portion of the plurality of first strands anneal to at least a portion of a first adaptor nucleic acid. Embodiment 43. The method of embodiment 41, wherein the one or more primers anneal to at least a portion of the plurality of first strands after the first strands have undergone cytosine conversion but not before the first strands have undergone cytosine conversion. Embodiment 44. The method of embodiment 41, wherein the one or more primers anneal to at least a portion of the plurality of first strands only if a cytosine of the first strand does not undergo cytosine conversion. Embodiment 45. The method according to any one of embodiments 38 to 44, wherein the amplification occurs after cytosine conversion. Embodiment 46 The method according to any one of embodiments 25 to 45, wherein the method further comprises enriching the plurality of second strands or their amplification products prior to detection. Embodiment 47. The method according to any one of embodiments 25 to 46, wherein the method further comprises enriching the plurality of first strands or their amplification products prior to detection. Embodiment 48. The method according to embodiment 46 or embodiment 47, wherein the method comprises separating one or more first strands or their amplification products from one or more second strands or their amplification products prior to detection. Embodiment 49. The method of embodiment 48, wherein the separating comprises: (a) combining one or more bait molecules with a plurality of first and second strands, wherein one or more bait molecules preferentially hybridize to one or more of the first strands or their amplification products, or to one or more of the second strands or their amplification products, thereby producing nucleic acid hybrids; and (b) isolating the nucleic acid hybrids. Embodiment 50. The method of embodiment 49, wherein the one or more bait molecules hybridize to one or more of the first strands or their amplification products but not to a plurality of the second strands or their amplification products. Embodiment 51. The method of embodiment 49 or embodiment 50, wherein one or more bait molecules hybridize to at least a portion of the first adaptor nucleic acid. Embodiment 52. The method of embodiment 49 or embodiment 50, wherein the separation occurs after a cytosine conversion treatment, and one or more bait molecules hybridize to one or more of the first strands or their amplification products if one or more cytosines of the first strand undergo cytosine conversion. Embodiment 53. The method of embodiment 49, wherein the one or more bait molecules hybridize to one or more of the second strands or their amplification products but not to a plurality of the first strands or their amplification products. Embodiment 54. The method of embodiment 53, wherein the one or more bait molecules hybridize to at least a portion of the second adaptor nucleic acid. 55. The separation is (a) combining one or more first bait molecules with a plurality of first and second strands, where the one or more first bait molecules preferentially hybridize to one or more of the first strands or their amplification products, thereby producing a first nucleic acid hybrid; (b) isolating the first nucleic acid hybrid; and (c) combining one or more second bait molecules with the plurality of first and second strands, where the one or more second bait molecules preferentially hybridize to one or more of the second strands or their amplification products, thereby producing second nucleic acid hybrids; (d) isolating the second nucleic acid hybrid. Embodiment 56. The method of embodiment 55, wherein one or more first bait molecules hybridize to at least a portion of a first adaptor nucleic acid and / or one or more second bait molecules hybridize to at least a portion of a second adaptor nucleic acid. Embodiment 57. The method according to any one of embodiments 25 to 56, wherein a plurality of first and second strands or their amplification products are detected together. Embodiment 58. The method according to any one of embodiments 25 to 56, wherein the first and second strands or their amplification products are detected separately. Embodiment 59. The method of any one of embodiments 25 to 58, wherein the method comprises subjecting a plurality of first single-stranded DNA fragments to two or more rounds of primer extension prior to cytosine conversion in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragments, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides that includes a cytosine analog that is resistant to cytosine conversion. Embodiment 60. The method of any one of embodiments 1, 2, 5-34, and 37-59, wherein the plurality of first strands are sequenced at a different sequencing depth than the plurality of second strands. Embodiment 61. The method of embodiment 60, wherein the plurality of second strands are sequenced at a sequencing depth that is at least 2 times, at least 4 times, at least 8 times, at least 10 times, at least 50 times, at least 100 times, or at least 1000 times higher than the sequencing depth of the plurality of first strands. Embodiment 62. The method of embodiment 60, wherein the plurality of first strands are sequenced at a sequencing depth that is at least 2 times, at least 4 times, at least 8 times, at least 10 times, at least 50 times, at least 100 times, or at least 1000 times higher than the sequencing depth of the plurality of second strands. Embodiment 63. The method of any one of embodiments 1 to 62, wherein after primer extension, a plurality of first strands are hybridized to a plurality of second strands in a plurality of double-stranded nucleic acids. Embodiment 64. The method of any one of embodiments 1 to 63, further comprising denaturing the plurality of double-stranded DNA fragments prior to the round of primer extension to provide a plurality of first single-stranded DNA fragments. Embodiment 65. The method of any one of embodiments 1 to 64, further comprising obtaining a plurality of single-stranded or double-stranded DNA fragments from the sample. Embodiment 66 The method of embodiment 65, further comprising obtaining a sample from the individual. Embodiment 67. The method of embodiment 66, wherein the individual has cancer, is suspected of having cancer, or is undergoing treatment for cancer. Embodiment 68. The method of embodiment 66, wherein the individual is being screened for cancer or cancer recurrence. Embodiment 69. The method of any one of embodiments 65 to 68, wherein the sample comprises tissue, cells, and / or nucleic acid from a cancer. Embodiment 70. The method of any one of embodiments 65 to 69, wherein the sample comprises tissue, cells, and / or nucleic acid from normal tissue. Embodiment 71. The method of any one of embodiments 65 to 70, wherein the sample comprises a tissue biopsy sample, a liquid biopsy sample, or a normal control. Embodiment 72. The method of embodiment 71, wherein the sample is derived from a tumor biopsy, a tumor specimen, or circulating tumor cells. Embodiment 73. The method of any one of embodiments 65 to 70, wherein the sample is a liquid biopsy sample and comprises blood, plasma, serum, cerebrospinal fluid, sputum, stool, urine, or saliva. Embodiment 74. The method of any one of embodiments 1 to 73, wherein the cytosine analogs resistant to cytosine conversion include 5-methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), 5-carboxylcytosine (5caC), 5-formylcytosine, 5-(beta-D-glucosylmethyl)cytosine (5gmC), 5-ethyl dCTP, 5-methyl dCTP, 5-fluoro dCTP, 5-bromo dCTP, 5-iodo dCTP, 5-chloro dCTP, 5-trifluoromethyl dCTP, or 5-aza dCTP. Embodiment 75. The method of any one of embodiments 1 to 74, wherein the cytosine conversion is by bisulfite treatment, TET-assisted bisulfite treatment, oxidized bisulfite treatment, APOBEC, or TET / beta-glucosyltransferase-assisted APOBEC treatment. Embodiment 76. The method according to any one of embodiments 1 to 75, wherein if unmethylated cytosines are present in the first strand, at least 80%, at least 85%, at least 90%, at least 95%, or 100% undergo cytosine conversion as a result of the cytosine conversion treatment. Embodiment 77. The method according to any one of embodiments 1 to 76, wherein if unmethylated cytosines are present in the first strand, about 80% to about 97% undergo cytosine conversion as a result of the cytosine conversion treatment. Embodiment 78. The method of any one of embodiments 1 to 77, wherein up to 20%, up to 15%, up to 10%, up to 5%, up to 2%, up to 1%, or up to 0.5% of the cytosine analogs of the second strand undergo cytosine conversion as a result of the cytosine conversion treatment. Embodiment 79. The method according to any one of embodiments 1 to 78, wherein about 0.5% to about 5% of the cytosine analogs of the second strand undergo cytosine conversion as a result of the cytosine conversion treatment. Embodiment 80 The method of any one of embodiments 1 to 79, wherein the nucleic acid polymerase is capable of incorporating cytosine analogs into the nucleic acid. Embodiment 81 The method of any one of embodiments 1 to 80, further comprising subjecting the plurality of first single-stranded DNA fragments to end repair prior to primer extension. Embodiment 82. The method of any one of embodiments 1, 2, 5-34, and 37-81, further comprising, after sequencing the plurality of first and second strands or their amplification products, comparing the sequences of the plurality of second strands with the sequences of the plurality of first strands. Embodiment 83. The method of any one of embodiments 1, 2, 5-34, and 37-81, further comprising, after sequencing the plurality of first and second strands or their amplification products, comparing the plurality of first strand and / or second strand sequences to a reference genome sequence. Embodiment 84. A method for detecting cancer in an individual, comprising detecting methylation levels and / or somatic mutations in a sample comprising a plurality of nucleic acids obtained from the individual according to a method according to any one of embodiments 1 to 83, wherein the methylation levels and / or somatic mutations detected in the sample identify the individual as having cancer. Embodiment 85. A method for detecting minimal residual disease in an individual who has been treated for cancer or is being treated for cancer, comprising detecting methylation levels and / or somatic mutations in a sample comprising a plurality of nucleic acids obtained from the individual according to a method according to any one of embodiments 1 to 83, wherein the methylation levels and / or somatic mutations detected in the sample identify the individual as having minimal residual disease or lack thereof. Embodiment 86. A method for screening an individual suspected of having cancer, comprising detecting methylation levels and / or somatic mutations in a sample comprising a plurality of nucleic acids obtained from the individual according to a method according to any one of embodiments 1 to 83, wherein the methylation levels and / or somatic mutations detected in the sample identify the individual as likely to have cancer. Embodiment 87. A method for determining the prognosis of an individual having cancer, comprising detecting methylation levels and / or somatic mutations in a sample comprising a plurality of nucleic acids obtained from the individual according to a method according to any one of embodiments 1 to 83, wherein the methylation levels and / or somatic mutations detected in the sample at least partially determine the prognosis of the individual. Embodiment 88. A method for predicting survival time of an individual having cancer, comprising detecting methylation levels and / or somatic mutations in a sample comprising a plurality of nucleic acids obtained from the individual according to a method according to any one of embodiments 1 to 83, wherein the methylation levels and / or somatic mutations detected in the sample are at least partially predictive of survival time of the individual. Embodiment 89. A method for predicting or detecting tumor burden in an individual having cancer, comprising detecting methylation levels and / or somatic mutations in a sample comprising a plurality of nucleic acids obtained from the individual according to a method according to any one of embodiments 1 to 83, wherein the methylation levels and / or somatic mutations detected in the sample at least partially predict or detect the tumor burden of the individual. Embodiment 90. A method for predicting the responsiveness of an individual having cancer to a treatment, comprising detecting methylation levels and / or somatic mutations in a sample comprising a plurality of nucleic acids obtained from the individual according to a method according to any one of embodiments 1 to 83, wherein the methylation levels and / or somatic mutations detected in the sample are used at least in part to predict the individual's responsiveness to the treatment. Embodiment 91. A method for monitoring the response of an individual undergoing treatment for cancer, comprising: (a) administering a treatment to an individual having cancer; (b) detecting in a sample comprising a plurality of nucleic acids obtained from the individual a methylation level and / or a somatic mutation according to any one of embodiments 1 to 83, wherein the methylation level and / or the somatic mutation detected in the sample is used, at least in part, to monitor a response to a treatment. Embodiment 92. A method for monitoring cancer in an individual, comprising: (a) detecting methylation levels and / or somatic mutations in a first sample comprising a plurality of nucleic acids obtained from an individual according to a method according to any one of embodiments 1 to 83; (b) detecting methylation levels and / or somatic mutations in a second sample comprising a plurality of nucleic acids obtained from the individual according to the method of any one of embodiments 1 to 83; (c) determining differences in methylation levels and / or somatic mutations between the first sample and the second sample, thereby monitoring cancer in the individual. Embodiment 93. A system comprising: one or more processors; and a memory configured to store one or more computer program instructions, the one or more computer program instructions, when executed by the one or more processors, obtaining a first plurality of sequence reads of one or more first nucleic acid molecules or amplification products thereof, the first nucleic acid molecules being subjected to a cytosine conversion treatment under conditions such that unmethylated cytosines, if present in the first nucleic acid molecule, are converted to cytosines; obtaining a second plurality of sequence reads of one or more second nucleic acid molecules or amplification products thereof, the second nucleic acid molecules being complementary to the first nucleic acid molecule prior to cytosine conversion and comprising cytosine analogs that are resistant to cytosine conversion; analyzing the first plurality of sequence reads for the presence or absence of methylation at one or more cytosines inferred based on a cytosine conversion or lack thereof; and The system is configured to analyze the second plurality of sequence reads for sequence information. Embodiment 94. The system of embodiment 93, wherein the one or more first nucleic acid molecules or their amplification products further comprise a first adaptor nucleic acid sequence, and the one or more computer program instructions, when executed by the one or more processors, are further configured to demultiplex the first and second plurality of sequence reads based at least in part on detection of the first adaptor nucleic acid sequence in the first plurality of sequence reads. Embodiment 95. The system of embodiment 94, wherein the first adaptor nucleic acid is attached to the 3' end of at least a portion of the plurality of first nucleic acid molecules. Embodiment 96 The system of embodiment 94, wherein the first adaptor nucleic acid is attached to the 5' end of at least a portion of the plurality of first nucleic acid molecules. Embodiment 97. A system described in any one of embodiments 94 to 96, wherein the first adapter nucleic acid contains one or more unmethylated cytosines that are converted during cytosine conversion. Embodiment 98. A system described in any one of embodiments 94 to 96, wherein the first adapter nucleic acid comprises one or more methylated cytosines. Embodiment 99. The system of embodiment 94, wherein the first adaptor nucleic acid is between the 5' and 3' ends of at least a portion of the plurality of first nucleic acid molecules. Embodiment 100. The system of any one of embodiments 93 to 99, wherein the one or more second nucleic acid molecules or their amplification products further comprise a second adaptor nucleic acid sequence, and the one or more computer program instructions, when executed by the one or more processors, are further configured to demultiplex the first and second plurality of sequence reads based at least in part on detection of the second adaptor nucleic acid sequence in the second plurality of sequence reads. Embodiment 101. The system of embodiment 100, wherein a second adaptor nucleic acid is attached to the 3' end of at least a portion of the plurality of second nucleic acid molecules. Embodiment 102. The system of embodiment 100, wherein a second adaptor nucleic acid is attached to the 5' end of at least a portion of the plurality of second nucleic acid molecules. Embodiment 103. The system of embodiment 100, wherein the second adaptor nucleic acid is between the 5' and 3' ends of at least a portion of the plurality of second nucleic acid molecules. Embodiment 104. A system described in any one of embodiments 94 to 103, wherein the first plurality of sequence reads are at least 2-fold, at least 4-fold, at least 8-fold, at least 10-fold, at least 50-fold, at least 100-fold, or at least 1000-fold more abundant than the second plurality of sequence reads. Embodiment 105. A system described in any one of embodiments 94 to 103, wherein the second plurality of sequence reads are at least 2-fold, at least 4-fold, at least 8-fold, at least 10-fold, at least 50-fold, at least 100-fold, or at least 1000-fold more abundant than the first plurality of sequence reads. Embodiment 106. A system according to any one of embodiments 94 to 105, wherein the first nucleic acid molecule is obtained from a sample prior to cytosine conversion. Embodiment 107. The system of embodiment 106, wherein the sample is from an individual having or suspected of having cancer. Embodiment 108. The system of embodiment 107, wherein the sample comprises tissue, cells, and / or nucleic acid from a cancer. Embodiment 109. A system described in embodiment 107 or embodiment 108, wherein the sample comprises tissue, cells, and / or nucleic acid from normal tissue. Embodiment 110. A system described in any one of embodiments 106 to 109, wherein the first and / or second plurality of sequence reads are obtained by sequencing, and optionally, the sequencing includes the use of massively parallel sequencing (MPS) technology, whole genome sequencing (WGS), whole exome sequencing, targeted sequencing, direct sequencing, or Sanger sequencing technology, and optionally, the massively parallel sequencing technology includes next-generation sequencing (NGS). Embodiment 111. A system described in any one of embodiments 106 to 110, further configured to generate a molecular profile for the sample based at least in part on the analysis, wherein one or more program instructions, when executed by one or more processors, are configured to: Embodiment 112. The system of embodiment 111, wherein treatment is administered to the individual based at least in part on the molecular profile. Embodiment 113. The system described in embodiment 111 or embodiment 112, wherein the molecular profile further includes results from a comprehensive genomic profiling (CGP) test, a gene expression profiling test, a cancer hotspot panel test, a DNA methylation test, a DNA fragmentation test, an RNA fragmentation test, or any combination thereof. Embodiment 114. A system described in any one of embodiments 111 to 113, wherein the molecular profile further includes results from a nucleic acid sequencing-based test. Embodiment 115. A system described in any one of embodiments 94 to 114, further configured to compare one sequence read of the second plurality of sequence reads with one sequence read of the first plurality of sequence reads, when one or more computer program instructions are executed by one or more processors. Embodiment 116. A system described in any one of embodiments 94 to 114, further configured to compare one or more sequence reads of the first and / or second plurality of sequence reads with a reference genome sequence, wherein one or more computer program instructions, when executed by one or more processors, are configured to: Embodiment 117. A non-transitory computer-readable storage medium comprising one or more programs executable by one or more computer processors for implementing a method, obtaining a first plurality of sequence reads of one or more first nucleic acid molecules or amplification products thereof, the first nucleic acid molecules having been subjected to a cytosine conversion treatment under conditions such that unmethylated cytosines, if present in the first nucleic acid molecule, are subjected to cytosine conversion; obtaining a second plurality of sequence reads of one or more second nucleic acid molecules or amplification products thereof, the second nucleic acid molecules being complementary to the first nucleic acid molecule prior to cytosine conversion and comprising cytosine analogs that are resistant to cytosine conversion; analyzing the first plurality of sequence reads for the presence or absence of methylation at one or more cytosines inferred based on a cytosine conversion or lack thereof; and analyzing the second plurality of sequence reads for sequence information. Embodiment 118. The non-transitory computer-readable storage medium of embodiment 117, wherein the one or more first nucleic acid molecules or their amplification products further comprise a first adaptor nucleic acid sequence, and the method further comprises demultiplexing the first and second plurality of sequence reads based at least in part on detection of the first adaptor nucleic acid sequence in the first plurality of sequence reads. Embodiment 119. The non-transitory computer-readable storage medium of embodiment 118, wherein the first adaptor nucleic acid is attached to the 3' end of at least a portion of the plurality of first nucleic acid molecules. Embodiment 120. The non-transitory computer-readable storage medium of embodiment 118, wherein the first adaptor nucleic acid is attached to a 5' end of at least a portion of the plurality of first nucleic acid molecules. Embodiment 121. A non-transitory computer-readable storage medium according to any one of embodiments 118 to 120, wherein the first adapter nucleic acid comprises one or more unmethylated cytosines that are converted during cytosine conversion. Embodiment 122. A non-transitory computer-readable storage medium according to any one of embodiments 118 to 120, wherein the first adaptor nucleic acid comprises one or more methylated cytosines. Embodiment 123. The non-transitory computer-readable storage medium of embodiment 118, wherein the first adaptor nucleic acid is between the 5' end and the 3' end of at least a portion of the plurality of first nucleic acid molecules. Embodiment 124. A non-transitory computer-readable storage medium according to any one of embodiments 117 to 123, wherein the one or more second nucleic acid molecules or their amplification products further comprise a second adaptor nucleic acid sequence, and the method further comprises demultiplexing the first and second plurality of sequence reads based at least in part on detection of the second adaptor nucleic acid sequence in the second plurality of sequence reads. Embodiment 125. The non-transitory computer-readable storage medium of embodiment 124, wherein a second adaptor nucleic acid is attached to the 3' end of at least a portion of the plurality of second nucleic acid molecules. Embodiment 126. The non-transitory computer-readable storage medium of embodiment 124, wherein a second adaptor nucleic acid is attached to the 5' end of at least a portion of the plurality of second nucleic acid molecules. Embodiment 127. The non-transitory computer-readable storage medium of embodiment 124, wherein the second adaptor nucleic acid is between the 5' and 3' ends of at least a portion of the plurality of second nucleic acid molecules. Embodiment 128. A non-transitory computer-readable storage medium according to any one of embodiments 117 to 127, wherein the first plurality of sequence reads are at least 2-fold, at least 4-fold, at least 8-fold, at least 10-fold, at least 50-fold, at least 100-fold, or at least 1000-fold more abundant than the second plurality of sequence reads. Embodiment 129. A non-transitory computer-readable storage medium described in any one of embodiments 117 to 127, wherein the second plurality of sequence reads are at least 2 times, at least 4 times, at least 8 times, at least 10 times, at least 50 times, at least 100 times, or at least 1000 times more abundant than the first plurality of sequence reads. Embodiment 130. A non-transitory computer-readable storage medium according to any one of embodiments 117 to 129, wherein the first nucleic acid molecule is obtained from a sample prior to cytosine conversion. Embodiment 131. A non-transitory computer-readable storage medium according to embodiment 130, wherein the sample is from an individual having or suspected of having cancer. Embodiment 132. A non-transitory computer-readable storage medium according to embodiment 131, wherein the sample comprises tissue, cells, and / or nucleic acid from a cancer. Embodiment 133. A non-transitory computer-readable storage medium according to embodiment 131 or embodiment 132, wherein the sample comprises tissue, cells, and / or nucleic acid from normal tissue. Embodiment 134. A non-transitory computer-readable storage medium described in any one of embodiments 130 to 133, wherein the first and / or second plurality of sequence reads are obtained by sequencing, and optionally, the sequencing includes the use of massively parallel sequencing (MPS) technology, whole genome sequencing (WGS), whole exome sequencing, targeted sequencing, direct sequencing, or Sanger sequencing technology, and optionally, the massively parallel sequencing technology includes next-generation sequencing (NGS). Embodiment 135. A non-transitory computer-readable storage medium according to any one of embodiments 130 to 134, wherein the method further comprises generating a molecular profile for the sample based at least in part on the analysis. Embodiment 136. The non-transitory computer-readable storage medium of embodiment 135, wherein treatment is administered to the individual based at least in part on the molecular profile. Embodiment 137. The non-transitory computer-readable storage medium of embodiment 135 or embodiment 136, wherein the molecular profile further comprises results from a comprehensive genomic profiling (CGP) test, a gene expression profiling test, a cancer hotspot panel test, a DNA methylation test, a DNA fragmentation test, an RNA fragmentation test, or any combination thereof. Embodiment 138. A non-transitory computer-readable storage medium according to any one of embodiments 135 to 137, wherein the molecular profile further comprises results from a nucleic acid sequencing-based test. Embodiment 139. A non-transitory computer-readable storage medium described in any one of embodiments 117 to 138, wherein the method further comprises comparing one sequence read of the second plurality of sequence reads with one sequence read of the first plurality of sequence reads. Embodiment 140. A non-transitory computer-readable storage medium described in any one of embodiments 117 to 138, wherein the method further comprises comparing one or more sequence reads of the first and / or second plurality of sequence reads to a reference genome sequence. Embodiment 141. A method for detecting genetic and epigenetic information in a single workflow, comprising: providing a plurality of first single-stranded DNA fragments; subjecting a plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragments, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides that include a cytosine analog that is resistant to TET-assisted pyridine borane treatment, thereby generating a plurality of first strands corresponding to the first single-stranded DNA fragments and a plurality of second strands that are complementary to the first strands, wherein the second strands include a cytosine analog that is resistant to TET-assisted pyridine borane treatment; subjecting a plurality of the first and second strands to TET-assisted pyridine borane treatment under conditions such that any methylated cytosines present in the first strand undergo cytosine conversion; detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands. Embodiment 142. A method for detecting genetic and epigenetic information in a single workflow, comprising: providing a plurality of first single-stranded DNA fragments; attaching a first adaptor nucleic acid to at least a portion of a plurality of first single-stranded DNA fragments; subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragment downstream of or at the first adapter nucleic acid, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides that include a cytosine analog that is resistant to TET-assisted pyridine borane treatment, thereby generating a plurality of first strands that include the first single-stranded DNA fragments with the first adapter nucleic acid, and a second single-stranded DNA fragment that is complementary to the first single-stranded DNA fragments, and a plurality of second strands that include a second adapter nucleic acid that is complementary to the first adapter nucleic acid, the second strands including a cytosine analog that is resistant to TET-assisted pyridine borane treatment; subjecting the plurality of first and second strands to TET-assisted pyridine borane treatment under conditions such that any methylated cytosines present in the first strand undergo cytosine conversion, wherein following cytosine conversion, the sequences of the first and second adaptor nucleic acids are no longer complementary; detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands. Embodiment 143. A method for detecting genetic and epigenetic information in a single workflow, comprising: providing a plurality of first single-stranded DNA fragments; subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer, the primer comprising: (i) a second adapter nucleic acid portion comprising a non-complementary 5' overhang that does not anneal to the first single-stranded DNA fragment; and (ii) a portion that anneals to at least a portion of the first single-stranded DNA fragment; (b) a nucleic acid polymerase; and (c) a mixture of nucleotides comprising a cytosine analog that is resistant to TET-assisted pyridine borane treatment, thereby generating a plurality of first strands comprising first single-stranded DNA fragments having a first adapter nucleic acid complementary to a portion of the second adapter nucleic acid portion that anneals to the first single-stranded DNA fragment, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragment and a second adapter nucleic acid, the second strands comprising a cytosine analog that is resistant to TET-assisted pyridine borane treatment; subjecting a plurality of the first and second strands to TET-assisted pyridine borane treatment under conditions such that any methylated cytosines present in the first strand undergo cytosine conversion; detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands. Embodiment 144. A method for detecting genetic and epigenetic information in a single workflow, comprising: providing a plurality of first single-stranded DNA fragments; subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer, the primer comprising (i) a second adapter nucleic acid portion that does not anneal to the first single-stranded DNA fragment and (ii) a portion that anneals to at least a portion of the first single-stranded DNA fragment, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides comprising a cytosine analog that is resistant to TET-assisted pyridine borane treatment, thereby generating a plurality of first strands comprising a first single-stranded DNA fragment having a first adapter nucleic acid complementary to the second adapter nucleic acid, and a plurality of second strands comprising a second single-stranded DNA fragment complementary to the first single-stranded DNA fragment and a second adapter nucleic acid non-complementary to the first adapter nucleic acid, the second strands comprising a cytosine analog that is resistant to TET-assisted pyridine borane treatment; subjecting the plurality of first and second strands to TET-assisted pyridine borane treatment under conditions such that any methylated cytosines present in the first strand undergo cytosine conversion, wherein following cytosine conversion, the sequences of the first and second adaptor nucleic acids are no longer complementary; detecting at least a portion of the plurality of first strands and at least a portion of the plurality of second strands.
[0180] The disclosures of all publications, patents, and patent applications referenced herein are each incorporated herein by reference in their entirety. To the extent that a reference incorporated by reference conflicts with the present disclosure, the present disclosure shall control. EXAMPLES
[0181] The present invention will be more fully understood by referring to the following examples. However, they should not be interpreted as limiting the scope of the present invention. It is understood that the examples and embodiments described herein are for illustrative purposes only, and various modifications or changes will be suggested to those skilled in the art in light thereof and are included within the spirit and scope of this application and the appended claims.
[0182] Example 1: Proof of concept of genetic and epigenetic information in a single workflow This example provides a proof-of-concept analysis of genetic (e.g., sequence) and epigenetic (e.g., methylation) information in a single workflow. Specifically, this example shows that the "genomic strand" (versus the sequence information) can be synthesized from the original DNA with high efficiency and protected from cytosine conversion.
[0183] Materials and Methods Gene libraries were prepared from 20 ng of methylated and unmethylated HCT116 human cell lines using the NEBNext Ultra II Library Preparation Kit according to standard procedures. Full-length methylated adapters containing a 4-base "work-stream" tag (GGCC) between the P7 sequence and the i7 index were ligated to the DNA. One round of primer extension using NEB Q5 polymerase and Zymo dNTP mix with 5-methyl-dCTP was performed, followed by enzymatic cytosine conversion using the NEBNext® Enzymatic Methyl-seq Kit. Finally, the converted libraries are PCR amplified with NEB Q5U polymerase and P5 / P7 primers. After PCR amplification, samples were normalized to 1 nM and subjected to NGS sequencing on a Novaseq, achieving an average coverage of 52x. An overview of the workflow used is provided in Figure 5A.
[0184] result Following the workflow shown in Figure 5A, the proportions of "genomic strands" (e.g., incorporating conversion-resistant cytosine analogs) and their amplification products, and "methylated strands" (e.g., free of cytosine analogs and containing any present DNA methylation) and their amplification products were determined by the percentage of identified reads. A 4 bp workstream index was used to demultiplex the reads to distinguish between genomic and methylated strands and their amplification products as a result of cytosine conversion.
[0185] As shown in Figure 5B, both the genomic strand and the methylated strand were identified via sequencing based on the sequence reads identified from each. The genomic strand was synthesized via the primer extension step with an efficiency of 80%, thereby indicating that this strand was synthesized with high efficiency during primer extension.
[0186] Figure 5C shows the results of cytosine conversion. The genomic strand was largely preserved after cytosine conversion with a protection efficiency of about 98.7%. In contrast, the methylated strand was preserved with an efficiency of about 96.8%. These results indicate that the genomic strand and its amplification product were protected from cytosine conversion, thereby preserving sequence information.
[0187] Example 2: Detection of genetic and epigenetic variants in cancer cfDNA in a single assay This example shows that cancer-associated methylation variants and gene variants can be detected in cfDNA using a single workflow methodology. Moreover, a single workflow from the same input mass of DNA can generate comparable signals compared to independent methods (e.g., WGS and WG methylation sequencing).
[0188] Materials and Methods Four cancer cfDNA samples with characterized gene variants and two non-cancerous cfDNA samples were analyzed. Using an input of 20ng cfDNA, each sample was tested with the following assays: standard WGS, standard WG methylation (using enzymatic methylation sequencing), and single workflow protocol. Standard WGS libraries were prepared using the NEBNext Ultra II Library Preparation Kit according to the procedure and PCR amplified using NEB Q5 polymerase and NEB Unique Dual Index Primer. Standard WG methylation libraries were prepared using the NEBNext® Enzymatic Methyl-seq Kit according to the standard procedure and PCR amplified using NEB Q5U polymerase and NEB Unique Dual Index Primer. Single workflow libraries were prepared using the protocol described herein (using enzymatic methylation sequencing for cytosine conversion and 5-mC for cytosine analogs) as described in Example 1. After PCR amplification, libraries were normalized to 1nM and subjected to NGS sequencing on Novaseq, achieving an average coverage of 109 times.
[0189] result First, we evaluated the protection efficiency between the genomic strand and the methylated strand in a single workflow from cfDNA. The genomic strand refers to the cytosine-converted protected copy of the original molecule generated with 5-mC. As shown in Figure 6, the cytosines were 99.05% protected from enzymatic conversion (average from six libraries), allowing the preservation of genetic information. The methylated strand is preserved and contains methylation information, as shown by a protection efficiency of 4.3% (average from six libraries).
[0190] Next, the synthesis of genomic strands was evaluated. The average percent reads identified for genomic strands and methylated strands across the six samples was 53.12 and 46.88, respectively (Figure 7). These results indicate that primer extension is highly efficient in synthesizing genomic strands from the original DNA molecules. Reads in the library were identified using a four-base workstream index that distinguishes between genomic strands and methylated strands and their amplification products.
[0191] Methylation of cancer cfDNA samples was analyzed via WG enzymatic methylation sequencing compared to the single workflow methodology. The cancer methylation score, which evaluates consensus methylation sites from individual DNA molecules, was strongly correlated in the single workflow method and the WG enzymatic methylation sequencing method, as shown in Figure 8. The best-fit line had a slope of 1.1 and R2 = 1. Higher correlation in methylation levels was observed between the single workflow methodology and the standard WG methylation methodology in smaller bins and functional regions, including 1 kb bins (Figure 9A), 10 kb bins (Figure 9B), 100 kb bins (Figure 9C), CpG islands (Figure 9D), and CpG shores (Figure 9E). Sex chromosomes were not included. No filtering was applied to the bins. The average methylation percentage was also preserved in the single workflow technique versus the standard WG methylation analysis. The Pearson correlation coefficients of the mean methylation fraction (AMF) in the single workflow compared to standard WG enzymatic methylation sequencing are shown in Figure 10. The results show high concordance between methodologies, with r > 0.975 in the smaller bin sizes and functional regions. These results indicate that the methylation analysis using the single workflow technique was highly concordant with the methylation analysis obtained using standard techniques such as WG enzymatic methylation sequencing.
[0192] Next, we characterized the analysis of genetic variants using a single workflow. Table A shows the variants analyzed. [Table 1]
[0193] High concordance in called allele frequencies was observed between the single workflow and WGS in four cancer cfDNA samples (Figure 11). Sixteen characterized variants with allele frequencies (AF) above 3% were shared between the single workflow and standard WG techniques. These results indicate that genomic signal from the single workflow methodology was preserved, with strong concordance in the detection of genetic variants.
[0194] In conclusion, the single workflow method allows for multi-omic detection of methylation and genetic information in one simplified workflow with limited input DNA required. Results from this example showed that the assay works at a technical level in cfDNA. Genomic strands were synthesized from original DNA / methylated DNA with high efficiency and were protected during cytosine conversion. Signals were also preserved in the single workflow compared to standard independent methods, as shown by the strong correlation in both cancer-associated methylation scores and genetic variant calls. The single workflow assay also allowed for C->T variant calling despite exposure of genomic strands to cytosine conversion conditions.
[0195] Example 3: Enrichment of genomic and / or methylation strand information in a single workflow assay This example shows that by using primers designed to specifically bind to the genomic or methylated strand during PCR amplification, both the genomic and methylated strands can be amplified during a single workflow assay.
[0196] Materials and Methods 20 ng of input DNA was used from methylated and unmethylated HCT116 human cell lines. Single workflow analysis was performed as described in Examples 1 and 2. Two primer conditions were tested by PCR: (1) standard single workflow primers as described above, and (2) a primer designed to specifically bind to the genomic strand and a P5 primer. The abundance of the genomic strand versus the methylated strand was determined using the % sequence reads as described above. Subsequent assays used a primer designed to specifically bind to the methylated strand and a P5 primer as well. The sequence used for the genomic strand specific primer was CAA GCA GAA GAC GGC ATA CGA GAT TT (SEQ ID NO: 1). The sequence used for the methylated strand specific primer was CAA GCA GAA GAC GGC ATA CGA GAT CC (SEQ ID NO: 2).
[0197] result Preferential amplification of the genomic strand was achieved using genomic strand-specific primers. Across the two samples, the average read% on the genomic strand was 98.7% and 1.3% on the methylated strand, as shown in Figure 12. However, using standard single workflow primers, the average read% on the genomic strand was 50.5% and 49.5% on the methylated strand.
[0198] Using primers specific for either the genomic or methylated strand, preferential amplification and enrichment of either strand can be achieved. Figure 13 shows that this strategy has been successfully used to selectively amplify either strand. Thus, specific primers can be used to amplify either or both the genomic and methylated strands, for example, for enrichment prior to sequencing.
[0199] This example shows that hybrid capture baits complementary to genomic regions of interest and capture baits complementary to cytosine-conserved methylated regions of interest can be captured together from a single workflow library.
[0200] A custom hybrid capture bait targeting actionable mutations in the genomic strand was obtained (2.2Mb). A separate set of custom baits complementary to the cytosine conversion methylation biomarker of interest was obtained separately (5.52Mb). The genomic panel was mixed with the custom methylation panel in an equimolar ratio to create a single-workflow hybrid capture panel. 10ng of cfDNA isolated from a healthy donor was input into the single-workflow library preparation as described in Examples 1 and 2. 500ng of the resulting library was captured using the custom single-workflow panel using the capture kit.
[0201] After sequencing, the captured libraries were aligned to their respective genomes or methylated genomes, fragments were de-duplicated by fragment start / end position, and the average specific coverage across the baited regions was quantified. High specific coverage across the targeted genomic regions (1598.9X) and methylated regions (1510.5X) was observed, as depicted in Figure 14A. Given the low cfDNA input mass, this result shows high capture efficiency of both strands in the combined capture from a single workflow library (approximately 50% given an input haploid genome equivalent of 3030). The capture also showed high on-target rates of both the genomic regions of interest and the methylated regions, as well as uniform coverage across the separate probe sets, as measured by the Fold-80 base penalty in Figure 14B.
Claims
1. 1. A method for detecting genetic and epigenetic information in a single workflow, comprising: providing a plurality of first single-stranded DNA fragments; subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragments, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides that contain a cytosine analog that is resistant to cytosine conversion, thereby generating a plurality of first strands corresponding to the first single-stranded DNA fragments and a plurality of second strands that are complementary to the first strands, wherein the second strands contain the cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that unmethylated cytosines present in the first strand undergo cytosine conversion; detecting methylation information from at least a portion of said plurality of first strands and detecting sequence information from at least a portion of said plurality of second strands; A method comprising:
2. 10. The method of claim 1, wherein the detecting is by sequencing, microarray, quantitative PCR (qPCR), digital droplet PCR (ddPCR), next generation sequencing (NGS), or molecular inversion probes.
3. 10. The method of claim 1, wherein the method further comprises subjecting the plurality of first strands and / or the plurality of second strands to amplification prior to detecting, and wherein the detecting comprises detecting at least a portion of the plurality of first and second strands and their amplification products, if present. (a) subjecting said plurality of second strands to amplification in the presence of one or more primers that anneal to at least a portion of said plurality of second strands but not to said plurality of first strands; and / or (b) subjecting the plurality of first strands to amplification in the presence of one or more primers that anneal to at least a portion of the plurality of first strands but not to the plurality of second strands. The method of claim 3, comprising:
5. 10. The method of claim 1, wherein the method further comprises enriching the plurality of second strands or their amplification products and / or enriching the plurality of first strands or their amplification products prior to detecting.
6. 6. The method of claim 5, wherein the method comprises separating one or more first strands or amplification products thereof from one or more second strands or amplification products thereof prior to detecting.
7. Separation, (1) (a) combining one or more bait molecules with the plurality of first and second strands, wherein the one or more bait molecules preferentially hybridize to one or more of the first strands or their amplification products, or to one or more of the second strands or their amplification products, thereby producing nucleic acid hybrids; and (b) isolating the nucleic acid hybrids; or (2) (a) combining one or more first bait molecules with the plurality of first and second strands, wherein the one or more first bait molecules preferentially hybridize to one or more of the first strands or their amplification products, thereby producing first nucleic acid hybrids; (b) isolating the first nucleic acid hybrids; (c) combining one or more second bait molecules with the plurality of first and second strands, wherein the one or more second bait molecules preferentially hybridize to one or more of the second strands or their amplification products, thereby producing second nucleic acid hybrids; and (d) isolating the second nucleic acid hybrids. The method of claim 6, comprising:
8. 2. The method of claim 1, wherein the method comprises subjecting the plurality of first single-stranded DNA fragments to two or more rounds of primer extension prior to cytosine conversion in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragments, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides that includes a cytosine analog that is resistant to cytosine conversion.
9. 10. The method of claim 1, further comprising attaching one or more nucleic acid adaptors to one or more of the first single-stranded DNA fragments.
10. 1. A method for detecting genetic and epigenetic information in a single workflow, comprising: providing a plurality of first single-stranded DNA fragments; attaching a first adaptor nucleic acid to at least a portion of the plurality of first single-stranded DNA fragments; subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer that anneals to at least a portion of the first single-stranded DNA fragment downstream of or at the first adaptor nucleic acid, (b) a nucleic acid polymerase, and (c) a mixture of nucleotides comprising a cytosine analog that is resistant to cytosine conversion, thereby generating a plurality of first strands comprising the first single-stranded DNA fragments with the first adaptor nucleic acid, and a plurality of second strands comprising second single-stranded DNA fragments complementary to the first single-stranded DNA fragments, and a second adaptor nucleic acid complementary to the first adaptor nucleic acid, wherein the second strands comprise the cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that if an unmethylated cytosine is present in the first strand, it undergoes cytosine conversion, wherein after cytosine conversion, the sequences of the first and second adaptor nucleic acids are no longer complementary; detecting methylation information from at least a portion of said plurality of first strands and detecting sequence information from at least a portion of said plurality of second strands; A method comprising:
11. (a) the first adapter nucleic acid comprises one or more unmethylated cytosines that are converted during cytosine conversion; or (b) the first adapter nucleic acid comprises one or more methylated cytosines and the primer comprises one or more unmethylated cytosines that are converted during cytosine conversion; The method of claim 10.
12. 1. A method for detecting genetic and epigenetic information in a single workflow, comprising: providing a plurality of first single-stranded DNA fragments; subjecting the plurality of first single-stranded DNA fragments to one round of primer extension in the presence of (a) a primer, the primer comprising: (i) a second adapter nucleic acid portion that does not anneal to the first single-stranded DNA fragment; and (ii) a portion that anneals to at least a portion of the first single-stranded DNA fragment; (b) a nucleic acid polymerase; and (c) a mixture of nucleotides comprising a cytosine analog that is resistant to cytosine conversion, thereby generating a plurality of first strands comprising the first single-stranded DNA fragments having a first adapter nucleic acid that is complementary to the second adapter nucleic acid, and a plurality of second strands comprising second single-stranded DNA fragments complementary to the first single-stranded DNA fragments and the second adapter nucleic acid that is non-complementary to the first adapter nucleic acid, the second strands comprising the cytosine analog that is resistant to cytosine conversion; subjecting the plurality of first and second strands to a cytosine conversion treatment under conditions such that if an unmethylated cytosine is present in the first strand, it undergoes cytosine conversion, wherein after cytosine conversion, the sequences of the first and second adaptor nucleic acids are no longer complementary; detecting methylation information from at least a portion of said plurality of first strands and detecting sequence information from at least a portion of said plurality of second strands; A method comprising:
13. providing a plurality of first single-stranded DNA fragments; the second adapter nucleic acid portion of the primer comprises a non-complementary 5' overhang that does not anneal to the first single-stranded DNA fragment, and the plurality of first strands comprise the first single-stranded DNA fragments having the first adapter nucleic acid complementary to the portion of the second adapter nucleic acid portion that anneals to the first single-stranded DNA fragment. The method of claim 12.
14. 11. The method of claim 10, further comprising demultiplexing sequence information from the first and second strands based on the first and / or second adapter nucleic acids.
15. 11. The method of claim 10, wherein the method further comprises subjecting the plurality of first strands and / or the plurality of second strands to amplification prior to detecting, and wherein the detecting comprises detecting at least a portion of the plurality of first and second strands and their amplification products, if present.
16. 2. The method of claim 1, wherein the plurality of first strands are sequenced at a different sequencing depth than the plurality of second strands.
17. 10. The method of claim 1, further comprising obtaining the plurality of single-stranded or double-stranded DNA fragments from a sample, wherein the sample comprises tissue, cells, and / or nucleic acids from cancer and / or normal tissue.
18. 2. The method of claim 1, wherein the cytosine analogs that are resistant to cytosine conversion comprise 5-methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), 5-carboxylcytosine (5caC), 5-formylcytosine, 5-(beta-D-glucosylmethyl)cytosine (5gmC), 5-ethyl dCTP, 5-methyl dCTP, 5-fluoro dCTP, 5-bromo dCTP, 5-iodo dCTP, 5-chloro dCTP, 5-trifluoromethyl dCTP, or 5-aza dCTP.
19. 2. The method of claim 1, wherein the cytosine conversion is by bisulfite treatment, TET-assisted bisulfite treatment, oxidized bisulfite treatment, APOBEC, or TET / beta-glucosyltransferase-assisted APOBEC treatment.
20. The method of claim 1 , wherein the nucleic acid polymerase is capable of incorporating the cytosine analog into a nucleic acid.
21. The method described in claim 1, wherein the plurality of first single-stranded DNA fragments are from a sample obtained from an individual.
22. Determining a methylation level of one or more genes in the sample based on the detected methylation information; detecting one or more mutations in one or more genes in the sample based on the detected sequence information; determining whether the individual has cancer based at least in part on the methylation levels and / or mutations determined in the sample; and 22. The method of claim 21, comprising:
23. The method described in claim 22, wherein the one or more mutations are somatic mutations.
24. Determining a methylation level of one or more genes in the sample based on the detected methylation information; detecting one or more mutations in one or more genes in the sample based on the detected sequence information; detecting a response by said individual to a treatment based at least in part on said methylation level and / or mutation determined in said sample; 22. The method of claim 21, comprising:
25. Determining a methylation level of one or more genes in the sample based on the detected methylation information; detecting one or more mutations in one or more genes in the sample based on the detected sequence information; detecting minimal residual disease or loss thereof in said sample based at least in part on said methylation and / or mutation determined in said sample; 22. The method of claim 21, comprising: