TET-assisted pyridine borane sequencing
TET-assisted pyridine borane sequencing (TAPS) addresses the limitations of bisulfite sequencing by introducing DHU residues and using tolerant polymerases, enhancing sequencing quality and coverage of highly methylated regions, particularly in low-input samples.
Patent Information
- Application Number
- JP2025519962
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-04-19
- Filing Date
- 2023-10-03
- Publication Date
- 2025-10-09
AI Technical Summary
Current bisulfite sequencing methods, such as bisulfite sequencing and its derivatives, face challenges with DNA degradation in low-input samples and reduced sequencing quality due to harsh chemical reactions, incomplete conversion of cytosines, and false detection of 5mC and 5hmC, which limits their application in clinical and basic research.
The development of TET-assisted pyridine borane sequencing (TAPS) methods that introduce dihydrouracil (DHU) residues into nucleic acids, using polymerases tolerant to DHU residues for amplification, followed by exponential amplification, to create sequencing libraries without bisulfite treatment.
TAPS methods improve sequencing quality and coverage of highly methylated regions, reducing DNA degradation and false positives, making them suitable for low-input samples and clinical applications.
Smart Images

Figure 2025533890000001_ABST
Abstract
Description
[Technical Field]
[0001] This disclosure provides compositions and methods related to TET-assisted pyridine borane sequencing (TAPS). In particular, this disclosure provides optimized methods for generating and sequencing TAPS libraries. [Background technology]
[0002] 5-Methylcytosine (5mC) and 5-hydroxymethylcytosine (5hmC) are two major epigenetic marks found in mammalian genomes. 5hmC is generated from 5mC by dioxygenases of the ten-eleven translocation (TET) family. TET further oxidizes 5hmC to 5-formylcytosine (5fC) and 5-carboxylcytosine (5caC), which are present in much lower abundance in mammalian genomes compared to 5mC and 5hmC (10-100 times lower than 5hmC). 5mC and 5hmC work together to play important roles in a wide range of biological processes, from gene regulation to normal development. Aberrant DNA methylation and hydroxymethylation have been associated with various diseases and are widely recognized as hallmarks of cancer. Therefore, determining 5mC and 5hmC in DNA sequences is not only important for basic research but also useful for clinical applications, including diagnosis and therapy.
[0003] The current representative and most widely used method for DNA methylation and hydroxymethylation analysis is bisulfite sequencing (BS) and its derivatives (e.g., TET-assisted bisulfite sequencing (TAB-Seq) and bisulfite sequencing (oxBS). Similarly, bisulfite sequencing is the most established method for assaying whole genome DNA methylation. Both of these methods use bisulfite treatment to convert unmethylated cytosines to uracil, while leaving 5mC and / or 5hmC intact. Through PCR amplification of bisulfite-treated DNA (which reads uracil as thymine), modification information for each cytosine can be obtained at single-base resolution (converting C to T provides the position of unmethylated cytosines). However, bisulfite sequencing has at least two major drawbacks: First, bisulfite treatment is a harsh chemical reaction that degrades over 90% of DNA due to depurination under the required acidic and thermal conditions. This degradation severely limits its application to low-input samples, such as circulating cell-free DNA and clinical samples, including single-cell sequencing. Second, bisulfite sequencing relies on the complete conversion of unmodified cytosines to thymines. Unmodified cytosines account for approximately 95% of all cytosines in the human genome. Converting all these positions to thymine reduces sequence complexity, reduces sequencing quality, decreases mapping rates, leads to uneven genome coverage, increases sequencing costs, and reduces the ability to call variants. Bisulfite sequencing methods are also prone to false detection of 5mC and 5hmC due to incomplete conversion of unmodified cytosines to thymine.
[0004] Sequencing of DNA samples that have been treated to change natural bases can be difficult, especially when using massively parallel next-generation sequencing (NGS).In particular, problems arise when highly methylated regions of interest are underrepresented.The present invention provides a solution to this problem. Summary of the Invention
[0005] Embodiments of the present disclosure include methods for sequencing libraries after introducing dihydrouracil (DHU) residues by methods such as TET-assisted pyridine borane sequencing (TAPS) and variants of TAPS, including TAPS with β-glycosylation blocking (TAPSβ) and chemically assisted pyridine borane sequencing (CAPS). According to these embodiments, the methods include introducing DHU residues into a nucleic acid sample and preparing a sequencing library by a synthesis step using a first polymerase or polymerase mixture that is tolerant to the DHU residues and / or the products resulting from the introduction of DHU residues and / or the TAPS process, followed by exponential amplification.
[0006] Thus, in some embodiments, the present invention provides methods for amplifying a target nucleic acid molecule comprising a dihydrouracil (DHU) residue, the method comprising: synthesizing one or more complementary strands of the target nucleic acid comprising a DHU residue using a first polymerase or polymerase mixture that is tolerant to the DHU residue and / or to products resulting from the introduction of the DHU residue and / or the TAPS process, to obtain a target nucleic acid mixture comprising the target nucleic acid comprising the DHU residue and one or more complementary strands; and exponentially amplifying the target nucleic acid mixture to obtain amplified target nucleic acids.
[0007] In some embodiments, the first polymerase or polymerase mixture is 5.0 x 10 -5In some embodiments, the first polymerase or polymerase mixture is selected from the group consisting of Bst3.0 polymerase, Sulfolobus polymerase IV, a combination of Bst3.0 polymerase and Sulfolobus polymerase IV, Klenow polymerase, Klenow exopolymerase, Polκ polymerase, Mu-mLV reverse transcriptase, SD polymerase, Tth polymerase, OneTaq polymerase, a combination of OneTaq and Tth polymerase, 5D4 polymerase, a 5D4 polymerase blend with Taq polymerase, and SD polymerase.
[0008] In some embodiments, the first polymerase is thermolabile. In some embodiments, the first polymerase is thermostable. In some embodiments, the step of exponentially amplifying the complementary strand of the target nucleic acid utilizes a first polymerase or polymerase mixture that is tolerant to DHU residues and / or products resulting from the introduction of DHU residues and / or the TAPS process.
[0009] In some embodiments, the step of exponentially amplifying the pre-amplified target nucleic acid utilizes a second polymerase or polymerase. In some embodiments, the second polymerase or polymerase is 5.0 x 10 5 In some embodiments, the second polymerase or polymerase mixture has an error rate of less than 1.0 x 10 -6 In some embodiments, the second polymerase has an error rate of less than 5.0 x 10. In some embodiments, the second polymerase is selected from the group consisting of GoTaq polymerase and KAPA HiFi Uracil+ polymerase. -5 Polymerases with error rates less than 0.05 are thermostable.
[0010] In some embodiments, the first polymerase and the second polymerase are provided in a master mix.
[0011] In some embodiments, synthesizing the complementary strand of a target nucleic acid containing a DHU residue with a first polymerase or polymerase mixture further includes synthesis in a buffer containing about 0.5-0.75 mM MnSO.
[0012] In some embodiments, the method further comprises quantifying the amplified target nucleic acid.
[0013] In some embodiments, the method further comprises sequencing the exponentially amplified target nucleic acid.
[0014] In some embodiments, the target nucleic acid comprising a DHU residue has a sequencing library adaptor ligated to each end. In some embodiments, the sequencing library adaptor comprises an index sequence. In some embodiments, the sequencing library adaptor comprises a sequence complementary to a sequencing primer. In some embodiments, the sequencing library adaptor comprises a sequence complementary to an index primer.
[0015] In some embodiments, the step of synthesizing a complementary strand of a target nucleic acid comprising a DHU residue further comprises annealing a forward and / or reverse primer(s) to the sequencing library adaptor.
[0016] In some embodiments, the step of exponentially amplifying the complementary strand of the target nucleic acid comprises annealing a library amplification primer to the pre-amplified target nucleic acid.
[0017] In some embodiments, the sequencing is performed by massively parallel sequencing.
[0018] In some preferred embodiments, the step of contacting the oxidized nucleic acid sample containing 5caC and / or 5fC with a borane reducing agent further comprises reacting the oxidized nucleic acid sample with the borane reducing agent in a reaction mixture comprising 45.0% to 52.5% DMSO by volume.
[0019] In some preferred embodiments, the step of contacting the oxidized nucleic acid sample containing 5caC and / or 5fC with a borane reducing agent further comprises reacting the oxidized nucleic acid sample with the borane reducing agent at a temperature between 45.0°C and 52.5°C.
[0020] In some preferred embodiments, the step of contacting the oxidized nucleic acid sample containing 5caC and / or 5fC with a borane reducing agent further comprises reacting the oxidized nucleic acid sample with the borane reducing agent for 45 to 60 minutes.
[0021] In some preferred embodiments, the present invention further provides a method for converting 5-carboxylcytosine (5caC) and / or 5-formylcytosine (5fC) to dihydrouracil (DHU), the method comprising contacting a nucleic acid sample containing 5caC and / or 5fC with a borane reducing agent in a reaction mixture containing 45.0% to 52.5% DMSO by volume. In some preferred embodiments, the method further comprises reacting the oxidized nucleic acid sample with the borane reducing agent at a temperature of 45.0°C to 52.5°C. In some preferred embodiments, the method further comprises reacting the oxidized nucleic acid sample with the borane reducing agent for 45 to 60 minutes.
[0022] In some preferred embodiments, the borane reducing agent comprises an agent selected from the group consisting of 2-picoline borane (pic-BH), borane, sodium borohydride, sodium cyanoborohydride, and sodium triacetoxyborohydride. In some preferred embodiments, the borane reducing agent comprises sodium borohydride. In some preferred embodiments, the borane reducing agent comprises sodium cyanoborohydride. In some preferred embodiments, the borane reducing agent comprises sodium triacetoxyborohydride. In some preferred embodiments, the borane reducing agent comprises 2-picoline borane.
[0023] In some preferred embodiments, the method includes contacting the nucleic acid sample with an oxidizing agent prior to contacting with the borane reducing agent. In some preferred embodiments, the oxidizing agent is a ten-eleven translocation (TET) enzyme. In some preferred embodiments, the TET enzyme includes human TET1, human TET2, human TET3, mouse TET1, mouse TET2, mouse TET3, Naegleria TET (NgTET), Coprinopsis cinerea (CcTET), or a derivative or analog thereof. In some preferred embodiments, the oxidizing agent includes a chemical oxidizing agent. In some preferred embodiments, the chemical oxidizing agent includes manganese oxide (MnO), potassium ruthenate (KRuO), potassium perruthenate (KRuO), or Cu(II) / TEMPO.
[0024] In some preferred embodiments, the method further comprises adding a protecting group to one or more modified cytosines in the nucleic acid sample.
[0025] In some preferred embodiments, the method further comprises sequencing the nucleic acid sample after contacting with the borane reducing agent to identify the converted cytosine bases. [Brief explanation of the drawings]
[0026] [Figure 1]Normalized GC bias of NGS sequencing for fully methylated lambda spike-ins of two TAPS-treated samples compared to two TAPS-untreated samples. [Figure 2] Schematic diagram of the complementary strand synthesis step before amplification. [Figure 3] Normalized GC bias of NGS sequencing for fully methylated lambda spike-in in the Bst3.0 complementary strand synthesis step (with or without denaturation step) prior to amplification. [Figure 4] Normalized GC bias of NGS sequencing for fully methylated lambda spike-in in the Bst3.0 complementary strand synthesis step before amplification using alternative buffer or spike-in options. [Figure 5] Normalized GC bias of NGS sequencing for fully methylated lambda spike-in in the Bst3.0+ / - Sulpholobus pol IV complementary strand synthesis step before amplification. [Figure 6] Normalized GC bias of NGS sequencing for fully methylated lambda spike-in in the Bst3.0+ / -WarmStart RTx complementary strand synthesis step prior to amplification. [Figure 7] Normalized GC bias of NGS sequencing for fully methylated lambda spike-ins in the Bst3.0 or M-MuLVRT complementary strand synthesis step before amplification. [Figure 8] Normalized GC bias of NGS sequencing for fully methylated lambda spike-in in the OneTaq™ and Tth complementary strand synthesis steps prior to amplification. [Figure 9] Normalized GC bias of NGS sequencing for fully methylated lambda spike-in using OneTaq™ and Tth complementary strand synthesis steps prior to amplification, or using OneTaq™ only, or Tth only. [Figure 10]Normalized GC bias of NGS sequencing for fully methylated lambda spike-ins using OneTaq™ and a Tth complementary strand synthesis step prior to amplification, or using Taq alone (both with 0.75 mM or 0 mM MnSO4). [Figure 11] Normalized GC bias of NGS sequencing for fully methylated lambda spike-in in the polymerase κ (kappa) complementary strand synthesis step before amplification. [Figure 12] Normalized GC bias of NGS sequencing of fully methylated lambda spike-ins during the DNA PolI Klenow fragment exo complementary strand synthesis step before amplification. [Figure 13] Normalized GC bias of NGS sequencing for fully methylated lambda spike-in in the SD polymerase complementary strand synthesis step before amplification, with or without denaturation. [Figure 14] Normalized GC bias of NGS sequencing for fully methylated lambda spike-in with pre-amplification 5D4 complementary strand synthesis step alone or as a spike-in option. [Figure 15] Normalized GC bias of NGS sequencing for fully methylated lambda spike-in in library amplification using Kapa Hifi Uracil+ as standard, 5D4 as spike-in option, or replacing Hifi U+ with 10:1 Taq:5D4. [Figure 16A] Average modification rate (16A) and depth (16B) of selected marker regions of high-coverage whole-genome sequencing of NA12878. Conditions shown are without a pre-amplification complementary strand synthesis step (control, red), and with an initial 1-minute 98°C denaturation step (98, green) and without an initial 1-minute 98°C denaturation step (no98, yellow) in Bst3.0 pre-amplification synthesis. [Figure 16B]Average modification rate (16A) and depth (16B) of selected marker regions of high-coverage whole-genome sequencing of NA12878. Conditions shown are without a pre-amplification complementary strand synthesis step (control, red), and with an initial 1-minute 98°C denaturation step (98, green) and without an initial 1-minute 98°C denaturation step (no98, yellow) in Bst3.0 pre-amplification synthesis. [Figure 17A] Average modification rate (17A) and normalized depth (17B, average depth indicated by dashed line) of selected marker regions from high-coverage whole-genome sequencing of NA12878. Conditions shown are no pre-amplification complementary strand synthesis step (KU_std, green), Bst3.0 pre-amplification complementary strand synthesis step (Kapa Hifi Uracil+ spiked with 0.75 mM MnSO(4)) (Bst_spike_Mn, yellow), and OneTaq and Tth pre-amplification complementary strand synthesis step (CS_std, red). [Figure 17B] Average modification rate (17A) and normalized depth (17B, average depth indicated by dashed line) of selected marker regions from high-coverage whole-genome sequencing of NA12878. Conditions shown are no pre-amplification complementary strand synthesis step (KU_std, green), Bst3.0 pre-amplification complementary strand synthesis step (Kapa Hifi Uracil+ spiked with 0.75 mM MnSO(4)) (Bst_spike_Mn, yellow), and OneTaq and Tth pre-amplification complementary strand synthesis step (CS_std, red). [Figure 18A] Average modification rate (18A) and normalized depth (18B) of selected marker regions from whole genome sequencing (WGS) of pooled normal cfDNA. Conditions shown are Kapa Hifi Uracil+ (Bst) spiked with Bst3.0 complementary strand synthesis step before amplification and SD polymerase complementary strand synthesis step before amplification (SD). [Figure 18B]Average modification rate (18A) and normalized depth (18B) of selected marker regions from whole genome sequencing (WGS) of pooled normal cfDNA. Conditions shown are Kapa Hifi Uracil+ (Bst) spiked with Bst3.0 complementary strand synthesis step before amplification and SD polymerase complementary strand synthesis step before amplification (SD). [Figure 19A] Average modification rates (19A) and normalized depth (19B) of selected marker regions from whole genome sequencing (WGS) of pooled normal cfDNA. Conditions shown are no pre-extension (KU), a Bst3.0 complementary strand synthesis step before amplification as a spike into Kapa Hifi Uracil+ (Bst), and a OneTaq™ and Tth complementary strand synthesis step (OTT) before amplification. [Figure 19B] Average modification rates (19A) and normalized depth (19B) of selected marker regions from whole genome sequencing (WGS) of pooled normal cfDNA. Conditions shown are no pre-extension (KU), a Bst3.0 complementary strand synthesis step before amplification as a spike into Kapa Hifi Uracil+ (Bst), and a OneTaq™ and Tth complementary strand synthesis step (OTT) before amplification. [Figure 20A] Average modification rates (left) and normalized depth (right) of select marker regions from hybridization capture target sequencing of pooled normal cfDNA. Conditions shown are no pre-extension (KU), a Bst3.0 complementary strand synthesis step (Bst) before amplification as spiked into Kapa Hifi Uracil+, and a OneTaq and Tth complementary strand synthesis step (OTT) before amplification. [Figure 20B]Average modification rates (left) and normalized depth (right) of select marker regions from hybridization capture target sequencing of pooled normal cfDNA. Conditions shown are no pre-extension (KU), a Bst3.0 complementary strand synthesis step (Bst) before amplification as spiked into Kapa Hifi Uracil+, and a OneTaq and Tth complementary strand synthesis step (OTT) before amplification. [Figure 21] Normalized GC bias of NGS sequencing for fully methylated lambda spike-in in the Bst3.0 complementary strand synthesis step before amplification with alternative buffer or spike-in options using the Swift BioScience Accel Methyl-Seq kit. [Figure 22] Normalized GC bias of NGS sequencing for fully methylated lambda spike-in in the Bst3.0 or DNA PolI Klenow fragment exo complementary strand synthesis step before amplification using the Claret Bioscience SRSLY kit. [Figure 23] Normalized GC bias of NGS sequencing for fully methylated lambda spike-in in the Bst3.0 complementary strand synthesis step before amplification using the Takara Bio EpiXplore kit. [Figure 24] Normalized coverage of selected marker regions with low methylation levels after amplification with TAPS and various polymerases (Kapa Hifi Uracil+, Bst and OTT). [Figure 25] Average conversion of low methylated select marker regions after amplification with TAPS and different polymerases (Kapa Hifi Uracil+, Bst and OTT). [Figure 26] Normalized GC bias of NGS sequencing for fully methylated lambda spike-in in the SeqAmp polymerase complementary strand synthesis step before amplification, in the Bst polymerase complementary strand synthesis step before amplification, or without a separate complementary strand synthesis step before amplification. [Figure 27]Normalized GC bias of NGS sequencing of fully methylated lambda spike-in with the Therminator polymerase complementary strand synthesis step before amplification, with the Bst polymerase complementary strand synthesis step before amplification, or without a separate complementary strand synthesis step before amplification. DETAILED DESCRIPTION OF THE INVENTION
[0027] Recently, we developed a bisulfite-free DNA methylation sequencing method called TET-assisted pyridine borane sequencing (TAPS and its variants, including TAPSβ and CAPS), described in PCT / US2019 / 012627, PCT / IB2020 / 056435, PCT / IB2021 / 000630, PCT / IB2021 / 051091, and PCT / IB2022 / 000420 (each of which is incorporated herein by reference in its entirety). TAPS is based on the use of a mild chemical reaction to directly detect DNA methylation and has demonstrated improved sequence quality, mapping rate, and coverage compared to bisulfite sequencing, while halving sequencing costs. The combination of direct methylation detection and the non-destructive nature of TAPS makes it useful for a variety of nucleic acid samples, including DNA obtained from organisms in the Monera (bacteria), protists, fungi, plant kingdoms, and animal kingdoms. Target nucleic acids can also be obtained from viruses. Nucleic acid samples can be obtained from a patient or subject, an environmental sample, or an organism of interest (e.g., both cells and circulating cell-free DNA (cfDNA obtained from tissues, cells, cell aggregates, blood, plasma, serum, organ secretions, semen (semen), vaginal secretions, cerebrospinal fluid (CSF), saliva, mucus, urine, stool, sweat, pancreatic juice, gastric secretions, gastric juice (gastric lavage), peritoneal fluid, synovial fluid, pleural fluid (pleural lavage), pericardial fluid, ascites, amniotic fluid, nasal secretions, ocular fluid, breast milk, or other bodily fluids containing the desired nucleic acid or cfDNA), DNA obtained from a biopsy, and DNA obtained from cells, secretions, or tissues from lymph glands, breast, liver, bile duct, pancreas, mouth, stomach, colon, rectum, esophagus, small intestine, appendix, duodenum, polyps, gallbladder, anus, prostate, endometrium, vagina, ovaries, cervix, skin, bladder, kidney, lung, and / or peritoneum). In other embodiments, the nucleic acid sample may be obtained from a sample that is cancerous or contains cancerous tissue or cells, or is suspected of being cancerous or containing cancerous tissue or cells. In some embodiments, the nucleic acid sample is obtained from a subject suffering from, suspected of suffering from, or being screened to determine the presence of a disease or disorder (e.g., cancer).In some embodiments, the nucleic acid sample is circulating cell-free DNA (cell-free DNA or cfDNA), for example, DNA found in blood, not present in cells.As those skilled in the art will recognize based on the present disclosure, cfDNA can be separated from body fluids using methods known in the art.Commercially available kits for separating cfDNA are available, including, for example, Circulating Nucleic Acid Kit (Qiagen).Nucleic acid samples can be obtained from enrichment steps, including but not limited to, antibody immunoprecipitation, chromatin immunoprecipitation, restriction enzyme digestion-based enrichment, hybridization-based enrichment, or chemical labeling-based enrichment.
[0028] As further described herein, the methods of the present invention provide improved amplification and sequencing of nucleic acid molecules containing DHU residues (preferably DHU residues introduced by the TAPS protocol), or nucleic acid molecules resulting from the TAPS protocol, or nucleic acid molecules containing by-products of the TAPS protocol. Without being limited to a particular theory, it is contemplated that the presence of DHU residues or other by-products introduced by the TAPS protocol reduces coverage of methylated regions of target nucleic acids during amplification and sequencing. The present invention addresses this problem. In certain embodiments, a first polymerase or polymerase mixture that tolerates DHU residues or other by-products of the TAPS protocol is utilized to generate complementary strands to the target nucleic acid in at least a first round of amplification, followed, optionally, by an exponential amplification step using a second polymerase or polymerase mixture. In some preferred embodiments, the first polymerase or polymerase mixture tolerates the presence of DHU residues in the target nucleic acid. In some preferred embodiments, the first polymerase or polymerase mixture is tolerant to products resulting from the introduction of DHU residues into the nucleic acid. In some preferred embodiments, the first polymerase or polymerase mixture is tolerant to products resulting from the TAPS process, hi some preferred embodiments, the first polymerase or polymerase mixture is tolerant to the presence of DHU residues in the target nucleic acid and / or is tolerant to products resulting from the introduction of DHU residues into the nucleic acid and / or is tolerant to products resulting from the TAPS process.
[0029] In some preferred embodiments, the first polymerase or polymerase mixture is 5.0 x 10 -5 and the second polymerase or polymerase mixture is characterized by having an error rate of greater than 5.0 x 10 -5The TAPS protocol is characterized by having an error rate of less than 1 / 2. In some preferred embodiments, the use of DHU and / or TAPS-resistant polymerases to generate complementary strands results in improved coverage of methylated (and therefore DHU-rich) regions of a biological target nucleic acid sample being processed by the TAPS protocol compared to the same protocol without complementary strand synthesis using DHU and / or TAPS-resistant polymerases. In some embodiments, the improvement in normalized GC bias relative to a fully methylated reference or target sequence with complementary strand synthesis by DHU and / or TAPS-resistant polymerases serves as a proxy for demonstrating improved coverage of methylated regions with DHU and / or TAPS-resistant polymerases compared to a control without complementary strand synthesis by DHU and / or TAPS-resistant polymerases. See Figure 1. Improved GC bias can lead to increased methylation due to reduced competition between DHU-containing and non-DHU-containing strands. In further embodiments, the resulting sequencing library is suitable for use in various sequencing methods, including NGS methods.
[0030] The results provided herein demonstrate that the methods of the invention provide improved sequencing coverage of underrepresented, highly methylated regions when standard library preparation and sequencing protocols are utilized. Because highly methylated regions are of clinical interest, the methods of the invention are useful, for example, for cancer diagnosis and biomarker discovery.
[0031] The section headings used in this section and throughout this disclosure are for organizational purposes only and are not intended to be limiting.
[0032] 1.Definition Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. In case of conflict, the present document, including definitions, will control. Preferred methods and materials are described below; however, methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present disclosure. All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entirety. The materials, methods, and examples disclosed herein are illustrative only and are not intended to be limiting.
[0033] As used herein, the terms "comprise(s)," "include(s)," "having," "has," "can," "contain(s)," and variations thereof are intended to be open-ended transitional phrases, terms, or words that do not exclude the possibility of additional acts or structures. The singular forms "a," "and," and "the" include plural references unless the context clearly dictates otherwise. The present disclosure contemplates other embodiments that "comprise," "consist," and "consist essentially of" the embodiments or elements presented herein, whether explicitly stated or not.
[0034] In describing ranges of values herein, each intervening value is expressly contemplated with the same precision. For example, for the range of 6 to 9, the numbers 7 and 8 are contemplated in addition to 6 and 9, and for the range of 6.0 to 7.0, the numbers 6.0, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, and 7.0 are expressly contemplated.
[0035] In describing ranges of values herein, each intervening value is expressly contemplated with the same precision. For example, for the range of 6 to 9, the numbers 7 and 8 are contemplated in addition to 6 and 9, and for the range of 6.0 to 7.0, the numbers 6.0, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, and 7.0 are expressly contemplated.
[0036] As used herein, "correlated with" refers to compared to.
[0037] As used herein, "methylation" refers to methylation of cytosine at the C5 or N4 position of cytosine, the N6 position of adenine, or other types of nucleic acid methylation. Typical in vitro DNA amplification methods do not preserve the methylation pattern of the amplification template, so in vitro amplified DNA is typically unmethylated. However, "unmethylated DNA" or "methylated DNA" can also refer to amplified DNA in which the original template was unmethylated or methylated, respectively.
[0038] Therefore, as used herein, " methylated nucleotide " or " methylated nucleotide base " refers to the presence of a methyl moiety on a nucleotide base, where the methyl moiety is not present in a recognized typical nucleotide base. For example, cytosine does not contain a methyl moiety on its pyrimidine ring, but 5-methylcytosine contains a methyl moiety at the 5th position of its pyrimidine ring. Therefore, cytosine is not a methylated nucleotide, and 5-methylcytosine is a methylated nucleotide.
[0039] As used herein, a "methylated nucleic acid molecule" refers to a nucleic acid molecule that contains one or more methylated nucleotides.
[0040] As used herein, the "methylation state," "methylation profile," "methylation status," and "methylation signature" of a nucleic acid molecule refer to the presence or absence of one or more methylated nucleotide bases in a nucleic acid molecule. For example, a nucleic acid molecule containing a methylated cytosine is considered to be methylated (e.g., the methylation state of the nucleic acid molecule is methylated). A nucleic acid molecule that does not contain any methylated nucleotides is considered to be unmethylated.
[0041] As used herein, "methylation frequency" or "percent (%) methylation" refers to the number of instances in which a molecule or locus is methylated relative to the number of instances in which the molecule or locus is unmethylated. Methylation state frequencies can be used to represent a population of individuals or a sample from a single individual. For example, a nucleotide locus with a methylation state frequency of 50% has 50% of its instances methylated and 50% of its instances unmethylated. Such frequencies can be used, for example, to represent the degree to which a nucleotide locus or nucleic acid region is methylated in a population of individuals or a collection of nucleic acids. Thus, if the methylation in a first population or pool of nucleic acid molecules differs from the methylation in a second population or pool of nucleic acid molecules, the methylation state frequency of the first population or pool will differ from the methylation state frequency of the second population or pool. Such frequencies can also be used, for example, to represent the degree to which a nucleotide locus or nucleic acid region is methylated in a single individual. For example, such frequencies can be used to represent the degree to which a group of cells from a tissue sample is methylated or unmethylated at a nucleotide locus or nucleic acid region.
[0042] As used herein, the term "error rate" as applied to a polymerase refers to the frequency of errors introduced by the polymerase during replication of a nucleic acid sequence. For example, 5×10 -5 The error rate of 10 replicates 5 This means that an average of five errors are introduced per base.
[0043] As used herein, the term "polymerase or polymerase mixture that is tolerant to DHU residues and / or the introduction of DHU residues into target nucleic acid molecules and / or products resulting from the TAPS process," which may be used interchangeably with the term "DHU- and / or TAPS-resistant polymerase or polymerase mixture," refers to a polymerase or polymerase mixture that provides improved coverage of methylated regions of a methylated target DNA sequence that has been processed by the TAPS, TAPSβ, or CAPS protocol, as compared to Taq polymerase and / or Kapa HiFi Uracil+ polymerase, as assayed by coverage of fully methylated lambda. In some embodiments, an assay for GC bias serves as a surrogate for coverage of methylated regions, where improved GC bias of one enzyme compared to a reference enzyme (e.g., Taq polymerase or KAPA HiFi Uracil+ polymerase) as determined by amplification and sequencing of a reference sequence (e.g., fully methylated lambda) indicates improved coverage of methylated regions in a biological sample.
[0044] As used herein, the term "improved coverage" with respect to methylated regions in a target sequence refers to the ability to maintain a more representative proportion of aligned sequence reads corresponding to highly methylated DNA fragments and a more representative proportion of aligned sequence reads corresponding to highly / unmethylated DNA fragments, such that coverage of highly methylated regions approaches average coverage across the genome and / or methylation signals are improved.
[0045] As used herein, the term "patient" or "subject" refers to an organism that is the subject of various tests provided by the present technology. The term "subject" includes animals, preferably mammals, including humans. In preferred embodiments, the subject is a primate. In even more preferred embodiments, the subject is human. Furthermore, with respect to diagnostic methods, preferred subjects are vertebrate subjects. Preferred vertebrates are warm-blooded, and preferred warm-blooded vertebrates are mammals. Preferred mammals are most preferably humans. As used herein, the term "subject" includes both human and animal subjects. Accordingly, veterinary uses are provided herein. As such, the present technology provides for the diagnosis of mammals, such as humans, as well as mammals of endangered importance, such as the Amur tiger, mammals of economic importance, such as animals raised on farms for human consumption, and / or animals of social importance to humans, such as animals kept as pets or in zoos. Examples of such animals include, but are not limited to, carnivores such as cats and dogs; swine such as pigs, hogs, and wild boars; ruminants and / or ungulates such as cows, oxen, sheep, giraffes, deer, goats, bison, and camels; pinnipeds; and horses.
[0046] 2. TET-assisted pyridine borane sequencing (TAPS) Embodiments of the present disclosure provide bisulfite-free base-resolution methods (e.g., TAPS and related methods TAPSβ and CAPS, collectively referred to as TAPS) for detecting 5-methylcytosine (5mC) and 5-hydroxymethylcytosine (5hmC) in sequences, including for use with blood samples (cellular DNA and cfDNA) and DNA obtained from biopsies. As disclosed in PCT / US2019 / 012627, U.S. Patent Publication No. 20200370114, U.S. Patent Publication No. 20210317519, PCT / IB2020 / 056435, PCT / IB2021 / 000630, PCT / IB2021 / 051091, and PCT / IB2022 / 000420 (each of which is incorporated by reference herein in its entirety), TAPS involves mild enzymatic and chemical reactions that directly and quantitatively detect 5mC and 5hmC at base resolution without affecting unmodified cytosine. The present disclosure also provides methods for detecting 5-formylcytosine (5fC) and 5-carboxylcytosine (5caC) at base resolution without affecting unmodified cytosine. Thus, the methods provided herein provide for mapping of 5mC, 5hmC, 5fC, and 5caC, overcoming the shortcomings of conventional methods such as bisulfite sequencing.
[0047] According to these embodiments, the methods of the present disclosure include converting 5mC and 5hmC (or only 5mC if 5hmC is protected) to 5caC and / or 5fC. In some embodiments, this step includes contacting the DNA or RNA sample with a ten-eleven translocation (TET) enzyme. TET enzymes are a family of enzymes that catalyze the transfer of an oxygen molecule to the C5 methyl group in 5mC, resulting in the formation of 5-hydroxymethylcytosine (5hmC). TET enzymes catalyze the oxidation of 5hmC to 5fC and the further oxidation of 5fC to form 5caC. TET enzymes useful in the methods of the present disclosure include one or more of human TET1, TET2, and TET3; mouse TET1, TET2, and TET3; Naegleria TET (NgTET); Coprinopsis cinerea (CcTET); the catalytic domain of mouse TET1 (mTET1CD), and derivatives or analogs thereof.
[0048] The disclosed methods can also include converting 5caC and / or 5fC in the nucleic acid sample to DHU. In some embodiments, this step includes contacting the DNA or RNA sample with a reducing agent (e.g., a borane reducing agent, such as pyridine borane, 2-picoline borane (pic-BH), borane, sodium borohydride, sodium cyanoborohydride, sodium triacetoxyborohydride, triethylamine borane, and tri(t-butyl)amine borane).
[0049] The present inventors have identified improved reaction conditions for converting 5fC and / or 5caC to DHU, which unexpectedly increase the conversion of 5fC and / or 5caC to DHU while minimizing false positive rates and bias and shortening reaction times. In some preferred embodiments, dimethyl sulfoxide (DMSO) is included in the reaction mixture at a concentration of 40.0% to 60.0% v / v, and ranges and values therein (e.g., 41.0% to 59.0% v / v, 42.0% to 58.0% v / v, 43.0% to 57.0% v / v, 44.0% to 56.0% v / v, 45.0% to 55.0% v / v, 46.0% to 54.0% v / v, 47.0% to 53.0% v / v, 48.0% to 52.0% v / v, 49.0% to 59.0% v / v, 45.0% to 52.0% v / v, 46.0% to 52.0% v / v, or 47.0% to 52.0%). In some particularly preferred embodiments, dimethyl sulfoxide (DMSO) is included in the reaction mixture at a concentration of 45.0% to 52.5% v / v. In some particularly preferred embodiments, DMSO is included in the reaction mixture at a concentration of 48.0% to 52.0% v / v. In some preferred embodiments, the reaction is carried out at a temperature of 40.0°C to 60.0°C, or any range or value therein (e.g., 41.0 to 59.0, 42.0 to 58.0, 43.0 to 57.0, 44.0 to 56.0, 45.0 to 55.0, 46.0 to 54.0, 47.0 to 53.0, 48.0 to 52.0, 49.0 to 59.0, 45.0 to 52.0, 46.0 to 52.0, or 47.0 to 52.0°C). In particularly preferred embodiments, the reaction is carried out at a temperature of 45.0°C to 52.5°C. In particularly preferred embodiments, the reaction is carried out at a temperature of 48.0°C to 52.0°C. In some preferred embodiments, the reaction time for the borane reduction step is 30 minutes to 90 minutes, and ranges and values therein (e.g., 35 to 85, 40 to 80, 45 to 75, 50 to 70, or 50 to 60 minutes). In particularly preferred embodiments, the reaction time for the borane reduction step is 45 minutes to 75 minutes. In particularly preferred embodiments, the reaction time for the borane reduction step is 45 minutes to 60 minutes.
[0050] In some embodiments, the step of converting 5hmC to 5fC comprises oxidizing 5hmC to 5fC by contacting the DNA with, for example, manganese oxide (MnO), potassium ruthenate (KRuO), potassium perruthenate (KRuO), and / or Cu(II) / TEMPO (copper(II) perchlorate and 2,2,6,6-tetramethylpiperidine-1-oxyl (TEMPO)). The 5fC in the DNA sample is then converted to DHU by a method disclosed herein (e.g., by a borane reaction).
[0051] Methods for identifying 5mC. In some embodiments, the methods of the present disclosure include identifying 5mC in a DNA sample (target DNA or whole genome) and providing a quantitative measure of the frequency of 5mC modification at each position in the DNA where the modification is identified. In some embodiments, the proportion of T at each transition position provides a quantitative level of 5mC at each position in the DNA. According to these embodiments, the methods for identifying 5mC may include the use of a protecting group. In other embodiments, the methods for identifying 5mC do not require the use of a protecting group.
[0052] When a protecting group is used to identify 5mC in DNA without including 5hmC, the 5hmC in the sample is blocked from conversion to 5caC and / or 5fC. In some embodiments, the 5hmC in the sample DNA is rendered unreactive in subsequent steps by adding a protecting group to the 5hmC. In one embodiment, the protecting group is a sugar containing a modified sugar, such as glucose or 6-azido-glucose (6-azido-6-deoxy-D-glucose). The sugar protecting group can be added to the hydroxymethyl group of 5hmC by contacting the DNA sample with a uridine diphosphate (UDP) sugar in the presence of one or more glucosyltransferase enzymes. In some embodiments, the glucosyltransferase is T4 bacteriophage β-glucosyltransferase (βGT), T4 bacteriophage α-glucosyltransferase (αGT), and derivatives and analogs thereof. βGT is an enzyme that catalyzes the chemical reaction in which a beta-D-glucosyl (glucose) residue is transferred from UDP-glucose to a 5-hydroxymethylcytosine residue in nucleic acids.
[0053] Methods for identifying 5hmC. In some embodiments, the methods of the present disclosure include identifying 5mC or 5hmC in a DNA sample (target DNA or whole genome). In some embodiments, the methods provide a quantitative measure of the frequency of 5mC or 5hmC modifications at each position in the DNA where the modification is identified. In some embodiments, the proportion of T at each transition position provides a quantitative level of 5mC or 5hmC at each position in the DNA. According to these embodiments, the methods for identifying 5mC or 5hmC provide the locations of 5mC and 5hmC but do not distinguish between the two cytosine modifications. Rather, both 5mC and 5hmC are converted to DHU. The presence of DHU can be detected directly, or the modified DNA (DHU converted to T) can be replicated using the methods of the present disclosure, for example. In some embodiments, the methods for identifying 5hmC include the use of a protecting group. In other embodiments, the methods for identifying 5hmC do not require the use of a protecting group.
[0054] Method for identifying 5mC and / or 5hmC. The present disclosure provides a method for identifying 5mC and 5hmC in DNA by performing a method for identifying 5mC in a first DNA sample and performing a method for identifying 5mC or 5hmC in a second DNA sample. In some embodiments, the first and second DNA samples are derived from the same DNA sample. For example, the first and second samples can be separate aliquots taken from a sample containing DNA (e.g., cellular DNA or cfDNA) to be analyzed.
[0055] Because 5mC and 5hmC (unprotected) are converted to 5fC and 5caC before conversion to DHU, 5fC and 5caC present in a DNA sample are detected as 5mC and / or 5hmC. However, given the extremely low levels of 5fC and 5caC in genomic DNA under normal conditions, this is often acceptable when analyzing methylation and hydroxymethylation in DNA samples. 5fC and 5caC signals can be eliminated by protecting 5fC and 5caC from conversion to DHU, for example, by hydroxylamine conjugation and EDC coupling, respectively. According to these embodiments, the method identifies the location and proportion of 5hmC in DNA by comparing the location and proportion of 5mC with the location and proportion of 5mC or 5hmC (combined). Alternatively, the location and frequency of 5hmC modification in DNA can be measured directly.
[0056] In some embodiments, identifying 5fC and / or 5caC provides the location of 5fC and / or 5caC but does not distinguish between these two cytosine modifications. Rather, both 5fC and 5caC are converted to DHU, which is detected by the methods described herein.
[0057] Methods for identifying 5caC. In some embodiments, the methods include identifying 5caC in a DNA sample (target DNA or whole genome) and providing a quantitative measure of the frequency of 5caC modification at each position in the DNA where the modification is identified. In some embodiments, the proportion of T at each transition position provides a quantitative level of 5caC at each position in the DNA. According to these embodiments, the methods for identifying 5caC may include the use of a protecting group. In other embodiments, the methods for identifying 5caC do not require the use of a protecting group.
[0058] In some embodiments, identification of 5caC in DNA can be achieved when 5fC is protected (and 5mC and 5hmC are not converted to DHU). In some embodiments, adding a protecting group to 5fC in a DNA sample involves contacting the DNA with an aldehyde-reactive compound (e.g., including hydroxylamine derivatives, hydrazine derivatives, and hydrazide derivatives). Hydroxylamine derivatives include ashydroxylamine, hydroxylamine hydrochloride, hydroxylammonium sulfate, hydroxylamine phosphate, O-methylhydroxylamine, O-hexylhydroxylamine, O-pentylhydroxylamine, O-benzylhydroxylamine, and in particular O-ethylhydroxylamine (EtONH), O-alkylated or O-arylated hydroxylamines, acids, or salts thereof. Hydrazine derivatives include N-alkylhydrazine, N-arylhydrazine, N-benzylhydrazine, N,N-dialkylhydrazine, N,N-diarylhydrazine, N,N-dibenzylhydrazine, N,N-alkylbenzylhydrazine, N,N-arylbenzylhydrazine, and N,N-alkylarylhydrazine. Hydrazide derivatives include toluenesulfonylhydrazide, N-acylhydrazide, N,N-alkylacylhydrazide, N,N-benzylacylhydrazide, N,N-arylacylhydrazide, N-sulfonylhydrazide, N,N-alkylsulfonylhydrazide, N,N-benzylsulfonylhydrazide, and N,N-arylsulfonylhydrazide.
[0059] Methods for identifying 5fC. In some embodiments, the methods include identifying 5fC in a DNA sample (target DNA or whole genome), and providing a quantitative measure of the frequency of 5fC modification at each position where the modification is identified in the DNA. In some embodiments, the proportion of T at each transition position provides a quantitative level of 5fC at each position of the DNA. According to these embodiments, the method for identifying 5fC may include the use of a protecting group. In other embodiments, the method for identifying 5fC does not require the use of a protecting group.
[0060] In some embodiments, adding a blocking group to 5caC in a DNA sample can be accomplished by (i) contacting the DNA sample with a coupling agent, e.g., a carboxylic acid derivatizing reagent, e.g., a carbodiimide derivative such as l-ethyl-3-(3-dimethylaminopropyl)carbodiimide (EDC) or N,N'-dicyclohexylcarbodiimide (DCC), and (ii) contacting the DNA sample with an amine, hydrazine, or hydroxylamine compound. Thus, for example, the DNA sample can be treated with EDC and then with benzylamine, ethylamine, or another amine to block 5caC by forming an amide that blocks conversion of 5caC to DHU (e.g., by borane reduction).
[0061] 3. Sequencing Library The present disclosure provides a method for obtaining methylation signature.In some embodiments, the method includes: separating DNA (for example, cellular DNA or cfDNA) from sample; preparing a sequencing library that comprises DNA; and carrying out TET-assisted pyridine borane sequencing (TAPS) on the sequencing library to obtain the methylation signature of DNA.In some embodiments, the methylation signature is a whole genome methylation signature.
[0062] In some embodiments, preparing a sequencing library includes ligating a sequencing adapter to the separated DNA to facilitate sequencing reactions. Sequencing adapters suitable for massively parallel sequencing technology may be used. The present invention is not limited to a specific sequencing technology. In some preferred embodiments, sequencing technologies such as those provided by Illumina or Nanopore may be used. For example, sequencing technologies suitable for use in the present invention include, but are not limited to, those described in U.S. Patent Publication No. 20100120098, U.S. Patent Publication No. 20120208705, U.S. Patent Publication No. 20120208724, WO2012 / 061832, and U.S. Patent Publication No. 2015 / 0368638 (each of which is incorporated herein by reference in its entirety).
[0063] In some embodiments, the adapter comprises one or more sites that can hybridize to a primer. In some embodiments, the adapter comprises at least a first primer site. In some embodiments, the adapter comprises at least a first primer site and a second primer site. The orientation of the primer sites in such embodiments can be such that the primer hybridizing to the first primer site and the primer hybridizing to the second primer site are in the same or different orientations. In one embodiment, the primer sequence in the linker can be complementary to a primer used for amplification. In another embodiment, the primer sequence is complementary to a primer used for sequencing.
[0064] In some embodiments, the linker can include a first primer site and a second primer site with a non-amplifiable site disposed therebetween. The non-amplifiable site is useful for blocking extension of a polynucleotide chain between the first primer site and the second primer site, where the polynucleotide chain hybridizes to one of the primer sites. The non-amplifiable site can also be useful for preventing concatemerization. Examples of non-amplifiable sites include nucleotide analogs, non-nucleotide chemical moieties, amino acids, peptides, and polypeptides. In some embodiments, the non-amplifiable site includes a nucleotide analog that does not significantly base pair with A, C, G, or T.
[0065] Some embodiments include a linker comprising a first primer site and a second primer site with a fragmentation site disposed therebetween. Other embodiments may use forked or Y-shaped adapter designs useful for directional sequencing, as described in U.S. Patent No. 7,741,463, which is incorporated herein by reference.
[0066] In some embodiments, the adapter may include an index or barcode sequence. In further preferred embodiments, the adapter may include a unique molecular identifier (UMI).
[0067] In some embodiments, a carrier nucleic acid or mixture of carrier nucleic acids (e.g., DNA) is added to the sequencing library before performing TAPS. The carrier nucleic acid can be any specific or non-specific DNA molecule (or nucleic acid derivative thereof) that enhances one or more aspects of DNA recovery from the sample.
[0068] As described above, in some preferred embodiments, a nucleic acid containing a DHU residue is subjected to a complementary strand synthesis step using a first polymerase or polymerase mixture that is tolerant of DHU residues and / or products resulting from the introduction of DHU residues and / or the TAPS process, and an exponential amplification step using a second polymerase or polymerase mixture. In some embodiments, the second polymerase or polymerase mixture may be the same as the first polymerase or polymerase mixture or may be a different polymerase or polymerase mixture from the first. In certain embodiments in which the second polymerase mixture is different from the first polymerase mixture, the second polymerase or polymerase mixture may also be tolerant of DHU residues and / or products resulting from the introduction of DHU residues into the target nucleic acid molecule and / or the TAPS process. In embodiments in which the same DHU- and / or TAPS-resistant polymerase is used for both complementary strand synthesis and exponential amplification, it will be understood that the initial complementary strand synthesis step may be part of the exponential amplification process. The complementary strand synthesis and amplification steps can be performed before or after incorporating a sequencing adapter sequence into the target nucleic acid. In some preferred embodiments, the first polymerase or polymerase mixture is 5.0 x 10 -5 In some preferred embodiments, the complementary strand synthesis step with a DHU and / or TAPS-resistant polymerase or polymerase mixture results in improved coverage of highly methylated (and thus DHU-rich) regions compared to the same protocol without the complementary strand synthesis step with a DHU and / or TAPS-resistant polymerase or polymerase mixture.
[0069] Suitable DHU and / or TAPS first polymerases and polymerase mixtures for use in the complementary strand synthesis step include, but are not limited to, Bst3.0 polymerase (New England Biolabs, Beverly, MA), Sulfolobus DNA polymerase IV (New England Biolabs, Beverly, MA), a combination of Bst3.0 polymerase and Sulfolobus DNA polymerase IV, Klenow polymerase (New England Biolabs, Beverly, MA), Klenow exopolymerase (ThermoFisher Scientific, Grand Island, NY), Polκ polymerase, Mu-mLV reverse transcriptase, SD polymerase (Bioron), Tth polymerase (Sigma-Aldrich, St. Louis, MO), OneTaq™ (New England Biolabs, Beverly, MA), OneTaq™ (New England Biolabs, Beverly, MA). a combination of Tth polymerase (Biolabs, Beverly, MA) with 5D4 polymerase, 5D4 polymerase, and a 5D4 polymerase blend with Taq polymerase. In some preferred embodiments, the first polymerase or polymerase blend is thermolabile. In other preferred embodiments, the first polymerase or polymerase blend is thermostable.
[0070] The polymerase utilized in the exponential amplification step may be any polymerase suitable for use in amplification and / or sequencing, and may be the same as or different from the first polymerase. In some preferred embodiments, the polymerase utilized in the exponential amplification step is different from the first polymerase and is referred to as a second polymerase or polymerase mixture. In some particularly preferred embodiments, the polymerase utilized in the exponential amplification step has a lower error rate than the first polymerase or polymerase mixture. In some preferred embodiments, the second polymerase or polymerase mixture is characterized as a high-fidelity polymerase. In some preferred embodiments, the second polymerase or polymerase mixture has an error rate of 5.0 x 10 or less. -5 In some preferred embodiments, the second polymerase or polymerase mixture is selected from Taq polymerase (e.g., GoTaq™ polymerase (Promega, Fitzburgh, Wis.) and engineered B-family polymerases (e.g., KAPA HiFi Uracil+ polymerase (Roche, Indianapolis, Ind.)). In some preferred embodiments, the first polymerase or polymerase mixture is thermostable.
[0071] In some preferred embodiments, the complementary strand synthesis step utilizes a forward primer and / or a reverse primer that anneals to a sequencing adapter. In some preferred embodiments, the exponential amplification step utilizes a sequencing primer that anneals to a region of the sequencing adapter. See Figure 2.
[0072] It is contemplated that DNA methylation signatures are useful for understanding fundamental biological processes and disease pathology, as well as disease detection. For example, methylation signatures / frequencies / markers, etc., may be useful in understanding and studying gene regulation, genomic imprinting, differentiation, development, gene-environment interactions (e.g., smoking, nutrition), aging, various diseases and conditions (e.g., autoimmune diseases, cancer, cardiovascular diseases, CNS diseases, congenital diseases, infectious diseases, metabolic diseases and conditions, NIPT-related tests, etc.), for cancer and other disease detection and diagnosis, and transplant monitoring. In some embodiments, as described herein, the method further includes identifying at least one methylation biomarker from the DNA methylation signature (e.g., whole-genome DNA methylation signature) and determining whether the methylation biomarker differs from a methylation biomarker of a reference sequence or a control sequence. In some embodiments, the methylation biomarker comprises a differentially methylated region (DMR). In some embodiments, the method further includes classifying the sample based on the DMR compared to a reference DMR. In some embodiments, the reference DMR corresponds to a non-disease control or a disease control.
[0073] In some embodiments, the method further comprises identifying at least one methylation biomarker from the DNA methylation signature and determining a tissue of origin corresponding to the methylation biomarker, as described herein. In some embodiments, the method further comprises classifying the sample based on the tissue of origin biomarker.
[0074] In some embodiments, and as described herein, the method further includes identifying a DNA fragmentation profile and determining whether the fragmentation profile is indicative of cancer. According to these embodiments, the DNA fragmentation profile can be determined from TAPS sequencing data (e.g., alignment positions of read pairs).
[0075] In some embodiments, the method further includes identifying at least one sequence variant in the DNA sample and determining whether the sequence variant is indicative of cancer. For example, in some embodiments, TAPS can also distinguish methylation from C-to-T genetic variants or single nucleotide polymorphisms (SNPs) and therefore can be used to detect genetic variants. In some embodiments, methylation and C-to-T SNPs can result in different patterns in TAPS. For example, methylation can result in a T / G read in the original top strand / original bottom strand and an A / C read in the complementary strand. In some embodiments, a C-to-T SNP can result in a T / A read in the original top strand / original bottom strand and the complementary strand. This further enhances the utility of TAPS in obtaining both methylation information and genetic variants, i.e., mutations, in a single experiment and sequencing run. This capability of the TAPS method disclosed herein provides for integration of genomic analysis, including epigenetic analysis, with a substantial reduction in sequencing costs, for example, by eliminating the need to perform standard whole genome sequencing (WGS).
[0076] According to the above embodiments, the methods of the present disclosure include the use of TAPS to generate information regarding methylation signatures, methylation biomarkers, DNA fragment profiles, DNA sequence information (e.g., variants), and tissue-of-origin information in a single experiment for diagnosing / detecting a disease or other condition in a subject (e.g., those provided as examples above). As one of skill in the art will recognize based on this disclosure, TAPS as disclosed herein can be used to generate any combination of methylation signatures, methylation biomarkers, DNA fragment profiles, DNA sequence information (e.g., variants), and tissue-of-origin information for diagnosing / detecting a disease or other condition in a subject (e.g., those provided as examples above). In some embodiments, a methylation signature can be obtained, and one or more of the methylation biomarkers, DNA fragment profiles, DNA sequence information (e.g., variants), and tissue-of-origin information can also be obtained and used to diagnose / detect a disease or other condition in a subject (e.g., those provided as examples above). In some embodiments, the methylation status of biomarkers can be obtained, and one or more of a methylation signature, a DNA fragment profile, DNA sequence information (e.g., variants), and tissue-of-origin information can also be obtained, and can be used to diagnose / detect a disease or other condition in a subject (e.g., those provided as examples above). In some embodiments, a DNA fragmentation profile can be obtained, and one or more of a methylation signature, a methylation biomarker, a DNA sequence information (e.g., variants), and tissue-of-origin information can also be obtained, and can be used to diagnose / detect a disease or other condition in a subject (e.g., those provided as examples above). In some embodiments, DNA sequence variants can be identified, and one or more of a methylation signature, a methylation biomarker, a DNA fragment profile, and tissue-of-origin information can also be obtained, and can be used to diagnose / detect a disease or other condition in a subject (e.g., those provided as examples above).In some embodiments, tissue of origin information can be obtained (e.g., from a whole genome DNA methylation signature), and one or more of a methylation signature, a methylation biomarker, a DNA fragment profile, and DNA sequence information (e.g., variants) can also be obtained and used to diagnose / detect a disease or other condition in a subject (e.g., those provided as examples above).
[0077] In some embodiments, performing TAPS on a sequencing library to obtain a whole genome methylation signature comprises identifying 5mC modifications of DNA and providing a quantitative measure of the frequency of the 5mC modifications. In some embodiments, performing TAPS on a sequencing library to obtain a whole genome methylation signature comprises identifying 5hmC modifications of DNA and providing a quantitative measure of the frequency of the 5hmC modifications. In some embodiments, performing TAPS on a sequencing library to obtain a whole genome methylation signature comprises identifying 5caC modifications of DNA and providing a quantitative measure of the frequency of the 5caC modifications. In some embodiments, performing TAPS on a sequencing library to obtain a whole genome methylation signature comprises identifying 5fC modifications of DNA and providing a quantitative measure of the frequency of the 5fC modifications.
[0078] As one of ordinary skill in the art will recognize based on the present disclosure, the methods described herein (e.g., TAPS) can be used to diagnose / detect any type of cancer. Types of cancer that can be detected / diagnosed using the methods of the present disclosure include, but are not limited to, lung cancer, melanoma, colon cancer, colorectal cancer, neuroblastoma, breast cancer, prostate cancer, renal cell carcinoma, transitional cell carcinoma, cholangiocarcinoma, brain cancer, non-small cell lung cancer, pancreatic cancer, liver cancer, stomach cancer, bladder cancer, esophageal cancer, mesothelioma, thyroid cancer, head and neck cancer, osteosarcoma, hepatocellular carcinoma, carcinoma of unknown primary, ovarian cancer, endometrial cancer, glioblastoma, Hodgkin's lymphoma, and non-Hodgkin's lymphoma. In some embodiments, types of cancer or metastatic forms of cancer that can be detected / diagnosed using the methods of the present disclosure include, but are not limited to, carcinoma, sarcoma, lymphoma, germ cell tumor, and blastoma. In some embodiments, the cancer is an invasive and / or metastatic cancer (e.g., stage II cancer, stage III cancer, or stage IV cancer). In some embodiments, the cancer is an early stage cancer (e.g., stage 0 cancer, stage I cancer) and / or is not an invasive and / or metastatic cancer.
[0079] According to these embodiments, the present disclosure provides methods for quantitatively identifying the location of one or more of 5mC, 5hmC, 5caC, and / or 5fC in a nucleic acid with base resolution, without affecting unmodified cytosines. In some embodiments, the nucleic acid is DNA. In some embodiments, the DNA is cfDNA (e.g., circulating cfDNA). In some embodiments, the nucleic acid is RNA. In some embodiments, the nucleic acid sample includes a target nucleic acid that is DNA or a target nucleic acid that is RNA. In some embodiments, the method is applied to the entire genome and is not limited to a specific target nucleic acid.
[0080] The nucleic acid can be any nucleic acid having a cytosine modification (i.e., 5mC, 5hmC, 5fC, and / or 5caC), but is not limited to a DNA fragment and / or genomic DNA. The nucleic acid can be a single nucleic acid molecule in a sample, or the entire population of nucleic acid molecules in a sample, or any portion thereof (the entire genome or a subset thereof). The nucleic acid can be natural nucleic acid from a source (e.g., a cell, tissue sample, etc.), or can be pre-converted into a form compatible with high-throughput sequencing, for example, by fragmentation, repair, and ligation with adapters for sequencing. Thus, the nucleic acid can include multiple nucleic acid sequences, such that the methods described herein can be used to generate libraries of target nucleic acid sequences that can be analyzed individually (e.g., by sequencing individual targets) or in groups (e.g., by high-throughput or next-generation sequencing methods).
[0081] Because the disclosed methods utilize mild enzymatic and chemical reactions that avoid the substantial degradation of nucleic acids associated with methods such as bisulfite sequencing, the disclosed methods are useful for the analysis of low input samples, such as, for example, circulating cell-free DNA and single-cell analysis.
[0082] In some embodiments, the DNA sample contains picogram amounts of DNA. In some embodiments, the DNA sample contains about 1 pg to about 900 pg of DNA, about 1 pg to about 500 pg of DNA, about 1 pg to about 100 pg of DNA, about 1 pg to about 50 pg of DNA, or about 1 to about 10 pg of DNA. In some embodiments, the DNA sample contains less than about 200 pg, less than about 100 pg of DNA, less than about 50 pg of DNA, less than about 20 pg of DNA, less than about 15 pg of DNA, less than about 10 pg of DNA, or less than about 5 pg of DNA.
[0083] In some embodiments, the DNA sample contains nanogram quantities of DNA. Sample DNA for use in the disclosed methods can be any quantity, including, but not limited to, DNA from a single cell or a bulk DNA sample. In some embodiments, the methods can be performed on DNA samples containing about 1 to about 500 ng of DNA, about 1 to about 200 ng of DNA, about 1 to about 100 ng of DNA, about 1 to about 50 ng of DNA, about 1 to about 10 ng of DNA, or about 2 to about 5 ng of DNA. In some embodiments, the DNA sample contains less than about 100 ng of DNA, less than about 50 ng of DNA, less than 40 ng of DNA, less than 30 ng of DNA, less than 20 ng of DNA, less than 15 ng of DNA, less than 5 ng of DNA, and less than 2 ng of DNA. In some embodiments, the DNA sample contains microgram quantities of DNA.
[0084] The methods of the present disclosure may also include amplifying the copy number of the modified nucleic acid by methods known in the art. When the modified nucleic acid is DNA, the copy number can be increased by, for example, PCR, cloning, and primer extension. The copy number of each target DNA can be amplified by PCR using primers specific to a particular target DNA sequence. Alternatively, multiple different modified target DNA sequences can be amplified by cloning them into a DNA vector using standard techniques. In some embodiments, the copy number of multiple different modified target DNA sequences is increased by PCR, for example, to generate a library for next-generation sequencing, in which double-stranded adapter DNA is pre-ligated to the sample DNA (or modified sample DNA) and PCR is performed using primers complementary to the adapter DNA.
[0085] In some embodiments, the method includes detecting the sequence of the modified nucleic acid. The modified target DNA or RNA contains DHU at the position where one or more of 5mC, 5hmC, 5fC, and 5caC were present in the unmodified target DNA or RNA. DHU acts as T in DNA replication and sequencing methods. Therefore, cytosine modifications can be detected by any direct or indirect method known in the art for identifying C→T changes. Such methods include sequencing methods (e.g., Sanger sequencing, microarrays, and next-generation sequencing). C to T transitions can also be detected by restriction enzyme analysis, where the C to T transition eliminates or introduces a restriction endonuclease recognition sequence.
[0086] Embodiments of the present disclosure also provide kits for identifying 5mC and 5hmC in DNA. Such kits include reagents for identifying 5mC and 5hmC using the methods described herein. The kits may also include reagents for identifying 5caC and 5fC using the methods described herein. In some embodiments, the kits include a TET enzyme, a borane reducing agent, and instructions for carrying out the method. In some embodiments, the borane reducing agent is selected from one or more of the group consisting of pyridine borane, 2-picoline borane (pic-BH), borane, sodium borohydride, sodium cyanoborohydride, and sodium triacetoxyborohydride. In further preferred embodiments, the kits include first and second polymerases or polymerase mixtures as described in detail above.
[0087] In some embodiments, the kit further comprises a 5hmC protecting group and a glycosyltransferase enzyme. In some embodiments, the protecting group added to the 5hmC is a sugar. In some embodiments, the sugar is a naturally occurring sugar or a modified sugar, e.g., glucose or a modified glucose. In some embodiments, the protecting group is added to the 5hmC by contacting the nucleic acid sample with a UDP linked to a sugar, e.g., UDP-glucose, or a UDP linked to a modified glucose, in the presence of a glucosyltransferase enzyme, e.g., T4 bacteriophage β-glucosyltransferase (βGT) and T4 bacteriophage α-glucosyltransferase (αGT), and derivatives and analogs thereof.
[0088] In some embodiments, the kit further comprises an oxidizing agent selected from manganese oxide (MnO), potassium ruthenate (KRuO), potassium perruthenate (KRuO), and / or Cu(II) / TEMPO (copper(II) perchlorate and 2,2,6,6-tetramethylpiperidine-1-oxyl (TEMPO)). In some embodiments, the kit comprises a reagent for protecting 5fC in a nucleic acid sample. In some embodiments, the kit comprises an aldehyde-reactive compound, including, for example, a hydroxylamine derivative, a hydrazine derivative, and a hydrazide derivative described herein. In some embodiments, the kit comprises a reagent for protecting 5caC described herein. In some embodiments, the kit comprises a reagent for isolating DNA or RNA. In some embodiments, the kit comprises a reagent for isolating low-input DNA from a sample, for example, cfDNA from blood, plasma, or serum.
[0089] In some embodiments, the methods of the present disclosure include treating a patient (e.g., a patient with cancer, a patient with early-stage cancer, or a patient suspected of having cancer). In some embodiments, the methods include determining a methylation signature provided herein and administering a treatment to the patient based on the results of determining the methylation signature. The treatment may include administering a pharmaceutical compound, a vaccine, performing surgery, imaging the patient, and / or performing another test. In some embodiments, the methods of the present disclosure can be used as part of clinical screening, methods for prognostic evaluation, methods for monitoring the outcome of therapy, methods for identifying patients most likely to respond to a particular therapeutic treatment, methods for imaging patients or subjects, and methods for drug screening and development.
[0090] Unless otherwise defined herein, scientific and technical terms used in connection with this disclosure shall have the meanings commonly understood by those of ordinary skill in the art. For example, any technical terms used in connection with cell and tissue culture, molecular biology, immunology, microbiology, genetics, and protein and nucleic acid chemistry and hybridization described herein, and these techniques, are well known and commonly used in the art. The meaning and scope of terms should be clear, but in the event of any potential ambiguity, the definitions provided herein shall take precedence over any dictionary or external definitions. Furthermore, unless otherwise required by context, singular terms shall include the plural, and plural terms shall include the singular. [Example]
[0091] 4. Working Example It will be readily apparent to those skilled in the art that other suitable modifications and adaptations of the methods of the present disclosure described herein are readily applicable and discernible, and may be made using suitable equivalents without departing from the scope of the present disclosure or the aspects and embodiments disclosed herein. Having described the present disclosure in detail, the same will be more clearly understood by reference to the following examples, which are intended merely to illustrate certain aspects and embodiments of the disclosure and should not be construed as limiting the scope of the disclosure. The disclosures of all journal references, U.S. patents, and publications mentioned herein are hereby incorporated by reference in their entirety.
[0092] The present disclosure has multiple aspects, illustrated by the following non-limiting examples.
[0093] Example 1 NGS libraries generated from TET-assisted pyridine borane sequencing (TAPS) show reduced coverage of DNA regions with a high density of methylated cytosines compared to average coverage (Figure 1). In Figure 1, the normalized GC bias metric of a reference fully methylated lambda DNA sequence serves as a surrogate for highly methylated DNA, and the difference in the curves for TAPS-treated and non-TAPS-treated fully methylated lambda DNA represents the difference in coverage of methylated regions found in biological samples. If clinically relevant regions of methylated biological samples are underrepresented, sensitivity decreases, necessitating increased sequencing to achieve sufficient coverage. To improve the sequencing efficiency of TAPS libraries, we have found that combining a typical library amplification method with an initial complementary strand synthesis step using DHU or other TAPS product-resistant polymerases improves coverage of highly methylated regions of DNA, as measured by the normalized GC bias.
[0094] Partial NGS sequencing library adapters (i5 and i7 partial Y-shaped adapters, shown in green in Figure 2) are ligated to the fragmented DNA either before or after TAPS treatment. During TAPS treatment, methylated cytosines (meC) are converted to dihydrouracil (DHU) bases by TET oxidation and borane reduction (shown in purple boxes in Figure 2).
[0095] During amplification of these libraries with index primers, polymerases insert an adenine opposite the DHU base, which is then converted to thymidine in the following amplification step. However, many polymerases do not prefer the DHU base or the introduction of DHU residues and / or replication to the products resulting from the TAPS process.
[0096] In the complementary strand synthesis step, this initial duplication beyond the DHU base is performed using an alternative polymerase. For the partial adapter shown in Figure 2, the complementary strand synthesis step utilizes one of a number of polymerases known to have low bias toward DHU (e.g., Bst3.0) and a reverse primer (full i7, shown in Figure 2b), allowing for the copying of library molecules via incubation at a constant temperature (with or without an initial denaturation step). This results in the conversion of the adapter-ligated DNA, now with the DHU base (Figure 2c). This DNA then undergoes a typical PCR amplification using a high-fidelity polymerase with excellent GC bias coverage (e.g., Kapa Hifi Uracil+), and a forward index primer and library amplification primer are added to generate sufficient DNA for sequencing. We have also found that a synthesis step utilizing both index primers (full-length i5 and i7) achieves the same results. Similar conditions are expected to be beneficial when amplifying libraries generated with full-length adapter-ligated DNA.
[0097] Exemplary reagents utilized in the sequencing methods described herein are shown in Tables 1-3. [Table 1] [Table 2] [Table 3]
[0098] Figures 3-14 provide sequencing results for fully methylated lambda spike-ins prepared using Kapa Hyperprep™ and various polymerases in the complementary strand synthesis and amplification steps. Bst3.0 polymerase, combined with several other polymerases and reverse transcriptases using various workflow options, demonstrates improved coverage uniformity. Figure 8 shows the improvement of OneTaq™ and Tth in the presence of OneTaq buffer and MnSO4. We further optimized the method and found similar improvements with OneTaq alone, Tth alone, or Taq polymerase alone and MnSO4, with 0.5-0.75 mM being optimal (Figures 9 and 10). Unless otherwise noted, OneTaq and Tth conditions have 0.75 mM MnSO4. Figures 11, 12, and 13 show the improvement when using polymerase κ, Klenow Exo, and SD polymerase, respectively.
[0099] We also found that the engineered polymerase 5D4 performed well under the complementary strand synthesis step, either alone or as a spike into Kapa Hifi Uracil+ (Figure 14). We emphasize that coverage of high DHU regions was also improved when 5D4 was used for library amplification (without the initial complementary strand synthesis step) in combination with Taq polymerase or as a spike into Kapa Hifi Uracil+ (Figure 15).
[0100] The inventors have further found that selected beneficial complementary strand synthesis conditions result in improved coverage of marker regions in high-coverage whole-genome sequencing, which are typically highly methylated and / or have a high density of CpG sites that result in lower-than-average coverage when sequenced under standard conditions. The Bst3.0 complementary strand synthesis step conditions (Figures 16A-B) and the OneTaq™ and Tth complementary strand synthesis step conditions (Figures 17A-B) are shown herein. The inventors also experience benefits when there is competition between methylated and unmethylated versions of the target sequence. This manifests as improved methylation signals in markers with low levels of methylation. See Figures 24 and 25. Figure 24 shows the normalized coverage of selected marker regions with low levels of methylation after amplification using TAPS and various polymerases (Kapa Hifi Uracil+, Bst, and OTT). FIG. 25 shows the conversion rate of the selectable marker region after amplification with TAPS and various polymerases (Kapa Hifi Uracil+, Bst, and OTT).
[0101] The initial complementary strand synthesis using SD polymerase also shows similar or better normalized coverage compared to the Bst complementary strand synthesis step at highly methylated marker regions in the whole-genome sequencing of NA12878 (Figure 18A-B). The lack of data points for methylation in some regions highlights the low coverage in these highly methylated regions; in this example, coverage is improved compared to without the initial complementary strand synthesis step, but there were not enough reads to reliably determine the average methylation.
[0102] Furthermore, the inventors show that improved coverage by complementary strand synthesis step conditions using Bst or OneTaq and Tth (OTT) is also seen with both whole genome sequencing (WGS) (Figures 19A - B - improved coverage markers b, d, e, h, i) or hybridization capture target sequencing (Figures 20A - B - improved coverage markers A - C, E) at selected highly methylated markers within a pool of normal cfDNA.
[0103] Selected complementary strand synthesis step options also showed improvement with the Accel - NGS Methyl - Seq DNA Library Kit from Swift BioSciences, the SRSLY NGS Library Kit from Claret Bioscience, and the EpiXplore™ Methylated DNA Kit from Takara Bio, as shown in Figures 21 - 23.
[0104] A number of other polymerases were screened that did not show improved normalized GC bias when used in the complementary strand synthesis step under the above test conditions. Examples of these polymerases include KAPA HiFi Uracil +, full - length Bst polymerase, Therminator polymerase, phi29 polymerase, AMV reverse transcriptase, Taq polymerase (NEB), NEB Q5U polymerase, NEB LongAmp Taq, Pyromark, Phusion U<SeqAmp, and ProtoScript™ reverse transcriptase. See, for example, Figures 26 and 27, which provide the results of the complementary strand synthesis step using SeqAmp and Therminator polymerase, respectively, compared to Bst polymerase.
[0105] Example 2 This example provides data related to the optimization of conditions for converting 5-carboxylcytosine (5caC) and / or 5-formylcytosine (5fC) residues in oxidized nucleic acid samples to dihydrouracil (DHU) residues using a borane reducing agent (e.g., pic-borane). These data demonstrate that the conversion of 5caC to dihydrouracil (DHU) using borane reduction chemistry under previously established conditions or those described in Nature Biotechnology (37) 424-429 (2019) is improved by including 45% to 52.5% v / v organic solvent (e.g., dimethyl sulfoxide (DMSO)) in the reaction mixture. Furthermore, the data demonstrate that improved results can be obtained by using increased solvent volumes in the reaction, as well as reaction temperatures of 45 to 52.5 °C and shortened reaction times of 45 to 60 min.
[0106] Compared to alternative borane reduction conditions, the use of a high concentration of solvent (e.g., DMSO) in the reaction offers various advantages. For example, previous conditions using 10% DMSO resulted in significant bias in library amplification after the borane reaction. Using reaction conditions with a high solvent concentration, as described herein, can effectively reduce this bias. Optimization of alternative borane reduction conditions has been attempted, including modifying conditions by increasing temperature, shortening reaction time, varying borane concentration, and using different buffer and pH conditions. However, these previous optimization efforts have not yielded the beneficial optimization of DMSO concentration demonstrated herein. For example, compared to conditions using elevated solvent levels, alternative conditions that only increase the reaction temperature result in an increased false positive rate, while conditions that only shorten the reaction time result in a decreased conversion rate. This maintains a low false positive rate and improved genome coverage. Thus, the data indicate that solvent concentration mediates the effects of shortening the reaction time (which generally decreases conversion rate) and increasing the reaction temperature (which generally increases false positive (FP) rate and bias).
[0107] Specifically, when the reaction time was changed (reduced from 2 hours to 1 hour) without changing other conditions (37°C, 10% DMSO), the conversion rate decreased from 92% to 85% (measured using spiked-in methylated pUC19). When the reaction time was reduced from 2 hours to 1 hour in conjunction with an increase in reaction temperature to 50°C, the conversion rate (measured using spiked-in methylated lambda template) increased to approximately 94%, but this increase was accompanied by an increase in the false positive rate to approximately 2%. This represents an approximately six-fold increase over the approximately 0.35% false positive rate observed under standard conditions (10% DMSO, 37°C, 2 hours). Unexpectedly, when the DMSO concentration was increased to 50% in reactions conducted at 50°C for 1 hour, the high conversion rate (average >94%) was maintained, while the false positive rate decreased to an average of less than approximately 0.3%. Data supporting these observations are provided in Table 4. GC bias plots (not shown) were generated for TAPS reactions performed with various polymerases (Hifi HotStart KAPA Uracil plus, Bst3.0, and OneTaq / Tth) in 10% DMSO at 37°C for 2 hours or 50% DMSO at 50°C for 1 hour. The plots for the reactions using 50% DMSO were flatter than those for the reactions using 10% DMSO. This demonstrates improved coverage and reduced GC bias for reactions containing 50% DMSO. Additional data further clarifying the optimal ranges for time, temperature, and solvent concentration are discussed in Table 4 below. [Table 4]
[0108] Unless otherwise noted, the experiments described herein compared optimized reaction conditions, varying DMSO concentration, reaction temperature, and reaction time, with baseline reaction conditions. The baseline reaction conditions utilized a reaction volume (50 μl) containing 50 ng of oxidized dsDNA, 100 mM buffer (5 μl) at pH 4.0, and 100 mM Pic-borane (5 μl) in DMSO (providing 10% v / v DMSO), with the reaction carried out at 37°C for 2 hours. Various values of solvent concentration, reaction time, and reaction temperature were evaluated to establish optimal ranges, as reported in the following table. Numerous experimental parameters were tested under various conditions, as reported in the following table.
[0109] Briefly, methylation conversion refers to the detection of C→T conversions after TAPS in fully methylated Lambda or partially methylated pUC19 spike-ins. The pUC19 DNA spike-in contains approximately 20% methylation, intended to represent actual conditions where less than 100% of the template is methylated. A false positive is defined as the detection of a C→T conversion in a 2-kb spike-in that is not fully methylated. GC bias describes the dependency between the number of fragments (read coverage) and GC content found in Illumina sequencing data. GC dropout is a metric related to the degree of sequencing bias in a sample; samples with greater GC bias have correspondingly higher GC dropout. Thus, for example, Lambda GC dropout represents the sequencing bias (most Cs are methylated) from a methylated Lambda DNA spike-in to a TAPS reaction.
[0110] The effect of increasing the DMSO concentration in the reaction mixture was evaluated. TAPS reactions were performed at 50°C for 1 hour using various DMSO concentrations (including 0%, 10%, 25%, 50%, 60%, and 75% v / v DMSO). The data are presented in Table 5. As can be seen, TAPS reactions utilizing 50% DMSO have comparable conversion levels as assayed by Lambda methylation to 10% DMSO, improved conversion levels as assayed by pUC19 methylation (more representative of real-world conditions), significantly improved false positive rates, and improved GC dropout metrics. The improvement in GC bias was also confirmed in the GC plots, with the curves for the 50% DMSO reactions being flatter than those for reactions containing lower amounts of DMSO. TAPS reactions containing 25% DMSO had good conversion rates but a relatively high false positive rate and a higher GC bias, as represented by the GC dropout metric. Conversion rates decreased when the DMSO concentration was increased to 60% or 75%. [Table 5]
[0111] The effect of increasing the reaction temperature using 10% or 50% (v / v) DMSO was further evaluated. Conversion rates and false positive rates were measured for various temperatures (e.g., 20°C, 37°C, 45°C, and 50°C) using 10% and 50% (v / v) DMSO and a 1-hour reaction time. The data are presented in Table 6. Conversion, false positive, and GC dropout metrics were also measured for higher temperatures (e.g., 50°C, 75°C, and 100°C) using 10% or 50% (v / v) DMSO, respectively. The data are presented in Table 7. Referring to Table 6, the data demonstrate that the use of increasing DMSO concentrations reduces the false positive rate. In particular, at reaction temperatures above 37°C, using 50% DMSO in the reaction mixture reduces the false positive rate compared to using 10% DMSO. Referring to Table 7, the data show that the false positive rate begins to increase at reaction temperatures above 50°C and that there is also a decrease in recoverable DNA. The use of 50% DMSO results in a decrease in GC bias at all temperatures evaluated, as indicated by the GC dropout metric. The effect on GC bias was confirmed by the GC bias plot (not shown), which showed a flatter curve when 50% DMSO was included in the reaction mixture. [Table 6] [Table 7]
[0112] The effect of reaction time was evaluated using 10% or 50% (v / v) DMSO. Conversion rates and false positive rates were measured for reactions carried out for 15 minutes or 1 hour at 37°C or 50°C using 10% or 50% (v / v) DMSO, respectively. Conversion rates and GC bias (expressed as a GC dropout metric) were also measured for reactions carried out for 1 hour, 2 hours, 6 hours, or 24 hours at a reaction temperature of 50°C using 50% (v / v) DMSO, respectively. The data are presented in Tables 8 and 9. Referring to Table 8, the data show that there is no benefit to increasing DMSO to shorten the reaction time, as evidenced by the low conversion rates for reactions carried out for 15 minutes at either 37°C or 50°C using either 10% or 50% (v / v) DMSO. Referring to Table 9, the data show that as the reaction time increases (2 hours, 6 hours, and 24 hours), the false positive rate increases and yields begin to decrease. The GC dropout metrics also indicate that longer reaction times are generally associated with increased GC bias. The effect of GC bias was confirmed in the GC bias plots (not shown), which showed flatter curves, especially for the 1- and 2-hour reactions, compared to the 6- and 24-hour reactions. [Table 8] [Table 9]
[0113] Additional experiments were conducted to determine the optimal ranges for DMSO concentration (35%, 40%, 45%, 47.5%, 50%, 52.5%, 55%, and 60% (v / v) DMSO), reaction temperature (45°C, 47.5°C, 50°C, 52.5°C, and 55°C), and reaction time (45 min, 50 min, 55 min, 1 h, and 2 h). The collated data are presented in Table 10. From these data, it can be determined that the optimal range of DMSO is 45% to 52.5% v / v. Outside this range, either a decrease in conversion rate or an increase in false positive rate is observed. Furthermore, it can be determined that the optimal reaction time is 45 min to 1 h. Outside this range, an increase in false positive rate is observed, while the conversion rate remains high. Finally, it can be determined that the optimal reaction temperature is 45°C to 52.5°C. Outside this range, an increased false positive rate is observed. [Table 10-1] [Table 10-2] Finally, the optimized conditions of 50% v / v DMSO and 1 hour reaction time at 50° C. (designated ESI-NEW in Table 11) were compared with several sets of conditions. Some of these have been previously utilized (designated in Table 10 as ESI-OLD (10% DMSO, 37°C, 1 hour), CS (8% DMSO, 37°C, 1 or 4 hours), and NB (Nature Biotechnology (37) 424-429 (2019) (pyridine borane (PyB) or Pic-borane (PicB) with 0% or 50% DMSO, 37°C or 50°C, 1, 3, or 16 hours)). The data are presented in Table 11. These data show that the optimized conditions (ESI-NEW) provided the highest conversion levels of any of the conditions tested, particularly for the pUC19 template (i.e., greater than 96% for Lambda methylation and greater than 9.5% for pUC19 methylation). The optimized conditions also provided the highest yields while maintaining low false positive rates and favorable GC dropout metrics.
Table 11
Claims
1. 1. A method for amplifying a target nucleic acid molecule containing dihydrouracil (DHU) residues, comprising: synthesizing one or more complementary strands of a target nucleic acid containing a DHU residue using a first polymerase or polymerase mixture that is tolerant to the DHU residue and / or the products resulting from the introduction of the DHU residue and / or the TAPS process, to obtain a target nucleic acid mixture containing the target nucleic acid containing a DHU residue and one or more complementary strands; and exponentially amplifying the target nucleic acid mixture to obtain amplified target nucleic acids.
2. the first polymerase or polymerase mixture is 5.0 x 10 -5 The method of claim 1 , wherein the method has an error rate of greater than 0.
05.
3. 3. The method of any one of claims 1 to 2, wherein the first polymerase or polymerase mixture is selected from the group consisting of Bst3.0 polymerase, Sulfolobus polymerase IV, a combination of Bst3.0 polymerase and Sulfolobus polymerase IV, Klenow polymerase, Klenow exopolymerase, Polκ polymerase, Mu-mLV reverse transcriptase, SD polymerase, Tth polymerase, OneTaq polymerase, a combination of OneTaq and Tth polymerase, 5D4 polymerase, a 5D4 polymerase blend with Taq polymerase, and SD polymerase.
4. The method of any one of claims 1 to 3, wherein the first polymerase is thermolabile.
5. The method of any one of claims 1 to 3, wherein the first polymerase is thermostable.
6. 4. The method of claim 1, wherein the step of exponentially amplifying the complementary strand of the target nucleic acid utilizes the first polymerase or polymerase mixture that is tolerant to the DHU residue and / or products resulting from the introduction of the DHU residue and / or a TAPS process.
7. 6. The method of any one of claims 1 to 5, wherein the step of exponentially amplifying the pre-amplified target nucleic acid utilizes a second polymerase or polymerase mixture that is different from the first polymerase or polymerase mixture.
8. the second polymerase or polymerase mixture is 5.0 x 10 -5 8. The method of claim 7, having an error rate of less than
9. the second polymerase or polymerase mixture is 1.0 x 10 -6 8. The method of claim 7, having an error rate of less than
10. 10. The method of any one of claims 7 to 9, wherein the second polymerase is selected from the group consisting of GoTaq polymerase and KAPA HiFi Uracil+ polymerase.
11. 5.0 x 10 -5 10. The method of any one of claims 7 to 9, wherein the polymerase having an error rate of less than 100 is thermostable.
12. The method of any one of claims 7 to 10, wherein the first polymerase and the second polymerase are provided in a master mix.
13. Synthesizing a complementary strand of the target nucleic acid containing a DHU residue with a first polymerase or polymerase mixture further comprises using about 0.5 to 0.75 mM MnSO 4 13. The method of any one of claims 1 to 12, comprising synthesis in a buffer comprising:
14. The method according to any one of claims 1 to 13, further comprising quantifying the amplified target nucleic acid.
15. The method of any one of claims 1 to 14, further comprising the step of sequencing the exponentially amplified target nucleic acid.
16. 16. The method of claim 15, wherein the target nucleic acid comprising a DHU residue has sequencing library adaptors ligated to each end.
17. 17. The method of claim 16, wherein the sequencing library adaptors comprise index sequences.
18. The method of any one of claims 16 to 17, wherein the sequencing library adaptor comprises a sequence complementary to a sequencing primer.
19. 19. The method of any one of claims 15 to 18, wherein the sequencing library adapter comprises a sequence complementary to an index primer.
20. 20. The method of any one of claims 15 to 19, wherein synthesizing a complementary strand of the target nucleic acid comprising a DHU residue further comprises annealing a forward primer and / or a reverse primer(s) to the sequencing library adaptor.
21. 21. The method of any one of claims 15 to 20, wherein the step of exponentially amplifying the complementary strand of the target nucleic acid comprises annealing library amplification primers to the pre-amplified target nucleic acid.
22. The method of any one of claims 15 to 21, wherein the sequencing is performed by massively parallel sequencing.
23. 23. The method of any one of claims 1 to 22, wherein the target nucleic acid molecule comprising DHU is produced by a process comprising contacting an oxidized nucleic acid sample comprising 5-carboxylcytosine (5caC) and / or 5-formylcytosine (5fC) with a borane reducing agent.
24. 24. The method of claim 23, wherein the borane reducing agent comprises an agent selected from the group consisting of 2-picoline borane (pic-BH3), borane, sodium borohydride, sodium cyanoborohydride, and sodium triacetoxyborohydride.
25. 25. The method of any one of claims 23-24, wherein the step of contacting the oxidized nucleic acid sample containing 5caC and / or 5fC with a borane reducing agent further comprises reacting the oxidized nucleic acid sample with the borane reducing agent in a reaction mixture comprising 45.0% to 52.5% DMSO by volume.
26. 26. The method of any one of claims 23-25, wherein the step of contacting the oxidized nucleic acid sample containing 5caC and / or 5fC with a borane reducing agent further comprises reacting the oxidized nucleic acid sample with the borane reducing agent at a temperature of from 45.0°C to 52.5°C.
27. 27. The method of any one of claims 23 to 26, wherein the step of contacting the oxidized nucleic acid sample containing 5caC and / or 5fC with a borane reducing agent further comprises reacting the oxidized nucleic acid sample with the borane reducing agent for 45 to 60 minutes.
28. A method for converting 5-carboxylcytosine (5caC) and / or 5-formylcytosine (5fC) to dihydrouracil (DHU), comprising contacting a nucleic acid sample containing 5caC and / or 5fC with a borane reducing agent in a reaction mixture containing 45.0% to 52.5% DMSO by volume.
29. 30. The method of claim 28, further comprising reacting the oxidized nucleic acid sample with the borane reducing agent at a temperature of 45.0°C to 52.5°C.
30. 30. The method of any one of claims 28 to 29, further comprising reacting the oxidized nucleic acid sample with the borane reducing agent for 45 to 60 minutes.
31. 31. The method of any one of claims 28 to 30, wherein the borane reducing agent comprises an agent selected from the group consisting of 2-picoline borane (pic-BH3), borane, sodium borohydride, sodium cyanoborohydride, and sodium triacetoxyborohydride.
32. 32. The method of claim 31 , wherein the borane reducing agent comprises sodium borohydride.
33. 32. The method of claim 31 , wherein the borane reducing agent comprises sodium cyanoborohydride.
34. 32. The method of claim 31 , wherein the borane reducing agent comprises sodium triacetoxyborohydride.
35. 32. The method of claim 31 , wherein the borane reducing agent comprises 2-picoline borane.
36. 36. The method of any one of claims 28 to 35, comprising contacting the nucleic acid sample with an oxidizing agent prior to contacting with the borane reducing agent.
37. 37. The method of claim 36, wherein the oxidizing agent is ten-eleven translocation (TET) enzyme.
38. 38. The method of claim 37, wherein the TET enzyme comprises human TET1, human TET2, human TET3, mouse TET1, mouse TET2, mouse TET3, Naegleria TET (NgTET), Coprinopsis cinerea (CcTET), or a derivative or analog thereof.
39. 37. The method of claim 36, wherein the oxidizing agent comprises a chemical oxidizing agent.
40. The chemical oxidant is manganese oxide (MnO 2 ), potassium ruthenate (K 2 RuO 4 ), potassium perruthenate (KRuO 4 40. The method of claim 39, comprising:
41. 41. The method of any one of claims 36 to 40, further comprising adding a protecting group to one or more modified cytosines in the nucleic acid sample.
42. 42. The method of any one of claims 28 to 41, further comprising sequencing the nucleic acid sample after contacting with the borane reducing agent to identify converted cytosine bases.