TET-assisted pyridine borane sequencing

Through TET-assisted pyridine borane sequencing (TAPS) method, the problems of existing sulfite sequencing methods in sequencing low-input samples and hypermethylated regions are solved, achieving more efficient DNA methylation sequencing.

CN120019160APending Publication Date: 2025-05-16EXACT SCIENCES INNOVATION LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202380070920.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-04-19
Filing Date
2023-10-03
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing sulfite sequencing methods have severe DNA degradation when processing low-input samples and rely on the conversion of unmodified cytosine, resulting in poor sequencing quality, low mapping and increased sequencing costs.

Method used

The library was sequenced by introducing dihydrouracil (DHU) residues using TET-assisted pyridineborane sequencing (TAPS) method. The method includes introducing DHU residues into a nucleic acid sample and synthesizing a sequencing library using a polymerase that is resistant to DHU residues, for exponential amplification.

Benefits of technology

The sequencing coverage of highly methylated regions is improved, and the shortcomings of sulfite sequencing methods in sequencing low-input samples and highly methylated regions are solved, reducing sequencing costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120019160A_ABST
    Figure CN120019160A_ABST
Patent Text Reader

Abstract

Methods for amplifying libraries after introduction of dihydrouracil (DHU) residues by methods such as TET assisted pyridine borane sequencing (TAPS) and variants of TAPS, including TAPS (TAPS beta) blocked by beta-glycosylation and chemically assisted pyridine borane sequencing (CAPS)), are described. The method comprises introducing DHU residues into a nucleic acid sample and preparing a sequencing library by reacting with a complementary strand synthesis step performed with a first polymerase or a mixture of polymerases that is tolerant to DHU residues and / or products resulting from the introduction of DHU residues and / or a TAPS process, followed by exponential amplification. Also described are improved methods for the conversion of oxidized nucleotide residues to DHU.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure provides compositions and methods related to TET-assisted pyridine borane sequencing (TAPS). Specifically, the present disclosure provides optimized methods for generating and sequencing TAPS libraries. Background Art

[0002] 5-Methylcytosine (5mC) and 5-hydroxymethylcytosine (5hmC) are two major epigenetic marks found in mammalian genomes. 5hmC is generated from 5mC by ten-eleven translocation (TET) family dioxygenases. TETs can further oxidize 5hmC to 5-formylcytosine (5fC) and 5-carboxylcytosine (5caC), which are much less abundant in mammalian genomes than 5mC and 5hmC (10- to 100-fold lower than 5hmC). Together, 5mC and 5hmC play crucial roles in a wide range of biological processes, from gene regulation to normal development. Aberrant DNA methylation and hydroxymethylation are associated with various diseases and are recognized hallmarks of cancer. Therefore, the determination of 5mC and 5hmC in DNA sequences is of great value not only for basic research but also for clinical applications, including diagnosis and treatment.

[0003] The gold standard and most widely used method for DNA methylation and hydroxymethylation analysis is currently sulfite sequencing (BS), and its derivative methods such as TET-assisted sulfite sequencing (TAB-Seq) and oxidative sulfite sequencing (oxBS). Similarly, sulfite sequencing is the most complete method for determining whole genome DNA methylation. All of these methods use sulfite treatment to convert unmethylated cytosine to uracil while keeping 5mC and / or 5hmC intact. By PCR amplification of sulfite-treated DNA, uracil is read as thymine, and the modification information of each cytosine can be inferred at a single base resolution (where C is converted to T to provide the position of unmethylated cytosine). However, sulfite sequencing has at least two major disadvantages. First, sulfite treatment is a violent chemical reaction that degrades more than 90% of DNA due to depurination under the required acidic and thermal conditions. This degradation severely limits its application to low-input samples, such as clinical samples and single-cell sequencing including circulating cell-free DNA. Second, bisulfite sequencing relies on the complete conversion of unmodified cytosine to thymine. Unmodified cytosine accounts for approximately 95% of the total cytosine in the human genome. Converting all of these positions to thymine severely reduces sequence complexity, leading to poor sequencing quality, low mapping rates, uneven genome coverage, and increased sequencing costs, as well as reduced ability to call variants. Bisulfite sequencing methods are also prone to false detection of 5mC and 5hmC due to incomplete conversion of unmodified cytosine to thymine.

[0004] Sequencing DNA samples that have been treated to modify naturally occurring bases can be difficult, especially when using massively parallel next generation sequencing (NGS) methods. In particular, underrepresentation of highly methylated regions of interest can cause problems. The present invention provides a solution to this problem. Summary of the invention

[0005] Embodiments of the present disclosure include methods for sequencing a library after the introduction of dihydrouracil (DHU) residues by methods such as TET-assisted pyridine borane sequencing (TAPS) and variants of TAPS, including TAPS blocked by β-glycosylation (TAPSβ) and chemically assisted pyridine borane sequencing (CAPS). According to these embodiments, the method includes introducing the DHU residue into a nucleic acid sample, and preparing a sequencing library by a synthesis step with a first polymerase or polymerase mixture that is tolerant to the DHU residue and / or products generated by the introduction of the DHU residue and / or the TAPS process, followed by exponential amplification.

[0006] Therefore, in some embodiments, the present invention provides a method for amplifying a target nucleic acid molecule containing dihydrouracil (DHU) residues, comprising: synthesizing one or more complementary chains of a target nucleic acid containing DHU residues with a first polymerase or a polymerase mixture that is tolerant to DHU residues and / or products produced by the introduction of DHU residues and / or a TAPS process to provide a target nucleic acid mixture comprising a target nucleic acid containing DHU residues and one or more complementary chains; and exponentially amplifying the target nucleic acid mixture to provide an amplified target nucleic acid.

[0007] In some embodiments, the error rate of the first polymerase or polymerase mixture is greater than 5.0×10 -5 In some embodiments, the first polymerase or polymerase mixture is selected from the group consisting of: Bst3.0 polymerase, Sulpholobus polymerase IV, a combination of Bst3.0 polymerase and Sulpholobus polymerase IV, Klenow polymerase, Klenow exopolymerase, Polκ polymerase, Mu-mLV reverse transcriptase, SD polymerase, Tth polymerase, OneTaq polymerase, a combination of OneTaq and Tth polymerase, 5D4 polymerase, a mixture of 5D4 polymerase and Taq polymerase, and SD polymerase.

[0008] In some embodiments, the first polymerase is thermolabile. In some embodiments, the first polymerase is thermostable. In some embodiments, the step of exponentially amplifying the complementary strand of the target nucleic acid utilizes a first polymerase or polymerase mixture that is tolerant to DHU residues and / or products produced by the introduction of DHU residues and / or the TAPS process.

[0009] In some embodiments, the step of exponentially amplifying the pre-amplified target nucleic acid utilizes a second polymerase or polymerase. In some embodiments, the error rate of the second polymerase or polymerase is less than 5.0×10 5 In some embodiments, the error rate of the second polymerase or polymerase mixture is less than 1.0×10 -6 In some embodiments, the second polymerase is selected from the group consisting of GoTaq polymerase and KAPA HiFi Uracil+ polymerase. In some embodiments, the error rate is less than 5.0X10 -5 The polymerase is thermostable.

[0010] In some embodiments, the first polymerase and the second polymerase are provided in a master mix.

[0011] In some embodiments, synthesizing the complementary strand of the target nucleic acid comprising DHU residues with the first polymerase or polymerase mixture further comprises performing the synthesis in a buffer comprising about 0.5-0.75 mM MnSO4.

[0012] In some embodiments, the method further comprises quantifying the amplified target nucleic acid.

[0013] In some embodiments, the method further comprises the step of sequencing the exponentially amplified target nucleic acid.

[0014] In some embodiments, the target nucleic acid comprising a DHU residue has a sequencing library adapter attached to each end. In some embodiments, the sequencing library adapter comprises an index sequence. In some embodiments, the sequencing library adapter comprises a sequence complementary to a sequencing primer. In some embodiments, the sequencing library adapter comprises a sequence complementary to an index primer.

[0015] In some embodiments, the step of synthesizing a complementary strand of the target nucleic acid comprising a DHU residue further comprises annealing one or more forward and / or reverse primers to a sequencing library adaptor.

[0016] In some embodiments, the step of exponentially amplifying the complementary strand of the target nucleic acid comprises annealing a library amplification primer to the pre-amplified target nucleic acid.

[0017] In some embodiments, sequencing is performed by massively parallel sequencing.

[0018] In some preferred embodiments, the step of contacting the oxidized nucleic acid sample comprising 5caC and / or 5fC with a borane reducing agent further comprises reacting the oxidized nucleic acid sample with the borane reducing agent in a reaction mixture comprising 45.0 vol% to 52.5 vol% DMSO.

[0019] In some preferred embodiments, the step of contacting the oxidized nucleic acid sample comprising 5caC and / or 5fC with a borane reducing agent further comprises reacting the oxidized nucleic acid sample with the borane reducing agent at a temperature of 45.0°C to 52.5°C.

[0020] In some preferred embodiments, the step of contacting the oxidized nucleic acid sample comprising 5caC and / or 5fC with a borane reducing agent further comprises reacting the oxidized nucleic acid sample with the borane reducing agent for a period of 45 to 60 minutes.

[0021] In some preferred embodiments, the present invention also provides a method for converting 5-carboxycytosine (5caC) and / or 5-formylcytosine (5fC) into dihydrouracil (DHU), comprising contacting a nucleic acid sample containing 5caC and / or 5fC with a borane reducing agent in a reaction mixture containing 45.0% to 52.5% by volume of DMSO. In some preferred embodiments, the method further comprises reacting the oxidized nucleic acid sample with the borane reducing agent at a temperature of 45.0° C. to 52.5° C. In some preferred embodiments, the method further comprises reacting the oxidized nucleic acid sample with the borane reducing agent for a period of 45 to 60 minutes.

[0022] In some preferred embodiments, the borane reducing agent comprises an agent selected from the group consisting of 2-picoline borane (pic-BH3), borane, sodium borohydride, sodium cyanoborohydride and sodium triacetoxyborohydride. In some preferred embodiments, the borane reducing agent comprises sodium borohydride. In some preferred embodiments, the borane reducing agent comprises sodium cyanoborohydride. In some preferred embodiments, the borane reducing agent comprises sodium triacetoxyborohydride. In some preferred embodiments, the borane reducing agent comprises 2-picoline borane.

[0023] In some preferred embodiments, the method includes contacting the nucleic acid sample with an oxidant before contacting with a borane reducing agent. In some preferred embodiments, the oxidant is a ten-eleven translocation (TET) enzyme. In some preferred embodiments, the TET enzyme includes human TET1, human TET2, human TET3, mouse TET1, mouse TET2, mouse TET3, Naegleria TET (NgTET), Coprinopsis cinerea (CcTET), or a derivative or analog thereof. In some preferred embodiments, the oxidant includes a chemical oxidant. In some preferred embodiments, the chemical oxidant includes manganese oxide (MnO2), potassium ruthenate (K2RuO4), potassium perruthenate (KRuO4) or Cu(II) / TEMPO.

[0024] In some preferred embodiments, the method further comprises adding a blocking group to one or more modified cytosines in the nucleic acid sample.

[0025] In some preferred embodiments, the method further comprises sequencing the nucleic acid sample after contacting with the borane reducing agent to identify the converted cytosine base. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 Normalized GC bias of fully methylated lambda-spiked NGS sequencing of two TAPS-treated samples compared to two non-TAPS-treated samples.

[0027] Figure 2 . Schematic diagram of the complementary chain synthesis steps before amplification.

[0028] Figure 3 Normalized GC bias for NGS sequencing of fully methylated lambda spike-ins using a Bst 3.0 complementary strand synthesis step with or without a denaturation step prior to amplification.

[0029] Figure 4 Normalized GC bias for NGS sequencing of fully methylated lambda spike-ins using a Bst 3.0 complementary strand synthesis step prior to amplification using either the separate buffer or spike-in option.

[0030] Figure 5 . Normalized GC bias for NGS sequencing of fully methylated lambda spike-ins using a Bst 3.0 + / - Sulpholobus pol IV complementary strand synthesis step prior to amplification.

[0031] Figure 6 . Normalized GC bias for NGS sequencing of fully methylated lambda spike-ins using Bst 3.0 + / - WarmStartRTx complementary strand synthesis step prior to amplification.

[0032] Figure 7 Normalized GC bias for NGS sequencing of fully methylated lambda spike-ins using either Bst 3.0 or M-MuLV RT complementary strand synthesis steps prior to amplification.

[0033] Figure 8 .Use OneTaq before amplification TM Normalized GC bias for NGS sequencing of fully methylated lambda spike-in at the Tth complementary strand synthesis step.

[0034] Fig. 9 .Use OneTaq before amplification TMand Tth complementary strand synthesis step, or OneTaq only TM , or Tth only for normalized GC bias of NGS sequencing of fully methylated lambda spike-ins.

[0035] Fig.10 .Use OneTaq before amplification TM Normalized GC bias for NGS sequencing of fully methylated lambda spikes with either the Tth complementary strand synthesis step or Taq alone (both with 0.75 mM or 0 mM MnSO4).

[0036] Fig.11 Normalized GC bias for NGS sequencing of fully methylated lambda spike-ins using a polymerase K (kappa) complementary strand synthesis step prior to amplification.

[0037] Fig.12 Normalized GC bias for NGS sequencing of fully methylated lambda spike-ins using a DNA Pol I Klenow fragment exo-complementary strand synthesis step prior to amplification.

[0038] Fig.13 Normalized GC bias for NGS sequencing of fully methylated lambda spike-ins using an SD complementary strand synthesis step with or without denaturation prior to amplification.

[0039] Fig.14 . Normalized GC bias for NGS sequencing of fully methylated lambda spike-ins using a 5D4 complementary strand synthesis step either alone or as a spike-in option prior to 1 amplification.

[0040] Fig.15 Normalized GC bias for NGS sequencing of fully methylated lambda spike-ins by library amplification using Kapa HiFi Uracil+ as standard, or using 5D4 as the spike-in option, or replacing HiFi U+ with 10:1 Taq:5D4.

[0041] Figures 16A-16B Average modification rate (16A) and depth (16B) of selected marker regions in high-coverage whole-genome sequencing of NA12878. The conditions shown are without a complementary strand synthesis step prior to amplification (control, red), and with (98, green) and without (no98, yellow) an initial 1-minute 98°C denaturation step prior to amplification with Bst3.0 synthesis.

[0042] Figures 17A-17BAverage modification rate (17A) and normalized depth (17B, average depth shown as dashed line) of selected marker regions from high-coverage whole-genome sequencing of NA12878. The conditions shown are no complementary strand synthesis step before amplification (KU_std, green), Bst3.0 complementary strand synthesis step before amplification as a spike into Kapa HiFi Uracil+ supplemented with 0.75mM MnS0(4) (Bst_spike_Mn, yellow), and OneTaq and Tth complementary strand synthesis step before amplification (CS_std, red).

[0043] Figures 18A-18B Average modification rates (18A) and normalized depths (18B) of selected marker regions from whole genome sequencing (WGS) sequencing of pooled normal cfDNA. The conditions shown are Bst3.0 complementary strand synthesis step (Bst) as a spike-in to Kapa HiFi Uracil+ prior to amplification, and SD polymerase complementary strand synthesis step (SD) prior to amplification.

[0044] Figures 19A-19B Average modification rates (19A) and normalized depths (19B) of selected marker regions from whole genome sequencing (WGS) of pooled normal cfDNA. Conditions shown are no pre-extension (KU), Bst3.0 complementary strand synthesis step as a spike-in to KapaHiFiUracil+ prior to amplification (Bst), and OneTaq TM and Tth complementary strand synthesis step (OTT).

[0045] Figures 20A-20B Average modification rate (left) and normalized depth (right) of selected tagged regions from hybridization capture targeted sequencing of pooled normal cfDNA. Conditions shown are no pre-extension (KU), Bst3.0 complementary strand synthesis step as a spike-in to Kapa HiFiUracil+ prior to amplification (Bst), and OneTaq and Tth complementary strand synthesis step prior to amplification (OTT).

[0046] Fig.21 Normalized GC bias for NGS sequencing of fully methylated lambda spike-ins using the Swift BioScience Accel Methyl-Seq Kit with a Bst 3.0 complementary strand synthesis step prior to amplification using either the buffer alone or spike-in options.

[0047] Fig. 22Normalized GC bias for NGS sequencing of fully methylated lambda spike-ins using the Claret Bioscience SRSLY kit with either Bst 3.0 or DNA polI Klenow fragment exo-complementary strand synthesis steps prior to amplification.

[0048] Fig.23 GC bias normalized for NGS sequencing of fully methylated lambda spike-ins using the Takara Bio EpiXplore kit using a BST 3.0 complementary strand synthesis step prior to amplification.

[0049] Fig.24 Normalized coverage of selected marker regions with low levels of methylation after TAPS and amplification with different polymerases (Kapa HiFi Uracil+, Bst, and OTT).

[0050] Fig.25 Average conversion rates of selected tagged regions with low levels of methylation after TAPS and amplification with different polymerases (Kapa HiFi Uracil+, Bst, and OTT).

[0051] Fig.26 Normalized GC bias for NGS sequencing of fully methylated lambda spikes using a SeqAmp polymerase complementary strand synthesis step prior to amplification, using a Bst polymerase complementary strand synthesis step prior to amplification, or no separate complementary strand synthesis step prior to amplification.

[0052] Fig. 27 Normalized GC bias for NGS sequencing of fully methylated lambda spikes using a Therminator polymerase complementary strand synthesis step prior to amplification, using a Bst polymerase complementary strand synthesis step prior to amplification, or no complementary strand synthesis step prior to amplification. DETAILED DESCRIPTION

[0053] Recently, TET-assisted pyridine borane sequencing (TAPS and its variants, including TAPSβ and CAPS), a sulfite-free DNA methylation sequencing method, has been developed, as described in PCT / US2019 / 012627, PCT / IB2020 / 056435, PCT / IB2021 / 000630, PCT / IB2021 / 051091, and PCT / IB2022 / 000420, each of which is incorporated herein by reference in its entirety. TAPS is based on the use of mild chemical reactions to directly detect DNA methylation, and improves sequence quality, mapping rate, and coverage compared to bisulfite sequencing, while reducing sequencing costs by half. Direct methylation detection combined with the non-destructive nature of TAPS allows it to be used for a variety of nucleic acid samples, including DNA obtained from organisms in the prokaryotes (bacteria), protists, fungi, plants, and animals. Target nucleic acids can also be obtained from viruses. Nucleic acid samples can be obtained from a patient or subject, from an environmental sample, or from a target organism (e.g., cells and circulating cell-free DNA (cfDNA obtained from tissues, cells, cell collections, blood, plasma, serum, organ secretions, semen (seminal fluid), vaginal secretions, cerebrospinal fluid (CSF), saliva, mucus, urine, feces, sweat, pancreatic juice, gastric secretions, gastric juice (gastric lavage fluid), ascites, synovial fluid, pleural fluid (pleural lavage fluid), pericardial fluid, peritoneal fluid, amniotic fluid, nasal fluid, optic nerve fluid, breast milk, or any other body fluid containing the desired nucleic acid or cfDNA), DNA obtained from a biopsy, and DNA obtained from cells, secretions, or cells, secretions, or tissues from the lymph glands, breast, liver, bile duct, pancreas, oral cavity, stomach, colon, rectum, esophagus, small intestine, appendix, duodenum, polyps, gall bladder, anus, prostate, endometrium, vagina, ovary, cervix, skin, bladder, kidney, lung, and / or peritoneum). In other embodiments, nucleic acid samples can be obtained from cancerous, containing cancerous tissue or cells, or suspected cancerous or suspected containing cancerous tissue or cells. In some embodiments, nucleic acid samples are obtained from subjects suffering from diseases or conditions (such as cancer), suspected of having diseases or conditions, or being screened to determine the presence of diseases or conditions. In some embodiments, nucleic acid samples are circulating cell free DNA (cell free DNA or cfDNA), such as DNA found in blood and not present in cells. As will be recognized by those of ordinary skill in the art based on the present disclosure, cfDNA can be separated from body fluids using methods known in the art. Commercial kits that can be used to separate cfDNA include, for example, circulating nucleic acid kits (Qiagen). DNA samples can be produced by enrichment steps, including but not limited to antibody immunoprecipitation, chromatin immunoprecipitation, enrichment based on restriction enzyme digestion, enrichment based on hybridization, or enrichment based on chemical markers.

[0054] As further described herein, the methods of the present invention provide improved amplification and sequencing of nucleic acid molecules containing DHU residues, preferably DHU residues introduced by the TAPS protocol, or nucleic acid molecules produced by the TAPS protocol, or nucleic acid molecules containing byproducts of the TAPS protocol. Without being bound by any particular theory, it is contemplated that the presence of DHU residues or other byproducts introduced by the TAPS protocol can result in reduced coverage of methylated regions in the target nucleic acid during amplification and sequencing. The present invention addresses this problem. In certain embodiments, a first polymerase or polymerase mixture that is tolerant to DHU residues or other byproducts of the TAPS protocol is used to generate a strand complementary to the target nucleic acid in at least a first round of amplification, followed by an exponential amplification step, optionally using a second polymerase or polymerase mixture. In some preferred embodiments, the first polymerase or polymerase mixture is tolerant to the presence of DHU residues in the target nucleic acid. In some preferred embodiments, the first polymerase or polymerase mixture is tolerant to products produced by introducing DHU residues into nucleic acids. In some preferred embodiments, the first polymerase or polymerase mixture is tolerant to products produced by the TAPS process. In some preferred embodiments, the first polymerase or polymerase mixture is tolerant to the presence of DHU residues in the target nucleic acid and / or to products produced by the introduction of DHU residues into the nucleic acid and / or to products produced by the TAPS process.

[0055] In some preferred embodiments, the first polymerase or polymerase mixture is characterized by having a molecular weight greater than 5.0×10 -5 and the second polymerase or polymerase mixture is characterized by having an error rate of less than 5.0X10 -5 In some preferred embodiments, the use of DHU and / or TAPS-resistant polymerases to produce complementary chains results in increased coverage of methylated (and therefore DHU-rich) regions of biological target nucleic acid samples that have been treated by the TAPS protocol, compared to the same protocol without using DHU and / or TAPS-resistant polymerases for complementary chain synthesis. In some embodiments, the improved normalized GC deviation of a fully methylated reference or target sequence that uses DHU and / or TAPS-resistant polymerases for complementary chain synthesis relative to a control that does not use DHU and / or TAPS-resistant polymerases for complementary chain synthesis is used as an indicator to demonstrate improved coverage of methylated regions by DHU and / or TAPS-resistant polymerases. See Figure 1 Considering that the improved GC bias may lead to higher methylation due to less competition between DHU-containing strands and non-DHU-containing strands. In further embodiments, the resulting sequencing library is suitable for a variety of sequencing methods, including NGS methods.

[0056] The results presented herein show that the methods of the invention provide improved sequencing coverage of highly methylated regions that are underrepresented when using standard library preparation and sequencing protocols. Because highly methylated regions have clinical significance, the methods of the invention can be used, for example, for cancer diagnosis and biomarker discovery.

[0057] The section headings as used in this section and the entire disclosure herein are for organizational purposes only and are not intended to be limiting.

[0058] 1. Definition

[0059] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those generally understood by those of ordinary skill in the art. In the event of a conflict, this document (including definitions) shall prevail. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure, preferred methods and materials are described below. All publications, patent applications, patents and other references mentioned herein are incorporated by reference in their entirety. The materials, methods and embodiments disclosed herein are illustrative only and are not intended to be restrictive.

[0060] As used herein, the terms "comprise," "include," "having," "has," "can," "contain," and variations thereof are intended to be open-ended conjunctions, terms, or words that do not exclude the possibility of additional actions or structures. Unless the context clearly dictates otherwise, the singular forms "a," "an," and "the" include plural referents. The present disclosure also contemplates other embodiments "comprising," "consisting of," and "consisting essentially of," whether or not explicitly stated.

[0061] For the recitation of numerical ranges herein, each intermediate number with the same degree of precision is expressly contemplated. For example, for a range of 6 to 9, the numbers 7 and 8 are included in addition to 6 and 9, and for a range of 6.0 to 7.0, the numbers 6.0, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9 and 7.0 are expressly included.

[0062] For the recitation of numerical ranges herein, each intermediate number with the same degree of precision is expressly contemplated. For example, for a range of 6 to 9, the numbers 7 and 8 are included in addition to 6 and 9, and for a range of 6.0 to 7.0, the numbers 6.0, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9 and 7.0 are expressly included.

[0063] "Related to," as used herein, means compared to.

[0064] As used herein, "methylation" refers to methylation of cytosine at position C5 or N4 of cytosine, methylation of cytosine at position N6 of adenine, or other types of nucleic acid methylation. In vitro amplified DNA is usually non-methylated because typical in vitro DNA amplification methods do not retain the methylation pattern of the amplified template. However, "unmethylated DNA" or "methylated DNA" may also refer to amplified DNA that is unmethylated or methylated in the original template, respectively.

[0065] Thus, as used herein, "methylated nucleotide" or "methylated nucleotide base" refers to the presence of a methyl moiety on a nucleotide base, wherein the methyl moiety is not present in recognized typical nucleotide bases. For example, cytosine does not contain a methyl moiety on its pyrimidine ring, but 5-methylcytosine contains a methyl moiety at position 5 of its pyrimidine ring. Thus, cytosine is not a methylated nucleotide, and 5-methylcytosine is a methylated nucleotide.

[0066] As used herein, a "methylated nucleic acid molecule" refers to a nucleic acid molecule containing one or more methylated nucleotides.

[0067] As used herein, the "methylation state", "methylation profile", "methylation status" and "methylation signature" of a nucleic acid molecule refers to the presence or absence of one or more methylated nucleotide bases in a nucleic acid molecule. For example, a nucleic acid molecule containing methylated cytosine is considered to be methylated (e.g., the methylation state of the nucleic acid molecule is methylated). A nucleic acid molecule that does not contain any methylated nucleotides is considered to be unmethylated.

[0068] As used herein, "methylation frequency" or "methylation percentage (%)" refers to the number of instances in which a molecule or locus is methylated relative to the number of instances in which the molecule or locus is unmethylated. The methylation state frequency can be used to describe a population of individuals or a sample from a single individual. For example, a nucleotide locus with a methylation state frequency of 50% is methylated in 50% of cases and unmethylated in 50% of cases. Such a frequency can be used, for example, to describe the degree of methylation of a nucleotide locus or a nucleic acid region in a population of individuals or a nucleic acid collection. Therefore, when the methylation in a first population or pool of nucleic acid molecules is different from the methylation in a second population or pool of nucleic acid molecules, the methylation state frequency of the first population or pool will be different from the methylation state frequency of the second population or pool. Such a frequency can also be used, for example, to describe the degree of methylation of a nucleotide locus or a nucleic acid region in a single individual. For example, such a frequency can be used to describe the degree of methylation or unmethylation at a nucleotide locus or a nucleic acid region in a group of cells from a tissue sample.

[0069] As used herein, the term "error rate" when applied to a polymerase refers to the frequency with which the polymerase introduces errors during the replication of a nucleic acid sequence. For example, an error rate of 5 x 10 -5 This means that every 10 copies 5 Each base will introduce an average of 5 errors.

[0070] As used herein, the term "polymerase or polymerase mixture that tolerates DHU residues and / or products resulting from the introduction of DHU residues into a target nucleic acid molecule and / or a TAPS process" is used interchangeably with the term "DHU and / or TAPS tolerant polymerase or polymerase mixture" to refer to a polymerase or polymerase mixture that provides better coverage of methylated regions of a methylated target DNA sequence treated with a TAPS, TAPSβ or CAPS protocol, as measured by coverage of fully methylated lambda, compared to Taq polymerase and / or KAPA HiFi Uracil+ polymerase. In some embodiments, a GC bias assay is used as a surrogate for methylated region coverage, wherein an enzyme, as determined by amplification and sequencing of a reference sequence (e.g., fully methylated lambda), has an increased GC bias compared to a reference enzyme (e.g., Taq polymerase or KAPA HiFi Uracil+ polymerase), indicating improved coverage of methylated regions in a biological sample.

[0071] As used herein, the term "improved coverage" when referring to methylated regions in a target sequence refers to [the ability to maintain the proportion of aligned sequence reads corresponding to highly methylated DNA fragments and the proportion of aligned sequence reads corresponding to less highly / non-methylated DNA fragments in a more representative manner, such that the coverage of highly methylated regions is closer to the average coverage of the entire genome, and / or the methylation signal is improved.]

[0072] As used herein, the term "patient" or "subject" refers to an organism to be subjected to the various tests provided by the technology. The term "subject" includes animals, preferably mammals, including humans. In a preferred embodiment, the subject is a primate. In an even more preferred embodiment, the subject is a human. Further with respect to the diagnostic method, the preferred subject is a vertebrate subject. Preferred vertebrates are warm-blooded; preferred warm-blooded vertebrates are mammals. Preferred mammals are most preferably humans. As used herein, the term "subject" includes both human and animal subjects. Therefore, veterinary therapeutic uses are provided herein. Therefore, the present technology provides diagnosis of mammals, such as humans, and those mammals that are important due to endangerment, such as Siberian tigers; mammals of economic importance, such as animals raised on farms for human consumption; and / or animals of social importance to humans, such as animals raised as pets or in zoos. Examples of such animals include, but are not limited to: carnivores such as cats and dogs; suids, including pigs, hogs, and wild boars; ruminants and / or ungulates such as cattle, bulls, sheep, giraffes, deer, goats, bison, and camels; pinnipeds; and horses.

[0073] 2.TET-assisted pyridine borane sequencing (TAPS)

[0074] Embodiments of the present disclosure provide a sulfite-free, base-resolution method (e.g., TAPS and related methods TAPSβ and CAPS, collectively referred to as TAPS) for detecting 5-methylcytosine (5mC) and 5-hydroxymethylcytosine (5hmC) in a sequence, including for use with DNA obtained from blood samples (cellular DNA as well as cfDNA) and biopsies. As disclosed in PCT / US2019 / 012627, U.S. Patent Publication 20200370114, U.S. Patent Publication 20210317519, PCT / IB2020 / 056435, PCT / IB2021 / 000630, PCT / IB2021 / 051091, and PCT / IB2022 / 000420 (each of which is incorporated herein by reference in its entirety), TAPS includes the use of mild enzymatic and chemical reactions to directly and quantitatively detect 5mC and 5hmC with base resolution without affecting unmodified cytosine. The present disclosure also provides methods for detecting 5-formylcytosine (5fC) and 5-carboxylcytosine (5caC) at base resolution without affecting unmodified cytosine. Therefore, the methods provided herein provide mapping of 5mC, 5hmC, 5fC and 5caC, and overcome the shortcomings of previous methods (such as bisulfite sequencing).

[0075] According to these embodiments, the methods of the present disclosure include a step of converting 5mC and 5hmC (or only 5mC if 5hmC is blocked) to 5caC and / or 5fC. In some embodiments, the step includes contacting a DNA or RNA sample with a ten-eleven translocation (TET) enzyme. TET enzymes are a family of enzymes that catalyze the transfer of oxygen molecules to the C5 methyl group on 5mC, thereby forming 5-hydroxymethylcytosine (5hmC). TET further catalyzes the oxidation of 5hmC to 5fC and the oxidation of 5fC to form 5caC. TET enzymes that can be used in the methods of the present disclosure include one or more of the following: human TET1, TET2, and TET3; mouse TET1, TET2, and TET3; Naegleria TET (NgTET); Coprinus cinereus (CcTET); the catalytic domain of mouse TET1 (mTET1CD); and derivatives or analogs thereof.

[0076] The method of the present disclosure may also include a step of converting 5caC and / or 5fC in the nucleic acid sample into DHU. In some embodiments, this step includes contacting the DNA or RNA sample with a reducing agent, including, for example, a borane reducing agent, such as pyridine borane, 2-methylpyridine borane (pic-BH3), borane, sodium borohydride, sodium cyanoborohydride, sodium triacetoxyborohydride, triethylamine borane, and tri(tert-butyl)amine borane.

[0077] The present inventors have identified improved reaction conditions for converting 5fC and / or 5caC to DHU. The improved reaction conditions unexpectedly increase the conversion of 5fC and / or 5caC to DHU while minimizing false positive rates and biases, while also providing shortened reaction times. In some preferred embodiments, the concentration of dimethyl sulfoxide (DMSO) included in the reaction mixture is 40.0% to 60.0% v / v and ranges and values ​​therein (e.g., 41.0% to 59.0% v / v, 42.0% to 58.0% v / v, 43.0% to 57.0% v / v, 44.0% to 56.0% v / v, 45.0% to 55.0% v / v, 46.0% to 54. 0% v / v, 47.0% to 53.0% v / v, 48.0% to 52.0% v / v, 49.0% to 59.0% v / v, 45.0% to 52.0% v / v, 46.0% to 52.0% v / v or 47.0% to 52.0%) In some particularly preferred embodiments, the concentration of dimethyl sulfoxide (DMSO) included in the reaction mixture is 45.0% to 52.5% v / v. In some particularly preferred embodiments, the concentration of dimethyl sulfoxide (DMSO) included in the reaction mixture is 48.0% to 52.0% v / v. In some preferred embodiments, the reaction is carried out at a temperature of 40.0°C to 60.0°C and ranges and values ​​therein (e.g., 41.0 to 59.0, 42.0 to 58.0, 43.0 to 57.0, 44.0 to 56.0, 45.0 to 55.0, 46.0 to 54.0%, 47.0 to 53.0, 48.0 to 52.0, 49.0 to 59.0, 45.0 to 52.0, 46.0 to 52.0, or 47.0 to 52.0°C). In some particularly preferred embodiments, the reaction is carried out at a temperature of 45.0°C to 52.5°C. In some particularly preferred embodiments, the reaction is carried out at a temperature of 48.0°C to 52.0°C. In some preferred embodiments, the reaction time of the borane reduction step is 30 minutes to 90 minutes and ranges and values ​​therein (e.g., 35 to 85, 40 to 80, 45 to 75, 50 to 70, or 50 to 60 minutes). In some particularly preferred embodiments, the reaction time of the borane reduction step is 45 minutes to 75 minutes. In some particularly preferred embodiments, the reaction time of the borane reduction step is 45 minutes to 60 minutes.

[0078] In some embodiments, the step of converting 5hmC to 5fC comprises oxidizing 5hmC to 5fC by contacting the DNA with, for example, manganese oxide (MnO2), potassium ruthenate (K2RuO4), potassium perruthenate (KRuO4), and / or Cu(II) / TEMPO (copper (II) perchlorate and 2,2,6,6-tetramethylpiperidin-1-oxyl (TEMPO)). 5fC in the DNA sample is then converted to DHU by a method disclosed herein (e.g., by a borane reaction).

[0079] Methods for identifying 5mC. In some embodiments, the methods of the present disclosure include identifying 5mC in a DNA sample (targeted DNA or whole genome) and providing a quantitative measurement of the frequency of 5mC modifications at each position where the modification is identified in the DNA. In some embodiments, the percentage of T at each transition position provides a quantitative level of 5mC at each position in the DNA. According to these embodiments, the method for identifying 5mC may include the use of a blocking group. In other embodiments, the method for identifying 5mC does not require the use of a blocking group.

[0080] When a blocking group is used to identify 5mC in DNA that does not contain 5hmC, 5hmC in the sample is blocked so that it will not be converted into 5caC and / or 5fC. In some embodiments, by adding a blocking group to 5hmC, 5hmC in the sample DNA is made unreactive to subsequent steps. In one embodiment, the blocking group is a sugar, including a modified sugar, such as glucose or 6-azido-glucose (6-azido-6-deoxy-D-glucose). By contacting the DNA sample with uridine diphosphate (UDP)-sugar in the presence of one or more glucosyltransferases, a sugar blocking group can be added to the hydroxymethyl group of 5hmC. In some embodiments, the glucosyltransferase is T4 bacteriophage β-glucosyltransferase (βGT), T4 bacteriophage α-glucosyltransferase (αGT) and its derivatives and analogs. βGT is an enzyme that catalyzes a chemical reaction in which β-D-glucosyl (glucose) residues are transferred from UDP-glucose to 5-hydroxymethylcytosine residues in nucleic acids.

[0081] Methods for identifying 5hmC. In some embodiments, the methods of the present disclosure include identifying 5mC or 5hmC in a DNA sample (targeted DNA or whole genome). In some embodiments, the method provides a quantitative measurement of the frequency of 5mC or 5hmC modifications at each position where the modification is identified in the DNA. In some embodiments, the percentage of T at each transition position provides a quantitative level of 5mC or 5hmC at each position in the DNA. According to these embodiments, the method for identifying 5mC or 5hmC provides the positions of 5mC and 5hmC, but does not distinguish between the two cytosine modifications. Instead, 5mC and 5hmC are both converted to DHU. The presence of DHU can be detected directly, or the modified DNA can be replicated, for example, by the methods of the present disclosure, wherein DHU is converted to T. In some embodiments, the method for identifying 5hmC includes the use of a blocking group. In other embodiments, the method for identifying 5hmC does not require the use of a blocking group.

[0082] Methods for identifying 5mC and / or identifying 5hmC. The present disclosure provides a method for identifying 5mC in DNA and identifying 5hmC, the method being performed by performing a method for identifying 5mC on a first DNA sample and performing a method for identifying 5mC or 5hmC on a second DNA sample. In some embodiments, the first and second DNA samples are derived from the same DNA sample. For example, the first and second samples can be separate aliquots taken from a sample containing DNA to be analyzed (e.g., cellular DNA or cfDNA).

[0083] Since 5mC and 5hmC (which are not blocked) are converted to 5fC and 5caC before conversion to DHU, any 5fC and 5caC present in the DNA sample will be detected as 5mC and / or 5hmC. However, given that the levels of 5fC and 5caC in genomic DNA under normal conditions are extremely low, this will generally be acceptable when analyzing methylation and hydroxymethylation in DNA samples. The 5fC and 5caC signals can be eliminated by protecting 5fC and 5caC from conversion to DHU, respectively, by, for example, hydroxylamine conjugation and EDC coupling. According to these embodiments, the method identifies the position and percentage of 5hmC in DNA by comparing the position and percentage of 5mC with the position and percentage of 5mC or 5hmC (together). Alternatively, the position and frequency of 5hmC modifications in DNA can be directly measured.

[0084] In some embodiments, identifying 5fC and / or 5caC provides the location of 5fC and / or 5caC, but does not distinguish between the two cytosine modifications. Instead, both 5fC and 5caC are converted to DHU, which is detected by the methods described herein.

[0085] Methods for identifying 5caC. In some embodiments, the methods include identifying 5caC in a DNA sample (targeted DNA or whole genome) and providing a quantitative measure of the frequency of 5caC modifications at each position where the modification is identified in the DNA. In some embodiments, the percentage of T at each transition position provides a quantitative level of 5caC at each position in the DNA. According to these embodiments, the method for identifying 5caC may include the use of a blocking group. In other embodiments, the method for identifying 5caC does not require the use of a blocking group.

[0086] In some embodiments, when 5fC is blocked (and 5mC and 5hmC are not converted to DHU), identification of 5caC in DNA can occur. In some embodiments, adding a blocking group to 5fC in a DNA sample includes contacting the DNA with an aldehyde-reactive compound, including, for example, hydroxylamine derivatives, hydrazine derivatives, and hydrazide derivatives. Hydroxylamine derivatives include hydroxylamine (ashydroxylamine); hydroxylamine hydrochloride; acid hydroxylammonium sulfate; hydroxylamine phosphate; O-methylhydroxylamine; O-hexylhydroxylamine; O-pentylhydroxylamine; O-benzylhydroxylamine; and in particular O-ethylhydroxylamine (EtONH2), O-alkylated or O-arylated hydroxylamine, its acid or salt. Hydrazine derivatives include N-alkylhydrazine, N-arylhydrazine, N-benzylhydrazine, N, N-dialkylhydrazine, N, N-diarylhydrazine, N, N-dibenzylhydrazine, N, N-alkylbenzylhydrazine, N, N-arylbenzylhydrazine and N, N-alkylarylhydrazine. The hydrazide derivatives include toluenesulfonyl hydrazide, N-hydrazide, N,N-alkylhydrazide, N,N-benzylhydrazide, N,N-arylhydrazide, N-sulfonylhydrazide, N,N-alkylsulfonylhydrazide, N,N-benzylsulfonylhydrazide and N,N-arylsulfonylhydrazide.

[0087] Method for identifying 5fC. In some embodiments, the method includes identifying 5fC in a DNA sample (targeted DNA or whole genome), and providing a quantitative measurement of the frequency of 5fC modifications at each position identified in the DNA. In some embodiments, the percentage of T at each transition position provides a quantitative level of 5fC at each position in the DNA. According to these embodiments, the method for identifying 5fC may include the use of a blocking group. In other embodiments, the method for identifying 5fC does not require the use of a blocking group.

[0088] In some embodiments, adding a blocking group to 5caC in a DNA sample can be accomplished by (i) contacting the DNA sample with a coupling agent, such as a carboxylic acid derivatizing agent, such as a carbodiimide derivative, such as 1-ethyl-3-(3-dimethylaminopropyl)carbodiimide (EDC) or N,N'-dicyclohexylcarbodiimide (DCC), and (ii) contacting the DNA sample with an amine, hydrazine, or hydroxylamine compound. Thus, for example, 5caC can be blocked by treating the DNA sample with EDC and then treating the DNA sample with benzylamine, ethylamine, or another amine to form an amide that prevents conversion of 5caC to DHU (e.g., by borane reduction).

[0089] 3. Sequencing Library

[0090] The present disclosure provides a method for obtaining a methylation signature. In some embodiments, the method includes isolating DNA (e.g., cells or cfDNA) from a sample; preparing a sequencing library comprising the DNA; and performing TET-assisted pyridine borane sequencing (TAPS) on the sequencing library to obtain a methylation signature of the DNA. In some embodiments, the methylation signature is a genome-wide methylation signature.

[0091] In some embodiments, preparing the sequencing library includes connecting the sequencing adapter to the separated DNA to facilitate the sequencing reaction. Sequencing adapters suitable for large-scale parallel sequencing technology can be used. The present invention is not limited to any specific sequencing technology. In some preferred embodiments, those provided by Illumina or Nanopore can be utilized. For example, sequencing technology suitable for the present invention includes but is not limited to those described in U.S. Patent Publication 20100120098, U.S. Patent Publication 20120208705, U.S. Patent Publication 20120208724, WO2012 / 061832 and U.S. Patent Publication 2015 / 0368638, each of which is incorporated herein by reference in its entirety.

[0092] In some embodiments, the adapter comprises one or more sites that can hybridize with a primer. In some embodiments, the adapter comprises at least a first primer site. In some embodiments, the adapter comprises at least a first primer site and a second primer site. The orientation of the primer sites in such embodiments can be such that the primer hybridized with the first primer site and the primer hybridized with the second primer site are in the same orientation or different orientations. In one embodiment, the primer sequence in the adapter can be complementary to the primer used for amplification. In another embodiment, the primer sequence is complementary to the primer used for sequencing.

[0093] In some embodiments, the joint can include the first primer site, the second primer site and the non-amplification site therebetween. The non-amplification site is used to block the extension of the polynucleotide chain between the first and second primer sites, wherein the polynucleotide chain hybridizes with one of the primer sites. The non-amplification site can also be used to prevent concatemers. The example of the non-amplification site includes nucleotide analogs, non-nucleotide chemical parts, amino acids, peptides and polypeptides. In some embodiments, the non-amplification site includes nucleotide analogs that are not obviously base paired with A, C, G or T.

[0094] Some embodiments include an adapter comprising a first primer site, a second primer site, and a fragmentation site therebetween. Other embodiments may use a forked or Y-shaped adapter design that can be used for directional sequencing, as described in U.S. Pat. No. 7,741,463, which is incorporated herein by reference.

[0095] In some embodiments, the adapter may include an index or barcode sequence. In further preferred embodiments, the adapter may include a unique molecular identifier (UMI).

[0096] In some embodiments, a carrier nucleic acid or a mixture of carrier nucleic acids (such as DNA) is added to the sequencing library prior to performing TAPS. The carrier nucleic acid can be any specific or non-specific DNA molecule (or its nucleic acid derivative) that enhances one or more aspects of recovering DNA from a sample.

[0097] As specified above, in some preferred embodiments, a nucleic acid containing a DHU residue is subjected to a complementary chain synthesis step with a first polymerase or polymerase mixture that is tolerant to the DHU residue and / or the products produced by the introduction of the DHU residue and / or the TAPS process, and an exponential amplification step with a second polymerase or polymerase mixture. In some embodiments, the second polymerase or polymerase mixture may be the same as the first polymerase or polymerase mixture, or may be a polymerase or polymerase mixture that is different from the first polymerase or polymerase mixture. In certain embodiments in which the second polymerase mixture is different from the first polymerase mixture, the second polymerase or polymerase mixture may also be tolerant to the DHU residue and / or the products produced by the introduction of the DHU residue into the target nucleic acid molecule and / or the TAPS process. In embodiments in which the same DHU and / or TAPS-tolerant polymerase is used for both complementary chain synthesis and exponential amplification, it should be understood that the initial complementary chain synthesis step may be part of the exponential amplification process. The complementary chain synthesis step and the amplification step may be performed before or after the incorporation of the sequencing adapter sequence into the target nucleic acid. In some preferred embodiments, the first polymerase or polymerase mixture is characterized by having a DNA sequence greater than 5.0X10 -5In some preferred embodiments, performing the complementary chain synthesis step using a DHU and / or TAPS-tolerant polymerase or polymerase mixture results in increased coverage of highly methylated (and therefore DHU-rich) regions compared to the same protocol without performing the complementary chain synthesis step using a DHU and / or TAPS-tolerant polymerase or polymerase mixture.

[0098] DHU and / or TAPS first polymerases and polymerase mixtures suitable for the complementary strand synthesis step include, but are not limited to, Bst 3.0 polymerase (New England Biolabs, Beverly, MA), Sulfolobus DNA polymerase IV (New England Biolabs, Beverly, MA), a combination of Bst 3.0 polymerase and Sulfolobus DNA polymerase IV, Klenow polymerase (New England Biolabs, Beverly, MA), Klenow exopolymerase (Thermo Fisher Scientific, Grand Island, NY), Pol κ polymerase, Mu-mLV reverse transcriptase, SD polymerase (Bioron), Tth polymerase (Sigma Aldrich, St. Louis, MO), OneTaq TM (New England Biolabs, Beverly, MA), OneTaq TM (New England Biolabs, Beverly, MA) and Tth polymerase, 5D4 polymerase and a mixture of 5D4 polymerase and Taq polymerase. In some preferred embodiments, the first polymerase or polymerase mixture is thermostable. In other preferred embodiments, the first polymerase or polymerase mixture is thermostable.

[0099] The polymerase used in the exponential amplification step can be any polymerase suitable for amplification and / or sequencing, and can be the same or different from the first polymerase. In some preferred embodiments, the polymerase used in the exponential amplification is different from the first polymerase and is represented as a second polymerase or polymerase mixture. In some particularly preferred embodiments, the polymerase used in the exponential amplification step has an error rate less than the first polymerase or polymerase mixture. In some preferred embodiments, the second polymerase or polymerase mixture is characterized as a high-fidelity polymerase. In some preferred embodiments, the second polymerase or polymerase mixture is characterized in that the error rate is less than 5.0X10 -5 In some preferred embodiments, the second polymerase or polymerase mixture is selected from Taq polymerase, such as GoTaq TMPolymerase (Promega, Fitchburg, WI), and engineered B-family polymerases, such as KAPA HiFi Uracil+ polymerase (Roche, Indianapolis, IN). In some preferred embodiments, the first polymerase or polymerase mixture is thermostable.

[0100] In some preferred embodiments, the complementary strand synthesis step utilizes forward and / or reverse primers that anneal to the sequencing adapter. In some preferred embodiments, the exponential amplification step utilizes sequencing primers that anneal to a region of the sequencing adapter. Figure 2 .

[0101] Considering that DNA methylation features can be used to understand basic biological processes and disease pathology and disease detection. For example, methylation features / frequency / markers, etc. can be used to understand and study gene regulation, genomic imprinting, differentiation, development, gene-environment interactions (such as smoking, nutrition), aging, various diseases and conditions (such as autoimmune diseases, cancer, cardiovascular disease, CNS disease, congenital diseases, infectious diseases, metabolic diseases and states, NIPT related tests, etc.), for detecting and diagnosing cancer and other diseases and for monitoring transplants. In some embodiments and as described herein, the method also includes identifying at least one methylation biomarker from a cfDNA whole genome methylation feature (such as a whole genome DNA methylation feature), and determining that the methylation biomarker is different from a methylation biomarker in a reference or control sequence. In some embodiments, the methylation biomarker comprises a differentially methylated region (DMR). In some embodiments, the method also includes classifying samples based on DMR compared to a reference DMR. In some embodiments, the reference DMR corresponds to a non-disease control or a disease control.

[0102] In some embodiments and as described herein, the method further comprises identifying at least one methylation biomarker from the DNA methylation signature and determining the tissue of origin corresponding to the methylation biomarker. In some embodiments, the method further comprises classifying the sample based on the tissue of origin biomarker.

[0103] In some embodiments and as described herein, the method further comprises identifying a DNA fragmentation profile, and determining whether the fragmentation profile is indicative of cancer.According to these embodiments, a DNA fragmentation profile may be determined from TAPS sequencing data (eg, read pair alignment positions).

[0104] In some embodiments, the method also includes identifying at least one sequence variant from the DNA sample, and determining whether the sequence variant indicates cancer. For example, in some embodiments, TAPS can also distinguish methylation from C to T genetic variants or single nucleotide polymorphisms (SNPs), and can therefore be used to detect genetic variants. In some embodiments, methylation and C to T SNPs can lead to different patterns in TAPS. For example, methylation can lead to T / G reads in the original top chain / original bottom chain, and A / C reads in the chains complementary to these. In some embodiments, C to T SNPs can lead to T / A reads in the original top chain / original bottom chain and in the chains complementary to these. This further increases the utility of TAPS in providing both methylation information and genetic variants and therefore providing mutations in one experiment and sequencing runs. This ability of the TAPS method disclosed herein provides the integration of genomic analysis with epigenetic analysis, and significantly reduces sequencing costs by eliminating the need to perform, for example, standard whole genome sequencing (WGS).

[0105] According to the above embodiments, the methods of the present disclosure include using TAPS in a single experiment to generate information related to methylation signatures, methylation biomarkers, DNA fragmentation profiles, DNA sequence information (such as variants), and tissue of origin information to diagnose / detect a subject's disease or other condition (such as those provided in the above examples). As will be recognized by those of ordinary skill in the art based on the present disclosure, TAPS as disclosed herein can be used to generate any combination of methylation signatures, methylation biomarkers, DNA fragmentation profiles, DNA sequence information (such as variants), and tissue of origin information to diagnose / detect a subject's disease or other condition (such as those provided in the above examples). In some embodiments, methylation signatures can be obtained, and one or more of methylation biomarkers, DNA fragmentation profiles, DNA sequence information (such as variants), and tissue of origin information can also be obtained and used to diagnose / detect a subject's disease or other condition (such as those provided in the above examples). In some embodiments, the methylation state of a biomarker can be obtained, and one or more of methylation signatures, DNA fragmentation profiles, DNA sequence information (such as variants), and tissue of origin information can also be obtained and used to diagnose / detect a subject's disease or other condition (such as those provided in the above examples). In some embodiments, a DNA fragmentation profile may be obtained, and one or more of methylation signatures, methylation biomarkers, DNA sequence information (such as variants), and tissue of origin information may also be obtained and used to diagnose / detect a disease or other condition in a subject (such as those provided in the examples above). In some embodiments, DNA sequence variants may be identified, and one or more of methylation signatures, methylation biomarkers, DNA fragmentation profiles, and tissue of origin information may also be obtained and used to diagnose / detect a disease or other condition in a subject (such as those provided in the examples above). In some embodiments, tissue of origin information may be obtained (e.g., from a genome-wide DNA methylation profile), and one or more of methylation signatures, methylation biomarkers, DNA fragmentation profiles, and DNA sequence information (such as variants) may also be obtained and used to diagnose / detect a disease or other condition in a subject (such as those provided in the examples above).

[0106] In some embodiments, performing TAPS on a sequencing library to obtain a genome-wide methylation signature includes identifying 5mC modifications in the DNA and providing a quantitative measurement of the frequency of the 5mC modifications. In some embodiments, performing TAPS on a sequencing library to obtain a genome-wide methylation signature includes identifying 5hmC modifications in the DNA and providing a quantitative measurement of the frequency of the 5hmC modifications. In some embodiments, performing TAPS on a sequencing library to obtain a genome-wide methylation signature includes identifying 5caC modifications in the DNA and providing a quantitative measurement of the frequency of the 5caC modifications. In some embodiments, performing TAPS on a sequencing library to obtain a genome-wide methylation signature includes identifying 5fC modifications in the DNA and providing a quantitative measurement of the frequency of the 5fC modifications.

[0107] Based on the disclosure, it will be appreciated by those of ordinary skill in the art that the methods described herein (e.g., TAPS) can be used to diagnose / detect any type of cancer. The types of cancer that can be detected / diagnosed using the methods disclosed herein include, but are not limited to, lung cancer, melanoma, colon cancer, colorectal cancer, neuroblastoma, breast cancer, prostate cancer, renal cell carcinoma, transitional cell carcinoma, bile duct cancer, brain cancer, non-small cell lung cancer, pancreatic cancer, liver cancer, gastric cancer, bladder cancer, esophageal cancer, mesothelioma, thyroid cancer, head and neck cancer, osteosarcoma, hepatocellular carcinoma, primary unknown cancer, ovarian cancer, endometrial cancer, glioblastoma, Hodgkin lymphoma (Hodgkin lymphoma) and non-Hodgkin lymphoma (non-Hodgkin lymphomas). In some embodiments, the types of cancer or the metastatic forms of cancer that can be detected / diagnosed by the methods disclosed herein include, but are not limited to, cancer, sarcoma, lymphoma, germ cell tumors and blastoma. In some embodiments, cancer is invasive and / or metastatic cancer (e.g., II stage cancer, III stage cancer or IV stage cancer). In some embodiments, the cancer is an early stage cancer (eg, stage 0 cancer, stage I cancer), and / or is not an invasive and / or metastatic cancer.

[0108] According to these embodiments, the present disclosure provides a method for quantitatively identifying the position of one or more of 5mC, 5hmC, 5caC and / or 5fC in nucleic acids at base resolution without affecting unmodified cytosine. In some embodiments, the nucleic acid is DNA. In some embodiments, the DNA is cfDNA (e.g., circulating cfDNA). In some embodiments, the nucleic acid is RNA. In some embodiments, the nucleic acid sample includes a target nucleic acid that is DNA or a target nucleic acid that is RNA. In some embodiments, the method is applied to the entire genome and is not limited to a specific target nucleic acid.

[0109] The nucleic acid can be any nucleic acid with a cytosine modification (i.e., 5mC, 5hmC, 5fC, and / or 5caC), but is not limited to DNA fragments and / or genomic DNA. The nucleic acid can be a single nucleic acid molecule in a sample, or it can be a whole population of nucleic acid molecules in a sample or any portion thereof (the whole genome or a subset thereof). The nucleic acid can be a natural nucleic acid from a source (such as a cell, a tissue sample, etc.), or it can be pre-converted into a high-throughput sequencing-ready form, for example, by fragmentation, repair, and connection to an adapter for sequencing. Therefore, the nucleic acid can comprise a plurality of nucleic acid sequences, so that the methods described herein can be used to generate a library of target nucleic acid sequences that can be analyzed individually (e.g., by determining the sequence of an individual target) or in groups (e.g., by a high-throughput or next-generation sequencing method).

[0110] Because the methods of the present disclosure utilize mild enzymatic and chemical reactions, avoiding substantial degradation of nucleic acids associated with methods such as bisulfite sequencing, the methods of the present disclosure can be used to analyze low-input samples, such as circulating cell-free DNA and for single-cell analysis.

[0111] In some embodiments, the DNA sample comprises a picogram amount of DNA. In some embodiments, the DNA sample comprises about 1pg to about 900pg DNA, about 1pg to about 500pg DNA, about 1pg to about 100pg DNA, about 1pg to about 50pg DNA, or about 1 to about 10pg DNA. In some embodiments, the DNA sample comprises less than about 200pg, less than about 100pg DNA, less than about 50pg DNA, less than about 20pg DNA, less than about 15pg DNA, less than about 10pg DNA, or less than about 5pg DNA.

[0112] In some embodiments, DNA sample comprises the DNA of nanogram amount.The sample DNA used in the method of the present disclosure can be any amount, including but not limited to the DNA from unicellular or a large amount of DNA samples.In some embodiments, the method can be carried out to the DNA sample comprising about 1 to about 500ng DNA, about 1 to about 200ng DNA, about 1 to about 100ng DNA, about 1 to about 50ng DNA, about 1 to about 10ng DNA, about 2 to about 5ng DNA.In some embodiments, DNA sample comprises less than about 100ng DNA, less than about 50ng DNA, less than 40ng DNA, less than 30ng DNA, less than 20ng DNA, less than 15ng DNA, less than 5ng DNA and less than 2ng DNA.In some embodiments, DNA sample comprises the DNA of microgram amount.

[0113] The method of the present disclosure may also include a step of amplifying the copy number of the modified nucleic acid by methods known in the art. When the modified nucleic acid is DNA, the copy number can be increased by, for example, PCR, cloning and primer extension. The copy number of a single target DNA can be amplified by PCR using primers specific to a specific target DNA sequence. Alternatively, multiple different modified target DNA sequences can be amplified by cloning into a DNA vector by standard techniques. In some embodiments, the copy number of multiple different modified target DNA sequences is increased by PCR to produce a library for next generation sequencing, wherein, for example, double-stranded adapter DNA has been pre-connected to sample DNA (or connected to modified sample DNA), and PCR is performed using primers complementary to the adapter DNA.

[0114] In some embodiments, the method includes a step of detecting the sequence of the modified nucleic acid. The modified target DNA or RNA contains DHU at the position where one or more of 5mC, 5hmC, 5fC and 5caC are present in the unmodified target DNA or RNA. DHU acts as T in DNA replication and sequencing methods. Therefore, cytosine modification can be detected by any direct or indirect method known in the art for identifying C to T conversion. Such methods include sequencing methods, such as Sanger sequencing, microarrays and next generation sequencing methods. C to T conversion can also be detected by restriction enzyme analysis, wherein C to T conversion abolishes or introduces restriction endonuclease recognition sequences.

[0115] Embodiments of the present disclosure also provide a kit for identifying 5mC and 5hmC in DNA. Such a kit includes reagents for identifying 5mC and 5hmC by the method described herein. The kit may also contain reagents for identifying 5caC and identifying 5fC by the method described herein. In some embodiments, the kit includes a TET enzyme, a borane reducing agent, and instructions for performing the method. In some embodiments, the borane reducing agent is selected from one or more of the group consisting of pyridine borane, 2-methylpyridine borane (pic-BH3), borane, sodium borohydride, sodium cyanoborohydride, and sodium triacetoxyborohydride. In a further preferred embodiment, the kit includes the first and second polymerases or polymerase mixtures as described in detail above.

[0116] In some embodiments, the kit also includes a 5hmC blocking group and a glucosyltransferase. In some embodiments, the blocking group added to 5hmC is a sugar. In some embodiments, the sugar is a naturally occurring sugar or a modified sugar, such as glucose or a modified glucose. In some embodiments, the blocking group is added to 5hmC by contacting the nucleic acid sample with a UDP (e.g., UDP-glucose) connected to a sugar or a UDP connected to a modified glucose in the presence of an enzyme glucosyltransferase, such as T4 bacteriophage β-glucosyltransferase (βGT) and T4 bacteriophage α-glucosyltransferase (αGT) and derivatives and analogs thereof.

[0117] In some embodiments, the kit further comprises an oxidizing agent selected from manganese oxide (MnO2), potassium ruthenate (K2RuO4), potassium perruthenate (KRuO4) and / or Cu(II) / TEMPO (copper (II) perchlorate and 2,2,6,6-tetramethylpiperidine-1-oxyl (TEMPO)). In some embodiments, the kit comprises a reagent for blocking 5fC in a nucleic acid sample. In some embodiments, the kit comprises an aldehyde-reactive compound, including, for example, hydroxylamine derivatives, hydrazine derivatives, and hydrazide derivatives as described herein. In some embodiments, the kit comprises a reagent for blocking 5caC as described herein. In some embodiments, the kit comprises a reagent for isolating DNA or RNA. In some embodiments, the kit comprises a reagent for isolating low input DNA from a sample, such as isolating cfDNA from blood, plasma, or serum.

[0118] In some embodiments, the method of the present disclosure includes treating a patient (e.g., a patient suffering from cancer, suffering from early stage cancer, or suspected of suffering from cancer). In some embodiments, the method includes determining a methylation signature as provided herein and administering treatment to the patient based on the results of determining the methylation signature. Treatment may include administering a pharmaceutical compound, a vaccine, performing surgery, imaging the patient, and / or performing another test. In some embodiments, the method of the present disclosure may be used as a part of a method for clinical screening, a prognostic assessment method, a method for monitoring therapy results, a method for identifying patients most likely to respond to a specific therapeutic treatment, a method for imaging a patient or subject, and a method for drug screening and development.

[0119] Unless defined otherwise in this article, the scientific and technical terms used in conjunction with the present disclosure should have the meaning commonly understood by those skilled in the art. For example, any nomenclature used in conjunction with cell and tissue culture, molecular biology, immunology, microbiology, genetics, and protein and nucleic acid chemistry and hybridization described herein is well-known and commonly used in the art. The meaning and scope of the term should be clear, but if there is any implicit ambiguity, the definition provided herein takes precedence over any dictionary or external definition. In addition, unless the context requires otherwise, singular terms should include plural terms and plural terms should include singular terms.

[0120] 4. Examples

[0121] It will be apparent to those skilled in the art that other suitable modifications and adaptations of the disclosed methods described herein are readily applicable and understandable, and can be made using suitable equivalents without departing from the scope of the disclosure or the aspects and embodiments disclosed herein. Having now described the disclosure in detail, it will be more clearly understood by reference to the following examples, which are intended only to illustrate some aspects and embodiments of the disclosure and should not be construed as limiting the scope of the disclosure. All journal references, U.S. patents, and published disclosures cited herein are hereby incorporated by reference in their entirety.

[0122] The present disclosure has several aspects, which are illustrated by the following non-limiting examples.

[0123] Example 1

[0124] NGS libraries generated by TET-assisted pyridine borane sequencing (TAPS) showed reduced coverage of DNA regions with high density of methylated cytosines relative to average coverage ( Figure 1 ).exist Figure 1 In the present study, the normalized GC bias metric referenced to a fully methylated lambda DNA sequence was used as a surrogate for hypermethylated DNA, and the difference in the curves for TAPS-treated and non-TAPS-treated fully methylated lambda DNA represents the difference in coverage of methylated regions seen in biological samples. Underrepresentation of clinically relevant regions of methylated biological samples can lead to reduced sensitivity and the need for increased sequencing to achieve adequate coverage. To improve sequencing efficiency of TAPS libraries, we found that using a polymerase that tolerates DHU or other TAPS products for the initial complementary strand synthesis step, when combined with typical library amplification techniques, improved coverage of DNA hypermethylated regions as measured by the normalized GC bias.

[0125] Some NGS sequencing library adapters ( Figure 2The Y-shaped i5 and i7 adapters (shown in green in Figure 2) were ligated to the fragmented DNA before or after TAPS treatment. During TAPS treatment, TET oxidation and borane reduction convert methylated cytosine (meC) to dihydrouracil (DHU) bases ( Figure 2 (shown in purple box in the figure).

[0126] During amplification of these libraries using index primers, the polymerase will insert an adenine opposite the DHU base, resulting in subsequent conversion to thymidine in the next amplification step. However, many polymerases do not favor the replication of the opposite DHU base, or the products generated by the introduction of the DHU residue and / or the TAPS process.

[0127] This initial replication via the DHU base is performed by an alternative polymerase during the complementary strand synthesis step. Figure 2 The partial adapter shown, the complementary strand synthesis step uses one of a variety of polymerases with a small bias towards DHU, such as Bst3.0, and a reverse primer (such as Figure 2 b) allowing the replication of the library molecules under constant temperature incubation (with or without an initial denaturation step). This results in the DHU base in the adapter-ligated DNA now being converted to adenine ( Figure 2 c). The DNA is then subjected to a typical PCR amplification using an ultra-high fidelity polymerase (such as Kapa HiFi Uracil+) with excellent GC bias coverage and the addition of a forward index primer and a library amplification primer to generate enough DNA for sequencing. We also found that the synthesis step using two index primers (full-length i5 and i7) achieved the same results. It is expected that similar conditions will be beneficial for amplifying libraries generated from DNA connected by full-length adapters.

[0128] Tables 1 to 3 provide exemplary reagents utilized in the sequencing methods described herein:

[0129] Table 1: Reagents for complementary strand synthesis

[0130]

[0131]

[0132] *Primer = complete reverse i7, or complete reverse i7 and forward i5

[0133] + If a denaturation step is included, a thermolabile polymerase must be added after the denaturation step

[0134] Table 2: Complementary strand synthesis incubation steps

[0135]

[0136] *Temperature depends on polymerase

[0137] + Typically 1 cycle, but 2-4 cycles have also been tested

[0138] Table 3: Subsequent processing

[0139]

[0140] * If intact reverse i7 and forward i5 are used in the complementary strand synthesis step, no additional primers are required

[0141] Figures 3 to 14 Provides Kapa Hyperprep TM Sequencing results for fully methylated lambda spikes prepared with various polymerases in the complementary strand synthesis and amplification steps. Bst 3.0 polymerase demonstrated improved coverage uniformity with multiple workflow options and in combination with several other polymerases and reverse transcriptases. Figure 8 Shown is the presence of OneTaq buffer and MnSO4. TM We further optimized the method and showed that the same improvement could be obtained using OneTaq alone, Tth alone, or Taq polymerase alone together with MnSO4, with 0.5-0.75 mM MnSO4 being found to be optimal ( Fig. 9 and 10 Unless otherwise stated, OneTaq and Tth conditions had 0.75 mM MnSO4. Fig.11 , 12 and 13 show improvements using polymerase K, Klenow exo-, and SD polymerases, respectively.

[0142] We also found that engineered polymerase 5D4 performed well under complementary strand synthesis step conditions either alone or as a spike in with Kapa HiFi Uracil+ ( Fig.14 We noted that 5D4 also improved coverage of high DHU regions when combined with Taq polymerase or as a spike-in to KapaHiFiUracil+ for library amplification (without an initial complementary strand synthesis step). Fig.15 ).

[0143] We further found that in high-coverage whole-genome sequencing, selection of favorable complementary strand synthesis conditions resulted in improved coverage of marked regions that are often highly methylated and / or have a high density of CpG sites, resulting in lower than average coverage when sequenced under standard conditions. Shown here are the complementary strand synthesis step conditions for Bst3.0 ( Fig.16A -B) and OneTaqTM and Tth complementary chain synthesis step conditions ( Fig.17A -B). We also see benefits from competition between the methylated and unmethylated versions of the target sequence. This manifests as an increase in the methylation signal for the less methylated marker. Fig.24 and 25 . Fig.24 Shown are normalized coverage of selected marker regions with low levels of methylation after TAPS and amplification with different polymerases (Kapa Hifi Uracil+, Bst, and OTT). Fig.25 Shown are the conversion rates of selected marker regions after TAPS and amplification with different polymerases (Kapa Hifi Uracil+, Bst, and OTT).

[0144] Initial complementary strand synthesis using SD polymerase also showed similar or greater normalized coverage than the Bst complementary strand synthesis step performed on highly methylated marker regions in the NA12878 whole genome sequencing ( Fig.18A -B). Missing data points for methylation in some regions highlight the low coverage of these highly methylated regions, and in this case, while coverage is improved relative to no initial complementary strand synthesis step, there are not enough reads to confidently determine average methylation.

[0145] Furthermore, we show that the improved coverage under complementary strand synthesis step conditions using Bst or OneTaq and Tth (OTT) is also achieved in whole genome sequencing (WGS) ( Fig.19A -B- Improved coverage markers b, d, e, h, i) and hybrid capture targeted sequencing ( Fig. 20A -B- Increased coverage was observed on selected hypermethylated markers within the normal cfDNA pool in both markers AC, E).

[0146] like Figure 21-23 As shown, the Accel-NGS Methyl-Seq DNA library kit from Swift BioSciences, the SRSLY NGS library kit from Claret Bioscience, and the EpiXplore TM For the Methylated DNA kit, the selected complementary strand synthesis step option also showed improvement.

[0147] Many other polymerases were screened that did not increase the normalized GC bias when used in the complementary strand synthesis step under the above test conditions. Examples of these polymerases include KAPA HiFi Uracil+, full-length Bst polymerase, Therminator polymerase, phi29 polymerase, AMV reverse transcriptase, Taq polymerase (NEB), NEB Q5U polymerase, NEBLongAmp Taq, Pyromark, Phusion U<SeqAmp, and ProtoScript TM Reverse transcriptase. See, for example, Fig.26 and 27 , which provide the results of the complementary strand synthesis step using SeqAmp and Therminator polymerases, respectively, compared to Bst polymerase.

[0148] Example 2

[0149] This example provides data related to the optimization of conditions for converting 5-carboxycytosine (5caC) and / or 5-formylcytosine (5fC) residues in an oxidized nucleic acid sample to dihydrouracil (DHU) residues via the use of a borane reducing agent (e.g., pic-borane). These data indicate that the conversion rate of 5caC to dihydrouracil (DHU) using borane reduction chemistry using previously established conditions or conditions described in Nature Biotechnology (37) 424-429 (2019) can be increased by adding 45% to 52.5% v / v of an organic solvent such as dimethyl sulfoxide (DMSO) to the reaction mixture. The data further indicate that in addition to using an increased amount of solvent in the reaction, improved results can also be obtained using a reaction temperature of 45 to 52.5°C and / or a shortened reaction time of 45 to 60 minutes.

[0150] Compared with alternative borane reduction conditions, using a higher concentration of solvent (such as DMSO) in the reaction has multiple benefits. For example, the previous conditions using 10% DMSO resulted in a significant deviation in library amplification after the borane reaction. As described in this article, the use of reaction conditions with increased solvent concentration can effectively reduce this deviation. Alternative borane reduction condition optimization has been attempted, including changing the conditions by increasing temperature, shortening reaction time, changing borane concentration, and using different buffer and pH conditions. However, these previous optimization efforts did not lead to the beneficial optimization of DMSO concentration shown in this article. For example, compared with the conditions using increased solvent levels, which can maintain a low false positive rate and increased genome coverage, the alternative conditions of increasing the reaction temperature alone can lead to an increase in the false positive rate, while the conditions of reducing the reaction time alone can lead to a decrease in the conversion rate. Therefore, the data show that the solvent concentration mediates the effects of reducing the reaction time (usually reducing the conversion rate) and increasing the reaction temperature (usually increasing the false positive (FP) rate and deviation).

[0151] Specifically, when only the reaction time was changed (shortened from 2 hours to 1 hour) without changing other conditions (37°C, with 10% DMSO), the conversion rate dropped from 92% to 85% (measured using spiked methylated pUC19). When the reaction time was reduced from 2 hours to 1 hour, while the reaction temperature was increased to 50°C, the conversion rate (measured using spiked methylated Lambda template) increased to about 94%, but this increase was also accompanied by an increase in the false positive rate to about 2%. That represents an increase of about 6 times from the false positive rate of about 0.35% observed under standard conditions (10% DMSO, 37°C, 2 hours). Unexpectedly, when the concentration of DMSO was increased to 50% in the reaction carried out at 50°C for 1 hour, the high conversion rate (average over 94%) was still maintained, while the false positive rate was reduced to an average of less than about 0.3%. Table 4 provides supporting data for these observations. GC bias graphs (not shown) were also generated for TAPS reactions using different polymerases (Hifi HotStart KAPA Uracil plus, Bst 3.0, and OneTaq / Tth) at 37°C for 2 hours with 10% DMSO or at 50°C for 1 hour with 50% DMSO. The graphs for the reactions using 50% DMSO are flatter than those for the reactions using 10% DMSO, demonstrating the increased coverage and decreased GC bias for the reactions containing 50% DMSO. Additional data that further defines the optimal ranges for time, temperature, and solvent concentrations are discussed after Table 4.

[0152] Table 4: Results of different time, temperature and DMSO concentration

[0153]

[0154] Unless otherwise stated, the experiments described herein compare optimized reaction conditions with base reaction conditions that vary DMSO concentration, reaction temperature, and reaction time. Base reaction conditions utilize a 50 μl reaction volume containing 50 ng of oxidized dsDNA, 100 mM buffer (pH 4.0) (5 μl), 100 mM Pic-borane in DMSO (5 μl) (providing 10% v / v DMSO), where the reaction is performed at 37°C for 2 hours. Multiple values ​​of solvent concentration, reaction time, and reaction temperature were evaluated to determine the optimal range and are reported in the table below. Multiple experimental parameters under different conditions were examined and reported in the table below.

[0155] In short, the methylation conversion rate represents the C->T conversion rate after TAPS detection in a fully methylated Lambda or partially methylated pUC19 spike-in. The pUC19 DNA spike contains approximately 20% methylation and is intended to represent real-world conditions where less than 100% of the template is methylated. A false positive is defined as the detection of a C->T conversion on a completely unmethylated 2kb spike-in. GC bias describes the dependency between fragment counts (read coverage) and GC content found in Illumina sequencing data. GC loss is a metric related to the degree of sequencing bias in a sample, whereby samples with greater GC bias have correspondingly higher GC loss. For example, Lambda GC loss therefore represents the sequencing bias of a methylated Lambda DNA spike-in to a TAPS reaction, in which most Cs are methylated.

[0156] The effect of increasing the DMSO concentration in the reaction mixture was evaluated. TAPS reactions were performed at 50°C for 1 hour using different DMSO concentrations (including 0%, 10%, 25%, 50%, 60% and 75% DMSO v / v). The data are presented in Table 5. As can be seen, the TAPS reaction using 50% DMSO has similar conversion levels to 10% DMSO (as determined by Lambda methylation), increased conversion levels (as determined by pUC19 methylation) (which is more representative of real time conditions), greatly increased false positive rates, and improved GC loss indicators. Improvements relative to GC bias were also observed in the GC graph, where the curve for the 50% DMSO reaction was flatter than the curve for the reaction containing a lower amount of DMSO. The TAPS reaction containing 25% DMSO had good conversion, but had a relatively high false positive rate and a high GC bias (as represented by the GC loss indicator). When the DMSO concentration was increased to 60% or 75%, the conversion rate decreased.

[0157]

[0158] The influence of increasing the reaction temperature and using 10% or 50% DMSO v / v was further evaluated. The conversion rate and false positive rate at different temperatures (including 20°C, 37°C, 45°C and 50°C), using 10% and 50% DMSO v / v and a reaction time of 1 hour were measured. The data are presented in Table 6. The conversion rate, false positive and GC loss index under higher temperatures (including 50°C, 75°C and 100°C, each with 10% or 50% DMSO v / v) were also measured. The data are presented in Table 7. With reference to Table 6, the data show that the false positive rate is reduced using the increased DMSO concentration. Specifically, at a reaction temperature higher than 37°C, the false positive rate is reduced compared to using 10% DMSO in the reaction mixture using 50% DMSO. With reference to Table 7, the data show that at a reaction temperature higher than 50°C, the false positive rate begins to increase, and the recoverable DNA is also reduced. As indicated by the GC loss index, the use of 50% DMSO also results in a reduction in GC bias at all evaluated temperatures. The effect on GC bias was confirmed by the GC bias graph (not shown), which demonstrated that the curve was flatter when 50% DMSO was included in the reaction mixture.

[0159] Table 6. Effect of increasing reaction temperature on TAPS reaction.

[0160]

[0161]

[0162] The effect of the time of the reaction with 10% or 50% DMSO v / v was evaluated. The conversion rate and false positive rate of the reaction were measured for each reaction of 15 minutes or 1 hour using 10% or 50% DMSO v / v at 37°C or 50°C. The conversion rate and GC deviation (expressed as GC loss index) of the reaction for each reaction of 1 hour, 2 hours, 6 hours or 24 hours using 50% DMSO v / v and at a reaction temperature of 50°C were also measured. The data are presented in Tables 8 and 9. With reference to Table 8, the data show that for shorter reaction times, increasing DMSO does not have any benefit, as demonstrated by the lower conversion rate of the reaction at 37°C or 50°C and using 10% or 50% DMSO v / v for 15 minutes. With reference to Table 9, the data show that the longer reaction times of 2 hours, 6 hours and 24 hours begin to increase the false positive rate and reduce the yield. The GC loss index also shows that longer reaction times are generally associated with increased GC deviations. The effect on GC bias was confirmed by the GC bias graph (not shown), which demonstrated flatter curves, especially for the 1 and 2 hour reactions compared to the 6 and 24 hour reactions.

[0163] Table 8. Effect of time on TAPS reaction at different DMSO concentrations.

[0164]

[0165]

[0166] Additional experiments were performed to further determine the optimal range of DMSO concentration (35%, 40%, 45%, 47.5%, 50%, 52.5%, 55% and 60% DMSO v / v were measured), reaction temperature (45°C, 47.5°C, 50°C, 52.5°C and 55°C were measured) and reaction time (45 minutes, 50 minutes, 55 minutes, 1 hour and 2 hours were measured). The collated data are presented in Table 10. From these data, it can be determined that the optimal range of DMSO is 45% to 52.5% v / v. Outside this range, a decrease in conversion or an increase in false positive rate is observed. The optimal reaction time can be further determined to be 45 minutes to 1 hour. Outside this range, an increase in false positive rate is observed, while the conversion rate remains high. Finally, the optimal reaction temperature can be determined to be 45°C to 52.5°C. Outside this range, an increase in false positive rate is observed.

[0167]

[0168]

[0169] Finally, the optimized conditions of 50% v / v DMSO and a reaction time of 1 hour at 50°C (designated ESI-NEW in Table 11) were compared with several sets of conditions, some of which had been used previously (designated ESI-OLD (10% DMSO, 37°C, 1 hour), CS (8% DMSO, 37°C, 1 or 4 hours), and NB (Nature 2010). Biotechnology (37) 424-429 (2019) (pyridine borane (PyB) or Pic-borane (PicB), 0 or 50% DMSO, 37°C or 50°C, 1, 3 or 16 hours). The data are presented in Table 11. These data show that the optimized conditions (ESI-NEW) provide the highest conversion levels of any tested conditions, especially for the pUC19 template test (i.e., greater than 96% for Lambda methylation and greater than 9.5% for pUC19 methylation). The optimized conditions also provide the highest yield while maintaining a low false positive rate and good GC loss indicators.

[0170]

[0171]

Claims

1. A method for amplifying a target nucleic acid molecule comprising dihydrouracil (DHU) residues, comprising: synthesizing one or more complementary strands of a target nucleic acid comprising a DHU residue using a first polymerase or polymerase mixture that is tolerant to the DHU residue and / or products resulting from the introduction of the DHU residue and / or the TAPS process to provide a target nucleic acid mixture comprising a target nucleic acid comprising a DHU residue and one or more complementary strands; and The target nucleic acid mixture is exponentially amplified to provide amplified target nucleic acids.

2. The method of claim 1, wherein the error rate of the first polymerase or polymerase mixture is greater than 5.0×10 -5 .

3. The method of any one of claims 1 to 2, wherein the first polymerase or polymerase mixture is selected from the group consisting of: Bst3.0 polymerase, Sulpholobus polymerase IV, a combination of Bst3.0 polymerase and Sulpholobus polymerase IV, Klenow polymerase, Klenow exopolymerase, Polκ polymerase, Mu-mLV reverse transcriptase, SD polymerase, Tth polymerase, OneTaq polymerase, a combination of OneTaq and Tth polymerase, 5D4 polymerase, a mixture of 5D4 polymerase and Taq polymerase, and SD polymerase.

4. The method of any one of claims 1 to 3, wherein the first polymerase is thermolabile.

5. The method of any one of claims 1 to 3, wherein the first polymerase is thermostable.

6. The method of any one of claims 1 to 3, wherein the step of exponentially amplifying the complementary strand of the target nucleic acid utilizes the first polymerase or polymerase mixture that is tolerant to DHU residues and / or products generated by the introduction of the DHU residues and / or the TAPS process.

7. The method of any one of claims 1 to 5, wherein the step of exponentially amplifying the pre-amplified target nucleic acid utilizes a second polymerase or polymerase mixture that is different from the first polymerase or polymerase mixture.

8. The method of claim 7, wherein the error rate of the second polymerase or polymerase mixture is less than 5.0×10 -5 .

9. The method of claim 7, wherein the error rate of the second polymerase or polymerase mixture is less than 1.0×10 -6 .

10. The method of any one of claims 7 to 9, wherein the second polymerase is selected from the group consisting of GoTaq polymerase and KAPA HiFi Uracil+ polymerase.

11. The method according to any one of claims 7 to 9, wherein the error rate is less than 5.0×10 -5 The polymerase is thermostable.

12. The method of any one of claims 7 to 10, wherein the first polymerase and the second polymerase are provided in a master mix.

13. The method of any one of claims 1 to 12, wherein synthesizing the complementary strand of the target nucleic acid comprising DHU residues with the first polymerase or polymerase mixture further comprises performing the synthesis in a buffer comprising about 0.5-0.75 mM MnSO4.

14. The method of any one of claims 1 to 13, further comprising quantifying the amplified target nucleic acid.

15. The method of any one of claims 1 to 14, further comprising the step of sequencing the exponentially amplified target nucleic acid.

16. The method of any one of claims 15, wherein the target nucleic acid comprising DHU residues has a sequencing library adapter attached to each end.

17. The method of claim 16, wherein the sequencing library adapter comprises an index sequence.

18. The method of any one of claims 16 to 17, wherein the sequencing library adaptor comprises a sequence complementary to a sequencing primer.

19. The method of any one of claims 15 to 18, wherein the sequencing library adapter comprises a sequence complementary to an index primer.

20. The method of any one of claims 15 to 19, wherein the step of synthesizing a complementary strand of a target nucleic acid comprising DHU residues further comprises annealing one or more forward and / or reverse primers to the sequencing library adapter.

21. The method of any one of claims 15 to 20, wherein the step of exponentially amplifying the complementary strand of the target nucleic acid comprises annealing a library amplification primer to the pre-amplified target nucleic acid.

22. The method of any one of claims 15 to 21, wherein the sequencing is performed by massively parallel sequencing.

23. The method of any one of claims 1 to 22, wherein the target nucleic acid molecule comprising DHU is produced by a process comprising contacting an oxidized nucleic acid sample comprising 5-carboxycytosine (5caC) and / or 5-formylcytosine (5fC) with a borane reducing agent.

24. The method of claim 23, wherein the borane reducing agent comprises an agent selected from the group consisting of 2-picoline borane (pic-BH3), borane, sodium borohydride, sodium cyanoborohydride, and sodium triacetoxyborohydride.

25. The method of any one of claims 23 to 24, wherein the step of contacting the oxidized nucleic acid sample comprising 5caC and / or 5fC with a borane reducing agent further comprises reacting the oxidized nucleic acid sample with the borane reducing agent in a reaction mixture comprising 45.0% to 52.5% DMSO by volume.

26. The method of any one of claims 23 to 25, wherein the step of contacting the oxidized nucleic acid sample comprising 5caC and / or 5fC with a borane reducing agent further comprises reacting the oxidized nucleic acid sample with the borane reducing agent at a temperature of 45.0°C to 52.5°C.

27. The method of any one of claims 23 to 26, wherein the step of contacting the oxidized nucleic acid sample comprising 5caC and / or 5fC with a borane reducing agent further comprises reacting the oxidized nucleic acid sample with the borane reducing agent for a period of 45 to 60 minutes.

28. A method for converting 5-carboxycytosine (5caC) and / or 5-formylcytosine (5fC) into dihydrouracil (DHU), comprising contacting a nucleic acid sample comprising 5caC and / or 5fC with a borane reducing agent in a reaction mixture comprising 45.0% to 52.5% by volume of DMSO.

29. The method of claim 28, further comprising reacting the oxidized nucleic acid sample with the borane reducing agent at a temperature of 45.0°C to 52.5°C.

30. The method of any one of claims 28 to 29, further comprising reacting the oxidized nucleic acid sample with the borane reducing agent for a period of 45 to 60 minutes.

31. The method of any one of claims 28 to 30, wherein the borane reducing agent comprises an agent selected from the group consisting of 2-picoline borane (pic-BH3), borane, sodium borohydride, sodium cyanoborohydride, and sodium triacetoxyborohydride.

32. The method of claim 31, wherein the borane reducing agent comprises sodium borohydride.

33. The method of claim 31, wherein the borane reducing agent comprises sodium cyanoborohydride.

34. The method of claim 31, wherein the borane reducing agent comprises sodium triacetoxyborohydride.

35. The method of claim 31, wherein the borane reducing agent comprises 2-methylpyridine borane.

36. The method of any one of claims 28 to 35, comprising contacting the nucleic acid sample with an oxidizing agent prior to contacting with the borane reducing agent.

37. The method of claim 36, wherein the oxidizing agent is a ten-eleven translocation (TET) enzyme.

38. The method of claim 37, wherein the TET enzyme comprises human TET1, human TET2, human TET3, mouse TET1, mouse TET2, mouse TET3, Naegleria TET (NgTET), Coprinus cinereus (CcTET), or a derivative or analog thereof.

39. The method of claim 36, wherein the oxidant comprises a chemical oxidant.

40. The method of claim 39, wherein the chemical oxidant comprises manganese oxide (MnO2), potassium ruthenate (K2RuO4), potassium perruthenate (KRuO4) or Cu(II) / TEMPO.

41. The method of any one of claims 36 to 40, further comprising adding a blocking group to one or more modified cytosines in the nucleic acid sample.

42. The method of any one of claims 28 to 41, further comprising sequencing the nucleic acid sample after contacting with the borane reducing agent to identify converted cytosine bases.

Citation Information

Patent Citations

  • Transposon end compositions and methods for modifying nucleic acids

    US20100120098A1

  • Linking sequence reads using paired code tags

    US20120208705A1

  • Linking sequence reads using paired code tags

    US20120208724A1

  • Methods and compositions for nucleic acid sequencing

    US20150368638A1

  • Bisulfite-free, base-resolution identification of cytosine modifications

    US20200370114A1