A single-base precision library construction method for whole-genome DNA cytosine hydroxymethylation modification
By improving the capture process and library preparation method, and utilizing APOBEC3A enzyme and streptavidin magnetic beads for capture, the problems of high false positive rate and high sequencing cost in DNA hydroxymethylation detection have been solved. A balance between single-base accuracy and low sequencing data volume has been achieved, making it suitable for routine double-stranded library preparation kits.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SANGON BIOTECH (SHANGHAI) CO LTD
- Filing Date
- 2023-02-13
- Publication Date
- 2026-06-02
AI Technical Summary
Existing DNA hydroxymethylation detection technologies have shortcomings in terms of single-base accuracy and data integrity, especially in terms of high false positive rate, high sequencing cost and poor versatility, which cannot simultaneously meet the requirements of high accuracy and low sequencing data volume.
By improving the capture process, employing APOBEC3A enzyme deamination treatment and streptavidin magnetic bead capture, combined with click chemistry and double-strand library preparation kits, the false positive rate was reduced and the sequencing depth and versatility were improved.
It achieves single-base precision detection, reduces the false positive rate to 1.5%, reduces sequencing costs by 80%, and is applicable to routine double-stranded library preparation kits, improving data reliability and sequencing depth.
Smart Images

Figure CN116145267B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of genomics, epigenetics, and molecular biology, and more specifically, to a method for single-base precision library construction using whole-genome DNA cytosine hydroxymethylation modification. Background Technology
[0002] Epigenetic modifications refer to changes in gene expression levels that are not caused by alterations in gene sequence, including DNA methylation, histone modifications, non-coding RNA, and genomic imprinting. These epigenetic changes play important roles in embryonic development, cell differentiation, tissue-specific protein expression, maintaining genome integrity, and chromosome stability by turning certain genes on or off. 5-hydroxymethylcytosine (5hmC) is a newly discovered epigenetic modification base, often referred to as the sixth base of human DNA, and exists at low levels in various cell types of mammals. 5hmC is the hydroxylated form of methylated cytosine (5mC), formed by the oxidation of 5mC by ten-eleven translocation (TET) family proteins.
[0003] In 1952, Wyatt et al. first discovered the epigenetic marker 5hmC in T-even lineage bacteriophages. In 1972, Penn et al. discovered 5hmC in the brain tissue of adult rats, mice, and frogs using chromatographic analysis, but it was initially thought to be due to DNA abnormalities caused by oxidation and therefore did not receive much attention. It wasn't until 2009 that multiple studies found 5hmC in mouse Purkinje cells and cerebellar granule cells, bringing it to widespread attention. In 2011, scientists discovered differences in the distribution of 5hmC among different tissues in the human body: 5hmC levels were high in the liver, kidneys, brain, and large intestine, relatively low in lung tissue, and the lowest levels were found in mammary glands, heart, and placental tissues. 5hmC plays an important role in cell differentiation and tissue development. In 2012, Freudenberg et al. found that in mouse embryonic stem cells, 5hmC levels gradually decreased with cell differentiation, leading to the loss of pluripotency in embryonic stem cells, indicating that 5hmC is closely related to embryonic stem cell differentiation and early individual development. In addition, Orr et al. found that when human brain tissue is in the developmental stage, the content of 5hmC is high in more differentiated parts such as the fetal cortex, while the content is low in the progenitor cell region around the ventricles, which suggests that 5hmC plays a specific role in the development of the central nervous system.
[0004] The level of 5hmC can serve as an indicator to distinguish tumor tissue from normal tissue, guiding tumor diagnosis and even providing epigenetic clues for tumor treatment. Yang et al. found that 5hmC levels were lower in human liver cancer, lung cancer, pancreatic cancer, breast cancer, and prostate cancer tissues than in surrounding normal tissues; Orr et al. also found lower 5hmC levels in malignant gliomas; furthermore, Lian et al.'s research showed that 5hmC was absent in melanoma tissue, thus serving as a basis for determining whether melanoma is present. Abnormal 5hmC levels can also disrupt the balance within hematopoietic cells, leading to hematopoietic cell proliferation and abnormal maturation, thereby developing into tumor cells. Pronier et al. found that the formation of various hematopoietic system tumors, such as myeloid neoplasms (MPN), myelodysplastic syndromes (MDS), and acute myeloid leukemia (AML), is due to the inhibition of TET2 expression in cells, leading to decreased 5hmC levels and interfering with myeloid cell development and differentiation, thus developing into myeloid tumor cells.
[0005] Currently, DNA hydroxymethylation detection technologies are mainly divided into two categories: precipitation detection technologies based on the antigen-antibody immunoassay principle and non-precipitation detection technologies that do not rely on the immunoassay principle. Antigen-antibody immunoassay precipitation detection technologies work by specifically treating 5hmC through chemical or enzymatic reactions, followed by precipitation techniques to specifically capture DNA fragments containing 5hmC. These methods mainly include J-binding protein (JBP) precipitation, glycosylation, periodate oxidation, and biotinylation (GLIB). Non-precipitation detection technologies mainly include mass spectrometry, quantitative hmC glucosylation assay, thin-layer chromatography, high-pressure liquid chromatography, and single-molecule real-time DNA sequencing (SMRT). In recent years, with the continuous innovation of high-throughput sequencing technologies and the continuous reduction in sequencing costs, they have been widely applied in various biological research fields.
[0006] While various high-throughput detection technologies for DNA hydroxymethylation exist, each has its advantages and disadvantages. Bisulfite-based methods, such as OxBS-seq and TAB-seq, can detect 5hmC at single-base resolution, providing comprehensive and accurate quantitative information for scientific and disease research. However, the simultaneous oxidation and bisulfite treatment leads to significant loss of genomic DNA, limiting their practicality in studies with limited sample sizes. In contrast, conversion sequencing methods using the AID / APOBEC family DNA deaminase APOBEC3A (A3A) (ACE-seq) effectively preserve DNA integrity while maintaining single-base accuracy. However, to meet analytical requirements at the whole-genome sequencing level, a sequencing depth of at least 30× (at least 90G of sequencing data for human samples) is required, undoubtedly increasing sequencing costs and the complexity of data storage and analysis. Furthermore, OxBS-seq requires the simultaneous construction of BS-seq libraries. While this combined sequencing method can achieve single-base resolution for 5hmC, it demands extremely large sample volumes and high sequencing costs, and places extremely high demands on the professional skills of both experimental operators and data analysts. Otherwise, biases are easily introduced, leading to reduced data reliability. Enrichment-based analysis methods, on the other hand, can effectively reduce sequencing costs while increasing sequencing depth. The classic hydroxymethylation capture method (5hmC-Seal) was pioneered by Professor He Chuan's team at the University of Chicago. Its process mainly includes: 1) DNA sample fragmentation, 2) end repair and addition of "A", 3) adapter ligation, 4) purification, 5) 5hmC labeling and click chemistry, 6) affinity enrichment, and 7) PCR amplification. This method, after five years of optimization, has advantages such as shorter processing time, ease of operation, lower sequencing data volume, and suitability for large-sample data analysis, but it cannot achieve single-base accuracy. In 2021, Xiaogang Li et al. applied deaminase A3A to the 5hmC-Seal method and invented a new method, DIP-CAB-Seq, which detects genome-wide DNA hydroxymethylation modifications with less data volume while maintaining single-base accuracy.
[0007] Traditional 5hmC-Seal and DIP-CAB-Seq methods are based on 5hmC capture, and the library enrichment and amplification steps still use magnetic bead amplification, resulting in a false positive rate of over 20%, which greatly affects the reliability of the data and increases the number of trial and error attempts in data result verification. At the same time, the implementation of DIP-CAB-Seq library construction requires hydroxymethylation modification of the adapter, which is not universal. In addition, the experimental results of the DIP-CAB-Seq method are unstable and have poor reproducibility.
[0008] In view of this, the present invention is proposed. Summary of the Invention
[0009] The purpose of this invention is to provide a single-base precision library construction method for whole-genome DNA cytosine hydroxymethylation modification to solve the aforementioned technical problems. This invention reduces the false positive rate to approximately 1.5% by modifying the capture process; simultaneously, through improvements to the library construction process, it can be applied to conventional double-stranded library construction kits, greatly improving versatility and facilitating increased market share; the data reproducibility rate can reach nearly 90%.
[0010] This invention is implemented as follows:
[0011] This invention provides a method for single-base precision library construction based on whole-genome DNA cytosine hydroxymethylation modification, comprising the following steps:
[0012] 5hmC labeling reaction; click chemistry reaction; sample DNA fragmentation; purification of fragmented products via magnetic beads; 5hmC fragmented DNA capture; washing of captured products; reduction reaction; APOBEC enzyme deamination reaction; purification of deamination products; double-strand conversion reaction; end repair and 3′ A addition reaction; adapter ligation; purification and double sorting of ligation products; library amplification; library purification;
[0013] 5hmC labeling reaction refers to the labeling of 5hmC on genomic DNA using a glycosylation labeling system;
[0014] The click chemistry reaction involves mixing and reacting DBCO-PEG3-SS-Biotin with the reaction product labeled with 5hmC.
[0015] DNA fragmentation of a sample involves using ultrasound to fragment the DNA in the sample.
[0016] 5hmC fragmented DNA capture uses streptavidin magnetic beads to capture fragmented DNA containing 5hmC.
[0017] The washing process for the captured products included: washing the magnetic beads three times with a buffer and then washing them one to three times with ddH2O.
[0018] The APOBEC enzyme deamination reaction deaminates unglycosylated C and 5mC to U;
[0019] The double-strand conversion reaction uses Klenow Fragment to convert the purified deamination product into double-stranded DNA.
[0020] On the one hand, the library construction method provided by this invention achieves single-base accuracy. Detection methods for assessing 5hmC levels based on antibody enrichment (hMeDIP-Seq) and glucosylation modification enrichment (hMe-Seal chemical labeling, JBP-1 precipitation, and GLIB precipitation) have certain drawbacks, namely, difficulty in accurate quantification and inability to achieve single-base resolution. However, this invention, by treating the genomic DNA sample with the APOBEC3A enzyme, effectively deaminates unglycosylated C and 5mC to U, while glycosylated 5hmC remains undeaminated and is still detected as C after sequencing, thus achieving single-base accuracy. Furthermore, the enzymatic treatment method causes less damage to the DNA sample than chemical reagent treatment, and yields a higher library yield with the same input.
[0021] Secondly, the library construction method provided by this invention can achieve high sequencing depth with a relatively low amount of sequencing data. The average level of hydroxymethylated cytosine in the whole genome is approximately 1% of cytosine. However, while detection methods based on bisulfite (TAB-Seq, OxBS-Seq) and APOBEC3A deaminase (ACE-seq) can achieve the 5hmC single-base resolution requirement for the whole genome, they require a huge amount of sequencing data and incur significant sequencing costs. This invention, by biotin-labeling glycosylated 5hmC and then using streptavidin magnetic beads for specific capture, pulls down only the DNA fragments containing 5hmC, achieving 30× sequencing data with 15Gb, reducing sequencing costs by more than 80%.
[0022] Thirdly, the database construction method provided by this invention improves data reliability:
[0023] (1) Reduce the false positive rate.
[0024] False positives have always been a common problem in capture experiments. Standard assays have shown that hMe-Seal capture false positives can reach as high as 20%. A series of exploratory experiments have confirmed that false positives mainly originate from DNA fragments without hmC sites becoming entangled on the macromolecule DBCO-PEG3-SS-Biotin, captured nucleic acid fragments, and capture magnetic beads. This invention reduces the false positive rate to approximately 1.5% through the following operational method.
[0025] To reduce the false positive rate of capture, the inventors adopted the following technical means: the DNA fragmentation process was set after the click chemical reaction (addition of DBCO-PEG3-SS-Biotin) and before the 5hmC fragmented DNA capture. Ultrasonic waves were used to oscillate and release the non-capture fragments wrapped on DBCO-PEG3-SS-Biotin, and the non-capture fragments were removed by magnetic bead purification, thereby reducing the false positive rate caused by non-capture fragments.
[0026] When using streptavidin-modified magnetic beads to capture and elute DNA fragments containing 5hmC, an additional water wash step is added after washing the magnetic beads with 1×B&W Buffer. This step aims to remove as many non-capture fragments as possible that are entangled with the DNA fragments captured by the magnetic beads, thereby further reducing the false positive rate caused by non-capture fragments.
[0027] Before library amplification, a reducing agent is used to "cleave" the disulfide bonds in the DBCO-PEG3-SS-Biotin molecule to remove the magnetic beads, further removing non-capture fragments entangled on the magnetic beads. This more thoroughly reduces the false positive rate caused by non-capture fragments.
[0028] (2) The captured DNA fragments were treated with APOBEC3A deaminase:
[0029] Treatment of fragments containing hmC sites with APOBEC3A deaminase only can improve the conversion efficiency of non-hydroxymethylated cytosine sites.
[0030] (3) The library preparation method provided by the present invention can effectively increase the sequencing depth to 60× (about 30G, which is much lower than ACE-seq), effectively reduce the insufficient data analysis coverage caused by sequencing depth, and improve the single base detection accuracy.
[0031] Fourthly, the library construction method provided by this invention is compatible with conventional double-stranded library construction kits. In recent years, detection methods for assessing 5hmC levels have failed to simultaneously meet the requirements of single-base accuracy and low sequencing data volume. Until 2021, the DIP-CAB-Seq detection method emerged, which could simultaneously meet these two indicators. However, this method requires the use of single-stranded library construction kits, which undoubtedly reduces its versatility and increases library construction costs. In contrast, this invention designs random primers to convert deaminated single-stranded DNA samples into double-stranded DNA using Klenow Fragment, followed by conventional double-stranded library construction, making it more universally applicable.
[0032] In a preferred embodiment of the present invention, the reaction system for the 5hmC labeling reaction includes: UDP-N3-Glu, genomic DNA, HEPES, MgCl2, and T4-β-glucosyltransferase (βGT). The 5hmC labeling reaction is incubated at 35-37°C for 1-2 hours. The reaction system for the 5hmC labeling reaction is used to protect 5hmC.
[0033] The reaction system for the 5hmC labeling reaction includes, for example, 34.25 μL genomic DNA, 2.5 μL HEPES, 1.25 μL MgCl2, 8 μL UDP-N3-Glu, and 4 μL βGT, and is incubated at 37°C for 2 h. The initial loading volume of genomic DNA is 500–1500 ng.
[0034] The concentrations of HEPES, MgCl2, and UDP-N3-Glu were 1M, 1M, and 1mM, respectively; the pH of HEPES was 8.0.
[0035] In a preferred embodiment of the present invention, the volume ratio of DBCO-PEG3-SS-Biotin to the 5hmC-labeled reaction product in the click chemistry reaction is 1:20-24; the click chemistry reaction is incubated at 35-37°C for 1-2 hours.
[0036] In one alternative embodiment, the concentration of DBCO-PEG3-SS-Biotin is 4-4.5 mM.
[0037] In a preferred embodiment of the present invention, the DNA fragmentation time is 15-38 min during the DNA fragmentation process.
[0038] In one optional implementation, the fragmentation time is 35-38 minutes, and the fragmented fragment ranges from 200-500 bp. Within this fragmentation time, high-quality DNA fragments of the target fragment can be obtained.
[0039] Post-fragmentation purification: The fragmented product was purified using magnetic beads, with a purification fold of 1.5×.
[0040] In a preferred embodiment of the present invention, 5hmC fragmented DNA capture includes: adding streptavidin magnetic beads that have been washed and resuspended in buffer to fragmented genomic DNA, mixing well, and incubating at room temperature for 20-30 min to capture 5hmC fragmented DNA. The amount of capture magnetic beads is 5 μL. The volume of 1×Buffer used to resuspend the magnetic beads is 25 μL.
[0041] In one alternative implementation, the buffer contains Tris-HCl, EDTA, and NaCl; in another alternative implementation, the buffer contains 10 mM Tris-HCl, 1 mM EDTA, and 2 M NaCl.
[0042] In a preferred embodiment of the present invention, the step of washing the capture product helps to wash away non-capture fragments (non-specific fragments) wrapped around the nucleic acid fragments, thereby reducing false positives.
[0043] In a preferred embodiment of the present invention, the reduction reaction includes: adding a reducing agent to resuspend the magnetic beads, and incubating by shaking at 25-27°C and 900-950 rpm; for example, the shaking time is 70 min.
[0044] In one alternative implementation, the reducing agent is a DTT solution.
[0045] In a preferred embodiment of the present invention, the APOBEC enzyme deamination reaction includes: adding formamide to the sample to be deamination treated for denaturation, and then mixing and incubating the denatured DNA with APOBEC Reaction buffer, BSA, APOBEC and ddH2O.
[0046] The denaturation conditions are 85°C for at least 10 minutes. For example, denaturation for 10-20 minutes.
[0047] In one alternative implementation, the incubation conditions are: 35-37°C, incubation for 2-3 hours.
[0048] Deamination reaction system: 20 μL denatured DNA, 10 μL 10×APOBEC Reaction buffer, 1.5 μL LBSA, 1.5 μL APOBEC and 67 μL nuclease-free water, incubated at 37°C for 3 h.
[0049] Purification and recovery of deamination product: The deamination product was purified and concentrated using column chromatography. The purification kit was the ZymoOligo Clean and Concentrator Kit. In a preferred embodiment of the invention, the reaction system for the double-strand conversion reaction included: the purified deamination product, Klenow Buffer, Random Primers, Solution I, and dNTPs. In an optional embodiment, the conditions for the double-strand conversion reaction included: incubation at 95-96°C for 5 min, followed by the addition of Klenow Fragment and incubation at 37°C for 60 min. The concentration of Klenow Fragment was 5 U / μL.
[0050] In a preferred embodiment of the present invention, the connector sequence is as shown in SEQ ID NO.1-2.
[0051] SEQ ID NO.1:
[0052] 5′-Phos-GATCGGAAGAGCACACGTCTGAACTCCAGT*C-3′
[0053] SEQ ID NO.2:
[0054] 5′-ACACTCTTTCCCTACACGACGCTCTTCCGATC*T-3′
[0055] End repair, phosphorylation, and 3′ A addition reaction include: incubation at 30°C and 72°C for 20 min each in the presence of Endprep buffer and Endprep enzyme to repair end gaps and add A to the 3′ end of the double-stranded transformed genomic DNA fragment.
[0056] Adapter ligation involves adding 30 μL of Ligation Enhancer, 5 μL of T4 DNA Ligase, and DNA Adaptors to ligate the repaired and A-tailed DNA.
[0057] The purification and double sorting of the ligation products included: purification and double sorting of the ligation products using magnetic beads; the magnetic beads used for purification were Hieff NGS DNA Selection Beads; the purification method was 0.6× magnetic bead purification, and the sorting methods were 0.7× and 0.2× magnetic bead sorting.
[0058] Library amplification includes: adding amplification enzyme, reaction buffer, and adapter primers containing the index to amplify the purified and sorted ligation products; the amplification cycle number is 11-15.
[0059] Library purification included: adding 0.9× Hieff NGS DNA Selection Beads, vortexing to mix, and incubating at room temperature for 5 min. Eluting with 20 μL of ddH2O. Before purification, add ddH2O to a final volume of 50 μL.
[0060] The present invention has the following beneficial effects:
[0061] This invention proposes a library construction method that, while ensuring data authenticity, integrates low data volume, single-base accuracy, and applicability to conventional double-stranded library construction kits. This method has the following advantages:
[0062] 1. High precision of single bases.
[0063] Currently, detection methods for assessing 5hmC levels based on antibody enrichment (hMeDIP-Seq) and glucosyl modification enrichment (hMe-Seal chemical labeling, JBP-1 precipitation, and GLIB precipitation) have certain limitations, namely, difficulty in accurate quantification and inability to achieve single-base resolution. While the ACE-seq enzyme conversion method can achieve single-base accuracy using APOBEC3A deaminase, its conversion rate is insufficient because it processes all DNA from the entire genome. This invention, however, only deaminates genomic DNA containing 5hmC using APOBEC3A, improving conversion rate while reducing DNA fragment damage caused by chemical conversion, thus obtaining a higher library yield with the same sample input.
[0064] 2. Low sequencing data volume can meet the requirements of high sequencing depth.
[0065] The average level of hydroxymethylated cytosine in the whole genome is about 1% of cytosine. While methods such as TAB-Seq and OxBS-Seq based on bisulfite, and ACE-seq based on APOBEC3A deaminase, can achieve the single-base resolution of 5hmC in the whole genome, they require a huge amount of sequencing data and incur significant sequencing costs. This invention, by biotinylating glycosylated 5hmC and using streptavidin magnetic beads for specific capture, pulls down only the DNA fragments containing 5hmC, achieving 30× of 15Gb sequencing data and reducing sequencing costs by more than 80%.
[0066] 3. Applicable to conventional double-stranded library preparation kits.
[0067] In recent years, detection methods for assessing 5hmC levels have failed to simultaneously meet the requirements of single-base accuracy and low sequencing data volume. Until 2021, the DIP-CAB-Seq detection method emerged, which could simultaneously meet these two requirements. However, this method requires the use of single-stranded library construction kits, which undoubtedly reduces its versatility and increases library construction costs. This invention, on the other hand, designs random primers to convert deaminated single-stranded DNA samples into double-stranded DNA using Klenow Fragment, followed by routine double-stranded library construction, making it more universally applicable.
[0068] 4. The database construction method provided by this invention improves data reliability:
[0069] (1) False positives have always been a common problem in capture experiments. Standard tests have shown that hMe-Seal capture false positives can reach as high as 20%. A series of exploratory experiments have confirmed that false positives mainly arise from DNA fragments without hmC sites entangled on the macromolecule DBCO-PEG3-SS-Biotin, captured nucleic acid fragments, and capture magnetic beads. This invention reduces the capture false positive rate to about 1.5% through the following operating method.
[0070] To reduce the false positive rate of capture, the inventors adopted the following technical means: the DNA fragmentation process was set after the click chemical reaction (addition of DBCO-PEG3-SS-Biotin) and before the 5hmC fragmented DNA capture. Ultrasonic waves were used to oscillate and release the non-capture fragments wrapped on DBCO-PEG3-SS-Biotin, and the non-capture fragments were removed by magnetic bead purification, thereby reducing the false positive rate caused by non-capture fragments.
[0071] When using streptavidin-modified magnetic beads to capture and elute DNA fragments containing 5hmC, an additional water wash step is added after washing the magnetic beads with 1×B&W Buffer. This step aims to remove as many non-capture fragments as possible that are entangled with the DNA fragments captured by the magnetic beads, thereby further reducing the false positive rate caused by non-capture fragments.
[0072] Before library amplification, a reducing agent is used to "cleave" the disulfide bonds in the DBCO-PEG3-SS-Biotin molecule to remove magnetic beads, further removing non-capture fragments entangled on the magnetic beads. This further reduces the false positive rate of non-capture fragments.
[0073] (2) The captured DNA fragments were treated with APOBEC3A deaminase:
[0074] Treatment of fragments containing hmC sites with APOBEC3A deaminase only can improve the conversion efficiency of non-hydroxymethylated cytosine sites.
[0075] (3) The library preparation method provided by the present invention can effectively increase the sequencing depth to 60× (about 30G, which is much lower than ACE-seq), effectively reduce the insufficient data analysis coverage caused by sequencing depth, and improve the single base detection accuracy. Attached Figure Description
[0076] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0077] Figure 1 This is an overall flowchart of the present invention;
[0078] Figure 2 This is the fragment detection result for Agilent 2100;
[0079] Figure 3 Comparison of results from different single-base detection accuracy methods;
[0080] Figure 4 Comparison of hmCG site consistency results for libraries constructed with DNA fragmentation prior to DNA fragmentation and methylation-modified adapters;
[0081] Figure 5 Statistical results of the number of mCG and hmCG sites detected by different methods;
[0082] Figure 6 This is a comparison chart of the results of the CACE-seq method and the ACE-seq detection method provided by this invention. Detailed Implementation
[0083] Reference will now be made to detailed embodiments of the present invention, one or more of which are described below. Each example is provided for explanation and not for limitation of the invention. In fact, it will be apparent to those skilled in the art that various modifications and variations can be made to the invention without departing from its scope or spirit. For example, features described or illustrated as part of one embodiment may be used in another embodiment to produce further embodiments.
[0084] Unless otherwise specified, the practice of this invention will employ conventional techniques of cell biology, molecular biology (including recombinant technologies), microbiology, biochemistry, and immunology, which are within the capabilities of those skilled in the art. This technique is well explained in the literature, such as *Molecular Cloning: A Laboratory Manual*, 2nd edition (Sambrook et al., 1989); *Oligonucleotide Synthesis* (edited by M.J. Gait, 1984); *Methods in Enzymology* (Academic Press, Inc.); *Handbook of Experimental Immunology* (edited by D.M. Weir and C.C. Blackwell); *Current Protocols in Molecular Biology* (edited by F.M. Ausubel et al., 1987); *PCR: The Polymerase Chain Reaction* (edited by Mullis et al., 1994); and *Current Protocols in Immunology* (edited by J.E. C. Olgan et al., 1991), each of which is explicitly incorporated herein by reference.
[0085] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Where specific conditions are not specified in the embodiments, conventional conditions or conditions recommended by the manufacturer shall apply. Reagents or instruments whose manufacturers are not specified are all conventional products that can be purchased commercially.
[0086] The features and performance of the present invention will be further described in detail below with reference to embodiments.
[0087] Reagent preparation:
[0088] The main components of 1×B&W Buffer include Tris-HCl (pH 7.5), EDTA, and NaCl, with concentrations of 5 mM, 0.5 mM, and 1 M, respectively.
[0089] The main components of 2×B&W Buffer include Tris-HCl (pH 7.5), EDTA and NaCl, with concentrations of 10 mM, 1 mM and 2 M, respectively.
[0090] Example 1
[0091] This embodiment provides a method for constructing a single-base precision capture library (CACE-seq) based on hydroxymethylation modification of mouse brain tissue samples. See the flowchart below. Figure 1 As shown.
[0092] Genomic DNA was extracted using a commercial tissue extraction kit (DNeasy Blood & Tissue Kits, Qiagen, 69504). Sample concentration and integrity (260 / 280 > 1.8, 260 / 230 > 1.7, main band intact) were assessed using Nanodrop, Qubit, and 1% agarose gel electrophoresis.
[0093] 1. 5hmC labeling reaction: Prepare the reaction system shown in the table below and incubate at 37℃ for 2 hours in a PCR instrument. Purify with 1.5×Hieff NGS DNA Selection beads and elute with 24 μL ddH2O.
[0094] reagents Volume (μL) DNA (500-1500ng) 34.25 1M HEPES 2.5 <![CDATA[1M MgCl2]]> 1.25 1mM UDP-N3-Glu 8 T4-βGT 4 Total volume 50
[0095] 2. Click chemical reaction: Add 1 μL of 4.5 mM DBCO-PEG3-SS-Biotin to the 5 hmC labeled reaction product and incubate at 37 °C for 2 h.
[0096] 3. Purification of biotin-containing products: Add ddH2O to the mixture after the click chemistry reaction is completed to a final volume of 50 μL, purify with 1.5× Hieff NGS DNA Selection beads, and elute with 50 μL ddH2O.
[0097] 4. DNA fragmentation: DNA was fragmented for 38 minutes using a QSONICA fragmentation instrument.
[0098] 5. Purification of fragmented products: Purify with 1.5× Hieff NGS DNA Selection beads and elute with 25 μL ddH2O.
[0099] 6. 5hmC Fragmented DNA Capture: Take 5 μL of resuspended streptavidin magnetic beads into a new PCR tube, magnetically aspirate, discard the supernatant, and wash the magnetic beads three times with 5 μL of 1×B&W Buffer. Add 25 μL of 2×B&W Buffer and 25 μL of purified DNA to the washed magnetic beads, mix well, and incubate at room temperature for 30 min.
[0100] 7. Washing the captured product: After incubation, magnetically aspirate the beads, discard the supernatant, and wash the magnetic beads three times with 50 μL 1×B&W Buffer, followed by washing the magnetic beads once with 100 μL ddH2O.
[0101] 8. DTT Reduction Reaction: Resuspend the magnetic beads in 20 μL of DTT and incubate at 25°C and 900 rpm for 70 min. Transfer the supernatant to a new PCR tube using magnetic aspirator, then add 30 μL of ddH2O to bring the volume to 50 μL. Purify using 1.5× Hieff NGS DNA Selection beads and elute with 16 μL of ddH2O.
[0102] 9. APOBEC enzyme deamination reaction: Add 4 μL of formamide to each DNA sample (16 μL) and denature at 85°C for at least 10 min; then add the reagents in the table below to the denatured sample and incubate at 37°C for 3 h.
[0103] reagents Volume (μL) Nuclease-free water 67 10×APOBEC Reaction buffer 10 BSA 1.5 APOBEC 1.5
[0104] 10. Purification of deamination products: The product from the previous reaction was purified and concentrated using the Zymo Oligo Clean and Concentrator Kit, followed by elution with 15 μL ddH2O.
[0105] 11. Double-stranded transformation reaction: Add the reagents in the table below, incubate at 95℃ for 5 min, then quickly place the PCR tube at 4℃ or on ice to cool for 5 min; then add 1 μL Klenow Fragment and react at 37℃ for 1 h.
[0106] reagents Volume (μL) DNA 12.25 10×Klenow Buffer 2.5 Random Primer 2 Solution I (100×) 0.25 10mM dNTP 2 Total volume 19
[0107] 12. Purification of the double-stranded transformation product: Add ddH2O to a final volume of 50 μL. Purify using 1.8× Hieff NGS DNA Selection beads, eluting with 25 μL ddH2O.
[0108] 13. Double-stranded DNA fragment end repair and 3′ end A addition reaction: End repair and 3′ end A addition were performed using a commercial double-stranded library construction kit (Sangon, N608380). The reaction system is shown in the table below, with reaction temperatures and times of 30℃ for 20 min and 72℃ for 20 min.
[0109] reagents Volume (μL) Converted dsDNA 25 Endprep Buffer 3 Endprep enzyme 2 Total volume 30
[0110] 14. Adapter ligation: Prepare the adapter ligation reaction system according to the table below. The reaction temperature and time are 20℃ and 30min. After the reaction, purify with 0.6×Hieff NGS DNA Selection beads and elute with 50μL ddH2O.
[0111] reagents Volume (μL) Suspend DNA 30 Fast Ligation Buffer 15 Fast Ligase 2.5 Short Adaptor 1.2 <![CDATA[ddH2O to Total volume]]> 50
[0112] The connector sequence is as follows:
[0113] SEQ ID NO.1:
[0114] 5′-Phos-GATCGGAAGAGCACACGTCTGAACTCCAGT*C-3′
[0115] SEQ ID NO.2:
[0116] 5′-ACACTCTTTCCCTACACGACGCTCTTCCGATC*T-3′
[0117] 15. Adapter product sorting: Double selection was performed using 0.7× and 0.2× Hieff NGS DNA Selection beads, and finally eluted with 13 μL ddH2O.
[0118] 16. Library Amplification: Add 15 μL of 2×Hot Start PCR mix and 2 μL of LUDI Adaptor to the sorted adapter products, mix thoroughly, and then proceed with the amplification reaction according to the following procedure:
[0119]
[0120] 17. Library purification: After PCR, add ddH2O to a final volume of 50 μL, purify with 0.9× Hieff NGS DNA Selectionbeads, and elute with 16 μL ddH2O.
[0121] 18. Library Quality Control and Sequencing: After library construction, the library concentration was detected using Qubit, and the fragment range was determined by 2% agarose gel electrophoresis or Agilent 2100 sequencing. After passing quality control, the library was sequenced using an Illumina sequencer. Quality control results are referenced... Figure 2 As shown, the results indicate that the library bands consist of a single main peak, ranging from 250 to 450 bp, and are ready for sequencing.
[0122] Comparative Example 1
[0123] This comparative example provides a method for constructing a hydroxymethylated single-base precision library (ACE-seq) of mouse brain tissue samples. The difference from Example 1 is that this comparative example first breaks down the DNA, then labels and deaminates it, and does not involve a 5hmC fragmentation DNA capture step.
[0124] Genomic DNA was extracted using a commercial tissue extraction kit (DNeasy Blood & Tissue Kits, Qiagen, 69504). The concentration and integrity of the samples were assessed using Nanodrop, Qubit, and 1% agarose gel electrophoresis (260 / 280 > 1.8, 260 / 230 > 1.7, main band intact).
[0125] 1. DNA fragmentation: DNA was fragmented for 38 minutes using a QSONICA fragmentation instrument.
[0126] 2. 5hmC labeling reaction: Prepare the reaction system shown in the table below and incubate it in a PCR instrument at 37℃ for 2 hours, then add water to a final volume of 50 μL. Purify with 1.5×Hieff NGS DNA Selection beads and elute with 16 μL ddH2O.
[0127] reagents Volume (μL) DNA (100ng) 15.6 10×Cutsmart Buffer 2.0 UDP-Glucose (2mM) 0.4 T4-βGT (10 U / μL) 2.0 Total volume 20
[0128] 3. APOBEC enzyme deamination reaction: Add 4 μL of formamide to each DNA sample (16 μL) and denature at 85°C for at least 10 min; then add the reagents in the table below to the denatured sample and incubate at 37°C for 3 h.
[0129]
[0130]
[0131] 4. Purification of deamination products: The product from the previous reaction was purified and concentrated using the Zymo Oligo Clean and Concentrator Kit, followed by elution with 15 μL ddH2O.
[0132] 5. Double-stranded transformation reaction: Add the reagents listed in the table below, incubate at 95°C for 5 min, then quickly place the PCR tube at 4°C or on ice to cool for 5 min; then add 1 μL of Klenow Fragment and react at 37°C for 1 h.
[0133] reagents Volume (μL) DNA 12.25 10×Klenow Buffer 2.5 Random Primer 2 Solution I (100×) 0.25 10mM dNTP 2 Total volume 19
[0134] 6. Purification of the double-stranded transformation product: Add ddH2O to a final volume of 50 μL. Purify using 1.8× Hieff NGS DNA Selection beads, eluting with 25 μL ddH2O.
[0135] 7. Double-stranded DNA fragment end repair and 3′ end A addition reaction: End repair and 3′ end A addition were performed using a commercial double-stranded library construction kit (Sangon, N608380). The reaction system is shown in the table below, with reaction temperatures and times of 30℃ for 20 min and 72℃ for 20 min.
[0136] reagents Volume (μL) Converted dsDNA 25 Endprep Buffer 3 Endprep enzyme 2 Total volume 30
[0137] 8. Connector connection: Prepare the connector connection for the reaction system according to the table below. The reaction temperature and time are 20℃ and 30min.
[0138] reagents Volume (μL) Suspend DNA 30 Fast Ligation Buffer 15 Fast Ligase 2.5 Short Adaptor 1.5 <![CDATA[ddH2O to Total volume]]> 50
[0139] The connector sequence is as follows:
[0140] 5′-Phos-GATCGGAAGAGCACACGTCTGAACTCCAGT*C-3′
[0141] 5′-ACACTCTTTCCCTACACGACGCTCTTCCGATC*T-3′
[0142] After the reaction, the DNA was purified using 0.6× Hieff NGS DNA Selection beads and eluted with 50 μL ddH2O.
[0143] 9. Adapter product sorting: Double selection was performed using 0.7× and 0.2× Hieff NGS DNA Selection beads, and finally eluted with 13 μL ddH2O.
[0144] 10. Library Amplification: Add 15 μL of 2×Hot Start PCR mix and 2 μL of LUDI Adaptor to the sorted adapter products, mix thoroughly, and then proceed with the amplification reaction according to the following procedure:
[0145]
[0146] 11. Library purification: After PCR, add ddH2O to a final volume of 50 μL, purify with 0.9× Hieff NGS DNA Selectionbeads, and elute with 16 μL ddH2O.
[0147] 12. Library Quality Control and Sequencing: After library construction, the library concentration was detected using Qubit, and the fragment range was detected by 2% agarose gel electrophoresis or Agilent 2100. After passing quality control, the library was sequenced using an Illumina sequencer.
[0148] Comparative Example 2
[0149] A method for constructing a single-base precision capture library of mouse brain tissue samples with DNA fragmentation first (similar to DIP-CAB-Seq). Compared to Example 1, DNA fragmentation, double-strand end repair, A addition and adapter (modified adapter) ligation are performed first, followed by labeling, click chemistry, capture and deamination reactions.
[0150] 1. Genomic DNA extraction, purification, and detection: Mouse brain tissue was extracted and purified using the DNeasy Blood & Tissue Kit (QIAGENGermantown, MD); the concentration and quality of sample DNA were detected using Qubit, Nanodrop (A260 / 280≥1.8 and 260 / 230≥1.7) and 1% agarose gel electrophoresis (single band).
[0151] 2. DNA fragmentation: DNA was fragmented for 38 minutes using a QSONICA fragmentation instrument.
[0152] 3. Double-stranded DNA fragment end repair and 3′ end A addition: Library construction was performed using the kit (NEB#E7120L). The reaction system was prepared according to the table below, and the reaction temperature and time were 20℃ for 30 min and 65℃ for 30 min.
[0153] reagents Volume (μL) Fragmented DNA 50 NEBNext Ultra II End Prep Reaction Buffer 7 NEBNext Ultra II End Prep Enzyme Mix 3 Total volume 60
[0154] 4. Connector connection: Prepare the connector connection reaction system according to the table below. The reaction temperature and time are 20℃ and 15min. After the reaction, purify with 110μL NEBNext Sample Purification Beads and elute with 35μL ddH2O.
[0155] reagents Volume (μL) End repaired / dA-tailing DNA 60 NEBNext Ultra II Ligation Master Mix 30 NEBNext Ligation Enhancer 1 NEBNext EM-seq Adaptor 2.5 Total volume 93.5
[0156] 5. 5hmC labeling reaction: Prepare the reaction system shown in the table below and incubate at 37℃ for 2 hours in a PCR instrument. Purify with 1.5×Hieff NGS DNA Selection beads and elute with 24 μL ddH2O.
[0157] reagents Volume (μL) DNA (500-1,500 ng) 34.25 1M HEPES 2.5 <![CDATA[1M MgCl2]]> 1.25 1mM UDP-N3-Glu 8 T4-βGT 4 Total volume 50
[0158] 6. Click chemical reaction: Add 1 μL of 4.5 mM DBCO-PEG3-SS-Biotin to the 5 hmC labeled reaction product and incubate at 37 °C for 2 h.
[0159] 7. Purification of biotin-containing products: Add ddH2O to the mixture after the click chemistry reaction is completed to a final volume of 50 μL, purify with 1.5× Hieff NGS DNA Selection beads, and elute with 25 μL ddH2O.
[0160] 8. 5hmC Fragmented DNA Capture: Take 5 μL of resuspended streptavidin magnetic beads into a new PCR tube, magnetically aspirate, discard the supernatant, and wash the magnetic beads three times with 5 μL of 1×B&W Buffer. Add 25 μL of 2×B&W Buffer and 25 μL of purified DNA to the washed magnetic beads, mix well, and incubate at room temperature for 30 min.
[0161] 9. Washing the captured product: After incubation, magnetically aspirate the beads, discard the supernatant, and wash the beads three times with 50 μL 1×B&W Buffer and once with 100 μL ddH2O.
[0162] 10. DTT Reduction Reaction: Resuspend the magnetic beads in 20 μL of DTT and incubate at 25°C and 900 rpm for 70 min. Transfer the supernatant to a new PCR tube using magnetic aspirator, then add 30 μL of ddH2O to bring the volume to 50 μL. Purify using 1.5× Hieff NGS DNA Selection beads and elute with 16 μL of ddH2O.
[0163] 11. APOBEC enzyme deamination reaction: Add 4 μL of formamide to each DNA sample (16 μL) and denature at 85°C for at least 10 min; then add the reagents in the table below to the denatured sample and incubate at 37°C for 3 h.
[0164] reagents Volume (μL) Nuclease-free water 67 10×APOBEC Reaction buffer 10 BSA 1.5 APOBEC 1.5
[0165] 12. Purification of deamination products: The deamination products were purified using 1×NEBNext Sample Purification Beads and eluted with 20 μL ddH2O.
[0166] 13. Library amplification: Perform the amplification reaction according to the reaction system and procedure in the table below:
[0167] reagents Volume (μL) Deaminated DNA 20 EM-seq Index Primer 5 NEBNext Q5U Master Mix 25 Total volume 50
[0168]
[0169]
[0170] 14. Library purification: The amplification product was purified with 0.9×NEBNext Sample Purification Beads and eluted with 21μLddH2O.
[0171] 15. Library Quality Control and Sequencing: After library construction, the library concentration was detected using Qubit, and the fragment range was detected by 2% agarose gel electrophoresis or Agilent 2100. After passing quality control, the library was sequenced using an Illumina sequencer.
[0172] Experimental Example 1
[0173] To verify the detection effectiveness of the method of this invention, we performed WGBS, OxBS, EM-seq, hMe-Seal, and ACE-seq tests on the same DNA sample at two different bioassay companies. Based on the sequencing results of WGBS, OxBS, and EM-seq, the ratio of mCG (mCG formed after methylation of the C nucleotide in CG dinucleotides) to hmCG (meta-hydroxymethylcytosine, a protein in the Tet enzyme family that further converts methylated cytosine) sites in the tested sample DNA was calculated to be approximately 4:1, with approximately 30,650,630 mCG sites and 7,662,657 hmCG sites, respectively. Figure 5 This is consistent with the results reported in the literature.
[0174] However, the number of hmCG sites detected by the two different biotech companies using the ACE-seq detection method was around 20,000,000, with a false positive rate of over 50%. This extremely high false positive rate, which cannot be accurately reflected by the conversion rate of the standard, may be due to low sequencing depth or APOBEC3A deaminase site bias.
[0175] Comparison with EM-seq data revealed that most false positive sites in hmCG detected by the ACE-seq method were undeamination methylated cytosine on CpG. Figure 3 (As shown in A).
[0176] The method of this invention reduces the total number of CpG sites requiring APOBEC3A deamination transformation to 51.84% by using UDP-N3-Glu glycosylation protection, biotin labeling, and streptavidin magnetic bead capture, and also increases the conversion rate from 56.82% to 90.65%. Figure 6 The capture area consistency was as high as 82.33%-92.06%. Figure 3 In the case of B), the number of captured hmCG sites is closer to the theoretical value, and the false positive rate is reduced to 15% or even lower. Figure 3 The C-means (which cannot guarantee that non-consistent sites are false positive sites) effectively reduces the trial and error time and cost of subsequent verification experiments.
[0177] The results of the hmCG site consistency comparison of libraries constructed using DNA fragmentation procedures and methylation-modified adapters are referenced. Figure 4 As shown.
[0178] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for single-base precision library construction based on whole-genome DNA cytosine hydroxymethylation modification, characterized in that, Its contents include the following steps performed in sequence: 5hmC labeling reaction; click chemistry reaction; sample DNA fragmentation; purification of fragmented products via magnetic beads; 5hmC fragmented DNA capture; washing of captured products; reduction reaction; APOBEC enzyme deamination reaction; purification of deamination products; double-strand conversion reaction; end repair and 3' end A addition reaction; adapter ligation; purification and double sorting of ligation products; library amplification; library purification; The 5hmC labeling reaction refers to the labeling of 5hmC on genomic DNA using a glycosylation labeling system; The click chemistry reaction involves mixing and reacting DBCO-PEG3-SS-Biotin with a 5hmC-labeled reaction product. The sample DNA fragmentation is performed using ultrasound to fragment the sample DNA, causing the non-captured fragments wrapped around DBCO-PEG3-SS-Biotin to vibrate and become free. The 5hmC fragmented DNA capture method uses streptavidin magnetic beads to capture fragmented DNA containing 5hmC. The washing and capturing products include: cleaning the magnetic beads 3 times with buffer and then cleaning the magnetic beads 1-3 times with ddH2O; The APOBEC enzyme deamination reaction involves deaminating unglycosylated C and 5mC to U; The double-strand conversion reaction is performed using Klenow Fragment to convert the purified deamination product into double-stranded DNA.
2. The method for single-base precision library construction based on whole-genome DNA cytosine hydroxymethylation modification according to claim 1, characterized in that, The reaction system for the 5hmC labeling reaction includes: UDP-N3-Glu, genomic DNA, HEPES, MgCl2, and T4-β-glucosyltransferase (βGT). The 5hmC labeling reaction is carried out by incubation at 35-37℃ for 1-2 h.
3. The method for single-base precision library construction based on whole-genome DNA cytosine hydroxymethylation modification according to claim 1, characterized in that, In the click chemistry reaction, the volume ratio of DBCO-PEG3-SS-Biotin to the 5hmC-labeled reaction product is 1:20-24; the click chemistry reaction is incubated at 35-37℃ for 1-2 h.
4. The method for single-base precision library construction based on whole-genome DNA cytosine hydroxymethylation modification according to claim 3, characterized in that, The concentration of DBCO-PEG3-SS-Biotin is 4-4.5 mM.
5. The method for single-base precision library construction based on whole-genome DNA cytosine hydroxymethylation modification according to claim 1, characterized in that, During the DNA fragmentation process, the DNA fragmentation time is 15-38 min.
6. The method for single-base precision library construction based on whole-genome DNA cytosine hydroxymethylation modification according to claim 5, characterized in that, The fragmentation time is 35-38 min, and the fragmented fragment range is 200-500 bp.
7. The method for single-base precision library construction based on whole-genome DNA cytosine hydroxymethylation modification according to claim 5, characterized in that, The 5hmC fragmented DNA capture process includes: adding streptavidin magnetic beads that have been washed and resuspended with buffer to fragmented genomic DNA, mixing well, and incubating at room temperature for 20-30 min to capture 5hmC fragmented DNA.
8. The method for single-base precision library construction based on whole-genome DNA cytosine hydroxymethylation modification according to claim 7, characterized in that, The buffer contains Tris-HCl, EDTA, and NaCl.
9. The method for single-base precision library construction based on whole-genome DNA cytosine hydroxymethylation modification according to claim 8, characterized in that, The buffer contains 10 mM Tris-HCl, 1 mM EDTA and 2 M NaCl.
10. The method for single-base precision library construction based on whole-genome DNA cytosine hydroxymethylation modification according to claim 8, characterized in that, The reduction reaction includes: adding a reducing agent to resuspend the magnetic beads, and incubating with shaking at 25-27°C and 900-950 rpm.
11. The method for single-base precision library construction based on whole-genome DNA cytosine hydroxymethylation modification according to claim 10, characterized in that, The reducing agent is a DTT solution.
12. The method for single-base precision library construction based on whole-genome DNA cytosine hydroxymethylation modification according to claim 1, characterized in that, The APOBEC enzyme deamination reaction includes: adding formamide to the sample to be deamination treated for denaturation, and then mixing and incubating the denatured DNA with APOBEC Reaction buffer, BSA, APOBEC and ddH2O.
13. The method for single-base precision library construction based on whole-genome DNA cytosine hydroxymethylation modification according to claim 12, characterized in that, The incubation conditions are: 35-37℃, incubation for 2-3 hours.
14. The method for single-base precision library construction based on whole-genome DNA cytosine hydroxymethylation modification according to claim 1, characterized in that, The conditions for the double-strand conversion reaction include: incubation at 95-96℃ for 5 min, followed by the addition of KlenowFragment and incubation at 37℃ for 60 min.
15. The method for single-base precision library construction based on whole-genome DNA cytosine hydroxymethylation modification according to claim 1, characterized in that, The connector sequence is shown in SEQ ID NO.1-2.