Glycosidase-assisted single base resolution sequencing method for trace modified bases in DNA
Through a glycosidase-assisted method, micro modifications in DNA are specifically cleaved and combined with linker ligation and PCR amplification technology, the problem of insufficient quantification, resolution and specificity of micro modification sequencing in DNA in the prior art is solved, and efficient and accurate single-base resolution sequencing is achieved.
Patent Information
- Application Number
- CN202510108577.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-16
AI Technical Summary
In the prior art, when performing whole-genome sequencing of trace modifications in DNA, there are problems such as poor quantitative ability, low resolution, low sensitivity, insufficient specificity, and limited localization analysis to specific modifications, making it difficult to achieve efficient and accurate single-base resolution sequencing.
Using a glycosidase-assisted method, a single-base gap at the ends of 5’-PO4 and 3’-OH were formed by specifically cleaving traces of uracil, 8-oxyguanine and AP sites, and a single-base resolution sequencing of trace-modified bases in DNA was achieved.
The enrichment and sequencing of micro-modified DNA with high sensitivity, high specificity and high resolution is achieved, which simplifies the operation process, improves research efficiency, and provides a powerful tool for disease mechanism research.
Smart Images

Figure CN120005984A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the field of biotechnology, and in particular to a glycosidase-assisted single-base resolution sequencing method for trace modified bases in DNA. Background Art
[0002] In addition to the four bases A, G, C, and T, DNA also contains a large number of unconventional nucleotides containing covalent modifications, namely modified nucleotides. The presence of modified nucleotides increases the ability of DNA to encode genetic information. DNA modification There are a variety of low-content trace modifications in nucleic acids, including uracil (Uracil, U), AP sites (apyrimidinic / apurinic sites, AP), 8-oxoguanine (8-Oxo-7,8-dihydro-2'-deoxyguanosine, 8OG) and hypoxanthine. These trace modifications can be produced by factors such as spontaneous hydrolysis of nucleic acids, cellular oxidative stress and exposure to environmental pollutants, or by enzyme catalysis such as deamination reactions, base excision reactions and nucleotide misincorporation in cells. Studies have shown that trace modifications increase the risk of DNA breaks and pose a major threat to the stability of the genome. For example, DNA breaks will occur during the repair of trace modifications such as AP sites and uracil, which will hinder the progress of DNA replication enzymes and affect the stability and integrity of the genome. The study also found that trace modifications actively participate in a variety of biological processes, including the occurrence and development of breast cancer, lung cancer, megaloblastic anemia, etc. Abnormal changes in trace modifications are an important cause of the occurrence of a variety of diseases.
[0003] In order to deeply study the mechanism of action of trace modifications in various biological processes, it is necessary to perform efficient and accurate single-base resolution sequencing of trace modifications on genes. At present, many whole-genome sequencing methods for trace modifications have been developed, such as chromatin immunoprecipitation or chemical labeling methods to enrich DNA containing trace modifications and combine with next-generation sequencing technology (NGS) to achieve the localization analysis of trace modifications in genomic DNA. However, antibody enrichment or chemical molecule enrichment methods have many defects, including poor quantitative ability, low resolution, low sensitivity, insufficient specificity, and localization analysis is limited to specific modifications. At present, some glycosidase-assisted breakpoint sequencing methods have also been reported, such as Ucaps-seq for detecting uracil modifications and CLAPS-seq for detecting 8OG. Although these methods all achieve single-base resolution sequencing of trace modifications, they all have certain defects and shortcomings. Since trace modifications are present in low concentrations in DNA, these methods still require enrichment of DNA containing trace modifications through antibody enrichment and other means. However, the enrichment methods for different trace modifications vary greatly, and it is necessary to develop targeted enrichment strategies for single modifications, which is difficult to develop. Therefore, it is necessary to develop a trace modification enrichment and post-sequencing analysis method that is simple to operate, highly accurate, and highly universal.
[0004] Since glycosidase UdgX-H109S can specifically recognize and cut uracil, glycosidase Pab-AGOG can specifically recognize and cut 8OG, and endonuclease APE1 can specifically recognize and cut AP sites, and single-strand breaks can be formed after cutting, and the 5' end of the breakpoint position is a phosphate. Based on this characteristic, the present application provides a glycosidase-assisted single-base resolution sequencing method for trace modified bases in DNA. Summary of the invention
[0005] In view of the above deficiencies in the prior art, the present invention provides a glycosidase-assisted sequencing method for site-specific micro-modified bases in DNA that achieves single-base resolution, high sensitivity, high specificity, and simple operation.
[0006] The principle of the present invention is shown in Figure 1In this method, firstly, the fragmented DNA is end-repaired and connected with adapters, so that both ends of the DNA double-strand are connected to adapter 1, and the ends of the connected product contain only hydroxyl groups; the trace modification on the DNA is specifically cut by glycosidase or APE1 enzyme (glycosidase specifically cuts uracil or 8OG; APE1 enzyme specifically cuts endogenous AP site), forming a single base gap between the 5'-PO4 and 3'-OH ends, and the DNA without trace modification will not be cut; then, extension is performed using adapter 1 as a template, and the extension of the DNA containing trace modification is terminated at the position before the trace modification and the 5' end of the breakpoint is phosphorus The truncated DNA containing the acid radical but not the DNA containing the trace modification is extended to obtain the full-length DNA with a hydroxyl group at the 5' end; under the action of DNA ligase, the extension product of the DNA containing only the trace modification can be connected to the 3' hydroxyl group of the second connector P5 with a hydroxyl group at the end; PCR amplification is performed using primers corresponding to the P5 connector, at which time only the extension product of the DNA containing the trace modification can be amplified, that is, the goal of specifically amplifying the DNA containing the trace modification is achieved, and the position connected to P5 in the amplified product is the previous position of the modified nucleoside; finally, the sequencing of the trace modification with single-base resolution is achieved through sequencing technology (for example: NGS). As long as the modification site forms a single-strand break at the 5' phosphate end under the treatment of the biological enzyme, the enrichment of the trace modified DNA can be achieved using the method of the present invention. This method has high universality and high applicability, and has the potential to become a powerful tool for sequencing trace modifications.
[0007] The principle of single-base resolution sequencing micro-modification is as follows: the connection position between the hydroxyl-terminated adapter P5 (the 3' and 5' ends of the adapter P5 are hydroxyl groups) and DNA is the position before the micro-modification. When processing actual samples, the DNA is first fragmented, and after end repair and adapter connection treatment, the ends of the genomic DNA are all free hydroxyl groups. Then, the control group uses DNA polymerase to extend the micro-modified uncut DNA sample. Since the control group cannot be connected to the adapter P5, no PCR product is generated in the control group at this time, and the original DNA sequence cannot be displayed (if the method of the present invention needs to be verified, the adapter P5-1 with a phosphate group at the 5' end can be used to connect the control group DNA, and the sequence of the DNA at a specific site can be obtained by PCR amplification and sequencing); the experimental group uses DNA polymerase to extend the micro-modified cut DNA sample, and connects the hydroxyl-terminated adapter P5 (the 5' end of the DNA sample is connected to the 3' end of the adapter P5), and the site information of the micro-modification is obtained by PCR and sequencing. Based on this, during the experiment, the sample was divided into two equal parts, one for the micro-modification to be cut (using glycosidase to specifically cut uracil or 8OG, or using APE1 enzyme to specifically cut the endogenous AP site), and the other was not cut. DNA polymerase was then used to extend the extension product at the same time, and the extension product was amplified by PCR and sequenced to obtain the corresponding DNA sequence. The difference between the linker P5 at the hydroxyl end and the genomic DNA connection sequence determined the location of the micro-modification.
[0008] To achieve the above purpose, the specific technical solutions of the present invention are as follows:
[0009] In a first aspect, the present invention provides a glycosidase-assisted single-base resolution sequencing method for trace modified bases in DNA, comprising the following steps:
[0010] (1) Fragment the DNA, then perform end repair and connect with adapter 1 to make the DNA ends free hydroxyl groups;
[0011] (2) dividing the DNA sample treated in step (1) into two equal parts, and treating at least one of the DNA samples with glycosidase and / or APE1 enzyme;
[0012] (3) Add the same DNA probe to the two DNA samples respectively, and use DNA polymerase to perform an extension reaction; a portion of the DNA probe used is completely complementary to adapter 1, and the other portion is not complementary to adapter 1 and is used for subsequent PCR amplification;
[0013] (4) Recover the extension product and connect the linker P5 with hydroxyl groups at the 3' and 5' ends, and analyze the trace modification based on PCR amplification and sequencing positioning. The extension of DNA containing trace modifications results in a truncated DNA that terminates at the position before the trace modification and has a phosphate group at the 5' end, while the extension of DNA without trace modifications results in a full-length DNA with a hydroxyl group at the 5' end; under the action of DNA ligase, only the extension product of the DNA where the trace modification is located can be connected to the linker P5 with a hydroxyl group at the end; PCR amplification is performed using primers corresponding to the P5 linker, at which time only the extension product of the DNA where the trace modification is located can be amplified, that is, the goal of specifically amplifying DNA containing trace modifications is achieved, and the position in the amplified product that is connected to P5 is the position before the modified nucleoside; finally, single-base resolution sequencing of the trace modification is achieved through sequencing.
[0014] Furthermore, in step (2), the DNA sample is treated, and a corresponding treatment method is selected according to the trace modified base to be sequenced; when the trace modification to be sequenced is uracil or 8OG, step (2) is specifically as follows: the DNA sample treated in step (1) is equally divided into two parts, one part is not treated with glycosidase (adding the corresponding protein buffer without adding protein), and the other part is treated with glycosidase, and the two DNA samples are then treated with APE1 enzyme; when the trace modification to be sequenced is an AP site (endogenous AP site), step (2) is specifically as follows: the DNA sample treated in step (1) is equally divided into two parts, one part is not treated with APE1 enzyme (adding the corresponding protein buffer without adding protein), and the other part is treated with APE1 enzyme.
[0015] When the trace modification to be sequenced is uracil or 8OG, in step (2), if no glycosidase is added and the trace modification in the DNA is not cut, then both the DNA containing the trace modification and the DNA without the trace modification are extended to produce full-length DNA with a hydroxyl group at the 5' end, and no connection product and PCR product are generated at this time; if active glycosidase is added, the trace modification on the DNA is cut to form a single base gap with a phosphate group at the 5' end, then the DNA containing the trace modification is extended to produce a truncated DNA that terminates at the position before the trace modification and has a phosphate group at the 5' end, and the DNA without the trace modification is not cut and is extended to produce a full-length DNA with a hydroxyl group at the 5' end. At this time, only the truncated DNA can be connected to the adapter with a hydroxyl group at the end, so the PCR product only includes the truncated DNA containing the trace modification, and the position connected to the adapter P5 is the position of the trace modification. The site information of the subsequent trace modification at a specific site: the connection position of the adapter P5 and the DNA is the position before the trace modification. When the trace modification to be sequenced is an AP site, in step (2), the principle of using APE1 enzyme to treat one of the DNA samples is the same as above.
[0016] Furthermore, when the trace modification to be sequenced is uracil or 8OG, the conditions for treating the DNA sample with the glycosidase are the appropriate conditions corresponding to the glycosidase; when the trace modification to be sequenced is uracil, the conditions are as follows: UdgX-H109S glycosidase is reacted with the DNA sample in UdgX-H109S buffer (50 mM Tris-HCl pH 8.0, 1 mM Na2EDTA, 1 mM DTT, 25 μg / mL BSA) at 37°C for 0.3-12 h; when the trace modification to be sequenced is 8OG, the conditions are as follows: Pab-AGOG glycosidase is reacted with the DNA sample in glycosidase buffer (20 mM Tris-HCl, 5 mM DTT, 8%glycerol, pH 8.0) at 37°C for 5 min-12 h.
[0017] Furthermore, when the trace modification to be sequenced is uracil or 8OG, after the glycosidase treatment, the conditions for treating the DNA sample with the APE1 enzyme are: adding the APE1 enzyme and the APE1 buffer to the reaction system, and reacting at 37° C. for 1-16 h.
[0018] Furthermore, when the trace modification to be sequenced is uracil or 8OG, the specific steps of using the glycosidase, and after the glycosidase treatment, using the APE1 enzyme to treat the DNA sample are as follows:
[0019] 1) When the trace modification to be sequenced is uracil, the following reaction system (20 μL) is used: 50 mM Tris-HCl pH 8.0, 1 mM Na2EDTA, 1 mM DTT, 25 μg / mL BSA, 8 pmol UdgX-H109S, 1-10 μg DNA sample, add water to make up to 10 μL; the reaction system is placed at 37 ℃ for 0.3-12 h; when the trace modification to be sequenced is 8OG, the following reaction system (20 μL) is used: 20 mM Tris-HCl, 5 mM DTT, 8% glycerol, pH 8.0, 0.2 μMPab-AGOG, 1-10 μg DNA sample, add water to make up to 20 μL; the reaction system is placed at 37 ℃ for 5 min-12 h.
[0020] 2) After the reaction in step 1) is completed, add 1 μL 10 U / μL APE1 enzyme and 2 μL 10 × APE1 buffer, and place the reaction system at 37 °C for 1-16 h.
[0021] Furthermore, when the trace modification to be sequenced is an AP site, the conditions for treating the DNA sample with the APE1 enzyme are as follows: reacting the APE1 enzyme with the DNA sample in an APE1 enzyme buffer (50 mM Potassium Acetate, 20 mM Tris-acetate, 10 mM Magnesium Acetate, 1 mM DTT, pH 7.9) at 37°C for 1-12 h.
[0022] Furthermore, when the trace modification to be sequenced is an AP site, the reaction system for treating the DNA sample with the APE1 enzyme is as follows (20 μL): 50 mM Potassium Acetate, 20 mM Tris-acetate, 10 mM Magnesium Acetate, 1 mM DTT, pH 7.9, 1 U APE1, 1-10 μg DNA sample, add water to make up to 20 μL; the reaction system is placed at 37°C for 1-12 h.
[0023] Furthermore, the DNA polymerase in step (3) is Taq DNA polymerase.
[0024] In a second aspect, the present invention provides a kit for localizing and detecting trace modifications in DNA, comprising glycosidase, APE1 enzyme, and DNA polymerase.
[0025] Furthermore, the kit also contains a reaction buffer of the corresponding protein or enzyme.
[0026] In a third aspect, the present invention provides the use of glycosidase or APE1 enzyme in the localization and detection of trace modifications of DNA.
[0027] In a fourth aspect, the present invention provides the use of glycosidase or APE1 enzyme in preparing a kit for locating and detecting trace modifications in DNA.
[0028] Compared with the prior art, the present invention is beneficial in that:
[0029] (1) Simple and easy to implement: The method provided by the present invention is simple in design, easy to understand and execute, and the entire experimental process is low-cost and has low instrument requirements, which greatly facilitates its promotion and application in various laboratory environments.
[0030] (2) Precise positioning: The method provided by the present invention can accurately sequence trace modifications in DNA, providing reliable data support for research.
[0031] (3) Rapid evaluation: The method provided by the present invention does not require complicated optimization steps. Through qPCR amplification, trace modifications at different sites on the genome can be quickly located, thereby improving research efficiency.
[0032] (4) Disease research tool: The method provided by the present invention can be used to analyze changes in key trace modification sites in biological samples, providing a powerful tool for in-depth exploration of the role of trace modifications in various diseases, and has important scientific and clinical value.
[0033] (5) The method provided by the present invention has high sensitivity, good specificity, simple operation and high efficiency. It can not only be used for single-base resolution sequencing of trace modifications in low-abundance biological samples, but also provides a new method for functional research of trace modifications in organisms. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a schematic diagram of the principle of glycosidase-assisted single-base resolution sequencing of trace modifications in DNA in the present invention;
[0035] Figure 2 is a graph showing the results of glycosidase activity detection in the present invention; wherein, Figure 2 A is a denaturing PAGE image showing that dsU:A and dsU:G are not cleaved when UdgX-H109S is not added, and that UdgX-H109S specifically cleaves dsU:A and dsU:G after adding UdgX-H109S; Figure 2 B is a denaturing PAGE image showing that ds8OG:A and ds8OG:C are not cleaved when Pab-AGOG is not added, and that Pab-AGOG specifically cleaves ds8OG:A and ds8OG:C after Pab-AGOG is added; Figure 2 C is a denaturing PAGE image showing that dsAP:A, dsAP:G, dsAP:C and dsAP:T are not cleaved when APE1 is not added, and APE1 specifically cleaves dsAP:A, dsAP:G, dsAP:C and dsAP:T after APE1 is added;
[0036] Figure 3 The sequencing result of the DNA containing trace modifications in the mixed chain after the DNA containing trace modifications and the DNA not containing trace modifications in the present invention are mixed in a ratio of 1:9; wherein, Figure 3 A is the agarose gel results of the UdgX-H109S enzyme-treated group and the untreated group after the uracil-modified and non-uracil-modified DNAs were mixed at a ratio of 1:9, and the corresponding Sanger sequencing results; Figure 3 B is the agarose gel results of the APE1 enzyme-treated group and the untreated group after the AP site-containing and AP site-free DNA were mixed at a ratio of 1:9, and the corresponding Sanger sequencing results; Figure 3 C is the agarose gel results of the Pab-AGOG enzyme treated group and the untreated group after the 8OG modified and non-8OG modified DNAs were mixed at a ratio of 1:9, and the corresponding Sanger sequencing results;
[0037] Figure 4 The enrichment of 329-bp dsDNA-U / A in the experimental group after mixing 329-bp dsDNA-U / A and 328-bp dsDNA-T / A at a ratio of 1:9;
[0038] Figure 5 This is the analysis result of the present invention using Sanger sequencing results to locate and detect uracil at Chr9:1061185001 and Chr1:226143862 sites in HEK293T cells. DETAILED DESCRIPTION
[0039] The technical solution of the present invention will be described clearly and completely below. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0040] The present invention provides a glycosidase-assisted single-base resolution sequencing method for trace modified bases in DNA (the principle is shown in Figure 1 ), which includes the following steps:
[0041] (1) Fragment the DNA, then perform end repair and connect with adapter 1 to make the DNA ends free hydroxyl groups;
[0042] (2) dividing the DNA sample treated in step (1) into two equal parts, and treating at least one of the DNA samples with glycosidase and / or APE1 enzyme;
[0043] (3) Adding the same DNA probe to the two DNA samples respectively, and performing an extension reaction with a DNA polymerase; a portion of the DNA probe used is completely complementary to adapter 1, and the other portion is not complementary to adapter 1 and is used for subsequent PCR amplification;
[0044] (4) Recover the extended product and connect it to the linker P5 with a free hydroxyl group at the end, and analyze the trace modification based on PCR amplification and Sanger sequencing.
[0045] In step (2), the DNA sample is treated, and a corresponding treatment method is selected according to the trace modified base to be sequenced; when the trace modification to be sequenced is uracil or 8OG, step (2) is specifically as follows: the DNA sample treated in step (1) is equally divided into two parts, one part is not treated with glycosidase (adding the corresponding protein buffer without adding protein), and the other part is treated with glycosidase, and the two DNA samples are then treated with APE1 enzyme; when the trace modification to be sequenced is an AP site, step (2) is specifically as follows: the DNA sample treated in step (1) is equally divided into two parts, one part is not treated with APE1 enzyme (adding the corresponding protein buffer without adding protein), and the other part is treated with APE1 enzyme.
[0046] Example 1
[0047] Detection of glycosidase activity in cleaving double-stranded DNA containing trace amounts of modification
[0048] Prepare the annealing double-stranded DNA reaction system (10 μL): 1 μL 100 mM Tris-HCl (pH 7.0), 0.5 μL 1 M NaCl, 4 μL 10 μM single-stranded DNA containing trace modifications, 4.5 μL 10 μM complementary single-stranded DNA without trace modifications, place the reaction system at 95 ℃ for 5 min, and then cool to 25 ℃ at a rate of 1 ℃ / 1 s to form double-stranded DNA.
[0049] 1) Detection of the cleavage activity of UdgX-H109S on DNA containing uracil modifications:
[0050] The reaction system was 20 μL: 50 mM Tris-HCl pH 8.0, 1 mM Na2EDTA, 1 mM DTT, 25 μg / mL BSA, 8 pmol UdgX-H109S, 2 pmol DNA sample to be tested, and water was added to make the total volume 20 μL; the reaction system was placed at 37°C for 20 min.
[0051] Annealed DNA sample to be tested:
[0052] dsU:A:5'-ATTCTCCTGCCTCAGCCTCC U GAGTAGCTGGGACTACAGGC-3' (there is a trace amount of modified base U between the 20th C base and the 21st G base in the sequence shown in SEQ ID NO.1); 3'-TAAGAGGACGGAGTCGGAGG A CTCATCGACCCTGATGTCCG-5' (SEQ ID NO. 2);
[0053] dsU:G:5'-ATTCTCCTGCCTCAGCCTCC U GAGTAGCTGGGACTACAGGC-3' (there is a trace amount of modified base U between the 20th C base and the 21st G base in the sequence shown in SEQ ID NO.3); 3'-TAAGAGGACGGAGTCGGAGG G CTCATCGACCCTGATGTCCG-5' (SEQ ID NO.4);
[0054] Denaturing PAGE results showed that 8 pmol of glycosidase UdgX-H109S could completely remove 2 pmol of dsU:A / dsU:G substrate.
[0055] 2) Detection of the cleavage activity of Pab-AGOG on 8OG-modified DNA:
[0056] The reaction system was 20 μL: 20 mM Tris-HCl, 5 mM DTT, 8% glycerol, pH 8.0, 0.2 μM Pab-AGOG, 2 pmol of the DNA sample to be tested, and water was added to make the total volume 20 μL; the reaction system was placed at 37 °C for 5 min.
[0057] Annealed DNA sample to be tested:
[0058] ds8OG:A:5'-ACAAC TACAT GTGTA ACAGT TCCTG CAT 8OG G GCGGC ATGAA CCGGAGGCCC ATCCT CA -3' (the 8-position of the G base at position 29 in the sequence shown in SEQ ID NO.5 is oxidized); 3'-TGTTG ATGTA CACAT TGTCA AGGAC GTA A C CGCCG TACTT GGCCT CCGGG TAGGA GT -5' (SEQID NO.6);
[0059] ds8OG:C:5'-ACAAC TACAT GTGTA ACAGT TCCTG CAT 8OG G GCGGC ATGAA CCGGAGGCCC ATCCT CA -3' (the 8-position of the G base at position 29 in the sequence shown in SEQ ID NO.7 is oxidized); 3'-TGTTG ATGTA CACAT TGTCA AGGAC GTA CC CGCCG TACTT GGCCT CCGGG TAGGA GT -5' (SEQ ID NO. 8;
[0060] Denaturing PAGE results showed that 0.2 μM glycosidase Pab-AGOG could completely remove 2 pmol of ds8OG:A / ds8OG:C substrate.
[0061] 3) Detection of APE1 cleavage activity on DNA containing AP sites:
[0062] The reaction system was 20 μL: 50 mM Potassium Acetate, 20 mM Tris-acetate, 10 mM Magnesium Acetate, 1 mM DTT, pH 7.9, 1 U APE1, 2 pmol of the DNA sample to be tested, and water was added to make the total volume 20 μL; the reaction system was placed at 37 °C for 60 min.
[0063] Annealed DNA sample to be tested:
[0064] dsAP:A:5'-ATTCT CCTGC CTCAG CCTCC AP GAGT AGCTG GGACT ACAGG C -3' (there is an AP site between the 20th C base and the 21st G base in the sequence shown in SEQ ID NO.9); 3'-TAAGAGGACG GAGTC GGAGG A CTCA TCGAC CCTGA TGTCC G -5' (SEQ ID NO. 10);
[0065] dsAP:G:5'-ATTCT CCTGC CTCAG CCTCC AP GAGT AGCTG GGACT ACAGG C -3' (there is an AP site between the 20th C base and the 21st G base in the sequence shown in SEQ ID NO.11); 3'-TAAGAGGACG GAGTC GGAGG G CTCA TCGAC CCTGA TGTCC G -5' (SEQ ID NO. 12);
[0066] dsAP:C:5'-ATTCT CCTGC CTCAG CCTCC APGAGT AGCTG GGACT ACAGG C -3' (there is an AP site between the 20th C base and the 21st G base in the sequence shown in SEQ ID NO.13); 3'-TAAGAGGACG GAGTC GGAGG C CTCA TCGAC CCTGA TGTCC G -5' (SEQ ID NO. 14);
[0067] dsAP: T:5'-ATTCT CCTGC CTCAG CCTCC AP GAGT AGCTG GGACT ACAGG C -3' (there is an AP site between the 20th C base and the 21st G base in the sequence shown in SEQ ID NO.15); 3'-TAAGAGGACG GAGTC GGAGG T CTCA TCGAC CCTGA TGTCC G -5' (SEQ ID NO. 16);
[0068] Denaturing PAGE results showed that 1 U APE1 could completely remove 2 pmol of dsAP:A / dsAP:G / dsAP:C / dsAP:T substrate;
[0069] Example 2
[0070] Glycosidase-assisted selective enrichment of DNA containing trace modifications and sequencing to detect the location of trace modifications
[0071] 1) End repair and ligation with adapter 1 with restriction endonuclease site (Hieff NGS® Fast-Pace End Repair / dA-Tailing Module (Cat. No. 12608ES24)), then add corresponding glycosidase to the 1 ug 235bp DNA mixture with and without micro-modification obtained by PCR amplification (the ratio of DNA with and without micro-modification is 1:9), and react at 37°C.
[0072] If detecting DNA containing uracil modification, use 235bp containing uracil modification 235bp-dsU:A: 5'-CGAGG CGTGA GGTCA CTTCA TCTGC GTCGA CGTGA GAAGG CGGTA ATACG GTTA UCCACA GAATCAGGGG ATAAC GCAGG AAAGA ACATG TGAGC AAAAG GCCAG CAAAA GGCCA GGAAC CGTAA AAAGGCCGCG TTGCT GGCGT TTTTC CATAG GCTCC GCCCC CCTGA CGAGC ATCAC AAAAA TCGAC GCTCAAGTCA GAGGT GGCGA AACCC GACAG GACTA TAAAG GCTCC -3' (there is a trace modified base U between the 54th base A and the 55th base C in the sequence shown in SEQ ID NO.17); 3'- GCTCC GCACT CCAGTGAAGT AGACG CAGCT GCACT CTTCC GCCAT TATGC CAAT A GGTGT CTTAG TCCCC TATTG CGTCCTTTCT TGTAC ACTCG TTTTC CGGTC GTTTT CCGGT CCTTG GCATT TTTCC GGCGC AACGA CCGCAAAAAG GTATC CGAGG CGGGG GGACT GCTCG TAGTG TTTTT AGCTG CGAGT TCAGT CTCCA CCGCTTTGGG CTGTC CTGAT ATTTC CGAGG -5' (SEQ ID NO.18) and 235 bp DNA without uracil modification 235 bp-dsT:A: 5'- CGAGG CGTGA GGTCA CTTCA TCTGC GTCGA CGTGA GAAGG CGGTAATACG GTTA T CCACA GAATC AGGGG ATAAC GCAGG AAAGA ACATG TGAGC AAAAG GCCAG CAAAAGGCCA GGAAC CGTAA AAAGG CCCGG TTGCT GGCGT TTTTC CATAG GCTCC GCCCC CCTGA CGAGCATCAC AAAAA TCGAC GCTCA AGTCA GAGGT GGCGA AACCC GACAG GACTA TAAAG GCTCC -3' (SEQ ID NO.19); 3'- GCTCC GCACT CCAGT GAAGT AGACG CAGCT GCACT CTTCC GCCATTATGC CAAT AGGTGT CTTAG TCCCC TATTG CGTCC TTTCT TGTAC ACTCG TTTTC CGGTC GTTTTCCGGT CCTTG GCATT TTTCC GGCGC AACGA CCGCA AAAAG GTATC CGAGG CGGGG GGACT GCTCGTAGTG TTTTT AGCTG CGAGT TCAGT CTCCA CCGCT TTGGG CTGTC CTGAT ATTTC CGAGG -5' (SEQ ID NO.20). At this time, react at 37°C for 20 minutes.
[0073] If detecting DNA containing 8OG modification, use 235bp containing 8OG modification 235bp-ds8OG:C: 5'- CGAGGCGTGA GGTCA CTTCA TCTGC GTCGA CGTGA GAAGG CGGTA ATACG GTTA 8OG CCACA GAATCAGGGG ATAAC GCAGG AAAGA ACATG TGAGC AAAAG GCCAG CAAAA GGCCA GGAAC CGTAA AAAGGCCGCG TTGCT GGCGT TTTTC CATAG GCTCC GCCCC CCTGA CGAGC ATCAC AAAAA TCGAC GCTCAAGTCA GAGGT GGCGA AACCC GACAG GACTA TAAAG GCTCC -3' (the 8-position of the G base at position 55 in the sequence shown in SEQ ID NO.21 is oxidized); 3'- GCTCC GCACT CCAGT GAAGT AGACG CAGCT GCACTCTTCC GCCAT TATGC CAAT CGGTGT CTTAG TCCCC TATTG CGTCC TTTCT TGTAC ACTCG TTTTCCGGTC GTTTT CCGGT CCTTG GCATT TTTCC GGCGC AACGA CCGCA AAAAG GTATC CGAGG CGGGGGGACT GCTCG TAGTG TTTTT AGCTG CGAGT TCAGT CTCCA CCGCT TTGGG CTGTC CTGAT ATTTCCGAGG -5’ (SEQ ID NO.22) and 235bp DNA without 8OG modification 235bp-dsG:C: 5’- CGAGG CGTGAGGTCA CTTCA TCTGC GTCGA CGTGA GAAGG CGGTA ATACG GTTA G CCACA GAATC AGGGG ATAACGCAGG AAAGA ACATG TGAGC AAAAG GCCAG CAAAA GGCCA GGAAC CGTAA AAAGG CCGCG TTGCTGGCGT TTTTC CATAG GCTCC GCCCC CCTGA CGAGC ATCAC AAAAA TCGAC GCTCA AGTCA GAGGTGGCGA AACCC GACAG GACTA TAAAG GCTCC -3’ (SEQ ID NO.23); 3’- GCTCC GCACT CCAGTGAAGT AGACG CAGCT GCACT CTTCC GCCAT TATGC CAAT C GGTGT CTTAG TCCCC TATTG CGTCCTTTCT TGTAC ACTCG TTTTC CGGTC GTTTT CCGGT CCTTG GCATT TTTCC GGCGC AACGA CCGCAAAAAG GTATC CGAGG CGGGG GGACT GCTCG TAGTG TTTTT AGCTG CGAGT TCAGT CTCCA CCGCTTTGGG CTGTC CTGAT ATTTC CGAGG -5’ (SEQ ID NO.24). At this time, react at 37 °C for 5 min.
[0074] If detecting DNA containing AP site, use 235bp-AP:A containing AP site: 5'- CGAGG CGTGAGGTCA CTTCA TCTGC GTCGA CGTGA GAAGG CGGTA ATACG GTTA AP CCACA GAATC AGGGG ATAACGCAGG AAAGA ACATG TGAGC AAAAG GCCAG CAAAA GGCCA GGAAC CGTAA AAAGG CCGCG TTGCTGGCGT TTTTC CATAG GCTCC GCCCC CCTGA CGAGC ATCAC AAAAA TCGAC GCTCA AGTCA GAGGTGGCGA AACCC GACAG GACTA TAAAG GCTCC -3' (there is an AP site between the 54th A base and the 55th C base in the sequence shown in SEQ ID NO.25); 3'- GCTCC GCACT CCAGT GAAGT AGACG CAGCTGCACT CTTCC GCCAT TATGC CAAT A GGTGT CTTAG TCCCC TATTG CGTCC TTTCT TGTAC ACCGTTTC CGGTC GTTTT CCGGT CCTTG GCATT TTTCC GGCGC AACGA CCGCA AAAAG GTATC CGAGGCGGGG GGACT GCTCG TAGTG TTTTT AGCTG CGAGT TCAGT CTCCA CCGCT TTGGG CTGTC CTGATATTTC CGAGG -5' (SEQ ID NO.26) and 235bp-dsT:A without AP site: 5'- CGAGGCGTGA GGTCA CTTCA TCTGC GTCGA CGTGA GAAGG CGGTA ATACG GTTA TCCACA GAATC AGGGGATAAC GCAGG AAAGA ACATG TGAGC AAAAG GCCAG CAAAA GGCCA GGAAC CGTAA AAAGG CCGCGTTGCT GGCGT TTTTC CATAG GCTCC GCCCC CCTGA CGAGC ATCAC AAAAA TCGAC GCTCA AGTCAGAGGT GGCGA AACCC GACAG GACTA TAAAG GCTCC -3' (SEQ ID NO.27); 3'- GCTCC GCACTCCAGT GAAGT AGACG CAGCT GCACT CTTCC GCCAT TATGC CAAT A GGTGT CTTAG TCCCC TATTGCGTCC TTTCT TGTAC ACTCG TTTTC CGGTC GTTTT CCGGT CCTTG GCATT TTTCC GGCGC AACGACCGCA AAAAG GTATC CGAGG CGGGG GGACT GCTCG TAGTG TTTTT AGCTG CGAGT TCAGT CTCCACCGCT TTGGG CTGTC CTGAT ATTTC CGAGG -5' (SEQ ID NO.28). Incubate at 37°C for 60 min. Detect DNA containing endogenous AP sites. After this step is completed, proceed directly to step 3).
[0075] 2) APE1 cleavage of AP sites: Add 1 μL 10 U / μL APE1 enzyme and 2 μL 10 × APE1 buffer to the reaction product treated with glycosidase, place the reaction system at 37°C, and react for 60 min to cleave the generated AP sites.
[0076] 3) Taq DNA polymerase extension: After the above reaction, 0.4 μM DNA probe (5'- CACGT AGCAC TTCAT GCAGT AGTAA TCGTC CTTTA TAGTC CTGTC GGGTT TCGCC ACCT -3' (SEQ ID NO.29)) was added to the sample. The reaction system includes: 1 × reaction buffer, 1 U Taq DNA polymerase, 0.2 mM dNTP mixture, 0.4 μM DNA probe, 1 ug DNA sample after reaction, and then react at 95 °C for 3 min to open the double strand, and then react at 63 °C for 10 min to anneal the DNA template and probe, and then react at 72 °C for 10 min to extend the template DNA, and add a dATP to the 3' end of the extension product.
[0077] 4) Adapter ligation: Adapter ligation system (20 μL): 2 μL 10 × T4 Ligase Buffer, 2 U T4 DNALigase, 100-500 ng post-reaction DNA, 5 μL 4 μM adapter P5 (5'- AATGA TACGG CGACC ACCGA GATCTACACT ATAGC CTACA CTCTT TCCCT ACACG ACGCT CTTCC GATCT -3' (SEQ ID NO.30); 3'-TTACT ATGCC GCTGG TGGCT CTAGA TGTGA TATCG GATGT GAGAA AGGGA TGTGC TGCGA GAAGGCTAGA -5' (SEQ ID NO.31)), add water to a total volume of 20 μL. React at 16 ℃ for 2 h for adapter ligation.
[0078] 5) PCR amplification and Sanger sequencing: The PCR system was 5 μL 10 × Taq DNA Polymerase Buffer, 1U Taq DNA Polymerase, 8 pmol 235-F (5'- AATGA TACGG CGACC ACCGA GATCT AC -3' (SEQ ID NO.32)), 8 pmol 235-R (5'- CACGT AGCAC TTCAT GCAGT AGTAA TCGTC C -3' (SEQ ID NO.33)), 2 μL template DNA, 0.2 mM dNTP, and water was added to the total volume of 50 μL. The PCR program was performed on a PCR instrument. The specific program for qPCR was: 95 °C, 3 min, (95 °C, 30 s; 65 °C, 60 s; 72 °C, 60 s) × 15 cycles.
[0079] The feasibility and accuracy of the method were analyzed by agarose gel electrophoresis and Sanger sequencing results.
[0080] When the method was tested for the enrichment of uracil-modified DNA, the results were Figure 3 A, the first lane without adding UdgX-H109S glycosidase has no band, and the second lane with adding UdgX-H109S glycosidase has a single band, which was sequenced by Sanger sequencing. The results showed that the DNA chain terminated before the uracil modification and was directly connected to the linker P5.
[0081] When the method was tested for the enrichment of DNA containing AP sites, the results were Figure 3 B, the first lane without adding APE1 enzyme has no band, while the second lane with adding APE1 enzyme has a single band, which was sequenced by Sanger sequencing. The results showed that the DNA chain terminated before the AP site and was directly connected to the linker P5.
[0082] When the method was tested for the enrichment of 8OG-modified DNA, the results were shown in Figure 3 C, the first lane without adding Pab-AGOG glycosidase has no band, and the second lane with adding Pab-AGOG glycosidase has a single band, which was sequenced by Sanger sequencing. The results showed that the DNA chain terminated before the 8OG modification and was directly connected to the linker P5.
[0083] The above results indicate that the glycosidase-assisted sequencing method for trace modifications can selectively enrich DNA containing trace modifications and terminate at the position before the trace modification.
[0084] Example 3
[0085] High-throughput sequencing was used to detect the accuracy of the glycosidase-assisted sequencing micro-modification method. The specific steps are as follows:
[0086] 1) First, two dsDNAs with basically the same length but completely different sequences were obtained by PCR. One DNA was a 329-bp dsDNA-U / A containing uracil modification (5'- GTGTG TGCGT GTGTG TGCGT GTGTG TGTT UGTGTGTGTGC GGTGT GAGGT ATGTG CCCCG CCACA AGTTC AGCGT GTCCA CTACG ACCTG CCCAT AGAGCTGTCA AGCGG AAGAT GGAGA ACCGT GATTA CCGGG GAACG CCAGG GATGA GATGG GGCGA GGGCGAGGGC GATGC CACCT ACGGC AAGCT GACCC TGAAG TTCAT CTGCA CCACC GGCAA GCTGC CCTCGTGACC ACCCT GACCT ACGGC GTGCA GTGCT TCAGC CGCTA CCCCG ACCAC ATGAA GCAGC ACGACTTCGC TCCCC TCTCT AAGGA AGTCG GGGAA GCGG -3’ (There is a minor modified base U between the T base at position 29 and the G base at position 30 in the sequence shown in SEQ ID NO. 34); 3’- CACAC ACGCA CACAC ACGCACACAC ACAAA CACAC ACACG CCACA CTCCA TACAC GGGGC GGTGT TCAAG TCGCA CAGGT GATGCTGGAC GGGTA TCTCG ACAGT TCGCC TTCTA CCTCT TGGCA CTAAT GGCCC CTTGC GGTCC CTACTCTACC CCGCT CCCGC TCCCG CTACG GTGGA TGCCG TTCGA CTGGG ACTTC AAGTA GACGT GGTGGCCGTT CGACG GGAGC ACTGG TGGGA CTGGA TGCCG CACGT CACGA AGTCG GCGAT GGGGC TGGTGTACTT CGTCG TGCTG AAGCG AGGGG AGAGA TTCCT TCAGC CCCTT CGCC -5’ (SEQ IDNO.35)6 is a 328-bp dsDNA-T / A(5'- TTCAA GTCCG CCATGCCCGT CGGAT GAGTT GACTT GCTGG CGTTT TTCCA TAGGC TCCGC CCCCC TGACG AGCAT CACAAAAATC GACGC TCAAG TCAGA GGTGG CGAAA CCCGA CAGGA CAAGG CTACG AAGCT TGGCG TAATCATGGT CATAG CTGTT TCCTG TGTGA AATTG TTATC CGCTC ACAAT TCCAC ACAAC ATACG AGCCGGAAGC ATAAA GTGTA AAGCC TGGGG TGCCT AATGA GTGAG CTAAC TCACA TTAAT TGCGT TGCGCTCACT GCCCG CTTTC CAGTC GGGAA ACCTG TCGTG CCAGC TGCAT TAATG AAT -3' (SEQ IDNO.36):3'- AAGTT CAGGC GGTAC GGGCA GCCTA CTCAA CTGAA CGACC GCAAA AAGGT ATCCGAGGCG GGGGG ACTGC TCGTA GTGTT TTTAG CTGCG AGTTC AGTCT CCACC GCTTT GGGCT GTCCTGTTCC GATGC TTCGA ACCGC ATTAG TACCA GTATC GACAA AGGAC ACACT TTAAC AATAG GCGAGTGTTA AGGTG TGTTG TATGC TCGGC CTTCG TATTT CACAT TTCGG ACCCC ACGGA TTACT CACTCGATTG AGTGT AATTA ACGCA ACGCG AGTGA CGGGC GAAAG GTCAG CCCTT TGGAC AGCAC GGTCGACGTA ATTAC TTA -5' (SEQ ID NO.37)). 329-bp dsDNA-U / A and 328-bp dsDNA-T / A were mixed in different proportions and divided equally into control group and experimental group for reaction and NGS analysis. The proportion of initial DNA can be calculated by adding inactivated UdgX-H109S to the control group; the experimental group was added with active UdgX-H109S, and its PCR product was only truncated DNA from 329-bp dsDNA-U / A. The calculation result was the proportion of DNA enriched with uracil modification. The enrichment of DNA containing uracil modification can be calculated by comparing the control group with the experimental group. Since the DNA in the control group cannot be connected to the adapter with a hydroxyl group at the 5' end, no PCR product is generated at this time, resulting in the inability to determine the proportion of initial DNA in the control group. Therefore, the adapter P5-1 with a phosphate group at the 5' end was used in the control group (the sequence of the adapter P5-1 is the same as that of the adapter P5, except that the 5' end of the adapter P5-1 is a phosphate group), but the adapter P5 with hydroxyl groups at both the 3' and 5' ends was still used in the experimental group. .
[0087] 2) The enrichment of 329-bp dsDNA-U / A in the experimental group was investigated after 329-bp dsDNA-U / A and 328-bp dsDNA-T / A were mixed in a ratio of 1:9. Analysis of the original 5G NGS sequencing data obtained showed that the control group and the experimental group had a matching rate of up to 95.69% and 96.5% with the original DNA sequence, respectively; only 6% of the DNA detected in the control group came from 329-bp dsDNA-U / A, and the remaining 94% of the DNA came from 328-bp dsDNA-T / A, which was not much different from the initial input ratio (such as Figure 4 As shown in A); in the experimental group, up to 99.6% of the DNA detected came from 329-bp dsDNA-U / A, and only 0.4% of the DNA came from 328-bp dsDNA-T / A. False positives in the background can be removed by subsequent bioinformatics analysis (such as Figure 4 B). Next, the enrichment of uracil-modified DNA was calculated by qPCR. According to formula 2 C T (处理组329-bp dsDNA-U / A比例)−C T (对照组329-bp dsDNA-U / A比例) It was calculated that the method of the present invention has a 1920-fold enrichment for 329-bp dsDNA-U / A (e.g. Figure 4 C), which shows that the method of the present invention can accurately and efficiently enrich DNA containing uracil modifications. In addition, the NGS analysis results show that the breakpoints of up to 99.28% of the DNA enriched in the experimental group are located before uracil, which shows that the method of the present invention can accurately sequence and analyze uracil on DNA (such as Figure 4 D).
[0088] Example 4
[0089] Sequencing was used to detect uracil modification at Chr9:1061185001 and Chr1:226143862 in HEK293T cells. The specific steps are as follows:
[0090] 1) First, APE1 was added to remove the influence of the natural AP site. During the experiment, 2000 ng of human embryonic kidney cell 293 (HEK293T) genomic DNA was extracted and divided into two equal parts, one for the control group and the other for the experimental group for genomic DNA library construction. Inactivated UdgX-H109S was used in the control group, and full-length DNA with a hydroxyl group at the 5' end was obtained after extension; active UdgX-H109S was added to the experimental group, and the uracil in the DNA was cut, so the DNA containing uracil modification was extended to obtain truncated DNA with a phosphate group at the 5' end, while the DNA without uracil modification was extended to obtain full-length DNA with a hydroxyl group at the 5' end. The full-length DNA was subsequently removed by connecting the adapter 2 (i.e., adapter P5) with a hydroxyl group at the end. However, no PCR product was generated in the control group at this time, and the original DNA sequence could not be displayed. Therefore, the control group's adapter was changed to adapter P5-1 with a phosphate group at the 5' end. At this time, the DNA sequence in the PCR product was the initial DNA sequence of the genome; the experimental group used adapter P5 with hydroxyl groups at both the 3' and 5' ends. The template for PCR amplification came from the truncated DNA modified with uracil, and the breakpoint was before uracil.
[0091] 2) Select Chr9:1061185001 containing uracil modification reported in the literature (e.g. Figure 5 A) and Chr1:226143862 (as Figure 5 B) site for detection. The Sanger sequencing results showed that the Chr9:1061185001 and Chr1:226143862 sites in the control group were both read as T; while in the experimental group, the Sanger sequencing results stayed at the Chr9:1061185000 and Chr1:226143861 positions, i.e., the position before the Chr9:1061185001 and Chr1:226143862 positions, which showed that the method of the present invention can accurately cut uracil on genomic DNA, and the sequence of the DNA in the control group is the initial DNA sequence. The above results prove that the method of the present invention can successfully enrich DNA containing uracil modification and achieve accurate sequencing of uracil modification in genomic DNA. Therefore, the method of the present invention can quickly and easily sequence and analyze the uracil modification level at any position of the genome.
[0092] In summary, the method provided by the present invention can be used for single-base resolution positioning and sequencing of trace modifications in low-abundance biological samples, and for analyzing changes in key trace modification sites in biological samples. It provides a powerful tool for in-depth exploration of the role of trace modifications in various diseases and has important scientific and clinical value.
[0093] The above specific embodiments describe the implementation of the present invention in detail, but the present invention is not limited to the specific details in the above embodiments. Within the scope of the claims and technical concept of the present invention, the technical solution of the present invention can be modified and changed in many simple ways, and these simple modifications all belong to the protection scope of the present invention.
Claims
1. A glycosidase-assisted single-base resolution sequencing method for trace modified bases in DNA, characterized in that: The following steps are included: (1) Fragment the DNA, then perform end repair and linker ligation to make the DNA ends free hydroxyl groups; (2) dividing the DNA sample treated in step (1) into two equal parts, and treating at least one of the DNA samples with glycosidase and / or APE1 enzyme; (3) Add DNA probes to the two DNA samples respectively and perform extension reaction with DNA polymerase; (4) Recover the extended product and connect it to the adapter P5 with a free hydroxyl group at the 5' end, and achieve single-base resolution sequencing and micro-modification based on PCR amplification and sequencing.
2. The method for sequencing trace modified bases in DNA with single base resolution assisted by glycosidase according to claim 1, characterized in that: In step (2), the DNA sample is processed, and a corresponding processing method is selected according to the trace modified bases to be sequenced: When the trace modification to be sequenced is uracil or 8OG, step (2) is specifically as follows: the DNA sample treated in step (1) is equally divided into two parts, one part is not treated with glycosidase, and the other part is treated with glycosidase, and the two DNA samples are then treated with APE1 enzyme; When the trace modification to be sequenced is an AP site, step (2) is specifically as follows: the DNA sample treated in step (1) is equally divided into two parts, one part is not treated with APE1 enzyme, and the other part is treated with APE1 enzyme.
3. The method for sequencing trace modified bases in DNA with single base resolution assisted by glycosidase according to claim 2, characterized in that: When the trace modification to be sequenced is uracil or 8OG, the conditions for treating the DNA sample with the glycosidase are: reacting the glycosidase with the DNA sample in a glycosidase buffer at 37°C.
4. The method of single-base resolution sequencing of trace modified bases in DNA assisted by glycosidase according to claim 2, characterized in that: After the glycosidase treatment, the conditions for using the APE1 enzyme to treat the DNA sample are as follows: adding the APE1 enzyme and the APE1 buffer to the reaction system and reacting at 37°C.
5. The method for sequencing trace modified bases in DNA with single base resolution assisted by glycosidase according to claim 3, characterized in that: When the trace modification to be sequenced is an AP site, the conditions for treating the DNA sample with the APE1 enzyme are: reacting the APE1 enzyme with the DNA sample in an APE1 enzyme buffer at 37°C.
6. A kit for localizing and detecting trace modifications in DNA, characterized in that: Contains glycosidase, APE1 enzyme, and DNA polymerase.
7. A kit for locating and detecting trace modifications in DNA according to claim 6, characterized in that: Also contains the reaction buffer for the corresponding protein or enzyme.
8. Application of glycosidase or APE1 enzyme in localizing and detecting trace modifications in DNA.
9. Use of glycosidase or APE1 enzyme in the preparation of a kit for locating and detecting trace modifications in DNA.