Construction method and application of single-stranded DNA library capable of efficiently capturing short fragments

The TAS-seq library construction method solves the challenges of library construction with low starting amounts and short DNA fragments through steps such as polyC tailing, linear amplification, and clip adapter ligation. It achieves efficient preservation of short DNA fragments, improves the sensitivity and accuracy of detection, and is applicable to various DNA sample types.

CN121344152APending Publication Date: 2026-01-16UNIV OF SCI & TECH OF CHINA +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511284125.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively handle low starting amounts and short fragments of single-stranded DNA when constructing sequencing libraries, leading to information loss and low detection efficiency, especially for difficult samples such as FFPE, cfDNA, and ctDNA.

Method used

The TAS-seq library construction method was used to construct a high-efficiency single-stranded DNA library by performing polyC tailing, linear amplification, clip adapter ligation, and PCR amplification on single-stranded DNA, combined with magnetic bead purification technology, which preserved short DNA fragments and improved library construction efficiency.

Benefits of technology

Under conditions of low starting amount and short fragment, the TAS-seq library preparation method can obtain more comprehensive information on rare DNA samples, improve detection sensitivity and accuracy, and is suitable for library preparation applications of pure single-stranded, pure double-stranded and mixed DNA. It has a wide range of applications and is simple and fast to operate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121344152A_ABST
    Figure CN121344152A_ABST
Patent Text Reader

Abstract

The invention relates to a construction method and application of a single-stranded DNA library. Specifically, linear amplification is carried out after polyC is tailed, the efficiency of joint connection is improved, and in the amplification process, biotin-labeled dCTP is added according to a certain proportion, so that short fragments in the library are reserved to the greatest extent. The method is suitable for low initial quantity, and library construction can also be carried out on rare samples with damaged fragments and serious degradation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of gene analysis and detection, and specifically to a method for constructing a single-stranded DNA library. Background Technology

[0002] Next-generation sequencing (NGS) is fundamental to current biological and biomedical research. Library construction is a crucial step in NGS research, and its quality and coverage directly impact the quality of subsequent data analysis. Low sample DNA quality and low starting amounts have long been bottlenecks limiting NGS applications. Many challenging samples, such as FFPE (formalin-fixed paraffin-embedded) samples, cfDNA, ctDNA, and ancient DNA, are characterized by short fragment lengths and internal DNA breaks. Using standard double-stranded DNA library construction methods requires patching the protruding ends of double-stranded DNA (dsDNA) molecules before adapter ligation, resulting in the loss of a portion of the original sequence (terminal motif). Furthermore, since adapters are used with double-stranded DNA ligation, double-stranded library construction ignores single-stranded DNA (ssDNA) molecules, excluding them from the final library. During double-stranded DNA library construction, adapter ligation is typically performed at a 10:1 ratio (adapter to DNA fragment). With very low starting DNA amounts, this ratio can lead to adapter dimer formation, resulting in the loss of valuable data during PCR library amplification. Adapter dimers can be purified and removed by optimizing the magnetic bead ratio to retain larger molecular weight DNA, but short double-stranded fragments will also be lost. Currently, several kits are available for short DNA library construction, but many problems remain. For example, the VAHTS ssDNALibrary Prep Kit for Illumina uses ligation at both the 3' and 5' ends, resulting in low ligation efficiency. Furthermore, excess adapters removed by magnetic bead sorting also remove short fragments, making it unsuitable for detecting short DNA fragments. Similarly, IDT's xGen™ ssDNA & Low-Input DNA Library Preparation Kit uses magnetic bead sorting, but it can only capture ssDNA of 200-350 nt in size, making it unsuitable for detecting shorter ssDNA. Summary of the Invention

[0003] To address the problems existing in the prior art, this invention proposes a single-stranded DNA library construction method that is more suitable for low starting amounts and short DNA fragments.

[0004] The purpose of this invention is to maximize the preservation of short fragments in the library and improve library construction efficiency when constructing single-stranded DNA libraries with short DNA fragments and low starting amounts.

[0005] Terminology Explanation:

[0006] cfDNA stands for cell-free DNA, which refers to DNA fragments that are free outside the cell.

[0007] ctDNA stands for circulating tumor DNA, which is DNA released into the bloodstream by tumor cells and circulates throughout the body.

[0008] UMI, or Unique Molecular Identifier, is a molecular barcode that can specifically identify molecules in a sample library and can also be used for error correction during sequencing.

[0009] An index is a tag or marker used in sequencing to distinguish samples.

[0010] bio-dCTP refers to biotinylated dCTP.

[0011] The present invention adopts the following technical solution:

[0012] On one hand, the present invention provides a method for constructing a single-stranded DNA library, the method comprising the following steps:

[0013] (1) PolyC tailing is performed on the fragment of the single-stranded DNA to obtain a tailing product, wherein the tail contains N bases C, preferably N is 10-15, more preferably N is 10;

[0014] (2) Linear amplification of the tailed product is performed using tailing primers, dCTP, and bio-dCTP to obtain a linear amplification product, wherein the linear amplification product includes an amplification strand labeled with bio-dCTP and the original template strand;

[0015] (3) Separate the bio-dCTP-tagged amplified strand from the original template strand;

[0016] (4) Connect the bio-dCTP-tagged amplification strand to the clamp connector to obtain the ligation product;

[0017] (5) The ligation product is subjected to PCR amplification to obtain a DNA library.

[0018] In some implementations, the ratio of dCTP to bio-dCTP is 9:1, 8:2, 7:3, 6:4 or 1:1, preferably 6:4.

[0019] In some embodiments, the sequence of the tailing primer is shown in SEQ ID NO.1.

[0020] In some embodiments, the clamp connector includes a UMI sequence, the length of which is preferably 8-12 bases, preferably 12 bases, wherein each base is independently A, T, C or G. The T base is included in the UMI sequence to prevent the formation of secondary structures between the connectors.

[0021] The splint connector is generated by annealing the forward and reverse primers of the splint connector. Preferably, the sequence of the forward primer of the splint connector is shown in SEQ ID NO.2, and the sequence of the reverse primer of the splint connector is shown in SEQ ID NO.3. In step (5), PCR amplification uses library amplification P5 primer and library amplification P7 primer. The library amplification P5 primer includes, from the 5' end to the 3' end, the following sequence: first fragment, index 2, and second fragment. The sequence of the first fragment is: AATGATACGGCGACCACCGAGATCTACAC (SEQ ID NO.4), and the sequence of the second fragment is: ACACTCTTTCCCTACACGACGCTCTTC (SEQ ID NO.5). The length of the index 2 sequence is at least 8 bases, wherein the bases are independently A, T, C, or G.

[0022] In some embodiments, the library amplification P7 primers comprise, from the 5' end to the 3' end, a third fragment, index 1, and a fourth fragment; the sequence of the third fragment is CAAGCAGAAGACGGCATACGAGAT (SEQ ID NO. 6), the sequence of the fourth fragment is GTGACTGGAGTTCAGACG (SEQ ID NO. 7), and the length of index 1 is at least 8 bases, wherein the bases are independently A, T, C, or G.

[0023] In some implementations, step (2) is followed by a purification step, which uses DNA-binding magnetic beads to purify, specifically, to remove free bio-dCTP.

[0024] In some implementations, step (4) is followed by a purification step, specifically, purification using streptavidin magnetic beads to purify the ligation product.

[0025] In some implementations, step (5) is followed by a purification step, using gel extraction or DNA sorting magnetic beads, specifically to purify the PCR amplification product.

[0026] In some embodiments, the DNA fragment is 30-500 nt in length, preferably 30-400 nt, more preferably 30-300 nt, and most preferably 30-200 nt.

[0027] In some implementations, the DNA fragment is 30-70 nt or 70-170 nt in length.

[0028] On the other hand, the present invention provides a DNA library, which is constructed by the above-described method, preferably a sequencing library.

[0029] On the other hand, the present invention provides a method for DNA detection, the method comprising constructing a library using the above-described method and sequencing or analyzing the library.

[0030] The DNA library preparation method of this invention is named TAS-seq library preparation (TAS-seq, Tailing Amplification Splint ligation sequencing). An exemplary flowchart of the TAS-seq library preparation method is shown below. Figure 1 .

[0031] Because double-stranded DNA can be denatured into single-stranded DNA, and various secondary structures of DNA, such as G quadruplexes and other non-classical B double helix forms, can also be denatured into single-stranded DNA, theoretically this invention can capture all types of DNA fragments and distinguish between positive and negative strand information.

[0032] Compared with the prior art, the present invention has the following technical effects:

[0033] Compared to currently widely used ssDNA library construction methods, this invention is more suitable for lower starting amounts of ssDNA, retaining at least 50 nt of insert fragments. It is more feasible for rare samples with damaged or severely degraded fragments, enabling more comprehensive acquisition of information and characteristics from rare DNA samples, such as DNA methylation detection, ancient DNA detection, and liquid biopsy. This invention, through steps such as tailing, linear amplification, and splint adapter ligation, maximizes the amount of original DNA template, improves purification efficiency, retains all DNA fragments, and enhances the capture efficiency of short DNA fragments. The DNA library construction method of this invention can be stably used with a starting amount of 1 ng of DNA, exhibits high sensitivity, no bias, wide applicability, and is fast and simple to operate. Library construction can be completed within one day after obtaining the DNA sample. It can be used for pure single-stranded DNA library construction applications (e.g., Example 1), pure double-stranded DNA library construction applications (e.g., Example 2), and mixed single- and double-stranded DNA library construction applications (e.g., Example 3), and can retain short ssDNA, thus having a wide range of applications. Attached Figure Description

[0034] Figure 1This is an exemplary flowchart of the TAS-seq library construction method. The blue and red segments represent Illumina connector sequences, the red segment represents the index sequence, the yellow segment represents the UMI sequence, and the green segment represents the annealing matching sequence of the P5 clamp connector.

[0035] Figure 2 This is an agarose gel image of a TAS-seq library constructed using specific primer sequences of 44 nt and 108 nt synthesized with 1 ng DNA.

[0036] Figure 3 The images show denaturing polyacrylamide gel electrophoresis (PAG) images of linear amplification products with different dCTP and bio-dCTP ratios during the linear amplification process in Example 2 (Figure A), and denaturing polyacrylamide gel electrophoresis images of P5 clip connector ligation products (Figure B).

[0037] Figure 4 These are library quality control images from TAS-seq library construction of ChIP samples. A represents library construction using input (genomic background) samples, B represents library construction using histone H3K4me3 samples, C represents library construction using histone H3K27ac samples, and D represents library construction using plasma cfDNA samples.

[0038] Figure 5 The signal distribution of histones H3K4me3 and H3K27ac in the gene body is shown, with TSS representing the transcription start site and TES representing the transcription termination site.

[0039] Figure 6 The images show the fragment distribution (A) and peak distribution (B) of plasma cfDNA detected using the TAS-seq library preparation method. Here, 30-70 refers to DNA fragments with lengths between 30 and 70, and 70-170 refers to DNA fragments with lengths between 70 and 170. Detailed Implementation

[0040] To make the objectives and technical solutions of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings.

[0041] The single-stranded DNA sample used for library construction in the method of the present invention can be single-stranded DNA from any biological sample, such as single-stranded DNA extracted from a biological sample using conventional methods in the art, or single-stranded DNA obtained from double-stranded DNA from a biological sample by denaturation or other means.

[0042] Suitable starting DNA samples for library construction using the method of this invention include histone ChIP samples, cfDNA samples, etc. Histone ChIP samples refer to histone-labeled DNA samples that have undergone chromatin immunoprecipitation (ChIP). The preparation method can specifically involve performing chromatin immunoprecipitation (ChIP) experiments on DNA from biological samples, followed by cell fixation, grinding, sonication fragmentation, overnight incubation with histone (e.g., H3K4me3 and H3K27ac) antibodies, decrosslinking, purification, and other procedures to finally obtain input samples (double-stranded DNA samples with genomic background) and histone-labeled double-stranded DNA samples.

[0043] cfDNA samples are generally derived from plasma and can be obtained using methods known to those skilled in the art. Phenol-chloroform extraction and magnetic bead extraction are preferred methods for extracting cfDNA samples to preserve short DNA fragments to the greatest extent possible.

[0044] Particularly suitable DNA samples include short single-stranded DNA fragments, such as single-stranded DNA with a length of 30-500 nt, preferably 30-400 nt, more preferably 30-300 nt, and most preferably 30-200 nt.

[0045] The length of the polyC tail used in the method of the present invention can be determined by those skilled in the art based on specific experimental conditions, preferably 10-15 C, more preferably 10 C.

[0046] In the polyC tailing reaction of this invention, TdT terminal transferase and T4 polynucleotide kinase can be used to catalyze the binding of deoxynucleotides to the 3' end of the DNA molecule. The role of T4 polynucleotide kinase is to remove any phosphate groups that may be present at the 3' end of the DNA fragment. Phosphate groups reduce tailing efficiency and decrease the overall library preparation efficiency. If the sample DNA fragment does not have phosphate groups at the 3' end, T4 polynucleotide kinase may not be added.

[0047] In the linear amplification of the method of the present invention, biotin labeling is used to label the template strand, preferably bio-dCTP.

[0048] Unless otherwise specified, the experimental methods described in the following tests are conventional methods; for tests where specific techniques or conditions are not specified, they shall be performed in accordance with the techniques or conditions described in the literature in this field or in accordance with the product instructions; unless otherwise specified, the reagents and materials described are commercially available.

[0049] Example 1: Constructing a library using ssDNA

[0050] In this embodiment, specific lengths of ssDNA, such as 44nt and 108nt, are used. Using specific lengths of synthetic DNA facilitates gel electrophoresis to check the efficiency of each library preparation step.

[0051] In this embodiment, TdT terminal transferase and T4 polynucleotide kinase were purchased from NEB; T4 DNA ligase, T4 DNA ligase buffer, and phanta mix were purchased from Novizan; Triton X-100, dGTP, dATP, dTTP, and dCTP were purchased from Yisheng Biotechnology; bio-dCTP was purchased from Jena Bioscience; 2G hot-start DNA polymerase and 2G buffer A were purchased from KAPA; RNase A and PEG4000 were purchased from Thermo Fisher Scientific; protein A / G magnetic beads, DNA smarter binding magnetic beads, and DNA sorting magnetic beads were purchased from Yisheng Biotechnology; streptavidin magnetic beads were purchased from NEB; H3K4me3 and H3K27ac antibodies were purchased from ABclonal; DNA binding column was purchased from Thermo Fisher Scientific; PFA fixative was purchased from Sinopharm Group; PMSF was purchased from Sangon Biotech; and BSA was purchased from BioFroxx.

[0052] Table 1. Primers and their sequences

[0053]

[0054] In SEQ ID NO.1, H represents a non-G base, and N is independently A, T, C, or G.

[0055] In SEQ ID NO.2 and 3, NNNTNNNNTNNN is a UMI sequence with a length of 12 bases, where N is independently A, T, C or G, and T is added to prevent secondary structures from forming between linkers; P indicates terminal phosphate modification; NH2C7 indicates C7 amino modification.

[0056] The P5 primers for library amplification consist of the following sequences from the 5' end to the 3' end: a first fragment, index 2, and a second fragment. The sequence of the first fragment is: AATGATACGGCGACCACCGAGATCTACAC (SEQ ID NO.4), and the sequence of the second fragment is: ACACTCTTTCCCTACACGACGCTCTTC (SEQ ID NO.5). The length of the index 2 sequence is at least 8 bases, wherein the bases are independently A, T, C, or G.

[0057] The P7 primers for library amplification consist of the following sequences from the 5' end to the 3' end: a third fragment, index 1, and a fourth fragment. The sequence of the third fragment is CAAGCAGAAGACGGCATACGAGAT (SEQ ID NO. 6), and the sequence of the fourth fragment is GTGACTGGAGTTCAGACG (SEQ ID NO. 7). The length of index 1 is at least 8 bases, wherein the bases are independently A, T, C, or G.

[0058] The primers and clip connectors of this invention were synthesized by Shanghai Sangon Biotech Co., Ltd., and purified by polyacrylamide gel electrophoresis (PAGE) or high performance liquid chromatography (HPLC).

[0059] The P5 splint adapter was obtained by annealing the P5 splint adapter forward primer (TAS P5 UMI forward primer) and the P5 splint adapter reverse primer (TAS P5 UMI reverse primer) in the following reaction system: annealing program: 95 ℃ for 5 min; cooling to 25 ℃ at -0.1 ℃ / s. A 25 μM P5 splint adapter was obtained using this system and stored at -20 ℃ for later use.

[0060] Table 2. Reaction system for obtaining P5 clamp joints

[0061]

[0062] ssDNA library construction includes the following steps:

[0063] 1) DNA polyC tailing reaction

[0064] Add the DNA sample to a PCR tube and denature at 95°C for 3 min in a polyC-tailing reaction system, then immediately place on ice for at least 3 min. Perform polyC-tailing of DNA using TdT terminal transferase and T4 polynucleotide kinase in the system shown in Table 3.

[0065] Table 3. Reaction system for DNA polyC tailing

[0066]

[0067] The PCR tubes were then placed in a PCR instrument, and the reaction program was set and run: incubation at 37 °C for 1 h (adding a certain number of C bases to the 3' end of the DNA fragment), followed by incubation at 75 °C for 10 min (heat-inactivating the enzyme). The tubes were then stored at 12 °C.

[0068] 2) Linear amplification

[0069] Linear amplification was performed to increase the amount of original DNA template, as described in Tables 2 and 3. PCR reactions were carried out in the linear amplification reaction system and stored by incubation on ice.

[0070] Table 4. Linear amplification reaction system

[0071]

[0072] Table 5. Linear Amplification PCR Reaction Procedure

[0073]

[0074] 3) Purification using magnetic beads

[0075] Purification was performed using a 2.2× DNA smarter combined with magnetic beads to remove free bio-dCTP and prevent it from occupying the magnetic beads and affecting subsequent capture efficiency. The beads were then eluted with 80 μL of enzyme-free water.

[0076] 4) Separate the amplification strand and the template strand.

[0077] 80 μL of the elution product from step 3) was denatured to separate the biotin-labeled amplified strand from the original template strand. Specifically, denaturation was performed at 95 °C for 3 min, followed immediately by placing on ice for 3 min.

[0078] 5) Streptavidin magnetic bead capture

[0079] The purpose of using streptavidin magnetic beads for capture is to remove excess linear amplification primers and template strands, while retaining all biotin-tagged (bio-dCTP) DNA fragments (of any length) that have been successfully linearly amplified.

[0080] First, wash the streptavidin beads: Place 10 μL of streptavidin beads in a centrifuge tube, add 400 μL of 1×B&W buffer, and mix by pipetting. Place the centrifuge tube on a magnetic rack and incubate at room temperature for 1 min to completely remove the supernatant. Repeat the washing process twice. Finally, resuspend the beads in 10 μL of 1×B&W buffer.

[0081] Table 6. 1 × B&W Buffer

[0082]

[0083] Then, using the cleaned streptavidin magnetic beads, incubate with shaking at room temperature for 2–4 h in the streptavidin magnetic bead capture system. Afterward, wash the streptavidin magnetic beads twice with 1×B&W buffer and 10 mM Tris-HCl (pH 7.4), respectively. Finally, resuspend the washed streptavidin magnetic beads in 7 μL of 10 mM Tris-HCl (pH 7.4).

[0084] Table 7. Streptavidin magnetic bead capture system

[0085]

[0086] 6) On-bead P5 splint connection

[0087] After thoroughly mixing by blowing in the reaction system connected by the clamp, place it on a vertical rotary mixer with the rotation speed set to 8 rpm and incubate at room temperature for 2-3 h.

[0088] Table 8. Reaction system with P5 clamp connection

[0089]

[0090] This step involves ligating the DNA fragment to a P5 splint adapter. The P5 splint adapter is formed by annealing the TAS P5 UMI forward primer (SEQ ID NO.2) and the TAS P5 UMI reverse primer (SEQ ID NO.3).

[0091] 7) Biotin affinity purification

[0092] Biotin affinity purification was performed using streptavidin magnetic beads to remove excess P5 splice adapters. The streptavidin magnetic beads were washed twice with 1×B&W buffer and 10 mM Tris-HCl (pH 7.4), respectively. Finally, the washed streptavidin magnetic beads were resuspended in 26 μL of 10 mM Tris-HCl (pH 7.4) for purification.

[0093] 8) Library expansion

[0094] The number of DNA libraries is increased by PCR exponential amplification. Specifically, P5 and P7 primers are used to ligate P5 and P7 ilkumina adapters to the DNA fragments (see [link to PCR exponential amplification]). Figure 1 The Illumina adapter includes an index sequence, an Illumina flow cell binding sequence, and a sequencing primer binding site.

[0095] The P5-terminal Illumina flow cell binding sequence is AATGATACGGCGACCACCGAGATCTACAC (SEQ ID NO.8).

[0096] The P7 terminal Illumina flow cell binding sequence is CAAGCAGAAGACGGCATACGAGAT (SEQ ID NO. 9).

[0097] The binding site for the P5 sequencing primers is: ACACTCTTTCCCTACACGACGCTCTTCCGATCT (SEQ ID NO.10).

[0098] The binding site for the P7 sequencing primers is: GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT (SEQ ID NO.11)

[0099] Table 9. Reaction system for library amplification

[0100]

[0101] Table 10. PCR reaction procedure for library amplification

[0102]

[0103] After preparing a 1.5% agarose gel, load the sample and perform electrophoresis at 120 V. Check the results after 40 min.

[0104] 9) Purification

[0105] Purification can be performed using 1× DNA sorting magnetic beads or by gel extraction. The DNA can be eluted with 20 μL of 10 mM Tris-HCl (pH 7.4) and then subjected to next-generation sequencing.

[0106] Under the condition of an initial amount of 1 ng DNA, such as Figure 2 As shown, this library construction method can retain 44 nt and 108 nt single-stranded DNA, and the gel image shows that the library is free of adapter contamination, indicating high library construction efficiency.

[0107] Example 2: Marking effect and P5 clamp connector connection efficiency of different dCTP and bio-dCTP ratios

[0108] The starting DNA was 50 nt in length. The first seven steps of the method described in Example 1 were used, the difference being the different dCTP:bio-dCTP ratios used during linear amplification: 7:3, 6:4, and 1:1. The results are as follows: Figure 3 As shown in Figure A, when the dCTP:bio-dCTP ratio is 6:4, grayscale analysis of the electrophoresis image shows that only a very small amount of amplification product remains in the streptavidin binding buffer. Figure 3The red arrow in section A indicates the remaining unlabeled amplification product (present in the streptavidin binding buffer); the labeled linear product has already bound to the streptavidin beads. When the dCTP:bio-dCTP ratio is 1:1, there is virtually no residual amplification product, indicating that labeling efficiency increases as the dCTP:bio-dCTP ratio decreases. Figure 3 B. Electrophoresis grayscale analysis revealed that the P5 clip adapter ligation efficiencies under different dCTP:bio-dCTP ratios were 7:3 (65%), 6:4 (53%), and 1:1 (50%), all showing high P5 adapter ligation efficiencies. However, this also suggests that ligation efficiency decreases as the dCTP:bio-dCTP ratio decreases. When using different DNA samples for library construction, a suitable dCTP:bio-dCTP ratio can be selected based on the DNA fragment length to achieve good library construction results. Considering both labeling efficiency and ligation efficiency, a 6:4 dCTP:bio-dCTP ratio is more effective when the DNA fragment is short (e.g., 50 nt). When the DNA fragment length is greater than or equal to 100 nt, higher dCTP:bio-dCTP ratios, such as 7:3, 8:2, and 9:1, can be used.

[0109] Example 3: Library construction using ChIP samples and detection of histone signal distribution in the gene body.

[0110] In this embodiment, C57BL / 6J mice were used and purchased from Hefei Jisai Biotechnology Co., Ltd. TdT terminal transferase and T4 polynucleotide kinase were purchased from NEB; T4 DNA ligase, T4 DNA ligase buffer, and phanta mix were purchased from Novizan; Triton X-100, dGTP, dATP, dTTP, and dCTP were purchased from Yisheng Biotechnology; bio-dCTP was purchased from Jena Bioscience; 2G hot-start DNA polymerase and 2G buffer A were purchased from KAPA; RNase A and PEG4000 were purchased from Thermo Fisher Scientific; protein A / G magnetic beads, DNA smarter binding magnetic beads, and DNA sorting magnetic beads were purchased from Yisheng Biotechnology; streptavidin magnetic beads were purchased from NEB; H3K4me3 and H3K27ac antibodies were purchased from ABclonal; DNA binding column was purchased from Thermo Fisher Scientific; PFA fixative was purchased from Sinopharm Group; PMSF was purchased from Sangon Biotech; and BSA was purchased from BioFroxx.

[0111] The preparation of chromatin immunoprecipitation (ChIP) related reagents is described in Table 11-17.

[0112] Table 11. Lysis buffer LB1 50 mL

[0113]

[0114] Table 12. Lysis buffer LB2 50mL

[0115]

[0116] Table 13. Lysis buffer LB3 10 mL

[0117]

[0118] Table 14. RIPA buffer (as wash buffer) 50 mL

[0119]

[0120] Table 15. TE buffer 50 mL

[0121]

[0122] Table 16. TBS buffer 10 mL

[0123]

[0124] Table 17. Elution buffer 10 mL

[0125]

[0126] ChIP sample preprocessing is performed according to the following steps:

[0127] 1. Remove the frozen testes of mice 17 days after birth from -80°C and place them in a culture dish containing pre-cooled 1×PBS. Remove the white membrane with tweezers.

[0128] 2. PFA fixation and crosslinking: Place the spermatogenic tubules (with the white membrane removed) into a 1.5 mL EP tube, add approximately 200 μL of PBS containing 1% PFA fixative, and mince the tissue with scissors; then transfer it to 10 mL of PBS containing 1% PFA fixative, and crosslink at room temperature for 10 min, inverting occasionally to ensure complete crosslinking; add 500 μL of 2.5 M glycine and 500 μL of 10% BSA, with a final concentration of 0.5% BSA to reduce tissue adhesion to the tube wall, and terminate crosslinking at room temperature for 5 min.

[0129] 3. Grinding: Centrifuge at 4000 g for 5 min at room temperature; discard the supernatant, resuspend in 1 mL of PBS containing 0.5% BSA, transfer to a homogenizer, and move the homogenizer up and down approximately 15 times; filter through a 70 μm filter into a 50 mL centrifuge tube, and rinse the filter with PBS containing 0.5% BSA; transfer the cell suspension to a 1.5 mL EP tube, centrifuge at 2000 g for 5 min at 4°C, and discard the supernatant.

[0130] 4. Lysis and membrane perforation: Resuspend cells in 1 mL of lysis buffer LB1 (add 10 μL of 200 mM PMSF), thoroughly disperse, lyse on ice for 5 min, centrifuge at 2000 g for 5 min at 4°C, and discard the supernatant.

[0131] Then resuspend the cells in 1 mL of lysis buffer LB2 (add 10 μL of 200 mM PMSF and 50 μL of 10% BSA), thoroughly disperse the cells, centrifuge at 2000 g for 5 min at 4 °C, and discard the supernatant; repeat this step.

[0132] Then resuspend the cell nuclei in 500 μL of lysis buffer LB3 (with an additional 10 μL of 200 mM PMSF and 50 μL of 10% BSA), thoroughly disperse the nuclei, and incubate at room temperature for 10-15 min.

[0133] 5. Ultrasonic interruption: Take 560 μL of the product from step 4 and use a non-contact ultrasonic instrument to perform at least 16 cycles of 30 s on and 30 s off.

[0134] 6. Decrosslinking: After sonication, 40 μL of sample can be taken out and 0.8 μL of 5 M NaCl, 1 μL of LNase A and 1 μL of proteinase K can be added respectively. Decrosslinking is performed at 37 ℃ for 1 h and 65 ℃ for 10 h.

[0135] 7. Purification: Elute with DNA sorting magnetic beads, perform 1.5% agarose gel electrophoresis, and check the sonication effect. Ideally, the fragment distribution should be between 200-500 bp. If it meets the requirements, proceed to the next step.

[0136] 8. Add the sonicated sample to a final volume of 1.8 mL using lysis buffer LB3, then add 200 μL of 10% Triton X-100, for a total volume of 2 mL. Aliquot this volume into two 1.5 mL EP tubes. Take 20 μL from each tube and mix them as the input sample. The remaining 980 μL in each tube are labeled ChIP sample A and ChIP sample B, respectively. The samples can be stored at -20°C.

[0137] 9. Antibody and magnetic bead pre-binding: First, wash protein A and protein G magnetic beads with 1 mL of 0.5% BSA, repeating the washing twice. Then, add 1 mL of 0.5% BSA and incubate at room temperature for 30 min to block the magnetic beads, reducing non-specific binding. After blocking, remove the supernatant and resuspend each bead in 250 μL of 0.5% BSA. Mix the two types of magnetic beads thoroughly and aliquot them into two new centrifuge tubes (tube A and tube B). Add 4 μg of antibody H3K4me3 to tube A and 4 μg of antibody H3K27ac to tube B. Add 0.5% BSA to each tube to a final volume of 250 μL and incubate at 4°C for 1–3 h.

[0138] Then collect the magnetic beads using a magnetic rack, discard the supernatant, and wash once with 1 mL of 0.5% BSA; discard the supernatant, and resuspend each sample in 60 μL of 0.5% BSA.

[0139] 10. Antigen-antibody binding: Add tube A and tube B (containing both protein A magnetic beads and protein G magnetic beads, and already bound with the corresponding antibodies) to the ChIP sample A and ChIP sample B obtained in step 8, respectively. Mix well and incubate overnight at 4 °C. No magnetic beads or antibodies are added to the input sample.

[0140] Collect the magnetic beads using a magnetic rack and discard the supernatant. Add 1 mL of pre-cooled RIPA buffer and rotate the tube to wash the magnetic beads, avoiding pipetting to prevent sample loss. Incubate at 4 °C for 5 min after each wash, repeating the washing process at least 5 times. Finally, wash once with TBS buffer to completely remove the supernatant.

[0141] 11. Decrosslinking: For the input sample (saved in step 8), ChIP sample A and ChIP sample B (step 10), bring the volume to 200 μL with elution buffer. Add 0.8 μL RNase A, 4 μL proteinase K and 12 μL 5M NaCl to each of the above samples. Incubate at 37 °C for 30 min and at 65 °C overnight to decrosslink and obtain preliminary DNA samples.

[0142] 12. Purification: The preliminary DNA sample obtained by purification using a DNA binding column was eluted with 23 μL of 0.1×TE buffer to obtain the DNA sample, which was then stored at -80℃.

[0143] Using the obtained DNA sample, a library was constructed using the library construction method described in Example 1, and then next-generation sequencing was performed.

[0144] The method for detecting histone signal distribution on the gene body is as follows: After next-generation sequencing of the library, bioinformatics analysis is performed on the sequencing results. Specifically, the cutadapt software (https: / / cutadapt.readthedocs.io / en / stable / ) is used to remove adapters and low-quality bases for library quality control. The bwa tool (https: / / bio-bwa.sourceforge.net / ) is used to align the mouse genome to obtain a BAM file. The deeptools bamCoverage tool is used to convert the BAM file to a BW file. The deeptools computeMetrix scale-regions tool (https: / / deeptools.readthedocs.io / en / latest / ) is used to calculate the signal distribution of the library on the gene body. The gene body refers to the distance between the transcription start site and the transcription termination site of each gene. Different histones have their own characteristic distribution patterns. For example, the signal of H3K4me3 is highly enriched specifically at the transcription start site, while there is almost no signal at the transcription termination site.

[0145] Figure 4 AC displays the quality control results of the ChIP sample libraries, representing the input, K4me3, and K27ac samples respectively. The fragment distribution peaks are between 300-500bp, and there is no adapter contamination (the adapter length is between 150-160bp, and there is no obvious signal in the quality control plot at this position), indicating that the library quality is qualified.

[0146] Bioinformatics analysis results as follows Figure 5 As shown, the distribution signals of H3K4me3 and H3K27ac on the gene body are displayed. The results show that the H3K4me3 signal is highly enriched specifically at transcription start sites, while there is almost no signal at transcription termination sites. Although the signal of H3K27ac at transcription start sites is lower than that of H3K4me3, it is still enriched at transcription start sites. The input signal is evenly distributed throughout the gene body. This indicates that the TAS-seq library construction method can construct libraries from genomically derived ChIP samples, and the obtained signal distribution characteristics conform to the distribution characteristics of histones, which is consistent with theoretical results.

[0147] Example 4: A method for constructing a library using cfDNA from plasma samples, and for detecting the distribution of short fragments of cfDNA on the genome.

[0148] In this embodiment, the sample is a plasma sample from an ovarian cancer patient, provided by Anhui Provincial Hospital, and the blood sample has passed ethical review.

[0149] Plasma sample pretreatment includes the following steps:

[0150] 1. Obtaining plasma samples

[0151] After obtaining fresh blood, centrifuge it within 4 hours to obtain plasma. The centrifugation conditions are as follows: centrifuge at 1400 g for 10 min at 4 ℃, take the upper layer and transfer it to a new centrifuge tube to remove the lower layer of blood cells; then centrifuge at 14000 g for 10 min at 4 ℃ to reduce cell DNA contamination and obtain plasma samples, which are then immediately stored at -80 ℃ for later use.

[0152] 2. Obtain cfDNA

[0153] First, remove the frozen plasma sample from -80 °C, thaw it on ice, and then add 12 μL of 5M NaCl, 10 μL of 0.5 M EDTA, 30 μL of 10% SDS and 10 μL of proteinase K (purchased from Thermo Fisher Scientific) per 500 μL, and incubate overnight at 60 °C.

[0154] 2. Purification of phenol, chloroform, and isoamyl alcohol

[0155] After overnight digestion, add at least three times the volume of DNA extraction buffer (phenol: chloroform: isoamyl alcohol (volume ratio) = 25:24:1), mix well and emulsify, then centrifuge at 10000 g for 10 min, and transfer the upper aqueous phase to a new centrifuge tube.

[0156] 3. Purification using magnetic beads

[0157] Next, add 2× DNA sorting magnetic beads A and 3× isopropanol, mix well, and incubate at room temperature for at least 10 min.

[0158] DNA sorting magnetic beads A is an improvement on the original Speed ​​beads (purchased from Cytiva): First, the original Speed ​​beads are washed twice with a buffer solution of 10 mM Tris-HCl pH 8.0, 1 mM EDTA pH 8.0, and 0.05% (v / v) Tween-20. Then, the magnetic beads are resuspended in a buffer solution of 20% (w / v) PEG8000, 2.5 M NaCl, 10 mM Tris-HCl pH 8.0, 1 mM EDTA pH 8.0, and 0.05% (v / v) Tween-20, and gently mixed to avoid generating a large number of air bubbles, thus obtaining DNA sorting magnetic beads A.

[0159] After incubation, place the centrifuge tube on a magnetic rack for at least 2 minutes. Once the liquid is clear, remove the liquid, keep the tube on the magnetic rack, and wash the magnetic beads with 80% ethanol. Repeat this step twice. Then air dry for 4-5 minutes, observe the surface of the magnetic beads for cracks, add preheated elution buffer (enzyme-free water or 0.1×TE), mix thoroughly, and incubate at room temperature for at least 5 minutes.

[0160] The tube is then placed on a magnetic rack, and the liquid is transferred to a new centrifuge tube, which becomes the DNA sample for library construction. It is then stored at -80°C.

[0161] Using the obtained DNA sample, a library was constructed using the library construction method described in Example 1, followed by next-generation sequencing. The fragment distribution characteristics of the cfDNA were detected through the following steps:

[0162] 1. After next-generation sequencing of the library, bioinformatics analysis was performed on the sequencing results. Specifically, the cutadapt software (https: / / cutadapt.readthedocs.io / en / stable / ) was used to remove adapters and low-quality bases for library quality control. The bwa tool (https: / / bio-bwa.sourceforge.net / ) was used to align the mouse genome to obtain a BAM file. The deeptools bamCoverage tool was used to convert the BAM file to a BW file. The deeptools bamPEFragmentSize tool (https: / / deeptools.readthedocs.io / en / latest / ) was used to calculate the fragment distribution characteristics of the entire cfDNA library. 2. To select cfDNA fragments within different length ranges, the steps are as follows: Use the samtools software and the awk command, such as `samtools view input.bam | \awk 'length($10) >= 30 && length($10) <= 70 {print$0}' | samtools view -bS filtered.sam > filtered.bam`, which selects cfDNA fragments with a length greater than 30 but less than 70. Similarly, modifying the range will select cfDNA fragments with a length greater than 70 but less than 170.

[0163] 3. To detect the peak distribution characteristics of cfDNA, the steps are as follows: Use MACS2 software (https: / / bioconductor.org / packages / release / bioc / html / ChIPseeker.html) to perform peak calling (i.e., the location of significant signal enrichment on the genome) on the cfDNA bam file to obtain the peak position bed file, and use ChIPseeker (https: / / pypi.org / project / MACS2 / ) to annotate the peak position bed file and draw a bar chart.

[0164] Figure 4 D shows the library quality control results of the plasma cfDNA sample (10ng DNA). The fragment distribution peaks are around 200bp and 300bp, and there is no adapter contamination (the adapter length is between 150-160bp, and there is no obvious signal in the quality control map at this position), indicating that there are no obvious problems with the quality of the constructed library. Figure 6 The fragment and peak distribution of plasma cfDNA were detected using the TAS-seq library preparation method. The fragment distribution results showed that the two main peaks of cfDNA fragment length distribution were 30-70bp and 70-170bp. The peak distribution results showed that the proportion of peaks in the promoter region of 30-70bp cfDNA was significantly increased compared with that of 70-170bp cfDNA, indicating that capturing short fragments can increase the capture of cfDNA distributed in the promoter region, which contains more information about transcriptional regulation.

[0165] sequence list

[0166] SEQ ID NO.1: GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTGGGGGGGGGGHN

[0167] SEQ ID NO.2: CCAGTGCTGCTCTACATNNNNNN-NH2C7

[0168] SEQ ID NO.3: P-TGTAGAGCAGCACTGGNNNTNNNNTNNNAGATCGGAAGAGCGTCGTGTAGGGAAAGA-NH2C7

[0169] SEQ ID NO.4: AATGATACGGCGACCACCGAGATCTACAC

[0170] SEQ ID NO.5:ACACTCTTTCCCTACACGAGCTCTTC

[0171] SEQ ID NO.6: CAAGCAGAAGACGGCATACGAGAT

[0172] SEQ ID NO.7: GTGACTGGAGTTCAGACG

[0173] SEQ ID NO.8: AATGATACGGCGACCACCGAGATCTACAC

[0174] SEQ ID NO.9: CAAGCAGAAGACGGCATACGAGAT

[0175] SEQ ID NO.10: ACACTCTTTCCCTACACGACCGCTTCCGATCT

[0176] SEQ ID NO.11: GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT

[0177] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for constructing a single-stranded DNA library, comprising the following steps: (1) polyC tailing on fragments of the single-stranded DNA to obtain a tailing product, wherein the tail comprises N bases of C, preferably N is 10-15, more preferably N is 10; (2) linear amplification on the tailing product using a tailing primer, dCTP and bio-dCTP to obtain a linear amplification product, wherein the linear amplification product comprises an amplified strand labeled by bio-dCTP and an original template strand; (3) separating the amplified strand labeled by bio-dCTP and the original template strand; (4) connecting the amplified strand labeled by bio-dCTP with a splinted adapter to obtain a connection product; (5) PCR amplification on the connection product to obtain a DNA library.

2. The method of claim 1, wherein, The sequence of the tailing primer is shown in SEQ ID NO.

1.

3. The method of claim 1, wherein, The splinted adapter comprises a UMI sequence, and the splinted adapter is generated by annealing a splinted adapter forward primer and a splinted adapter reverse primer, preferably the sequence of the splinted adapter forward primer is shown in SEQ ID NO. 2, and the sequence of the splinted adapter reverse primer is shown in SEQ ID NO. 3; The PCR amplification in step (5) uses a library amplification P5 primer and a library amplification P7 primer, wherein the library amplification P5 primer comprises, in order from 5' end to 3' end, a first segment, an index 2 and a second segment, and the sequence of the first segment is shown in SEQ ID NO. 4, and the sequence of the second segment is shown in SEQ ID NO.

5.

4. The method of claim 3, wherein, The library amplification P7 primer comprises, in order from 5' end to 3' end, a third segment, an index 1 and a fourth segment, and the sequence of the third segment is shown in SEQ ID NO. 6, and the sequence of the fourth segment is shown in SEQ ID NO.

7.

5. The method of claim 1, wherein, The step (2) further comprises a purification step, and the purification is performed using DNA binding magnetic beads.

6. The method of claim 1, wherein, The step (4) further comprises a purification step, and the purification is performed using streptavidin magnetic beads.

7. The method of claim 1 wherein, The step (5) further comprises a purification step, and the purification is performed using gel purification or DNA sorting magnetic beads.

8. The method of claim 1, wherein, The length of the DNA fragment is 30-500 nt, preferably 30-400 nt, more preferably 30-300 nt, and most preferably 30-200 nt. 9.A DNA library, which is obtained by the method of any one of claims 1-8, preferably the DNA library is a sequencing library.

10. A method of DNA detection, characterized by, The method comprises constructing a library by the method of any one of claims 1-8, and sequencing or analyzing the library.