5'-ligation-based single-stranded DNA-specific high-throughput sequencing method

By using the 5' stem ring linker structure and DNA polymerase without exonuclease activity to supplement the double-stranded DNA ends, adding a 3' terminal poly(dA) tail, and connecting the 5' stem ring linker, the problem of being unable to specifically sequence single-stranded DNA in the prior art is solved, and efficient and accurate single-stranded DNA sequencing is achieved.

WO2025167161A1PCT designated stage Publication Date: 2025-08-14INSTITUTE OF BASIC MEDICAL SCIENCES CHINESE ACADEMY OF MEDICAL SCIENCES
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/123892
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-06
Filing Date
2024-10-10
Publication Date
2025-08-14

AI Technical Summary

Technical Problem

Existing high-throughput sequencing technologies cannot specifically distinguish and sequence single-stranded DNA, resulting in the inability to efficiently sequence single-stranded DNA.

Method used

Using a 5' stem ring link structure, the 5' end of double-stranded DNA is supplemented by DNA polymerase without exonuclease activity, and the 3' terminal poly(dA) tail is added, the 5' stem ring link is connected, and amplified to form a single-stranded DNA sequencing library.

Benefits of technology

High-throughput sequencing of single-stranded DNA is achieved, ensuring the accuracy and efficiency of sequencing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024123892_14082025_PF_FP_ABST
    Figure CN2024123892_14082025_PF_FP_ABST
Patent Text Reader

Abstract

Provided is a 5'-ligation-based single-stranded DNA-specific high-throughput sequencing method, relating to the technical field of biology. The present invention specifically comprises a method for preparing a 5'-end-ligation-based single-stranded DNA sequencing library. The method comprises the following steps of treating a sample to be detected: step 1) using a DNA polymerase without exonuclease activity to fill in the 5' end of dsDNA; step 2) adding a tail to the 3' end; (3) ligating a 5' stem loop adapter having undergone renaturation treatment, wherein the structure of the 5' stem loop adapter is an overhang-a random base region-a first stem region-a loop region-a second stem region, and the first stem region and the second stem region form a double strand by means of renaturation; and step 4) carrying out amplification on a ligated product. The amplified product constitutes a single-stranded DNA sequencing library.
Need to check novelty before this filing date? Find Prior Art

Description

A high-throughput sequencing method for single-stranded DNA based on 5' ligation Technical Field

[0001] The invention belongs to the field of biotechnology, and in particular relates to a high-throughput sequencing method for single-stranded DNA specificity based on 5' connection. Background Art

[0002] High-throughput sequencing technology represents a revolutionary change from traditional Sanger sequencing (also known as first-generation sequencing). It sequences hundreds of thousands to millions of nucleic acid molecules at a time, hence its name, next-generation sequencing (NGS). High-throughput sequencing technology has rapidly advanced, analyzing and identifying genes that contribute to numerous normal and pathogenic traits in humans, animals, and plants, unlocking previously unknown insights into genetics and developmental biology at the genome-wide level.

[0003] Numerous studies have shown that single-stranded DNA is generated in a variety of DNA-related biological processes, and free single-stranded DNA has recently been found in human plasma. To explore the biological significance of single-stranded DNA, the development of specific, sensitive, and robust sequencing methods is essential. Currently reported sequencing methods are unable to specifically distinguish single-stranded from double-stranded DNA in samples, making them incapable of sequencing single-stranded DNA.

[0004] The high-throughput sequencing process is generally divided into four steps. The first step in the sequencing process is DNA library preparation, which involves introducing sequencing adapters at both ends of the target DNA to prepare for subsequent steps. Library preparation determines the direction of the entire sequencing process and is the root cause of the differences between various sequencing methods. To achieve specific sequencing of single-stranded DNA, library preparation methods still need to be improved.

[0005] Summary of the Invention

[0006] The present invention aims to establish a highly specific, stable, and sensitive single-stranded DNA high-throughput sequencing technology. Specifically, the present invention provides the following technical solutions:

[0007] In one aspect, the present invention provides a 5' stem-loop linker having a structure of: protruding end-random base region-first stem region-loop region-second stem region, wherein the first stem region and the second stem region form a double strand through annealing.

[0008] Preferably, the 5' end or the 3' end of the 5' stem-loop linker is a hydroxyl group.

[0009] Preferably, both the 5' end and the 3' end of the 5' stem-loop linker are hydroxyl groups and are free of modification.

[0010] Preferably, the first stem region and the second stem region are complementary and connected to each other by hydrogen bonds through the corresponding relationship of different bases to form the stem region, and a double helix structure is formed by annealing, so that the 5' stem-loop linker forms a stem-loop structure.

[0011] Preferably, the complementary pair comprises at least 75%, 80%, 85%, 90%, 95% or complete complementarity.

[0012] Preferably, the length of the random base region is 1-12 nt.

[0013] Preferably, the length of the random base region is 6 nt.

[0014] Preferably, the length of the protruding end is 0-20 nt, specifically including 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nt; more preferably 12 nt.

[0015] Preferably, the length of the first stem region or the second stem region is 8-25 nt, specifically including 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 nt; more preferably 11 nt.

[0016] Preferably, the lengths of the first stem region and the second stem region may be the same or different.

[0017] Preferably, the length of the loop region is 0-50nt, specifically including 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50nt; preferably, the length of the loop region is 5-50nt; more preferably, it is 38nt.

[0018] Preferably, the 5' stem-loop linker has the sequence shown in SEQ ID NO.1.

[0019] In some embodiments, a molecular tag, such as N10, may be further provided in the loop region. The molecular tag may be provided at any position in the loop region.

[0020] Preferably, the 5' stem-loop linker connected to the molecular tag has the sequence shown in SEQ ID NO. 2 or SEQ ID NO. 3.

[0021] In another aspect, the present invention provides a method for preparing a single-stranded DNA sequencing library based on 5' end ligation, the method comprising the following steps of processing a sample to be tested:

[0022] Step 1) using a DNA polymerase without exonuclease activity to fill in the 5' end of the double-stranded DNA to obtain a treated first reaction product;

[0023] Step 2) performing 3'-end tailing on the first reaction product to obtain a second reaction product;

[0024] Step 3) ligating the second reaction product to the annealed 5' stem-loop linker as described above to obtain a ligation product;

[0025] Step 4) amplifying the ligation product. The amplified product constitutes a single-stranded DNA sequencing library.

[0026] Specifically, the absence of exonuclease activity refers to the absence of 5'→3' exonuclease activity and / or 3'→5' exonuclease activity.

[0027] Preferably, the amount of the DNA polymerase without exonuclease activity is not limited, and any concentration can fully fill the 5' overhanging end, for example, 0.25-1.5 U / μL.

[0028] Preferably, the DNA polymerase without exonuclease activity comprises KF enzyme (Klenow Fragment-3'→5'exo - ).

[0029] Specifically, the KF enzyme is an N-terminal truncated fragment of Escherichia coli DNA polymerase I, which does not have 5'→3' exonuclease activity or 3'→5' exonuclease activity.

[0030] Preferably, the tailing in step 2) includes adding a poly(dA) tail, or adding a poly(dT) tail, a poly(dC) tail, and a poly(dG) tail.

[0031] Preferably, terminal transferase and dNTP tailing are used.

[0032] Preferably, the reagents used for the tailing may be terminal transferase and dATP, thereby adding a poly (dA) tail.

[0033] Preferably, the terminal transferase is TdT. Specifically, TdT, whose Chinese name is Terminal Deoxynucleotidyl Transferase and English name is Terminal Deoxynucleotidyl Transferase, is a template-independent DNA polymerase that can catalyze the addition of dNTPs to the 3' hydroxyl end of oligonucleotides, single-stranded or double-stranded DNA.

[0034] Preferably, the molar ratio of dATP / DNA in the reaction system of step 2) can be adjusted according to the total amount of DNA in the system, for example, greater than 100 or 100-5000. Specific dATP / DNA ratios such as 100, 500, 1000, and 5000 all have good tailing effects.

[0035] Preferably, before the tailing is performed in step 2), the method further comprises the step of dephosphorylating or purifying the first reaction product obtained after the treatment in step 1).

[0036] Preferably, the dephosphorylation treatment is achieved using phosphatase such as recombinant shrimp alkaline phosphatase (rSAP), which can remove the influence of dNTP in the system.

[0037] Preferably, the DNA purification method is conventionally known in the art, for example, phenol chloroform and isoamyl alcohol are used for purification in the specific examples.

[0038] Preferably, the molar ratio of the 5' stem-loop linker to the DNA is greater than 12.5:1, more preferably greater than 50:1, and most preferably 50:1 or greater than 50:1. The DNA is the DNA obtained after the treatment in step 2).

[0039] Preferably, before the connection in step 3), the method further comprises the step of phosphorylating the second reaction product obtained after the treatment in step 2), which can improve the connection effect.

[0040] Preferably, T4 polynucleotide kinase is used for phosphorylation. Preferably, a purification step is further included after the phosphorylation.

[0041] Preferably, the 5' stem-loop linker ligation is achieved by T4 DNA ligase.

[0042] Preferably, the annealing in step 3) refers to the process by which the two complementary strands of the denatured DNA fully or partially return to their native double helical structure under appropriate conditions. After thermal denaturation of the DNA, the temperature is slowly lowered to gradually cool the DNA and maintain it within a certain range below the Tm value, allowing the denatured single-stranded DNA to regain its double helical structure. This annealing is also called annealing.

[0043] Preferably, before amplification in step 4), the step of purifying the single-stranded DNA ligation product is also included.

[0044] Preferably, when the tailing is to add a poly(dA) tail, Oligo d(T) can be used. 25 Purification by magnetic beads.

[0045] Preferably, primers containing an index sequence (Index) can be used in the amplification in step 4). By using primers containing an index sequence to amplify the ligation product, data from different sample sources during sequencing can be distinguished and the sample source can be identified.

[0046] Specifically, the "tag sequence" can be introduced into the 5' end or the 3' end of the ligation product in step 3). Methods for designing tag sequences are well known to those skilled in the art, as are designing primers containing tag sequences (indexes).

[0047] The "amplification" described in the present invention can be divided into two categories: variable temperature amplification and constant temperature amplification. Variable temperature amplification mainly includes the classic polymerase chain reaction (PCR) and ligase chain reaction (LCR), while constant temperature amplification includes strand displacement amplification (SDA), rolling circle amplification (RCA), loop mediated amplification (LAMP), helicase-dependent isothermal DNA amplification (HDA), nucleic acid sequence-based amplification (NASBA), and transcription-based amplification system (TAS). In the specific embodiment, PCR is used as an example for library construction.

[0048] On the other hand, the present invention provides a high-throughput sequencing method for a single-stranded DNA sequencing library prepared based on the method described above.

[0049] Preferably, the method comprises the steps of preparing a library according to the aforementioned method, and sequencing the sequencing library.

[0050] Preferably, the sequencing can be performed by a high-throughput sequencing platform including Illumina NovaSeq, HiSeq X Ten, Illumina HiSeq, Illumina MiSeq, PacBio Sequel, 10×Genomics and MGISEQ-2000.

[0051] In another aspect, the present invention provides a kit for preparing a library for high-throughput sequencing of single-stranded DNA, wherein the kit comprises the aforementioned 5' stem-loop linker.

[0052] Preferably, the kit further comprises a DNA polymerase without exonuclease activity.

[0053] Preferably, the DNA polymerase is selected from KF enzyme.

[0054] Preferably, the kit further comprises one or more of a dephosphorylation reagent, a tailing reagent, a DNA purification reagent, and an amplification reagent.

[0055] Preferably, the dephosphorylation agent is selected from dephosphorylases.

[0056] Preferably, the dephosphorylase is selected from SAP.

[0057] Preferably, the tailing reagent comprises terminal transferase and or nucleotides.

[0058] Preferably, the terminal transferase is selected from TdT.

[0059] Preferably, the nucleotide is selected from dATP.

[0060] Preferably, the amplification reagents include at least one forward primer and at least one reverse primer.

[0061] Preferably, the primer comprises a tag sequence.

[0062] Preferably, the reverse primer comprises a tag sequence.

[0063] Preferably, the amplification reagent includes three reverse primers.

[0064] Preferably, the sequences of the primers are shown as SEQ ID NO. 4-7 respectively.

[0065] On the other hand, the present invention provides a device or system for preparing a single-stranded DNA sequencing library based on 5' end ligation, the device comprising: a 5' protruding end polishing unit, a 3' end tailing unit, a 5' stem-loop junction ligation unit and an amplification unit.

[0066] Preferably, the polishing unit for the 5' protruding end is used to polish the 5' end of dsDNA using a DNA polymerase without exonuclease activity.

[0067] Preferably, the 3'-terminal tailing unit is used to add a tail to the 3'-terminal, preferably poly(dA).

[0068] Preferably, the connecting unit of the 5' stem-loop linker is used to connect a 5' stem-loop linker, and the structure of the 5' stem-loop linker is as described above.

[0069] Preferably, the length of the random base region is 6 nt, the length of the protruding end is 12 nt, the length of the first stem region or the second stem region is 11 nt, and the length of the loop region is 38 nt.

[0070] In the present invention, the term "DNA" as used herein may refer to any polymer comprising deoxyribonucleotides, including but not limited to modified or unmodified DNA. Those skilled in the art will appreciate that the source of genomic DNA is not particularly limited and may be obtained from any possible route, including direct commercial acquisition, direct acquisition from other laboratories, or direct extraction from a sample. According to an embodiment of the present invention, the single-stranded DNA molecule may be obtained by reverse transcription of RNA. According to one embodiment of the present invention, the single-stranded DNA molecule may be a cDNA molecule obtained by reverse transcription of RNA. According to an embodiment of the present invention, the single-stranded DNA molecule may be obtained by denaturing a double-stranded DNA sample. According to a specific embodiment of the present invention, the single-stranded DNA molecule may be obtained by heat denaturing a double-stranded DNA sample. The amount of the single-stranded DNA described herein is not particularly limited, and the length of the single-stranded DNA molecule described herein is not particularly limited. Single-stranded DNA greater than 20 nt is preferred, preferably 20-80 or 80-100 nt, specifically including 20, 40, 79, and 80 nt.

[0071] Alternatively, the single-stranded DNA molecule may be obtained by reverse transcription of RNA; alternatively, the single-stranded DNA molecule may be a cDNA molecule obtained by reverse transcription of RNA; alternatively, the single-stranded DNA molecule may be obtained by extraction from a target test sample. Preferably, the test sample includes peripheral blood, tissue, blood, serum, plasma, urine, saliva, semen, milk, cerebrospinal fluid, tears, sputum, mucus, lymph, cytosol, ascites, pleural effusion, amniotic fluid, bladder washing fluid and bronchoalveolar lavage fluid, etc., taken from animals (especially humans); the test sample may also be taken from bacterial culture, bacterial colonies, virus suspension, environmental concentrate, food, raw materials, water samples or water concentrates, etc. The test sample contains single-stranded DNA molecules and may also contain other non-target nucleic acids. The test sample contains at least 1 fmol of single-stranded DNA molecules. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] FIG1 is a schematic diagram showing a double-stranded DNA having a 5' overhanging end being filled into a blunt end.

[0073] FIG2 is a schematic diagram of the reaction principle of single-stranded DNA 5′ end-specific ligation mediated by a stem-loop linker.

[0074] FIG3 is a flow chart of the present invention's single-stranded DNA specific sequencing based on 5' end ligation.

[0075] FIG4 is a schematic diagram of the DNA structure and the result diagram for verifying that the 5' protruding end is filled, wherein lane 1 is a blank control and lane 7 is an equal amount of positive control.

[0076] FIG5 is a graph showing the results of verifying the tailing effect under different dATP / DNA conditions.

[0077] FIG6 is a diagram showing the results of verifying the connection effects under different linker / DNA ratios.

[0078] FIG7 is a diagram showing the verification results of the specific ligation between the 5' stem-loop linker and single-stranded DNA.

[0079] FIG8 is a schematic diagram of synthesizing single-stranded DNA and verifying the quality of the library prepared by the present invention by high-throughput sequencing.

[0080] FIG9 is a diagram showing the detection results of library quality.

[0081] FIG10 is a statistical result diagram of the sequencing data volume and alignment rate of samples to be tested at different gradients.

[0082] FIG11 is a graph showing the distribution analysis results of the number of reads for samples tested at different gradients.

[0083] FIG12 is a graph showing the quality of the cfDNA sample library preparation. DETAILED DESCRIPTION

[0084] The present invention will be further described below with reference to specific embodiments, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solutions and inventive concepts of the present invention, should be covered by the scope of protection of the present invention.

[0085] Unless otherwise specified, the materials and reagents used in the following examples can be obtained from commercial sources.

[0086] The present invention is based on the principle of 5' end-linked single-stranded DNA specific sequencing:

[0087] 1. Fill in the double-stranded DNA with 5' protruding ends to make them blunt ends

[0088] As shown in Figure 1 (left), there are three possible 5'-end configurations for double-stranded DNA: 3'-end overhang, blunt end, and 5'-end overhang. Due to the directional nature of DNA synthesis, double-stranded DNA with 5'-end overhangs can be trimmed to blunt ends. After trimming, the 5'-end configurations of the double-stranded DNA in the system are reduced to two (Figure 1 (right)), facilitating subsequent specific ligation.

[0089] Conventional DNA polymerases have exonuclease activity. In order to ensure the integrity of single-stranded DNA in the system during the filling process, we selected KF enzyme (Klenow fragment 3'→5'exo - , KF exo - KF enzyme is an N-terminal truncated fragment of Escherichia coli DNA polymerase I and does not contain 5'→3' exonuclease activity or 3'→5' exonuclease activity.

[0090] 2. Use 5' protruding stem-loop linkers to achieve specific ligation of single-stranded DNA

[0091] After patching, the 5'-end configurations of the double-stranded DNA in the sample are reduced to two. Therefore, we designed a stem-loop adapter with a 5' overhang. This steric hindrance between the stem-loop adapter's overhang and the two 5'-end configurations of the double-stranded DNA prevents the adapter from ligating to the double-stranded DNA, thereby specifically ligating the single-stranded DNA in the system.

[0092] The 5' stem-loop linker described in this invention contains a 5' terminal sequence, a random base (N6), a unique molecular identifier (UMI), and a 3' terminal sequence. The N6 is designed to allow for non-preferential pairing with the 5' end of single-stranded DNA, while the UMI provides unique identification and classification for each linker, thereby distinguishing individual sequences. Furthermore, to prevent self-ligation, both the 5' and 3' ends of the stem-loop linker are hydroxyl groups and remain unmodified (Figure 2).

[0093] Example 1: Feasibility of single-stranded DNA specific sequencing based on 5' end ligation

[0094] Step 1: dsDNA 5' end filling

[0095] First, the DNA sample is 5'-end filled to reduce the complexity of the 5' end of the dsDNA in the sample and prepare for subsequent specific ligation.

[0096] Table 1. Reaction system

[0097] Specifically, prepare the reaction system as shown in Table 1, where the amount of single-stranded DNA in the sample is not particularly limited. According to specific embodiments of the present invention, the amount of single-stranded DNA is ≥ 1 fmol, specifically ≥ 25 pg, thereby achieving high efficiency and accuracy in sequencing library construction. Mix thoroughly, centrifuge briefly, and incubate in a PCR instrument at 37°C for 30 minutes. After the reaction is completed, immediately store at 4°C or on ice.

[0098] To verify whether the 5' overhangs of double-stranded DNA can be filled by KF enzyme, we synthesized double-stranded DNA with 10-nt 5' overhangs ending in two A bases. Only when the 5' overhangs were filled could biotin-labeled dUTP be incorporated, allowing detection by streptavidin.

[0099] Specifically, we used 10 pmol of the above double-stranded DNA, added different gradients of KF enzyme (0.25-1.5 U / μL) for end filling, then performed urea denaturing gel electrophoresis, and used a chemiluminescence biotin-labeled nucleic acid detection kit for detection.

[0100] The experimental results showed that the 5' overhanging ends of double-stranded DNA could be fully filled (Figure 4), where lane 1 was a negative control without KF enzyme, and lane 7 was a positive control with an equal amount of biotin-labeled double-stranded DNA.

[0101] Step 2: Dephosphorylation

[0102] Since there will be excess dNTPs in the system after the previous step of the filling reaction, in order to avoid it affecting the subsequent DNA 3' end tailing reaction, it is necessary to use recombinant shrimp alkaline phosphatase (rSAP) to dephosphorylate it.

[0103] Table 2. Reaction system

[0104] Using the reaction system shown in Table 2, add rSAP to the reaction mixture from step 1. Mix thoroughly, centrifuge briefly, and incubate in a PCR instrument at 37°C for 60 minutes. Immediately store at 4°C or on ice after the reaction.

[0105] Step 3: DNA purification

[0106] The reaction product from step 2 was purified using phenol-chloroform-isoamyl alcohol (DNA extraction reagent) to remove the enzyme and buffer in the reaction system and prevent the loss of small fragments in the sample. Finally, 40 μL of ddH₂O was added to dissolve the reaction product for the subsequent tailing reaction.

[0107] Step 4: Adding poly(dA) to the 3' end

[0108] TdT (Terminal Transferase) is used to form a polyadenylated deoxynucleotide (Poly(dA)) tail (also referred to as a Poly-dA tail in this invention) at the 3' end of the reaction product. The Poly(dA) tail can be used as a 3' primer in subsequent steps and can also be used for specific enrichment and purification of the product.

[0109] Table 3, reaction system:

[0110] After adding the above reactants to the reaction system described in Table 3, the amount of dATP can be adjusted based on the DNA content in the sample. However, according to specific embodiments of the present invention, a dATP / DNA molar ratio between 100 and 5000 ensures sufficient poly(dA) tailing. After thorough mixing, centrifuge briefly and incubate in a PCR instrument at 37°C for 15 minutes. Immediately after completion of the reaction, store at 4°C or on ice.

[0111] To verify whether TdT can adequately add poly(dA) tails of appropriate length to single-stranded DNA, we synthesized 80-nt single-stranded DNA (ss80) and conducted a gradient experiment varying the dATP / DNA molar ratio. We used 5 pmol of ss80 and added varying amounts of dATP at molar ratios of 100:1, 500:1, 1000:1, and 5000:1. TdT was then added to initiate 3'-end tailing. The resulting poly(dA) tail length was determined by urea-denaturing gel electrophoresis.

[0112] The results showed that when the dATP / DNA molar ratio was greater than 100, TdT was able to adequately add a poly(dA) tail of appropriate length (greater than 20 nt) to single-stranded DNA. A poly(dA) tail longer than 20 nt facilitated subsequent magnetic bead purification and PCR amplification. The results are shown in Figure 5, where M indicates a marker.

[0113] Step 5: DNA 5' end phosphorylation

[0114] T4 DNA ligase-mediated ligation requires a 5' phosphate group on the target DNA. Therefore, prior to ligating the 5' end adapter, the reaction product must be phosphorylated at the 5' end using T4 Polynucleotide Kinase (T4 PNK).

[0115] Table 4. Reaction system

[0116] Prepare the above system, mix thoroughly, centrifuge briefly, and incubate at 37°C in a PCR instrument for 30 minutes. Immediately store at 4°C or on ice after the reaction.

[0117] Step 6: Column purification

[0118] The reaction product was purified and recovered using a GeneJET Purification Kit (Thermo Scientific K0702) according to the kit instructions to remove the enzyme and buffer in the reaction system. Finally, 40 μL of ddH2O was added to dissolve the reaction product for subsequent ligation reactions.

[0119] Step 7: Specific ligation mediated by 5' stem-loop linker

[0120] T4 DNA ligase is used to connect the purified and recovered product to the renatured 5' end adapter. After the adapter is renatured, it can form a stem-loop structure (containing an 11bp complementary pairing region (stem region) and a 38nt non-complementary pairing region (loop region) in the middle). The 5' end of the stem-loop adapter is an 18nt protruding end containing 6 random deoxynucleotides NNNNNN (N6), which can both complementarily pair with single-stranded DNA and prevent the adapter from connecting to double-stranded DNA. At the same time, to prevent the stem-loop adapter from connecting to each other, the 5' end of the adapter is a hydroxyl group instead of a phosphate group. In addition, the loop region can also contain 10 random deoxynucleotide sequences N10 as a molecular tag UMI (Unique molecular identifier) ​​to accurately distinguish DNA templates from different sources in the future (as shown in Figure 2).

[0121] The sequence of the adapter is shown in SEQ ID NO.1. Subsequent experiments were conducted using two sets of adapters connected with UMIs, UMI-LM and UMI-LR. The sequence of UMI-LM is shown in SEQ ID NO.2, and the sequence of UMI-LR is shown in SEQ ID NO.3. The difference between UMI-LM and UMI-LR lies in the location of the loop region where the UMI sequence is located. The UMI sequence of UMI-LM is located near the 5' end, while the UMI sequence of UMI-LR is located near the 3' end. Both UMI-LM and UMI-LR have high ligation efficiency. The sequence information is as follows:

[0122] (1) Refolding of the 5' end linker

[0123] Dilute the 5' stem-loop linker solution with ddH2O, add 10× annealing buffer (100 mM Tris-HCl (pH 8.0), 500 mM NaCl) during dilution, and then renature according to the renaturation procedure in the table below.

[0124] Table 5. Refolding procedure

[0125] (2) 5' end adapter ligation

[0126] Table 6. Ligation reaction system

[0127] Mix thoroughly and centrifuge briefly. Incubate in a PCR instrument at 16°C for 2 hours, then at 75°C for 20 minutes to inactivate the ligase. Immediately store at 4°C or on ice after the reaction.

[0128] Experimental Results 1. To determine the conditions for efficient ligation of adapters to target DNA, we experimented with the adapter / DNA molar ratio. We mixed the adapter UMI LM with biotinylated single-stranded DNA of varying lengths: ss20 (20 nt), ss40 (40 nt), and ss79 (79 nt) at varying molar ratios (12.5:1, 25:1, and 50:1) and then performed ligation reactions. After the reactions were completed, urea-polyacrylamide electrophoresis was performed and biotin detection was performed. The results showed that when the adapter-to-DNA molar ratio was 50:1, the adapter efficiently ligated to target DNA of varying lengths (Figure 6).

[0129] Experimental Results 2. We experimentally validated the specific ligation of single-stranded DNA (ssDNA) mediated by a 5' stem-loop linker. We designed and synthesized a double-stranded DNA (ds95) with a 3' overhang, a double-stranded DNA (ds80) with a 5' overhang, and a smaller single-stranded DNA (ss50) for specificity verification. To maintain consistency and the possibility of ligation with the stem-loop linker, the overhang of ds80 and ss50 consisted of six random deoxyribonucleic acid residues.

[0130] To verify ligation specificity, we mixed ss50 with ds95 and ds80, performed end-filling reactions, and then tested the specific ligation reaction. To test the specific ligation of stem-loop adapters to single-stranded DNA, we used biotinylated UMI LM adapters. The specificity of the adapter-ssDNA ligation was verified by measuring the size of the ligation products (the ligation product from the adapter to ss50 is the smallest, while non-specific ligation products (i.e., the ligation product from the adapter ds80 to ds95) are larger).

[0131] As shown in Figure 7, lane 1 is a control without ligase, and lane 7 is a blank control containing only the biotinylated linker. Lanes 2, 3, and 4 represent individual reactions with ds80, ds95, and ss50, respectively. No ligation products are observed in lanes 2 and 3, while ligation products between ss50 and the linker are observed in lane 4. Lanes 5 and 6 represent reactions with a 1:1 mixture of ss50, ds80, and ds95, respectively, and the linker. It can be seen that even in the mixed single- and double-stranded DNA system, the 5' stem-loop linker only ligates to ss50, with no nonspecific ligation products between the linker and ds80 or ds95. Furthermore, the amount of ligation product in lanes 4, 5, and 6 is virtually identical, indicating that the linker is fully ligated to the single-stranded DNA ss50 even in the mixed system.

[0132] The above results indicate that the 5' stem-loop linker can specifically and efficiently connect single-stranded DNA in the system, thereby enabling specific sequencing of single-stranded DNA by constructing a library and performing high-throughput sequencing according to conventional techniques.

[0133] Step 8: Oligo d(T)25 magnetic bead purification

[0134] Magnetic beads (Oligo d(T)25 Magnetic Beads, NEB S1419S) were used to purify and recover the target single-stranded DNA product that had been connected to the stem-loop adapter, while removing excess 5' end adapters.

[0135] The following operations were modified and optimized according to the magnetic beads instructions:

[0136] 1. Add 100 μL of Lysis / Binding Solution to a 200 μL tube, add 20 μL of magnetic bead suspension, vortex briefly, and incubate at room temperature for 2 minutes.

[0137] 2. Place the centrifuge tube on the magnetic rack, incubate for 2 minutes, and discard the supernatant.

[0138] 3. Add 20 μL of ddH2O to 80 μL of sample mixture, then add to the equilibrated magnetic beads and mix well.

[0139] 4. Place the centrifuge tube in the PCR instrument, heat the sample at 65℃ for 5 minutes, then quickly cool it to 4℃. Take out the centrifuge tube when the temperature reaches 4℃.

[0140] 5. Incubate at room temperature for 10 minutes.

[0141] 6. Place the centrifuge tube on the magnetic rack, incubate for 5 minutes, and discard the supernatant.

[0142] 7. Add 100 μL of Wash Buffer 1 (WB 1) to the centrifuge tube and mix well by pipetting.

[0143] 8. Place the centrifuge tube on the magnetic rack, incubate for 2 minutes, and discard the supernatant.

[0144] 9. Add 100 μL of Wash Buffer 2 (WB 2) to the centrifuge tube and mix well by pipetting.

[0145] 10. Place the centrifuge tube on the magnetic rack, incubate for 2 minutes, and discard the supernatant.

[0146] 11. Add 100 μL of low salt buffer (LSB) to the centrifuge tube and mix well by pipetting.

[0147] 12. Place the centrifuge tube on the magnetic rack, incubate for 2 minutes, and discard the supernatant.

[0148] 13. Add 36 μL of ddH2O and mix well by pipetting.

[0149] 14. Place the centrifuge tube in a PCR instrument, heat the sample at 95°C for 5 minutes, and then cool it to 25°C. When the temperature reaches 25°C, remove the centrifuge tube and place the microcentrifuge tube on the magnetic rack and incubate for 2 minutes.

[0150] 15. Transfer the supernatant to a new tube and use immediately or store at -20°C.

[0151] Step 9: PCR amplification

[0152] The single-stranded DNA in the product purified and recovered in the previous step contains both a 3' terminal Poly (dA) tail and a 5' terminal adapter. PCR amplification using this as a template can yield a specific single-stranded DNA library.

[0153] Table 7, primer sequences:

[0154] Table 8, reaction system:

[0155] Table 9. Reaction Procedure

[0156] Step 10: Clip Selection

[0157] Before sequencing, it is necessary to remove residual DNA polymerase, dNTP mixture, inorganic salts, and excess primers in the reaction system. Therefore, we used VAHTS DNA Clean beads (Vazyme, N411-01-AA) and followed the instructions to purify and recover the PCR product, ultimately obtaining a library that can be directly sequenced.

[0158] Example 2: High-throughput sequencing to verify sensitivity and stability

[0159] To confirm its feasibility, sensitivity, and stability, we randomly selected and synthesized 200 single-stranded DNA libraries of 60-100 nt from the Escherichia coli genome, then prepared them into different gradients, constructed the libraries, and sequenced them for analysis (see Figure 8 for a schematic diagram).

[0160] We constructed libraries of different single-stranded DNA libraries (1 fmol to 640 fmol) using our method and then sequenced them on the Nova-seq 6000. Based on the different initial DNA amounts used for library construction, we adjusted the PCR cycle numbers shown in Table 9 to balance the library DNA obtained in each gradient. From low to high initial DNA amounts, the PCR cycle numbers were 30, 20, 18, 16, and 14, with three replicates for each group.

[0161] Figure 9 shows the capillary electrophoresis results of the obtained library. The results show that this method can generate high-quality sequencing libraries.

[0162] The analysis results shown in Figure 10 demonstrate that the amount of clean reads obtained from libraries with different gradients is comparable and stable. Furthermore, when the filtered data were mapped to the genome, the mapping rates were also extremely high, with the lowest mapping rate exceeding 98.5% for the 1 fmol group. The mapping rates were also very stable across the different groups.

[0163] Example 3: High-throughput sequencing verifies lack of bias

[0164] In order to further analyze whether this method has obvious bias, we analyzed the distribution of the filtered data obtained in Example 2.

[0165] Figure 11 shows the distribution of read counts for 200 single-stranded DNA sequences from sequencing results at different gradients. The results show a similar "S"-shaped distribution of reads across all gradients (the integrated results are shown in the lower right corner). Furthermore, reads were detected for each single-stranded DNA sequence. These results demonstrate that this method has no apparent bias, is stable, and exhibits extremely high sensitivity.

[0166] Example 4: This method can be efficiently applied to prepare cfDNA sample libraries from cell culture media

[0167] To further analyze the efficiency and specificity of this method for detecting single-stranded DNA in real biological samples, we collected culture medium from non-small cell lung cancer A549 cells, extracted cfDNA samples, and constructed libraries according to the above process (steps 1 to 10). UMI-LM and UMI-LR adapters were used, respectively, and 18 PCR cycles were used in step 9. Capillary electrophoresis results demonstrated that both stem-loop adapter-based library construction methods used in this invention consistently produced high-quality libraries (Figure 12).

[0168] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should be considered to fall within the scope of protection of the present invention.

Claims

1. A 5' stem-loop linker, wherein the structure of the 5' stem-loop linker is: protruding end-random base region-first stem region-loop region-second stem region, wherein the first stem region and the second stem region form a double strand through annealing.

2. The 5' stem-loop linker according to claim 1, characterized in that The length of the random base region is 1-12 nt, the length of the protruding end is 0-20 nt, the length of the first stem region or the second stem region is 8-25 nt, and the length of the loop region is 0-50 nt; Preferably, the length of the random base region is 3-12 nt, the length of the first stem region and the second stem region are 8-25 nt respectively, and the length of the loop region is 5-50 nt; Preferably, the length of the random base region is 6 nt, the length of the protruding end is 12 nt, the length of the first stem region or the second stem region is 11 nt, and the length of the loop region is 38 nt; Preferably, the 5' stem-loop linker has the sequence shown in SEQ ID NO.1; Preferably, the loop region is further provided with a molecular tag; preferably, the molecular tag is N10; preferably, the 5' stem-loop linker has a sequence shown in SEQ ID NO.2 or SEQ ID NO.

3.

3. The 5' stem-loop linker according to claim 1, characterized in that The 5' end and / or the 3' end of the 5' stem-loop linker is a hydroxyl group; Preferably, the 5' end and the 3' end contain no modification.

4. A method for preparing a single-stranded DNA sequencing library based on 5' end ligation, the method comprising the following steps of processing a sample to be tested: Step 1) using a DNA polymerase without exonuclease activity to fill in the 5' end of the double-stranded DNA to obtain a treated first reaction product; Step 2) performing 3'-end tailing on the first reaction product to obtain a second reaction product; Step 3) ligating the second reaction product to the renatured 5' stem-loop linker according to any one of claims 1 to 3 to obtain a ligation product; Step 4) amplify the ligation product. Preferably, the DNA polymerase having no exonuclease activity is KF enzyme.

5. The method according to claim 4, characterized in that: Step 2) tailing using terminal transferase and dNTPs; Preferably, the molar ratio of dNTP to DNA is ≥100:1; Preferably, the terminal transferase is selected from TdT; the dNTP is selected from dATP; Preferably, before the tailing in step 2), the method further comprises the step of dephosphorylating or purifying the first reaction product obtained after the treatment in step 1); Preferably, the phosphorylation is performed using a phosphatase; Preferably, the phosphatase is selected from SAP.

6. The method according to claim 4, characterized in that: The molar ratio of the 5' stem-loop linker to the DNA obtained after treatment in step 2) is ≥50:1; Preferably, before the ligation in step 3), the method further comprises the step of phosphorylating the second reaction product obtained after the treatment in step 2); Preferably, phosphorylation is performed using T4 polynucleotide kinase.

7. The method according to claim 4, characterized in that Before amplification in step 4), the method further includes a step of purifying the single-stranded DNA ligation product; Preferably, magnetic beads are used to purify the single-stranded DNA ligation products.

8. A high-throughput sequencing method for a single-stranded DNA sequencing library prepared by the method according to any one of claims 4 to 7; Preferably, the sequencing can be performed by a high-throughput sequencing platform including Illumina NovaSeq, HiSeq X Ten, Illumina HiSeq, Illumina MiSeq, PacBio Sequel, 10×Genomics and MGISEQ-2000. 9 . A kit for preparing a library for high-throughput sequencing of single-stranded DNA, comprising the 5′ stem-loop linker according to any one of claims 1 to 3 .

10. The kit according to claim 9, characterized in that The kit also includes a DNA polymerase without exonuclease activity; Preferably, the DNA polymerase is selected from KF enzyme.

11. The kit according to claim 9 or 10, characterized in that The kit further comprises one or more of a dephosphorylation reagent, a tailing reagent, a DNA purification reagent, and an amplification reagent.

12. The kit according to claim 11, characterized in that The dephosphorylation agent is selected from dephosphorylase; Preferably, the dephosphorylase is selected from SAP.

13. The kit according to claim 11, characterized in that The tailing reagent includes terminal transferase and or nucleotides; Preferably, the terminal transferase is selected from TdT. Preferably, the nucleotide is selected from dATP.

14. The kit according to claim 11, characterized in that The amplification reagent includes at least one forward primer and at least one reverse primer; Preferably, the primer comprises a tag sequence; Preferably, the reverse primer comprises a tag sequence; Preferably, the amplification reagent comprises three reverse primers; Preferably, the sequences of the primers are shown as SEQ ID NO. 4-7 respectively.

15. A system or apparatus for preparing a single-stranded DNA sequencing library based on 5' end ligation, the system or apparatus comprising: 5' protruding end finishing unit, 3' end tailing unit, 5' stem-loop junction ligation unit and amplification unit; The 5' protruding end polishing unit is used to polish the 5' end of dsDNA using a DNA polymerase without exonuclease activity; The 3' end tailing unit is used to add tail to the 3' end; The connecting unit of the 5' stem-loop linker is used to connect the 5' stem-loop linker according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Construction method and application of single-stranded sequencing library

    CN109536579A

  • Construction method of single-stranded nucleic acid molecule sequencing library

    CN115197998A

  • Capturing and library building method based on single-chain connection and application

    CN116200478A

  • Specific high-throughput sequencing method for single-stranded DNA based on 5 'connection

    CN117701679A

  • Methods, Compositions, and Kits for Preparing Nucleic Acid Libraries

    US20200115736A1