Splicing type library building method based on complementary tail sequence annealing and collaborative extension connection reaction

By employing complementary tail-order annealing and synergistic extension ligation reactions, the problems of low well utilization and limited read length in short-fragment nucleic acid libraries in third-generation sequencing have been solved, achieving efficient multi-molecule splicing and platform compatibility, thereby improving sequencing quality and efficiency.

CN121931099APending Publication Date: 2026-04-28THE FIRST AFFILIATED HOSPITAL ZHEJIANG UNIV COLLEGE OF MEDICINE +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
THE FIRST AFFILIATED HOSPITAL ZHEJIANG UNIV COLLEGE OF MEDICINE
Filing Date
2025-12-31
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In third-generation sequencing, short nucleic acid libraries result in low well utilization, limited read length, complex operation, and high cost. Existing library preparation processes rely on specific primers or homologous sequences, making it difficult to achieve multi-molecule cascade splicing.

Method used

Complementary tail annealing and synergistic extension ligation reactions are employed to achieve multi-round continuous splicing of short fragments through multiple temperature cycles in the same reaction system. The synergistic action of DNA polymerase and DNA ligase is utilized to form continuous long concatenator molecules, which are compatible with third-generation sequencing adapters through end-modification primers.

Benefits of technology

It significantly improved average read length and sequencing well utilization, simplified the library preparation process, reduced costs, and improved base calling accuracy and read consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121931099A_ABST
    Figure CN121931099A_ABST
Patent Text Reader

Abstract

The invention discloses a splicing type library building method based on complementary tail sequence annealing and collaborative extension connection reaction. Dividing a nucleic acid library to be detected into two parts, and respectively adding complementary tail sequences to obtain two parts of reaction products; mixing the two reaction products after the complementation of the tail sequence, performing denaturation cooling treatment, performing specific annealing on the complementation tail sequence to form an inter-fragment bridging structure, and performing annealing-extension-connection in the solution after the reaction to form a continuous controller long-chain molecule so as to obtain a double-chain product; the method comprises the following steps: adding a primer with end repair, adding polymerase, ligase and a nucleotide substrate at the same time, carrying out annealing treatment to obtain a preliminary library, and finally carrying out magnetic bead purification treatment and linker connection to obtain a final spliced library. According to the method, the problems of limited read length and low hole utilization rate in the construction of the third-generation sequencing library are solved, the average read length and the sequencing hole utilization rate are remarkably improved, the library construction process is simplified, and the cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of nucleic acid sequencing library construction technology, and in particular to a splicing library construction method based on complementary tail annealing and cooperative extension ligation reaction. Background Technology

[0002] The advent of third-generation sequencing (TGS) has enabled researchers to obtain complete nucleic acid sequence information directly at the single-molecule level. This allows for comprehensive analysis of genome structure, detection of complex alternative splicing events, and identification of gene fusions, structural variations, and epigenetic modifications. Representative platforms include PacBio (SMRT sequencing) and Oxford Nanopore (nanopore sequencing). Their latest systems, which balance long read lengths with high accuracy, have been widely used in whole-genome sequencing, full-length transcriptome sequencing, and diversity population studies.

[0003] However, whether in genomic or transcriptomic libraries, when samples in short fragment form (such as fragmented genomic DNA, single-cell cDNA libraries, barcode batch libraries, degraded samples, or PCR products) are directly used for third-generation sequencing, common technical bottlenecks exist: the fragments are generally short, resulting in short sequencing well (or nanopore) occupancy time, low well utilization, and limited output per unit cost; molecules read from different wells or reaction sites are not repetitive, lacking single-molecule-level repetitive calibration information, affecting read consistency and base recall accuracy; to adapt to the adapter structure and stability requirements of long-read platforms, existing library preparation processes require multiple rounds of end-revision, ligation, and purification, which are complex, time-consuming, and costly in terms of reagents; traditional splicing strategies usually rely on preset primer sequences or homology recognition regions, limiting the ligation sites and splicing times, making it difficult to achieve multi-molecule cascade splicing, thus limiting further improvements in average read length and sequencing well utilization.

[0004] Therefore, there is an urgent need for a novel library construction method that can achieve multi-molecule continuous splicing, end-modification, and platform compatibility through complementary sequence annealing and enzymatic synergistic reactions without relying on specific primers or homologous sequences, so as to significantly improve read length, sequencing well utilization, and overall data quality at the genomic and transcriptomic levels. Summary of the Invention

[0005] To address the problems existing in the background technology, mainly targeting the issues of limited read length and low well utilization in the construction of third-generation sequencing libraries, this invention aims to provide a method for splicing and constructing libraries applicable to multiple types of nucleic acid libraries, including genomics and transcriptomics. Through complementary tail-driven spontaneous annealing and synergistic extension ligation using DNA polymerase and DNA ligase, multiple rounds of continuous splicing of short fragments are achieved in the same reaction system. Platform compatibility is achieved through end-modification primer extension, thereby significantly improving the average read length and sequencing well utilization, simplifying the library construction process, and reducing costs.

[0006] The method of the present invention includes at least the following steps: 1) Document partitioning and end-order bonus: The nucleic acid library to be tested was divided into two portions and complementary tail sequences were added to each portion to obtain two reaction products; 2) Mixed denaturation and complementary annealing: The two reaction products after the complementary tail sequence were mixed, denatured and cooled, and then the complementary tail sequence was specifically annealed to form a bridging structure between fragments to obtain the post-reaction solution. 3) Collaborative Extension and Connection: After the reaction, the bridging nucleic acid fragments are continuously linked by multiple cycles of annealing, extension and ligation in the solution to form continuous concatenator long chain molecules and obtain double-stranded products. Specifically, by setting a program in the temperature control instrument, the reaction system is automatically subjected to multiple temperature cycles between the annealing temperature (4–8℃) and the extension / bonding temperature (approximately 37℃), thereby achieving multiple annealing-extension-bonding processes.

[0007] 4) End-modified primer extension and platform compatibility: To the double-stranded product, end-banding primers were added, and DNA polymerase, DNA ligase, and nucleotide substrate were added simultaneously for annealing to obtain a preliminary library. 5) Purification and adapter connection: The preliminary library was purified by magnetic beads and connected with adapters to obtain the final assembled library.

[0008] Step 1) specifically involves dividing the nucleic acid library to be tested into two portions and adding complementary tail sequences to the 3' ends of each portion. The complementary tail sequences are either a complementary combination of a poly(A) tail sequence and a poly(T) tail sequence, or a complementary combination of a poly(G) tail sequence and a poly(C) tail sequence. Preferably, one portion is given a poly(A) tail and the other a poly(T) tail; alternatively, a poly(G) / poly(C) tail can be used, with one portion given a poly(G) tail and the other a poly(C) tail; or other complementary tail sequences.

[0009] The complementary tail addition is performed using terminal deoxyribonucleotidyl transferase TdT or an equivalent tailing method, with a tail length of 10-50 nt, at a reaction temperature of 37℃ and a reaction time of 10-30 min. The nucleic acid library to be tested is a DNA library or a cDNA library.

[0010] The type and length of the complementary tail sequence are adjusted within the range of 6-80 nt to control annealing stability and splicing depth.

[0011] Step 2) is an annealing step, specifically: the two reaction products are mixed, denatured at 95°C for 2-3 minutes, and then slowly cooled to 8-4°C, so that the complementary tail sequence of the two reaction products is specifically linked to form an inter-fragment bridging structure.

[0012] Step 3) specifically involves simultaneously adding DNA polymerase and DNA ligase to the post-reaction solution of the same reaction system, along with nucleotide substrate, and reacting at 37°C for 30 min–2 h to perform a synergistic extension and ligation reaction. This process allows the annealing region to be completed first by DNA polymerase extension. The annealing region refers to the locally double-stranded region formed by the complementary tail annealing of the two reaction products in step 2, and then covalently linked by phosphodiester bonds catalyzed by DNA ligase. The products after extension and ligation form new continuous DNA molecules, which still contain complementary tail structures. Under temperature-controlled conditions, by adjusting the reaction system from the extension / ligation temperature to a temperature range suitable for complementary tail annealing, the annealing region can be formed again.

[0013] Therefore, the annealing-extension-linking process can be repeated in the same reaction system to achieve multiple cycles and gradually form continuous long-chain concatenator molecules.

[0014] Preferably, the DNA ligase is T4 or T7 DNA Ligase, and the nucleotide substrate is dATP or dTTP.

[0015] The double-stranded product obtained after synergistic extension ligation still retains incomplete sticky tails at both ends. Step 4) specifically involves: adding end-repair primers with a 3' end of 15dT and end-repair primers with a 3' end of 15dA sequentially to the double-stranded product, and simultaneously adding DNA polymerase, DNA ligase, and nucleotide substrate to the same system, and reacting at 37°C for 10-30 min each to carry out synergistic extension ligation reaction, achieving primer annealing-extension-ligation at both ends, and obtaining a preliminary library compatible with the adapter structure of third-generation sequencing.

[0016] Before adding DNA polymerase, DNA ligase and nucleotide substrate, an optional annealing step can be performed: denature at 95°C for 2-3 minutes and then slowly cool to 8-4°C to allow the double-stranded product to be specifically ligated with the complementary tail sequence of the end-modified primer to form an inter-segment bridging structure.

[0017] The combinations of the 3' end-modified primer at 15dT and the 3' end-modified primer at 15dA are respectively SEQ ID No. 1 and SEQ ID No. 2, and SEQ ID No. 3 and SEQ ID No. 4.

[0018] Step 5) specifically involves: after purifying the product of the preliminary library with magnetic beads to remove free primers and short fragments, ligating the product with adapters using the Nanopore or PacBio platform to finally obtain a spliced ​​library that can be directly used for third-generation sequencing.

[0019] The 5' end of the modified primer integrates a sequencing primer binding site, a barcode, or a directional marker. The sequencing primer binding site is specifically a primer binding sequence common to third-generation sequencing platforms, including but not limited to sequencing primer binding sites required by Nanopore or PacBio platforms. The barcode is specifically a specific nucleotide sequence used for sample differentiation, with a length of 4–24 nt. The directional marker is specifically a nucleotide sequence or structural marker used to differentiate the direction of the inserted fragment.

[0020] The method can be packaged into a kit containing a tailing enzyme, DNA polymerase, DNA ligase, end-repair primers, nucleotides, and a buffer system. The buffer system refers to a reaction buffer suitable for DNA annealing, extension, and ligation reactions. The buffer system includes at least one of a pH buffer, divalent metal ions, salt ions, and a stabilizer, preferably Mg²⁺. 2+ or Mn 2+ .

[0021] The method of this invention is applicable to various types of nucleic acid libraries, including but not limited to: genomic DNA libraries, transcriptome cDNA libraries, amplified product libraries, barcode mixed libraries, etc., and is compatible with third-generation sequencing platforms (such as Oxford Nanopore and PacBioSMRT).

[0022] The beneficial effects of this invention are: This invention utilizes complementary tail sequence programmable annealing and synergistic enzyme reactions to achieve multi-round continuous cascade splicing of short nucleic acid fragments in solution without relying on specific primer sites or homologous arms. The number of splices and the final read length can be controlled by continuous variables such as tail sequence length, reaction time, and enzyme amount, theoretically possessing the potential for "infinite connection" expansion. Through the directional participation of end-modified primers, the resulting long chains can be rapidly compatible with third-generation sequencing adapter structures.

[0023] The method of this invention significantly improves the average read length and well occupancy time, thereby increasing well utilization and output per unit cost. Introducing multi-fragment information within a single long read facilitates single-molecule consistency calibration, improves base calling accuracy and downstream assembly robustness, and reduces translocation and contamination risks by using a single-tube sequential reaction method, thus lowering operational complexity and reagent costs. The method also exhibits good versatility and platform compatibility for both genomic and transcriptomic libraries. Attached Figure Description

[0024] Figure 1 This is a schematic diagram illustrating the principle of the splicing library construction method based on complementary tail-order annealing and cooperative extension linkage reaction of the present invention. Figure 2 This is a graph showing the fragment length distribution of the 293T cell single-cell transcriptome cDNA library in Example 1 before splicing and after splicing using the method of this invention. Figure 3 This is a graph showing the statistical results of the read length distribution after sequencing the full-length transcript cDNA spliced ​​library of mouse brain tissue on the Nanopore platform in Example 2. Figure 4 This is an agarose gel electrophoresis analysis result of the DNA fragment splicing effect under different annealing-extension-ligation cycle numbers in Example 3; Figure 5 This is a graph showing the statistical results of the read length distribution of the full-length transcript cDNA library of mouse brain tissue that was not spliced ​​in Comparative Example 2 after sequencing on the Nanopore platform. Figure 6 This is a graph showing the fragment length distribution of the 293T cell single-cell transcriptome cDNA library constructed using the USER enzymatic method in Comparative Example 3. Detailed Implementation

[0025] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0026] The embodiments of the present invention are as follows: Example 1: Assembly and library construction of a 293T cell single-cell transcriptome cDNA library This embodiment uses a single-cell transcriptome cDNA library derived from the human embryonic kidney cell line 293T (HEK293T) as the starting sample to illustrate the specific application of the splicing library construction method based on complementary tail annealing and co-extension ligation reaction of the present invention in real biological samples.

[0027] (1) TdT tailing reaction The obtained double-stranded cDNA library was divided into two aliquots, each approximately 250 ng. Terminal transferase TdT (5–50 U, preferably 10–20 U per reaction) and its appropriate buffer were added to each aliquot. One aliquot was treated with dATP, and the other with dTTP, both at a final concentration of 0.2 mM. The mixture was incubated at 37 °C for 10–30 min to induce poly(A) and poly(T) tails at the 3′ ends of the cDNA molecules, respectively. After the reaction, the aliquots were incubated at 75 °C for 5–10 min to inactivate TdT. The tailing length can be controlled within the range of 10–50 nt by adjusting the reaction time and substrate concentration, thereby obtaining two types of tailed products with complementary tails.

[0028] (2) Complementary tail-order annealing The two tailing reaction products were mixed, and the poly(A) tail and poly(T) tail were brought into contact in the same reaction system. The mixture was denatured at 95 °C for 2–3 min to fully unwind the double-stranded structure. Subsequently, the reaction system was slowly cooled to 8–4 °C at a rate of 0.5 °C / s using a preset cooling program. This allowed the complementary tails to undergo specific base pairing, forming a bridging structure between the fragments and generating local double-stranded annealing regions, providing reaction sites for subsequent extension and linkage reactions.

[0029] (3) Synergistic extension connection reaction In the same reaction system after the above annealing reaction, DNA polymerase (1–10 U, preferably 2–5 U) and DNA ligase (50–400 U, preferably 100–200 U) are added, along with dATP and dTTP (both at a final concentration of 0.2 mM). The reaction is carried out at 37°C for 30 min–2 h, allowing the annealed region formed by complementary tail annealing to be extended and completed by DNA polymerase, followed by the formation of phosphodiester bonds between adjacent DNA fragments catalyzed by DNA ligase, thus achieving covalent ligation between the fragments. The products after extension and ligation form new continuous DNA molecules, which still contain tail structures capable of complementary pairing. By alternating the reaction system between 37°C and 8–4°C for 5 cycles under temperature control, new annealed regions can be formed multiple times and gradually extended and ligated, ultimately forming continuous concatenator long-chain molecules. After the reaction, the reaction is carried out at 75°C for 5–10 min to inactivate DNA polymerase and DNA ligase. In this embodiment, the DNA polymerase is Bst DNA polymerase, and the DNA ligase is T4 DNA Ligase.

[0030] (4) End-modified primer extension-ligation In the concatenator double-stranded product obtained after the co-extension ligation reaction, end-repair primers with 3′ 15dT and 3′ 15dA ends were added sequentially. After denaturation at 95 °C for 2–3 min, the mixture was cooled to 8–4 °C using a pre-programmed cooling procedure, allowing the end-repair primers to undergo complementary annealing with the incomplete sticky tails remaining at both ends of the double-stranded product. Subsequently, DNA polymerase (1–10 U, preferably 2–5 U) and DNA ligase (50–400 U, preferably 100–200 U) were added to the same reaction system, along with dATP and dTTP (both at a final concentration of 0.2 mM). The mixture was reacted at 37 °C for 10–30 min to complete the extension and covalent ligation of both ends of the double-stranded product, thereby obtaining a structurally stable preliminary spliced ​​library compatible with third-generation sequencing adapters. In this embodiment, the end-modification primer sequence with 3′ end 15dT is NB01-F (SEQ ID No. 1): 5′-CACAAAGACACCGACAACTTTCTTTTTTTTTTTTTTTTTTT-3′, and the end-modification primer sequence with 3′ end 15dA is NB01-R SEQ ID No. 1 (SEQ ID No. 2): 5′-AAGAAAGTTGTCGGTGTCTTTGTGAAAAAAAAAAAAAAA-3′.

[0031] (5) Quality inspection of spliced ​​document library The fragment length distribution of the 293T cell single-cell transcriptome cDNA splicing library constructed using the method described in this embodiment was detected by capillary electrophoresis analysis, and the results are as follows: Figure 2 As shown, the main peak of the unassembled library was concentrated in the approximately 500 bp region, while the library processed by the method of this invention showed a clear distribution characteristic of migrating towards longer fragments, with the main peak concentrated in the approximately 15 kb region. Long fragment signals of tens of thousands of bases in length could also be observed, indicating that the method of this invention can effectively promote the splicing of cDNA fragments and the formation of long concatenator molecules, significantly increasing the average fragment length of third-generation sequencing libraries.

[0032] Example 2: Assembly and library construction of a full-length cDNA library from mouse brain tissue This embodiment uses a full-length transcript cDNA library derived from mouse brain tissue as the starting sample to illustrate the specific application of the splicing library construction method based on complementary tail annealing and co-extension ligation reaction of the present invention in samples from complex tissue sources.

[0033] (1) TdT tailing reaction The obtained double-stranded cDNA library was divided into two aliquots, each approximately 250 ng. Terminal transferase TdT (5–50 U, preferably 10–20 U per reaction) and its appropriate buffer were added to each aliquot. One aliquot was treated with dATP, and the other with dTTP, both at a final concentration of 0.2 mM. The mixtures were reacted at 37 °C for 10–30 min to induce poly(A) and poly(T) tails at the 3′ ends of the cDNA molecules, respectively. After the reaction, the aliquots were inactivated at 75 °C for 5–10 min. Subsequently, to remove free enzyme and unreacted nucleotides, both tailing products were purified using magnetic beads, with the volume of magnetic beads being 1.5 times the volume of the reaction mixture. The purified tailing products were then eluted with 50 µL of pure water.

[0034] (2) Complementary tail-order annealing The two purified tailing reaction products were mixed, and the poly(A) tail and poly(T) tail were brought into contact in the same reaction system. The mixture was denatured at 95 °C for 2–3 min to fully unwind the double-stranded structure. Subsequently, the reaction system was slowly cooled to 8–4 °C at a rate of 0.5 °C / s using a preset cooling program. This allowed the complementary tails to undergo specific base pairing, forming a bridging structure between the fragments and generating local double-stranded annealing regions, providing reaction sites for subsequent extension and ligation reactions.

[0035] (3) Synergistic extension connection reaction In the same reaction system after the above annealing reaction, DNA polymerase (1–10 U, preferably 2–5 U) and DNA ligase (50–400 U, preferably 100–200 U) are added, along with dATP and dTTP (both at a final concentration of 0.2 mM). The reaction is carried out at 37°C for 30 min–2 h, allowing the annealed region formed by complementary tail annealing to be extended and completed by DNA polymerase, followed by the formation of phosphodiester bonds between adjacent DNA fragments by DNA ligase, thus achieving covalent ligation between the fragments. The products after extension and ligation form new continuous DNA molecules, which still contain complementary tail structures. By alternating the reaction system between 37°C and 8–4°C for a total of 8 cycles under temperature control, new annealed regions can be formed multiple times and gradually extended and ligated, ultimately forming continuous concatenator long-chain molecules. After the reaction, the reaction is carried out at 75°C for 5–10 min to inactivate DNA polymerase and DNA ligase. The obtained concatenator double-stranded product was then purified using magnetic beads, with the volume of the magnetic beads being 1.5 times the volume of the reaction system. Finally, it was eluted with 50 µL of pure water to obtain the purified spliced ​​product. In this embodiment, Bst DNA polymerase was used as the DNA polymerase, and T4 DNA ligase was used as the DNA ligase.

[0036] (4) End-modified primer extension-ligation In the concatenator double-stranded product obtained after co-extension ligation and purification, end-repair primers with a 3′ end of 15 dT and 3′ end of 15 dA were added sequentially. After denaturation at 95 °C for 2–3 min, the mixture was cooled to 8–4 °C using a pre-programmed cooling procedure, allowing the end-repair primers to undergo complementary annealing with the incomplete sticky tails remaining at both ends of the double-stranded product. Subsequently, DNA polymerase (1–10 U, preferably 2–5 U) and DNA ligase (50–400 U, preferably 100–200 U) were added to the same reaction system, along with dATP and dTTP (both at a final concentration of 0.2 mM). The reaction was carried out at 37 °C for 10–30 min to complete the extension and covalent ligation of the two ends of the double-stranded product. After the reaction, the DNA polymerase and DNA ligase were inactivated at 75 °C for 5–10 min. The reaction product was then purified using magnetic beads, with the volume of the magnetic beads being 1.5 times the volume of the reaction system. Finally, the product was eluted with 50 µL of pure water to obtain a preliminary assembled library with a stable structure and compatibility with third-generation sequencing adapters. In this embodiment, the primer sequence with a 3′ 15dT end is NB02-F (SEQ ID No. 3): 5′-ACAGACGACTACAAACGGAATCGATTTTTTTTTTTTTTT-3′, and the primer sequence with a 3′ 15dA end is NB02-R (SEQ ID No. 4): 5′-TCGATTCCGTTTGTAGTCGTCTGTAAAAAAAAAAAAAAA-3′.

[0037] (5) Motor protein adapter connection The assembled library after primer end-modification was ligated with motor protein adapters according to the Nanopore platform library construction procedure. Native adapters, ligation buffer, and fast DNA ligase were added to a 50 µL reaction system, and the reaction was carried out at room temperature for 20 min to complete adapter ligation. After the reaction, unligated adapters were removed by magnetic bead purification, followed by washing with long fragment buffer, and finally elution with 50 µL elution buffer to obtain the sequencing-ready library product. In this embodiment, the Nanopore #EXP-NBA114 kit was used to ligate motor protein adapters into the primer-end-modified product.

[0038] (6) Sequencing data quality control The full-length mouse brain tissue transcript cDNA splicing library constructed using the method described in this embodiment was sequenced on the Nanopore platform. Length statistical analysis was performed on the obtained sequencing data, and the read length distribution results are as follows: Figure 3As shown in the figure. The results show that the sequencing reads of the assembled library exhibit a clear long fragment distribution characteristic, with the main peak of the length distribution concentrated in the approximately 20 kb region, indicating that the assembled library constructed by the method of this invention can obtain long read data suitable for third-generation sequencing.

[0039] In summary, both Examples 1 and 2 are based on a splicing library construction technique using complementary tail annealing and synergistic extension ligation reactions. By performing end-capping of double-stranded cDNA fragments, complementary tail annealing, and synergistic extension and ligation reactions in the same reaction system, continuous covalent splicing of multiple cDNA fragments is achieved, forming long concatenator molecules. This method can repeatedly execute the annealing-extension-ligation process under temperature-controlled conditions, effectively extending the length of the library molecules and obtaining long-read libraries suitable for third-generation sequencing platforms. Experimental results show that the method of this invention has advantages such as simple operation process, low cost of required raw materials (enzymes, primers, etc.), convenient acquisition, high splicing efficiency, and stable library structure. It is applicable to transcriptome cDNA samples from different sources and types, and can stably obtain long-fragment libraries with a length of 10,000 base pairs, providing high-quality input materials for third-generation sequencing. While maintaining the same overall technical route, Examples 1 and 2 made corresponding adjustments to the specific process flow according to the different sample sources and complexity. In Example 1, a single-cell transcriptome cDNA library derived from 293T cells was used as the starting sample. The sample composition was relatively homogeneous, and no additional purification operations were introduced after each key step in the reaction process. This example primarily served to verify the feasibility and splicing effect of the method of the present invention in the construction of single-cell transcriptome libraries. In contrast, Example 2 used a full-length transcriptome cDNA library derived from mouse brain tissue as the starting sample. This sample had higher complexity and a wider range of fragment lengths. Therefore, magnetic bead purification was introduced after steps such as tailing, co-extension ligation, and primer end-revision reactions to further remove free primers, enzymes, and short fragments, thereby improving the stability of subsequent splicing and sequencing processes. These differences demonstrate that the method of the present invention can flexibly adjust process conditions according to different sample types, exhibiting good adaptability and versatility.

[0040] Example 3: Comparison of DNA fragment splicing effects under different extension and ligation cycles This embodiment uses Escherichia coli (E. coli) Escherichia coliTwo DNA fragments, specifically amplified from the genome, were used as starting samples: one approximately 750 bp in length and the other approximately 1800 bp in length. These two DNA fragments were treated as independent samples for repeated experiments. For the approximately 750 bp DNA fragment, dATP and dTTP were added during the reaction; for the approximately 1800 bp DNA fragment, dGTP and dCTP were added.

[0041] For each independent sample, the following reaction grouping conditions were set: (1) initial fragment group without any treatment; (2) extension-only group with only DNA polymerase added for extension reaction; (3) extension + ligation (1 cycle) group with DNA polymerase and DNA ligase added and annealed for 1 cycle; (4) extension + ligation (5 cycles) group with DNA polymerase and DNA ligase added and annealed for 5 cycles; (5) extension + ligation (10 cycles) group with DNA polymerase and DNA ligase added and annealed for 10 cycles. After each independent sample completed the above reaction, the obtained DNA product was analyzed by agarose gel electrophoresis to assess the trend of DNA fragment length changes under different cycle numbers.

[0042] The fragment size distribution of DNA products obtained from each independent sample under different reaction conditions is as follows: Figure 4 As shown. For the starting DNA fragment sample with a length of approximately 750 bp, the fragment length distribution of the starting fragment group and the extension-only group was basically the same, and no obvious long fragment signal was observed. In the extension + ligation (1 cycle) group, splicing products with a length higher than the starting fragment began to appear, with a length of approximately 2–3 times that of the starting fragment. In the extension + ligation (5 cycles) group and the extension + ligation (10 cycles) group, the fragment length was further increased, forming longer DNA splicing products.

[0043] For starting DNA fragments of approximately 1800 bp, the same trend was observed: extension alone was insufficient to significantly alter the fragment length distribution. However, with the introduction of extension and ligation and a gradual increase in the number of cycles, the length of the spliced ​​products showed a gradual increasing trend, resulting in longer DNA fragments. Under both conditions, the added dATP+dTTP and dGTP+dCTP could be used as nucleotide substrates in the extension and ligation reactions, respectively, and the resulting fragment lengths showed consistent trends.

[0044] Comparative Example 1: Direct third-generation sequencing results of an unspliced ​​library of full-length transcript cDNA from mouse brain tissue This comparative example uses a full-length transcript cDNA library derived from mouse brain tissue, as the starting sample, in Example 2. The library was constructed directly using third-generation sequencing without splicing to evaluate the sequencing results of the library without splicing.

[0045] Specifically, the full-length transcript cDNA library was constructed using the Nanopore Amplification-Free Barcoding Kit-24 V14 (SQK-NBD114.24). First, the cDNA was repaired and a dA tail was added using NEBNext FFPE DNA Repair Mixture and the NEBNext End Repair / dA Tail Addition Module. Then, the dT end of the barcode adapter was ligated to the dA tail of the cDNA molecule. After obtaining cDNA samples with different barcode sequences, the samples were mixed, and the sticky ends of the barcode adapter were further ligated to the third-generation sequencing adapter to obtain an unspliced ​​sequencing library.

[0046] The unspliced ​​cDNA library of full-length transcripts from mouse brain tissue constructed using the above comparative method was sequenced on the Nanopore platform, and the length statistical analysis of the obtained sequencing data was performed. The results of the read length distribution are as follows: Figure 5 As shown in the figure, the sequencing reads of the unassembled library were mainly concentrated in the shorter fragment range, with the main peak of the length distribution concentrated in the approximately 1 kb region, indicating that the average read length of the obtained library was limited without splicing.

[0047] Comparative Example 3: 293T cell single-cell transcriptome cDNA library was constructed using the USER enzymatic method for tandem library preparation. The uracil-specific excision reagent (USER) consists of uracil DNA glycosylase (UDG) and endonuclease VIII (EndoVIII), which can recognize and excise uracil bases in DNA molecules, thereby creating gaps or sticky ends at predetermined locations.

[0048] This comparative example uses the 293T cell single-cell transcriptome cDNA library used in Example 1 as the starting sample. Referring to the USER enzyme-based cDNA tandem library construction method (High-throughput RNAisoform sequencing using programmed cDNA concatenation) reported by Al'Khafaji et al., based on the known sequences of the amplification adapters in the single-cell transcriptome cDNA library, five pairs of dUTP-containing PCR primers were designed to amplify the cDNA library, introducing uracil bases at predetermined positions in the amplification products. Subsequently, a USER enzyme mixture was added to the reaction system to cleave the dUTPs, forming ligable sticky ends at both ends of the cDNA molecules. Multiple cDNA fragments were then ligated using DNA ligase to construct a USER enzyme-based tandem library.

[0049] The fragment length distribution of the tandem library constructed using the USER enzymatic method described above was detected by capillary electrophoresis analysis, and the results are as follows: Figure 6 As shown, the main peak of the single-cell transcriptome cDNA library fragments before splicing was concentrated in the approximately 500 bp region; after USER enzymatic tandem processing, the overall distribution of library fragment lengths shifted towards longer fragments, and no obvious single main peak was observed. Furthermore, a distribution feature showing a several-fold increase in library fragment length relative to the starting library was detected, with an increase of approximately 5-fold. This increase corresponds to the number of 5 pairs of PCR primers containing dUTP used in this comparative example.

[0050] Further analysis revealed that the USER enzymatic tandem library construction strategy relies on amplification adapters with known sequences in the starting library and requires the design of multiple pairs of dUTP-containing PCR primers for different library samples. The degree of tandem assembly is limited by the number of primers designed and the number of preset cleavage sites. In contrast, the method of this invention does not rely on specific sequence sites. By introducing polynucleotide tails into any double-stranded DNA library and performing controlled annealing-extension-ligation cycles in the same reaction system, continuous splicing of cDNA fragments can be achieved through multiple cycles under temperature-controlled conditions, thus demonstrating a different technical approach in terms of process flexibility and scalability.

[0051] The above specific embodiments are used to explain and illustrate the present invention, but not to limit the present invention. Any modifications and changes made to the present invention within the spirit and scope of the claims shall fall within the protection scope of the present invention.

[0052] The above description is only a preferred embodiment of the present invention. Therefore, all equivalent changes or modifications made to the structure, features and principles described in the claims of this patent application are included in the scope of this patent application.

[0053] The gene sequence involved in this invention is as follows: SEQ ID No.1; Name: NB01-F, an end-modified primer with 15dT at the 3' end DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct CACAAAGACACCGACAACTTTCTTTTTTTTTTTTTTTTT SEQ ID No.2; Name: NB01-R, a 3' end-modified primer with 15dA extension. DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct AAGAAAGTTGTCGGTGTCTTTGTGAAAAAAAAAAAAAAA SEQ ID No. 3; Name: NB02-F, a primer with 15dT at the 3' end DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct ACAGACGACTACAAACGGAATCGATTTTTTTTTTTTTTTT SEQ ID No.4; Name: NB02-R, a 3' end-modified primer with 15dA extension. DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct TCGATTCCGTTTGTAGTCGTCTGTAAAAAAAAAAAAAAA.

Claims

1. A splicing-type library construction method based on complementary tail-order annealing and cooperative extension linkage reaction, characterized in that: 1) Document partitioning and end-order bonus: The nucleic acid library to be tested was divided into two portions and complementary tail sequences were added to each portion to obtain two reaction products. 2) Mixed denaturation and complementary annealing: The two reaction products after the complementary tail sequence were mixed, denatured and cooled, and then the complementary tail sequence was specifically annealed to form a bridging structure between fragments to obtain the post-reaction solution. 3) Collaborative Extension and Connection: After the reaction, the bridging nucleic acid fragments are continuously linked by multiple cycles of annealing, extension and ligation in the solution to form continuous concatenator long chain molecules and obtain double-stranded products. 4) End-modified primer extension and platform compatibility: To the double-stranded product, end-banding primers were added, and DNA polymerase, DNA ligase, and nucleotide substrate were added simultaneously for annealing to obtain a preliminary library. 5) Purification and adapter connection: The preliminary library was purified by magnetic beads and connected with adapters to obtain the final assembled library.

2. The splicing-type library construction method based on complementary tail-order annealing and cooperative extension linkage reaction according to claim 1, characterized in that: Step 1) specifically involves dividing the nucleic acid library to be tested into two parts and adding complementary tail sequences to the 3' ends of each part. The complementary tail sequences are either complementary combinations of poly(A) tail sequences and poly(T) tail sequences or complementary combinations of poly(G) tail sequences and poly(C) tail sequences.

3. The splicing-type library construction method based on complementary tail-order annealing and cooperative extension linkage reaction according to claim 2, characterized in that: The complementary tail addition is performed using terminal deoxyribonucleotidyl transferase TdT or an equivalent tailing method, with a tail length of 10-50 nt, a reaction temperature of 37℃, and a time of 10-30 min.

4. The splicing library construction method based on complementary tail-order annealing and cooperative extension linkage reaction according to claim 1, characterized in that: The length of the complementary tail sequence is adjusted within the range of 6-80 nt.

5. The splicing-type library construction method based on complementary tail-order annealing and cooperative extension linkage reaction according to claim 1, characterized in that: Step 2) specifically involves mixing the two reaction products, denaturing them at 95°C for 2-3 minutes, and then cooling them to 8-4°C to allow the complementary tail sequence of the two reaction products to specifically connect and form an inter-segment bridging structure.

6. The splicing-type library construction method based on complementary tail-order annealing and cooperative extension linkage reaction according to claim 1, characterized in that: Step 3) specifically involves: simultaneously adding DNA polymerase and DNA ligase to the solution after the reaction, and supplementing with nucleotide substrate, and reacting at 37°C for 30 min to 2 h to carry out a synergistic extension and ligation reaction, so that the DNA polymerase first extends and completes the reaction, and then the DNA ligase catalyzes the formation of phosphodiester bonds to achieve covalent ligation.

7. The splicing library construction method based on complementary tail-order annealing and cooperative extension linkage reaction according to claim 6, characterized in that: The DNA ligase used is T4 or T7 DNALigase, and the nucleotide substrate is dATP or dTTP.

8. The splicing library construction method based on complementary tail-order annealing and cooperative extension linkage reaction according to claim 1, characterized in that: Step 4) specifically involves adding an end-repair primer with a 3' end of 15dT and an end-repair primer with a 3' end of 15dA to the double-stranded product, and simultaneously adding DNA polymerase, DNA ligase, and nucleotide substrate. The mixture is then reacted at 37°C for 10-30 minutes each to perform a synergistic extension ligation reaction, resulting in a preliminary library compatible with the adapter structure of third-generation sequencing.

9. A splicing-type library construction method based on complementary tail-order annealing and cooperative extension linkage reaction as described in claim 8, characterized in that: Before adding DNA polymerase, DNA ligase and nucleotide substrate, an annealing step is performed: denature at 95℃ for 2-3 min and then cool to 8-4℃ to allow the double-stranded product to be specifically ligated with the complementary tail sequence of the end-modified primer to form an inter-segment bridging structure.

10. The splicing library construction method based on complementary tail-order annealing and cooperative extension linkage reaction according to claim 1, characterized in that: Step 5) specifically involves: after purifying the preliminary library with magnetic beads to remove free primers and short fragments, ligating the library with adapters using the Nanopore or PacBio platform to obtain a spliced ​​library that can be directly used for third-generation sequencing.