Next-generation sequencing method based on Tn5 transposase and DNA library rapid construction kit
By preparing transposon complexes using Tn5 transposase and mixing them with sequencing adapter sequences, combined with PCR amplification and magnetic bead purification, the problems of cumbersome operation and high cost in next-generation sequencing methods are solved, achieving efficient and accurate DNA library construction and sequencing.
Patent Information
- Application Number
- CN202511516225.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-23
- Publication Date
- 2025-11-21
AI Technical Summary
Existing second-generation sequencing methods based on Tn5 transposases are cumbersome to operate, have poor compatibility with low starting DNA samples, suffer from sequence bias affecting library randomness, are costly, and are difficult to adapt to various sequencing needs.
The transposon complex was prepared using Tn5 transposase and mixed with the sequencing adapter sequence for DNA fragmentation and adapter addition. Combined with PCR amplification and sorting, the sequencing data was purified using magnetic beads and then processed and analyzed.
It simplifies the library construction process, improves sequencing efficiency and data management flexibility, ensures the accuracy and widespread application of high-throughput sequencing, reduces costs, and is adaptable to various sequencing platforms and scenarios.
Smart Images

Figure CN120989224A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of sequencing technology, specifically to a second-generation sequencing method based on Tn5 transposase and a rapid DNA library construction kit. Background Technology
[0002] High-throughput sequencing (HTS), also known as next-generation sequencing (NGS) or second-generation sequencing, is a revolutionary change from traditional Sanger sequencing (first-generation sequencing). Sanger sequencing, due to its high accuracy, is considered the "gold standard" for gene detection. However, Sanger sequencing has low throughput, obtaining only one sequence of 700-1000 bp in length per run, which cannot meet the needs of large-scale sequencing. It also requires high concentration and purity of the starting template DNA, is difficult to handle contaminated or specially structured samples, has a short effective read length for short PCR products, and is susceptible to environmental interference, leading to sequencing failures. Compared to traditional Sanger sequencing, high-throughput sequencing can obtain hundreds of thousands to millions of nucleic acid molecules in a single run, significantly improving sequencing efficiency and is widely used in fields such as genome sequencing, transcriptome sequencing, and epigenetic analysis.
[0003] In next-generation sequencing (NGS) library construction is the process of processing DNA samples into fragmented DNA suitable for sequencing. Traditional library construction methods typically include DNA fragmentation, end repair, adapter ligation, and PCR amplification. Tn5 transposase, derived from the transposon Tn5 in *E. coli*, has unique functions that have led to its widespread application in NGS library construction. By integrating DNA fragmentation and adapter ligation, Tn5 transposase can cleave double-stranded DNA (dsDNA) and insert specific adapter sequences at the cleavage site, thus achieving a one-step completion of DNA fragmentation and adapter ligation. Tn5 transposase requires a low amount of starting DNA, has a fast reaction rate, and can complete DNA fragmentation and adapter ligation in a short time.
[0004] Currently, Tn5 transposase-based DNA library construction methods have been widely used in next-generation sequencing, but some problems still exist. For example, the operation steps are cumbersome, requiring multiple reactions and failing to completely simplify the process; the compatibility with low starting amounts of DNA samples is poor; sequence bias exists during DNA fragmentation and adapter ligation, affecting the randomness of the library; the cost is high, with high reagent costs, making large-scale application difficult; and the versatility is insufficient, making it difficult to adapt to various sequencing needs, platforms, and application scenarios. Summary of the Invention
[0005] In view of the shortcomings of the existing technology, the purpose of this invention is to provide a second-generation sequencing method based on Tn5 transposase, a rapid DNA library construction kit for second-generation sequencing, and a method for constructing DNA libraries for second-generation sequencing.
[0006] To achieve the above objectives, the present invention adopts the following technical solution:
[0007] One aspect of the present invention provides a second-generation sequencing method based on Tn5 transposase, including constructing a second-generation sequencing DNA library, purifying the second-generation sequencing DNA library, sequencing with a second-generation sequencer, and processing and analyzing sequencing data;
[0008] The method for constructing a second-generation sequencing DNA library comprises the following steps:
[0009] (1) Preparation of transposable complex: Mix Tn5 transposase with the sequencing adapter sequence shown in SEQ ID NO:1-20, add transposable reaction buffer to the mixture, and incubate for 5-10 minutes to allow Tn5 transposase and adapter to form a stable complex. The prepared transposable complex can be stored at -20℃ or used directly in the next step.
[0010] (2) Sample addition: Prepare the DNA sample to be tested, add the transposon complex to the reaction tube containing the DNA sample to be tested, incubate at 37°C for 15 minutes to allow the transposon complex to fragment the DNA and add adapters to both ends, add stop solution to terminate the transposon reaction and form DNA fragments with adapters.
[0011] (3) PCR amplification: Set up a PCR reaction system, use P5 and N5 as primers to amplify DNA fragments with N5 adapters, and use P7 and N7 as primers to amplify DNA fragments with N7 adapters. Ensure that the PCR reaction system contains high-fidelity DNA polymerase, dNTPs and other necessary components, and perform PCR amplification according to the set reaction conditions.
[0012] (4) Fragment sorting: The amplified DNA fragments are sorted by gel electrophoresis or size selection column to remove DNA fragments that do not conform to the target length range, forming DNA fragments to be sequenced;
[0013] (5) Quality control: The sorted DNA fragments are subjected to quality control checks. The quality of the library is checked by electrophoresis or quantitative fluorescence method. The concentration of the library is adjusted to meet the requirements of the sequencing platform and the library construction is completed.
[0014] Preferably, the second-generation sequencing DNA library purification involves: recovering the library DNA using magnetic beads or column purification methods, and quantifying the library using quality control standards.
[0015] Preferably, the second-generation sequencer sequencing is performed by mixing different purified DNA libraries according to the effective concentration and target data volume requirements, and then using a sequencing platform to obtain sequencing data.
[0016] Preferably, the sequencing data processing and analysis follows the following specific route:
[0017] 1) Raw data preprocessing: Use the quality control software FastP to remove low-quality reads and cut out the sequencing adapter sequences in the reads to generate high-quality Clean Data;
[0018] 2) Alignment and filtering: Clean Data was aligned back to the reference genome using Bowtie 2 alignment software, PCR duplicates were removed using Sambamba, and reads with an alignment quality greater than 30 were extracted using Samtools.
[0019] 3) Standardization and visualization: Use deeptools to standardize the number of reads to RPKM values and generate bigwig files for signal peak visualization (visualization tool IGV).
[0020] 4) Advanced analysis: Calculate the Pearson correlation of reads between samples at the genome level (deeptools tool).
[0021] The HOMER tool was used to identify regions of significant enrichment of reads on the genome (Peak Calling).
[0022] Calculate the FRiP value to evaluate the signal-to-noise ratio (software: bedtools, samtools);
[0023] Evaluate the overlap between enriched peak regions obtained by different library construction methods (R packages ChIPseeker and ChIPpeakAnno).
[0024] 5) Visualization and report generation: Use IGV to display the signal peak diagram in the bigwig file and generate a comprehensive report.
[0025] Preferably, the PCR amplification reaction conditions are: pre-denaturation for 5 min; then denaturation at 95-98℃ for 45 s, annealing at 60-65℃ for 1.5 min, extension at 72℃ for 1 min, and amplification for 15-20 cycles.
[0026] Preferably, the DNA polymerase is a high-fidelity DNA polymerase.
[0027] Preferably, the DNA fragment to be sequenced is 300-600 bp in length.
[0028] Preferably, the starting template amount of the DNA sample to be tested is 1-50 ng.
[0029] Another aspect of the present invention provides a rapid DNA library construction kit for next-generation sequencing based on Tn5 transposase, the kit comprising: Tn5 transposase, transposition reaction buffer, stop solution, fragmentation reaction buffer, sequencing adapter sequence, DNA polymerase, DNA polymerase reaction buffer, positive control and negative control; wherein the sequencing adapter sequence is an N5 adapter sequence with an index SEQ ID NO:1-8 and an N7 adapter sequence with an index SEQ ID NO:9-20.
[0030] The beneficial effects achieved by this invention are as follows: This invention uses Tn5 transposase for paired-end index library construction and sequencing, which simplifies the library construction process, ensures the efficiency and accuracy of high-throughput sequencing, utilizes the unique characteristics of Tn5 transposase to achieve DNA fragmentation and adapter addition in a single step, and improves sequencing efficiency and data management flexibility by introducing paired-end indexes. Each index sequence ensures high specificity and traceability. At the same time, the sequencing method can work effectively with low starting template amounts, expanding the application range. In addition, Tn5 transposase has relatively uniform cutting characteristics, which can reduce sequence bias and ensure broader genome coverage and higher data reliability. Attached Figure Description
[0031] Figure 1 This is a schematic diagram of the Tn5 transposase-based next-generation sequencing DNA library construction method used in this embodiment of the invention. Detailed Implementation
[0032] The present application will now be described in further detail through specific embodiments. In these embodiments, many details are described to facilitate a better understanding of the present application. However, those skilled in the art will readily recognize that some features may be omitted in different situations, or may be replaced by other elements, materials, or methods. In some cases, certain operations related to the present application are not shown or described in the specification to avoid obscuring the core parts of the application with excessive description. For those skilled in the art, detailed description of these related operations is not necessary; they can fully understand the related operations based on the descriptions in the specification and general technical knowledge in the art.
[0033] This invention provides a Tn5 transposase-based next-generation sequencing method, a rapid DNA library construction kit for next-generation sequencing, and a method for constructing DNA libraries for next-generation sequencing. The principle of this Tn5 transposase-based next-generation sequencing DNA library construction method is as follows: Figure 1As shown, the sequencing adapter and OE adapter sequences are first integrated into a coated adapter, which is then incubated with transposase in vitro to form a DNA-binding transposase dimer. Subsequently, this complex is co-incubated with the DNA to be sequenced and broken down to form a sequenceable DNA fragment. Finally, an enzyme is added to fill in the nicks created during the insertion of the complex, and the library is constructed by PCR amplification. In addition, an index is introduced during PCR amplification to achieve large-scale parallel DNA sequencing.
[0034] The Tn5 transposase-based next-generation sequencing method of this application includes constructing a next-generation sequencing DNA library, purifying the next-generation sequencing DNA library, sequencing with a next-generation sequencer, and processing and analyzing sequencing data. All three steps are combined into constructing a next-generation sequencing DNA library and sequencing.
[0035] Next-generation sequencing DNA library construction sequencing technology route:
[0036] (1) Preparation of transposable complex: Mix Tn5 transposase with the sequencing adapter sequence shown in SEQ ID NO:1-20, add transposable reaction buffer to the mixture, and incubate for 5-10 minutes to allow Tn5 transposase and adapter to form a stable complex. The prepared transposable complex can be stored at -20℃ or used directly in the next step.
[0037] (2) Sample addition: Prepare the DNA sample to be tested, add the transposon complex to the reaction tube containing the DNA sample to be tested, incubate at 37°C for 15 minutes to allow the transposon complex to fragment the DNA and add adapters to both ends, add stop solution to terminate the transposon reaction and form DNA fragments with adapters.
[0038] (3) PCR amplification: Set up a PCR reaction system, use P5 and N5 as primers to amplify DNA fragments with N5 adapters, and use P7 and N7 as primers to amplify DNA fragments with N7 adapters. Ensure that the PCR reaction system contains high-fidelity DNA polymerase, dNTPs and other necessary components (including PCR reaction buffer, MgCl2 and bovine serum albumin). Perform PCR amplification according to the set reaction conditions.
[0039] (4) Fragment sorting: The amplified DNA fragments are sorted by gel electrophoresis or size selection column to remove DNA fragments that do not conform to the target length range, forming DNA fragments to be sequenced;
[0040] (5) Quality control: The sorted DNA fragments are subjected to quality control checks. The quality of the library is checked by electrophoresis or quantitative fluorescence method. The concentration of the library is adjusted to meet the requirements of the sequencing platform and the library construction is completed.
[0041] (6) Second-generation sequencing DNA library purification: The library DNA is recovered using magnetic beads or column purification methods, and the library is quantified using quality control standards.
[0042] (7) Second-generation sequencer sequencing: According to the effective concentration and the required amount of data to be sequenced, different purified DNA libraries are mixed and sequenced using a sequencing platform to obtain sequencing data.
[0043] Sequencing data processing and analysis technology roadmap:
[0044] 1) Raw data preprocessing: Use the quality control software FastP to remove low-quality reads and cut out the sequencing adapter sequences in the reads to generate high-quality Clean Data;
[0045] 2) Alignment and filtering: Clean Data was aligned back to the reference genome using Bowtie 2 alignment software, PCR duplicates were removed using Sambamba, and reads with an alignment quality greater than 30 were extracted using Samtools.
[0046] 3) Standardization and visualization: Use deeptools to standardize the number of reads to RPKM values and generate bigwig files for signal peak visualization (visualization tool IGV).
[0047] 4) Advanced analysis: Calculate the Pearson correlation of reads between samples at the genome level (deeptools tool).
[0048] The HOMER tool was used to identify regions of significant enrichment of reads on the genome (Peak Calling).
[0049] Calculate the FRiP value to evaluate the signal-to-noise ratio (software: bedtools, samtools);
[0050] Evaluate the overlap between enriched peak regions obtained by different library construction methods (R packages ChIPseeker and ChIPpeakAnno).
[0051] 5) Visualization and report generation: Use IGV to display the signal peak diagram in the bigwig file and generate a comprehensive report.
[0052] Low-quality reads refer to sequencing reads that do not meet high-quality standards in one or more aspects.
[0053] Low-quality reads include: ① High base identification error rate: The quality score (Q-value) of each base reflects the probability that the base is correctly identified. Generally, reads with Q-values below a certain threshold (e.g., Q20, representing 99% accuracy; or Q30, representing 99.9% accuracy) are considered low-quality reads; ② Excessive unidentified bases (N): If a read contains a large number of unidentified bases (labeled N), it indicates that bases at these positions were not accurately identified during sequencing, and such reads are also considered low-quality; ③ Reads that are too short: For some applications, if the read length is significantly shorter than expected, it may not provide sufficient information for downstream analysis, and such reads may also be considered low-quality; ④ Adapter contamination: When a sequencing read contains sequencing adapter sequences that should not be present, this may be due to incomplete fragmentation or non-specific binding during amplification. These reads containing adapter sequences should also be marked as low-quality and removed.
[0054] During sequencing data processing and analysis, the FASTP software tool is used to automatically filter out reads that do not meet the standards based on a set quality threshold (e.g., average Q value less than Q20). Reagents for rapid construction of next-generation sequencing DNA libraries based on Tn5 transposase include:
[0055] Tn5 transposase, transposition reaction buffer, stop solution, fragmentation reaction buffer, sequencing adapter sequence, DNA polymerase, DNA polymerase reaction buffer, positive control and negative control; the sequencing adapter sequence is an N5 adapter sequence with an index SEQ ID NO:1-8 and an N7 adapter sequence with an index SEQ ID NO:9-20.
[0056] In addition, nucleic acids extracted using existing mature nucleic acid extraction technologies can all be used in this library construction kit. After sample extraction, the product concentration was quantitatively detected using "Qubit 4.0" to determine the initial library construction concentration. PCR amplification products from positive samples were used to prepare strong positive control samples (1x10⁻¹²). 6 (copy / μl) and critical positive control (1x10) 2 (Copies / μl), using laboratory-grade pure water as an anion control.
[0057] Example: A second-generation sequencing method based on Tn5 transposase, the specific steps of which are as follows:
[0058] (1) Preparation of transposon complex: Tn5 transposase from Illumina Nextera XT DNA Library Prep Kit was mixed with sequencing adapters, transposon reaction buffer was added to the mixture, and the mixture was incubated at 37°C for 5 minutes. Then it was immediately placed on ice to cool and form a stable transposon complex, which was stored at -20°C. The sequencing adapter sequences included N5 adapter sequences with indexes as shown in SEQ ID NO:1-8 and N7 adapter sequences with indexes as shown in SEQ ID NO:9-20.
[0059] SEQ ID NO:1-8:5'-AATGATACGGCGACCACCGAGATCTACACNNNNNNNNTCGTCGGCAGCGTC-3'
[0060] SEQ ID NO:9-20:5'-CAAGCAGAAGACGGCATACGAGATNNNNNNNNGTCTCGTGGGCTCGG-3'
[0061] In the sequencing adapter sequence, [NNNNNNNN] is the index sequence.
[0062] Mixed specifications: 5 μL Tn5 transposase, 5 μL sequencing adapters (each Index sequence), 10 μL transposation reaction buffer, add water to a total volume of 50 μL;
[0063] (2) Adding the sample: Prepare a DNA sample to be tested with a starting template amount of 10 ng. Add the transposon complex to the reaction tube containing the DNA sample to be tested, adjust the total volume to 50 μL, and incubate at 37°C for 15 minutes. After incubation, add 5 μL of stop solution (EDTA solution) to terminate the transposon reaction and form a DNA fragment with adapter.
[0064] (3) PCR amplification: Set up the PCR reaction system, including 1 μL of high-fidelity DNA polymerase, 5 μL of 10× DNA polymerase reaction buffer (components: 500 mM pH 8.8 Tris-HCl, 100 mM KCl, 15 mM MgCl2, 100 μg / mL BSA) (final concentration: 1×), 1 μL of dNTPs (10 mM each), 2.5 μL each of primers (P5 and N5 or P7 and N7, 10 μM) (final concentration: 0.5 μM), 5 μL of DNA fragment with adapter, and add water to a total volume of 50 μL; the PCR program is set to pre-denaturation for 5 minutes; then denature at 95℃ for 45 seconds, anneal at 60℃ for 1.5 minutes, extend at 72℃ for 1 minute, for a total of 20 cycles, and finally extend at 72℃ for 5 minutes, and maintain at 4℃;
[0065] The PCR amplification primers include P5 primer and P7 primer. The P5 primer is complementary to the N5 adapter sequence, and the P7 primer is complementary to the N7 adapter sequence. The P5 primer sequence is shown in SEQ ID NO:21, and the P7 primer sequence is shown in SEQ ID NO:22.
[0066] SEQ ID NO:21: 5'-AATGATACGGCGACCACCGAGATCTACAC-3'
[0067] SEQ ID NO:22:5'-CAAGCAGAAGACGGCATACGAGAT-3'
[0068] (4) Fragment sorting: Mix the amplified DNA fragment with SPRIselect magnetic beads at a volume ratio of 1:0.8, incubate at room temperature for 5 minutes, separate the magnetic beads using a magnetic rack, remove the supernatant, wash the magnetic beads twice with 80% ethanol, and aspirate the liquid after standing for 30 seconds each time. Dry the magnetic beads at room temperature for 5 minutes, and finally elute the target DNA fragment (length between 300-600 bp, confirmed by gel electrophoresis) with TE buffer to form the DNA fragment to be sequenced.
[0069] (5) Quality control: Perform quality control checks on the sorted DNA fragments, use Qubit to determine the library concentration, adjust the library concentration to meet the requirements of the sequencing platform, and complete the library construction;
[0070] (6) Purification of second-generation sequencing DNA library: Refer to (4) fragment sorting, use magnetic beads to recover library DNA as needed, and quantify the library using quality control standards.
[0071] (7) Second-generation sequencer sequencing: According to the effective concentration and the required amount of data to be sequenced, different purified DNA libraries are mixed and loaded onto the Illumina MiSeq or NovaSeq platform for sequencing to obtain sequencing data.
[0072] (8) Sequencing data processing and analysis:
[0073] 1) Raw data preprocessing: Use the quality control software FastP to remove low-quality reads and cut out the sequencing adapter sequences in the reads to generate high-quality Clean Data;
[0074] 2) Alignment and filtering: Clean Data was aligned back to the reference genome using Bowtie 2 alignment software, PCR duplicates were removed using Sambamba, and reads with an alignment quality greater than 30 were extracted using Samtools.
[0075] 3) Standardization and visualization: Use deeptools to standardize the number of reads to RPKM values and generate bigwig files for signal peak visualization (visualization tool IGV).
[0076] 4) Advanced analysis: Calculate the Pearson correlation of reads between samples at the genome level (deeptools tool).
[0077] The HOMER tool was used to identify regions of significant enrichment of reads on the genome (Peak Calling).
[0078] Calculate the FRiP value to evaluate the signal-to-noise ratio (software: bedtools, samtools);
[0079] Evaluate the overlap between enriched peak regions obtained by different library construction methods (R packages ChIPseeker and ChIPpeakAnno).
[0080] 5) Visualization and Report Generation: Use IGV to display signal peaks in bigwig files and generate comprehensive reports that include charts, statistics, and biological interpretations.
[0081] The N5 and N7 adapter sequences with indexes in the sequencing adapter sequences of this embodiment include the following sequences:
[0082] Table 1 Sequencing adapter sequence listing
[0083] Primer name Primer sequence Index sequence SEQ ID NO N501 5'-AATGATACGGCGACCACCGAGATCTACAC TAGATCGC TCGTCGGCAGCGTC-3']] TAGATCGC 1 N502 5'-AATGATACGGCGACCACCGAGATCTACAC CTCTCTAT TCGTCGGCAGCGTC-3']] CTCTCTAT 2 N503 5'-AATGATACGGCGACCACCGAGATCTACAC TATCCTCT TCGTCGGCAGCGTC-3']] TATCCTCT 3 N504 5'-AATGATACGGCGACCACCGAGATCTACAC AGAGTAGA TCGTCGGCAGCGTC-3']] AGAGTAGA 4 N505 5'-AATGATACGGCGACCACCGAGATCTACAC GTAAGGAG TCGTCGGCAGCGTC-3']] GTAAGGAG 5 N506 5'-AATGATACGGCGACCACCGAGATCTACAC ACTGCATA TCGTCGGCAGCGTC-3']] ACTGCATA 6 N507 5'-AATGATACGGCGACCACCGAGATCTACAC AAGGAGTA TCGTCGGCAGCGTC-3']] AAGGAGTA 7 N508 5'-AATGATACGGCGACCACCGAGATCTACAC CTAAGCCT TCGTCGGCAGCGTC-3']] CTAAGCCT 8 N701 5'-CAAGCAGAAGACGGCATACGAGAT TCGCCTTA GTCTCGTGGGCTCGG-3' TCGCCTTA 9 N702 5-CAAGCAGAAGACGGCATACGAGAT CTAGTACG GTCTCGTGGGCTCGG-3']] CTAGTACG 10 N703 5'-CAAGCAGAAGACGGCATACGAGAT TTCTGCCT GTCTCGTGGGCTCGG-3' TTCTGCCT 11 N704 5'-CAAGCAGAAGACGGCATACGAGAT GCTCAGGA GTCTCGTGGGCTCGG-3' GCTCAGGA 12 N705 <![CDATA[5'-CAAGCAGAAGACGGCATACGAGAT AGGAGTCC GTCTCGTGGGCTCGG-3']]> AGGAGTCC 13 N706 <![CDATA[5-CAAGCAGAAGACGGCATACGAGAT CATGCCTA GTCTCGTGGGCTCGG-3']]> CATGCCTA 14 N707 <![CDATA[5-CAAGCAGAAGACGGCATACGAGAT GTAGAGAG GTCTCGTGGGCTCGG-3']]> GTAGAGAG 15 N708 <![CDATA[5-CAAGCAGAAGACGGCATACGAGAT CCTCTCTG GTCTCGTGGGCTCGG-3']]> CCTCTCTG 16 N709 <![CDATA[5-CAAGCAGAAGACGGCATACGAGAT AGCGTAGC GTCTCGTGGGCTCGG-3']]> AGCGTAGC 17 N710 <![CDATA[5'-CAAGCAGAAGACGGCATACGAGAT CAGCCTCG GTCTCGTGGGCTCGG-3']]> CAGCCTCG 18 N711 <![CDATA[5'-CAAGCAGAAGACGGCATACGAGAT TGCCTCTT GTCTCGTGGGCTCGG-3']]> TGCCTCTT 19 N712 <![CDATA[5'-CAAGCAGAAGACGGCATACGAGAT TCCTCTAC GTCTCGTGGGCTCGG-3']]> TCCTCTAC 20
[0084] In the table, the underlined sequences are the corresponding index sequences.
[0085] Performance testing: 1. Document library build success rate
[0086] Two DNA samples from different sources were selected: human genomic DNA and mouse genomic DNA. The initial template amount for each sample was 10 ng. The experimental group was operated according to the steps provided in the example, while the control group used the conventional method of Illumina TruSeq DNAPCR-Free Library Prep Kit.
[0087] Library construction: Tn5 transposase, sequencing adapter, and transposition reaction buffer were mixed and incubated at 37°C for 5 minutes, then cooled to ice for storage; the transposition complex was added to a reaction tube containing 10 ng of the DNA sample to be tested, and the total volume was adjusted to 50 μL, and incubated at 37°C for 15 minutes; then the stop solution was added to stop the reaction; the PCR reaction system was set up and amplification was performed according to the specified PCR program; the target size DNA fragment was recovered using the SPRIselect magnetic bead method.
[0088] Control group library construction: The procedure was performed according to the instructions of the Illumina TruSeq DNA PCR-Free Library Prep Kit.
[0089] Qubit quantitative analysis was performed on the library of each sample to confirm the quality and concentration of the library. The success rate of library construction for each method was recorded (number of successfully constructed samples / total number of samples). If the library concentration reached the expected standard and there was no significant degradation, it was considered to be successfully constructed.
[0090] The test results are shown in Table 2. The experimental group (the method of the example) successfully constructed the library in all attempts, while the control group (the traditional method) only had about one-third of the experiments successful. This indicates that the Tn5 transposase-based method has a higher library construction success rate.
[0091] Table 2 Comparison of library construction quality and success rate between the examples and traditional methods
[0092] Sample type Number of experiments Experimental library quality (Qubit concentration, ng / μL) Was the experimental group successfully constructed? Control group library quality (Qubit concentration, ng / μL) Whether the control group was successfully constructed Human Genome 1 20 yes 18 yes Human Genome 2 21 yes 17 no Human Genome 3 22 yes 16 no Mouse genome 1 19 yes 17 yes Mouse genome 2 18 yes 15 no Mouse genome 3 20 yes 16 no
[0093] 2. Sensitivity and Specificity
[0094] Different concentrations of DNA samples were selected (starting amounts of 0.1 ng, 1 ng, 5 ng, 10 ng, and 50 ng, and diluted to a total volume of 50 μL). The experimental group followed the steps in the example, while the control group used the traditional method with the Illumina TruSeq DNA PCR-FreeLibrary Prep Kit.
[0095] For each selected concentration of DNA sample, libraries were constructed using experimental and control methods, respectively. After library construction, all samples underwent quality control checks to ensure the libraries met sequencing requirements. The constructed libraries were sequenced using Illumina MiSeq or NovaSeq platforms. The sequencing data underwent preliminary processing, including removing low-quality reads and adapter sequences to generate high-quality Clean Data. The Clean Data was aligned back to the reference genome using alignment software (BWA), and the effective data volume and alignment rate were calculated. Bioinformatics tools (samtools) were used to check for non-specific insertions or other background noise.
[0096] The test results are shown in Table 3. Compared with the control group (traditional method), the experimental group (method of the example) showed higher effective data volume and alignment rate under different starting DNA amounts, while maintaining a low non-specific insertion ratio, improving the quality and reliability of sequencing data, and demonstrating better sensitivity and specificity.
[0097] Table 3. Average effective data volume, alignment rate, and nonspecific insertion ratio at different concentrations.
[0098] Initial DNA amount (ng) Number of experiments Effective data volume of the experimental group (M reads) Number of valid data in the control group (M reads) Comparison rate of experimental group (%) Control group comparison rate (%) Nonspecific insertion rate (%) in the experimental group Nonspecific insertion rate (%) in the control group 0.1 1 8 6 85 78 0.5 0.6 0.1 2 9 5 86 77 0.6 0.7 0.1 3 10 6 84 79 0.5 0.6 1 1 15 12 90 85 0.3 0.4 1 2 16 11 91 84 0.4 0.5 1 3 15 12 90 85 0.3 0.4 5 1 20 16 92 87 0.2 0.3 5 2 21 17 93 88 0.2 0.3 5 3 20 16 92 87 0.2 0.3 10 1 25 20 93 88 0.1 0.2 10 2 26 21 94 89 0.1 0.2 10 3 25 20 93 88 0.1 0.2 50 1 30 24 94 90 0.05 0.1 50 2 31 25 95 91 0.05 0.1 50 3 30 24 94 90 0.05 0.1
[0099] 3. Repeatability and Reproducibility
[0100] A representative DNA sample (human genomic DNA) was selected, starting with 10 ng. Experiments were conducted in two different laboratory environments, ensuring that equipment and reagent batches were kept as consistent as possible in each environment. Experiments were performed by operators with varying levels of experience to assess the reproducibility of the method. For the selected sample, libraries were constructed using both the experimental group (example) and the control group (traditional library construction method), with each sample undergoing three independent experiments. The detailed steps described in the example were followed, including transposon preparation, sample addition, PCR amplification, fragment sorting, and quality control. The constructed libraries were sequenced using the Illumina MiSeq platform.
[0101] The sequencing data undergoes preliminary processing, including removing low-quality reads and adapter sequences to generate high-quality Clean Data. Alignment software (such as BWA) is used to align the Clean Data back to the reference genome, and the effective data volume and alignment rate are calculated. Bioinformatics tools (samtools) are used to check for non-specific insertions or other background noise. The effective data volume, alignment rate, and proportion of non-specific insertions are recorded for each experiment using each method to ensure consistency and reliability of the results.
[0102] The test results are shown in Table 4. The experimental group (example) showed high repeatability and reproducibility when the method was performed independently by different operators under different laboratory conditions. Its average effective data volume was higher than that of the control group (traditional method). At the same time, it maintained a high comparison rate and a low non-specific insertion ratio, indicating that the method not only performed well in a single experiment, but also provided stable and reliable results in multiple experiments across laboratories and operators.
[0103] Table 4. Comparison of repeatability and reproducibility between the examples and the traditional methods.
[0104] laboratory Operators method Effective data volume (Mreads) Comparison rate (%) Nonspecific insertion rate (%) Lab A Operator 1 experimental group 25 94 0.1 Lab A Operator 2 experimental group 26 95 0.1 Lab A Operator 3 experimental group 25 94 0.1 Lab B Operator 1 experimental group 24 93 0.1 Lab B Operator 2 experimental group 25 94 0.1 Lab B Operator 3 experimental group 24 93 0.1 Lab A Operator 1 control group 20 88 0.3 Lab A Operator 2 control group 19 87 0.4 Lab A Operator 3 control group 20 88 0.3 Lab B Operator 1 control group 19 87 0.4 Lab B Operator 2 control group 20 88 0.3 Lab B Operator 3 control group 19 87 0.4
[0105] 4. Effectiveness of advanced analysis
[0106] A set of widely accepted standard or simulated datasets (such as data provided by the ENCODE project) were selected, and data obtained from the experimental group (example) and the control group (traditional library construction method) were analyzed separately. The HOMER tool was used to process the data generated by both methods, identifying significant enriched regions (Peaks), and recording the number and location of Peaks identified in each sample. The FRiP (Fraction of Reads in Peaks) values for each method were calculated using bedtools and samtools. First, the aligned SAM / BAM files were converted to BED format, and their intersection with the Peaks file was calculated. The FRiP value for each sample was recorded, along with the number of Peaks and the FRiP value for each method, to evaluate the performance of different methods in advanced analysis.
[0107] Two different samples (Sample 1 and Sample 2) were used, and the experimental and control methods were used for analysis, respectively. The test results are shown in Table 5. The experimental method showed a significant advantage in peak calling and FRiP value calculation when performing advanced analysis. The experimental method identified more peaks than the traditional method in the control group, while maintaining a higher FRiP value, indicating that it has higher sensitivity and specificity in detecting significantly enriched regions on the genome.
[0108] Table 5. Number of Peaks and FRiP Values for Sample 1
[0109] method Number of experiments Number of Peaks FRiP value (%) experimental group 1 12,473 73.2 experimental group 2 12,608 75.9 experimental group 3 12,394 74.1 control group 1 10,189 62.3 control group 2 10,321 66.1 control group 3 10,412 64.8
[0110] Table 6. Number of Peaks and FRiP Values for Sample 2
[0111] method Number of experiments Number of Peaks FRiP value (%) experimental group 1 12,501 73.5 experimental group 2 12,610 76.0 experimental group 3 12,389 74.0 control group 1 10,190 62.4 control group 2 10,320 66.0 control group 3 10,410 64.7
[0112] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the claims of the present invention.
Claims
1. A second-generation sequencing method based on Tn5 transposase, characterized in that, This includes constructing second-generation sequencing DNA libraries, purifying second-generation sequencing DNA libraries, sequencing using a second-generation sequencer, and processing and analyzing sequencing data. The specific steps for constructing the next-generation sequencing DNA library are as follows: (1) Preparation of transposable complex: Mix Tn5 transposase with the sequencing adapter sequence shown in SEQ ID NO:1-20, add transposable reaction buffer to the mixture, and incubate for 5-10 minutes to allow Tn5 transposase and adapter to form a stable complex. The prepared transposable complex can be stored at -20℃ or used directly in the next step. (2) Sample addition: Prepare the DNA sample to be tested, add the transposon complex to the reaction tube containing the DNA sample to be tested, incubate at 37°C for 15 minutes to allow the transposon complex to fragment the DNA and add adapters to both ends, add stop solution to terminate the transposon reaction and form DNA fragments with adapters. (3) PCR amplification: Set up a PCR reaction system, use P5 and N5 as primers to amplify DNA fragments with N5 adapters, and use P7 and N7 as primers to amplify DNA fragments with N7 adapters. The PCR reaction system contains high-fidelity DNA polymerase, dNTPs and other necessary components. Perform PCR amplification according to the PCR amplification reaction conditions. (4) Fragment sorting: The amplified DNA fragments are sorted by gel electrophoresis or size selection column to remove DNA fragments that do not conform to the target length range, forming DNA fragments to be sequenced; (5) Quality control: The sorted DNA fragments are subjected to quality control checks. The quality of the library is checked by electrophoresis or quantitative fluorescence method. The concentration of the library is adjusted to meet the requirements of the sequencing platform and the library construction is completed.
2. The next-generation sequencing method based on Tn5 transposase according to claim 1, characterized in that, The purification of the second-generation sequencing DNA library involves: recovering the library DNA using magnetic beads or column purification methods, and quantifying the library using quality control standards.
3. The next-generation sequencing method based on Tn5 transposase according to claim 1, characterized in that, The second-generation sequencer sequencing process involves mixing different purified DNA libraries according to the required effective concentration and target data volume, and then using a sequencing platform to perform sequencing to obtain sequencing data.
4. The next-generation sequencing method based on Tn5 transposase according to claim 1, characterized in that, The specific route for sequencing data processing and analysis is as follows: 1) Raw data preprocessing: Use the quality control software FastP to remove low-quality reads and cut out the sequencing adapter sequences in the reads to generate high-quality Clean Data; 2) Alignment and filtering: Clean Data was aligned back to the reference genome using Bowtie 2 alignment software, PCR duplicates were removed using Sambamba, and reads with an alignment quality greater than 30 were extracted using Samtools. 3) Standardization and visualization: Use deeptools to standardize the number of reads to RPKM values and generate bigwig files for signal peak visualization; 4) Advanced analysis: Calculate Pearson correlations of reads across the genome using the deeptools tool; The HOMER tool was used to identify regions of significant enrichment of reads on the genome; the FRiP values were calculated using the software bedtools and samtools to assess the signal-to-noise ratio; and the overlap between enriched peak regions obtained by different library preparation methods was evaluated using the R packages ChIPseeker and ChIPpeakAnno. 5) Visualization and Report Generation: Use IGV visualization to display signal peak diagrams in the bigwig file and generate comprehensive reports.
5. The next-generation sequencing method based on Tn5 transposase according to claim 1, characterized in that, The PCR amplification reaction conditions are as follows: pre-denaturation for 5 min; then denaturation at 95-98℃ for 45 s, annealing at 60-65℃ for 1.5 min, extension at 72℃ for 1 min, and amplification for 15-20 cycles.
6. The next-generation sequencing method based on Tn5 transposase according to claim 1, characterized in that, The DNA polymerase is a high-fidelity DNA polymerase.
7. The next-generation sequencing method based on Tn5 transposase according to claim 1, characterized in that, The DNA fragment to be sequenced is 300-600 bp in length.
8. The next-generation sequencing method based on Tn5 transposase according to claim 1, characterized in that, The initial template amount of the DNA sample to be tested is 1-50 ng.
9. A rapid DNA library construction kit for next-generation sequencing based on Tn5 transposase, characterized in that, The kit contains reagents for constructing next-generation sequencing DNA libraries, comprising: Tn5 transposase, transposition reaction buffer, stop solution, fragmentation reaction buffer, sequencing adapter sequences, DNA polymerase, DNA polymerase reaction buffer, positive control and negative control; the sequencing adapter sequences are N5 adapter sequences with index SEQ ID NO:1-8 and N7 adapter sequences with index SEQ ID NO:9-20.
Citation Information
Patent Citations
Preparation method of accurate quantitative ATAC-seq library and kit
CN112410403A
Method for constructing microscale sample m6A modification detection library under assistance of Tn5 transposase and application thereof
CN113061648A
Second-generation sequencing method and library construction method
CN113337590A
High-efficiency high-throughput gene sequencing data processing system
CN119446266A
Method of constructing sequencing library
US20170341051A1
Cited By
Plasmid sequencing method
CN114507903A