A targeted nanopore sequencing method based on hairpin adapter circularization and tandem repeat correction, circular template and kit
Patent Information
- Application Number
- CN202611081811.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-21
- Publication Date
- 2026-09-25
AI Technical Summary
然而,现有方法在靶向富集和建库过程中通常将PCR作为扩增手段,存在以下问题:PCR的指数扩增极易产生气溶胶污染,且气溶胶产物通常不含接头序列,与后续实验体系的模板无法有效区分,导致假阳性风险高;扩增在连接接头之前进行,无法在源头引入唯一分子标签,难以区分原始分子和扩增重复,也无法消除扩增和测序错误累积带来的定量偏差
(1)扩增前即引入UMI,且RCA使UMI与靶序列共同串联重复,可在单个读段内对UMI本身进行多次读取校准,彻底消除扩增或测序错误导致的UMI误读,实现精准的分子去重和绝对定量。
Abstract
Description
Technical Field
[0001] This invention relates to the field of gene sequencing technology, specifically to a targeted nucleic acid sequencing method, and more particularly to a method for high-accuracy targeted sequencing by generating tandem repeat sequences using hairpin adapter circularization and rolling circle amplification, and then performing nanopore sequencing and tandem repeat consistency correction. Background Technology
[0002] Nanopore sequencing technology has advantages such as long read lengths, real-time detection, and portable equipment, but its accuracy in single-base recognition is relatively low, which limits its application in targeted sequencing fields that require high-precision detection.
[0003] Generating tandem repeat copies of target molecules using rolling circle amplification (RCA) and then sequencing and calibrating multiple copies of the same molecule is an effective strategy for improving the accuracy of nanopore sequencing. However, existing methods typically use PCR as the amplification method during targeted enrichment and library construction, which has the following problems: exponential amplification by PCR is prone to aerosol contamination, and aerosol products usually do not contain adapter sequences, making them indistinguishable from the template in subsequent experimental systems, resulting in a high risk of false positives; amplification occurs before adapter ligation, failing to introduce a unique molecular tag at the source, making it difficult to distinguish the original molecule from the amplified repeat, and failing to eliminate quantitative bias caused by the accumulation of amplification and sequencing errors. In addition, conventional RCA strategies do not integrate a unique molecular tag into the tandem repeat structure, making it impossible to calibrate the tag itself during sequencing, leading to inaccurate molecule counting.
[0004] Therefore, there is an urgent need for a sequencing method that can significantly improve the accuracy of nanopore targeted sequencing, completely avoid aerosol cross-contamination, and achieve precise molecular quantification. Summary of the Invention
[0005] The purpose of this invention is to provide a method that can significantly improve the accuracy of nanopore targeted sequencing, completely avoid aerosol contamination, and achieve precise molecular counting.
[0006] To achieve the above objectives, the present invention provides a targeted nucleic acid sequencing method, comprising the following steps: (1) Fragment, end repair, phosphorylation and A-tailing of the nucleic acid sample to be tested are performed to obtain linear double-stranded nucleic acid fragments with dA tails at both ends; (2) A hairpin connector with a pre-annealed stem-loop structure is provided, wherein the hairpin connector is a single-stranded DNA, wherein its 5' end and 3' end are anti-complementary to form a double-stranded stem, and the middle loop region contains a unique molecular tag (UMI) composed of a random sequence, and the 5' end of the hairpin connector has a phosphate group; (3) The hairpin connector is connected to the linear double-stranded nucleic acid fragment obtained in step (1) under the action of ligase, so that the two strands of the double-stranded nucleic acid fragment are circularized through the hairpin connector to form a single-stranded covalently closed circular template. (4) Using the single-stranded circular template as a template, and using RCA primers complementary to the hairpin connector circular region sequence, rolling circle amplification is performed under the action of strand displacement polymerase to generate a linear RCA product. The RCA product is a long chain containing multiple tandem repeat sequences, and each repeat unit contains the target nucleic acid sequence and the unique molecular tag. (5) Using the linear RCA product as a template, superbranched rolling circle amplification was performed using target-specific primers and universal primers complementary to the hairpin adapter sequence to obtain the HRCA product; (6) Perform nanopore sequencing on the HRCA product; (7) Identify tandem repeat units from nanopore sequencing reads, extract the target sequence and unique molecular tag of each repeat unit, obtain a high-accuracy consistent sequence through multicopy consistency correction, and use the corrected unique molecular tag for molecular counting and deduplication.
[0007] The core advantage of this invention is: (1) UMI is introduced before amplification, and RCA makes UMI and target sequence repeat in tandem. UMI itself can be read and calibrated multiple times in a single read segment, completely eliminating UMI misreading caused by amplification or sequencing errors, and achieving accurate molecular deduplication and absolute quantification.
[0008] (2) Long tandem repeat structures allow for self-consistency correction using repeat units within the same read segment, significantly reducing random errors in nanopore sequencing and obtaining highly accurate consistent sequences.
[0009] (3) All amplifications are performed after the hairpin adapter is connected. The first round of RCA is linear amplification, and the product is a very long DNA chain that is not easy to form aerosols. Even if a small amount of aerosol is generated, both ends of the aerosol have the full-length sequence of the hairpin adapter, which cannot be effectively replicated or generate characteristic ends in subsequent amplifications. It can be filtered by bioinformatics, thus eliminating aerosol cross-contamination at both the physical and informational levels.
[0010] (4) The HRCA step is performed using target-specific primers and universal primers, similar to the specific amplification of PCR. However, because the template is a tandem repeat structure, the product is tandem and branched, which further improves the sequencing throughput and signal intensity.
[0011] (5) This method uniformly processes DNA and RNA (reverse transcription into double-stranded cDNA) samples, and the process is simple. Detailed Implementation Example 1: Fragmentation, terminal repair, phosphorylation, and addition of an A tail; Using a commercially available one-step fragmentation, end-repair, and A-tailing kit (such as Novizan ND627), genomic DNA or double-stranded cDNA is processed to obtain fragmented linear double-stranded DNA with dA tails at both ends. The fragment length is controlled at 300-500 bp, and then purified for later use.
[0012] Hairpin connector annealing and ring-forming connection: A synthetic hairpin connector, the nucleotide sequence of which is shown in SEQ ID NO:1: 5'-GATCGGAAGAGCACACGTCTNNNNNNATCGTACGTTATTTATTTACGGCGCTAGCTAGCTAGCTAGCTAGACGTGTGCTCTTCCGATCT-3' NNNNNN is a 6-base random sequence that constitutes a unique molecular tag (UMI). The single-stranded DNA was dissolved in annealing buffer, heated at 95°C for 5 minutes, and then slowly cooled to room temperature to allow it to self-anneal and form a stem-loop structure: the reverse complementary regions at both ends of the sequence pair to form a double-stranded stem, the middle region remains a single-stranded loop, and the 5' end carries a phosphate group.
[0013] The annealed hairpin connectors were mixed with the DNA fragments obtained in step 1 and ligated in a T4 DNA ligase system. The 5' phosphate of the hairpin connector was linked to the 3'-OH of one strand of the DNA fragment, and the 3'-OH of the hairpin connector was linked to the 5' phosphate of the other strand of the DNA fragment, thereby connecting the two strands of the double-stranded DNA fragment into two independent single-stranded covalently closed circular DNA molecules.
[0014] Rolling circle amplification (RCA) Using purified single-stranded circular DNA as a template, RCA primers (nucleotide sequence shown in SEQ ID NO:2: 5'-AGCTAGCTAGCTAGCGCCGTA-3', with a phosphate thioester modification between the two nucleotides at the 3' end), phi29 DNA polymerase, and dNTPs were added. The mixture was reacted at 30°C for rolling circle amplification, producing an ultra-long RCA product containing hundreds or thousands of tandem repeats. The RCA product was then purified.
[0015] Hyperbranched rolling circle amplification (HRCA) Using purified RCA products as templates, target-specific forward primers (designed according to the target gene) and universal reverse primers (such as 5'-AGACGTGTGCTCTTCCGATC-3', complementary to the upstream sequence of the hairpin adapter loop region) were added, and HRCA amplification was performed under the action of Bst DNA polymerase. Because the RCA products contain a large number of target sequences and universal primer binding sites arranged in tandem, the reaction produces a large number of branched, tandemly repeated products of varying lengths, forming diffuse bands. The HRCA products were then purified.
[0016] Nanopore sequencing The HRCA product is fragmented into 3-5K fragments, the ends are repaired, and sequencing adapters are ligated for single-molecule sequencing on a nanopore sequencer.
[0017] Data Analysis After base identification of the raw sequencing electrical signal, the boundaries of tandem repeat units are located by searching for hairpin connector feature sequences. Each read is split into consecutive repeat units, and the target sequence fragment and UMI sequence of each unit are extracted. Multiple sequence alignment is performed on all target sequence fragments within the same read to generate a consensus sequence, eliminating random errors in nanopore sequencing. The UMI sequence is corrected by majority voting to obtain an accurate UMI. The corrected UMI is used to remove duplicates and count the consensus sequence, achieving high-accuracy qualitative and absolute quantitative analysis of the target molecules.
[0018] Example 2 (Accuracy Verification) Using standards with known mutation frequencies as samples, the method of this invention and conventional nanopore targeted sequencing methods were employed for detection. After tandem repeat consistency correction, the base accuracy of this invention was improved from the conventional ~90% to >99.9%, and it successfully detected 0.1% low-frequency mutations. Conventional methods, however, could not reliably detect these low-frequency mutations due to sequencing errors.
[0019] The above description is only a preferred embodiment of the present invention. All equivalent changes and modifications made within the scope of the claims of the present invention should be included in the scope of the present invention.
Claims
1. A targeted nucleic acid sequencing method, characterized in that, Includes the following steps: (1) Fragment, end repair, phosphorylation and A-tailing of the nucleic acid sample to be tested are performed to obtain linear double-stranded nucleic acid fragments with dA tails at both ends; (2) Provide a hairpin connector with a pre-annealed stem-loop structure, wherein the hairpin connector is a single-stranded DNA, the 5' end and the 3' end are anti-complementary to form a double-stranded stem, the middle loop region contains a unique molecular tag composed of random sequences, and the 5' end has a phosphate group; (3) The hairpin connector is connected to the linear double-stranded nucleic acid fragment under the action of ligase, so that the two strands of the double-stranded nucleic acid fragment are circularized through the hairpin connector to form a single-stranded covalently closed circular template. (4) Using the single-stranded circular template as a template, and using RCA primers complementary to the hairpin connector circular region sequence, rolling circle amplification is performed under the action of strand displacement polymerase to generate a linear RCA product containing multiple tandem repeat sequences, each repeat unit containing the target nucleic acid sequence and the unique molecular tag; (5) Using the linear RCA product as a template, superbranched rolling circle amplification was performed using target-specific primers and universal primers complementary to the hairpin adapter sequence to obtain the HRCA product; (6) Perform nanopore sequencing on the HRCA product; (7) Identify tandem repeat units from sequencing reads, extract the target sequence and unique molecular tag of each repeat unit, obtain the consistent sequence through multicopy consistency correction, and use the corrected unique molecular tag to count molecules.
2. The method according to claim 1, characterized in that, The nucleotide sequence of the hairpin connector is shown in SEQ ID NO:1, where N is any one of A, T, C, and G, and NNNNNN is the unique molecular tag.
3. The method according to claim 1, characterized in that, The nucleotide sequence of the RCA primer is shown in SEQ ID NO:
2.
4. The method according to claim 3, characterized in that, The 3' terminal nucleotides of the RCA primers are modified with thiophosphate.
5. The method according to claim 1, characterized in that, Step (3) also includes digesting the uncircularized linear nucleic acid molecule with an exonuclease and then denaturing the ligation product by heat treatment to allow it to unwind and self-circulate.
6. The method according to claim 1, characterized in that, The strand displacement polymerase is phi29 DNA polymerase.
7. The method according to claim 1, characterized in that, The 3' end of the target-specific primer is modified with thiophosphate.
8. The method according to claim 1, characterized in that, The nanopore sequencing was performed using Oxford Nanopore sequencing.
9. The method according to claim 1, characterized in that, The fragment length obtained in step (1) is 100-1000 bp, preferably 300-500 bp.
10. A single-stranded circular DNA template, characterized in that, It is a covalently closed single-chain molecule containing the following structural units: (i) A single-stranded sequence of a nucleic acid fragment or its cDNA fragment derived from the sample to be tested; (ii) A hairpin connector sequence containing a unique molecular tag, covalently connected to both ends of the single-stranded sequence, the hairpin connector sequence being shown in SEQ ID NO:1, wherein NNNNNN is the unique molecular tag.
11. The single-stranded circular DNA template according to claim 10, characterized in that, The template is prepared by the method described in any one of claims 1-9.
12. A kit for targeted nanopore sequencing, characterized in that, Include: (i) A hairpin connector as shown in SEQ ID NO:1, wherein NNNNNN is a random molecular tag; (ii) RCA primers as shown in SEQ ID NO:2; (iii) Target-specific primers and universal primers complementary to the hairpin connector sequence; (iv) DNA polymerases with strand displacement activity.
13. The reagent kit according to claim 12, characterized in that, It also includes the DNA polymerase with strand displacement activity as phi29 DNA polymerase, SD DNA polymerase, and bst DNA polymerase, and the 3' end of the RCA primer is modified with phosphate thioester.