Tandem multicopy nucleic acid molecule construction method and use thereof

By performing A-addition to the 3' end of template nucleic acid molecules and strand substitution reactions, tandem multi-copy nucleic acid molecules were constructed, solving the problem of high error rates in long-read sequencing and achieving efficient and accurate sequencing library construction and data analysis.

WO2025251175A9PCT designated stage Publication Date: 2026-04-23MGI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
MGI TECH CO LTD
Filing Date
2024-06-03
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Existing long-read sequencing technologies suffer from high error rates, especially Oxford nanopore sequencing and Pacific Bio's real-time sequencing, which affect their application in medical research and clinical diagnosis. Furthermore, existing methods for generating multi-copy nucleic acid molecules are inefficient and suffer from significant template loss.

Method used

By adding an A to the 3' end of the template nucleic acid molecule, ligating the adapter, performing a strand displacement reaction, denaturation, and annealing, a tandem multi-copy nucleic acid molecule can be constructed, simplifying the reaction system and improving sequencing accuracy.

Benefits of technology

It significantly reduces sequencing error rates, improves sequencing accuracy, enhances the compatibility and efficiency of sequencing library construction, and is suitable for constructing high-quality sequencing libraries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024097087_23042026_PF_FP_ABST
    Figure CN2024097087_23042026_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a tandem multicopy nucleic acid molecule construction method and the use thereof. Said construction method comprises: performing A-tailing treatment on the 3' end of a template nucleic acid molecule, so as to obtain a nucleic acid molecule having a 3' A-tail; performing linker ligation treatment on the nucleic acid molecule having the 3' A-tail; performing first strand displacement reaction treatment on a linker ligation treatment product; performing denaturation and annealing treatment on a product of the first strand displacement reaction; and in the presence of a strand displacement primer, a product of the denaturation and annealing treatment undergoing a second strand displacement reaction treatment. In the linker ligation treatment product, the 5' end of the nucleic acid molecule having the 3' A-tail is linked to the 3' end of a linker sequence, and the 3' end of the nucleic acid molecule having the 3' A-tail is not linked to the 5' end sequence of the linker sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Methods for constructing tandem multicopy nucleic acid molecules and their applications Technical Field

[0001] This application relates to the field of biotechnology, specifically to a method for constructing tandem multicopy nucleic acid molecules and its application, and more specifically, to a method for constructing tandem multicopy nucleic acid molecules and its application, nucleic acid molecules, a method for constructing nucleic acid libraries, sequencing libraries, sequencing methods, and reagent kits. Background Technology

[0002] In the field of life science research, sequencing technology has become one of the most commonly used and important research methods. Since the Human Genome Project, gene sequencing has extensively influenced the research methods in life sciences, and the genomes of various model species are continuously being sequenced and analyzed in laboratories around the world. Sequencing technologies can be classified into short-read sequencing and long-read sequencing based on read length. Different read length sequencing technologies have different advantages in different applications, with long-read sequencing mainly represented by Oxford Nanopore Technology's nanopore sequencing technology and Pacific Biosciences' single-molecule real-time sequencing technology. Long-read sequencing can produce sequencing fragments of sufficient length, greatly promoting the development of fields such as genome assembly and variant detection. However, both Oxford nanopore sequencing and single-molecule real-time sequencing have extremely high error rates (close to 10%), affecting the accuracy of analytical results and limiting their application in medical research and clinical diagnosis.

[0003] To address the high error rate of long-read sequencing, Pacific Life Sciences developed Hi-Fi sequencing technology. This technology uses dumbbell-shaped adapters attached to both ends of the template to perform multiple cyclic sequencing runs on the positive and negative strands of the same molecule, followed by cross-correction to generate a consistent sequence, improving sequencing accuracy to 99.99%. Oxford Nanopore developed 2D sequencing technology that simultaneously sequences the positive and negative sequences of a single strand. It corrects sequencing errors in two separate nanopore sequencing runs, but the probability of the positive and negative strands entering the same nanopore for sequencing simultaneously (<60%) is low, and its sequencing quality only reaches 95%, far from the requirements for clinical testing. Generating a consistent sequence remains the best method for improving nanopore sequencing accuracy. Since nanopore sequencing cannot perform cyclic sequencing, it requires generating physically multiple copies of the molecule to obtain a consistent sequence. Currently, methods based on circularization and rolling circle amplification (RoBA) are used to prepare physically tandem multiple copies of the molecule to generate a consistent sequence and correct sequencing errors. However, this method is cumbersome, has low circularization efficiency, and suffers from significant template loss.

[0004] Therefore, methods for generating multi-copy nucleic acid molecules still need improvement.

[0005] Summary of the Invention

[0006] The purpose of this application is to address at least one of the problems of the prior art. To this end, this application proposes a method for constructing tandem repeat multiple-copy nucleic acid molecules.

[0007] Specifically, this application provides the following technical solution:

[0008] In a first aspect, this application proposes a method for constructing tandemly repeated multiple-copy nucleic acid molecules. According to embodiments of this application, the method includes: adding an A to the 3' end of a template nucleic acid molecule to obtain a nucleic acid molecule with a 3'-A tail; ligating the nucleic acid molecule with the 3'-A tail to a linker; performing a first-strand substitution reaction on the linker ligation product; denaturing and annealing the first-strand substitution reaction product; and performing a second-strand substitution reaction on the denatured and annealed product in the presence of a strand substitution primer; wherein, in the linker ligation product, the 5' end of the nucleic acid molecule with the 3'-A tail is connected to the 3' end of the linker sequence, and the 3' end of the nucleic acid molecule with the 3'-A tail is not connected to the 5' end sequence of the linker sequence. The aforementioned method can be used to conveniently and efficiently construct tandemly repeated multiple-copy nucleic acid molecules. In some examples of this application, this method is even more efficient for constructing long double-stranded DNA molecules.

[0009] In a second aspect of this application, a nucleic acid molecule is proposed. According to embodiments of this application, the nucleic acid molecule is constructed using the method described in the first aspect of this application. In some examples of this application, high-quality sequencing libraries can be efficiently constructed based on the aforementioned nucleic acid molecule.

[0010] In a third aspect, this application proposes a method for constructing a nucleic acid library. According to embodiments of this application, the method includes: adding an A to the 3' end of a template nucleic acid molecule to obtain a nucleic acid molecule with a 3'-A tail; ligating the nucleic acid molecule with the 3'-A tail to a linker; performing a first-strand substitution reaction on the linker ligation product; denaturing and annealing the first-strand substitution reaction product; performing a second-strand substitution reaction on the denatured and annealed product in the presence of a strand substitution primer; and constructing a sequencing library from the second-strand substitution reaction product to obtain the nucleic acid library; wherein, in the linker ligation product, the 5' end of the nucleic acid molecule with the 3'-A tail is connected to the 3' end of the linker sequence, and the 3' end of the nucleic acid molecule with the 3'-A tail is not connected to the 5' end sequence of the linker sequence. In some examples of this application, this method can rapidly construct high-quality sequencing libraries and has stronger compatibility for subsequent sequencing and data analysis.

[0011] In a fourth aspect, this application provides a sequencing library. According to embodiments of this application, the sequencing library is constructed using the method described in the third aspect of this application. In some examples of this application, sequencing using the above-described sequencing library can significantly reduce the sequencing error rate.

[0012] In a fifth aspect of this application, a sequencing method is proposed. According to embodiments of this application, the method includes: performing sequencing analysis on the sequencing library described in the fourth aspect or a sequencing library constructed using the method described in the third aspect to obtain a nucleic acid sequence to be tested. In some examples of this application, this method can generate high-quality sequencing data and significantly reduce the sequencing error rate.

[0013] In a sixth aspect, this application provides a kit. According to embodiments of this application, the kit is used to implement the above-described tandem multiple-copy nucleic acid molecule construction method, nucleic acid library construction method, or sequencing method. In some examples of this application, the aforementioned kit can be used for efficient and portable preparation of the above-described nucleic acid molecules.

[0014] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0015] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:

[0016] Figure 1 is a schematic diagram of the PacBio HiFi sequencing principle provided in the present application.

[0017] Figure 2 is a schematic diagram of the Oxford Nanopore sequencing principle provided in the present application.

[0018] Figure 3 is a schematic diagram of the neck ring connector and primers provided in a specific embodiment of this application;

[0019] Figure 4 is a schematic diagram of the DNA molecule construction process provided in a specific embodiment of this application;

[0020] Figure 5 is a schematic diagram of the DNA molecular nanopore library preparation process provided in the specific embodiments of this application.

[0021] The above figures are illustrative only and are not intended to be limiting. In the figures, for illustrative purposes, the dimensions of some components may be exaggerated and not drawn to scale. These dimensions and relative dimensions do not necessarily correspond to an actual representation in practice. Detailed Implementation

[0022] In this document, unless otherwise stated, the singular forms “a,” “an,” etc., include plural referents (more than one); “a group” or “a plurality” refers to two or more.

[0023] In this document, unless otherwise stated, the terms “comprising” or “including” are open-ended expressions, meaning they include the contents specified in this invention but do not exclude other aspects.

[0024] In this document, unless otherwise stated, the terms “first,” “second,” “third,” “fourth,” etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated; features defined with “first,” “second,” etc., may explicitly or implicitly include one or more of the stated features.

[0025] In this article, unless otherwise stated, nucleotide sequences are written from left to right in the 5' to 3' direction.

[0026] Currently, the mainstream long-read sequencing methods (also known as third-generation sequencing) include PacBio HiFi sequencing and Oxford Nanopore sequencing. The principle of PacBio HiFi sequencing, as shown in Figure 1, involves the continuous cycling of polymerase on a circular library to achieve repeated sequencing of the same sequence, thereby reducing the base recognition error rate. The principle of Oxford Nanopore sequencing, as shown in Figure 2, involves attaching a double-stranded linker containing motor proteins to one end of the target sequence and a neck-loop adapter to the other end, enabling the unwinding and sequencing of double-stranded DNA. During sequencing, the end containing the motor proteins is sequenced first. After the positive strand is sequenced, the direction is changed through the neck-loop adapter, and the negative strand is sequenced. Therefore, each target sequence is sequenced a maximum of twice.

[0027] Compared to PacBio HiFi, Oxford Nanopore sequencing has a lower ability to sequence the same sequence multiple times. To address this issue, a method using tandem repeats has been proposed. First, the double-stranded DNA fragment in the sample is circularized to obtain single-stranded circular DNA. This single-stranded circular DNA is then used as a template for rolling circle amplification, generating tandem repeats. Subsequently, the standard Nanopore adapter ligation step is performed. However, this method for generating tandem repeats has several problems. For example, the efficiency of circularizing linear double-stranded DNA into single-stranded circular DNA is very low (typically no more than 30%). This means that the circularization step results in significant sample loss, leading to a substantial loss of template sequence information, severely impacting sequencing efficiency and accuracy.

[0028] To improve sequencing efficiency and accuracy, this application proposes a method for constructing nucleic acid molecules, comprising:

[0029] Step 1: Add an A tail to the 3' end of the template nucleic acid molecule to obtain a nucleic acid molecule with a 3'-A tail;

[0030] In some examples of this application, the template nucleic acid molecule is double-stranded DNA. The aforementioned double-stranded DNA is selected from at least one of full-length cDNA, mitochondrial DNA, and chromosomal DNA. In some preferred examples of this application, the aforementioned double-stranded DNA is not less than 100 bp.

[0031] The aforementioned double-stranded DNA can be obtained through reverse transcription of RNA or through the synthesis of single-stranded DNA.

[0032] In some examples of this application, for template nucleic acid molecules with non-pointed ends, the method further includes: performing end repair processing on the template nucleic acid molecules.

[0033] Step 2: Connect nucleic acid molecules with a 3'-A tail using adapters;

[0034] In some examples of this application, the adapter is selected from a loop adapter, which includes a loop structure region and a double-strand structure region. In some examples of this application, the loop adapter includes at least a portion of the sequence of the strand substitution primer. In a specific example of this application, the structure of the loop adapter is shown in Figure 3. The loop adapter provides a starting point for a strand-side reaction, enabling DNA polymerase to perform a strand substitution reaction along the template strand.

[0035] In some examples of this application, the adapter includes a fixed sequence, which includes a primer-binding sequence and / or a molecular tag sequence (UMI). In some examples of this application, the aforementioned molecular tag sequence is located in the neck loop region or the double-stranded region. The molecular tag sequence is used to determine the copy number of the DNA molecule; the fixed sequence is used to form the neck loop structure and improve adapter stability; the primer-binding sequence is used to perform complementary pairing with the strand substitution primer and guide polymerase binding.

[0036] Step 3: Perform a first-strand substitution reaction on the adapter ligation product; wherein, in the adapter ligation product, the 5' end of the double-stranded DNA molecule with blunt ends and 3'-A tail is connected to the 3' end of the adapter sequence, and the 3' end of the double-stranded DNA molecule with blunt ends and 3'-A tail is not connected to the 5' end sequence of the adapter sequence.

[0037] In some examples of this application, the strand substitution reaction is carried out in the presence of a strand substitution polymerase. The aforementioned strand substitution polymerase includes at least one of Bst DNA polymerase, Bsu DNA polymerase, Phi29 DNA polymerase, and the Klenow fragment (exo-).

[0038] Step 4: Denature and anneal the product of the first chain substitution reaction;

[0039] In some examples of this application, the chain substitution primer may or may not carry a neck loop sequence (Figure 1).

[0040] Step 5: In the presence of the chain substitution primer, the denatured and annealed products are subjected to a second chain substitution reaction.

[0041] The above method requires only one primer to construct the entire tandem repeat multiple copy nucleic acid molecule. The reaction system has a single component, making system preparation simpler, and the obtained products can be used in subsequent steps. Furthermore, the obtained tandem repeat multiple copy nucleic acid molecules are more compatible with the construction and sequencing of long-length sequencing libraries.

[0042] For ease of understanding, referring to Figure 4, the above nucleic acid molecule construction method is described in detail using double-stranded DNA as a template nucleic acid molecule as a specific example.

[0043] 1. End repair and 3' A-tail addition (for TA cloning) are performed on the template double-stranded DNA molecule to obtain a double-stranded DNA molecule with blunt ends and a 3'-A tail. The end repair and 3'-A tail steps can be performed using routine laboratory methods, and this application does not impose specific limitations.

[0044] It should be noted that the end repair step is not used for double-stranded DNA molecules with blunt ends.

[0045] 2. The obtained double-stranded DNA molecules with blunt ends and 3'-A tails are ligated using adapters (e.g., neck loop adapters). In the adapter ligation product, the 5' end of the double-stranded DNA molecule with blunt ends and 3'-A tails is linked to the 3' end of the adapter sequence, while the 3' end of the double-stranded DNA molecule with blunt ends and 3'-A tails is not linked to the 5' end of the adapter sequence. Those skilled in the art will understand that, through the complementary pairing type of the DNA molecules, a neck loop-double-stranded DNA molecule-neck loop molecule is formed.

[0046] 3. In the presence of strand displacement DNA polymerase, a strand displacement reaction is performed on the neck loop-double-stranded DNA molecule-neck loop molecule to obtain a double-stranded DNA molecule with complete adapters at both ends.

[0047] 4. Denaturation and annealing reactions are performed on double-stranded DNA molecules with complete adapters to convert them into single-stranded molecules with complete adapters at both ends.

[0048] 5. In the presence of second-strand replacement primers and strand replacement DNA polymerase, the denatured and annealed products were subjected to a second-strand replacement reaction.

[0049] In another aspect of this application, a nucleic acid molecule is proposed. This nucleic acid molecule is obtained through any of the example methods described above. In some examples of this application, constructing sequencing libraries based on the aforementioned nucleic acid molecule can significantly improve library construction efficiency.

[0050] In some examples of this application, the nucleic acid molecule is double-stranded DNA.

[0051] In another aspect of this application, a method for constructing a nucleic acid library is proposed, the method comprising:

[0052] Step 1: Add an A tail to the 3' end of the template nucleic acid molecule to obtain a nucleic acid molecule with a 3'-A tail;

[0053] In some examples of this application, the template nucleic acid molecule is double-stranded DNA. The aforementioned double-stranded DNA is selected from at least one of full-length mRNA, mitochondrial DNA, and chromosomal DNA.

[0054] In some examples of this application, the double-stranded DNA is selected from at least one of full-length cDNA, mitochondrial DNA, and chromosomal DNA. In some preferred examples of this application, the aforementioned double-stranded DNA is not less than 100 bp.

[0055] The aforementioned double-stranded DNA can be obtained through reverse transcription of RNA or through the synthesis of single-stranded DNA.

[0056] In some examples of this application, for template nucleic acid molecules with non-pointed ends, the method further includes: performing end repair processing on the template nucleic acid molecules.

[0057] Step 2: Connect nucleic acid molecules with a 3'-A tail using adapters;

[0058] In some examples of this application, the adapter is selected from a loop adapter, which includes a loop structure region and a double-strand structure region. In some examples of this application, the loop adapter includes at least a portion of the sequence of the strand substitution primer. In one example of this application, the loop adapter structure is shown in Figure 3. The loop adapter provides a starting point for a strand-side reaction, enabling DNA polymerase to perform a strand substitution reaction along the template strand.

[0059] In some examples of this application, the adapter includes a fixed sequence, which includes a primer-binding sequence and / or a molecular tag sequence (UMI). In some examples of this application, the aforementioned molecular tag sequence is located in the neck loop region or the double-stranded region. The molecular tag sequence is used to determine the copy number of the DNA molecule; the fixed sequence is used to form the neck loop structure and improve adapter stability; the primer-binding sequence is used to perform complementary pairing with the strand substitution primer and guide polymerase binding.

[0060] Step 3: Perform a first-strand substitution reaction on the adapter ligation product; wherein, in the adapter ligation product, the 5' end of the double-stranded DNA molecule with blunt ends and 3'-A tail is connected to the 3' end of the adapter sequence, and the 3' end of the double-stranded DNA molecule with blunt ends and 3'-A tail is not connected to the 5' end sequence of the adapter sequence.

[0061] In some examples of this application, the strand substitution reaction is carried out in the presence of a strand substitution polymerase. The aforementioned strand substitution polymerase includes at least one of Bst DNA polymerase, Bsu DNA polymerase, Phi29 DNA polymerase, and the Klenow fragment (exo-).

[0062] Step 4: Denature and anneal the product of the first chain substitution reaction;

[0063] In some examples of this application, the chain substitution primer may or may not carry a neck loop sequence (Figure 1).

[0064] Step 5: In the presence of the chain substitution primer, the denatured and annealed products are subjected to a second chain substitution reaction.

[0065] Step 6: Perform sequencing library construction processing on the product of the second-strand substitution reaction to obtain the nucleic acid library.

[0066] In some examples of this application, the sequencing library construction process includes phosphorylation of the product of the second strand substitution reaction. The phosphorylation is performed in the presence of a phosphokinase. In one specific example of this application, the phosphokinase is selected from T4 phosphokinase.

[0067] In some examples of this application, the sequencing library construction process further includes: adding an alpha (A) to the 3' end of the product of the second strand substitution reaction. This alpha addition provides binding sites for the sequencing adapters, enabling nucleic acid molecules to be effectively ligated to them. Furthermore, the alpha addition can reduce ligation bias, improve ligation efficiency, and increase the success rate of library construction.

[0068] In some examples of this application, the sequencing library construction process further includes: ligating the product of the second strand substitution reaction to a sequencing adapter. After the above phosphorylation and A addition treatment, the double-stranded DNA molecules are further ligated to the sequencing adapter under the action of a ligase to obtain a sequencing library.

[0069] In some examples of this application, the sequencing adapter is a nanopore sequencing adapter. In some preferred examples of this application, the nanopore sequencing adapter carries a motor protein. In some preferred examples of this application, the nanopore sequencing adapter also carries a tether protein.

[0070] The sequencing libraries obtained using this method avoid the generation of intermediate products, and the resulting final products are of a single composition, all of which can be used for subsequent sequencing analysis steps.

[0071] For ease of understanding, referring to Figure 5, the above nucleic acid library construction method is described in detail using double-stranded DNA as a template nucleic acid molecule as a specific example.

[0072] 1. End repair and 3' A-tail addition (for TA cloning) are performed on the template double-stranded DNA molecule to obtain a double-stranded DNA molecule with blunt ends and a 3'-A tail. The end repair and 3'-A tail steps can be performed using routine laboratory methods, and this application does not impose specific limitations.

[0073] It should be noted that the end repair step is not used for double-stranded DNA molecules with blunt ends.

[0074] 2. The obtained double-stranded DNA molecules with blunt ends and 3'-A tails are ligated using adapters (e.g., neck loop adapters). In the adapter ligation product, the 5' end of the double-stranded DNA molecule with blunt ends and 3'-A tails is linked to the 3' end of the adapter sequence, while the 3' end of the double-stranded DNA molecule with blunt ends and 3'-A tails is not linked to the 5' end of the adapter sequence. Those skilled in the art will understand that, through the complementary pairing type of the DNA molecules, a neck loop-double-stranded DNA molecule-neck loop molecule is formed.

[0075] 3. In the presence of strand displacement DNA polymerase, a strand displacement reaction is performed on the neck loop-double-stranded DNA molecule-neck loop molecule to obtain a double-stranded DNA molecule with complete adapters at both ends.

[0076] 4. Denaturation and annealing reactions are performed on double-stranded DNA molecules with complete adapters to convert them into single-stranded molecules with complete adapters at both ends.

[0077] 5. In the presence of second-strand replacement primers and strand replacement DNA polymerase, the denatured and annealed products were subjected to a second-strand replacement reaction.

[0078] 7. Phosphorylate and add A to the 3' end of the double-stranded DNA molecule obtained in step 6.

[0079] 8. Add sequencing adapters (containing motor proteins) to the double-stranded DNA molecules obtained in step 7 to obtain the sequencing library.

[0080] In another aspect of this application, a sequencing library is proposed, which is constructed using the methods described in any of the examples above. In some examples of this application, sequencing of the above-described sequencing library can reduce sequencing errors.

[0081] In another aspect, this application proposes a sequencing method, which includes: performing sequencing analysis on the sequencing library described above or a sequencing library constructed by the above method to obtain the nucleic acid sequence to be tested. In some examples of this application, this method can generate high-quality sequencing data and significantly reduce the sequencing error rate.

[0082] In some examples of this application, the sequencing analysis process includes: determining the number of copies of each original DNA molecule included in the original sequencing reads based on the fixed sequence and molecular tag sequence in the adapter sequence, and dividing the original sequencing reads into DNA sequences with the corresponding number of copies.

[0083] In some examples of this application, the sequencing analysis process includes: performing mutual correction on the multiple copies of DNA sequence contained in the sequencing reads to obtain corrected DNA sequences. This correction eliminates sequencing errors.

[0084] In some examples of this application, the sequencing analysis further includes: combining the corrected DNA sequences to determine a homologous sequence; and comparing the homologous sequence with a reference sequence to determine the nucleic acid sequence to be tested.

[0085] It should be noted that the aforementioned reference sequences are established sequences. These can be pre-assembled DNA and / or RNA sequences determined by oneself, or publicly available DNA and / or RNA sequences determined by others. They can be any reference template from the biological category of the sample source individual / target individual, such as all or at least a portion of publicly available genome assembly sequences from the same biological category. If the sample source individual or target individual is human, its genome reference sequence (also called the reference genome or reference chromosome set) can be selected from human reference genomes provided by the UCSC, NCBI, or ENSEMBL databases, such as HG19, HG38, GRCh36, GRCh37, GRCh38, etc. Those skilled in the art can understand the correspondence between the above reference genome versions through the database descriptions and select the version to use. Furthermore, a resource library containing more reference sequences can be pre-configured. For example, before comparison, sequences that are closer to or have more specific characteristics can be selected or assembled based on factors such as the target individual's sex, ethnicity, and region to serve as reference sequences, which helps to obtain more accurate sequence analysis results subsequently.

[0086] In some examples of this application, the sequencing analysis further includes: performing bioinformatics analysis on the obtained nucleic acid sequence to be tested, such as gene function prediction, translation and structural analysis of protein encoding, etc.

[0087] In another aspect, this application provides a kit for implementing the above-described tandem multiple-copy nucleic acid molecule construction method, nucleic acid library construction method, or sequencing method. In some examples of this application, the aforementioned nucleic acid molecules can be rapidly prepared based on the kit, thereby enabling nucleic acid library construction and / or sequencing.

[0088] In some examples of this application, the aforementioned kit includes at least one of adapter, DNA polymerase, DNA ligase, dNTP, and strand displacement polymerase.

[0089] In some examples of this application, the kit further includes: sequencing reagents and instructions. The foregoing instructions describe the corresponding steps for implementing the above-described tandem multiple-copy nucleic acid molecule construction method, nucleic acid library construction method, or sequencing method.

[0090] Embodiments of the present invention will now be described in more detail, examples of which are illustrated in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the invention. Where specific techniques or conditions are not specified in the embodiments, they are performed in accordance with techniques or conditions described in the literature in the art or according to the product instructions. Reagents or instruments used, unless otherwise specified, are all commercially available conventional products.

[0091] Example 1: Construction and sequencing of DNA nanopore sequencing libraries

[0092] The experimental protocol of this embodiment is as follows: A single product of human mitochondria was amplified, with a length of approximately 16 kb. The DNA was used to prepare libraries according to the method of this application and conventional methods. The obtained sequencing libraries were sequenced on a nanopore sequencer, and the accuracy of the sequencing was evaluated.

[0093] The conventional nanopore library was prepared strictly according to the Nanopore Ligation sequencing gDNA V14-human sample (N50 10kb) on PromethION (SQK-LSK114) instruction manual.

[0094] 1. Add A to the end

[0095] Terminal addition of A reaction system: as shown in Table 1;

[0096] Table 1. Terminal A Addition Reaction System

[0097] Reaction conditions: Place the above reaction system on a PCR instrument and incubate at 75℃ for 20 min. After the reaction, purify the product using 1.0×AMPure magnetic beads, and finally dissolve the purified product in 13 μl of elution buffer.

[0098] Of the reagents mentioned above, Taq hot-start DNA polymerase and 10X standard Taq reaction buffer were from NEB (product number M0495S); dATP solution was from NEB (product number N0440S); and AMPure XP Beads were from Beckman Coulter (product number A63882).

[0099] 2. Joint connection and chain displacement reaction:

[0100] 2.1 Connector Connection Reaction

[0101] Prepare the adapter ligation reaction system using the DNA obtained in step 1 according to Table 2;

[0102] Table 2 Joint Connection Reaction System

[0103] In the above reagents, the neck ring connector sequence is:

[0104] The underlined N part is the molecular tag, and the bolded part is the primer binding site; T4 DNA ligase and 10X DNA Ligase Reaction Buffer are from NEB, catalog number M0202V;

[0105] 2.2 Joint connection reaction conditions:

[0106] The above reaction system was placed on a PCR instrument and reacted at 20°C for 15 min. After the reaction, the product was purified using 1.0×AMPure magnetic beads, and finally dissolved in 15 μl of elution buffer.

[0107] 2.3 Chain displacement reaction system

[0108] Prepare a chain displacement reaction system by connecting the reaction products with the connector according to Table 3;

[0109] Table 3 Chain displacement reaction system

[0110] All the reagents mentioned above were from NEB: Bst 2.0 DNA Polymerase and 10X Isothermal Amplification Buffer (Catalog No. M0537S); 100mM MgSO4 (Catalog No. B1003S); and 10mM dNTP Mix (Catalog No. N0447).

[0111] 2.4 Chain displacement reaction conditions

[0112] The above reaction system was placed on a PCR instrument and reacted at 65°C for 10 min. After the reaction, the product was purified using 1.0×AMPure magnetic beads, and finally dissolved in 20 μl of elution buffer.

[0113] 3. Denaturation and chain displacement reactions

[0114] The purified product obtained in step 2 was placed on a PCR instrument and reacted at 95°C for 3 min, followed by 65°C for 3 min. After the reaction, the product was placed in an ice box, and the chain displacement reaction system was prepared according to Table 4.

[0115] Table 4 Chain displacement reaction system

[0116] The Primer sequence is: TCGTAGCCATGTCGTTC (SEQ ID NO:2); the reagents in Table 4 are all from NEB: Bst 2.0 DNA Polymerase and 10X Isothermal Amplification Buffer (catalog number M0537S); 100mM MgSO4 (catalog number B1003S); and 10mM dNTP Mix (catalog number N0447).

[0117] The above reaction system was placed on a PCR instrument and reacted at 65℃ for 20 min. After the reaction, the product was purified using 1.0×AMPure magnetic beads, and finally dissolved in 15 μl of elution buffer.

[0118] 4' and 5' end phosphorylation and 3' end A-tailing

[0119] Prepare the phosphorylation and A-tailing reaction system for the DNA obtained in step 3 according to Table 5;

[0120] Table 5 Phosphorylation and A-tailing reaction systems

[0121] The above reagents are from NEB: T4 phosphokinase and T4 PNK Reaction Buffer (10X) are catalog number M0201V; Taq hot-start DNA polymerase is from NEB, catalog number M0495S; dATP solution is from NEB, catalog number N0440S.

[0122] The above reaction system was placed on a PCR instrument and incubated at 37°C for 15 min; then at 65°C for 15 min. After the reaction, the product was purified using 1.0×AMPure magnetic beads, and finally dissolved in 60 μl of elution buffer.

[0123] 5. Addition of nanopore sequencing adapters

[0124] Prepare nanopore sequencing adapter ligation systems using the DNA obtained in step 4 according to Table 6.

[0125] Table 6. Nanopore sequencing adapter ligation system

[0126] Of the reagents mentioned above, NEB Next Quick T4 DNA Ligase is from NEB Corporation, catalog number E6056; Ligation Buffer and Adapter are from Nanopore Corporation, catalog number SQK-LSK114.

[0127] The above reaction system was placed on a PCR instrument and incubated at 25°C for 10 min; then at 65°C for 20 min. After the reaction, the product was purified using 0.4X AMPure magnetic beads, and finally dissolved in 25 μl of elution buffer. The library was quantified using the HS Qubit ssDNA kit.

[0128] 6. Sequencing

[0129] The obtained library was sequenced on a PromethION Flow Cell instrument using the R10.4.1 flow cell, following the standard instruction manual. The Ligation Sequencing Kit V14 (SQK-LSK114) was used.

[0130] 7. Information Analysis

[0131] The data obtained by conventional methods are compared using Minimap and visualized using IGV; the data obtained by the invention method are first subjected to consistency sequence generation, then compared using Minimap and visualized using IGV.

[0132] 8. Results Analysis

[0133] Table 7. Statistics on the length of raw sequencing data

[0134] The above results show that, compared with conventional methods, the raw sequencing reads obtained based on the method of this application are longer. The proportion of reads longer than 32kb represents the proportion of data with more than 2 sequencing runs (including positive and negative strands) in 93.3%; the proportion of reads longer than 64kb represents the proportion of data with more than 4 sequencing runs (including positive and negative strands) in 66.2%, which also indicates that the proportion of tandem repeats with more than 2 copies is 66.2%; there are also reads longer than 128kb, accounting for 30.6%.

[0135] Example 2: Construction and sequencing of RNA nanopore sequencing libraries

[0136] The experimental protocol for this embodiment is as follows: 10,000 copies of the COVID-19 standard (Twist Synthetic SARS-CoV-2 RNA controls, cat#102024:Control 2 MN908947.3) were added. The DNA was prepared into libraries according to both the inventive method and conventional methods. The obtained sequencing libraries were then sequenced using a nanopore sequencer. The 3' end fragment of the COVID-19 standard was approximately 5 kb in length. The accuracy of the sequencing was evaluated.

[0137] The conventional nanopore library was operated strictly in accordance with the Nanopore cDNA-PCR Sequencing Kits instruction manual.

[0138] 1. cDNA synthesis

[0139] Oligod T was used to synthesize full-length cDNA from total RNA and then performed template conversion. The system (Table 8) and conditions are as follows;

[0140] Table 8 Template Conversion System

[0141] The TSO sequence is: GACATGGCTACGATCCGACTT (SEQ ID NO:3); wherein, the SEQ ID NO:3 sequence contains 3 riboguanosines (rGrGr+G) at the end;

[0142] Oligo dT sequence is: GCATCCATTAGTTAGGCTAGTTTTTTTTTTTTTTTTTTVN (SEQ ID NO: 4);

[0143] The above reaction system was placed on a PCR instrument and incubated at 42°C for 60 min. After the reaction, the product was purified using 1.0×AMPure magnetic beads, and finally dissolved in 21 μl of elution buffer. Reverse transcriptase was purchased from Thermo Fisher Scientific, catalog number 18090200.

[0144] 2. PCR reaction

[0145] The PCR reaction system is shown in Table 9;

[0146] Table 9 PCR Reaction System

[0147] The PCR reaction procedure is shown in Table 10;

[0148] Table 10 PCR reaction procedure

[0149] After the reaction, the product was purified using 1.0×AMPure magnetic beads, and finally dissolved in 13 μl of elution buffer. The PCR amplification enzyme was obtained from NEB, cat#M0533S;

[0150] PCR Primer1: GACATGGCTACGATCCGACTT (SEQ ID NO: 5);

[0151] PCR Primer2: GCATCCATTAGTTAGGCTAG (SEQ ID NO: 6).

[0152] 3. Joint connection and chain displacement reaction:

[0153] Prepare the adapter ligation reaction system using the DNA obtained in step 2 according to Table 11;

[0154] Table 11 Joint Connection Reaction System

[0155] The neck ring connector sequence is as follows: The underlined N part is the molecular tag, and the bold part is the primer binding site;

[0156] The T4 DNA ligase and 10X DNA Ligase Reaction Buffer mentioned above are from NEB, catalog number M0202V;

[0157] The above reaction system was placed on a PCR instrument and incubated at 20°C for 15 min. After the reaction was complete, the product was purified using 1.0×AMPure magnetic beads, and finally dissolved in 15 μl of elution buffer.

[0158] Prepare a chain displacement reaction system using the purified product according to Table 12;

[0159] Table 12 Chain displacement reaction system

[0160] All the above reagents were from NEB: Bst 2.0 DNA Polymerase and 10X Isothermal Amplification Buffer (Catalog No. M0537S); 100mM MgSO4 (Catalog No. B1003S); and 10mM dNTP Mix (Catalog No. N0447).

[0161] The above reaction system was placed on a PCR instrument and incubated at 65°C for 10 min. After the reaction was complete, the product was purified using 1.0×AMPure magnetic beads, and finally dissolved in 20 μl of elution buffer.

[0162] 4. Denaturation and chain displacement reactions

[0163] Place the purified product from step 3 onto a PCR instrument and incubate at 95°C for 3 minutes, then at 65°C for 3 minutes. After the reaction is complete, place the container in an ice box.

[0164] Prepare strand displacement reaction systems for denatured DNA according to Table 13;

[0165] Table 13 Chain displacement reaction system

[0166] The sequence of Primer: TCGTAGCCATGTCGTTC (SEQ ID NO:8);

[0167] All the above reagents were from NEB: Bst 2.0 DNA Polymerase and 10X Isothermal Amplification Buffer (Catalog No. M0537S); 100mM MgSO4 (Catalog No. B1003S); and 10mM dNTP Mix (Catalog No. N0447).

[0168] The above reaction system was placed on a PCR instrument and incubated at 65°C for 20 min. After the reaction, the product was purified using 1.0×AMPure magnetic beads, and finally dissolved in 15 μl of elution buffer.

[0169] 5. Phosphorylation at the 5' end and the addition of an A tail at the 3' end

[0170] Prepare the phosphorylation and A-tailing reaction system using the DNA obtained in step 4 according to Table 14.

[0171] Table 14 Phosphorylation and A-tailing reaction systems

[0172] The above reagents are from NEB: T4 phosphokinase and T4 PNK Reaction Buffer (10X) are catalog number M0201V; Taq hot-start DNA polymerase is from NEB, catalog number M0495S; dATP solution is from NEB, catalog number N0440S.

[0173] The above reaction system was placed on a PCR instrument and incubated at 37°C for 15 min; then at 65°C for 15 min. After the reaction, the product was purified using 1.0×AMPure magnetic beads, and finally dissolved in 60 μl of elution buffer.

[0174] 6. Addition of nanopore sequencing adapters

[0175] Prepare nanopore sequencing adapter ligation systems using the DNA obtained in step 5 according to Table 15.

[0176] Table 15 Nanopore sequencing adapter ligation system

[0177] Of the reagents mentioned above, NEB Next Quick T4 DNA Ligase is from NEB Corporation, catalog number E6056; Ligation Buffer and Adapter are from Nanopore Corporation, catalog number SQK-LSK114.

[0178] The above reaction system was placed on a PCR instrument and incubated at 25°C for 10 min; then at 65°C for 20 min. After the reaction, the product was purified using 0.4X AMPure magnetic beads, and finally dissolved in 25 μl of elution buffer. The library was quantified using the HS Qubit ssDNA kit.

[0179] 7. Sequencing

[0180] The obtained library was sequenced on a PromethION Flow Cell instrument. The R10.4.1 flow cell was selected, and the operation was performed according to the standard instruction manual using Ligation Sequencing Kit V14 (SQK-LSK114).

[0181] 8. Information Analysis

[0182] The data obtained by conventional methods are compared using Minimap and visualized using IGV; the data obtained by the invention method are first subjected to consistency sequence generation, then compared using Minimap and visualized using IGV.

[0183] 9. Results Analysis

[0184] Table 16 Statistics on the length of raw sequencing data

[0185] The above results show that, compared with conventional methods, the raw sequencing reads obtained based on the method of this application are longer. The proportion of reads longer than 10kb represents the proportion of data with more than 2 sequencing runs (including positive and negative strands) in 92.3% of the data; the proportion of reads longer than 20kb represents the proportion of data with more than 4 sequencing runs (including positive and negative strands) in 64.1% of the data, which also indicates that the proportion of tandem repeats with more than 2 copies is 64.1%; there are also reads longer than 128kb, indicating that the proportion of tandem repeats with more than 4 copies is also 36.1%.

[0186] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0187] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.

Claims

1. A method for constructing tandem multiple-copy nucleic acid molecules, characterized in that, include: The template nucleic acid molecule is treated with A-addition at the 3' end to obtain a nucleic acid molecule with a 3'-A tail; Nucleic acid molecules with a 3'-A tail are ligated with adapters; The joint connection product is subjected to a first-chain substitution reaction treatment. The product of the first-chain substitution reaction was denatured and annealed. In the presence of chain substitution primers, the denatured and annealed products were subjected to a second chain substitution reaction. In the adapter ligation product, the 5' end of the nucleic acid molecule with the 3'-A tail is connected to the 3' end of the adapter sequence, while the 3' end of the nucleic acid molecule with the 3'-A tail is not connected to the 5' end sequence of the adapter sequence.

2. The method of claim 1, wherein, The template nucleic acid molecule is double-stranded DNA; Optionally, the double-stranded DNA is selected from at least one of full-length cDNA, mitochondrial DNA, and chromosomal DNA.

3. The method of claim 2, wherein, The double-stranded DNA was obtained through reverse transcription of RNA.

4. The method of claim 2, wherein, The double-stranded DNA molecule is obtained by synthesizing a second strand of single-stranded DNA.

5. The method of claim 1, wherein, Prior to the A-addition process, the procedure further includes: performing end-repair processing on the template nucleic acid molecule.

6. The method of claim 1, wherein, The connector is selected from the neck ring type connector.

7. The method of claim 6, wherein, The neck-loop connector includes at least a portion of the sequence of the chain-displacement primer.

8. The method of claim 6, wherein, The connector includes a molecular tag sequence.

9. The method of claim 1, wherein, The chain displacement reaction is carried out in the presence of chain displacement polymerase; Optionally, the strand substitution polymerase includes at least one of Bst DNA polymerase, Bsu DNA polymerase, Phi29 DNA polymerase, and Klenow fragment (exo-).

10. A nucleic acid molecule, characterized in that, It is obtained by constructing the method according to any one of claims 1 to 9.

11. The nucleic acid molecule of claim 10, wherein The nucleic acid molecule is double-stranded DNA.

12. A method of nucleic acid library construction, comprising, include: The template nucleic acid molecule is treated with A-addition at the 3' end to obtain a nucleic acid molecule with a 3'-A tail; Nucleic acid molecules with a 3'-A tail are ligated with adapters; The joint connection product is subjected to a first-chain substitution reaction treatment. The product of the first-chain substitution reaction was denatured and annealed. In the presence of chain substitution primers, the denatured and annealed products were subjected to a second chain substitution reaction. The product of the second-strand substitution reaction is subjected to sequencing library construction to obtain the nucleic acid library; In the adapter ligation product, the 5' end of the nucleic acid molecule with the 3'-A tail is connected to the 3' end of the adapter sequence, while the 3' end of the nucleic acid molecule with the 3'-A tail is not connected to the 5' end sequence of the adapter sequence.

13. The method of claim 12, wherein, The template nucleic acid molecule is double-stranded DNA; Optionally, the double-stranded DNA is selected from at least one of full-length cDNA, mitochondrial DNA, and chromosomal DNA.

14. The method of claim 13, wherein, The double-stranded DNA was obtained through reverse transcription of RNA.

15. The method of claim 13, wherein, The double-stranded DNA molecule is obtained by synthesizing a second strand of single-stranded DNA.

16. The method of claim 12, wherein, Prior to the A-addition process, the procedure further includes: performing end-repair processing on the template nucleic acid molecule.

17. The method of claim 12, wherein, The connector is selected from the neck ring type connector.

18. The method of claim 17, wherein, The neck-loop connector includes at least a portion of the sequence of the chain-displacement primer.

19. The method of claim 17, wherein, The connector includes a molecular tag sequence.

20. The method of claim 12, wherein, The chain displacement reaction is carried out in the presence of chain displacement polymerase; Optionally, the strand substitution polymerase includes at least one of Bst DNA polymerase, Bsu DNA polymerase, Phi29 DNA polymerase, and Klenow fragment (exo-).

21. The method of claim 12, wherein, The sequencing library construction process includes: phosphorylation of the product of the second strand substitution reaction.

22. The method of claim 21, wherein, The phosphorylation treatment was performed in the presence of phosphokinase.

23. The method of claim 21, wherein, The sequencing library construction process further includes: adding an A to the 3' end of the product of the second strand substitution reaction.

24. The method of claim 23, wherein, The sequencing library construction process further includes: performing sequencing adapter ligation on the product of the second strand substitution reaction.

25. The method of claim 23, wherein, The sequencing adapter is a nanopore sequencing adapter; Preferably, the nanopore sequencing adapter carries a motor protein; More preferably, the nanopore sequencing adapter also carries a tether protein.

26. A sequencing library, wherein, It is obtained by constructing the method according to any one of claims 12 to 25.

27. A sequencing method, comprising: include: The sequencing library described in claim 26 is subjected to sequencing analysis to obtain the nucleic acid sequence to be tested.

28. The method of claim 27, wherein, The sequencing analysis process includes: Based on the fixed sequence and molecular tag sequence in the adapter sequence, the number of copies of each original DNA molecule included in the original sequencing sequence is determined, and the original sequencing sequence is divided into DNA sequences with the corresponding number of copies.

29. The method of claim 28, wherein, The sequencing analysis process includes: The DNA sequences obtained from sequencing with multiple copy numbers are cross-corrected to obtain aligned and corrected DNA sequences.

30. The method of claim 28 or 29, wherein, The sequencing analysis further includes: Based on the corrected DNA sequence, a consistent sequence is determined by combining the sequences. Based on the reference sequence, the consistency sequence is compared to determine the nucleic acid sequence to be tested.

31. A kit comprising, The kit is used to implement the tandem multiple copy nucleic acid molecule construction method according to any one of claims 1 to 9, the nucleic acid library construction method according to any one of claims 12 to 25, or the sequencing method according to any one of claims 27 to 30.