Tandem multicopy nucleic acid molecule construction method and use thereof
By performing A-addition to the 3' end of template nucleic acid molecules and strand substitution reactions, tandem repeat multicopy nucleic acid molecules are constructed, solving the problems of high error rate and low efficiency in long read sequencing technology, and realizing efficient and accurate sequencing library construction and data analysis.
Patent Information
- Application Number
- PCT/CN2024/097087
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-03
- Publication Date
- 2025-12-11
AI Technical Summary
Existing long-read sequencing technologies suffer from high error rates. In particular, Oxford Nanopore sequencing has a low ability to sequence the same sequence multiple times, and existing methods are inefficient at generating tandem repeats of molecules, which limits sequencing accuracy and efficiency.
By adding an A to the 3' end of the template nucleic acid molecule, ligating the adapter, performing strand displacement reaction, denaturation and annealing, tandem repeat multicopy nucleic acid molecules can be constructed, simplifying the reaction system and improving the compatibility and efficiency of sequencing library construction.
It significantly reduces sequencing error rate, improves sequencing accuracy and efficiency, enhances sequencing library compatibility, increases read length and copy number, and improves sequencing data quality.
Smart Images

Figure CN2024097087_11122025_PF_FP_ABST
Abstract
Description
Method for constructing tandem multi-copy nucleic acid molecules and application thereof TECHNICAL FIELD
[0001] The present application relates to the field of biotechnology, in particular, the present application relates to a method for constructing tandem multi-copy nucleic acid molecules and application thereof, more particularly, the present application relates to a method for constructing tandem multi-copy nucleic acid molecules and application thereof, a nucleic acid molecule, a nucleic acid library construction method, a sequencing library, a sequencing method and a kit. BACKGROUND
[0002] In the field of life science research, sequencing technology has become one of the most commonly used and most important research means. Since the human genome project was launched, gene sequencing has widely affected the research methods of life sciences, and the genomes of various model species are being determined and analyzed in laboratories around the world. Sequencing technology can be classified into short read sequencing and long read sequencing according to the sequencing read length. Different read length sequencing technologies have different advantages in different applications, among which long read sequencing is mainly represented by Oxford Nanopore Technology's nanopore sequencing technology and Pacific Biosciences' single molecule real-time sequencing technology. Long read sequencing can generate sequencing fragments of sufficient length, greatly promoting the development of analysis fields such as genome assembly and variant detection. However, both Oxford Nanopore sequencing and single molecule real-time sequencing have a very high error rate (close to 10%), which affects the accuracy of the analysis results and limits their application in medical research and clinical diagnosis.
[0003] To solve the problem of high error rate in long read sequencing, Pacific Biosciences developed HIFI sequencing technology, which connects dumbbell-shaped adapters to the ends of the template, performs multiple cycle sequencing on the positive and negative strands of the same molecule, and then corrects each other to form consistent sequences, which can improve the sequencing accuracy to 99.99%. Oxford Nanopore developed 2D sequencing technology, which simultaneously sequences the positive and negative sequences of a strand. The sequencing errors in the two sequencing processes are corrected by two independent nanopore sequencing, but the probability of the positive and negative strands entering the same nanopore one after the other is relatively low <60%, and the sequencing quality can only reach 95%, which is still far from the requirements of clinical detection. Generating consistent sequences is still the best way to solve the accuracy of nanopore sequencing. Nanopore sequencing cannot be cyclically sequenced, and needs to generate physically multiple copy molecules to sequence to obtain consistent sequences. Currently, there are methods based on circularization and roll circle amplification (Roll Circle Amplification) to prepare physically tandem multiple copy molecules to generate sequencing consistent sequences and correct sequencing errors in the sequencing process, but this method has a relatively cumbersome procedure, low circularization efficiency and large template loss.
[0004] Therefore, the method for generating multiple copy nucleic acid molecules still needs to be improved.
[0005] SUMMARY
[0006] The present application aims to solve at least one of the problems in the prior art. To this end, the present application provides a method for constructing a tandem repeat multi-copy nucleic acid molecule.
[0007] Specifically, the present application provides the following technical solutions:
[0008] In a first aspect of the present application, the present application provides a method for constructing a tandem multi-copy nucleic acid molecule. According to an embodiment of the present application, the method comprises: performing 3' end A-tailing on a template nucleic acid molecule to obtain a nucleic acid molecule with a 3' end-A tail; performing linker ligation on the nucleic acid molecule with a 3' end-A tail; performing first strand displacement reaction on the product of the linker ligation; performing denaturation and annealing on the product of the first strand displacement reaction; performing second strand displacement reaction on the product of the denaturation and annealing in the presence of a strand displacement primer; wherein, in the product of the linker ligation, the 5' end of the nucleic acid molecule with a 3' end-A tail is connected to the 3' end of the linker sequence, and the 3' end of the nucleic acid molecule with a 3' end-A tail is not connected to the 5' end sequence of the linker sequence. The foregoing method can be used to construct a tandem repeat multi-copy nucleic acid molecule simply and efficiently. In some examples of the present application, the method is used to construct a long fragment double-stranded DNA molecule more efficiently.
[0009] In a second aspect of the present application, the present application provides a nucleic acid molecule. According to an embodiment of the present application, the nucleic acid molecule is obtained by the method of the first aspect of the present application. In some examples of the present application, the foregoing nucleic acid molecule can be used to construct a high-quality sequencing library efficiently.
[0010] In a third aspect of the present application, the present application provides a method for constructing a nucleic acid library. According to an embodiment of the present application, the method comprises: performing 3' end A-tailing on a template nucleic acid molecule to obtain a nucleic acid molecule with a 3' end-A tail; performing linker ligation on the nucleic acid molecule with a 3' end-A tail; performing first strand displacement reaction on the product of the linker ligation; performing denaturation and annealing on the product of the first strand displacement reaction; performing second strand displacement reaction on the product of the denaturation and annealing in the presence of a strand displacement primer; performing sequencing library construction on the product of the second strand displacement reaction to obtain the nucleic acid library; wherein, in the product of the linker ligation, the 5' end of the nucleic acid molecule with a 3' end-A tail is connected to the 3' end of the linker sequence, and the 3' end of the nucleic acid molecule with a 3' end-A tail is not connected to the 5' end sequence of the linker sequence. In some examples of the present application, the method can be used to construct a high-quality sequencing library quickly, and is more compatible with subsequent sequencing and data analysis.
[0011] In a fourth aspect of the present application, a sequencing library is provided. According to embodiments of the present application, the sequencing library is obtained by the method of the third aspect of the present application. In some examples of the present application, sequencing with the sequencing library described above can significantly reduce the sequencing error rate.
[0012] In a fifth aspect of the present application, a sequencing method is provided. According to embodiments of the present application, the method comprises: performing sequencing analysis processing on the sequencing library of the fourth aspect or the sequencing library constructed by the method of the third aspect to obtain the nucleic acid sequence to be detected. In some examples of the present application, based on the method, high-quality sequencing data can be sequenced, and the sequencing error rate is significantly reduced.
[0013] In a sixth aspect of the present application, a kit is provided. According to embodiments of the present application, the kit is used to implement the above-mentioned tandem multi-copy nucleic acid molecule construction method or nucleic acid library construction method or sequencing method. In some examples of the present application, the aforementioned kit can be used for efficient and portable preparation of the above-mentioned nucleic acid molecule.
[0014] Additional aspects and advantages of the present application will be in part apparent and in part pointed out hereinafter. BRIEF DESCRIPTION OF DRAWINGS
[0015] The above and / or additional aspects and advantages of the present application will become apparent and be readily appreciated from the following description, including the appended drawings, wherein:
[0016] FIG. 1 is a schematic diagram of the principle of PacBio HiFi sequencing provided in the summary of the present application;
[0017] FIG. 2 is a schematic diagram of the principle of Oxford Nanopore sequencing provided in the summary of the present application;
[0018] FIG. 3 is a schematic diagram of a neck ring type linker and a primer provided in the detailed description of the present application;
[0019] FIG. 4 is a schematic diagram of a DNA molecule construction process provided in the detailed description of the present application;
[0020] FIG. 5 is a schematic diagram of a DNA molecule nanopore library preparation process provided in the detailed description of the present application.
[0021] The above figures are merely schematic and non-limiting. In the drawings, the size of some components can be exaggerated and not drawn to scale for the purpose of explanation. The size and relative sizes of the components are not necessarily in accordance with the true reduction of the application when implemented. DETAILED DESCRIPTION
[0022] In the present document, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. "A set" or "a plurality" means two or more.
[0023] In the present document, the terms "comprising" or "including," or any other
[0024] In the present document, the terms "first," "second," "third," "fourth" and the like in the description and in the claims, are used for distinguishing between similar elements and not necessarily for describing a sequential or chronological order. It is to be understood that the terms so used are interchangeable under appropriate circumstances and are not to be construed as limiting.
[0025] In the present document, nucleotide sequences are written in the 5' to 3' direction from left to right, unless otherwise indicated.
[0026] Currently, the mainstream long-read sequencing method (also known as third-generation sequencing) includes PacBio HiFi sequencing and Oxford Nanopore sequencing. Among them, the principle of PacBio HiFi sequencing is shown in Figure 1, which repeatedly sequences the same sequence multiple times by polymerase cycling on a circular library, thereby reducing the base recognition error rate; the principle of Oxford Nanopore sequencing is shown in Figure 2, which realizes the melting and sequencing of double-stranded DNA by connecting a double-linker with a motor protein on one end of the target sequence and a neck-loop linker on the other end. In the sequencing process, the end with the motor protein is sequenced first, and after the positive strand is sequenced, the direction is changed through the neck-loop linker to sequence the negative strand. Therefore, each target sequence is sequenced at most twice.
[0027] Compared with PacBio HiFi, the ability of Oxford Nanopore sequencing to sequence the same sequence multiple times is lower. To solve this problem, some people have proposed using tandem repeat multiple copies. First, the double-stranded DNA fragments of the sample are circularized to obtain single-stranded circular DNA, and then tandem repeat multiple copies are generated by using this as a template for rolling circle amplification. Subsequently, the conventional adapter addition step of Nanopore is performed. However, this method of generating tandem repeat multiple copies has some problems, such as the efficiency of circularizing linear double-stranded DNA into single-stranded circular DNA is very low (usually no more than 30%), which means that a large amount of sample will be lost in the circularization step, resulting in a large amount of template sequence information being lost, which seriously affects the efficiency and accuracy of sequencing.
[0028] In order to improve the efficiency and accuracy of sequencing, in one aspect of the present application, the present application proposes a nucleic acid molecule construction method, which comprises:
[0029] Step one: 3' end A-tailing of the template nucleic acid molecule to obtain a nucleic acid molecule with a 3' end-A tail;
[0030] In some examples of the present application, the template nucleic acid molecule is double-stranded DNA. The aforementioned double-stranded DNA is selected from at least one of full-length cDNA, mitochondrial DNA and chromosomal DNA. In some preferred examples of the present application, the aforementioned double-stranded DNA is not less than 100 bp.
[0031] In some examples of the present application, the aforementioned double-stranded DNA can be obtained by reverse transcription of RNA or by double-strand synthesis of single-stranded DNA.
[0032] In some examples of the present application, for the template nucleic acid molecule with non-blunt ends, further comprising: end repair treatment of the template nucleic acid molecule.
[0033] Step two: linker ligation treatment of the nucleic acid molecule with a 3' end-A tail;
[0034] In some examples of the present application, the linker is selected from a necking type linker, which includes a necking structure region and a double-stranded structure region. In some examples of the present application, the necking type linker includes at least part of the sequence of the strand displacement primer. In a specific example of the present application, the structure of the necking type linker is shown in Figure 3. The necking type linker provides a starting point for strand side reaction, so that the DNA polymerase can perform strand displacement reaction along the template strand.
[0035] In some examples of the present application, the linker includes a fixed sequence, which includes a primer binding sequence and / or a molecular tag sequence (UMI). In some examples of the present application, the aforementioned molecular tag sequence is located in the necking structure region or the double-stranded structure region. The molecular tag sequence is used to determine the number of DNA molecule copies; the fixed sequence is used to form a necking structure, improve the stability of the linker; the primer binding sequence is used for complementary pairing with the strand displacement primer, guiding the binding of the polymerase.
[0036] Step three: first strand displacement reaction treatment of the linker ligation treatment product; wherein in the linker ligation treatment product, the 5' end of the double-stranded DNA molecule with blunt ends and a 3'-A tail is connected to the 3' end of the linker sequence, and the 3' end of the double-stranded DNA molecule with blunt ends and a 3'-A tail is not connected to the 5' end sequence of the linker sequence;
[0037] In some examples of the present application, the strand displacement reaction is performed in the presence of a strand displacement polymerase. The aforementioned strand displacement polymerase includes at least one of Bst DNA polymerase, Bsu DNA polymerase, Phi29 DNA polymerase and Klenow fragment (exo-).
[0038] Step four: denaturation and annealing treatment is performed on the first strand displacement reaction product;
[0039] In some examples of the present application, the strand displacement primer can carry a neck portion sequence or not carry a neck portion sequence (Figure 1).
[0040] Step five: second strand displacement reaction treatment is performed on the denaturation and annealing treatment product in the presence of a strand displacement primer.
[0041] The aforementioned method only needs one primer to complete the construction of the tandem repeat multi-copy nucleic acid molecule, the reaction system components are single, the system preparation is more simple, and the obtained product can be used in subsequent steps. Moreover, the obtained tandem repeat multi-copy nucleic acid molecule is more compatible for long-length sequencing library construction and sequencing.
[0042] For the convenience of understanding, with reference to Figure 4, the aforementioned nucleic acid molecule construction method is described in detail taking a double-stranded DNA as a template nucleic acid molecule as a specific example.
[0043] 1. End repair and 3' end A addition (for TA cloning) are performed on the template double-stranded DNA molecule to obtain a double-stranded DNA molecule with a blunt end and a 3' end-A tail. The end repair and 3' end A addition steps can be performed by using conventional laboratory methods, which are not specifically limited in the present application.
[0044] It should be noted that the end repair step is not used for the double-stranded DNA molecule with a blunt end.
[0045] 2. The obtained double-stranded DNA molecule with a blunt end and a 3' end-A tail is subjected to adapter ligation (such as a neck adapter). In the adapter ligation treatment product, the 5' end of the double-stranded DNA molecule with a blunt end and a 3' end-A tail is connected to the 3' end of the adapter sequence, and the 3' end of the double-stranded DNA molecule with a blunt end and a 3' end-A tail is not connected to the 5' end sequence of the adapter sequence. It can be understood by those skilled in the art that, through the complementary pairing type of the DNA molecule, a neck-double-stranded DNA molecule-neck molecule is formed.
[0046] 3. The neck-double-stranded DNA molecule-neck molecule is subjected to a strand displacement reaction in the presence of a strand displacement DNA polymerase to obtain a double-stranded DNA molecule with complete adapters at both ends.
[0047] 4. Denature and anneal the double-stranded DNA molecule with complete adaptors to make the double-stranded DNA molecule into single-stranded DNA, and the two ends have complete adaptors.
[0048] 5. Perform second strand displacement reaction on the denatured and annealed product in the presence of second strand displacement primer and strand displacement DNA polymerase.
[0049] In another aspect of the present application, the present application provides a nucleic acid molecule. The nucleic acid molecule is obtained by any of the above-mentioned example methods. In some examples of the present application, based on the aforementioned nucleic acid molecule, the sequencing library construction can significantly improve the library construction efficiency.
[0050] In some examples of the present application, the nucleic acid molecule is double-stranded DNA.
[0051] In another aspect of the present application, the present application provides a nucleic acid library construction method, which comprises:
[0052] Step 1: Perform 3' end A treatment on the template nucleic acid molecule to obtain a nucleic acid molecule with a 3' end-A tail;
[0053] In some examples of the present application, the template nucleic acid molecule is double-stranded DNA. The aforementioned double-stranded DNA is selected from at least one of full-length mRNA, mitochondrial DNA and chromosomal DNA.
[0054] In some examples of the present application, the double-stranded DNA is selected from at least one of full-length cDNA, mitochondrial DNA and chromosomal DNA. In some preferred examples of the present application, the aforementioned double-stranded DNA is not less than 100 bp.
[0055] Wherein, the aforementioned double-stranded DNA can be obtained by reverse transcription of RNA, or can be obtained by double-strand synthesis of single-stranded DNA.
[0056] In some examples of the present application, for the template nucleic acid molecule with non-blunt ends, further comprising: performing end repair treatment on the template nucleic acid molecule.
[0057] Step 2: Perform adaptor ligation treatment on the nucleic acid molecule with 3' end-A tail;
[0058] In some examples of the present application, the adaptor is selected from a neck ring type adaptor, which comprises a neck ring structure region and a double-stranded structure region. In some examples of the present application, the neck ring type adaptor comprises at least part of the sequence of the strand displacement primer. In one example of the present application, the structure of the neck ring type adaptor is shown in Figure 3. The neck ring type adaptor provides a starting point for strand displacement reaction, so that the DNA polymerase can perform strand displacement reaction along the template strand.
[0059] In some examples of the present application, the adaptor comprises a fixed sequence, and the fixed sequence comprises a primer binding sequence and / or a molecular tag sequence (UMI). In some examples of the present application, the molecular tag sequence is located in the neck structure region or the double-stranded structure region. The molecular tag sequence is used to determine the number of DNA molecule copies, the fixed sequence is used to form a neck structure and improve the stability of the adaptor, and the primer binding sequence is used to complementarily pair with the strand displacement primer and guide the binding of the polymerase.
[0060] Step three: performing a first strand displacement reaction on the adaptor ligation product, wherein in the adaptor ligation product, the 5' end of the double-stranded DNA molecule with blunt ends and a 3'-A tail is connected to the 3' end of the adaptor sequence, and the 3' end of the double-stranded DNA molecule with blunt ends and a 3'-A tail is not connected to the 5' end sequence of the adaptor sequence.
[0061] In some examples of the present application, the strand displacement reaction is performed in the presence of a strand displacement polymerase. The strand displacement polymerase includes at least one of Bst DNA polymerase, Bsu DNA polymerase, Phi29 DNA polymerase, and Klenow fragment (exo-).
[0062] Step four: performing denaturation and annealing treatment on the first strand displacement reaction product.
[0063] In some examples of the present application, the strand displacement primer can carry a neck portion sequence or not carry a neck portion sequence (Figure 1).
[0064] Step five: performing a second strand displacement reaction on the denaturation and annealing treatment product in the presence of a strand displacement primer.
[0065] Step six: performing sequencing library construction treatment on the second strand displacement reaction product to obtain the nucleic acid library.
[0066] In some examples of the present application, the sequencing library construction treatment comprises: performing a phosphorylation treatment on the second strand displacement reaction product. The phosphorylation treatment is performed in the presence of a phosphokinase. In a specific example of the present application, the phosphokinase is selected from T4 phosphokinase.
[0067] In some examples of the present application, the sequencing library construction treatment further comprises: performing a 3' end A-tailing treatment on the second strand displacement reaction product. Through the A-tailing treatment, a binding site is provided for the sequencing adaptor, so that the nucleic acid molecule can be effectively connected to the sequencing adaptor. In addition, the A-tailing treatment can also reduce the occurrence of ligation bias, improve the ligation efficiency, and increase the success rate of library construction.
[0068] In some examples of the present application, the sequencing library construction process further comprises: performing a sequencing adapter ligation process on the second strand displacement reaction process product. After the above phosphorylation and A-tailing process, the double-stranded DNA molecules are further ligated with sequencing adapters under the action of ligase to obtain a sequencing library.
[0069] In some examples of the present application, the sequencing adapter is a nanopore sequencing adapter. In some preferred examples of the present application, the nanopore sequencing adapter carries a motor protein. In some preferred examples of the present application, the nanopore sequencing adapter further carries a tether protein.
[0070] The sequencing library obtained based on the method avoids the generation of intermediate products, and the final product obtained is single in composition and can be used in subsequent sequencing analysis steps.
[0071] For ease of understanding, with reference to FIG. 5, the above nucleic acid library construction method is described in detail taking a double-stranded DNA as a template nucleic acid molecule as a specific example.
[0072] 1. End repair and 3' end A-tailing (for TA cloning) are performed on the template double-stranded DNA molecules to obtain double-stranded DNA molecules with blunt ends and 3' end-A tails. The end repair and 3' end A-tailing steps can be performed using conventional laboratory methods, which are not specifically limited in the present application.
[0073] It should be noted that the end repair step is not used for double-stranded DNA molecules with blunt ends.
[0074] 2. The obtained double-stranded DNA molecules with blunt ends and 3' end-A tails are subjected to adapter ligation (such as neck ring adapter). In the adapter ligation process product, the 5' end of the double-stranded DNA molecule with blunt ends and 3'-A tail is connected to the 3' end of the adapter sequence, and the 3' end of the double-stranded DNA molecule with blunt ends and 3'-A tail is not connected to the 5' end sequence of the adapter sequence. Those skilled in the art can understand that, through the complementary pairing type of the DNA molecule, a neck ring-double-stranded DNA molecule-neck ring molecule is formed.
[0075] 3. A strand displacement reaction is performed on the neck ring-double-stranded DNA molecule-neck ring molecule in the presence of a strand displacement DNA polymerase to obtain double-stranded DNA molecules with complete adapters at both ends.
[0076] 4. Denaturation and annealing are performed on the double-stranded DNA molecules with complete adapters to make the double-stranded DNA molecules single-stranded and have complete adapters at both ends.
[0077] 5. A second strand displacement reaction is performed on the denaturation and annealing process product in the presence of a second strand displacement primer and a strand displacement DNA polymerase.
[0078] 7. Phosphorylating and 3' end A-tailing the double-stranded DNA molecules obtained in step 6.
[0079] 8. Adding sequencing adapters (containing a motor protein) to the double-stranded DNA molecules obtained in step 7 to obtain a sequencing library.
[0080] In yet another aspect of the present application, the present application provides a sequencing library, which is obtained by the method of any of the above examples. In some examples of the present application, sequencing the above sequencing library can reduce sequencing errors.
[0081] In yet another aspect of the present application, the present application provides a sequencing method, which comprises: sequencing and analyzing the above sequencing library or the sequencing library obtained by the above method to obtain a nucleic acid sequence to be detected. In some examples of the present application, based on the high-quality sequencing data obtained by the method, the sequencing error rate can be significantly reduced.
[0082] In some examples of the present application, the sequencing and analyzing comprises: determining the number of copies of each original DNA molecule included in the original sequencing reads based on the fixed sequence and the molecular tag sequence in the adapter sequence, and dividing the original sequencing reads into DNA sequences with corresponding copy numbers.
[0083] In some examples of the present application, the sequencing and analyzing comprises: correcting the multiple copies of DNA sequences contained in the sequencing reads with each other to obtain corrected DNA sequences. By correction, sequencing errors are eliminated.
[0084] In some examples of the present application, the sequencing and analyzing further comprises: combining the corrected DNA sequences to determine a consensus sequence; and aligning the consensus sequence based on a reference sequence to determine a nucleic acid sequence to be detected.
[0085] It should be noted that the above-mentioned reference sequence (reference, ref) is a determined sequence, which can be a DNA and / or RNA sequence determined and assembled by oneself, or a DNA and / or RNA sequence determined and published by others, or a reference template of any biological category to which the sample source individual or the target individual belongs, for example, all or at least part of the published genome assembly sequence of the same biological category. If the sample source individual or the target individual is a human, the genome reference sequence (also referred to as a reference genome or a reference chromosome) can be selected from the human reference genomes provided by the UCSC, NCBI or ENSEMBL database, such as HG19, HG38, GRCh36, GRCh37, GRCh38, etc. Those skilled in the art can understand the correspondence of the above-mentioned reference genome versions by referring to the database instructions and select the version to be used. Further, a library containing more reference sequences can also be pre-configured, for example, before alignment, a sequence closer or more characteristic in some aspects is selected or determined and assembled as a reference sequence according to the gender, race, region, etc. of the target individual, which is helpful to obtain more accurate sequence analysis results subsequently.
[0086] In some examples of the present application, the sequencing analysis further comprises: performing bioinformatics analysis on the obtained nucleic acid sequence to be tested, such as gene function prediction, translation and structure analysis of encoded proteins, etc.
[0087] In still another aspect of the present application, the present application provides a kit for realizing the above-mentioned tandem multi-copy nucleic acid molecule construction method or nucleic acid library construction method or sequencing method. In some examples of the present application, the above-mentioned nucleic acid molecule can be quickly prepared based on the foregoing kit, so as to perform nucleic acid library construction and / or sequencing.
[0088] In some examples of the present application, the foregoing kit comprises: at least one of a linker, a DNA polymerase, a DNA ligase, a dNTP and a strand displacement polymerase.
[0089] In some examples of the present application, the kit further comprises: a sequencing reagent and an instruction. The foregoing instruction records the corresponding steps for realizing the above-mentioned tandem multi-copy nucleic acid molecule construction method or nucleic acid library construction method or sequencing method.
[0090] Embodiments of the present application will be described in more detail below, examples of which are shown in the accompanying drawings. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present application, and cannot be understood as a limitation of the present application. If a specific technique or condition is not specified in the embodiments, the technique or condition described in the literature in the art or according to the product instruction is used. If the manufacturer of the reagent or instrument is not specified, it is a conventional product that can be obtained by purchase.
[0091] Example 1: DNA nanopore sequencing library construction and sequencing
[0092] The experimental scheme of this example is as follows: a single product of human mitochondria is amplified, with a length of about 16 kb, and DNA library preparation is performed according to the method of the present application and the conventional method, respectively. The sequencing library obtained is sequenced on a nanopore sequencer, and the accuracy of sequencing is evaluated.
[0093] Among them, the conventional nanopore library is strictly operated according to the Nanopore company Ligation sequencing gDNA V14-human sample (N50 10kb) on PromethION (SQK-LSK114) operation instruction.
[0094] 1. Terminal A addition
[0095] The terminal A addition reaction system is shown in Table 1;
[0096] Table 1 Terminal A addition reaction system
[0097] Reaction conditions: place the above reaction system in a PCR instrument at 75°C for 20 min. After the reaction, purify with 1.0x AMPure magnetic beads, and finally dissolve the purified product in 13 μl elution buffer;
[0098] Among the above reagents, Taq hot start DNA polymerase and 10X standard Taq reaction buffer are from NEB company, item number M0495S; dATP solution is from NEB company, item number N0440S; AMPure XP Beads magnetic beads are from Beckman company, item number A63882.
[0099] 2. Adapter ligation and strand displacement reaction:
[0100] 2.1 Adapter ligation reaction
[0101] The DNA obtained in step 1 is prepared into an adapter ligation reaction system according to Table 2;
[0102] Table 2 Adapter ligation reaction system
[0103] Among the above reagents, the neck ring adapter sequence is:
[0104] Among them, the underlined N part is a molecular tag, and the bold part is a primer binding position; T4 DNA ligase and 10X DNA Ligase Reaction Buffer are from NEB, item number M0202V;
[0105] 2.2 Ligation reaction condition:
[0106] The above reaction system was placed in a PCR instrument, and reacted at 20°C for 15 min. After the reaction, 1.0x AMPure magnetic beads were used for purification, and finally the purified product was dissolved in 15 μl elution buffer.
[0107] 2.3 Strand displacement reaction system
[0108] The ligation reaction product was prepared according to Table 3 to prepare the strand displacement reaction system;
[0109] Table 3 Strand displacement reaction system
[0110] The above reagents are from NEB company, Bst 2.0 DNA Polymerase and 10X Isothermal Amplification Buffer, item number M0537S; 100mM MgSO4, item number B1003S; 10mM dNTP Mix, item number N0447;
[0111] 2.4 Strand displacement reaction condition
[0112] The above reaction system was placed in a PCR instrument, and reacted at 65°C for 10 min. After the reaction, 1.0x AMPure magnetic beads were used for purification, and finally the purified product was dissolved in 20 μl elution buffer.
[0113] 3, denaturation and strand displacement reaction
[0114] The purified product obtained in step 2 was placed in a PCR instrument, and reacted at 95°C for 3 min and at 65°C for 3 min. After the reaction, it was placed in an ice box, and a strand displacement reaction system was prepared based on Table 4;
[0115] Table 4 Strand displacement reaction system
[0116] Among them, the sequence of Primer is: TCGTAGCCATGTCGTTC (SEQ ID NO: 2); the reagents in Table 4 are from NEB company, Bst 2.0 DNA Polymerase and 10X Isothermal Amplification Buffer, item number M0537S; 100mM MgSO4, item number B1003S; 10mM dNTP Mix, item number N0447;
[0117] The above reaction system was placed in a PCR instrument, and reacted at 65°C for 20 min. After the reaction, 1.0x AMPure magnetic beads were used for purification, and finally the purified product was dissolved in 15 μl elution buffer.
[0118] 4. 5' end phosphorylation and 3' end A-tailing
[0119] The DNA obtained in step 3 was prepared into a phosphorylation and A-tailing reaction system according to Table 5;
[0120] Table 5 Phosphorylation and A-tailing reaction system
[0121] The above reagents are from NEB Company, T4 Phosphokinase and T4 PNK Reaction Buffer (10X) with item number M0201V; Taq hot start DNA polymerase is from NEB Company with item number M0495S; dATP solution is from NEB Company with item number N0440S;
[0122] The above reaction system was placed in a PCR instrument, 37°C, 15 min; 65°C, 15 min. After the reaction, 1.0X AMPure magnetic beads were used for purification, and finally the purified product was dissolved in 60 μl elution buffer.
[0123] 5. Nanopore sequencing adapter addition
[0124] The DNA obtained in step 4 was prepared into a nanopore sequencing adapter ligation system according to Table 6;
[0125] Table 6 Nanopore sequencing adapter ligation system
[0126] Among the above reagents, NEB Next Quick T4 DNA Ligase is from NEB Company with item number E6056; Ligation Buffer and Adapter are from Nanopore Company with item number SQK-LSK114;
[0127] The above reaction system was placed in a PCR instrument, 25°C, 10 min; 65°C, 20 min. After the reaction, 0.4X AMPure magnetic beads were used for purification, and finally the purified product was dissolved in 25 μl elution buffer. The library was quantified using HS Qubit ssDNA kit.
[0128] 6. Sequencing on machine
[0129] The obtained library was sequenced on a PromethION Flow Cell instrument, the chip was selected as R10.4.1 flow cell, and the operation was performed according to the standard instruction manual, Ligation Sequencing Kit V14 (SQK-LSK114).
[0130] 7. Information analysis
[0131] Data obtained by the conventional method were aligned using Minimap and visualized using IGV; data obtained by the inventive method were first subjected to consensus sequence generation, then the obtained data were aligned using Minimap and visualized using IGV.
[0132] 8. Result analysis
[0133] Table 7 Length statistics of raw sequencing data
[0134] The above results show that, compared with the conventional method, the raw sequencing read length obtained based on the method of the application is longer, and the read length greater than 32 kb accounts for 93.3%, representing the proportion of data with more than 2 times of sequencing (including positive and negative strands); the read length greater than 64 kb accounts for 66.2%, representing the proportion of more than 4 times (including positive and negative strands), which also indicates that the copy number of the obtained tandem repeat is greater than 2 times, and the proportion is 66.2%; there is also a read length greater than 128 kb, and the proportion reaches 30.6%.
[0135] Example 2: RNA nanopore sequencing library construction and sequencing
[0136] The experimental scheme of this example is as follows: take the new crown standard product (Twist Synthetic SARS-CoV-2 RNA controls, cat#102024: Control 2 MN908947.3), 10000 copies are put in, and the DNA is prepared into a library according to the inventive method and the conventional method respectively, and the obtained sequencing library is sequenced on a nanopore sequencer, the new crown standard product 3 end fragment length is about 5 kb, and the accuracy of sequencing is evaluated.
[0137] Among them, the conventional nanopore library is strictly operated according to the operation instruction of Nanopore company cDNA-PCR Sequencing Kits.
[0138] 1. cDNA synthesis
[0139] Oligod T is used for full-length cDNA synthesis of total RNA, and template conversion is performed, and the system (Table 8) and conditions are as follows:
[0140] Table 8 Template conversion system
[0141] The TSO sequence is: GACATGGCTACGATCCGACTT (SEQ ID NO: 3); wherein the end of the SEQ ID NO: 3 sequence contains three riboguanosines (rGrGr+G);
[0142] Oligo dT sequence: GCATCCATTAGTTAGGCTAGTTTTTTTTTTTTTTTTTTVN (SEQ ID NO: 4);
[0143] The above reaction system was placed on a PCR instrument, 42°C, 60 min. After the reaction, 1.0x AMPure magnetic beads were used for purification, and finally the purified product was dissolved in 21 μl elution buffer. Reverse transcriptase was purchased from thermofisher company, item number 18090200.
[0144] 2. PCR reaction
[0145] The PCR reaction system is shown in Table 9;
[0146] Table 9 PCR reaction system
[0147] The PCR reaction program is shown in Table 10;
[0148] Table 10 PCR reaction program
[0149] After the reaction, 1.0x AMPure magnetic beads were used for purification, and finally the purified product was dissolved in 13 μl elution buffer. PCR amplification enzyme from NEB, cat#M0533S;
[0150] PCR Primer 1: GACATGGCTACGATCCGACTT (SEQ ID NO: 5);
[0151] PCR Primer 2: GCATCCATTAGTTAGGCTAG (SEQ ID NO: 6).
[0152] 3. Linker ligation and strand displacement reaction:
[0153] The DNA obtained in step 2 was prepared into a linker ligation reaction system according to Table 11;
[0154] Table 11 Linker ligation reaction system
[0155] The loop linker sequence is: The underlined N part is a molecular tag, and the bold part is the primer binding position;
[0156] The above T4 DNA ligase and 10X DNA Ligase Reaction Buffer are from NEB, item number M0202V;
[0157] Put the above reaction system on PCR instrument, 20°C, 15 min. After the reaction, use 1.0x AMPure magnetic beads for purification, and finally dissolve the purified product in 15 μl elution buffer.
[0158] Prepare the chain displacement reaction system according to Table 12 for the purified product;
[0159] Table 12 Chain displacement reaction system
[0160] The above reagents are from NEB company, Bst 2.0 DNA Polymerase and 10X Isothermal Amplification Buffer, item number M0537S; 100 mM MgSO4, item number B1003S; 10 mM dNTP Mix, item number N0447;
[0161] Put the above reaction system on PCR instrument, 65°C, 10 min. After the reaction, use 1.0x AMPure magnetic beads for purification, and finally dissolve the purified product in 20 μl elution buffer.
[0162] 4, denaturation and chain displacement reaction
[0163] Put the purified product of step 3 on PCR instrument, 95°C, 3 min, 65°C, 3 min. After the reaction, place it in an ice box.
[0164] Prepare the chain displacement reaction system according to Table 13 for the denatured DNA;
[0165] Table 13 Chain displacement reaction system
[0166] The sequence of Primer: TCGTAGCCATGTCGTTC (SEQ ID NO: 8);
[0167] The above reagents are from NEB company, Bst 2.0 DNA Polymerase and 10X Isothermal Amplification Buffer, item number M0537S; 100 mM MgSO4, item number B1003S; 10 mM dNTP Mix, item number N0447;
[0168] Put the above reaction system on PCR instrument, 65°C, 20 min. After the reaction, use 1.0x AMPure magnetic beads for purification, and finally dissolve the purified product in 15 μl elution buffer;
[0169] 5, 5' end phosphorylation and 3' end A tailing
[0170] The DNA obtained from step 4 was prepared into a phosphorylation and A tail adding reaction system according to Table 14
[0171] Table 14 Phosphorylation and A tail adding reaction system
[0172] The above reagents are from NEB, T4 polynucleotide kinase and T4 PNK Reaction Buffer (10X) with the item number M0201V; Taq hot start DNA polymerase is from NEB with the item number M0495S; dATP solution is from NEB with the item number N0440S;
[0173] The above reaction system was placed in a PCR instrument, 37℃, 15min; 65℃, 15min. After the reaction, 1.0X AMPure magnetic beads were used for purification, and finally the purified product was dissolved in 60μl elution buffer.
[0174] 6. Nanopore sequencing adapter addition
[0175] The DNA obtained from step 5 was prepared into a nanopore sequencing adapter ligation system according to Table 15;
[0176] Table 15 Nanopore sequencing adapter ligation system
[0177] Among the above reagents, NEB Next Quick T4 DNA Ligase is from NEB with the item number E6056; Ligation Buffer and Adapter are from Nanopore with the item number SQK-LSK114;
[0178] The above reaction system was placed in a PCR instrument, 25℃, 10min; 65℃, 20min. After the reaction, 0.4X AMPure magnetic beads were used for purification, and finally the purified product was dissolved in 25μl elution buffer. The library was quantified using HS Qubit ssDNA kit.
[0179] 7. Sequencing on machine
[0180] The obtained library was sequenced on a PromethION Flow Cell instrument, the chip was selected as R10.4.1 flow cell, and the operation was performed according to the standard instruction manual, Ligation Sequencing Kit V14 (SQK-LSK114).
[0181] 8. Information analysis
[0182] The data obtained by the conventional method is aligned by Minimap and visualized by IGV; the data obtained by the inventive method is first subjected to consistent sequence generation, then aligned by Minimap, and visualized by IGV.
[0183] 9. Result analysis
[0184] Table 16 Length statistics of raw sequencing data
[0185] The above results show that, compared with the conventional method, the raw sequencing read length obtained based on the method of the present application is longer, and the read length greater than 10 kb accounts for 92.3%, representing the proportion of data with more than 2 times of sequencing (including positive and negative strands); the read length greater than 20 kb accounts for 64.1%, representing the proportion of more than 4 times (including positive and negative strands), which also indicates that the copy number of the tandem repeat obtained is greater than 2 times, and the proportion is 64.1%; there is also a read length greater than 128 kb, indicating that the proportion of the copy number of the tandem repeat obtained is greater than 4 times, which also reaches 36.1%.
[0186] In the description of the present specification, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples. In addition, different embodiments or examples described in the present specification and the features of different embodiments or examples can be combined and combined by those skilled in the art without contradiction.
[0187] Although the embodiments of the present application have been shown and described above, it should be understood that the above embodiments are exemplary and should not be construed as limiting the present application, and those skilled in the art can make changes, modifications, replacements and variations to the above embodiments within the scope of the present application.
Claims
1. A method for constructing tandem multiple-copy nucleic acid molecules, characterized in that, The method comprises: performing 3' end A-tailing on a template nucleic acid molecule to obtain a nucleic acid molecule having a 3' end-A tail; performing adaptor ligation on the nucleic acid molecule having the 3' end-A tail; performing a first strand displacement reaction on the adaptor ligation product; performing denaturation and annealing on the first strand displacement reaction product; performing a second strand displacement reaction on the denaturation and annealing product in the presence of a strand displacement primer; wherein, in the adaptor ligation product, the 5' end of the nucleic acid molecule having the 3' end-A tail is connected to the 3' end of the adaptor sequence, and the 3' end of the nucleic acid molecule having the 3' end-A tail is not connected to the 5' end sequence of the adaptor sequence.
2. The method of claim 1, wherein, The template nucleic acid molecule is double-stranded DNA. Optionally, the double-stranded DNA is selected from at least one of full-length cDNA, mitochondrial DNA and chromosomal DNA.
3. The method of claim 2, wherein, The double-stranded DNA is obtained by reverse transcription of RNA.
4. The method of claim 2, wherein, The double-stranded DNA molecule is obtained by double-stranded synthesis of single-stranded DNA.
5. The method of claim 1, wherein, Further comprising, before the A-tailing, performing end repair on the template nucleic acid molecule.
6. The method of claim 1, wherein, The adaptor is selected from a neck-lace type adaptor.
7. The method of claim 6, wherein, The neck-lace type adaptor comprises at least part of the sequence of the strand displacement primer.
8. The method of claim 6, wherein, The adaptor comprises a molecular tag sequence.
9. The method of claim 1, wherein, The strand displacement reaction is performed in the presence of a strand displacement polymerase. Optionally, the strand displacement polymerase comprises at least one of Bst DNA polymerase, Bsu DNA polymerase, Phi29 DNA polymerase and Klenow fragment (exo-).
10. A nucleic acid molecule, characterized in that, The nucleic acid library is constructed by the method of any one of claims 1-9.
11. The nucleic acid molecule of claim 10, wherein The nucleic acid molecule is double-stranded DNA.
12. A method of nucleic acid library construction, comprising, The method comprises: performing 3' end A-tailing on a template nucleic acid molecule to obtain a nucleic acid molecule having a 3' end-A tail; performing adaptor ligation on the nucleic acid molecule having the 3' end-A tail; performing a first strand displacement reaction on the adaptor ligation product; performing denaturation and annealing on the first strand displacement reaction product; performing a second strand displacement reaction on the denaturation and annealing product in the presence of a strand displacement primer; performing sequencing library construction on the second strand displacement reaction product to obtain the nucleic acid library; wherein, in the adaptor ligation product, the 5' end of the nucleic acid molecule having the 3' end-A tail is connected to the 3' end of the adaptor sequence, and the 3' end of the nucleic acid molecule having the 3' end-A tail is not connected to the 5' end sequence of the adaptor sequence.
13. The method of claim 12, wherein, The template nucleic acid molecule is double-stranded DNA. Optionally, the double-stranded DNA is selected from at least one of full-length cDNA, mitochondrial DNA and chromosomal DNA.
14. The method of claim 13, wherein, The double-stranded DNA is obtained by reverse transcription of RNA.
15. The method of claim 13, wherein, The double-stranded DNA molecule is obtained by double-stranded synthesis of single-stranded DNA.
16. The method of claim 12, wherein, Further comprising, before the A-tailing, performing end repair on the template nucleic acid molecule.
17. The method of claim 12, wherein, The adaptor is selected from a neck-lace type adaptor.
18. The method of claim 17, wherein, The neck-lace type adaptor comprises at least part of the sequence of the strand displacement primer.
19. The method of claim 17, wherein, The adaptor comprises a molecular tag sequence.
20. The method of claim 12, wherein, The strand displacement reaction is performed in the presence of a strand displacement polymerase. Optionally, the strand displacement polymerase comprises at least one of Bst DNA polymerase, Bsu DNA polymerase, Phi29 DNA polymerase and Klenow fragment (exo-).
21. The method of claim 12, wherein, The sequencing library construction process comprises: performing a phosphorylation process on the second strand displacement reaction process product.
22. The method of claim 21, wherein, The phosphorylation process is performed in the presence of a phosphorylase.
23. The method of claim 21, wherein, The sequencing library construction process further comprises: performing a 3' end A-tailing process on the second strand displacement reaction process product.
24. The method of claim 23, wherein, The sequencing library construction process further comprises: performing a sequencing adaptor ligation process on the second strand displacement reaction process product.
25. The method of claim 23, wherein, The sequencing adaptor is a nanopore sequencing adaptor. Preferably, the nanopore sequencing adaptor carries a motor protein. More preferably, the nanopore sequencing adaptor further carries a tether protein.
26. A sequencing library, wherein, The sequencing library is constructed by the method of any one of claims 12-25.
27. A sequencing method, comprising: The method comprises: performing a sequencing analysis process on the sequencing library of claim 26 to obtain the nucleic acid sequence to be detected.
28. The method of claim 27, wherein, The sequencing analysis process comprises: determining the number of copies of each original DNA molecule included in the original sequencing sequence based on the fixed sequence and the molecular tag sequence in the adaptor sequence, and dividing the original sequencing sequence into DNA sequences with corresponding copy numbers.
29. The method of claim 28, wherein, The sequencing analysis process comprises: performing mutual correction processing between the multiple copy number DNA sequences obtained by sequencing to obtain aligned and corrected DNA sequences.
30. The method of claim 28 or 29, wherein, The sequencing analysis process further comprises: combining the corrected DNA sequences to determine a consensus sequence; performing an alignment process on the consensus sequence based on a reference sequence to determine the nucleic acid sequence to be detected.
31. A kit comprising, The kit is used to implement the tandem multi-copy nucleic acid molecule construction method of any one of claims 1-9 or the nucleic acid library construction method of any one of claims 12-25 or the sequencing method of any one of claims 27-30.
Citation Information
Patent Citations
Amplification of polynucleotides by rolling circle amplification
CN101415838A
Compositions and methods for nucleic acid sequencing
CN102084001A
Systems and methods for clonal replication and amplification of nucleic acid molecules for genomic and therapeutic applications
CN106460065A
Double chain nucleic acid fragment joint adding method, library constructing method and kit
CN108060191A
Methods and compositions for analyzing nucleic acid
CN111954720A