Design method of two-dimensional DNA molecular tag and amplification primer in whole plasmid sequencing
By designing two-dimensional DNA molecular tags and amplification primers, combined with nanopore sequencing technology, the problems of low throughput and high cost in whole plasmid sequencing are solved, and high throughput and low-cost full-length plasmid sequencing are achieved.
Patent Information
- Application Number
- CN202510289669.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art has problems of low throughput and high cost in whole plasmid sequencing, making it difficult to achieve high throughput and low cost full-length plasmid sequencing.
The design method of two-dimensional DNA molecular tags and amplification primers is adopted, combined with nanopore sequencing technology, high-throughput plasmid sequencing is achieved. This method performs high-throughput sequencing by designing specific amplification primers and two-dimensional DNA molecular tags, linearizing plasmids and adding tags.
The application scenarios of the nanopore sequencing platform have been greatly expanded, high-throughput full plasmid sequencing has been achieved, and the reagent consumables, time and labor costs have been reduced.
Smart Images

Figure CN120060443A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of molecular biology, and particularly relates to a method for designing two-dimensional DNA molecular tags and amplification primers in whole plasmid sequencing. Background Art
[0002] As the most commonly used gene vector in modern biotechnology, plasmids play a key role in the fields of scientific research and medicine. Plasmid construction is a core ability for researchers to synthesize new proteins, which includes the gene encoding of related proteins and the attachments for providing expression or genetic trait control for specific stress options. These genetic toolkits enable fine control, but more importantly, all features are currently verified and correct so that the experiment can succeed.
[0003] Recently, research from GenScript has shown that up to 50% of laboratory plasmids may have design or sequence errors, which may have a negative impact on their intended applications. As a cloning service provider, the team had the opportunity to process and evaluate a large number of plasmids from academic and industrial laboratories around the world. Through this first large-scale quality assessment, the study found an unexpectedly high error rate. Specifically, approximately 15% of the plasmids had obvious design errors, and approximately 35% of the plasmids had sequence errors in the functional regions. Especially in adeno-associated virus (AAV) transfer plasmids, approximately 40% of the inverted terminal repeats (ITRs) had mutations compared to the wild type. Overall, the team estimated that 45 - 50% of laboratory-made plasmids contained undetected design and / or sequence errors, which may damage their intended applications. In addition, because customers perform certain self-checks before submitting plasmids, the true scale of the problem may be even larger, and the 50% error rate may be underestimated. Although the importance of plasmids is self-evident, a global quality assessment of these laboratory-made plasmids is still lacking. Checking for sequence errors in plasmids by sequencing can help researchers identify and correct such problems in their research, but sequencing also requires a certain period of time and cost.
[0004] As a first-generation sequencing method, Sanger sequencing has always been the gold standard method for sequencing due to its extremely high accuracy. However, for long-sequence nucleic acid samples to be sequenced, it is necessary to repeat the primer design, primer synthesis, and sequencing steps multiple times, resulting in a long experimental time and high cost. In addition, first-generation sequencing cannot accurately obtain the sequence in the case of low-frequency mutations and mutation deletions.
[0005] The 1.5-generation sequencing technology of Beijing Tsingke overcomes these limitations. It is developed based on NGS molecular barcoding technology and error elimination algorithms. It is a sequencing technology that can accurately analyze long fragments over 1 kb without primers. The single sequencing can reach 20 kb, and it has the advantages of the accuracy of Sanger first-generation sequencing and the high throughput of second-generation sequencing. However, the 1.5-generation sequencing of Beijing Tsingke has a very long cycle (10 working days) and is still relatively expensive.
[0006] KBSeq of Shanghai Sangon uses second-generation sequencing technology to construct a library for plasmid sequencing. Compared with first-generation sequencing, KBSeq does not require the design and synthesis of sequencing primers. The single sequencing can obtain a plasmid sequence with a total length of up to 30 kb at most. The sequencing cycle is relatively short, but the cost is still high (5 - 7 working days, 38 - 98 yuan / sample for 2 - 30 kb plasmids). And KBSeq inherits the disadvantages of second-generation sequencing. The sequencing effect on characteristic regions such as high GC, polymers, and repetitive regions is not good. Plasmids with such characteristic regions may not be sequenced to the full length.
[0007] The research group of Zhang Yong from the Institute of Zoology, Chinese Academy of Sciences and the research group of Ruan Jue from the Shenzhen Institute of Agricultural Genomics, Chinese Academy of Agricultural Sciences developed a PacBio library construction technology LILAP (low-input, low-cost, and amplification-free library-generation method for PacBio sequencing) based on Tn5 transposase with low DNA usage (100 ng), no amplification, and low cost. LILAP uses Tn5 transposase and PacBio sequencing adapters with hairpin structures to form a dimer transposase complex, achieving DNA fragmentation and sequencing adapter ligation in one step. All processes are carried out in a single test tube, minimizing DNA loss during the library construction process and simplifying the library construction process. Tags can also be added to the sequencing adapters for multiplexed library construction and sequencing. However, its library construction cost is still relatively high (up to 10 US dollars / sample).
[0008] ONT in the UK uses MinION™ sequencing chips on MinION or GridION™ sequencing devices, paired with the EPI2ME analysis platform for sequence assembly in plasmid construction. The nanopore sequencing technology used in ONT's MinION™ sequencing chips belongs to long-read sequencing, which is friendly to plasmid repeat regions and can perform highly accurate, flexible, and safe sequencing of full-length plasmids. The results can be obtained within hours without sending construction information to a third party for verification. Moreover, complete sequence data can be obtained from a single experiment, eliminating the need for multiple technologies to confirm the accuracy of the construction. However, the experimental process of ONT plasmid sequencing requires plasmid extraction, and currently, only kits with a maximum throughput of 96 tags are provided, with a library construction cost of $8.39 per sample. Compared with other existing technologies, whole-plasmid sequencing based on nanopore sequencing technology only relies on its sequencing platform to shorten the sequencing cycle and does not solve the problems of low throughput and high cost in existing technologies.
[0009] Therefore, in view of the sequencing characteristics of nanopore sequencing technology, combined with sequences such as expression or genetic characteristic control elements and low-repeat sequences contained in plasmids that can set highly specific primers, taking the sequencing samples stored in 96-well plates as an example, it is particularly important to develop a design method for two-dimensional DNA molecular tags and amplification primers applied to whole-plasmid sequencing using the design method of two-dimensional DNA molecular tags and the corresponding sequencing method disclosed in the patent with the application number CN202410505390.1. Summary of the Invention
[0010] Aiming at the deficiencies of existing whole-plasmid sequencing technologies, the present invention provides a design method for two-dimensional DNA molecular tags and amplification primers in whole-plasmid sequencing.
[0011] To solve the above technical problems, the present invention provides the following technical solutions: A design method for two-dimensional DNA molecular tags and amplification primers in whole-plasmid sequencing, wherein the design method is selected from one of the following (Ⅰ) to (Ⅲ): (Ⅰ) The ligation combination of two-dimensional DNA molecular tags and amplification primers; (Ⅱ) The ligation combination of two-dimensional DNA molecular tags with sticky ends of double restriction enzyme sites, sticky ends of double restriction enzyme sites, and amplification primers; (Ⅲ) The ligation combination of two-dimensional DNA molecular tags and homologous arms; The amplification primers are characteristic elements of the plasmid itself, low-repeat sequences, or single-stranded DNA fragments with a GC percentage content of 30 - 80%; the two-dimensional DNA molecular tags include tag 1 and tag 2.
[0012] Furthermore, the characteristic elements include resistance genes or expression regulation genes contained in the plasmid itself.
[0013] Further, the design method (Ⅰ) includes the following steps: Determine the forward primer sequence and the reverse primer sequence of the amplification primer; Connect tag 1 and the forward primer sequence into a single-stranded DNA fragment as the forward primer fragment, and connect tag 2 and the reverse primer sequence into a single-stranded DNA fragment as the reverse primer fragment.
[0014] Further, the design method (Ⅱ) includes the following steps: (A) Determine the forward primer sequence and the reverse primer sequence of the amplification primer, Connect the protection base of the restriction enzyme site, the base sequence of the restriction enzyme site and the forward primer sequence into a single-stranded DNA fragment as the forward amplification primer, and connect the protection base of the restriction enzyme site, the base sequence of another restriction enzyme site and the reverse amplification primer sequence into a single-stranded DNA fragment as the reverse amplification primer, and then obtain a linear plasmid DNA with restriction enzyme sites at both ends by PCR amplification; (B) Digest the obtained linear plasmid DNA with two restriction endonucleases to obtain a linear plasmid DNA with sticky ends containing different restriction enzyme sites at both ends; (C) Add the reverse complementary sequence of one of the sticky ends of the restriction enzyme site in step (B) to the ends of the double-stranded DNA of tag 1 and its complementary strand, and add the reverse complementary sequence of the other sticky end of the restriction enzyme site in step (B) to the ends of the double-stranded DNA of tag 2 and its complementary strand, and then ligate them to both ends of the digested linear plasmid DNA obtained in step (B) respectively by ligase to obtain a linear plasmid DNA containing tag 1 and tag 2.
[0015] Further, the restriction enzyme sites in step (1) and step (3) are the cleavage sites of the two restriction endonucleases in step (2).
[0016] Further, the protection base of the restriction enzyme site is 1 to 5 bases to ensure the effective binding and digestion of the restriction endonuclease with its recognition site.
[0017] Further, the design method (Ⅲ) includes the following steps: (a) Determine the forward primer sequence and the reverse primer sequence of the amplification primer, and determine the corresponding homologous arm sequences according to the sequences at both ends of the linear plasmid amplified by the forward and reverse primers, which are homologous arm sequence 1 and homologous arm sequence 2 respectively; (c) Connect tag 1 and homologous arm sequence 1, synthesize double-stranded DNA 1 containing the reverse complementary strand, connect tag 2 and homologous arm sequence 2, synthesize double-stranded DNA 2 containing the reverse complementary strand, and ligate them to both ends of the linear DNA amplified by the forward and reverse primers in step (a) respectively by homologous arm recombination.
[0018] A high-throughput sequencing method for whole plasmids. According to the above-mentioned design method of two-dimensional DNA molecular tags and amplification primers in whole plasmid sequencing, linear plasmid DNA containing two-dimensional DNA molecular tags is obtained, and then high-throughput plasmid sequencing is carried out using nanopore sequencing technology.
[0019] The present invention has the following beneficial effects: 1. The design method of two-dimensional DNA molecular tags and amplification primers in whole plasmid sequencing provided by the present invention combines the method of amplification primers and two-dimensional DNA molecular tags, and applies the nanopore sequencing platform to high-throughput full-length sequencing of mutant library plasmids, high-throughput verification of full-length gene synthesis plasmids, high-throughput construction of full-length plasmid sequencing, sequence identification of known sequence information plasmids or plasmids with incomplete information stored in molecular biology research institutions or enterprises, etc. It greatly expands the sequencing application scenarios of the nanopore sequencing platform, making it no longer limited to genomic sequencing and no longer limited to sequencing of at most 96 plasmid mixed samples. At the same time, classification sequencing is carried out according to plasmid information and sequencing throughput, making high-throughput whole plasmid sequencing simple and orderly.
[0020] 2. The present invention designs different specific amplification primers based on known elements, low-repeat sequences or other DNA sequences with appropriate GC% contained in plasmid vectors, and uses different amplification primers to achieve efficient PCR amplification of multiple types of samples of the same type of plasmid or different types of plasmids.
[0021] 3. Through the coordinated design of two-dimensional DNA molecular tags and amplification primers used for plasmid linearization, the present invention provides three methods for linearizing plasmids and adding tags (i.e., the design methods (I), (II), and (III) of two-dimensional DNA molecular tags and amplification primers described in the invention content). Selecting the corresponding kits for each of the three methods can add tags to obtain linear plasmid DNA containing tags. The linearization of plasmids and the addition of tags are not limited to the above three methods.
[0022] 4. By connecting different types of amplification primers and two-dimensional DNA molecular tags, the amplification products from different sources can be mixed into one sample, and subsequent library construction operations based on nanopore sequencing technology can be carried out to complete sequencing. While achieving ultra-high-throughput whole plasmid sequencing, it greatly reduces reagent consumables, time, and labor costs. Description of the Drawings
[0023] Figure 1 Schematic diagram of the design principle of two-dimensional DNA molecular tags and primers for whole plasmid sequencing; Figure 2 Schematic diagram of the design principle of two-dimensional DNA molecular tags and primers in plasmids containing characteristic elements; Figure 3Schematic diagram of linear plasmid DNA synthesized during whole plasmid sequencing of plasmids containing characteristic elements; Figure 4 Schematic diagram of the design principle of two-dimensional DNA molecular tags, sticky ends of double digestion sites, and primers in Example 2; Figure 5 Schematic diagram of the principle of whole plasmid sequencing of two-dimensional DNA molecular tags and homologous arm sequences in Example 3; Figure 6 Electrophoresis test result diagram of Example 5 of the present invention. Detailed implementation manners
[0024] Specifically, the "two-dimensional DNA molecular tag" in the present invention has the same meaning as the "two-dimensional DNA molecular tag" mentioned in the patent with the application number CN202410505390.1 that has been published. Specifically, the meaning of the two-dimensional DNA molecular tag is disclosed in this patent: The two-dimensional DNA molecular tag includes tag 1 and tag 2. Tag 1 and tag 2 are DNA base sequences with the same or different lengths. The DNA bases are any one of the four bases A, T, G, and C. The lengths of tag 1 and tag 2 are selected according to the requirements of sequencing throughput, and the number of bases selected is m. The number of repetitions of tag 1 and tag 2 is n. Finally, the lengths of tag 1 and tag 2 are the product of m and n. The number of repetitions n is at least 3 times. G and C do not exist simultaneously in the bases of tag 1 and tag 2. Tag 1 and tag 2 respectively contain three bases A, T, G or A, T, C with different arrangements. Tag 1 represents the well number; specifically, tag 1-1 represents the molecular tag corresponding to the first well position on the microplate, and so on. Tag 1-96 represents the molecular tag corresponding to the 96th well position on the microplate. Tag 2 represents the plate number; specifically, tag 2-1 represents the molecular tag corresponding to the first plate in a group of microplates, and so on. Tag 2-n represents the molecular tag corresponding to the nth plate in this group of microplates.
[0025] At the same time, the above patent discloses the application method of molecular tags in nanopore high-throughput sequencing, and its steps are as follows: (1) Synthesize the sequence of tag 1 together with the upstream of the forward primer, so that tag 1 becomes a part of the forward primer; synthesize the sequence of tag 2 together with the upstream of the reverse primer, so that tag 2 becomes a part of the reverse primer; (2) Add the forward primer and reverse primer in step (1) to the corresponding well positions of the high-throughput sequencing samples (96-well plate, 384-well plate or 1536-well plate) respectively, and then perform PCR amplification to obtain PCR amplification products; (3) Mix the PCR amplification products of the multiple samples to be tested obtained in step (2), and then ligate the sequencing adapters provided by the nanopore sequencing equipment manufacturer to perform on-machine sequencing; (4) Analyze the sequencing results to determine the base sequence of the sample to be tested. The mixed PCR amplification products include multiple samples, each sample corresponding to a set of tag sequence combinations including tag 1 and tag 2, and the tag sequence combinations corresponding to different samples are different; The sequencing results are a mixture of multiple sequencing fragments, and the samples actually sequenced are determined by determining the tag sequence combinations of each sequencing fragment.
[0026] Based on the two-dimensional DNA molecular tags disclosed above, the present application provides a design method for two-dimensional DNA molecular tags and amplification primers, that is, the content involved in step (1) of the application method of the above molecular tags in nanopore high-throughput sequencing: adding the sequence of tag 1 upstream of the forward primer for synthesis together, making tag 1 a part of the forward primer; adding the sequence of tag 2 upstream of the reverse primer for synthesis together, making tag 2 a part of the reverse primer; this part of the content.
[0027] The design method of two-dimensional DNA molecular tags and amplification primers in whole plasmid sequencing is selected from one of the following (Ⅰ) to (Ⅲ): (Ⅰ)The ligation combination of two-dimensional DNA molecular tags and amplification primers; (Ⅱ)The ligation combination of two-dimensional DNA molecular tags, sticky ends of double restriction enzyme sites and amplification primers; (Ⅲ)The ligation combination of two-dimensional DNA molecular tags and homologous arms; The amplification primer is a characteristic element of the plasmid itself, a low-repeat sequence or a single-stranded DNA fragment with a GC percentage content of 30-80%.
[0028] The amplification primers of the present invention generally come from other sequences with appropriate GC% content such as characteristic elements and low-repeat sequences to ensure their Tm values and specificities. The specific design principle of whole plasmid sequencing primers is as Figure 1 shown. For plasmids with the same amplification primers, the sequencing sequences are directly assigned to different plasmids through two-dimensional DNA molecular tags. For different types of plasmids with different amplification primers, whole plasmid amplification can be carried out synchronously. The obtained sequencing sequences are first assigned to the corresponding types of plasmids through different amplification primers, and then the sequencing sequences are assigned to the corresponding plasmids according to the two-dimensional DNA molecular tags. As Figure 2As shown in the figure, plasmid of type A has kanamycin resistance, and the amplification primers Vs_Kan-F / R designed according to the kanamycin resistance gene can be used. Plasmid of type B is induced by lactose for expression, and the amplification primers Vs_LacI-F / R designed according to the LacI gene can be used. For example, when there are hundreds or thousands of plasmid DNAs in the same batch to be sequenced, such as gene mutant libraries, gene synthesis verification, etc., but the plasmid DNA sequences only differ in the inserted genes at the multiple cloning sites, or even only in the mutation sites on the genes, then these hundreds or thousands of plasmid DNAs can uniformly select the amplification primers designed according to the vector plasmid DNA sequence for amplification to obtain linearized plasmid DNAs containing different inserted gene sequences. If the known information of the plasmid to be sequenced or the strain carrying the plasmid to be tested is single or very little, such as the known information is only that the strain is resistant to kanamycin or the inducer used is IPTG or lactose, then Vs_Kan-F / R or Vs_LacI-F / R can be used to amplify the unknown plasmid respectively to obtain linearized plasmids. The above application scenarios fully reflect the practicability of the amplification primer design scheme in the present invention. After amplification, the amplification products that have passed the electrophoresis detection are mixed in equal proportions, and A tails and adapters are added to obtain linear plasmid DNAs containing sequencing adapters for sequencing on the machine. After amplification with different amplification primers, linear plasmid DNAs containing sequencing adapters are obtained, as specifically shown in Figure 3 as shown.
[0029] Definition of terms used in the present invention: Unless otherwise specified, the initial definitions provided for the terms in this article apply to the terms throughout the specification; for terms not specifically defined in this article, meanings that can be given to them by those skilled in the art should be given according to the disclosed content and context. The reagents and equipment used in the present invention are all known products or obtained by purchasing commercially available products. The technical solutions of the present invention will be clearly and completely described below in conjunction with specific embodiments.
[0030] Example 1 The connection combination of two-dimensional DNA molecular tags and amplification primers. For this method of tag addition, in this example, the plasmid to be detected shows kanamycin (kan) resistance as a characteristic element of the plasmid itself. According to this characteristic, the amplification primers Vs_kan-f1 and r1 for tag addition are designed at the resistance gene, and examples are as follows.
[0031] In this example, the sequence of tag 1 and the forward primer are referred to as the "forward primer fragment", and the sequence of tag 2 and the reverse primer are referred to as the "reverse primer fragment".
[0032] The base sequence of Vs_kan-f1 (forward primer fragment) is: 5’-GAGTAGATGAGTAGATGAGTAGAT GAATCGCAGACCGATACCAGGATC-3’ Among them, "5’-GAGTAGATGAGTAGATGAGTAGAT-3" is the base sequence of tag 1-1 contained in Vs_kan-f1, and "5’-GAATCGCAGACCGATACCAGGATC-3" is the forward primer sequence for amplification contained in Vs_kan-f1; The base sequence of Vs_kan-r1 (reverse primer fragment) is: 5’-GGATTGATGGATTGATGGATTGAT CGACTCGTCCAACATCAATACAACC -3’ Among them, "5’-GGATTGATGGATTGATGGATTGAT-3’" is the base sequence of tag 2-1 contained in Vs_kan-r1, and "5’-CGACTCGTCCAACATCAATACAACC-3’" is the reverse primer sequence for amplification contained in Vs_kan-r1.
[0033] Example 2 Connection combination of two-dimensional DNA molecular tags, sticky ends of double restriction enzyme sites, and amplification primers In the method of adding tags, the plasmid in this example contains a lactose-inducible expression regulatory gene (LacI). The schematic diagram of the design principle of two-dimensional DNA molecular tags, sticky end sequences of double restriction enzyme sites, and amplification primers is as Figure 4 shown. Taking Vs_LacI-f1 / r1 as an example, (A) Determine the forward primer sequence and reverse primer sequence of the amplification primer: The base sequence of Vs_LacI-f1 is: 5’-NNNNN CTCGAG CAACCAGCATCGCAGTGGG -3’. Among them, "5’-CAACCAGCATCGCAGTGGG-3’" is the base sequence of the forward primer designed according to the LacI gene contained in Vs_LacI-f1, "5’-CTCGAG—3’" is the base sequence of the XhoI restriction enzyme cleavage site attached to the 5’ end of the forward amplification primer contained in Vs_LacI-f1, and "NNNNN" is the restriction enzyme cleavage site protection base. Considering comprehensively in combination with the primer design principle, it is generally 1 to 5 bases. The purpose of the restriction enzyme cleavage site protection base is to ensure the effective binding and cleavage of the restriction endonuclease to its recognition site.
[0034] The base sequence of Vs_LacI-r1 is: 5’-NNNNNAGATCT CCAACGATCAGATGGCGCTG -3’, Among them, "5’-CCAACGATCAGATGGCGCTG-3’" is the base sequence of the reverse primer designed according to the LacI gene contained in Vs_LacI-r1, "5’-AGATCT-3’" is the base sequence of the BglII restriction enzyme cleavage site attached to the 5' end of the reverse amplification primer contained in Vs_LacI-r1, and "NNNNN" is the base for protecting the restriction enzyme cleavage site. Considering the primer design principle comprehensively, it is generally 1 to 5 bases to ensure the effective binding and cleavage of the restriction endonuclease to the recognition site.
[0035] The linear DNA fragment obtained after PCR amplification is as follows:
[0036] (B) The linear plasmid DNA with different sticky ends at both ends after digestion with XhoI and BglII is as follows:
[0037] (C) Tag 1-1 with a sticky end containing the XhoI restriction enzyme cleavage site and its complementary strand:
[0038] Tag 2-1 with a sticky end containing the BglII restriction enzyme cleavage site and its complementary strand:
[0039] Add the tag with a sticky end containing the restriction enzyme cleavage site to the ends of the linear plasmid DNA with different sticky ends obtained in step (B) above. If using the T4 DNA ligase from NEB company, the 20 μL ligation system: 2 μL 10X T4 DNA Ligase Buffer (10X T4 DNA ligase buffer), 0.02 pmol of linear plasmid DNA with different sticky ends, 0.06 pmol of tag 1 with a sticky end, 0.06 pmol of tag 2 with a sticky end, 1 μL of T4 DNA ligase, and Nuclease-free water (ultrapure water without nuclease contamination) to make up 20 μL. Ligation program: 2 h at room temperature.
[0040] Taking Vs_LacI-f1 / r1 as an example, the sequence schematic diagram is as follows:
[0041] Example 3 The meaning of the homologous arm: The homologous arm refers to the sequence that is exactly the same as both sides of the target gene sequence. The homologous arm 1 sequence contained in tag 1 is exactly the same as the 5' end sequence of the target gene, and the homologous arm 2 sequence contained in tag 2 is exactly the same as the 3' end sequence of the target gene. It is the region where the tag recognizes the target gene and recombination occurs.
[0042] The ligation combination of two-dimensional DNA molecule tags and homologous arms, in the tag addition method, the addition of forward and reverse amplification primers and tags containing homologous arms, as Figure 5 shown. Taking Vs_P15Aori-f1 and r1 as examples, (a) Determine the forward primer sequence and reverse primer sequence of the amplification primer The base sequence of the forward primer Vs_P15Aori-f1 is: 5’ -CGCGTTTGTCTCATTCCACGC -3’, The base sequence of the reverse primer Vs_P15Aori-r1 is: 5’ -GCCATAACAGCGGAATGACACCG -3’, The linear DNA fragment obtained by PCR amplification:
[0043] (b) The base sequence of tag 1-1 containing homologous arm 1 of the Vs_P15Aori-f1 sequence and its complementary strand are:
[0044] Among them, 5’-GAGTAGATGAGTAGATGAGTAGAT-3’ is the base sequence of tag 1-1, 5’-CGCGTTTGTCTCATTCCACGC-3’ is the base sequence of homologous arm 1; The base sequence of tag 2-1 containing homologous arm 2 of the Vs_P15Aori-r1 sequence and its complementary strand are:
[0045] Among them, 5’-GGATTGATGGATTGATGGATTGAT-3’ is the base sequence of tag 2-1, and 5’-GCCATAACAGAGGAAGACACCG-3’ is the base sequence of homologous arm 2; The structure of the linear DNA fragment containing tag 1 and tag 2 obtained after splicing is as follows:
[0046] Example 4 Bacterial liquid culture Transform the prepared mutant library DNA into Escherichia coli BL21(DE3), plate it on LB solid medium containing the appropriate antibiotic, and culture it overnight at 37°C. Select a plate with an appropriate colony density, pick colonies into a 96-well microplate, and dispense 200 μL of LB liquid medium containing the corresponding antibiotic into each well in advance. Culture it overnight on a shaker at 180 rpm and 30°C. Transfer 100 μL of the bacterial liquid to a new 96-well microplate for sequencing.
[0047] Example 5 PCR amplification Prepare a PCR reaction system, which includes: 1x reaction mix, 1.25 U DNA polymerase, 0.2 μM F primer, 0.2 μM R primer, and 4 μL of the overnight culture bacterial liquid diluted 5-fold.
[0048] PCR program: (1) Pre-denaturation at 95°C for 5 min; (2) Pre-denaturation at 98°C for 1 min; (3) Denaturation at 98°C for 10 s; (4) Annealing at 55°C - 72°C for 10 s and extension at 72°C for 1 min (6 kb - 12 kb / min); steps (3) - (4) are repeated 20 - 30 times; (5) Continue to extend at 72°C for 5 min and store at 4°C.
[0049] Sampling and detection: Test by 1% agarose gel electrophoresis, and the electrophoresis test results are as Figure 6 shown. Among them, the electrophoresis detection sample: the PCR product after the PCR program runs (the linear plasmid to be tested with a label). After obtaining the amplified band, mix and store it for preparing the sequencing library.
[0050] Example 6 Purify the mixed linear plasmids to be tested with labels Purify the reaction product with 1x magnetic beads (BEAVER BEADS). Transfer 100 µL of the PCR product mixture finally obtained in Example 5 into a 1.5 mL EP tube, add 100 µL of resuspended magnetic beads, vortex and mix well, then centrifuge instantaneously. Incubate at room temperature for 5 - 10 min to allow the DNA to bind fully to the magnetic beads. Place the EP tube on a magnetic stand until the magnetic beads are separated from the liquid phase, and remove the supernatant. Add 200 µL of 80% ethanol to wash the magnetic beads, remove the supernatant after washing for about 30 - 60 s, repeat the washing once, and try to remove the residual liquid in the tube with a pipette. After the magnetic beads are air-dried (about 30 s), add 50 µL of NFW (Nuclease-free water) to resuspend the magnetic beads, incubate at room temperature for 5 min, place the EP tube on the magnetic stand until the magnetic beads are separated from the liquid phase and the liquid phase is clear and colorless, about 1 min. Transfer the supernatant to a new 1.5 mL Ep tube, and measure the library concentration: 56.6 ng / µL.
[0051] Example 7 End repair / dA tail addition and adapter ligation End repair / dA tail addition: Use NEBNext Ultra Ⅱ end repair enzyme and NEBNext FFPE repair enzyme to perform end repair and dA addition on the purified nucleic acid sample in the previous step. 60 µL reaction system: 3.5 µL of NEBNext Ultra Ⅱ end repair reaction buffer, 3 µL of NEBNext Ultra Ⅱ end repair enzyme mixture, 3.5 µL of NEBNext FFPE repair buffer, 2 µL of NEBNext FFPE repair mixture, and take 1000 - 1500 ng of nucleic acid sample. Incubate at 20 °C for 5 min and at 65 °C for 5 min. Purify the reaction product with 1x magnetic beads, and the steps are the same as the purification in S2.
[0052] Using the Ligation Sequencing Kit V14 (the chip model corresponding to the ligation kit: SQK-LSK114), configure a 100 µL reaction system according to the kit instructions: 25 µL of ligation buffer (LNB), 10 µL of NEBNext quick T4 ligase, 5 µL of sequencing adapter F (LA), and 60 µL of the nucleic acid treated by end repair above. Incubate at room temperature for 10 min. Add 40 µL of resuspended magnetic beads, vortex and mix well, then centrifuge instantaneously, and incubate at room temperature for 5 min to allow the DNA to bind fully to the magnetic beads. Place the EP tube on the magnetic rack until the magnetic beads are separated from the liquid phase, and remove the supernatant. Add 200 µL of wash buffer SFB to wash the magnetic beads, remove the supernatant after washing for about 30 - 60 s, repeat the washing once, and try to remove the residual liquid in the tube using a pipette. After the magnetic beads are air-dried (about 30 s), add 15 µL of NFW to resuspend the magnetic beads, incubate at room temperature for 5 - 10 min, place the EP tube on the magnetic rack until the magnetic beads are separated from the liquid phase and the liquid phase is clear and colorless, about 1 min. Transfer the supernatant to a new 1.5 mL Ep tube, and measure the library concentration: 26.8 ng / µL.
[0053] Example 8 Library loading Library loading: For the requirements of loading samples, 37.5 µL of sequencing buffer II (SBII), 25.5 µL of loading beads II (LBII). SBII and LBII are two tubes of reagents, both from the Ligation Sequencing Kit V14; 72 ng of the above DNA library with adapters added, and NFW is added to make up 75 µL. Use the ONT sequencing chip and sequencer to sequence and analyze the library. The specific operation process refers to the ONT product manual. The MinKNOW software is used for real-time data acquisition, and then the data is analyzed such as basecalling to obtain the sequencing results.
[0054] Example 9. Data splitting and analysis: The data is assigned to the corresponding plasmid types according to different amplification primers, and the data obtained by basecalling is split into each well of each plate according to the Barcode representing the plate number and well number. Use alignment data comparison to obtain single consensus sequences.
[0055] It should be understood that the disclosed invention is not limited to the specific methods, schemes, and substances described, as these can vary. In addition, for the molecular tags with the number of bases determined in the present invention, all possible molecular tags can be presented in full through a programmed procedure, and then the molecular tags that meet the requirements can be screened and synthesized together with the primer sequences. It should also be understood that the terms used herein are for the purpose of describing specific embodiments only and are not intended to limit the scope of the present invention, which is limited only by the appended claims.
[0056] Those skilled in the art will also recognize, or be able to ascertain, many equivalents to the specific embodiments of the invention described herein using no more than routine experimentation. These equivalents are also encompassed by the appended claims.
Claims
1. A method for designing two-dimensional DNA molecular tags and amplification primers in whole plasmid sequencing, characterized in that: The design method is selected from one of the following (I) to (III): (I) A connection combination of a two-dimensional DNA molecular tag and an amplification primer; (II) A connection combination of a two-dimensional DNA molecular tag, a sticky end with double restriction sites, and an amplification primer; (III) The connection combination of two-dimensional DNA molecular tags and homology arms; The amplification primer is a characteristic element of the plasmid itself, a low-repetitive sequence or a single-stranded DNA fragment with a GC percentage of 30-80%; the two-dimensional DNA molecular label includes label 1 and label 2.
2. The method for designing two-dimensional DNA molecular tags and amplification primers in whole plasmid sequencing according to claim 1, characterized in that: The characteristic elements include resistance genes or expression regulatory genes contained in the plasmid itself.
3. The method for designing two-dimensional DNA molecular tags and amplification primers in whole plasmid sequencing according to claim 1, characterized in that: The design method (I) comprises the following steps: Determine the forward primer sequence and reverse primer sequence of the amplification primer; The label 1 and the forward primer sequence are connected to form a single-stranded DNA fragment as the forward primer fragment, and the label 2 and the reverse primer sequence are connected to form a single-stranded DNA fragment as the reverse primer fragment.
4. The method for designing two-dimensional DNA molecular tags and amplification primers in whole plasmid sequencing according to claim 1, characterized in that: The design method (II) comprises the following steps: (A) Determine the forward primer sequence and reverse primer sequence of the amplification primer. The specific method is as follows: The protective base of the restriction site, the base sequence of the restriction site and the forward primer sequence are connected to form a single-stranded DNA fragment as a forward amplification primer, and the protective base of the restriction site, the base sequence of another restriction site and the reverse amplification primer sequence are connected to form a single-stranded DNA fragment as a reverse amplification primer, and then PCR amplification is performed to obtain a linear plasmid DNA containing restriction sites at both ends; (B) The obtained linear plasmid DNA is digested with two restriction endonucleases to obtain linear plasmid DNA with sticky ends containing different restriction endonucleases at both ends; (C) Add the reverse complementary sequence of the sticky end of one restriction site in the above step (B) to the end of the double-stranded DNA of tag 1 and its complementary chain, and add the reverse complementary sequence of the sticky end of another restriction site in the above step (B) to the end of the double-stranded DNA of tag 2 and its complementary chain, and then connect them to the two ends of the linearized plasmid DNA obtained after restriction digestion in the above step (B) by ligase to obtain linearized plasmid DNA containing tag 1 and tag 2.
5. The method for designing two-dimensional DNA molecular tags and amplification primers in whole plasmid sequencing according to claim 4, characterized in that: The protective bases of the restriction site are 1 to 5 bases to ensure that the restriction endonuclease effectively binds to its recognition site and cuts.
6. The method for designing two-dimensional DNA molecular tags and amplification primers in whole plasmid sequencing according to claim 1, characterized in that: The design method (III) comprises the following steps: (a) determining the forward primer sequence and the reverse primer sequence of the amplification primers, and determining the corresponding homology arm sequences according to the sequences at both ends of the linear plasmid obtained by amplification of the forward and reverse primers, which are homology arm sequence 1 and homology arm sequence 2 respectively; (b) Label 1 and homology arm sequence 1 are connected to synthesize double-stranded DNA 1 containing a complementary chain, and label 2 and homology arm sequence 2 are connected to synthesize double-stranded DNA 2 containing a complementary chain, which are respectively connected to the two ends of the linear DNA amplified in the above step (a) through homology arm recombination.
7. A high-throughput sequencing method for obtaining full-length sequences of multiple different plasmids, characterized in that: According to the design method of two-dimensional DNA molecular tags and amplification primers in whole plasmid sequencing described in any one of claims 1 to 6 above, linear DNA samples of multiple different plasmids containing two-dimensional DNA molecular tags are obtained, and then their full lengths are sequenced using nanopore sequencing technology.
Citation Information
Patent Citations
Two-dimensional DNA molecular tag and application thereof in nanopore high-throughput sequencing
CN118166083A