A method for constructing a DNA barcode next-generation sequencing library and application thereof

By employing two rounds of PCR amplification and purification, a DNA barcode next-generation sequencing library was constructed, which solved the problems of low throughput and high cost in mixed sample analysis and enabled accurate species identification and determination of mixing ratios.

CN116200832BActive Publication Date: 2026-02-17MGI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111440727.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-30
Publication Date
2026-02-17
Estimated Expiration
2041-11-30

AI Technical Summary

Technical Problem

Existing DNA barcoding next-generation sequencing technology suffers from problems such as low throughput, high cost, low data utilization, and difficulty in determining the pooling ratio in mixed sample analysis.

Method used

A two-round PCR amplification method was used. In the first round of amplification, a random sequence tag was added, and in the second round of amplification, a library tag and sequencing primer binding sequence were added. The DNA barcode region was amplified by specific primers to construct a DNA barcode next-generation sequencing library, which was then purified.

Benefits of technology

It enables accurate identification of mixed samples and determination of mixing ratios, simplifies the operation process, reduces costs, and is suitable for the analysis of mixed samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0003382768700000051
    Figure BDA0003382768700000051
  • Figure BDA0003382768700000061
    Figure BDA0003382768700000061
  • Figure BDA0003382768700000071
    Figure BDA0003382768700000071
Patent Text Reader

Abstract

The application discloses a DNA barcode second-generation sequencing library construction method and application thereof. The method is simple and rapid, and can be used for preparing a sequencing library of a barcode DNA through two rounds of PCR, and can be used for assembling full-length DNA barcodes, accurately identifying species and determining a mixed ratio, and is especially suitable for mixed sample analysis. The sequencing library can be directly used for second-generation sequencing, is especially suitable for mixed sample identification, and has simple and rapid operation and low cost. The application has important application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of bioinformatics, specifically relating to a method for constructing a DNA barcode next-generation sequencing library and its application. Background Technology

[0002] DNA barcoding is a molecular biology technique that uses a recognized, relatively short DNA sequence from the genome for species identification. It is widely used in the identification of different biological groups and serves as an auxiliary and effective supplement to traditional morphological identification methods. This method involves screening universal barcodes, establishing a barcode database and identification platform, and using bioinformatics methods to analyze and compare DNA data to identify species. For example, the mitochondrial cytochrome C oxidase subunit I (CO I) gene fragment is a commonly used marker sequence for animal group identification; genes with faster evolutionary rates, such as Cytb, can be used for molecular tracing of subspecies, sister species, or different geographical populations; analysis of 16S rRNA genes can yield taxonomic characteristics of various bacteria and is often used for microbial community composition analysis; my country was the first to propose the ITS2 sequence as a universal sequence for identifying medicinal plants and established a molecular identification system for medicinal plants with ITS2 as the core and psbA-trnH as supplementary sequences.

[0003] DNA barcode identification typically involves designing primers based on conserved sequences at both ends of the barcode DNA for PCR amplification, followed by Sanger sequencing or next-generation sequencing, and then aligning the DNA sequences to a database for species identification. However, Sanger sequencing has relatively low throughput and can only analyze single species, making it unsuitable for mixed sample types. The general library construction process based on next-generation sequencing includes amplification of the target fragment, purification of the amplified product, fragmentation of the purified product (including ultrasonic fragmentation, enzymatic fragmentation, etc.), end repair, 3'-end A addition, and adapter ligation. This entire process is complex and relatively expensive. Furthermore, due to the small sequence differences between species and the presence of conserved regions, data utilization is low, and identifying mixed samples is relatively difficult, making it impossible to determine the mixing ratio. Summary of the Invention

[0004] The purpose of this invention is to construct DNA barcode next-generation sequencing libraries that can be used for mixed sample analysis or multi-species analysis.

[0005] This invention provides the first protection of a method for constructing a DNA barcode next-generation sequencing library, comprising the following steps:

[0006] (1) Using the DNA to be tested as a template, the first round of PCR amplification was performed using primer pairs to obtain the first round of amplicon;

[0007] The primer pair consists of a forward primer and a reverse primer;

[0008] The forward primer consists of a first universal sequence, a random sequence tag, and a first specific sequence from the 5' end to the 3' end; the reverse primer consists of a second specific sequence.

[0009] (2) After completing step (1), the first round of amplicon is amplified by a second round of PCR using primer combinations to obtain the second round of amplicon;

[0010] The primer combination consists of primer 1, primer 2 and primer 3;

[0011] Primer 1 includes the first universal sequence; primer 2 consists of the second specific sequence, the second universal sequence, and a random sequence from the 5' end to the 3' end; primer 3 consists of the library tag sequencing primer binding sequence, the library tag sequence, and the second universal sequence from the 5' end to the 3' end.

[0012] (3) After completing step (2), take the second round of amplicon, purify it, and obtain the DNA barcode next-generation sequencing library.

[0013] In the above method, the first round of PCR amplification is used for specific amplification with random sequence tags. The second round of PCR amplification is used for library amplification.

[0014] In the above method, the forward primer and the reverse primer specifically amplify the DNA barcode region in the template through the first specific sequence and the second specific sequence (a unique tag is added to the 5' end of each amplicon by a random sequence tag).

[0015] In the above method, during the second round of PCR amplification, primer 2 is anchored to the reverse strand of the second specific sequence of the first round amplicons via a second specific sequence. Based on the foldability of the DNA strand in three-dimensional space, the random sequence is complementary to the bases at any position in the DNA barcode region and extended. Primer 1 is anchored to the reverse strand of the first universal sequence of the first round amplicons via a first universal sequence and extended. Products of unequal fragment sizes are obtained by amplification using primers 1 and 2. Each amplification product has a first universal sequence and a unique tag originating from a different template at the 5' end, and a second universal sequence at the 3' end. Primer 3 is anchored to the reverse strand of the second universal sequence of the above products via a second universal sequence and extended, adding a library tag sequence and a library tag sequencing primer binding sequence to the above products.

[0016] In the above method, both the first universal sequence and the second universal sequence are sequencing universal adapter sequences and are from a different source than the template.

[0017] In the above method, in step (1), the random sequence tag can be a base N with a length of 10-20bp (such as 10-15bp, 15-20bp, 10bp, 15bp or 20bp), where N is any one of A, T, G and C.

[0018] In the above method, during step (1) of the first round of PCR amplification, the forward and reverse primers are amplified in 2-4 cycles (e.g., 2, 3, 4) to obtain the DNA barcode region sequence.

[0019] In the above method, in step (2), the random sequence can be a base N with a length of 6-15bp (such as 6-10bp, 10-15bp, 6bp, 10bp or 15bp), where N is any one of A, T, G and C.

[0020] In the above method, the purpose of purification in step (3) is to remove short DNA fragments.

[0021] In any of the methods described above, the second-generation sequencing library is adapted for paired-end sequencing.

[0022] In any of the methods described above, the DNA to be tested may be the DNA of a species or a sample.

[0023] The present invention also protects a kit for constructing DNA barcode next-generation sequencing libraries of species or samples, comprising any of the above-described forward primers, any of the above-described reverse primers, any of the above-described primer 1, any of the above-described primer 2, and any of the above-described primer 3.

[0024] The kit may specifically consist of the forward primer, the reverse primer, primer 1, primer 2 and primer 3.

[0025] The application of any of the above-described construction methods or reagent kits in the identification of species or samples also falls within the scope of protection of this invention.

[0026] The application of any of the above-described construction methods or reagent kits in detecting species purity or sample purity also falls within the scope of protection of this invention.

[0027] The application of any of the above-described construction methods or kits in analyzing the mixing ratio of mixed species or mixed samples is also within the scope of protection of this invention.

[0028] The species mentioned above may be a single species or two or more species.

[0029] The samples mentioned above can be single samples or mixed samples.

[0030] Any of the species mentioned above may be a plant. The plant may be a medicinal plant.

[0031] When the DNA to be tested is plant DNA, the species is plant, or the sample originates from a plant, the internal transcribed spacer region 2 (ITS2) is a commonly used DNA barcode sequence for plant identification, located between the 5.8S and 28S rRNA sequences. It exhibits extremely extensive sequence polymorphism in most eukaryotes, reflecting recent evolutionary characteristics and species-level differences. By designing a primer pair targeting the conserved regions of 5.8S and 28S rRNA flanking ITS2, the ITS2 sequences of different plant species can be amplified, allowing for species identification based on the ITS2 sequence. The sequences flanking ITS2 are conserved in all plants; the first specific sequence is the conserved region of 5.8S rRNA, and the second specific sequence is the conserved region of 28S rRNA. A single primer pair can amplify the ITS2 region of all plants.

[0032] The medicinal plants mentioned above can be ginseng, American ginseng, and / or Panax notoginseng. In this case, the nucleotide sequences of the forward primer, reverse primer, primer 1, primer 2, and primer 3 are shown sequentially as SEQ ID NO:1—SEQ ID NO:5.

[0033] Experiments have shown that preparing samples by mixing ginseng, American ginseng, and / or Panax notoginseng in different proportions, and then preparing DNA barcode next-generation sequencing libraries according to the method provided in this invention, is effective. The library is then sequenced, and the first 10 bp UIDs of reads1 are extracted to group reads1 and reads2. The sequences are then assembled, and the assembled sequences are compared with a database. The species identification accuracy is 100%, and the similarity between the assembled sequences and the target species is 100%. The proportion of the mixed samples is calculated based on the number of UIDs of the same species, and the proportion is generally accurate. Therefore, the method provided in this invention can accurately identify species and determine the mixing ratio, and is suitable for the analysis of mixed samples.

[0034] This invention provides a method for constructing a DNA barcode next-generation sequencing library. This method involves two rounds of PCR to prepare the sequencing library from barcode DNA. The method is simple, rapid, and facilitates the assembly of full-length DNA barcodes. It can accurately identify species and determine mixing ratios, making it particularly suitable for the analysis of mixed samples. This sequencing library can be directly used for next-generation sequencing, especially for the identification of mixed samples. The operation is simple, rapid, and cost-effective. This invention has significant application value. Attached Figure Description

[0035] Figure 1 This is a flowchart illustrating the method for constructing a DNA barcode next-generation sequencing library.

[0036] Figure 2The structures of the forward primer, reverse primer, primer 1, primer 2, and primer 3 for two-round specific amplification are shown.

[0037] Figure 3 The results of product fragment distribution detected by Agilent 2100 in Example 2 are shown. Detailed Implementation

[0038] The present invention will now be described in further detail with reference to specific embodiments. The given embodiments are merely illustrative of the invention and not intended to limit its scope. The embodiments provided below can serve as a guide for further improvements by those skilled in the art and do not constitute a limitation on the invention in any way.

[0039] Unless otherwise specified, the experimental methods used in the following examples are conventional methods, performed according to the techniques or conditions described in the literature in this field or according to the product instructions. Unless otherwise specified, the materials and reagents used in the following examples are commercially available.

[0040] Example 1: Establishment of a DNA barcode next-generation sequencing library

[0041] This invention, through extensive experimentation, establishes a method for constructing DNA barcode next-generation sequencing libraries. The specific steps are as follows:

[0042] 1. Synthesize primer pairs for the first round of specific amplification.

[0043] The primer pairs used in the first round of amplification consist of a forward primer and a reverse primer.

[0044] The forward primer consists of a first universal sequence, a random sequence tag, and a first specific sequence, from the 5' end to the 3' end. The reverse primer contains a second specific sequence. Both the forward and reverse primers specifically amplify the DNA barcode region in the template using the first and second specific sequences. The random sequence tag is a 10-20 bp N base, where N can be any one of A, T, G, or C.

[0045] 2. Perform the first round of PCR amplification.

[0046] The template was subjected to the first round of PCR amplification to obtain the first round of amplicon. During the first round of PCR amplification, the forward and reverse primers were used for two cycles of amplification to obtain the DNA barcode region sequence, and a unique tag was added to the 5' end of each amplicon using a random sequence tag.

[0047] The first round of PCR amplification is used for specific amplification with random sequence tags.

[0048] 3. Synthesize the primer combination for the second round of amplification.

[0049] The primer combination for the second round of amplification consists of primer 1, primer 2, and primer 3.

[0050] Primer 1 includes a first universal sequence. Primer 2 consists of a second specific sequence, a second universal sequence, and a random sequence from the 5' end to the 3' end. Primer 3 consists of a library tag sequencing primer binding sequence, a library tag sequence, and a second universal sequence from the 5' end to the 3' end.

[0051] Both the first and second universal sequences are universal adapter sequences for complete sequencing (such as MGI / Illumina and Proton platform adapters) and are not from the unknown sequence (i.e., the template). The random sequence is a 6-15 bp long N base, where N can be any of A, T, G, and C.

[0052] Primer 2 is anchored to the reverse strand of the second specific sequence of the first-round amplicon via a second specific sequence. Based on the foldability of the DNA strand in three-dimensional space, the random sequence pairs complementaryly with bases at any position in the DNA barcode region and extends, simultaneously amplifying with primer 1 to obtain products of unequal fragment sizes. Each amplified product has a first universal sequence and a tag sequence identical to the template at the 5' end, and a second universal sequence at the 3' end. Primer 3 adds a library tag and a library tag sequencing primer binding sequence to the above products via the second universal sequence to obtain a sequencing library.

[0053] 4. Perform a second round of PCR amplification.

[0054] The first-round amplicons were subjected to a second round of PCR amplification to obtain the second-round amplicons.

[0055] During the second round of PCR amplification, primer 2 is anchored to the reverse strand of the second specific sequence of the first round amplicons via a second specific sequence. Based on the foldability of the DNA strand in three-dimensional space, the random sequence is complementary to the bases at any position in the DNA barcode region and extended. Primer 1 is anchored to the reverse strand of the first universal sequence of the first round amplicons via a first universal sequence and extended. Amplification with primers 1 and 2 yields products of unequal fragment sizes. Each amplified product has the first universal sequence and a unique tag originating from a different template at the 5' end, and the second universal sequence at the 3' end. Primer 3 is anchored to the reverse strand of the second universal sequence of the above products via the second universal sequence and extended. A library tag sequence and a library tag sequencing primer binding sequence are added to the above products.

[0056] The second round of PCR amplification was used for library amplification.

[0057] 5. The second-round amplicon is purified to remove fragments that are too short, resulting in a DNA barcode next-generation sequencing library.

[0058] This second-generation sequencing library is adapted for paired-end sequencing.

[0059] The flowchart for constructing the above DNA barcode next-generation sequencing library is shown below. Figure 1 .

[0060] The structures of the forward and reverse primers for the first round of specific amplification are shown in [reference needed]. Figure 2 Left middle image.

[0061] The structures of primers 1, 2, and 3 in the primer combination for the second round of amplification are shown below. Figure 2 The image is shown in the middle right corner.

[0062] Example 2: Identification of mixed medicinal plant samples using the method established in Example 1.

[0063] I. Primer Design and Synthesis

[0064] The internal transcribed spacer region 2 (ITS 2) is a commonly used DNA barcode sequence for plant identification, located between the 5.8S and 28S rRNA. Because ITS 2 does not incorporate into mature ribosomes, it experiences very little natural selection pressure during evolution, thus tolerating greater variation. It exhibits extensive sequence polymorphism in most eukaryotes, reflecting recent evolutionary characteristics and species-level differences. A single primer pair can be designed to amplify the ITS 2 sequences of different plant species, targeting the conserved regions of 5.8S and 28S rRNA flanking ITS 2, allowing for species identification based on the ITS 2 sequence. The flanking sequences of ITS 2 are conserved in all plants; the first specific sequence is the conserved region of 5.8S rRNA, and the second specific sequence is the conserved region of 28S rRNA. A single primer pair can amplify the ITS 2 region of all plants. Primer pairs for the first round of specific amplification of the ITS 2 fragment and primer combinations for the second round of amplification were designed and synthesized according to the method in Example 1. The specific nucleotide sequences of the primers are shown in Table 1.

[0065] Table 1

[0066]

[0067]

[0068] Note: A single underscore represents the first general sequence, a double underscore represents the first specific sequence, a wavy line represents the second specific sequence, a dashed line represents the second general sequence, and N represents any one of A, T, G, and C, which can form random sequence tags, random sequences, and library tag sequences.

[0069] II. Identification of mixed samples of different medicinal plants using the method established in Example 1

[0070] 1. Preparation of genomic DNA from mixed samples

[0071] (1) Take 40mg of commercially available ginseng, American ginseng and Panax notoginseng slices respectively, and pulverize them using a TissueLyser II instrument (Qiagen, Cat No. / ID:85300) to obtain ginseng powder, American ginseng powder and Panax notoginseng powder.

[0072] (2) Ginseng powder and Panax notoginseng powder were mixed at mass ratios of 3:1, 1:1 and 1:3 respectively to obtain mixed sample 1, mixed sample 2 and mixed sample 3. American ginseng powder and Panax notoginseng powder were mixed at mass ratios of 3:1, 1:1 and 1:3 respectively to obtain mixed sample 4, mixed sample 5 and mixed sample 6 respectively.

[0073] (3) Use the Meiji Plant Sample DNA Rapid Extraction Reagent (Cat No. / ID:MD5118) to extract genomic DNA from the mixed samples (mixed sample 1, mixed sample 2, mixed sample 3, mixed sample 4, mixed sample 5 or mixed sample 6) to obtain mixed sample genomic DNA.

[0074] 2. Establishment of a pooled sample DNA barcoding next-generation sequencing library

[0075] (1) Prepare the first round PCR reaction system. The first round PCR reaction system is 50 μl, consisting of 5 μl of mixed sample genomic DNA (mixed sample 1 genomic DNA, mixed sample 2 genomic DNA, mixed sample 3 genomic DNA, mixed sample 4 genomic DNA, mixed sample 5 genomic DNA or mixed sample 6 genomic DNA), 25 μl of KAPA 2G Multiplex PCR Kit (Cat No. / ID is KK5802), 2.5 μl of forward primer aqueous solution (concentration of 10 μM) as shown in Table 1, 2.5 μl of reverse primer aqueous solution (concentration of 10 μM) as shown in Table 1, and 15 μl of molecular-grade water.

[0076] (2) After completing step (1), take the first round of PCR reaction system and perform the first round of PCR amplification to obtain PCR amplification product 1.

[0077] The reaction program was as follows: 95℃ for 3 min, 1 cycle; 98℃ for 20 s, 60℃ for 1 min, 72℃ for 30 s, 2 cycles; 72℃ for 5 min, 1 cycle.

[0078] (3) After completing step (2), add 60 μl of XP beads (AgencourtAMPureXP magnetic beads from Beckman Coulter, catalog number A63881) to the PCR amplification product 1 for purification, and then dissolve it with 19 μl of TE buffer to obtain the reaction product.

[0079] (4) Prepare the second round PCR reaction system. The second round PCR reaction system is 50 μl, consisting of 19 μl of reactants, 25 μl of KAPA 2G Multiplex PCR Kit, 2.5 μl of primer 1 aqueous solution (concentration of 10 μM) as shown in Table 1, 1 μl of primer 2 aqueous solution (concentration of 10 μM) as shown in Table 1, and 2.5 μl of primer 3 aqueous solution (concentration of 10 μM) as shown in Table 1.

[0080] (5) After completing step (4), take the second round of PCR reaction system and perform the second round of PCR amplification to obtain PCR amplification product 2.

[0081] The reaction program was as follows: 95℃ for 3 min, 1 cycle; 98℃ for 20 s, 60℃ for 15 s, 72℃ for 30 s, 30 cycles; 72℃ for 5 min, 1 cycle.

[0082] (6) After completing step (5), add 40 μl of XP beads to the PCR amplification product 2 for purification, and then dissolve it with 20 μl of TE buffer to obtain the product. This product is the mixed sample DNA barcode next-generation sequencing library.

[0083] Each pooled sample was replicated three times. This means that each pooled sample yielded three DNA barcode next-generation sequencing libraries.

[0084] The size of the products was measured using an Agilent 2100 microscope. Some results are shown below. Figure 3 The results showed that the amplified product was a diffuse band with a main peak of approximately 350 bp, as expected.

[0085] The concentration of the products was quality checked using qubit 2.0. The results showed that the concentration of the second-round PCR amplification products was greater than 10 ng / μL, which was within the acceptable range and met expectations.

[0086] 3. Sequencing

[0087] All products obtained in step 2 were standardized and mixed in equal amounts to obtain a mixed library. The mixed library was then subjected to parallel sequencing to obtain sequencing results; the sequencing platform was MGISEQ-2000, and the sequencing type was PE100. The read sequencing results were analyzed, including filtering adapter primer sequences, grouping paired-end reads by tags, assembly, alignment, species identification, and detection according to their mixing ratios.

[0088] The analysis results are shown in Tables 2 and 3 (mixed samples 1-1, 1-2, and 1-3 are triplicate replicates of mixed sample 1; mixed samples 2-1, 2-2, and 2-3 are triplicate replicates of mixed sample 2, and so on). The results indicate that the method established in Example 1 can prepare DNA barcode next-generation sequencing libraries of varying lengths. These libraries can be directly sequenced. The first 10 bp UIDs of reads1 are extracted to group reads1 and reads2, followed by sequence assembly. The assembled sequences are compared with a database, achieving 100% species identification accuracy and 100% similarity to the target species. The proportion of mixed samples is calculated based on the number of UIDs of the same species, and the proportion is correct. Therefore, the DNA barcode next-generation sequencing library prepared using the method established in Example 1 can accurately identify species and determine the mixing ratio, making it suitable for the analysis of mixed samples.

[0089] Table 2

[0090]

[0091]

[0092] Table 3. Assembly Sequence

[0093]

[0094]

[0095]

[0096]

[0097]

[0098]

[0099]

[0100]

[0101]

[0102] The present invention has been described in detail above. For those skilled in the art, the invention can be practiced in a wide range of ways with equivalent parameters, concentrations, and conditions without departing from its spirit and scope, and without requiring unnecessary experiments. Although specific embodiments have been given, it should be understood that further modifications can be made to the invention. In summary, according to the principles of the invention, this application is intended to include any changes, uses, or improvements to the invention, including changes made using conventional techniques known in the art that depart from the scope disclosed herein. Some of the essential features can be applied within the scope of the following appended claims. <110> Shenzhen BGI Genomics Co., Ltd. <120> A method for constructing DNA barcode next-generation sequencing libraries and its application <160> 41 <170> PatentIn version 3.5 <210> 1 <211> 50 <212> DNA <213> Artificial sequence <400> 1 acatggctac gatccgactt nnnnnnnnnn atgcgatact tggtgtgaat 50 <210> 2 <211> twenty one <212> DNA <213> Artificial sequence <400> 2 gacgcttctc cagactacaa t 21 <210> 3 <211> 29 <212> DNA <213> Artificial sequence <400> 3 cacagaacga catggctacg atccgactt 29 <210> 4 <211> 47 <212> DNA <213> Artificial sequence <400> 4 gacgcttctc cagactacaa tgaccgcttg gcctccgact tnnnnnn 47 <210> 5 <211> 54 <212> DNA <213> Artificial sequence <400> 5 agccaaggag ttnnnnnnnn nnttgtcttc ctaagaccgc ttggcctccg actt 54 <210> 6 <211> 497 <212> DNA <213> Artificial sequence <400> 6 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacgcatcg cgtcgccccc 120 caacccatca ctcccttgcg ggagttgagg cggaggggcg gataatggcc tcccgtgtct 180 caccgcgcgg ttggcccaaa tgcgagtcct tggcgatgga cgtcacgaca agtggtggtt 240 gtaaaaagcc ctcttctcat gtcgtgcggt gacccgtcgc cagcaaaagc tctcatgacc 300 ctgttgcgcc gtcctcgacg tgcgctccga ccgcgacccc aggtcaggcg ggactacccg 360 ctgagtttaa gcatatcaat aagcggagga aaagaaactt acaaggattc ccctagtaac 420 ggcgagcgaa ccgggaatag cccagcttga aaatcgggcg accccgtcgt ccgaattgta 480 gtctgaaaaa agcgtca 497 <210> 7 <211> 497 <212> DNA <213> Artificial sequence <400> 7 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacgcatcg cgtcgccccc 120 caacccatca ttccctcgcg ggagtcgatg cggaggggcg gataatggcc tccctgtct 180 caccgcgcgg ttggcccaaa tgcgagtcct tggcgatgga cgtcacgaca agtggtggtt 240 gttaaaagcc ctcttctcat gtcgtgcggt gacccgtcgc cagcaaaagc tctcatgacc 300 ctgttgcgct gtcctcgacg cgcgctccga ccgcgacccc aggtcaggcg ggactacccg 360 ctgagtttaa gcatatcaat aagcggagga aaagaaactt acaaggattc ccctagtaac 420 ggcgagcgaa ccgggaatag cccagcttga aaattgggcg acctcgtcgt ccgaattgta 480 gtctagaaaa gcgtcaa 497 <210> 8 <211> 497 <212> DNA <213> Artificial sequence <400> 8 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacgcatcg cgtcgccccc 120 caacccatca ctcccttgcg ggagttgagg cggaggggcg gataatggcc tcccgtgtct 180 caccgcgcgg ttggcccaaa tgcgagtcct tggcgatgga cgtcacgaca agtggtggtt 240 gtaaaaagcc ctcttctcat gtcgtgcggt gacccgtcgc cagcaaaagc tctcatgacc 300 ctgttgcgcc gtcctcgacg tgcgctccga ccgcgacccc aggtcaggcg ggactacccg 360 ctgagtttaa gcatatcaat aagcggagga aaagaaactt acaaggattc ccctagtaac 420 ggcgagcgaa ccgggaatag cccagcttga aaatcgggcg accccgtcgt ccgaattgta 480 gtctgaaaaa agcgtca 497 <210> 9 <211> 497 <212> DNA <213> Artificial sequence <400> 9 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacccatcg cgtcgccccc 120 caacccatca ttccctcgcg ggagtcgatg cggaggggcg gataatggcc tcccgtgtct 180 caccgcgcgg ttggccaaa tgcgagtcct tggcgatgga cgtcacgaca agtggtggtt 240 gttaaaagcc ctcttctcat gtcgtgcggt gaccgtcgc cagcaaagc tctcatgacc 300 ctgttgcgct gtcctcgacg cgcgctccga ccgcgacccc aggtcaggcg ggactacccg 360 ctgagtttaa gcatatcaat aagcggagga aaagaaactt acaaggattc ccctagtaac 420 ggcgagcgaa ccgggaatag cccagcttga aaattgggcg acctcgtcgt ccgaattgta 480 gtctagaaaa gcgtcaa 497 <210> 10 <211> 497 <212> DNA <213> Artificial Sequence <400> 10 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacccatcg cgtcgccccc 120 caacccatca ctcccttgcg ggagttgagg cggaggggcg gataatggcc tcccgtgtct 180 caccgcgcgg ttggcccaaa tgcgagtcct tggcgatgga cgtcacgaca agtggtggtt 240 gtaaaaagcc ctcttctcat gtcgtgcggt gacccgtcgc cagcaaaagc tctcatgacc 300 ctgttgcgcc gtcctcgacg tgcgctccga ccgcgacccc aggtcaggcg ggactacccg 360 ctgagtttaa gcatatcaat aagcggagga aaagaaactt acaaggattc ccctagtaac 420 ggcgagcgaa ccgggaatag cccagcttga aaatcgggcg accccgtcgt ccgaattgta 480 gtctgaaaaa agcgtca 497 <210> 11 <211> 497 <212> DNA <213> Artificial sequence <400> 11 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacgcatcg cgtcgccccc 120 caacccatca ttccctcgcg ggagtcgatg cggaggggcg gataatggcc tcccgtgtct 180 caccgcgcgg ttggcccaaa tgcgagtcct tggcgatgga cgtcacgaca agtggtggtt 240 gttaaaagcc ctcttctcat gtcgtgcggt gacccgtcgc cagcaaaagc tctcatgacc 300 ctgttgcgct gtcctcgacg cgcgctccga ccgcgacccc aggtcaggcg ggactacccg 360 ctgagtttaa gcatatcaat aagcggagga aaagaaactt acaaggattc ccctagtaac 420 ggcgagcgaa ccgggaatag cccagcttga aaattgggcg acctcgtcgt ccgaattgta 480 gtctagaaaa gcgtcaa 497 <210> 12 <211> 497 <212> DNA <213> Artificial sequence <400> 12 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacgcatcg cgtcgccccc 120 caacccatca ctcccttgcg ggagttgagg cggaggggcg gataatggcc tcccgtgtct 180 caccgcgcgg ttggcccaaa tgcgagtcct tggcgatgga cgtcacgaca agtggtggtt 240 gtaaaaagcc ctcttctcat gtcgtgcggt gacccgtcgc cagcaaaagc tctcatgacc 300 ctgttgcgcc gtcctcgacg tgcgctccga ccgcgacccc aggtcaggcg ggactacccg 360 ctgagtttaa gcatatcaat aagcggagga aaagaaactt acaaggattc ccctagtaac 420 ggcgagcgaa ccgggaatag cccagcttga aaatcgggcg accccgtcgt ccgaattgta 480 gtctgaaaaa agcgtca 497 <210> 13 <211> 497 <212> DNA <213> Artificial sequence <400> 13 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacgcatcg cgtcgccccc 120 caacccatca ttccctcgcg ggagtcgatg cggaggggcg gataatggcc tccctgtct 180 caccgcgcgg ttggcccaaa tgcgagtcct tggcgatgga cgtcacgaca agtggtggtt 240 gttaaaagcc ctcttctcat gtcgtgcggt gacccgtcgc cagcaaaagc tctcatgacc 300 ctgttgcgct gtcctcgacg cgcgctccga ccgcgacccc aggtcaggcg ggactacccg 360 ctgagtttaa gcatatcaat aagcggagga aaagaaactt acaaggattc ccctagtaac 420 ggcgagcgaa ccgggaatag cccagcttga aaattgggcg acctcgtcgt ccgaattgta 480 gtctagaaaa gcgtcaa 497 <210> 14 <211> 497 <212> DNA <213> Artificial sequence <400> 14 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacgcatcg cgtcgccccc 120 caacccatca ctcccttgcg ggagttgagg cggaggggcg gataatggcc tcccgtgtct 180 caccgcgcgg ttggcccaaa tgcgagtcct tggcgatgga cgtcacgaca agtggtggtt 240 gtaaaaagcc ctcttctcat gtcgtgcggt gacccgtcgc cagcaaaagc tctcatgacc 300 ctgttgcgcc gtcctcgacg tgcgctccga ccgcgacccc aggtcaggcg ggactacccg 360 ctgagtttaa gcatatcaat aagcggagga aaagaaactt acaaggattc ccctagtaac 420 ggcgagcgaa ccgggaatag cccagcttga aaatcgggcg accccgtcgt ccgaattgta 480 gtctgaaaaa agcgtca 497 <210> 15 <211> 497 <212> DNA <213> Artificial sequence <400> 15 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacccatcg cgtcgccccc 120 caacccatca ttccctcgcg ggagtcgatg cggaggggcg gataatggcc tcccgtgtct 180 caccgcgcgg ttggccaaa tgcgagtcct tggcgatgga cgtcacgaca agtggtggtt 240 gttaaaagcc ctcttctcat gtcgtgcggt gaccgtcgc cagcaaagc tctcatgacc 300 ctgttgcgct gtcctcgacg cgcgctccga ccgcgacccc aggtcaggcg ggactacccg 360 ctgagtttaa gcatatcaat aagcggagga aaagaaactt acaaggattc ccctagtaac 420 ggcgagcgaa ccgggaatag cccagcttga aaattgggcg acctcgtcgt ccgaattgta 480 gtctagaaaa gcgtcaa 497 <210> 16 <211> 497 <212> DNA <213> Artificial Sequence <400> 16 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacccatcg cgtcgccccc 120 caacccatca ctcccttgcg ggagttgagg cggaggggcg gataatggcc tcccgtgtct 180 caccgcgcgg ttggcccaaa tgcgagtcct tggcgatgga cgtcacgaca agtggtggtt 240 gtaaaaagcc ctcttctcat gtcgtgcggt gacccgtcgc cagcaaaagc tctcatgacc 300 ctgttgcgcc gtcctcgacg tgcgctccga ccgcgacccc aggtcaggcg ggactacccg 360 ctgagtttaa gcatatcaat aagcggagga aaagaaactt acaaggattc ccctagtaac 420 ggcgagcgaa ccgggaatag cccagcttga aaatcgggcg accccgtcgt ccgaattgta 480 gtctgaaaaa agcgtca 497 <210> 17 <211> 497 <212> DNA <213> Artificial sequence <400> 17 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacgcatcg cgtcgccccc 120 caacccatca ttccctcgcg ggagtcgatg cggaggggcg gataatggcc tcccgtgtct 180 caccgcgcgg ttggcccaaa tgcgagtcct tggcgatgga cgtcacgaca agtggtggtt 240 gttaaaagcc ctcttctcat gtcgtgcggt gacccgtcgc cagcaaaagc tctcatgacc 300 ctgttgcgct gtcctcgacg cgcgctccga ccgcgacccc aggtcaggcg ggactacccg 360 ctgagtttaa gcatatcaat aagcggagga aaagaaactt acaaggattc ccctagtaac 420 ggcgagcgaa ccgggaatag cccagcttga aaattgggcg acctcgtcgt ccgaattgta 480 gtctagaaaa gcgtcaa 497 <210> 18 <211> 497 <212> DNA <213> Artificial sequence <400> 18 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacgcatcg cgtcgccccc 120 caacccatca ctcccttgcg ggagttgagg cggaggggcg gataatggcc tcccgtgtct 180 caccgcgcgg ttggcccaaa tgcgagtcct tggcgatgga cgtcacgaca agtggtggtt 240 gtaaaaagcc ctcttctcat gtcgtgcggt gacccgtcgc cagcaaaagc tctcatgacc 300 ctgttgcgcc gtcctcgacg tgcgctccga ccgcgacccc aggtcaggcg ggactacccg 360 ctgagtttaa gcatatcaat aagcggagga aaagaaactt acaaggattc ccctagtaac 420 ggcgagcgaa ccgggaatag cccagcttga aaatcgggcg accccgtcgt ccgaattgta 480 gtctgaaaaa agcgtca 497 <210> 19 <211> 497 <212> DNA <213> Artificial sequence <400> 19 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacgcatcg cgtcgccccc 120 caacccatca ttccctcgcg ggagtcgatg cggaggggcg gataatggcc tccctgtct 180 caccgcgcgg ttggcccaaa tgcgagtcct tggcgatgga cgtcacgaca agtggtggtt 240 gttaaaagcc ctcttctcat gtcgtgcggt gacccgtcgc cagcaaaagc tctcatgacc 300 ctgttgcgct gtcctcgacg cgcgctccga ccgcgacccc aggtcaggcg ggactacccg 360 ctgagtttaa gcatatcaat aagcggagga aaagaaactt acaaggattc ccctagtaac 420 ggcgagcgaa ccgggaatag cccagcttga aaattgggcg acctcgtcgt ccgaattgta 480 gtctagaaaa gcgtcaa 497 <210> 20 <211> 497 <212> DNA <213> Artificial sequence <400> 20 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacgcatcg cgtcgccccc 120 caacccatca ctcccttgcg ggagttgagg cggaggggcg gataatggcc tcccgtgtct 180 caccgcgcgg ttggcccaaa tgcgagtcct tggcgatgga cgtcacgaca agtggtggtt 240 gtaaaaagcc ctcttctcat gtcgtgcggt gacccgtcgc cagcaaaagc tctcatgacc 300 ctgttgcgcc gtcctcgacg tgcgctccga ccgcgacccc aggtcaggcg ggactacccg 360 ctgagtttaa gcatatcaat aagcggagga aaagaaactt acaaggattc ccctagtaac 420 ggcgagcgaa ccgggaatag cccagcttga aaatcgggcg accccgtcgt ccgaattgta 480 gtctgaaaaa agcgtca 497 <210> 21 <211> 497 <212> DNA <213> Artificial sequence <400> 21 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacccatcg cgtcgccccc 120 caacccatca ttccctcgcg ggagtcgatg cggaggggcg gataatggcc tcccgtgtct 180 caccgcgcgg ttggccaaa tgcgagtcct tggcgatgga cgtcacgaca agtggtggtt 240 gttaaaagcc ctcttctcat gtcgtgcggt gaccgtcgc cagcaaagc tctcatgacc 300 ctgttgcgct gtcctcgacg cgcgctccga ccgcgacccc aggtcaggcg ggactacccg 360 ctgagtttaa gcatatcaat aagcggagga aaagaaactt acaaggattc ccctagtaac 420 ggcgagcgaa ccgggaatag cccagcttga aaattgggcg acctcgtcgt ccgaattgta 480 gtctagaaaa gcgtcaa 497 <210> 22 <211> 497 <212> DNA <213> Artificial Sequence <400> 22 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacccatcg cgtcgccccc 120 caacccatca ctcccttgcg ggagttgagg cggaggggcg gataatggcc tcccgtgtct 180 caccgcgcgg ttggcccaaa tgcgagtcct tggcgatgga cgtcacgaca agtggtggtt 240 gtaaaaagcc ctcttctcat gtcgtgcggt gacccgtcgc cagcaaaagc tctcatgacc 300 ctgttgcgcc gtcctcgacg tgcgctccga ccgcgacccc aggtcaggcg ggactacccg 360 ctgagtttaa gcatatcaat aagcggagga aaagaaactt acaaggattc ccctagtaac 420 ggcgagcgaa ccgggaatag cccagcttga aaatcgggcg accccgtcgt ccgaattgta 480 gtctgaaaaa agcgtca 497 <210> 23 <211> 497 <212> DNA <213> Artificial sequence <400> 23 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacgcatcg cgtcgccccc 120 caacccatca ttccctcgcg ggagtcgatg cggaggggcg gataatggcc tcccgtgtct 180 caccgcgcgg ttggcccaaa tgcgagtcct tggcgatgga cgtcacgaca agtggtggtt 240 gttaaaagcc ctcttctcat gtcgtgcggt gacccgtcgc cagcaaaagc tctcatgacc 300 ctgttgcgct gtcctcgacg cgcgctccga ccgcgacccc aggtcaggcg ggactacccg 360 ctgagtttaa gcatatcaat aagcggagga aaagaaactt acaaggattc ccctagtaac 420 ggcgagcgaa ccgggaatag cccagcttga aaattgggcg acctcgtcgt ccgaattgta 480 gtctagaaaa gcgtcaa 497 <210> 24 <211> 497 <212> DNA <213> Artificial Sequence <400> 24 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacccatcg cgtcgccccc 120 caacccatca ctccttgcg ggagtggacg cggaggggag gataatggcc tcccgtgtct 180 caccgcgcgg tggcccaaa tgccctcctt ggcgatggat gtcacgacaa gtggtggttg 240 taaaaagccc tctctcatg tcttgcggtg accctattgc cagcaaagc tctcatgacc 300 ctgttgcgcc gtcctccacg tgcgctccga ccgcgaccc aggtcaggcg ggactacccg 360 ctgagtttaa acatcaat aagcggatga aaaaaaactt acaaggattc ccctaaaaac 420 ggcgagcgaa ccgggaatag cccatcttga aaatcgggcg accccatcgt ccgtcattgt 480 agtctgagaa gcgtcaa 497 <210> 25 <211> 497 <212> DNA <213> Artificial sequence <400> 25 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacgcatcg cgtcgccccc 120 caacccatca ttccctcgcg ggagtcgatg cggaggggcg gataatggcc tcccgtgtct 180 caccgcgcgg ttggcccaaa tgcgagtcct tggcgatgga cgtcacgaca agtggtggtt 240 gttaaaagcc ctcttctcat gtcgtgcggt gacccgtcgc cagcaaaagc tctcatgacc 300 ctgttgcgct gtcctcgacg cgcgctccga ccgcgacccc aggtcaggcg ggactacccg 360 ctgagtttaa gcatatcaat aagcggagga aaagaaactt acaaggattc ccctagtaac 420 ggcgagcgaa ccgggaatag cccagcttga aaattgggcg acctcgtcgt ccgaattgta 480 gtctagaaaa gcgtcaa 497 <210> 26 <211> 497 <212> DNA <213> Artificial Sequence <400> 26 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacccatcg cgtcgccccc 120 caacccatca ctccttgcg ggagtggacg cggaggggag gataatggcc tcccgtgtct 180 caccgcgcgg tggcccaaa tgccctcctt ggcgatggat gtcacgacaa gtggtggttg 240 taaaaagccc tctctcatg tcttgcggtg accctattgc cagcaaagc tctcatgacc 300 ctgttgcgcc gtcctccacg tgcgctccga ccgcgaccc aggtcaggcg ggactacccg 360 ctgagtttaa acatcaat aagcggatga aaaaaaactt acaaggattc ccctaaaaac 420 ggcgagcgaa ccgggaatag cccatcttga aaatcggcg acccatcgt ccgtcattgt 480 agtctgagaa gcgtcaa 497 <210> 27 <211> 497 <212> DNA <213> Artificial Sequence <400> 27 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacccatcg cgtcgccccc 120 caacccatca ttccctcgcg ggagtcgatg cggaggggcg gataatggcc tcccgtgtct 180 caccgcgcgg ttggccaaa tgcgagtcct tggcgatgga cgtcacgaca agtggtggtt 240 gttaaaagcc ctcttctcat gtcgtgcggt gaccgtcgc cagcaaagc tctcatgacc 300 ctgttgcgct gtcctcgacg cgcgctccga ccgcgacccc aggtcaggcg ggactacccg 360 ctgagtttaa gcatatcaat aagcggagga aaagaaactt acaaggattc ccctagtaac 420 ggcgagcgaa ccgggaatag cccagcttga aaattgggcg acctcgtcgt ccgaattgta 480 gtctagaaaa gcgtcaa 497 <210> 28 <211> 497 <212> DNA <213> Artificial Sequence <400> 28 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacccatcg cgtcgccccc 120 caacccatca ctccttgcg ggagtggacg cggaggggag gataatggcc tcccgtgtct 180 caccgcgcgg tggcccaaa tgccctcctt ggcgatggat gtcacgacaa gtggtggttg 240 taaaaagccc tctctcatg tcttgcggtg accctattgc cagcaaagc tctcatgacc 300 ctgttgcgcc gtcctccacg tgcgctccga ccgcgaccc aggtcaggcg ggactacccg 360 ctgagtttaa acatcaat aagcggatga aaaaaaactt acaaggattc ccctaaaaac 420 ggcgagcgaa ccgggaatag cccatcttga aaatcggcg acccatcgt ccgtcattgt 480 agtctgagaa gcgtcaa 497 <210> 29 <211> 497 <212> DNA <213> Artificial Sequence <400> 29 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacccatcg cgtcgccccc 120 caacccatca ttccctcgcg ggagtcgatg cggaggggcg gataatggcc tcccgtgtct 180 caccgcgcgg ttggccaaa tgcgagtcct tggcgatgga cgtcacgaca agtggtggtt 240 gttaaaagcc ctcttctcat gtcgtgcggt gaccgtcgc cagcaaagc tctcatgacc 300 ctgttgcgct gtcctcgacg cgcgctccga ccgcgacccc aggtcaggcg ggactacccg 360 ctgagtttaa gcatatcaat aagcggagga aaagaaactt acaaggattc ccctagtaac 420 ggcgagcgaa ccgggaatag cccagcttga aaattgggcg acctcgtcgt ccgaattgta 480 gtctagaaaa gcgtcaa 497 <210> 30 <211> 497 <212> DNA <213> Artificial Sequence <400> 30 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacccatcg cgtcgccccc 120 caacccatca ctccttgcg ggagtggacg cggaggggag gataatggcc tcccgtgtct 180 caccgcgcgg tggcccaaa tgccctcctt ggcgatggat gtcacgacaa gtggtggttg 240 taaaaagccc tctctcatg tcttgcggtg accctattgc cagcaaagc tctcatgacc 300 ctgttgcgcc gtcctccacg tgcgctccga ccgcgaccc aggtcaggcg ggactacccg 360 ctgagtttaa acatcaat aagcggatga aaaaaaactt acaaggattc ccctaaaaac 420 ggcgagcgaa ccgggaatag cccatcttga aaatcgggcg accccatcgt ccgtcattgt 480 agtctgagaa gcgtcaa 497 <210> 31 <211> 497 <212> DNA <213> Artificial sequence <400> 31 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacgcatcg cgtcgccccc 120 caacccatca ttccctcgcg ggagtcgatg cggaggggcg gataatggcc tcccgtgtct 180 caccgcgcgg ttggcccaaa tgcgagtcct tggcgatgga cgtcacgaca agtggtggtt 240 gttaaaagcc ctcttctcat gtcgtgcggt gacccgtcgc cagcaaaagc tctcatgacc 300 ctgttgcgct gtcctcgacg cgcgctccga ccgcgacccc aggtcaggcg ggactacccg 360 ctgagtttaa gcatatcaat aagcggagga aaagaaactt acaaggattc ccctagtaac 420 ggcgagcgaa ccgggaatag cccagcttga aaattgggcg acctcgtcgt ccgaattgta 480 gtctagaaaa gcgtcaa 497 <210> 32 <211> 497 <212> DNA <213> Artificial Sequence <400> 32 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacccatcg cgtcgccccc 120 caacccatca ctccttgcg ggagtggacg cggaggggag gataatggcc tcccgtgtct 180 caccgcgcgg tggcccaaa tgccctcctt ggcgatggat gtcacgacaa gtggtggttg 240 taaaaagccc tctctcatg tcttgcggtg accctattgc cagcaaagc tctcatgacc 300 ctgttgcgcc gtcctccacg tgcgctccga ccgcgaccc aggtcaggcg ggactacccg 360 ctgagtttaa acatcaat aagcggatga aaaaaaactt acaaggattc ccctaaaaac 420 ggcgagcgaa ccgggaatag cccatcttga aaatcggcg acccatcgt ccgtcattgt 480 agtctgagaa gcgtcaa 497 <210> 33 <211> 497 <212> DNA <213> Artificial Sequence <400> 33 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacccatcg cgtcgccccc 120 caacccatca ttccctcgcg ggagtcgatg cggaggggcg gataatggcc tcccgtgtct 180 caccgcgcgg ttggccaaa tgcgagtcct tggcgatgga cgtcacgaca agtggtggtt 240 gttaaaagcc ctcttctcat gtcgtgcggt gaccgtcgc cagcaaagc tctcatgacc 300 ctgttgcgct gtcctcgacg cgcgctccga ccgcgacccc aggtcaggcg ggactacccg 360 ctgagtttaa gcatatcaat aagcggagga aaagaaactt acaaggattc ccctagtaac 420 ggcgagcgaa ccgggaatag cccagcttga aaattgggcg acctcgtcgt ccgaattgta 480 gtctagaaaa gcgtcaa 497 <210> 34 <211> 497 <212> DNA <213> Artificial Sequence <400> 34 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacccatcg cgtcgccccc 120 caacccatca ctccttgcg ggagtggacg cggaggggag gataatggcc tcccgtgtct 180 caccgcgcgg tggcccaaa tgccctcctt ggcgatggat gtcacgacaa gtggtggttg 240 taaaaagccc tctctcatg tcttgcggtg accctattgc cagcaaagc tctcatgacc 300 ctgttgcgcc gtcctccacg tgcgctccga ccgcgaccc aggtcaggcg ggactacccg 360 ctgagtttaa acatcaat aagcggatga aaaaaaactt acaaggattc ccctaaaaac 420 ggcgagcgaa ccgggaatag cccatcttga aaatcggcg acccatcgt ccgtcattgt 480 agtctgagaa gcgtcaa 497 <210> 35 <211> 497 <212> DNA <213> Artificial Sequence <400> 35 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacccatcg cgtcgccccc 120 caacccatca ttccctcgcg ggagtcgatg cggaggggcg gataatggcc tcccgtgtct 180 caccgcgcgg ttggccaaa tgcgagtcct tggcgatgga cgtcacgaca agtggtggtt 240 gttaaaagcc ctcttctcat gtcgtgcggt gaccgtcgc cagcaaagc tctcatgacc 300 ctgttgcgct gtcctcgacg cgcgctccga ccgcgacccc aggtcaggcg ggactacccg 360 ctgagtttaa gcatatcaat aagcggagga aaagaaactt acaaggattc ccctagtaac 420 ggcgagcgaa ccgggaatag cccagcttga aaattgggcg acctcgtcgt ccgaattgta 480 gtctagaaaa gcgtcaa 497 <210> 36 <211> 497 <212> DNA <213> Artificial Sequence <400> 36 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacccatcg cgtcgccccc 120 caacccatca ctccttgcg ggagtggacg cggaggggag gataatggcc tcccgtgtct 180 caccgcgcgg tggcccaaa tgccctcctt ggcgatggat gtcacgacaa gtggtggttg 240 taaaaagccc tctctcatg tcttgcggtg accctattgc cagcaaagc tctcatgacc 300 ctgttgcgcc gtcctccacg tgcgctccga ccgcgaccc aggtcaggcg ggactacccg 360 ctgagtttaa acatcaat aagcggatga aaaaaaactt acaaggattc ccctaaaaac 420 ggcgagcgaa ccgggaatag cccatcttga aaatcgggcg accccatcgt ccgtcattgt 480 agtctgagaa gcgtcaa 497 <210> 37 <211> 497 <212> DNA <213> Artificial sequence <400> 37 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacgcatcg cgtcgccccc 120 caacccatca ttccctcgcg ggagtcgatg cggaggggcg gataatggcc tcccgtgtct 180 caccgcgcgg ttggcccaaa tgcgagtcct tggcgatgga cgtcacgaca agtggtggtt 240 gttaaaagcc ctcttctcat gtcgtgcggt gacccgtcgc cagcaaaagc tctcatgacc 300 ctgttgcgct gtcctcgacg cgcgctccga ccgcgacccc aggtcaggcg ggactacccg 360 ctgagtttaa gcatatcaat aagcggagga aaagaaactt acaaggattc ccctagtaac 420 ggcgagcgaa ccgggaatag cccagcttga aaattgggcg acctcgtcgt ccgaattgta 480 gtctagaaaa gcgtcaa 497 <210> 38 <211> 497 <212> DNA <213> Artificial Sequence <400> 38 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacccatcg cgtcgccccc 120 caacccatca ctccttgcg ggagtggacg cggaggggag gataatggcc tcccgtgtct 180 caccgcgcgg tggcccaaa tgccctcctt ggcgatggat gtcacgacaa gtggtggttg 240 taaaaagccc tctctcatg tcttgcggtg accctattgc cagcaaagc tctcatgacc 300 ctgttgcgcc gtcctccacg tgcgctccga ccgcgaccc aggtcaggcg ggactacccg 360 ctgagtttaa acatcaat aagcggatga aaaaaaactt acaaggattc ccctaaaaac 420 ggcgagcgaa ccgggaatag cccatcttga aaatcggcg acccatcgt ccgtcattgt 480 agtctgagaa gcgtcaa 497 <210> 39 <211> 497 <212> DNA <213> Artificial Sequence <400> 39 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacccatcg cgtcgccccc 120 caacccatca ttccctcgcg ggagtcgatg cggaggggcg gataatggcc tcccgtgtct 180 caccgcgcgg ttggccaaa tgcgagtcct tggcgatgga cgtcacgaca agtggtggtt 240 gttaaaagcc ctcttctcat gtcgtgcggt gaccgtcgc cagcaaagc tctcatgacc 300 ctgttgcgct gtcctcgacg cgcgctccga ccgcgacccc aggtcaggcg ggactacccg 360 ctgagtttaa gcatatcaat aagcggagga aaagaaactt acaaggattc ccctagtaac 420 ggcgagcgaa ccgggaatag cccagcttga aaattgggcg acctcgtcgt ccgaattgta 480 gtctagaaaa gcgtcaa 497 <210> 40 <211> 497 <212> DNA <213> Artificial Sequence <400> 40 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacccatcg cgtcgccccc 120 caacccatca ctccttgcg ggagtggacg cggaggggag gataatggcc tcccgtgtct 180 caccgcgcgg tggcccaaa tgccctcctt ggcgatggat gtcacgacaa gtggtggttg 240 taaaaagccc tctctcatg tcttgcggtg accctattgc cagcaaagc tctcatgacc 300 ctgttgcgcc gtcctccacg tgcgctccga ccgcgaccc aggtcaggcg ggactacccg 360 ctgagtttaa acatcaat aagcggatga aaaaaaactt acaaggattc ccctaaaaac 420 ggcgagcgaa ccgggaatag cccatcttga aaatcggcg acccatcgt ccgtcattgt 480 agtctgagaa gcgtcaa 497 <210> 41 <211> 497 <212> DNA <213> Artificial Sequence <400> 41 atgcgatact tggtgtgaat tgcagaatcc cgtgaaccat cgagtctttg acgcaagttg 60 cgcccgaagc cattaggccg agggcacgtc tgcctgggcg tcacccatcg cgtcgccccc 120 caacccatca ttccctcgcg ggagtcgatg cggaggggcg gataatggcc tcccgtgtct 180 caccgcgcgg ttggccaaa tgcgagtcct tggcgatgga cgtcacgaca agtggtggtt 240 gttaaaagcc ctcttctcat gtcgtgcggt gaccgtcgc cagcaaagc tctcatgacc 300 ctgttgcgct gtcctcgacg cgcgctccga ccgcgacccc aggtcaggcg ggactacccg 360 ctgagtttaa gcatatcaat aagcggagga aaagaaactt acaaggattc ccctagtaac 420 ggcgagcgaa ccgggaatag cccagcttga aaattgggcg acctcgtcgt ccgaattgta 480 gtctagaaaa gcgtcaa 497

Claims

1. A method for constructing a DNA barcode next-generation sequencing library, comprising the following steps: (1) using a DNA to be detected as a template, performing first round PCR amplification with a primer pair to obtain a first round amplicon; the primer pair is composed of a forward primer and a reverse primer; the forward primer comprises, from 5' end to 3' end, a first universal sequence, a random sequence tag and a first specific sequence; the reverse primer is a second specific sequence; the random sequence tag is composed of several bases N; N is any one of A, T, G and C; (2) after step (1) is completed, performing second round PCR amplification on the first round amplicon with a primer combination to obtain a second round amplicon; the primer combination is composed of primer 1, primer 2 and primer 3; the primer 1 comprises the first universal sequence; the primer 2 comprises, from 5' end to 3' end, the second specific sequence, a second universal sequence and a random sequence; the primer 3 comprises, from 5' end to 3' end, a library tag sequencing primer binding sequence, a library tag sequence and the second universal sequence; (3) after step (2) is completed, taking the second round amplicon, purifying to obtain a DNA barcode next-generation sequencing library.

2. The construction method of claim 1, wherein: the forward primer and the reverse primer specifically amplify the DNA barcode region in the template through the first specific sequence and the second specific sequence.

3. The construction method of claim 1, wherein: the first universal sequence and the second universal sequence are both sequencing universal adapter sequences and are different from the template. 4.The method according to claim 1, wherein: in step (1), the random sequence tag is 10-20 bp in length; in step (2), the random sequence is 6-15 bp in length.

5. The construction method of claim 1, wherein: in step (1), the forward primer and the reverse primer are used to obtain the DNA barcode region sequence through 2-4 cycles of amplification in the first round PCR amplification. 6.A kit for constructing a DNA barcode next-generation sequencing library of a species or a sample, comprising the forward primer, the reverse primer, the primer 1, the primer 2 and the primer 3 in the method according to any one of claims 1-5.

7. The kit of claim 6, wherein: the kit is composed of the forward primer, the reverse primer, the primer 1, the primer 2 and the primer 3. 8.Use of the method according to any one of claims 1-5 or the kit according to claim 6 or 7 in identifying a species or a sample. 9.Use of the method according to any one of claims 1-5 or the kit according to claim 6 or 7 in detecting the purity of a species or a sample. 10.Use of the method according to any one of claims 1-5 or the kit according to claim 6 or 7 in analyzing the mixing ratio of mixed species or mixed samples.

Citation Information

Patent Citations

  • Multi-sample mixed sequencing method and kit

    CN102181533A

  • Primer group for constructing sequencing library with known flanking sequences and method thereof

    CN113088561A