High-throughput molecular identification method for macrofungi based on second-generation sequencing technology

Through second-generation sequencing technology and high-throughput methods, the problems of high sequencing cost and high failure rate in large fungal sample identification are solved, and efficient and low-cost molecular identification of large batch samples are achieved, which simplifies experimental steps and cycles.

CN115094129BActive Publication Date: 2025-09-02MICROBIOLOGY INST OF SHAANXI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210798898.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-06
Publication Date
2025-09-02
Estimated Expiration
2042-07-06

AI Technical Summary

Technical Problem

The prior art has problems such as high sequencing cost, high failure rate and complicated experimental steps in the identification of large fungal samples. Especially due to sequencing failure caused by impurity of ITS gene fragments, it is difficult to achieve efficient molecular identification of large batches of samples.

Method used

Using second-generation sequencing technology, ITS1 and ITS2 sequences were amplified by PCR and added Barcode tag sequences were mixed sequencing. High-throughput sequencing was performed by combining the Illumina HiSeq platform, and fastq-multx and DADA2 were used for sequence splitting and splicing, and species annotation was used for UNITE database to achieve efficient sample molecular identification.

Benefits of technology

High-throughput molecular identification of large fungal samples is achieved, which reduces experimental costs, improves success rates, reduces experimental steps and cycles, and avoids the complicated work caused by sequencing failure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0003733245500000031
    Figure BDA0003733245500000031
  • Figure BDA0003733245500000041
    Figure BDA0003733245500000041
  • Figure BDA0003733245500000051
    Figure BDA0003733245500000051
Patent Text Reader

Abstract

The present invention discloses a high-throughput macrofungal molecular identification method based on second-generation sequencing technology. The method comprises the following steps: using extracted macrofungal genomic DNA as a template, amplifying ITS1 and ITS2 sequences by PCR, and obtaining PCR products; using the PCR products as templates, adding barcode tag sequences to the ITS1 and ITS2 sequences by combining specific primer sequences, and obtaining ITS1 and ITS2 PCR products with barcode tag sequences; mixing all ITS1 and ITS2 PCR products with barcode tag sequences, and then recovering them through gel excision and library construction to obtain a mixed PCR product sample; performing double-end sequencing on the mixed PCR product sample using an Illumina HiSeq PE250 sequencing platform to obtain a high-throughput sequencing result; and performing adapter sequence prediction and removal, sequence splitting, barcode tag sequence deletion, sequence splicing, and species annotation on the high-throughput sequencing result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of second-generation high-throughput sequencing and data analysis, and specifically relates to a high-throughput large fungal molecular identification method based on second-generation sequencing technology. Background Art

[0002] In the investigation, identification and phylogenetic study of large fungal field resources, the traditional method is generally to extract the genomic DNA of a single large fungal fruiting body sample and perform PCR amplification on its ribosomal gene transcribed spacer (ITS) sequence. The obtained PCR product is then sequenced using Sanger sequencing technology, and finally the bidirectional sequences are manually spliced ​​one by one to obtain the ITS sequence information. However, for most large fungal fruiting body samples, since they have multiple copies or have parasitic, associated or symbiotic relationships with other fungi, the obtained ITS gene fragments are impure, which in turn leads to the occurrence of double peaks in the sequencing process and causes sequencing failure.

[0003] For such samples, the ITS gene fragment must be linked to a T-vector via TA cloning. After transformation, single-clone colony selection, shake flask culture, and plasmid extraction, the ITS sequence can be sequenced using Sanger sequencing. Sequencing large quantities of fungal samples is expensive, and samples that fail sequencing require further TA cloning to complete the sequencing, which is complex, labor-intensive, and time-consuming.

[0004] Next-generation sequencing, also known as high-throughput sequencing, offers advantages such as low sequencing cost, deep sequencing depth, and large data volumes, and holds enormous potential and application value in sample testing. Amplicon (ITS or 16S) sequencing is a key technology in high-throughput sequencing applications. It is a highly targeted method for sequencing specific PCR products or captured fragments, and can be used to analyze target regions within specific environments.

[0005] Currently, amplicon sequencing is widely used in the classification of species such as bacteria, fungi, and archaea, as well as in the research of soil, marine, and intestinal microorganisms. It has the advantages of low template requirements, simple operation, strong specificity, high sensitivity, and short sequencing time. Compared with high-throughput sequencing platforms such as 454 and IonTorrent, the Illumina HiSeq sequencing platform is a type of short-read sequencing platform with higher sequencing throughput and lower cost. Its sequencing length can meet the requirements for the classification and identification of large fungi.

[0006] Compared with traditional Sanger sequencing, the second-generation high-throughput sequencing technology based on amplicon enrichment has the advantages of low sequencing cost, high result accuracy and low experimental intensity. Therefore, the application of this technology is becoming more and more extensive. However, there is still a lack of methods or tools for batch identification of large fungal samples using second-generation high-throughput sequencing technology. Summary of the Invention

[0007] In view of this, the main purpose of the present invention is to provide a high-throughput large fungal molecular identification method based on second-generation sequencing technology.

[0008] To achieve the above object, the technical solution of the present invention is achieved as follows:

[0009] The present invention provides a high-throughput molecular identification method for macrofungi based on second-generation sequencing technology, which is as follows:

[0010] Therefore, the extracted macrofungal genomic DNA was used as a template to amplify ITS1 and ITS2 sequences by PCR to obtain PCR products;

[0011] Using the PCR product as a template, adding barcode tag sequences to the ITS1 and ITS2 sequences through a combination of specific primer sequences to obtain ITS1 and ITS2 PCR products with barcode tag sequences;

[0012] All ITS1 and ITS2 PCR products with barcode tag sequences were mixed separately, and then recovered by gel excision and library construction to obtain mixed PCR product samples;

[0013] The mixed PCR product sample was subjected to double-end sequencing using the Illumina HiSeq PE250 sequencing platform to obtain high-throughput sequencing results;

[0014] The high-throughput sequencing results are subjected to adapter sequence prediction and removal, sequence splitting, Barcode tag sequence deletion, sequence splicing and species annotation.

[0015] In the above scheme, the PCR amplification of ITS1 and ITS2 sequences was performed with the following reaction procedure: 95°C for 3 min; 95°C for 20 s, 47°C for 20 s, 72°C for 30 s, 35 cycles; 72°C for 5 min; and storage at 4°C.

[0016] In the above scheme, the barcode tag sequence is added to the ITS1 and ITS2 sequences by a specific primer sequence combination. The PCR reaction procedure is: 95°C for 3 minutes; 95°C for 20 seconds, 47°C for 20 seconds, 72°C for 30 seconds, 10 cycles; 95°C for 20 seconds, 59°C for 20 seconds, 72°C for 30 seconds, 25 cycles; 72°C for 5 minutes; and storage at 4°C.

[0017] In the above scheme, the Barcode tag sequence is shown in the table:

[0018]

[0019]

[0020]

[0021] In the above scheme, the universal primer sequence is ACCHGCGGARGGATCATTAC, TTYDCNRCGTTCTTCATCGT, GTGAVTCATCRARTNTTTGA or CCTBCSCTTANTDATATGCC.

[0022] In the above scheme, the sequence of the ITS1 product with the Barcode tag sequence is shown in the table:

[0023]

[0024] In the above scheme, the sequence of the ITS2 product with the Barcode tag sequence is shown in the table:

[0025]

[0026] In the above scheme, the high-throughput sequencing results are subjected to the following steps: prediction and removal of adapter sequences, sequence splitting, deletion of barcode tag sequences, sequence splicing, and species annotation:

[0027] Prediction and removal of adapter sequences: Use fastp to predict and remove adapter sequences, and set the sliding window length to 16;

[0028] Sequence splitting: Use fastq-multx to split the sequencing results of the mixed sample according to the barcode tag sequence unique to each large fungal sample. Place the split sequence files into different folders according to the sample ID and build manifests for each;

[0029] Deletion and sequence splicing of barcode tag sequences: Use DADA2 to delete and splice the barcode tag sequences of each macrofungal sample sequence;

[0030] Species Annotation: All sequences were annotated with species using the trained annotator.

[0031] In the above scheme, the annotator is obtained after training using the UNITE database with a large amount of large fungal ITS sequence information as a training set and using the Bayesian classification algorithm.

[0032] Compared with the existing technology, the present invention can complete the molecular identification of up to 2500 (50×50) large fungal samples at a time, which greatly reduces the experimental cost compared to performing first-generation Sanger sequencing on each sample separately; the success rate is high: it avoids sequencing failure (double peaks) caused by impure ITS sequence fragments in first-generation sequencing; it avoids the complicated work caused by TA cloning due to sequencing failure (double peaks), reduces the experimental steps, saves experimental materials, and shortens the experimental cycle. DETAILED DESCRIPTION

[0033] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0034] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, article, or device comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, article, or device comprising the element.

[0035] The present invention provides a high-throughput molecular identification method for macrofungi based on second-generation sequencing technology, which is implemented by the following steps:

[0036] Step 1: Using the extracted macrofungal genomic DNA as a template, amplify the ITS1 and ITS2 sequences by PCR to obtain PCR products;

[0037] Specifically, in the amplification of ITS1 and ITS2 sequences by PCR, the PCR reaction procedure is: 95°C for 3 min; 95°C for 20 s, 47°C for 20 s, 72°C for 30 s, 35 cycles; 72°C for 5 min; and storage at 4°C.

[0038] PCR amplification was performed using universal primers for ITS1 (ITS1F: 5′-ACCHGCGGARGGATCATTAC-3′; ITS2R: 5′-TTYDCNRCGTTCTTCATCGT-3′) and ITS2 (gITS7F: 5′-GTGAVTCATCRARTNTTTGA-3′; ITS4R: 5′-CCTBCSCTTANTDATATGCC-3′) as amplification primers, and Rapid Taq Master Mix enzyme purchased from Vazyme (Nanjing Novozyme Biotechnology Co., Ltd.).

[0039] The PCR amplification system was 20 μL and consisted of 10 μL 2× Rapid Taq Master Mix, 1 μL upstream primer, 1 μL downstream primer, 1 μL DNA, and 7 μL ddH 2 O. In this reaction system, the DNA was from different macrofungal samples, and the primer concentration was 10 μM.

[0040] Step 2: Using the PCR product as a template, a barcode tag sequence is added to the ITS1 and ITS2 sequences through a combination of specific primer sequences to obtain ITS1 and ITS2 PCR products with barcode tag sequences;

[0041] Specifically, the PCR reaction program was: 95°C for 3 min; 95°C for 20 s, 47°C for 20 s, 72°C for 30 s, 10 cycles; 95°C for 20 s, 59°C for 20 s, 72°C for 30 s, 25 cycles; 72°C for 5 min; and storage at 4°C.

[0042] The PCR amplification system was 30 μL, consisting of 15 μL 2× Rapid Taq Master Mix, 1.5 μL upstream primer, 1.5 μL downstream primer, 2 μL first-round PCR product, and 10 μL ddH2O.

[0043] The Barcode tag sequence is shown in the table:

[0044]

[0045]

[0046]

[0047]

[0048] The universal primer sequence is ACCHGCGGARGGATCATTAC, TTYDCNRCGTTCTTCATCGT, GTGAVTCATCRARTNTTTGA or CCTBCSCTTANTDATATGCC.

[0049] The sequence of the ITS1 product with the Barcode tag sequence is shown in the table:

[0050]

[0051] The sequence of the ITS2 product with the Barcode tag sequence is shown in the table:

[0052] Step 3: All ITS1 and ITS2 PCR products with barcode tag sequences were mixed separately, and then after gel excision, recovery, and library construction, a mixed PCR product sample was obtained;

[0053] Step 4: Perform double-end sequencing on the mixed PCR product sample using the Illumina HiSeq PE250 sequencing platform to obtain high-throughput sequencing results;

[0054] Step 5: The high-throughput sequencing results are subjected to the following steps: prediction and removal of adapter sequences, sequence splitting, deletion of barcode tag sequences, sequence splicing, and species annotation.

[0055] Specifically, prediction and removal of the adapter sequence: use fastp to predict and remove the adapter sequence, and the length of the sliding window is set to 16;

[0056] Sequence splitting: Use fastq-multx to split the sequencing results of the mixed sample according to the barcode tag sequence unique to each large fungal sample. Place the split sequence files into different folders according to the sample ID and build manifests for each;

[0057] Deletion and sequence splicing of barcode tag sequences: Use DADA2 to delete and splice the barcode tag sequences of each macrofungal sample sequence;

[0058] Species Annotation: All sequences were annotated with species using the trained annotator.

[0059] The annotator is obtained after training using the UNITE database with a large amount of large fungal ITS sequence information added as a training set and using a Bayesian classification algorithm.

[0060] Example 1:

[0061] In this example, 100 large fungal samples (sample numbers M001-M100) were used as an example to perform second-generation high-throughput sequencing using the Barcode-specific primer markers provided by the present invention. The specific process is as follows:

[0062] (1) DNA extraction from large fungal samples

[0063] DNA was extracted using the Ezup column-type fungal genomic DNA extraction kit (Shanghai Sangon Biotechnology Co., Ltd.), and the operation method was carried out according to the kit extraction instructions.

[0064] (2) First round of PCR reaction

[0065] The first round of PCR amplification was performed to enrich the ITS1 and ITS2 target regions of the macrofungal samples. Using the extracted macrofungal DNA as template, ITS1 (ITS1F: 5'-ACCHGCGGARGGATCATTAC-3'; ITS2R: 5'-TTYDCNRCGTTCTTCATCGT-3') and ITS2 (gITS7F: 5'-GTGAVTCATCRARTNTTTGA-3'; ITS4R: 5'-CCTBCSCTTANTDATATGCC-3') universal primers were used as amplification primers, respectively. PCR amplification was performed using Rapid Taq Master Mix enzyme purchased from Vazyme (Nanjing Novozyme Biotechnology Co., Ltd.).

[0066] The PCR amplification system was 20 μL and consisted of 10 μL 2× Rapid Taq Master Mix, 1 μL upstream primer, 1 μL downstream primer, 1 μL DNA, and 7 μL ddH 2 O. In this reaction system, the DNA was from different macrofungal samples, and the primer concentration was 10 μM.

[0067] Reaction program: 95°C for 3 min; 95°C for 20 s, 47°C for 20 s, 72°C for 30 s, 35 cycles; 72°C for 5 min; and storage at 4°C.

[0068] Perform 1% agarose gel electrophoresis, adjust the voltage to 120V, electrophoresis for 25 minutes, and observe the electrophoresis results using a gel imaging system.

[0069] (3) Adding Barcode tag sequence

[0070] Using the first-round PCR reaction product as a template, a specific primer sequence combination was used for PCR amplification, and a second-round PCR amplification was performed to add a barcode tag sequence to the ITS sequence, as shown in the table below.

[0071]

[0072]

[0073] The PCR amplification system was 30 μL, consisting of 15 μL 2× Rapid Taq Master Mix, 1.5 μL upstream primer, 1.5 μL downstream primer, 2 μL first-round PCR product, and 10 μL ddH2O.

[0074] Reaction program: 95°C for 3 min; 95°C for 20 s, 47°C for 20 s, 72°C for 30 s, 10 cycles; 95°C for 20 s, 59°C for 20 s, 72°C for 30 s, 25 cycles; 72°C for 5 min; storage at 4°C.

[0075] (4) Sample mixing and PCR product purification

[0076] Combine the second-round PCR products from all samples and run electrophoresis on a 1% agarose gel at 120 V for 25 min. Observe the results using a gel imaging system. Cut the gel containing the target fragment and recover it using the SanPrep Column-Based DNA Gel Recovery Kit (Shanghai Sangon Biotechnology Co., Ltd.). Refer to the manufacturer's instructions for specific procedures. Check the concentration and purity of the recovered product using a OneDrop-1000 and store at -20°C until needed.

[0077] (5) Library construction and sequencing

[0078] The PCR products recovered above were sent to the Xi'an branch of Beijing Qingke Biotechnology Co., Ltd. for library construction and sequencing. The mixed PCR product samples were sequenced with double ends using the Illumina HiSeq PE250 sequencing platform, and the sequencing depth of the samples was 3G.

[0079] (6) Use the script to automatically process the high-throughput sequencing results, as follows:

[0080] ① Use fastp to predict and remove the adapter sequence, and set the sliding window length (-w) to 16;

[0081] ② Sequence splitting: Use fastq-multx to split the sequencing results of the mixed sample according to the Barcode tag sequence unique to each large fungal sample. Place the split sequence files into different folders according to the sample ID, and construct manifests for each.

[0082] ③ Deletion of barcode tag sequences and sequence splicing: Use DADA2 to delete the barcode tag sequences and sequence splicing for each large fungal sample sequence.

[0083] ④ Species annotation: All sequences were annotated with species using a trained annotator. The annotator was trained using the UNITE database with a large amount of large fungal ITS sequence information as the training set and a Bayesian classification algorithm.

[0084] (7) Data collation

[0085] Organize the molecular identification results.

[0086] Comparative Example 1:

[0087] Samples M001-M100 were sequenced using first-generation sequencing technology. The specific steps are as follows:

[0088] PCR amplification was performed using the extracted macrofungal sample DNA as a template and ITS universal primers (ITS4: 5′-TCCTCCGCTTATTGATATG-3′; ITS5: 5′-GGAAGTAAAAGTCGTAACAAGG-3′) as amplification primers using Rapid Taq Master Mix enzyme purchased from Vazyme (Nanjing Novozyme Biotechnology Co., Ltd.).

[0089] The PCR amplification system was 20 μL and consisted of 10 μL 2× Rapid Taq Master Mix, 1 μL upstream primer, 1 μL downstream primer, 1 μL DNA, and 7 μL ddH 2 O. In this reaction system, the DNA was from different macrofungal samples, and the primer concentration was 10 μM.

[0090] Reaction procedure: 95°C for 3 min; 95°C for 20 s, 48°C for 20 s, 72°C for 30 s, 35 cycles; 72°C for 5 min; and storage at 4°C.

[0091] The PCR products were sent to the Xi'an branch of Beijing Qingke Biotechnology Co., Ltd. for first-generation Sanger sequencing. The sequencing primers were the primers used for PCR amplification. The sequencing was required to be bidirectional. The successfully sequenced sequences were spliced ​​one by one to obtain the ITS sequence of the large fungal sample.

[0092] The ITS sequences obtained by the first-generation Sanger sequencing were uploaded to NCBI BLASTN for sequence alignment to obtain the molecular identification results of macrofungi (see the table below for details).

[0093]

[0094] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention.

Claims

1. A high-throughput molecular identification method for large fungi based on second-generation sequencing technology, characterized in that: The method is: The ITS1 and ITS2 sequences were amplified by PCR using the extracted macrofungal genomic DNA as template to obtain PCR products. Using the PCR product as a template, adding barcode tag sequences to the ITS1 and ITS2 sequences through a combination of specific primer sequences to obtain ITS1 and ITS2 PCR products with barcode tag sequences; All ITS1 and ITS2 PCR products with barcode tag sequences were mixed separately, and then recovered by gel excision and library construction to obtain mixed PCR product samples; The mixed PCR product sample was subjected to double-end sequencing using the Illumina HiSeq PE250 sequencing platform to obtain high-throughput sequencing results; The high-throughput sequencing results are used to predict and remove adapter sequences, split sequences, delete barcode tag sequences, splice sequences, and annotate species; Species annotation was performed on all sequences using the trained annotator; The annotator is obtained after training using the UNITE database with a large amount of large fungal ITS sequence information as a training set and using the Bayesian classification algorithm; The Barcode tag sequence is shown in the table: The universal primer sequences are ACCHGCGGARGGATCATTAC, TTYDCNRCGTTCTTCATCGT, GTGAVTCATCRARTNTTTGA, or CCTBCSCTTANTDATATGCC.

2. The high-throughput macrofungal molecular identification method based on second-generation sequencing technology according to claim 1, characterized in that: In the PCR amplification of ITS1 and ITS2 sequences, the PCR reaction procedure is: 95°C for 3 min; 95°C for 20 s, 47°C for 20 s, 72°C for 30 s, 35 cycles; 72°C for 5 min; and storage at 4°C.

3. The high-throughput macrofungal molecular identification method based on second-generation sequencing technology according to claim 2, characterized in that: The PCR reaction procedure for adding barcode tag sequences to ITS1 and ITS2 sequences by using a specific primer sequence combination is as follows: 95°C for 3 min; 95°C for 20 s, 47°C for 20 s, 72°C for 30 s, 10 cycles; 95°C for 20 s, 59°C for 20 s, 72°C for 30 s, 25 cycles; 72°C for 5 min; and storage at 4°C.

4. The high-throughput macrofungal molecular identification method based on second-generation sequencing technology according to claim 3, characterized in that: The sequence of the ITS1 product with the Barcode tag sequence is shown in the table:

5. The high-throughput macrofungal molecular identification method based on second-generation sequencing technology according to claim 4, characterized in that: The sequence of the ITS2 product with the Barcode tag sequence is shown in the table:

6. The high-throughput macrofungal molecular identification method based on second-generation sequencing technology according to any one of claims 1 to 5, characterized in that: The high-throughput sequencing results are subjected to the following steps: prediction and removal of adapter sequences, sequence splitting, deletion of barcode tag sequences, and sequence splicing: Prediction and removal of adapter sequences: Use fastp to predict and remove adapter sequences, and set the sliding window length to 16; Sequence splitting: Use fastq-multx to split the sequencing results of the mixed sample according to the barcode tag sequence unique to each large fungal sample. Place the split sequence files into different folders according to the sample ID and build manifests for each; Deletion and sequence splicing of barcode tag sequences: Use DADA2 to delete the barcode tag sequences and splice the sequences of each large fungal sample.

Citation Information

Patent Citations

  • Marker based on high-throughput sequencing and method and kit for capturing one or multiple specific genes of multiple samples

    CN105524983A

  • Method for detecting pathogenic fungi of maize leaves based on high-throughput sequencing technology, and application

    CN111808936A