Primer for simultaneously identifying weeds and diseases carried by weeds based on three-generation sequencing and application

Through the application of third-generation sequencing technology and primers, the gap in the detection of weeds and pathogens in grain quarantine has been solved, and the rapid and accurate identification of weeds and diseases has been achieved, reducing the risk of disease transmission.

CN120758654APending Publication Date: 2025-10-10ANIMAL AND PLANT & FOOD DETECTION CENTER JIANGSU ENTRY EXIT INSPECTION AND QUARANTINE BUREAU
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510895216.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

In grain quarantine, existing technologies fail to effectively detect imported weeds and the pathogens they carry, posing a risk of pathogen introduction. Traditional detection methods also have problems such as double peaks and different sequencing.

Method used

Using primers and kits based on third-generation sequencing technology, through PCR amplification, library construction and high-throughput sequencing, combined with OTU cluster analysis and species annotation, rapid identification of weeds and the diseases they carry can be achieved.

Benefits of technology

It achieves accurate identification of weeds and disease carriers, avoids the defects of traditional detection methods, and can quickly and effectively identify weed species and pathogenic fungi that are difficult to identify morphologically, thereby reducing the risk of crop diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120758654A_ABST
    Figure CN120758654A_ABST
Patent Text Reader

Abstract

The invention relates to the field of molecular biology, in particular to the field of gene detection, and more particularly relates to a primer for simultaneously identifying weeds and diseases carried by the weeds based on three-generation sequencing and application of the primer. According to the method, a mode of simultaneously detecting the weeds and the carried diseases is provided for the first time, so that the problems of double peaks, different sequencing and the like which often occur in a traditional detection mode can be avoided, and meanwhile, species information of the weeds and the carried diseases can be obtained. Particularly, the method has important significance on amaranth weeds and the like which are difficult to identify in morphology. By utilizing the technical scheme disclosed by the invention, the species of weeds and plant pathogenic fungi carried by the weeds can be quickly and effectively identified, and disease risks and ecological threats of crops and other plants caused by the weeds and the pathogenic fungi carried by the weeds are avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of molecular biology, in particular to the field of gene detection, and more specifically to primers and applications for simultaneously identifying weeds and the diseases they carry based on third-generation sequencing. Background Art

[0002] my country is a major grain importer, and the safety of imported grain is crucial to national security. Therefore, the quarantine and identification of imported grain is a crucial task for quarantine professionals. Currently, grain quarantine generally focuses on two areas: invasive weeds and fungal plant diseases carried by grain.

[0003] Third-generation sequencing (NGS) is a novel sequencing technology that combines the advantages of high throughput, rapid speed, long read length, and low cost. Currently, NGS is divided into two main categories based on their technical principles. One is single-molecule fluorescence sequencing, exemplified by Helicos' SMS technology and Pacific Bioscience's SMRT technology. This technology uses fluorescently labeled deoxynucleotides, and then uses a microscope to record changes in fluorescence intensity in real time. When the fluorescently labeled deoxynucleotide is incorporated into a DNA strand, its fluorescence is detected simultaneously along the DNA strand. Once the deoxynucleotide forms a chemical bond with the DNA strand, the fluorescent group is cleaved by the DNA polymerase, and the fluorescence disappears. This fluorescently labeled deoxynucleotide does not affect the activity of the DNA polymerase, and after fluorescence cleavage, the synthesized DNA strand is identical to the native DNA strand. The other is nanopore sequencing, exemplified by Oxford Nanopore in the UK. This novel nanopore sequencing method uses electrophoresis to drive individual molecules through a nanopore, achieving sequencing. Because the diameter of the nanopore is very small, it only allows a single nucleic acid polymer to pass through. The charge properties of individual ATCG bases are different, and the type of base passing through can be detected by the difference in electrical signals, thereby achieving sequencing.

[0004] The third-generation sequencing technology mentioned in the present invention refers to the first type of third-generation sequencing method based on the principle of single-molecule fluorescence sequencing technology. Summary of the Invention

[0005] The inventors of this invention have discovered that currently, grain quarantine efforts have yet to address the presence of weeds mixed with grain, much less the potential for these weeds to carry pathogens and the types of pathogens they may harbor. If weeds and the pathogens they may carry are allowed to enter the country unchecked, these weed- and pathogen-laden grains could become a primary source of infection for crop pathogens during the following growing season, or serve as intermediate hosts, spreading plant diseases.

[0006] So, what foreign weeds might be present in imported grain? Do these weed seeds carry pathogens? How can we quickly detect and identify these weeds and the pathogens they carry? These are all questions worth pondering and urgently awaiting exploration by plant quarantine professionals.

[0007] Currently, these issues are largely unresolved in the field of imported grain quarantine. Quarantine workers urgently need to conduct research on the detection of plant pathogens carried by weeds in imported grain to fill this gap and prevent the risk of pathogens being introduced into my country through weed vectors.

[0008] In order to solve the above technical problems, the present invention discloses primers for simultaneously identifying weeds and the diseases they carry based on third-generation sequencing, including primer F (SEQ ID NO: 1): 5'-GGAAGTAAAAGTCGTAACAAGG-3' and primer R (SEQ ID NO: 2): 5'-TCCTCCGCTTATTGATATGC-3'.

[0009] The present invention further discloses the application of the primers. One of the application modes is a kit for simultaneously identifying weeds and the diseases they carry based on third-generation sequencing, wherein the kit includes the primers.

[0010] At the same time, the kit further contains PCR amplification reagents, sequencing library construction reagents and gene purification reagents.

[0011] The weeds referred to here include but are not limited to: Amaranthus quinquefolius, Amaranthus serrata, wild oats, Brassica napus, wild cabbage, Brassica rapa, Bromus oleracea, Bromus oleracea, Bromus chinensis, Bromus serrata, Bromus chinensis, Chenopodium album, Festuca australis, Falobasidium stepposum, Prototheca serrata, Lentil, Lolium multiflorum, Lolium perenne, Polygonum amphibians, Polygonum jojoba, Polygonum sorrel, Persicaria sp. P2241, Polygonum indigofera, Polygonum aviculare, Radish, Wild Radish, and Eurasian celery.

[0012] The diseases include but are not limited to infections with Alternaria alternata, Aspergillus flavus, Epichloe occultans, Cylindrospermum leaf spot, and Candida.

[0013] In addition, the present invention further discloses another application of the above primers, namely, a method for simultaneously identifying weeds and the diseases they carry based on third-generation sequencing, comprising the following steps: S1: Collect weeds intercepted from imported grains and extract DNA; S2: Establish PCR amplification system and amplification procedure; S3: Purify the PCR amplification products, quantify the PCR amplification products, and perform homogenization treatment on the PCR products based on the quantitative results; S4: Construction of PacBio library; S5: high-throughput sequencing using PacBio program; S6: Use sequencing data to identify weeds and disease-carrying organisms.

[0014] In some embodiments, the DNA extraction method in step S1 involves first surface disinfecting the weeds to be tested, then adding them to a culture medium containing an antibiotic and incubating them at room temperature with shaking. All weeds are then ground into a powder using liquid nitrogen, and genomic DNA from the weed sample mixture is extracted according to the procedures in the DNeasy Plant Mini Kit instructions. Surface disinfection can be performed by first surface disinfecting all weed samples in a sterilized beaker (or other container). In a specific embodiment, the antibiotic used is penicillin-streptomycin (100X), with 1 mL of penicillin-streptomycin added per 100 mL of culture medium. Furthermore, in a specific embodiment, the shaking incubation at room temperature is performed at a rotation speed of 150 rpm, a room temperature of 24°C, and a time of 10 hours. Specifically, the conical flask containing the sample can be placed in a shaker and incubated at a rotation speed of 150 rpm and a temperature of 24°C for 10 hours.

[0015] In some embodiments, the PCR amplification procedure is: a. 95°C, 5 min; b. 95℃, 30s; 58℃, 30s, 72℃, 45s; c. 72°C, 10 min, stop.

[0016] In some embodiments, the PCR amplification reaction system is:

[0017] In some embodiments, the PCR amplification products are quantitatively quantified using fluorescence quantification.

[0018] In some embodiments, the PCR products are homogenized by mixing them in proportion according to the sample sequencing amount requirement, so that the product of each PCR amplification product and its proportion is equal.

[0019] In some embodiments, constructing a PacBio library comprises the following steps: S4-1: Connect the third-generation sequencing "Y"-shaped adapter; S4-2: Magnetic bead screening to remove linker self-ligated fragments; S4-3: Enrichment and amplification of the library template using PCR amplification; S4-4: Denaturation with sodium hydroxide to obtain single-stranded DNA fragments and complete the construction of the third-generation sequencing PacBio library.

[0020] In some embodiments, high-throughput sequencing using the PacBio program comprises the following steps: S5-1: primer annealing to template; S5-2: Form a complex between fluorescently labeled dNTPs and enzyme + DNA template. The enzyme here can be NEB T4 DNA Rapid Ligase, and the DNA template refers to a DNA fragment in the PacBio library. S5-3: Collect the fluorescence signal emitted by the fluorescent dNTP after being irradiated by laser; S5-4: The enzyme reaction extends the chain and causes the fluorescent group on the dNTP to fall off; S5-5: The polymerization reaction continues and sequencing is completed simultaneously.

[0021] In some embodiments, determining information about weeds and pest-carrying organisms using sequencing data includes the following steps: S6-1: The data of each sample is distinguished according to the index sequence. The extracted data is saved in fastq format. Each sample of PE data has two files, fq1 and fq2, which contain reads at both ends of the sequencing, and the sequences correspond one to one in order; S6-2: After the PacBio data was downloaded, the instrument's built-in SMRTLINK (v11) was used to obtain the consensus CCS sequence (circular consensus sequencing); S6-3: Obtain statistical analysis results of community structure using OTU cluster analysis method and species annotation.

[0022] The OTU clustering analysis method can be based on 98.65% similarity OTU clustering, which is the international default standard for the full-length 16s Strain level, or the DADA2 OTU clustering method based on Qiime2.

[0023] Specifically, in some embodiments, the OTU cluster analysis steps are as follows: S6-3-1: Extract non-repeating sequences from the optimized sequence; S6-3-2: Remove single sequences without duplication; S6-3-3: OUT clustering was performed on non-repetitive sequences according to 97% similarity. Chimeras were removed during the clustering process to obtain representative OUT sequences. S6-3-4: Map all optimized sequences to the OUT representative sequence, select sequences with a similarity of more than 97% to the OUT representative sequence, and generate the OUT table, which includes: 01.original / : original non-flattened result; 02.normalize / : the result of the draw; otu_table.xls: Statistics table of sequence numbers in otu of each sample; otu_rep.fasta: otu representative sequence in fasta format; otu_table.biom: biom format otu table.

[0024] In some embodiments, species annotation is performed using the uclust algorithm and the RDP Classifier algorithm on OTU representative sequences at a 97% similarity level. The community composition of each sample is calculated at each taxonomic level: domain, kingdom, phylum, class, order, family, genus, and species.

[0025] In some embodiments, the confidence threshold of the uclust algorithm analysis is 0.8, and in some embodiments, the confidence threshold of the RDP Classifier algorithm is 0.7.

[0026] This invention proposes, for the first time, a method for simultaneous detection of weeds and the diseases they carry. This method avoids the problems of double peaks and sequencing discrepancies often encountered in traditional detection methods, while also providing information on the species of both weeds and the diseases they carry. This is particularly valuable for morphologically challenging weeds such as Amaranth. The technical solution disclosed in this invention enables rapid and effective identification of weeds and the species of the plant pathogens they carry, mitigating the disease risks and ecological threats to crops and other plants caused by weeds and their pathogens. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] 图1 This is the experimental result of validating the specific fluorescence PCR method for barley leaf spot pathogen. DETAILED DESCRIPTION

[0028] For a better understanding of the present invention, the present invention is further described below with reference to specific examples. The experimental methods used in the following examples are conventional experimental methods unless otherwise specified; the materials and reagents used are commercially available reagents and materials unless otherwise specified. Example 1

[0029] In this example, weeds in imported barley are taken as an example.

[0030] This batch of imported barley was from France. The weeds were separated and sent to experts for identification. The weed information is shown in Table 1. Table 1:

[0031] The nucleic acid of the above weeds was extracted according to the following method: First, sterilize all weed samples in a sterilized beaker (or other container) and add them to an Erlenmeyer flask containing culture medium and antibiotics. The antibiotic of choice is penicillin-streptomycin (100X), with 1 mL of penicillin-streptomycin added per 100 mL of culture medium. The Erlenmeyer flask containing the samples is then placed in a shaker at 150 rpm and 24°C for 10 hours. Given the small number of weeds, this step allows for the effective enrichment of pathogens carried by the weeds. Then, grind all weeds into a powder using liquid nitrogen. Genomic DNA from the weed sample pool is extracted according to the instructions in the DNeasy Plant Mini Kit. The pooled DNA sample is stored at -20°C until further use.

[0032] The mixed DNA was amplified by PCR using TransGen AP221-02: TransStart Fastpfu DNA Polymerase and ABI GeneAmp® 9700 PCR instrument.

[0033] All samples were analyzed under formal experimental conditions, with three replicates per sample. PCR products from the same sample were mixed and analyzed by 2% agarose gel electrophoresis. PCR products were recovered by gel excision using the AxyPrep DNA Gel Recovery Kit (AXYGEN) and eluted with Tris-HCl. Detection was then performed by 2% agarose gel electrophoresis.

[0034] PCR was performed using TransStart Fastpfu DNA Polymerase in a 20 μl reaction system.

[0035] The PCR amplification procedure is: a. 95°C, 5 min; b. 95℃, 30s; 58℃, 30s, 72℃, 45s; c. 72°C, 10 min, stop.

[0036] PCR products were quantified using the QuantiFluor™-ST Blue Fluorescence Quantitation System (Promega). The mixture was then homogenized by adjusting the mixing ratio according to the sequencing requirements for each sample. After homogenization, the gDNA amount was 1-2 ng.

[0037] Next, we constructed the SMRTbell Express TPK2.0 library. First, we ligated the Y-shaped adapters used in the third-generation sequencing library construction. Magnetic bead screening was then used to remove adapter-ligated fragments. PCR amplification was then performed to enrich the library template. Finally, sodium hydroxide was used for denaturation to obtain single-stranded DNA fragments. The PCR amplification procedure and reaction system used here were identical to those previously described.

[0038] Then, perform PacBio sequencing as follows: 1) Primer annealing to template; 2) Fluorescently labeled dNTPs form a complex with the enzyme + DNA template and bind briefly; 3) Fluorescent dNTPs are irradiated by laser light, emitting fluorescence, and the fluorescence signal is collected; 4) The enzyme reaction process extends the chain while simultaneously causing the fluorescent group on the dNTP to fall off; 5) The polymerization reaction continues and sequencing is performed simultaneously.

[0039] During library construction, 0.6× AMpure PB was used for purification.

[0040] The sequencing results were further analyzed.

[0041] First, the data of each sample is distinguished according to the index sequence, and the extracted data is saved in fastq format. Each sample of PE data has two files, fq1 and fq2, which contain the reads at both ends of the sequencing, and the sequences correspond one to one in order.

[0042] Fastq is a file format used in high-throughput sequencing technologies that reflects the base quality of sequenced data. Each read contains four lines of information: the first and third lines consist of a file identifier and a read name (ID) (the first line begins with "@" and the third line begins with "+"; the ID in the third line can be omitted, but the "+" cannot). The second line contains the base sequence, and the fourth line contains the sequencing quality value corresponding to each base in the sequence content of the second line.

[0043] Then, the PacBio data was optimized. After the PacBio data was downloaded from the instrument, the consistent CCS sequence (circular consensus sequencing) was obtained using SMRTLINK (v11) provided by the instrument. The accuracy of the CCS sequence reached the QV20 (99% accuracy) level.

[0044] The effective sequence statistics table of this sample is displayed in the directory 00.Datastat / , where the sample tested this time has Sequences: 34605; Bases (bp): 23484410; Average Length (bp): 678.64.

[0045] Furthermore, we performed OUT clustering and species annotation, and the results are in the directory 01.OTU.tax / 01.OTU / .

[0046] In the OUT cluster analysis, non-repetitive sequences are first extracted from the optimized sequences to reduce the amount of redundant calculations in the intermediate analysis process. Then, single sequences without repetitions are removed. Then, OTU clustering is performed on non-repetitive sequences (excluding single sequences) according to 97% similarity. Chimeras are removed during the clustering process to obtain representative sequences of OTUs.

[0047] Map all optimized sequences to the OTU representative sequence, select sequences with a similarity of more than 97% to the OTU representative sequence, and generate an OTU table in the result directory 01.OTU.tax / 01.OTU / . The descriptions of each table are as follows: 01.original / : original non-flattened result; 02.normalize / : the result of the draw; otu_table.xls: Statistics table of sequence numbers in otu of each sample; otu_rep.fasta: otu representative sequence in fasta format; otu_table.biom: biom format otu table.

[0048] Based on the results of the 02.normalize / method, we further performed species annotation. To obtain the species classification information corresponding to each OTU, we used the uclust algorithm to perform taxonomic analysis on the OTU representative sequences at a 97% similarity level. We also calculated the community composition of each sample at each taxonomic level: domain, kingdom, phylum, class, order, family, genus, and species. The comparison database is as follows: Silva (Release138.1 http: / / www.arb-silva.de); RDP (Release 11.5 http: / / rdp.cme.msu.edu / ); Greengene (Release 13.8 http: / / greengenes.secondgenome.com / ); Unite (Release 8.2 http: / / unite.ut.ee / index.php).

[0049] Software and algorithms: The uclust algorithm (http: / / www.drive5.com / usearch / manual / uclust_algo.html) uses a confidence threshold of 0.8. The RDP Classifier (version 2.2 http: / / sourceforge.net / projects / rdp-classifier / ) uses a confidence threshold of 0.7.

[0050] The results are stored in the directory: 01.OTU.tax / 01.OTU / , where: 01.original / : original non-flattened result; 02.normalize / : the result of the draw; otu_taxa_table.xls: Statistics of sequence numbers in each sample otu and OTU species annotation information; otu_taxa_table.biom : otu species classification table in biom format; tax_summary_a / : statistics of sample sequences at each taxonomic level; tax_summary_r / : Statistics table of relative abundance percentage of sample sequences at each taxonomic level.

[0051] According to the species annotation, the number of species annotated to each classification level (Phylum, Class, Order, Family, Genus, Species) for each sample was counted. The results are shown in Table 3: Table 3: Taxonomy ZC-F Phylum 8 Class 12 Order 26 Family 33 Genus 47 Species 81 Furthermore, we see the species composition analysis structure in the directory 03.Community / , as shown in Table 4. Table 4:

[0052] The table shows the species component names and the number of times the characteristic sequences were identified.

[0053] As shown in Table 4, a total of 81 weed and plant pathogenic fungi were detected to the species level in French barley weed testing. Of these, 32 were valid detections with read counts of 5 or higher, 26 weed species, and 7 plant pathogenic fungi were detected to the species level. The results of next-generation sequencing and morphological identification of weeds to the species level were generally consistent, and even morphologically challenging Amaranth weeds were well distinguished to the species level. Testing for pathogens carried in weed seeds revealed multiple pathogens, including Ramularia collo-cygn (marked in red), a pathogen of particular concern. Ramularia collo-cygn is one of the dangerous plant pathogens prohibited under bilateral trade agreements for imported barley between my country and the United States and other countries.

[0054] Furthermore, in order to verify the accuracy of the method of the present invention, we performed a fluorescence PCR method specific for Ramularia collo-cygni from barley (Hordeum vulgare) on French weed nucleic acid samples according to the method of the reference (JMG Taylor, LJ Paterson and ND Havis, A quantitative real-time PCR assay for the detection of Ramularia collo-cygni from barley (Hordeum vulgare). Applied Microbiology, 50 (2010) 493-499). The results are as follows: 图1 As shown, Ramularia collo-cygn was detected in weeds in France.

[0055] The above is a specific embodiment of the present invention. It should be noted that those skilled in the art can make several improvements and modifications without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. Primers for simultaneous identification of weeds and their disease carriers based on third-generation sequencing, including primer F: 5'-GGAAGTAAAAGTCGTAACAAGG-3' and primer R: 5'-TCCTCCGCTTATTGATATGC-3'.

2. Use of the primers according to claim 1 in the simultaneous identification of weeds and the diseases they carry based on third-generation sequencing.

3. The use according to claim 2, characterized in that: A kit for simultaneously identifying weeds and the diseases they carry based on third-generation sequencing, comprising the primers described in claim 1.

4. The use according to claim 3, characterized in that: The kit also contains PCR amplification reagents, sequencing library construction reagents, and gene purification reagents.

5. The use according to claim 2, characterized in that: The weeds to be tested are first surface disinfected, then added to a culture medium containing antibiotics and cultured under shaking at room temperature.

6. The use according to claim 5, characterized in that: The antibiotic is penicillin-streptomycin (100X).

7. The use according to claim 6, characterized in that: The amount of antibiotic added was 1 mL of penicillin and streptomycin per 100 mL of culture medium.