Method for identifying invasive plants in soil environment DNA sample by using chloroplast genome data
By constructing a whole gene database of invasive plant chloroplasts and PCR amplification, combined with high-throughput sequencing technology, the accuracy of invasive plant species identification in mixed soil environmental DNA samples was solved, and efficient detection of multiple invasive plants was achieved.
Patent Information
- Application Number
- CN202510472585.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art is difficult to effectively use high-throughput sequencing technology to accurately analyze invasive plant species in mixed soil environmental DNA samples, especially due to the huge amount of DNA data brought by animal and microbial sheddings.
By constructing a whole gene database of chloroplasts invasive plants, animal and microbial sheddings were screened out, and only the extracted eDNA was PCR amplified, and the sequencing was performed using high-throughput sequencing technology, and comparing it with the chloroplast gene database to identify invasive plant species.
Accurate and comprehensive analysis of invasive plant species in soil environmental DNA samples is achieved, identification efficiency and accuracy are improved, chloroplast gene information can be detected in a variety of invasive plants, and the shortcomings of a single sample are made up.
Smart Images

Figure CN120290769A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of gene identification, and particularly relates to a method for identifying invasive plants in soil environmental DNA samples by using chloroplast genome data. Background Art
[0002] Samples contain abundant environmental DNA (eDNA), which comes from the shed materials (including shed cells, individual remains, pollen, etc.) of animals, plants, and microorganisms living in their vicinity. The DNA data brought by these shed materials is huge and highly mixed, which poses a severe challenge to the effective application of high-throughput sequencing technology.
[0003] Therefore, how to use high-throughput sequencing technology to analyze the invasive plant species in this region from the DNA in mixed samples is a current problem in this field. Summary of the Invention
[0004] To solve the problems existing in the above-mentioned prior art, the present invention provides a method for identifying invasive plants in soil environmental DNA samples by using chloroplast genome data. Only the PCR amplicons of the chloroplast genome are performed on the extracted eDNA, and the shed materials of animals and microorganisms are screened out, thereby simplifying the DNA data volume and being able to accurately and comprehensively complete the analysis of invasive plant species in the region.
[0005] The specific technical solution adopted by the present invention is as follows:
[0006] A method for identifying invasive plants in soil environmental DNA samples by using chloroplast genome data, characterized in that it includes the following steps:
[0007] S1. Construct a chloroplast whole-gene database of invasive plants;
[0008] S2. Dispersedly set a plurality of sample points within the demarcated area, and collect samples at each sample point;
[0009] S3. Respectively extract eDNA from the samples at each sample point;
[0010] S4. Perform PCR amplification on the extracted eDNA to obtain amplicon products;
[0011] S5. Use high-throughput sequencing technology to sequence the amplicon products to obtain high-throughput data;
[0012] S6. Use the chloroplast whole-gene database of invasive plants to compare with the high-throughput data to obtain the types of invasive plant species.
[0013] In the described step S1, the method for constructing the database includes establishing the complete chloroplast genome data of known invasive plant species; for invasive plant species lacking complete chloroplast genome data, integrating the chloroplast gene fragments of the invasive plant species in the public database as the draft chloroplast genome of the invasive plant species; if there is complete chloroplast genome data of a related species of the invasive plant species, using the complete chloroplast genome data of its related species as a substitute for the invasive plant species; if there is no complete chloroplast genome data of a related species of the invasive plant species, using the chloroplast genome data of its congeneric species as a substitute.
[0014] In the described step S2, the sample collected is the topsoil.
[0015] In the described step S3, the extraction process of eDNA from the sample is as follows:
[0016] S3a. Mix the soil samples collected from the same sampling point;
[0017] S3b. Grind the mixed soil samples;
[0018] S3c. Add cell lysate to the soil samples to release DNA;
[0019] S3d. Adsorb DNA using a filter membrane;
[0020] S3e. Purify the adsorbed DNA once using a buffer;
[0021] S3f. Purify the adsorbed DNA twice using a buffer;
[0022] S3g. Evaporate the buffer to obtain dry DNA;
[0023] S3h. Add TLE solution to dissolve the dry DNA to obtain a pure DNA solution.
[0024] Alternatively, the sample collected in the described step S2 is surface plant debris.
[0025] In the described step S3, use a plant DNA extraction kit to extract eDNA from the surface plant debris sample.
[0026] In the described step S5, mix the amplicon products of the same sample and perform amplicon sequencing to obtain high-throughput data, and then perform DNA paired-end sequence data splicing to obtain the spliced data of the high-throughput data.
[0027] In the described step S4, perform PCR amplification on the chloroplast DNA in the extracted eDNA.
[0028] The beneficial effects of the present invention are:
[0029] The present invention establishes the complete chloroplast genome data of known invasive plant species; for invasive plant species lacking complete chloroplast genome data, the chloroplast gene fragments of the invasive plant species in the public database are integrated as the draft chloroplast genome of the invasive plant species; if there is complete chloroplast genome data of a related species of the invasive plant species, the complete chloroplast genome data of its related species is used as a substitute for the invasive plant species; if there is no complete chloroplast genome data of a related species of the invasive plant species, the chloroplast genome data of its congeneric species is used as a substitute. By constructing a database of complete chloroplast genomes of invasive plants, finally, 100% of the invasive plant species in the List of Invasive Plants in China are covered;
[0030] Chloroplast genes are relatively stable in the environment and can exist in the environment such as soil for a long time. Even under relatively harsh environmental conditions, they can better preserve their genetic information. Under conditions such as soil drought and high temperature, chloroplast gene fragments released by invasive plants into the topsoil or plant debris on the ground can still be detected for species identification, which has high feasibility and reliability;
[0031] Chloroplast genes are suitable for high-throughput sequencing technology, which can sequence and analyze a large number of samples simultaneously and quickly obtain a large amount of data. Through one high-throughput sequencing experiment, the chloroplast gene information of multiple invasive plants in soil eDNA samples can be detected, which is conducive to identifying multiple invasive plant species that may exist in the environment and helps improve the identification efficiency and accuracy;
[0032] The samples in the present invention include topsoil and plant debris samples on the ground. By collecting topsoil samples and plant debris samples on the ground, the deficiencies in a single sample can be made up, and the identification of invasive species can be more accurate and comprehensive. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is a schematic diagram of the operation process of the present invention;
[0034] Figure 2 It is a location map of sampling points;
[0035] Figure 3 It is a schematic diagram of the extraction process of eDNA in soil samples; DETAILED DESCRIPTION OF THE INVENTION
[0036] The present invention will be further described below in conjunction with the drawings and specific embodiments:
[0037] Example 1. The present invention relates to a method for identifying invasive plants in soil environmental DNA samples using chloroplast genome data. As Figure 1 shown, it includes the following steps:
[0038] S1. Construct a chloroplast whole - genome database of invasive plants. The chloroplast whole - genome database of invasive plants is constructed based on chloroplast whole - genomes and chloroplast fragments. The list of invasive plants refers to "List of Invasive Plants in China" published by Ma Jinshuang et al., which contains 397 invasive plant species. The specific operation for constructing the database is as follows: establish 230 chloroplast whole - genome data of 201 invasive plant species; for invasive plant species lacking chloroplast whole - genome data, by integrating existing chloroplast gene fragments of the invasive plant species in public databases, integrate these chloroplast gene fragments of different genes (including removing duplicates of the same fragments and concatenating different gene fragments) as the chloroplast genome draft of the species, including 1446 chloroplast gene fragments of 104 invasive plant species; if there is a species with existing chloroplast whole - genome data that is phylogenetically closely related to the invasive species, its closely related species can be used as a substitute; if there is no published chloroplast whole - genome of a closely related species for the invasive plant, the chloroplast genome of its congeneric species is used as a substitute, including 163 chloroplast whole - genomes of 146 invasive plants of different species within the same genus. Finally, available data of 397 invasive plant species are obtained, covering 100% of the invasive plant species in this list.
[0039] S2. In this embodiment, 8 relatively scattered sampling points are selected as sample plots for collecting samples throughout the entire territory of Hebei Province. As Figure 2 shown, each sampling point is YD011, YD034, YD058, YD065, YD105, YD116, YD145, YD152; at each sampling point, a 5m * 5m quadrat is selected, and surface soil samples are collected at five scattered positions within the quadrat.
[0040] S3. Extract eDNA from the surface soil samples of each sampling point respectively; as Figure 3 shown, the extraction process of eDNA from the surface soil samples is as follows:
[0041] S3a. Mix the 5 soil samples collected from the same sampling point as one sample;
[0042] S3b. Grind the mixed soil samples;
[0043] S3c. Add cell lysis buffer to the soil samples to promote DNA release; this cell lysis buffer is selected from the cell lysis buffer in the PowerSoil DNA Isolation Kit ( DNA Isolation Kit) of MOBIO;
[0044] S3d. Centrifuge the solution at 10000g for 10 minutes, transfer the supernatant to a new centrifuge tube, adsorb DNA using a DNA affinity filter membrane, incubate at room temperature for 2 minutes, and then centrifuge at 10000g for 10 minutes;
[0045] S3e. Purify the adsorbed DNA once with buffer and centrifuge at 10,000 g for 10 minutes;
[0046] S3f. Purify the adsorbed DNA twice with buffer and centrifuge at 10,000 g for 10 minutes;
[0047] S3g. Evaporate the buffer to dryness to obtain dry DNA;
[0048] S3h. Add 100 μL of TLE solution (Tris-LiCl-EDTA solution) to dissolve the DNA, centrifuge at 10,000 g for 10 minutes to obtain a pure DNA solution.
[0049] S4. Perform PCR amplification on the more hypervariable chloroplast genome regions in the extracted eDNA. The hypervariable chloroplast genome regions are: ndhF1, ndhF2, ndhF3, ndhF4, ndhF5, ndhF6, rpoB1, rpoB2, rpoB3, rpoB4, and rpoB5.
[0050] The primer sequence information for the hypervariable regions is shown in Table 1.
[0051] Table 1 Primer sequence information for hypervariable regions
[0052]
[0053] For the specific amplification procedure, refer to the literature Yanlei Liu, Chao Xu, Yuzhe Sun, Ping Wu, Xun Chen, Wenpan Dong, Xueying Yang* and Shiliang Zhou*. Method for quick DNA barcode reference library construction. 2021. Ecology and Evolution. 11, 11627 - 11638. After amplification, these amplicon products are recovered by gel cutting for later use.
[0054] S5. Mix the purified amplicon products of the same sample into one sample and send it to Novogene in Beijing for PE250 amplicon sequencing. The high-throughput data after sequencing is first quality-controlled using the NGS QC Tool Kit, and then the Pair end double-ended sequence data is assembled using the FLASH software based on the overlapping regions of the forward and reverse sequences. The assembled data is used for subsequent analysis of invasive plant species;
[0055] S6. The directly utilized data after splicing is used to implement the analysis of the quantity and abundance of species contained in the data by means of the assembly-free AFRAID method. The process of the AFRAID method refers to the literature: Yanlei Liu, Kai Chen, Lihu Wang, Xinqiang Yu, Chao Xu, Zhili Suo, Shiliang Zhou*, Shuo Shi* and Wenpan Dong*. Assembly-free reads accurate identification (AFRAID) approach outperforms other methods of DNA barcoding in the walnut family (Juglandaceae). 2024. Plant Diversity. 10, 2.
[0056] The types of invasive plant species are obtained by comparing the whole chloroplast gene database of invasive plants with high-throughput data.
[0057] Example 2. The difference between this example and Example 1 is that the sample collected in step S2 is surface plant debris; in step S3, a plant DNA extraction kit (Baypure magnetic bead method polysaccharide polyphenol plant genomic DNA extraction kit, product number PSDM-64) is used to extract eDNA from the surface plant debris sample, and the operation process refers to the instruction manual of this kit.
[0058] The surface plant debris covers the upper part of the topsoil layer. The samples of Example 1 and Example 2 can be collected simultaneously. The specific operation is to vertically insert a sampling cylinder into the soil at each sampling position, pull out the sampling cylinder. The lower part of the sample collected in the sampling cylinder is surface soil, and the upper part of the sample collected in the sampling cylinder is surface plant debris.
[0059] Through the high-throughput sequencing of Example 1 and Example 2, a total of 2,597,088 high-throughput data were finally obtained, and the specific situation is shown in Table 2.
[0060] Table 2 Sample and high-throughput data acquisition situation
[0061]
[0062]
[0063] In addition, during sampling, on-site investigations of the plants within and around the quadrat were carried out as a control, and the situation of invasive plants found is shown in Table 3.
[0064] Table 3 Situation of invasive plants found in the surface plant investigation
[0065]
[0066] The obtained high-throughput data was identified using the chloroplast genome data in the existing database.
[0067] The known invasive plant situations reflected by the soil eDNA in Example 1 are shown in Table 4:
[0068] Table 4 Invasive plants found in soil eDNA
[0069]
[0070] The known invasive plant situations reflected by the surface plant debris in Example 2 are shown in Table 5:
[0071] Table 5 Invasive plants found in surface plant debris
[0072]
[0073] The above data indicates that, compared with the surface plant debris samples, the soil eDNA shows a poorer reflection of Bidens pilosa. The surface plant debris samples not only cover all the invasive plants surveyed on the surface, but even include more invasive plant species that are invisible to the naked eye or overlooked. In summary, we achieved a qualitative assessment of plant invasive species using environmental samples, specifically the eDNA of surface plant debris samples.
Claims
1. A method for identifying invasive plants in soil environmental DNA samples using chloroplast genome data, characterized in that, Including the following steps: S1. Construct a complete chloroplast gene database of invasive plants; S2. Dispersedly set multiple sample points within the demarcated area and collect samples at each sample point; S3. Respectively extract eDNA from the samples at each sample point; S4. Perform PCR amplification on the extracted eDNA to obtain amplicon products; S5. Use high-throughput sequencing technology to sequence the amplicon products to obtain high-throughput data; S6. Use the complete chloroplast gene database of invasive plants to compare with the high-throughput data to obtain the types of invasive plant species.
2. The method for identifying invasive plants in a soil environmental DNA sample using chloroplast genome data according to claim 1, wherein: In the step S1, the construction method of the database includes establishing the complete chloroplast genome data of known invasive plant species; for invasive plant species lacking complete chloroplast genome data, integrate the chloroplast gene fragments of the invasive plant species in the public database as the draft chloroplast genome of the invasive plant species; if there is complete chloroplast genome data of a related species of the invasive plant species, use the complete chloroplast genome data of its related species as a substitute for the invasive plant species; if there is no complete chloroplast genome data of a related species of the invasive plant species, use the chloroplast genome data of its congeneric species as a substitute.
3. The method for identifying invasive plants in soil environmental DNA samples using chloroplast genome data according to claim 1, characterized in that: The sample collected in the step S2 is topsoil.
4. The method for identifying invasive plants in soil environmental DNA samples using chloroplast genome data according to claim 3, wherein: In the step S3, the extraction process of eDNA from the sample is as follows: S3a. Mix the soil samples collected at the same sample point; S3b. Grind the mixed soil samples; S3c. Add cell lysate to the soil samples to release DNA; S3d. Adsorb DNA using a filter membrane; S3e. Purify the adsorbed DNA once using a buffer; S3f. Purify the adsorbed DNA twice using a buffer; S3g. Evaporate the buffer to obtain dry DNA; S3h. Add TLE solution to dissolve the dry DNA to obtain a pure DNA solution.
5. The method for identifying invasive plants in soil environmental DNA samples using chloroplast genome data according to claim 1, wherein: The sample collected in the step S2 is surface plant debris.
6. The method for identifying invasive plants in a soil environmental DNA sample using chloroplast genome data according to claim 5, characterized in that: In the step S3, use a plant DNA extraction kit to extract eDNA from the surface plant debris sample.
7. The method for identifying invasive plants in soil environmental DNA samples using chloroplast genome data according to claim 1, characterized in that: In the step S5, mix the amplicon products of the same sample and perform amplicon sequencing to obtain high-throughput data, and then perform DNA paired-end sequence data splicing to obtain the spliced data of the high-throughput data.
8. A method for identifying invasive plants in soil environmental DNA samples using chloroplast genome data according to claim 1, characterized in that: In the step S4, perform PCR amplification on the chloroplast DNA in the extracted eDNA.