Cactus plant species identification method based on high-throughput sequencing data
By screening specific variant regions based on high-throughput sequencing data and designing specific primers, the problems of low resolution and high cost in the identification of cactus species have been solved, achieving high-precision and low-cost species identification, which is suitable for rapid detection in medicinal material markets and customs ports.
Patent Information
- Application Number
- CN202511226353.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-29
- Publication Date
- 2025-10-17
AI Technical Summary
The existing methods for identifying cactus species have problems such as low resolution of universal DNA barcode segments, lack of standardized extraction methods for chloroplast genome variable regions, and high cost of super barcodes, which limit their application.
Based on high-throughput sequencing data, specific variant regions were systematically screened, specific primers were designed, and species identification was performed by amplifying and sequencing the specific variant regions and comparing them with chloroplast genome databases.
It significantly improves the precision and accuracy of species identification of cacti, reduces testing costs and technical barriers, and makes rapid and low-cost on-site testing possible.
Smart Images

Figure CN120796570A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of bioinformatics, and specifically relates to a plant molecular identification method based on bioinformatics, and more specifically to a cactus plant species identification method based on high-throughput sequencing data. Background Art
[0002] Cactus Opuntia ) belongs to the Cactaceae family and is one of the most species-rich genera in the family. It is native to America and has been widely introduced. This genus of plants not only has unique morphology and adaptability to arid environments, but also has important medicinal value. In traditional medicine, prickly pear cactus ( O. ficus-indica ), Single-thorn Cactus ( O. monacantha ) and cactus ( O. dillenii Different parts of various cactus plants, such as Opuntia spp., are widely used to clear heat and detoxify, cool blood and stop bleeding, and promote qi and blood circulation, treating a variety of infectious, inflammatory, and traumatic conditions. However, species diversity within the genus is high, and some medicinal species share similar or even controversial morphological characteristics, which can easily lead to confusion or misuse, affecting efficacy and safety. Therefore, developing rapid and accurate identification techniques is crucial to ensure the effective utilization and quality control of medicinal cactus resources.
[0003] At present, the identification of Chinese medicinal material germplasm mainly relies on three technical paths: morphological identification, chemical component analysis, and DNA molecular identification. Traditional morphological identification methods have long been used as a basic means of quality control of natural medicines because of their convenience and intuitiveness. They are widely used in traditional Chinese medicine, Ayurveda medicine, and European and American herbal medicine systems. However, due to the destruction of traits during processing, reliance on empirical judgment, and the widespread existence of intraspecific variation and interspecific convergence, the accuracy of morphological methods in distinguishing powder preparations, extracts, and related species is limited. Analytical methods based on chemical components, such as high performance liquid chromatography (HPLC), can be used to identify medicinal materials, but because the content of chemical components is greatly affected by environmental and processing factors, and the detection process is complex and costly, it is difficult to meet the needs of efficient and standardized identification. To make up for the limitations of morphological and chemical methods, identification technologies based on DNA molecular markers have rapidly emerged in recent years. DNA barcode technology uses interspecific differences in short fragment sequences to achieve rapid and accurate identification of medicinal plants. Currently, rbcL 、 matK 、 ITS 、 trnH–psbA DNA barcodes such as nucleotide barcodes have been widely used as standard or auxiliary barcodes for the molecular identification of botanical medicinal materials and their adulterants. With the advancement of high-throughput sequencing and bioinformatics technologies, super-barcode strategies have significantly improved the resolution and phylogenetic inference capabilities of species identification in complex taxa by integrating multiple gene regions or complete chloroplast genomes. These strategies have been validated in the identification of various medicinal plants.
[0004] However, the germplasm resource identification of Cactaceae is facing severe challenges. There are a large number of species in the genus, and the morphologies are highly similar. In addition, there is a wide range of hybridization and phenotypic plasticity, which leads to a large number of varieties and subspecies classification units being described. Traditional DNA barcodes (such as rbcL 、 matK 、 ITS ) have insufficient resolution in the identification of Cactaceae species. Especially when facing morphologically similar, frequently hybridized or recently diverged closely related species, a single barcode fragment is difficult to effectively distinguish. High-throughput sequencing technology (such as genome skimming) makes it possible to obtain complete chloroplast genomes, providing a technical basis for developing super barcodes and high-variation molecular markers. Some studies have used chloroplast genome data to construct phylogenetic trees, but have not systematically mined and standardized specific high-variation regions as identification tools. The current problems mainly include: 1) low resolution of universal DNA barcode segments; 2) lack of standardized extraction of variation regions from chloroplast genomes; 3) super barcodes have large amounts of information, but are high in cost and limited in application.
[0005] Therefore, it is urgent to develop a method for rapid and low-cost identification of Cactaceae species based on high-throughput sequencing data, systematic screening of specific high-variation regions, and application. SUMMARY
[0006] In order to solve the problems of low resolution of universal DNA barcode segments, lack of standardized extraction of variation regions from chloroplast genomes, and high cost and limited application of super barcodes in the existing identification methods of Cactaceae species, the present application provides a method for identifying Cactaceae species based on high-throughput sequencing data.
[0007] The technical scheme adopted by the present application is: a method for identifying Cactaceae species based on high-throughput sequencing data, comprising the following steps: S1. Obtain a chloroplast genome database of Cactaceae plants and specific variation regions, and design specific primers according to the specific variation regions; S2. Extract DNA from the sample to be tested, and use specific primers to amplify and sequence the specific variation regions of the DNA of the sample to be tested; S3. Compare the sequencing results with the standard sequences in the chloroplast genome database to identify the species.
[0008] As a preferred embodiment, the step S1 comprises: S1-1. Collect a plurality of Cactaceae samples, respectively extract DNA, and perform genome-level sequencing. After assembly and annotation, obtain the chloroplast genome corresponding to the sample, and construct a chloroplast genome database; S1-2. Perform multi-sequence alignment and nucleotide polymorphism analysis on the obtained chloroplast genome to obtain a specific mutation region; S1-3. Design specific primers according to the obtained specific mutation region.
[0009] Preferably, the screening method of the specific mutation region comprises: Based on the chloroplast genome database, perform serial sequence alignment on each gene and spacer region to construct a corresponding matrix; According to the obtained matrix, calculate the nucleotide diversity Pi value of the protein coding gene CDS sequence and the intergenic spacer IGS sequence, and screen a specific mutation region according to the nucleotide diversity Pi value.
[0010] Preferably, the nucleotide diversity Pi value of the specific mutation region is not less than 0.015.
[0011] Preferably, the specific mutation region comprises infA 、 atpE-trnM-CAU 、 trnE-UUC-trnT-GGU 、 clpP- trnG-UCC_1 、 trnQ-UUG-rps16_1 、 rpoC1_1-rpoB 、 psbN-psbT 、 rps8-infA 、 atpA-atpF_2 Preferably, the specific mutation region consists of infA 、 atpE-trnM-CAU 、t rnE-UUC-trnT-GGU 、 clpP- trnG-UCC_1 、 trnQ-UUG-rps16_1 、 rpoC1_1-rpoB 、 psbN-psbT 、 rps8-infA 、 atpA-atpF_2 .
[0012] Preferably, the specific primers comprise at least one of the sequences shown in SEQ ID NO. 1-9.
[0013] The present application also provides a primer set for identifying Cactus plant species, which targets a specific mutation region comprising infA 、 atpE-trnM-CAU 、 trnE-UUC-trnT-GGU 、 clpP-trnG-UCC_ 1 、 trnQ-UUG-rps16_1 、 rpoC1_1-rpoB 、 psbN-psbT 、 rps8-infA 、 atpA-atpF_2 .
[0014] As preferred, the primer set comprises specific primers comprising at least one of the sequences shown in SEQ ID NO. 1-9. Preferably, the specific primers consist of the sequences shown in SEQ ID NO. 1-9.
[0015] The present application also provides a kit for identifying Cactus plant species, comprising the primer set for identifying Cactus plant species. The kit can also comprise reaction reagents for PCR amplification, which are well known to those skilled in the art.
[0016] The present application also provides the use of the primer set for identifying Cactus plant species and the kit for identifying Cactus plant species in identifying Cactus plant species. The use comprises extracting DNA of a sample to be tested, using specific primers targeting specific variable regions to amplify and sequence specific variable regions of DNA of the sample to be tested, and comparing the sequencing results with standard sequences in a chloroplast genome database to identify the species.
[0017] The present application has the following beneficial effects: Based on high-throughput sequencing data, the present application screens specific variable regions in chloroplast genomes of Cactus plants, and develops a standardized rapid identification process for Cactus plant species by combining Pi value screening and specific variable region markers. Compared with traditional DNA barcode fragments, the identification accuracy of targeting specific variable regions is significantly improved, thereby effectively distinguishing species with high morphological overlap and improving identification accuracy. At the same time, this method greatly reduces the technical threshold and cost. By designing specific primers targeting specific variable regions to amplify DNA of the sample to be tested by PCR, whole genome sequencing and assembly are avoided, the cost and period of single sample detection are reduced, and the method is compatible with conventional PCR platforms, making it possible to perform on-site rapid detection in scenarios such as medicinal material markets and customs ports. BRIEF DESCRIPTION OF DRAWINGS
[0018] Figure 1 FIG. 1 is a diagram of identification results of the present application based on specific variable regions of a species to be tested and a genomic database. DETAILED DESCRIPTION
[0019] The present application can be implemented or applied in other different specific embodiments, and the details in the specification can be modified or changed based on different views and applications without departing from the spirit of the present application. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict. In the embodiments of the present application, the methods used are conventional methods, and the reagents used can be obtained from commercial channels.
[0020] To improve the accuracy of Opuntia germplasm identification, this embodiment takes 36 representative species of Opuntia (Table 1) as the research object, and develops a super DNA barcode identification technology based on the complete chloroplast genome sequence. Through the standardized sequencing, assembly and annotation process, the nucleotide variation pattern of the chloroplast genome of Opuntia is systematically analyzed. By using the complete chloroplast genome as a super barcode, the species resolution is expected to be significantly improved.
[0021] Table 1 Experimental materials and sources Species name Species number Source information XRZ-19 Chenshan Botanical Garden OQ613378 NCBI XRZ-22 Chenshan Botanical Garden OQ613380 NCBI XRZ-41 Chenshan Botanical Garden XRZ-20 Chenshan Botanical Garden OQ613382 NCBI OQ613383 NCBI OQ613384 NCBI XRZ-15 Chenshan Botanical Garden OQ613387 NCBI OQ613388 NCBI XRZ-10 Chenshan Botanical Garden OK448352 NCBI OQ613390 NCBI XRZ-23 Chenshan Botanical Garden XRZ-27 Chenshan Botanical Garden OQ613393 NCBI XRZ-01 Chenshan Botanical Garden OQ613395 NCBI OQ613396 NCBI XRZ-28 Chenshan Botanical Garden OQ613397 NCBI XRZ-24 Chenshan Botanical Garden XRZ-38 Chenshan Botanical Garden OQ613398 NCBI XRZ-31 Chenshan Botanical Garden MN114084 NCBI XRZ-11 Chenshan Botanical Garden OQ613403 NCBI XRZ-13 Chenshan Botanical Garden OQ613405 NCBI XRZ-30 Chenshan Botanical Garden XRZ-45 Chenshan Botanical Garden OQ613407 NCBI XRZ-07 Chenshan Botanical Garden XRZ-08 Chenshan Botanical Garden Specifically, the Opuntia plant species identification method based on high-throughput sequencing data provided in this embodiment mainly includes the following steps: S1-1. Collect multiple Opuntia samples, respectively extract DNA, genome skimming, after assembly and annotation, obtain the chloroplast genome of the corresponding sample, and construct a chloroplast genome database; S1-2. Based on the chloroplast genome database, perform sequence alignment of each gene and intergenic region in series, and construct the corresponding matrix; according to the obtained matrix, calculate the nucleotide diversity Pi value of the protein coding gene CDS sequence and the intergenic region IGS sequence, and screen the specific variation region according to the nucleotide diversity Pi value; S1-3. Design specific primers according to the obtained specific variation region; S2. Extract the DNA of the sample to be tested, and use specific primers to amplify and sequence the specific variation region of the DNA of the sample to be tested; S3. Compare the sequencing results with the standard sequence in the chloroplast genome database to identify the species.
[0022] Further, in step S1-1, high-quality genomic DNA is extracted from the epidermal tissue of the cactus stem by using the improved CTAB method, and the specific steps are as follows: (1) Sample pretreatment: take 100 mg of fresh stem epidermal tissue, freeze quickly with liquid nitrogen, and then grind thoroughly into powder with a mortar; (2) Lysis treatment: transfer the tissue powder to a 1.5 mL centrifuge tube, add 400 μL of FP1 lysis buffer (containing 1% β-mercaptoethanol) and 6 μL of RNase A (10 mg / mL), vortex for 1 min to mix evenly, and incubate at room temperature for 10 min; (3) Protein precipitation: add 130 μL of FP2 precipitation buffer, shake vigorously for 1 min, then centrifuge at 12,000 rpm for 5 min (room temperature), and collect the supernatant into a new centrifuge tube; (4) DNA precipitation: 0.7 volume of pre-cooled isopropanol was added to the supernatant, mixed gently, and centrifuged at 12,000 rpm for 2 min (4°C), and the DNA precipitate at the bottom of the tube was retained; (5) Purification by washing: 600 μL of 70% pre-cooled ethanol was added, vortexed for 5 min, and then centrifuged at 12,000 rpm for 2 min (4°C). This washing step was repeated twice. (6) Dry and dissolve: The inverted centrifuge tube was dried on sterile filter paper at room temperature for 5-10 min to remove residual ethanol. Finally, 50 μL of TE buffer (10 mM Tris-HCl, 1 mM EDTA, pH 8.0) was added, and the mixture was incubated at 65°C for 30 min, with gentle mixing to promote DNA dissolution.
[0023] Further, in step S1-1, the DNA sequencing library construction and sequencing analysis method of the chloroplast genome is as follows. First, the genomic DNA is extracted by conventional methods, and its quality is systematically evaluated, including visual observation of sample clarity to exclude color abnormalities or foreign matter contamination, detection of A260 / A280 ratio by Nanodrop spectrophotometer to evaluate purity, accurate determination of DNA concentration and total amount by Qubit fluorescence quantifier, and verification of DNA integrity and fragment distribution by 0.8% agarose gel electrophoresis. After obtaining qualified samples, the DNA is controllably broken by fragmentation enzyme, and then the DNA end is repaired by end repair kit to complete the dA tail structure. Then the sequencing adapter is connected to the two ends of the DNA fragment by T4 DNA ligase, and the target length (usually 300-500 bp) of DNA fragment is obtained by magnetic bead screening. According to the experimental requirements, selective PCR amplification can be performed to enrich the library, and the amplification product needs to be quantified again by Qubit and agarose gel electrophoresis to verify the concentration and fragment distribution. The qualified library is denatured to form single-stranded DNA, and then circular structure is generated by circularization reaction, and specific enzyme digestion is used to digest residual linear DNA to improve the quality of the library. Finally, according to the sequencing throughput requirement, the library is pooled, and DNBSEQ high-throughput sequencing platform is used to complete double-end sequencing to obtain raw sequencing data. This standardized process ensures that the library construction quality meets the requirements of subsequent bioinformatics analysis.
[0024] Further, in step S1-1, the chloroplast genome assembly method is as follows. The GetOrganelle software (v1.7.7) is used for chloroplast genome assembly. The raw sequencing data is double-end sequencing result (sample_1.fastq.gz and sample_2.fastq.gz), and iterative extension is performed by setting the angiosperm chloroplast reference database (-F embplant_pt) as the seed template. The execution command parameters include: the initial word size is set to 79 (-w 79), the number of iterations is limited to 15 rounds (-R 15), the k-mer gradient parameter group is set to 21, 45, 65, 85, 105, and 127 (-k 21, 45, 65, 85, 105, 127), and the number of parallel computing threads is 30 (-t 30). The contig extension state is dynamically monitored during running, and if extension stagnation or circular structure unclosed phenomenon occurs, the word size parameter is gradually adjusted (±5 units) to optimize sequence extension specificity. The final output result is saved in the output-plastome directory, which is verified for single circular structure integrity by the Bandage visualization tool, and the assembly accuracy is verified by Blastn comparison of the chloroplast genomes of related species.
[0025] Further, in step S1-1, the chloroplast genome annotation method is as follows. The chloroplast genome annotation is completed based on the Geneious Prime software, and a reference genome guided iterative annotation strategy is adopted. First, the circular genome sequence obtained by assembly is globally aligned with the related species Cactus chloroplast reference genome, and the collinearity structure is calibrated by using the MAFFT plug-in. The annotation information of protein coding genes, tRNA and rRNA is automatically transferred by using the "Annotate from Reference" function in the software, and the boundary positions of core functional genes such as O. pycnantha , ndh 、 rbcL 、 matK are focused on. For the fuzzy region of the intergenic region (IR / SC) boundary, secondary verification is performed in combination with the GeSeq online annotation platform (https: / / chlorobox.mpimp-golm.mpg.de / GenBank2Sequin.html). The tRNA recognition adopts the Aragorn algorithm, and the bacterial / chloroplast genetic code table is set for scanning. Finally, the "Circular Viewer" module is used to visualize the circular genome map, and the physical position overlap of the annotation characteristics is manually checked to ensure that the annotation of the splicing site of the rps12 and other split genes conforms to the typical structural characteristics of plant chloroplasts. The annotation result is exported as a standard GenBank format file, including CDS, exon, misc_feature and other biological feature annotation levels.
[0026] A total of 36 complete Opuntia chloroplast genomes were obtained in this study, with a length of 121295-153768 bp, consisting of a large single copy region (LSC, 87,248-101,497 bp), a small single copy region (SSC, 4,094-33.343 bp), and two inverted repeat regions (IR, 1,022-30625 bp). The total GC content of Opuntia chloroplast genomes ranged from 35.9% to 36.8%. We observed that the number of genes varied from 107 to 139, resulting in a range of 75 to 91 protein-coding genes, 30 to 37 tRNAs, and 4 to 8 rRNAs. In Opuntia, most genes are usually present as a single copy in a species, but some genes are randomly present as single or double copies in the 36 species. These genes include 5 ndh genes ( ndhA, ndhB, ndhF, ndhG, ndhH ), 3 rpl genes ( rpl2, rpl23, rpl32 ), 4 rps genes ( rps7, rps12, rpl15, rpl19 ), 8 tRNA genes ( trnE-UUC, trnF-GAA, trnI- CAU, trnI-GAU, trnL-CAA, trnN-GUU, trnQ-UUG, trnV-GAC ), 4 rRNA genes ( rrn23, rrn16, rrn5, rrn4.5 ), and ycf genes ( ycf1, ycf2 ). Among the 36 Opuntia species, 8 protein body-coding genes ( rpl16, petD, petB, atpF, rpoC1, rps16, ndhB, ndhA ) have one intron. The r ps12 and ycf3 genes have two introns. Notably, ycf2 genes are missing in many species. Meanwhile, ndhJ and ndhK genes are missing in O. polyacantha . The reduction in gene copy number and the loss of ndh , rpl and rps genes appear to have occurred independently in different Opuntia lineages. Further, in step S1-2, the nucleotide polymorphism analysis method is as follows. The chloroplast genome nucleotide substitution rate analysis is used to study the frequency of nucleotide substitution between different sequences in the chloroplast genome, reflecting the rate of genome evolution. This analysis helps to reveal the genetic differences between different species or different populations of the same species, genome stability, and adaptability, etc. The nucleotide substitution rate (Pi) is usually obtained by calculating the nucleotide difference between multiple sequences. The higher the Pi value, the greater the variability of the region, the faster the rate of genome evolution, and vice versa, which means that the region is more conservative and may be under stronger selective pressure. To explore highly variable genetic markers of chloroplast genomes, protein-coding gene (CDS) and non-coding region (nCDS) sequences were extracted from 37 Cereus species. The extraction process was achieved by Python script in CPStools software. MAFFT was used for sequence alignment of each gene and intergenic region to construct the corresponding matrix. The nucleotide diversity (Pi) of all protein-coding genes (CDS) and non-coding regions (introns and intergenic regions) was calculated by DNASP, with a window length of 1000 bp and a step size of 300 bp.
[0027] The results show that there are significant differences in the Pi values of different genes and regions in the chloroplast genomes of the analyzed Cereus plants (Table 2). Among all genes and regions, infA has the highest Pi value (0.02113), indicating that this gene evolves the fastest in Cereus plants and may be subject to weaker purifying selection or positive selection. Next is clpP (0.01094). In contrast, the Pi values of multiple genes and regions are 0, including petN, psbM and psbZ These genes or regions are completely conserved in the analyzed Cereus species, with no detected nucleotide variation, indicating that they may be subject to extremely strong purifying selection and are essential for maintaining the basic functions of chloroplasts. Among the remaining genes, the Pi values range from 0.00022 ( atpH ) to 0.00848 ( petL ). Generally, genes encoding ribosomal proteins (such as rps and rpl genes) and genes involved in the core process of photosynthesis (such as psa, psb, pet and atp genes) usually have lower Pi values, indicating that they are subject to stronger purifying selection and evolve more slowly. Some genes encoding NADH dehydrogenase subunits (such as ndhThe Pi values of the genes are relatively high, indicating that they may have relatively fast evolutionary rates. To evaluate the evolutionary rates of different intergenic spacer (IGS) regions in the chloroplast genomes of Cactaceae plants, the nucleotide polymorphism (Pi) values of 66 IGS regions were also calculated in this embodiment. Compared with coding regions, IGS regions are usually under weaker selective pressure, and thus their nucleotide substitution rates can reflect the background mutation rate and evolutionary rate of the genomic region. The results show that there are significant differences in the Pi values of different IGS regions in the chloroplast genomes of Cactaceae plants analyzed, indicating that their evolutionary rates differ greatly. Among all the IGS regions, atpE- trnM-CAU The IGS region between trnT-UGU and trnL-UAA has the highest Pi value (0.02818), indicating that this region has the fastest evolutionary rate in Cactaceae plants and may have accumulated more mutations. Next are trnE-UUC-trnT-GGU (0.02267) and clpP-trnG-UCC_1 (0.02060). In contrast, the Pi values of three IGS regions are 0, including rpl22-rps3, psaB-psaA, psbL-psbF and psbF-psbE . These regions are completely conserved in the Cactaceae species analyzed, and no nucleotide variation is detected, indicating that they may be subject to certain functional or structural constraints, or that mutations in these regions have a large negative impact on the fitness of the plant. Therefore, this embodiment takes Pi value ≥ 0.015 as the threshold to screen infA , atpE-trnM-CAU , t rnE-UUC- trnT-GGU , clpP-trnG-UCC_1 , trnQ-UUG-rps16_1 , rpoC1_1-rpoB , psbN-psbT , rps8-infA , atpA-atpF_2 as specific variable regions.
[0028] Table 2. Nucleotide diversity Pi value analysis Gene Pi IGS Pi 0.02113 0.02818 0.01094 0.02267 0.00848 0.0206 0.00808 0.01679 0.00767 0.01617 0.00762 0.01607 0.0076 0.01598 0.00718 0.01572 0.00565 0.012 0.00503 0.01154 0.0049 0.00992 0.00482 0.00985 0.0048 0.00941 0.00477 0.00877 0.00456 0.00867 0.00449 0.0075 0.00431 0.00746 0.0043 0.00723 0.00417 0.00708 0.00416 0.00697 0.00375 0.00681 0.00334 0.00664 0.00325 0.00553 0.00325 0.00545 0.00289 0.00533 0.00287 0.00528 0.00267 0.00522 0.0026 0.00517 0.00249 0.00514 0.00248 0.00506 0.00243 0.00474 0.00228 0.00458 0.00218 0.00448 0.00217 0.0044 0.00214 0.00425 0.00208 0.00422 0.00191 0.00419 0.00186 0.00403 0.00177 0.00402 0.00171 0.00385 0.00168 0.00372 0.00156 0.00372 0.00154 0.00324 0.00148 0.00323 0.00146 0.00322 0.00136 0.00317 0.00134 0.00313 0.00133 0.00313 0.00132 0.00305 0.00113 0.0028 0.00107 0.00274 0.00102 0.00273 0.00097 0.0027 0.00095 0.00257 0.00094 0.00253 0.0009 0.00243 0.00081 0.00233 0.00068 0.00222 0.00047 0.00212 0.00046 0.00098 0.00041 0.00091 0.00027 0.00085 0.00025 0 0.00022 0 0 0 0 0 0 Further, in step S1-3, specific primers are designed according to the obtained specific variable regions infA , atpE-trnM-CAU , t rnE-UUC- trnT-GGU , clpP-trnG-UCC_1 , trnQ-UUG-rps16_1 , rpoC1_1-rpoB , psbN-psbT , rps8-infA , atpA-atpF_2 respectively, and the specific primers include sequences SEQ ID NO. 1-9. With the specific primers shown in SEQ ID NO. 1-9, a primer set for identifying Cactaceae plant species is composed.
[0029] S2. Extract DNA from the sample to be tested, and use the specific primers described above to amplify and sequence the specific mutation region of the sample DNA. The following 10 Opuntia species were selected as samples in this example, and the sample names were added with “_X” as an identifier for identification. Opuntia basilaris, O. ficus-indica, O. humifusa, O. pilifera, O. pycnantha, O. quimilo, O. sulphurea, O. scheerii, O. azurea and O. phaeacantha .
[0030] S3. Compare the sequencing results with the standard sequences in the chloroplast genome database to identify the species. In order to clearly distinguish the test samples from the reference sequences in the database, “_X” was added to the sample name of the test species. The identification results are as follows Figure 1 .
[0031] Among the 10 species analyzed, O. humifusa, O. azurea and O. ficus-indica The topological structure of the phylogenetic tree of three species has low support, indicating that there is a large uncertainty in the phylogenetic position. Overall, the identification success rate of the identification method constructed in this example is about 60% in these 10 species. Due to the low support rate of some species, it is speculated that it may be related to the assembly quality of the chloroplast genome data, which also suggests that the integrity and accuracy of the sample data need to be improved in subsequent research to improve the reliability of identification.
[0032] The above-described embodiments are merely preferred embodiments of the present application and do not limit the scope of the present application. Without departing from the design spirit of the present application, various modifications and improvements to the technical solutions of the present application made by those skilled in the art shall fall within the scope of protection of the present application.
Claims
1. A method for identifying cactus species based on high-throughput sequencing data, characterized in that: The following steps are involved: S1. Obtain a chloroplast genome database and specific variant regions of cacti, and design specific primers based on these regions. S2. Extracting the test sample DNA, using specific primers to amplify and sequence the specific variant region of the test sample DNA; S3. Compare the sequencing results with the standard sequences in the chloroplast genome database to identify the species.
2. The method according to claim 1, wherein The step S1 comprises: S1-1. Collect multiple cactus samples, extract DNA, sequence them at the genome level, assemble and annotate the corresponding chloroplast genomes, and construct a chloroplast genome database. S1-2. Perform multiple sequence alignment and nucleotide polymorphism analysis on the obtained chloroplast genomes to screen for specific variant regions. S1-3. Design specific primers based on the obtained specific variant regions.
3. The method according to any one of claims 1 or 2, characterized in that The screening method for the specific variable region comprises: Based on the chloroplast genome database, a tandem sequence alignment was performed on each gene and spacer region to construct a corresponding matrix; Based on the obtained matrix, the nucleotide diversity Pi values of the protein-coding gene CDS sequence and the intergenic region IGS sequence were calculated, and specific variable regions were screened based on the nucleotide diversity Pi values.
4. The method according to claim 3, wherein The nucleotide diversity Pi value of the specific variable region is not less than 0.
015.
5. The method according to claim 1, wherein The specific variable region includes infA 、 atpE-trnM-CAU , t rnE-UUC-trnT-GGU 、 clpP-trnG-UCC_1 、 trnQ-UUG-rps16_1 、 rpoC1_1-rpoB 、 psbN-psbT 、 rps8-infA 、 atpA-atpF_2 At least one of .
6. The method according to claim 1, wherein The specific primer includes at least one of the sequences shown in SEQ ID NO. 1-X.
7. A primer set for identifying species of cactus plants, characterized in that: The primer set targets a specific variable region, and the specific variable region includes infA 、 atpE-trnM-CAU , t rnE-UUC-trnT-GGU 、 clpP-trnG-UCC_1 、 trnQ-UUG-rps16_1 、 rpoC1_1-rpoB 、 psbN-psbT 、 rps8-infA 、 atpA-atpF_2 At least one of .
8. The primer set according to claim 7, wherein The primer set includes specific primers, and the specific primers include at least one of the sequences shown in SEQ ID NO. 1-X.
9. A kit for identifying species of cactus plants, characterized in that: The primer set comprising the primer set for identifying cactus plant species according to any one of claims 7 or 8.
10. Use of the primer set for identifying species of the genus Cactus according to any one of claims 7 or 8, or the kit for identifying species of the genus Cactus according to claim 9 in identifying species of the genus Cactus.