Agropyron cristatum variety identification method based on chloroplast genome high-throughput sequencing
By using high-throughput chloroplast genome sequencing technology, combined with phylogenetic analysis and specific marker screening, the taxonomic confusion in the identification of ice grass species has been resolved, enabling rapid and accurate species identification and improving the accuracy and reliability of identification.
Patent Information
- Application Number
- CN202511365919.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-24
- Publication Date
- 2025-10-31
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The identification of ice plant species is subject to taxonomic confusion, and traditional morphological identification methods are difficult to accurately distinguish between different species. The identification process is complicated by morphological variations and natural hybridization.
A high-throughput sequencing method based on chloroplast genomes was adopted, including sample collection, DNA extraction, library construction and sequencing, data processing, chloroplast genome assembly, genome annotation, phylogenetic analysis and variant identification. Combined with phylogenetic tree, principal component analysis and specific marker screening, accurate identification of ice plant species was achieved.
It enables rapid and accurate identification of ice grass species, improves the accuracy and reliability of identification, overcomes the interference of morphological variations and hybridization, and ensures the scientific nature and repeatability of experimental operations.
Smart Images

Figure CN120866503A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of plant taxonomy technology, specifically a method for identifying species of Icegrass based on high-throughput sequencing of chloroplast genomes. Background Technology
[0002] Wheatgrass (Triticum aestivum) is a perennial herb belonging to the genus Triticum in the family Poaceae. According to the latest taxonomic treatment, wheatgrass species are naturally distributed across Eurasia, and exist in diploid, tetraploid, or hexaploid forms. Despite varying ploidy levels, they all possess only the P genome. Their inflorescence axes are continuous and not separated at the nodes; each node has a spikelet. The spikelets are laterally compressed, containing 3 to 16 florets. The glumes are shorter than the florets and asymmetrically keeled. The asymmetrically keeled lemma has a ridge on its back and an acute to awned apex.
[0003] There are approximately 13 species of ice grass worldwide. The grasslands of Northwest China possess rich germplasm resources, including five species. Agropyron cristatum (L.) Gaertn., Agropyron desertorum (Fisch.) Schult., Agropyron michnoi Roshev. Agropyron sibiricum (Willd.) P. Beauv., Agropyron mongolicum Keng], four variants [ Agropyron cristatum var. pectinatum (M. Bieb.) Roshev. ex B. Fedtsch., Agropyron cristatum var. pluriflorum HL Yang, Agropyron desertorum var. pilosiusculum Melderis, Agropyron mongolicum var. villosum HL Yang] and a variant ( Agropyron sibiricum f.pubiflorum Roshev. is recorded in the Flora of China.
[0004] Icegrass plants possess strong resistance to cold, drought, and salinity, making them an important component of ecosystems in arid and semi-arid regions. In ecological restoration, they effectively stabilize soil, prevent erosion, and improve the ecological environment. As a valuable forage resource, they have high nutritional value and palatability, serving as excellent fodder for livestock such as cattle and sheep. Accurate identification of icegrass is fundamental to the effective collection, preservation, and utilization of its germplasm resources. However, due to extensive morphological variation, natural hybridization, and artificial propagation activities, this genus suffers from taxonomic confusion, complicating species identification. Summary of the Invention
[0005] To address the aforementioned technical problems, this invention provides a method for identifying *Agropyron cristatum* species based on high-throughput chloroplast genome sequencing, characterized by comprising the following steps: (a) Sample collection and DNA extraction: Healthy leaf samples of the genus *Agropyron* were collected, and total DNA was extracted using the CTAB method. (b) Library construction and sequencing: Total DNA was fragmented to 350bp, and DNA sequencing libraries were generated by fragment purification, end repair, 3'-A tailing and adapter ligation. 150bp paired-end sequencing was performed using the DNBSEQ-T7 platform to obtain raw sequencing data. (c) Data processing and chloroplast genome assembly: FastP was used to filter the raw sequencing data for quality, and GetOrganelle software was used to assemble the chloroplast genome de novo; (d) Genome annotation and feature analysis: The assembled chloroplast genome was annotated using CpGAVAS2; (e) Phylogenetic analysis: Multiple sequence alignment was performed using MAFFT, nucleotide diversity analysis was performed using DnaSP, and maximum likelihood phylogenetic trees were constructed using PhyloSuite and IQ-tree; (f) Variation identification and analysis: Based on the maximum likelihood phylogenetic tree constructed in step (e), species with clear evolutionary relationships were selected as reference genomes. SNPs and InDel variations were identified by comparison with the reference genomes. Principal component analysis was performed using PLINK, and species-specific markers were screened using vcfR. (g) Comprehensive judgment: Based on the phylogenetic tree topology, PCA clustering results, and the presence or absence of specific markers, the species of ice grass samples are identified.
[0006] Moreover, in step (a) sample collection and DNA extraction, the samples are collected from wild populations in northern China, with adjacent individuals spaced at least 100 meters apart.
[0007] Furthermore, in step (c) data processing and chloroplast genome assembly, the parameters for quality filtering are: quality threshold Q5, a maximum of 50 low-quality bases and 15 ambiguous bases (N) allowed per read, a minimum retained read length of 150 bp, and a maximum allowed number of differences in the overlapping region of paired reads of 1.
[0008] Furthermore, in step (e) of the comparison and phylogenetic analysis, nucleotide diversity was calculated using DnaSP software with a window size of 400 bp and a step size of 200 bp.
[0009] Furthermore, in step (e) of the comparison and phylogenetic analysis, rice ( Oryza sativa The L.) chloroplast genome was used as the outgroup. The model used to construct the maximum likelihood tree was GTR+F+I+R9 or TVM+F+I+R4. The branch support rate was calculated by the UltraFast Bootstrap algorithm and repeated 1000 times.
[0010] Moreover, in step (f) variation identification analysis, the screening criteria for species-specific markers are: fixed presence in the target species (frequency f=1.0), complete absence in the control species (f=0.0), and verification in at least 3 independent phylogenetic lineages.
[0011] Beneficial effects: 1. Due to the wide range of morphological variations, natural hybridization, and artificial propagation activities in the genus *Agropyron*, taxonomy is confused, and traditional morphological identification methods are difficult to accurately distinguish between different species. This invention utilizes high-throughput chloroplast genome sequencing technology to identify *Agropyron* species at the molecular level. It can quickly and accurately identify the species of *Agropyron* samples without being affected by morphological variations and hybridization, effectively solving the limitations of traditional identification methods and improving the accuracy of identification.
[0012] 2. This invention constructs a phylogenetic tree, performs principal component analysis (PCA), and screens species-specific markers. It then uses a combination of phylogenetic tree topology, PCA clustering results, and the presence or absence of specific markers to determine the species of *Agropyron cristatum*. This multi-dimensional analysis method can more comprehensively and accurately reflect the genetic relationships and differences among *Agropyron cristatum* species, greatly improving the reliability of species identification.
[0013] 3. This invention details the entire operational process from sample collection and DNA extraction, library construction and sequencing, data processing and chloroplast genome assembly, to genome annotation and feature analysis, comparative and phylogenetic analysis, variant identification analysis, and comprehensive judgment. It also clearly defines the key parameters for each step, ensuring the scientific and standardized nature of the experimental operation, enabling different laboratories to replicate experimental results, and improving the credibility of the research.
[0014] 4. In the process of data processing and analysis, a variety of advanced bioinformatics tools and models were used, such as fastp, GetOrganelle, CpGAVAS2, MAFFT, DnaSP, PhyloSuite, IQ-tree, PLINK, vcfR, etc. These tools and models have high accuracy and reliability in their respective fields and can provide strong technical support for the identification of ice grass species. Attached Figure Description
[0015] Figure 1Five species and two varieties of the genus *Agrostis* ( Agropyron Morphological characteristics of the plant in its natural habitat, including details of the whole plant and spikelets, AC: rhizome ice plant ( A. michnoi DE: Ice Grass ( A. cristatum FG: Desert Ice Plant ( A. desertorum ); HI: Hairy Sand Ice Plant ( A. desertorum var. pilosiusculum JL: Sand purslane ( A. mongolicum ); MN: *Phragmites australis* ( A. mongolicum var. villosum OQ: Siberian ice grass ( A. sibiricum ); Figure 2 A chloroplast genome map of the genus *Agropyron*; Figure 3 Sliding window analysis of chloroplast genome nucleotide diversity in 48 *Agropyron* samples; Figure 4 A maximum likelihood phylogenetic tree of the genus *Agropyron* constructed based on the complete chloroplast genome; Figure 5 A maximum likelihood phylogenetic tree of the genus *Agropyron* constructed based on the chloroplast hypervariable region; Figure 6 Principal component analysis of genetic structure based on SNPs in the chloroplast genome of 48 *Agropyron* samples. Detailed Implementation
[0016] Example 1: Sample Collection and Identification From July to October 2024, 48 samples were collected from wild populations in seven regions of northern China (Xinjiang, Qinghai, Gansu, Inner Mongolia, Ningxia, Hebei, and Liaoning), representing five species and two varieties of the genus *Agrostis*. The specific species are as follows: A. cristatum , A. desertorum , A. michnoi , A. sibiricum , A. mongolicum , A. cristatum var. pectinatum and A. desertorum var. pilosiusculumAll sampling activities were conducted only outside protected areas (near nature reserves). The sampling strategy ensured the acquisition of genetically distinct individuals, obtaining single samples from healthy leaves of different plants, and maintaining a distance of at least 100 meters between adjacent individuals to avoid population redundancy. For each sample, a voucher sample was carefully collected and preserved at the Herbarium of the Mengcao Seed Industry Center (Hohhot, China), providing a verifiable taxonomic record. Formal taxonomic identification of all samples was conducted by Dr. Zhang Zhongshuai (State Key Laboratory of Forest Tree Genetics and Breeding, Chinese Academy of Forestry), primarily referencing the Flora of China. Morphological identification of the Flora of China confirmed that all samples matched their target taxa. Figure 1 Representative images of key diagnostic features (such as spikelet morphology and leaf characteristics) for each taxonomic unit are displayed, ensuring visual confirmation of correct classification. Sample collection information is shown in Table 1.
[0017] Table 1 Sample Collection Records and Geographic Information Sample number Latin name abbreviation nation City Collection location code longitude latitude 2407070642 AC0642 China Ulanqab NM-WLCB 112.61 41.03 2409231134 AC1134 China Haiyan QH-HY 100.64 37.13 2024071538 AC1538 China Alxa Left Banner NM-ALSZQ 105.72 38.85 202410011734 AC1734 China Urumqi XJ-WLMQ 87.21 43.33 2407061910 AC1910 China wuchuan NM-WC 111.80 41.21 2407061955 AC1955 China wuchuan NM-WC 111.71 41.20 2024102041 AC2041 China Zhangye GS-ZY 99.60 38.91 2024102065 AC2065 China Zhangye GS-ZY 99.60 38.91 2024102169 AC2169 China Mulei Kazakh Autonomous County XJ-CJ 90.47 43.82 2024071502 var. AP1502 China Alxa Left Banner NM-ALSZQ 105.72 38.85 2024102039 var. AP2039 China Gongliu XJ-YL 82.91 43.23 2024102040 var. AP2040 China Gongliu XJ-YL 82.91 43.23 2024080202 AS0202 China Hulunbuir NM-HLE 120.01 49.30 202410206 AS0206 China Chifeng NM-CF 117.89 43.30 202410209 AS0209 China Mulei Kazakh Autonomous County XJ-CJ 90.47 43.82 2024080517 AS0517 China Xilin Gol League NM-MD 116.34 44.27 2407221058 AS1058 China Chifeng NM-WNTQ 119.34 42.94 2407201318 AS1318 China Chengde HB-CD 117.48 42.34 2407131916 AS1916 China Urad Front Banner NM-WLTQQ 109.83 40.82 2407061956 AS1956 China wuchuan NM-WC 111.71 41.20 2024102020 AS2020 China Wuyuan NM-WY 108.33 41.12 2024102038 AS2038 China Wuyuan NM-WY 108.33 41.12 2024102060 AS2060 China Mulei Kazakh Autonomous County XJ-CJ 90.47 43.82 2024102062 AS2062 China Chifeng NM-CF 117.89 43.30 202407275 AS7275 China Hulunbuir NM-BDG 119.31 49.88 202410205 var. AM0205 China Chifeng NM-CF 117.85 43.27 2024080507 var. AM0507 China Xilin Gol League NM-HTL 116.21 43.38 2407060959 var. AM0959 China Hohhot NM-HHHT 111.84 40.97 2407061229 var. AM1229 China Hohhot NM-HHHT 111.84 40.97 2407201312 var. AM1312 China Chengde HB-CD 117.48 42.34 2407201324 var. AM1324 China Chengde City HB-CD 117.48 42.34 2407211337 var. AM1337 China Keshiketeng Banner NM-CF 117.83 43.26 2407211812 var. AM1812 China Chifeng NM-WNTQ 119.07 43.04 2024102011 var. AM2011 China Xilin Gol League NM-DL 116.44 42.21 2408030712 AR0712 China Fuxin LN-FX 122.22 42.70 2408031010 AR1010 China Tongliao NM-TL 122.22 42.74 202410203 AK0203 China Xilin Gol League NM-DL 116.44 42.21 2024080506 AK0506 China Xilin Gol League NM-HTL 116.21 43.38 2407130937 AK0937 China Ejin Horo Banner NM-YJHL 109.69 39.56 2407130945 AK0945 China Ejin Horo Banner NM-YJHL 109.69 39.56 2408181557 AK1557 China Baotou NM-GY 110.09 40.73 2024102010 AK2010 China Xilin Gol League NM-DL 116.44 42.21 2024102018 AK2018 China Xilin Gol League NM-DL 116.44 42.21 2024102019 AK2019 China Xilin Gol League NM-DL 116.44 42.21 2024102024 AK2024 China Xilin Gol League NM-DL 116.44 42.21 2024072906 AK2906 China Yinchuan NX-YC 106.30 38.46 2407131150 AB1150 China Ordos NM-YJHL 109.65 39.59 2407131152 AB1152 China Ordos NM-YJHL 109.65 39.59
[0018] Example 2 Sampling and DNA Extraction Total DNA was extracted from leaves using a modified cetyltrimethylammonium bromide (CTAB) method. After successful DNA quality assessment, genomic DNA was lysed into ~350 bp fragments using a Covaris lysis instrument. Sequencing libraries were generated through fragment purification, end repair, 3'-A tailing, and adapter ligation. Paired-end sequencing (150 bp read length) was performed on the DNBSEQ-T7 platform using next-generation high-throughput sequencing.
[0019] Example 3: Chloroplast Genome Assembly and Annotation Raw data was filtered using fastp (v0.20.0). The base quality filtering threshold was set to Q5, allowing a maximum of 50 low-quality bases and 15 ambiguous bases (N) per read, with a minimum retained read length of 150 bp. For paired-end data merging, the maximum cardinality difference in overlapping regions was set to 1 (≤10% difference percentage). Chloroplast genome assembly was performed using GetOrganille (v1.7.7.1), with at least 10 reads aligned to the reference genome. BLAST+ (v2.16.0) (https: / / blast.ncbi.nlm.nih.gov / Blast.cgi) was used for alternative annotation of the assembled sequences. The assembled genome was annotated using the CpGAVAS2 online tool. Accession numbers for the annotated chloroplast genomes were available in the NCBIGenBank database (Table 2).
[0020] The chloroplast genomes of 48 individuals from five species and two varieties of *Agropyron cristatum* showed high conservation, exhibiting a typical four-part structure, including a large single-copy region (LSC), a small single-copy region (SSC), and two inverted repeat regions (IR). Figure 2 Genes transcribed clockwise are located outside the outer circle, and genes transcribed counterclockwise are located inside the inner circle. The outermost circle displays the gene map, with genes distinguished by color according to their functional categories. The inner circle shows the genomic regions: LSC (large single copy region), SSC (small single copy region), and two IR regions: IRA and IRB. The black ring in the innermost circle indicates GC content. The genome size range and GC content are labeled at the center of the map. The total length of the sequenced chloroplast genome is 135321-135564 bp, with a length variation of 0.18%. The LSC region is 79590-79707 bp, the SSC region is 12634-12782 bp, and the IR region is 42966-43096 bp. The genes in the 48 chloroplast genomes are highly homogeneous, with each genome containing 132 genes: 84 protein-coding genes, 40 tRNA genes, and 8 rRNA genes, indicating no significant gene loss or gain. The GC content of all samples was 38%, reflecting the evolutionary stability of the chloroplast genome base composition. Notably, one sample from Xinjiang... A. cristatum The SSC region of sample (AC1734) is significantly shorter than that of other samples (12634 bp vs. 12761±19 bp).
[0021] Table 2. Species sample information and GenBank accession list for the genus *Agropyron*. Species name abbreviation GenBank (L.) Gaertn.0642 AC0642 PV030929 (L.) Gaertn. 1134 AC1134 PV030931 (L.) Gaertn. 1538 AC1538 PV030933 (L.) Gaertn. 1734 AC1734 PV030934 (L.) Gaertn. 1910 AC1910 PV030936 (L.) Gaertn. 1955 AC1955 PV030937 (L.) Gaertn.2041 AC2041 PV030938 (L.) Gaertn.2065 AC2065 PV030940 (L.) Gaertn.2169 AC2169 PV030954 var.(M. Bieb.) Roshev. ex B. Fedtsch.1502 AP1502 PV030941 var.(M. Bieb.) Roshev. ex B. Fedtsch.2039 AP2039 PV030955 var.(M. Bieb.) Roshev. ex B. Fedtsch.2040 AP2040 PV030956 (Fisch.) Schult. 0202 AS0202 PV030926 (Fisch.) Schult. 0206 AS0206 PV020957 (Fisch.) Schult. 0209 AS0209 PV030927 (Fisch.) Schult. 0517 AS0517 PV030942 (Fisch.) Schult. 1058 AS1058 PV030943 (Fisch.) Schult. 1318 AS1318 PV030944 (Fisch.) Schult. 1916 AS1916 PV030945 (Fisch.) Schult. 1956 AS1956 PV030952 (Fisch.)Schult. 2020 AS2020 PV030946 (Fisch.)Schult. 2038 AS2038 PV030947 (Fisch.) Schult. 2060 AS2060 PV030939 (Fisch.) Schult. 2062 AS2062 PV020958 (Fisch.) Schult. 7275 AS7275 PV030948 var.Melderis0205 AM0205 PQ845757 var.Melderis 0507 AM0507 PV030928 var.Melderis 0959 AM0959 PV030949 var.Melderis 1229 AM1229 PV030950 var.Melderis 1312 AM1312 PV030932 var.Melderis 1324 AM1324 PV030953 var.Melderis 1337 AM1337 PV030951 var.Melderis 1812 AM1812 PV030935 var.Melderis 2011 AM2011 PV020956 Roshev0712 AR0712 PV030930 Roshev1010 AR1010 PV021113 Keng0203 AK0203 PV021110 Keng0506 AK0506 PV021111 Keng0937 AK0937 PV021112 Keng 0945 AK0945 PV020955 Keng 1557 AK1557 PV021116 2010 AK2010 PV021117 Keng 2018 AK2018 PV021118 Keng 2019 AK2019 PV021119 Keng 2024 AK2024 PV021120 Keng 2906 AK2906 PV020959 (Willd.) P. Beauv. 1150 AB1150 PV021114 (Willd.) P. Beauv. 1152 AB1152 PV021115
[0022] Example 4: Genomic Comparison and Nucleotide Polymorphism Analysis The chloroplast genome was aligned with MAFFT (v7.526), and nucleotide diversity (Pi) was analyzed by DnaSP (v6.12.03) with a sliding window of 400 bp and a step size of 200 bp.
[0023] The Pi value (nucleotide diversity) represents the average number of nucleotide differences between any two randomly selected DNA sequences in a population, and is a key indicator for measuring DNA sequence polymorphism within or between species. A high Pi value indicates abundant nucleotide variation in a genomic region, which may be related to increased mutation rates, selection pressure, or recombination events; a low Pi value indicates a more conserved sequence with low nucleotide diversity.
[0024] Nucleotide diversity in the chloroplast genome of *Agropyron cristatum* is unevenly distributed, with distinct peaks at specific locations (e.g., approximately 56444 bp and 105329 bp), reflecting high nucleotide variation at these sites, which may be related to functional adaptation or evolutionary pressure. For example, psbA , rbcL-psaI , petA-psbJ and psbE-petL Regions near isogens showed relatively high Pi values (Pi > 0.002). The SSC region also contained sites with elevated Pi values (Pi > 0.002), such as... rpl32-trnL and rpl32 area( Figure 3 The X-axis represents the midpoint position (bp) of each sliding window, and the Y-axis displays the nucleotide diversity value of the corresponding window.
[0025] Example 5: Phylogeny and Nucleotide Diversity of *Agropyron cristatum* For phylogenetic analysis, 48 *Agropyron cristatum* chloroplast genomes and homologous sequences of relevant species were analyzed from NCBI, including... Phyllostachys edulis (Carrière) J. Houz., Giant field grass Roth, Oats green L., Brachypodium sylvaticum (Huds.) P. Beauv., Perennial ryegrass L., Honeybee rough Trin., Phalaris reed L., Meadow grass L., Hairy Stipa L., Aegilops tauschii Coss., Siberian Elymus L., Wheatgrass (Gaertner)Nevski, Barley L., Leymus secalinus (Georgi) Tzvelev, Psathyrostachys rush (Fisch.) Nevski, Rye cereal L., Summer wheat L is used. Rice green L is an outgroup.
[0026] Maximum likelihood (ML) trees were constructed using PhyloSuite and IQ-tree, applying the GTR+F+I+R9 model. Branch support was evaluated using UltraFast guidance (1000 replicates, maximum titer = 1000, minimum cor.coef = 0.90). After manually extracting hypervariable regions, a second ML tree was constructed using PhyloSuite with TVM+F+I+R4 in IQ-tree, with identical parameters. MEGA was used to calculate phylogenetic information sites for the entire chloroplast genome and hypervariable regions of *Ceratophyllum demersum*.
[0027] A maximum likelihood phylogenetic tree based on complete chloroplast genomes showed that 48 samples from 5 species and 2 varieties formed a strongly supported monophyletic group (bootstrap=100). While the monophyletic nature of the genus *Agropyron* was highly supported, phylogenetic relationships within the genus remained unresolved. A notable feature was the small difference between all *Agropyron* chloroplast genomes, reflected in short internal branching in the phylogenetic analysis. Figure 4 The numbers on the branches represent bootstrapping support rates (see Table 3 for information on 18 relevant species from NCBI). A phylogenetic tree based on high-variable regions is shown below. Figure 5 As shown (the numbers above the branches represent the bootstrap support rate; information on 18 related species from NCBI is shown in Table 3).
[0028] Table 3 Information on 18 relevant species from NCBI sources used for phylogenetic analysis scientific name GenBank registration number (Career) J. Houz. NC_051817 L. NC_027468 (Huds.) P. Beauv. NC_058601 L. NC_009950 Trin. NC_050212 L. NC_027481 L. NC_057962 L. OP681436 Coss. NC_022133 L. NC_058919 (Gaertner) Nevsky NC_059978 L. NC_008590 (Georgi) Tzvelev MZ595321 (Fisch.) Nevski NC_043838 L. NC_021761 L. NC_002762 (G. Forst.) R. Br. OM307679 (Hack.) Oh! OM683284 L. (outgroup) NC_031333 Within the genus *Agropyron*, only 178 out of 135,978 valid loci in the chloroplast genome (0.13%) contain phylogenetic information. In the hypervariable region (out of a total of 14,992 valid loci), only 68 loci contain phylogenetic information, representing 0.45% of these regions. None of these five species form monophyletic clades; instead, they are nested within samples from other species. These two varieties— A. cristatum var. scallop and A. desert var. hairy They also failed to form independent monophyletic groups. Instead, within the genus *Agrostis*, geographically close specimens typically form well-supported clades: for example, three specimens from Duolun County, Inner Mongolia... A. mongolica Individuals; two in Barkol, Xinjiang A. desert Individual plants; two plants in Keshiketeng Banner, Inner Mongolia A. desert and a plant A. desert var. hairy .
[0029] Example 6 Principal Component Analysis of SNPs Genetic variations, including single nucleotide polymorphisms (SNPs) and insertions / deletions, were systematically identified using reference-guided scanning (AS0202). Position-specific SNP detection, gap-based insertion / deletion identification, and contiguous variant merging were employed. The integrated variant set was formatted according to the VCF specification and processed using BCFtools (v1.14). Population structure analysis was performed using PCA in PLINK (v1.9), and the results were visualized using ggplot2 (v3.4.0) in R (v4.2.1). Species diagnostic markers were identified using vcfR (v1.13.0) under stringent criteria: 1) target species fixed (f=1.0), 2) control species absent (f=0.0), and 3) validated in ≥3 phylogenetic lineages.
[0030] Based on chloroplast genome variations from 48 *Agropyron cristatum* accessions, 439 polymorphic sites were identified, including 364 single nucleotide polymorphisms (SNPs) and 75 insertion / deletion variants (InDels). No species-specific SNPs or insertions / deletions were found. Principal component analysis (PCA) was performed using only SNP variants. The first two principal components (PC1 and PC2) explained 7.7% and 7.2% of the total genetic variance, respectively. However, *Agropyron cristatum* species did not separate on the first and second principal components but showed extensive overlap, such as... Figure 6 As shown, principal component 1 (PC1) and principal component 2 (PC2) explained 7.7% and 7.2% of the genetic variation, respectively. Sample points are colored by taxonomic unit, and ellipses represent the clustering range of each population's 95% confidence interval. Abbreviations: AC: Icewort ( A. cristatum AP: *Agropyron cristatum* ( A. cristatum var. scallop AS: Sand plant ( A. desert AM: *Agropyron cristatum* (a type of barnyard grass) A. desert var. hairy AR: Rhizome Ice Plant ( A. michnoi ), AK: Mongolian Ice Grass ( A. mongolica ), AB: Siberian ice grass ( A. sibiricum ).
Claims
1. A method for identifying *Agropyron cristatum* species based on high-throughput chloroplast genome sequencing, characterized in that, Includes the following steps: (a) Sample collection and DNA extraction: Healthy leaf samples of the genus *Agropyron* were collected, and total DNA was extracted using the CTAB method. (b) Library construction and sequencing: Total DNA was fragmented to 350bp, and DNA sequencing libraries were generated by fragment purification, end repair, 3'-A tailing and adapter ligation. 150bp paired-end sequencing was performed using the DNBSEQ-T7 platform to obtain raw sequencing data. (c) Data processing and chloroplast genome assembly: FastP was used to filter the raw sequencing data for quality, and GetOrganelle software was used to assemble the chloroplast genome de novo; (d) Genome annotation and feature analysis: The assembled chloroplast genome was annotated using CpGAVAS2; (e) Phylogenetic analysis: Multiple sequence alignment was performed using MAFFT, nucleotide diversity analysis was performed using DnaSP, and maximum likelihood phylogenetic trees were constructed using PhyloSuite and IQ-tree; (f) Variation identification and analysis: Based on the maximum likelihood phylogenetic tree constructed in step (e), species with clear evolutionary relationships were selected as reference genomes. SNPs and InDel variations were identified by comparison with the reference genomes. Principal component analysis was performed using PLINK, and specific markers were screened using vcfR. (g) Comprehensive determination: Based on the topological structure of the maximum likelihood phylogenetic tree constructed in step (e), the principal component analysis clustering results in step (f), and the results of species-specific markers, the species of ice grass samples are identified.
2. The method for identifying *Agropyron cristatum* species based on high-throughput chloroplast genome sequencing as described in claim 1, characterized in that, In step (a) sample collection and DNA extraction, the samples are collected from wild populations in northern China, with adjacent individuals spaced at least 100 meters apart.
3. The method for identifying *Agropyron cristatum* species based on high-throughput chloroplast genome sequencing as described in claim 1, characterized in that... In step (c) data processing and chloroplast genome assembly, the parameters for quality filtering are: quality threshold Q5, a maximum of 50 low-quality bases and 15 ambiguous bases allowed per read, a minimum retained read length of 150 bp, and a maximum allowed number of differences in the overlapping region of paired reads of 1.
4. The method for identifying *Agropyron cristatum* species based on high-throughput chloroplast genome sequencing as described in claim 1, characterized in that... In step (e) of the comparison and phylogenetic analysis, the nucleotide diversity analysis uses a window size of 400 bp and a step size of 200 bp.
5. The method for identifying *Agropyron cristatum* species based on high-throughput chloroplast genome sequencing as described in claim 1, characterized in that... In step (e) of the comparison and phylogenetic analysis, the rice chloroplast genome was used as the outgroup, and the model used to construct the maximum likelihood tree was GTR+F+I+R9 or TVM+F+I+R4. The branch support rate was calculated by the UltraFast Bootstrap algorithm and repeated 1000 times.
6. The method for identifying *Agropyron cristatum* species based on high-throughput chloroplast genome sequencing as described in claim 1, characterized in that, In step (f) variation identification analysis, the screening criteria for species-specific markers are: fixed presence in the target species, complete absence in the control species, and validation in at least three independent phylogenetic lineages.
Citation Information
Patent Citations
Method for rapidly identifying variety of vector mixed stump
CN112143818A
SSR primer of endangered species thuja sutchuenensis and application of SSR primer
CN112226531A
Camellia oleifera variety identification method
CN117457075A
SNP (Single Nucleotide Polymorphism) site special for maternal line source of Mongolian agropyron cristatum, molecular marker and application of SNP site
CN117987586A
Yuling hickory chloroplast genome and application thereof in germplasm identification
CN118360425A