Salvia chinensis chloroplast genome and application thereof
By providing the complete chloroplast genome sequence of Sagebrush and its application, and using high-throughput sequencing and phylogenetic analysis, the problem of difficulty in identifying closely related species of Sagebrush was solved, efficient and accurate species identification and phylogenetic research were achieved, and the accuracy and reliability of plant classification were improved.
Patent Information
- Application Number
- CN202510868949.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-10-10
AI Technical Summary
Existing technologies make it difficult to accurately identify Sage sinensis and its closely related species. The length of traditional DNA barcode sequences is limited, and the accuracy is insufficient, especially in identifying closely related species or varieties, which affects the reliability of plant classification and phylogenetic research.
This paper provides the complete chloroplast genome sequence of Salvia sinensis and its applications. The chloroplast genome is obtained through high-throughput sequencing and splicing assembly. Combining sequence alignment and phylogenetic analysis, a phylogenetic tree is constructed to achieve efficient and accurate species identification and phylogenetic relationship analysis.
It significantly improves the accuracy of species identification of Sage sinensis and the reliability of phylogenetic research, can clearly distinguish the branches of closely related species, solves the problem of identification difficulties in traditional methods, and is suitable for variety identification and kinship analysis of other plants.
Smart Images

Figure CN120758515A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of gene technology, and more particularly to a chloroplast genome of Salvia sinensis and an application thereof. Background Art
[0002] Chinese sage (Salvia chinensis) is an annual or perennial herb in the genus Salvia, Lamiaceae. It is widely distributed in East China, Hubei, Sichuan, Guangxi, Guangdong, and Hunan. Also known as Shijianchuan, Chinese sage has the properties of clearing heat and detoxifying, promoting blood circulation, and relieving pain. Pharmacological experiments have shown that the main chemical components of Chinese sage have anti-tumor activity (Zeng Juan, 2014). Salvia species are numerous and share considerable morphological similarities. Traditional classification and identification rely primarily on morphological characteristics, which can lead to identification errors due to differences in environment and developmental stage. Furthermore, existing molecular identification techniques typically use ribosomal ITS sequences or partial chloroplast genome fragments (such as rbcL and matK) as DNA barcodes. However, these conventional DNA barcodes are limited in length and lack the ability to distinguish closely related species, making it particularly difficult to accurately distinguish closely related species or varieties.
[0003] The chloroplast genome is relatively independent of the nuclear genome, inherited maternally, and characterized by conserved structure and abundant single-copy sequences. A complete chloroplast genome (cpDNA) is typically 150,000–160,000 base pairs in length, containing a typical circular tetrad structure consisting of large and small single-copy regions (LSC and SSC) and a pair of inverted repeats (IRs). Its genomic composition is relatively stable. Compared to DNA barcodes derived from single gene fragments, whole chloroplast genomes encompass more genomic and non-coding region variation, providing a large number of sites of interspecies variation and offering unique advantages in species identification. It has been reported that whole chloroplast genome sequences can be used as "super DNA barcodes" for plant species identification, significantly improving the accuracy and reliability of identification of closely related species (Yuan et al., 2023; Li et al., 2015). Incorporating chloroplast genomes into species identification and phylogenetic studies can overcome the difficulties encountered by traditional methods in identifying closely related species, providing new technical tools for plant classification, germplasm conservation, and genetic diversity research.
[0004] However, the complete chloroplast genome sequence of S. sinensis has not yet been published, and genetic databases lack chloroplast genome information for this species. This has, to a certain extent, limited the accurate identification of S. sinensis and its related species and the study of their phylogenetic relationships.
[0005] Therefore, there is an urgent need to determine the complete chloroplast genome sequence of Salvia sinensis and apply it to species identification and phylogenetic analysis to fill the gaps in existing technology and improve the accuracy of molecular identification of Salvia plants in the Lamiaceae family and the reliability of phylogenetic studies. Summary of the Invention
[0006] In view of this, the present invention provides a chloroplast genome of Salvia sinensis and applications thereof.
[0007] In order to achieve the above object, the present invention adopts the following technical solutions:
[0008] The chloroplast genome of Sagebrush sphenanthera, the nucleotide sequence of the chloroplast genome of Sagebrush sphenanthera is a full-length sequence connected in sequence by sequences such as SEQ ID NO.1, SEQ ID NO.2, SEQ ID NO.3, SEQ ID NO.4, SEQ ID NO.5, SEQ ID NO.6, SEQID NO.7, SEQ ID NO.8, SEQ ID NO.9, SEQ ID NO.10, SEQ ID NO.11, SEQ ID NO.12, SEQ ID NO.13, SEQ ID NO.14, SEQ ID NO.15 and SEQ ID NO.16.
[0009] Another object of the present invention is to provide an application of the above-mentioned Salvia sinensis chloroplast genome, wherein the application is any one of the following:
[0010] (1) Identification of Salvia sinensis varieties;
[0011] (2) Application in reconstructing the phylogenetic tree of Salvia plants.
[0012] Another object of the present invention is to provide a method for identifying a species of Salvia sinensis, comprising the following steps:
[0013] (1) The total DNA of the plant sample to be identified was extracted using the CTAB method;
[0014] (2) Perform high-throughput paired-end sequencing on the extracted DNA to obtain sequencing data covering the chloroplast genome;
[0015] (3) Splicing and assembling the sequencing data to obtain the chloroplast genome sequence of the sample to be tested;
[0016] (4) performing sequence alignment on the assembled chloroplast genome sequence of the sample to be tested and the chloroplast genome sequence of S. sinensis (the full-length sequence connected in sequence by sequences such as SEQ ID NO.1-SEQ ID NO.16). When the similarity index reaches 99.9% or more, the sample to be tested is determined to be S. sinensis.
[0017] Preferably, the splicing and assembling method in step (3) specifically includes:
[0018] a. De novo assembly of high-throughput sequencing data was performed using SPAdes 3.14.0 software to generate preliminary assembly fragments, with the k-mer parameter set to 95;
[0019] b. Use Geneious version 2025.1.1 software to confirm the sequence and connection relationship of each fragment, complete the assembly of the chloroplast genome, and manually proofread the sequence;
[0020] c. Use CPGAVAS2 software and Sequin tools to perform functional annotation and sequence verification on the assembled chloroplast genome to obtain the complete chloroplast genome sequence of the sample to be tested, and use CPGAVAS2 online software to draw an annotation map of the chloroplast genome.
[0021] Another object of the present invention is to provide a method for reconstructing a phylogenetic tree of the genus Salvia, using the above-mentioned chloroplast genome sequence of S. sinensis and the chloroplast genome sequences of other Salvia species to perform phylogenetic analysis, comprising the following steps:
[0022] (1) MAFFT software was used to perform a multiple sequence alignment of the chloroplast genome sequences of S. sinensis and other Salvia species;
[0023] (2) ModelFinder software was used to construct a phylogenetic tree based on the results of multiple sequence alignment, and a bootstrap test was performed to evaluate the reliability of the branches, thereby obtaining a phylogenetic relationship diagram of the genus Salvia.
[0024] Beneficial Effects: This study provides the first complete chloroplast genome sequence of S. sinensis, filling a gap in chloroplast genome data for this species in gene banks. This sequence can be used as a specific molecular marker for S. sinensis identification, significantly improving the accuracy and reliability of species identification.
[0025] The complete chloroplast genome sequence of S. sinensis was applied to the phylogenetic analysis of the Salvia genus, clarifying its phylogenetic position within the genus. The results showed that the phylogenetic tree constructed using the complete chloroplast genome has a higher resolution, clearly distinguishing closely related species. This is of great significance for the taxonomic identification and genetic evolutionary relationship determination of the Salvia genus.
[0026] The identification method provided by this invention features a simple, efficient, and highly versatile process. By obtaining the complete chloroplast genome through a single high-throughput sequencing run and performing sequence alignment, it can rapidly identify whether a sample is Salvia sinensis, avoiding the cumbersome experimental steps and subjective judgments required by traditional methods. This method is also applicable to the identification of other plant species and phylogenetic analysis, and has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0028] Figure 1 The accompanying figure is a circular map of the chloroplast genome of Sage sinensis of the present invention; the positions of the large single copy region (LSC), small single copy region (SSC) and inverted repeat region (IR) as well as the distribution of each functional gene are marked therein, the direction of the arrow indicates the transcription direction of the gene, and the inner circle shows the GC content distribution.
[0029] Figure 2 The attached figure is a phylogenetic tree of the genus Salvia based on the complete chloroplast genome sequence.
[0030] Figure 3 The attached figure is a fragment diagram of the comparison results of the chloroplast genome sequences of the test samples A and B with the reference sequence.
[0031] Figure 4 The attached figure shows the phylogenetic relationship between the test sample A and the test sample B and the species of the genus Salvia. DETAILED DESCRIPTION
[0032] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0033] The embodiments of the present invention disclose a chloroplast genome of Salvia sinensis and its application. The reagents used in the present invention, unless otherwise specified, are all commercially available and their sources are not specifically limited. The methods involved, unless otherwise specified, are all conventional methods and will not be described in detail here.
[0034] Example 1 Obtaining the Chloroplast Genome of Salvia sinensis
[0035] 1. Collection of Sagebrush and DNA extraction: The present invention adopts the CTAB method to extract total DNA. Several fresh young leaves of Sagebrush are quickly frozen with liquid nitrogen and ground into fine powder. The ground tissue powder is placed in a pre-cooled extraction buffer (containing CTAB, PVP and other ingredients to remove polyphenol and polysaccharide impurities), incubated in a 65°C water bath for 30 minutes, and gently shaken to mix. Subsequently, the DNA precipitate is recovered by chloroform-isoamyl alcohol extraction and isopropanol precipitation. The DNA is washed with 70% ethanol and dissolved in TE buffer. The quality and concentration of the DNA are detected using an ultramicro spectrophotometer and agarose gel electrophoresis, and then stored in a -20°C refrigerator for use.
[0036] 2. High-throughput sequencing library construction: Take the qualified DNA samples mentioned above and prepare the library according to the high-throughput sequencing library construction protocol. Fragment the genomic DNA into approximately 350bp fragments, connect sequencing adapters, amplify and enrich the target fragments by PCR, and construct a double-end sequencing library. The library is sequenced on a high-throughput sequencing platform to obtain 2×150bp PE (paired-end) sequencing read length data. Ensure that the data volume generated for each individual is approximately 5-10Gb.
[0037] 3. Chloroplast Genome Assembly and Gene Annotation: High-quality sequence reads were input into SPAdes (version 3.14.0) for de novo assembly. Various k-mer parameters were tried during the assembly process to optimize the results; in this example, k = 95 was used for assembly. Next, Geneious (version 2025.1.1) was used to confirm the sequence and connectivity of each fragment, complete the assembly of the chloroplast genome, and manually proofread the sequence for errors. Gene annotation and sequence verification were performed on the assembled S. sinensis chloroplast genome using CPGAVAS2 and Sequin software.
[0038] The results showed that the chloroplast genome of S. sinensis is 151,458 bp in size (the full-length sequence consisting of SEQ ID NO.1, SEQ ID NO.2, SEQ ID NO.3, SEQ ID NO.4, SEQ ID NO.5, SEQ ID NO.6, SEQ ID NO.7, SEQ ID NO.8, SEQ ID NO.9, SEQ ID NO.10, SEQ ID NO.11, SEQ ID NO.12, SEQ ID NO.13, SEQ ID NO.14, SEQ ID NO.15, and SEQ ID NO.16), with a total GC content of 38.02%. It has a circular tetrad structure and contains a large single copy region (82,768 bp), a small single copy region (17,576 bp), and a pair of inverted repeats (25,557 bp) ( Figure 1 The genome encodes 114 unique genes (Table 1), including 80 protein-coding genes, 30 transfer RNA genes, and 4 ribosomal RNA genes. The circular map of the chloroplast genome of S. Figure 1 shown.
[0039] Table 1 Annotated gene information of the chloroplast genome of S. sinensis
[0040]
[0041]
[0042] Example 2 Phylogenetic Analysis Application
[0043] The assembled chloroplast genome sequence of S. sinensis was aligned with the chloroplast genomes of 65 Salvia species downloaded from the NCBI website (Table 2) using the multiple sequence alignment software MAFFT (version 7.311) [with the complete chloroplast genome of Mentha longifolia (GenBank ID: NC032054) as the outgroup].
[0044] Table 2 Chloroplast genome sequences of 65 Salvia species
[0045]
[0046]
[0047] Based on the comparison results, TVM+F+I+G4 was selected as the best model by ModelFinder software based on the BIC (Bayesian information criterion) standard. Phylogenetic analysis was performed using the phylogenetic tree construction software MEGA (version 7.0.26) to construct a phylogenetic tree using the maximum likelihood method. The resulting phylogenetic tree was used to determine the relationship and evolutionary branch structure between S. sinensis and other Salvia species. Figure 2 As shown in the data, S. sinensis clustered into a branch with Salvia bowleyana and Salvia dabieshanensis, and the bootstrap support rate of its differentiation node was 100, which clearly revealed the phylogenetic position of S. sinensis in the genus Salvia.
[0048] Example 3 Identification of S. sinensis Varieties Using the S. sinensis Chloroplast Genome
[0049] 1. DNA extraction of sample A and sample B: genomic DNA of sample A and sample B were extracted using the CTAB method.
[0050] 2. High-throughput sequencing: perform 150bp read-length paired-end sequencing on a high-throughput sequencing platform, and set the sequencing data volume for each sample to be approximately 5-10Gb.
[0051] 3. Chloroplast genome assembly: The specific steps are the same as step 3 of Example 1. The complete chloroplast genomes of sample A and sample B are assembled from the sequencing data.
[0052] 4. Sequence alignment and similarity determination: The chloroplast genome sequences of sample A and sample B were aligned with the reference chloroplast genome of S. sinensis (the full-length sequence consisting of sequences such as SEQ ID NO.1-SEQ ID NO.16 connected in sequence) (see Appendix Figure 3 The comparison results showed that the similarity index of sample A was 99.81%, which was lower than the judgment threshold of 99.9%, so sample A was not S. sinensis. The similarity index of sample B was 99.99%, which was higher than the threshold of 99.9%, so sample B was judged to be S. sinensis.
[0053] 5. Phylogenetic tree verification: According to the phylogenetic analysis method, the chloroplast genomes of sample A, sample B and representative species of Salvia such as Salvia sinensis were reconstructed (see Appendix Figure 4 The results showed that sample A and Salvia miltiorrhiza clustered in the same branch with a branch support rate of 100, while sample B and Sage sinensis clustered in the same branch with a branch support rate of 100, further verifying the accuracy of the sequence alignment results.
[0054] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0055] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.