Sea salt brassica napobrassica chloroplast gene for identifying sea salt brassica napobrassica variety and application thereof
By using the whole genome sequence of chloroplasts from *Zanthoxylum bungeanum* as a DNA barcode, combined with next-generation sequencing technology and sequence alignment methods, the accuracy problem of *Zanthoxylum bungeanum* variety identification was solved, achieving efficient and low-cost species identification and phylogenetic research.
Patent Information
- Application Number
- CN202511363254.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-23
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies make it difficult to accurately distinguish sea salt mustard greens from other mustard varieties based on morphological characteristics. The low degree of conservation of nuclear genome sequences leads to insufficient accuracy in species identification.
The complete chloroplast genome sequence of *Gnaphalium affine* was used as a DNA barcode. The complete chloroplast genome sequence was obtained through next-generation sequencing technology, and the sequence was compared using MAFFT software. When the similarity index reached 99.99% or higher, the sample was identified as *Gnaphalium affine*.
It improves the accuracy and efficiency of identification of Haiyan pickled mustard greens, provides more variation sites and genetic information, reduces detection costs and time, and supports the protection and breeding of Haiyan pickled mustard green germplasm resources.
Smart Images

Figure CN120945109A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of plant genetics and breeding technology, specifically relating to a chloroplast gene of *Hemiberlesia lataniae* used for identifying varieties of *Hemiberlesia lataniae* and its application. Background Technology
[0002] Haiyan turnip is a specialty of Haiyan County, Jiaxing City, Zhejiang Province. Its main characteristic is its enlarged tuber, which is widely used in pickling. However, online classifications of Haiyan turnip contain some scientific errors, such as stating that "turnip is also mustard greens, or rutabaga." The species of rutabaga has the Latin name... Brassica rapa L. Actually, turnips are a different species from sea salt mustard greens. Turnips are morphologically similar to mustard greens and radishes, and the morphological characteristics of these three have prompted scholars to differentiate them. Furthermore, the mustard greens to which sea salt mustard greens belong are not a basic species, but rather a hybrid of black mustard and brassica. Black mustard and brassica diverged 20 million years ago. The hybrid mustard greens have been cultivated into a series of varieties, and are sometimes treated as variants or even species by taxonomists; these are the mustard type of the three types of brassica.
[0003] Cultivated mustard species can be divided into four subgroups: Integrifolia, Juncea, Napiformis, and Tsatsai. Integrifolia includes leaf mustard, cut-leaf mustard, large-petiole mustard, and head mustard; Juncea includes oil-seed mustard; Napiformis includes root mustard; and Tsatsai includes multishoot mustard and stem mustard. Therefore, according to this classification, the sea salt mustard should be Napiformis. Napiformis was previously also identified as... Brassica juncea subsp. napiformis .
[0004] Morphological characteristics alone are no longer sufficient for systematic classification among different species of root mustard. With the rapid development of molecular biology, molecular typing of plants at the genomic level, i.e., classifying Brassica species using molecular methods, is an effective approach. However, the nuclear genome has a low degree of conservation, resulting in low accuracy when used for species classification. Therefore, there is an urgent need to provide a new sequence for species identification among mustard species. Summary of the Invention
[0005] The purpose of this invention is to provide a chloroplast gene for identifying the Haiyan kohlrabi variety, thereby filling the gap in existing genome databases that do not contain the complete chloroplast genome sequence of Haiyan kohlrabi, and to use this chloroplast genome sequence as a DNA barcode for molecular-level identification of Haiyan kohlrabi.
[0006] This invention provides a chloroplast gene of *Caulis Chinensis* for identifying *Caulis Chinensis* varieties from Haiyan. The sequence of the chloroplast gene is shown in SEQ ID NO.1, and the chloroplast gene is the whole genome sequence of the chloroplasts of *Caulis Chinensis*.
[0007] The second aspect of the present invention provides the application of the chloroplast gene of *Gnaphalium affine*, which is used for the variety identification of *Gnaphalium affine*.
[0008] Preferably, the method for identifying the variety of sea salt pickled mustard greens is as follows: Leaves were collected from the samples to be identified, and total genomic DNA was extracted to ensure that the leaves contained chloroplasts; Second-generation genome DNA sequencing; The sequences obtained from sequencing are spliced and assembled to obtain the chloroplast genome sequence of the sample to be identified; The chloroplast genome sequence of the sample to be identified is compared with the chloroplast gene of *Gnaphalium affine*. When the similarity index of the sample to be identified reaches 99.99% or higher, the sample to be identified is determined to be *Gnaphalium affine*.
[0009] Preferably, the method for extracting the total genomic DNA is the CTAB method.
[0010] Preferably, after extracting the total genomic DNA, RNase is added to the total genomic DNA to remove RNA from it.
[0011] Preferably, the sequence alignment process is as follows: The chloroplast genome sequence of the sample to be identified was compared with the chloroplast gene of *Gnaphalium affine* shown in SEQ ID NO.1 using MAFFT software in auto mode.
[0012] Preferably, the similarity index is calculated as follows: ;in, B The similarity index is expressed in % (%). M Total alignment length, in bp; N The number of differential bases is expressed in bp.
[0013] The third aspect of the present invention provides the application of the chloroplast gene of *Brassica oleracea*, which is used to construct a phylogenetic tree of *Brassica* species.
[0014] Compared with the prior art, the beneficial effects of the present invention are: This invention provides a chloroplast gene for identifying the *Caulis Haiyanensis* cultivar. haiyanensis*. The sequence of the chloroplast gene is shown in SEQ ID NO.1, and the chloroplast gene is the chloroplast genome of *Caulis Haiyanensis*. This invention utilizes next-generation sequencing technology to sequence, assemble, and annotate the chloroplast genome of *Caulis Haiyanensis*, thereby obtaining the complete chloroplast genome sequence. The obtained chloroplast genome sequence is then compared with the chloroplast genome sequences of other materials to be identified to determine whether the material to be identified is *Caulis Haiyanensis*.
[0015] Compared to markers such as SSR and RLFP, chloroplast genome sequences, as super DNA barcodes for species identification, provide more variation sites and genetic information with higher resolution. Chloroplast genome sequences can be obtained completely through low-depth next-generation sequencing, offering advantages in terms of lower cost and time.
[0016] This invention discloses for the first time the chloroplast genome sequence of *Brassica oleracea*, and applies it to the variety identification and phylogenetic study of *Brassica oleracea*, elucidating its phylogenetic position within the *Brassica* genus. The results of this study are of great significance for the future conservation of *Brassica oleracea* varieties, as well as for the breeding and conservation of *Brassica* species.
[0017] This invention utilizes molecular systematics and phylogenetic genomics methods to establish phylogenetic relationships among Brassica species through chloroplast genomes. Based on this, the whole chloroplast genome is used as a super DNA barcode for the Brassica genus to identify mustard varieties, providing a theoretical basis and technical support for the protection and species identification of germplasm resources of Haiyan mustard greens. This invention is of great significance for the breeding, genetic diversity evaluation, protection and utilization of Haiyan mustard greens.
[0018] The molecular identification method provided by this invention is characterized by accurate results and simple operation when applied to the identification of *Hemiberlesia lataniae* species. It only requires next-generation sequencing (NGS) and does not require traditional PCR testing. Currently, the cost of NGS is very low, and chloroplast genome sequencing can certainly become an important identification method for *Hemiberlesia lataniae* germplasm in practice.
[0019] This invention also provides the application of the above-mentioned *Brassica oleracea* chloroplast genome in the identification of *Brassica oleracea* and the reconstruction of the phylogenetic tree of known *Brassica* species, providing technical support for the protection of *Brassica oleracea* germplasm resources and species identification.
[0020] This invention uses the entire chloroplast genome of *Gnaphalium affine* as a DNA barcode to identify varieties of *Gnaphalium affine*. Because the chloroplast genome contains more genetic information, the identification results are more accurate. Attached Figure Description
[0021] Figure 1 The morphological characteristics of the sea salt pickled mustard greens.
[0022] Figure 2 This is a diagram of the chloroplast genome structure of *Zanthoxylum bungeanum*.
[0023] Figure 3 Neighbor-joining phylogenetic trees were constructed for the chloroplast whole genome sequences of *Hemiberlesia lataniae* and 116 other publicly available *Brassica*, *Radix*, and *Capsella* species.
[0024] Figure 4 A phylogenetic tree was constructed based on the maximum likelihood of the chloroplast genomes of *Hemiberlesia lataniae* and 116 other species of the genera *Brassica*, *Radix*, and *Capsella*. Detailed Implementation
[0025] The present invention will be further illustrated below with specific embodiments, but these embodiments do not limit the scope of the invention. Modifications or substitutions to the details and form of the technical solutions of the present invention may be made without departing from the spirit and scope of the invention, but all such modifications or substitutions fall within the protection scope of the present invention.
[0026] The inventive concept of this invention is as follows: With societal development, humanity's exploration of biodiversity has deepened, and the demand for the rational utilization of biological resources and rapid species identification has increased. To address this need, DNA barcoding technology has emerged. A DNA barcode is a standardized DNA sequence used for species identification, similar to a product barcode, to distinguish different species. For example, a fragment approximately 680 bp in the mitochondrial cytochrome oxidase subunit 1 gene can be used to identify over 200 species of lepidopteran insects. Each species possesses its own specific DNA sequence; by comparing a sample sequence with a reference sequence in a database, the species to which the sample belongs can be determined.
[0027] The core of DNA barcoding technology lies in its ability to rapidly and accurately identify the species to which an unknown sample belongs, which has wide applications in plant identification, ecological assessment, biodiversity monitoring, biosafety detection, and conservation biology research. With ongoing research, DNA barcoding technology has evolved from single-sequence encoding to the genome level, leading to the concept of "super barcodes." Although many researchers have proposed different plant DNA barcoding schemes, the short sequences and limited information of conventional barcodes result in insufficient ability to identify recently diverged species. Using the entire chloroplast genome sequence as a super DNA barcode for species identification can significantly improve identification efficiency, especially for closely related species, and there are already many successful precedents.
[0028] Most conventional plant DNA barcode sequences, with a small portion originating from nuclear DNA, are derived from chloroplast DNA. The chloroplast genome of angiosperms is a circular double-stranded genome capable of semi-independent replication and transcription. The chloroplast genome is largely maternally inherited, structurally conserved, and its long evolutionary history has united it with the mitochondrial and nuclear genomes to maintain species uniqueness, thus providing information for phylogenetic analysis. The chloroplast genome is smaller than the nuclear genome, exhibiting higher interspecific variation and lower intraspecific variation. Compared to conventional DNA barcodes, the whole chloroplast genome serves as a super DNA barcode for species identification, providing more variation sites and genetic information with greater resolution.
[0029] Taking Brassica plants as an example, at the molecular level, previous researchers have mostly used conventional molecular markers to study the interspecific relationships of individual species within the genus. The types of markers reported include RFLP, RAPD, SSR, etc. These markers mainly utilize the diverse molecular evolutionary information of the nuclear genome and are verified in multiple species through PCR.
[0030] The present invention aims to utilize the obtained whole chloroplast genome of *Gnaphalium affine* as a DNA barcode for the identification of *Gnaphalium affine* varieties.
[0031] To enable those skilled in the art to better understand and implement the technical solutions of this invention, the invention will be further described below with reference to specific embodiments. Unless otherwise specified, all reagents used in this invention are commercially available, and all methods used are conventional techniques in the art.
[0032] The list of abbreviations for this invention is shown in Table 1, and the list of species names is shown in Table 2.
[0033] Table 1 List of Abbreviations Table 2. Species Name Comparison Table Example 1 The chloroplast genes of the *Heiyan kohlrabi* variety were identified as follows: This invention provides a genome sequence of chloroplasts from sea salt Chinese cabbage, the nucleotide sequence of which is shown in SEQ ID NO.1.
[0034] The genome sequence of chloroplasts from *Zanthoxylum bungeanum* leaves in Haiyan was obtained as follows: 1) Total DNA extraction from samples: Fresh leaves of *Ziziphus jujuba* were collected from germplasm resources preserved by the Haiyan Agricultural and Rural Affairs Bureau in Jiaxing City, Zhejiang Province. Total DNA was extracted from approximately 10g of fresh leaf samples using the CTAB method. The samples were dried using a vacuum dryer to remove residual ethanol, and the DNA precipitate was air-dried. 100μL of EB buffer was added to dissolve the precipitate, followed by digestion with 3μL of 100mg / mL RNase at 37℃ for 30min to dissolve the precipitate and digest the RNA. Finally, the samples were quality checked using Nanodrop, Qbuit, and electrophoresis. Once the quality was deemed acceptable, the total DNA was obtained. See the Haiyan *Ziziphus jujuba* sample for details. Figure 1 .
[0035] 2) The extracted total DNA was subjected to high-throughput next-generation sequencing on the Illumina platform. Total DNA was randomly fragmented using a Covaris ultrasonic disruptor, followed by end repair, A-tailing, sequencing adapter addition, purification, and PCR amplification to complete library preparation. After library testing, different libraries were pooled into flow cells according to the effective concentration and target data volume requirements. After cBOT clustering, sequencing was performed using the Illumina high-throughput sequencing platform to obtain 2×150bp paired-end fragments; ensuring at least 1Gb of data was generated per individual.
[0036] 3) Sequencing Data Analysis: Since raw sequencing data may contain low-quality sequences, adapter sequences, etc., it needs to be filtered to obtain clean reads before chloroplast genome assembly. Raw sequencing data filtering uses FASTP software with default parameters: reads with more than 5% N bases are removed; reads with 50% or more low-quality bases are removed (bases with a quality value of 5 or less are considered low-quality); reads with adapter contamination are removed. Chloroplast genome assembly is performed using GetOrganelle software with default parameters. CPGAVAS2 software is used for gene annotation and mapping of the chloroplast genome to obtain the complete *Gnaphalium affine* chloroplast genome. A chloroplast genome annotation map is drawn using DOGMA software. Structural analysis and composition analysis of the chloroplast genome are performed.
[0037] 4) Results: The full-length chloroplast genome sequence of *Ziziphus julibrissin* is 153,483 bp, containing a large single copy region of 70,199 bp, a small single copy region of 17,774 bp, and two inverted repeat sequences, each 26,212 bp. The chloroplast genome encodes a total of 130 genes, including 85 protein-coding genes, 8 rRNA genes, and 37 tRNA genes. Figure 2 The chloroplast genome map of *Hemiberlesia lataniae* was drawn using OGDRAW software.
[0038] Example 2 The identification of chloroplast genes in *Hemiberlesia lataniae* varieties and their applications are detailed below: 1. Application of the chloroplast genome sequence of *Ziziphus jujuba* in the identification of *Ziziphus jujuba* varieties.
[0039] Specific implementation methods for identifying varieties of Haiyan pickled mustard greens: This invention purchased a sample of Chinese cabbage seeds from a randomly selected vegetable market in Haiyan City, Jiaxing City, for seed variety identification and testing. According to the vendor, the seeds were harvested from Chinese cabbage grown locally in Haiyan and were self-saved seeds.
[0040] 1) Sow the seeds in a greenhouse and take samples of the leaves that have grown after 30 days.
[0041] 2) Total DNA was extracted from approximately 10 grams of fresh leaves using the CTAB method. The leaves were dried in a vacuum dryer to remove residual ethanol, and the DNA precipitate was dried. 100 μL of EB buffer was added to dissolve the precipitate, and 3 μL of 100 mg / mL RNase was added for digestion. The precipitate was dissolved and the RNA was digested at 37°C for 30 min. Finally, the DNA was tested using Nanodrop, Qbuit, and electrophoresis.
[0042] 3) After completion, measure the concentration of the extracted DNA to confirm that it meets the initial concentration requirements for DNA sequencing library construction.
[0043] 4) Construct a DNA sequencing library for the leaf DNA according to the instructions of the DNA sequencing library preparation kit.
[0044] 5) Perform sequencing on the constructed DNA sequencing library using an Illumina sequencer; obtain sequencing data, which is 1Gb of 150bp sequenced data from both ends.
[0045] 6) The DNA extraction, library construction, and sequencing steps described above were completed by a research service company.
[0046] 7) Use Fastp software to perform quality control on the raw sequencing data, filter out low-quality sequencing data, and remove adapter portions from the sequences.
[0047] 8) The filtered sequencing data were assembled using GetOrangelle software to obtain the chloroplast genome sequence of the material, which was 153484 bp in length.
[0048] 9) The chloroplast genome sequence of this material was aligned with the chloroplast genome sequence of *Ziziphus jujuba* (sea salt turnip) obtained in this invention, as shown in SEQ ID NO.1, using MAFFT software in auto mode. Since the chloroplast genome is a circular DNA sequence, the start position of the linear sequence can be freely adjusted. Therefore, based on the alignment results, the start position of the chloroplast genome sequence of this material was adjusted to ensure the maximum alignment length. The overall alignment length was 153484 bp.
[0049] 10) Calculate the similarity index between the assembled species chloroplast genome sequence and the *Ziziphus jujuba* chloroplast genome sequence in this invention.
[0050] The similarity index is calculated as follows: ;in, B The similarity index is expressed in % (%). M Total alignment length, in bp; N The number of differential bases is expressed in bp.
[0051] When the similarity index between the two reaches 99.99% or higher, the variety is identified as Haiyan turnip. More preferably, when the chloroplast genome nucleotide sequence of the assembled sample is compared with the chloroplast genome nucleotide sequence of Haiyan turnip, if the sequences are completely identical, that is, the similarity index is 100%, the sample is definitely identified as Haiyan turnip.
[0052] The number of differing nucleotides between the two sequences was counted, and there were a total of 3 different bases. Therefore, the similarity index between the two sequences is: =99.99805%>99.99%. Therefore, it is determined that the seeds purchased from the Haiyan vegetable market, which were claimed to be Haiyan turnips, are indeed Haiyan turnips.
[0053] This embodiment demonstrates how to use the chloroplast genome sequence of *Zanthoxylum bungeanum* obtained by the present invention to identify some *Zanthoxylum bungeanum* varieties on the market.
[0054] 2. Application of *Brassica napus* in sea salt and in reconstructing the phylogenetic tree of known *Brassica* species, with specific implementation methods: To investigate the phylogenetic position of *Brassica rapa*, the chloroplast genome sequence of *Brassica rapa* obtained in this invention and the chloroplast genome sequences of 116 representative Brassica species downloaded from the NCBI website were used for multiple sequence alignment in auto mode using MAFFT software.
[0055] Using ModelFinder software, the optimal model was selected based on the BIC standard. Using IQTREE2 software, the maximum likelihood method was employed with 1000 replicates to reconstruct the phylogenetic tree, resulting in a phylogenetic tree of the Brassica genus based on the chloroplast genome. Using MEGA software with 1000 replicates, a neighbor-joining phylogenetic tree was constructed. See [link to MEGA software]. Phylogenetic trees based on the complete chloroplast genome using both neighbor-joining and maximum likelihood methods are also constructed. Figure 3 and Figure 4 The genetic relationship between Haiyan pickled mustard greens and other mustard varieties was determined.
[0056] See results Figure 3 and Figure 4 Based on the phylogenetic analysis of the chloroplast genome sequence of this invention, it was found that *Gnaphalium affine* is related to three... Brassica juncea subsp. integrifolia The sample and the other three mustard greens ( Brassica juncea The samples clustered together, indicating that the chloroplast genome sequence of *Gnaphalium affine* was similar to that of mustard (*Gnaphalium affine*). Brassica juncea The high similarity between the chloroplast genomes of the two species confirms that the sea salt turnip is a type of mustard green. Brassica juncea One of them. It was discovered that sea salt turnips and... B. juncea subsp. integrifolia They are closely related, but their morphological characteristics are significantly different. Further morphological analysis confirms their species classification as... B. juncea var. napiformis .
[0057] Three Brassica juncea subsp. integrifolia The GenBank numbers are KX68665.1, KX681663.1 and KX681662.1, respectively.
[0058] Figure 3 This is a neighbor-joining phylogenetic tree constructed using the whole chloroplast genome sequences of *Hemiberlesia lataniae* and 116 other publicly disclosed *Brassica*, *Radix*, and *Capsella* species described in this invention. The NCBI GenBank accession numbers for these 116 chloroplast genomes are provided in parentheses. This neighbor-joining phylogenetic tree was reconstructed based on complete chloroplast genome alignment data and statistical analysis was performed using 1000 bootstrap replicates. Bootstrap support is represented by node color: red indicates support below 50%, dark green indicates support between 50% and 70%, light green indicates support above 70%, and gray indicates data unavailable.
[0059] Figure 4This is a maximum likelihood phylogenetic tree constructed using the chloroplast genome sequences of *Hemiberlesia lataniae* and 116 other species from the genera *Brassica*, *Radix*, and *Capsella* described in this invention. The NCBI GenBank accession numbers for these 116 chloroplast genomes are provided in parentheses. This maximum likelihood phylogenetic tree was reconstructed based on complete chloroplast genome alignment data, and statistical analysis was performed using 1000 bootstrap replicates. Bootstrap support is represented by node color: red indicates support below 50%, dark green indicates support between 50% and 70%, light green indicates support above 70%, and gray indicates data unavailable.
[0060] SEQ ID NO.1 is 153483 bp in length. For specific sequence information, please refer to the computer-readable vector of the nucleotide or amino acid sequence listing. Specifically, 1 bp to 4248 bp are found in Sequence 1 of the computer-readable vector of the nucleotide or amino acid sequence listing, 4249 bp to 8531 bp are found in Sequence 2, and so on.
[0061] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0062] The embodiments described above are merely examples of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention.
Claims
1. A chloroplast gene for identifying varieties of *Hemiberlesia lataniae*, characterized in that, The sequence of the chloroplast gene of *Gnaphalium affine* is shown in SEQ ID NO.
1. The chloroplast gene of *Gnaphalium affine* is the whole genome sequence of the chloroplast of *Gnaphalium affine*.
2. The application of the chloroplast gene in sea salt mustard greens as described in claim 1, characterized in that, The chloroplast gene of *Gnaphalium affine* was used for variety identification of *Gnaphalium affine*.
3. The application as described in claim 2, characterized in that, The method for identifying the variety of sea salt pickled mustard greens is as follows: Leaves were collected from the samples to be identified, and total genomic DNA was extracted to ensure that the leaves contained chloroplasts; Second-generation genome DNA sequencing; The sequences obtained from sequencing are spliced and assembled to obtain the chloroplast genome sequence of the sample to be identified; The chloroplast genome sequence of the sample to be identified is compared with the chloroplast gene of *Gnaphalium affine*. When the similarity index of the sample to be identified reaches 99.99% or higher, the sample to be identified is determined to be *Gnaphalium affine*.
4. The application as described in claim 3, characterized in that, The total genomic DNA was extracted using the CTAB method.
5. The application as described in claim 3, characterized in that, After extracting the total genomic DNA, RNase needs to be added to the total genomic DNA to remove the RNA from it.
6. The application as described in claim 3, characterized in that, The sequence alignment process is as follows: The chloroplast genome sequence of the sample to be identified was compared with the chloroplast gene of *Gnaphalium affine* shown in SEQ ID NO.1 using MAFFT software in auto mode.
7. The application as described in claim 3, characterized in that, The similarity index is calculated as follows: ;in, B The similarity index is expressed in % (%). M Total alignment length, in bp; N The number of differential bases is expressed in bp.
8. The application of the chloroplast gene in sea salt mustard greens as described in claim 1, characterized in that, The chloroplast genes of the sea salt Chinese cabbage were used to construct a phylogenetic tree of Brassica species.