Carya hunanensis chloroplast genome and application in germplasm identification
By constructing the chloroplast genome of Yuanling walnut and applying an improved detection method, the problems of long detection cycles and inaccurate identification in existing technologies have been solved, enabling rapid and accurate seedling identification and phylogenetic research, which has important significance for germplasm conservation and breeding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-24
- Publication Date
- 2026-03-03
AI Technical Summary
Current testing technologies can only distinguish Yuanling walnuts based on fruit morphology and yield. The testing cycle is long and limited by time and plant growth and development stages, making it difficult to achieve rapid and accurate seedling identification.
This paper provides a method for the chloroplast genome of *Juglans regia* from Yuanling and its application, including total DNA extraction using a modified CTAB method, paired-end sequencing, chloroplast genome assembly and splicing, alignment with reference sequences, data processing using software such as SPAdes, Sequencher, Geneious, and Plann, and construction of a phylogenetic tree to achieve rapid and accurate germplasm identification.
It enables rapid and accurate seedling identification, removes the limitations of time and growth stage, provides higher detection efficiency and accuracy, and supports phylogenetic research and germplasm conservation of Juglandaceae plants.
Smart Images

Figure CN118360425B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of gene technology, specifically to a chloroplast genome of Yuanling walnut and a method for identifying germplasm using it. Background Technology
[0002] With societal development, the demand for exploring biodiversity, scientifically utilizing biological resources, and rapidly identifying species has become increasingly urgent, leading to the emergence of DNA barcoding technology. Initially, DNA barcoding referred to a standardized DNA sequence used for species identification. Herbert et al., the fathers of DNA barcoding, successfully identified two hundred lepidopteran insect species using a 680bp fragment from the mitochondrial cytochrome oxidase subunit 1 (COI) gene, and first proposed the concept of using DNA fragments to distinguish species, similar to how barcodes distinguish products. Each species possesses a unique DNA sequence, distinguishing it from other species; this is the fundamental principle of DNA barcoding technology. Therefore, by comparing the DNA sequence of a sample with a reference sequence library, the species and even variety of the sample can be identified.
[0003] As research has deepened, the concept of DNA barcoding has evolved from the initial conventional DNA barcodes with single sequences to genome-level superbarcodes. Although numerous scholars have proposed various DNA barcodes for plant identification, the short length and limited information of conventional barcodes result in poor ability to identify recently diverged species. Using the entire chloroplast genome sequence as a superDNA barcode for species identification could significantly improve identification efficiency, especially for closely related species; several beneficial attempts have already been made in this area. The chloroplast genome of angiosperms is a circular double-stranded structure capable of independent replication and transcription. The chloroplast genome is largely maternally inherited, structurally conserved, and its long evolutionary history has integrated it with the mitochondrial and nuclear genomes, collectively maintaining species uniqueness and providing information for phylogenetic analysis.
[0004] Most conventional plant DNA barcode sequences, with the exception of a small portion being nuclear DNA, originate from chloroplast DNA. The chloroplast genome is smaller than the nuclear genome, exhibiting higher interspecific variation and lower intraspecific variation. Compared to conventional DNA barcodes, the entire chloroplast genome serves as a super DNA barcode for species identification, providing more variation sites and genetic information with greater resolution. Some scholars believe that the variation sites contained in the entire plant chloroplast genome are comparable to the identification capabilities of the COI gene in animals. The core purpose of DNA barcoding technology is to determine the species to which an unknown sample belongs. It can be used in any field requiring knowledge of biological species, such as medicinal plant identification, ecological assessment, biodiversity monitoring, biosafety monitoring, and conservation biology research. Currently, DNA barcoding is widely used in fields such as species classification, resource conservation, systematics and evolution, and biodiversity research.
[0005] "Yuanling Hickory" is a superior single plant selected from wild hickory resources in Hunan Province. It boasts numerous advantages, including high yield and excellent quality. Its promotion and application can effectively address the industrial issues of hickory in Hunan Province and even nationwide. Assembling the chloroplast genome of "Yuanling Hickory" and analyzing its phylogenetic position is of great significance for future research on the protection, phylogeny, and breeding of hickory plants. Furthermore, the applied chloroplast genome sequence can enable rapid and accurate identification of "Yuanling Hickory" seedlings, which is crucial for protecting seedling intellectual property rights and promoting the high-quality development of the Hunan hickory industry in my country. Summary of the Invention
[0006] To address the aforementioned technical challenges, overcome the limitations of existing detection technologies that rely solely on fruit morphology and yield for differentiation, shorten the detection cycle, and remove the constraints imposed by time and plant growth stages, this solution provides a chloroplast genome of Yuanling walnut and its application in germplasm identification.
[0007] To achieve the above objectives, the present invention first provides a chloroplast genome of *Juglans regia*, the nucleotide sequence of which is shown in SEQ ID NO.1.
[0008] Based on a general inventive concept, this invention also provides an application of the chloroplast genome of *Juglans regia* in germplasm identification, comprising the following steps:
[0009] S1. Total DNA was extracted from the young leaves of the test sample using a modified CTAB method.
[0010] S2. Perform paired-end sequencing on the total DNA extracted in step S1;
[0011] S3, splicing and assembly of chloroplast genomes;
[0012] S4. The chloroplast genome nucleotide sequence of the assembled sample is compared with the chloroplast genome nucleotide sequence of Yuanling walnut. Only when the similarity index reaches 99.9% or higher can it be identified as Yuanling walnut.
[0013] Preferably, the sequencing depth in step S2 is ≥30×.
[0014] Preferably, the method for assembling and annotating the chloroplast genome in step S3 is as follows:
[0015] (1) De novo assembly of high-throughput sequencing data was performed using SPAdes 3.10.1 software, with the k-mer parameter set to 95;
[0016] (2) Based on the published pecan chloroplast genome sequence as a reference sequence, the Sequencher v5.4 software was used to select fragments belonging to the chloroplast genome from the assembled data for preliminary assembly;
[0017] (3) Using Geneious R 10.2.3 software, the preliminary assembly results were mapped with the original reads to complete the splicing and assembly of the chloroplast genome, and manual proofreading was performed;
[0018] (4) The chloroplast genome was annotated and proofread using Plann and Sequin software, thus obtaining the complete chloroplast genome of Yuanling walnut. The chloroplast genome annotation map was drawn using the online software DOGMA.
[0019] The present invention has the following beneficial effects:
[0020] (1) This invention discloses for the first time the chloroplast genome of *Juglans regia* and applies it to the phylogenetic study of Juglandaceae plants, thus elucidating the phylogenetic position of *Juglans regia* within the Juglandaceae family. The results of this study are of great significance for the germplasm conservation and new variety breeding of *Juglans regia* in Hunan.
[0021] (2) This scheme overcomes the shortcomings of existing detection technologies that can only distinguish between fruits based on morphological characteristics and yield, shortens the detection cycle, removes the limitations imposed by time and plant growth and development stages, and achieves rapid and accurate identification of seedling sources. The molecular identification method proposed in this application provides accurate results with good specificity; it is characterized by simple operation, high detection efficiency, accurate detection, and good repeatability.
[0022] (3) The super DNA barcode of the present invention was used to analyze the phylogenetic relationship of species in the Juglandaceae family and to identify seedlings of suspected Yuanling walnut. This shows that the super DNA barcode of the present invention can be used for phylogenetic research of Juglandaceae plants and can accurately identify seedlings or grafted seedlings of Yuanling walnut. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 Example 1: Chloroplast genome map of *Walnut husk* leaves from Yuanling;
[0025] Figure 2 Example 1: A phylogenetic tree of Juglandaceae plants based on chloroplast genome sequences;
[0026] Figure 3 This is a diagram showing the nucleotide sequence alignment of the chloroplast genome of sample A in Experiment Example 1 with that of Yuanling walnut leaves;
[0027] Figure 4 This is the phylogenetic tree of sample A in Experiment Example 1, along with Yuanling hickory and common Hunan hickory. Detailed Implementation
[0028] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0029] The following embodiments are used to illustrate the present invention, but are not intended to limit the scope of the invention. Any modifications or substitutions made to the methods, steps, or conditions of the present invention without departing from the spirit and essence of the invention are within the scope of the invention.
[0030] Unless otherwise specified, the technical means used in the embodiments are conventional means well known to those skilled in the art; unless otherwise specified, the reagents used in the embodiments are all commercially available.
[0031] Example 1
[0032] (1) Extraction of total DNA from chloroplast genome
[0033] DNA was extracted from the young leaves of *Juglans regia* in Yuanling using a modified CTAB method. The young leaves were ground into powder in liquid nitrogen. 0.1 g of the powder was transferred to a 2 mL microcentrifuge tube, and 630 μL of 1.5x CTAB buffer (preheated to 65°C) was added. 70 μL of anhydrous ethanol was added, and the tube was heated in a 65°C water bath for 30 min, shaking four times during the process. The microcentrifuge tube was then cooled, and 700 μL of chloroform:isoamyl alcohol (24:1) was added and mixed thoroughly. The mixture was centrifuged at 10,000 rpm for 10 min. The supernatant was transferred to a new microcentrifuge tube, and 2 / 3 of the supernatant volume of isopropanol was added. The mixture was mixed, the precipitate was allowed to stand, and the tube was centrifuged at 4,000 rpm for 3 min. The supernatant was collected, or the filamentous DNA was directly picked out and washed, and 500 μL of [unspecified ingredient] was added. Wash twice with 76% ethanol, air dry in a clean bench, add 100 μL of TEL buffer or ddH2O to dissolve DNA, add 20 μg / L ribonuclease, digest at 37℃ for 30 min, add an equal volume of chloroform:isoamyl alcohol = 24:1 and mix well, centrifuge at 1000 rpm for 10 min, add 2 volumes of anhydrous ethanol, mix well, precipitate, wash twice with 500 μL of ethanol, air dry in a clean bench, add 50 μL of ddH2O to dissolve DNA.
[0034] (2) Total DNA paired-end sequencing
[0035] After the genomic DNA of the sample passed the test, the DNA was fragmented using mechanical fragmentation (ultrasound). Then, the 400-600 bp DNA fragments were purified, end-repaired, 3′-A-added, and sequencing adapters were ligated. Fragment size selection was performed by agarose gel electrophoresis, followed by PCR amplification to form a sequencing library. The constructed library underwent library quality control. The quality-controlled libraries were then subjected to paired-end (PE) sequencing on the Illumina NovaSeq 6000 platform with a read length of 150 bp, ensuring a sequencing depth of at least 30× (no less than 10 Gb of data).
[0036] (3) Sequencing data quality control
[0037] The raw data is filtered using the fastp (version 0.20.0, https: / / github.com / OpenGene / fastp) software. This includes removing sequencing adapters and primer sequences from the reads, filtering out reads with an average quality value less than Q5, and filtering out reads with more than 5 consecutive undetected bases. The high-quality reads obtained after this series of quality control measures are called Clean Data.
[0038] (4) Chloroplast genome assembly
[0039] De novo assembly of high-throughput sequencing data was performed using SPAdes 3.10.1 software, with the k-mer parameter set to 95. Based on the published walnut chloroplast genome sequence as a reference, fragments belonging to the chloroplast genome were selected from the assembled data using Sequencher v5.4 software for preliminary assembly. The preliminary assembly results were mapped to the original reads using Geneious R 10.2.3 software to complete the splicing and assembly of the chloroplast genome, followed by manual proofreading. The chloroplast genome was annotated and proofread using Plann and Sequin software, thus obtaining the complete Yuanling walnut chloroplast genome. A chloroplast genome annotation map was drawn using the online software DOGMA (Dual Organellar Genome Annotator), as shown below. Figure 1 As shown.
[0040] (5) Chloroplast genome annotation
[0041] The CDS of chloroplasts were annotated using prodigal v2.6.3 (https: / / www.github.com / hyattpd / Prodigal), rRNA was predicted using hmmer v3.1b2 (http: / / www.hmmer.org / ), and tRNA was predicted using aragornv1.2.38 (http: / / 130.235.244.92 / ARAGORN / ).
[0042] The genome size of the chloroplasts of *Juglans regia* is 160,142 bp, with a total GC content of 36.22%. It contains one large single-copy region of 89,719 bp (GC content 33.82%), one small single-copy region of 18,755 bp (GC content 29.92%), and two inverted repeat sequences of 25,834 bp (GC content 42.67%). A total of 131 genes were annotated in the *Juglans regia* chloroplast genome, including 85 protein-coding genes, 37 transfer RNA genes, 8 ribosomal RNA genes, and 1 pseudogene, as shown in Table 1 below.
[0043] Table 1. Annotated gene information of the chloroplast genome of *Juglans regia* from Yuanling.
[0044]
[0045] Note: * indicates that the gene has 1 intron; ** indicates that the gene has 2 introns; # indicates a pseudogene; the number in parentheses indicates the copy number of the gene.
[0046] (6) Phylogenetic analysis methods
[0047] The assembled chloroplast genome of Yuanling walnut was merged with the chloroplast genome sequences of 16 species from the Juglandaceae and 3 species from the Rosaceae family published in GenBank (Table 2), and the sequences were aligned and manually adjusted using MAFFT (v7) and BioEdit software, respectively.
[0048] Table 2 shows the chloroplast genome sequences of the 19 species used for phylogenetic analysis.
[0049]
[0050]
[0051] Phylogenetic analysis was performed using the whole chloroplast genome. Circular sequences were assigned to the same starting point. Sequences between species were aligned using MAFFT software (v7.427, --auto mode). The aligned data were then used to construct phylogenetic trees using RAxML v8.2.10 (https: / / cme.h-its.org / exelixis / software.html) software, with the GTRGAMMA model selected. The maximum likelihood (ML) method was used, with 1000 replicates.
[0052] The results are as follows Figure 2 As shown, *Carya yuanlingensis* clusters with other species in the *Carya* genus and is most closely related to *Carya guizhouensis*. The self-development support rate of its differentiation node is 100%, indicating high credibility.
[0053] Experimental Example 1
[0054] Identification of suspected Yuanling walnut samples
[0055] Sample A, suspected to be a seedling of "Yuanling Mountain Walnut" collected in the wild, was identified. The specific steps are as follows:
[0056] (1) Extraction of total DNA from leaf sample A: Modified CTAB method, same as step (1) in Example 1.
[0057] (2) Genome sequencing: Illumina HiSeq PE150 was used for paired-end sequencing, the same as step (2) in Example 1;
[0058] (3) Sequencing data quality control: Same as step (3) in Example 1.
[0059] (4) Chloroplast genome splicing and assembly: Same as step (4) in Example 1.
[0060] (5) The chloroplast genome of the assembled sample A to be identified was compared with the nucleotide sequence (super barcode) of Yuanling walnut obtained in Example 1.
[0061] The results are as follows Figure 3 As shown, the chloroplast genome length of sample A is 160139 bp, and the sequence identity with the super barcode is 99.99%, which indicates that sample A is Yuanling walnut.
[0062] The chloroplast genome of sample A was combined with the nucleotide sequence of the super DNA barcode and the chloroplast genome sequences of four other common Hunan walnuts. Multiple sequence alignment was performed using MAFFT software (v7.427, --auto mode). The aligned data were then used to construct a phylogenetic tree using RAxML v8.2.10 (https: / / cme.h-its.org / exelixis / software.html) software, with the GTRGAMMA model selected and the maximum likelihood (ML) method employed. The bootstrap value was set to 1000 replicates.
[0063] The results are as follows Figure 4 As shown, test sample A clustered with the super DNA barcode Yuanling walnut, and the self-development support rate of its differentiation node was 100%, indicating high reliability. Phylogenetic analysis also proved that test sample A is Yuanling walnut.
[0064] The above description is merely a preferred embodiment of the present invention, and the scope of protection of the present invention is not limited to the above embodiments. For those skilled in the art, any improvements and modifications obtained without departing from the technical concept of the present invention should also be considered within the scope of protection of the present invention.
[0065] The sequence of SEQ ID NO.1 is as follows:
[0066] <? xml version="1.0" encoding="UTF-8"? >
[0067] <! DOCTYPE ST26SequenceListing PUBLIC"- / / WIPO / / DTD Sequence Listing1.3 / / EN""ST26SequenceListing_V1_3.dtd">
[0068] <ST26SequenceListing dtdVersion="V1_3" fileName="The Chloroplast Genome of Carya hunanensis and Its Application in Germplasm Identification.xml" softwareName="WIPOSequence" softwareVersion="2.3.0" productionDate="2024-02-22">
[0069] <applicantfilereference> Hunan Botanical Garden< / applicantfilereference>
[0070] <ApplicantName languageCode="zh">Hunan Botanical Garden
[0071] <applicantnamelatin> Hunan Botanical Garden< / applicantnamelatin>
[0072] <InventionTitle languageCode="zh">The Chloroplast Genome of Carya hunanensis and Its Application in Germplasm Identification
[0073] <sequencetotalquantity> 1< / sequencetotalquantity>
[0074] <SequenceData sequenceIDNumber="1">
[0075] <insdseq>
[0076] <INSDSeq_length> 160142< / INSDSeq_length>
[0077] <INSDSeq_moltype> DNA< / INSDSeq_moltype>
[0078] <INSDSeq_division> PAT< / INSDSeq_division>
[0079] <INSDSeq_feature-table>
[0080] <insdfeature>
[0081] <INSDFeature_key>source< / INSDFeature_key>
[0082] <INSDFeature_location>1..160142< / INSDFeature_location>
[0083] <INSDFeature_quals>
[0084] <insdqualifier>
[0085] <INSDQualifier_name>mol_type< / INSDQualifier_name>
[0086] <INSDQualifier_value>other DNA< / INSDQualifier_value>
[0087] < / insdqualifier>
[0088] <INSDQualifier id="q2">
[0089] <INSDQualifier_name>organism< / INSDQualifier_name>
[0090] <INSDQualifier_value>Artificial Sequence< / INSDQualifier_value>
[0091]
[0092] < / INSDFeature_quals>
[0093] < / insdfeature>
[0094] < / INSDSeq_feature-table>
[0095]
[0096] < / insdseq>
[0097]
[0098]
Claims
1. An application of the chloroplast genome of *Juglans regia* in germplasm identification of *Juglans regia*, characterized in that... Includes the following steps: S1. Total DNA was extracted from the young leaves of the test samples using a modified CTAB method: S2. Perform paired-end sequencing on the total DNA extracted in step S1; S3, splicing and assembly of chloroplast genomes; S4. The chloroplast genome nucleotide sequence of the assembled sample is compared with the chloroplast genome nucleotide sequence of Yuanling walnut. Only when the similarity index reaches 99.99% or higher can it be identified as Yuanling walnut. The nucleotide sequence of the chloroplast genome of the Yuanling walnut is shown in SEQ ID NO.
1.
2. The application according to claim 1, characterized in that, The sequencing depth in step S2 is ≥30×.
3. The application according to claim 1, characterized in that, The specific method for assembling and annotating the chloroplast genome in step S3 is as follows: (1) De novo assembly of high-throughput sequencing data was performed using SPAdes 3.10.1 software, with the k-mer parameter set to 95; (2) Based on the published pecan chloroplast genome sequence as a reference sequence, the Sequencher v5.4 software was used to select fragments belonging to the chloroplast genome from the assembled data for preliminary assembly; (3) Using Geneious R 10.2.3 software, the preliminary assembly results were mapped with the original reads to complete the splicing and assembly of the chloroplast genome, and manual proofreading was performed; (4) The chloroplast genome was annotated and proofread using Plann and Sequin software, thus obtaining the complete chloroplast genome of Yuanling walnut. The chloroplast genome annotation map was drawn using the online software DOGMA.
Citation Information
Patent Citations
Application of cerasus buergeriana chloroplast genome in variety identification
CN114807330A
Prunus padus chloroplast genome and application in variety identification
CN115044695A