Assembled and annotated chloroplast genome of pteris amurensis
By assembling and annotating the chloroplast genome of *Miaofeng* fern, the problem of insufficient research on single fern genomes has been solved, enriching the fern gene database, promoting the understanding of fern evolutionary history and environmental adaptation mechanisms, and applying it to biodiversity conservation and plant ecological restoration.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG FOREST & GRASS GERMPLASM RESOURCE CENT (SHANDONG YAOXIANG FOREST FARM)
- Filing Date
- 2026-01-23
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies for fern genome studies, especially for individual rock ferns, are relatively rare. The lack of systematic assembly and annotation limits our understanding of fern evolutionary history and environmental stress adaptation mechanisms.
An assembled and annotated chloroplast genome of *Dryopteris melanogaster* is provided, including nucleotide sequences, protein-coding gene families, tRNA and rRNA genes, assembled and annotated using nanopore single-molecule sequencing and specific software to ensure the integrity and accuracy of the genome.
This research enriched the fern gene database, deciphered the genetic code of *Miaofengyan fern*, and provided a foundation for exploring the evolutionary history and environmental adaptation mechanisms of ferns. It can be applied to biodiversity conservation, plant ecological restoration, and horticulture.
Smart Images

Figure CN121896233A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of molecular biology and plant genetics, and more specifically to an assembled and annotated chloroplast genome of *Miaofengyan fern*. Background Technology
[0002] *Woodsia oblonga* Ching & S. H. Wu is a small fern belonging to the genus *Woodsia* in the family Woodsiaceae. The plant typically grows to a height of 7-18 cm. Its rhizome is short and ascending or erect, densely covered with brown, membranous, lanceolate scales along with the base of the petiole. *Woodsia oblonga* grows in shady rock crevices on mountain slopes at altitudes of 200-1800 meters, primarily distributed in Miaofeng Mountain (Beijing), Ji County and Beidaihe (Hebei), Taishan and Kunyu Mountain (Shandong), and Song County (Henan). Ferns are incredibly diverse and, as a core component of Earth's plant life, play a vital role in ecosystems and the evolutionary history of terrestrial plants. Ferns not only maintain the carbon-oxygen balance of ecosystems through photosynthesis and respiration, but also contribute to beautifying living environments, supporting agricultural production, and promoting pharmaceutical research.
[0003] Chloroplasts are the site of photosynthesis in higher plants and algae, converting light energy into chemical energy. They are unique energy converters and, as semi-autonomous organelles within plant cells, are subject to dual regulation by nuclear genes and their own genes.
[0004] There is relatively little research on ferns both domestically and internationally. The main research areas include new species records, species nomenclature standards, fern classification and floristic analysis, spore morphology studies, and chloroplast gene studies. Genomic studies on individual rock ferns are also rare.
[0005] Therefore, studying the assembled and annotated chloroplast genome of *Miaofengyan fern* is an urgent problem to be solved. Summary of the Invention
[0006] Specifically addressing the shortcomings of existing technologies, this invention provides an assembled and annotated chloroplast genome of *Fernopteris yunnanensis*, which is a single circular molecule. The total length of the *Fernopteris yunnanensis* chloroplast genome is 149,180 bp, and the GC content is 43.14%. This invention enriches the fern gene database, laying the foundation for exploring the evolutionary history of ferns and their adaptation mechanisms to environmental stress. It deciphers the genetic code of *Fernopteris yunnanensis* at the molecular level and can be applied to biodiversity conservation, plant ecological restoration, and horticulture.
[0007] The objective of this invention is achieved through the following technical solution: This invention provides an assembled and annotated chloroplast genome of *Fragaria melanogaster*, wherein the chloroplast genome of *Fragaria melanogaster* is a single circular molecule; the total length of the chloroplast genome of *Fragaria melanogaster* is 149180 bp; the nucleotide sequence of the chloroplast genome of *Fragaria melanogaster* is shown in SEQ ID NO.1; and the chloroplast genome of *Fragaria melanogaster* has a GC content of 43.14%.
[0008] In some specific embodiments of the present invention, the chloroplast genome of *Dryopteris melanogaster* includes tRNA genes, rRNA genes, and a protein-coding gene family; the number of tRNA genes is 29, the number of rRNA genes is 4, and the number of protein-coding gene families is 17.
[0009] In some specific embodiments of the present invention, the protein-coding gene family includes NADH dehydrogenase subunit genes, photosystem I subunit genes, photosystem II subunit genes, cytochrome b / f complex subunit genes, ATP synthase subunit genes, ribulose-1,5-bisphosphate carboxylase / oxygenase large subunit genes, DNA-dependent RNA polymerase genes, ribosome large subunit genes, ribosome small subunit genes, maturation enzyme genes, type C cytochrome synthase genes, membrane protein genes, translation initiation factor genes, protease genes, acetyl-CoA-carboxylase subunit genes, prochlorophyll reductase subunit genes, and conserved open reading frame genes.
[0010] In some specific embodiments of the present invention, the NADH dehydrogenase subunit genes include ndhA, ndhB, ndhC, ndhD, ndhE, ndhF, ndhG, ndhH, ndhI, ndhJ, and ndhK; The photosystem I subunit genes include psaA, psaB, psaC, psaI, and psaJ; The photosystem II subunit genes include psbA, psbB, psbC, psbD, psbE, psbF, psbH, psbI, psbJ, psbK, psbL, psbM, psbN, psbT, psbZ, and ycf3; The cytochrome b / f complex subunit genes include petA, petB, petD, petG, petL, and petN; The ATP synthase subunit genes include atpA, atpB, atpE, atpF, atpH, and atpI; The ribulose-1,5-bisphosphate carboxylase / oxygenase large subunit gene includes rbcL; The DNA-dependent RNA polymerase genes include rpoA, rpoB, rpoC1, and rpoC2; The ribosomal large subunit genes include rpl2, rpl14, rpl16, rpl20, rpl21, rpl22, rpl23, rpl32, rpl33 and rpl36; The ribosomal small subunit genes include rps2, rps3, rps4, rps7, rps8, rps11, rps12, rps14, rps15, rps16, rps18 and rps19; The maturation enzyme gene includes matK; The c-type cytochrome synthase gene includes ccsA; The membrane protein gene includes cemA; The translation initiation factor gene includes infA; The protease gene includes clpP; The subunit gene of the acetyl-CoA-carboxylase includes accD; The protochlorophyll reductase subunit genes include chlB, chlL, and chlN; The conserved open reading frame genes include ycf1, ycf2, ycf4, and ycf12; In some specific embodiments of the present invention, the nucleotide sequence of ndhA is shown in SEQ ID NO.2; the nucleotide sequence of ndhB is shown in SEQ ID NO.3; the nucleotide sequence of ndhC is shown in SEQ ID NO.4; the nucleotide sequence of ndhD is shown in SEQ ID NO.5; the nucleotide sequence of ndhE is shown in SEQ ID NO.6; the nucleotide sequence of ndhF is shown in SEQ ID NO.7; the nucleotide sequence of ndhG is shown in SEQ ID NO.8; the nucleotide sequence of ndhH is shown in SEQ ID NO.9; the nucleotide sequence of ndhI is shown in SEQ ID NO.10; the nucleotide sequence of ndhJ is shown in SEQ ID NO.11; and the nucleotide sequence of ndhK is shown in SEQ ID NO.12. The nucleotide sequence of psaA is shown in SEQ ID NO.13; the nucleotide sequence of psaB is shown in SEQ ID NO.14; the nucleotide sequence of psaC is shown in SEQ ID NO.15; the nucleotide sequence of psaI is shown in SEQ ID NO.16; and the nucleotide sequence of psaJ is shown in SEQ ID NO.17. The nucleotide sequence of psbA is shown in SEQ ID NO. 18; the nucleotide sequence of psbB is shown in SEQ ID NO. 19; the nucleotide sequence of psbC is shown in SEQ ID NO. 20; the nucleotide sequence of psbD is shown in SEQ ID NO. 21; the nucleotide sequence of psbE is shown in SEQ ID NO. 22; the nucleotide sequence of psbF is shown in SEQ ID NO. 23; the nucleotide sequence of psbH is shown in SEQ ID NO. 24; the nucleotide sequence of psbI is shown in SEQ ID NO. 25; the nucleotide sequence of psbJ is shown in SEQ ID NO. 26; the nucleotide sequence of psbK is shown in SEQ ID NO. 27; the nucleotide sequence of psbL is shown in SEQ ID NO. 28; the nucleotide sequence of psbM is shown in SEQ ID NO. 29; the nucleotide sequence of psbN is shown in SEQ ID NO. 30; the nucleotide sequence of psbT is shown in SEQ ID NO. 31; and the nucleotide sequence of psbZ is shown in SEQ ID NO. 18. As shown in NO.32; the nucleotide sequence of ycf3 is shown in SEQ ID NO.33; The nucleotide sequence of petA is shown in SEQ ID NO.34; the nucleotide sequence of petB is shown in SEQ ID NO.35; the nucleotide sequence of petD is shown in SEQ ID NO.36; the nucleotide sequence of petG is shown in SEQ ID NO.37; the nucleotide sequence of petL is shown in SEQ ID NO.38; and the nucleotide sequence of petN is shown in SEQ ID NO.39. The nucleotide sequence of atpA is shown in SEQ ID NO. 40; the nucleotide sequence of atpB is shown in SEQ ID NO. 41; the nucleotide sequence of atpE is shown in SEQ ID NO. 42; the nucleotide sequence of atpF is shown in SEQ ID NO. 43; the nucleotide sequence of atpH is shown in SEQ ID NO. 44; the nucleotide sequence of atpI is shown in SEQ ID NO. 45; and the nucleotide sequence of rbcL is shown in SEQ ID NO. 46. The nucleotide sequence of rpoA is shown in SEQ ID NO.47; the nucleotide sequence of rpoB is shown in SEQ ID NO.48; the nucleotide sequence of rpoC1 is shown in SEQ ID NO.49; and the nucleotide sequence of rpoC2 is shown in SEQ ID NO.50. The nucleotide sequence of rpl2 is shown in SEQ ID NO. 51; the nucleotide sequence of rpl14 is shown in SEQ ID NO. 52; the nucleotide sequence of rpl16 is shown in SEQ ID NO. 53; the nucleotide sequence of rpl20 is shown in SEQ ID NO. 54; the nucleotide sequence of rpl21 is shown in SEQ ID NO. 55; the nucleotide sequence of rpl22 is shown in SEQ ID NO. 56; the nucleotide sequence of rpl23 is shown in SEQ ID NO. 57; the nucleotide sequence of rpl32 is shown in SEQ ID NO. 58; the nucleotide sequence of rpl33 is shown in SEQ ID NO. 59; and the nucleotide sequence of rpl36 is shown in SEQ ID NO. 60. The nucleotide sequence of rps2 is shown in SEQ ID NO. 61; the nucleotide sequence of rps3 is shown in SEQ ID NO. 62; the nucleotide sequence of rps4 is shown in SEQ ID NO. 63; the nucleotide sequence of rps7 is shown in SEQ ID NO. 64; the nucleotide sequence of rps8 is shown in SEQ ID NO. 65; the nucleotide sequence of rps11 is shown in SEQ ID NO. 66; the nucleotide sequence of rps12 is shown in SEQ ID NO. 67; the nucleotide sequence of rps14 is shown in SEQ ID NO. 68; the nucleotide sequence of rps15 is shown in SEQ ID NO. 69; the nucleotide sequence of rps16 is shown in SEQ ID NO. 70; the nucleotide sequence of rps18 is shown in SEQ ID NO. 71; and the nucleotide sequence of rps19 is shown in SEQ ID NO. 72. The nucleotide sequence of matK is shown in SEQ ID NO.73; the nucleotide sequence of ccsA is shown in SEQ ID NO.74; the nucleotide sequence of cemA is shown in SEQ ID NO.75; the nucleotide sequence of infA is shown in SEQ ID NO.76; the nucleotide sequence of clpP is shown in SEQ ID NO.77; and the nucleotide sequence of accD is shown in SEQ ID NO.78. The nucleotide sequence of chlB is shown in SEQ ID NO. 79; the nucleotide sequence of chlL is shown in SEQ ID NO. 80; the nucleotide sequence of chlN is shown in SEQ ID NO. 81; the nucleotide sequence of ycf1 is shown in SEQ ID NO. 82; the nucleotide sequence of ycf2 is shown in SEQ ID NO. 83; the nucleotide sequence of ycf4 is shown in SEQ ID NO. 84; and the nucleotide sequence of ycf12 is shown in SEQ ID NO. 85.
[0011] In some specific embodiments of the present invention, the alanine Ala codon to GCU, arginine Arg codon to AGA, and leucine Leu codon to UUA codon in the chloroplast genome of the Miaofeng rock fern have the strongest preference, with an RSCU value of 1.51; while the serine Serine codon to AGC codon has the weakest preference, with an RSCU value of 0.53.
[0012] In some specific embodiments of the present invention, the chloroplast genome of *Miaofengyan fern* includes repetitive sequences; the repetitive sequences include simple repetitive sequences, tandem repetitive sequences, and scattered repetitive sequences; The number of simple repeating sequences is 37; the number of tandem repeating sequences is 23; and the number of scattered repeating sequences is 59 pairs.
[0013] In some specific embodiments of the present invention, the simple repeating sequence includes a single nucleotide repeating sequence, a dinucleotide repeating sequence, a trinucleotide repeating sequence, and a tetranucleotide repeating sequence; The number of mononucleotide repeat sequences is 22; the number of dinucleotide repeat sequences is 8; the number of trinucleotide repeat sequences is 1; and the number of tetranucleotide repeat sequences is 6.
[0014] In some specific embodiments of the present invention, the scattered repeating sequence includes palindromic repeats, forward repeats, reverse repeats, and complementary repeats; The number of palindromic repetitions is 21; the number of forward repetitions is 23; the number of reverse repetitions is 9; and the number of complementary repetitions is 6. The longest length of the palindromic repeat sequence is 22502 bp; The longest sequence in the positive repeat is 57 bp.
[0015] In some specific embodiments of the present invention, the Miaofeng rock fern is derived from a Miaofeng rock fern plant located in Linyi City, Shandong Province, at latitude 35.5419444N and longitude 117.8558333E.
[0016] The beneficial effects achieved by this invention are as follows: This invention provides an assembled and annotated chloroplast genome of *Miaofengyan fern*, enriching the fern gene database, laying the foundation for exploring the evolutionary history of ferns and their adaptation mechanisms to environmental stress, deciphering the genetic code of *Miaofengyan fern* at the molecular level, and applying it to biodiversity conservation, plant ecological restoration, and horticulture. Attached Figure Description
[0017] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 Images of the Miaofeng Rock Fern plant and its habitat; among them... Figure 1 A in the image represents the young leaves of the Miaofengyan fern plant; Figure 1 B in the diagram represents the habitat of the Miaofeng rock fern. Figure 2 This is a structural diagram of the chloroplast genome of *Adiantum melanogaster*. Figure 3 A diagram showing the codon bias analysis of the chloroplast genome of *Adiantum melanogaster*. Figure 4 A diagram showing the repetitive sequence analysis of the chloroplast genome of *Adiantum melanogaster*. Figure 4 In the diagram, A represents a simple repeat sequence analysis. Figure 4 B in the diagram represents the analysis of tandem repeat sequences and scattered repeat sequences. Detailed Implementation
[0019] The present invention will be further described in detail below with reference to specific embodiments. The following embodiments are not intended to limit the present invention, but only to illustrate the present invention. Unless otherwise specified, the experimental methods used in the following embodiments are generally performed under conventional conditions. Unless otherwise specified, the materials and reagents used in the following embodiments are commercially available.
[0020] Example 1: Chloroplast genome sequencing of *Miaofengyan fern* On August 4, 2024, experimental material of *Miaofeng rock fern* was collected from rock crevices in Guimeng Mountain, Pingyi County, Linyi City (elevation 730.7m, 35.5419444N, 117.8558333E). Fresh, intact young leaves were selected and immediately placed in liquid nitrogen for rapid freezing, then stored in an ultra-low temperature freezer at -80°C. Images of the *Miaofeng rock fern* plant and its habitat are shown below. Figure 1 As shown, the experimental material in this embodiment is the fresh tender leaves of *Dryopteris melanoxylon*, such as... Figure 1 As shown in A, the living environment of the Miaofeng rock fern is as follows: Figure 1As shown in B in the diagram.
[0021] Genomic DNA was obtained using a DNA extraction kit from TIANGEN, and sequencing and library construction were performed using nanopore single-molecule sequencing technology to obtain raw sequence data. The nucleotide sequence of the chloroplast genome of the Miaofengyan fern is shown in SEQ ID NO.1.
[0022] Example 2: Assembly and annotation of raw chloroplast genomic DNA sequence data from *Fern simonii*. The raw sequence data of the chloroplast genome DNA obtained from *Adiantum melanogaster* were used to assemble a circular chloroplast genome using the default parameters of GetOrganelle software (v1.7.7.1). Preliminary genome annotation was performed using CPGAVAS2 software, and the chloroplast genome map was visualized using OGDRAW software. tRNAs were annotated using tRNAscan-SE software, and rRNAs were annotated using BLASTN software. All annotation errors were manually corrected using CPGView and Apollo software, referencing closely related species.
[0023] Results: The assembled and annotated chloroplast genome of *Miaofengyan* fern was obtained through software visualization. The structural diagram of the *Miaofengyan* fern chloroplast genome is shown below. Figure 2As shown, the main structure of the chloroplast genome of *Adiantum melanogaster* is a single circular molecule with a total length of 149,180 bp and a GC content of 43.14%. The chloroplast genome of *Adiantum melanogaster* was annotated, and the annotated chloroplast-coding genes are shown in Table 1. A total of 84 unique protein-coding genes (4 of which are multiple copies), 29 tRNA genes (6 of which are multiple copies), and 4 rRNA genes (4 of which are multiple copies) were annotated. Protein-coding genes comprise 17 gene families: 11 NADH dehydrogenase subunit genes (ndhA, ndhB, ndhC, ndhD, ndhE, ndhF, ndhG, ndhH, ndhI, ndhJ, ndhK); 5 photosystem I subunit genes (psaA, psaB, psaC, psaI, psaJ); 16 photosystem II subunit genes (psbA, psbB, psbC, psbD, psbE, psbF, psbH, psbI, psbJ, psbK, psbL, psbM, psbN, psbT, psbZ, ycf3); and 6 cytochrome genes. The gene sequence includes: β / F complex subunit genes (petA, petB, petD, petG, petL, petN); 6 ATP synthase subunit genes (atpA, atpB, atpE, atpF, atpH, atpI); 1 ribulose-1,5-bisphosphate carboxylase / oxygenase large subunit gene (rbcL); 4 DNA-dependent RNA polymerase genes (rpoA, rpoB, rpoC1, rpoC2); and 10 ribosomal large subunit genes (rpl2, rpl14, rpl16, rpl20, rpl21, rpl22, rpl23, rpl32, rpl33, rpl36). Twelve ribosomal small subunit genes (rps2, rps3, rps4, rps7, rps8, rps11, rps12, rps14, rps15, rps16, rps18, rps19); one maturation enzyme gene (matK); one c-type cytochrome synthase gene (ccsA); one membrane protein gene (cemA); one translation initiation factor gene (infA); one protease gene (clpP); one acetyl-CoA-carboxylase subunit gene (accD); three protochlorophyll reductase subunit genes (chlB, chlL, chlN); and four conserved open reading frame genes (ycf1, ycf2, ycf4, ycf12).
[0024] The nucleotide sequence of ndhA is shown in SEQ ID NO.2; the nucleotide sequence of ndhB is shown in SEQ ID NO.3; the nucleotide sequence of ndhC is shown in SEQ ID NO.4; the nucleotide sequence of ndhD is shown in SEQ ID NO.5; the nucleotide sequence of ndhE is shown in SEQ ID NO.6; the nucleotide sequence of ndhF is shown in SEQ ID NO.7; the nucleotide sequence of ndhG is shown in SEQ ID NO.8; the nucleotide sequence of ndhH is shown in SEQ ID NO.9; the nucleotide sequence of ndhI is shown in SEQ ID NO.10; the nucleotide sequence of ndhJ is shown in SEQ ID NO.11; and the nucleotide sequence of ndhK is shown in SEQ ID NO.12. The nucleotide sequence of psaA is shown in SEQ ID NO.13; the nucleotide sequence of psaB is shown in SEQ ID NO.14; the nucleotide sequence of psaC is shown in SEQ ID NO.15; the nucleotide sequence of psaI is shown in SEQ ID NO.16; and the nucleotide sequence of psaJ is shown in SEQ ID NO.17. The nucleotide sequence of psbA is shown in SEQ ID NO. 18; the nucleotide sequence of psbB is shown in SEQ ID NO. 19; the nucleotide sequence of psbC is shown in SEQ ID NO. 20; the nucleotide sequence of psbD is shown in SEQ ID NO. 21; the nucleotide sequence of psbE is shown in SEQ ID NO. 22; the nucleotide sequence of psbF is shown in SEQ ID NO. 23; the nucleotide sequence of psbH is shown in SEQ ID NO. 24; the nucleotide sequence of psbI is shown in SEQ ID NO. 25; the nucleotide sequence of psbJ is shown in SEQ ID NO. 26; the nucleotide sequence of psbK is shown in SEQ ID NO. 27; the nucleotide sequence of psbL is shown in SEQ ID NO. 28; the nucleotide sequence of psbM is shown in SEQ ID NO. 29; the nucleotide sequence of psbN is shown in SEQ ID NO. 30; the nucleotide sequence of psbT is shown in SEQ ID NO. 31; and the nucleotide sequence of psbZ is shown in SEQ ID NO. 18. As shown in NO.32; the nucleotide sequence of ycf3 is shown in SEQ ID NO.33; The nucleotide sequence of petA is shown in SEQ ID NO.34; the nucleotide sequence of petB is shown in SEQ ID NO.35; the nucleotide sequence of petD is shown in SEQ ID NO.36; the nucleotide sequence of petG is shown in SEQ ID NO.37; the nucleotide sequence of petL is shown in SEQ ID NO.38; and the nucleotide sequence of petN is shown in SEQ ID NO.39. The nucleotide sequence of atpA is shown in SEQ ID NO. 40; the nucleotide sequence of atpB is shown in SEQ ID NO. 41; the nucleotide sequence of atpE is shown in SEQ ID NO. 42; the nucleotide sequence of atpF is shown in SEQ ID NO. 43; the nucleotide sequence of atpH is shown in SEQ ID NO. 44; the nucleotide sequence of atpI is shown in SEQ ID NO. 45; and the nucleotide sequence of rbcL is shown in SEQ ID NO. 46. The nucleotide sequence of rpoA is shown in SEQ ID NO.47; the nucleotide sequence of rpoB is shown in SEQ ID NO.48; the nucleotide sequence of rpoC1 is shown in SEQ ID NO.49; and the nucleotide sequence of rpoC2 is shown in SEQ ID NO.50. The nucleotide sequence of rpl2 is shown in SEQ ID NO. 51; the nucleotide sequence of rpl14 is shown in SEQ ID NO. 52; the nucleotide sequence of rpl16 is shown in SEQ ID NO. 53; the nucleotide sequence of rpl20 is shown in SEQ ID NO. 54; the nucleotide sequence of rpl21 is shown in SEQ ID NO. 55; the nucleotide sequence of rpl22 is shown in SEQ ID NO. 56; the nucleotide sequence of rpl23 is shown in SEQ ID NO. 57; the nucleotide sequence of rpl32 is shown in SEQ ID NO. 58; the nucleotide sequence of rpl33 is shown in SEQ ID NO. 59; and the nucleotide sequence of rpl36 is shown in SEQ ID NO. 60. The nucleotide sequence of rps2 is shown in SEQ ID NO. 61; the nucleotide sequence of rps3 is shown in SEQ ID NO. 62; the nucleotide sequence of rps4 is shown in SEQ ID NO. 63; the nucleotide sequence of rps7 is shown in SEQ ID NO. 64; the nucleotide sequence of rps8 is shown in SEQ ID NO. 65; the nucleotide sequence of rps11 is shown in SEQ ID NO. 66; the nucleotide sequence of rps12 is shown in SEQ ID NO. 67; the nucleotide sequence of rps14 is shown in SEQ ID NO. 68; the nucleotide sequence of rps15 is shown in SEQ ID NO. 69; the nucleotide sequence of rps16 is shown in SEQ ID NO. 70; the nucleotide sequence of rps18 is shown in SEQ ID NO. 71; and the nucleotide sequence of rps19 is shown in SEQ ID NO. 72. The nucleotide sequence of matK is shown in SEQ ID NO.73; the nucleotide sequence of ccsA is shown in SEQ ID NO.74; the nucleotide sequence of cemA is shown in SEQ ID NO.75; the nucleotide sequence of infA is shown in SEQ ID NO.76; the nucleotide sequence of clpP is shown in SEQ ID NO.77; and the nucleotide sequence of accD is shown in SEQ ID NO.78. The nucleotide sequence of chlB is shown in SEQ ID NO. 79; the nucleotide sequence of chlL is shown in SEQ ID NO. 80; the nucleotide sequence of chlN is shown in SEQ ID NO. 81; the nucleotide sequence of ycf1 is shown in SEQ ID NO. 82; the nucleotide sequence of ycf2 is shown in SEQ ID NO. 83; the nucleotide sequence of ycf4 is shown in SEQ ID NO. 84; and the nucleotide sequence of ycf12 is shown in SEQ ID NO. 85.
[0025] Table 1. Genes encoding chloroplasts of *Miaofengyan fern*
[0026] Note: (×2) represents the number of copies of the gene.
[0027] Example 3: Codon bias in the chloroplast genomic DNA of *Miaofengyan fern* Protein-coding sequences from the genome were extracted using Phylosuite software. Mega 7.0 software was used to analyze codon bias in protein-coding genes of the chloroplast genome, and RSCU values were calculated. An RSCU value > 1 indicates that the codon is frequently used and is the preferred codon; an RSCU value < 1 indicates low codon selectivity; and an RSCU value = 1 indicates no codon bias.
[0028] Results: Eukaryotic genomes contain 64 codons encoding 20 amino acids and 3 stop codons. The amino acid-codon correspondence exhibits degeneracy: except for tryptophan (Trp, corresponding to UGG) and methionine (Met, corresponding to AUG), which are encoded by a single codon, other amino acids can be encoded by multiple synonymous codons (e.g., alanine (Ala) corresponds to GCU, GCC, GCA, and GCG). The frequency of synonymous codon usage varies significantly between different species and even between different organelles (such as mitochondria). This non-uniform distribution is called codon usage bias and is considered an adaptive strategy developed during long-term evolution to maintain intracellular translation efficiency and stability. Therefore, in genome research, indicators such as relative synonymous codon usage (RSCU) are often used to analyze codon usage bias.
[0029] Codon bias analysis was performed on 84 unique PCGs of chloroplasts from the Miaofengyan fern, such as... Figure 3 As shown in Table 2, the codon usage of each amino acid is illustrated. Asparagine (Asn), aspartic acid (Asp), cysteine (Cys), glutamine (Gln), glutamic acid (Glu), histidine (His), lysine (Lys), phenylalanine (Phe), and tyrosine (Tyl) are encoded by only two codons, and all exhibit codon shifts. Furthermore, the RSCU values for the start codon (Met) AUG and tryptophan (Trp) UGG are both 1. Among all amino acids, alanine (Ala) shows the strongest preference for the GCU codon, arginine (Arg) for the AGA codon, and leucine (Leu) for the UUA codon, all with RSCU values of 1.51; serine (Ser) shows the weakest preference for the AGC codon, with an RSCU value of 0.53.
[0030] Table 2. Relative synonymous codon usage of each amino acid pair in the chloroplast genome of *Dryopteris yunnanensis*.
[0031]
[0032] Example 4: Repetitive sequences of chloroplast genomic DNA from *Miaofengyan fern* The DNA sequence in biological cells contains many repeating sequences, which are identical or complementary segments appearing at different locations in the genome. These repeating sequences play an important role in gene regulation. Microsatellite repeats, tandem repeats, and sporadic repeats were identified using MISA (https: / / webblast.ipk-gatersleben.de / misa / ), TRF (https: / / tandem.bu.edu / trf / trf.unix.help.html), and the REPuter web server (https: / / bibiserv.cebitec.uni-bielefeld.de / reputer / ). The results were visualized using Excel (2021) software and the Circos package.
[0033] Results: Simple sequence repeats (SSRs) are DNA fragments in the genome consisting of repeated short base units (generally no more than 6 bp). They are a common and efficient molecular marker. The diagram of repeat sequence analysis of the chloroplast genome of *Dryopteris melanogaster* is shown below. Figure 4 As shown, a total of 37 SSRs were found in the chloroplast genome of *Adiantum melanogaster*, such as... Figure 4 As shown in Figure A, this includes 22 mononucleotide repeats, 8 dinucleotide repeats, 1 trinucleotide repeat, and 6 tetranucleotide repeats. No pentamer or hexamer SSRs were detected in this chloroplast genome. Monomeric and dimeric SSRs accounted for 81.08% of the total SSRs, and cytosine (C) mononucleotide repeats accounted for 36.36% (8 repeats) of the 22 monomeric SSRs.
[0034] Tandem repeat sequences (also known as satellite DNA) are core repeat units of 2–200 base pairs, repeated tandemly multiple times. They are typically described as long, short, and medium-length, and are widely found in eukaryotic and prokaryotic genomes. The chloroplast genome of *Dryopteris yunnanensis* contains 23 tandem repeat sequences with a matching degree greater than 84% and a length between 2 and 24 bp, such as… Figure 4 As shown in B. Simultaneously, 59 pairs of scattered repeating sequences with a length greater than or equal to 30 were observed, such as... Figure 4 As shown in B, there are 21 pairs of palindromic repeats, the longest of which is 22502 bp, 23 pairs of forward repeats, the longest of which is 57 bp, 9 pairs of reverse repeats, and 6 pairs of complementary repeats.
[0035] The above specific embodiments further illustrate the purpose, technical solution and beneficial effects of this application. It should be understood that the above are only specific embodiments of this application and are not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made on the basis of the technical solution of this application should be included within the scope of protection of this application.
Claims
1. An assembled and annotated chloroplast genome of *Miaofengyan fern*, characterized in that, The chloroplast genome of *Fragaria melanogaster* is a single circular molecule; the total length of the chloroplast genome of *Fragaria melanogaster* is 149180 bp; the nucleotide sequence of the chloroplast genome of *Fragaria melanogaster* is shown in SEQ ID NO.1; the chloroplast genome of *Fragaria melanogaster* has a GC content of 43.14%.
2. The assembled and annotated chloroplast genome of *Miaofengyan fern* according to claim 1, characterized in that, The chloroplast genome of the Miaofengyan fern includes tRNA genes, rRNA genes, and a family of protein-coding genes; The number of tRNA genes is 29, the number of rRNA genes is 4, and the number of protein-coding gene families is 17.
3. The assembled and annotated chloroplast genome of *Miaofengyan fern* according to claim 2, characterized in that, The protein-coding gene family includes NADH dehydrogenase subunit genes, photosystem I subunit genes, photosystem II subunit genes, cytochrome b / f complex subunit genes, ATP synthase subunit genes, ribulose-1,5-bisphosphate carboxylase / oxygenase large subunit genes, DNA-dependent RNA polymerase genes, ribosome large subunit genes, ribosome small subunit genes, maturation enzyme genes, c-type cytochrome synthase genes, membrane protein genes, translation initiation factor genes, protease genes, acetyl-CoA-carboxylase subunit genes, prochlorophyll reductase subunit genes, and conserved open reading frame genes.
4. The assembled and annotated chloroplast genome of *Miaofengyan fern* according to claim 3, characterized in that, The NADH dehydrogenase subunit genes include ndhA, ndhB, ndhC, ndhD, ndhE, ndhF, ndhG, ndhH, ndhI, ndhJ, and ndhK; The photosystem I subunit genes include psaA, psaB, psaC, psaI, and psaJ; The photosystem II subunit genes include psbA, psbB, psbC, psbD, psbE, psbF, psbH, psbI, psbJ, psbK, psbL, psbM, psbN, psbT, psbZ, and ycf3; The cytochrome b / f complex subunit genes include petA, petB, petD, petG, petL, and petN; The ATP synthase subunit genes include atpA, atpB, atpE, atpF, atpH, and atpI; The ribulose-1,5-bisphosphate carboxylase / oxygenase large subunit gene includes rbcL; The DNA-dependent RNA polymerase genes include rpoA, rpoB, rpoC1, and rpoC2; The ribosomal large subunit genes include rpl2, rpl14, rpl16, rpl20, rpl21, rpl22, rpl23, rpl32, rpl33 and rpl36; The ribosomal small subunit genes include rps2, rps3, rps4, rps7, rps8, rps11, rps12, rps14, rps15, rps16, rps18 and rps19; The maturation enzyme gene includes matK; The c-type cytochrome synthase gene includes ccsA; The membrane protein gene includes cemA; The translation initiation factor gene includes infA; The protease gene includes clpP; The subunit gene of the acetyl-CoA-carboxylase includes accD; The protochlorophyll reductase subunit genes include chlB, chlL, and chlN; The conserved open reading frame genes include ycf1, ycf2, ycf4, and ycf12.
5. The assembled and annotated chloroplast genome of *Miaofengyan fern* according to claim 4, characterized in that, The nucleotide sequence of ndhA is shown in SEQ ID NO.2; the nucleotide sequence of ndhB is shown in SEQ ID NO.3; the nucleotide sequence of ndhC is shown in SEQ ID NO.4; the nucleotide sequence of ndhD is shown in SEQ ID NO.5; the nucleotide sequence of ndhE is shown in SEQ ID NO.6; the nucleotide sequence of ndhF is shown in SEQ ID NO.7; the nucleotide sequence of ndhG is shown in SEQ ID NO.8; the nucleotide sequence of ndhH is shown in SEQ ID NO.9; the nucleotide sequence of ndhI is shown in SEQ ID NO.10; the nucleotide sequence of ndhJ is shown in SEQ ID NO.11; and the nucleotide sequence of ndhK is shown in SEQ ID NO.
12. The nucleotide sequence of psaA is shown in SEQ ID NO.13; the nucleotide sequence of psaB is shown in SEQ ID NO.14; the nucleotide sequence of psaC is shown in SEQ ID NO.15; the nucleotide sequence of psaI is shown in SEQ ID NO.16; and the nucleotide sequence of psaJ is shown in SEQ ID NO.
17. The nucleotide sequence of psbA is shown in SEQ ID NO. 18; the nucleotide sequence of psbB is shown in SEQ ID NO. 19; the nucleotide sequence of psbC is shown in SEQ ID NO. 20; the nucleotide sequence of psbD is shown in SEQ ID NO. 21; the nucleotide sequence of psbE is shown in SEQ ID NO. 22; the nucleotide sequence of psbF is shown in SEQ ID NO. 23; the nucleotide sequence of psbH is shown in SEQ ID NO. 24; the nucleotide sequence of psbI is shown in SEQ ID NO. 25; the nucleotide sequence of psbJ is shown in SEQ ID NO. 26; the nucleotide sequence of psbK is shown in SEQ ID NO. 27; the nucleotide sequence of psbL is shown in SEQ ID NO. 28; the nucleotide sequence of psbM is shown in SEQ ID NO. 29; the nucleotide sequence of psbN is shown in SEQ ID NO. 30; the nucleotide sequence of psbT is shown in SEQ ID NO. 31; and the nucleotide sequence of psbZ is shown in SEQ ID NO.
18. As shown in NO.32; the nucleotide sequence of ycf3 is shown in SEQ ID NO.33; The nucleotide sequence of petA is shown in SEQ ID NO.34; the nucleotide sequence of petB is shown in SEQ ID NO.35; the nucleotide sequence of petD is shown in SEQ ID NO.36; the nucleotide sequence of petG is shown in SEQ ID NO.37; the nucleotide sequence of petL is shown in SEQ ID NO.38; and the nucleotide sequence of petN is shown in SEQ ID NO.
39. The nucleotide sequence of atpA is shown in SEQ ID NO. 40; the nucleotide sequence of atpB is shown in SEQ ID NO. 41; the nucleotide sequence of atpE is shown in SEQ ID NO. 42; the nucleotide sequence of atpF is shown in SEQ ID NO. 43; the nucleotide sequence of atpH is shown in SEQ ID NO. 44; the nucleotide sequence of atpI is shown in SEQ ID NO. 45; and the nucleotide sequence of rbcL is shown in SEQ ID NO.
46. The nucleotide sequence of rpoA is shown in SEQ ID NO.47; the nucleotide sequence of rpoB is shown in SEQ ID NO.48; the nucleotide sequence of rpoC1 is shown in SEQ ID NO.49; and the nucleotide sequence of rpoC2 is shown in SEQ ID NO.
50. The nucleotide sequence of rpl2 is shown in SEQ ID NO. 51; the nucleotide sequence of rpl14 is shown in SEQ ID NO. 52; the nucleotide sequence of rpl16 is shown in SEQ ID NO. 53; the nucleotide sequence of rpl20 is shown in SEQ ID NO. 54; the nucleotide sequence of rpl21 is shown in SEQ ID NO. 55; the nucleotide sequence of rpl22 is shown in SEQ ID NO. 56; the nucleotide sequence of rpl23 is shown in SEQ ID NO. 57; the nucleotide sequence of rpl32 is shown in SEQ ID NO. 58; the nucleotide sequence of rpl33 is shown in SEQ ID NO. 59; and the nucleotide sequence of rpl36 is shown in SEQ ID NO.
60. The nucleotide sequence of rps2 is shown in SEQ ID NO. 61; the nucleotide sequence of rps3 is shown in SEQ ID NO. 62; the nucleotide sequence of rps4 is shown in SEQ ID NO. 63; the nucleotide sequence of rps7 is shown in SEQ ID NO. 64; the nucleotide sequence of rps8 is shown in SEQ ID NO. 65; the nucleotide sequence of rps11 is shown in SEQ ID NO. 66; the nucleotide sequence of rps12 is shown in SEQ ID NO. 67; the nucleotide sequence of rps14 is shown in SEQ ID NO. 68; the nucleotide sequence of rps15 is shown in SEQ ID NO. 69; the nucleotide sequence of rps16 is shown in SEQ ID NO. 70; the nucleotide sequence of rps18 is shown in SEQ ID NO. 71; and the nucleotide sequence of rps19 is shown in SEQ ID NO.
72. The nucleotide sequence of matK is shown in SEQ ID NO.73; the nucleotide sequence of ccsA is shown in SEQ ID NO.74; the nucleotide sequence of cemA is shown in SEQ ID NO.75; the nucleotide sequence of infA is shown in SEQ ID NO.76; the nucleotide sequence of clpP is shown in SEQ ID NO.77; and the nucleotide sequence of accD is shown in SEQ ID NO.
78. The nucleotide sequence of chlB is shown in SEQ ID NO. 79; the nucleotide sequence of chlL is shown in SEQ ID NO. 80; the nucleotide sequence of chlN is shown in SEQ ID NO. 81; the nucleotide sequence of ycf1 is shown in SEQ ID NO. 82; the nucleotide sequence of ycf2 is shown in SEQ ID NO. 83; the nucleotide sequence of ycf4 is shown in SEQ ID NO. 84; and the nucleotide sequence of ycf12 is shown in SEQ ID NO.
85.
6. The assembled and annotated chloroplast genome of *Miaofengyan fern* according to claim 1, characterized in that, In the chloroplast genome of the Miaofengyan fern, alanine (Ala) showed the strongest preference for the GCU codon, arginine (Arg) for the AGA codon, and leucine (Leu) for the UUA codon, with an RSCU value of 1.51; while serine (Ser) showed the weakest preference for the AGC codon, with an RSCU value of 0.
53.
7. The assembled and annotated chloroplast genome of *Miaofengyan fern* according to claim 1, characterized in that, The chloroplast genome of the Miaofengyan fern includes repetitive sequences; the repetitive sequences include simple repetitive sequences, tandem repetitive sequences, and scattered repetitive sequences. The number of simple repeating sequences is 37; the number of tandem repeating sequences is 23; and the number of scattered repeating sequences is 59 pairs.
8. The assembled and annotated chloroplast genome of *Miaofengyan fern* according to claim 7, characterized in that, The simple repeating sequences include mononucleotide repeating sequences, dinucleotide repeating sequences, trinucleotide repeating sequences, and tetranucleotide repeating sequences; The number of mononucleotide repeat sequences is 22; the number of dinucleotide repeat sequences is 8; the number of trinucleotide repeat sequences is 1; and the number of tetranucleotide repeat sequences is 6.
9. The assembled and annotated chloroplast genome of *Miaofengyan fern* according to claim 7, characterized in that, The scattered repeating sequences include palindromic repeats, forward repeats, reverse repeats, and complementary repeats; The number of palindromic repetitions is 21; the number of forward repetitions is 23; the number of reverse repetitions is 9; and the number of complementary repetitions is 6. The longest length of the palindromic repeat sequence is 22502 bp; The longest sequence in the positive repeat is 57 bp.
10. The assembled and annotated chloroplast genome of *Miaofengyan fern* according to claim 1, characterized in that, The Miaofeng rock fern is derived from a plant species of Miaofeng rock fern located in Linyi City, Shandong Province, at latitude 35.5419444N and longitude 117.8558333E.