SNP (Single Nucleotide Polymorphism) molecular marker related to oil content of common camellia oleifera seed kernel and application of SNP molecular marker
By developing the SNP molecular marker Chr21_164424223 for chromosome 21 of common Camellia oleifera, the problem of rapid identification of high oil content kernels in existing technologies has been solved, enabling rapid and accurate breeding screening at the seedling stage, saving costs and time.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- RES INST OF SUBTROPICAL FORESTRY CHINESE ACAD OF FORESTRY
- Filing Date
- 2026-03-16
- Publication Date
- 2026-05-19
AI Technical Summary
Existing technologies make it difficult to quickly and accurately identify and select common camellia seeds with high oil content, resulting in long breeding cycles, high costs, and susceptibility to environmental influences.
A SNP molecular marker Chr21_164424223 located on chromosome 21 of Camellia oleifera was developed. Through PCR amplification and genotyping analysis, the oil content of the seed kernels was identified and screened at an early stage.
It enables rapid and accurate identification of kernel oil content during the seedling stage, saving breeding costs, improving selection efficiency, and shortening the breeding cycle.
Smart Images

Figure SMS_1 
Figure SMS_2 
Figure SMS_3
Abstract
Description
Technical Field
[0001] This invention relates to the field of molecular biology, and in particular to SNP molecular markers related to the oil content of common camellia seeds and their applications. Background Technology
[0002] Common tea oil ( Camellia oleifera Abel., belonging to the genus Camellia in the family Theaceae. Camellia Camellia seed oil (L.) is one of the world's four major woody oilseeds. Rich in nutrients, camellia seed oil is a high-quality edible oil with over 90% unsaturated fatty acids, primarily oleic and linoleic acids. It is also rich in vitamin E, squalene, sterols, and other nutrients, giving it high nutritional and health value.
[0003] In recent years, significant progress has been made in the genetic breeding of Camellia oleifera, with the completion of multiple Camellia oleifera genome sequencing and mapping projects, and substantial advancements in the analysis of seed oil synthesis and accumulation mechanisms. Building upon this foundation, identifying and screening molecular markers for early selection of high-oil-content germplasm, and constructing corresponding detection systems, are effective ways to accelerate the breeding process of superior high-oil-content Camellia oleifera varieties.
[0004] Compared to traditional breeding techniques, marker-assisted selection can significantly shorten the breeding cycle, with particularly pronounced advantages for economic forest breeding where fruit production is the primary objective. Effective and precise molecular markers are the foundation and key to marker-assisted breeding.
[0005] Therefore, developing molecular markers related to the oil content of common camellia kernels is of great significance for molecular marker-assisted breeding and genetic improvement of oil yield in common camellia. Summary of the Invention
[0006] To address the aforementioned technical problems, this invention develops molecular markers related to the oil content of common camellia seeds and their applications.
[0007] Since *Camellia oleifera* is a typical outcrossing species, linkage disequilibrium (LD) typically decreases rapidly within a small range, LD mapping of important traits can be performed. All transcripts of *Camellia oleifera* kernels are used as regions for marker development in this invention. Given a natural population of *Camellia oleifera* that exhibits significant genetic variation, the development of SNP molecular markers significantly correlated with variations in the oil content of *Camellia oleifera* kernels can be effectively carried out.
[0008] The development process of the SNP molecular markers of this invention is basically as follows: (1) Collect camellia germplasm resources extensively in the entire distribution area of common camellia and establish a natural population of common camellia with widely separated kernel oil content.
[0009] (2) Fully mature seeds of 221 common camellia oleifera germplasms from natural populations were collected, and the oil content of the kernels was determined by Soxhlet extraction.
[0010] (3) Kernels from 221 common camellia oleifera plants in a natural population during the high-speed oil synthesis period were collected. Total RNA was extracted using the RNAprepPure Polysaccharide and Polyphenol Plant Total RNA Extraction Kit (centrifuge column type, TIANGEN kit code no. DP441). A cDNA library was constructed for each sample and analyzed using Illumina HiSeq. TM Second-generation transcriptome sequencing was performed on the 4000 platform.
[0011] (4) Using the genome of common Camellia oleifera 'Changlin 40' (Zhu, H., Wang, F., Xu, Z., Wang, G., Hu, L., Cheng, J., Ge, X., Liu, J., Chen, W., Li, Q., Xue, F., Liu, F., Li, W., Wu, L., Cheng, X., Tang, X., Yang, C., Lindsey, K., Zhang, X., Ding, F., Hu, H., Hu, X. and Jin, S. (2024) The complex hexaploid oil-Camellia genometraces back its phylogenomic history and multi-omics analysis of Camellia oilbiosynthesis. Plant Biotechnol. J., https: / / doi.org / 10.1111 / pbi.14412.) as the reference sequence, the SNP sites of the transcriptome sequences of 221 samples obtained in (3) were analyzed by multiple sequence alignment.
[0012] (5) SNP data were strictly filtered according to the following principles: each locus had only 2 alleles; genotype deletion rate ≤20%; minimum allele frequency ≥5%; SNP quality value ≥100; number of homozygous genotype samples exceeded 10; heterozygous genotype rate ≤70%. The software bcftools v1.9 (http: / / www.htslib.org / doc / bcftools.html) used in the process is publicly available and free.
[0013] (6) Input the genotype data of the population into GCTA v1.25.2 (Jian Y, S Hong L, Goddard ME, Visscher PM, 2011. GCTA: a tool for genome-wide complex trait analysis. American Journal of Human Genetics 88, 76-82.) software to perform principal component analysis (PCA).
[0014] (7) Input the genotype data of the population, the data of the first 10 principal components (PC), the phenotypic data of the oil content of the kernel, and the Kinship matrix data into the IIIVmrMLM software, and use the unified mixed linear model (MLM) method to analyze the linkage disequilibrium of SNP molecular markers and the oil content trait of common camellia kernel.
[0015] Using the above technical measures, this invention ultimately obtained a result that was highly significantly correlated with the oil content of ordinary camellia seed kernels ( P <10 -5 The SNP molecular marker Chr21_164424223 located on chromosome 21 (see Table 1 for details) contributes 3.54% to phenotypic variation.
[0016] Table 1 SNP molecular marker information
[0017] Based on this, the present invention proposes the following technical solution.
[0018] In a first aspect, the present invention provides an SNP molecular marker related to the oil content of common camellia kernels. The SNP molecular marker is located at base 164424223 on chromosome 21 of common camellia, with a base polymorphism of C / T and a genome version number of Changlin 40 V1.0.
[0019] In some implementations, the SNP molecular marker site has a genotype of C / C, corresponding to high oil content; a genotype of C / T, corresponding to candidate high oil content; and a genotype of T / T, corresponding to low oil content.
[0020] In some embodiments, the SNP molecular marker is located at position 103 of the nucleotide sequence shown in SEQ ID NO.1, and the base polymorphism is C / T.
[0021] SEQ ID NO.1: ATGGACTGGATTCTTGAATAGCAAATAGTGAAGCTCAAACAAGGAGATGAAAGGAGAATCATCTGGGGCTGATCATTGGGATCTCCATAGGGGTGGTAATTGGCGTGCTTTTGGCTATATTTGCACTGTTTTGCGTTAGGTACCATAGGAGACATTCACAGATAGGGAATAGCAGTTCTCGGAGGGCGGCAACTATCCCCATTCGTACTAACGGTGCTGATTCTTGTATAGTATTATCAGA Secondly, the present invention provides primers for amplifying the SNP molecular marker.
[0022] Preferably, the primers are those shown in SEQ ID NO.2 and SEQ ID NO.3.
[0023] SEQ ID NO.2: ATGGACTGGATTCTTGAATAGCA SEQ ID NO.3: ATCAGCACCGTTAGTACGAATG The SNP molecular markers related to the oil content of common Camellia oleifera kernels of the present invention can be obtained by PCR amplification using primers as shown in SEQ ID NO.2-3 with common Camellia oleifera genomic DNA as a template.
[0024] Thirdly, the present invention provides a kit for identifying the oil content of common camellia seed kernels, which contains the aforementioned primers.
[0025] In some implementations, the kit is a PCR kit.
[0026] Fourthly, the present invention provides the use of the SNP molecular marker, the primer, or the kit in at least one of the following aspects: (1) Identification of the oil content phenotype of common camellia seeds; (2) Identification, improvement or molecular marker-assisted breeding of common camellia oleifera germplasm resources; (3) Early prediction of oil content in common camellia seeds; (4) Screening common camellia with high oil content.
[0027] Preferably, the target trait for the identification, improvement, or molecular marker-assisted breeding of the common camellia oleifera germplasm resource is the oil content of the common camellia oleifera kernel.
[0028] Fifthly, the present invention provides a method for identifying the oil content of common camellia seed kernels, comprising: Using the genomic DNA of the common camellia oleifera seed as a template, PCR amplification was performed using primers that amplified the SNP molecular markers. The genotype of the common camellia oleifera seed kernel oil content phenotype was identified based on the genotype of the SNP molecular marker loci.
[0029] Preferably, the identification method includes: (1) Extract genomic DNA from the common camellia oleifera to be identified; (2) Using the genomic DNA as a template, perform PCR amplification using the primers; (3) Analyze the genotype of the SNP molecular markers in the PCR amplification products, and determine the phenotype of the oil content of the kernel of the common camellia to be identified based on the genotype.
[0030] In some implementations, the PCR amplification reaction program is as follows: 94-95℃, 3-5 min; 94-95℃, 15-30 s, 65-69℃, 40-60 s, 38-45 cycles; 67-70℃, 3-6 min.
[0031] The preferred reaction procedure for the PCR amplification reaction is as follows: 95℃, 3 min, 1 cycle for pre-denaturation; 95℃, 15 s for denaturation; 68℃, 45 s for extension, 40 cycles; 68℃, 5 min, 1 cycle for complete extension.
[0032] In some implementations, the SNP molecular marker site has a genotype of C / C, corresponding to high oil content; a genotype of C / T, corresponding to candidate high oil content; and a genotype of T / T, corresponding to low oil content.
[0033] In some implementations, the Camellia oleifera to be identified can be any Camellia oleifera breeding material, including individuals from natural populations or sexual populations.
[0034] Preferably, the genomic DNA of common Camellia oleifera is extracted using the TaKaRa MiniBEST Plant Genomic DNA Extraction Kit (TaKaRa, Dalian, China).
[0035] In some implementations, after the PCR amplification, the obtained PCR product is detected and recovered by agarose gel electrophoresis.
[0036] In some embodiments, the concentration of agarose gel in the agarose gel electrophoresis is 1.2%. Gel recovery is performed using the AxyPrep DNA Gel Recovery Kit (AxyGEN, Code No. AP-GX-50).
[0037] In some implementations, the genotype of the SNP molecular marker can be analyzed using conventional techniques in the art, such as sequencing, for example, sequencing using SEQ ID NO.2-3 as sequencing primers.
[0038] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention develops a single SNP molecular marker highly correlated with the oil content of common Camellia oleifera kernels, which can explain 3.54% of the variance in the oil content phenotypic variance. In conventional selection breeding of common Camellia oleifera, the identification of kernel oil content requires 5-6 years of seedling cultivation, which is time-consuming and labor-intensive. In contrast, the SNP molecular marker of this invention has a clear location, a convenient and rapid detection method, and can be identified and screened at the seedling stage, greatly saving production and breeding costs, improving selection efficiency, and is unaffected by environmental factors, making it more targeted and requiring less work. This invention can accelerate the breeding process of common Camellia oleifera with high kernel oil content and has broad application prospects. Detailed Implementation
[0039] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention. The oil content described in this invention is the oil content of the seed kernel.
[0040] The 221 natural population materials used in the following examples, all of which were common Camellia oleifera clones, were collected and evaluated by the Woody Oilseed Breeding and Cultivation Research Group of the Subtropical Forestry Research Institute of the Chinese Academy of Forestry and are preserved in the germplasm resource nursery of Dongfanghong Forest Farm, Wucheng District, Jinhua, Zhejiang Province.
[0041] Unless otherwise specified, all techniques or conditions used in the following examples are conventional methods or performed in accordance with techniques or conditions described in the literature in this field, or in accordance with the product instructions. Reagents and instruments used without specified manufacturers are all conventional products that can be purchased from legitimate channels.
[0042] This invention relates to molecular biology experiments. Unless otherwise specified, reference can be made to the book *Molecular Cloning* (J. Sambrook, E.F. Fritsch, and T. Maniatis, Science Press, 1994). This book and its subsequent editions are the most commonly used and guiding reference books for those skilled in the art when performing experiments related to molecular biology. In addition, depending on the experimental purpose, those skilled in the art complete the corresponding experiments under the guidance of the operating manuals accompanying various commercial reagent kits or entrust them to specialized companies, such as primer synthesis and gene sequencing.
[0043] Example 1: Construction and trait determination of an isolated population of oil content from common Camellia oleifera seed kernels. In this embodiment, 221 natural populations of common Camellia oleifera germplasm resources were collected from a nursery. Their origins covered most of my country's main Camellia oleifera producing areas, including Zhejiang, Hunan, Jiangxi, Guangxi Zhuang Autonomous Region, Fujian, and Guangdong provinces. After the fruits of all 221 individuals were fully mature (5% of the fruits split open), seeds were collected, and the oil content of the kernels was determined using Soxhlet extraction. The operational steps are as follows: (1) Prepare medium-speed filter paper packs and place them in an aluminum box. Dry them at 105°C until constant weight is achieved. Record the weight of the aluminum box and the filter paper packs. W 1 ).
[0044] (2) Remove the hard seed coat from an appropriate amount of common camellia seeds, dry them at 105℃ to constant weight, crush them with a pulverizer, pack them into a filter paper bag and wrap them up, and record the total weight of the aluminum box, filter paper bag and sample. W 2 ).
[0045] (3) Using a Swiss Buchi Soxhlet extractor B-811LSV, the weighed sample filter paper package was placed in an extraction flask, and about 100 ml of petroleum ether was added. Extraction was carried out for 6 hours, and the petroleum ether was recovered. The filter paper package (containing residue) was placed in an aluminum box and dried at 105°C until constant weight was achieved. The weights of the aluminum box, filter paper package, and residue were recorded. W 3 ).
[0046] Kernel oil content = [( W 2 - W 3 ) / ( W 2 - W 1 )]×100% The results of the oil content determination of common camellia seeds showed that the oil content of natural population seeds exhibited a normal distribution, indicating that this trait has quantitative characteristics.
[0047] Example 2: Transcriptome sequencing and polymorphic site identification of seed kernels during the rapid lipid synthesis period 1. Total RNA extraction from the kernels of 221 common Camellia oleifera clones during the high-speed oil synthesis period: Total RNA was extracted from immature kernels of various common Camellia oleifera clones using the RNAprep Pure Polysaccharide and Polyphenol Plant Total RNA Extraction Kit (centrifugal column type, TIANGEN Kit Code No. DP441).
[0048] 2. Transcriptome sequencing: Total RNA from each sample was tested for purity and concentration, and ribosomal RNA was removed to maximize the retention of all coding RNA and ncRNA. The resulting RNA was randomly fragmented into short segments, and cDNA first-strand was synthesized using the fragmented RNA as a template with six-base random hexamers. Then, buffer, dNTPs (dUTP instead of dTTP), RNase H, and DNA polymerase I were added to synthesize cDNA second-strand. The cDNA was purified using a QiaQuick PCR kit and eluted with EB buffer. End repair, addition of base A, and the addition of sequencing adapters were performed, followed by degradation of the second strand by UNG (Uracil-N-Glycosylase). Fragment size selection was performed using agarose gel electrophoresis, followed by PCR amplification. Finally, the constructed sequencing library was analyzed using Illumina HiSeq. TM Second-generation transcriptome sequencing was performed on the 4000 platform.
[0049] 3. Polymorphic site identification: To ensure data quality, the clean reads obtained after initial filtering are further filtered more rigorously to obtain high-quality clean reads for subsequent information analysis. The filtering steps are as follows: (1) Remove reads containing connectors; (2) Remove reads that are all A bases; (3) Remove reads containing more than 10% N; (4) Remove low-quality reads (the number of bases with a quality value of Q≤20 accounts for more than 50% of the total reads).
[0050] Tophat v2.1.1 (Trapnell C, Roberts A, Goff L, et al., 2012. Differential gene and transcript expression analysis of RNA-seq experiments with TopHat and Cufflinks. Nature protocols 7, 562-78.) software was used to align high-quality reads from each sample to a hexaploid reference genome sequence (Zhu, H., Wang, F., Xu, Z., Wang, G., Hu, L., Cheng, J., Ge, X., Liu, J., Chen, W., Li, Q., Xue, F., Liu, F., Li, W., Wu, L., Cheng, X., Tang, X., Yang, C., Lindsey, K., Zhang, X., Ding, F., Hu, H., Hu, X. and Jin, S. (2024) The complex hexaploid oil-Camellia genometraces back its phylogenomic history and multi-omics analysis of Camellia oilbiosynthesis. Plant Biotechnol. J., https: / / doi.org / 10.1111 / pbi.14412.). Unaligned sequences were removed, and the remaining sequences were used to identify SNP sites using bcftools v1.9 software (http: / / www.htslib.org / doc / bcftools.html). The identified SNP sites underwent rigorous filtering to obtain high-quality SNP data. The filtering criteria are as follows: (1) There are only 2 alleles at each locus; (2) Genotype deletion rate ≤ 20%; (3) Minimum allele frequency (MAF) ≥ 5%; (4) SNP quality value ≥ 100; (5) The number of homozygous genotype samples is greater than 10; (6) The heterozygous genotype sample rate is ≤70%.
[0051] Example 3: Screening of SNP molecular markers related to oil content in common camellia seed kernels 1. Group structure analysis: Principal component analysis (PCA) was performed on the natural population of Camellia oleifera using GCTA v1.25.2 (Jian Y, S Hong L, Goddard ME, Visscher PM, 2011.GCTA: a tool for genome-wide complex trait analysis. American Journal of Human Genetics 88, 76-82.). The first 10 principal components (PCs) were used as fixed effects for subsequent association analysis (Table 2).
[0052] Table 2. Top 10 PC values of selected individuals in a natural population.
[0053] 2. Association Analysis: All SNP locus data, the top 10 PC values, phenotypic data (see Example 1), and Kinship matrix data were imported into IIIVmrMLM software. MLM analysis was used to analyze the linkage disequilibrium between SNPs and kernel oil content. SNP molecular markers significantly associated with kernel oil content were screened. After multiple validation corrections, one locus with a highly significant association with kernel oil content was detected. P <10 -5 (Table 1) This locus is located in the intergenic region of chromosome 21 and contributes 3.54% to the difference in oil content (Table 1).
[0054] Example 4: Application of the SNP molecular markers of the present invention in high-oil breeding of common camellia oleifera (1) Select a common Camellia oleifera hybrid F1 generation family as material (the maternal parent is 'Changlin 53' and the paternal parent is 'Changlin 40', both of which are nationally approved improved varieties with improved variety numbers 'Guo S-SC-CO-012-2008' and 'Guo S-SC-CO-011-2008' respectively), and collect young leaves to extract total genomic DNA according to the method in Example 2.
[0055] (2) Genomic DNA was amplified by PCR using the primers shown in SEQ ID NO.2-3. The reaction system is shown in Table 3.
[0056] Table 3 Reaction System
[0057] The PCR amplification program was as follows: 95℃, 3 min, 1 cycle for pre-denaturation; 95℃, 15 s for denaturation, 68℃, 45 s for extension, 40 cycles; 68℃, 5 min, 1 cycle for complete extension.
[0058] (3) PCR amplification products were subjected to gel detection, purification, recovery, sequencing, and genotyping. Gel detection and purification were performed according to the instructions of the AxyPrep DNA Gel Recovery Kit (AxyGEN, Code No. AP-GX-50). DNA was recovered from the gel, and the corresponding amplification primers were used as sequencing primers. The nucleotide sequence of the amplification products was determined by Sanger sequencing, and the genotype of each SNP site on the sequencing peak map was interpreted using Chromas software.
[0059] (4) Identify the genotype of the Chr21_164424223 locus in all individuals. Compare the relationship between the genotype of this locus and the oil content. If the genotype is C / C, the ordinary camellia is a high-oil camellia; if the genotype is C / T, the ordinary camellia is a candidate for high-oil camellia; if the genotype is T / T, the ordinary camellia is a low-oil camellia.
[0060] (5) Collect fully mature seeds from all F1 generation individuals and determine the oil content of their kernels according to the method in Example 1.
[0061] The results are shown in Table 4. Among the individuals with the high oil content genotype at Chr21_164424223, 77.27% of the individuals had a kernel oil content higher than the population average (38.98%). Among the individuals with the candidate high oil content genotype, 50% of the individuals had a kernel oil content higher than the population average (38.98%), while the kernel oil content of individuals with the low oil content genotype was lower than the population average.
[0062] This demonstrates that the SNP molecular markers of the present invention are effective in assisting the selection of common camellia with high oil content. They can be used for early identification or auxiliary identification, which can greatly save production costs, improve selection efficiency, and accelerate the breeding process of high-oil common camellia.
[0063] Table 4. Kernel oil content and genotype data of F1 individual plants
[0064] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A SNP molecular marker associated with the oil content of common camellia seed kernels, characterized in that, The SNP molecular marker is located at base 164424223 on chromosome 21 of Camellia oleifera, with a base polymorphism of C / T and a genome version number of Changlin 40V1.
0.
2. The SNP molecular marker according to claim 1, characterized in that, The SNP molecular marker sites have a genotype of C / C, corresponding to high oil content; a genotype of C / T, corresponding to candidate high oil content; and a genotype of T / T, corresponding to low oil content.
3. The SNP molecular marker according to claim 2, characterized in that, The SNP molecular marker is located at position 103 of the nucleotide sequence shown in SEQ ID NO.1, and the polymorphism of the base is C / T.
4. Primers for amplifying the SNP molecular marker as described in any one of claims 1 to 3.
5. The primer according to claim 4, characterized in that, The primers are shown in SEQ ID NO.2 and SEQ ID NO.
3.
6. A reagent kit for identifying the oil content of common camellia seed kernels, characterized in that, It contains the primers described in claim 4 or 5.
7. The use of the SNP molecular marker according to any one of claims 1 to 3, or the primer according to claim 4 or 5, or the kit according to claim 6, in at least one of the following aspects: (1) Identification of the oil content phenotype of common camellia seeds; (2) Identification, improvement or molecular marker-assisted breeding of common camellia oleifera germplasm resources; (3) Early prediction of oil content in common camellia seeds; (4) Screening common camellia with high oil content.
8. A method for identifying the oil content of common camellia seed kernels, characterized in that, include: Using the genomic DNA of the common camellia oleifera to be tested as a template, PCR amplification reaction was performed using primers for amplifying the SNP molecular markers described in any one of claims 1 to 3, and the genotype of the common camellia oleifera seed kernel oil content phenotype was identified based on the genotype of the SNP molecular marker loci.
9. The identification method according to claim 8, characterized in that, The PCR amplification reaction program is as follows: 94~95℃, 3~5min; 94~95℃, 15~30s, 65~69℃, 40~60s, 38~45 cycles; 67~70℃, 3~6min.
10. The identification method according to claim 8 or 9, characterized in that, The SNP molecular marker sites have a genotype of C / C, corresponding to high oil content; a genotype of C / T, corresponding to candidate high oil content; and a genotype of T / T, corresponding to low oil content.