A method for identifying oil-tea camellia varieties

By constructing a DNA fingerprint map of Camellia oleifera, using SNP and Indel marker technology, and combining PCA, Structure and phylogenetic tree analysis, the problem of imperfect Camellia oleifera variety identification system was solved, rapid and accurate variety identification was achieved, and the interests of growers and market regulations were protected.

CN117457075BActive Publication Date: 2025-09-26SICHUAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311239808.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-25
Publication Date
2025-09-26
Estimated Expiration
2043-09-25

AI Technical Summary

Technical Problem

Existing technologies lack a clear, convenient and fast system for identifying oil tea varieties, which leads to illegal vendors passing inferior products off as good ones, damaging the interests of growers and confusing the market identification and naming systems.

Method used

By constructing a DNA fingerprint of Camellia oleifera, using SNP and Indel marker technology, and combining PCA, Structure and phylogenetic tree analysis, a method for identifying Camellia oleifera varieties was developed, including DNA extraction, sequencing, data quality control, fingerprint library construction and variety identification.

Benefits of technology

It has achieved accurate identification of tea oil varieties, protected the interests of growers, standardized the market, and provided a fast and accurate means of variety identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117457075B_ABST
    Figure CN117457075B_ABST
Patent Text Reader

Abstract

The present invention provides a method for identifying oil-tea camellia varieties, belonging to the field of bioinformatics technology. The present invention utilizes SLAF-seq technology to conduct preliminary development of molecular markers for oil-tea camellia materials, obtains SNPs polymorphisms of oil-tea camellias, and analyzes their phylogenetic evolution and genetic diversity. The present invention constructs DNA fingerprints for major commercially available varieties and develops related identification technologies. This method can identify the major oil-tea camellia varieties currently on the market and can be further promoted and utilized in areas such as market circulation and cultivation management of oil-tea camellias, thereby safeguarding the interests of the majority of growers and regulating the market.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of bioinformatics, and in particular to a method for identifying oil-tea camellia varieties. Background Art

[0002] Camellia oleifera Abel. refers to a group of plants in the Theaceae family and genus Camellia whose seeds are high in oil and have economic value for cultivation. Camellia oleifera has a long history of cultivation, and through long-term artificial and natural breeding, a rich collection of camellia germplasm resources has been developed. Germplasm resources are the foundation of all breeding work and the guarantee of plant life and reproduction. However, there are still deficiencies in the precise identification and utilization of camellia germplasm resources. Constructing DNA fingerprints of camellia varieties can overcome the limitations of morphological identification alone and is crucial for the precise identification of germplasm resources and gene discovery. Currently, with the rapid development of molecular biotechnology, DNA molecular marker technology is becoming increasingly mature and has been widely used in a variety of crops.

[0003] Molecular markers are genetic markers based on nucleotide sequence variations within genetic material between individuals, and are a direct reflection of genetic polymorphism at the DNA level. The first generation of molecular marker technology is represented by RFLP marker technology, which detects the insertion and deletion of genomic bases through hybridization of specific probes, thereby comparing the differences (i.e. polymorphism) at the DNA level of different varieties (individuals). The comparison of multiple probes can establish the evolutionary and taxonomic relationships of organisms. The second generation of molecular marker technology is a molecular marker developed based on PCR technology according to the polymorphism of microsatellite sequences in the population, including SSR, ISSR and other molecular marker technologies. The third generation of molecular markers is a new generation of molecular marker technology based on high-throughput sequencing, mainly including SNPs molecular markers, Indel, SV marker technology, etc.

[0004] With the continuous development of the oil-tea ecosystem, farmers' demand for oil-tea cultivation is also increasing. However, there is currently no clear, convenient, and rapid variety identification system on the market. Unscrupulous vendors exploit this to pass off inferior products as high-quality ones, severely damaging the interests of growers and creating significant difficulties in standardizing the identification and naming system for oil-tea varieties. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for identifying oil-tea camellia varieties.

[0006] In order to achieve the above-mentioned object of the invention, the present invention provides the following technical solutions:

[0007] The present invention provides a method for identifying oil-tea camellia varieties, comprising the following steps:

[0008] (1) Take leaves of known oil-tea tea varieties, extract DNA to construct a library, and sequence to obtain raw sequencing data;

[0009] (2) Perform data quality control and filtering on the original sequencing data, use the Camellia oleifera individual as the reference genome, construct an index for the reference genome, align the sequence to the reference genome, and obtain the alignment file;

[0010] (3) performing SNP and Indel identification on the comparison file described in step (2) to obtain the stored variation information of each individual, and merging and typing the stored variation information;

[0011] (4) The sites within 5 bp before and after the SNP were filtered out, and Qual By Depth, RMS Mapping Quality, and Fisher Strand were used to filter the missing rate and minor allele frequency to obtain the filtered genomic SNP data;

[0012] (5) Constructing a Camellia oleifera fingerprint library based on the genomic SNP site data after filtering in step (4);

[0013] (6) Take the oil-tea camellia to be tested, repeat steps (1) to (5), and obtain the data analysis results after data analysis;

[0014] The oil-tea camellia variety is identified by combining the analysis results of step (6) and the fingerprint library of step (5).

[0015] Preferably, the data quality control filtering in step (2) is specifically as follows: each read is trimmed according to the quality value, a window is moved from the 3' end value to the beginning of the read, and bases with an average quality value of less than 20 in the window are removed; a window is moved from the 5' end to the end of the read, and bases with an average quality value of less than 20 in the window are removed, the average quality value of the bases in the sliding window is set to 20, and reads with a length of less than 36 are filtered out.

[0016] Preferably, the reference genes are indexed using BWA software and SAMtools software.

[0017] Preferably, step (4) filters out reads with Qual By Depth < 2.0, RMS Mapping Quality < 40.0, and QUAL < 30.0, and uses Fisher's exact test to determine whether the current variant has a strand-specific tendency, i.e., FS > 60.0.

[0018] Preferably, the missing rate filtering in step (4) is specifically as follows: input 0.5 at --max-missing for filtering; the allele frequency filtering is specifically as follows: input 0.05 at --maf for filtering.

[0019] Preferably, before step (5) constructing the Camellia oleifera fingerprint library, PCA analysis, Structure analysis, phylogenetic tree analysis and ancestral component analysis are performed on the SNP data.

[0020] Preferably, PCA analysis is performed using PLINK software to obtain three files ending with bim, nosex, and bed; PCA calculation is then performed using PLINK software to obtain two files ending with eigenval and eigenvec, which are visualized using the ggplot2 package of the R language; Structure analysis is performed using Admixturev1.30 software, and a Q file containing ancestral information and a log file containing the k value, i.e., CVerror, are obtained from VCF file 1. The optimal K value is found by drawing a k-CV curve, and the Q file is input into the R language ggplot2 package to draw a structure diagram.

[0021] Preferably, MEGA software is used to calculate and draw the NJ evolutionary tree, the initial value of bootstrap is set to 1000 times, FigTreev1.44 is used to modify and beautify the NJ evolutionary tree, and the coloring of each branch of the NJ tree is drawn by the result of Admixture obtained by the structure analysis.

[0022] By adopting the above technical solution, the present invention has the following beneficial effects: the present invention constructs a DNA fingerprint map for the main commercially available varieties and develops relevant identification technologies, which can identify the main oil-tea varieties currently available on the market, and can be further promoted and utilized in the fields of market circulation, cultivation management, etc. of oil-tea, so as to safeguard the interests of the majority of growers and regulate the market. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 This is the SLAF test flow chart;

[0024] Figure 2 Flowchart for bioinformatics data analysis;

[0025] Figure 3 This is the NJ evolutionary tree of the oil-tea camellia varieties in Example 1;

[0026] Figure 4 The PCA analysis results of the camellia varieties in Example 1 are shown below;

[0027] Figure 5This is the population structure analysis of Camellia oleifera in Example 1 (K=2-9);

[0028] Figure 6 This is the population structure analysis of Camellia oleifera in Example 1 (K=10-18);

[0029] Figure 7 For example 1, the core germplasm screening of Camellia oleifera was performed;

[0030] Figure 8 The fingerprints of 26 Camellia oleifera strains in Example 1 are shown in FIG.

[0031] Figure 9 This is a comparison table of DNA fingerprints between varieties in Example 1;

[0032] Figure 10 The physical DNA fingerprints of the 25 oil-tea camellia varieties in Example 1;

[0033] Figure 11 is the NJ evolutionary tree of the 13 Camellia oleifera strains in Example 2;

[0034] Figure 12 The PCA analysis results of 13 Camellia oleifera strains in Example 2 are shown;

[0035] Figure 13 This is the population structure analysis of 13 oil-tea camellia plants in Example 2;

[0036] Figure 14 This is the NJ evolutionary tree of the five Camellia oleifera strains in Example 3;

[0037] Figure 15 The PCA analysis results of the five camellia oil plants in Example 3 are shown below;

[0038] Figure 16 This is the population structure analysis of the five oil-tea camellia plants in Example 3;

[0039] Figure 17 is the NJ phylogenetic tree of the 13 Camellia oleifera strains in Example 4;

[0040] Figure 18 The PCA analysis results of 13 Camellia oleifera strains in Example 4 are shown below;

[0041] Figure 19 This is the population structure analysis of 13 Camellia oleifera plants in Example 4;

[0042] Figure 20 is the NJ evolutionary tree of the five Camellia oleifera strains in Example 5;

[0043] Figure 21 The PCA analysis results of the five camellia oil plants in Example 5 are shown below;

[0044] Figure 22This is the population structure analysis of the five camellia oil plants in Example 5;

[0045] Figure 23 is the NJ phylogenetic tree of the five Camellia oleifera strains in Example 6;

[0046] Figure 24 The PCA analysis results of the five camellia oil plants in Example 6 are shown below;

[0047] Figure 25 This is the population structure analysis of the five camellia oil plants in Example 6;

[0048] Figure 26 This is the NJ evolutionary tree of the two Camellia oleifera varieties in Example 7;

[0049] Figure 27 The PCA analysis results of the two camellia oleifera varieties in Example 7 are shown below;

[0050] Figure 28 This is the population structure analysis of the two Camellia oleifera varieties in Example 7;

[0051] Figure 29 This is the NJ evolutionary tree of the four Camellia oleifera varieties in Example 8;

[0052] Figure 30 The PCA analysis results of the four camellia oleifera varieties in Example 8 are shown below;

[0053] Figure 31 This is the population structure analysis of the four oil-tea camellia varieties in Example 8. DETAILED DESCRIPTION

[0054] The present invention provides a method for identifying oil-tea camellia varieties, comprising the following steps:

[0055] (1) Take leaves of known oil-tea tea varieties, extract DNA to construct a library, and sequence to obtain raw sequencing data;

[0056] (2) Perform data quality control and filtering on the original sequencing data, use the Camellia oleifera individual as the reference genome, construct an index for the reference genome, align the sequence to the reference genome, and obtain the alignment file;

[0057] (3) performing SNP and Indel identification on the comparison file described in step (2) to obtain the stored variation information of each individual, and merging and typing the stored variation information;

[0058] (4) The sites within 5 bp before and after the SNP were filtered out, and Qual By Depth, RMS Mapping Quality, and Fisher Strand were used to filter the missing rate and minor allele frequency to obtain the filtered genomic SNP data;

[0059] (5) Constructing a Camellia oleifera fingerprint library based on the genomic SNP site data after filtering in step (4);

[0060] (6) Take the oil-tea camellia to be tested, repeat steps (1) to (5), and obtain the data analysis results after data analysis;

[0061] The oil-tea camellia variety is identified by combining the analysis results of step (6) and the fingerprint library of step (5).

[0062] In the present invention, it is preferred to collect healthy, clean leaves that are free of pests and diseases, with young, tender green leaves being the most preferred. The fresh weight of leaves collected from each plant is preferably 4-6g, more preferably 5g. Dirt and dust are removed from the surface of the leaves. Fresh leaves from the same plant are placed in a ziplock bag, which is then sealed with silica gel and stored for later use.

[0063] The method extracts leaf DNA, constructs a library, selects the most suitable enzyme digestion scheme, performs enzyme digestion on the genomic DNA of each qualified sample, selects RsaI and HaeIII enzymes to digest the DNA sample and recover enzyme-digested fragments of 414 to 464 bp (SLAF tags), performs 3'-end A addition treatment on the recovered enzyme-digested fragments, connects to dual-index sequencing adapters, performs PCR amplification, purifies, mixes samples, cuts gel to select target fragments, and performs PE126 bp sequencing using Illumina HiSeq after the library passes quality inspection.

[0064] The present invention preferably uses High-quality DNA was extracted using a QIAGEN Genomic (QIAGEN, Dusseldorf, Germany) kit. The principles of the enzyme digestion protocol were as follows: (1) the proportion of enzyme-digested fragments located in repetitive sequences was as low as possible; (2) the enzyme-digested fragments were distributed as evenly as possible across the genome; (3) the length of the enzyme-digested fragments was consistent with the specific experimental system; and (4) the number of enzyme-digested fragments (SLAF tags) obtained ultimately met the expected number of tags.

[0065] Nipponbare rice (Oryza sativa L. japonica) was used as a control to evaluate the digestion efficiency of RsaI and HaeIII, thereby determining the accuracy and effectiveness of the experimental process. Reads generated by sequencing each sample were clustered based on sequence similarity. Reads clustered together were derived from the same SLAF fragment (SLAF tag). A SLAF tag with sequence differences (SNPs and Indels) between samples was defined as a polymorphic SLAF tag. SLAF tags on each Camellia oleifera chromosome were counted as raw data.

[0066] After obtaining the raw data, we first performed quality control filtering on the raw sequencing data using fastp v0.2011 software. Each read was trimmed based on its quality score. A sliding window was run from the 3' end of the read to the beginning, removing bases with an average quality score of less than 20 (-320). A sliding window was run from the 5' end of the read to the end, removing bases with an average quality score of less than 20 (-520). The sliding window was then set to a base average quality score tail of 20 (-M20). Furthermore, a length filter was performed, discarding reads with a length of less than 36 (-136).

[0067] Camellia oleifera individuals were used as the reference genome, and the reference genome was indexed using SAMtools v1.12 and BWA v.0717-r1188 software. Sequences were aligned to the reference genome using BWA software. The input parameter (-R) was set to the reads header, which was enclosed in quotation marks. Because the same sample may include multiple sequencing results, the alignment results from different lanes, different libraries, or different samples were merged into the same file for processing. RG was used to mark and distinguish them. At the same time, the number of threads (-t20) was set to improve the alignment efficiency. The resulting SAM was converted to bam format using SAMtools software and sorted to obtain the alignment file.

[0068] Picard v1.119 software was used to mark duplicates to remove PCR duplicate sequences and sequence duplications caused by high-density flowcells. Since the basic unit of the Illumina sequencer is the flowcell, the sequencing reaction occurs and proceeds on the flowcell. The high-density flowcell significantly improves the sequencing throughput, which will cause the problem of duplicate sequence reading.

[0069] SAMtools software was used to construct an index for the bam format files that had been sorted to facilitate subsequent searches for variant sites.

[0070] Using GTAKv4.12 software, the BAM-formatted alignment file obtained above was input to identify variants, search for SNPs and indels, and generate a VCF file storing variant information for each individual. The generated individual VCF files were then merged and typed. The Variant Call Format (VCF) is a file format used to record variant information, primarily consisting of an annotation section beginning with a "#" and a body section without a "#" prefix. ALT represents the variant allele; if multiple alleles exist, they are separated by commas. The base type and base number here refer to the individual base type number for SNPs.

[0071] SNPs were marked and obtained using GATK software (-select-type SNPs) and filtered. Bcftools was then used to filter out sites 5 bp before and after the SNP. Qual By Depth, RMS Mapping Quality, and Fisher Strand were then used to eliminate the influence of sequencing depth and determine the true quality value of the site. Missing rate and minor allele frequency were filtered to obtain filtered genomic SNP data. The filtering filtered out reads with Qual By Depth < 2.0, RMS Mapping Quality < 40.0, and QUAL < 30.0. Fisher's exact test was used to determine whether the current variant had a strand-specific tendency, i.e., FS > 60.0. The missing rate filter was specifically configured as follows: 0.5 was entered in --max-missing for filtering. The allele frequency filter was specifically configured as follows: 0.05 was entered in --maf for filtering to obtain filtered genomic SNP data.

[0072] Sequencing depth (Qual By Depth) represents the ratio of the total number of bases sequenced to the genome size and is used as an indicator to evaluate the sequencing quantity;

[0073] Genotype Quality (GQ) represents the possibility of the genotype existing at the site. The larger the value, the greater the possibility of the genotype.

[0074] Minor allele frequency (MAF) refers to the frequency of the second most common genotype in a given population.

[0075] The filtered genomic SNP data were analyzed and fingerprints were generated. The data analysis included principal component analysis (PCA), structure analysis, phylogenetic tree analysis, and ancestral component analysis.

[0076] PCA analysis is a method of data dimensionality reduction. For a large genome sequencing sample, the number of SNPs identified is huge and the information is redundant. The process of PCA analysis is to extract key information from the SNP sites in order to effectively characterize and distinguish different samples through a small number of variables. In population genetics, different subgroups are clustered mainly based on the degree of SNP difference (or divergence) of different materials, and the results can be used to verify each other with the results of other clustering methods. In the present invention, based on the detected and filtered high-quality single nucleotide polymorphisms (SNPs) vcf files, PCA analysis was performed using PLINK software. The vcf files were first converted to PLINK format to obtain three files ending with map, nosex, and ped. The map file contains the name and location information of the variation information, nosex contains the ID information of the individual, and the ped file contains the phenotypic and genotypic information. Then, PCA is calculated using PLINK software, resulting in two files ending with eigenval and eigenvec. The eigenval file represents the weight of each PCA, and the other records the eigenvectors, which are used to draw the coordinate axes and visualized using the ggplot2 package in R.

[0077] Population structure analysis, also known as population stratification, refers to the presence of subgroups with different gene frequencies in the population under study, in which the materials within the same subgroup are closely related, and the subgroups are more distantly related. Population structure analysis can quantify the number of ancestors of the population under study and infer the bloodline origin of each sample. It is a population cluster analysis method that is currently widely used and helps to understand the evolutionary process of the material. Therefore, based on SNP data, admixture software is used to analyze the population structure of Camellia oleifera materials. For the study population, the number of subgroups (K value) is set to 2-9 for clustering, and the clustering results are cross-validated to determine the optimal number of clusters based on the valley value of the cross-validation error rate. In the present invention, Structure analysis is performed using Admixture v1.30 software, and a Q file containing ancestral information and a log file containing the k value, i.e., CVerror, are obtained through VCF file information. The optimal K value is found by drawing the k-CV curve, and the Q file is input to draw the structure diagram through the R language ggplot2 package.

[0078] Phylogenetic tree is a phylogenetic branch classification method that can be used to show the evolutionary relationship between species with a common ancestor. The phylogenetic tree is constructed based on SNPs data, usually using the neighbor-joining method (NJ). It is based on the principle of minimum evolution and can simultaneously give the topological structure and branch length. Ancestral component analysis is used to explore the degree of mixing of ancestral components of individuals and the degree of gene exchange between groups. It is more reasonable to estimate the optimal number of ancestors, that is, to divide the group into several subgroups. The number of subgroups here is called K. The range of K can usually be preset to 2 to n. For each K value simulation result, the software will calculate a CV error value and a maximum likelihood value. For the structure simulated by each K value, it is generally believed that the CVerror value is the smallest and the K value is the best. In the present invention, MEGA software was used to calculate and draw the neighbor-joining (NJ) evolutionary tree (bootstrap 1000 times); FigTree v144 was used to modify and beautify the evolutionary tree. At the same time, the coloring of each branch of the NJ tree was drawn by the results of the Admixture analysis of the Structure analysis. The judgment standard was that when one ancestor of the individual accounted for more than 90%, it was judged as a homozygous individual, otherwise it was a heterozygous individual.

[0079] Prepare a fingerprint of Camellia oleifera, compare the sample to be tested with the fingerprint library, and identify the Camellia oleifera variety based on the above analysis.

[0080] The technical solutions provided by the present invention are described in detail below with reference to the embodiments, but they should not be construed as limiting the scope of protection of the present invention.

[0081] Example 1 Establishment of Camellia oleifera fingerprint

[0082] This project used SLAF-seq (Specific-Locus Amplified Fragment Sequencing) technology to conduct preliminary molecular marker development on 104 Camellia oleifera (26 varieties). The SLAF-seq molecular marker polymorphisms were obtained and their genetic evolution and diversity were analyzed. The 26 Camellia oleifera varieties are as follows:

[0083] Jiang'an 1 (JA-1), Jiang'an 54 (JA-54), Jiang'an 80 (JA-80);

[0084] Tsui Ping 15 (CP-15), Tsui Ping 16 (CP-16), Tsui Ping 37 (CP-37);

[0085] Xulong 1 (XL-1), Xulong 2 (XL-2), Xulong 3 (XL-3);

[0086] Changlin 3 (CL-3), Changlin 4 (CL-4), Changlin 40 (CL-40), Changlin 53 (CL-53),

[0087] Chuanrong 50 (CR-50), Chuanrong 55 (CR-55), Chuanrong 153 (CR-153), Chuanrong 156 (CR-156), Chuanrong 447 (CR-447)

[0088] Darling 1 (DL-1), Darling 22 (DL-22), Chuanfu 53 (CF-53), Chuanfu 227 (CF-227), Huajin (HJ), Huaxin (HX), Huashuo (HS), Xianglin 210 (XL-210).

[0089] Based on the existing published information on genome size and GC content, the Camellia angustifolia genome was selected as the reference genome for enzyme digestion experiments. The specific genome information is shown in Table 1:

[0090] Table 1 Reference genome information

[0091] Reference genome Genome size (Mb) GC (%) Assembly accuracy Camellia oleifera 3,001 37.63 Chromosome

[0092] 1. Take the 26 varieties of Camellia oleifera leaves, select 4 samples for each variety, and digest the genomic DNA of each qualified sample according to the selected optimal enzyme digestion scheme (RsaI and HaeIII). After digestion, the 414-464bp enzyme-digested fragment (SLAF tag) is recovered. The obtained enzyme-digested fragment is treated with 3' end A, connected with Dual-index sequencing adapter, PCR amplified, purified, mixed, and gel-cut to select the target fragment. After the library passes the quality inspection, PE126bp sequencing is performed using Illumina HiSeq. The specific operation process is shown in Figure 1 .

[0093] 2. Bioinformatics Analysis

[0094] Dual-index was used to identify the raw sequencing data and obtain reads for each sample. Fastp v0.2011 software was first used to perform quality control filtering on the raw sequencing data. Each read was trimmed based on its quality value. A sliding window was run from the 3' end of the read to the beginning, removing bases with an average quality value less than 20 (-320). A sliding window was run from the 5' end of the read to the end, removing bases with an average quality value less than 20 (-520). The average quality value of the bases in the sliding window was set to 20 (-M20). Length filtering was also performed, discarding reads with a length of less than 36 (-136).

[0095] After filtering the sequencing reads for adapters, the sequencing quality and data volume are assessed. The data is used to evaluate the efficiency of the enzyme digestion protocol and to determine the accuracy and effectiveness of the experimental process. Based on bioinformatics analysis, genome-wide SNP markers are developed for the population, and representative, high-quality SNPs within the population are used for population polymorphism analysis. The specific operation is as follows: use GTAKv4.12 software, input the BAM format alignment file obtained above, identify variants, find SNPs and Indels, obtain the VCF file storing the variant information of each individual, and then merge and type the generated individual VCF; use GATK software to mark SNPs (-select-type SNPs) and filter: then use Bcftools to filter out the sites 5bp before and after the SNP, and then to eliminate the influence of sequencing depth and determine the true quality value of the site, perform QualBy Depth filtering, filter out reads with QD < 2.0, perform quality value filtering, filter out reads with QUAL < 30.0, perform RMS Mapping Quality filtering, filter out reads with MQ < 40.0, and filter Fisher Strand at the same time, that is, use Fisher's exact test to determine whether the current variant has a chain-specific tendency (FS> 60.0). In order to reduce the influence of low-quality sites and reduce running time, use VCFTOOLS software to perform missing rate (--max-missing 0.5) and minor allele frequency (--maf 0.05). The bioinformatics analysis process is as follows Figure 2 shown.

[0096] 3. Sequencing data quality statistics

[0097] Using the Camellia angustifolia reference genome as the reference, an average of 490,759 SLAF tags were generated per sample, with an average SLAF depth of 10.23x. Sequencing data for each sample were analyzed, including read count, Q30, and GC content. A total of 775.32 Mbreads were obtained, with an average Q30 of 93.58% and an average GC content of 42.65%. This data was further analyzed.

[0098] 4. Genetic evolution analysis

[0099] The phylogenetic tree of each sample was constructed using MEGAX software based on the neighbor-joining method. Figure 3 As shown. Figure 3It can be seen that the sampling of Changlin 40, Changlin 53, Xulong 1, Changlin 4, Xulong 2, Huaxin, Huashuo, Cuiping 37, Jiang'an 80, Chuanfu 53, Chuanfu 227, Xianglin 210, Xulong 3, Changlin 3, Jiang'an 1, Cuiping 16, Cuiping 15, and Jiang'an 54, these 18 varieties are pure, and the bud collection nursery should be relatively pure and uncontaminated.

[0100] Dalin 1 (1-3) and Dalin 1-4 are likely Dalin 22; Chuanrong 156 (2-4), Chuanrong 50 (2-4), Chuanrong 156-1, and Chuanrong 50-1 are likely sample mixes. The sampling of these three varieties is relatively pure and can be used for fingerprint database construction and further analysis. Chuanrong 447 (3-4), Chuanrong 447 (1-2); Dalin 22 (1-2), Dalin 22 (3-4); Chuanrong 153 (2, 4), and Chuanrong 153 (1, 3) have sampling issues and are severely contaminated, requiring careful database construction. Chuanrong 55 and Huajin have very mixed sampling and are completely unusable, so they will not be considered for subsequent database construction (i.e., only the 25 varieties excluding Chuanrong 55 will be used for fingerprint database construction).

[0101] Xulong 1, Changlin 4, and Dalin 1 are very close, indicating a high degree of genetic similarity. Xulong 3 and Xianglin 210 are also very close; Huajin is likely a hybrid of Huaxin and Huashuo.

[0102] 5. Principal Component Analysis

[0103] This example uses EIGENSOFT software to perform principal component analysis based on SNP data to obtain the clustering of samples. Through PCA analysis, we can know which samples are relatively close and which samples are relatively distant, which can assist in evolutionary analysis. The PCA clustering is as follows Figure 4 As shown. Figure 4 It can be seen that Xulong 3 and Xianglin 210 are clustered into one category (mixed with Chuanrong 153), Changlin 4, Dalin 1 and Xulong 1 are clustered into one category, indicating that their relationship is relatively close, while the rest are clustered into one category (distant from the above five varieties), and the results are basically consistent with the NJ tree analysis.

[0104] 6. Group structure analysis

[0105] Based on SNP data, admixturc software was used to analyze the population structure of Camellia oleifera. For the study population, the number of subgroups (K value) was set to 1-18 for clustering, and the clustering results were cross-validated to determine the optimal number of clusters based on the valley value of the cross-validation error rate. The analysis results are as follows: Figure 5 、 Figure 6 As shown. Figure 5 and Figure 6As can be seen, the ancestral composition of most populations is independent. However, populations Xulong 3 and Xianglin 210, as well as populations Xulong 1, Changlin 4, and Dalin 1, almost consistently clustered together during K = 2-18, with subpopulation composition similarity near 100%. Therefore, it is suspected that they are closely related, sharing a common parent. Huajin-1 and Huajin-4 consistently clustered with Huashuo during K = 2-18, while Huajin-2 consistently clustered with Huaxin during K = 2-18, suggesting that these populations may have descended from Huashuo and Huaxin, respectively. Chuanrong 55 is highly mixed, so it was not considered in the subsequent fingerprinting.

[0106] 7. Construction of electronic DNA fingerprint

[0107] 7.1 Core Collection Screening

[0108] Core collections represent the widest possible range of genetic diversity across a particular species and its closely related wild species, encompassing morphological characteristics, geographic distribution, genes, and genotypes. This has important academic and practical implications for promoting germplasm exchange, utilization, and genebank management. Core Hunter II can extract diverse, representative, and minimally redundant subsets from a large collection of germplasm resources to construct core collections or mini-core collections. Based on genetic variation marker (SNP) data and weighted processing using multiple assessment measures (e.g., Modified Rogers distance, Shannons Diversity Index), selected materials exhibit high diversity, high representativeness, and high allelic richness. This analysis employed Core Hunter II software, combined with the weighted indices Modified Rogers distance (0.7) and Shannons Diversity Index (0.3), to screen for a gradient of 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, and 0.9 of the total germplasm resource. The gene coverage (CV) of the screened materials was evaluated, and 31 materials were finally determined as core germplasm, while the expected core germplasm was 26, indicating that there may be some contamination in the collected samples.

[0109] 7.2 Fingerprint Construction

[0110] This example is based on the developed SNP molecular markers and adopts the difference comparison method. After filtering according to the following method, a total of 25 variable sites are obtained, and a DNA fingerprint of Camellia oleifera is constructed based on these sites.

[0111] Screening principles:

[0112] (1) Markers are evenly distributed across the genome;

[0113] (2) There are no other SNPs in the 10 bp before and after the marker site, and the 200 bp segment before and after the marker is blast-specific;

[0114] (3) discarding population sites with a maf lower than 5% and a site deletion rate higher than 50%;

[0115] (4) discarding sites with polymorphic information content (PIC) less than 0.3;

[0116] (5) The site is located in the gene region (upstream and downstream 10 kb, 3 / 5' UTR region, exon region, intron region).

[0117] The fingerprints constructed are as follows Figure 8 shown.

[0118] 8. Construction of physical DNA fingerprint

[0119] The 25 selected electronic fingerprints were further screened through preliminary experiments such as primer design, PCR amplification, and Sanger sequencing. Finally, eight valid fingerprint loci were selected for further physical DNA fingerprint construction. Annotation information for the eight candidate loci is shown in Table 2.

[0120] Table 2 Annotation information of candidate sites

[0121]

[0122]

[0123] 8.1 Sample Selection

[0124] For the 25 varieties except Chuanrong 55 (because of sampling reasons, each individual of Huajin was considered as a variety for construction), three biological replicates were randomly selected for each variety based on the NJ cluster tree results for further experiments.

[0125] 8.2 DNA extraction

[0126] For all selected samples, 20 mg of dry insect-free leaves were taken to extract DNA using a plant DNA rapid extraction kit (Nanjing Novozymes Biotechnology Co., Ltd.). The concentration and purity of the DNA were detected by ultraviolet spectrophotometry, and the extracted DNA was stored at -20°C.

[0127] 8.3 PCR amplification and Sanger sequencing

[0128] A total of 8 pairs of site SNP-specific primers were selected as shown in Table 9 and synthesized by Shanghai Bioengineering Technology Service Co., Ltd. 2X Reaction Mix, Golden DNA Polymerase (containing Mg 2+) were purchased from Beijing Tiangen Biochemical Technology Co., Ltd. The reaction used a 15 μl system: 1.5 μl of template DNA (20 ng μl⁻¹), 7.5 μl of 2X Reaction Mix, 0.3 μl of Golden DNA Polymerase, 0.6 μl of each upstream and downstream primer (10 pmol μl⁻¹), and the volume was brought to 15 μl with ddH₂O. The reaction procedure was as follows: 94°C pre-denaturation for 5 min; 35 cycles of denaturation at 94°C for 30 s, annealing at 58°C for 30 s, and extension at 72°C for 1 min; and finally, extension at 72°C for 7 min. The cells were stored at 4°C and sent to Beijing Qingke Biotechnology Co., Ltd. for Sanger sequencing.

[0129] Table 3 Specific primers for 8 DNA fingerprint loci

[0130]

[0131]

[0132] 8.4 Data Sequencing Analysis

[0133] Data from successful Sanger sequencing were interpreted in SnapGene 6.0.2 according to the following rules: A locus with a base content greater than 50 and a standard peak shape was interpreted as a single base; this interpretation rule also applied to multiple peaks. For peaks with a content less than 50 but distinct characteristics, interpretation was performed based on three biological replicates. For any discrepancies among the three biological replicates, the genotype of the majority individual was used as the standard, while the genotype of the minority individual was also recorded.

[0134] 8.5 Physical DNA fingerprint of Camellia oleifera

[0135] After site-by-site interpretation, the physical DNA fingerprints of the following 25 oil-tea camellia varieties were finally obtained: Figure 9 、 Figure 10 As shown, from the physical fingerprint map we can know:

[0136] 1. The allele types of Xianglin 210 and Xulong 3, and Changlin 4 and Xulong 1, are completely identical, and remain consistent after additional testing. This suggests that these two varieties may be closely related, or even derived from a common parent. Furthermore, Dalin 1 is highly similar to Changlin 4 and Xulong 1, suggesting that Dalin 1 may be closely related to these two varieties, necessitating careful identification.

[0137] 2. After comparing Huajin-1 and Huajin-4 and combining them with the previous basic group structure analysis, it was found that they might have been mixed up with Huashuo, while Huajin-2 was mixed up with Huaxin. At the same time, Chuanrong 153-3 was suspected to be Chuanrong 156 after analysis.

[0138] 3. The consistency of the newly sent Huajin samples is very good, which basically confirms that the previous Huajin samples were taken incorrectly. Therefore, the physical DNA fingerprint map was constructed using the newly sent samples, and the construction results are as follows: Figure 10 .

[0139] 4. When tested with newly submitted ASUS, Huaxin and confused Huajin, the fingerprint map can accurately distinguish them, indicating that the physical fingerprint map is highly reliable and can be used in further production practice.

[0140] Example 2

[0141] In this example, SLAF-seq (Specific-Locus Amplified Fragment Sequencing) technology was used to preliminarily develop molecular markers for 13 superior Camellia oleifera strains (numbered GY01 to GY13) from Guangyuan City, obtain the SLAF-seq molecular marker polymorphism of Camellia oleifera, and analyze its genetic evolution and diversity.

[0142] Based on the existing published information on genome size and GC content, the Camellia angustifolia genome was finally selected as the reference genome for enzyme digestion experiments. The specific information of the genome was the same as in Example 1.

[0143] 1. Determine the enzyme digestion protocol

[0144] The genomic DNA of each qualified sample was digested using the selected optimal enzyme digestion protocol (RsaI and HaeIII). After digestion, the 414-464 bp digested fragment (SLAF tag) was recovered. The resulting digested fragments were 3′-end amplified, ligated with dual-index sequencing adapters, amplified by PCR, purified, mixed, and gel-cleaved to select the target fragment. After the library passed quality inspection, PE 126 bp sequencing was performed using an Illumina HiSeq. The specific operation process was the same as in Example 1.

[0145] 2. Bioinformatics Analysis

[0146] Same as Example 1

[0147] 3. Genetic evolution analysis

[0148] The phylogenetic tree of each sample was constructed using MEGAX software based on the neighbor-joining method. Figure 11 As shown. Figure 11As can be seen, the 13 individuals GY01-GY13 are not similar to any other varieties in the existing database. GY forms an independent branch on the phylogenetic tree and is separated from the individuals in the database at the base of the tree. Therefore, these 13 individuals are considered to be a potential new genetic germplasm resource (not included in the Camellia oleifera fingerprint database). Within these 13 individuals, GY06, GY12, GY07, and GY05 are located on the same branch (based on branch length and GY07 as the standard, the similarity rates of GY05, GY06, and GY012 to GY07 are 88.57%, 80.90%, and 89.28%, respectively); GY02, GY01, and GY04 are located on the same branch (based on GY01 as the standard, the similarity rates of GY02 ​​and GY04 to GY01 are 75.84% and 67.90%, respectively, while the similarity between GY02 ​​and GY04 is 89.54%).

[0149] Using EIGENSOFT software, we performed principal component analysis based on SNP data to obtain the clustering of samples. Through PCA analysis, we can know which samples are relatively close and which samples are relatively distant, which can assist in evolutionary analysis. The PCA clustering is as follows: Figure 12 The PCA results show that the 13 individuals tested did not cluster into the same category as the samples of the previously established varieties. Within the 13 materials, individuals GY06, GY12, GY07, and GY05 were similar; GY02, GY09, and GY04 were similar; GY13 and GY11 were similar. The remaining individuals were quite different and had significant differences from other individuals on the PC1 dimension, but little difference on the PC2 dimension. This indicates that the 13 individuals tested exhibited strong heterogeneity from the individuals in the Camellia oleifera fingerprint library.

[0150] 4. This is based on SNP data and uses admixture software to analyze the population structure of Camellia oleifera materials. For the research population, the number of subpopulations (K value) was set to 2-9 for clustering, and the clustering results were cross-validated to determine the optimal number of clusters based on the trough value of the cross-validation error rate. The cross-validation error rate showed that the optimal number of subpopulations K for the Camellia oleifera population was 1, indicating that the optimal number of subpopulations for the Camellia oleifera population was 1. This may be because the sequencing data was simplified for genome sequencing and the kinship of these Camellia oleifera varieties is very close. Therefore, the conclusions of the population structure analysis are for reference only.

[0151] 5. Group structure analysis results are as follows Figure 13As shown, population structure analysis revealed that the ancestral composition of most populations was independent. Between K = 2 and 9, GY04, GY01, GY09, and GY02 ​​were almost always clustered together; GY13 and GY11 were almost always clustered together except when K = 9; and GY05, GY06, GY07, GY12, GY03, and GY10 were also distinct from other varieties in the database. Therefore, the individuals tested through ancestral composition analysis were not recorded in the Camellia oleifera fingerprint database and are therefore considered potential new germplasm resources.

[0152] The "new" germplasm resources described in this paper refer to the varieties included in the existing fingerprint library (the fingerprint library will be continuously improved in the future, adding new varieties and removing inaccurate ones to achieve even greater accuracy). In fact, these "new" germplasm resources were formed decades ago, far longer than the time of these varieties. Furthermore, the identification results here are based on sequence, not phenotypes. Therefore, phenotypes should be considered in subsequent breeding to achieve optimal breeding results.

[0153] 6. Based on the above analysis, the identification conclusion is: The 13 superior camellia oleifera strains submitted for testing are distinct from all varieties in the existing camellia oleifera fingerprint database and can be used as superior strains for cultivation of new germplasm resources. Within these 13 individuals, GY02, GY01, and GY04 are similar (GY02 and GY04 are highly similar); GY13 and GY11 are highly similar; and GY06, GY12, GY07, and GY05 are similar. However, since these superior strains are not closely related to any strains in the database, it is recommended to consult historical literature for verification. Furthermore, comprehensive consideration of phenotypic characteristics is necessary for the selection of superior strains and new germplasm resources.

[0154] Example 3

[0155] In this example, Specific-Locus Amplified Fragment Sequencing (SLAF-seq) technology was used to preliminarily develop molecular markers for five Camellia oleifera materials (numbered A1, A2, A3, A4, and A5) from Gongxing Town, Jiange County, Guangyuan City. SLAF-seq molecular marker polymorphisms of Camellia oleifera were obtained, and its genetic evolution and diversity were analyzed.

[0156] Based on the existing published information on genome size and GC content, the Camellia angustifolia genome was selected as the reference genome for enzyme digestion experiments. The specific genome information is the same as in Example 1:

[0157] 1. Determine the enzyme digestion protocol

[0158] The genomic DNA of each qualified sample was digested using the selected optimal enzyme digestion protocol (RsaI and HaeIII). After digestion, the 414-464 bp digested fragment (SLAF tag) was recovered. The resulting digested fragment (SLAF tag) was 3′-end amplified, ligated with a dual-index sequencing adapter, amplified by PCR, purified, mixed, and gel-cleaved to select the target fragment. After the library passed quality inspection, it was sequenced. The specific operation process was the same as in Example 1.

[0159] 2. Bioinformatics Analysis

[0160] Same as Example 1.

[0161] 3. Using MEGA X software, the phylogenetic tree of each sample was constructed based on the neighbor-joining method. The results are as follows: Figure 14 As shown. Figure 14 It can be seen that A1, A2, A5, and A4 are similar to Changlin No. 4 in the library, and A3 is similar to Changlin No. 40 in the oil-tea fingerprint library.

[0162] 4. This example uses EIGENSOFT software to perform principal component analysis based on SNP data to obtain the clustering of samples. PCA analysis can be used to determine which samples are relatively close and which samples are relatively distant, which can assist in evolutionary analysis. The PCA clustering is shown in Figure 4. Figure 15 As shown. Figure 15 It can be seen that A1, A2, A5, and A4 tested are similar to Changlin No. 4 in the Camellia oleifera fingerprint library, and A3 is similar to Changlin No. 40 in the Camellia oleifera fingerprint library.

[0163] 5. Based on SNP data, admixture software was used to analyze the population structure of the Camellia oleifera material. For the study population, the number of subpopulations (K value) was set to 2-9 for clustering, and the clustering results were cross-validated to determine the optimal number of clusters based on the trough value of the cross-validation error rate. The cross-validation error rate showed that the optimal subpopulation number K value for the Camellia oleifera population was 1, indicating that the optimal subpopulation number for the Camellia oleifera population was 1. This may be because the sequencing data was simplified for genome sequencing and the kinship of these Camellia oleifera varieties was very close. Therefore, the conclusions of the population structure analysis are for reference only.

[0164] Group structure analysis Figure 16 As shown by Figure 16 It can be seen that the ancestral components of most groups are independent. When K=2-9: A1, A2, A5, A4 have the same genetic components as Changlin 4, and A3 has the same genetic components as Changlin 40.

[0165] 6. Based on the above analysis, the identification conclusion is: after comprehensive analysis of the phylogenetic tree, PCA analysis and admixture ancestral component analysis, the five camellia oil tree individuals submitted for testing are of the Changlin variety, and the genetic components of A1, A2, A5, and A4 are the same as those of Changlin No. 4, and the genetic components of A3 are the same as those of Changlin No. 40.

[0166] Example 4

[0167] In this example, SLAF-seq (Specific-Locus Amplified Fragment Sequencing) technology was used to preliminarily develop molecular markers for 13 superior Camellia oleifera strains (numbered HL010, HL006, HL005, HL004, HL008, HL016, HL014, HL011, HL012, HL007, HL013, HL015, and HL018) from Huili City, Liangshan Yi Autonomous Prefecture. SLAF-seq molecular marker polymorphisms were identified for Camellia oleifera, and its genetic evolution and diversity were analyzed.

[0168] Based on the existing published information on genome size and GC content, the Camellia angustifolia genome was selected as the reference genome for enzyme digestion experiments. The specific genome information is the same as in Example 1:

[0169] 1. Determine the enzyme digestion protocol

[0170] The genomic DNA of each qualified sample was digested using the selected optimal enzyme digestion protocol (RsaI and HaeIII). After digestion, the 414-464 bp digested fragment (SLAF tag) was recovered. The resulting digested fragment (SLAF tag) was 3′-end amplified, ligated with a dual-index sequencing adapter, amplified by PCR, purified, mixed, and gel-cleaved to select the target fragment. After the library passed quality inspection, it was sequenced. The specific operation process was the same as in Example 1.

[0171] 2. Bioinformatics Analysis

[0172] Same as Example 1.

[0173] 3. Use MEGAX software and construct the phylogenetic tree of each sample based on the neighbor-joining method, such as Figure 17 As shown. Figure 17It can be seen that the eight individuals HL010, HL006, HL005, HL004, HL008, HL016, HL014, and HL013 are not similar to other varieties in the existing library and form independent primary branches (they separated from Cuiping 16, Jiang'an 1, and Cuiping 15 on the main branch very early). Therefore, the above seven individuals are considered to be a potential new genetic germplasm resource; the two individuals HL011 and HL012 also cluster together and form independent primary branches (they separated from Changlin 40 on the main branch very early). Therefore, these two individuals are also considered to be potential new genetic germplasm resources; HL007 is close to Dalin 1-4 (similarity 90.67%, calculated by NJ branch length), and HL015 (86.15%) and HL018 (75.81%) are close to Chuanrong 447.

[0174] 4. Use EIGENSOFT software to perform principal component analysis based on SNP data to obtain the clustering of samples. Through PCA analysis, we can know which samples are relatively close and which samples are relatively distant, which can assist in evolutionary analysis. PCA clustering is as follows Figure 18 As shown. Figure 18 It can be seen that the samples of HL010, HL006, HL005, HL004, HL008, HL016, HL014, HL013 / HL011, and HL012 submitted for testing do not cluster into the same category as the varieties in the early database construction; HL007, HL013, HL015, and HL018 are close to the varieties in the database construction, and the results are basically consistent with the NJ tree analysis.

[0175] 5. Based on SNP data, admixture software was used to analyze the population structure of the Camellia oleifera material. For the study population, the number of subpopulations (K value) was set to 2-9 for clustering, and the clustering results were cross-validated to determine the optimal number of clusters based on the trough value of the cross-validation error rate. The cross-validation error rate showed that the optimal subpopulation number K value for the Camellia oleifera population was 1, indicating that the optimal subpopulation number for the Camellia oleifera population was 1. This may be because the sequencing data was simplified for genome sequencing and the kinship of these Camellia oleifera varieties was very close. Therefore, the conclusions of the population structure analysis are for reference only.

[0176] Group structure analysis Figure 19 As shown by Figure 19The ancestral composition of most populations is independent. For K = 2-9, the following combinations: HL010, HL006, HL005, HL004, HL008, HL016, and HL014 are inconsistent with the composition of other varieties in the Camellia oleifera fingerprint library; HL007 has an ancestral composition similar to that of Dalin 1-4; HL018 has an ancestral composition similar to that of Chuanrong 447; and HL011 and HL012 are separated from the other lines. Therefore, based on ancestral composition analysis, HL010, HL006, HL005, HL004, HL008, HL016, HL014, HL013, HL011, HL012, and HL015 are considered potential new germplasm resources.

[0177] 6. Based on the above analysis, the identification conclusions are: HL010, HL006, HL005, HL004, HL008, HL016, HL014, HL013, HL015, HL011, and HL012 are distinct from all varieties in the existing fingerprint database and are therefore suitable for cultivation as superior new germplasm resources. HL018 is similar to Chuanrong 447-2, but also has some differences. HL007 is a superior strain similar to Dalin 1-4. These identification results need to be considered in conjunction with phenotypic analysis before selecting superior strains and new germplasm resources.

[0178] Example 5

[0179] In this example, Specific-Locus Amplified Fragment Sequencing (SLAF-seq) technology was used to preliminarily develop molecular markers for five Camellia oleifera materials (B1, B2, B3, B4, and B5) from Guangyuan Tiancheng Agricultural Development Co., Ltd. SLAF-seq molecular marker polymorphisms were obtained for Camellia oleifera and its genetic evolution and diversity were analyzed.

[0180] Based on the existing published information on genome size and GC content, the Camellia angustifolia genome was selected as the reference genome for enzyme digestion experiments. The specific genome information is the same as in Example 1:

[0181] 1. Determine the enzyme digestion protocol

[0182] The genomic DNA of each qualified sample was digested using the selected optimal enzyme digestion protocol (RsaI and HaeIII). After digestion, the 414-464 bp digested fragment (SLAF tag) was recovered. The resulting digested fragment (SLAF tag) was 3′-end amplified, ligated with a dual-index sequencing adapter, amplified by PCR, purified, mixed, and gel-cleaved to select the target fragment. After the library passed quality inspection, it was sequenced. The specific operation process was the same as in Example 1.

[0183] 2. Bioinformatics Analysis

[0184] Same as Example 1.

[0185] 3. Using MEGAX software, the phylogenetic tree of each sample was constructed based on the neighbor-joining method. The results are as follows: Figure 20 As shown. Figure 20 It can be seen that B3, B4, and B5 are similar to Changlin No. 4 in the library, B1 is similar to Changlin No. 40 in the oil-tea camellia fingerprint library, and B2 is similar to Chuanrong 153.

[0186] 4. Use EIGENSOFT software to perform principal component analysis based on SNP data to obtain the clustering of samples. Through PCA analysis, we can know which samples are relatively close and which samples are relatively distant, which can assist in evolutionary analysis. PCA clustering is as follows Figure 21 As shown. Figure 21 It can be seen that B3, B4, and B5 tested are similar to Changlin No. 4 in the Camellia oleifera fingerprint library, B1 is similar to Changlin No. 40 in the Camellia oleifera fingerprint library, and B2 is similar to Chuanrong 153.

[0187] 5. Based on SNP data, admixture software was used to analyze the population structure of the Camellia oleifera material. For the study population, the number of subpopulations (K value) was set to 2-9 for clustering, and the clustering results were cross-validated to determine the optimal number of clusters based on the trough value of the cross-validation error rate. The cross-validation error rate showed that the optimal subpopulation number K value for the Camellia oleifera population was 1, indicating that the optimal subpopulation number for the Camellia oleifera population was 1. This may be because the sequencing data was simplified for genome sequencing and the kinship of these Camellia oleifera varieties was very close. Therefore, the conclusions of the population structure analysis are for reference only.

[0188] Group structure analysis Figure 22 As shown by Figure 22 It can be seen that the ancestral components of most groups are independent. When K=2-9: B3, B4, B5 have the same genetic components as Changlin 4, B1 has the same genetic components as Changlin 40, and B2 is similar to Chuanrong 153.

[0189] 6. Based on the above analysis, the identification conclusion is: B1, B3, B4, and B5 are Changlin varieties. B2 may have been sampled incorrectly due to a sampling error and is similar to CR-153. Furthermore, B3, B4, and B5 share genetic components with Changlin 4, B1 shares genetic components with Changlin 40, and B2 is similar to CR-153. The identification of sample sampling errors through the Camellia oleifera fingerprint further demonstrates the reliability of the Camellia oleifera fingerprint.

[0190] Example 6

[0191] In this example, Specific-Locus Amplified Fragment Sequencing (SLAF-seq) technology was used to preliminarily develop molecular markers for five Camellia oleifera materials (numbered C1, C2, C3, C4, and C5) from Fenghuang Village, Gongxing Town, Jiange County, Guangyuan City. SLAF-seq molecular marker polymorphisms were obtained for Camellia oleifera and its genetic evolution and diversity were analyzed.

[0192] Based on the existing published information on genome size and GC content, the Camellia angustifolia genome was selected as the reference genome for enzyme digestion experiments. The specific genome information is the same as in Example 1:

[0193] 1. Determine the enzyme digestion protocol

[0194] The genomic DNA of each qualified sample was digested using the selected optimal enzyme digestion protocol (RsaI and HaeIII). After digestion, the 414-464 bp digested fragment (SLAF tag) was recovered. The resulting digested fragment (SLAF tag) was 3′-end amplified, ligated with a dual-index sequencing adapter, amplified by PCR, purified, mixed, and gel-cleaved to select the target fragment. After the library passed quality inspection, it was sequenced. The specific operation process was the same as in Example 1.

[0195] 2. Bioinformatics Analysis

[0196] Same as Example 1.

[0197] 3. Using MEGA X software, the phylogenetic tree of each sample was constructed based on the neighbor-joining method. The results are as follows: Figure 20 As shown. Figure 23 It can be seen that C1, C2, and C5 are close to Changlin No. 4 in the warehouse, C4 is close to Changlin No. 40 in the warehouse, and C3 is close to Changlin No. 53.

[0198] 4. Use EIGENSOFT software to perform principal component analysis based on SNP data to obtain the clustering of samples. Through PCA analysis, we can know which samples are relatively close and which samples are relatively distant, which can assist in evolutionary analysis. PCA clustering is as follows Figure 24 As shown. Figure 24 It can be seen that the tested C1, C2, and C5 are similar to Changlin No. 4, C3 is similar to Changlin No. 53, and C4 is similar to Changlin No. 40.

[0199] 5. Based on SNP data, admixture software was used to analyze the population structure of the Camellia oleifera material. For the study population, the number of subpopulations (K value) was set to 2-9 for clustering, and the clustering results were cross-validated to determine the optimal number of clusters based on the trough value of the cross-validation error rate. The cross-validation error rate showed that the optimal subpopulation number K value for the Camellia oleifera population was 1, indicating that the optimal subpopulation number for the Camellia oleifera population was 1. This may be because the sequencing data was simplified for genome sequencing and the kinship of these Camellia oleifera varieties was very close. Therefore, the conclusions of the population structure analysis are for reference only.

[0200] Group structure analysis Figure 25 As shown by Figure 25 It can be seen that the ancestral components of most groups are independent. When K = 2-9: C1, C2, C5 have the same genetic components as Changlin 4, C3 has the same genetic components as Changlin 53, and C4 has the same genetic components as Changlin 40.

[0201] 6. Based on the above analysis, the identification conclusion is: after comprehensive analysis of the phylogenetic tree, PCA analysis and admixture ancestral component analysis, the five camellia oleifera individuals sent for testing are of the Changlin variety, and the genetic components of C1, C2, and C5 are the same as those of Changlin No. 4, the genetic components of C3 are the same as those of Changlin No. 53, and the genetic components of C4 are the same as those of Changlin No. 40.

[0202] Example 7

[0203] In this example, Specific-Locus Amplified Fragment Sequencing (SLAF-seq) technology was used to preliminarily develop molecular markers for two tea cultivars (Changlin 4 and Changlin 53) from Jisen Agriculture Co., Ltd. The polymorphism of SLAF-seq molecular markers in tea cultivars was obtained, and their genetic evolution and diversity were analyzed.

[0204] Based on the existing published information on genome size and GC content, the Camellia angustifolia genome was selected as the reference genome for enzyme digestion experiments. The specific genome information is the same as in Example 1:

[0205] 1. Determine the enzyme digestion protocol

[0206] The genomic DNA of each qualified sample was digested using the selected optimal enzyme digestion protocol (RsaI and HaeIII). After digestion, the 414-464 bp digested fragment (SLAF tag) was recovered. The resulting digested fragment (SLAF tag) was 3′-end amplified, ligated with a dual-index sequencing adapter, amplified by PCR, purified, mixed, and gel-cleaved to select the target fragment. After the library passed quality inspection, it was sequenced. The specific operation process was the same as in Example 1.

[0207] 2. Bioinformatics Analysis

[0208] Same as Example 1.

[0209] 3. Use MEGAX software and construct the phylogenetic tree of each sample based on the neighbor-joining method, such as Figure 23 As shown. Figure 26 The Changlin 53 sample submitted for testing is consistent with the fingerprint database. While Changlin 4 is not clustered with the database, it still forms a single branch. It also highly aligns with Dalin 1, Xunong 1, and Changlin in the Camellia oleifera fingerprint database. Therefore, combined with subsequent analysis, the sample submitted for testing was identified as Changlin 4 and Changlin 53.

[0210] 4. Use EIGENSOFT software to perform principal component analysis based on SNP data to obtain the clustering of samples. Through PCA analysis, we can know which samples are relatively close and which samples are relatively distant, which can assist in evolutionary analysis. PCA clustering is as follows Figure 27 As shown. Figure 27 It can be seen that the Changlin No. 4 and Changlin No. 53 submitted for testing are consistent with those in the fingerprint database.

[0211] 5. Based on SNP data, admixture software was used to analyze the population structure of the Camellia oleifera material. For the study population, the number of subpopulations (K value) was set to 2-9 for clustering, and the clustering results were cross-validated to determine the optimal number of clusters based on the trough value of the cross-validation error rate. The cross-validation error rate showed that the optimal subpopulation number K value for the Camellia oleifera population was 1, indicating that the optimal subpopulation number for the Camellia oleifera population was 1. This may be because the sequencing data was simplified for genome sequencing and the kinship of these Camellia oleifera varieties was very close. Therefore, the conclusions of the population structure analysis are for reference only.

[0212] Group structure analysis Figure 28 As shown, the conclusion of population structure analysis shows that the ancestral components of most populations are independent. When K=2-9: the principal components of Changlin 4 and Changlin 53 submitted for testing are consistent with those of Changlin 4 and Changlin 53 in the atlas library. Therefore, the Changlin 4 and Changlin 53 submitted for testing are Changlin 4 and Changlin 53 in the library.

[0213] 6. Based on the above analysis, the identification concludes that the Changlin 53 and Changlin 4 varieties submitted by Jisen Agriculture for testing are all in-stock varieties and are consistent with the commercially available varieties. Furthermore, combined with previous testing results, the Changlin 40, Changlin 53, Changlin 3, and Changlin 4 varieties submitted by Yijisen Agriculture are all in-stock varieties and consistent with the commercially available varieties. However, there are some issues with the mixing of seedlings in the company's nursery. Therefore, it is recommended that the company manage its nursery more effectively and clearly define the planting zones for oil-tea camellia varieties.

[0214] Example 8

[0215] In this example, this project used SLAF-seq (Specific-Locus Amplified Fragment Sequencing) technology to preliminarily develop molecular markers for 12 oil-tea camellia materials (4 varieties: Changlin 40, Changlin 3, Changlin 4, and Changlin 53), obtained the SLAF-seq molecular marker polymorphism of oil-tea camellia, and analyzed its genetic evolution and diversity.

[0216] Based on the existing published information on genome size and GC content, the Camellia angustifolia genome was selected as the reference genome for enzyme digestion experiments. The specific genome information is the same as in Example 1:

[0217] 1. Determine the enzyme digestion protocol

[0218] The genomic DNA of each qualified sample was digested using the selected optimal enzyme digestion protocol (RsaI and HaeIII). After digestion, the 414-464 bp digested fragment (SLAF tag) was recovered. The resulting digested fragment (SLAF tag) was 3′-end amplified, ligated with a dual-index sequencing adapter, amplified by PCR, purified, mixed, and gel-cleaved to select the target fragment. After the library passed quality inspection, it was sequenced. The specific operation process was the same as in Example 1.

[0219] 2. Bioinformatics Analysis

[0220] Same as Example 1.

[0221] 3. Using MEGAX software, based on the neighbor-joining method, 1000 random replicates were performed to construct a phylogenetic tree for each sample. Figure 29 As shown, the reproducibility of the tested Changlin 40, Changlin 3, and Changlin 4 is consistent, while one sample of Changlin 53 has a problem. Furthermore, based on phylogenetic relationships, the tested Changlin 40, Changlin 3, and Changlin 53 are consistent with the strains previously established in the database; Changlin 4 is similar to the established varieties (because Changlin 4, Dalin 1, and Xulong 1 are closely related, the Changlin 4 tested here can be considered Changlin 4).

[0222] 4. Use EIGENSOFT software to perform principal component analysis based on SNP data to obtain the clustering of samples. Through PCA analysis, we can know which samples are relatively close and which samples are relatively distant, which can assist in evolutionary analysis. PCA clustering is as follows Figure 30 The PCA results showed that the samples of Changlin 40, Changlin 3, and Changlin 53, which were tested, clustered together with the samples of the cultivars in the early stage of the database; Changlin 4 was similar to the cultivars in the database, and the results were basically consistent with the NJ tree analysis.

[0223] 5. Therefore, based on the SNP data, admixture software was used to analyze the population structure of the Camellia oleifera material. For the research population, the number of subpopulations (K value) was set to 1-18 for clustering, and the clustering results were cross-validated to determine the optimal number of clusters based on the valley value of the cross-validation error rate. The cross-validation error rate showed that the optimal subpopulation number K value for the Camellia oleifera population was 1, indicating that the optimal subpopulation number for the Camellia oleifera population was 1. This may be because the sequencing data was simplified for genome sequencing and the kinship of these Camellia oleifera varieties was very close. Therefore, the conclusions of the population structure analysis are for reference only.

[0224] The conclusions of group structure analysis are as follows Figure 31 As shown, Figure 31 The results indicate that most populations have independent ancestral components. The four newly tested Changlin populations almost consistently clustered together when K = 2-9, and the subpopulation composition similarity was nearly 100%. Therefore, it is suspected that they are closely related and share a common parent. One sample from Changlin 53 was sampled as Changlin 4, indicating some mixing in the nursery, which requires further refinement.

[0225] 6. Based on the above analysis, the identification conclusion is: the Changlin No. 4, Changlin No. 53, Changlin No. 3, and Changlin No. 40 varieties submitted for testing and identification are basically consistent with the resource library established in the early stage, and can be confirmed as the Changlin series of tea oil plant lines.

[0226] In summary, the present invention provides a method for identifying camellia cultivars, constructs DNA fingerprints for major commercially available varieties, and develops related identification technologies capable of identifying these major camellia cultivars. In the examples of the present invention, camellia fingerprints were constructed using 26 camellia cultivars. It should be noted that the camellia cultivar identification method of the present invention includes, but is not limited to, the varieties described in the examples.

[0227] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. A method for identifying oil-tea camellia varieties, characterized in that: The following steps are involved: (1) Take leaves of known tea varieties, extract DNA, construct a library, and sequence to obtain raw sequencing data; (2) Perform data quality control and filtering on the original sequencing data, use the Camellia oleifera individual as the reference genome, build an index for the reference genome, align the sequence to the reference genome, and obtain the alignment file; (3) performing SNP and Indel identification on the comparison file described in step (2), obtaining the stored variation information of each individual, and merging and typing the stored variation information; (4) Filter out the sites 5 bp before and after the SNP, and perform Qual By Depth, RMS Mapping Quality, and Fisher Strand filtering to filter the missing rate and minor allele frequency to obtain the filtered genomic SNP site data; (5) Constructing a Camellia oleifera fingerprint map based on the genomic SNP site data; The primers used to construct the Camellia oleifera fingerprint library include the sequences shown in SEQ ID NO.1 to SEQ ID NO.16; (6) Take the oil-tea camellia to be tested, repeat steps (1) to (5), and obtain the data analysis results after data analysis; The camellia oleifera variety is identified by combining the analysis results of step (6) and the camellia oleifera fingerprint library of step (5).

2. The identification method according to claim 1, wherein The data quality control filtering in step (2) is specifically as follows: each read is trimmed according to the quality value, a window is moved from the 3' end value to the beginning of the read, and the bases with an average quality value less than 20 in the window are removed; a window is moved from the 5' end to the end of the read, and the bases with an average quality value less than 20 in the window are removed, the average quality value of the bases in the sliding window is set to 20, and reads with a length of less than 36 are filtered out.

3. The identification method according to claim 1, wherein BWA software and SAMtools software were used to construct an index for the reference genes.

4. The identification method according to claim 1, wherein Step (4) filtered out reads with Qual By Depth < 2.0, RMS Mapping Quality < 40.0, and QUAL < 30.0, and used Fisher’s exact test to determine whether the current variant had a strand-specific tendency, i.e., FS > 60.

0.

5. The identification method according to claim 1, wherein The specific method for filtering the missing rate in step (4) is to enter 0.5 in --max-missing for filtering; the specific method for filtering the minor allele frequency is to enter 0.05 in --maf for filtering.

6. The identification method according to claim 1, wherein Before constructing the Camellia oleifera fingerprint library in step (5), PCA analysis, structure analysis, phylogenetic tree analysis and ancestral component analysis are performed on the SNP site data.

7. The identification method according to claim 6, characterized in that PCA analysis was performed using PLINK software, resulting in three files ending in bim, nosex, and bed. PCA calculations were then performed using PLINK software, resulting in two files ending in eigenval and eigenvec, which were visualized using the R language ggplot2 package. Structure analysis was performed using Admixture v1.30 software. A Q file containing ancestral information and a log file containing the k value, or CV error, were obtained from VCF file 1. The optimal K value was found by plotting the k-CV curve, and the Q file was input into the R language ggplot2 package to draw the structure diagram.

8. The identification method according to claim 6, characterized in that MEGA software was used to calculate and draw the NJ evolutionary tree, with the initial value of bootstrap set to 1000 times. FigTree v1.44 was used to modify and beautify the NJ evolutionary tree. The coloring of each branch of the NJ tree was drawn by the results of Admixture obtained from the Structure analysis.

9. The identification method according to claim 1, characterized in that The Camellia oleifera fingerprint library is shown in the following table: 。

Citation Information

Patent Citations

  • Non-hybrid offspring identification method based on simplified genome sequencing and SNP minor allele frequency

    CN111826429A

  • DNA fragment related to content of linoleic acid in oil-tea camellia seed oil and application of DNA fragment

    CN113604593A