Construction method and application of camellia whole genome liquid phase chip
Through magnetic bead-mediated CTAB method and high-throughput sequencing technology, combined with a variety of genetic analysis software, a whole-genome liquid phase chip of camellia was constructed, which solved the problem that the existing technology was difficult to deeply analyze the genetic relationship of camellia, and achieved systematic genetic analysis and molecular breeding guidance.
Patent Information
- Application Number
- CN202510657128.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-08-22
AI Technical Summary
Existing single analytical means or local research methods are difficult to meet the requirements of systematic and in-depth genetic analysis and molecularly assisted breeding of camellia.
High-quality genomic DNA was prepared by magnetic bead-mediated CTAB method, fragmented and ligated with linkers, and hybrid probes were captured by streptomycin-affinity labeled magnetic beads. After enriching the target fragments, high-throughput sequencing and strict data filtering were performed, and data processing and analysis were performed in combination with a variety of genetic analysis software.
Comprehensively cover camellia genome information, systematically analyze genetic relations and population structure, guide molecular assisted breeding, and promote the improvement and optimization of camellia varieties.
Smart Images

Figure CN120519609A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of liquid-phase chips, and specifically to a method for constructing a whole-genome liquid-phase chip for camellia and its application. Background Art
[0002] Camellia plays an important role in plant research and horticulture. With the increasing demand for the study of its genetic characteristics and the cultivation of excellent varieties, there is an urgent need for a technology that can comprehensively obtain genomic information and effectively analyze its genetic relationships, population structure, and trait associations. Existing single analysis methods or local research methods are difficult to meet the requirements for systematic and in-depth genetic analysis and molecular-assisted breeding of camellia. Therefore, the development of a technical system integrating the construction of genomic liquid-phase chips and multi-dimensional genetic analysis has become the research focus in this field. Summary of the Invention
[0003] Aiming at the deficiencies of the existing technology, the present invention provides a method for constructing a whole-genome liquid-phase chip for camellia and its application, solving the problem that it is difficult for the existing technology to meet the requirements for systematic and in-depth genetic analysis and molecular-assisted breeding of camellia through single analysis methods or local research methods.
[0004] To achieve the above objectives, the present invention is realized through the following technical solutions:
[0005] A method for constructing a whole-genome liquid-phase chip for camellia, comprising the following steps:
[0006] Collect camellia samples, and prepare high-quality genomic DNA from fresh leaves using the magnetic bead-mediated CTAB method, where the DNA satisfies 1.8 < OD260 / 280 < 2.0 and OD260 / 230 > 2.0;
[0007] Fragment the genomic DNA into fragments of 200 - 300 bp, and construct a library through end repair, ligation of adapters, and Pre-PCR amplification;
[0008] Hybridize the probe with the target region, capture the hybridized probe with streptavidin-labeled magnetic beads, and perform Post-PCR amplification after enriching the target fragment;
[0009] Sequence the target region using the DNBSEQ-T7 high-throughput sequencing platform, and convert the original image data into RawReads;
[0010] Filter RawReads: Remove reads with adapters, N content exceeding 10%, and bases with a quality value lower than 10 accounting for more than 50% to obtain CleanReads;
[0011] CleanReads were mapped to the reference genome using bwa software. Individual VCF calls were made based on the HaplotypeCaller method of GATKv.4.1.2.0. Population SNP calls were made by merging VCFs, and GATK hard filtering was used to exclude false positive variants. The parameters were QD < 2.0, MQ < 40.0, FS > 60.0, SOR > 3.0, MQRankSum < -12.5, QUAL < 30.0, and ReadPosRankSum < -8.0.
[0012] Preferably, the camellia samples include samples used for testing SNP array efficiency, leaf trait determination, genome resequencing, and multi-source genetic structure and phylogeny analysis.
[0013] Preferably, the construction method also includes a site screening process:
[0014] Based on the resequencing data of Camellia varieties, quality filtering, redundancy filtering, probe design evaluation, LD screening and merging with genetic linkage map data to remove duplicates were performed in sequence to finally obtain the target SNPs.
[0015] Preferably, the quality filtering parameters include deletion rate <0.2, minimum quality score ≥50, and minimum allele frequency >0.15.
[0016] Preferably, the redundancy filtering requires that the GC content of the 100 bp sequence upstream and downstream of the SNP site is 40%-60% without N bases and is a single copy in the whole genome.
[0017] Preferably, the present invention also provides an application of a camellia whole-genome liquid phase chip for phylogenetic analysis of camellia, comprising: calculating genetic distances using VCF2Dis, constructing a neighbor-joining tree based on TreeBest software, setting the Bootstrap value to 1000, and visualizing using iTOL online software.
[0018] Preferably, for the population structure analysis of Camellia: the population structure is inferred using Admixture and visualized using pophelper; the principal components are calculated using Plink2, and the first two principal components are clustered using R language ggplot2.
[0019] Preferably, ImageJ is used to measure the leaf tip length, leaf margin serration length, leaf length, leaf width, and leaf area, wherein the leaf tip length is the distance from the top of the leaf tip along the leaf vein to the inscribed circle, and the leaf margin serration length is the shortest distance from the top of the serration to the bottom edge at one-third of the leaf blade.
[0020] Preferably, for camellia genome-wide association study GWAS: an EMMAX mixed linear model is used, the genetic similarity matrix is used as the random effect variance-covariance matrix, and significant SNPs are screened through a Manhattan plot.
[0021] Preferably, for genomic selection analysis of Camellia GS: four models are used: GBLUP, BayesA, BayesB and BayesC, among which the GBLUP model is y = Xβ + Zμ + ε, which is implemented by the mmer function of the R package sommer; the Bayes model is y = Xβ + Wα m +∈, implemented by the BGLR function of the R package BGLR;
[0022] Where y is the phenotypic value vector, β is the fixed effect vector, X is the association matrix of β, W is the SNP marker genotype score matrix, α m is the SNP marker random effect, ∈ is the error effect;
[0023] The prediction ability was evaluated using 5-fold cross-validation, and the heritability H was estimated based on GBLUP. 2 , and integrate GWAS significant SNPs.
[0024] The present invention provides a method for constructing a Camellia whole-genome liquid phase chip and its application. It has the following beneficial effects:
[0025] 1. This invention constructs a liquid phase chip by preparing high-quality genomic DNA from a variety of camellia samples, library construction, probe hybridization capture, high-throughput sequencing, and rigorous data filtering and analysis. This fully covers the genome information of camellia, laying a solid foundation for in-depth research on the genetic characteristics of camellia, and can systematically and completely obtain camellia genome-related data.
[0026] 2. This invention covers a variety of applications such as phylogenetic analysis and population structure analysis. It can analyze the genetic relationships and population structure of camellia from multiple dimensions, providing a comprehensive technical means for exploring the genetic diversity and evolutionary relationships of camellia, and helping to deeply understand the genetic background and evolutionary process of camellia species.
[0027] 3. By using genomic selection analysis and whole-genome association analysis, the present invention can mine genetic information related to camellia traits, provide guidance for molecular-assisted breeding of camellia, help screen camellia varieties with excellent traits, promote the scientific development of camellia breeding, and promote the improvement and optimization of camellia varieties. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 The present invention is a flow chart of a method for constructing a camellia whole genome liquid phase chip.
[0029] Figure 2 This is a pie-shaped diagram of the distribution of SNP annotation types in the whole genome of Camellia japonica according to the present invention;
[0030] Figure 3 This is a schematic diagram of the chromosome density distribution of SNP sites in the whole genome of Camellia japonica according to the present invention;
[0031] Figure 4 Schematic diagram of the distribution histogram of the number of SNPs on each chromosome of Camellia japonica according to the present invention;
[0032] Figure 5 This is a pie-shaped diagram of the distance distribution between SNP sites in the whole genome of Camellia japonica according to the present invention;
[0033] Figure 6 This is a schematic diagram of the correlation analysis between the chromosome length and the number of SNPs in Camellia japonica according to the present invention;
[0034] Figure 7 This is a statistical diagram of the quality of camellia liquid phase chip sequencing data of the present invention;
[0035] Figure 8 This is a schematic diagram of verifying the coverage efficiency of the target area of the Camellia liquid phase chip of the present invention;
[0036] Figure 9 Schematic diagram of the Camellia phylogeny, principal component analysis, and comprehensive population structure analysis of the present invention;
[0037] Figure 10 This is a schematic diagram of measuring properties of Camellia japonica leaves of the present invention;
[0038] Figure 11 This is a Manhattan diagram of the genome-wide association analysis of camellia leaf traits of the present invention;
[0039] Figure 12 Schematic diagram of the chromosome 11 gene-log10(p) association analysis of the present invention;
[0040] Figure 13 Schematic diagram of the distribution and significance comparison of trait values of genotypes AA, AT and TT in the first measurement, the second measurement and the average level of the present invention;
[0041] Figure 14 This is a Manhattan diagram of the multi-locus genome-wide association analysis of camellia leaf traits of the present invention;
[0042] Figure 15 Schematic diagram of the Camellia genomic selection model prediction accuracy and related trait analysis of the present invention;
[0043] Figure 16 This is the site screening process intention of the present invention. DETAILED DESCRIPTION
[0044] The technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.
[0045] Please refer to the attached Figure 1 - attached Figure 16 , the embodiment of the present invention provides a method for constructing a whole-genome liquid-phase chip of camellia, including the following steps:
[0046] Collect camellia samples, and use the magnetic bead-mediated CTAB method to prepare high-quality genomic DNA from fresh leaves. The DNA satisfies 1.8 < OD260 / 280 < 2.0 and OD260 / 230 > 2.0;
[0047] Fragment the genomic DNA into fragments of 200-300 bp, and construct a library through end repair, ligation of adapters, and Pre-PCR amplification;
[0048] Hybridize the probe with the target region, capture the hybridized probe by streptavidin-labeled magnetic beads, and perform Post-PCR amplification after enriching the target fragment;
[0049] Sequence the target region using the DNBSEQ-T7 high-throughput sequencing platform, and convert the original image data into RawReads;
[0050] Filter RawReads: Remove reads with adapters, N content exceeding 10%, and bases with a quality value lower than 10 accounting for more than 50% to obtain CleanReads;
[0051] Use the bwa software to map CleanReads to the reference genome, perform individual VCF calls based on the HaplotypeCaller method of GATK v.4.1.2.0, perform population SNP calls by merging VCFs, and use GATK hard filtering to exclude false-positive variations with parameters QD < 2.0, MQ < 40.0, FS > 60.0, SOR > 3.0, MQRankSum < -12.5, QUAL < 30.0, ReadPosRankSum < -8.0.
[0052] The camellia samples include samples for testing the efficiency of SNP arrays, genetic structure and phylogenetic analysis, as well as leaf trait determination and genome resequencing (for conducting GWAS and GS analyses).
[0053] Specifically, camellia resource samples are preserved in the Camellia Germplasm Resource Conservation Center of the Research Institute of Subtropical Forestry, Chinese Academy of Forestry (RISF) (119.96°E, 30.06°N), of which 15 camellia samples are used to test the efficiency of the SNP array; 220 germplasms are used for leaf trait determination and genome resequencing, and GWAS and GS analysis are carried out. In order to conduct genetic structure and phylogenetic analysis, 54 new camellia varieties were collected from Guangdong Province, Guangxi Zhuang Autonomous Region, Hunan Province and Yunnan Province and preserved in RISF. The genomic DNA of all samples was prepared from fresh leaves using the magnetic bead-mediated method with cetyltrimethylammonium bromide (CTAB), and high-quality DNA (1.8 <OD260 / 280<2.0,OD260 / 230> 2.0) for target sequencing.
[0054] Please see the attached Figure 2 -Attached Figure 6 , the chip features of the present invention are:
[0055] Contains 20,900 SNPs, of which 54.58% are located in genic regions (including 0.14% missense mutations, 14.78% synonymous mutations, and 36.00% introns), and 45.42% are located in non-genic regions (including 20.72% UTR3, 19.08% UTR5, 5.10% upstream and downstream regions, and 2.20% intergenic regions).
[0056] SNPs were evenly distributed on the 15 chromosomes of Camellia japonica. The average distance between adjacent SNPs was 123.404 kb, and only 5.82% of the distances were greater than 500 kb. The number of SNPs was positively correlated with chromosome length (r = 0.826, p = 1.47 × 10 -4 ), among which the longest chromosome 1 has the most sites and the shortest chromosome 15 has the least sites.
[0057] Please see the attached Figure 7 and attached Figure 8 , the chip efficiency of the present invention was verified: a total of 11,797,304 SNP sites were detected in 15 individuals; after filtering, 378,125 valid sites remained (minimum allele frequency > 0.05, deletion rate < 0.1, only biallelic sites were retained);
[0058] 97.2% (20,313) of the probes were detected in the typed samples;
[0059] On average, 98.26% of captured and sequenced reads mapped to the reference genome;
[0060] The average target area coverage was 96.37%, with only D-1 and Y-27A falling below 95%. The lowest coverage was for Y-27A, at 89.91%.
[0061] The average sequencing depth of the 15 samples was 83;
[0062] Please see the attached Figure 16 , the construction method also includes the site screening process:
[0063] Based on the resequencing data of Camellia varieties, quality filtering, redundant filtering probe design, LD screening, and merging with genetic linkage map data to remove duplicates were performed in sequence to finally obtain the target SNPs.
[0064] Quality filtering parameters included missingness rate <0.2, minimum quality score ≥50, and minimum allele frequency >0.15.
[0065] Redundancy filtering requires that the GC content of the 100bp sequence upstream and downstream of the SNP site is 40%-60%, there is no N base, and it is a single copy in the whole genome.
[0066] Specifically, based on the resequencing data of 220 representative camellia varieties, the specific process of screening 21K SNP sites for constructing liquid phase arrays is as follows:
[0067] The number of SNP sites initially obtained through resequencing data (28,033,254).
[0068] Quality filtering: missing rate <0.2, minimum quality score ≥50, and minimum allele frequency >0.15.
[0069] Redundancy filtering: The GC content within 100 bp (201 bp) upstream and downstream of the SNP site was greater than 40% and less than 60%. Sequences within 100 bp upstream and downstream of the SNP site were extracted and subjected to whole-genome blast alignment. Sequences with a similarity greater than 60% and an alignment length greater than 60% (approximately 120 bp) were considered to be a single copy, and only single-copy sites were retained. No N sequences were present within 100 bp upstream and downstream of the SNP site. The number of remaining sites: 498,377.
[0070] Evaluate sites suitable for probe design: select 150nt sequences on both sides of the site, and use a window size of 100bp and a step size of 1bp to select all possible sequences within it, and select sequences with a GC content between 40% and 60%, no repeated sequences, no N bases, and a whole genome copy of 1. Sort by GC content, and give priority to sequences with a GC content close to 50% for probe design. If there are multiple suitable probe sequences, one is randomly selected. This process is performed by a Perl script developed by the team. Number of remaining sites: 186,026
[0071] LD screening: Number of remaining loci: 20,476.
[0072] Merge with genetic map loci: 4255 linkage maps were combined with the LD filtered file to remove duplicates, compared to the original data, and after probe evaluation, only 424 loci remained. Number of loci remaining after integration: 20,900.
[0073] Through the above steps, 20,900 SNP sites were finally screened out for the construction of liquid phase array.
[0074] The present invention also provides an application of a camellia whole-genome liquid phase chip for phylogenetic analysis of camellia, comprising: calculating genetic distances using VCF2Dis, constructing a neighbor-joining tree based on TreeBest software, setting a Bootstrap value of 1000, and visualizing the data using iTOL online software.
[0075] Please see the attached Figure 9 , used for camellia population structure analysis: Admixture was used to infer the population structure and pophelper was used to visualize it; Plink2 was used to calculate the principal components, and R language ggplot2 was used to perform cluster analysis on the first two principal components.
[0076] Specifically, phylogenetic analysis used VCF2Dis to calculate genetic distances. Based on these distances, a neighbor-joining (NJ) tree was constructed using TreeBest software with a bootstrap value of 1000. The phylogenetic tree was plotted using the online software iTOL. Population structure analysis used Admixture and pophelper for visualization. Principal component analysis used Plink2 to calculate principal components, and the first two principal components were visualized using ggplot2 in R.
[0077] Genotyping and population analysis were performed using the whole-genome Camellia 21K array. A total of 69 Camellia cultivars were selected for array analysis based on representative origins, 15 of which were used for the array evaluation described above. PCA analysis using 60,370 valid SNPs (missing rate = 0, MAF ≥ 0.15) revealed four major population clusters, with all azalea hybrids clustering with the original azalea cultivars.
[0078] Further analysis of phylogeny and population structure was conducted to dissect the relationships among the samples. The results showed that, taking into account the recorded information of cultivated varieties, seven subgroups (K = 7) combined with phylogenetic relationships provided the best classification of the samples. For example, Groups 1 and 4 were primarily composed of red Camellia lines, including the wild species 'Naidong' Y24-C and the single red Camellia Y43-B. In Group 7, Vietnamese clasping tea hybrids clustered with two rhododendron red Camellia hybrids, C_62 / 63, and one red Camellia hybrid, C_76, reflecting the close relationship between the parental lines.
[0079] Among them, each K value has an independent label and is automatically sorted according to the proportion of ancestors.
[0080] Please see the attached Figure 10 , used for Camellia leaf phenotypic statistics: ImageJ was used to measure leaf tip length, leaf margin serration, leaf length, leaf width, and leaf area, where the leaf tip length was the distance from the tip of the leaf along the vein to the inscribed circle, and the leaf margin serration length was the shortest distance from the top of the serration to the bottom edge at one-third of the leaf blade.
[0081] Specifically, ImageJ was used to accurately measure the leaf tip length LTL, leaf margin serration length LMSL, leaf length LL, leaf width LW, and leaf area LA, and GWAS analysis was performed using the two measurement results and the mean of the 21k chip pair.
[0082] Method for measuring leaf tip length: find the turning point of the leaf tip (the place where the leaf tip meets the leaf edge), draw an inscribed circle with tangent lines between the two points, and the distance from the top of the leaf tip along the direction of the leaf vein to the circle is the length of the leaf tip. Record the average of the three values.
[0083] Method for measuring the length of leaf edge serrations: select three serrations at one-third of the leaf to measure their length, then measure the shortest distance from the top of the serration to the bottom of the serration as the serration length, and finally record the average of the nine values for the three leaves.
[0084] Please see the attached Figure 11 -Attached Figure 13 , used for genome-wide association analysis (GWAS) of camellia: the EMMAX mixed linear model was used, the genetic similarity matrix was used as the random effect variance-covariance matrix, and significant SNPs were screened through the Manhattan plot.
[0085] Specifically, GWAS analysis was performed using a mixed linear model (MLM) in EMMAX. To avoid false positives, a pairwise genetic similarity matrix (KinshipMatrix) calculated based on a simple matching coefficient was used as the variance-covariance matrix for the random effects. Manhattan plots of all SNPs were evaluated to identify the most significant SNPs for subsequent analysis.
[0086] To evaluate the capability of the array for GWAS analysis, a GWAS analysis was performed for two representative traits, leaf tip length and leaf margin serration length, using the mean of two measurements. Three SNPs significantly associated with leaf tip length and one SNP significantly associated with leaf margin serration length were identified at a threshold of 1e-4. Functional annotation of genes within the 200 kb region upstream and downstream of the locus was performed. SNPChr11_103989725 was significantly associated with leaf tip length (p=5.18×10-5). The three haplotypes at this locus (AA / AT / TT) showed significant differences in both the two measurements and the mean. The AA haplotype was associated with the longest leaf tip, followed by the AT haplotype, and finally the TT haplotype. Twelve genes were annotated within the 200 kb region upstream and downstream of the locus. Among them, gene EVM0035365.1, located 8.291 bp from the locus, encodes a protein phosphatase and belongs to the PP2C family. This study reveals that multiple genes in the PP2C family have direct or indirect effects on the proliferation of stem cells within the plant meristem.
[0087] Please see the attached Figure 14-15 , used for genomic selection analysis of Camellia genomics: four models, GBLUP, BayesA, BayesB, and BayesC, were used. The GBLUP model was y = Xβ + Zμ + ε, implemented by the mmer function of the R package sommer; the Bayes model was y = Xβ + Wα m +∈, implemented by the BGLR function of the R package BGLR;
[0088] Where y is the phenotypic value vector, β is the fixed effect vector, X is the association matrix of β, W is the SNP marker genotype score matrix, α m is the SNP marker random effect, ∈ is the error effect;
[0089] The prediction ability was evaluated using 5-fold cross-validation, and the heritability H was estimated based on GBLUP. 2 , and integrate GWAS significant SNPs.
[0090] Specifically, genomic prediction of five leaf-related traits was performed based on 220 samples. Four GS models were selected, and GBLUP was implemented using the 'mmer' function in the sommer package in R. gBLUP model:
[0091] y=Xβ+Zμ+∈
[0092] where y is the phenotypic value vector, β is the fixed effect vector, μ is the random effect vector, X and Z are the occurrence matrices of the fixed and random effects, respectively, and ∈ is the residual.
[0093] BayesA, BayesB, and BayesC are implemented using the 'BGLR' function in the R package BGLR version 1.1.0. The general Bayesian model is:
[0094] y=Xβ+Wα m +∈
[0095] Where: y is the phenotypic value vector, β is the fixed effect vector, X is the association matrix of β, W is the SNP marker genotype score matrix, α m is the SNP marker random effect, and ∈ is the error effect.
[0096] The reference population was randomly divided into five subsets using a 5-fold cross-validation method. Four of these were used as training populations, and the remaining one was used as a validation population. This process was repeated 100 times. The predictive ability (PA) was assessed by calculating the Pearson correlation coefficient between the phenotypic values of the validation population and the predicted genomic breeding values (GEBVs). The GBLUP model was used to estimate the variance components of each trait and to estimate the broad-sense heritability (H2). PA / √H was used. 2 Evaluate prediction accuracy (PC).
[0097] The population was analyzed using the multilocus analysis model 3VmrMLM, based on the IIIvmrMLM software package, to identify quantitative trait nucleotides (QTNs) significantly associated with the target trait. This model employed a compressed variance components mixed model approach to estimate additive and dominant effects, as well as their interactions with environmental and penetrant effects, ensuring comprehensive control for all potential polygenic backgrounds. The model threshold was set to LOD = 3.
[0098] To evaluate the performance of the array in genomics prediction, five leaf-related traits were selected. Using five-fold cross-validation, 220 samples were randomly divided into training and validation cohorts for genomics analysis. Four different models (GBLUP, BayesA, BayesB, and BayesC) and five different marker subset sizes (1k, 5k, 10k, 15k, and 21k, respectively) were investigated for their predictive power. The results showed that all models had low predictive power at 1k markers, but provided valuable genomic predictions at 5k markers, reaching a plateau at 15k markers. For the three traits (leaf margin serration length LMSL), leaf width LW, and leaf tip length LTL), some models performed better at 5k markers than at 21k markers. Although some differences were not significant, using fewer markers significantly reduced computational time. Among the four GS models based on the 21k marker set, the Bayes C model showed higher prediction accuracy for leaf tip and leaf length; the GBLUP model showed higher prediction accuracy for leaf margin serration and leaf area, and the Bayes B model for leaf width. The heritability of five leaf-related traits was calculated based on the GBLUP model. Leaf width had the highest heritability of 0.436, which was considered high. Leaf tip and leaf area had medium heritabilities of 0.349 and 0.312, respectively. Leaf margin serration and leaf length had lower heritabilities of 0.258 and 0.29, respectively, which were considered low. Pearson correlation analysis showed a significant positive correlation between the heritability of the traits and their GS prediction ability (r = 0.849, P < 0.05). Given its short runtime and high efficiency, as well as its good performance across all traits, GBLUP was used for subsequent analyses and model optimization.
[0099] Based on different traits, significant SNPs from two GWAS models (MLM and HVMrMLM) were combined and added as fixed effects to the models to evaluate the impact of adding a priori SNPs on model prediction accuracy. The results showed that the PA and PC of all trait models were significantly improved (24.7%-64.7%). Among them, the model optimization for the LA trait was the most significant, with an improvement of 64.7% (PA: 0.184-0.303; PC: 0.362-0.597). Therefore, the results show that effective GS prediction requires the integration of key loci from GWAS results.
[0100] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A method for constructing a Camellia whole genome liquid phase chip, characterized in that: It includes the following steps: Collect camellia samples, and prepare high-quality genomic DNA from fresh leaves using the magnetic bead-mediated CTAB method, where the DNA satisfies 1.8 < OD260 / 280 < 2.0 and OD260 / 230 > 2.0; Fragment the genomic DNA into fragments of 200 - 300 bp, and construct a library through end repair, adapter ligation, and Pre-PCR amplification; Hybridize the probe with the target region, capture the hybridized probe using streptavidin-labeled magnetic beads, and perform Post-PCR amplification after enriching the target fragment; Sequence the target region using the DNBSEQ-T7 high-throughput sequencing platform, and convert the original image data into RawReads; Filter RawReads: Remove reads with adapters, where the proportion of bases with N content exceeding 10% and quality value lower than 10 is more than 50%, to obtain CleanReads; Map CleanReads to the reference genome using the bwa software, perform individual VCF calling based on the HaplotypeCaller method of GATK v.4.1.2.0, perform population SNP calling by merging VCFs, and use GATK hard filtering to exclude false-positive variations with parameters QD < 2.0, MQ < 40.0, FS > 60.0, SOR > 3.0, MQRankSum < -12.5, QUAL < 30.0, ReadPosRankSum < -8.
0.
2. The method for constructing a Camellia whole genome liquid phase chip according to claim 1, characterized in that: The camellia samples include samples for testing SNP array efficiency, leaf trait determination, genome re-sequencing, and multi-source genetic structure and phylogenetic analysis.
3. The method for constructing a Camellia whole genome liquid phase chip according to claim 1, characterized in that: The construction method further includes a locus screening process: Based on the camellia variety re-sequencing data, perform quality filtering, redundancy filtering, probe design evaluation, LD screening, and merging and deduplication with genetic linkage map data in sequence, and finally obtain target SNPs.
4. The method for constructing a Camellia whole genome liquid phase chip according to claim 3, characterized in that: The quality filtering parameters include a missing rate < 0.2, a minimum quality score ≥ 50, and a minimum allele frequency > 0.
15.
5. The method for constructing a Camellia whole genome liquid phase chip according to claim 3, characterized in that: The redundancy filtering requires that the GC content of the 100 bp sequence upstream and downstream of the SNP locus is 40% - 60%, without N bases and is a single copy in the whole genome.
6. An application of a Camellia whole genome liquid phase chip, used in the method for constructing a Camellia whole genome liquid phase chip according to any one of claims 1 to 5, characterized in that: For the phylogenetic analysis of camellia, it includes: calculating the genetic distance using VCF2Dis, constructing a neighbor-joining tree based on the TreeBest software, setting the Bootstrap value to 1000, and visualizing through the iTOL online software.
7. The use of a Camellia whole genome liquid phase chip according to claim 6, characterized in that: For the population structure analysis of camellia: Infer the population structure using Admixture and visualize through pophelper; calculate the principal components using Plink2, and perform clustering analysis on the first two principal components using the R language ggplot2.
8. The use of a Camellia whole genome liquid phase chip according to claim 6, characterized in that: [[ID= 9. The use of a Camellia whole genome liquid phase chip according to claim 6, characterized in that: For the genome-wide association study of camellia genotype GWAS: the EMMAX mixed linear model was used, the genetic similarity matrix was used as the random effect variance-covariance matrix, and significant SNPs were screened by Manhattan plot.
10. The use of a Camellia whole genome liquid phase chip according to claim 6, characterized in that: For genomic selection analysis of Camellia genomics: four models were used: GBLUP, BayesA, BayesB, and BayesC. The GBLUP model was y = Xβ + Zμ + ε, implemented by the mmer function of the R package sommer; the Bayes model was y = Xβ + Wα m +∈, implemented by the BGLR function of the R package BGLR; Where y is the phenotypic value vector, β is the fixed effect vector, X is the association matrix of β, W is the SNP marker genotype score matrix, α m is the SNP marker random effect, ∈ is the error effect; The prediction ability was evaluated using 5-fold cross-validation, and the heritability H was estimated based on GBLUP. 2 , and integrate GWAS significant SNPs.