Tea tree precise breeding phenotype prediction method based on GWAS and eQTL analysis
Through the tea tree precision breeding phenotype prediction method based on GWAS and eQTL analysis, the problems of tea tree genome complexity and long phenotype identification cycle are solved, and efficient early phenotype prediction and breeding efficiency are achieved.
Patent Information
- Application Number
- CN202510138734.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-06-27
AI Technical Summary
The tea tree has a huge genome and high heterozygous degree, which leads to difficulty in screening excellent phenotypes. The traditional breeding methods are inefficient, and the tea tree has a long growth cycle and a low hybrid fruiting rate.
The tea tree precision breeding phenotype prediction method based on GWAS and eQTL analysis allows efficient early phenotype prediction by collecting genomic data and phenotype data, preprocessing data, building learning models, and conducting model training and testing.
It achieves efficient early phenotypic prediction, improves the selection accuracy during breeding, shortens the breeding cycle, screens out individuals with excellent characteristics, and accelerates the selection and breeding of tea tree varieties.
Smart Images

Figure CN120220800A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of tea tree breeding, and specifically relates to a method for predicting tea tree precise breeding phenotypes based on GWAS and eQTL analysis. Background Art
[0002] The tea tree (Camellia sinensis) is a perennial crop with important economic value. Key agronomic traits such as the budding period, resistance, and aroma components directly affect the yield and quality of tea. However, due to the large and highly heterozygous tea tree genome, it is difficult to screen for excellent phenotypes. At the same time, the tea tree has a long growth cycle and a low hybridization seed setting rate, making traditional breeding methods inefficient. Therefore, developing a technical method that can efficiently predict adult phenotypes at an early stage of the tea tree has become the key to improving breeding efficiency.
[0003] Based on genome-wide association analysis (GWAS) and expression quantitative trait locus (eQTL) analysis, and combined with machine learning algorithms, the present invention proposes an efficient method for predicting tea tree phenotypes, which can quickly screen out excellent genotypes of target traits, greatly improve breeding efficiency, and shorten the breeding cycle. Summary of the Invention
[0004] In order to make up for the deficiencies of the prior art, the present invention provides a method for predicting tea tree precise breeding phenotypes based on GWAS and eQTL analysis, which directly uses genotype information for the construction of a phenotype prediction model, thereby realizing efficient early phenotype prediction to improve the selection accuracy in the breeding process.
[0005] The technical problems to be solved by the present invention can be achieved through the following technical solutions:
[0006] The method for predicting tea tree precise breeding phenotypes based on GWAS and eQTL analysis specifically includes the following steps:
[0007] Step 1: Collect data and perform data annotation;
[0008] Step 2: Preprocess the annotated data;
[0009] Step 3: Construct a learning model;
[0010] Step 4: Model training and testing.
[0011] Further, the data in Step 1 includes genomic data and phenotypic data. The genomic data is the resequencing data of collected tea germplasm resources. Through GWAS and eQTL analysis, important loci related to multiple target agronomic traits and aroma components are identified. Under the condition of setting the threshold to 1e-2, key molecular markers affecting the phenotype are initially obtained. The phenotypic data is the record and precise detection of multiple adult agronomic phenotypes and various aroma metabolite components of the tea tree.
[0012] Furthermore, the specific process of data preprocessing in Step 2 is as follows:
[0013] Step 21, data cleaning and standardization: Clean, denoise, and standardize the genomic data and phenotypic data;
[0014] Step 22, feature screening: Through the GWAS and eQTL results, screen out the gene loci and expression features that are most valuable for phenotypic prediction, optimize the feature input, and improve the prediction accuracy and model stability;
[0015] Step 23, dimensionality reduction processing: Perform dimensionality reduction on the features to reduce the computational complexity and avoid overfitting.
[0016] Furthermore, in Step 22, the specific process of feature screening is as follows:
[0017] Extract single nucleotide polymorphism (SNP) locus information from the genotype data;
[0018] Read the significant major loci screened by GWAS and eQTL from the trait-SNP association file;
[0019] Through genome-wide association analysis, screen out the SNP loci significantly associated with the target trait (based on the significance threshold, P value < 0.01). At the same time, use expression quantitative trait locus eQTL analysis to further screen out the gene expression loci related to the target trait.
[0020] Furthermore, in Step 23, the specific process of dimensionality reduction processing is as follows:
[0021] Use GWAS and eQTL analysis to screen out the SNP loci significantly associated with the target trait, significantly reducing the feature dimension of the genotype data;
[0022] Integrate important covariates by combining the kinship matrix to further optimize the feature input;
[0023] Fill in the missing values and standardize the screened feature matrix to eliminate the dimension difference and noise interference;
[0024] Some algorithms (such as lightweight gradient boosting machine, random forest regression, gradient boosting regression) automatically ignore the features with lower contributions through the built-in feature selection function, further realizing the dynamic optimization of the feature dimension and ensuring the efficiency and generalization ability of the model.
[0025] Further, in step 3, the learning model adopts six machine learning algorithms, including support vector regression, light gradient boosting machine, random forest regression, linear regression, gradient boosting regression, and Bayesian ridge model. A population kinship matrix is added on the basis of the algorithms, and multiple models are comprehensively compared. Finally, the best-performing model, the Bayesian ridge plus population kinship matrix algorithm, is selected.
[0026] Further, in step 3, combining the key locus information identified by GWAS and eQTL, a feature combination is designed, and the model parameters are optimized through 10-fold cross-validation to make the accuracy of the model in predicting adult phenotypes at the seedling stage reach the optimal level; among them, the feature combination is designed as follows:
[0027] Step 31: Eliminate highly collinear loci, select key SNP loci with a significance level (P<0.01) and eQTL loci that significantly regulate target traits to ensure the independence of input features;
[0028] Step 32: Combine the screened SNP loci with eQTL and the kinship matrix to form the final input feature matrix;
[0029] Step 33: Perform Z-score standardization on the feature matrix to ensure that different features have the same dimension during model training.
[0030] Further, the specific content of step 4 includes:
[0031] Step 41: Division of the training set and the test set: Divide the data into a training set and a test set. The test set does not participate in model training and is used to verify the generalization ability of the model;
[0032] Step 42: Performance evaluation: Use 10-fold cross-validation, mean square error (MSE), and coefficient of determination (R2) indicators to evaluate the prediction effect of the model, and prove the advantage of the learning model in breeding efficiency by predicting new populations.
[0033] Compared with the prior art, the present invention has the following advantages:
[0034] (1) Based on the screening of gene loci by GWAS and eQTL analysis, the present invention directly uses genotype information for the construction of a phenotype prediction model, thereby realizing efficient early phenotype prediction and improving the selection accuracy in the breeding process; and by integrating genomic data and phenotype data, a phenotype prediction model is constructed using feature optimization and machine learning techniques, effectively solving the problems of the complexity of the tea tree genome and the long phenotype identification cycle.
[0035] (2) The method of the present invention can identify individuals with excellent characteristics in the seedling stage, thereby accelerating the selection and breeding of tea varieties, greatly shortening the breeding cycle, and is suitable for high-efficiency modern agricultural breeding; and by using 139 core germplasm resources of tea trees to verify the model prediction, 10 varieties with multiple excellent gene loci were successfully screened out. These varieties performed well in terms of germination period, quality, yield and resistance; this not only provides a precise direction for the improvement of tea varieties in theory, but also proves in practice that the model screening results are highly consistent with market demand.
[0036] (3) The method of the present invention screens out the superior genotypes that contribute the most to the phenotype. The BR_K (Bayesian Ridge plus population kinship matrix) model provides a basis for molecular marker chip detection for tea tree breeding, significantly shortens the breeding cycle, improves breeding efficiency, and provides a new technical path for the innovation and rapid breeding of tea varieties. The BR_K model was applied to the budding period and cold resistance analysis of 120 germplasm resources, further verifying its applicability and prediction accuracy in different phenotypes. The model can accurately reflect the performance changes of tea trees under different environmental conditions, indicating its wide applicability and potential in breeding, and is expected to promote the development of intelligent breeding technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 is a flow chart of the method of the present invention;
[0038] Figure 2 The determination coefficient R of multiple trait characteristics in the BR_K model was compared using a threshold of 0.01 in the experiment of this invention. 2 (A) and mean square error (MSE) (B) values; BF: budding period; HEX: hexadecanoic acid; DEC: linal; LIN: cis-linalool oxide; LDS: field disease index; AL: indoor anthracnose lesion size; LON: β-ionone; FDI: freeze damage index; BN: number of branches;
[0039] Figure 3 The top 10 cultivars with the highest comprehensive scores of the traits predicted by the experiment of the present invention are shown in the figure. Different colors represent the comprehensive scores assigned to various traits. HD: Huangyan; GRG1: Georgia 1; JM1: Jiaming 1; SC: Shanbuzhong; ZFC: Zaofengchun; SRG: Junhe Zaosheng; ZBD: Zaobaidian; SCZ: Shucha Zao; MKH: Muzhiyuan Zaosheng; CNYHZ: Chuannong Huangya Zao;
[0040] Figure 4In the specific experiment of the present invention, the germination period (A) and cold resistance phenotype (B) of 120 tea germplasm resources were predicted and analyzed using the training model (BR_K), *P<0.05; **P<0.01; ***P<0.001; ns indicates no significant difference (non-parametric test);
[0041] Figure 5 These are the top 10 key loci ranked by the BR_K model of the present invention for predicting the effect values of different traits. Specific embodiments
[0042] In order to make the objectives, technical solutions and advantages of the present invention clearer, the following further details the present invention in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0043] As Figure 1 shown, the precise breeding phenotype prediction method for tea trees based on GWAS and eQTL analysis specifically includes the following steps:
[0044] Step 1: Collect data and perform data annotation. The data includes genomic data and phenotypic data.
[0045] Genomic data: Resequencing data of 151 tea germplasm resources were collected, and important loci related to multiple target agronomic traits and aroma components were identified through GWAS and eQTL analysis. Under the condition of setting the threshold to 1e-2, 138,669 key molecular markers affecting the phenotype were initially obtained.
[0046] Phenotypic data: Record and accurately detect multiple adult agronomic phenotypes of tea trees (mainly including stress resistance, growth and development, germination period) and various aroma metabolite components.
[0047] Step 2: Preprocess the annotated data. The specific content is as follows:
[0048] (1) Data cleaning and standardization: Clean, denoise and standardize the genomic and phenotypic data to ensure the quality of the data.
[0049] (2) Feature screening: Through the GWAS and eQTL results, screen out the gene loci and expression features most valuable for phenotype prediction, optimize the feature input, and improve the prediction accuracy and model stability. The specific process is as follows:
[0050] ① Extract single nucleotide polymorphism (SNP) locus information from the genotype data;
[0051] ② Read the significant major effect loci screened by GWAS and eQTL from the trait-SNP association file;
[0052] ③Through genome-wide association studies (GWAS), SNP loci significantly associated with the target trait were screened (based on a significance threshold, P value < 0.01). Meanwhile, using expression quantitative trait locus (eQTL) analysis, gene expression loci related to the target trait were further screened out.
[0053] (3) Dimensionality reduction: Methods such as [specific methods] were used to reduce the dimensionality of the features, reducing computational complexity and avoiding overfitting. The specific process is as follows:
[0054] ①Using GWAS and eQTL analysis to screen SNP loci significantly associated with the target trait, significantly reducing the feature dimension of the genotype data;
[0055] ②Combining the kinship matrix to integrate important covariates to further optimize the feature input;
[0056] ③Filling in missing values and standardizing the screened feature matrix to eliminate dimensional differences and noise interference;
[0057] Some algorithms (such as LightGBM, Random Forest Regression, Gradient Boosting Regression) automatically ignore features with lower contributions through built-in feature selection functions, further achieving dynamic optimization of the feature dimension and ensuring the efficiency and generalization ability of the model.
[0058] Step 3: Construct a learning model.
[0059] The learning model of the present invention uses six machine learning algorithms, including Support Vector Regression (SVR), LightGBM, Random Forest Regression (RFC), Linear Regression (LR), Gradient Boosting Regression (GBR), and Bayesian Ridge (BR) model; a kinship matrix of the population was added to the algorithms, multiple models were comprehensively compared, and finally the best-performing model, the Bayesian Ridge plus kinship matrix of the population (BR_K) algorithm, was selected. Inputting the kinship matrix as a covariate together with the genotype data into the model can control the influence of the genetic background on trait prediction and improve the prediction accuracy.
[0060] Support Vector Regression (SVR): Using a kernel function (RBF kernel) to map the features to a high-dimensional space and finding an optimized hyperplane for regression, suitable for small-sample prediction problems.
[0061] LightGBM: An efficient implementation based on Gradient Boosting Decision Trees (GBDT), iteratively generating weak learners and improving prediction performance, with a built-in feature selection function, suitable for processing high-dimensional data.
[0062] Random Forest Regression (RFC): An ensemble method based on decision trees, constructing multiple decision trees by randomly sampling data and features and voting to output the prediction result; able to evaluate the importance of features and effectively handle non-linear relationships.
[0063] Linear Regression (LR): A simple linear model that uses the least squares method to fit the linear relationship between phenotypes and features, suitable for predicting traits with strong linear correlations.
[0064] Gradient Boosting Regression (GBR): Adopts the method of gradually optimizing the loss function to improve regression accuracy, and has good modeling effects for non-linear complex relationships.
[0065] Bayesian Ridge Regression: Introduces Bayesian priors on the basis of ridge regression, balances model complexity and prediction performance, is suitable for high-dimensional small-sample data, and can effectively prevent overfitting.
[0066] Then, combining the key locus information identified by GWAS and eQTL, design feature combinations, and optimize the model parameters through 10-fold cross-validation to make the accuracy of the model in predicting adult phenotypes at the seedling stage reach the optimal. The main evaluation indicators include mean squared error (MSE), R 2 value and Pearson correlation coefficient.
[0067] The feature combination design is as follows:
[0068] ① Eliminate highly collinear loci, select key SNP loci with a significance level (P < 0.01) and eQTL loci that significantly regulate target traits to ensure the independence of input features.
[0069] ② Combine the screened SNP loci with eQTL and the kinship matrix to form the final input feature matrix.
[0070] ③ Perform Z-score standardization on the feature matrix to ensure that different features have the same dimension during model training.
[0071] The specific steps for optimizing model parameters through 10-fold cross-validation are as follows:
[0072] ① Randomly divide the data into 10 equal parts, use 9 parts of the data as the training set and 1 part of the data as the validation set each time, and loop 10 times to ensure that each part of the data is used as the validation set;
[0073] ② For each machine learning model, optimize the key hyperparameters as follows:
[0074] SVR: Adjust the penalty parameter and kernel function parameter;
[0075] RFC: Optimize the number of trees and the maximum depth;
[0076] LGBM: Adjust the learning rate and the depth of the tree;
[0077] ③ Calculate the mean squared error (MSE), coefficient of determination (R 2 2) and Pearson correlation coefficient for each validation, and determine the optimal parameter combination of the model by averaging the evaluation metrics of each fold;
[0078] ④ Comprehensively compare the prediction performance of each model and select the optimal model.
[0079] Step 4: Model training and testing.
[0080] (1) Training set and test set division: Divide the data into a training set and a test set. The test set does not participate in model training and is used to verify the generalization ability of the model.
[0081] (2) Performance evaluation: Use indicators such as 10-fold cross-validation, mean squared error (MSE), coefficient of determination (R2) to evaluate the prediction effect of the model, and prove the advantage of this method in breeding efficiency by predicting a new population. The specific steps of performance evaluation are as follows:
[0082] ① 10-fold cross-validation: Randomly divide the data set into 10 non-overlapping subsets (folds); in each round of validation, select 1 of these subsets as the validation set, and the remaining 9 subsets as the training set; repeat the training and validation process 10 times, and each subset serves as the validation set once; finally, record the model performance indicators of each validation, and calculate the average value and standard deviation of all folds.
[0083] ② Mean squared error
[0084] Measure the average squared error between the predicted value and the true value, reflecting the overall deviation of the model prediction. The specific formula is as follows:
[0085]
[0086] where y i is the actual value, is the model predicted value, n is the number of samples, and the smaller the MSE, the better the model prediction effect.
[0087] ③ Coefficient of determination
[0088] Reflect the fitting degree of the model to the data, and the value range is [0, 1]. The closer the value is to 1, the better the model fitting effect.
[0089]
[0090] where is the mean of the actual values. When R 2 is negative, it indicates that the model prediction ability is poor.
[0091] ④ New population prediction verification
[0092] a) Further verify the applicability of the model in actual breeding by using the trained model to predict a new breeding population.
[0093] b) New data prediction: Input the genotype data and kinship matrix of the new population, and output the corresponding phenotypic prediction results.
[0094] c) Model stability verification: Observe the fluctuation of the prediction results of the new population to ensure the consistency of model performance among different samples.
[0095] Step 5. Experimental verification.
[0096] (1) Prediction accuracy
[0097] In multiple cross-validations, the learning model generally maintained high accuracy, and the average prediction accuracy for multiple major phenotypes (such as yield and disease resistance) exceeded 80%. Especially in the prediction of the key trait of tea plants - the budding period, the accuracy exceeded 90%. In addition, the average mean square error of the learning model for these phenotypes ranged from 0.029 to 0.232, far exceeding the statistical effect of traditional breeding methods in the early prediction of complex traits of tea plants, as Figure 2 shown.
[0098] (2) Optimal parent selection
[0099] To optimize the breeding of tea plant varieties, this application verified the model prediction using 139 released core germplasm resources of tea plants. The key goal of tea breeding is to cultivate varieties that are both high-quality and early-germinating, while also considering yield and resistance. Through weighted analysis of traits such as the budding period, quality, yield, and resistance, 10 varieties with multiple excellent gene loci were finally selected. As Figure 3 shown, these varieties include Huang Dan (HD), Georgia No. 1 (GRG1), Jia Ming No. 1 (JM1), Shanbu Variety (SC), Early Spring (ZFC), Junhe Zaosheng (SRG), Early White Tip (ZBD), Shucha Early (SCZ), Muyuan Zaosheng (MKH), and Chuannong Huangya Early (CNYHZ). These varieties are highly consistent with industry standards, fully demonstrating that the screening results meet market demands.
[0100] (3) Population phenotypic prediction
[0101] The learning model of this application was also applied to the analysis of the budding period and cold resistance of 120 germplasm resources to further test its prediction accuracy and wide applicability. As Figure 4 shown, the results indicate that the learning model can accurately reflect the changes in the budding period and cold resistance, demonstrating its powerful prediction ability and the potential to optimize trait selection in breeding programs.
[0102] (4) Integration of excellent genotypes
[0103] As Figure 5 shown, for different traits, the BR_K (Bayesian Ridge plus population relationship matrix) model screened out the top 10 excellent genotypes with the greatest impact on the phenotype. These genotypes can be applied to molecular marker chip detection, thus significantly shortening the breeding cycle of tea trees and significantly improving the breeding efficiency.
[0104] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for predicting tea plant precision breeding phenotypes based on GWAS and eQTL analysis, characterized in that: The specific steps include: Step 1: Collect data and label it; Step 2: Preprocess the labeled data; Step 3: Build a learning model; Step 4: Model training and testing.
2. The method for tea plant precision breeding phenotype prediction based on GWAS and eQTL analysis according to claim 1, characterized in that: The data in step 1 include genomic data and phenotypic data. The genomic data are resequencing data collected from tea germplasm resources. Through GWAS and eQTL analysis, important sites related to multiple target agronomic traits and aroma components are identified. When the threshold is set to 1e-2, key molecular markers affecting the phenotype are preliminarily obtained; the phenotypic data are to record and accurately detect multiple adult agronomic phenotypes of tea trees and multiple aroma metabolite components.
3. The method for tea plant precision breeding phenotype prediction based on GWAS and eQTL analysis according to claim 2, characterized in that: The specific process of data preprocessing in step 2 is: Step 21, data cleaning and standardization: cleaning, denoising and standardization of genomic data and phenotypic data; Step 22: Feature screening: Through GWAS and eQTL results, the most valuable gene loci and expression features for phenotype prediction are screened to optimize feature input; Step 23: Reduce the dimension of features to avoid overfitting.
4. The method for tea plant precision breeding phenotype prediction based on GWAS and eQTL analysis according to claim 3, characterized in that: In step 22, the specific process of feature screening is as follows: Extract single nucleotide polymorphism (SNP) site information from genotype data; Read the significant main effect sites screened by GWAS and eQTL from the trait-SNP association file; Through whole-genome association analysis, SNP sites significantly associated with the target traits were screened. Based on the significance threshold, P value <0.01, at the same time, expression quantitative trait locus eQTL analysis was used to further screen gene expression sites related to the target traits.
5. The method for tea plant precision breeding phenotype prediction based on GWAS and eQTL analysis according to claim 3, characterized in that: In step 23, the specific process of dimensionality reduction is as follows: GWAS and eQTL analysis were used to screen SNP sites significantly associated with target traits, significantly reducing the characteristic dimensions of genotype data; The kinship matrix was combined to integrate important covariates and further optimize feature input; The screened feature matrix is then filled with missing values and standardized to eliminate dimensional differences and noise interference.
6. The method for tea plant precision breeding phenotype prediction based on GWAS and eQTL analysis according to claim 1, characterized in that: The learning model in step 3 uses six machine learning algorithms, including support vector regression, lightweight gradient boosting machine, random forest regression, linear regression, gradient boosting regression and Bayesian Ridge model. The population kinship matrix is added to the algorithm, and multiple models are comprehensively compared to finally select the Bayesian Ridge plus population kinship matrix algorithm with the best prediction effect.
7. The method for tea plant precision breeding phenotype prediction based on GWAS and eQTL analysis according to claim 6, characterized in that: In step 3, the key loci information identified by GWAS and eQTL were combined to design feature combinations, and the model parameters were optimized through 10-fold cross validation; the feature combination design was as follows: Step 31: Eliminate highly collinear sites, select key SNP sites with P < 0.01 and eQTL sites that significantly regulate target trait genes, and ensure the independence of input features; Step 32, combining the screened SNP loci with the eQTL and the kinship matrix to form a final input feature matrix; Step 33: Perform Z-score standardization on the feature matrix to ensure that different features have the same dimension during model training.
8. The method for tea plant precision breeding phenotype prediction based on GWAS and eQTL analysis according to claim 1, characterized in that: The specific contents of step 4 include: Step 41: Divide the data into training set and test set. The test set does not participate in model training and is used to verify the generalization ability of the model. Step 42, performance evaluation: Use 10-fold cross validation, mean square error, and determination coefficient indicators to evaluate the prediction effect of the model, and prove the advantage of the learning model in breeding efficiency by predicting new populations.
Citation Information
Cited By
Efficient breeding decision support system and method based on Chinese trumpet creeper
CN120994983A
Method for improving stress resistance of secondary metabolites of tea trees
CN121687184A