Phenotype prediction method and system for fusing genome and phenotype group

Through the prediction method of fusing genome and phenotype group, a genome selection and phenotype group selection model was established, and the weight calculation method was used to solve the problem of insufficient accuracy of traditional prediction models in the prediction of low heritability complex traits, achieving higher prediction accuracy and stability.

CN120108491APending Publication Date: 2025-06-06NANJING AGRICULTURAL UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510144669.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Traditional genome selection models have insufficient accuracy in predicting complex traits with low heritability, and phenotypic selection mainly relies on indirect estimation of auxiliary traits, making it difficult to capture the genetic background of complex traits.

Method used

The phenotypic prediction method of fusion genome and phenotypic group is adopted. By obtaining the genotype data and phenotypic data of crops, a genomic selection prediction model and phenotypic group selection prediction model were established respectively, and the weight of the prediction results was calculated using the Deoptim or FastW method, and the weight sum was performed to obtain the phenotypic prediction results.

Benefits of technology

It significantly improves the accuracy of crop phenotype prediction and improves the stability of cross-environment prediction. Compared with a single genome selection or phenotype selection model, the prediction accuracy is improved and the calculation time is shortened.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108491A_ABST
    Figure CN120108491A_ABST
Patent Text Reader

Abstract

The invention discloses a genome and phenotype group fused phenotype prediction method and system, and the method comprises the steps: firstly obtaining genotype data and phenotype data of a to-be-detected crop, then respectively building a genome selection prediction model and a phenotype group selection prediction model, and carrying out the prediction through the genotype data and the phenotype data, obtaining a genome selection prediction result and a phenotype group selection prediction result; and finally, calculating the weights of the genome selection prediction result and the phenotype group selection prediction result by using a Deoptim method or a FastW method, and carrying out weighted summation to obtain the phenotype prediction result of the crop to be detected. According to the method, the advantages of genome and phenotype group prediction are fused, the prediction precision is improved, and the calculation time is shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to crop phenotype prediction, in particular to a phenotype prediction method and system integrating genome and phenotype group. Background Art

[0002] Food production is a strategic industry that ensures the safety of the country and the people. With the rapid development of life sciences, omics sequencing technologies are emerging one after another, providing key technical support for understanding complex life forms. The traditional breeding research paradigm is shifting from experience-driven (a small number of genes) to a data-driven paradigm based on large-scale omics sequencing technology (gene networks).

[0003] Genomic research has made breakthrough progress, and multidisciplinary disciplines such as next-generation sequencing technology and genetics are widely used in modern crop breeding. With the gradual maturity of technologies such as double haploid induction, the cost of crop inbred line production has been greatly reduced and the number has increased significantly. If you rely on field phenotypic testing and manually select hundreds of thousands of material combinations, the cost is so high that it is almost impossible to complete. With the help of nonlinear feature extraction and automatic modeling technology of machine / deep learning, ideal genome artificial intelligence screening and design can significantly improve breeding accuracy and efficiency.

[0004] Genomic selection (GS) is a method for predicting breeding values ​​based on whole-genome molecular markers, and is also an important technology for crop breeding. However, traditional genomic selection models often show insufficient accuracy in predicting complex traits with low heritability. Phenomic selection (PS) is regarded as a supplement and alternative to GS. It can utilize the genetic correlation between traits to achieve the purpose of accurate prediction. However, PS mainly relies on indirect estimation of auxiliary traits, which makes it difficult to capture the genetic background of complex traits. Summary of the invention

[0005] Purpose of the invention: The purpose of the present invention is to provide a phenotypic prediction method and system that integrates genome and phenotypic group data to address the limitations of using GS and PS alone, and to significantly improve the accuracy of crop phenotype prediction and the stability of cross-environment prediction by integrating the prediction results of genome data and phenotypic group data.

[0006] Technical solution: The phenotype prediction method of the present invention that integrates genome and phenotype group includes the following steps:

[0007] (1) Obtaining genotype data and phenotypic data of the crops to be tested;

[0008] (2) establishing a genome selection prediction model and a phenotype group selection prediction model respectively, using the genotype data and phenotype data in step (1) to perform predictions, and obtaining genome selection prediction results and phenotype group selection prediction results;

[0009] (3) The Deoptim method is used to calculate the weights of the genome selection prediction results and the phenotypic group selection prediction results, and the weighted sum of the genome selection prediction results and the phenotypic group selection prediction results is obtained to obtain the phenotypic prediction results of the crops to be tested.

[0010] Furthermore, in step (2), the genome selection prediction model and the phenotype group selection prediction model are both machine learning models or deep learning models.

[0011] Furthermore, in step (2), the genome selection prediction model and the phenotype group selection prediction model are both RF models.

[0012] The phenotype prediction system integrating genome and phenotype group of the present invention comprises:

[0013] A data acquisition unit, used to acquire genotype data and phenotype data of the crop to be tested;

[0014] A prediction unit, used to establish a genome selection prediction model and a phenotype group selection prediction model respectively, and use the genotype data and phenotype data in the data acquisition unit to make predictions to obtain genome selection prediction results and phenotype group selection prediction results;

[0015] The fusion unit is used to calculate the weights of the genome selection prediction results and the phenotypic group selection prediction results using the Deoptim method, and to perform weighted summation of the genome selection prediction results and the phenotypic group selection prediction results to obtain the phenotypic prediction results of the crops to be tested.

[0016] Another phenotype prediction method integrating genome and phenotype group according to the present invention comprises the following steps:

[0017] (1) Obtaining genotype data and phenotypic data of the crops to be tested;

[0018] (2) establishing a genome selection prediction model and a phenotype group selection prediction model respectively, using the genotype data and phenotype data in step (1) to perform predictions, and obtaining genome selection prediction results and phenotype group selection prediction results;

[0019] (3) using the FastW method to calculate the weights of the genomic selection prediction results and the phenotypic group selection prediction results, and performing weighted summation of the genomic selection prediction results and the phenotypic group selection prediction results to obtain the phenotypic prediction results of the tested crop; wherein the weights of the genomic selection prediction results and the phenotypic group selection prediction results calculated using the FastW method include:

[0020]

[0021] Among them, w g represents the weight of the genomic selection prediction result, w p represents the weight of the prediction result of phenotype group selection, h 2 represents the heritability of the target trait, and cor represents the absolute value of the correlation coefficient among the auxiliary traits with the highest correlation with the target trait.

[0022] Furthermore, in step (2), the genome selection prediction model and the phenotype group selection prediction model are both machine learning models or deep learning models.

[0023] Furthermore, in step (2), the genome selection prediction model and the phenotype group selection prediction model are both RF models.

[0024] Another phenotype prediction system integrating genome and phenotype group according to the present invention comprises:

[0025] A data acquisition unit, used to acquire genotype data and phenotype data of the crop to be tested;

[0026] A prediction unit, used to establish a genome selection prediction model and a phenotype group selection prediction model respectively, and use the genotype data and phenotype data in the data acquisition unit to make predictions to obtain genome selection prediction results and phenotype group selection prediction results;

[0027] The fusion unit is used to calculate the weights of the genomic selection prediction results and the phenotypic group selection prediction results using the FastW method, and perform weighted summation of the genomic selection prediction results and the phenotypic group selection prediction results to obtain the phenotypic prediction results of the tested crop; wherein the weights of the genomic selection prediction results and the phenotypic group selection prediction results calculated using the FastW method include:

[0028]

[0029] Among them, w g represents the weight of the genomic selection prediction result, w p represents the weight of the prediction result of phenotype group selection, h 2 represents the heritability of the target trait, and cor represents the absolute value of the correlation coefficient among the auxiliary traits with the highest correlation with the target trait.

[0030] The electronic device described in the present invention includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is loaded into the processor, the phenotype prediction method based on the fused genome and phenotype group is implemented.

[0031] The electronic device described in the present invention includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is loaded into the processor, another phenotype prediction method of the fusion genome and phenotype group is implemented.

[0032] Beneficial effects: Compared with the prior art, the advantages of the present invention are: (1) the present invention utilizes the advantages of the genome and phenotype groups, and fuses the prediction results of the two, so that the prediction accuracy is improved compared with the single GS model and PS model; (2) the present invention uses the heritability of traits and the correlation coefficient between traits to construct a weight allocation method FastW, and the weights calculated by the method have a very high correlation with the weights obtained by multiple rounds of iterations of the global optimization algorithm (DEoptim), and the calculation time is greatly shortened. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 This is a comparison chart of the accuracy of different model fusion strategies in Example 1 of the present invention;

[0034] Figure 2 It is a correlation analysis diagram of weights calculated by the DEoptim algorithm and the FastW method in Example 2 of the present invention;

[0035] Figure 3 It is a time chart of weight calculation by the DEoptim algorithm and the FastW method in this embodiment 2. DETAILED DESCRIPTION

[0036] The technical solution of the present invention is further described below in conjunction with the accompanying drawings.

[0037] Example 1

[0038] The phenotype prediction method of the fusion genome and phenotype group includes the following steps.

[0039] (1) Data preparation

[0040] This example conducts experiments on four species (rice, corn, wheat, and soybean), and the data sets are from the following sources:

[0041] Rice comes from the Chinese Plant Gene Research Center (Wuhan), and this rice material represents the genetic diversity of cultivated varieties. The dataset provides genotypes and phenotypes of 529 rice materials, covering 10 agronomic traits, including HD (Heading Date), PH (Plant Height), NP (Number of Panicles), NEP (Number of Effective Panicles), YD (Yield), GW (Grain Weight), SL (Spikelet Length), GLH (Grain Length), GWH (Grain Width), and GT (Grain Thickness).

[0042] Maize comes from the Institute of Crop Science, Zhejiang University. The data set originally contained 326 maize materials. After screening the genome sequencing data and filtering out varieties with low marker numbers and less phenotypic data, 244 materials were finally obtained with genotype and phenotypic data, of which the phenotypic data covered seven agronomic traits, including DA (Day to Anthesis), EH (Ear Height), LL (Leaf Length), LW (Leaf Width), LLA (Lower Leaf Angle), PH (Plant Height), and ULA (Upper Leaf Angle).

[0043] Wheat comes from the wheat gene bank of the International Maize and Wheat Improvement Center (CIMMYT), which contains the genotypes and phenotypes of 2,000 Iranian wheat (Triticum aestivum) local varieties, covering 8 agronomic traits, including GL (Grain Length), GW (Grain Width), GH (Grain Hardness), TKW (Thousand Kernel Weight), TW (Test Weight), SDS (Sodium Dodecyl Sulfate), GP (Grain Protein), and PH (Plant Height).

[0044] The SoyNAM dataset (hereinafter referred to as Soybean) comes from the Agricultural Research Service of the United States Department of Agriculture and includes 4,312 soybean materials. In order to meet the needs of subsequent model mobility analysis, this study selected 1,260 soybean materials from three years and three locations. These materials include genotype and phenotypic data, covering seven agronomic traits, including PH (Plant Height), YD (Yield), ME (Moisture), PT (Protein), OIL (Oil), FR (Fiber), and GW (Grain Weight).

[0045] (2) Experimental methods

[0046] This example uses 5 machine learning models (RF, Lasso, SVM, XGBoost and LightGBM) and 1 deep learning model (DNNGP) to perform genome selection prediction and phenotypic group selection prediction on data sets of 4 species. Each model evaluates its performance using a ten-fold cross-validation approach. The data set is randomly divided into ten subsets: nine subsets are used for training and validation, and the remaining subsets are used for testing. This process is repeated ten times to ensure a robust evaluation of model performance.

[0047] The above 6 models are introduced as follows:

[0048] RF is an ensemble learning method that combines the predictions of multiple decision trees to improve the accuracy and stability of the model. Due to the integration of multiple decision trees, RF is robust to noise and outliers in the training data. RF is commonly used in GS and PS to extract important features. We run RF using the R package ranger, with all parameters set to default values.

[0049] Lasso is a regression analysis method mainly used for variable selection and regularization, especially for high-dimensional data sets such as genomic data. Lasso can automatically select variables in regression models, facilitate the identification of key predictive features, and eliminate features with low correlation by shrinking coefficients to zero, thereby improving the accuracy and efficiency of prediction. We run Lasso using the R package glmnet, with all parameters set to default values.

[0050] SVM is a supervised learning algorithm for classification and regression tasks. Its core idea is to find an optimal decision boundary to distinguish different feature classes. SVM has demonstrated its robust performance under high-dimensional data and noise by introducing kernel functions. We used the R package e1071 to run SVM, and all parameters were set to default values.

[0051] XGBoost is an efficient gradient boosting algorithm that is widely used in classification and regression tasks. Unlike traditional gradient boosting methods, it introduces regularization terms to reduce overfitting and improve model generalization. XGBoost is computationally efficient and uses parallel computing, cache optimization, and sparse matrix processing techniques to accelerate training. In addition, it can automatically handle missing values, which is critical for genomic and phenotypic group predictions because missing data is common in these data sets. We run XGBoost using the R package xgboost, with all parameters set to default values.

[0052] LightGBM is very efficient in handling high-dimensional features and large-scale structured data, which makes it extremely valuable in future big data-driven genomic prediction. By improving the selection of split points and gain calculation in LightGBM, the model accuracy can be improved and sparse matrices and large-scale features can be handled efficiently. The key difference between LightGBM and other decision tree models is that LightGBM grows leaf by leaf rather than layer by layer, a strategy that has been shown to improve model accuracy. We run LightGBM using the R package lightgbm with all parameters set to default values.

[0053] DNNGP is a genomic prediction model using deep convolutional neural networks. It achieves prediction accuracy comparable to machine learning models and demonstrates high accuracy among existing deep learning models. By introducing PCA for dimensionality reduction in DNNGP, computational time and memory consumption can be reduced while improving prediction accuracy and efficiency. Compared with other deep learning models such as DeepGS, DLGWAS, and SoyDNGP, DNNGP's dimensionality reduction technology integrates genomic data.

[0054] (3) Result Fusion

[0055] The genome prediction model and phenotype group prediction model are trained and predicted respectively, and the DEoptim algorithm is used for global optimization to merge the prediction results. The DEoptim algorithm optimizes the model weight distribution according to the prediction accuracy to ensure that the prediction results remain stable and close to the true value. The phenotype data consists of n×m 1 The matrix M 1 Indicates that n is the sample size, m is 1 is the number of phenotypes. The genomic data consists of n×m 2 The matrix M 2 Indicates that m 2 is the number of SNPs. 1 Input into the model to generate the result y 1 , while M 2 Input into the model to generate the result y 2Then, the DEoptim algorithm is used to optimize the prediction and assign weights w 1 and w 2 , and get the final result r 1 , where r 1 =w 1 ×y 1 +w 2 ×y 2 .

[0056] In machine learning models, M 1 and M 2 are independently input into the model for training and prediction, generating y 1 and 2 Then, the DEoptim algorithm assigns the optimal weights and obtains the final prediction r 1 In the deep learning model, since DNNGP needs to reduce the dimension of genomic data, M 2 After dimensionality reduction, it becomes n×m 4 The matrix M 4 , where m 4 is the number of features after dimensionality reduction using Plink. 1 and M 4 As input, it generates the result y 1 and 4 The DEoptim algorithm assigns weights w to the genomic and phenotypic group predictions. 1 and w 4 , and get the final result r 2 , where r 2 =w 1 ×y 1 +w 4 ×y 4 .

[0057] The accuracy of the above models was compared on four datasets. To account for the differences in traits among the datasets, three agronomic traits were randomly selected from each dataset to evaluate the prediction accuracy of each method. In the rice dataset, the selected traits were plant height (PH), yield (YD), and grain weight (GW); in the maize dataset, leaf height (LH), ear height (EH), and plant height (PH) were selected; in the wheat dataset, test weight (TW), grain protein (GP), and grain hardness (GH) were selected; in the soybean dataset, yield (YD), plant protein (PT), and grain weight (GW) were selected.

[0058] To maintain consistency in training and testing data between models, and because deep learning models require validation sets to prevent overfitting, all datasets were uniformly divided into training, validation, and test sets using a ratio of 8:1:1. The five machine learning models used training and test sets, while the deep learning models used training, validation, and test sets. The correlation coefficient between the predicted and true values ​​was used to evaluate model accuracy.

[0059] like Figure 1 Shown is the accuracy comparison of different model fusion strategies of this embodiment, that is, the prediction results of genome prediction, phenotypic group prediction, and fusion of genome prediction and phenotypic group prediction are compared using one model. In the figure, DNNGPg, DNNGPp, and DNNGP_R respectively represent genome prediction, phenotypic group prediction, and result fusion prediction using the DNNGP model, and the same symbol combination rules apply to other models. Points of the same color in the figure represent the prediction accuracy of different traits in the same species. The top horizontal line of each box represents the upper quartile, the middle horizontal line represents the median, and the bottom horizontal line represents the lower quartile. The vertical lines extending from the box represent the maximum and minimum values, and the points outside the range of these vertical lines represent outliers. The accuracy "Accuracy" in the figure represents the correlation coefficient between the predicted value and the true value.

[0060] The results show that the prediction accuracy after data fusion is significantly higher than that of the single GS and PS models, among which the model with the highest accuracy is RF_R, which is 57.03% and 14.06% higher than RFg and RFp respectively. In addition, the application of the invention on a variety of crops proves the universality and efficiency of the invention.

[0061] Example 2

[0062] (1) Data preparation

[0063] This example conducts experiments on four species (rice, corn, wheat, and soybean), and the source of the data set is the same as that of Example 1, which will not be repeated here.

[0064] (2) Experimental methods

[0065] This example uses 5 machine learning models (RF, Lasso, SVM, XGBoost and LightGBM) and 1 deep learning model (DNNGP) to perform genome selection prediction and phenotypic group selection prediction on data sets of 4 species. Each model evaluates its performance using a ten-fold cross-validation approach. The data set is randomly divided into ten subsets: nine subsets are used for training and validation, and the remaining subsets are used for testing. This process is repeated ten times to ensure a robust evaluation of model performance.

[0066] The above 6 models are the same as those in Example 1 and will not be described in detail here.

[0067] (3) Result Fusion

[0068] The genome prediction model and the phenotype group prediction model are trained and predicted respectively. The DEoptim algorithm in Example 1 requires time and resources to perform multiple rounds of iterations. Therefore, this example uses the FastW method to calculate the weight w of the genome prediction result. g and the weight w of the phenotype group prediction results p , and finally get the crop phenotype prediction result pred_final:

[0069]

[0070] pred_final = w g *pred g +w p *pred p ;

[0071] pred g Indicates the genome prediction result, pred p represents the phenotype group prediction result, h 2 The heritability of the target trait is calculated using the mixed linear model of the GCTA (Genome-wide Complex Trait Analysis) software to obtain the broad heritability of each trait h 2 ; cor represents the absolute value of the correlation coefficient with the target trait among the auxiliary traits, and the correlation coefficient between traits is calculated using the Pearson correlation coefficient.

[0072] The accuracy of the above models was compared on four datasets. To account for the differences in traits among the datasets, three agronomic traits were randomly selected from each dataset to evaluate the prediction accuracy of each method. In the rice dataset, the selected traits were plant height (PH), yield (YD), and grain weight (GW); in the maize dataset, leaf height (LH), ear height (EH), and plant height (PH) were selected; in the wheat dataset, test weight (TW), grain protein (GP), and grain hardness (GH) were selected; in the soybean dataset, yield (YD), plant protein (PT), and grain weight (GW) were selected.

[0073] like Figure 2The figure shows the correlation analysis of the weights calculated by the DEoptim algorithm of Example 1 and the FastW method of this example. Both the genome selection prediction and the phenotypic group selection prediction adopt the RF model. DEoptim_w represents the weights assigned by DEoptim, while FastW_w represents the weights calculated based on heritability and correlation coefficient. Green represents the weight values ​​of the three traits LH (leaf length), EH (ear height) and PH (plant height) in corn under two weight distribution methods, and their correlation coefficients are 0.96, 0.99 and 0.99 respectively; blue represents the weight values ​​of the three traits YD (yield), PT (protein) and GW (grain weight) in soybean under two weight distribution methods, and their correlation coefficients are 0.77, 0.99 and 0.94 respectively; gray represents the weight values ​​of the three traits PH (plant height), YD (yield) and GW (grain width) in rice under two weight distribution methods, and their correlation coefficients are 0.98, 0.86 and 0.98 respectively; yellow represents the weight values ​​of the three traits TW (test weight), GP (grain protein) and GH (grain hardness) in wheat under two weight distribution methods, and their correlation coefficients are 0.97, 0.82 and 0.96 respectively. G_w represents the weight of genome prediction, and P_w represents the weight of phenotypic prediction.

[0074] The results show that the weights calculated by FastW are highly correlated with the weights assigned by the global optimization algorithm DEoptim, with an average correlation coefficient of 0.93, indicating that the FastW method of this embodiment is comparable to the global optimization algorithm in terms of accuracy.

[0075] like Figure 3 The figure shows the time analysis of the weight calculation of the DEoptim algorithm in Example 1 and the FastW method in this example. The ordinate of each model in the figure represents the time required to run the fusion of the genome and phenotype group data. The results show that the time consumed by DEoptim is 2-3 times that of FastW, indicating that the FastW method in this example is more efficient.

Claims

1. A phenotype prediction method integrating genome and phenotype groups, characterized in that: The steps include: (1) Obtaining genotype data and phenotypic data of the crops to be tested; (2) establishing a genome selection prediction model and a phenotype group selection prediction model respectively, using the genotype data and phenotype data in step (1) to perform predictions, and obtaining genome selection prediction results and phenotype group selection prediction results; (3) The Deoptim method is used to calculate the weights of the genome selection prediction results and the phenotypic group selection prediction results, and the weighted sum of the genome selection prediction results and the phenotypic group selection prediction results is obtained to obtain the phenotypic prediction results of the crops to be tested.

2. The phenotype prediction method of the fusion genome and phenotype group according to claim 1, characterized in that: In step (2), the genome selection prediction model and the phenotype group selection prediction model are both machine learning models or deep learning models.

3. The phenotype prediction method of the fusion genome and phenotype group according to claim 2, characterized in that: In step (2), the genome selection prediction model and the phenotype group selection prediction model are both RF models.

4. A phenotype prediction system based on the method of claim 1 that integrates genome and phenotype group, characterized in that: include: A data acquisition unit, used to acquire genotype data and phenotype data of the crop to be tested; A prediction unit, used to establish a genome selection prediction model and a phenotype group selection prediction model respectively, and use the genotype data and phenotype data in the data acquisition unit to make predictions to obtain genome selection prediction results and phenotype group selection prediction results; The fusion unit is used to calculate the weights of the genome selection prediction results and the phenotypic group selection prediction results using the Deoptim method, and to perform weighted summation of the genome selection prediction results and the phenotypic group selection prediction results to obtain the phenotypic prediction results of the crops to be tested.

5. A phenotype prediction method integrating genome and phenotype groups, characterized in that: The steps include: (1) Obtaining genotype data and phenotypic data of the crops to be tested; (2) establishing a genome selection prediction model and a phenotype group selection prediction model respectively, using the genotype data and phenotype data in step (1) to perform predictions, and obtaining genome selection prediction results and phenotype group selection prediction results; (3) using the FastW method to calculate the weights of the genome selection prediction results and the phenotypic group selection prediction results, and performing weighted summation of the genome selection prediction results and the phenotypic group selection prediction results to obtain the phenotypic prediction results of the crop to be tested; The FastW method is used to calculate the weights of the genomic selection prediction results and the phenotypic group selection prediction results, including: Among them, w g represents the weight of the genomic selection prediction result, w p represents the weight of the prediction result of phenotype group selection, h 2 represents the heritability of the target trait, and cor represents the absolute value of the correlation coefficient among the auxiliary traits with the highest correlation with the target trait.

6. The phenotype prediction method of the fusion genome and phenotype group according to claim 5, characterized in that: In step (2), the genome selection prediction model and the phenotype group selection prediction model are both machine learning models or deep learning models.

7. The phenotype prediction method of the fusion genome and phenotype group according to claim 6, characterized in that: In step (2), the genome selection prediction model and the phenotype group selection prediction model are both RF models.

8. A phenotype prediction system based on the method of claim 5 that integrates genome and phenotype groups, characterized in that: include: A data acquisition unit, used to acquire genotype data and phenotype data of the crop to be tested; A prediction unit, used to establish a genome selection prediction model and a phenotype group selection prediction model respectively, and use the genotype data and phenotype data in the data acquisition unit to make predictions to obtain genome selection prediction results and phenotype group selection prediction results; A fusion unit, used to calculate the weights of the genome selection prediction results and the phenotypic group selection prediction results using the FastW method, and perform weighted summation of the genome selection prediction results and the phenotypic group selection prediction results to obtain the phenotypic prediction results of the crop to be tested; The FastW method is used to calculate the weights of the genomic selection prediction results and the phenotypic group selection prediction results, including: Among them, w g represents the weight of the genomic selection prediction result, w p represents the weight of the prediction result of phenotype group selection, h 2 represents the heritability of the target trait, and cor represents the absolute value of the correlation coefficient among the auxiliary traits with the highest correlation with the target trait.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the computer program is loaded into a processor, the phenotype prediction method of the fusion genome and phenotype group according to any one of claims 1 to 3 is implemented.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the computer program is loaded into a processor, the phenotype prediction method of the fusion genome and phenotype group according to any one of claims 5 to 7 is implemented.

Citation Information

Cited By

  • Crop phenotype prediction method and crop phenotype prediction device

    CN121237200A