Methods and applications for constructing genome-wide selection models based on multi-trait phenotype modeling

By combining machine learning models with genome-wide selection models, the relationships between multiple traits are captured, overcoming the limitations of single-trait prediction in existing technologies and achieving more efficient and accurate breeding selection.

CN118248207BActive Publication Date: 2025-10-31INSTITUTE OF CROP SCIENCE CHINESE ACADEMY OF AGRICULTURAL SCIENCES +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410240643.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-04
Publication Date
2025-10-31
Estimated Expiration
2044-03-04

AI Technical Summary

Technical Problem

Existing genome-wide selection models primarily predict individual traits, neglecting the genetic and environmental correlations between multiple traits. This results in limited prediction accuracy and selection effectiveness. Furthermore, the use of selection index methods has limited effectiveness in predicting target traits with high heritability and incurs significant computational burden.

Method used

Establish a genome-wide selection model based on multi-trait phenotypic modeling, capture linear or non-linear relationships between traits through machine learning models, and combine the genome-wide selection model with machine learning phenotypic models to predict and select target traits.

Benefits of technology

It improves the prediction accuracy and selection efficiency of target traits, is applicable to a variety of traits, reduces computing power requirements, provides more reliable breeding decision support, and is applicable to the fields of animal and plant breeding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118248207B_ABST
    Figure CN118248207B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of plant and animal breeding technology, and discloses a method and system for constructing a genome-wide selection model based on a multi-trait phenotypic model. The multi-trait model, established based on phenotypic data, uses the predicted values ​​or estimated breeding values ​​obtained from the genome selection model as input data for a machine learning phenotypic prediction model to predict the final phenotypic value for strain selection. A multi-trait machine learning phenotypic model is established to capture the linear or non-linear relationships between plant phenotypic traits. Based on this, the predicted values ​​and estimated breeding values ​​of each trait obtained from the genome selection model are combined. This invention improves the prediction accuracy of target traits, accelerates the breeding process, increases the selection efficiency of target traits, and saves breeding costs; it has wide applications in the fields of agricultural plant and animal breeding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of plant and animal breeding technology, and in particular relates to a method and application for constructing a genome-wide selection model based on multi-trait phenotypic modeling. Background Technology

[0002] Genomic selection (GS) is a method that uses genome-wide markers to assess an individual's genetic makeup and obtain estimated breeding values. It is primarily based on linkage disequilibrium, assuming that each quantitative trait locus (QTL) is in linkage disequilibrium with at least one marker and can explain most of the genetic variance. By constructing predictive models, GS can calculate estimated breeding values, enabling prediction and selection in early individuals, effectively shortening generation intervals, improving selection accuracy and breeding efficiency, accelerating the breeding process, and saving costs. Furthermore, GS demonstrates good predictive performance when dealing with complex traits with low heritability and difficult measurement, truly realizing the guidance of genomic technology in breeding practices. Although GS was first proposed in animal breeding, it has now made significant progress in plant breeding and is widely used in crops such as Arabidopsis thaliana, maize, wheat, barley, and soybean.

[0003] Statistical models, as the core of genome-wide prediction (GS), significantly influence the accuracy and efficiency of GS. Based on the chosen model, GS can be broadly categorized into direct methods, indirect methods, and GS based on machine learning or deep learning. Direct methods treat individuals as random effects, using the kinship matrix constructed from the genetic information of the training and validation populations as the variance-covariance matrix. Variance components are estimated iteratively, and then a mixed linear model is solved to obtain the estimated breeding value for the individual to be predicted. Direct methods are represented by GBLUP (Genomic Best Linear Unbiased Prediction), which is computationally efficient but has slightly lower accuracy. Indirect methods estimate marker effects in the training population and then sum these effects using genotype information from the prediction population to obtain the estimated breeding value for each individual. Indirect methods are represented by rrBLUP (Ridge Regression Best Linear Unbiased Prediction), BayesA, BayesB, BayesC, and Bayesian LASSO, but these methods are computationally intensive and relatively slow. Genomics-based (GS) models, utilizing machine learning and deep learning, can learn the relationships between genotypes and phenotypes in a training population from existing samples and infer phenotypic values ​​from the genotype data of a validation population. Representative methods include CropGBM (Crop Genomic Breeding Machine), DeepGS, and DNNGP (Deep Neural Network for Genomic Prediction). These methods consider multiple interactions and correlations between features, offering high computational efficiency and accuracy.

[0004] With continuous improvement and optimization of genomic selection (GS) statistical models, their stability and variety have increased. However, existing technologies still have some problems and limitations. Most GS models currently predict and select for single traits, neglecting the genetic basis between multiple related traits. Combining multiple phenotypic models can not only obtain genetic correlations between traits but also environmental correlations, potentially improving the accuracy of phenotypic prediction. While using selection index (SI) to assist in predicting target traits can construct a comprehensive index for joint selection of multiple traits using genetic correlations, this method has limited predictive power for highly heritable target traits. Furthermore, using auxiliary traits unrelated to the target trait risks reducing the predictive power of genomic selection for the target trait, requiring strict selection criteria for auxiliary traits. In addition, the computational demands increase with the number of auxiliary traits. Therefore, it is urgent to develop a genome-wide selection model with higher prediction accuracy, lower computational requirements, and more stable prediction results.

[0005] Based on the above analysis, the problems and shortcomings of the existing technology are as follows:

[0006] (1) Most current GS models mainly predict and select for a single trait, ignoring the genetic and environmental correlations between multiple traits. This method of predicting a single trait cannot make full use of the genetic associations between multiple traits, thus limiting the accuracy of prediction and the effectiveness of selection.

[0007] (2) Selection index can be used to combine the genetic correlation of multiple traits for joint selection. This method has limited effectiveness in predicting target traits with high heritability and has strict requirements for the selection of auxiliary traits. The more auxiliary traits there are, the greater the computational burden. Summary of the Invention

[0008] To address the current technological gaps, this invention provides a method and application for constructing a genome-wide selection model based on multi-trait phenotype modeling.

[0009] This invention establishes a machine learning model based on multi-trait phenotypic data to capture linear or non-linear relationships between traits. Then, the predicted values ​​or estimated breeding values ​​of multiple traits obtained from the GS prediction model are input into the multi-trait phenotypic model to predict the phenotypic phenotype of the target trait.

[0010] Furthermore, the method for constructing a genome-wide selection model based on a multi-trait phenotype model includes the following steps:

[0011] Step 1: Establish a multi-trait machine learning phenotypic model;

[0012] Step two involves combining genome-wide selection models with machine learning phenotypic models.

[0013] Furthermore, step one above primarily corresponds to the order in which the phenotypic data input into the machine learning phenotypic model and the predicted and estimated breeding values ​​of each trait obtained from the genome-wide selection model are arranged, to ensure data consistency and the effectiveness of capturing linear or nonlinear relationships between traits. The machine learning model can be selected from linear regression, logistic regression, support vector machines, decision trees, or random forests, depending on the target trait.

[0014] Building a machine learning multi-trait phenotypic model includes the following steps:

[0015] (1) Identify and clean up missing and outliers in the phenotypic data;

[0016] (2) Randomly rearrange the order of plant sample data in the phenotypic data;

[0017] (3) Separate the target trait from the phenotypic data and normalize the remaining traits as feature values ​​of the target trait;

[0018] (4) The phenotypic values ​​of each plant sample data after normalization are used as input data. The parameters and hyperparameters of the machine learning multi-phenotypic model are set according to the actual settings. The performance measurement is reflected by comparing the phenotypic prediction results with the target phenotypic values. The parameters are further adjusted, and the phenotypic prediction model with the best performance measurement is saved.

[0019] Furthermore, combining machine learning phenotypic prediction models with genome-wide selection models specifically includes the following steps ( Figure 1 ):

[0020] 1) Select a whole-genome selection model to predict all traits separately. The whole-genome model can be GBLUP, BayesA, BayesB, BayesC, Bayesian LASSO, DNNGP or CropGBM, etc.

[0021] 2) Perform data cleaning, processing, and normalization on the phenotypic predictions or estimated breeding values;

[0022] 3) The predicted or estimated breeding values ​​after the above data processing are represented as characteristic values ​​of the target trait;

[0023] 4) Load the phenotypic prediction model and input the feature values ​​for prediction. The model makes accurate predictions based on the input feature values ​​and outputs the final phenotypic prediction value that is close to the target trait.

[0024] The purpose of this invention is to provide a method for constructing a genome-wide selection model based on multi-trait phenotype modeling, comprising:

[0025] Data preprocessing unit: responsible for receiving and processing phenotypic data, including identifying and cleaning missing and outlier values ​​in the data, and randomly rearranging the order of plants;

[0026] Data normalization unit: used to normalize non-target traits in phenotypic data as feature values ​​and separate out target traits;

[0027] Machine learning training unit: Based on the normalized trait feature values, the machine learning phenotypic model is trained and saved by setting appropriate model parameters and hyperparameters;

[0028] Genomic selection model unit: Genetic evaluation of individuals is performed using whole-genome markers to obtain estimated genomic breeding values;

[0029] Model integration unit: combines machine learning phenotypic models and genomic selection models to perform data cleaning, processing and normalization on predicted values ​​and estimated breeding values;

[0030] Furthermore, the machine learning training unit adjusts the model based on the performance metrics of the predicted phenotypic values ​​and the target phenotypic values, and saves the phenotypic prediction model with the best performance metrics.

[0031] Furthermore, the model integration unit uses a phenotypic prediction model, inputs normalized feature values ​​for accurate prediction, and outputs a final phenotypic prediction value that is close to the target trait.

[0032] The purpose of this invention is to provide a method that combines machine learning algorithms and genome-wide selection models to improve the accuracy of prediction and selection efficiency of target traits.

[0033] Another objective of this invention is to provide a method for genome-wide selection using multi-trait phenotypic data, thereby achieving more scientific and efficient breeding selection.

[0034] Based on the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solution to be protected by this invention are as follows:

[0035] First, this invention comprehensively utilizes multi-trait data for joint analysis, effectively leveraging the genetic and environmental correlations between traits. This invention uses multi-trait phenotypic data to establish a machine learning phenotypic model, capturing linear or non-linear relationships between traits. Then, the predicted values ​​or estimated breeding values ​​of each trait obtained from a genome-wide selection model are input into the machine learning phenotypic model to predict the final phenotypic value and conduct line selection. This integrated model aims to simultaneously utilize the correlations between multiple traits, as well as between genotypes and traits, combining machine learning and genome-wide selection to predict target traits and guide line selection. This model has lower computational requirements, more stable prediction results, and provides more reliable scientific evidence for breeding decision-makers. In practical applications, different combinations of machine learning models and different genome-wide selection models can be selected. The final prediction performance shows that this method is applicable to most traits.

[0036] Secondly, this invention utilizes multi-trait phenotypic data to provide a method for constructing a genome-wide selection model based on a multi-trait phenotypic model, improving the prediction accuracy and selection efficiency of target traits, and is applicable to most traits. This invention can be widely applied in the fields of agricultural animal and plant breeding. In plant breeding fields such as crops, forest trees, and fruits and vegetables, it can be used to predict and select individuals or strains with superior genetic characteristics. Similarly, in animal breeding fields such as poultry, fish, and livestock, it can improve the performance of target traits based on related traits, increase genetic gain, and accelerate the genetic improvement process.

[0037] Third, the technical solution of this invention fills a technological gap in the industry both domestically and internationally: Currently, many genome-wide selection models mainly focus on predicting single traits. Although a few methods can utilize the genetic correlations between multiple traits, they place higher demands on the selection of auxiliary traits. Moreover, the addition of auxiliary traits significantly increases the demand for computing power. This invention, by comprehensively utilizing the associations between multiple traits, constructs a machine learning phenotypic model to capture the linear or nonlinear relationships between traits. It has lower computing power requirements, allowing for the selection of high-performance models based on actual computing power, and eliminates the need for rigorous selection of auxiliary traits. Compared with existing methods, this invention is not only more flexible but also more efficient in processing large-scale data, more practical, and has a wider range of applications.

[0038] The technical solution of this invention solves a long-standing technical problem that people have long desired to solve but have never been able to: the whole genome selection model is constantly developing, but it still faces many challenges, one of which is prediction accuracy. This invention proposes a method for constructing a whole genome selection model based on a multi-trait phenotype model that combines machine learning models with whole genome selection. This method not only effectively improves prediction accuracy and selection efficiency, but is also applicable to the prediction or selection of most traits.

[0039] Fourth, the significant technological advancements brought about by the method for constructing a genome-wide selection model based on multi-trait phenotypic modeling in this invention are reflected in the following aspects:

[0040] 1. Integrating multi-source data to enhance prediction accuracy

[0041] By combining the predicted values ​​of various traits and estimated breeding values ​​obtained from a genomic selection model with a machine learning phenotypic prediction model, this method can make full use of existing genetic information and phenotypic data. This ensemble approach helps improve the accuracy of the final phenotypic prediction, thereby increasing breeding efficiency and success rate.

[0042] 2. Capturing complex relationships between traits

[0043] The constructed multi-trait machine learning phenotypic models can capture both linear and non-linear relationships between traits, which is often difficult to achieve in traditional genomic selection models. By identifying these complex trait associations, breeders can more accurately predict and select lines with desired combinations of traits.

[0044] 3. Refinement of data preprocessing and model optimization

[0045] Detailed data preprocessing steps (such as identifying and cleaning missing and outliers in phenotypic data, random rearrangement, etc.) and parameter tuning of the machine learning model ensure the robustness and generalization ability of the model. Such refinement helps improve the model's adaptability to new data and the accuracy of predicting phenotypic values.

[0046] 4. Improve the efficiency of breeding decision-making

[0047] By utilizing machine learning models for phenotypic prediction, this method can quickly and accurately predict the final phenotypic values ​​of a strain, thus providing breeders with scientific decision support. This efficient predictive capability can significantly shorten the breeding cycle and accelerate the development and promotion of superior varieties.

[0048] 5. Flexibility and scalability

[0049] This construction method is not only applicable to plant breeding, but can also be adapted and applied to other biological breeding fields as needed. Its flexibility and scalability make this method a promising candidate for widespread application.

[0050] The method for constructing a genome-wide selection model based on a multi-trait phenotype model has brought significant technological advancements in improving prediction accuracy, capturing complex associations between traits, optimizing data processing and model training, enhancing breeding decision-making efficiency, and increasing the flexibility and scalability of the method, providing a highly efficient and accurate selection tool for modern breeding. Attached Figure Description

[0051] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0052] Figure 1 This is a schematic diagram illustrating the combination of genome selection and machine learning multi-trait model provided in an embodiment of the present invention;

[0053] Figure 2 This is a schematic diagram illustrating the prediction accuracy of soybean yield-related traits provided in this embodiment of the invention; GBLUP: Genomic best linear unbiased prediction; SI: Selection Index; CropGBM: Crop Genomic Breeding Machine; MT-GBLUP: Multi-Trait GBLUP based on a multi-trait phenotype model combined with the GBLUP method; MT-CropGBM: Multi-Trait CropGBM based on a multi-trait phenotype model combined with the CropGBM method;

[0054] Figure 3This is a schematic diagram illustrating the selection efficiency of yield-related traits using different methods provided in the embodiments of the present invention; PS: phenotypic selection; CropGBM: genomic selection; MT-GBLUP: whole-genome selection based on a multi-trait phenotypic model combined with the GBLUP method; MT-CropGBM: whole-genome selection based on a multi-trait phenotypic model combined with the CropGBM method. Detailed Implementation

[0055] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0056] To address the problems existing in the prior art, the technical solution adopted in this invention is as follows:

[0057] Limitations of single-trait prediction: Traditional genome-wide selection models primarily predict individual traits, neglecting the interactions between multiple traits. This invention, by establishing a multi-trait phenotype model, considers the interactions between multiple traits, including linear and non-linear relationships, thus providing a more comprehensive genetic assessment.

[0058] Limitations in data processing and integration: Existing technologies have shortcomings in processing and integrating phenotypic data and predictions obtained from genomic selection models. The data preprocessing steps in this invention (such as data cleaning and normalization) ensure the consistency and accuracy of the input data, providing a reliable foundation for subsequent model training and prediction.

[0059] Low accuracy in model training and prediction: Traditional GS models have limitations in accuracy and stability. This invention combines machine learning techniques to optimize the model training process and adjusts parameters based on performance metrics of predicted phenotypic values ​​and target trait values, thereby improving the model's accuracy and stability.

[0060] The technical effects and significant technological advancements brought about by this invention in solving the problems of the prior art are as follows:

[0061] Improved accuracy of multi-trait joint prediction: By considering the interaction between traits, this invention improves the prediction accuracy in the breeding process, making strain selection more precise and efficient.

[0062] Optimization of data processing: Refined data preprocessing steps ensure high quality of model input, thereby improving the reliability of the final prediction.

[0063] Enhanced model flexibility and adaptability: By selecting different machine learning models and combinations of different genome-wide selection models, this invention can adapt to different phenotypic data and genetic backgrounds, providing a more flexible and adaptable breeding evaluation tool.

[0064] Significant improvement in breeding efficiency: Using the method of this invention, strain selection can be carried out quickly and accurately in the early stages, thereby effectively shortening the breeding cycle and saving time and costs.

[0065] This invention provides a specific application example of a genome-wide selection model for soybean yield based on a multi-trait model with 75 traits. This example combines a random forest model with GBLUP and CropGBM respectively to construct two models: MT-GBLUP (Multi-Trait GBLUP) and MT-CropGBM (Multi-Trait CropGBM), which predict and select for single-plant grain weight, single-plant pod number, single-plant grain number, and 100-grain weight. The construction steps are as follows:

[0066] Step 1: Build a random forest phenotypic model based on multi-trait data

[0067] (1) Identify and clean up missing and outlier values ​​in data containing multiple strains, each with 75 traits, to improve data quality and the accuracy of machine learning phenotypic model predictions;

[0068] (2) The order of soybean lines in the multi-line phenotypic data with 75 traits was randomly rearranged to eliminate or reduce potential sorting bias.

[0069] (3) The single-plant grain weight, single-plant pod number, single-plant grain weight and 100-grain weight in the phenotypic data are respectively taken as the target traits of the random forest phenotypic model, and the remaining traits are taken as the feature values ​​of the target traits for normalization processing. This improves the processing efficiency of data in the machine learning phenotypic model, ensures that all features have similar scales, thereby speeding up the model training speed and improving the algorithm performance.

[0070] (4) Input the normalized soybean line phenotypic values ​​into the phenotypic model, set the initial values ​​of the machine learning model parameters and hyperparameters according to the actual situation and train the model. Based on the performance measurement of the predicted phenotypic values ​​and the target phenotypic values, further adjust the parameters to achieve the optimal prediction results of the model, and then save the phenotypic prediction model with the best performance measurement.

[0071] Step 2: The random forest multi-trait phenotypic model is combined with GBLUP and CropGBM respectively. This process not only enhances the accuracy of prediction and improves the efficiency of key target trait selection, but also provides a more reliable scientific basis for breeding decisions.

[0072] (1) Using GBLUP and CropGBM, the traits of 75 soybean lines were predicted to obtain estimated breeding values ​​and predicted values.

[0073] (2) Similar to step one, perform data cleaning, processing and normalization on the predicted value and the estimated breeding value. The predicted value or estimated breeding value after completing the above data processing operations is expressed as the characteristic value of the target trait.

[0074] (3) Load the phenotypic prediction model and input the feature values ​​for prediction. The model makes accurate predictions based on the input feature values ​​and outputs the final phenotypic prediction value that is close to the target trait for soybean variety selection.

[0075] 1. A genome-wide selection model for soybean yield based on a multi-trait phenotypic model

[0076] This embodiment constructs a multi-trait model based on the random forest algorithm, utilizing phenotypic data other than yield-related traits. Estimated breeding values ​​or phenotypic values ​​from GBLUP and CropGBM are used as input features to the multi-trait model. Results show that MT-GBLUP (Multi-Trait GBLUP) and MT-CropGBM (Multi-Trait CropGBM) improve the prediction accuracy and selection efficiency of the target trait. This indicates that the combination of multi-trait models and genotype selection holds promise for playing a significant role in breeding, helping to optimize crop yield and improve agricultural production efficiency.

[0077] 1.1 Phenotypic analysis of 75 agronomic traits in soybean

[0078] The test materials were soybean nested association mapping (NAM) populations, with 'Zhongdou 41' as the common maternal parent and the remaining 35 paternal parents derived from major and local varieties across various regions. The NAM populations, comprising 2,455 lines, were planted in Jingzhou City, Hubei Province in the summer of 2020. Seventy-five soybean agronomic traits were collected during the planting period and after harvest, according to soybean germplasm resource data standards (Appendix 1). Data on leaf type, grain type, and grain color were collected using the Wanshen LA_S root analysis system and the SC-G automatic seed testing and dry grain weight analyzer, while protein and oil content data were obtained using the built-in model scanning of the Boton (DA7200) scanner.

[0079] Outliers were removed from the phenotypic data using the 3-sigma principle. The Shapiro-Wilk method was then used for normality testing, and descriptive statistics and Pearson correlation analysis were performed using R packages such as Psych and Hmisc. Descriptive statistical results showed that the variation range of the 75 traits was between 0.02 and 5925.24, with 16 traits having a coefficient of variation greater than 30%, classifying them as highly variable traits (Appendix Table 2). Specifically, the variation ranges for single-plant grain weight, single-plant pod number, single-plant grain number, and 100-grain weight were 4.36–54.13, 14.40–147.80, 19.25–263.40, and 11.50–30.51, respectively, with coefficients of variation of 32.78%, 36.97%, 34.52%, and 15.58%. Skewness, kurtosis, and the Shapiro-Wilk normality test results showed that all 75 traits approximately followed a normal distribution (Appendix Table 2). Correlation analysis showed that there were significant positive correlations (P<0.01) among grain weight per plant, number of grains per plant, and number of pods per plant, with correlations ranging from 0.77 to 0.88. However, the 100-grain weight was negatively correlated with both the number of grains and pods per plant, at -0.31 and -0.25, respectively, while the correlation with grain weight per plant was 0.19. Similarly, grain weight, number of grains, and number of pods per plant were highly correlated with the number of pods on branches, the number of pods on the main stem, branch length, and the number of effective branches, but the 100-grain weight was negatively correlated with these traits (Appendix Table 3).

[0080] 1.2 Predictive accuracy of soybean yield-related traits

[0081] Genotypic data were obtained from the F8 generation microarray data of the NAM population. The raw data were converted to VCF (Variant CallFormat) format, and PLINK software was used to screen for markers with a completeness of ≥80% (--geno 0.20) and a minor allele frequency of <0.05 (--maf0.05), leaving 103,966 SNP markers. Missing genotypes were then filled in using Beagle 5.4 software.

[0082] To evaluate the predictive capabilities of different methods, this study used GBLUP, SI, CropGBM, MT-GBLUP (Multi-Trait GBLUP), and MT-CropGBM (Multi-Trait CropGBM) to predict four yield-related agronomic traits based on phenotypic and genotypic data of yield-related traits. The Pearson correlation coefficient between the phenotypic values ​​in the validation set and the estimated breeding value / predicted value was used as the prediction accuracy. GBLUP employed a five-fold cross-validation method with random seeds, 100 iterations, and calculated the average prediction accuracy across multiple iterations. SI, building upon GBLUP, used four yield traits in rotation as target traits and combined traits other than target traits as auxiliary traits to construct a selection index and calculate the estimated breeding value for the target traits. CropGBM divided the dataset into an 80% training set and a 20% validation set for prediction. MT-GBLUP and MT-CropGBM are based on a random forest model constructed using multi-trait phenotypic data. They use the estimated breeding values ​​or phenotypic prediction values ​​calculated by GBLUP and CropGBM, respectively, to comprehensively predict four yield-related agronomic traits.

[0083] In the GBLUP prediction results, the prediction accuracies for single-plant grain weight, single-plant grain number, single-plant pod number, and 100-grain weight were 0.53, 0.59, 0.66, and 0.82, respectively. Figure 2 Compared to GBLUP, the SI prediction accuracy for grain weight per plant and number of pods per plant improved by 0.06 and 0.04, respectively, while the grain number per plant and 100-grain weight both decreased by 0.02. Similarly, only the prediction accuracy for the number of pods per plant improved by CropGBM by 0.03, while the prediction accuracy for the other three traits decreased to varying degrees (0.02–0.06). In contrast to CropGBM, MT-GBLUP improved the prediction accuracy for grain weight per plant, grain number per plant, and number of pods per plant to 0.72, 0.70, and 0.82, respectively. Furthermore, MT-CropGBM showed further improvements (…). Figure 2 Compared to GBLUP, the improvements in all four traits ranged from 0.03 to 0.25. The reported yield GS prediction accuracy ranged from 0.26 to 0.72, while MT-CropGBM (0.78–0.87) showed a more significant improvement (Table 1 and...). Figure 2The comparison of the results of the four prediction methods shows that CropGBM performed the worst on its own, but its predictive ability was significantly improved when combined with a multi-trait phenotypic model. While SI also utilizes the interrelationships between phenotypes, the improvement was lower, and the accuracy of some traits decreased. MT-GBLUP and MT-CropGBM, on the other hand, performed well in most traits, showing significant improvements. This indicates that the method proposed in this invention, which combines a multi-trait model with GS, is expected to provide strong support for breeding work and has good predictive performance on yield-related traits.

[0084] Table 1. Accuracy of GS prediction for soybean yield in the literature.

[0085]

[0086] Note: PKHS: Regenerated Hilbert Space

[0087] 1.3 Selection efficiency for soybean yield

[0088] To evaluate the selection effectiveness of different methods for high-yielding soybean lines, high-yielding soybeans were screened using criteria of 28g grain weight per plant, 22g 100-grain weight, 70 pods per plant, and 130 grains per plant, and the selection efficiency was calculated using the following formula:

[0089] Phenotypic selection:

[0090] CropGBM, MT-GBLUP, and MT-CropGBM:

[0091] Where y is the selection efficiency, x t x represents the total number of all varieties planted in Jingzhou City, Hubei Province over 20 years. r To determine the number of high-product lines obtained based on phenotypic selection analysis, x m x represents the number of high product lines obtained from CropGBM, MT-GBLUP, or MT-CropGBM predictive values. s This represents the number of lines that match the high-product lines obtained from phenotypic selection analysis, as derived using other prediction methods.

[0092] Based on phenotypic selection, 1072 lines exceeded the set high-yield thresholds for single-plant grain weight, 1115 lines for single-plant grain number, 1007 lines for single-plant pod number, and 993 lines for 100-grain weight, with selection efficiencies of 43.67%, 45.42%, 41.02%, and 40.45%, respectively. Similarly, based on the high-yield threshold analysis, CropGBM predicted 1244, 1230, 1096, and 926 high-yield lines for single-plant grain weight, single-plant grain number, single-plant pod number, and 100-grain weight, respectively. Among these, 864, 875, 783, and 744 were consistent with the high-yield lines selected phenotypically, with selection efficiencies of 69.45%, 71.14%, 71.44%, and 80.35%, respectively. Figure 3 Compared to CropGBM, the selection efficiency of the four yield traits in MT-GBLUP showed a decreasing trend, with the reduction ranging from 0.94% to 6.57%. In contrast to MT-CBLUP, MT-CropGBM exhibited the best selection efficiency, with selection efficiencies for grain weight per plant, number of grains per plant, number of pods per plant, and 100-grain weight reaching 72.92%, 74.39%, 76.43%, and 81.11%, respectively. Figure 3 The multi-trait random forest model combined with CropGBM proposed in this invention (MT-CropGBM) can improve resource utilization efficiency and achieve higher production benefits by selecting high-product lines.

[0093] Appendix 1. Classification, abbreviations and units of 75 agronomic traits

[0094]

[0095]

[0096] Continued from Appendix 1

[0097]

[0098]

[0099]

[0100]

[0101]

[0102] Appendix 2. Descriptive statistical analysis of 75 agronomic traits

[0103] Continued from Appendix 2

[0104]

[0105]

[0106]

[0107] Note: SWPP: Grain weight per plant; PNPP: Number of pods per plant; SNPP: Number of grains per plant; 100-SW: 100-grain weight; FTM: Flowering to maturity period; MT: Maturity period; FT: Flowering period; IN: Number of nodes on the main stem; SD: Stem diameter; PH: Plant height; POH: Pod formation habit; LG: Lodging tendency; GH: Growth habit; BL: Branch length; BIN: Number of nodes on branches; BD: Branch diameter; EBN: Number of effective branches; BA: Branch angle; MPN: Number of moldy pods; MPNR: Mold pod ratio; FSPNPP: Number of four pods per plant; BPN: Number of pods on branches; FS R: Percentage of four pods per plant; PW: Pod width; PL: Pod length; FPH: Bottom pod height; SPN: Number of pods on the main stem; SP: Grain circumference; SL: Grain length; B: Blue; R: Red; G: Green; SW: Grain width; SA: Grain area; SLWR: Grain length-to-width ratio; TMLL: Upper middle leaf length; TMLP: Upper middle leaf circumference; TSLP: Upper lateral leaf circumference; TSLL: Upper lateral leaf length; TMLA: Upper middle leaf area; TSLW: Upper lateral leaf width; TMLW: Upper middle leaf width; TSLA: Upper lateral leaf area; TMLLWR: The leaf length-to-width ratio of the upper middle leaves; TSLLWR: leaf length-to-width ratio of the upper lateral leaves; MMLP: leaf perimeter of the middle middle leaves; MMLA: leaf area of ​​the middle middle leaves; MSLL: leaf length of the middle lateral leaves; MSLP: leaf perimeter of the middle lateral leaves; MMLL: leaf length of the middle middle leaves; MMLLWR: leaf length-to-width ratio of the middle middle leaves; MSLW: leaf width of the middle lateral leaves; MSLA: leaf area of ​​the middle lateral leaves; MMLW: leaf width of the middle middle leaves; MSLLWR: leaf length-to-width ratio of the middle lateral leaves; BSLP: leaf perimeter of the lower lateral leaves; BMLL: leaf length of the lower middle leaves; BMLP: leaf length of the lower middle leaves. Leaf circumference; BSLL: Lower lateral leaf length; BMLW: Lower middle leaf width; BMLA: Lower middle leaf area; BSLW: Lower lateral leaf width; BSLA: Lower lateral leaf area; BSLLWR: Lower lateral leaf length-to-width ratio; BMLLWR: Lower middle leaf length-to-width ratio; PLM: Middle petiole length; PLB: Lower petiole length; PLT: Upper petiole length; PAT: Upper petiole angle; PAM: Middle petiole angle; PAB: Lower petiole angle; PC: Protein content; OC: Oil content; POC: Total protein and fat; FC: Color.

[0108] Appendix 3: Correlation between single-plant grain weight, single-plant grain number, single-plant pod number, and 100-grain weight with other agronomic traits

[0109]

[0110]

[0111]

[0112] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions, and improvements made by those skilled in the art within the scope of the technology disclosed in the present invention, and within the spirit and principles of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for constructing a genome-wide selection model based on a multi-trait phenotype model, characterized in that, The method for constructing the genome-wide selection model based on a multi-trait phenotype model includes the following steps: Step 1: Establish a machine learning multi-trait phenotypic model; Step two: Combine machine learning phenotypic models with genomic selection models; Building a machine learning multi-trait phenotypic model specifically includes the following steps: (1) Identify and clean up missing and outliers in the phenotypic data; (2) Randomly rearrange the order of plants in the phenotypic data; (3) Separate the target trait from the phenotypic data and normalize the remaining traits as feature values ​​of the target trait; (4) The normalized phenotypic values ​​of each plant are used as input. The machine learning phenotypic model is trained according to the actual model parameters and hyperparameters of the machine learning model. Based on the performance measurement of the predicted phenotypic values ​​and the target phenotypic values, the parameters are further adjusted. Then the phenotypic prediction model with the best performance measurement is saved. The phenotypic prediction model is a random forest phenotypic model; Combining machine learning phenotypic models with genomic selection models specifically includes the following steps: 1) Select a whole-genome selection model to predict all traits separately. The whole-genome selection model is either GBLUP or CropGBM. 2) Perform data cleaning, processing, and normalization on the predicted and estimated breeding values; 3) The predicted or estimated breeding values ​​after completing the above data processing operations are represented as characteristic values ​​of the target trait; 4) Load the phenotypic prediction model and input the feature values ​​for prediction. The model makes accurate predictions based on the input feature values ​​and outputs the final phenotypic prediction value that is close to the target trait.

2. A system for constructing a genome-wide selection model based on a multi-trait phenotype model using the construction method of claim 1, characterized in that, include: Data preprocessing unit: responsible for receiving and processing phenotypic data, including identifying and cleaning missing and outlier values ​​in the data, and randomly rearranging the order of plants; Data normalization unit: used to normalize non-target traits in phenotypic data as feature values ​​and separate out target traits; Machine learning training unit: Based on the normalized trait feature values, machine learning phenotypic models are trained by setting appropriate model parameters and hyperparameters; Genomic selection model unit: Genetic evaluation of individuals is performed using whole-genome markers to obtain estimated genomic breeding values; Model integration unit: combines machine learning phenotypic models and genomic selection models to perform data cleaning, processing and normalization on predicted values ​​and estimated breeding values; The machine learning phenotypic model is a random forest phenotypic model, and the whole-genome selection model is GBLUP or CropGBM.

3. The genome-wide selection model based on a multi-phenotypic model constructed by the method for constructing a genome-wide selection model based on a multi-phenotypic model as described in claim 1.

4. The application of the genome-wide selection model based on a multi-trait phenotype model as described in claim 3 in the field of plant breeding.

Citation Information

Patent Citations

  • Whole-genome selective breeding method and apparatus

    CN111524545A

  • Rice grain cadmium accumulation character prediction device and early warning system based on whole genome selection research

    CN115579057A