Soybean yield phenotype quantitative prediction method fusing machine learning and gene effect
Through the quantitative prediction method of soybean yield phenotypes that integrate machine learning and genetic effects, the problem that traditional models cannot accurately predict soybean yield phenotypes in multiple environments is solved, and the accurate prediction and genetic explanatory improvement of soybean yield phenotypes are achieved.
Patent Information
- Application Number
- CN202510580072.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-07
AI Technical Summary
Traditional soybean growth simulation models cannot accurately predict soybean yield phenotypes with different genetic backgrounds in multiple environments, and the genetic interpretability of genetic parameters is insufficient.
The quantitative prediction method of soybean yield phenotypes using a fusion machine learning algorithm and genetic effects is adopted. Through genome-wide association analysis and multiple machine learning algorithms, the phenotype sensitive parameters of soybean yield are obtained to achieve accurate prediction of yield phenotypes of different soybean varieties under different environments.
The genetic interpretability of the soybean yield phenotype simulation model is improved, accurate prediction of soybean yield phenotypes of different environments and genotypes is achieved, and the difficulty of parameter correction and labor cost is reduced.
Smart Images

Figure CN120108512A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of agricultural information technology and relates to a soybean yield phenotype quantitative prediction method integrating machine learning and gene effect, which is used for quantitatively predicting soybean yield phenotypes with different gene data under different planting environments. Background Art
[0002] Soybean Glycine max (L.) Merr.] is rich in unsaturated fatty acids, high-quality protein and a variety of trace elements, has high nutritional value and economic value, is one of the important oil and feed grains, and is directly related to the supply security of important agricultural products such as edible oil, meat, eggs and milk.
[0003] The soybean yield phenotype is affected by the interaction between genotype and environment (G×E), and is the result of environmental conditions and the interaction between genes and environment (G×E). Genetic information determines the variety characteristics of soybeans. Researchers can accurately identify the QTL, SNP or gene itself that causes the variation of target traits at the genome level, providing a strong guarantee for screening the genetic information of soybean yield phenotype. At present, the QTL related to soybean roots, stems, yield, nutrients, and biotic and abiotic stresses have been located using whole genome association analysis methods, which has promoted the mining and identification of soybean yield-related genes. However, traditional genomic prediction models do not take into account the influence of environmental effects and gene-environment interaction effects, resulting in the inability to accurately predict soybean yield phenotypes under multiple environments and different genetic backgrounds.
[0004] The soybean growth simulation model integrates the interaction effects of genotype (G) × environmental conditions (E) × management technology (M), which can quantitatively describe and predict the soybean growth and development process and its dynamic relationship with the environment and technology. Therefore, the soybean growth simulation model can be used to predict the yield phenotype of soybean under different environments, providing agricultural producers with better planting and management decisions. However, it is difficult to obtain genetic parameters in traditional soybean growth simulation models, and the relationship and mechanism between soybean yield genetic parameters and genome are still unclear. The genetic interpretability of genetic parameters needs further study.
[0005] In summary, there is an urgent need for a quantitative prediction method for soybean yield phenotype that integrates machine learning and genetic effects to improve the genetic interpretability of soybean yield phenotype simulation models, accelerate the phenotypic identification of soybean lines, and provide a theoretical basis and effective tools for the screening and evaluation of soybean genotype varieties in different ecological zones. Summary of the invention
[0006] The technical problem to be solved by the present invention is to provide a quantitative prediction method for soybean yield phenotype that integrates machine learning algorithms and genetic effects. According to the genetic data and sensitive parameters of soybean model yield phenotype obtained by analysis, combined with whole genome association analysis and machine learning algorithms, accurate prediction of yield phenotypes of different soybean varieties under different environments can be achieved.
[0007] The purpose of the present invention is achieved by adopting the following technical solutions: A soybean yield phenotype quantitative prediction method integrating machine learning and genetic effects, comprising the following steps: Step 1: Collect soybean-related data, including: variety information, genetic data, meteorological data and soil data of the planting site, field management measures and yield formation data; Step 2: Divide the acquired soybean-related data into environmental training sets and environmental verification sets according to the planting location and planting year; Step 3: Based on the soil data, meteorological data and soybean growth simulation model of the planting site, the sensitivity of the genetic parameters of the yield phenotype of the soybean growth simulation model is analyzed using the genetic parameter sensitivity analysis method; Step 4: According to the sensitivity information of the yield phenotypic genetic parameters of the soybean growth simulation model, the yield phenotypic genetic parameters of the soybean growth simulation model are divided into three combinations: all sensitivity parameter combinations, all extreme sensitivity parameter combinations and growth-related sensitivity parameter combinations; Step 5: Based on the collected soybean-related data, the divided environmental training set, the environmental verification set and the soybean growth simulation model, the optimization algorithm is used to obtain the numerical information of the sensitive genetic parameters of different soybean varieties, that is, the optimal yield genetic parameter combination; Step 6: The obtained soybean optimal yield genetic parameter combination and gene data are used for association analysis using multiple genome-wide association analysis GWAS methods with different matrices to obtain a significant SNP marker data set of sensitive genetic parameters; Step 7: Divide the soybean optimal yield genetic parameter combination and sensitive genetic parameter significant SNP marker data set into a variety training set and a variety verification set according to a fixed ratio of the number of varieties; Step 8: Using the machine learning algorithm, cross-validate the variety training set, and use the variety validation set as an external validation to obtain a soybean yield phenotypic genetic parameter prediction model that integrates gene effects and machine learning algorithms; based on the soybean yield phenotypic genetic parameter prediction model that integrates gene effects and machine learning algorithms, obtain soybean yield phenotypic genetic parameters with gene effects; Step 9: Input the obtained soybean yield phenotype genetic parameters with gene effects into the soybean growth simulation model to achieve accurate prediction of soybean yield phenotype.
[0008] Preferably, step 1 comprises the following steps: (1) Data on soybean flowering, seeding, poding, initial maturity, maturity, number of pods per plant, number of seeds per plant, 100-seed weight, and plot yield were collected based on field trials; (2) Record the management measures of the soybean experimental field, including sowing, fertilization, irrigation, weeding, spraying, thinning, harvesting, and seed testing; (3) Use gene sequencing technology to sequence the whole genome of soybean varieties and obtain genetic data of soybean varieties; (4) Collect and organize soil and meteorological data of the test area.
[0009] Preferably, step 2 comprises the following steps: The collected soybean growth, yield and meteorological data were analyzed, and the soybeans were divided into environmental training set and environmental verification set according to the planting location and planting year.
[0010] Preferably, step 3 comprises the following steps: The collected meteorological data and soil data were input into the soybean growth simulation model, and the Latin hypercube single factor LH-OAT algorithm and the extended Fourier amplitude sensitivity test EFAST were used to conduct sensitivity analysis on the yield phenotypic genetic parameters of the soybean growth simulation model. The yield phenotypic genetic parameters of the soybean growth simulation model were sorted from high to low according to the sensitivity analysis values, and divided into all sensitivity parameter combinations, all extremely sensitive parameter combinations and growth-related sensitivity parameter combinations.
[0011] Preferably, step 5 comprises the following steps: (1) Based on the collected soybean-related data and soybean growth simulation model, the meteorological data, soil data, soybean growth and yield phenotypic data of the environmental training set are input into the soybean growth simulation model to provide the necessary conditions for the correction of the yield phenotypic genetic parameters of the soybean growth simulation model; (2) The calling code of the soybean growth simulation model is combined with the differential evolution algorithm DE and the Markov chain Monte Carlo algorithm MCMC to calculate the simulated yield results of the yield phenotypic genetic parameters of the soybean growth simulation model, and then the yield phenotypic genetic parameters of the soybean growth simulation model are corrected to obtain the optimal yield genetic parameter combination of the soybean growth simulation model for each soybean variety; (3) The environmental validation set was input into the soybean growth simulation model to test the prediction of the optimal yield genetic parameter combination.
[0012] Preferably, step 6 comprises the following steps: (1) The obtained soybean yield all sensitive genetic parameter combinations, all extremely sensitive genetic parameter combinations, growth-related sensitive genetic parameter combinations and gene data were respectively used for association analysis using 7 genome-wide association analysis GWAS methods with different matrices (P matrix + K matrix; Q matrix + K matrix); (2) The obtained significant SNP markers of sensitive genetic parameters were combined to obtain a dataset of significant SNP markers of all soybean yield sensitive genetic parameters, a dataset of significant SNP markers of all extremely sensitive genetic parameters, and a dataset of significant SNP markers of growth-related sensitive genetic parameters.
[0013] Preferably, step 7 comprises the following steps: The soybean all yield sensitivity genetic parameter significant SNP marker data set, all extreme sensitivity genetic parameter significant SNP marker data set, growth-related sensitivity genetic parameter significant SNP marker data set and soybean growth and yield data were randomly divided into variety training set and variety verification set in a ratio of 8:2 according to the number of varieties.
[0014] Preferably, step 8 comprises the following steps: (1) The obtained significant SNP marker data sets of all sensitive genetic parameter combinations, all extremely sensitive genetic parameter combinations, and growth-related sensitive genetic parameter combinations and all sensitive genetic parameter combinations, all extremely sensitive genetic parameter combinations, and growth-related sensitive genetic parameter combinations are used as input one-to-one, and the machine learning algorithm is used for algorithm learning and training; (2) A ten-fold cross-validation method was used for the variety training set. The purpose of this method is to test the stability of the constructed algorithm. (3) The variety validation set is used as an external validation set to test the constructed algorithm in order to verify the generalization of the constructed algorithm; (4) Through machine learning algorithm calculations, the yield prediction results of the soybean yield phenotype model that integrates machine learning and genetic effects for all sensitive genetic parameter combinations, all extremely sensitive genetic parameter combinations, and growth-related sensitive genetic parameter combinations were obtained. The yield prediction results obtained by different machine learning algorithms and sensitivity parameter combinations were compared, and the model with the highest yield prediction accuracy was recorded. Finally, the optimal soybean yield phenotype genetic parameter prediction model with genetic effects was obtained.
[0015] Preferably, the machine learning algorithm is a BayesC algorithm, a LASSO algorithm, a rrBLUP algorithm, a random forest algorithm, a support vector machine regression algorithm or a gradient boosting decision tree XGBoost algorithm.
[0016] The beneficial effects brought by the present invention are: (1) The present invention constructs a genetic parameter prediction model of gene effect through multiple machine learning algorithms, achieves the purpose of predicting soybean yield genetic parameters through gene information, effectively reduces the parameter correction problem in the soybean yield phenotypic model used alone, saves test time and labor costs; and increases the genetic interpretability of the soybean yield phenotypic model; (2) The present invention divides soybean yield phenotypic genetic parameters into different groups according to their sensitivity, and combines a variety of whole-genome association analysis methods and two different matrix (covariate) combinations, effectively increasing the number of significant SNP markers obtained, and for the first time provides a variety of processing schemes for the construction of a gene effect-soybean growth simulation model, providing technical support for the development of crop models; (3) The present invention solves the problem of integrating genetic effects and soybean growth simulation models. It can use genetic information and soybean yield phenotype models to accurately predict soybean yield phenotypes under different climatic conditions in different regions, providing digital support for quantitative evaluation of yield phenotypes and formulation of breeding strategies. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a flow chart of the method in the present invention.
[0018] Figure 2 This is a climate data analysis chart of the soybean growth environment in the present invention.
[0019] Figure 3 This is a sensitivity analysis diagram of genetic parameters of the soybean growth simulation model in the present invention.
[0020] Figure 4 This is a diagram of the soybean yield simulation results (environmental training set) using the soybean growth simulation model in the present invention.
[0021] Figure 5 This is a diagram of the soybean yield prediction results (environmental validation set) using the soybean growth simulation model in the present invention.
[0022] Figure 6 This is a diagram of the genetic data analysis of soybean varieties in the present invention.
[0023] Figure 7 This is the Q+K matrix analysis diagram of the whole genome association analysis in the present invention.
[0024] Figure 8 This is the P+K matrix analysis diagram of the whole genome association analysis in the present invention.
[0025] Fig. 9 This is a diagram of the prediction results of the soybean yield genetic parameter variety training set and validation set that integrates multiple machine learning and gene effects in the present invention.
[0026] Fig.10This is a diagram of the prediction results of the soybean yield environment training set that integrates multiple machine learning algorithms and genetic effects in the present invention.
[0027] Fig.11 This is a diagram of the prediction results of the soybean yield environment validation set that integrates multiple machine learning algorithms and genetic effects in the present invention. DETAILED DESCRIPTION
[0028] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the protection scope of the present invention.
[0029] Taking the main soybean varieties in Northeast China and Huanghuaihai region as the research objects, the method of the present invention is specifically described, comprising the following steps: like Figure 1 As shown, a soybean yield phenotype quantitative prediction method integrating machine learning and genetic effects includes the following steps: Step 1: As shown in Table 1, 75 main varieties in Northeast China and Huanghuaihai region were obtained, and the experiment was planted in 12 ecological points in Northeast China and Huanghuaihai region from 2009 to 2011. The soybean phenotypic data included growth period (sowing period, initial flowering period, maturity period) and yield phenotypic data (yield, number of pods per plant, number of grains per plant, grain weight per plant, 100-grain weight); soybean gene data included 3.38 million SNP markers obtained by second-generation sequencing (the reference gene line was Williams 82).
[0030] Table 1 Planting locations of main soybean varieties in Northeast China and Huanghuaihai Place latitude longitude Heihe 50.14 N 127.29°E Zhalantun 48.00°N 122.47°E Jiusan Research Institute 47.20 N 123.57 E Helen 47.28°N 126.57°E Jiamusi 46.47°N 130.22 E Harbin 45.44°N 126.36°E Changchun 43.54°N 125.19 E Jilin 43.52°N 126.33 E Tieling 42.28°N 123.51°E Changping 40.22 N 116.20 E Jinan 36.67 N 117.00°E Xuchang 34.02°N 113.82 E Step 2: Divide the obtained soybean-related data into environmental training sets and environmental verification sets according to planting areas and planting years. Figure 2 As shown, the data obtained in Northeast China and Huanghuaihai region are divided into four regions: Northeast North, Northeast Central, Northeast South and Huanghuaihai region; the daily maximum temperature, daily minimum temperature, daily radiation value and daily precipitation in the obtained meteorological data are sorted and analyzed according to the four regions, and the representative planting environment combination of each region is extracted as the validation set, and the other planting environment combinations are used as the training set; Step 3: Based on the soil data, meteorological data and soybean growth simulation model of the planting site, the sensitivity of the genetic parameters of the yield phenotype of the soybean growth simulation model was analyzed using the genetic parameter sensitivity analysis method. Figure 3As shown, the meteorological data, soil data and soybean field planting management measures of different soybean planting locations were input into the soybean yield phenotypic model DSSAT-CROPGRO-Soybean model; the soybean genetic parameters in the DSSAT-CROPGRO-Soybean model were subjected to sensitivity analysis using the LH-OAT sensitivity analysis method and the EFAST sensitivity analysis method; (1) Among them, the LH-OAT sensitivity analysis method is the simplest and most direct method for global sensitivity analysis of model parameters. First, according to the LH sampling idea, the entire parameter space is divided into K layers. Then, a random sample is taken from each layer to generate an LH sampling parameter group (a parameter set containing m parameters). Then, according to the OAT idea, m small parameter changes are made to each LH sampling parameter group (only one parameter is changed each time), and the change of the objective function before and after each small change is calculated. The LH-OAT sensitivity analysis method is completed by the written code.
[0031] (2) The extended Fourier amplitude sensitivity test (EFAST) method decomposes the variance of model parameters to model output variables and divides the sensitivity of parameters into two categories: ① the sensitivity of a single parameter to the model output variable, measured by the first-order sensitivity index; ② the sensitivity of the interaction between parameters to the model output variable, measured by the total sensitivity index. EFAST sensitivity analysis was performed using SimLab sensitivity analysis software.
[0032] Step 4: According to the values of sensitivity analysis, the genetic parameters of soybean growth simulation model are arranged from high to low, and divided into three groups: all sensitivity parameters, all extreme sensitivity parameters and growth-related sensitivity parameters; They are (1) all sensitive parameters of soybean yield: CSDL, PPSEN, EM-FL, FL-SH, FL-SD, SD-PM, LFMAX, SLAVR, SIZLF, WTPSD, SFDUR, PODUR, THRSH; all extremely sensitive parameters of soybean yield: CSDL, PPSEN, EM-FL, FL-SH, FL-SD; sensitive parameters related to soybean yield growth: LFMAX, SLAVR, SIZLF, WTPSD, SFDUR, PODUR, THRSH; (2) All sensitivity parameters of soybean pod number: CSDL, PPSEN, EM-FL, FL-SH, FL-SD, SD-PM, LFMAX, SLAVR, SIZLF, WTPSD, SFDUR, SDPDV, PODUR, THRSH; All extreme sensitivity parameters of soybean pod number: CSDL, PPSEN, EM-FL, FL-SH, FL-SD, WTPSD, PODUR; Soybean pod number growth-related sensitivity parameters: LFMAX, SLAVR, SIZLF, WTPSD, SFDUR, SDPDV, PODUR, THRSH; (3) All sensitive parameters of soybean grain number: CSDL, PPSEN, EM-FL, FL-SH, FL-SD, SD-PM, LFMAX, SLAVR, SIZLF, WTPSD, SFDUR, PODUR; All extremely sensitive parameters of soybean grain number: CSDL, PPSEN, EM-FL, FL-SH, FL-SD, WTPSD; Sensitivity parameters related to soybean grain number growth: LFMAX, SLAVR, SIZLF, WTPSD, SFDUR, PODUR.
[0033] Step 5: Based on the collected soybean-related data and soybean growth simulation model, the optimization algorithm is used to obtain the sensitive genetic parameter information of different soybean varieties. Figure 4 As shown, the differential evolution algorithm (DE) was used to correct soybean yield genetic parameters and obtain the optimal genetic parameter combination in the training set; Figure 5 As shown, the environmental data selected by the validation set are input into the model for validation to test the generalization ability of the optimal genetic parameter combination.
[0034] Among them, the differential evolution algorithm (DE) searches and optimizes through differential mutation operations and probabilistic crossover operations to obtain the optimal parameter combination. First, given the parameter correction range, the algorithm randomly extracts two parameter arrays, adds the weighted difference vector to the array to generate a new parameter vector; mixes the parameters of the mutation vector with the parameters of another predetermined target vector according to certain rules to generate sub-individuals; then selects the better one from the target individual (simulated value) and the test individual (measured value) as the next generation, gradually narrows the range, and finally outputs the optimal parameter combination.
[0035] Step 6: The acquired soybean variety genetic parameters and gene data are used for association analysis using multiple genome-wide association analysis (GWAS) methods with different matrices to obtain significant SNP markers of sensitive genetic parameters; like Figure 6 As shown, firstly, the genetic data was quality controlled, then the linkage disequilibrium of the soybean population was calculated, and then the population structure and principal component analysis of the soybean population was further performed; like Figure 7 and Figure 8 As shown, seven GWAS analysis methods, including IIIVmrMLM, mrMLM, FASTmrMLM, FASTmrEMMA, pLARmEB, pKWmEB and ISIS EM-BLASSO, were used. P matrix + K matrix and Q matrix + K matrix were used as covariates to analyze the quality-controlled genetic data and screen out significant SNP markers related to genetic parameters.
[0036] Step 7: Divide the soybeans into variety training set and variety verification set using a random sampling method according to the ratio of 8:2 of the number of varieties; Step 8: Combine the obtained significant SNP markers with the machine learning algorithm, and use a combination of cross-validation and external validation to obtain the soybean yield phenotypic genetic parameters of the fusion gene effect; As shown in Table 2 , the significant SNP markers were converted into digital formats of 0, 1, and 2 for use in computer algorithm learning; Table 2 Gene information conversion code correspondence table Gene information Convert code AA 0 Heterozygous 1 aa 2 like Fig. 9 As shown, the three sets of significant SNP marker data sets and soybean yield phenotype measured data were used as input, and the machine learning algorithm codes (BayesC algorithm, LASSO algorithm, rrBLUP algorithm, random forest algorithm (RF), support vector machine regression algorithm (SVR) and gradient boosting decision tree (XGBoost) algorithm) were used for algorithm learning and training; (1) First, the algorithm is learned using the variety training set, the optimal hyperparameter combination of the machine learning algorithm is obtained through the Bayesian algorithm, and the stability of the constructed algorithm is tested using the ten-fold cross-validation method; (2) After the machine learning algorithm completes data learning, the genetic data of the variety training set is used to predict the genetic parameters of gene effects; (3) Next, the genetic information of the variety validation set is used to predict the genetic parameters of gene effects. The variety validation set is used as an external validation set to test the generalization of the constructed algorithm, thereby obtaining the optimal combination value of the genetic parameters of soybean yield phenotype that integrates the gene effects.
[0037] (3.1) Among them, the ridge regression best linear unbiased prediction (rrBLUP) method is a multiple genomic information prediction method based on the Bayesian framework. It simulates variety parameters by estimating the intercept and the effect value of SNP markers, and further constructs a linear regression model by learning the input features to complete parameter prediction; (3.2) The operation principle of BayesC is based on the Bayesian statistical method, which models the gene effect as sampling from a prior distribution with zero mean, effectively reducing noise and improving prediction accuracy. In the learning of genetic data, BayesC can run stably and produce high-quality prediction results. (3.3) LASSO (Least Absolute Shrinkage and Selection Operator) is a linear regression method that aims to improve the prediction accuracy of the model by introducing an L1 regularization term while performing feature selection. LASSO uses the L1 penalty of the regression coefficient to control the complexity of the model, thereby reducing overfitting and further improving the prediction effect of gene effect parameters.
[0038] (3.4) The Random Forest (RF) algorithm is a machine learning algorithm based on ensemble learning. It first generates a series of decision trees. Each decision tree is trained on an independent data set. Sample perturbation is added by sampling with replacement, and an attribute perturbation is introduced. As the number of learning samples increases, the random forest will gradually converge, and finally the optimal genetic parameters of gene effects are obtained through analysis. (3.5) Support Vector Machines (SVM) is a binary classification model that aims to find a hyperplane to segment samples. The principle of segmentation is to maximize the interval. Support vector machines are nonlinear methods. When the sample size is relatively small, they can grasp the nonlinear relationship between data and features, avoid the problems of neural network structure selection and local optimality, and obtain the best prediction results. (3.6) XGBoost (eXtreme Gradient Boosting) is a boosting tree model whose basic component is the decision tree. It can perform a second-order Taylor expansion on the loss function, increasing accuracy while being able to customize the loss function. XGBoost adds a regularization term to the objective function to penalize complex models, reduce model variance, and reduce overfitting, thereby obtaining the best prediction results.
[0039] Step 9: Input the obtained fusion gene effect and soybean yield phenotype genetic parameters of the machine learning algorithm into the soybean growth simulation model respectively to achieve accurate prediction of soybean yield phenotype.
[0040] like Fig.10 and Fig.11As shown, the six genetic parameters of soybean yield phenotype gene effects predicted by machine learning obtained in step 8 and the meteorological data, soil data and management measures collected in step 1 were input into the DSSAT-CROPGRO-Soybean model, and the yield phenotype prediction results of different soybean varieties in different planting locations were obtained through model calculation.
[0041] Six machine learning algorithms were selected from the groups of gene effect-yield phenotype all sensitive genetic parameters, gene effect-yield phenotype all extremely sensitive genetic parameters and gene effect-yield phenotype growth-related sensitivity genetic parameters. The combination with the lowest RMSE, nRMSE and MAF among the 18 methods was the optimal combination of the constructed fusion model. In this study, the gene effect-yield phenotype growth-related sensitivity genetic parameter method fused with the SVR algorithm had the best simulation prediction effect.
[0042] The above are only preferred embodiments of the present invention. It should be pointed out that, for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A quantitative prediction method for soybean yield phenotype integrating machine learning and genetic effects, characterized in that: The following steps are involved: Step 1: Collect soybean-related data, including: variety information, genetic data, meteorological data and soil data of the planting site, field management measures and yield formation data; Step 2: Divide the acquired soybean-related data into environmental training sets and environmental verification sets according to the planting location and planting year; Step 3: Based on the soil data, meteorological data and soybean growth simulation model of the planting site, the sensitivity of the genetic parameters of the yield phenotype of the soybean growth simulation model is analyzed using the genetic parameter sensitivity analysis method; Step 4: According to the sensitivity information of the yield phenotypic genetic parameters of the soybean growth simulation model, the yield phenotypic genetic parameters of the soybean growth simulation model are divided into three combinations: all sensitivity parameter combinations, all extreme sensitivity parameter combinations and growth-related sensitivity parameter combinations; Step 5: Based on the collected soybean-related data, the divided environmental training set, the environmental verification set and the soybean growth simulation model, the optimization algorithm is used to obtain the numerical information of the sensitive genetic parameters of different soybean varieties, that is, the optimal yield genetic parameter combination; Step 6: The obtained soybean optimal yield genetic parameter combination and gene data are used for association analysis using multiple genome-wide association analysis GWAS methods with different matrices to obtain a significant SNP marker data set of sensitive genetic parameters; Step 7: Divide the soybean optimal yield genetic parameter combination and sensitive genetic parameter significant SNP marker data set into a variety training set and a variety verification set according to a fixed ratio of the number of varieties; Step 8: Using the machine learning algorithm, cross-validate the variety training set, and use the variety validation set as an external validation to obtain a soybean yield phenotypic genetic parameter prediction model that integrates gene effects and machine learning algorithms; based on the soybean yield phenotypic genetic parameter prediction model that integrates gene effects and machine learning algorithms, obtain soybean yield phenotypic genetic parameters with gene effects; Step 9: Input the obtained soybean yield phenotype genetic parameters with gene effects into the soybean growth simulation model to achieve accurate prediction of soybean yield phenotype.
2. The method for quantitatively predicting soybean yield phenotype by integrating machine learning and gene effects according to claim 1, characterized in that: The step 1 comprises the following steps: (1) Data on soybean flowering, seeding, poding, initial maturity, maturity, number of pods per plant, number of seeds per plant, 100-seed weight, and plot yield were collected based on field trials; (2) Record the management measures of the soybean experimental field, including sowing, fertilization, irrigation, weeding, spraying, thinning, harvesting, and seed testing; (3) Use gene sequencing technology to sequence the whole genome of soybean varieties and obtain genetic data of soybean varieties; (4) Collect and organize soil and meteorological data of the test area.
3. The method for quantitatively predicting soybean yield phenotype by integrating machine learning and gene effects according to claim 1, characterized in that: The step 2 comprises the following steps: The collected soybean growth, yield and meteorological data were analyzed, and the soybeans were divided into environmental training set and environmental verification set according to the planting location and planting year.
4. The method for quantitatively predicting soybean yield phenotype by integrating machine learning and gene effects according to claim 1, characterized in that: The step 3 comprises the following steps: The collected meteorological data and soil data were input into the soybean growth simulation model, and the Latin hypercube single factor LH-OAT algorithm and the extended Fourier amplitude sensitivity test EFAST were used to conduct sensitivity analysis on the yield phenotypic genetic parameters of the soybean growth simulation model. The yield phenotypic genetic parameters of the soybean growth simulation model were sorted from high to low according to the sensitivity analysis values, and divided into all sensitivity parameter combinations, all extremely sensitive parameter combinations and growth-related sensitivity parameter combinations.
5. The method for quantitatively predicting soybean yield phenotype by integrating machine learning and gene effects according to claim 1, characterized in that: The step 5 comprises the following steps: (1) Based on the collected soybean-related data and soybean growth simulation model, the meteorological data, soil data, soybean growth and yield phenotypic data of the environmental training set are input into the soybean growth simulation model to provide the necessary conditions for the correction of the yield phenotypic genetic parameters of the soybean growth simulation model; (2) The calling code of the soybean growth simulation model is combined with the differential evolution algorithm DE and the Markov chain Monte Carlo algorithm MCMC to calculate the simulated yield results of the yield phenotypic genetic parameters of the soybean growth simulation model, and then the yield phenotypic genetic parameters of the soybean growth simulation model are corrected to obtain the optimal yield genetic parameter combination of the soybean growth simulation model for each soybean variety; (3) The environmental validation set was input into the soybean growth simulation model to test the prediction of the optimal yield genetic parameter combination.
6. The method for quantitatively predicting soybean yield phenotype by integrating machine learning and gene effects according to claim 1, characterized in that: The step 6 comprises the following steps: (1) The obtained soybean yield all sensitive genetic parameter combinations, all extremely sensitive genetic parameter combinations, growth-related sensitive genetic parameter combinations and gene data were respectively used for association analysis using 7 genome-wide association analysis GWAS methods with different matrices; (2) The obtained significant SNP markers of sensitive genetic parameters were combined to obtain a dataset of significant SNP markers of all soybean yield sensitive genetic parameters, a dataset of significant SNP markers of all extremely sensitive genetic parameters, and a dataset of significant SNP markers of growth-related sensitive genetic parameters.
7. The method for quantitatively predicting soybean yield phenotype by integrating machine learning and gene effects according to claim 1, characterized in that: The step 7 comprises the following steps: The soybean all yield sensitivity genetic parameter significant SNP marker data set, all extreme sensitivity genetic parameter significant SNP marker data set, growth-related sensitivity genetic parameter significant SNP marker data set and soybean growth and yield data were randomly divided into variety training set and variety verification set in a ratio of 8:2 according to the number of varieties.
8. The method for quantitatively predicting soybean yield phenotype by integrating machine learning and gene effects according to claim 1, characterized in that: The step 8 comprises the following steps: (1) The obtained significant SNP marker data sets of all sensitive genetic parameter combinations, all extremely sensitive genetic parameter combinations, and growth-related sensitive genetic parameter combinations and all sensitive genetic parameter combinations, all extremely sensitive genetic parameter combinations, and growth-related sensitive genetic parameter combinations are used as input one-to-one, and the machine learning algorithm is used for algorithm learning and training; (2) A ten-fold cross validation method was used for the variety training set; (3) Use the variety validation set as an external validation set to test the constructed algorithm; (4) Through machine learning algorithm calculations, the yield prediction results of the soybean yield phenotype model that integrates machine learning and genetic effects for all sensitive genetic parameter combinations, all extremely sensitive genetic parameter combinations, and growth-related sensitive genetic parameter combinations were obtained. The yield prediction results obtained by different machine learning algorithms and sensitivity parameter combinations were compared, and the model with the highest yield prediction accuracy was recorded. Finally, the optimal soybean yield phenotype genetic parameter prediction model with genetic effects was obtained.
9. The method for quantitatively predicting soybean yield phenotype by integrating machine learning and gene effects according to claim 8, characterized in that The machine learning algorithm is a BayesC algorithm, a LASSO algorithm, a rrBLUP algorithm, a random forest algorithm, a support vector machine regression algorithm or a gradient boosting decision tree XGBoost algorithm.
Citation Information
Patent Citations
Methods of identifying immunomodulatory genes
CN113748126A
Machine learning method for estimating genotype specific parameters of soybean phenological period simulation model by using SNP (Single Nucleotide Polymorphism) marker
CN117854590A
Quantitative prediction method, system and device based on crop growth period phenotype and regional adaptability thereof
CN118171785A
Improved molecular breeding methods
WO2016069078A1
Breeding cross-generation phenotype prediction method and system based on ensemble learning, and electronic device
WO2024212036A1