A method for quantitatively predicting soybean yield phenotypes by integrating machine learning and gene effects

By combining machine learning and genetic effects, combined with genome-wide association analysis and multiple machine learning algorithms, the problem of difficulty in obtaining genetic parameters in the soybean yield phenotype simulation model is solved, and the accurate prediction of soybean yield phenotype in different environments is achieved, improving the genetic interpretability and prediction accuracy of the model.

CN120108512BActive Publication Date: 2025-07-04NANJING AGRICULTURAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510580072.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-07-04
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

It is difficult to obtain genetic parameters in traditional soybean growth simulation models, and it is impossible to accurately predict soybean yield phenotypes in multiple environments, and the influence of gene-environmental interaction effects is not considered.

Method used

Using a method of fusion machine learning and genetic effects, the phenotype sensitive parameters of soybean yield are obtained through genome-wide association analysis and multiple machine learning algorithms, combined with soybean growth simulation models, and accurate predictions are achieved in different environments.

Benefits of technology

It improves the genetic interpretability of the soybean yield phenotype simulation model, reduces the difficulty of parameter correction, saves test time and labor costs, provides digital support, and provides a theoretical basis for the screening and evaluation of soybean varieties in different ecological areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108512B_ABST
    Figure CN120108512B_ABST
Patent Text Reader

Abstract

The present invention proposes a method for quantitatively predicting soybean yield phenotypes by integrating machine learning and gene effects, including: taking soybean yield phenotypes as key phenotypic traits, dividing the genetic parameters in the soybean growth simulation model into three groups by using a genetic parameter sensitivity analysis method; comprehensively using various genome-wide association analysis techniques of different matrices (covariates) to obtain significant single nucleotide polymorphism (SNP) markers of soybean yield phenotype genetic parameters and analyze the genetic mechanism of soybean yield phenotypes; obtaining soybean yield phenotype sensitivity genetic parameters with gene effects according to multiple machine learning algorithms, thereby establishing a soybean yield phenotype gene effect simulation algorithm. The present invention can accurately predict the yield phenotypes and their change trends of different soybean varieties in different growth environments without relying on traditional parameter tuning tools and manual measurements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of agricultural information technology, and relates to a method for quantitatively predicting soybean yield phenotypes by integrating machine learning and gene effects, which is used to quantitatively predict soybean yield phenotypes with different gene data under different planting environments. Background Art

[0002] Soybean Glycine max (L.) Merr.] contains rich unsaturated fatty acids, high-quality proteins and various trace elements, has high nutritional value and economic value, and is one of the important oilseeds and feed grains, which is directly related to the supply security of important agricultural products such as edible oil, meat, eggs, and milk.

[0003] Soybean yield phenotypes are affected by the interaction between genotype and environment (G×E), and are the result of environmental conditions and the interaction between genes and the environment (G×E). Among them, gene information determines the variety characteristics of soybeans. Researchers can accurately identify QTL, SNP or genes themselves that cause variation in target traits at the genomic level, providing a strong guarantee for screening gene information of soybean yield phenotypes. At present, QTLs related to soybean roots, stems, yield, nutritional components, and biotic and abiotic stresses have been mapped using genome-wide association analysis methods, which has promoted the mining and identification of genes related to soybean yield. However, traditional genomic prediction models do not consider the effects of environmental effects and gene-environment interaction effects, resulting in the inability to accurately predict soybean yield phenotypes with different genetic backgrounds in multiple environments.

[0004] Soybean growth simulation models integrate the interaction effects of genotype (G) × environmental conditions (E) × management techniques (M), and can quantitatively describe and predict the soybean growth and development process and its dynamic relationship with the environment and technology. Therefore, the yield phenotypes of soybeans in different environments can be predicted through soybean growth simulation models, providing better planting and management decisions for agricultural producers. However, it is relatively difficult to obtain genetic parameters in traditional soybean growth simulation models, and the relationship and mechanism of action between soybean yield genetic parameters and genomes are still unclear, and the genetic interpretability of genetic parameters needs further study.

[0005] In summary, there is an urgent need for a method for quantitatively predicting soybean yield phenotypes by integrating machine learning and gene effects, which is used to improve the genetic interpretability of soybean yield phenotype simulation models, accelerate the phenotypic identification of soybean lines, and provide a theoretical basis and effective tool for screening and evaluating soybean genotype varieties in different ecological regions. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to provide a quantitative prediction method for soybean yield phenotype that integrates machine learning algorithms and gene effects. According to the obtained gene data and soybean model yield phenotype sensitive parameters, combined with genome-wide association analysis and machine learning algorithms, accurate prediction of the yield phenotypes of different soybean varieties in different environments is achieved.

[0007] The object of the present invention is achieved by adopting the following technical solutions:

[0008] A quantitative prediction method for soybean yield phenotype that integrates machine learning and gene effects, comprising the following steps:

[0009] Step 1: Collect soybean-related data, including: variety information, gene data, meteorological data and soil data of the planting location, field management measures, and yield formation data;

[0010] Step 2: Divide the obtained soybean-related data into an environmental training set and an environmental validation set according to the planting location and planting year;

[0011] Step 3: According to the soil data, meteorological data of the planting location, and the soybean growth simulation model, use the genetic parameter sensitivity analysis method to analyze the sensitivity of the yield phenotype genetic parameters of the soybean growth simulation model;

[0012] Step 4: According to the sensitivity information of the yield phenotype genetic parameters of the soybean growth simulation model, divide the yield phenotype genetic parameters of the soybean growth simulation model into three combinations: all sensitivity parameter combinations, all extremely sensitivity parameter combinations, and growth-related sensitivity parameter combinations;

[0013] Step 5: According to the collected soybean-related data, the divided environmental training set, environmental validation set, and soybean growth simulation model, use an optimization algorithm to obtain the numerical information of the sensitivity genetic parameters of different soybean varieties, that is, the optimal yield genetic parameter combination;

[0014] Step 6: Use multiple genome-wide association analysis (GWAS) methods of different matrices to perform association analysis on the obtained optimal yield genetic parameter combination of soybeans and gene data to obtain a sensitivity genetic parameter significant SNP marker dataset;

[0015] Step 7: Divide the optimal yield genetic parameter combination of soybeans and the sensitivity genetic parameter significant SNP marker dataset into a variety training set and a variety validation set according to a fixed ratio of the number of varieties;

[0016] Step 8: Use machine learning algorithms to perform cross-validation on the variety training set and use the variety validation set as an external validation for testing to obtain a prediction model for soybean yield phenotypic genetic parameters that integrates gene effects and machine learning algorithms; obtain soybean yield phenotypic genetic parameters with gene effects based on the prediction model for soybean yield phenotypic genetic parameters that integrates gene effects and machine learning algorithms.

[0017] Step 9: Input the obtained soybean yield phenotypic genetic parameters with gene effects into the soybean growth simulation model respectively to achieve accurate prediction of soybean yield phenotypes.

[0018] Preferably, step 1 includes the following steps:

[0019] (1) Collect data on the initial flowering stage, initial grain stage, initial pod stage, early maturity stage, maturity stage, number of pods per plant, number of grains per plant, 100-seed weight, and plot yield of soybeans according to field trials.

[0020] (2) Record the management measures of the soybean experimental field, including sowing, fertilization, irrigation, weeding, spraying pesticides, thinning seedlings, harvesting, and variety testing.

[0021] (3) Use gene sequencing technology to perform whole-genome sequencing on soybean varieties to obtain gene data of soybean varieties.

[0022] (4) Collect and organize soil data and meteorological data of the experimental area.

[0023] Preferably, step 2 includes the following steps:

[0024] Analyze the collected soybean growth, yield data, and meteorological data, and divide soybeans into an environmental training set and an environmental validation set according to the planting location and planting year.

[0025] Preferably, step 3 includes the following steps:

[0026] Input the collected meteorological data and soil data into the soybean growth simulation model, and use the Latin hypercube single-factor one-by-one test LH-OAT algorithm and the extended Fourier amplitude sensitivity test EFAST to perform sensitivity analysis on the yield phenotypic genetic parameters of the soybean growth simulation model; sort the sensitivity analysis values of the yield phenotypic genetic parameters of the soybean growth simulation model from high to low, and divide all sensitivity parameter combinations, all extremely sensitive parameter combinations, and growth-related sensitivity parameter combinations.

[0027] Preferably, step 5 includes the following steps:

[0028] (1) According to the collected soybean-related data and the soybean growth simulation model, input the meteorological data, soil data, and soybean growth and yield phenotype data of the environmental training set into the soybean growth simulation model to provide the necessary conditions for calibrating the yield phenotype genetic parameters of the soybean growth simulation model;

[0029] (2) Use the call code of the soybean growth simulation model combined with the differential evolution algorithm DE and the Markov chain Monte Carlo algorithm MCMC to calculate the simulated yield results of the yield phenotype genetic parameters of the soybean growth simulation model, and then calibrate the yield phenotype genetic parameters of the soybean growth simulation model to obtain the optimal yield genetic parameter combination of the soybean growth simulation model for each soybean variety;

[0030] (3) Input the environmental validation set into the soybean growth simulation model to test the prediction of the optimal yield genetic parameter combination.

[0031] Preferably, step 6 includes the following steps:

[0032] (1) Respectively associate all the sensitive genetic parameter combinations, all the extremely sensitive genetic parameter combinations, and the growth-related sensitive genetic parameter combinations of the obtained soybean yield with the gene data using 7 genome-wide association analysis GWAS methods of different matrices (P matrix + K matrix; Q matrix + K matrix);

[0033] (2) Take the union of the obtained significant SNP markers of the sensitivity genetic parameters to obtain the significant SNP marker datasets of all the yield sensitivity genetic parameters, all the extremely sensitivity genetic parameters, and the growth-related sensitivity genetic parameters of soybeans.

[0034] Preferably, step 7 includes the following steps:

[0035] Use the method of random sampling to divide the significant SNP marker datasets of all the yield sensitivity genetic parameters, all the extremely sensitivity genetic parameters, and the growth-related sensitivity genetic parameters of soybeans and the soybean growth and yield data into a variety training set and a variety validation set according to the ratio of 8:2 of the variety quantity.

[0036] Preferably, step 8 includes the following steps:

[0037] (1) Take the significant SNP marker datasets of all the sensitive genetic parameter combinations, all the extremely sensitive genetic parameter combinations, and the growth-related sensitive genetic parameter combinations and all the sensitive genetic parameter combinations, all the extremely sensitive genetic parameter combinations, and the growth-related sensitive genetic parameter combinations one by one as inputs, and use machine learning algorithms for algorithm learning and training;

[0038] (2) The ten-fold cross-validation method is adopted for the variety training set, and the purpose of this method is to test the stability of the constructed algorithm;

[0039] (3) The variety validation set is used as an external validation set to test the constructed algorithm, and the purpose is to test the generalization of the constructed algorithm;

[0040] (4) Through machine learning algorithm operations, the yield prediction results of the soybean yield phenotype model that fuses machine learning and gene effects for all sensitive genetic parameter combinations, all extremely sensitive genetic parameter combinations, and growth-related sensitive genetic parameter combinations are obtained. The yield prediction results obtained by different machine learning algorithms and sensitivity parameter combinations are compared, and the model with the highest yield prediction accuracy is recorded. Finally, the optimal soybean yield phenotype genetic parameter prediction model with gene effects is obtained.

[0041] Preferably, the machine learning algorithm is BayesC algorithm, LASSO algorithm, rrBLUP algorithm, random forest algorithm, support vector machine regression algorithm or gradient boosting decision tree XGBoost algorithm.

[0042] The beneficial effects brought by the present invention are:

[0043] (1) The present invention constructs a genetic parameter prediction model with gene effects through multiple machine learning algorithms, realizes the purpose of predicting soybean yield genetic parameters through gene information, effectively reduces the parameter calibration problem in the soybean yield phenotype model used alone, saves experimental test time and labor costs; and increases the genetic interpretability of the soybean yield phenotype model;

[0044] (2) The present invention divides different groups according to the sensitivity of soybean yield phenotype genetic parameters, and combines multiple genome-wide association analysis methods and two different matrix (covariate) combinations, effectively increasing the number of significant SNP markers obtained, and for the first time providing multiple processing schemes for the construction of the gene effect-soybean growth simulation model, providing technical support for the development of crop models;

[0045] (3) The present invention solves the problem of the fusion of gene effects and soybean growth simulation models, and can accurately predict soybean yield phenotypes under different climate conditions in different regions by using gene information and soybean yield phenotype models, providing digital support for the quantitative evaluation of yield phenotypes and the formulation of breeding strategies. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 It is the flow chart of the method in the present invention.

[0047] Figure 2 It is the data analysis chart of the soybean growth environment climate in the present invention.

[0048] Figure 3This is the sensitivity analysis diagram of genetic parameters of the soybean growth simulation model in the present invention.

[0049] Figure 4 This is the diagram of the simulated soybean yield results (environmental training set) using the soybean growth simulation model in the present invention.

[0050] Figure 5 This is the diagram of the predicted soybean yield results (environmental validation set) using the soybean growth simulation model in the present invention.

[0051] Figure 6 This is the data analysis diagram of soybean variety genes in the present invention.

[0052] Figure 7 This is the analysis diagram of the Q+K matrix of genome-wide association analysis in the present invention.

[0053] Figure 8 This is the analysis diagram of the P+K matrix of genome-wide association analysis in the present invention.

[0054] Figure 9 This is the prediction result diagram of the soybean yield genetic parameter variety training set and validation set by integrating multiple machine learning and gene effects in the present invention.

[0055] Figure 10 This is the prediction result diagram of the soybean yield environmental training set by integrating multiple machine learning algorithms and gene effects in the present invention.

[0056] Figure 11 This is the prediction result diagram of the soybean yield environmental validation set by integrating multiple machine learning algorithms and gene effects in the present invention. Detailed implementation manners

[0057] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and should not be used to limit the protection scope of the present invention.

[0058] Taking the main soybean cultivars in Northeast China and the Huang-Huai-Hai region as the research objects, the method of the present invention will be specifically described, including the following steps:

[0059] As Figure 1 shown, a phenotypic quantitative prediction method for soybean yield by integrating machine learning and gene effects includes the following steps:

[0060] Step 1: As shown in Table 1, 75 main cultivars in Northeast China and the Huang-Huai-Hai region were obtained and planted at 12 ecological sites in Northeast China and the Huang-Huai-Hai region from 2009 to 2011. The soybean phenotypic data include growth periods (sowing date, initial flowering date, maturity date) and yield phenotypic data (yield, number of pods per plant, number of seeds per plant, seed weight per plant, 100-seed weight); the soybean gene data are 3.38 million SNP markers obtained by second-generation sequencing (the reference gene line is Williams 82).

[0061] Table 1 Planting locations of main cultivated soybean varieties in Northeast China and the Huang-Huai-Hai region

[0062] Location Latitude Longitude Heihe 50.14 N 127.29°E Zhalantun 48.00°N 122.47°E Jiusan Research Institute 47.20 N 123.57 E Hailun 47.28°N 126.57°E Jiamusi 46.47°N 130.22 E Harbin 45.44°N 126.36°E Changchun 43.54°N 125.19 E Jilin 43.52°N 126.33 E Tieling 42.28°N 123.51°E Changping 40.22 N 116.20 E Jinan 36.67 N 117.00°E Xuchang 34.02°N 113.82 E

[0063] Step 2: Divide the obtained soybean-related data into an environmental training set and an environmental validation set according to the planting region and planting year. As Figure 2 shown, the data obtained from Northeast China and the Huang-Huai-Hai region are divided into four regions: northern Northeast China, central Northeast China, southern Northeast China, and the Huang-Huai-Hai region; the daily maximum temperature, daily minimum temperature, daily radiation value, and daily precipitation in the obtained meteorological data are sorted and analyzed according to the four regions, and the representative planting environment combinations in each region are extracted as the validation set, and other planting environment combinations are used as the training set;

[0064] Step 3: According to the soil data, meteorological data, and soybean growth simulation model at the planting location, use the genetic parameter sensitivity analysis method to analyze the sensitivity of the yield phenotypic genetic parameters of the soybean growth simulation model. As Figure 3 shown, the sorted meteorological data, soil data, and soybean field planting management measures at different soybean planting locations are input into the soybean yield phenotypic model DSSAT-CROPGRO-Soybean model; use the LH-OAT sensitivity analysis method and the EFAST sensitivity analysis method to perform sensitivity analysis on the soybean genetic parameters in the DSSAT-CROPGRO-Soybean model;

[0065] (1) Among them, the LH-OAT sensitivity analysis method is the simplest and most direct method for global sensitivity analysis of model parameters. First, according to the LH sampling idea, the entire parameter space is divided into K layers, and then a random sampling is performed once from each layer to generate an LH sampling parameter group (a parameter set containing m parameters). Then, according to the OAT idea, m minor parameter changes are made to each LH sampling parameter group (only one parameter is changed each time), and the change of the objective function before and after each minor change is calculated; the LH-OAT sensitivity analysis method is completed by the written code.

[0066] (2) The Extended Fourier Amplitude Sensitivity Test (EFAST) decomposes the variance of model parameters on the model output variables and classifies the parameter sensitivities into two categories: ① the sensitivity of a single parameter to the model output variable, measured by the first-order sensitivity index; ② the sensitivity of the interaction between parameters to the model output variable, measured by the total sensitivity index. The EFAST sensitivity analysis is completed using the SimLab sensitivity analysis software.

[0067] Step 4: According to the numerical values of the sensitivity analysis arranged from high to low, the genetic parameters of the soybean growth simulation model are divided into three groups: all sensitive parameters, all extremely sensitive parameters, and growth-related sensitive parameters.

[0068] Respectively: (1) All sensitive parameters of soybean yield: CSDL, PPSEN, EM-FL, FL-SH, FL-SD, SD-PM, LFMAX, SLAVR, SIZLF, WTPSD, SFDUR, PODUR, THRSH; All extremely sensitive parameters of soybean yield: CSDL, PPSEN, EM-FL, FL-SH, FL-SD; Growth-related sensitive parameters of soybean yield: LFMAX, SLAVR, SIZLF, WTPSD, SFDUR, PODUR, THRSH;

[0069] (2) All sensitive parameters of soybean pod number: CSDL, PPSEN, EM-FL, FL-SH, FL-SD, SD-PM, LFMAX, SLAVR, SIZLF, WTPSD, SFDUR, SDPDV, PODUR, THRSH; All extremely sensitive parameters of soybean pod number: CSDL, PPSEN, EM-FL, FL-SH, FL-SD, WTPSD, PODUR; Growth-related sensitive parameters of soybean pod number: LFMAX, SLAVR, SIZLF, WTPSD, SFDUR, SDPDV, PODUR, THRSH;

[0070] (3) All sensitive parameters of soybean grain number: CSDL, PPSEN, EM-FL, FL-SH, FL-SD, SD-PM, LFMAX, SLAVR, SIZLF, WTPSD, SFDUR, PODUR; All extremely sensitive parameters of soybean grain number: CSDL, PPSEN, EM-FL, FL-SH, FL-SD, WTPSD; Growth-related sensitive parameters of soybean grain number: LFMAX, SLAVR, SIZLF, WTPSD, SFDUR, PODUR.

[0071] Step 5: According to the collected soybean-related data and the soybean growth simulation model, use the optimization algorithm to obtain the sensitivity genetic parameter information of different soybean varieties. Such as Figure 4As shown in the figure, the differential evolution algorithm (DE) is used to correct the genetic parameters of soybean yield, and the optimal genetic parameter combination in the training set is obtained; as Figure 5 shown, the environmental data selected from the validation set is input into the model for verification to test the generalization ability of the optimal genetic parameter combination.

[0072] Among them, the differential evolution algorithm (DE) searches and optimizes through the mutation operation in the form of difference and the crossover operation of probabilistic selection to obtain the optimal parameter combination. First, given the parameter correction range, the algorithm randomly extracts two parameter arrays, and adds the weighted difference vector to the array to generate a new parameter vector; the parameters of the mutant vector and the parameters of another pre-determined target vector are mixed according to certain rules to generate sub-individuals; then the better one is selected as the next generation between the target individual (simulated value) and the trial individual (measured value), gradually narrowing the range, and finally the optimal parameter combination is output.

[0073] Step 6: Use various genome-wide association analysis (GWAS) methods with different matrices to perform association analysis on the obtained soybean variety genetic parameters and gene data to obtain significant SNP markers for sensitive genetic parameters;

[0074] As Figure 6 shown, first perform quality control on the gene data, then calculate the linkage disequilibrium of the soybean population, and further perform population structure and principal component analysis on the soybean population;

[0075] As Figure 7 and Figure 8 shown, use a total of 7 GWAS analysis methods including IIIVmrMLM, mrMLM, FASTmrMLM, FASTmrEMMA, pLARmEB, pKWmEB, and ISIS EM-BLASSO, and use the P matrix + K matrix and the Q matrix + K matrix as covariates to analyze the quality-controlled gene data respectively, and screen out significant SNP markers related to genetic parameters.

[0076] Step 7: Divide the soybeans into a variety training set and a variety validation set according to the ratio of 8:2 of the variety quantity by the method of random extraction;

[0077] Step 8: Combine the obtained significant SNP markers with machine learning algorithms, and use a method combining cross-validation and external validation to obtain the phenotypic genetic parameters of soybean yield with integrated gene effects;

[0078] As shown in Table 2, convert the significant SNP markers into a digital format of 0, 1, 2 for the learning of computer algorithms;

[0079] Table 2 Corresponding table of gene information conversion codes

[0080] Gene information Conversion code AA 0 Heterozygous 1 aa 2

[0081] As Figure 9 shown, the three groups of significant SNP marker datasets obtained and the measured data of soybean yield phenotypes are used as inputs, and algorithm learning and training are carried out through the written machine learning algorithm codes (BayesC algorithm, LASSO algorithm, rrBLUP algorithm, random forest algorithm (RF), support vector machine regression algorithm (SVR), and gradient boosting decision tree (XGBoost) algorithm);

[0082] (1) First, use the variety training set for algorithm learning, obtain the optimal hyperparameter combination of the machine learning algorithm through the Bayesian algorithm, and use the ten-fold cross-validation method to test the stability of the constructed algorithm;

[0083] (2) After the machine learning algorithm completes data learning, use the gene data of the variety training set to predict the genetic parameters of gene effects;

[0084] (3) Next, use the gene information of the variety validation set to predict the genetic parameters of gene effects. The variety validation set is used as an external validation set to test the generalization of the constructed algorithm, so as to obtain the optimal combined value of the genetic parameters of soybean yield phenotypes that integrate gene effects.

[0085] (3.1) Among them, the ridge regression best linear unbiased prediction (rrBLUP) method is a multiple genomic information prediction method based on the Bayesian framework. It simulates variety parameters by estimating the intercept and the effect values of SNP markers, and further constructs a linear regression model through learning the input features to complete parameter prediction;

[0086] (3.2) The operation principle of BayesC is based on the Bayesian statistical method. It models gene effects as being sampled from a prior distribution with zero mean, effectively reducing noise and improving prediction accuracy; in the learning of gene data, BayesC can run stably and produce high-quality prediction results;

[0087] (3.3) LASSO (Least Absolute Shrinkage and Selection Operator) is a linear regression method aimed at improving the prediction accuracy of the model by introducing the L1 regularization term while performing feature selection; LASSO uses the L1 penalty of the regression coefficient to control the complexity of the model, thereby reducing overfitting and further improving the prediction effect of gene effect parameters.

[0088] (3.4) The Random Forest (RF) algorithm is a machine learning algorithm based on ensemble learning. First, a series of decision trees are generated, each of which is trained on an independent data set. Sampling with replacement is used to add sample perturbations, and at the same time, an attribute perturbation is introduced. As the number of learning samples increases, the random forest will gradually converge, and finally, the optimal genetic effect genetic parameters are obtained through analysis.

[0089] (3.5) Support Vector Machines (SVM) is a binary classification model. Its purpose is to find a hyperplane to segment the samples, and the principle of segmentation is to maximize the margin. SVM is a non-linear method. When the sample size is relatively small, it can capture the non-linear relationship between data and features, avoid the problems of neural network structure selection and local optimality, and obtain the optimal prediction results.

[0090] (3.6) XGBoost (eXtreme Gradient Boosting) is a boosting tree model, and its basic component is a decision tree. It can perform a second-order Taylor expansion on the loss function, increasing the accuracy while being able to customize the loss function. XGBoost adds a regularization term to the objective function to penalize complex models, reduce the model variance, and reduce overfitting, thus obtaining the optimal prediction results.

[0091] Step 9: Input the obtained genetic parameters of the fusion gene effect and the soybean yield phenotype of the machine learning algorithm into the soybean growth simulation model respectively to achieve accurate prediction of the soybean yield phenotype.

[0092] As Figure 10 and Figure 11 shown, input the genetic parameters of the soybean yield phenotype gene effect predicted by the 6 machine learning methods obtained in Step 8, the meteorological data, soil data, and management measures collected in Step 1 into the DSSAT-CROPGRO-Soybean model, and obtain the yield phenotype prediction results of different soybean varieties at different planting locations through model operation.

[0093] Select the combination with the lowest RMSE, nRMSE, and MAF among the 18 methods, which are the 6 machine learning algorithms within the group of all sensitivity genetic parameters of gene effect - yield phenotype, all extremely sensitivity genetic parameters of gene effect - yield phenotype, and growth-related sensitivity genetic parameters of gene effect - yield phenotype, as the optimal combination of the constructed fusion model. In the research object of this study, the method of the growth-related sensitivity genetic parameters of gene effect - yield phenotype that fuses the SVR algorithm has the best simulation prediction effect.

[0094] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A method for quantitatively predicting soybean yield phenotypes by integrating machine learning and gene effects, characterized in that, It includes the following steps: Step 1: Collect soybean-related data, including: variety information, gene data, meteorological data and soil data of the planting location, field management measures and yield formation data; Step 2: Divide the obtained soybean-related data into an environmental training set and an environmental validation set according to the planting location and planting year; Step 3: According to the soil data, meteorological data of the planting location and the soybean growth simulation model, use the genetic parameter sensitivity analysis method to analyze the sensitivity of the yield phenotypic genetic parameters of the soybean growth simulation model; Step 4: According to the sensitivity information of the yield phenotypic genetic parameters of the soybean growth simulation model, divide the yield phenotypic genetic parameters of the soybean growth simulation model into three combinations: all sensitivity parameter combinations, all extremely sensitivity parameter combinations and growth-related sensitivity parameter combinations; Step 5: According to the collected soybean-related data, the divided environmental training set, environmental validation set and soybean growth simulation model, use the optimization algorithm to obtain the sensitivity genetic parameter numerical information of different soybean varieties, that is, the optimal yield genetic parameter combination; Step 6: Use multiple genome-wide association analysis GWAS methods with different matrices to perform association analysis on the obtained optimal soybean yield genetic parameter combination and gene data to obtain a sensitivity genetic parameter significant SNP marker dataset; Step 7: Divide the optimal soybean yield genetic parameter combination and the sensitivity genetic parameter significant SNP marker dataset into a variety training set and a variety validation set according to a fixed ratio of the number of varieties; Step 8: Use machine learning algorithms to perform cross-validation on the variety training set and use the variety validation set as an external validation for testing to obtain a soybean yield phenotypic genetic parameter prediction model that combines gene effects and machine learning algorithms; Obtain soybean yield phenotypic genetic parameters with gene effects based on the soybean yield phenotypic genetic parameter prediction model that combines gene effects and machine learning algorithms; Step 9: Input the obtained soybean yield phenotypic genetic parameters with gene effects into the soybean growth simulation model respectively to achieve accurate prediction of the soybean yield phenotype.

2. The quantitative prediction method for soybean yield phenotype integrating machine learning and gene effects according to claim 1, wherein The said Step 1 includes the following steps: (1) Collect data on the initial flowering stage, initial grain stage, initial pod stage, early maturity stage, maturity stage, number of pods per plant, number of grains per plant, 100-grain weight, and plot yield of soybeans according to field trials; (2) Record the management measures of the soybean experimental field, including sowing, fertilizing, irrigating, weeding, spraying pesticides, thinning, harvesting, and variety testing; (3) Use gene sequencing technology to perform whole-genome sequencing on soybean varieties to obtain gene data of soybean varieties; (4) Collect and organize soil data and meteorological data of the experimental area.

3. A method for quantitatively predicting soybean yield phenotype by integrating machine learning and gene effects according to claim 1, characterized in that, The said Step 2 includes the following steps: Analyze the collected soybean growth, yield data and meteorological data, and divide the soybeans into an environmental training set and an environmental validation set according to the planting location and planting year.

4. A method for quantitatively predicting soybean yield phenotypes by integrating machine learning and gene effects according to claim 1, characterized in that The said Step 3 includes the following steps: The collected meteorological data and soil data are input into the soybean growth simulation model, and the Latin hypercube single factor is used to test the LH-OAT algorithm and the extended Fourier amplitude sensitivity test EFAST one by one for the sensitivity analysis of the yield phenotypic genetic parameters of the soybean growth simulation model; according to the sensitivity analysis values of the yield phenotypic genetic parameters of the soybean growth simulation model, they are sorted from high to low, and all sensitivity parameter combinations, all extremely sensitive parameter combinations, and growth-related sensitivity parameter combinations are divided.

5. A quantitative prediction method for soybean yield phenotype integrating machine learning and gene effects according to claim 1, characterized in that Step 5 includes the following steps: (1) According to the collected soybean-related data and the soybean growth simulation model, the meteorological data, soil data, soybean growth, and yield phenotypic data of the environmental training set are input into the soybean growth simulation model to provide necessary conditions for the calibration of the yield phenotypic genetic parameters of the soybean growth simulation model; (2) Use the call code of the soybean growth simulation model to combine the differential evolution algorithm DE and the Markov chain Monte Carlo algorithm MCMC to calculate the simulated yield results of the yield phenotypic genetic parameters of the soybean growth simulation model, and then calibrate the yield phenotypic genetic parameters of the soybean growth simulation model to obtain the optimal yield genetic parameter combination of the soybean growth simulation model for each soybean variety; (3) Input the environmental validation set into the soybean growth simulation model to test the prediction of the optimal yield genetic parameter combination.

6. The quantitative prediction method for soybean yield phenotype integrating machine learning and gene effects according to claim 1, characterized in that Step 6 includes the following steps: (1) The obtained soybean yield all sensitive genetic parameter combinations, all extremely sensitive genetic parameter combinations, and growth-related sensitive genetic parameter combinations are respectively associated with gene data using 7 genome-wide association analysis GWAS methods with different matrices; (2) Take the union of the obtained sensitivity genetic parameter significant SNP markers to obtain the soybean all yield sensitivity genetic parameter significant SNP marker dataset, all extremely sensitivity genetic parameter significant SNP marker dataset, and growth-related sensitivity genetic parameter significant SNP marker dataset.

7. A quantitative prediction method for soybean yield phenotype integrating machine learning and gene effects according to claim 1, characterized in that Step 7 includes the following steps: The soybean all yield sensitivity genetic parameter significant SNP marker dataset, all extremely sensitivity genetic parameter significant SNP marker dataset, growth-related sensitivity genetic parameter significant SNP marker dataset, and soybean growth and yield data are randomly selected and divided into a variety training set and a variety validation set according to a ratio of 8:2 of the number of varieties.

8. A method for quantitatively predicting the soybean yield phenotype by integrating machine learning and gene effects according to claim 1, characterized in that, Step 8 includes the following steps: (1) The obtained significant SNP marker datasets of all sensitive genetic parameter combinations, all extremely sensitive genetic parameter combinations, and growth-related sensitive genetic parameter combinations and all sensitive genetic parameter combinations, all extremely sensitive genetic parameter combinations, and growth-related sensitive genetic parameter combinations are used as inputs one by one, and machine learning algorithms are used for algorithm learning and training; (2) The ten-fold cross-validation method is used for the variety training set; (3) The variety validation set is used as an external validation set to test the constructed algorithm; (4) Through the operation of machine learning algorithms, obtain the yield prediction results of the soybean yield phenotypic model that integrates machine learning and gene effects for all sensitive genetic parameter combinations, all extremely sensitive genetic parameter combinations, and growth-related sensitive genetic parameter combinations. Compare the yield prediction results obtained by different machine learning algorithms and sensitivity parameter combinations, record the model with the highest yield prediction accuracy, and finally obtain the optimal soybean yield phenotypic genetic parameter prediction model with gene effects.

9. A method for quantitatively predicting soybean yield phenotype by integrating machine learning and gene effects according to claim 8, characterized in that The machine learning algorithms are BayesC algorithm, LASSO algorithm, rrBLUP algorithm, random forest algorithm, support vector machine regression algorithm, or gradient boosting decision tree XGBoost algorithm.

Citation Information

Patent Citations

  • Machine learning method for estimating genotype specific parameters of soybean phenological period simulation model by using SNP (Single Nucleotide Polymorphism) marker

    CN117854590A

  • Quantitative prediction method, system and device based on crop growth period phenotype and regional adaptability thereof

    CN118171785A