A crop phenotype prediction method and a crop phenotype prediction device
By constructing a prediction model that integrates feature matrices and combining parental genotypes and multi-environmental phenotypic data, the problems of long breeding cycles and poor accuracy in existing technologies have been solved, enabling rapid and accurate crop phenotypic prediction and improving breeding efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INSTITUTE OF CROP SCIENCE CHINESE ACADEMY OF AGRICULTURAL SCIENCES
- Filing Date
- 2025-12-03
- Publication Date
- 2026-04-24
AI Technical Summary
Existing wheat breeding methods require phenotypic identification at multiple locations over many years, resulting in long breeding cycles, high costs, and low efficiency. Genomic selection methods cannot take into account environmental differences, leading to poor prediction accuracy, especially in unknown environments.
By determining the genotype and phenotypic data of the parents, a fusion feature matrix is constructed using a trained prediction model combined with multi-environmental data to predict the crop phenotype under the current environment.
It enables efficient and accurate crop phenotypic prediction across environments, shortens the breeding cycle, reduces costs, and improves breeding efficiency and accuracy.
Smart Images

Figure CN121237200B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of agricultural technology, specifically to a crop phenotypic prediction method and a crop phenotypic prediction device. Background Technology
[0002] Current breeding methods for crops such as wheat generally employ multi-year, multi-location phenotypic identification. This requires planting multiple lines of recombinant inbred lines in various environments to determine the crop's phenotype under different conditions. This results in a long breeding cycle, high costs, and low overall breeding efficiency. Some technologies use genomic selection methods, predicting breeding values through crop genomic markers, which can shorten the breeding cycle and improve efficiency to some extent. However, because genomic selection methods cannot consider phenotypic differences of the same genotype under different environments, their predictive accuracy is poor when predicting crop phenotypes in unknown environments. Summary of the Invention
[0003] To overcome the problems existing in related technologies, an exemplary embodiment of this disclosure provides a crop phenotypic prediction method in a first aspect for predicting crop phenotypes across environments. The crop phenotypic prediction method includes: determining the genotype of a parent; determining the phenotypic data of the parent in the current environment; inputting the genotype of the parent and the phenotypic data of the parent into a trained prediction model to determine a predicted value of the crop phenotype in the current environment, wherein the prediction model is trained based on the genotypes of multiple lines of the recombinant inbred line of the parent and the phenotypic data of the multiple lines of the recombinant inbred line in multiple environments.
[0004] In some embodiments, determining the phenotypic data of the parent in the current environment includes: planting the parent in the current environment; and determining the phenotypic data of the parent planted in the current environment.
[0005] In some embodiments, the prediction model is trained by the following method: determining the genotype matrix of multiple lines of the recombinant inbred line; determining the phenotypic data of the multiple lines of the recombinant inbred line in each environment based on multiple environments, wherein the phenotypic data includes at least one of the following: grain yield, thousand-grain weight, number of grains per ear, number of ears per unit area, and plant height; determining the plasticity characteristics of each line based on the phenotypic data of the multiple lines of the recombinant inbred line in each environment; and training the model based on the plasticity characteristics and the genotype matrix to determine the completed prediction model.
[0006] In some embodiments, determining the genotype matrix of multiple lines of the recombinant inbred line includes: determining all genotyping data of the multiple lines of the recombinant inbred line; and based on all the genotyping data, screening genetic markers of the multiple lines of the recombinant inbred line to determine the genotype matrix.
[0007] In some embodiments, determining the phenotypic data of multiple lines of the recombinant inbred line in each environment based on multiple environments includes: planting multiple lines of the recombinant inbred line of the parent multiple times in multiple environments; determining the phenotypic data of each line after each planting in multiple environments; and determining the average value of the phenotypic data determined by each line after multiple plantings in the same environment as the phenotypic data of the current line in the current environment based on the phenotypic data.
[0008] In some embodiments, determining the plasticity characteristic of each line based on phenotypic data of multiple lines of the recombinant inbred line in each environment includes: determining the population mean of the phenotypic data of multiple lines of the recombinant inbred line in each environment; fitting a linear regression model of the crop phenotype of each line with the population mean as the independent variable and the phenotypic data of multiple lines of the recombinant inbred line in each environment as the dependent variable; and determining the plasticity characteristic based on the linear regression model of the crop phenotype, wherein the plasticity characteristic is a feature vector composed of the slopes of the linear regression models of all crop phenotypes.
[0009] In some embodiments, the step of training the model based on the plasticity features and the genotype matrix to determine the completed prediction model includes: constructing a fusion feature matrix based on the plasticity features and the genotype matrix; and training the model based on the fusion feature matrix to determine the completed prediction model.
[0010] In some embodiments, the step of training the model based on the fused feature matrix to determine the completed prediction model includes: inputting the fused feature matrix into multiple candidate models for training, wherein the multiple candidate models include at least two of the following: minimum absolute shrinkage and selection algorithm model, ridge regression model, partial least squares regression model, support vector regression model, and random forest model; verifying the multiple candidate models after training using a 10-fold crossover and / or leave-one-out environment method, and determining the verification result, wherein the verification result includes: prediction correlation coefficient, and / or root mean square error; and determining one of the candidate models as the completed prediction model based on the verification result.
[0011] Secondly, this disclosure also provides a crop phenotypic prediction device for performing the crop phenotypic prediction method as described in the first aspect, wherein the crop phenotypic prediction device includes: a genotype determination module for determining the genotype of the parent; a phenotypic data processing module for determining the phenotypic data of the parent; and a crop phenotypic prediction module for determining the phenotypic prediction value of the crop based on the genotype of the parent and the phenotypic data.
[0012] Thirdly, this disclosure also provides a computer-readable storage medium storing a program for performing the crop phenotypic prediction method as described in the first aspect.
[0013] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this disclosure. According to the crop phenotypic prediction method provided by this disclosure, the predicted phenotypic value of the recombinant inbred line population under any environment can be directly determined by the prediction model based on the genotype of the parents and their phenotypic characteristics under any environment. Through this disclosure, the influence of parental genotype and environmental factors on crop phenotypic effects can be comprehensively considered through the prediction model, thereby predicting crop phenotypic values more quickly and accurately. The crop phenotypic prediction method provided by this disclosure only requires the genotype of the parents and their phenotypic data under the current environment to determine the predicted phenotypic value of their recombinant inbred line population under the current environment, without the need for planting crops and conducting phenotypic identification at multiple locations over many years. This effectively shortens the breeding cycle and reduces costs, and has high accuracy and reliability in predicting crop phenotypic effects under unknown environments. Attached Figure Description
[0014] This disclosure can be better understood by describing exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, in which:
[0015] Figure 1 This is a flowchart illustrating a crop phenotypic prediction method according to a published exemplary embodiment;
[0016] Figure 2 This is a flowchart illustrating a crop phenotypic prediction method according to a published exemplary embodiment;
[0017] Figure 3 This is a flowchart illustrating a crop phenotypic prediction method according to a published exemplary embodiment;
[0018] Figure 4 This is a flowchart illustrating a crop phenotypic prediction method according to a published exemplary embodiment;
[0019] Figure 5 This is a flowchart illustrating a crop phenotypic prediction method according to a published exemplary embodiment;
[0020] Figure 6 This is a flowchart illustrating a crop phenotypic prediction method according to a published exemplary embodiment;
[0021] Figure 7 This is a scatter plot of measured and predicted values of crop grain count per ear, as shown in an exemplary embodiment of a published document.
[0022] Figure 8 This is a scatter plot of measured and predicted values of the number of ears per unit area of crop, as shown in an exemplary embodiment of a public document.
[0023] Figure 9 This is a scatter plot of measured and predicted values of crop thousand-grain weight, as shown in an exemplary embodiment of a published document.
[0024] Figure 10 It is a scatter plot of measured and predicted crop heights according to an exemplary embodiment of a public document;
[0025] Figure 11 It is a scatter plot of measured and predicted crop grain yields according to an exemplary embodiment of a public document;
[0026] Figure 12 This is a schematic diagram illustrating the relationship between the test set proportion and prediction accuracy of crop ear grain number data according to an exemplary embodiment of a published document;
[0027] Figure 13 This is a schematic diagram illustrating the relationship between the test set proportion and prediction accuracy of crop ear count per unit area data, based on an exemplary disclosed invention.
[0028] Figure 14 This is a schematic diagram illustrating the relationship between the test set proportion and prediction accuracy of crop thousand-grain weight data according to an exemplary disclosed invention.
[0029] Figure 15 This is a schematic diagram illustrating the relationship between the test set proportion and prediction accuracy of crop plant height data according to an exemplary embodiment of a public document.
[0030] Figure 16 This is a schematic diagram illustrating the relationship between the test set proportion and prediction accuracy of crop grain yield data according to an exemplary disclosed invention.
[0031] Figure 17 This is a schematic diagram illustrating, according to an exemplary disclosed invention, the prediction of yield traits of 10% of test lines in a population under unknown conditions using phenotypic data from 90% of individuals. Detailed Implementation
[0032] The following describes specific embodiments of this disclosure. It should be noted that, in order to maintain brevity, this specification cannot provide a detailed description of all features of the actual embodiments. It should be understood that, in the actual implementation of any embodiment, just as in any engineering or design project, various specific decisions are often made to achieve the developer's specific goals and to meet system-related or business-related constraints, and this can change from one embodiment to another. Furthermore, it is understood that although the efforts made in this development process may be complex and lengthy, for those skilled in the art related to the content of this disclosure, changes in design, manufacturing, or production based on the technical content disclosed herein are merely conventional technical means and should not be construed as insufficient content of this disclosure.
[0033] For crop breeding, especially wheat breeding, the traditional method involves repeatedly planting the parent lines and recombinant inbred lines formed through multiple generations of self-pollination under various conditions. Phenotypic data, such as grain yield, thousand-grain weight, and number of grains per ear, are directly obtained from these parent lines and their recombinant inbred lines. However, this multi-year, multi-location phenotypic identification method is time-consuming, costly, and inefficient due to the need for repeated plantings of the parent lines and their recombinant inbred lines under different conditions. Current technology can employ genomic selection, which uses genomic markers from the parent lines and their recombinant inbred lines to determine the phenotypic data of the recombinant inbred lines formed through multiple generations of self-pollination under these parent lines. This data can then be used as a basis for breeding, thus shortening the breeding cycle to some extent. However, current genomic selection methods use models built on crop phenotypic data from a single environment, without considering the plasticity of crop phenotypes, i.e. the biological characteristic that crops with the same genotype can exhibit phenotypic differences in different environments. As a result, genomic selection methods have poor accuracy in predicting crop phenotypes, especially in predicting crop phenotypes in unknown environments, which affects the breeding process and leads to low accuracy and efficiency in breeding.
[0034] To address the aforementioned problems, exemplary embodiments of this disclosure provide a crop phenotype prediction method for predicting crop phenotypes across environmental contexts, such as... Figure 1 As shown, the crop phenotype prediction method includes steps S110 to S130.
[0035] Step S110: Determine the genotypes of the parents. First, the genotypes of the two crops used as parents during breeding can be determined using an SNP (single nucleotide polymorphism) chip. SNP chips can quickly detect SNP markers in the parent genome. SNP markers are genetic markers formed by single nucleotide variations in the genome; therefore, detection can determine the genotype data of the parents. Specifically, a wheat 50K SNP chip can be used to genotype the parents, thereby determining their genotypes and obtaining all SNP markers in the parent genome. Furthermore, after obtaining all SNP markers in the parent genome, quality control can be performed on the SNP markers to screen and identify high-quality SNP markers. In addition, a genotype matrix can be constructed based on all high-quality SNP markers for subsequent input into a prediction model.
[0036] Step S120: Determine the phenotypic data of the parents in the current environment. The current environment refers to the planting environment after breeding with this set of parents. This step eliminates the need for high-precision and continuous environmental monitoring; only the phenotypic data of the parents in the current environment needs to be determined. In subsequent steps, based on the phenotypic data and genotype of the parents in the current environment, the predicted phenotypic values of the crop bred from this set of parents can be determined. The phenotypic data of the parents may include, but is not limited to, the following traits: grain yield, thousand-grain weight, number of grains per ear, number of ears per unit area, plant height, etc.
[0037] In some embodiments, step S120, determining the phenotypic data of the parents in the current environment, may include planting the parents in the current environment and determining the phenotypic data of the parents planted in the current environment. Firstly, two parents can be directly planted in the current environment, and after the parents mature, measurements can be directly taken on the mature parents to determine their phenotypic data in the current environment. In some embodiments, after planting the parents, the mean phenotypic value of the two parents in the current environment can be determined. Using this mean as the independent variable and the phenotypic data of the two parents in the current environment as the dependent variable, a linear regression model is fitted. The slope of the linear regression model is used as the plasticity index of the parents' phenotypic characteristics in the current environment. Based on the plasticity index, an eigenvector is constructed, which represents the plasticity characteristic value of the parents.
[0038] Step S130: Input the genotype and phenotypic data of the parents into the trained prediction model to determine the predicted phenotypic value of the crop in the current environment. The prediction model is trained based on the genotypes of multiple lines of the recombinant inbred lines of the parents and the phenotypic data of these lines in multiple environments. By inputting the genotypes and phenotypic data of the parents in the current environment into the trained prediction model, the model can directly output the predicted phenotypic value of the offspring that stably inherits the parental traits obtained through breeding with this set of parents, when planted in the current environment. In other words, the predicted phenotypic value of the recombinant inbred line population of this set of parents under the current planting conditions. Alternatively, the genotype information and phenotypic data of the parents can be fused, and the fused data can be input into the trained prediction model. The prediction model then outputs the predicted phenotypic value of the offspring crops of this set of parents. Specifically, the genotype matrix of the parents can be fused with their plasticity eigenvalues to determine the fusion feature matrix of the parents. This fusion feature matrix can include multiple SNP markers of the parents, the environment in which the parents were planted, and information on their phenotypic data. By inputting the fusion feature matrix of the parents into the prediction model, the model can be provided with information such as the parents' genotype, phenotypic data, and the environment in which the parents were planted. This allows the prediction model to forecast the phenotypic data of offspring bred from this set of parents of the same specifications, when planted in the current environment. Through step S130, only the parents' genes and the phenotypic data determined by planting the parents in the current environment need to be input to quickly and accurately determine the phenotypic data of offspring bred from this set of parents when planted in the current environment. Therefore, only the genotypes of the parents and their phenotypic data under the current environment need to be determined as input to the model. Only planting and data collection of the parents under the current environment are required; there is no need for hybridization breeding of the parents or planting of the hybrid offspring lines. This effectively reduces the time required for the breeding process, shortens the breeding cycle, and lowers breeding costs, thus significantly improving breeding efficiency. The prediction model provided in this embodiment can be trained based on the genotype data of multiple lines in the recombinant inbred lines of the parents, and the phenotypic data determined after planting these multiple lines in multiple environments. Specifically, a group of parents can be hybridized and then subjected to multiple generations of self-pollination to determine the recombinant inbred line population of this group of parents, which includes multiple lines. The genotypes of the parents and the multiple lines in the recombinant inbred line population, i.e., the SNP markers of the parents and the multiple lines in the recombinant inbred line population, can be used as training samples for the prediction model. Meanwhile, multiple lines in the recombinant inbred line population can be planted in multiple different environments, and the phenotypic data of each line in different environments can be determined and used as training samples for the predictive model.By using genotypes of multiple lines of recombinant inbred lines from the parent plants, as well as phenotypic data of these lines in various environments, as training samples for the model, the influence of both environment and genotype on crop phenotype can be considered simultaneously during model training. This allows the trained model to determine the phenotype of the recombinant inbred line population in a corresponding environment based on the parent genotype and its phenotype in any given environment. This enables cross-environmental prediction of crop phenotypes without the need for direct monitoring of environmental data, effectively reducing data acquisition costs in the breeding process.
[0039] Therefore, the crop phenotypic prediction method provided in this embodiment can achieve cross-environment crop phenotypic prediction by combining the genotype of the parent and its phenotypic data under the current environment, and inputting both into a prediction model trained with multi-environment data. Compared with existing genome selection methods that build models based solely on phenotypic data under a single environment, this embodiment significantly improves the accuracy and stability of predictions by introducing a model trained with phenotypic data from multiple environments. By using model training samples containing multi-environment phenotypic information in this embodiment, the model can infer the phenotypic characteristics of recombinant inbred lines of the parents in the corresponding environment, even when the input only contains phenotypic information of the parents under any environment. This eliminates the need for high-precision monitoring or additional modeling of environmental factors, simplifying the data collection and processing process. Thus, high-precision phenotypic prediction under cross-environmental conditions can be achieved, reducing prediction bias caused by environmental changes. Experiments show that compared with traditional single-environment prediction models, this embodiment can effectively improve the accuracy of crop phenotypic prediction results.
[0040] Furthermore, this method eliminates the need for multi-location, multi-year planting trials and progeny population validation of parents. It only requires a single acquisition of the parents' genotype and single-environment phenotypic data to quickly obtain prediction results, significantly reducing the number of field trials and breeding cycles, thus improving breeding efficiency and resource utilization. Simultaneously, the prediction model provided in this embodiment maintains high generalization ability across multiple environmental scenarios, demonstrating good adaptability and robustness to different crop varieties, climate zones, or planting conditions. Therefore, the crop phenotypic prediction method of this embodiment can effectively improve the accuracy and efficiency of cross-environment crop phenotypic prediction, achieving efficient and reliable inference of crop traits under complex environmental conditions.
[0041] In some embodiments, such as Figure 2 As shown, the prediction model is trained by the following method, including steps S210 to S240.
[0042] Step S210: Determine the genotype matrix of multiple lines of the recombinant inbred line. Genotyping data can be determined by using SNP chips to genotype multiple lines of the recombinant inbred line. This genotyping data may include multiple SNP markers from the multiple lines of the recombinant inbred line. Based on the genotyping data, the genotype matrix of the multiple lines of the recombinant inbred line can be further determined. The genotype matrix can characterize genes with relatively obvious genetic characteristics in the recombinant inbred line.
[0043] Step S220: Based on multiple environments, determine the phenotypic data of multiple lines of the recombinant inbred line in each environment. The phenotypic data includes at least one of the following: grain yield, thousand-grain weight, number of grains per ear, number of ears per unit area, and plant height. First, multiple different environments suitable for crop cultivation can be identified. Under each environment, determine the phenotypic data of multiple lines in the recombinant inbred line population obtained through parental hybridization. Phenotypic data are data that characterize crop yield, traits, etc., and may include: grain yield, thousand-grain weight, number of grains per ear, number of ears per unit area, and plant height. Through phenotypic data, the traits and yield of each line under different environments can be characterized.
[0044] Step S230: Based on the phenotypic data of multiple lines of the recombinant inbred line in each environment, determine the plasticity characteristics of each line. Since crops exhibit phenotypic plasticity—that is, crops with the same genotype may show phenotypic differences when grown in different environments—the plasticity characteristics of each line can be determined based on the phenotypic data of multiple lines of the recombinant inbred line in multiple different environments. Based on the plasticity characteristics of each line of the recombinant inbred line, the relationship between the phenotypic data of each line and the environment can be determined.
[0045] Step S240: Train the model based on plasticity features and the genotype matrix to determine the completed prediction model. Plasticity features and the genotype matrix can be used as training samples for the model to train and obtain a completed prediction model. The prediction model can determine the phenotypic prediction values of the hybrid offspring of the parents based on the genotype and phenotypic data of the parents, thus accurately predicting the yield and traits of the hybrid offspring of the parents under unknown environments. Here, the unknown environment can be an environment where continuous and high-precision environmental monitoring data cannot be obtained. Training the model based on plasticity features and the genotype matrix allows the model to learn the relationship between environment and phenotype in plasticity features, as well as the genotypes of the parents and their recombinant inbred lines. Therefore, the completed prediction model can comprehensively consider the influence of genotype and environmental factors on the phenotype when determining the predicted value of the crop phenotype, thereby achieving accurate and rapid prediction of crop phenotypes.
[0046] According to the prediction model training method provided in this embodiment, by introducing plasticity features and genotype matrices as joint training inputs, the accuracy and generalization ability of the prediction model in cross-environment crop phenotypic prediction can be significantly improved. Unlike traditional genome selection models that are trained based on only single environmental phenotypic data, this embodiment comprehensively considers the impact of environmental differences on crop trait performance during model training. This allows the model to automatically learn the relationship between genotype and environment and their interaction on phenotype, thereby effectively improving the model's accuracy, generalization, and robustness in cross-environment prediction. This enables efficient and accurate prediction of crop phenotypes, significantly improving the efficiency and success rate of intelligent breeding.
[0047] Through steps S210 to S240, the model can not only identify the variation patterns of key traits such as yield and plant height of each recombinant inbred line under different environments, but also quantify the crop's sensitivity to environmental responses through plasticity characteristic parameters. This method enables the prediction model to maintain high phenotypic prediction accuracy when facing unknown or changing environmental conditions. Specifically, the model can deeply correlate the genetic characteristics reflected in the genotype matrix with the environmental response patterns represented by the plasticity characteristics, thereby enabling rapid prediction of the phenotypic characteristics of offspring under target environments after inputting new parental genotypes and a small amount of environmental information.
[0048] In some embodiments, such as Figure 3 As shown, step S210, which determines the genotype matrix of multiple lines of recombinant inbred lines, may include steps S211 and S212.
[0049] Step S211: Determine all genotyping data for multiple lines of the recombinant inbred line. Genotyping of multiple lines in the recombinant inbred line can be performed using a microarray, thereby determining all genotyping data for each line. The genotyping data may include multiple genetic markers, i.e., SNP markers. Specifically, genotyping of multiple lines of the recombinant inbred line can be performed using a wheat 50K SNP microarray, thus obtaining all genotyping data for multiple lines of the recombinant inbred line more efficiently and accurately. Furthermore, genotyping of the parents can also be performed simultaneously, thereby determining the genotype matrix of the parents and multiple lines of the recombinant inbred line.
[0050] Step S212 involves screening genetic markers from multiple lines of recombinant inbred lines based on all genotyping data to determine the genotype matrix. Since some low-quality genetic markers exist in the genotyping data, to improve the quality of the genotype matrix and the accuracy of the prediction model, genetic markers in all genotyping data can be screened to determine a genotype matrix containing high-quality genetic markers. Step S212 allows for quality control of all genotyping data, removing low-quality SNP markers. Specifically, SNP markers can be screened based on their deletion rate in multiple lines of recombinant inbred lines, minor allele frequencies, and chromosomal location. SNP markers with a deletion rate greater than 10% in multiple lines of recombinant inbred lines can be deleted. Since multiple lines of recombinant inbred lines are permanently homozygous populations obtained through multiple generations of self-pollination after crossing two parents, genotypes differ among these lines. Therefore, SNP markers with a deletion rate greater than 10% in multiple lines of recombinant inbred lines can be considered low-quality SNPs due to their high deletion rate. These SNP markers can be deleted to improve the data quality of the genotype matrix. Similarly, SNP markers with a minor allele frequency less than 0.05 can be deleted. Minor allele frequency represents the proportion of the second most common allele in multiple lines of recombinant inbred lines. When the minor allele frequency of an SNP marker is less than 0.05, it can be considered to provide limited information. Therefore, SNP markers with a minor allele frequency less than 0.05 can be deleted to reduce the amount of invalid information in the genotype matrix. Furthermore, SNP markers with unclear chromosomal locations can also be deleted. In the process of genotyping multiple strains and obtaining genotyping data, some SNP markers in the genotyping data cannot be accurately located on a specific chromosome, or their positions on the chromosome are ambiguous or not unique. Deleting these SNP markers with unclear chromosome locations can effectively avoid errors in subsequent SNP marker analysis and model training, improve the accuracy of the genotype matrix data, effectively reduce the amount of invalid data, and thus improve the efficiency of subsequent data calculation and model training.
[0051] According to this embodiment, by genotyping multiple lines of recombinant inbred lines and their parents, high-density SNP marker data covering the entire genome can be obtained efficiently and in one go, ensuring that the genotypic information of each line is comprehensive and has high resolution. Through rigorous quality control and screening of all genotyping data, SNP markers with a deletion rate greater than 10%, a minor allele frequency less than 0.05, and unclear chromosomal locations are deleted, significantly improving the effective information ratio and analytical accuracy of the genotype matrix. This process effectively removes noisy markers and redundant information, resulting in higher accuracy and stability of the genotype matrix. This embodiment effectively improves the accuracy and usability of genetic information in the genotype matrix, significantly increasing the training efficiency and prediction reliability of subsequent phenotypic prediction models.
[0052] In some embodiments, such as Figure 4 As shown, step S220, based on multiple environments, determines the phenotypic data of multiple lines of the recombinant inbred line in each environment, and may include steps S221 to S223.
[0053] Step S221 involves planting multiple varieties of the recombinant inbred lines of the parents multiple times under various environments. Based on the parents, the parents can be hybridized to obtain a population of recombinant inbred lines derived from the hybridization. This population can include multiple different varieties. Based on the obtained multiple varieties, multiple varieties of the recombinant inbred lines can be planted under various environments, and for each environment, multiple varieties of the recombinant inbred lines can be planted multiple times, thereby effectively reducing phenotypic data errors in subsequent data processing. A Latin square design can be used to arrange the planting of multiple varieties of the recombinant inbred lines. Specifically, for a population of recombinant inbred lines derived from a hybridization of the two parents, Zhongmai 578 and Jimai 22, 262 F1 lines can be selected. 2:7 Using the progeny lines as samples, the aforementioned 262 recombinant inbred lines can be planted in eight different environments. These eight environments may include locations such as Shijiazhuang, Dezhou, and Xinxiang, and planting can be carried out at these locations in different years. In each of the eight environments, 262 lines can be planted, and each line can be planted three times within the same environment to obtain more phenotypic data, thereby reducing the phenotypic error of the 262 lines.
[0054] Step S222 involves determining the trait data for each variety after each planting under multiple environments. Based on the multiple crop varieties planted, the trait data for each variety can be determined after the crop matures. The trait data for each variety can include specific data on the yield of that variety under the current environment, such as grain yield, number of grains per ear, thousand-grain weight, number of ears per unit area, or plant height.
[0055] Step S223: Based on the trait data, the average value of the trait data determined from multiple plantings of each strain under the same environment is determined as the phenotypic data of the current strain in the current environment. Since each strain can be planted multiple times under the same environment, the average value of multiple trait data determined from multiple plantings under the same environment can be taken based on all the trait data of each strain. This average value is then used as the phenotypic data of the strain in the current environment, effectively reducing the error in monitoring the phenotypic data of each strain, improving the accuracy of the phenotypic data of each strain, and resulting in a prediction model trained based on the phenotypic data of the aforementioned multiple strains having higher prediction accuracy.
[0056] According to the method provided in this embodiment, by repeatedly planting multiple lines of recombinant inbred lines and scientifically arranging the planting locations of different lines using a Latin square design, systematic errors caused by environmental heterogeneity can be effectively reduced. This design, while ensuring randomness and balance, can minimize the interference of non-genetic factors such as differences in soil fertility and microclimate changes on phenotypes, thereby improving the comparability and representativeness of phenotypic data. By accurately measuring traits such as grain yield, number of grains per ear, thousand-grain weight, number of ears per unit area, and plant height for each line under different environments, phenotypic changes in crops under environmental variations can be fully captured, providing a complete data foundation for phenotypic analysis. Furthermore, by averaging the phenotypic data obtained from multiple plantings under the same environment, noise interference in phenotypic measurements can be effectively reduced. This results in higher stability and reliability of phenotypic data for each line in each environment. Furthermore, it effectively enhances the reliability and scientific rigor of the prediction model, providing high-precision data support for breeding.
[0057] In some embodiments, such as Figure 5 As shown, step S230, which determines the plasticity characteristics of each line based on phenotypic data of multiple lines of recombinant inbred lines in each environment, may include steps S231 to S233.
[0058] Step S231: Determine the population mean of phenotypic data for multiple lines of the recombinant inbred line under each environment. Based on the phenotypic data of multiple lines of the recombinant inbred line under each environment, the population mean can be determined. The population mean can be the average of the phenotypic data of multiple lines of the recombinant inbred line under the same environment. Specifically, the grain yield of multiple lines in the recombinant coordinate system can be extracted under the same environment. The grain yields determined by planting all the above lines under this environment are added together and the average is taken, thereby determining the population mean of grain yield.
[0059] Step S232: Using the population mean as the independent variable and the phenotypic data of multiple lines of the recombinant inbred line in each environment as the dependent variable, fit a linear regression model for the crop phenotype of each line. Since the population mean can represent the average phenotype of multiple lines under the same environment, it can be used as the independent variable, and the phenotypic data of multiple lines in each environment can be used as the dependent variable to fit a linear regression model. Because there are multiple lines, it is necessary to fit a linear regression model between the phenotypic data of each line and the corresponding population mean. Specifically, for the i-th line in the recombinant inbred line population, the model can be fitted as follows: The slope of the linear regression model is... The intercept of the linear regression model is .in, It can be a very small number, through This avoids the intercept of the linear regression model being zero. Specifically, It can be 10 -10 .
[0060] Step S233: Based on the linear regression model of crop phenotypes, determine the plasticity characteristics, where the plasticity characteristics are feature vectors composed of the slopes of the linear regression models of all crop phenotypes. Based on the multiple linear regression models determined in step S232, the plasticity characteristics of each line can be determined. Specifically, for the i-th line in the recombinant inbred line population, the model can be fitted as follows: The slope of the linear regression model is... The intercept of the linear regression model is Among them, the slope can be... This serves as the phenotypic plasticity index for the strain. Based on the phenotypic plasticity indices of the aforementioned multiple strains, the plasticity characteristics of the recombinant inbred line can be determined. Specifically, the plasticity characteristics can be a feature vector jointly composed of the plasticity indices of multiple strains. Through the plasticity characteristics, the relationship between the phenotype of each strain of the recombinant inbred line and the environment can be determined.
[0061] According to this embodiment, by determining the mean phenotypic data of the recombinant inbred line population under each environment, the relationship between the overall phenotypic characteristics of the recombinant inbred line population and the environment can be effectively reflected, providing a standardized basis of independent variables for subsequent linear regression. This processing method can effectively eliminate systematic errors between environments. This embodiment uses the population mean as the independent variable and the actual phenotypic data of each line as the dependent variable for linear regression fitting, which can intuitively determine the trend of individual line phenotypic changes with the environment. Through regression analysis provided in this embodiment, the phenotypic changes of each recombinant inbred line under different environments can be accurately quantified, enabling the prediction model to consider both genotype effects and environmental plasticity effects, thereby achieving higher accuracy and stability in cross-environment prediction of crop phenotypic characteristics.
[0062] In some embodiments, such as Figure 6 As shown, step S240, which involves training the model based on plasticity features and the genotype matrix to determine the completed prediction model, may include steps S241 and S242.
[0063] Step S241: Construct a fusion feature matrix based on plasticity features and the genotype matrix. During model training, it is necessary to train the model based on plasticity features and the genotype matrix. Therefore, the plasticity features and the genotype matrix can be fused first to construct a fusion feature matrix, which can then be used for subsequent model training.
[0064] Step S242: Train the model based on the fused feature matrix to determine the completed prediction model. The fused feature matrix can be used as a sample for model training. Since the fused feature matrix includes information related to plasticity features and the genotype matrix, the prediction model trained based on the fused feature matrix can determine the phenotypic prediction value of crops based on genotype and environmental factors. This enables cross-environment phenotypic prediction while maintaining high prediction accuracy and stability. Furthermore, the model training method provided in this embodiment can maintain high prediction accuracy even with a small number of training samples.
[0065] According to this embodiment, a fusion feature matrix can be constructed based on plasticity characteristics and the genotype matrix, thereby integrating relevant data on genotype, environment, and phenotype. This fusion feature matrix comprehensively reflects the genetic characteristics of each line in a recombinant inbred line population and their differences in plasticity under different environments. Compared to traditional models trained solely on genotype or environmental characteristics, this embodiment effectively improves the model's feature utilization efficiency and learning depth, providing richer representational information for subsequent predictions. This embodiment achieves synergistic modeling of gene effects and environmental effects, enabling the trained prediction model to stably and accurately predict crop phenotypes under different environments, significantly improving the reliability of cross-environment phenotype prediction and the efficiency of the breeding process.
[0066] In some embodiments, step S242, training the model based on the fused feature matrix to determine the completed prediction model, may include the following steps: The fused feature matrix is input into multiple candidate models for training. These candidate models include at least two of the following: a minimum absolute shrinkage and selection algorithm model, a ridge regression model, a partial least squares regression model, a support vector regression model, and a random forest model. The same fused feature matrix can be input into multiple different candidate models to train different candidate models respectively. Based on the trained candidate models, the model with the highest prediction accuracy and stability among the candidate models is selected as the prediction model. The candidate models include, but are not limited to, at least two of the following models: a minimum absolute shrinkage and selection algorithm model, a ridge regression model, a partial least squares regression model, a support vector regression model, and a random forest model.
[0067] Validate multiple trained candidate models using 10-fold cross-validation and / or leave-one-out-of-environment methods to determine the validation results. The validation results include the prediction correlation coefficient and / or root mean square error. For the trained candidate models, their prediction accuracy can be validated to determine the validation results. Specifically, the 10-fold cross-validation method can be used to evaluate the performance of the candidate models and determine the validation results. Alternatively, the leave-one-out-of-environment method can be used to evaluate the performance of the candidate models and determine the validation results. Specifically, evaluating candidate models using the leave-one-out-of-environment method allows for a more accurate assessment of the candidate models' cross-environment predictive capabilities. Validating multiple trained candidate models using 10-fold cross-validation and / or leave-one-out-of-environment methods determines the validation results of the candidate models. The validation results can include the prediction correlation coefficient and the root mean square error, where the prediction correlation coefficient is the correlation coefficient between the actual observed values of the phenotype and the predicted values determined by the candidate models.
[0068] Based on the validation results, a candidate model is selected as the prediction model for completion of training. According to the validation results determined in step S2422, the candidate model with a higher prediction correlation coefficient or a smaller root mean square error can be selected as the prediction model for completion of training, based on the prediction correlation coefficient or root mean square error of the candidate model's predicted crop yield. This allows for the selection of candidate models with higher prediction accuracy through validation results. Specifically, validation results can be determined separately for each phenotypic data, thereby determining the prediction correlation coefficient of each candidate model for different phenotypic data, such as the prediction correlation coefficient for grain yield, the prediction correlation coefficient for thousand-grain weight, etc.
[0069] According to this embodiment, by inputting the fused feature matrix into multiple different candidate models and training them, a candidate model with higher prediction accuracy and efficiency can be selected as the prediction model based on breeding needs. The prediction model is then further optimized, and the final determined prediction model maintains high prediction accuracy and generalization stability in multi-trait prediction. This embodiment enables the final determined prediction model to not only have high-precision prediction capabilities in a single environment but also higher accuracy and reliability in cross-environment prediction tasks. Compared to traditional single-model training schemes, this effectively improves the accuracy of crop phenotypic prediction and the scientific rigor of breeding decisions based on the prediction model determined in this embodiment.
[0070] In some embodiments, a population of recombinant inbred lines derived from a cross between two parents, Zhongmai 578 and Jimai 22, can be used, selecting 262 F1 lines from it. 2:7 The F1 strain was planted using a Latin square design in eight different environments, with each strain being planted three times in the same environment. These eight environments could include Shijiazhuang, Dezhou, and Xinxiang in 2021 and 2022. In Xinxiang, irrigation and water-saving treatments could be implemented. The F1 strains were then planted... 2:7Phenotypic data were collected from the lines to determine the phenotypic data of multiple recombinant inbred lines. This phenotypic data included grain yield, thousand-grain weight, number of grains per spike, number of spikes per unit area, and plant height. To ensure the accuracy of the phenotypic data and reduce the impact of errors, the phenotypic data obtained from three repeated plantings of the same line under the same environment were averaged to improve the accuracy of subsequent analyses. A wheat 50K SNP chip was used to genotype the parents and 262 recombinant inbred lines, identifying all SNP markers for multiple recombinant inbred lines. Subsequently, SNP marker quality control was performed, removing low-quality SNP markers. Specifically, based on all SNP markers, SNP markers with a deletion rate greater than 10% in the population, SNP markers with a minor allele frequency less than 0.05, and markers with unclear chromosome locations were removed, ultimately obtaining 1503 high-quality SNP markers. A 262×1503 genotype matrix was then constructed based on these SNP markers.
[0071] Subsequently, the plasticity characteristics of each strain can be extracted. First, for each environment, the phenotypic mean of each trait for all strains within the recombinant inbred line population can be calculated as the population mean for that environment. Taking plant height as an example, the population mean of the environment can be used as the independent variable, and the phenotypic value of each strain in each environment, i.e., plant height, can be used as the dependent variable to determine a linear regression equation for each strain. For the i-th strain in the recombinant inbred line population, the model can be fitted as follows: The slope of the linear regression model is... The slope can be Defined as the phenotypic plasticity index of the strain, and finally, based on the above process, linear regression models were fitted to 262 strains respectively to determine 262 plasticity characteristic indices, which constitute 262×1 plasticity characteristic values.
[0072] The genotype matrix M (262×1503) and the plasticity eigenvalue β (262×1) can be fused to construct a fused feature matrix. Five candidate models can be selected: minimum absolute shrinkage and selection algorithm, ridge regression, partial least squares regression, support vector regression, and random forest. The fusion feature matrix is input into each of these models to train them. After training, the candidate models are evaluated using ten-fold cross-validation and leave-one-out methods. By determining the prediction correlation coefficient and root mean square error of each candidate model, the accuracy and precision of the predictions are assessed. Based on the breeding requirements, a suitable candidate model is selected as the prediction model.
[0073] The predictive correlation coefficients of the same candidate model differ for different phenotypes, indicating that the same candidate model does not perform equally well in detecting different crop traits. Taking a predictive model as an example... Figures 7 to 11 The image shows a scatter plot of predicted values for grain yield, thousand-grain weight, number of grains per ear, number of ears per unit area, and plant height under unknown environments, determined by a prediction model based on phenotypic data of individuals under seven known environments. Figures 7 to 11 As shown, the deviation between the predicted and actual phenotypic data can be determined. Therefore, the accuracy of the model's predictions for each phenotypic data can be determined. Figures 7 to 11 The closer the scatter points are to the dashed line in the graph, the higher the accuracy and precision of the predicted phenotypic data determined by the prediction model. Figures 12 to 16 As shown, it is possible to determine the predicted relevance (i.e., the change in the predicted relevance coefficient) and the root mean square error between the predicted value and the actual phenotypic data under different proportions of the test set to the entire training dataset. Based on... Figures 12 to 16 As shown, for the five phenotypic data—grain yield, thousand-grain weight, number of grains per ear, number of ears per unit area, and plant height—the highest correlation coefficient and the lowest root mean square error were observed when 10% of the data was used as the test set and the remaining 90% as the training set. Therefore, it can be determined that the larger the training population and the smaller the test set, the higher the prediction accuracy and precision of the predictive model for each phenotypic data. The predictive model was evaluated using a ten-fold cross-validation method, such as... Figure 17 As shown, the correlation coefficients between observed and predicted values of grain yield range from 0.268 to 0.659; the prediction accuracy for 1000-grain weight and plant height is relatively high, with correlation coefficients exceeding 0.763 and 0.886, respectively; the prediction performance for grain number per ear is moderate, with correlation coefficients ranging from 0.498 to 0.778; while the prediction accuracy for ear number, which has the lowest heritability, varies greatly, with correlation coefficients between 0.084 and 0.513. Therefore, based on the evaluation results of each candidate model, a suitable candidate model can be selected according to the breeding needs. Specifically, the suitable model may differ for different categories of phenotypic data. For example, for ear number and grain weight, the minimum absolute shrinkage and selection algorithm model has better prediction accuracy; for ear number per unit area and plant height, the ridge regression model has higher prediction accuracy; and for grain yield, the partial least squares regression model has higher prediction accuracy.
[0074] Finally, the established prediction model can be tested. For unknown environments, a leave-one-environment approach can be used, employing seven environments as the training set to predict plant height in the remaining unknown environment. The prediction correlation coefficient ranges from 0.763 to 0.886. For unknown crop lines, family-based cross-validation can be used. 10% of the lines are randomly selected as the validation set, and the remaining 90% are used for training. The average phenotype of these unknown lines across all environments is predicted, repeated 50 times. The average prediction correlation coefficient is greater than 0.739. When both the crop line and environment are unknown, a complete environment and 10% of the lines can be selected simultaneously, and the remaining data can be used for training. In this scenario, the prediction correlation coefficient for plant height still remains greater than 0.722.
[0075] Furthermore, genome-wide association analysis (GWAS) can be performed using the GAPIT model. A significantly associated SNP locus was detected at 485.7 Mb on wheat chromosome 5B. Competitive allele-specific polymerase chain reaction (KASP) markers were developed for this locus and validated in an independent natural population containing 166 varieties. The results showed that varieties carrying superior alleles exhibited significantly lower plant height variation across different environments, i.e., lower plasticity, with a variation range of less than 0.01, demonstrating the application value of this locus in regulating phenotypic stability and its potential use in marker-assisted selection.
[0076] Based on the same inventive concept, this disclosure also provides a crop phenotypic prediction device for performing the crop phenotypic prediction method as described in any of the foregoing embodiments, wherein the crop phenotypic prediction device may include: a genotype determination module, a phenotypic data processing module, and a crop phenotypic prediction module.
[0077] The genotyping module is used to determine the genotypes of the parents. The genotyping module can execute step S110 to determine the genotypes of the parents. The genotyping module can identify multiple SNP markers in the parents.
[0078] The phenotypic data processing module is used to determine the phenotypic data of the parents. The phenotypic data processing module can execute step S120 to obtain the phenotypic performance data of the parents under the target environment, such as yield, thousand-grain weight, number of grains per ear and other trait information. It can also perform standardization processing and outlier removal on the raw phenotypic data to improve the accuracy of the data and reduce errors.
[0079] The crop phenotypic prediction module is used to determine the predicted phenotypic value of a crop based on the genotype and phenotypic data of its parents. Based on the parental genotype determined by the genotype determination module and the phenotypic data of the parents after planting in a specific environment determined by the phenotypic data processing module, the module uses a pre-trained prediction model to output the predicted phenotypic value of the corresponding crop under the target environment.
[0080] The crop phenotypic prediction device provided in this embodiment can efficiently acquire and store genotypic information from different parents, providing stable genetic basis data for subsequent predictions. The phenotypic data processing module can standardize, fuse, and normalize phenotypic observation data under multiple environments, reducing bias caused by inconsistent data distribution under different environmental conditions, thereby improving the robustness of the input data. The crop phenotypic prediction module can automatically output phenotypic prediction results for the target crop under different environments based on the input of parental genotypic and phenotypic data by calling the trained prediction model. Compared to traditional prediction schemes that rely on a single environment or linear modeling, this device can comprehensively consider the influence of genotype and environment, effectively improving cross-environment phenotypic prediction performance.
[0081] Based on the same inventive concept, this disclosure also provides a computer-readable storage medium storing a program for performing the crop phenotypic prediction method as described in any of the foregoing embodiments. One embodiment of this disclosure provides an electronic device. The electronic device includes one or more processors, memory, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components are interconnected via different buses and can be mounted on a common motherboard or otherwise mounted as needed. The processor can process instructions executed within the electronic device, including instructions stored in or on memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple electronic devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). The figures show an example of a single processor.
[0082] The processor can be a central processing unit, a network processor, or a combination thereof. The processor may further include hardware chips. These hardware chips can be application-specific integrated circuits (ASICs), programmable logic devices (PLDs), or combinations thereof. The programmable logic devices can be complex programmable logic devices (CLPs), field-programmable gate arrays (FPGAs), general-purpose array logic (GDAs), or any combination thereof.
[0083] The memory stores instructions executable by at least one processor to cause the at least one processor to perform the crop phenotypic prediction method shown in the above embodiments.
[0084] The memory may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device. Furthermore, the memory may include high-speed random access memory (RAM) and non-transient memory, such as at least one disk storage device, flash memory, or other non-transient solid-state storage device. In some alternative embodiments, the memory may include memory remotely located relative to the processor, which can be connected to the electronic device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks (LANs), mobile communication networks, and combinations thereof. The memory may include volatile memory, such as random access memory (RAM); it may also include non-volatile memory, such as flash memory, hard disks, or solid-state drives (SSDs); and it may include combinations of the above types of memory. The electronic device also includes input devices and output devices. The processor, memory, input devices, and output devices may be connected via a bus or other means.
[0085] Input devices can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the electronic device, such as touchscreens, keypads, mice, trackpads, touchpads, joysticks, one or more mouse buttons, trackballs, joysticks, etc. Output devices may include display devices, auxiliary lighting devices (e.g., LEDs), and haptic feedback devices (e.g., vibration motors). The aforementioned display devices include, but are not limited to, liquid crystal displays, light-emitting diodes, displays, and plasma displays. In some alternative embodiments, the display device may be a touchscreen.
[0086] This application uses specific terms to describe embodiments of the application. For example, "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of the application. Therefore, it should be emphasized and noted that "an embodiment," "one embodiment," or "an alternative embodiment" mentioned twice or more in different locations in this specification do not necessarily refer to the same embodiment. Furthermore, certain features, structures, or characteristics in one or more embodiments of the application can be appropriately combined. Similarly, it should be noted that, in order to simplify the description of the disclosure in this application and thus aid in the understanding of one or more embodiments, the foregoing description of the embodiments of the application sometimes combines multiple features into one embodiment, drawing, or description thereof. In fact, the features of an embodiment are fewer than all the features of a single embodiment disclosed above.
[0087] The basic concepts have been described above. Obviously, for those skilled in the art, the above disclosure is merely illustrative and does not constitute a limitation of this application. Although not explicitly stated herein, those skilled in the art may make various modifications, improvements, and corrections to this application. Such modifications, improvements, and corrections are suggested in this application, and therefore remain within the spirit and scope of the embodiments of this application.
Claims
1. A method for predicting crop phenotypes, characterized in that, A method for predicting crop phenotypes across environments, comprising: Determine the genotypes of the parents; Determine the phenotypic data of the parents under the current environment; The genotype and phenotypic data of the parent lines are input into the trained prediction model to determine the predicted value of the crop's phenotypic phenotype in the current environment. The prediction model is trained based on the genotypes of multiple lines of the recombinant inbred line of the parent lines and the phenotypic data of multiple lines of the recombinant inbred line in multiple environments. The prediction model is trained using the following method: Determine the genotype matrix of multiple lines of the recombinant inbred line; Based on multiple environments, phenotypic data of multiple lines of the recombinant inbred line in each environment are determined, wherein the phenotypic data includes at least one of the following: grain yield, thousand-grain weight, number of grains per ear, number of ears per unit area, and plant height. Based on the phenotypic data of multiple lines of the recombinant inbred line in each environment, the plasticity characteristics of each line are determined; Based on the plasticity features and the genotype matrix, the model is trained to determine the completed prediction model; The step of determining the phenotypic data of multiple lines of the recombinant inbred line in each environment based on multiple environments includes: Multiple varieties of recombinant inbred lines of the parent were obtained by multiple plantings of the parent in multiple environments, wherein the multiple varieties of the recombinant inbred lines were derived from parental hybridization; The trait data for each strain were determined after each planting under multiple environments; Based on the trait data, the average value of the trait data determined by multiple plantings of each strain under the same environment is determined as the phenotypic data of the current strain in the current environment.
2. The crop phenotypic prediction method according to claim 1, characterized in that, Determining the phenotypic data of the parent in the current environment includes: The parent plants are cultivated under the current conditions; To determine the phenotypic data of the parents grown under the current environment.
3. The crop phenotypic prediction method according to claim 1, characterized in that, Determining the genotype matrix of multiple lines of the recombinant inbred line includes: Determine the complete genotyping data of multiple lines of the recombinant inbred line; Based on all the genotyping data, genetic markers of multiple lines of the recombinant inbred line were screened to determine the genotype matrix.
4. The crop phenotypic prediction method according to claim 1, characterized in that, Based on the phenotypic data of multiple lines of the recombinant inbred line in each environment, the plasticity characteristics of each line are determined, including: Determine the population mean of the phenotypic data of multiple lines of the recombinant inbred line under each environment; Using the population mean as the independent variable and the phenotypic data of multiple lines of the recombinant inbred line in each environment as the dependent variable, a linear regression model of the crop phenotype of each line is fitted. The plasticity feature is determined based on the linear regression model of the crop phenotype, wherein the plasticity feature is a feature vector composed of the slopes of the linear regression models of all the crop phenotypes.
5. The crop phenotypic prediction method according to claim 1, characterized in that, The step of training the model based on the plasticity feature and the genotype matrix to determine the completed prediction model includes: Based on the plasticity feature and the genotype matrix, a fusion feature matrix is constructed; The model is trained based on the fused feature matrix to determine the completed prediction model.
6. The crop phenotypic prediction method according to claim 5, characterized in that, The step of training the model based on the fused feature matrix to determine the completed prediction model includes: The fused feature matrix is input into multiple candidate models for training. The multiple candidate models include at least two of the following: minimum absolute shrinkage and selection algorithm model, ridge regression model, partial least squares regression model, support vector regression model, and random forest model. The trained candidate models are validated using a 10-fold crossover and / or leave-one-out environment method to determine the validation results, wherein the validation results include: prediction correlation coefficient and / or root mean square error; Based on the verification results, one of the candidate models is selected as the prediction model that has completed training.
7. A crop phenotypic prediction device, characterized in that, For performing the crop phenotypic prediction method as described in any one of claims 1-6, the crop phenotypic prediction device comprises: A genotype determination module is used to determine the genotype of the parents; The phenotypic data processing module is used to determine the phenotypic data of the parents; The crop phenotypic prediction module is used to determine the predicted phenotypic value of the crop based on the genotype and phenotypic data of the parents.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program for performing the crop phenotypic prediction method as described in any one of claims 1-6.
Citation Information
Patent Citations
High-efficiency multi-environment whole genome prediction method and system based on deep learning
CN118098344A
Genome prediction method and device based on genotype and environment interaction heterogeneous graph
CN118471327A
Phenotypic character prediction method and device, storage medium and electronic equipment
CN118506864A
Crop phenological period prediction method, system and device based on whole genome prediction
CN120412711A