Method and system for predicting crop phenotype utilizing crop genotype data augmentation based on genotype position information
By training crop phenotype prediction models with a diverse and realistic dataset augmented using genotype location information and SNP data, the method enhances prediction accuracy and reliability, addressing the challenges of collecting high-dimensional genotype data for crop improvement.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- KOREA ELECTRONICS TECH INST
- Filing Date
- 2025-11-28
- Publication Date
- 2026-06-18
AI Technical Summary
Current machine learning models for predicting crop phenotypes face challenges due to the difficulty in collecting large-scale, diverse, and realistic genotype data, which is complex and costly, and struggle to effectively learn local correlations in high-dimensional data.
A method and system for predicting crop phenotypes by training a model with a diverse, realistic training dataset constructed by augmenting crop genotype data using genotype location information and SNP data, incorporating SNP positional dependency to reflect actual biological characteristics.
Improves the accuracy and reliability of crop phenotype prediction models, enabling efficient breeding for traits like disaster resistance and high yield, contributing to agricultural productivity and food security.
Smart Images

Figure KR2025020061_18062026_PF_FP_ABST
Abstract
Description
Method and System for Predicting Crop Phenotypes Using Crop Genotype Data Augmentation Based on Genotype Location Information
[0001] The present invention relates to machine learning application technology in the fields of agriculture and biotechnology, and more specifically, to a method for predicting crop phenotype data from crop genome data based on machine learning.
[0002] Machine learning technology is gaining increasing importance in fields such as agriculture and biotechnology. In particular, models that predict phenotypes based on the genetic characteristics of crops can contribute to accelerating the speed of crop improvement. However, learning genotype data is difficult due to a lack of samples and high-dimensional features, and existing models have limitations in effectively learning the local correlations of genotype data.
[0003] In other words, the collection and analysis of current genotype data are highly complex and costly, and securing large-scale data is difficult due to their high-dimensional nature. In particular, while digital breeding aimed at improving crop traits requires accurate and large amounts of genotype data, there are limitations in collecting only the desired data due to environmental factors and the complexity of gene expression.
[0004] Therefore, it is necessary to explore ways to improve the accuracy of machine learning-based crop phenotype prediction models by securing diverse, realistic, and sufficient training datasets.
[0005] The present invention has been devised to solve the aforementioned problems, and the objective of the present invention is to provide a method for predicting crop phenotypes by training a crop phenotype prediction model with a diverse, realistic, and sufficient training dataset constructed by augmenting crop genotype data using genotype location information and SNP (Single Nucleotide Polymorphism) data, and using the same.
[0006] A method for training a crop phenotype prediction model according to an embodiment of the present invention for achieving the above objective comprises: a step of constructing a training dataset of a crop phenotype prediction model, which is a machine learning model that predicts phenotype data composed of crop traits from crop genome data; a step of inputting the genome data of the constructed training dataset into a crop phenotype prediction model to be trained, thereby predicting phenotype data from the crop genome data; and a step of updating the parameters of the crop phenotype prediction model so as to reduce the difference between the predicted phenotype data and the phenotype data of the constructed training dataset.
[0007] The construction step may include: a step of acquiring genomic data including genotype data for a crop and location data for each genotype; a step of acquiring phenotypic data for a crop; a step of preprocessing the acquired genomic data and phenotypic data; and a step of combining the preprocessed genomic data and phenotypic data to construct a training dataset for a crop phenotypic prediction model, which is a machine learning model that predicts crop phenotypic data from crop genomic data.
[0008] The construction step may further include: a step of extracting embedding vectors from preprocessed genomic data; a step of generating new genotype data different from the genotype data of the preprocessed genomic data from the extracted embedding vectors; and a step of adding a new training dataset by combining the generated new genotype data with the location data and preprocessed phenotype data of the preprocessed genomic data.
[0009] Genotype data can be SNP (Single Nucleotide Polymorphism) data.
[0010] The preprocessing step may involve converting the nucleotide sequences of the genotype data into integers, by mapping the allele combinations of the genotype data to their respective integer values.
[0011] The preprocessing step may involve categorizing each characteristic data of the phenotypic data by dividing it into multiple intervals.
[0012] The extraction step may include: a step of extracting an embedding vector from genotype data of genomic data; a step of extracting an embedding vector from location data of genomic data; and a step of combining the extracted embedding vectors into a single embedding vector.
[0013] The generation step may include: inputting the combined embedding vector into a convolution layer to extract features; and inputting the extracted features into an FCL to average pool the features to generate new genotype data.
[0014] The extraction step is performed using an embedding network that extracts embedding vectors from preprocessed genomic data, and the generation step is performed using a genotype generation network that generates new genotype data from the extracted embedding vectors; the parameters of the embedding network and the genotype generation network can be updated and trained in a direction that reduces the loss between the new genotype data and the genotype data of the preprocessed genomic data.
[0015] According to another aspect of the present invention, a crop phenotype prediction model learning system is provided, comprising: a storage unit in which a training dataset of a crop phenotype prediction model is constructed, which is a machine learning model that predicts phenotype data composed of traits of a crop from genomic data of a crop; and a processor that inputs the genomic data of the constructed training dataset into a crop phenotype prediction model to be trained, predicts phenotype data from the crop genomic data, and updates the parameters of the crop phenotype prediction model so as to reduce the difference between the predicted phenotype data and the phenotype data of the constructed training dataset.
[0016] According to another aspect of the present invention, a method for predicting a crop phenotype is provided, comprising: a step of acquiring genomic data of a crop; a step of inputting the acquired genomic data into a crop phenotype prediction model, which is a machine learning model that predicts phenotype data composed of traits of a crop from the genomic data of a crop, to predict phenotype data from the crop genomic data; and a step of outputting the predicted phenotype data. The crop phenotype prediction model is characterized by being trained by: a step of constructing a training dataset; a step of inputting the genomic data of the constructed training dataset into a crop phenotype prediction model to be trained to predict phenotype data from the crop genomic data; and a step of updating the parameters of the crop phenotype prediction model so as to reduce the difference between the predicted phenotype data and the phenotype data of the constructed training dataset.
[0017] According to another aspect of the present invention, a crop phenotype prediction system is provided, comprising: an acquisition unit for acquiring genomic data of a crop; a processor for predicting phenotype data from crop genomic data by inputting the acquired genomic data into a crop phenotype prediction model, which is a machine learning model for predicting phenotype data composed of traits of a crop from crop genomic data; and an output unit for outputting the predicted phenotype data. The crop phenotype prediction model is characterized by constructing a training dataset, inputting the genomic data of the constructed training dataset into a crop phenotype prediction model to learn the phenotype data from crop genomic data, and learning the parameters of the crop phenotype prediction model by updating such that the difference between the predicted phenotype data and the phenotype data of the constructed training dataset is reduced.
[0018] As described above, according to the embodiments of the present invention, by training a crop phenotype prediction model with a diverse, realistic, and sufficient training dataset constructed by augmenting crop genotype data using genotype location information and SNP data, the prediction accuracy of the crop phenotype prediction model can be improved.
[0019] In addition, according to embodiments of the present invention, by reflecting SNP location dependency when augmenting genotype data so that gene mutations occur only at realistically significant locations, it is possible to further contribute to improving the accuracy of crop phenotype prediction models through genotype data augmentation that corresponds to actual biological characteristics.
[0020] And according to the embodiments of the present invention, the prediction accuracy and reliability of a crop phenotype prediction model can be increased together by using a training dataset constructed through data augmentation that has statistical characteristics similar to existing genotype data.
[0021] Furthermore, according to the embodiments of the present invention, crop phenotype prediction models with high reliability and accuracy enable the prediction of crop phenotypes in various environments, thereby ultimately opening the way to more efficiently breed crops with specific traits such as disaster resistance, disease resistance, and high yield, and contributing to the improvement of agricultural productivity and the strengthening of food security by developing varieties resistant to environmental changes.
[0022] FIG. 1 is a flowchart of a crop phenotype prediction model learning method according to an embodiment of the present invention,
[0023] FIG. 2 is a flowchart of a method for constructing a training dataset for a crop phenotype prediction model according to an embodiment of the present invention,
[0024] Figures 3 and 4 are diagrams illustrating the structure of a genotype data augmentation model.
[0025] FIG. 5 is a flowchart of a genetic data augmentation model learning method according to another embodiment of the present invention,
[0026] FIG. 6 is a flowchart of a crop phenotype prediction method according to another embodiment of the present invention,
[0027] FIG. 7 shows experimental results using a crop phenotype prediction model trained according to an embodiment of the present invention,
[0028] FIG. 8 is a hardware configuration diagram of a crop phenotype prediction system according to another embodiment of the present invention.
[0029] The present invention will be described in more detail below with reference to the drawings.
[0030] In an embodiment of the present invention, a method and system for predicting crop phenotypes utilizing crop genotype data augmentation based on genotype location information are presented. This technology improves the prediction accuracy of a crop phenotype prediction model by training the model with a diverse, realistic, and sufficient training dataset constructed by augmenting crop genotype data using genotype location information and SNP data.
[0031] In particular, when augmenting genotype data, SNP positional dependency is reflected so that gene mutations occur only at realistically significant locations, thereby performing genotype data augmentation that corresponds to actual biological characteristics.
[0032] FIG. 1 is a diagram illustrating the flow of a crop phenotype prediction model learning method according to an embodiment of the present invention. A crop phenotype prediction model is a machine learning model that predicts phenotype data composed of crop traits from crop genome data.
[0033] To train a crop phenotype prediction model, a training dataset for the crop phenotype prediction model is first constructed (S110). The training dataset combines the crop genome data, which is the input to the crop phenotype prediction model, and the crop phenotype data, which is the correct answer to be output from the crop phenotype prediction model. The method for constructing the training dataset will be described in detail later with reference to FIG. 2.
[0034] The genomic data of the training dataset constructed in the next step S110 is input into a crop phenotype prediction model to be trained (S120), and the crop phenotype prediction model is made to predict phenotype data from the input crop genomic data (S130).
[0035] Subsequently, the phenotype data predicted by the crop phenotype prediction model in step S130 is compared with the ground truth phenotype data of the training dataset constructed in step S110, and the parameters of the crop phenotype prediction model are updated so that the difference between the two is reduced (S140).
[0036] The construction of the training dataset in step S110 will be described in detail below with reference to FIG. 2. FIG. 2 is a diagram illustrating the flow of a method for constructing a training dataset for a crop phenotype prediction model according to another embodiment of the present invention.
[0037] To build a training dataset, as described above, genomic data and phenotypic data of the target crop are first obtained (S111).
[0038] Genomic data includes genotype data for crops and location data on the genomes of each genotype. Genomic data can be configured to include only SNP (Single Nucleotide Polymorphism) data, that is, genotype data in which a single base is found to be mutable, rather than all genotype data of the crop. More specifically, it can be composed of SNP data in which the proportion of missing data is less than a certain percentage (e.g., 20%) and the minimum allele frequency is greater than a certain percentage (e.g., 5%).
[0039] Phenotype data is data that represents the traits of a crop. In the case of fruit and vegetable crops, phenotype data can be composed of traits such as fruit weight, fruit width, fruit length, hardness, and sugar content.
[0040] Meanwhile, in step S111, genomic and phenotypic data are acquired separately for each crop variety.
[0041] The genomic data and phenotypic data obtained in the next step S111 are preprocessed (S112). The preprocessing process includes mapping the genomic data and phenotypic data into integer values.
[0042] Since positional data among genomic data is acquired in a numerical state, it is required to convert the nucleotide sequences of genotype data in genomic data into integers. Meanwhile, since genotype data consists of SNP data, allele combinations (AA, AT, TT or GG, GC, CC) are converted into integers by mapping them to their respective integer values. For example, genotype AA can be converted to 0, AT to 1, TT to 2, ..., and so on to convert them into integers.
[0043] The trait data constituting the phenotype data consists of numerical values, and since they are continuous data, they may contain outliers that can affect prediction accuracy. To address this, the preprocessing in step S112 categorizes each trait data into multiple intervals. For example, in the case of weight data, ~20 is categorized as 0, 20~40 as 1, 40~60 as 2, 60~80 as 3, and 80~ as 4.
[0044] Then, the genomic data and phenotypic data preprocessed in step S112 are combined to form a training dataset for a crop phenotypic prediction model (S113).
[0045] Then, an embedding vector for the genomic data preprocessed in step S112 is extracted (S114), and new genotype data different from the genotype data of the genomic data obtained in step S111 is generated from the extracted embedding vector (S115).
[0046] In the following step S115, new genotype data is combined with location data acquired in step S111 to form genomic data, and the genomic data is combined with phenotype data acquired / preprocessed in steps S111 / S112 to add a new training dataset (S116). This augments the training dataset.
[0047] In step S116, only the genotype data from the existing genomic data is replaced with the new genotype data generated in step S115. The newly generated genomic data is genomic data in which the genotype data has been modified at practically significant locations; while it is genomic data in which genetic mutations have occurred, it possesses characteristics similar to the genotype structure of the actual genomic data.
[0048] Meanwhile, the embedding vector extraction step (S114) and the new genotype data generation step (S115) are performed based on machine learning. The structure of the genotype data augmentation model, which is a machine learning model for performing steps S114 and S115, is shown in FIGS. 3 and 4.
[0049] As illustrated in FIGS. 3 and 4, the genotype data augmentation model is configured to include an embedding network (210) and a genotype generation network (220), and the genotype generation network (220) is configured to include a 1D convolution layer (221) and an FC (Fully Connected) layer (222).
[0050] The embedding network (210) is configured to extract embedding vectors of preprocessed genomic data. Genomic data is input into the embedding network (210) separated by species (BP1150, BP1151, BP1152, BP1155).
[0051] Genome data consists of location data and genotype data, so the embedding network (210) extracts an embedding vector from the location data and extracts an embedding vector from the genotype data, and concatenates the two extracted embedding vectors into one embedding vector and outputs it.
[0052] The genotype generation network (220) generates new genotype data from the embedding vectors extracted and output from the embedding network (210). To this end, the 1D convolution layer (221) of the genotype generation network (220) extracts high-dimensional features from the embedding vectors extracted from the embedding network (210), and the FC layer (222) generates new genotype data by mean pooling the high-dimensional features extracted by the 1D convolution layer (221).
[0053] The genotype generation network (220) is a structure in which a 1D convolution layer (221) learns local patterns and an FC layer (222) learns global features to reduce model complexity and prevent overfitting, thereby strengthening the patterns of the entire data.
[0054] Hereinafter, a method for training a genotype data augmentation model composed of an embedding network (210) and a genotype generation network (220) will be described in detail with reference to FIG. 5. FIG. 5 is a flowchart of a method for training a genotype data augmentation model according to another embodiment of the present invention.
[0055] To train a genotype data augmentation model, first, genomic data of the target crop is acquired (S310), and the acquired genomic data is preprocessed (S320). The preprocessing includes a process of mapping the genomic data to integer values.
[0056] In the next step S320, the preprocessed genomic data is input into the embedding network (210) of the genotype data augmentation model to extract the embedding vector of the preprocessed genomic data (S330).
[0057] Then, the embedding vector extracted in step S330 is input into a genotype generation network (220) to generate new genotype data (S340).
[0058] Subsequently, the parameters of the embedding network (210) and the genotype generation network (220) are updated in a way that reduces the loss between the new genotype data generated in step S340 and the genotype data of the preprocessed genome data in step S320 (S350).
[0059] The cross-entropy loss function can be utilized as an applicable loss function in the S350 stage, and the specific loss function can be configured as follows.
[0060]
[0061] Here, is the genotype data of the genome data preprocessed in step S320, and is new genotype data generated in step S340. The above loss function enables the reliability and accuracy of a crop phenotype prediction model trained on a training dataset constructed from genotype data by augmenting it with genotype data that has statistical characteristics similar to the existing genotype data.
[0062] Hereinafter, a method for predicting a crop phenotype using a model trained according to FIG. 1 will be described in detail with reference to FIG. 6. FIG. 6 is a diagram illustrating the flow of a crop phenotype prediction method according to another embodiment of the present invention.
[0063] To predict a crop phenotype, the genomic data of the target crop for which phenotype data is to be predicted is first obtained (S410).
[0064] Genome data obtained in the next step S410 is input into a crop phenotype prediction model (S420), and the crop phenotype prediction model predicts phenotype data from the input crop genome data (S430).
[0065] Then, the phenotypic data predicted in step S430 is output as phenotypic data for the target crop (S440).
[0066] Figure 7 presents the results of a phenotypic data prediction experiment conducted on 192 tomato varieties using a crop phenotypic prediction model trained according to the learning method of an embodiment of the present invention.
[0067] In the experiment, the crop phenotype prediction model was constructed based on various machine learning models. The machine learning models that served as the basis for the crop phenotype prediction model are EN (Elastic Net), LLAR (Lasso Least Angle Regression), LASSO (Lasso), ADA (AdaBoost), LR (Linear Regression), BR (Bayesian Ridge), RIDGE (Ridge Regression), CAT (CatBoost), HUBER (Huber Regressor), KNN (K-Neighbors), and RF (Random Forest).
[0068] When using the training dataset constructed according to the genotype data augmentation presented in the embodiment of the present invention, the R in predicting characteristics such as fruit weight and firmness of tomatoes is higher than when not doing so. 2 It was confirmed that the score improved by more than 20%, and the Mean Absolute Error (MAE) and Root Mean Square Error (RMSE) decreased, indicating a significant improvement in prediction accuracy. Accordingly, it can be seen that the construction of a training dataset and the training of a model using it, according to an embodiment of the present invention, can effectively improve the prediction performance of a crop phenotype prediction model even with limited data.
[0069] FIG. 8 is a diagram illustrating the hardware configuration of a crop phenotype prediction system according to another embodiment of the present invention. The crop phenotype prediction system according to an embodiment of the present invention can be implemented as a computing system comprising a communication unit (510), an output unit (520), a processor (530), an input unit (540), and a storage unit (550) as illustrated.
[0070] The communication unit (510) is a communication interface for connecting to an external network or external device, and is configured to acquire genomic data and phenotypic data of a target crop required for building a learning dataset in relation to an embodiment of the present invention, and to acquire genomic data of a target crop for prediction.
[0071] The output unit (520) is an output means for displaying the result of an operation performed by the processor (530), and the input unit (540) is a user interface that receives user commands and transmits them to the processor (530).
[0072] The processor (530) trains a genotype data augmentation model according to the procedure illustrated in FIG. 5 above, constructs a training dataset for a crop phenotype prediction model according to the procedure illustrated in FIG. 2 above using the trained genotype data augmentation model, and trains a crop phenotype prediction model according to the procedure illustrated in FIG. 1 above. Additionally, the processor (530) predicts crop phenotype data from crop genome data according to the procedure illustrated in FIG. 6 using the trained crop phenotype prediction model.
[0073] The storage unit (550) provides storage space necessary for the processor (530) to function and operate. In relation to an embodiment of the present invention, the storage unit (550) may store a learning dataset, a genotype data augmentation model, a crop phenotype prediction model, etc.
[0074] Up to now, preferred embodiments of a crop phenotype prediction method and system utilizing crop genotype data augmentation based on genotype location information have been described in detail.
[0075] In the above embodiment, the crop phenotype prediction model was trained using a diverse, realistic, and sufficient training dataset constructed by augmenting crop genotype data using genotype location information and SNP data, thereby improving the prediction accuracy of the crop phenotype prediction model.
[0076] In particular, by reflecting SNP positional dependency when augmenting genotype data to ensure that gene mutations occur only at realistically significant locations, it is possible to augment with genotype data that corresponds to actual biological characteristics. This allows for increased reliability and accuracy of the crop phenotype prediction model using a training dataset constructed through data augmentation that has statistical characteristics similar to existing genotype data.
[0077] Meanwhile, it goes without saying that the technical concept of the present invention may also be applied to a computer-readable recording medium containing a computer program that enables the device and method according to the present embodiment to perform their functions. Furthermore, the technical concept according to various embodiments of the present invention may be implemented in the form of computer-readable code recorded on a computer-readable recording medium. A computer-readable recording medium may be any data storage device that can be read by a computer and store data. For example, a computer-readable recording medium may be a ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical disk, hard disk drive, etc. Additionally, computer-readable code or a program stored on a computer-readable recording medium may be transmitted through a network connected between computers.
[0078] Furthermore, although preferred embodiments of the present invention have been illustrated and described above, the present invention is not limited to the specific embodiments described above. Various modifications are possible by those skilled in the art without departing from the essence of the invention as claimed in the claims, and such modifications should not be understood individually from the technical spirit or perspective of the present invention.
Claims
1. A step of constructing a training dataset for a crop phenotype prediction model, which is a machine learning model that predicts phenotype data composed of crop traits from crop genome data; A step of inputting genomic data from a constructed training dataset into a crop phenotype prediction model to be trained, thereby predicting phenotype data from the crop genomic data; A method for training a crop phenotype prediction model characterized by including the step of updating the parameters of the crop phenotype prediction model so that the difference between the predicted phenotype data and the phenotype data of the constructed training dataset is reduced.
2. In Claim 1, The construction phase is, A step of acquiring genomic data including genotype data for crops and location data for each genotype; Step of acquiring phenotypic data for crops; A step of preprocessing acquired genomic data and phenotypic data; A method for training a crop phenotype prediction model, characterized by including the step of constructing a training dataset for a crop phenotype prediction model, which is a machine learning model that predicts crop phenotype data from crop genomic data by combining preprocessed genomic data and phenotype data.
3. In Claim 2, The construction phase is, A step of extracting an embedding vector from preprocessed genomic data; A step of generating new genotype data different from the genotype data of the preprocessed genome data from the extracted embedding vector; A method for training a crop phenotype prediction model, characterized by further including the step of adding a new training dataset by combining newly generated genotype data with location data and preprocessed genomic data and preprocessed phenotype data.
4. In Claim 3, Genotype data, A method for training a crop phenotype prediction model characterized by SNP (Single Nucleotide Polymorphism) data.
5. In Claim 4, The preprocessing step is, A method for training a crop phenotype prediction model characterized by converting the base sequence of genotype data into an integer, wherein the allele combinations of the genotype data are mapped to respective integer values to convert them into integers.
6. In Claim 3, The preprocessing step is, A crop phenotype prediction model learning method characterized by dividing each trait data of phenotype data into multiple intervals and categorizing them.
7. In Claim 3, The extraction step is, A step of extracting an embedding vector from genotype data of genomic data; A step of extracting an embedding vector from location data of genomic data; A method for training a crop phenotype prediction model characterized by including the step of combining extracted embedding vectors into a single embedding vector.
8. In Claim 7, The generation step is, A step of inputting the combined embedding vector into a convolution layer to extract features; A method for training a crop phenotype prediction model characterized by including the step of inputting extracted features into an FCL and generating new genotype data by mean pooling the features.
9. In Claim 8, The extraction step is, It is performed using an embedding network that extracts embedding vectors from preprocessed genomic data, and The generation step is, It is performed using a genotype generation network that generates new genotype data from extracted embedding vectors, and Embedding networks and genotype generation networks are, A method for training a crop phenotype prediction model characterized by updating and learning parameters in a direction that reduces the loss between new genotype data and preprocessed genotype data.
10. A storage unit in which a training dataset of a crop phenotype prediction model, which is a machine learning model that predicts phenotype data composed of crop traits from crop genome data, is constructed; and A crop phenotype prediction model learning system characterized by including a processor that inputs genomic data of a constructed training dataset into a crop phenotype prediction model to be trained, predicts phenotype data from the crop genomic data, and updates the parameters of the crop phenotype prediction model so as to reduce the difference between the predicted phenotype data and the phenotype data of the constructed training dataset.
11. Step of acquiring crop genome data; A step of predicting phenotype data from crop genome data by inputting the acquired genome data into a crop phenotype prediction model, which is a machine learning model that predicts phenotype data composed of crop traits from crop genome data; Includes a step of outputting predicted phenotype data; and Crop phenotype prediction models are, Step of building the training dataset; A step of inputting genomic data from a constructed training dataset into a crop phenotype prediction model to be trained, thereby predicting phenotype data from the crop genomic data; A crop phenotype prediction method characterized by being trained by the step of updating the parameters of a crop phenotype prediction model so that the difference between the predicted phenotype data and the phenotype data of a constructed training dataset is reduced.
12. Acquisition unit for acquiring crop genome data; A processor that predicts phenotype data from crop genome data by inputting acquired genome data into a crop phenotype prediction model, which is a machine learning model that predicts phenotype data composed of crop traits from crop genome data; Includes an output unit that outputs predicted phenotype data; and Crop phenotype prediction models are, Construct a training dataset, Input the genomic data of the constructed training dataset into a crop phenotype prediction model to be trained, and predict phenotype data from the crop genomic data, A crop phenotype prediction system characterized by the parameters of a crop phenotype prediction model being trained by updating to reduce the difference between the predicted phenotype data and the phenotype data of a constructed training dataset.