A deep learning-based regional-scale crop yield near-real-time prediction method

By combining a deep learning LSTM model with real-time phenological information, the problem of low crop yield estimation accuracy was solved, and near real-time prediction of crop yield during the growing season was achieved, thus improving prediction accuracy.

CN115730523BActive Publication Date: 2026-03-24HUAZHONG AGRI UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-28
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing technologies lack sufficient consideration of crop growth mechanisms and processes, resulting in low yield estimation accuracy and making it difficult to achieve near real-time prediction of crop yield during the growing season.

Method used

By employing the Long Short-Term Memory (LSTM) method in deep learning and combining it with real-time phenological information, a nonlinear relationship model between crop yield and related variables is constructed. Yield-related variables are obtained through remote sensing data and the model is trained to achieve near real-time yield prediction.

Benefits of technology

It achieves near real-time accurate prediction of crop yield, improving the accuracy of yield estimation and the predictive ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115730523B_ABST
    Figure CN115730523B_ABST
Patent Text Reader

Abstract

The application discloses a kind of regional scale crop yield near real-time prediction method based on deep learning, comprising: obtaining the source data of target pixel yield related variable, calculating the yield related variable of target area;Source data is processed into corresponding feature variable, and model training set and test set are generated;According to the characteristics of crop yield accumulation, construct a suitable LSTM model, the time series of feature variable in training set is imported into the model for training, and a yield prediction model is obtained;With R 2 And root mean square error (RMSE) as evaluation standard, the feature variable of test set is imported into yield prediction model for verification, and the optimal model is obtained.The application solves the restriction of existing yield estimation method due to the complex nonlinear relationship between yield and environmental variables changes with crop phenology, and also solves the problem that yield cannot be accurately predicted in near real time in the current method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of yield prediction technology, specifically to a near real-time prediction method for regional-scale crop yields based on deep learning. Background Technology

[0002] The formation and accumulation of crop yield is a complex process. Understanding this process hinges on identifying the crop's growth stages, the relationship between crop parameters and yield at different growth stages, and the response of crop growth, yield accumulation, and environmental factors. At different growth stages, the impact of environmental factors on yield varies significantly. For example, appropriate drought during the corn seedling stage can promote root growth and development; drought during the corn silking stage can affect pollination and cause barren tips; and insufficient carbohydrate supply during the corn milk stage can easily lead to the termination of grain growth.

[0003] Accurate understanding of crop yield information at the regional scale is crucial for macroeconomic regulation of food production and agricultural policy. Current research methods still lack sufficient consideration of crop growth mechanisms and processes, thus limiting the accuracy of yield estimation. Furthermore, previous methods primarily estimated yields after the harvest season, while real-time yield prediction during the growing season remains a challenge. How to combine crop growth mechanisms and processes to achieve near-real-time yield prediction during the growing season remains a key research challenge and bottleneck.

[0004] There are currently no effective solutions to these problems. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides a near-real-time crop yield prediction method at the regional scale based on deep learning. By combining near-real-time crop phenological estimation information and near-real-time crop area extraction results, and fully considering the key factors affecting crop yield formation and accumulation at different phenological stages, as well as the differences in the impact of these factors on crop yield at different phenological stages, this invention solves the bottleneck problem of yield estimation accuracy in existing methods. It also addresses the difficulty of current methods in achieving near-real-time intra-season yield prediction. This invention utilizes the Long Short-Term Memory (LSTM) method in deep learning, combined with real-time phenological information, to construct the nonlinear relationship and variation relationship between crop yield and related variables at different phenological stages, achieving near-real-time accurate prediction of crop yield within the growing season.

[0006] This invention provides a near real-time prediction method for regional-scale crop yield based on deep learning, comprising: acquiring source data of yield-related variables for target pixels; acquiring yield-related variables for the target region, including phenology, single-band, water stress, temperature, and radiative transfer variables; processing the source data into corresponding feature variables to generate a model training set and a test set; constructing a suitable LSTM model based on the characteristics of cumulative crop yield; importing the time series of feature variables from the generated training set into the model for training to obtain a yield prediction model; and using R... 2 Using the root mean square error (RMSE) as the evaluation criterion, the test set is imported into the yield prediction model for validation to obtain the optimal model; finally, the yield-related variables of the crop yield region to be measured are input into the model to obtain the yield.

[0007] Furthermore, the yield-related variables of the target crop are obtained, including: extracting target crop pixels through target crop mask data, and obtaining near real-time yield-related variables from the start of the crop growing season to the present time after preprocessing the source data of yield-related variables obtained through remote sensing, including key phenological periods, single-band reflectance, water stress, temperature, and radiation transfer variables.

[0008] Furthermore, the preprocessing of the source data for yield-related variables includes: first, unifying the spatiotemporal resolution of the source data, calculating the source data as characteristic variables, i.e., yield-related variables; then, integrating the yield-related variables according to phenological stages; and finally, calculating the average value of the yield-related variables at the regional scale of the yield statistics.

[0009] Furthermore, an LSTM crop yield estimation model is constructed, including setting and optimizing the structure of the LSTM model based on the characteristics of cumulative crop yield and considering the performance and accuracy of the model.

[0010] Furthermore, the training of the LSTM model includes: taking the processed yield-related variables as input and crop yield statistics as output, generating training set samples to train the crop yield estimation model, adjusting the training parameters according to the actual situation to ensure model performance, obtaining the optimized parameters of the model, and obtaining the yield estimation model.

[0011] Furthermore, real-time crop yield estimation includes: inputting yield-related variables of the crop yield region to the trained yield estimation model, obtaining the output prediction result, i.e., obtaining the real-time crop yield estimation result.

[0012] Beneficial effects

[0013] Compared with the prior art, the present invention has the following beneficial effects:

[0014] This invention provides a near real-time prediction method for regional-scale crop yield based on deep learning. By combining real-time phenological information, it effectively fits the complex nonlinear relationship between yield and yield-related variables in yield prediction. At the same time, it considers the differences in the impact of characteristic variables of key phenological stages of crops on yield, thus achieving near real-time prediction of crop yield. Attached Figure Description

[0015] Figure 1 This is a flowchart illustrating an example of a near real-time prediction method for regional-scale crop yield based on deep learning, as described in this invention.

[0016] Figure 2 Diagram showing key phenological stages;

[0017] Figure 3 A data processing roadmap;

[0018] Figure 4 This is a model for estimating and predicting LSTM output. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] Please see Figure 1 A near real-time prediction method for regional-scale crop yield based on deep learning includes the following steps:

[0021] Step 1: Based on the characteristics of crop yield accumulation process, determine the yield-related variables of the target crop, obtain the source data and calculate the yield-related variables after preprocessing, including phenology, single-band, water stress, temperature, and radiation transfer variables.

[0022] Step 2: Construct a deep learning-based crop yield estimation model and optimize its structure;

[0023] Step 3: Using the yield-related variables obtained in Step 1 as input data and historical crop yield statistics as output, generate a training set sample to train the crop yield estimation model, obtain the model's optimization parameters, obtain the trained model, and validate the model using a validation set, with the coefficient of determination (R²) used as the output. 2 The model is evaluated using two main performance metrics: root mean square error (RMSE) and root mean square error (RMSE), resulting in an optimized model.

[0024] Step four: Use the yield-related variables of the crop yield region to be measured as input to the model obtained in step three, and output the prediction results to obtain real-time crop estimation results.

[0025] In this invention example, the source data of yield-related variables for maize pixels in the target area are first obtained. These variables, including phenology, single-band, water stress, temperature, and radiative transfer variables, are acquired through remote sensing. The source data is then processed into corresponding feature variables to generate a model training set and a test set. A suitable LSTM model is constructed based on the characteristics of cumulative crop yield. The time series data of the feature variables from the generated training set are imported into the model for training, resulting in a yield prediction model. Using R... 2 Using the root mean square error (RMSE) as an evaluation criterion, the test set is imported into the yield prediction model for validation to obtain the optimal model.

[0026] In this embodiment of the invention, step 1 mainly includes the following parts:

[0027] 1. Obtain source data for production-related variables in the target region;

[0028] 2. Mask extraction of the target crop;

[0029] 3. Preprocess the source data.

[0030] The present invention is implemented using the US Corn Belt as an example, including six states: Iowa, Indiana, Illinois, eastern Nebraska, southern Minnesota, and northwestern Ohio.

[0031] Specifically, obtain source data for production-related variables in the target region:

[0032] Among the yield-related variables used in this invention example, the phenological-related variables include: four key phenological stages (emergence stage (VE) to silking stage (ES), silking stage to milk stage (SD), milk stage to waxy stage (DD), waxy stage to full maturity stage (DM)), and the key phenological stages are as follows: Figure 2Five sets of yield-related variables were identified through analysis: single-band related variables, water stress related variables, temperature related variables, radiative transfer related variables, and phenological period related variables (Table 1). The single-band related variables include reflectance in the red band, near-infrared band, shortwave infrared 1, and shortwave infrared 2. Water stress related variables include average precipitation, ET / PET, ET (evapotranspiration), and PET (potential evapotranspiration). Temperature related variables include KDD (high-temperature day), GDD (growing degree day), average maximum temperature, and average minimum temperature. Radiative transfer related variables include WDRVI, PARin, PARpotential, PARin×WDRVI, PARpotential×WDRVI, and PARratio. Phenological related variables include the near-real-time estimated lengths and start times of the four key phenological periods, with the data collected from 2001 to 2019.

[0033] Among them, the single-band variables are derived from the 8-day composite surface reflectance products MOD09Q1 and MOD09A1, synthesized from data from the Moderate-resolution Imaging Spectroradiometer (MODIS) sensor, including the red (620nm-670nm), near-infrared (841nm-876nm), shortwave infrared 1 (1230nm-1250nm), and shortwave infrared 2 (1628nm-1652nm) bands. The resolution of the red and near-infrared bands is 250m, and the resolution of the shortwave infrared 1 and shortwave infrared 2 bands is 500m. Among the water stress-related variables, evapotranspiration is derived from the 8-day composite global evapotranspiration product MOD09Q1 and MOD09A1 from MODIS. 16A2 provides evapotranspiration (ET), potential evapotranspiration (PET), latent heat flux, and potential latent heat flux data at a resolution of 500 meters; precipitation data comes from the PRISM meteorological dataset, which provides estimates of seven key climate elements: precipitation, minimum temperature, maximum temperature, mean dew point, minimum vapor pressure difference, maximum vapor pressure difference on the horizontal surface, and global shortwave total solar radiation; temperature-related variables are calculated from the maximum and minimum temperatures provided by the PRISM global spatially continuous meteorological dataset, with the formulas for calculating KDD (high temperature day) and GDD (growth degree day) as follows: ,

[0034] ,

[0035] Where Tmax is the maximum temperature and Tmin is the minimum temperature; among the radiative transfer-related variables, the variable directly reflecting solar radiation is calculated from the hourly shortwave radiation product dataset (NLDAS-2) provided by the North American Land Data Assimilation System (NLDAS) and the daily net radiation product dataset (CERES_NETFLUX) provided by the Cloud and Earth Radiation Energy System (CERES), with (PARin, PARpotential, PARratio) as calculated by the formula: PARin (Photosynthetically Active Radiation) is the photosynthetically active radiation, PARpotential (Potential Photosynthetically Active Radiation) is the potential photosynthetically active radiation, and PARratio is the ratio between them; WDRVI reflects the growth status of vegetation. WDRVI is constructed using band 1 (620nm-670nm) and band 2 (841nm-876nm) of the MODIS 8-day synthesized surface reflectance product MOD09Q1, with the following formula:

[0036] ,in The value is set to 0.1; all variables are shown in Table 1 below:

[0037]

[0038] Table 1

[0039] Specifically, a mask is used to extract the target crop and related variables:

[0040] In this invention example, the DeepCropMapping (DCM) crop planting area near real-time extraction method proposed by Zhejiang University is adopted. The crop classification data (Cropland Data Layer, CDL) released by the U.S. Department of Agriculture and the National Agricultural Statistics Service (NASS) is used as training data to construct a maize planting area extraction model and extract maize pixels in the target area.

[0041] Specifically, data preprocessing is performed on the source data:

[0042] In this invention example, the obtained yield-related variables are resampled to 250m and a unified temporal resolution is applied. Then, according to the four key phenological periods (ES, SD, DD, and DM), the data with unified spatiotemporal resolution are integrated into four groups, and the average value of each yield-related variable for all maize pixels in each county is calculated. The data processing roadmap is as follows. Figure 3

[0043] In this embodiment of the invention, step 2 mainly includes the following parts:

[0044] 1. Set the time step and number of layers of the LSTM model according to the characteristics of cumulative output;

[0045] 2. Determine the number of hidden units in the fixed batch and LSTM layers to ensure model performance;

[0046] This example uses an LSTM model with one input layer, two LSTM layers, and one output layer. Since crop yield accumulation mainly involves four key phenological stages, the time step of the LSTM model is set to 4 to capture the cumulative effect of each variable group at different growth stages. The input layer of the model aggregates the variables into averages at the four key stages, and the output is the county-level maize yield. This invention example implements an LSTM yield estimation and prediction model based on PyTorch. To improve the performance of the LSTM model, a fixed batch size of 500, two unidirectional LSTM layers with 450 hidden units each, and a gradient descent-based optimizer (Adam) with a learning rate of 0.001 are used for parameter optimization. The LSTM yield estimation and prediction model is as follows: Figure 4 .

[0047] In this embodiment of the invention, step 3 mainly includes the following parts:

[0048] 1. Generate a training set to train the model;

[0049] 3. Validate the model.

[0050] This invention uses preprocessed yield-related variables from the beginning of the crop growing season to different growth stages as model inputs, and statistically reported yield data as model outputs to form training and test sets. In this example, county-level maize yields from the USDA National Bureau of Statistics reports for the study area in the United States from 2001 to 2019 are used for model training and validation. Cross-validation is used to optimize the training and test sets. Data from 2001 to 2019 is used as the training set, and data from one specific year is used as the test set. The results are then validated by comparing predicted and reported yields. This study uses two main model performance evaluation metrics: the coefficient of determination (R²). 2 ) and root mean square error (RMSE). The formula for calculating RMSE is: ,in This refers to the crop yield predicted by an LSTM model based on multi-source remote sensing data and other auxiliary data. The crop yield is reported based on ground surveys. Finally, a yield estimation model is trained.

[0051] In this embodiment of the invention, step 4 mainly includes the following parts:

[0052] By inputting yield-related variables from the start of the crop growing season to the present time into the yield estimation model trained on the region to be tested, the output prediction result is obtained, that is, the real-time prediction result of crop yield is obtained.

[0053] This invention compares the model's predictions with those reported by NASS. The results show that the proposed prediction model has good estimation accuracy from the maize reproductive stage (i.e., the silking stage), with an RMSE of 0.93 Mh / ha. As crop growth information increases in subsequent growth stages, the accuracy of the yield model continues to improve. Therefore, the maize yield prediction model proposed in this invention has reliable near-real-time yield prediction capabilities at the regional scale.

[0054] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0055] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A near real-time prediction method for regional-scale crop yield based on deep learning, characterized in that, Includes the following steps: Step 1: Obtain yield-related variables of the target crop: Based on the characteristics of the crop yield accumulation process, determine the yield-related variables of the target crop, obtain the source data and preprocess it, and calculate the yield-related variables from the crop growing season to different growth stages, including phenology, single band, water stress, temperature, and radiation transfer variables. Step 2: Construction of a near real-time yield prediction model based on deep learning: Construct a crop yield estimation model based on deep learning and optimize the model structure; Step 3: Training the near real-time crop yield prediction model based on deep learning: Using the yield-related variables obtained in Step 1 as input data and historical crop yield statistics as output, a training set sample is generated to train the crop yield estimation model, obtain the optimized parameters of the model, obtain the trained model, and validate the model using a validation set, with the coefficient of determination (R²) used as the basis for the model's performance. 2 The model is evaluated using two main performance metrics: root mean square error (RMSE) and root mean square error (RMSE), to obtain the optimized model. Step 4: Near-real-time prediction of crop yield: Use the near-real-time time series data of the relevant characteristic variables of the crop yield region since the crop growing season as the input of the model obtained in Step 3, and output the prediction results to obtain the near-real-time prediction results of the crop. Based on the characteristics of crop yield accumulation process, yield-related variables of the target crop are determined. After obtaining source data and preprocessing, yield-related variables are calculated, including phenology, single-band, water stress, temperature, and radiation transfer variables. Step one involves obtaining yield-related variables for the target crop, including: The target crop pixels are extracted by masking data. After obtaining the source data of yield-related variables through remote sensing, the yield-related variables since the crop growing season are obtained through preprocessing, including key phenological periods, single-band reflectance, water stress, temperature, and radiative transfer variables. To unify the spatiotemporal resolution of the source data, the source data is calculated as characteristic variables, i.e., yield-related variables. Then, the yield-related variables are integrated according to phenological stages. Finally, the average value of the yield-related variables on the regional scale of the yield statistics is calculated.

2. The method for near real-time prediction of regional-scale crop yield based on deep learning according to claim 1, characterized in that, Step two involves constructing a near real-time output prediction model based on deep learning, including: Construct an LSTM deep learning model and optimize the model structure by using the key phenological period length as the LSTM time step.

3. The method for near real-time prediction of regional-scale crop yield based on deep learning according to claim 1, characterized in that, Step three, training the near real-time crop yield prediction model based on deep learning, includes: Using processed yield-related variables as input and historical crop yield statistics as output, a training set is generated to train the crop yield estimation model. The training parameters are adjusted according to the actual situation to ensure model performance, and the optimized parameters of the model are obtained to obtain the yield estimation model.

4. The method for near real-time prediction of regional-scale crop yield based on deep learning according to claim 1, characterized in that, Step four, near real-time crop yield prediction, includes: The yield estimation model is trained by inputting near-real-time yield-related variables of the crop to be measured from the beginning of the crop growing season to the present time, and the prediction results are calculated and output, that is, the real-time prediction results of crop yield are obtained.

Citation Information

Patent Citations

  • Crop yield estimation method based on deep space-time feature joint learning

    CN111027752A

  • Crop yield prediction method based on remote sensing and ensemble learning

    CN114819298A

  • KR20220144218A