Crop yield prediction method and system based on space-time geographically weighted regression
By combining spatiotemporal geographic weighted regression and LSTM model, using Moran eigenvectors to construct a crop yield prediction model, the problems of insufficient spatial representation and neglected spatiotemporal characteristics in the existing methods are solved, and high-precision and widely applicable crop yield prediction are achieved.
Patent Information
- Application Number
- CN202510765496.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-10
AI Technical Summary
The existing crop yield prediction methods have insufficient spatial representation and limited generalization capabilities in large-scale promotion and application, and have failed to fully consider the spatiotemporal characteristics of crop growth, which makes it difficult to improve the prediction accuracy.
A method based on spatiotemporal geographic weighted regression is adopted, combining long and short-term memory model (LSTM) and Moran eigenvectors to build a crop yield prediction model. Through spatiotemporal data fitting and nonlinear mapping, spatial autocorrelation, spatiotemporal heterogeneity and time dependence are integrated to improve prediction accuracy and generalization capabilities.
It achieves high prediction accuracy and strong generalization capabilities in different regions and years, breaks through the limitations of traditional methods, and significantly improves the accuracy and efficiency of crop yield prediction.
Smart Images

Figure CN120278347A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of crop yield prediction, and specifically relates to a crop yield prediction method and system based on spatio-temporal geographically weighted regression. Background Art
[0002] There are generally three categories of existing crop yield prediction methods: (1) field sampling-based on-site investigation methods, such as agronomic sampling methods, harvest yield measurement methods, etc.; (2) prediction methods based on remote sensing data, mainly using vegetation indices (such as the normalized difference vegetation index NDVI, enhanced vegetation index EVI, etc.) for statistical analysis with measured yield data; (3) statistical data-based modeling methods, which predict by establishing the relationship between various explanatory variables such as meteorology, soil, and management and crop yield.
[0003] The above methods have their own advantages and disadvantages in practical applications. The field sampling-based on-site investigation method has the advantage of being able to obtain high-precision point data, but there are problems such as limited sampling points, insufficient spatial representativeness, and it is difficult to achieve large-scale popularization and application due to factors such as equipment, cost, and operation efficiency. The prediction method based on remote sensing data has the advantages of large-scale and rapid data acquisition, but faces technical challenges such as long revisit periods, insufficient spatial resolution of some sensors, and being easily restricted by weather conditions. These two methods also have a common limitation: usually, predictions are made for specific plots, and their applicability to different regions and different years is poor, and the generalization ability is limited.
[0004] The statistical data-based modeling method can usually achieve crop yield predictions over a relatively large range, but there are still significant limitations: existing methods do not fully consider the key spatio-temporal characteristics that change during the crop growth process in the modeling process, and spatio-temporal characteristics have an important impact on the crop growth process. Existing methods often use simple assumptions or directly ignore these complex relationships, resulting in the prediction accuracy of the model being difficult to further improve. Summary of the Invention
[0005] The purpose of this application is to provide a crop yield prediction method and system based on spatio-temporal geographically weighted regression. When constructing a crop yield prediction model, this application effectively integrates spatio-temporal features, which can further improve the generalization ability and prediction accuracy of the crop yield prediction model.
[0006] On the one hand, the crop yield prediction method based on spatio-temporal geographically weighted regression provided by this application includes: Obtain the county-level environmental data of the target area during the growth cycle of the target crop and the county-level yield data of the target crop, where the county-level environmental data includes the impact factor data that affects the yield of the target crop; Perform differentiation and factor detection on each impact factor data in the county-level environmental data respectively, and screen out the key explanatory variables from the impact factors; Build a crop yield prediction model and train it using county-level data of key explanatory variables and county-level yield data of the target crop; Use the trained crop yield prediction model to predict the county-level yield of the target crop; The above crop yield prediction model is constructed based on a spatio-temporal geographically weighted regression model with key explanatory variables as independent variables and the target crop yield as the dependent variable, and includes: a first LSTM model and a second LSTM model; wherein, the first LSTM model is trained to fit the spatio-temporal weights of each independent variable in the spatio-temporal geographically weighted regression model and output a spatio-temporal weight sequence; the second LSTM model is trained to non-linearly fit the weighted values obtained by multiplying the independent variable sequence by the spatio-temporal weight sequence to obtain the predicted value of the county-level yield of the target crop.
[0007] In some specific embodiments, the county-level environmental data includes multiple types of agricultural meteorological data, geographical data, and target crop cultivated land data.
[0008] In some specific embodiments, the mathematical expression of the crop yield prediction model is as follows: ; Wherein: represents the crop yield of the county i ; represents the i th p explanatory variable of the county weighted term, is the explanatory variable spatio-temporal weight, p = 1, 2,... P , P is the number of explanatory variables; represents non-linear fitting.
[0009] In some specific embodiments, the crop yield prediction method of the present application further includes: Calculate a centered geographical adjacency matrix according to the geographical polygon data of the target area, extract Moran eigenvectors from the geographical adjacency matrix using the Moran eigenvector spatial filtering method, and take all or part of the Moran eigenvectors as characteristic variables; wherein, the geographical adjacency matrix is a matrix describing the adjacent relationship between counties in the target area; Perform differentiation and factor detection on each characteristic variable respectively, and screen out key characteristic variables; When constructing the crop yield prediction model, the independent variables further include key characteristic variables.
[0010] When the independent variables further include key characteristic variables, the mathematical expression of the crop yield prediction model is as follows: ; Among them: represents the crop yield of the county region; i of the county region; represents the county region i of the p th explanatory variable weighted term, is the explanatory variable spatiotemporal weight, p = 1, 2,... P , P is the number of explanatory variables; represents the county region i of the q th characteristic variable weighted term, is the characteristic variable spatiotemporal weight, q = 1, 2,... Q , Q is the number of characteristic variables; represents non - linear fitting.
[0011] In some specific embodiments, the first LSTM model sequentially includes: An input layer, which is used to receive a data sequence composed of spatiotemporal data. Each group of spatiotemporal data in the data sequence includes a time node and the county region location data under this time node; One or several LSTM layers, the number of layers of which is determined according to simulation experiments, and which is used to fit the spatiotemporal weight sequence corresponding to the independent variable sequence under the input spatiotemporal data and output it; One or several fully - connected layers, the number of layers of which is determined according to simulation experiments, and which is used to map the output of the LSTM layer to the feature space; An output layer, which is used to output the spatiotemporal weight sequence in the feature space.
[0012] In some specific embodiments, the second LSTM model sequentially includes: An input layer, which is used to receive the independent variable sequence and the spatiotemporal weight sequence output by the first LSTM model, and multiply the independent variable sequence and the spatiotemporal weight sequence correspondingly to obtain a weighted value sequence; One or several LSTM layers, the number of layers of which is determined according to simulation experiments, and which is used to perform non - linear fitting on the weighted values in the weighted value sequence to obtain the target crop yield prediction value and output it; One or several fully - connected layers, the number of layers of which is determined according to simulation experiments, and which is used to map the output of the LSTM layer to the target space; An output layer, which is used to output the target crop yield prediction value in the target space.
[0013] In some specific embodiments, predicting the county-level yield of a target crop using a trained crop yield prediction model includes: Inputting the target year, county location data, and the county-level data corresponding to the independent variables in the target year, and the crop yield prediction model outputs the predicted value of the county-level yield of the target crop in the target year.
[0014] On the other hand, the crop yield prediction system based on spatio-temporal geographically weighted regression provided in this application includes: The first module, which is constructed based on the LSTM model and is trained to fit the spatio-temporal weight sequence corresponding to the independent variable sequence according to the input target year and county location data; The second module, which is constructed based on the LSTM model and is trained to learn the non-linear mapping relationship between the weighted values in the weighted value sequence and the county-level yield of the target crop, and perform non-linear fitting on the weighted values in the weighted value sequence in the target year based on the non-linear mapping relationship to obtain the predicted value of the county-level yield of the target crop in the target year; wherein, the weighted value sequence is obtained by multiplying the independent variable sequence by the spatio-temporal weight sequence; the independent variable sequence includes the county-level data of the above key explanatory variables.
[0015] In some specific embodiments, the independent variable sequence further includes the data of the above key feature variables.
[0016] Compared with the prior art, this application has the following advantages and beneficial effects: This application combines the spatio-temporal geographically weighted regression model with the ability to analyze spatial heterogeneity and the long short-term memory model LSTM that can extract time series features. At the same time, the long short-term memory model LSTM is also used to perform non-linear fitting on the weighted terms in the spatio-temporal geographically weighted regression model, breaking through the limitation that the spatio-temporal geographically weighted regression model can only handle linear relationships, thereby significantly improving the prediction accuracy and generalization ability of the crop yield prediction model. In the preferred solution of this application, the Moran eigenvector containing spatial structure information is introduced as an independent variable into the prediction model to achieve the deep fusion of spatial autocorrelation, spatio-temporal heterogeneity, and time dependence, thereby further improving the performance of the prediction model.
[0017] The model constructed in this application shows high prediction accuracy and strong generalization ability for different regions and different years, and can effectively overcome the limitations of current crop prediction methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a schematic flowchart of the crop yield prediction method of this application; Figure 2 It is a schematic diagram of the structural principle of the first LSMT model and the second LSMT model in the embodiment. SPECIFIC EMBODIMENTS
[0019] To make the objectives, technical solutions, and advantages of this application clearer, the following will, in conjunction with the accompanying drawings, clearly and completely describe the technical solutions of this application. Obviously, the specific embodiments described are only a part of the specific embodiments of this application, rather than all of them. All other specific embodiments obtained by those of ordinary skill in the art based on the specific embodiments in this application without creative efforts belong to the scope of protection of this application.
[0020] The following will, in conjunction with Figure 1 , provide the specific implementation process of the crop yield prediction method based on spatio-temporal geographically weighted regression of this application. The specific steps are as follows: S100: Obtain the county-level environmental data of the target area during the growth cycle of the target crop and the county-level yield data of the target crop. Among them, the county-level environmental data includes the impact factor data that affects the yield of the target crop; The target crop is the crop to be predicted. The crop growth cycle generally includes the sowing period, the growth period, and the maturity period. In this specific embodiment, the target crop is selected as corn, whose sowing period is from March to April, the growth period is from May to August, and the maturity period is from September to November.
[0021] In this specific embodiment, the county-level environmental data includes agricultural meteorological data, geographical data, and target crop cultivated land data. The agricultural meteorological data further includes daily rainfall, daily maximum temperature, daily minimum temperature, daily average temperature, daily average dew point temperature, daily maximum vapor pressure deficit, daily minimum vapor pressure deficit, and daily temperature difference; the geographical data includes altitude, longitude, and latitude; the target crop cultivated land data includes the cultivated land area of the target crop.
[0022] In this specific embodiment, the initial environmental data obtained is 4 km resolution data, so it also includes preprocessing the initial environmental data into county-level scale data. Specifically, first use the fishnet tool in the Geographic Information System (GIS) to establish a 4 km resolution fishnet model for the target area, then match the initial environmental data with the fishnet model based on spatial location, and finally aggregate the initial environmental data to the county-level scale.
[0023] S200: Conduct differentiation and factor detection on each impact factor data in the county-level environmental data respectively, and screen out the key explanatory variables from the impact factors; In this specific embodiment, a geographical detector is selected for differentiation and factor detection. Specifically, use R in the language of geodetector package for geographical detector analysis, including: using factor _ detector ( ) function to calculate the q value and pValues, retain the influencing factors that are significant at the significance level as key explanatory variables. The significance level is usually set to 0.05. The key explanatory variables retained in this specific embodiment are listed in Table 1 below.
[0024] Table 1 Key explanatory variables screened in this specific embodiment
[0025] A key explanatory variable refers to an influencing factor that has a significant impact on the target variable. In this application, the target variable refers to crop yield. The main purpose of this step is to retain key information while reducing the data dimension. Subsequently, constructing a prediction model based on the screened key explanatory variables helps improve the interpretability of the prediction model.
[0026] In a preferred embodiment of this specific embodiment, Moran eigenvectors that can better reflect spatial structure information are also introduced as feature variables in the prediction model to further improve the spatial autocorrelation of the prediction model.
[0027] The extraction method of Moran eigenvectors is as follows: Calculate the centered geographic adjacency matrix according to the geographic polygon data of the target area, and use the Moran eigenvector spatial filtering method to extract Moran eigenvectors from the geographic adjacency matrix, and take all or part of the Moran eigenvectors as feature variables.
[0028] Geographic polygon data is a type of spatial data in geographic information systems. It is composed of a series of ordered coordinate points connected to describe an area with a clear boundary in the geographic space. The geographic adjacency matrix is a matrix used to describe the adjacent relationship between regions in the geographic space. It is a binary matrix, and the elements in the matrix represent the adjacency relationship between spatial units and . In this application, the spatial unit corresponds to a county, that is, the geographic adjacency matrix described in this application describes the adjacent relationship between counties in the target area; when spatial units and are adjacent, then = 1; otherwise, = 0.
[0029] The centered geographic adjacency matrix is obtained by centering the geographic adjacency matrix. The centering process is as follows: Denote the geographic adjacency matrix as matrix C , the size of matrix C is n*n, Define the projection matrix M = I - J / n , I is the identity matrix of size n*n , J is n*nAll-ones matrix of size; Using the projection matrix M Project the matrix C onto the centered space, that is, obtain the centered geographical adjacency matrix. Specifically, the centered geographical adjacency matrix is MCM matrix.
[0030] Moran Eigenvector Spatial Filtering (MESF for short) is a method for processing spatial data. In this application, Moran eigenvectors are extracted from the geographical adjacency matrix by using Moran Eigenvector Spatial Filtering, thereby converting spatial structure information into feature variables.
[0031] Considering that there are a large number of extracted Moran eigenvectors, only some Moran eigenvectors can be taken as feature variables. Specifically, select some Moran eigenvectors that can better reflect the spatial structure from the Moran eigenvectors as feature variables. Generally speaking, Moran eigenvectors with larger absolute values of eigenvalues can better reflect the spatial structure. The screening rules for Moran eigenvectors can be set according to the eigenvalues. For example, take Moran eigenvectors with absolute values of eigenvalues greater than a preset threshold as feature variables, and the preset threshold is set in advance; or sort the Moran eigenvectors according to the absolute values of eigenvalues, and take several Moran eigenvectors with the largest absolute values of eigenvalues as feature variables.
[0032] In this specific embodiment, Moran eigenvectors are screened according to the absolute value of the ratio of the eigenvalue to the maximum eigenvalue, and Moran eigenvectors with the absolute value of the ratio greater than the ratio threshold are taken as feature variables. The ratio threshold is an empirical value, which is set to 0.5 in this specific embodiment, but is not limited to 0.5 and can be adjusted in actual applications.
[0033] S300: Construct a crop yield prediction model, and train it using the county-level data of key explanatory variables and the county-level yield data of the target crop; The constructed crop yield prediction model is constructed based on a spatio-temporal geographically weighted regression model with key explanatory variables as independent variables and the target crop yield as the dependent variable, and it includes: a first LSTM model and a second LSTM model; the first LSTM model is trained to fit the spatio-temporal weights of each independent variable in the spatio-temporal geographically weighted regression model and output a spatio-temporal weight sequence; the second LSTM model is trained to non-linearly fit the weighted values obtained by multiplying the independent variable sequence and the spatio-temporal weight sequence correspondingly, and obtain and output the predicted value of the county-level yield of the target crop.
[0034] In the preferred scheme of this specific embodiment, the independent variables further include key feature variables, which are screened by differentiating and factor detecting the foregoing feature variables. The mathematical expression of the crop yield prediction model constructed in this preferred scheme is as follows: (1) In formula (1): represents the dependent variable of position i , which represents the crop yield of the county in this application i ; represents the weighted term of the i th p explanatory variable of position , is the spatio-temporal weight of the explanatory variable , p = 1, 2,... P , P and is the number of explanatory variables; i represents the weighted term of the q th feature variable of position , is the spatio-temporal weight of the feature variable q = 1, 2,... Q , Q and is the number of feature variables;
[0035] The first LSTM model is trained to fit the spatio-temporal weight sequence corresponding to the independent variable sequence according to the input spatio-temporal data; the second LSTM model is trained to non-linearly fit the weighted value obtained by multiplying the independent variable sequence by the spatio-temporal weight sequence output by the first LSTM model, and output the predicted crop yield value.
[0036] The spatio-temporal weight sequence is composed of the spatio-temporal weights of each independent variable. Multiplying the independent variable sequence by the spatio-temporal weight sequence means multiplying each independent variable in the independent variable sequence by its corresponding spatio-temporal weight.
[0037] In this specific embodiment, both the first LSTM model and the second LSTM model are constructed using the torch library of Python.
[0038] The specific structure of the first LSTM model includes an input layer, one or several LSTM layers, one or several fully connected layers, and an output layer connected in sequence, where: An input layer, which is used to receive a data sequence composed of spatio-temporal data and process the data sequence into a sequence with a dimension of [B, T, C]; each group of spatio-temporal data in the data sequence includes a time node and the position data corresponding to the time node, B represents the batch size, T represents the length of the data sequence, and C represents spatio-temporal data, including longitude, latitude, and year; longitude and latitude represent the position data of the county, and specifically, the longitude and latitude data of the centroid of the county can be used to represent the position data of the county; according to the application scenario of the present application, the position data in each group of spatio-temporal data in the above data sequence is the same; One or several LSTM layers, preferably 1 - 3 layers, specifically determined according to simulation experiments, which are used to fit the spatio-temporal weight sequence corresponding to the independent variable sequence under the input spatio-temporal data and output it; One or several fully connected layers, preferably 3 - 7 layers, specifically determined according to simulation experiments, which are used to map the output of the LSTM layer to the feature space; An output layer, which is used to output the spatio-temporal weight sequence in the feature space.
[0039] As a preferred solution, a Dropout layer is connected after each LSTM layer and each fully connected layer to prevent overfitting.
[0040] The specific structure of the second LSTM model also includes an input layer, one or several LSTM layers, one or several fully connected layers, and an output layer connected in sequence, where: An input layer, which is used to receive the independent variable sequence and the spatio-temporal weight sequence output by the first LSTM model, and multiply the independent variable sequence and the spatio-temporal weight sequence correspondingly to obtain a weighted value sequence; One or several LSTM layers, preferably 1 - 3 layers, specifically determined according to simulation experiments, which are used to perform non-linear fitting on the weighted values in the weighted value sequence to obtain the target crop yield prediction value and output it; One or several fully connected layers, preferably 3 - 7 layers, specifically determined according to simulation experiments, which are used to map the output of the LSTM layer to the target space; An output layer, which is used to output the target crop yield prediction value in the target space.
[0041] As a preferred solution, a Dropout layer is connected after each LSTM layer and each fully connected layer to prevent overfitting.
[0042] Please refer to Figure 2 , which shows the principle schematic of the first LSTM model and the second LSTM model, where the number of LSTM layers and fully connected layers shown in the first LSTM model and the second LSTM model is 1, ( T -1), T , ( T+1) The moment corresponds to different years. Represents the input spatio-temporal data. , where and represent the position i position data, represents the time node; s and c respectively represent the hidden layer and the long-term memory node of the LSTM layer. s and c The first digit of the superscript of s and c is "1", representing the s and c of the first LSTM model; the first digit of the superscript is "2", representing the s and c of the second LSTM model; Represents the spatio-temporal weight sequence output by the first LSTM model; Represents the independent variable sequence; Represents the position i predicted value of the target crop county-level yield.
[0043] In this specific embodiment, the training of the crop yield prediction model includes: Construct a data set, including independent variable data and corresponding target crop county-level yield data in different years. In this specific embodiment, the independent variable data comes from county-level environmental data, specifically including data corresponding to key explanatory variables in county-level environmental data, and may also include data corresponding to key feature variables; randomly divide the data set into a training set, a validation set, and a test set, accounting for 50%, 25%, and 25% of the data set respectively; Use the training set to train the crop yield prediction model. During the training process, the LSTM layer of the first LSTM model is used to learn the spatio-temporal coupling characteristics of the spatio-temporal data sequence, and fit the spatio-temporal weight sequence based on the spatio-temporal coupling characteristics. The LSTM layer of the second LSTM model is used to learn the non-linear mapping relationship between the weighted value sequence and the target crop yield. At the end of each training cycle, input the validation set into the trained crop yield prediction model to predict the crop yield prediction value, and calculate the training loss and validation loss on the validation set; adjust the learning rate according to the validation loss, and save the model parameters with the smallest validation loss; Load the model parameters with the smallest validation loss, use the test set to test the crop yield prediction model, and calculate the evaluation index.
[0044] Adopt the coefficient of determination R 2 , mean absolute error MAE, Root Mean Square Error RMSE and running time t Evaluate the performance of the above-trained crop yield prediction model, and the evaluation index data are listed in Table 2. R 2 It is used to evaluate the model fitting effect. The closer its value is to 1, the better the model fitting effect. MAE and RMSE Both are used to evaluate the prediction accuracy of the model. MAE and RMSE The smaller the value, the higher the prediction accuracy of the model. t is the time for the model to complete a prediction task. t The smaller the value, the higher the computational efficiency of the model.
[0045] Meanwhile, the above dataset is also used to train and evaluate the traditional spatio-temporal geographically weighted regression model. The evaluation indexes are listed in Table 2. The independent variables and dependent variables of the traditional spatio-temporal geographically weighted regression model are the same as those of the model of this application. The traditional spatio-temporal geographically weighted regression model is based on linear fitting, that is, the dependent variable is predicted through linear fitting between the weighted values, and the spatio-temporal weights are determined by using the bisquare kernel function. It can be seen from Table 2 that the prediction accuracy and computational efficiency of the model of this application are significantly better than those of the traditional spatio-temporal geographically weighted regression model, especially the computational efficiency has been significantly improved.
[0046] Table 2 Evaluation Indexes of Crop Yield Prediction Model
[0047] The spatio-temporal geographically weighted regression model has spatio-temporal analysis capabilities. The long short-term memory model can simultaneously consider the long-term and short-term characteristics of time series data (such as agricultural meteorological data, etc.), and can effectively capture the temporal dynamic characteristics in the crop growth process. This application combines the spatio-temporal geographically weighted regression model and the long short-term memory model LSTM, which can more accurately capture the complex dynamic characteristics of crop yield, thereby improving the robustness and accuracy of prediction. In the preferred solution of this specific embodiment, the Moran eigenvector containing spatial structure information is introduced as an independent variable into the prediction model, which can reveal the spatial distribution characteristics of crop yield, and can further realize the deep fusion of spatial autocorrelation, spatio-temporal heterogeneity and time dependence, thereby further improving the performance of the prediction model.
[0048] S400: Use the trained crop yield prediction model to predict the county-level yield of the target crop.
[0049] In this specific embodiment, input the target year, county location data, and the county-level data corresponding to the independent variable in the target year, and the crop yield prediction model outputs the predicted value of the county-level yield of the target crop in the target year.
[0050] Note that the above is only a preferred embodiment of the present application and the technical principles applied. Those skilled in the art will understand that the present application is not limited to the specific embodiments described herein, and various obvious changes, re-adjustments and substitutions can be made by those skilled in the art without departing from the protection scope of the present application. Therefore, although the present application has been described in more detail through the above embodiments, the present application is not limited to the above embodiments. Without departing from the concept of the present application, more other equivalent embodiments can also be included, all of which fall within the protection scope of the present application.
Claims
1. A crop yield prediction method based on spatio-temporal geographically weighted regression, characterized in that, Including: Obtain the county-level environmental data of the target area and the county-level yield data of the target crop during the growth cycle of the target crop. Among them, the county-level environmental data includes the impact factor data that affects the yield of the target crop; Perform differentiation and factor detection on each impact factor data in the county-level environmental data respectively, and screen out the key explanatory variables from the impact factors; Construct a crop yield prediction model, and use the county-level data of the key explanatory variables and the county-level yield data of the target crop for training; Use the trained crop yield prediction model to predict the county-level yield of the target crop; The crop yield prediction model is constructed based on a spatio-temporal geographically weighted regression model with the key explanatory variables as independent variables and the target crop yield as the dependent variable, and includes: a first LSTM model and a second LSTM model; among them, the first LSTM model is trained to fit the spatio-temporal weights of each independent variable in the spatio-temporal geographically weighted regression model and output a spatio-temporal weight sequence; the second LSTM model is trained to non-linearly fit the weighted values obtained by multiplying the independent variable sequence and the spatio-temporal weight sequence correspondingly to obtain the predicted value of the county-level yield of the target crop.
2. The crop yield prediction method according to claim 1, characterized in that: The county-level environmental data includes multiple types of agricultural meteorological data, geographical data, and cultivated land data of the target crop.
3. The crop yield prediction method according to claim 1, characterized in that: The mathematical expression of the crop yield prediction model is as follows: ; Wherein: represents the crop yield of the county i ; represents the i th p explanatory variable weighted term of is the spatio-temporal weight of the explanatory variable ; p = 1, 2,... P , P is the number of explanatory variables; represents non-linear fitting.
4. The crop yield prediction method according to claim 1, characterized in that, further Including: Calculate a centralized geographical adjacency matrix according to the geographical polygon data of the target area, extract Moran eigenvectors from the geographical adjacency matrix by using the Moran eigenvector spatial filtering method, and take all or part of the Moran eigenvectors as feature variables; among them, the geographical adjacency matrix is a matrix describing the adjacent relationship between counties in the target area; Perform differentiation and factor detection on each feature variable respectively, and screen out the key feature variables; When constructing the crop yield prediction model, the independent variable also includes the key feature variables.
5. The crop yield prediction method according to claim 4, characterized in that: When the independent variable also includes the key feature variables, the mathematical expression of the crop yield prediction model is as follows: ; Wherein: represents the crop yield of the county i ; represents the i th p explanatory variable of the county weighted term, is the spatio-temporal weight of the explanatory variable ; p = 1, 2,... P , P is the number of explanatory variables; represents the i th q characteristic variable of the county weighted term, is the spatio-temporal weight of the characteristic variable ; q = 1, 2,... Q , Q is the number of characteristic variables; represents non-linear fitting.
6. The crop yield prediction method according to claim 1, characterized in that: The first LSTM model sequentially includes: An input layer, which is used to receive a data sequence composed of spatio-temporal data. Each group of spatio-temporal data in the data sequence includes a time node and the county location data under that time node; One or several LSTM layers, the number of layers of which is determined according to simulation experiments, and which is used to fit the spatio-temporal weight sequence corresponding to the independent variable sequence under the input spatio-temporal data and output it; One or several fully connected layers, the number of layers of which is determined according to simulation experiments, and which is used to map the output of the LSTM layer to the feature space; An output layer, which is used to output the spatio-temporal weight sequence in the feature space.
7. The crop yield prediction method according to claim 1, characterized in that: The second LSTM model sequentially includes: An input layer, which is used to receive the independent variable sequence and the spatio-temporal weight sequence output by the first LSTM model, and multiply the independent variable sequence and the spatio-temporal weight sequence correspondingly to obtain a weighted value sequence; One or several LSTM layers, the number of layers of which is determined according to simulation experiments, are used to perform non-linear fitting on the weighted values in the weighted value sequence to obtain the predicted value of the target crop yield and output it; One or several fully connected layers, the number of layers of which is determined according to simulation experiments, are used to map the output of the LSTM layer to the target space; An output layer is used to output the predicted value of the target crop yield in the target space.
8. The crop yield prediction method according to claim 1 or 4, characterized in that: Using the trained crop yield prediction model to predict the county-level yield of the target crop, including: Inputting the target year, county location data, and the county-level data corresponding to the independent variables in the target year, the crop yield prediction model outputs the predicted value of the county-level yield of the target crop in the target year.
9. A crop yield prediction system based on spatio-temporal geographically weighted regression, characterized in that, Including: The first module is constructed based on the LSTM model and is trained to fit the spatio-temporal weight sequence corresponding to the independent variable sequence according to the input target year and county location data; The second module is constructed based on the LSTM model and is trained to learn the non-linear mapping relationship between the weighted values in the weighted value sequence and the county-level yield of the target crop, and perform non-linear fitting on the weighted values in the weighted value sequence of the target year based on the non-linear mapping relationship to obtain the predicted value of the county-level yield of the target crop in the target year; the weighted value sequence is obtained by multiplying the independent variable sequence and the spatio-temporal weight sequence correspondingly; The independent variable sequence includes the county-level data of the key explanatory variables, and the key explanatory variables are determined by the following method: performing differentiation and factor detection on the data of each influencing factor that affects the target crop in the county-level environmental data, and screening out the key explanatory variables from the influencing factors.
10. The crop yield prediction system according to claim 9, characterized in that: The independent variable sequence further includes the data of the key feature variables, and the key feature variables are determined by the following method: Calculating a centralized geographic adjacency matrix according to the geographic polygon data of the target area, extracting Moran eigenvectors from the geographic adjacency matrix by using the Moran eigenvector spatial filtering method, and taking all or part of the Moran eigenvectors as the feature variables; wherein, the geographic adjacency matrix is a matrix describing the adjacent relationship between counties in the target area; Performing differentiation and factor detection on each feature variable respectively, and screening out the key feature variables.
Citation Information
Patent Citations
Crop yield estimation method based on deep space-time feature joint learning
CN111027752A
Soybean yield prediction method based on machine learning fused satellite and weather data
CN113537645A
Regional scale crop yield near-real-time prediction method based on deep learning
CN115730523A
Large-area adaptive crop yield prediction method and system based on attention mechanism
CN118070951A
KR20210114751A
Cited By
Precise wheat fertilization method based on soil spatial variability
CN121549155A