A spatially adaptive crop yield prediction method and system
By introducing a spatially adaptive prediction network model in crop yield prediction, using the channel attention mechanism and multi-head self-attention mechanism, the problems of correlation differences between bands of multispectral remote sensing images and the attention weight shift of traditional models are solved, and higher prediction accuracy and adaptability are achieved.
Patent Information
- Application Number
- CN202410121863.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-29
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-01-29
AI Technical Summary
Existing crop yield prediction techniques ignore the differences in correlation between bands when processing multispectral remote sensing images in different regions, and traditional recursive neural network structures lead to attention weight shifts, resulting in low prediction accuracy.
A spatially adaptive crop yield prediction method is proposed. By constructing a prediction network model of the multi-spectral feature fusion layer and the crop growth timing feature extraction layer, the channel attention mechanism is used to capture the correlation between multi-spectral images, and the timing features of crop growth are extracted through the multi-head self-attention mechanism.
It improves the accuracy and reliability of crop yield prediction, can better adapt to spatial heterogeneity in different regions, and significantly improves the accuracy and consistency of prediction results.
Smart Images

Figure CN117952264B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of agricultural crop yield prediction, and in particular to a spatially adaptive crop yield prediction method and system. Background Art
[0002] Existing crop yield prediction technologies mainly include but are not limited to the following four categories:
[0003] (1) The main core of the method based on traditional statistical analysis is to collect and analyze a large amount of local data and make predictions based on this. However, this method has obvious disadvantages. First, data collection is time-consuming and costly, especially when it comes to large-scale data collection. In addition, due to the intervention of human factors in the data collection and statistical process, these data often have errors, resulting in inaccurate predictions. For example, data entry errors, sampling bias, and other statistical problems may lead to deviations in the final prediction results.
[0004] (2) Models that rely on a deep understanding of the principles of crop growth. Prediction methods based on crop growth mechanisms attempt to predict crop yields by deeply understanding each stage of crop growth from planting to harvesting. The core of this method is the comprehensive consideration of various factors of crop growth, such as soil, climate, water, light, etc. But the problem is that crop growth is a very complex process that is affected and interacted with by many factors. Even experienced researchers find it difficult to fully understand and accurately simulate this process. Therefore, many models in the past tend to oversimplify the various factors of crop growth, which leads to low prediction accuracy in real scenarios.
[0005] (3) Crop yield prediction based on traditional machine learning methods. With the development of machine learning technology, more and more researchers have adopted these methods in the study of crop yield prediction. The use of machine learning can manage and analyze complex data patterns, facilitating the identification and understanding of the complex interrelationships between various influencing factors without the need for a prior in-depth understanding of the basic mechanisms of crop growth. Machine learning-based models have the ability to automatically adjust and optimize parameters without the need for large amounts of experimental data and complex parameter calibration and model validation calculations. This is in stark contrast to traditional crop mechanical models. Typical traditional machine learning techniques for crop yield prediction include random forest regression, support vector regression, multi-layer regression, Lasso regression, and boosted regression trees. Compared with previous yield prediction methods, machine learning models are able to more accurately capture the complex nonlinear relationships in crop growth and show excellent performance in crop yield prediction. However, traditional machine learning models also have defects, such as the need to manually select and design features, which means that the performance of the model depends largely on the quality of feature selection.
[0006] (4) Crop yield prediction based on deep learning methods. Deep learning models can automatically learn and extract meaningful features, which greatly simplifies the process of feature selection and engineering. Furthermore, deep learning models are very good at processing large-scale, high-dimensional and complex data, enabling them to extract effective features from large, multi-source and heterogeneous data, which is more in line with the data characteristics that affect crop growth. Many current studies have verified the advantages of using deep learning models for crop yield prediction. At present, the main input data types of crop yield prediction models based on deep learning are tabular records of crop field growth environment and remote sensing time series image data. Considering factors such as accessibility, cost and error, most studies focus on selecting one of these two input data types. Models that use crop field growth environment records as input include data such as climate variables, precipitation, soil conditions, seed genotypes, and geographic locations. These factors together create the field growth environment of crops. The prior art uses a DNN network to extract features from heterogeneous data from different sources to predict corn yield. The prior art designs a CNN-RNN framework that uses one-dimensional convolution to separately capture the dependence of crop growth on soil variables at different depths and the temporal dependence on climate variables. They then used an RNN model to capture the dynamic changes in crop yield over time after genetic improvement of crop genotypes. The existing technology further improved the CNN-RNN model by integrating graph neural networks (GNNs), allowing crop growth characteristics from neighboring regions to be aggregated, thereby enhancing the model's fit to different geographical characteristics and improving its robustness in different regions. These studies have confirmed the great potential of deep learning in yield prediction, significantly outperforming traditional machine learning methods such as LASSO, RF, and decision trees. However, the diversity of variables affecting crop growth and the inconsistency of the types of input variables proposed by different researchers make it difficult to intuitively compare the performance of these models. In addition, data are usually obtained through sensors or weather stations and are usually recorded digitally. Factors such as low sampling density and sensor collection errors lead to reduced accuracy of field data, further limiting the performance of the model.
[0007] Compared with crop field data collected on site, multispectral remote sensing images are not only easier to obtain and cheaper, but also have a unified data format. This feature allows the same satellite product to provide images with the same resolution and error distribution for different regions, thereby facilitating direct comparisons between model performances and enhancing the portability of models across regions. Significant progress has been made in this area in recent years. For example, the prior art proposed a model based on convolutional neural networks (CNNs) and long short-term memory networks (LSTMs), and introduced Gaussian processes. They used MODIS remote sensing image sequences as input and successfully fitted crop yields. Similarly, the prior art used a 3D-CNN model to simulate the spatial spectral changes and temporal dependence during crop growth, while the prior art designed a DACM network that can capture the differences in crop growth in different regions and times, thereby accurately fitting the spatial heterogeneity of crop growth. These models have performed well in practical applications, providing strong evidence for the feasibility of predicting crop yields based on remote sensing images.
[0008] Existing models often encounter difficulties when predicting crop yields over large areas. The reason for this is that due to different geographical locations, the climate, soil and other environmental variables in each region have unique effects on crop growth and yield. We call this diversity of crop growth patterns due to geographical factors "spatial heterogeneity." At present, although some models have noticed this heterogeneity and tried to fit it, the actual operation is often based on overly simplified methods, resulting in many core influencing factors being ignored. In addition, due to the design and structure of these models, they may still produce deviations when considering spatial heterogeneity, thus affecting the final prediction accuracy.
[0009] The two main challenges we face are:
[0010] (1) Ignoring the differences in correlations between multispectral satellite image channels: Multispectral images from different regions may reflect unique correlations between spectral channels, which may be related to many key factors of crop growth. Ignoring these correlations may prevent the model from fully capturing the specificity between different regions, thereby affecting the accuracy of predictions.
[0011] (2) Limited by the traditional recurrent neural network structure: When trying to fit the impact of crop growth stages in different regions on final yield, some models may be constrained by the traditional recurrent neural network structure, resulting in the problem of attention weight shift. This structural limitation may prevent the model from deeply understanding and representing the dynamic changes of crop growth, further reducing the accuracy and reliability of the prediction. Summary of the invention
[0012] The purpose of the present invention is to overcome the problem that the above-mentioned prior art ignores the differences in correlation between multispectral remote sensing image bands in different regions, and the previous model's attention weight offset to the crop growth stage causes the model to have low accuracy in yield prediction in a large area, and proposes a new spatial adaptive prediction network.
[0013] Specifically, the present invention proposes a spatially adaptive crop yield prediction method, which includes:
[0014] Step 1: construct a spatial adaptive prediction network model including a multispectral feature fusion layer and a crop growth temporal feature extraction layer, and obtain a multispectral image sequence marked with actual crop yields as training data;
[0015] Step 2: The multispectral feature fusion layer reduces the dimension of each image in the training data to obtain reduced-dimensional features, performs average pooling and maximum pooling on the reduced-dimensional features to obtain two sets of pooling vectors, obtains attention weights according to the pooling vectors, and uses the attention weights to weight the reduced-dimensional features to generate two new feature maps, extracts the features of the new feature maps, and obtains crop growth features;
[0016] Step 3, inputting an input sequence consisting of crop growth features of each picture in the training data into the crop growth temporal feature extraction layer, the crop growth temporal feature extraction layer extracts the temporal features of the crop growth process from the input sequence of each time period using multi-head self-attention, predicts the crop yield according to the temporal features, obtains the crop yield prediction value, constructs a loss function according to the crop yield prediction value and the actual crop yield, and trains the spatial adaptive prediction network model to obtain a crop yield prediction model;
[0017] Step 4: input the spectral image sequence of the crop yield area to be predicted into the crop yield prediction model to obtain the crop yield prediction result of the area.
[0018] The spatially adaptive crop yield prediction method, wherein the training data in step 1 is a multispectral image sequence X = {X 1 ,X 2 ,...,X t}, where each image X i ∈R b×h×w represents the multispectral image of the area at time t, where b is the number of bands of the remote sensing image, and h and w are the vertical and horizontal pixel counts, respectively;
[0019] The step 2 includes: converting the training data into a reduced dimension feature of [b,n] dimension, where b represents the number of bands, n represents the number of boxes, and each box represents a pixel count within a specific intensity range; applying average pooling and maximum pooling to the reduced dimension feature in n dimension to obtain two sets of pooled vectors of dimension [b,1]; normalizing the pooled vectors after being processed by a shared two-layer MLP to obtain the attention weight, which is used to weight the reduced dimension feature to generate a feature map of dimension [b,n], and connecting it with the reduced dimension feature to form a new feature map of [3,b,n], and inputting the new feature map into a 2D convolution module for feature extraction to obtain the crop growth feature.
[0020] The spatially adaptive crop yield prediction method, wherein step 3 comprises:
[0021] The multi-head self-attention executes the self-attention mechanism in parallel in multiple different weight matrix subspaces, and splices the outputs of the self-attention mechanisms of all subspaces together to obtain the time series feature; wherein the self-attention mechanism is used to calculate the weighted sum of all parts for each part of the input sequence, and the self-attention mechanism includes:
[0022] For each part of the input sequence, a learnable weight matrix W is used Q , W K and W V Construct a query (Q), key (K), and value (V):
[0023] Q=XW Q
[0024] K=XW K
[0025] V=XW V
[0026] The score is calculated using the dot product of the query and all keys, then divided by the square root of the key dimension, and the Softmax function is applied to each part of the input sequence to calculate the weight Weight:
[0027]
[0028] Multiply the weight Weight by the value V to get the output of the self-attention mechanism.
[0029] The spatially adaptive crop yield prediction method, wherein the convolution kernel size of the 2D convolution module is 3*3.
[0030] The present invention also proposes a spatially adaptive crop yield prediction system, which includes:
[0031] The initial module is used to construct a spatial adaptive prediction network model including a multispectral feature fusion layer and a crop growth temporal feature extraction layer, and obtain a multispectral image sequence marked with actual crop yields as training data;
[0032] A training module is used to reduce the dimension of the features of each image in the training data through the multispectral feature fusion layer to obtain reduced dimension features, perform average pooling and maximum pooling on the reduced dimension features to obtain two groups of pooling vectors, obtain attention weights according to the pooling vectors, and use the attention weights to weight the reduced dimension features to generate two new feature maps, extract features of the new feature maps, and obtain crop growth features; input an input sequence consisting of crop growth features of each image in the training data into the crop growth temporal feature extraction layer, and the crop growth temporal feature extraction layer uses multi-head self-attention to extract the temporal features of the crop growth process from the input sequence in each time period, predict crop yield according to the temporal features, and obtain a crop yield prediction value; a loss function is constructed according to the crop yield prediction value and the actual crop yield to train the spatial adaptive prediction network model to obtain a crop yield prediction model;
[0033] The prediction module is used to input the spectral image sequence of the crop yield area to be predicted into the crop yield prediction model to obtain the crop yield prediction result of the area.
[0034] The spatially adaptive crop yield prediction system, wherein the training data in the initial module is a multispectral image sequence X={X 1 ,X 2 ,...,X t}, where each image X i ∈R b×h×w represents the multispectral image of the area at time t, where b is the number of bands of the remote sensing image, and h and w are the vertical and horizontal pixel counts, respectively;
[0035] The training module includes: converting the training data into a reduced dimension feature of [b,n] dimension, where b represents the number of bands, n represents the number of boxes, and each box represents a pixel count within a specific intensity range; applying average pooling and maximum pooling to the reduced dimension feature in n dimension to obtain two sets of pooling vectors of dimension [b,1]; the pooling vector is normalized after being processed by a shared two-layer MLP to obtain the attention weight, which is used to weight the reduced dimension feature to generate a feature map of dimension [b,n], and connect it with the reduced dimension feature to form a new feature map of [3,b,n], and input the new feature map into a 2D convolution module for feature extraction to obtain the crop growth feature.
[0036] The spatially adaptive crop yield prediction system, wherein the training module comprises:
[0037] The multi-head self-attention executes the self-attention mechanism in parallel in multiple different weight matrix subspaces, and splices the outputs of the self-attention mechanisms of all subspaces together to obtain the time series feature; wherein the self-attention mechanism is used to calculate the weighted sum of all parts for each part of the input sequence, and the self-attention mechanism includes:
[0038] For each part of the input sequence, a learnable weight matrix W is used Q , W K and W V Construct a query (Q), key (K), and value (V):
[0039] Q=XW Q
[0040] K=XW K
[0041] V=XW V
[0042] The score is calculated using the dot product of the query and all keys, then divided by the square root of the key dimension, and the Softmax function is applied to each part of the input sequence to calculate the weight Weight:
[0043]
[0044] Multiply the weight Weight by the value V to get the output of the self-attention mechanism.
[0045] The spatially adaptive crop yield prediction system, wherein the convolution kernel size of the 2D convolution module is 3*3.
[0046] The present invention also proposes a server, which includes the spatially adaptive crop yield prediction device.
[0047] The present invention also proposes a storage medium for storing a computer program for executing the spatially adaptive crop yield prediction method.
[0048] It can be seen from the above scheme that the advantages of the present invention are:
[0049] (1) Statistical evaluation indicators:
[0050] We are Figure 4Our prediction results for county-level soybean yield are reported in , focusing on three key metrics: RMSE, Corr, and MAPE. To reduce the random fluctuations in the performance of the model in different years, we conducted experiments for multiple years from 2009 to 2015. Each row records the prediction performance of different models in the same year. The results show that except for 2009 and 2013, when our model's RMSE is slightly lower than DACM, our model outperforms the competing methods in other years, with an average RMSE reduction of 9% compared to the state-of-the-art models. In addition, our proposed model performs well on Corr in different years, and the average Corr is the highest among all methods. On the MAPE metric, our model shows a clear superiority, exceeding the baseline by 10% in some years. In addition, the average MAPE for all the years studied is about 9% lower than the state-of-the-art models. Our model achieves state-of-the-art results on all three evaluation metrics, indicating that our method captures the spatial heterogeneity in the crop growth process more accurately than the competing methods.
[0051] (2) Distribution Similarity Measurement
[0052] In this study, we aimed to implement a model that can adaptively and accurately fit crop yields over a large area. To verify whether our model achieved its goal, we compared the distribution of yield values predicted by the model with the distribution of actual yield values, hoping that the predicted distribution would be as close to the actual distribution as possible. First, we estimated the probability density functions of the actual and predicted yield values separately using kernel density estimation. Figure 2 In , we visualized the probability density functions of the actual and predicted values for 2014. In addition, we Figure 2 The overlap between the predicted values and the actual values of different models is marked in the figure to quantitatively evaluate the fit of the model to the yield distribution over a large range. The results show that the yield distribution predicted by our model is closer to the distribution of actual yield values than all other methods, and the overlap of our proposed method is more than 7 percentage points higher than that of the most competitive model. Figure 2 The results demonstrate that our model fits the spatial heterogeneity of crop growth well and can adaptively and accurately predict crop yields over a large area.
[0053] (3) Error map analysis
[0054] To more effectively demonstrate the ability of our model to adaptively fit crop yields in different regions, we Figure 3The prediction errors of various models in 2015 are shown in Figure 1. In this figure, we can clearly see the distribution of prediction errors of different models in various counties, and the darker the color, the larger the prediction error in the county. It is obvious that our method keeps the prediction error within a very small range in most counties, and our method outperforms other competing methods in almost all regions. This shows that our method is suitable for adaptively predicting crop yields in a wide range of areas. This not only highlights the robustness and accuracy of our model, but also its ability to adapt to different geographical conditions and crop growth patterns. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 It is a flow chart of the present invention;
[0056] Figure 2 It is the probability density function graph of the actual value and the predicted value;
[0057] Figure 3 Prediction error plots for various models;
[0058] Figure 4 This is a diagram showing the prediction results of county-level soybean yields according to the present invention. DETAILED DESCRIPTION
[0059] In order to meet the above challenges, the present invention introduces a spatial adaptive prediction network (SAPN) model. This model consists of two core modules: a multispectral feature fusion module and a crop growth heterogeneity time series feature extraction module. The multispectral feature fusion module uses a channel attention mechanism to capture the correlation differences between multispectral images in different regions. By generating new feature channels and then fusing them through a convolution module, it provides a detailed understanding of the spectral properties of different regions. The crop growth heterogeneity time series feature extraction module uses adaptive position encoding to get rid of the sequential execution method of the traditional recursive network model. This innovation solves the problem of attention weight bias, enhances the calculation speed, and improves the accuracy of the yield prediction task.
[0060] In order to achieve the above technical effects, the present invention includes the following key technical points:
[0061] Key point 1: Multispectral feature fusion: The channel attention mechanism is used to capture the correlation differences between multispectral images in different regions. By generating new feature channels and then fusing them through convolution modules, it provides a detailed understanding of the spectral properties of different regions.
[0062] Key point 2: Extraction of temporal features of crop growth; the use of adaptive position encoding breaks away from the sequential execution method of the traditional recursive network model, solves the problem of attention weight bias, enhances the calculation speed, and improves the ultimate robustness of the model.
[0063] In order to make the above features and effects of the present invention more clearly and understandably described, embodiments are given below and described in detail with reference to the accompanying drawings. This specification discloses one or more embodiments that include the features of the present invention. The disclosed embodiments are only for illustration. The scope of protection of the present invention is not limited to the disclosed embodiments, and the present invention is defined by the attached claims.
[0064] The overall process of our method is as follows Figure 1 As shown. Given a multispectral image sequence X = {X1, X2, ..., Xt} covering a specific area of interest. Each image Xi∈Rb×h×w represents a multispectral image of the area at time t. Where b is the number of bands of the remote sensing image, while h and w are the vertical and horizontal pixel counts, respectively. Our goal is to learn a spatial adaptive prediction network (SAPN) model that maps the function F: It is a prediction of crop yield in the area covered by the multispectral image sequence X. We hope to predict the value As close as possible to the actual average crop yield Y in the region, which indicates that the learned mapping function F is correct.
[0065] (1) Multispectral feature fusion stage
[0066] The process of this stage is as follows Figure 1As shown on the left. The input data of our model is processed using the histogram algorithm to convert the remote sensing image of each band into a feature vector of length n. The original input remote sensing image has a total of b bands, so the reduced feature is of dimension [b,n], where b represents the number of bands and n represents the number of boxes, each box representing the pixel count within a specific intensity range. First, we apply average pooling and maximum pooling to the reduced features in n dimensions to obtain two sets of pooled vectors of dimension [b,1]. Then, these vectors are processed by a shared two-layer MLP, and the output of the MLP is normalized by the Softmax function to obtain the attention weights. Next, these attention weights are used to weight the original reduced input, that is, the output of the MLP layer is used as the attention weight, and is weighted and added to the feature vector obtained by the histogram algorithm to produce two new feature vectors of dimension [b,n]. Next, the two new feature maps are expanded in dimension and concatenated with the reduced-dimensional features to form a [3, b, n] feature map; specifically, the expansion includes two new feature vectors of dimension [b, n] each expanding in a new dimension to [1, b, n], and the reduced-dimensional features are upscaled and concatenated with the two new dimensions [1, b, n] to form a [3, b, n] feature map. Finally, the connected feature channels are fed into a two-dimensional convolution module (2D-Conv) for feature extraction, completing the operation of the multispectral feature fusion module. That is, after generating a new [3, b, n] feature map, the features are input into a 2D convolution module for feature extraction. Here, we use the newly generated dimensions as channels and input them into the convolution module. The size of the 2D convolution kernel is set to 3*3.
[0067] (2) Extraction of temporal characteristics of crop growth
[0068] This phase mainly consists of Figure 1 The crop growth temporal feature module in the right part is completed. After the multi-spectral feature fusion in the first stage, the output of the first stage is used as the input of the second stage. Here, multi-head self-attention is mainly used to extract the temporal features of the crop growth process from the crop growth features in different time periods. The self-attention mechanism allows the model to consider other parts of the input sequence when encoding a specific part of the input sequence. It calculates the weighted sum of all parts for each part of the input. The mechanism can be described as follows:
[0069] 1. Construction of query, key and value: For each part of the input sequence, use the learnable weight matrix W Q , W K and W V Construct a query (Q), key (K), and value (V):
[0070] Q=XWQ
[0071] K=XW K
[0072] V=XW v
[0073] 2. Calculate the score using the dot product of the query and all keys, then divide by the square root of the key dimension, and apply the Softmax function to each part of the input sequence to calculate the weight:
[0074]
[0075] 3. Multiply the weights obtained above by the value to get the output of the self-attention mechanism. In order to make the model more expressive, we use multi-head attention, which means that the above process will be performed in parallel in multiple different weight matrix subspaces, and finally the outputs of these subspaces will be spliced together.
[0076] When processing time series data, such as monitoring crop growth characteristics at different time periods, traditional RNN models may encounter certain limitations, such as long-term dependencies and execution efficiency issues. Our invention introduces learnable positional encoding to overcome these challenges. Positional encoding can capture temporal information in the input sequence, helping the model understand the growth of crops at different stages. Compared with the step-by-step calculation of RNN, positional encoding allows the model to operate in parallel on the entire sequence, thereby enhancing computational efficiency. Learnable positional encoding enables the model to adaptively learn temporal patterns in the sequence instead of relying on fixed mathematical functions.
[0077] The temporal characteristics of crop growth are as follows Figure 1 As shown in the process on the right, the fused time series features are finally input into a multi-layer perceptron MLP, and the output result is the yield prediction value.
[0078] The following is a system embodiment corresponding to the above method embodiment. This embodiment can be implemented in conjunction with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment. In order to reduce repetition, they are not repeated here. Accordingly, the relevant technical details mentioned in this embodiment can also be applied in the above embodiment.
[0079] The present invention also proposes a spatially adaptive crop yield prediction system, which includes:
[0080] The initial module is used to construct a spatial adaptive prediction network model including a multispectral feature fusion layer and a crop growth temporal feature extraction layer, and obtain a multispectral image sequence marked with actual crop yields as training data;
[0081] A training module is used to reduce the dimension of the features of each image in the training data through the multispectral feature fusion layer to obtain reduced dimension features, perform average pooling and maximum pooling on the reduced dimension features to obtain two groups of pooling vectors, obtain attention weights according to the pooling vectors, and use the attention weights to weight the reduced dimension features to generate two new feature maps, extract features of the new feature maps, and obtain crop growth features; input an input sequence consisting of crop growth features of each image in the training data into the crop growth temporal feature extraction layer, and the crop growth temporal feature extraction layer uses multi-head self-attention to extract the temporal features of the crop growth process from the input sequence in each time period, predict crop yield according to the temporal features, and obtain a crop yield prediction value; a loss function is constructed according to the crop yield prediction value and the actual crop yield to train the spatial adaptive prediction network model to obtain a crop yield prediction model;
[0082] The prediction module is used to input the spectral image sequence of the crop yield area to be predicted into the crop yield prediction model to obtain the crop yield prediction result of the area.
[0083] The spatially adaptive crop yield prediction system, wherein the training data in the initial module is a multispectral image sequence X={X 1 ,X 2 ,...,X t}, where each image X i ∈R b×h×w represents the multispectral image of the area at time t, where b is the number of bands of the remote sensing image, and h and w are the vertical and horizontal pixel counts, respectively;
[0084] The training module includes: converting the training data into a reduced dimension feature of [b,n] dimension, where b represents the number of bands, n represents the number of boxes, and each box represents a pixel count within a specific intensity range; applying average pooling and maximum pooling to the reduced dimension feature in n dimension to obtain two sets of pooling vectors of dimension [b,1]; the pooling vector is normalized after being processed by a shared two-layer MLP to obtain the attention weight, which is used to weight the reduced dimension feature to generate a feature map of dimension [b,n], and connect it with the reduced dimension feature to form a new feature map of [3,b,n], and input the new feature map into a 2D convolution module for feature extraction to obtain the crop growth feature.
[0085] The spatially adaptive crop yield prediction system, wherein the training module comprises:
[0086] The multi-head self-attention executes the self-attention mechanism in parallel in multiple different weight matrix subspaces, and splices the outputs of the self-attention mechanisms of all subspaces together to obtain the time series feature; wherein the self-attention mechanism is used to calculate the weighted sum of all parts for each part of the input sequence, and the self-attention mechanism includes:
[0087] For each part of the input sequence, a learnable weight matrix W is used Q , W K and W V Construct a query (Q), key (K), and value (V):
[0088] Q=XW Q
[0089] K=XW K
[0090] V=XW V
[0091] The score is calculated using the dot product of the query and all keys, then divided by the square root of the key dimension, and the Softmax function is applied to each part of the input sequence to calculate the weight Weight:
[0092]
[0093] Multiply the weight Weight by the value V to get the output of the self-attention mechanism.
[0094] The spatially adaptive crop yield prediction system, wherein the convolution kernel size of the 2D convolution module is 3*3.
[0095] The present invention also proposes a server, which includes the spatially adaptive crop yield prediction device.
[0096] The present invention also proposes a storage medium for storing a computer program for executing the spatially adaptive crop yield prediction method.
[0097] Although the embodiments of the present invention have been disclosed as above, they are not limited to the applications listed in the specification and the implementation modes, and they can be fully applied to various fields suitable for the present invention. For those familiar with the art, additional modifications can be easily implemented. Therefore, without departing from the general concept defined by the claims and the scope of equivalents, the present invention is not limited to the specific details and the illustrations shown and described herein.
Claims
1. A spatially adaptive crop yield prediction method, characterized in that: include: Step 1: construct a spatial adaptive prediction network model including a multispectral feature fusion layer and a crop growth temporal feature extraction layer, and obtain a multispectral image sequence marked with actual crop yields as training data; Step 2: The multispectral feature fusion layer reduces the dimension of each image in the training data to obtain reduced-dimensional features, performs average pooling and maximum pooling on the reduced-dimensional features to obtain two sets of pooling vectors, obtains attention weights according to the pooling vectors, and uses the attention weights to weight the reduced-dimensional features to generate two new feature maps, extracts the features of the new feature maps, and obtains crop growth features; Step 3, inputting an input sequence consisting of crop growth features of each picture in the training data into the crop growth temporal feature extraction layer, the crop growth temporal feature extraction layer extracts the temporal features of the crop growth process from the input sequence of each time period using multi-head self-attention, predicts the crop yield according to the temporal features, obtains the crop yield prediction value, constructs a loss function according to the crop yield prediction value and the actual crop yield, and trains the spatial adaptive prediction network model to obtain a crop yield prediction model; Step 4, inputting the spectral image sequence of the crop yield area to be predicted into the crop yield prediction model to obtain the crop yield prediction result of the area; The training data is a multispectral image sequence X={X 1 ,X 2 ,...,X t }, where each image X i ∈R b×h×w represents the multispectral image of the area at time t, where b is the number of bands of the remote sensing image, and h and w are the vertical and horizontal pixel counts, respectively; The step 2 includes: converting the training data into a reduced dimension feature of [b,n] dimension, where b represents the number of bands, n represents the number of boxes, and each box represents a pixel count within a specific intensity range; applying average pooling and maximum pooling to the reduced dimension feature in n dimension to obtain two sets of pooled vectors of dimension [b,1]; normalizing the pooled vectors after being processed by a shared two-layer MLP to obtain the attention weight, which is used to weight the reduced dimension feature to generate a feature map of dimension [b,n], and connecting it with the reduced dimension feature to form a new feature map of [3,b,n], and inputting the new feature map into a 2D convolution module for feature extraction to obtain the crop growth feature.
2. The spatially adaptive crop yield prediction method according to claim 1, characterized in that: Step 3 includes: The multi-head self-attention executes the self-attention mechanism in parallel in multiple different weight matrix subspaces, and splices the outputs of the self-attention mechanisms of all subspaces together to obtain the time series feature; wherein the self-attention mechanism is used to calculate the weighted sum of all parts for each part of the input sequence, and the self-attention mechanism includes: For each part of the input sequence, a learnable weight matrix W is used Q , W K and W V Construct a query (Q), key (K), and value (V): Q=XW Q K=XW K V=XW V The score is calculated using the dot product of the query and all keys, then divided by the square root of the key dimension, and the Softmax function is applied to each part of the input sequence to calculate the weight Weight: Multiply the weight Weight by the value V to get the output of the self-attention mechanism.
3. The spatially adaptive crop yield prediction method according to claim 1, characterized in that: The convolution kernel size of this 2D convolution module is 3*3.
4. A spatially adaptive crop yield prediction system, characterized in that: include: The initial module is used to construct a spatial adaptive prediction network model including a multispectral feature fusion layer and a crop growth temporal feature extraction layer, and obtain a multispectral image sequence marked with actual crop yields as training data; A training module is used to reduce the dimension of the features of each image in the training data through the multispectral feature fusion layer to obtain reduced dimension features, perform average pooling and maximum pooling on the reduced dimension features to obtain two groups of pooling vectors, obtain attention weights according to the pooling vectors, and use the attention weights to weight the reduced dimension features to generate two new feature maps, extract features of the new feature maps, and obtain crop growth features; input an input sequence consisting of crop growth features of each image in the training data into the crop growth temporal feature extraction layer, and the crop growth temporal feature extraction layer uses multi-head self-attention to extract the temporal features of the crop growth process from the input sequence in each time period, predict crop yield according to the temporal features, and obtain a crop yield prediction value; a loss function is constructed according to the crop yield prediction value and the actual crop yield to train the spatial adaptive prediction network model to obtain a crop yield prediction model; A prediction module, used for inputting a spectral image sequence of a crop yield area to be predicted into the crop yield prediction model to obtain a crop yield prediction result of the area; In the initial module, the training data is a multispectral image sequence X={X 1 ,X 2 ,...,X t }, where each image X i ∈R b×h×w represents the multispectral image of the area at time t, where b is the number of bands of the remote sensing image, and h and w are the vertical and horizontal pixel counts, respectively; The training module includes: converting the training data into a reduced dimension feature of [b,n] dimension, where b represents the number of bands, n represents the number of boxes, and each box represents a pixel count within a specific intensity range; applying average pooling and maximum pooling to the reduced dimension feature in n dimension to obtain two sets of pooling vectors of dimension [b,1]; the pooling vector is normalized after being processed by a shared two-layer MLP to obtain the attention weight, which is used to weight the reduced dimension feature to generate a feature map of dimension [b,n], and connect it with the reduced dimension feature to form a new feature map of [3,b,n], and input the new feature map into a 2D convolution module for feature extraction to obtain the crop growth feature.
5. The spatially adaptive crop yield prediction system according to claim 4, characterized in that: This training module includes: The multi-head self-attention executes the self-attention mechanism in parallel in multiple different weight matrix subspaces, and splices the outputs of the self-attention mechanisms of all subspaces together to obtain the time series feature; wherein the self-attention mechanism is used to calculate the weighted sum of all parts for each part of the input sequence, and the self-attention mechanism includes: For each part of the input sequence, a learnable weight matrix W is used Q , W K and W V Construct a query (Q), key (K), and value (V): Q=XW Q K=XW K V=XW V The score is calculated using the dot product of the query and all keys, then divided by the square root of the key dimension, and the Softmax function is applied to each part of the input sequence to calculate the weight Weight: Multiply the weight Weight by the value V to get the output of the self-attention mechanism.
6. The spatially adaptive crop yield prediction system according to claim 4, characterized in that: The convolution kernel size of this 2D convolution module is 3*3.
7. A server, characterized in that: A spatially adaptive crop yield prediction system comprising the system described in any one of claims 4-6.
8. A storage medium for storing a computer program for executing the spatially adaptive crop yield prediction method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Network model and device for plant identification, and electronic equipment
CN115471745A
Winter wheat yield prediction method, device and equipment based on multi-modal canopy image
CN116307105A