Estimation model construction method fusing remote sensing data and meteorological data

By constructing a yield estimation model based on deep feature fusion of multi-source data, the problem of feature mixing caused by modal differences between remote sensing data and meteorological data was solved, and the high accuracy and stability of winter wheat yield prediction were improved.

CN120635706APending Publication Date: 2025-09-12ZHENGZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510751105.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

When fusing remote sensing data with meteorological data, the existing winter wheat yield prediction model ignores the modal differences between the two, resulting in mixed feature expressions, affecting the model's prediction ability and making it difficult to fully mine their respective effective information.

Method used

A yield estimation model based on deep feature fusion of multi-source data is constructed, including a global spatiotemporal feature extraction module, a feature extraction module based on time series decomposition, and a cross-modal feature fusion module, which respectively extract the key features of remote sensing and meteorological data, and enhance the model's sensitivity to yield change trends through a cross-modal gated fusion strategy.

Benefits of technology

It has significantly improved the accuracy of winter wheat yield estimation and model stability, and enhanced adaptability to complex climate changes and prediction accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635706A_ABST
    Figure CN120635706A_ABST
Patent Text Reader

Abstract

The invention discloses a yield estimation model construction method fusing remote sensing data and meteorological data, and aims to effectively integrate multi-source data features to improve the yield estimation precision of winter wheat. According to the method, firstly, a winter wheat yield estimation remote sensing data set and a meteorological data set are constructed, and a solid data basis is provided for subsequent research. A multi-source data depth feature fusion yield estimation model is constructed, key feature representation of remote sensing data and meteorological data is comprehensively extracted, deep correlation between remote sensing features and meteorological features is effectively modeled, and therefore the sensitivity of the model to the yield change trend is enhanced. And finally, training and evaluating the model through the constructed data set, comparing the performance of the model by adopting three indexes of R, RMSE and MAE, and selecting an optimal yield estimation model. Experimental results show that the method is remarkably improved in the aspects of yield estimation precision and model stability, and a feasible technical scheme is provided for large-scale winter wheat yield estimation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of crop yield estimation, and in particular to a method for constructing a yield estimation model that integrates remote sensing data and meteorological data. Background Art

[0002] Food security is a core cornerstone of national stability and sustainable social development, and ensuring food supply has become a topic of great concern to the international community. Winter wheat, as one of my country's major grain crops, accounts for approximately one-fifth of the country's total grain output and plays a vital role in the country's food security system. Therefore, accurate pre-harvest forecasting of winter wheat yields is of great practical significance for optimizing agricultural production management, formulating scientific and rational food policies, and responding to climate change and extreme weather events.

[0003] In recent years, the rapid development of remote sensing technology has provided strong data support for crop growth monitoring and yield forecasting. With its wide coverage, high spatiotemporal resolution, and non-destructive monitoring capabilities, remote sensing data can effectively characterize crop growth status, assess the impact of pests and diseases, and be used for crop yield estimation. Remote sensing variables such as vegetation indices, spectral reflectance, thermal infrared data, and radar backscatter characteristics have been widely used in crop yield prediction research. Furthermore, breakthroughs in deep learning technology have provided a new research paradigm for agricultural data analysis. Compared to traditional machine learning methods, deep learning can automatically extract high-dimensional features through multi-layer neural networks, reducing manual intervention and thus improving the accuracy and generalization of data analysis. In the field of crop yield prediction, researchers have attempted to apply deep learning architectures such as convolutional neural networks, recurrent neural networks, and Transformers to remote sensing image analysis to enhance the accuracy of crop growth assessment and yield prediction. However, remote sensing data primarily reflects crop growth status and cannot directly represent the impact of environmental factors such as temperature and precipitation on crop yield. As a result, prediction models that rely solely on remote sensing data still have certain limitations when adapting to complex climate change.

[0004] To address this issue, recent research has shifted toward multi-source data fusion, leveraging the complementary information from different data sources to improve the stability and reliability of crop yield prediction. In particular, the integration of remote sensing and meteorological data has become a research hotspot for crop yield prediction. Meteorological data can provide detailed information about the crop growing environment, including key climatic factors such as temperature variations and precipitation patterns, which have a direct impact on crop yield. Therefore, incorporating meteorological data into remote sensing data not only addresses the shortcomings of remote sensing data in representing environmental variables but also enhances the predictive capabilities of the model, thereby improving the accuracy and stability of crop yield prediction. In recent years, researchers both domestically and internationally have made significant progress in multi-source data fusion. Tian et al. proposed a winter wheat yield prediction method based on LSTM, combining meteorological data with remote sensing indices and utilizing a multi-time-step input strategy to improve prediction accuracy. Experimental results show that this method outperforms both BPNN and SVM, demonstrating greater adaptability and robustness across diverse sampling regions and climate change conditions. Zhou et al. used random forest, support vector machine and LASSO methods to predict county-level yields in China's three major wheat-growing regions based on climate variables and remote sensing data. The results showed that moisture-related climate variables were better than temperature variables, and SVM had the best prediction performance, further verifying the effectiveness of multi-source data fusion in large-scale wheat yield prediction.

[0005] Current winter wheat yield estimation methods based on multi-source data primarily employ data-level fusion. This involves processing remote sensing and meteorological data to the same spatiotemporal resolution, then overlaying them and feeding them into a deep learning network for unified feature extraction and modeling. However, this fusion approach ignores the modal differences between remote sensing and meteorological data, resulting in mixed feature representations and impacting the model's predictive power. Remote sensing data primarily reflects crop spatial spectral characteristics, such as vegetation indices, spectral reflectance, and texture information, while meteorological data describes temporal variations in environmental factors such as precipitation, temperature, and sunshine duration. Due to significant differences in data structure, feature representation, and timescale between the two, directly using the same network for feature extraction can lead to interference between remote sensing and meteorological features, making it difficult for the model to fully exploit the effective information from each. Furthermore, under a unified network architecture, the model may favor learning patterns from one data type while ignoring important features from the other, limiting the accuracy and generalization of yield predictions. Therefore, how to construct an independent feature extraction network based on the different modal characteristics of remote sensing data and meteorological data, and adopt a reasonable feature fusion strategy to fully mine the information of multi-source heterogeneous data, is an important challenge facing current winter wheat yield estimation research. Summary of the Invention

[0006] The present invention aims to provide a method for constructing a yield estimation model that integrates remote sensing data and meteorological data, aiming to effectively integrate multi-source data features, thereby improving the accuracy of winter wheat yield estimation. Specifically, first, a remote sensing dataset for winter wheat yield estimation and a meteorological dataset for winter wheat yield estimation are constructed to provide a solid data foundation for subsequent research. Secondly, a multi-source data deep feature fusion yield estimation model is constructed. The multi-source data deep feature fusion yield estimation model consists of three modules: a global spatiotemporal feature extraction module, a feature extraction module based on time series decomposition, and a cross-modal feature fusion module. It aims to fully extract the key feature representations of remote sensing data and meteorological data, effectively model the deep correlation between remote sensing data features and meteorological data features, and enhance the model's sensitivity to yield change trends. Finally, the constructed dataset will be trained and evaluated through the constructed model using R 2 The optimal yield estimation model is selected by evaluating the three indicators of RMSE, MAE, and RMSE. Compared with existing methods, the proposed method significantly improves the yield estimation accuracy and model stability.

[0007] The present invention is achieved through the following technical solutions: The method for constructing a yield estimation model integrating remote sensing data and meteorological data comprises the following steps: Step S1: Construction of winter wheat yield estimation dataset This step mainly involves the acquisition of remote sensing data, meteorological data and crop yield data in the study area. Remote sensing data uses MODIS remote sensing images as the main data source, and the selected products include surface reflectance data MOD09A1 and land cover classification product MCD12Q1. Meteorological data comes from the ERA5-Land dataset, and the selected meteorological variables include average 2-meter temperature, maximum 2-meter temperature, minimum 2-meter temperature and precipitation. Crop yield data comes from the statistical yearbooks of various provinces, cities and counties. Next, all data are preprocessed: first, the remote sensing data are masked and cropped, and matched with the county-level yield data to construct a remote sensing data yield estimation dataset; then, the meteorological data are cropped and spatially weighted averaged, and also matched with the county-level yield data to construct a meteorological data yield estimation dataset. Finally, the constructed remote sensing yield estimation dataset and meteorological yield estimation dataset are divided into training set and test set; Step S2: Construction of a yield estimation model based on deep feature fusion of multi-source data The multi-source data deep feature fusion yield estimation model (MSDFFM) proposed in this paper consists of three main modules. The first is the global spatiotemporal feature extraction module (GSTFEN), which is used to extract spatial, spectral and temporal features from remote sensing images to accurately characterize the growth status information of crops. The second is the feature extraction module based on time series decomposition (TDFEM), which is used to extract trend terms and seasonal terms of meteorological data respectively, ensuring that meteorological features can effectively reflect environmental changes at different time scales, and thus capture the impact of long-term climate change and periodic environmental fluctuations on crop growth. Finally, the cross-modal feature fusion module (CMGFM) aims to give full play to the complementary advantages of remote sensing data and meteorological data. A cross-modal gated fusion module is designed, which combines the cross-attention mechanism with the gating mechanism to realize the dynamic fusion of remote sensing features and meteorological features. The cross-attention mechanism is used to capture the deep correlation between different modal features, while the gating mechanism effectively suppresses data redundancy and information interference by adaptively adjusting the contribution weights of each modal feature. Finally, through feature fusion, the model's sensitivity to yield change trends is enhanced, thereby improving the accuracy of winter wheat yield estimation; Step S3: Yield estimation model training and prediction The production estimation data set constructed in step S1 will be trained and evaluated through the model established in step S2, with the focus on selecting the optimal production estimation model and comparing and analyzing the performance of different models under multiple indicators. At this stage, it is first necessary to systematically screen the various models or parameter configurations obtained during the training process. The specific method is to evaluate them through a unified test set. The data in this test set has not participated in the training or parameter adjustment of any model, and its diversity and complexity can more truly reflect the generalization ability of the model in actual application scenarios. In order to verify the performance advantages of the production estimation model proposed in the present invention, it is compared and analyzed with many widely used production estimation models, and the coefficient of determination (R 2 ), root mean square error (RMSE) and mean absolute error (MAE) are used to evaluate the prediction accuracy of the model.

[0008] The specific steps of step S1 are: This paper uses MODIS remote sensing imagery as the primary data source, with selected products including surface reflectance data MOD09A1 and land cover classification data MCD12Q1. Data acquisition, processing, and export are all implemented on the Google Earth Engine platform using JavaScript. Meteorological data is derived from the ERA5-Land dataset, with selected meteorological variables including average temperature at 2 meters, maximum temperature at 2 meters, minimum temperature at 2 meters, and precipitation, obtained from the Copernicus climate data repository. Crop yield data is sourced from statistical yearbooks of various provinces, cities, and counties.

[0009] During the data preprocessing stage, the present invention ensures that remote sensing images can cover the entire growing period of winter wheat. When selecting MODIS images, the time range is set to early October each year to the end of June of the following year, and image data is extracted at intervals of 8 days to form a time series data of 30 time steps. During the data extraction process, the county is used as the basic unit, and the crop mask provided by the MCD12Q1 dataset is combined to identify the winter wheat planting area. Given the uncertainty of the MCD12Q1 mask data, it may not be possible to accurately define the planting range of winter wheat. Therefore, the data is further optimized to improve the accuracy of the image data.

[0010] This study selected the main winter wheat-producing areas as the research target. According to the China Agricultural and Rural Information Network, winter wheat planting areas have little overlap with soybean and rice distribution, effectively eliminating interference from other crops. Furthermore, considering winter wheat is a winter crop, its growth cycle overlaps only briefly with other major crops (such as corn, soybeans, and peanuts). Therefore, a dataset based on its specific growth cycle can be constructed, further minimizing the impact of other crops on the data and ensuring that remote sensing data more accurately reflects the growth status of winter wheat.

[0011] During the remote sensing data processing stage, the study area was screened based on the winter wheat yield data, and counties with low yields were eliminated to ensure the representativeness of the data. Subsequently, a 2×2 threshold was set on the GEE platform to eliminate areas with too few valid pixels, because when the number of winter wheat pixels is too low, the neural network model may find it difficult to effectively extract features, and may even mistakenly identify it as noise data, thereby affecting the model performance. Therefore, the standard size of the unified image block in the present invention is 32×32 pixels, and concentrated, contiguous areas with a large coverage area are given priority, and scattered planting areas are eliminated. When the number of winter wheat pixels is small, zero-value filling is used to ensure the integrity of the data. The remote sensing dataset finally constructed consists of image blocks of 30 time steps, and the data dimensions are (30, 32, 32, 7), where 30 is the time step, 32×32 is the spatial size, and 7 is the number of spectral bands.

[0012] To ensure that meteorological data fully covers the entire winter wheat growth period and is consistent with the time steps of remote sensing imagery, the present invention sets the image time of the first time step of each year's remote sensing data as the start time and the image time of the last time step as the end time when downloading meteorological data, thereby obtaining 240 days of meteorological data each year. During data processing, gridded meteorological data are clipped according to the county boundaries of China's major winter wheat-producing areas to ensure that the extracted data accurately corresponds to the study area. Subsequently, a spatially weighted average method is used to calculate county-level daily meteorological values, in which the meteorological data of each grid is weighted and summed according to its area to reflect the impact of different grids on the overall meteorological conditions of the county. When the areas of all grids are consistent, the arithmetic mean is directly taken for calculation. This method can effectively ensure the representativeness of county-level meteorological data, so that it accurately reflects the overall climate characteristics of the study area. Finally, a daily meteorological data time series from 2003 to 2022 was constructed based on the winter wheat growth cycle. The data dimensions for each year are (240, 4), where 240 is the time step and 4 is the number of characteristic variables.

[0013] Finally, the collected winter wheat yield data were matched with remote sensing data and meteorological data to construct a remote sensing dataset and meteorological dataset for winter wheat yield estimation from 2003 to 2022, with the data from 2003 to 2018 used as the training set and the data from 2019 to 2022 used as the test set.

[0014] The specific steps of step S2 are: The multi-source data deep feature fusion estimation model proposed in this paper mainly consists of three modules: a global spatiotemporal feature extraction module, a feature extraction module based on time series decomposition, and a cross-modal gated fusion module.

[0015] S2.1. First, we construct a global spatiotemporal feature extraction module. This module is primarily used to extract the spatial-spectral-temporal features of remote sensing data to accurately characterize crop growth status. This module consists of three components: a spatial-spectral feature extraction module, a coupled attention fusion module, and a temporal encoder module. The spatial-spectral feature extraction module utilizes a CNN-Transformer dual-branch architecture to extract local and global spatial spectral features from multispectral remote sensing imagery; the coupled attention fusion module dynamically adjusts and fuses features from the CNN and ViT; and the temporal encoder module learns the long-range temporal dependencies between the various growth stages of winter wheat in long-term imagery.

[0016] In the spatial-spectral feature extraction module, the CNN branch is used to extract local feature information of remote sensing data. This paper uses ResNet50 as the feature extraction backbone in the CNN branch. The Transformer branch is designed to extract global feature information of remote sensing data. This paper uses ViT as the backbone network in the Transformer branch. This process can be expressed as: F c =ResNet50(X) F t =ViT(X) Where X represents the input remote sensing data, F c represents the features extracted by the CNN branch, F t Represents the features extracted by the Transformer branch.

[0017] The coupled attention fusion module aims to aggregate local features from the CNN branch and global features from the ViT branch, which consists of a local-to-global interaction process and a global-to-local interaction process. The two processes can be expressed as: C t =Add(Softmax(Q c K t T )V t , F c ) T c =Add(Softmax(Q t K c T )V c , F t ) Where, F c and F t represents the local features from the CNN branch and the global features from the ViT branch. c , K c , V c and Q t , K t , V t Represents F c and F t The generated query matrix, key matrix and value matrix. C t and T c They represent the CNN features that integrate global feature information and the ViT features that integrate local feature information, respectively.

[0018] After the two processes of local-to-global interaction and global-to-local interaction, the CNN features that integrate global feature information and the ViT features that integrate local feature information are concatenated to obtain the final fusion feature. The specific expression is as follows: CT=Concat(C t , T c ) Where CT represents the fusion feature after fusing the local features and global features of remote sensing data.

[0019] The temporal encoder module, designed with reference to the Transformer encoder architecture, is used to model the correlation of the extracted spatial-spectral feature maps. This encoder receives a feature vector that compresses the spatial-spectral information as input. It then passes through a multi-head self-attention layer to calculate the attention coefficient of the image features between each two time steps, encoding the temporal information. Finally, a multi-layer perceptron layer optimizes the feature representation. The above process can be expressed as: X r =TemporalEncoder(CT) Where, X r Represents the spatial-spectral-temporal features extracted by the global spatiotemporal feature extraction module in remote sensing data.

[0020] S2.2. Build a feature extraction module based on time series decomposition. This module extracts trend and seasonal features from meteorological data to capture the impact of long-term climate change and cyclical environmental fluctuations on crop growth. This module consists of three components: a hybrid time series decomposition block, a seasonal encoder, a trend encoder, and a multilayer perceptron layer.

[0021] The hybrid time series decomposition block decomposes meteorological data to extract its trend and seasonal characteristics. This block uses several different kernels of Avgpool(·) to separate the different trend and seasonal components. To distinguish the contributions of different component features to the final result, each component feature is assigned a different weight and a weighted sum is performed. The specific expression of the hybrid time series decomposition block is as follows: X s =XX t Where, X t ,X s ∈R I×d , respectively represent the trend component and seasonal component after mixed decomposition, Represents the components extracted by different kernel functions, α k Indicates its weight.

[0022] The seasonal encoder and trend encoder share the same structural design, aiming to extract temporal features at different scales from seasonal and trend components to improve the accuracy of winter wheat yield estimation. The encoder consists of N stacked encoder layers, each containing two sublayers. The first sublayer is external attention, while the second is a simple fully connected feedforward network. Residual connections and layer normalization are applied to each sublayer to ensure gradient stability and accelerate model training convergence. The specific expression is: S=Seasonal Encoder(X s ) T=Trend Encoder(X t ) Where S represents the trend feature output after the trend component passes through N encoders, and T represents the seasonal feature output after the seasonal component passes through N encoders.

[0023] Then, the trend features and seasonal features are integrated through the multi-layer perceptron layer to obtain the feature representation of meteorological data. The specific expression is: X m =MLP(S+T) Where, X m Represents the meteorological data features extracted by the feature extraction module based on time series decomposition.

[0024] S2.3. Construct a cross-modal gated fusion module, which includes four key parts: feature alignment projection, cross-modal cross-attention interaction, gated feature fusion, and regression prediction.

[0025] This module first analyzes the remote sensing feature X r and meteorological characteristics X m Perform feature alignment projection and map them to the same latent space to eliminate the dimension mismatch problem. The specific expression is as follows: H r =W r X r +b r , H m =W m X m +b m Where H r , is the aligned feature, W r , W m is the linear projection weight, b r , b m is the bias term.

[0026] Subsequently, a cross-modal attention mechanism is adopted, with remote sensing features as queries (Q) and meteorological features as keys (K, V). The attention mechanism captures the regulatory effect of meteorological factors on remote sensing information and generates enhanced remote sensing features. The specific expression is as follows: Q=H r W Q , K=H m W K , V=H m W V H attn =Attention(Q, K, V) Where W Q , W K , W V is a learnable parameter, H attn It is a cross-modal attention feature.

[0027] Next, the contribution ratio of the original remote sensing features and the attention-enhanced features is adaptively adjusted through the gated fusion mechanism to construct a more robust fusion feature. The specific expression is as follows: G=σ(W g [H r ;H attn ]+b g ) H f =G⊙H r +(1-G)⊙H attn Where G is the gate weight (calculated by Sigmoid), W g is the fusion parameter, b g is the bias term, H f is a fusion feature. When meteorological conditions significantly affect crop growth, the gating coefficient G approaches 0, enhancing the role of climate regulation. When remote sensing phenotypic characteristics dominate, G approaches 1, preserving the integrity of the original remote sensing information.

[0028] Finally, the fused feature H f and meteorological characteristics m Further splicing and deep MLP regression are used to achieve the final prediction of winter wheat yield. The specific expression is as follows: H final =[H f ;H m ] Y=W2·ReLU(W1·LayerNorm(H final )+b1)+b2 Where W1 and W2 are the regression layer weights, b1 and b2 are biases, and Y is the final output prediction value.

[0029] The beneficial effects of the present invention are: In view of the different characteristics of remote sensing data and meteorological data, the present invention constructs a global spatiotemporal feature extraction network and a feature extraction network based on time series decomposition to extract the key features of remote sensing data and meteorological data, respectively. It also adopts a cross-modal gating fusion strategy to fuse the features of two different modalities, which can effectively improve the accuracy of winter wheat yield estimation and provide a feasible technical solution for large-scale winter wheat yield prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 It is a flow chart of the present invention; Figure 2 This is a structural diagram of the multi-source data deep feature fusion production estimation model of the present invention; Figure 3 This is a structural diagram of the global spatiotemporal feature extraction module of the present invention; Figure 4 This is a structural diagram of a feature extraction module based on time series decomposition of the present invention; Figure 5 This is a structural diagram of the cross-modal gating fusion module of the present invention; Figure 6 This is a scatter plot of the actual and predicted yields of winter wheat from the multi-source data deep feature fusion yield estimation model of the present invention and other different experiments. DETAILED DESCRIPTION

[0031] The present invention will be further described below with reference to the accompanying drawings and examples.

[0032] like Figure 1 As shown, the present invention provides a method for constructing a yield estimation model that integrates remote sensing and meteorological data. This method includes three steps: constructing a winter wheat yield estimation dataset, building a yield estimation model using deep feature fusion of multi-source data, and training and predicting the yield estimation model. Each step combines the characteristics of remote sensing and meteorological data to improve the accuracy, stability, and generalization of winter wheat yield estimation.

[0033] The method for constructing a yield estimation model integrating remote sensing data and meteorological data comprises the following steps: Step S1: Construction of winter wheat yield estimation dataset This step primarily involves acquiring remote sensing data, meteorological data, and crop yield data for the study area. Remote sensing data uses MODIS imagery as the primary source, with selected products including surface reflectance data (MOD09A1) and land cover classification product (MCD12Q1). Meteorological data comes from the ERA5-Land dataset, with selected meteorological variables including mean 2-meter temperature, maximum 2-meter temperature, minimum 2-meter temperature, and precipitation. Crop yield data is sourced from statistical yearbooks of provinces, cities, and counties. Next, all data are preprocessed: first, the remote sensing data is masked and cropped, and then matched with county-level yield data to construct a remote sensing data yield estimation dataset. Subsequently, the meteorological data is cropped and spatially weighted averaged, and similarly matched with county-level yield data to construct a meteorological data yield estimation dataset. Finally, these remote sensing and meteorological yield estimation datasets are divided into training and test sets. Based on the winter wheat planting area in each county, this study selected Henan, Shandong, Hebei, Shanxi, Anhui, Jiangsu, Yunnan, Sichuan, Shaanxi, Hubei, and Xinjiang as the research area, covering 784 counties. The study area covers the main winter wheat planting areas in my country to ensure the representativeness and regional applicability of the data.

[0034] The specific steps of step S1 are: This paper uses MODIS remote sensing imagery as the primary data source, with selected products including surface reflectance data MOD09A1 and land cover classification data MCD12Q1. Data acquisition, processing, and export are all implemented on the Google Earth Engine platform using JavaScript. Meteorological data is derived from the ERA5-Land dataset, with selected meteorological variables including average temperature at 2 meters, maximum temperature at 2 meters, minimum temperature at 2 meters, and precipitation, obtained from the Copernicus climate data repository. Crop yield data is sourced from statistical yearbooks of various provinces, cities, and counties.

[0035] During the data preprocessing stage, the present invention ensures that remote sensing images can cover the entire growing period of winter wheat. When selecting MODIS images, the time range is set to early October each year to the end of June of the following year, and image data is extracted at intervals of 8 days to form a time series data of 30 time steps. During the data extraction process, the county is used as the basic unit, and the crop mask provided by the MCD12Q1 dataset is combined to identify the winter wheat planting area. Given the uncertainty of the MCD12Q1 mask data, it may not be possible to accurately define the planting range of winter wheat. Therefore, the data is further optimized to improve the accuracy of the image data.

[0036] This study selected the main winter wheat-producing areas as the research target. According to the China Agricultural and Rural Information Network, winter wheat planting areas have little overlap with soybean and rice distribution, effectively eliminating interference from other crops. Furthermore, considering winter wheat is a winter crop, its growth cycle overlaps only briefly with other major crops (such as corn, soybeans, and peanuts). Therefore, a dataset based on its specific growth cycle can be constructed, further minimizing the impact of other crops on the data and ensuring that remote sensing data more accurately reflects the growth status of winter wheat.

[0037] During the remote sensing data processing stage, the study area was screened based on the winter wheat yield data, and counties with low yields were eliminated to ensure the representativeness of the data. Subsequently, a 2×2 threshold was set on the GEE platform to eliminate areas with too few valid pixels, because when the number of winter wheat pixels is too low, the neural network model may find it difficult to effectively extract features, and may even mistakenly identify it as noise data, thereby affecting the model performance. Therefore, the standard size of the unified image block in the present invention is 32×32 pixels, and concentrated, contiguous areas with a large coverage area are given priority, and scattered planting areas are eliminated. When the number of winter wheat pixels is small, zero-value filling is used to ensure the integrity of the data. The remote sensing dataset finally constructed consists of image blocks of 30 time steps, and the data dimensions are (30, 32, 32, 7), where 30 is the time step, 32×32 is the spatial size, and 7 is the number of spectral bands.

[0038] To ensure that meteorological data fully covers the entire winter wheat growth period and is consistent with the time steps of remote sensing imagery, the present invention sets the image time of the first time step of each year's remote sensing data as the start time and the image time of the last time step as the end time when downloading meteorological data, thereby obtaining 240 days of meteorological data each year. During data processing, gridded meteorological data are clipped according to the county boundaries of China's major winter wheat-producing areas to ensure that the extracted data accurately corresponds to the study area. Subsequently, a spatially weighted average method is used to calculate county-level daily meteorological values, in which the meteorological data of each grid is weighted and summed according to its area to reflect the impact of different grids on the overall meteorological conditions of the county. When the areas of all grids are consistent, the arithmetic mean is directly taken for calculation. This method can effectively ensure the representativeness of county-level meteorological data, so that it accurately reflects the overall climate characteristics of the study area. Finally, a daily meteorological data time series from 2003 to 2022 was constructed based on the winter wheat growth cycle. The data dimensions for each year are (240, 4), where 240 is the time step and 4 is the number of characteristic variables.

[0039] Finally, the collected winter wheat yield data were matched with remote sensing data and meteorological data to construct a remote sensing dataset and meteorological dataset for winter wheat yield estimation from 2003 to 2022, with the data from 2003 to 2018 used as the training set and the data from 2019 to 2022 used as the test set.

[0040] Step S2: Construction of a yield estimation model based on deep feature fusion of multi-source data like Figure 2 As shown in the figure, the proposed multi-source data deep feature fusion yield estimation model (MSDFFM) comprises a global spatiotemporal feature extraction module (GSTFEN), a time series decomposition-based feature extraction module (TDFEM), and a cross-modal feature fusion module (CMGFM). The global spatiotemporal feature extraction module (GSTFEN) extracts spatial, spectral, and temporal features from remote sensing imagery to accurately characterize crop growth status. The time series decomposition-based feature extraction module (TDFEM) extracts trend and seasonal features from meteorological data, ensuring that meteorological features effectively reflect environmental changes at different time scales and thus capturing the impact of long-term climate change and cyclical environmental fluctuations on crop growth. The cross-modal feature fusion module (CMGFM) aims to leverage the complementary advantages of remote sensing and meteorological data. A cross-modal gated fusion module is designed, combining a cross-attention mechanism with a gating mechanism to achieve dynamic fusion of remote sensing and meteorological features. The cross-attention mechanism captures deep correlations between features from different modalities, while the gating mechanism effectively suppresses data redundancy and information interference by adaptively adjusting the contribution weights of each modal feature. Ultimately, feature fusion enhances the model's sensitivity to yield trends, thereby improving the accuracy of winter wheat yield estimation.

[0041] The specific steps of step S2 are: The multi-source data deep feature fusion estimation model proposed in the present invention includes a global spatiotemporal feature extraction module, a feature extraction module based on time series decomposition, and a cross-modal gated fusion module.

[0042] S2.1, first build a global spatiotemporal feature extraction module, such as Figure 3 As shown in the figure, the global spatiotemporal feature extraction module is used to extract the spatial, spectral, and temporal features of remote sensing data to accurately characterize crop growth status. The global spatiotemporal feature extraction module consists of a spatial-spectral feature extraction module, a coupled attention fusion module, and a temporal encoder module. The spatial-spectral feature extraction module utilizes a CNN-Transformer dual-branch architecture to extract local and global spatial spectral features from multispectral remote sensing imagery. The coupled attention fusion module dynamically adjusts and fuses features from the CNN and ViT. The temporal encoder module learns the long-range temporal dependencies between various growth stages of winter wheat in long-term imagery.

[0043] In the spatial-spectral feature extraction module, the CNN branch is used to extract local feature information of remote sensing data. This paper uses ResNet50 as the feature extraction backbone in the CNN branch. The Transformer branch is designed to extract global feature information of remote sensing data. This paper uses ViT as the backbone network in the Transformer branch. This process can be expressed as: F c =ResNet50(X) F t =ViT(X) Where X represents the input remote sensing data, F c represents the features extracted by the CNN branch, F t Represents the features extracted by the Transformer branch.

[0044] The coupled attention fusion module aims to aggregate local features from the CNN branch and global features from the ViT branch, which consists of a local-to-global interaction process and a global-to-local interaction process. The two processes can be expressed as: C t =Add(Softmax(Q c K t T )V t , F c ) T c =Add(Softmax(Q t K c T )V c , F t ) Where, F c and F t represents the local features from the CNN branch and the global features from the ViT branch. c , K c , V c and Q t , K t , V t Represents F c and F t The generated query matrix, key matrix and value matrix. C t and T c They represent the CNN features that integrate global feature information and the ViT features that integrate local feature information, respectively.

[0045] After the two processes of local-to-global interaction and global-to-local interaction, the CNN features that integrate global feature information and the ViT features that integrate local feature information are concatenated to obtain the final fusion feature. The specific expression is as follows: CT=Concat(C t , T c ) Where CT represents the fusion feature after fusing the local features and global features of remote sensing data.

[0046] The temporal encoder module, designed with reference to the Transformer encoder architecture, is used to model the correlation of the extracted spatial-spectral feature maps. This encoder receives a feature vector that compresses the spatial-spectral information as input. It then passes through a multi-head self-attention layer to calculate the attention coefficient of the image features between each two time steps, encoding the temporal information. Finally, a multi-layer perceptron layer optimizes the feature representation. The above process can be expressed as: X r =TemporalEncoder(CT) Where, X r Represents the spatial-spectral-temporal features extracted by the global spatiotemporal feature extraction module in remote sensing data.

[0047] S2.2, build a feature extraction module based on time series decomposition, such as Figure 4 As shown in Figure 2, this module is used to extract trend and seasonal features from meteorological data, capturing the impact of long-term climate change and periodic environmental fluctuations on crop growth. This module consists of three components: a hybrid time series decomposition block, a seasonal encoder, a trend encoder, and a multilayer perceptron layer.

[0048] The hybrid time series decomposition block decomposes meteorological data to extract its trend and seasonal characteristics. This block uses several different kernels of Avgpool(·) to separate the different trend and seasonal components. To distinguish the contributions of different component features to the final result, each component feature is assigned a different weight and a weighted sum is performed. The specific expression of the hybrid time series decomposition block is as follows: X s =XX t Where, X t , X s ∈R I×d , respectively represent the trend component and seasonal component after mixed decomposition, Represents the components extracted by different kernel functions, α k Indicates its weight.

[0049] The seasonal encoder and trend encoder share the same structural design, aiming to extract temporal features at different scales from seasonal and trend components to improve the accuracy of winter wheat yield estimation. The encoder consists of N stacked encoder layers, each containing two sublayers. The first sublayer is external attention, while the second is a simple fully connected feedforward network. Residual connections and layer normalization are applied to each sublayer to ensure gradient stability and accelerate model training convergence. The specific expression is: S=Seasonal Encoder(X s ) T=Trend Encoder(X t ) Where S represents the trend feature output after the trend component passes through N encoders, and T represents the seasonal feature output after the seasonal component passes through N encoders.

[0050] Then, the trend features and seasonal features are integrated through the multi-layer perceptron layer to obtain the feature representation of meteorological data. The specific expression is: X m =MLP(S+T) Where, X m Represents the meteorological data features extracted by the feature extraction module based on time series decomposition. S2.3, build a cross-modal gating fusion module, such as Figure 5 As shown in the figure, the cross-modal gated fusion module includes four key parts: feature alignment projection, cross-modal cross-attention interaction, gated feature fusion, and regression prediction.

[0051] Cross-modal gated fusion module for remote sensing feature X r and meteorological characteristics X m Perform feature alignment projection and map them to the same latent space to eliminate the dimension mismatch problem. The specific expression is as follows: H r =W r X r +b r , H m =W m X m +b m Where H r , is the aligned feature, W r , W m is the linear projection weight, b r , b m is the bias term.

[0052] Subsequently, a cross-modal attention mechanism is adopted, with remote sensing features as queries (Q) and meteorological features as keys (K, V). The attention mechanism captures the regulatory effect of meteorological factors on remote sensing information and generates enhanced remote sensing features. The specific expression is as follows: Q=H r W Q , K=H m W K , V=H m W V H attn =Attention(Q, K, V) Where W Q , W K , W V is a learnable parameter, H attn It is a cross-modal attention feature.

[0053] Next, the contribution ratio of the original remote sensing features and the attention-enhanced features is adaptively adjusted through the gated fusion mechanism to construct a more robust fusion feature. The specific expression is as follows: G=σ(W g [H r ;H attn ]+b g ) H f =G⊙H r +(1-G)⊙H attn Where G is the gate weight (calculated by Sigmoid), W g is the fusion parameter, b g is the bias term, H f is a fusion feature. When meteorological conditions significantly affect crop growth, the gating coefficient G approaches 0, enhancing the role of climate regulation. When remote sensing phenotypic characteristics dominate, G approaches 1, preserving the integrity of the original remote sensing information.

[0054] Finally, the fused feature H f and meteorological characteristics m Further splicing and deep MLP regression are used to achieve the final prediction of winter wheat yield. The specific expression is as follows: H final =[H f ;H m ] Y=W2·ReLU(W1·LayerNorm(H final )+b1)+b2 Where W1 and W2 are the regression layer weights, b1 and b2 are biases, and Y is the final output prediction value.

[0055] Step S3: Yield estimation model training and prediction The production estimation data set constructed in step S1 will be trained and evaluated through the model established in step S2, with the focus on selecting the optimal production estimation model and comparing and analyzing the performance of different models under multiple indicators. At this stage, it is first necessary to systematically screen the various models or parameter configurations obtained during the training process. The specific method is to evaluate them through a unified test set. The data in this test set has not participated in the training or parameter adjustment of any model, and its diversity and complexity can more truly reflect the generalization ability of the model in actual application scenarios. In order to verify the performance advantages of the production estimation model proposed in the present invention, it is compared and analyzed with many widely used production estimation models, and the coefficient of determination (R 2 ), root mean square error (RMSE) and mean absolute error (MAE) are used to evaluate the prediction accuracy of the model.

[0056] The specific steps of step S3 are:

[0057] The yield estimation model proposed in this invention and the comparative yield estimation models are all implemented based on the PyTorch deep learning framework. During the experiment, the basic learning rate was set to 0.001, and the ReduceLROnPlateau mechanism was used for dynamic adjustment. The total number of training rounds was set to 100 rounds. All deep learning models used the Adam optimizer, in which the weight decay coefficient was set to 1e-5 and the batch size was set to 32. In addition, all experiments were trained in the environment of 64-bit Intel Core i7-6900K CPU and NVIDIA GeForce 1080 GPU (8GB video memory). The experiment used R 2 , RMSE and MAE are used to evaluate the estimation ability of the model. The calculation formula is as follows: Where N represents the number of samples in the crop dataset, Y i and represent the true value of output and its mean, respectively, and and Represent the model predicted output value and its mean respectively. For model evaluation, R 2 The closer the value is to 1, the smaller the RMSE and MAE values ​​are, and the higher the accuracy of the winter wheat yield estimation model is.

[0058] The specific experimental process is as follows: First, the effectiveness of the Global Spatiotemporal Feature Extraction Module (GSTFEN) and the Time Series Decomposition-based Feature Extraction Module (TDFEM) in extracting features from remote sensing and meteorological data is evaluated. Second, based on these evaluation results, the effectiveness of the multi-source data deep feature fusion yield estimation model is further evaluated.

[0059] In order to verify the effectiveness of the GSTFEN module designed by the present invention, the present invention removed the meteorological data feature extraction module and the cross-modal gated fusion module, and only retained the rest of the model to predict winter wheat yield. In addition, the present invention used the winter wheat yield estimation remote sensing dataset to conduct a comparative experiment on the GSTFEN module and four current mainstream remote sensing data yield estimation models (CNN, ViT, LSTM and CNN-LSTM). The experimental results are shown in Table 1. From the analysis of the data in the table, it can be seen that GSTFEN performed significantly better than other yield estimation methods in all years, with its annual average RMSE, MAE and R 2 The results reached 0.591 t / ha, 0.475 t / ha, and 0.848, respectively. Compared with CNN, LSTM, ViT, and CNN-LSTM, GSTFEN reduced the RMSE by 38.31%, 29.22%, 39.26%, and 16.05%, respectively, and the MAE by 37.42%, 28.79%, 39.10%, and 15.78%, respectively, significantly improving prediction accuracy. This result demonstrates that the GSTFEN module can effectively and deeply mine information reflecting winter wheat growth characteristics from remote sensing data and accurately grasp the impact of these characteristics on winter wheat yield, thereby significantly improving the accuracy of yield estimation. Table 1 GSTFEN module effectiveness evaluation

[0060] In order to verify the effectiveness of the TDFEM module designed by the present invention, the remote sensing data feature extraction module and the cross-modal gated fusion module were removed, and only the other parts of the model were retained for winter wheat yield prediction. In addition, the present invention selected a meteorological data set for winter wheat yield estimation and conducted a comparative experiment on the TDFEM module with three current mainstream meteorological data yield estimation models (LSTM, GRU and Transformer). The experimental results are shown in Table 2. As can be seen from the data in the table, TDFEM performed well in all evaluation indicators, with its annual average RMSE, MAE and R 2The TDFEM model achieved 0.612 t / ha, 0.489 t / ha, and 0.838 t / ha, respectively, outperforming all compared models. Compared with the LSTM, GRU, and Transformer models, the TDFEM model achieved RMSE reductions of 30.84%, 29.57%, and 8.38%, respectively, and MAE reductions of 30.93%, 28.65%, and 7.74%, respectively. These results demonstrate that the TDFEM model can effectively extract meteorological features and accurately characterize the impact of meteorological factors on winter wheat yield, thereby improving forecast accuracy. Table 2 TDFEM module effectiveness evaluation

[0061] Finally, to verify the effectiveness of the proposed MSDFFM model, this section designed four comparative experiments to evaluate the impact of single-source versus multi-source data fusion on forecasting performance. These four groups of experiments are as follows: 1) Group R: Yield forecasting using GSTFEN using only remote sensing data; 2) Group M: Yield forecasting using TDFEM using only meteorological data; 3) Group R+M (Concat): Replacing the CMGFM module in MSDFFM with Concat splicing to evaluate the impact of directly concatenating features on forecasting performance; and 4) Group R+M (MSDFFM): Yield forecasting using the MSDFFM model to evaluate the impact of the CMGFM module on forecasting performance.

[0062] The experimental results are shown in Table 3. From the analysis of the data in the table, it can be seen that the R+M (MSDFFM) group shows excellent performance in all evaluation indicators. Its annual average RMSE, MAE and R 2 The results were 0.508t / ha, 0.408t / ha and 0.888 respectively. Compared with the other experimental groups (R group, M group and R+M (Concat) group), the RMSE of the R+M (MSDFFM) group was reduced by 14.04%, 16.99% and 8.14%, respectively, and the MAE was reduced by 14.11%, 16.22% and 7.9%, respectively. This result shows that the multi-source data deep feature fusion model (MSDFFM) proposed in this paper can effectively explore the deep correlation between remote sensing data and meteorological data, enhance the model's perception of the dynamic changes in winter wheat yield, and thus significantly improve the accuracy of yield prediction. Table 3 Comparison of yield estimation performance in different experiments

[0063] like Figure 6The figure shows a scatter plot of actual and predicted winter wheat yields for the multi-source data deep feature fusion yield estimation model presented in this paper and other experiments. The scatter plot reveals that the other yield estimation models have large prediction errors in high-yield and low-yield areas, while the R+M (MSDFFM) predictions are highly consistent with actual yields, with the majority of data points distributed near the 1:1 line, with only a few discrete points. This demonstrates that the R+M (MSDFFM) model exhibits a higher correlation coefficient and outperforms other models, indicating that it can better capture the changing trends in winter wheat yield.

Claims

1. A method for constructing a yield estimation model integrating remote sensing data and meteorological data, characterized in that: The following steps are involved: Step S1: constructing a winter wheat yield estimation dataset for future use; Step S2: Constructing a yield estimation model based on deep feature fusion of multi-source data: The multi-source data deep feature fusion yield estimation model MSDFFM includes a global spatiotemporal feature extraction module GSTFEN, a time series decomposition-based feature extraction module TDFEM, and a cross-modal feature fusion module CMGFM. The global spatiotemporal feature extraction module GSTFEN is used to extract the spatial-spectral-temporal features of remote sensing images to accurately characterize the growth status of crops. The time series decomposition-based feature extraction module TDFEM is used to extract trend and seasonal features of meteorological data, respectively, to ensure that meteorological features can effectively reflect environmental changes at different time scales, so as to capture the impact of long-term climate change and periodic environmental fluctuations on crop growth. The cross-modal feature fusion module CMGFM combines the cross-attention mechanism with the gating mechanism to achieve dynamic fusion of remote sensing and meteorological features. The cross-attention mechanism is used to capture the deep correlation between different modal features, while the gating mechanism effectively suppresses data redundancy and information interference by adaptively adjusting the contribution weight of each modal feature. Ultimately, the fused features can enhance the model's sensitivity to yield change trends, thereby improving the accuracy of winter wheat yield estimation. Step S3: Production estimation model training and prediction: The yield estimation dataset constructed in step S1 will be trained and evaluated by the model established in step S2.

2. The method for constructing a yield estimation model integrating remote sensing data and meteorological data according to claim 1, characterized in that: In step S1, remote sensing data, meteorological data, and crop yield data within the study area are obtained, and the remote sensing data, meteorological data, and crop yield data are preprocessed. A remote sensing data yield estimation dataset is constructed by masking and cropping the remote sensing data and matching it with county-level yield data. A meteorological data yield estimation dataset is constructed by cropping, spatially weighted averaging, and matching the meteorological data with county-level yield data. Subsequently, the remote sensing yield estimation dataset and the meteorological yield estimation dataset are divided into a training set and a test set.

3. The method for constructing a yield estimation model integrating remote sensing data and meteorological data according to claim 1, characterized in that: The specific steps of step S2 are: S2.

1. Construct a global spatiotemporal feature extraction module. This module is used to extract the spatial-spectral-temporal features of remote sensing data to accurately characterize crop growth. The module includes a spatial-spectral feature extraction module, a coupled attention fusion module, and a temporal encoder module. The spatial-spectral feature extraction module uses a CNN-Transformer dual-branch structure to extract local and global spatial spectral features from multispectral remote sensing images. A coupled attention fusion module is used to dynamically adjust and fuse features from CNN and ViT; a temporal encoder module is used to learn the long-range temporal dependencies between different stages of the winter wheat growth period in long time series images; S2.

2. Construct a feature extraction module based on time series decomposition. This module is used to extract trend and seasonal features of meteorological data, respectively, to capture the impact of long-term climate change and periodic environmental fluctuations on crop growth. The feature extraction module based on time series decomposition includes a hybrid time series decomposition block, a seasonal encoder, a trend encoder, and a multi-layer perceptron layer. S2.

3. Construct a cross-modal gated fusion module, which includes feature alignment projection, cross-modal cross-attention interaction, gated feature fusion, and regression prediction. The cross-modal gated fusion module first performs the remote sensing feature X r and meteorological characteristics X m Perform feature alignment projection and map them to the same latent space to eliminate the dimension mismatch problem; the specific expression is as follows: H r =W r X r +b r ,H m =W m X m +b m Where, is the aligned feature, W r ,W m is the linear projection weight, b r ,b m is the bias term; Subsequently, a cross-modal attention mechanism is adopted, with remote sensing features as queries (Q) and meteorological features as keys (K, V). The attention mechanism is used to capture the regulatory effect of meteorological factors on remote sensing information and generate enhanced remote sensing features. The specific expression is as follows: Q=H r W Q ,K=H m W K ,V=H m W V H attn =Attention(Q,K,V) Where W Q ,W K ,W V is a learnable parameter, H attn It is a cross-modal attention feature; The contribution ratio of the original remote sensing features and the attention-enhanced features is adaptively adjusted through the gated fusion mechanism to construct a more robust fusion feature. The specific expression is as follows: G=σ(W g [H r ;H attn ]+b g ) H f =G⊙H r +(1-G)⊙H attn Where G is the gate weight, calculated by Sigmoid, and W g is the fusion parameter, b g is the bias term, H f is a fusion feature; when meteorological conditions have a significant impact on crop growth, the gating coefficient G approaches 0, enhancing the role of climate regulation characteristics; when remote sensing phenotypic characteristics dominate, G approaches 1, preserving the integrity of the original remote sensing information; Finally, the fused feature H f and meteorological characteristics m Further splicing and deep MLP regression are used to achieve the final prediction of winter wheat yield; the specific expression is as follows: H final =[H f ;H m ] <h2 style=";text-align:left;direction:ltr">Y = W2 ReLU(W1 LayerNorm(H<h2 style=";text-align:left;direction:ltr"> final <h2 style=";text-align:left;direction:ltr"> )+b1)+b2 Where W1 and W2 are the regression layer weights, b1 and b2 are biases, and Y is the final output prediction value.

4. The method for constructing a yield estimation model integrating remote sensing data and meteorological data according to claim 3, characterized in that: In S2.1, in the spatial spectrum feature extraction module, the CNN branch is used to extract local feature information of remote sensing data; ResNet50 is used as the feature extraction backbone in the CNN branch; and the Transformer branch is designed to extract global feature information of remote sensing data; ViT is used as the backbone network in the Transformer branch; the process is expressed as: F c =ResNet50(X) F t =ViT(X) Where X represents the input remote sensing data, F c represents the features extracted by the CNN branch, F t Represents the features extracted by the Transformer branch; The coupled attention fusion module aims to aggregate local features from the CNN branch and global features from the ViT branch. It consists of a local-to-global interaction process and a global-to-local interaction process. The two processes are expressed as: C t =Add(Softmax(Q t K t T )V t ,F c ) T c =Add(Softmax(Q t K c T )V c ,F t ) Where, F c and F t represents the local features from the CNN branch and the global features from the ViT branch; Q c , K c , V c and Q t , K t , V t Respectively represent F c and F t The generated query matrix, key matrix and value matrix; C t and T c They represent CNN features that integrate global feature information and ViT features that integrate local feature information respectively; After the two processes of local-to-global interaction and global-to-local interaction, the CNN features that integrate global feature information and the ViT features that integrate local feature information are concatenated to obtain the final fusion feature. The specific expression is as follows: CT=Concat(C t ,T c ) Where CT represents the fusion feature after fusing the local features and global features of remote sensing data; The temporal encoder module is designed based on the Transformer encoder architecture and is used to model the correlation of the extracted spatial-spectral feature maps. The encoder receives a feature vector that compresses the spatial-spectral information as input. Then, a multi-head self-attention layer calculates the attention coefficient of the image features between each two time steps to encode the temporal information. Finally, a multi-layer perceptron layer optimizes the feature representation. The above process can be expressed as: X r =TemporalEncoder(CT) Where, X r Represents the spatial-spectral-temporal features extracted by the global spatiotemporal feature extraction module in remote sensing data.

5. The method for constructing a yield estimation model integrating remote sensing data and meteorological data according to claim 3, characterized in that: In S2.2, the hybrid time series decomposition block is used to decompose meteorological data to extract its trend and seasonal characteristics; The hybrid time series decomposition block uses several different kernels of Avgpool(·) to separate different trend components and seasonal components. At the same time, in order to distinguish the contribution of different component features to the final result, different weights are assigned to each component feature and a weighted sum is performed. The specific expression of the hybrid time series decomposition block is as follows: X s =X-X t Where, X t ,X s ∈R I×d , respectively represent the trend component and seasonal component after mixed decomposition, Represents the components extracted by different kernel functions, α k Indicates its weight; The seasonal encoder and trend encoder use the same structural design, aiming to extract temporal features of different scales from seasonal and trend components to improve the accuracy of winter wheat yield estimation. The encoder consists of N stacked encoder layers, each containing two sublayers. The first sublayer is external attention, while the second sublayer is a simple fully connected feedforward network. Residual connections and layer normalization are applied to each sublayer to ensure gradient stability and improve the convergence speed of model training. The specific expression is: S=Seasonal Encoder(X s ) T=Trend Encoder(X t ) Where S represents the trend feature output by the trend component after passing through N encoders, and T represents the seasonal feature output by the seasonal component after passing through N encoders; Then, the trend features and seasonal features are integrated through a multi-layer perceptron layer to obtain the feature representation of meteorological data; The specific expression is: X m =MLP(S+T) Where, X m Represents the meteorological data features extracted by the feature extraction module based on time series decomposition.