MODIS Land Surface Temperature Compensation Method Based on Spatiotemporal Feature Fusion
Through the MODIS surface temperature complement method based on spatiotemporal feature fusion, the feature information of one-dimensional time series and three-dimensional spatiotemporal sequence is used to extract and fusion using neural network models, solving the problem of low complement accuracy in the existing technology and achieving higher complement accuracy and robustness.
Patent Information
- Application Number
- CN202411930959.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-12-26
AI Technical Summary
The existing surface temperature data complementary technology has limitations and fails to effectively capture the complexity of the data, especially when processing sparse data and long time series, resulting in inaccurate complementary results.
The MODIS surface temperature complement method based on spatiotemporal feature fusion is adopted. Through one-dimensional time series and three-dimensional spatiotemporal sequence as input features, the neural network model is used for feature extraction and fusion, including the 1D-CNN-BiNbeats module and the 3D-CNN module, combined with the spatiotemporal interactive fusion module and the multi-dimensional feature synthesis module to obtain the complement value results.
It significantly improves the complement value accuracy, can more effectively capture nonlinear features and long-term dependencies in the data, improves the accuracy and robustness of complement value results, especially when processing long-term series data.
Smart Images

Figure CN119377891B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of surface temperature data reconstruction, and particularly relates to a method for filling missing values of MODIS surface temperature based on spatio-temporal feature fusion. Background Art
[0002] Surface temperature is an important indicator of the Earth's surface characteristics, directly affecting multiple fields such as climate change, ecosystem, crop growth, and water resource management. The change of surface temperature not only reflects the dynamic changes of the natural environment but also is closely related to human activities. MODIS can obtain various important parameters such as surface temperature, vegetation cover, ocean, and climate on a global scale with a daily frequency, having a relatively high spatial resolution (about 250 meters to 1 kilometer) and temporal resolution.
[0003] MODIS surface temperature data provides rich spatio-temporal information and is widely used in fields such as climate change research, ecological monitoring, agricultural production, and urban heat island effect. However, due to factors such as climate conditions, sensor errors, and cloud cover, there are often missing values or outliers in MODIS data, which poses challenges to data analysis and application. Therefore, it is particularly important to perform effective filling processing on MODIS data.
[0004] There are multiple deficiencies in the current filling techniques for surface temperature data, which to a certain extent restrict the filling accuracy and efficiency. The main defects of the existing technologies are as follows:
[0005] (1) Limitations of traditional filling methods: Traditional filling techniques usually assume that data is spatially continuous when dealing with sparse data. However, this assumption does not always hold, especially in the presence of obvious noise or missing values. Traditional methods fail to effectively capture the complexity in the data, resulting in a large deviation between the filling results and the actual situation.
[0006] (2) Deficiencies of simple models: Many filling models based on simple neural networks (such as basic RNN or 1D CNN) fail to fully capture the complexity of time series data. Such basic neural networks often encounter problems of gradient vanishing or explosion in processing long sequence data, leading to poor learning effects of the model in a relatively long time range. This limitation makes the performance of the model far from satisfactory when filling in a long time scale, affecting the reliability of filling.
[0007] (3) Lack of spatio-temporal information fusion: Many current value filling methods fail to effectively utilize spatial information when dealing with geographical data. Traditional time series models often only focus on changes in the time dimension and ignore the temperature changes of surrounding pixels. For example, when filling in the MODIS land surface temperature data, simple time series models do not take into account the influence of other pixels in the neighborhood. This way of lacking spatial information fusion leads to unreasonable value filling results.
[0008] Therefore, there is an urgent need for a MODIS land surface temperature value filling method based on spatio-temporal feature fusion to solve the above problems. Summary of the Invention
[0009] To solve the above technical problems, the present invention proposes a MODIS land surface temperature value filling method based on spatio-temporal feature fusion. This method uses one-dimensional time series and three-dimensional spatio-temporal series as input features, fully considering the characteristic information of land surface temperature in time and space. The model consists of two parallel branch modules and two multi-dimensional feature fusion modules. Each branch module is used to process different types of data and has many significant advantages.
[0010] The present invention provides a MODIS land surface temperature value filling method based on spatio-temporal feature fusion, including:
[0011] Obtain the data to be filled in;
[0012] Input the data to be filled in into the spatio-temporal feature fusion network model to obtain the filling result. Among them, the spatio-temporal feature fusion network model is obtained by training with a training set. The training set is historical land surface temperature data. The spatio-temporal feature fusion network model is constructed using a neural network. The neural network is used to extract one-dimensional time series features and three-dimensional spatio-temporal series features, and fuse the one-dimensional time series features and three-dimensional spatio-temporal series features to obtain the filling result.
[0013] Optionally, obtaining the training set includes:
[0014] Obtain the original land surface temperature data;
[0015] Perform splicing, cropping, and reprojection processing on the original land surface temperature data to obtain the training set.
[0016] Optionally, the spatio-temporal feature fusion network model includes: a one-dimensional time series feature extraction module, a three-dimensional spatio-temporal series feature extraction module, a spatio-temporal interaction fusion module, and a multi-dimensional feature synthesis module;
[0017] The one-dimensional time series feature extraction module is used to extract the one-dimensional feature information of the land surface temperature;
[0018] The three-dimensional spatio-temporal sequence feature extraction module is used to extract three-dimensional feature information of surface temperature;
[0019] The spatio-temporal interaction fusion module is used to fuse the one-dimensional feature information and the three-dimensional feature information;
[0020] The multi-dimensional feature synthesis module is used to synthesize the fused one-dimensional feature information and three-dimensional feature information.
[0021] Optionally, the extraction of the one-dimensional feature information of surface temperature includes:
[0022] Using a 1D-CNN unit to extract one-dimensional time series features from the data to be filled;
[0023] Using a bidirectional N-Beats unit to capture the basic pattern of the one-dimensional time series features and learn long-term dependencies;
[0024] According to the long-term dependencies, obtain the one-dimensional feature information.
[0025] Optionally, the N-Beats unit includes: an input layer, a basic block, a stacking layer, and an output layer;
[0026] The input layer is used to input the one-dimensional time series features;
[0027] The basic block is used to capture the information flow in the time series from the past to the present, extract the forward dependence of the sequence, and capture the reverse information flow from the future to the present, extract the backward dependence of the sequence, and obtain a prediction result according to the forward dependence and the backward dependence;
[0028] The stacking layer is used to correct the prediction result;
[0029] The output layer is used to output the corrected prediction result.
[0030] Optionally, the three-dimensional spatio-temporal sequence feature extraction module includes: an input layer, a convolutional layer, an activation layer, a pooling layer, and an output layer;
[0031] The input layer is used to input the neighborhood spatio-temporal sequence data of the missing value points;
[0032] The convolutional layer is used to extract three-dimensional features of the spatio-temporal sequence data through three-dimensional convolution operations;
[0033] The activation layer is used to perform non-linear processing on the three-dimensional features using an activation function;
[0034] The pooling layer is used to reduce the feature dimension and extract representative three-dimensional features;
[0035] The output layer is used to output the three-dimensional feature information.
[0036] Optionally, fusing the one-dimensional feature information and the three-dimensional feature information includes:
[0037] Unifying the feature dimensions of the one-dimensional feature information and the three-dimensional feature information through a convolution operation to obtain dimension-unified features;
[0038] Fusing the dimension-unified features to obtain fused feature information.
[0039] Optionally, performing a reduction process on the fused one-dimensional feature information and three-dimensional feature information;
[0040] Performing a dimensionality reduction operation on the reduced one-dimensional feature information and three-dimensional feature information to extract the reduced one-dimensional feature and three-dimensional feature;
[0041] Mapping the reduced one-dimensional feature and three-dimensional feature to obtain a one-dimensional mapped feature and a three-dimensional mapped feature;
[0042] Combining the one-dimensional mapped feature and the three-dimensional mapped feature to obtain a complementary value result.
[0043] Optionally, when training the spatio-temporal feature fusion network model, use the Adam optimization algorithm to update the parameters of the spatio-temporal feature fusion network model.
[0044] Compared with the prior art, the present invention has the following advantages and technical effects:
[0045] (1) Improving the complementary value accuracy: By using the 1DCNN-BiNbeats module to process one-dimensional time series data, the present invention can more effectively capture the non-linear features and long-term dependence relationships in the data. The design of this model enables it to perform excellently in processing long time series data, thereby significantly improving the accuracy of the complementary value. Especially when performing complementary value on MODIS land surface temperature data, this model can better fit the actual temperature change trend, making the complementary value result more reliable.
[0046] (2) Effectively integrating spatio-temporal information: The model used in the present invention takes one-dimensional time series and three-dimensional spatio-temporal series as inputs, and has two input processing modules, namely the 1DCNN-BiNbeats module and the 3D-CNN module, fully considering the spatio-temporal characteristics of land surface temperature data. The model not only considers the one-dimensional time series data of the target point, but also integrates the data within its surrounding neighborhood. In this way, the complementary value result is more reasonable, can reflect the influence of the surrounding environment on the target point, and improves the accuracy and coherence of the complementary value result.
[0047] (3) Feature fusion and restoration ability: The uniqueness of the present invention lies in its adoption of a multi-dimensional feature fusion and restoration strategy during the value filling process. During the value filling process, the three-dimensional features extracted by the 3D-CNN module and the one-dimensional features extracted by the 1DCNN-BiNbeats module are fused and restored through the spatio-temporal interaction fusion module, which more effectively extracts important information in the time domain and spatial domain. A loss function will be generated during this process. This way of feature fusion and restoration enhances information exchange between modules and reduces information loss, ultimately improving the accuracy of the value filling result. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The drawings forming a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:
[0049] Figure 1 is a flowchart of the MODIS land surface temperature filling method based on spatio-temporal feature fusion according to an embodiment of the present invention;
[0050] Figure 2 is a schematic diagram of the spatio-temporal feature fusion network according to an embodiment of the present invention;
[0051] Figure 3 is a structural diagram of the spatio-temporal interaction fusion module according to an embodiment of the present invention;
[0052] Figure 4 is a structural diagram of the multi-dimensional feature synthesis module according to an embodiment of the present invention;
[0053] Figure 5 is a diagram of the land surface temperature filling result in Yutian area according to an embodiment of the present invention;
[0054] Figure 6 is a diagram of the land surface temperature filling result in Wenchuan area according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine with the embodiments to detail this application.
[0056] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from here.
[0057] The following explains the professional terms involved in the embodiments:
[0058] MODIS Data: MODIS (Moderate Resolution Imaging Spectroradiometer) is a sensor on an Earth observation satellite launched by NASA (National Aeronautics and Space Administration). It can capture spectral images of the Earth's surface at moderate resolution. MODIS provides multispectral data covering the atmosphere, land, and ocean, and is widely used in areas such as environmental monitoring, climate research, vegetation analysis, and disaster assessment. Its data can be used for spatio-temporal series analysis to help scientists and decision-makers better understand the Earth system and its changes.
[0059] Deep Learning: Deep learning is a branch of machine learning that uses multi-layer neural networks to automatically extract and learn features from data by simulating the working mode of human brain neurons. It can process large-scale datasets and is widely applied in fields such as image recognition, natural language processing, and speech recognition. Deep learning optimizes model parameters through the backpropagation algorithm, enabling it to perform excellently in complex tasks and thus outperforming traditional machine learning methods in many application scenarios.
[0060] CNN: Convolutional Neural Network (CNN) is a deep learning model particularly suitable for image processing and computer vision tasks. It extracts features through convolutional layers, reduces the number of parameters using local receptive fields and weight sharing, and improves computational efficiency. CNN usually consists of multiple convolutional layers, pooling layers, and fully connected layers, and can automatically learn and identify important patterns and features in images. It is widely used in fields such as image classification, object detection, and image generation.
[0061] N-BEATS: N-BEATS (Neural Basis Expansion Analysis for Time Series) is a deep learning-based time series prediction model. It processes time series data through a multi-layer fully connected network and uses a block structure to capture different patterns. N-BEATS does not rely on prior knowledge and can automatically learn the trends and seasonality of time series. This model generates predictions through a feedforward neural network and optimizes performance using a regression loss function. It is suitable for various time series prediction tasks, showing excellent performance and high flexibility.
[0062] The existing LST reconstruction methods mainly include the following categories: reconstruction methods based on spatial information, reconstruction methods based on temporal information, and methods that integrate spatio-temporal information. Spatial information reconstruction realizes interpolation according to the spatial correlation between missing pixels and their adjacent clear-sky pixels, such as methods like spline function, regression tree analysis, inverse distance weighting, and Kriging interpolation. Such methods do not require additional reference information and are easy to implement. However, due to being greatly affected by surface cover changes and surface elevation changes, there are problems such as unclear reconstructed images and poor accuracy in this type of method. Temporal reconstruction, also known as multi-temporal information reconstruction or time series information reconstruction, its principle is to reconstruct missing pixels using the land surface temperature of adjacent dates at the same location, mainly including: harmonic analysis method, singular spectrum analysis method, method based on daily temperature cycle model, method based on physical model, method based on robust regression of multi-temporal information, method based on SG filtering, etc. Although the reconstruction method based on temporal filtering can smooth the daily variation of land surface temperature, when the missing value data in adjacent periods increases, the reconstruction effect will deteriorate. The most basic strategy of the reconstruction method that integrates spatio-temporal information is to apply spatial and temporal reconstruction methods in sequence: generally, filling values in the time domain first, and then filling values in the spatial domain, so as to realize the reconstruction of land surface temperature data. Although this type of method comprehensively considers time and space information during the reconstruction process, it does not achieve true spatio-temporal joint modeling, which restricts the improvement of the accuracy of the reconstruction results.
[0063] In addition to the above methods, in recent years, reconstruction methods based on machine learning have received extensive attention and been widely used by scholars due to their strong learning ability, high robustness, etc., such as artificial neural networks, deep convolutional networks, and long short-term memory networks. Nowadays, in the research of time series prediction and filling, the most commonly used are the CNN model and the LSTM model. CNN can automatically extract spatial features in data through convolutional layers and is suitable for multi-dimensional data such as images and time series. Due to the local connection and shared weights of convolutional operations, CNN is usually more efficient than fully connected layers, especially when dealing with large-scale data. However, when dealing with time series data, CNN is difficult to capture the long-term dependencies of the data, especially when the sequence is long. The LSTM model is designed specifically to capture long-term dependencies and can effectively handle the long-term relationships in time series data. LSTM performs well when dealing with rapidly changing or irregular data and is suitable for dynamic and complex sequence data. However, the recursive structure of the LSTM model makes the training and inference processes relatively slow, especially for long sequence data, consuming a large amount of computing resources. Due to its recursive nature, LSTM is difficult to perform parallel computing during training, which limits its scalability. LSTM usually has more parameters, which may lead to longer training time and greater model storage requirements. The currently proposed N-Beats model, different from the recursive structure of LSTM, adopts a fully connected feed-forward neural network structure, which makes the model easier to perform parallel computing, improving the training speed and efficiency. And N-Beats can handle long time series and performs well in feature extraction. Especially in the case of no time dependence, it can effectively avoid the problem of gradient disappearance in long sequence learning. Moreover, N-Beats learns the features of time series by stacking multiple blocks, without the need to manually select features or perform additional preprocessing, which simplifies the modeling process. However, these models lack an efficient feature fusion mechanism during the training process and usually rely only on a single feature for filling, unable to effectively capture the complex relationships between data. This lack of comprehensive feature extraction and fusion makes the filling results vulnerable to interference from noise and outliers, reducing the robustness and generalization ability of the model.
[0064] In view of the above problems, the present invention proposes a MODIS land surface temperature filling method based on spatio-temporal feature fusion, as Figure 1 shown, which specifically includes the following steps:
[0065] Obtain the data to be filled;
[0066] Input the data with missing values into the spatio-temporal feature fusion network model to obtain the filling result. The spatio-temporal feature fusion network model is obtained by training with a training set, which is historical land surface temperature data. The spatio-temporal feature fusion network model is constructed using a neural network. The neural network is used to perform spatio-temporal interaction fusion on the data with missing values and conduct multi-dimensional feature synthesis to obtain the filling result.
[0067] Specifically, this method uses one-dimensional time series and three-dimensional spatio-temporal series as input features, fully fuses the feature information of both, and comprehensively extracts the feature information of land surface temperature in time and space. The model consists of two parallel branch modules and two feature fusion modules. Each branch module is used to process different types of data. The 1D-CNN-BiNbeats module consists of a 1D-CNN model and a bidirectional N-Beats model, which is used to process one-dimensional time series and deeply extract the feature information of land surface temperature in the time series. The 3D-CNN module extracts the feature information of land surface temperature in time and space through a 3D-CNN model. It should be noted that a multi-dimensional feature fusion module is introduced in this model to extract and fuse the features of each branch, thereby enhancing the information exchange between modules and reducing information loss. This model can efficiently mine the one-dimensional time features and three-dimensional spatio-temporal features of land surface temperature at the same time, and based on the multi-dimensional feature fusion module, it realizes the filling of land surface temperature using multi-dimensional features and obtains a land surface temperature filling result with higher accuracy.
[0068] Furthermore, obtaining the training set includes:
[0069] Obtain the original land surface temperature data;
[0070] Perform splicing, cropping, and reprojection processing on the original land surface temperature data to obtain the training set.
[0071] Specifically, the MODIS data product used in the present invention is in HDF format. The MRT (MODIS Projection Tool) software provided by NASA needs to be used to perform splicing, cropping, and reprojection processing on the original image. Since the projection of MODIS data is generally Sinusoidal projection, which is not suitable for conventional data viewing methods, the obtained data needs to be reprojected. The present invention uses the MOD11A2 data from 2003 to 2015 with a period of 8 days, and the size is 900×900×598. The MOD11A2 data comes from the Terra satellite, is less affected by factors such as cloud cover, and the data is more accurate.
[0072] Furthermore, the spatio-temporal feature fusion network model includes: a 1D CNN-BiNbeats module, a 3D-CNN module, a spatio-temporal interaction fusion module, and a multi-dimensional feature synthesis module;
[0073] 1D CNN-BiNbeats module for extracting one-dimensional feature information of land surface temperature;
[0074] 3D-CNN module for extracting three-dimensional feature information of land surface temperature;
[0075] Spatio-temporal interaction and fusion module for fusing one-dimensional feature information and three-dimensional feature information;
[0076] Multi-dimensional feature synthesis module for synthesizing the fused one-dimensional feature information and three-dimensional feature information.
[0077] Specifically, as Figure 2 shown, the spatio-temporal feature fusion network consists of a 1D CNN-BiNbeats module, a 3D-CNN module and two feature fusion modules. Among them, the 1D CNN-BiNbeats module is good at processing one-dimensional sequence data and is used to extract the feature information of land surface temperature in the time series; the 3D-CNN module extracts the three-dimensional features of land surface temperature through the 3D-CNN model to better mine the spatio-temporal feature information of land surface temperature; in addition, the present invention designs a multi-dimensional feature fusion module and a fusion loss, and this module can extract and fuse the features of each branch to strengthen the information exchange between different branches. Among them, Fl: data 1 feature, F1': fused data 1 feature, f1: branch 1 output feature, F2: data 2 feature, F2': fused data 2 feature, f2: branch 2 output feature, ConvlD: 1x1 convolution Pooling Layer: pooling layer, L rec : fusion reduction loss, Bi-BB: bidirectional basic block, Conv3D: 3x3 convolution L final : filling value accuracy loss.
[0078] Furthermore, the extraction of the one-dimensional feature information of land surface temperature includes:
[0079] Using 1D-CNN units to extract the data to be filled for one-dimensional time series feature extraction;
[0080] Using bidirectional N-Beats units to capture the basic patterns of one-dimensional time series features and learn long-term dependence relationships;
[0081] According to the long-term dependence relationship, obtain one-dimensional feature information.
[0082] Specifically, the 1D Convolutional Neural Network (1D-CNN) is a deep learning architecture specifically designed to process one-dimensional data (such as time series or text data). Through local connection and weight sharing, this model can effectively extract the features of the input data and is applicable to tasks such as classification, regression, and sequence prediction. In 1D-CNN, the convolution operation extracts local features of the input data through a convolution kernel. Let the input vector be X, and the convolution kernel be , then the convolution operation can be expressed as:
[0083]
[0084] where, is the output of the convolutional layer, is the size of the convolution kernel, represents the position of the current output element, represents the position of the element within the convolution kernel or window. The output of the convolutional layer usually undergoes processing by an activation function to introduce non-linear characteristics. The commonly used activation function is ReLU (Rectified Linear Unit), which is defined as:
[0085]
[0086] To reduce the feature dimension and computational complexity, 1D-CNN uses a pooling layer. The max pooling operation is defined as:
[0087]
[0088] where the window size determines the range of the pooling operation. Finally, the extracted features are passed to the fully connected layer for integration. Let the input be the feature vector , and the weight matrix be , the output of the fully connected layer can be expressed as:
[0089]
[0090] where, is the bias term. Through the above structure, 1D-CNN can effectively extract temporal features from one-dimensional data, thereby improving the performance of the model in related tasks. This model shows significant advantages when processing input data with temporal dependencies.
[0091] Furthermore, the N-Beats unit includes: an input layer, a basic block, a stacking layer, and an output layer;
[0092] The input layer is used to input one-dimensional temporal features;
[0093] The basic block is used to capture the information flow from the past to the current in the time series, extract the forward dependencies of the sequence, capture the reverse information flow from the future to the current, extract the backward dependencies of the sequence, and obtain the prediction result based on the forward and backward dependencies;
[0094] The stacking layer is used to correct the prediction result;
[0095] The output layer is used to output the corrected prediction result.
[0096] Specifically, the N-Beats (Neural Basis Expansion Analysis Time Series Forecasting) model is an innovative time series forecasting method, mainly composed of a multi-layer stacked structure block to efficiently extract basic features such as trends and seasonality of the sequence. The key to the design of this model lies in the stacking of multiple basic blocks (BasicBlock), each block consisting of a fully connected neural network, and improving the prediction effect in a recursive manner. The model structure is as follows:
[0097] Input layer: Input time series data The model divides it into segments of appropriate size for prediction.
[0098] Basic Block: Each basic block contains multiple layers of fully connected neural networks, and the output is divided into two parts: "forward prediction" and "backward reconstruction", which are used to generate future values and reconstruct the current input sequence respectively. Forward output : Used to predict future time points, backward output : Used to reconstruct the input, help the model learn the basic patterns of the time series, and form a low-dimensional representation of the data.
[0099] Stacking structure: The model recursively corrects the prediction error layer by layer through the stacking of multiple basic blocks, and accumulates the prediction results of each layer in turn to improve the prediction accuracy.
[0100] Output layer: The final forward prediction result is obtained by adding the forward outputs of each basic block.
[0101] Suppose the incoming time series is , and the forward and backward outputs of each basic block are denoted as and . The final predicted value and the reconstructed value are defined as follows:
[0102]
[0103]
[0104] Among them, is the number of base blocks, represents the th base block. The N-BEATS model weights and fuses the prediction results of all base blocks to generate the final prediction result:
[0105]
[0106] Among them, and respectively represent the weighting coefficients of the forward and backward outputs.
[0107] The present invention improves the N-Beats model and constructs a bidirectional N-Beats layer. The forward N-Beats layer captures the information flow from the past to the present in the time series, extracts the forward dependence of the sequence, and the backward N-Beats layer captures the reverse information flow from the future to the present, extracts the backward dependence of the sequence. That is, the base block is improved into a bidirectional base block, and finally, through concat connection, the corresponding one-dimensional features are output. The bidirectional structure can enhance the model's understanding of the overall time series information, thereby possibly improving the prediction accuracy.
[0108] In summary, the specific method of the 1D CNN-BiNbeats module is as follows: First, take the one-dimensional time series where the missing value points are located as the input and input it into the 1D-CNN model for feature extraction to reduce redundant data, making the data volume smaller and the data features unchanged. Then, use the bidirectional N-beats model to learn the basic patterns, temporal features, and long-term dependence relationships of the data, capture the information flow from the past to the present and the reverse information flow from the future to the present in the time series, extract the forward dependence and backward dependence of the sequence, and finally output the corresponding one-dimensional features through concat connection.
[0109] Furthermore, the 3D-CNN module includes: an input layer, a convolutional layer, an activation layer, a pooling layer, and an output layer;
[0110] The input layer is used to input the neighborhood spatio-temporal sequence data of the missing value points;
[0111] The convolutional layer is used to extract the three-dimensional features of the spatio-temporal sequence data through three-dimensional convolution operations;
[0112] The activation layer is used to perform non-linear processing on the three-dimensional features using an activation function;
[0113] The pooling layer is used to reduce the feature dimension and extract representative three-dimensional features;
[0114] The output layer is used to output three-dimensional feature information.
[0115] Specifically, the 3D-CNN module receives the spatio-temporal sequence data of the 3×3 neighborhood of the missing value points through the 3D-CNN model, that is, the dimension of the input data is 3×3×598. This module extracts spatio-temporal features through convolutional operations on the neighborhood data. Specifically, the 3D-CNN module uses its three-dimensional convolutional kernel to extract features in both the time and space dimensions, thereby capturing the deep features of the data. This process helps to construct a richer three-dimensional feature representation, providing a more accurate basis for the subsequent value filling task. Finally, the three-dimensional features output by the model will be used to fuse with the results of other modules, further improving the accuracy and robustness of value filling.
[0116] 3D-CNN (3D Convolutional Neural Network) is a deep learning architecture designed specifically for processing three-dimensional data, capable of effectively extracting spatio-temporal features. In the present invention, the 3D-CNN model is used to process the spatio-temporal sequence data of the 3×3×598 neighborhood of the missing value points to enhance the model's understanding of complex temporal patterns. Its model structure is as follows:
[0117] Input layer: The input data is , representing the spatio-temporal sequence of the 3×3 neighborhood of the missing value points.
[0118] Convolutional layer: Extracts features through three-dimensional convolutional operations. The convolutional operation can be expressed as:
[0119] (8)
[0120] Among them, is the three-dimensional convolutional kernel, are the sizes of the convolutional kernel in the space and time dimensions respectively, is the index of the output tensor, representing the position in the three-dimensional space, is the index of the convolutional kernel, used to traverse each element of the convolutional kernel.
[0121] Activation function: The output of the convolutional layer usually goes through an activation function to introduce non-linearity. The commonly used activation function is ReLU, and its definition is:
[0122]
[0123] Among them, is the output tensor of the activation layer.
[0124] Pooling layer: To reduce the feature dimension and extract more representative features, 3D-CNN uses a pooling layer. The formula for max pooling is:
[0125]
[0126] Among them: is the output tensor of the pooling layer, is the size of the pooling window, which determines the scope of the pooling operation.
[0127] Output layer: Finally, after multiple layers of convolution and pooling, the output feature representation is for subsequent value filling tasks.
[0128] The 3D-CNN model can effectively capture the spatio-temporal features around the missing value points and improve the ability to understand the data. This enables the model to more accurately predict the missing data points during the value filling process, thereby improving the overall prediction performance. By using three-dimensional convolution operations, 3D-CNN not only focuses on the correlation of the spatial neighborhood but also considers the dynamic changes in the time dimension, providing a richer information basis for the value filling task.
[0129] Furthermore, fusing one-dimensional feature information and three-dimensional feature information includes:
[0130] Unify the feature dimensions of the one-dimensional feature information and the three-dimensional feature information through convolution operations to obtain dimension-unified features;
[0131] Fuse the dimension-unified features to obtain the fused feature information.
[0132] Specifically, the specific steps of the spatio-temporal interaction fusion module: Embed this module in two modules (1D CNN-BiNbeats module and 3D-CNN module). This module focuses on processing features of different dimensions and unifies the feature dimensions of each branch through convolution operations to ensure that these features can be effectively fused in the same space. Subsequently, these features with unified dimensions are fused by means of feature addition (Add operation). Feature addition means adding two features element-wise, increasing the information volume without changing the dimension. The fused features are passed back to their respective branch networks through convolutional kernels for further processing. It should be noted that a multi-dimensional feature fusion loss (see Figure 3 ). In Figure 3 , Fl: Feature of Data 1, F1': Feature of Fused Data 1, F2: Feature of Data 2, F2': Feature of Fused Data 2, Conv: Convolutional layer, Kernel: Convolutional kernel, ADD: Addition, Concat: Concatenation, L rec: Fusion reduction loss. The specific calculation method is as follows: First, concatenate (concat) two input features (F1, F2) and two output features (F1', F2') respectively, and then calculate the loss during the fusion process through a loss function. The feature fusion loss function designed in the present invention aims to ensure the integrity and similarity of information of features in different dimensions during the fusion process, enhance the synergistic effect between branch features, so that the fused features are more accurate. By introducing a spatio-temporal interaction fusion module, features in different dimensions or from different sources are effectively fused, the feature expression ability is improved, the robustness of the model is enhanced, and information loss is effectively reduced. Each branch can not only independently process its own data, but also share and interact information with other branches during the fusion process, effectively absorbing useful information from other branches. This design not only enhances the feature interaction between branches, but also reduces information loss, realizes more efficient feature fusion, and improves the accuracy and generalization ability of the final filling result.
[0133] Further, the synthesis of the fused one-dimensional feature information and three-dimensional feature information includes:
[0134] Perform reduction processing on the fused one-dimensional feature information and three-dimensional feature information;
[0135] Perform dimensionality reduction operations on the reduced one-dimensional feature information and three-dimensional feature information, and extract the reduced one-dimensional feature and three-dimensional feature;
[0136] Perform mapping on the reduced one-dimensional feature and three-dimensional feature to obtain one-dimensional mapped feature and three-dimensional mapped feature;
[0137] Synthesize the one-dimensional mapped feature and three-dimensional mapped feature to obtain the filling result.
[0138] Specifically, the specific steps of the multi-dimensional feature synthesis module are as follows. After each branch network completes feature extraction and fusion, the model performs dimensionality reduction operations on the output features (f1, f2) from the two branches through a convolutional layer to extract features, as Figure 4 shown, where fl: the output feature of branch 1, f2: the output feature of branch 2, Conv: convolutional layer, Add: addition, FC: fully connected layer, L final : Filling accuracy loss. Then, further mapping is performed on the features processed by the convolutional layer through a fully connected layer to improve the feature expression ability. Finally, the two feature elements output by the fully connected layer are added to obtain the final filling result, and a loss for evaluating the filling result is set . During the training process of the present invention and both use the mean square error (MSE for short) as the loss function, which is a commonly used numerical standard, and the calculation formula is as follows:
[0139]
[0140] represents the i-th true value, represents the i-th predicted value, is the total number of data points.
[0141] These two loss functions together constitute the loss function of the present invention, and the total loss can be expressed as:
[0142]
[0143] This design not only strengthens the expression of features in each dimension, but also realizes the unified processing of multi-dimensional features at the overall level of the model, learns more spatio-temporal information and multi-dimensional features, thereby improving the performance of the model in complex tasks, and enabling the model to more accurately capture the spatio-temporal variation characteristics of surface temperature.
[0144] Furthermore, when training the spatio-temporal feature fusion network model, the Adam optimization algorithm is used to update the parameters of the spatio-temporal feature fusion network model.
[0145] The present invention also includes: model training and model testing and result evaluation;
[0146] Model training:
[0147] The input of the model data is one-dimensional time series and three-dimensional spatio-temporal series data. Among them, the size of the MODIS data used is 900×900×598, and 70% of the data set is selected as the training set when training the model. During the training process, first set the model hyperparameters: learning rate: 0.001; epoch: 100; batch size: 64, optimizer: Adam; loss function: MSE; Dropout parameter: 0.5.
[0148] This model first passes the data from the input layer to the output layer through layer-by-layer calculations between network layers to obtain the prediction result, and then adjusts the network parameters based on the gradient descent algorithm to reduce the error index between the prediction result and the reference value. In this study, the Adam optimization algorithm is used during the training process. The Adam optimizer can automatically adjust the values of the learning rate and momentum according to the current training status, better cope with different data dimensions and model complexities, and thus significantly improve the training effect of the model. Compared with other optimization algorithms, the Adam optimizer obtains momentum by calculating the mean of the gradients. By introducing momentum, the Adam optimizer can consider the influence of previous gradients when updating parameters, and thus proceed more stably during the training process. The adaptive learning rate dynamically adjusts the size of the learning rate and the update step of the parameters by calculating the variance of the gradients to meet the update requirements of different parameters. The update principle of the Adam optimizer is as follows:
[0149]
[0150] Among them, represents the loss function; represents the current gradient; 、 and 、 represent the mean and variance of the gradients at the current and previous moments respectively; and represent the correction results of the estimated values; represents the learning rate, with a default value of 0.001; and are the decay exponents of the mean and variance of the gradients respectively; represents the smoothing term, used to prevent the denominator from being zero; t is the number of iterations, represents the updated parameter.
[0151] Model testing and result evaluation:
[0152] Select 30% of the dataset as the test set, input it into the trained model, and use two performance indicators, the correlation coefficient (R) and the root mean square error (RMSE), to evaluate the inversion accuracy of the spatio-temporal feature fusion network model. The root mean square error (RMSE), as a commonly used indicator, is often used to measure the magnitude of the error between the predicted data and the reference data; while the correlation coefficient (R) is an indicator used to evaluate the degree of association between two variables, and its value range is between -1 and 1. The calculation methods of the root mean square error (RMSE) and the correlation coefficient (R) are respectively:
[0153]
[0154]
[0155] Among them, is the actual observed value, is the model predicted value, represents the total number of sample points, and are the average predicted value and the average actual observed value respectively.
[0156] The following takes Yutian area in Xinjiang and Wenchuan area in Sichuan as examples to elaborate on this embodiment in detail:
[0157] The Yutian area in Xinjiang (central position 79.9°E, 37.1°N) and the Wenchuan area in Sichuan (central position 103.4°E, 31.0°N) compare the method of this embodiment with the following three methods: the surface temperature filling method based on the CNN model (CNN method), the surface temperature filling method based on the N-beats model (N-beats method), and the surface temperature filling method based on the hybrid model with the best current filling effect (SCLSTM method).
[0158] To evaluate the reconstruction effect of the model, the "removal-reconstruction-comparison" strategy was used to conduct qualitative and quantitative analyses of the prior art on the reconstruction effects of the method of this embodiment and the comparison methods. The "removal-reconstruction-comparison" strategy is to first remove some existing values from the complete LST image and fill them with zeros, then use the method of this embodiment and the comparison methods to reconstruct the missing value pixels respectively, and finally compare and analyze the reconstructed LST values with the original data. Taking the image of the 30th period in 2008 in the Yutian area as an example, the following are the comparison images and accuracy comparison tables before and after filling values according to the method of this embodiment. From Figure 5 it can be seen that the texture of the image after filling values is delicate and natural, without obvious boundary effects.
[0159] As can be seen from Table 1, the filling value result of the method of this embodiment is the best, the correlation coefficient is the highest at 0.951, and the RMSE is the smallest at 0.903, all of which are better than the comparison methods.
[0160] Table 1 Comparison results of the method of this embodiment with the CNN method, N-beats method, and SCLSTM method (Yutian area)
[0161]
[0162] The following are the filling value result diagram and accuracy comparison table of the 15th period image in 2007 in the Wenchuan area:
[0163] From Figure 6It can be seen that the image texture after value filling is delicate and natural, without obvious boundary effects. As can be seen from Table 2, the value filling result of the method in this embodiment is the best, the highest correlation coefficient is 0.952, and the minimum RMSE is 0.709, still all better than the comparative methods, indicating that the method in this embodiment has universality.
[0164] Table 2 Comparison results of the method in this embodiment with the CNN method, N-beats method, and SCLSTM method (Wenchuan area)
[0165]
[0166] Through the example application of this technical solution, the following two conclusions can be drawn:
[0167] (1) The method in this embodiment can better explore the characteristics of land surface temperature in the spatio-temporal sequence, making the reconstructed image have no obvious boundary effects, the texture information is more natural, and the accuracy of the reconstructed LST is more stable and accurate.
[0168] (2) The method in this embodiment has strong regional applicability and can achieve good reconstruction effects in areas with less missing values, areas with more missing values, areas with flat terrain, areas with large terrain undulations, and areas with different climate conditions, proving the regional practicality of this embodiment.
[0169] The above is only a preferred specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in this application should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. The MODIS surface temperature compensation method based on spatiotemporal feature fusion is characterized by: include: Get the value data to be supplemented; The data to be supplemented is input into a spatiotemporal feature fusion network model to obtain a supplementation result, wherein the spatiotemporal feature fusion network model is obtained by training a training set, the training set is historical surface temperature data, the spatiotemporal feature fusion network model is constructed using a neural network, the neural network is used to extract one-dimensional time series features and three-dimensional spatiotemporal series features, and the one-dimensional time series features and the three-dimensional spatiotemporal series features are subjected to feature fusion to obtain the supplementation result; The spatiotemporal feature fusion network model includes: a one-dimensional time series feature extraction module, a three-dimensional spatiotemporal sequence feature extraction module, a spatiotemporal interactive fusion module and a multi-dimensional feature synthesis module; The one-dimensional time series feature extraction module is used to extract one-dimensional feature information of the surface temperature; The three-dimensional spatiotemporal sequence feature extraction module is used to extract three-dimensional feature information of surface temperature; The spatiotemporal interactive fusion module is used to fuse the one-dimensional feature information and the three-dimensional feature information; The multi-dimensional feature synthesis module is used to synthesize the fused one-dimensional feature information and three-dimensional feature information; The step of synthesizing the fused one-dimensional feature information and three-dimensional feature information includes: Performing restoration processing on the fused one-dimensional feature information and three-dimensional feature information; Performing a dimensionality reduction operation on the restored one-dimensional feature information and three-dimensional feature information to extract the reduced one-dimensional features and three-dimensional features; Mapping the one-dimensional features and the three-dimensional features after dimensionality reduction to obtain one-dimensional mapping features and three-dimensional mapping features; The one-dimensional mapping feature and the three-dimensional mapping feature are synthesized to obtain a complementation result.
2. The MODIS surface temperature compensation method based on spatiotemporal feature fusion according to claim 1 is characterized in that: Acquiring the training set includes: Get raw surface temperature data; The original surface temperature data is spliced, cropped and reprojected to obtain the training set.
3. The MODIS surface temperature compensation method based on spatiotemporal feature fusion according to claim 1 is characterized in that: Extracting one-dimensional characteristic information of surface temperature includes: Extracting the to-be-complemented value data using a 1D-CNN unit to perform one-dimensional time series feature extraction; Using bidirectional N-Beats units to capture the basic patterns of the one-dimensional time series features and learn long-term dependencies; The one-dimensional feature information is obtained according to the long-term dependency relationship.
4. The MODIS surface temperature compensation method based on spatiotemporal feature fusion according to claim 3 is characterized in that: The N-Beats unit includes: an input layer, a basic block, a stacking layer and an output layer; The input layer is used to input the one-dimensional time series feature; The basic block is used to capture the information flow from the past to the present in the time series, extract the forward dependency of the sequence and capture the reverse information flow from the future to the present, extract the backward dependency of the sequence, and obtain the prediction result according to the forward dependency and the backward dependency; The stacked layer is used to correct the prediction result; The output layer is used to output the corrected prediction result.
5. The MODIS surface temperature compensation method based on spatiotemporal feature fusion according to claim 1 is characterized in that: The three-dimensional spatiotemporal sequence feature extraction module includes: an input layer, a convolution layer, an activation layer, a pooling layer and an output layer; The input layer is used to input the neighborhood spatiotemporal series data of missing points; The convolution layer is used to extract the three-dimensional features of the spatiotemporal sequence data through a three-dimensional convolution operation; The activation layer is used to perform nonlinear processing on the three-dimensional features using an activation function; The pooling layer is used to reduce feature dimensions and extract representative three-dimensional features; The output layer is used to output the three-dimensional feature information.
6. The MODIS surface temperature compensation method based on spatiotemporal feature fusion according to claim 1 is characterized in that: Fusion of the one-dimensional feature information and the three-dimensional feature information includes: Unifying the feature dimensions of the one-dimensional feature information and the three-dimensional feature information through a convolution operation to obtain dimensionally unified features; The unified features of the dimensions are fused to obtain fused feature information.
7. The MODIS surface temperature compensation method based on spatiotemporal feature fusion according to claim 1 is characterized in that: When training the spatiotemporal feature fusion network model, the Adam optimization algorithm is used to update the parameters of the spatiotemporal feature fusion network model.
Citation Information
Patent Citations
Virtual test multivariable space-time sequence generation method
CN118552747A