A Sea Surface Temperature Prediction Method Based on Multi-Scale Frequency Domain Enhancement
Through the combination of multi-scale frequency domain enhancement module and convolutional long and short-term memory network module, the problems of low accuracy and poor stability of sea surface temperature prediction in the existing technology are solved, and more efficient sea surface temperature prediction is achieved, especially the capture of complex spatial and temporal changes at sea and land junctions and the improvement of stability over the entire sea area.
Patent Information
- Application Number
- CN202510660120.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-22
AI Technical Summary
Existing sea surface temperature prediction methods are difficult to effectively capture the spatial and temporal dynamics synergistic evolution law of ocean temperature field, resulting in low prediction accuracy and poor stability, especially when processing high-dimensional and multi-scale spatiotemporal data, it is impossible to adaptively focus key areas.
The multi-scale frequency domain enhancement sea surface temperature prediction method is adopted, through the combination of the multi-scale frequency domain enhancement module and the convolutional long and short-term memory network module, the multi-scale cyclic expansion convolution module and the frequency domain space joint attention module are used to capture the complex spatial and temporal changes at the sea and land junction, and improve the local prediction accuracy and prediction stability within the entire sea area.
It significantly improves the local accuracy of sea surface temperature prediction and prediction stability over the entire sea area, can control the frequency information type more finely, weaken short-term turbulence interference, strengthen the characteristics of long-term climate events, and improve the training stability and prediction accuracy of the model.
Smart Images

Figure CN120181132B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of sea surface temperature prediction, and particularly relates to a sea surface temperature prediction method based on multi-scale frequency domain enhancement. Background Art
[0002] Sea Surface Temperature (SST), as a core parameter of the interaction between the ocean and the atmosphere, plays a crucial role in the global weather and climate system. It is closely related to the formation conditions of extreme weather and affects the structure and function of the marine ecosystem. Predicting SST can not only prevent meteorological disasters such as droughts and floods in a timely manner, but also ensure the safety and stability of the marine ecosystem, and is also helpful for studying climate events such as the Southern Oscillation and the Indian Ocean Dipole. It plays a decisive role in fields such as tropical cyclone generation, El Niño phenomenon prediction, and fishery resource distribution, and has now become a key indicator for marine environmental monitoring and climate research.
[0003] Current SST prediction methods are mainly divided into two types, namely numerical models based on physical mechanisms and numerical-driven methods based on statistics. The former relies on reliable physical prior knowledge and has a clear physical interpretability, but involves many parameters and has a high computational cost; the latter is represented by traditional machine learning algorithms such as Support Vector Machine (SVM) and Random Forest. Although it can mine statistical laws through historical data, it is difficult to capture non-linear spatio-temporal coupling features when dealing with high-dimensional and multi-scale SST spatio-temporal data, and traditional methods cannot adaptively focus on key regions and are difficult to extract long-term dependence relationships of SST, resulting in low prediction accuracy.
[0004] In recent years, deep learning technology has opened up a new way for SST prediction. Typical time series models such as Long Short-Term Memory Network (LSTM) and its spatial extension variant ConvLSTM can model time dependence through the memory gating mechanism, but the local receptive field of its convolution kernel also limits the ability to capture large-scale spatial correlations. These methods often decouple spatio-temporal features, first extract spatial patterns through convolution, and then model time evolution through the network. This serial architecture is difficult to depict the co-evolution law of spatio-temporal dynamics in the ocean temperature field, and blindly stacking feature extraction layers (ConvLSTM layers) not only has low efficiency, but also leads to data redundancy, thus limiting the prediction accuracy. In order to more comprehensively analyze the spatio-temporal law of SST data, it is necessary to develop a deep learning model that can simultaneously consider time and space features and can effectively process multi-channel information. Summary of the Invention
[0005] To solve the above technical problems, the present invention provides a sea surface temperature prediction method based on multi-scale frequency domain enhancement, which effectively captures the complex spatio-temporal variation characteristics at the land-sea boundary through a multi-scale frequency domain enhancement mechanism, thereby improving the local prediction accuracy while significantly enhancing the prediction stability across the entire sea area.
[0006] To achieve the above object, the technical solution of the present invention is as follows:
[0007] A sea surface temperature prediction method based on multi-scale frequency domain enhancement, comprising the following steps:
[0008] Step 1, preprocess the sea surface temperature dataset obtained from the ERA5 reanalysis data of the European Centre for Medium-Range Weather Forecasts and divide it into a training set, a validation set, and a test set;
[0009] Step 2, input the data of the training set into the deep learning model MFE-ConvLSTM for training, use the validation set to adjust the parameters, and finally evaluate the model accuracy with the data of the test set;
[0010] The deep learning model MFE-ConvLSTM consists of two parts: a multi-scale frequency domain enhancement module and a convolutional long short-term memory network module;
[0011] The multi-scale frequency domain enhancement module includes three parallel branches, and each branch contains a main branch and a sub-branch; each main branch is sequentially composed of a multi-scale cyclic dilated convolution module and a frequency domain-spatial joint attention module; the three main branches are respectively configured with cyclic dilated convolutions with dilation rates of 1, 2, and 4; the sub-branch adjusts the channel dimension through a 3D convolution kernel with a kernel size of 1×1×1; the output of the sub-branch and the main branch are added element by element, and then feature fusion and non-linear enhancement are performed through a ReLU activation function; the channel features output by the three parallel branches are concatenated along the channel dimension to form a composite multi-scale feature tensor, and then a 3D convolution kernel with a size of 1×1×1 is used to project the concatenated high-dimensional features to the target dimension, and after ReLU activation, it is input into a 3D convolution kernel with a size of 1×1×1 and then output;
[0012] The convolutional long short-term memory network module includes two key parts: a ConvLSTM2D layer and a 1×1 two-dimensional convolutional layer; inside the ConvLSTM2D layer, multiple ConvLSTM units are recursively expanded along the time dimension to process the input sequence, and each ConvLSTM unit contains an input gate, a forget gate, an output gate, and a candidate memory unit, and the output channel is 24, and the gating and state update are both achieved through a convolution operation with a kernel size of 3×3 and a stride of 1, and the temporal modeling is completed while maintaining the input spatial resolution; the 1×1 convolutional layer is used to achieve the fusion and output of multi-channel features;
[0013] Step 3: After preprocessing the sea surface temperature data of the sea area to be predicted, input it into the qualified model for sea surface temperature prediction.
[0014] In the above solution, the multi-scale cyclic dilated convolution module expands the data dimension to three dimensions, uses cyclic data padding at the start and end of the time series, and sets the cyclic padding only at the beginning and end of the time dimension; the padding amount is calculated as follows:
[0015] ;
[0016] In the formula, is the dilation rate, which controls the sampling interval between elements inside the convolution kernel; is the convolution kernel size, that is, the kernel size for expanding the receptive field in the time dimension.
[0017] In the above solution, in the parallel structure of the multi-scale cyclic dilated convolution module, weight normalization is used to wrap the multi-scale cyclic dilated convolution module to normalize the weights of the convolution kernel, and the normalized weights are decomposed into a direction vector and a scaling factor :
[0018] ;
[0019] In the formula, is the normalized weight, is the unnormalized weight vector, representing the direction of the weight, is the learnable scaling factor, representing the magnitude of the weight, is the L2 norm of
[0020] In the above solution, the frequency-domain spatial joint attention module is composed of a channel attention module and a spatial attention module, and feature calibration is achieved through frequency-domain - spatial-domain joint modeling;
[0021] The channel attention module includes a frequency-domain feature selection module and a weight generation module. The frequency-domain feature selection module, i.e., the DCT layer, constructs a multi-spectrum filter through discrete cosine transform basis functions to select different frequency components, obtains the frequency-domain coordinates from a predefined mapping table, and dynamically generates filter weights; the weight generation module performs feature compression and weight recovery through two fully connected layers;
[0022] The spatial attention module consists of two convolutional layers. The first convolutional layer has a kernel size of 3×3, and the number of output channels is the original number of channels divided by the spatial compression rate. The ReLU activation function is used to capture local spatial context and reduce the channel dimension. The second convolutional layer has a kernel size of 3×3 and 1 output channel. The Sigmoid activation function is used to generate a spatial attention mask. There is a batch normalization layer after each convolutional layer. This weight map is multiplied element-wise with the feature map processed by the channel attention module to achieve re-weighting of the features, and finally a fused output feature map is generated.
[0023] In a further technical solution, in the frequency domain feature selection module, after the input features retain the original information through residual connection, they enter the frequency domain feature selection module and the weight generation module in sequence for feature enhancement: First, the DCT layer is used to screen key frequency components, and then the first fully connected layer is used for dimensionality reduction, with the number of output channels being the original number of channels divided by the channel compression rate, and ReLU activation is used; then the second fully connected layer is used to restore the channel dimension, and the unnormalized attention weights are output; finally, the Sigmoid function is applied to generate the channel attention map, and through the feature reshaping operation, it is broadcast to the size of the original feature map. After the input features are extracted with multi-spectrum components by the DCT layer, the channel attention weights are generated through the compression-expansion structure, and finally multiplied element-wise with the original features to achieve re-weighting of the features.
[0024] In the above solution, the formula of the frequency domain spatial joint attention module is as follows:
[0025] ;
[0026] ;
[0027] ;
[0028] Among them, is the input tensor, is the channel attention module, is the spatial attention module, is the output of the channel attention module, is the output of the spatial attention module, is the output of the frequency domain spatial joint attention module; is element-wise multiplication. In a further technical solution, the discrete cosine transform basis function is as follows:
[0029] ;
[0030] In the formula, is the frequency index, is the height and width of the feature map, is the spatial position.
[0031] In a further technical solution, the formula of the channel attention module is as follows:
[0032] ;
[0033] The formula of the spatial attention module:
[0034] ;
[0035] In the formula, is the output of the channel attention module, represents the output of the spatial attention module, represents the Sigmoid activation function, and both represent the Conv2D function with a convolution kernel size of 3×3. The former has a stride of 1, and the latter has a stride of , and both use the same padding method; B represents the batch normalization layer.
[0036] In a further technical solution, in the ConvLSTM cell, the forget gate determines which information in the cell state at the previous time step needs to be forgotten, and the output of the forget gate is a value between 0 and 1, indicating the proportion of retention or discard at each position; the input gate controls the degree of update of the input at the current time step to the cell state, and the output of the input gate determines which parts in the candidate cell state will be written into the current cell state ; the output gate determines which information should be extracted from the cell state at the current time step for the hidden state , that is, the final hidden state is the result obtained by filtering a part of the cell state through the output gate.
[0037] In a further technical solution, in the ConvLSTM cell, the output of the forget gate is:
[0038] ;
[0039] In the formula, is the output of the forget gate, is the Sigmoid activation function, is the weight matrix input to the forget gate, is the input at the current time step, is the weight matrix from the hidden state of the previous time step to the forget gate, is the hidden state of the previous time step, is the weight matrix from the cell state of the previous time step to the forget gate, is the cell state of the previous time step, is the bias term of the forget gate, is the convolution operation, is the Hadamard product; the output of the input gate is:
[0040] ;
[0041] In the formula, is the output of the input gate, is the weight matrix input to the input gate, the weight matrix from the hidden state of the previous time step to the input gate, is the weight matrix from the cell state of the previous time step to the input gate, is the bias term of the input gate;
[0042] The candidate cell state is:
[0043] ;
[0044] In the formula, is the candidate cell state, is the hyperbolic tangent activation function, is the weight matrix input to the candidate cell state, is the weight matrix from the hidden state of the previous time step to the candidate cell state, is the bias term of the candidate cell state;
[0045] The cell state of the current time step is:
[0046] ;
[0047] In the formula, is the cell state of the current time step;
[0048] The output of the output gate is:
[0049] ;
[0050] In the formula, is the output of the output gate, is the weight matrix input to the output gate, is the weight matrix from the hidden state of the previous time step to the output gate, is the weight matrix from the cell state of the current time step to the output gate, is the bias term of the output gate;
[0051] The hidden state at the current time step is:
[0052] ;
[0053] In the formula, is the hidden state at the current time step.
[0054] By the above technical solution, a sea surface temperature prediction method based on multi-scale frequency domain enhancement provided by the present invention has the following beneficial effects:
[0055] 1. The present invention adopts a multi-scale cyclic dilated convolution module (CDIL-T). Through the cyclic padding strategy in the time dimension and the exponential dilation rate design, it can capture the daily variations, weekly fluctuations, and seasonal cycle characteristics of the ocean temperature field simultaneously, and can effectively improve the time feature extraction ability of the model in the 20-day prediction task;
[0056] 2. The present invention proposes a frequency domain space joint attention module (FCSA). Through the discrete cosine transform (DCT) frequency domain decomposition and hierarchical convolution space focusing, it realizes high-frequency noise suppression and low-frequency signal enhancement. It not only weakens the short-term turbulence interference, but also strengthens the characteristics of long-term climate events such as ENSO, enabling the model to more precisely control the type of frequency information used for attention calculation, and thus better capture the basic structure and gradual change characteristics of the sea surface temperature;
[0057] 3. The present invention introduces a weight normalization dynamic adjustment strategy. By decoupling the convolutional kernel direction vector and the scaling factor, the model can not only greatly compress the number of parameters while ensuring the accuracy, but also enable the convolutional branches with different dilation rates to automatically learn the feature importance weights, and the training stability is greatly improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art.
[0059] Figure 1 is a schematic flow chart of a sea surface temperature prediction method based on multi-scale frequency domain enhancement disclosed in an embodiment of the present invention;
[0060] Figure 2 is the overall architecture diagram of the MFE-ConvLSTM model;
[0061] Figure 3 is the structure diagram of the MFEM module;
[0062] Figure 4 is the structure diagram of the CDIL-T module with different dilation rates; (a) is the structure diagram with a dilation rate of 1; (b) is the structure diagram with a dilation rate of 2; (c) is the structure diagram with a dilation rate of 4;
[0063] Figure 5 It is the structural diagram of the FCSA attention mechanism module;
[0064] Figure 6 It is the schematic diagram of the convolutional long short-term memory network module;
[0065] Figure 7 They are the average RMSE and MAE predicted by each model for 1 - 20 days. (a) is RMSE; (b) is MAE;
[0066] Figure 8 They are the RMSE error distribution maps predicted by different models for the 20th day; (a) is LSTM; (b) is ConvLSTM; (c) is MFE - ConvLSTM;
[0067] Figure 9 They are the MAE error distribution maps predicted by different models for the 20th day; (a) is LSTM; (b) is ConvLSTM; (c) is MFE - ConvLSTM. Specific implementation manners
[0068] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.
[0069] The present invention provides a sea surface temperature prediction method based on multi-scale frequency domain enhancement, as Figure 1 shown, including the following steps:
[0070] Step 1: Preprocess the sea surface temperature dataset obtained from the ERA5 reanalysis data of the European Centre for Medium-Range Weather Forecasts, and divide it into a training set, a validation set, and a test set.
[0071] The SST data used in this embodiment is sourced from the ERA5 reanalysis dataset of the European Centre for Medium-Range Weather Forecasts (ECMWF). The data time span is from January 1, 2014 to December 31, 2023, and the spatial coverage range is from longitude to and latitude to of the sea area. The original data time resolution is hourly. To meet the needs of medium- and long-term sea temperature prediction, the present invention preprocesses the data through the following standardization process:
[0072] a: For the original hourly continuous observation data, the arithmetic mean method is used to calculate the daily average sea temperature value to retain the seasonal and synoptic scale change characteristics while eliminating short-term fluctuations;
[0073] b: For the invalid observation units marked as "NaN" in the ERA5 data (mainly distributed in the nearshore shallow water areas and extreme weather periods), a zero-value replacement strategy is adopted to avoid the gradient anomaly problem caused by missing values in the model;
[0074] c: Convert the thermodynamic temperature of the daily-averaged SST dataset to degrees Celsius to eliminate the potential bias of the absolute temperature scale on model training. Subsequently, uniformly convert the data format to 32-bit floating-point type for subsequent normalization and sliding window construction processing:
[0075] ;
[0076] d: Divide the preprocessed SST reanalysis dataset into continuous subsets in chronological order. Divide the data from January 1, 2014 to December 31, 2022 into a training set and a validation set according to a ratio of 8:2. The training set is used for optimizing model parameters, and the validation set is used for hyperparameter tuning and early stopping mechanism monitoring; Divide the data for one year from January 1, 2023 to December 31, 2023 into a test set for evaluating the generalization performance of the model. The dimension of the model input tensor , including the SST sequence with consecutive time steps, where is the batch size, is the time step length (observation window size), is the spatial dimension, is the number of channels (set to 1 for the univariate SST prediction task), and the dimension of the label tensor , representing the true SST value one day after the end time step of the input sequence. This design captures spatio-temporal evolution features through a five-dimensional input tensor and realizes a single-step prediction task through a four-dimensional output tensor, thereby achieving accurate SST prediction;
[0077] e: After dividing the dataset, map the values to the interval [0,1] through linear normalization. The normalization formula is:
[0078] ;
[0079] In the formula, represents the SST value in the original data, represents the normalized SST value, represents the minimum value of SST in the original data, represents the maximum value of SST in the original data.
[0080] Step 2: Input the data of the training set into the deep learning model MFE-ConvLSTM for training, use the validation set to adjust the parameters, and finally evaluate the model accuracy with the data of the test set.
[0081] S1. Construct the deep learning model MFE-ConvLSTM
[0082] As Figure 2 shown, the deep learning model MFE-ConvLSTM consists of two parts: a multi-scale frequency domain enhancement module and a convolutional long short-term memory network module.
[0083] 1. Multi-scale frequency domain enhancement module
[0084] As Figure 3 shown, the multi-scale frequency domain enhancement module (Multi-Scale Frequency Enhanced Mechanism, MFEM) includes three parallel branches, and each branch contains a main branch and a sub-branch; each main branch is sequentially connected by a multi-scale cyclic dilated convolutional module (CDIL-T) wrapped by weight normalization and a frequency domain-spatial joint attention module (FCSA); the three main branches are respectively configured with cyclic dilated convolutions with dilation rates of 1, 2, and 4 to capture spatio-temporal features at different time scales through exponentially expanding the receptive field. The number of channels of each group of outputs is 64, 128, and 256 in sequence to achieve complementary expression of multi-scale features. The FCSA attention mechanism, as the post-processing of each CDIL-T, realizes the collaborative enhancement of local spatio-temporal feature extraction and global dynamic calibration through the cascade optimization of frequency domain channel attention and spatial attention. The sub-branch adjusts the channel dimension through a 3D convolutional kernel with a size of 1×1×1 to achieve residual connections with different numbers of channels. The sub-branch adds element-wise to the output of the main branch, retains the key features of the original input, and the output result realizes feature fusion and non-linear enhancement with the ReLU activation function. The channel features output by the three parallel branches are concatenated along the channel dimension to form a composite multi-scale feature tensor. Subsequently, a 3D convolutional kernel with a size of 1×1×1 is used to project the concatenated high-dimensional features to the target dimension, reduce redundant information and retain key spatio-temporal patterns. After being activated by ReLU and then input into a 3D convolutional kernel with a size of 1×1×1, the computational efficiency of the model can be effectively optimized.
[0085] First, the input data with 1 channel is sent into Figure 4The three CDIL-T modules shown in (a), (b), and (c) in the figure each independently process the timing characteristics of different dilation rates. After being processed by each module, the data output channels sequentially become 64, 128, and 256. The increasing number of channels is combined with the dilation rate expansion, enabling different parallel blocks to focus on complementary spatio-temporal patterns and avoiding redundant features. Then, the data output from each block is sent to the FCSA attention mechanism module, which performs channel weighting on different frequency components simultaneously to retain multi-scale features. As the post-processing of each CDIL-T module, FCSA can achieve the cascaded optimization of local convolution combined with global attention. Subsequently, the residual output is added element-wise to the data output from the attention mechanism module, activated by the ReLU function, and then concatenated in the channel dimension. The total number of channels is compressed to 64 through a 3D convolution with a kernel size of 1×1×1, projecting the multi-scale features into a unified low-dimensional space while reducing the number of parameters for the convenience of subsequent module processing. Finally, the output activated by the ReLU function is compressed to 1 in terms of the total number of channels through a 3D convolution with a kernel size of 1×1×1 and then sent to the convolutional long short-term memory network module to reduce the computational burden of the model.
[0086] (1) Multi-scale Circular Dilated Convolution Module (CDIL-T)
[0087] The CDIL-CNN (Circular Dilated Convolutional Neural Network) architecture adopts a special design of splicing both ends of a one-dimensional long sequence. The receptive field is enlarged through dilated convolutions with different dilation rates, and the ring structure is used to reduce feature loss, enabling subsequent modules to obtain more global feature information. However, the original CDIL-CNN architecture is mainly designed for one-dimensional data and has limited applications. In order to adapt to the particularity of temporal information, the present invention improves on the basis of the CDIL-CNN module by combining a parallel structure, and proposes a multi-scale circular dilated convolution module (CDIL-T). This module expands the data dimension to three dimensions, and uses circular data padding at the start and end of the time series, with circular padding only at the head and tail in the time dimension, maintaining the periodic characteristics of the sea surface temperature data. The setting of the padding amount is as follows:
[0088] ;
[0089] In the formula, is the dilation rate, which controls the sampling interval between elements inside the convolutional kernel; is the convolutional kernel size, that is, the kernel size for expanding the receptive field in the time dimension. Dilated convolution expands the receptive field through spaced sampling, but may cause loss of edge information in the input sequence. By using circular padding with the head and tail connected, it can ensure that the convolutional kernel can completely cover the start and end regions of the input sequence, thereby keeping the length of the output in the time dimension consistent with the input.
[0090] Ocean temperature data has periodicity (such as seasonal changes, day-night temperature differences). Compared with zero-padding in traditional convolution, it will destroy the edge continuity. The circular convolution scheme performs circular padding in the time dimension to connect the head and tail of the sequence, retaining the complete periodic pattern of the time series and avoiding edge mutations during prediction. At the same time, a 3D dilated convolutional kernel is used, and the convolutional kernel size is , and only the dilation setting of dilated convolution is performed in the time dimension. The dilation rate follows exponential growth, is the hidden layer index, enabling the model to capture both short-term local fluctuations and long-term periodic patterns. Furthermore, while capturing changes at weekly, semi-monthly, monthly, and annual scales, it retains local spatial features (such as temperature correlations in adjacent sea areas), avoiding the destruction of locality by dilated convolution. In terms of parameters, the traditional Inception multi-branch structure increases the number of channels through parallel convolutional layers to capture features at different time scales, but its number of parameters grows linearly with the number of channels, resulting in high computational complexity and significant parameter redundancy. In addition, stacking convolutional layers can only linearly expand the receptive field and are difficult to efficiently integrate global features across time scales. In contrast, CDIL-T uses a single convolutional kernel combined with a circular padding mechanism to cover the complete time series period and dynamically adjusts the dilation rate using dilated convolution, achieving exponential expansion of the receptive field without increasing the number of parameters, thereby greatly improving the computational efficiency.
[0091] In the circular dilated convolution parallel structure of the present invention, weight normalization is used to wrap the CDIL-T module to normalize the weights of the convolutional kernel, enhancing the training stability. By decomposing the normalized weights into a direction vector ( ) and a scaling factor ( ), the direction and magnitude of the weights are explicitly decoupled:
[0092] ;
[0093] In the formula, is the normalized weight, is the unnormalized weight vector, representing the direction of the weight, is the learnable scaling factor, representing the magnitude of the weight, is 's L2 norm.
[0094] This decoupling makes the gradient direction more stable during backpropagation, especially suitable for the gradient aggregation scenario of multiple branches in deep parallel structures, and can effectively alleviate the gradient dispersion problem caused by the receptive field jump of dilated convolutions. The scaling factor g, as a learnable parameter, enables each convolutional kernel to have the ability to adaptively adjust the feature response intensity. Before the attention mechanism intervenes, this design allows the convolutional branches with different dilation rates to automatically learn the feature importance weights, forming complementary enhancement with the subsequent attention mechanism. In the parallel multi-branch structure, the convolutional kernels with different dilation rates are calibrated through independent weight normalization layers, which can automatically balance the magnitude differences of features at different scales.
[0095] (2) FCSA module
[0096] The present invention proposes a frequency-domain spatial joint attention module (Frequency Channel-Spatial Attention, FCSA) based on multi-spectrum analysis and spatial feature enhancement. Its core consists of channel attention (FrequencyChannel Attention, FCA) and spatial attention (Spatial Attention, SA), and realizes feature calibration through frequency-domain to spatial-domain joint modeling (see Figure 5 ). It enables the model to more finely control the type of frequency information used for attention calculation, thereby better capturing the basic structure and gradual change features of sea surface temperature.
[0097] Channel attention usually extracts global information based on global average pooling (GAP). Global average pooling assumes that all spatial positions are equally important and learns channel weights through fully connected layers, but it may ignore the importance differences of different frequency components, and time-domain averaging cannot distinguish high-frequency (vortices, fronts) and low-frequency (ENSO) signals, and cannot effectively capture the frequency component changes in signals or satellite images, treating all types of information equally, especially those low-frequency components that reflect large-scale structures or slow change trends and the noise in the image. The FCSA module extracts feature information from the frequency domain perspective by introducing multi-spectrum DCT transform (discrete cosine transform). GAP is just the lowest frequency component of DCT, eliminating a large amount of other effective frequency domain information, and the features lack diversity. DCT enhances the channel feature representation by introducing multiple frequency components instead of GAP. The input is mapped to the frequency domain through DCT basis functions, explicitly separating different frequency components, and extracting features from the frequency domain perspective. The frequency domain compression (channel compression rate, r) and parameter sharing mechanism (DCT filter weight reuse) can maintain good performance while effectively reducing FLOPs compared to traditional channel attention.
[0098] The channel attention module is divided into a frequency-domain feature selection module and a weight generation module. The frequency-domain feature selection module, i.e., the DCT layer, constructs a multi-spectrum filter through discrete cosine transform (DCT) basis functions to select different frequency components, obtains the frequency-domain coordinates from a predefined mapping table, and dynamically generates filter weights. The weight generation module performs feature compression and weight recovery through two fully connected layers. After the input features retain the original information through residual connections, they sequentially enter the frequency-domain feature selection module and the weight generation module for feature enhancement. First, the DCT layer filters the key frequency components, and then the first fully connected layer reduces the dimension, with the output channels being the original number of channels divided by the channel compression ratio (the channel compression ratio is set to 16 in the experiment), and ReLU activation is used. Then, the second fully connected layer restores the channel dimension, and the unnormalized attention weights are output. Finally, the Sigmoid function is applied to generate the channel attention map, and through the feature reshaping operation, it is broadcast to the size of the original feature map. After the input features are extracted with multi-spectrum components by the DCT layer, the channel attention weights are generated through a compression-expansion structure, and finally, they are multiplied element-wise with the original features to achieve feature reweighting. By this method, the spectral information smoothing problem of traditional global average pooling (GAP) is overcome.
[0099] The spatial attention module uses a hierarchical convolutional structure to generate spatial dimension weights, expands the receptive field through cascaded convolutions, and balances local details and global context modeling. This part consists of two convolutional layers. The first convolutional layer has a kernel size of 3×3, and the output channels are the original number of channels divided by the spatial compression ratio (the spatial compression ratio is set to 2 in the experiment), and the ReLU activation function is used to capture local spatial context and reduce the channel dimension to save computational effort. The second convolutional layer has a kernel size of 3×3, and the output channels are 1, and the Sigmoid activation function is used to generate the spatial attention mask to highlight important spatial regions. There is a batch normalization (Batch Normalization, abbreviated as BN) layer behind each convolutional layer to improve the stability and convergence speed of the model. This weight map is then multiplied element-wise with the feature map processed by the channel attention to achieve feature reweighting, and finally, a fused output feature map is generated. The channel attention focuses on the importance of frequency-domain features, and the spatial attention strengthens the spatial position correlation. The two achieve joint calibration, strengthen the important feature response, suppress redundant information, and improve the model's feature selective attention ability.
[0100] In the sea surface temperature prediction task, FCSA significantly improves the model's ability to model complex ocean thermal features by introducing frequency-domain feature enhancement and dynamic frequency selection.
[0101] The FCSA formula is as follows:
[0102] ;
[0103] ;
[0104] ;
[0105] Among them, is the input tensor, is the channel attention module, is the spatial attention module, is the output of the channel attention module, is the output of the spatial attention module, is the output of the frequency-domain and spatial joint attention module; is the element-wise multiplication.
[0106] For the input of the model, the frequency-domain feature selection module realizes frequency-domain feature selection through the discrete cosine transform (DCT). The experiment realizes the screening of different frequency components through a predefined mapping table, and dynamically calculates the weight values position by position according to the frequency selection strategy. Define the DCT basis function:
[0107] ;
[0108] In the formula, is the frequency index, is the height and width of the feature map, is the spatial position.
[0109] After grouping the input feature map by channels, perform element-wise multiplication with the predefined DCT filter, and then sum along the spatial dimension. Define the weighted sum of frequency-domain features:
[0110] ;
[0111] In the formula, represents the channel grouping index, each group contains channels ( is the number of frequency components).
[0112] Concatenate the spectral values of different frequency components into a complete spectral vector. Define the spectral concatenation operation:
[0113] ;
[0114] In the formula, through operation, concatenate the spectral values of each group of channels to generate the global frequency feature .
[0115] Map the spectral vector to the attention weights through two fully connected layers, and use the Sigmoid activation function to limit the weights within the range of [0,1]. Define the attention weight generation formula:
[0116] ;
[0117] In the formula, represents the Sigmoid activation function, represents two fully connected layers.
[0118] Multiply the attention weights with the original input feature map channel by channel to enhance important features and suppress redundant information. Define the feature map weighting formula:
[0119] ;
[0120] In the formula, is the original input tensor, is the attention weight, represents channel-by-channel multiplication.
[0121] Obtain the channel attention module formula:
[0122] ;
[0123] Spatial attention module formula:
[0124] ;
[0125] In the formula, is the output of the channel attention module, represents the output of the spatial attention module, represents the Sigmoid activation function, and both represent the Conv2D function with a convolution kernel size of 3×3. The former has a stride of 1, and the latter has a stride of , and both use the same padding method; B represents the batch normalization layer.
[0126] Activation function formula:
[0127] ReLU: ;
[0128] Sigmoid: ;
[0129] In the formula, represents the input SST data, represents the activation function.
[0130] 2. Convolutional Long Short-Term Memory Network Module
[0131] The convolutional long short-term memory network module uses the ConvLSTM2D layer as the core component, and its network architecture includes two key parts: the ConvLSTM2D layer and the 1×1 convolutional layer (see Figure 6 ). Inside the ConvLSTM2D layer, multiple ConvLSTM units are recursively unfolded along the time dimension to process the input sequence. Each unit contains an input gate, a forget gate, an output gate, and a candidate memory unit. The output channels are set to 24, and both the gating and state updates are achieved through convolutional operations with a kernel size of 3×3 and a stride of 1, which completes the temporal modeling while maintaining the input spatial resolution; the 1×1 convolutional layer realizes the efficient fusion of multi-channel features. Compared with the limitation of traditional LSTM that can only process one-dimensional time series signals, ConvLSTM retains the two-dimensional spatial structure information while modeling in the time dimension through the local receptive field characteristics of the convolutional kernel. The input gate, forget gate, and output gate of ConvLSTM are all implemented through convolutional operations (time dimension + spatial dimension), enabling the network to capture spatio-temporal correlations synchronously. This characteristic makes it suitable for tasks with significant spatio-temporal correlations such as sea surface temperature prediction. Through the feature transfer mechanism of consecutive time steps, ConvLSTM can generate multi-channel feature maps containing spatio-temporal dynamic features, greatly improving its ability to extract spatial features compared with the flattened time series processing method of ordinary LSTM. The subsequent 1×1 convolutional layer then performs a non-linear combination of the multi-channel feature maps output by ConvLSTM through a cross-channel parameter sharing mechanism, projects the high-dimensional features to the single-channel space while maintaining the spatial resolution, effectively eliminating redundant features and enhancing the representation ability of key information, reflecting the superiority of spatio-temporal joint modeling. Finally, the future 20-day SST is predicted through rolling prediction.
[0132] ;
[0133] In the formula, is the output of the forget gate, is the Sigmoid activation function, is the weight matrix input to the forget gate, is the input at the current time step, is the weight matrix from the hidden state of the previous time step to the forget gate, is the hidden state of the previous time step, is the weight matrix from the cell state of the previous time step to the forget gate, is the cell state of the previous time step, is the bias term of the forget gate, is the convolutional operation, is the Hadamard product.
[0134] ;
[0135] wherein, is the output of the input gate, is the weight matrix input to the input gate, is the weight matrix from the hidden state of the previous time step to the input gate, is the weight matrix from the cell state of the previous time step to the input gate, is the bias term of the input gate.
[0136] ;
[0137] wherein, is the candidate cell state, is the hyperbolic tangent activation function, is the weight matrix input to the candidate cell state, is the weight matrix from the hidden state of the previous time step to the candidate cell state, is the bias term of the candidate cell state.
[0138] ;
[0139] wherein, is the cell state at the current time step.
[0140] ;
[0141] wherein, is the output of the output gate, is the weight matrix input to the output gate, is the weight matrix from the hidden state of the previous time step to the output gate, is the weight matrix from the cell state at the current time step to the output gate, is the bias term of the output gate.
[0142] ;
[0143] wherein, is the hidden state at the current time step.
[0144] In ConvLSTM, the forget gate determines which information in the cell state of the previous time step needs to be forgotten, and the output of the forget gate is a value between 0 and 1, representing the proportion of retention or discard at each position; the input gate controls the update degree of the input to the cell state at the current time step, and the output of the input gate determines which parts in the candidate cell state will be written into the current cell state ; the output gate determines the hidden state at the current time step Which information should be extracted from the cell state i.e., the final hidden state is the result obtained by filtering a part of the cell state through the output gate.
[0145] S2. Model Training and Testing
[0146] The deep learning model is trained using the training set, the parameters are adjusted using the validation set, and finally the model is evaluated using the test set
[0147] Spatio-temporal sequence samples are constructed through the sliding window algorithm. With a 40-day time step as the input, the SST distribution on the 41st day is predicted. Then, the SST on the 41st day is added to the end of this 40-day sequence, and the SST on the 1st day is removed to form a new 40-day sequence to predict the SST data on the 42nd day. This cycle is repeated to achieve rolling prediction of the SST data for the next 20 days, ensuring effective modeling of time continuity and spatial correlation. A comparison chart of the predicted values and true values of each model for the 20th day is plotted, and the average RMSE and average MAE for predicting the next 1 - 20 days are calculated. The RMSE and MAE formulas are as follows:
[0148] ;
[0149] ;
[0150] where represents the predicted SST value, represents the actually observed SST value, refers to the number of grid points in this dataset. The smaller the values of RMSE and MAE, the better the prediction performance of the model.
[0151] The deep learning model of the present invention adopts a systematic hyperparameter optimization strategy to achieve efficient training and stable convergence of the model on complex spatio-temporal data. In this embodiment, the batch size is set to 4. The noise perturbation brought by mini-batch gradient descent can make the loss surface escape from local minima. The training epoch size is 90. The optimal model is saved through a callback function, and the weight file is stored in the HDF5 format to ensure the reproducibility of the experiment. For the characteristics of the regression task, the mean squared error (MSE) is selected as the loss function. Combining data normalization processing can effectively eliminate the interference of dimension differences on gradient updates. The ReLU and Sigmoid activation functions are used to effectively alleviate the gradient vanishing problem and constrain the predicted values within a reasonable temperature range. A learning rate decay strategy is introduced for adaptive adjustment of the learning rate. When the validation set loss has not decreased for 5 consecutive epochs, the learning rate is reduced by a factor Perform attenuation, and at the same time set early stopping as a regularization method. If the validation set loss has not improved for 20 consecutive epochs, terminate the training and roll back to the historical optimal weights. This strategy dynamically evaluates the generalization ability of the model and intervenes in a timely manner before the risk of overfitting increases, enabling the model to achieve better prediction results.
[0152] Evaluate the generalization ability and prediction accuracy of the model on the test set:
[0153] Figure 7 In (a) and (b), the average RMSE and MAE comparisons of different models predicting for 1 - 20 days are shown respectively. Table 1 calculates the average RMSE and MAE of different models predicting for 1 day, 5 days, 10 days, 15 days, and 20 days on the test set. From Figure 7 and Table 1, it can be seen that from LSTM, ConvLSTM to MFE - ConvLSTM, the average RMSE and average MAE of predicting the 20th day are continuously decreasing, and the prediction effect is getting better and better.
[0154] Table 1 Average RMSE and MAE
[0155]
[0156] Figure 8 and Figure 9 are the RMSE distribution map and MAE distribution map of different models predicting the 20th day. The RMSE (MAE) distribution map is an error distribution map drawn by averaging the RMSE (MAE) of predicting the 20th day for each point on the test set. The experimental results show that Figure 8 and Figure 9 the traditional LSTM shown in (a) and Figure 8 and Figure 9 the ConvLSTM model shown in (b) in have significant error accumulation phenomena in the prediction of the land - edge sea area. In contrast, the MFE - ConvLSTM of the present invention ( Figure 8 and Figure 9 shown in (c)) model exhibits better prediction performance and achieves better results in the prediction of the land - edge sea area; in terms of the prediction accuracy of the entire sea area, the error of the model is lower. This model effectively captures the complex spatio - temporal change characteristics at the land - sea junction through a multi - scale frequency - domain enhancement mechanism, thereby improving the local prediction accuracy while significantly enhancing the prediction stability across the entire sea area.
[0157] Step 3, after pre - processing the sea surface temperature data of the sea area to be predicted, input it into the qualified model for sea surface temperature prediction.
[0158] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A sea surface temperature prediction method based on multi-scale frequency domain enhancement, characterized in that It includes the following steps: Step 1: Preprocess the sea surface temperature dataset obtained from the ERA5 reanalysis data of the European Centre for Medium-Range Weather Forecasts, and divide it into a training set, a validation set, and a test set; Step 2: Input the data of the training set into the deep learning model MFE-ConvLSTM for training, use the validation set to adjust the parameters, and finally evaluate the model accuracy with the data of the test set; The deep learning model MFE-ConvLSTM consists of two parts: a multi-scale frequency domain enhancement module and a convolutional long short-term memory network module; The multi-scale frequency domain enhancement module includes three parallel branches, and each branch contains a main branch and a sub-branch; each main branch is sequentially composed of a multi-scale cyclic dilated convolution module and a frequency domain-spatial joint attention module; the three main branches are respectively configured with cyclic dilated convolutions with dilation rates of 1, 2, and 4; the sub-branch adjusts the channel dimension through a 3D convolution kernel with a kernel size of 1×1×1; the output of the sub-branch and the main branch are added element by element, and then feature fusion and non-linear enhancement are performed through the ReLU activation function; the channel features output by the three parallel branches are concatenated along the channel dimension to form a composite multi-scale feature tensor, and then a 3D convolution kernel with a size of 1×1×1 is used to project the concatenated high-dimensional features to the target dimension, and after ReLU activation, it is input to a 3D convolution kernel with a size of 1×1×1 and then output; The convolutional long short-term memory network module includes two key parts: a ConvLSTM2D layer and a 1×1 two-dimensional convolutional layer; inside the ConvLSTM2D layer, multiple ConvLSTM units are recursively unfolded through the time dimension to process the input sequence, and each ConvLSTM unit contains an input gate, a forget gate, an output gate, and a candidate memory unit, and the output channels are 24, and the gating and state updates are all achieved through convolution operations with a kernel size of 3×3 and a stride of 1, and the temporal modeling is completed while maintaining the input spatial resolution; the 1×1 convolutional layer is used to achieve the fusion and output of multi-channel features; Step 3: After preprocessing the sea surface temperature data of the sea area to be predicted, input it into the qualified model for sea surface temperature prediction.
2. The sea surface temperature prediction method based on multi-scale frequency domain enhancement according to claim 1, wherein, The multi-scale cyclic dilation convolution module extends the data dimension to three dimensions, uses cyclic data padding at the start and end of the time series, and sets the cyclic padding to be performed only at the head and tail in the time dimension; the padding amount is calculated as follows: ; where is the dilation rate, which controls the sampling interval between elements inside the convolutional kernel; is the convolutional kernel size, that is, the kernel size for expanding the receptive field in the time dimension.
3. A sea surface temperature prediction method based on multi-scale frequency domain enhancement according to claim 1, characterized in that In the parallel structure of the multi-scale cyclic dilated convolution module, weight normalization is used to wrap the multi-scale cyclic dilated convolution module to normalize the weights of the convolution kernels, and the normalized weights are decomposed into a direction vector and a scaling factor : ; In the formula, is the normalized weight, is the unnormalized weight vector representing the direction of the weight, is the learnable scaling factor representing the magnitude of the weight, is the L2 norm of 4. A sea surface temperature prediction method based on multi-scale frequency domain enhancement according to claim 1, characterized in that, The frequency domain-spatial joint attention module is composed of a channel attention module and a spatial attention module, and realizes feature calibration through frequency domain-spatial joint modeling; The channel attention module includes a frequency domain feature selection module and a weight generation module. The frequency domain feature selection module, that is, the DCT layer, constructs a multi-spectrum filter through discrete cosine transform basis functions to select different frequency components, obtains the frequency domain coordinates from a predefined mapping table, and dynamically generates filter weights; the weight generation module performs feature compression and weight recovery through two fully connected layers; The spatial attention module consists of two convolutional layers. The first convolutional layer has a kernel size of 3×3, and the number of output channels is the original number of channels divided by the spatial compression rate. The ReLU activation function is used to capture local spatial context and reduce the channel dimension. The second convolutional layer has a kernel size of 3×3, and the number of output channels is 1. The Sigmoid activation function is used to generate a spatial attention mask. There is a batch normalization layer after each convolutional layer. This weight map is multiplied element-wise with the feature map processed by the channel attention module to achieve re-weighting of features, and finally a fused output feature map is generated.
5. A sea surface temperature prediction method based on multi-scale frequency domain enhancement according to claim 4, characterized in that In the frequency domain feature selection module, after the input features retain the original information through residual connections, they enter the frequency domain feature selection module and the weight generation module in sequence for feature enhancement: First, the DCT layer is used to filter out key frequency components, and then the first fully connected layer is used for dimensionality reduction, with the number of output channels being the original number of channels divided by the channel compression rate, and ReLU activation is used; then the second fully connected layer is used to restore the channel dimension and output unnormalized attention weights; finally, the Sigmoid function is applied to generate a channel attention map, and through feature reshaping operations, it is broadcast to the size of the original feature map. After the input features are extracted with multi-spectrum components by the DCT layer, channel attention weights are generated through a compression-expansion structure, and finally they are multiplied with the original features channel by channel to achieve re-weighting of features.
6. The method for predicting sea surface temperature based on multi-scale frequency domain enhancement according to claim 1, wherein, The formula of the frequency domain spatial joint attention module is as follows: ; ; ; Among them, is the input tensor, is the channel attention module, is the spatial attention module, is the output of the channel attention module, is the output of the spatial attention module, is the output of the frequency-domain and spatial joint attention module; is element-wise multiplication.
7. A sea surface temperature prediction method based on multi-scale frequency domain enhancement according to claim 4, characterized in that The discrete cosine transform basis function is shown as follows: ; wherein, is the frequency index, is the height and width of the feature map, is the spatial position.
8. A sea surface temperature prediction method based on multi-scale frequency domain enhancement according to claim 4, characterized in that, The formula of the channel attention module is as follows: ; The formula of the spatial attention module: ; Wherein, is the output of the channel attention module, represents the output of the spatial attention module, represents the Sigmoid activation function, and both represent the Conv2D function with a convolution kernel size of 3×3. The former has a stride of 1, and the latter has a stride of , and both use the same padding method; B represents the batch normalization layer.
9. A sea surface temperature prediction method based on multi-scale frequency domain enhancement according to claim 1, characterized in that In the said ConvLSTM cell, the forget gate determines which information in the cell state of the previous time step needs to be forgotten, and the output of the forget gate is a value between 0 and 1, representing the proportion of retention or discard at each position; is a value between 0 and 1, representing the proportion of retention or discard at each position; The input gate controls the input at the current time step The degree of update to the cell state, the output of the input gate Determines which parts of the candidate cell state will be written into the current cell state ; the output gate determines the hidden state at the current time step which information should be extracted from the cell state i.e., the final hidden state is the result obtained after filtering a part of the cell state through the output gate.
10. A sea surface temperature prediction method based on multi-scale frequency domain enhancement according to claim 9, characterized in that, In the ConvLSTM cell, the output of the forget gate is: ; wherein, is the output of the forget gate, is the Sigmoid activation function, is the weight matrix input to the forget gate, is the input at the current time step, is the weight matrix from the hidden state of the previous time step to the forget gate, is the hidden state of the previous time step, is the weight matrix from the cell state of the previous time step to the forget gate, is the cell state of the previous time step, is the bias term of the forget gate, is the convolution operation, is the Hadamard product; The output of the input gate is: ; In the formula, is the output of the input gate, is the weight matrix input to the input gate, is the weight matrix from the hidden state of the previous time step to the input gate, is the weight matrix from the cell state of the previous time step to the input gate, is the bias term of the input gate; The candidate cell state is: ; In the formula, is the candidate cell state, is the hyperbolic tangent activation function, is the weight matrix input to the candidate cell state, is the weight matrix from the hidden state at the previous time step to the candidate cell state, is the bias term of the candidate cell state; The cell state at the current time step is: ; In the formula, represents the cell state at the current time step; The output of the output gate is: ; Wherein, is the output of the output gate, is the weight matrix input to the output gate, is the weight matrix from the hidden state at the previous time step to the output gate, is the weight matrix from the cell state at the current time step to the output gate, is the bias term of the output gate; The hidden state at the current time step is: ; wherein, is the hidden state at the current time step.
Citation Information
Patent Citations
Sea surface temperature prediction method fused with space-time granularity context neural network
CN117172355A
Short-term photovoltaic power prediction method based on three-dimensional meteorological data multi-source fusion
CN117424232A