A method and device for correcting sea surface temperature prediction value based on space-time axial attention
By employing a sea surface temperature (SST) prediction method based on spatiotemporal axial attention, and through feature extraction and attention computation using a convolutional input layer and encoder, the method addresses the problem of insufficient prediction accuracy in existing methods, achieving higher accuracy and stability in SST prediction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NORTH CHINA UNIVERSITY OF TECHNOLOGY
- Filing Date
- 2026-03-31
- Publication Date
- 2026-07-03
AI Technical Summary
Existing sea surface temperature prediction methods suffer from insufficient prediction accuracy, limited generalization ability, and lack of sensitivity to spatial relationships when faced with complex nonlinear and high-dimensional spatiotemporal data, resulting in prediction bias and discontinuity, which makes it difficult to meet operational needs.
A sea surface temperature (SST) prediction correction method based on spatiotemporal axial attention is adopted. Feature extraction and position encoding are performed through a convolutional input layer, the encoder performs attention calculations in the time, longitude, and latitude dimensions, and the decoder outputs the SST prediction results within the target time period, thereby enhancing the feature representation capability and accuracy.
It improves the accuracy and stability of sea surface temperature prediction, effectively integrates features from different dimensions, and enhances the accuracy of predicted values and the model's ability to be applied across regions.
Smart Images

Figure CN122332785A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of marine environment prediction technology, and in particular to a method and apparatus for correcting sea surface temperature predictions based on spatiotemporal axial attention. Background Technology
[0002] Currently, as sea surface temperature (SST) forecasting systems rapidly develop towards greater precision and heterogeneity, an increasing number of intelligent systems are emerging, collecting ocean data in real time through various large-scale sensor networks. To fully utilize the performance of these new technologies and achieve intelligent and precise forecasting, higher requirements are being placed on SST sensing networks, including high-precision sensing and accurate forecasting.
[0003] Currently, numerical models are the primary means of ocean temperature prediction. However, due to limitations in initial conditions, boundary conditions, external forcing, and simplification of physical processes, their outputs often deviate significantly from observed data. Therefore, bias correction based on numerical model outputs has become an important way to improve forecast accuracy. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a method and apparatus for correcting sea surface temperature predictions based on spatiotemporal axial attention, so as to improve the accuracy of sea surface temperature prediction.
[0005] According to one aspect of the present invention, a method for correcting sea surface temperature predictions based on spatiotemporal axial attention is provided, the method comprising: Acquire target sea surface temperature data, wherein the target sea surface temperature data includes at least temperature data, spatial information, and time information, and the spatial information includes longitude information and latitude information; The target sea surface temperature data is input into a pre-trained target sea surface temperature prediction correction model to obtain the target sea surface temperature prediction result output by the target sea surface temperature prediction correction model. The target sea surface temperature prediction correction model includes a convolutional input layer, an encoder, and a decoder. The convolutional input layer is used to extract features, encode time, and encode location on the target sea surface temperature data to obtain a target feature vector. The encoder is used to perform time-dimension attention calculation, longitude-dimension attention calculation, and latitude-dimension attention calculation on the target feature vector to obtain a first target output vector; The decoder is configured to output a target sea surface temperature prediction result within a target time period based on the first target output vector and the historical prediction value output by the decoder, wherein the historical prediction value is the sea surface temperature prediction value output by the decoder based on the first target output vector.
[0006] According to another aspect of the present invention, a sea surface temperature prediction correction device based on spatiotemporal axial attention is provided, the device comprising: The acquisition module is used to acquire target sea surface temperature data, wherein the target sea surface temperature data includes at least temperature data, spatial information and time information, and the spatial information includes longitude information and latitude information; The prediction module is used to input the target sea surface temperature data into a pre-trained target sea surface temperature prediction correction model and obtain the target sea surface temperature prediction result output by the target sea surface temperature prediction correction model. The target sea surface temperature prediction correction model includes a convolutional input layer, an encoder, and a decoder. The convolutional input layer is used to extract features, encode time, and encode location on the target sea surface temperature data to obtain a target feature vector. The encoder is used to perform time-dimension attention calculation, longitude-dimension attention calculation, and latitude-dimension attention calculation on the target feature vector to obtain a first target output vector; The decoder is configured to output a target sea surface temperature prediction result within a target time period based on the first target output vector and the historical prediction value output by the decoder, wherein the historical prediction value is the sea surface temperature prediction value output by the decoder based on the first target output vector.
[0007] According to another aspect of the present invention, an electronic device is provided, comprising: Processor; and Stored program memory, The program includes instructions that, when executed by the processor, cause the processor to perform any of the above-described methods for correcting sea surface temperature predictions based on spatiotemporal axial attention.
[0008] According to another aspect of the present invention, a non-transient computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to perform any of the above-described sea surface temperature prediction correction methods based on spatiotemporal axial attention.
[0009] One or more technical solutions provided in this invention embodiment, after acquiring target sea surface temperature (SST) data, input the target SST data into a pre-trained target SST prediction correction model. This target SST prediction correction model includes a convolutional input layer, an encoder, and a decoder. The convolutional input layer can perform feature extraction, temporal encoding, and positional encoding on the target SST data to obtain a target feature vector. The encoder performs attention calculations on the target feature vector in three dimensions to obtain a first target output vector. The decoder outputs the target SST prediction result within the target time period based on the first target output vector and the historical prediction values output by the decoder. By applying this invention embodiment, temporal and positional encoding are performed on the extracted features during the model input stage, thereby enhancing the expression of the temporal and positional information contained in the original data. The encoder performs attention calculations on the target feature vector in the spatiotemporal, longitude, and latitude dimensions respectively to obtain the first target output vector, achieving effective fusion of features from different dimensions, further improving feature expression capabilities, and thus improving the accuracy of the SST prediction value output based on these features. Attached Figure Description
[0010] Further details, features, and advantages of the invention are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which: Figure 1 A flowchart illustrating the sea surface temperature prediction correction method based on spatiotemporal axial attention provided by the present invention; Figure 2 A schematic diagram of a training process for the sea surface temperature prediction correction model provided by the present invention; Figure 3 A schematic flowchart of the spatiotemporal axis attention mechanism provided by the present invention; Figure 4 A schematic diagram of a structure of the sea surface temperature prediction correction model provided by the present invention; Figure 5 A diagram showing the comparison of the effects of correcting models for different sea surface temperature predictions; Figure 6 A schematic diagram of a sea surface temperature prediction correction device based on spatiotemporal axial attention provided by the present invention; Figure 7 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present invention is shown. Detailed Implementation
[0011] Embodiments of the present invention will now be described in more detail with reference to the accompanying drawings. While some embodiments of the invention are shown in the drawings, it should be understood that the invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the invention. It should be understood that the accompanying drawings and embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the invention.
[0012] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.
[0013] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the following description. It should be noted that the concepts of "first", "second", etc., mentioned in this invention are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.
[0014] It should be noted that the terms "a" and "a plurality of" used in this invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0015] The names of the messages or information exchanged between the multiple devices in the embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0016] Existing sea surface temperature prediction correction techniques can be broadly classified into three categories: The first category consists of traditional statistical methods, such as Model Output Statistics (MOS) and Gaussian process regression. These methods correct for errors by establishing a linear or statistical relationship between the numerical model output and the observed values. Their advantages are simplicity and strong interpretability, but their ability to fit complex nonlinear relationships is limited.
[0017] The second category is machine learning methods, such as support vector machines and random forests. These methods can uncover nonlinear mapping relationships and improve correction capabilities to some extent compared to statistical methods. However, they are insufficient in feature extraction and dependency modeling capabilities when faced with large-scale, high-dimensional spatiotemporal data.
[0018] The third category comprises deep learning methods, including convolutional neural networks (CNNs), recurrent neural networks (such as LSTM and ConvLSTM), and Transformer-based models. These methods have demonstrated superior performance in weather and ocean forecasting, automatically extracting complex features and effectively modeling nonlinear spatiotemporal relationships. However, convolutional structures are typically adept at capturing local features and struggle to efficiently model long-range dependencies; recurrent structures are inefficient in modeling long sequences; and the standard Transformer, when processing regular grid-like data, lacks sensitivity to spatial relationships and struggles to fully capture the physical coupling between latitude and longitude dimensions.
[0019] While existing sea surface temperature (SST) prediction correction techniques have improved the accuracy of numerical models to some extent, they still have several shortcomings. Traditional statistical methods mostly rely on linear or weakly nonlinear assumptions, making it difficult to handle the highly nonlinear and complex spatiotemporal coupling relationships in ocean dynamic processes, and thus often fail in rapidly changing or extreme ocean environments. Machine learning methods, such as support vector machines and random forests, have overcome the limitations of linear models and can describe a certain degree of nonlinear mapping, but their modeling capabilities are mainly concentrated in low-dimensional feature spaces, making it difficult to effectively utilize the high-dimensional spatiotemporal structure of large-scale gridded SST data, thus limiting prediction accuracy and generalization ability. Deep learning methods, such as convolutional neural networks (CNNs) and convolutional long short-term memory networks (ConvLSTMs), have advantages in extracting local spatial features and temporal dependencies, but the receptive field of convolutional operations is limited, lacking the ability to capture long-range spatial dependencies and cross-dimensional interactions, while recurrent structures are inefficient when processing long-term series and suffer from gradient decay problems.
[0020] While the Transformer architecture, widely adopted in recent years, performs exceptionally well in natural language processing and time series modeling, its direct application to regular grid-type sea surface temperature (SST) data presents significant limitations. Firstly, standard location encoding schemes primarily target one-dimensional sequences, failing to adequately represent the spatial continuity and directionality between latitude and longitude grid points, resulting in insufficient sensitivity to spatial structure. Secondly, the computational complexity of global self-attention mechanisms increases quadratically with the input size, incurring substantial computational and storage overhead in large-scale grid fields, thus limiting practical business applications. Existing methods attempt to mitigate these issues through multi-dimensional attention mechanisms, but these are largely derived from video analysis or general time series prediction tasks, lacking optimization for the characteristics of physical field data like SST. Cross-dimensional interaction information is often not fully utilized, leading to continued bias in prediction results. Furthermore, most existing methods rely entirely on data-driven optimization without introducing reasonable physical constraints. Although this may improve error metrics, prediction results may exhibit local discontinuities or unevenness, or even generate anomalous fields that violate physical laws, severely impacting the scientific interpretability and business reliability of the results. At the same time, existing methods lack the ability to generalize across different regions, seasons, or longer time scales, and the models have poor stability when applied across regions, making it difficult to meet the needs of operational forecasting.
[0021] Based on this, the present invention provides a method and apparatus for correcting sea surface temperature (SST) predictions based on spatiotemporal axial attention. The method for correcting SST predictions based on spatiotemporal axial attention provided by the present invention can be applied to any electronic device with SST prediction function, such as a computer, server, industrial control computer, etc. The following describes the solution of the present invention with reference to the accompanying drawings: Figure 1 A flowchart illustrating the sea surface temperature prediction correction method based on spatiotemporal axial attention provided by the present invention may include the following steps: S101. Obtain target sea surface temperature data, wherein the target sea surface temperature data includes at least temperature data, spatial information and time information, and the spatial information includes longitude information and latitude information; S102. Input the target sea surface temperature data into a pre-trained target sea surface temperature prediction correction model, and obtain the target sea surface temperature prediction result output by the target sea surface temperature prediction correction model, wherein the target sea surface temperature prediction correction model includes a convolutional input layer, an encoder, and a decoder. The convolutional input layer is used to extract features, encode time, and encode location on the target sea surface temperature data to obtain a target feature vector. The encoder is used to perform time-dimension attention calculation, longitude-dimension attention calculation, and latitude-dimension attention calculation on the target feature vector to obtain a first target output vector; The decoder is configured to output a target sea surface temperature prediction result within a target time period based on the first target output vector and the historical prediction value output by the decoder, wherein the historical prediction value is the sea surface temperature prediction value output by the decoder based on the first target output vector.
[0022] In this embodiment of the invention, after acquiring the target sea surface temperature (SST) data, the target SST data is input into a pre-trained target SST prediction correction model. This model includes a convolutional input layer, an encoder, and a decoder. The convolutional input layer performs feature extraction, temporal encoding, and positional encoding on the target SST data to obtain a target feature vector. The encoder performs attention calculations on the target feature vector in three dimensions to obtain a first target output vector. The decoder outputs the target SST prediction result within the target time period based on the first target output vector and the historical prediction values output by the decoder. By applying this embodiment of the invention, temporal and positional encoding are performed on the extracted features during the model input stage, thereby enhancing the expression of the temporal and positional information contained in the original data. The encoder performs attention calculations on the target feature vector in the spatiotemporal, longitude, and latitude dimensions respectively to obtain the first target output vector, achieving effective fusion of features from different dimensions, further improving feature representation capabilities, and thus improving the accuracy of the SST prediction value output based on these features.
[0023] The target sea surface temperature (SST) data can be raw SST data, typically referring to seawater temperature data collected at the ocean surface or at different depths. It can also include temperature statistics such as daily average, weekly average, and monthly average temperatures. Raw SST data can be collected using any feasible method, such as measurements from buoys or survey vessels, or retrieved from radiation signals from satellite sensors. The collected raw SST data can also contain relevant collection information, such as time and location information. The time information can include a timestamp of the data collection, and the location information can include the latitude and longitude of the data collection location. As one possible implementation, the ocean area can be divided into multiple grids according to a preset resolution when collecting raw SST data. The latitude and longitude information of the raw SST data can be the latitude and longitude data of the observation point where the raw SST data was collected, or it can be the latitude and longitude data of the grid.
[0024] The collected raw sea surface temperature (SST) data can be input into a pre-trained target SST prediction correction model for SST prediction correction. This model can output SST prediction results for a target time period based on the raw SST data. The target time period can be set according to the actual application scenario, such as one day or one week. The SST prediction correction model is pre-trained based on historical SST data, as one possible implementation method. Figure 2 As shown, the sea surface temperature prediction correction model can be pre-trained through the following steps: S11. Obtain the training dataset, which includes multiple historical sea surface temperature data; The training dataset can be obtained from publicly available datasets or from historically collected sea surface temperature (SST) data. The methods for collecting historical SST data are explained above and will not be repeated here. As one possible implementation, historical SST data collected before the current raw SST data collection time can be stored in a database, thus using all or part of the historical SST data in the database as the training dataset. Each historical SST data storage contains a corresponding index of actual SST data within a target time period. This index can include the storage location of the actual SST data within the target time period, etc. The actual SST data can be used as label data for the historical SST data in subsequent model training.
[0025] S12. Convert the historical sea surface temperature data into a four-dimensional grid structure, and slice the four-dimensional grid structure data in the time dimension according to the preset latitude and longitude window size to obtain multiple local spatial samples with overlapping areas, and construct a historical sample set.
[0026] Historical sea surface temperature (SST) data is typically presented in a flat structure, containing only time, latitude and longitude, and various feature values. However, it fails to reflect temporal periodicity and spatial neighborhood relationships. To better represent the temporal, spatial, and feature information contained within SST data, this flat historical SST data can be converted into a four-dimensional grid structure. This four-dimensional grid structure includes a time dimension, a longitude dimension, a latitude dimension, and a feature dimension. These four dimensions form a four-dimensional spatiotemporal data cube. Filling this cube with historical SST data yields the corresponding four-dimensional grid structure data.
[0027] One possible implementation is to extract timestamps from historical sea surface temperature (SST) data. For example, all historical SST data can be deduplicated to obtain a timestamp sequence. Each timestamp can then be converted into a time feature vector containing a period and a trend vector. The period information includes month, week, and hour information, while the trend information can be the number of days since the prediction start time, which refers to the start time of the predicted SST data to be output. For instance, a timestamp can be converted into a vector in the form of [week 1, hour 0, month 1, day 0].
[0028] Secondly, longitude and latitude information can be extracted from historical sea surface temperature data to create a spatial index. Specifically, the longitude and latitude in the historical sea surface temperature data can be deduplicated and sorted according to preset rules (e.g., latitude from south to north, longitude from west to east), and each longitude and latitude can be assigned an integer index. For example: Latitude list: [20°N, 21°N, 22°N] → Indices 0, 1, 2 (a total of L latitudes); Longitude list: [120°E, 121°E, 122°E] → Indices 0, 1, 2 (a total of O longitudes); Based on the dimensions above, initialize an empty four-dimensional array with the shape: [time T × latitude L × longitude O × feature F]. For example, if time T = 48 (2 days, 1 time step every 6 hours); latitude L = 3 (20°N, 21°N, 22°N); longitude O = 3 (120°E, 121°E, 122°E); and feature F = 2 (sea surface temperature, salinity), then the initialized cube shape is [48 × 3 × 3 × 2], with each position temporarily empty.
[0029] After establishing the initial cube, flattened data can be filled into it. Specifically, based on the timestamp and latitude / longitude information of each record, the corresponding index can be found, and the feature values can be filled into the corresponding positions in the cube. For example, historical sea surface temperature data is 2023-01-01 00:00, 20°N, 120°E, 25.3, 34.1. Its corresponding time index is 0 (corresponding to the first timestamp); latitude index is 0 (corresponding to 20°N); longitude index is 0 (corresponding to 120°E). Therefore, the features "25.3, 34.1" can be filled into the [0, 0, 0, 0] and [0, 0, 0, 1] positions of the cube.
[0030] In this way, by transforming the scattered one-dimensional list into a grid structure like a cube, with each dimension having a specific meaning, the time dimension is a time step arranged in sequence, preserving the order and period of time; the longitude and latitude dimensions are arranged in geographical order, preserving the structure that adjacent longitude and latitude are spatial neighbors.
[0031] Four-dimensional spatiotemporal cubes typically cannot be directly used in model processing. Therefore, they can be broken down into multiple local spatiotemporal blocks as training samples. Each sample contains continuous time and local space, allowing the model to learn the sea surface temperature evolution patterns within local spatiotemporal regions. One possible implementation is to slide along the time dimension with a pre-set sliding step size. At each time segment, a slice is taken from the spatial plane with a specified latitude and longitude window size, generating multiple local spatial samples with overlapping areas. Based on the above example, if the time window size is set to 3, the time sliding step size to 1, and the spatial window size to 3×3 (i.e., a local region of 3 latitudes + 3 longitudes), and the spatial sliding step size is 1, taking one sample for each latitude / longitude step, then the model can first slide along the time axis with a step size, taking out continuous time segments containing 3 time points each time step is slid, such as time steps 0, 1, 2, 1, 2, 3, until all time steps are slid through, resulting in multiple time segments. Within each time segment, the spatial dimension slides. Specifically, for each time segment, a local spatial region is sampled on a latitude × longitude spatial plane according to the spatial sliding step size. For example, the first spatial sampling: latitude 0-2, longitude 0-2 (corresponding to 20°N~22°N, 120°E~122°E); the second spatial sampling: latitude 0-2, longitude 1-3 (corresponding to 20°N~22°N, 121°E~123°E); the third spatial sampling: latitude 1-3, longitude 0-2 (corresponding to 21°N~23°N, 120°E~122°E), and so on, until the entire spatial plane is sampled, resulting in multiple local spatial samples. These local spatial samples constitute the historical sample set, which is used to train the initial sea surface temperature prediction correction model.
[0032] By using the above techniques, multiple local spatial samples with time dependence and spatial locality are constructed. Each local spatial sample contains continuous time steps and is a local spatial region, which enables the model to learn the characteristics of sea temperature change over time and the correlation characteristics of sea temperature between adjacent sea areas. In addition, each sample has the same spatiotemporal window size, which facilitates batch training of the model.
[0033] S13. Input the historical sample set into the initial sea surface temperature prediction correction model, and obtain the sea surface temperature prediction value output by the sea surface temperature correction model; The sea surface temperature prediction and correction model used in this invention can be extended and improved based on the Transformer architecture and constructed using an overall Encoder–Decoder framework. Specifically, the sea surface temperature prediction and correction model may include a convolutional input layer, an encoder, and a decoder.
[0034] The convolutional input layer is used to extract features from the input data. As one possible implementation, this convolutional input layer may include a spatial feature extraction module, a location encoding module, a temporal feature embedding module, and a feature fusion module. The spatial feature extraction module is used to extract features from the historical sample set through convolution to obtain the first sea surface temperature feature; The location encoding module is used to encode the first sea surface temperature feature using a location encoding matrix to obtain the location feature; The time feature embedding submodule is used to extract timestamps from the target sea surface temperature data according to preset time index information at different levels, so as to obtain time features at different levels; for the time features at different levels, each time feature is embedded by the time embedding vector corresponding to the time feature at different levels to obtain time embedding features. The feature fusion module is used to fuse the location features and the time embedding features to obtain a first feature vector.
[0035] The spatial feature extraction module consists of several convolutional layers. These layers can convolve each local spatial sample in the historical sample set to extract the original features of the historical sample set. Specifically, each local spatial sample can be mapped to the same hidden dimension, which is a model hyperparameter and can be set according to the actual application scenario. For example, when the shape of the local spatial sample input to the model is [B, T, H, W, C_in] (batch size is B, time step is T, spatial size is H×W, and the original feature dimension of each grid point is C_in), after the convolution operation, the spatial dimension is transformed into a spatial sequence of S = H′× W′, making the tensor shape [B, T, S, d_model], where d_model is the set hidden dimension.
[0036] The original features can be input into the positional encoding module, which contains a learnable positional encoding matrix. The shape of this matrix can be determined based on the dimensions of the original features; for example, the dimensions of the positional encoding matrix could be [1, 1, S, d_model]. As one possible implementation, the positional encoding matrix is updated during model training via backpropagation, without requiring a predefined analytical function, and is continuously updated as a trainable parameter during model training. In actual computation, the positional encoding matrix Ps is added element-wise to the spatial feature sequence through a broadcast mechanism, thereby introducing a unique and learnable position vector for each spatial location: Xs = Z + Ps, where Z is the spatial feature sequence extracted by convolution, and Xs is the output after adding the positional encoding.
[0037] The temporal feature embedding module contains a learnable embedding matrix and uses this matrix to perform temporal embedding on local spatial samples to obtain corresponding embedded features. As mentioned above, the timestamps in the local spatial samples are converted into a unified format, which includes different levels of time information, such as year, month, and day information. The temporal feature embedding module can train the embedding matrix separately for different levels of time information to fully learn each time information. As a possible implementation, discrete temporal features at different levels in the local spatial samples can be extracted, such as year index, month index, and day index. For each level of discrete temporal feature, an independent embedding vector can be constructed. The embedding vectors are arranged in the order of timestamp hierarchy to form the embedding matrix. Training the embedding matrix completes the training of each embedding vector. The initial value of the embedding matrix can be randomly generated and automatically updated during model training through gradient backpropagation. The time index in the input sample is used as the query key to retrieve multiple embedding vectors corresponding to time attributes from the corresponding embedding matrix, thus forming a set of time embedding sub-vectors with consistent dimensions.
[0038] The first feature vector is obtained by concatenating the temporal embedding features and the positional features. In one possible embodiment, the feature fusion module can fuse the positional features and the embedding features through vector addition or concatenation. For example, all embedding features can be concatenated and input into a linear mapping layer to restore the specified embedding dimension. The fused temporal embedding vector is then added to or concatenated with the feature vector of the original input data in either the channel or position dimension. This allows the input data to possess explicit temporal structure information before entering the backbone network, enhancing the model's ability to capture temporal correlations during the feature learning stage.
[0039] Throughout the training process, both the temporal embedding matrix and the parameters of the linear mapping layer participate in end-to-end optimization updates, enabling the embedding process to automatically learn the time-periodic features and trend information most relevant to the task objective. Through this temporal feature embedding method, the model can obtain a more accurate and detailed temporal characterization in a high-dimensional feature space, thereby effectively improving the accuracy and stability of downstream prediction tasks.
[0040] The first feature vector obtained by feature fusion can be input into an encoder for further feature extraction. This encoder may include a first spatiotemporal axis attention module, a first 3D convolutional feedforward module, and a second 3D convolutional feedforward module. The first spatiotemporal axis attention module is used to perform a first attention calculation on the first feature vector in the latitudinal dimension based on the trainable model latitudinal length to obtain a first attention feature; perform an attention calculation on the first attention feature in the longitude dimension based on the trainable model longitude length to obtain a second attention feature; perform an attention calculation on the second attention feature in the time dimension based on the trainable time steps to obtain a third attention feature; and fuse the first feature vector and the third attention feature to obtain a fused attention feature.
[0041] The first spatiotemporal axis attention module can perform attention calculations on target features through a spatiotemporal axis attention mechanism. Since sea surface temperature (SST) data has strong spatiotemporal coupling characteristics, an attention mechanism oriented towards the spatiotemporal dimension can be used to construct the internal dependency structure of the SST data. Specifically, the spatiotemporal axis attention mechanism can perform attention calculations on local spatial samples from both temporal and spatial dimensions, while maintaining global connectivity and context awareness by combining temporal and spatial features.
[0042] In natural sea surface temperature field data, spatial local correlations are particularly significant. To enhance the model's ability to express spatial structure, we designed a strategy of first aggregating local features along the longitude and latitude axes. Specifically, firstly, on the longitude axis, attention weights are calculated along the longitude dimension for all points at each time step and latitude position to extract local spatial features. Subsequently, the calculation results are reshaped into a tensor of shape [B,T,H,C] (where... The batch size is represented by [B, T, H, C], the time step is represented by [T, H, C], and attention is calculated for all points at the same dimensional position along the dimensional axis to capture the spatial dependencies in the dimensional direction. Finally, the result is reshaped back into [B, T, H, C] and the spatiotemporal feature relationships between time series are sequentially coupled and fused along the time axis.
[0043] In the stacking process of the three-axis attention mechanism, information is transferred through residual connections. The features extracted from the previous axis are used as input for the attention calculation of the next axis, ensuring that the attention modeling of the next axis is based on a contextual environment that incorporates the semantic information of the previous axis. In this way, after multiple layers of processing along the time axis, the model gradually accumulates and fuses spatiotemporal features along the longitude, latitude, and time axes, achieving efficient hierarchical fusion and expression of spatial and temporal information.
[0044] For example, the spatiotemporal axis attention mechanism can perform spatiotemporal dimension attention calculation on local spatial samples using the following formula: (1) (2) (3) (4) Among them, X (1) ,X (2) ,X (3) , representing the intermediate results of the features after attention calculation on the dimension, longitude, and time axis, respectively; the time step is T, the latitude length is W, and the longitude length is H, all derived from the model hyperparameter settings. X out The first line represents the final feature representation output after three-axis fusion, i.e., the fused attention feature; the second line represents the reshaping operation that unfolds the input features into a one-dimensional sequence along the specified axis in order to perform serialized attention calculation; the third line represents the normalization operation of the LayerNorm layer, which is used to stabilize the feature distribution and improve the convergence of the model; the fourth line represents the axial attention sub-modules in the width, height and time dimensions, respectively, and all of them perform attention calculation through the standard multi-head attention mechanism.
[0045] Using the above techniques, independent multi-head self-attention layers are defined on the longitude, latitude, and time axes, respectively. Each axis's attention module is a one-dimensional position-sensitive structure, capable of capturing long-range dependencies in the corresponding dimension. In each dimension's attention mechanism, each attention head calculates different weight representations, and the multi-head results are finally concatenated and integrated to construct a feature representation with high-dimensional spatiotemporal awareness capabilities, providing richer spatiotemporal information support for downstream bias correction tasks.
[0046] like Figure 3 As shown, Figure 3 A schematic diagram of a spatiotemporal axis attention mechanism provided by the present invention: The input of the first spatiotemporal axis attention module is a spatiotemporal cube, which consists of spatiotemporal features and embedded features. The shape of both features is [T, W, H, 512], that is, time (T) × longitude (W) × latitude (H) × feature dimension (512).
[0047] To fully capture the dependencies of spatiotemporal data across different dimensions, the first spatiotemporal axis attention module is split into three parallel attention branches, calculating attention along the latitude, longitude, and time axes, respectively. Specifically, the latitudinal multi-head attention transforms the input cube [T, W, H, 512] into [T×W, H, 512], effectively merging time and longitude into a single dimension. This allows attention to be calculated only along the latitude (H) direction, enabling the model to learn the relationships between points at different latitudes at the same time and longitude. For each feature block of [T×W, H, 512], the multi-head attention mechanism calculates the latitudinal relationships, resulting in n_head attention outputs a1, a2, a3…an. The outputs from multiple heads are then concatenated to restore the feature dimension to 512. This can be achieved by normalizing the attention outputs using a softmax layer before concatenation.
[0048] The meridional multi-head attention branch transforms the input cube [T, W, H, 512] into [T×H, W, 512], merging time and latitude into a single dimension, allowing attention to be calculated only in the longitude (W) direction. Here, the association in the longitude direction is also calculated through the multi-head attention mechanism, and the output dimension after splicing is 512.
[0049] The time-axis multi-head attention branch transforms the input cube [T, W, H, 512] into [H×W, T, 512], merging latitude and longitude into a single dimension, allowing attention to be calculated only in the time (T) direction. The multi-head attention mechanism is used to calculate the correlation in the time direction, and the concatenated output has a dimension of 512.
[0050] The attention outputs of the three axes are first added and fused together, and then concatenated with the original first feature vector to finally obtain a fused feature that includes latitudinal, longitudinal, temporal, and original features.
[0051] By employing the above techniques, attention for each axis focuses on only one dimension, resulting in more efficient computation and avoiding the high complexity of global attention. Simultaneously capturing spatial (latitude and longitude) and temporal dependencies makes it highly suitable for modeling spatiotemporal data such as sea surface temperature. Using multi-head attention for each axis allows the model to learn finer-grained relationships in different subspaces.
[0052] The fused features of the first spatiotemporal axis attention output can be input to the first 3D convolutional feedforward module. The first 3D convolutional feedforward module is used to perform 3D convolution on the fused attention features and the first feature vector to obtain the first convolutional feature. The second 3D convolutional feedforward module is used to perform convolution on the first convolutional feature and the target attention feature to obtain the first target output vector.
[0053] 3D convolutional feedforward networks are deep learning models that combine feedforward neural networks (FFNs) with 3D convolution kernels. The core of these models is used to process three-dimensional spatial structure data or temporal-spatial fusion data with a time dimension. Through sliding computation of the 3D convolution kernel on a three-dimensional tensor, local spatiotemporal / spatial features are automatically extracted. Finally, tasks such as classification and regression are completed through fully connected layers. The first 3D convolutional feedforward module can perform 3D convolution on the target attention features output by the first spatiotemporal axis attention module to extract the fused features. As one possible implementation, the input to the first 3D convolutional feedforward module can be the target attention output by the first spatiotemporal axis attention module and the first feature vector obtained by fusing temporal and spatial features. Specifically, the first feature vector and the target attention features can be residually concatenated and input into the first 3D convolutional feedforward module to avoid network degradation and enhance feature reuse.
[0054] The first convolutional feature output by the first 3D feedforward convolutional module can be input into the second 3D feedforward convolutional module. Specifically, a normalization layer can be connected between the first and second 3D feedforward convolutional modules. The normalization layer can normalize the first convolutional feature output by the first 3D feedforward convolutional module, outputting the normalized feature as [B, T, H, W, d_model]. This normalized feature and the fused feature are then input into the second 3D feedforward convolutional module for further feature extraction to obtain the first target output vector. The dimensions of the first target output vector are [B, d_model, T, H, W].
[0055] The first target output vector can be input to the decoder for decoding. The decoder's input can be the first target output vector and historical prediction values, which are the previous prediction results output by the decoder based on the first target output vector. Sea surface temperature (SST) data prediction typically requires predicting SST data results within a preset time period; that is, the prediction result is usually a SST data sequence, which contains multiple SST data prediction values arranged in chronological order. Therefore, the historical prediction values output based on the first target output vector can be used as part of the decoder's input.
[0056] The decoder may include a second spatiotemporal axis attention module, a third 3D convolutional feedforward module, a cross attention module, and an output module; The second spatiotemporal axis attention module is used to perform time dimension attention calculation, longitude dimension attention calculation, and latitude dimension attention calculation on the historical prediction values to obtain the second target output vector; the attention calculation process here can be referred to the above description, and will not be repeated here.
[0057] The third 3D convolutional feedforward module is used to perform 3D convolution on the second target output vector to obtain the second convolutional feature. The cross-attention module is used to perform cross-attention calculation based on the second convolutional features and the first target output vector to obtain the target output features.
[0058] Cross-attention is a variant of the attention mechanism. Its core feature is that it allows the model to use one set of queries to focus on another set of different keys and values, thereby establishing a relationship between the two sets of sequences. In this invention, the cross-attention module can perform cross-attention calculation on the first target output vector and the second convolutional features output by the third 3D feedforward convolution module to obtain the attention output.
[0059] The output module is used to output the target sea surface temperature prediction result based on the target output features.
[0060] In one possible embodiment, normalization layers can be included between the third 3D feedforward convolutional module and the cross-attention module, as well as between the cross-attention module and the output module, to shape the feature dimensions into the target dimension. For example, during the model's spatial-temporal embedding stage, the input dimension is [B,T,H,W,C], and the output dimension is [B,T,S,d_model]; after entering the encoder, the dimension changes in the spatiotemporal axis attention mechanism as follows: the input latitude axis attention is reshaped to [B...]. T The output of [H,W,d_model] is [B,T,H,W,d_model]; the input longitude axis attention reshaping is [B T The output [W,H,d_model] is [B,T,H,W,d_model]; after being reshaped by the time axis input, it becomes [B H The output of [W,T,d_model] is [B,T,H,W,d_model]. The feedforward neural network [B,d_model, T, H, W] outputs [B,d_model, T, H, W], and the normalization layer outputs [B, T, H, W,d_model]. The inputs and outputs of each part in the decoder are the same as those in the encoder. Finally, the linear function reduces [B,T,H,W,d_model] to [B,T,H,W,C].
[0061] like Figure 4 As shown, Figure 4This is a schematic diagram of a sea surface temperature (SST) prediction correction model provided by the present invention. The model can include an encoder and a decoder. The core task of the encoder is to transform historical spatiotemporal data into a high-dimensional feature representation. Specifically, it includes input preprocessing, core feature extraction, and output. Input preprocessing includes spatial encoding and embedding. Spatial encoding is used to generate a learnable location encoding vector based on latitude and longitude, allowing the model to understand spatial neighborhood relationships. Temporal feature embedding is used to transform timestamps into high-dimensional vectors containing periods and trends. The sum of the two is used as the initial input to the encoder. The core feature extraction mechanism includes a multi-head attention mechanism along the spatiotemporal axes: attention is calculated on the time, longitude, and latitude axes respectively to capture spatiotemporal local correlations. A 3D convolutional feedforward network extracts spatiotemporal local features using 3D convolutions, and then enhances the feature representation through a feedforward network. Normalization layers are used to stabilize training and avoid gradient fluctuations. Residual connections sum the input and output of each submodule to alleviate gradient vanishing and enhance feature reuse. The high-dimensional features output by the encoder will serve as the key and value for cross-attention in the subsequent decoder.
[0062] The core task of the decoder is to generate future spatiotemporal sequences based on historical features and the current prediction state. Input preprocessing is similar to that of the encoder, including spatial encoding, learnable positional encoding, temporal feature embedding, and residual connections to ensure the consistency of input features. The core generation layer uses a spatiotemporal axis multi-head attention mechanism to capture the spatiotemporal correlations within the future sequence (self-attention). A 3D convolutional feedforward network extracts local spatiotemporal features of the future sequence. The cross-attention layer uses the decoder's current state as the query and the historical features output by the encoder as the key and value, dynamically focusing on information in the history most relevant to the current prediction. For example, when predicting the future sea surface temperature of a certain sea area, it automatically associates similar spatiotemporal patterns from the past. Normalization layers and residual connections are used to stabilize training, ensuring smooth gradient propagation to the output. Finally, a linear layer generates the prediction output (such as a 7-day sea surface temperature grid).
[0063] The sea surface temperature prediction correction model provided by this invention achieves multimodal feature fusion through spatial encoding, temporal embedding, and raw data fusion, enabling the model to simultaneously understand spatial location, temporal period, and raw observation values. It captures long-range dependencies through spatiotemporal multi-head attention and captures local spatiotemporal patterns through 3D convolution, taking into account both global and local features. When generating future predictions, the decoder can dynamically reference the most critical information from history, improving prediction accuracy.
[0064] S14. Calculate the target loss function value based on the actual sea surface temperature value corresponding to each historical sample in the historical sample set and the predicted sea surface temperature value.
[0065] S15. Adjust the parameters of the sea surface temperature prediction correction model based on the target loss function value until the target loss function value converges.
[0066] To further enhance the physical plausibility and spatial continuity of the model's prediction results, a Laplacian physical constraint term can be introduced into the objective loss function. The Laplacian term is typically used to describe the smoothness or rate of change of physical quantities in space, reflecting the degree of change of scalar fields (such as temperature and pressure) in the spatial dimension, and is commonly found in partial differential equations (PDEs). This invention, based on the idea of Physical Information Neural Networks (PINNs), embeds physical laws into the deep learning model training process, and constrains the spatial gradient of the predicted sea surface temperature field by introducing a Laplacian regularization term. Specifically, in the original mean squared error (MSE) loss function, a Laplacian physical constraint term is introduced. MSE Based on this, construct the joint loss function L total The expression is defined as formulas (5), (6), and (7).
[0067] (5) (6) (7) In equations (5) to (7) This represents the Laplace operator value at the spatial point (x, y). It describes the "average difference," or "curvature," of that point relative to its neighborhood. Indicates the grid spacing at latitude and longitude resolution; It is the Laplace physical constraint loss; This represents the weighting coefficient, used to balance data errors and physical constraints; N represents the preset number of grid points.
[0068] Based on the latitude and longitude information contained in ocean reanalysis data, the Laplace term can be calculated using the temperature difference between adjacent grid points. Specifically, for each grid point, the temperature difference between it and its neighboring grid points (upper, lower, left, and right) is calculated, and the Laplace operator value is obtained by approximating it using a second-order central difference. This operator reflects the spatial smoothness of the local temperature field, consistent with the physical characteristics of ocean thermal diffusion processes.
[0069] In the spatiotemporal correction of sea surface temperature (SST), the model not only relies on input factors such as historical temperature, salinity, and current velocity, but also needs to consider the spatial interactions and constraints of these physical quantities. By introducing a Laplace term, the model can obtain additional physical guidance during the optimization process, making the predicted SST distribution smoother and more continuous, and consistent with the laws of heat diffusion, thereby improving the physical consistency and spatial stability of the prediction results.
[0070] After calculating the target loss function value based on the actual sea surface temperature data corresponding to the predicted value and the training data, the parameters in the sea surface temperature prediction correction model can be adjusted by gradient descent and other parameter adjustment methods until the target loss function value converges. The sea surface temperature prediction correction model when the target loss function value converges is taken as the target sea surface temperature prediction correction model. Here, convergence can be achieved when the target loss function value is less than the preset loss threshold, or when the change in the target loss function value is less than the preset difference.
[0071] The process of the target sea surface temperature prediction correction model for processing the raw sea surface temperature data can be found in the training section. Here, we will only provide a brief explanation and will not elaborate further.
[0072] In one possible embodiment, the above method further includes: converting the target sea surface temperature data into four-dimensional grid structure data, wherein the four-dimensional grid structure is a cubic structure composed of time, latitude, longitude and temperature data; The four-dimensional grid structure data is sliced in the time dimension according to the preset latitude and longitude window size to obtain multiple local spatial data with overlapping areas, which are used as target input data. The step of inputting the target sea surface temperature data into a pre-trained target sea surface temperature prediction correction model includes: The target input data is input into a pre-trained target sea surface temperature prediction correction model.
[0073] In one possible embodiment, the convolutional input layer includes: a spatial feature extraction module, a position encoding module, a temporal feature embedding module, and a feature fusion module; The spatial feature extraction module is used to extract features from the target input data through convolution to obtain the target sea surface temperature features; The location encoding module is used to encode the target sea surface temperature features using a pre-trained location encoding matrix to obtain location features. The time feature embedding submodule is used to extract timestamps from the target sea surface temperature data according to preset time index information at different levels, so as to obtain time features at different levels; for the time features at different levels, the time features are embedded by time embedding vectors corresponding to the time features at different levels obtained through pre-training, so as to obtain time embedding features. The feature fusion module is used to fuse the location features and the temporal embedding features to obtain the target feature vector.
[0074] In one possible embodiment, the encoder includes: a first spatiotemporal axis attention module, a first 3D convolutional feedforward module, and a second 3D convolutional feedforward module. The first spatiotemporal axis attention module is used to perform a first attention calculation on the target feature vector in the latitudinal dimension based on the pre-trained model latitudinal length to obtain a first attention feature; perform an attention calculation on the first attention feature in the longitude dimension based on the pre-trained model longitude length to obtain a second attention feature; perform an attention calculation on the second attention feature in the time dimension based on the pre-trained time steps to obtain a third attention feature; and fuse the target feature vector and the third attention feature to obtain the target attention feature. The first 3D convolutional feedforward module is used to perform 3D convolution on the target attention features and the target feature vector to obtain the first convolutional feature; The second 3D convolutional feedforward module is used to perform 3D convolution on the first convolutional features and the target attention features to obtain the first target output vector.
[0075] In one possible embodiment, the decoder includes: a second spatiotemporal axis attention module, a third 3D convolutional feedforward module, a cross-attention module, and an output module; The second spatiotemporal axis attention module is used to perform time dimension attention calculation, longitude dimension attention calculation, and latitude dimension attention calculation on the historical prediction values to obtain the second target output vector; The third 3D convolutional feedforward module is used to perform 3D convolution on the second target output vector to obtain the second convolutional feature. The cross-attention module is used to perform cross-attention calculation based on the second convolutional feature and the first target output vector to obtain the target output feature; The output module is used to output the target sea surface temperature prediction result based on the target output features.
[0076] In this embodiment of the invention, the original flattened data is reconstructed into a four-dimensional spatiotemporal cube through spatiotemporal data reconstruction and a sliding window sampling mechanism. Based on this, training samples are generated by combining the time step and local spatial regions using a spatiotemporal sliding window mechanism. This method can preserve the spatiotemporal local structure and enhance the model's learning effect on sea surface temperature evolution patterns.
[0077] Furthermore, this invention employs a spatiotemporal axis attention mechanism, decomposing the attention mechanism into sequential execution of latitude axis attention → longitude axis attention → time axis attention. It models long-range dependencies separately on each axis, avoiding the computational complexity caused by multidimensional coupling while ensuring the consistency and globality of extracted features. This axis-specific strategy addresses the problems of the classic Transformer's insensitivity to spatial dimensions and insufficient interaction between time and space when processing spatiotemporal sequences.
[0078] By using physical constraints based on the Laplacian operator, a Laplacian operator regularization term is introduced into the loss function to constrain the smoothness and continuity of the prediction results in the spatial dimension. This method avoids the model generating overly discrete or physically inconsistent correction results, thus improving the physical consistency and interpretability of the results.
[0079] Traditional methods (such as MOS and Gaussian process regression) are often based on linear or weakly nonlinear relationships, which cannot characterize the highly nonlinear and complex coupling features of ocean dynamic processes. This invention overcomes the linear assumption of traditional statistical methods by utilizing deep learning, especially axial attention mechanisms, to model long-range dependencies in latitude, longitude, and time dimensions, thus avoiding the limitations of representation caused by linear assumptions.
[0080] Traditional machine learning methods (such as SVM and random forest) can learn nonlinear relationships, but they are insufficient in feature extraction and spatiotemporal dependency modeling when faced with high-dimensional, large-scale spatiotemporal data. This invention, through spatiotemporal data cube reconstruction and sliding window sampling, combined with spatiotemporal axial attention, explicitly captures multi-scale spatiotemporal dependencies, significantly improving the ability to model complex sea surface temperature evolution patterns.
[0081] Existing deep learning methods often rely solely on data-driven approaches and lack physical consistency constraints, potentially leading to predictions that do not conform to natural laws. This invention introduces a Laplace physical constraint term into the Transformer correction framework, ensuring that the prediction results remain spatially smooth and continuous, better aligning with the physical laws governing ocean temperature evolution.
[0082] like Figure 5 As shown, Figure 5 This section describes the correction effects of different models on the 7-day predictions of the MACOM (Mass Conservation Ocean Model, also known as the Mazu Ocean Model). Figure 5 As can be seen, since the error between the predicted and actual values was small in the first three days, the correction effects of each model did not differ significantly in this stage, but the spatiotemporal correction model still showed some improvement. From the fourth day onwards, the prediction error increased significantly, and other models still struggled to effectively correct the error. However, the spatiotemporal correction model provided by this invention demonstrated a correction capability significantly superior to other models, indicating that the model proposed in this invention has a clear advantage in improving the accuracy of MACOM predictions.
[0083] In summary, this invention achieves high-precision and physically consistent correction of sea surface temperature prediction by introducing learnable spatial coding, spatiotemporal axial attention, and Laplace physical constraints, with significantly better results than existing statistical, machine learning, and deep learning methods.
[0084] Based on the same inventive concept, the present invention also provides a sea surface temperature prediction correction device based on spatiotemporal axial attention, such as... Figure 6 As shown, the device may include: The acquisition module 601 is used to acquire target sea surface temperature data, wherein the target sea surface temperature data includes at least temperature data, spatial information and time information, and the spatial information includes longitude information and latitude information; The prediction module 602 is used to input the target sea surface temperature data into a pre-trained target sea surface temperature prediction correction model and obtain the target sea surface temperature prediction result output by the target sea surface temperature prediction correction model. The target sea surface temperature prediction correction model includes a convolutional input layer, an encoder, and a decoder. The convolutional input layer is used to extract features, encode time, and encode location on the target sea surface temperature data to obtain a target feature vector. The encoder is used to perform time-dimension attention calculation, longitude-dimension attention calculation, and latitude-dimension attention calculation on the target feature vector to obtain a first target output vector; The decoder is configured to output a target sea surface temperature prediction result within a target time period based on the first target output vector and the historical prediction value output by the decoder, wherein the historical prediction value is the sea surface temperature prediction value output by the decoder based on the first target output vector.
[0085] In one possible embodiment, the acquisition module is further configured to convert the target sea surface temperature data into four-dimensional grid structure data, wherein the four-dimensional grid structure is a cubic structure composed of time, latitude, longitude and temperature data; The four-dimensional grid structure data is sliced in the time dimension according to the preset latitude and longitude window size to obtain multiple local spatial data with overlapping areas, which are used as target input data. The step of inputting the target sea surface temperature data into a pre-trained target sea surface temperature prediction correction model includes: The target input data is input into a pre-trained target sea surface temperature prediction correction model.
[0086] In one possible embodiment, the convolutional input layer includes: a spatial feature extraction module, a position encoding module, a temporal feature embedding module, and a feature fusion module; The spatial feature extraction module is used to extract features from the target input data through convolution to obtain the target sea surface temperature features; The location encoding module is used to encode the target sea surface temperature features using a pre-trained location encoding matrix to obtain location features. The time feature embedding submodule is used to extract timestamps from the target sea surface temperature data according to preset time index information at different levels, so as to obtain time features at different levels; for the time features at different levels, the time features are embedded by time embedding vectors corresponding to the time features at different levels obtained through pre-training, so as to obtain time embedding features. The feature fusion module is used to fuse the location features and the temporal embedding features to obtain the target feature vector.
[0087] In one possible embodiment, the encoder includes: a first spatiotemporal axis attention module, a first 3D convolutional feedforward module, and a second 3D convolutional feedforward module. The first spatiotemporal axis attention module is used to perform a first attention calculation on the target feature vector in the latitudinal dimension based on the pre-trained model latitudinal length to obtain a first attention feature; perform an attention calculation on the first attention feature in the longitude dimension based on the pre-trained model longitude length to obtain a second attention feature; perform an attention calculation on the second attention feature in the time dimension based on the pre-trained time steps to obtain a third attention feature; and fuse the target feature vector and the third attention feature to obtain the target attention feature. The first 3D convolutional feedforward module is used to perform 3D convolution on the target attention features and the target feature vector to obtain the first convolutional feature; The second 3D convolutional feedforward module is used to perform 3D convolution on the first convolutional features and the target attention features to obtain the first target output vector.
[0088] In one possible embodiment, the decoder includes: a second spatiotemporal axis attention module, a third 3D convolutional feedforward module, a cross-attention module, and an output module; The second spatiotemporal axis attention module is used to perform time dimension attention calculation, longitude dimension attention calculation, and latitude dimension attention calculation on the historical prediction values to obtain the second target output vector; The third 3D convolutional feedforward module is used to perform 3D convolution on the second target output vector to obtain the second convolutional feature. The cross-attention module is used to perform cross-attention calculation based on the second convolutional feature and the first target output vector to obtain the target output feature; The output module is used to output the target sea surface temperature prediction result based on the target output features.
[0089] In one possible embodiment, the target sea surface temperature prediction correction model is pre-trained through the following steps: Obtain a training dataset, which includes multiple historical sea surface temperature data; The historical sea surface temperature data is converted into a four-dimensional grid structure, and the four-dimensional grid structure data is sliced in the time dimension according to a preset latitude and longitude window size to obtain multiple local spatial samples with overlapping areas, thus constructing a historical sample set. The historical sample set is input into the initial sea surface temperature prediction correction model, and the sea surface temperature prediction value output by the sea surface temperature correction model is obtained. The target loss function value is calculated based on the actual sea surface temperature value corresponding to each local spatial sample in the historical sample set and the predicted sea surface temperature value. The target loss function includes Laplace physical constraints. The parameters of the sea surface temperature prediction correction model are adjusted based on the target loss function value until the target loss function value converges, thus obtaining the target sea surface temperature prediction correction model.
[0090] In one possible embodiment, the Laplace physical constraint is expressed by the following formula:
[0091]
[0092] in, 2 (x,y) represents the Laplacian operator value at the spatial point (x,y), and h represents the grid spacing at latitude and longitude resolution; L Laplace It is the Laplace physical constraint loss, where N represents the preset number of grid points.
[0093] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in this invention comply with relevant laws and regulations and do not violate public order and good morals.
[0094] An exemplary embodiment of the present invention also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to cause the electronic device to perform a method according to an embodiment of the present invention.
[0095] An exemplary embodiment of the present invention also provides a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform a method according to an embodiment of the present invention.
[0096] An exemplary embodiment of the present invention also provides a computer program product, including a computer program, wherein, when executed by a computer's processor, the computer program is used to cause the computer to perform a method according to an embodiment of the present invention.
[0097] refer to Figure 7 The present invention will now be described in the form of a structural block diagram of an electronic device 700 that can serve as a server or client of the present invention, which is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0098] like Figure 7 As shown, the electronic device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. The RAM 703 may also store various programs and data required for the operation of the electronic device 700. The computing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0099] Multiple components in electronic device 700 are connected to I / O interface 705, including: input unit 706, output unit 707, storage unit 708, and communication unit 709. Input unit 706 can be any type of device capable of inputting information to electronic device 700. Input unit 706 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device. Output unit 707 can be any type of device capable of presenting information and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 708 may include, but is not limited to, disk and optical disk. Communication unit 709 allows electronic device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth™ devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.
[0100] The computing unit 701 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above. For example, in some embodiments, any of the above-described sea surface temperature prediction correction methods based on spatiotemporal axial attention can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 700 via ROM 702 and / or communication unit 709. In some embodiments, the computing unit 701 can be configured by any other suitable means (e.g., by means of firmware) to perform any of the above-described sea surface temperature prediction correction methods based on spatiotemporal axial attention.
[0101] The program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0102] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0103] As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0104] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0105] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0106] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.
Claims
1. A method for correcting sea surface temperature predictions based on spatiotemporal axial attention, characterized in that, The method includes: Acquire target sea surface temperature data, wherein the target sea surface temperature data includes at least temperature data, spatial information, and time information, and the spatial information includes longitude information and latitude information; The target sea surface temperature data is input into a pre-trained target sea surface temperature prediction correction model to obtain the target sea surface temperature prediction result output by the target sea surface temperature prediction correction model. The target sea surface temperature prediction correction model includes a convolutional input layer, an encoder, and a decoder. The convolutional input layer is used to extract features, encode time, and encode location on the target sea surface temperature data to obtain a target feature vector. The encoder is used to perform time-dimension attention calculation, longitude-dimension attention calculation, and latitude-dimension attention calculation on the target feature vector to obtain a first target output vector; The decoder is configured to output a target sea surface temperature prediction result within a target time period based on the first target output vector and the historical prediction value output by the decoder, wherein the historical prediction value is the sea surface temperature prediction value output by the decoder based on the first target output vector.
2. The method according to claim 1, characterized in that, The method further includes: The target sea surface temperature data is converted into four-dimensional grid structure data, wherein the four-dimensional grid structure is a cubic structure composed of time, latitude, longitude and temperature data; The four-dimensional grid structure data is sliced in the time dimension according to the preset latitude and longitude window size to obtain multiple local spatial data with overlapping areas, which are used as target input data. The step of inputting the target sea surface temperature data into a pre-trained target sea surface temperature prediction correction model includes: The target input data is input into a pre-trained target sea surface temperature prediction correction model.
3. The method according to claim 2, characterized in that, The convolutional input layer includes: a spatial feature extraction module, a position encoding module, a temporal feature embedding module, and a feature fusion module; The spatial feature extraction module is used to extract features from the target input data through convolution to obtain the target sea surface temperature features; The location encoding module is used to encode the target sea surface temperature features using a pre-trained location encoding matrix to obtain location features. The time feature embedding submodule is used to extract timestamps from the target sea surface temperature data according to preset time index information at different levels, so as to obtain time features at different levels; for the time features at different levels, the time features are embedded by time embedding vectors corresponding to the time features at different levels obtained through pre-training, so as to obtain time embedding features. The feature fusion module is used to fuse the location features and the temporal embedding features to obtain the target feature vector.
4. The method according to claim 3, characterized in that, The encoder includes: a first spatiotemporal axial attention module, a first 3D convolutional feedforward module, and a second 3D convolutional feedforward module. The first spatiotemporal axis attention module is used to perform a first attention calculation on the target feature vector in the latitudinal dimension based on the pre-trained model latitudinal length to obtain a first attention feature; perform an attention calculation on the first attention feature in the longitude dimension based on the pre-trained model longitude length to obtain a second attention feature; perform an attention calculation on the second attention feature in the time dimension based on the pre-trained time steps to obtain a third attention feature; and fuse the target feature vector and the third attention feature to obtain the target attention feature. The first 3D convolutional feedforward module is used to perform 3D convolution on the target attention features and the target feature vector to obtain the first convolutional feature; The second 3D convolutional feedforward module is used to perform 3D convolution on the first convolutional features and the target attention features to obtain the first target output vector.
5. The method according to claim 4, characterized in that, The decoder includes: a second spatiotemporal axis attention module, a third 3D convolutional feedforward module, a cross attention module, and an output module; The second spatiotemporal axis attention module is used to perform time dimension attention calculation, longitude dimension attention calculation, and latitude dimension attention calculation on the historical prediction values to obtain the second target output vector; The third 3D convolutional feedforward module is used to perform 3D convolution on the second target output vector to obtain the second convolutional feature. The cross-attention module is used to perform cross-attention calculation based on the second convolutional feature and the first target output vector to obtain the target output feature; The output module is used to output the target sea surface temperature prediction result based on the target output features.
6. The method according to any one of claims 1-5, characterized in that, The target sea surface temperature prediction correction model is pre-trained through the following steps: Obtain a training dataset, which includes multiple historical sea surface temperature data; The historical sea surface temperature data is converted into a four-dimensional grid structure, and the four-dimensional grid structure data is sliced in the time dimension according to a preset latitude and longitude window size to obtain multiple local spatial samples with overlapping areas, thus constructing a historical sample set. The historical sample set is input into the initial sea surface temperature prediction correction model, and the sea surface temperature prediction value output by the sea surface temperature correction model is obtained. The target loss function value is calculated based on the actual sea surface temperature value corresponding to each local spatial sample in the historical sample set and the predicted sea surface temperature value. The target loss function includes Laplace physical constraints. The parameters of the sea surface temperature prediction correction model are adjusted based on the target loss function value until the target loss function value converges, thus obtaining the target sea surface temperature prediction correction model.
7. The method according to claim 6, characterized in that, The Laplace physical constraint is expressed by the following formula: in, 2 (x,y) represents the Laplacian operator value at the spatial point (x,y), and h represents the grid spacing at latitude and longitude resolution; L laplace It is the Laplace physical constraint loss, where N represents the preset number of grid points.
8. A sea surface temperature prediction correction device based on spatiotemporal axial attention, characterized in that, The device includes: The acquisition module is used to acquire target sea surface temperature data, wherein the target sea surface temperature data includes at least temperature data, spatial information and time information, and the spatial information includes longitude information and latitude information; The prediction module is used to input the target sea surface temperature data into a pre-trained target sea surface temperature prediction correction model and obtain the target sea surface temperature prediction result output by the target sea surface temperature prediction correction model. The target sea surface temperature prediction correction model includes a convolutional input layer, an encoder, and a decoder. The convolutional input layer is used to extract features, encode time, and encode location on the target sea surface temperature data to obtain a target feature vector. The encoder is used to perform time-dimension attention calculation, longitude-dimension attention calculation, and latitude-dimension attention calculation on the target feature vector to obtain a first target output vector; The decoder is configured to output a target sea surface temperature prediction result within a target time period based on the first target output vector and the historical prediction value output by the decoder, wherein the historical prediction value is the sea surface temperature prediction value output by the decoder based on the first target output vector.
9. An electronic device, comprising: processor; as well as Stored program memory, The program includes instructions that, when executed by the processor, cause the processor to perform the method according to any one of claims 1-7.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-7.