MSTKAN model ENSO prediction method based on attention mechanism

By introducing multi-axis spatiotemporal attention mechanism and KAN network in ENSO prediction, the MSTKAN model solves the shortcomings of the existing technology in capturing complex nonlinear relationships and spatial directional characteristics, significantly improving the accuracy of ENSO prediction.

CN120105016AActive Publication Date: 2025-06-06TAIYUAN UNIVERSITY OF TECHNOLOGY
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510270468.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-06
Estimated Expiration
2045-03-07

AI Technical Summary

Technical Problem

Existing deep learning networks are difficult to effectively capture complex nonlinear relationships and spatial directional features in ENSO prediction, resulting in performance degradation in long-term prediction.

Method used

The MSTKAN model based on attention mechanism is adopted, and the introduction of multi-axis spatiotemporal attention module and Kolmogorov-Arnold network (KAN) is used to capture complex dependencies and directional features in ENSO data more flexible and precisely.

Benefits of technology

The MSTKAN model outperforms existing models in the long-term and short-term prediction accuracy of ENSO, providing more reliable and accurate prediction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120105016A_ABST
    Figure CN120105016A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of ENSO prediction methods, and particularly relates to an MSTKAN model ENSO prediction method based on an attention mechanism, and the method comprises the following steps: obtaining the monthly data of three important ocean and atmospheric variables, namely global upper ocean temperature abnormality, latitudinal wind stress and longitudinal wind stress; the primary processing of the data mainly comprises the following steps; constructing a prediction model; carrying out network training optimization by adopting an Adam algorithm; and training and verifying the model. The MSTKAN can capture fine-grained features in different directions through a multi-axis space-time attention mechanism of the MSTKAN, and inherent directional dependence in ENSO data is fully considered. And a Kolmogorov-Arnold network is introduced to replace a traditional linear layer, and a learnable nonlinear function is used to more flexibly and accurately approach to output features, so that the ability of the model in capturing a complex nonlinear relationship is enhanced. By improving the architecture, the MSTKAN exceeds the existing mainstream model in long-term and short-term prediction accuracy of ENSO, and a more reliable prediction result is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of ENSO prediction methods, and specifically relates to an ENSO prediction method based on an MSTKAN model of an attention mechanism. Background Art

[0002] Accurate prediction of real-time ocean-atmosphere conditions remains a long-standing challenge in climate research and has important scientific and economic significance. Take the El Niño-Southern Oscillation (ENSO) as an example. This is the most significant ocean-atmosphere phenomenon in the tropical Pacific Ocean. It occurs on an interannual scale and is mainly manifested as basin-wide sea surface temperature (SST) anomalies and the associated atmospheric circulation anomalies. Studies have shown that ENSO has a profound impact on the climate system and society through remote interactions between the atmosphere and the ocean. For example, the occurrence of ENSO events often triggers extreme weather on a global scale and even affects global crop yields in the following year. Therefore, in the past few decades, a large amount of research has been devoted to understanding and predicting ENSO.

[0003] With the deepening of observation and understanding of the El Nino-Southern Oscillation (ENSO), its simulation and prediction have also made significant progress. Physics-based dynamic models have always been the core tools for understanding the ENSO process and making predictions, but the inadequacy of these models in expressing key processes has led to deviations in climate system simulations and made long-term ENSO predictions challenging. At present, it is still extremely difficult to achieve ENSO predictions for more than one year relying on traditional physical models.

[0004] However, the rapid development of deep learning (DL) algorithms and their innovative applications in earth sciences have provided new perspectives for improving the modeling of weather and climate phenomena. Unlike models based on physical theories, data-driven DL models automatically capture the intrinsic relationship between input predictors and output predictors through neural networks without relying on explicit physical laws, thereby significantly improving the modeling accuracy of nonlinear systems, including ENSO prediction. For example, DL models have successfully achieved The sea surface temperature (SST) index is predicted more than 15 months in advance and effectively alleviates the spring predictability barrier (SPB). In addition, the DL model can more comprehensively predict the SST and atmospheric anomalies associated with ENSO, and generate more reasonable and robust results by considering the spatiotemporal dependencies of multiple anomaly fields.

[0005] In recent years, deep learning, especially neural networks (CNN, RNN, LSTM), has made significant progress in ENSO prediction, especially in capturing spatiotemporal dependencies and improving prediction accuracy. However, these models face challenges in balancing long-range dependencies and local feature extraction. To this end, Transformer-based models use self-attention mechanisms to solve the problem of long-range dependency modeling, but still face difficulties in fine-grained local feature extraction. Most spatiotemporal models focus on global dependencies and ignore the east-west and north-south spatial dependencies in ENSO data, which are crucial for accurate ENSO prediction. In addition, the traditional linear prediction layer cannot effectively capture the complex nonlinear relationships in ENSO evolution, especially in long-term climate trend prediction, and the model performance usually decreases. MSTKAN replaces the traditional linear layer by introducing the Kolmogorov-Arnold network (KAN), uses learnable nonlinear functions to more flexibly and accurately approximate the output features, thereby enhancing the model's ability to capture complex nonlinear relationships. Summary of the invention

[0006] In response to the technical problems existing in the prediction lead time of the above-mentioned deep learning network, the present invention provides an ENSO prediction method based on the MSTKAN model of the attention mechanism.

[0007] In order to solve the above technical problems, the technical solution adopted by the present invention is:

[0008] An ENSO prediction method based on the MSTKAN model of attention mechanism includes the following steps:

[0009] S1. Obtain monthly data of three important ocean and atmospheric variables: global upper ocean temperature anomaly, zonal wind stress and meridional wind stress;

[0010] S2. Preliminary data processing mainly includes: replacing missing values ​​and invalid values; setting the required latitude and longitude range; interpolating to a regular grid and normalizing; constructing input suitable for deep learning model processing; dividing the processed data into training set, validation set, and test set in turn;

[0011] S3. Construct prediction model: The main body of the MSTKAN model is divided into three parts: multi-layer convolutional encoder, multi-axis spatiotemporal attention module, and KAN prediction network. The multi-axis spatiotemporal attention module is divided into spatial attention module Intra-SA and temporal attention module Inter-TA, which are responsible for extracting features in both time and space dimensions and capturing complex dependencies in data. The KAN prediction network uses KAN to replace the traditional linear layer, and two KAN layers are used here.

[0012] S4. Adopt Adam algorithm to optimize network training, including: setting learning rate to 1e-5, using customized learning rate change strategy, adopting Warm-up strategy, gradually increasing learning rate at the beginning of training, and then gradually decreasing learning rate as training progresses; improving the stability and efficiency of model training;

[0013] S5, training and validating the model: Use the model constructed in S3 and the training set and validation set obtained in S2 to train and validate the model. Propose a joint loss function based on the characteristics of El Niño and ocean and atmospheric variables. Save the best model in the validation process based on the evaluation indicators. Use the test set to experimentally test the effectiveness of the proposed model.

[0014] The method for preliminary processing of input data in S2 is:

[0015] First, the long-term trend and seasonal cycle are removed to calculate the monthly anomalies. The data are then interpolated to a regular grid with a latitude of 1 degree and a longitude of 2 degrees, and a latitude of 0.5 degrees from 5°S to 5°N and 1 degree in other areas. The land grid and missing values ​​are assigned to zero. The wind stress and temperature anomalies are then normalized by the spatial mean standard deviation. Finally, the nine layers of data are spliced ​​along the layer axis and input into the model.

[0016] The network structure of the MSTKAN model in S3 is:

[0017] The first module is Multi-layer Convolutional Encoder, which consists of three layers of convolution;

[0018] The second module is Multi-Axis Spatiotemporal Attention Structure, which stacks several layers, each of which contains a spatial attention module Intra-SA and a temporal attention module Inter-SA;

[0019] The third module, KAN Prediction Head, consists of global average pooling and two KAN layers.

[0020] The construction method of the multi-layer convolution encoder in S3 is: using a multi-layer convolution encoder module to further process the input data, including three convolution layers Conv, using a Leaky ReLU activation function, and applying Dropout to prevent overfitting.

[0021] The construction method of the multi-axis spatiotemporal attention module in S3 is as follows: Let X represent a set of input features, with a size of Where T, C, H and W are time, number of channels, height and width respectively.

[0022] The construction method of the spatial attention module Intra-SA is:

[0023] The Intra-SA block contains two parallel branches: horizontal attention (H-SA) and vertical attention (V-SA); in H-SA, the query, key and value are first generated, represented as and For multi-head attention, the number of heads of multi-head attention is set to M = 8; next, ,and Reshape into a 2D tensor with dimensions express H horizontal blocks; the attention features of H-SA are Indicates; afterwards, Reshape Back Then concatenate all the attention features to obtain Symmetrically, V-SA calculates vertical attention in the vertical direction to generate attention features Similarly, all V-SA features are reshaped and concatenated as The features of horizontal attention and vertical attention are concatenated to form the output of the spatial attention module.

[0024] The construction method of the temporal attention module Inter-TA is:

[0025] The Inter-TA block contains two parallel branches: horizontal attention (H-TA) and vertical attention (V-TA); in H-TA, X h First, it is segmented into H non-overlapping and co-localized horizontal regions. i∈{1,2,...,H}; Next, according to Generate queries, keys, and values Next, we will Reshape into a 2D tensor with dimensions Then, horizontal attention features are generated in the time dimension Afterwards, Reshape And concatenate the features of all heads along the channel dimension to generate Then, the H attention outputs Fold into represents the weighted features on the horizontal plane; symmetrically, vertical attention is applied to X v To obtain the weighted features of V-TA The features of horizontal attention and vertical attention are concatenated to form the output of the temporal attention module.

[0026] The method of the joint loss function used in S5 is:

[0027] Loss_var: This loss function is used to calculate the root mean square error (RMSE) between the variables predicted by the model and the true value; and average and sum over multiple dimensions to get the overall variable loss; the formula is as follows:

[0028]

[0029] Where: y pred represents the predicted value; y true represents the true value; B represents the batch size; T represents the time step; H and W represent the height and width respectively;

[0030] Loss_nino: This loss function is used to calculate the root mean square error (RMSE) between the variables predicted by the model and the true value; the formula is as follows:

[0031]

[0032] Where: n pred Represents the predicted value; n true represents the true value; N represents the number of samples; T represents the number of time steps;

[0033] Combine_loss: The combined loss function combines the two loss functions to comprehensively consider the error; the formula is as follows:

[0034] combine_loss=loss_var+loss_nino.

[0035] The evaluation index in S5 is:

[0036] Correlation Coefficient (Corr); The correlation coefficient is used to measure the linear correlation between the predicted value and the true value; the formula is as follows:

[0037]

[0038] Where: y true is the true value; y pred is the predicted value; and They are y true and pred The mean of ; n is the number of samples;

[0039] Root Mean Squared Error (RMSE) RMSE is used to measure the average error between the predicted value and the true value; the formula is as follows:

[0040]

[0041] Where: y true is the true value; y pred is the predicted value; n is the number of samples;

[0042] Mean Absolute Error (MAE) MAE is used to measure the average absolute error between the predicted value and the true value; the formula is as follows:

[0043]

[0044] Where: y true is the true value; y pred is the predicted value; n is the number of samples.

[0045] Compared with the prior art, the present invention has the following beneficial effects:

[0046] The MSTKAN of the present invention can capture fine-grained features in different directions through its multi-axis spatiotemporal attention mechanism, fully considering the inherent directional dependence in ENSO data. And by introducing the Kolmogorov-Arnold network (KAN), replacing the traditional linear layer, using learnable nonlinear functions, more flexibly and accurately approximating the output features, thereby enhancing the model's ability to capture complex nonlinear relationships. Through the improved architecture, MSTKAN surpasses existing mainstream models in the long-term and short-term prediction accuracy of ENSO, providing more reliable prediction results. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the implementation methods of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the implementation methods or the description of the prior art. Obviously, the drawings in the following description are only exemplary, and for ordinary technicians in this field, other implementation drawings can be derived from the provided drawings without creative work.

[0048] The structures, proportions, sizes, etc. illustrated in this specification are only used to match the contents disclosed in the specification so as to facilitate understanding and reading by persons familiar with the technology. They are not used to limit the conditions under which the present invention can be implemented, and therefore have no substantial technical significance. Any structural modification, change in proportion or adjustment of size shall still fall within the scope of the technical contents disclosed in the present invention without affecting the effects and purposes that can be achieved by the present invention.

[0049] Figure 1 This is the framework diagram of the MSTKAN model of the present invention

[0050] Figure 2 For the MSTKAN model Plots evaluating the index’s predictive skill;

[0051] Figure 3 A seasonal assessment graph of the skills associated with the present invention;

[0052] Figure 4 This is a comparison chart of the prediction and reality of the ENSO event of the present invention. DETAILED DESCRIPTION

[0053] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. These descriptions are only to further illustrate the features and advantages of the present invention, rather than to limit the claims of the present invention. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0054] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0055] This embodiment provides an ENSO prediction method based on the MSTKAN model of the attention mechanism. Figure 1-4 As shown, the following steps are included:

[0056] 1. Collect monthly data on ocean and atmospheric variables at different time points

[0057] The study area covers 92°E to 330°E, 20°S to 20°N. Three key variables were selected - namely, meridional and zonal wind stress, and temperature anomalies at seven depths (5m, 20m, 40m, 60m, 90m, 120m and 150m) of the upper ocean.

[0058] 2. Process the collected raw flux meteorological data

[0059] Monthly anomalies are calculated by removing long-term trends and seasonal changes in the climate state. The anomaly data in the selected longitude and latitude regions are interpolated onto a regular grid. The longitudinal resolution is 2 degrees, and the latitudinal resolution is 0.5 degrees within the range of 5 degrees south latitude to 5 degrees north latitude, and 1 degree in other areas. All land grids and missing values ​​are assigned zero. Normalization is performed by normalizing wind stress and temperature anomalies by the spatial mean standard deviation (SD) to eliminate the influence of magnitude differences during training. Finally, all normalized data are concatenated along the layer axis to construct a dataset containing nine layers and provide it as input to the model.

[0060] 3. Build the model

[0061] The MSTKAN model ENSO prediction method based on the attention mechanism is established; the basic architecture of the model includes the following layers: multi-layer convolutional encoder, multi-axis spatiotemporal attention network, KAN prediction module. The multi-axis spatiotemporal attention module can be divided into the spatial attention module Intra-SA and the temporal attention module Inter-TA. The KAN prediction network uses KAN to replace the traditional linear layer, and two KAN layers are used here.

[0062] The multi-layer convolution encoder module further processes the input data, including three convolutional layers Conv, which gradually extracts input features, uses a 3×3 convolution kernel, a step size of 1, and sets padding = 1; reduces the spatial resolution, increases the number of channels, and finally adjusts the input variables to a shape suitable for processing by the multi-axis spatiotemporal attention module, uses the Leaky ReLU activation function, and applies Dropout to prevent overfitting.

[0063] The multi-axis spatiotemporal attention network contains Intra-SA and Inter-TA modules. The Inter-SA block contains two parallel branches: horizontal attention (H-SA) and vertical attention (V-SA). Let X represent a set of input features with size Where T, C, H and W are time, number of channels, height and width respectively. For each time point t∈{1,2,...,T}, first normalized using layer normalization and split along the channel to obtain the input features of H-SA and V-SA and The formula is as follows:

[0064]

[0065] Among them, C′=C / 2.

[0066] In H-SA, we first generate the query, key, and value, represented as and For multi-head attention, the formula is as follows:

[0067]

[0068] in, and m∈1,2,...,M,represents the linear transformation matrices of query, key and value respectively, and M=8 is the number of heads of multi-head attention. and Reshape into a 2D tensor with dimensions Attentional Characteristics of H-SA The calculation is as follows: Afterwards, Reshape Back Then concatenate all the attention features to obtain Symmetrically, V-SA computes vertical attention in the vertical direction to generate attention features Similarly, all V-SA features are reshaped and concatenated as Final spatial attention module SA output Generates the following:

[0069]

[0070] For the temporal attention module Intra-TA, spatiotemporal attention is applied to the co-localized regions of a set of input sequences to extract temporal information. It has two parallel branches: horizontal attention (H-TA) and vertical attention (V-TA). We will Split into X h and Layer normalization is performed to generate the input features of H-TA and V-TA, and the formula is as follows:

[0071] (X h ,X v )=Split(LN(X)).

[0072] In H-TA, X h First, it is segmented into H non-overlapping and co-localized horizontal regions. i∈{1,2,...,H}. Next, according to Generate queries, keys, and values Next, Reshape into a 2D tensor with dimensions Then, attention features are generated in the time dimension The formula is as follows:

[0073]

[0074] Afterwards, Reshape And concatenate the features of all heads along the channel dimension to generate Then, these H attention outputs Fold into Represents the weighted features in the horizontal direction. Symmetrically, attention is applied to X v To obtain the weighted features of V-TA Generate the final temporal attention module TA output using the same process as that used to generate the final spatial attention module Intra-SA output By interweaving the multi-head spatial attention module SA and the temporal attention module TA in the MSTKAN network, it can better capture the long-distance dependencies of longitude and latitude regions, and the network can learn richer features to increase the credibility of the prediction results.

[0075] The KAN prediction network consists of a global average pooling and two KAN layers. It converts the features after encoding and attention processing into the final output prediction. It first reshapes the input into a shape suitable for processing, aggregates information through global average pooling, then generates the predicted output through KAN, and finally reshapes it into a shape that meets the output requirements.

[0076] 4. Model Training

[0077] In this example, the input data dimensions are [12, 9, 51, 120]. The target data dimensions are [20, 9, 51, 120]. In order to achieve the best training effect, the training batch size is set to 2, the number of iterations is set to 50, and the early stopping strategy is adopted. Finally, the trained model is saved and used to obtain the prediction results of the model on the test set.

[0078] 5. Model Evaluation

[0079] Correlation Coefficient (Corr). The correlation coefficient is used to measure the linear correlation between the predicted value and the true value. The formula is as follows:

[0080]

[0081] Where: y true is the true value. pred is the predicted value. and They are y true and pred The mean of . n is the number of samples.

[0082] Root Mean Squared Error (RMSE) RMSE is used to measure the average error between the predicted value and the true value. The formula is as follows:

[0083]

[0084] Where: y true is the true value, y pred is the predicted value and n is the number of samples.

[0085] Mean Absolute Error (MAE) MAE is used to measure the average absolute error between the predicted value and the true value. The formula is as follows:

[0086]

[0087] Where: y true is the true value, y pred is the predicted value and n is the number of samples.

[0088] Only the preferred embodiments of the present invention are described in detail above, but the present invention is not limited to the above embodiments. Various changes can be made within the knowledge scope of ordinary technicians in this field without departing from the purpose of the present invention, and various changes should be included in the protection scope of the present invention.

Claims

1. A MSTKAN model ENSO prediction method based on attention mechanism, characterized in that: The following steps are involved: S1. Obtain monthly data of three important ocean and atmospheric variables: global upper ocean temperature anomaly, zonal wind stress and meridional wind stress; S2. Preliminary data processing mainly includes: replacing missing values ​​and invalid values; setting the required latitude and longitude range; interpolating to a regular grid and normalizing; constructing input suitable for deep learning model processing; dividing the processed data into training set, validation set, and test set in turn; S3. Construct prediction model: The main body of the MSTKAN model is divided into three parts: multi-layer convolutional encoder, multi-axis spatiotemporal attention module, and KAN prediction network. The multi-axis spatiotemporal attention module is divided into spatial attention module Intra-SA and temporal attention module Inter-TA, which are responsible for extracting features in both time and space dimensions and capturing complex dependencies in data. The KAN prediction network uses KAN to replace the traditional linear layer, and two KAN layers are used here. S4. Adopt Adam algorithm to optimize network training, including: setting learning rate to 1e-5, using customized learning rate change strategy, adopting Warm-up strategy, gradually increasing learning rate at the beginning of training, and then gradually decreasing learning rate as training progresses; improving the stability and efficiency of model training; S5, training and validating the model: Use the model constructed in S3 and the training set and validation set obtained in S2 to train and validate the model. Propose a joint loss function based on the characteristics of El Niño and ocean and atmospheric variables. Save the best model in the validation process based on the evaluation indicators. Use the test set to experimentally test the effectiveness of the proposed model.

2. The ENSO prediction method based on the MSTKAN model of the attention mechanism according to claim 1 is characterized in that: The method for preliminary processing of input data in S2 is: First, the long-term trend and seasonal cycle are removed to calculate the monthly anomalies. The data are then interpolated to a regular grid with a latitude of 1 degree and a longitude of 2 degrees, and a latitude of 0.5 degrees from 5°S to 5°N and 1 degree in other areas. The land grid and missing values ​​are assigned to zero. The wind stress and temperature anomalies are then normalized by the spatial mean standard deviation. Finally, the nine layers of data are spliced ​​along the layer axis and input into the model.

3. The ENSO prediction method based on the MSTKAN model of the attention mechanism according to claim 1 is characterized in that: The network structure of the MSTKAN model in S3 is: The first module is Multi-layer Convolutional Encoder, which consists of three layers of convolution; The second module is Multi-Axis Spatiotemporal Attention Structure, which stacks several layers, each of which contains a spatial attention module Intra-SA and a temporal attention module Inter-SA; The third module, KAN Prediction Head, consists of global average pooling and two KAN layers.

4. The ENSO prediction method based on the MSTKAN model of the attention mechanism according to claim 1 is characterized in that: The construction method of the multi-layer convolution encoder in S3 is: using a multi-layer convolution encoder module to further process the input data, including three convolution layers Conv, using a Leaky ReLU activation function, and applying Dropout to prevent overfitting.

5. The ENSO prediction method based on the MSTKAN model of the attention mechanism according to claim 1 is characterized in that: The construction method of the multi-axis spatiotemporal attention module in S3 is as follows: Let X represent a set of input features, with a size of Where T, C, H and W are time, number of channels, height and width respectively.

6. The ENSO prediction method based on the MSTKAN model of the attention mechanism according to claim 5 is characterized in that: The construction method of the spatial attention module Intra-SA is: The Intra-SA block contains two parallel branches: horizontal attention (H-SA) and vertical attention (V-SA); in H-SA, the query, key and value are first generated, represented as and For multi-head attention, the number of heads of multi-head attention is set to M = 8; next, and Reshape into a 2D tensor with dimensions express H horizontal blocks; the attention features of H-SA are Indicates; afterwards, Reshape Back Then concatenate all the attention features to obtain Symmetrically, V-SA calculates vertical attention in the vertical direction to generate attention features Similarly, all V-SA features are reshaped and concatenated as The features of horizontal attention and vertical attention are concatenated to form the output of the spatial attention module.

7. The ENSO prediction method based on the MSTKAN model of the attention mechanism according to claim 5 is characterized in that: The construction method of the temporal attention module Inter-TA is: The Inter-TA block contains two parallel branches: horizontal attention (H-TA) and vertical attention (V-TA); in H-TA, X h First, it is segmented into H non-overlapping and co-localized horizontal regions. Next, according to Generate queries, keys, and values Next, we will Reshape into a 2D tensor with dimensions Then, horizontal attention features are generated in the time dimension Afterwards, Reshape And concatenate the features of all heads along the channel dimension to generate Then, the H attention outputs Fold into represents the weighted features on the horizontal plane; symmetrically, vertical attention is applied to X v To obtain the weighted features of V-TA The features of horizontal attention and vertical attention are concatenated to form the output of the temporal attention module.

8. The ENSO prediction method based on the MSTKAN model of the attention mechanism according to claim 1 is characterized in that: The method of the joint loss function used in S5 is: Loss_var: This loss function is used to calculate the root mean square error (RMSE) between the variables predicted by the model and the true value; and average and sum over multiple dimensions to get the overall variable loss; the formula is as follows: Where: y pred represents the predicted value; y true represents the true value; B represents the batch size; T represents the time step; H and W represent the height and width respectively; Loss_nino: This loss function is used to calculate the root mean square error (RMSE) between the variables predicted by the model and the true value; the formula is as follows: Where: n pred Represents the predicted value; n true represents the true value; N represents the number of samples; T represents the number of time steps; Combine_loss: The combined loss function combines the two loss functions to comprehensively consider the error; the formula is as follows: combine_loss=loss_var+loss_nino.

9. The ENSO prediction method based on the MSTKAN model of the attention mechanism according to claim 1, characterized in that: The evaluation index in S5 is: Correlation Coefficient (Corr); The correlation coefficient is used to measure the linear correlation between the predicted value and the true value; the formula is as follows: Where: y true is the true value; y pred is the predicted value; and They are y true and pred The mean of ; n is the number of samples; Root Mean Squared Error (RMSE) RMSE is used to measure the average error between the predicted value and the true value; the formula is as follows: Where: y true is the true value; y pred is the predicted value; n is the number of samples; Mean Absolute Error (MAE) MAE is used to measure the average absolute error between the predicted value and the true value; the formula is as follows: Where: y true is the true value; y pred is the predicted value; n is the number of samples.

Citation Information

Patent Citations

  • Method for solving object relationship question-answering task in video by utilizing multiple interaction attention mechanism

    CN110727824A

  • Small target detection method and system based on visual attention mechanism

    CN117274661A

  • ENSO event spatio-temporal prediction method and device fusing multi-source data and medium

    CN117688978A

  • Marine weather prediction method and system based on space-time hybrid convolution attention

    CN118296453A

  • ENSO prediction method and model for guiding attention based on content

    CN118521009A