Sea surface temperature prediction method based on depth time index attention model
By adopting a deep time index attention model and a meta-learning framework in sea surface temperature prediction, combining time index and dual attention mechanism, the problems of weak extrapolation ability and long training time in the existing technology are solved, and higher prediction accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510467901.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art has weak extrapolation capabilities and a long training time in sea surface temperature prediction, making it difficult to effectively capture the long-term dependence and nonlinear characteristics of time series.
The sea surface temperature prediction method based on the deep time index attention model is adopted, combined with the time index model and the dual attention mechanism, the model parameters are optimized through the meta-learning framework to improve the generalization ability and prediction accuracy of the model.
By introducing time index features and dual attention mechanisms, the model can more effectively capture the complex patterns and long-term dependencies of the time series, improving the accuracy and robustness of sea surface temperature predictions, and reducing training time.
Smart Images

Figure CN119989005A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of ocean prediction technology, and in particular to a sea surface temperature prediction method based on a deep time index attention model. Background Art
[0002] Marine meteorological forecasting is one of the core tasks of marine engineering. Sea surface temperature (SST), as a physical quantity representing the hot and cold state of the ocean surface, is an important variable affecting the marine ecosystem, typhoon path prediction, and El Niño phenomenon prediction. High-precision prediction of SST can help evaluate the marine ecosystem, accurately predict marine natural disasters, and maintain shipping safety. Therefore, accurate SST prediction is crucial. At present, there are two main methods for sea surface temperature prediction: numerical prediction and data-driven. Numerical prediction is based on solving physical equations based on ocean thermodynamics, which has extremely high requirements for input parameters and a large amount of calculation. The latter mainly includes statistical model methods and artificial intelligence methods. Statistical model methods include ARIMA, support vector machine (SVM), etc. The disadvantage is that it is difficult to capture long-term dependencies and nonlinear characteristics. Artificial intelligence methods mainly include Transformer, Informer, etc., but the training time is long and lacks sufficiently high extrapolation performance. Summary of the invention
[0003] Purpose of the invention: The purpose of the present invention is to provide a sea surface temperature prediction method based on a deep time index attention model, which combines the time index model with the dual attention mechanism, introduces a meta-learning framework to deal with the long-term dependency problem of time series, obtains the optimal parameters of the model, and enables the model to have better generalization ability, thus solving the problems of weak extrapolation ability and long training time of current deep neural networks when predicting time series.
[0004] Technical solution: The present invention provides a method for predicting sea surface temperature based on a deep time index attention model, comprising the following steps: (1) Using the interpolated sea surface temperature (OISST) dataset provided by the open source website, we selected the sea surface temperature (SST) data of five observation points in the South China Sea; (2) Preprocess the collected one-dimensional sea temperature data and the corresponding time index data and divide the data set into training set, validation set and test set; (3) Construct and train a SST prediction model based on deep time index attention, which includes: Fourier feature transform layer, implicit neural representation INR layer and dual attention layer; (4) The test set is input into the trained model for prediction, and the SST prediction results are denormalized and finally compared with the observed values.
[0005] Furthermore, in step (1), the preprocessing includes normalizing the time index data and standardizing the predicted SST data corresponding to the time index.
[0006] Furthermore, step (2) includes the following steps: (21) The time index data of the training set is normalized, with the one-dimensional time index data as input and the corresponding predicted SST data as output; the formula for normalizing the time index is as follows: ; in, Represents the normalized index of the tth day, with a dimension of 1; The total number of days for time index data; (22) Standardize the SST data of the training set: ; in, Represents the standardized SST data, with a dimension of 1; Indicates t The SST data of the day, represents the standard deviation of SST data, Represents the mean of SST data.
[0007] Furthermore, step (3) is as follows: the Fourier feature transform layer transforms the normalized time index into high-frequency and low-frequency features through multi-scale Fourier transform; the INR layer performs nonlinear transformation on the Fourier features through the multi-layer perceptron MLP to extract temporal dependencies; the dual attention layer combines the sliding exponential average EMA and the dynamic projection mechanism to enhance the focus on key features.
[0008] Furthermore, the Fourier feature transform layer formula is as follows: ; in, represents the normalized time index; S represents the number of preset scale parameters, which is used to control the coverage of different frequency ranges; Indicates S The frequency scale is converted to T The one-dimensional time feature vector of is expanded into the dimension Matrix of T is the size of the lookback window; is the temporal feature of the Fourier encoding output.
[0009] Furthermore, the INR layer formula is as follows: ; in, is the input of the INR layer, Temporal features for Fourier encoding; The function represents the output of the Fourier feature transform layer, T is the size of the lookback window, S is the number of scale parameters in the Fourier feature layer; For the k The hidden layer feature representation of the layer, Indicates k The weight matrix of the layer, d is the number of hidden layer neurons of INR, Indicates k The bias of the layer, The activation function represents the nonlinear modeling capability. It is k The output of the layer; , It is the last layer in INR ( K layer), It is K The input of the layer, It represents the hidden layer features after the INR layer is transformed by MLP. is the time index of the lookback window.
[0010] Furthermore, the dual attention layer is implemented as follows: the hidden features are mapped into query, key and value matrices; the formula is as follows: ; in, is the hidden layer feature output by the INR layer, T Indicates the size of the lookback window, d is the number of neurons in the hidden layer in INR, which extracts local nonlinear features through MLP; , , Respectively represent h The query, key, and value weight matrices of each header; H is the total number of attention heads; , , Respectively represent h The query vector, key vector, and value vector of the attention head; Exponential moving average is introduced and applied to query matrix and key matrix. The formula is as follows: ; in, is the sliding average coefficient, is the parameter of weight reduction, Represents the sliding average result of the previous time step; is the query and key matrix after sliding average; Next, the sliding average result is dynamically projected, and the formula is as follows: ; in, is the bond matrix after sliding exponential averaging, is the value matrix, is the weight matrix of dynamic projection, is the bias vector of dynamic projection, The function normalizes along the feature dimension and generates dynamic projection weights. Einstein summation convention function, representing matrix multiplication, addition and transposition operations; are the keys and values after dynamic projection; Finally, the attention weights are calculated: ; in, is the dot product of the query and the key, is the key after dynamic projection, is the dot product result of the query and the key, indicating the similarity between time steps; is the scaling factor; For the h The attention weight matrix of each head; Calculate the h The output features of the attention head are as follows: ; in, is the value after dynamic projection, through The attention weights are weighted summed up to get h Subspace context features of attention heads; For the h The output features of the attention heads; h After calculating the output features of each head one by one, multi-head splicing and linear transformation are performed: ; in, For the h The output features of each head, is the splicing function; is a linear transformation matrix used to fuse multi-subspace information. The final output context feature is the global context information enhanced by the attention mechanism; finally, the concatenated feature is obtained.
[0011] Furthermore, in step (3), the model training process is as follows: a meta-learning framework is adopted to divide the parameters into basis parameters and meta-parameters; wherein, the inner loop optimizes the basis parameters through ridge regression based on the backtracking window; and the outer loop globally optimizes the meta-parameters through back-propagation.
[0012] An electronic device described in the present invention includes a memory, a processor, and a computer program stored in the memory, and the processor implements the steps of any one of the methods when executing the program.
[0013] A computer-readable storage medium according to the present invention stores a computer program, and when the program is executed by a processor, the steps of any one of the methods are implemented.
[0014] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages: by introducing time index features, a sea surface temperature prediction model based on deep time index attention is constructed. The Fourier feature layer captures the high-frequency and low-frequency features of the time series through multiple scale parameters, and enhances the ability to express complex time patterns; the implicit neural representation (INR) layer gradually extracts high-order nonlinear representations of time features through multi-layer perceptrons and ReLU activation functions, and enhances the modeling ability of the model for complex time dependencies; the dual attention layer captures multiple dependencies of the time series in parallel through a multi-head attention mechanism, introduces sliding exponential averaging and dynamic projection, enhances the global modeling ability of contextual information, and focuses on important features, while smoothing noise and improving the robustness of prediction; introduces a meta-learning framework to train model parameters, solves the basis parameters through the inner loop ridge regression, improves the generalization ability of the model for unseen data, and globally optimizes the meta-parameters through the outer loop back propagation to ensure consistent performance in different time windows, avoid overfitting, and improve the prediction accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is a flow chart of the present invention; Figure 2 Structural diagram of introducing a meta-learning framework for the deep time-indexed attention model of our invention. DETAILED DESCRIPTION
[0016] The technical solution of the present invention is further described below in conjunction with the accompanying drawings.
[0017] like Figure 1 As shown, an embodiment of the present invention provides a sea surface temperature prediction method based on a deep time index attention model, comprising the following steps: (1) The optimal interpolation sea surface temperature (OISST) dataset provided by the National Oceanic and Atmospheric Administration (NOAA) of the United States was used to select sea surface temperature (SST) data from five observation points in the South China Sea with a temporal resolution of 1 day and a spatial resolution of 0.25°×0.25°. The time span is from January 1, 2004 to December 31, 2016, a total of 4749 days. The longitude of the five observation points selected is 113.125, and the latitudes are 11.625, 13.625, 15.625, 17.625, and 19.625, respectively.
[0018] (2) Preprocess the collected one-dimensional SST data and the corresponding time index data, including normalizing the time index data and normalizing the predicted SST data corresponding to the time index. The data set is divided into training set, validation set and test set in a ratio of 7:1:2. The following steps are included: (21) Normalize the time index data of the training set. The one-dimensional time index data is the input, and the corresponding predicted SST data is the output. Normalize the time index: ; in, Represents the normalized index of the tth day, with a dimension of 1; The total number of days for time index data; (22) Standardize the SST data of the training set: ; in, Represents the standardized SST data, with a dimension of 1; Indicates t The SST data of the day, represents the standard deviation of SST data, Represents the mean of SST data.
[0019] (3) Construct an SST prediction model based on deep time index attention. The model consists of a Fourier feature generation layer, an implicit neural representation (INR) layer, and a dual attention layer. The details are as follows: The normalized time index training data is first input into a Fourier feature layer to generate multi-scale Fourier features; then, the Fourier features are nonlinearly transformed by applying 5 sets of multi-layer perceptrons (MLP) through the implicit neural representation layer to generate hidden features; finally, the important features in the hidden features are enhanced through the dual attention layer, large weight values are assigned, and a context feature matrix is generated. The dual attention layer includes a sliding exponential average (EMA) and a dynamic projection mechanism. The time index model is defined as a model that makes predictions only through the time index features of future time steps: , is the SST data, Represents the time index corresponding to each SST data value. This model can be summarized as a function fitting problem, where the input of the function is the time index of the time series and the output is the SST forecast value. The network is parameterized as a neural network and the weights of the neural network are updated in a completely data-driven manner. The network input is the time coordinate, Divide the parameter space into basis parameters and meta parameters , the basis parameters are defined as the bias and weight of the last layer of MLP, and the meta parameters are defined as the parameters of other layers of MLP and the dual attention layer. The meta-learning framework is adopted, the input data is the context feature matrix, and the output data is the SST value of the training set. The training process includes two stages: the inner loop and the outer loop. Among them, the inner loop calculates the weight of the last layer of MLP through the ridge regression method in the backtracking window and bias , the outer loop optimizes the global parameters by back propagating the prediction error. If the termination condition (maximum number of training rounds or triggering early stopping) is not reached, the inner and outer loops continue, otherwise the training ends. It includes the following steps: (31) Perform Fourier feature transform on the normalized time index training set data obtained by preprocessing, and the formula is as follows: ; in, represents the normalized time index; S represents the number of preset scale parameters, which is used to control the coverage of different frequency ranges; Indicates S The frequency scale is converted to T The one-dimensional time feature vector of is expanded into the dimension Matrix of T is the size of the lookback window; is the temporal feature of the Fourier encoding output. In this experiment, the scale parameter is set to [0.01, 0.1, 1, 5, 10, 20, 50, 100]. S= 8 scales. Represents a random matrix of the Sth scale, the elements in the matrix are Medium sampling, It is the default S The deviation parameter of each scale determines the projection matrix For each scale S, and the projection matrix Performs matrix multiplication and then applies sine and cosine functions element-wise. represents the output feature vector, where , is the number of sine and cosine features generated at each scale; T Indicates the sequence length of the model input. This test task is to predict the single-point SST value of 5 days, and the model's backtracking window size is set to 20, so T is 20.
[0020] (32) The high-dimensional Fourier features obtained by (31) Input to the INR module. The INR module consists of K The specific definition of the INR module is as follows: ; in, is the input of the INR layer, Temporal features for Fourier encoding; The function represents the output of the Fourier feature transform layer, T is the size of the lookback window, S is the number of scale parameters in the Fourier feature layer; For the k The hidden layer feature representation of the layer, Indicates k The weight matrix of the layer, d is the number of hidden layer neurons of INR, Indicates k The bias of the layer, The activation function represents the nonlinear modeling capability. It is k The output of the layer; , It is the last layer in INR ( K layer), It is K The input of the layer, It represents the hidden layer features after the INR layer is transformed by MLP. is the time index of the lookback window. The first layer of MLP transforms the Fourier features from The dimension is mapped to d dimension, and the subsequent layers gradually extract high-order nonlinear features through the ReLU activation function. The data dimension of MLP output is , T Indicates the sequence length of the model input. It represents the hidden state of the kth layer in the INR layer, and gradually extracts the nonlinear representation of the time features through the multi-layer ReLU network, compresses the high-dimensional Fourier features to d dimensions, and retains the temporal dependency. is the final output function of the INR module, is the time index of the lookback window, are the weight matrix and bias of the last layer of MLP, is the input value of the last layer.
[0021] (33) The hidden features obtained by (32) Input to the dual attention layer, mapped to H subspaces ( H =8), generate query, key and value matrices. The formula is as follows: ; in, It is the hidden layer feature output by the INR layer, that is, the local nonlinear feature extracted by MLP; , , Respectively represent h The query, key, and value weight matrices of the heads; the input features are mapped to the subspace with dimensions , the dimension of the weight matrix is . Multiply the mapped weight matrix and the hidden layer features to get the query matrix , key matrix , value matrix , respectively represent the query feature, key feature, and value feature of the subspace.
[0022] Exponential moving average is introduced and applied to query matrix and key matrix. The formula is as follows: ; in, is the sliding average coefficient, is the parameter of weight reduction, Represents the sliding average result of the previous time step; is the query and key matrix after sliding average; Next, the sliding average result is dynamically projected, and the formula is as follows: ; in, is the bond matrix after sliding exponential averaging, is the value matrix, is the weight matrix of dynamic projection, is the bias vector of dynamic projection, The function normalizes along the feature dimension and generates dynamic projection weights. Einstein summation convention function, representing matrix multiplication, addition and transposition operations; are the keys and values after dynamic projection; Finally, the attention weights are calculated: ; in, is the dot product of the query and the key, is the key after dynamic projection, is the dot product result of the query and the key, indicating the similarity between time steps; is the scaling factor; For the h The attention weight matrix of each head; Calculate the h The output features of the attention head are as follows: ; in, is the value after dynamic projection, through The attention weights are weighted summed up to get h Subspace context features of attention heads; For the h The output features of the attention heads; h After calculating the output features of each head one by one, multi-head splicing and linear transformation are performed: ; in, For the h The output features of each head, is the splicing function; is a linear transformation matrix used to fuse multi-subspace information. The final output context feature is the global context information enhanced by the attention mechanism; finally, the concatenated feature is obtained.
[0023] (34) The context features obtained in step (33) As the input of the meta-learning framework. In the meta-learning framework, the parameter space of the model is divided into meta-parameters and basis parameters. Including weights of multi-layer MLP , the weight of the dual attention layer . Base parameters is the weight of the last layer of MLP and For the training set data, the lookback window and the prediction window are divided. The length of the lookback window is defined as , the length of the prediction window is , that is, predicting the sea surface temperature for the next L days by using the time index data of the past T days. Let the current time step be t, and the interval of the backtracking window is from arrive , as the training set training parameters, the prediction window interval is from t to , used to verify the predictive performance of the model, and the training model parameters are divided into two stages: inner loop and outer loop.
[0024] In the inner loop stage, for each time window t, the optimal basis parameters are solved based on the prediction loss function and : ; in, Represents meta parameters, represents the basis parameters, is a time index, indicating the j time steps, Indicates the first j The SST observations at time steps. L represents the selected loss function. The calculation formula is: ; in, Represents context features, is the weight matrix, and b is the bias. Expand the loss function and apply the sum function to calculate the sum of the losses of all time steps in the lookback window. In order to prevent overfitting, the L2 regularization term is introduced to obtain the formula of the loss function: ; in It means calculating the mean square error between the predicted value and the true value. is the regularization coefficient, is the L2 norm of W, which is used to prevent overfitting so that W will not take too large a value, thereby improving the generalization ability of the model. If the context features and true values of the backtracking window are expressed in matrix form, let , then the formula for updating the basis parameters can be expressed as: ; The above parameter update problem can be transformed into a ridge regression problem, and its closed-form solution is: ; in, is the context feature matrix containing each time point in the lookback window, contains the true SST value at each time point, is the identity matrix. is the function for calculating the arithmetic mean. The optimal basis parameters are calculated and .
[0025] In the outer loop stage, the loss parameter is calculated by the derivative chain rule The gradient of , updates the meta-parameters to optimize the global generalization ability. The formula for updating the meta-parameters is: ; in, It means that the basic parameters obtained by the input inner loop are used to obtain the SST prediction value. The context features generated for the meta-parameters. The gradient of the global loss to the original parameters is calculated by the chain rule: ; because yes The function of gradient calculation needs to be considered right Dependencies. Implicit calculation through the automatic differentiation framework in pytorch. Apply gradient descent to update the parameters: ; in, is the learning rate, which controls the step size of parameter update. The termination condition is: when the outer loop reaches the number of training rounds or the change in global loss is less than the given threshold ( 1e-5 ), the training is stopped and the final parameter space is generated to obtain the trained prediction model.
[0026] (4) Input the test set into the trained model for prediction, denormalize the SST prediction results, and finally compare them with the observed values to calculate the prediction error. The details are as follows: (41) Time index of the test set Input into the trained model and get the predicted value after the model calculation , and the predicted value is de-standardized, the formula is as follows: ; in, The mean and standard deviation of the SST data during the training phase. is the predicted value after denormalization.
[0027] (42) In order to comprehensively evaluate the prediction performance of the model, the root mean square error (MSE) and mean absolute error (MAE) are used as evaluation indicators. These two indicators can reflect the prediction accuracy and stability of the model from different angles. The specific formula is: ; Where N is the total number of predicted samples, is the predicted value after denormalization, is the observed value.
[0028] The present invention selects five observation points, whose longitudes and latitudes are P1 (113.125°N, 11.625°N), P2 (113.125°N, 13.625°N), P3 (113.125°N, 15.625°N), P4 (113.125°N, 17.625°N), and P5 (113.125°N, 19.625°N). The error results of the four deep learning models and the present invention are compared, and the root mean square error (RMSE) and mean absolute error (MAE) of each model at the five points P1, P2, P3, P4, and P5 of the South China Sea test set are calculated, as shown in Table 1; and the average running time of each model per round at each observation point is recorded (in seconds), as shown in Table 2. The results show that the predicted values obtained by the model of the present invention at all observation points are closer to the observed values than those of other models, and the running time of the present invention is the shortest, achieving good time performance.
[0029] Table 1 Evaluation indicators of different models at five observation points (unit: °C) ; Table 2 Average running time per round of different models at five observation points (unit: s) .
Claims
1. A sea surface temperature prediction method based on a deep time index attention model, characterized in that: The following steps are involved: (1) Using the interpolated sea surface temperature (OISST) dataset provided by the open source website, we selected the sea surface temperature (SST) data of five observation points in the South China Sea; (2) Preprocess the collected one-dimensional sea temperature data and the corresponding time index data and divide the data set into training set, validation set and test set; (3) Construct and train a SST prediction model based on deep time index attention, which includes: Fourier feature transform layer, implicit neural representation INR layer and dual attention layer; (4) The test set is input into the trained model for prediction, and the SST prediction results are denormalized and finally compared with the observed values.
2. The method for predicting sea surface temperature based on a deep time index attention model according to claim 1, characterized in that: In step (1), the preprocessing includes normalizing the time index data and standardizing the predicted SST data corresponding to the time index.
3. The method for predicting sea surface temperature based on a deep time index attention model according to claim 1, characterized in that: Step (2) includes the following steps: (21) The time index data of the training set is normalized, with the one-dimensional time index data as input and the corresponding predicted SST data as output; the formula for normalizing the time index is as follows: ; in, Represents the normalized index of the tth day, with a dimension of 1; The total number of days for time index data; (22) Standardize the SST data of the training set: ; in, Represents the standardized SST data, with a dimension of 1; Indicates t The SST data of the day, represents the standard deviation of SST data, Represents the mean of SST data.
4. The method for predicting sea surface temperature based on a deep time index attention model according to claim 1, characterized in that: Step (3) is as follows: the Fourier feature transform layer transforms the normalized time index through multi-scale Fourier transform to generate high-frequency and low-frequency features; the INR layer performs nonlinear transformation on the Fourier features through the multi-layer perceptron MLP to extract temporal dependencies; the dual attention layer combines the sliding exponential average EMA and the dynamic projection mechanism to enhance the focus on key features.
5. The method for predicting sea surface temperature based on a deep time index attention model according to claim 4, characterized in that: The Fourier feature transform layer formula is as follows: ; in, represents the normalized time index; S represents the number of preset scale parameters, which is used to control the coverage of different frequency ranges; Indicates S The frequency scale is converted to T The one-dimensional time feature vector of is expanded into the dimension Matrix of T is the size of the lookback window; is the temporal feature of the Fourier encoding output.
6. A method for predicting sea surface temperature based on a deep time index attention model according to claim 4, characterized in that: The INR tier formula is as follows: ; in, is the input of the INR layer, Temporal features for Fourier encoding; The function represents the output of the Fourier feature transform layer, T is the size of the lookback window, S is the number of scale parameters in the Fourier feature layer; For the k The hidden layer feature representation of the layer, Indicates k The weight matrix of the layer, d is the number of hidden layer neurons of INR, Indicates k The bias of the layer, The activation function represents the nonlinear modeling capability. It is k The output of the layer; , is the weight matrix and bias of the last layer in INR, It is K The input of the layer, It represents the hidden layer features after the INR layer is transformed by MLP. is the time index of the lookback window.
7. The method for predicting sea surface temperature based on a deep time index attention model according to claim 4, characterized in that: The dual attention layer is implemented as follows: map the hidden features into query, key, and value matrices; the formula is as follows: ; in, is the hidden layer feature output by the INR layer, T Indicates the size of the lookback window, d is the number of neurons in the hidden layer of INR, which extracts local nonlinear features through MLP; , , Respectively represent h The query, key, and value weight matrices of each header; H is the total number of attention heads; , , Respectively represent h The query vector, key vector, and value vector of the attention head; Exponential moving average is introduced and applied to query matrix and key matrix. The formula is as follows: ; in, is the sliding average coefficient, is the parameter of weight reduction, Represents the sliding average result of the previous time step; is the query and key matrix after sliding average; Next, the sliding average result is dynamically projected, and the formula is as follows: ; in, is the bond matrix after sliding exponential averaging, is the value matrix, is the weight matrix of dynamic projection, is the bias vector of dynamic projection, The function normalizes along the feature dimension and generates dynamic projection weights. Einstein summation convention function, representing matrix multiplication, addition and transposition operations; are the keys and values after dynamic projection; Finally, the attention weights are calculated: ; in, is the dot product of the query and the key, is the key after dynamic projection, is the dot product result of the query and the key, indicating the similarity between time steps; is the scaling factor; For the h The attention weight matrix of each head; Calculate the h The output features of the attention head are as follows: ; in, is the value after dynamic projection, through The attention weights are weighted summed up to get h Subspace context features of attention heads; For the h The output features of the attention heads; h After calculating the output features of each head one by one, multi-head splicing and linear transformation are performed: ; in, For the h The output features of each head, is the splicing function; is a linear transformation matrix used to fuse multi-subspace information. The final output context feature is the global context information enhanced by the attention mechanism; finally, the concatenated feature is obtained.
8. The method for predicting sea surface temperature based on a deep time index attention model according to claim 4, characterized in that: In step (3), the model training process is as follows: a meta-learning framework is used to divide the parameters into base parameters and meta-parameters; the inner loop optimizes the base parameters through ridge regression based on the backtracking window; and the outer loop globally optimizes the meta-parameters through back propagation.
9. An electronic device comprising a memory, a processor and a computer program stored in the memory, characterized in that: When the processor executes the program, the steps of the method according to any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Power transformer oil temperature prediction method and system, storage medium and equipment
CN118260718A
Medium and long term sea wave significant wave height prediction method based on EmaDform
CN118364230A
Depth time index load prediction method based on meta-learning
CN118966436A
Sea surface temperature prediction method based on adaptive space-time diagram neural network
CN119128534A
Cascaded local implicit transformer for arbitrary-scale super-resolution
US20240070809A1
Cited By
Seawater state multi-factor forecasting method based on mixed attention mechanism
CN120850015A