Sea ice concentration prediction method based on multilayer stacked ConvLSTM
Through the multi-layer stacked ConvLSTM neural network model, combined with the encoder and decoder structure, the problems of large amount of calculation and poor forecasting timeliness in marine environmental forecasts are solved, and high-precision prediction of sea ice density is achieved, especially in medium and short-term predictions, the prediction efficiency is significantly improved.
Patent Information
- Application Number
- CN202510596132.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-08-15
AI Technical Summary
The prior art has problems such as large amount of calculation and poor forecasting timeliness in marine environment forecasting, especially in the prediction of sea ice density, which is difficult to effectively capture the errors of nonlinear relationships, resulting in insufficient forecasting accuracy.
The multi-layer stacked ConvLSTM neural network model is adopted, combined with the encoder and decoder structure, and the spatial characteristics of ocean data are captured through convolutional operations and maintained long-term spatial dependence, so as to predict sea ice density.
It significantly improves the accuracy and accuracy of sea ice density prediction, especially in medium and short-term prediction, which can effectively capture the spatial characteristic relationships and temporal evolution patterns in the ocean area, improving prediction efficiency.
Smart Images

Figure CN120493737A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of polar ocean environmental factor prediction, and specifically relates to a sea ice density prediction method based on multi-layer stacked ConvLSTM. Background Art
[0002] Ocean forecasts are mainly divided into numerical forecasts and statistical forecasts. Numerical forecasts can accurately predict meteorological phenomena based on current meteorological conditions using mathematical models of the ocean. However, numerical forecasts also have limitations such as large computational complexity, poor performance in periodic forecasts, and poor forecast timeliness. On the contrary, statistical models perform well in ocean forecasts, regardless of the time interval, as long as the data has sufficient correlation. With the rapid development of artificial intelligence technology, it has demonstrated significant superiority in the field of data feature extraction, and neural network methods based on statistical analysis are constantly developing in SIC forecasting. Driven by massive ocean spatiotemporal series data, deep learning methods can quickly and efficiently analyze and process the spatiotemporal characteristics of ocean data. It can not only deeply explore the dynamic characteristics of ocean elements, but also construct an accurate ocean element prediction and forecast data model. To this end, the present invention proposes a sea ice density prediction method based on a multi-layer stacked ConvLSTM. Summary of the Invention
[0003] The purpose of the present invention is to provide a sea ice density prediction method based on a multi-layer stacked ConvLSTM. First, the MLS-ConvLSTM network is applied to the short- and medium-term forecast of Arctic sea ice. This method can well restore the spatial information characteristics of sea ice, effectively reduce the error caused by nonlinear relationships, and improve the accuracy of statistical forecasts.
[0004] The technical solutions adopted by the present invention are as follows:
[0005] A sea ice density prediction method based on multi-layer stacked ConvLSTM includes the following steps:
[0006] Step 1: Obtain polar ocean reanalysis data and perform preprocessing to eliminate land points;
[0007] Step 2: Establish an MLS-ConvLSTM neural network model and perform simulation;
[0008] Step 3: Determine the prediction effect of the MLS-ConvLSTM neural network model, optimize the parameters of the MLS-ConvLSTM neural network model, and output the trained MLS-ConvLSTM neural network model;
[0009] Step 4: Use the MLS-ConvLSTM neural network model to analyze the ocean data to complete the prediction of polar sea ice density.
[0010] Preferably, in step 1, the data preprocessing and analysis method is as follows:
[0011] Step 101: During the data reading process, the start and end time of the experimental data and the longitude and latitude information of the selected sea area are set to define the time and space range of the data required for the experiment;
[0012] Step 102: Check the current year. If it is a leap year, set February of that year to 29 days; otherwise, set February of that year to 28 days.
[0013] Step 103: Reading the geographical location, sea ice density and other information of the selected target sea area on a certain date from the database, and storing the information in a new storage file created for the target sea area;
[0014] Step 104: Process all the files in the database that participate in the experiment in order, and repeat steps 102 and 103 until the data reading reaches the deadline;
[0015] Step 105: Reduce the dimension of the read data by one, and remove the dimensional information representing the sea ice density.
[0016] Preferably, in step 2, the ConvLSTM neural network model is an optimized multi-layer stacked ConvLSTM network architecture, which integrates two major components: an encoder and a decoder;
[0017] The encoder is a composite structure integrating multiple convolutional layers and ConvLSTM units. The convolution operation captures key information in spatial feature maps of different scales. At the same time, the ConvLSTM unit maintains the long-term spatiotemporal dependencies of the input tensor.
[0018] The decoder, as a mirror-symmetric structure of the encoder, is also constructed by multiple layers of ConvLSTM and deconvolution operations. The working mechanism of the ConvLSTM unit is consistent with that of the encoder. The deconvolution operation is the opposite of the convolution operation. It is responsible for converting the encoder output tensor into the target output tensor, completing the reverse reconstruction of the information.
[0019] Preferably, the step 3 comprises the following steps:
[0020] Step 301: Compare the prediction results of the MLS-ConvLSTM network with those of traditional methods visually to determine the prediction effect.
[0021] Step 302: Selecting the root mean square error (RMSE) and the Pearson coefficient (PCC) as evaluation indicators to determine the prediction effect of the MLS-ConvLSTM network method;
[0022] Step 303: Adjust the parameters of the MLS-ConvLSTM neural network model according to the accuracy of the forecast effect, repeat steps 2 and 3, and output the trained MLS-ConvLSTM neural network model.
[0023] Preferably, in step 301, an intuitive visual comparison is made between the forecast results output by the MLS-ConvLSTM neural network model and the forecast results obtained by linear regression and empirical orthogonal function decomposition; this includes drawing spatial distribution maps of key indicators such as sea ice coverage and thickness, and observing the changing trends and differences of different forecasting methods in time series, so as to preliminarily judge the advantages of the LS-ConvLSTM neural network model in capturing the dynamic changes of sea ice.
[0024] Preferably, in step 302, the network loss is calculated by the root mean square error method (RMSE), which calculates the deviation between the predicted value and the label value, and uses Y i represents the measured value, f(x i ) represents the network prediction value, and its specific formula is:
[0025]
[0026] Preferably, in step 302, X is used t represents the obtained data matrix, γ0 is the product of the matrix standard deviation, and the specific formula for calculating the autocorrelation coefficient is as follows:
[0027]
[0028] Preferably, in step 303, during the network training process, the network parameters are continuously adjusted by means of back-propagation algorithm optimization to minimize the value of the loss function, and the correlation coefficient is used as a statistical correlation measure between the predicted value and the true label value; when the correlation coefficient is high, the predicted output of the model is highly consistent with the true label, that is, the prediction accuracy of the model is high; conversely, the prediction accuracy of the model is low, and the model parameters are adjusted at this time, and the prediction and verification are repeated until the prediction accuracy of the model meets the requirements, and finally the trained MLS-ConvLSTM neural network model is output.
[0029] The technical effects achieved by the present invention are:
[0030] The purpose of this invention is to address the urgent need for marine environmental protection for various marine operating platforms, including offshore platforms, underwater unmanned and manned vehicles, and ships. To achieve this goal, the present invention focuses on exploring complex and variable ocean spatiotemporal series data derived from multiple channels, with highly nonlinear characteristics and high coupling. By deeply analyzing the inherent characteristics of this data, an advanced prediction method, the MLS-ConvLSTM method, is proposed. This method is based on the improvement and optimization of the traditional convolutional long short-term memory neural network. By introducing a multi-layer stacking mechanism, this method structure not only deepens the model's learning ability but also significantly enhances its ability to capture and analyze complex spatiotemporal dynamic characteristics. At the same time, the network fully utilizes the convolution operation in the ConvLSTM, enabling the network to effectively process and analyze the spatial information of sea ice concentration, thereby accurately capturing the spatial characteristic relationships of SIC within the ocean area and its temporal evolution pattern. Compared with traditional methods, this improved design demonstrates significant advantages in integrating time series information with spatial distribution characteristics, providing a solid theoretical foundation and technical support for the accurate prediction of sea ice concentration. Using the MLS-ConvLSTM method of the present invention, the ConvLSTM unit is able to maintain the long-term and spatial dependencies of the input tensor. Furthermore, as the number of stacked layers increases, the network captures richer information across all spatial dimensions. Therefore, the performance of medium- and short-term SIC predictions is significantly improved by using a multi-layer stacked ConvLSTM network. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 This is a flow chart of a sea ice density prediction method based on multi-layer stacked ConvLSTM of the present invention;
[0032] Figure 2 It is the Greenland Sea area representing the data domain of the present invention;
[0033] Figure 3 It is the network architecture of the multi-layer stacked convolutional long short-term memory network of the present invention;
[0034] Figure 4 This is the internal structure diagram of the ConvLSTM network of the present invention;
[0035] Figure 5 It is a comparison of the prediction results of the three methods of the present invention;
[0036] Figure 6 It is a comparison of the prediction results of different forecast durations of the present invention;
[0037] Figure 7 This is the flow chart of data reading and grid transformation using Python language of the present invention;
[0038] Figure 8It is a flow chart of the MLS-ConvLSTM prediction method of the present invention;
[0039] Figure 9 This is the ConvLSTM structure diagram of the present invention;
[0040] Figure 10 It is the deconvolution principle diagram of the present invention;
[0041] Figure 11 Schematic diagram of multi-channel convolution of the present invention. DETAILED DESCRIPTION
[0042] In order to make the purpose and advantages of the present invention more clearly understood, the present invention is described in detail below with reference to the following examples. It should be understood that the following text is only used to describe one or more specific embodiments of the present invention and does not strictly limit the scope of protection of the present invention.
[0043] like Figures 1-11 As shown, a sea ice density prediction method based on multi-layer stacked ConvLSTM includes the following steps: Step 1: obtain polar sea area reanalysis data and perform preprocessing to eliminate land points;
[0044] In actual use, the present invention selected a series of data sets from the global ocean physical reanalysis product suite provided by the Copernicus Marine Environment Monitoring Service (CMEMS) as the verification sea area, as input and target data, covering key indicators such as sea surface temperature (SST), sea surface salinity (SSS) and sea ice concentration (SIC). These data mainly rely on the current real-time global prediction capabilities of the CMEMS system. The core of the model architecture lies in the NEMO platform, which is driven by the recent re-analysis results of the ERA-Interim dataset (later converted to the ERA5 dataset) of the European Centre for Medium-Range Weather Forecasts (ECMWF) at the surface level. Our research focuses on the waters near the Greenland Sea, with a specific coordinate range of 20°W to 10°W and 65°N to 75°N. The data adopted spans ten years from 2010 to 2019. Figure 2 A map of the Greenland Sea representing the study area is shown;
[0045] For data processing, Greenland Sea SIC data from 2010 to 2019 were selected. The entire dataset spans a ten-year period. 80% of the samples were assigned as training sample sets, and the remaining 20% as validation sample sets. Specifically, the data from 2010 to 2017 constituted the training sample set for neural network parameter training, while the data from the remaining years served as validation sample sets for evaluating network performance. It is worth noting that both training and validation samples consist of two parts: an input sequence and a target sequence. If the input sequence ranges from A+1 to A+B, the corresponding target sequence range is A+B+1 to A+B+L, where the length of the input sequence is B and the length of the target sequence is L. During the training phase, the training domain was meticulously divided into a 20×20 grid system, with each grid cell covering a 0.5°×0.5° geographic area. To improve training efficiency and model performance, all input data was normalized before training. This step aims to eliminate the dimensional differences between different features, thereby ensuring that the model can more effectively learn and identify potential patterns in the data.
[0046] Obtain reanalysis data and perform preprocessing such as eliminating land points. The flow chart is as follows: Figure 7 shown.
[0047] Preferably, the data is preprocessed, such as filling missing values, smoothing outliers, and performing spatial and temporal interpolation, to provide a clean, continuous, and consistent dataset for subsequent analysis. The data preprocessing and analysis methods are as follows:
[0048] Step 101: During the data reading process, the start and end time of the experimental data and the longitude and latitude information of the selected sea area are set to define the time and space range of the data required for the experiment;
[0049] Step 102: Check the current year. If it is a leap year, set February of that year to 29 days; otherwise, set February of that year to 28 days.
[0050] Step 103: Reading the geographical location, sea ice density and other information of the selected target sea area on a certain date from the database, and storing the information in a new storage file created for the target sea area;
[0051] Step 104: Process all the files in the database that participate in the experiment in order, and repeat steps 102 and 103 until the data reading reaches the deadline;
[0052] Step 105: Reduce the dimension of the read data by one, and remove the dimensional information representing the sea ice density.
[0053] Step 2: Establish an MLS-ConvLSTM neural network model and perform simulation;
[0054] In step 2, the ConvLSTM neural network model is an optimized multi-layer stacked ConvLSTM network architecture that integrates two major components: the encoder and the decoder.
[0055] The encoder, a composite structure integrating multiple convolutional layers and ConvLSTM units, demonstrates excellent spatiotemporal information extraction capabilities. The convolution operation efficiently captures key information from spatial feature maps at different scales. Furthermore, the ConvLSTM unit not only effectively maintains the long-term spatiotemporal dependencies of the input tensor, but also successfully avoids the problem of vanishing gradients during training.
[0056] The decoder, as a mirror-symmetrical structure of the encoder, is also constructed by multi-layer ConvLSTM and deconvolution operations. The mechanism of the ConvLSTM unit is consistent with that of the encoder. The deconvolution operation is the opposite of the convolution operation. It is responsible for converting the encoder output tensor into the target output tensor, completing the reverse reconstruction of the information. The overall architecture of the multi-layer stacked ConvLSTM network is intuitively demonstrated as follows: Figure 3 as well as Figure 9 As shown, the monomer unit structure of the multi-layer superimposed ConvLSTM network is as follows Figure 4 As shown, the deconvolution principle diagram is as follows Figure 10 As shown, multi-channel convolution is as follows Figure 11 shown.
[0057] The ConvLSTM network structure mainly includes input gate, forget gate, output gate and memory unit. The forget gate determines the cell state C of the previous time step. t-1 Which fragments should be retained to the cell state C of the current time step t The equation of the forget gate is as follows:
[0058] F t =σ(W x,f *X t +W h,f *H t-1 +b f ) (1)
[0059] The input gate controls the current time step input X t Keep to C t The weight of the input gate is as follows:
[0060] I t =σ(W x,i *X t +W h,i *H t-1 +b i ) (2)
[0061] The cell state at this time The equation is:
[0062]
[0063] The output gate selectively passes C t Output H to ConvLSTM t The equation for the output gate is as follows:
[0064] O t =σ(W x,o *X t +W h,o *H t-1 +b o ) (4)
[0065] Finally, the output of the ConvLSTM unit at time t is H t :
[0066]
[0067] The above are the main formulas involved in the calculation process of ConvLSTM network, where * is the convolution operation, W is the convolution parameter, b is the bias coefficient, σ is the sigmoid activation function, X t is the input at time t, C t is the cell state at time t, H t is the hidden state at time t.
[0068] The convolutional long short-term memory neural network (Convolutional Long Short Term Memory, ConvLSTM) used in this paper integrates the convolution operation into the recurrent network calculation process of the long short-term memory neural network (Long Short Term Memory, LSTM). This network not only retains the dynamic gate mechanism in the LSTM network to solve the long-term sequence dependency problem, but also replaces the vector dot product operation with the matrix Hadamard product operation during the recurrent calculation process, so that the network can simultaneously extract the spatiotemporal features of the sequence. ConvLSTM internally contains an input gate, a forget gate, an output gate, and a memory unit. The spatiotemporal sequence data feature extraction and analysis process is as follows:
[0069] First, ConvLSTM is based on the hidden state H of the previous time step t-1 With the input X at the current time step t Splicing along the channel layer as the input of the forget gate unit, using Sigmoid as the activation function, to obtain a matrix F in the interval [0,1] tIt is used to characterize the probability of neuron information being forgotten in ConvLSTM. The greater the probability of the cell state being retained from 0 to 1, the greater the probability of the cell state being retained from the current time step input X. t Extract new neuron information and obtain candidate neuron information under the input through the input gate unit And the probability matrix I that determines whether the candidate information is retained t ; Then, the neuron information of the current time step is updated through F t and old neuron information C t-1 The dot product operation determines the useless state information that needs to be discarded, and I t and The dot product operation is used to obtain the new neuron information that needs to be added. The neuron information C of the current time step can be completed by adding the two results. t Update; Finally, through H t-1 and X t The output neuron information is obtained through the output gate unit as the probability O of the hidden layer state t and the neuron information C after activation using the tanh function t Multiply them together to get the hidden layer output H at the current moment t , H t It contains multiple hidden channels, each corresponding to the different spatiotemporal characteristics of the ocean environment at the current moment. This process is formalized as follows:
[0070]
[0071]
[0072] Where i, f, c, and o are the input gate, forget gate, cell state, and output gate, respectively. W and b are the weight coefficient and bias weight of the convolutional long short-term memory neural network, respectively. The subscripts i, f, c, and o represent the gate units to which W and b belong, respectively. The subscripts x and h are respectively in the corresponding gate units and the input X. t and the hidden state H t-1 W;I for Hadamard product t ,F t ,C t ,O t are the output matrices of the input gate, forget gate, cell state, and output gate at the current time step respectively; σ and tanh represent the Sigmoid function and the hyperbolic tangent function respectively; and * represent vector product and matrix Hadamard product respectively; C represents the number of hidden channels, Represents the hidden state matrix of the cth channel at time t, where the (N, M)th element in the matrix is represented as
[0073] The convolutional long short-term memory neural network uses the Back Propagation Through Time (BPTT) algorithm to update the weight coefficients and bias weights. The hidden layer output H of ConvLSTM is calculated through the above forward process. t Based on the target result, the network layer and time-direction errors are reversely calculated, and the gradient value corresponding to each weight parameter is calculated. The gradient descent algorithm is used to gradually optimize the weight coefficient until the error is minimized and the gradient value is 0. The convolutional long short-term memory neural network method described above can extract the spatiotemporal characteristics of the SIC spatiotemporal sequence and output them as the hidden layer state of each spatiotemporal feature.
[0074] Step 3: Determine the prediction effect of the MLS-ConvLSTM neural network model, optimize the parameters of the MLS-ConvLSTM neural network model, and output the trained MLS-ConvLSTM neural network model. The ConvLSTM network training process is as follows: Figure 8 As shown;
[0075] Preferably, step 3 comprises the following steps:
[0076] Step 301: Compare the prediction effect of the MLS-ConvLSTM network with the traditional method in an intuitive visual way, where the ConvLSTM network training process is as follows: Figure 8 The training process of the ConvLSTM network for polar SIC prediction is as follows: first, the target prediction object is determined to be the spatiotemporal variation characteristics of sea ice density in the waters near the Greenland Sea. Second, by setting the historical data length, input data length, and forecast duration parameters, a three-dimensional training dataset with temporal correlation is constructed and divided into a training set and a validation set. Next, the multi-layer convolutional loop structure parameters of the ConvLSTM network are configured, and the standardized training set data is input into the network for iterative training. The network weight parameters are dynamically updated using the backpropagation algorithm by calculating the loss function between the network output value and the actual label. Simultaneously, the loss function value and prediction accuracy index are calculated using the validation set to achieve a quantitative evaluation of the network performance. Finally, through multiple iterative optimizations, the trained prediction model is obtained, and the spatiotemporal evolution forecast results of the sea ice density in the sea area are output. By integrating spatiotemporal feature extraction with sequence prediction capabilities, this method significantly improves the forecast accuracy of sea ice parameter changes and the model generalization performance.
[0077] Using geographic polar coordinate projection mapping, a multi-dimensional comparative analysis of the forecast results of a multi-layer stacked convolutional long short-term memory network (MLS-ConvLSTM) was conducted with EOF+LSTM, CNN+LSTM methods, and ground-truth satellite data. Specifically, polar coordinate mapping and spatial interpolation were applied to the forecast data for the target sea area to generate a standardized sea ice density distribution map. A multi-model comparative verification framework was established to compare the high-resolution forecast fields output by the MLS-ConvLSTM with the forecast results of EOF+LSTM and CNN+LSTM, respectively. The accuracy of the forecasts was visually compared with traditional methods. In particular, the MLS-ConvLSTM, through its cascaded stacking design of multi-layer convolutional loop structures and its cross-layer gradient optimization mechanism, achieved superior spatial resolution in its forecast results compared to traditional methods.
[0078] Step 302: Select the root mean square error (RMSE) and the Pearson coefficient (PCC) as evaluation indicators to determine the forecast effect of the MLS-ConvLSTM network method; in order to more accurately and comprehensively evaluate the forecast effect of each model, the present invention selects two commonly used evaluation indicators: RMSE and PCC. RMSE measures the square root of the average of the sum of the squares of the differences between the forecast value and the actual observation value, which can reflect the size of the forecast error; while PCC evaluates the accuracy of the forecast by calculating the proportion of correct forecasts. The combined use of these two indicators not only takes into account the accuracy of the forecast, but also the reliability of the forecast. These two indicators can judge the deviation between the forecast data and the actual data, and can be used as a basis for scientifically evaluating the forecast effect of each forecast method;
[0079] Step 303: Adjust the parameters of the MLS-ConvLSTM neural network model according to the accuracy of the forecast effect, repeat steps 2 and 3, and output the trained MLS-ConvLSTM neural network model.
[0080] Preferably, in step 301, the forecast results output by the MLS-ConvLSTM neural network model are visually compared with those obtained by linear regression and empirical orthogonal function decomposition. This includes plotting the spatial distribution of key indicators such as sea ice coverage and thickness, and observing the changing trends and differences in the time series of different forecasting methods, thereby preliminarily determining the advantages of the LS-ConvLSTM neural network model in capturing the dynamic changes of sea ice. By comparing the differences in the detailed feature changes between the real image and the output images of each forecasting method, a preliminary judgment can be made on the performance of the MLS-ConvLSTM and traditional forecasting methods.
[0081] More accurately determine the prediction effect of the MLS-ConvLSTM network method, and select RMSE and PCC as evaluation indicators
[0082] In model training, loss (Loss) and Pearson coefficient (PCC) are the main basis for updating network parameters. This experiment calculates network loss through the root mean square error method (RMSE). The loss value here is obtained by comparing the output value predicted by the model with the given label value (ie, the true value or expected value) one by one, calculating the difference between the two, then accumulating the squares of these differences, and obtaining the average of the cumulative sum, and finally taking the square root of this average. It is used to characterize how big the gap is between the predicted value and the actual sea ice data indicated by the label value, and serves as a reference for network parameter adjustment. The RMSE method effectively comprehensively considers the overall prediction performance of the model on the entire data set, especially for those prediction results with larger deviations, giving higher penalty weights. Such a design helps the model to continuously reduce prediction errors and improve prediction accuracy during training, and ultimately achieve more accurate and reliable predictions of actual sea ice data. The root mean square error method obtains the deviation between the predicted value and the label value, and uses Y i represents the measured value, f(x i ) represents the network prediction value, and its specific formula is:
[0083]
[0084] The Pearson coefficient (PCC) can be understood as a numerical metric used to rigorously evaluate the accuracy of predictions. It directly reflects the degree of consistency between the model's predicted output and the true label (or expected value). The autocorrelation coefficient, used to characterize the accuracy of predictions after a partial model training phase, can be understood as a numerical value used to evaluate the accuracy of predictions; the closer it is to 1, the better the prediction. A PCC close to 1 may not be achieved for predictions with a sufficiently small loss. During model training, the loss function is designed to be as small as possible to indicate the degree of error between the model's predicted output and the true label. However, a sufficiently small loss value does not always guarantee a Pearson correlation coefficient (PCC) close to 1. This is because the loss function may focus on measuring the total amount or distribution of error, while the autocorrelation coefficient focuses more on assessing the proportion of correctly classified (or accurate) predictions in the prediction results. However, when the trained network can obtain a PCC close to 1 in the prediction, it means that the network model training effect is good, which means that the model has successfully learned the inherent rules of the data set during the training process and its prediction performance has reached a high level. t represents the obtained data matrix, γ0 is the product of the matrix standard deviation, and the specific formula for calculating the autocorrelation coefficient is as follows:
[0085]
[0086] Adjusting and optimizing network parameters is a core step in network training. Its fundamental purpose is to simultaneously achieve two key objectives: minimizing the loss function and improving the correlation coefficient between predicted values and true labels. These two objectives together constitute a dual perspective for optimizing network performance. During network training, network parameter modifications are aimed at minimizing the loss while simultaneously improving the correlation coefficient between predicted values and true labels. The loss function is a key metric that measures the difference between the model's predicted output and the true label. A smaller loss function generally indicates a smaller prediction error, meaning the model's performance is closer to ideal. Therefore, during network training, continuously adjusting network parameters to minimize the loss function through optimization techniques such as the backpropagation algorithm is an important way to improve model performance. The correlation coefficient, in this context, can be understood as a measure of the statistical correlation between the predicted value and the true label value. A high correlation coefficient indicates a high degree of consistency between the model's predicted output and the true label, indicating a high degree of model prediction accuracy. Conversely, a low correlation coefficient indicates a low degree of model prediction accuracy. In this case, the model parameters should be adjusted, and prediction and verification should be repeated until the model's prediction accuracy meets the requirements. Finally, the trained MLS-ConvLSTM neural network model is output. Therefore, during network training, in addition to focusing on reducing the loss function, appropriate regularization methods, data augmentation techniques, and early stopping strategies are also necessary to prevent overfitting and ensure that the model maintains high prediction accuracy even on the test set.
[0087] Based on the above results, we choose to use a 3-layer stacked ConvLSTM network with a 7-day input to study the performance of medium-term forecasting. The results are recorded in Table x.
[0088] Table 1. Results for different forecast periods
[0089]
[0090] Table 1 presents an analysis of the effect of the number of stacking layers on prediction performance, reporting prediction results for input sequence lengths of 3, 5, and 7 days. The results show that within a specific range, the validation set loss decreases and the prediction performance significantly improves compared to the input sequence. This suggests that as the number of stacking layers increases, the model may extract richer information from multiple spatial dimensions, thereby achieving better prediction results. Furthermore, when the input sequence length increases, the ConvLSTM network with a three-layer stacking structure demonstrates better performance. While maintaining the same prediction period, the model's predictive ability increases with increasing input sequence length.
[0091] Step 4: Use the MLS-ConvLSTM neural network model to analyze the ocean data to complete the prediction of polar sea ice density.
[0092] In actual experiments, the present invention used two classic SIC prediction networks, EOF+LSTM and CNN+LSTM, to compare the prediction performance of different networks. The architecture of the above three networks is a single-layer stacked network, in which the LSTM units and convolutional long short-term memory network units in the single-layer network have 64 hidden states. In addition, the input sequence length of all networks is 7 days. To explore the prediction timeliness performance of the encoding-decoding architecture, we set up all experiments with prediction periods of 7, 15, and 20 days, respectively. The prediction performance results are shown in Table 2.
[0093] Table 2. Results for different networks and different forecast periods
[0094]
[0095] Table 2 shows the impact of different network methods on prediction performance and provides clear experimental results, which is helpful for multi-dimensional analysis of SIC prediction effects. It is worth noting that as spatiotemporal models, CNN+LSTM and convolutional long short-term memory networks perform better than EOF+LSTM networks. This also confirms that EOF+LSTM and CNN+LSTM with too many redundant connections make it less likely for the network to capture local spatial features, while convolutional long short-term memory networks are effective in capturing these features. In addition, this work analyzes the prediction timeliness of the encoding-decoding architecture. The results of different forecast timelinesses show that as the forecast timeliness increases, the RMSE loss gradually increases and PCC(r) gradually decreases. Compared with other methods, the convolutional long short-term memory network shows superiority in medium-term predictions. We also tried to stack the above networks to obtain better prediction results. As the number of stacking layers increases, the convolutional long short-term memory network shows great success, while other methods perform mediocrely. Prediction performance is as follows Figure 5 shown.
[0096] according to Figure 5 The results show that all three networks perform well in terms of sea ice segmentation. However, the one-layer convolutional long short-term memory network can roughly describe the internal distribution of sea ice compared to the other two networks. Furthermore, the SIC prediction performance of each network gradually deteriorates with increasing forecast period.
[0097] To investigate the impact of network parameters on short- and medium-term SIC predictions, we varied the input sequence length, network architecture, and prediction period. For experiments exploring the network's input sequence length, we analyzed performance with input sequence lengths of 3, 5, and 7 days and a prediction period of 7 days. Similarly, we increased the number of stacked layers of the convolutional long short-term memory network and sequentially varied the network architecture to explore performance. The convolutional long short-term memory network in a single-layer network had 64 hidden states, the two-layer convolutional long short-term memory network had 64 and 128 hidden states, and the three-layer convolutional long short-term memory network had 64, 128, and 128 hidden states per layer, respectively. The results of these experiments are reported in Table 3.
[0098] Table 3. Results of different input sequence lengths and network structures
[0099]
[0100] Table 3 shows the impact of the number of stacked layers on forecasting performance and provides forecasting results for varying input sequence lengths (3, 5, and 7 days). The results show that, within a certain range, validation sample loss decreases and forecasting performance significantly improves. It is likely that as the number of layers increases, more information across different spatial dimensions is extracted, leading to better forecasting performance. Furthermore, the 3-layer stacked convolutional long short-term memory network exhibits better performance as the input sequence length increases. Given the same forecast period, the model performs better as the input sequence length increases. Based on these results, we selected a 3-layer stacked convolutional long short-term memory network with a 7-day input to investigate its medium-term forecasting performance. The results are reported in Table 4.
[0101] Table 4. Results for different forecast periods
[0102]
[0103]
[0104] Table 4 shows the performance of a three-layer stacked convolutional long short-term memory network at different forecast periods when the input sequence length is 7 days. The network still has good learning and generalization capabilities for medium-term forecasts, and its performance decreases as the number of forecast periods increases.
[0105] Intuitively represent the actual performance of the prediction results. For example, the 15-day prediction values and true values from 2019.03.30 to 2019.04.13 are selected to further evaluate the network. Figure 6 .like Figure 2 As shown in the figure, the three-layer stacked convolutional long short-term memory network performs well in predicting sea ice extension and sea ice density within the ocean area. Moreover, the performance gradually decreases as the forecast time increases, with the 7-day forecast performing better.
[0106] The foregoing is merely a preferred embodiment of the present invention. It should be noted that those skilled in the art may make various improvements and modifications without departing from the principles of the present invention, and such improvements and modifications are also within the scope of protection of the present invention. Structures, devices, and operating methods not specifically described or explained herein shall, unless otherwise specified or limited, be implemented in accordance with conventional means in the art.
Claims
1. A sea ice density prediction method based on multi-layer stacked ConvLSTM, characterized by: The following steps are included Step 1: Obtain polar ocean reanalysis data and perform preprocessing to eliminate land points; Step 2: Establish an MLS-ConvLSTM neural network model and perform simulation; Step 3: Determine the prediction effect of the MLS-ConvLSTM neural network model, optimize the parameters of the MLS-ConvLSTM neural network model, and output the trained MLS-ConvLSTM neural network model; Step 4: Use the MLS-ConvLSTM neural network model to analyze the ocean data to complete the prediction of polar sea ice density.
2. The method for predicting sea ice density based on a multi-layer stacked ConvLSTM according to claim 1, characterized in that: In step 1, the data preprocessing and analysis methods are as follows: Step 101: During the data reading process, the start and end time of the experimental data and the longitude and latitude information of the selected sea area are set to define the time and space range of the data required for the experiment; Step 102: Check the current year. If it is a leap year, set February of that year to 29 days; otherwise, set February of that year to 28 days. Step 103: Reading the geographical location, sea ice density and other information of the selected target sea area on a certain date from the database, and storing the information in a new storage file created for the target sea area; Step 104: Process all the files in the database that participate in the experiment in order, and repeat steps 102 and 103 until the data reading reaches the deadline; Step 105: Reduce the dimension of the read data by one, and remove the dimensional information representing the sea ice density.
3. The method for predicting sea ice density based on a multi-layer stacked ConvLSTM according to claim 1, characterized in that: In step 2, the ConvLSTM neural network model is an optimized multi-layer stacked ConvLSTM network architecture, which integrates two major components: an encoder and a decoder; The encoder is a composite structure integrating multiple convolutional layers and ConvLSTM units, and the convolution operation captures key information in spatial feature maps of different scales; At the same time, the ConvLSTM unit maintains the long-term spatiotemporal dependencies of the input tensor; The decoder, as a mirror-symmetric structure of the encoder, is also constructed by multiple layers of ConvLSTM and deconvolution operations. The working mechanism of the ConvLSTM unit is consistent with that of the encoder. The deconvolution operation is the opposite of the convolution operation. It is responsible for converting the encoder output tensor into the target output tensor, completing the reverse reconstruction of the information.
4. The method for predicting sea ice density based on a multi-layer stacked ConvLSTM according to claim 1, characterized in that: The step 3 comprises the following steps: Step 301: Compare the prediction results of the MLS-ConvLSTM network with those of traditional methods visually to determine the prediction effect. Step 302: Selecting the root mean square error (RMSE) and the Pearson coefficient (PCC) as evaluation indicators to determine the prediction effect of the MLS-ConvLSTM network method; Step 303: Adjust the parameters of the MLS-ConvLSTM neural network model according to the accuracy of the forecast effect, repeat steps 2 and 3, and output the trained MLS-ConvLSTM neural network model.
5. The method for predicting sea ice density based on a multi-layer stacked ConvLSTM according to claim 4, characterized in that: In step 301, the forecast results output by the MLS-ConvLSTM neural network model are visually compared with the forecast results obtained by linear regression and empirical orthogonal function decomposition; This includes drawing spatial distribution maps of key indicators such as sea ice coverage and thickness, and observing the changing trends and differences of different forecasting methods in time series, so as to preliminarily judge the advantages of the LS-ConvLSTM neural network model in capturing the dynamic changes of sea ice.
6. The method for predicting sea ice density based on a multi-layer stacked ConvLSTM according to claim 4, characterized in that: In step 302, the network loss is calculated by the root mean square error method, which calculates the deviation between the predicted value and the label value, and uses Y i represents the measured value, f(x i ) represents the network prediction value, and its specific formula is:
7. The method for predicting sea ice density based on a multi-layer stacked ConvLSTM according to claim 4, characterized in that: In step 302, use X t represents the obtained data matrix, γ0 is the product of the matrix standard deviation, and the specific formula for calculating the autocorrelation coefficient is as follows:
8. The method for predicting sea ice density based on a multi-layer stacked ConvLSTM according to claim 4, characterized in that: In step 303, during the network training process, the network parameters are continuously adjusted by means of back-propagation algorithm optimization to minimize the value of the loss function, and the correlation coefficient is used as a statistical correlation measure between the predicted value and the true label value; when the correlation coefficient is high, the predicted output of the model is highly consistent with the true label, that is, the prediction accuracy of the model is high; conversely, the prediction accuracy of the model is low, and the model parameters are adjusted at this time, and the prediction and verification are repeated until the prediction accuracy of the model meets the requirements, and finally the trained MLS-ConvLSTM neural network model is output.