A Global Ocean Environment Forecasting Method Based on Multi-Level Feature Aggregation
By building a global marine environment forecasting system with multi-level feature aggregation, the problems of high computing costs and low forecast accuracy in the existing technology are solved, and high-resolution global marine environment forecasting is achieved, which improves forecast accuracy and reduces calculation costs.
Patent Information
- Application Number
- CN202311650549.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-05
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-12-05
AI Technical Summary
The existing numerical model method driven by differential equations is too costly to calculate and the forecasting speed is slow. The existing data-driven deep learning marine environment forecast system is difficult to process 1/12 degree of high-resolution global data, and the forecasting accuracy needs to be improved.
Build a global marine environment forecasting system based on multi-level feature aggregation, including marine feature extraction module, marine feature aggregation module and global marine regional reduction module. Through multi-level feature extraction and aggregation, the problems of insufficient feature representation capabilities and high computing costs can be alleviated, and forecast accuracy can be improved.
On the premise of ensuring real-time forecasting, the accuracy of marine environmental forecasting is improved and the calculation cost is reduced, and high-resolution global marine environmental forecasting is achieved.
Smart Images

Figure CN117849903B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of ocean environmental forecasting, and particularly to a global ocean environmental forecasting method based on multi-level feature aggregation for optimizing forecasting accuracy and achieving medium- and long-term forecasting. Background Art
[0002] Ocean environmental forecasting refers to the process of predicting and estimating important elements in each layer of the global ocean environment (each layer corresponds to a depth vertically downward from the sea surface, for example, the first layer corresponds to a depth of 0.5 meters, and the second layer corresponds to a depth of 1.5 meters), such as temperature, salinity, flow velocity, sea surface height, etc., through technical means such as data collection and analysis and model calculation. It is a key tool for addressing important challenges such as climate change, ocean disasters, and ecosystem sustainability. Forecasting accuracy and real-time performance are the core concerns in this field.
[0003] In ocean environmental forecasting, performance evaluation is divided into two key aspects: accuracy and real-time performance. Accuracy reflects the precision of the forecasting method, while real-time performance emphasizes the rapid response ability of the method. For ocean environmental forecasting applied to tasks such as navigation, fisheries, ocean resource management, and natural disaster warning, real-time performance is crucial. If forecasting information cannot be provided in a timely manner in an emergency, it may lead to significant losses and dangers.
[0004] Currently, the development of ocean environmental forecasting methods mainly focuses on two directions: numerical model forecasting methods driven by differential equations and big data-driven ocean environmental forecasting methods.
[0005] Numerical model prediction methods driven by differential equations, such as the UK Met Office's FOAM (Forecast Ocean Assimilation Model) system, the PSY3 and PSY4 systems of the French GLORYS12 center, the GIOPS (Global Ice Ocean Prediction System - Canadian Operational Network of Coupled Environmental Prediction Systems, GIOPS) system of Environment Canada, and the BLK (The Ocean Model Analysis and Prediction System, BLK) system of the Australian Bureau of Meteorology, use mathematical equations and computer simulations of ocean processes to generate ocean environmental prediction data. However, numerical model prediction methods driven by differential equations usually require a large amount of computing time, resulting in insufficient timeliness of predictions. Moreover, numerical model prediction methods depend on human understanding and mastery of the laws of the ocean environment. However, at present, human understanding of the ocean system is still very limited. Therefore, the prediction level cannot fully meet the needs of social development, and the accuracy still needs to be improved.
[0006] The big data-driven ocean environmental prediction method draws on deep learning technology and makes full use of satellite remote sensing data, sensor observation data, and historical information in the ocean database to generate real-time predictions. Due to the development of GPU computing power, the big data-driven ocean environmental prediction method has more obvious advantages in real-time prediction compared with the numerical model prediction method driven by differential equations.
[0007] The deep learning forecasting method driven by big data has first achieved great progress in the field of weather forecasting. The literature "Bi, K., Xie, L., Zhang, H. et al. Accurate medium-range global weather forecasting with 3D neural networks. Nature 619, 533–538 (2023)." (the paper by Bi, K., Xie, L., etc.: Accurate Medium-Range Global Weather Forecasting with 3D Neural Networks) introduced a high-resolution global weather forecasting system - Pangu, based on 3D-Transformer. This forecasting system can generate 7-day weather forecasting results at a spatial resolution of 0.25°, and its evaluation indicators on 80% of weather variables exceed those of the Integrated Forecasting System-High Resolution (IFS-HRES), an advanced numerical weather forecasting system operated by the European Centre for Medium-Range Weather Forecasts (ECMWF). Pangu has realized a deep learning-based weather forecasting model, which has the advantages of fast inference speed and high forecasting accuracy. Due to the great success of the deep learning forecasting method driven by big data in the field of weather forecasting, recently there have also been works on using the deep learning method driven by big data to forecast the ocean environment. The literature "Xiong, W., Xiang, Y., Wu, H., Zhou, S., Sun, Y., Ma, M., & Huang, X. (2023). AI-GOMS: Large AI-Driven Global Ocean Modeling System. ArXiv, abs / 2308.03152." (the paper by Xiong, W., Xiang, Y., etc.: Large AI-Driven Global Ocean Modeling System: AI-GOMS) is an ocean environment forecasting system based on an autoencoder-based model with Fourier operators. This system takes the HYCOM global reanalysis ocean data with a spatial resolution of 0.25° as input and conducts daily global forecasts for the next 30 days. However, it does not use the evaluation criteria of the Intercomparison and Validation Task Team (IVTT), an internationally authoritative organization for evaluating ocean forecasting systems, to comprehensively evaluate the model performance, making it impossible to directly compare the forecasting performance of this system with that of numerical forecasting models.
[0008] Pangu and AI-GOMS have proven that deep learning has significant performance advantages over traditional numerical models in the fields of meteorological forecasting and marine environmental forecasting, but they can only forecast global data with a spatial resolution of 0.25°. The spatial resolution of ocean reanalysis data is as high as 1 / 12 degree, which will result in a large computational overhead during model training. In addition, ocean data has defects such as few observation sites, a single observation channel, and insufficient feature characterization capabilities, which makes the reanalysis data have a certain gap compared with the actual observation values. Only through standardized methods and unified standard evaluation can the model performance be comprehensively and accurately evaluated. These important issues have always hindered the development of data-driven deep learning methods for marine environmental forecasting.
[0009] How to improve the real-time nature and accuracy of forecasts and reduce computing costs while alleviating the lack of feature characterization capabilities for high-resolution ocean data remains a technical problem that needs to be urgently solved by relevant technicians in this field. Summary of the invention
[0010] The technical problem to be solved by the present invention is that the existing numerical model method driven by differential equations has too high computational cost and slow prediction speed, and the existing data-driven deep learning marine environment prediction system is difficult to process 1 / 12 degree high-resolution global data and the prediction accuracy needs to be improved. A global marine environment prediction method based on multi-level feature aggregation is proposed. Under the premise of ensuring the real-time forecast, a multi-level feature aggregation network is used to extract multi-scale features, alleviate the problems of insufficient feature representation ability and high computational cost, improve the prediction accuracy, and achieve high-resolution global marine environment forecast.
[0011] To solve the above technical problems, the technical solution of the present invention is as follows: construct a global ocean environmental forecasting system based on multi-level feature aggregation. The system consists of an ocean feature extraction module, an ocean feature aggregation module, and a global ocean area restoration module. Construct the data set required for the global ocean environmental forecasting system, and divide the data set into a training set, a validation set, and a test set. Align the time resolution and spatial resolution of the training set, validation set, and test set, extract the key layer data, and finally perform normalization processing. Then use the training set to train the ocean feature extraction module, ocean feature aggregation module, and global ocean area restoration module in the global ocean environmental forecasting system. After one round of training, use the validation set to test the forecasting accuracy of the trained global ocean environmental forecasting system. If the network weight parameters of this round make the current forecasting accuracy optimal, then save the weight parameters of the trainable modules (ocean feature extraction module, ocean feature aggregation module, and global ocean area restoration module) in the network of this round. After the last round of training, a trained global ocean environmental forecasting system with the best forecasting performance can be obtained; finally, use the trained global ocean environmental forecasting system with the best forecasting performance to perform global ocean environmental forecasting based on the global ocean environmental data of a certain day input by the user. The input global ocean environmental data is a three-dimensional grid data containing elements of each layer, and the elements include: three sea surface elements (sea surface wind field, sea surface height, sea surface temperature), as well as seawater temperature, salinity, and flow velocity from the sea surface to the 33rd layer (the flow velocity includes two components, namely the eastward flow velocity of seawater and the northward flow velocity of seawater). Each layer corresponds to a depth vertically downward from the sea surface. For example, the first layer corresponds to a depth of 0.5 meters, the second layer corresponds to a depth of 1.5 meters, and the 33rd layer corresponds to a depth of 643.5 meters. The input format is V×H×W, where V represents the total number of elements containing elements of each layer (V = 3 + number of layers × 4), H represents the number of grid points in the latitude direction of the grid, and W represents the number of grid points in the longitude direction of the grid), and obtain the global ocean environmental forecasting results of the target elements (the target elements include sea surface height, sea surface temperature, temperature, salinity, and flow velocity from the 1st to the 33rd layer) for the corresponding number of days.
[0012] The technical solution of the present invention includes the following steps:
[0013] In the first step, construct a global ocean environmental forecasting system based on multi-level feature aggregation. As Figure 1 shown, the ocean environmental forecasting system consists of an ocean feature extraction module, an ocean feature aggregation module, and a global ocean area restoration module.
[0014] The ocean feature extraction module is connected to the ocean feature aggregation module and the global ocean area restoration module. The ocean feature extraction module receives the global ocean environment data input by the user (i.e., three-dimensional grid data containing various ocean layer elements), extracts ocean features from the global ocean environment data, and sends the feature maps of the ocean features to the ocean feature aggregation module and the global ocean area restoration module. The ocean feature extraction module consists of a two-dimensional convolutional neural network (see the paper "Dosovitskiy A, Beyer L, Kolesnikov A, et al. An image is worth 16x16 words: Transformers for image recognition at scale [J]. arXiv preprint arXiv:2010.11929, 2020." by Dosovitskiy A, Beyer L, etc.: Transformers for large-scale image recognition) and a Layer Normalization layer (see the paper "Ba J L, Kiros J R, Hinton G E. Layer normalization [J]. arXiv preprint arXiv:1607.06450, 2016." by Lin T Ba J L, Kiros J R, etc.: Hierarchical normalization). The two-dimensional convolutional neural network divides the three-dimensional grid data into patches, each patch having a size of 6×6 pixels, extracts the ocean features of the global ocean environment data, and sends the feature maps of the ocean features to the first Layer Normalization layer. The stride and the convolutional kernel size of the two-dimensional convolutional neural network are both 6. The first Layer Normalization layer normalizes the feature maps of the ocean features to ensure the stability of training, and sends the normalized feature maps of the ocean features to the ocean feature aggregation module.
[0015] The ocean feature aggregation module is connected to the ocean feature extraction module and the global ocean area restoration module. The ocean feature aggregation module consists of 5 Spatial Information Extraction (SIE) networks with non-shared weights (these 5 networks are respectively denoted as the first, second, third, fourth, and fifth SIE networks), an upsampling module, and a downsampling module. The first SIE network and the fifth SIE network have the same structure, both consisting of a local SIE network and a global SIE network; the second SIE network to the fourth SIE network have the same structure, all consisting of two consecutive local SIE networks and a global SIE network. The first SIE network is composed of the first local SIE network and the first global SIE network. The first local SIE network consists of a window multi-head self-attention layer and a Multilayer Perceptron (MLP). The window multi-head self-attention layer receives the feature map of the normalized ocean features from the ocean feature extraction module, performs hierarchical normalization on the feature map of the normalized ocean features, divides the hierarchically normalized feature map into multiple windows of size 7×7 (each window contains 7×7 patches), uses the window multi-head self-attention mechanism (see the literature "Liu Z, Lin Y, Cao Y, et al. Swin transformer: Hierarchical vision transformer using shifted windows [C] / / Proceedings of the IEEE / CVF international conference on computer vision. 2021: 10012-10022." by Liu Z, Lin Y, etc.: Hierarchical vision transformer using shifted windows) and residual connection operations to obtain the feature map enhanced by the attention mechanism, and sends the feature map enhanced by the attention mechanism to the MLP. The MLP consists of two linear layers, receives the feature map enhanced by the attention mechanism from the window multi-head self-attention layer, performs hierarchical normalization, linear projection, and residual connection operations on it, obtains the feature map that fuses the local spatial information within the window, and sends the feature map that fuses the local spatial information within the window to the first global SIE network. The first global SIE network consists of a feature grouping network, a group feature fusion network, and a group feature propagation network. The feature grouping network of the first global SIE network receives the feature map that fuses the local spatial information within the window from the first local SIE network, uses the cross-attention mechanism to group and aggregate the features in the feature map that fuses the local spatial information within the window to obtain group features, and sends the group features to the group feature fusion network.The main part of the group feature fusion network is composed of MLP-mixer (see the paper "Tolstikhin IO, Houlsby N, Kolesnikov A, et al. Mlp-mixer: An all-mlp architecture for vision [J]. Advances in neural information processing systems, 2021, 34: 24261-24272." by Tolstikhin IO, Houlsby N, etc.). MLP-mixer can achieve a good balance between model performance and computational overhead. The group feature fusion network receives group features from the feature grouping network, exchanges group feature information and updates the group features, so that each group feature can contain the current group information and reflect the global information to a certain extent, obtaining the updated group features, and sending the updated group features to the group feature propagation network. The group feature propagation network uses the cross-attention mechanism to propagate the updated group features to all features in the feature map that fuses the local spatial information within the fusion window, and uses the depthwise separable convolutional neural network to perform convolutional feature extraction on the feature map of the local spatial information within the window, obtaining the feature map after the first-stage feature enhancement, and sending the feature map after the first-stage feature enhancement to the downsampling module. The local SIE network structures in the second to fifth SIE networks are the same as the first local SIE network structure, and the global SIE network structures in the second to fifth SIE networks are the same as the first global SIE network.
[0016] The downsampling module is connected to the first SIE network and the second SIE network and consists of a linear layer. The downsampling module receives the feature map enhanced in the first stage from the first SIE network, combines every 2×2, that is, 4 adjacent patches in the feature map enhanced in the first stage together, and performs hierarchical normalization after splicing in the channel dimension to obtain a channel-spliced feature map with a channel number 4 times that of the feature map enhanced in the first stage. The downsampling module performs a linear transformation in the channel dimension of the channel-spliced feature map to halve the channel number of the channel-spliced feature map, obtaining a downsampled feature map. The height and width of the downsampled feature map are half of those of the feature map enhanced in the first stage, and the channel number is 2 times that of the feature map enhanced in the first stage, that is, the downsampling ratio becomes twice that of the first SIE network. Then, the downsampled feature map is sent to the second SIE network. Two consecutive local SIE networks (i.e., the second local SIE network and the third local SIE network) and a global SIE network (i.e., the second global SIE network) of the second SIE network enhance the downsampled feature map to obtain a feature map enhanced in the second stage and send the feature map enhanced in the second stage to the third SIE network; two consecutive local SIE networks (i.e., the fourth local SIE network and the fifth local SIE network) and a global SIE network (i.e., the third global SIE network) of the third SIE network enhance the feature map enhanced in the second stage to obtain a feature map enhanced in the third stage and send the feature map enhanced in the third stage to the fourth SIE network; two consecutive local SIE networks (i.e., the sixth local SIE network and the seventh local SIE network) and a global SIE network (i.e., the fourth global SIE network) of the fourth SIE network enhance the feature map enhanced in the third stage to obtain a feature map enhanced in the fourth stage and send the feature map enhanced in the fourth stage to the upsampling module.
[0017] The upsampling module is connected to the fourth SIE network and the fifth SIE network and consists of two linear layers. The upsampling module receives the feature map enhanced in the fourth stage from the fourth SIE network. The first linear layer of the upsampling module performs a linear transformation in the channel dimension of the feature map enhanced in the fourth stage to obtain a channel-transformed feature map with a channel number 2 times that of the feature map enhanced in the fourth stage; change the shape of the channel-transformed feature map so that the height and width are 2 times those of the channel-transformed feature map and the channel number is half of that of the channel-transformed feature map, obtaining a shape-transformed feature map and sending the shape-transformed feature map to the second linear layer of the upsampling module. The second linear layer of the upsampling module fuses the channel-domain information of the shape-transformed feature map without changing the size of the channel dimension to obtain an upsampled feature map and sends the upsampled feature map to the fifth SIE network. The height and width of the upsampled feature map are 2 times those of the feature map enhanced in the fourth stage, and the channel number is half of that of the feature map enhanced in the fourth stage.
[0018] The fifth SIE network performs fifth-stage feature enhancement on the upsampled feature map to obtain a feature map with fifth-stage feature enhancement, and sends the feature map with fifth-stage feature enhancement to the global ocean area restoration module.
[0019] The local and global SIE networks in the first to fifth SIE networks enable each SIE network to have local and global spatial perception capabilities and enhance the feature representation ability. To better capture the dynamic changes of ocean water bodies and further reduce the computational amount, a masking mechanism is added to the window multi-head self-attention layer in the local SIE network and the feature grouping network in the global SIE network, so that the attention score of the land area is 0. Hierarchical design and multi-scale feature extraction of ocean environmental elements are achieved through the downsampling module and the upsampling module, and multi-scale information of the global ocean can be captured. The ocean feature aggregation module extracts multi-scale features of global and local ocean area information through five stages, and sends the feature map with fifth-stage feature enhancement to the global ocean area restoration module to improve the prediction accuracy of the global ocean environmental forecasting system.
[0020] The global ocean area restoration module consists of a second Layer Normalization layer and a two-dimensional transposed convolutional neural network, and is connected to the fifth SIE network of the ocean feature extraction module and the ocean feature aggregation module. It receives the feature map with fifth-stage feature enhancement from the ocean feature aggregation module and the feature map of the normalized ocean features from the ocean feature extraction module. The second Layer Normalization layer of the global ocean area restoration module first concatenates the feature map with fifth-stage feature enhancement and the feature of the normalized ocean features in the channel dimension to obtain a channel concatenated feature map, then normalizes the channel concatenated feature map, and sends the normalized channel concatenated feature map to the two-dimensional transposed convolutional neural network layer. The stride and convolution kernel size of the two-dimensional transposed convolutional neural network layer are both 6. The normalized channel concatenated feature map is upsampled 6 times using a 6×6 transposed convolution with a stride of 6 to obtain the predicted three-dimensional grid data. The height and width of the predicted three-dimensional grid data are the same as those of the user-input global ocean environmental data, and the channel dimension is the number of total target elements (including sea surface height, sea surface temperature, temperature, salinity, and flow velocity from layer 1 to layer 33), and the predicted three-dimensional grid number is the global ocean environmental forecasting result.
[0021] Second step, construct the training set, validation set, and test set, the method is:
[0022] 2.1 Download GLORYS12 global ocean reanalysis data, ERA5 wind field reanalysis data, and GHR sea surface temperature (SST) satellite data, and perform quality control on the data, the method is:
[0023] 2.1.1 Download GLORYS12 global ocean reanalysis data (see the paper by Jean-Michel L, Eric G, Romain BB, et al., "The Copernicus global 1 / 12° oceanic and sea ice GLORYS12 reanalysis" [J]. Frontiers in Earth Science, 2021, 9: 698876.), ERA5 wind field reanalysis data (see the paper by Hersbach H, Bell B, Berrisford P, et al., "The ERA5 global reanalysis" [J]. Quarterly Journal of the Royal Meteorological Society, 2020, 146(730): 1999-2049.), and GHR sea surface temperature (SST) data (see the paper by Martin M, Dash P, Ignatov A, et al., "Group for High Resolution Sea Surface temperature (GHRSST) analysis fields inter-comparisons. Part 1: A GHRSST multi-product ensemble (GMPE)" [J]. Deep Sea Research Part II: Topical Studies in Oceanography, 2012, 77: 21-30.). Collect data with a time span from 1993 to 2020 from these three types of data as a dataset for ocean environmental forecasting. The GLORYS12 global ocean reanalysis data has a spatial resolution of 1 / 12°, a spatial span of -180° to 180° in the longitude direction and -80° to 90° in the latitude direction, a time resolution of 1 day, and includes reanalysis data of multiple elements (temperature, salinity, flow velocity, sea level height. The flow velocity includes two components: the eastward flow velocity of seawater and the northward flow velocity of seawater), and each element contains multiple layers of data, with each layer corresponding to a specific depth vertically downward from the sea surface.The present invention only forecasts the seawater temperature, salinity, flow velocity, and sea level height. Therefore, first, five elements, namely thetao (temperature), so (salinity), uo (eastward seawater flow velocity), vo (northward seawater flow velocity), and zos (sea surface height) in the GLORYS12 data, are selected. For temperature, salinity, eastward seawater flow velocity, and northward seawater flow velocity, data from the 1st to 33rd layers (from a depth of 0.5 meters vertically downward from the sea surface to a depth of 643.5 meters) are extracted; for the sea surface height, data from the 1st layer are extracted. The sea surface wind has a significant impact on the marine environment. Therefore, the ERA5 reanalysis data is also used as a component of the dataset for marine environment forecasting. The spatial resolution of the ERA5 wind field reanalysis data is 1 / 4 degree, with a spatial span of -180 degrees to 180 degrees in the longitude direction and -90 degrees to 90 degrees in the latitude direction. The time resolution is 1 hour, including two single-layer elements, u10 (eastward wind speed at 10 meters) and v10 (northward wind speed at 10 meters). The spatial resolution of the sea surface temperature data of GHR is 1 / 20 degree, with a spatial span of -180 degrees to 180 degrees in the longitude direction and -90 degrees to 90 degrees in the latitude direction. The time resolution is 1 day, and it only contains satellite data of the sea surface temperature (analysed_sst).
[0024] 2.1.2 The downloaded GLORYS12 global ocean reanalysis data, ERA5 wind field reanalysis data, and GHR sea surface temperature (SST) satellite data are all in nc format. Before formally processing the data, it is necessary to first perform quality control on the data. Using the xarray library in python, the above data is read starting from January 1, 1993. If the data is missing, damaged, or some elements in the data are missing (such data are "dirty" data with problems), then the "dirty" data is written to the log. Then, the good data corresponding to the "dirty" data recorded in the log is downloaded again to ensure that there are no such problems with the data before it is received by the global ocean environment forecasting system based on multi-level feature aggregation;
[0025] 2.2 The spatial resolutions and time resolutions of the three datasets, namely GLORYS12 global ocean reanalysis data, ERA5 wind field reanalysis data, and GHR sea surface temperature satellite data, are different. Let the number of days from 1993 to 2020 be D, and let the original datasets of the three datasets, GLORYS12 global ocean reanalysis data, ERA5 wind field reanalysis data, and GHR sea surface temperature satellite data, be M, E, and G. Next, the time and spatial resolutions of M, E, and G are aligned. The method is as follows:
[0026] 2.2.1 Align the time resolution of the ERA5 wind field reanalysis data with the GLORYS12 global ocean reanalysis data and the GHR sea surface temperature satellite data. The method is as follows:
[0027] 2.2.1.1 Let the variable d = 1, and initialize the ERA5 wind field reanalysis dataset E1 after time resolution alignment to be empty;
[0028] 2.2.1.2 Extract the data at specific time points from the ERA5 wind field reanalysis data on the d-th day in E, that is, select the data at 00, 06, 12, and 18 hours of the day, and take the average of the sea surface wind field data at these four times as the daily average data of the wind field on the d-th day, and put the daily average data of the wind field on the d-th day into E1;
[0029] 2.2.1.3 If d ≤ D, let d = d + 1, and go to 2.2.1.2; if d > D, obtain the ERA5 wind field reanalysis dataset E1 after time resolution alignment, and go to 2.2.2;
[0030] 2.2.2 Align the spatial resolutions of the two datasets M and E1 with G. The method is as follows:
[0031] 2.2.2.1 Let d = 1, and initialize the reanalysis dataset E2 of ERA5 and the GHR sea surface temperature dataset G1 after spatial resolution alignment to be empty;
[0032] 2.2.2.2 Perform bilinear interpolation on the daily average wind field data in E1 and the GHR sea surface temperature data on the d-th day in G respectively, to obtain the wind field data on the d-th day with a spatial resolution of 1 / 12 degree, and put it into E2; and obtain the sea surface temperature data on the d-th day with a spatial resolution of 1 / 12 degree, and put it into G1;
[0033] 2.2.2.3 If d ≤ D, let d = d + 1, and go to 2.2.2.2; if d > D, obtain the ERA5 wind field reanalysis dataset E2 with a spatial resolution of 1 / 12 degree and the GHR sea surface temperature dataset G1, and go to 2.2.3;
[0034] 2.2.3 Align the spatial ranges of the two datasets E2 and G1 with that of M. The method is as follows:
[0035] 2.2.3.1 Let d = 1, and initialize the set of the ERA5 wind field data E3 and the GHR sea surface temperature data G2 after spatial range alignment to be empty;
[0036] 2.2.3.2 Select the part of the daily average wind field data and the GHR sea surface temperature data on the d-th day in E2 and G1 from -180 degrees to 180 degrees in the longitude direction and from -80 degrees to 90 degrees in the latitude direction, and put the daily average wind field data on the d-th day in this spatial range into E3; put the GHR sea surface temperature data on the d-th day in this spatial range into G2;
[0037] 2.2.3.3 If d ≤ D, let d = d + 1, and go to 2.2.3.2; if d > D, obtain the re - analyzed ERA5 wind field dataset E3 after spatial - range alignment and the sea - surface temperature dataset G2 of GHR, and go to 2.4;
[0038] 2.4 After aligning the time resolution and spatial resolution of the data, the data in E3, G2, and M are still in nc format. To facilitate data reading in the ocean environmental forecasting system, it is necessary to convert the nc format to npy format. The method is as follows:
[0039] 2.4.1 Let the variable d = 1, and initialize the dataset S in npy format to be empty;
[0040] 2.4.2 Read the GLORYS12 global ocean re - analysis data, ERA5 wind field re - analysis data, and GHR sea - surface temperature data on the d - th day from M, E3, and G2 respectively. Sequentially extract the data of the 1st to 33rd layers of the thetao, so, uo, vo, zos elements in the GLORYS12 data, the u10, v10 elements in the ERA5 wind field data, and the analysed_sst element in GHR; then splice the above - selected element data together, save it as an npy file, and then put it into S;
[0041] 2.4.3 If d ≤ D, let d = d + 1, and go to 2.4.2; if d > D, the generation of npy data is completed, and a new dataset S is obtained, and go to 2.5.
[0042] 2.5 Divide S into a training set, a validation set, and a test set, and calculate the mean and standard deviation of the training - set data, and normalize the data. The method is as follows:
[0043] 2.5.1 Use the mv command to move the data from 1993 to 2017 in S to the training set train, move the data in 2018 to the validation set val, and move the data from 2019 to 2020 to the test set test;
[0044] 2.5.2 According to the overall mean and standard deviation of the training - set train data, use the standardization method to standardize the data in the training set train to obtain the standardized training set D M . The standardization method is as follows:
[0045] 2.5.2.1 Let the variable d = 1, take the data on the d - th day in train as X d , and initialize the standardized training set D M to be empty;
[0046] 2.5.2.2 Standardize X d The specific process is shown in formula (1):
[0047]
[0048] where μ is the overall mean of the training set train data, and σ is the overall standard deviation. After standardizing X d put it into D M ;
[0049] 2.5.2.3 If d ≤ D, let d = d + 1, and go to 2.5.2.2; if d > D, the standardization of the training set train data is completed, and go to 2.5.3;
[0050] 2.5.3 According to the overall mean and standard deviation of the training set data (for the training process of the model, the data of the validation set and the test set are unknown, so only the overall mean and standard deviation of the training set data can be used), use the standardization method described in step 2.5.2 to standardize the data in the validation set val, and obtain the standardized validation set D V ;
[0051] 2.5.4 According to the overall mean and standard deviation of the training set data, standardize the data in the test set test, and obtain the standardized test set D T .
[0052] In the third step, use the gradient backpropagation method to train the global ocean environmental forecasting system constructed in the first step, and obtain the network weight parameters with the best forecasting effect on the validation set D v . The method is as follows:
[0053] 3.1 Initialize the network weight parameters of each module in the ocean environmental forecasting system. Use a normal distribution with a mean of 0 and a variance of 0.01 to initialize the network weight parameters of the ocean feature extraction module, the ocean feature aggregation module, and the global ocean area restoration module.
[0054] 3.2 Set the training parameters of the ocean environmental forecasting system. Set the warm-up learning rate to 5×10 -8 , the warm-up training step to 3, and the learning rate learning_rate to 5×10 -5 . Select AdamW (see the paper "Loshchilov I, Hutter F. Decoupled weight decay regularization [J]. arXiv preprint arXiv:1711.05101, 2017." by Loschchilov I, Hutter F: Decoupled weight decay regularization) as the model training optimizer. The hyperparameters of this model training optimizer are β1 = 0.9, β2 = 0.95, and "weight decay" = 1×10 -3The batch size (mini_batch_size) of network training is 1. The maximum number of training steps (maxepoch) is 50.
[0055] 3.3 Train the ocean environmental forecasting system by taking the square of the difference between the forecasting result output by the ocean environmental forecasting system during one training and the true value as the loss value (loss), and use gradient backpropagation to update the network weight parameters until the loss value reaches the threshold or the number of training steps reaches maxepoch. For each round of training, if the forecasting accuracy of the current system on the validation set D v is the best, save the network weight parameters of the current round. The method is as follows:
[0056] 3.3.1 Let the training step epoch = 1. One pass through all the data in the training set is one epoch, and initialize the batch number N b = 1;
[0057] 3.3.2 The ocean feature extraction module reads the N M -th batch from D b , a total of B (0 ≤ B ≤ 16 and B is a positive integer) three-dimensional grid data, and records these B three-dimensional grid data in matrix form I train , and I train contains B matrices of H×W×V. Here, H represents the height of the input grid data, W represents the width of the input grid data, and V represents the channel dimension (V = 3 + number of layers × 4, that is, 3 sea surface elements, as well as the temperature, salinity, eastward velocity of seawater, and northward velocity of seawater at each layer).
[0058] 3.3.3 The ocean feature extraction module uses the ocean feature extraction method to extract features from I train , and sends the normalized feature map X containing the ocean features of I train to the ocean feature aggregation module and the global ocean area restoration module. The method is as follows:
[0059] 3.3.3.1 The two-dimensional convolutional neural network in the ocean feature extraction module extracts the ocean features of I train , and obtains the feature map containing ocean features. The method is: use a 6×6 convolution with a stride of 6 to perform patch division on the B three-dimensional grid data of I train , the size of each patch is 6×6, and the representation dimension of each patch is 416, to obtain the feature map containing ocean features. Then send the feature map containing ocean features to the first Layer Normalization layer;
[0060] 3.3.3.2 The first Layer Normalization layer normalizes the feature maps containing ocean features at the level of each three-dimensional grid data to obtain the normalized feature map X containing ocean features of I train The normalization operation helps to improve the training effect of the network. The first Layer Normalization layer sends X to the ocean feature aggregation module and the global ocean area restoration module. The resolution of X is The number of channels of X is 416;
[0061] 3.3.4 The ocean feature aggregation module receives X from the ocean feature extraction module, generates the feature-enhanced feature map X5 of the fifth stage, and sends X5 to the global ocean area restoration module. The method is as follows:
[0062] 3.3.4.1 The first SIE network receives X from the ocean feature extraction module, and uses a feature aggregation method based on local and global spatial information extraction to perform local spatial self-attention enhancement and global information fusion on X to obtain the first-stage feature-enhanced feature map X 1 of X. The method is as follows:
[0063] 3.3.4.1.1 The first local SIE network of the first SIE network uses a local feature extraction method to extract features from X. The method is as follows: The window multi-head self-attention layer of the first local SIE network performs hierarchical normalization on X, divides X into multiple 7×7 windows, that is, each window contains 7×7 patches. Then the window multi-head self-attention layer of the first local SIE network parallelly uses the window multi-head self-attention mechanism to perform local spatial self-attention enhancement on X within each window, and then performs a residual connection to obtain the feature map after the residual connection During the calculation process of the window multi-head self-attention mechanism, the attention weights of the features corresponding to the land areas in X are set to 0 using the mask mechanism. Therefore, the land part of the area will not affect the ocean area during the calculation process. Then the window multi-head self-attention layer of the first local SIE network sends the feature map to the multi-layer perceptron of the first local SIE network. The multi-layer perceptron of the first local SIE network first performs hierarchical normalization on , and then uses a linear mapping to fuse in the channel dimension, and then performs a residual connection operation to obtain the feature map that fuses the local spatial information within the window and sends to the first global SIE module.
[0064] 3.3.4.1.2 The first global SIE network of the first SIE network receives from the first local SIE network and uses a feature enhancement method to perform on Extract and fuse global information to obtain a feature map X with enhanced features in the first stage that combines local and global spatial information 1 , and the method is as follows:
[0065] 3.3.4.1.2.1 The first feature grouping network randomly initializes a set of features G, where G contains N features, and the dimension of each feature is d (here it is 416). The first feature grouping network first uses the multi-head cross-attention mechanism to aggregate and update the features in, obtaining the updated set of features G′, where each feature in G′ aggregates the information of a cluster of patches with similar semantics. During the calculation of the multi-head cross-attention mechanism, a masking mechanism is used to mask the features corresponding to the land part in, so that the land part does not participate in the calculation of the cross-attention mechanism. Then the first feature grouping network sends the set of features G′ to the first set of feature fusion networks.
[0066] 3.3.4.1.2.2 The first set of feature fusion networks receives G′ from the first feature grouping network, exchanges information among each feature in G′, and then fuses and updates the information of the set of features G′ to obtain the updated set of features The first MLP layer of the first set of feature fusion networks fuses on the G′ spatial domain to obtain a set of features with fused spatial information The second MLP layer fuses on the channel domain to obtain the updated set of features The first set of feature fusion networks sends the updated set of features to the first set of feature propagation networks.
[0067] 3.3.4.1.2.3 The first set of feature propagation networks receives from the first set of feature fusion networks and receives from the first local SIE network. Using the multi-head cross-attention mechanism, taking the features in as the query vector, taking the set of features as the key vector and value vector, according to the different features in , different weights are assigned to each feature in the set of features , and then the features in are weighted and summed to obtain the global spatial feature U, and U is sent to the first set of feature propagation networks.
[0068] 3.3.4.1.2.4 The feed-forward neural network of the first set of feature propagation networks concatenates U with the features in to obtain the concatenated feature map X 1′ , and for X 1′Perform a linear transformation, and then obtain the feature map X after residual connection through the residual connection. 1″ ; The depthwise separable convolutional neural network performs convolutional feature extraction on X 1″ to obtain the feature map X of the first-stage feature enhancement that fuses local and global spatial information. 1 . The resolution of X 1 is X 1 has 416 channels. The first group of feature propagation networks of the first SIE network sends X 1 to the downsampling module.
[0069] 3.3.4.2 The downsampling module receives X 1 from the first group of feature propagation networks of the first SIE network, and performs downsampling on X 1 using the patch combination splicing and channel-domain linear transformation method to obtain the downsampled feature map X D , and sends X D to the second SIE network. The method is as follows:
[0070] 3.3.4.2.1 Combine each 2×2 adjacent patch in X 1 , and perform hierarchical normalization operations after splicing these 4 patches in the channel domain to obtain a channel-spliced feature map with 4 times the number of channels of X 1 (the channel dimension of the spliced feature map is 4×416);
[0071] 3.3.4.2.2 Perform a linear transformation on the channel dimension of the channel-spliced feature map to halve the number of channels of the channel-spliced feature map (i.e., 2×416) to obtain the downsampled feature map X D . The resolution of the feature map of X D is X D The number of channels of the feature map is 832. The downsampling ratio of the downsampled feature map X D is twice that of X 1 , thus forming a hierarchical structure design to achieve multi-scale feature extraction and also reduce the computational overhead.
[0072] 3.3.4.2.3 The downsampling module sends X D to the second SIE network.
[0073] 3.3.4.3 The second SIE network receives the downsampled feature map X D sent by the downsampling module. The second local SIE network and the third local SIE network use the local feature extraction method described in step 3.3.4.1.1 to perform operations on X D Perform feature extraction twice. The second global SIE network uses the feature enhancement method described in step 3.3.4.1.2 to enhance the features of X after feature extraction twice D Extract and fuse the global information for feature enhancement to obtain the feature map X with feature enhancement in the second stage 2 , and send X 2 to the third SIE network. The resolution of the feature map of X 2 is X 2 . The number of channels of the feature map of X is 832. The second SIE network only enhances the features of X 1 without changing the resolution of X 1 .
[0074] 3.3.4.4 The third SIE network receives X 2 . The fourth local SIE network and the fifth local SIE network of the third SIE network use the feature extraction method described in step 3.3.4.1.1 to perform feature extraction twice on X 2 . The third global SIE network uses the local feature enhancement method described in step 3.3.4.1.2 to extract and fuse the global information for feature enhancement of X after feature extraction twice 2 to obtain the feature map X with feature enhancement in the third stage 3 , and send X 3 to the fourth SIE network. The resolution of the feature map of X 3 is X 3 . The number of channels of the feature map of X is 832. The third SIE network only enhances the features of X 2 without changing the resolution of X 2 .
[0075] 3.3.4.5 The fourth SIE network receives X 3 . The sixth local SIE network and the seventh local SIE network of the fourth SIE network use the feature extraction method described in step 3.3.4.1.1 to perform feature extraction twice on X 3 . The fourth global SIE network uses the local feature enhancement method described in step 3.3.4.1.2 to extract and fuse the global information for feature enhancement of X after feature extraction twice 3 to obtain the feature map X with feature enhancement in the fourth stage 4 , and send X 4 to the upsampling module. The resolution of the feature map of X 4 is X 4 . The number of channels of the feature map of X is 832. The fourth SIE network only enhances the features of X 3 without changing the resolution of X 3 .
[0076] 3.3.4.6 The upsampling module receives X from the fourth SIE network 4 , and performs upsampling on X 4 using the upsampling method to obtain the upsampled feature map X U , and sends X U to the fifth SIE network. The method is as follows:
[0077] 3.3.4.6.1 The first linear layer of the upsampling module performs a linear transformation on X 4 in the channel domain, changing the number of channels to twice that of X 4 (i.e., the number of channels is 1664) to obtain the channel-transformed feature map.
[0078] 3.3.4.6.2 The first linear layer of the upsampling module changes the shape of the channel-transformed feature map so that the height and width are twice that of the channel-transformed feature map, i.e., the feature map resolution is and the number of channels is 1 / 4 of the channel-transformed feature map (i.e., the number of channels is 416) to obtain the shape-transformed feature map.
[0079] 3.3.4.6.3 The first linear layer of the upsampling module performs hierarchical normalization on the shape-transformed feature map to obtain the hierarchically normalized shape-transformed feature map, and sends the hierarchically normalized shape-transformed feature map to the second linear layer.
[0080] 3.3.4.6.4 The second linear layer of the upsampling module fuses the information in the channel domain of the hierarchically normalized shape-transformed feature map without changing the size of the channel dimension to obtain the upsampled feature map X U , and sends X U to the fifth SIE network.
[0081] 3.3.4.7 The fifth SIE network receives X from the upsampling module U . The eighth local SIE network of the fifth SIE network uses the local feature extraction method described in step 3.3.4.1.1 to perform feature extraction on X U . The fifth global SIE network uses the feature enhancement method described in step 3.3.4.1.2 to perform global information extraction and fusion feature enhancement on the X after feature extraction U to obtain the feature map X 5 after feature enhancement in the fifth stage. X 5 is also a matrix. Send X 5 to the global ocean area reduction module.
[0082] 3.3.5 The second Layer Normalization layer of the global ocean area reduction module receives X from the ocean feature aggregation module 5 , receives X from the ocean feature extraction module, and sends X5 Concatenate with X in the channel dimension to obtain a concatenated channel feature map, and normalize the concatenated channel feature map to obtain a normalized concatenated channel feature map, and send the normalized concatenated channel feature map to the two-dimensional transposed convolutional neural network of the global ocean area restoration module.
[0083] 3.3.4.6 The upsampling module receives X from the fourth SIE network 4 , and uses the upsampling method to upsample X 4 to obtain an upsampled feature map X U , and send X U to the fifth SIE network, the method is:
[0084] 3.3.4.6.1 The first linear layer of the upsampling module performs a linear transformation on X 4 in the channel domain, changing the number of channels to twice that of X 4 (i.e., the number of channels is 1664) to obtain a channel-transformed feature map.
[0085] 3.3.4.6.2 Change the shape of the channel-transformed feature map so that the height and width are twice that of the channel-transformed feature map, that is, the feature map resolution is and the number of channels is 1 / 4 of the channel-transformed feature map (i.e., the number of channels is 416) to obtain a shape-transformed feature map.
[0086] 3.3.4.6.3 The first linear layer sends the shape-transformed feature map to the second linear layer after hierarchical normalization.
[0087] 3.3.4.6.4 The second linear layer fuses the information in the channel domain without changing the size of the channel dimension to obtain the upsampled feature map X U , and send X U to the fifth SIE network.
[0088] 3.3.4.6 The fifth SIE network receives X from the upsampling module U , and the eighth local SIE network, the ninth local SIE network and the fifth global SIE network of the fifth SIE network extract and enhance the features of X 3 to obtain a feature map X after feature enhancement in the fifth stage 5 , X 5 is also 's matrix. Then send X 5 to the global ocean area restoration module.
[0089] 3.3.5 The second Layer Normalization layer of the global ocean area restoration module receives X from the ocean feature aggregation module 5 and, receives X from the ocean feature extraction module, and sends X5 The two feature maps, namely, and X, are concatenated in the channel dimension to obtain a channel concatenated feature map, and the channel concatenated feature map is normalized to obtain a normalized channel concatenated feature map, and the normalized channel concatenated feature map is sent to the two-dimensional transposed convolutional neural network of the global ocean area restoration module.
[0090] 3.3.6 The two-dimensional transposed convolutional neural network upsamples the resolution of the normalized channel concatenated feature map by 6 times. The height of the three-dimensional grid data is equal to the height H of I train The width is equal to the width W of I train The number of channels is the number of elements C to be predicted out (C out = 2 + number of layers × 4, including sea surface temperature, sea surface height, and temperature, salinity, eastward velocity of seawater, and northward velocity of seawater at each layer), and the three-dimensional grid data is the generated global ocean environmental element forecast result The data format is [H×W×C out , It represents the forecast results of sea surface temperature, sea surface height, and temperature, salinity, flow velocity (the flow velocity includes two components, eastward flow velocity of seawater and northward flow velocity of seawater) from layer 1 to layer 33 starting from the t-th day.
[0091] 3.3.7 Design the total loss function of the ocean environmental forecasting system As shown in formula (2):
[0092]
[0093] where represents the forecast value of the c-th (c ≤ C out ) element at the grid point with coordinates [i, j] (i ≤ H, j ≤ W) on the longitude-latitude grid. represents Y t+τ the true value of the c-th element at the grid point with coordinates [i, j] on the longitude-latitude grid.
[0094] 3.3.8 Let N b = N b + 1. If go to 3.3.2; if (|D M | represents the total number of three-dimensional grid data in the training set D M ), go to the fourth step.
[0095] Fourth step, use the validation set to verify the forecasting accuracy of the global ocean environmental forecasting system after the end of the current epoch, and retain the network weight parameters with the best performance as the network weight parameters of the global ocean environmental forecasting system. The method is:
[0096] 4.1 Input the data in the standardized validation set D obtained in step 2.5.3 into the ocean environmental forecasting system after the end of the current epoch; V into the ocean environmental forecasting system after the end of the current epoch;
[0097] 4.2 Let the variable v = 1, and let V be the total number of all three-dimensional grid data in the validation set;
[0098] 4.3 The ocean feature extraction module reads the v-th three-dimensional grid data D V from D, and uses the ocean feature extraction method described in step 3.3.3 to extract features from D v , and obtains the normalized feature map X v containing the ocean features of D v . Send X v to the ocean feature aggregation module and the global ocean area reduction module. X v is v a matrix with a resolution size of and 416 channels;
[0099] 4.4 The first SIE network in the ocean feature aggregation module receives X v , and uses the feature aggregation method based on local and global spatial information extraction described in step 3.3.4.1 to perform local spatial self-attention enhancement and global information fusion on X v , and obtains the feature map v with enhanced features in the first stage of D Send to the downsampling module;
[0100] 4.5 The downsampling module receives from the first SIE network and uses the patch combination splicing and channel domain linear transformation method described in step 3.3.4.2 to perform downsampling on , and obtains the downsampled feature map of the th Send to the second SIE network. is a matrix with a resolution size of
[0101] 4.6 The second SIE network receives from the downsampling module. The second local SIE network and the third local SIE network of the second SIE network use the local feature extraction method described in step 3.3.4.1.1 to perform 2 times of feature extraction on , and the second global SIE network uses the feature enhancement method described in step 3.3.4.1.2 to perform feature enhancement on Perform global information extraction and fusion feature enhancement to obtain the feature map after the second-stage feature enhancement Send To the third SIE network Is Of the matrix, with a resolution size of The number of channels is 832;
[0102] 4.7 The third SIE network receives the The fourth local SIE network and the fifth local SIE network of the third SIE network use the feature extraction method described in step 3.3.4.1.1 for Perform feature extraction 2 times, and the third global SIE network uses the local feature enhancement method described in step 3.3.4.1.2 for the After 2 times of feature extraction, perform global information extraction and fusion feature enhancement to obtain the feature map after the third-stage feature enhancement Send To the fourth SIE network Is Of the matrix, with a resolution size of The number of channels is 832;
[0103] 4.8 The fourth SIE network receives the The sixth local SIE network and the seventh local SIE network of the fourth SIE network use the feature extraction method described in step 3.3.4.1.1 for Perform feature extraction 2 times, and the fourth global SIE network uses the local feature enhancement method described in step 3.3.4.1.2 for the After 2 times of feature extraction, perform global information extraction and fusion feature enhancement to obtain the feature map after the fourth-stage feature enhancement Send To the upsampling module Is Of the matrix, with a resolution size of The number of channels is 832;
[0104] 4.9 The upsampling module receives the feature map Use the upsampling method described in step 3.3.4.6 for Perform upsampling to obtain D v Of the upsampled feature map Send To the fifth SIE network Is Of the matrix, with a resolution size of The number of channels is 416;
[0105] 4.10 The fifth SIE network receives the upsampled feature map The eighth local SIE network uses the local feature extraction method described in step 3.3.4.1.1 to perform feature extraction, and the fifth global SIE network uses the feature enhancement method described in step 3.3.4.1.2 to perform global information extraction and fusion feature enhancement on the v feature map after feature extraction, obtaining the feature map after the fifth-stage feature enhancement of D Send it to the global ocean area restoration module;
[0106] 4.11 The two-dimensional transposed convolutional neural network of the global ocean area restoration module receives and upsamples it by 6 times, outputting a three-dimensional grid data. The height of this three-dimensional grid data is equal to the height H of the three-dimensional grid data in the training set, the width is equal to the width W of the three-dimensional grid data in the training set, and the number of channels is the number of elements C to be predicted out . The output three-dimensional grid data is the forecast result of the global ocean environmental elements It takes the t-th day as the starting time, and is the forecast result of the sea surface temperature, sea surface height, and the temperature, salinity, and flow velocity (the flow velocity includes two components: the eastward flow velocity of seawater and the northward flow velocity of seawater) from the 1st to the 33rd layers on the (t + τ)-th day;
[0107] 4.12 The second Layer Normalization layer of the global ocean area restoration module uses the root mean square error (RMSE) to measure the forecast performance of the ocean environmental forecasting system for v with D as the input and the forecast performance for the τ-th day later. The calculation method of RMSE is shown in formula (3):
[0108]
[0109] where DM represents the inverse normalization operation, represents the forecast value of the c-th (c ≤ C out ) element at the grid point with coordinates [i, j] (i ≤ H, j ≤ W) on the longitude-latitude grid, represents the true value of the c-th element at the grid point with coordinates [i, j] on the longitude-latitude grid.
[0110] 4.13 Let v = v + 1. If v ≤ V, go to 4.3; if v > V, it means that the set of global ocean environmental forecast results after the end of the current epoch has been obtained, go to 4.14;
[0111] 4.14 The second Layer Normalization layer calculates the mean of the RMSE of the ocean environmental element prediction ensemble, obtains the average RMSE of the system on the validation set, and proceeds to 4.15;
[0112] 4.15 If epoch = 1, save the network weight parameters after the end of the current epoch and directly proceed to 4.16; if epoch = 3, set learning_rate = 5×10 -5 , if the average RMSE of the system on the validation set after the end of the current epoch is less than the average RMSE after the end of the previous epoch, then update the network weight parameters to the network weight parameters after the end of the current epoch and proceed to 4.16; if the average RMSE of the system on the validation set after the end of the current epoch is not less than the average RMSE after the end of the previous epoch, directly proceed to 4.16; if epoch = 2 or epoch > 3, if the average RMSE of the system on the validation set after the end of the current epoch is less than the average RMSE after the end of the previous epoch, then update the network weight parameters to the network weight parameters after the end of the current epoch and proceed to 4.16; if the average RMSE of the system on the validation set after the end of the current epoch is not less than the average RMSE after the end of the previous epoch, directly proceed to 4.16;
[0113] 4.16 Let epoch = epoch + 1. If epoch ≤ maxepoch, proceed to 3.3.2; if epoch > maxepoch, it means the training is completed, obtain the trained global ocean environmental prediction system, and proceed to the fifth step;
[0114] Fifth step, use the trained ocean environmental prediction system to predict the global ocean environmental data I input by the user. The method is as follows:
[0115] 5.1 According to the overall mean and standard deviation of the training set train data (since the overall mean and standard deviation of the data input by the user are unknown, the overall mean and standard deviation of the training set are used as a substitute), use the standardization method described in step 2.5.2 to standardize the global ocean environmental data I input by the user. I is a three-dimensional grid data, that is, a three-dimensional matrix in the format of [H×W×V], and obtain the standardized matrix I nor , and input I nor into the ocean feature extraction module;
[0116] 5.2 The ocean feature extraction module receives I nor , and uses the ocean feature extraction method described in step 3.3.3 to extract features from I nor and obtain the normalized feature map X containing the ocean features of I nor I , send X I to the ocean feature aggregation module.
[0117] 5.3 The first SIE network in the ocean feature aggregation module receives X I , and uses the feature aggregation method of local and global spatial information extraction described in step 3.3.4.1 to I perform local spatial self-attention enhancement and global information fusion on X to obtain I nor the feature map after the first-stage feature enhancement of Send to the downsampling module;
[0118] 5.4 The downsampling module receives and uses the patch combination splicing and channel domain linear transformation method described in 3.3.4.2 to perform downsampling on to obtain the downsampled feature map of Send to the second SIE network. is a matrix of, with a resolution size of and 832 channels;
[0119] 5.5 The second SIE network receives from the downsampling module and performs the second-stage feature extraction on : The second local SIE network and the third local SIE network of the second SIE network use the local feature extraction method described in step 3.3.4.1.1 to perform 2 times of feature extraction on, and the second global SIE network uses the feature enhancement method described in step 3.3.4.1.2 to perform global information extraction and fusion feature enhancement on the after 2 times of feature extraction to obtain I the feature map of the second-stage feature enhancement of nor Send to the third SIE network. is a matrix of with a resolution size of and 832 channels;
[0120] 5.6 The third SIE network receives The fourth local SIE network and the fifth local SIE network of the third SIE network use the feature extraction method described in step 3.3.4.1.1 to perform 2 times of feature extraction on, and the third global SIE network uses the local feature enhancement method described in step 3.3.4.1.2 to perform local feature enhancement on the Extract and fuse global information to enhance features, obtaining I nor The feature map after the third-stage feature enhancement of Send To the fourth SIE network. Is The matrix of The number of channels is 832;
[0121] 5.7 The fourth SIE network receives The sixth local SIE network and the seventh local SIE network of the fourth SIE network use the feature extraction method described in step 3.3.4.1.1 to Perform feature extraction twice, and the fourth global SIE network uses the local feature enhancement method described in step 3.3.4.1.2 for the After twice of feature extraction, perform global information extraction, fusion, and feature enhancement to obtain I nor The feature map after the fourth-stage feature enhancement of Send To the upsampling module. Is The matrix of The number of channels is 832;
[0122] 5.8 The upsampling module receives Use the upsampling method described in step 3.3.4.6 to Perform upsampling to obtain The upsampled feature map of Send To the fifth SIE network. Is The matrix of The number of channels is 416; Use the upsampling method described in step 3.3.4.6 to Perform upsampling to obtain D v The upsampled feature map of
[0123] 5.9 The fifth SIE network receives The eighth local SIE network uses the local feature extraction method described in step 3.3.4.1.1 to Perform feature extraction, and the fifth global SIE network uses the feature enhancement method described in step 3.3.4.1.2 for the After feature extraction, perform global information extraction, fusion, and feature enhancement, and the fifth global SIE network extracts global information to obtain I nor The feature map after the fifth-stage feature enhancement of Send To the global ocean area restoration module;
[0124] 5.10 The two-dimensional transposed convolutional neural network of the global ocean area reduction module receives For Upsample by a factor of 6 and output a three-dimensional grid data. The height of the output three-dimensional grid data is equal to H, the width is equal to W, and the number of channels is the number of elements C to be predicted out (C out is the number of elements that the user needs to forecast). The output three-dimensional grid data is the global ocean environment forecast result of the global ocean environment data I input by the user It takes the t-th day as the starting time and is the forecast result of the sea surface temperature, sea surface height, and temperature, salinity, and flow velocity (the flow velocity includes two components: the eastward flow velocity of seawater and the northward flow velocity of seawater) of the 1st to 33rd layers on the (t + τ)-th day;
[0125] Step 6, end.
[0126] The following beneficial effects can be achieved by adopting the present invention:
[0127] The present invention proposes a global ocean environment forecasting method based on multi-level feature aggregation. The present invention adopts a local spatial information extraction network and a global spatial information extraction network to alleviate the problems of insufficient feature representation ability and high computational cost, and improve the forecasting accuracy of high-resolution data. The following effects can be achieved by adopting the present invention:
[0128] 1. The present invention constructs a global ocean environment forecasting system integrating an ocean feature extraction module, an ocean feature aggregation module, and a global ocean area reduction module. On the basis of ensuring that the forecasting method can quickly meet the real-time requirement, by using the local spatial self-attention enhancement and global information fusion ability of the ocean feature extraction module, the aggregation method and network structure suitable for ocean features are designed, and a large improvement in forecasting accuracy is achieved. By using the evaluation process and evaluation method of IVTT to conduct experiments on the present invention, the forecasting accuracy of the present invention is greatly improved compared with the numerical forecasting method described in the background art.
[0129] 2. The ocean feature aggregation module of the present invention uses the SIE network to enhance the feature representation ability multiple times, and uses the hierarchical design of the overall model to enhance the multi-scale representation ability of features; the local SIE network of the present invention uses a masked window multi-head self-attention layer to make the forecasting task pay more attention to ocean water bodies and realize local spatial information aggregation; the SIE network uses a cross-attention mechanism and an MLP-mixer to realize the fusion of all spatial information, resulting in a significant improvement in forecasting accuracy.
[0130] 3. In the third step of the present invention, the ocean environmental forecasting system is trained using the gradient backpropagation method. In the fourth step, a validation set is used to verify the forecasting accuracy of the ocean environmental forecasting system after the end of the current epoch, and the network weight parameters with the best performance are retained as the network weight parameters of the global ocean environmental forecasting system, so that the trained global ocean environmental forecasting system is the most suitable system for prediction, and the forecasting accuracy is greatly improved in the fifth step. BRIEF DESCRIPTION OF THE DRAWINGS
[0131] Figure 1 It is a logical structure diagram of the ocean environmental forecasting system constructed in the first step of the present invention.
[0132] Figure 2 It is the overall flowchart of the present invention.
[0133] Figure 3 It is an example diagram of the forecasting result during the test of the effect of the present invention (taking the forecasting results of the sea surface temperature on the seventh day, the salinity of the first layer, and the seawater velocity of the first layer as examples). Figure 3 (a) shows a rendering of the forecasting result of the sea surface temperature on January 8, 2019, obtained by taking the sea surface temperature data on January 1, 2019 as the input of the global ocean environmental forecasting system; Figure 3 (b) shows a rendering of the forecasting result of the salinity of the first layer on January 8, 2019, obtained by taking the salinity data of the first layer on January 1, 2019 as the input of the global ocean environmental forecasting system; Figure 3 (c) shows a rendering of the forecasting result of the eastward seawater velocity of the first layer on January 8, 2019, obtained by taking the eastward seawater velocity data of the first layer on January 1, 2019 as the input of the global ocean environmental forecasting system; Figure 3 (d) shows a rendering of the forecasting result of the northward seawater velocity of the first layer on January 8, 2019, obtained by taking the northward seawater velocity data of the first layer on January 1, 2019 as the input of the global ocean environmental forecasting system. DETAILED DESCRIPTION OF THE INVENTION
[0134] The following is a description of specific examples of the present invention with reference to the accompanying drawings. As Figure 2 shown, the present invention includes the following steps:
[0135] In the first step, a global ocean environmental forecasting system based on multi-level feature aggregation is constructed. As Figure 1 shown, the ocean environmental forecasting system consists of an ocean feature extraction module, an ocean feature aggregation module, and a global ocean area reduction module.
[0136] The ocean feature extraction module is connected to the ocean feature aggregation module and the global ocean area restoration module. The ocean feature extraction module receives the global ocean environment data input by the user (i.e., three-dimensional grid data containing various layer elements), extracts ocean features from the global ocean environment data, and sends the feature maps of the ocean features to the ocean feature aggregation module and the global ocean area restoration module. The ocean feature extraction module consists of a two-dimensional convolutional neural network (see the paper "Dosovitskiy A, Beyer L, Kolesnikov A, et al. An image is worth 16x16 words: Transformers for image recognition at scale [J]. arXiv preprint arXiv:2010.11929, 2020." by Dosovitskiy A, Beyer L, etc.: Transformers for large-scale image recognition) and a first LayerNormalization layer (see the paper "Ba J L, Kiros J R, Hinton G E. Layer normalization [J]. arXiv preprint arXiv:1607.06450, 2016." by Lin T Ba J L, Kiros J R, etc.: Hierarchical normalization). The two-dimensional convolutional neural network divides the three-dimensional grid data into patches, each patch being 6×6 pixels in size, extracts the ocean features of the global ocean environment data, and sends the feature maps of the ocean features to the first Layer Normalization layer. The stride and the convolutional kernel size of the two-dimensional convolutional neural network are both 6. The first LayerNormalization layer normalizes the feature maps of the ocean features to ensure the stability of training, and sends the normalized feature maps of the ocean features to the ocean feature aggregation module.
[0137] The ocean feature aggregation module is connected to the ocean feature extraction module and the global ocean area restoration module. The ocean feature aggregation module consists of five Spatial Information Extraction (SIE) networks with non-shared weights (these five networks are respectively denoted as the first, second, third, fourth, and fifth SIE networks), an upsampling module, and a downsampling module. The first SIE network and the fifth SIE network have the same structure, both consisting of a local SIE network and a global SIE network; the second SIE network to the fourth SIE network have the same structure, all consisting of two consecutive local SIE networks and a global SIE network. The first SIE network is composed of the first local SIE network and the first global SIE network. The first local SIE network is composed of a window multi-head self-attention layer and a Multilayer Perceptron (MLP). The window multi-head self-attention layer receives the feature map of the normalized ocean features from the ocean feature extraction module, performs hierarchical normalization on the feature map of the normalized ocean features, divides the hierarchically normalized feature map into multiple 7×7 windows, and uses the window multi-head self-attention mechanism (see the literature "Liu Z, Lin Y, Cao Y, et al. Swin transformer: Hierarchical vision transformer using shifted windows [C] / / Proceedings of the IEEE / CVF International conference on computer vision. 2021: 10012-10022." by Liu Z, Lin Y, etc.: Hierarchical vision transformer using shifted windows) and residual connection operations to obtain the attention mechanism-enhanced feature map, and sends the attention mechanism-enhanced feature map to the multilayer perceptron. The multilayer perceptron consists of two linear layers, receives the attention mechanism-enhanced feature map from the window multi-head self-attention layer, performs hierarchical normalization, linear projection, and residual connection operations on it, obtains the feature map that fuses the local spatial information within the window, and sends the feature map that fuses the local spatial information within the window to the first global SIE network. The first global SIE network is composed of a feature grouping network, a group feature fusion network, and a group feature propagation network. The feature grouping network of the first global SIE network receives the feature map that fuses the local spatial information within the window from the first local SIE network, uses the cross-attention mechanism to group and aggregate the features in the feature map that fuses the local spatial information within the window to obtain group features, and sends the group features to the group feature fusion network.The main part of the group feature fusion network is composed of MLP-mixer (see the paper by Tolstikhin IO, Houlsby N, Kolesnikov A, et al., "Mlp-mixer: An all-mlp architecture for vision" in "Advances in neural information processing systems, 2021, 34: 24261-24272."). MLP-mixer can achieve a good balance between model performance and computational overhead. The group feature fusion network receives group features from the feature grouping network, exchanges group feature information and updates the group features, so that each group feature can contain the current group information and to a certain extent reflect the global information, obtaining the updated group features, and sending the updated group features to the group feature propagation network. The group feature propagation network uses the cross-attention mechanism to propagate the updated group features to all features in the feature map that fuses local spatial information within the fusion window, and uses the depthwise separable convolutional neural network to perform convolutional feature extraction on the feature map of local spatial information within the window, obtaining the feature map after the first-stage feature enhancement, and sending the feature map after the first-stage feature enhancement to the downsampling module. The local SIE network structures in the second to fifth SIE networks are the same as the first local SIE network structure, and the global SIE network structures in the second to fifth SIE networks are the same as the first global SIE network.
[0138] The downsampling module is connected to the first SIE network and the second SIE network and consists of a linear layer. The downsampling module receives the feature map after the first-stage feature enhancement from the first SIE network, combines every 2×2 adjacent patches in the feature map after the first-stage feature enhancement, and performs hierarchical normalization operations after splicing in the channel dimension to obtain a channel-spliced feature map with 4 times the number of channels of the feature map after the first-stage feature enhancement. The downsampling module performs a linear transformation in the channel dimension of the channel-spliced feature map, halves the number of channels of the channel-spliced feature map, and obtains a downsampled feature map. The height and width of the downsampled feature map are half of those of the feature map after the first-stage feature enhancement, and the number of channels is 2 times that of the feature map after the first-stage feature enhancement, that is, the downsampling ratio becomes twice that of the first SIE network. Then, the downsampled feature map is sent to the second SIE network. Two consecutive local SIE networks (i.e., the second local SIE network and the third local SIE network) and a global SIE network (i.e., the second global SIE network) of the second SIE network enhance the downsampled feature map to obtain a feature map with second-stage feature enhancement and send the feature map with second-stage feature enhancement to the third SIE network; two consecutive local SIE networks (i.e., the fourth local SIE network and the fifth local SIE network) and a global SIE network (i.e., the third global SIE network) of the third SIE network enhance the feature map with second-stage feature enhancement to obtain a feature map with third-stage feature enhancement and send the feature map with third-stage feature enhancement to the fourth SIE network; two consecutive local SIE networks (i.e., the sixth local SIE network and the seventh local SIE network) and a global SIE network (i.e., the fourth global SIE network) of the fourth SIE network enhance the feature map with third-stage feature enhancement to obtain a feature map with fourth-stage feature enhancement and send the feature map with fourth-stage feature enhancement to the upsampling module.
[0139] The upsampling module is connected to the fourth SIE network and the fifth SIE network and consists of two linear layers. The upsampling module receives the feature map with fourth-stage feature enhancement from the fourth SIE network. The first linear layer of the upsampling module performs a linear transformation in the channel dimension of the feature map with fourth-stage feature enhancement to obtain a channel-transformed feature map with 2 times the number of channels of the feature map with fourth-stage feature enhancement; change the shape of the channel-transformed feature map so that the height and width are 2 times those of the channel-transformed feature map and the number of channels is half of that of the channel-transformed feature map to obtain a shape-transformed feature map and send the shape-transformed feature map to the second linear layer of the upsampling module. The second linear layer of the upsampling module fuses the channel-domain information of the shape-transformed feature map without changing the size of the channel dimension to obtain an upsampled feature map and sends the upsampled feature map to the fifth SIE network. The height and width of the upsampled feature map are 2 times those of the feature map with fourth-stage feature enhancement, and the number of channels is half of that of the feature map with fourth-stage feature enhancement.
[0140] The fifth SIE network performs fifth-stage feature enhancement on the upsampled feature map to obtain a feature map with fifth-stage feature enhancement, and sends the feature map with fifth-stage feature enhancement to the global ocean area restoration module.
[0141] The local and global SIE networks in the first to fifth SIE networks enable each SIE network to have local and global spatial perception capabilities and enhance the feature representation ability. To better capture the dynamic changes of ocean water bodies and further reduce the computational load, a masking mechanism is added to the window multi-head self-attention layer in the local SIE network and the feature grouping network in the global SIE network, so that the attention score of the land area is 0. Hierarchical design and multi-scale feature extraction of ocean environmental elements are achieved through the downsampling module and the upsampling module, and multi-scale information of the global ocean can be captured. The ocean feature aggregation module extracts multi-scale features of global and local ocean area information through five stages and sends the feature map with fifth-stage feature enhancement to the global ocean area restoration module to improve the prediction accuracy of the global ocean environmental forecasting system.
[0142] The global ocean area restoration module consists of a second Layer Normalization layer and a two-dimensional transposed convolutional neural network, which is connected to the fifth SIE network of the ocean feature extraction module and the ocean feature aggregation module. It receives the feature map with fifth-stage feature enhancement from the ocean feature aggregation module and the feature map of the normalized ocean features from the ocean feature extraction module. The second Layer Normalization layer of the global ocean area restoration module first concatenates the feature map with fifth-stage feature enhancement and the feature of the normalized ocean features in the channel dimension to obtain a channel concatenated feature map, then normalizes the channel concatenated feature map, and sends the normalized channel concatenated feature map to the two-dimensional transposed convolutional neural network layer. The stride and convolutional kernel size of the two-dimensional transposed convolutional neural network layer are both 6. The normalized channel concatenated feature map is upsampled 6 times using a 6×6 transposed convolution with a stride of 6 to obtain the predicted three-dimensional grid data. The height and width of the predicted three-dimensional grid data are the same as those of the user-input global ocean environmental data, and the channel dimension is the total number of target elements (including sea surface height, sea surface temperature, temperature, salinity, and flow velocity from layer 1 to layer 33). The predicted three-dimensional grid data is the global ocean environmental forecasting result.
[0143] Step 2: Construct the training set, validation set, and test set. The method is as follows:
[0144] 2.1 Download GLORYS12 global ocean reanalysis data, ERA5 wind field reanalysis data, and GHR sea surface temperature (SST) satellite data, and perform quality control on the data. The method is as follows:
[0145] 2.1.1 Use GLORYS12 global ocean reanalysis data (see the paper by Jean-Michel L, Eric G, Romain BB, et al., "The Copernicus global 1 / 12° oceanic and sea ice GLORYS12 reanalysis" [J]. Frontiers in Earth Science, 2021, 9: 698876.), ERA5 wind field reanalysis data (see the paper by Hersbach H, Bell B, Berrisford P, et al., "The ERA5 global reanalysis" [J]. Quarterly Journal of the Royal Meteorological Society, 2020, 146(730): 1999 - 2049.) and GHR's sea surface temperature (SST) data (see the paper by Martin M, Dash P, Ignatov A, et al., "Group for High Resolution Sea Surface temperature (GHRSST) analysis fields inter-comparisons. Part 1: A GHRSST multi-product ensemble (GMPE)" [J]. Deep Sea Research Part II: Topical Studies in Oceanography, 2012, 77: 21 - 30.). Collect data from 1993 to 2020 as a dataset for ocean environmental forecasting. The GLORYS12 global ocean reanalysis data has a spatial resolution of 1 / 12°, with a spatial span from -180° to 180° in the longitude direction and from -80° to 90° in the latitude direction. The time resolution is 1 day, including reanalysis data of multiple elements (temperature, salinity, flow velocity, sea level height. The flow velocity includes two components: the eastward flow velocity of seawater and the northward flow velocity of seawater), and each element contains 50 layers of data, with each layer corresponding to a specific depth vertically downward from the sea surface.The present invention only forecasts the seawater temperature, salinity, flow velocity, and sea level height. Therefore, first, five elements, namely thetao (temperature), so (salinity), uo (eastward seawater flow velocity), vo (northward seawater flow velocity), and zos (sea surface height), in the GLORYS12 data are selected. For temperature, salinity, eastward seawater flow velocity, and northward seawater flow velocity, data from the 1st to the 33rd layers are extracted; for the sea surface height, data from the 1st layer is extracted. The sea surface wind has a significant impact on the marine environment. Therefore, the ERA5 reanalysis data is also used as a component of the dataset for marine environment forecasting. The spatial resolution of the ERA5 wind field reanalysis data is 1 / 4 degree, with a spatial span from -180 degrees to 180 degrees in the longitude direction and from -90 degrees to 90 degrees in the latitude direction. The time resolution is 1 hour, including two single-layer elements, u10 (eastward wind speed at 10 meters) and v10 (northward wind speed at 10 meters). The spatial resolution of the sea surface temperature data of GHR is 1 / 20 degree, with a spatial span from -180 degrees to 180 degrees in the longitude direction and from -90 degrees to 90 degrees in the latitude direction. The time resolution is 1 day, and it only contains satellite data of the sea surface temperature (analysed_sst).
[0146] 2.1.2 The downloaded GLORYS12 global ocean reanalysis data, ERA5 wind field reanalysis data, and GHR's sea surface temperature (SST) satellite data are all in nc format. Before formally processing the data, it is necessary to first perform quality control on the data. Using the xarray library in python, the above data is read starting from January 1, 1993. If the data is missing, damaged, or some elements in the data are missing, it is written to the log. Then, the problematic "dirty" data files recorded in the log are downloaded again to ensure that there are no such problems with the data before it is received by the global ocean environment forecasting system based on multi-level feature aggregation;
[0147] 2.2 The spatial resolutions and time resolutions of the three datasets, namely GLORYS12 global ocean reanalysis data, ERA5 wind field reanalysis data, and GHR sea surface temperature satellite data, are different. Let the number of days from 1993 to 2020 be D, and let the original datasets of the three datasets, GLORYS12 global ocean reanalysis data, ERA5 wind field reanalysis data, and GHR sea surface temperature satellite data, be M, E, and G. Next, the time and spatial resolutions of M, E, and G are aligned. The method is as follows:
[0148] 2.2.1 Align the time resolution of the ERA5 wind field reanalysis data with the other two datasets. The method is as follows:
[0149] 2.2.1.1 Let the variable d = 1, and initialize the ERA5 wind field reanalysis dataset E1 with aligned time resolution as empty;
[0150] 2.2.1.2 Extract the data at specific time points from the ERA5 wind field reanalysis data on the d-th day in E, that is, select the data at 00, 06, 12, and 18 hours of a day, and take the mean of the sea surface wind field data at these four times as the daily average data of the wind field on the d-th day, and put the daily average data of the wind field on the d-th day into E1;
[0151] 2.2.1.3 If d ≤ D, let d = d + 1, and go to 2.2.1.2; if d > D, obtain the ERA5 wind field reanalysis data set E1 with aligned time resolution, and go to 2.2.2;
[0152] 2.2.2 Align the spatial resolutions of the two data sets M and E1 with G. The method is as follows:
[0153] 2.2.2.1 Let d = 1, and initialize the reanalysis data set E2 of ERA5 with aligned spatial resolution and the GHR sea surface temperature data set G1 to be empty;
[0154] 2.2.2.2 Perform bilinear interpolation on the daily average wind field data on the d-th day in E1 and the GHR sea surface temperature data in G respectively to obtain the wind field data on the d-th day with a spatial resolution of 1 / 12 degree, and put it into E2; and obtain the sea surface temperature data on the d-th day with a spatial resolution of 1 / 12 degree, and put it into G1;
[0155] 2.2.2.3 If d ≤ D, let d = d + 1, and go to 2.2.2.2; if d > D, obtain the ERA5 wind field reanalysis data set E2 with a spatial resolution of 1 / 12 degree and the GHR sea surface temperature data set G1, and go to 2.2.3;
[0156] 2.2.3 Align the two data sets E2 and G1 with the spatial range of M. The method is as follows:
[0157] 2.2.3.1 Let d = 1, and initialize the set of the ERA5 wind field data E3 and the GHR sea surface temperature data G2 after aligning the spatial range to be empty;
[0158] 2.2.3.2 Select the part of the daily average wind field data on the d-th day in E2 and G1 and the GHR sea surface temperature data from -180 degrees to 180 degrees in the longitude direction and from -80 degrees to 90 degrees in the latitude direction, and put the daily average wind field data on the d-th day in this spatial range into E3; put the GHR sea surface temperature data on the d-th day in this spatial range into G2;
[0159] 2.2.3.3 If d ≤ D, let d = d + 1, and go to 2.2.3.2; if d > D, obtain the ERA5 wind field reanalysis data set E3 with aligned spatial range and the GHR sea surface temperature data set G2, and go to 2.4;
[0160] 2.4 After aligning the temporal and spatial resolutions of the data, the data in E3, G2, and M are still in nc format. To facilitate the data reading of the ocean environmental forecasting system, the nc format files need to be converted to npy format. The method is as follows:
[0161] 2.4.1 Let the variable d = 1, and initialize the dataset S in npy format to be empty;
[0162] 2.4.2 Read the GLORYS12 global ocean reanalysis data, ERA5 wind field reanalysis data, and GHR sea surface temperature data for the d-th day from M, E3, and G2 respectively. Sequentially extract the data of the 1st to 33rd layers of the thetao, so, uo, vo, and zos elements in the GLORYS12 data, the u10 and v10 elements in the ERA5 wind field data, and the analysed_sst element in GHR; then splice the above selected element data, save it as an npy file, and then put it into S;
[0163] 2.4.3 If d ≤ D, let d = d + 1, and go to 2.4.2; if d > D, the generation of npy data is completed, and a new dataset S is obtained, then go to 2.5.
[0164] 2.5 Divide S into a training set, a validation set, and a test set, and calculate the mean and standard deviation of the training set data, and normalize the data. The method is as follows:
[0165] 2.5.1 Use the mv command to move the data from 1993 to 2017 in S to the training set train, move the data in 2018 to the validation set val, and move the data from 2019 to 2020 to the test set test;
[0166] 2.5.2 According to the overall mean and standard deviation of the training set train data, use the standardization method to standardize the data in the training set train to obtain the standardized training set D M . The standardization method is as follows:
[0167] 2.5.2.1 Let the variable d = 1, take the data of the d-th day in train as X d , and initialize the standardized training set D M to be empty;
[0168] 2.5.2.2 Standardize X d as shown in the specific process of formula (1):
[0169]
[0170] where μ is the overall mean of the training set train data, and σ is the overall standard deviation. The standardized Xd Put it into D M ;
[0171] 2.5.2.3 If d ≤ D, let d = d + 1, and go to 2.5.2.2; if d > D, the standardization of the training set train data is completed, and go to 2.5.3;
[0172] 2.5.3 According to the overall mean and standard deviation of the training set data (for the training process of the model, the data of the validation set and the test set are unknown, so only the overall mean and standard deviation of the training set data can be used), use the standardization method described in step 2.5.2 to standardize the data in the validation set val, and obtain the standardized validation set D V ;
[0173] 2.5.4 According to the overall mean and standard deviation of the training set data, standardize the data in the test set test, and obtain the standardized test set D T .
[0174] The third step is to use the gradient backpropagation method to train the ocean environmental forecasting system constructed in the first step, and obtain the network weight parameters with the best forecasting effect on the validation set D v . The method is as follows:
[0175] 3.1 Initialize the network weight parameters of each module in the ocean environmental forecasting system. Use a normal distribution with a mean of 0 and a variance of 0.01 to initialize the network weight parameters of the ocean feature extraction module, the ocean feature aggregation module, and the global ocean area restoration module.
[0176] 3.2 Set the training parameters of the ocean environmental forecasting system. Set the warm-up learning rate to 5×10 -8 , the warm-up training step size to 3, and the learning rate learning_rate to 5×10 -5 . Select AdamW (see the paper "Loshchilov I, Hutter F. Decoupled weight decay regularization [J]. arXiv preprint arXiv:1711.05101, 2017." by Loschchilov I, Hutter F: Decoupled Weight Decay Regularization) as the model training optimizer. The hyperparameters of this model training optimizer are β1 = 0.9, β2 = 0.95, and "weight decay" = 1×10 -3 . The batch size (mini_batch_size) of network training is 1. The maximum training step size (maxepoch) is 50.
[0177] 3.3 Train the ocean environment forecasting system by taking the square of the difference between the forecasting result output by the ocean environment forecasting system during one training and the true value as the loss value, and using gradient backpropagation to update the network weight parameters until the loss value reaches the threshold or the training step reaches maxepoch. In each round of training, if the forecasting accuracy of the current system is the best on the validation set D v save the network weight parameters of the current round. The method is as follows:
[0178] 3.3.1 Let the training step epoch = 1. One pass through all the data in the training set is one epoch, and initialize the batch number N b = 1;
[0179] 3.3.2 The ocean feature extraction module reads the N M th batch from D b , a total of B (0 ≤ B ≤ 16 and B is a positive integer) three-dimensional grid data, and records these B three-dimensional grid data in matrix form I train , and I train contains B matrices of H×W×V. Here, H represents the height of the input grid data, W represents the width of the input grid data, and V represents the channel dimension (V = 3 + number of layers × 4, that is, 3 sea surface elements, as well as the temperature, salinity, eastward velocity of seawater, and northward velocity of seawater at each layer).
[0180] 3.3.3 The ocean feature extraction module uses the ocean feature extraction method to extract features from I train , and sends the normalized feature map X containing the ocean features of I train to the ocean feature aggregation module and the global ocean area reduction module. The method is as follows:
[0181] 3.3.3.1 The two-dimensional convolutional neural network in the ocean feature extraction module extracts the ocean features of I train , and obtains the feature map containing the ocean features. The method is: use a 6×6 convolution with a stride of 6 to perform patch division on the B three-dimensional grid data of I train , the size of each patch is 6×6, and the representation dimension of each patch is 416, to obtain the feature map containing the ocean features. Then send the feature map containing the ocean features to the first Layer Normalization layer;
[0182] 3.3.3.2 The first Layer Normalization layer normalizes the feature map containing the ocean features at the level of each three-dimensional grid data, and obtains the one containing I trainThe normalized feature map X of ocean features. The normalization operation helps improve the training effect of the network. The first Layer Normalization layer sends X to the ocean feature aggregation module and the global ocean area restoration module. The resolution of X is The number of channels of X is 416;
[0183] 3.3.4 The ocean feature aggregation module receives X from the ocean feature extraction module, generates the feature enhanced feature map X5 of the fifth stage, and sends X5 to the global ocean area restoration module. The method is as follows:
[0184] 3.3.4.1 The first SIE network receives X from the ocean feature extraction module, and uses a feature aggregation method based on local and global spatial information extraction to perform local spatial self-attention enhancement and global information fusion on X, obtaining the first stage feature enhanced feature map X of X 1 . The method is as follows:
[0185] 3.3.4.1.1 The first local SIE network of the first SIE network uses a local feature extraction method to extract features from X. The method is: the window multi-head self-attention layer of the first local SIE network performs hierarchical normalization on X, divides X into multiple 7×7 windows, that is, each window contains 7×7 patches. Then the window multi-head self-attention layer of the first local SIE network parallelly uses the window multi-head self-attention mechanism to perform local spatial self-attention enhancement on X within each window, and then performs a residual connection to obtain the feature map after the residual connection During the calculation process of the window multi-head self-attention mechanism, the mask mechanism is used to set the attention weights of the features corresponding to the land areas in X to 0. Therefore, the land part of the area will not affect the ocean area during the calculation process. Then the window multi-head self-attention layer of the first local SIE network sends the feature map to the multi-layer perceptron of the first local SIE network. The multi-layer perceptron of the first local SIE network first performs hierarchical normalization on , and then uses a linear mapping to perform fusion on in the channel dimension, and then performs a residual connection operation to obtain the feature map that fuses the local spatial information within the window and sends to the first global SIE module.
[0186] 3.3.4.1.2 The first global SIE network of the first SIE network receives from the first local SIE network, and uses a feature enhancement method to perform extraction and fusion of global information on to obtain the first stage feature enhanced feature map X that fuses local and global spatial information 1 , and the method is:
[0187] 3.3.4.1.2.1 The first feature grouping network randomly initializes a set of group features G, where G contains N features, and the dimension of each feature is d (here it is 416). The first feature grouping network first uses the multi-head cross-attention mechanism to aggregate and update the feature groups in, obtaining the updated group features G′, where each feature in G′ aggregates the information of a cluster of patches with similar semantics. During the calculation of the multi-head cross-attention mechanism, a masking mechanism is used to mask the features corresponding to the land part in, so that the land part does not participate in the calculation of the cross-attention mechanism. Then the first feature grouping network sends the group features G′ to the first group feature fusion network.
[0188] 3.3.4.1.2.2 The first group feature fusion network receives G′ from the first feature grouping network, exchanges information between each feature in G′, and then fuses and updates the G′ group feature information to obtain the updated group features The first MLP layer of the first group feature fusion network fuses the group features that obtain the fused spatial information in the G′ spatial domain The second MLP layer performs fusion on in the channel domain to obtain the updated group features The first group feature fusion network sends the updated group features to the first group feature propagation network.
[0189] 3.3.4.1.2.3 The first group feature propagation network receives from the first group feature fusion network and receives from the first local SIE network and adopts the multi-head cross-attention mechanism to use the features in as the query vectors, use the group features as the key vectors and value vectors, and according to the different features in assign different weights to each feature in the group features , and then perform weighted summation on the features in according to the weights to obtain the global spatial feature U, and send U to the first group feature propagation network.
[0190] 3.3.4.1.2.4 The feed-forward neural network of the first group feature propagation network concatenates U with the features in to obtain the concatenated feature map X 1′ , and performs a linear transformation on X 1′ , and then obtains the feature map X after residual connection through residual connection 1″ ; the depthwise separable convolutional neural network performs operations on X 1″Perform convolutional feature extraction to obtain the feature map X of the first-stage feature enhancement that fuses local and global spatial information 1 X 1 has a resolution of X 1 has 416 channels. The first group of feature propagation networks of the first SIE network sends X 1 to the downsampling module.
[0191] 3.3.4.2 The downsampling module receives X 1 from the first group of feature propagation networks of the first SIE network, and performs downsampling on X 1 using the patch combination splicing and channel-domain linear transformation method to obtain the downsampled feature map X D , and sends X D to the second SIE network. The method is as follows:
[0192] 3.3.4.2.1 Combine each 2×2 adjacent patch in X 1 , and perform hierarchical normalization operations after splicing these 4 patches in the channel domain to obtain a channel-spliced feature map with 4 times the number of channels of X 1 (the channel dimension of the spliced feature map is 4×416);
[0193] 3.3.4.2.2 Perform a linear transformation on the channel dimension of the channel-spliced feature map to halve the number of channels of the channel-spliced feature map (i.e., 2×416) to obtain the downsampled feature map X D . X D has a feature map resolution of X D has a feature map channel number of 832. The downsampling ratio of the downsampled feature map X D is twice that of X 1 , thus forming a hierarchical structure design to achieve multi-scale feature extraction and also reduce the computational overhead.
[0194] 3.3.4.2.3 The downsampling module sends X D to the second SIE network.
[0195] 3.3.4.3 The second SIE network receives the downsampled feature map X D sent by the downsampling module. The second local SIE network and the third local SIE network perform 2 times of feature extraction on X D using the local feature extraction method described in step 3.3.4.1.1. The second global SIE network performs the extraction and fusion of global information and feature enhancement on X D after 2 times of feature extraction using the feature enhancement method described in step 3.3.4.1.2 to obtain the feature map X 2, send X 2 to the third SIE network. X 2 The resolution of the feature map of X 2 is 832 for the number of channels of the feature map of X 1 , and the second SIE network only performs feature enhancement on X 1 without changing the resolution of X
[0196] 3.3.4.4 The third SIE network receives X 2 , and the fourth local SIE network and the fifth local SIE network of the third SIE network use the feature extraction method described in step 3.3.4.1.1 to perform 2 times of feature extraction on X 2 , and the third global SIE network uses the local feature enhancement method described in step 3.3.4.1.2 to perform global information extraction and fusion feature enhancement on X 2 after 2 times of feature extraction, and obtains the feature map X 3 of the third-stage feature enhancement, and sends X 3 to the fourth SIE network. X 3 The resolution of the feature map of X 3 is 832 for the number of channels of the feature map of X 2 , and the third SIE network only performs feature enhancement on X 2 without changing the resolution of X
[0197] 3.3.4.5 The fourth SIE network receives X 3 , and the sixth local SIE network and the seventh local SIE network of the fourth SIE network use the feature extraction method described in step 3.3.4.1.1 to perform 2 times of feature extraction on X 3 , and the fourth global SIE network uses the local feature enhancement method described in step 3.3.4.1.2 to perform global information extraction and fusion feature enhancement on X 3 after 2 times of feature extraction, and obtains the feature map X 4 of the fourth-stage feature enhancement, and sends X 4 to the upsampling module. X 4 The resolution of the feature map of X 4 is 832 for the number of channels of the feature map of X 3 , and the fourth SIE network only performs feature enhancement on X 3 without changing the resolution of X
[0198] 3.3.4.6 The upsampling module receives X 4 from the fourth SIE network, and uses the upsampling method to perform upsampling on X 4 , and obtains the upsampled feature map X U , and sends X USend it to the fifth SIE network by:
[0199] 3.3.4.6.1 The first linear layer of the upsampling module performs a linear transformation in the channel domain of X 4 to change the number of channels to twice that of X 4 (i.e., the number of channels is 1664), obtaining a channel-transformed feature map.
[0200] 3.3.4.6.2 The first linear layer of the upsampling module changes the shape of the channel-transformed feature map such that the height and width are twice that of the channel-transformed feature map, i.e., the feature map resolution is and the number of channels is 1 / 4 of the channel-transformed feature map (i.e., the number of channels is 416), obtaining a shape-transformed feature map.
[0201] 3.3.4.6.3 The first linear layer of the upsampling module performs hierarchical normalization on the shape-transformed feature map to obtain a hierarchically normalized shape-transformed feature map, and sends the hierarchically normalized shape-transformed feature map to the second linear layer.
[0202] 3.3.4.6.4 The second linear layer of the upsampling module fuses the information in the channel domain of the hierarchically normalized shape-transformed feature map without changing the size of the channel dimension, obtaining an upsampled feature map X U , and sends X U to the fifth SIE network.
[0203] 3.3.4.7 The fifth SIE network receives X from the upsampling module U , and the eighth local SIE network of the fifth SIE network uses the local feature extraction method described in step 3.3.4.1.1 to perform feature extraction on X U , and the fifth global SIE network uses the feature enhancement method described in step 3.3.4.1.2 to perform global information extraction and fusion feature enhancement on the X after feature extraction U to obtain a feature map X after feature enhancement in the fifth stage 5 , and X 5 is also a matrix. Send X 5 to the global ocean area reduction module.
[0204] 3.3.5 The second Layer Normalization layer of the global ocean area reduction module receives X from the ocean feature aggregation module 5 , receives X from the ocean feature extraction module, and concatenates X 5 and X in the channel dimension to obtain a channel-concatenated feature map, and normalizes the channel-concatenated feature map to obtain a normalized channel-concatenated feature map, and sends the normalized channel-concatenated feature map to the two-dimensional transposed convolutional neural network of the global ocean area reduction module.
[0205] 3.3.4.6 The upsampling module receives X from the fourth SIE network 4 , and uses the upsampling method to upsample X 4 to obtain the upsampled feature map X U , and sends X U to the fifth SIE network. The method is as follows:
[0206] 3.3.4.6.1 The first linear layer of the upsampling module performs a linear transformation on X 4 in the channel domain, changing the number of channels to twice that of X 4 (i.e., the number of channels is 1664) to obtain the channel-transformed feature map.
[0207] 3.3.4.6.2 Change the shape of the channel-transformed feature map so that the height and width are twice that of the channel-transformed feature map, i.e., the feature map resolution is and the number of channels is 1 / 4 of the channel-transformed feature map (i.e., the number of channels is 416) to obtain the shape-transformed feature map.
[0208] 3.3.4.6.3 The first linear layer sends the shape-transformed feature map to the second linear layer after hierarchical normalization.
[0209] 3.3.4.6.4 The second linear layer fuses the information in the channel domain without changing the size of the channel dimension to obtain the upsampled feature map X U , and sends X U to the fifth SIE network.
[0210] 3.3.4.6 The fifth SIE network receives X from the upsampling module U . The eighth local SIE network, the ninth local SIE network, and the fifth global SIE network of the fifth SIE network perform feature extraction and enhancement on X 3 to obtain the feature map X after feature enhancement in the fifth stage 5 , and X 5 is also a matrix. Then X 5 is sent to the global ocean area reduction module.
[0211] 3.3.5 The second Layer Normalization layer of the global ocean area reduction module receives X from the ocean feature aggregation module 5 and, receives X from the ocean feature extraction module, concatenates the two feature maps X 5 and X in the channel dimension to obtain the channel-concatenated feature map, normalizes the channel-concatenated feature map to obtain the normalized channel-concatenated feature map, and sends the normalized channel-concatenated feature map to the two-dimensional transposed convolutional neural network of the global ocean area reduction module.
[0212] 3.3.6 The two-dimensional transposed convolutional neural network upsamples the resolution of the normalized channel-concatenated feature map by 6 times. The height of the three-dimensional grid data is equal to the height H of I train The width is equal to the width W of I train The number of channels is the number of elements C to be predicted out (C out = 2 + number of layers × 4, including sea surface temperature, sea surface height, and temperature, salinity, eastward velocity of seawater, and northward velocity of seawater at each layer), and the three-dimensional grid data is the forecast result of global ocean environmental elements The data format is [H×W×C out , It represents the forecast results of sea surface temperature, sea surface height, and temperature, salinity, flow velocity (the flow velocity includes two components: eastward flow velocity of seawater and northward flow velocity of seawater) from the 1st to the 33rd layer at the (t + τ)-th day starting from the t-th day
[0213] 3.3.7 Design the total loss function of the ocean environmental forecasting system As shown in formula (2):
[0214]
[0215] where represents The predicted value of the c-th (c ≤ C out ) element at the grid point with coordinates [i, j] (i ≤ G, j ≤ W) on the longitude-latitude grid represents Y t+τ The true value of the c-th element at the grid point with coordinates [i, j] on the longitude-latitude grid
[0216] 3.3.8 Let N b = N b + 1. If go to 3.3.2; if (|D M | represents the total number of three-dimensional grid data in the training set D M ), go to the fourth step
[0217] Fourth step, use the validation set to verify the forecasting accuracy of the ocean environmental forecasting system after the end of the current epoch, and retain the network weight parameters with the best performance as the network weight parameters of the global ocean environmental forecasting system. The method is as follows:
[0218] 4.1 Input the data in the standardized validation set D V obtained in step 2.5.3 into the ocean environmental forecasting system after the end of the current epoch;
[0219] 4.2 Let the variable v = 1, and let V be the total number of all three-dimensional grid data in the validation set;
[0220] 4.3 The ocean feature extraction module extracts from D V reads the v-th three-dimensional grid data D v , and uses the ocean feature extraction method described in step 3.3.3 to extract features from D v to obtain a normalized feature map X of the ocean features containing D v , and sends X v to the ocean feature aggregation module and the global ocean area reduction module. X v is v a matrix with a resolution size of and 416 channels;
[0221] 4.4 The first SIE network in the ocean feature aggregation module receives X v , and uses the feature aggregation method based on local and global spatial information extraction described in step 3.3.4.1 to perform local spatial self-attention enhancement and global information fusion on X v to obtain a feature map with enhanced features in the first stage of D v , sends to the downsampling module;
[0222] 4.5 The downsampling module receives from the first SIE network and uses the patch combination splicing and channel domain linear transformation method described in step 3.3.4.2 to perform downsampling on to obtain a downsampled feature map of the th and sends to the second SIE network. is a matrix with a resolution size of
[0223] 4.6 The second SIE network receives from the downsampling module The second local SIE network and the third local SIE network of the second SIE network use the local feature extraction method described in step 3.3.4.1.1 to perform feature extraction on twice, and the second global SIE network uses the feature enhancement method described in step 3.3.4.1.2 to perform global information extraction and fusion feature enhancement on after 2 times of feature extraction to obtain a feature map with enhanced features in the second stage and sends to the third SIE network. is matrix with a resolution size of The number of channels is 832;
[0224] 4.7 The third SIE network receives what is sent by the second SIE network The fourth local SIE network and the fifth local SIE network of the third SIE network use the feature extraction method described in step 3.3.4.1.1 to perform feature extraction twice, and the third global SIE network uses the local feature enhancement method described in step 3.3.4.1.2 for the after the two - time feature extraction to perform global information extraction and fusion feature enhancement, obtaining the feature map after the third - stage feature enhancement Send to the fourth SIE network. is matrix with a resolution size of The number of channels is 832;
[0225] 4.8 The fourth SIE network receives what is sent by the third SIE network The sixth local SIE network and the seventh local SIE network of the fourth SIE network use the feature extraction method described in step 3.3.4.1.1 to perform feature extraction twice, and the fourth global SIE network uses the local feature enhancement method described in step 3.3.4.1.2 for the after the two - time feature extraction to perform global information extraction and fusion feature enhancement, obtaining the feature map after the fourth - stage feature enhancement Send to the up - sampling module. is matrix with a resolution size of The number of channels is 832;
[0226] 4.9 The up - sampling module receives the feature map and uses the up - sampling method described in step 3.3.4.6 to perform up - sampling, obtaining the up - sampled feature map of D v Send to the fifth SIE network. is matrix with a resolution size of The number of channels is 416;
[0227] 4.10 The fifth SIE network receives the up - sampled feature map The eighth local SIE network uses the local feature extraction method described in step 3.3.4.1.1 to Feature extraction is performed. The fifth global SIE network uses the feature enhancement method described in step 3.3.4.1.2 to perform feature enhancement on the features after extraction to extract and fuse global information for feature enhancement, obtaining D v the feature map after the fifth-stage feature enhancement of D Send it to the global ocean area restoration module;
[0228] 4.11 The two-dimensional transposed convolutional neural network of the global ocean area restoration module receives Perform upsampling by a factor of 6 and output a three-dimensional grid data. The height of the three-dimensional grid data is equal to the height H of the three-dimensional grid data in the training set, the width is equal to the width W of the three-dimensional grid data in the training set, and the number of channels is the number of elements C to be predicted out The output three-dimensional grid data is the forecast result of the global ocean environmental elements It takes the t-th day as the starting time and is the forecast result of the sea surface temperature, sea surface height, and temperature, salinity, and flow velocity (the flow velocity includes two components: the eastward flow velocity of seawater and the northward flow velocity of seawater) from the 1st to the 33rd layers on the (t + τ)-th day;
[0229] 4.12 The second Layer Normalization layer of the global ocean area restoration module uses the root mean square error (RMSE) to measure the forecasting performance of the ocean environmental forecasting system for the τ days after taking D v as the input. The calculation method of RMSE is shown in formula (3):
[0230]
[0231] where DM represents the inverse normalization operation, represents the forecast value of the c-th (c ≤ C out ) element at the grid point with coordinates [i, j] (i ≤ H, j ≤ W) on the longitude-latitude grid, represents the true value of the c-th element at the grid point with coordinates [i, j] on the longitude-latitude grid.
[0232] 4.13 Let v = v + 1. If v ≤ V, go to 4.3; if v > V, it means that the set of global ocean environmental forecasting results after the end of the current epoch has been obtained, go to 4.14;
[0233] 4.14 The second Layer Normalization layer calculates the mean of the RMSE of the ocean environmental element forecasting set to obtain the average RMSE of the system on the validation set, and go to 4.15;
[0234] 4.15 If epoch = 1, directly save the network weight parameters after the end of the current epoch, and go to 4.16; if epoch = 3, set learning_rate = 5×10 -5 , if the average RMSE of the system on the validation set after the end of the current epoch is less than the average RMSE after the end of the previous epoch, update the network weight parameters to the network weight parameters after the end of the current epoch, and go to 4.16; if the average RMSE of the system on the validation set after the end of the current epoch is not less than the average RMSE after the end of the previous epoch, directly go to 4.16; if epoch = 2 or epoch > 3, if the average RMSE of the system on the validation set after the end of the current epoch is less than the average RMSE after the end of the previous epoch, update the network weight parameters to the network weight parameters after the end of the current epoch, and go to 4.16; if the average RMSE of the system on the validation set after the end of the current epoch is not less than the average RMSE after the end of the previous epoch, directly go to 4.16;
[0235] 4.16 Set epoch = epoch + 1. If epoch ≤ maxepoch, go to 3.3.2; if epoch > maxepoch, it means the training is over, and the trained global ocean environmental forecasting system is obtained. Go to the fifth step;
[0236] Fifth step, use the trained ocean environmental forecasting system to forecast the global ocean environmental data I input by the user. The method is as follows:
[0237] 5.1 According to the overall mean and standard deviation of the training set train data (the overall mean and standard deviation of the data input by the user are unknown, so the overall mean and standard deviation of the training set are used as a substitute), use the standardization method described in step 2.5.2 to standardize the global ocean environmental data I input by the user. I is a three-dimensional grid data, that is, a three-dimensional matrix with the format [H×W×V], and the standardized matrix I nor , and input I nor into the ocean feature extraction module;
[0238] 5.2 The ocean feature extraction module receives I nor , and uses the ocean feature extraction method described in step 3.3.3 to extract features from I nor to obtain the normalized feature map X containing the ocean features of I nor , and send X I to the ocean feature aggregation module. I
[0239] 5.3 The first SIE network in the ocean feature aggregation module receives X I , the feature aggregation method for local and global spatial information extraction described in step 3.3.4.1 is used to process X I Perform local spatial self-attention enhancement and global information fusion to obtain I nor The feature map after the first-stage feature enhancement of Send To the downsampling module;
[0240] 5.4 The downsampling module receives Use the patch combination splicing and channel-domain linear transformation method described in 3.3.4.2 to Perform downsampling to obtain The downsampled feature map of Send To the second SIE network. Is Matrix of The number of channels is 832;
[0241] 5.5 The second SIE network receives from the downsampling module Perform the second-stage feature extraction on The second local SIE network and the third local SIE network of the second SIE network use the local feature extraction method described in step 3.3.4.1.1 to Perform 2 times of feature extraction. The second global SIE network uses the feature enhancement method described in step 3.3.4.1.2 to After 2 times of feature extraction, perform global information extraction and fusion feature enhancement to obtain I nor The feature map of the second-stage feature enhancement of Send To the third SIE network. Is Matrix of The number of channels is 832;
[0242] 5.6 The third SIE network receives The fourth local SIE network and the fifth local SIE network of the third SIE network use the feature extraction method described in step 3.3.4.1.1 to Perform 2 times of feature extraction. The third global SIE network uses the local feature enhancement method described in step 3.3.4.1.2 to After 2 times of feature extraction, perform global information extraction and fusion feature enhancement to obtain I nor The feature map of the third-stage feature enhancement of Send To the fourth SIE network. Is The matrix has a resolution size of The number of channels is 832;
[0243] 5.7 The fourth SIE network receives The sixth local SIE network and the seventh local SIE network of the fourth SIE network use the feature extraction method described in step 3.3.4.1.1 to perform feature extraction twice. The fourth global SIE network uses the local feature enhancement method described in step 3.3.4.1.2 to perform global information extraction and fusion feature enhancement on the nor feature map after two times of feature extraction to obtain the Send to the upsampling module. is The matrix has a resolution size of The number of channels is 832;
[0244] 5.8 The upsampling module receives Use the upsampling method described in step 3.3.4.6 to perform upsampling to obtain the upsampled feature map of Send to the fifth SIE network. is The matrix has a resolution size of The number of channels is 416; Use the upsampling method described in step 3.3.4.6 to perform upsampling to obtain D v the upsampled feature map of
[0245] 5.9 The fifth SIE network receives The eighth local SIE network uses the local feature extraction method described in step 3.3.4.1.1 to perform feature extraction. The fifth global SIE network uses the feature enhancement method described in step 3.3.4.1.2 to perform global information extraction and fusion feature enhancement on the nor feature map after feature extraction to obtain the Send to the global ocean area restoration module;
[0246] 5.10 The two-dimensional transposed convolutional neural network of the global ocean area restoration module receives For Upsample by a factor of 6 and output a three-dimensional grid data. The height of the output three-dimensional grid data is equal to H, the width is equal to W, and the number of channels is the number of elements C to be predicted. out (C out (is the number of elements to be predicted specified by the user). The output three-dimensional grid data is the global ocean environmental prediction result of the global ocean environmental data I input by the user. It is the prediction result of the sea surface temperature, sea surface height, and the temperature, salinity, and flow velocity (the flow velocity includes two components: the eastward flow velocity and the northward flow velocity of seawater) of the 1st to 33rd layers starting from the t-th day and predicting the (t + τ)-th day.
[0247] Step 6, end.
[0248] The prediction data of the present invention is evaluated by the Intercomparison and Validation Task Team (GOV IV-TT) of the Forecast System Intercomparison and Validation Working Group under the Global Ocean Data Assimilation Experiment (GODAE) Ocean View framework. The time span of the prediction data is from 2019 to 2020, and the predicted elements include sea level anomaly SLA (Sea Level Anomaly), sea surface temperature SST (Sea Surface Temperature), the seawater temperature, salinity, and flow velocity (including two components: the eastward flow velocity and the northward flow velocity of seawater) of the 1st to 33rd layers.
[0249] Among them, GODAE started in 1997 and it is the only international organization that coordinates and promotes the development of global and regional ocean analysis and forecasting systems. GODAE OceanView is the inheritance and extension of GODAE. Since its establishment in 2009, it has been committed to continuously consolidating and promoting the development of international operational oceanography through international cooperation. GOV focuses on different scientific issues by establishing different task teams. Among them, the Intercomparison and Validation Task Team (IVTT) was established in 2010 to jointly promote the scientific verification and comparison of operational oceanography systems. Its activities include formulating standards to evaluate analysis and forecast fields, establishing global and regional comparison plans, etc. Under a unified framework, the strengths and weaknesses of international ocean analysis and forecasting systems are judged by means of mutual comparison and quantitative evaluation. The evaluation principles of GODA for hindcast and forecast products of the ocean include Consistency, Quality, and Performance, that is, verifying whether the results of the system are consistent with the existing understanding of ocean circulation and climate characteristics, qualitatively analyzing the differences between the "optimal values" of the system and the ocean true values, and the short-term forecasting ability of each system.
[0250] The evaluation method used is based on four types of "Classes". Class 1 to Class 3 criteria are mostly applied to the assessment of the system's climatology and hindcast: Class 1 criteria are mainly used to evaluate ocean processes such as coherent spatial structures or major current systems, ocean fronts, and mesoscale eddies; Class 2 criteria focus on the comparison with the time series of moored buoys and the statistical analysis between time series; Class 3 criteria mainly evaluate the spatial distribution and temporal variation of volume transport, heat transport, and eddy kinetic energy. While Class 4 criteria focus on the assessment of the system's forecast performance and forecasting skills, and evaluate the near-real-time accuracy of the system by comparing with all available ocean observations (in-situ drifting buoys or satellite data). The establishment of the Class 4 inter-comparison framework is based on the voluntary actions of IVTT participants, that is: one of the operational agencies collects observational data, conducts quality corrections, provides an initial file (in NetCDF format) containing the observations and the forecast fields of its own operational ocean forecast system (OOFS), and uploads it to the USGODAE server. Other participants download this initial file, interpolate their OOFS forecast fields into similar data files according to this file format, and then upload them to the USGODAE server. Finally, the server contains all the OOFS results participating in the Class 4 inter-comparison. Currently, the UK Meteorological Office voluntarily provides the initial files of sea surface temperature (SST), sea level anomaly (SLA), and T / S profiles (Temperature / Salinity Profiles), and Environment Canada provides the initial file of sea ice concentration.
[0251] The experimental environment is centos7.6 (a version of the Linux system), equipped with a central processing unit of Intel Xeon5218 CPU * 2 of the Intel Xeon series, with a processing frequency of 2.30 GHz. In addition, it is equipped with eight NVIDIA A100-SMX (80G) graphics processors, with a core frequency of 1.4 GHz and a video memory capacity of 80 GB.
[0252] The process of testing the present invention is as follows: The forecast results output by the global ocean environmental forecasting system based on multi-level feature aggregation are interpolated to the observation data points in the IVTT standard through two methods of bilinear interpolation and cubic spline interpolation. The interpolated results are combined with other agencies providing forecast data according to the forecast time, and the root mean square error RMSE between the forecast data and the observation data is calculated every day. Finally, the calculation results for the two years from 2019 to 2020 are obtained.
[0253] First, define the performance evaluation indicators for ocean forecasting methods. In this experiment, the IVTT Class4 standard evaluation method under the Global Ocean Data Assimilation Experiment (GODAE) OceanView framework is used to evaluate the root mean square error (RMSE), a key indicator. RMSE is an indicator used to measure the prediction accuracy of a prediction model on continuous data. It represents the difference between the predicted value and the actual observed value. The smaller the RMSE value, the higher the prediction accuracy.
[0254] The experimental results of the present invention are respectively analyzed and compared with the prediction results of four authoritative ocean forecasting agencies of the operational oceanography systems, namely FOAM of the UK Meteorological Office, PSY3 and PSY4 of the French GLORYS12 Center, GIOPS-CONCEPTS of the Canadian Environment Center, and BLUElink OceanMAPS of the Australian Meteorological Office (see the paper by Ryan A G, Regnier C, Divakaran P, et al., "GODAE OceanView Class 4 forecast verification framework: global ocean inter-comparison" [J]. Journal of Operational Oceanography, 2015, 8(sup1): s98 - s111. The paper by Ryan AG, Regnier C, etc.). Since the AI-GOMS system has not publicly reported the forecast evaluation results under the IVTT Class4 standard and cannot process data with a spatial resolution of 1 / 12 degree, it does not have the significance for horizontal comparison.
[0255] The forecast evaluation results of each forecasting system for SLA in terms of the RMSE indicator are shown in Table 1. Table 1 shows the evaluation results of the present invention and four numerical forecasting systems, namely FOAM (which can only forecast up to the 6th day), BLK (which can only forecast up to the 8th day), GIOPS, and PSY4, within a ten-day forecast period. Compared with FOAM, the forecast error of the present invention is reduced by an average of 6.56%; compared with PSY4, the forecast error of the present invention is reduced by an average of 4.67%; compared with GIOPS, the forecast error of the present invention is reduced by an average of 4.93%; compared with BLK, the forecast error of the present invention is reduced by an average of 5.03%. Therefore, the forecast accuracy of the present invention for SLA is better than that of other numerical forecasting systems (the smaller the RMSE value, the higher the forecast accuracy).
[0256] Table 1 Forecast RMSE of each forecasting system for SLA
[0257] Forecast model 1 day 2 days 3 days 4 days 5 days 6 days 7 days 8 days 9 days 10 days FOAM 0.061 0.062 0.063 0.064 0.065 0.066 - - - - BLK 0.069 0.070 0.071 0.071 0.072 0.073 0.074 0.069 - - GIOPS 0.067 0.068 0.069 0.070 0.071 0.072 0.073 0.074 0.075 0.076 PSY4 0.058 0.059 0.060 0.061 0.062 0.063 0.064 0.065 0.066 0.067 The present invention 0.057 0.059 0.061 0.058 0.061 0.060 0.062 0.063 0.064 0.064
[0258] The prediction evaluation results of each prediction system for SST in terms of the RMSE index are shown in Table 2. Table 2 shows the evaluation results of the present invention compared with four numerical prediction systems, namely FOAM (which can only predict up to the 6th day), BLK (which can only predict up to the 7th day), GIOPS, and PSY4, within a ten-day prediction cycle. It can be seen from the experimental results that the prediction accuracy of the present invention for SST is always higher than that of the two numerical prediction systems, PSY4 and GIOPS.
[0259] Table 2 Prediction RMSE of Each Prediction System for SST
[0260] Forecast model 1 day 2 days 3 days 4 days 5 days 6 days 7 days 8 days 9 days 10 days FOAM 0.388 0.456 0.504 0.542 0.576 0.607 - - - - BLK 0.505 0.535 0.563 0.589 0.615 0.641 0.669 - - - GIOPS 0.601 0.616 0.632 0.648 0.664 0.681 0.697 0.718 0.739 0.763 PSY4 0.693 0.703 0.713 0.752 0.739 0.749 0.768 0.790 0.813 0.889 The present invention 0.556 0.580 0.644 0.637 0.650 0.675 0.689 0.703 0.728 0.733
[0261] Since only PSY4 provides the prediction data of the eastward seawater velocity (uo) and the northward seawater velocity (vo) at a depth of 15 m underwater, the present invention is only compared with PSY4 in terms of the prediction of flow velocity. The comparison of the evaluation results of the eastward seawater velocity and the northward seawater velocity at a depth of 15 m underwater of each prediction system in terms of RMSE is shown in Tables 3 and 4. It can be seen from the experimental results that within a 10-day cycle, the prediction accuracy of the present invention for the eastward seawater velocity and the northward seawater velocity at a depth of 15 m underwater is always higher than that of PSY4.
[0262] Table 3 Prediction RMSE of Each Prediction System for uo at a Depth of 15 m Underwater
[0263] Forecast model 1 day 2 days 3 days 4 days 5 days 6 days 7 days 8 days 9 days 10 days PSY4 0.2025 0.2049 0.2071 0.2093 0.2116 0.2136 0.2157 0.2177 0.2195 0.2214 The present invention 0.1925 0.1945 0.1954 0.1963 0.1977 0.1978 0.1948 0.1977 0.1985 0.1989
[0264] Table 4 Prediction RMSE of Each Prediction System for vo at a Depth of 15 m Underwater
[0265] Forecast model 1 day 2 days 3 days 4 days 5 days 6 days 7 days 8 days 9 days 10 days PSY4 0.1969 0.1990 0.2011 0.2031 0.2052 0.2072 0.2092 0.2113 0.2132 0.2149 The present invention 0.1868 0.1881 0.1890 0.1891 0.1898 0.1900 0.1886 0.1896 0.1899 0.1908
[0266] The prediction evaluation results of each prediction system for the temperature of layers 1 to 33 in terms of the RMSE index are shown in Table 5. Table 5 shows the evaluation results of the present invention compared with four numerical prediction systems, namely FOAM (which can only predict up to the 6th day), BLK (which can only predict up to the 7th day), GIOPS, and PSY4, within a ten-day prediction cycle. It can be seen from the experimental results that the average prediction accuracy of the present invention for the temperature of layers 1 to 33 is always higher than that of other numerical prediction systems.
[0267] Table 5 Average Prediction RMSE of Each Prediction System for the Temperature of Layers 1 to 33
[0268]
[0269]
[0270] The prediction evaluation results of each prediction system for the salinity of layers 1 to 33 in terms of the RMSE index are shown in Table 6. Table 6 shows the evaluation results of the present invention compared with four numerical prediction systems, namely FOAM (which can only predict up to the 6th day), BLK (which can only predict up to the 7th day), GIOPS, and PSY4, within a ten-day prediction cycle. From the experimental results, it can be seen that the average prediction accuracy of the present invention for the salinity of layers 1 to 33 is always higher than that of the two numerical prediction systems, PSY4 and GIOPS.
[0271] Table 6 Average prediction RMSE of each prediction system for the salinity of layers 1 to 33
[0272] Forecast model 1 day 2 days 3 days 4 days 5 days 6 days 7 days 8 days 9 days 10 days FOAM 0.1066 0.1071 0.1075 0.1076 0.1077 0.1089 - - - - BLK 0.1390 0.1398 0.1407 0.1405 0.1416 0.1424 0.1429 - - - GIOPS 0.1141 0.1146 0.1151 0.1158 0.1167 0.1173 0.1181 0.1188 0.1195 0.1202 PSY4 0.1108 0.1109 0.1121 0.1124 0.1135 0.1137 0.1145 0.1147 0.1157 0.1158 The present invention 0.0995 0.1018 0.1033 0.1053 0.1053 0.1065 0.1061 0.1075 0.1091 0.1094
[0273] Figure 3 For an example diagram of the prediction results in the above test process (taking the prediction results of the sea surface temperature on the 7th day, the salinity of the first layer, and the seawater velocity of the first layer as examples). Figure 3 (a) shows a rendering of the prediction result of the sea surface temperature on January 8, 2019, obtained by using the sea surface temperature data on January 1, 2019, as the input of the global ocean environment prediction system of the present invention, indicating that the present invention can predict the sea surface temperature for 7 days and beyond. The prediction error of the sea surface temperature on the 7th day is shown in Table 2, which is 0.062, proving that the present invention can effectively predict the sea surface temperature; Figure 3 (b) shows a rendering of the prediction result of the salinity data of the first layer on January 8, 2019, obtained by using the salinity data of the first layer on January 1, 2019, as the input of the global ocean environment prediction system of the present invention, indicating that the present invention can predict the salinity for 7 days and beyond. The average prediction error of the salinity of layers 1 to 33 on the 7th day is shown in Table 6, which is 0.1061, proving that the present invention can effectively predict the salinity; Figure 3 (c) shows a rendering of the prediction result of the eastward velocity of the seawater of the first layer on January 8, 2019, obtained by using the eastward velocity data of the seawater of the first layer on January 1, 2019, as the input of the global ocean environment prediction system; Figure 3 (d) shows a rendering of the prediction result of the northward velocity of the seawater of the first layer on January 8, 2019, obtained by using the northward velocity data of the seawater of the first layer on January 1, 2019, as the input of the global ocean environment prediction system. Figure 3 (c) and (d) indicate that the present invention can predict the seawater velocity for 7 days and beyond. The prediction error of the eastward velocity of the seawater at a depth of 15 meters on the 7th day is shown in Table 3, which is 0.1948; the prediction error of the northward velocity of the seawater at a depth of 15 meters on the 7th day is shown in Table 4, which is 0.1886, proving that the present invention can effectively predict the seawater velocity.
Claims
1. A global ocean environmental forecasting method based on multi-level feature aggregation, characterized in that It includes the following steps: In the first step, a global ocean environmental forecasting system based on multi-level feature aggregation is constructed; the ocean environmental forecasting system consists of an ocean feature extraction module, an ocean feature aggregation module, and a global ocean area restoration module; The ocean feature extraction module is connected to the ocean feature aggregation module and the global ocean area restoration module. The ocean feature extraction module receives the global ocean environmental data input by the user, which is three-dimensional grid data containing the elements of each layer of the ocean, extracts the ocean features from the global ocean environmental data, and sends the feature maps of the ocean features to the ocean feature aggregation module and the global ocean area restoration module; The ocean feature extraction module consists of a two-dimensional convolutional neural network and a first Layer Normalization layer; the two-dimensional convolutional neural network divides the three-dimensional grid data into blocks, that is, patches, extracts the ocean features of the global ocean environmental data, and sends the feature maps of the ocean features to the first Layer Normalization layer; the first Layer Normalization layer normalizes the feature maps of the ocean features and sends the normalized feature maps of the ocean features to the ocean feature aggregation module; The ocean feature aggregation module is connected to the ocean feature extraction module and the global ocean area restoration module; the ocean feature aggregation module consists of 5 spatial information extraction networks with non-shared weights, namely the SIE network, the upsampling module, and the downsampling module; the first SIE network and the fifth SIE network have the same structure, both consisting of a local SIE network and a global SIE network; the second SIE network to the fourth SIE network have the same structure, all consisting of two consecutive local SIE networks and a global SIE network; the first SIE network is composed of the first local SIE network and the first global SIE network; the first local SIE network is composed of a window multi-head self-attention layer and a multi-layer perceptron MLP; the window multi-head self-attention layer receives the feature map of the normalized ocean features from the ocean feature extraction module, performs hierarchical normalization on the feature map of the normalized ocean features, divides the hierarchically normalized feature map into multiple windows, uses the window multi-head self-attention mechanism and residual connection operations to obtain the attention mechanism-enhanced feature map, and sends the attention mechanism-enhanced feature map to the multi-layer perceptron; the multi-layer perceptron consists of two linear layers, receives the attention mechanism-enhanced feature map from the window multi-head self-attention layer, performs hierarchical normalization, linear projection, and residual connection operations on it, obtains the feature map that fuses the local spatial information within the window, and sends the feature map that fuses the local spatial information within the window to the first global SIE network; the first global SIE network is composed of a feature grouping network, a group feature fusion network, and a group feature propagation network; the feature grouping network of the first global SIE network receives the feature map that fuses the local spatial information within the window from the first local SIE network, uses the cross-attention mechanism to group and aggregate the features in the feature map that fuses the local spatial information within the window to obtain group features, and sends the group features to the group feature fusion network; the main part of the group feature fusion network is composed of MLP-mixer; the group feature fusion network receives the group features from the feature grouping network, exchanges group feature information and updates the group features, so that each group feature can reflect global information to a certain extent while containing the current group information, obtains the updated group features, and sends the updated group features to the group feature propagation network; The group feature propagation network uses the cross-attention mechanism to propagate the updated group features to all the features in the feature map that fuses the local spatial information within the window, and uses the depthwise separable convolutional neural network to perform convolutional feature extraction on the feature map that fuses the local spatial information within the window to obtain the feature map after the first-stage feature enhancement, and sends the feature map after the first-stage feature enhancement to the downsampling module; the local SIE network structures in the second SIE network to the fifth SIE network are the same as the first local SIE network structure, and the global SIE network structures in the second SIE network to the fifth SIE network are the same as the first global SIE network; The downsampling module is connected to the first SIE network and the second SIE network and consists of a linear layer. The downsampling module receives the feature map after the first-stage feature enhancement from the first SIE network, combines every 2×2 adjacent patches in the feature map after the first-stage feature enhancement together, performs a hierarchical normalization operation after concatenating them in the channel dimension, and obtains a channel-concatenated feature map with 4 times the number of channels of the feature map after the first-stage feature enhancement. The downsampling module performs a linear transformation on the channel dimension of the channel-concatenated feature map, halves the number of channels of the channel-concatenated feature map, and obtains a downsampled feature map. The height and width of the downsampled feature map are half of those of the feature map after the first-stage feature enhancement, and the number of channels is 2 times that of the feature map after the first-stage feature enhancement, that is, the downsampling ratio becomes twice that of the first SIE network. Then, the downsampled feature map is sent to the second SIE network. The second local SIE network, the third local SIE network, and the second global SIE network of the second SIE network enhance the downsampled feature map to obtain a feature map after the second-stage feature enhancement, and send the feature map after the second-stage feature enhancement to the third SIE network. The fourth local SIE network, the fifth local SIE network, and the third global SIE network of the third SIE network enhance the feature map after the second-stage feature enhancement to obtain a feature map after the third-stage feature enhancement, and send the feature map after the third-stage feature enhancement to the fourth SIE network. The sixth local SIE network, the seventh local SIE network, and the fourth global SIE network of the fourth SIE network enhance the feature map after the third-stage feature enhancement to obtain a feature map after the fourth-stage feature enhancement, and send the feature map after the fourth-stage feature enhancement to the upsampling module; The upsampling module is connected to the fourth SIE network and the fifth SIE network and consists of two linear layers. The upsampling module receives the feature map after the fourth-stage feature enhancement from the fourth SIE network; The first linear layer of the upsampling module performs a linear transformation on the channel dimension of the feature map after the fourth-stage feature enhancement to obtain a channel-transformed feature map with 2 times the number of channels of the feature map after the fourth-stage feature enhancement. The shape of the channel-transformed feature map is changed so that the height and width are 2 times those of the channel-transformed feature map, and the number of channels is half of that of the channel-transformed feature map, obtaining a shape-transformed feature map, and sending the shape-transformed feature map to the second linear layer of the upsampling module. The second linear layer of the upsampling module fuses the channel-domain information of the shape-transformed feature map without changing the size of the channel dimension to obtain an upsampled feature map, and sends the upsampled feature map to the fifth SIE network. The height and width of the upsampled feature map are 2 times those of the feature map after the fourth-stage feature enhancement, and the number of channels is half of that of the feature map after the fourth-stage feature enhancement; The fifth SIE network performs a fifth-stage feature enhancement on the upsampled feature map to obtain a feature map after the fifth-stage feature enhancement, and sends the feature map after the fifth-stage feature enhancement to the global ocean area restoration module; A masking mechanism is added to the window multi-head self-attention layer in the local SIE network among the first to fifth SIE networks and the feature grouping network in the global SIE network, making the attention score of the land area 0; The global ocean area restoration module consists of a second Layer Normalization layer and a two-dimensional transposed convolutional neural network, and is connected to the fifth SIE network of the ocean feature extraction module and the ocean feature aggregation module. It receives the feature map of the fifth-stage feature enhancement from the ocean feature aggregation module and the feature map of the normalized ocean features from the ocean feature extraction module; The second Layer Normalization layer of the global ocean area restoration module concatenates the feature map of the fifth-stage feature enhancement and the normalized ocean features in the channel dimension to obtain a channel-concatenated feature map, and then normalizes the channel-concatenated feature map and sends the normalized channel-concatenated feature map to the two-dimensional transposed convolutional neural network layer; the two-dimensional transposed convolutional neural network layer upsamples the normalized channel-concatenated feature map by 6 times using transposed convolution to obtain the predicted three-dimensional grid data. The height and width of the predicted three-dimensional grid data are the same as those of the user-input global ocean environment data, and the channel dimension is the number of total target elements. The predicted three-dimensional grid data is the global ocean environment forecast result; Step 2: Construct the training set, validation set, and test set. The method is as follows: 2.1 Download the GLORYS12 global ocean reanalysis data, ERA5 wind field reanalysis data, and GHR sea surface temperature satellite data, and perform quality control on the data. The method is as follows: 2.1.1 Download the GLORYS12 global ocean reanalysis data, ERA5 wind field reanalysis data, and GHR sea surface temperature data, and collect the data with a time span from 1993 to 2020 as the dataset for ocean environment forecasting; select five elements in the GLORYS12 data, namely thetao (temperature), so (salinity), uo (eastward seawater velocity), vo (northward seawater velocity), and zos (sea surface height). For temperature, salinity, eastward seawater velocity, and northward seawater velocity, extract the data from layers 1 to 33, and extract the data of the first layer for the sea surface height; also use the ERA5 reanalysis data as a component of the dataset for ocean environment forecasting. The spatial resolution of the ERA5 wind field reanalysis data is 1 / 4 degree, the spatial span is -180 degrees to 180 degrees in the longitude direction and -90 degrees to 90 degrees in the latitude direction, the time resolution is 1 hour, and it includes two single-layer elements, namely u10 (10-meter eastward wind speed) and v10 (10-meter northward wind speed); the spatial resolution of the GHR sea surface temperature data is 1 / 20 degree, the spatial span is -180 degrees to 180 degrees in the longitude direction and -90 degrees to 90 degrees in the latitude direction, the time resolution is 1 day, and it only includes sea surface temperature satellite data; 2.1.2 Perform quality control on the data: Use the xarray library in Python to read the GLORYS12 global ocean reanalysis data, ERA5 wind field reanalysis data, and GHR sea surface temperature satellite data starting from January 1, 1993. If the data is "dirty" data with problems, that is, data is missing, damaged, or some elements in the data are missing, then write the "dirty" data to the log; then re-download the good data corresponding to the "dirty" data. 2.2 Let the number of days from 1993 to 2020 be D, and let the original datasets of the three datasets of GLORYS12 global ocean reanalysis data, ERA5 wind field reanalysis data, and GHR sea surface temperature satellite data be M, E, and G. Align the time and space resolutions of M, E, and G. The method is as follows: 2.2.1 Align the time resolution of the ERA5 wind field reanalysis data with the GLORYS12 global ocean reanalysis data and the GHR sea surface temperature satellite data to obtain the ERA5 wind field reanalysis dataset E1 with aligned time resolution. 2.2.2 Align the space resolutions of the two datasets M and E1 with G to obtain the ERA5 wind field reanalysis dataset E2 with a space resolution of 1 / 12 degree and the sea surface temperature dataset G1 of GHR. 2.2.3 Align the spatial ranges of the two datasets E2 and G1 with that of M to obtain the ERA5 wind field reanalysis dataset E3 with aligned spatial range and the sea surface temperature dataset G2 of GHR. 2.4 Convert the data in E3, G2, and M from nc format to npy format to obtain the dataset S in npy format. 2.5 Divide S into a training set, a validation set, and a test set, and calculate the mean and standard deviation of the training set data to normalize the data. The method is as follows: 2.5.1 Use the mv command to move the data from 1993 to 2017 in S to the training set train, move the data in 2018 to the validation set val, and move the data from 2019 to 2020 to the test set test. 2.5.2 Standardize the data in the training set train according to the overall mean and standard deviation of the training set train data to obtain the standardized training set D M ; The standardization method is as follows: 2.5.2.1 Let the variable d = 1, and take the data of the d-th day in train as X d , and initialize the standardized training set D M to be empty; 2.5.2.2 Standardize X d The specific process is shown in Formula (1) as follows: where μ is the overall mean of the training set train data, and σ is the overall standard deviation; the standardized X d is put into D M ; 2.5.2.3 If d ≤ D, let d = d + 1, and go to 2.5.2.2; if d > D, the standardization of the training set train data is completed, and go to 2.5.
3. 2.5.3 Standardize the data in the validation set val using the standardization method described in step 2.5.2 according to the overall mean and standard deviation of the training set data, and obtain the standardized validation set D V ; 2.5.4 Standardize the data in the test set test according to the overall mean and standard deviation of the training set data to obtain the standardized test set D T ; Step 3: Use the gradient backpropagation method to train the Global Ocean Environment Forecasting System to obtain the network weight parameters with the best forecasting effect on the validation set D v The method is as follows: 3.1 Initialize the network weight parameters of each module in the ocean environmental forecasting system; use a normal distribution with a mean of 0 and a variance of 0.01 to initialize the network weight parameters of the ocean feature extraction module, the ocean feature aggregation module, and the global ocean area reduction module. 3.2 Set the training parameters of the ocean environmental forecasting system, including the warm-up learning rate, the warm-up training step, the learning rate learning_rate, the batch size mini_batch_size for network training; the maximum training step maxepoch; select AdamW as the model training optimizer, and set the hyperparameters β1, β2, and "weight decay" of the model training optimizer. 3.3 Train the ocean environmental forecasting system by taking the square of the difference between the forecast result output by the ocean environmental forecasting system during one training and the true value as the loss value loss, and using gradient backpropagation to update the network weight parameters until the loss value reaches the threshold or the training step reaches maxepoch; in each training round, if the forecasting accuracy of the current system on the validation set D v is the best, save the network weight parameters of the current round; the method is as follows: 3.3.1 Let the training step epoch = 1. One pass through all the data in the training set is one epoch. Initialize the batch number N b = 1; 3.3.2 The ocean feature extraction module reads the N M -th batch from D b , with a total of B three-dimensional grid data. These B three-dimensional grid data are denoted in matrix form as I train . I train contains B matrices of H×W×V; where H represents the height of the input grid data, W represents the width of the input grid data, V represents the channel dimension, V = 3 + number of layers × 4, that is, 3 sea surface elements, as well as the temperature, salinity, eastward velocity of seawater, and northward velocity of seawater at each layer. B is a positive integer; 3.3.3 The ocean feature extraction module uses an ocean feature extraction method to extract features from I train and sends the normalized feature map X containing the ocean features of I train to the ocean feature aggregation module and the global ocean area reduction module. The method is as follows: 3.3.3.1 The two-dimensional convolutional neural network in the ocean feature extraction module extracts the ocean features of I train to obtain a feature map containing ocean features. The method is as follows: Use convolution to perform patch division on B three-dimensional grid data of I train to obtain a feature map containing ocean features; then send the feature map containing ocean features to the first Layer Normalization layer; 3.3.3.2 The first Layer Normalization layer normalizes the feature maps containing ocean features at each level of the three-dimensional grid data to obtain the normalized feature map X containing ocean features with I train , and the normalization operation helps to improve the training effect of the network; the first Layer Normalization layer sends X to the ocean feature aggregation module and the global ocean area restoration module; the resolution of X is The number of channels of X is 416; 3.3.4 The ocean feature aggregation module receives X from the ocean feature extraction module, generates the feature map X5 with enhanced features in the fifth stage, and sends X5 to the global ocean area reduction module. The method is as follows: 3.3.4.1 The first SIE network receives X from the ocean feature extraction module, and uses a feature aggregation method based on the extraction of local and global spatial information to perform local spatial self-attention enhancement and global information fusion on X, obtaining a feature map X with enhanced features in the first stage of X 1 ; The method is as follows: 3.3.4.1.1 The first local SIE network of the first SIE network uses a local feature extraction method to extract features from X. The method is as follows: The window multi-head self-attention layer of the first local SIE network performs hierarchical normalization on X and divides X into multiple windows; then the window multi-head self-attention layer of the first local SIE network parallelly uses the window multi-head self-attention mechanism to enhance the local spatial self-attention of X within each window, and then performs a residual connection to obtain the feature map after the residual connection. During the calculation process of the window multi-head self-attention mechanism, the attention weights of the features corresponding to the land areas in X are set to 0 using the masking mechanism. Then, the window multi-head self-attention layer of the first local SIE network sends it to the multi-layer perceptron of the first local SIE network; the multi-layer perceptron of the first local SIE network first performs hierarchical normalization, then uses a linear mapping to fuse in the channel dimension, and then performs a residual connection operation to obtain the feature map that fuses the local spatial information within the window. And sends it to the first global SIE module. 3.3.4.1.2 The first global SIE network of the first SIE network receives from the first local SIE network and uses a feature enhancement method to extract and fuse global information to obtain a feature map X of the first-stage feature enhancement that fuses local and global spatial information 1 , and the method is as follows: 3.3.4.1.2.1 The first feature grouping network randomly initializes a set of group features G, where G contains N features, and the dimension of each feature is d, and d is 416. The first feature grouping network first uses the multi-head cross-attention mechanism to aggregate and update the feature groups in, and obtain the updated group features G ′ , and each feature in G ′ aggregates the information of a cluster of patches with similar semantics; during the calculation of the multi-head cross-attention mechanism, a masking mechanism is used to mask the features corresponding to the land part in, so that the land part does not participate in the calculation of the cross-attention mechanism; then the first feature grouping network sends the group features G ′ to the first group feature fusion network. 3.3.4.1.2.2 The first group of feature fusion networks receive G from the first feature grouping network ′ , perform information exchange among each feature in G ′ , then fuse and update the group feature information of G ′ to obtain the updated group feature The first MLP layer of the first group of feature fusion networks performs fusion on the group feature that obtains the fused spatial information in the G ′ spatial domain The second MLP layer performs fusion on the channel domain to obtain the updated group feature The first group of feature fusion networks send the updated group feature to the first group of feature propagation networks; 3.3.4.1.2.3 The first group of feature propagation networks receive from the first group of feature fusion networks and receive from the first local SIE network Adopt a multi-head cross-attention mechanism to use the features in as query vectors, and use the group features as key vectors and value vectors. According to the different features in, assign different weights to each feature in the group features and then perform weighted summation on the features in to obtain the global spatial feature U, and send U to the first group of feature propagation networks; 3.3.4.1.2.4 The feedforward neural network of the first group of feature propagation networks concatenates U with the features in to obtain the concatenated feature map X 1′ , and performs a linear transformation on X 1′ , and then obtains the feature map X 1″ after residual connection through the residual connection; the depthwise separable convolutional neural network performs convolutional feature extraction on X 1″ to obtain the feature map X 1 of the first-stage feature enhancement that fuses local and global spatial information; the resolution of X 1 is X 1 has 416 channels; the first group of feature propagation networks of the first SIE network sends X 1 to the downsampling module; 3.3.4.2 The downsampling module receives X from the first set of feature propagation networks of the first SIE network 1 , and performs downsampling on X using the patch combination splicing and channel domain linear transformation method 1 to obtain the downsampled feature map X D , and sends X D to the second SIE network by the following method: 3.3.4.2.1 Combine each 2×2 adjacent patch in X 1 together, and perform hierarchical normalization after splicing these 4 patches in the channel domain to obtain a channel-spliced feature map with 4 times the number of channels of X 1 The channel dimension of the spliced feature map is 4×416; 3.3.4.2.2 Perform a linear transformation on the channel dimension of the channel concatenated feature map, halve the number of channels of the channel concatenated feature map, and obtain the downsampled feature map X D ; X D The feature map resolution of X D The number of channels of the feature map of X is 832; the downsampled feature map X D The downsampling ratio of X is 1 twice that of X; 3.3.4.2.3 The downsampling module sends X D to the second SIE network; 3.3.4.3 The second SIE network receives the downsampled feature map X sent by the downsampling module D , and the second local SIE network and the third local SIE network use the local feature extraction method described in step 3.3.4.1.1 to process X D for two times of feature extraction. The second global SIE network uses the feature enhancement method described in step 3.3.4.1.2 to perform global information extraction and fusion feature enhancement on X after two times of feature extraction D to obtain the feature map X with feature enhancement in the second stage 2 , and send X 2 to the third SIE network; the resolution of the feature map of X 2 is X 2 The number of channels of the feature map of X is 832. The second SIE network only performs feature enhancement on X 1 without changing the resolution of X 1 ; 3.3.4.4 The third SIE network receives X 2 , the fourth local SIE network and the fifth local SIE network of the third SIE network adopt the feature extraction method described in step 3.3.4.1.1 to perform feature extraction on X 2 twice, and the third global SIE network adopts the local feature enhancement method described in step 3.3.4.1.2 to perform global information extraction and fusion feature enhancement on X after the two-time feature extraction 2 to obtain the feature map X with feature enhancement in the third stage 3 , and send X 3 to the fourth SIE network; the resolution of the feature map of X 3 is X 3 the number of channels of the feature map of X is 832, and the third SIE network only performs feature enhancement on X 2 without changing the resolution of X 2 ; 3.3.4.5 Fourth SIE Network Receives X 3 , the sixth local SIE network and the seventh local SIE network of the fourth SIE network use the feature extraction method described in step 3.3.4.1.1 to perform feature extraction on X 3 twice, and the fourth global SIE network uses the local feature enhancement method described in step 3.3.4.1.2 to perform global information extraction and fusion feature enhancement on X after the two - time feature extraction 3 to obtain the feature map X with enhanced features in the fourth stage 4 , and send X 4 to the up - sampling module; the resolution of the feature map of X 4 is X 4 the number of channels of the feature map of X is 832, and the fourth SIE network only performs feature enhancement on X 3 without changing the resolution of X 3 ; 3.3.4.6 The upsampling module receives X from the fourth SIE network 4 , and performs upsampling on X 4 using the upsampling method to obtain the upsampled feature map X U , and sends X U to the fifth SIE network by the following method: 3.3.4.6.1 The first linear layer of the upsampling module performs a linear transformation in the channel domain of X 4 to double the number of channels to that of X 4 and obtains a channel-transformed feature map; 3.3.4.6.2 The first linear layer of the upsampling module changes the shape of the channel-transformed feature map, such that the height and width are twice that of the channel-transformed feature map, i.e., the feature map resolution is The number of channels is 1 / 4 of that of the channel-transformed feature map, obtaining a shape-transformed feature map; The first linear layer of the upsampling module in 3.3.4.6.3 performs hierarchical normalization on the shape-transformed feature map to obtain the shape-transformed feature map after hierarchical normalization, and sends the shape-transformed feature map after hierarchical normalization to the second linear layer; The second linear layer of the upsampling module in 3.3.4.6.4 fuses the information in the channel domain of the shape-transformed feature map after hierarchical normalization, without changing the size of the channel dimension, to obtain the upsampled feature map X U , and sends X U to the fifth SIE network; 3.3.4.7 The fifth SIE network receives X from the upsampling module U , and the eighth local SIE network of the fifth SIE network uses the local feature extraction method described in step 3.3.4.1.1 to extract features from X U . The fifth global SIE network uses the feature enhancement method described in step 3.3.4.1.2 to extract and fuse global information for the X after feature extraction U to obtain the feature map X after feature enhancement in the fifth stage 5 , and X 5 is also a matrix; send X 5 to the global ocean area restoration module; 3.3.5 The second Layer Normalization layer of the global ocean area restoration module receives X from the ocean feature aggregation module 5 , receives X from the ocean feature extraction module, and concatenates X 5 and X in the channel dimension to obtain a channel concatenated feature map, normalizes the channel concatenated feature map to obtain a normalized channel concatenated feature map, and sends the normalized channel concatenated feature map to the two-dimensional transposed convolutional neural network of the global ocean area restoration module; 3.3.4.6 The upsampling module receives X from the fourth SIE network 4 , and performs upsampling on X 4 using the upsampling method to obtain the upsampled feature map X U , and sends X U to the fifth SIE network by the following method: 3.3.4.6.1 The first linear layer of the upsampling module performs a linear transformation in the channel domain of X 4 and doubles the number of channels to obtain a channel-transformed feature map; 4 3.3.4.6.2 Change the shape of the channel-transformed feature map so that the height and width are twice that of the channel-transformed feature map, i.e., the feature map resolution is The number of channels is 1 / 4 of the channel-transformed feature map to obtain a shape-transformed feature map; The first linear layer in 3.3.4.6.3 performs hierarchical normalization on the shape-transformed feature map and then sends it to the second linear layer; 3.3.4.6.4 The second linear layer fuses the information in the channel domain to obtain the upsampled feature map X without changing the size of the channel dimension. U , and send X U to the fifth SIE network; 3.3.4.6 The fifth SIE network receives X from the upsampling module U , and the eighth local SIE network, the ninth local SIE network, and the fifth global SIE network of the fifth SIE network perform feature extraction and enhancement on X 3 to obtain the feature map X after feature enhancement in the fifth stage 5 , and X 5 is also a matrix of; then send X 5 to the global ocean area restoration module; 3.3.5 The second Layer Normalization layer of the global ocean area reduction module receives X from the ocean feature aggregation module 5 In addition, it receives X from the ocean feature extraction module, and combines X 5 and the two feature maps of X in the channel dimension to obtain a channel concatenated feature map, and normalizes the channel concatenated feature map to obtain a normalized channel concatenated feature map, and sends the normalized channel concatenated feature map to the two-dimensional transposed convolutional neural network of the global ocean area reduction module; 3.3.6 The two-dimensional transposed convolutional neural network upsamples the resolution of the normalized channel-concatenated feature map by 6 times to obtain the predicted three-dimensional grid data. The height of the three-dimensional grid data is equal to the height H of I train The width is equal to the width W of I train The number of channels is the number of elements C to be predicted out , C out = 2 + number of layers × 4, including sea surface temperature, sea surface height, and temperature, salinity, eastward velocity of seawater, and northward velocity of seawater at each layer. The three-dimensional grid data is the generated global ocean environmental element forecast result The data format is [H×W×C out , It represents the forecast results of sea surface temperature, sea surface height, and temperature, salinity, and flow velocity from layer 1 to layer 33 at the (t + τ)-th day starting from the t-th day; the flow velocity includes two components: eastward flow velocity of seawater and northward flow velocity of seawater 3.3.7 Design of the total loss function for the ocean environmental forecasting system As shown in formula (2): wherein denotes the predicted value of the c-th element at the grid point with coordinates [i, j] on the latitude-longitude grid; denotes Y t+τ the true value of the c-th element at the grid point with coordinates [i, j] on the latitude-longitude grid, where i ≤ H, j ≤ W, and c ≤ C out ; 3.3.8 Let N b = N b + 1, if go to 3.3.2; if |D M | represents the total number of three-dimensional grid data in the training set D M go to the fourth step; In the fourth step, use the validation set to verify the prediction accuracy of the global ocean environmental forecasting system after the end of the current epoch, and retain the network weight parameters with the best performance as the network weight parameters of the global ocean environmental forecasting system; the method is: 4.1 Input the data in the standardized validation set D obtained in step 2.5.3 into the ocean environmental forecasting system after the end of the current epoch; V 4.2 Let the variable v = 1, and let V be the total number of all three-dimensional grid data in the validation set; 4.3 Ocean feature extraction module reads from D V the v-th three-dimensional grid data D v , and uses the ocean feature extraction method described in step 3.3.3 to extract features from D v , obtaining a normalized feature map X of the ocean features containing D v . Then it sends X v to the ocean feature aggregation module and the global ocean area reduction module; X v is v a matrix with a resolution size of and 416 channels; 4.4 The first SIE network in the ocean feature aggregation module receives X v , and uses the feature aggregation method based on the extraction of local and global spatial information described in step 3.3.4.1 to perform local spatial self-attention enhancement and global information fusion on X v to obtain D v The feature map after the first-stage feature enhancement of Send to the downsampling module; 4.5 downsampling module receives from the first SIE network Use the patch combination splicing and channel domain linear transformation method described in step 3.3.4.2 to perform downsampling to obtain the downsampled feature map Send to the second SIE network; is matrix, with a resolution size of and 832 channels; 4.6 The second SIE network receives from the downsampling module The second local SIE network and the third local SIE network of the second SIE network adopt the local feature extraction method described in step 3.3.4.1.1 to perform feature extraction twice. The second global SIE network adopts the feature enhancement method described in step 3.3.4.1.2 to perform global information extraction and fusion feature enhancement on the after twice of feature extraction, and obtains the feature map after feature enhancement in the second stage Send to the third SIE network; Yes is a matrix of, with a resolution size of The number of channels is 832; 4.7 The third SIE network receives the one sent by the second SIE network The fourth local SIE network and the fifth local SIE network of the third SIE network adopt the feature extraction method described in step 3.3.4.1.1 to perform feature extraction twice. The third global SIE network adopts the local feature enhancement method described in step 3.3.4.1.2 to extract and fuse the global information for feature enhancement of the result after the two-time feature extraction, and obtain the feature map after feature enhancement in the third stage Send to the fourth SIE network; Yes is the matrix of, with a resolution size of The number of channels is 832; 4.8 The fourth SIE network receives what is sent by the third SIE network The sixth local SIE network and the seventh local SIE network of the fourth SIE network adopt the feature extraction method described in step 3.3.4.1.1 to perform feature extraction twice. The fourth global SIE network adopts the local feature enhancement method described in step 3.3.4.1.2 to extract and fuse global information for feature enhancement on the result after the two - time feature extraction, obtaining the feature map after feature enhancement in the fourth stage Send it to the up - sampling module; It is a matrix of, with a resolution size of and the number of channels is 832; 4.9 The upsampling module receives the feature map Perform upsampling on using the upsampling method described in steps 3.3.4.6 to obtain D v The upsampled feature map of Send to the fifth SIE network; is a matrix of with a resolution size of and 416 channels; 4.10 The fifth SIE network receives the upsampled feature map The eighth local SIE network uses the local feature extraction method described in step 3.3.4.1.1 for feature extraction, and the fifth global SIE network uses the feature enhancement method described in step 3.3.4.1.2 for the extraction and fusion of global information for feature enhancement after feature extraction, obtaining D v the feature map after the fifth-stage feature enhancement Send it to the global ocean area restoration module; 4.11 The two-dimensional transposed convolutional neural network of the global ocean area restoration module receives For Upsample by a factor of 6 and output a three-dimensional grid data. The height of the three-dimensional grid data is equal to the height H of the three-dimensional grid data in the training set, the width is equal to the width W of the three-dimensional grid data in the training set, and the number of channels is the number of elements C to be predicted. out ; The output three-dimensional grid data is the prediction result of the global ocean environmental elements. It takes the t-th day as the starting time and is the prediction results of the sea surface temperature, sea surface height, and temperature, salinity, and flow velocity of layers 1 to 33 on the (t + τ)-th day; 4.12 The second Layer Normalization layer of the global ocean area reduction module uses the root mean square error (RMSE) to measure the forecasting performance of the ocean environmental forecasting system for the τ-th day after taking D v as the input; the calculation method of RMSE is shown in formula (3): where DM represents the inverse normalization operation, represents the predicted value of the c-th element at the grid point with coordinates [i, j] on the latitude-longitude grid, represents the true value of the c-th element at the grid point with coordinates [i, j] on the latitude-longitude grid; 4.13 Let v = v + 1. If v ≤ V, go to 4.3; if v > V, it means that the set of global ocean environmental forecasting results after the end of the current epoch has been obtained, go to 4.14; 4.14 The second Layer Normalization layer calculates the mean of the RMSE of the ocean environmental element forecasting set, and obtains the average RMSE of the system on the validation set, then go to 4.15; 4.15 If epoch = 1, save the network weight parameters after the end of the current epoch, and directly go to 4.16; if epoch = 3, set learning_rate = 5×10 -5 , if the average RMSE of the system on the validation set after the end of the current epoch is less than the average RMSE after the end of the previous epoch, update the network weight parameters to the network weight parameters after the end of the current epoch, and go to 4.16; if the average RMSE of the system on the validation set after the end of the current epoch is not less than the average RMSE after the end of the previous epoch, directly go to 4.16; if epoch = 2 or epoch > 3, if the average RMSE of the system on the validation set after the end of the current epoch is less than the average RMSE after the end of the previous epoch, update the network weight parameters to the network weight parameters after the end of the current epoch, and go to 4.16; if the average RMSE of the system on the validation set after the end of the current epoch is not less than the average RMSE after the end of the previous epoch, directly go to 4.16; 4.16 Let epoch = epoch + 1. If epoch ≤ maxepoch, go to 3.3.2; if epoch > maxepoch, it means that the training is over and the trained global ocean environmental forecasting system is obtained, go to the fifth step; In the fifth step, use the trained ocean environmental forecasting system to forecast the global ocean environmental data I input by the user; the method is: 5.1 Standardize the global ocean environmental data I input by the user using the standardization method described in step 2.5.2 according to the overall mean and standard deviation of the training set train data; I is a three-dimensional grid data, that is, a three-dimensional matrix in the format of [H×W×V], and obtain the standardized matrix I nor , and input I nor into the ocean feature extraction module; 5.2 The ocean feature extraction module receives I nor , and uses the ocean feature extraction method described in step 3.3.3 to extract features from I nor , obtaining a normalized feature map X nor of the ocean features containing I I , and sends X I to the ocean feature aggregation module; 5.3 The first SIE network in the ocean feature aggregation module receives X I , and uses the feature aggregation method for local and global spatial information extraction described in step 3.3.4.1 to X I to perform local spatial self-attention enhancement and global information fusion, obtaining I nor the feature map after the first-stage feature enhancement of Send to the downsampling module; 5.4 The downsampling module receives Perform patch combination splicing and channel domain linear transformation as described in 3.3.4.2 on to perform downsampling and obtain the downsampled feature map of Send to the second SIE network; is a matrix of with a resolution size of and 832 channels; 5.5 The second SIE network receives from the downsampling module and performs feature extraction in the second stage: The second local SIE network and the third local SIE network of the second SIE network use the local feature extraction method described in step 3.3.4.1.1 to perform feature extraction twice, and the second global SIE network uses the feature enhancement method described in step 3.3.4.1.2 to perform extraction and fusion of global information for feature enhancement, obtaining the feature map with feature enhancement in the second stage of I nor ; Then it is sent to the third SIE network; is a matrix of with a resolution size of and 832 channels; 5.6 Third SIE Network Receiving The fourth local SIE network and the fifth local SIE network of the third SIE network use the feature extraction method described in step 3.3.4.1.1 to perform feature extraction twice. The third global SIE network uses the local feature enhancement method described in step 3.3.4.1.2 to extract and fuse global information for feature enhancement on the result after two times of feature extraction, obtaining nor the feature map after feature enhancement in the third stage of I Send to the fourth SIE network; is matrix of, with a resolution size of and the number of channels is 832; 5.7 Fourth SIE Network Receiving The sixth local SIE network and the seventh local SIE network of the fourth SIE network use the feature extraction method described in step 3.3.4.1.1 to perform feature extraction twice. The fourth global SIE network uses the local feature enhancement method described in step 3.3.4.1.2 to perform global information extraction and fusion feature enhancement on the after the two - time feature extraction, and obtain the feature map after the fourth - stage feature enhancement of I nor Send it to the up - sampling module; will is a matrix of, with a resolution size of and the number of channels is 832; 5.8 The upsampling module receives Perform upsampling on using the upsampling method described in step 3.3.4.6 to obtain the upsampled feature map of Send to the fifth SIE network; is a matrix with a resolution size of and 416 channels; perform upsampling on using the upsampling method described in step 3.3.4.6 to obtain the upsampled feature map of D v 5.9 Fifth SIE Network Receiving The eighth partial SIE network uses the partial feature extraction method described in step 3.3.4.1.1 to perform feature extraction. The fifth global SIE network uses the feature enhancement method described in step 3.3.4.1.2 to perform global information extraction and fusion feature enhancement on the result after feature extraction, and the fifth global SIE network extracts global information to obtain the feature map after feature enhancement in the fifth stage of I nor Send it to the global ocean area restoration module; 5.10 The two-dimensional transposed convolutional neural network of the global ocean area restoration module receives Perform Upsample by a factor of 6 and output a three-dimensional grid data; the height of the output three-dimensional grid data is equal to H, the width is equal to W, and the number of channels is the number of elements C to be predicted out ; the output three-dimensional grid data is the global ocean environmental forecast result of the global ocean environmental data I input by the user It takes the t-th day as the starting time and is the forecast result of the sea surface temperature, sea surface height, and temperature, salinity, and flow velocity of layers 1 to 33 on the (t + τ)-th day; In the sixth step, end.
2. The global ocean environment forecasting method based on multi-level feature aggregation according to claim 1, characterized in that The size of the patch divided by the two-dimensional convolutional neural network from the three-dimensional grid data is 6×6 pixels, the representation dimension of each patch is 416, the stride and the convolutional kernel size of the two-dimensional convolutional neural network are both 6; the stride and the convolutional kernel size of the two-dimensional transposed convolutional neural network layer are both 6, and a 6×6 transposed convolution with a stride of 6 is used.
3. The global ocean environmental forecasting method based on multi-level feature aggregation according to claim 1, characterized in that The method for aligning the time resolution of the ERA5 wind field reanalysis data with the GLORYS12 global ocean reanalysis data and the GHR sea surface temperature satellite data in step 2.2.1 is: 2.2.1.1 Let the variable d = 1, and initialize the ERA5 wind field reanalysis data set E1 after time resolution alignment to be empty; 2.2.1.2 Extract the data at specific time points from the ERA5 wind field reanalysis data on the d-th day in E, that is, select the data at 00, 06, 12, and 18 o'clock in a day, and take the mean of the sea surface wind field data at these four times as the daily average data of the wind field on the d-th day, and put the daily average data of the wind field on the d-th day into E1; 2.2.1.3 If d ≤ D, let d = d + 1, go to 2.2.1.2; if d > D, obtain the ERA5 wind field reanalysis data set E1 after time resolution alignment, go to 2.2.
2.
4. A method for global ocean environmental forecasting based on multi-level feature aggregation as described in claim 1, wherein the method for aligning the spatial resolutions of the two data sets M and E1 with G in step 2.2.2 is: 2.2.2.1 Let \(d = 1\), initialize the reanalysis dataset \(E2\) of ERA5 after spatial resolution alignment and the GHR sea surface temperature dataset \(G1\) to be empty; 2.2.2.2 Perform bilinear interpolation on the daily average wind field data and GHR sea surface temperature data on the \(d\)-th day in \(E1\) and \(G\) respectively, to obtain the wind field data on the \(d\)-th day with a spatial resolution of 1 / 12 degree, and put it into \(E2\); and obtain the sea surface temperature data on the \(d\)-th day with a spatial resolution of 1 / 12 degree, and put it into \(G1\); 2.2.2.3 If \(d\leq D\), let \(d = d + 1\), go to 2.2.2.2; if \(d > D\), obtain the ERA5 wind field reanalysis dataset \(E2\) and the GHR sea surface temperature dataset \(G1\) with a spatial resolution of 1 / 12 degree.
5. A global ocean environment forecasting method based on multi-level feature aggregation as claimed in claim 1, wherein the method for aligning the two datasets \(E2\) and \(G1\) with the spatial range of \(M\) in step 2.2.3 is: 2.2.3.1 Let \(d = 1\), initialize the set of ERA5 wind field data \(E3\) and GHR sea surface temperature data \(G2\) after spatial range alignment to be empty; 2.2.3.2 Select the part of the daily average wind field data and GHR sea surface temperature data on the \(d\)-th day in \(E2\) and \(G1\) from -180 degrees to 180 degrees in the longitude direction and from -80 degrees to 90 degrees in the latitude direction. Put the daily average wind field data on the \(d\)-th day in this spatial range into \(E3\); put the GHR sea surface temperature data on the \(d\)-th day in this spatial range into \(G2\); 2.2.3.3 If \(d\leq D\), let \(d = d + 1\), go to 2.2.3.2; if \(d > D\), obtain the ERA5 wind field reanalysis dataset \(E3\) and the GHR sea surface temperature dataset \(G2\) after spatial range alignment.
6. The global ocean environmental forecasting method based on multi-level feature aggregation according to claim 1, wherein The method for converting the data in \(E3\), \(G2\) and \(M\) from nc format to npy format in step 2.4 is: 2.4.1 Let the variable \(d = 1\), initialize the dataset \(S\) in npy format to be empty; 2.4.2 Read the GLORYS12 global ocean reanalysis data, ERA5 wind field reanalysis data, and GHR sea surface temperature data on the \(d\)-th day from \(M\), \(E3\) and \(G2\) respectively. Sequentially extract the data of the 1st to 33rd layers of the \(\theta o\), \(so\), \(uo\), \(vo\), \(zos\) elements in the GLORYS12 data, the \(u10\), \(v10\) elements in the ERA5 wind field data, and the \(analysed\_sst\) element in \(GHR\); then splice the above selected element data together, save it as an npy file, and then put it into \(S\); 2.4.3 If \(d\leq D\), let \(d = d + 1\), go to 2.4.2; if \(d > D\), the generation of npy data is completed, and a new dataset \(S\) is obtained.
7. The global ocean environment forecasting method based on multi-level feature aggregation according to claim 1, characterized in that 3.2 The method for setting the training parameters of the marine environmental forecasting system is as follows: Set the warm-up learning rate to 5×10 -8 , the warm-up training step size to 3, and the learning rate learning_rate to 5×10 -5 ; Set the hyperparameters of the model training optimizer: β1 to 0.9, β2 to 0.95, and "weight decay" to 1×10 -3 ; The batch size mini_batch_size for network training is 1; The maximum training step size maxepoch is 50.
8. The global ocean environmental forecasting method based on multi-level feature aggregation according to claim 1, wherein In step 3.3.2, \(B\) satisfies \(0\leq B\leq16\).
9. The global ocean environmental forecasting method based on multi-level feature aggregation according to claim 1, characterized in that When the window multi-head self-attention layer of the first local SIE network in step 3.3.4.1.1 performs hierarchical normalization on \(X\), \(X\) is divided into multiple \(7\times7\) windows, that is, each window contains \(7\times7\) patches.
Citation Information
Patent Citations
Marine fish image recognition method based on deep learning
CN111814881A
Edge sea high-temporal-spatial-resolution subsurface temperature field reconstruction method based on U-net framework
CN115114841A
Multi-modal image fusion method based on multi-scale feature extraction
CN116071282A
Three-dimensional ocean temperature and salt field forecasting method, system and equipment based on deep learning
CN116306318A