Model training method, mesoscale vortex recognition method and application
By improving the U-Net network and constructing a new loss function, and combining multi-source data and gradient difference constraints, the problem of insufficient generalization ability of the mesoscale eddy detection algorithm under complex sea conditions was solved, and higher detection accuracy and robustness were achieved.
Patent Information
- Application Number
- CN202511055540.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-07
AI Technical Summary
In existing technologies, the detection algorithms for mesoscale eddies lack generalization ability under complex sea conditions, and artificial intelligence methods suffer from low detection accuracy due to data defects.
By combining multi-source data, improving the U-Net network and Softmax classification layer, introducing ECA attention mechanism and residual units, constructing a new loss function, and adding gradient difference constraint terms, we can improve the feature extraction and recognition performance of the model.
It significantly improved the identification performance of mesoscale eddies, enhanced the model's adaptability to complex sea conditions, and improved detection accuracy.
Smart Images

Figure CN120913090A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of mesoscale eddy intelligent detection, and in particular to a model training method, a mesoscale eddy identification method and application. BACKGROUND
[0002] As a typical mesoscale dynamic process, the diameter of mesoscale eddy can be from tens of kilometers to hundreds of kilometers, and the life cycle usually lasts for several days to several months. Due to the obvious dynamic characteristics of mesoscale eddy, especially its huge kinetic energy (about more than 90% of the total kinetic energy of the global ocean), mesoscale eddy plays an important role in the material transport and energy conversion of the ocean. Therefore, mesoscale eddy has become one of the research focuses of oceanography. Mesoscale eddy can be divided into cyclonic eddy (CE) and anticyclonic eddy (AE). In the northern hemisphere, cyclonic eddy rotates counterclockwise, and the geostrophic effect produces a divergent flow field, resulting in the sea surface height inside the eddy being lower than that outside the eddy, so cyclonic eddy can be identified by the low anomaly of sea surface height (SSH). Similarly, anticyclonic eddy rotates clockwise in the northern hemisphere, producing a convergent flow field, resulting in the SSH inside the eddy being higher than that outside the eddy, so anticyclonic eddy can be identified by the high anomaly of SSH.
[0003] At present, the sea surface identification algorithm of mesoscale eddy can be roughly divided into two categories: traditional detection method and artificial intelligence detection method. The traditional method is mainly constructed by fluid dynamics theory, including Okubo-Weiss (OW) algorithm based on physical characteristics, winding angle (WA) algorithm based on geometric characteristics and vector geometry (VG) algorithm. Among them, the OW algorithm starts from the physical characteristics of mesoscale eddy, and detects eddy by parameterizing physical field information such as SSH field. The WA algorithm finds the eddy core in the sea level anomaly (SLA), and expands outward to find the closed stream line or sea level height contour, and the outermost closed stream line or sea level height contour is the boundary of the eddy. The VG algorithm intuitively defines mesoscale eddy as a region that meets certain constraints, and the feature of this region is that the velocity vector rotates clockwise or counterclockwise around the center, so as to realize the identification of eddy. However, the above traditional methods need to set threshold according to human experience, and there is a significant limitation in generalization ability, so they cannot adapt to complex sea conditions well.
[0004] In recent years, with the development of artificial intelligence methods, they have also been used in the detection of mesoscale eddies, such as the EddyNet algorithm based on the U-Net model, the EddyYolo algorithm for vortex multi-target detection, and the pyramid scene parsing network (PSPNet) algorithm for semantic and detail fusion, which can effectively identify sea surface eddies. The detection of mesoscale eddies based on SSH data is the most common method. Duo et al. designed and optimized an OEDNet target detection model, which takes the deep residual network and feature pyramid network as the main structure, and can automatically identify and locate eddies. Dong Ziyi et al. based on the sea level anomaly data obtained by satellite inversion, made a series of improvements to the U-Net network. They embedded a convolutional attention mechanism to enhance the attention of the feature extraction stage to the region with high class discrimination. However, since artificial intelligence is purely data-driven, if the data itself has defects, these defects will affect the detection results of mesoscale eddies, resulting in a decrease in the detection accuracy of mesoscale eddies.
[0005] Therefore, how to avoid the defects of artificial intelligence algorithms and improve the detection accuracy of mesoscale eddies is a problem to be solved at present. SUMMARY
[0006] In view of the above problems, the embodiments of the present application provide a model training method and a mesoscale eddy recognition method and device for improving the detection accuracy of an artificial intelligence model, so as to overcome the above problems or at least partially solve the above problems.
[0007] In a first aspect, the embodiments of the present application provide a model method, which comprises:
[0008] Obtaining multi-source sample data and pre-processing the multi-source sample data to obtain a sample mesoscale eddy dataset;
[0009] Inputting the sample mesoscale eddy dataset into a mesoscale eddy detection model to output a sample classification category of the mesoscale eddy;
[0010] Constructing a total loss function of the mesoscale eddy detection model, and determining a model total loss value based on the sample classification category;
[0011] Based on the model total loss value, adjusting the model parameters of the mesoscale eddy detection model to obtain a target mesoscale eddy detection model.
[0012] Further, the mesoscale eddy detection model comprises an input layer, an improved U-Net network and a Softmax classification layer; wherein:
[0013] The input layer is used to input the sample remote sensing data and the sample reanalysis data in the sample mesoscale eddy dataset into the improved U-Net network through the ECA attention mechanism channel respectively;
[0014] The improved U-Net network is used to extract feature information of input data, and to segment a mesoscale vortex and an anticyclonic vortex;
[0015] The Softmax classification layer is used to determine a predicted classification category of the mesoscale vortex according to the segmentation result, and the predicted classification category is used to indicate that the mesoscale vortex is a cyclonic vortex or an anticyclonic vortex.
[0016] Further, the improved U-Net network adopts an encoding-decoding structure, and a first residual unit is introduced in each of three down-sampling stages of an encoder, and a second residual unit is introduced in each of three up-sampling stages of a decoder, the first residual unit includes two groups of batch normalization layers, activation layers, dropout layers, convolution layers and residual connections, and a maximum pooling layer, and the second residual unit includes two groups of batch normalization layers, activation layers, dropout layers, convolution layers and residual connections; the decoder fuses the up-sampled feature maps and the convolution layer outputs in the corresponding down-sampling stages of the encoder through a connection module in each of the three up-sampling stages, and then inputs the fusions to the second residual unit.
[0017] Further, the inputting of the sample mesoscale vortex dataset into the mesoscale vortex detection model and the outputting of a sample classification category of the mesoscale vortex include:
[0018] The input layer respectively outputs remote sensing data weight feature maps and reanalysis data weight feature maps of the sample remote sensing data and the sample reanalysis data through the ECA attention mechanism channel, and obtains a first fused weight feature map by fusing the two;
[0019] The remote sensing data weight feature maps and the reanalysis data weight feature maps are simultaneously down-sampled three times in the encoder, and the remote sensing data down-sampled maps and the reanalysis data down-sampled maps generated by each down-sampling are fused to obtain an overall output feature map;
[0020] The overall output feature map is input into a Softmax classification layer to output a final predicted classification category of the mesoscale vortex.
[0021] Further, the constructing of a total loss function of the mesoscale vortex detection model and the determining of a model total loss value include:
[0022] According to the predicted classification category and the sample label, a Dice coefficient of a cyclonic vortex, an anticyclonic vortex and a non-vortex is determined respectively;
[0023] The Dice coefficients of the cyclonic vortex, the anticyclonic vortex and the non-vortex are weightedly averaged to obtain a weighted average Dice coefficient, and then 1 is subtracted from the weighted average Dice coefficient to obtain a Dice coefficient loss value;
[0024] A gradient loss constraint between the sample label and the predicted classification category is determined.
[0025] A total loss function is established based on the Dice coefficient loss and the gradient loss constraint, and a model total loss value is determined according to the total loss function:
[0026] loss = dice_coef_loss + a x gdiff loss
[0027] Wherein, loss represents the model total loss after adding the gradient difference constraint term, gdiff loss represents the gradient loss constraint, and a represents an adjustable weight.
[0028] In a second aspect, the embodiments of the present application provide a mesoscale vortex identification method, and the method comprises:
[0029] Obtaining an image to be detected;
[0030] Inputting the image to be detected into a target mesoscale vortex detection model to identify a mesoscale vortex classification category; wherein the target mesoscale vortex detection model is obtained by the method of any one of claims 1-5;
[0031] Based on the mesoscale vortex classification category, determining the mesoscale vortex category in the image to be detected.
[0032] In a third aspect, the embodiments of the present application provide a mesoscale vortex identification device, and the device comprises:
[0033] A data acquisition module is configured to acquire multi-source sample data, and pre-process the multi-source sample data to obtain a sample mesoscale vortex data set;
[0034] An input-output module is configured to input the sample mesoscale vortex data set into a mesoscale vortex detection model, and output a sample classification category of the mesoscale vortex;
[0035] A loss determination module is configured to construct a total loss function of the mesoscale vortex detection model, and determine a model total loss value;
[0036] A parameter adjustment module is configured to adjust model parameters of the mesoscale vortex detection model based on the model total loss value, to obtain a target mesoscale vortex detection model.
[0037] In a fourth aspect, the embodiments of the present application provide an electronic device, which comprises a memory, a processor, and a computer program stored in the memory, and the processor executes the computer program to implement the model training method or the mesoscale vortex identification method according to any one of the above.
[0038] In a fifth aspect, an embodiment of the present application provides a readable storage medium, the readable storage medium storing a program or instructions, the program or instructions being executed by a processor to implement the model training method or the mesoscale vortex identification method according to any one of the above.
[0039] Specific beneficial effects are that:
[0040] First, the present application uses data from different sources, such as remote sensing data and reanalysis data, to make up for the shortcomings of single-source data through multi-source data.
[0041] Second, the present application constructs an improved U-Net network model, which introduces a channel attention mechanism for the input of multi-source data, so as to adaptively pay attention to the importance of different channel features and amplify the feature information of the identified object to highlight its main factors. At the same time, multi-source data is fused at each level in the model to ensure that early data features are not lost.
[0042] Third, the present application constructs a new loss function, adds a gradient difference constraint term in the loss function, and further strengthens the training features, thereby further increasing the feature difference of the model.
[0043] In summary, the present application can significantly improve the recognition performance of the model for mesoscale vortices, and has important significance for intelligent detection of marine mesoscale vortices. BRIEF DESCRIPTION OF DRAWINGS
[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0045] Figure 1 is a flowchart of a model training method provided by an embodiment of the present application;
[0046] Figure 2a is a visualization diagram of original sea surface height data SSH in an input area, Figure 2b is a mesoscale vortex result diagram obtained by using the PET method and the Shapely tool;
[0047] Figure 3 is a U-Net model architecture diagram;
[0048] Figure 4 is a mesoscale vortex detection model architecture diagram;
[0049] Figure 5 is a structure diagram of an ECA module;
[0050] Figure 6 Training a model body module architecture;
[0051] Figure 7 Constructing a total loss function process;
[0052] Figure 8 Training a remote sensing data loss function curve without gradient constraints;
[0053] Figure 9 Training a model accuracy curve;
[0054] Figure 10 (a)-(l) are comparative results of mesoscale vortex detection and classification under different conditions; wherein, Figure 10 (a) is an SSH visualization graph (rs, weight a=0); Figure 10 (b) is a true value (rs, weight a=0); Figure 10 (c) is a model prediction result (rs, weight a=0); Figure 10 (d) is an SSH visualization graph (rs+ra, weight a=0); Figure 10 (e) is a true value (rs+ra, weight a=0); Figure 10 (f) is a model prediction result (rs+ra, weight a=0); Figure 10 (g) is an SSH visualization graph (rs, weight a=0.8); Figure 10 (h) is a true value (rs, weight a=0.8); Figure 10 (i) is a model prediction result (rs, weight a=0.8); Figure 10 (j) is an SSH visualization graph (rs+ra, weight a=0.6); Figure 10 (k) is a true value (rs+ra, weight a=0.6); Figure 10 (l) is a model prediction result (rs+ra, weight a=0.6). DETAILED DESCRIPTION
[0055] Exemplary embodiments of the present application will be described herein below with reference to the accompanying drawings. Although exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided so that the present application can be more thoroughly understood, and so that the scope of the present application can be accurately conveyed to those skilled in the art.
[0056] Embodiment One
[0057] Reference Figure 1 , Figure 1 A flowchart of a model training method provided by the embodiment of the present application, the model training method comprising the following steps:
[0058] Step 1, obtaining multi-source sample data, and preprocessing the multi-source sample data to obtain a mesoscale eddy dataset.
[0059] Optionally, the step 1 comprises the following sub-steps:
[0060] Sub-step 101: obtaining sample remote sensing data and sample reanalysis data respectively;
[0061] In order to avoid the shortcomings of single-source data, in the embodiments of the present application, multi-source sample data is obtained, which can include two types of remote sensing data and reanalysis data. Among them, the remote sensing data adopts Global Ocean Gridded Sea Level Anomalies, which comes from the Copernicus Marine Environment Monitoring Service Center (CMEMS), and the time range starts from January 1, 1993, and the spatial resolution is 1 / 4°; the reanalysis data adopts Global Ocean Reanalysis (GLORYS), which also comes from the Copernicus Marine Environment Monitoring Service Center, and the time range starts from January 1, 1993, and the spatial resolution is 1 / 12°. In the embodiments of the present application, for the two types of multi-source data, the selected range is the South China Sea region (104°E-122°E, 2°N-25°N), and the surface sea level height (SSH) data of this region is extracted. The sample remote sensing data is the satellite altimeter remote sensing data, and the sample reanalysis data is the global ocean reanalysis data.
[0062] Sub-step 102: labeling the eddy region for each sample remote sensing data obtained to obtain labeled sample remote sensing data;
[0063] First, in the embodiments of the present application, the PET (py-eddy-tracker) method is used to identify mesoscale eddies in each sample remote sensing data, preparing for the next step of making data set labels. The PET method is an eddy detection and tracking method proposed by Mason et al. in 2014, which is often used for label making of mesoscale eddy data sets. This method is based on SSH data to identify mesoscale eddies. First, the sea level anomaly SLA contour at 1 cm interval on the horizontal plane is calculated, then it is determined whether the contour is closed, and it needs to meet the following conditions: (1) the shape error of the mesoscale eddy is ≤55%, wherein the shape error is defined as the ratio of the area of the closed contour to the area of the fitted circle; (2) it is determined that the mesoscale eddy contains a certain number of pixels; (3) for the anti-gyre (gyre), only the pixels with SLA value higher (lower) than the current SLA interval value of the eddy are included; (4) for the anti-gyre (gyre), only one local maximum (minimum) is allowed; (5) the amplitude value range of the mesoscale eddy: 1 cm≤A≤150 cm.
[0064] Secondly, after obtaining the identification results of mesoscale eddies in the South China Sea region by the PET method, the Shapely tool is used to make vortex region labels. The Shapely package is a Python tool package specially used for two-dimensional plane image calculation. After obtaining the contours of the gyre and anti-gyre, the research area is converted to pixel values by the Shapely tool package, and each pixel value point is judged. If the pixel value point is within the gyre contour, it is marked as “2”. If the pixel value point is within the anti-gyre contour, it is marked as “1”, and the rest of the background and land area is marked as “0”. The vortex labeling effect is shown in Figures 2a-2b Figure 2a For the visualization of the original sea level height data SSH in the input area, Figure 2b For the mesoscale eddy data set obtained by using the PET method and the Shapely tool.
[0065] Sub-step 103: based on the labeled sample remote sensing data and the unlabeled sample reanalysis data, a sample mesoscale eddy data set is constructed, and the sample mesoscale eddy data set is divided into a training set and a test set.
[0066] After labeling the sample remote sensing data, such as finally labeling 4018 sea level height data graphs from 2010 to 2020, based on all the labeled sample remote sensing data and the unlabeled sample reanalysis data, a sample mesoscale eddy data set is constructed, and then the sample mesoscale eddy data set is divided into a training set and a test set according to the proportion, such as 3652 sea level height data graphs from 2010 to 2019 as the training set, and 366 sea level height data graphs in 2020 as the test set.
[0067] Step 2: input the mesoscale vortex dataset in the sample into the mesoscale vortex detection model, and output the predicted classification category of the mesoscale vortex;
[0068] The U-Net architecture is a classical image segmentation model based on a convolutional neural network (CNN), and its core features are a symmetrical U-shaped encoder-decoder architecture and a skip connection mechanism, so it benefits from the skip connection from the shrinkage path (downsampling) to the expansion path (upsampling), and can also incorporate the feature information of the early stage into the calculation, as shown in Figure 3 Therefore, the present application takes the U-Net model as the basic model for mesoscale vortex recognition, and optimizes the traditional U-Net model on this basis to obtain the mesoscale vortex detection model.
[0069] Firstly, the data source is expanded from the original single-source SSH data to multi-source SSH information of remote sensing and reanalysis data, and an attention mechanism is added for adjustment; secondly, the multi-source data is fused at each level to further improve the detection capability of the model for mesoscale vortex; finally, the gradient difference of the loss function is also trained to strengthen the feature extraction capability of the model. The architecture of the model is shown in Figure 4
[0070] The mesoscale vortex detection model comprises an input layer, an improved U-Net network and a Softmax classification layer, the input layer is used to obtain data in the sample mesoscale vortex dataset, and the input data is input to the improved U-Net network through the ECA attention mechanism channel, the improved U-Net network adopts an encoding-decoding structure, and a first residual unit is introduced in each of the three down-sampling stages of the encoder, and a second residual unit is introduced in each of the three up-sampling stages of the decoder, the first residual unit comprises two groups of batch normalization layers, activation layers, dropout layers, convolution layers (3x3) and residual connections, and a maximum pooling layer, and the second residual unit comprises two groups of batch normalization layers, activation layers, dropout layers, convolution layers (3x3) and residual connections; the decoder fuses the up-sampled feature maps in the three up-sampling stages with the convolution layer output in the corresponding down-sampling stage of the encoder through a connection module (contact operation) and then inputs the fused features to the second residual unit.
[0071] The ECA attention mechanism channel (as shown in Figure 4 the black arrow) is realized based on the ECA module, and is designed to enhance the expression capability of the model in different channels by weighting the features of different channels, as shown in Figure 5 The remote sensing data and the reanalysis data are input into the improved U-Net network as input data respectively through the ECA attention mechanism channel. Specifically, the remote sensing data or the reanalysis data first enter a one-dimensional convolution layer (1D Conv) for feature extraction and output attention weights, and then the output (attention weights) of the convolution layer is compressed to between 0 and 1 through a Sigmoid activation function, so as to be used for weighting the feature map, and finally a weight feature map is obtained. The weight feature map calculated through the ECA module is input into the improved U-Net network.
[0072] The data then enter the main part of the detection model. The improved U-Net network is obtained by adding a residual unit to the U-Net network. After introducing the residual unit, the model adds the feature information of the previous residual block to the next residual block, effectively avoiding the problems of gradient disappearance and vortex feature information loss caused by deep network while deepening the network depth. The design of the mesoscale vortex detection model still follows the 3-level full 32 filter architecture of the original U-Net model, which specifically includes three stages of encoding down-sampling paths, and each stage of the down-sampling path contains a convolution layer, an activation layer, a pooling layer and a dropout layer. In the up-sampling stage of the decoder, deconvolution (such as Figure 4 is used to restore the original resolution, and a total of three times are performed from right to left (blue arrow) as Figure 6 Figure 6 is shown, Figure 4 is the working process of the mesoscale vortex detection model, in which the training unit corresponds to the convolution, pooling, BN and other operations in
[0073] The encoder path has three stages, and the residual unit module of each stage mainly includes two groups of batch normalization (BN), ReLU activation function layer, Dropout layer and 3x3 convolution layer. After connecting with the input feature map (i.e. the weight feature map output by ECA) and calculating the residual, the input feature map resolution is halved through a 2x2 max pooling layer. Among them, the 3x3 convolution layer is mainly used to extract the feature information of the vortex input image, and the Dropout layer can ignore some features to avoid overfitting.
[0074] On the contrary, the decoder path also has three stages, and at each stage, the up-sampled feature map is first connected with the corresponding level of the encoder convolution layer through the Concatenate function, and then returned to the original resolution of the image before performing the corresponding feature extraction. Such an improvement further improves the semantic segmentation accuracy of the atmospheric vortex and anticyclonic vortex.
[0075] Optionally, step 2 includes the following sub-steps:
[0076] Sub-step 201: input the mesoscale vortex dataset in the sample into the input layer, and the sample remote sensing data and the sample reanalysis data enter two ECA attention mechanism channels respectively, output the remote sensing data weight feature map and the reanalysis data weight feature map, and perform convolution fusion on the remote sensing data weight feature map and the reanalysis data weight feature map to obtain a first fusion weight feature map (such as the black / grey block in Figure 4 , and connect the first fusion weight feature map with the remote sensing data weight feature map and the reanalysis data weight feature map again through the concat operation (such as the black plus sign in Figure 4 , to ensure that the original information will not be lost after convolution;
[0077] Sub-step 202: simultaneously perform three times of down-sampling on the remote sensing data weight feature map and the reanalysis data weight feature map in the encoder, and perform fusion on the remote sensing data down-sampling map and the reanalysis data down-sampling map generated after each time of down-sampling to obtain a first fusion feature map, a second fusion feature map and a third fusion feature map in sequence. Similarly, the first fusion feature map, the second fusion feature map and the third fusion feature map are connected with the corresponding remote sensing data down-sampling map and the reanalysis data down-sampling map respectively to ensure that the original information will not be lost;
[0078] Sub-step 203: as shown in the intermediate feature map in Figure 4 , the decoder performs up-sampling on the third fusion feature map, the second fusion feature map, the first fusion feature map and the first fusion weight feature map in reverse direction (blue arrow in Figure 4 ) through deconvolution to obtain a total output feature map;
[0079] Sub-step 204: input the total output feature map into the Softmax classification layer to output the final prediction classification category of the mesoscale vortex; the prediction classification category is used to indicate whether the mesoscale vortex is a cyclonic vortex or an anticyclonic vortex.
[0080] Step 3: construct a total loss function of the mesoscale vortex detection model, and determine a total loss value of the model.
[0081] Optionally, step 3 includes the following sub-steps:
[0082] Sub-step 301: calculate the Dice coefficients of the cyclonic vortex, the anticyclonic vortex and the non-vortex respectively according to the prediction classification category and the sample label;
[0083] The Dice coefficient is a commonly used index for measuring the similarity between two sets. In the image segmentation task, the Dice coefficient can evaluate the overlapping degree between the binary segmentation result output by the model and the actual segmentation result. The calculation formula of the Dice coefficient is shown in the following (1):
[0084]
[0085] where A represents the actual segmentation result pixel set, B represents the true segmentation result pixel, |A∩B| represents the number of pixels at the intersection of A and B sets, |A| and |B| represent the number of pixels in A and B, respectively.
[0086] The Dice coefficient value ranges from 0 to 1. The closer the value is to 1, the higher the degree of overlap between the model prediction result and the actual result, indicating better segmentation performance. In the training process of the deep learning model, the Dice coefficient is used as a loss function or evaluation index to help optimize the performance of the model in the task.
[0087] Sub-step 302: Weighted average the Dice coefficients of the cyclonic vortex, anti-cyclonic vortex and non-vortex to obtain a weighted average Dice coefficient, and then subtract 1 from the weighted average Dice coefficient to obtain a Dice coefficient loss value;
[0088] Sub-step 303: Calculate the gradient difference of the current mesoscale vortex segmentation result, and determine according to the gradient difference;
[0089] The gradient difference information of the loss function can provide additional calculation constraints for the recognition of mesoscale vortices, help to mine more feature information, and thus improve the accuracy and reliability of the recognition result. First, calculate the gradient difference of the current recognition result, and return two tensors representing the gradient difference along the y-axis and x-axis. Then square sum the two tensors to represent the sum of the gradients along the x-axis and y-axis, as shown in equation (2):
[0090] gdiff=dx×dx+dy×dy (2)
[0091] where dx represents the gradient difference along the x-axis, dy represents the gradient difference along the y-axis, and gdiff represents the sum of the gradients of the predicted classification categories along the x-axis and y-axis.
[0092] Sub-step 304: Calculate the gradient loss constraint between the sample label and the predicted classification category;
[0093] Then calculate the gradient loss constraint between the true label (gdiff_true) and the prediction result (gdiff_pred). The squared difference between the true label result and the prediction result needs to be calculated, and the average value is returned as the gradient difference constraint loss, as shown in equation (3).
[0094]
[0095] where gdiff true,i represents the loss gradient sum of the i-th true label, gdiff pred,i represents the i-th mesoscale vortex segmentation result, and N represents the total number of samples.
[0096] Sub-step 305: establishing a total loss function based on the Dice coefficient loss and the gradient loss constraint, and determining a model total loss value according to the total loss function;
[0097] The above two loss results are combined to obtain a total loss function of the mesoscale vortex detection model, as shown in formula (4).
[0098] loss = dice_coef_loss + a x gdiff loss (4)
[0099] Wherein, loss represents the model total loss after adding the gradient difference constraint term, gdiff loss represents the gradient loss constraint, and a represents the weight of the latter, which can be adjusted to obtain the best effect. The whole loss function structure diagram is shown in Figure 7 .
[0100] Step 4: adjusting the model parameters of the mesoscale vortex detection model based on the model total loss value to obtain a target mesoscale vortex detection model.
[0101] The model parameters are parameters in the improved U-Net model, including learning rate, batch size, weight decay, etc.
[0102] In order to avoid overfitting, the model total loss value on the test set is monitored, and when the validation set loss does not obviously decrease within 30 consecutive epochs, the training process will be automatically terminated, and the model with the best performance on the validation set is reserved as the target mesoscale vortex detection model. Alternatively, when the number of times of adjusting the model parameters is greater than or equal to the preset iteration number, the adjustment of the model parameters is stopped, and the model obtained after the last parameter adjustment is taken as the target mesoscale vortex detection model.
[0103] Simulation case
[0104] In order to verify the effect of the target mesoscale vortex detection model constructed by the application, the following ablation experiments are performed.
[0105] 1) Remote sensing data training result
[0106] The U-Net model with ECA but without gradient constraint is trained for input of pure remote sensing SSH data, and the results of 200 iteration cycles are shown in Table 1. It can be seen that when the cycle period is 200, the training loss and validation loss are small, and the highest classification accuracy is 0.9268, indicating that the network model meets the requirements of intelligent identification of mesoscale vortex. After training, the model's curve is used to visualize the loss value change of the model during training. By observing the trend of training loss and validation loss, we can understand the training state of the model, and whether there are problems such as overfitting or underfitting. The Matplotlib library is used to plot the loss value graph within the training period, and the title and x and y axes are set on the graph; the training loss (train_loss) and validation loss (val_loss) curves are plotted on the y axis in a semi-log scale, and the training time is plotted on the x axis. The loss function curve is shown in Figure 8
[0107] Table 1. Training results of optimized U-Net architecture for remote sensing SSH
[0108]
[0109] 2) Training results of loss function gradient difference constraint
[0110] In order to verify that the performance of the model after adding the gradient difference constraint is improved, multiple comparison tests are established to analyze the results. For different gradient difference constraint weights, the corresponding training results are analyzed and compared to obtain the best constraint condition. Therefore, the weight value of the model with gradient difference constraint is determined respectively, and the weight a is an arithmetic sequence from 0.1 to 1.0 with a difference of 0.1. The model is trained for 200 cycles, and the results shown in Table 2 are obtained.
[0111] Compared with the original model, the model with gradient difference constraint has significantly improved accuracy in identifying mesoscale vortex when increasing the weight a of gradient loss constraint, which also shows that further feature extraction through the gradient difference of the loss function enables the model to have high precision, low loss, and good generalization ability, so that better results can be achieved in the identification of mesoscale vortex. Especially when the weight a of gradient loss constraint is 0.8, the performance of the model recognition reaches the current optimal level. At the same time, the Dice coefficient is high, all exceeding 0.8, which also shows that the segmentation results and model performance of the model are very good.
[0112] However, it is also found that after adding the gradient difference constraint term, the convergence speed of the entire model becomes slower and slower, especially when a takes the values of 0.9 and 1.0, the entire model no longer converges, and the model automatically ends training at step 101, unable to detect mesoscale vortex, at which time the experiment fails.
[0113] Table 2. Training results of the optimized U-Net architecture for pure remote sensing SSH
[0114]
[0115] 3) Training results with loss function gradient difference constraint combined with multi-source data
[0116] Since deep learning models are data-driven, if only one type of data is used, the training results will be affected if the data itself has defects. Therefore, if the input data comes from different sources, the defects of single-source data can be compensated.
[0117] In this experiment, the SSH part of satellite altimeter remote sensing data and global ocean reanalysis data was extracted for data fusion. After convolution with down-sampling of the two types of data, fusion and training were performed at each layer of the model. At the same time, multiple comparison experiments were also established to analyze the results: the gradient difference constraint term was still added, and the weight value with the gradient difference constraint was determined respectively, the weight a was determined as an arithmetic sequence from 0.1 to 1.0 with a difference of 0.1, and the model was trained for 200 cycles to obtain the results shown in Table 3. Like the experimental results in 2.2, the model with the gradient difference constraint significantly improves the accuracy of identifying mesoscale eddies as the weight a of the gradient loss constraint increases, and the addition of multi-source data makes the data more reliable, so the identification results of mesoscale eddies can further achieve better results compared to the experimental results in 2.2.
[0118] As can be seen from Table 3, as the weight a increases, the identification accuracy also further improves. When a is 0.6 and 0.7, the training classification accuracy and validation classification accuracy of the model are the highest, which are 94.46% and 93.88% respectively, which have greatly improved compared to the original model. At the same time, the Dice coefficient is also high, still all above 0.8, which also shows that the segmentation results and model performance of the model are very good. In addition, although the gradient difference constraint term is added in the multi-source data model, the convergence speed of the entire model becomes slower and slower, but when a takes the value of 0.9 and 1.0, the entire model can still continue to converge, but the training classification accuracy and validation classification accuracy have decreased, indicating that the model trained at this time has begun to overfit, and is not the best proportion configuration.
[0119] Table 3. Training results of the optimized U-Net architecture for remote sensing SSH + reanalysis SSH data
[0120]
[0121] The training results of Tables 1-3 are summarized to obtain the model training accuracy curve as shown in Figure 9 Table 4.Figure 9 In the figure, rs represents remote sensing data, ra represents reanalysis data, and the abscissa represents the weight a of the gradient loss constraint. It can be seen from the figure that for the same model architecture, the accuracy of the multi-source data rs+ra is higher than that of the single-source data rs in the training set and the validation set. At the same time, the curve shows an overall upward trend, indicating that after increasing the weight of the gradient loss constraint, the accuracy of the model training will also increase. When the weight a is 0.6 and 0.7, the training classification accuracy and the validation classification accuracy of the model are the highest, which is the optimal configuration.
[0122] Figure 10 The comparison figure of the detection results of the four groups of mesoscale eddies is shown in the figure. In the figure, (a, b, c) are the SSH visualization, real label value and model prediction result of the single-source data rs and the model without gradient difference constraint, (d, e, f) are the SSH visualization, real label value and model prediction result of the multi-source data rs+ra and the model without gradient difference constraint, (g, h, i) are the SSH visualization, real label value and model prediction result of the single-source data rs and the model with gradient difference constraint weight a=0.8, and (j, k, l) are the SSH visualization, real label value and model prediction result of the multi-source data rs+ra and the model with gradient difference constraint weight a=0.6. By comparing the detection classification figures, it can be seen that the model prediction of (c) and (f) does not fully extract the image features due to the lack of gradient difference constraint, and misses some eddies and loses some accuracy. The gradient difference of the loss function is increased in (i) and (l), and more eddies are detected based on the data. Obviously, the model with loss gradient constraint has higher consistency between the prediction figure and the real label figure of the mesoscale eddy identification result, can better fit the mesoscale eddy characteristics, and make highly accurate prediction. This shows that due to the addition of the gradient difference constraint in the loss function, the training effect of the intelligent model is better, so adding the gradient difference constraint of the loss function in a certain range can significantly improve the recognition performance of the model.
[0123] In order to further verify the characteristics of the model, it is also compared with the classical encoding and decoding network U-Net, SegNet and pyramid scene analysis network PSPNet. The experimental results are shown in Table 4. It can be seen that compared with other semantic segmentation models, the model proposed in the present application shows certain advantages in eddy detection and classification effect.
[0124] Table 4. Comparison of mesoscale eddy detection accuracy of different models
[0125]
[0126] Example Two
[0127] The mesoscale eddy identification method provided in the present embodiment can comprise:
[0128] Step 1: obtaining an image to be detected;
[0129] The image to be detected refers to one of a remote sensing image or a reanalysis image;
[0130] Step 2: inputting the image to be detected into a target mesoscale vortex detection model, and outputting a mesoscale vortex classification category; the target mesoscale vortex detection model is obtained according to the model training method described above;
[0131] Step 3: determining a mesoscale vortex category in the image to be detected based on the mesoscale vortex classification category.
[0132] In the embodiments of the present application, the mesoscale vortex classification category can be displayed in the output image of the target mesoscale vortex detection model, the mesoscale vortex position can be identified using bright colors, and the first classification category corresponding to the mesoscale vortex can be identified using different numbers or different colors. For example, 1 can represent an anticyclonic vortex, and 2 can represent a cyclonic vortex. On this basis, the mesoscale vortex category in the image to be detected can be determined according to the first classification category.
[0133] Embodiment three
[0134] The model training device provided in the embodiments of the present application can include:
[0135] The data acquisition module is configured to acquire multi-source sample data, and pre-process the multi-source sample data to obtain a sample mesoscale vortex data set;
[0136] The input and output module is configured to input the sample mesoscale vortex data set into a mesoscale vortex detection model, and output a sample classification category of the mesoscale vortex;
[0137] The loss determination module is configured to construct a total loss function of the mesoscale vortex detection model, and determine a model total loss value;
[0138] The parameter adjustment module is configured to adjust model parameters of the mesoscale vortex detection model based on the model total loss value, and obtain a target mesoscale vortex detection model.
[0139] The apparatus in the embodiments of the present application can be an electronic device or a component in an electronic device, for example, an integrated circuit or a chip. The electronic device can be a terminal or other device than a terminal. For example, the electronic device can be a GPU BOX, a mobile phone, a tablet computer, a notebook computer, a palm computer, a vehicle-mounted electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), and the like. The electronic device can also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, a self-service machine, and the like. The embodiments of the present application are not limited in this regard.
[0140] The apparatus in the embodiments of the present application can be an electronic device or a component in an electronic device, for example, an integrated circuit or a chip. The electronic device can be a terminal or other device than a terminal. For example, the electronic device can be a GPU BOX, a mobile phone, a tablet computer, a notebook computer, a palm computer, a vehicle-mounted electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), and the like. The electronic device can also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, a self-service machine, and the like. The embodiments of the present application are not limited in this regard.
[0141] The embodiments of the present application also provide a computer readable storage medium, which stores a computer program / instruction. The computer program / instruction is executed by a processor to implement the steps of the mesoscale vortex identification method or the model training method disclosed in the embodiments of the present application.
[0142] The embodiments of the present application also provide a computer program product. The computer program product is run on an electronic device to enable the processor to implement the steps of the mesoscale vortex identification method or the model training method disclosed in the embodiments of the present application.
[0143] The embodiments of the present application also provide an electronic device, which includes a memory, a processor, and a computer program stored in the memory. The processor executes the computer program to implement the model training method or the mesoscale vortex identification method according to any one of the above embodiments.
[0144] Each of the embodiments in the present specification is described in a progressive manner, and each embodiment focuses on the difference from other embodiments. The same or similar parts between the embodiments can be referred to each other.
[0145] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, apparatuses, electronic device, and computer program products according to the embodiments of the present application. It is understood that each flow and / or block in the flowcharts and / or block diagrams, and a combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminals to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminals generate a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks. Figure 1 an apparatus that implements the functions specified in the flowcharts and / or block diagrams.
[0146] These computer program instructions can also be stored in a computer-readable memory that can direct the computer or other programmable data processing terminals to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including instruction apparatus, which implements the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks. Figure 1 an apparatus that implements the functions specified in the flowcharts and / or block diagrams.
[0147] These computer program instructions can also be loaded into a computer or other programmable data processing terminal, so that a series of operation steps are performed on the computer or other programmable terminal to produce a computer-implemented process, so that the instructions executed on the computer or other programmable terminal provide a process for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks. Figure 1 an apparatus that implements the functions specified in the flowcharts and / or block diagrams.
[0148] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to the embodiments once they know the basic inventive concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present application.
[0149] Finally, it is to be understood that the phraseology or terminology such as "first" and "second" etc. used herein is merely intended to differentiate one entity or operation from another entity or operation, without necessarily requiring or implying any actual such relationship or order between such entities or operations. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0150] The above describes in detail a model training method, a mesoscale vortex identification method and device provided by the present application. The principles and implementation manners of the present application are described by using specific examples. The above description of the embodiments is only used to help understand the method and core idea of the present application. Meanwhile, for those skilled in the art, the specific implementation manners and application ranges can be changed according to the idea of the present application. In summary, the content of the specification should not be understood as a limitation of the present application.
Claims
1. A model training method, characterized in that, The method comprises: acquiring multi-source sample data, and preprocessing the multi-source sample data to obtain a sample mesoscale vortex data set; inputting the sample mesoscale vortex data set into a mesoscale vortex detection model to output a sample classification category of the mesoscale vortex; constructing a total loss function of the mesoscale vortex detection model, and determining a model total loss value based on the sample classification category; adjusting model parameters of the mesoscale vortex detection model based on the model total loss value to obtain a target mesoscale vortex detection model.
2. The method of claim 1, wherein, The mesoscale vortex detection model comprises an input layer, an improved U-Net network and a Softmax classification layer; wherein: the input layer is used for inputting sample remote sensing data and sample reanalysis data in the sample mesoscale vortex data set into the improved U-Net network through an ECA attention mechanism channel respectively; the improved U-Net network is used for extracting feature information of the input data, and segmenting a cyclonic vortex and an anticyclonic vortex; the Softmax classification layer is used for determining a predicted classification category of the mesoscale vortex according to a segmentation result, and the predicted classification category is used for indicating that the mesoscale vortex is a cyclonic vortex or an anticyclonic vortex.
3. The method of claim 2, wherein, The improved U-Net network adopts an encoding-decoding structure, and a first residual unit is introduced in each of three down-sampling stages of an encoder, and a second residual unit is introduced in each of three up-sampling stages of a decoder; the first residual unit comprises two groups of batch normalization layers, an activation layer, a dropout layer, a convolution layer, residual connection and a maximum pooling layer, and the second residual unit comprises two groups of batch normalization layers, an activation layer, a dropout layer and a convolution layer; in each of the three up-sampling stages, the decoder fuses the up-sampled feature map and the convolution layer output in the corresponding down-sampling stage of the encoder through a connection module, and then inputs the up-sampled feature map and the convolution layer output into the second residual unit.
4. The method of claim 3, wherein, The input layer outputs remote sensing data weight feature maps and reanalysis data weight feature maps of the sample remote sensing data and the sample reanalysis data through the ECA attention mechanism channel respectively, fuses the two to obtain a first fused weight feature map; the remote sensing data weight feature maps and the reanalysis data weight feature maps are down-sampled three times in the encoder, and the remote sensing data down-sampled maps and the reanalysis data down-sampled maps generated in each down-sampling are fused to obtain an overall output feature map; the overall output feature map is input into the Softmax classification layer to output a final predicted classification category of the mesoscale vortex. The method comprises:
5. The method of claim 4, wherein, determining a Dice coefficient of a cyclonic vortex, an anticyclonic vortex and a non-vortex according to the predicted classification category and a sample label respectively; weighting and averaging the Dice coefficients of the cyclonic vortex, the anticyclonic vortex and the non-vortex to obtain a weighted average Dice coefficient, and then subtracting the weighted average Dice coefficient from 1 to obtain a Dice coefficient loss value; determining a gradient loss constraint between the sample label and the predicted classification category; A total loss function is established based on a Dice coefficient loss and a gradient loss constraint, and a model total loss value is determined according to the total loss function: loss = dice_coef_loss + a x gdiff loss wherein loss represents the total loss of the model after adding the gradient difference constraint term, gdiff loss represents the gradient loss constraint, and a represents an adjustable weight.
6. A mesoscale vortex identification method, characterized by, The method comprises: An image to be detected is acquired; The image to be detected is input into a target mesoscale vortex detection model to identify a mesoscale vortex classification category; wherein the target mesoscale vortex detection model is obtained by the method of any one of claims 1-5; Based on the mesoscale vortex classification category, a mesoscale vortex category in the image to be detected is determined.
7. A model training apparatus characterized by comprising: The device comprises: A data acquisition module is configured to acquire multi-source sample data, and pre-process the multi-source sample data to obtain a sample mesoscale vortex dataset; An input-output module is configured to input the sample mesoscale vortex dataset into a mesoscale vortex detection model, and output a sample classification category of the mesoscale vortex; A loss determination module is configured to construct a total loss function of the mesoscale vortex detection model, and determine a model total loss value; A parameter adjustment module is configured to adjust model parameters of the mesoscale vortex detection model based on the model total loss value, to obtain a target mesoscale vortex detection model.
8. A mesoscale vortex identification apparatus, characterized by, The device comprises: An acquisition module is configured to acquire an image to be detected; An identification module is configured to input the image to be detected into a target mesoscale vortex detection model to identify a mesoscale vortex classification category; wherein the target mesoscale vortex detection model is obtained by the model training method of any one of claims 1-5; An output module is configured to determine a mesoscale vortex category in the image to be detected based on the mesoscale vortex classification category.
9. An electronic device, comprising: A computer program product comprises a memory, a processor, and a computer program stored on the memory, wherein the processor executes the computer program to implement the model training method or the mesoscale vortex identification method of any one of the preceding claims.
10. A readable storage medium, characterized by, A readable storage medium stores a program or instructions, which are executed by a processor to implement the model training method or the mesoscale vortex identification method of any one of the preceding claims.