Red tide water body monitoring method based on improved Unit + + network
By applying an improved Unet++ network in red tide monitoring and combining drones to collect data, the existing red tide monitoring methods are solved, and red tide monitoring with high accuracy and high timeliness are achieved, and monitoring, early warning and governance of red tide disasters are supported.
Patent Information
- Application Number
- CN202510080995.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-16
AI Technical Summary
The existing red tide monitoring methods have high cost, limited monitoring range, poor timeliness, low recognition accuracy and rely on manual intervention, which is difficult to meet the actual use needs.
The red tide water body monitoring method based on the improved Unet++ network is adopted, and the water body information is collected through drones, and the image feature extraction and recognition is performed using the encoder and decoder of the improved Unet++ network. Combined with residual connection, spatial pyramid pooling module and extrusion excitation module, the network structure is optimized to improve the accuracy and timeliness of red tide type recognition.
It improves the accuracy and timeliness of red tide type identification, can supervise and predict the occurrence of red tides, reduces monitoring costs, and enhances the early warning capabilities for ecological disasters such as red tides.
Smart Images

Figure CN120014492A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of data processing for supervision and prediction, and in particular relates to a red tide water body monitoring method based on an improved Unet++ network. Background Art
[0002] Red tide is a harmful ecological phenomenon in which certain phytoplankton, protozoa or bacteria in seawater proliferate or aggregate explosively under certain environmental conditions, causing the water to change color. The occurrence of red tide has a significant impact on coastal fisheries and tourism, and timely monitoring is crucial.
[0003] At present, the red tide monitoring methods mainly include on-site monitoring, satellite remote sensing monitoring and ship-based monitoring. Although on-site monitoring can obtain more accurate red tide information, it is costly, has a limited monitoring range and poor timeliness. Satellite remote sensing monitoring can cover a larger sea area, but is limited by factors such as spatiotemporal resolution and cloud coverage. The accuracy of identifying small-scale and local red tide types is not ideal, and it is difficult to provide accurate and real-time monitoring. The current remote sensing monitoring algorithm mainly relies on the threshold segmentation method, and the determination of its threshold is limited by time and space, requiring manual intervention, making its large-scale application challenging. Ship-based monitoring also faces the problems of high cost and lack of flexibility. Therefore, the existing red tide water extraction methods are difficult to meet actual use needs. Summary of the invention
[0004] To solve the above problems, the present invention proposes a red tide water body monitoring method based on an improved Unet++ network, which overcomes the limitations of existing red tide monitoring methods, improves the accuracy and timeliness of red tide type identification, can supervise and predict the occurrence of red tides, and provide strong support for the monitoring, early warning and control of red tide disasters.
[0005] To achieve the above object, the present invention adopts the following technical solutions: A red tide water body extraction method based on an improved Unet++ network comprises the following steps: S1. Using a drone to collect information about the water body to obtain drone image data, and preprocessing the drone image data to obtain a data set; S2. Down-sample the input drone image using the encoder of the improved Unet++ network to extract image features; the encoder includes 5 encoding layers, the first encoding layer to the fourth encoding layer each include a convolution layer and a pooling layer, and the fifth encoding layer only includes a convolution layer but does not include a pooling layer; The specific process of step S2 is: S21, taking the data set as input, training the improved Unet++ network; the first encoding layer of the encoder of the improved Unet++ network adds the convolution output feature image and the input feature image to complete the residual mapping, and performs a squeeze excitation module operation on the added feature image to suppress channels of no interest; S22, after the convolution layer of the first encoding layer, the feature image size is 512×512, and the number of channels becomes 32; on the output result of the convolution layer, the 2×2 maximum pooling method is used to downsample the size of the output feature image to 256×256; S23, passing the feature image output by the first coding layer through the second coding layer, the size of the output feature image of the second coding layer is 128×128, and the number of channels is changed to 64; passing the feature image output by the second coding layer through the third coding layer, the size of the output feature image of the third coding layer is 64×64, and the number of channels is 128; passing the feature image output by the third coding layer through the fourth coding layer, the size of the output feature image of the fourth coding layer is 32×32, and the number of channels is changed to 256; passing the feature image output by the fourth coding layer through the fifth coding layer, the size of the output feature image of the fifth coding layer remains unchanged, and the number of channels is changed to 512; S3, using the spatial pyramid pooling module at the end of the encoder, by introducing multiple convolution operations with different dilation rates, to capture contextual information of different scales and generate feature images with different receptive fields; S4, using the decoder of the improved Unet++ network to gradually restore the feature image size output by the encoder to the original drone image size, and output the feature image recognition result; The decoder in step S4 includes four decoding layers and four corresponding upsampling modules; the upsampling module uses a 4×4 transposed convolution layer to perform an upsampling operation to restore the feature image size to twice the size of the input feature image, and the number of output channels is the number of channels of the previous level encoder; the decoding layer includes a Residual-SE module, the number of output channels is the number of encoder channels of the same resolution, and the input of each decoding layer includes the upsampled output from the previous layer decoder, the encoding layer output of the same resolution and all the decoding layer outputs of lower resolutions. Feature images at different levels are spliced in the channel dimension to form an input feature image containing multi-scale information; at the end of the decoder, a spatial pyramid pooling module is used, and a 1×1 convolution kernel is used to reduce the channel dimension of the feature image output by the last decoding layer to 2, and the Sigmoid function is used to convert the output value into the probability of whether it is an abnormal water body; then the argmax function is used to find the index of the maximum value on the specified dimension in the tensor, where the specified dimension is the dimension of the channel, and finally the pixel point with an index value of 1 is returned as the red tide area, and finally the red tide semantic segmentation result image is obtained; S5, for the training set in the data set of step S1, repeat steps S2 to S4 to train the improved Unet++ network to obtain a trained improved Unet++ network; S6. Input the test set into the trained improved Unet++ network to obtain the results of red tide water body anomaly recognition, which is used to monitor red tide water bodies.
[0006] Preferably, the specific process of preprocessing the drone image data in step S1 is: S11. Use Labelme annotation software to annotate the red tide water body in the drone image, and convert the annotation result into a binary image to obtain a data set; wherein the data set includes the drone image and the corresponding label; S12. Using the sliding window partitioning method, the drone images and labels were cropped into a size of 512×512 pixels with a repetition rate of 20%, and samples with background pixels accounting for more than 90% of the total number of single sample pixels were removed to reduce the impact of the imbalance of positive and negative samples on the classifier performance; S13. Divide the drone images and their corresponding labels into training set, validation set and test set according to the ratio of 7:2:1.
[0007] Preferably, in step S13, data amplification is performed on the divided training set by flipping, mirroring and rotating to expand the data volume of the training set.
[0008] Preferably, the specific process of step S21 is: S211, the encoder further includes a Residual-SE module, which includes two convolutional layers and a squeeze excitation module, the first convolutional layer is followed by a batch normalization layer and a ReLU nonlinear activation function, and the second convolutional layer is followed by a batch normalization layer, which is used to improve the training speed, stability and nonlinear ability of the improved Unet++ network; the step size of the convolutional layer is 1, and the convolution kernel size is 3×3; a convolution with a step size of 1 and a convolution kernel size of 1×1 is applied to the input feature image, and the output feature image adjusted by the 1×1 convolution is added to the feature image after the two layers of convolution to complete the residual mapping; S212, the output feature image obtained by residual mapping is subjected to a squeeze excitation module and a ReLU nonlinear activation function to convert the feature image into a channel descriptor; wherein the squeeze excitation module is subjected to global average pooling; S213, the channel descriptor first passes through the first fully connected layer to reduce the dimension from C to C / r, where r is the scaling factor and the activation function is ReLU; then passes through the second fully connected layer to restore the dimension to the original C, and generates the weight of each channel through the Sigmoid activation function; S214, multiplying the calculated channel weight coefficient by the input feature image to achieve recalibration of the feature image, so as to adjust the response of each channel according to the corresponding weight.
[0009] Preferably, the specific process of step S3 is: S31. A spatial pyramid pooling module is used at the end of the encoder. The spatial pyramid pooling module includes a 1×1 convolution, a global average pooling layer, and three ordinary convolution layers with dilation rates of 6, 12, and 18 respectively. S32, concatenating the convolution outputs of different expansion rates and the results of global average pooling to form an enhanced feature representation for integrating information from different receptive fields; S33. Use 1×1 convolution to integrate the spliced feature images, reduce the feature dimension, and output the final feature image.
[0010] Preferably, in step S4, a hybrid loss function is used to calculate the loss, and the calculation formula of the hybrid loss function is: , where represents the mixed loss function; () represents Dice coefficient loss; () represents Lovász Softmax loss; p represents the probability predicted by the improved Unet++ network; y represents the true label of the corresponding pixel.
[0011] After adopting the above technical solution, the present invention has the following beneficial effects: 1. The present invention optimizes the traditional Unet++ network by introducing residual connections, spatial pyramid pooling modules and squeeze excitation modules to obtain an improved Unet++ network, which solves the problem that the traditional model has poor adaptability to the characteristics of red tide water bodies. The traditional model is difficult to automatically adjust the feature extraction strategy, and the combination of multiple modules of the present invention can optimize network parameters according to data characteristics during training and recognition, and adapt to different scenarios. Compared with static models, it can flexibly focus on key features according to changes in task importance, such as large areas or local anomalies, to improve recognition accuracy. The improved Unet++ network of the present invention enhances the ability to extract red tide image features of different scales, morphologies and image quality changes, thereby improving the ability to extract red tide water bodies, and can monitor and predict the occurrence of red tides, providing strong support for the monitoring, early warning and governance of red tide disasters.
[0012] 2. Compared with traditional satellite remote sensing monitoring and on-site manual monitoring, the present invention can not only cover a wider range of sea areas by combining drone monitoring, but also provide real-time feedback of water body abnormal information, greatly reducing the monitoring cost, and is not affected by factors such as clouds. In addition, based on the flexibility of drones, data can be collected under different climatic conditions and time periods, effectively improving the comprehensiveness and richness of monitoring data, and enhancing the early warning capabilities of ecological disasters such as red tides.
[0013] 3. The present invention adopts a mixed loss function, which can effectively alleviate the imbalance problem of positive and negative samples, making the model more accurate in processing abnormal water bodies, reducing the risk of overfitting and improving classification accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 It is a schematic diagram of the process of the present invention; Figure 2 It is a process framework diagram of the present invention; Figure 3 It is a structural schematic diagram of image feature extraction of the present invention; Figure 4 It is a structural schematic diagram of the Residual-SE module of the present invention; Figure 5 It is a structural schematic diagram of the extrusion excitation module of the present invention; Figure 6 It is a structural schematic diagram of the spatial pyramid pooling module of the present invention; Figure 7 This is a diagram of the extraction results of the drone image of the present invention. DETAILED DESCRIPTION
[0015] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0016] like Figures 1 to 7 As shown, a red tide water body monitoring method based on an improved Unet++ network includes the following steps: S1. Using a drone to collect information about the water body to obtain drone image data, and preprocessing the drone image data to obtain a data set; The specific process of preprocessing the drone image data in step S1 is as follows: S11. Use Labelme annotation software to annotate the red tide water body in the drone image, and convert the annotation result into a binary image to obtain a data set; wherein the data set includes the drone image and the corresponding label; S12. Using the sliding window partitioning method, the drone images and labels were cropped into a size of 512×512 pixels with a repetition rate of 20%, and samples with background pixels accounting for more than 90% of the total number of single sample pixels were removed to reduce the impact of the imbalance of positive and negative samples on the classifier performance; S13, dividing the drone images and corresponding labels into a training set, a validation set, and a test set according to a ratio of 7:2:1; In step S13, the divided training set is subjected to data augmentation by flipping, mirroring and rotating to expand the data volume of the training set; S2. Down-sample the input drone image using the encoder of the improved Unet++ network to extract image features; the encoder includes 5 encoding layers, the first encoding layer to the fourth encoding layer each include a convolution layer and a pooling layer, and the fifth encoding layer only includes a convolution layer but does not include a pooling layer; The specific process of step S2 is: S21, taking the data set as input, training the improved Unet++ network; the first encoding layer of the encoder of the improved Unet++ network adds the convolution output feature image and the input feature image to complete the residual mapping, and performs a squeeze excitation module operation on the added feature image to suppress channels of no interest; The specific process of step S21 is: S211, the encoder further includes a Residual-SE module, which includes two convolutional layers and a squeeze excitation module, the first convolutional layer is followed by a batch normalization layer and a ReLU nonlinear activation function, and the second convolutional layer is followed by a batch normalization layer, which is used to improve the training speed, stability and nonlinear ability of the improved Unet++ network; the step size of the convolutional layer is 1, and the convolution kernel size is 3×3; a convolution with a step size of 1 and a convolution kernel size of 1×1 is applied to the input feature image, and the output feature image adjusted by the 1×1 convolution is added to the feature image after the two layers of convolution to complete the residual mapping; S212, the output feature image obtained by residual mapping is subjected to a squeeze excitation module and a ReLU nonlinear activation function to convert the feature image into a channel descriptor; wherein the squeeze excitation module is subjected to global average pooling; S213, the channel descriptor first passes through the first fully connected layer to reduce the dimension from C to C / r, where r is the scaling factor and the activation function is ReLU; then passes through the second fully connected layer to restore the dimension to the original C, and generates the weight of each channel through the Sigmoid activation function; S214, multiplying the calculated channel weight coefficient by the input feature image to achieve recalibration of the feature image, so as to adjust the response of each channel according to the corresponding weight; S22, after the convolution layer of the first encoding layer, the feature image size is 512×512, and the number of channels becomes 32; on the output result of the convolution layer, the 2×2 maximum pooling method is used to downsample the size of the output feature image to 256×256; S23, passing the feature image output by the first coding layer through the second coding layer, the size of the output feature image of the second coding layer is 128×128, and the number of channels is changed to 64; passing the feature image output by the second coding layer through the third coding layer, the size of the output feature image of the third coding layer is 64×64, and the number of channels is 128; passing the feature image output by the third coding layer through the fourth coding layer, the size of the output feature image of the fourth coding layer is 32×32, and the number of channels is changed to 256; passing the feature image output by the fourth coding layer through the fifth coding layer, the size of the output feature image of the fifth coding layer remains unchanged, and the number of channels is changed to 512; S3, using the spatial pyramid pooling module at the end of the encoder, by introducing multiple convolution operations with different dilation rates, to capture contextual information of different scales and generate feature images with different receptive fields; The specific process of step S3 is: S31. A spatial pyramid pooling module is used at the end of the encoder. The spatial pyramid pooling module includes a 1×1 convolution, a global average pooling layer, and three ordinary convolution layers with dilation rates of 6, 12, and 18 respectively. S32, concatenating the convolution outputs of different expansion rates and the results of global average pooling to form an enhanced feature representation for integrating information from different receptive fields; S33, integrating the spliced feature images using 1×1 convolution, reducing the feature dimension, and outputting the final feature image; S4, using the decoder of the improved Unet++ network to gradually restore the feature image size output by the encoder to the original drone image size, and output the feature image recognition result; The decoder in step S4 includes four decoding layers and four corresponding upsampling modules; the upsampling module uses a 4×4 transposed convolution layer to perform an upsampling operation to restore the feature image size to twice the size of the input feature image, and the number of output channels is the number of channels of the previous level encoder; the decoding layer includes a Residual-SE module, the number of output channels is the number of encoder channels of the same resolution, and the input of each decoding layer includes the upsampled output from the previous layer decoder, the encoding layer output of the same resolution and all the decoding layer outputs of lower resolutions. Feature images at different levels are spliced in the channel dimension to form an input feature image containing multi-scale information; at the end of the decoder, a spatial pyramid pooling module is used, and a 1×1 convolution kernel is used to reduce the channel dimension of the feature image output by the last decoding layer to 2, and the Sigmoid function is used to convert the output value into the probability of whether it is an abnormal water body; then the argmax function is used to find the index of the maximum value on the specified dimension in the tensor, where the specified dimension is the dimension of the channel, and finally the pixel point with an index value of 1 is returned as the red tide area, and finally the red tide semantic segmentation result image is obtained; In step S4, a hybrid loss function is used to calculate the loss. The calculation formula of the hybrid loss function is: , where represents the mixed loss function; () represents Dice coefficient loss; () represents Lovász Softmax loss; p represents the probability predicted by the improved Unet++ network; y represents the true label of the corresponding pixel; S5, for the training set in the data set of step S1, repeat steps S2 to S4 to train the improved Unet++ network to obtain a trained improved Unet++ network; S6. Input the test set into the trained improved Unet++ network to obtain the results of red tide water body anomaly recognition, which is used to monitor red tide water bodies.
[0017] Performance Test: In order to verify the effectiveness of the present invention, the segmentation results of the red tide water color anomaly in the test set were evaluated using four evaluation indicators: overall classification accuracy, F1 score, Miou and Kappa coefficient.
[0018] Overall Classification Accuracy refers to the ratio of the number of correctly predicted samples to the total number of samples. The calculation formula is: , where is the overall classification accuracy; TP is the number of correctly predicted water color abnormal areas; TN is the number of correctly predicted non-water color abnormal areas; FP is the number of areas predicted as water color abnormal areas that are actually non-water color abnormal areas; FN is the number of areas predicted as non-water color abnormal areas that are actually water color abnormal areas.
[0019] The F1 score is the harmonic mean of precision and recall.
[0020] Precision (P): , where P is the precision; Recall (R): , where R is the recall rate; Miou is the average of the ratio of the intersection to the union of each category. For each category i, , , where is the ratio of the intersection and union of the i-th category; TP i The number of water color anomaly areas predicted correctly for the i-th category; FP i FN is the number of areas predicted as abnormal water color areas for the i-th category but actually non-abnormal water color areas; i The number of areas predicted as non-abnormal water color areas for the i-th category that are actually abnormal water color areas; n is the number of categories.
[0021] The Kappa coefficient is an indicator to measure the consistency of classification. It takes into account the expectation of random consistency and is calculated as follows: , where is the proportion of the actual classification that is consistent with the predicted classification; is the proportion of agreement under random classification.
[0022] The test results are as follows: the overall classification accuracy, F1 score, Miou and Kappa coefficients are 0.923, 0.923, 0.833 and 0.850, respectively, which proves that the present invention can effectively extract red tides.
[0023] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed by the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A red tide water body monitoring method based on an improved Unet++ network, characterized in that: The following steps are involved: S1. Using a drone to collect information about the water body to obtain drone image data, and preprocessing the drone image data to obtain a data set; S2. Down-sample the input drone image using the encoder of the improved Unet++ network to extract image features; the encoder includes 5 encoding layers, the first encoding layer to the fourth encoding layer each include a convolution layer and a pooling layer, and the fifth encoding layer only includes a convolution layer but does not include a pooling layer; The specific process of step S2 is: S21, taking the data set as input, training the improved Unet++ network; the first encoding layer of the encoder of the improved Unet++ network adds the convolution output feature image and the input feature image to complete the residual mapping, and performs a squeeze excitation module operation on the added feature image to suppress channels of no interest; S22, after the convolution layer of the first encoding layer, the feature image size is 512×512, and the number of channels becomes 32; on the output result of the convolution layer, the 2×2 maximum pooling method is used to downsample the size of the output feature image to 256×256; S23, passing the feature image output by the first coding layer through the second coding layer, the size of the output feature image of the second coding layer is 128×128, and the number of channels is changed to 64; passing the feature image output by the second coding layer through the third coding layer, the size of the output feature image of the third coding layer is 64×64, and the number of channels is 128; passing the feature image output by the third coding layer through the fourth coding layer, the size of the output feature image of the fourth coding layer is 32×32, and the number of channels is changed to 256; passing the feature image output by the fourth coding layer through the fifth coding layer, the size of the output feature image of the fifth coding layer remains unchanged, and the number of channels is changed to 512; S3, using the spatial pyramid pooling module at the end of the encoder, by introducing multiple convolution operations with different dilation rates, to capture contextual information of different scales and generate feature images with different receptive fields; S4, using the decoder of the improved Unet++ network to gradually restore the feature image size output by the encoder to the original drone image size, and output the feature image recognition result; The decoder in step S4 includes four decoding layers and four corresponding upsampling modules; the upsampling module uses a 4×4 transposed convolution layer to perform an upsampling operation to restore the feature image size to twice the size of the input feature image, and the number of output channels is the number of channels of the previous level encoder; the decoding layer includes a Residual-SE module, the number of output channels is the number of encoder channels of the same resolution, and the input of each decoding layer includes the upsampled output from the previous layer decoder, the encoding layer output of the same resolution and all the decoding layer outputs of lower resolutions. Feature images at different levels are spliced in the channel dimension to form an input feature image containing multi-scale information; at the end of the decoder, a spatial pyramid pooling module is used, and a 1×1 convolution kernel is used to reduce the channel dimension of the feature image output by the last decoding layer to 2, and the Sigmoid function is used to convert the output value into the probability of whether it is an abnormal water body; then the argmax function is used to find the index of the maximum value on the specified dimension in the tensor, where the specified dimension is the dimension of the channel, and finally the pixel point with an index value of 1 is returned as the red tide area, and finally the red tide semantic segmentation result image is obtained; S5, for the training set in the data set of step S1, repeat steps S2 to S4 to train the improved Unet++ network to obtain a trained improved Unet++ network; S6. Input the test set into the trained improved Unet++ network to obtain the results of red tide water body anomaly recognition, which is used to monitor red tide water bodies.
2. A red tide water body monitoring method based on an improved Unet++ network as claimed in claim 1, characterized in that: The specific process of preprocessing the drone image data in step S1 is as follows: S11. Use Labelme annotation software to annotate the red tide water body in the drone image, and convert the annotation result into a binary image to obtain a data set; wherein the data set includes the drone image and the corresponding label; S12. Using the sliding window partitioning method, the drone images and labels were cropped into a size of 512×512 pixels with a repetition rate of 20%, and samples with background pixels accounting for more than 90% of the total number of single sample pixels were removed to reduce the impact of the imbalance of positive and negative samples on the classifier performance; S13. Divide the drone images and their corresponding labels into training set, validation set and test set according to the ratio of 7:2:
1.
3. A red tide water body monitoring method based on an improved Unet++ network as claimed in claim 2, characterized in that: In step S13, the divided training set is augmented by flipping, mirroring and rotating to increase the data volume of the training set.
4. A red tide water body monitoring method based on an improved Unet++ network as claimed in claim 1, characterized in that: The specific process of step S21 is: S211, the encoder further includes a Residual-SE module, which includes two convolutional layers and a squeeze excitation module, the first convolutional layer is followed by a batch normalization layer and a ReLU nonlinear activation function, and the second convolutional layer is followed by a batch normalization layer, which is used to improve the training speed, stability and nonlinear ability of the improved Unet++ network; the step size of the convolutional layer is 1, and the convolution kernel size is 3×3; a convolution with a step size of 1 and a convolution kernel size of 1×1 is applied to the input feature image, and the output feature image adjusted by the 1×1 convolution is added to the feature image after the two layers of convolution to complete the residual mapping; S212, the output feature image obtained by residual mapping is subjected to a squeeze excitation module and a ReLU nonlinear activation function to convert the feature image into a channel descriptor; wherein the squeeze excitation module is subjected to global average pooling; S213, channel descriptor first passes through the first fully connected layer to reduce the dimension from C to C / r, where r is the scaling factor and the activation function is ReLU; then passes through the second fully connected layer to restore the dimension to the original C, and generates the weight of each channel through the Sigmoid activation function; S214, multiplying the calculated channel weight coefficient by the input feature image to achieve recalibration of the feature image, so as to adjust the response of each channel according to the corresponding weight.
5. The red tide water body monitoring method based on the improved Unet++ network as claimed in claim 1, characterized in that: The specific process of step S3 is: S31. A spatial pyramid pooling module is used at the end of the encoder. The spatial pyramid pooling module includes a 1×1 convolution, a global average pooling layer, and three ordinary convolution layers with dilation rates of 6, 12, and 18 respectively. S32, concatenating the convolution outputs of different expansion rates and the results of global average pooling to form an enhanced feature representation for integrating information from different receptive fields; S33. Use 1×1 convolution to integrate the spliced feature images, reduce the feature dimension, and output the final feature image.
6. A red tide water body monitoring method based on an improved Unet++ network as claimed in claim 1, characterized in that: In step S4, a hybrid loss function is used to calculate the loss. The calculation formula of the hybrid loss function is: , where represents the mixed loss function; () represents Dice coefficient loss; () represents Lovász Softmax loss; p represents the probability predicted by the improved Unet++ network; y represents the true label of the corresponding pixel.