Coastal zone marine environment pollution intelligent monitoring method based on deep learning
The coastal marine environment pollution monitoring method improves detection accuracy and adaptability by employing a YOLOv11 network with mixed-channel spatial attention and auxiliary enhancement modules, optimizing the model with a geometric-sensitive loss function for robust pollution identification across varying conditions and scales.
Patent Information
- Application Number
- CN202510374295.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-15
AI Technical Summary
The existing deep learning-based coastal marine environmental pollution monitoring method has defects in accuracy and average accuracy, cannot adapt to different lighting and weather conditions, and is not effective in multi-scale object detection.
The coastal marine environmental pollution monitoring model based on the YOLOv11 network is adopted, combining the hybrid channel space attention module, split-intensive multi-branch module and split-cascaded dense multi-branch module, an auxiliary enhancement detection network is added, and the training effect is optimized using geometrically sensitive auxiliary loss function to improve the robustness and detection accuracy of the model.
High-precision detection under complex backgrounds and multi-scale targets is achieved, which enhances the adaptability of the model and the ability to identify small targets, and improves the accuracy and robustness of the detection.
Smart Images

Figure CN120318678A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of unmanned intelligent monitoring of the coastal marine environment, and particularly relates to an intelligent monitoring method for coastal marine environmental pollution based on deep learning. Background Art
[0002] As one of the countries with the longest coastlines in the world, China has vast coastal zone resources, spanning multiple seas such as the East China Sea, the Yellow Sea, the Bohai Sea, and the South China Sea. The coastal zone is not only an important fishery resource area in China but also contains rich tourism resources. Numerous natural landscapes such as bays, beaches, and islands attract a large number of tourists, driving the rapid growth of the coastal economy. However, with the rapid development of China's economy in recent years, the environment of the coastal zone is facing increasingly severe challenges. The pollution problem in the nearshore marine area is becoming increasingly prominent, seriously affecting the balance of the marine ecosystem. Therefore, it is very necessary to monitor the marine environmental pollution situation around the coastal zone, detect pollution in a timely manner, and dispose of it.
[0003] Traditional object detection algorithms usually rely on manually designed features. These features usually require the experience of domain experts and cannot adaptively learn the optimal features from the data. In recent years, deep learning-based object detection algorithms, including R-CNN, SSD, and YOLO series, have become leading algorithms in the field of object detection. Through the training of deep networks, more high-dimensional features can be captured to adapt to different changes in illumination, scale, angle, etc. Therefore, the deep learning method has surpassed the traditional method in the accuracy of object detection, especially in large-scale datasets.
[0004] Although the above deep learning-based object detection methods can achieve the monitoring of coastal marine environmental pollution, their performance has certain defects in terms of accuracy and mean average precision. Summary of the Invention
[0005] The purpose of the present invention is to overcome the defects of the prior art and propose an intelligent monitoring method for coastal marine environmental pollution based on deep learning.
[0006] In view of this, the present invention proposes an intelligent monitoring method for coastal marine environmental pollution based on deep learning. The method includes:
[0007] Preprocess the collected coastal zone images;
[0008] Input the preprocessed coastal zone images into a pre-trained coastal marine environmental pollution monitoring model to obtain the location and category of the pollution;
[0009] The coastal zone marine environmental pollution monitoring model is based on the YOLOv11 network, including a backbone network, a neck network, and a head network. Among them, in the backbone network, a shallow information perception module with a hybrid channel spatial attention module is used to extract key features of coastal zone images and obtain key feature maps. In the neck network, a split dense multi-branch module and a split cascaded dense multi-branch module are used to further fuse and enhance the key features.
[0010] During the training phase of the coastal zone marine environmental pollution monitoring model, an auxiliary enhancement detection network is added to the head network, and a geometric sensitivity auxiliary loss function is used to optimize the training effect.
[0011] Preferably, the preprocessing includes: adjusting the image size and normalizing the pixel values of the image.
[0012] Preferably, the processing process of the shallow information perception module includes:
[0013] Performing convolution processing on the input features to obtain a feature map, splitting it along the channel into a first feature map and a second feature map, passing the second feature map through two hybrid channel spatial attention modules with the same structure, then sequentially splicing the first feature map, the second feature map, the output of the first hybrid channel spatial attention module, and the output of the second hybrid channel attention module along the channel, and outputting the splicing result after one convolution.
[0014] Preferably, the hybrid channel spatial attention module includes: a hybrid local channel attention module and a spatial attention module. The processing process of the hybrid channel spatial attention module includes: sequentially passing the input features through the hybrid local channel attention module and the spatial attention module, and then element-wise adding the input features and the output of the spatial attention module.
[0015] The processing process of the hybrid local channel attention module includes:
[0016] Performing two convolution operations on a feature tensor with an input shape of (C, W, H) and then performing a local average pooling operation to convert it into a feature vector of size (C, K s , K s ) to extract local spatial information, where C is the number of channels, W is the width of the feature map, H is the height of the feature map, K s is the length of the feature map after the local average pooling operation, and K s ×K s is the size of the feature map after local average pooling;
[0017] Sending the local spatial information into the global information branch and the local information branch respectively to calculate their respective attention weights;
[0018] The attention weights output from the two branches are added element-wise and then de-pooled to restore the hybrid local channel attention weights with the shape of (C, W, H).
[0019] Multiply the input feature element-wise with the hybrid local channel attention weights.
[0020] Preferably, the processing process of the global information branch includes:
[0021] The input feature is globally average pooled into a one-dimensional feature, and then the channel dependence relationship under the global information is extracted through one-dimensional convolution to obtain the channel attention weights under the global information. The channel attention weights under the global information are restored to the same size as the input feature through de-pooling operation.
[0022] The processing process of the local information branch includes:
[0023] Reshape the input feature into a one-dimensional feature vector, and then through one-dimensional convolution, reshape the result after convolution into the same size as the input feature.
[0024] Preferably, the processing process of the spatial attention module includes:
[0025] Perform max pooling and average pooling on the input feature vector respectively to obtain two feature maps, concatenate the two feature maps in the channel dimension, perform convolution on the concatenated feature vector and use the Sigmoid activation function for normalization processing to obtain the spatial attention map; multiply the input feature vector element-wise with the spatial attention map.
[0026] Preferably, the processing process of the split dense multi-branch module includes:
[0027] Perform convolution processing on the input feature to obtain a feature map, split it along the channel into a third feature map and a fourth feature map, pass the fourth feature map through two aggregation dense layer modules with the same structure, and then concatenate the third feature map, the fourth feature map, the output of the first aggregation dense layer module and the output of the second aggregation dense layer module in order along the channel, and output the concatenated result after one convolution;
[0028] The processing process of the aggregation dense layer module includes:
[0029] The feature obtained by continuously performing two convolutions on the input feature is copied into two paths. One path of the feature is sent to four parallel fully connected layers after average pooling, and then the four parallel fully connected layers are concatenated in order along the channel. After passing through one fully connected layer, it is multiplied element-wise with the other path of the feature.
[0030] Preferably, the processing process of the split cascaded dense multi-branch module includes:
[0031] Perform convolution processing on the input features, split the feature map obtained by convolution into a fifth feature map and a sixth feature map along the channels, pass the sixth feature map through two cascaded aggregation dense layer modules, and then concatenate the fifth feature map, the sixth feature map, the output of the first cascaded aggregation dense layer module, and the output of the second cascaded aggregation dense layer module in sequence along the channels, and output after performing convolution on the concatenation result;
[0032] The processing process of the cascaded aggregation dense layer module includes:
[0033] Send the input features into two branches respectively. The first branch passes the input features through convolution and two aggregation dense layer modules in sequence. The second branch performs a convolution operation, and then concatenates the outputs of the first branch and the second branch in sequence along the channels, and outputs after performing convolution on the concatenated result.
[0034] Preferably, in the training stage of the coastal zone marine environmental pollution monitoring model, an auxiliary enhancement detection network is added to the head network. The auxiliary enhancement detection network includes three auxiliary detection heads with the same structure, which are used to perform additional classification and regression tasks on the intermediate feature layer to enhance the regression ability of the model; the auxiliary detection head includes: an auxiliary head regression branch and an auxiliary head classification branch.
[0035] Preferably, the geometric sensitivity auxiliary loss function L is:
[0036] L=(L cls +L siou +L dfl )(1 + λ)
[0037] Among them, L cls is the classification loss, L siou is the bounding box loss, L dfl is the distribution focal loss, and λ is the auxiliary loss weight coefficient;
[0038] The bounding box loss L siou is the SCYLLA loss with angle loss, distance loss, shape loss, and intersection over union loss between the predicted box and the ground truth box, and satisfies the following formula:
[0039]
[0040] Among them, L IOU is the intersection over union loss, Δ is the distance loss, and Ω is the shape loss;
[0041] The angle loss Λ satisfies the following formula:
[0042]
[0043] Among them, σ is the distance between the center point of the predicted box and the center point of the ground truth box, ch The vertical distance between the center point of the predicted box and the center point of the ground truth box
[0044] The distance loss Δ satisfies the following formula:
[0045]
[0046] where ρ x is the normalized squared difference between the abscissa of the center point of the predicted box and the abscissa of the center point of the ground truth box, and ρ y is the normalized squared difference between the ordinate of the center point of the predicted box and the ordinate of the center point of the ground truth box, and γ = 2 - Λ;
[0047] The shape loss Ω satisfies the following formula:
[0048]
[0049] where ω w is the normalized difference between the width of the predicted box and the width of the ground truth box, and ω h is the normalized difference between the height of the predicted box and the height of the ground truth box, and θ represents the degree of attention to the shape loss;
[0050] The intersection over union loss IoU between the predicted box and the ground truth box is:
[0051]
[0052] where |B ∩ B GT | is the intersection part of the predicted box and the ground truth box, and |B ∪ B GT | is the union part of the predicted box and the ground truth box.
[0053] Compared with the prior art, the advantages of the present invention are as follows:
[0054] 1. Strong robustness and adaptability: The coastal zone marine environmental pollution monitoring model designed and implemented by the present invention has strong robustness and adaptability, can adapt to different weather and lighting conditions, and can accurately detect in the coastal zone under different environments.
[0055] 2. High-precision recognition: By introducing an attention mechanism, adding a detection head with an auxiliary head and improving the loss function during the model training stage, the coastal zone marine environmental pollution monitoring model has higher detection accuracy under complex backgrounds and multi-scale targets.
[0056] 3. Multi-scale target extraction ability: The coastal zone marine environmental pollution monitoring model can effectively process targets of different scales, especially for small target objects, and can maintain a high accuracy rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1It is an intelligent monitoring invention method for coastal zone marine environmental pollution based on deep learning of the present invention;
[0058] Figure 2 It is the network structure of an intelligent monitoring method for coastal zone marine environmental pollution based on deep learning of the present invention;
[0059] Figure 3 It is the structure diagram of the shallow information perception module;
[0060] Figure 4 It is the structure diagram of the hybrid local channel attention module and the spatial attention module;
[0061] Figure 5 It is the structure diagram of the split dense multi-branch module;
[0062] Figure 6 It is the structure diagram of the aggregation dense layer module;
[0063] Figure 7 It is the structure diagram of the split cascaded dense multi-branch module;
[0064] Figure 8 It is the structure diagram of the cascaded aggregation dense layer module;
[0065] Figure 9 It is the comparison diagram of the auxiliary detection head network structure and the main detection head network structure;
[0066] Figure 10 It is the schematic diagram of the angle loss;
[0067] Figure 11 It is the schematic diagram of the distance loss. Detailed implementation manners
[0068] Before describing the embodiments of the present application in detail, first briefly describe the technical concept of the present application: Traditional detection algorithms for intelligent monitoring of coastal zone marine environmental pollution based on deep learning have deficiencies in aspects such as detection accuracy, detail recognition ability, multi-scale target extraction ability, and adaptability. Therefore, an intelligent monitoring method for coastal zone marine environmental pollution based on deep learning provided by the present application constructs a monitoring model for coastal zone marine environmental pollution based on deep learning, trains it using the preprocessed coastal zone image data, and uses the trained monitoring model for coastal zone marine environmental pollution to detect the coastal zone marine environment image to obtain the location and category of the pollution. The method proposed by the present application can effectively improve the adaptability, detection accuracy, detail recognition ability, and multi-scale target extraction ability of coastal zone marine environmental pollution monitoring.
[0069] Specifically, the method of the present invention includes:
[0070] Step 1) Uniformly adjust the size of the obtained training samples to a fixed size, and normalize the pixel values of the images to between 0 and 1.
[0071] Step 2) Construct a deep learning-based coastal zone marine environmental pollution monitoring model;
[0072] Step 3) Use the training samples to train the coastal zone marine environmental pollution monitoring model;
[0073] Step 4) Input the coastal zone image into the trained coastal zone marine environmental pollution monitoring model to obtain the location and category of the pollution.
[0074] The construction of the deep learning-based coastal zone marine environmental pollution monitoring model includes:
[0075] Use the model backbone network to extract the key features of the coastal zone image and obtain the key feature map.
[0076] Use the model neck network to further fuse and enhance the features of the key features.
[0077] Use the model head network to generate the bounding boxes for object detection and class predictions for the fused and enhanced features.
[0078] Use the geometric sensitivity auxiliary loss to optimize the training effect of the model.
[0079] The construction of the deep learning-based coastal zone marine environmental pollution monitoring model includes:
[0080] In the backbone network of the deep learning-based coastal zone marine environmental pollution monitoring model, a shallow information perception module with a hybrid channel spatial attention module is adopted.
[0081] In the neck network of the deep learning-based coastal zone marine environmental pollution monitoring model, a split dense multi-branch module and a split cascaded dense multi-branch module are adopted.
[0082] For the deep learning-based coastal zone marine environmental pollution monitoring model, during the training stage of the model, on the basis of the existing main detection head in the head network, an auxiliary enhanced detection network is added. During the inference stage of the model, only the main detection head is used for object detection.
[0083] The loss function of the deep learning-based coastal zone marine environmental pollution model is the geometric sensitivity auxiliary loss.
[0084] The technical solutions of the present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0085] Example 1
[0086] As Figure 1 and Figure 2As shown in the figure, an embodiment of the present invention provides an intelligent monitoring method for coastal zone marine environmental pollution based on deep learning, which mainly includes the following steps:
[0087] Step 1) Uniformly adjust the size of the obtained training samples to a fixed size, and normalize the pixel values of the images to between 0 and 1.
[0088] Step 2) Construct a coastal zone marine environmental pollution monitoring model based on deep learning.
[0089] Step 3) Obtain a coastal zone image dataset and perform annotation processing to obtain training samples.
[0090] Step 4) Use the training samples to train the coastal zone marine environmental pollution monitoring model.
[0091] Step 5) Input the coastal zone image into the trained coastal zone marine environmental pollution monitoring model to obtain the location and category of the pollution.
[0092] Specifically, the coastal zone marine environmental pollution monitoring model based on deep learning includes:
[0093] In the backbone network of the coastal zone marine environmental pollution monitoring model based on deep learning, a shallow information perception module with a hybrid channel spatial attention mechanism is adopted.
[0094] In the neck network of the coastal zone marine environmental pollution monitoring model based on deep learning, a split dense multi-branch module and a split cascaded dense multi-branch module are adopted.
[0095] During the model training stage, an auxiliary enhancement detection network is added to the head network of the coastal zone marine environmental pollution monitoring model based on deep learning. During the model inference stage, the auxiliary enhancement detection network is removed, and only the main detection head performs object detection. There are three main detection heads with the same structure in the head network, and each main detection head includes: a main head regression branch and a main head classification branch.
[0096] The loss function of the coastal zone marine environmental pollution monitoring model based on deep learning is a geometric sensitivity auxiliary loss.
[0097] The shallow information perception module is as Figure 3 shown. This module first performs convolutional processing on the input features, and then splits the feature map along the channels into feature Figure 1 and feature Figure 2 . Subsequently, feature Figure 2 is passed through two hybrid channel spatial attention modules. Then, feature Figure 1 , feature Figure 2, the outputs of the first hybrid channel spatial attention module and the second hybrid channel attention module are concatenated along the channels in sequence. Finally, the concatenated result is output after one convolution.
[0098] The hybrid channel spatial attention module consists of a hybrid local channel attention module and a spatial attention module. The specific operation is to pass the input feature through the hybrid local channel attention module and the spatial attention module in sequence, and finally add the input feature and the output of the spatial attention module element-wise. The network structures of the hybrid local channel attention module and the spatial attention module are as Figure 4 shown.
[0099] The hybrid local channel attention module first performs two convolution operations on the input feature tensor with the shape of (C, W, H), and then performs local average pooling operation to convert it into a feature vector with the size of (C, K s , K s ) to extract local spatial information. Here, C is the number of channels, W is the width of the feature map, H is the height of the feature map, and K s is the length of the feature map after local average pooling operation. K s ×K s is the size of the feature map after local average pooling. Here, K s = 5. The local spatial information is sent to the global information branch and the local information branch respectively to calculate their respective attention weights. Then, the attention weights output by the two branches are added element-wise and then anti-pooled to restore the shape and size of the hybrid local channel attention weight to (C, W, H). Finally, the input feature and the hybrid local channel attention weight are multiplied element-wise.
[0100] The global information branch specifically includes: globally averaging the input feature into a one-dimensional feature, then extracting the channel dependence relationship under the global information through one-dimensional convolution to obtain the channel attention weight under the global information, and finally restoring the channel attention weight under the global information to the same size as the input feature through anti-pooling operation.
[0101] The local information branch is characterized in that the local information branch specifically includes: reshaping the input feature into a one-dimensional feature vector, then passing the reshaped one-dimensional feature vector through one-dimensional convolution, and finally reshaping the result after convolution into the same shape and size as the input feature.
[0102] The spatial attention module first performs max - pooling and average - pooling on the input feature vector respectively, and concatenates the two resulting feature maps in the channel dimension. Then, it performs convolution on the concatenated feature vector and uses the Sigmoid activation function to normalize the result of the convolution to obtain the spatial attention map. Finally, the input feature vector is multiplied element - by - element with the spatial attention map.
[0103] The split dense multi - branch module is as Figure 5 shown. This module first performs convolution on the input feature, and then splits the feature map obtained by convolution along the channel into feature Figure 3 and feature Figure 4 . Subsequently, feature Figure 4 is passed through two aggregation dense layer modules. Then, feature Figure 3 , feature Figure 4 , the output of the first aggregation dense layer module, and the output of the second aggregation dense layer module are concatenated in order along the channel. Finally, the concatenated result is output after one convolution.
[0104] The aggregation dense layer module is as Figure 6 shown. Specifically, the input feature is convolved twice continuously and then copied into branch one and branch two. Subsequently, the feature of branch two is average - pooled and then fed into four parallel fully - connected layers. Then, the four parallel fully - connected layers are concatenated in order along the channel, passed through another fully - connected layer, and finally multiplied element - by - element with the output of branch one.
[0105] The split - cascaded dense multi - branch module is as Figure 7 shown. This module first performs convolution on the input feature, and then splits the feature map obtained by convolution along the channel into feature Figure 5 and feature Figure 6 . Subsequently, feature Figure 6 is passed through two cascaded aggregation dense layer modules. Then, feature Figure 5 , feature Figure 6 , the output of the first cascaded aggregation dense layer module, and the output of the second cascaded aggregation dense layer module are concatenated in order along the channel. Finally, the concatenated result is output after one convolution.
[0106] The cascaded aggregation dense layer module is as Figure 8 shown. This module first sends the input feature into branch one and branch two for processing respectively. Branch one passes the input feature through convolution and two aggregation dense layer modules in order. Branch two only performs one convolution operation. Subsequently, the outputs of branch one and branch two are concatenated in order along the channel, and the concatenated result is convolved again and then output.
[0107] The auxiliary enhancement detection network consists of a first auxiliary detection head, a second auxiliary detection head, and a third auxiliary detection head, and each auxiliary detection head has the same network structure. The main detection network consists of a first main detection head, a second main detection head, and a third main detection head. The relationship between the network structure of the auxiliary detection head and the network structure of the main detection head is as Figure 9 shown. Specifically: The main head regression branch and the main head classification branch together constitute the main detection head. The auxiliary head regression branch and the auxiliary head classification branch together constitute the auxiliary detection head. In the training mode, the three main detection heads respectively process the feature maps of the 16th, 19th, and 22nd layers. These layers usually include higher-level semantic information and are suitable for more refined object detection. The three auxiliary detection heads respectively process the feature maps of the 10th, 13th, and 16th layers, aiming to perform additional classification and regression tasks on the intermediate feature layers to enhance the regression ability of the model. During the training process, the main detection head and the auxiliary detection head will respectively output regression predictions and classification predictions, and calculate the corresponding losses. Finally, the total loss is the weighted sum of the losses of the main detection head and the auxiliary detection head, where the loss weight of the auxiliary detection head is 0.25. In the inference stage of the model, only the main detection head is used to process the feature maps. Specifically, it includes dynamically generating anchor point coordinates, decoding the regression prediction into bounding box coordinates through the distribution focal loss, and outputting the final prediction result containing the detection box coordinates and class probabilities.
[0108] The geometrically sensitive auxiliary loss L is specifically:
[0109] L = L cls + L siou + L dfl + λ * L auxiliary
[0110] where L cls is the classification loss, L siou is the bounding box loss we adopted, which is the SCYLLA loss with angular loss, distance loss, shape loss, and intersection over union loss between the predicted box and the ground truth box, L dfl is the distribution focal loss, λ is the auxiliary loss weight coefficient, and L auxiliary is the auxiliary head loss. In one embodiment, λ takes 0.25.
[0111] The auxiliary head loss L auxiliary The formula is specifically:
[0112] L auxiliary = L cls + L siou + L dfl
[0113] The specific formula for the SCYLLA intersection over union loss with angular loss, distance loss, shape loss, and intersection over union loss between the predicted box and the ground truth box is:
[0114]
[0115] Where L IOU is the intersection over union loss, Δ is the distance loss, and Ω is the shape loss. The angle loss is embodied in the distance loss.
[0116] The angle loss Λ is as Figure 10 shown, and the formula is as follows:
[0117]
[0118] where σ is the distance between the center point of the predicted bounding box and the center point of the ground truth bounding box, c h is the vertical distance between the center point of the predicted bounding box and the center point of the ground truth bounding box, c w is the horizontal distance between the center point of the predicted bounding box and the center point of the ground truth bounding box.
[0119] The distance loss Δ is as Figure 11 shown, and the formula is as follows:
[0120]
[0121] where t is a variable parameter, is the normalized squared difference between the abscissas of the center points of the predicted bounding box and the ground truth bounding box, is the normalized squared difference between the ordinates of the center points of the predicted bounding box and the ground truth bounding box. and are the abscissa and ordinate of the ground truth bounding box respectively. and are the abscissa and ordinate of the predicted bounding box respectively. C x is the width of the minimum bounding rectangle of the ground truth bounding box and the predicted bounding box, C y is the height of the minimum bounding rectangle of the ground truth bounding box and the predicted bounding box. γ = 2 - Λ, where Λ is the angle loss.
[0122] The shape loss Ω has the following formula:
[0123]
[0124] where t is a variable parameter, is the normalized difference between the width of the predicted bounding box and the width of the ground truth bounding box, is the normalized difference between the height of the predicted bounding box and the height of the ground truth bounding box. w and h are the width and height of the predicted bounding box respectively. w gt and h gt are the width and height of the ground truth bounding box respectively. The magnitude of θ controls the degree of attention to the shape loss.
[0125] The intersection over union loss IoU between the predicted bounding box and the ground truth bounding box has the following formula:
[0126]
[0127] where |B∩B GT | is the intersection part of the predicted box and the ground truth box, and |B∪B GT | is the union part of the predicted box and the ground truth box.
[0128] Embodiment 2
[0129] Embodiment 2 of the present invention proposes an intelligent monitoring system for coastal marine environmental pollution based on deep learning, which is implemented based on the method of Embodiment 1. The system includes:
[0130] A preprocessing module for preprocessing the collected coastal zone images;
[0131] A monitoring output module for inputting the preprocessed coastal zone images into a pre-trained coastal marine environmental pollution monitoring model to obtain the location and category of the pollution;
[0132] The coastal marine environmental pollution monitoring model is based on the YOLOv11 network and includes a backbone network, a neck network, and a head network; among them, in the backbone network, a shallow information perception module with a hybrid channel spatial attention module is used to extract the key features of the coastal zone images to obtain a key feature map; in the neck network, a split dense multi-branch module and a split cascaded dense multi-branch module are used to further fuse and enhance the key features;
[0133] In the training stage of the coastal marine environmental pollution monitoring model, a geometric sensitivity auxiliary loss function is used to optimize the training effect.
[0134] It should be noted that in the embodiments of the above system, the included modules are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional modules are only for the convenience of mutual distinction and do not limit the protection scope of the present invention.
[0135] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the embodiments, those of ordinary skill in the art should understand that any modification or equivalent replacement of the technical solutions of the present invention does not depart from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. An intelligent monitoring method for marine environmental pollution in the coastal zone based on deep learning, the method comprising: Preprocessing the collected coastal zone images; Inputting the preprocessed coastal zone images into a pre-trained coastal zone marine environmental pollution monitoring model to obtain the location and category of the pollution; The coastal zone marine environmental pollution monitoring model is based on the YOLOv11 network and includes a backbone network, a neck network, and a head network; wherein, in the backbone network, a shallow information perception module with a hybrid channel spatial attention module is used to extract key features of the coastal zone images and obtain a key feature map; in the neck network, a split dense multi-branch module and a split cascaded dense multi-branch module are used to further fuse and enhance the key features; During the training phase of the coastal zone marine environmental pollution monitoring model, an auxiliary enhancement detection network is added to the head network, and a geometric sensitivity auxiliary loss function is used to optimize the training effect.
2. The intelligent monitoring method for coastal marine environmental pollution based on deep learning according to claim 1, characterized in that, The preprocessing includes: adjusting the image size and normalizing the pixel values of the image.
3. The intelligent monitoring method for coastal marine environmental pollution based on deep learning according to claim 1, wherein The processing process of the shallow information perception module includes: Performing convolution processing on the input features to obtain a feature map, splitting it along the channel into a first feature map and a second feature map, passing the second feature map through two hybrid channel spatial attention modules with the same structure, and then sequentially splicing the first feature map, the second feature map, the output of the first hybrid channel spatial attention module, and the output of the second hybrid channel attention module along the channel, and outputting the splicing result after one convolution.
4. The intelligent monitoring method for coastal marine environmental pollution based on deep learning according to claim 3, characterized in that, The hybrid channel spatial attention module includes: a hybrid local channel attention module and a spatial attention module, and the processing process of the hybrid channel spatial attention module includes: sequentially passing the input features through the hybrid local channel attention module and the spatial attention module, and then performing element-wise addition of the input features and the output of the spatial attention module; The processing process of the hybrid local channel attention module includes: After performing two convolution operations on the input feature tensor of shape (C, W, H), a local average pooling operation is performed, and it is converted into a feature vector of size (C, K s , K s ) to extract local spatial information, where C is the number of channels, W is the width of the feature map, H is the height of the feature map, and K s is the length of the feature map after the local average pooling operation, and K s × K s is the size of the feature map after local average pooling; Sending the local spatial information into a global information branch and a local information branch respectively, and calculating their respective attention weights; Performing element-wise addition on the attention weights output by the two branches and then performing unpooling to restore them into a hybrid local channel attention weight with a shape size of (C, W, H); Performing element-wise multiplication of the input features and the hybrid local channel attention weight.
5. The intelligent monitoring method for coastal marine environmental pollution based on deep learning according to claim 4, wherein The processing process of the global information branch includes: Performing global average pooling on the input features to obtain a one-dimensional feature, then extracting the channel dependence relationship under the global information through one-dimensional convolution to obtain the channel attention weight under the global information, and restoring the channel attention weight under the global information to the same size as the input features through unpooling operations; The processing process of the local information branch includes: Reshaping the input features into a one-dimensional feature vector, then performing one-dimensional convolution, and reshaping the result after convolution into the same size as the input features.
6. The intelligent monitoring method for coastal marine environmental pollution based on deep learning according to claim 4, characterized in that The processing process of the spatial attention module includes: Perform max pooling and average pooling on the input feature vectors respectively to obtain two feature maps. Concatenate the two feature maps in the channel dimension, perform convolution on the concatenated feature vectors, and use the Sigmoid activation function for normalization processing to obtain the spatial attention map; Multiply the input feature vectors element-wise with the spatial attention map.
7. The intelligent monitoring method for coastal marine environmental pollution based on deep learning according to claim 1, wherein The processing process of the split dense multi-branch module includes: Perform convolution processing on the input features to obtain a feature map, split it along the channel into a third feature map and a fourth feature map, pass the fourth feature map through two aggregation dense layer modules with the same structure, and then concatenate the third feature map, the fourth feature map, the output of the first aggregation dense layer module, and the output of the second aggregation dense layer module in sequence along the channel, and output the concatenated result after one convolution; The processing process of the aggregation dense layer module includes: Copy the features obtained by performing two consecutive convolutions on the input features into two paths. One path of features is sent to four parallel fully connected layers after average pooling, and then the outputs of the four parallel fully connected layers are concatenated in sequence along the channel. After passing through one fully connected layer, multiply it element-wise with the other path of features.
8. The intelligent monitoring method for coastal marine environmental pollution based on deep learning according to claim 1, characterized in that The processing process of the split cascaded dense multi-branch module includes: Perform convolution processing on the input features, split the convolution-obtained feature map along the channel into a fifth feature map and a sixth feature map, pass the sixth feature map through two cascaded aggregation dense layer modules, and then concatenate the fifth feature map, the sixth feature map, the output of the first cascaded aggregation dense layer module, and the output of the second cascaded aggregation dense layer module in sequence along the channel, and output the concatenated result after one convolution; The processing process of the cascaded aggregation dense layer module includes: Send the input features into two branches respectively. The first branch passes the input features through convolution and two aggregation dense layer modules in sequence. The second branch performs one convolution operation, and then concatenate the outputs of the first branch and the second branch in sequence along the channel, and output the concatenated result after convolution again.
9. The intelligent monitoring method for coastal marine environmental pollution based on deep learning according to claim 1, characterized in that, In the training stage of the coastal zone marine environmental pollution monitoring model, an auxiliary enhancement detection network is added to the head network. The auxiliary enhancement detection network includes three auxiliary detection heads with the same structure, which are used to perform additional classification and regression tasks on the intermediate feature layers to enhance the regression ability of the model; The auxiliary detection head includes: an auxiliary head regression branch and an auxiliary head classification branch.
10. The intelligent monitoring method for coastal marine environmental pollution based on deep learning according to claim 1, wherein The geometric sensitivity auxiliary loss function L is: L = (L cls + L siou + L dfl )(1 + λ) Among them, L cls is the classification loss, L siou is the bounding box loss, L dfl is the distribution focal loss, and λ is the auxiliary loss weight coefficient; Bounding box loss L siou It is the SCYLLA loss with angular loss, distance loss, shape loss, and the intersection over union loss between the predicted box and the ground truth box, satisfying the following formula: Among them, L IOU is the intersection over union loss, Δ is the distance loss, and Ω is the shape loss; The angle loss Λ satisfies the following formula: Among them, σ is the distance between the center point of the predicted bounding box and the center point of the ground truth bounding box, and c h is the vertical distance between the center point of the predicted bounding box and the center point of the ground truth bounding box; The distance loss Δ satisfies the following formula: Among them, ρ x is the normalized mean squared error between the abscissa of the center point of the predicted box and the abscissa of the center point of the ground truth box, and ρ y is the normalized mean squared error between the ordinate of the center point of the predicted box and the ordinate of the center point of the ground truth box, γ = 2 - Λ; The shape loss Ω satisfies the following formula: Among them, ω w is the normalized difference between the width of the predicted box and the width of the ground truth box, and ω h is the normalized difference between the height of the predicted box and the height of the ground truth box, where θ represents the degree of attention to the shape loss; The intersection over union loss IoU between the predicted box and the ground truth box is: Among them, |B∩B GT | is the intersection part of the predicted box and the ground truth box, and |B∪B GT | is the union part of the predicted box and the ground truth box.