UUV underwater fishing net visual identification method suitable for dark environment

By introducing sparse connectivity and deformable convolution attention mechanisms on UUVs, combined with SCBAM, CoT and SEAM modules, the problem of low accuracy of fishing net recognition in underwater dim environments is solved, and efficient and accurate fishing net detection is achieved.

CN120339816APending Publication Date: 2025-07-18HARBIN UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510405910.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In underwater dim environments, it is difficult for UUV to accurately identify fishing nets, especially when the light is weak and there are interferences of objects similar to the fishing nets around it, and the prior art is difficult to improve detection accuracy.

Method used

Introduce a sparse connectivity and deformable convolution attention mechanism, combining SCBAM, CoT and SEAM modules to optimize feature interaction and boundary fitting, and enhance the recognition ability of fishing nets.

Benefits of technology

In dim environments, the accuracy and recall of fishing net detection are significantly improved, reaching 92.00% and 82.10%, and the computing efficiency is optimized to achieve high-precision fishing net recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339816A_ABST
    Figure CN120339816A_ABST
Patent Text Reader

Abstract

According to the UUV underwater fishing net visual identification method suitable for the dark environment, the UUV can accurately identify the fishing net in the underwater dark environment. According to the method, an attention mechanism based on sparse connectivity and deformable convolution is introduced, and optimization is performed specially for a visual environment with weak underwater light, so that the fishing net detection performance of the UUV in a dark environment is enhanced. Besides, the introduced SCBAM module can efficiently extract the feature information of the fishing net, meanwhile, the CoT module and the SEAM module enable the UUV to accurately identify the fishing net when an object similar to the fishing net exists around through deeper feature interaction, and therefore the high-precision detection effect is achieved in the dark underwater environment. The result shows that the accuracy of the method is 92.00%, the recall rate is 82.10%, and the mAP is 87.30%. Compared with other models, the detection of the fishing net in the dark environment has the highest precision, and an effective solution is provided for the UUV to identify the fishing net in the dark environment.
Need to check novelty before this filing date? Find Prior Art

Description

(1) Technical Field

[0001] The present invention belongs to the field of computer vision, and particularly relates to image recognition technology, especially a UUV underwater fishing net vision recognition method applicable to dim environments. (2) Background Art

[0002] Unmanned Underwater Vehicles (UUVs) are important tools for ocean exploration. They can be applied in a wide range of tasks in ocean surveys, resource exploration, and the military field. However, underwater fishing nets are potential lethal hazards that UUVs often face when performing tasks in complex and unknown ocean environments. Therefore, developing application software such as image processors and vision processors to identify fishing nets through technologies such as image recognition, image enhancement, and image discrimination is the key to improving the intelligence and autonomous survival ability of underwater vehicles before they are deployed to sea.

[0003] Inspired by the human visual system, object recognition in robots and other intelligent systems is usually carried out with the help of optical sensors. However, due to factors such as light refraction, water flow interference, and suspended particles in the underwater environment, the image contrast is often low and blurred, which poses challenges to vision-based target detection algorithms. Fishing nets have complex textures and are easily combined with the underwater environment, making it particularly difficult to accurately identify and locate them using the YOLO algorithm. Due to the fine structure and large size of fishing nets, the YOLO algorithm is difficult to capture the overall shape and detailed features of fishing nets in a single-scale field of view. This may cause small targets, such as a part of the fishing net, to be ignored, wasting computing resources or losing the key features of large-sized fishing nets, affecting the positioning and recognition accuracy. In addition, overlapping targets in underwater scenes, such as multiple layers or partially stacked fishing nets, increase the recognition difficulty. Since the YOLO algorithm has limited ability to handle overlapping targets, it is difficult to distinguish closely connected or partially occluded targets.

[0004] In order to solve the problem of low accuracy in identifying fishing nets, the attention mechanism CBAM is introduced in the paper "Lightweight Fish Detection Method Based on Improved YOLOv8" to improve the detection accuracy of the network model, but the improvement is not much; in the patent "A Method for Underwater Fishing Net Identification Based on Improved YOLOv7 Model", the MobileOne module is added to the Backbone network of the YOLOv7 model to improve the model lightweight, reducing the number of network parameters, calculation amount, and reasoning time, but the model is still complex; in the paper "Light-YOLO: A Study of a Lightweight YOLOv8n-Based Method for Underwater Fishing Net In the patent "A method for underwater blue-green laser fishing net recognition and positioning based on semantic segmentation", the trained fishing net semantic segmentation network is used to perform semantic segmentation on the near image and the far image respectively, and the image to be recognized and the segmentation map corresponding to the image to be recognized are superimposed on the channel dimension to obtain the superimposed image to be recognized. Although the recognition accuracy is improved, the method is too complicated.

[0005] In order to solve the above problems, the present invention proposes a UUV underwater fishing net visual recognition method suitable for dim environments. The method can effectively improve the detection accuracy of fishing nets when the underwater light is weak and there are interferences from objects similar to fishing nets around. (III) Summary of the invention

[0006] The purpose of the present invention is to provide a model that can accurately identify fishing nets when the underwater light is weak and there are objects similar to fishing nets around, so as to achieve safe and efficient underwater operations. The present invention introduces an attention mechanism based on sparse connectivity and deformable convolution, which is optimized for the visual conditions of weak underwater light. This attention mechanism enhances the detection performance by focusing on the key visual features of the fishing net, and the introduced SCBAM, CoT and SEAM modules further improve the recognition accuracy of UUV for fishing nets when there are objects similar to fishing nets around through deeper feature interactions.

[0007] To achieve the above object, the present invention adopts the following technical solution:

[0008] S1. Generate a set of uniformly distributed reference points on the feature map and use the offset network to learn the offsets of these reference points based on the query features.

[0009] S2. For each query vector q, use the offset network O to calculate the offset vector ΔP of all reference points obtained from the query feature.offset :

[0010] ΔP offset = O(q) (1)

[0011] Wherein, O(q) represents the new offset of each reference point relative to the original position;

[0012] S3. In the feature sampling and deformation stage, the offset network takes the query feature as input, and based on the learned offset, samples the feature map through bilinear interpolation to obtain the deformed key and value:

[0013]

[0014] Wherein, K deformed is the deformed key, V deformed is the deformed value, (m) represents the m-th attention head, where F and G are functions that map the deformed position back to the feature map to extract the key and value respectively;

[0015] S4. After calculating the offset corresponding to each reference point, calculate the relative position deviation from the deformed point, enhance the multi-head attention mechanism, and output the transformed feature representation. The multi-head attention calculation follows the multi-head attention calculation principle in the standard transformer module, where the keys and values used have undergone the above deformation operations, and the relative position offset calculated from the deformed point is introduced to enhance the attention mechanism. The formula is as follows:

[0016]

[0017] Wherein, ζ is a softmax function used to normalize the attention weight of each position, d is half of the feature dimension in each attention head. By projecting the query, key, and value into separate subspaces and performing independent attention calculations, the model can capture different pattern dependencies simultaneously. DA2D further introduces the ability to dynamically adjust the positions of the key points and values, enabling the UUV to accurately identify fishing nets in a dim environment;

[0018] S5. Embed SCBAM (Super Convolutional Block Attention Module) in the feature extraction layer of YOLOv8. SCBAM creatively proposes a minimum pooling operation on the basis of the CBAM module, breaking through the limitations of the traditional separation of channel attention and spatial attention, and achieving triple pooling fusion while maintaining computational efficiency. SCBAM is composed of a channel attention module (Channel Attention Module, CAM) and a spatial attention module (Spatial Attention Module, SPA). The output M of the channel attention module C(F) can be calculated by the following formula:

[0019]

[0020] In the formula, F is the input feature map, AvgPool, MaxPool, and MinPool respectively represent the global average pooling, max pooling, and min pooling operations, MLP represents the multi-layer perceptron, σ represents the Sigmoid activation function, W0 and W1 are the weights in the 3-layer fully connected network, and the output M of the spatial attention module S (F) can be calculated by the following formula:

[0021] M S (F) = δ(f 7*7 ([AvgPool(F); MaxPool(F); MinPool(F)])) (6)

[0022] In the formula, f 7*7 represents a 7*7 convolution operation, [AvgPool(F); MaxPool(F); MinPool(F)] represents concatenating the average pooling, max pooling, and min pooling results along the channel axis. Through this module, the UUV can identify the fishing net when there are objects similar to the fishing net around it;

[0023] S6. Adopt the CoT (contextual transformer, CoT) module. This module constructs a new architecture containing four context dimensions, which not only breaks through the limitations of the traditional local and global dichotomy but also solves the problem of non-linear fusion of spatio-temporal information in dynamic scenarios. For the input feature X, define three variables K (Keys), Q (Query), and V (Values). Perform context encoding on the input Key through a 3×3 convolution to obtain the static context feature information K1 between local adjacent keys. Then concatenate K1 with Q, and generate the attention matrix A through two consecutive 1×1 convolutions:

[0024] A = [K1,Q]W θ W δ (7)

[0025] S7. Multiply the matrix A by V to obtain the dynamic context feature information K2:

[0026] K2 = V*A (8)

[0027] S8. Add the local static context feature information K1 and the dynamic context feature information K2 as the output to obtain Y:

[0028] Y = K1 + K2 (9)

[0029] S9. Incorporate the SEAM module into the YOLO detection framework to enhance the spatial transformation invariance of the model. SEAM achieves covariant regularization by constructing a dual-network structure with shared weights to ensure that self-supervision is provided for network learning from various transformed Class Activation Maps (CAMs).

[0030] R ER = ||F(A(I)) - A(F(I))||1 (10)

[0031] In the formula, F(·) represents the network, A(·) represents the affine transformation, and I is the input image.

[0032] S10. Introduce a Pixel Correlation Module (PCM) at the end of the SEAM network. The PCM captures the context information of each pixel through a self-attention mechanism to further refine the CAMs. The PCM uses the cosine distance to measure the feature similarity between pixels and calculates the affinity through the normalization of the inner product in the feature space, optimizing the CAMs to more accurately fit the target boundary, so that the UUV can accurately identify the fishing net in a dim environment with interference from objects similar to fishing nets around it.

[0033] The present invention has the following beneficial effects:

[0034] (1) Introduce a fast and efficient feature processing module DA2D based on sparse linkage and variant convolution. Through this module, the present invention can dynamically adjust the position of the attention sampling points according to the input content, enabling the model to focus on the local features of the fishing net, reducing the computational overhead, effectively improving the accuracy rate, increasing the accuracy rate from 87.60% to 89.50%, the recall rate from 78.10% to 79.80%, and the mAP from 83.70% to 86.20%.

[0035] (2) Propose the SCBAM module. The SCBAM creatively proposes a minimum pooling operation on the basis of the CBAM module, breaking through the limitations of the traditional separation of channel attention and spatial attention, achieving triple pooling fusion while maintaining computational efficiency. By combining channel attention and spatial attention, double refinement of the input features is realized. This design enables the model to simultaneously focus on which channels and which spatial positions are meaningful, thereby improving the representation ability and decision-making accuracy of the model, enabling the network to accurately capture the important features of the target, and thus achieving accurate identification of the fishing net when there is interference from objects similar to fishing nets around it, increasing the accuracy rate from 89.50% to 90.70%, the recall rate from 79.80% to 80.50%, and the mAP from 86.20% to 86.90%.

[0036] (3) The CoT module is proposed. This module constructs a new architecture that includes four context dimensions, not only breaking through the limitations of the traditional local and global dichotomy, but also solving the problem of non-linear fusion of spatio-temporal information in dynamic scenarios, effectively improving the model's ability to identify and model fishing nets, enhancing the model's recognition accuracy, so as to achieve accurate identification of fishing nets in dim environments, increasing the accuracy rate from 90.70% to 91.50%, the recall rate from 80.50% to 81.30%, and the mAP from 86.90% to 87.10%.

[0037] (4) The SEAM module is introduced. By constructing a twin network structure with shared weights, it optimizes the boundary fitting of fishing nets, enhances the model's sensitivity to context information, uses covariant regularization technology to ensure the consistency of CAMs under different transformations, and uses the PCM pixel correlation module to refine CAMs, increasing the accuracy rate from 91.50% to 92.00%, the recall rate from 81.30% to 82.10%, and the mAP from 87.10% to 87.30%. (IV) BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 is the flow chart of the present invention;

[0039] Figure 2 is the flow chart of the DA2D module;

[0040] Figure 3 is the flow chart of the SCBAM module;

[0041] Figure 4 is the diagram of the channel attention mechanism module;

[0042] Figure 5 is the diagram of the spatial attention mechanism module;

[0043] Figure 6 is the diagram of the CoT module;

[0044] Figure 7 is the diagram of the PCM module;

[0045] Figure 8 is the experimental environment of the present invention;

[0046] Figure 9 The comparison diagrams of the present invention, YOLOv7-tiny and YOLOv8n. (V) DETAILED IMPLEMENTATION MANNER

[0047] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and experimental examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. The overall flowchart of the system of the present invention is as Figure 1 shown.

[0048] S1. Generate a set of uniformly distributed reference points on the feature map, and use the offset network to learn the offsets of these reference points based on the query features;

[0049] S2. For each query vector q, use the offset network O to calculate the offset vectors ΔP of all reference points obtained from the query features offset :

[0050] ΔP offset = O(q) (1)

[0051] In the formula, O(q) represents the new offset of each reference point relative to the original position;

[0052] S3. In the feature sampling and deformation stage, the offset network takes the query features as input, and based on the learned offsets, samples the feature map through bilinear interpolation to obtain the deformed keys and values:

[0053]

[0054] In the formula, K deformed is the deformed key, V deformed is the deformed value, (m) represents the m-th attention head, where F and G are functions that map the deformed positions back to the feature map to extract keys and values respectively;

[0055] S4. After calculating the offsets corresponding to each reference point, calculate the relative position deviation from the deformed points, enhance the multi-head attention mechanism, and output the transformed feature representation. The multi-head attention calculation follows the multi-head attention calculation principle in the standard transformer module, where the keys and values used have undergone the above deformation operations, and the relative position offsets calculated from the deformed points are introduced to enhance the attention mechanism. The formula is as follows:

[0056]

[0057] In the formula, ζ is a softmax function used to normalize the attention weights at each position, d is half of the feature dimension in each attention head. By projecting the query, key, and value into separate subspaces and performing independent attention calculations, the model can capture different pattern dependencies simultaneously. DA2D further introduces the ability to dynamically adjust the positions of key points and values, enabling the UUV to accurately identify fishing nets in a dim environment. The flowchart of DA2D is asFigure 2 as shown

[0058] S5. Embed SCBAM (Super Convolutional Block Attention Module) into the feature extraction layer of YOLOv8. SCBAM creatively proposes the minimum pooling operation on the basis of the CBAM module, breaking through the limitations of the separation of traditional channel attention and spatial attention, and achieving triple pooling fusion while maintaining computational efficiency. The flow chart of SCBAM is as Figure 3 shown. SCBAM consists of a channel attention module (Channel Attention Module, CAM) and a spatial attention module (Spatial Attention Module, SPA). The output M C (F) of the channel attention module can be calculated by the following formula:

[0059]

[0060] In the formula, F is the input feature map, AvgPool, MaxPool, and MinPool respectively represent global average pooling, max pooling, and minimum pooling operations, MLP represents a multi-layer perceptron, σ represents the Sigmoid activation function, and W0 and W1 are the weights in the 3-layer fully connected network. The flow chart of the channel attention module is as Figure 4 shown. The output M S (F) of the spatial attention module can be calculated by the following formula:

[0061] M S (F) = δ(f 7*7 ([AvgPool(F); MaxPool(F); MinPool(F)])) (6)

[0062] In the formula, f 7*7 represents a 7*7 convolution operation, and [AvgPool(F); MaxPool(F); MinPool(F)] represents concatenating the average pooling, max pooling, and minimum pooling results along the channel axis. This module enables the UUV to recognize the fishing net when there are objects similar to the fishing net around it. The flow chart of the spatial attention module is as Figure 5 shown;

[0063] S6. Adopt the CoT (Contextual Transformer) module. This module constructs a new architecture containing four context dimensions, which not only breaks through the limitations of the traditional local and global dichotomy, but also solves the problem of non-linear fusion of spatio-temporal information in dynamic scenarios. The flow chart of the CoT module is as Figure 6As shown, for the input feature X, three variables K (Keys), Q (Query), and V (Values) are defined. The input Key is contextually encoded through a 3×3 convolution to obtain the static context feature information K1 between local adjacent keys. Then, K1 is concatenated with Q, and two consecutive 1×1 convolutions are used to generate the attention matrix A:

[0064] A = [K1, Q]W θ W δ (7)

[0065] S7. Multiply the matrix A by V to obtain the dynamic context feature information K2:

[0066] K2 = V * A (8)

[0067] S8. Add the local static context feature information K1 and the dynamic context feature information K2 as the output to obtain Y:

[0068] Y = K1 + K2 (9)

[0069] S9. Incorporate the SEAM module into the YOLO detection framework to enhance the spatial transformation invariance of the model. SEAM achieves covariant regularization by constructing a dual-network structure with shared weights to ensure self-supervision for network learning from various transformed Class Activation Maps (CAMs):

[0070] R ER = ||F(A(I)) - A(F(I))||1 (10)

[0071] where F(·) represents the network, A(·) represents the affine transformation, and I is the input image;

[0072] S10. Introduce a Pixel Correlation Module (PCM) at the end of the SEAM network. The flowchart of the PCM module is as Figure 7 shown. It captures the context information of each pixel through a self-attention mechanism to further refine the CAMs. The PCM uses the cosine distance to measure the feature similarity between pixels and calculates the affinity through the normalization of the inner product in the feature space to optimize the CAMs to more accurately fit the target boundary, enabling the UUV to accurately identify fishing nets in a dim environment with interference from objects similar to fishing nets around;

[0073] Table 1 Training Environment and Hardware Platform Parameters

[0074]

[0075] Table 2 Some key parameters set during model training

[0076]

[0077] To further evaluate the effectiveness of the present invention in detecting fishing nets, taking YOLOv8n as the baseline network, ablation experiments were conducted on each improved module, mainly using precision, recall, and mean average precision as reference indicators. The specific experimental results are shown in Table 3 below:

[0078] Table 3 Ablation experiments

[0079]

[0080]

[0081] To highlight the excellent performance of the present invention, it was first compared with the traditional lightweight YOLO series. The specific data is shown in Table 4:

[0082] Table 4 Comparative experiments of traditional YOLO series

[0083]

[0084] In the comparative analysis of the traditional YOLO series models, the present invention demonstrated excellent performance under lightweight design. The model achieved an accuracy of 92.00% and a recall of 82.10%, exceeding all other comparison models including YOLOv7-tiny. The experimental environment of the present invention is as Figure 8 shown. The comparison graphs of the present invention, YOLOv7-tiny, and YOLOv8n for identifying fishing nets are as Figure 9 shown. It can be seen from the graph that YOLOv7-tiny and YOLOv8n identify patterns similar to fishing nets as fishing nets, while the present invention can accurately identify fishing nets in the presence of interference from objects similar to fishing nets.

[0085] Although the frame rate (FPS) of 121.95 of the present invention is lower than that of 833.33 of some other models such as YOLOv5n, the FPS provided by the present invention is sufficient to meet the requirements of most real-time application scenarios, especially in video surveillance and mobile devices, while maintaining high accuracy and low computational requirements.

[0086] After comparing the traditional lightweight YOLO series, comparisons were continued with other YOLO-based improvements and other object detection models. The specific data is shown in Table 5:

[0087] Table 5 Comparative experiments of other models

[0088]

[0089] Compared with other models, the present invention not only achieves the highest mAP, reaching 87.30%, but also has the fewest parameters, only 4.40M, and the GFLOPs is only 6.10, showing high accuracy in fishing net detection. In addition, the present invention adopts the CoT and SEAM modules, enhancing the interaction between features and enabling the model to better understand context information at different levels. Combining the deformable convolution DA2D and the sparse connectivity feature module, the present invention not only outperforms many competitors in terms of accuracy, but also has an inference speed of 121.95FPS, second only to 227.69FPS of YOLOX. However, it achieves extremely high computational efficiency while maintaining high accuracy, making it an ideal choice that combines both accuracy and efficiency.

[0090] Implementing the above specific implementation schemes further illustrates the invention purpose, technical solutions and beneficial effects of the present invention. The above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting the protection scope of the present invention. Those of ordinary skill in the art should understand that any modification and equivalent replacement of the technical solutions of the present invention are included in the protection scope of the present invention.

Claims

1. A visual recognition method for UUV underwater fishing nets applicable to dim environments, characterized in that, It includes the following steps: S1. Generate a set of uniformly distributed reference points on the feature map, and use the offset network to learn the offsets of these reference points based on the query feature; S2. For each query vector q, use the offset network O to calculate the offset vectors ΔP of all reference points obtained from the query features offset : ΔP offset = O(q) (1) In the formula, O(q) represents the new offset of each reference point relative to the original position; S3. In the feature sampling and deformation stage, the offset network takes the query feature as the input. Based on the learned offsets, the feature map is sampled through bilinear interpolation to obtain the deformed keys and values: where K deformed is the transformed key, V deformed is the transformed value, (m) represents the m-th attention head, where F and G are functions that map the transformed positions back to the feature map to extract the key and value respectively; S4. After calculating the offsets corresponding to each reference point, calculate the relative position deviation from the deformed points, enhance the multi-head attention mechanism, and output the transformed feature representation. The multi-head attention calculation follows the multi-head attention calculation principle in the standard transformer module, where the keys and values used have undergone the above deformation operations, and the relative position offsets calculated from the deformed points are introduced to enhance the attention mechanism. The formula is as follows: In the formula, ζ is a softmax function used to normalize the attention weights at each position, and d is half of the feature dimension in each attention head. By projecting the query, key, and value into separate subspaces and performing independent attention calculations, the model can simultaneously capture different pattern dependencies. DA2D further introduces the ability to dynamically adjust the positions of the key points and values, enabling the UUV to accurately identify fishing nets in a dim environment; S5. Embed SCBAM (Super Convolutional Block Attention Module) in the feature extraction layer of YOLOv8. SCBAM creatively proposes a minimum pooling operation on the basis of the CBAM module, breaking through the limitations of the separation of traditional channel attention and spatial attention, and achieving triple pooling fusion while maintaining computational efficiency. SCBAM consists of a channel attention module (Channel Attention Module, CAM) and a spatial attention module (Spatial Attention Module, SPA). The output M C (F) can be calculated by the following formula: Wherein, F is the input feature map, AvgPool, MaxPool, and MinPool respectively represent global average pooling, max pooling, and min pooling operations, MLP represents a multi-layer perceptron, σ represents the Sigmoid activation function, W0 and W1 are the weights in the 3-layer fully connected network, and the output M of the spatial attention module S (F) can be calculated by the following formula: M S (F) = δ(f 7*7 ([AvgPool(F); MaxPool(F); MinPool(F)])) (6) where f 7*7 represents a 7*7 convolution operation, and [AvgPool(F); MaxPool(F); MinPool(F)] represents concatenating the average pooling, maximum pooling, and minimum pooling results along the channel axis. Through this module, the UUV can identify a fishing net when there are objects similar to a fishing net around it; S6. Adopt the CoT (contextual transformer, CoT) module. This module constructs a new architecture containing four context dimensions, which not only breaks through the limitations of the traditional local and global dichotomy but also solves the problem of non-linear fusion of spatio-temporal information in dynamic scenarios. For the input feature X, define three variables K (Keys), Q (Query), and V (Values). The input Key is contextually encoded through a 3×3 convolution to obtain the static context feature information K1 between local adjacent keys. Then, K1 is concatenated with Q, and two consecutive 1×1 convolutions are used to generate the attention matrix A: A = [K1, Q]W θ W δ (7) S7. Multiply the matrix A by V to obtain the dynamic context feature information K2: K2 = V * A (8) S8. Add the local static context feature information K1 and the dynamic context feature information K2 as the output to obtain Y: Y = K1 + K2 (9) S9. Add the SEAM module to the YOLO detection framework to enhance the spatial transformation invariance of the model. SEAM realizes covariant regularization by constructing a dual-network structure with shared weights to ensure self-supervision for network learning from various transformed Class Activation Maps (CAMs): R ER = ||F(A(I)) - A(F(I))||1 (10) In the formula, F(·) represents the network, A(·) represents the affine transformation, and I is the input image; S10. Introduce a Pixel Correlation Module (PCM) at the end of the network in SEAM. Capture the context information of each pixel through the self-attention mechanism to further refine the CAMs. The PCM uses the cosine distance to measure the feature similarity between pixels and calculates the affinity through the normalization of the inner product in the feature space, optimizing the CAMs to more accurately fit the target boundary, so that the UUV can accurately identify the fishing net in a dim environment and when there are interferences from objects similar to fishing nets around it.