Underwater target detection method and device
By integrating the multi-scale expansion attention mechanism into the C2PSA module of the YOLOv11 network model and introducing the Slim-Neck structure into the neck network, the detection accuracy and computational complexity problems of the existing underwater target detection methods in complex optical interference and small target dense distribution environments are solved, and more efficient underwater target detection performance is achieved.
Patent Information
- Application Number
- CN202510524309.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-05-27
AI Technical Summary
Existing underwater target detection methods perform poorly in underwater environments with complex optical interference and dense distribution of small targets, especially in low-contrast images and complex backgrounds with low detection accuracy and high computational complexity, making it difficult to deploy on edge devices.
The multi-scale expansion attention mechanism (MSDA) is incorporated into the backbone network C2PSA module of the YOLOv11 network model to form the C2PSA_MSDA module, and the Slim-Neck network is introduced into the neck network, including the GSConv module and the VoV-GSCSPC module, combined with the EMA module to enhance the feature extraction capability of the detection head.
By introducing the MSDA attention mechanism and Slim-Neck network structure, the model's detection ability of multi-scale targets is improved, the attention to key areas is enhanced, the interference of complex background information is suppressed, the detection and classification capabilities are improved, and the model complexity and computing overhead are reduced, which is suitable for real-time detection on edge devices.
Smart Images

Figure CN120047810A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of target detection, and in particular to an underwater target detection method and device. Background Art
[0002] With the advancement of science and the development of technology, underwater object detection (UOD) has become a key technology for marine resource development and sustainable resource management.
[0003] Among them, the complexity of the underwater environment poses a severe challenge to detection technology. First, there is relatively serious optical interference in the actual underwater environment. Light attenuation, scattering and color distortion lead to low image contrast and difficulty in distinguishing targets from backgrounds. Second, most marine organisms occupy less than 1.65% of the image area, and dense distribution easily leads to target overlap and missed detection. In addition, small targets have significant dense distribution characteristics. UDD data set statistics show that 35% of the targets are less than 32×32 pixels in size, and the spatial distribution density reaches 4.7 per square meter. In addition, edge devices such as underwater robots are limited by computing power, making it difficult to strike a balance between accuracy and efficiency.
[0004] At present, underwater target detection methods are usually divided into two categories, one is based on traditional target detection methods, and the other is based on deep learning target detection methods. Traditional target detection often uses sonar equipment (such as side scan sonar, multi-beam sonar) to obtain underwater target images, combined with optical image processing technology, and analyzes the target shape through echo signals. In recent years, deep learning technology has developed rapidly, ushering in a paradigm shift in UOD, and various machine learning and deep learning technologies have been applied to the field of underwater real-time target detection. However, the accuracy of existing methods, such as YOLOv8 and Faster R-CNN, in underwater scenes is 35% to 45% lower than that in terrestrial applications, mainly because they cannot handle low-contrast images and small targets, and the missed detection rate for complex backgrounds, low-resolution scenes, and small target scenes is as high as 43%. For example, in low-light conditions, scallops, as a key indicator of ocean health, are often misclassified as rocks or garbage, resulting in abnormal detection. In addition, existing methods rely too much on image enhancement technology, resulting in the loss of some details, high computational complexity, and difficulty in deploying on edge devices.
[0005] Based on this, the YOLO (You Only Look Once) algorithm emerged in the existing technology, which overturned the target detection paradigm. By converting the detection task into a single-stage regression problem, it achieved a balance between real-time detection and high precision for the first time. On this basis, YOLOv11 adopts the Dynamic Sparse Convolution and Hierarchical Decoupled Head architecture, achieving an average precision (AP) of 68.9% on the MS COCO dataset, while increasing the inference speed to 165 frames per second (FPS), which is significantly better than previous models such as YOLOv10.
[0006] However, despite its excellent performance in general scenarios, YOLOv11 still faces significant challenges when directly applied to underwater target detection. For example, the model suffers from performance degradation under complex optical interference, feature confusion due to dense distribution of small targets, high computational overhead, and insufficient real-time performance. Summary of the invention
[0007] Based on this, it is necessary to provide an underwater target detection method and device to improve the underwater target detection performance in response to the above technical problems.
[0008] A method for underwater target detection, comprising: Obtaining the original image containing the underwater target, and generating a data set based on the original image; Obtain a YOLOv11 network model, wherein the network model includes: an input end, a backbone network, a neck network, a head network, and an output end connected in sequence; a multi-scale expansion attention mechanism is integrated into the C2PSA module of the backbone network to form a C2PSA_MSDA module, and a network model to be trained is obtained; the C2PSA_MSDA module includes: a Conv module, one or more PSA_MSDA modules, a Concat module, and a Conv module connected in sequence, and the PSA_MSDA module includes: an MSDA module and two Conv modules connected in sequence; Inputting the data set into the network model to be trained, training the network model to be trained, and obtaining a trained network model; Use the trained network model to detect underwater targets.
[0009] In one embodiment, after forming the C2PSA_MSDA module, a VoV-GSCSPC module is added after each Concat module connected to the head network in the neck network to form a Slim-Neck network to obtain a network model to be trained; The VoV-GSCSPC module includes: a first Conv module, a GSBottleneckC module, a Concat module and a second Conv module connected in sequence. The input end of the VoV-GSCSPC module is connected to the first Conv module and is connected to the Concat module through the third Conv module. The output end of the VoV-GSCSPC module is connected to the second Conv module.
[0010] In one embodiment, the GSbottleneckC module includes two GSConv modules connected in series, one GSConv module is connected to the first Conv module, the other GSConv module is connected to the Concat module, and the first Conv module is further connected to the Concat module via a DWConv module.
[0011] In one embodiment, after forming the Slim-Neck network, an EMA module is added before each Detect module of the head network to obtain a network model to be trained.
[0012] In one embodiment, the operation process of the MSDA module includes: Generate query, key and value through linear projection according to the input feature map; Divide the channel into multiple independent heads, combine the query, key and value to obtain the query, key and value of each head, and perform SWDA operation on each head to extract the key and value of the local space; According to the key and value of the local space, the attention score is calculated and weighted summed; According to the weighted summation result, the outputs of all heads are concatenated and the features are aggregated through a linear layer.
[0013] In one embodiment, the attention score is calculated based on the key and value of the local space, and a weighted sum is performed, including: Calculate the attention score based on the key in the local space; According to the value of the local space and the attention score, a weighted sum is performed to obtain the result of the weighted sum.
[0014] In one embodiment, the operation process of the VoV-GSCSPC module includes:
[0015]
[0016]
[0017] In the formula, For passing The output features after the operation, For the segmentation operation, the input channel is divided into two parts, namely and , is the input feature, is the main branch feature after processing, For stack operations, To perform n times of GSbottleneckC processing, is the main branch of the channel, To output the result, For splicing operation, Another branch of the channel, is a 1×1 convolution operation, is the weight of the convolution.
[0018] In one embodiment, the operation process of the GSConv module includes:
[0019]
[0020]
[0021] In the formula, For depthwise separable convolution , is the depthwise convolution, is the original feature, is a 1×1 convolution operation, is the point-by-point convolution kernel, For standard convolution , is the convolution operation, is the grouped shuffle convolution, is the Swish / Mish nonlinear activation function, and are different channel attention coefficients.
[0022] In one embodiment, the operation process of the EMA module includes: The input feature map is divided into multiple groups, and the output of the 1×1 branch and the 3×3 branch in each group of feature maps is subjected to 2D global average pooling to obtain a global feature vector; According to the global eigenvector, the attention map is calculated through matrix dot product; According to the attention map and the original features, the enhanced features are obtained.
[0023] An underwater target detection device, comprising: An acquisition module is used to acquire the original image containing the underwater target and generate a data set based on the original image; A modeling module is used to obtain a YOLOv11 network model, wherein the network model includes: an input end, a backbone network, a neck network, a head network, and an output end connected in sequence; a multi-scale expansion attention mechanism is integrated into the C2PSA module of the backbone network to form a C2PSA_MSDA module, and a network model to be trained is obtained; the C2PSA_MSDA module includes: a Conv module, one or more PSA_MSDA modules, a Concat module, and a Conv module connected in sequence, and the PSA_MSDA module includes: an MSDA module and two Conv modules connected in sequence; A training module, used for inputting the data set into the network model to be trained, training the network model to be trained, and obtaining a trained network model; The detection module is used to detect underwater targets using the trained network model.
[0024] The above-mentioned underwater target detection method and device combine the MSDA attention mechanism with the backbone network and integrate it into the C2PSA module of YOLOv11, which can help the model improve the detection ability of multi-scale targets, enhance the focus on key areas, suppress the interference of complex background information, and help improve the overall detection and classification capabilities of the model; considering that the introduction of the attention mechanism may make the model more complex, the lightweight convolution GSConv module and VoV-GSCSPC module that reduce the computational complexity are introduced to construct a Slim-Neck network to improve the efficiency of convolution operations, and at the same time more effectively integrate the feature map information of different stages. On the basis of VoV-GSCSPC, the deep separable convolution (DWConv module) and fixed channel ratio are introduced to further reduce the model complexity without losing accuracy, speed up the detection speed, and make the network more adaptable to the requirements of actual detection tasks for detection speed and lightweight; the EMA and Detect modules are integrated to form an enhanced detection head, reorganize the channel dimension and batch dimension, and use cross-dimensional interactions to capture pixel-level relationships. The feature loss in lightweight design is compensated by channel attention recalibration, while reducing computational overhead and retaining the key information of each channel, improving the model's ability to process features. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 A schematic diagram of a flow chart of an underwater target detection method in an embodiment; Figure 2 A network structure diagram of a C2PSA_MSDA module in one embodiment; Figure 3 is an operation process diagram of the MSDA module in one embodiment; Figure 4is a network structure diagram of a VoV-GSCSPC module in one embodiment; Figure 5 is a diagram of the operation process of the GSConv module in one embodiment; Figure 6 is an operation process diagram of an EMA module in one embodiment; Figure 7 A network structure diagram of YOLOv11-MSE in one embodiment; Figure 8 A diagram showing a sample UDD data set in an embodiment; Fig. 9 A schematic diagram of the comparison results between the present application and the YOLO series network in one embodiment; Fig.10 This is one of the comparison graphs of detection performance of the YOLO series model on different target categories in one embodiment; Fig.11 The second comparison chart of the detection performance of the YOLO series models on different target categories in one embodiment; Fig.12 The third comparison chart of detection performance of YOLO series models on different target categories in one embodiment; Fig.13 This is a fourth comparison chart of the detection performance of the YOLO series models on different target categories in one embodiment; Fig.14 This is one of the schematic diagrams of ablation experiment results in one embodiment; Fig.15 This is a second schematic diagram of ablation experiment results in an embodiment; Fig.16 This is a third schematic diagram of ablation experiment results in one embodiment; Fig.17 This is a fourth schematic diagram of ablation experiment results in one embodiment; Fig.18 is a graph showing the relationship between the average precision and the computational complexity of an ablation experiment in one embodiment; Fig.19 One of the comparison graphs of the model loss of the YOLOv11 model and the present application during the training process in one embodiment; Fig. 20 The second comparison diagram of the model loss of the YOLOv11 model in one embodiment and the present application during the training process; Fig.21 This is one of the comparison diagrams of the detection results of different models in one embodiment; Fig. 22 The second comparison diagram of the detection results of different models in an embodiment; Fig.23This is the third comparison chart of the detection results of different models in an embodiment. DETAILED DESCRIPTION
[0026] In order to make the purpose, technical solutions and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.
[0027] In addition, the descriptions of "first", "second", etc. in this application are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, "multiple groups" means at least two groups, such as two groups, three groups, etc., unless otherwise clearly and specifically defined.
[0028] In this application, unless otherwise clearly specified and limited, the terms "connection", "fixation", etc. should be understood in a broad sense. For example, "fixation" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection, an electrical connection, a physical connection, or a wireless communication connection; it can be a direct connection, or an indirect connection through an intermediate medium, or it can be the internal connection of two elements or the interaction relationship between two elements, unless otherwise clearly defined. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0029] In addition, the technical solutions between the various embodiments of the present application can be combined with each other, but it must be based on the fact that ordinary technicians in the field can implement it. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such combination of technical solutions does not exist and is not within the scope of protection required by this application.
[0030] This application provides an underwater target detection method, such as Figure 1 The flowchart shown, in one embodiment, includes: Step 101, obtaining an original image containing an underwater target, and generating a data set based on the original image.
[0031] In this step, how to obtain the original image containing the underwater target and how to generate a data set based on the original image are both existing technologies and will not be described in detail here.
[0032] Step 102, obtain the YOLOv11 network model, the network model includes: an input end, a backbone network, a neck network, a head network and an output end connected in sequence; a multi-scale expansion attention mechanism is integrated into the C2PSA module of the backbone network to form a C2PSA_MSDA module, and a network model to be trained is obtained; the C2PSA_MSDA module includes: a Conv module, one or more PSA_MSDA modules, a Concat module and a Conv module connected in sequence, and the PSA_MSDA module includes: an MSDA module and two Conv modules connected in sequence.
[0033] Specifically: The multi-scale expansion attention mechanism is integrated into the C2PSA module of the backbone network to form the C2PSA_MSDA module and obtain the network model to be trained.
[0034] like Figure 2 As shown, the C2PSA_MSDA module includes: a Conv module, one or more PSA_MSDA modules, a Concat module and a Conv module connected in sequence, and the PSA_MSDA module includes: an MSDA module and two Conv modules connected in sequence.
[0035] like Figure 3 As shown, the operation process of the MSDA module includes: According to the input feature map , generating queries, keys, and values through linear projection:
[0036] Divide the channel into n There are independent heads, and the number of channels in each head is , combining the query, key, and value to get the query, key, and value for each header:
[0037] Perform SWDA operation on each head (this operation is the existing technology) and use the void rate Extract the keys and values of the local space:
[0038]
[0039] According to the key of the local space, the attention score is calculated:
[0040]
[0041] According to the value of the local space and the attention score, a weighted sum is performed to obtain the result of the weighted sum:
[0042] According to the result of weighted summation, the output of all heads is Concatenate and aggregate features through a linear layer:
[0043] In the formula, B is the batch size, C is the number of channels (Channels), H is the height of the feature map, W is the width of the feature map, For query, is the key, For the value, In order to achieve feature space transformation through the fully connected layer, For the i Individual queries, For the i The key of the head, For the i The value of the head, For the segmentation operation, the channel is divided into n A separate head, is the key of the local space, is the value of the local space, To expand the local region into a column vector, is the size of the convolution kernel, is the attention score, is the number of channels or dimensions of each head, is the result of weighted summation, is the output feature map of the MSDA module, For the splicing operation, specifically here the output of all heads is spliced.
[0044] Preferably: The multi-scale dilated attention mechanism is integrated into the C2PSA module of the backbone network to form a C2PSA_MSDA module. A VoV-GSCSPC module is added after each Concat module connected to the head network in the neck network to form a Slim-Neck network, thus obtaining the network model to be trained.
[0045] like Figure 4As shown, the VoV-GSCSPC module includes: a first Conv module, a GSBottleneckC module, a Concat module and a second Conv module connected in sequence, the input end of the VoV-GSCSPC module is connected to the first Conv module and is connected to the Concat module through the third Conv module, and the output end of the VoV-GSCSPC module is connected to the second Conv module; the GSBottleneckC module includes two GSConv modules connected in series, one GSConv module is connected to the first Conv module, and the other GSConv module is connected to the Concat module, and the first Conv module is also connected to the Concat module through a DWConv module.
[0046] The operation process of the VoV-GSCSPC module includes:
[0047]
[0048]
[0049] In the formula, For passing The output features after the operation, For the split operation, the input channel is divided into two parts, namely and , is the input feature, is the main branch feature after processing, For stack operations, To perform n times of GSbottleneckC processing, is the main branch of the channel, The output result is obtained by concatenating the processed main branch features with the other branch of the channel and completing feature fusion through 1×1 convolution. For the splicing operation, here is specifically and To splice, Another branch of the channel, It is a 1×1 convolution operation, which is specifically used here Perform a 1×1 convolution operation on the concatenated features. is the weight of the convolution; The operation process of the GSConv module includes:
[0050]
[0051]
[0052] In the formula, For depthwise separable convolution , is the depthwise convolution, is the original feature, is a 1×1 convolution operation, is the point-by-point convolution kernel, For standard convolution , is the convolution operation, is the grouped shuffled convolution, is the Swish / Mish nonlinear activation function, and are different channel attention coefficients (specifically, learnable channel attention coefficients).
[0053] like Figure 5 As shown in the figure, the GSConv module adopts a dual-branch structure, which reduces the theoretical computational amount (FLOPs) by about 60% compared with the standard convolution, and compensates for the feature loss of DSC through the SC branch. Specifically: the input feature map first passes through the standard convolution layer, then the depth-separable convolution layer, and finally, the outputs of the two convolution layers are concatenated together, and the concatenated feature map is shuffled to optimize the feature representation between channels.
[0054] Further preferably: The multi-scale dilated attention mechanism is integrated into the C2PSA module of the backbone network to form a C2PSA_MSDA module. A VoV-GSCSPC module is added after each Concat module connected to the head network in the neck network to form a Slim-Neck network. An EMA module is added before each Detect module (Detect detection head) of the head network to obtain the network model to be trained.
[0055] like Figure 6 As shown in the figure, the operation process of the EMA module includes: The input feature map Divided into G groups, each group , where the 1×1 branch generates channel attention weights through 1D global average pooling and 1×1 convolution , the 3×3 branch extracts multi-scale spatial features through 3×3 convolution , and perform 2D global average pooling on the outputs of the 1×1 branch and the 3×3 branch to obtain the global feature vector:
[0056]
[0057] Based on the global feature vector, the attention map is calculated through matrix dot product to capture pixel-level relationships:
[0058] According to the attention map and the original features, the enhanced features are obtained:
[0059] In the formula, is the global feature vector output by the 1×1 branch, is the global feature vector output by the 3×3 branch, is the height of the feature map, is the width of the feature map, The original feature After being processed by a 1×1 convolutional layer, the spatial position The output eigenvalue at The original feature After being processed by the 3×3 convolutional layer, the spatial position The output eigenvalue at is the attention map, is the normalized attention weight, For the enhanced features, The original feature.
[0060] In this step, a multi-scale dilated attention mechanism, namely the MSDA attention mechanism, is introduced into the backbone network, and a C2PSA_MSDA module is proposed, which can effectively aggregate semantic information of different scales within the attended receptive field and effectively reduce the redundancy of the self-attention mechanism without the need for complex operations and additional computational costs. Specifically: After the SPPF module of the backbone network in the YOLOv11 network model, the C2PSA module is introduced to enhance the spatial attention in the feature map, thereby allowing the model to focus on the key areas in the image more effectively. Through spatial feature aggregation, YOLOv11 can focus on specific areas of interest, thereby improving the detection accuracy of objects of different sizes and positions; on this basis, the MSDA attention mechanism is introduced in the C2PSA module to separate the channels of the feature map into multiple heads. Each head uses different dilation rates to process different feature subsets in parallel, making full use of the sparsity of the self-attention mechanism at different scales, reducing computational redundancy while maintaining performance, and will not cause a quadratic increase in computational complexity and memory usage, thus expanding the practical application potential of the model; in addition, the MSDA attention mechanism divides the channels of the feature map into different heads and applies different dilation rates and receptive fields in each head. The dilation rate of MSDA is empirically set to [1, 2, 3], covering small, medium and large receptive fields to capture multi-scale contextual information. This strategy enables the model to capture image features at multiple scales, which are then integrated and fed into a linear layer for feature aggregation, allowing the model to understand images at different scales, thereby improving the overall understanding of the image content. Therefore, the C2PSA_MSDA module can not only capture local details, but also perceive contextual information in a wider area, thereby enhancing the expressiveness of the model and solving the problem that the traditional C2PSA module has limitations in adapting to targets of different scales. It is particularly suitable for multi-scale target positioning and classification tasks in complex scenes, especially in the optimization process of underwater target detection models in complex environments, such as simultaneously identifying targets with significant scale differences such as sea urchins, scallops and various types of fish. The present application has significant advantages.
[0061] Furthermore, although the introduction of the attention mechanism can significantly improve the model's ability to detect targets of different scales and focus on key areas, this operation is often accompanied by an increase in the complexity of the model. Therefore, introducing the GSconv module and the VoV-GSCSPC module in the neck network to form a Slim-Neck network can reduce computational complexity and inference latency while ensuring that the accuracy of the model is not affected. In other words, the computational process and network architecture are simplified as much as possible without sacrificing detection accuracy to reduce the complexity of the model. Specifically: In the backbone network of the convolutional neural network (CNN), the input image usually undergoes a similar conversion process: spatial information is gradually transferred to channel information. In this process, the compression of the spatial dimensions (width and height) of the feature map and the expansion of the channel dimension may lead to the loss of semantic information. Channel-intensive convolution calculation can maximize the retention of implicit connections between channels, while channel-sparse convolution completely severs these connections. Therefore, the GSConv module is integrated in the neck network to maintain these connections as much as possible with low time complexity. At the same time, the GSBottleneckC module is designed to further enhance the network's ability to process features and improve the model's learning ability by stacking GSConv modules. In addition, based on the GSBottleneckC module, a one-time aggregation strategy is adopted to construct a cross-level fractional network GSCSPC (cross-level fractional The VoV-GSCSPC module builds a path aggregation feature pyramid network, uses different structural design schemes to improve feature utilization efficiency and network performance, and realizes deep fusion with multi-scale information. The deep separable convolution (DWConv module) and fixed channel ratio are introduced in the VoV-GSCSPC module, which significantly reduces the number of floating-point operations (FLOPs) and parameters. Its lightweight characteristics provide an efficient solution for edge computing. When the model is carried on platforms with limited computing resources such as underwater robots and underwater monitoring equipment, it avoids problems such as large memory usage and slow computing speed. It can realize real-time target detection in resource-constrained underwater environments, thereby expanding the deployment scope of the YOLOv11 model.
[0062] Furthermore, in pursuit of lightweight, Slim-Neck simplifies the model complexity through GSConv modules and lightweight bottleneck structures. However, the underwater environment has characteristics such as low light, water scattering, and target blur. The features of small targets are already weak. After adding Slim-Neck, the number of channels may be reduced and the convolution operation may be simplified, resulting in insufficient extraction of detailed features such as the edges of underwater targets, resulting in missing key information. This is an important challenge for underwater target detection. In order to make up for the missing information and enhance feature extraction, an efficient multi-scale attention (EMA) module is integrated with the Detect detection head. The multi-scale attention mechanism of EMA can focus on the key features of underwater targets by allocating attention weights. At the same time, the attention mechanism of EMA can learn to suppress underwater noise patterns. By identifying noise feature patterns in multi-scale features, low attention weights are allocated, the true features of the target are retained, and Slim-Neck is alleviated. The problem of insufficient anti-noise ability caused by the decrease in complexity is solved, and the robustness of the model in complex underwater environments is improved; the EMA module is designed to retain the information of each channel while reducing the computational overhead. It reshapes the batch dimension of some channels and groups the channel dimension into multiple sub-features so that the spatial semantic features are evenly distributed in each feature group. In addition, the EMA module recalibrates the channel weights in each parallel branch by encoding global information and captures pairwise relationships at the pixel level through cross-dimensional interactions.
[0063] Step 103: input the data set into the network model to be trained, train the network model to be trained, and obtain a trained network model.
[0064] In this step, how to train the network model to be trained and obtain the trained network model belongs to the existing technology and will not be described in detail here.
[0065] Step 104: Use the trained network model to perform underwater target detection.
[0066] In this step, how to use the trained network model to perform underwater target detection belongs to the existing technology and will not be described in detail here.
[0067] It should be noted that the Conv module, the Concat module, the first Conv module, the DWConv module, the second Conv module and the third Conv module are all prior arts, wherein the structures and functions of the first Conv module, the second Conv module and the third Conv module are exactly the same as those of the Conv module.
[0068] In this embodiment, the YOLOv11 network model is improved to obtain a lightweight adaptive detection model based on YOLOv11s, namely YOLOv11-MSE. Figure 7As shown in the figure, YOLOv11-MSE is used to optimize underwater detection performance and further improve the detection accuracy in defect detection: First, MSDA (Multi-Scale Dilated Attention) is introduced in the C2PSA module in the backbone to integrate the multi-scale dilated attention mechanism into the backbone network, so as to dynamically capture the contextual features of multi-scale targets, enhance contextual feature extraction, suppress background noise interference, and minimize computational redundancy; secondly, a Slim-Neck network based on the GSConv module and the VoV-GSCSPC module is designed, and effective and efficient feature fusion of feature maps at different stages is achieved through a hybrid convolution strategy, which significantly reduces the model complexity without losing accuracy and speeds up the detection speed of the model; finally, an efficient multi-scale attention module (EMA) is introduced in the detection head to capture pixel-level relationships through cross-dimensional interactions, enhance the model's key feature expression capabilities, and suppress environmental noise.
[0069] The above underwater target detection method combines the MSDA attention mechanism with the backbone network and integrates it into the C2PSA module of YOLOv11, which can help the model improve the detection ability of multi-scale targets, enhance the focus on key areas, suppress the interference of complex background information, and help improve the overall detection and classification capabilities of the model; considering that the introduction of the attention mechanism may make the model more complex, the lightweight convolution GSConv module and VoV-GSCSPC module that reduce the computational complexity are introduced, and the Slim-Neck network is constructed to improve the efficiency of the convolution operation, while more effectively integrating the feature map information of different stages. Based on VoV-GSCSPC, the deep separable convolution (DWConv module) and fixed channel ratio are introduced to further reduce the model complexity without losing accuracy, speed up the detection speed, and make the network more adaptable to the requirements of actual detection tasks for detection speed and lightweight; the EMA and Detect modules are integrated to form an enhanced detection head, reorganize the channel dimension and batch dimension, use cross-dimensional interactions to capture pixel-level relationships, and compensate for feature loss in lightweight design through channel attention recalibration. At the same time, the computational overhead is reduced and the key information of each channel is retained, thereby improving the model's ability to process features.
[0070] It should be understood that although Figure 1 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.
[0071] In a specific embodiment, the Underwater Detection Dataset (UDD) is used to detect underwater targets in open seawater environments. This groundbreaking dataset contains 2,227 high-resolution 4K images collected under real ocean conditions, annotated with three benthic species: sea cucumbers, sea urchins, and scallops. As the first public dataset to achieve centimeter-level accuracy in complex open sea environments, UDD meets the urgent need for high-quality training data required to improve underwater target recognition.
[0072] The notable features of UDD include: excellent spatial resolution of up to 3840×2160 pixels, which can preserve the fine morphological features of marine organisms; and good ecological validity through natural underwater lighting conditions and real biological distribution. It is worth noting that there is a significant inter-class imbalance in this dataset, with sea urchins accounting for 90.5% of the annotations, sea cucumbers and scallops accounting for 7.6% and 1.9% respectively. In addition, more than 90% of the targets occupy <1.654% of the image area, which poses a major challenge to small target detection.
[0073] The current research challenges facing UDD applications are mainly reflected in three dimensions: (1) The high cost of underwater imaging systems limits data acquisition, which is manifested in the limited size of data sets and differences in data quality across depth gradients.
[0074] (2) There is a serious class imbalance problem in the dataset, and the detection effect of minority class samples (such as sea cucumbers and scallops) is poor. In addition, the objects in underwater images are usually small and densely distributed, which increases the difficulty of detection.
[0075] (3) The computational efficiency requirements of real-time detection on embedded robot platforms require efficient underwater target detection while maintaining detection accuracy.
[0076] The UDD dataset annotates three types of underwater organisms, namely sea cucumbers, sea urchins, and scallops. Figure 8 shown.
[0077] The image characteristics and detection challenges of the three types of targets are shown in Table 1.
[0078] Table 1: Image characteristics and detection challenges of three types of targets in the UDD dataset
[0079] Through the analysis of three types of targets, the common challenges of underwater target detection are summarized: (1) Optical environment interference: Underwater light attenuation and scattering cause image blur and color distortion, which reduces the recognition of target features and affects the detection performance of YOLOv11.
[0080] (2) Small targets and dense distribution: The three types of targets are generally small in size and highly concentrated. The existing detection network is not capable of perceiving the features of small targets, and target differentiation and positioning in dense scenes become technical difficulties.
[0081] (3) Computing resource constraints: The computing power of the edge devices carried by underwater robots is limited. The real-time deployment of high-precision detection models requires a balance between model complexity and reasoning efficiency, which places higher requirements on lightweight network design.
[0082] The UDD dataset is used to train the model. The dataset has a total of 2227 images. The three types of underwater targets on each image are annotated with precise bounding boxes, which provide detailed information about the location, size and category of the target. In addition, the above dataset is randomly divided into a training set and a validation set in a ratio of 8:2.
[0083] The experiment is based on the Pytorch deep learning framework and runs in the Anaconda virtual environment. The specific environment configuration and hyperparameter settings belong to the existing technology.
[0084] The model is evaluated by precision, recall, mean average precision (mAP, including mAP50 and mAP50-95) and computational complexity (GFLOPs), respectively verifying the practical application effect of the proposed YOLOv11-MSE model. It should be noted that the specific definitions and corresponding calculation methods of precision, recall, mean average precision and computational complexity belong to the prior art and will not be repeated here.
[0085] 1. Compare the proposed YOLOv11-MSE network with other networks in the YOLO series. The selected YOLO series networks include YOLOv5, YOLOv9, YOLOv10 and YOLOv11. The results are shown in Table 2 and Fig. 9 shown.
[0086] Table 2: Comparison results between this application and the YOLO series network
[0087] By systematically comparing the performance of the YOLO series networks on the UDD dataset, the effectiveness of the improved YOLOv11-MSE network model is verified. Fig. 9 As shown in the figure, YOLOv11-MSE has achieved the best results in several key indicators such as accuracy, recall, mAP50, and computational complexity. Compared with the performance of different YOLOs in GFLOPs, YOLOv6 has a GFLOPs of up to 44, which has a higher complexity. In contrast, YOLOv5, YOLOv8, YOLOv9, YOLOv10, YOLOv11, and YOLOv12 perform significantly better than YOLOv6 in terms of GFLOPs. Specifically, YOLOv5 has a GFLOPs of 23.8, YOLOv8 has a GFLOPs of 28.4, YOLOv9 has a GFLOPs of 26.7, YOLOv10 has a GFLOPs of 24.5, and YOLOv11 has a GFLOPs of 21.3. YOLOv12 has the lowest computational complexity, with a GFLOPs of only 21.2. From the analysis of the algorithm evolution trend, we can see It can be seen that developers continuously optimize the network architecture and reduce computational complexity during the iteration of the YOLO series networks, so that the network can better adapt to the needs of real-time reasoning and low-power devices; the histogram drawn by the training indicators of this application fluctuates relatively smoothly, but YOLOv11 is significantly better than other versions in all aspects, and is only slightly lower than the latest version of the YOLO series, YOLOv12, in terms of accuracy and GFLOPs. This application takes into account the lightweight and effectiveness of YOLOv11 and selects it as the basis for improvement; after the improvement of this application, YOLOv11 exceeds other network results in the three major indicators of accuracy, recall rate, and mAP50, and achieves the best, while greatly reducing the complexity of the model, which is 6.57% lower than YOLOv11, indicating that this application breaks through the accuracy-efficiency balance bottleneck of the traditional YOLO architecture through innovative module design improvements.
[0088] 2. In order to objectively evaluate the detection performance advantages of the proposed model in class imbalance scenarios, a horizontal comparison experiment was conducted on three types of underwater targets, focusing on the breakthrough progress of small target detection. The evaluation indicators include precision, recall rate, mAP50 and mAP50-95. The detection performance comparison of the YOLO series models on different target categories is shown in Figures 10 to 11. Fig.13 shown.
[0089] It can be seen that the model proposed in this application shows advantages in detecting specific targets in multiple indicators. Although it has not fully surpassed the indicators that some models are good at, it has made significant breakthroughs in small target detection under class imbalance scenarios, especially in scallop detection, achieving significant surpassing of YOLO11, reflecting the model's adaptability to complex scenarios. Specifically, for scallop targets that account for less than 2% of the UDD dataset, the model optimizes the local feature discrimination ability through the multi-scale dilated attention mechanism (MSDA), and combines the EMA module to enhance the feature robustness in a noisy environment, achieving an accuracy and mAP50 improvement of 60.62% and 10.16% respectively compared with the baseline model, effectively alleviating the problem of missed detection of small targets caused by class imbalance; for sea cucumber target detection, although sea cucumber samples account for only 7.6%, the model uses the Slim-Neck lightweight feature fusion strategy and cross-dimensional attention calibration, and the recall rate is improved by 14.2% compared with the optimal comparison model, significantly reducing the missed detection rate under complex backgrounds; and for sea urchin detection, which is dominated by samples (90.5%), the model of this application maintains its competitiveness through global context modeling while introducing a lightweight design to balance computational efficiency. While maintaining the stability of the mAP50-95 indicator, the computational complexity is reduced by 6.57%, verifying the effectiveness of the accuracy-efficiency trade-off. It is worth noting that this application shows significant generalization ability under extreme class imbalance conditions; taking scallop as an example, its mAP50 improvement is an order of magnitude higher than other categories, proving that the multi-scale attention mechanism and cross-dimensional interaction strategy can effectively capture weak features and suppress background interference. In addition, while reducing the computational complexity by 19.1%, this application only causes a slight decrease of 0.012 in mAP50-95, highlighting the superiority of lightweight design.
[0090] In summary, this application has achieved a leap-forward improvement in small target detection performance through modular innovation, especially in low-proportion and high-difficulty targets (such as scallops), providing a new solution for real-time detection of complex underwater scenes.
[0091] 3. In order to intuitively evaluate the impact of different modules on the accuracy of the model, this application conducted a series of ablation experiments based on YOLOv11 to prove the effectiveness of each module. First, YOLOv11s was used as the baseline model to conduct experiments to obtain benchmark detection results as a reference for subsequent experiments. Subsequently, a series of improvements were made to the baseline model to verify its effectiveness. First, the C2PSA module in the YOLOv11 backbone network was optimized, and other parts were kept unchanged to ensure the purity of the experimental results. Similarly, after verifying all the improvement methods one by one, these methods were combined to form the final YOLOv11-MSE model to further verify its effectiveness. The specific ablation experiment results are shown in Table 3 and Figures 14 to 15. Fig.17 shown.
[0092] Table 3: Ablation experiment results
[0093] It can be seen that each model has a large oscillation on the accuracy curve, and the characteristics are analyzed in combination with the underwater small target detection scene. The pixel ratio of underwater small targets (such as sea cucumbers, scallops, etc.) is small, and the feature information is sparse, so the model has natural difficulties in extracting their features. In addition, due to the light attenuation and interference of suspended particles in the underwater environment, the visual distinction between the target and the background is low. The model is easily affected by noise during the learning process, resulting in blurred judgment boundaries for small targets, which in turn causes fluctuations in detection accuracy (Precision). At the same time, the number of samples of categories such as sea urchins in the dataset is significantly higher than that of sea cucumbers and scallops. The gradient contribution of minority class samples during training is insufficient, and the model's detection accuracy for minority classes is prone to an "overfitting-underfitting" cycle, exacerbating Precision oscillations. As shown in Table 3, compared with exp1 and exp2, after adding G2PSA_MSDA alone, the model only increases GFLOPs by 0.47%, and all other indicators are improved, including Precision by 3.46%, Recall by 0.33%, mAP50 by 1.65%, and mAP50-95 from 0.299 to 0.302, which is the best among all models, indicating that its multi-scale dynamic attention mechanism effectively enhances the feature representation ability of underwater small targets. Slim-Neck reduces GFLOPs from 21.3 to 19.5, a decrease of 19.1%, by cascading GSConv and VoV-GSCSPC modules while maintaining a decrease of mAP50-95 of less than or equal to 0.012. Its lightweight efficiency exceeds similar solutions, verifying the effectiveness of structural lightweight design. After adding the EMA module to epx4, Recall increased from 0.61 to 0.695, an increase of 12.23%, and mAP50 reached 0.686, indicating that the module enhances the stability and generalization ability of model training, reflecting that EMA enhances the detection ability of difficult samples (such as sea cucumbers and scallops) through feature recalibration.
[0094] As shown in Figure 18, after adding C2PSA_MSDA, Slim-Neck and EMA, the detection accuracy and precision reached a new breakthrough, with mAP50 increasing from 0.666 to 0.689, an increase of 3.45%, and Precision also significantly increasing by 9.67%, both of which ranked the best among all models. In addition, the model GFLOPs was compressed by 6.57%. In summary, YOLOv11-MSE successfully solved the three core contradictions in underwater target detection: (1) the contradiction between the lack of small target features and computational complexity; (2) the contradiction between class imbalance and high false detection rate; (3) the contradiction between model lightweight and accuracy.
[0095] 4. To further verify the optimization effect of this application on model convergence, Figures 19 and Fig. 20 A comparison chart of the model loss of the YOLOv11 model and the present application during the training process is presented. The comparison shows that the loss value of the YOLOv11 model is higher than that of the improved model of the present application, whether in the training stage or the verification stage. Specific data show that the final loss values of the present application model on the training set and the verification set are 1.18515 and 1.32397 respectively, while the loss values of the YOLOv11 model on the training set and the verification set are 1.18886 and 1.33960. This result fully demonstrates that the improved model has achieved significant improvements in training stability and convergence performance.
[0096] 5. In order to intuitively verify the effectiveness of this application compared with other models, we discussed it from the perspective of qualitative analysis. We selected three representative detection images in the test set as cases to test the actual performance of different YOLO series models in underwater target detection tasks. Figure 21 to Figure 23 The comparison chart of detection results of different models is shown. The five rows in the figure correspond to the prediction results of the original image, YOLOv10, YOLOv11, YOLOv12 and the model YOLOv11-MSE of the present application.
[0097] In order to verify the detection performance of this application in practical applications, a series of experiments were performed using the test set retained in the data set to evaluate the adaptability of this application in dynamic hydrological conditions and underwater environments with different visibility. The core purpose of these experiments is to evaluate the performance of this application in real-world environments. Fig.21 As shown in the figure, in an environment with bright and clear lighting conditions, several sea urchins and sea cucumbers are located at the edge of the image. YOLOv11 fails to identify these creatures, while YOLOv11-MSE successfully detects them, and the confidence is significantly improved, showing a strong edge detection capability. Fig. 22 As shown in the figure, in a blurry and dark light environment, due to the low visibility underwater, YOLOv11 misidentifies the end of the seaweed in the water as a sea cucumber, and the black shadow on the edge as a sea urchin. Other YOLO series models also have serious misjudgments and cannot distinguish between seaweed and sea cucumbers with similar colors and shapes. However, YOLOv11-MSE successfully detected the sea cucumber, proving the effectiveness of its improved attention mechanism under blurry and dark light conditions. Fig.23 As shown, although there is no Fig. 22 It is so dark that target detection is still affected to a certain extent. For scallops trapped in the mud on the seabed, YOLOv11 and YOLOv12 both missed detections and misidentified seaweed as sea cucumbers. However, YOLOv11-MSE identified and located the target more accurately.
[0098] Experiments have shown that when this application is compared with YOLOv5, YOLOv8, YOLOv9, YOLOv10, YOLOv11 and the latest version of YOLOv12, the accuracy, recall rate and mAP50 of YOLOv11-MSE detection are the highest, and the computational complexity (GFLOPs) is the lowest at 19.9, which is 6.57% lower than that of YOLOv11. This application has improved the detection accuracy of various types of targets, especially for unbalanced targets such as scallops, and the model has the best overall performance, which further verifies the feasibility and practicality of YOLOv11-MSE in the field of underwater target detection. In addition, the effectiveness of each module was verified through ablation experiments, that is, the C2PSA module was improved in the original network structure, the cross-level local network VoV-GSCSPC module was designed using GSConv and one-time aggregation methods, and Slim-Neck was constructed. The enhanced detection head was further improved, the EMA attention mechanism was integrated, the noise was dynamically suppressed, and the discriminative features were amplified, especially for challenging categories such as scallops and sea cucumbers, so that the model performance was the best. Finally, compared with YOLOv11, the recall rate (R) of YOLOv11-MSE increased by 1.31%, the precision rate (P) increased by 9.67%, and the mAP50 increased by 3.45%.
[0099] In summary, experiments based on the UDD dataset show that compared with the baseline model YOLOv11, YOLOv11-MSE achieves 9.67% and 3.45% improvements in detection accuracy and mean average precision (mAP50), respectively, while reducing computational complexity by 6.57%. Ablation experiments further verify the collaborative optimization effect of each module, especially in the category imbalance scenario, the detection accuracy of rare category targets (such as scallops) is significantly improved, and its accuracy and mAP50 are increased by 60.62% and 10.16%, respectively. In addition, a large number of experiments conducted on the underwater detection dataset (UDD) show that YOLOv11-MSE achieves a mean average precision (mAP50) of 68.9% under a threshold of 0.5, which is 3.45% higher than YOLOv11, and the amount of computation (GFLOPs) is reduced by 6.57%. This application provides an efficient solution for edge computing scenarios such as underwater robots and ecological monitoring through lightweight design and resistance to optical degradation, and sets a new benchmark for the balance between accuracy, efficiency and adaptability in underwater target detection.
[0100] The present application also provides an underwater target detection device, which in one embodiment includes: an acquisition module, a modeling module, a training module and a detection module, wherein: An acquisition module is used to acquire the original image containing the underwater target and generate a data set based on the original image; A modeling module is used to obtain a YOLOv11 network model, wherein the network model includes: an input end, a backbone network, a neck network, a head network, and an output end connected in sequence; a multi-scale expansion attention mechanism is integrated into the C2PSA module of the backbone network to form a C2PSA_MSDA module, and a network model to be trained is obtained; the C2PSA_MSDA module includes: a Conv module, one or more PSA_MSDA modules, a Concat module, and a Conv module connected in sequence, and the PSA_MSDA module includes: an MSDA module and two Conv modules connected in sequence; A training module, used for inputting the data set into the network model to be trained, training the network model to be trained, and obtaining a trained network model; The detection module is used to detect underwater targets using the trained network model.
[0101] For the specific definition of an underwater target detection device, please refer to the definition of an underwater target detection method above, which will not be repeated here. Each module in the above device can be implemented in whole or in part by software, hardware and a combination thereof. Each of the above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.
[0102] The contents not described in detail in this specification belong to the prior art known to professional and technical personnel in this field.
[0103] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0104] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A method for underwater target detection, characterized in that: include: Obtaining the original image containing the underwater target, and generating a data set based on the original image; Obtain a YOLOv11 network model, wherein the network model includes: an input end, a backbone network, a neck network, a head network, and an output end connected in sequence; a multi-scale expansion attention mechanism is integrated into the C2PSA module of the backbone network to form a C2PSA_MSDA module, and a network model to be trained is obtained; the C2PSA_MSDA module includes: a Conv module, one or more PSA_MSDA modules, a Concat module, and a Conv module connected in sequence, and the PSA_MSDA module includes: an MSDA module and two Conv modules connected in sequence; Inputting the data set into the network model to be trained, training the network model to be trained, and obtaining a trained network model; Use the trained network model to detect underwater targets.
2. The underwater target detection method according to claim 1, characterized in that: After forming the C2PSA_MSDA module, a VoV-GSCSPC module is added after each Concat module connected to the head network in the neck network to form a Slim-Neck network, thereby obtaining a network model to be trained; The VoV-GSCSPC module includes: a first Conv module, a GSBottleneckC module, a Concat module and a second Conv module connected in sequence. The input end of the VoV-GSCSPC module is connected to the first Conv module and is connected to the Concat module through the third Conv module. The output end of the VoV-GSCSPC module is connected to the second Conv module.
3. The underwater target detection method according to claim 2, characterized in that: The GSbottleneckC module includes two GSConv modules connected in series, one GSConv module is connected to the first Conv module, the other GSConv module is connected to the Concat module, and the first Conv module is further connected to the Concat module via a DWConv module.
4. The underwater target detection method according to claim 3, characterized in that: After forming the Slim-Neck network, an EMA module is added before each Detect module of the head network to obtain a network model to be trained.
5. The underwater target detection method according to any one of claims 1 to 4, characterized in that: The operation process of the MSDA module includes: Generate query, key and value through linear projection according to the input feature map; Divide the channel into multiple independent heads, combine the query, key and value to obtain the query, key and value of each head, and perform SWDA operation on each head to extract the key and value of the local space; According to the key and value of the local space, the attention score is calculated and weighted summed; According to the weighted summation result, the outputs of all heads are concatenated and the features are aggregated through a linear layer.
6. The underwater target detection method according to claim 5, characterized in that: According to the key and value of the local space, the attention score is calculated and weighted summation is performed, including: Calculate the attention score based on the key in the local space; According to the value of the local space and the attention score, a weighted sum is performed to obtain the result of the weighted sum.
7. An underwater target detection method according to any one of claims 2 to 4, characterized in that: The operation process of the VoV-GSCSPC module includes: In the formula, For passing The output features after the operation, For the segmentation operation, the input channel is divided into two parts, namely and , is the input feature, is the main branch feature after processing, For stack operations, To perform n times of GSbottleneckC processing, is the main branch of the channel, To output the result, For splicing operation, Another branch of the channel, is a 1×1 convolution operation, is the weight of the convolution.
8. An underwater target detection method according to claim 3 or 4, characterized in that: The operation process of the GSConv module includes: In the formula, For depthwise separable convolution , is the depthwise convolution, is the original feature, is a 1×1 convolution operation, is the point-by-point convolution kernel, For standard convolution , is the convolution operation, is the grouped shuffled convolution, is the Swish / Mish nonlinear activation function, and are different channel attention coefficients.
9. The underwater target detection method according to claim 4, characterized in that: The operation process of the EMA module includes: The input feature map is divided into multiple groups, and the output of the 1×1 branch and the 3×3 branch in each group of feature maps is subjected to 2D global average pooling to obtain a global feature vector; According to the global eigenvector, the attention map is calculated through matrix dot product; According to the attention map and the original features, the enhanced features are obtained.
10. An underwater target detection device, characterized in that: include: An acquisition module is used to acquire the original image containing the underwater target and generate a data set based on the original image; A modeling module is used to obtain a YOLOv11 network model, wherein the network model includes: an input end, a backbone network, a neck network, a head network, and an output end connected in sequence; a multi-scale expansion attention mechanism is integrated into the C2PSA module of the backbone network to form a C2PSA_MSDA module, and a network model to be trained is obtained; the C2PSA_MSDA module includes: a Conv module, one or more PSA_MSDA modules, a Concat module, and a Conv module connected in sequence, and the PSA_MSDA module includes: an MSDA module and two Conv modules connected in sequence; A training module, used for inputting the data set into the network model to be trained, training the network model to be trained, and obtaining a trained network model; The detection module is used to detect underwater targets using the trained network model.
Citation Information
Patent Citations
Underwater biological detection method based on improved YOLOv7 network
CN119068323A
Marine organism detection method, device and equipment and storage medium
CN119274205A
Cited By
Mechanical part defect detection method and system based on improved YOLOv12 model
CN120279020A
Wheat basal stem rot identification method, system, equipment and medium
CN121214232A
A river pollution outlet detection method, system and device
CN122598005A
A river pollution outlet detection method, system and device
CN122598005B