Underwater blurred image target detection method and device
By improving the feature extraction and fusion network of the YOLO11 model, the problem of low detection accuracy in underwater blurred images is solved, achieving efficient and accurate underwater target detection. It is suitable for the special characteristics of the underwater environment and is suitable for mobile deployment.
Patent Information
- Application Number
- CN202511344476.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-09-19
AI Technical Summary
Underwater image target detection suffers from problems such as low detection accuracy, poor precision, and high false detection and false negative rates. In particular, the detection performance of the YOLO11 model is limited in underwater blurred image environments.
The feature extraction and feature fusion networks of the YOLO11 model are improved by replacing the CBS module with a hybrid pooling convolution module, adding hybrid pooling operations, optimizing the SPPF module, introducing a P2 connection path in the feature fusion network, and using BCE and DFL/CIOU loss functions for training, while optimizing the weighting of the loss function.
It improves the accuracy and efficiency of target detection in underwater blurred images, reduces the false negative rate, has a small number of model parameters, is easy to deploy on mobile devices, has strong practicality and generalization ability, and its detection performance has been verified on public datasets.
Smart Images

Figure CN120833548A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of image recognition, and particularly relates to an underwater blurred image target detection method and device. BACKGROUND
[0002] Unlike the conventional target detection environment, due to the existence of a large amount of suspended medium in the underwater environment and the influence of absorption and scattering of light underwater, the underwater image has problems such as detail loss, imaging blur, low contrast, and serious image degradation. In addition, underwater organisms are often densely distributed, and have problems such as small targets, stacking, and overlapping. Therefore, underwater image target detection has problems such as low detection accuracy, poor accuracy, false detection, and high miss detection rate.
[0003] In recent years, with the rapid development of deep learning technology, the target detection technology has made great progress. The YOLO series is one of the classic models in the field of target detection, and is known for its high real-time performance, fast detection speed, and small parameter quantity. YOLO11 is an iterative version of YOLOV8 and is one of the most advanced YOLO models, with higher accuracy and detection efficiency. However, when it comes to the detection of underwater blurred images, the detection performance of YOLO11 is limited by the underwater environment, and cannot well play the effectiveness of detection. SUMMARY
[0004] The present application provides an underwater blurred image target detection method and device, aiming to solve the problem of detecting blurred images in underwater scenes and improve the accuracy and efficiency of detection.
[0005] To achieve the above technical purposes, the present application adopts the following technical solutions: An underwater blurred image target detection method, comprising: Step 1, collect an underwater blurred image dataset, and perform class labeling and bounding box labeling on the target objects in the images to obtain a labeled dataset; Step 2, build a neural network model, and use the labeled dataset to train the built neural network model to obtain a target detection model; The neural network model is improved from YOLO11n, and the improvement includes: (1) replacing the 3rd, 4th, and 5th CBS modules of the feature extraction network of YOLO11n with a hybrid pooling convolution module, and (2) improving the SPPF module of the feature extraction network of YOLO11n: adding a hybrid pooling operation between the 1st CBS module and the 1st max pooling of the SPPF module and as a residual input branch of Concat in the SPPF module; Step 3, using the trained target detection model, performing target detection on the underwater blurred image to be detected, and outputting an image with a bounding box and a target class in the bounding box.
[0006] Further, the mixed-pooling convolution module is expressed as: ; Wherein, Y represents the output feature map of the mixed-pooling convolution module, X represents the input feature map of the mixed-pooling convolution module, SiLU represents the activation function, BN represents the batch normalization operation, conv2d represents the two-dimensional convolution operation, Maxpool represents the maximum pooling operation, and Avgpool represents the average pooling operation.
[0007] Further, the improvement of the neural network model on YOLO11n also includes: for the feature map P2 obtained by the first C3K2 module of the feature extraction network of YOLO11n, adding a connection path at the lowermost layer of the feature fusion network of YOLO11n: processing and splicing the feature map P2 and the last obtained feature map from top to bottom, and then performing fusion processing from bottom to top.
[0008] Further, based on the improved feature fusion network of YOLO11n, the feature fusion network includes two parts of a top-down path and a bottom-up path. The top-down path starts from the up-sampling operation of the topmost output X5 of the feature extraction network, and the up-sampling operation includes, from input to output, nearest neighbor up-sampling, a Concat module, and a C3K2 module, and generates feature maps X4, X3, and X2 through three up-sampling operations; wherein the Concat module of each up-sampling operation performs feature fusion on the feature map obtained by the nearest neighbor up-sampling and the corresponding size feature maps P4, P3, and P2 obtained by the feature extraction network through horizontal connection. The bottom-up path starts from the down-sampling operation of the lowermost output of the top-down path, and the down-sampling operation includes, from input to output, a CBS module, a Concat module, and a C3K2 module, and generates feature maps F3, F4, and F5 through three down-sampling operations, as inputs of the detection module; wherein the Concat module of each down-sampling operation performs feature fusion on the feature map obtained by the CBS module and the corresponding size feature maps X3, X4, and X5 of the top-down path through horizontal connection.
[0009] Further, the loss function of training the neural network model to obtain the target detection model adopts BCE as the classification loss function of the target class, and adopts DFL and CIOU as the loss function of the bounding box regression, and then the classification loss function and the regression loss function are weighted to obtain the overall loss function, which is expressed as: ; In the formula, the overall loss function is represented as , , respectively represent BCE classification loss, DFL regression loss, CIOU regression loss, respectively are , , corresponding weights.
[0010] An underwater blurred image target detection device comprises: A data collection and labeling module is configured to collect an underwater blurred image dataset, label target objects in the images in terms of categories and bounding boxes, and obtain a labeled dataset. A model training module is configured to train a neural network model that has been built using the labeled dataset, and obtain a target detection model. The neural network model is improved from YOLO11n, and the improvement includes: (1) replacing the third, fourth and fifth CBS modules of the feature extraction network of YOLO11n with a hybrid pooling convolution module, and (2) improving the SPPF module of the feature extraction network of YOLO11n: adding a hybrid pooling operation between the first CBS module and the first max pooling of the SPPF module and using it as a residual input branch of Concat in the SPPF module. A target detection module is configured to configure the target detection model obtained through training to perform target detection on an underwater blurred image to be detected, and output an image with a bounding box and a target category in the bounding box.
[0011] The underwater blurred image target detection method and device provided by the application have a small number of model parameters, can fully adapt to the particularity of the underwater environment, realize efficient and accurate detection of underwater blurred images, have a low rate of missed detection, and can effectively complete the task requirements of underwater blurred image detection. Compared with existing underwater image detection technologies, the application has the following advantages: (1) The network model has a small number of parameters, high efficiency, and is easy to deploy on mobile terminals.
[0012] (2) The application designs a hybrid pooling convolution and feature fusion path optimization, improves the feature extraction capability of the model, effectively avoids the loss of important detail information due to the complexity of the underwater environment, and further improves the performance of the target detection model.
[0013] (3) It has strong practicality and generalization ability. The application achieves good results on the public dataset URPC2020, and the detection performance of underwater blurred images is fully verified. BRIEF DESCRIPTION OF DRAWINGS
[0014] Figure 1 It is a network structure diagram for underwater blurred image target detection of the embodiment of the application.
[0015] Figure 2 Structure diagram of a mixed-pooling convolution (MAConv) structure for an embodiment of the present application.
[0016] Figure 3 Structure diagram of an improved SPPF (SPPF_improved) obtained by improving the SPPF for an embodiment of the present application.
[0017] Figure 4 Comparison diagram of detection effects of an embodiment of the present application and YOLO11n on a first picture in a URPC2020 test set.
[0018] Figure 5 Comparison diagram of detection effects of an embodiment of the present application and YOLO11n on a second picture in a URPC2020 test set. DETAILED DESCRIPTION
[0019] The embodiments of the present application are described in detail below, which are based on the technical solutions of the present application, and give detailed implementation manners and specific operation processes, and further explain the technical solutions of the present application.
[0020] The embodiment provides a method for detecting underwater blurred image targets, which can effectively enhance the feature extraction capability of underwater targets, fully fuse detailed information, and achieve higher detection precision and efficiency. Referring to Figure 1 , the method comprises the following steps.
[0021] Step 1, a real underwater scene underwater blurred image dataset is constructed, including a model training set, a verification set and a test set, the training set is used for model training, the verification set is used for verifying the precision of the trained model, and the test set is used for testing the generalization precision of the model.
[0022] The dataset constructed in the embodiment is a target detection dataset URPC2020 opened by an underwater target detection competition-optical event, which contains 4 main categories of sea cucumbers, sea urchins, starfish and scallops, a total of 7543 pictures, including 5543 training set images, 800 test set A images and 1200 test set B images. During the experiment, the training set images are randomly divided into an experimental training set and an experimental verification set according to 8:2. The resolution of the images includes multiple scales such as 720*405, 1920*1080 and 3840*2160.
[0023] Step 2.1, a neural network model for detecting underwater blurred image targets is designed, which is improved from YOLO11n and includes three parts of a feature extraction network, a feature fusion network and a detection module.
[0024] The feature extraction network is implemented as follows: for the training set data, first, the input image is preprocessed, including image scaling, normalization, mosaic data enhancement. The image scaling operation adjusts the input images of different sizes to a uniform size by the LetterBox scaling method, while maintaining the aspect ratio to avoid distortion, and the other parts are filled with gray, while adjusting the channel number; the normalization linearly maps the image pixel value from [0, 255] to [0, 1], and normalizes it to a distribution with a mean of 0 and a standard deviation of 1, accelerating model convergence; the core idea of this method is to randomly combine multiple (usually four) pictures into a new training picture, thereby enriching the background information and target diversity, and improving the generalization ability of the model. The implementation process of the Mosaic data enhancement method is as follows:
[0025] (1) Randomly select pictures: randomly select four pictures from the original data set, and randomly select picture splicing reference point coordinates (x, y).
[0026] (2) Random scaling and cropping: randomly scale and randomly crop each picture so that they can fit the splicing canvas and determine the position of each picture in the final combined picture.
[0027] (3) Picture splicing: place the four pictures in the four quadrants (top left, top right, bottom left, bottom right) of a large canvas, and adjust the position of each picture according to the randomly generated reference point position.
[0028] (4) Adjust the label: adjust the bounding box label of the target in each picture according to the coordinates of the spliced new picture.
[0029] Through the above preprocessing operation, a gray image with a size of 640*640*3 can be obtained as the input of the network, and then two CBS convolution downsampling and a C3K2 module are used to obtain a shallow feature map P2. If only ordinary CBS convolution is used for downsampling, the details of the target and part of the local features may be lost, and underwater targets often have high similarity to the background and are densely distributed, and the input image noise is obvious. In order to better extract image features, the present application improves the use of a hybrid pooling convolution MAConv and a C3K2 module as a group of feature extraction modules. The hybrid pooling convolution processes the input feature map in parallel through maximum pooling and average pooling, which can not only preserve the local features such as edges and textures of objects, but also smooth local fluctuations, suppress noise, provide global smooth features, highlight background information, and form a more comprehensive feature representation, thereby improving the underwater target detection accuracy. The structure of the hybrid pooling convolution MAConv is shown in Figure 2 , which can be expressed by the formula: ; Wherein Y represents the output feature map of the mixed pooling convolution module, X represents the input feature map of the mixed pooling convolution module, SiLU represents an activation function, BN represents a batch normalization operation, conv2d represents a two-dimensional convolution operation, Maxpool represents a maximum pooling operation, and Avgpool represents an average pooling operation.
[0030] Through the three groups of feature extraction modules, feature maps P3, P4 and P5 containing high semantic information can be sequentially obtained, wherein P2, P3 and P4 are added to the feature fusion path of the feature fusion network through Concat operation.
[0031] In order to make the output of the feature extraction network contain more rich positioning information and semantic information, the spatial pyramid pooling layer SPPF is further improved in the application, after the channel dimension reduction through the CBS convolution, the above-mentioned mixed pooling operation is introduced to increase a residual input branch of Concat, so as to integrate more fine features, make the prediction box closer to the real box, and improve the detection precision of the model, the structure of SPPF_improved is shown in Figure 3 . Then P5 continues to perform deep feature extraction through the improved spatial pyramid pooling layer SPPF_improved and the multi-head attention module C2PSA, and capture object information of different scales, so as to obtain a multi-scale feature representation X5 as the output of the feature extraction network and as the input of the feature fusion network.
[0032] The feature fusion network adopts a path aggregation network PAN as a feature extraction method, and a P2 connection path is added, the P2 feature map is obtained by processing the original input picture through two CBS convolutions and a C3K2 module, has a large resolution, and contains more rich position information, by adding the P2 connection path, the high semantic feature map is integrated with more accurate position information, so that more detail information is retained, and the problem of insufficient information fusion in the forward propagation process of the network is effectively solved. The feature fusion network is divided into an up-down path and a down-up path. The up-down path starts from the output X5 of the feature extraction network and performs an upsampling operation, the upsampling operation includes making the input sequentially perform nearest neighbor upsampling, Concat, C3K2 processing, and through three upsampling operations, feature maps X4, X3 and X2 are sequentially generated, wherein Concat represents that the feature maps P4, P3 and P2 of the corresponding size are fused through horizontal connection. The down-up path starts from the output X2 of the up-down path and performs a downsampling operation, in this process, the information of each layer of feature map is fused with the feature map of the corresponding size of the up-down path, and finally the feature maps F3, F4 and F5 of the same size as P3, P4 and P5 are obtained as the input of the detection module.
[0033] The detection module is implemented by taking the output feature maps F3, F4, and F5 as the inputs of the three detection heads, obtaining the model's original prediction output and calculating the corresponding loss function to guide model optimization. The loss function uses BCE as the classification loss function, DFL and CIOU as the regression loss function, and the total loss is obtained through a weighted sum, which can be expressed as follows: Finally, the model’s raw prediction output is converted into the final underwater target detection output through post-processing operations.
[0034] Of the total loss 、 、 Represent BCE classification loss, DFL regression loss, and CIOU regression loss respectively. They are 、 、 The corresponding weights can be adjusted according to actual needs. Specifically: ; ; ; in, Indicates the number of target categories, Indicates the The true label value of the class sample, Indicates the model The predicted probability of class samples.
[0035] represents the true label value of the sample, 、 Represents the distance to the true label The nearest two integers, Indicates that the model prediction output is equal to The probability of Indicates that the model prediction output is equal to probability.
[0036] Indicates the distance between the predicted box and the center point of the real box, is the diagonal distance of the minimum circumscribed rectangle, Relative value indicating the aspect ratio of the box. is the smoothing coefficient, and IOU is the intersection-over-union ratio, which is an important indicator to measure the similarity between the predicted value and the true value. It represents the ratio of the common part of the predicted box and the true value box area to the total area of the two. The larger the IOU value, the closer the predicted structure is to the true value, and the better the model performance.
[0037] The post-processing operation mainly includes three key steps: (1) Eliminate low-probability prediction boxes by confidence filtering; (2) Decode the bounding box, convert the offset in the distribution form to the actual coordinates; (3) Remove overlapping redundant boxes using non-maximum suppression (NMS) and restore the coordinates to the original image size; (4) Finally output the bounding box, confidence and class information of each target.
[0038] Step 2.2, training a neural network to obtain a target detection model, including training the underwater blurred image target detection network designed in step 2; the experimental environment of the embodiment is shown in Table 1.
[0039] During the training process, some hyperparameters are set as follows: the input picture size is 640*640, the number of worker threads is set to 8, the batch size is 32, SGD is used as the optimizer, the initial learning rate is 0.01, the number of iterations is 150, no pre-trained weights are used, and other parameters are default values.
[0040] ; Step 3, verify the target detection model obtained by training on the test set, and the embodiment specifically tests on the URPC2020 public data set.
[0041] The embodiment uses precision (Precision), recall (Recall), and mean average precision (mAP) as performance evaluation indicators of the model. The calculation formula is as follows: ; ; ; Where TP (True Positive) represents the number of samples that the model predicts as positive and is actually positive; FP (False Positive) represents the number of samples that the model predicts as positive but is actually negative; TN (True Negative) represents the number of samples that the model predicts as negative and is actually negative; FN (False Negative) represents the number of samples that the model predicts as negative but is actually positive. AP i represents the average precision of the i-th class, and the value is equal to the area surrounded by the i-th class precision-recall curve and the coordinate axis, which represents the precision performance of the model at various recall rates.
[0042] To facilitate understanding of the effectiveness of the technical effects of the present application, the ablation experiment and the comparative experiment of the present application are shown in Tables 2 and 3.
[0043] ; ; According to the results in Tables 2 and 3, it can be known that the improved method proposed in the application effectively improves the performance of the model, compared with the model detection method based on the original YOLO11n, the recall rate of the model based on the improved YOLO11n network structure in the application is increased by 3.5%, mAP50 is increased by 2.3%, mAP50-95 is increased by 3%, and the parameter amount is only increased by 0.58M. The results show that the method of the application effectively improves the missed detection and repeated detection of targets in the underwater scene, and significantly improves the recognition accuracy and positioning precision of underwater targets. Figure 4 and Figure 5 Figures 1 and 2 are comparison diagrams of the target detection effects of the method of the application and YOLO11n on two different underwater blurred images, respectively, wherein, Figure 4 (a) and Figure 5 (a) is an original underwater blurred image, Figure 4 (b) and Figure 5 (b) is the target detection result of the original underwater blurred image based on the YOLO11n network structure, Figure 4 (c) and Figure 5 (c) is the target detection result of the original underwater blurred image based on the improved YOLO11n network structure in the method of the application, wherein the boxes of different colors represent different class labels.
[0044] The above embodiments are preferred embodiments of the application, and those skilled in the art can also make various transformations or improvements on the basis of the above embodiments, and these transformations or improvements should be within the scope of protection of the application without departing from the general concept of the application.
Claims
1. An underwater blurred image object detection method, characterized in that, The application relates to an underwater fuzzy image target detection method. Step 1, collecting an underwater fuzzy image dataset, performing category labeling and boundary box labeling on target objects in the images to obtain a labeled dataset; Step 2, building a neural network model and training the built neural network model using the labeled dataset to obtain a target detection model; The neural network model is improved based on YOLO11n, and the improvement includes the following two aspects: (1) replacing the third, fourth and fifth CBS modules of the feature extraction network of YOLO11n with a mixed pooling convolution module, and (2) improving the SPPF module of the feature extraction network of YOLO11n: adding a mixed pooling operation between the first CBS module and the first maximum pooling of the SPPF module and taking the mixed pooling operation as a residual input branch of Concat in the SPPF module; Step 3, using the trained target detection model to perform target detection on the underwater fuzzy image to be detected, and outputting an image with a boundary box and a target category in the boundary box.
2. The underwater blurred image object detection method of claim 1, wherein, The expression of the mixed pooling convolution module is as follows: ; Wherein Y represents the output feature map of the mixed pooling convolution module, X represents the input feature map of the mixed pooling convolution module, SiLU represents an activation function, BN represents a batch normalization operation, conv2d represents a two-dimensional convolution operation, Maxpool represents a maximum pooling operation, and Avgpool represents an average pooling operation. 3.The underwater blurred image object detection method of claim 1, wherein, The improvement of the neural network model on YOLO11n also includes the following: for the feature map P2 obtained by the first C3K2 module of the feature extraction network of YOLO11n, a connection path is added at the lowermost layer of the feature fusion network of YOLO11n: the feature map P2 and the last obtained feature map from top to bottom are processed and spliced, and then fusion processing is performed from bottom to top.
4. The underwater blurred image object detection method of claim 3, wherein, The improved feature fusion network based on YOLO11n includes two parts, i.e. a top-down path and a bottom-up path. The top-down path starts from the up-sampling operation of the topmost output X5 of the feature extraction network, and the up-sampling operation includes, from input to output, a nearest neighbor up-sampling, a Concat module and a C3K2 module, and the feature maps X4, X3 and X2 are generated through three up-sampling operations; the Concat module of each up-sampling operation performs feature fusion on the feature map obtained through the nearest neighbor up-sampling and the feature map P4, P3 and P2 of the corresponding size obtained through the feature extraction network through horizontal connection; The bottom-up path starts from the down-sampling operation of the lowermost output of the top-down path, and the down-sampling operation includes, from input to output, a CBS module, a Concat module and a C3K2 module, and the feature maps F3, F4 and F5 are generated through three down-sampling operations and taken as the input of the detection module; The Concat module of each down-sampling operation performs feature fusion on the feature map obtained through the CBS module and the feature maps X3, X4 and X5 of the corresponding size of the top-down path through horizontal connection.
5. The underwater blurred image object detection method of claim 1, wherein, The loss function of training the neural network model to obtain the target detection model adopts BCE as the classification loss function of the target class, adopts DFL and CIOU as the loss function of the boundary box regression, and then weights the classification loss function and the regression loss function to obtain the overall loss function, which is represented as: ; In the formula, The loss function represents the whole, 、 、 BCE classification loss, DFL regression loss, CIOU regression loss respectively, Respectively 、 、 The corresponding weight.
6. An underwater blurred image object detection device, characterized by, The method comprises the following steps: a data collection and labeling module, configured to: collect an underwater fuzzy image dataset, and perform class labeling and boundary box labeling on the target in the image to obtain a labeled dataset; a model training module, configured to: train the neural network model built to obtain the target detection model using the labeled dataset; wherein the neural network model is improved from YOLO11n, and the improvement comprises: (1) replacing the third, fourth and fifth CBS modules of the feature extraction network of YOLO11n with a hybrid pooling convolution module, and (2) improving the SPPF module of the feature extraction network of YOLO11n: adding a hybrid pooling operation between the first CBS module and the first maximum pooling of the SPPF module as a residual input branch of Concat in the SPPF module; a target detection module, configured to: configure the target detection model trained to perform target detection on an underwater fuzzy image to be detected, and output an image with a boundary box and a target class in the boundary box.
Citation Information
Patent Citations
Underwater target detection method, device, equipment and medium
CN117115632A
Target detection method, device and equipment based on underwater blurred image
CN119380005A
Cited By
Unmanned aerial vehicle monitoring data processing method and system based on intelligent construction site
CN121280956A