Underwater target detection method based on improved YOLOv9

By improving YOLOv9's underwater target detection method, combining multi-scale branch convolution and attention mechanism, the problems of missed and missed objects in underwater target detection are solved, and the detection accuracy and robustness are improved.

CN120147609APending Publication Date: 2025-06-13YANSHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510204315.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

When detecting underwater targets, the existing technology faces the problems of missing and mis-detection of small objects, and the complex underwater environment increases the difficulty of detecting small objects.

Method used

The underwater object detection method based on improved YOLOv9 is adopted, and the loss function is optimized to improve detection accuracy by converting RepConvN in RepNCSPELAN4 into multi-scale branch convolution, combining position attention, spatial attention, channel attention and feature assistance modules.

Benefits of technology

Effectively extract the characteristic information of the target, suppress the impact of the underwater environment on detection performance, improve the detection accuracy, and demonstrate strong robustness and generalization ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147609A_ABST
    Figure CN120147609A_ABST
Patent Text Reader

Abstract

The invention discloses an underwater target detection method based on improved YOLOv9, and belongs to the technical field of underwater target detection, and the underwater target detection method comprises the following steps: S1, RepConvN in RepNCSPELAN4 is converted into Multi-scale Branching Convolution (MSBC), the original average pooling operation in the MSBC is changed into an Integrating pool (Intpool) module in which the maximum pooling and the average pooling are combined, and finally, an Officient-CSPELAN4 backbone structure is obtained; s2, providing a feature auxiliary module located between the neck and the head by combining position attention, space attention and channel attention; s3, adopting an inner-assisted shape-iou loss function, and embedding the Officient-CSPELAN4 backbone and the feature auxiliary module into the YOLOv9c model, so as to obtain an improved YOLOv9c model; s4, inputting the image of the UCPR2020 underwater data set into the improved YOLOv9c model for training, and verifying the performance of the model; and S5, detecting the public data set PASCAL VOC by using the trained improved YOLOv9c model, and verifying the generalization performance of the trained improved YOLOv9c model. According to the invention, the efficiency and reliability of underwater target detection are significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of underwater target detection, and particularly relates to an underwater target detection method based on improved YOLOv9. Background Art

[0002] The ocean is rich in resources. With the increasing demand for ocean exploration and development, the rapid development of computer vision technology provides important support for this field. Currently, underwater target detection technology has been widely applied in many fields, such as the positioning and tracking of underwater targets, the salvage of marine objects, the aquaculture and development of fisheries, etc. Underwater optical images have high resolution and rich information, so they are suitable for underwater target detection. However, the complexity of the underwater environment leads to problems such as color distortion, low contrast, and unclear feature information in underwater images. Therefore, how to process the visual information presented by the images and quickly and stably detect the targets is a key issue.

[0003] With the development of deep learning, many object detection methods based on neural networks have been proposed, and these methods have high-precision and strong generalization performance. However, detecting underwater targets still faces two major challenges. 1) The problem of missed detection and false detection of small objects is relatively prominent. Since underwater targets are small, blurred, and stacked with each other, higher requirements are imposed on the extraction of target feature information; in addition, considering that the downsampling operation will lose the position information of shallow targets, how to extract information more effectively is a major challenge for underwater target detection. 2) The complex underwater environment increases the detection difficulty of small objects. The attenuation coefficient of light in water is inconsistent, resulting in uneven underwater illumination; in addition, compared with the road surface environment, the underwater background has a higher similarity with the target and there are more suspended substances in the water. Summary of the Invention

[0004] In order to solve the problem that the existing technology has relatively prominent problems of missed detection and false detection of small objects when detecting underwater targets; and at the same time to solve the problem that the complex underwater environment in the existing technology increases the detection difficulty of small objects when detecting underwater targets. The present invention provides an underwater target detection method based on improved YOLOv9, which can effectively obtain the feature information of the targets and suppress the interference information underwater, highlighting the ability of feature expression and improving the accuracy of underwater target detection.

[0005] The technical solution adopted by the present invention for an underwater target detection method based on improved YOLOv9 is as follows:

[0006] An underwater target detection method based on improved YOLOv9, characterized in that it includes the following steps,

[0007] S1. Convert the RepConvN in RepNCSPELAN4 into a multi-scale branching convolution (MSBC), and change the original average pooling operation in the MSBC to an Integrated pool (Intpool) module that combines max pooling and average pooling, finally obtaining the Efficient-CSPELAN4 backbone structure;

[0008] S2. By combining position attention, spatial attention, and channel attention, propose a feature assistance module located between the neck and the head;

[0009] S3. Adopt an inner-assisted shape-iou loss function, and embed the Efficient-CSPELAN4 backbone and the feature assistance module into the YOLOv9c model to obtain an improved YOLOv9c model;

[0010] S4. Input the images of the UCPR2020 underwater dataset into the improved YOLOv9c model for training and verify the model performance;

[0011] S5. Use the trained improved YOLOv9c model to detect the public dataset PASCAL VOC and verify the generalization performance of the trained improved YOLOv9c model.

[0012] A further improvement of the technical solution of the present invention lies in: the specific steps of the Intpool module in the step S1 are as follows:

[0013] S1.1. Input the feature map X 2C×H×W , and then perform a Split operation on it to obtain and Subsequently, perform average pooling and max pooling on X 1 and X 2 respectively.

[0014] S1.2. Perform a Concat operation on the pooling results in the channel dimension to obtain the output feature map Y 2C×H×W

[0015]

[0016] where ⊕ is the concatenation in the channel dimension.

[0017] A further improvement of the technical solution of the present invention lies in that: the position attention in step S2 captures the correlation between any two positions without involving learning parameters, optimizes the weight of the target according to the correlation, and establishes the connection between spatial attention and channel attention; spatial attention focuses on the effective regions in the feature map in the spatial dimension and is a supplement to position attention; channel attention establishes the dependence relationship between channels through one-dimensional convolution, while ensuring light weight, retaining the feature information of the original channels.

[0018] A further improvement of the above technical solution of the present invention lies in that: step S2 includes the following steps:

[0019] S2.1. Input feature X ∈ C×H×W After passing through the Shape(S) module, it is transformed into feature a ∈ C×HW , the Trsnspose(T) module swaps the feature dimensions to obtain a dimension of HW×C, and the above two-dimensional modules are multiplied matrix by matrix to obtain feature b ∈ C×C , and then Softmax normalization processing is performed to obtain a weight matrix, where the Softmax operation formula is as follows,

[0020]

[0021] where, b i,j and a i,j respectively represent the values of the i-th row and j-th column of the relationship matrices b and a;

[0022] S2.2. After applying the weight on the feature map a and then performing the S operation, a dimension same as the input feature is obtained, and the intermediate feature c ∈ C×H×W is obtained through a residual connection, and the feature c is input into the channel attention module (ECA), and the feature map is compressed through average pooling in the channel dimension;

[0023] S2.3. Replace the common fully connected layer with a 1*k convolution. The size k of the one-dimensional convolution kernel is related to the number of adjacent channel information captured, and a channel attention weight value f ∈ C is generated through Sigmoid operation,

[0024] f = Sigmoid(conv1d(avgpool(c)))

[0025] Applying the weight value f on the input feature X to obtain feature d ∈ C×H×W ;

[0026] S2.4. Perform average pooling and max pooling operations on the input feature respectively, and splice the obtained results on the channel, and then obtain the feature weighted coefficient u ∈ H×W of the spatial information through convolution. The corresponding formula is as follows,

[0027] u = Sigmoid(f 2 / 1 [avgpool(x):maxpool(x)])

[0028] where [:] represents the concat operation of channel concatenation, f 2 / 1 represents a convolutional operation with an input size of 2 and an output size of 1;

[0029] Applying the spatial feature adjustment coefficient u to the feature c gives the feature e ∈ C×H×W , and finally adding the information with the feature d at the corresponding position to obtain the output feature P.

[0030] A further improvement of the technical solution of the present invention is that: the Inner-IoU in the step S3 is expressed as follows,

[0031]

[0032] union = (w gt *h gt )*(ratio) 2 +(w*h)*(ratio) 2 -inter

[0033]

[0034] where, represents the center coordinates of the actual box, (x c ,y c ) represents the center coordinates of the predicted box, the height and width of the actual box are represented by h gt and w gt respectively, the height and width of the predicted box are represented by h and w, ratio is a scale factor with a set range of [0.5, 1.5], Ratio is centered on 1, and the variation range is 0.5;

[0035] The formula of Shape-IoU is as follows,

[0036]

[0037] In the formula, ww and hh represent the weight coefficients on the width and height respectively, scale is a scaling factor related to the target size of the dataset, and the corresponding bounding box regression loss is as follows,

[0038] L Shape-Iou = 1 - Iou + distance shape + 0.5×Ω shape

[0039] Therefore, the Shape-iou loss function based on inner assistance is as follows:

[0040] Loss = 1 - Iou inner + distance shape + 0.5 × Ω shape 。

[0041] A further improvement of the technical solution of the present invention lies in: the method for verifying the model performance in step S4 includes: according to the output result of the improved YOLOv9c model, dividing the output result into true positives (TP), true negatives (TN), false positives (FP), and false negatives (FN), calculating the precision, average precision, and recall rate metrics based on the quantities of TP, TN, FP, and FN, and evaluating the performance of the improved YOLOv9c model according to the above metrics.

[0042] A further improvement of the technical solution of the present invention is that in step S4, the UCPR2020 underwater dataset is divided into four categories, namely sea cucumbers, starfish, sea urchins, and scallops. The epoch for training the improved YOLOv9c model is set to 200, the initial learning rate is 0.01, the input image size is 640×640, and the batch size is 4.

[0043] A further improvement of the technical solution of the present invention lies in: the method for verifying the generalization performance of the trained improved YOLOv9c model in step S5 includes:

[0044] Using the PASCAL VOC 2007 and PASCAL VOC 2012 training sets to train the model, testing the model with the PASCAL VOC2007 test set, and obtaining the detection results. At the same time, a comparative experiment is conducted with other detectors on the VOC dataset.

[0045] Due to the adoption of the above technical solution, the technical progress achieved by the present invention includes:

[0046] The present invention improves the detection performance of the model for small targets by introducing a multi-scale branch convolution module MSBC, using convolution kernels of different sizes for feature extraction, expanding the perception field, and strengthening the description of small target features.

[0047] The present invention optimizes the intpool structure, enables the network to obtain texture and edge information, strengthens the distinction between the background and the foreground, highlights the detailed information, and solves the limitation of single pooling.

[0048] The present invention enhances the feature expression ability of small targets by designing a feature assistance module, selectively enhancing multi-scale aggregated features, and suppressing the weight of background information.

[0049] The improved loss function of the present invention reduces the influence of sample size and sample shape on bounding box regression and speeds up the convergence rate of the model.

[0050] Through the above design, the present invention effectively extracts the feature information of the target and suppresses the influence of the underwater environment on the detection performance, thereby improving the detection accuracy. The detection results on the VOC dataset show that the model also exhibits strong robustness and generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 is a schematic diagram of the overall process of an underwater target detection method based on improved YOLOv9 of the present invention;

[0052] Figure 2 is a framework diagram of the improved YOLOv9c model of an underwater target detection method based on improved YOLOv9 of the present invention;

[0053] Figure 3 is a network structure diagram of the feature assistance module of an underwater target detection method based on improved YOLOv9 of the present invention;

[0054] Figure 4 is a P-R curve diagram of the original YOLOv9c model of an underwater target detection method based on improved YOLOv9 of the present invention;

[0055] Figure 5 is a P-R curve diagram of the improved YOLOv9c model of an underwater target detection method based on improved YOLOv9 of the present invention;

[0056] Figure 6 is a detection effect diagram of an underwater target detection method based on improved YOLOv9 of the present invention in the UCPR2020 underwater dataset. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the specific embodiments and with reference to the accompanying drawings. In the following description, the description of well-known structures and technologies is omitted to avoid unnecessarily confusing the concepts of the present invention.

[0058] The present invention provides an underwater target detection method based on improved YOLOv9, referring to Figures 1-6 as shown, including the following steps:

[0059] S1. Convert RepConvN in RepNCSPELAN4 into a multi-scale branching convolution (MSBC), and change the original average pooling operation in MSBC to an Integrated pool (Intpool) module that combines max pooling and average pooling. Finally, obtain the Efficient-CSPELAN4 backbone structure.

[0060] The specific steps for constructing the Intpool module are as follows:

[0061] S1.1. Input the feature map X 2C×H×W , and then perform a Split operation on it to obtain and Subsequently, perform average pooling and max pooling on X 1 and X 2 respectively.

[0062] S1.2. Perform a Concat operation on the pooling results in the channel dimension to obtain the output feature map Y 2C×H×W

[0063]

[0064] where is the concatenation in the channel dimension.

[0065] On the one hand, due to the small size and mutual stacking of underwater targets, the texture of the collected images is unclear. On the other hand, as the network depth increases, in the downsampling convolution operation, it is easy to cause the loss of feature information, which in turn increases the proportion of useless information and affects the performance of the subsequent aggregation network and the accuracy of the detection output. In order to obtain more effective feature information, the present invention proposes an Efficient-CSPELAN4 (ECSPE) structure to replace the original module RepNCSPELAN4. Specifically, a multi-scale branching convolution (MSBC) with four branches based on the Inception network is proposed. During the convolution operation, each branch uses convolution kernels of different sizes for feature extraction. Through multi-scale feature fusion, features of different sizes can be obtained, the perception field can be extended, and the description of small target features can be strengthened, thereby improving the detection performance of the model for small targets. The advantage of the multi-scale branching convolution is that it can extract image features more adequately without increasing the inference time cost, and improves the situation of information loss caused by the increase of the convolutional depth network.

[0066] For the original backbone network, the expressive ability of backbone features can be improved by increasing the number of RepConvNs, but this will increase the complexity of the network and thus affect the inference speed. The multi-scale branched convolution has different types of branches and can adaptively select different types of convolutions for feature extraction, enhancing the feature extraction ability of the model. Specifically, the input feature map X passes through a two-dimensional convolution with a kernel of 1, and the channels of the input features are linearly superimposed at each projection position, adjusting the number of channels while retaining the original image plane structure. For Conv3X3, a large receptive field can cover more pixels and better capture context information. Additionally, during the inference stage, the multi-scale branched convolution is reparameterized so that the four branches are equivalent to a convolution block, without increasing the inference time while retaining the trained parameters, only slightly increasing the network parameters during the training stage.

[0067] Detail information is crucial for detecting small targets. To address the problem that max pooling only retains the maximum value information within the effective region, the Integrated pool (Intpool) module is proposed in the MSBC structure. The Intpool structure combines average pooling and max pooling, solving the limitations of single pooling. It not only highlights the difference between the object and the background at the edge but also increases the attention weight of the position information, retaining more detail information.

[0068] S2. By combining position attention, spatial attention, channel attention, and shortcut connections, a feature assistance module located between the neck and the head is proposed.

[0069] After passing through the pyramid aggregation network, the feature layer fuses multi-scale feature information and is finally fed into the detection head for object classification and regression of the object box. For object detection, the detection head is the part that finally produces the results and is directly related to the quality and convergence speed of the detection model. During the model training process, due to the weak feature representation of small targets, the small proportion of pixels, and the complexity of the underwater environment, it is impossible to effectively obtain the optimization weights of small targets, affecting the detection accuracy. Therefore, before detection, it is particularly important to enhance the distinguishability between the foreground and the background of the multi-scale features and enhance the feature expression ability of the foreground. Additionally, there are semantic differences in the feature maps of different scales, and the fused feature layer may produce a confusion effect, leading to network confusion in localization and recognition. To mitigate this negative impact, the present invention designs a feature assistance module. The feature assistance module effectively aggregates three-dimensional features by establishing the association between spatial features and channel features. At the same time, the feature assistance module can also capture long-range features, which is beneficial to selectively enhancing the multi-scale aggregated features and suppressing the weight of background information, thereby enhancing the feature expression ability of small targets.

[0070] Inspired by attention, the feature assistance module combines positional attention, spatial attention, channel attention, and shortcut connections. Without involving learning parameters, positional attention captures the correlation between any two positions, optimizes the weights of the targets according to the correlation, and establishes the connection between spatial attention and channel attention. Spatial attention focuses on the effective regions in the feature map in the spatial dimension and is a supplement to positional attention. Channel attention establishes the dependence between channels through one-dimensional convolution, while retaining the feature information of the original channels while ensuring light weight. The feature assistance module is a convenient and flexible structure that does not require setting any parameters.

[0071] The structure of the feature assistance module is shown in Figure 3 as follows. The input feature X ∈ C×H×W is transformed into feature a ∈ C×HW after passing through the Shape(S) module. The Trsnspose(T) module transposes the feature dimensions to obtain a dimension of HW×C. The matrices of the above two dimensions are multiplied to obtain feature b ∈ C×C . Then, Softmax normalization is performed to obtain the weight matrix, and the Softmax operation formula is as follows.

[0072]

[0073] where b i,j and a i,j represent the values of the i-th row and j-th column of the relationship matrices b and a, respectively. After applying the weights on the feature map a and then performing the S operation, the same dimension as the input feature is obtained. Through the residual connection, the intermediate feature c ∈ C×H×W is obtained. The feature c is input into the channel attention module (ECA), and the feature map is compressed through average pooling in the channel dimension.

[0074] To reduce the parameters of the model and make the model more lightweight, the commonly used fully connected layer is cancelled and replaced with a 1*k convolution. The size k of the one-dimensional convolution kernel is related to the number of adjacent channel information captured. Finally, a channel attention weight f ∈ C is generated through the Sigmoid operation. The expression formula is as follows.

[0075] f = Sigmoid(conv1d(avgpool(c)))

[0076] Applying the weight f on the input feature X obtains feature d ∈ C×H×WIn addition, in order to further capture the effective features inside the feature layer and cooperate with the position attention to fuse the information in the spatial dimension, the present invention adopts spatial attention. Spatial attention obtains the feature adjustment coefficient in the spatial dimension by compressing the channel size. Specifically, it performs average pooling and max pooling operations on the input features respectively. And the obtained results are concatenated on the channel, and then convolved to obtain the feature weighting coefficient u of the spatial information ∈ H×W The corresponding formula is as follows.

[0077] u = Sigmoid(f 2 / 1 [avgpool(x):maxpool(x)])

[0078] In the formula, [:] represents the concat operation of channel concatenation, and f 2 / 1 represents a convolution operation with an input size of 2 and an output size of 1. Applying the spatial feature adjustment coefficient u to the feature c to obtain the feature e ∈ C×H×W , and finally adding the information to the feature d at the corresponding position to obtain the output feature P.

[0079] S3. Adopt the inner-assisted shape-iou loss function, and embed the Efficient-CSPELAN4 backbone and the feature assistance module into the YOLOv9c model to obtain the improved YOLOv9c model.

[0080] The positioning accuracy of object detection depends to a large extent on the loss function of bounding box regression. Therefore, bounding box regression plays a crucial role in the field of object detection. The CIOU loss function used by YOLOv9c takes into account the shape information of the target box. By introducing a correction factor, the loss is more robust to target boxes of different shapes, and also takes into account the diagonal distance of the target box, which helps to improve the accuracy of the object detection model when positioning the target, so it is more sensitive to the position prediction of the target box. However, CIOU ignores the influence of the inherent attributes such as the shape and scale of its own bounding box on regression. In addition, for larger target objects, in the case of the same position offset, the IOU change of smaller target objects is relatively large, which also fully shows that the IOU-based loss function is very sensitive to the position movement of small objects. This reduces the overall detection accuracy of the object detector. To solve the above deficiencies, the present invention adopts the inner-assisted shape-iou loss function. Because there is only a scale difference between the auxiliary bounding box and the actual bounding box, and the change trend of the IoU value of the auxiliary bounding box is consistent with the change trend of the actual bounding box during the regression process, so in the existing IOU loss, adding the auxiliary bounding box can not only reflect the quality of the actual bounding box regression result, but also accelerate the convergence speed of the model. The specific representation of Inner-IoU is as follows:

[0081]

[0082] union = (w gt * h gt ) * (ratio) 2 + (w * h) * (ratio) 2 - inter

[0083]

[0084] where represents the center coordinates of the actual box, (x c , y c ) represents the center coordinates of the predicted box. The height and width of the actual box are represented by h gt and w gt respectively. The height and width of the predicted box are represented by h and w. Generally, ratio is a scale factor with a setting range of [0.5, 1.5]. Ratio is centered on 1 and the variation range is 0.5. When ratio is less than 1, the size of the auxiliary bounding box is smaller than that of the actual bounding box, and its effective regression range is smaller than the IoU loss, which has a gain for the regression of high IoU cases. When ratio is greater than 1, the size of the auxiliary bounding box is larger than that of the actual bounding box, expanding the effective regression range, which has a gain for the regression of low IoU cases. Shape - IoU takes into account the influence of the differences in the shapes and scales of the self - bounding boxes in the regression samples on the IoU value. Specifically, for the bounding box regression samples with the same shape and the same deviation, compared with the regression samples of larger sizes, the IoU value of the regression samples of smaller - sized bounding boxes is more significantly affected by the shape of the actual box. For the bounding box regression samples with the same size and the same deviation, the deviation in the short - side direction of the bounding box has a more significant impact on the IoU than the deviation in the long - side direction of the bounding box. The formula of Shape - IoU is as follows:

[0085]

[0086] where ww and hh represent the weight coefficients in width and height respectively, and scale is a scaling factor related to the target size of the dataset. The corresponding bounding box regression loss is as follows:

[0087] L Shape-Iou = 1 - Iou + distance shape + 0.5×Ω shape

[0088] Therefore, the Shape - iou loss function based on inner assistance is:

[0089] Loss = 1 - Iou inner + distance shape + 0.5×Ω shape .

[0090] S4. Input the images of the UCPR2020 underwater dataset into the improved YOLOv9c model for training and verify the model performance.

[0091] Specifically, the UCPR2020 underwater dataset is divided into four categories, namely sea cucumbers, starfish, sea urchins, and scallops. The dataset contains a total of 5,455 images, which were taken at the marine ranch in Zhangzidao, Dalian, China. There are 4,418 images in the training set, 491 images in the validation set, and 546 images in the test set. The epoch for training the model is set to 200, and the initial learning rate is 0.01. The input image size is 640×640, and the batch size is 4.

[0092] Meanwhile, the above method for verifying the model performance includes: according to the output results of the improved YOLOv9c model, the output results are divided into true positives (TP), true negatives (TN), false positives (FP), and false negatives (FN). Calculate the precision, average precision, and recall rate metrics based on the quantities of TP, TN, FP, and FN, and evaluate the performance of the improved YOLOv9c model according to the above metrics.

[0093] 1: Ablation experiments on the UCPR2020 dataset

[0094] Under the condition of the same experimental conditions and parameter settings, verify the roles of position attention (PA), spatial attention (SA), and channel attention (CA) in the feature assistance module respectively. Similarly, verify the performance of ECSPE and regression loss in the network. The results are shown in Table 1.

[0095] Table 1 Ablation experiments

[0096]

[0097] In the above experiment, the yolov9c model was used as the baseline, and its mAp value on the underwater dataset was 85.62%. When the feature assistance module only included position attention features, the mAp increased by 0.28% compared to the baseline model. When the feature assistance module only included spatial attention features, the mAp increased by 0.45% compared to the baseline. When the gating unit only fused channel attention features, the mAp increased by 0.53% compared to the baseline. When the feature assistance module included both position attention and spatial attention, the mAp increased by 0.98% compared to the baseline. When the feature assistance module included both position attention and channel attention, the mAp increased by 0.79% compared to the baseline. When the feature assistance module included both spatial attention and channel attention, the mAp increased by 1.06% compared to the baseline. When the feature assistance module included all three attention features, the mAp increased by 1.12% compared to the baseline model. On the basis of including the feature assistance module, the ECSPE structure was continued to be added, and the mAp increased by 1.26% compared to the baseline model. Finally, the mAp of the network proposed in the present invention reached 87.24%, an increase of 1.62% compared to the baseline model. Through a series of ablation experiments, it can be clearly seen the positive contributions of each improvement point to the overall detection effect, thus proving the effectiveness of these improvement points.

[0098] 2: Comparative experiments on the UCPR2020 dataset

[0099] Then, the improved YOLOv9c model was trained using the UCPR2020 underwater dataset and compared with other detectors. These included one-stage detectors: SSD, YOLO5, YOLOX, YOLOv7, YOLOv8, RetinaNet and two-stage detectors: Faster R-CNN, Cascade R-CNN and other detectors with outstanding performance: DETR, ATSS, FCOS. The performance metrics included map0.5 and detection speed FPS. The results showed that the detection accuracy of SSD512 was not high, its mAP achieved 67.33%, but the detection speed could meet the real-time requirements, with an FPS of 16.4. The mAP values of Cascade R-CNN, RetinaNet, and DETR all exceeded 77.5%. The mAP of YOLOv5, YOLOX, YOLOv7, and YOLOv8 exceeded 81%, among which YOLOv5 had the highest FPS, reaching 38.4. The mAP of the model proposed in the present invention reached 87.24%, significantly superior to other comparison algorithms. Compared with YOLOv9, the amp of the model in the present invention increased by 1.62%, and the FPS only decreased by 0.1. The detailed results are shown in Table 2 Table 2 Comparative experiments

[0100]

[0101]

[0102] S5. Use the trained improved YOLOv9c model to detect the public dataset PASCAL VOC, and verify the generalization performance of the trained improved YOLOv9c model.

[0103] Specifically, train with the PASCAL VOC 2007 and PASCAL VOC 2012 datasets, and test on the PASCAL VOC2007 test set to verify the generalization performance of the model of the present invention. The comparison models are divided into single-stage detectors and two-stage detectors. Among them, SA-FPN in the two-stage detectors has the highest accuracy, with an mAP of 79.1%, and R-FCN has the fastest detection speed, with an FPS of 11. Among the single-stage detectors, SSD300 with VGGNet as the backbone has the fastest detection speed, with an FPS of 46. The mAP of YOLOv9 reaches 86.3%, and the FPS is 7.64. The model of the present invention sacrifices some detection speed compared with YOLOv9c but achieves higher detection accuracy, with an mAP of 87.1%, which is the highest accuracy among all detectors. As shown in Table 3

[0104] Table 3 Comparative experiments on the VOC dataset

[0105]

[0106] In the above embodiments, the present invention provides an underwater target detection method based on improved YOLOv9. The present invention improves the detection performance of the model for small targets by introducing a multi-scale branch convolution module MSBC, using convolution kernels of different sizes for feature extraction, expanding the perception field, and strengthening the description of small target features. By optimizing the intpool structure, the network obtains texture and edge information, strengthens the distinction between the background and the foreground, highlights the detail information, and solves the limitation of single pooling. By designing a feature assistance module, selectively enhancing the multi-scale aggregation features, and suppressing the weight of the background information, the feature expression ability of small targets is enhanced. The improved loss function reduces the influence of the sample size and sample shape on the bounding box regression, and speeds up the convergence speed of the model. Through the above designs, the feature information of the target is effectively extracted, and the influence of the underwater environment on the detection performance is suppressed, thereby improving the detection accuracy. The detection results on the VOC dataset show that the model also exhibits strong robustness and generalization ability.

[0107] The above-described embodiments are only descriptions of the preferred embodiments of the present invention, and do not limit the concept and scope of the present invention. Without departing from the design concept of the present invention, various modifications and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope of the present invention. The technical content claimed by the present invention has been fully recorded in the claims.

Claims

1. An underwater target detection method based on improved YOLOv9, characterized in that: The following steps are included: S1. Convert RepConvN in RepNCSPELAN4 into Multi-scale branching convolution (MSBC), and change the original average pooling operation in MSBC into an Integrated pool (Intpool) module that combines maximum pooling and average pooling, and finally obtain the backbone structure of Efficient-CSPELAN4. S2, by combining position attention, spatial attention, and channel attention, a feature auxiliary module located between the neck and the head is proposed; S3, using the inner-assisted shape-iou loss function, and embedding the Efficient-CSPELAN4 backbone and feature-assisted module into the YOLOv9c model to obtain an improved YOLOv9c model; S4. Input the images of the UCPR2020 underwater dataset into the improved YOLOv9c model for training and verify the model performance; S5. Use the trained improved YOLOv9c model to test the public dataset PASCAL VOC to verify the generalization performance of the trained improved YOLOv9c model.

2. The underwater target detection method based on improved YOLOv9 according to claim 1, characterized in that: The specific steps of the Intpool module in step S1 are: S1.

1. Input feature map X 2C×H×W , and then perform a Split operation on it to get and Then, average pooling and maximum pooling are performed on X1 and X2 respectively; S1.

2. Perform a Concat operation on the pooling result in the channel dimension to obtain the output feature map Y 2C×H×W in It is the splicing in the channel dimension.

3. The underwater target detection method based on improved YOLOv9 according to claim 1, characterized in that: The position attention in step S2 captures the correlation between any two positions without involving learning parameters, optimizes the weight of the target based on the correlation, and establishes the connection between spatial attention and channel attention; spatial attention focuses on the effective area in the feature map in the spatial dimension, which is a supplement to the position attention; channel attention establishes the dependency between channels through one-dimensional convolution, retaining the feature information of the original channel while ensuring lightweight.

4. The underwater target detection method based on improved YOLOv9 according to claim 2, characterized in that: The step S2 comprises the following steps: S2.1、Input feature X∈ C×H×W After passing through the Shape(S) module, it is transformed into feature a∈ C×HW , the Trsnspose(T) module swaps the feature dimensions to obtain the dimension of HW×C, and the modules of the above two dimensions perform matrix multiplication to obtain the feature b∈ C×C , and then perform Softmax normalization to obtain the weight matrix, where the Softmax calculation formula is as follows: Among them, b i,j and a i,j Represent the values ​​of the i-th row and j-th column of the relationship matrix b and a respectively; S2.2, after applying the weights on the feature map a, perform the S operation to obtain the same dimension as the input feature, and obtain the intermediate feature c∈ through the residual link C×H×W , input feature c into the channel attention module (ECA) and compress the feature map by average pooling in the channel dimension; S2.

3. Replace the commonly used fully connected layer with a 1*k convolution. The size of the one-dimensional convolution kernel k is related to the amount of information captured from adjacent channels. A channel attention weight f∈ is generated through Sigmoid operation. C , f=Sigmoid(conv1d(avgpool(c))) Apply weight f to input feature X to get feature d∈ C×H×W ; S2.4, perform average pooling and maximum pooling operations on the input features respectively, and concatenate the results on the channel, and then obtain the feature weighting coefficient u∈ of the spatial information through convolution. H×W , the corresponding formula is as follows, u=Sigmoid(f 2 / 1 [avgpool(x):maxpool(x)]) Where [:] represents the concat operation of channel concatenation, f 2 / 1 Represents a convolution operation with an input size of 2 and an output size of 1; Apply the spatial feature adjustment coefficient u to feature c to obtain feature e∈ C×H×W , and finally add the information to feature d at the corresponding position to get the output feature P.

5. The underwater target detection method based on improved YOLOv9 according to claim 1, characterized in that: In step S3, Inner-IoU is expressed as follows: union=(w gt *h gt )*(ratio) 2 +(w*h)*(ratio) 2 -inter in, Indicates the center coordinates of the actual box, (x c ,y c ) represents the center coordinate of the predicted box, and the height and width of the actual box are represented by h gt and w gt Indicates that the height and width of the prediction box are represented by h and w, and ratio is a scale factor with a setting range of [0.5,1.5]. Ratio is centered at 1 and varies by 0.

5. The formula of Shape-IoU is as follows, Where ww and hh represent the weight coefficients on width and height respectively, scale is the scaling factor, which is related to the target size of the dataset. The corresponding bounding box regression loss is as follows, L Shape-Iou =1-Iou+distance shape +0.5×Ω shape Therefore, the inner-assisted Shape-iou loss function is: Loss=1-Iou inner +distance shape +0.5×Ω shape 。 6. The underwater target detection method based on improved YOLOv9 according to claim 1, characterized in that: The method for verifying the model performance in step S4 includes: according to the output results of the improved YOLOv9c model, the output results are divided into true positive examples (TP), true negative examples (TN), false positive examples (FP) and false negative examples (FN), and the precision, average precision and recall rate indicators are calculated according to the number of TP, TN, FP and FN, and the performance of the improved YOLOv9c model is evaluated according to the above indicators.

7. The underwater target detection method based on improved YOLOv9 according to claim 1, characterized in that: In step S4, the UCPR2020 underwater dataset is divided into four categories, namely, sea cucumbers, starfish, sea urchins and scallops. The epoch for training the improved YOLOv9c model is set to 200, the initial size of the learning rate is 0.01, the input image size is 640×640, and the batch size is 4.

8. The underwater target detection method based on improved YOLOv9 according to claim 1, characterized in that: The method for verifying the generalization performance of the trained improved YOLOv9c model in step S5 includes: The PASCAL VOC 2007 and PASCAL VOC 2012 training sets are used to train the model, and the PASCAL VOC 2007 test set is used to test the model and obtain the detection results. At the same time, comparative experiments are performed with other detectors on the VOC dataset.