Underwater target detection method and system based on improved Yolov11
By improving the YOLOv11 network, designing new modules and structures, and using NWD activation function, the problems of poor identification accuracy and high computational complexity in underwater target detection are solved, and efficient and robust underwater target detection are achieved.
Patent Information
- Application Number
- CN202510092977.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
AI Technical Summary
In the prior art, underwater target detection has problems such as poor identification accuracy, small underwater target miss detection, and large number of model parameters and difficult to deploy.
By improving the YOLOv11 network, the C3k2_Faster_EMA module, deep convolution feature pyramid structure and RSCD detection head are designed, and the NWD activation function is used to reduce the computational complexity and improve the detection accuracy.
It realizes the reduction of computational complexity while maintaining or improving detection accuracy, and solves the problem that traditional methods have high computational complexity and the inability of a single structural feature extractor to capture differences between different objects.
Smart Images

Figure CN120014235A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of computer vision technology, relates to underwater target detection technology, and specifically relates to an underwater target detection method and system based on improved Yolov11. Background Art
[0002] With the rapid development of social economy, resource consumption has also increased dramatically, forcing humans to seek better ways to obtain resources. The ocean occupies 71% of the earth's area. It has many precious resources. If you want to obtain marine resources, the first thing to do is to detect various types of information about marine resources. Therefore, underwater target detection tasks come into being. Underwater target detection is of great significance to marine resource exploration, geological surveys, etc. It is also an effective means to study and protect marine life and resources.
[0003] In recent years, deep learning technology has developed rapidly and can automatically extract features through end-to-end learning, which greatly improves the effect of object detection and overcomes the limitations of traditional methods. Object detection technology based on deep learning can be divided into two categories: two-stage detection and one-stage detection. Two-stage detection methods, such as R-CNN, Fast-R-CNN, and Faster-R-CNN, first generate candidate boxes and then classify these boxes. Although this type of method has high accuracy, it has a large computational overhead. In contrast, one-stage detection methods (such as EfficientDet, SSD, YOLO, etc.) have faster processing speeds and higher detection efficiency while ensuring high accuracy. In particular, the YOLO series has been widely used in practical applications. In the article "Improved YOLOv7 Underwater Target Detection Algorithm" published in "Computer Engineering and Applications, 2024, 60(06): 89-99", Liang Xiuman et al. proposed an improved algorithm model based on YOLOv7 to address problems such as low visibility and color distortion underwater. This model incorporates a multi-information stream fusion attention mechanism to improve the detection accuracy when the image is blurred. In "Electronics", an underwater target detection model based on Yolov7 was proposed, and an enhancement branch was constructed to expand the Yolov7 feature extraction, thereby improving the detection accuracy. However, these methods use stacking network layers to improve network accuracy, and the introduction of enhanced feature extraction operation units does not achieve lightweight. While improving model detection accuracy, the computational complexity is also very high. In addition, the scale of underwater organisms varies greatly, and traditional single-structure feature extractors are difficult to capture the differences between different types of objects. There is still room for improvement in detection accuracy, and it is very challenging to apply them to underwater detection equipment. Therefore, it is necessary to discover excellent models with smaller computational complexity for application on underwater target detection platforms. Summary of the invention
[0004] Purpose of the invention: In order to overcome the shortcomings of the prior art, such as poor recognition accuracy, missed detection of small underwater targets, and large number of model parameters that are difficult to deploy, an underwater target detection method and system based on improved Yolov11 are provided.
[0005] Technical solution: To achieve the above purpose, the present invention provides an underwater target detection method based on improved Yolov11, comprising the following steps:
[0006] S1: Divide the collected data set into training set and test set in proportion;
[0007] S2: Improve the Yolov11 network;
[0008] S3: Train and test the improved Yolov11 network through the training set and test set respectively, and use the trained Yolov11 network as the underwater target detection model;
[0009] S4: Obtain underwater target detection results through the underwater target detection model.
[0010] Furthermore, the improvements to the Yolov11 network in step S2 include: designing a C3k2_Faster_EMA module and replacing the Yolov11 backbone C3K2 module; proposing a deep convolutional feature pyramid structure and replacing the SPPF module in Yolov11; designing an RSCD detection head and replacing the original Yolov11 detection head; and using an NWD activation function to replace the activation function in Yolov11.
[0011] Furthermore, the specific implementation steps of designing the C3k2_Faster_EMA module in step S2 and replacing the Yolov11 backbone C3K2 module include:
[0012] A1: Send the input feature map X to Pconv (partial convolution), the specific formula is:
[0013]
[0014] Among them, Y i is the feature value in the output image, which is 512×512 in the present invention, i is the size of the output feature map, X k is the feature value in the input feature map X, which is 512×512 in the present invention, W k is the weight in the convolution kernel, the weight range is 0 to 1, M k It is the value in the mask image, and its value can only be 0 and 1. 1 means the feature is known, and 0 means the pixel feature is unknown. The feature corresponding to the position of 0 in the mask image does not participate in the calculation;
[0015] A2: Input the result Y after Pconv (partial convolution) into the MLP layer, where Y is Y i It is obtained by splicing along the horizontal axis. The specific formula is:
[0016] h mlp =Conv1×1(Y,dim)
[0017] X mlp =Conv1×1(h mlp ,dim)
[0018] A3: The result X after the MLP layer mlp Input into the DropPath layer, the specific formula is:
[0019] x drop =DropPath(x mlp )
[0020] A4: Pass the result x through the DropPath layer drop The final result is obtained by sending it into the exponential moving average model (EMA), and its specific formula is:
[0021] x att =EMA(x drop ).
[0022] Furthermore, the specific implementation steps of proposing a deep convolutional feature pyramid structure and replacing the SPPF module in Yolov11 in step S2 include:
[0023] B1: According to the bottom to the top of the deep convolution feature pyramid structure, 3×3 convolution operations are used to generate feature maps of different dimensions P0, P1, P2, P3, P4, P5, P6; the formula is:
[0024] P l =Conv(P l-1 ),l∈{1,2,3,4,5,6}
[0025] Among them, P l is the feature map of this layer, P l-1 It is the output of the previous layer; the P6 feature map is the smallest and the P0 feature map is the largest. As the number of layers increases, the size of the feature map gradually decreases;
[0026] B2: Apply a shared convolution operation Conv at each layer share , and concatenate them according to the channel dimension as follows:
[0027] P l ′=Contact(Conv share (P l)), l∈{1,2,3,4,5,6}
[0028] B3: The result P of splicing the shared convolution kernel l 'After a downsampling module, it is added pixel by pixel with P5 and then goes through a 3×3 convolution to obtain the output result
[0029]
[0030] Furthermore, the specific implementation steps of designing the RSCD detection head and replacing the original Yolov11 detection head in step S2 include:
[0031] C1: Input three feature maps x j , processed by a separate 1×1 convolution and normalization layer, and then further processed by a 3×3 shared convolution. The specific formula is as follows:
[0032]
[0033] Among them, Conv_GN is a convolutional layer with group normalization; the feature map x i After 1×1 convolution and group normalization operations, a new feature map is obtained.
[0034] C2: All processed feature maps are fused through a shared convolutional layer:
[0035]
[0036] Among them, Share_Conv is a shared convolution module;
[0037] C3: Feature map after shared convolution The convolution operation is sent to the separation regression branch to obtain Conv_Reg and the classification branch Conv_Cls. Then, each uses a 1×1 convolution with independent weights to complete the first-stage prediction task, and uses the Scale layer to scale the output Conv_Reg of the separation regression branch, dynamically adjust the importance of the feature map, and help locate underwater targets of different sizes.
[0038] Furthermore, the expression of the NWD activation function in step S2 is:
[0039]
[0040] in, is the center coordinate of the prediction box, (w a ,h a ) is the prediction box width and height, is the coordinate of the center position of the real frame, (w b ,h b ) are the actual frame width and height, C is a user-defined constant whose value range is the positive real number domain.
[0041] The present invention also provides an underwater target detection system based on the improved Yolov11, including a peripheral interface, a memory and a processor;
[0042] Peripheral interface, used to realize data input and output in the process of data interaction with various external devices, so as to ensure the effective connection and communication between the system and external devices;
[0043] a memory for storing computer program instructions capable of being executed on the processor;
[0044] The processor is used to execute the steps of an underwater target detection method based on improved Yolov11 when running computer program instructions.
[0045] The present invention provides a lightweight enhanced feature extraction unit, which aims to reduce the computational complexity and solve the problem of low efficiency caused by high computational complexity in the prior art while maintaining or improving the detection accuracy.
[0046] The present invention provides a multi-structure feature extraction method, which can effectively cope with the differences in scale changes of underwater organisms, enhance detection accuracy through a multi-level feature extraction method, and solve the problem that the existing single structure feature extractor cannot handle the differences between different objects.
[0047] Beneficial effects: Compared with the prior art, the present invention has the following advantages:
[0048] 1. The present invention designs a C3k2_Faster_EMA module, and uses the C3k2_Faster_EMA module to replace the C3K2 module. This design makes the weight update in the training process smoother, reduces the training instability caused by excessive learning rate or gradient explosion, and avoids oscillation and gradient disappearance.
[0049] 2. The present invention designs an RSCD detection head and uses the RSCD detection head to replace the YOLOv11 detection head. This design improves the ability of the detection head to capture details and reduces the number of parameters and the amount of calculation of the algorithm.
[0050] 3. The present invention designs a deep convolutional feature pyramid structure, which modifies the SPPF structure in Yolov11 into a deep convolutional feature pyramid structure. This design can not only capture multi-scale features, but also avoid excessive parameter redundancy by sharing convolutional layers, making the network more capable of extracting diverse and recognizable features from complex images. Compared with SPPF, this improved structure can perform better when dealing with small objects, complex backgrounds and other problems.
[0051] 4. In order to improve the accuracy and robustness of network training, the present invention adopts the NWD loss function and optimizes the measurement method for small targets, so that the network's detection performance is further improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 It is a model framework diagram of the method of the present invention;
[0053] Figure 2 It is a schematic diagram of the structure of C3k2_Faster_EMA in the present invention;
[0054] Figure 3 Schematic diagram of the structure of the deep convolution feature pyramid in the present invention;
[0055] Figure 4 Schematic diagram of the RSCD detection head in the present invention;
[0056] Figure 5 It is a curve diagram of the number of iterations-accuracy of training the model proposed by the present invention;
[0057] Figure 6 It is a curve diagram of the number of iterations-loss value for training the model proposed in the present invention;
[0058] Figure 7 It is a set of visualization results predicted by using the optimal weight parameters. DETAILED DESCRIPTION
[0059] The present invention is further explained below in conjunction with the accompanying drawings and specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, various equivalent forms of modifications to the present invention by those skilled in the art all fall within the scope defined by the claims attached to this application.
[0060] like Figure 1 As shown, this embodiment provides an underwater target detection method based on improved Yolov11, comprising the following steps:
[0061] S1: Obtain a marine resource dataset, taking the URPC dataset of the Underwater Robot Perception Challenge as an example, which contains 5543 underwater optical images. The target organisms in the dataset cover four categories: sea cucumbers, sea urchins, scallops, and starfish. The dataset is randomly divided into a training set and a test set with a division ratio of 8:2;
[0062] S2: Improve the Yolov11 network;
[0063] Reference Figure 1 , the improvements to the Yolov11 network include: designing the C3k2_Faster_EMA module and replacing the Yolov11 backbone C3K2 module. By introducing the self-developed Faster_EMA structure, the model size is greatly reduced, and the network's ability to capture detailed features is enhanced; a deep convolutional feature pyramid structure is proposed to replace the SPPF module in Yolov11, and the detection ability of targets of different scales is improved through multi-scale feature fusion; the RSCD detection head is designed and replaced with the original Yolov11 detection head, and the number of parameters and computational complexity are significantly reduced through shared convolution and scale-aware mechanisms; the NWD activation function is used to replace the activation function in Yolov11. The NWD activation function can dynamically adjust the activation response of different features, enhance the model's learning ability for small targets, and improve the network's detection accuracy. Through these improvements, the YOLOv11 network has been significantly improved in detection performance, computational efficiency, and feature expression capabilities, forming an efficient and robust target detection framework.
[0064] In this embodiment, after completing the preprocessing of the data set, a yaml configuration file is constructed, the paths of the training set and the test set are written into the data.yaml file, and the configuration file information is adjusted. In the official yaml file, the C3k2_Faster_EMA module is used to replace the C3K2 module, the deep convolution feature pyramid module is used to replace the SPPF module in Backbone, and the RSCD detection head is used to replace the original detection head. The number of training rounds is adjusted to 300 rounds, the batch-size is adjusted to 4, and the category of the detection object is adjusted to 10.
[0065] like Figure 2 As shown in the figure, the specific implementation steps of designing the C3k2_Faster_EMA module and replacing the Yolov11 trunk C3K2 module include:
[0066] A1: Send the input feature map X to Pconv (partial convolution), the specific formula is:
[0067]
[0068] Among them, Y iis the feature value in the output image. In this embodiment, the value is 512×512, i is the size of the output feature map, and X k is the feature value in the input feature map X, which is 512×512 in this embodiment, and W k is the weight in the convolution kernel, the weight range is 0 to 1, M k It is the value in the mask image, and its value can only be 0 and 1. 1 means the feature is known, and 0 means the pixel feature is unknown. The feature corresponding to the position of 0 in the mask image does not participate in the calculation;
[0069] A2: Input the result Y after Pconv (partial convolution) into the MLP layer, where Y is Y i It is obtained by splicing along the horizontal axis. The specific formula is:
[0070] h mlp =Conv 1×1 (Y, dim)
[0071] X mlp =Conv 1×1 (h mlp ,dim)
[0072] A3: The result X after the MLP layer mlp Input into the DropPath layer, the specific formula is:
[0073] x drop =DropPath(x mlp )
[0074] A4: Pass the result X through the DropPath layer drop The final result is obtained by sending it into the exponential moving average model (EMA), and its specific formula is:
[0075] X att =EMA(X drop )
[0076] like Figure 3 As shown in the figure, the specific implementation steps of proposing a deep convolutional feature pyramid structure and replacing the SPPF module in Yolov11 include:
[0077] B1: According to the bottom to the top of the deep convolution feature pyramid structure, 3×3 convolution operations are used to generate feature maps of different dimensions P0, P1, P2, P3, P4, P5, P6; the formula is:
[0078] P l =Conv(P l-1 ),l∈{1,2,3,4,5,6}
[0079] Among them, P lis the feature map of this layer, P l-1 It is the output of the previous layer; the P6 feature map is the smallest and the P0 feature map is the largest. As the number of layers increases, the size of the feature map gradually decreases;
[0080] B2: Apply a shared convolution operation Conv at each layer share , and concatenate them according to the channel dimension as follows:
[0081] P l ′=Contact(Conv share (P l )), l∈{1,2,3,4,5,6}
[0082] B3: The result P of splicing the shared convolution kernel l 'After a downsampling module, it is added pixel by pixel with P5 and then goes through a 3×3 convolution to obtain the output result
[0083]
[0084] like Figure 4 As shown in the figure, the specific implementation steps of designing the RSCD detection head and replacing the original Yolov11 detection head include:
[0085] C1: Input three feature maps x i , are processed by a separate 1×1 convolution and normalization layer, and then further processed by a 3×3 shared convolution, where the three input feature maps x i , x0 is Figure 1 The output feature map of the C3K2 module pointing to the upper side of the RSCD module, x1 is Figure 1 The middle point is the output feature map of the C3K2 module in the middle of the RSCD module, and x2 is Figure 1 The output feature map of the Conv module in the RSCD module is as follows:
[0086]
[0087] Among them, Conv_GN is a convolutional layer with group normalization; the feature map x i After 1×1 convolution and group normalization operations, a new feature map is obtained.
[0088] C2: All processed feature maps are fused through a shared convolutional layer:
[0089]
[0090] Among them, Share_Conv is a shared convolution module;
[0091] C3: Feature map after shared convolution The convolution is sent to the separation regression branch to obtain Conv_Reg and the classification branch Conv_Cls; then, each uses a 1×1 convolution with independent weights to complete the first-stage prediction task, and the output Conv_Reg of the separation regression branch is scaled using the Scale layer. First, the Scale layer initializes the weights for each channel of the feature map and initializes these weights to 1, indicating that the feature map is not scaled in the initial state. When the network is trained, the Scale layer updates the original parameters based on the current task loss to help locate underwater targets of different sizes.
[0092] Use the NWD activation function to replace the activation function in Yolov11. The expression of the NWD activation function is:
[0093]
[0094] in, is the center coordinate of the prediction box, (w a ,h a ) is the prediction box width and height, is the coordinate of the center position of the real frame, (w b ,h b ) are the actual frame width and height, C is a user-defined constant whose value range is the positive real number domain.
[0095] S3: Train and test the improved Yolov11 network through the training set and test set respectively, and use the trained Yolov11 network as the underwater target detection model;
[0096] In this embodiment, according to the set training hyperparameters, after 300 rounds of forward derivation, a set of optimal weight parameters is obtained, which is saved for obtaining prediction results;
[0097] Figure 5 The relationship between the number of iterations and the prediction accuracy during model training is shown. As the number of iterations increases, the model prediction accuracy continues to improve.
[0098] Figure 6 The relationship between the number of model iterations and the loss value is shown. As the number of iterations increases, the loss value gradually decreases. Since the loss value reflects the difference between the model predicted label and the true label, the lower the loss value, the smaller the difference between the two. This also indirectly reflects that as the number of iterations increases, the model prediction accuracy is also constantly improving.
[0099] S4: Obtain underwater target detection results through the underwater target detection model;
[0100] The trained model is used to test the URPC dataset and the final performance of the model is tested. Figure 7 A set of results predicted by the model is shown. The model proposed by the present invention is then compared with the benchmark model. The following table compares the results before and after lightweighting.
[0101]
[0102] As can be seen from the above table, the number of parameters of the improved Yolov11 model has decreased by 0.37M, a decrease of 14.3%. At the same time, mAP@0.5 has increased by 0.3%, the recall rate has increased by 4.1%, and the mAP has also increased by 1.4%. The experimental results show that the model proposed in this invention is superior to the benchmark model in various performance indicators, demonstrating its potential in practical applications.
[0103] This embodiment also provides an underwater target detection system based on improved Yolov11, including a peripheral interface, a memory and a processor;
[0104] Peripheral interface, used to realize data input and output in the process of data interaction with various external devices, so as to ensure the effective connection and communication between the system and external devices;
[0105] a memory for storing computer program instructions capable of being executed on the processor;
[0106] The processor is used to execute the steps of an underwater target detection method based on improved Yolov11 when running computer program instructions.
Claims
1. An underwater target detection method based on improved Yolov11, characterized in that: The steps include: S1: Divide the collected data set into training set and test set in proportion; S2: Improve the Yolov11 network; S3: Train and test the improved Yolov11 network through the training set and test set respectively, and use the trained Yolov11 network as the underwater target detection model; S4: Obtain underwater target detection results through the underwater target detection model.
2. The underwater target detection method based on improved Yolov11 according to claim 1 is characterized in that: The improvements to the Yolov11 network in step S2 include: designing a C3k2_Faster_EMA module and replacing the Yolov11 backbone C3K2 module; proposing a deep convolutional feature pyramid structure and replacing the SPPF module in Yolov11; designing an RSCD detection head and replacing the original Yolov11 detection head; and using an NWD activation function to replace the activation function in Yolov11.
3. The underwater target detection method based on improved Yolov11 according to claim 2 is characterized in that: The specific implementation steps of designing the C3k2_Faster_EMA module in step S2 and replacing the Yolov11 trunk C3K2 module include: A1: Send the input feature map X to Pconv, the specific formula is: Among them, Y i is the eigenvalue in the output image, i is the size of the output feature map, X k is the feature value in the input feature map X, W k is the weight in the convolution kernel, the weight range is 0 to 1, M k It is the value in the mask image, and its value can only be 0 and 1. 1 means the feature is known, and 0 means the feature is unknown. The feature corresponding to the position 0 in the mask image does not participate in the calculation; A2: Input the result Y after Pconv into the MLP layer, where Y is Y i It is obtained by splicing along the horizontal axis. The specific formula is: h mlp =Conv 1×1 (Y,dim) X mlp =Conv 1×1 (h mlp ,dim) A3: The result X after the MLP layer mlp Input into the DropPath layer, the specific formula is: X drop =DropPath(X mlp ) A4: Pass the result X through the DropPath layer drop The final result is obtained by sending it into the exponential moving average model (EMA), and its specific formula is: X att =EMA(X drop )。 4. The underwater target detection method based on improved Yolov11 according to claim 2 is characterized in that: The specific implementation steps of proposing a deep convolution feature pyramid structure and replacing the SPPF module in Yolov11 in step S2 include: B1: According to the bottom to the top of the deep convolution feature pyramid structure, 3×3 convolution operations are used to generate feature maps of different dimensions P0, P1, P2, P3, P4, P5, P6; the formula is: P l =Conv(P l-1 ),l∈{1,2,3,4,5,6} Among them, P l is the feature map of this layer, P l-1 It is the output of the previous layer; the P6 feature map is the smallest and the P0 feature map is the largest. As the number of layers increases, the size of the feature map gradually decreases; B2: Apply a shared convolution operation Conv at each layer share , and concatenate them according to the channel dimension as follows: P′ l =Contact(Conv share (P l )),l∈{1,2,3,4,5,6} B3: The result P of splicing the shared convolution kernel l 'After a downsampling module, it is added pixel by pixel with P5 and then goes through a 3×3 convolution to obtain the output result 5. The underwater target detection method based on improved Yolov11 according to claim 2 is characterized in that: The specific implementation steps of designing the RSCD detection head and replacing the original Yolov11 detection head in step S2 include: C1: Input three feature maps x j , processed by a separate 1×1 convolution and normalization layer, and then further processed by a 3×3 shared convolution. The specific formula is as follows: Among them, Conv_GN is a convolutional layer with group normalization; the feature map x j After 1×1 convolution and group normalization operations, a new feature map is obtained. C2: All processed feature maps are fused through a shared convolutional layer: Among them, Share_Conv is a shared convolution module; C3: Feature map after shared convolution The convolution operation is sent to the separation regression branch to obtain Conv_Reg and the classification branch Conv_Cls. Then, each uses a 1×1 convolution with independent weights to complete the first-stage prediction task, and uses the Scale layer to scale the output Conv_Reg of the separation regression branch, dynamically adjust the importance of the feature map, and help locate underwater targets of different sizes.
6. The underwater target detection method based on improved Yolov11 according to claim 2 is characterized in that: The expression of the NWD activation function in step S2 is: <h2 style=";text-align:left;direction:ltr">NWD(N<h2 style=";text-align:left;direction:ltr"> a <h2 style=";text-align:left;direction:ltr"> ,N<h2 style=";text-align:left;direction:ltr"> b <h2 style=";text-align:left;direction:ltr"> )=exp in, is the center coordinate of the prediction box, (w a ,h a ) is the prediction box width and height, is the coordinate of the center position of the real frame, (w b ,h b ) are the actual frame width and height, C is a user-defined constant whose value range is the positive real number domain.
7. An underwater target detection system based on improved Yolov11, characterized in that: Includes peripheral interfaces, memory and processor; The peripheral interface is used to implement data input and output during data interaction with various external devices, thereby ensuring effective connection and communication between the system and the external devices; The memory is used to store computer program instructions that can be executed on the processor; The processor is used to execute the steps of an underwater target detection method based on improved Yolov11 according to any one of claims 1 to 6 when running the computer program instructions.
Citation Information
Cited By
Traditional Chinese medicine dispensing authenticity real-time detection method, system and equipment and storage medium
CN120259686A
Method, system, device and storage medium for real-time detection of authenticity of traditional Chinese medicine preparations
CN120259686B
Lightweight identification method for cavity diseases in road
CN121074632A