Lightweight arbitrary direction ship detection method based on improved YOLO11 network

By improving the YOLO11 network, the introduction of MSDA, GSConv and CSPStage modules is solved, and the problem of insufficient ship detection accuracy and efficiency in remote sensing images is achieved, and high-precision and fast ship detection in any direction is achieved.

CN120032258APending Publication Date: 2025-05-23FUJIAN CHUANZHENG COMM COLLEGE
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411849187.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-16
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art has problems with insufficient detection accuracy and efficiency in ship detection in remote sensing images, especially when ship detection is detected in different directions, background pixels and interference factors are easily introduced, and the detection speed is not enough to meet the real-time requirements.

Method used

Using the improved YOLO11 network, by introducing MSDA module, GSConv module and CSPStage module into the Neck network, a lightweight and efficient neck network is designed to achieve efficient fusion of multi-scale features and accuracy of target detection.

Benefits of technology

It significantly improves the detection accuracy and speed of ships in remote sensing images, enhances the perception ability of ships with small sizes and arbitrary directions, reduces missed and missed detection, and improves the robustness and real-timeness of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032258A_ABST
    Figure CN120032258A_ABST
Patent Text Reader

Abstract

The invention discloses a lightweight arbitrary direction ship detection method based on an improved YOLO11 network. A Backbone network performs feature extraction and initial fusion on a remote sensing image containing a ship image to generate an initial feature map; the MSDA module performs modeling on a long-distance pixel dependency relationship among different scale features of the initial feature map to obtain an expanded feature map; the GSConv module carries out standard convolution and depth separable convolution operation on the feature map of the corresponding stage to obtain a convolution feature map; the CSPStage module performs multi-scale feature fusion on the feature maps of the corresponding stages to obtain a fused feature map; and the Head network performs target classification and regression processing on the fusion feature map in the corresponding stage to obtain a detection image containing the ship position. According to the invention, the improved YOLO11 network is constructed to carry out ship image detection in any direction, and the detection precision and speed of the ship in the remote sensing image are obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of image processing, and in particular relates to a lightweight arbitrary-direction ship detection method based on an improved YOLO11 network. Background Art

[0002] Target detection in optical remote sensing images is one of the basic tasks in the field of aerial and satellite image analysis, which aims to classify and locate the targets of interest in the image and effectively extract important information about the target to be detected. Therefore, accurate detection of ships in optical remote sensing images has important practical significance and safety value in the fields of fisheries, maritime search and rescue, and maritime traffic management.

[0003] At present, the detection of ships in remote sensing images is generally achieved by using a lightweight remote sensing image ship detection algorithm based on deep learning or an arbitrary direction target detection algorithm. Although lightweight remote sensing image ship detection based on deep learning has the advantages of high accuracy, high efficiency and can be deployed on edge devices, it is easy to introduce background pixels and interference factors in ship detection in different directions; in particular, for overlapping ship targets, the method based on horizontal frame detection will have the problem of overlapping detection frames, which seriously affects the visualization effect of ship detection and is not conducive to subsequent data analysis. Arbitrary direction target detection has achieved excellent results in the task of arbitrary direction ship detection in remote sensing images, but it still has shortcomings in detection speed and is difficult to meet the real-time requirements of remote sensing scenes. Summary of the invention

[0004] The purpose of the present invention is to solve the shortcomings of the above-mentioned background technology and provide a lightweight arbitrary direction ship detection method based on an improved YOLO11 network that can take into account both detection accuracy and efficiency.

[0005] The technical solution adopted by the present invention is: a lightweight arbitrary-direction ship detection method based on an improved YOLO11 network, wherein the improved YOLO11 network includes a Backbone network, a Neck network and a Head network, wherein the Neck network includes an MSDA module and multiple GSConv modules and multiple CSPStage modules located at different data processing stages;

[0006] The method comprises:

[0007] The Backbone network extracts features and initially fuses remote sensing images containing ship images to generate initial feature maps;

[0008] The MSDA module models the long-distance pixel dependencies between different scale features of the initial feature map to obtain the expanded feature map;

[0009] The GSConv module performs standard convolution and depth-separable convolution operations on the feature maps of the corresponding stage to obtain convolution feature maps;

[0010] The CSPStage module performs multi-scale feature fusion on the feature maps of the corresponding stages to obtain a fused feature map;

[0011] The Head network performs target classification and regression processing on the fused feature map of the corresponding stage to obtain a detection image containing the position of the ship.

[0012] Furthermore, the MSDA module includes

[0013] The first Linear module is used to perform linear projection on the initial feature map to obtain a first feature map;

[0014] The SWDA module is used to dilate different heads of the channels of the first feature map using different dilation rates to obtain a plurality of second feature maps of different scales;

[0015] The Concat submodule is used to concatenate multiple second feature maps of different scales to obtain an overall third feature map;

[0016] The second Linear module is used to perform linear processing on the third feature map to obtain the expanded feature map.

[0017] Furthermore, the GSConv module includes

[0018] The standard Conv layer is used to perform regular convolution operations on the corresponding feature maps to obtain standard feature maps;

[0019] The DWConv layer is used to perform a depth-separable convolution operation on the standard feature map to obtain a convolution feature map;

[0020] The Concat layer is used to perform a cascade operation on the standard feature map and the convolutional feature map to obtain a comprehensive feature map;

[0021] The Shuffle layer is used to redistribute and sort the comprehensive feature map to obtain the convolution feature map.

[0022] Furthermore, the CSPStage module includes

[0023] A first Conv block is used to perform a convolution operation on the corresponding feature map to obtain a fourth feature map;

[0024] Simplify Rep block, used to train and infer the fourth feature map to obtain the fifth feature map

[0025] A second Conv block is used to perform a convolution operation on the fifth feature map to obtain a sixth feature map;

[0026] an element-by-element addition module, configured to add the fourth feature map and the sixth feature map to obtain a seventh feature map;

[0027] The Concat block is used to perform a cascade operation on the fourth feature map and the seventh feature map to obtain the fused feature map.

[0028] Furthermore, the Backbone network includes a first Conv module, a second Conv module, a first C3k2 module, a third Conv module, a second C3k2 module, a fourth Conv module, a third C3k2 module, a fifth Conv module, a fourth C3k2 module and an SPPF module which are connected in sequence.

[0029] Furthermore, the Neck network also includes multiple Concat modules and Upsample modules located at different data processing stages, the Concat module is used to perform a concatenation operation on the corresponding feature map, and the Upsample module is used to perform an upsampling operation on the corresponding feature map.

[0030] Further, the Neck network includes an MSDA module, a first GSConv module, a second GSConv module, a first Concat module, a first CSPStage module, a first Upsample module, a third GSConv module, a second Concat module, a second CSPStage module, a second Upsample module, a third Concat module and a third CSPStage module connected in sequence;

[0031] The first GSConv module is cascaded with the first CSPStage module, and the third Concat module is cascaded with the second C3k2 module in the Backbone network;

[0032] The Head network includes a first Detect module, and the first Detect module is connected to a third CSPStage module.

[0033] Further, the Neck network also includes a fourth GSConv module, a fourth Concat module, and a fourth CSPStage module sequentially connected to the third CSPStage module, and the fourth Concat module is cascaded with the second CSPStage module;

[0034] The Head network includes a second Detect module, and the second Detect module is connected to a fourth CSPStage module.

[0035] Furthermore, the Neck network also includes a sixth GSConv module and a fifth GSConv module, a fifth Concat module, and a fifth CSPStage module sequentially connected to the fourth CSPStage module, and the sixth GSConv module is connected between the second CSPStage module and the fifth Concat module;

[0036] The Head network includes a third Detect module, and the third Detect module is connected to the fifth CSPStage module.

[0037] The beneficial effects of the present invention are:

[0038] The present invention constructs an improved YOLO11 network to detect ship images in any direction, which significantly improves the detection accuracy and speed of ships in remote sensing images.

[0039] The present invention improves the Neck network in the original YOLOv11 network and designs a lightweight and efficient Neck network. The entire improved YOLO11 network has the advantages of high precision and light weight, and effectively meets the needs of real-time detection of ships in any orientation in remote sensing images.

[0040] The present invention adds an MSDA module in the Neck network, and models the long-distance pixel dependency relationship between features of different scales through the MSDA module, thereby eliminating redundant query block information, enriching the multi-scale feature semantic information of ships, enhancing the perception of small-sized and arbitrarily oriented ships, reducing missed detections and false detections, improving detection accuracy and robustness, and significantly improving the detection accuracy of ships in remote sensing images.

[0041] The present invention introduces a lightweight convolution module GSConv into the Neck network, which can retain the hidden dependencies between channels to the maximum extent and realize the interaction of local feature information, thereby improving the accuracy of target detection.

[0042] The present invention replaces the C3k2 module in the original Neck network with a CSPStage module. The CSPStage module can capture the underlying semantic information and spatial information of the feature map, realize the efficient fusion of multi-scale features, improve the feature utilization efficiency, and reduce the computing resource requirements of the network. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 The schematic diagram is a principle diagram of improving the YOLO11 network in the present invention.

[0044] Figure 2 It is a schematic diagram of the principle of the MSDA module of the present invention.

[0045] Figure 3Schematic diagram of the principle of the GSConv module of the present invention.

[0046] Figure 4 It is a schematic diagram of the principle of the CSPStage module of the present invention.

[0047] Figure 5 Schematic diagram of the comparison of heat maps generated by different network detection methods (in the figure: (a) is the input image; (b) is the image output by YOLO11n; (c) is the image output by YOLO11n-MSDA; (d) is the image output by YOLO11N-MSDA-CSPstage; (e) is the image output by the detection method of the present invention).

[0048] Figure 6 Schematic diagram of ship detection results of the detection method of the present invention. DETAILED DESCRIPTION

[0049] The specific embodiments of the present invention are further described below in conjunction with the accompanying drawings. It should be noted that the description of these embodiments is used to help understand the present invention, but does not constitute a limitation of the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0050] The present invention provides a lightweight arbitrary direction ship detection method based on an improved YOLO11 network, such as Figure 1 As shown, the improved YOLO11 network includes a Backbone network, a Neck network and a Head network, and the Neck network includes an MSDA module and multiple GSConv modules, multiple CSPStage modules, multiple Concat modules and multiple Upsample modules located at different data processing stages;

[0051] The method comprises:

[0052] The Backbone network extracts features and initially fuses remote sensing images containing ship images to generate initial feature maps;

[0053] The MSDA module models the long-distance pixel dependencies between different scale features of the initial feature map to obtain the expanded feature map;

[0054] The GSConv module performs standard convolution and depth-separable convolution operations on the feature maps of the corresponding stage to obtain convolution feature maps;

[0055] The CSPStage module performs multi-scale feature fusion on the feature maps of the corresponding stages to obtain a fused feature map;

[0056] The Concat module is used to concatenate corresponding feature maps;

[0057] The Upsample module is used to perform upsampling operations on the corresponding feature maps;

[0058] The Head network performs target classification and regression processing on the fused feature map of the corresponding stage to obtain a detection image containing the position of the ship.

[0059] The present invention significantly improves the detection accuracy and speed of ships in remote sensing images by constructing an improved YOLO11 network to detect ship images in any direction. It can be understood that the above-mentioned improved YOLO11 network is an improvement on the original YOLO11n network. The MSDA (multi-scale dilated attention) module, GSConv (lightweight convolution) module and CSPStage (cross-stage partial connection) module are introduced into the original Neck network to design a lightweight and effective neck network (Neck network). The MSDA module constructs long-distance pixel dependencies on the feature map output by the Backbone network, enriching the contextual semantic information of the multi-scale feature map;

[0060] The GSConv module performs channel compression at the neck, reducing the number of feature map channels and reducing the computational burden while maintaining the integrity of feature information; the CSPStage module further captures the underlying semantic and spatial information of the feature map and achieves efficient fusion of multi-scale features.

[0061] In some embodiments, since the scale of ships in optical remote sensing images is small and the features are not obvious, conventional convolutional neural networks can only construct local dependencies of features, ignoring the impact of long-distance pixel dependencies on ship detection. Although the visual converter can use the global attention mechanism to establish long-distance context dependencies of ship features between image blocks, it takes twice the computational cost. Therefore, in order to balance the computational complexity and feature dependency, the present invention introduces the MSDA module in the Neck network.

[0062] The structure of the MSDA module is as follows Figure 2 As shown in Figure 1, it includes the first Linear module, the SWDA module, the Concat submodule and the second Linear module. The complete names of the modules are not shown in the figure. The MSDA module adopts a multi-head design, that is, the channels of the feature map are divided into n different heads. For a given feature map X∈RH×W×C(i.e., the feature map output by the SPPF module), X is the feature map, R is a real number set, H, W, C are the height, width, and number of channels of the feature map respectively, the first feature map is obtained by linear projection through the first Linear module, and the corresponding query Q, key value K, and value V are obtained; the SWDA (sliding window dilated attention) module uses different dilation rates to perform multi-scale sliding window dilated attention operations (i.e., dilation processing) on ​​different heads of the channels of the first feature map, which is used to capture semantic information at different scales and obtain multiple second feature maps of different scales. The Concat submodule splices multiple second feature maps of different scales (i.e., the outputs of multiple different heads) to obtain the overall third feature map, and the second Linear module performs linear processing (i.e., feature aggregation) on the third feature map to obtain the dilated feature map. This process can be described by formulas (1)-(2).

[0063] h i =SWDA(Q i ,K i ,V i ,r i ),1≤i≤n (1)

[0064] X 4 =Linear(Concat[h 1 ,…,h n ]) (2)

[0065] In the formula, h i represents the second feature map output by the SWDA module of the i-th head, r i represents the expansion rate of the ith head, X 4 Represents the expanded feature map output by the second Linear module.

[0066] The MSDA module uses the local and sparsity properties of the self-attention mechanism to model the long-distance pixel dependencies between features of different scales, thereby eliminating redundant query block information and enriching the multi-scale feature semantic information of ships. The addition of the MSDA module not only improves the computational efficiency, but also expands the range of the receptive field, enabling the network to capture complex features more effectively. In addition, the MSDA module can suppress background interference, adaptively aggregate multi-scale features, enhance the perception of small-sized and arbitrarily oriented ships, reduce missed detections and false detections, improve detection accuracy and robustness, and significantly improve the detection accuracy of ships in remote sensing images.

[0067] In some embodiments, the ship detection task in remote sensing images often requires real-time performance, which poses a challenge to the lightweight of the detection network. However, the Conv blocks in the neck of the original YOLO11n network lose some semantic information during feature transfer and fusion. In addition, the Conv blocks introduce too many parameters during the training process, making it difficult to deploy the ship detection algorithm on mobile platforms with limited resources. Therefore, the present invention introduces a lightweight convolutional GSConv module in the neck network, and its structure is as shown in Figure 3 shown.

[0068] The GSConv module effectively improves the performance of object detection by introducing the ideas of group sparse convolution and group dense convolution. Compared with traditional convolution modules, the GSConv module more effectively retains detailed information when processing large-scale objects, thereby improving the accuracy of object detection. Specifically, the GSConv module consists of a standard Conv (convolution) layer, a DWConv layer, a Concat layer, and a Shuffle layer. The full names of each module are not shown in the figure. The standard Conv layer performs a regular convolution operation on the feature map with channel C1 to generate a standard feature map with channel C2 / 2. Subsequently, the DWConv layer performs a DWConv (depthwise separable convolution) operation on the standard feature map with channel number C2 / 2, and uses the per-channel convolution and pointwise convolution techniques to obtain a convolutional feature map with channel number C2 / 2. The convolutional feature map has the advantages of fewer parameters and higher computational efficiency compared with the standard feature map, but contains less feature information. Therefore, it is necessary to perform a concatenation operation on the two feature maps with channel number C2 / 2 through the Concat layer to obtain a comprehensive feature map with channel number C2 to compensate for the feature information lost during the spatial compression and channel expansion processes. Finally, the Shuffle layer performs a shuffle (random permutation) operation on the features of channel C2, and through uniform mixing, fully mixes the feature information from the standard Conv layer and the DWConv layer, and maximally retains the hidden dependencies between channels to achieve local feature information interaction.

[0069] In some embodiments, in view of the problems of low utilization rate of ship features and insufficient fusion expression ability in optical remote sensing images, the idea of RepGFPN is introduced. The C3k2 module in the neck of the original YOLO11n algorithm is replaced with a CSPStage module to achieve efficient fusion of multi-scale features, which improves the feature utilization efficiency and reduces the computational resource requirements of the network. The CSPStage module uses a Conv block to reduce the dimension of the feature channels, and then constructs a network structure with a reparameterization mechanism through a Simplify Rep block. The simplified Rep structure is trained using a convolution with a kernel of 3×3 and inferred using a convolution with a kernel of 1×1. The structure of the CSPStage module is as shown in Figure 4As shown in the figure, it consists of a first Conv block, a Simplify Rep block, a second Conv block, an element-by-element addition module, a Concat block and a third Conv block. The complete names of the modules are not shown in the figure.

[0070] For input features of different scales, the CSPStage module first combines multi-scale feature information and divides it into two layers for processing. The upper layer uses only one convolution block with a kernel of 1×1. In contrast, the lower layer uses the Simplify Rep module and a convolution block with a kernel of 3×3 to complete the staged feature fusion through data transmission to maintain the fluidity of multi-scale feature information. Finally, the feature information of the upper and lower layers is fused through a cascade operation, and the feature fusion result is output through a Conv block.

[0071] It can be understood that according to the description of the various parts of the above improved Y OLO11 network, the composition of the Backbone network, Neck network and Head network is as follows, and the complete names of the modules are not shown in the figure:

[0072] The Backbone network includes a first Conv module, a second Conv module, a first C3k2 module, a third Conv module, a second C3k2 (feature extraction) module, a fourth Conv module, a third C3k2 module, a fifth Conv module, a fourth C3k2 module and an SPPF (Spatial pyramid pooling fast) module connected in sequence.

[0073] The Neck network includes an MSDA module, a first GSConv module, a second GSConv module, a first Concat module, a first CSPStage module, a first Upsample module, a third GSConv module, a second Concat module, a second CSPStage module, a second Upsample module, a third Concat module, a third CSPStage module, a fourth GSConv module, a fourth Concat module, a fourth CSPStage module, a fifth GSConv module, a fifth Concat module, a fifth CSPStage module, and a sixth GSConv module, which are connected in sequence, and the sixth GSConv module is connected between the second CSPStage module and the fifth Concat module.

[0074] The Head network includes a first Detect module, a second Detect module and a third Detect module. The first Detect module is connected to the third CSPStage module, the second Detect module is connected to the fourth CSPStage module, and the third Detect module is connected to the fifth CSPStage module.

[0075] Efficient connection strategies (such as skip connections and residual connections) are introduced at appropriate positions in the Neck network to ensure effective gradient transfer and efficient feature transfer. The specific connections are: the first GSConv module is cascaded with the first CSPStage module, the third Concat module is cascaded with the second C3k2 module in the Backbone network; the fourth Concat module is cascaded with the second CSPStage module.

[0076] After the improved YOLO11 network is constructed, it is trained through a specific training set. The training set can be selected from the HRSC2016 dataset described below. The improved YOLO11 network after training is used to detect ships in remote sensing images, which has good detection efficiency and accuracy.

[0077] It is understandable that in order to verify the performance of the detection method of the present invention, the HRSC2016 data set is selected for verification. The HRSC2016 data set contains 2976 ship targets at any angle and 1061 nearshore background images and sea background images, most of which are nearshore background images. The ship samples in the remote sensing image are annotated using a rotating frame, the image resolution is about 1150×780 pixels, and the storage format is .bmp. The image background of this data set (including ports, sea surfaces, small islands, thin clouds, etc.) is complex and diverse, the scale of ships varies greatly, and the ships at the port terminal are densely arranged, which is close to the real application scene. Therefore, the performance test of the ship detection of the present invention is carried out using the HRSC2016 data set, which can effectively verify the practicality of the method proposed in the present invention in the task of marine ship detection. The remote sensing image ship data set is randomly divided into a training set, a validation set, and a test set in a ratio of 8:1:1.

[0078] It should be noted that the improved YOLO11 network is trained through the above training set. The improved YOLO11 network is initialized without using the pre-trained network weights. When the loss function is smooth, the improved YOLO11 network training is completed. The specific hyperparameter information of the network used for training is shown in Table 1.

[0079] Table 1 Network specific hyperparameter information

[0080]

[0081] It should be noted that in order to objectively evaluate the detection performance of the improved Y OLO11 network of the present invention, the present invention introduces precision, recall, mAP50, network parameters, network size and FLOP as evaluation indicators, and verifies the accuracy of ship detection through the validation set.

[0082] Precision refers to the ratio of the number of positive ship samples predicted by the network to the total number of predicted samples. The precision calculation is shown in formula (3):

[0083]

[0084] Among them, TP represents the sample predicted to be positive for the ship and actually is also positive for the ship, and FP represents the sample predicted to be positive for the ship and actually is negative for the ship.

[0085] Recall rate refers to the ratio of the number of positive samples of ships correctly predicted by the network to the number of all positive samples in the data set. The calculation formula of recall rate is shown in formula (4):

[0086]

[0087] In the formula, FN represents samples that are predicted to be negative but are actually positive.

[0088] The average precision (AP) represents the area formed by the PR curve with the recall rate as the x-axis and the precision rate as the y-axis. The calculation is shown in formula (5):

[0089]

[0090] The mean average precision (mAP) is the result of weighted average of the AP values ​​of all sample categories, which is used to measure the detection performance of the network for all ship categories. The calculation of mAP is shown in formula (6):

[0091]

[0092] Among them, AP j Represents the AP value of the jth category, and M represents the number of categories in the training dataset. mAP@0.5 represents the average accuracy when the IoU of the detection network is set to 0.5. mAP@0.5:0.95 represents the average accuracy when the IoU value is in the range of 0.5 to 0.95.

[0093] In the application, in order to verify the enhanced effect of each improved module in the Neck network on ship detection in optical remote sensing images, an ablation experiment was conducted on a test set containing only the ship category dataset (i.e., one module was added at a time, and the performance was evaluated one by one), using the same dataset and experimental settings. The ablation experiment results are shown in Table 2.

[0094] As can be seen from Table 2, when the network size, FLOPs and number of parameters are the same as the original YOLO11n network, the introduction of the MSDA module increases mAP@0.5 and mAP@0.5:0.95 by about 1%, proving the effectiveness of the MSDA module in building dependencies between different scale features and capturing multi-scale semantic information. After the introduction of the CSPStage module, the network size increased by 2.3MB, FLOPs increased by 2G, and the number of parameters increased by 1.2×106, but the precision and recall rates were improved, with mAP@0.5 increased by 1.3% and mAP@0.5:0.95 increased by 0.8%. The CSPStage module has been proven to improve feature utilization and achieve efficient fusion of multi-scale features. After the introduction of the GSConv module, FLOPs was reduced by 0.3G, the network size was reduced by 0.3MB, the number of parameters was 3.5×106, the precision rate reached 99.4%, the recall rate reached 96.1%, the mAP@0.5 reached 97.6%, and the mAP@0.5:0.95 reached 90.1%. Therefore, the method proposed in the present invention has the advantages of high precision and light weight, and more effectively meets the needs of real-time detection of ships in any orientation in remote sensing images.

[0095] Table 2 Detection effect of the improved points of the method of the present invention

[0096]

[0097] In the application, in order to intuitively and easily reflect the feature map area that the network pays attention to, the gradient weighted class activation map is introduced to generate the heat map of YOLO11n and the method proposed in the present invention. The redder the area in the heat map, the higher the network's attention. On the contrary, the area with low attention is represented by blue. The heat map comparison results of the image detection results of different methods are shown in the figure. Figure 5 shown.

[0098] from Figure 5It can be seen that for marine ship images, although the YOLO11n network can effectively focus on large-sized ships, it is difficult to accurately perceive small-sized ships and mistake waves for hulls. With the introduction of the MSDA module, CSPStage module and GSConv module, the network focuses on small ships in increasingly precise areas, effectively suppressing background interference. For blurred remote sensing images, the YOLO11n network only obtains partial information about ship features and fails to capture ship edge information. The method proposed in the present invention completely extracts ship information and effectively distinguishes the coastline from the hull, indicating that the improved YOLO11 network can accurately detect ships in low-visibility weather such as fog. For close-range ship images, the YOLO11n network mistakenly regards two ships as targets, and the area of ​​concern includes the aisle between the ships. The improved YOLO11 network of the present invention can avoid the interference problem caused by too close distance. It can be seen that the method proposed in the present invention can effectively deal with the problem of too close distance between ships and realize accurate detection and positioning of ships of multiple sizes and arbitrary angles.

[0099] In the application, the present invention introduces an attention mechanism to guide the network training in the neck network, and verifies the influence of multiple different attention modules on the ship detection performance. The experimental results for different attention modules are shown in Table 3. Although the GAM module and the Biformer module have improved in mAP@0.5 and mAP@0.5:0.95, the network size, the number of detection reagents and the amount of parameters are larger than the method proposed in this article, and are not suitable for deployment on edge platforms with limited resources. The test results of the EMA module, LSK module, and MSDA module are similar in terms of network size, detection time, and the amount of parameters, but the MSDA module has the highest score in mAP@0.5. Therefore, the present invention adds an MSDA module to the neck of the YOLO11 network for ship detection, which can not only keep the network lightweight, but also improve the accuracy of ship detection.

[0100] Table 3 Performance comparison of different attention modules

[0101]

[0102] In the application, in order to further verify the superiority of the detection method of the present invention, a comparative experiment was conducted with the current mainstream OBB detection algorithm. Gliding Vertex, R3Det, ​​Oriented RCNN, RoI-Transformer, H2RBox-v2, PSC, YOLOv8n-OBB and YOLO11n-OBB algorithms were introduced to compare the detection performance of ships in any direction in remote sensing images. The experimental results are shown in Table 4.

[0103] As can be seen from Table 4, although Gliding Vertex, R3Det, ​​Oriented RCNN, RoI-Transformer, H2RBox-v2 and PSC methods have achieved good mAP@0.5 and mAP@0.5:0.95 in remote sensing image ship detection, their network parameters are too many, resulting in excessive storage space. Although the YOLO series network has low FLOP, it has the advantages of small network size and few parameters. In the case of the YOLOv8n-OBB network with a network size of 7.5MB and a parameter of 3.1M, mAP@0.5 reached 93.7% and mAP@0.5:0.95 reached 86.2%. YOLO11n-OBB achieved better performance under the conditions of smaller network size and fewer parameters. Although the method proposed in the present invention is slightly larger than YOLO11n-OBB in terms of network size and parameter amount, it achieved the best performance, with mAP@0.5 reaching 97.6% and mAP@0.5:0.95 reaching 90.1%. It can be seen that the method proposed in the present invention achieves high-precision detection performance for ships in any direction at the cost of a smaller storage space, and has the advantages of light structure and fast detection.

[0104] Table 4 Performance comparison of different detection methods

[0105]

[0106] In order to more clearly perceive the improvement effect of the present invention, we conducted image-by-image detection on the test set, and selected representative data in a complex background with large scale differences, dense distribution, and low distinction from the background for visualization experiments. The ship detection results are shown in the figure below. Figure 6 shown.

[0107] Experimental results show that compared with the existing detection methods, the proposed method exhibits superior ship detection performance, reaching 97.6% mAP@0.5 and 90.1% mAP@0.5:0.95, and the detection speed meets the requirements of real-time ship detection.

[0108] It should be understood that the specific order or hierarchy of steps in the disclosed process is an example of an exemplary method. Based on design preferences, it should be understood that the specific order or hierarchy of steps in the process can be rearranged without departing from the scope of protection of the present disclosure. The attached method claims present the elements of the various steps in an exemplary order and are not intended to be limited to the specific order or hierarchy described.

[0109] In order to make the description of the present disclosure more detailed and complete, the above description is provided for the implementation and specific embodiments of the present invention; however, this is not the only form of implementing or using the specific embodiments of the present invention. The implementation covers the features of multiple specific embodiments and the method steps and sequences used to construct and operate these specific embodiments. However, other specific embodiments can also be used to achieve the same or equivalent functions and step sequences.

[0110] Those skilled in the art may also appreciate that the various illustrative logic blocks, units, and steps listed in the embodiments of the present invention may be implemented by electronic hardware, computer software, or a combination of the two. To clearly demonstrate the interchangeability of hardware and software, the various illustrative components, units, and steps described above have generally described their functions. Whether such functions are implemented by hardware or software depends on the specific application and the design requirements of the entire system. Those skilled in the art may use various methods to implement the described functions for each specific application, but such implementation should not be understood as exceeding the scope of protection of the embodiments of the present invention.

[0111] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed by the present invention should be included in the protection scope of the present invention. The contents not described in detail in this specification belong to the prior art known to professional and technical personnel in the field.

Claims

1. A lightweight arbitrary-direction ship detection method based on an improved YOLO11 network, characterized in that: The improved YOLO11 network includes a Backbone network, a Neck network and a Head network, wherein the Neck network includes an MSDA module and multiple GSConv modules and multiple CSPStage modules located at different data processing stages; The method comprises: The Backbone network extracts features and initially fuses remote sensing images containing ship images to generate initial feature maps; The MSDA module models the long-distance pixel dependencies between different scale features of the initial feature map to obtain the expanded feature map; The GSConv module performs standard convolution and depth-separable convolution operations on the feature maps of the corresponding stage to obtain convolution feature maps; The CSPStage module performs multi-scale feature fusion on the feature maps of the corresponding stages to obtain a fused feature map; The Head network performs target classification and regression processing on the fused feature map of the corresponding stage to obtain a detection image containing the position of the ship.

2. The lightweight arbitrary-direction ship detection method based on the improved YOLO11 network according to claim 1 is characterized in that: The MSDA module includes The first Linear module is used to perform linear projection on the initial feature map to obtain a first feature map; The SWDA module is used to dilate different heads of the channels of the first feature map using different dilation rates to obtain a plurality of second feature maps of different scales; The Concat submodule is used to concatenate multiple second feature maps of different scales to obtain an overall third feature map; The second Linear module is used to perform linear processing on the third feature map to obtain the expanded feature map.

3. The lightweight arbitrary-direction ship detection method based on the improved YOLO11 network according to claim 1, characterized in that: The GSConv module includes The standard Conv layer is used to perform regular convolution operations on the corresponding feature maps to obtain standard feature maps; The DWConv layer is used to perform a depth-separable convolution operation on the standard feature map to obtain a convolution feature map; The Concat layer is used to perform a cascade operation on the standard feature map and the convolutional feature map to obtain a comprehensive feature map; The Shuffle layer is used to redistribute and sort the comprehensive feature map to obtain the convolution feature map.

4. The lightweight arbitrary-direction ship detection method based on the improved YOLO11 network according to claim 1, characterized in that: The CSPStage module includes A first Conv block is used to perform a convolution operation on the corresponding feature map to obtain a fourth feature map; Simplify Rep block, used to train and infer the fourth feature map to obtain the fifth feature map A second Conv block is used to perform a convolution operation on the fifth feature map to obtain a sixth feature map; an element-by-element addition module, configured to add the fourth feature map and the sixth feature map to obtain a seventh feature map; The Concat block is used to perform a cascade operation on the fourth feature map and the seventh feature map to obtain the fused feature map.

5. The lightweight arbitrary-direction ship detection method based on the improved YOLO11 network according to claim 1, characterized in that: The Backbone network includes a first Conv module, a second Conv module, a first C3k2 module, a third Conv module, a second C3k2 module, a fourth Conv module, a third C3k2 module, a fifth Conv module, a fourth C3k2 module and an SPPF module which are connected in sequence.

6. The lightweight arbitrary-direction ship detection method based on the improved YOLO11 network according to claim 1, characterized in that: The Neck network also includes multiple Concat modules and Upsample modules located at different data processing stages, the Concat module is used to perform a splicing operation on the corresponding feature map, and the Upsample module is used to perform an upsampling operation on the corresponding feature map.

7. The lightweight arbitrary-direction ship detection method based on the improved YOLO11 network according to claim 6 is characterized in that: The Neck network includes an MSDA module, a first GSConv module, a second GSConv module, a first Concat module, a first CSPStage module, a first Upsample module, a third GSConv module, a second Concat module, a second CSPStage module, a second Upsample module, a third Concat module and a third CSPStage module connected in sequence; The first GSConv module is cascaded with the first CSPStage module, and the third Concat module is cascaded with the second C3k2 module in the Backbone network; The Head network includes a first Detect module, and the first Detect module is connected to a third CSPStage module.

8. The lightweight arbitrary-direction ship detection method based on the improved YOLO11 network according to claim 7, characterized in that: The Neck network also includes a fourth GSConv module, a fourth Concat module, and a fourth CSPStage module sequentially connected to the third CSPStage module, and the fourth Concat module is cascaded with the second CSPStage module; The Head network includes a second Detect module, and the second Detect module is connected to a fourth CSPStage module.

9. The lightweight arbitrary-direction ship detection method based on the improved YOLO11 network according to claim 8, characterized in that: The Neck network also includes a sixth GSConv module and a fifth GSConv module, a fifth Concat module, and a fifth CSPStage module sequentially connected to the fourth CSPStage module, and the sixth GSConv module is connected between the second CSPStage module and the fifth Concat module; The Head network includes a third Detect module, and the third Detect module is connected to the fifth CSPStage module.

Citation Information

Cited By

  • Target detection method, device and equipment

    CN120953789A

  • Field operation safe wearing detection method based on improved YOLOv11 algorithm

    CN122024179A