A sea surface ship small target detection method based on improved YOLOv11

CN122821082APending Publication Date: 2026-09-25HARBIN ENG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610841752.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-11
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

然而海上目标检测作为特定领域应用,面临比通用场景更为严峻的挑战,海面镜面反射、波浪、雾霾及多变光照导致目标外观剧烈变化,产生大量噪声干扰;近处船只与远处小型浮标尺寸差异巨大,且远端船舶及小型漂浮目标在图像中往往呈现为极小像素区域,小目标特征信息极其匮乏;船舶等目标具有细长、方向多变等特点,对边界框回归提出更高要求

Benefits of technology

首先,通过采用基于部分卷积的C3k2_PConv模块替代传统卷积结构,在保持特征提取能力的同时显著降低了计算开销,有效抑制了复杂海面背景中的冗余特征干扰。其次,在骨干网络末端引入跨维度注意力机制,能够自适应地增强船舶目标的特征响应并抑制波浪、反光等背景噪声,提升了目标与背景的区分度。再者,采用基于内容感知的动态上采样算子替代固定插值方法,能够根据图像内容动态调整采样位置,精准还原船舶边缘细节,避免了小目标特征在上采样过程中的丢失。最后,通过增设高分辨率P2检测头,弥补了深层网络对小目标感知能力不足的缺陷,实现了从微小船舶到大型船只的全尺度覆盖。上述改进协同作用,显著提升了海面船舶小目标的检测精度与鲁棒性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821082A_ABST
    Figure CN122821082A_ABST
Patent Text Reader

Abstract

The application discloses a sea surface ship small target detection method based on an improved YOLOv11, and comprises the following steps: obtaining a sea surface ship image, and obtaining an input feature map based on the sea surface ship image; a backbone network of the improved YOLOv11 network is used to perform feature extraction on the input feature map, wherein a C3k2_PConv module is used to perform spatial feature extraction on the input feature map to obtain a spatial feature map, and a cross-dimension attention mechanism module at the end of the backbone network is used to perform feature enhancement on the spatial feature map to obtain an enhanced feature map; the enhanced feature map is input into a neck network, a dynamic upsampling operator based on content perception is used to perform multi-scale feature reconstruction on the enhanced feature map, and a multi-scale reconstructed feature map is obtained; the multi-scale reconstructed feature map is input into a detection head, target detection is performed through a multi-scale detection head system formed by an added high-resolution P2 detection layer and an original detection layer, and a detection result of a sea surface ship small target is output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of target detection technology, and in particular relates to a method for detecting small targets on sea surface vessels based on an improved YOLOv11. Background Technology

[0002] With the rapid development of the global marine economy and intelligent shipping, automatic ship target detection technology has become a key sensing foundation for building intelligent maritime regulatory systems and supporting autonomous ship navigation. This technology is of significant practical importance for improving waterway monitoring efficiency, ensuring maritime navigation safety, and promoting the development and utilization of marine resources. Against this backdrop, developing an intelligent method that adapts to the complex and ever-changing marine environment and achieves efficient and reliable detection has become an important research direction in the maritime field.

[0003] Target detection technology, especially single-stage detectors represented by the YOLO series, has made significant progress. From YOLOv1 to the latest YOLOv11, the core improvements revolve around multi-scale feature fusion and efficient network architecture, such as using feature pyramid networks and adaptive anchor boxes to enhance the detection capability of targets at different scales. Meanwhile, the introduction of the Transformer architecture allows the model to enhance its modeling of global contextual information through self-attention mechanisms, providing a new approach to handling complex scenes. However, maritime target detection, as a domain-specific application, faces more severe challenges than in general scenarios. Surface reflections, waves, fog, and variable lighting cause drastic changes in target appearance, generating significant noise interference. Nearby vessels and distant small buoys differ greatly in size, and distant vessels and small floating targets often appear as extremely small pixel areas in images, resulting in a severe lack of feature information for small targets. Targets such as ships are often slender and have variable orientations, placing higher demands on bounding box regression. While single-stage target detectors, represented by YOLOv11, perform excellently in general scenarios, they still face significant technical bottlenecks when handling multi-scale targets at sea.

[0004] To address the aforementioned challenges, this invention proposes a multi-scale target detection method for ships on the sea surface based on an improved YOLOv11. Summary of the Invention

[0005] To address the aforementioned technical problems, this invention proposes a method for detecting small targets on the sea surface based on an improved YOLOv11, thereby resolving the issues present in the prior art.

[0006] To achieve the above objectives, this invention provides a method for detecting small targets on the sea surface based on an improved YOLOv11, comprising: Acquire images of ships on the sea surface, and obtain an input feature map based on the images of ships on the sea surface; The improved YOLOv11 backbone network is used to extract features from the input feature map. Specifically, the C3k2_PConv module, which is constructed based on partial convolution, is used to extract spatial features from the input feature map to obtain a spatial feature map. The cross-dimensional attention mechanism module at the end of the backbone network is used to enhance the spatial feature map to obtain an enhanced feature map. The enhanced feature map is input into the neck network of the improved YOLOv11 network, and a content-aware dynamic upsampling operator is used to perform multi-scale feature reconstruction on the enhanced feature map to obtain a multi-scale reconstructed feature map. The multi-scale reconstructed feature map is input into the detection head of the improved YOLOv11 network. Target detection is performed through a multi-scale detection head system composed of the added high-resolution P2 detection layer and the original detection layer, and the detection results of small targets such as ships on the sea surface are output.

[0007] Optionally, the process of acquiring images of ships on the sea surface and obtaining an input feature map based on the images of ships on the sea surface includes: Based on the collected raw video data of the sea surface, images of ships on the sea surface are extracted frame by frame from the raw video data of the sea surface; The ship targets in the sea surface ship image are labeled with categories and locations using a labeling tool to obtain labeled image data; Based on the labeled image data, the images are divided into a training image set, a validation image set, and a test image set according to a preset ratio. The images in the training image set or the test image set are used as input to obtain the input feature map.

[0008] Optionally, the process of extracting spatial features from the input feature map using the C3k2_PConv module based on partial convolution to obtain the spatial feature map includes: Based on the input feature map, preliminary convolution and normalization processing are performed using CBS units to obtain a preliminary integrated feature map; The preliminary integrated feature map is split along the channel dimension to obtain the first branch feature map and the second branch feature map; Based on the first branch feature map, the identity mapping method is used to directly transmit the feature map, and the original information is preserved in the identity mapping feature map. Based on the second branch feature map, several stacked C3k_PConv units are used sequentially to perform partial convolution and depth feature transformation processing to obtain a depth transformation feature map. The identity mapping feature map and the depth transformation feature map are concatenated along the channel dimension to obtain the concatenated feature map. The concatenated feature map is fused and convolved using CBS units to output a spatial feature map.

[0009] Optionally, the process of using a cross-dimensional attention mechanism module at the end of the backbone network to perform feature enhancement on the spatial feature map to obtain an enhanced feature map includes: Based on the spatial feature map, it is divided equally along the channel dimension to obtain the first attention branch feature map and the second attention branch feature map. Based on the first attention branch feature map, the features are grouped by channel and then multi-scale spatial dependency capture processing is performed using depthwise separable convolution with different kernel sizes to obtain a multi-receptive field feature map. Based on the multi-receptive field feature map, mean pooling is performed in both row and column dimensions to obtain pooled features, which are then processed by Sigmoid gating to generate a spatial attention weight map. At the same time, channel attention processing based on window downsampling is combined to obtain the first enhanced feature map. Based on the second attention branch feature map, the feature group is split into channel sub-branch and spatial sub-branch. The channel sub-branch is processed by global average pooling to obtain the channel weight map, and the spatial sub-branch is processed by group normalization to obtain the spatial weight map. Based on the channel weight map and the spatial weight map, channel-space attention collaborative enhancement processing is performed to obtain the second enhanced feature map. Based on the first enhanced feature map and the second enhanced feature map, a channel global gating mechanism is used to generate dynamic weights for each channel to obtain a weighted fusion feature map; The weighted fused feature map is shuffled through channels to output an enhanced feature map.

[0010] Optionally, the process of inputting the enhanced feature map into the neck network of the improved YOLOv11 network and performing multi-scale feature reconstruction on the enhanced feature map using a content-aware dynamic upsampling operator includes: S1. Obtain the low-resolution feature map to be upsampled in the enhanced feature map as input; S2. A 1×1 convolutional layer is used to perform convolution processing on the low-resolution feature map, and the two-dimensional offset of each spatial location is learned from the convolution result to obtain the dynamic sampling offset. S3. Based on the dynamic sampling offset and the predefined initial sampling grid, perform superposition processing to generate dynamically adjusted sampling coordinates; S4. Based on the sampling coordinates, the low-resolution feature map is resampled using a grid sampling operation to generate a high-resolution output feature map; S5. Repeat S1-S4 to perform dynamic upsampling processing on feature maps of different scales.

[0011] Optionally, the multi-scale feature reconstruction process further includes a cross-level feature fusion mechanism, specifically including: Obtain the high-level semantic features output by the backbone network, and directly inject the high-level semantic features into the corresponding multi-scale fusion node in the neck network, bypassing the intermediate attention layer. The high-level semantic features and the spatial features in the multi-scale fusion nodes are fused using a channel splicing method to obtain cross-level fusion features; Multi-level features output by C3k2_PConv modules at different levels in the neck network are obtained, and the multi-level features are directly concatenated and fused to obtain dense interactive features. Multi-scale feature reconstruction is performed based on the cross-level fusion features and the dense interaction features.

[0012] Optionally, the process of target detection using a multi-scale detection head system composed of an added high-resolution P2 detection layer and the original detection layer includes: A feature map with a downsampling rate of 1 / 4 is extracted from the shallow layer of the backbone network, and a high-resolution detection branch P2 is constructed based on the 1 / 4 feature map; Obtain the P3, P4, and P5 detection branches in the original detection layer with downsampling rates of 1 / 8, 1 / 16, and 1 / 32, respectively. The P2 high-resolution detection branch, together with the P3, P4, and P5 detection branches, constitutes a four-head detection system. The multi-scale reconstructed feature maps are input into the P2 detection branch, P3 detection branch, P4 detection branch and P5 detection branch respectively for target prediction processing, and the detection results of small targets of ships on the sea surface are output.

[0013] The present invention also provides a computer, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described thereon.

[0014] The present invention also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described thereon.

[0015] Compared with the prior art, the present invention has the following advantages and technical effects: First, by replacing the traditional convolutional structure with a C3k2_PConv module based on partial convolution, computational overhead is significantly reduced while maintaining feature extraction capabilities, effectively suppressing redundant feature interference in complex sea surface backgrounds. Second, a cross-dimensional attention mechanism is introduced at the end of the backbone network, which adaptively enhances the feature response of ship targets and suppresses background noise such as waves and reflections, improving the distinction between targets and background. Third, a content-aware dynamic upsampling operator is used instead of a fixed interpolation method, which dynamically adjusts the sampling position according to the image content, accurately restoring ship edge details and avoiding the loss of small target features during the upsampling process. Finally, by adding a high-resolution P2 detection head, the deficiency of deep networks in perceiving small targets is compensated for, achieving full-scale coverage from small ships to large vessels. The synergistic effect of these improvements significantly enhances the detection accuracy and robustness of small targets on the sea surface. Attached Figure Description

[0016] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a flowchart of an embodiment of the present invention; Figure 2 This is a schematic diagram of the CBS structure according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the SPPF structure according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the C2PSA structure according to an embodiment of the present invention; Figure 5 This is a network structure diagram of the small target detection algorithm for ships on the sea surface based on the improved YOLOv11 in an embodiment of the present invention; Figure 6 This is a schematic diagram of the C3k2_PConv structure according to an embodiment of the present invention; Figure 7 This is a schematic diagram of the Dysample structure according to an embodiment of the present invention; Figure 8 This is a schematic diagram of the Shuffle_Attention structure according to an embodiment of the present invention; Figure 9 This is a schematic diagram of the SCSA_Attention structure according to an embodiment of the present invention; Figure 10 This is a schematic diagram of the MSA_Attention structure according to an embodiment of the present invention. Detailed Implementation

[0017] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0018] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0019] Example 1 This invention focuses on the highly challenging area of ​​visual scene understanding and small target detection at sea. Based on the YOLO11 detection framework, which has the advantages of both accuracy and real-time performance, this invention addresses the challenges of high proportion of small targets and difficulty in feature extraction in multi-scale target detection at sea. By improving the C3K2 module and the traditional upsampling module, adding a small target detection head, and proposing a cross-dimensional MSA mechanism, the accuracy of small target ship target detection at sea is improved.

[0020] like Figure 1 As shown, this embodiment provides a method for detecting small targets on the sea surface based on an improved YOLOv11. This method replaces some standard convolutional structures by constructing a lightweight C3k2_PConv module in the backbone network; introduces a content-aware DySample dynamic upsampling operator in the feature fusion network to improve the quality of multi-scale feature reconstruction; adds a dedicated high-resolution P2 detection layer in the detection head to enhance direct perception of small targets; and embeds a cross-dimensional attention mechanism module (MSA) specifically designed for maritime scenes at the end of the backbone network to suppress background interference such as waves and reflections. A cross-level feature fusion mechanism is designed to further improve the model's robustness in representing fine-grained targets against a maritime background. Validation was performed on a self-built dataset of sea surface vessels containing 70% small targets, effectively improving the network's ability to detect multi-scale targets at sea, especially small vessels.

[0021] This invention proposes a small target detection algorithm for ships at sea based on an improved YOLOv11. It addresses the shortcomings of current detection models in terms of accuracy and robustness when dealing with practical challenges such as the high proportion of small targets at sea, strong background interference, and drastic scale changes. The algorithm achieves a performance breakthrough by constructing a dedicated dataset and four core improvements. (1) A lightweight C3k2_PConv module is introduced to optimize the backbone network, and partial convolution is used to extract spatial non-redundant features; (2) The DySample dynamic resampling operator is integrated to accurately restore the edge details of ships by dynamically adjusting the upsampling kernel, thereby improving the reconstruction quality of detailed information during feature pyramid fusion; (3) A dedicated P2 high-resolution small target detection head is added to extract shallower fine-grained feature maps, thereby compensating for the lack of perception of small ship features by deep networks after multiple downsampling, and significantly improving the detection accuracy of the algorithm in extreme multi-scale scenarios; (4) A cross-dimensional marine scene attention mechanism (MSA) is designed to realize cross-dimensional spatial and channel interaction using the MSA module, effectively suppressing sea surface reflection and wave clutter, and enhancing the distinction between targets and background; (5) A cross-level feature fusion mechanism (CDFF) is designed to allow high-order semantic features to retain their original integrity and directly act on the prediction head, thereby improving the discriminative spatial distribution of features. Figure 1 As shown, it includes the following steps: Step 1: Obtain the image dataset from the video of ships on the sea surface, that is, extract the images of ships on the sea surface frame by frame from the original video data of the sea surface. Step 2: Based on the target features, the ship targets in each image are labeled. That is, the labeling tool is used to label the category and location of the ship targets in the sea surface ship images to obtain labeled image data. Based on the labeled image data, the images are divided into training image set, validation image set and test image set according to a preset ratio. The target detection dataset is divided into training set, validation set and test set in a ratio of 7:2:1. The images in the training image set or test image set are used as input to obtain the input feature map. Step 3: Construct an improved YOLOv11 multi-scale target detection model for the sea surface; Step 4: Train the improved YOLOv11 multi-scale ship target detection model built in Step 3; Step 5: Use the improved YOLOv11 multi-scale ship target detection model trained in Step 4 to perform target detection on the test set, and verify it through the target detection evaluation index; Step 6: Write the trained weight parameters into detect.py, use Python to build a framework to run the improved model, and verify that the multi-scale ship target detection model on the sea surface is more accurate than the original model and meets the actual needs.

[0022] As a specific implementation method of this embodiment, the specific implementation process of the small target detection method for ships on the sea surface based on the improved YOLOv11 includes: acquiring images of ships on the sea surface, and obtaining an input feature map based on the images; using the backbone network of the improved YOLOv11 network to extract features from the input feature map, wherein the C3k2_PConv module based on partial convolution is used to extract spatial features from the input feature map to obtain a spatial feature map, and the cross-dimensional attention mechanism module at the end of the backbone network is used to enhance the spatial feature map to obtain an enhanced feature map; inputting the enhanced feature map into the neck network of the improved YOLOv11 network, and using a content-aware dynamic upsampling operator to reconstruct multi-scale features from the enhanced feature map to obtain a multi-scale reconstructed feature map; inputting the multi-scale reconstructed feature map into the detection head of the improved YOLOv11 network, and performing target detection through a multi-scale detection head system composed of an added high-resolution P2 detection layer and the original detection layer, and outputting the detection result of small targets for ships on the sea surface.

[0023] To address the challenges of high proportion of small targets, complex backgrounds, and insufficient detection accuracy of traditional algorithms in surface vessel detection, this invention proposes a small target detection algorithm for surface vessels based on an improved YOLOv11. Building upon YOLOv11, the algorithm undergoes several structural improvements: the C3k2_PConv module is introduced to enhance local feature modeling capabilities, such as... Figure 6 As shown, it reduces computational overhead while suppressing redundant features in complex backgrounds; it incorporates DySample dynamic sampling technology, such as... Figure 7 As shown, the reconstruction quality of small target details is improved during feature pyramid fusion; a new P2 detection head is added to capture high-resolution fine-grained features, compensating for the shortcomings of traditional detection heads in missing small targets; a cross-dimensional maritime scene attention mechanism (MSA) is designed, such as... Figure 10 As shown, with the cross-level feature fusion mechanism (CDFF), such as Figure 5 As shown, this study focuses on the key features of ship targets, suppresses background interference such as waves and clouds, and enhances the relevance of feature representation. The trained network model is validated using a self-made small-target ship image dataset, and evaluation metrics such as Precision, Recall, mAP, and FPS are employed.

[0024] Furthermore, the process of establishing a dataset of small target images of ships on the sea surface includes: This invention constructs a self-made MSDD dataset, specifically designed for the task of detecting ships on the sea surface. The dataset contains 2198 high-quality real-world sea scene images, accurately annotating 6281 ship targets, covering five common ship types: cargo ships, yachts, sailboats, other vessels, and speedboats. The class distribution is balanced, effectively avoiding class bias problems in model training. Combining the characteristics of real-world sea detection scenes, targets with a pixel percentage of less than 0.1% are defined as micro-targets. This type of target accounts for as much as 68.7% of the dataset, focusing on the challenging scenario of detecting small target ships on the sea surface. The dataset features realistic scenes, standardized annotations, and typical target features, and is adapted for training with an improved YOLOv11 network, providing high-quality data support for research on the detection of micro-targets of ships on the sea surface.

[0025] Furthermore, the C3k2_PConv basic module includes: In the improved YOLO11 architecture, the C3k2_PConv module is the core unit for achieving a balance between feature extraction and lightweight design in the backbone and head networks. The original C3k2 structure is reconstructed by introducing the design concept of Partial Convolution (PConv). For example... Figure 6 As shown, this module first performs preliminary integration of input features through a CBS unit to obtain a preliminary integrated feature map. Then, the preliminary integrated feature map is split into two branches along the channel dimension to obtain a first branch feature map and a second branch feature map. The first branch feature map is directly passed to retain the original information to obtain an identity mapping feature map that retains the original information. The second branch feature map is then subjected to deep feature transformation through multiple stacked C3k_PConv units. The deep transformation feature map is obtained by cascading channel splitting and partial convolution, which realizes refined feature extraction. The features of the two branches are concatenated along the channel dimension to obtain a concatenated feature map, and then fused through a CBS unit to obtain a spatial feature map. This allows C3k2_PConv to efficiently learn rich feature representations while controlling computational cost.

[0026] In maritime small target vessel detection tasks, the sea surface environment is complex and the scale of the vessels varies greatly. C3k2_PConv, utilizing its unique channel stripping and feature reconstruction mechanism, can accurately capture the local detailed textures of distant small vessels while maintaining robust representation of the global outline of nearby large ships. Combined with the newly added P2 detection head and DySample dynamic upsampling operator in the network, it ensures the integrity of key feature information during cross-scale transmission, effectively mitigating feature loss of small targets in complex sea conditions (such as wave interference and light refraction), and providing solid modular support for building a high-performance maritime intelligent monitoring system.

[0027] Furthermore, the Dysample basic module includes: the DySample module is a key improvement over the traditional fixed sampling method of YOLO11, such as... Figure 7 As shown, this paper replaces fixed upsampling strategies such as bilinear interpolation with dynamic offset sampling. First, a lightweight 1×1 convolutional layer processes the low-resolution input feature map, learning a two-dimensional offset (Δx, Δy) for each spatial location. Then, these offsets are combined with a predefined initial sampling grid to generate dynamically adjusted sampling coordinates. Finally, through grid sampling operations, the original features are resampled based on these adaptive coordinates to generate a high-resolution output feature map. This breaks the constraint of fixed sampling positions in traditional upsampling operators, enabling the model to dynamically adjust its receptive field according to the content of the sea surface scene.

[0028] When dealing with the task of detecting small targets on the sea surface, traditional upsampling methods often lead to blurred target edges and loss of details, which is a core challenge for detecting tiny targets. However, DySample's dynamic offset mechanism can generate an offset pointing towards the edge in the ship outline area based on the feature content, thereby effectively enhancing edge sharpness and preserving the fine texture information that is crucial for tiny targets. Combined with the newly added P2 detection head in the network, it significantly enhances the network's ability to distinguish and locate small-scale targets. At the same time, the background interference such as waves and reflections that are widespread in the maritime scene have repetitive texture features. Through the semantically aware offset learned by DySample, it can generate a sampling strategy that tends to be averaged in the background area, avoiding the amplification effect of abnormal textures, thereby naturally suppressing the interference of background noise during the feature reconstruction process.

[0029] Furthermore, the addition of a small target detection layer involves the following: In the original three-scale detection head architecture of YOLOv11, the model predicts on feature maps with downsampling rates of 1 / 8, 1 / 16, and 1 / 32, respectively. While these mid-to-deep feature layers can capture semantic information of the target, high-magnification downsampling leads to the loss of pixel-level details of small targets. To address this issue, this invention adds a P2 high-resolution detection head for micro-targets. Its core principle is to extract feature maps with a downsampling rate of 1 / 4 from a shallower layer of the backbone network, constructing a dedicated detection branch that preserves the high-resolution spatial information of the original image, forming a four-head detection system of P2, P3, P4, and P5. This detection head employs a relatively shallower network layer design to prevent the destruction of detailed information by excessive nonlinear transformations. It also features denser, smaller-scale anchor box priors and is matched with a dynamic label allocation strategy optimized for small targets, ensuring that a large number of micro-targets can obtain sufficient positive sample supervision signals.

[0030] In the task of detecting small targets at sea, small targets have a low pixel ratio in the sea scene and are easily masked by background noise such as wave reflection and fog. The P2 high-resolution feature can capture the core pixel features of such targets, solving the problem of loss of small target features in the transmission of deep networks. The detailed features reconstructed by DySample in the feature pyramid fusion process provide high-quality input features for the P2 head, which directly uses these features for prediction, improving the detection accuracy of the model for small targets. P2 and the original detection head form a downsampling level coverage of "4 times to 32 times", which can be adapted to large, medium, small and extremely small sea targets, eliminating the blind spot of the original system for extremely small targets.

[0031] Furthermore, the multi-dimensional maritime scene attention module includes: the multi-dimensional scene attention mechanism (MSA) is a core improvement addressing the insufficient capture of small target features and weak background noise suppression capabilities at sea, such as... Figure 10 As shown, a hybrid architecture of "dual-branch attention + global dynamic gating + channel shuffling" is constructed to achieve refined feature enhancement and efficient fusion. The input spatial feature map is first divided into two branches along the channel dimension to obtain the first attention branch feature map and the second attention branch feature map. The first branch groups the first attention branch feature map by channel and captures the spatial dependence of different receptive fields through depthwise separable convolutions with four different kernel sizes (3, 5, 7, and 9) to obtain a multi-receptive field feature map, which is adapted to the feature perception needs of targets of different scales at sea. The multi-receptive field feature map is then used to generate spatial attention weights through mean pooling and sigmoid gating in the row / column dimension to enhance the spatial position features of the target. Finally, combined with channel attention based on window downsampling, the feature representation of key channels is further improved to obtain the first enhanced feature map.

[0032] The second attention branch feature map of the second branch is obtained by splitting the features into channel sub-branches and spatial sub-branches by grouping. Channel weights are generated by global average pooling and spatial weights are generated by group normalization, respectively. The collaborative enhancement of channel-space attention is achieved in a lightweight way to obtain the second enhanced feature map.

[0033] Meanwhile, the module is designed with a Channel-wise Global Gate (CGG). Based on the global context information of the input features of the first and second enhanced feature maps, a dynamic weight of 0-1 is generated for each channel. The outputs of the two branches are adaptively weighted to obtain a weighted fusion feature map. Finally, the enhanced feature map is output by breaking the feature separation between branches through a channel shuffling operation, which promotes cross-branch information interaction and ensures the integrity of the features.

[0034] In marine environments, wave clutter, sunlight reflection, and sea fog often cause the features of small target vessels to be submerged by the background, resulting in a large number of missed or false detections. The multi-scale convolutional array of MSA can provide a more targeted local receptive field for extremely small vessels; the global dynamic gating mechanism enables MSA to dynamically adjust the attention strategy according to the statistical characteristics of the entire image, finding the optimal feature enhancement scheme for foggy weather, nighttime, or strong reflective conditions, greatly improving the robustness of the model under changing sea conditions.

[0035] In the overall network architecture, the MSA module is strategically placed at the end of the backbone network. It enhances the quality of features before they enter the multi-scale fusion process, effectively preventing background interference from being amplified and propagated in subsequent upsampling and stitching operations. Simultaneously, the features at the end of the backbone network already possess strong semantic information, allowing MSA to more accurately determine which regions are targets and which are background based on these semantic clues, thus enabling more precise attention allocation. This effectively alleviates the recognition dilemma of small targets in deep networks caused by feature sparsity. Through cross-dimensional feature completion, it significantly improves the algorithm's accuracy and environmental robustness in capturing small targets in wide ocean areas. The design of this module effectively alleviates the recognition dilemma of small targets in deep networks caused by feature sparsity. Through cross-dimensional feature completion, it significantly improves the algorithm's accuracy and environmental robustness in capturing small targets in wide ocean areas.

[0036] Furthermore, the cross-level feature fusion mechanism includes: addressing key issues in small target ship detection in maritime scenarios, such as low accuracy in small target recognition and insufficient fusion of high-level semantics and low-level detail features, this invention proposes a cross-level feature fusion mechanism (Cross-Dimensional Feature Fusion, CDFF) based on the YOLO11 neck network architecture, such as... Figure 5 As shown, this method overcomes the limitations of feature propagation and fusion in the original network, providing a more efficient feature representation scheme for small target detection.

[0037] The original YOLO11 neck network uses a traditional FPN-PAN structure. High-level semantic features enhanced by SMA (MaritimeScene Attention) and C2PSA attention modules at the backbone network ends can only be passed layer by layer to the fusion nodes in the neck via a single top-down linear path. This results in the gradual dilution of high-level semantic information during multiple upsampling and convolution operations, and a lack of direct interaction with the multi-scale spatial features of the neck, making it difficult to effectively balance the capture of fine-grained details of small-scale ships with global semantic discrimination capabilities. To address these bottlenecks, the CDFF mechanism achieves a significant improvement in feature fusion performance through two core innovative designs, such as... Figure 5As shown, firstly, a direct backbone-neck connection is constructed, breaking the original linear propagation constraint. This allows high-level semantic features output from the backbone network to bypass the intermediate attention layer and directly inject into the corresponding multi-scale fusion nodes in the neck. Preliminary fusion of semantic and spatial information is achieved through concat concatenation, effectively mitigating information loss. Secondly, a cross-layer feature interaction path is introduced within the neck, allowing direct concatenation and fusion of features output from C3k2_PConv modules at different levels in the neck network. This constructs a dense feature interaction network, further strengthening the deep coupling between high-level semantic features and low-level detailed features, bridging the representation gap between them. This design not only improves the efficiency and richness of feature fusion but also enhances the model's ability to locate and identify small-scale ships in maritime scenes. Simultaneously, it optimizes feature propagation efficiency and reduces information redundancy. Compared to the original YOLO11 neck network architecture, it achieves significant improvements in detection accuracy and robustness.

[0038] Furthermore, the small target detection algorithm for ships on the sea surface includes: This invention proposes a small target detection algorithm for ships on the sea surface based on an improved YOLOv11, the overall network architecture of which is as follows: Figure 5 As shown, it adopts the efficient single-stage detection framework of YOLOv11, which consists of two main parts: the backbone network and the head network. It constructs a small target detection model for ships in the marine environment to address the core challenges of large differences in ship target size, high proportion of small targets and strong background interference in the marine scene.

[0039] The backbone network uses CBS basic modules as its basic units, such as Figure 2 As shown, the network consists of a cascaded convolutional layer (Conv), a batch normalization layer (BN), and a SiLU activation function, responsible for extracting local spatial features and performing nonlinear transformations. Based on this, the network uses a C2PSA module, such as... Figure 4 As shown, the integration of convolution and self-attention mechanisms enables the modeling of long-range dependencies while capturing local features, enhancing the understanding of complex sea surface contexts. To improve feature extraction efficiency and reduce computational overhead, we introduce a lightweight C3k2_PConv module to replace part of the original C2k3 structure, such as... Figure 6 As shown, this module employs a partial convolution mechanism, performing 3×3 convolution on only one-quarter of the channels while maintaining the identity mapping for the remaining channels. This significantly reduces the number of parameters and computational cost while preserving the integrity of the original features to the greatest extent possible. The backbone network ends with an SPPF module, such as... Figure 3 As shown, multi-receptive-field contextual information aggregation is achieved through three cascaded max-pooling layers, enhancing the model's adaptability to scale changes. Simultaneously, a cross-dimensional maritime scene attention mechanism, MSA, is designed at the end of the backbone network, such as... Figure 10As shown, this module utilizes a dual-branch heterogeneous attention mechanism, ShuffleAttention, as... Figure 8 As shown, with SCSA, as Figure 9 As shown, the collaborative design of global dynamic gating effectively suppresses background noise such as waves and reflections, and enhances the feature response of small targets, providing cleaner and more focused feature input for subsequent feature fusion.

[0040] In the feature fusion process, the head network introduces the DySample dynamic upsampling operator to replace the traditional interpolation method, such as... Figure 7 As shown, sampling offsets are dynamically generated based on the local semantic content of the feature maps, achieving content-aware adaptive upsampling. Building upon the existing P3, P4, and P5 detection heads, a new high-resolution P2 detection head is added, targeting micro-targets. This branch directly utilizes feature maps with a shallow 1 / 4 downsampling rate from the backbone network to achieve direct perception and precise localization of pixel-level micro-ships. The entire head network, through multiple C3k2_PConv feature transformations, DySample upsampling, and cross-layer feature fusion, ultimately forms a four-head detection system from P2 to P5, achieving full-scale coverage from micro-to-large ships.

[0041] As a specific implementation of this embodiment, the process of inputting the enhanced feature map into the neck network of the improved YOLOv11 network and using a content-aware dynamic upsampling operator to reconstruct the enhanced feature map at multiple scales includes: S1, obtaining the low-resolution feature map to be upsampled in the enhanced feature map as input; S2, performing convolution processing on the low-resolution feature map using a 1×1 convolution layer, learning the two-dimensional offset of each spatial location from the convolution result to obtain the dynamic sampling offset; S3, performing superposition processing based on the dynamic sampling offset and a predefined initial sampling grid to generate dynamically adjusted sampling coordinates; S4, performing resampling processing on the low-resolution feature map using a grid sampling operation based on the sampling coordinates to generate a high-resolution output feature map; S5, repeating S1-S4 to perform dynamic upsampling processing on feature maps at different scales.

[0042] Furthermore, the algorithm file and parameter configuration process for small target detection of ships on the sea surface based on the improved YOLOv11 include: This invention addresses the challenges of multi-scale perception and small target recognition in the marine environment. It constructs a dedicated dataset, improves the YOLOv11 algorithm, refines the feature extraction module, and designs an attention mechanism. Experiments verify the effectiveness of the proposed algorithm. Specific experimental and verification steps are as follows: Step 1: Configure the network training configuration file. First, preprocess and configure the network using the self-built MSDD dataset. The MSDD dataset contains 2198 high-quality real-world sea surface images, accurately annotating 6281 ship targets, covering five common ship types: cargo ship, yacht, sailboat, other ship, and speedboat, with a balanced class distribution. Based on the characteristics of actual maritime scenes, targets with a pixel percentage less than 0.1% are defined as micro-targets; micro-targets account for as much as 68.7% of the dataset.

[0043] In the YOLOv11 data configuration file, add ['Cargo ship', 'Yacht', 'Sailboat', 'Other ship', 'Speedboat'] sequentially to the category list (names). Simultaneously, add the implementation code for custom modules such as C3k2_PConv, DySample, and MSA to Addmodule and complete their registration in __init__.py. When parsing the configuration file in tasks.py, ensure that the newly added modules can be correctly called. In the model configuration file, adjust the network layer parameters accordingly based on the improved network structure (introducing C3k2_PConv, DySample, P2 detector head, and MSA module), and add output channels and anchor point settings for the P2 detector layer. Run the model building script to verify that the network structure modifications are correct.

[0044] Step 2: Configure the network training environment. This experimental environment uses the Sugon cloud computing service system, equipped with 2×HygonC86 7185 32-core CPUs, 256GB of memory, and 2×NVIDIA V100 GPUs, and features a 100GB IB high-speed computing network, belonging to the xhhgnormal02 computing queue. The model is built on the deep learning framework PyTorch, using Python version 3.8.0. The specific hardware configuration for the network experiment is shown in Table 1. The software environment configuration for the network experiment is shown in Table 2. The parameter settings for the improved YOLOv11 model are shown in Table 3.

[0045] Table 1 Table 2 Table 3 Furthermore, the validation based on the target detection evaluation indicators includes: Step 1: Selection of Target Detection Evaluation Metrics. To objectively evaluate the improved YOLO11 model's detection performance in the marine environment, this study uniformly used the constructed Small Target Detection Dataset (MSDD) for model training, validation, and performance testing. In the experiments, precision (P), recall (R), mean average precision (mAP), and frames per second (FPS) were selected as core quantitative evaluation metrics. During target matching, the Intersection over Union (IoU) threshold was set to 0.5. mAP, as a global metric measuring the model's comprehensive ability to locate and classify targets, is defined as the arithmetic mean of the average precision (AP) for each class, i.e., the average area enclosed by the precision curves at different recall levels. The mathematical definitions of the relevant metrics are as follows: In the formula, TP represents the number of correctly detected positive samples, FP represents the number of incorrectly detected positive samples, FN represents the number of missed positive samples, and m is the total number of target categories identified.

[0046] Step 2: Experimental Results Analysis. To verify the effectiveness of the improved target detection algorithm proposed in this invention, the experimental results are shown in Table 4.

[0047] Table 4 All models were trained for 400 epochs using 640×640×3 resolution images as input under a unified training strategy. Experiments compared YOLOv8, YOLOv9, YOLOv10, YOLOv11, YOLOv12, and the improved YOLOv11 method for detecting small targets on the sea surface proposed in this invention (Ours). Experimental results show that the proposed method for detecting small targets on the sea surface achieves a precision (P) of 95%, a 5.9% improvement over YOLOv11, while simultaneously increasing the recall (R) to 85.9%, a 7.5% improvement over YOLOv11, thus achieving high reliability in sea surface target detection.

[0048] mAP of the algorithm of this invention 50The accuracy reached 92.5%, representing improvements of 5.3, 6.1, 7.1, 6.6, and 6.7 percentage points compared to YOLOv8, YOLOv9, YOLOv10, YOLOv11, and YOLOv12 models, respectively. Simultaneously, the mAP50-95 reached 64%, a 3.9% improvement over YOLOv11's 60.1%, directly validating the proposed multi-scale feature enhancement and structure optimization strategy. This strategy effectively mines the deep features of surface vessel targets, addressing the core issue of insufficient feature extraction for small surface targets in traditional YOLOv11. In summary, the improved method proposed in this invention demonstrates outstanding detection accuracy, providing a superior technical solution for the precise detection of small surface vessel targets in the marine environment.

[0049] This invention is based on a self-built ship detection dataset containing over 2000 real marine images, with small targets accounting for over 70%. Using the YOLOv11 target detection framework as a foundation, a cross-dimensional marine scene attention mechanism (MSA) is designed. The C3k2_PConv structure and MSA are integrated into the backbone network, and the DySample dynamic resampling operator is introduced into the neck network. A P2 fine-grained small target detection layer is added at the detection end, constructing a method for detecting small targets on the sea surface based on an improved YOLOv11. Its advantages are: (1) It solves the problem of lacking a targeted, publicly available benchmark dataset with a large proportion of small targets in marine ship detection research; (2) It effectively alleviates the problem of missed and false detections of small ship targets in marine images due to weak features and strong background interference; (3) While improving accuracy, it ensures the practical efficiency of the algorithm and maintains a low parameter count to meet the lightweight requirements of maritime monitoring.

[0050] This invention is based on the YOLOv11 target detection framework. It designs a lightweight C3k2_PConv module to optimize the feature extraction efficiency of the backbone network, introduces the DySample dynamic upsampling operator in the feature fusion stage to enhance the multi-scale information reconstruction capability, adds a P2 high-resolution detection layer dedicated to small targets in the detection head, and embeds a cross-dimensional maritime scene attention mechanism (MSA) at the end of the backbone network to construct a small target detection algorithm for ships on the sea surface based on the improved YOLOv11.

[0051] The present invention also provides a computer, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described thereon.

[0052] The present invention also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described thereon.

[0053] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for detecting small targets on the sea surface of ships based on an improved YOLOv11, characterized in that, Includes the following steps: Acquire images of ships on the sea surface, and obtain an input feature map based on the images of ships on the sea surface; The improved YOLOv11 backbone network is used to extract features from the input feature map. Specifically, the C3k2_PConv module, which is constructed based on partial convolution, is used to extract spatial features from the input feature map to obtain a spatial feature map. The cross-dimensional attention mechanism module at the end of the backbone network is used to enhance the spatial feature map to obtain an enhanced feature map. The enhanced feature map is input into the neck network of the improved YOLOv11 network, and a content-aware dynamic upsampling operator is used to perform multi-scale feature reconstruction on the enhanced feature map to obtain a multi-scale reconstructed feature map. The multi-scale reconstructed feature map is input into the detection head of the improved YOLOv11 network. Target detection is performed through a multi-scale detection head system composed of the added high-resolution P2 detection layer and the original detection layer, and the detection results of small targets such as ships on the sea surface are output.

2. The method for detecting small targets on the sea surface based on the improved YOLOv11 according to claim 1, characterized in that, The process of acquiring images of ships on the sea surface and obtaining input feature maps based on these images includes: Based on the collected raw video data of the sea surface, images of ships on the sea surface are extracted frame by frame from the raw video data of the sea surface; The ship targets in the sea surface ship image are labeled with categories and locations using a labeling tool to obtain labeled image data; Based on the labeled image data, the images are divided into a training image set, a validation image set, and a test image set according to a preset ratio. The images in the training image set or the test image set are used as input to obtain the input feature map.

3. The method for detecting small targets on the sea surface based on improved YOLOv11 according to claim 1, characterized in that, The process of extracting spatial features from the input feature map using the C3k2_PConv module, which is based on partial convolution, includes: Based on the input feature map, preliminary convolution and normalization processing are performed using CBS units to obtain a preliminary integrated feature map; The preliminary integrated feature map is split along the channel dimension to obtain the first branch feature map and the second branch feature map; Based on the first branch feature map, the identity mapping method is used to directly transmit the feature map, and the original information is preserved in the identity mapping feature map. Based on the second branch feature map, several stacked C3k_PConv units are used sequentially to perform partial convolution and depth feature transformation processing to obtain a depth transformation feature map. The identity mapping feature map and the depth transformation feature map are concatenated along the channel dimension to obtain the concatenated feature map. The concatenated feature map is fused and convolved using CBS units to output a spatial feature map.

4. The method for detecting small targets on the sea surface based on the improved YOLOv11 according to claim 3, characterized in that, The process of using a cross-dimensional attention mechanism module at the end of the backbone network to perform feature enhancement on the spatial feature map to obtain an enhanced feature map includes: Based on the spatial feature map, it is divided equally along the channel dimension to obtain the first attention branch feature map and the second attention branch feature map. Based on the first attention branch feature map, the features are grouped by channel and then multi-scale spatial dependency capture processing is performed using depthwise separable convolution with different kernel sizes to obtain a multi-receptive field feature map. Based on the multi-receptive field feature map, mean pooling is performed in both row and column dimensions to obtain pooled features, which are then processed by Sigmoid gating to generate a spatial attention weight map. At the same time, channel attention processing based on window downsampling is combined to obtain the first enhanced feature map. Based on the second attention branch feature map, the feature group is split into channel sub-branch and spatial sub-branch. The channel sub-branch is processed by global average pooling to obtain the channel weight map, and the spatial sub-branch is processed by group normalization to obtain the spatial weight map. Based on the channel weight map and the spatial weight map, channel-space attention collaborative enhancement processing is performed to obtain the second enhanced feature map. Based on the first enhanced feature map and the second enhanced feature map, a channel global gating mechanism is used to generate dynamic weights for each channel to obtain a weighted fusion feature map; The weighted fused feature map is shuffled through channels to output an enhanced feature map.

5. The method for detecting small targets on the sea surface based on the improved YOLOv11 according to claim 1, characterized in that, The process of inputting the enhanced feature map into the neck network of the improved YOLOv11 network and performing multi-scale feature reconstruction on the enhanced feature map using a content-aware dynamic upsampling operator includes: S1. Obtain the low-resolution feature map to be upsampled in the enhanced feature map as input; S2. A 1×1 convolutional layer is used to perform convolution processing on the low-resolution feature map, and the two-dimensional offset of each spatial location is learned from the convolution result to obtain the dynamic sampling offset. S3. Based on the dynamic sampling offset and the predefined initial sampling grid, perform superposition processing to generate dynamically adjusted sampling coordinates; S4. Based on the sampling coordinates, the low-resolution feature map is resampled using a grid sampling operation to generate a high-resolution output feature map; S5. Repeat S1-S4 to perform dynamic upsampling processing on feature maps of different scales.

6. The method for detecting small targets on the sea surface based on improved YOLOv11 according to claim 1, characterized in that, The multi-scale feature reconstruction process also includes a cross-level feature fusion mechanism, specifically including: Obtain the high-level semantic features output by the backbone network, and directly inject the high-level semantic features into the corresponding multi-scale fusion node in the neck network, bypassing the intermediate attention layer. The high-level semantic features and the spatial features in the multi-scale fusion nodes are fused using a channel splicing method to obtain cross-level fusion features; Multi-level features output by C3k2_PConv modules at different levels in the neck network are obtained, and the multi-level features are directly concatenated and fused to obtain dense interactive features. Multi-scale feature reconstruction is performed based on the cross-level fusion features and the dense interaction features.

7. The method for detecting small targets on the sea surface based on improved YOLOv11 according to claim 1, characterized in that, The process of target detection using a multi-scale detection head system consisting of an added high-resolution P2 detection layer and the original detection layer includes: A feature map with a downsampling rate of 1 / 4 is extracted from the shallow layer of the backbone network, and a high-resolution detection branch P2 is constructed based on the 1 / 4 feature map; Obtain the P3, P4, and P5 detection branches in the original detection layer with downsampling rates of 1 / 8, 1 / 16, and 1 / 32, respectively. The P2 high-resolution detection branch, together with the P3, P4, and P5 detection branches, constitutes a four-head detection system. The multi-scale reconstructed feature maps are input into the P2 detection branch, P3 detection branch, P4 detection branch and P5 detection branch respectively for target prediction processing, and the detection results of small targets of ships on the sea surface are output.

8. A computer comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in claim 1.

9. A storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in claim 1.