Sea target instance segmentation method based on dual attention mechanism and multi-scale feature fusion

Through the feature fusion of the Swin Transformer backbone network and the dual-attention mechanism, combined with bottom-up path aggregation and dual-branch feature extraction of the SOLOv2 architecture, the accuracy and real-time problems of marine target instance segmentation in complex marine environments are solved, and efficient target segmentation effects are achieved.

CN120599262APending Publication Date: 2025-09-05DALIAN MARITIME UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510739104.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing marine target instance segmentation methods have low instance segmentation accuracy in complex marine environments, especially under conditions such as strong light and fog, where it is difficult to effectively distinguish between targets and backgrounds. Traditional network models also find it difficult to balance detailed information and semantic information, resulting in low instance segmentation accuracy and making it difficult to meet the real-time processing needs of marine intelligent equipment.

Method used

A multi-scale feature extraction module based on the Swin Transformer backbone network is adopted, combined with a feature enhancement module with channel and spatial dual attention mechanisms, a bottom-up path aggregation module is designed, and a dual-branch feature extraction module based on the SOLOv2 architecture is used to achieve maritime target instance segmentation.

Benefits of technology

It significantly improves the segmentation consistency of target boundaries, enhances the instance differentiation ability of dense small targets, reduces the missed detection rate, and meets the needs of marine intelligent equipment for real-time target perception.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599262A_ABST
    Figure CN120599262A_ABST
Patent Text Reader

Abstract

The invention discloses a sea target instance segmentation method based on a double attention mechanism and multi-scale feature fusion. The method comprises the following steps: S1, obtaining a sea target instance segmentation data set; s2, establishing a sea target instance segmentation model, and training the sea target instance segmentation model; the marine target instance segmentation model comprises a multi-scale feature extraction module; a feature enhancement module; a bottom-up path aggregation module; a double-branch feature extraction module established based on the SOLOv2 architecture and used for extracting category branch features and mask features and obtaining an instance segmentation result based on the category branch features and the mask features; and S3, obtaining an instance segmentation result based on the trained marine target instance segmentation model. According to the method, a marine target instance segmentation model is designed for multiple challenges (sea wave reflection interference, multi-scale target coexistence and easy loss of small target details) of target segmentation in a marine environment, and marine scene characteristics are fully adapted through collaborative design among the modules.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a method for marine target instance segmentation based on a dual attention mechanism and multi-scale feature fusion. Background Art

[0002] As a key technology in the field of computer vision, maritime target instance segmentation can perform pixel-level recognition and individual differentiation of targets such as ships and buoys in complex marine environments. It has important application value in fields such as marine resource exploration and ship intelligent navigation.

[0003] Currently, deep learning-based instance segmentation methods are mainly divided into two-stage and single-stage methods. Two-stage methods, such as Mask R-CNN, although they have high detection accuracy, are computationally complex, slow inference speed, and rely on high-precision candidate region screening, making it difficult to meet real-time processing requirements. Single-stage methods, such as SOLOv2, use dynamic convolution kernels to generate masks. This method divides the input image into a uniform grid, with each grid responsible for generating an instance mask for the corresponding position, avoiding redundant calculations of region proposals and significantly improving inference efficiency. Therefore, single-stage methods perform relatively well in conventional scenarios. However, in dense ocean multi-target scenarios, the scales of adjacent ship targets are similar, and this method is prone to mask overlap, which in turn leads to failure in instance differentiation. At the same time, existing instance segmentation methods typically use ResNet as the backbone network. However, in ocean scenes, due to the complex wave texture, light reflection, and background diversity, ResNet has difficulty effectively distinguishing between targets and backgrounds, and has not optimized the feature extraction method and attention mechanism accordingly for the ocean environment. For example, the traditional spatial attention module has difficulty effectively suppressing the high-frequency interference of wave texture, resulting in the existing model's insufficient ability to focus on the edge of the hull; the traditional channel attention mechanism has limited effect on enhancing the metal reflective features of ships under complex lighting conditions, resulting in blurred target boundaries under strong light conditions caused by wave reflections, and easy omission of small targets in foggy scenes. In addition, ocean targets exhibit significant multi-scale characteristics due to differences in shooting distance. Traditional network models find it difficult to balance detailed information and semantic information during the feature fusion process, resulting in low instance segmentation accuracy and difficulty meeting the perception needs of marine intelligent equipment for complex scenes. Summary of the Invention

[0004] The present invention provides a method for instance segmentation of marine targets based on a dual attention mechanism and multi-scale feature fusion, so as to overcome the technical problem of low instance segmentation accuracy in complex scenes such as strong light and foggy weather in the existing technology.

[0005] In order to achieve the above object, the technical solution of the present invention is:

[0006] A method for marine target instance segmentation based on dual attention mechanism and multi-scale feature fusion, the specific steps include:

[0007] S1: Obtain the marine object instance segmentation dataset;

[0008] S2: establishing a maritime object instance segmentation model, and training the maritime object instance segmentation model based on the maritime object instance segmentation dataset to obtain a trained maritime object instance segmentation model;

[0009] The maritime target instance segmentation model includes:

[0010] A multi-scale feature extraction module based on the Swin Transformer backbone network is used to extract target features of different scales in the marine target instance segmentation dataset;

[0011] A feature enhancement module based on the channel and spatial dual attention mechanism is used to enhance the target features output by the multi-scale feature extraction module;

[0012] The bottom-up path aggregation module is used to perform cross-level feature fusion processing on the features output by the feature enhancement module;

[0013] A dual-branch feature extraction module based on the SOLOv2 architecture is used to extract the category branch features and mask features from the features output by the bottom-up path aggregation module, and obtain instance segmentation results based on the category branch features and mask features;

[0014] S3: Performing actual marine target instance segmentation based on the trained marine target instance segmentation model to obtain an instance segmentation result.

[0015] Furthermore, the multi-scale feature extraction module established based on the Swin Transformer backbone network includes: a first feature extraction module, a second feature extraction module, a third feature extraction module, a fourth feature extraction module, a first feature integration module, a second feature integration module, a third feature integration module and a fourth feature integration module;

[0016] The first feature extraction module is used to extract visual features of the marine target instance segmentation dataset and transmit them to the second feature extraction module and the first feature integration module respectively;

[0017] The second feature extraction module is used to extract the first-level semantic features of the visual features and transmit them to the third feature extraction module and the second feature integration module respectively;

[0018] The third feature extraction module is used to extract the second level semantic features of the first level semantic features and transmit them to the fourth feature extraction module and the third feature integration module respectively;

[0019] The fourth feature extraction module is used to extract the global semantic features of the second-level semantic features and transmit them to the fourth feature integration module;

[0020] The fourth feature integration module is used to perform dimension adjustment processing on the global semantic feature to obtain the first initial scale feature and transmit it to the third feature integration module;

[0021] The third feature integration module is used to perform dimension adjustment processing on the second layer semantic features to obtain a second initial scale feature, and transmit it to the second feature integration module, and at the same time add the first initial scale feature and the second initial scale feature to obtain a first added scale feature;

[0022] The second feature integration module is used to perform dimension adjustment processing on the first-level semantic features to obtain a third initial scale feature, and transmit it to the fourth feature integration module, and at the same time add the second initial scale feature and the third initial scale feature to obtain a second added scale feature;

[0023] The first feature integration module is configured to perform dimension adjustment processing on the visual feature to obtain a fourth initial scale feature, and add the third initial scale feature and the fourth initial scale feature to obtain a third added scale feature.

[0024] Furthermore, each feature extraction module includes at least two consecutive first Swin Transformer blocks of a multi-head self-attention module W-MSA with a regular window and a second Swin Transformer block of a multi-head self-attention module SW-MSA with a shifted window;

[0025] The calculation formula for two consecutive Swin Transformer blocks is:

[0026]

[0027] Where z l-1 is the input feature of the first Swin Transformer block; is the output feature of the multi-head self-attention module with a regular window; z l is the output feature of the first Swin Transformer block; is the output feature of the multi-head self-attention module with a shift window; z l+1 is the output feature of the second Swin Transformer block; MLP represents the processing process of the multi-layer perceptron module; LN is the layer normalization processing process.

[0028] Furthermore, each feature integration module includes a 1×1 convolution layer, a 2x upsampling layer, and a 3×3 convolution layer;

[0029] The 1×1 convolutional layer is used to adjust the number of channels of the input features;

[0030] The 2x upsampling layer is used to perform a 2x upsampling operation on the features after the number of channels is adjusted, so as to improve the spatial resolution of the features after the number of channels is adjusted;

[0031] The 3×3 convolutional layer is used to adjust the number of channels of the features processed by improving the spatial resolution.

[0032] Furthermore, the feature enhancement module established based on the channel and spatial dual attention mechanism includes a first enhancement module, a second enhancement module, a third enhancement module and a fourth enhancement module;

[0033] The first enhancement module is used to enhance the third added scale feature to obtain a first enhanced feature;

[0034] The second enhancement module is used to perform enhancement processing on the second added scale feature to obtain a second enhanced feature;

[0035] The third enhancement module is used to enhance the first added scale feature to obtain a third enhanced feature;

[0036] The fourth enhancement module is used to enhance the first initial scale feature to obtain a fourth enhanced feature;

[0037] Each enhancement module includes: a channel attention module and a spatial attention module;

[0038] The calculation formula of the channel attention module is:

[0039] M c (F)=σ(MLP(AvgPool(F))+MLP(MaxPool(F))) (2)

[0040] Where F represents the input feature map, σ represents the Sigmoid activation function; AvgPool represents the average pooling operation; MaxPool represents the maximum pooling operation; MLP represents the processing process of the multi-layer perceptron module;

[0041] The calculation formula of the spatial attention module is:

[0042] M s (F)=σ(f 7×7 ([AvgPool(F);MaxPool(F)])) (3)

[0043] Among them, f 7×7 Indicates the use of a convolution operation with a convolution kernel size of 7×7;

[0044] The enhanced features are obtained based on the output of the channel attention module and the spatial attention module, which are expressed as:

[0045]

[0046] in, Represents an element-wise multiplication operation.

[0047] Furthermore, the calculation formula of the bottom-up path aggregation module includes:

[0048] P2=φ(M2) (5)

[0049]

[0050] P6=MaxPool(P5) (7)

[0051] Wherein, M2 represents the first enhanced feature, P2 represents the first fused feature obtained by the bottom-up path aggregation module fusing the first enhanced feature; i=3,4,5 Represent the second enhancement feature, the third enhancement feature and the fourth enhancement feature respectively; P i=3,4,5 represents the second fused feature, the third fused feature, and the fourth fused feature obtained by the bottom-up path aggregation module by performing feature fusion based on the second enhanced feature, the third enhanced feature, and the fourth enhanced feature respectively; Indicates the downsampling operation on the fusion feature of the i-1 layer; Concat(·) indicates the splicing operation; φ(·) indicates the fusion operation, involving a 1×1 convolution layer and a 3×3 convolution layer; P6 is the fifth fusion feature obtained by performing maximum pooling on the fourth fusion feature by the bottom-up path aggregation module.

[0052] Furthermore, the dual-branch feature extraction module established based on the SOLOv2 architecture includes: a category branch and a mask branch;

[0053] The category branches are used to extract category branch features from the first to fifth fusion features, respectively, and are expressed as:

[0054] C b =CategoryHead(P b ) (8)

[0055] Among them, P b is the b-th fusion feature output by the bottom-up path aggregation module, b = 2, 3, 4, 5, 6; CategoryHead(·) represents the category branch convolution head, C b The category branch feature of each spatial position in the fusion feature output by the category branch, wherein the category branch feature is the category probability;

[0056] The mask branch is used to extract the mask features from the first to fifth fusion features respectively, which are expressed as:

[0057] M b =MaskHead(P b ) (9)

[0058] Among them, MaskHead(·) represents the mask branch convolution head, M b The mask feature map of the fused features output by the mask branch.

[0059] Furthermore, the step of obtaining an instance segmentation result based on the category branch feature and the mask feature includes:

[0060] The predicted category score of each spatial location point in the fusion feature is calculated based on the category branch feature Expressed as:

[0061]

[0062] Among them, Conv 3×3 represents a 3×3 convolution operation, and σ represents a Sigmoid activation function;

[0063] The mask features are associated and calculated through dynamic convolution operation to obtain the candidate instance mask Expressed as:

[0064]

[0065] Among them, DynamicConv(·) represents the operation of dynamically generating masks, θ (i,j) is the dynamic convolution parameter corresponding to the i', j'th position in the mask feature;

[0066] According to the predicted category score of each spatial location point in the fusion feature Filter out the spatial location points whose predicted category scores meet the set threshold;

[0067] A matrix non-maximum suppression method is used to perform parallel operations on the spatial location points and candidate instance masks whose predicted category scores meet the set threshold, and the suppression probability corresponding to the candidate instance mask is calculated;

[0068] Multiplying the predicted category score of the spatial position point that meets the set threshold by the corresponding suppression probability to obtain the confidence score of the corresponding candidate instance mask. The confidence score is used to measure the credibility of the candidate instance mask as the final instance;

[0069] All candidate instance masks are sorted in descending order according to the confidence score, and a confidence threshold is set. Only candidate instance masks with a confidence score higher than the confidence score threshold are retained to generate a final instance set, which is the instance segmentation result of the marine target.

[0070] Beneficial effects: The present invention establishes a marine target instance segmentation model, extracts target features of different scales in the marine target instance segmentation dataset through a multi-scale feature extraction module based on the Swin Transformer backbone network, and combines it with a feature enhancement module based on the channel and spatial dual attention mechanism to enhance the feature distinction between the target and the complex background, strengthen the semantic information of key areas such as the edge of the hull, and significantly improve the segmentation consistency of the target boundary. It solves the boundary fuzziness problem of traditional methods in strong light and foggy scenes, and at the same time improves the instance distinction ability of dense small targets, effectively reduces the missed detection rate, and is particularly suitable for complex marine scenes with multiple overlapping targets. A bottom-up path aggregation module is designed, which retains the detailed information of targets such as buoys and small ships through the cross-level fusion of low-level high-resolution features and high-level semantic features. A dual-branch feature extraction module based on the SOLOv2 architecture is adopted to significantly improve the reasoning efficiency while ensuring the segmentation accuracy. The present invention meets the needs of marine intelligent equipment for real-time target perception, thereby providing a technical basis for tasks such as marine resource exploration activities and intelligent navigation systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0072] Figure 1 This is a flowchart of a method for marine target instance segmentation based on dual attention mechanism and multi-scale feature fusion;

[0073] Figure 2 1 is a diagram showing the overall architecture of a maritime target instance segmentation model according to an embodiment of the present invention;

[0074] Figure 3 Schematic diagram of the structure of a multi-scale feature extraction module in an embodiment of the present invention;

[0075] Figure 4 Schematic diagram of the structure of two consecutive Swin Transformer blocks in an embodiment of the present invention;

[0076] Figure 5Schematic diagram of the structure of the channel attention module and the spatial attention module in an embodiment of the present invention;

[0077] Figure 6 This is a specific workflow diagram of BPAM in an embodiment of the present invention;

[0078] Figure 7 This is a visualization result diagram of multi-category ship segmentation in an embodiment of the present invention;

[0079] Figure 8 and Figure 9 This is a comparison chart of multiple algorithm instance segmentation in an embodiment of the present invention. DETAILED DESCRIPTION

[0080] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0081] This embodiment provides a method for marine target instance segmentation based on dual attention mechanism and multi-scale feature fusion. Figure 1 As shown, the specific steps include:

[0082] S1: Obtain the marine object instance segmentation dataset;

[0083] Specifically, the maritime object instance segmentation dataset described in this embodiment contains 7,938 high-resolution images covering 12 typical maritime target categories, including sailboats, fishing boats, and lighthouses. The data sources are a combination of publicly available internet images, maritime datasets such as SMD, samples filtered from general datasets, and field footage. The dataset covers a wide range of sea conditions and lighting conditions, including sunny, foggy, and rainy days. In constructing the maritime object instance segmentation dataset, image quality was ensured through image processing steps such as removing damaged / duplicate images and uniform perspective filtering. The processed images were annotated with polygonal masks using the Labelme tool and cross-verified by professionals. This provides a diverse marine scene training benchmark for the maritime object instance segmentation model, further enhancing its robustness in real-world marine environments.

[0084] S2: Establish Figure 2 The maritime object instance segmentation model shown is trained based on the maritime object instance segmentation dataset to obtain a trained maritime object instance segmentation model;

[0085] The maritime target instance segmentation model includes:

[0086] A multi-scale feature extraction module based on the Swin Transformer backbone network is used to extract target features of different scales in the marine target instance segmentation dataset;

[0087] Specifically, if Figure 3 As shown in the figure, in this embodiment, Swin Transformer is used as the backbone network. In the task of marine target instance segmentation, Swin Transformer shows significant advantages. Its self-attention mechanism can effectively distinguish the wave texture and target contour in the complex ocean background, improve the segmentation accuracy of the marine target instance segmentation model, and process multi-scale features through a hierarchical structure, thereby effectively reducing the computational complexity while maintaining accuracy.

[0088] Specifically, in this embodiment, the multi-scale feature extraction module outputs feature maps of four different scales.

[0089] A feature enhancement module based on the channel and spatial dual attention mechanism is used to enhance the target features output by the multi-scale feature extraction module;

[0090] The bottom-up path aggregation module (BPAM) is used to perform cross-level feature fusion processing on the features output by the feature enhancement module, realize the interaction between low-level visual features and high-level semantic features, and improve the fusion effect of features at different levels;

[0091] A dual-branch feature extraction module based on the SOLOv2 architecture is used to extract the category branch features and mask features from the features output by the bottom-up path aggregation module, and obtain instance segmentation results based on the category branch features and mask features;

[0092] S3: Performing actual marine target instance segmentation based on the trained marine target instance segmentation model to obtain an instance segmentation result.

[0093] In a specific embodiment, the multi-scale feature extraction module established based on the Swin Transformer backbone network includes: a first feature extraction module, a second feature extraction module, a third feature extraction module, a fourth feature extraction module, a first feature integration module, a second feature integration module, a third feature integration module and a fourth feature integration module;

[0094] The first feature extraction module is used to extract low-level visual features of the marine target instance segmentation dataset, such as edges, textures, color gradients and other features, and transmit them to the second feature extraction module and the first feature integration module respectively;

[0095] The second feature extraction module is used to extract the first-level semantic features of visual features, such as contour, shape and other features, and transmit them to the third feature extraction module and the second feature integration module respectively;

[0096] The third feature extraction module is used to extract the second-level semantic features of the first-level semantic features, such as features such as object parts, and transmit them to the fourth feature extraction module and the third feature integration module respectively;

[0097] The fourth feature extraction module is used to extract global semantic features of the second-level semantic features, such as complete target, scene understanding and other features, and transmit them to the fourth feature integration module;

[0098] The fourth feature integration module is used to perform dimension adjustment processing on the global semantic feature to obtain the first initial scale feature and transmit it to the third feature integration module;

[0099] The third feature integration module is used to perform dimension adjustment processing on the second-level semantic features to obtain a second initial scale feature, and transmit it to the second feature integration module, and at the same time add the first initial scale feature and the second initial scale feature to obtain a first added scale feature;

[0100] The second feature integration module is used to perform dimension adjustment processing on the first-level semantic features to obtain a third initial scale feature, and transmit it to the fourth feature integration module, and at the same time add the second initial scale feature and the third initial scale feature to obtain a second added scale feature;

[0101] The first feature integration module is used to perform dimensionality adjustment processing on the low-level visual features to obtain a fourth initial scale feature, and add the third initial scale feature and the fourth initial scale feature to obtain a third added scale feature.

[0102] Specifically, this embodiment uses the hierarchical feature extraction capability of Swin Transformer, and based on the four-level feature pyramid structure and channel attention mechanism, it takes into account the overall contours of large-scale targets such as ships and lighthouses and the precise segmentation of fine structures such as anchor chains and portholes, significantly improving the segmentation robustness of multi-scale targets, and can adaptively enhance the discriminative features of targets of different scales, thereby improving the segmentation accuracy of multi-scale targets.

[0103] Specifically, the image (with a size of H×W×3) input to the multi-scale feature extraction module first undergoes image patching. The main purpose of image patching is to divide the input image into non-overlapping small patches (patches). In this way, Swin Transformer can reduce the size of the data for subsequent processing, thereby significantly reducing memory usage and computational overhead. In this embodiment, through image patching, the size of the input image becomes In the first stage (Stage 1), the Linear Embedding layer is used to After linear transformation of the image, the size is In the second stage (Stage 2), the Patch Merging layer is used to After downsampling the image, the size is In the third stage (Stage 3), the PatchMerging layer is used to merge the image of size After downsampling the image, the size is In the fourth stage (Stage 4), the size of the image is After downsampling the image, the size is image.

[0104] In a specific embodiment, Figure 4 As shown, each feature extraction module includes at least two consecutive first Swin Transformer blocks of a multi-head self-attention module W-MSA with a regular window and a second Swin Transformer block of a multi-head self-attention module SW-MSA with a shifted window;

[0105] The calculation formula for two consecutive Swin Transformer blocks is:

[0106]

[0107] Where z l-1 is the input feature of the first Swin Transformer block; is the output feature of the multi-head self-attention module with a regular window; z l is the output feature of the first Swin Transformer block; is the output feature of the multi-head self-attention module with a shift window; z l+1 is the output feature of the second Swin Transformer block; MLP represents the processing process of the multi-layer perceptron module; LN is the layer normalization processing process.

[0108] Specifically, in this embodiment, these Swin Transformer blocks are alternately stacked in the forward propagation to build an effective information expression and fusion path, wherein W-MSA calculates self-attention within a fixed window, which can reduce the amount of computation, while SW-MSA enhances the information interaction between windows by shifting the window position, breaking the limitations brought by spatial division. Specifically, the traditional Vision Transformer adopts the global multi-head self-attention mechanism MSA, and its computational complexity increases with the square of the input image size, making it difficult to process high-resolution images. The W-MSA mechanism of Swin Transformer divides the image in the form of non-overlapping local windows, and calculates self-attention independently within each window. For an input image with h×w patches, the computational complexities of the global multi-head self-attention mechanism MSA and W-MSA are:

[0109] Ω(MSA)=4hwC 2 +2(hw) 2 C (2)

[0110] Ω(W-MSA)=4hwC 2 +2M 2 hwC (3)

[0111] As can be seen from the above formula, the former is quadratically related to the number of patches hw, while the latter is linearly related when M is a fixed value (in this embodiment, the default setting is 7), which significantly reduces the processing cost of high-resolution images.

[0112] In addition, the Swin Transformer block also introduces a relative position bias in the self-attention calculation, which enables this model to better capture the spatial relationship between local structures. The calculation formula is:

[0113]

[0114] Among them, Q, K, V are query, key and value matrices, d is the query / key dimension, and B is the learnable relative position bias matrix.

[0115] In a specific embodiment, each feature integration module includes a 1×1 convolution layer, a 2x upsampling layer, and a 3×3 convolution layer;

[0116] The 1×1 convolutional layer is used to adjust the number of channels of the input features. The formula is:

[0117] F n '=Conv 1×1 (F n ) (5)

[0118] Among them, Conv 1×1 (·) represents a 1×1 convolution operation, F n Input feature map for the nth layer;

[0119] The 2x upsampling layer is used to perform a 2x upsampling operation on the features after adjusting the number of channels to improve the spatial resolution of the features after adjusting the number of channels, so that it can maintain rich detail information and prepare richer spatial information for subsequent convolution operations. The upsampled feature map F n ” means:

[0120] F n ”=Upsample ×2 (F n ') (6)

[0121] Among them, Upsample ×2 (·) indicates a 2x upsampling operation;

[0122] The 3×3 convolutional layer is used to adjust the number of channels of the feature after the spatial resolution is improved to make it consistent with the number of channels of other feature maps, and finally obtain a feature map with 256 channels, which is expressed as:

[0123] F n ”'=Conv 3×3 (F n ”) (7)

[0124] Among them, Conv 3×3 Represents a 3×3 convolution operation.

[0125] Specifically, the feature integration module is a feature pyramid module (FPN module), which is used to adjust the input feature maps of different scales into feature maps with a uniform number of channels for subsequent processing.

[0126] In a specific embodiment, the feature enhancement module established based on the channel and spatial dual attention mechanism includes a first enhancement module, a second enhancement module, a third enhancement module and a fourth enhancement module;

[0127] The first enhancement module is used to enhance the third added scale feature to obtain a first enhanced feature;

[0128] The second enhancement module is used to perform enhancement processing on the second added scale feature to obtain a second enhanced feature;

[0129] The third enhancement module is used to enhance the first added scale feature to obtain a third enhanced feature;

[0130] The fourth enhancement module is used to enhance the first initial scale feature to obtain a fourth enhanced feature;

[0131] like Figure 5 As shown, each enhancement module includes: a channel attention module (CAM) and a spatial attention module (SAM);

[0132] Specifically, the channel attention module strengthens the channel relationship of features by generating a channel attention map, focusing on the meaningful information in the input image. It first aggregates spatial information through parallel maximum pooling and average pooling operations, and then generates a channel attention map through a multi-layer perceptron module and a Sigmoid activation function. The calculation formula is:

[0133] M c (F)=σ(MLP(AvgPool(F))+MLP(MaxPool(F))) (8)

[0134] Among them, F represents the input feature map, σ represents the Sigmoid activation function; AvgPool represents the average pooling operation; MaxPool represents the maximum pooling operation; MLP represents the processing process of the multi-layer perceptron module. The multi-layer perceptron module includes a shared fully connected layer (Shared Dense Layer). Based on the shared fully connected layer, different tasks can reuse the same Dense Layer to process some features (such as FC1 and FC2). This sharing mechanism can reduce the number of parameters.

[0135] Specifically, the spatial attention module exploits the spatial relationship of features by generating a spatial attention map, focusing on the location information of the target. It applies maximum pooling and average pooling operations to the output of the channel attention module along the channel axis, concatenates the two feature maps through tensor operations, and processes them through a convolutional layer and a sigmoid activation function to highlight the key target area and reduce the influence of background noise. The calculation formula of the spatial attention module is:

[0136] M s (F)=σ(f 7×7 ([AvgPool(F),MaxPool(F)])) (9)

[0137] Among them, f 7×7 Indicates the use of convolution (Conv2D) operation with a convolution kernel size of 7×7;

[0138] The enhanced features are obtained based on the output of the channel attention module and the spatial attention module, which are expressed as:

[0139]

[0140] in, Represents the element-wise multiplication operation, which multiplies the input feature with the two attention maps sequentially to obtain the enhanced feature.

[0141] Specifically, this embodiment integrates the Dual Attention Module (DAM) into the marine target instance segmentation model, such as Figure 5 As shown, DAM enhances feature representation by sequentially applying channel attention mechanism and spatial attention mechanism, helping our offshore object instance segmentation model to more accurately identify and focus on the target area.

[0142] In a specific embodiment, the calculation formula of the bottom-up path aggregation module includes:

[0143] P2=φ(M2) (11)

[0144]

[0145] P6=MaxPool(P5) (13)

[0146] Wherein, M2 represents the first enhanced feature, P2 represents the first fused feature obtained by the bottom-up path aggregation module fusing the first enhanced feature; i=3,4,5 Represent the second enhancement feature, the third enhancement feature and the fourth enhancement feature respectively; P i=3,4,5 represents the second fused feature, the third fused feature, and the fourth fused feature obtained by the bottom-up path aggregation module by performing feature fusion based on the second enhanced feature, the third enhanced feature, and the fourth enhanced feature respectively; Indicates the downsampling operation (Downsample) of the fusion feature of the i-1 layer; Concat(·) indicates the splicing operation; φ(·) indicates the fusion operation, involving a 1×1 convolution layer and a 3×3 convolution layer; P6 is the fifth fusion feature obtained by the bottom-up path aggregation module performing maximum pooling on the fourth fusion feature.

[0147] Specifically, if Figure 6 As shown in the figure, it is the specific workflow of BPAM. Its working process starts from the base layer, and P2 is directly generated by M2 through the convolution operation. Subsequently, for higher-level enhanced features (M3 to M5), the module concatenates the input of the current layer with the output of the lower layer that has been downsampled and resized. The concatenated features are then processed by a fusion operation to obtain the output of the layer. Finally, in order to cover a wider receptive field, a feature map P6 is generated by applying a maximum pooling operation to P5. Ultimately, BPAM generates a set of multi-scale feature maps containing {P2, P3, P4, P5, P6}. These feature maps combine rich semantic information and fine spatial details, and are simultaneously fed into the subsequent dual-branch feature extraction module.

[0148] In a specific embodiment, a dual-branch feature extraction module established based on the SOLOv2 architecture includes: a category branch and a mask branch;

[0149] Specifically, this embodiment uses a dual-branch feature extraction module based on the SOLOv2 architecture, eliminates redundant region proposal generation steps, and combines matrix non-maximum suppression to optimize the post-processing process. This significantly improves inference efficiency while ensuring segmentation accuracy.

[0150] The category branches are used to extract category branch features from the first to fifth fusion features, respectively, and are expressed as:

[0151] C b =CategoryHead(P b ) (14)

[0152] Among them, P b is the b-th fusion feature output by the bottom-up path aggregation module, b = 2, 3, 4, 5, 6; CategoryHead(·) represents the category branch convolution head, C b The category branch feature of each spatial position in the fusion feature output by the category branch, wherein the category branch feature is the category probability;

[0153] The mask branch is used to extract the mask features from the first to fifth fusion features respectively, which are expressed as:

[0154] M b =MaskHead(P b ) (15)

[0155] Among them, MaskHead(·) represents the mask branch convolution head, M b The mask feature map of the fused features output by the mask branch.

[0156] In a specific embodiment, the step of obtaining an instance segmentation result based on the category branch feature and the mask feature includes:

[0157] The predicted category score of each spatial location point in the fusion feature is calculated based on the category branch feature Expressed as:

[0158]

[0159] Among them, Conv 3×3 represents a 3×3 convolution operation, and σ represents a Sigmoid activation function;

[0160] Specifically, in this embodiment, the category branch is responsible for performing category prediction at each spatial location to ensure that objects of different sizes can be covered at different scales.

[0161] The mask features are associated and calculated through dynamic convolution operation to obtain the candidate instance mask Expressed as:

[0162]

[0163] Among them, DynamicConv(·) represents the operation of dynamically generating masks, θ (i,j) is the dynamic convolution parameter corresponding to the i', j'th position in the mask feature;

[0164] According to the predicted category score of each spatial location point in the fusion feature Filter out the spatial location points whose predicted category scores meet the set threshold;

[0165] A matrix non-maximum suppression method (Matrix NMS) is used to perform parallel operations on the spatial location points and candidate instance masks whose predicted category scores meet the set threshold, and the suppression probability corresponding to the candidate instance mask is calculated;

[0166] Specifically, the suppression probability is used to measure the spatial redundancy relationship between each spatial location point that meets the set threshold and other spatial location points, and dynamically adjust the category score of the spatial location point that meets the set threshold;

[0167] Multiplying the predicted category score of the spatial position point that meets the set threshold by the corresponding suppression probability to obtain the confidence score of the corresponding candidate instance mask. The confidence score is used to measure the credibility of the candidate instance mask as the final instance;

[0168] All candidate instance masks are sorted in descending order by confidence score, and a confidence threshold is set. Redundant masks with low confidence scores are removed, and only candidate instance masks with scores above the confidence threshold are retained to generate a final instance set, reducing duplicate predictions. The final instance set is the instance segmentation result of the marine target. Each instance in the final instance set includes a mask image and a corresponding predicted class label.

[0169] The method proposed in this embodiment addresses the multiple challenges of target segmentation in marine environments (wave reflection interference, coexistence of multi-scale targets, and easy loss of small target details) by innovatively integrating the Swin Transformer, DAM, and BPAM modules in a scenario-based manner: the hierarchical window attention mechanism of the Swin Transformer can effectively model the long-range spatial relationship between ships and wave textures, overcoming the defect of the traditional CNN local receptive field in insufficient modeling of complex backgrounds; the DAM module can suppress invalid high-frequency noise in strong light reflection areas through channel-spatial dual-dimensional attention weighting, while enhancing discriminative features such as hull outlines and portholes; the BPAM module adopts a bottom-up feature fusion path to transfer the anchor chain and buoy details in the underlying high-resolution features to the high-level semantic features, solving the problem of feature annihilation caused by downsampling of small targets at a distance; finally, combined with the dual-branch feature extraction module, while ensuring the accuracy of dense ship instance distinction, it can meet the rigid requirements of the marine monitoring system for real-time processing (<100ms / frame). In this embodiment, the collaborative design between the modules fully adapts to the characteristics of the marine scene, forming a complete technical chain from feature extraction, enhancement, fusion to instance segmentation.

[0170] like Figure 7 The figure shows the multi-object segmentation capability of the maritime object instance segmentation model in complex ocean scenes. The figure annotates the detection boxes and masks of 12 types of vessels, including freightboats, ferries, cruise ships, drills, fishboats, sailboats, beacons, warships, buoys, submarines, and speedboats. In the dense port scene, freight ships (yellow boxes) and ferries (green boxes) experience no overlapping false detections. The model can fully segment the details of the submarine's (pink box) periscope (<30×30 pixels), and the localization error of the buoy's tiny structure (blue box) is less than 2 pixels. As can be seen from the color-coded instance masks, the maritime object instance segmentation model maintains high segmentation accuracy despite strong light reflections (highlight areas on the fishing boat's hull) and fog interference (blurred areas on the cruise ship's outline), verifying the DAM module's robustness to complex lighting conditions.

[0171] like Figure 8 and Figure 9As shown in the figure, the performance differences between the maritime target instance segmentation model and mainstream algorithms such as SOLOv2(ResNet101), SOLOv2(ResNet50), BlendMask(ResNet101), BlendMask(ResNet50), CondInst(ResNet101), CondInst(ResNet50), Box(ResNet101), Box(ResNet50), YOLACT(ResNet101), YOLACT(ResNet50), Mask R-CNN(ResNet101) and Mask R-CNN(ResNet50) are shown. In the same test scenario, traditional methods such as Mask R-CNN (ResNet50) produce mask overlap for densely packed cargo ships, and YOLACT (ResNet50) exhibits broken segmentation in the sailboat mast area. However, the maritime target instance segmentation model accurately separates adjacent ship instances (blue box intersection-over-union ratio reaches 92%) through cross-layer feature fusion of the BPAM module, and repairs the mask discontinuity problem caused by dynamic surges (hull edge integrity is improved by 95%). The comparison results highlight the comprehensive advantages of the method proposed in this embodiment in terms of anti-interference, instance differentiation accuracy, and edge computing adaptability.

[0172] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for marine target instance segmentation based on dual attention mechanism and multi-scale feature fusion, characterized by: The specific steps include: S1: Obtain the marine object instance segmentation dataset; S2: establishing a maritime object instance segmentation model, and training the maritime object instance segmentation model based on the maritime object instance segmentation dataset to obtain a trained maritime object instance segmentation model; The maritime target instance segmentation model includes: A multi-scale feature extraction module based on the Swin Transformer backbone network is used to extract target features of different scales in the marine target instance segmentation dataset; A feature enhancement module based on the channel and spatial dual attention mechanism is used to enhance the target features output by the multi-scale feature extraction module; The bottom-up path aggregation module is used to perform cross-level feature fusion processing on the features output by the feature enhancement module; A dual-branch feature extraction module based on the SOLOv2 architecture is used to extract the category branch features and mask features from the features output by the bottom-up path aggregation module, and obtain instance segmentation results based on the category branch features and mask features; S3: Performing actual marine target instance segmentation based on the trained marine target instance segmentation model to obtain an instance segmentation result.

2. The method for marine target instance segmentation based on dual attention mechanism and multi-scale feature fusion according to claim 1 is characterized in that: The multi-scale feature extraction module based on the Swin Transformer backbone network includes: a first feature extraction module, a second feature extraction module, a third feature extraction module, a fourth feature extraction module, a first feature integration module, a second feature integration module, a third feature integration module and a fourth feature integration module; The first feature extraction module is used to extract visual features of the marine target instance segmentation dataset and transmit them to the second feature extraction module and the first feature integration module respectively; The second feature extraction module is used to extract the first-level semantic features of the visual features and transmit them to the third feature extraction module and the second feature integration module respectively; The third feature extraction module is used to extract the second level semantic features of the first level semantic features and transmit them to the fourth feature extraction module and the third feature integration module respectively; The fourth feature extraction module is used to extract the global semantic features of the second-level semantic features and transmit them to the fourth feature integration module; The fourth feature integration module is used to perform dimension adjustment processing on the global semantic feature to obtain the first initial scale feature and transmit it to the third feature integration module; The third feature integration module is used to perform dimension adjustment processing on the second layer semantic features to obtain a second initial scale feature, and transmit it to the second feature integration module, and at the same time add the first initial scale feature and the second initial scale feature to obtain a first added scale feature; The second feature integration module is used to perform dimension adjustment processing on the first-level semantic features to obtain a third initial scale feature, and transmit it to the fourth feature integration module, and at the same time add the second initial scale feature and the third initial scale feature to obtain a second added scale feature; The first feature integration module is configured to perform dimension adjustment processing on the visual feature to obtain a fourth initial scale feature, and add the third initial scale feature and the fourth initial scale feature to obtain a third added scale feature.

3. The method for marine target instance segmentation based on dual attention mechanism and multi-scale feature fusion according to claim 2 is characterized in that: Each feature extraction module includes at least two consecutive first Swin Transformer blocks of a multi-head self-attention module W-MSA with a regular window and a second Swin Transformer block of a multi-head self-attention module SW-MSA with a shifted window; The calculation formula for two consecutive Swin Transformer blocks is: Where z l-1 is the input feature of the first Swin Transformer block; is the output feature of the multi-head self-attention module with a regular window; z l is the output feature of the first Swin Transformer block; is the output feature of the multi-head self-attention module with a shift window; z l+1 is the output feature of the second Swin Transformer block; MLP represents the processing process of the multi-layer perceptron module; LN is the layer normalization processing process.

4. The method for marine target instance segmentation based on dual attention mechanism and multi-scale feature fusion according to claim 3 is characterized in that: Each feature integration module consists of a 1×1 convolution layer, a 2x upsampling layer, and a 3×3 convolution layer; The 1×1 convolutional layer is used to adjust the number of channels of the input features; The 2x upsampling layer is used to perform a 2x upsampling operation on the features after the number of channels is adjusted, so as to improve the spatial resolution of the features after the number of channels is adjusted; The 3×3 convolutional layer is used to adjust the number of channels of the features processed by improving the spatial resolution.

5. The method for marine target instance segmentation based on dual attention mechanism and multi-scale feature fusion according to claim 4 is characterized in that: The feature enhancement module based on the channel and space dual attention mechanism includes a first enhancement module, a second enhancement module, a third enhancement module and a fourth enhancement module; The first enhancement module is used to enhance the third added scale feature to obtain a first enhanced feature; The second enhancement module is used to perform enhancement processing on the second added scale feature to obtain a second enhanced feature; The third enhancement module is used to enhance the first added scale feature to obtain a third enhanced feature; The fourth enhancement module is used to enhance the first initial scale feature to obtain a fourth enhanced feature; Each enhancement module includes: a channel attention module and a spatial attention module; The calculation formula of the channel attention module is: M c (F)=σ(MLP(AvgPool(F))+MLP(MaxPool(F))) (2) Where F represents the input feature map, σ represents the Sigmoid activation function; AvgPool represents the average pooling operation; MaxPool represents the maximum pooling operation; MLP represents the processing process of the multi-layer perceptron module; The calculation formula of the spatial attention module is: M s (F)=σ(f 7×7 ([AvgPool(F);MaxPool(F)])) (3) Among them, f 7 × 7 Indicates the use of a convolution operation with a convolution kernel size of 7×7; The enhanced features are obtained based on the output of the channel attention module and the spatial attention module, which are expressed as: in, Represents an element-wise multiplication operation.

6. The method for marine target instance segmentation based on dual attention mechanism and multi-scale feature fusion according to claim 5 is characterized in that: The calculation formula of the bottom-up path aggregation module is include: P2=φ(M2) (5) P6=MaxPool(P5) (7) Wherein, M2 represents the first enhanced feature, P2 represents the first fused feature obtained by the bottom-up path aggregation module fusing the first enhanced feature; i=3,4,5 Represent the second enhancement feature, the third enhancement feature and the fourth enhancement feature respectively; P i=3,4,5 represents the second fused feature, the third fused feature, and the fourth fused feature obtained by the bottom-up path aggregation module by performing feature fusion based on the second enhanced feature, the third enhanced feature, and the fourth enhanced feature respectively; Indicates the downsampling operation on the fusion feature of the i-1 layer; Concat(·) indicates the splicing operation; φ(·) indicates the fusion operation, involving a 1×1 convolution layer and a 3×3 convolution layer; P6 is the fifth fusion feature obtained by performing maximum pooling on the fourth fusion feature by the bottom-up path aggregation module.

7. The method for marine target instance segmentation based on dual attention mechanism and multi-scale feature fusion according to claim 6 is characterized in that: The dual-branch feature extraction module based on the SOLOv2 architecture includes: category branch and mask branch; The category branches are used to extract category branch features from the first to fifth fusion features, respectively, and are expressed as: C b =CategoryHead(P b ) (8) Among them, P b is the b-th fusion feature output by the bottom-up path aggregation module, b = 2, 3, 4, 5, 6; CategoryHead(·) represents the category branch convolution head, C b The category branch feature of each spatial position in the fusion feature output by the category branch, wherein the category branch feature is the category probability; The mask branch is used to extract the mask features from the first to fifth fusion features respectively, which are expressed as: M b =MaskHead(P b ) (9) Among them, MaskHead(·) represents the mask branch convolution head, M b The mask feature map of the fused features output by the mask branch.

8. The method for marine target instance segmentation based on dual attention mechanism and multi-scale feature fusion according to claim 7 is characterized in that: The steps of obtaining instance segmentation results based on category branch features and mask features include: The predicted category score of each spatial location point in the fusion feature is calculated based on the category branch feature Expressed as: Among them, Conv 3×3 represents a 3×3 convolution operation, and σ represents a Sigmoid activation function; The mask features are associated and calculated through dynamic convolution operation to obtain the candidate instance mask Expressed as: Among them, DynamicConv(·) represents the operation of dynamically generating masks, θ (i,j) is the dynamic convolution parameter corresponding to the i', j'th position in the mask feature; According to the predicted category score of each spatial location point in the fusion feature Filter out the spatial location points whose predicted category scores meet the set threshold; A matrix non-maximum suppression method is used to perform parallel operations on the spatial location points and candidate instance masks whose predicted category scores meet the set threshold, and the suppression probability corresponding to the candidate instance mask is calculated; Multiplying the predicted category score of the spatial position point that meets the set threshold by the corresponding suppression probability to obtain the confidence score of the corresponding candidate instance mask. The confidence score is used to measure the credibility of the candidate instance mask as the final instance; All candidate instance masks are sorted in descending order according to the confidence scores, and a confidence threshold is set. Only candidate instance masks with a confidence score higher than the confidence score threshold are retained to generate a final instance set, which is the instance segmentation result of the marine target.

Citation Information

Cited By

  • Offshore multi-target tracking detection method for unmanned ship

    CN121305346A