A Multi-Scene Ship Detection and Segmentation Method Based on Hybrid Attention
By introducing hybrid attention mechanisms and the use of feature pyramids in the Mask R-CNN model, the problem of ship detection and segmentation under complex background and noise influence is solved, achieving higher detection and segmentation accuracy, especially in multi-scenario and small-sized ship detection.
Patent Information
- Application Number
- CN202211295062.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-21
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-10-21
AI Technical Summary
The prior art is difficult to effectively detect and divide ships under the influence of complex backgrounds and noise, especially small ships, resulting in insufficient recognition accuracy and real-time.
Using an improved Mask R-CNN model based on a hybrid attention mechanism, a feature pyramid is obtained through FPN, and a channel and spatial attention mechanism are added to the bottom-up path to enhance feature extraction capabilities and segmentation effects.
It significantly improves the accuracy of ship detection and segmentation in different complex environments, especially in multi-scenario and small-sized ship detection, which can better deal with complex background and noise interference.
Smart Images

Figure CN115631427B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence technology, and in particular relates to a multi-scenario ship detection and segmentation method. Background Art
[0002] High-resolution remote sensing image ship target detection is a research hotspot in the field of remote sensing image processing, and is widely used in military reconnaissance and civilian monitoring. As the main means of earth observation, high-resolution satellite remote sensing can achieve sub-meter accuracy. Ships, ports and other targets can be manually interpreted from remote sensing images. However, as the amount of remote sensing image data acquired continues to grow, the manual interpretation method is inefficient and difficult to cope with the growing amount of data. Driven by this demand, the demand for high-performance remote sensing image automatic interpretation systems is urgent.
[0003] Automated ship detection can help obtain information about ship distribution. Its role in maritime surveillance is becoming increasingly important and can help control illegal fishing and cargo transportation. In recent years, the synthesis of satellite remote sensing images for ship detection has become an important research direction. As deep learning shines in computer vision tasks such as target detection and image segmentation, using deep learning to process remote sensing image classification, target detection, segmentation and other tasks has become a mainstream method. Existing remote sensing image semantic segmentation based on deep learning is mostly based on FCN, DeconvNet, U-Net, DeepLab and other network improvements. However, most of the current recognition and segmentation methods focus on target detection and semantic segmentation of remote sensing images.
[0004] At present, the research methods of ship target detection can be roughly divided into two categories: one is the ship detection method based on traditional image processing technology, which detects targets through image segmentation, feature extraction and classifier classification; the other is based on deep learning technology, which uses algorithms such as convolutional neural networks for ship detection, which has become a research hotspot. Traditional image processing methods are widely used and the technology is relatively mature. They can solve the detection and segmentation tasks of remote sensing images to a certain extent. However, with the increasing resolution of remote sensing images, the increasing coverage area, and the increasing richness of information such as geometric textures, traditional methods are gradually unable to cope with the detection and segmentation tasks of remote sensing images, and are limited in practical applications. Although the method based on deep learning has good performance in recognition accuracy and real-time performance, many algorithms have certain pertinence. The ships in remote sensing images are not perfect bird's-eye views, and there will be certain angles and some shadows, which will affect the extraction of some features. Even some ships are small in size and are easily ignored during feature extraction, resulting in non-recognition. Therefore, it is urgent to design a recognition algorithm that can overcome the influence of different complex backgrounds and noises, and can detect and segment small ships. Summary of the invention
[0005] To overcome the deficiencies of the prior art, the present invention provides a multi-scene ship detection and segmentation method based on hybrid attention, an improved model based on Mask R-CNN. In Mask R-CNN, FPN is used to obtain a feature pyramid. After FPN, a bottom-up path is added to shorten the path for the underlying features to be transmitted to the high-level network, so as to retain more shallow features, improve the detection and segmentation effects of small targets, and by using channel and spatial attention mechanisms in the bottom-up path, the feature map can better respond to regions and suppress background regions, thereby improving the detection and segmentation effects.
[0006] The technical solutions adopted by the present invention to solve its technical problems include the following steps:
[0007] Step 1: Data preprocessing;
[0008] Select a remote sensing image ship dataset, divide the dataset into a training set and a test set; perform instance segmentation format annotation on the images in the dataset, and then perform data augmentation to expand the dataset;
[0009] Step 2: Feature extraction;
[0010] Use ResNet101 as the backbone network for initialization, and then use the FPN network for feature extraction to obtain a feature pyramid {P 2 , P 3 , P 4 , P 5}, by utilizing the high resolution of shallow low-level features and the high semantic information of deep high-level features, fusing features of different layers to improve the ability of the network to extract target features, specifically as follows:
[0011] Step 2-1: Initialize the network model using a pre-trained ResNet101 model, which is divided into 5 stages. These 5 stages respectively correspond to the output of feature maps at different scales, used to establish the feature pyramid {P 2 , P 3 , P 4 , P 5} of the FPN network;
[0012] Step 2-2: The FPN network uses the methods of bottom-up path, top-down path and lateral connection to fuse shallow features and deep features;
[0013] Step 2-2-1: Bottom-up path;
[0014] Bottom-up is the process of the convolutional network; for the ResNet101 model, the feature activation outputs of the last residual structure in each stage are used, and these outputs are represented as feature maps {C 2 , C 3 , C 4 , C 5}, corresponding to the outputs of conv2, conv3, conv4, and conv5;
[0015] Step 2-2-2: Top-down path and lateral connections;
[0016] Through the bottom-up path, FPN obtains four groups of feature maps. To combine these four groups of feature maps with different characteristics, FPN uses the strategy of top-down and lateral connections:
[0017] P 5 is obtained by convolving the feature map C 5 using a 1×1 convolutional kernel with 256 channels; for the feature map P i+1 , i = 4, 3, 2, perform 2x upsampling. The upsampled feature map is fused with the corresponding feature map C i through element-wise addition operation to generate a new fused feature map, and finally obtain the feature pyramid {P 2 , P 3 , P 4 , P 5};
[0018] Step 3: Feature enhancement;
[0019] Add a bottom-up structure to the Mask R-CNN network, receive the feature pyramid obtained by the FPN network, shorten the transmission path between the shallow features and the top layer, enhance the transmission of the shallow features in the network, so that more shallow features can be utilized, and add channel and spatial attention mechanisms to further enhance the feature extraction ability of the network, thereby obtaining a new feature pyramid {N 2 , N 3 , N 4 , N 5}, and then send it to the subsequent network layer, specifically as follows:
[0020] Step 3-1: Introduce channel attention mechanism and spatial attention mechanism to assign weights to image features; Assume that the feature map F = {F 1 , F 2 , …, F C} is obtained through the FPN network, where W and H represent the width and height of the feature map F respectively, C represents the number of channels of the feature map F, and F i ∈ R W×HRepresents the feature of the $i$-th channel of the feature map $F$;
[0021] Step 3-1-1: Channel attention mechanism;
[0022] For each channel feature $F$ of the feature map $F$ i Perform average pooling operation, pooling the feature map of size $W\times H$ in each channel into a feature map of size $1\times1$, obtaining a channel feature vector $V$:
[0023] $V = [V$ 1 , $V$ 2 , …, $V$ C , $V\in\mathbb{R}$ C
[0024] where $V$ i represents the feature obtained after average pooling for the $i$-th channel; then use a convolutional kernel of size $1\times1$ to learn aggregated features from each channel $V$ i and perform the Softmax operation on the learned features to make the sum of all channel values equal to 1, achieving weight allocation for each channel;
[0025] Step 3-1-2: Spatial attention mechanism
[0026] For the feature map $F$, the set of spatial points is defined as $L = \{(x, y)|x = 1, \ldots, W; y = 1, \ldots, H\}$, where $(x, y)$ represents the coordinates of the spatial point; the process of generating the spatial attention map is as follows:
[0027] First, use a $1\times1$ convolution to integrate the feature map $F$ with $C$ channels into a single feature map $M$ with 1 channel:
[0028] $M = W$ S * $F + b$ S , $V\in\mathbb{R}$ C
[0029] where $W$ S represents the convolutional kernel, * represents the convolution operation, and $b$ S represents the bias; at this time, the feature map $M$ fuses the information of all channels in the feature map $F$;
[0030] Then perform the Softmax operation on the space of the feature map $M$. Let $M(l)$ represent the feature vector at the spatial position $l$, then the attention weight obtained at $l$ is:
[0031]
[0032] where the sum of the weights of all spatial points is 1, $\sum$ l∈L $a$ S(l) = 1; Assign weights to each pixel point in space;
[0033] Step 3-2: Combine the spatial attention mechanism and the channel attention mechanism to obtain a channel-first then spatial attention unit CA-SA; This attention unit is formed by sequentially connecting the channel attention unit CA and the spatial attention unit SA in series. The feature map F first passes through the channel attention unit to obtain the channel-level attentional feature F ca , and then F ca is sent to the spatial attention unit to obtain the final channel-spatial attentional feature F csa ;
[0034] Step 3-3: Add a Bottom-up structure to the Mask R-CNN network and add an attention mechanism, specifically as follows: To generate the feature map N i+1 , i = 2, 3, 4, N i (W×H×C) is downsampled by a convolutional kernel of size 3×3, with 256 channels and a convolutional stride of 2, reducing the height and width of the feature map by half, making it the same size as P i+1 (W×H×C). The downsampled feature map is added element-wise to P i+1 to obtain the fused feature map N i_half , and then the fused feature map N i_half is sent to the attention unit module CA-SA to obtain the attentional feature map N i ; Finally, a new feature pyramid {N 2 , N 3 , N 4 , N 5} is obtained and then sent to the subsequent grid layer;
[0035] Step 4: Object detection and segmentation;
[0036] On top of the set of feature maps finally produced by the backbone network, use RPN to generate multiple region proposals, and then generate the corresponding ROI windows for these region proposals respectively. After RoIAlign, the corresponding feature information is obtained, which is then used for subsequent object classification, object box localization, and object box Mask map recognition;
[0037] Step 4-1: The extracted features enter the RPN network to generate region proposals; This layer uses softmax to determine whether the anchors belong to the foreground or background, and then uses bounding box regression to correct the anchors to obtain accurate proposals;
[0038] Step 4-2: According to the candidate boxes obtained in Step 4-1 and the feature pyramid extracted in Step 3-3, extract the feature maps corresponding to the candidate boxes through the mapping relationship, and enter the RoI pooling layer to pool each feature box into a fixed size of H×W;
[0039] Step 4-3: Input the features obtained in Step 4-2 into the Mask branch and the fully connected layer respectively; after inputting into the Mask branch, obtain the Mask result. Through two parallel fully connected layers, output the classification score of this area and perform bounding box regression on the candidate boxes again to obtain the final position of the detection box.
[0040] Preferably, the remote sensing image ship dataset is the remote sensing image ship dataset of Airbus for semantic segmentation tasks.
[0041] Preferably, the data augmentation process is to randomly flip the training images horizontally.
[0042] The beneficial effects of the present invention are as follows:
[0043] The present invention proposes an end-to-end deep learning network for ship detection and segmentation in remote sensing satellite images. Compared with the baseline model and other methods, the present invention significantly improves the overall ship detection and segmentation accuracy in different complex environments and can well handle ship detection in multiple scenarios. In addition, the method of the present invention further improves the detection and segmentation accuracy of small ships. Description of the Drawings
[0044] Figure 1 This is an example of the annotation format conversion of the embodiment of the present invention. (a) Original image, (b) Annotation in semantic segmentation format, (c) Annotation in instance segmentation format.
[0045] Figure 2 This is the overall architecture diagram of ResNet101 in the embodiment of the present invention.
[0046] Figure 3 This is the schematic diagram of the FPN structure in the embodiment of the present invention.
[0047] Figure 4 This is the illustration of channel attention (CA) and spatial attention (SA) in the embodiment of the present invention.
[0048] Figure 5 This is the illustration of the Bottom-up-AM structure in the embodiment of the present invention.
[0049] Figure 6 This is the schematic diagram of the "lateral connection" feature fusion of the Bottom-up-AM in the embodiment of the present invention.
[0050] Figure 7 Samples for ship detection and segmentation in the embodiments of the present invention: (a) original image, (b) ground truth, (c) results of the Mask R-CNN baseline model, (d) results of Mask R-CNN + S-NMS, (e) results of the Mask Scoring R-CNN model; (f) results of the PANET model; (g) results of the SCRDet model; (h) results of Mask R-CNN_B-CA-SA (this method);
[0051] Figure 8 Schematic structural diagram of the method for multi-scene ship detection and segmentation based on hybrid attention of the present invention Detailed implementation manners
[0052] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0053] The method for instance segmentation of remote sensing images based on deep learning of the present invention realizes pixel-level segmentation of the target while achieving target detection, and in view of the problems of complex background features and large noise interference in remote sensing images, an attention module is introduced, and the attention mechanism is used to suppress interference features such as noise in the complex background; in view of the problem of poor segmentation and detection and recognition effects of small ships in remote sensing images, an attention "bottom-up" multi-scale feature fusion structure is constructed to improve the network feature extraction ability and improve the detection and recognition effects of small ships in remote sensing images.
[0054] A method for multi-scene ship detection and segmentation based on hybrid attention, characterized by comprising the following steps:
[0055] Step A: A data preprocessing module selects a suitable data set and preprocesses the data set to prepare for the training and testing of subsequent models.
[0056] Step A-1: Use the remote sensing image ship data set created by Airbus for exploring the segmentation of ship targets on the sea surface in remote sensing images, which can be used for semantic segmentation tasks. This data set contains ship remote sensing images in different regions and various scenarios (sea-land background, cloud occlusion, clear targets), with a total of more than 150,000 images. However, most of the images do not contain ships and need to be screened, and some images with poor effects are deleted. Finally, about 42,500 images are left, and the data set is divided into a training set and a test set.
[0057] Step A-2: The original data set is encoded in the RLE format, which is a labeling method for semantic segmentation tasks. In order to be applicable to instance segmentation tasks and facilitate training, the original labeling method needs to be converted. An example of the labeling format conversion is as Figure 1 shown.
[0058] Step A-3: Data augmentation processing. The training images are randomly flipped horizontally to expand the dataset size.
[0059] Step B: Feature extraction module. Initialize using ResNet101 as the backbone network, and then use the FPN network for feature extraction to obtain the feature pyramid {P 2 , P 3 , P 4 , P 5}}. By leveraging the high resolution of shallow low-level features and the high semantic information of deep high-level features, fuse the features of different layers to improve the network's ability to extract target features.
[0060] Step B-1: Initialize the network model using a pre-trained ResNet101 model, which is divided into 5 Stages. These 5 stages respectively correspond to the output of feature maps at different scales, used to build the feature pyramid of the FPN network. As Figure 2 shown.
[0061] Step B-2: The FPN network uses the "bottom-up" pathway, "top-down" pathway, and "lateral connection" methods to fuse shallow features and deep features.
[0062] (1) Bottom-up path
[0063] The bottom-up is the process of the convolutional network. For ResNet101, use the feature activation output of the last residual structure in each stage. Represent the outputs of these residual modules as {C 2 , C 3 , C 4 , C 5}}, corresponding to the outputs of conv2, conv3, conv4, conv5, and note that they have a stride (i.e., receptive field) of {4, 8, 16, 32} pixels relative to the input image. Considering memory occupancy, conv1 is not included in the pyramid.
[0064] (2) Top-down path and lateral connection
[0065] Through the bottom-up path, the FPN obtains four groups of feature maps. The shallow feature maps such as C 2 contain more texture information, while the deep feature maps such as C 5Contains more semantic information. To combine these four groups of feature maps with different characteristics, FPN uses a top-down and lateral connection strategy, as Figure 3 shown.
[0066] P 5 is obtained by convolving the feature map C 5 using a 1×1 convolutional kernel with 256 channels. For other levels of {P 3 , P 4 , P 5}, the feature map P i+1 (W×H×C) (i = 4, 3, or 2) is upsampled by a factor of 2. Then the upsampled feature map is fused with the corresponding feature map C i (2W×2H×C) (after passing through a 1×1 convolutional layer to change the number of channels to 256) through element-wise addition to generate a new fused feature map, and finally obtain the feature pyramid {P 2 , P 3 , P 4 , P 5}, which is sent to the subsequent network.
[0067] Step C: Feature Enhancement Module. Add a Bottom-up structure to the Mask R-CNN network. Receive the feature pyramid obtained from the FPN network, shorten the transmission path from the shallow features to the top layer, enhance the transmission of shallow features in the network, so that more shallow features can be utilized, and add channel and spatial attention mechanisms to further enhance the feature extraction ability of the network, thereby obtaining a new feature pyramid {N 2 , N 3 , N 4 , N 5}, and then send it to the subsequent network layer.
[0068] Step C-1: Introduce channel attention mechanism and spatial attention mechanism to assign weights to image features. Assume that the feature map F = {F 1 , F 2 , …, F C} is obtained through the convolutional network, where W and H represent the width and height of the feature map F respectively, C represents the number of channels of the feature map F, and F i ∈R W×H represents the feature of the i-th channel of the feature map F.
[0069] (1) Channel Attention Mechanism
[0070] For each channel feature F of the feature map F iPerform average pooling operation on it, and pool the feature map of size W×H in each channel into a feature vector of size 1×1. Then, a channel feature vector V is obtained:
[0071] V = [V 1 , V 2 , …, V C , V ∈ R C
[0072] where V i represents the feature obtained after average pooling in the i-th channel. Then, use a 1×1 convolutional kernel to learn aggregated features from each channel V i , and then perform the Softmax operation on the learned features to make the sum of all channel values equal to 1, thus realizing the weight assignment for each channel. As Figure 4 is the process of the channel attention mechanism.
[0073] (2) Spatial attention mechanism
[0074] For the feature map F, the set of spatial points is defined as L = {{x, y}|x = 1, …, W; y = 1, …, H}, where (x, y) represents the coordinates of the spatial point. The process of generating the spatial attention map is as follows:
[0075] First, use a 1×1 convolution to integrate the feature map F with C channels into a feature map M with 1 channel:
[0076] M = W S *F + b S , V ∈ R C
[0077] where, W S represents the convolutional kernel, * represents the convolution operation, and b S represents the bias. At this time, the feature map M integrates the information of all channels in the convolutional feature F.
[0078] Then, perform the Softmax operation on the space of the feature map M. Let M(l) represent the feature vector at the spatial position l, then the attention weight obtained at l is:
[0079]
[0080] where, the sum of the weights of all spatial points is 1, ∑ l∈L a S (l) = 1. In this way, the weight assignment for each pixel point in the space is completed. As Figure 4 is the process of the spatial attention mechanism.
[0081] Step C-2: Design an attention unit based on the spatial attention mechanism and the channel attention mechanism. Combine the spatial attention mechanism and the channel attention mechanism to obtain CA-SA (channel-first then spatial attention unit). This attention unit is formed by cascading CA (channel attention unit) and SA (spatial attention unit) in sequence, as Figure 5 shown. The feature map F first passes through the channel attention unit to obtain the channel-level attention feature F ca , and then F ca is sent to the spatial attention unit to obtain the final channel-spatial attention feature F csa .
[0082] Step C-3: Compared with natural images, remote sensing images are taken from a high altitude and cover a wide range, and the targets in remote sensing images are very small. To effectively detect small targets in remote sensing images, a Bottom-up structure is added to the Mask R-CNN network to shorten the transmission path from shallow features to the top layer, enhance the transmission of shallow features in the network, and add an attention mechanism to further enhance the feature extraction ability of the network. Figure 5 is the Bottom-up structure with the attention mechanism added, called Bottom-up-AM (Bottom-up Structure with Attention Module). Specifically as follows: N 2 is the same as P 2 . To generate the feature map N i+1 (i = 2, 3, or 4), as Figure 6 shown, N i (W×H×C) is downsampled by a convolutional kernel with a size of 3×3, 256 channels, and a convolutional stride of 2, reducing the height and width of the feature map by half, which is the same size as P i+1 (W×H×C). The downsampled feature map is added element-wise to P i+1 to obtain the fused feature map N i_half , and then the fused feature map N i_half is sent to the attention unit module (CA-SA) to obtain the attention feature map N i . Finally, a new feature pyramid {N 2 , N 3 , N 4 , N 5} is obtained and then sent to the subsequent grid layer.
[0083] Step D: Target detection and segmentation module. Based on the set of feature maps finally generated by the backbone network, the RPN is used to generate multiple region proposals, and then these region proposals are respectively used to generate corresponding ROI windows. After RoIAlign, the corresponding feature information is obtained, which can be used for subsequent target classification, target box localization, and target box Mask map recognition, etc.
[0084] Step D-1: The extracted features enter the RPN network to generate region proposals. This layer uses softmax to determine whether the anchors belong to the foreground or background, and then uses bounding box regression to correct the anchors to obtain accurate proposals.
[0085] Step D-2: According to the candidate boxes obtained in Step D-1 and the feature maps extracted in Step C-3, the feature maps corresponding to the candidate boxes are extracted through the mapping relationship, and each feature box is pooled into a fixed size of H×W in the RoI pooling layer.
[0086] Step D-3: The features obtained in Step D-2 are respectively passed into the Mask branch and the fully connected layer. After passing into the Mask branch, the Mask result is obtained. Through two parallel fully connected layers, the classification score of this region is output, and the bounding box regression is performed on the candidate box again to obtain the final accurate position of the detection box. The specific experimental results of this method are as Figure 7 shown. Specific embodiments:
[0088] The structural schematic diagram of this method is as Figure 8 shown. Now, in combination with Figure 8 the specific implementation process of this method will be described in detail. The specific steps are as follows:
[0089] Step A: Data preprocessing module.
[0090] Select the remote sensing image ship dataset of Airbus as the experimental data, and screen out 42,500 images from it as the experimental dataset of this method. Then, 3000 images are randomly selected from the dataset as the test set, and the remaining images are used as the training set.
[0091] The original dataset is encoded in the RLE format. For the convenience of training, it is converted to the COCO annotation format.
[0092] In order to make the training model have better performance, the training images are randomly flipped horizontally to enrich the image training set.
[0093] Step B: Feature Extraction Module.
[0094] It is initialized with a pre-trained ResNet-101 model. The batch size is set to 1, the initial learning rate is set to 0.001. Stochastic Gradient Descent (SGD) is used as the optimizer during training, with its weight decay set to 0.0001 and the momentum to 0.9. The maximum number of iterations for the entire model training is set to 350,000 times. The Feature Pyramid Network (FPN) of the ResNet-based feature extraction network consists of three parts: bottom-up, lateral connections, and top-down.
[0095] The bottom-up network is based on ResNet. The convolutional process of ResNet is divided into five stages. Starting from the second stage, denoted as Conv2, Conv3, Conv4, Conv5, which are represented by {C 2 , C 3 , C 4 , C 5}. For each convolution of a module, the size of the feature map becomes 1 / 2 of the original. The output of each stage is used as the input for the lateral connections.
[0096] In the lateral connections of the bottom-up ResNet, {C 2 , C 3 , C 4 , C 5} respectively perform a 1×1 convolution to keep the same number of output channels and fuse with the top-down network.
[0097] The top-down starts from the highest layer for nearest neighbor upsampling and fuses with the lateral connections. The fused features are processed to output several Feature Maps of different sizes. M5 is downsampled to get P6. After two layers of fusion, four feature maps of different sizes are obtained, denoted as {P 2 , P 3 , P 4 , P 5} from bottom to top in sequence. Next, the extracted features are input into the Bottom-up structure.
[0098] Step C: Feature Enhancement Module.
[0099] Each feature map N i (of size W×H) first undergoes a convolution operation with a 3×3 convolution kernel and a convolution stride of 2 to reduce the spatial size of the feature map, obtaining a new feature map, which may be called Then, through lateral connections, each element in the feature map and the feature map N′ iPerform an element-wise addition operation. After the addition and fusion of the feature maps, a convolution operation with a convolution kernel size of 3×3 and a convolution stride of 2 is used to eliminate the overlapping shadows of the fusion and generate a new feature map N i+1 , and finally, the feature map is sent to the attention unit module (CA-SA) to obtain the attention feature map, that is, a new feature pyramid {N 2 ,N 3 ,N 4 ,N 5}.
[0100] Step D: Target detection and segmentation module.
[0101] Set a predetermined number of ROIs for each point in the feature map set, and finally obtain multiple optimal candidate ROIs. Then, the candidate ROIs are sent into the RPN network for binary classification of foreground or background and bounding box regression to filter out some candidate ROIs. Perform ROIAlign operation on the remaining ROIs, that is, first map the pixels of the original image and the feature map, and then map the feature map and the fixed features. Finally, perform N-class classification, position prediction, and target box regression on the ROIs. Predict the Mask and perform semantic segmentation through an improved segmentation network to obtain a binary mask image.
Claims
1. A multi-scenario ship detection and segmentation method based on hybrid attention, characterized in that, it includes the following steps: Step 1: Data preprocessing; Select the remote sensing image ship dataset, divide the dataset into a training set and a test set; perform instance segmentation format annotation on the images in the dataset, and then perform data augmentation to expand the dataset; Step 2: Feature extraction; Initialize using ResNet101 as the backbone network, and then use the FPN network for feature extraction to obtain the feature pyramid {P 2 , P 3 , P 4 , P 5}. By leveraging the high resolution of shallow low-level features and the high semantic information of deep high-level features, fuse the features of different layers to improve the network's ability to extract target features, as follows: Step 2-1: Initialize the network model using a pre-trained ResNet101 model, which is divided into 5 stages. These 5 stages respectively correspond to the output of feature maps at different scales and are used to build the feature pyramid {P 2 , P 3 , P 4 , P 5} of the FPN network; Step 2-2: The FPN network uses the methods of bottom-up path, top-down path and lateral connection to fuse shallow features and deep features; Step 2-2-1: Bottom-up path; Bottom-up is the process of the convolutional network; for the ResNet101 model, the feature activation outputs of the last residual structure in each stage are used, and these outputs are represented as feature maps {C 2 , C 3 , C 4 , C 5}, corresponding to the outputs of conv2, conv3, conv4, and conv5; Step 2-2-2: Top-down path and lateral connection; Through the bottom-up path, the FPN obtains four groups of feature maps. In order to combine these four groups of feature maps with different characteristics, the FPN uses the strategy of top-down and lateral connection: P 5 is convolved from the feature map C 5 using a 1×1 convolutional kernel with 256 channels; for the feature map P i+1 , i = 4, 3, 2, perform 2x upsampling, and the upsampled feature map is fused with the corresponding feature map C i through element-wise addition to generate a new fused feature map, and finally obtain the feature pyramid {P 2 , P 3 , P 4 , P 5}; Step 3: Feature enhancement; Add a Bottom-up structure to the Mask R-CNN network, receive the feature pyramid obtained by the FPN network, shorten the transmission path between the shallow features and the top layer, enhance the transmission of the shallow features in the network, enable more shallow features to be utilized, and add channel and spatial attention mechanisms to further enhance the feature extraction ability of the network, so as to obtain a new feature pyramid {N 2 , N 3 , N 4 , N 5}, and then send it to the subsequent network layer, specifically as follows: Step 3-1: Introduce the channel attention mechanism and the spatial attention mechanism to assign weights to the image features; Assume that the feature map F = {F 1 , F 2 ,..., F C} is obtained through the FPN network, where W and H represent the width and height of the feature map F respectively, C represents the number of channels of the feature map F, and F i ∈ R W×H represents the feature of the i-th channel of the feature map F; Step 3-1-1: Channel attention mechanism; For each channel feature F of the feature map F i perform an average pooling operation to pool the feature map of size W×H in each channel into a feature vector V of size 1×1, obtaining a channel feature vector V: V = [V 1 , V 2 ,..., V C , V ∈ R C Among which V i represents the feature obtained after average pooling for the i-th channel; then a 1×1 convolutional kernel is used to learn aggregated features from each channel V i and then a Softmax operation is performed on the learned features to make the sum of all channel values equal to 1, achieving weight allocation for each channel; Step 3-1-2: Spatial attention mechanism For the feature map F, the set of spatial points is defined as L = {{x, y}|x = 1,..., W; y = 1,..., H}, where (x, y) represents the coordinates of the spatial point; the process of generating the spatial attention map is as follows: First, use a 1×1 convolution to integrate the feature map F with C channels into a feature map M with 1 channel: M = W S *F + b s , V ∈ R C Among them, W S represents the convolution kernel, * represents the convolution operation, and b s represents the bias; at this time, the feature map M integrates the information of all channels in the feature map F; Then perform the Softmax operation on the space of the feature map M. Let M(l) represent the feature vector at the spatial position l, then the attention weight obtained at l is: Among them, the sum of the weights of all points in space is 1, ∑ l∈L a S (l) = 1; complete the assignment of weights to each pixel point in space; Step 3-2: Combine the spatial attention mechanism and the channel attention mechanism to obtain a channel-first then spatial attention unit CA-SA; this attention unit is formed by serially connecting the channel attention unit CA and the spatial attention unit SA in sequence. The feature map F first passes through the channel attention unit to obtain the channel-level attentional feature F ca , and then F ca is fed into the spatial attention unit to obtain the final channel-spatial attentional feature F csa ; Step 3-3: Add a Bottom-up structure to the MaskR-CNN network and add an attention mechanism as follows: To generate the feature map N i+1 , where i = 2, 3, 4, N i (W×H×C) is downsampled by a convolutional kernel of size 3×3, with 256 channels and a convolutional stride of 2, reducing the height and width of the feature map by half, making it the same size as P i+1 (W×H×C). The downsampled feature map is added element-wise to P i+1 to obtain the fused feature map N i_half . Then, the fused feature map N i_half is sent to the attention unit module CA-SA to obtain the attention feature map N i . Finally, a new feature pyramid {N 2 , N 3 , N 4 , N 5} is obtained and then sent to the subsequent grid layer; Step 4: Object detection and segmentation; On the set of feature maps finally generated by the backbone network, use the RPN to generate multiple region proposals, and then generate the corresponding ROI windows for these region proposals respectively. After RoIAlign, the corresponding feature information is obtained, which is then used for subsequent object classification, object box localization and object box Mask map recognition; Step 4-1: The extracted features enter the RPN network to generate region proposals; this layer judges whether the anchors belong to the foreground or the background through softmax, and then uses bounding box regression to correct the anchors to obtain accurate proposals; Step 4-2: According to the candidate boxes obtained in Step 4-1 and the feature pyramid extracted in Step 3-3, extract the feature maps corresponding to the candidate boxes through the mapping relationship, and enter the RoI pooling layer to pool each feature box to a fixed size of H×W; Step 4-3: Transmit the features obtained in Step 4-2 into the Mask branch and the fully connected layer respectively; after entering the Mask branch, the Mask result is obtained. Through two parallel fully connected layers, the classification score of this region is output and the bounding box regression is performed on the candidate box again to obtain the final position of the detection box.
2. A multi-scenario ship detection and segmentation method based on hybrid attention according to claim 1, characterized in that, The remote sensing image ship dataset is the remote sensing image ship dataset of Airbus for semantic segmentation tasks.
3. A multi-scenario ship detection and segmentation method based on hybrid attention according to claim 1, characterized in that the data augmentation process is to randomly horizontally flip the training images.
Citation Information
Patent Citations
Optical remote sensing image ship detection and segmentation method based on deep learning
CN112507777A
Ship instance segmentation method based on spatial distribution attention
CN114581670A