Attention-Enhanced Arbitrary-Direction Dense Ship Target Detection Method

By improving the RetinaNet network structure, BiFPN and dual-branch attention enhancement module were introduced, and combined with the rotary detection box, the false alarm and missed detection problems of ship target detection in SAR images are solved, and the precise positioning and direction estimation of ship targets is achieved, and the detection accuracy is improved.

CN116168240BActive Publication Date: 2025-07-29XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310070828.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-19
Publication Date
2025-07-29
Estimated Expiration
2043-01-19

AI Technical Summary

Technical Problem

The existing SAR image target detection methods are difficult to effectively distinguish ship targets and backgrounds under complex backgrounds, resulting in false alarms and missed detection. The traditional methods have low detection accuracy under complex conditions of clutter backgrounds, and deep learning-based methods cannot accurately locate ship target positions.

Method used

Using an arbitrary intensive ship target detection method based on attention enhancement, the weighted bidirectional feature fusion network BiFPN and a dual-branch attention enhancement module are introduced, and the feature extraction capability is improved and the direction estimation is performed, and the RetinaNet network structure is improved.

Benefits of technology

Improve detection performance in complex coastal scenarios, reduce mis-checking, achieve accurate positioning and direction estimation of ship targets, and improve detection accuracy and information utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116168240B_ABST
    Figure CN116168240B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for detecting arbitrary-direction dense ship targets based on attention enhancement, which relates to the technical field of image processing and includes: obtaining an original image; extracting multi-scale features of the original image based on a backbone feature extraction module to obtain a plurality of feature maps; based on a dual-branch attention enhancement module, learning the importance degrees of each channel and each space in at least part of the feature maps, and learning the importance degrees of each position information in at least part of the feature maps to obtain an output feature map based on the dual-branch attention enhancement module; based on a weighted bidirectional feature fusion network, screening and fusing the information of the output feature map through cross-scale connection operations to obtain an enhanced feature map; and detecting the enhanced feature map based on the classification structure and the bounding box regression structure in the detector. The present invention can improve the network's ability to extract important features and achieve accurate extraction and direction estimation of targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and in particular relates to a method for detecting dense ship targets in any direction based on attention enhancement. Background Art

[0002] Synthetic aperture radar (SAR), as an active microwave imaging sensor, possesses the unique capability of all-day, all-weather Earth observation. It has become one of the primary methods for Earth observation and holds a crucial position in ocean exploration. Ship target detection, a fundamental function of marine ship management systems, plays a vital role in furthering ship target identification and tracking. Therefore, research on SAR ship target detection is of great significance.

[0003] In existing technologies, the main challenge in SAR image target detection research is to extract the target region of interest from SAR images and remove false alarms caused by environmental clutter and artificial clutter. Existing mainstream SAR ship target detection methods can be divided into traditional model-driven detection algorithms and data-driven deep learning detection algorithms. Traditional SAR target detection algorithms are mainly represented by constant false alarm rate target detection methods based on the statistical distribution of background clutter and salient target detection algorithms based on visual attention models. Deep learning-based detection algorithms are divided into two-stage detection algorithms, such as the R-CNN series, and one-stage detection algorithms, such as the YOLO series. However, in complex backgrounds such as islands, ports, and bays, the clutter scattering intensity of SAR images is high, the clutter background is non-uniform, and the distribution of ship targets is diverse. Traditional methods based on constant false alarm rate (CFAR) have difficulty selecting a suitable clutter background model and cannot be well applied to ship target detection at sea under multi-scale and complex background clutter conditions, resulting in a large number of false alarms and missed alarms. Deep learning-based detection algorithms usually use deep learning detection algorithms applied in the optical field, such as Yolov5 and ReDet. Most of these methods use horizontal boxes for detection, resulting in the inclusion of most background pixels in the detection box. This makes it impossible to accurately locate the position of ship targets, which is not conducive to target detection.

[0004] Therefore, there is an urgent need to improve the defects in the prior art. Summary of the Invention

[0005] In order to solve the above problems existing in the prior art, the present invention provides a method for detecting dense ship targets in any direction based on attention enhancement. The technical problem to be solved by the present invention is achieved through the following technical solutions:

[0006] In a first aspect, the present invention provides a method for detecting dense ship targets in any direction based on attention enhancement, comprising:

[0007] Obtain the original image;

[0008] Based on the backbone feature extraction module, extract multi-scale features of the original image to obtain multiple feature maps;

[0009] Based on the dual-branch attention enhancement module, learn the importance of each channel and each space in at least part of the feature maps respectively, as well as learn the importance of each position information in at least part of the feature maps, to obtain the first-branch attention feature map and the second-branch attention feature map, and merge the first-branch attention feature map and the second-branch attention feature map to obtain the output feature map based on the dual-branch attention enhancement module;

[0010] Based on the weighted bidirectional feature fusion network, through cross-scale connection operations, screen and fuse the information of the output feature map to obtain the enhanced feature map;

[0011] Based on the classification structure and bounding box regression structure in the detector, detect the enhanced feature map.

[0012] Advantages of the present invention:

[0013] An arbitrary-direction dense ship target detection method based on attention enhancement provided by the present invention aims to solve problems such as poor detection performance and missed and false detections in complex coastal scenarios. At the same time, it uses a rotated detection box instead of a horizontal detection box to achieve direction estimation of the target while effectively distinguishing the target area from the background area; it uses the weighted bidirectional feature fusion network BiFPN instead of the PANet network, and uses adaptive adjustment of feature weights to obtain more context information and global information, improving the information utilization rate; it uses the dual-branch attention enhancement module to fully strengthen the roles of spatial attention information, channel attention information, and position information, and enhance the network's ability to extract important features.

[0014] The following will further elaborate on the present invention in conjunction with the drawings and embodiments. Description of the Drawings

[0015] Figure 1 is a structural diagram of an arbitrary-direction dense ship target detection method based on attention enhancement provided by an embodiment of the present invention;

[0016] Figure 2 is a schematic diagram of a ResNet residual block network structure provided by an embodiment of the present invention;

[0017] Figure 3 is a schematic diagram of a dual-branch attention enhancement module provided by an embodiment of the present invention;

[0018] Figure 4 is a schematic diagram of a channel attention module provided by an embodiment of the present invention;

[0019] Figure 5 It is a schematic diagram of the spatial attention model provided by an embodiment of the present invention;

[0020] Figure 6 It is a schematic diagram of the coordinate attention model provided by an embodiment of the present invention;

[0021] Figure 7 It is a schematic diagram of the BiFPN network structure provided by an embodiment of the present invention;

[0022] Figure 8 (a) It is a result diagram of the true annotation position of the ship target provided by an embodiment of the present invention;

[0023] Figure 8 (b) It is a result diagram of the experimental comparison of the detection performance provided by an embodiment of the present invention;

[0024] Figure 8 (c) It is a result diagram of the experimental results of the detection performance provided by an embodiment of the present invention. Detailed implementation manners

[0025] The following further describes the present invention in detail with reference to specific embodiments, but the implementation manners of the present invention are not limited thereto.

[0026] The sea remote sensing SAR image scene is relatively complex. There are a large number of clutter false alarms, especially in the nearshore area, and most of the ship targets in the sea are small targets in the far sea. The disadvantages of the existing technologies are that it is difficult for the detection algorithm to extract the target features of the SAR image, it is difficult to effectively obtain the feature information of small targets, and the detection accuracy is not high in the case of complex backgrounds and large target scale differences, and false alarms and missed detections are likely to occur. In addition, most of the existing ship target detection methods use horizontal detection frames for detection, resulting in most of the background pixels being included in the detection frame, and the position of the ship target cannot be accurately located, which is not conducive to the detection work of the target.

[0027] As an excellent single-stage target detection algorithm, RetinaNet has achieved very remarkable detection effects on large optical datasets such as ImageNet, PASCAL VOC, and MSCOCO. However, due to the large differences between SAR images and optical images, there is strong clutter interference in the sea remote sensing SAR images in actual complex scenes, so direct application results in poor detection performance.

[0028] In view of this, aiming at the problems of complex SAR image background, large target scale difference, and high detection false alarm rate, the present invention provides a method for detecting arbitrary-direction dense ship targets based on attention enhancement. This detection method is improved on the basis of the RetinaNet horizontal detection network. First, a weighted bidirectional feature fusion network BiFPN is introduced to replace the PANet network, and the feature weights are adaptively adjusted to obtain more context information and global information, improving the information utilization rate. Second, a dual-branch attention enhancement module is proposed to fully strengthen the roles of spatial attention information, channel attention information, and position information, further enhancing the network's feature extraction ability. At the same time, the structure of the Feature Pyramid Networks (FPN) is modified, and the proposed dual-branch attention enhancement module is placed before the BiFPN network, and an attention mechanism is used between each layer of the backbone network and the BiFPN network. Finally, a rotated detection box is used to estimate the direction of the target, reducing the overlapping problem of detected objects caused by horizontal detection boxes, making the detection box more accurately locate the target, and being more conducive to the detection of densely arranged targets.

[0029] Please refer to Figure 1 as shown in Figure 1 Figure 1 is a structural diagram of a method for detecting arbitrary-direction dense ship targets based on attention enhancement provided by an embodiment of the present invention. A method for detecting arbitrary-direction dense ship targets based on attention enhancement provided by the present invention includes:

[0030] S101. Obtain the original image.

[0031] Specifically, in this embodiment, the original image is obtained by a Synthetic Aperture Radar (SAR). The obtained SAR image has a relatively complex scene, especially in the nearshore area, there are a large number of clutter false alarms, and the false alarm rate of ship detection is high.

[0032] S102. Based on the backbone feature extraction module, extract multi-scale features of the original image to obtain multiple feature maps.

[0033] Specifically, in this embodiment, the backbone feature extraction module undertakes the main task of extracting multi-scale features of the image. A backbone feature extraction module with good performance is crucial for the extraction effect of the detection target. Generally speaking, the depth of the backbone feature extraction network directly affects the performance of the model. However, as the network depth increases, the network performance degrades. Among them, the ResNet residual network alleviates the gradient disappearance and degradation problems brought by deep neural networks when the number of layers is too deep by using internal residual blocks with skip connections. The skip connection method introduced in its structure enables the information of the previous residual block to flow directly into the next residual block, improving the information flow. Please refer to Figure 2As shown Figure 2 is a schematic diagram of the ResNet residual block network structure provided by an embodiment of the present invention.

[0034] The residual block network structure can be divided into a downsampling residual block and a normal residual block; among them, the downsampling residual block corresponding to a stride of 2 is applicable to the case where the number of input and output channels is different, and a convolutional layer is added to the left bypass branch, and this convolutional layer plays a role in matching the dimensional difference between the input and output; the normal residual block corresponding to a stride of 1 is applicable to the case where the number of corresponding input and output channels is the same, and the input can be directly added to the output of the skip connection. In this embodiment, for the detection task of the ship dataset, the ResNet50 network is selected as the backbone feature extraction network. After passing through the backbone feature extraction network, 5 different sizes of feature maps are generated, and the C3 / C4 / C5 feature maps with downsampling strides of 8 / 16 / 32 are selected as the input feature layers of the next-level deep feature fusion network.

[0035] S103. Based on the dual-branch attention enhancement module, learn the importance degrees of each channel and each space in at least part of the feature maps respectively, and learn the importance degrees of each position information in at least part of the feature maps, obtain a first-branch attention feature map and a second-branch attention feature map, and merge the first-branch attention feature map and the second-branch attention feature map to obtain an output feature map based on the dual-branch attention enhancement module.

[0036] Specifically, in this embodiment, in order to make the target detection network pay more attention to useful feature information and suppress invalid features and noises, a dual-branch attention enhancement module is proposed, which respectively includes a first-branch attention enhancement module that combines channel and spatial information and a second-branch attention enhancement module that captures direction and position perception; the first-branch attention enhancement module includes a channel attention model and a spatial attention model, which learn the importance degrees of different channels and different spaces of the feature maps. In addition, considering that in the implementation of the target detection algorithm, the network needs to focus more on the position area where the target is located, an additional coordinate attention model included in the second-branch attention enhancement module is added to capture the information of direction and position perception; compared with the single-branch attention enhancement module, using the dual-branch attention enhancement module can capture more feature information of the target, further strengthening the roles of spatial attention, channel attention information, and position information, so that the network learns multi-faceted feature information of the target. Please refer to Figure 3 As shown Figure 3 is a schematic diagram of the dual-branch attention enhancement module provided by an embodiment of the present invention.

[0037] (1) Channel attention module

[0038] Specifically, please refer to Figure 4 As shown Figure 4It is a schematic diagram of the channel attention module provided by an embodiment of the present invention. In this embodiment, inspired by the CBAM attention mechanism, using convolution to replace the fully connected layer in the CBAM channel attention mechanism will have better cross-channel information acquisition ability; in addition, global average pooling (GlobalAvgpool) is used to aggregate context information, and global max pooling (Global Maxpool) is used to eliminate useless information in the feature map.

[0039] At least part of the feature map F is processed by the first global max pooling layer to obtain the max pooling feature map F max ; at least part of the feature map F is processed by the first global average pooling layer to obtain the average pooling feature map F avg ;

[0040] The max pooling feature map F max is processed by the first dynamic convolution layer to aggregate the in-channel neighborhood information of the max pooling feature map F max ; the average pooling feature map F avg is processed by the second dynamic convolution layer to aggregate the in-channel neighborhood information of the average pooling feature map F avg ;

[0041] The features in the max pooling feature map F max after being processed by the first dynamic convolution layer are added to the features in the average pooling feature map F avg after being processed by the second dynamic convolution layer one by one, and the addition result is processed by the first Sigmoid activation function to obtain the channel attention feature map M c ;

[0042] The features in at least part of the feature map F are multiplied by the features in the channel attention feature map M c one by one to obtain the channel attention generation result F c .

[0043] Among them, the expression of the average pooling feature map F avg is:

[0044] F avg = AdaptiveAvgPool(F);

[0045] The expression of the max pooling feature map F max is:

[0046] F max = AdaptiveMaxPool(F);

[0047] The expression of the channel attention feature map M c is:

[0048]

[0049]

[0050] Channel attention generation result F c The expression of is as follows:

[0051]

[0052] In the above formula, AdaptiveAvgPool is an adaptive average pooling kernel, and AdaptiveMaxPool is an adaptive maximum pooling kernel. is a one-dimensional dynamic convolution kernel with a convolution kernel size of k, σ is the first Sigmoid activation function. is element-wise multiplication, C is the number of feature channels of the feature map, and odd is the odd value closest to the result.

[0053] It should be noted that the sizes of the first dynamic convolution layer and the second dynamic convolution layer are k, and 1*1 convolution is performed using the dynamic convolution kernel.

[0054] (2) Spatial attention model

[0055] Specifically, please refer to Figure 5 as shown in Figure 5 is a schematic diagram of the spatial attention model provided by the embodiment of the present invention. In this embodiment, the channel attention generation result F c is processed through the second max pooling layer to obtain the max pooling feature map F1'; the channel attention generation result F c is processed through the second average pooling layer to obtain the average pooling feature map F2'.

[0056] The max pooling feature map F1' and the average pooling feature map F2' are concatenated by channel to obtain the feature map F', and after the feature map F' is processed through the dilated convolution layer, it is then processed by the second Sigmoid activation function to obtain the feature weights M of each pixel point in the feature map F c ; s ;

[0057] The feature weights M of each pixel point in the feature map F c are multiplied by the channel attention generation result F s to obtain the first branch attention feature map F1. c ;

[0058] Among them, the expression of the feature map F' is:

[0059] F' = concat[AvgPool(F c ) ; MaxPool(F c )];

[0060] The feature weight M of each pixel point s The expression is as follows:

[0061]

[0062] The expression of the first branch attention feature map F1 is as follows:

[0063]

[0064] In the above formula, AvgPool is average pooling, MaxPool is max pooling, and concat is data splicing processing by channel. has a size of 3×3, and the dilation rate of the dilated convolution is 3.

[0065] It should be noted that the size of the dilated convolution layer is 3×3.

[0066] (3) Coordinate attention model

[0067] Specifically, please refer to Figure 6 as shown in Figure 6 which is a schematic diagram of the coordinate attention model provided by the embodiment of the present invention. In this embodiment, considering that the original CA (Coordinate Attention, CA) does not use global max pooling to eliminate useless information, therefore, global max pooling processing is added on the basis of the CA attention mechanism to balance the feature information of the SAR image.

[0068] At least part of the feature map F is processed through a global average pooling layer in the horizontal direction and a global average pooling layer in the vertical direction to obtain an average pooling result in the horizontal direction and an average pooling result in the vertical direction; at least part of the feature map F is processed through a global max pooling layer in the horizontal direction and a global max pooling layer in the vertical direction to obtain a max pooling result in the horizontal direction and a max pooling result in the vertical direction;

[0069] Among them, the expressions of the average pooling result in the horizontal direction, the max pooling result, the average pooling result in the vertical direction, and the max pooling result are respectively:

[0070]

[0071]

[0072]

[0073]

[0074] Among them, x c is the feature map related to the C-th channel, is the average pooling result in the horizontal direction, is the max pooling result in the horizontal direction, is the average pooling result in the vertical direction, is the max pooling result in the vertical direction.

[0075] The average pooling result in the horizontal direction and the average pooling result in the vertical direction are concatenated, and are sequentially processed through the first convolutional layer and the third non-linear activation function to obtain the feature map F A ; The max pooling result in the horizontal direction and the max pooling result in the vertical direction are concatenated, and are sequentially processed through the second convolutional layer and the fourth non-linear activation function to obtain the feature map F M ;

[0076] Among them, the feature map F A has the following expression:

[0077]

[0078] The feature map F M has the following expression:

[0079]

[0080] Among them, the sizes of the third convolutional layer, the fourth convolutional layer, the fifth convolutional layer, and the sixth convolutional layer are 1×1, that is, C 1×1 is a 1×1 convolutional process, and δ is a non-linear activation function;

[0081] The feature map F A is separated into a first horizontal direction feature map and a first vertical direction feature map; The feature map F M is separated into a second horizontal direction feature map and a second vertical direction feature map;

[0082] The first horizontal direction feature map is sequentially processed through the third convolutional layer and the fifth Sigmoid activation function to obtain the first attention weight The first vertical direction feature map is sequentially processed through the fourth convolutional layer and the sixth Sigmoid activation function to obtain the second attention weight The second horizontal direction feature map is sequentially processed through the fifth convolutional layer and the seventh Sigmoid activation function to obtain the third attention weight The second vertical direction feature map is sequentially processed through the sixth convolutional layer and the eighth Sigmoid activation function to obtain the fourth attention weight

[0083] The expression of the first attention weight is:

[0084]

[0085] The expression of the second attention weight is as follows:

[0086]

[0087] The expression of the third attention weight is as follows:

[0088]

[0089] The expression of the fourth attention weight is as follows:

[0090]

[0091] Add the first attention weight and the third attention weight to obtain the first result; add the second attention weight and the fourth attention weight to obtain the second result; multiply the first result, the second result and the feature map F to obtain the second branch attention feature F2;

[0092] Among them, the expression of the second branch attention feature is as follows:

[0093]

[0094] The output result of the double-branch attention enhancement module is as follows:

[0095] F out = F1 + F2.

[0096] S104. Based on the weighted bidirectional feature fusion network, through cross-scale connection operations, screen and fuse the information of the output feature map to obtain an enhanced feature map.

[0097] Specifically, please refer to Figure 7 as shown in Figure 7 which is a schematic diagram of the BiFPN network structure provided by an embodiment of the present invention. In this embodiment, feature representations at different levels represent different semantic information, and their contributions to the network are inconsistent. Considering this, this embodiment introduces a weighted bidirectional feature fusion network BiFPN to replace the original PANet to achieve screening and fusion of feature information. Through cross-scale connection operations, low-level detail information is more easily mapped to high-level semantic information, enhancing feature propagation and reuse.

[0098] In this embodiment, taking node 6 as an example, the expression of the fusion structure is as follows:

[0099]

[0100]

[0101] Among them, is the intermediate feature of the top-down path, is the output feature of the bottom-up path, ω i and ω i ' are adaptive weights, Conv is the convolution operation, Resize is the upsampling or downsampling operation, and ε is a very small value.

[0102] S105. Detect the enhanced feature map based on the classification structure and the bounding box regression structure in the detector.

[0103] Specifically, in this embodiment, most detectors are composed of a classification branch, a bounding box regression branch, and a quality evaluation branch in the final prediction stage. However, this embodiment may lead to inconsistency between the training and testing stages, resulting in negative samples with low classification scores being ranked in front of certain positive samples. Considering the above problems, a classification structure and a bounding box regression structure are used for detector design.

[0104] The loss function of the detector is:

[0105]

[0106] Among them, c x,y is the classification score, is the true class label of the target, t x,y =(x c ,y c ,w,h,t) is the predicted bounding box position by the bounding box regression structure, is the true bounding box position, N pos is the number of positive samples, λ is a hyperparameter, (x,y)∈pos means this sample point is a positive sample, L cls is the loss function of the classification structure, L reg is the loss function of the bounding box regression structure.

[0107] The expression of the loss function of the classification structure is:

[0108]

[0109] Among them, α is the balance factor, which controls the weight of positive samples in the overall loss, γ is the modulation coefficient, and (x,y)∈neg means this sample point is a positive sample.

[0110] The expression of the loss function of the bounding box regression structure is:

[0111]

[0112] Among them, r is the difference between the predicted bounding box position by the bounding box regression structure and the true bounding box position.

[0113] In summary, the method for detecting arbitrary - direction dense ship targets based on attention enhancement provided in this embodiment aims to solve problems such as poor detection performance, missed detection, and false detection in complex coastal scenarios. At the same time, it uses a rotated detection box instead of a horizontal detection box to achieve direction estimation of the target while effectively distinguishing the target area from the background area; it uses a weighted bidirectional feature fusion network BiFPN to replace the PANet network, and adaptively adjusts the feature weights to obtain more context information and global information, improving the information utilization rate; it uses a double - branch attention enhancement module to fully strengthen the roles of spatial attention information, channel attention information, and position information, and enhance the network's ability to extract important features.

[0114] In an optional embodiment of the present invention, the following simulation experiments are carried out for verification.

[0115] (1) Experimental data and parameters

[0116] The simulation experiment of the present invention is implemented using the publicly available SSDD + dataset in China. The SSDD + dataset changes the target annotation box from a horizontal border to a rotated border on the basis of the SSDD dataset, which is convenient for performing rotated target detection tasks. The dataset is divided into a training set and a test set according to the 8:2 division rule. The algorithm is implemented based on the pytorch target detection framework, uses the Adam random gradient descent algorithm as the optimizer, the maximum number of training iterations Max iteration = 800, the initial learning rate lr = 0.0001, and the experiment runs on a computer equipped with an NVIDIA T40c GPU. At the same time, the model is pre - trained on the optical dataset ImageNet.

[0117] (2) Comparison of detection performance

[0118] To verify the effectiveness of the proposed ship target detection method, some classic target detection network frameworks are selected as comparative experiments for method verification, and the improved network is compared with the original RetinaNet network. The comparison results are shown in Table 1.

[0119] Table 1 Performance comparison between the method of the present invention and other methods

[0120] Index R3Det ReDet RetinaNet The method of the present invention mAP 80.75 85.42 83.01 87.59

[0121] It can be seen from Table 1 that the detection method proposed in the above embodiments has a relatively high AP value. Compared with the basic RetinaNet network, the average precision AP value has increased by 4.58%. This is due to the enhancement of the network feature extraction ability by the weighted bidirectional feature fusion network and the improved attention module. Thus, it can be seen that the detection method proposed in the above embodiments has obvious detection performance advantages. In order to further intuitively display the improvement effect of the proposed detection algorithm, several test images are selected to verify the RetinaNet method and the improved method of this paper. Please refer to Figure 8 (b) to 8(c) as shown. Figure 8 (a) is the result map of the true annotation position of the ship target provided by the embodiment of the present invention. Figure 8 (b) is a result map of the experimental comparison of the detection performance provided by the embodiment of the present invention. Figure 8 (c) is the experimental result map of the detection performance provided by the embodiment of the present invention.

[0122] Please continue to refer to Figure 8 (a) to Figure 8 (c) as shown. Targets in three cases, namely small-scale targets in the open sea, large-scale targets with variable berthing angles, and dense targets near the shore, are selected. From Figure 8 (a) to Figure 8 (c), it can be seen that for the detection of small targets in the open sea, the method proposed in the above embodiments maintains the good detection performance of the original detection method; for large-scale targets with variable berthing angles, the RetinaNet method has a missed alarm situation in the detection result, while the method proposed in the present invention effectively solves the missed alarm problem, proving that the method proposed in the present invention has strong feature extraction ability; for dense targets near the shore, although both methods have a certain degree of missed detection, the method proposed in the present invention effectively reduces the detection false alarm and detects more ship targets, further indicating that the method proposed in this embodiment can make the network pay more attention to useful target information while effectively suppressing the clutter background, improving the detection rate.

[0123] Compared with the classical RetinaNet method, the average precision AP value of the model detection proposed in the present invention has increased by 4.55%. While enhancing the target features, it suppresses the background clutter, effectively improving the SAR ship detection performance in complex scenes, making the detection of ship targets in SAR images more robust and accurate.

[0124] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant are intended to cover non-exclusive inclusion, so that an article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the article or device comprising said element. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The orientation or positional relationship indicated by "above", "below", "left", "right", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation on the present invention.

[0125] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.

[0126] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. An arbitrary-direction dense ship target detection method based on attention enhancement, characterized in that Including: Obtain the original image; Based on the backbone feature extraction module, extract the multi-scale features of the original image to obtain multiple feature maps; Based on the dual-branch attention enhancement module, respectively learn the importance of each channel and each space in at least part of the feature maps, as well as learn the importance of each position information in at least part of the feature maps, to obtain the first-branch attention feature map and the second-branch attention feature map, and merge the first-branch attention feature map and the second-branch attention feature map to obtain the output feature map based on the dual-branch attention enhancement module; The dual-branch attention enhancement module includes a first-branch attention enhancement module that combines channel and spatial information, and a second-branch attention enhancement module that captures direction and position perception; wherein, the first-branch attention enhancement module includes a channel attention model and a spatial attention model; the second-branch attention enhancement module includes a coordinate attention model; Based on the weighted bidirectional feature fusion network, through cross-scale connection operations, screen and fuse the information of the output feature map to obtain an enhanced feature map; Based on the classification structure and bounding box regression structure in the detector, detect the enhanced feature map.

2. The method for detecting arbitrary-direction dense ship targets based on attention enhancement according to claim 1, wherein The backbone feature extraction module is a ResNet50 residual network, and the residual blocks of the ResNet50 residual network include downsampling residual blocks and ordinary residual blocks.

3. The method for detecting arbitrary - direction dense ship targets based on attention enhancement according to claim 1, wherein The channel attention model includes a first global max pooling layer, a first global average pooling layer, a first dynamic convolution layer, a second dynamic convolution layer, and a first Sigmoid activation function; The process of learning the importance of each channel in at least part of the feature maps includes: At least a part of the feature map is processed by the first global max pooling layer to obtain a max pooling feature map ; at least a part of the feature map is processed by the first global average pooling layer to obtain an average pooling feature map ; The maximum pooling feature map is processed by the first dynamic convolution layer to aggregate the maximum pooling feature map intra-channel neighborhood information; the average pooling feature map is processed by the second dynamic convolution layer to aggregate the average pooling feature map intra-channel neighborhood information; The features in the max-pooling feature map processed by the first dynamic convolution layer are added to the features in the average-pooling feature map processed by the second dynamic convolution layer one by one, and the addition result is processed by the first Sigmoid activation function to obtain a channel attention feature map ; Multiply at least some of the features in the feature map by the features in the channel attention feature map one by one to obtain the channel attention generation result .

4. The method for detecting arbitrary - direction dense ship targets based on attention enhancement according to claim 3, characterized in that, The spatial attention model includes a second max pooling layer, a second average pooling layer, an atrous convolution layer, and a second Sigmoid activation function; The process of learning the importance of each space in at least part of the feature maps includes: The channel attention generation result is processed by the second max pooling layer to obtain a max pooling feature map ; The channel attention generation result is processed by the second average pooling layer to obtain an average pooling feature map ; The maximum pooling feature map and the average pooling feature map are subjected to data splicing processing by channel to obtain a feature map , and after the feature map is processed by the dilated convolutional layer and then processed by the second Sigmoid activation function, the feature weights of each pixel point in the feature map are obtained ; Multiply the feature weights of each pixel point in the feature map by the channel attention generation result to obtain the first branch attention feature map .

5. The method for detecting arbitrary-direction dense ship targets based on attention enhancement according to claim 1, characterized in that The coordinate attention model includes a global average pooling layer and a global max pooling layer in the horizontal direction, and a global average pooling layer and a global max pooling layer in the vertical direction, a first convolution layer, a second convolution layer, a third non-linear activation function, a fourth non-linear activation function, a third convolution layer, a fourth convolution layer, a fifth convolution layer, a sixth convolution layer, a fifth Sigmoid activation function, a sixth Sigmoid activation function, a seventh Sigmoid activation function, and an eighth Sigmoid activation function; The process of learning the importance of each position information in at least part of the feature maps includes: At least part of the feature maps are processed by a global average pooling layer in the horizontal direction and a global average pooling layer in the vertical direction to obtain an average pooling result in the horizontal direction and an average pooling result in the vertical direction; at least part of the feature maps are processed by a global max pooling layer in the horizontal direction and a global max pooling layer in the vertical direction to obtain a max pooling result in the horizontal direction and a max pooling result in the vertical direction; Concatenate the average pooling result in the horizontal direction and the average pooling result in the vertical direction, and sequentially process them through the first convolutional layer and the third non-linear activation function to obtain a feature map ; Concatenate the max pooling result in the horizontal direction and the max pooling result in the vertical direction, and sequentially process them through the second convolutional layer and the fourth non-linear activation function to obtain a feature map ; Separate the feature map into a first horizontal direction feature map and a first vertical direction feature map ; Separate the feature map into a second horizontal direction feature map and a second vertical direction feature map ; The first horizontal direction feature map is successively processed by the third convolutional layer and the fifth Sigmoid activation function to obtain the first attention weight ; the first vertical direction feature map is successively processed by the fourth convolutional layer and the sixth Sigmoid activation function to obtain the second attention weight ; the second horizontal direction feature map is successively processed by the fifth convolutional layer and the seventh Sigmoid activation function to obtain the third attention weight ; the second vertical direction feature map is successively processed by the sixth convolutional layer and the eighth Sigmoid activation function to obtain the fourth attention weight ; Add the first attention weight to the third attention weight to obtain a first result; add the second attention weight to the fourth attention weight to obtain a second result; multiply the first result, the second result, and the feature map to obtain the second branch attention feature .

6. The method for detecting arbitrary-direction dense ship targets based on attention enhancement according to claim 1, characterized in that The loss function of the detector is: ; Among them, is the classification score, is the true class label of the target, is the predicted bounding box position by the bounding box regression structure, is the true bounding box position, is the number of positive samples, is a hyperparameter, indicates that this sample point is a positive sample, is the classification structure loss function, is the bounding box regression structure loss function.

7. The method for detecting arbitrary-direction dense ship targets based on attention enhancement according to claim 6, wherein The expression of the loss function of the classification structure is: ; Among them, is the balance factor, which controls the weight of the positive samples in the overall loss. is the modulation coefficient. indicates that this sample point is a positive sample.

8. The method for detecting arbitrary - direction dense ship targets based on attention enhancement according to claim 6, wherein, The expression of the loss function of the bounding box regression structure is: ; Among them, is the difference between the predicted border position of the bounding box regression structure and the true border position.

Citation Information

Patent Citations

  • Aerial image multi-scale target detection method based on spatial pyramid attention driving

    CN111401201A

  • Remote sensing ship target detection method based on deformation attention pyramid

    CN115115601A