SAR Ship Detection Method Based on Coordinate Attention and Long-Short Range Context

By adopting detection methods based on coordinate attention and long and short-range context in SAR ship detection, the problems of noise interference and difficulty in detection of small targets are solved, and higher detection performance and accuracy are achieved.

CN115147720BActive Publication Date: 2025-06-17CHONGQING INNOVATION CENTER OF BEIJING INSTITUTE OF TECHNOLOGY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210718888.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-23
Publication Date
2025-06-17
Estimated Expiration
2042-06-23

AI Technical Summary

Technical Problem

The prior art has noise interference and small target detection in SAR ship detection, resulting in unsatisfactory detection performance.

Method used

The detection method based on coordinate attention and long and short-distance context is adopted, and the detection performance is improved through feature extraction network, long and short-distance context collaborative extraction network and PAN pyramid feature fusion network, combined with YOLOX anchor-free frame decoupling detection head.

Benefits of technology

Effectively suppress noise interference, improve the accuracy and performance of small-object detection, and improve the detection accuracy and generalization performance of the detection model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115147720B_ABST
    Figure CN115147720B_ABST
Patent Text Reader

Abstract

The present invention provides a SAR ship detection method based on coordinate attention and long- and short-range context, including: obtaining a SAR ship image to be detected, where the SAR ship image to be detected contains corresponding ships; constructing a feature extraction network according to the coordinate attention mechanism, inputting the SAR ship image to be detected into the feature extraction network to obtain a feature map enhanced by coordinate attention; constructing a long- and short-range context collaborative extraction network according to the long- and short-range context information, inputting the feature map enhanced by coordinate attention into the long- and short-range context collaborative extraction network to obtain a feature map strengthened by context; performing feature fusion on the feature map strengthened by context through a PAN pyramid feature fusion network to obtain a fused feature map; inputting the fused feature map into a YOLOX anchor-free decoupled detection head to obtain the ship position and ship category. The present invention can alleviate image noise interference and can accurately detect small targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of radar ship detection, and in particular to a SAR ship detection method based on coordinate attention and long and short distance context. Background Art

[0002] Synthetic Aperture Radar (SAR) has the imaging characteristics of all-weather and all-day. With the development of airborne and spaceborne satellites in recent years, SAR has been widely used in military and civilian fields. SAR ship detection, as a basic maritime task, has important value in aspects such as maritime traffic control, fishery management, and maritime emergency rescue. Target detection is an important link in the SAR ship detection task. Through a series of algorithms, ship targets on the image will be automatically located and identified, and the target detection performance is the key factor affecting the ship detection accuracy.

[0003] Due to the unique imaging mechanism of SAR, there are many speckle noises in the image, making it difficult to distinguish the target from the background and noise. Therefore, compared with optical remote sensing images, the processing of SAR images is more difficult. Because the resolution of SAR images is relatively low, the ship target scale is small, and there is little visual information, it is difficult to extract discriminative features. Moreover, the boundary is blurred and it is easily affected by environmental factors, resulting in difficulty for the detection model to accurately locate and identify. Existing methods such as Feature Pyramid Network (FPN) and Path Aggregation Network (PAN) fuse the feature maps of adjacent levels from bottom to top and from top to bottom, combining the rich semantic information in the deep feature maps with the spatial fine-grained features in the shallow feature maps to highlight the semantic characteristics of small targets in the high-resolution feature maps. However, this method still cannot avoid the loss of information features caused by multiple downsamplings of small targets during the feature extraction process. In addition, existing methods mostly use the anchor box mechanism to predict the target position. A large number of redundant anchor boxes are densely distributed in the input image, which not only brings a heavy burden to the network calculation, but also the artificially set hyperparameters may lead to difficult model convergence, resulting in unsatisfactory detection performance.

[0004] In summary, when the existing technology is used for SAR ship detection, there are problems such as it is difficult to accurately locate and identify ship targets in SAR images due to speckle noises and small scales; it is impossible to avoid the information loss caused by multiple downsamplings, resulting in difficulty in improving the detection performance; the anchor box mechanism not only increases the computational redundancy, but also makes the training process more complex.

[0005] Therefore, there is an urgent need for a SAR ship detection method that can alleviate image noise interference and accurately detect small targets. Summary of the Invention

[0006] Based on this, it is necessary to provide a SAR ship detection method based on coordinate attention and long- and short-range context for the above technical problems.

[0007] A SAR ship detection method based on coordinate attention and long- and short-range context includes the following steps: obtaining a SAR ship image to be measured, where the SAR ship image to be measured contains corresponding ships; constructing a feature extraction network according to the coordinate attention mechanism, inputting the SAR ship image to be measured into the feature extraction network, and obtaining a feature map enhanced by coordinate attention; constructing a long- and short-range context collaborative extraction network according to the long- and short-range context information, inputting the feature map enhanced by coordinate attention into the long- and short-range context collaborative extraction network, and obtaining a feature map enhanced by context; performing feature fusion on the feature map enhanced by context through a PAN pyramid feature fusion network to obtain a fused feature map; and inputting the fused feature map into a YOLOX anchor-free decoupled detection head to obtain the ship position and ship category.

[0008] In one embodiment, the step of constructing a feature extraction network according to the coordinate attention mechanism, inputting the SAR ship image to be measured into the feature extraction network, and obtaining a feature map enhanced by coordinate attention specifically includes: constructing a feature extraction network according to the coordinate attention mechanism and inputting the SAR ship image to be measured into the feature extraction network; in the feature extraction network, performing a two-fold downsampling on the SAR ship image to be measured through a convolution operation with a convolution kernel size of 3×3 and a stride of 2, halving the dimension of the downsampled image through two convolution operations with a convolution kernel size of 1×1 respectively to obtain an initial feature map, and respectively sending it into a first processing path and a second processing path; in the first processing path, introducing the initial feature map into a coordinate attention residual block to obtain a first feature map; in the second processing path, denoting the initial feature map as a second feature map; stacking the first feature map and the second feature map channel-wise, and performing a convolution operation with a convolution kernel size of 1×1 on the stacked feature map to refine the features, so as to obtain feature maps enhanced by coordinate attention at different levels.

[0009] In one embodiment, the step of introducing the initial feature map into a coordinate attention residual block to obtain a first feature map in the first processing path specifically includes: reducing the number of channels of the initial feature map through a 1×1 convolution; performing depth convolution on the feature map with reduced channels through a convolution operation with a convolution kernel size of 3×3; inputting the depth-convolved feature map into a coordinate attention module, and restoring the number of channels through a 1×1 convolution to obtain a restored feature map; and adding the restored feature map and the initial feature map element-wise to obtain a first feature map.

[0010] In one embodiment, the long - short distance context collaborative extraction network is constructed based on the long - short distance context information, and the feature map enhanced by coordinate attention is input into the long - short distance context collaborative extraction network to obtain the context - enhanced feature map. Specifically, it includes: constructing the long - short distance context collaborative extraction network according to the long - short distance context information; inputting the feature map enhanced by coordinate attention into the long - short distance context collaborative extraction network; in the long - short distance context collaborative extraction network, inputting the feature map enhanced by coordinate attention into two parallel non - linear calculation modules, where the non - linear calculation module includes a long - distance context module and a short - distance context module, to obtain a long - distance context feature map and a short - distance context feature map; splicing the long - distance context feature map and the short - distance context feature map according to the way of interlacing in corresponding channels, and fusing each pair of adjacent spliced long - distance context feature maps and short - distance context feature maps into one feature map through a 1×1 grouped convolution; mapping the fused feature map to the range between 0 and 1 through the Sigmoid function to obtain the long - short distance hybrid context weight map; summing the long - short distance hybrid context weight map and the feature map enhanced by coordinate attention to obtain the context - enhanced feature map.

[0011] In one embodiment, the long - distance context is captured by a dilated depth convolution with a kernel size of 5×5 and a dilation rate of 5 and a 1×1 depth convolution; the short - distance context is captured by a 1×1 depth convolution and a dilated depth convolution with a kernel size of 3×3 and a dilation rate of 3.

[0012] In one embodiment, the fusion of the context - enhanced feature map through the PAN pyramid feature fusion network specifically includes: sending the context - enhanced feature map into the PAN pyramid feature fusion network, and refining the position information and semantic information of the context - enhanced feature map through the bottom - up and top - down information flow to obtain the fused feature map.

[0013] In one embodiment, the input of the fused feature map into the YOLOX anchor - free decoupled detection head to obtain the ship position and ship category specifically includes: inputting the fused feature map into the YOLOX anchor - free decoupled detection head to obtain the target classification feature map, the target box position regression feature map, and the target box confidence regression feature map; obtaining the ship position and ship category according to the target classification feature map, the target box position regression feature map, and the target box confidence regression feature map.

[0014] Compared with the prior art, the advantages and beneficial effects of the present invention are as follows: By obtaining the SAR ship image to be measured, and the corresponding ship is included in the SAR image to be measured, a feature extraction network is constructed according to the coordinate attention mechanism, and the SAR ship image to be measured is input into the feature extraction network to obtain a feature map enhanced by coordinate attention, so as to strengthen the focusing ability on small targets and suppress the interference of background noise; A long-short distance context collaborative extraction network is constructed according to the long-short distance context information, and the feature map enhanced by coordinate attention is input into the long-short distance context collaborative extraction network to obtain a feature map enhanced by context, which can simultaneously collect environmental information in cross-regions and adjacent regions, enrich the significant features of small targets, and improve the detection performance of small targets; Through the PAN pyramid feature fusion network, the feature map enhanced by context is fused to obtain a fused feature map, which can simultaneously perform cross-level transfer fusion of position information and semantic information, enriching the feature expression of small targets; The fused feature map is input into the YOLOX anchor-free decoupled detection head to obtain the ship position and ship category, improving the ship target detection performance of the SAR image, and improving the detection accuracy and generalization performance of the detection model. Description of the Drawings

[0015] Figure 1 It is a schematic flow chart of a SAR ship detection method based on coordinate attention and long-short distance context in one embodiment;

[0016] Figure 2 It is a schematic network structure diagram of a SAR ship detection method based on coordinate attention and long-short distance context in one embodiment;

[0017] Figure 3 It is a schematic principle diagram of a feature extraction network in one embodiment;

[0018] Figure 4 It is a schematic principle diagram of a long-short distance context collaborative extraction network in one embodiment. Detailed Embodiments

[0019] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below through specific embodiments in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0020] In one embodiment, as Figures 1 to 4 shown, a SAR ship detection method based on coordinate attention and long-short distance context is provided, including the following steps:

[0021] Step S101, obtain the SAR ship image to be measured, and the corresponding ship is included in the SAR ship image to be measured.

[0022] Specifically, an SAR image to be measured is obtained by a synthetic aperture radar, and the corresponding ship is included in the SAR ship image to be measured.

[0023] Step S102, construct a feature extraction network according to the coordinate attention mechanism, input the SAR ship image to be measured into the feature extraction network, and obtain a feature map enhanced by coordinate attention.

[0024] Specifically, construct a feature extraction network according to the coordinate attention mechanism, reconstruct the residual structure in the feature extraction network, adopt depthwise separable convolution to reduce the computational amount, input the SAR ship image to be measured into the feature extraction network, perform feature extraction from shallow to deep, and obtain feature maps enhanced by coordinate attention at different levels, so as to strengthen the focusing ability on small targets while reducing the computational amount and suppressing the interference of background noise.

[0025] Among them, the coordinate attention mechanism decomposes the channel attention into two 1D feature encoding processes. These two processes respectively aggregate features along two spatial directions. In this way, the coordinate attention can capture long-range correlations along one spatial direction, and at the same time can retain accurate position information along the other spatial direction. The obtained feature maps are respectively encoded into a pair of direction-aware and position-sensitive attention maps, and the attention maps can be complementarily applied to the input feature maps to enhance the representation of the object of interest, that is, the ship.

[0026] Step S103, construct a long-short distance context collaborative extraction network according to the long-short distance context information, input the feature map enhanced by coordinate attention into the long-short distance context collaborative extraction network, and obtain a feature map strengthened by context.

[0027] Specifically, construct according to the long-short distance context information. Through the long-short distance context collaborative extraction network, perform feature extraction on the feature maps enhanced by coordinate attention at different levels, and obtain a feature map strengthened by context, which can utilize receptive fields of different sizes, collect environmental information in both cross-regional and adjacent regions at the same time, enrich the significant features of small targets, and thus improve the detection performance of small targets.

[0028] Among them, in the long-short distance context collaborative extraction network, two different-sized receptive fields are implemented by adopting a pair of dilated convolutions with different dilation rates, respectively extract the cross-regional and adjacent environmental features of the target, and perform effective fusion. Among them, the receptive field refers to the size of the area on the input image to which the pixel points on the feature map output by each layer of the convolutional neural network are mapped back, that is, the size of a point on the feature map relative to the original image, and it is also the area of the input image that the features of the convolutional neural network can see.

[0029] Step S104, perform feature fusion on the feature map strengthened by context through the PAN pyramid feature fusion network to obtain a fused feature map.

[0030] Specifically, through the PAN pyramid feature fusion network, using bottom-up and top-down information flows, the rich semantic features and high-resolution position features in the context-enhanced feature map are fully fused to obtain the fused feature map. In the PAN pyramid feature fusion network, top-down feature fusion is performed first, and then bottom-up feature fusion is carried out, reducing the number of feature maps that the features need to pass through, thereby improving the feature fusion speed and having a good fusion effect.

[0031] Step S105: Input the fused feature map into the YOLOX anchor-free decoupled detection head to obtain the ship position and ship category.

[0032] Specifically, the YOLOX anchor-free decoupled detection head divides the task into two subtasks, including a classification subtask and a localization subtask. After inputting the fused feature map into the YOLOX anchor-free decoupled detection head, target localization and target classification are respectively performed to obtain the ship position and ship category in the SAR image, thereby improving the detection effect.

[0033] In this embodiment, by obtaining the SAR ship image to be measured, and the SAR image to be measured contains the corresponding ship, a feature extraction network is constructed according to the coordinate attention mechanism, and the SAR ship image to be measured is input into the feature extraction network to obtain a feature map enhanced by coordinate attention, thereby strengthening the focusing ability on small targets and suppressing the interference of background noise; a long-short distance context collaborative extraction network is constructed according to the long-short distance context information, and the feature map enhanced by coordinate attention is input into the long-short distance context collaborative extraction network to obtain a context-enhanced feature map, which can simultaneously collect the environmental information of cross-regions and adjacent regions, enrich the significant features of small targets, and improve the detection performance of small targets; through the PAN pyramid feature fusion network, feature fusion is performed on the context-enhanced feature map to obtain the fused feature map, which can simultaneously perform cross-level transfer fusion of position information and semantic information, enriching the feature expression of small targets; the fused feature map is input into the YOLOX anchor-free decoupled detection head to obtain the ship position and ship category, improving the ship target detection performance of the SAR image and the detection accuracy and generalization performance of the detection model.

[0034] Among them, step S102 specifically includes: constructing a feature extraction network according to the coordinate attention mechanism and inputting the SAR ship image to be measured into the feature extraction network; in the feature extraction network, performing a two-fold downsampling on the SAR ship image to be measured through a convolution operation with a convolution kernel size of 3×3 and a stride of 2, and respectively halving the dimensions of the downsampled image through two convolution operations with a convolution kernel size of 1×1 to obtain an initial feature map, and respectively sending it into a first processing path and a second processing path; in the first processing path, introducing the initial feature map into a coordinate attention residual block to obtain a first feature map; in the second processing path, denoting the initial feature as a second feature map; stacking the first feature map and the second feature map in channels, and passing through a convolution operation with a convolution kernel size of 1×1 to refine the stacked feature map to obtain a coordinate attention enhanced feature map at different levels.

[0035] As Figure 3 shown, input the SAR image to be measured into the feature extraction network, perform a two-fold sampling through a convolution operation with a convolution kernel size of 3×3 and a stride of 2, and then respectively halve the dimensions of the downsampled SAR image to be measured through two convolution operations with a convolution kernel size of 1×1 to obtain an initial feature map, and respectively send it into two different processing paths, namely the first processing path and the second processing path.

[0036] Among them, in the first processing path, the processing process of the initial feature map is: reducing the number of channels of the initial feature map through a 1×1 convolution; performing depth convolution on the feature map with reduced channels through a convolution operation with a convolution kernel size of 3×3; inputting the depth-convolved feature map into the coordinate attention module, restoring the number of channels through a 1×1 convolution to obtain a restored feature map; adding the restored feature map and the initial feature map element by element to obtain a first feature map.

[0037] Specifically, introduce the initial feature map A into the coordinate attention residual block, that is, first reduce the number of channels through a 1×1 convolution, then perform a 3×3 depth convolution, enter the coordinate attention calculation module to obtain an attention feature map, restore the number of channels of the attention feature map through a 1×1 convolution, and finally add the obtained feature map and the initial feature map A element by element to obtain a first feature map, completing the calculation of the coordinate attention calculation residual block.

[0038] In the second processing path, denote the initial feature map B as the second feature map.

[0039] Finally, stack the first feature map and the second feature map channel - by - channel, and then perform a convolution operation with a convolution kernel size of 1×1 on the stacked feature map to refine the features of the merged feature map, obtain feature maps with enhanced coordinate attention at different levels, thereby strengthening the focusing ability on small targets, alleviating the problem of information loss caused by multiple downsamplings, and being able to suppress the interference of background noise, and further improving the target detection effect.

[0040] Among them, step S103 specifically includes: constructing a long - short - range context collaborative extraction network according to the long - short - range context information; inputting the feature map with enhanced coordinate attention into the long - short - range context collaborative extraction network; in the long - short - range context collaborative extraction network, input the feature map with enhanced coordinate attention into two parallel non - linear calculation modules. The non - linear calculation module includes a long - range context module and a short - range context module to obtain a long - range context feature map and a short - range context feature map; according to the way of interleaving in corresponding channels, splice the long - range context feature map and the short - range context feature map, and through a 1×1 grouped convolution, fuse each pair of adjacent spliced long - range context feature maps and short - range context feature maps into one feature map; map the fused feature map to the range between 0 and 1 through the Sigmoid function to obtain a long - short - range mixed context weight map; sum the long - short - range mixed context weight map and the feature map with enhanced coordinate attention to obtain a context - enhanced feature map.

[0041] As Figure 4 shown, input the feature maps with enhanced coordinate attention at different levels into the long - short - range context collaborative extraction network, and send the input feature maps into two parallel non - linear calculation modules respectively. The two parallel non - linear calculation modules are a long - range context module and a short - range context module respectively, so as to obtain a long - range context feature map and a short - range context feature map.

[0042] Among them, the long - range context is captured by a dilated depth convolution with a convolution kernel size of 5×5 and a dilation rate of 5 and a 1×1 depth convolution; the short - range context is captured by a 1×1 depth convolution and a dilated depth convolution with a convolution kernel size of 3×3 and a dilation rate of 3.

[0043] After obtaining the long- and short-range context feature maps, the long-range context feature map and the short-range context feature map are concatenated according to the sequential interleaving of corresponding channels, and each pair of adjacent concatenated long-range context feature maps and short-range context feature maps are fused into a single feature map through a 1×1 grouped convolution; and the fused feature map is mapped to the range between 0 and 1 through the Sigmoid function to obtain the long- and short-range mixed context weight map; the long- and short-range mixed context weight map is summed with the corresponding coordinate attention enhanced feature map to obtain the context-reinforced feature map, achieving the simultaneous acquisition of cross-region and neighboring region environmental information, enriching the salient feature map of small targets, and improving the detection performance for small targets.

[0044] Among them, step S104 specifically includes: sending the context-reinforced feature map into the PAN pyramid feature fusion network, and refining the position information and semantic information of the context-reinforced feature map through the bottom-up and top-down information flows to obtain the fused feature map.

[0045] Specifically, sending the context-reinforced feature map into the PAN pyramid feature fusion network, and fully fusing the position information and semantic information of the context-reinforced feature map through the bottom-up and top-down information flows to obtain the fused feature map, realizing the cross-level transfer and fusion of position information and semantic information simultaneously, and enriching the feature expression of small targets.

[0046] Among them, step S105 specifically includes: inputting the fused feature map into the YOLOX anchor-free decoupled detection head to obtain the object classification feature map, the object box position regression feature map, and the object box confidence regression map; obtaining the ship position and ship class according to the object classification feature map, the object box position regression feature map, and the object box confidence regression map.

[0047] Specifically, after fusion, feature maps at different levels are obtained, and the fused feature maps are respectively sent into the YOLOX anchor-free decoupled detection head. Through the YOLOX anchor-free decoupled detection head, the object classification feature map, the object box position regression feature map, and the object box confidence regression map are obtained. According to the object classification feature map, the corresponding ship class information can be obtained, and according to the object box position regression feature map, the position information of the corresponding ship can be obtained. At the same time, the confidence of the output result can also be judged according to the object box confidence regression map, facilitating subsequent processing based on the ship position and ship classification, improving the ship target detection performance of the SAR image, and enhancing the detection accuracy and generalization performance.

[0048] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a computer storage medium (ROM / RAM, magnetic disk, optical disk) and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. Therefore, the present invention is not limited to any specific combination of hardware and software.

[0049] The above content is a further detailed description of the present invention in combination with specific implementation manners. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A SAR ship detection method based on coordinate attention and long- and short-range context, characterized in that, Including the following steps: Obtain a SAR ship image to be measured, where the SAR ship image to be measured contains a corresponding ship; Construct a feature extraction network according to the coordinate attention mechanism, and input the SAR ship image to be measured into the feature extraction network to obtain a feature map enhanced by coordinate attention, including: constructing a feature extraction network according to the coordinate attention mechanism, and inputting the SAR ship image to be measured into the feature extraction network; in the feature extraction network, perform a two-fold downsampling on the SAR ship image to be measured through a convolution operation with a convolution kernel size of 3×3 and a stride of 2, and halve the dimensions of the downsampled image through two convolution operations with a convolution kernel size of 1×1 respectively to obtain an initial feature map, and send them into a first processing path and a second processing path respectively; Among them, in the first processing path, introduce the initial feature map into a coordinate attention residual block to obtain a first feature map, including: reduce the number of channels of the initial feature map through a 1×1 convolution; perform depth convolution on the feature map with reduced channels through a convolution operation with a convolution kernel size of 3×3; input the depth-convolved feature map into a coordinate attention module, and restore the number of channels through a 1×1 convolution to obtain a restored feature map; add the restored feature map and the initial feature map element by element to obtain a first feature map; In the second processing path, record the initial feature map as a second feature map; stack the first feature map and the second feature map on the channels, and perform a convolution operation with a convolution kernel size of 1×1 on the stacked feature map to refine the features and obtain feature maps enhanced by coordinate attention at different levels; Construct a long-short distance context collaborative extraction network according to the long-short distance context information, and input the feature map enhanced by coordinate attention into the long-short distance context collaborative extraction network to obtain a context-reinforced feature map; Through a PAN pyramid feature fusion network, perform feature fusion on the context-reinforced feature map to obtain a fused feature map; Input the fused feature map into a YOLOX anchor-free decoupled detection head to obtain the ship position and ship category.

2. The SAR ship detection method based on coordinate attention and long- and short-range context according to claim 1, characterized in that, The step of constructing a long-short distance context collaborative extraction network according to the long-short distance context information, inputting the feature map enhanced by coordinate attention into the long-short distance context collaborative extraction network, and obtaining a context-reinforced feature map specifically includes: Construct a long-short distance context collaborative extraction network according to the long-short distance context information; Input the feature map enhanced by coordinate attention into the long-short distance context collaborative extraction network; In the long-short distance context collaborative extraction network, input the feature map enhanced by coordinate attention into two parallel non-linear calculation modules, where the non-linear calculation module includes a long-distance context module and a short-distance context module, to obtain a long-distance context feature map and a short-distance context feature map; According to the way of sequentially interspersing corresponding channels, splice the long-distance context feature map and the short-distance context feature map, and fuse each pair of adjacent spliced long-distance context feature maps and short-distance context feature maps into a feature map through a 1×1 grouped convolution; The fused feature map is mapped to the range of 0 to 1 through the Sigmoid function to obtain the long- and short-range hybrid context weight map; The long- and short-range hybrid context weight map and the feature map enhanced by coordinate attention are summed to obtain the context-reinforced feature map.

3. The SAR ship detection method based on coordinate attention and long- and short-range context according to claim 2, characterized in that, The long-range context is captured by a dilated depth convolution with a convolution kernel size of 5×5 and a dilation rate of 5 and a 1×1 depth convolution; the short-range context is captured by a 1×1 depth convolution and a dilated depth convolution with a convolution kernel size of 3×3 and a dilation rate of 3.

4. The SAR ship detection method based on coordinate attention and long- and short-range context according to claim 1, characterized in that, The context-reinforced feature map is fused through the PAN pyramid feature fusion network, which specifically includes: The context-reinforced feature map is fed into the PAN pyramid feature fusion network, and through the bottom-up and top-down information flow, the position information and semantic information of the context-reinforced feature map are refined to obtain the fused feature map.

5. The SAR ship detection method based on coordinate attention and long- and short-range context according to claim 1, characterized in that, The fused feature map is input into the YOLOX anchor-free decoupled detection head to obtain the ship position and ship class, which specifically includes: The fused feature map is input into the YOLOX anchor-free decoupled detection head to obtain the target classification feature map, the target box position regression feature map, and the target box confidence regression map; According to the target classification feature map, the target box position regression feature map, and the target box confidence regression map, the ship position and ship class are obtained.

Citation Information

Patent Citations

  • Improved YOLOv3 target detection method based on attention mechanism

    CN112508014A

  • Lightweight aircraft detection method based on improved Yolov4-tiny

    CN113780211A