SAR Image Target Detection Method Based on Feature Fusion and Cross-Layer Connection
By adopting feature fusion and cross-layer connection methods in SAR image object detection, the problem of low detection accuracy in the prior art is solved, and the object detection effect with higher accuracy and lower false alarm rate is achieved.
Patent Information
- Application Number
- CN202211446544.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-18
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-11-18
AI Technical Summary
When existing SAR image object detection algorithms face weather, environmental interference and increase in data volume, it is difficult to effectively detect targets, resulting in false alarms or missed detection, and the detection effect is poor.
SAR image object detection method based on feature fusion and cross-layer connection is adopted, feature maps of different scales are extracted through the downsampling module, and attention feature fusion and cross-scale feature fusion are used to generate more accurate object detection results.
It improves the detection accuracy and ability of small objects in SAR images, reduces the occurrence of false alarms and missed detection, and meets the higher detection accuracy requirements in practical applications.
Smart Images

Figure CN115713689B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of target detection, and particularly relates to a SAR image target detection method based on feature fusion and cross-layer connection. Background Art
[0002] SAR (Synthetic Aperture Radar) uses active microwave imaging and has a certain penetration effect. It can effectively detect various camouflaged targets, is suitable for reconnaissance targets under various harsh battlefield conditions, and can provide high-resolution images under various conditions. Therefore, target detection of SAR images plays an important role in military strategic deployment, ocean management and other fields.
[0003] Existing SAR target detection algorithms mainly include constant false alarm rate detection algorithms, multi-image feature fusion algorithms, support vector machines, Bayesian classifiers, etc. With the interference of weather and environment and the increase of data volume, the solution of these traditional algorithms for detection models will become very difficult, which will cause false alarms or missed detections of targets and reduce the detection effect of targets.
[0004] In recent years, deep learning technology has developed rapidly. Since deep learning technology can achieve target detection without the need for time-consuming and laborious artificial feature design, it has promoted the research on SAR image target detection based on deep learning. In related technologies, Su Juan et al. proposed improvements in three aspects: transfer learning, shallow feature enhancement, and data augmentation in the paper "Improved SSD Algorithm for Small Target Ship Detection in SAR Images", which improved the detection accuracy compared with the original SSD detection accuracy. Hu Changhua, in the paper "Small Target Ship Detection in SAR Images Based on Deep Convolutional Neural Network", introduced a balance term of target size into the loss function, making small targets have lower loss function values, so that small targets are easier to detect and the detection effect is improved.
[0005] However, the above methods cannot meet the high-precision requirements for target detection in practical applications. Therefore, those skilled in the art urgently need to propose a SAR image small target detection method with high-precision detection effect. Summary of the Invention
[0006] In order to solve the above problems existing in the prior art, the present invention provides a SAR image target detection method based on feature fusion and cross-layer connection. The technical problems to be solved by the present invention are realized through the following technical solutions:
[0007] An embodiment of the present invention provides a SAR image target detection method based on feature fusion and cross-layer connection, including:
[0008] Obtain an image to be detected, where the image to be detected is a SAR image;
[0009] Input the image to be detected into a pre-trained detection model, so that the detection model downsamples the image to be detected to obtain first-class feature maps of different scales;
[0010] After performing attention feature fusion on the first-class feature maps to obtain second-class feature maps, perform cross-scale feature fusion on the first-class feature maps and the second-class feature maps to obtain third-class feature maps;
[0011] Use the third-class feature maps to perceive the targets in the image to be detected and obtain target detection results.
[0012] In an embodiment of the present invention, the detection model includes a downsampling module, and the downsampling module includes a first downsampling layer, a second downsampling layer, a third downsampling layer, and a fourth downsampling layer. The first downsampling layer, the second downsampling layer, the third downsampling layer, and the fourth downsampling layer all include a convolutional layer with a convolutional kernel size of 3×3 and a stride of 2.
[0013] In an embodiment of the present invention, the step of inputting the image to be detected into a pre-trained detection model so that the detection model downsamples the image to be detected to obtain first-class feature maps of different scales includes:
[0014] After inputting the image to be detected into the pre-trained detection model, the first downsampling layer, the second downsampling layer, the third downsampling layer, and the fourth downsampling layer respectively perform 4-fold downsampling, 8-fold downsampling, 16-fold downsampling, and 32-fold downsampling on the image to be detected to obtain feature maps F1, F2, F3, and F4.
[0015] In an embodiment of the present invention, the detection model further includes a first feature fusion module;
[0016] The step of performing attention feature fusion on the first-class feature maps to obtain second-class feature maps includes:
[0017] The first feature fusion module sequentially inputs the feature map F4 into a convolutional layer with a convolutional kernel size of 1×1 and three serially connected max-pooling layers with a size of 5×5 to obtain the feature map F5 in the second-class feature maps. The size of the feature map F5 is the same as that of the feature map F4;
[0018] The first feature fusion module performs 2-fold upsampling on the feature map F5 and then fuses it with the feature map F3 to obtain the feature map F6;
[0019] The first feature fusion module performs 2-fold upsampling on the feature map F6 and then fuses it with the feature map F2 to obtain the feature map F7;
[0020] After the first feature fusion module performs 2x upsampling on the feature map F7, it is fused with the feature map F1 to obtain the feature map F8.
[0021] In an embodiment of the present invention, the step of performing 2x upsampling on the feature map F5 and then fusing it with the feature map F3 to obtain the feature map F6 includes:
[0022] Perform 2x upsampling on the feature map F5;
[0023] Use the attention mechanism to process the upsampled feature map F5 and the feature map F3 respectively;
[0024] Perform channel concatenation on the feature map F5 and the feature map F3 processed by the attention mechanism to obtain the feature map F6.
[0025] In an embodiment of the present invention, the attention mechanism processes the feature map F3 according to the following steps:
[0026] Use the residual network to process the feature map F3 to obtain the feature map F3';
[0027] Perform global average pooling on the feature map F3' in the horizontal and vertical directions respectively to obtain the horizontal direction feature map and the vertical direction feature map
[0028] Perform convolution on the horizontal direction feature map and the vertical direction feature map respectively, and perform channel concatenation on the obtained first sub-feature map and second sub-feature map to obtain the third sub-feature map;
[0029] After performing non-linear processing on the third sub-feature map, split it to obtain the horizontal direction feature map and the vertical direction feature map
[0030] Perform convolution and normalization on the horizontal direction feature map and the vertical direction feature map respectively to obtain the horizontal direction attention weight and the vertical direction attention weight;
[0031] Multiply the feature map F3', the horizontal direction attention weight, and the vertical direction attention weight to obtain the feature map F3 processed by the attention mechanism.
[0032] In an embodiment of the present invention, the detection model further includes a second feature fusion module;
[0033] The steps of performing cross-scale feature fusion on the first type of feature map and the second type of feature map to obtain the third type of feature map include:
[0034] The second feature fusion module sequentially inputs the feature map F8 into a convolutional layer with a kernel size of 1×1 and three max pooling layers with a size of 5×5 connected in series to obtain the feature map F9 in the third type of feature map, and the size of the feature map F9 is the same as that of the feature map F8;
[0035] After the second feature fusion module performs 2-fold downsampling on the feature map F9, it performs feature fusion with the feature map F2 and the feature map F7 to obtain the feature map F10;
[0036] After the second feature fusion module performs 2-fold downsampling on the feature map F10, it performs feature fusion with the feature map F3 and the feature map F6 to obtain the feature map F11;
[0037] After the second feature fusion module performs 2-fold downsampling on the feature map F11, it performs feature fusion with the feature map F5 to obtain the feature map F12.
[0038] In an embodiment of the present invention, the steps of performing 2-fold downsampling on the feature map F9 and then performing feature fusion with the feature map F2 and the feature map F7 to obtain the feature map F10 include:
[0039] Perform 2-fold downsampling on the feature map F9;
[0040] Use the attention mechanism to process the downsampled feature map F9, the feature map F2, and the feature map F7 respectively;
[0041] Perform channel splicing on the feature map F9, the feature map F2, and the feature map F7 that have been processed by the attention mechanism to obtain the feature map F10.
[0042] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0043] The present invention provides a SAR target detection method based on feature fusion and cross-layer connection. By performing 4-fold, 8-fold, 16-fold, and 32-fold downsampling, the shallow detail features of the image to be detected are extracted, and at the same time, the shallow detail features are fused with the deep semantic features, which can improve the detection ability of the target. In addition, the present invention also introduces an attention mechanism in the process of feature map fusion, which can further improve the detection accuracy of the target in the image to be detected.
[0044] The following will further describe the present invention in detail with reference to the drawings and embodiments. Description of the Drawings
[0045] Figure 1It is a flowchart of a SAR image target detection method based on feature fusion and cross-layer connection provided by an embodiment of the present invention;
[0046] Figure 2 It is a schematic diagram of a SAR image target detection method based on feature fusion and cross-layer connection provided by an embodiment of the present invention;
[0047] Figure 3 It is another schematic diagram of a SAR image target detection method based on feature fusion and cross-layer connection provided by an embodiment of the present invention;
[0048] Figure 4 It is a flowchart of the training process of the detection model provided by an embodiment of the present invention;
[0049] Figure 5 It is a precision curve graph in the training process of the detection model provided by an embodiment of the present invention;
[0050] Figure 6 It is a mean average precision curve graph in the training process of the detection model provided by an embodiment of the present invention;
[0051] Figure 7 It is a loss curve graph in the training process of the detection model provided by an embodiment of the present invention;
[0052] Figure 8a It is a schematic diagram of the target detection result provided by an embodiment of the present invention;
[0053] Figure 8b It is another schematic diagram of the target detection result provided by an embodiment of the present invention. Specific embodiments
[0054] The following further describes the present invention in detail with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.
[0055] Figure 1 It is a flowchart of a SAR image target detection method based on feature fusion and cross-layer connection provided by an embodiment of the present invention, Figure 2 It is a schematic diagram of a SAR image target detection method based on feature fusion and cross-layer connection provided by an embodiment of the present invention. Please refer to Figure 1-2 , an embodiment of the present invention provides a SAR image target detection method based on feature fusion and cross-layer connection, including:
[0056] S1. Obtain the image to be detected, and the image to be detected is a SAR image;
[0057] S2. Input the image to be detected into a pre-trained detection model so that the detection model downsamples the image to be detected to obtain first-class feature maps of different scales;
[0058] S3. After performing attention feature fusion on the first type of feature maps to obtain the second type of feature maps, perform cross-scale feature fusion on the first type of feature maps and the second type of feature maps to obtain the third type of feature maps;
[0059] S4. Use the third type of feature maps to perceive the targets in the image to be detected and obtain the target detection results.
[0060] In this embodiment, since the image to be detected is a SAR image, the input end of the detection model can be a single-channel input. After the image to be detected is input into the above detection model, first perform downsampling on the image to be detected to extract the detailed features of the targets included in the SAR image at the shallow layer.
[0061] Optionally, the detection model includes a downsampling module, and the downsampling module includes a first downsampling layer, a second downsampling layer, a third downsampling layer, and a fourth downsampling layer. In this embodiment, the first downsampling layer, the second downsampling layer, the third downsampling layer, and the fourth downsampling layer all include convolutional layers with a convolutional kernel size of 3×3 and a stride of 2.
[0062] Please continue to refer to Figure 2 , in the above step S2, the step of inputting the image to be detected into the pre-trained detection model to enable the detection model to perform downsampling on the image to be detected to obtain the first type of feature maps at different scales includes:
[0063] After the image to be detected is input into the pre-trained detection model, the first downsampling layer, the second downsampling layer, the third downsampling layer, and the fourth downsampling layer respectively perform 4-fold downsampling, 8-fold downsampling, 16-fold downsampling, and 32-fold downsampling on the image to be detected to obtain the feature map F1, the feature map F2, the feature map F3, and the feature map F4.
[0064] Figure 3 is another schematic diagram of the SAR image target detection method based on feature fusion and cross-layer connection provided by the embodiments of the present invention. As Figure 3 shown, in step S3, the above detection model further includes a first feature fusion module;
[0065] The step of performing attention feature fusion on the first type of feature maps to obtain the second type of feature maps includes:
[0066] The first feature fusion module sequentially inputs the feature map F4 into a convolutional layer with a convolutional kernel size of 1×1 and three serially connected max-pooling layers with a size of 5×5 to obtain the feature map F5 in the second type of feature maps, and the size of the feature map F5 is the same as the size of the feature map F4;
[0067] After the first feature fusion module performs 2-fold upsampling on the feature map F5, it fuses with the feature map F3 to obtain the feature map F6;
[0068] After the first feature fusion module performs 2x upsampling on the feature map F6, it fuses with the feature map F2 to obtain the feature map F7;
[0069] After the first feature fusion module performs 2x upsampling on the feature map F7, it fuses with the feature map F1 to obtain the feature map F8.
[0070] In the above step S4, the detection model further includes a second feature fusion module;
[0071] The step of performing cross-scale feature fusion on the first type of feature map and the second type of feature map to obtain the third type of feature map includes:
[0072] The second feature fusion module sequentially inputs the feature map F8 into a convolutional layer with a kernel size of 1×1 and three 5×5 max pooling layers connected in series to obtain the feature map F9 in the third type of feature map, and the size of the feature map F9 is the same as the size of the feature map F8;
[0073] After the second feature fusion module performs 2x downsampling on the feature map F9, it performs feature fusion with the feature map F2 and the feature map F7 to obtain the feature map F10;
[0074] After the second feature fusion module performs 2x downsampling on the feature map F10, it performs feature fusion with the feature map F3 and the feature map F6 to obtain the feature map F11;
[0075] After the second feature fusion module performs 2x downsampling on the feature map F11, it performs feature fusion with the feature map F5 to obtain the feature map F12.
[0076] Specifically, please continue to refer to Figure 3, taking the case where the size of the image to be detected is 640×640 as an example, after the image to be detected is downsampled by 4 times, 8 times, 16 times, and 32 times respectively, a feature map F1 with a size of 160×160, a feature map F2 with a size of 80×80, a feature map F3 with a size of 40×40, and a feature map F4 with a size of 20×20 are obtained; further, the size of the feature map F5 in the second type of feature map is the same as that of the feature map F4 in the first type of feature map. After the feature map F5 is upsampled by 2 times, it is fused with the feature map F3 to obtain a feature map F6 with a size of 40×40. Then, after the feature map F6 is upsampled by 2 times, it is fused with the feature map F2 to obtain a feature map F7 with a size of 80×80. Next, after the feature map F7 is upsampled by 2 times, it is fused with the feature map F1 to obtain a feature map F8 with a size of 160×160; the size of the feature map F9 in the third type of feature map is the same as that of the feature map F8. After the feature map F9 is downsampled by 2 times, it is fused with the feature maps F2 and F7 to obtain a feature map F10 with a size of 80×80. Then, the feature map F10 is downsampled by 2 times and fused with the feature maps F3 and F6 to obtain a feature map F11 with a size of 40×40. Finally, F11 is downsampled by 2 times and fused with the feature map F5 to obtain a feature map F12 with a size of 20×20.
[0077] From the above process, it can be seen that after the processing of the first feature fusion module and the second feature fusion module, feature maps F12 with a size of 20×20, F11 with a size of 40×40, F10 with a size of 80×80, and F9 with a size of 160×160 can be obtained. The receptive fields of the feature maps F9, F10, F11, and F12 mapped to the image to be detected with a size of 640×640 are 4×4, 8×8, 16×16, and 32×32 respectively, corresponding to the smallest, smaller, medium, and larger targets. Therefore, the detection model can perceive the pixel points of different-sized targets through these four scales of feature maps, thereby generating corresponding detection frames, and finally only retaining the detection frame with the highest confidence as the detection result of the target to be detected.
[0078] In the above step S4, the step of fusing the upsampled feature map F5 by 2 times with the feature map F3 to obtain the feature map F6 includes:
[0079] Upsample the feature map F5 by 2 times;
[0080] Process the upsampled feature map F5 and the feature map F3 respectively using the attention mechanism;
[0081] Perform channel splicing on the feature map F5 and the feature map F3 processed by the attention mechanism to obtain the feature map F6.
[0082] Optionally, the feature map F3 is processed using an attention mechanism according to the following steps:
[0083] S41. Process the feature map F3 using a residual network to obtain a feature map F3';
[0084] S42. Perform global average pooling on the feature map F3' in the horizontal and vertical directions respectively to obtain a horizontal direction feature map and a vertical direction feature map
[0085] S43. Convolve the horizontal direction feature map and the vertical direction feature map respectively, and perform channel concatenation on the obtained first sub-feature map and second sub-feature map to obtain a third sub-feature map;
[0086] S44. After performing non-linear processing on the third sub-feature map, split it to obtain a horizontal direction feature map F x2 and a vertical direction feature map F y2 ;
[0087] S45. Convolve and normalize the horizontal direction feature map F x1 and the vertical direction feature map F y1 respectively to obtain a horizontal direction attention weight and a vertical direction attention weight;
[0088] S46. Multiply the feature map F3', the horizontal direction attention weight, and the vertical direction attention weight to obtain the feature map F3 processed by the attention mechanism.
[0089] In this embodiment, first, the feature map F3 is processed using a residual network, and then the obtained feature map F3' is subjected to global average pooling. Exemplarily, pooling kernels (H, 1) and (1, W) are respectively used to encode the features of the feature map F3' in the horizontal and vertical directions, as shown in formulas (1) and (2):
[0090]
[0091]
[0092] wherein, represents the feature map of the c-th channel in the h-th row, represents the feature map of the c-th channel in the w-th column, W and H respectively represent the width and height of the feature map F3', and F c represents the c-th feature of the feature map F3'.
[0093] It can be seen that after global average pooling, two feature maps in the horizontal and vertical directions, namely F x1 and Fy1 , the attention mechanism adopted in this embodiment not only focuses on the channel direction but also on the spatial direction features, which can ensure that the spatial information in the other direction is not lost while searching for information in a single direction, facilitating more accurate positioning of the target in the image to be detected by the detection model.
[0094] Further, perform convolution and channel splicing on the horizontal direction feature map and the vertical direction feature map to obtain a third sub-feature map, and perform normalization and non-linear processing on the third sub-feature map. Exemplarily, the RelU non-linear activation function can be used for non-linearization, and the calculation method is shown in formula (3):
[0095]
[0096] where F([z h ,z w ) represents the feature map f obtained after convolution and fusion of the horizontal direction feature map and the vertical direction feature map .
[0097] In the above step S45, normalization processing is performed through two 1×1 convolutions and Sigmoid activation function operations, and the attention weights in the horizontal direction and the vertical direction are output. The calculation process is shown in formulas (4) and (5):
[0098]
[0099]
[0100] where and respectively represent the horizontal direction feature map and the vertical direction feature map obtained by performing the Split operation on the feature map f.
[0101] Finally, multiply the horizontal attention weight η w , the vertical direction attention weight η w by the feature map F3’ to obtain the feature map F3 processed by the attention mechanism.
[0102] Optionally, the step of performing 2-fold downsampling on the feature map F9 and then performing feature fusion with the feature map F2 and the feature map F7 to obtain the feature map F10 includes:
[0103] Perform 2-fold downsampling on the feature map F9;
[0104] Use the attention mechanism to process the downsampled feature map F9, the feature map F2, and the feature map F7 respectively;
[0105] The feature map F9, the feature map F2, and the feature map F7 processed by the attention mechanism are concatenated in channels to obtain the feature map F10.
[0106] It should be understood that after the first type of feature map is generated in the present invention, the second type of feature map and the third type of feature map can be generated through feature fusion. Before each feature fusion, the attention mechanism can be used to process the feature maps to be fused. Since the processing process is similar to the processing process of the feature map F3, it will not be elaborated here.
[0107] The above SAR image target detection method based on feature fusion and cross-layer connection network will be further described through simulation experiments below.
[0108] First, a training set, a validation set, and a test set for SAR image target detection are established. The above data sets can all select the public data set SSDD, which contains 1,160 SAR images and 2,456 ships. 70% of the SAR images are randomly selected as the training set, 20% of the SAR images are used as the validation set, and 10% of the SAR images are used as the test set, and the target is set to the ship category.
[0109] Next, the above-established training set, validation set, and test set are preprocessed, and the Mosaic algorithm is used to achieve data augmentation of SAR images. The Mosaic algorithm stitches 4 images by means of random scaling, random cropping, and random arrangement, which is beneficial to enriching the background and small targets of the detected objects. At the same time, when calculating batch normalization, the data of four images will be calculated at one time, so that a relatively good effect can be achieved without a very large mini-batch size. At the same time, the diverse target samples make the trained model have strong generalization ability.
[0110] Figure 4 It is a flowchart of the detection model training process provided by an embodiment of the present invention. Exemplarily, as Figure 4 shown, the detection model is trained according to the following steps:
[0111] First, set the training parameters. Table 1 shows some network hyperparameters of the detection model. The random optimization algorithm Adam is used for training. The size of the training batch is set to Batchsize = 8, the momentum parameter momentum = 0.937, the initial learning rate lr0 = 0.01, the termination learning rate lrf = 0.1, and the training iteration times Epoch = 300.
[0112] Table 1
[0113]
[0114] Figure 5It is the precision curve graph during the training process of the detection model provided by the embodiment of the present invention. Figure 6 It is the mean average precision curve graph during the training process of the detection model provided by the embodiment of the present invention. Figure 7 It is the loss curve graph during the training process of the detection model provided by the embodiment of the present invention. Further, the preprocessed training set and validation set are sequentially input into the constructed detection model, and the size of the SAR image is adaptively scaled, as Figure 5-7 shown. According to the change trends of the loss, precision, and recall rate of the training set and the loss, precision, and mean average precision of the validation set, the learning rate and the number of iterations are adjusted until the precision change and the loss change gradually tend to a stable state, and the final learning rate and the number of iterations are determined.
[0115] After completing the training of the detection model according to the determined learning rate and the number of iterations, the preprocessed test set is input into the trained detection model. The detection effect is as Figure 7 shown. The detection model is evaluated from the precision and the mean average precision, and it is judged whether the evaluation result of the detection model meets the actual application requirements. If it meets the actual application requirements, the model is used for the detection of ship targets in SAR images. As shown in 8a - 8b, the detected target ship is in the target box; otherwise, the depth and width of the detection model are corrected, and the adjusted detection model is retrained.
[0116] The detection results of the object detection method provided by the present invention are compared with those of the SAR image object detection methods based on YOLOv3, SSD, and Faster - RCNN on the SSDD dataset. The mean average precision is increased by 3.3%, 10.8%, and 14.4% respectively, and the precision and the mean average precision reach 95.9% and 98% respectively.
[0117] As can be seen from the above embodiments, the beneficial effects of the present invention are as follows:
[0118] The present invention provides a SAR object detection method based on feature fusion and cross - layer connection. By performing 4 - fold, 8 - fold, 16 - fold, and 32 - fold downsampling to extract the shallow - layer detail features of the image to be detected, and at the same time fusing the shallow - layer detail features with the deep - layer semantic features, the detection ability for targets can be improved. In addition, the present invention introduces an attention mechanism in the process of feature map fusion, which can further improve the detection precision of targets in the image to be detected.
[0119] In the description of the present invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality of" means two or more unless otherwise specifically defined.
[0120] The descriptions with reference to terms such as "an embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.
[0121] Although the present application has been described in conjunction with various embodiments herein, however, in the process of implementing the claimed present application, those skilled in the art can understand and achieve other variations of the disclosed embodiments by viewing the accompanying drawings, the disclosure, and the appended claims.
[0122] The above content is a further detailed description of the present invention in conjunction with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited only to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A method for SAR image target detection based on feature fusion and cross-layer connection, characterized in that, Including: Obtain the image to be detected, where the image to be detected is a SAR image; Input the image to be detected into a pre-trained detection model, so that the detection model downsamples the image to be detected to obtain first-class feature maps of different scales; After performing attention feature fusion on the first-class feature maps to obtain second-class feature maps, perform cross-scale feature fusion on the first-class feature maps and the second-class feature maps to obtain third-class feature maps; Use the third-class feature maps to perceive the targets in the image to be detected and obtain target detection results; The detection model includes a downsampling module, and the downsampling module includes a first downsampling layer, a second downsampling layer, a third downsampling layer, and a fourth downsampling layer; The step of inputting the image to be detected into a pre-trained detection model so that the detection model downsamples the image to be detected to obtain first-class feature maps of different scales includes: after inputting the image to be detected into the pre-trained detection model, the first downsampling layer, the second downsampling layer, the third downsampling layer, and the fourth downsampling layer respectively perform 4-fold downsampling, 8-fold downsampling, 16-fold downsampling, and 32-fold downsampling on the image to be detected to obtain feature map F1, feature map F2, feature map F3, and feature map F4; The detection model further includes a first feature fusion module; The step of performing attention feature fusion on the first-class feature maps to obtain second-class feature maps includes: The first feature fusion module sequentially inputs the feature map F4 into a convolutional layer with a convolution kernel size of 1×1 and 3 serially connected max-pooling layers with a size of 5×5 to obtain the feature map F5 in the second-class feature maps, and the size of the feature map F5 is the same as that of the feature map F4; after the first feature fusion module performs 2-fold upsampling on the feature map F5, it fuses with the feature map F3 to obtain the feature map F6; after the first feature fusion module performs 2-fold upsampling on the feature map F6, it fuses with the feature map F2 to obtain the feature map F7; after the first feature fusion module performs 2-fold upsampling on the feature map F7, it fuses with the feature map F1 to obtain the feature map F8; The step of performing 2-fold upsampling on the feature map F5 and then fusing it with the feature map F3 to obtain the feature map F6 includes: performing 2-fold upsampling on the feature map F5; using an attention mechanism to process the upsampled feature map F5 and the feature map F3 respectively; concatenating the channels of the feature map F5 and the feature map F3 processed by the attention mechanism to obtain the feature map F6; The detection model further includes a second feature fusion module; The steps of performing cross-scale feature fusion on the first type of feature map and the second type of feature map to obtain the third type of feature map include: The second feature fusion module sequentially inputs the feature map F8 into a convolutional layer with a kernel size of 1×1 and three max-pooling layers with a size of 5×5 connected in series to obtain the feature map F9 in the third type of feature map, and the size of the feature map F9 is the same as that of the feature map F8; after the second feature fusion module performs 2-fold downsampling on the feature map F9, it performs feature fusion with the feature map F2 and the feature map F7 to obtain the feature map F10; after the second feature fusion module performs 2-fold downsampling on the feature map F10, it performs feature fusion with the feature map F3 and the feature map F6 to obtain the feature map F11; after the second feature fusion module performs 2-fold downsampling on the feature map F11, it performs feature fusion with the feature map F5 to obtain the feature map F12; The steps of performing 2-fold downsampling on the feature map F9 and then performing feature fusion with the feature map F2 and the feature map F7 to obtain the feature map F10 include: performing 2-fold downsampling on the feature map F9; using the attention mechanism to process the downsampled feature map F9, the feature map F2, and the feature map F7 respectively; concatenating the feature map F9, the feature map F2, and the feature map F7 that have been processed by the attention mechanism in the channel dimension to obtain the feature map F10.
2. The method for SAR image target detection based on feature fusion and cross-layer connection according to claim 1, characterized in that, The first downsampling layer, the second downsampling layer, the third downsampling layer, and the fourth downsampling layer all include convolutional layers with a kernel size of 3×3 and a stride of 2.
3. The method for SAR image target detection based on feature fusion and cross-layer connection according to claim 2, characterized in that, The feature map F3 is processed by the attention mechanism according to the following steps: The feature map F3 is processed by the residual network to obtain the feature map F3'; Perform global average pooling on the feature map F3’ in the horizontal and vertical directions respectively to obtain the horizontal direction feature map and the vertical direction feature map Convolve the horizontal direction feature map and the vertical direction feature map respectively, and perform channel splicing on the obtained first sub-feature map and second sub-feature map to obtain a third sub-feature map; After performing non-linear processing on the third sub-feature map, split it to obtain a horizontal-direction feature map and a vertical-direction feature map Convolve and normalize the horizontal direction feature map and the vertical direction feature map respectively to obtain the attention weights in the horizontal direction and the attention weights in the vertical direction; The feature map F3', the horizontal attention weight, and the vertical attention weight are multiplied to obtain the feature map F3 processed by the attention mechanism.