SAR target detection method based on wavelet decomposition and attention feature fusion

Through wavelet decomposition and Bayesian shrinkage threshold denoising combined with the dual-branch structure of ConvNeXt and Swin Transformer, the problem of noise and clutter interference in SAR images is solved, and efficient object detection effect is achieved.

CN120374959AActive Publication Date: 2025-07-25NAT UNIV OF DEFENSE TECH
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510828572.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-07-25
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

When processing SAR images, the existing SAR object detection method based on deep learning is difficult to effectively distinguish high-frequency noise from the target edge, resulting in spectrum confusion and gradient update direction conflict in the feature learning process, and it is impossible to take into account noise suppression and target feature fidelity. The existing convolutional neural network and Transformer detection framework are limited in performance in SAR images.

Method used

The high-frequency coefficients are denoised by wavelet decomposition and Bayesian shrinkage threshold denoising methods. Combined with the dual-branch structure of ConvNeXt and Swin Transformer, feature adaptive fusion is performed through FGAM, and target detection is performed using FPN and RPN.

Benefits of technology

It effectively suppresses spot noise and background clutter interference in SAR images, improves the accuracy and robustness of target detection, reduces the calculation amount and improves the anti-clutter interference performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374959A_ABST
    Figure CN120374959A_ABST
Patent Text Reader

Abstract

The invention relates to an SAR target detection method based on wavelet decomposition and attention feature fusion, and the method comprises the steps: carrying out the wavelet decomposition of an original SAR image, obtaining a low-frequency coefficient and a high-frequency coefficient, and carrying out the denoising of the high-frequency coefficient through the Bayesian threshold shrinkage; a double-branch structure of ConvNeXt and a double-branch structure of Swin Transform are used for extracting high-frequency features and low-frequency features at different levels respectively, and feature self-adaptive fusion is carried out on each level based on FGAM; inputting the fusion features of each level into the FPN for feature fusion again to obtain a plurality of feature maps with the same number of channels; and inputting each feature map into the RPN to obtain a target detection result. By explicitly utilizing frequency domain priori knowledge, the phenomena of leak detection and false alarm in a complex scene are effectively inhibited.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of radar signal processing, and particularly to a SAR target detection method based on wavelet decomposition and attention feature fusion. Background Art

[0002] Since Synthetic Aperture Radar (SAR) has the ability of all-weather observation, the target detection method based on SAR images has become an important means for current reconnaissance and monitoring, and has important application value in military and civilian fields. However, the acquisition and transmission processes of SAR images are often interfered by multiplicative speckle noise and complex background clutter. These interferences usually come from various links such as image acquisition, processing, encoding, and transmission, and will affect the signal quality in visible or invisible forms, thereby severely restricting the accuracy and robustness of target detection. Specifically, the multiplicative speckle noise will reduce the local contrast, while the complex ground clutter will lead to an increase in the false alarm rate. These factors significantly increase the detection difficulty of targets such as airplanes and ships.

[0003] In recent years, SAR target detection methods based on deep learning have made breakthroughs in the optical field, and detectors such as Faster R-CNN, YOLO, and DETR have emerged. However, when directly migrating the existing detector framework into SAR images for target detection, the performance is often limited. The reason is that SAR images have significant frequency-domain separability. Among them, the high-frequency components (HF) are mainly composed of target edges, textures, and multiplicative speckle noise, and their energy distribution shows local sharp characteristics; while the low-frequency components (LF) reflect the global structure of the scene, background clutter, and large-scale ground object features, with slow-changing energy characteristics. Deep learning methods realize target representation through end-to-end feature learning, but its implicit "black box" feature coupling mechanism faces essential challenges in the SAR scenario, that is, the undifferentiated fusion of high-frequency noise and low-frequency semantics in the feature space, resulting in the model being difficult to balance noise suppression and target feature fidelity. In addition, the existing detection frameworks based on Convolutional Neural Network (CNN) or Transformer generally adopt a single-branch architecture to directly process the original data. This design will cause the limited receptive field of the convolutional kernel to be difficult to effectively distinguish high-frequency noise and target edges, resulting in spectral confusion in the feature learning process. At the same time, the same layer of the network needs to model high-frequency details and low-frequency semantics simultaneously, resulting in conflicting gradient update directions and weakening the model representation ability. Summary of the Invention

[0004] The object of the present invention is to provide a new SAR target detection strategy. By designing a wavelet denoising method based on Bayesian shrinkage threshold, a dual-branch feature extraction structure composed of ConvNeXt and Swin Transformer, and a feature adaptive fusion method based on FGAM, the optimization of the SAR image feature extraction and fusion process is realized, and the problems of speckle noise and background clutter interference existing in SAR images in complex scenarios are solved.

[0005] To achieve the above object, the present invention provides a SAR target detection method based on wavelet decomposition and attention feature fusion, including the following steps: S1. Perform wavelet decomposition on the original SAR image to obtain low-frequency coefficients and high-frequency coefficients, and use Bayesian threshold shrinkage to denoise the high-frequency coefficients; S2. Use the dual-branch structure of ConvNeXt and Swin Transformer to extract high-frequency features and low-frequency features at different levels respectively, and perform feature adaptive fusion based on FGAM at each level; S3. Input the fusion features of each level into FPN to perform feature fusion again to obtain feature maps with the same number of channels; S4. Input each feature map into RPN respectively to obtain the target detection result.

[0006] Preferably, the step S1 includes: Construct the following model for the original SAR image: (1); Wherein, y, x, and e respectively represent the wavelet coefficients of the original SAR image, the wavelet coefficients of the real SAR image, and the wavelet coefficients of the noise, and m and n respectively represent the coordinates of the corresponding coefficient matrices in the horizontal and vertical directions; Take the logarithm of formula (1), and the expression is as follows: (2); Wherein, Y, X, and N respectively represent the wavelet coefficients of the original SAR image, the wavelet coefficients of the real SAR image, and the wavelet coefficients of the noise after logarithmic transformation; Perform wavelet decomposition on the model to obtain the low-frequency coefficient LL and high-frequency coefficients of the SAR image. The high-frequency coefficients include the horizontal high-frequency coefficient LH, the vertical high-frequency coefficient HL, and the diagonal high-frequency coefficient HH. Threshold denoising processing is performed on all high-frequency coefficients respectively. Specifically, the part of the high-frequency coefficients less than the threshold is set to zero to remove the noise, and the new high-frequency coefficients are obtained by adding the denoised high-frequency coefficients.

[0007] Preferably, in the step S1, the threshold used in the denoising process is calculated according to the Bayesian shrinkage method, and the expression is as follows: (4); Wherein, G represents the horizontal high-frequency coefficient LH, the vertical high-frequency coefficient HL or the diagonal high-frequency coefficient HH, represents the threshold value of the corresponding high-frequency coefficient, represents the variance of the wavelet coefficients of the real SAR image, represents the variance of the corresponding high-frequency coefficient after wavelet decomposition, and the expression is as follows: (5); Wherein, represents the median function, represents the matrix of the corresponding high-frequency coefficient, that is, the horizontal high-frequency coefficient matrix , the vertical high-frequency coefficient matrix or the diagonal high-frequency coefficient matrix .

[0008] Preferably, the step S2 includes: Construct a dual-branch structure composed of ConvNeXt and Swin Transformer, extract high-frequency features through ConvNeXt, and extract low-frequency features through Swin Transformer; Construct a feature adaptive fusion network based on FGAM; Add the extracted high-frequency features and low-frequency features to obtain the spatial attention feature map X and input it into FGAM. After the feature map X passes through the parallel spatial attention module and channel attention module, the corresponding weights and are output respectively, and the expressions are as follows: (6); (7); Wherein, represents the channel connection operation, represents the ReLU function activation operation, represents the convolution operation with a convolution kernel size of , , respectively represent the features in the spatial dimension after the global average pooling operation and the global maximum pooling operation, represents the features in the channel dimension after the global average pooling operation; Based on the weights and After the addition, the feature map X is subjected to channel connection operation, and then the channel shuffle operation, 7×7 convolution operation, and Sigmoid function activation operation are performed in sequence to obtain the spatial attention feature map W with the feature weight adjusted and optimized; Based on the feature weights in the feature map W, high-frequency features are added through weighted summation and residual connection. and low frequency characteristics After fusion, the added features are projected through a 1×1 convolutional layer to obtain the final fused features. , the expression is as follows: (8); Among them, W is the weight assigned in the feature map W, and I is a matrix whose elements are all 1.

[0009] It should be noted that feature fusion can integrate feature information of different levels in high-frequency features and low-frequency features, enrich the description of the target object, and thus strengthen the network model's discrimination of target attributes; Generally speaking, the simplest feature fusion method is to achieve it through element-by-element addition, but this method is not flexible enough and cannot effectively capture the relationship between features; Some studies also use splicing to achieve feature fusion, which helps the model learn more complex information by increasing the dimension of features, but may also introduce some irrelevant information; More importantly, these two types of feature fusion methods do not take into account the problem of mismatched receptive fields of features with different frequency characteristics. For example, a pixel in a high-frequency feature may come from a pixel area in a low-frequency feature. Therefore, simple addition or splicing cannot solve the problem of receptive field matching before feature fusion; FGAM fully mixes the channel attention weights and spatial attention weights through channel shuffling to promote information interaction, enhances the weight of important information in the feature map, and weakens the weight of unimportant information, which is beneficial to improve the target detection accuracy. Furthermore, FGAMF further refines the feature map according to the frequency characteristics of the input features, guiding the model to focus more on the key areas of concern, thereby improving the model's anti-clutter interference performance.

[0010] Preferably, in the dual-branch structure, Swin Transformer selects the Base version, and ConvNeXt selects the Base version.

[0011] Preferably, the dual-branch structure has four feature extraction levels, and the corresponding channel numbers are 128, 256, 512 and 1024 respectively.

[0012] Preferably, step S3 comprises: After performing convolution operations with a 1×1 convolution kernel on multiple fusion features from different levels respectively, interpolate and add them to the convolution operation results of the previous level in sequence. Then, perform convolution operations with a 3×3 convolution kernel on each result of the addition operation, and output the corresponding feature maps; For those without a previous level for interpolation and addition operations, directly take the results of their convolution operations, perform convolution operations with a 3×3 convolution kernel, output the corresponding feature maps, and repeat the convolution operation with a 3×3 convolution kernel on this feature map, and output the corresponding feature maps.

[0013] The present invention has at least the following beneficial effects: 1. Aiming at the speckle noise characteristics of SAR images, the present invention introduces a Bayesian shrinkage threshold denoising method based on wavelet transform. The SAR image is decomposed into low-frequency coefficients and high-frequency coefficients through wavelet transform, where the low-frequency represents large-scale structural information, and the high-frequency contains detailed features and residual noise. The speckle noise contained in the high-frequency coefficients is regarded as Gaussian distribution, and the Bayesian shrinkage method is adopted to calculate the threshold for adaptive denoising. This method can better simulate the characteristics of speckle noise and remove the noise.

[0014] 2. Aiming at the clutter interference problem of SAR images, the present invention introduces a feature adaptive fusion method based on FGAM attention mechanism. This method can fully consider the problem of receptive field matching between high-frequency features and low-frequency features, and adaptively allocate appropriate weights for the input features according to their frequency characteristics. It not only reduces the information loss in the feature fusion process, ensures the integrity of high-frequency features and low-frequency features, but also focuses on enhancing the feature response in the target-related areas, thereby improving the clutter interference resistance performance of the model. 3. Aiming at the different information contained in the high-frequency coefficients and low-frequency coefficients obtained by wavelet decomposition, the present invention designs a dual-branch structure of ConvNeXt and Swin Transformer to extract high-frequency features and low-frequency features respectively. Among them, Swin Transformer adopts a window attention mechanism and a hierarchical feature processing method, enabling the model to obtain global attention while greatly reducing the computational complexity of the model, reducing the computational complexity to a linear relationship with the image size. And ConvNeXt combines the advantages of traditional convolutional neural networks and Transformer models, aiming to provide more efficient computational performance. At the same time, through depth convolution, the model can process input features in a finer-grained manner and retain a high accuracy; combined with the CNN branch focusing on high-frequency detail learning and the Transformer branch capturing low-frequency global features, it can avoid feature competition and maximize the utilization of prior knowledge.

[0015] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood from the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements.

[0017] Figure 1 It is an architecture diagram of a SAR target detection method based on wavelet decomposition and attention feature fusion for Embodiment 1 of the present invention; Figure 2 It is a result comparison diagram of threshold denoising using the Bayesian shrinkage method in Embodiment 1 of the present invention, where (a) is the original image and (b) is the denoised image; Figure 3 For Figure 2 the operation flow chart from (a) to (b) in Figure 4 It is an architecture diagram of a feature adaptive fusion network based on FGAM in Embodiment 1 of the present invention; Figure 5 For Figure 4 the architecture diagram of the FGAM network in Figure 6 It is an architecture diagram of the FPN network in Embodiment 1 of the present invention; Figure 7 It is a result comparison diagram of detecting SAR ship targets in Embodiment 2 of the present invention, where (a) uses the YOLOv8 algorithm and (b) uses the method of the present invention; Figure 8 It is a result comparison diagram of detecting SAR vehicle targets in Embodiment 2 of the present invention, where (a) uses the YOLOv8 algorithm and (b) uses the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] The following provides a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist understanding, and they should be regarded as merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0019] Embodiment 1 The idea of the technical solution of the present invention is as Figure 1As shown in the figure, the input SAR image is decomposed into high-frequency coefficients and low-frequency coefficients by using the wavelet transform method. The Bayesian shrinkage method is adopted to remove the noise in the high-frequency coefficients, and the high-frequency coefficients after noise removal are weighted and recombined into new high-frequency coefficients. For the high-frequency coefficients and low-frequency coefficients obtained after wavelet decomposition and denoising, the present invention designs a dual-branch structure composed of a ConvNeXt branch and a Swin Transformer branch, aiming to utilize the powerful feature extraction capabilities of ConvNeXt for high-frequency features and Swin Transformer for low-frequency features to fully extract the target information and improve the target detection accuracy. Based on the Feature-guided Attention Mechanism (FGAM), the extracted features are adaptively fused to more reasonably optimize the weight allocation and reduce the information loss in the feature fusion process. After the fused features pass through the Feature Pyramid Networks (FPN) and the Region Proposal Network (RPN), the target detection results are output.

[0020] The technical details of the present invention will be described in more detail below in combination with derivations and formulas.

[0021] S1. Perform wavelet decomposition on the original SAR image to obtain low-frequency coefficients and high-frequency coefficients, and use Bayesian threshold shrinkage to denoise the high-frequency coefficients. Specifically: Since the speckle noise of the SAR image is multiplicative noise, the following model can be constructed for the original SAR image: (1); Among them, y, x, and e respectively represent the wavelet coefficients of the original SAR image, the wavelet coefficients of the real SAR image, and the wavelet coefficients of the noise, and m and n respectively represent the coordinates of the corresponding coefficient matrices in the horizontal and vertical directions; For the convenience of subsequent processing, take the logarithm of Equation (1) to convert the multiplicative noise into additive noise. The expression is as follows: (2); Among them, Y, X, and N respectively represent the wavelet coefficients of the original SAR image, the wavelet coefficients of the real SAR image, and the wavelet coefficients of the noise after logarithmic transformation; Since there is independence between the input original SAR image and the noise, from Equation (2), we can obtain: (3); Among them, 、 、 respectively represent the variances of the wavelet coefficients of the original SAR image, the wavelet coefficients of the true SAR image, and the wavelet coefficients of the noise; After taking the logarithm operation on the multiplicative noise model that follows the Gamma distribution, its distribution can be approximated as a Gaussian distribution; Perform wavelet decomposition on the model to obtain the low-frequency coefficient LL and high-frequency coefficients of the SAR image. Among them, the high-frequency coefficients include the horizontal high-frequency coefficient LH, the vertical high-frequency coefficient HL, and the diagonal high-frequency coefficient HH. Threshold denoising is performed on all high-frequency coefficients respectively. Specifically, the part of the high-frequency coefficients that is less than the threshold is set to zero to remove the noise, and the new high-frequency coefficients are obtained by adding the denoised high-frequency coefficients; In the above denoising process, the threshold is calculated according to the principle of Bayesian shrinkage method, and the expression is as follows: (4); Among them, G represents the horizontal high-frequency coefficient LH, the vertical high-frequency coefficient HL, or the diagonal high-frequency coefficient HH, represents the threshold corresponding to the high-frequency coefficient, represents the variance of the wavelet coefficients of the true SAR image, which is obtained by the formula for calculating the matrix variance, represents the variance of the corresponding high-frequency coefficient after wavelet decomposition, and the expression is as follows: (5); Among them, represents the median function, represents the matrix of the corresponding high-frequency coefficient, that is, the horizontal high-frequency coefficient matrix , the vertical high-frequency coefficient matrix or the diagonal high-frequency coefficient matrix .

[0022] To verify the effectiveness of the noise removal method in step S1, an original SAR image is selected from the dataset for denoising. The low-frequency coefficient and the new high-frequency coefficient obtained by wavelet decomposition are recombined, and restored through wavelet inverse transform to obtain the denoised SAR image. The result is as Figure 2 shown. The presence of speckle noise will cause the originally clear contours or textures in the image to be submerged by the noise, resulting in a significant decrease in the image contrast. However, the SAR image obtained after wavelet decomposition denoising has a higher contrast and obvious denoising effect. The entire process of decomposition - denoising - recombination - restoration is as Figure 3 shown.

[0023] S2. Use ConvNeXt and Swin Transformer to extract high-frequency features and low-frequency features at different levels respectively, and perform feature adaptive fusion based on FGAM at each level. Specifically: Construct a dual-branch structure composed of ConvNeXt and Swin Transformer; among them, Swin Transformer is a general backbone network designed based on the Transformer architecture, and includes four versions: Tiny (T), Small (S), Base (B), and Large (L). In the present invention, in order to balance the working performance and computational efficiency of the model, Swin Transformer-B is selected as the backbone network for low-frequency feature extraction; ConvNeXt is an efficient convolutional neural network (CNN) model, which also includes four versions: T, S, B, and L. In order to correspond to Swin Transformer, ConvNeXt-B is selected as the backbone network for high-frequency feature extraction in the present invention; Construct as Figure 4 shown in the FGAM-based feature adaptive fusion network, that is Figure 1 the FGAMF module in Figure 5 shown, and the extracted high-frequency features and low-frequency features are added to obtain the spatial attention feature map X and input into FGAM. After the feature map X passes through the parallel spatial attention module and channel attention module, the corresponding weights and are output respectively, and the expressions are as follows: (6); (7); Among them, represents the channel connection operation, represents the ReLU function activation operation, represents the convolution operation with a convolution kernel size of , , respectively represent the features in the spatial dimension after the global average pooling operation and the global maximum pooling operation, represents the features in the channel dimension after the global average pooling operation; Based on the result of adding the weights and , perform a channel connection operation on the feature map X, and then successively pass through the channel shuffle operation, the 7×7 convolution operation, and the Sigmoid function activation operation to obtain the spatial attention feature map W with the feature weights adjusted and optimized; Based on the feature weights in the feature map W, the high-frequency features and low-frequency features After fusion, the added features can be projected through a 1×1 convolutional layer to obtain the final fused features. , and the expression is as follows: (8); Among them, W is the weight assigned in the feature map W, and I is a matrix with all elements being 1; In the present invention, the extraction of high-frequency features and low-frequency features by the double-branch structure is set to four different levels, and they respectively correspond to the number of channels of 128, 256, 512, and 1024, that is, four fused features C1 - C4 are output.

[0024] S3. Input the fused features of each level into the FPN for feature fusion again to obtain multiple feature maps with the same number of channels. Specifically: Since the FPN structure can achieve the mixing between "coarse-grained" features and "fine-grained" features by promoting cross-level feature fusion, the FPN is introduced in the present invention; the key to the FPN to achieve feature fusion lies in the upsampling operation and the downsampling operation. Among them, upsampling refers to performing an interpolation operation on the input to increase the resolution, and downsampling refers to reducing the data dimension through a convolutional operation from the original data. And the original SAR image in the present invention outputs a list of gradually decreasing feature maps after the above-mentioned wavelet decomposition, Bayesian threshold denoising, and double-branch structure, which is equivalent to having completed the downsampling process; Therefore, as Figure 6 shown, after the fused features C1 - C4 from four different levels pass through a convolutional operation with a convolutional kernel of 1×1, they are then interpolated and added to the convolutional operation results of the previous level in turn to obtain the corresponding addition operation results M1 - M4. Among them, M4 has no previous level for interpolation addition operation, and directly takes the convolutional operation result of C4; For M1 - M3, a convolutional operation with a convolutional kernel of 3×3 is performed to output the corresponding feature maps P1 - P3. For M4, a convolutional operation with a convolutional kernel of 3×3 is performed to output the corresponding feature map P4. P4 repeats a convolutional operation with a convolutional kernel of 3×3 to output the corresponding feature map P5; the number of channels of the feature maps P1 - P5 is the same.

[0025] S4. Input each feature map into the RPN respectively to obtain the object detection results. Specifically: Input the 5 feature maps P1 - P5 with the same number of channels into the detection heads of 5 identical RPNs respectively. Each branch performs foreground and background classification and bbox regression. Calculate the IoU value according to the bbox obtained by regression and the ground truth in the classification label to obtain the loss value. The model will backpropagate the gradient according to the loss in multiple epochs (that is, a complete traversal of the entire training dataset by the model), and iteratively optimize the weights of the object detection model to make the predicted bbox and category close to the ground truth.

[0026] Example 2 In this example, the SAR target detection method in Example 1 was verified. By collecting data on the spot in Changde, Hunan, a dataset for this experiment was constructed. This dataset contains a total of 7,270 SAR images, and the data was divided according to the ratio of training set: test set: validation set = 8:1:1.

[0027] In this example, a self-made SAR dataset was used as the target detection dataset, and YOLOv8 was used as the comparative experimental algorithm to perform SAR target detection on the method of the present invention and the YOLOv8 algorithm. The results are as Figure 7 and Figure 8 shown. Specifically: in Figure 7 , YOLOv8 misidentified a ship as a bridge, while the method of the present invention identified it correctly. In Figure 8 , both the completeness rate and the precision rate of YOLOv8 in vehicle recognition were significantly inferior to those of the method of the present invention. It can be seen that the method of the present invention has significant advantages in the detection of small targets in SAR images.

[0028] The above has described the embodiments of the present disclosure. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other ordinary skill in the art in the technical field to understand the disclosed embodiments.

[0029] The above are only optional embodiments of the present disclosure and are not used to limit the present disclosure. For those skilled in the art, the present disclosure can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A SAR target detection method based on wavelet decomposition and attention feature fusion, characterized in that It includes the following steps: S1. Perform wavelet decomposition on the original SAR image to obtain low-frequency coefficients and high-frequency coefficients, and use Bayesian threshold shrinkage to denoise the high-frequency coefficients; S2. Use the dual-branch structure of ConvNeXt and Swin Transformer to extract high-frequency features and low-frequency features at different levels respectively, and perform feature adaptive fusion based on FGAM at each level; S3. Input the fusion features of each level into FPN for feature fusion again to obtain multiple feature maps with the same number of channels; S4. Input each feature map into RPN respectively to obtain the target detection results.

2. The SAR target detection method based on wavelet decomposition and attention feature fusion according to claim 1, wherein The step S1 includes: Construct the following model for the original SAR image: (1); where y, x, and e represent the wavelet coefficients of the original SAR image, the wavelet coefficients of the real SAR image, and the wavelet coefficients of the noise respectively, and m and n represent the coordinates of the corresponding coefficient matrices in the horizontal and vertical directions; Take the logarithm of formula (1), and the expression is as follows: (2); where Y, X, and N represent the wavelet coefficients of the original SAR image, the wavelet coefficients of the real SAR image, and the wavelet coefficients of the noise after logarithmic transformation respectively; Perform wavelet decomposition on the model to obtain the low-frequency coefficient LL and high-frequency coefficients of the SAR image. The high-frequency coefficients include the horizontal high-frequency coefficient LH, the vertical high-frequency coefficient HL, and the diagonal high-frequency coefficient HH. Threshold denoising is performed on all high-frequency coefficients respectively. Specifically, the part of the high-frequency coefficients that is less than the threshold is set to zero to remove the noise, and the new high-frequency coefficients are obtained by adding the denoised high-frequency coefficients.

3. The SAR target detection method based on wavelet decomposition and attention feature fusion according to claim 2, wherein In the step S1, the threshold used in the denoising process is calculated according to the Bayesian shrinkage method, and the expression is as follows: (4); where G represents the horizontal high-frequency coefficient LH, the vertical high-frequency coefficient HL, or the diagonal high-frequency coefficient HH, represents the threshold corresponding to the high-frequency coefficient, represents the variance of the wavelet coefficients of the real SAR image, represents the variance of the corresponding high-frequency coefficient after wavelet decomposition, and the expression is as follows: (5); Among them, represents the median function, represents the matrix corresponding to the high-frequency coefficients, that is, the horizontal high-frequency coefficient matrix , the vertical high-frequency coefficient matrix or the diagonal high-frequency coefficient matrix .

4. The SAR target detection method based on wavelet decomposition and attention feature fusion according to claim 3, wherein The step S2 includes: Construct a dual-branch structure composed of ConvNeXt and Swin Transformer, extract high-frequency features through ConvNeXt, and extract low-frequency features through Swin Transformer; Construct a feature adaptive fusion network based on FGAM; The extracted high-frequency features and low-frequency features are added to obtain the spatial attention feature map X and input into FGAM. After the feature map X passes through the parallel spatial attention module and channel attention module, the corresponding weights and are respectively output. The expression is as follows: (6); (7); Among them, represents the channel connection operation, represents the ReLU function activation operation, represents the convolution operation with a convolution kernel size of ; , respectively represent the features in the spatial dimension after the global average pooling operation and the global maximum pooling operation, represents the feature in the channel dimension after the global average pooling operation; Based on the weight and After adding the results, perform a channel connection operation on the feature map X, and then successively pass through a channel shuffle operation, a 7×7 convolution operation, and a Sigmoid function activation operation to obtain a spatial attention feature map W with optimized feature weights; Based on the feature weights in the feature map W, the high-frequency features are fused with the low-frequency features through weighted summation and residual connection. The added features are then projected through a 1×1 convolutional layer to obtain the final fused features and low-frequency features The expression is as follows: The expression is as follows: (8); where W is the weight assigned in the feature map W, and I is a matrix with all elements being 1.

5. The SAR target detection method based on wavelet decomposition and attention feature fusion according to claim 4, wherein In the dual-branch structure, the Base version of Swin Transformer is selected, and the Base version of ConvNeXt is selected.

6. The SAR target detection method based on wavelet decomposition and attention feature fusion according to claim 4, characterized in that, The number of feature extraction levels of the dual-branch structure is four levels, and they correspond to the number of channels of 128, 256, 512, and 1024 respectively.

7. The SAR target detection method based on wavelet decomposition and attention feature fusion according to claim 4, wherein The step S3 includes: After respectively performing convolution operations with a convolution kernel of 1×1 on multiple fusion features from different levels, then interpolate and add them to the convolution operation results of the previous level in turn. Perform convolution operations with a convolution kernel of 3×3 on each addition operation result obtained, and output the corresponding feature maps; For those without an upper level for interpolation and addition operations, directly take the convolution operation results and perform convolution operations with a convolution kernel of 3×3, output the corresponding feature maps, and repeat the convolution operation with a convolution kernel of 3×3 on this feature map, and output the corresponding feature maps.

Citation Information

Patent Citations

  • SAR (Synthetic Aperture Radar) image change detection method based on adaptive weight and high frequency threshold

    CN106296655A

  • SAR image target recognition method based on wavelet threshold denoising combined with convolutional neural network

    CN108898155A

  • SAR (Synthetic Aperture Radar) target detection method based on multi-scale perception and Transform auxiliary attention generation

    CN118781320A

  • Target detection method for view angle of unmanned aerial vehicle

    CN119048730A

  • SAR image ship target detection method based on multi-attention fusion

    CN119360196A