Fine-grained target detection method, device and equipment for SAR (Synthetic Aperture Radar) image

By adopting texture enhancement feature extraction network and space-frequency feature interaction module in SAR image detection, combined with gated attention dynamic fusion and mask supervision module, the problem of difficult to distinguish noise from objects and using aircraft geometry prior knowledge in the prior art is solved, and high-precision and robust aircraft target detection is achieved.

CN120070877APending Publication Date: 2025-05-30NAT UNIV OF DEFENSE TECH

Patent Information

Application Number
CN202510556194.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When processing aircraft targets, existing SAR image detection methods are difficult to effectively distinguish high-frequency noise from real objects, accurately capture subtle features, and do not fully utilize the prior knowledge of aircraft geometry, which increases the risk of semantic confusion among categories.

Method used

Textured enhanced feature extraction network is used to extract image features through fractional-order Gabor convolution units and residual convolution units, and fuse them through gated attention dynamic fusion mechanism. At the same time, the space-frequency feature interaction module is introduced for hierarchical multi-scale feature fusion, and category-specific shape prior knowledge is integrated through a lightweight mask supervision module.

Benefits of technology

It improves the accuracy and robustness of target detection, effectively reduces intra-class variance and inter-class ambiguity, and improves the fine-grained detection performance of aircraft targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070877A_ABST
    Figure CN120070877A_ABST
Patent Text Reader

Abstract

The invention relates to a fine-grained target detection method, device and equipment for an SAR (Synthetic Aperture Radar) image, and the method comprises the steps: carrying out the feature extraction of an SAR target image through employing a texture enhancement feature extraction network, and obtaining a multi-scale feature image, and enabling the texture enhancement feature extraction network to be obtained through the optimization of a feature extraction block in a CSPDarknet network backbone network. The optimized feature extraction block comprises a fractional order Gabor convolution unit and a residual convolution unit which are used for respectively extracting image features, and a fusion unit which is used for performing gating attention dynamic fusion on the image features extracted by the fractional order Gabor convolution unit and the residual convolution unit; and fusing global contexts in the multi-scale feature image by using the spatial frequency feature interaction unit to obtain a multi-scale fused feature image, and outputting position information and category prediction information of the target according to the multi-scale fused feature image by using a classification branch and a regression branch in the detection unit. By adopting the method, the precision and robustness of target detection can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of SAR target detection, and in particular to a fine-grained target detection method, device, and equipment for SAR images. Background Art

[0002] Synthetic Aperture Radar (SAR), as an active microwave remote sensing technology with all-weather and all-time imaging capabilities, has been widely used in the field of target detection and recognition. Thanks to the rapid development of SAR imaging technology, high-resolution SAR images are easier to obtain than ever before, which provides opportunities for fine-grained target detection. Fine-grained target detection aims to detect and classify highly similar sub-categories within the same category, and the detection and classification of aircraft, as typical targets, are of great significance.

[0003] However, existing methods for detecting and classifying aircraft mainly focus on designing efficient multi-scale feature fusion modules and attention mechanisms to cope with the discrete distribution and scale changes of aircraft. In addition, Strong Scattering Points (SSPs), which are high-intensity echoes caused by specific geometric features of aircraft, have been proven to have distinctiveness and stability under SAR imaging. Some studies have proposed calculating discrete priors for different aircraft categories and estimating the target pose angle to guide the model to learn discriminative features, thereby improving the classification accuracy of aircraft target slices. Although the above methods based on convolutional neural networks have achieved success in detecting and classifying aircraft from high-resolution SAR images, they still have some limitations. Specifically, current methods mainly rely on spatial information to construct the network, which limits their ability to effectively distinguish high-frequency noise components from real objects and accurately capture subtle features. In addition, due to the inherent challenges in extracting robust scattering structure patterns related to the target in large-scale SAR images, these methods are usually limited to processing aircraft target slices. For aircraft detection in large-scale SAR images, existing detection networks often fail to fully utilize the prior knowledge of aircraft geometry to comprehensively encode category-specific discriminative features, thus increasing the risk of semantic confusion between categories, especially in complex scenarios. Summary of the Invention

[0004] Based on this, it is necessary to provide a fine-grained target detection method, device, and equipment for SAR images that can achieve high precision and high robustness in response to the above technical problems.

[0005] A fine-grained target detection method for SAR images, the method includes: Obtain a SAR target image to be detected; The texture-enhanced feature extraction network is used to extract features from the SAR target image to obtain a multi-scale feature image. Among them, the texture-enhanced feature extraction network is obtained by optimizing the feature extraction block in the backbone network of the CSPDarknet network. The optimized feature extraction block includes a fractional-order Gabor convolution unit and a residual convolution unit that respectively extract image features, and a fusion unit that dynamically fuses the image features extracted by these two units through gated attention; The spatial frequency feature interaction unit is used to fuse the global context in the multi-scale feature image to obtain a multi-scale fusion feature image; The detection unit is used to perform target detection according to the multi-scale fusion feature image. Among them, the classification branch and the regression branch in the detection unit respectively output the position information and the category prediction information of the target.

[0006] In one embodiment, in the fractional-order Gabor convolution unit: The input image is divided into a preset number of groups of images, and each group of images is processed through a convolutional fractional-order Gabor kernel to construct multiple directional features; After splicing the directional features, the output features of the fractional-order Gabor convolution unit are obtained.

[0007] In one embodiment, the convolutional fractional-order Gabor kernel is expressed as: ; Among them, ; In the above formula, and are the coordinates in the spatial domain and the fractional domain respectively, represents a transformation kernel, represents the transformation angle, represents the transformation order, is a Gaussian window function, and the superscript " - " represents the complex conjugate. U and V respectively represent the number of samples in the fractional domain, and are the sampling intervals, H and W are the height and width of the image respectively, represents the input image.

[0008] In one embodiment, the residual convolution unit is expressed as: ; In the above formula, represents the output features of the residual convolution unit, and respectively represent convolutional layers with convolutional kernels of 3×3 and 5×5, represents the input image of the residual convolutional unit.

[0009] In one embodiment, in the fusion unit, fusion weights are adaptively generated based on the output features of the fractional-order Gabor convolutional unit and the residual convolutional unit, and the output features of the fusion unit are obtained according to the fusion weights and the output features; Among them, the formula for adaptively generating fusion weights is as follows: ; In the above formula, represents the Softmax function, represents global average pooling, 、 respectively represent the output features of the fractional-order Gabor convolutional unit and the residual convolutional unit, 、 respectively represent the corresponding fusion weights of the output features of the fractional-order Gabor convolutional unit and the residual convolutional unit.

[0010] In one embodiment, a fine-grained object detection model is constructed according to the texture enhancement feature extraction network, the spatial frequency feature interaction unit, and the detection unit; The SAR target image to be detected is input into the fine-grained object detection model to achieve object detection.

[0011] In one embodiment, when training the fine-grained object detection model: A mask prediction unit and a target type shape mask library are set in the detection unit, where the target type shape mask library includes various types of target shape mask maps; The mask prediction unit generates a corresponding target mask prediction map according to the large-scale fusion feature image in the multi-scale fusion feature image; The target shape mask map corresponding to the target in the original training image is obtained from the target type shape mask library; By constructing a loss function according to the target shape mask and the target mask prediction map, category-aware semantic constraints are performed during the training process.

[0012] In one embodiment, the mask prediction unit: Successively uses a convolutional layer, upsampling, and skip connections on the large-scale fusion feature image to gradually decode features and restore the spatial resolution, and obtains the target mask prediction map with the same size as the original training image.

[0013] The present application also provides a fine-grained target detection device for SAR images, and the device includes: an SAR image acquisition module, configured to acquire an SAR target image to be detected; a multi-scale feature extraction module, configured to perform feature extraction on the SAR target image by using a texture-enhanced feature extraction network to obtain a multi-scale feature image, wherein the texture-enhanced feature extraction network is obtained by optimizing a feature extraction block in the backbone network of the CSPDarknet network, and the optimized feature extraction block includes a fractional-order Gabor convolution unit and a residual convolution unit that respectively extract image features, and a fusion unit that dynamically fuses the image features extracted by these two units through gated attention; a multi-scale feature fusion module, configured to fuse the global context in the multi-scale feature image by using a spatial frequency feature interaction unit to obtain a multi-scale fusion feature image; a target detection module, configured to perform target detection according to the multi-scale fusion feature image by using a detection unit, wherein the classification branch and the regression branch in the detection unit respectively output the position information of the target and the category prediction information.

[0014] A computer device includes a memory and a processor. When the processor executes a computer program, the following steps are implemented: acquire an SAR target image to be detected; perform feature extraction on the SAR target image by using a texture-enhanced feature extraction network to obtain a multi-scale feature image, wherein the texture-enhanced feature extraction network is obtained by optimizing a feature extraction block in the backbone network of the CSPDarknet network, and the optimized feature extraction block includes a fractional-order Gabor convolution unit and a residual convolution unit that respectively extract image features, and a fusion unit that dynamically fuses the image features extracted by these two units through gated attention; fuse the global context in the multi-scale feature image by using a spatial frequency feature interaction unit to obtain a multi-scale fusion feature image; perform target detection according to the multi-scale fusion feature image by using a detection unit, wherein the classification branch and the regression branch in the detection unit respectively output the position information of the target and the category prediction information.

[0015] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, the following steps are implemented: acquire an SAR target image to be detected; The texture-enhanced feature extraction network is used to extract features from the SAR target image to obtain a multi-scale feature image. Among them, the texture-enhanced feature extraction network is obtained by optimizing the feature extraction block in the backbone network of the CSPDarknet network. The optimized feature extraction block includes a fractional-order Gabor convolution unit and a residual convolution unit that respectively extract image features, and a fusion unit that dynamically fuses the image features extracted by these two units through gated attention; The spatial frequency feature interaction unit is used to fuse the global context in the multi-scale feature image to obtain a multi-scale fused feature image; The detection unit is used to perform target detection according to the multi-scale fused feature image. Among them, the classification branch and the regression branch in the detection unit respectively output the position information of the target and the category prediction information.

[0016] The above-mentioned fine-grained target detection method, device and equipment for SAR images use the texture-enhanced feature extraction network to extract features from the SAR target image to obtain a multi-scale feature image. The texture-enhanced feature extraction network is obtained by optimizing the feature extraction block in the backbone network of the CSPDarknet network. The optimized feature extraction block includes a fractional-order Gabor convolution unit and a residual convolution unit that respectively extract image features, and a fusion unit that dynamically fuses the image features extracted by these two units through gated attention. The spatial frequency feature interaction unit is used to fuse the global context in the multi-scale feature image to obtain a multi-scale fused feature image. The classification branch and the regression branch in the detection unit are used to output the position information of the target and the category prediction information according to the multi-scale fused feature image. Using this method can effectively improve the accuracy and robustness of target detection. Description of the Drawings

[0017] Figure 1 It is a schematic flowchart of a fine-grained target detection method for SAR images in an embodiment; Figure 2 It is a schematic structural diagram of an optimized feature extraction block in an embodiment; Figure 3 It is a schematic structural diagram of a spatial frequency feature interaction unit in an embodiment; Figure 4 It is a schematic diagram of the process of generating a target mask prediction map in an embodiment; Figure 5 It is a schematic diagram of the overall framework of a fine-grained target detection model in an embodiment; Figure 6 It is a schematic flowchart of the process of generating a target mask image in an embodiment; Figure 7 It is a schematic block diagram of the structure of a schematic flowchart device in an embodiment; Figure 8 It is the internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0018] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0019] Existing SAR image detection and classification methods for aircraft mainly focus on designing multi-scale feature fusion modules and attention mechanisms to deal with the discrete distribution and scale changes of aircraft. Strong scattering points (SSPs) are discriminative and stable. Some studies improve the classification accuracy by calculating discrete priors and estimating attitude angles. Although the methods based on convolutional neural networks are successful, they have limitations: First, current methods mostly rely on spatial information to construct networks, making it difficult to effectively distinguish high-frequency noise from real objects and accurately capture subtle features; Second, since it is challenging to extract robust SSPs, usually only aircraft target slices are processed; Third, in large-scale SAR image detection, existing networks do not fully utilize the prior knowledge of aircraft geometric shapes to encode discriminative features, increasing the risk of semantic confusion between categories in complex scenarios.

[0020] To address the above problems, in the present application, as Figure 1 shown, a fine-grained object detection method for SAR images is proposed, including the following specific steps: Step S100: Obtain a SAR target image to be detected.

[0021] Step S110: Use a texture-enhanced feature extraction network to extract features from the SAR target image to obtain a multi-scale feature image. Among them, the texture-enhanced feature extraction network is obtained by optimizing the feature extraction blocks in the backbone network of the CSPDarknet network. The optimized feature extraction blocks include a fractional-order Gabor convolution unit and a residual convolution unit that respectively extract image features, and a fusion unit that dynamically fuses the image features extracted by these two units through gated attention.

[0022] Step S120: Use a spatial frequency feature interaction unit to fuse the global context in the multi-scale feature image to obtain a multi-scale fused feature image.

[0023] Step S130: Use a detection unit to perform object detection based on the multi-scale fused feature image. Among them, the classification branch and the regression branch in the detection unit respectively output the position information of the target and the category prediction information.

[0024] In step S100, the target object in the SAR target image to be detected is an aircraft. Different from ship and vehicle targets with rectangular geometries in SAR images, aircraft present a discrete appearance composed of strong scattering points. This phenomenon can be attributed to the complex geometry of the aircraft, resulting in more complex scattering mechanisms such as multiple scattering, cavity scattering, edge diffraction, and reflection. Compared with the engine, the fuselage with a smooth surface is prone to reflection, resulting in weaker backscattering characteristics. In addition, aircraft in SAR images have significant azimuth sensitivity. Even for the same type of aircraft, its visual appearance and scattering characteristics will vary significantly at different azimuth angles. Moreover, the high similarity between different types of aircraft may cause semantic ambiguity, making fine-grained aircraft detection in synthetic aperture radar (SAR) images an extremely challenging task.

[0025] For SAR fine-grained target detection with the target being an aircraft, in this application, a texture-enhanced backbone network based on the fractional-order Gabor transform (FrGT) is proposed. By combining the fractional-order Fourier transform principle with multi-directional Gabor filtering, it adaptively captures the discriminative features of the target while effectively suppressing the interference of speckle noise and background clutter. Then, a spatial-frequency feature interaction module (SF2IM) is introduced to perform hierarchical multi-scale feature fusion, so as to extract comprehensive and rich context information. Finally, a lightweight mask supervision module is proposed to integrate category-specific shape prior knowledge through differentiable constraint learning, thereby effectively reducing the intra-class variance and inter-class ambiguity.

[0026] Furthermore, this method includes three parts that sequentially process the SAR target image, namely multi-scale feature extraction using the texture-enhanced feature extraction network in step S110, context feature fusion of the multi-scale features using the spatial-frequency feature interaction unit in step S120, and finally target detection using the detection unit based on the multi-scale fusion features in step S130.

[0027] In step S110, based on the feature extraction backbone network in the CSPDarknet network, the feature extraction blocks therein are optimized using the Gabor transform to obtain texture-enhanced semantic features.

[0028] Considering that in synthetic aperture radar (SAR) images, various types of aircraft exhibit unique texture patterns, which are important distinguishing features. In addition, aircraft in SAR images usually span multiple scales and orientations. However, traditional convolutional neural networks are difficult to capture multi-scale and multi-directional patterns due to their isotropic convolutional kernels. The Gabor transform (GT) has been proven to be advantageous in scenarios involving scale and rotation changes. The fractional Fourier transform (FrFT) has been shown to effectively mitigate the impact of Doppler frequency shift.

[0029] In this embodiment, the optimized feature extraction block structure is as Figure 2 shown. The texture enhancement module based on FrGT is implemented using a dual-branch structure. The input feature map is split into two independent feature maps, which are processed by a convolutional fractional Gabor unit (FrGU) and a residual convolutional unit (RCU) respectively. The FrGU is responsible for extracting rotation-invariant texture features, that is, focusing on extracting the features of the image, while the RCU focuses on capturing the spatial patterns of the image. Finally, a gated attention dynamic fusion mechanism adaptively fuses the feature maps from the convolutional FrGU and RCU. This enables the model to better focus on learning discriminative feature representations while minimizing the interference of complex backgrounds.

[0030] In this embodiment, in the fractional Gabor convolutional unit: the input image is divided into a preset number of groups of images, and each group of images is processed through a convolutional fractional Gabor kernel to construct multiple directional features. After splicing the directional features, the output features of the fractional Gabor convolutional unit are obtained.

[0031] Specifically, to better interpret complex scenes containing semantic changes, the fractional Gabor transform (FrGT) is used to modulate the convolutional kernel of the convolutional neural network (CNN) to extract multi-scale and multi-directional features. To modulate Conv2D, the standard FrGT is extended to a two-dimensional discrete space. For an image , the convolutional FrGT, that is, the convolutional fractional Gabor kernel, can be expressed as: ; where, ; In the above formula, and are the coordinates in the spatial domain and the fractional domain respectively, represents a transform kernel, represents the transform angle, represents the transform order, is a Gaussian window function, and the superscript " -" denotes the complex conjugate, U and V respectively represent the number of samples in the fractional domain, and is the sampling interval, H and W are respectively the height and width of the image, represents the input image.

[0032] As Figure 2 shown, in the convolutional FrGT unit, the input image is divided into v groups , and all these groups are processed through convolutional fractional Gabor kernels to construct multi-directional features , and finally, the output feature is obtained through feature concatenation . The process is expressed as: ; In the above formula, represents a set of FrGK (fractional Gabor kernels) with different directions and scales, is a learned k×k convolutional kernel with i input channels and o output channels, represents channel concatenation.

[0033] In this embodiment, the residual convolutional unit is used to capture spatial features complementary to the texture enhancement features generated by the fractional Gabor convolutional unit. The residual convolutional unit is expressed as: ; In the above formula, represents the output feature of the residual convolutional unit, and respectively represent convolutional layers with 3×3 and 5×5 kernels, represents the input image of the residual convolutional unit, where .

[0034] In this embodiment, in the fusion unit, based on the output features of the fractional Gabor convolutional unit and the residual convolutional unit, the fusion weights are adaptively generated through gated attention dynamic fusion (GADF), and the output feature of the fusion unit is obtained according to the fusion weights and the output features. Among them, the formula for adaptively generating the fusion weights is as follows: ; In the above formula, represents the Softmax function, represents global average pooling, 、 respectively represent the output features of the fractional Gabor convolutional unit and the residual convolutional unit, 、 respectively represent the corresponding fusion weights of the output features of the fractional-order Gabor convolution unit and the residual convolution unit.

[0035] Furthermore, the fused feature is expressed as: ; In this embodiment, in the optimized feature extraction block, the input feature map is divided into two feature maps, which are processed in the convolutional fractional-order Gabor unit (FrGU) and the residual convolution unit (RCU) respectively. The FrGU extracts multi-directional features, while the RCU is used to capture the spatial patterns of the image. Finally, the feature maps of the convolutional FrGU and RCU are adaptively fused through a gated attention dynamic fusion mechanism, which can better focus on learning discriminative feature representations while reducing the interference of complex backgrounds. Finally, a texture enhancement feature extraction network is constructed by stacking the texture enhancement module based on FrGT (i.e., the optimized feature extraction block) and the downsampling layer to capture rich image features. Fast Spatial Pyramid Pooling (Fast-SPP) is incorporated into the top-level feature extraction layer to effectively capture multi-scale context information.

[0036] The texture enhancement feature extraction network in step S110 focuses on local texture representation, while in step S120, the relationship between the discrete scattering points of the aircraft is modeled through the spatial frequency feature interaction unit.

[0037] Global information plays a crucial role in SAR aircraft detection. Although the method based on convolutional neural network (CNN) has limitations in global information modeling, the method based on transformer performs excellently in capturing long-range dependencies, but its computational requirements are high. To solve this problem, a Spatial-Frequency Feature Interaction Module (SF2IM) is proposed, which can effectively fuse global context information with extremely low parameters and computational costs.

[0038] As Figure 3 shown, in the SF2IM, the multi-scale feature images (P 2 、P 3 、P 4 ) extracted from the texture enhancement feature extraction network are processed. C2PSA (Cross-Stage Pyramid Attention) is applied to the large-scale feature map to highlight key features and suppress redundant features. Then, hierarchical feature interaction is achieved by stacking Spatial Frequency Blocks (SFBs) equipped with upsampling and downsampling operations.

[0039] Furthermore, as Figure 3 shown, in the SFB, the input feature map is expanded through a 1×1 convolutional layer and normalized through LayerNorm, while Figure 3Po in 2 、Po 3 、Po 4 represent the output features after SFB processing. The processed features are split along the channel dimension and processed in parallel in different paths. Among them, the spatial domain path uses a 3×3 convolutional layer to capture local spatial patterns. In contrast, the frequency domain path combines frequency domain processing to establish a global receptive field, so as to comprehensively capture the mutual relationships.

[0040] Specifically, the fast Fourier transform (FFT) is used to transform the spatial features into the frequency domain, where the global information is encoded as frequency components in the frequency domain. A global learnable filter is used to capture the key frequency features. Subsequently, the frequency features are transformed back to the spatial domain through the inverse fast Fourier transform (IFFT). Finally, the spatial features and frequency features are fused through concatenation and layer normalization (LayerNorm) to capture the relationships between the discrete features of the aircraft. This dual-branch architecture can model local patterns and global dependencies simultaneously, ensuring a comprehensive feature representation for SAR-based aircraft detection.

[0041] Specifically, if the input feature map of SF2IM is , then the output feature map of SF2IM can be expressed as: ; In the above formula, represents the spatial feature extraction through the standard convolutional layer, which retains the local structural pattern. and represent the fast Fourier transform (FFT) operator and the inverse fast Fourier transform (IFFT) operator respectively, represents the feature extraction in the frequency domain.

[0042] In this embodiment, a fine-grained object detection model is constructed according to the above texture enhancement feature extraction network, spatial frequency feature interaction unit and detection unit, and the SAR target image to be detected is input into the fine-grained object detection model to achieve object detection.

[0043] In this embodiment, when training the fine-grained object detection model, the category-aware semantic consistency based on the shape mask is used as the constraint condition for detection refinement. Among them, the purpose of the category-aware semantic constraint based on the shape mask is to use prior knowledge to guide the model to better capture discriminative features. The category-aware semantic constraint based on the shape mask works together with the bounding box regression and classification branches in the proposed method to more accurately capture the intra-class consistency features and inter-class discriminative features.

[0044] In this embodiment, in order to implement category-aware semantic constraints based on shape masks, a mask prediction unit and a target type shape mask library are set in the detection unit. Among them, the target type shape mask library includes various types of target shape mask maps. The mask prediction unit generates a corresponding target mask prediction map according to the large-scale fusion feature image in the multi-scale fusion feature image, obtains the target shape mask map corresponding to the target in the original training image from the target type shape mask library, and finally constructs a loss function according to the target shape mask and the target mask prediction map to perform category-aware semantic constraints during the training process.

[0045] Specifically, before training, corresponding mask maps, that is, target shape mask maps, are constructed according to optical images of various aircraft categories, and various target shape mask maps are formed into a target type shape mask library. During the training process, the target shape mask map consistent with the target category in the original training image is extracted as the label map. The target shape mask map is created by using the aircraft shape mask extracted from the optical image in cooperation with the oriented bounding box (OBB) in the SAR image.

[0046] Such as Figure 4 shown, the mask prediction unit is a fully convolutional architecture that integrates hierarchical features from the backbone network through convolutional layers, upsampling, and skip connections to gradually restore the spatial resolution. Finally, a semantic category map with the same size as the original image is obtained.

[0047] Specifically, in the mask prediction unit, the large-scale fusion feature image is successively decoded by convolutional layers, upsampling, and skip connections to gradually restore the spatial resolution, and a target mask prediction map with the same size as the original training image is obtained.

[0048] Furthermore, in the mask prediction unit, after upsampling the large-scale fusion feature image P02 and concatenating it with the multi-scale feature image P0, it passes through a convolutional layer, then after upsampling and concatenating it with the multi-scale feature image P1, it passes through a convolutional layer, and then after upsampling and convolutional layers, the corresponding target mask prediction map is obtained.

[0049] In this embodiment, for the target instance B, the semantic constraint loss function constructed by the target shape mask and the target mask prediction map is expressed as: ; In the above formula, represents the number of pixels within the region of the object instance B, represents the one-hot encoded vector of the true label, represents the predicted probability that the nth pixel in the semantic category map belongs to the category c.

[0050] During the training phase, the optimal weights can only be obtained when the detection results are consistent with the set semantic constraints. That is to say, semantic constraints can effectively reduce false positives and improve the overall performance of the algorithm.

[0051] In this embodiment, end-to-end training is carried out through an object detection task (responsible for bounding box regression and class prediction) and an auxiliary task (responsible for semantic conditional constraints based on shape masks). Therefore, the total loss function can be expressed as: ; In the above formula, are the corresponding weight coefficients, and represent the classification loss and the bounding box regression loss respectively, represents the cross-entropy loss for semantic conditional constraints. For the object detection task, the classification loss adopts binary cross-entropy, and the bounding box regression loss includes the distribution focal loss (DFL) and the rotation IOU loss.

[0052] In this embodiment, an algorithm flow for the training phase is also provided, as shown in Table 1: Table 1

[0053] Furthermore, after training is completed, the mask prediction unit and the target type shape mask library in the measurement unit are removed.

[0054] As Figure 5 shown, it is a schematic diagram of the overall framework of the fine-grained object detection model.

[0055] In this paper, the effectiveness of this method is also demonstrated through simulation experiment results. First, an experimental dataset named SAR-RADD was constructed to illustrate the effectiveness of this method. A total of 66 single-polarization panoramic SAR images with a resolution of 1 meter were used. This dataset covers 16 typical airport scenes at home and abroad and 13 aircraft categories. At the same time, SAR-RADD follows the DOTA format. The original panoramic SAR images were adaptively cropped into 1024×1024 pixel samples through a sliding window, and adjacent windows overlapped by 512 pixels. In SAR-RADD, aircraft targets were annotated by experienced SAR image interpretation experts using the roLabelImg software. RoLabelImg is a graphical annotation tool that generates corresponding annotation files containing the position and category information of each aircraft. This dataset contains 1562 image samples, the training set contains 1164 images, and the test set consists of 398 images.

[0056] Furthermore, high-resolution optical images can provide detailed visual information of the aircraft. Based on the visual features of high-resolution optical images, different types of aircraft were manually annotated to construct a library of aircraft target shape masks with the nose uniformly facing the vertical direction. Since the appearances of different types of aircraft are similar cross structures, but there are obvious differences in component characteristics (wing sweep angle, engine configuration and distribution, and nose shape). These structural similarities and differences provide an effective basis for CNN to extract key discriminative features of aircraft targets. For the labeled SAR images, the orientation, geometric size, and spatial position parameters of the objects can be extracted by parsing the obb-based annotation files. Then, using affine transformations (including rotation, scaling, and translation), the masks in the aircraft target shape mask library are adaptively matched with the aircraft in the SAR images. Subsequently, coarse-grained masks are generated. These coarse-grained masks describe the aircraft target contours but may contain jagged edges due to the scaling operation. An edge smoothing algorithm is used to eliminate the jagged effect caused by scaling. Finally, a fine mask image is formed. The flowchart of mask generation is shown in Figure 6 shown.

[0057] Furthermore, the comparison with the mainstream algorithms is shown in Table 2. Table 2 presents the detection performance of the proposed method and the mainstream object detection algorithms for 13 types of aircraft in the SAR-RADD dataset. DenoDet is specifically designed for SAR image interpretation, emphasizing the saliency of the target while suppressing noise, thus improving the detection and recognition performance of aircraft. Therefore, it performs very well among the mainstream object detection methods. However, the proposed method achieves the optimal detection and classification performance, with an mAP of 79.3%, which is 4.7% and 6.9% better than DenoDet (74.6%) and YOLOv11 (72.4%) respectively. This demonstrates the superiority of the proposed method in the fine-grained detection task of aircraft targets in SAR images, as shown in Table 2:

[0058] In the above-mentioned fine-grained object detection method for SAR images, a texture enhancement backbone network based on fractional-order Gabor transform (FrGT) is first proposed, which combines the principle of fractional-order Fourier transform with multi-directional Gabor filtering to adaptively capture the discriminative features of the target while effectively suppressing the interference of speckle noise and background clutter. Then, a spatial-frequency feature interaction module (SF2IM) is introduced to perform hierarchical multi-scale feature fusion, so as to extract comprehensive and rich context information. Finally, a lightweight mask supervision module is proposed, which integrates category-specific shape prior knowledge through differentiable constraint learning, thereby effectively reducing the intra-class variance and inter-class ambiguity. This method innovatively designs a fractional-order Gabor convolutional network based on mask supervision for fine-grained aircraft detection in SAR images. By considering the scattering characteristics, frequency-domain knowledge, and category-specific shape prior knowledge of aircraft targets, this network achieves high-precision fine-grained aircraft detection. At the same time, by combining frequency-domain information with convolutional neural networks, a texture enhancement method based on fractional-order Gabor transform (FrGT) and SF2IM is developed to enhance the integrity and distinctiveness of feature extraction. The texture enhancement based on fractional-order Gabor convolution (FrGT) aims to adaptively enhance the saliency of multi-scale and multi-directional component features while suppressing the interference of background clutter. Then, SF2IM extracts more representative features of the target to improve the performance of fine-grained aircraft detection. A lightweight mask supervision module is also proposed to solve the problems of intra-class differences and inter-class ambiguity by introducing category-specific shape constraints, thereby improving the fine-grained classification effect of SAR images.

[0059] It should be understood that although Figure 1 the steps in the flowchart of Figure 1 are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover,

[0060] In one embodiment, as Figure 7 shown, a fine-grained object detection device for SAR images is provided, including: an SAR image acquisition module 200, a multi-scale feature extraction module 210, a multi-scale feature fusion module 220, and an object detection module 230, where: The SAR image acquisition module 200 is configured to acquire a SAR target image to be detected; The multi-scale feature extraction module 210 is configured to extract features from the SAR target image by using a texture-enhanced feature extraction network to obtain a multi-scale feature image. Among them, the texture-enhanced feature extraction network is obtained by optimizing the feature extraction blocks in the backbone network of the CSPDarknet network. The optimized feature extraction block includes a fractional-order Gabor convolutional unit and a residual convolutional unit that respectively extract image features, and a fusion unit that performs gated attention dynamic fusion on the image features extracted by these two units; The multi-scale feature fusion module 220 is configured to fuse the global context in the multi-scale feature image by using a spatial frequency feature interaction unit to obtain a multi-scale fusion feature image; The target detection module 230 is configured to perform target detection according to the multi-scale fusion feature image by using a detection unit. Among them, the classification branch and the regression branch in the detection unit respectively output the position information of the target and the category prediction information.

[0061] For the specific limitations of the fine-grained target detection device for SAR images, reference can be made to the limitations of the fine-grained target detection method for SAR images in the above text, which will not be elaborated here. Each module in the above-mentioned fine-grained target detection device for SAR images can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.

[0062] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 8 shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a fine-grained target detection method for SAR images. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the outer shell of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0063] Those skilled in the art can understand that Figure 8 The structure shown in Figure 8 is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0064] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented: Obtain a SAR target image to be detected; Use a texture-enhanced feature extraction network to extract features from the SAR target image to obtain a multi-scale feature image. Among them, the texture-enhanced feature extraction network is obtained by optimizing the feature extraction blocks in the backbone network of the CSPDarknet network. The optimized feature extraction blocks include a fractional-order Gabor convolution unit and a residual convolution unit that respectively extract image features, and a fusion unit that dynamically fuses the image features extracted by these two units through gated attention; Use a spatial frequency feature interaction unit to fuse the global context in the multi-scale feature image to obtain a multi-scale fused feature image; Use a detection unit to perform target detection according to the multi-scale fused feature image. Among them, the classification branch and the regression branch in the detection unit respectively output the position information of the target and the category prediction information.

[0065] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: Obtain a SAR target image to be detected; Use a texture-enhanced feature extraction network to extract features from the SAR target image to obtain a multi-scale feature image. Among them, the texture-enhanced feature extraction network is obtained by optimizing the feature extraction blocks in the backbone network of the CSPDarknet network. The optimized feature extraction blocks include a fractional-order Gabor convolution unit and a residual convolution unit that respectively extract image features, and a fusion unit that dynamically fuses the image features extracted by these two units through gated attention; Use a spatial frequency feature interaction unit to fuse the global context in the multi-scale feature image to obtain a multi-scale fused feature image; Use a detection unit to perform target detection according to the multi-scale fused feature image. Among them, the classification branch and the regression branch in the detection unit respectively output the position information of the target and the category prediction information.

[0066] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0067] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0068] The above-described embodiments merely represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application should be subject to the appended claims.

Claims

1. A fine-grained target detection method for SAR images, characterized in that: The method comprises: Acquire a SAR target image to be detected; Using a texture enhancement feature extraction network to extract features from the SAR target image to obtain a multi-scale feature image, wherein the texture enhancement feature extraction network is obtained by optimizing the feature extraction block in the CSPDarknet backbone network, and the optimized feature extraction block includes a fractional-order Gabor convolution unit and a residual convolution unit for respectively extracting image features, and a fusion unit for performing gated attention dynamic fusion of the image features extracted by the two units; Using a spatial frequency feature interaction unit to fuse the global context in the multi-scale feature image to obtain a multi-scale fused feature image; A detection unit is used to perform target detection based on the multi-scale fusion feature image, wherein a classification branch and a regression branch in the detection unit respectively output location information and category prediction information of the target.

2. The fine-grained target detection method for SAR images according to claim 1, characterized in that: In the fractional-order Gabor convolution unit: The input image is divided into a preset number of groups of images, and each group of images is processed by convolution with a fractional-order Gabor kernel to construct multi-directional features; After concatenating the features in each direction, the output features of the fractional-order Gabor convolution unit are obtained.

3. The fine-grained target detection method for SAR images according to claim 2, characterized in that: The convolution fractional-order Gabor kernel is expressed as: ; in, ; In the above formula, and are the coordinates in the spatial domain and the fractional domain respectively, represents a transformation kernel, Indicates the transformation angle, represents the transformation order, is a Gaussian window function, the superscript " - " represents the complex conjugate, U and V represent the number of samples in the fractional domain, and is the sampling interval, H and W are the height and width of the image respectively, Represents the input image.

4. The fine-grained target detection method for SAR images according to claim 2, characterized in that: The residual convolution unit is expressed as: ; In the above formula, represents the output feature of the residual convolution unit, and They represent convolution layers with 3×3 and 5×5 kernels respectively. Represents the input image of the residual convolution unit.

5. The fine-grained target detection method for SAR images according to claim 2, characterized in that: In the fusion unit, a fusion weight is adaptively generated based on the output features of the fractional-order Gabor convolution unit and the residual convolution unit, and the output features of the fusion unit are obtained according to the fusion weight and the output features; Among them, the adaptive generation of fusion weights uses the following formula: ; In the above formula, represents the Softmax function, represents global average pooling, , Respectively represent the output features of the fractional-order Gabor convolution unit and the residual convolution unit, , They respectively represent the corresponding fusion weights of the output features of the fractional-order Gabor convolution unit and the residual convolution unit.

6. The fine-grained target detection method for SAR images according to any one of claims 1 to 5, characterized in that: Constructing a fine-grained target detection model based on the texture enhancement feature extraction network, the spatial frequency feature interaction unit and the detection unit; The SAR target image to be detected is input into a fine-grained target detection model to achieve target detection.

7. The fine-grained target detection method for SAR images according to claim 6, characterized in that: When training the fine-grained object detection model: The detection unit is provided with a mask prediction unit and a target type shape mask library, wherein the target type shape mask library includes multiple types of target shape mask images; The mask prediction unit generates a corresponding target mask prediction image according to the large-scale fusion feature image in the multi-scale fusion feature image; Obtaining a target shape mask image corresponding to the target in the original training image from the target type shape mask library; By constructing a loss function according to the target shape mask map and the target mask prediction map, category-aware semantic constraints are performed during the training process.

8. The fine-grained target detection method for SAR images according to claim 7, characterized in that: The mask prediction unit: The large-scale fusion feature image is sequentially subjected to convolutional layers, upsampling and skip connections to gradually decode features and restore spatial resolution, thereby obtaining the target mask prediction image of the same size as the original training image.

9. A fine-grained target detection device for SAR images, characterized in that: The device comprises: A SAR image acquisition module is used to acquire a SAR target image to be detected; A multi-scale feature extraction module, used for extracting features from the SAR target image using a texture enhancement feature extraction network to obtain a multi-scale feature image, wherein the texture enhancement feature extraction network is obtained by optimizing the feature extraction block in the CSPDarknet network backbone network, and the optimized feature extraction block includes a fractional-order Gabor convolution unit and a residual convolution unit for respectively extracting image features, and a fusion unit for performing gated attention dynamic fusion of the image features extracted by the two units; A multi-scale feature fusion module, used to fuse the global context in the multi-scale feature image using a spatial frequency feature interaction unit to obtain a multi-scale fused feature image; The target detection module is used to use the detection unit to perform target detection based on the multi-scale fusion feature image, wherein the classification branch and regression branch in the detection unit respectively output the target's position information and category prediction information.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Multi-resolution SAR (Synthetic Aperture Radar) target identification method and device based on scale perception domain adaptation

    CN115578633A

  • Large-scale urban village extraction method based on token mask mechanism

    CN116310628A

  • Hyperspectral remote sensing image classification method based on self-attention context network

    US20230260279A1

Cited By

  • Optical guidance SAR (Synthetic Aperture Radar) target detection method based on frequency domain enhancement and dynamic mask

    CN120374604A

  • Video cross-modal pedestrian re-identification method based on frequency domain perception and space-time aggregation

    CN120496132A