SAR image target intelligent detection method based on sparse DETR model

Through the sparse DETR model of the SAR image target intelligent detection method, the lightweight backbone network and Transformer encoder sparse feature processing are used to solve the problem of difficult distinction between targets and backgrounds in SAR images, and achieve high-precision moving target recognition in complex electromagnetic environments.

CN120655898APending Publication Date: 2025-09-16CHINESE PEOPLES LIBERATION ARMY UNIT 96901
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510763794.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-16

Smart Images

  • Figure CN120655898A_ABST
    Figure CN120655898A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of target detection, and particularly discloses an SAR image target intelligent detection method based on a sparse DETR model, and the method comprises the steps: obtaining a training image data set of a moving target; the method comprises the following steps: performing feature extraction on an input image through a lightweight backbone network to form sparse Query representation; the sparse features are input into a Transform encoder with a deformable attention mechanism, and modeling is carried out; the output of the encoder is sent to the Transform decoder; parameters of the SAR image moving target intelligent detection model are iteratively optimized, and training is completed; and inputting a to-be-detected moving target image into the trained model, and outputting a target detection bounding box, a classification label and a corresponding confidence coefficient in combination with the drawing script to obtain a moving target detection result. According to the invention, background, interference, noise and moving target objects can be accurately distinguished, and accurate identification of the moving target in a complex electromagnetic environment is realized. According to the method, the target feature detection precision is improved, and meanwhile, the unnecessary calculation burden is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of target detection, and specifically discloses an intelligent target detection method for SAR images based on a sparse DETR model. Background Art

[0002] With the recent development of Synthetic Aperture Radar (SAR) technology, its application in target detection has garnered widespread attention. SAR, an active radar, acquires high-resolution images of targets by transmitting pulse signals and receiving echoes. Compared to optical imaging, SAR offers the advantages of all-weather operation and weather resistance. However, due to the complex radar physics and environmental interference, SAR image processing and target recognition face numerous challenges. SAR image processing involves preprocessing and enhancing raw data to improve image quality and target recognition accuracy. Target recognition is a core task in SAR image processing, encompassing target detection, segmentation, and recognition. Target detection methods are typically based on pixel values, texture, or shape, while segmentation separates targets from the background for easier analysis. Target recognition methods achieve automatic classification through matching against a database. However, SAR still has many limitations. For example, the background, influenced by various factors, constantly fluctuates randomly and complexly. When using SAR to detect moving targets, these fluctuations produce echoes with complex distributions. Consequently, this background clutter poses challenges to moving target recognition.

[0003] The application of SAR image processing and moving target detection technologies still faces several challenges and difficulties. First, target information in SAR images is often heavily intertwined with background clutter, making it difficult to distinguish target edges and details. Second, due to differences in radar physics and parameter settings, SAR images acquired by different radar systems vary significantly. Finally, SAR images differ significantly from optical images in terms of texture and lighting.

[0004] Traditional moving target detection methods rely mainly on manually designed features and classifiers. However, these methods often fail to achieve ideal detection results when processing SAR images affected by complex and variable interference. Summary of the Invention

[0005] Aiming at the problem of SAR moving target detection in complex electromagnetic environments, the present invention proposes an intelligent target detection method for SAR images based on a sparse DETR model, which includes the following steps:

[0006] Step 1: Obtain a training image dataset of moving targets, and input the images and label data into an intelligent moving target detection model for SAR images, where the intelligent moving target detection model for SAR images is built based on a sparse DETR model.

[0007] Step 2: The input image is passed through a lightweight backbone network for feature extraction, capturing the contextual information of moving objects in the image and outputting a multi-scale feature map. Position embedding is also introduced to preserve spatial location information. The extracted features are fed into a scoring network to calculate the significance score for each spatial location. Only the tokens with the highest scores are selected to form a sparse query representation.

[0008] Step 3: The sparse features are input into the Transformer encoder with a deformable attention mechanism for modeling;

[0009] Step 4: The encoder output is fed into the Transformer decoder, where deformable cross-attention is used to cross-learn features between the query and the encoder output. This is then fed into the feedforward network to improve feature representation. The reinforced tokens output by the decoder are fed into the detection head to generate target categories and bounding box predictions for each query. The predictions are aligned with the ground truth using a bipartite graph matching algorithm, and the classification loss and bounding box regression loss are calculated. This overall loss is then used for backpropagation to optimize network parameters.

[0010] Step 5: Continue to perform steps 2 to 4, iteratively optimize the parameters of the SAR image moving target intelligent detection model until the loss function converges or the change amplitude is less than a set threshold, thereby completing the training of the SAR image moving target intelligent detection model;

[0011] Step 6: Input the moving target image to be detected into the trained SAR image moving target intelligent detection model, and combine it with the drawing script to output the target detection bounding box, classification label and corresponding confidence to obtain the moving target detection result.

[0012] Preferably, in step 1, the sparse DETR model includes a lightweight backbone network, a Transformer-based encoder-decoder and a detection head.

[0013] Preferably, in step 1, after obtaining a training image dataset of a moving target, images in the training image dataset are annotated and different interference types are added to construct a SAR image dataset in a complex electromagnetic environment.

[0014] Preferably, in step 2,

[0015] The lightweight backbone network is a ResNet-50 network, and the input image is subjected to an initial convolutional layer and a maximum pooling operation to extract underlying features;

[0016] The image features are sequentially passed through the four stages of the ResNet-50 network. Each stage contains multiple residual modules, and each module consists of multiple convolutional layers and residual connections.

[0017] After obtaining the multi-scale feature map, the output two-dimensional feature map is flattened into a sequence form, and the feature vector at each position corresponds to a spatial position in the image;

[0018] Constructing a two-dimensional position encoding matrix that matches the length and number of channels of the sequence to embed the spatial position information of each token, wherein the position encoding adopts a fixed position encoding constructed by sine and cosine functions or a trainable position encoding based on learnable parameters;

[0019] The positional encoding is concatenated with the original feature vector to obtain an input sequence containing spatial perception capabilities, which is used as the input of the Transformer encoder.

[0020] The position-encoded features are input into the scoring network, and some significant tokens are selected for refinement, i.e., encoder tokens are sparsely processed.

[0021] Preferably, in step 3, the sparse query tokens selected by the scoring network are used as input, and each encoder layer performs the following operations in sequence:

[0022] Through the deformable self-attention module, each token adaptively selects a local area as the focus range according to its position and samples contextual information from the original feature map;

[0023] Perform feature conversion and nonlinear mapping on the updated token representation through a feedforward neural network;

[0024] Repeatedly stack xN in multiple encoder layers e The i-th encoder layer updates the feature x by i-1 :

[0025]

[0026] Where DefAttn is deformable attention, LN is layer normalization, and FFN is feedforward network.

[0027] Preferably, in step 4, the Transformer decoder receives the sparse tokens output by the encoder and a set of learnable object queries. Each layer of the decoder first models the dependencies between different queries through a self-attention module, and then performs a deformable cross-attention operation to align the queries with the tokens output by the encoder, guiding each query to focus on the potential target area.

[0028] Compared with the prior art, the present invention has the following advantages:

[0029] Compared to traditional convolutional neural networks, the sparse DETR model employed in the deep learning portion of the present method utilizes a Transformer architecture. Through this architecture, the present method leverages a self-attention mechanism to capture long-range contextual relationships between targets and backgrounds in an image. This enables the present method to learn more significant features within and between images, accurately distinguishing between background, interference, noise, and moving objects, and achieving precise recognition of moving targets in complex electromagnetic environments.

[0030] The sparse query mechanism proposed by the method of the present invention selectively processes significant encoder tokens, so that the method of the present invention only focuses on key feature areas, thereby allowing the method of the present invention to focus on areas where the target exists and ignore background areas affected by noise, thereby improving the accuracy of target feature detection while reducing unnecessary computational burden. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 This is a flow chart of the sparse DETR model in the present invention;

[0032] Figure 2 Schematic diagram of the module structure of the sparse DETR model in the present invention;

[0033] Figure 3 Schematic diagram of the visualization results of test images under different interference conditions. DETAILED DESCRIPTION

[0034] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0035] In a specific embodiment of the present invention, the present invention constructs a SAR image target intelligent detection model based on a sparse DETR model. Figure 1 As shown, after the image input is extracted by the backbone network and embedded with position, the scoring network first selects key positions to generate sparse queries, which are then fed into an encoder consisting of a deformable self-attention and feedforward network. The encoding result is input into the decoder (including self-attention and deformable cross-attention). Finally, the detection head predicts the bounding box and category, supplemented by auxiliary head supervision and bipartite graph matching to optimize the overall detection performance. The method of the present invention includes:

[0036] Step 1: Obtain a training image dataset of moving targets, perform image annotation processing on it, and add different interference types to the original dataset to construct a SAR image dataset in a complex electromagnetic environment. Select a certain batch size of images and their label data as the input of the subsequent moving target intelligent detection model;

[0037] Step 2: The input image first passes through a lightweight backbone network for feature extraction, efficiently capturing the contextual information of moving objects in the image and outputting a multi-scale feature map. Position embedding is also introduced to preserve spatial location information. The extracted features are then fed into a scoring network, which calculates a saliency score for each spatial location and selects only a small number of tokens with the highest scores to form a sparse query representation.

[0038] Step 3: The sparse features are input into the Transformer encoder with a deformable attention mechanism for deep modeling.

[0039] The sparse tokens are input into the Transformer encoder, which consists of a deformable self-attention mechanism and a feedforward neural network, for feature modeling, gradually refining the representation of the sparse tokens. The encoder consists of multiple repeated layers stacked together.

[0040] Step 4: The encoder output is fed into the Transformer decoder, which uses self-attention to model the internal dependencies between queries.

[0041] Deformable cross attention performs feature cross-learning on the query and encoder output, guiding the query to focus on areas where the target may be located. This is then fed into the feedforward network to further enhance feature expression capabilities.

[0042] The enhanced tokens output by the decoder are fed into the detection head to generate the target category and bounding box prediction for each query. At the same time, auxiliary heads are introduced in the encoder and decoder stages to provide additional supervision signals for the intermediate features.

[0043] Finally, all prediction results are aligned with the real annotations through a bipartite graph matching algorithm (such as Hungarian matching), and the classification loss and bounding box regression loss are calculated. The overall loss is used for backpropagation to optimize the parameters of the entire network.

[0044] Step 5: Continue to execute steps 2 to 4, iteratively optimize the model parameters until the loss function converges or the change amplitude is less than the set threshold, complete the model training, and achieve model establishment.

[0045] Step 6: Input the moving target image to be detected into the trained moving target intelligent detection model, and combine it with the drawing script to output the target detection bounding box, classification label and corresponding confidence to obtain the moving target detection result.

[0046] In a specific embodiment of the present invention, Figure 2 As shown in FIG, the sparse DETR model in the present invention includes: a lightweight backbone network, a Transformer-based encoder-decoder, and a detection head that makes the final prediction output.

[0047] In a specific embodiment of the present invention, in step 1 of the present invention, obtaining a training image dataset of a moving target includes:

[0048] The dataset SRSDD-v1.0 (Remote Sensing, Lei, S.; Lu, D.; Qiu, X.; Ding, C. SRSDD-v1.0: A High-Resolution SAR Rotation Ship Detection Dataset. RemoteSens. 2021, 13, 5104. https: / / doi.org / 10.3390 / rs13245104) was used. The original images in this dataset were acquired from the GF-3 SAR satellite in spotlight (SL) mode, with a resolution of 1 meter in both range and azimuth, and both HH and VV polarization modes. 666 1024×1024 pixel slices were cropped from the original images. Each slice was adjusted for brightness and contrast to facilitate annotation. The dataset contains 2884 targets across six categories, with 63.1% of the data set containing scenes with complex backgrounds and high levels of interference.

[0049] The present invention aims to improve the shortcomings of the SRSDD-v1.0 dataset to enhance the effect of model training. Since the size of a single image in this dataset is relatively large, the proportion of moving targets in the image is relatively small. If the original dataset is used directly for training, it will be more difficult for the model to learn the target features, and the accuracy and convergence speed of the model training will also be affected. Therefore, the present invention first divides each image in the SRSDD-v1.0 dataset into 16 slices of 256 pixels × 256 pixels on average, and then performs subsequent processing on the sliced ​​dataset.

[0050] The present invention further adds different interference types to the data set to construct a SAR image data set in a complex electromagnetic environment.

[0051] In a specific embodiment of the present invention, a simulated background clutter data set is constructed. Since background clutter is affected by various factors, the echo of electromagnetic waves emitted by the synthetic aperture radar to the background will contain various interference information. In view of the randomness of background clutter, a probability distribution model is usually used to model the background clutter. The Weibull distribution can effectively simulate the amplitude statistical characteristics of background clutter, and compared to another K distribution commonly used to simulate background clutter, the calculation of the Weibull distribution is simpler. Therefore, the present invention uses the Weibull distribution, which has an equally good effect on background clutter simulation and is easy to calculate, to generate simulated background clutter interference. The probability density function of the Weibull distribution is as follows:

[0052]

[0053] Here, λ > 0 is the scale parameter, and k > 0 is the shape parameter, which controls the shape of the tail of the distribution function. Noise was added to 36.9% of the data in the dataset to simulate the effects of background clutter on SAR images. The specific steps are: First, a noise map with the same size as the original image and amplitude conforming to a Weibull distribution was generated, with parameters k = 0.8 and λ = 1.6. The noise map was then overlaid on the original image to simulate the effects of additive noise background clutter on SAR imaging, thereby simulating background clutter interference.

[0054] In a specific embodiment of the present invention, a simulated active interference data set is constructed. Active interference to SAR refers to the influence of interference signals actively generated by interference equipment on the SAR system. Active interference will reduce the resolution of SAR images and blur the target features, thereby affecting the target detection capability of the SAR system. Active interference can be roughly divided into suppression interference and deception interference. Among them, suppression interference refers to covering the radar echo by emitting high-power noise signals, thereby making it difficult for synthetic aperture radar to detect the target information to be detected; deception interference is also called simulated interference, which mainly uses a jammer to intercept SAR signals and forward and generate interference signals with frequencies, waveforms and modulation methods similar to or related to the radar signals emitted by the SAR system, thereby achieving the purpose of deceiving the SAR system and affecting its normal detection work.

[0055] Suppressive jamming is the easiest and most widely used form of jamming. For SAR, suppressive jamming enhances the background noise, blurring the SAR image and even making it impossible to identify useful targets. Since suppressive jamming has no correlation with the SAR signal, its effectiveness depends on the power of the noise signal transmitted by the jammer. Therefore, Gaussian noise can be used to simulate suppressive jamming. The specific steps for adding simulated suppressive jamming to a dataset are: A Gaussian distribution function with a mean of 0 and a variance of 0.01 is used to randomly generate a noise pattern that conforms to a Gaussian distribution and has the same size as the unjammed image in the dataset. The original image and the noise pattern are then superimposed to simulate suppressive active jamming during SAR imaging. Coherent noise refers to jamming signals generated at the same frequency and phase, which makes the target's echo signal difficult to distinguish. Coherent noise can be used to simulate deceptive jamming.

[0056] In one embodiment of the present invention, a combined interference dataset is constructed. In actual operating environments, synthetic aperture radars are not affected by a single type of interference, but rather by a complex mix of these various interference types. To simulate the complex electromagnetic environment in which synthetic aperture radars operate, the aforementioned methods for adding simulated interference to datasets should be combined to construct a dataset under the combined influence of multiple interferences. This improves the accuracy and robustness of the model's detection in these complex electromagnetic environments.

[0057] In a specific embodiment of the present invention, in step 2 of the present invention, the input image is subjected to feature extraction and position encoding by a lightweight backbone network, including:

[0058] The image feature extraction module uses the ResNet-50 network as the backbone network to extract multi-scale semantic features from the input image. First, the input image passes through the initial convolution layer and maximum pooling operation to extract the underlying features; then, the image features pass through the four stages of ResNet-50 (stage1 to stage4) in sequence. Each stage contains multiple residual blocks (Residual Block), and each module is composed of multiple convolution layers and residual connections, which effectively alleviates the gradient disappearance and enhances the feature transfer capability. Among them, as the depth of the network increases, the spatial resolution of the feature map gradually decreases and the semantic information gradually increases, and finally a high-level semantic feature map is output in stage4. The above multi-scale features can be used as the input of the subsequent Transformer encoder, and the position information is integrated for subsequent attention mechanisms and target detection tasks.

[0059] After obtaining the multi-scale feature map, the model further introduces a positional encoding mechanism to embed the position information into the feature representation. First, the two-dimensional feature map output by the backbone network is flattened into a sequence form, and the feature vector at each position corresponds to a spatial position in the image. Subsequently, a two-dimensional position encoding matrix that matches the length and number of channels of the sequence is constructed to embed the spatial position information of each token. The position encoding can be a fixed position encoding constructed by sine and cosine functions, or a trainable position encoding based on learnable parameters. Among them, the sine and cosine encoding method performs multi-frequency sine and cosine transforms on each position index to generate a vector with unique spatial position information; while the learnable position encoding regards each position as an independent parameter and automatically learns the optimal spatial expression through training. Finally, the position encoding is spliced ​​with the original feature vector to obtain an input sequence containing spatial perception capabilities, which is used as the input of the Transformer encoder, thereby realizing the modeling and utilization of spatial position information in the image.

[0060] The position-encoded features are input into the scoring network to select a small number of significant tokens for refinement, thereby reducing the complexity of attention calculation in the subsequent encoder and retaining key information, which is called encoder tokens sparsification.

[0061] The encoder tokens sparsification process is as follows: Encoder tokens sparsification refers to selectively refining only a small part of the encoder tokens, which are obtained from the backbone network feature map X based on certain criteria. feat The feature map X is not updated during this process. feat The value of will not be changed when passing through the encoder. This method proposes a method for measuring X feat The scoring network of the importance of each token in And give the set of ρ-significant regions Definition: For a given retention ratio ρ, ρ-significant region set The set of the top ρ% tokens with the highest scores is the set of ρ-significant regions. The relationship between the set of ρ-significant regions and the number of elements in the query set is as follows:

[0062]

[0063] In a specific embodiment of the present invention, in step 3 of the present invention, the sparsely-sampled features are input into a Transformer encoder with a deformable attention mechanism for deep modeling, including:

[0064] In the encoder stage, the sparse query tokens selected by the scoring network are first taken as input. Each encoder layer performs the following operations in sequence: First, through the deformable self-attention module, each token adaptively selects a local area as the focus range according to its position, samples context information from the original feature map, and realizes efficient local feature interaction; then, the updated token representation is subjected to feature transformation and nonlinear mapping through the feedforward neural network (FFN) to improve its expressive power. The whole process is repeated in multiple encoder layers stacked xN e times to gradually refine the semantic representation of the token. The i-th encoder layer updates the feature x by i-1 :

[0065]

[0066] Here, DefAttn refers to deformable attention, LN refers to layer normalization, and FFN refers to feedforward network. The values ​​of unselected tokens also pass through the encoder layer, so when updating selected tokens, they can be referenced as keys. This means that unselected tokens can pass information to selected tokens while minimizing computational cost, and they do not lose their value information in the process.

[0067] During training, some encoder layer outputs are fed into auxiliary heads to perform classification and bounding box regression supervision on intermediate features. The introduction of auxiliary heads provides the model with richer gradient signals, guiding early layers to learn more discriminative representations, accelerating model convergence and improving overall performance. Ultimately, the encoder outputs sparser and more semantically rich token representations, which serve as input to the decoder and provide strong support for object detection tasks.

[0068] In a specific embodiment of the present invention, in step 4 of the present invention, the output of the encoder is fed into the Transformer decoder and the detection head including:

[0069] In the decoder stage, the model receives the sparse tokens output by the encoder and a set of learnable object queries as input. Each decoder layer first models the dependencies between different queries through the self-attention module, and then performs a deformable cross-attention operation to align the query with the tokens output by the encoder, guiding each query to focus on the potential target area. By continuously iterating xN dThe present invention proposes a method to determine the significant tokens using the cross attention map of the Transformer decoder, so as to obtain the feature map X extracted by the backbone network. feat Obtain a significant set of tokens

[0070] As training progresses, the Transformer decoder gradually increases its attention to a subset of encoder output tokens that are beneficial for object detection, so the decoder's cross-attention map can be used for saliency evaluation. In addition, the method determines salient tokens by training a scoring network that predicts saliency pseudo-true values ​​defined by the decoder's cross-attention map and uses it to determine which encoder tokens should be further dynamically refined.

[0071] In a specific embodiment of the present invention, in order to determine the encoder X feat The method aggregates the decoder cross attention between all object queries and encoder outputs to generate a map of the same size as the backbone network's feature map, which is defined as the decoder cross attention map (DAM). In the case of deformable attention, for each encoder token, the corresponding value of the DAM is obtained by accumulating the attention weights of the decoder object queries whose attention offset points to the encoder output tokens.

[0072] In a specific embodiment of the present invention, in order to train a scoring network to determine significant tokens, the method binarizes the DAM so that only the top ρ% encoder tokens ranked highest by attention weight are retained, thereby finding a small number of encoder tokens that are most referenced by the decoder. The binarized DAM indicates whether each encoder token is included in the top ρ% encoder tokens with the most references. Subsequently, the model sets up a 4-layer scoring network g to predict the probability of given encoder tokens being included in the top ρ% tokens with the most references, and trains the network by minimizing the binary cross entropy (BCE) loss between the prediction and the binarized DAM. The specific loss function is shown as follows:

[0073]

[0074] Where, represents the binarized DAM value of the i-th encoder tokens. Empirical observations show that DAM optimization is very stable even in the early stages of training.

[0075] In a specific embodiment of the present invention, the present invention adopts an encoder auxiliary head optimization strategy. The auxiliary head receives the token representation features from the intermediate layer, and after the feature transformation is performed by a lightweight feedforward network, it outputs the category prediction and bounding box regression results respectively. The prediction results are matched with the real annotations to calculate the classification loss and regression loss, and are used as auxiliary supervision signals to participate in back propagation together with the backbone loss, thereby strengthening the inter-layer collaboration and feature guidance during the network training process. Since the number of encoder tokens is much larger than the decoder tokens, in order to reduce the computational cost, other DETR models only add auxiliary detection heads to the decoder layer. However, in the present invention, only part of the encoder tokens are refined by the encoder, so adding an auxiliary head only to the sparse encoder will not cause too much extra burden. Applying the auxiliary detection head and Hungarian loss on selected tokens can alleviate the gradient vanishing problem and stabilize the convergence of the deep encoder, and correspondingly improve the detection performance.

[0076] The final moving target detection results are obtained through the classification network (detection head). To achieve a one-to-one correspondence between the predicted results and the true annotated targets, a bipartite graph matching strategy based on the Hungarian algorithm is adopted. Specifically, a cost matrix is ​​first constructed based on the classification error and bounding box regression error between the predicted output and the annotated box. The predicted results and the true target are regarded as nodes on both sides of a bipartite graph, and the weight of the edge is the matching cost. Subsequently, by solving the minimum cost matching of this bipartite graph, the optimal one-to-one assignment relationship is determined, and loss calculation and backpropagation are performed accordingly, thereby improving the stability and accuracy of target detection training.

[0077] The method and application effect of the present invention are described and evaluated in detail below through a specific embodiment.

[0078] 1. Experimental environment

[0079] This invention uses Linux (kernel 5.4+) as the operating system, supports Pytorch 1.12.1+cu113, and the hardware configuration includes 32GB of memory, 8GB of hard disk space, and NVIDIA RTX 3090GPU to ensure good operating performance.

[0080] 2. Hyperparameter settings

[0081] Table I Model hyperparameter settings

[0082] Hyperparameters Parameter value Hyperparameters Parameter value lr 0.0002 lr_drop 40 lr_backbone 2e-05 clip_max_norm 0.1 lr_linear_proj_mult 0.1 enc_layers 6 batch_size 16 dec_layers 6 weight_decay 0.0001 set_cost_class 2 epochs 200 set_cost_bbox 5 dim_feedforward 1024 set_cost_giou 2 hidden_dim 256 cls_loss_coef 2 num_queries 300 bbox_loss_coef 5 dropout 0.1 giou_loss_coef 2 dec_n_points 4 focal_alpha 0.25 enc_n_points 4 mask_prediction_coef 1 num_workers 8 seed 42

[0083] 3. Evaluation indicators

[0084] This paper uses mAP50 as an evaluation metric for moving target recognition methods. mAP50 is a common evaluation metric in the field of target detection, measuring the average accuracy of a model at an Intersection-in-Union (IoU) threshold of 0.5. IoU refers to the ratio of the intersection area of ​​the predicted bounding box to the union area of ​​the true bounding box. When calculating mAP50, the area enclosed by the PR curve and the coordinate axes at an IoU threshold of 0.5 is calculated for each category. Finally, the average of these values ​​across categories is taken to obtain the average accuracy of the model as a whole.

[0085] 4. Quantification results

[0086] Table II Quantitative results of the model on different datasets

[0087] Dataset mAP50 SRSDD-v1.0 86.206% SRSDD-v1.0+ Background Clutter 84.866% SRSDD-v1.0+ Suppressive Active Jamming 84.076% SRSDD-v1.0+ deceptive active jammer 82.931% SRSDD-v1.0+ combined interference 80.574%

[0088] Table II shows the quantization results of this model on different data sets. It can be seen from this table that the model proposed in this invention can effectively suppress the impact of interference on target detection in complex electromagnetic environments containing different types of interference. The quantization results show that this model has good anti-interference ability, proving that the Transformer architecture proposed in this model can effectively obtain the long-distance contextual relationship between the target and the background, thereby achieving accurate distinction between the target and the noise. For the case of suppressive active interference or background passive interference with complex noise, the quantization results confirm that the sparse encoder token method proposed in this invention can effectively overcome background noise interference, thereby achieving an improvement in model detection accuracy.

[0089] 5. Visualize the results

[0090] like Figure 3 As shown, the present invention maintains high accuracy in identifying moving targets in complex electromagnetic environments containing different interference types. The visualization results intuitively demonstrate that the present invention can accurately identify and classify targets even in complex electromagnetic environments with low contrast between the target and the background. This demonstrates that the proposed sparse encoder tokens enable the model to focus on the target area and eliminate the influence of redundant background information.

[0091] 6. Conclusion

[0092] In response to the problem of the mixing of target information and background clutter in SAR images, the present invention proposes a transformer-based sparse DETR model, which is characterized by the sparsification of encoder tokens and the search for significant encoder tokens. Encoder token sparsification refers to obtaining a small number of tokens from feature maps that meet certain standards, and the encoder only focuses on refining these tokens. Encoder token sparsification helps the model filter out noise caused by interference in SAR images and enables the model to focus on the target area in the image, thereby improving the accuracy of model detection and training speed. Finding significant encoder tokens specifically involves introducing a scoring network that predicts the pseudo-true situation of significance defined by the decoder cross-attention map, and based on this the network determines which tokens should be further refined. It enables the model to automatically focus on important feature areas near the target, reduce the impact of noise areas on target detection, and improve the accuracy and convergence speed of the model.

[0093] The present invention trains the model and verifies the effect on a constructed SAR image dataset with multiple interferences. The quantitative results obtained show that the network structure used in the present invention can improve the accuracy of SAR image moving target recognition under the influence of complex electromagnetic environments. The experiment proves the effectiveness of the innovative structure proposed by the present invention using the network. The visualization experiment results show that the model used in the present invention can accurately detect moving targets in SAR images under multiple interferences. However, it can be seen from the visualization results that the model still has certain limitations in detecting the boundary between the target and the background. While detecting the target and making a correct prediction of its category, there is sometimes a gap between the predicted boundary box and the real box. Therefore, the future work direction of the present invention is to study methods to improve the model's ability to learn target boundary features while ensuring the existing results, and further improve the model's ability to detect target boundaries.

[0094] The above content is only for explaining the technical idea of ​​the present invention and cannot be used to limit the protection scope of the present invention. Any changes made on the basis of the technical solution in accordance with the technical idea proposed by the present invention shall fall within the protection scope of the claims of the present invention.

[0095] It should be noted that the terms "first", "second", etc. in the description, claims, and drawings of the present invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way are interchangeable where appropriate so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or that are inherent to these processes, methods, products, or apparatus.

Claims

1. An intelligent target detection method for SAR images based on a sparse DETR model, characterized in that: The steps include: Step 1: Obtain a training image dataset of moving targets, and input the images and label data into an intelligent moving target detection model for SAR images, where the intelligent moving target detection model for SAR images is built based on a sparse DETR model. Step 2: The input image is passed through a lightweight backbone network for feature extraction, capturing the contextual information of moving objects in the image and outputting a multi-scale feature map. Position embedding is also introduced to preserve spatial location information. The extracted features are fed into a scoring network to calculate the significance score for each spatial location. Only the tokens with the highest scores are selected to form a sparse query representation. Step 3: The sparse features are input into the Transformer encoder with a deformable attention mechanism for modeling; Step 4: The output of the encoder is fed into the Transformer decoder, and the deformable cross attention performs feature cross-learning between the query and the encoder output; The input is fed into the feedforward network to improve the feature expression capability; the enhanced tokens output by the decoder are input into the detection head to generate the target category and bounding box prediction corresponding to each query; The prediction results are aligned with the real annotations through a bipartite graph matching algorithm, and the classification loss and bounding box regression loss are calculated. The overall loss is used for backpropagation to optimize the network parameters. Step 5: Continue to perform steps 2 to 4, iteratively optimize the parameters of the SAR image moving target intelligent detection model until the loss function converges or the change amplitude is less than a set threshold, thereby completing the training of the SAR image moving target intelligent detection model; Step 6: Input the moving target image to be detected into the trained SAR image moving target intelligent detection model, and combine it with the drawing script to output the target detection bounding box, classification label and corresponding confidence to obtain the moving target detection result.

2. The method according to claim 1, characterized in that In step 1, the sparse DETR model includes a lightweight backbone network, a Transformer-based encoder-decoder, and a detection head.

3. The method according to claim 1, characterized in that In step 1, after obtaining the training image dataset of moving targets, the images in the training image dataset are annotated and different interference types are added to construct a SAR image dataset in a complex electromagnetic environment.

4. The method according to claim 1, wherein In step 2, The lightweight backbone network is a ResNet-50 network, and the input image is subjected to an initial convolutional layer and a maximum pooling operation to extract underlying features; The image features are sequentially passed through the four stages of the ResNet-50 network. Each stage contains multiple residual modules, and each module consists of multiple convolutional layers and residual connections. After obtaining the multi-scale feature map, the output two-dimensional feature map is flattened into a sequence form, and the feature vector at each position corresponds to a spatial position in the image; Constructing a two-dimensional position encoding matrix that matches the length and number of channels of the sequence to embed the spatial position information of each token, wherein the position encoding adopts a fixed position encoding constructed by sine and cosine functions or a trainable position encoding based on learnable parameters; The positional encoding is concatenated with the original feature vector to obtain an input sequence containing spatial perception capabilities, which is used as the input of the Transformer encoder. The position-encoded features are input into the scoring network, and some significant tokens are selected for refinement, i.e., encoder tokens are sparsely processed.

5. The method according to claim 1, characterized in that In step 3, the sparse query tokens selected by the scoring network are used as input, and each encoder layer performs the following operations in sequence: Through the deformable self-attention module, each token adaptively selects a local area as the focus range according to its position and samples contextual information from the original feature map; Perform feature conversion and nonlinear mapping on the updated token representation through a feedforward neural network; Repeatedly stack xN in multiple encoder layers e The i-th encoder layer updates the feature x by i-1 : Where DefAttn is deformable attention, LN is layer normalization, and FFN is feedforward network.

6. The method according to claim 1, characterized in that In step 4, the Transformer decoder receives the sparse tokens output by the encoder and a set of learnable object queries. Each layer of the decoder first models the dependencies between different queries through the self-attention module, and then performs a deformable cross-attention operation to align the queries with the tokens output by the encoder, guiding each query to focus on the potential target area.