SAR (Synthetic Aperture Radar) image vehicle target identification method based on self-attention guide diffusion model
Through the self-attention-guided diffusion model, noise interference is simulated and the target area weight is dynamically adjusted. Combined with the reverse diffusion and denoising process, the robustness problem of SAR image target recognition is solved, and the recognition accuracy and robustness are improved.
Patent Information
- Application Number
- CN202510208733.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-07-01
AI Technical Summary
In the prior art, under low signal-to-noise ratio conditions, SAR image object recognition methods are difficult to effectively extract robust features, resulting in a decrease in recognition accuracy, especially under complex backgrounds and noise interference.
The self-attention-guided diffusion model is adopted to simulate noise interference through the diffusion process, dynamically adjust the target area weights in combination with the self-attention mechanism, and restore the target information through the reverse diffusion and denoising process, and use a full-connection layer network for target recognition.
It significantly improves the target recognition accuracy and robustness under low signal-to-noise ratio conditions, enhances the noise processing capability, reduces calculation overhead, and is suitable for high real-time application scenarios.
Smart Images

Figure CN120236205A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of radar, and particularly relates to a method for identifying vehicle targets in SAR images based on a self-attention guided diffusion model. Background Art
[0002] Synthetic Aperture Radar (SAR) is an active microwave detection sensor. With the continuous maturity of imaging technology, the field of radar target perception based on SAR images has been continuously developed. SAR target recognition realizes the identification of target categories or models, and plays an important role in aspects such as target dynamic monitoring and environmental situation analysis. However, SAR images obtained in actual complex environments usually contain noise interference, resulting in changes in information such as target structural characteristics.
[0003] Traditional target recognition methods often rely on manually designed feature extraction techniques. These methods are limited by changes in the shape, pose, and environmental conditions of the target, and cannot accurately identify the target types in SAR images with low signal-to-noise ratio, resulting in limited recognition performance. Currently, with the rapid development of deep learning, neural networks have powerful feature mining and non-linear fitting capabilities, and target recognition methods based on neural networks have gradually been applied in SAR target recognition. Although these methods have powerful representation capabilities when dealing with high-dimensional data, traditional neural network methods are a data-driven approach, and are still prone to problems such as a decrease in recognition accuracy when facing complex backgrounds and noise interference. The model trained with high signal-to-noise ratio data has limited generalization ability for low signal-to-noise ratio test data, and the recognition performance is poor. Therefore, how to utilize the deep feature learning ability of neural networks and at the same time mine robust features related to target characteristics to perform robust target recognition on low signal-to-noise ratio SAR images is a technical problem that urgently needs to be solved currently. Summary of the Invention
[0004] In order to solve the problems existing in the prior art, the present invention provides a method for identifying vehicle targets in SAR images based on a self-attention guided diffusion model. This method combines the self-attention mechanism with the diffusion model, uses the diffusion process to learn noise-robust features, and the self-attention mechanism guides the diffusion model to learn features reflecting target structural characteristics, realizing robust feature extraction and target recognition of noisy SAR images, and improving the noise-robust ability of existing SAR target recognition methods.
[0005] In order to achieve the above object, the present invention is realized through the following technical solutions:
[0006] The present invention is a method for identifying vehicle targets in SAR images based on a self-attention guided diffusion model. This SAR target robust recognition method includes:
[0007] Step 1: Obtain the low signal-to-noise ratio SAR image to be recognized;
[0008] Step 2: Input the SAR image to be recognized into the trained SAR target recognition network based on the self-attention guided diffusion model to obtain the class prediction result of the image to be recognized.
[0009] Among them, the SAR target recognition network based on the self-attention guided diffusion model includes four parts: the diffusion process, self-attention mechanism guided learning, reverse diffusion and denoising, and target recognition. The diffusion process simulates the noise interference in the SAR image by gradually adding noise to the image and prepares the noise image for the subsequent denoising process. The self-attention mechanism guided learning process dynamically adjusts the weights of the target regions by calculating the correlations between regions in the image, so that the self-attention guided diffusion model can automatically focus on the target regions, suppress the irrelevant background, and achieve the effect of noise suppression. The reverse diffusion and denoising process gradually removes the noise added in the diffusion process and restores the target region information of the image. The target recognition process uses a fully connected layer network, outputs the class of the target through the softmax layer, inputs the denoised image into the target recognition network, performs target prediction, and obtains the recognition result.
[0010] A further improvement of the present invention is that: in the diffusion process, the input SAR image X0 is diffused, and noise is gradually added to obtain the noise image X t :
[0011]
[0012] Among them, α t is the noise coefficient of the diffusion process, and ε t represents the standard Gaussian noise.
[0013] A further improvement of the present invention is that: the self-attention mechanism guides the diffusion process, enabling the model to automatically focus on the target regions and reduce the influence of background noise. The self-attention mechanism dynamically adjusts the weights by calculating the correlations between regions of the input image and strengthens the representation of the target regions. The specific process is as follows:
[0014] Set the feature map obtained by passing the input image X0 through the convolutional layer which is mapped to the query vector key vector and value vector That is:
[0015] Q = W q X K = W k X V = W v X
[0016] Among them, W q, W k , W v respectively represent the convolutional kernel weights for generating queries, keys, and values. Then, the self-attention weight matrix is calculated
[0017]
[0018] where A ij represents the correlation between input images i and j. In this way, the self-attention mechanism module can dynamically calculate the relationships between different regions. Finally, the output Z after self-attention is obtained through weighted summation:
[0019] Z = AV
[0020] where Z represents the weighted feature map and serves as the input for the next step of reverse diffusion.
[0021] A further improvement of the present invention lies in that: the reverse diffusion and denoising processes use the denoising network D to restore a clear image
[0022]
[0023] where θ represents the parameters of the denoising network, and the denoising network D is a deep neural network composed of convolutional layers, deconvolutional layers, and self-attention mechanism modules.
[0024] A further improvement of the present invention lies in that: the structure of the denoising network includes:
[0025] Convolutional layer 1: The input image size is [C×H×W], the convolutional kernel size is set to 3×3, and the output channels are 64;
[0026] Convolutional layer 2: The input channels are 64, the convolutional kernel size is 3×3, and the output channels are 128;
[0027] Convolutional layer 3: The input channels are 128, the convolutional kernel size is 3×3, and the output channels are 256;
[0028] Self-attention mechanism module: This module weights the input feature map, calculates the correlation between each pixel, and readjusts the feature map;
[0029] Deconvolutional layer 1: The input channels are 256, the convolutional kernel size is 3×3, and the output channels are 128;
[0030] Deconvolutional layer 2: The input channels are 128, the convolutional kernel size is 3×3, and the output channels are 64;
[0031] Deconvolutional layer 3: The input channels are 64, the convolutional kernel size is 3×3, and the output channels are 3;
[0032] The output image is the restored target image
[0033] A further improvement of the present invention lies in that: the target recognition network includes a three-layer MLP network, which are successively connected by a first fully connected layer, a Relu activation function, a second fully connected layer, a Relu activation function, a third fully connected layer, and a softmax layer. The number of nodes set in the fully connected layers are: 1024, 256, and the number C of input data categories respectively. The output of the third fully connected layer is mapped to the corresponding label through softmax to obtain a predicted probability vector
[0034] A further improvement of the present invention lies in that: the loss function of the target recognition network is expressed as
[0035]
[0036] where C represents the total number of target categories, y c represents the true label, represents the predicted label, and L c represents the cross-entropy loss function between the true label and the predicted label of the input SAR image
[0037] A further improvement of the present invention lies in that: the training method of the SAR image vehicle target recognition method based on the self-attention guided diffusion model includes
[0038] S1. Set the time step t of the diffusion process, set the number of iterations to Q, set the learning rate ρ of the stochastic gradient algorithm, and set the batchsize size m
[0039] S2. Randomly initialize the parameters in the SAR target recognition network constructed based on the self-attention guided diffusion model using a normal distribution
[0040] S3. Select m SAR images from the training dataset to form a training sample group. According to the total number of samples M, a total of M / m training sample groups are obtained
[0041] S4. Calculate the loss function L of a training group sample in each iteration process using the loss function of the SAR target recognition network based on the self-attention guided diffusion model. Use the stochastic gradient descent algorithm to optimize the loss function L of each training sample group in turn m to update the parameters in the recognition model network, and complete the network parameter training of one iteration process. Among them, the loss function L m is expressed as m as
[0042]
[0043] where represents the i-th SAR image in the m-th training sample group, represents the SAR image output after the denoising network performs the diffusion process and reverse diffusion denoising. C represents the total number of target categories, represents the true label of the i-th SAR image in the m-th training sample group, represents the predicted label of the i-th SAR image in the training sample group. γ represents an adjustable weight parameter used to adjust the weight between the denoising network and the recognition network.
[0044] S5. Continuously repeat the process of S1 - S4 until the maximum number of iterations Q is reached, terminate the iteration, and obtain the trained SAR target recognition network based on the self-attention guided diffusion model.
[0045] S6. Input the low signal-to-noise ratio SAR image sample X * in the test dataset into the above-mentioned denoising network, and output the denoised SAR image
[0046] S7. Input the test image X to be predicted * into the denoising network combined with self-attention for feature extraction, and then input the extracted features into the above-mentioned target recognition network to obtain the category prediction label vector y of the test sample X * ; * ;
[0047] S8. Based on the category prediction label vector y * , determine the category to which the SAR image sample X * belongs to the dimension with the maximum probability value, and complete the category prediction of the SAR image sample X * to be predicted.
[0048] The beneficial effects of the present invention are as follows:
[0049] Through an end-to-end deep neural network framework, the present invention seamlessly integrates four core modules: the diffusion process, self-attention mechanism-guided learning, reverse diffusion and denoising, and target recognition, effectively avoiding the feature adaptation problem caused by the independence of each module in traditional methods, and improving the stability and accuracy of the recognition system.
[0050] The present invention simulates the noise interference in SAR images through the diffusion model process, learns noise-robust features, and combines the reverse diffusion model to gradually remove noise and restore the information of the target area, significantly enhancing the target recognition ability under low signal-to-noise ratio conditions.
[0051] The self-attention mechanism module of the present invention can dynamically adjust the weights of target regions, automatically focus on the target regions by calculating the correlations between image regions, effectively suppress the interference of background noise, and significantly improve the accuracy of feature extraction compared with traditional fixed-weight methods.
[0052] The present invention has significant advantages in terms of processing efficiency. Through a unified and optimized end-to-end design, it reduces the repeated calculations caused by module separation in traditional two-stage methods, not only improving the recognition accuracy but also significantly reducing the computational overhead, and is particularly suitable for application scenarios with high real-time requirements.
[0053] Compared with existing single denoising or enhancement techniques, the present invention combines the synergistic mechanism of diffusion and reverse diffusion, significantly enhancing the robustness of noise processing and providing a clear and reliable feature basis for target recognition.
[0054] This method is not only applicable to SAR image processing under different noise levels and complex backgrounds, but also has strong generalization ability and robustness, providing broad application prospects for fields such as national defense monitoring, remote sensing imaging, and environmental observation. Brief Description of the Drawings
[0055] Figure 1 is a schematic flow diagram of the present invention.
[0056] Figure 2 is the overall block diagram of the present invention.
[0057] Figure 3 is a schematic diagram of a diffusion process of the present invention.
[0058] Figure 4 is a schematic diagram of a self-attention mechanism module of the present invention;
[0059] Figure 5 is a schematic diagram of a reverse diffusion and reconstruction process of the present invention. Detailed Embodiments
[0060] The following will disclose the embodiments of the present invention with reference to the drawings. For the sake of clarity, many practical details will be described together in the following narrative. However, it should be understood that these practical details are not used to limit the present invention. That is to say, in some embodiments of the present invention, these practical details are not necessary.
[0061] As Figure 1 and Figure 2 , the present invention provides a method for identifying vehicle targets in SAR images based on a self-attention-guided diffusion model. The method for identifying vehicle targets in SAR images includes:
[0062] Step 1, obtain a low signal-to-noise ratio SAR image to be recognized;
[0063] Step 2: Input the SAR image to be recognized into the trained SAR target recognition network based on the self-attention guided diffusion model to obtain the class prediction result of the image to be recognized;
[0064] Among them, the SAR target recognition network based on the self-attention guided diffusion model includes four parts: the diffusion process, the self-attention mechanism guided learning, the reverse diffusion and denoising, and the target recognition. The diffusion process adds noise to the image step by step to simulate the noise interference in the SAR image and prepares the noise image for the subsequent denoising process. The self-attention mechanism guided learning process calculates the correlation between regions in the image and dynamically adjusts the weights of the target regions, enabling the model to automatically focus on the target regions, suppress the irrelevant background, and achieve the effect of noise suppression. The reverse diffusion and denoising process gradually removes the noise added in the diffusion process and restores the target region information of the image. The target recognition process uses a fully connected layer network, outputs the class of the target through the softmax layer, inputs the denoised image into the target recognition network for target prediction, and obtains the recognition result.
[0065] As Figure 3 shown, the diffusion process of the present invention gradually adds noise to the input SAR image to simulate the possible noise interference in the SAR image. The addition of noise uses standard Gaussian noise and completes the diffusion of the image within a certain time step to obtain the noise image X t . Specifically, for the input SAR image X0, its diffusion process can be described by the following formula:
[0066]
[0067] Among them, α t is the noise coefficient of the diffusion process, and ε t represents the standard Gaussian noise.
[0068] The diffusion process can help the model adapt to the low signal-to-noise ratio environment, improve the robustness to noise, and at the same time provide the image with noise for the subsequent denoising operation for the model to learn how to recover the target information from the noise during training. The diffusion process not only enhances the model's processing ability in the noise environment but also enables the model to better extract the target information in the face of high noise interference.
[0069] As Figure 4 shown, the self-attention mechanism is used to calculate the correlation between regions in the image and dynamically adjust the weights of each region according to the correlation. Through guided learning, the model can automatically focus on the target regions, suppress the interference of the irrelevant background, and thus enhance the representation ability of the target regions. The specific process is as follows:
[0070] Set the feature map obtained by the input image X0 passing through the convolutional layer It is mapped to a query vector through a convolutional layer key vector and value vector That is:
[0071] Q = W q X K = W k X V = W v X
[0072] where W q , W k , W v represent the convolutional kernel weights for generating the query, key, and value respectively. Then calculate the self-attention weight matrix
[0073]
[0074] where A ij represents the correlation between input images i and j. In this way, the self-attention mechanism module can dynamically calculate the relationships between different regions. Finally, the output Z after self-attention is obtained through weighted summation:
[0075] Z = AV
[0076] where Z represents the weighted feature map and serves as the input for the next step of reverse diffusion.
[0077] Under low signal-to-noise ratio conditions, background noise will affect the recognition of the target. The self-attention mechanism enhances the features of the target region and suppresses the interference of the background, thereby improving the model's focusing ability and recognition accuracy for the target. The self-attention mechanism can effectively increase the weight of the important region in the image, that is, the target region, so that the model can more accurately identify the target.
[0078] As Figure 5 shown, the purpose of the reverse diffusion and denoising process of the present invention is to gradually remove the noise added during the diffusion process and restore the clear image of the target region This process uses a denoising network D to restore the target region information of the image, and its diffusion process can be described by the following formula:
[0079]
[0080] where θ represents the parameters of the denoising network, and the denoising network D is a deep neural network composed of a convolutional layer, a transposed convolutional layer, and a self-attention mechanism module.
[0081] Specifically, the structure of the denoising network is as follows:
[0082] Convolutional layer 1: The input image size is [C×H×W], the convolutional kernel size is set to 3×3, and the output channels are 64;
[0083] Convolutional layer 2: The number of input channels is 64, the size of the convolutional kernel is 3×3, and the number of output channels is 128;
[0084] Convolutional layer 3: The number of input channels is 128, the size of the convolutional kernel is 3×3, and the number of output channels is 256;
[0085] Self-attention mechanism module: This module weights the input feature map, calculates the correlation between each pixel, and re-adjusts the feature map;
[0086] Transposed convolutional layer 1: The number of input channels is 256, the size of the convolutional kernel is 3×3, and the number of output channels is 128;
[0087] Transposed convolutional layer 2: The number of input channels is 128, the size of the convolutional kernel is 3×3, and the number of output channels is 64;
[0088] Transposed convolutional layer 3: The number of input channels is 64, the size of the convolutional kernel is 3×3, and the number of output channels is 3;
[0089] The output image is the restored target image
[0090] The process of reverse diffusion and denoising ensures that the noise is gradually removed, while being able to retain the target features in the original image, improving the accuracy of the final target recognition. The denoising effect is very crucial because only the denoised image can provide sufficiently clear target information for target recognition.
[0091] The target recognition network includes a three-layer MLP network, which are successively connected by the first fully connected layer, the Relu activation function, the second fully connected layer, the Relu activation function, the third fully connected layer, and the softmax layer. The number of nodes set in the fully connected layers are: 1024, 256, and the number C of input data categories. The output of the third fully connected layer is mapped to the corresponding label through softmax to obtain the predicted probability vector.
[0092] The loss function L of the target recognition module c is expressed as:
[0093]
[0094] where C represents the total number of target categories, y c represents the true label, represents the predicted label, and L c represents the cross-entropy loss function between the true label and the predicted label of the input SAR image.
[0095] These four modules interact with each other to jointly achieve the robust recognition of SAR images under low signal-to-noise ratio conditions. The diffusion process introduces noise to enhance robustness, the self-attention mechanism ensures the focusing on the target area, the reverse diffusion and denoising processes gradually recover the target information, and the target recognition completes the final classification task. Overall, this design that combines noise modeling, feature focusing, denoising, and classification improves the accuracy and robustness of SAR target recognition in complex environments.
[0096] The present invention also provides a training method for a method of recognizing vehicle targets in SAR images based on a self-attention guided diffusion model. Before training, it is first necessary to generate a training data set and a test data set.
[0097] Specifically, the training method for the robust recognition network of SAR targets based on the self-attention guided diffusion model includes:
[0098] S1. Set the time step t of the diffusion process, set the number of iterations as Q, set the learning rate ρ of the stochastic gradient algorithm, and set the batchsize as m;
[0099] S2. Randomly initialize the parameters in the SAR target recognition network constructed based on the self-attention guided diffusion model by using the normal distribution;
[0100] S3. Select m SAR images from the training data set to form a training sample group. According to the total number of samples M, a total of M / m training sample groups are obtained;
[0101] S4. Calculate the loss function L of a training group of samples in each iteration process by using the loss function of the SAR target recognition network based on the self-attention guided diffusion model m , and use the stochastic gradient descent algorithm to optimize the loss function L of each of the above-mentioned training sample groups m in turn to update the parameters in the recognition model network, and complete the network parameter training of one iteration process;
[0102] Specifically, select m SAR images from the training data set to form a minibatch, and calculate the loss function L of the network constructed by the samples in each minibatch m :
[0103]
[0104] Use the stochastic gradient descent algorithm to optimize the above-mentioned objective function, train the recognition network parameters, and complete the training process of one epoch.
[0105] S5. Continuously repeat the process of S1 - S4 until the maximum number of iterations Q is reached, terminate the iteration, and obtain a trained SAR target recognition network based on the self - attention - guided diffusion model.
[0106] After obtaining the trained SAR target recognition network based on the self - attention - guided diffusion model, test the trained network.
[0107] S6. Input the low - SNR SAR image sample X in the test dataset * into the denoising network, and output the denoised SAR image
[0108] S7. Input the test image X to be predicted * into the denoising network with self - attention for feature extraction, and then input the extracted features into the target recognition network to obtain the class prediction label vector y of the test sample X * ; * ;
[0109] S8. Based on the class prediction label vector y * , determine the class to which the SAR image sample X * belongs to the dimension with the maximum probability value, and complete the class prediction of the SAR image sample X to be predicted * .
[0110] The above is only the implementation manner of the present invention and is not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.
Claims
1. A vehicle target recognition method for SAR images based on a self-attention guided diffusion model, characterized by: The SAR image vehicle target recognition method specifically comprises the following steps: Step 1: Obtain a low signal-to-noise ratio SAR image to be identified; Step 2: input the low signal-to-noise ratio SAR image to be identified into a trained SAR target recognition network based on a self-attention guided diffusion model to obtain a category prediction result of the image to be identified; The SAR target recognition network based on the self-attention guided diffusion model includes four parts: diffusion process, self-attention mechanism guided learning, back diffusion and denoising, and target recognition. The diffusion process simulates the noise interference in the low signal-to-noise ratio SAR image by gradually adding noise to the low signal-to-noise ratio SAR image, and prepares the noise image for the subsequent denoising process. The self-attention mechanism guided learning process dynamically adjusts the weight of the target area by calculating the correlation between the regions in the low signal-to-noise ratio SAR image, so that the self-attention guided diffusion model can automatically focus on the target area, suppress irrelevant background, and achieve the effect of noise suppression. The back diffusion and denoising process gradually removes the noise added in the diffusion process and restores the target area information of the low signal-to-noise ratio SAR image. The target recognition process adopts a fully connected layer network, outputs the target category through the softmax layer, and inputs the denoised image into the SAR target recognition network based on the self-attention guided diffusion model to perform target prediction and obtain a recognition result.
2. The SAR image vehicle target recognition method based on the self-attention guided diffusion model according to claim 1 is characterized in that: The diffusion process diffuses the input low signal-to-noise ratio SAR image X0 and gradually adds noise to obtain a noisy image X t : Among them, α t is the noise coefficient of the diffusion process, ε t represents standard Gaussian noise.
3. The SAR image vehicle target recognition method based on the self-attention guided diffusion model according to claim 1 is characterized in that: The self-attention mechanism guides the learning process by calculating the correlation between regions in the low signal-to-noise ratio SAR image and dynamically adjusting the weight of the target region, so that the self-attention guided diffusion model can automatically focus on the target region and suppress irrelevant background to achieve the effect of noise suppression. Specifically, the process includes the following: First, set the input low signal-to-noise ratio SAR image X0 to pass through the convolution layer to obtain the feature map Feature Map Mapped to query vector through convolutional layer Key Vector Sum value vector Right now: Q=W q X K=W k X V=W v X Among them, W q ,W k ,W v Represents the convolution kernel weights for generating queries, keys, and values respectively; Secondly, calculate the self-attention weight matrix Among them, A ij Represents the correlation between input images i and j. In this way, the self-attention mechanism module can dynamically calculate the relationship between different regions; Finally, the output feature map Z after self-attention is obtained by weighted summation: Z=AV Wherein, Z represents the weighted feature map, which serves as the input of the reverse diffusion and denoising process.
4. The SAR image vehicle target recognition method based on the self-attention guided diffusion model according to claim 1 is characterized in that: The reverse diffusion and denoising process gradually removes the noise added in the diffusion process and restores the target area information of the low signal-to-noise ratio SAR image. Specifically, the reverse diffusion and denoising process uses the denoising network D to restore a clear image. Among them, θ represents the parameters of the denoising network.
5. The SAR image vehicle target recognition method based on the self-attention guided diffusion model according to claim 4 is characterized in that: The denoising network D is a deep neural network composed of a convolutional layer, a deconvolutional layer, and a self-attention mechanism module. The specific structure includes: Convolutional layer 1: The input image size is [C×H×W], the convolution kernel size is set to 3×3, and the output channel is 64; Convolutional layer 2: The input channel is 64, the convolution kernel size is 3×3, and the output channel is 128; Convolutional layer 3: input channels are 128, convolution kernel size is 3×3, and output channels are 256; Self-attention mechanism module: The self-attention mechanism module weights the input feature map, calculates the correlation between each pixel, and rescales the feature map; Deconvolution layer 1: input channel is 256, convolution kernel size is 3×3, and output channel is 128; Deconvolution layer 2: input channel is 128, convolution kernel size is 3×3, and output channel is 64; Deconvolution layer 3: input channel is 64, convolution kernel size is 3×3, and output channel is 3; The output image is the restored target image 6. The SAR image vehicle target recognition method based on the self-attention guided diffusion model according to claim 1 is characterized by: The target recognition network includes a three-layer MLP network, which is connected in sequence to a first fully connected layer, a Relu activation function, a second fully connected layer, a Relu activation function, a third fully connected layer and a softmax layer, wherein the number of nodes set in the fully connected layer is 1024, 256 and the number of input data categories C, respectively. The output of the third fully connected layer is mapped to the corresponding label through softmax to obtain a predicted probability vector.
7. The SAR image vehicle target recognition method based on the self-attention guided diffusion model according to claim 1 is characterized by: The loss function of the target recognition network is expressed as: Among them, C represents the total number of target categories, y c represents the true label, represents the predicted label, L c Represents the cross entropy loss function between the true label and the predicted label of the input low SNR SAR image.
8. The SAR image vehicle target recognition method based on the self-attention guided diffusion model according to claim 7 is characterized by: In step 2, the training of the SAR target recognition network based on the self-attention guided diffusion model specifically includes the following steps: Step 2.1, set the time step t of the diffusion process, set the number of iterations to Q, set the learning rate ρ of the stochastic gradient algorithm, and set the batch size m; Step 2.2, using normal distribution to randomly initialize the parameters of the constructed SAR target recognition network based on the self-attention guided diffusion model; Step 2.3, select m SAR images from the training data set to form a training sample group, and according to the total number of samples M, a total of M / m training sample groups are obtained; Step 2.4: Using the loss function L of the SAR target recognition network based on the self-attention guided diffusion model c Calculate the loss function L of a training set sample in each iteration m , using the stochastic gradient descent algorithm to calculate the loss function L of each training sample group in turn m Optimize to update the parameters in the SAR target recognition network based on the self-attention guided diffusion model and complete the network parameter training in an iterative process; Step 2.4: Repeat steps 2.1 to 2.4 until the maximum number of iterations Q is reached, terminate the iteration, and obtain a trained SAR target recognition network based on the self-attention guided diffusion model.
9. The SAR image vehicle target recognition method based on the self-attention guided diffusion model according to claim 8 is characterized in that: In step 2.4, the loss function L m It is expressed as: in, represents the i-th SAR image in the m-th training sample group, It means that after the denoising network The SAR image output after diffusion process and reverse diffusion denoising, C represents the total number of target categories, represents the true label of the i-th SAR image in the m-th training sample group, represents the predicted label of the i-th SAR image in the training sample group, and γ represents an adjustable weight parameter, which is used to adjust the weight between the denoising network and the recognition network.
10. The SAR image vehicle target recognition method based on the self-attention guided diffusion model according to claim 1 is characterized by: In step 2, the denoised image is input into the SAR target recognition network based on the self-attention guided diffusion model to perform target prediction and obtain the recognition result, which specifically includes the following steps: First, the low signal-to-noise ratio SAR image sample X in the test dataset is * Input to the denoising network and output the denoised SAR image Secondly, the test image X to be predicted * The feature extraction is then input into the SAR target recognition network based on the self-attention guided diffusion model to obtain the test sample X. * The class prediction label vector y * ; Finally, the label vector y is predicted based on the category * , the low signal-to-noise ratio SAR image sample X * The category to which the dimension with the maximum probability value belongs is determined, and the SAR image sample X to be predicted is completed. * Category prediction.