Domain generalization heterogenous SAR image target identification method based on feature separation and diffusion model
Through the domain generalization method of feature separation and diffusion model, the problem of target recognition in heterologous SAR images is solved, and the accurate recognition and generalization performance improvement in images of different source sources is achieved.
Patent Information
- Application Number
- CN202510438006.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-08-01
AI Technical Summary
The prior art is difficult to effectively identify targets in heterologous SAR images, mainly due to the heterologous data distribution caused by different imaging platforms and times, making it difficult for the model to accurately extract target features.
The domain generalization method based on feature separation and diffusion model is adopted, image features are extracted through the backbone network and domain-specific and domain-invariant features are separated. The diffusion model is used for denoising optimization, and target recognition is combined with a classifier.
It improves the accuracy of heterologous SAR image target recognition and the generalization ability of the model, and can accurately identify targets in SAR images from different sources, improving recognition accuracy and robustness.
Smart Images

Figure CN120411802A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of aperture radar image target recognition, in particular to a domain generalization heterogeneous SAR image target recognition method based on feature separation and diffusion model. Background Art
[0002] Automatic and rapid target recognition in Synthetic Aperture Radar (SAR) imagery is crucial for improving the efficiency and value of remote sensing applications and is a current frontier in SAR image interpretation. Compared to traditional methods, deep learning models are more capable of recognizing high-level semantic information in images and can be effectively used in downstream classification tasks.
[0003] However, efficient performance relies on extensive data collection and distribution consistency across the training and testing domains. Once the datasets come from different devices or have varying imaging quality (heterogeneity), the final results are significantly degraded. Due to the low resolution of SAR images, the pixel information of ground targets is highly diluted, making it sometimes difficult to manually determine their category. This results in fewer reliable samples available for training, making it easy for the model to overfit. Furthermore, low-resolution SAR images contain less semantic information, which greatly increases the difficulty of feature decoupling.
[0004] Heterogeneous SAR images primarily originate from various imaging platforms, such as satellite-based, airborne, and ground-based SAR systems. Due to differences in radar equipment parameters, orbits, flight altitudes, and speeds carried by these platforms, the resulting synthetic aperture radar images exhibit significant variations in imaging resolution, viewing angle, and target scattering characteristics. Furthermore, varying imaging times, including daytime and seasonal variations, can significantly alter the electromagnetic scattering characteristics of the same target in SAR images. The combined effect of these factors results in a highly heterogeneous distribution of SAR image data from different sources. When faced with heterogeneous SAR images that differ significantly from the training data, traditional models struggle to accurately adapt and extract effective target features due to the significant variations in target geometry, texture, and scattering intensity across these image sources.
[0005] Therefore, a domain-generalized target recognition method for heterogeneous SAR images is urgently needed. Summary of the Invention
[0006] The present invention aims to provide a domain generalized heterogeneous SAR image target recognition method based on feature separation and diffusion model to overcome the defects of the prior art.
[0007] The technical solution adopted by the present invention to achieve the above-mentioned purpose is: a domain generalization heterogeneous SAR image target recognition method based on feature separation and diffusion model, comprising the following steps:
[0008] Construct an object recognition model composed of a backbone network and a classifier;
[0009] Input SAR images in multiple source domain datasets into the backbone network in sequence. Extract all-element features through the backbone network, including image features in multiple source domains; and separate the image features into domain-specific features and domain-invariant features; remove the domain-specific features, thus leaving the domain-invariant features for subsequent classification; denoise the additive features through a diffusion model, thereby optimizing the backbone network in the form of gradient backpropagation; the classifier further classifies the obtained domain-invariant features to output the object category, completing the training of the object recognition model;
[0010] Obtain SAR images in real time and get the object category through the trained object recognition model.
[0011] The backbone network adopts a classification loss function, and the construction of the classification loss function includes the following steps:
[0012] In the backbone network, m domain-specific classifiers are trained by minimizing the classification loss function and maximizing the uncertainty loss, which are respectively
[0013]
[0014] where, respectively represent the domain-specific classifier classification loss function and the uncertainty loss function, is a mathematical expectation calculator, θ ,
[0018] , , , ,
[0016] , , , , , , , , , i ,
[0017] ,
[0014] , , , ,
[0015] is parameters of, x is the data, used to represent the input SAR image, y is the target class label, k and i are both the ordinals of the source domains, and k≠i, referring to different source domains, j represents the ordinal of the data-label pair in a certain source domain, respectively represent the i-th source domain and the k-th source domain, uses cross-entropy classification loss, is the uncertainty loss, and its form is as follows:
[0015]
[0016] where, C is the total number of target categories, p is the probability; the domain-specific classifier first predicts the image, and then obtains the probabilities of all categories; the category with the lowest probability is called the most unlikely category, and the image will be labeled as this category; then, train the domain-specific classifier with the uncertainty loss to predict the most unlikely category:
[0017]
[0018] where, Represents the cross - entropy classification loss; after training is completed, fix the parameters θ of the domain - specific classifier i , and use the domain - specific classifier to learn domain - independent features;
[0019] Use the encoder - decoder network within the backbone network to map the image to a new feature space, and the output is input to the domain - specific classifier and the new domain - invariant classifier; the source domain is used to maximize the uncertainty loss, as follows:
[0020]
[0021] where, represents the source - domain uncertainty loss function, E is the encoder - decoder network within the backbone network, θ e is the parameter of E;
[0022] Add the reconstruction loss L in the encoder - decoder network of the backbone network r :
[0023]
[0024] where, represents the encoder - decoder network loss function, is the pixel - level L - 2 norm reconstruction loss function;
[0025] Train the domain - invariant classifier by minimizing the following classification loss for all source - domain output images:
[0026]
[0027] where, represents the domain - invariant classifier loss function, θ f is the parameter of the domain - invariant classifier F di ;
[0028] The classification loss function is as follows
[0029]
[0030] where, λ1, λ2, and λ3 are trade - off hyperparameters that control the weights of the corresponding losses.
[0031] For the training of the diffusion model and the classifier, the construction of the loss function is as follows:
[0032]
[0033] where, represents the diffusion model loss function, is the additive noise feature, ∈ t is the noise added to the additive noise feature, sampled from standard Gaussian noise; t is the diffusion step number, used to characterize the magnitude of the added noise, T is the total number of diffusion steps, f0 is a parameter of the diffusion model; φ is a parameter of the diffusion model, ∈ φ represents the noise of the parameter φ, which is a function of the additive noise feature and the diffusion step number t;
[0034]
[0035] where, represents the classifier loss function, is the cross-entropy loss function, α is a parameter of the classifier, G α is the classifier.
[0036] After the diffusion model is trained, multi-step denoising is performed on the noise to obtain the enhanced feature
[0037]
[0038] where, n is the sampled noise, is the filtered noise, α t and σ t are the set hyperparameters, corresponding to the amount of noise removed in the denoising process, t is the diffusion step number, used to characterize the magnitude of the added noise;
[0039] The filtered noise is adjusted and modified through the following formula:
[0040]
[0041] where, is the denoising noise adjusted according to the required generation category, is the adjustment direction obtained by the classifier according to the required category c, v is the adjustment value, p α represents the probability of the parameter α;
[0042] For the corresponding class label c of the denoised enhanced sample, the diffusion model is fine-tuned using the cross-entropy objective function:
[0043]
[0044] where, represents the denoising loss function, is the cross-entropy loss function, f0 is a parameter of the diffusion model, is the feature of the diffusion model, G ω represents the classifier; ω represents the parameter of the classifier;
[0045] The loss function is as follows:
[0046]
[0047] Among them, μ1, μ2, and μ3 are trade-off hyperparameters.
[0048] A domain generalization heterologous SAR image target recognition system based on feature separation and diffusion model, comprising:
[0049] A target recognition model construction module, configured to construct a target recognition model composed of a backbone network and a classifier;
[0050] A model training module, configured to sequentially input SAR images in multiple source domain datasets into the backbone network, extract full-element features through the backbone network, including image features in multiple source domains; separate the image features into domain-specific features and domain-invariant features; remove the domain-specific features, thereby leaving the domain-invariant features for subsequent classification; denoise the additive features through a diffusion model, thereby optimizing the backbone network in the form of gradient backpropagation; the classifier further classifies the obtained domain-invariant features to output the target category, and completes the training of the target recognition model;
[0051] A target recognition module, configured to obtain SAR images in real time and obtain the target category through the trained target recognition model.
[0052] A domain generalization heterologous SAR image target recognition device based on feature separation and diffusion model, comprising a memory and a processor; the memory is used for storing a computer program; the processor is used for, when executing the computer program, implementing the domain generalization heterologous SAR image target recognition method as described above.
[0053] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the domain generalization heterologous SAR image target recognition method as described above is implemented.
[0054] The present invention has the following beneficial effects and advantages:
[0055] 1. The present invention can enable the target recognition model to mine generalized target feature representations from limited training data, overcome the differences in heterologous data, and accurately recognize targets in various SAR images from different sources.
[0056] 2. Through the joint training of the diffusion model and the classifier, and through the joint optimization of constructing the loss function, after the diffusion model is trained, multi-step denoising is performed on the noise, and according to the corresponding class labels, the loss function is constructed. This makes the data generated by the diffusion model more conducive to the training of the classifier, forming a virtuous cycle and improving the class accuracy of the generated samples; promoting the collaborative optimization of the generated samples and the classification task, improving the generalization ability and robustness of the model, and thus improving the target recognition accuracy.
[0057] 2. The present invention combines the "plug-in" diffusion model with the proposed domain generalization method to achieve rapid feature broadening. The diffusion model is used to learn the representation method to generate diverse new samples for feature generalization, thereby enriching the data distribution of the source domain. In addition, the classifier is also used to guide the diffusion model to generate samples with correct semantics, thereby improving the generalization performance of the model in the target domain. Brief Description of the Drawings
[0058] Figure 1 The domain generalization SAR image target recognition network structure diagram of the present invention;
[0059] Figure 2 Four types of target images in three heterologous image datasets of the present invention;
[0060] Figure 3 The feature visualization diagram of the present invention for identifying targets in the target domain using t-SNE technology. Detailed Embodiment
[0061] The following further elaborates on the present invention in conjunction with the drawings and embodiments.
[0062] Regarding the heterologous problem in SAR image recognition, the present invention improves the performance through a domain generalization diffusion model. Therefore, relevant work reviews are respectively carried out on two parts: SAR image target recognition by the diffusion model and domain transfer SAR image target recognition.
[0063] A. SAR Image Target Recognition Based on Diffusion Model
[0064] Existing technologies attempt to directly apply the diffusion model to target feature extraction and recognition in SAR images. By learning a large number of target samples from SAR images, the diffusion model can capture specific feature patterns of the targets, thereby achieving accurate classification of the targets.
[0065] B. Domain Transfer SAR Image Target Recognition
[0066] Domain adaptation and domain generalization technologies aim to solve the heterologous problem and improve the SAR target recognition ability of the model in different domains.
[0067] In summary, there are few research results on using domain generalization strategies to solve the problem of cross-source SAR image target recognition. We are considering a more unique strategy to improve the generalization performance by introducing diffusion models.
[0068] In the practical application of synthetic aperture radar (SAR) image target recognition, it is often encountered that the training data and test data come from different data domains. This difference in data distribution will lead to a decrease in the accuracy rate during testing. When we define these sample sets with different data distributions as domains, the domain generalization strategy can well solve the above problems when the target domain is invisible. This paper develops a novel domain generalization framework: First, learn domain-invariant features by actively removing domain-specific features from the input images; Second, introduce an implicit diffusion model based on a classifier in the feature space. After training in the feature space, this model can conditionally generate accurate and rich source domain features to achieve fine control and diverse generation of features, thereby improving the model's generalization ability for the target domain. Experimental verification shows that the accuracy of the proposed method exceeds that of existing advanced methods.
[0069] In the domain generalization problem, there are N visible labeled source domains for training, each with a different distribution of data and label pairs. Next, the trained model will be tested on a significantly different and unseen target domain in which the data distribution is inconsistent with all source domains, that is The main goal of domain generalization is to train a domain-independent feature extractor and classifier to generalize the model trained in the source domain to the unseen target domain. To achieve data augmentation, an additional diffusion model needs to be trained with the classifier to guide the generation of diffusion model conditions. The algorithm framework of the present invention is as Figure 1 shown and includes the following steps:
[0070] SAR images in multiple source domain datasets are sequentially input into the backbone network;
[0071] First, extract full-element features through the backbone network, including image features in multiple source domains;
[0072] Separate the image features into domain-specific features and domain-invariant features through the domain-specific classifier and related encoder network within the backbone network;
[0073] Remove the relevant domain-specific features, thus leaving the domain-invariant features for subsequent classification;
[0074] The diffusion model can accurately denoise the additive features, thereby optimizing the above-mentioned domain generalization performance of the backbone network in the form of gradient backpropagation;
[0075] The classifier further classifies the domain-invariant features obtained by iteration into corresponding target categories.
[0076] Through the above steps, the training of the target recognition network composed of the backbone network and the classifier is completed;
[0077] Input the real-time obtained SAR image into the trained target recognition network to obtain the target category.
[0078] Domain-invariant features for generalization
[0079] Based on the above analysis, it is necessary to focus on mining the domain-specific features and domain-invariant features of heterologous SAR images, and then focus on retaining the domain-invariant features to ensure the generalization performance of the model.
[0080] Specifically, it is necessary to train m domain-specific classifiers Among them, the classifier only uses the domain-specific features of the source domain to distinguish images. The domain-specific classifier F ds should not use domain-invariant features as input. In other words, we hope to be able to effectively distinguish images of different categories in the source domain, but should be difficult to distinguish images of different categories in any other domain. The domains excluding the source domain are used to maximize the uncertainty of classification, or make it more difficult to classify. In the domains excluding the source domain, the classification performance should be similar to random prediction.
[0081] Specifically, the cross-entropy loss objective function is used to minimize the classification loss in the design of the classifier based on the cross-entropy loss objective function. Specifically, the classifier is trained by minimizing the classification loss function and maximizing the uncertainty loss, which are respectively
[0082]
[0083] where is the mathematical expectation calculator, θ i is the parameter of, x is the data, y is the label, k and i are both the ordinals of the source domain, and k≠i, referring to different source domains, j represents the ordinal of the data-label pair in a certain source domain, using the cross-entropy classification loss, is the uncertainty loss, and its form is as follows
[0084]
[0085] Where C is the total number of classes and p is the probability. The minimum possible loss is an alternative to the entropy loss. The classifier first predicts the image and then obtains the probabilities for all classes. The class with the lowest probability is called the least likely class. The image is then labeled with this class. Then, we train the classifier to predict the least likely class in the following form
[0086]
[0087] where is the estimate of the j-th label in the k-th source domain. After training, we fix the parameters θ of these domain-specific classifiers i and use these classifiers to learn domain-invariant features.
[0088] To retain the domain-invariant features learned by the domain-specific classifiers, an encoder-decoder network within the backbone network is used to map the image to a new feature space, and the output of the encoder-decoder network is input to the domain-specific classifier and the new domain-invariant classifier. The source domain in this step is used to maximize the uncertainty loss, as follows
[0089]
[0090] where E is the encoder-decoder network and θ e are the parameters of E. Freeze the parameters θ i and train the parameters θ e of the encoder-decoder network E. Maximizing the uncertainty loss forces the decoder to output a feature image e i = E(x i ) that contains fewer domain-specific features than the input image, where i represents the ordinal number of the data (here, the image). In this way, the encoder-decoder network can remove the domain-specific features from the input image x and retain the domain-independent features in the output image e. To maintain the overall similarity between the input and output images, a loss function the reconstruction loss L r
[0091]
[0092] where is the pixel-level L-2 norm reconstruction loss function. In addition, the domain-invariant classifier is trained by minimizing the classification loss
[0093]
[0094] where θ f is the parameter of the domain-invariant classifier F di The classification loss formula (7) also updates the encoder-decoder network to prevent the encoder-decoder network from losing domain-invariant features due to the uncertainty loss. If it is difficult to distinguish domain-specific features from domain-invariant features, the uncertainty loss may also remove domain-invariant features.
[0095] Therefore, in the design of separating domain-invariant features and domain-specific features and achieving generalization performance, the complete training objective is written in the following form
[0096]
[0097] where λ1, λ2, and λ3 are trade-off hyperparameters that control the weights of the corresponding losses.
[0098] Diffusion model based on generalized conditioning
[0099] Specifically, the features extracted in the previous section are used to train the diffusion model D φ , and an additional classifier G α is trained, which can guide the direction of the noise features. To accurately denoise the additive features, noise is added to the loss function of the diffusion model to match the filtered noise with the L-2 norm loss. The final loss function is as follows
[0100]
[0101] where represents the loss function of the diffusion model, is the additive noise feature, ε t is the noise added to the additive noise feature, sampled from standard Gaussian noise. t is the diffusion step, representing different magnitudes of added noise, T is the total number of diffusion steps, and f0 is the parameter of the diffusion model. φ is the parameter of the diffusion model, and ∈ φ represents the noise of the parameter φ, which is a function of the additive noise feature and the diffusion step t.
[0102] In addition, an additional classifier is trained to guide the generation of the diffusion model. This classifier guides the denoising direction by accurately classifying the additive features during the denoising process, that is, training with the cross-entropy loss
[0103]
[0104] where the cross-entropy loss function is used.
[0105] In addition, after obtaining a trained diffusion model, multi-step denoising of the noise can obtain enhanced features
[0106]
[0107] where n is the sampling noise, is the filtered noise, and α t and σ t are hyperparameters set according to experience and correspond to the amount of noise removed during the denoising process. Gaussian white noise is denoised through multiple steps to obtain enhanced features
[0108] To ensure that the enhanced features are still semantically reasonable, the denoising process is supervised using pre-trained classifiers that control the categories generated by the diffusion model, i.e., adjusting and modifying the filtered noise. Further cropping is performed to filter the noise, and supervised learning is carried out, finally forming the following form
[0109]
[0110] where is the denoised noise adjusted according to the desired generated category, is the adjustment direction obtained by the classifier according to the desired category c, υ is the adjustment value, and p α represents the probability of the parameter α.
[0111] The diffusion model guided by the classifier is used to generate enhanced samples of the desired category. Combining the corresponding category labels c of the enhanced samples, the model is further fine-tuned using a simple cross-entropy objective function to obtain better generalization results
[0112]
[0113] where, represents the denoising loss function, is the cross-entropy loss function, f0 are the parameters of the diffusion model, are the features of the diffusion model, and G ω represents the classifier.
[0114] Combining the above design steps, the overall objective of training can be written in the following form
[0115]
[0116] where μ1, μ2, and μ3 are trade-off hyperparameters.
[0117] Experimental Results and Analysis
[0118] A. Dataset and Experimental Settings
[0119] To verify the effectiveness of the proposed method, two heterologous SAR ship image datasets, OpenSAR and FUSAR, and one optical remote sensing dataset, FGSC, were used. All these images in the two datasets were from different platforms and regions, with different resolutions and acquisition dates, and were collected by different types of vessels and imaging preprocessing methods. In the two SAR datasets, due to different pre-classification criteria, the same category might have different names, and categories with the same name might contain different types. OpenSAR was captured by the ESA Sentinel-1 satellite, which is a C-band radar with VV and VH polarizations. FUSAR was imaged by China's Gaofen-3 satellite, which is a C-band radar with VV and HH polarizations. Sentinel-1 and Gaofen-3 satellites have many differences in detailed technical parameters, so OpenSAR and FUSAR are two SAR image datasets with different joint probability distributions. FGSC is a remote sensing image dataset mainly obtained from Google Earth and Gaofen-2 satellite, and four ship categories overlapping with the above two datasets were selected, namely cargo ships, fishing boats, tankers, and others. The four ship categories corresponding to the two datasets are shown in Figure 2 It can be seen that the ship images corresponding to these categories are very different.
[0120] In the encoder-decoder part of the network and the diffusion model, U-Net was used as the backbone architecture, while ResNet-50 was used as the backbone architecture in the domain generalization setting. Stochastic gradient descent with a momentum of 0.9 was selected as the optimization method for all training processes. The batch size was set to 16, and the training epochs were set to 500. The initial learning rate of the network was 0.001, while the initial learning rate of the diffusion model was 10 -6 . According to experience, the values of parameters λ1, λ2, and λ3 were set to 1, 0.5, and 0.3 respectively. At the same time, the values of parameters μ1, μ2, and μ3 were set to 0.5, 0.2, and 0.1 respectively.
[0121] Without loss of generality, for each dataset, the proposed experimental scheme used two of the domains as the source domain and the remaining one as the target domain. The source dataset was divided into a training set and a validation set.
[0122] B. Analysis of Comparative Experiment Results
[0123] To demonstrate the superiority of the proposed method in the task of heterologous SAR image target recognition, a comparative experiment was first conducted by comparing it with the existing state-of-the-art (SOTA) methods. The following five SOTA methods were selected as the comparison methods, namely soft segmentation randomization (SSR), two-level domain alignment (TDA), decoupled domain-invariant features (DDF), multi-source domain generalization (MDG), and contrast-enhanced domain generalization (CDG). In Table 1, the recognition performances of the three domain generalization tasks are shown respectively.
[0124] Table 1 Heterogeneous SAR Image Target Recognition Performance Table for Three Domain Generalization Tasks (%)
[0125]
[0126] It can be intuitively seen from the quantitative results that the proposed method has the most superior accuracy. In the three domain generalization tasks, the average recognition accuracy of the proposed method is at least 4% higher than that of other SOTA methods. This proves that the method of feature decoupling and removing domain-specific features while retaining domain-general features, as well as the strategy of feature enhancement through the diffusion model, is absolutely effective for the conversion between remote sensing images from different sources. In fact, some other problems can also be reflected from the experimental results. When the target domain is an optical remote sensing image (the source domains are two SAR image datasets), the accuracy of target recognition is lower than the other two cases (the source domains are one SAR image dataset and one optical remote sensing dataset). It is also well known that it is more difficult to extract features from SAR images in nature. Therefore, when SAR images do not participate in training as the target domain, the generalization fitting effect of the model on them is poor. At the same time, when using OpenSAR as the target domain, the target recognition accuracy is lower than when using FUSAR as the target domain, indicating that the data in OpenSAR is more complex and has more interference.
[0127] C. Ablation Study and Related Analysis
[0128] 1) Diffusion Model: An important idea of the domain generalization strategy in this paper is to introduce a diffusion model to achieve feature enhancement and generalization, thereby improving the generalization performance and further improving the target recognition accuracy. Therefore, this example designs an ablation experiment to verify the target recognition accuracy with or without the diffusion model. Taking the FGSC in Table 2 as the target domain, it can also be clearly seen that it is the diffusion model that enables the strategy proposed in this paper to increase the target recognition accuracy by 5.6%.
[0129] Table 2 Target Recognition Results with / without Diffusion Model (%)
[0130]
[0131] 2) Backbone network: In the above experiments, ResNet-50 was used as the backbone network for object recognition, and ResNet was also used as the backbone network for other parameters. Therefore, in this ablation experiment, two backbone networks, ResNet-34 and ResNet-101, were selected for comparison to verify the experimental device used. As can be clearly seen from Table 3 (taking the target domain as OpenSAR as an example), when ResNet-34 was used as the backbone network, the loss was far from converging, resulting in poor accuracy performance. When ResNet-101 was used as the backbone network, the object recognition accuracy of the domain generalization task was not much different from that when ResNet-50 was used as the backbone network. Since the structure of ResNet-101 is relatively complex and the complexity of model operation will also increase, ResNet-50 was selected as the backbone network in this paper.
[0132] Table 3 Object recognition results of different backbone networks (%)
[0133]
[0134] 3) Feature visualization: To prove that the method in this paper has better intra-class and inter-class relationships, the class spacing is larger while the samples are closer. Taking the third task as an example, Figure 3 (a) - (f) in it show the results of visualizing the target domain features through t-distributed stochastic neighbor embedding (t-SNE). It can be seen that the proposed method can not only make up for the above points that need to be verified, but also has a certain feature dispersion, that is, the enhanced features are distributed in a wider feature space, enabling the model to learn more feature information, thereby improving the classification accuracy and generalization ability of the model, and further improving the accuracy of object recognition.
[0135] To solve the problems of model generalization and migration ability brought by heterogeneous SAR image data, a novel domain generalization framework was developed. First, domain-invariant features were learned by actively removing domain-specific features from the input images. In addition, considering that the high-quality and diverse semantic features generated by the diffusion model can train the model, a plug-and-play feature enhancement diffusion model was designed and combined with the domain generalization method.
[0136] Extensive experiments show that compared with all existing domain generalization methods, the framework of the present invention achieves more superior recognition performance.
Claims
1. A method for target recognition of heterogeneous SAR images based on feature separation and diffusion models, characterized in that Including the following steps: Construct a target recognition model consisting of a backbone network and a classifier; Input SAR images in multiple source domain datasets into the backbone network in sequence, and extract all-element features through the backbone network, including image features in multiple source domains; And separate the image features into domain-specific features and domain-invariant features; Remove the domain-specific features, thus leaving the domain-invariant features for subsequent classification; denoise the additive features through a diffusion model, thus optimizing the backbone network in the form of gradient backpropagation; the classifier further classifies the obtained domain-invariant features to output the target category, and complete the training of the target recognition model; Obtain SAR images in real time, and obtain the target category through the trained target recognition model.
2. The domain generalization heterologous SAR image target recognition method based on the feature separation and diffusion model according to claim 1, characterized in that, The backbone network adopts a classification loss function, and the construction of the classification loss function includes the following steps: In the backbone network, m domain-specific classifiers i = 1,... m are trained by minimizing a classification loss function and maximizing an uncertainty loss, respectively, as Among them, respectively represent the domain-specific classifier classification loss function and the uncertainty loss function, is the mathematical expectation calculator, θ i is 's parameter, x is the data, used to represent the input SAR image, y is the class label of the target, k and i are both the ordinals of the source domain, and k≠i, referring to different source domains, j represents the ordinal of the data-label pair in a certain source domain, S i 、S k respectively represent the i-th source domain and the k-th source domain, uses the cross-entropy classification loss, is the uncertainty loss, and its form is as follows: Among them, C is the total number of categories of the target, and p is the probability; the domain-specific classifier first predicts the image and then obtains the probabilities of all categories; the category with the lowest probability is called the most unlikely category, and the image will be labeled as this category; then, with the uncertainty loss train the domain-specific classifier to predict the most unlikely category: Among them, represents the cross-entropy classification loss; after training is completed, the parameters θ of the domain-specific classifier are fixed i , and the domain-specific classifier is used to learn domain-independent features; The encoder-decoder network within the backbone network is used to map the image to a new feature space, and the output is input to the domain-specific classifier and the new domain-invariant classifier; the source domain is used to maximize the uncertainty loss, as shown below: Among them, represents the source domain uncertainty loss function, E is the encoder-decoder network within the backbone network, and θ e is the parameter of E; Add the reconstruction loss L to the encoder-decoder network of the backbone network r : Among them, represents the encoder-decoder network loss function, is the pixel-level L-2 norm reconstruction loss function; Train a domain-invariant classifier by minimizing the following classification loss of all source domain output images: Among them, represents the domain-invariant classifier loss function, and θ f is the parameter of the domain-invariant classifier F di ; Classification loss function As follows Where λ1, λ2, and λ3 are trade-off hyperparameters that control the corresponding loss weights.
3. The domain generalization heterologous SAR image target recognition method based on the feature separation and diffusion model according to claim 1, characterized in that, For the training of the diffusion model and the classifier, the construction of the loss function is as follows: Among them, represents the diffusion model loss function, is the additive noise feature, ∈ t is the noise added to the additive noise feature, sampled from standard Gaussian noise; t is the diffusion step number, used to characterize the magnitude of the added noise, T is the total number of diffusion steps, f0 is the parameter of the diffusion model; φ is the parameter of the diffusion model, ∈ φ represents the noise of the parameter φ, which is a function of the additive noise feature and the diffusion step number t; Among them, represents the classifier loss function, is the cross-entropy loss function, α is the parameter of the classifier, and G α is the classifier.
4. The method for target recognition of heterogeneous SAR images for domain generalization based on the feature separation and diffusion model according to claim 1, wherein, After the training of the diffusion model is completed, perform multi-step denoising on the noise to obtain enhanced features where n is the sampling noise, is the filtered noise, α t and σ t are the set hyperparameters, corresponding to the amount of noise removed during the denoising process, and t is the diffusion step, used to characterize the magnitude of the added noise; Adjust and modify the filtered noise, which is achieved through the following formula: Among them, is the denoising noise adjusted according to the required generation category, is the adjustment direction obtained by the classifier according to the required category c, v is the adjustment value, and p α represents the probability of the parameter α; For the corresponding class label c of the denoised enhanced samples, use the cross-entropy objective function to fine-tune the diffusion model: Among them, represents the denoising loss function, is the cross-entropy loss function, f0 are the parameters of the diffusion model, are the features of the diffusion model, G ω represents the classifier; w represents the parameters of the classifier; The loss function is as follows: Where μ1, μ2, and μ3 are trade-off hyperparameters.
5. A domain generalization heterologous SAR image target recognition system based on feature separation and diffusion model, characterized in that Including: A target recognition model construction module, used to construct a target recognition model consisting of a backbone network and a classifier; A model training module, used to input SAR images in multiple source domain datasets into the backbone network in sequence, and extract all-element features through the backbone network, including image features in multiple source domains; And separate the image features into domain-specific features and domain-invariant features; Remove the domain-specific features, thus leaving the domain-invariant features for subsequent classification; denoise the additive features through a diffusion model, thus optimizing the backbone network in the form of gradient backpropagation; the classifier further classifies the obtained domain-invariant features to output the target category, and complete the training of the target recognition model; A target recognition module, used to obtain SAR images in real time, and obtain the target category through the trained target recognition model.
6. A domain generalization heterologous SAR image target recognition device based on feature separation and diffusion model, characterized in that, Including a memory and a processor; the memory is used to store a computer program; the processor is used to, when executing the computer program, implement the domain generalization heterologous SAR image target recognition method based on feature separation and diffusion model as described in any one of claims 1-4.
7. A computer-readable storage medium, characterized in that, A computer program is stored on the storage medium, and when the computer program is executed by the processor, the domain generalization heterologous SAR image target recognition method based on feature separation and diffusion model as described in any one of claims 1-4 is implemented.