Domain generalization SAR target recognition method based on pure simulation data training, terminal equipment and storage medium
By using the domain generalization method of pure simulation data training in SAR target recognition, a diverse simulation data domain is constructed, and a new intermediate domain image is generated using the generative style transformation model, the SAR image target recognition accuracy problem is solved in the absence of actual measured data, and an efficient and low-cost recognition effect is achieved.
Patent Information
- Application Number
- CN202510009255.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-03
AI Technical Summary
The prior art is difficult to achieve accurate target recognition of SAR images without actual measured data, especially in the fields of military reconnaissance and non-cooperative target monitoring, obtaining measured data is expensive and difficult to obtain.
The domain generalized SAR target recognition method based on pure simulation data training is adopted. By obtaining simulated images under different noise, background and imaging parameters, multiple source domains are constructed, and a new intermediate domain image is generated using a generative style transformation model to expand the diversity of the simulation data domain, thereby training a target recognition model that can adapt to the data distribution of unknown domains.
It realizes that the accuracy of SAR image target recognition without actual measured data is improved, eliminates dependence on actual measured data, reduces training costs, and is suitable for high-cost and unavailable measured data scenarios.
Smart Images

Figure CN119942267A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of synthetic aperture radar technology, and in particular to a domain generalized SAR target recognition method based on pure simulation data training, a terminal device and a storage medium. Background Art
[0002] In the field of remote sensing, synthetic aperture radar (SAR) is mainly used as a radar imaging system that generates high-resolution images using a wide frequency range. Compared with optical images, SAR plays an important role in surveillance and reconnaissance because it can produce images under all-weather and long-range conditions. However, SAR images show complex electromagnetic scattering phenomena, including various scattering mechanisms due to substructures on the target, clutter or interference signals, and speckle noise. Unlike optical images, SAR images contain intensity and phase information but lack color information. Therefore, human visual understanding of SAR images is quite difficult. Accordingly, manual analysis of large SAR image streams requires a lot of human resources, which has prompted the development of SAR automatic target recognition (ATR).
[0003] In recent years, with the development of deep learning technology, data-driven SAR target recognition has made great progress. In order to obtain a neural network model that can achieve effective target recognition on a specific SAR image, the most common method is to first obtain and annotate a large amount of homologous training data with the same sensor parameters, environmental background, and observation angle, and then design a network structure on the training data, and then solidify the network model through model training and parameter tuning. However, SAR images have different imaging mechanisms from optical images, and the target features are not obvious and have poor identifiability. In addition, SAR images are highly sensitive to the motion conditions and target posture of the imaging platform (such as airborne, satellite, missile-borne, etc.). In order to ensure good target recognition effects, data collection must be carried out again whenever the imaging platform or application conditions change, resulting in extremely high experimental costs. More importantly, in military fields such as military reconnaissance, regional surveillance, and precision strike, due to the non-cooperative and low-exposure characteristics of the target, it becomes more difficult or even impossible to obtain a large amount of homologous target data under real conditions.
[0004] With the development of related technologies such as computer-aided design (CAD) and electromagnetic simulation, it has become possible to generate a large number of high-quality SAR images through simulation. Simulation can flexibly and efficiently modify the motion conditions and target posture of the SAR imaging platform, and can easily obtain detailed label information of the target. Therefore, training deep models on simulation data is becoming a feasible option. At present, some researchers have tried to use SAR image samples obtained by electromagnetic simulation to support the classification task of measured samples. There have been many studies on SAR image simulation technology. Among them, the classic method is to establish a 3D model of the target, combine electromagnetic calculation and computer graphics methods to obtain SAR simulation images. With the development of SAR simulation technology, for some specific targets, simulated images can achieve a more realistic visual effect. However, experimental results show that it is difficult for neural networks trained directly using pure simulated images to accurately classify real images directly, because there are inevitably certain differences between SAR simulated images and measured images in terms of background texture, target structure, etc., that is, there is a significant domain offset between the simulation domain and the measured domain.
[0005] Most researchers have tried to solve the main problem of domain shift through domain adaptation algorithms, that is, to achieve cross-domain target recognition by aligning the features of labeled source domain simulation data and unlabeled or lightly labeled target domain measured data. However, domain adaptation algorithms still require a large amount of measured data to participate in training, which is not fully consistent with the constraints of actual application scenarios, especially non-cooperative targets and military applications. How to propose an effective and universal method to generalize the model to unseen measured data when only simulation data is available is an urgent problem to be solved.
[0006] CN113762203A discloses a cross-domain adaptive SAR image classification method, device and equipment based on simulated data. Its unsupervised domain adaptive learning algorithm must use labeled source domain simulated data and unlabeled target domain measured data at the same time during training. The fundamental principle is to calculate the gap between the features extracted by the two domains during training, and optimize the feature extractor through training to minimize the gap between the two domains. Therefore, it is necessary to obtain measured data from the target domain to participate in the training process. However, it is very difficult and costly to obtain measured SAR data for training in actual application scenarios, especially under the constraints of non-cooperative targets, military applications, etc., measured data cannot be obtained for model training at all, and can only be trained using pure simulated data. Therefore, this solution does not fully meet actual needs. Summary of the invention
[0007] The technical problem to be solved by the present invention is to provide a domain generalized SAR target recognition method, terminal device and storage medium based on pure simulation data training to improve the accuracy of domain generalized SAR image target recognition based on pure simulation data training.
[0008] In order to solve the above technical problems, the technical solution adopted by the present invention is: a domain generalization SAR target recognition method based on pure simulation data training, comprising the following steps:
[0009] S1, obtaining simulated images with different noises, different backgrounds, and different imaging parameters to form multiple source domains for training;
[0010] S2. Use the generative style transfer model to convert simulated images of different source domains to generate new intermediate domain images, and add the new intermediate domain images to the corresponding source domains to obtain new source domains;
[0011] S3. Using the new source domain as input of a target recognition model, training the target recognition model, and obtaining a recognition model.
[0012] The present invention is based on a domain generalization learning algorithm, and only needs to use labeled source domain simulation data during training, without using target domain data. The fundamental principle of the present invention is to generate diversified simulation data as the source domain, and synthesize a new domain by mixing the underlying feature statistics of training samples from different source domains, so that the model learns domain-invariant features rather than features of a specific domain during training, which enables the model to adapt to the data distribution of unknown domains, thereby achieving accurate classification of measured SAR images. Therefore, by adopting the solution of the present invention, pure simulation data can be used for training, which better meets actual needs.
[0013] The present invention obtains simulated SAR images under different noise, background, imaging parameters and other conditions to form multiple source domains that can be used for training. Considering that it still costs money to obtain simulated SAR images and the domain diversity may be insufficient, the present invention uses a generative style transfer model to convert between the obtained simulated images of different domains, generate new simulated images and add them as source domains that can be used for training, so as to expand the diversity of the simulated data domain. The present invention does not need to obtain measured SAR images at all during the training process, which greatly removes the restrictions.
[0014] The specific implementation process of step S2 includes:
[0015] Classify simulated SAR images with similar parameters and construct different simulated SAR image domains;
[0016] Using the different simulated SAR image domains as inputs of a style transfer network, and training the style transfer network;
[0017] The trained style transfer network is used to convert between different simulated SAR image domains to obtain new intermediate domain images.
[0018] The style transfer network adopts the CUT network.
[0019] In step S3, the specific implementation process of training the target recognition model includes:
[0020] The SAR simulated images with consistent parameters are classified into the same source domain and given the same domain labels 1 to N. The generated intermediate domain images are given another domain label N+1, and each simulated image is given a corresponding category label. The simulated images with domain labels and category labels are input into the target recognition model, and the mixed feature statistics are calculated by Mixstyle. The mixed feature statistics are used to replace the single domain features, and the target recognition model is trained according to the domain invariant features.
[0021] During training, the present invention inputs the given domain labels and category labels of SAR simulated images from different domains into the network, and calculates the mixed feature statistics through the Mixstyle module to replace the original single-domain features, so that the trained convolutional neural network has the ability to extract domain-invariant features, and then trains the classifier based on the domain-invariant features, so that the overall network has the ability to classify SAR images in a domain-generalized manner. During testing, the Mixstyle module is disabled, and the domain-invariant features in the measured images are directly extracted through the trained convolutional neural network to achieve target classification. Finally, it is achieved as a whole that a neural network model that can accurately classify measured images is trained using only simulated images.
[0022] The method of the present invention further comprises:
[0023] S4. Acquire a measured SAR image to be identified in the target domain, input the measured SAR image into a recognition model, and classify and identify the target category in the measured SAR image.
[0024] The target classification model adopts Resnet50; the Resnet50 includes multiple cascaded blocking blocks, wherein the second blocking block and the third blocking block, the third blocking block and the fourth blocking block, and the fourth blocking block and the fifth blocking block are each connected by a Mixstyle module.
[0025] As an inventive concept, the present invention also provides a terminal device, including a memory, a processor, and a computer program stored in the memory; the processor executes the computer program to implement the steps of the above method.
[0026] As an inventive concept, the present invention also provides a computer-readable storage medium having a computer program / instruction stored thereon; the computer program / instruction implements the steps of the above method when executed by a processor.
[0027] As an inventive concept, the present invention also provides a computer program product, including a computer program / instruction; when the computer program / instruction is executed by a processor, the steps of the above method are implemented.
[0028] Compared with the prior art, the beneficial effects of the present invention are as follows: the present invention does not need to use the measured data of the target domain at all during the training process, eliminating the limitations of other domain adaptive algorithms, and is far superior to general general classification and recognition methods in recognition performance, thereby improving the accuracy of domain generalized SAR image target recognition based on pure simulation data training. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 It is a flowchart of a domain generalization SAR target recognition method based on pure simulation data training in one embodiment;
[0030] Figure 2 A schematic diagram of a process of expanding a simulation data domain using a generative style transfer model in one embodiment;
[0031] Figure 3 A schematic diagram of cross-domain target recognition data distribution in one embodiment;
[0032] Figure 4 A schematic diagram showing a comparison between domain generalization and domain adaptation processes in an embodiment;
[0033] Figure 5 Schematic diagram of a residual learning module of ResNet in one embodiment;
[0034] Figure 6 A schematic diagram of the ResNet50 network structure in which a domain-style mixing module is added in one embodiment. DETAILED DESCRIPTION
[0035] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0036] Example 1
[0037] Data-driven SAR target recognition model training relies on a large amount of annotated homologous measured data. However, SAR images are not only expensive to acquire, but measured data of non-cooperative targets are almost impossible to obtain, which limits the practical application of data-driven SAR target classification algorithms.
[0038] In view of the above problems, in one embodiment, Figure 1 As shown, this embodiment proposes a simulation data model training method based on domain generalization, comprising the following steps:
[0039] Step S1, obtaining simulated images under different noise, background, imaging parameters and other conditions to form multiple source domains that can be used for training.
[0040] Specifically, for an invisible measured image domain, it is assumed that multiple corresponding simulated image domains (such as multiple noise coefficients, imaging parameters, etc.) can be obtained, which is completely feasible for the increasingly developed and improved SAR simulation system. The simulation method can flexibly and efficiently modify the motion conditions of the SAR imaging platform and the target posture, and can easily obtain detailed label information of the target. However, even if the simulated SAR images obtained under exactly the same parameter conditions are still different from the measured SAR images, such as the strong scattering points of the target body and the texture information of the background. These differences are caused by the accuracy of the target CAD model and the electromagnetic scattering calculation. Therefore, it is difficult for the neural network trained directly using the simulated image to directly classify the real image accurately.
[0041] Step S2: Considering that it still costs money to obtain simulated images, the domain diversity may be insufficient. Using the generative style transfer model, the simulated images of different domains are converted to generate new simulated images that can be added as source domains for training, so as to expand the diversity of the simulated data domain.
[0042] Specifically, we use available simulated SAR images to classify images with similar parameters and construct several simulated SAR image domains. These data are used as training sets to train image-to-image style transfer networks such as the CUT network. The trained style transfer network is used to convert between different simulated SAR image domains to generate new intermediate domain images. The newly generated images enrich the style diversity of the domain while retaining the invariant features of the target domain, making the training data more sufficient and enhancing the robustness of the target recognition network to style transformation.
[0043] Step S3, input the source domain image into the target recognition model for iterative training. By mixing the underlying feature statistics of training samples from different domains, a new domain can be implicitly synthesized, increasing the diversity of the source domain at the instance-level feature level, thereby improving the generalization of the training model and enabling the model to extract domain-invariant features. After training until all loss functions converge, a target recognition model with domain generalization capability is obtained.
[0044] Specifically, during training, SAR simulated images from different domains are given domain labels and category labels and input into the network. The Mixstyle module is used to calculate the mixed feature statistics to replace the original single-domain features, so that the trained convolutional neural network has the ability to extract domain-invariant features. Then, the classifier is trained based on the domain-invariant features, so that the overall network has the ability to classify SAR images in a domain-generalized manner. During testing, the Mixstyle module is disabled, and the domain-invariant features in the measured images are directly extracted through the trained convolutional neural network to achieve target classification. Ultimately, a neural network model that can accurately classify measured images is trained using only simulated images.
[0045] Step S4, obtaining a measured SAR image to be identified in the target domain, inputting the measured SAR image into the trained domain generalization target recognition model to classify and identify the target category in the measured SAR image.
[0046] Specifically, the measured SAR images in the target domain are tested. In this step, the parameters of the feature extractor and classifier trained in the above steps are disabled when testing the measured SAR images in the target domain. The mixed feature statistics are not calculated, and the gradient backpropagation is not performed to update the model parameters. The recognition rate can be calculated based on the categories predicted by the model and the label information of the measured SAR images in the target domain.
[0047] In this embodiment, the core idea of the method is: first, use simulation technology and generative models to obtain a variety of simulated or synthetic images, and construct a source domain with diverse styles but consistent target features. Then, in the process of training the model with simulated data, the domain generalization technology is used to enable the model to adapt to the data distribution of the unknown domain, so as to finally achieve accurate classification of the measured data. Specifically, the method uses pure simulation data to train the model and learns domain-invariant features through the Mixstyle module. The Mixstyle module synthesizes a new domain by mixing the underlying feature statistics of training samples from different domains, thereby improving the generalization ability of the model. During the training process, the model learns domain-invariant features rather than features of a specific domain, which enables the model to adapt to the data distribution of the unknown domain, thereby achieving accurate classification of measured SAR images.
[0048] In one embodiment, step S2 includes: taking a group of SAR simulation images with less clutter noise as style domain A, and taking another group of SAR simulation images with greater clutter noise as style domain B, and using them as training data for the generative style transfer model CUT. The training model converts between style domains A and B, so that the network can generate intermediate domain data between style domains A and B. The generated intermediate domain data is supplemented to the simulation data as training data for the domain generalization target recognition model, which expands the domain style diversity of the training data for clutter noise.
[0049] like Figure 2 As shown in the figure, CUT is a GAN-based image-to-image style transfer model that does not require pairing. It combines the ideas of contrastive learning and generative adversarial learning, uses the maximum mutual information of input and output image blocks to replace the cycle consistency loss, uses the information noise contrast loss (infoNCE Loss) as the estimated contrast loss function, and only needs to train one generator and discriminator. In this example, we obtained two sets of unpaired synthetic aperture radar (SAR) images: The generator G of CUT is divided into two parts: encoder and decoder. CUT first uses the encoder G l,enc Decompose the input image X into multiple feature maps Z l =H l (G l,enc (X)), where l represents the layer index, H l It is a two-layer multi-layer perceptron network. Then use the decoder G l,dec Mapping feature maps to generated images The generated image still needs to be encoded to obtain feature encoding Used to calculate the PatchNCE loss. PatchNCE loss is used to compare the generated images Each image block in is matched with its corresponding image block in the source image X, so that the features at the corresponding positions are as consistent as possible, while other image blocks in the source image X are regarded as negative samples, so that the feature gaps at different positions are as large as possible. NCE refers to noise contrast estimation. This loss function encourages the encoder to learn a representation that makes corresponding image blocks closer in the feature space. At the same time, it ensures that the distance between non-corresponding image blocks is farther, thereby maximizing mutual information. PatchNCE loss is specifically defined as:
[0050]
[0051] Among them, l and s represent the layer index and the spatial position within the layer respectively. is the feature vector of the image block corresponding to the query key, is the stack of feature vectors from negative image patches. stands for the information noise contrast loss (InfoNCE) loss function, which is used to maximize mutual information in contrastive learning. The contrastive learning problem consists of three signals: query key, positive sample and negative sample. What needs to be done is to make the query key and positive sample signal correlated and contrasted with the negative sample. First, the query key, positive sample and N negative samples are mapped into K-dimensional vectors respectively. in It means that the nth negative sample normalizes these samples to the unit ball to prevent the space from expanding or collapsing. This forms an N+1 classification problem. The cross entropy loss is calculated as follows, where τ is a proportional hyperparameter, often called a temperature coefficient, which indicates the probability that a positive sample is selected.
[0052]
[0053] In the image generation part of GAN, CUT still uses the discriminator adversarial loss of GAN. In this example, the cross entropy loss is used to ensure that the generated image is as similar as possible to the image in the target domain. The loss of this part is:
[0054]
[0055] Among them, D is the discriminator. In this example, a two-layer patchGAN binary classifier is used. y is the real image from the target domain, and x is the input image from the source domain.
[0056] The final optimization objective adds a consistency loss L PatchNCE (G,H,Y) to minimize Avoid unnecessary changes in the target characteristics of the generated image by the generator. Therefore, the total loss consists of three parts: adversarial loss, contrast loss, and consistency loss:
[0057] L cut =L GAN (G,D,X,Y)+L PatchNCE (G,H,X)+L PatchNCE (G,H,Y)
[0058] CUT focuses on the image patches inside the image rather than the entire image, which can achieve more accurate target feature preservation and efficient learning with limited data. At the same time, using image patches from the same input image as negative samples instead of other images reduces the need for large datasets and simplifies the learning process. The proposed model consists of a single generator and discriminator, so it is relatively simple compared to other methods and is more suitable for data augmentation tasks with small samples.
[0059] In one embodiment, if Figure 3The goal of domain generalization is to learn a model from one or several different but related domains (training sets) and generalize well on unseen test domains. The goal is to obtain a generalized prediction function h after training with several source domain data, so that its test error on unseen targets is minimized. The source domain refers to the domain where data is easy to obtain and a large amount of labeled data is available. In this example, it refers to simulated and synthetic SAR images, while the target domain refers to the domain where data is difficult to obtain, only a small amount of data is available or even completely invisible. In this example, it refers to measured SAR images. So far, researchers in SAR automatic target recognition have used domain adaptation methods to reduce domain differences. Domain adaptation is a method that involves training with source domain data and a small amount of target domain data at the same time, using the alignment of two domains at the feature level, or using a generative model to convert domains at the image level. Although previous studies have achieved high classification accuracy in the target domain, they still rely on the use of target domain knowledge during training, and are easily limited when applied to scenarios where target domain data cannot be obtained. To eliminate this limitation, we focus on the application of domain generalization technology in SAR automatic target recognition. The difference between domain adaptation and domain generalization is as follows: Figure 4 shown.
[0060] In this embodiment, in step S3, although the convolutional neural network (CNN) has shown excellent ability in learning the discriminative features of SAR images, the network's generalization ability to unseen domains is often poor. For this reason, this application introduces a new method of instance-level feature statistics based on probabilistic mixing of cross-source domain training samples with reference to research in the field of optics, motivated by the observation that the visual domain is closely related to image style. Such style information can be captured by the bottom layer of CNN, and the proposed style mixing is processed here. By mixing the underlying feature statistics of training samples from different domains, new domains can be implicitly synthesized, increasing the domain diversity of the source domain, thereby improving the generalization of the training model.
[0061] In style transfer models, instance-specific mean and standard deviation are widely used to normalize feature tensors, which has been proven to effectively remove image style information. This operation is called instance normalization (IN). Instance normalization eliminates differences between instances by subtracting the mean and dividing by the standard deviation on each channel of the feature map of each instance, thereby achieving style transfer. The formula is as follows:
[0062]
[0063] in, represents the input feature map, B, C, H, and W represent the batch size, number of channels, and feature map height and width, respectively. are learnable affine transformation parameters, are the mean and standard deviation of the feature map of each instance on each channel, respectively, and the calculation formula is as follows:
[0064]
[0065] In order to achieve arbitrary style transfer, Huang & Belongie proposed Adaptive Instance Normalization (AdaIN). AdaIN replaces γ and β in instance normalization with the feature statistics of style input γ, thereby achieving conversion between different styles. Its formula is as follows:
[0066]
[0067] Among them, y represents the style input, Still represents the mean and standard deviation on each channel.
[0068] This application draws on the idea of AdaIN, but applies it to domain generalization tasks. Unlike image generation decoders, this method introduces a plug-and-play module (Mixstyle) to regularize CNN training by perturbing the style information of source domain training instances, and inserts it between CNN layers such as CNN classifiers without explicitly generating new styles of images. The goal of this method is to synthesize new domains by mixing the underlying feature statistics of training samples from different domains, thereby improving the generalization ability of the model. More specifically, MixStyle mixes the feature statistics of two instances with a random convex weight to simulate the style of the new domain. The implementation steps are as follows:
[0069] First, generate a mixed reference batch: given a domain label, randomly select two instances from the input batch x in two different domains and swap their positions, then shuffle each batch to generate a reference batch Next, calculate the mixed feature statistics: use the Beta distribution to randomly sample instance weights λ and calculate the mean and standard deviation of the mixture according to the following formula:
[0070]
[0071] Among them, λ~Beta(α,α) means sampling from Beta distribution, α is a hyperparameter generally set to 0.1. Finally, apply the mixed feature statistics and apply the mixed feature statistics to the style normalized input batch x to get the final output:
[0072]
[0073] For the domain generalization task from the SAR image simulation domain to the measured domain, this application uses Resnet50 as the basic structure of the classification model. Resnet is the abbreviation of Residual Network. This series of networks is widely used in the field of image target classification and as part of the backbone classical neural network for computer vision tasks. Its residual structure is as follows Figure 5 As shown. ResNet is divided into 5 stages (Stage), among which Stage 0 has a relatively simple structure and can be regarded as a preprocessing of the input rather than feature extraction. The last 4 stages are composed of Bottleneck (BTNK) containing residual blocks, and the structure is relatively similar. Stage 1 contains 3 Bottlenecks, and the remaining 3 stages include 4, 6, and 3 Bottlenecks respectively. The output of each Stage 1-4 can be regarded as a level of style features. This application inserts the Mixstyle module after Stage 1-3 so that the network can learn domain-invariant features. The reason why the module is not inserted after Stage 4 is that it is closest to the prediction layer and tends to capture semantic content (ie, label-sensitive) information rather than style. And Stage 4 is followed by an average pooling layer, which essentially forwards the mean vector to the prediction layer, thereby forcing the mean vector to capture label-related information. Therefore, mixing statistical information in Stage 4 will destroy the inherent label space, and the overall network structure is as follows. Figure 6 shown.
[0074] For an invisible measured image domain, it is assumed that multiple corresponding simulated image domains (such as multiple noise coefficients, imaging parameters, etc.) can be obtained. This is completely feasible for the increasingly developed and improved SAR simulation system. During training, the SAR simulated images from different domains are given domain labels and category labels and input into the network. The mixed feature statistics are calculated by Mixstyle to replace the original single domain features, so that the trained convolutional neural network has the ability to extract domain invariant features. Then, the classifier is trained based on the domain invariant features, so that the overall network has the ability to generalize SAR image classification. During testing, the Mixstyle module is disabled, and the domain invariant features in the measured image are directly extracted through the trained convolutional neural network to achieve target classification. Finally, it is achieved that a neural network model that can accurately classify measured images can be trained using only simulated images.
[0075] In a verification example, the method of this example is experimentally verified and analyzed to prove its effectiveness, which is carried out from four aspects: experimental data introduction, experimental setting, result comparison and summary analysis.
[0076] (1) Introduction of experimental data
[0077] This experiment used the SAMPLE dataset and the S2M-5 dataset of measured and simulated images made by comparing it with the standard dataset and using the MSTAR dataset. The following will introduce these two datasets in detail.
[0078] The first one used is the SAMPLE dataset, which is a publicly released dataset of measured and simulated images of a SAR vehicle target. Its imaging conditions are similar to the standard operating conditions SOC of the MSTAR dataset. It is imaged in the X-band and HH polarization mode, with an image resolution of 0.3m and an image size of 128×128 pixels. The publicly released SAMPLE dataset contains two sets of data with different styles but completely consistent target parameters (denoted as 'decibel' and 'qpm'). They both include 10 types of vehicle targets, and the data are sampled at azimuth angles of 10° to 80° and elevation angles of 14° to 17°. And for each measured image, the dataset provides a simulated image with matching imaging conditions. Table 1 shows the number of measured and simulated images of each type of target in the publicly released SAMPLE dataset, as well as the specific models corresponding to each type of target.
[0079] Table 1 SAMPLE dataset 10 categories of targets
[0080]
[0081] The second dataset is the self-built measured and simulated image pair dataset S2M-5, which is derived from the SAMPLE and MSTAR datasets and contains 5 types of simulated and real vehicle targets. Specifically, our team collected 5 types of data with the same categories and models as the SAMPLE dataset in the MSTAR dataset as the real set of the self-built dataset, and the simulation set directly used the data in the SAMPLE dataset. The azimuth range of the simulation data is still between 10° and 80°, and the pitch range is between 14° and 17°; the azimuth range of the measured data is between 0° and 90°, and the pitch is limited to 15° and 17°. Unlike the SAMPLE data, there is no strict one-to-one correspondence between the simulated images and the real images of the self-built dataset, and the differences in factors such as target contours, background, and noise are greater, making the domain offset from the simulation domain to the measured domain larger than that of the SAMPLE data. Table 2 shows the number of measured images and simulated images of each type of target in the self-built S2M-5 dataset, as well as the specific models corresponding to each type of target.
[0082] Table 2 Five categories of targets in the self-built dataset S2M-5
[0083]
[0084] (2) Experimental setup
[0085] Only simulated images are used for training on both datasets, and tested on measured images. In the domain generalization task, for the SAMPLE dataset, in this example, the simulated datasets from the qpm and decibel groups are used as two visible source domains, namely the training set, and the measured data of qpm is used as the invisible target domain, namely the test set. For the training of the S2M-5 dataset, the simulated data from SAMPLE are also used as two source domains, but the measured data from MSTAR is used as the target domain. The size of the input image is resized to 256×256 pixels and cropped from the center to 224×224 pixels before inputting into the network for training and testing. In order to further enrich the domain diversity, random rotation of 0-90° and random color perturbation are additionally added as data augmentation methods during the training process. The optimizer uses the Adam optimizer, the initial learning rate is set to 0.001, the batch size is set to 8, and the training rounds are 30 rounds. In order to reduce the impact of randomness, this example sets ten different random seeds for each algorithm to train ten times, and calculates their average accuracy and standard deviation in the test set as the evaluation criteria of the algorithm.
[0086] (3) Results comparison
[0087] This example first selected several classic deep classification models for comparison to reflect the difficulty of deep model training caused by the domain gap between simulation and measured data. It also selected a generalized learning method RN18+ designed for SAR target recognition for comparison to reflect the superiority of the algorithm proposed in this example. All models were trained only on simulated images and tested on measured images. The hyperparameters of all models were adjusted to the optimal state. The results are shown in Table 3.
[0088] Table 3 Comparison of classification and recognition results
[0089]
[0090] (4) Summary and analysis
[0091] From the comparison of the results in Table 3, it can be found that when only simulated images are used for training, due to the existence of domain gap, classic traditional classification models such as VGG16, ResNet18, and ResNet50 generally perform poorly on measured data. This is reflected in the SAMPLE data, and is even more obvious on the S2M-5 dataset with a larger domain gap. Networks specially designed for SAR image classification, such as A-ConvNet, have improved performance compared to traditional models, but the accuracy is still not ideal. Although the generalization training method RN18+ designed for simulation to measured data can achieve an accuracy of 89.35% on the SAMPLE dataset, it has an accuracy of only 36.35% on the S2M-5 data with a larger domain gap and closer to the actual situation, indicating that this method has poor generalization and cannot be used in practice. The method based on domain generalization technology proposed in this example not only achieved an average accuracy of 90.04% on the SAMPLE dataset, with a smaller variance and more stability, but also had an average accuracy of 82.14% on the S2M-5 dataset, which fully demonstrated the generalization and superiority of the method proposed in this example.
[0092] Embodiment 2 of the present invention provides a terminal device corresponding to the above-mentioned embodiment 1. The terminal device may be a processing device for a client, such as a mobile phone, a laptop computer, a tablet computer, a desktop computer, etc., to execute the method of the above-mentioned embodiment.
[0093] The terminal device of this embodiment includes a memory, a processor, and a computer program stored in the memory; the processor executes the computer program in the memory to implement the steps of the method in the above-mentioned embodiment 1.
[0094] In some implementations, the memory may be a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk memory.
[0095] In some other implementations, the processor may be a central processing unit (CPU), a digital signal processor (DSP), or other general-purpose processors of various types, which are not limited herein.
[0096] Example 3
[0097] Embodiment 3 of the present invention provides a computer-readable storage medium corresponding to the above embodiment 1, on which a computer program / instruction is stored. When the computer program / instruction is executed by a processor, the steps of the method of the above embodiment 1 are implemented.
[0098] Computer readable storage media can be tangible devices that hold and store instructions used by instruction execution devices. Computer readable storage media can be, for example, but not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any combination thereof.
[0099] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of complete hardware embodiments, complete software embodiments, or embodiments in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code. The scheme in the embodiments of the present application can be implemented in various computer languages, for example, object-oriented programming language Java and literal scripting language JavaScript, etc.
[0100] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0101] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0102] Although the preferred embodiments of the present application have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0103] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. A domain generalization SAR target recognition method based on pure simulation data training, characterized in that: The following steps are involved: S1, obtaining simulated images with different noises, different backgrounds, and different imaging parameters to form multiple source domains for training; S2. Use the generative style transfer model to convert simulated images of different source domains to generate new intermediate domain images, and add the new intermediate domain images to the corresponding source domains to obtain new source domains; S3. Using the new source domain as input of a target recognition model, training the target recognition model, and obtaining a recognition model.
2. The domain generalized SAR target recognition method based on pure simulation data training according to claim 1 is characterized in that: The specific implementation process of step S2 includes: Classify simulated SAR images with similar parameters and construct different simulated SAR image domains; Using the different simulated SAR image domains as inputs of a style transfer network, and training the style transfer network; The trained style transfer network is used to convert between different simulated SAR image domains to obtain new intermediate domain images.
3. The domain generalized SAR target recognition method based on pure simulation data training according to claim 2 is characterized in that: The style transfer network adopts the CUT network.
4. The domain generalized SAR target recognition method based on pure simulation data training according to claim 1 is characterized in that: In step S3, the specific implementation process of training the target recognition model includes: The SAR simulated images with consistent parameters are classified into the same source domain and given the same domain labels 1 to N. The generated intermediate domain images are given another domain label N+1, and each simulated image is given a corresponding category label. The simulated images with domain labels and category labels are input into the target recognition model, and the mixed feature statistics are calculated by Mixstyle. The mixed feature statistics are used to replace the single domain features, and the target recognition model is trained according to the domain invariant features.
5. The domain generalized SAR target recognition method based on pure simulation data training according to any one of claims 1 to 4, characterized in that: Also includes: S4. Obtain a measured SAR image to be identified in the target domain, input the measured SAR image into a recognition model, and classify and identify the target category in the measured SAR image.
6. The domain generalized SAR target recognition method based on pure simulation data training according to any one of claims 1 to 4, characterized in that: The target classification model adopts Resnet50; the Resnet50 includes multiple cascaded blocking blocks, wherein the second blocking block and the third blocking block, the third blocking block and the fourth blocking block, and the fourth blocking block and the fifth blocking block are each connected by a Mixstyle module.
7. A terminal device comprising a memory, a processor and a computer program stored in the memory; characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 6.
8. A computer-readable storage medium having a computer program / instruction stored thereon; characterized in that: When the computer program / instructions are executed by a processor, the steps of the method described in any one of claims 1 to 6 are implemented.
9. A computer program product comprising a computer program / instructions; characterized in that: When the computer program / instructions are executed by a processor, the steps of the method of any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Cross-domain adaptive SAR image classification method and device based on simulation data, and equipment
CN113762203A
Heterogeneous and heterogenous SAR target identification method based on domain adaptation
CN114529766A
SAR small sample target detection method
CN114821294A
Weak supervision real-time target detection method based on progressive diversified domain migration
CN115565005A
Cross-pitch-angle SAR (Synthetic Aperture Radar) image target identification method and device and computer equipment
CN116863234A