Methods, apparatus and computer equipment for target recognition in SAR images across elevation angles

By simulating image realization processing and joint distribution domain adaptation, the target recognition model is optimized, which solves the problem of insufficient model generalization ability in target recognition of SAR images across elevation angles and improves recognition accuracy.

CN116863234BActive Publication Date: 2025-11-14NAT UNIV OF DEFENSE TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310859839.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-13
Publication Date
2025-11-14
Estimated Expiration
2043-07-13

AI Technical Summary

Technical Problem

Existing deep learning methods suffer from insufficient model generalization ability in target recognition of SAR images across elevation angles, especially when the training and testing data distributions are mismatched under different elevation angle imaging conditions, resulting in a significant drop in recognition performance.

Method used

By processing simulated images to generate images under corresponding pitch angle conditions, and combining prior information from the simulation data, the CycleGAN method is used to reduce the difference between the simulated and measured images. Supervised and unsupervised domain adaptation training is performed to optimize the target recognition model.

Benefits of technology

It improves the accuracy of target recognition in SAR images across elevation angles and enhances the model's generalization ability under different elevation angle conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116863234B_ABST
    Figure CN116863234B_ABST
Patent Text Reader

Abstract

This application relates to a method, apparatus, and computer device for target recognition in SAR images across elevation angles. The method includes: realigning simulated images, aligning and adapting the distribution domains between measured SAR images, aligning and adapting the edge distribution domains between generated images after realignment processing, and optimizing a target recognition model. This target recognition model is then used to classify target categories in the measured SAR images within the target domain. This method significantly outperforms other contrast domain adaptation methods in terms of recognition performance, improving the accuracy of target recognition in SAR images across elevation angles.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of synthetic aperture radar technology, and in particular to a method, apparatus and computer equipment for target recognition in SAR images across elevation angles. Background Technology

[0002] Synthetic Aperture Radar (SAR) is an active microwave imaging sensor that can penetrate clouds, rain, snow, and smoke, providing all-weather, all-day imaging capabilities.

[0003] With the continuous advancement of artificial intelligence, the close integration of deep learning and SAR image interpretation is beneficial for improving the performance of SAR image target recognition tasks. Convolutional neural networks (CNNs) are the most commonly used models in deep learning image classification and recognition tasks, but they are highly dependent on training data. When there is a significant difference in feature distribution between the training and test data, the network's generalization performance drops sharply. Target recognition methods based on deep neural networks mainly use CNNs as the backbone network and then embed modules that improve SAR target recognition performance into the network, such as electromagnetic scattering feature modules, interpretability modules, attention mechanism modules, pseudo-label modules, and feature fusion modules.

[0004] Existing deep learning methods have achieved good performance in SAR target recognition, but they typically only use labeled source domain data for supervised training. This approach requires training and testing data to come from the same or similar probability distributions; however, when this condition is difficult to meet, the performance of the neural network model drops significantly. SAR images are highly sensitive to imaging condition parameters. Even when imaged under the same sensor, different imaging condition parameters can lead to a mismatch between the distribution of training and testing data, which significantly reduces the model's generalization ability. The elevation angle of a SAR image refers to the angle between the radar signal and the horizontal direction of the ground. Different elevation angles affect the resolution, geometric distortion, and target recognition capabilities of SAR images. In vehicle target recognition tasks, due to the significant differences in imaging of the same target at different elevation angles, cross-elevation angle target detection is a highly challenging task, especially across large elevation angles. How to solve the problem of SAR image target recognition under different imaging conditions, represented by cross-elevation angle SAR image target recognition, is of great research significance.

[0005] Simulation technology can directly map imaging conditions and target 3D models into SAR images. Therefore, simulated SAR images can reflect the direct impact of imaging conditions on image features, and some studies have introduced simulated SAR images into the SAR image target recognition process. However, these studies all consider the case where the imaging differences are not significant, such as SOC (Surface Area Code), and do not take into account the differences between simulated SAR images and measured SAR images. The combination of deep learning and domain adaptation methods is currently a research hotspot in transfer learning. Classical methods use a domain adaptation regularization term to align the edge feature probability distributions of the source and target domains. However, due to the small amount of SAR target recognition data and the large imaging differences, simply aligning the edge distributions is insufficient to significantly improve recognition performance. Summary of the Invention

[0006] Therefore, it is necessary to provide a method, apparatus, and computer equipment for target recognition in SAR images across elevation angles to address the aforementioned technical problems.

[0007] A method for target recognition in SAR images across elevation angles, the method comprising:

[0008] The simulated SAR images and measured SAR images under different elevation angle conditions are processed to make the simulated images realistic, and generated images under the corresponding elevation angle conditions are generated. The measured SAR images include measured SAR images in the source domain and the target domain, and the generated images include generated images in the source domain and the target domain.

[0009] The generated images of the source and target domains, as well as the measured SAR images of the source and target domains, are input into the target recognition model for iterative training to obtain the trained target classification model. The training process includes two stages: the first training stage: supervised classification training and unsupervised domain adaptation training with aligned edge distribution are performed using the generated images of the source and target domains; the second training stage: supervised classification training and unsupervised domain adaptation training with aligned joint distribution are performed using the measured SAR images of the source and target domains.

[0010] Acquire a measured SAR image of the target domain, and input the measured SAR image into a trained target classification model to classify the target categories in the measured SAR image.

[0011] A target recognition device for SAR images across elevation angles, the device comprising:

[0012] The simulation image realization processing module is used to perform simulation image realization processing on the acquired simulation SAR images and measured SAR images under different elevation angle conditions, and generate generated images under the corresponding elevation angle conditions; the measured SAR images include measured SAR images in the source domain and the target domain, and the generated images include generated images in the source domain and the target domain.

[0013] The target recognition model training module is used to input generated images of the source and target domains, as well as measured SAR images of the source and target domains, into the target recognition model for iterative training to obtain the trained target classification model. The training process includes two stages: the first training stage: supervised classification training and unsupervised domain adaptation training with aligned edge distributions are performed using generated images of the source and target domains; the second training stage: supervised classification training and unsupervised domain adaptation training with aligned joint distributions are performed using measured SAR images of the source and target domains.

[0014] The target recognition module is used to acquire measured SAR images of the target domain and input the measured SAR images into a trained target classification model to classify the target categories in the measured SAR images.

[0015] A computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:

[0016] The simulated SAR images and measured SAR images under different elevation angle conditions are processed to make the simulated images realistic, and generated images under the corresponding elevation angle conditions are generated. The measured SAR images include measured SAR images in the source domain and the target domain, and the generated images include generated images in the source domain and the target domain.

[0017] The generated images of the source and target domains, as well as the measured SAR images of the source and target domains, are input into the target recognition model for iterative training to obtain the trained target classification model. The training process includes two stages: the first training stage: supervised classification training and unsupervised domain adaptation training with aligned edge distribution are performed using the generated images of the source and target domains; the second training stage: supervised classification training and unsupervised domain adaptation training with aligned joint distribution are performed using the measured SAR images of the source and target domains.

[0018] Acquire a measured SAR image of the target domain, and input the measured SAR image into a trained target classification model to classify the target categories in the measured SAR image.

[0019] The aforementioned method, apparatus, and computer equipment for target recognition in SAR images across elevation angles include: realignment processing of simulated images, alignment and joint distribution domain adaptation between measured SAR images, alignment and edge distribution domain adaptation between generated images after realignment processing, and optimization of the target recognition model. This target recognition model is then used to classify target categories in the measured SAR images within the target domain. This method significantly outperforms other contrast domain adaptation methods in terms of recognition performance, improving the accuracy of target recognition in SAR images across elevation angles. Attached Figure Description

[0020] Figure 1 This is a flowchart illustrating a cross-elevation angle SAR image target recognition method in one embodiment;

[0021] Figure 2 This is a schematic diagram of the measured SAR image recognition process for the target domain in another embodiment;

[0022] Figure 3 This is a schematic diagram of the simulation image realization process in another embodiment;

[0023] Figure 4 Here is a data flow graph for CycleGAN in another embodiment;

[0024] Figure 5 This is a flowchart illustrating the first training phase in another embodiment;

[0025] Figure 6 This is a framework diagram of a classic domain adaptation method based on adversarial learning in another embodiment;

[0026] Figure 7 This is a schematic diagram of the second-stage training process in another embodiment;

[0027] Figure 8 This is the residual learning module of ResNet in another embodiment;

[0028] Figure 9 This is a SAR image target recognition process based on a ResNet-18 network in another embodiment;

[0029] Figure 10 This is a structural block diagram of a cross-elevation angle SAR image target recognition device in one embodiment;

[0030] Figure 11 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0032] Significant differences still exist between SAR images obtained by the same sensor under different imaging conditions, especially in the target outline, shadow areas, and background information when imaging at different elevation angles. To address this issue, this application proposes a target recognition method for SAR images across elevation angles based on Simulation Image and Alignment Joint Distribution Domain Adaptation (SJDDA). This method combines prior information from simulation data with domain adaptation techniques to improve the generalization ability of the neural network model under different elevation angle conditions.

[0033] Addressing the significant differences between image backgrounds and target bodies at different elevation angles, this application proposes a cross-elevation angle SAR image target recognition method based on domain adaptation. The core idea is as follows: First, considering the difficulty in obtaining complete measured SAR datasets, simulation technology directly reflects the relationship between imaging conditions and SAR images. Complete SAR simulation images are relatively easy to obtain. By reducing the domain offset between simulated SAR images under different imaging conditions, the differences caused by imaging conditions can be indirectly reduced. Second, there are inherent differences between simulated SAR images and measured SAR images. To eliminate the impact of these differences, simulated SAR images are processed to realistically represent them, generating images that are more similar to measured SAR images. Then, the realistically represented generated images are introduced into the recognition process. On one hand, supervised classification training of the generated images in the source and target domains and unsupervised domain adaptation training with aligned edge distributions are performed. On the other hand, supervised classification training of the measured SAR images in the source and target domains and unsupervised domain adaptation training with aligned joint distributions are performed. Finally, the trained model is tested on the measured SAR images in the target domain (the test set) to verify its effectiveness.

[0034] In one embodiment, such as Figure 2 As shown, a target recognition method for SAR images across elevation angles is provided, which includes the following steps:

[0035] Step 100: Based on the simulated SAR images and measured SAR images obtained under different elevation angle conditions, perform simulation image realization processing to generate generated images under the corresponding elevation angle conditions; the measured SAR images include measured SAR images in the source domain and the target domain, and the generated images include generated images in the source domain and the target domain.

[0036] Specifically, simulated SAR images obtained under the same elevation angle still exhibit certain differences from measured SAR images, such as strong scattering points of the target and texture information of the background. These differences are caused by factors such as the accuracy of the target CAD model and electromagnetic scattering calculations. Therefore, the inherent differences in the dataset can lead to a decrease in the classification performance of neural networks. To reduce the differences between simulated and measured SAR images, the CycleGAN method is used to achieve realization of the simulated images.

[0037] It is worth noting that in addition to the CycleGAN method, there are also graph-to-graph translation networks similar to pix2pix for the realignment of simulated SAR images.

[0038] Step 102: Input the generated images of the source and target domains, as well as the measured SAR images of the source and target domains, into the target recognition model for iterative training to obtain the trained target classification model. The training process includes two stages: the first training stage: using the generated images of the source and target domains for supervised classification training and unsupervised domain adaptation training with aligned edge distributions; the second training stage: using the measured SAR images of the source and target domains for supervised classification training and unsupervised domain adaptation training with aligned joint distributions.

[0039] Specifically, in the first training phase, the generated images after realignment are subjected to domain adaptation to align edge distributions. Since the realigned target generated images and the measured SAR images in the target domain have similar probability feature distributions, and the simulated SAR images already carry category label information during imaging, the target domain generated images are used for supervised training to optimize the feature extractor and classifier parameters. The purpose of this step is to reduce imaging differences caused by different elevation angles through domain adaptation between generated images, making the classifier parameters more adaptable to the data feature probability distribution of the target domain image. This will improve the accuracy of predictions of the measured SAR images in the target domain during the second training phase.

[0040] In the second training phase, domain adaptation is performed on the source domain measured SAR images and the target domain measured SAR images, using an aligned joint distribution. Compared to the classic domain adaptation method that aligns edge distributions, the aligned joint distribution domain adaptation method aligns the two domains as a whole while performing fine-grained alignment by category, and can extract the "multi-peak structure" of category information in the features. The purpose of this step is to improve the model's ability to recognize different elevation angles, which is reflected in the feature extractor's ability to extract common features from SAR images at different elevation angles.

[0041] This method comprises two phases. The first phase is the simulation of SAR image realization, i.e., step 100. The second phase is the domain adaptation phase, in which the first and second training phases alternate within each iteration (step 102) for alternating optimization. After training is complete, the method proceeds to the testing phase.

[0042] Step 104: Obtain the measured SAR image of the target domain, and input the measured SAR image into the trained target classification model to classify the target categories in the measured SAR image.

[0043] Specifically, the model is tested using measured SAR images of the target domain. In this step, the parameters of the feature extractor and classifier trained in the first and second training phases are fixed. Gradient backpropagation is not performed to update the model parameters when testing the measured SAR images of the target domain. The recognition rate can be calculated using the class predicted by the model and the label information inherent in the measured SAR images of the target domain.

[0044] The actual SAR image recognition process in the target domain is as follows: Figure 2 As shown. Measured SAR images of the target domain are acquired and input into the trained target classification model. Feature extractors are used to extract features from the measured SAR images, and then global average pooling is applied to obtain the target domain features f. T Target domain features f T After passing through the classifier, the target classification prediction result is obtained.

[0045] The aforementioned target recognition method for SAR images across elevation angles includes: realignment processing of simulated images, alignment and joint distribution domain adaptation between measured SAR images, alignment and edge distribution domain adaptation between generated images after realignment processing, and optimization of the target recognition model. This target recognition model is then used to classify target categories in the measured SAR images within the target domain. This method significantly outperforms other contrast domain adaptation methods in terms of recognition performance, improving the accuracy of target recognition in SAR images across elevation angles.

[0046] In one embodiment, step 100 includes: using the simulated SAR image and the measured SAR image at the first elevation angle as the source domain data and target domain data at the first elevation angle, respectively; inputting the source domain data and target domain data at the first elevation angle into CycleGAN to perform realization processing on the simulated SAR image, generating a generated image at the first elevation angle; the generated image at the first elevation angle is a generated image using the simulated SAR image at the first elevation angle as the original image and the measured SAR image at the first elevation angle as the target image; the CycleGAN method is also used to perform realization processing on the simulated SAR images and measured SAR images under other elevation angle conditions to obtain generated images under other elevation angle conditions.

[0047] Specifically, a schematic diagram of the process for realizing simulated SAR images is shown below. Figure 3 As shown, a 17° simulated SAR image is used as the source domain data, and a 17° measured SAR image is used as the target domain data. These are input into CycleGAN. After model training, a 17° generated image is generated, using the 17° simulated SAR image as the source image and the 17° measured SAR image as the target image. This generated image possesses similar content information (target outline, shadows, etc.) to the 17° simulated SAR image, and similar background texture, brightness, and other features, as well as the target region scattering intensity, to the 17° measured SAR image. The same operation is performed between the 30° and 45° simulated SAR images and the measured SAR images, respectively.

[0048] The purpose of employing an image-to-image style transfer method in the simulation image realization module of this application is to reduce the visual style differences between heterogeneous SAR images from different sensors. Specifically, after training, an intermediate domain SAR image with a visual style similar to the target domain SAR image is generated, using the source domain SAR image as the image to be transferred. The latter has a more similar probability distribution in the feature space to the target domain image. Previous research on image-to-image style transfer methods mainly involved supervised learning in the form of image pairs, such as the classic pix2pix method. However, heterogeneous SAR image target recognition struggles to meet such stringent requirements. On one hand, constructing SAR datasets in the form of image pairs (where only one imaging difference exists between two SAR images, while other imaging conditions remain consistent) is difficult. On the other hand, in practical applications, the unknown nature of test data makes supervised domain adaptation impractical. Therefore, unsupervised learning methods that do not rely on image pair training data are more suitable for SAR target recognition tasks; the CycleGAN method is a typical example.

[0049] For the CycleGAN task, let the source domain be X, which contains N samples, and be represented as... The target domain is denoted as Y, which contains M samples, and is represented as follows: The probability distributions of the two are represented as x ~ p, respectively. data (x) and y~p data (y). For example... Figure 4 As shown, the CycleGAN method contains two generators and two discriminators, denoted as G1, G2, and D, respectively. X D YThe function of G1 is to map samples from the source domain X to the target domain Y, and the function of G2 is to map samples from the target domain Y back to the source domain X. The images G1(X) and G2(Y) generated by the two generators constitute the transferred image domain. This domain has similar content information to the original image, a similar visual style to the target image, and a probability distribution in the feature space that is closer to the target image. When the source domain is determined, the images generated from that source domain in the transferred image domain form the intermediate domain. For example, if the source domain is X, the intermediate domain is G1(X); if the source domain is Y, the intermediate domain is G2(Y). Discriminator D X The task is to distinguish whether the input samples come from the source domain X or the transferred image domain G2(Y), and similarly, the discriminator D... Y The task is to determine whether the input sample comes from the target domain Y or the transferred image domain G1(X).

[0050] CycleGAN needs to optimize two objective functions, such as... Figure 4 As shown, the first term is the adversarial loss L. GAN Its function is to ensure that the generated intermediate domain image and the target image have a similar visual style (image style coupled with information such as brightness, contrast, and texture), that is, they have relatively similar probability feature distributions; the second term is the cycle consistency loss L. CYC Its function is to ensure that the generated intermediate domain image has the same content information (target placement, target outline, shadow position, etc.) as the source image. Among them, L... GAN It is discriminator D X and D Y The data flow is calculated separately as follows: Figure 4 As shown by the solid and dashed lines, the specific formula is as follows:

[0051]

[0052]

[0053] Cyclic Consistent Loss L CYC The goal is to reduce the difference between the reconstructed image and the original image. The reconstructed image is obtained by inputting the corresponding intermediate domain image into another generator. The data flow is as follows: Figure 4 The image is shown by a double-dotted line. The reconstructed images are represented by G2(G1(X)) and G1(G2(Y)), respectively. The reconstructed images maintain consistency with the original images in terms of content information and target effect, and are represented as follows:

[0054] x→G1(x)→G2(G1(x))≈x (3)

[0055] y→G2(y)→G1(G2(y))≈y (4)

[0056] To maintain consistency between the reconstructed image and the original image, we can approximate the difference between the images, ||G2(G1(x))-x||1 and ||G1(G2(y))-y||1, to be as small as possible. Therefore, the cycle consistency loss is defined as follows:

[0057]

[0058] Therefore, the overall objective function of CycleGAN can be expressed as:

[0059] L CycleGAN (G1,G2,D X D Y ) = L GAN (G1,D Y ,X,Y)+L GAN (G2,D X ,Y,X)+αL CYC (G1,G2) (6)

[0060] Where α is the control cycle consistency loss L CYC The larger the value of the weight factor, the more the model will focus on preserving the original target content information of the image. In the process of image style transfer, this application needs to ensure that the image after style transfer retains the original content information as much as possible, so α needs to be set to a large value, which is 10 by default.

[0061] Since the objective function of classic GANs is unstable during training, this application replaces the negative log-likelihood loss function with the least mean square loss function in order to obtain a more stable training process and more reliable training results.

[0062]

[0063]

[0064] The overall objective function of the pixel layer migration module is expressed as follows:

[0065] L PLT (G1,G2,D X D Y ) = L LSGAN (G1,D Y ,X,Y)+L LSGAN (G2,D X ,Y,X)+αL CYC (G1,G2) (9)

[0066] In one embodiment, the first-stage training model includes: an object recognition model and a first domain discriminator; the object recognition model includes a feature extractor and a classifier; supervised classification training and unsupervised domain adaptation training with aligned edge distributions are performed using generated images of the source and target domains, including: inputting the generated images of the source and target domains into the feature extractor for feature extraction, and performing global average pooling layer processing on the extracted feature maps to obtain source domain generated image features and target domain generated image features; inputting the source domain generated image features and target domain generated image features into the first domain discriminator for domain label discrimination, and obtaining a first domain discrimination prediction through adversarial learning between the feature extractor and the first domain discriminator; performing supervised training based on the first domain discrimination prediction and the domain labels in the generated images to optimize the parameters of the feature extractor and the first domain discriminator; inputting the target domain generated image features into the classifier to obtain an object classification prediction; and performing supervised training based on the object classification prediction and the category labels in the generated images of the target domain to optimize the parameters of the feature extractor and the classifier.

[0067] Specifically, the flowchart for the first phase of training is as follows: Figure 5 As shown. This step mainly involves extracting features from the source domain generated image (obtained by realizing a simulated SAR image with imaging conditions consistent with the measured SAR image in the source domain) and the target domain generated image (obtained by realizing a simulated SAR image with imaging conditions consistent with the measured SAR image in the target domain). and The first domain discriminator is input to perform domain label discrimination. Through the adversarial learning process between the feature extractor and the first domain discriminator, the features obtained by the feature extractor have both target class discriminability and domain invariance.

[0068] In one embodiment, the first domain discriminator includes a binary classifier; the loss function of the first domain discriminator is:

[0069]

[0070] Where L(·) represents the cross-entropy loss function, and h(·) represents the mapping function that maps features in the feature space to labels in the domain label space, which includes a gradient reversal layer. Indicates the generated image from the source domain. This indicates that the target domain generates an image, and the source domain's domain label value l is set. i The value of the target domain's domain label is 0. t The value is 1.

[0071] The classification loss function for generating images from the target domain is:

[0072]

[0073] in, A classification loss function is generated for the target domain to represent the image. Generate an image for the target domain. Generate category labels for images in the target domain.

[0074] Specifically, among the features extracted by neural networks, shallow features have strong transferability and mainly include features such as edges, textures, and brightness. High-level features have poor transferability and mainly contain task-related semantic information. Reducing the differences between probability distributions in the feature space can significantly improve the transferability of the corresponding layer features. For example... Figure 6 As shown, the purpose of the feature alignment module is to further bring the features f extracted by the feature extractor about the source and target domains closer together. S and f T The probability distribution is used to improve the model's ability to recognize cross-domain targets.

[0075] In domain adaptation methods, adversarial learning-based methods effectively narrow the feature probability distribution. Classic adversarial learning-based domain adaptation methods mainly consist of two parts: a domain discriminator and a gradient reversal layer (GRL). Typically, the domain discriminator is a binary classifier whose main purpose is to determine whether the features extracted from the feature extractor come from the source or target domain. At this stage, the "adversarial" idea of ​​adversarial learning is reflected in the fact that the optimization objective of the domain discriminator is to minimize the domain discrimination loss L. D The optimization objective of the feature extractor is to maximize the domain discrimination loss L. D To achieve end-to-end adversarial learning rather than staged training, a gradient inversion layer is added between the domain discriminator and the feature extractor. Its function is to perform no additional processing on the features input to the domain discriminator during forward propagation, and to invert the gradients from the domain discriminator input to the feature extractor during backward propagation. In summary, the loss function of the first domain discriminator is described as follows:

[0076]

[0077] In the formula, L(·) represents the cross-entropy loss function, and h(·) represents the mapping function that maps features in the feature space to labels in the domain label space, which includes a gradient inversion layer. Indicates the generated image from the source domain. This indicates that the target domain will generate the image. The source domain's domain label value is set. i The value of the target domain's domain label is 0. t The value is 1.

[0078] Obtaining the domain discrimination loss in the first domain discriminator Simultaneously, the convolutional neural network generates images based on the target domain. and its category tags The classification loss was calculated. As shown in formula (11).

[0079] The main difference between the second-domain discriminator and the first-domain discriminator lies in the input content; the former inputs a feature vector f. S With category prediction vector The vector after the cross product, and the eigenvector f T With category prediction vector The vector obtained after the cross product, where the latter is the input feature vector. and

[0080] In one embodiment, the second-stage training model includes: a target recognition model and a second-domain discriminator; the target recognition model includes a feature extractor and a classifier; supervised classification training and unsupervised domain adaptation training with aligned joint distribution are performed using measured SAR images of the source and target domains, including: inputting the measured SAR images of the source and target domains into the feature extractor for feature extraction, and performing global average pooling layer processing on the extracted features to obtain source domain features and target domain features; inputting the source domain features and target domain features into two classifiers respectively to obtain source domain target classification prediction and target domain classification prediction; and performing a cross product operation between the source domain features and the source domain target classification prediction. The target domain features and target domain classification predictions are cross-producted to obtain new source domain features and new target domain features. These new source domain features and new target domain features are then flattened into feature vectors and input into a second domain discriminator for domain label discrimination. Through adversarial learning between the feature extractor and the second domain discriminator, a second domain discrimination prediction is obtained. Supervised training is then performed based on the second domain discrimination prediction and the domain labels in the measured SAR images to optimize the parameters of the feature extractor and the second domain discriminator. Finally, supervised training is performed based on the source domain target classification prediction, the target domain classification prediction, and the category labels in the measured SAR images of the source and target domains to optimize the parameters of the feature extractor and the classifier.

[0081] Specifically, the flowchart for the second phase of training is as follows: Figure 7 As shown. This step mainly involves extracting features f from the measured SAR images in the source and target domains. S and f T The corresponding target classification prediction vector obtained and A cross product operation is performed to obtain a new feature map, which is then flattened into a feature vector and input into the second domain discriminator for domain label discrimination. Through the adversarial learning process between the feature extractor and the second domain discriminator, the features obtained by the feature extractor simultaneously possess target class discriminability and domain invariance.

[0082] The main function of the second domain discriminator is to align the joint probability distribution of the source and target domains through adversarial learning, while reducing the domain offset between the source and target domains and acquiring fine-grained information about the target category in the target domain. This information is acquired during the training phase, which is beneficial for discriminating the target domain image during the testing phase. The biggest difference between the first and second domain discriminators is that the first domain discriminator aligns the edge distribution rather than the joint distribution during domain adaptation. In other words, the first domain discriminator does not acquire fine-grained category information of the target domain generated image, because supervised recognition training using the target domain generated image can better obtain fine-grained category information of the target domain image.

[0083] In one embodiment, the second domain discriminator employs a joint probability distribution and adaptation method; the loss function of the second domain discriminator is:

[0084]

[0085] In the formula, L(·) represents the cross-entropy loss function, and k(·) represents the mapping function that maps features in the feature space to labels in the domain label space, which includes a gradient reversal layer. This represents a sample image from the source domain. The sample image represents the target domain, and the domain label value l of the source domain is set. i The value of the target domain's domain label is 0. t =1;

[0086] The loss for classification and prediction of measured SAR images in the source domain is:

[0087]

[0088] in, This is a measured SAR image of the source region. For the category labels of the measured SAR images in the domain, The loss is used for classification prediction of measured SAR images in the source domain.

[0089] In one embodiment, the training process of the target recognition model is divided into two stages, wherein the optimization objective of the first stage is:

[0090] L CycleGAN (G1,G2,D X D Y (15)

[0091] The optimization objective for the second stage is:

[0092]

[0093] in, Generate an image classification prediction loss for the target domain. The loss for classification and prediction of measured SAR images in the source domain is... The loss function of the first domain discriminator The loss function of the second-domain discriminator uses three weighting factors: α, β, and γ. Preferably, α is set to 1 by default, β to 1 by default, and γ to 0.5 by default.

[0094] Specifically, domain adaptation methods can generally be divided into three categories based on their alignment methods:

[0095] (1) Marginal Distribution Domain Adaptation

[0096] Marginal probability distribution domain adaptation (also known as marginal distribution domain adaptation) aligns the marginal distributions (i.e., the distribution of data in each feature dimension) of the source and target domains. Its core idea is to achieve alignment by minimizing the distance between the two domains. As shown in equation (13), where: Dis mar (D s D t P(X) represents a measure of marginal distribution dissimilarity. s P(X) represents the source domain sample distribution. t ) represents the target domain sample distribution.

[0097]

[0098] (2) Conditional Distribution Domain Adaptation

[0099] Conditional probability distribution domain adaptation (also known as conditional distribution domain adaptation) aligns the conditional distributions (i.e., the output conditional probability distributions given a feature vector) of the source and target domains. Its core idea is to learn a conditional transformation function to map the conditional distribution of the source domain to the conditional distribution of the target domain. As shown in equation (14), where: Dis con (D s D t P(Y) represents a measure of the conditional distribution variance. s |X s P(Y) represents the conditional distribution of the source domain samples. t |X t ) represents the conditional distribution of the target domain samples.

[0100]

[0101] (3) Joint Distribution Domain Adaptation

[0102] Joint probability distribution domain adaptation (also known as joint distribution domain adaptation) is a method that aligns the joint distribution of the source and target domains. Its core idea is to learn a new joint distribution from which both the source and target domains can obtain similar samples, and where the classifier performs well. As shown in equation (15), where: Dis mar+con (D s D t ) represents a measure of joint distribution difference.

[0103]

[0104] To achieve the effect of aligning the joint distribution, this application will input the feature vector f. S With category prediction vector The vector after the cross product, and the eigenvector f T With category prediction vector The vector resulting from the cross product is used as a new feature vector and input into the second domain discriminator. The loss function of the second domain discriminator is described by equation (13).

[0105] Obtaining the domain discrimination loss in the second domain discriminator Simultaneously, the convolutional neural network also uses measured SAR images from the source domain. and its category tags The classification loss was calculated. As shown in formula (14).

[0106] In summary, the entire training process is divided into two phases. The optimization objective of the first phase is L. CycleGAN (G1,G2,D X D Y The optimization objective for the second stage is shown in equation (16).

[0107] In one embodiment, the target recognition model includes a feature extractor and a classifier; wherein the feature extractor is the feature extraction part of ResNet, and the classifier includes a fully connected layer.

[0108] Specifically, ResNet's greatest contribution is embedding residual modules into convolutional neural networks, effectively solving the gradient vanishing problem that occurs during the optimization process of deep networks. This allows ResNet to continuously improve its nonlinear fitting ability by increasing network depth. Therefore, ResNet is used as a feature extractor or directly as the backbone network in many tasks based on deep convolutional neural networks. The residual learning module consists of shallow networks and embedded self-mapping layers, such as... Figure 8As shown, x is the input; the output of x after passing through two weight layers and ReLU is F(x), called the residual function; the total output is H(x). Simple convolutional neural networks like AlexNet and VGG have difficulty directly learning the total output H(x). ResNet solves this problem by introducing a residual learning model. It learns the residual function F(x) and then indirectly learns the total output H(x) through a simple mapping function H(x) = F(x) + x, making it approximate the desired output.

[0109] As a preferred method, ResNet-18 is used as the feature extractor, such as Figure 9 As shown, the input SAR target image x is reduced to a feature map after nine convolution operations. The feature map obtained from the SAR image dimensionality reduction is then input into a Global Average Pooling Layer (GAP) for pooling, flattened into a feature vector, and then input into a Fully Connected Layer (FC) to obtain the class prediction f(x) of the input image. The FC is considered a classifier. The target classification loss is calculated from the class prediction f(x) of the model network and the target image label y. The cross-entropy loss function is chosen, so the target classification loss function L... C The definition is as follows:

[0110]

[0111] Where L(·) represents the cross-entropy loss function, f(·) represents the mapping function that maps the image input space to the output space of the class prediction, and x n For image samples, y n Here, N represents the label corresponding to the image, and N is the number of image samples.

[0112] It should be understood that, although Figure 1 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 1 At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.

[0113] In a verification embodiment, the method is experimentally verified to demonstrate its effectiveness. The verification is presented from five aspects: experimental data introduction, experimental data setup, hyperparameter settings, comparison methods, and experimental results and analysis.

[0114] (1) Introduction of Experimental Data

[0115] This experiment used the MSTAR dataset and an electromagnetic scattering simulation dataset created based on this standard dataset. The following sections will introduce these two datasets in detail.

[0116] 1) MSTAR dataset

[0117] MSTAR (Moving and Stationary Target Acquisition and Recognition) is a synthetic aperture radar (SAR) image dataset for stationary ground targets, released by the U.S. Department of Defense and Sandia National Laboratories. In the early 1990s, this dataset was used for target recognition of military vehicles. Later, DARPA released the dataset, and MSTAR has been widely used in SAR image target recognition research. To this day, the MSTAR dataset remains the recognized benchmark dataset in the field of SAR vehicle target recognition.

[0118] The MSTAR dataset primarily consists of SAR slice images of stationary vehicles, covering 10 categories of vehicle target images at four different elevation angles: 15°, 17°, 30°, and 45°. Figure 10 Four types of targets were showcased: 2S1 (self-propelled howitzer), BRDM2 (armored reconnaissance vehicle), T72 (tank), and ZSU234 (self-propelled anti-aircraft gun). The images in the MSTAR dataset have a resolution of 0.3m, are imaged in the X-band, and are polarized in HH mode.

[0119] The original SAR slice images in the MSTAR dataset are not of uniform size. The regions of interest are the targets in the center of the slice images and their shadow regions. Using the same training image size helps in training the neural network model. Therefore, we performed a center cropping operation on the original dataset to extract 128×128 pixel slice images containing various targets as training and testing data. All MSTAR data used later are images that have been center cropped.

[0120] 2) Electromagnetic scattering simulation dataset

[0121] The electromagnetic scattering simulation dataset was generated using Fastem-AMBIENTER software. This simulation technique is based on the rapid generation of high-resolution single-look SAR images using ray tracing, completing monocentric imaging with a single ray tracing operation. The imaging mode is circular monocentric 2D SAR imaging. The simulation data includes four vehicle target categories (2S1, BRDM2, T72, and ZSU234) based on two terrain backgrounds (dry sand and sparse dry grassland). Radar parameter settings include a center frequency of 9.6 GHz, a bandwidth of 591 MHz, an azimuth scan width of 1 degree, range and azimuth resolutions of 0.3 m, HH polarization, and a range and azimuth pixel spacing of 0.2 m.

[0122] The simulation conditions were compared with the Standard Operating Conditions (SOC) and Extended Operating Conditions (EOC) of the MSTAR dataset. Under SOC, the imaging conditions for training and testing data differed very little, with elevation angles of 15° and 17° respectively. Under EOC, the imaging conditions for training and testing data differed significantly, specifically in three categories: EOC-1 represented large elevation angles, for example, the elevation angle for the training data was 17°, while for the testing data it was 30°; EOC-2 involved different vehicle configurations, i.e., the addition or removal of some components on the target, such as removing the tank fuel tank from a T-62; EOC-3 involved different vehicle versions and functions, such as changing a vehicle from a transport vehicle to a reconnaissance vehicle.

[0123] The simulation data used in this embodiment corresponds to the EOC-1 condition, including three pitch angles (17°, 30° and 45°), and includes four types of targets: 2S1, BRDM2, T72 and ZSU234. Each type contains 360 images with azimuth angles of 0-359°, with azimuth angle intervals of 1°. The imaging background is a grassland background, for a total of 17280 images.

[0124] (2) Experimental data setup

[0125] To verify the effectiveness of SJDDA, this embodiment will conduct experiments using the two heterogeneous SAR image datasets mentioned above. Considering that the image differences are small under SOC conditions, and existing neural network methods can achieve a recognition performance of about 99%, this embodiment will not reproduce the classic SOC classification experiment.

[0126] Table 1 shows the number of images of four vehicle types at 17°, 30°, and 45° downward angles.

[0127]

[0128] Primarily considering the EOC scenario, cross-experiments are conducted between elevation angles of 17°, 30°, and 45°. When the measured data at 17° elevation angle is used as the training set and the data at 30° elevation angle is used as the test set, this experiment is denoted as "17°→30°". Similar experiments include "30°→17°", "17°→45°", "45°→17°", "30°→45°", and "45°→30°". Each experiment uses corresponding simulated SAR images. For example, according to the domain-adaptive data partitioning pattern, in the "17°→30°" experiment, the measured SAR image at 17° is used as the training set along with the source domain measured SAR image, and the simulated SAR image at 17° is used as the training set along with the source domain simulated SAR image. The measured SAR image at 30° is used as the test set along with the target domain measured SAR image, and the simulated SAR image at 30° is used as the training set along with the target domain simulated SAR image. Table 1 shows the various measured targets at each elevation angle. The simulated SAR images are 360 ​​omnidirectional images of each type of target at each elevation angle, all of which were used in the training.

[0129] (3) Hyperparameter settings

[0130] The simulation image and alignment-based joint distribution domain adaptation method involves two separate training steps. The first step trains the simulation image realization module CycleGAN, setting its training epochs to 200. Other key parameters are shown in Table 2. The image size generated by this module is 256×256 pixels.

[0131] Table 2. CycleGAN Training Hyperparameters

[0132]

[0133] Table 3 Training Hyperparameter Settings

[0134]

[0135] The hyperparameters of the Simulated Image and Alignment Joint Distribution Domain Adaptation (SJDDA) method are shown in Table 3, including optimizations for the first and second training phases. The initial learning rate (lr) was set to 0.01, the batch size (batch_size) was set to 8, and a total of 60 training epochs were performed. No additional data augmentation methods were introduced during training. The trained backbone network can then be used to test on the test set. During training, all randomization processes were controlled by random seeds, including network initialization, the training process, and the randomization process of CUDNN. Random seeds 1-5 were selected for 5 experiments.

[0136] The experimental environment was as follows: the training and testing machine had a motherboard model of B360M, a CPU model of Intel Core i7-8700, a graphics card model of NVIDIA GeForce RTX 2080ti, a memory capacity of 32G, and an operating system of Ubuntu 18.04.

[0137] (4) Comparison Method

[0138] In this embodiment, ResNet-18 is used as the backbone network, which is also known as the source-only method in the domain adaptation field. All the comparison methods below use ResNet-18 as the backbone network of the model.

[0139] Considering that the target recognition tasks involved in this embodiment are all in heterogeneous scenarios, the comparison methods used are all domain adaptation methods, specifically DAN, JAN, DANN, ADDA, and CDAN. DAN: After the feature extractor, multi-kernel maximum mean difference (MK-MMD) is used on the extracted features to reduce domain bias. JAN: Compared to DAN aligning edge distributions, JAN introduces the class vector predicted by the classifier, reducing domain bias along with the features, aiming to align the joint distribution between the source and target domains. DANN: The main innovation of DANN includes a domain discriminator and a gradient inversion layer. The latter allows the entire network to maximize the domain discriminant loss through the feature extractor. Minimizing the domain discriminant loss enables end-to-end processing, achieving adversarial learning and making the features extracted by the feature extractor domain invariant. ADDA: Based on DANN, it adopts a step-by-step training approach, separating the source domain feature extractor from the target domain feature extractor, making the parameter distribution of the latter closer to the distribution of the target domain data. CDAN: Similar to JAN, CDAN builds upon DANN by performing a cross product operation between the classifier's category prediction vector and the features. The resulting vector is then input into the domain discriminator for adversarial training. The advantage of this step is that it can extract the "multi-peak structure" of category information from the features, thereby improving the model's classification performance while ensuring domain adaptation performance.

[0140] (5) Analysis of experimental results

[0141] 1) Analysis of Simulation Image Realization Results

[0142] This embodiment analyzes the results of SAR image realization. Firstly, the realized images at 17° and 30° elevation angles are visually very similar to the measured SAR images. However, it can be seen that there are differences between the simulated SAR images and the measured SAR images in terms of strong scattering points and target contours. Since this embodiment sets the weight factor of the cycle consistency loss to 10 in the SAR image realization step, the realized image maintains the same content information as the original image. It is evident that these differences are difficult to reduce using pixel-level domain adaptation methods. In the experimental results at 45° elevation angle, a significant error is that the position of the target shadow changes after style transfer, and the shadow shape is altered. Compared to the realization results at 17° and 30° elevation angles, the realization results at 45° elevation angle appear less reliable. However, the realized image is closer to the measured SAR image in terms of background texture and brightness. Domain adaptation between simulated SAR images can also reduce the domain offset between the source and target domain measured data to some extent.

[0143] 2) Classification recognition rate

[0144] To verify the effectiveness of the method in this application, six sets of experiments were conducted under the conditions of 17°→30°, 30°→17°, 17°→45°, 45°→17°, 30°→45°, and 45°→30°. The results were compared with other domain adaptation methods. The experimental results are shown in Table 4. The numerical values ​​represent "mean recognition rate ± standard deviation," which is obtained from the results of five fixed random seed trials. Since all random processes are fixed by the random seed, this value can be accurately reproduced. The Average value represents the mean of the average recognition rates of the six experiments, used to measure the overall recognition performance of the model. Table 5 shows the recognition rates obtained by SJDDA under 10 random seeds (1-10), with the experiment being 17°→45°. Based on Tables 4 and 5, the following conclusions can be drawn:

[0145] (i) Differences in pitch angle lead to significant variations in image domain adaptation performance. Experiments show that the domain adaptation classification performance is highest for 17°→30° and 30°→17°, followed by 30°→45° and 45°→30°, with the lowest being 17°→45° and 45°→17°. The greater the pitch angle difference, the lower the performance of the domain adaptation method. For the source-only method, using the same 17° pitch angle source domain data, the recognition performance for 30° pitch angle data is 98.43%, but for 45° pitch angle data, the recognition performance is only 50.19%. This demonstrates that, regardless of whether it's a source-only or domain adaptation method, the generalization performance of the depth model decreases with increasing differences in imaging conditions.

[0146] (ii) The proposed method directly aligns the joint distribution between measured SAR images, mining knowledge from the target domain data. The recognition performance reaches 90.51%, surpassing all comparable methods. This demonstrates that aligning conditional distributions, rather than aligning edge distributions between the source and target domains, further narrows the inter- and intra-class differences between them. Domain adaptation using simulated SAR images further reduces imaging differences caused by varying operating conditions. Simultaneously, supervised optimization of the classifier using realistic target domain simulated SAR images enables the convolutional neural network to directly capture fine-grained target category information, improving overall model performance. The recognition performance at 17°→45° and 30°→45° is relatively poor compared to other experiments. This is attributed to two reasons: the data feature probability distribution difference is greater at 45°, making the results less reliable than those at 17° and 30° when realising simulated SAR images at 45°.

[0147] Table 4 Target recognition rate (%) at different pitch angles

[0148]

[0149] (iii) In the domain adaptation tasks of 17°→45° and 45°→17°, SJDDA achieved a recognition rate of 88.00% and 98.71%, respectively, which is 10.89% and 7.02% higher than the highest recognition rate of the classic unsupervised domain adaptation methods. In the 30°→45° task, it is 7.95% higher than the second place, and in the 45°→30° task, it is 3.77% higher than the second place. In the experiment of 17°→30°, it also achieved the highest recognition rate of 100%, and in the experiment of 30°→17°, it is only 0.02% lower than the first place. After 60 rounds of training, SJDDA can stably achieve the best performance. This indicates that, with sufficient training, feature-based domain adaptation methods are sufficient to improve the model's insufficient generalization ability when the pitch angle difference is small. However, when the pitch angle difference increases, the feature distribution becomes highly differentiated, and classical unsupervised domain adaptation cannot completely align the feature distribution. Our proposed method reduces the difference across pitch angles by performing domain adaptation between simulated SAR images. Aligning the joint distribution of features between measured data allows the network to adapt to target domain data with greater differences while transferring knowledge from the source domain to the target domain. Therefore, our method achieves higher classification accuracy under conditions of greater pitch angle difference. As shown in Table 5, our method achieved a recognition rate of 98.01% in experiments with a random seed of 9, demonstrating high recognition performance through suitable random initialization for SJDDA.

[0150] Table 5. Recognition rate of SJDDA under 10 random seeds, 17°→45° experiment.

[0151]

[0152] 3) Ablation experiments and recognition rates of various targets

[0153] First, ablation experiments were conducted on the three steps of the SJDDA training process to verify the effectiveness of each step. The first and second training phases are used in combination; that is, without the first training phase, the simulation of the SAR image in step 101 is meaningless. The ablation experiment results are shown in Table 6. The results show that each of the proposed steps is effective individually, and the combination of all three maximizes the recognition rate.

[0154] Table 6 shows the average recognition rate of SJDDA in ablation experiments.

[0155]

[0156] 4) Feature visualization

[0157] This embodiment employs the t-SNE dimensionality reduction method to visualize the features of the domain adaptation layers in Source-only, DAN, JAN, DANN, CDAN, and SJDDA networks in two-dimensional space for domain adaptation tasks with pitch angle changes of 17°→30°, 17°→45°, and 30°→45°. In this application, the domain adaptation layer for all compared methods is set as the layer preceding the network output layer.

[0158] For the target domain data feature distribution of models trained solely on source domain data, there are significant distribution differences across data at different pitch angles. The target domain data feature distribution differs considerably from the source domain data feature distribution, resulting in clustered target domain features with poor separability and thus the worst classification performance. In comparative experiments, the adversarial learning-based domain adaptation methods DANN and CDAN can narrow the source and target domain data feature distributions from 17° to 30° and from 30° to 45°, but their effect is still minimal at 17° to 45°. Specifically, the DAN method significantly reduces the separability of the source domain data during adaptation, and the domain adaptation loss term increases the source domain generalization error. This indicates that the domain adaptation loss term acts as a regularization constraint during network training. Although narrowing the source and target domain distributions can improve performance to some extent, it damages the task classification performance of the source domain data, resulting in the continued existence of the target domain generalization error. The feature visualization results of this application for all six sets of experiments show that, compared with other comparative methods, the method of this application has the best alignment between the source domain and the target domain data. It has completed the alignment task in almost all six sets of experiments. The alignment between 17°→45° and 30°→45° is relatively poor. Although it has not achieved fine alignment in each category, the overall alignment effect has been significantly improved compared with the comparative methods, which supports the recognition rate results.

[0159] In one embodiment, such as Figure 10 As shown, a target recognition device for SAR images across elevation angles is provided, comprising: a simulated image realization processing module, a target recognition model training module, and a target recognition module, wherein:

[0160] The simulation image realization processing module is used to perform simulation image realization processing on the acquired simulation SAR images and measured SAR images under different elevation angle conditions, and generate generated images under the corresponding elevation angle conditions; the measured SAR images include measured SAR images in the source domain and the target domain, and the generated images include generated images in the source domain and the target domain.

[0161] The target recognition model training module is used to input generated images of the source and target domains, as well as measured SAR images of the source and target domains, into the target recognition model for iterative training to obtain the trained target classification model. The training process includes two stages: the first training stage: supervised classification training and unsupervised domain adaptation training with aligned edge distributions are performed using generated images of the source and target domains; the second training stage: supervised classification training and unsupervised domain adaptation training with aligned joint distributions are performed using measured SAR images of the source and target domains.

[0162] The target recognition module is used to acquire measured SAR images of the target domain and input the measured SAR images into the trained target classification model to classify the target categories in the measured SAR images.

[0163] In one embodiment, the simulated SAR image realization processing module is further configured to use the simulated SAR image and the measured SAR image at the first elevation angle as the source domain data and target domain data of the first elevation angle, respectively; input the source domain data and target domain data of the first elevation angle into CycleGAN to perform realization processing on the simulated SAR image, generating a generated image at the first elevation angle; the generated image at the first elevation angle is a generated image with the simulated SAR image at the first elevation angle as the original image and the measured SAR image at the first elevation angle as the target image; the CycleGAN method is also used to perform simulated image realization processing on the simulated SAR image and the measured SAR image under other elevation angle conditions to obtain generated images under other elevation angle conditions.

[0164] In one embodiment, the first-stage training model includes: an object recognition model and a first domain discriminator; the object recognition model includes a feature extractor and a classifier; the object recognition model training module is used to input generated images of the source domain and the target domain into the feature extractor for feature extraction, and perform global average pooling layer processing on the extracted feature maps to obtain source domain generated image features and target domain generated image features; input the source domain generated image features and target domain generated image features into the first domain discriminator for domain label discrimination, and obtain a first domain discrimination prediction through adversarial learning between the feature extractor and the first domain discriminator; perform supervised training based on the first domain discrimination prediction and the domain labels in the generated images to optimize the parameters of the feature extractor and the first domain discriminator; input the target domain generated image features into the classifier to obtain an object classification prediction; and perform supervised training based on the object classification prediction and the category labels in the target domain generated images to optimize the parameters of the feature extractor and the classifier.

[0165] In one embodiment, the first-domain discriminator in the first-stage training model includes a binary classifier; the loss function of the first-domain discriminator is shown in Equation (10). The classification loss function of the target domain generated image is shown in Equation (11).

[0166] In one embodiment, the second-stage training model includes: a target recognition model and a second-domain discriminator; the target recognition model includes a feature extractor and a classifier; the target recognition model training module is used to input measured SAR images of the source domain and the target domain into the feature extractor for feature extraction, and to perform global average pooling layer processing on the extracted features to obtain source domain features and target domain features; the source domain features and target domain features are respectively input into two classifiers to obtain source domain target classification prediction and target domain classification prediction; the source domain features and source domain target classification prediction are cross-producted, and the target domain features and target domain classification prediction are cross-producted. A cross product operation is performed to obtain new source domain features and new target domain features. These features are then flattened into feature vectors and input into a second domain discriminator for domain label discrimination. Through adversarial learning between the feature extractor and the second domain discriminator, a second domain discrimination prediction is obtained. Supervised training is then performed based on the second domain discrimination prediction and the domain labels in the measured SAR images to optimize the parameters of the feature extractor and the second domain discriminator. Supervised training is then performed based on the source domain target classification prediction, the target domain classification prediction, and the category labels in the measured SAR images of the source and target domains to optimize the parameters of the feature extractor and the classifier.

[0167] In one embodiment, the second domain discriminator employs a joint probability distribution and adaptation method during the second-stage training process; the loss function of the second domain discriminator is shown in Equation (13). The source domain measured SAR image classification prediction loss is shown in Equation (14).

[0168] In one embodiment, the training process of the target recognition model in the target recognition model training module is divided into two stages, wherein the optimization objective of the first stage is shown in Equation (15); and the optimization objective of the second stage is shown in Equation (16).

[0169] In one embodiment, the target recognition model training module includes a feature extractor and a classifier; wherein the feature extractor is the feature extraction part of the ResNet network, and the classifier includes a fully connected layer.

[0170] Specific limitations regarding the cross-elevation angle SAR image target recognition device can be found in the limitations of the cross-elevation angle SAR image target recognition method described above, and will not be repeated here. Each module in the aforementioned cross-elevation angle SAR image target recognition device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0171] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 11 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a cross-elevation angle SAR image target recognition method. The display screen can be a liquid crystal display (LCD) or an e-ink display. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the computer device casing, or an external keyboard, touchpad, or mouse.

[0172] Those skilled in the art will understand that Figure 11 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0173] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps in the above method embodiment.

[0174] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0175] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for target recognition in SAR images across elevation angles, characterized in that, The method includes: Based on the simulated SAR images and measured SAR images obtained under different elevation angle conditions, the simulated images are processed to become realistic, generating generated images under the corresponding elevation angle conditions; the measured SAR images include measured SAR images in the source domain and the target domain, and the generated images include generated images in the source domain and the target domain. The generated images of the source and target domains, as well as the measured SAR images of the source and target domains, are input into the target recognition model for iterative training to obtain the trained target classification model. The training process includes two stages: the first training stage: supervised classification training and unsupervised domain adaptation training with aligned edge distributions are performed using the generated images of the source and target domains; the second training stage: supervised classification training and unsupervised domain adaptation training with aligned joint distributions are performed using the measured SAR images of the source and target domains. Within each iteration, the first and second training stages are performed alternately and optimized alternately until a preset stopping condition is met. Acquire a measured SAR image of the target domain, and input the measured SAR image into a trained target classification model to classify the target categories in the measured SAR image.

2. The method according to claim 1, characterized in that, Based on the simulated SAR images and measured SAR images obtained under different elevation angle conditions, the simulated images are processed to achieve realism, generating generated images under the corresponding elevation angle conditions, including: The simulated SAR image and the measured SAR image at the first elevation angle are used as the source domain data and target domain data at the first elevation angle, respectively. The source domain data and target domain data of the first elevation angle are input into CycleGAN to perform realization processing on the simulated SAR image and generate the generated image of the first elevation angle; the generated image of the first elevation angle is a generated image with the simulated SAR image of the first elevation angle as the original image and the measured SAR image of the first elevation angle as the target image. The CycleGAN method is also used to realize the realization of simulated SAR images under other elevation angle conditions, and generated images under other elevation angle conditions are obtained.

3. The method according to claim 1, characterized in that, The first-stage training model includes: the target recognition model and the first domain discriminator; the target recognition model includes a feature extractor and a classifier; Supervised classification training and unsupervised domain fitting training to align edge distributions are performed using generated images from the source and target domains, including: The generated images of the source domain and the target domain are input into the feature extractor for feature extraction, and the extracted feature maps are processed by a global average pooling layer to obtain the features of the generated image of the source domain and the generated image of the target domain. The source domain generated image features and the target domain generated image features are input into the first domain discriminator for domain label discrimination. Through adversarial learning between the feature extractor and the first domain discriminator, the first domain discrimination prediction is obtained. The parameters of the feature extractor and the first domain discriminator are optimized by performing supervised training based on the first domain discriminant prediction and the domain labels in the generated image. The target domain generated image features are input into the classifier to obtain the target classification prediction; The parameters of the feature extractor and the classifier are optimized by supervised training based on the target classification prediction and the class labels in the generated image of the target domain.

4. The method according to claim 3, characterized in that, The first domain discriminator includes a binary classifier; The loss function of the first-domain discriminant is: in, Represents the cross-entropy loss function. This represents a mapping function that maps features in the feature space to labels in the label space, and includes a gradient inversion layer. Indicates the generated image from the source domain. Indicates that the target domain generates an image. This is the domain label value of the source domain, and its value is 0. This is the domain label value for the target domain, and its value is 1. The classification loss function for generating images from the target domain is: in, A classification loss function is generated for the target domain to represent the image. Generate an image for the target domain. Generate category labels for images in the target domain.

5. The method according to claim 1, characterized in that, The second-stage training model includes: the target recognition model and the second-domain discriminator; the target recognition model includes a feature extractor and a classifier. Supervised classification training and unsupervised domain adaptation training with aligned joint distribution are performed using measured SAR images from the source and target domains, including: The measured SAR images of the source and target domains are input into the feature extractor for feature extraction, and the extracted features are processed by a global average pooling layer to obtain source domain features and target domain features. The source domain features and the target domain features are respectively input into two classifiers to obtain source domain target classification prediction and target domain classification prediction; Perform a cross product operation between the source domain features and the source domain target classification prediction, and perform a cross product operation between the target domain features and the target domain classification prediction to obtain new source domain features and new target domain features; The new source domain features and the new target domain features are flattened into feature vectors and then input into the second domain discriminator for domain label discrimination. Through adversarial learning between the feature extractor and the second domain discriminator, the second domain discrimination prediction is obtained. The parameters of the feature extractor and the second domain discriminator are optimized by supervised training based on the second domain discrimination prediction and the domain labels in the measured SAR image. The parameters of the feature extractor and the classifier are optimized through supervised training based on the source domain target classification prediction, the target domain classification prediction, and the category labels in the measured SAR images of the source and target domains.

6. The method according to claim 5, characterized in that, The second domain discriminator employs a joint probability distribution domain adaptation method; The loss function of the second domain discriminator is: In the formula, Represents the cross-entropy loss function. This represents a mapping function that maps features in the feature space to labels in the label space, and includes a gradient inversion layer. This represents a sample image from the source domain. The sample image represents the target domain, and the domain label value of the source domain is set. The value is 0, which is the domain label value of the target domain. =1; The loss for classification and prediction of measured SAR images in the source domain is: in, This is a measured SAR image of the source region. For the category labels of the measured SAR images in the domain, The loss is used for classification prediction of measured SAR images in the source domain.

7. The method according to claim 1, characterized in that, The training process of the target recognition model is divided into two stages, where the optimization objective of the first stage is: The optimization objective for the second stage is: in, Generate an image classification prediction loss for the target domain. The loss for classification and prediction of measured SAR images in the source domain is... The loss function of the first domain discriminator The loss function of the second domain discriminator There are 3 weighting factors.

8. The method according to claim 1, characterized in that, The target recognition model includes a feature extractor and a classifier; wherein the feature extractor is the feature extraction part of a ResNet network, and the classifier includes a fully connected layer.

9. A target recognition device for SAR images across elevation angles, characterized in that, The device includes: The simulation image realization processing module is used to perform simulation image realization processing on the acquired simulation SAR images and measured SAR images under different elevation angle conditions, and generate generated images under the corresponding elevation angle conditions; the measured SAR images include measured SAR images in the source domain and the target domain, and the generated images include generated images in the source domain and the target domain. The target recognition model training module is used to input generated images of the source and target domains, as well as measured SAR images of the source and target domains, into the target recognition model for iterative training to obtain the trained target classification model. The training process includes two stages: the first training stage: supervised classification training and unsupervised domain adaptation training with edge distribution alignment using generated images of the source and target domains; the second training stage: supervised classification training and unsupervised domain adaptation training with joint distribution alignment using measured SAR images of the source and target domains. Within each iteration, the first and second training stages are performed alternately and optimized until a preset stopping condition is met. The target recognition module is used to acquire measured SAR images of the target domain and input the measured SAR images into a trained target classification model to classify the target categories in the measured SAR images.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Cross-domain adaptive SAR image classification method and device based on simulation data, and equipment

    CN113762203A